No training on your data means a vendor processes your content to deliver the result you asked for and never feeds it back into training, fine-tuning, evaluating or benchmarking a machine learning model. Your documents are an input to the service, not raw material for the product.
That sentence costs a vendor nothing to write, which is why security reviews stopped accepting it. Below are the 12 questions that test it, and Parseur's answers to them, including the ones where the honest answer is "ask us and hold us to it in writing."
Key Takeaways
- No training on your data means your content gets processed, never recycled into a model.
- Any vendor can say it, so security reviews have moved on to the parts you can check: subprocessors named, retention quoted in days, and the anonymised-data carve-out closed.
- Training is where the leaks live. Model inversion can reconstruct documents out of a trained model, cross-tenant contamination carries one customer's patterns into another customer's output, and GDPR erasure stops working once a document has become a model weight.
- Parseur never reuses customer data to train its AI models and never sells it, is EU-hosted and GDPR-native, and lets you set the document retention window from one day to unlimited.
- Parseur is SOC 2 Type II compliant, with the report available through the trust center. HIPAA is still in progress, which means not HIPAA certified today. Parseur is also not a zero data retention vendor, and this page explains why no document parser you would actually want can be.
Why "We Never Train on Your Data" Stopped Meaning Anything
Every vendor says it now. Which is exactly the problem.
A reviewer who reads "we never train on your data" on a marketing page has learned nothing they can put in a risk register. The sentence has no scope, names no subprocessors, quotes no retention number, and carries nobody's signature.
The gap between that sentence and the contract turns out to be wide. DataGrail's 2026 Privacy and AI Trends Report, which analysed 2,400 popular business software providers, found that 63.6% of vendors advertising AI capabilities do not disclose a third-party AI subprocessor anywhere in their legal documentation.
A few of those vendors genuinely have no subprocessor to disclose. The rest are asking you to accept an unnamed supply chain on trust, and from the outside you cannot tell the two groups apart. That is the entire job of the questions below.
So the useful question is no longer whether a vendor has a no-training policy. It is what survives when you ask them to write it down.
Processing Versus Training, and Why Only One Is Reversible
Two things can happen to a document you upload, and the difference is whether anything is left behind when the job is done.
- Processing: your data is used in real time to complete a task, such as pulling the line items off a supplier invoice. When the task finishes, the data does not re-enter the system for future learning.
- Training: your data is fed into a model to improve it, so your inputs become part of what the model knows. There is no undo. Getting your content back out means retraining the model, and nobody has budgeted for that.
Blurring that line created three problems that outlived the blur. Confidential documents could be reconstructed out of a model that had memorised them. Ownership turned murky, because a model trained on your contracts has absorbed your intellectual property. And regulated industries lost the ability to answer the most basic audit question there is, which is where the client data actually went.
A no-training policy removes the ambiguity. Your data is processed for the task, then it stops.
What Leaks When a Model Trains on Your Documents
"No undo" is the kind of phrase you skim past. The mechanism behind it has a name. Model inversion is the reconstruction of training data by querying or analysing a trained model. Nothing gets breached. The model simply answers, and what it answers with was shaped by everything it absorbed.
Some material memorises better than others, and business documents sit near the top of the list. An invoice is not prose. It is a dense little grid of identifiers, amounts and relationships, which is exactly the kind of structure a model holds on to.
Cross-tenant contamination is when one customer's documents influence the output another customer receives. Duller name, wider blast radius. No access control is bypassed, nothing unusual shows up in anyone's logs, and your contract structures and pricing logic stop being exclusively yours the moment they join a shared pool.
The third failure is the one Legal finds first. GDPR gives a person the right to erasure, which is a five-minute job while their data is a file and an impossible one once it is a model weight. Delete the source and the model still knows. Untraining is not a feature anybody ships.
None of it trips an alarm, which is why it tends to surface late. Roughly 40% of organizations reported an AI-related privacy incident in 2024-2025, mostly leaks through prompts, logs or over-permissive APIs rather than attacks, according to Protecto. Kiteworks' 2025 AI governance survey supplies the volume: 26% of organizations say more than 30% of what their staff put into public AI tools is private or sensitive.
Which is why this belongs in an architecture review rather than a privacy policy. The twelve questions below are how you find out which one your vendor is answering from.
Zero Data Retention, and Why Parseur Does Not Claim It
Zero data retention (ZDR) means a provider processes your request in memory and writes nothing to disk. No stored inputs, no stored outputs, no logs of content. It is the strictest posture available, and procurement teams increasingly ask for it by name instead of asking about training.
It is also oversold constantly, so it pays to know what ZDR tends to leave out.
- Billing metadata survives. Timestamps, token counts and latency have to be written down somewhere, or nobody can invoice you.
- Abuse and safety monitoring can sit outside the arrangement entirely, which is a human review path you were never told about.
- Not every endpoint is ZDR-eligible. Calling one that is not is a quiet decision to step outside your own agreement, usually taken by an engineer who never read it.
Here is the part most vendor pages skip. Parseur is not a zero data retention product, and no document parser that lets you re-run an extraction or review last month's results honestly can be. Reprocessing a document requires the document to still be there. What Parseur gives you instead is a retention window you control, from one day up to unlimited depending on your plan, with automatic deletion the moment the window closes.
If a vendor offers you a full document history and zero data retention in the same breath, ask which one they are about to take away.
The 12 Questions to Ask a Document Extraction Vendor About AI Training and Retention

Copy this into your vendor questionnaire. Each one has a short answer if the vendor has done the work, and a long pause if it has not.
- Do you use our uploaded documents, extracted fields, metadata, corrections or outputs to train, fine-tune, evaluate or benchmark any model?
- Does that answer also cover anonymised, de-identified, aggregated or derived versions of our data?
- Which third-party AI, OCR or cloud subprocessors receive our documents? Name each one and say where it runs.
- Are those subprocessors contractually barred from using our data for training or service improvement?
- How long is each type of data kept: source documents, extracted fields, processing logs, error logs, caches and backups?
- Can we configure the retention window ourselves, and what is the shortest one you offer?
- Can we delete data on demand, and what is the deletion SLA in writing?
- Does deletion reach backups, and how long until they are overwritten?
- Can your employees or contractors view our documents, under what circumstances, and is that access logged?
- Is customer data ever copied into development, testing, QA or demo environments?
- Where is our data stored and processed, and can it ever leave that region?
- Will these commitments live in the signed DPA or MSA, and will they override your online terms if the two ever disagree?
A vendor that answers all twelve in one smooth paragraph is guessing. A vendor that answers with names, numbers and a clause reference has been asked before.
Parseur's Answers, Including the Awkward One

Parseur is a document parser. You send it supplier invoices, it sends back structured fields, and it does not train on customer data. Here is what happens to those invoices, mapped to the questions above.
| Question | Parseur's answer |
|---|---|
| 1. Model training | Customer data is never reused to train Parseur's AI models, and is never sold. Every extraction runs to produce your output and nothing else. |
| 5 and 6. Retention | The document retention period is configurable from one day up to unlimited, depending on your plan. Documents are deleted automatically when the window closes, and on demand whenever you ask. |
| 11. Location | EU-hosted and GDPR-native, with data residency maintained in the EU. The hosting data centre is ISO 27001 certified, which is the data centre's certificate rather than a Parseur certification. |
| 12. Roles and paper | You are the data controller, Parseur is the data processor, and the Data Processing Agreement governs what that allows. |
| Regulations | Aligned with EU GDPR, UK GDPR, California CCPA and CPRA, and Singapore PDPA. Detail on Privacy and GDPR at Parseur. |
| Certifications | Parseur is SOC 2 Type II compliant. HIPAA is still in progress. Parseur is not HIPAA certified today. |
| Security controls | Encryption in transit and at rest, role-based permissions and audit trails, plus regular third-party security testing. |
| 2, 3, 4, 7, 8, 9 and 10 | Not answered on this page. Ask us, and put the answer in the signed agreement rather than a web page we could quietly edit next quarter. |
Writing "in progress" where a logo would look better costs a little swagger and buys a lot of credibility. A reviewer who catches one overstated certification stops believing the other eleven answers, and they are right to.
Two of those rows you can verify before you talk to anybody. The Data Processing Agreement is a document you can read in full before signing, and the retention window is a setting you can see in your own account rather than a promise about one.
Your documents come in, structured fields go out, and nothing from them ends up in the model. The files sit in a retention window you control, and they leave when you say so. That is an architecture answer, and architecture is the only kind of answer worth signing. - Sylvain, CTO at Parseur
What OpenAI, Microsoft, Anthropic, AWS and Google Actually Commit To
Enterprise buyers dragged the whole market toward explicit guarantees, and the shape of a good answer is settled now.
- OpenAI splits by product. Consumer ChatGPT data may be used to improve models, while enterprise and API tiers are excluded by default with opt-out controls documented.
- Microsoft leans on isolation. Copilot for Microsoft 365 and Azure OpenAI exclude customer content from foundation model training and publish a default retention window for transient processing data.
- Anthropic commits publicly to not training on enterprise data.
- Google Cloud and AWS both publish training exclusions for Document AI and Textract in their data governance terms, though AWS requires you to apply an explicit organisation-level opt-out policy rather than defaulting to it. Excluded by default and excluded once someone remembers to click are not the same commitment.
Notice what they have in common. The answer lives in a governance document rather than a blog post, and it arrives with a named default. That is the bar, and it is the bar this page is trying to clear.
Why This Question Now Blocks Purchases
Buyers are not asking out of curiosity. A 2025 survey by The Futurum Group found that 52% of organizations prioritize AI vendor technical expertise, with data handling and privacy controls close behind at 51%. Over the same period, the 2025 Investment Management Compliance Testing Survey found that 57% of compliance officers identified AI usage as their top compliance concern, and PwC's 2025 AI agent survey found 28% of business leaders rank lack of trust in AI agents as a top barrier to wider deployment.
The public mood underneath those numbers is colder still. Roughly 70% of adults do not trust companies to use AI responsibly, and over 80% expect some level of misuse of their data, according to Protecto. Your customers have already assumed the worst, which is why a mishandling incident costs reputation long before it costs a fine.
Which is the real argument for writing all of this down. A no-training policy is not a feature you add to a comparison table. It is the thing that decides whether the security reviewer signs, and reviewers do not sign paragraphs. They sign clauses.
So send Parseur the twelve questions, then send the identical twelve to everyone else on your shortlist. Start with the Data Processing Agreement and the retention settings on your plan. The comparison is the whole point, and we would rather be read next to the others than instead of them.
Last updated on



