An AI vendor DPA has to do everything a standard SaaS DPA does, and then close five gaps the older format was never built for: model training use, the difference between inference and training, foundation model subprocessor chains, prompt and log retention, and how embeddings or vector stores get deleted. If your AI vendor handed you a DPA that reads like a copy-pasted SaaS template, check it against all five gaps.
Key takeaway
A DPA written before generative AI existed cannot protect you from what generative AI actually does with your data.
The Facts
- GDPR Article 28 requires a written DPA with every processor. This applies to AI vendors exactly the same as it applies to any SaaS tool. There is no carve-out for “the tool works differently.”
- Most AI vendors’ default commercial terms permit using your data to improve their models unless you explicitly negotiate that out. This is not true of most traditional SaaS tools.
- A DPA that only says “no training” is incomplete. It also needs to confirm that the underlying foundation model provider is bound by the same commitment.
- 72 hours is the regulatory clock GDPR runs on for breach notification. Many AI vendor drafts leave the vendor’s own notification window undefined, which quietly shifts risk onto you.
Where AI DPAs and Standard SaaS DPAs Actually Diverge
Most legal teams already know how to review a SaaS DPA. The problem is that AI vendors get reviewed against that same checklist, and the checklist has blind spots. Here is where the two documents part ways.
| Area | Standard SaaS DPA | AI Vendor DPA |
| Data use | Covers storage, processing, and hosting | Must also restrict training, fine-tuning, and model improvement |
| Retention | Covers stored records and backups | Must also cover prompts, outputs, chat logs, embeddings, and caches |
| Subprocessors | Lists infrastructure and support vendors | Must also disclose the foundation model provider actually running inference |
| Data flow | Mostly static and defined at signing | Can include live inference calls to third-party model APIs, sometimes per request |
| Deletion | Deletion from databases and backups | Must also cover model artifacts, vector stores, and any fine-tuned data |
| Processing purpose | “Deliver the service” | Must separate inference (generating an output) from training (improving the model), since these are legally distinct activities |
If a vendor’s DPA does not draw the line between inference and training, assume both are happening on your data until they confirm otherwise in writing.
Who Is the Controller and Who Is the Processor
In nearly every enterprise AI deployment:
- You are the data controller.
- The AI vendor is the data processor.
Your DPA needs to say this explicitly and needs to prohibit the vendor from acting like a controller, meaning it cannot reuse your data for its own purposes, including improving its own models, without your written consent.
The Mandatory Clauses, Updated for AI
| Clause | Why it matters | AI-specific addition |
| Purpose limitation | Prevents reuse beyond your use case | Must name inference and training as separate purposes |
| No-training guarantee | Ensures data is not used to train models | Must apply to prompts, outputs, uploads, and metadata, and must survive termination |
| Subprocessor disclosure | Shows where data may flow | Must name the foundation model provider, not just infrastructure vendors |
| Security measures | Required under GDPR Article 32 | Must cover inference pipelines and logging systems, not just storage |
| Data retention and deletion | Enables the right to be forgotten | Must include prompt logs, embeddings, and vector stores, not just database rows |
| Audit and cooperation | Required for regulatory inquiries | Should include the right to confirm the vendor’s model provider settings, such as zero-retention API tiers |
| Breach notification | Defines timelines and responsibilities | Vendor’s own notification window should be set at 24 to 48 hours so you can still hit your 72-hour regulatory deadline |
| Cross-border transfers | Ensures lawful international processing | Must account for the model provider’s processing location, which may differ from the vendor’s own |
| Liability cap | Determines what you can actually recover if something goes wrong | Should carve out breaches and IP claims tied to AI outputs from the general liability cap, not bury them inside it |
If any row is missing or vague, treat the DPA as incomplete for AI use.
Check the Terms of Service Too, Not Just the DPA
This trips up more teams than it should. AI vendors often put their right to use your data for “service improvement” in the Terms of Service, not the DPA. The DPA looks clean. The ToS is where the training rights actually live.
When the two documents conflict, the DPA typically takes priority, but only if your legal team actually reads both and only if the DPA explicitly says it overrides conflicting terms elsewhere in the agreement. Do not review the DPA in isolation. Pull the ToS and any acceptable use policy into the same review.
Language That Should Stop You From Signing
Watch for:
- “We may use data to improve our services”
- “Anonymized or aggregated use” with no definition of what that means technically
- No stated retention period
- No deletion SLA
- No named list of subprocessors
- No audit or verification rights
- Silence on which company actually runs the underlying model
Key takeaway
Vague language always resolves in the vendor’s favor, never yours.
How a No-Training Clause Should Actually Be Written
A real no-training clause is explicit and unconditional. It should state that:
- Your data is not used to train, fine-tune, or improve models, including third-party models the vendor relies on.
- This applies to prompts, uploads, chat logs, outputs, and metadata, not just the files you upload.
- The restriction survives contract termination.
Without all three, you cannot honestly tell your own customers or a regulator that your data was not reused.
The Part Most DPAs Skip: The Model Chain Underneath Your Vendor
Here is the risk that standard SaaS review checklists miss entirely. Your AI vendor is often not the company actually running the model. Your prompt frequently gets passed to a separate foundation model provider, such as OpenAI, Anthropic, or Google, sitting underneath your vendor as an undisclosed subprocessor.
A DPA that stops at “we don’t train on your data” is not enough if it does not also confirm that the foundation model provider under the hood is contractually held to the same no-training and retention rules. Ask your vendor directly:
- Which model provider or providers actually process my data during inference
- Whether the vendor uses a zero-data-retention or enterprise API tier with that provider
- Whether that provider’s own terms are attached to, or referenced by, your DPA
If your vendor cannot answer these three questions clearly, the no-training clause you signed may not mean what you think it means.
Here is why this matters more than it sounds. Once your data goes into a training run, it stops being a discrete record you can point to and delete. It becomes statistical weights spread across the model. You cannot un-train a model the way you can delete a database row. This is exactly why the no-training clause has to be airtight before any data changes hands, not fixed after the fact.
A Working Checklist You Can Use Today
Before you sign, confirm the DPA answers all of these:
- Does it name inference and training as separate, separately restricted purposes
- Does the no-training clause explicitly cover prompts, outputs, and metadata, and does it survive termination
- Does it name the actual foundation model provider, not just the vendor’s own company name
- Does it set a retention limit on logs, prompts, and outputs, not only on stored files
- Does it address deletion of embeddings, vector stores, and any fine-tuned artifacts
- Does it set the vendor’s own breach notification window at 24 to 48 hours
- Does it give you audit rights, or at minimum a current SOC 2 Type 2 report on request
- Does it specify the cross-border transfer mechanism for both the vendor and its model provider
Anything left blank is a negotiation item, not a formality to skip.
Validating the DPA After You Sign
Signing is not the finish line. Confirm that the paper matches reality:
- Map each clause to an actual product control
- Confirm deletion and retention workflows exist and actually run on schedule
- Verify access controls and logging match what the DPA describes
- Check that the vendor’s disclosed subprocessors match what is actually in production
- Include the vendor in your Data Protection Impact Assessments and ongoing risk reviews
A DPA only protects you if the system underneath it actually enforces what the document promises.
How CustomGPT.ai Structures Its DPA
CustomGPT.ai’s DPA is built around this exact set of AI-specific gaps, not the generic SaaS template. In practice that looks like: each bot runs in its own isolated data silo, so nothing you upload is shared across bots, even within your own account. Data is encrypted with AES-256 at rest and SSL in transit. CustomGPT.ai runs on OpenAI’s API inside a private AWS VPC, and OpenAI’s own policy confirms API data is not used for model training. You can delete source files immediately after processing if you do not want them retained at all. The platform is SOC 2 Type 2 certified and supports SAML 2.0 for enterprise identity management.
One important plan detail: the DPA is an Enterprise-tier feature, executed through a signed form, and CustomGPT.ai does not customize it on a case-by-case basis. If a DPA is a requirement for your deployment, that is worth confirming at the Enterprise level early, not after you have already built on a lower tie
Review a DPA built for AI, not retrofitted
SOC 2 Type 2 GDPR readySAML SSO
Trust Center →If you are currently reviewing AI vendors and want to see what a DPA built for this looks like in practice, you can try CustomGPT.ai free or talk to our team about Enterprise plans and DPA requirements.
What a Strong AI DPA Actually Buys You
- Faster legal sign-off, since your team is not stuck negotiating clauses that should have been there from the start
- Lower regulatory exposure
- Easier trust conversations with your own customers
- Room to actually scale AI adoption instead of stalling it in review
A well-written DPA turns from a blocker into something your legal team can approve quickly and confidently.
Conclusion
A standard SaaS DPA was never built to handle model training, inference pipelines, or a foundation model provider sitting underneath your vendor. An AI vendor DPA needs to explicitly restrict training, distinguish inference from training, disclose the real model provider chain, cover prompt and log retention, and support deletion all the way down to embeddings and fine-tuned artifacts. If any of that is missing, the agreement is a SaaS DPA wearing an AI vendor’s logo.
Frequently Asked Questions
Does my company need a DPA with every AI vendor?
Yes, if the vendor processes any personal data belonging to your customers, employees, or prospects. GDPR Article 28 makes this mandatory, and CCPA/CPRA imposes an equivalent requirement in the US. This applies regardless of the vendor’s size or how the tool works internally.
What should I look for in a DPA with an AI vendor?
Explicit limits on data use, a clear no-training guarantee that survives termination, disclosure of the actual foundation model provider, defined retention and deletion covering prompts and embeddings, and audit rights. A strong AI DPA should make each point testable.
How is an AI DPA different from a standard SaaS DPA?
A SaaS DPA covers storage, processing, and hosting. An AI DPA has to additionally cover model training, the distinction between inference and training, the foundation model provider underneath the vendor, and deletion of AI-specific artifacts like embeddings and vector stores.
Who is the controller and who is the processor in an AI deployment?
In most enterprise setups, you are the controller and the AI vendor is the processor. The DPA should say this explicitly and prohibit the vendor from reusing your data for its own purposes, including model improvement.
What is the difference between inference and training, and why does my DPA need to separate them?
Inference is the AI generating an output for you in real time. Training is using data to improve the underlying model. These are legally distinct activities. A DPA that only bans “training” may still permit your data to be logged, reviewed, or reused in ways that fall outside that narrow definition.
Why does the foundation model provider matter if I only contracted with the vendor?
Many AI vendors route your prompts to a separate company’s model, such as OpenAI, Anthropic, or Google. If your DPA does not confirm that no-training and retention commitments apply to that provider too, your data protection stops at the vendor and does not extend to where your data actually goes.
What language in an AI vendor’s DPA should be a red flag?
Phrases like “we may use data to improve our services,” undefined “anonymized or aggregated” use, no stated retention period, no deletion SLA, no named subprocessor list, and no audit rights. Each of these quietly shifts risk back onto you.
Does CCPA or CPRA require anything similar for US-based teams?
Yes. Under CCPA and CPRA, a written agreement with a service provider is required before personal information can be shared, and it must restrict the vendor from using that data outside the scope of the services, which includes model training.
How should a no-training clause actually be worded?
It should be explicit and unconditional: data is not used to train, fine-tune, or improve models, this applies to prompts, uploads, outputs, and metadata, and the restriction survives contract termination.
How do I check that a DPA is actually being enforced, not just signed?
Map each clause to a real product control, confirm deletion and retention workflows run as described, verify the disclosed subprocessor list matches production reality, and include the vendor in your Data Protection Impact Assessments.

Arooj Ejaz is the Marketing Operations Lead at CustomGPT.ai, where she works on content, growth operations, and go-to-market programs for AI agent and chatbot solutions.