SecurityDraft
Free resource · published 31 July 2026 · SecurityDraft is a trading name of Zebtech Ltd (UK)

The AI-CAIQ Starter Template

A founder's worksheet for enterprise AI security review: the 20 questions enterprise security teams ask AI vendors most often — what each one is really asking, the evidence you should gather before answering, and an honest answer pattern for each.

Scope — read this first. SecurityDraft (a trading name of Zebtech Ltd) drafts prepared responses assembled from the vendor's real documentation. The vendor reviews, verifies, corrects and signs off every answer, and owns the finished pack. SecurityDraft does not attest that any statement is true, does not audit or test controls, and issues no certification or assurance of any kind. We are a drafting-and-mapping service; the truth of each answer is, and remains, the vendor's own representation to its buyer. This line is printed on the cover of every deliverable and in the terms, and marketing copy never uses "certified / audited / assured / guaranteed compliant." The same principle applies to this free template: every answer you write with it is your representation to your buyer.
Free to use. This template is free for any purpose — use it internally, adapt it, share it with your team. No sign-up wall on the content, no licence strings. If it saves your deal, we'd love to hear about it, but you owe us nothing.

What the AI-CAIQ is, and why your buyer suddenly cares

Enterprise security teams have assessed cloud vendors for years using the Cloud Security Alliance's Consensus Assessment Initiative Questionnaire (CAIQ), the questionnaire companion to the CSA Cloud Controls Matrix (CCM). In 2025 the CSA extended that model to AI: the AI Controls Matrix (AICM) — v1.0 released 10 July 2025; current version v1.1, released June 2026 — and its questionnaire companion, the AI-CAIQ. The AICM builds on the CCM's control domains and adds AI-specific controls, spanning 18 domains and 247 control objectives (v1.1) covering model provenance, training data, AI supply chain, model security and more.

Because it is free, public, and published by the body enterprises already trust for cloud assessment, the AI-CAIQ is becoming the default instrument — buyers either send it directly or borrow heavily from it in their own AI vendor questionnaires. Answer it well once, and you have answered most of what any enterprise will ask.

Attribution & where to get the real thing. CAIQ, CCM, AICM and AI-CAIQ are publications of the Cloud Security Alliance and are available free from cloudsecurityalliance.org. This worksheet is an independent study aid produced by SecurityDraft. It is not affiliated with, endorsed by, or a substitute for the CSA or its official instruments. When a buyer sends you the full AI-CAIQ, answer the official document — use this worksheet to prepare. We deliberately describe control areas in plain language rather than quoting CSA control IDs; the official numbering belongs to the CSA instrument itself.

How to use this worksheet

Below are ten control domains and the twenty questions that, in our review of published AI vendor questionnaires, appear most often in one wording or another. For each question you get three things:

What it's really askingThe risk behind the question — answer this and you satisfy the reviewer; answer only the literal words and you get a follow-up round.
Evidence to gatherThe documents and facts to collect before writing anything. An answer with a named source survives review; an answer from memory invites challenge.
Answer patternA fill-in-the-blanks skeleton for a strong, honest answer. [Bracketed italics] are yours to complete — with facts you have verified, not hopes.
Four rules that decide whether your answers survive security review:
  1. Never guess. Every answer is a representation your company is making to a counterparty in a commercial process. If you don't know, find out or say "we will confirm" — an invented answer discovered later can cost the deal and the relationship.
  2. Answer the risk, not the vibe. Reviewers are pattern-matching for specific risks. Marketing language ("enterprise-grade", "bank-level") scores zero; specifics ("AES-256 at rest, TLS 1.2+ in transit, keys in [KMS]") score.
  3. An honest "no" with a compensating control beats a false "yes". "No, we do not yet hold [certification]; we compensate with [control] and have an assessment scheduled for [date]" is a passing answer in most reviews. A "yes" that unravels is a failed review.
  4. Build one canonical answer set. Answer once, precisely, and reuse everywhere — every questionnaire, trust page and sales call should say exactly the same thing. Divergent answers across documents are what reviewers are trained to catch.
Domain 1 of 10

AI governance & accountability

Q1. Do you have a formal AI governance policy, and a named individual accountable for AI risk?

What it's really asking

Is anyone actually in charge, or is AI risk nobody's job? The reviewer wants evidence that AI-specific risk is owned at a level that can say no to a feature.

Evidence to gather
  • Your AI use / AI development policy (even a two-page one — dated and approved beats aspirational and long).
  • The name and role of the accountable owner (CTO, co-founder — a person, not a committee).
  • Where AI risk is reviewed (e.g. a recurring engineering or leadership meeting) and how exceptions are approved.
Answer pattern

Yes. Our AI governance policy (v[X], approved [date]) covers [acceptable use, model changes, data handling in AI features, incident escalation]. Accountability for AI risk sits with [name, role]. Material AI changes are reviewed at [forum/cadence], and exceptions require [role] approval.

Q2. Do you maintain an inventory of the AI models and systems used in your product, and how are material changes reviewed before release?

What it's really asking

If your model provider swaps a model under you — or you swap it yourselves — will anyone notice, assess it, and tell the customer? Undocumented model changes are a top emerging worry for AI buyers.

Evidence to gather
  • A model inventory: each model/AI component, its provider and version, what it does in the product, what data reaches it.
  • Your change-management path for model swaps / prompt or fine-tune changes (PR review, eval gate, release notes).
  • Whether and how customers are notified of material model changes.
Answer pattern

We maintain a model inventory covering [N] AI components: [model — provider — function — data exposure, per row]. Model or prompt changes follow our standard change process ([review + evaluation gate]) and material changes are recorded in [release notes / changelog]; customers are notified of changes that alter data handling via [channel].

Domain 2 of 10

Model provenance & training data

Q3. Which models does your product rely on, who developed them, and do you train or fine-tune models yourselves?

What it's really asking

What is actually inside the box — and how much of the AI risk is yours versus inherited from a foundation-model provider? Vague answers here undermine every later answer.

Evidence to gather
  • The exact models/providers in production (e.g. "hosted API from [provider], model family [X]") — and any self-hosted or open-weight models with their sources.
  • Whether you fine-tune, and on what data; whether any customer data has ever been used in training or fine-tuning.
  • Your provider's published documentation on their models (terms, model cards, security pages) — you'll cite it repeatedly.
Answer pattern

Our product uses [model(s)] from [provider(s)], accessed via [hosted API / self-hosted deployment]. We [do / do not] fine-tune: [if yes: on what data, with what rights; if no: say so plainly]. No customer data is used to train or fine-tune any model [or state the precise, contracted exception].

Q4. What data was used to train or fine-tune the models you control, and how do you ensure you have rights to use it?

What it's really asking

Will we inherit an IP or privacy problem from your training data? For foundation models you don't control, the reviewer expects you to point at the provider — accurately, not evasively.

Evidence to gather
  • For models you fine-tune or train: a description of the dataset(s), their origin, and the rights basis (owned, licensed, permissively licensed, synthetic).
  • For third-party foundation models: the provider's own published statements on training data — reference them, don't restate them as your own knowledge.
  • Any data-provenance records you keep (dataset versions, collection dates, licences).
Answer pattern

For models we control: [dataset description, origin, rights basis, dataset version/date]. For the third-party foundation model(s) we use, training-data practices are documented by the provider at [reference]; we do not have independent visibility beyond the provider's published statements and contractual terms, and we say so rather than speculate.

Domain 3 of 10

Data handling in prompts & outputs

Q5. Is our data used to train models — yours, or any third party's?

What it's really asking

The single most common AI question in 2026 procurement. The reviewer wants a flat, checkable "no" (or a precisely scoped "yes") covering every model in the chain, including your providers' default settings.

Evidence to gather
  • Your foundation-model provider's terms for API/business traffic — specifically whether API inputs/outputs are used for training, and any opt-out/zero-retention configuration you have actually enabled.
  • Your own policy and the technical setting that enforces it (account configuration, contract clause, DPA reference).
  • Your DPA language on this point, if you have it.
Answer pattern

No. Customer data submitted to our product is not used to train models by us, and our model provider(s) [provider] contractually do not train on API traffic under [terms reference / enterprise agreement]. We have additionally [enabled zero-data-retention / opted out of X], configured on [date]. This commitment appears in our DPA at [clause].

Q6. How long are prompts, outputs and any derived data retained, where are they stored, and how are they deleted?

What it's really asking

Where does our data actually live once your AI feature has touched it — including logs, caches, embeddings and your provider's side — and can you make it go away?

Evidence to gather
  • A data-flow map for AI features: prompt → your services → provider → output → storage/logs/caches/vector indices.
  • Retention period at each hop, including your provider's stated API retention window, and storage regions.
  • Your deletion process (on request and at contract end) and how it reaches derived data such as embeddings.
Answer pattern

Prompts and outputs are retained in [store, region] for [period] for [purpose: support / abuse prevention], then deleted automatically. Derived data ([embeddings / indices]) is retained [period/policy]. Our provider retains API data for [provider window, per their published policy]. Deletion on request completes within [N days] and covers [all listed stores].

Domain 4 of 10

Foundation-model & sub-processor supply chain

Q7. Which sub-processors — including AI/foundation-model providers — process our data, and under what contractual terms?

What it's really asking

Who else are we really trusting when we trust you? Model providers are sub-processors, and reviewers now check that your sub-processor list admits it.

Evidence to gather
  • A complete sub-processor list — cloud, model provider(s), analytics, support tooling — with purpose, data categories and region for each.
  • DPAs / data-processing terms in place with each (especially the model provider).
  • How customers are notified of sub-processor changes.
Answer pattern

Our sub-processors are: [name — purpose — data categories — region — terms, per row, model provider(s) included]. Each is engaged under a data-processing agreement; the list is published at [URL / on request] and changes are notified via [channel] with [N days] notice.

Q8. What happens if your model provider deprecates a model, degrades, or has an outage — and what does that do to your service to us?

What it's really asking

Concentration risk and continuity: is a third party's model roadmap a single point of failure for the service we're buying?

Evidence to gather
  • Your fallback design: alternative model/provider, degraded non-AI mode, queuing — whatever is actually built (only claim what exists).
  • Provider deprecation policy and your tested migration path between model versions (your eval suite is the evidence).
  • How AI-feature availability is reflected in your SLA/status page.
Answer pattern

Model dependencies are isolated behind [abstraction layer]. On provider outage the product [fails over to X / degrades to Y with the following functionality preserved]. Model version migrations are validated against our evaluation suite ([scope]) before rollout. AI-feature availability is covered by [SLA/status page reference].

Domain 5 of 10

Identity & access control

Q9. How is internal access to customer data — including prompts and outputs — restricted, logged and reviewed?

What it's really asking

Can your engineers read our prompts? AI products pipe unusually sensitive free-text through logs and debugging tools, and reviewers know it.

Evidence to gather
  • Your access model: role-based access, least privilege, who can access production data and under what approval.
  • MFA/SSO enforcement for staff; access-review cadence and the last review date.
  • Whether prompts/outputs appear in internal logs or debugging tools, and the controls on that (redaction, restricted access, retention).
Answer pattern

Production access is role-based and least-privilege: [N] staff hold it, granted by [process] and reviewed [cadence; last review date]. MFA is enforced on all staff accounts; access is via SSO. Customer prompts/outputs in operational logs are [redacted / access-restricted to role X / retained N days], and all production data access is logged.

Q10. Does your product enforce tenant isolation, and can AI features be configured or disabled per tenant?

What it's really asking

Two things: architectural separation between customers, and whether a cautious enterprise can adopt your product while limiting the AI parts — an increasingly common procurement ask.

Evidence to gather
  • How tenancy is enforced (separate schemas/databases, row-level scoping, per-tenant keys) — including in vector stores and caches used by AI features.
  • Which AI features can be toggled/scoped per tenant, and how (admin console, contract flag).
  • Any per-tenant data-residency or model-choice options, if genuinely offered.
Answer pattern

Tenant data is isolated by [mechanism], including AI-specific stores ([vector index / cache scoping]). AI features can be [disabled / configured] per tenant via [mechanism]; with AI features off, [what the product still does].

Domain 6 of 10

Logging & monitoring

Q11. What do you log about AI interactions, and do those logs contain our data?

What it's really asking

Whether your observability stack is a shadow copy of the customer's data — and whether you can support an investigation without hoarding content forever.

Evidence to gather
  • What is logged per AI request (metadata only vs. full prompt/output content) and where those logs go.
  • Log retention, access restrictions, and whether content logging can be reduced for a given customer.
  • Third-party observability/analytics tools that see this data (they belong on the sub-processor list, Q7).
Answer pattern

Per AI request we log [metadata: timestamps, model, token counts, latency, error codes]; prompt/output content is [not logged / logged to a restricted store, retained N days, accessible to role X only]. Logs live in [system, region] with retention of [period]. [Content logging can be disabled per tenant / is fixed — say which.]

Q12. How do you monitor model behaviour in production — quality drift, misuse, anomalous or unsafe output?

What it's really asking

Once the demo is over and the model is live, would you actually notice if it started behaving badly — before we do?

Evidence to gather
  • Production metrics you actually track (error/refusal rates, output-length or classifier anomalies, user feedback signals, provider incident alerts).
  • Alerting thresholds and the escalation path when a model misbehaves (who is paged; can you roll back a model/prompt?).
  • Abuse/misuse detection on inputs (rate limits, abuse-pattern detection) — whatever genuinely exists.
Answer pattern

We monitor [metrics] against [thresholds], alerting to [on-call rota]. Regressions trigger [rollback of model/prompt version via mechanism]. User-reported output issues route to [channel] and are triaged within [time]. Misuse controls include [rate limiting / abuse detection].

Domain 7 of 10

Evaluation, testing & release gates

Q13. How do you evaluate model accuracy, safety and fitness for purpose before release — and on every change?

What it's really asking

Do you have engineering discipline around a probabilistic component, or do you ship on vibes? This is where AI-literate reviewers separate serious vendors from wrappers.

Evidence to gather
  • Your evaluation suite: what it tests (task accuracy, refusal correctness, harmful-output checks, regression cases from production issues), how many cases, how often it runs.
  • Release criteria: what score/behaviour blocks a release; who signs off.
  • Records of a recent eval run (date, result) — reviewers love a concrete example.
Answer pattern

Every model, prompt or retrieval change runs our evaluation suite: [N] cases covering [task accuracy / safety / regression scenarios], run in CI. Releases are blocked unless [criteria]; [role] signs off. Last full run: [date, result]. Production incidents are converted into new evaluation cases.

Q14. Have you red-teamed the system — prompt injection, jailbreaks, data exfiltration — and how were findings handled?

What it's really asking

Has anyone adversarial ever attacked this system on purpose? "Our provider red-teams the model" is only half an answer — the reviewer means your application: your prompts, tools and retrieval.

Evidence to gather
  • Any internal or third-party adversarial testing of your application: scope, date, method (even a structured internal exercise counts — describe it honestly as such).
  • Findings and their remediation status; retest evidence.
  • Your provider's published safety/red-teaming documentation, cited as theirs, layered under yours.
Answer pattern

Our application layer was adversarially tested on [date] by [internal team / named firm], covering [prompt injection via user input and retrieved content, jailbreaks, data-exfiltration attempts, tool abuse]. [N] findings were identified; [N] are remediated [dates], [N] accepted with rationale. Model-level safety testing is documented by [provider, reference]. Next exercise: [date].

Domain 8 of 10

Model & application security

Q15. What controls mitigate prompt injection and unsafe handling of model output?

What it's really asking

The reviewer is checking you know the failure modes specific to LLM applications — untrusted text steering the model, and model output being trusted downstream (rendered, executed, or fed to tools).

Evidence to gather
  • Input-side controls: separation of instructions from untrusted content, sanitisation of retrieved documents, limits on what user content can trigger.
  • Output-side controls: output treated as untrusted (escaped before rendering, never executed), restricted tool/function scopes, human-in-the-loop for consequential actions.
  • How your design maps to recognised guidance you actually follow (e.g. OWASP's LLM application risk list) — cite only if you have genuinely used it.
Answer pattern

Untrusted content (user input, retrieved documents) is [delimited / sanitised / never concatenated into system instructions]. Model output is treated as untrusted: [escaped before rendering; never executed; tool calls restricted to allow-listed actions with scopes X]. Consequential actions require [user confirmation]. Design reviewed against [guidance, date].

Q16. How do you prevent one customer's data appearing in another customer's outputs?

What it's really asking

Cross-tenant leakage through the AI path — retrieval pulling from the wrong tenant, shared caches, or fine-tuning on pooled data. This is the AI-specific twist on Q10, and reviewers ask it separately on purpose.

Evidence to gather
  • Tenant scoping in retrieval/RAG (per-tenant indices or enforced filters — which, exactly) and in any response caches.
  • Confirmation that no cross-customer training/fine-tuning occurs (ties back to Q3/Q5).
  • Tests you run that specifically probe cross-tenant leakage (eval cases, integration tests).
Answer pattern

Retrieval is scoped per tenant by [per-tenant index / mandatory tenant-ID filter enforced at query layer]; caches are [tenant-keyed / disabled]. No model is trained or fine-tuned on pooled customer data. Cross-tenant isolation is exercised by [tests, frequency], most recently [date].

Domain 9 of 10

Incident response & vulnerability management

Q17. Describe your security incident response process and customer notification commitments. Do they cover AI-specific incidents?

What it's really asking

When something goes wrong, will we find out from you or from the press? "AI-specific" is the modern addition: does your definition of an incident include model-mediated data exposure or harmful output at scale?

Evidence to gather
  • Your incident response plan (roles, severity levels, escalation) and the date it was last tested or exercised.
  • Contractual notification window for incidents affecting customer data (e.g. within 48/72 hours) and the channel.
  • Whether your incident taxonomy explicitly includes AI events: cross-tenant output leakage, mass harmful output, model compromise, provider-side incidents.
Answer pattern

Incidents follow our IR plan (v[X], last exercised [date]): [severity levels, on-call, escalation]. Customers are notified of incidents affecting their data within [hours] via [channel]. The plan explicitly classifies AI-specific events ([examples]) and covers incidents originating at our model provider, which we track via [provider status/notification channel].

Q18. How do you manage vulnerabilities and penetration testing — including the AI components of your stack?

What it's really asking

Standard security hygiene, extended to the parts of your stack a classic pentest scope used to skip: model endpoints, retrieval pipelines, tool integrations.

Evidence to gather
  • Dependency and infrastructure scanning cadence and tooling; patch SLAs by severity.
  • Your most recent penetration test: date, scope (did it include AI features?), firm, and whether a summary/attestation letter is shareable under NDA.
  • Your vulnerability disclosure route (security.txt, security@ address).
Answer pattern

Dependencies and images are scanned [cadence/tooling]; remediation SLAs are [critical: X days; high: Y]. Our last penetration test was [date] by [firm], scoped to include [AI features: prompt injection, model endpoints, RAG pipeline]; a summary is available under NDA. Vulnerabilities can be reported via [route].

Domain 10 of 10

Compliance posture & frameworks

Q19. Which certifications or attestations do you hold (e.g. SOC 2, ISO/IEC 27001, ISO/IEC 42001)? If none, what is your roadmap?

What it's really asking

A maturity proxy — and a trap for overclaiming. Reviewers verify certificates. An honest "not yet, here's the roadmap and here's what we do meanwhile" passes far more reviews than founders expect; a claimed certification that turns out to be "in progress" fails instantly.

Evidence to gather
  • Exactly what you hold: certificate/report type, issuer, scope, date. (SOC 2 Type I vs Type II matters; say which.)
  • If in progress: the stage you can prove (engagement letter signed, audit window booked) and the expected date.
  • The controls you operate today that the certification would attest to — your compensating story.
Answer pattern

We currently hold [certification, scope, date — or "no third-party certifications"]. [If in progress: "Our SOC 2 Type I audit window is booked for [date] with [firm]."] Meanwhile we operate [the material controls: encryption, access control, logging, IR — cross-reference your other answers], and this questionnaire documents them. We are happy to walk your team through any control directly.

Q20. How does your product align with the EU AI Act, NIST AI RMF, or ISO/IEC 42001? What is your system's EU AI Act risk classification?

What it's really asking

Whether you have done the homework on the frameworks their own governance team reports against. They want your reasoned self-assessment and mapping — not a legal opinion, and not a shrug.

Evidence to gather
  • Your self-assessed EU AI Act position: whether your use case plausibly falls in a high-risk category or is a general-purpose/limited-risk deployment, and the reasoning. (The Act's obligations phase in through 2026–2027 — note which apply to a company like yours, and take legal advice for a definitive classification; a questionnaire answer is not the place to improvise one.)
  • A simple mapping of your existing controls to NIST AI RMF functions (Govern / Map / Measure / Manage) and, if relevant, to ISO/IEC 42001's AI-management-system themes.
  • Transparency measures you already provide (AI-generated content labelling, user disclosure) — concrete beats abstract.
Answer pattern

Our self-assessment (dated [date], reviewed by [role/counsel]): the product is [classification + one-sentence reasoning] under the EU AI Act as currently phased in; applicable obligations ([e.g. transparency to users]) are met by [measures]. Our controls map to NIST AI RMF as follows: [Govern → Q1–2; Map → Q3–8; Measure → Q13–14; Manage → Q12, Q17]. This is our own assessment, not legal advice to you or your buyer.

After the worksheet: the full instrument

These twenty answers cover the high-frequency core, and — because the rules above force named sources — you will have assembled most of the evidence the full CSA AI-CAIQ asks for. When a buyer sends the complete instrument (or their own variant, or a SIG-Lite), work domain by domain from your canonical answer set: download the official AI-CAIQ free from the Cloud Security Alliance, and answer every control from evidence, never from memory.

If you'd rather not spend the next two weeks on this. SecurityDraft turns an AI startup's own documentation into the finished pack — a completed AI-CAIQ, a public trust page, and an EU AI Act / ISO 42001 / NIST AI RMF mapping appendix — done-for-you at a fixed price from £900 (ex-VAT), typically in days. You review, correct and sign off every answer, and you own every word: we draft; the claims stay yours. We don't audit or certify anything — and nothing in this template or our paid packs is an attestation, audit, certification or assurance of any control.