Procurement field guide · Checked August 13, 2026

Do not buy the demo. Evaluate the whole AI vendor relationship.

A capable model is only one part of an AI purchase. The decision also includes data use, security, output quality, human review, billing, support, contract promises, integrations, and the cost of leaving. This checklist turns those questions into evidence you can compare.

Michael Okeje

AI procurement research and vendor-risk analysis · Last updated August 13, 2026

Start with the use case, not the vendor shortlist

A vendor can be excellent for one workflow and a poor fit for another. Before opening a comparison spreadsheet, write what the system will do, what it will not do, what information it will see, what action it may take, and what a human will review. A drafting assistant that uses public information deserves a different assessment from an agent that changes customer records or influences employment, credit, health, or access decisions.

This order also prevents a common procurement mistake: allowing a compelling demo to define the problem. The demo usually shows the happy path with curated inputs. Your evaluation should include your actual cases, awkward inputs, permissions, review burden, volume, and failure recovery. If a vendor cannot let you test the workflow honestly, that is evidence about the fit.

The six-part AI vendor checklist

Use the questions below in the same order for every vendor. Ask for a written answer and record the product tier, region, model, date, and person who answered.

1. Use case and scope

  • What exact workflow is the tool supporting?
  • Who owns the outcome and who is affected by mistakes?
  • Is the product a draft assistant, recommendation system, agent, or system of record?
  • What is explicitly out of scope for this purchase?
  • What evidence would make us reject the tool?
Ask them to show: A one-sentence use case, data map, risk tier, owner, and acceptance test.

2. Data and privacy

  • What prompts, files, outputs, feedback, logs, and metadata are collected?
  • Are inputs used for training, product improvement, abuse monitoring, or human review?
  • Where are data and backups processed and stored?
  • What are the retention and deletion controls?
  • Which subprocessors can access the data?
  • Can the vendor support our required data-processing and privacy terms?
Ask them to show: Current privacy policy, data-processing terms, retention settings, subprocessor list, and written answers for the purchased tier.

3. Security and access

  • How are users authenticated and deprovisioned?
  • Are MFA, SSO, role permissions, audit logs, encryption, and network controls available?
  • How does the vendor manage vulnerabilities and incidents?
  • Can our organization restrict tools, exports, connectors, and administrator actions?
  • What independent security reports or certifications apply to this product?
Ask them to show: Security documentation, attestation or report where applicable, access-control configuration, incident process, and contract commitments.

4. Quality and evaluation

  • What does the vendor measure, and what does it not claim to measure?
  • Can we test representative and difficult cases before purchase?
  • How are hallucinations, omissions, bias, unsafe actions, and citations handled?
  • What happens when the model, retrieval source, or tool behavior changes?
  • Can we export traces or results for our own review?
Ask them to show: A shared test set, rubric, results by task type, failure examples, model/version record, and monitoring plan.

5. Cost and commercial terms

  • What is the real billing unit: seat, credit, minute, token, export, task, or API call?
  • What happens to retries, failed runs, overages, storage, and support?
  • Can usage be capped, alerted, approved, and audited?
  • What are minimum commitments, renewal terms, price-change rights, and cancellation rules?
  • What internal work is needed for integration, review, training, and support?
Ask them to show: A three-scenario total-cost model: normal, high-volume, and difficult-case usage, with all implementation and review costs.

6. Resilience and exit

  • Can we export source data, prompts, configurations, outputs, metadata, and audit records in usable formats?
  • How long does export and deletion take, and can it be verified?
  • Can another provider or our internal process reproduce the core workflow?
  • What happens if the vendor changes its model, API, price, region, or terms?
  • What support exists during migration or a service outage?
Ask them to show: An exit test, export sample, fallback process, contract language, and named owner for offboarding.

The vendor questionnaire is not the contract

A sales answer can clarify a product, but it may not bind the provider. The FTC's vendor-security guidance recommends spelling expectations out in writing and verifying compliance. Treat material promises about training, retention, security, uptime, incident notice, subprocessors, deletion, and support as contract questions. If a promise matters to the risk decision, it should not live only in a call recording or a sales email.

Translate claims into tests

“Enterprise security” should become a list of controls and evidence. “Accurate” should become representative cases and a rubric. “Your data is private” should become a data-flow and contract review.

Translate tests into terms

If a requirement is essential, state who must do what, by when, with what remedy if it fails. Confirm which documents govern when the marketing page, help center, order form, and contract differ.

Record the exception

If the product cannot meet a requirement, decide whether the use case can change, a compensating control is sufficient, or the vendor should be rejected. Do not bury an unresolved exception in a risk register nobody owns.

Keep the checked date

Vendor terms, models, subprocessors, and prices change. Store the date and product tier beside each answer, and define what change triggers a re-review.

A weighted scoring sheet that does not let a cheap demo win

Use weights that reflect the consequence of failure. For a low-risk writing tool, usability and cost may carry more weight. For a workflow handling customer records, security, privacy, evaluation, and exit should be hard gates rather than points that a slick interface can compensate for.

AreaSuggested weightScore 1-5 when evidence is sufficientAutomatic concern
Workflow fit and quality25%Representative test cases pass the required rubric.Vendor only permits a scripted demo.
Data and privacy20%Data use, retention, access, and deletion are clear for the tier.Answer depends on an undocumented setting.
Security and controls20%Access, logging, incident, and evidence match the risk.No usable access or incident path.
Total cost and support15%Three usage scenarios are affordable and support is defined.No cap or visibility on variable usage.
Integration and operations10%The team can implement, monitor, and own it.Critical knowledge remains vendor-only.
Exit and resilience10%Data, fallback, and migration are tested.No export or service-change plan.

A score is a decision aid, not a substitute for judgment. Set hard gates for data misuse, unauthorized access, unacceptable quality, and legal or policy constraints. A vendor should not pass because its strengths elsewhere mathematically outweigh a non-negotiable failure.

The five questions buyers forget

1

What happens when the vendor silently changes the model or retrieval behavior?

2

Can we prove what the system saw, produced, and did after a customer or regulator asks?

3

What is the cost of a difficult customer, long document, repeated retry, or failed integration?

4

Who is responsible for reviewing an output that is plausible but wrong?

5

How do we continue the workflow during an outage, account suspension, price increase, or vendor exit?

These questions move the decision beyond feature comparison. They also connect directly to the site's AI governance guide and AI evaluation guide, which cover ownership, evidence, monitoring, and incident response after the purchase.

Run an exit test before the tool becomes infrastructure

Export a sample of the records you would need to preserve: source files, prompts, configurations, outputs, feedback, metadata, and audit events. Ask a second person to use the export without the vendor's dashboard. If the workflow cannot be reconstructed, the business has already accepted lock-in.

Then write the fallback process. It may be slower, manual, or less elegant, but it should preserve the essential customer or operational outcome. Identify the credentials to revoke, integrations to disable, data to delete, customers to notify, and staff who can resume the work. An exit plan is not a prediction that the vendor will fail; it is an acknowledgement that a critical process should have a recovery path.

Export a representative dataset and configuration
Verify the export is readable outside the product
Document the manual or alternate workflow
Test a small fallback run
Record deletion and retention evidence
Set review triggers for price, model, terms, or ownership changes

A practical 14-day buying process

  1. Days 1-2: define the workflow, data, risk, owner, baseline, and rejection criteria.
  2. Days 3-5: shortlist vendors and request the same written answers, product tier, terms, and evidence from each.
  3. Days 6-8: run the same representative test set and record output quality, edits, latency, cost, and failures.
  4. Days 9-10: review privacy, security, subprocessors, access, incident, support, and contract commitments.
  5. Days 11-12: calculate normal, high-volume, and difficult-case cost; run a small export and fallback test.
  6. Days 13-14: score the evidence, document exceptions, negotiate material terms, and make a scale, pilot, redesign, or reject decision.

If the decision is to run a controlled trial, use the 30-day AI pilot plan rather than opening the tool to the whole organization. If the company is building its own agent, see how to build an AI agent and AI agent business models.

Frequently asked questions

What should I ask an AI vendor before buying?

Ask what data the vendor collects, stores, uses for training, and shares; where it is processed; who can access it; how outputs are evaluated; what happens when the model or subprocessor changes; how costs are calculated; what support and incident commitments exist; and how you can export and delete your data if you leave. Ask for evidence and contract language, not only marketing answers.

How do you evaluate the security of an AI tool?

Start with the data and actions the tool will handle, then assess access controls, encryption, authentication, logging, incident response, subprocessors, retention, vulnerability management, and relevant independent reports or attestations. The right level of diligence depends on the use case. A tool handling public brainstorming material does not need the same approval as one processing employee, customer, health, or financial data.

Can an AI vendor use my data to train its models?

The answer depends on the vendor, product tier, account settings, contract, and jurisdiction. Do not infer the policy from a product slogan. Read the current terms, privacy documentation, data-processing terms, and enterprise agreement, and ask what happens to prompts, files, feedback, logs, and outputs. Record the checked date because these terms can change.

What is AI vendor exit risk?

Exit risk is the difficulty and cost of leaving a vendor without losing data, workflow continuity, quality, or customer commitments. It includes proprietary formats, undocumented prompts, embedded integrations, model-specific behavior, long export times, minimum commitments, price changes, and lack of a fallback. Evaluate the exit path before the tool becomes part of a critical process.

Should a small business complete an AI vendor risk assessment?

Yes, but the assessment can be proportionate. A small business can begin with a one-page record of the use case, data, permissions, vendor claims, cost, quality test, owner, contract terms, incident contact, and exit plan. More sensitive or consequential uses deserve deeper review. The goal is to make an informed decision, not to recreate a large enterprise procurement department.

How do I compare two AI vendors fairly?

Give both vendors the same representative tasks, data boundary, success rubric, volume assumptions, and support questions. Compare accepted outcomes, errors, review time, latency, total cost, contract terms, controls, and exit options. Do not compare a polished demo from one vendor with an unconfigured trial from another or treat a benchmark score as proof of fit for your workflow.

Don't stop here

What to read next

Hand-picked guides our readers explore right after this one.