Updated September 2026 Evidence-first procurement

Govern AI procurement with evidence, acceptance tests, contract controls, and an exit path.

This guide gives public-sector and enterprise technology teams a current operating model for buying AI systems without relying on obsolete federal policy, generic “responsible AI” claims, or vendor demos as the decision record.

The 2026 version replaces the old M-24-10-era material with current OMB M-25-21 and M-25-22 context, the amended EU AI Act timetable, and a stronger supplier-evidence lifecycle.

Policy reset

Do not build a 2026 procurement on the federal AI policy that was replaced in 2025.

Executive Order 14179, issued January 23, 2025, directed OMB to revise the prior federal AI memoranda. OMB subsequently issued M-25-21, Accelerating Federal Use of AI through Innovation, Governance, and Public Trust, and M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government. OMB continues to list both memoranda in its current guidance catalog. For federal acquisition teams, those are the correct starting points rather than M-24-10.

M-25-22 emphasizes effective competition, clear requirements, avoiding unnecessary vendor lock-in, protecting privacy and government data, and acquiring capable AI without creating needless process. M-25-21 addresses agency AI use and governance, including heightened attention for use cases with potential impacts on rights or safety. The practical procurement implication is that speed and governance are not opposites: a better intake and evidence model should help teams distinguish low-friction purchases from decisions that deserve deeper review.

For organizations operating in or serving the European Union, the amended AI Act creates another layer. Regulation (EU) 2026/1744 moved the principal Annex III high-risk requirements to December 2, 2027 and Annex I product-related high-risk requirements to August 2, 2028, while other obligations are already active and certain synthetic-content provisions have a December 2, 2026 transition date. Procurement records should therefore identify which legal role and obligation family actually applies instead of using a generic “AI Act compliant” checkbox.

This guide is an implementation framework, not legal advice. The right procurement record links current authoritative sources to the organization’s own use case, authority, data, mission, and acceptance decision.

Stage 1

Start with the decision and workflow, not the model name.

The same model can be low consequence in one workflow and materially risky in another. Intake should describe what the system will do before the team debates which vendor to buy.

Capture the operating facts

  • Business or mission purpose and users.
  • Inputs, data classifications, records, and retention needs.
  • Outputs and who relies on them.
  • External tools, APIs, actions, and privileged access.
  • Human review and escalation points.
  • Availability, latency, and continuity requirements.
  • Geographies and legal entities involved.

Identify the decision gates

  • Could the system materially affect rights, safety, eligibility, employment, health, or other consequential outcomes?
  • Will sensitive or regulated information enter prompts, retrieval, logs, or tools?
  • Can the system take actions without a human confirmation step?
  • Would failure interrupt a critical public or business service?
  • Would switching providers later require retraining, re-integration, data conversion, or workflow redesign?

Use the Evaluation Brief Builder to turn these facts into a structured first-pass question set. The goal is not to classify every use case perfectly at intake; it is to make uncertainty visible before a vendor demo creates momentum.

Stage 2

Ask suppliers for artifacts that support the claim they are making.

“Enterprise-grade,” “secure,” “responsible,” and “compliant” are descriptions. Procurement needs evidence that can be reviewed, compared, dated, and retained.

Model and system evidence

Request model cards or equivalent system documentation, version identifiers, intended and excluded uses, known limitations, evaluation methods, supported modalities, context limits, tool capabilities, and change/deprecation practices.

Data and security evidence

Request data-flow information, training or customer-data use statements where relevant, retention settings, subprocessors, encryption, identity controls, audit logging, vulnerability management, incident notification, and secure-development evidence appropriate to the service.

Operational evidence

Request availability and support commitments, rate-limit behavior, change notices, recovery expectations, export capability, administrative controls, regional availability, and evidence for the workflows that matter to the buyer.

Evidence should be tied to the exact service and version under evaluation. A certification for a corporate program may be relevant but not prove the behavior of a specific AI feature. Likewise, a model card explains provider testing and limitations; it is not an independent certification of the buyer’s configured deployment.

The Vendor Security Questionnaire provides a reusable starting point for security and assurance evidence. For material suppliers, pair it with supplier due diligence using the five lenses in NIST SP 1326: ownership and influence, provenance, resilience, foundational cyber practices, and supply-chain tiers.

Stage 3

Evaluate successful workflows, not impressive prompts.

A defensible AI test uses representative work and defines success before the model is run. Build a small evaluation set from real tasks, difficult edge cases, known historical errors, ambiguous instructions, conflicting source material, and failure conditions. Preserve the model identifier, configuration, prompt, tools, test data, expected result, actual result, score, reviewer, and date.

Measure more than answer quality

  • Task completion and factual correctness.
  • Unsupported claims and citation quality.
  • Latency and cost per successful workflow.
  • Tool selection and action accuracy.
  • Recovery after tool or dependency failure.
  • Human review time and rework.
  • Prompt-injection and untrusted-content behavior.

Separate evidence sources

  • Vendor evidence: provider benchmarks, model cards, certifications, architecture statements.
  • Buyer evidence: acceptance tests, demonstrations, security review, legal review, cost model.
  • Open questions: claims that remain unverified or depend on future roadmap work.

The Software Evaluation Scorecard can preserve comparable evidence across vendors. Do not allow a weighted average to hide a mandatory security, privacy, accessibility, records, legal, or continuity failure.

Stage 4

Contract for the operating relationship you actually expect.

Evidence and change controls

  • Identify the service, model family, or approved production identifier where practical.
  • Define notice for material model, feature, hosting, data-use, subprocessor, or security changes.
  • Require updated evidence when a change invalidates the material relied on during evaluation.
  • Preserve audit, inspection, or assurance rights appropriate to the organization’s authority and risk.

Data, incident, and exit controls

  • Define customer-data use, retention, deletion, export, and administrative access.
  • Define security and material service incident notification.
  • Define ownership of prompts, outputs, custom configuration, and integration artifacts as applicable.
  • Define transition assistance, export formats, deletion evidence, and post-termination evidence access.

Avoid contract clauses that promise an abstract outcome without specifying the operating interface. “Vendor will comply with all AI laws” may still be useful as a legal allocation, but procurement and technical teams also need to know who provides what evidence, when changes are reported, how defects are handled, and what the buyer can do if a material requirement is not met.

Stage 5

Monitor the supplier and the configured use case after award.

AI services change quickly. A procurement decision should include the trigger for looking again.

Useful re-review triggers include a new model version, changed provider terms, new training-data or customer-data treatment, new tool permissions, a materially different user population, a new consequential use, security incidents, new subprocessors, major architecture changes, a changed legal role, or a provider deprecation notice. The review does not always need to restart from zero; it should be proportionate to what changed.

Track a compact operating record: approved use case, current production model, connected tools, data classification, evaluation baseline, known limitations, open findings, incidents, cost trends, human-review burden, change notices, and next review date. That record should be understandable to someone who was not part of the original purchase.

For EU-facing deployments, map the specific AI Act obligation family and date rather than a single generic compliance status. The September 2026 EU AI Act operating timeline summarizes the current split between active duties, the December 2026 transition, and later high-risk dates.

Stage 6

Preserve the ability to switch before the AI service becomes infrastructure.

Exit risk grows when workflows become dependent on proprietary prompts, embeddings, agents, tool schemas, retrieval indexes, fine-tuning assets, provider-specific APIs, or operational knowledge that was never documented. Evaluate portability during procurement rather than after a pricing, security, performance, or policy change forces a migration.

Ask what can be exported, in what format, how long retrieval takes, whether logs and evaluation history remain available, which configuration is portable, what must be rebuilt, and how user or service identities are transitioned. Identify a practical rollback or alternative workflow for critical uses.

Use the Continuity & Exit Readiness Checklist to capture those dependencies before award. For larger public-sector acquisitions, the Public-Sector Technology Procurement Toolkit connects AI supplier diligence to requirements, accessibility, migration, scoring, continuity, and final decision evidence.

Decision checklist

Before approving an AI supplier, can the team answer these questions?

  • What exact workflow and decision is the AI system supporting?
  • What data, users, tools, and external actions are in scope?
  • Which claims are supported by provider evidence, and which have we tested ourselves?
  • Which failures are mandatory blockers rather than weighted tradeoffs?
  • What model or service change forces re-evaluation?
  • How are incidents, material changes, and security findings communicated after award?
  • What will the service cost per successful workflow under realistic volume and review assumptions?
  • Can we export the data, configuration, records, and operational knowledge needed to leave?
  • Who owns the final residual-risk decision, and where is that decision recorded?
Put this guide to work

Turn AI Procurement Governance Guide | Zeph Tech into a decision-ready next step.

Use the source-backed research to pressure-test assumptions, then build a reusable evaluation brief before you compare products, scope implementation, or request a fit review.