Capability progression
Testing is useful when new systems may materially exceed prior state-of-the-art capabilities or cross risk-relevant capability levels.
The UK's AI Security Institute (AISI) conducts real pre-deployment evaluations of advanced AI systems and receives privileged access to leading models through close collaboration with AI companies. But AISI's current published material does not describe a blanket statutory process in which every frontier model must complete a fixed government approval, notification, or pass/fail authorization before UK deployment.
Zeph Tech's earlier February 2026 briefing overstated AISI's role as a mandatory deployment regulator. This maintained guide replaces that briefing with the current evidence from AISI's own publications.
AISI's mission is to give governments empirical understanding of advanced AI capabilities and risks. Its work includes building and running evaluations, conducting foundational research, developing evaluation infrastructure, collaborating with model developers, and sharing technical evidence with governments and other partners.
AISI's own materials state that leading AI companies agreed to collaborate with governments on testing the next generation of advanced models. The institute has pre-deployment access to leading systems and has published examples of testing models before public release. That is materially different from a legal requirement that every model must obtain AISI authorization.
The distinction matters for governance. A buyer, developer, policymaker, or auditor should not describe an AISI test as a statutory approval, market authorization, or deployment license unless a separate current legal instrument explicitly creates that effect.
AISI has said that it does not have capacity to evaluate every released model. Its published approach focuses on the most advanced systems and selects models using estimates of the risk that a system may possess harmful capabilities. Inputs can include factors such as training compute and expected accessibility, while AISI has also said it intends to improve its risk-estimation methods beyond simple proxies.
This means a single compute threshold should not be presented as the current legal boundary for AISI evaluation. Compute can be one signal in prioritization, but the institute's published evaluation approach is broader and risk-focused.
Testing is useful when new systems may materially exceed prior state-of-the-art capabilities or cross risk-relevant capability levels.
How a model is exposed, what tools it can use, and who can access it can affect real-world risk and therefore evaluation priority.
AISI has discussed evaluation across the AI lifecycle, including significant post-deployment changes such as tool integrations, fine-tuning access, longer context windows, and other changes that alter capability.
New scaffolds, elicitation methods, jailbreaks, or agent frameworks can change the effective capabilities of an already-deployed model and create reasons for renewed testing.
Structured tasks and benchmarks can provide repeatable signals about capabilities in specific domains and help identify areas that need deeper investigation.
Models can be placed in tool-enabled environments and asked to complete multi-step objectives so evaluators can observe capabilities that emerge through planning and interaction.
Domain experts interact with systems to test safeguards, uncover failure modes, and explore capabilities that static benchmarks may not reveal.
Evaluations can ask whether access to an AI system meaningfully increases a person's ability to perform a harmful or otherwise risk-relevant task compared with existing non-AI tools.
Published AISI work has compared pre-release models with reference models to understand relative capability changes rather than treating a raw score as a universal pass/fail verdict.
AISI has described designing evaluation methods and expected reporting structures before model access so time pressure around imminent releases does not eliminate methodological review and quality assurance.
A published example is the joint UK/U.S. evaluation of OpenAI's o1 model. The institutes received limited pre-deployment access, tested the model across cyber capabilities, biological capabilities, and software/AI development, compared it with reference systems, and shared findings with OpenAI before public release.
AISI has also published work based on pre-release snapshots supplied through ongoing collaborations with other developers. Its later research demonstrates that the institute's evaluation program extends beyond a single release checkpoint into safeguards, alignment, misuse, autonomy, and other risk-relevant research.
These examples support a factual statement that AISI performs pre-deployment testing. They do not support claims that AISI universally authorizes deployment, imposes a fixed notification period, or operates a statutory pass/fail licensing scheme for every frontier model.
An evaluation measures the tasks, scaffolds, prompts, tools, safeguards, model snapshot, and threat assumptions actually tested. Results should not be generalized beyond that scope without evidence.
Post-training changes, tool access, deployment configuration, fine-tuning, safeguards, model updates, and third-party scaffolds can materially change the system users interact with.
A model can pass a test suite because the capability is absent, because the evaluation failed to elicit it, or because the tested conditions differ from real deployment. Interpret negative results with methodological limits visible.
Technical evaluation can inform government decisions. Whether a developer has a legal duty, deployment restriction, disclosure requirement, or penalty exposure must come from the relevant law or current regulatory instrument—not from the existence of an AISI evaluation alone.
Maintain representative model snapshots, evaluation interfaces, safeguard controls, logging, environment isolation, reproducible configurations, and contacts who can explain training/deployment differences.
Record exactly which snapshot and system configuration an external evaluator tested. Do not market findings from one model version as proof about materially different later systems.
Ask whether a model or related system was externally evaluated, by whom, under what access conditions, on which version, against which risks, and whether the deployment you will use differs from the evaluated setup.
Track technical evaluation evidence, internal deployment approval, and external legal/regulatory obligations as distinct records. A positive result in one category does not automatically satisfy the others.
Trigger renewed testing when model capability, safeguards, tools, context length, fine-tuning, deployment scope, or attack techniques change enough to invalidate prior evidence.
AISI's role and UK AI policy can evolve. Monitor AISI and current UK government sources before stating that collaboration practices have become legal requirements or that a new statutory regime has taken effect.
The AI Model Evaluation Operations guide provides an organization-level pattern for representative datasets, acceptance thresholds, reproducibility, human review, change gates, and evidence retention.
AISI changed its name from the AI Safety Institute to the AI Security Institute on 14 February 2025. Older publications may use the former name.
Follow the next implementation topic without returning to search.
Evaluate model quality, risk, and evidence before deployment
Continue readingEstablish accountable AI governance and evidence practices
Continue readingEvaluate AI vendors with governance and evidence controls
Continue readingUse the source-backed research to pressure-test assumptions, then build a reusable evaluation brief before you compare products, scope implementation, or request a fit review.