On 27 July 2026, Microsoft AI announced MAI-Cyber-1-Flash inside MDASH. Microsoft paired a compact cyber model with GPT-5.4 escalation so most tasks use the cheaper specialist and exceptional cases use a larger model.
This release brief was checked against first-party material on 10 August 2026. The date above is the public announcement date, not the date a repository was created or a third-party provider added the model. Where access or weights arrived later, that distinction is recorded below.
Release record
| Field | Verified detail |
|---|---|
| Announcement | 27 July 2026 |
| Availability or weight release | 27 July 2026 through MDASH for verified defenders |
| Release type | restricted cybersecurity model and multi-agent system |
| Access | verified-defender access through MDASH; no open weights or unrestricted API |
| Architecture | a compact code-heavy model derived from MAI-Thinking-1 and routed within a 100-plus-agent harness |
| Maximum stated context | the MDASH task and model limits disclosed to customers |
What changed
The release is a model-routing case study. Microsoft says MAI-Cyber-1-Flash handles up to 90% of MDASH tasks, with the remaining 10% escalated to GPT-5.4. That combined system reaches about 96% on CyberGym while halving cost relative to Microsoft’s previous best MDASH offering. The benchmark therefore belongs to the model, historical security data and harness together—not to the compact checkpoint alone.
The practical comparison is therefore not simply whether MAI-Cyber-1-Flash inside MDASH has the largest headline score. Teams need to ask whether its architecture, access terms, latency, tool behaviour and evaluation setup match the workload they actually intend to run. A model can lead one harness while losing on cost, refusal behaviour, multilingual quality or repeatability in another.
Benchmarks worth retaining
| Evaluation | Reported result | How to read it |
|---|---|---|
| CyberGym | 95.95% for MDASH with MAI-Cyber-1-Flash and GPT-5.4 | Combined system result |
| Task coverage | up to 90% handled by the compact model | Routing share claimed by Microsoft |
| Cost | 50% saving versus the prior best MDASH offering | System-level economic claim |
These are release-time results, not independently reproduced guarantees. The headline 95.95% is for MDASH combining MAI-Cyber-1-Flash and GPT-5.4, not a standalone model score, and access is restricted to a governed defensive environment. Scores should remain attached to the disclosed effort setting, agent harness, tool access, timeout, context-management policy and judge model. Moving a number into a procurement sheet without those conditions creates false comparability.
Architecture and access
MAI-Cyber-1-Flash inside MDASH is described as a compact code-heavy model derived from MAI-Thinking-1 and routed within a 100-plus-agent harness with the MDASH task and model limits disclosed to customers of stated context. Its access position at verification time is verified-defender access through MDASH; no open weights or unrestricted API. That wording matters: open weights, source-available weights, an API, a product preview and a research demonstration give adopters very different rights and different levels of reproducibility.
Before deployment, record the exact model identifier or checkpoint, inference stack, quantisation, reasoning setting, region, price schedule and supplier terms. If the release uses a custom licence, read the licence itself rather than relying on the word “open” in launch copy. If it is API-only, preserve the dated documentation and change-notice route because the served snapshot can change without a downloadable artefact.
What an evaluation should test next
For MAI-Cyber-1-Flash inside MDASH, a credible internal gate should include:
- a frozen set of representative tasks with pass, fail and abstain criteria;
- a matched baseline using the same tools, timeout, prompt budget and reviewer rubric;
- repeated runs to expose variance rather than reporting a single best attempt;
- latency, token use and total task cost alongside task success;
- adversarial, multilingual and long-context cases relevant to the real deployment; and
- rollback evidence showing the previous model can be restored safely.
The wider model change-control guide explains how to keep model, prompt, tool and corpus changes reconstructable. The AI dependency inventory guide covers the release and supplier records needed after deployment.
AIEngine verdict
MAI-Cyber-1-Flash is a major applied-model release. Security buyers should evaluate routing precision, proof quality, tenant isolation and remediation safety rather than comparing the combined score to standalone models.
This is a launch assessment, not a certification. Benchmark leadership is useful evidence of where to test; it is not authorization to place the model in a high-impact workflow without domain evaluation, security review and an accountable owner.
Primary sources
Image provenance
Hero image: Microsoft AI official release artwork. The locally served WebP is a crop of the first-party release or model-card asset recorded in the repository provenance manifest.



