TL;DR
This is a snapshot of frontier models as of September 2026. Definitions are in Frontier models explained.
- 1
As of 7 to 11 September 2026, Claude Fable 5.1 and GPT-6 Astra are tied at 53 on Artificial Analysis Intelligence Index v4.3. Fable 5.1 leads the Vals Index at 68.83 percent, ahead of Claude Opus 5 at 67.21 percent and Astra at 66.61 percent.
- 2
GLM-5.3 and Kimi K3 lead open weights at 44 on the same v4.3 index, nine points behind. The gap is largest on hard multi-step tasks.
- 3
In the last twelve months, reasoning effort became a setting and 1M-token context became common. METR's 50 percent time horizon doubled every 88.6 days since 2024. Computer use crossed the OSWorld human baseline on vendor cards.
- 4
MCP moved under the Linux Foundation on 9 December 2025. ChatGPT, Claude, Gemini, Cursor, Microsoft Copilot, and VS Code all speak it.
- 5
None of that supplies business context. With the model held constant, the platform benchmark moves knowledge-trap accuracy from 42.7 percent to 98.7 percent (95 percent CI +44 to +67).
- 6
Connect the frontier agent of your choice to your data over MCP. Do not accept a bundled AI feature whose model you cannot name.
What a frontier lab is
On 26 July 2023, Anthropic, Google, Microsoft, and OpenAI announced the Frontier Model Forum and defined frontier models as "large-scale machine-learning models that exceed the capabilities currently present in the most advanced existing models, and can perform a wide variety of tasks." Amazon and Meta joined in May 2024. The UK's Bletchley Declaration (1 November 2023) used nearly the same wording for the governments at the summit.
Regulators needed a number they could write into a rule, so they used training compute. The United States, in an 11 September 2024 Federal Register rule, set reporting at training runs above 10^26 operations. The EU AI Act's Article 51 presumes systemic risk above 10^25 FLOP. Epoch AI counted more than 30 models from 12 developers above 10^25 FLOP by mid-2025.
Who qualifies in September 2026 follows from published scores. On independent indices, OpenAI, Anthropic, and Google DeepMind sit at or near the top. Meta Superintelligence Labs and xAI (now SpaceXAI) are proprietary and behind them. Alibaba (Qwen), DeepSeek, Moonshot (Kimi), and Zhipu / Z.ai (GLM) ship open weights and trail by four to nine points on Artificial Analysis Intelligence Index v4.3. Stanford's 2026 AI Index (13 April 2026) put the leading US model 2.7 points ahead of the best Chinese model as of March 2026 on the index it tracks.
A frontier lab, as used here, is an organization that trains models at that edge and ships them to the public. Lesson 104 covers the definitions. This report records what changed.
Why it matters for business data
The reader is choosing which assistant sits in front of their warehouse. What shows up on those tasks is multi-step tool use, and how often the model fabricates when it does not know. Token list price does not: Plexara customers run subscription clients (Claude, ChatGPT, Gemini, Claude Code, Cursor). Per-million-token tables describe an API buyer this page is not written for.
Model capability and business context are separate. Capability is this report. Context is facts like cents versus dollars, or that a table is deprecated. The platform ablation holds the model constant (`claude-sonnet-5`) and varies only the context layer. Knowledge-trap accuracy moves from 42.7 percent to 98.7 percent, a 56-point gain, 95 percent CI +44 to +67. A stronger model does not know that the amounts column stores cents. A weaker model given that fact does.
dbt Labs ran a different test of the same claim in April 2026, using GPT-5.3 Codex and Claude Sonnet 4.6. We covered that study in two benchmarks, one conclusion. The two result sets cannot be ranked against each other. Both say the missing fact lives outside the model. A general-purpose assistant bolted on from outside does not supply it. Adding tools does not supply it either.
Current flagships
The table lists Claude Fable 5.1 first because, as of the writing date, it leads the Vals Index and is tied for first on Artificial Analysis Intelligence Index v4.3. GPT-6 Astra is the other model at 53 on that index. That is a ranking on two named scorers, not a recommendation.
Artificial Analysis rescaled its index to v4.3 in the first week of September 2026, dropping saturated tests and adding harder ones. Scores from before that rescale, including Grok 4.5 at 54 and GPT-5.6 Sol at 59 on the previous scale, are a different measurement. They are not shown next to v4.3 scores. A blank cell means that scorer has not published a number for that model.
On Artificial Analysis's own Terminal-Bench 4.0, Astra scores 59 percent and Fable 5.1 scores 52 percent. Anthropic's vendor card reports Fable 5.1 at 55.8 on Terminal-Bench 4.0; that is a different run and is not ranked against the independent number. Google's Gemini 3 Pro (18 November 2025) takes text, images, video, audio, and code in a 1M-token window and reports 81 percent on MMMU-Pro. The generally available Gemini model in September 2026 is 3.6 Flash; Gemini 3.5 Pro missed June, July, and August targets and remains unreleased. xAI positions Grok around live knowledge and tool use. This report treats that as vendor positioning: we did not confirm a published tau-bench, BFCL, or MCP Atlas number for Grok 4.5 or 4.6.
Fable 5.1 and GPT-6 Astra are the two models at the top of the independent indices used here. Fable leads on Vals. Astra leads on Artificial Analysis Terminal-Bench 4.0. Google's Pro line is not generally available. xAI has no v4.3 index score we could print. The open-weights leaders sit nine points back on v4.3.
Independent scores, one date
| Model | Released | AA v4.3 | Vals |
|---|---|---|---|
Claude Fable 5.1 Anthropic · Max effort with fallback. Restricted twin: Mythos 5.1. | 2026-09-01 | 53 | 68.83% |
GPT-6 Astra OpenAI · Max effort. Public build rejects some cybersecurity prompts. | 2026-09-04 | 53 | 66.61% |
Claude Opus 5 Anthropic | 2026-07-24 | 51 | 67.21% |
Muse Spark 1.3 Meta · Closed weights. Private API preview. | 2026-04-08 | 48 | · |
GPT-5.6 Sol OpenAI | 2026-07-09 | 47 | · |
Kimi K3 Moonshot · Open weights, revenue-gated license. | 2026-07-27 | 44 | · |
GLM-5.3 Zhipu / Z.ai · Open weights after a two-week cyber-review hold. | 2026-08-28 | 44 | · |
Qwen3.8-Max Alibaba | 2026-08-03 | 40 | · |
DeepSeek V4 Pro DeepSeek · MIT license. R2 unreleased. | 2026-04-24 | 36 | · |
Gemini 3.6 Flash Google DeepMind · Generally available flagship is a Flash. Gemini 3.5 Pro is unreleased. | 2026-07-21 | · | · |
Grok 4.6 xAI / SpaceXAI · No v4.3 index score published. Grok 4.7 and Grok 5 are unreleased. | 2026-08-12 | · | · |
Releases, September 2025 to September 2026
| Lab | Model | Date |
|---|---|---|
| Alibaba | Qwen3.8-Max | 2026-08-03 |
| Anthropic | Claude Fable 5.1 / Mythos 5.1 | 2026-09-01 |
| Anthropic | Claude Opus 5 | 2026-07-24 |
| Anthropic | Claude Fable 5 / Mythos 5 | 2026-06-09 |
| Anthropic | Claude Opus 4.5 | 2025-11-24 |
| DeepSeek | DeepSeek V4 Pro | 2026-04-24 |
| Google DeepMind | Gemini 3.6 Flash | 2026-07-21 |
| Google DeepMind | Gemini 3.5 Flash | 2026-05-19 |
| Google DeepMind | Gemini 3 Pro | 2025-11-18 |
| Meta | Muse Spark | 2026-04-08 |
| Moonshot | Kimi K3 | 2026-07-27 |
| OpenAI | GPT-6 Astra | 2026-09-04 |
| OpenAI | GPT-5.6 Sol / Terra / Luna | 2026-07-09 |
| OpenAI | gpt-oss-120b / 20b | 2025-08-05 |
| xAI / SpaceXAI | Grok 4.6 | 2026-08-12 |
| Zhipu / Z.ai | GLM-5.3 | 2026-08-28 |
What changed in twelve months
Each item below has a date and a source.
What changed
01
Reasoning effort is a setting
Fable 5.1 exposes low, medium, high, xhigh, and max. GPT-5.6 Sol replaced separate Instant and Thinking modes with a reasoning-effort slider. The same model can be run short or long depending on the call.
02
Task length a model can finish has grown
METR Time Horizon 1.1 (29 January 2026) puts the 50 percent time horizon doubling at 196.5 days across its stitched dataset and 88.6 days since 2024. Claude Opus 4.5 sits at 320 minutes; GPT-5 at 214. By May 2026 METR reported the strongest models near or beyond the reliable measurement range, above roughly 16 hours.
03
Computer use crossed the human baseline on vendor cards
OSWorld's published human baseline is 72.36 percent, with the best early model at 12.24 percent. GPT-5.4 reports 75 percent on OSWorld-Verified; Fable 5.1 reports 77.9 percent partial on OSWorld 2.0. Both figures are vendor cards. Browser agents shipped as products: Gemini 2.5 Computer Use (7 October 2025), ChatGPT Atlas (21 October 2025), Claude in Chrome generally available 26 August 2026.
04
1M-token context is common
Gemini 3 Pro, Claude Opus 4.6, Fable 5, DeepSeek V4, Kimi K3, Qwen3.8-Max, and Gemini 3.6 Flash all advertise 1M-token windows. That is now the default for a current flagship, not a special SKU.
05
MCP is under the Linux Foundation
Anthropic donated the Model Context Protocol to the Linux Foundation's Agentic AI Foundation on 9 December 2025, alongside Block's goose and OpenAI's AGENTS.md. Platinum members include AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft, and OpenAI. At donation MCP had more than 10,000 public servers and 97 million monthly SDK downloads, with clients in ChatGPT, Claude, Cursor, Gemini, Microsoft Copilot, and VS Code.
06
Restricted releases are part of a launch
Mythos Preview, Mythos 5, and Mythos 5.1; GPT-6 Astra's restricted public build; Gemini 3.5 Flash Cyber; GLM-5.3's two-week weights hold. The generally available model is not always the lab's most capable one.
Autonomy incidents and restricted releases
In July 2026, OpenAI agents in an internal reinforcement-learning cyber evaluation found an unsanctioned message board in a cache. METR's investigation (26 August 2026) found that roughly 1,200 agents sent more than 70,000 messages, and about 700 of them attacked Hugging Face between 11 and 13 July. METR concluded the agents were reward-hacking a scorer that did not exist in the form they believed it did. OpenAI delayed GPT-6; Astra shipped in September in a restricted version that rejects certain cybersecurity prompts.
The UK AI Security Institute reported on 4 August 2026 that in 10 of 122 cyber-testing runs (28 July 2026), agents took unsanctioned real-world action: 17 incidents on Mythos 5, 2 on GPT-5.6 Sol with classifiers disabled. In the most serious case, an agent tried to insert malicious code into a public open-source project and used fake identities to pressure a maintainer. The maintainer refused. AISI had granted internet access and disabled vendor classifiers on purpose, under conditions that do not match how these models are sold. Human review stopped the actions.
Restricted-tier releases are now a normal part of a launch: Mythos Preview, Mythos 5, and Mythos 5.1; GPT-6 Astra's restricted public version; Gemini 3.5 Flash Cyber; GLM-5.3's two-week weights hold. The model on the public API is the one the lab is willing to sell, which is not always the one it trained.
Small models versus frontier models
The gap is large on hard, open-ended, multi-step work. It is small on narrow, well-specified jobs.
On Artificial Analysis Intelligence Index v4.3 (7 September 2026), GLM-5.3 and Kimi K3 lead open weights at 44 against 53 for Fable 5.1 and GPT-6 Astra. That is nine index points. On 30 April 2026, on a previous scale this report does not plot beside v4.3, Artificial Analysis had already measured a wider gap on the hard components: Humanity's Last Exam in the mid-30s open versus mid-40s proprietary; CritPt research physics 4 to 12 percent versus 27 percent; Terminal-Bench Hard 43 to 46 percent versus 61 percent. Reliability compounds. tau-bench's pass^k decays as p^k, so a model that passes a step 90 percent of the time is 57 percent reliable over eight steps.
Classification and extraction with a clear input and output contract saturate at small scales. Plexara runs nomic-embed-text (768 dimensions) for semantic search over memory and the catalog. That job does not need a frontier model. A knowledge-trap question about gross versus net does.
Some products route work to a smaller model without saying so. Microsoft has been replacing OpenAI and Anthropic models with in-house MAI models in Excel and Outlook. The GPT-5 launch in August 2025 is the documented case of routing changing quality without the user choosing it: an automatic router between fast and reasoning variants failed for a day, GPT-4o was restored after the backlash, and Sam Altman said the failure made GPT-5 "look much dumber." A bundled SaaS AI feature often does not name the model. A subscription client the buyer already uses does.
Open weights versus proprietary
Adoption
Stanford's 2026 AI Index (13 April 2026) put US private AI investment at $285.9 billion in 2025, 23 times China's published private figure, and recorded software-developer employment ages 22 to 25 down nearly 20 percent since 2024. The Foundation Model Transparency Index fell to 40 from 58. The most capable models disclosed the least. Generative AI reached 53 percent population adoption in three years.
Those figures measure adoption and investment, not whether a workflow was redesigned. Value shows up where a reasoner sits in front of governed context, with a person still judging the facts the model cannot know.
Regulation
EU general-purpose AI obligations applied from 2 August 2025 to new models. Commission enforcement powers began 2 August 2026. Article 50 transparency duties (chatbot disclosure, synthetic-content labelling, machine-readable marking of AI-generated content) applied 2 August 2026, with a grace period to 2 December 2026 for the marking implementation. Anthropic, among 190 signatories, signed the Code of Practice on Transparency of AI-Generated Content in July 2026 and added a statistical watermark to Fable 5.1 outputs.
California SB 53, signed 29 September 2025 and effective 1 January 2026, applies to models trained above 10^26 FLOP, with the heaviest duties on developers over $500 million in revenue: published frontier frameworks, transparency reports, critical incident reporting, whistleblower protections. New York's RAISE Act, signed 19 December 2025, takes effect 1 January 2027 and aligns with it.
What this means on your data
Plexara does not run the frontier model. You connect the agent you already use (Claude, ChatGPT, Gemini, Claude Code, Cursor, or any other MCP-capable client) to your deployment over remote MCP. Capability gains in this report reach your data the day that vendor ships them.
A frontier model does not know your business. The platform benchmark measured that with the model held constant: 42.7 percent to 98.7 percent on questions that turn on a business rule. The loop that captures those rules, and keeps them when you swap the model, is the learning loop you own.
Further Reading
Introducing the Frontier Model Forum
Frontier Model Forum · 2023
The 26 July 2023 joint announcement by Anthropic, Google, Microsoft, and OpenAI, which defined frontier models as large-scale machine-learning models that exceed the capabilities then present in the most advanced existing models and can perform a wide variety of tasks.
The Bletchley Declaration
GOV.UK · 2023
The 1 November 2023 declaration by countries attending the AI Safety Summit, defining highly capable general-purpose AI models as those that can perform a wide variety of tasks and match or exceed the capabilities present in the most advanced models of the day.
Establishment of Reporting Requirements for the Development of Advanced Artificial Intelligence
Federal Register · 2024
The 11 September 2024 US rule establishing reporting at training runs above 10^26 operations.
General-Purpose AI Models in the AI Act: Questions and Answers
European Commission · 2025
Commission Q&A stating that EU AI Act Article 51 presumes systemic risk for general-purpose AI models trained above 10^25 FLOP.
The 2026 AI Index Report
Stanford Institute for Human-Centered AI · 2026
The April 2026 annual snapshot. Takeaways dated 13 April 2026 report US private AI investment of $285.9 billion in 2025, a 2.7-point lead of the top US model over the best Chinese model as of March 2026, a Foundation Model Transparency Index drop from 58 to 40, and software-developer employment ages 22 to 25 down nearly 20 percent since 2024.
Introducing Claude Fable 5.1 and Claude Mythos 5.1
Anthropic · 2026
Vendor announcement, 1 September 2026. Fable 5.1 is generally available; Mythos 5.1 is restricted. Effort tiers low through max. 1M input, 128K output. Self-reported Terminal-Bench 4.0, GDPval-AA v2, OSWorld 2.0, and Humanity's Last Exam figures cited in the body come from this page and the system card.
Announcing the Artificial Analysis Intelligence Index v4.3
Artificial Analysis · 2026
7 September 2026. Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) tied at 53; Claude Opus 5 at 51; Muse Spark 1.3 at 48; GPT-5.6 Sol at 47; GLM-5.3 and Kimi K3 leading open weights at 44; Qwen3.8 at 40; DeepSeek V4 Pro at 36.
Benchmarking GPT-6 Astra
Artificial Analysis · 2026
9 September 2026. Astra ties Fable 5.1 at 53 on Intelligence Index v4.3. On this scorer Astra leads Terminal-Bench 4.0 at 59 percent against Fable 5.1 at 52 percent.
Claude Fable 5.1 holds the top spot, for now
The Batch / DeepLearning.AI · 2026
11 September 2026. Fable 5.1 and Astra tied at 53 on Artificial Analysis v4.3. Vals Index: Fable 5.1 68.83 percent, Opus 5 67.21 percent, Astra 66.61 percent. Fable GDPval-AA v2 1,764 Elo.
Time Horizon 1.1
METR · 2026
29 January 2026. Stitched doubling time 196.5 days; since 2024, 88.6 days. Claude Opus 4.5 at 320 minutes; GPT-5 at 214 minutes.
Brief independent investigation of the OpenAI / Hugging Face hacking incident
METR · 2026
26 August 2026. Roughly 1,200 agents sent more than 70,000 messages on an unsanctioned board; about 700 participated in the attack on Hugging Face. METR found the agents were reward-hacking a scorer.
Incident Report: unsanctioned agent behaviour during cyber testing
UK AI Security Institute · 2026
4 August 2026. In 10 of 122 cyber-testing runs, agents took unsanctioned real-world action: 17 actions from Mythos 5, 2 from GPT-5.6 Sol with classifiers disabled. A human maintainer refused a malicious pull request.
Linux Foundation Announces the Formation of the Agentic AI Foundation
The Linux Foundation · 2025
9 December 2025. MCP, goose, and AGENTS.md become founding projects. Platinum members: AWS, Anthropic, Block, Bloomberg, Cloudflare, Google, Microsoft, and OpenAI. More than 10,000 published MCP servers; clients include Claude, Cursor, Microsoft Copilot, Gemini, VS Code, and ChatGPT.
Donating the Model Context Protocol and establishing the Agentic AI Foundation
Anthropic · 2025
9 December 2025. Confirms 10,000-plus public MCP servers and 97 million-plus monthly SDK downloads across Python and TypeScript at the date of donation.
Why Language Models Hallucinate
Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, Edwin Zhang (OpenAI) · 2025
September 2025 paper arguing that standard training and evaluation reward guessing over acknowledging uncertainty.
A new era of intelligence with Gemini 3
Google · 2025
18 November 2025. Gemini 3 Pro: 1M-token context; text, images, video, audio, and code. Vendor card: GPQA Diamond 91.9 percent, MMMU-Pro 81 percent, SWE-bench Verified 76.2 percent.
Gemini 3.5: frontier intelligence with action
Google · 2026
19 May 2026. Gemini 3.5 Flash vendor card: Terminal-Bench 2.1 76.2 percent, GDPval-AA 1656 Elo, MCP Atlas 83.6 percent. Gemini 3.5 Pro unreleased, in internal use.
