Leading Frontier AI Models: Capability Ceiling (September 2026)

What sits at the top of the field right now, hosted or downloadable, with every licence, price and index score re-read from a primary source today.

Updated September 27, 2026

This page answers one question: what sits at the capability ceiling right now, regardless of licence, price, or whether the weights can be downloaded at all. It is the opening entry in a monthly series, and the ranking will be rebuilt from primary sources each month rather than edited in place.

Two neighbouring pages answer different questions, and for most readers of this site they are the better starting point. The open-weight ranking covers only models whose weights can be downloaded, and treats the licence as a first-class constraint on what you may then do with them. The VRAM tier index maps card memory to a model and quantization that will actually load. This page sets both constraints aside on purpose. Most of what follows will not run on consumer hardware, and one entry cannot be bought at any price without an invitation.

How this ranking was built

Every hosted model below was read from its vendor’s own model or pricing page on 27 September 2026. Every open-weight entry was read from the Hugging Face model API the same day: cardData.license and license_name for the licence, gated for the access flag, and the byte-exact safetensors.total for the parameter count. No figure here is quoted from recollection.

The ranking spine is the Artificial Analysis Intelligence Index, a third-party composite rather than a vendor’s own scorecard. The version is load-bearing. The methodology page states v4.1.1, built from nine evaluations weighted across agents (34 percent), coding (24 percent), scientific reasoning (24 percent) and general tasks (18 percent), including GDPval-AA v2, Terminal-Bench v2.1, SciCode, GPQA Diamond and Humanity’s Last Exam. The leaderboard table itself does not print a version string, so that version number comes from the methodology page and not from the table. Scores produced under different index versions sit on different scales and cannot be compared with each other.

The leaderboard lists a separate row for each reasoning-effort setting, so one model appears several times at different scores. The table below takes each model’s highest score among the rows read today.

The field, August 2026

Index (v4.1.1)ModelMade byWeightsLicence or access modelBest at
63Claude Opus 5AnthropicHosted (legacy)Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry; $5 / $25 per MTok; superseded by Opus 5.5 at $4 / $20Complex agentic coding and enterprise work
62Claude Fable 5AnthropicHosted (legacy)Same surfaces; $10 / $50 per MTok; superseded by Fable 5.1 at the same priceLong-running agents at the highest generally available capability
61GPT-5.6 SolOpenAIHostedOpenAI API; $4 / $20 per MTok short context, $8 / $30 long contextComplex professional work across a 1.05M-token context
60Kimi K3Moonshot AIOpenCustom kimi-k3 licence, ungatedThe only downloadable model inside the index top ten
58Qwen3.8 MaxAlibabaHostedAlibaba Cloud Model Studio qwen3.8-max (text and image/video understanding)Access route now verifiable from Alibaba’s own Model Studio list
57Muse Spark 1.2MetaNOT VERIFIEDNOT VERIFIEDNOT VERIFIED
57GPT-5.6 TerraOpenAIHostedOpenAI API; $2 / $12 per MTok short context, $4 / $18 long contextBalancing intelligence against cost
56Grok 4.5xAIHostedxAI API; $2 / $6 per MTok below 200k tokens, $4 / $12 at or aboveA 500k context at flagship quality
55Claude Sonnet 5AnthropicHosted (current)$2 / $10 per MTok; the introductory $2 / $10 price that ran through 31 August 2026 was made the permanent standard, the scheduled increase to $3 / $15 was cancelledSpeed with near-top intelligence
Not scoredClaude Mythos 5AnthropicHostedInvitation only, Project Glasswing; $10 / $50 per MTok; a Mythos 5.1 at the same price has since appeared alongsideDefensive cybersecurity work, per Anthropic’s own framing
Not scoredDeepSeek-V4-ProDeepSeekOpenMIT, ungated1.60T parameters under a plain permissive licence
Not scoredInklingThinking MachinesOpenApache-2.0, ungatedThe largest Apache-2.0 model found today; text, image and audio input
Not scoredGLM-5.2Z.aiOpenMIT, ungatedPermissive frontier weights at 753B
Not scoredK-EXAONE 2.0 750B A37BLG AI ResearchOpenApache-2.0, ungatedMultilingual work and long-context retrieval
Not scoredGemini 3.5 FlashGoogleHostedGemini API; $1.50 / $9.00 per MTokSustained agentic and coding performance, per Google’s own wording

The rows marked “Not scored” are grouped rather than ranked. No index score for any of them appeared in the top fifteen read today, and inventing a position for them would be exactly the error this page exists to avoid. Their placement reflects nothing except that they were not scored in the table consulted.

The ceiling is rented, not owned

Fourteen of the fifteen index positions read today belong to models whose weights could not be found on Hugging Face today. The single exception is Kimi K3, which reports 2,779,931,837,184 parameters in safetensors.total and is ungated on Hugging Face under a custom licence named kimi-k3 in its card data. Anyone may download it. Almost nobody can serve it: at 2.78T total parameters it is a multi-node deployment, not a workstation one.

The top of the market is a rental market, and its one downloadable entry is downloadable in the way a container ship is purchasable.

The licence still buys something real even when the hardware does not. Weights under MIT or Apache-2.0 can be inspected, fine-tuned by anyone with a cluster, and served by a provider other than the lab that trained them. DeepSeek-V4-Pro returns mit with gated: false; Inkling returns apache-2.0 with a licence link pointing at apache.org. Those are portability guarantees, not self-hosting guarantees, and the two are routinely confused.

One entry sits outside both categories. Anthropic’s documentation states that Claude Mythos 5 shares Claude Fable 5’s specifications and pricing but is offered in limited availability to approved customers in Project Glasswing, invitation only, with no self-serve sign-up. A model whose capability cannot be independently measured because it cannot be independently purchased is a real feature of the current field, and no public leaderboard covers it.

What this means if you came here for local AI

Nothing on the scored half of this table will run on a graphics card you own. That is not a reason to ignore it, and it is not a reason to buy hardware.

The gap between the ceiling and what fits in 24GB is one you close with task selection, not with money. A model scoring in the fifties on this composite is being measured on agentic coding, scientific reasoning and long-context retrieval. Summarisation, extraction, classification, rewriting and retrieval-augmented answering are not what separates these systems, and those are the jobs most self-hosted deployments actually run.

The recommendation is therefore a split one. Use the VRAM tier index to pick something that loads on your card and handles routine work locally, where the data never leaves the machine. Keep an API key for the few tasks where the ceiling genuinely matters, and read the open-weight ranking before putting any downloadable model into production, because its licence column is where the expensive surprises live.

Price is a useful check on how much of the ceiling you need. Claude Sonnet 5 now sits at $2 / $10 per MTok - the introductory price that ran through 31 August 2026 was made the permanent standard price, and the scheduled increase to $3 / $15 on 1 September 2026 was cancelled by Anthropic in a footnote on the pricing page. Claude Fable 5 still reads at $10 / $50 per MTok. The index separates Sonnet 5 from Fable 5 by seven points, so the cost ratio for that capability step is now wider than it was in August. GPT-5.6 Luna is listed at $0.20 / $1.20 (short context). The curve from adequate to best is steep at the top and flat for a long way below it.

What was searched for and not found

Absences here are stated with the exact query behind them.

Qwen3.8 Max weights. The Artificial Analysis table lists Qwen3.8 Max at 58. The Hugging Face API still returns HTTP 401 with {"error":"Invalid username or password."} for Qwen/Qwen3.8-Max on 27 September 2026, and that response is the same as a control path known not to exist, so the flagship Max weights remain gated or unpublished. Alibaba’s own Model Studio list, re-pulled on 27 September 2026, now names qwen3.8-max as the most capable text-generation model, with qwen3.8-flash and qwen3.7-plus listed alongside it - the August pass’s claim that “Alibaba’s own Model Studio list … contains no 3.8 entry at all” no longer holds. The Qwen account on Hugging Face now also hosts Qwen/Qwen3.8-27B (safetensors.total 27,781,427,952, Apache-2.0, ungated, last modified 2026-08-14), so the Qwen3.8 generation has shipped open weights for the 27B variant even though the Max remains hosted-only. The largest current public Qwen weights remain Qwen3.5-397B-A17B.

Muse Spark 1.2. This model appears in the leaderboard at 57, attributed to Meta, and could not be corroborated against any Meta source. A Hugging Face search for Muse-Spark returns nothing; the meta-llama account has published nothing newer than 28 April 2025; the facebook account’s recent uploads are vision and reconstruction models. Meta’s pages at ai.meta.com and llama.com did not return a readable model listing to an automated request. The row is marked NOT VERIFIED rather than described.

Frontier weights from Anthropic and OpenAI. No Anthropic model account on Hugging Face returns repositories. The openai account’s most recent language models remain gpt-oss-120b and gpt-oss-20b from 4 August 2025. meta-llama/Llama-5 returns the same not-found response as the control. The xai-org account tops out at grok-2 from 22 August 2025, so Grok 4.5 is hosted only.

A Gemini entry in the top fifteen. None appeared in the rows read today. Google’s documentation describes gemini-3.5-flash as its “Most intelligent model for sustained frontier performance on agentic and coding tasks”, while listing gemini-3.1-pro-preview as a preview. Absence from one third-party table on one day is weak evidence and is recorded as such.

A parameter-count discrepancy. Inkling’s model card states “975B total, 41B active”, while the API reports 952,377,623,626 in safetensors.total. Both figures are given rather than picking one.

All figures read from vendor documentation and the Hugging Face model API on 27 September 2026. Benchmark positions come from a third-party index whose version is recorded above; vendor capability claims quoted here are self-reported. Prices and model lineups change weekly.