LLM Benchmark Tracker

Snapshot date: September 27, 2026
Top Intelligence Index
53
Claude Fable 5.1 & GPT-6 Astra (tied, max effort) — Artificial Analysis
Top SWE-bench Verified
96%
Claude Opus 5 — Anthropic
Top GPQA Diamond
96.0%
GPT-6 Astra — OpenAI
Cheapest Capable Model
$0.22 / $0.66
DeepSeek-V4.1-Flash — 90.9% GPQA, $/1M in-out
Axis scale: 0–65 · Artificial Analysis Intelligence Index (raw points, best reported effort tier per model)

Computed from this dashboard's own data (not a third-party ranking): AA Intelligence Index divided by a blended $/1M-token price (weighted 3 input : 1 output, the usual approximation for a typical request). Only models with both an index score and confirmed pricing can be ranked — 9 of 19 this week.

Model Lab Weights AA Index Arena Elo SWE-bench Verified GPQA Speed (tok/s) Context $/1M in $/1M out Cost/Task Enterprise Spend Share Released
vs.

No public benchmark scores individual model versions (e.g. GPT-6 Astra vs. GPT-5.6 Sol) on ESG, environmental, or ethical-AI grounds the way the Model Comparison tab scores intelligence or coding — so this section is per lab, not per model, and the same rating applies to every model that lab ships. It combines three independent, differently-scoped, real sources rather than one invented "ESG score": Stanford HAI's Foundation Model Transparency Index (disclosure & accountability practices), the SINK Project (environmental/sustainability performance from public data only, a newer rater we could not independently verify beyond its own methodology page — weight it below the academic FMTI), and the EU AI Act's Article 51 systemic-risk presumption: models trained above 1025 FLOP are presumed "systemic risk" under the Act. As of Epoch AI's April 2026 compute report, OpenAI, Google, Anthropic and Meta are named as operating at that scale (obligations phase in through August 2026); we could not obtain a complete model-by-model registry, so this is a lab-level presumption, not a confirmed per-model designation. A dash means the lab simply hasn't been publicly assessed by that index — not that it scored zero.

Lab FMTI Transparency /100 (Dec 2025) SINK Environmental /100 EU AI Act Art. 51 Notes
* Exact numeric score not extractable from the available source excerpt; directional finding only — see the FMTI source for the primary figure.

Sourced from the Ramp AI Index, built from real corporate-card and token-spend transactions on Ramp's platform (not surveys or self-reported usage) — the closest thing to an independent "corporate penetration" figure for AI. It publishes two different metrics; don't conflate them. Business adoption (below) is lab-level: the share of U.S. businesses on Ramp with an active paid subscription to that lab. Enterprise spend share, in the Model Comparison tab, is model-level: each model's share of the dollars Ramp tracks flowing to AI token/subscription spend, published only for the models large enough to rank in Ramp's own top-10 that month — a blank cell there means "not in Ramp's tracked top 10," not zero. Ramp's customer base skews U.S., mid-market and tech-heavy, so treat this as a real but non-representative sample, not global market share.

Anthropic
43.8%
of U.S. Ramp businesses with a paid subscription (Sep 2026, +0.34pp MoM)
OpenAI
39.8%
of U.S. Ramp businesses with a paid subscription (Sep 2026, +0.09pp MoM)
Frontier models' token share
45%
Opus/Fable/Sol-class models' share of tracked token volume — down from a 53% August peak as cheaper standard models pick up share
Open-weight adoption
3.6%
of all Ramp businesses use open-weight models via routing platforms (likely understated per Ramp's own methodology note)
Other labs in this table (Google DeepMind, xAI, Alibaba, Z.AI, Moonshot AI, DeepSeek, StepFun, Meta, Shanghai AI Laboratory, InclusionAI) aren't broken out by name in Ramp's monthly public update, so no adoption figure is shown for them here — that's a reporting gap in the source, not a measured zero.

A different lens on "adoption": OpenRouter measures actual API tokens routed across its platform (prompt + completion), skewing toward individual developers and smaller teams rather than Ramp's corporate-card SMB sample. Ranking is by raw token volume, which reflects usage, not quality — a verbose model processes more tokens for the same task. Only the models from this table that appear in OpenRouter's current top ranks are shown; the rest of OpenRouter's top 5 that week (Hy4 Preview, Space Bunny Alpha, GPT-5.6 Luna) aren't tracked elsewhere in this dashboard so aren't listed below.

  • Artificial Analysis Intelligence Index — two trackers disagree, materially. Fetched directly from artificialanalysis.ai, the top of the index is a near-tie between Claude Fable 5.1 and GPT-6 Astra (53 points each, max-effort configs), with GPT-5.6 Sol well back in 13th (47). BenchLM's mirror of the same benchmark (benchlm.ai/benchmarks/artificialanalysis) shows a completely different order and a different scale: GPT-5.6 Sol first at 58.9%, GPT-6 Astra fifth at 52.8%, and scores expressed as percentages rather than raw points. This isn't a rounding difference — it's a different ranking of the same models. We did not average the two. The chart and "Top Intelligence Index" tile use the direct artificialanalysis.ai figures; treat the BenchLM percentages as a separate, disagreeing read until one tracker corrects.
  • SWE-bench Verified coverage lags the newest releases. The BenchLM SWE-bench Verified leaderboard tops out at Claude Opus 5 (96%), but several models released in the past three weeks — Claude Fable 5.1, GPT-6 Astra, Muse Spark 1.3, Gemini 3.8 Flash, GLM-5.3, Kimi K3, DeepSeek-V4.1-Flash — simply aren't on that leaderboard yet. Their SWE-bench Verified cells are blank (em dash), not zero; the tracker hasn't run them, not that they scored poorly.
  • SWE-bench Pro is not comparable to SWE-bench Verified, and its leaderboard looks stale. Morph's SWE-bench Pro leaderboard shows a substantially harder task set — even the top model there (Claude Opus 4.5, 45.9%) scores far below Verified leaders — and, as of this snapshot, lists no model released after roughly Q2 2026. We're not plotting SWE-bench Pro numbers alongside Verified numbers because the two benchmarks aren't on the same scale and mixing them would misstate difficulty.
  • Pricing is blended vs. split-rate depending on source. Artificial Analysis's own leaderboard reports "cost per task," not per-token input/output pricing. The $/1M in and $/1M out figures come from llm-stats.com's per-model pages instead; where we could not confirm a split rate (e.g. GPT-5.6 Sol, Step 5 Preview), the cells are left blank rather than guessed. The "blended price" used for the Best Value ranking and Efficient Frontier chart on the Overview tab is our own 3:1 input:output approximation of those same split rates — a modeling assumption, not a reported figure.
  • Effort-tier variants aren't separate models. Artificial Analysis lists the same underlying model multiple times at different reasoning-effort settings (low/medium/high/xhigh/max), each with its own index score and cost. The chart and table collapse each model to its single best-reported effort tier to avoid double-counting; if you need the low-cost/low-effort variant's numbers specifically, go to the source.
  • Four columns added Sep 21–27, uneven coverage. Weights, Arena Elo, Speed (tok/s), and Cost/Task were added over recent refreshes. Weights (open vs. proprietary) is confirmed for nearly every row; a few Chinese-lab entries are marked "not individually confirmed" because we're inferring from that lab's release pattern rather than a direct source, and Step 5 Preview is proprietary today with StepFun's own announcement promising open weights on Oct 15, 2026. Arena Elo comes from arena.ai's live text leaderboard and only covers the subset of models it currently ranks. Speed and Cost/Task are populated for only 5 of 19 models; repeated fetch failures on the source site prevented filling in the rest.
  • Agentic / tool-use benchmarks were evaluated and deliberately left out. We checked BFCL v4 and τ³-bench, two real public agentic-tool-use leaderboards, as a candidate KPI. Neither is usable here: both leaderboards are dominated by small/niche models and are missing nearly every frontier model in this table, and τ³-bench's own page states "most rows are provider self-reports rather than independent verification." Forcing that data in would misrepresent frontier-model agentic ability rather than measure it, so we skipped it rather than publish numbers we don't trust.
  • Two different Ramp metrics look similar but aren't. "Business adoption" (% of Ramp businesses paying a lab) and "spend share" (% of tracked AI dollars going to a lab or model) are separate measurements with separate figures — some secondary write-ups report the two under similar-sounding headlines with different numbers for the same lab in the same month. We used ramp.com's own September 2026 index page for the lab-level figures and a source citing Ramp's per-model spend breakdown for Enterprise Spend Share; we did not average or reconcile them. Only 5 of 19 models have a spend-share figure, because Ramp itself only publishes a top-10 ranking each month.
  • OpenRouter usage measures volume, not quality — and it's a third, different metric again. It's neither "who's paying" (Ramp) nor "who scores best" (AA Index) but "how many tokens got routed," which rewards verbose models and heavy automated/batch usage alongside genuine popularity. We list it separately and don't merge it into the Enterprise Spend Share column.
  • The EU AI Act systemic-risk flag is a lab-level presumption, not a confirmed model registry. We could not find a public, complete list of the specific model versions formally designated under Article 51 — only a secondary source naming the labs (OpenAI, Google, Anthropic, Meta) reported by Epoch AI to have crossed the 1025 FLOP compute threshold as of April 2026. Treat the "Presumed in-scope" tag in Governance as directional, not a legal determination.
  • The Best Value ranking and Efficient Frontier chart are our own calculations, not a cited third-party metric. Both divide the AA Index by a 3:1-weighted blended price we compute from the $/1M columns. They only include the 9 of 19 models with both an index score and confirmed split pricing — narrower coverage than the main table.
  • Tracking history is one snapshot deep. Week-over-week deltas started this refresh cycle; a real trend line needs several more weekly snapshots to be meaningful. See the Tracking History card on the Overview tab.
  • There is no per-model ESG, green, or ethics score — anywhere. We looked. No tracker rates individual model versions on environmental or ethical grounds the way Artificial Analysis rates intelligence. The Governance tab substitutes real, cited, company-level ratings rather than an invented per-model number. Treat SINK's scores as lower-confidence than FMTI's; several labs aren't assessed by either index, and a handful of FMTI's exact company scores couldn't be extracted from the source excerpt — those are shown as directional findings (marked *) rather than numbers we didn't actually see.