Computed from this dashboard's own data (not a third-party ranking): AA Intelligence Index divided by a blended $/1M-token price (weighted 3 input : 1 output, the usual approximation for a typical request). Only models with both an index score and confirmed pricing can be ranked — 9 of 19 this week.
| Model | Lab | Weights | AA Index | Arena Elo | SWE-bench Verified | GPQA | Speed (tok/s) | Context | $/1M in | $/1M out | Cost/Task | Enterprise Spend Share | Released |
|---|
No public benchmark scores individual model versions (e.g. GPT-6 Astra vs. GPT-5.6 Sol) on ESG, environmental, or ethical-AI grounds the way the Model Comparison tab scores intelligence or coding — so this section is per lab, not per model, and the same rating applies to every model that lab ships. It combines three independent, differently-scoped, real sources rather than one invented "ESG score": Stanford HAI's Foundation Model Transparency Index (disclosure & accountability practices), the SINK Project (environmental/sustainability performance from public data only, a newer rater we could not independently verify beyond its own methodology page — weight it below the academic FMTI), and the EU AI Act's Article 51 systemic-risk presumption: models trained above 1025 FLOP are presumed "systemic risk" under the Act. As of Epoch AI's April 2026 compute report, OpenAI, Google, Anthropic and Meta are named as operating at that scale (obligations phase in through August 2026); we could not obtain a complete model-by-model registry, so this is a lab-level presumption, not a confirmed per-model designation. A dash means the lab simply hasn't been publicly assessed by that index — not that it scored zero.
| Lab | FMTI Transparency /100 (Dec 2025) | SINK Environmental /100 | EU AI Act Art. 51 | Notes |
|---|
Sourced from the Ramp AI Index, built from real corporate-card and token-spend transactions on Ramp's platform (not surveys or self-reported usage) — the closest thing to an independent "corporate penetration" figure for AI. It publishes two different metrics; don't conflate them. Business adoption (below) is lab-level: the share of U.S. businesses on Ramp with an active paid subscription to that lab. Enterprise spend share, in the Model Comparison tab, is model-level: each model's share of the dollars Ramp tracks flowing to AI token/subscription spend, published only for the models large enough to rank in Ramp's own top-10 that month — a blank cell there means "not in Ramp's tracked top 10," not zero. Ramp's customer base skews U.S., mid-market and tech-heavy, so treat this as a real but non-representative sample, not global market share.
A different lens on "adoption": OpenRouter measures actual API tokens routed across its platform (prompt + completion), skewing toward individual developers and smaller teams rather than Ramp's corporate-card SMB sample. Ranking is by raw token volume, which reflects usage, not quality — a verbose model processes more tokens for the same task. Only the models from this table that appear in OpenRouter's current top ranks are shown; the rest of OpenRouter's top 5 that week (Hy4 Preview, Space Bunny Alpha, GPT-5.6 Luna) aren't tracked elsewhere in this dashboard so aren't listed below.