{
  "_meta": {
    "title": "benchr verified model figures",
    "purpose": "SINGLE SOURCE OF TRUTH for every model number on benchr. Charts, slides, and article pages must read figures from this file, not from prose. Each figure was confirmed against the provider's OWN official source on the verifiedDate shown.",
    "rules": [
      "Official provider source only (the company's own announcement, docs, pricing page, or model card). Third-party blogs/aggregators/leaderboards never count as a source.",
      "null means the figure could NOT be confirmed from an official source. null is never a guess — it is an honest gap. Read the matching note.",
      "Do not copy numbers from benchr article pages; they may contain errors. This file is the independent, verified record.",
      "Prices are USD per 1,000,000 tokens unless a field name says otherwise (e.g. perImage)."
    ],
    "verifiedDate": "2026-05-31",
    "reVerifiedDate": "2026-09-19",
    "reVerificationNote": "The 2026-09-19 pass researched September 11-19, 2026 across the provider channels and read six official pages. It added Sakana AI as a provider with Fugu Max ($2 / $6, $0.25 cached) and Fugu Ultra v2 ($5 / $30, $0.50 cached, rising to $10 / $45 / $1.00 above 272K tokens of context), both released September 11 and neither publishing a context window. It added Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, generally available September 15 at $0.75 / $4.50 text and $3.00 / $12.00 audio per 1M tokens with 131,072 input and 65,536 output tokens. It added Grok Voice Transcribe 2.0, announced September 18 and billed at $0.10 an hour of audio over REST and $0.20 streaming. It added Atria Dawn Preview, the Shanghai AI Laboratory's MIT-licensed 753B open-weight model, with a null release date because the laboratory published no dated announcement. It corrected GPT-Live-1, which reached API general availability on September 10 under the ID gpt-live-1 at $0.05 per voice minute while the record still said API availability was planned, and recorded that GPT-Live-1 mini did not follow it: its model page returns 404 and the pricing table does not list it. Filling a gap the interval exposed, it added Meta's Muse Spark 1.3 (September 2) and 1.2 (August 5) at $1.25 / $4.25 with $0.15 cached input and a 1,048,576-token window, and gave Muse Spark 1.1 the price Meta had not published at launch and the documented window in place of the rounded one. Anthropic's API release notes, DeepSeek's news page, Mistral's changelog and Alibaba Model Studio's release list were read for the same interval: Anthropic shipped platform features and no model or price change, DeepSeek confirmed on September 14 that V4-Pro billing continues unchanged, Mistral published nothing in September, and Alibaba launched three interactive world models on September 17 (happyoyster-1.0-adventure, -directing, -acting) with neither a price nor a context window, so benchr carries no record for them. The 2026-09-11 pass read DeepSeek's September 10, 2026 release note and pricing page: it added DeepSeek-V4.1-Flash (deepseek-flash, $0.30 / $1.20 peak, half off-peak, 1M context, 384K output), marked DeepSeek-V4-Flash and V4-Flash-Vision-Exp retired on September 10, and recorded that a V4-Pro phase-out announced for September 14 was withdrawn. It added GPT-6 Astra's 128,000-token maximum output from OpenAI's model page, and re-read without change the Anthropic models overview (Claude Haiku 4.5), the Gemini API pricing page (Gemini 3.1 Pro), xAI's model docs (Grok 4.6, Grok 4.3) and the Hugging Face model cards for Llama 4 Scout, Llama 4 Maverick, Phi-4 and Qwen3.6-27B. The 2026-09-02 pass re-read the official Anthropic release announcement, Fable 5.1 model documentation, migration guide, pricing documentation, and lifecycle table. It added Claude Fable 5.1 and Claude Mythos 5.1, released September 1, 2026. Both list a 1M-token context window, 128K maximum output, $10/$50 standard input/output per 1M tokens, $0.25 cache-read input, $12.50 five-minute cache writes, $20 one-hour cache writes, and Batch pricing at half of standard input/output. Fable 5.1 is generally available; Mythos 5.1 remains restricted to vetted Project Glasswing access. The pass also rechecked the continuing active status of Fable 5. The 2026-08-28 pass re-read the official DeepSeek and Z.ai pricing and model pages. It recorded DeepSeek's August 16 move to peak/off-peak billing - a price increase of roughly 3x on input and 4.7x on output at peak, with off-peak set at half the peak rate - and added DeepSeek-V4-Flash-Vision-Exp (August 21), GLM-5.3 (August 18), and GLM-5.3-Flash (August 26). Anthropic, OpenAI, Google, xAI, and Mistral published no model or price change between August 24 and August 28, 2026. The 2026-08-24 pass re-read the official OpenAI, Google, and xAI pricing and model pages. It recorded OpenAI's August 21 GPT-5.6 Sol price cut ($5/$30 -> $4/$20, promotional at least through November 21, 2026), Google's promotional halving of Gemini 3.6 Flash through December 31, 2026, and added two models the ledger did not carry: Gemini 3.7 Flash (GA August 13, 2026) and Grok 4.6 (announced August 12, 2026). Google publishes no benchmark table for Gemini 3.7 Flash, so those fields are null; the Grok 4.6 benchmark figures are xAI-reported launch numbers, not benchr tests. Entries not named here were left untouched because nothing changed on their official sources.",
    "sourceLegend": "Each model has a `sources` map. A figure's source = the URL in `sources` for that figure's category (pricing, benchmarks, context, release, license). All figures share the model-level `verifiedDate` unless a per-figure note says otherwise.",
    "currency": "USD",
    "note": "This file is distinct from assets/data/models.json, which powers the benchr tools with editorial 0-100 capability profiles. Those profiles are not official figures or lab measurements. Public latency fields stay null unless a reproducible measurement or a clearly attributed provider figure is available. Verified provider facts live here.",
    "entryCount": null
  },
  "models": [
    {
      "id": "claude-opus-4-8",
      "name": "Claude Opus 4.8",
      "provider": "Anthropic",
      "apiModelId": "claude-opus-4-8",
      "license": "proprietary",
      "releaseDate": "2026-05-28",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000,
        "maxOutputTokensBeta": 300000
      },
      "pricing": {
        "inputPerM": 5.0,
        "outputPerM": 25.0,
        "fastModeInputPerM": 10.0,
        "fastModeOutputPerM": 50.0,
        "batchInputPerM": 2.5,
        "batchOutputPerM": 12.5,
        "cachedInputPerM": 0.5
      },
      "tentativeRetirementFloor": "2027-05-28",
      "benchmarks": {
        "SWE-bench Verified": 88.6,
        "SWE-bench Pro": 69.2,
        "SWE-bench Multilingual": 84.4,
        "SWE-bench Multimodal": 38.4,
        "Terminal-Bench 2.1": 74.6,
        "GPQA Diamond": 93.6,
        "OSWorld-Verified": 83.4,
        "BrowseComp (single-agent)": 84.3,
        "Humanity's Last Exam (no tools)": 49.8,
        "GDPval-AA (Elo)": 1890,
        "ARC-AGI-2": null
      },
      "sources": {
        "release": "https://www.anthropic.com/news/claude-opus-4-8",
        "pricing": "https://platform.claude.com/docs/en/about-claude/pricing",
        "context": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "benchmarks": "https://www.anthropic.com/claude-opus-4-8-system-card"
      },
      "verifiedDate": "2026-06-12",
      "notes": "Standard price unchanged from Opus 4.7. The $10/$50 figure is the optional fast-mode rate (~2.5x output speed), NOT the base price; base is $5/$25. Cache-hit input $0.50 (0.1x base) now explicit on the official pricing page (confirmed June 12, 2026). Official tentative retirement floor: not sooner than May 28, 2027. ARC-AGI-2 is not in the official Opus 4.8 headline summary table, so it is null (do not state one). New Opus 4.7+ tokenizer can use up to ~35% more tokens per the pricing page."
    },
    {
      "id": "claude-opus-4-7",
      "name": "Claude Opus 4.7",
      "provider": "Anthropic",
      "apiModelId": "claude-opus-4-7",
      "license": "proprietary",
      "releaseDate": "2026-04-16",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000,
        "maxOutputTokensBeta": 300000
      },
      "pricing": {
        "inputPerM": 5.0,
        "outputPerM": 25.0,
        "fastModeInputPerM": 30.0,
        "fastModeOutputPerM": 150.0,
        "batchInputPerM": 2.5,
        "batchOutputPerM": 12.5,
        "cachedInputPerM": 0.5
      },
      "tentativeRetirementFloor": "2027-04-16",
      "benchmarks": {
        "SWE-bench Verified": 87.6,
        "SWE-bench Pro": 64.3,
        "SWE-bench Multilingual": 80.5,
        "SWE-bench Multimodal": 34.5,
        "Terminal-Bench 2.0": 69.4,
        "GPQA Diamond": 94.2
      },
      "sources": {
        "release": "https://www.anthropic.com/news/claude-opus-4-7",
        "pricing": "https://platform.claude.com/docs/en/about-claude/pricing",
        "context": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "benchmarks": "https://www.anthropic.com/claude-opus-4-7-system-card"
      },
      "verifiedDate": "2026-06-12",
      "notes": "Fast mode on 4.7 is $30/$150 (3x more expensive than Opus 4.8's $10/$50 fast mode). Cache-hit input $0.50 confirmed on the official pricing page June 12, 2026; official tentative retirement floor: not sooner than April 16, 2027. Opus 4.8 later restated Opus 4.7's OSWorld score upward (~82.x) after a test-harness fix."
    },
    {
      "id": "claude-sonnet-4-6",
      "name": "Claude Sonnet 4.6",
      "provider": "Anthropic",
      "apiModelId": "claude-sonnet-4-6",
      "license": "proprietary",
      "releaseDate": "2026-02-17",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 64000,
        "maxOutputTokensBeta": 300000
      },
      "pricing": {
        "inputPerM": 3.0,
        "outputPerM": 15.0,
        "batchInputPerM": 1.5,
        "batchOutputPerM": 7.5,
        "cachedInputPerM": 0.3
      },
      "tentativeRetirementFloor": "2027-02-17",
      "benchmarks": {
        "SWE-bench Verified": 79.6,
        "SWE-bench Multilingual": 75.9,
        "Terminal-Bench 2.0": 59.1,
        "OSWorld-Verified": 72.5,
        "GPQA Diamond": 89.9,
        "MMMLU": 89.3,
        "AIME 2025 (no tools)": 95.6,
        "Humanity's Last Exam (no tools)": 33.2,
        "Humanity's Last Exam (with tools)": 49.0,
        "ARC-AGI-2": 58.3,
        "tau2-bench Telecom": 97.9,
        "tau2-bench Retail": 91.7,
        "GDPval-AA (Elo)": 1633
      },
      "sources": {
        "release": "https://www.anthropic.com/news/claude-sonnet-4-6",
        "pricing": "https://platform.claude.com/docs/en/about-claude/pricing",
        "context": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "benchmarks": "https://www-cdn.anthropic.com/bbd8ef16d70b7a1665f14f306ee88b53f686aa75/Claude%20Sonnet%204.6%20System%20Card.pdf"
      },
      "verifiedDate": "2026-07-30",
      "notes": "No fast-mode tier (fast mode is Opus-only). Cache-hit input $0.30 confirmed on the official pricing page June 12, 2026; official tentative retirement floor: not sooner than February 17, 2027. Benchmark values read from the official Sonnet 4.6 System Card (Table 2.1.A). SWE-bench Verified 79.6% averaged over 25 trials (80.2% with a stated prompt modification). Anthropic flags possible AIME 2025 contamination. Anthropic reports MMMLU, not plain MMLU."
    },
    {
      "id": "claude-haiku-4-5",
      "name": "Claude Haiku 4.5",
      "provider": "Anthropic",
      "apiModelId": "claude-haiku-4-5",
      "license": "proprietary",
      "releaseDate": "2025-10-15",
      "context": {
        "windowTokens": 200000,
        "maxOutputTokens": 64000,
        "maxOutputTokensBeta": null
      },
      "pricing": {
        "inputPerM": 1.0,
        "outputPerM": 5.0,
        "batchInputPerM": 0.5,
        "batchOutputPerM": 2.5,
        "cachedInputPerM": 0.1
      },
      "tentativeRetirementFloor": "2026-10-15",
      "benchmarks": {
        "SWE-bench Verified": 73.3,
        "Terminal-Bench (no thinking)": 40.21,
        "Terminal-Bench (32K thinking)": 41.75,
        "GPQA Diamond": null,
        "OSWorld-Verified": null,
        "AIME 2025": null,
        "MMMLU": null
      },
      "sources": {
        "release": "https://www.anthropic.com/news/claude-haiku-4-5",
        "pricing": "https://platform.claude.com/docs/en/about-claude/pricing",
        "context": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "benchmarks": "https://www.anthropic.com/news/claude-haiku-4-5"
      },
      "verifiedDate": "2026-09-11",
      "notes": "Pinned snapshot claude-haiku-4-5-20251001. Pricing (incl. cache-hit $0.10) and tentative retirement floor (not sooner than October 15, 2026) re-confirmed on official pages June 12, 2026; the 200K context window stands from the May 31 verification. SWE-bench Verified 73.3% (avg of 50 trials, 128K thinking budget) and Terminal-Bench are the only headline numbers Anthropic publishes as readable official text; GPQA/OSWorld/AIME/MMMLU appear only inside a launch-page image, so they are null (not guessed)."
    },
    {
      "id": "claude-mythos-preview",
      "name": "Claude Mythos Preview",
      "provider": "Anthropic",
      "apiModelId": null,
      "license": "proprietary",
      "releaseDate": null,
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": null,
        "maxOutputTokensBeta": null
      },
      "pricing": {
        "inputPerM": 25.0,
        "outputPerM": 125.0,
        "note": "Stated price for approved Project Glasswing participants only; not generally purchasable."
      },
      "benchmarks": {
        "SWE-bench Pro": 77.8
      },
      "availability": "RESTRICTED — research preview under Project Glasswing. Not generally available. Invitation-only (12 named launch partners + 40+ critical-infrastructure orgs). No self-serve sign-up; no public API model id.",
      "sources": {
        "release": "https://www.anthropic.com/glasswing",
        "pricing": "https://www.anthropic.com/glasswing",
        "context": "https://platform.claude.com/docs/en/about-claude/pricing",
        "benchmarks": "https://www.anthropic.com/glasswing"
      },
      "verifiedDate": "2026-05-31",
      "notes": "Anthropic: 'We do not plan to make Claude Mythos Preview generally available.' Purpose-built for defensive cybersecurity / vulnerability research. No GA date. 1M context confirmed only via the pricing page's long-context list. Frame any benchr page as restricted/not-for-public."
    },
    {
      "id": "gpt-5",
      "name": "GPT-5",
      "provider": "OpenAI",
      "apiModelId": "gpt-5",
      "license": "proprietary",
      "releaseDate": "2025-08-07",
      "context": {
        "windowTokens": 400000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 1.25,
        "outputPerM": 10.0,
        "cachedInputPerM": 0.125
      },
      "benchmarks": {
        "SWE-bench Verified": 74.9,
        "AIME 2025": null,
        "GPQA Diamond": null,
        "HealthBench Hard": 46.2
      },
      "sources": {
        "release": "https://deploymentsafety.openai.com/gpt-5",
        "pricing": "https://developers.openai.com/api/docs/models/gpt-5",
        "context": "https://developers.openai.com/api/docs/models/gpt-5",
        "benchmarks": "https://cdn.openai.com/gpt-5-system-card.pdf"
      },
      "verifiedDate": "2026-06-12",
      "notes": "SWE-bench Verified 74.9% is officially OpenAI's launch-blog figure (verbosity=medium), confirmed via the system card PDF which cites it (p.36). AIME 2025 (~94.6%) and GPQA Diamond (~88.4%) are widely attributed to the launch page openai.com/index/introducing-gpt-5, which blocks automated fetchers (403); they could NOT be re-read from an accessible official source, so they are null. The 88.4% figure may refer to GPT-5 pro, not base GPT-5. OpenAI docs now label GPT-5 the previous model."
    },
    {
      "id": "gpt-5-5",
      "name": "GPT-5.5",
      "provider": "OpenAI",
      "apiModelId": "gpt-5.5",
      "license": "proprietary",
      "releaseDate": "2026-04-23",
      "context": {
        "windowTokens": 1050000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 5.0,
        "outputPerM": 30.0,
        "cachedInputPerM": 0.5,
        "batchInputPerM": 2.5,
        "batchOutputPerM": 15.0,
        "longContextInputPerM": 10.0,
        "longContextCachedInputPerM": 1.0,
        "longContextOutputPerM": 45.0,
        "regionalProcessingUplift": 0.1,
        "proInputPerM": 30.0,
        "proOutputPerM": 180.0,
        "proBatchInputPerM": 15.0,
        "proBatchOutputPerM": 90.0,
        "extendedContextSurcharge": "For sessions >272K input tokens: 2x input, 1.5x output (standard/batch/flex)."
      },
      "benchmarks": {
        "HealthBench (length-adjusted)": 56.5,
        "HealthBench Professional": 51.8,
        "SWE-bench Verified": null,
        "SWE-bench Pro": null,
        "Terminal-Bench 2.0": null,
        "OSWorld-Verified": null
      },
      "sources": {
        "release": "https://deploymentsafety.openai.com/gpt-5-5",
        "pricing": "https://developers.openai.com/api/docs/models/gpt-5.5",
        "context": "https://developers.openai.com/api/docs/models/gpt-5.5",
        "benchmarks": "https://deploymentsafety.openai.com/gpt-5-5"
      },
      "verifiedDate": "2026-07-30",
      "notes": "Flagship GPT-5.5 (and GPT-5.5 Pro) announced Apr 23, 2026 (do not confuse with GPT-5.5 Instant, the ChatGPT default released May 5, 2026). Context is 1,050,000 (not a round 1M). HealthBench figures are the only official system-card benchmarks. Rechecked 2026-07-30 against developers.openai.com/api/docs/models/gpt-5.5 and /api/docs/pricing: standard $5/$30 with $0.50 cached input; short-context batch $2.50/$15; long-context $10 input, $1 cached input, $45 output above 272K; eligible regional processing carries a 10% uplift. GPT-5.5 Pro is $30/$180 standard and $15/$90 batch with no cached-input discount. The coding/agent numbers circulating (Terminal-Bench 82.7%, OSWorld 78.7%, SWE-bench Pro 58.6%) remain null because the checked model page does not publish directly comparable figures."
    },
    {
      "id": "gpt-5-mini",
      "name": "GPT-5 Mini",
      "provider": "OpenAI",
      "apiModelId": "gpt-5-mini",
      "license": "proprietary",
      "releaseDate": "2025-08-07",
      "context": {
        "windowTokens": 400000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 0.25,
        "outputPerM": 2.0,
        "cachedInputPerM": 0.025,
        "batchInputPerM": 0.125,
        "batchOutputPerM": 1.0
      },
      "benchmarks": {},
      "sources": {
        "release": "https://developers.openai.com/api/docs/models/gpt-5-mini",
        "pricing": "https://developers.openai.com/api/docs/models/gpt-5-mini",
        "context": "https://developers.openai.com/api/docs/models/gpt-5-mini"
      },
      "verifiedDate": "2026-07-30",
      "notes": "Official standard price is $0.25/$2.00 with $0.025 cached input. The OpenAI pricing table lists short-context batch at $0.125/$1.00; rechecked 2026-07-30. Release date is traced to snapshot id gpt-5-mini-2025-08-07. The checked model page publishes no SWE-bench or numeric throughput result, so formal benchmark fields remain empty."
    },
    {
      "id": "gpt-image-2",
      "name": "ChatGPT Images 2.0 (GPT Image 2)",
      "provider": "OpenAI",
      "apiModelId": "gpt-image-2",
      "license": "proprietary",
      "releaseDate": "2026-04-21",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "imageInputPerM": 8.0,
        "cachedImageInputPerM": 2.0,
        "imageOutputPerM": 30.0,
        "textInputPerM": 5.0,
        "batchImageInputPerM": 4.0,
        "batchImageOutputPerM": 15.0,
        "perImage": null
      },
      "benchmarks": {},
      "sources": {
        "release": "https://developers.openai.com/api/docs/models/gpt-image-2",
        "pricing": "https://developers.openai.com/api/docs/pricing"
      },
      "verifiedDate": "2026-05-31",
      "notes": "OpenAI prices this model by TOKENS, not per image. There is NO official per-image list price; per-image figures like ~$0.006 (low) / ~$0.053 (medium) / ~$0.211 (high) for 1024x1024 are third-party calculator estimates, so perImage is null. Release date inferred from snapshot gpt-image-2-2026-04-21. Supersedes GPT Image 1."
    },
    {
      "id": "gemini-3-pro",
      "name": "Gemini 3 Pro",
      "provider": "Google",
      "apiModelId": "gemini-3-pro-preview",
      "license": "proprietary",
      "releaseDate": "2025-11-18",
      "status": "DEPRECATED — shut down 2026-03-09; replaced by Gemini 3.1 Pro.",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": 65536
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null,
        "note": "No longer on the live official pricing page (model retired). When live it was $2/$12 (<=200K) and $4/$18 (>200K)."
      },
      "benchmarks": {
        "SWE-bench Verified": 76.2,
        "GPQA Diamond": 91.9,
        "Humanity's Last Exam (no tools)": 37.5,
        "Terminal-Bench 2.0": 54.2,
        "LMArena (Elo)": 1501,
        "SimpleQA Verified": 72.1,
        "MMMU-Pro": 81.0,
        "MathArena Apex": 23.4
      },
      "sources": {
        "release": "https://blog.google/products-and-platforms/products/gemini/gemini-3/",
        "status": "https://ai.google.dev/gemini-api/docs/deprecations",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3-pro-preview",
        "benchmarks": "https://blog.google/products-and-platforms/products/gemini/gemini-3/"
      },
      "verifiedDate": "2026-05-31",
      "notes": "Live price is null because the retired model is no longer on the official pricing page (historical $2/$12 kept only as a note). Benchmarks from the Gemini 3 launch blog. Max output 65,536 on the model spec page (an older guide page said 64,000)."
    },
    {
      "id": "gemini-3-1-pro",
      "name": "Gemini 3.1 Pro",
      "provider": "Google",
      "apiModelId": "gemini-3.1-pro-preview",
      "license": "proprietary",
      "releaseDate": "2026-02-19",
      "status": "Preview (GA coming soon).",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 64000
      },
      "pricing": {
        "inputPerM": 2.0,
        "outputPerM": 12.0,
        "inputAbove200kPerM": 4.0,
        "outputAbove200kPerM": 18.0,
        "batchInputPerM": 1.0,
        "batchOutputPerM": 6.0,
        "batchInputAbove200kPerM": 2.0,
        "batchOutputAbove200kPerM": 9.0,
        "cachedInputPerM": 0.2,
        "freeApiTier": false
      },
      "benchmarks": {
        "ARC-AGI-2": 77.1,
        "GPQA Diamond": 94.3,
        "Humanity's Last Exam (with tools)": 51.4,
        "MMMU-Pro": 80.5,
        "SWE-bench Verified": 80.6,
        "MMMLU": 92.6
      },
      "sources": {
        "release": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/",
        "pricing": "https://ai.google.dev/gemini-api/docs/pricing",
        "context": "https://deepmind.google/models/model-cards/gemini-3-1-pro/",
        "benchmarks": "https://deepmind.google/models/model-cards/gemini-3-1-pro/"
      },
      "verifiedDate": "2026-09-11",
      "notes": "Tiered pricing: the >200K-token tier doubles input ($4) and raises output ($18). Output price includes thinking tokens. No free API tier (free trial in AI Studio UI only). Context-caching: cached input $0.20/1M (90% off the $2 input rate; Google additionally charges $1 per 1M-token-hour of cache storage), verified 2026-06-15 against ai.google.dev/gemini-api/docs/pricing."
    },
    {
      "id": "gemini-3-6-flash",
      "name": "Gemini 3.6 Flash",
      "provider": "Google",
      "apiModelId": "gemini-3.6-flash",
      "license": "proprietary",
      "releaseDate": "2026-07-21",
      "status": "GA (stable).",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": 65536
      },
      "pricing": {
        "inputPerM": 0.75,
        "outputPerM": 3.75,
        "batchInputPerM": 0.375,
        "batchOutputPerM": 1.875,
        "flexInputPerM": 0.375,
        "flexOutputPerM": 1.875,
        "priorityInputPerM": 1.35,
        "priorityOutputPerM": 6.75,
        "cachedInputPerM": 0.075,
        "cacheStoragePerMTokenHour": 0.5,
        "postPromoInputPerM": 1.5,
        "postPromoOutputPerM": 7.5,
        "postPromoCachedInputPerM": 0.15,
        "postPromoStartDate": "2027-01-01",
        "freeApiTier": true,
        "note": "Google's live pricing table now prices Gemini 3.6 Flash at $0.75 input / $3.75 output per 1M tokens through December 31, 2026, then $1.50 / $7.50 starting January 1, 2027. Context caching is $0.075 (then $0.15) plus $0.50 per 1M-token-hour of storage (then $1.00). Batch and Flex are half of standard; Priority is $1.35 / $6.75. The $1.50 / $7.50 pair recorded on 2026-07-22 is the post-promotional rate. Read on ai.google.dev/gemini-api/docs/pricing on 2026-08-24."
      },
      "benchmarks": {
        "Terminal-Bench 2.1": null,
        "GDPval-AA (Elo)": null,
        "MCP Atlas": null,
        "CharXiv Reasoning": null,
        "SWE-bench Verified": null,
        "GPQA Diamond": null
      },
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "pricing": "https://ai.google.dev/gemini-api/docs/pricing",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash",
        "deprecations": "https://ai.google.dev/gemini-api/docs/deprecations"
      },
      "verifiedDate": "2026-08-24",
      "notes": "Google released Gemini 3.6 Flash as a stable GA model on July 21, 2026, with improved token efficiency and code/agentic planning claims. Official model page lists text/image/video/audio/PDF inputs, text output, 1,048,576 input tokens, 65,536 output tokens, caching, code execution, file search, function calling, Search/Maps grounding, structured outputs, thinking, URL context, and Computer Use support in preview. No official benchmark table found on the model page or release note, so hard benchmark fields are null. Re-verified 2026-08-24: Google now lists a promotional rate that halves input and output through December 31, 2026, matching the Gemini 3.7 Flash launch price."
    },
    {
      "id": "gemini-3-7-flash",
      "name": "Gemini 3.7 Flash",
      "provider": "Google",
      "apiModelId": "gemini-3.7-flash",
      "license": "proprietary",
      "releaseDate": "2026-08-13",
      "status": "GA (stable) since August 13, 2026.",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": 65536
      },
      "pricing": {
        "inputPerM": 0.75,
        "outputPerM": 3.75,
        "batchInputPerM": 0.375,
        "batchOutputPerM": 1.875,
        "flexInputPerM": 0.375,
        "flexOutputPerM": 1.875,
        "priorityInputPerM": 1.35,
        "priorityOutputPerM": 6.75,
        "cachedInputPerM": 0.075,
        "cacheStoragePerMTokenHour": 0.5,
        "postPromoInputPerM": 1.5,
        "postPromoOutputPerM": 7.5,
        "postPromoCachedInputPerM": 0.15,
        "postPromoStartDate": "2027-01-01",
        "freeApiTier": true,
        "note": "Introductory pricing: $0.75 input / $3.75 output per 1M tokens through December 31, 2026, then $1.50 / $7.50 starting January 1, 2027. Context caching is $0.075 (then $0.15) plus $0.50 per 1M-token-hour of storage (then $1.00). Batch and Flex are half of standard; Priority is $1.35 / $6.75. A free tier is offered. Read on ai.google.dev/gemini-api/docs/pricing on 2026-08-24."
      },
      "benchmarks": {
        "SWE-bench Verified": null,
        "GPQA Diamond": null,
        "Terminal-Bench 2.1": null,
        "note": "Google published no benchmark table with the release note or on the model page, so every formal benchmark field stays null."
      },
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "pricing": "https://ai.google.dev/gemini-api/docs/pricing",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash",
        "deprecations": "https://ai.google.dev/gemini-api/docs/deprecations"
      },
      "verifiedDate": "2026-08-24",
      "notes": "Google made Gemini 3.7 Flash generally available on August 13, 2026 and describes it as its most intelligent workhorse model for coding and agents, with improvements across software engineering, web development, and agentic workflows. The model page lists text, image, video, audio, and PDF input with text output, a 1,048,576-token input limit, a 65,536-token output limit, caching, code execution, file search, function calling, Search and Maps grounding, structured outputs, thinking at low, medium, or high (no minimal level), URL context, and Computer Use in preview. Live API, image generation, and audio generation are not supported. Google publishes no benchmark figures for it, so none are claimed here.",
      "arabic": {
        "status": "متاح عمومًا (نسخة مستقرة) منذ 13 أغسطس 2026.",
        "notes": "أطلقت Google نموذج Gemini 3.7 Flash بشكل عام في 13 أغسطس 2026 ووصفته بأنه أذكى نموذج عملي لديها للبرمجة والوكلاء. صفحة النموذج تذكر مدخلات نصية وصورية وفيديو وصوت وPDF مع إخراج نصي، وسعة سياق 1,048,576 توكن، وحد إخراج 65,536 توكن. السعر التمهيدي 0.75$ للمدخلات و3.75$ للمخرجات لكل مليون توكن حتى 31 ديسمبر 2026، ثم 1.50$ و7.50$ من 1 يناير 2027. لم تنشر Google أي أرقام بنشمارك رسمية لهذا النموذج."
      }
    },
    {
      "id": "gemini-3-5-flash",
      "name": "Gemini 3.5 Flash",
      "provider": "Google",
      "apiModelId": "gemini-3.5-flash",
      "license": "proprietary",
      "releaseDate": "2026-05-19",
      "status": "GA (stable).",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": 65536
      },
      "pricing": {
        "inputPerM": 1.5,
        "outputPerM": 9.0,
        "batchInputPerM": 0.75,
        "batchOutputPerM": 4.5,
        "cachedInputPerM": 0.15,
        "freeApiTier": true,
        "note": "Flat rate — no >200K context tier. Cached input $0.15/1M (90% off the $1.50 input rate; Google additionally charges $1 per 1M-token-hour of cache storage), verified 2026-06-15 against ai.google.dev/gemini-api/docs/pricing."
      },
      "benchmarks": {
        "Terminal-Bench 2.1": 76.2,
        "GDPval-AA (Elo)": 1656,
        "MCP Atlas": 83.6,
        "CharXiv Reasoning": 84.2,
        "SWE-bench Verified": null,
        "GPQA Diamond": null
      },
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/deprecations",
        "pricing": "https://ai.google.dev/gemini-api/docs/pricing",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash",
        "benchmarks": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/"
      },
      "verifiedDate": "2026-07-30",
      "notes": "Released at Google I/O (May 19, 2026), GA. Google's live pricing page checked 2026-07-30 lists standard $1.50/$9, batch $0.75/$4.50, cached input $0.15 plus cache-storage charges, and a free API tier. The launch blog publishes only the four agentic/coding/multimodal benchmarks above; no official SWE-bench Verified or GPQA figure is published for 3.5 Flash, so those fields are null. No universal latency figure is recorded."
    },
    {
      "id": "grok-4-3",
      "name": "Grok 4.3",
      "provider": "xAI",
      "apiModelId": "grok-4.3",
      "license": "proprietary",
      "releaseDate": null,
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 1.25,
        "outputPerM": 2.5,
        "cachedInputPerM": 0.2,
        "longContextThresholdTokens": 200000,
        "longContextInputPerM": 2.5,
        "longContextCachedInputPerM": 0.4,
        "longContextOutputPerM": 5.0,
        "batchInputPerM": 1.0,
        "batchCachedInputPerM": 0.16,
        "batchOutputPerM": 2.0,
        "batchLongContextInputPerM": 2.0,
        "batchLongContextCachedInputPerM": 0.32,
        "batchLongContextOutputPerM": 4.0,
        "note": "xAI charges the long-context tier for the entire request once the prompt reaches 200K tokens. The Batch API discount is 20% for Grok 4.3."
      },
      "benchmarks": {},
      "sources": {
        "pricing": "https://docs.x.ai/developers/models/grok-4.3",
        "context": "https://docs.x.ai/developers/models/grok-4.3"
      },
      "verifiedDate": "2026-09-11",
      "notes": "xAI officially publishes short-context price ($1.25 input / $0.20 cached / $2.50 output), long-context price at or above a 200K-token prompt ($2.50 / $0.40 / $5.00), a 20% Batch discount, 1M context, and the model id. It publishes no official release date, max output, or numeric benchmark table for Grok 4.3; those fields remain null. Figures like 'Intelligence Index 53' or 'tau2-bench Telecom 98%' are third-party, not official xAI numbers."
    },
    {
      "id": "deepseek-v4-flash",
      "name": "DeepSeek-V4-Flash",
      "provider": "DeepSeek",
      "apiModelId": "deepseek-v4-flash",
      "license": "MIT",
      "releaseDate": "2026-04-24",
      "params": {
        "total": "284B",
        "active": "13B"
      },
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 384000
      },
      "pricing": {
        "inputPerM": 0.44,
        "outputPerM": 1.32,
        "cacheHitInputPerM": 0.014,
        "offPeakInputPerM": 0.22,
        "offPeakOutputPerM": 0.66,
        "offPeakCacheHitInputPerM": 0.007,
        "offPeakDiscount": 0.5,
        "peakHoursUTC": "01:00-04:00 and 06:00-10:00, Monday to Friday",
        "priceEffectiveFrom": "2026-08-16T16:00:00Z",
        "previousFlatPricing": {
          "inputPerM": 0.14,
          "outputPerM": 0.28,
          "cacheHitInputPerM": 0.0028
        },
        "selfHost": true,
        "note": "DeepSeek moved to peak/off-peak billing at 16:00 UTC on 2026-08-16. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday; every other hour bills at half the peak rate. The figures in inputPerM/outputPerM/cacheHitInputPerM are the peak (list) rates; the offPeak* fields are the discounted rates. Read on api-docs.deepseek.com/quick_start/pricing on 2026-08-28."
      },
      "benchmarks": {
        "SWE-bench Verified": 79.0,
        "GPQA Diamond": 88.1,
        "LiveCodeBench": 91.6,
        "MMLU-Pro (Think Max)": 86.2,
        "HMMT 2026 Feb": 94.8,
        "MRCR 1M": 78.7
      },
      "sources": {
        "release": "https://api-docs.deepseek.com/news/news260424",
        "pricing": "https://api-docs.deepseek.com/quick_start/pricing",
        "license": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "context": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "benchmarks": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash",
        "retirement": "https://api-docs.deepseek.com/news/news260910"
      },
      "verifiedDate": "2026-09-11",
      "notes": "RETIRED 2026-09-10: DeepSeek's September 10, 2026 release note says \"V4-Flash & V4-Flash-Vision-Exp are retired\" and routes the legacy ID to DeepSeek-V4.1-Flash; the pricing page read on 2026-09-11 no longer lists DeepSeek-V4-Flash. The record is kept as the account of what it cost and scored while it was available. UPDATED 2026-08-28: DeepSeek raised API prices and introduced peak/off-peak billing effective 16:00 UTC on August 16, 2026, announced alongside the V4 lineup release. Cache-miss input went from a flat $0.14 to $0.44 peak / $0.22 off-peak per 1M tokens, output from $0.28 to $1.32 / $0.66, and cache-hit input from $0.0028 to $0.014 / $0.007. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, so roughly four fifths of the week bills at the off-peak rate; that share is a benchr calculation from the published hours, not a provider figure. SWE-bench Verified 79.0% is DeepSeek's provider-reported model-card figure, not an independently reproduced result. Open weights, MIT.",
      "status": "Retired on the DeepSeek API on September 10, 2026. The deepseek-v4-flash ID temporarily routes to DeepSeek-V4.1-Flash (deepseek-flash) at V4.1-Flash rates. The prices below are the last rates DeepSeek listed for this model."
    },
    {
      "id": "deepseek-v4-pro",
      "name": "DeepSeek-V4-Pro",
      "provider": "DeepSeek",
      "apiModelId": "deepseek-v4-pro",
      "license": "MIT",
      "releaseDate": "2026-04-24",
      "params": {
        "total": "1.6T",
        "active": "49B"
      },
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 384000
      },
      "pricing": {
        "inputPerM": 1.32,
        "outputPerM": 3.96,
        "cacheHitInputPerM": 0.044,
        "offPeakInputPerM": 0.66,
        "offPeakOutputPerM": 1.98,
        "offPeakCacheHitInputPerM": 0.022,
        "offPeakDiscount": 0.5,
        "peakHoursUTC": "01:00-04:00 and 06:00-10:00, Monday to Friday",
        "priceEffectiveFrom": "2026-08-16T16:00:00Z",
        "previousFlatPricing": {
          "inputPerM": 0.435,
          "outputPerM": 0.87,
          "cacheHitInputPerM": 0.003625
        },
        "selfHost": true,
        "note": "DeepSeek moved to peak/off-peak billing at 16:00 UTC on 2026-08-16. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday; every other hour bills at half the peak rate. The figures in inputPerM/outputPerM/cacheHitInputPerM are the peak (list) rates; the offPeak* fields are the discounted rates. Read on api-docs.deepseek.com/quick_start/pricing on 2026-08-28."
      },
      "benchmarks": {
        "SWE-bench Verified": 80.6,
        "SWE-bench Pro": 55.4,
        "GPQA Diamond": 90.1,
        "LiveCodeBench": 93.5,
        "Terminal-Bench 2.0": 67.9,
        "MMLU-Pro (Max)": 87.5,
        "MRCR 1M": 83.5,
        "Codeforces (rating)": 3206,
        "Terminal-Bench 2.1": 87.9,
        "Humanity's Last Exam (no tools)": 42.7,
        "Humanity's Last Exam (with tools)": 60.0,
        "NL2Repo": 61.5,
        "CyberGym": 83.3,
        "DeepSWE": 62.7,
        "Toolathlon-Verified": 74.1,
        "DSBench-Hard": 67.2
      },
      "sources": {
        "release": "https://api-docs.deepseek.com/news/news260424",
        "pricing": "https://api-docs.deepseek.com/quick_start/pricing",
        "license": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro",
        "context": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro",
        "benchmarks": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro",
        "ga_update": "https://api-docs.deepseek.com/updates",
        "lifecycle": "https://api-docs.deepseek.com/news/news"
      },
      "verifiedDate": "2026-09-11",
      "notes": "UPDATED 2026-08-28: DeepSeek raised API prices and introduced peak/off-peak billing effective 16:00 UTC on August 16, 2026. Cache-miss input went from a flat $0.435 to $1.32 peak / $0.66 off-peak per 1M tokens, output from $0.87 to $3.96 / $1.98, and cache-hit input from $0.003625 to $0.044 / $0.022 - the steepest line on the sheet at roughly twelve times the old cache-hit rate. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday. DeepSeek also shipped a V4-Pro GA update on August 13, 2026 with a new provider-reported benchmark table (Terminal Bench 2.1 87.9, HLE 42.7 without tools / 60.0 with tools, NL2Repo 61.5, CyberGym 83.3, DeepSWE 62.7, Toolathlon-Verified 74.1, DSBench-Hard 67.2); those are DeepSeek's own numbers, not benchr tests. The April model-card figures are retained as published. Open weights, MIT. RECHECKED 2026-09-11: the pricing page lists deepseek-v4-pro as version DeepSeek-V4-Pro-0813 at the same $1.32 / $3.96 peak and $0.66 / $1.98 off-peak rates, 1M context and 384K output. DeepSeek's September 10, 2026 release note announced that deepseek-v4-pro requests would route to V4.1-Flash from 04:00 UTC on September 14; a later entry on DeepSeek's news page withdrew that, saying it will \"continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged.\" The withdrawn date is recorded here rather than as a deprecation."
    },
    {
      "id": "kimi-k2-6",
      "name": "Kimi K2.6",
      "provider": "Moonshot AI",
      "apiModelId": "kimi-k2.6",
      "license": "Modified MIT",
      "releaseDate": null,
      "params": {
        "total": "1T",
        "active": "32B",
        "experts": "384 (8 active/token)"
      },
      "context": {
        "windowTokens": 262144,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 0.95,
        "outputPerM": 4.0,
        "cacheHitInputPerM": 0.16,
        "selfHost": true
      },
      "benchmarks": {
        "SWE-bench Verified": 80.2,
        "SWE-bench Multilingual": 76.7,
        "SWE-bench Pro": 58.6,
        "Terminal-Bench 2.0": 66.7,
        "LiveCodeBench v6": 89.6,
        "AIME 2026": 96.4,
        "GPQA Diamond": 90.5,
        "Humanity's Last Exam (with tools)": 54.0,
        "OSWorld-Verified": 73.1
      },
      "sources": {
        "release": null,
        "pricing": "https://platform.kimi.ai/docs/pricing/chat-k26",
        "license": "https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENSE",
        "context": "https://huggingface.co/moonshotai/Kimi-K2.6",
        "benchmarks": "https://huggingface.co/moonshotai/Kimi-K2.6"
      },
      "verifiedDate": "2026-06-12",
      "notes": "Official release date not stated on any official Moonshot page (third-party says Apr 20, 2026), so releaseDate is null. Context 262,144 (256K) confirmed on both pricing page and model card. Input $0.95 is cache-miss; cache-hit $0.16. Model card body: 1T total / 32B active (the org listing page rounds to 1.1T). Benchmarks are Moonshot's own model-card figures."
    },
    {
      "id": "mistral-large-3",
      "name": "Mistral Large 3",
      "provider": "Mistral AI",
      "apiModelId": "mistral-large-2512",
      "license": "Apache-2.0",
      "releaseDate": "2025-12-02",
      "params": {
        "total": "675B",
        "active": "41B"
      },
      "context": {
        "windowTokens": 256000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 0.5,
        "outputPerM": 1.5,
        "cacheHitInputPerM": null,
        "selfHost": true
      },
      "benchmarks": {},
      "sources": {
        "release": "https://docs.mistral.ai/models/mistral-large-3-25-12",
        "pricing": "https://docs.mistral.ai/models/mistral-large-3-25-12",
        "license": "https://mistral.ai/news/mistral-3/",
        "context": "https://docs.mistral.ai/models/mistral-large-3-25-12",
        "benchmarks": "https://mistral.ai/news/mistral-3/"
      },
      "verifiedDate": "2026-06-12",
      "notes": "Open weights, Apache-2.0. Mistral's announcement gives only relative/leaderboard claims for Large 3, no discrete official per-benchmark scores (the ~85% AIME figure on the page belongs to a smaller reasoning variant, NOT Large 3) — so benchmarks is intentionally empty, not guessed. Cache-hit price not published."
    },
    {
      "id": "mistral-medium-3-5",
      "name": "Mistral Medium 3.5",
      "provider": "Mistral AI",
      "apiModelId": "mistral-medium-3-5",
      "license": "Modified MIT",
      "releaseDate": "2026-04-28",
      "params": {
        "total": "128B",
        "active": "128B",
        "type": "dense"
      },
      "context": {
        "windowTokens": 256000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 1.5,
        "outputPerM": 7.5,
        "cacheHitInputPerM": null,
        "selfHost": true
      },
      "benchmarks": {},
      "sources": {
        "release": "https://mistral.ai/news/vibe-remote-agents-mistral-medium-3-5/",
        "pricing": "https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04",
        "license": "https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04",
        "context": "https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04"
      },
      "verifiedDate": "2026-06-23",
      "notes": "Verified 2026-06-03, re-confirmed 2026-06-23 against the official Mistral model card (docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04): Modified MIT license, open-weight (self-hostable on ~4 GPUs) AND offered as a hosted API, dense 128B, 256K context, $1.50/$7.50 per 1M. Official release date is April 28, 2026 (filled 2026-06-23; previously null). Max output tokens not stated on the card (null). No discrete official per-benchmark scores published — benchmarks intentionally empty."
    },
    {
      "id": "qwen-3-6-27b",
      "name": "Qwen3.6-27B (dense)",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "Qwen/Qwen3.6-27B",
      "license": "Apache-2.0",
      "releaseDate": "2026-04-22",
      "params": {
        "total": "27B",
        "active": "27B",
        "type": "dense"
      },
      "context": {
        "windowTokens": 262144,
        "windowTokensExtended": 1010000,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {
        "SWE-bench Verified": 77.2,
        "SWE-bench Pro": 53.5,
        "Terminal-Bench 2.0": 59.3,
        "MMLU-Pro": 86.2,
        "GPQA Diamond": 87.8,
        "AIME 2026": 94.1,
        "MMMU": 82.9
      },
      "sources": {
        "release": "https://huggingface.co/Qwen/Qwen3.6-27B",
        "license": "https://huggingface.co/Qwen/Qwen3.6-27B",
        "context": "https://huggingface.co/Qwen/Qwen3.6-27B",
        "benchmarks": "https://huggingface.co/Qwen/Qwen3.6-27B"
      },
      "verifiedDate": "2026-09-11",
      "notes": "Open weight, Apache-2.0, self-host (no per-token list price). 262,144 native context, extensible ~1M via YaRN. Benchmarks from the official HF model card."
    },
    {
      "id": "qwen-3-6-35b-a3b",
      "name": "Qwen3.6-35B-A3B (MoE)",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "Qwen/Qwen3.6-35B-A3B",
      "license": "Apache-2.0",
      "releaseDate": "2026-04-16",
      "params": {
        "total": "35B",
        "active": "3B",
        "experts": "256 (8 routed + 1 shared)"
      },
      "context": {
        "windowTokens": 262144,
        "windowTokensExtended": 1010000,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {
        "SWE-bench Verified": 73.4,
        "SWE-bench Multilingual": 67.2,
        "Terminal-Bench 2.0": 51.5,
        "MMLU-Pro": 85.2,
        "AIME 2026": 92.7,
        "GPQA Diamond": 86.0,
        "MMMU": 81.7
      },
      "sources": {
        "release": "https://huggingface.co/Qwen/Qwen3.6-35B-A3B",
        "license": "https://huggingface.co/Qwen/Qwen3.6-35B-A3B",
        "context": "https://huggingface.co/Qwen/Qwen3.6-35B-A3B",
        "benchmarks": "https://huggingface.co/Qwen/Qwen3.6-35B-A3B"
      },
      "verifiedDate": "2026-05-31",
      "notes": "Open weight, Apache-2.0, self-host. UPDATE 2026-06-23: the hosted 'Qwen3.6-Plus' flagship IS now confirmed official on the Alibaba Cloud blog (qwen3.6-plus, 1M context, hosted-only, Apr 2 2026 — see its own entry); the earlier 'could not confirm' note is resolved. These two open-weight variants (27B, 35B-A3B) remain the ones benchr's tools track."
    },
    {
      "id": "llama-4-scout",
      "name": "Llama 4 Scout",
      "provider": "Meta",
      "apiModelId": "meta-llama/Llama-4-Scout-17B-16E",
      "license": "Llama 4 Community License",
      "releaseDate": "2025-04-05",
      "params": {
        "total": "109B",
        "active": "17B",
        "experts": "16"
      },
      "context": {
        "windowTokens": 10000000,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {
        "MMLU-Pro (0-shot)": 74.3,
        "GPQA Diamond": 57.2,
        "MMMU": 73.4,
        "MathVista": 73.7,
        "LiveCodeBench": 32.8
      },
      "sources": {
        "release": "https://ai.meta.com/blog/llama-4-multimodal-intelligence/",
        "license": "https://www.llama.com/llama4/license/",
        "context": "https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct",
        "benchmarks": "https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E"
      },
      "verifiedDate": "2026-09-11",
      "notes": "Open weight; $0 to self-host (Meta sells no API). Community License: >700M MAU needs a separate Meta license; not OSI-approved. 10M-token context. Benchmarks from Meta's official instruction-tuned model card."
    },
    {
      "id": "llama-4-maverick",
      "name": "Llama 4 Maverick",
      "provider": "Meta",
      "apiModelId": "meta-llama/Llama-4-Maverick-17B-128E-Instruct",
      "license": "Llama 4 Community License",
      "releaseDate": "2025-04-05",
      "params": {
        "total": "400B",
        "active": "17B",
        "experts": "128"
      },
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {
        "MMLU-Pro (0-shot)": 80.5,
        "GPQA Diamond": 69.8,
        "LiveCodeBench": 43.4,
        "MGSM": 92.3
      },
      "sources": {
        "release": "https://ai.meta.com/blog/llama-4-multimodal-intelligence/",
        "license": "https://www.llama.com/llama4/license/",
        "context": "https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct",
        "benchmarks": "https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct"
      },
      "verifiedDate": "2026-09-11",
      "notes": "Open weight; $0 to self-host. 1M-token context, 128 experts. Llama 4 Behemoth (288B active / ~2T total) was only ever previewed as 'still training' and was never released — do not list specs for it as a usable model."
    },
    {
      "id": "phi-4",
      "name": "Phi-4",
      "provider": "Microsoft",
      "apiModelId": "microsoft/phi-4",
      "license": "MIT",
      "releaseDate": "2024-12-12",
      "params": {
        "total": "14B",
        "active": "14B",
        "type": "dense"
      },
      "context": {
        "windowTokens": 16000,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {
        "GPQA Diamond": 56.1,
        "MMLU": 84.8,
        "HumanEval": 82.6,
        "MATH": 80.4
      },
      "sources": {
        "release": "https://huggingface.co/microsoft/phi-4",
        "license": "https://huggingface.co/microsoft/phi-4",
        "context": "https://huggingface.co/microsoft/phi-4",
        "benchmarks": "https://www.microsoft.com/en-us/research/publication/phi-4-technical-report/"
      },
      "verifiedDate": "2026-09-11",
      "notes": "Verified 2026-06-03: MIT-licensed, 14B dense, 16K context, self-host only (no Microsoft per-token API; available on Azure AI Foundry + Hugging Face). Released 2024-12-12. Benchmarks from the Phi-4 technical report / model card."
    },
    {
      "id": "claude-fable-5",
      "name": "Claude Fable 5",
      "provider": "Anthropic",
      "apiModelId": "claude-fable-5",
      "license": "proprietary",
      "releaseDate": "2026-06-09",
      "status": "Active. Anthropic's current lifecycle table says Claude Fable 5 will not retire before June 9, 2027.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 10.0,
        "outputPerM": 50.0,
        "cachedInputPerM": 1.0,
        "note": "Prompt caching = 90% input discount per the launch announcement. Included free on Pro/Max/Team/seat-based Enterprise June 9-22, 2026; usage credits from June 23."
      },
      "tentativeRetirementFloor": "2027-06-09",
      "benchmarks": {
        "SWE-bench Pro": null,
        "SWE-bench Verified": null,
        "GPQA Diamond": null,
        "note": "No Fable-only score is recorded here. Anthropic's public launch graphic combines Fable 5 and Mythos 5 and reports the higher configuration per row, so its 80.3 SWE-bench Pro figure must not be represented as a Fable-only result."
      },
      "sources": {
        "release": "https://www.anthropic.com/news/claude-fable-5-mythos-5",
        "pricing": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "context": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "benchmarks": "https://www-cdn.anthropic.com/2f9323abbcc4abe219577539efe19a623c9ca2bd/Claude%20Fable%205%20%26%20Claude%20Mythos%205%20System%20Card.pdf",
        "status": "https://platform.claude.com/docs/en/about-claude/model-deprecations"
      },
      "verifiedDate": "2026-09-02",
      "notes": "Generally available from June 9, 2026. Anthropic describes Fable 5 and Mythos 5 as two configurations of the same underlying model: Fable is the general-use configuration with added safeguards in high-risk biology and cybersecurity domains, while Mythos is restricted to approved Project Glasswing partners. A U.S. export-control directive on June 12 required Anthropic to block access by foreign nationals; because nationality could not be verified per request, Anthropic suspended both models globally. The controls were lifted June 30 and Fable returned globally July 1. Fable requires 30-day data retention for safety monitoring. Combined Fable/Mythos benchmark graphics are not treated as Fable-only scores. Rechecked September 2, 2026: Fable 5 remains Active; Fable 5.1 is an additional API model rather than a retirement notice for this version."
    },
    {
      "id": "claude-fable-5-1",
      "name": "Claude Fable 5.1",
      "provider": "Anthropic",
      "apiModelId": "claude-fable-5-1",
      "license": "proprietary",
      "releaseDate": "2026-09-01",
      "status": "Generally available through the Claude API and Anthropic's supported cloud platforms. Adaptive thinking is always on.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 10.0,
        "outputPerM": 50.0,
        "cachedInputPerM": 0.25,
        "cacheWrite5mPerM": 12.5,
        "cacheWrite1hPerM": 20.0,
        "batchInputPerM": 5.0,
        "batchOutputPerM": 25.0,
        "note": "Provider-published standard API rates. Cache reads are $0.25 per 1M tokens; five-minute and one-hour cache writes are $12.50 and $20.00 per 1M tokens. Batch halves standard input and output pricing."
      },
      "tentativeRetirementFloor": "2027-09-01",
      "benchmarks": {
        "Terminal-Bench 4.0": 55.8,
        "Terminal-Bench-Science 0.1": 52.6,
        "OSWorld 2.0 (partial success)": 77.9,
        "OSWorld 2.0 (strict success)": 41.7,
        "Humanity's Last Exam (no tools)": 60.9,
        "Humanity's Last Exam (with tools)": 65.0,
        "note": "All populated results are provider-reported in Anthropic's September 1 announcement, not benchr measurements. The announcement describes its own evaluation configuration and safeguards; do not compare these figures to differently configured third-party runs as if they were interchangeable."
      },
      "sources": {
        "release": "https://www.anthropic.com/claude-fable-and-mythos-5-1",
        "api": "https://platform.claude.com/docs/en/models/fable-5-1/overview",
        "pricing": "https://platform.claude.com/docs/en/models/fable-5-1/overview",
        "context": "https://platform.claude.com/docs/en/models/fable-5-1/overview",
        "benchmarks": "https://www.anthropic.com/claude-fable-and-mythos-5-1",
        "status": "https://platform.claude.com/docs/en/about-claude/model-deprecations"
      },
      "verifiedDate": "2026-09-02",
      "notes": "Released September 1, 2026. Fable 5.1 is a new API ID, not an alias for Fable 5. Anthropic documents that tool_choice type any and named-tool forcing return HTTP 400; auto and none are valid. Manual or disabled thinking configuration is not supported because adaptive thinking is always on. Assistant prefills are unsupported. Fable 5.1 is not supported on Priority Tier. Anthropic states that 30-day data retention normally applies; zero-data-retention access requires authorization."
    },
    {
      "id": "claude-mythos-5",
      "name": "Claude Mythos 5",
      "provider": "Anthropic",
      "apiModelId": "claude-mythos-5",
      "license": "proprietary",
      "releaseDate": "2026-06-09",
      "status": "Restricted — Project Glasswing / trusted access only; not generally available.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 10.0,
        "outputPerM": 50.0,
        "cachedInputPerM": 1.0,
        "note": "Same listed pricing as Fable 5 (incl. $1 cache-hit input); access is restricted to approved Glasswing organizations."
      },
      "benchmarks": {
        "note": "Shares the launch benchmark table with Fable 5 (scores shown are the higher of the two, within 1-3 points). Anthropic calls its cybersecurity capabilities the strongest of any model in the world."
      },
      "sources": {
        "release": "https://www.anthropic.com/news/claude-fable-5-mythos-5",
        "pricing": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "context": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "status": "https://www.anthropic.com/news/redeploying-fable-5"
      },
      "verifiedDate": "2026-09-02",
      "notes": "Successor to Claude Mythos Preview inside Project Glasswing. Anthropic says Mythos 5 shares the underlying model with Fable 5 but has fewer safeguards and remains limited to a small group of trusted partners; there is no self-serve access. The June 12 export-control directive caused a global suspension. Anthropic says the U.S. government approved restoration to a set of U.S. organizations on June 26 and access was restored July 1, while expansion to additional Glasswing partners remained in progress. Listed pricing and context are provider figures for approved access, not evidence of general availability. Rechecked September 2, 2026: Mythos 5 remains a restricted model; Mythos 5.1 is a separate restricted API ID."
    },
    {
      "id": "claude-mythos-5-1",
      "name": "Claude Mythos 5.1",
      "provider": "Anthropic",
      "apiModelId": "claude-mythos-5-1",
      "license": "proprietary",
      "releaseDate": "2026-09-01",
      "status": "Restricted to vetted Project Glasswing / trusted access. It is not self-service or generally available.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 10.0,
        "outputPerM": 50.0,
        "cachedInputPerM": 0.25,
        "cacheWrite5mPerM": 12.5,
        "cacheWrite1hPerM": 20.0,
        "batchInputPerM": 5.0,
        "batchOutputPerM": 25.0,
        "note": "Anthropic lists the same published API price card and cache schedule as Fable 5.1, but access is restricted. A listed price is not evidence of public availability."
      },
      "benchmarks": {
        "note": "Anthropic's launch announcement reports Mythos 5.1 results alongside Fable 5.1. Because access is restricted and the release tables distinguish configurations, this ledger does not infer cross-model or public-availability conclusions from them."
      },
      "sources": {
        "release": "https://www.anthropic.com/claude-fable-and-mythos-5-1",
        "api": "https://platform.claude.com/docs/en/models/fable-5-1/overview",
        "pricing": "https://platform.claude.com/docs/en/models/fable-5-1/overview",
        "context": "https://platform.claude.com/docs/en/models/fable-5-1/overview",
        "status": "https://www.anthropic.com/claude/mythos"
      },
      "verifiedDate": "2026-09-02",
      "notes": "Released September 1, 2026. Anthropic says Mythos 5.1 has the same headline context, output limit, price card, and adaptive-thinking model surface as Fable 5.1, while keeping access restricted to vetted users. It should not be presented as a public Claude tier or an available self-service fallback."
    },
    {
      "id": "gpt-5-4",
      "name": "GPT-5.4",
      "provider": "OpenAI",
      "apiModelId": "gpt-5.4",
      "license": "proprietary",
      "releaseDate": "2026-03-05",
      "context": {
        "windowTokens": 1050000,
        "maxOutputTokens": 128000,
        "note": "The current OpenAI API model page lists a 1,050,000-token context window and a 128,000-token maximum output. Inputs beyond the standard 272K window carry the long-context surcharge described in OpenAI's launch material."
      },
      "pricing": {
        "inputPerM": 2.5,
        "outputPerM": 15.0,
        "cachedInputPerM": 0.25,
        "note": "Long-context surcharge above 272K input per the official pricing page."
      },
      "benchmarks": {
        "GDPval (wins or ties)": 83.0,
        "SWE-Bench Pro (Public)": 57.7,
        "OSWorld-Verified": 75.0,
        "BrowseComp": 82.7,
        "Toolathlon": 54.6,
        "SWE-bench Verified": null,
        "note": "All populated figures are from OpenAI's GPT-5.4 launch evaluation table. OpenAI publishes SWE-Bench Pro (Public), not a SWE-bench Verified figure, for GPT-5.4."
      },
      "sources": {
        "release": "https://openai.com/index/introducing-gpt-5-4/",
        "pricing": "https://developers.openai.com/api/docs/models/gpt-5.4",
        "context": "https://developers.openai.com/api/docs/models/gpt-5.4",
        "benchmarks": "https://openai.com/index/introducing-gpt-5-4/"
      },
      "verifiedDate": "2026-07-29",
      "notes": "Released March 5, 2026 as GPT-5.4 Thinking + GPT-5.4 Pro ('most capable and efficient frontier model for professional work'); GPT-5.4 mini and nano followed March 17 (mini available to free tier, nano API-only). Built-in computer use; tuned for finance workflows (launched alongside ChatGPT for Excel, which it powered at launch). benchr removed an earlier gpt-5-4 entry on June 1, 2026 as 'unverified' — that removal was wrong; the model is real and was re-added after this June 10 verification. Superseded as OpenAI's flagship by GPT-5.5 (April 23, 2026)."
    },
    {
      "id": "claude-opus-4-6",
      "name": "Claude Opus 4.6",
      "provider": "Anthropic",
      "apiModelId": "claude-opus-4-6",
      "license": "proprietary",
      "releaseDate": "2026-02-05",
      "status": "Active (official lifecycle state on the Anthropic deprecations table, June 12, 2026); superseded as the newest Opus by 4.7 (April 2026) and 4.8 (May 2026) but fully supported.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 5.0,
        "outputPerM": 25.0,
        "fastModeInputPerM": 30.0,
        "fastModeOutputPerM": 150.0,
        "batchInputPerM": 2.5,
        "batchOutputPerM": 12.5,
        "cachedInputPerM": 0.5
      },
      "tentativeRetirementFloor": "2027-02-05",
      "benchmarks": {
        "note": "Launch benchmarks not re-verified for this entry; cite from the official announcement if needed."
      },
      "sources": {
        "release": "https://www.anthropic.com/news/claude-opus-4-6",
        "pricing": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "context": "https://platform.claude.com/docs/en/about-claude/models/overview"
      },
      "verifiedDate": "2026-06-12",
      "notes": "Released February 5, 2026. First Opus with the 1M-token context window (beta at launch); extended thinking AND adaptive thinking. The official deprecations table (June 12, 2026) lists it ACTIVE at $5/$25, 1M context, 128K max output, with a tentative retirement floor of February 5, 2027; fast mode ($30/$150, same rate as Opus 4.7) and $0.50 cache-hit input are on the official pricing page. An earlier benchr audit wrongly flagged 'Opus 4.6' as a non-existent model — it is real; the Opus lane runs 4.6 (Feb) → 4.7 (Apr) → 4.8 (May)."
    },
    {
      "id": "grok-4-20",
      "name": "Grok 4.20",
      "provider": "xAI",
      "apiModelId": "grok-4.20",
      "license": "proprietary",
      "releaseDate": "2026-03-10",
      "status": "Superseded as xAI's flagship by Grok 4.3 (April 2026); still listed in the xAI docs.",
      "context": {
        "windowTokens": 1000000
      },
      "pricing": {
        "inputPerM": 1.25,
        "outputPerM": 2.5,
        "cachedInputPerM": 0.2,
        "longContextThresholdTokens": 200000,
        "longContextInputPerM": 2.5,
        "longContextCachedInputPerM": 0.4,
        "longContextOutputPerM": 5.0,
        "batchInputPerM": 1.0,
        "batchCachedInputPerM": 0.16,
        "batchOutputPerM": 2.0,
        "batchLongContextInputPerM": 2.0,
        "batchLongContextCachedInputPerM": 0.32,
        "batchLongContextOutputPerM": 4.0,
        "note": "Current xAI rates. Long-context pricing applies to the whole request at 200K prompt tokens; Batch API is 20% below the applicable standard tier."
      },
      "benchmarks": {
        "note": "Not re-verified for this entry."
      },
      "sources": {
        "release": "https://docs.x.ai/developers/models/grok-4.20",
        "pricing": "https://docs.x.ai/developers/models/grok-4.20",
        "context": "https://docs.x.ai/developers/models/grok-4.20"
      },
      "verifiedDate": "2026-07-30",
      "notes": "Public beta February 17, 2026; second beta March 3; API release March 10. Current docs identify grok-4.20-0309-reasoning as the snapshot and grok-4.20 as an alias, with 1M context and short-context rates of $1.25 input / $0.20 cached / $2.50 output. At a prompt of 200K tokens or more, the whole request uses $2.50 / $0.40 / $5.00; Batch is 20% lower. Launch-era $2/$6 and 2M figures are historical, not current."
    },
    {
      "id": "qwen-3-5-397b-a17b",
      "name": "Qwen3.5-397B-A17B",
      "provider": "Alibaba (Qwen team)",
      "apiModelId": null,
      "license": null,
      "releaseDate": "2026-02-16",
      "status": "Previous generation — superseded by the Qwen3.6 series (April 2026).",
      "context": {
        "windowTokens": 262144,
        "extendedWindowTokens": 1010000
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null,
        "note": "Hosted-API pricing not confirmed on an official Alibaba page; third-party claims of $0.10/M input are unverified."
      },
      "benchmarks": {
        "note": "Not re-verified for this entry."
      },
      "sources": {
        "release": "https://github.com/QwenLM/Qwen3.6",
        "context": "https://huggingface.co/Qwen"
      },
      "verifiedDate": "2026-06-10",
      "notes": "Qwen3.5 flagship: 397B total / 17B active MoE (512 experts, 10 routed + 1 shared), hybrid Gated Delta Network + MoE architecture, native 262,144-token context extensible to ~1,010,000. The Qwen3.5 family spans nine sizes (Small 0.8B-9B, Medium 27B/35B-A3B/122B-A10B, flagship 397B-A17B); most open-weight under Apache 2.0 — the flagship's exact license was not re-verified, so it is left null. Qwen3.6 (April 2026) is the current series and what benchr's tools track."
    },
    {
      "id": "kimi-k2-7-code",
      "name": "Kimi K2.7-Code",
      "provider": "Moonshot AI",
      "apiModelId": "kimi-k2.7-code",
      "license": "Modified MIT",
      "releaseDate": null,
      "params": {
        "total": "1T",
        "active": "32B"
      },
      "context": {
        "windowTokens": 262144,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 0.95,
        "outputPerM": 4.0,
        "cacheHitInputPerM": 0.19,
        "selfHost": true
      },
      "benchmarks": {
        "Kimi Code Bench v2": 62.0,
        "Program Bench": 53.6,
        "MLS Bench Lite": 35.1,
        "Kimi Claw 24/7 Bench": 46.9,
        "MCP Atlas": 76.0,
        "MCP Mark Verified": 81.1
      },
      "sources": {
        "pricing": "https://platform.kimi.ai/docs/pricing/chat-k27-code",
        "license": "https://huggingface.co/moonshotai/Kimi-K2.7-Code",
        "context": "https://platform.kimi.ai/docs/models",
        "benchmarks": "https://huggingface.co/moonshotai/Kimi-K2.7-Code"
      },
      "verifiedDate": "2026-06-23",
      "notes": "Coding-specialist model built on Kimi K2.6 (Moonshot says ~30% fewer thinking tokens). Open weights, Modified MIT, 1T total / 32B active MoE, 256K (262,144) context confirmed on platform.kimi.ai/docs and the official Hugging Face model card. Input $0.95 is cache-miss; cache-hit input $0.19; output $4.00 per 1M. No official release date on any Moonshot page (the HF card shows only a relative 'updated' timestamp), so releaseDate is null. A higher-throughput hosted variant — Kimi K2.7-Code-Highspeed (~180 tok/s, $1.90/$0.38/$8.00) — is a separate entry. Benchmarks intentionally empty pending official model-card figures verified this session."
    },
    {
      "id": "kimi-k2-7-code-highspeed",
      "name": "Kimi K2.7-Code-Highspeed",
      "provider": "Moonshot AI",
      "apiModelId": "kimi-k2.7-code-highspeed",
      "license": "Modified MIT",
      "releaseDate": null,
      "params": {
        "total": "1T",
        "active": "32B"
      },
      "context": {
        "windowTokens": 262144,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 1.9,
        "outputPerM": 8.0,
        "cacheHitInputPerM": 0.38,
        "selfHost": true,
        "note": "High-throughput HOSTED serving tier of Kimi K2.7-Code (~180 tok/s output, up to ~260 tok/s short-context) at double the base rates. The same open weights underlie it; the premium buys faster Moonshot-hosted inference, not a different model."
      },
      "benchmarks": {},
      "sources": {
        "pricing": "https://platform.kimi.ai/docs/pricing/chat-k27-code",
        "context": "https://platform.kimi.ai/docs/models"
      },
      "verifiedDate": "2026-06-23",
      "notes": "Hosted high-speed variant of Kimi K2.7-Code, confirmed on platform.kimi.ai/docs (models + chat-k27-code pricing) June 23, 2026: cache-miss input $1.90, cache-hit $0.38, output $8.00 per 1M — exactly 2x the base K2.7-Code rates. Same 256K (262,144) context and 1T/32B architecture."
    },
    {
      "id": "qwen-3-6-plus",
      "name": "Qwen3.6-Plus",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "qwen3.6-plus",
      "license": "proprietary",
      "releaseDate": "2026-04-02",
      "status": "Hosted-only flagship on Alibaba Cloud Model Studio / DashScope. No open weights.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 0.5,
        "inputAbove256kPerM": 2.0,
        "outputPerM": 3.0,
        "outputAbove256kPerM": 6.0,
        "note": "Official Alibaba Cloud Model Studio per-token pricing (Singapore/international USD), verified 2026-06-28: input $0.5 (0-256K) / $2.0 (256K-1M), output $3.0 / $6.0. Resolves the prior 'could not confirm a per-token price' gap. Hosted-only, no open weights."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://www.alibabacloud.com/blog/qwen3-6-plus-towards-real-world-agents_603005",
        "context": "https://www.alibabacloud.com/help/en/model-studio/text-generation-model",
        "pricing": "https://www.alibabacloud.com/help/en/model-studio/model-pricing"
      },
      "verifiedDate": "2026-06-28",
      "notes": "Hosted-only agent flagship on Model Studio, 1M default context, released April 2, 2026. RESOLVES the earlier benchr gap that 'Qwen3.6-Plus could not be confirmed' — it is real and official, it was simply hosted-only and never on the open-weight repo. Superseded within the series by Qwen3.7-Plus (June 3, 2026). UPDATE 2026-06-28: per-token pricing IS now published on Model Studio (tiered rates recorded above)."
    },
    {
      "id": "qwen-3-6-max-preview",
      "name": "Qwen3.6-Max-Preview",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "qwen3.6-max-preview",
      "license": "proprietary",
      "announcementDate": "2026-04-22",
      "releaseDate": null,
      "status": "Currently listed in Alibaba Cloud Model Studio as a hosted preview. Alibaba's April 22 post contains conflicting 'coming soon' and 'available' wording, so the exact first API-availability date is not claimed. No open weights.",
      "context": {
        "windowTokens": 262144,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 1.3,
        "outputPerM": 7.8,
        "tieredInputPerM": {
          "0-128K": 1.3,
          "128K-256K": 2.0
        },
        "tieredOutputPerM": {
          "0-128K": 7.8,
          "128K-256K": 12.0
        },
        "note": "Tiered by prompt length on the Model Studio pricing table, Singapore region, read 2026-09-08: $1.30 input / $7.80 output up to 128K, $2.00 / $12.00 from 128K to 256K. The headline fields carry the first tier."
      },
      "benchmarks": {},
      "sources": {
        "announcement": "https://www.alibabacloud.com/blog/qwen3-6-max-preview-smarter-sharper-still-evolving_603055",
        "availability": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "context": "https://www.alibabacloud.com/blog/qwen3-6-max-preview-smarter-sharper-still-evolving_603055",
        "pricing": "https://www.alibabacloud.com/help/en/model-studio/model-pricing"
      },
      "verifiedDate": "2026-09-08",
      "notes": "Alibaba announced the hosted Max preview on April 22, 2026. The same official post says both 'coming soon' and 'available through the API', so April 22 is retained only as an announcement date. Current Model Studio pricing documentation lists qwen3.6-max-preview, confirming present availability. The announcement's evaluation notes use a 256K context. Proprietary, no open weights."
    },
    {
      "id": "qwen-3-7-plus",
      "name": "Qwen3.7-Plus",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "qwen3.7-plus",
      "license": "proprietary",
      "releaseDate": "2026-06-03",
      "status": "Hosted-only multimodal flagship on Alibaba Cloud Model Studio / DashScope. No open weights.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 65536,
        "note": "Alibaba's Qwen3.7-Plus launch page publishes a Model Studio configuration example with a 1,000,000-token context window and 65,536 max tokens. The Model Studio pricing page independently exposes a 256K-1M input tier."
      },
      "pricing": {
        "inputPerM": 0.4,
        "inputAbove256kPerM": 1.2,
        "outputPerM": 1.6,
        "outputAbove256kPerM": 4.8,
        "note": "Official Alibaba Cloud Model Studio per-token pricing (Singapore/international USD), verified 2026-06-28: input $0.4 (0-256K) / $1.2 (256K-1M), output $1.6 / $4.8. Resolves the prior 'not officially published' gap. Hosted-only (DashScope/Model Studio), no open weights. A promotional discount was running on Model Studio at verification time; the list rates are recorded here."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://www.alibabacloud.com/blog/qwen3-7-plus-multimodal-agent-intelligence_603206",
        "context": "https://www.alibabacloud.com/blog/qwen3-7-plus-multimodal-agent-intelligence_603206",
        "pricing": "https://www.alibabacloud.com/help/en/model-studio/model-pricing"
      },
      "verifiedDate": "2026-06-28",
      "notes": "Multimodal (text + image/video input) hosted agent flagship, API-only via Model Studio/DashScope, no open weights. Released June 3, 2026; Alibaba describes it as a comprehensive upgrade over Qwen3.6-Plus. UPDATE 2026-06-28: per-token pricing IS now published on Model Studio (tiered rates recorded above); the prior 'not published' note is resolved. The open-weight Qwen line benchr's tools track is separate (Qwen3.6-27B / 35B-A3B)."
    },
    {
      "id": "gpt-5-6-sol",
      "name": "GPT-5.6 Sol",
      "provider": "OpenAI",
      "apiModelId": "gpt-5.6-sol",
      "license": "proprietary",
      "releaseDate": "2026-07-09",
      "status": "GENERAL AVAILABILITY as of 2026-07-09. OpenAI's official API changelog announces GPT-5.6 Sol/Terra/Luna on July 9, and the live API models/pricing pages list the model IDs, prices, 1.05M context window, and 128K max output. This supersedes the July 8 preview-only correction.",
      "context": {
        "windowTokens": 1050000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 4.0,
        "outputPerM": 20.0,
        "cachedInputPerM": 0.4,
        "cacheWritePerM": 5.0,
        "longContextInputPerM": 8.0,
        "longContextOutputPerM": 30.0,
        "longContextCachedInputPerM": 0.8,
        "longContextCacheWritePerM": 10.0,
        "batchInputPerM": 2.0,
        "batchOutputPerM": 10.0,
        "batchCachedInputPerM": 0.2,
        "flexInputPerM": 2.0,
        "flexOutputPerM": 10.0,
        "fastInputPerM": 8.0,
        "fastOutputPerM": 40.0,
        "priorityInputPerM": null,
        "priorityOutputPerM": null,
        "priceStatus": "Promotional pricing effective 2026-08-21, read live on developers.openai.com/api/docs/pricing on 2026-08-24.",
        "note": "OpenAI cut Sol's standard rate on August 21, 2026: input $5.00 -> $4.00 (20% lower), output $30.00 -> $20.00 (33% lower), cached input $0.50 -> $0.40. Short-context cache writes are $5.00. The long-context tier (above 272K input tokens) is $8.00 input / $0.80 cached / $10.00 cache write / $30.00 output. Batch and Flex are half of standard ($2.00 / $10.00); Flex lists no long-context row. Fast mode replaced Priority processing and bills at 2x standard ($8.00 / $40.00 short context, $16.00 / $60.00 long context), so the priority fields are null rather than carrying stale numbers. OpenAI's changelog states the promotional rate is available at least through November 21, 2026."
      },
      "benchmarks": {
        "Terminal-Bench 2.1": 88.8,
        "Terminal-Bench 2.1 (ultra mode)": 91.9,
        "SWE-bench Verified": 89.8,
        "GPQA Diamond": 91.2,
        "note": "Terminal-Bench 2.1 figures are from the June 26 preview system card. SWE-bench Verified (89.8%) and GPQA Diamond (91.2%) were added to the preview system-card page on 2026-07-01; GA followed on 2026-07-09. All benchmark figures here are provider-reported."
      },
      "sources": {
        "release": "https://developers.openai.com/api/docs/changelog",
        "pricing": "https://developers.openai.com/api/docs/pricing",
        "benchmarks": "https://deploymentsafety.openai.com/gpt-5-6-preview",
        "context": "https://developers.openai.com/api/docs/models"
      },
      "verifiedDate": "2026-08-24",
      "notes": "UPDATED 2026-07-13: GPT-5.6 is now officially generally available, per OpenAI's July 9 API changelog and live models/pricing pages. The prior July 8 correction was accurate at that time but is now superseded. OpenAI's live models page lists a 1.05M context window and 128K max output for this tier; pricing is listed on the official pricing page. Re-verified 2026-08-24 after OpenAI's August 21 price cut; the $5/$30 rate recorded on 2026-07-13 is now historical and is logged in history.json."
    },
    {
      "id": "gpt-5-6-terra",
      "name": "GPT-5.6 Terra",
      "provider": "OpenAI",
      "apiModelId": "gpt-5.6-terra",
      "license": "proprietary",
      "releaseDate": "2026-07-09",
      "status": "GENERAL AVAILABILITY as of 2026-07-09. OpenAI's official API changelog announces GPT-5.6 Sol/Terra/Luna on July 9, and the live API models/pricing pages list the model IDs, prices, 1.05M context window, and 128K max output. This supersedes the July 8 preview-only correction.",
      "context": {
        "windowTokens": 1050000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 2.0,
        "outputPerM": 12.0,
        "cachedInputPerM": 0.2,
        "longContextInputPerM": 4.0,
        "longContextOutputPerM": 18.0,
        "longContextCachedInputPerM": 0.4,
        "priceStatus": "Price cut effective 2026-07-30, confirmed live on developers.openai.com/api/docs/pricing 2026-08-08.",
        "note": "OpenAI cut Terra's standard price 20% from $2.50/$15 to $2.00/$12 per 1M tokens effective July 30, 2026 (cached input $0.20, down from $0.25); long-context tier (above the short-context threshold) is $4.00/$18.00, cached $0.40. Batch/flex pricing was not restated after the cut; treat the prior 50%-of-standard batch note as stale until re-confirmed.",
        "priorityInputPerM": 5.0,
        "priorityOutputPerM": 30.0
      },
      "benchmarks": {
        "Terminal-Bench 2.1": 84.3,
        "SWE-bench Verified": 85.2,
        "GPQA Diamond": 88.0,
        "note": "Terminal-Bench 2.1 unchanged from preview: Terra 84.3% (ties Claude Fable 5; just above GPT-5.5 83.4%). SWE-bench Verified (85.2%) and GPQA Diamond (88.0%) added at GA. OpenAI frames Terra as GPT-5.5-class quality at roughly half the cost — the GA figures support that pitch."
      },
      "sources": {
        "release": "https://developers.openai.com/api/docs/changelog",
        "pricing": "https://developers.openai.com/api/docs/pricing",
        "benchmarks": "https://deploymentsafety.openai.com/gpt-5-6-preview",
        "context": "https://developers.openai.com/api/docs/models"
      },
      "verifiedDate": "2026-08-08",
      "notes": "UPDATED 2026-08-08: OpenAI cut Terra's API price 20% effective July 30, 2026 ($2.50/$15 to $2.00/$12), part of the same pricing update that cut Luna 80%. Confirmed on the live official pricing docs. Context/output limits unchanged at 1.05M/128K."
    },
    {
      "id": "gpt-5-6-luna",
      "name": "GPT-5.6 Luna",
      "provider": "OpenAI",
      "apiModelId": "gpt-5.6-luna",
      "license": "proprietary",
      "releaseDate": "2026-07-09",
      "status": "GENERAL AVAILABILITY as of 2026-07-09. OpenAI's official API changelog announces GPT-5.6 Sol/Terra/Luna on July 9, and the live API models/pricing pages list the model IDs, prices, 1.05M context window, and 128K max output. This supersedes the July 8 preview-only correction.",
      "context": {
        "windowTokens": 1050000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 0.2,
        "outputPerM": 1.2,
        "cachedInputPerM": 0.02,
        "longContextInputPerM": 0.4,
        "longContextOutputPerM": 1.8,
        "longContextCachedInputPerM": 0.04,
        "priceStatus": "Price cut effective 2026-07-30, confirmed live on developers.openai.com/api/docs/pricing 2026-08-08.",
        "note": "OpenAI cut Luna's standard price 80% from $1.00/$6.00 to $0.20/$1.20 per 1M tokens effective July 30, 2026 (cached input $0.02, down from $0.10); long-context tier (above the short-context threshold) is $0.40/$1.80, cached $0.04. Batch/flex pricing was not restated after the cut; treat the prior 50%-of-standard batch note as stale until re-confirmed.",
        "priorityInputPerM": 2.0,
        "priorityOutputPerM": 12.0
      },
      "benchmarks": {
        "Terminal-Bench 2.1": 82.5,
        "SWE-bench Verified": 79.8,
        "GPQA Diamond": 82.0,
        "note": "Terminal-Bench 2.1 unchanged from preview: Luna 82.5% (below GPT-5.5 83.4%, above Claude Opus 4.8 78.9%). SWE-bench Verified (79.8%) and GPQA Diamond (82.0%) added at GA. Luna is the cheapest, fastest tier and has the same 1.05M context window as Sol and Terra."
      },
      "sources": {
        "release": "https://developers.openai.com/api/docs/changelog",
        "pricing": "https://developers.openai.com/api/docs/pricing",
        "benchmarks": "https://deploymentsafety.openai.com/gpt-5-6-preview",
        "context": "https://developers.openai.com/api/docs/models"
      },
      "verifiedDate": "2026-08-08",
      "notes": "UPDATED 2026-08-08: OpenAI cut Luna's API price 80% effective July 30, 2026 ($1.00/$6.00 to $0.20/$1.20), the same pricing update that cut Terra 20%. Confirmed on the live official pricing docs. Context/output limits unchanged at 1.05M/128K."
    },
    {
      "id": "mistral-ocr-4",
      "name": "Mistral OCR 4",
      "provider": "Mistral AI",
      "apiModelId": "mistral-ocr-4-0",
      "license": "proprietary",
      "releaseDate": "2026-06-23",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "perThousandPagesStandard": 4.0,
        "perThousandPagesBatch": 2.0,
        "perThousandPagesDocumentAI": 5.0,
        "inputPerM": null,
        "outputPerM": null,
        "note": "Priced PER PAGE, not per token: $4 per 1,000 pages (standard API), $2 per 1,000 pages (Batch API, 50% off), $5 per 1,000 pages (Document AI / annotated). Verified 2026-06-28 on mistral.ai/pricing and the official news post. Token fields are null because this is a document-OCR model, not a per-token chat LLM. Alias mistral-ocr-latest points to mistral-ocr-4-0."
      },
      "benchmarks": {
        "OlmOCRBench": 85.2,
        "OmniDocBench": 93.07,
        "note": "Mistral's own announcement figures; also claims a 72% average human-preference win rate vs tested competitors."
      },
      "sources": {
        "release": "https://mistral.ai/news/ocr-4/",
        "pricing": "https://mistral.ai/pricing",
        "context": "https://docs.mistral.ai/models/model-cards/ocr-4-0",
        "benchmarks": "https://mistral.ai/news/ocr-4/"
      },
      "verifiedDate": "2026-06-28",
      "notes": "State-of-the-art document-intelligence / OCR model released 2026-06-23: paragraph-level bounding boxes, typed-block classification (titles, tables, equations, signatures), inline confidence scores; PDF/DOC/PPT/OpenDocument across 170 languages. GA via Mistral Studio, Amazon SageMaker, and Microsoft Foundry; single-container self-hosting for enterprise data-sovereignty. A different category from benchr's per-token chat set (Large 3, Medium 3.5) — kept in the verified record but out of the chat-comparison tool."
    },
    {
      "id": "mistral-small-4",
      "name": "Mistral Small 4",
      "provider": "Mistral AI",
      "apiModelId": "mistral-small-4",
      "license": "Apache-2.0",
      "releaseDate": "2026-03-16",
      "params": {
        "total": "119B",
        "active": "6B",
        "experts": "128 (4 active/token)"
      },
      "context": {
        "windowTokens": 256000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 0.15,
        "outputPerM": 0.6,
        "cacheHitInputPerM": null,
        "selfHost": true
      },
      "benchmarks": {},
      "sources": {
        "release": "https://mistral.ai/news/mistral-small-4/",
        "pricing": "https://mistral.ai/pricing",
        "license": "https://huggingface.co/mistralai/Mistral-Small-4-119B-2603",
        "context": "https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03"
      },
      "verifiedDate": "2026-06-28",
      "notes": "Open-weight (Apache-2.0) hybrid MoE — 119B total / 6B active (128 experts, 4 active/token), 256K context, text + image input, with a configurable reasoning-effort toggle. Hosted at $0.15/$0.60 per 1M on mistral.ai/pricing (verified 2026-06-28) and self-hostable. Released 2026-03-16 (a pre-existing gap in benchr's Mistral set; added 2026-06-28). No discrete official per-benchmark scores recorded this session, so benchmarks is intentionally empty."
    },
    {
      "id": "qwen-3-7-max",
      "name": "Qwen3.7-Max",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "qwen3.7-max",
      "license": "proprietary",
      "announcementDate": "2026-05-21",
      "releaseDate": null,
      "status": "Currently available through Alibaba Cloud Model Studio / DashScope. Alibaba's May 21 announcement said the API was still coming soon, so the exact first-availability date is not claimed. No open weights.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 65536,
        "note": "Alibaba's Qwen3.7-Max launch page publishes a Model Studio configuration example with a 1,000,000-token context window and 65,536 max tokens. Evaluation notes elsewhere on the page use smaller run-specific contexts; those settings do not replace the published model limit."
      },
      "pricing": {
        "inputPerM": 2.5,
        "outputPerM": 7.5,
        "note": "Official Singapore/international list pricing is $2.50/1M input and $7.50/1M output. Alibaba currently labels a limited-time 50% discount, but promotions are excluded from benchr's baseline because the documentation does not provide a durable end date. Verified 2026-07-30."
      },
      "benchmarks": {
        "SWE-bench Verified": 80.4,
        "SWE-bench Pro": 60.6,
        "SWE-bench Multilingual": 78.3,
        "Terminal-Bench 2.0 (Terminus)": 69.7,
        "MCP-Mark": 60.8,
        "MCP-Atlas": 76.4,
        "BFCL V4": 75.0
      },
      "benchmarkDetails": {
        "sourceType": "provider-reported",
        "evaluationSettings": "Alibaba reports an internal agent scaffold for SWE-bench, a Harbor/Terminus-2 harness with five runs for Terminal-Bench 2.0, and named public/internal agent benchmarks. These are not independent benchr test results.",
        "sourceDate": "2026-05-21"
      },
      "sources": {
        "announcement": "https://www.alibabacloud.com/blog/qwen3-7-the-agent-frontier_603154",
        "availability": "https://www.alibabacloud.com/help/en/model-studio/models",
        "context": "https://www.alibabacloud.com/blog/qwen3-7-the-agent-frontier_603154",
        "pricing": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "benchmarks": "https://www.alibabacloud.com/blog/qwen3-7-the-agent-frontier_603154"
      },
      "verifiedDate": "2026-07-30",
      "notes": "Alibaba publicly introduced Qwen3.7-Max on May 21, 2026, while explicitly saying Model Studio API access was 'coming soon'. Current Model Studio documentation lists the live qwen3.7-max API across six regions and maps the rolling alias to qwen3.7-max-2026-05-20. A snapshot identifier is not treated as proof of the public availability date, so releaseDate remains null until Alibaba documents that date."
    },
    {
      "id": "grok-build-0-1",
      "name": "Grok Build 0.1",
      "provider": "xAI",
      "apiModelId": "grok-build-0.1",
      "license": "proprietary",
      "releaseDate": "2026-05-29",
      "status": "Agentic coding model in public beta via the xAI API.",
      "context": {
        "windowTokens": 256000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 1.0,
        "outputPerM": 2.0,
        "cachedInputPerM": 0.2,
        "longContextThresholdTokens": 200000,
        "longContextInputPerM": 2.0,
        "longContextCachedInputPerM": 0.4,
        "longContextOutputPerM": 4.0,
        "note": "Current xAI rates: $1.00 input / $0.20 cached / $2.00 output below 200K prompt tokens; $2.00 / $0.40 / $4.00 for the whole request at or above 200K. The Batch API currently provides no discount for this model."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://x.ai/news/grok-build-0-1",
        "pricing": "https://docs.x.ai/developers/pricing",
        "context": "https://docs.x.ai/developers/models/grok-build-0.1"
      },
      "verifiedDate": "2026-07-30",
      "notes": "xAI's agentic coding model reached the API in public beta on May 29, 2026. The official model and pricing pages list a 256K context window, $1.00 input / $0.20 cached / $2.00 output below 200K prompt tokens, and $2.00 / $0.40 / $4.00 at or above 200K. xAI publishes a provider speed claim in the launch post, but benchr does not treat that as a universal latency measurement or store it as a benchmark."
    },
    {
      "id": "qwen-agentworld-35b-a3b",
      "name": "Qwen-AgentWorld-35B-A3B",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "Qwen/Qwen-AgentWorld-35B-A3B",
      "license": "Apache-2.0",
      "releaseDate": "2026-06-24",
      "params": {
        "total": "35B",
        "active": "3B"
      },
      "context": {
        "windowTokens": 262144,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null,
        "note": "Open weights, download-only (Hugging Face / ModelScope). Not offered as a hosted/priced API on Alibaba Cloud Model Studio, so there is no per-token list price."
      },
      "benchmarks": {
        "note": "Evaluated on the team's own AgentWorldBench (agent-environment-simulation). No independent third-party score yet."
      },
      "sources": {
        "release": "https://github.com/QwenLM/Qwen-AgentWorld",
        "license": "https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B",
        "context": "https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B",
        "benchmarks": "https://github.com/QwenLM/Qwen-AgentWorld"
      },
      "verifiedDate": "2026-06-28",
      "notes": "Open-weight (Apache-2.0) MoE 'language world model' for agentic environment simulation across 7 domains (MCP, Search, Terminal, SWE, Android, Web, OS), released 2026-06-24 on the official QwenLM GitHub. 35B total / 3B active, 256K context, built on Qwen3.5-35B-A3B-Base. A specialized agent-simulation model, NOT a general-purpose chat LLM — kept in the verified record but out of the chat-comparison tool. The open-weight general Qwen line benchr's tools track is separate (Qwen3.6-27B / 35B-A3B)."
    },
    {
      "id": "claude-sonnet-5",
      "name": "Claude Sonnet 5",
      "provider": "Anthropic",
      "apiModelId": "claude-sonnet-5",
      "license": "proprietary",
      "releaseDate": "2026-07-01",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 2.0,
        "outputPerM": 10.0,
        "cachedInputPerM": 0.2,
        "batchInputPerM": 1.0,
        "batchOutputPerM": 5.0,
        "note": "Anthropic's pricing page, read 2026-08-31: $2/$10 per 1M (cache-hit $0.20) is the standard price. The increase to $3/$15 that had been scheduled for September 1, 2026 was cancelled."
      },
      "tentativeRetirementFloor": "2027-07-01",
      "benchmarks": {
        "SWE-bench Verified": 89.4,
        "SWE-bench Pro": 71.8,
        "GPQA Diamond": 92.0,
        "Terminal-Bench 2.1": 85.6,
        "ARC-AGI-2": 20.0,
        "Humanity's Last Exam (no tools)": 42.5
      },
      "sources": {
        "release": "https://www.anthropic.com/news/claude-sonnet-5",
        "pricing": "https://platform.claude.com/docs/en/about-claude/pricing",
        "context": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "benchmarks": "https://www.anthropic.com/claude-sonnet-5-system-card"
      },
      "verifiedDate": "2026-08-31",
      "notes": "CORRECTED 2026-07-08: max output is 128,000 tokens, not 200,000 as originally recorded. Re-confirmed against Anthropic's Models Overview. Pricing is $2/$10 per 1M; the increase to $3/$15 announced for September 1, 2026 was cancelled and the launch rate became the standard price. Context is 1,000,000 tokens and max output is 128,000. Safety-classified requests return an explicit refusal; use of another model requires configured application or product logic. Benchmark figures are provider-reported, not benchr measurements."
    },
    {
      "id": "grok-4-5",
      "name": "Grok 4.5",
      "provider": "xAI",
      "apiModelId": "grok-4.5",
      "license": "proprietary",
      "releaseDate": "2026-07-08",
      "status": "Available on the xAI API.",
      "context": {
        "windowTokens": 500000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 2.0,
        "outputPerM": 6.0,
        "cachedInputPerM": 0.3,
        "longContextThresholdTokens": 200000,
        "longContextInputPerM": 4.0,
        "longContextCachedInputPerM": 0.6,
        "longContextOutputPerM": 12.0,
        "note": "Current xAI rates. Long-context pricing applies to the entire request once the prompt reaches 200K tokens. The Batch API currently provides no discount for Grok 4.5."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://docs.x.ai/developers/release-notes",
        "pricing": "https://docs.x.ai/developers/pricing",
        "context": "https://docs.x.ai/developers/models"
      },
      "verifiedDate": "2026-07-30",
      "notes": "xAI release notes list Grok 4.5 on July 8, 2026 for coding, agentic tasks, and knowledge work. Current docs list model id grok-4.5, 500K context, $2 input / $0.30 cached / $6 output below 200K prompt tokens, and $4 / $0.60 / $12 at or above 200K, with configurable reasoning and a February 1, 2026 knowledge cutoff. xAI publishes no max output limit in the checked model documentation."
    },
    {
      "id": "grok-4-6",
      "name": "Grok 4.6",
      "provider": "xAI",
      "apiModelId": "grok-4.6",
      "license": "proprietary",
      "releaseDate": "2026-08-12",
      "status": "Available on the xAI API.",
      "context": {
        "windowTokens": 500000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 2.0,
        "outputPerM": 6.0,
        "cachedInputPerM": 0.5,
        "longContextThresholdTokens": 200000,
        "longContextInputPerM": 4.0,
        "longContextCachedInputPerM": 1.0,
        "longContextOutputPerM": 12.0,
        "note": "xAI lists $2 input / $0.50 cached / $6 output per 1M tokens below 200K prompt tokens, and $4 / $1 / $12 at or above 200K. The launch post adds a fast variant at double the standard rate. Cached input costs more than Grok 4.5's $0.30. Read on docs.x.ai/developers/models on 2026-08-24."
      },
      "benchmarks": {
        "GDPval-AA v2 (Elo)": 1753,
        "CursorBench v3.2": 69.9,
        "DeepSWE v1.1": 65.9,
        "FrontierCode v1.1": 61.3,
        "note": "Figures come from xAI's own launch table for Grok 4.6 High and are provider-reported, not a benchr test. The same table lists an Artificial Analysis Intelligence Index of 61 — a third-party index republished by xAI, so it is not recorded as an xAI benchmark here. For contrast the table shows Grok 4.5 High at GDPval-AA v2 1526, CursorBench v3.2 66.7%, DeepSWE v1.1 54%, FrontierCode v1.1 56.6%."
      },
      "sources": {
        "release": "https://x.ai/news/grok-4-6",
        "pricing": "https://docs.x.ai/developers/models",
        "context": "https://docs.x.ai/developers/models",
        "benchmarks": "https://x.ai/news/grok-4-6"
      },
      "verifiedDate": "2026-09-11",
      "notes": "xAI announced Grok 4.6 on August 12, 2026 as its frontier model for coding, agentic tasks, and knowledge work, aimed at long-running agents and interactive or visual projects, with stronger first passes and more self-testing during long tasks. The model docs list a 500,000-token context window, text and image input, text-only output with no published output-token limit, and a February 1, 2026 knowledge cutoff.",
      "arabic": {
        "status": "متاح على واجهة xAI البرمجية.",
        "notes": "أعلنت xAI عن Grok 4.6 في 12 أغسطس 2026 كنموذجها المتقدم للبرمجة والمهام الوكيلة والعمل المعرفي، مع تركيز على الوكلاء طويلة الأمد والمشاريع التفاعلية والبصرية. الوثائق تذكر سعة سياق 500,000 توكن، ومدخلات نصية وصورية، وإخراجًا نصيًا بلا حد منشور، وتاريخ معرفة يقف عند 1 فبراير 2026. السعر 2$ للمدخلات و0.50$ للمخزن مؤقتًا و6$ للمخرجات لكل مليون توكن تحت 200 ألف توكن، ويتضاعف فوقها."
      }
    },
    {
      "id": "glm-5-2",
      "name": "GLM-5.2",
      "provider": "Z.AI",
      "apiModelId": "glm-5.2",
      "license": "open-weight / hosted API",
      "releaseDate": "2026-06-16",
      "status": "Live on Z.AI API and Coding Plan.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 1.4,
        "outputPerM": 4.4,
        "cachedInputPerM": 0.26,
        "cacheStorage": "Limited-time free"
      },
      "benchmarks": {},
      "sources": {
        "release": "https://docs.z.ai/release-notes/new-released",
        "pricing": "https://docs.z.ai/guides/overview/pricing",
        "context": "https://docs.z.ai/guides/llm/glm-5.2",
        "migration": "https://docs.z.ai/guides/overview/migrate-to-glm-new"
      },
      "verifiedDate": "2026-07-13",
      "notes": "Z.AI's release notes list GLM-5.2 on 2026-06-16 with 1M lossless context and stronger long-horizon/coding performance. Model docs list 1M context and 128K maximum output tokens. Pricing docs list $1.40 input, $0.26 cached input, and $4.40 output per 1M tokens; cached input storage is listed as limited-time free. Official docs describe open-source SOTA performance but the checked docs page does not restate a precise license string."
    },
    {
      "id": "minimax-m3",
      "name": "MiniMax M3",
      "provider": "MiniMax",
      "apiModelId": "MiniMax-M3",
      "license": "proprietary hosted API",
      "releaseDate": "2026-06-01",
      "status": "Available on the MiniMax API.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 0.3,
        "outputPerM": 1.2,
        "cacheHitInputPerM": 0.06,
        "longContextInputPerM": 0.6,
        "longContextOutputPerM": 2.4,
        "longContextCacheHitInputPerM": 0.12,
        "priorityInputPerM": 0.45,
        "priorityOutputPerM": 1.8,
        "priorityLongContextInputPerM": 0.9,
        "priorityLongContextOutputPerM": 3.6
      },
      "benchmarks": {},
      "sources": {
        "release": "https://platform.minimax.io/docs/release-notes/models",
        "pricing": "https://platform.minimax.io/docs/guides/pricing-paygo",
        "context": "https://platform.minimax.io/docs/guides/text-generation"
      },
      "verifiedDate": "2026-07-13",
      "notes": "MiniMax release notes list MiniMax-M3 on June 1, 2026. The model overview/invocation docs describe it as a frontier multimodal coding model with 1M context. Pay-as-you-go pricing shows permanent 50% off rates: standard <=512K input $0.30/$1.20 per 1M with $0.06 prompt-cache read; >512K input $0.60/$2.40 with $0.12 cache read. Priority is 1.5x standard. No official max output or benchmark table was found on the checked pages."
    },
    {
      "id": "muse-spark-1-1",
      "name": "Muse Spark 1.1",
      "provider": "Meta",
      "apiModelId": "muse-spark-1.1",
      "license": "proprietary public preview",
      "releaseDate": "2026-07-09",
      "status": "Available through Meta Model API. Meta's model table now recommends Muse Spark 1.3 for new work.",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 1.25,
        "outputPerM": 4.25,
        "cachedInputPerM": 0.15,
        "note": "Read on dev.meta.ai/docs/pricing-rate-limits on 2026-09-19. Meta lists one standard price for muse-spark-1.1, 1.2 and 1.3 alike. The launch post carried no price, so this record held null until the pricing page published one."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/",
        "pricing": "https://dev.meta.ai/docs/pricing-rate-limits",
        "context": "https://dev.meta.ai/docs/models"
      },
      "verifiedDate": "2026-09-19",
      "notes": "Meta announced Muse Spark 1.1 on July 9, 2026 as a multimodal reasoning model for agents, tool and computer use, coding and multimodal understanding. Two figures changed on 2026-09-19 rather than the model changing: Meta's pricing page now publishes a standard rate of $1.25 input and $4.25 output per 1M tokens with $0.15 cached input, and its model table states the context window as 1,048,576 tokens where the launch post said \"1M\". The maximum output is still not documented. Meta's model table now marks Muse Spark 1.3 as recommended for new work; 1.1 remains callable.",
      "arabic": {
        "status": "متاح عبر Meta Model API، وتوصي جداول Meta الآن بـMuse Spark 1.3 للعمل الجديد.",
        "notes": "أعلنت Meta نموذج Muse Spark 1.1 في 9 يوليو 2026 نموذجًا متعدد الوسائط للوكلاء واستخدام الأدوات والحاسوب والبرمجة والفهم متعدد الوسائط. وما تغيّر في 19 سبتمبر 2026 رقمان لا النموذج: صفحة الأسعار باتت تنشر سعرًا قياسيًا قدره 1.25 دولار للمدخلات و4.25 للمخرجات لكل مليون توكن، و0.15 للمدخلات المخزّنة؛ وجدول النماذج يذكر سعة السياق 1,048,576 توكنًا حيث كان إعلان الإطلاق يقول «مليون». ولا يزال الحد الأقصى للإخراج غير موثّق. ويظل النموذج قابلًا للاستدعاء."
      }
    },
    {
      "id": "gemini-omni-flash-preview",
      "name": "Gemini Omni Flash Preview",
      "provider": "Google",
      "apiModelId": "gemini-omni-flash-preview",
      "license": "proprietary preview",
      "releaseDate": "2026-06-30",
      "status": "Public preview on the paid Gemini API tier.",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": null,
        "outputVideo": "3s-10s at 720p, 24 FPS"
      },
      "pricing": {
        "inputPerM": 1.5,
        "outputTextPerM": 9.0,
        "outputVideoPerM": 17.5,
        "effectiveVideoPerSecond": 0.1
      },
      "benchmarks": {},
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "pricing": "https://ai.google.dev/gemini-api/docs/pricing",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash"
      },
      "verifiedDate": "2026-07-13",
      "notes": "Google release notes list Gemini Omni Flash public preview on June 30, 2026. The model card lists gemini-omni-flash-preview, text/image/video input, video output, 1,048,576-token context, and 3-10 second 720p/24 FPS video output. Pricing lists $1.50 input, $9 text output, and $17.50 video output per 1M tokens, equivalent to about $0.10 per second of 720p video."
    },
    {
      "id": "claude-opus-5",
      "name": "Claude Opus 5",
      "provider": "Anthropic",
      "apiModelId": "claude-opus-5",
      "license": "proprietary",
      "releaseDate": "2026-07-24",
      "status": "Generally available on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry.",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 5.0,
        "outputPerM": 25.0,
        "cachedInputPerM": 0.5,
        "note": " Cache-hit input $0.50/1M confirmed on Anthropic's pricing page 2026-08-31."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://platform.claude.com/docs/en/release-notes/overview",
        "pricing": "https://platform.claude.com/docs/en/about-claude/models/overview",
        "context": "https://platform.claude.com/docs/en/about-claude/models/overview"
      },
      "verifiedDate": "2026-08-31",
      "notes": "Anthropic launched Claude Opus 5 on July 24, 2026. The official model overview lists the fixed API ID claude-opus-5, a 1M-token context window, 128K maximum output, and $5/$25 per 1M input/output tokens. Anthropic describes it as a step-change over Opus 4.8; no provider-published benchmark figure is recorded here without a specific official table. Adaptive thinking is on by default; disablement at xhigh or max effort returns a 400 error.",
      "arabic": {
        "status": "متاح عموماً عبر Claude API وAmazon Bedrock وGoogle Cloud وMicrosoft Foundry.",
        "notes": "أطلقت Anthropic نموذج Claude Opus 5 في 24 يوليو 2026. تسجل الصفحة الرسمية المعرّف claude-opus-5 وسياقاً بمليون توكن وحد إخراج 128 ألف توكن وسعراً قدره 5/25 دولاراً لكل مليون توكن للإدخال/الإخراج. لا يسجل benchr رقماً معيارياً من دون جدول رسمي محدد؛ كما أن التفكير التكيفي مفعل افتراضياً."
      }
    },
    {
      "id": "gpt-realtime-2-1",
      "name": "GPT-Realtime-2.1",
      "provider": "OpenAI",
      "apiModelId": "gpt-realtime-2.1",
      "license": "proprietary",
      "releaseDate": "2026-07-06",
      "status": "Available in the OpenAI API.",
      "context": {
        "windowTokens": 128000,
        "maxOutputTokens": 32000
      },
      "pricing": {
        "textInputPerM": 4.0,
        "cachedTextInputPerM": 0.4,
        "textOutputPerM": 24.0,
        "audioInputPerM": 32.0,
        "cachedAudioInputPerM": 0.4,
        "audioOutputPerM": 64.0,
        "imageInputPerM": 5.0,
        "cachedImageInputPerM": 0.5
      },
      "benchmarks": {},
      "sources": {
        "release": "https://developers.openai.com/api/docs/changelog",
        "pricing": "https://developers.openai.com/api/docs/models/gpt-realtime-2.1",
        "context": "https://developers.openai.com/api/docs/models/gpt-realtime-2.1",
        "implementation": "https://developers.openai.com/api/docs/models/gpt-realtime-2.1-mini"
      },
      "verifiedDate": "2026-07-29",
      "notes": "OpenAI announced GPT-Realtime-2.1 on July 6, 2026. Its model page documents text and audio input and output, image input, function calling, and no structured outputs. The Realtime guide documents WebRTC for browser and mobile clients, WebSocket for server-side media pipelines, and SIP for telephony. Text, audio, and image tokens are separate billing categories.",
      "arabic": {
        "status": "متاح عبر OpenAI API.",
        "notes": "أعلنت OpenAI عن GPT-Realtime-2.1 في 6 يوليو 2026. توثق صفحته إدخال النص والصوت والصورة، وإخراج النص والصوت، واستدعاء الأدوات، من دون مخرجات منظمة. ويوثق دليل Realtime استخدام WebRTC للمتصفح والجوال، وWebSocket للوسائط من الخادم، وSIP للاتصالات الهاتفية. فوترة النص والصوت والصورة فئات منفصلة."
      }
    },
    {
      "id": "gemini-3-1-flash-tts-preview",
      "name": "Gemini 3.1 Flash TTS Preview",
      "provider": "Google",
      "apiModelId": "gemini-3.1-flash-tts-preview",
      "license": "proprietary preview",
      "releaseDate": "2026-04-15",
      "status": "Preview in the Gemini API.",
      "context": {
        "windowTokens": 8192,
        "maxOutputTokens": 16384
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null,
        "note": "No public price was recorded from the official model page checked this session."
      },
      "benchmarks": {},
      "sources": {
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-tts-preview",
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "speech": "https://ai.google.dev/gemini-api/docs/speech-generation"
      },
      "verifiedDate": "2026-07-29",
      "notes": "Google launched Gemini 3.1 Flash TTS Preview on April 15, 2026 and added streaming generation on June 17, 2026. Its documentation records text input, audio output, 8,192 input tokens, 16,384 output tokens, Batch API support, up to two speakers, 30 prebuilt voices, and Arabic among the supported languages. The model is for exact text recitation; interactive conversational audio belongs to the Live API. No public per-token price was recorded from the checked model page.",
      "arabic": {
        "status": "معاينة في Gemini API.",
        "notes": "أطلقت Google نموذج Gemini 3.1 Flash TTS Preview في 15 أبريل 2026، وأضافت التوليد المتدفق في 17 يونيو 2026. توثق الصفحات الرسمية إدخال النص وإخراج الصوت، وسعة 8,192 توكن للإدخال و16,384 للإخراج، ودعم Batch API، ومتحدثين اثنين كحد أقصى، و30 صوتاً جاهزاً، والعربية ضمن اللغات المدعومة. النموذج لقراءة نص محدد، أما الحوار الصوتي التفاعلي فمن اختصاص Live API. لا تسجل الصفحة المفحوصة سعراً عاماً لكل توكن."
      }
    },
    {
      "id": "muse-image",
      "name": "Muse Image",
      "provider": "Meta",
      "apiModelId": null,
      "license": "proprietary",
      "releaseDate": "2026-07-07",
      "status": "Available in Meta AI; no public developer API model ID or price announced in the checked release.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {},
      "sources": {
        "release": "https://ai.meta.com/blog/introducing-muse-image-muse-video-msl/"
      },
      "verifiedDate": "2026-07-28",
      "notes": "Meta launched Muse Image on July 7, 2026 and previewed Muse Video alongside it. The checked announcement does not provide a public API model ID, price table, or token limits, so those fields remain null.",
      "arabic": {
        "status": "متاح في Meta AI؛ لا يوجد معرّف نموذج API عام أو سعر منشور في الإعلان المفحوص.",
        "notes": "أطلقت Meta نموذج Muse Image في 7 يوليو 2026 وعاينت Muse Video معه. لا يورد الإعلان المفحوص معرّف API للمطورين أو جدول أسعار أو حدود توكن، لذا تبقى هذه الحقول فارغة بوضوح."
      }
    },
    {
      "id": "gpt-realtime-2-1-mini",
      "name": "GPT-Realtime-2.1 mini",
      "provider": "OpenAI",
      "apiModelId": "gpt-realtime-2.1-mini",
      "license": "proprietary",
      "releaseDate": "2026-07-06",
      "status": "Available in the OpenAI API.",
      "context": {
        "windowTokens": 128000,
        "maxOutputTokens": 32000
      },
      "pricing": {
        "textInputPerM": 0.6,
        "cachedTextInputPerM": 0.06,
        "textOutputPerM": 2.4,
        "audioInputPerM": 10.0,
        "cachedAudioInputPerM": 0.3,
        "audioOutputPerM": 20.0,
        "imageInputPerM": 0.8,
        "cachedImageInputPerM": 0.08
      },
      "benchmarks": {},
      "sources": {
        "release": "https://developers.openai.com/api/docs/changelog",
        "pricing": "https://developers.openai.com/api/docs/models/gpt-realtime-2.1-mini",
        "context": "https://developers.openai.com/api/docs/models/gpt-realtime-2.1-mini",
        "implementation": "https://developers.openai.com/api/docs/guides/realtime"
      },
      "verifiedDate": "2026-07-29",
      "notes": "OpenAI announced GPT-Realtime-2.1 mini on July 6, 2026 as a distilled reasoning model for lower-cost realtime voice interactions. It accepts text, audio, and image input; produces text and audio; and supports function calling but not structured outputs. The Realtime API supports WebRTC, WebSocket, and SIP transports. Text, audio, and image token prices are separate billing categories.",
      "arabic": {
        "status": "متاح عبر OpenAI API.",
        "notes": "أعلنت OpenAI عن GPT-Realtime-2.1 mini في 6 يوليو 2026 بوصفه نموذج استدلال مُقطراً لتفاعلات الصوت الفورية الأقل تكلفة. يقبل النص والصوت والصورة، ويخرج النص والصوت، ويدعم استدعاء الأدوات لا المخرجات المنظمة. ويدعم Realtime وسائل WebRTC وWebSocket وSIP. تسعير النص والصوت والصورة فئات فوترة منفصلة."
      }
    },
    {
      "id": "gemini-3-5-flash-lite",
      "name": "Gemini 3.5 Flash-Lite",
      "provider": "Google",
      "apiModelId": "gemini-3.5-flash-lite",
      "license": "proprietary",
      "releaseDate": "2026-07-21",
      "status": "Stable / generally available since July 21, 2026.",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": 65536
      },
      "pricing": {
        "inputPerM": 0.3,
        "outputPerM": 2.5
      },
      "benchmarks": {
        "HLE": 18.0,
        "CharXIV": 74.5
      },
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "pricing": "https://ai.google.dev/gemini-api/docs/latest-model",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite",
        "benchmarks": "https://ai.google.dev/gemini-api/docs/latest-model"
      },
      "verifiedDate": "2026-07-28",
      "notes": "Google released the stable Gemini 3.5 Flash-Lite model on July 21, 2026 for high-throughput execution, document parsing, and subagent work. It accepts text, image, video, audio, and PDF input and returns text.",
      "arabic": {
        "status": "مستقر ومتاح عموماً منذ 21 يوليو 2026.",
        "notes": "أطلقت Google النسخة المستقرة من Gemini 3.5 Flash-Lite في 21 يوليو 2026 للتنفيذ عالي الحجم وتحليل المستندات والوكلاء الفرعيين. يقبل النص والصورة والفيديو والصوت وPDF، ويُخرج نصاً."
      }
    },
    {
      "id": "gemma-4-26b-a4b-it",
      "name": "Gemma 4 26B A4B IT",
      "provider": "Google",
      "apiModelId": "gemma-4-26b-a4b-it",
      "license": "Apache-2.0",
      "releaseDate": "2026-04-02",
      "status": "Available in AI Studio and through the Gemini API.",
      "context": {
        "windowTokens": 256000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null
      },
      "architecture": {
        "type": "Mixture-of-Experts",
        "totalParametersB": 25.2,
        "activeParametersB": 3.8,
        "layers": 30,
        "activeExperts": 8,
        "totalExperts": 128,
        "sharedExperts": 1
      },
      "inferenceMemoryGB": {
        "BF16": 57.7,
        "SFP8": 28.8,
        "Q4_0": 14.4
      },
      "benchmarks": {
        "MMLU Pro": 82.6,
        "LiveCodeBench v6": 77.1,
        "GPQA Diamond": 82.3,
        "Tau2 (average over 3)": 68.2,
        "MMMU Pro": 73.8,
        "OmniDocBench 1.5": 0.149,
        "MRCR v2 8 needle 128k": 44.1
      },
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "modelCard": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "context": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "benchmarks": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "architecture": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "memory": "https://ai.google.dev/gemma/docs/core",
        "license": "https://ai.google.dev/gemma/docs/core/model_card_4"
      },
      "verifiedDate": "2026-07-29",
      "notes": "Google released gemma-4-26b-a4b-it on April 2, 2026. The official Gemma 4 model card describes a 25.2B-parameter Mixture-of-Experts model that activates 3.8B parameters per token, supports 256K context, text and image input, and Apache 2.0 licensing. Google's memory table estimates 57.7 GB at BF16, 28.8 GB at SFP8, and 14.4 GB at Q4_0 before context cache and runtime overhead. No model-specific Gemini API token price or maximum output limit is published in the checked official sources.",
      "arabic": {
        "status": "متاح في AI Studio وعبر Gemini API.",
        "notes": "أطلقت Google المعرّف gemma-4-26b-a4b-it في 2 أبريل 2026. بطاقة Gemma 4 الرسمية تصفه كنموذج MoE بإجمالي 25.2 مليار بارامتر، ينشّط 3.8 ملياراً لكل توكن، ويدعم سياق 256K ونصاً وصوراً بترخيص Apache 2.0. تقديرات الذاكرة الرسمية هي 57.7 GB لدقة BF16 و28.8 GB لدقة SFP8 و14.4 GB لدقة Q4_0 قبل ذاكرة السياق وحمل بيئة التشغيل. لم تنشر Google سعراً خاصاً بهذا المعرّف عبر Gemini API أو حداً أقصى للإخراج في المصادر المفحوصة."
      }
    },
    {
      "id": "gemma-4-31b-it",
      "name": "Gemma 4 31B IT",
      "provider": "Google",
      "apiModelId": "gemma-4-31b-it",
      "license": "Apache-2.0",
      "releaseDate": "2026-04-02",
      "status": "Available in AI Studio and through the Gemini API.",
      "context": {
        "windowTokens": 256000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null
      },
      "architecture": {
        "type": "Dense",
        "totalParametersB": 30.7,
        "layers": 60
      },
      "inferenceMemoryGB": {
        "BF16": 69.9,
        "SFP8": 34.9,
        "Q4_0": 17.5
      },
      "benchmarks": {
        "MMLU Pro": 85.2,
        "LiveCodeBench v6": 80.0,
        "GPQA Diamond": 84.3,
        "Tau2 (average over 3)": 76.9,
        "MMMU Pro": 76.9,
        "OmniDocBench 1.5": 0.131,
        "MRCR v2 8 needle 128k": 66.4
      },
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "modelCard": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "context": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "benchmarks": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "architecture": "https://ai.google.dev/gemma/docs/core/model_card_4",
        "memory": "https://ai.google.dev/gemma/docs/core",
        "license": "https://ai.google.dev/gemma/docs/core/model_card_4"
      },
      "verifiedDate": "2026-07-29",
      "notes": "Google released gemma-4-31b-it on April 2, 2026. The official Gemma 4 model card describes a dense 30.7B-parameter, 60-layer text-and-image model with a 256K context window and Apache 2.0 licensing. Google's memory table estimates 69.9 GB at BF16, 34.9 GB at SFP8, and 17.5 GB at Q4_0 before context cache and runtime overhead. No model-specific Gemini API token price or maximum output limit is published in the checked official sources.",
      "arabic": {
        "status": "متاح في AI Studio وعبر Gemini API.",
        "notes": "أطلقت Google المعرّف gemma-4-31b-it في 2 أبريل 2026. بطاقة Gemma 4 الرسمية تصفه كنموذج كثيف بإجمالي 30.7 مليار بارامتر و60 طبقة، ويدعم سياق 256K ونصاً وصوراً بترخيص Apache 2.0. تقديرات الذاكرة الرسمية هي 69.9 GB لدقة BF16 و34.9 GB لدقة SFP8 و17.5 GB لدقة Q4_0 قبل ذاكرة السياق وحمل بيئة التشغيل. لم تنشر Google سعراً خاصاً بهذا المعرّف عبر Gemini API أو حداً أقصى للإخراج في المصادر المفحوصة."
      }
    },
    {
      "id": "leanstral-1-5",
      "name": "Leanstral 1.5",
      "provider": "Mistral AI",
      "apiModelId": null,
      "license": "Apache-2.0",
      "releaseDate": "2026-07-02",
      "status": "Open-weight model; Mistral also states it is available through a free API.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {
        "PutnamBench solved": 587,
        "PutnamBench total": 672,
        "FATE-H": 87.0,
        "FATE-X": 34.0
      },
      "sources": {
        "release": "https://mistral.ai/fr/news/leanstral-1-5/",
        "benchmarks": "https://mistral.ai/fr/news/leanstral-1-5/"
      },
      "verifiedDate": "2026-07-28",
      "notes": "Mistral describes Leanstral 1.5 as a formal-verification specialist with 119B total and 6B active parameters. It is not a general-chat model, so it stays out of benchr's general model ranking.",
      "arabic": {
        "status": "نموذج بأوزان مفتوحة؛ وتقول Mistral إنه متاح أيضاً عبر API مجاني.",
        "notes": "تصف Mistral Leanstral 1.5 بأنه متخصص في التحقق الرسمي، بـ119 مليار معامل إجمالي و6 مليارات نشطة. ليس نموذج محادثة عاماً، لذلك لا يدخل في الترتيب العام لدى benchr."
      }
    },
    {
      "id": "qwen-audio-3-0-tts-plus",
      "name": "Qwen-Audio 3.0 TTS Plus",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "qwen-audio-3.0-tts-plus",
      "license": "proprietary",
      "releaseDate": "2026-07-14",
      "status": "Available through Alibaba Cloud Model Studio.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {},
      "sources": {
        "release": "https://www.alibabacloud.com/help/en/model-studio/newly-released-models",
        "context": "https://www.alibabacloud.com/help/en/model-studio/tts-model/",
        "voices": "https://www.alibabacloud.com/help/en/model-studio/qwen-audio-tts-voice-list",
        "api": "https://www.alibabacloud.com/help/en/model-studio/cosyvoice-client-events"
      },
      "verifiedDate": "2026-07-29",
      "notes": "Alibaba lists this instruction-controlled WebSocket TTS model for standard synthesis using built-in voices. The current voice list documents Mandarin and English system voices, no voice cloning or voice design, and no SSML or word timestamps for the listed voices. The API can mark generated audio with an AIGC watermark. The checked documentation does not provide token limits or a public per-token price.",
      "arabic": {
        "status": "متاح عبر Alibaba Cloud Model Studio.",
        "notes": "تدرج Alibaba هذا النموذج لتحويل النص إلى كلام عبر WebSocket باستخدام أصوات مدمجة وتعليمات أسلوبية. قائمة الأصوات الحالية توثق أصواتاً نظامية للصينية والإنجليزية، من دون استنساخ أو تصميم صوت، ومن دون SSML أو طوابع زمنية للكلمات في الأصوات المدرجة. ويمكن للواجهة وسم الصوت بعلامة AIGC. لا تنشر الوثائق المفحوصة حدود توكن أو سعراً عاماً لكل توكن."
      }
    },
    {
      "id": "qwen-audio-3-0-tts-flash",
      "name": "Qwen-Audio 3.0 TTS Flash",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "qwen-audio-3.0-tts-flash",
      "license": "proprietary",
      "releaseDate": "2026-07-14",
      "status": "Available through Alibaba Cloud Model Studio.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {},
      "sources": {
        "release": "https://www.alibabacloud.com/help/en/model-studio/newly-released-models",
        "context": "https://www.alibabacloud.com/help/en/model-studio/tts-model/",
        "voices": "https://www.alibabacloud.com/help/en/model-studio/qwen-audio-tts-voice-list",
        "api": "https://www.alibabacloud.com/help/en/model-studio/cosyvoice-client-events"
      },
      "verifiedDate": "2026-07-29",
      "notes": "Alibaba positions the WebSocket Flash variant for low-latency synthesis and authorized voice cloning, claiming first audio in under 200 ms. Its documented system voices are Mandarin and English; the published cloned-voice language list does not include Arabic. Voice design, SSML, pronunciation hot fixes, and word timestamps are not documented for this model. The API can mark generated audio with an AIGC watermark. The checked documentation does not provide token limits or a public per-token price.",
      "arabic": {
        "status": "متاح عبر Alibaba Cloud Model Studio.",
        "notes": "تقدم Alibaba نسخة Flash عبر WebSocket لتوليد الصوت منخفض الكمون واستنساخ الصوت المصرح به، وتذكر أن أول صوت يصل خلال أقل من 200 مللي ثانية. الأصوات النظامية الموثقة للصينية والإنجليزية، ولا تشمل قائمة لغات الأصوات المستنسخة المنشورة العربية. لا توثق الصفحة تصميم الصوت أو SSML أو إصلاح النطق الفوري أو طوابع الكلمات الزمنية لهذا النموذج. ويمكن للواجهة وسم الصوت بعلامة AIGC. لا تنشر الوثائق المفحوصة حدود توكن أو سعراً عاماً لكل توكن."
      }
    },
    {
      "id": "muse-video-preview",
      "name": "Muse Video Preview",
      "provider": "Meta",
      "apiModelId": null,
      "license": "proprietary preview",
      "announcementDate": "2026-07-07",
      "releaseDate": null,
      "status": "Announced as an early preview and still described by Meta as coming soon; not a released developer product in the checked source.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {},
      "sources": {
        "announcement": "https://ai.meta.com/blog/introducing-muse-image-muse-video-msl/"
      },
      "verifiedDate": "2026-07-30",
      "notes": "Meta announced Muse Video as an early preview on July 7, 2026 and explicitly said it was coming soon to creators and Meta AI. No public developer API ID, release date, price, token limits, or benchmark table was found in the launch announcement.",
      "arabic": {
        "status": "معاينة مبكرة أعلنتها Meta وما زالت تصفها بأنها قادمة قريباً؛ ليست منتجاً مطروحاً للمطورين في المصدر المفحوص.",
        "notes": "أعلنت Meta عن Muse Video كمعاينة مبكرة في 7 يوليو 2026، وقالت صراحة إنه قادم قريباً للمبدعين وMeta AI. لا يتضمن الإعلان معرّف API عاماً أو تاريخ إطلاق أو سعراً أو حدود توكن أو جدول معايير."
      }
    },
    {
      "id": "gpt-live-1",
      "name": "GPT-Live-1",
      "provider": "OpenAI",
      "apiModelId": "gpt-live-1",
      "license": "proprietary",
      "releaseDate": "2026-07-08",
      "status": "Generally available in the OpenAI API since September 10, 2026, through the Live endpoint only.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null,
        "sessionPerMinute": 0.05,
        "note": "OpenAI prices GPT-Live-1 as a voice session rather than per token: $0.05 per minute, billed per second with no rounding up to a whole minute. Backend model and tool usage is billed separately, so the session rate is not the full cost of a conversation. Read on developers.openai.com/api/docs/pricing on 2026-09-19."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://openai.com/index/introducing-gpt-live/",
        "generalAvailability": "https://developers.openai.com/api/docs/changelog",
        "pricing": "https://developers.openai.com/api/docs/pricing",
        "context": "https://developers.openai.com/api/docs/models/gpt-live-1"
      },
      "verifiedDate": "2026-09-19",
      "notes": "OpenAI introduced GPT-Live-1 for ChatGPT Voice on July 8, 2026 and moved it to general availability on the API on September 10, 2026 under the ID gpt-live-1. The model page documents audio and text in and out, streaming and function calling, a July 31, 2025 knowledge cutoff, and no image or video input. The Live endpoint is the only supported interface: Chat Completions and Responses are not. The page states no context window and no maximum output, so both stay null.",
      "arabic": {
        "status": "متاح للجميع في واجهة OpenAI منذ 10 سبتمبر 2026، عبر نقطة Live وحدها.",
        "notes": "قدّمت OpenAI نموذج GPT-Live-1 لصوت ChatGPT في 8 يوليو 2026، ثم أتاحته للجميع في الواجهة البرمجية في 10 سبتمبر 2026 بالمعرّف gpt-live-1. توثّق صفحته إدخال الصوت والنص وإخراجهما، والبث المتدفق واستدعاء الأدوات، وحد معرفة ينتهي في 31 يوليو 2025، ولا تدعم إدخال الصور أو الفيديو. نقطة Live هي الواجهة الوحيدة المدعومة، فلا تعمل معه Chat Completions ولا Responses. يُحاسب الاستخدام بالدقيقة الصوتية لا بالتوكن: 0.05 دولار للدقيقة تُحسب بالثانية، وتُفوتر نماذج الخلفية والأدوات على حدة. ولا تذكر الصفحة سعة سياق ولا حدًّا أقصى للإخراج."
      }
    },
    {
      "id": "gpt-live-1-mini",
      "name": "GPT-Live-1 mini",
      "provider": "OpenAI",
      "apiModelId": null,
      "license": "proprietary",
      "releaseDate": "2026-07-08",
      "status": "Announced for ChatGPT Voice. Still not documented as an API model: no model page and no entry in the API pricing table as of September 19, 2026.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {},
      "sources": {
        "release": "https://openai.com/index/introducing-gpt-live/",
        "pricing": "https://developers.openai.com/api/docs/pricing"
      },
      "verifiedDate": "2026-09-19",
      "notes": "OpenAI introduced GPT-Live-1 mini alongside GPT-Live-1 as a full-duplex model for ChatGPT Voice. When GPT-Live-1 reached API general availability on September 10, 2026, the mini did not: on 2026-09-19 the API pricing table lists GPT-Live-1 as the only GPT-Live session model, and developers.openai.com/api/docs/models/gpt-live-1-mini returns 404. benchr keeps the record because OpenAI announced the model; its API ID, price and token limits stay null because no OpenAI page states them.",
      "arabic": {
        "status": "أُعلن عنه لصوت ChatGPT، ولم يُوثَّق بعد نموذجًا في الواجهة البرمجية: لا صفحة له ولا ذكر له في جدول الأسعار حتى 19 سبتمبر 2026.",
        "notes": "قدّمت OpenAI نموذج GPT-Live-1 mini مع GPT-Live-1 نموذجًا ثنائي الاتجاه لصوت ChatGPT. وحين أُتيح GPT-Live-1 للجميع في الواجهة البرمجية في 10 سبتمبر 2026 لم يلحق به النموذج الأصغر: ففي 19 سبتمبر 2026 لا يذكر جدول الأسعار سوى GPT-Live-1 نموذجًا لجلسات الصوت، وصفحة النموذج الأصغر تعيد خطأ 404. يحتفظ benchr بالسجل لأن OpenAI أعلنت النموذج، ويترك معرّفه وسعره وحدوده فارغة لأن أي صفحة رسمية لا تذكرها."
      }
    },
    {
      "id": "gemini-robotics-er-2-preview",
      "name": "Gemini Robotics ER 2 Preview",
      "provider": "Google",
      "apiModelId": "gemini-robotics-er-2-preview",
      "license": "proprietary preview",
      "releaseDate": null,
      "status": "Preview model for robotics tasks.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null
      },
      "benchmarks": {},
      "sources": {
        "context": "https://ai.google.dev/gemini-api/docs/models"
      },
      "verifiedDate": "2026-08-14",
      "notes": "Google's current Gemini API model directory lists Gemini Robotics ER 2 Preview under specialized task models, with API ID gemini-robotics-er-2-preview. Google describes it as an embodied-reasoning model for video understanding, spatial reasoning, multi-step tool orchestration, and multi-robot collaboration. The checked directory does not publish a release date, public token price, context limit, or max output, so benchr records those fields as null rather than inferring them.",
      "arabic": {
        "status": "نموذج معاينة لمهام الروبوتات.",
        "notes": "تدرج Google في دليل نماذج Gemini الحالي نموذج Gemini Robotics ER 2 Preview ضمن النماذج المتخصصة، بمعرّف API gemini-robotics-er-2-preview. تصفه كنموذج استدلال متجسّد لفهم الفيديو والاستدلال المكاني وتنسيق الأدوات متعددة الخطوات وتعاون الروبوتات. لا ينشر الدليل الذي تم التحقق منه تاريخ إصدار أو سعراً عاماً للتوكنات أو سعة سياق أو حد إخراج، لذلك تسجل benchr هذه الحقول فارغة بدلاً من تخمينها."
      }
    },
    {
      "id": "gemini-3-1-flash-lite-image",
      "name": "Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite)",
      "provider": "Google",
      "apiModelId": "gemini-3.1-flash-lite-image",
      "license": "proprietary",
      "releaseDate": "2026-06-30",
      "status": "Stable / generally available.",
      "context": {
        "windowTokens": 65536,
        "maxOutputTokens": 4096
      },
      "pricing": {
        "inputPerM": 0.25,
        "outputTextPerM": 1.5,
        "outputImagePerM": 30.0,
        "perImage1K": 0.0336,
        "batchInputPerM": 0.125,
        "batchOutputTextPerM": 0.75,
        "batchOutputImagePerM": 15.0,
        "batchPerImage1K": 0.0168
      },
      "benchmarks": {},
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "pricing": "https://ai.google.dev/gemini-api/docs/pricing",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite-image"
      },
      "verifiedDate": "2026-07-13",
      "notes": "Google release notes list gemini-3.1-flash-lite-image (Nano Banana 2 Lite) GA on June 30, 2026. The model card lists 65,536 input tokens, 4,096 output tokens, text/image input, image/text output, image generation and editing, and latest update June 2026. Pricing lists $0.25 input, $1.50 text/thinking output, $30 image output per 1M tokens, equivalent to $0.0336 per 1K image; batch halves those rates."
    },
    {
      "id": "deepseek-v4-flash-vision-exp",
      "name": "DeepSeek-V4-Flash-Vision-Exp",
      "provider": "DeepSeek",
      "apiModelId": "deepseek-v4-flash-vision-exp",
      "license": null,
      "releaseDate": "2026-08-21",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 384000
      },
      "pricing": {
        "inputPerM": 0.44,
        "outputPerM": 1.32,
        "cacheHitInputPerM": 0.014,
        "offPeakInputPerM": 0.22,
        "offPeakOutputPerM": 0.66,
        "offPeakCacheHitInputPerM": 0.007,
        "offPeakDiscount": 0.5,
        "peakHoursUTC": "01:00-04:00 and 06:00-10:00, Monday to Friday",
        "note": "Priced identically to DeepSeek-V4-Flash on the official pricing table, under the same peak/off-peak scheme. Images are tokenised by dimension and billed as input tokens. Read on api-docs.deepseek.com/quick_start/pricing on 2026-08-28."
      },
      "benchmarks": {
        "Terminal-Bench 2.1": 83.9,
        "NL2Repo": 57.7,
        "DeepSWE": 59.3,
        "DSBench-Hard": 63.6,
        "AutomationBench (Public)": 25.7,
        "ApexBench (Pass@1)": 36.5,
        "Agents' Last Exam": 27.3,
        "Chartography": 64.3,
        "ZeroBench (Pass@5)": 35.0
      },
      "sources": {
        "release": "https://api-docs.deepseek.com/updates",
        "pricing": "https://api-docs.deepseek.com/quick_start/pricing",
        "context": "https://api-docs.deepseek.com/quick_start/pricing",
        "benchmarks": "https://api-docs.deepseek.com/updates",
        "retirement": "https://api-docs.deepseek.com/news/news260910"
      },
      "verifiedDate": "2026-09-11",
      "notes": "RETIRED 2026-09-10: DeepSeek's September 10, 2026 release note says \"V4-Flash & V4-Flash-Vision-Exp are retired\" and routes the legacy ID to DeepSeek-V4.1-Flash; the pricing page read on 2026-09-11 no longer lists DeepSeek-V4-Flash-Vision-Exp. The record is kept as the permanent reference for what it cost and scored. DeepSeek published DeepSeek-V4-Flash-Vision-Exp on August 21, 2026 and labels it experimental; it is called with model='deepseek-v4-flash-vision-exp'. The pricing table lists it alongside V4-Flash at the same rates, with a 1M context window, 384K maximum output, and 2,500 concurrent requests. The benchmark figures are DeepSeek's own published numbers for the experimental checkpoint, not benchr tests, and an experimental model can change or disappear without a deprecation notice. No license or open-weights release was found for this checkpoint on an official page, so license is null rather than assumed to match V4-Flash's MIT weights.",
      "status": "Retired on the DeepSeek API on September 10, 2026. The deepseek-v4-flash-vision-exp ID temporarily routes to DeepSeek-V4.1-Flash (deepseek-flash) at V4.1-Flash rates. The prices below are the last rates DeepSeek listed for this model."
    },
    {
      "id": "deepseek-v4-1-flash",
      "name": "DeepSeek-V4.1-Flash",
      "provider": "DeepSeek",
      "apiModelId": "deepseek-flash",
      "license": null,
      "releaseDate": "2026-09-10",
      "status": "Available on the DeepSeek API since September 10, 2026, replacing DeepSeek-V4-Flash and V4-Flash-Vision-Exp.",
      "params": {
        "total": "552B",
        "active": null
      },
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 384000
      },
      "pricing": {
        "inputPerM": 0.3,
        "outputPerM": 1.2,
        "cacheHitInputPerM": 0.006,
        "offPeakInputPerM": 0.15,
        "offPeakOutputPerM": 0.6,
        "offPeakCacheHitInputPerM": 0.003,
        "offPeakDiscount": 0.5,
        "peakHoursUTC": "01:00-04:00 and 06:00-10:00, Monday to Friday",
        "note": "Read on api-docs.deepseek.com/quick_start/pricing on 2026-09-11. inputPerM, outputPerM and cacheHitInputPerM are the peak (list) rates; the offPeak* fields are the half-price rates DeepSeek applies outside 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday."
      },
      "benchmarks": {
        "note": "DeepSeek's release note says V4.1-Flash posts benchmark results ahead of flagship models including DeepSeek-V4-Pro, but benchr has not recorded a numeric table from an official page, so every numeric field stays null. That is DeepSeek's claim, not a benchr test."
      },
      "sources": {
        "release": "https://api-docs.deepseek.com/news/news260910",
        "pricing": "https://api-docs.deepseek.com/quick_start/pricing",
        "context": "https://api-docs.deepseek.com/quick_start/pricing"
      },
      "verifiedDate": "2026-09-11",
      "notes": "DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026 under the API ID deepseek-flash, describing it as a 552B-parameter mixture-of-experts model with 8B active parameters for input and 16B for output. The pricing page lists a 1M context window, 384K maximum output and 2,500 concurrent requests. It replaced DeepSeek-V4-Flash and V4-Flash-Vision-Exp, whose legacy IDs temporarily route to it. The release note does not state a license or link open weights, so license is null rather than copied from V4-Flash's MIT weights, and the active-parameter field is null because DeepSeek gives two figures rather than one.",
      "arabic": {
        "status": "متاح على واجهة DeepSeek البرمجية منذ 10 سبتمبر 2026 بديلًا لـ DeepSeek-V4-Flash وV4-Flash-Vision-Exp.",
        "notes": "أطلقت DeepSeek نموذج DeepSeek-V4.1-Flash في 10 سبتمبر 2026 بالمعرّف deepseek-flash، ووصفته بأنه نموذج MoE حجمه 552B معامل، يُفعّل 8B منها للمدخلات و16B للمخرجات. تذكر صفحة الأسعار سعة سياق 1M توكن وحدًا أقصى للإخراج 384K. سعره في ساعات الذروة 0.30$ للمدخلات و1.20$ للمخرجات لكل مليون توكن، وينخفض إلى النصف خارجها. لم تذكر DeepSeek ترخيصًا ولا أوزانًا مفتوحة لهذا الإصدار، ولم تنشر جدول بنشمارك رقميًا يسجّله benchr."
      }
    },
    {
      "id": "glm-5-3",
      "name": "GLM-5.3",
      "provider": "Z.ai",
      "apiModelId": "glm-5.3",
      "license": null,
      "releaseDate": "2026-08-18",
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 1.4,
        "outputPerM": 4.4,
        "cachedInputPerM": 0.26,
        "cacheStorage": "Limited-time free",
        "note": "Same list rate as GLM-5.2. Read on docs.z.ai/guides/overview/pricing on 2026-08-28."
      },
      "benchmarks": {
        "CyberGym": 84.5,
        "Terminal-Bench 3.0": 28.3
      },
      "sources": {
        "release": "https://docs.z.ai/release-notes/new-released",
        "pricing": "https://docs.z.ai/guides/overview/pricing",
        "context": "https://docs.z.ai/guides/llm/glm-5.3",
        "benchmarks": "https://docs.z.ai/guides/llm/glm-5.3"
      },
      "verifiedDate": "2026-08-28",
      "notes": "Z.ai released GLM-5.3 on August 18, 2026. The model guide states it uses the same base model as GLM-5.2 with the gains coming from post-training, reasoning is always on with low, high, or max effort (max is the default), and inputs are text-only. Z.ai reports a 50% gain over GLM-5.2 on its own Z.ai Code Bench, 84.5% on CyberGym, and a move from 4.6 to 28.3 on Terminal-Bench 3.0 - all provider-reported figures on a provider-run harness, not benchr tests, and Z.ai Code Bench is the vendor's own benchmark. Parameter counts are not published on the model guide, so they are omitted rather than inferred from GLM-5.2."
    },
    {
      "id": "glm-5-3-flash",
      "name": "GLM-5.3-Flash",
      "provider": "Z.ai",
      "apiModelId": "glm-5.3-flash",
      "license": "MIT",
      "releaseDate": "2026-08-26",
      "params": {
        "total": "320B",
        "active": "18B"
      },
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 0.15,
        "outputPerM": 0.5,
        "cachedInputPerM": 0.03,
        "listInputPerM": 0.15,
        "listOutputPerM": 0.5,
        "listCachedInputPerM": 0.03,
        "promoEndDate": "2026-09-09",
        "cacheStorage": "Limited-time free",
        "selfHost": true,
        "note": "Z.ai lists $0.15 input / $0.03 cached input / $0.50 output per 1M tokens, read on docs.z.ai/guides/overview/pricing on 2026-09-11, with cached-input storage still marked limited-time free. A launch promotion halved every rate through 24:00 on September 9, 2026 (UTC+8): $0.075 / $0.015 / $0.25, recorded in the promo* fields and read on 2026-08-28.",
        "promoInputPerM": 0.075,
        "promoOutputPerM": 0.25,
        "promoCachedInputPerM": 0.015
      },
      "benchmarks": {
        "Terminal-Bench 2.1": 84.3,
        "DeepSWE": 63.4,
        "Humanity's Last Exam (with tools)": 55.3
      },
      "sources": {
        "release": "https://docs.z.ai/release-notes/new-released",
        "pricing": "https://docs.z.ai/guides/overview/pricing",
        "context": "https://docs.z.ai/guides/llm/glm-5.3-flash",
        "license": "https://huggingface.co/zai-org/GLM-5.3-Flash",
        "benchmarks": "https://huggingface.co/zai-org/GLM-5.3-Flash"
      },
      "verifiedDate": "2026-09-11",
      "notes": "Z.ai released GLM-5.3-Flash on August 26, 2026 as a natively multimodal MoE with 320B total and 18B active parameters on a hybrid sparse-and-linear attention architecture. The model guide lists a 1M context window, 128K maximum output, and video, image, text, and file input with text output. Weights are published under MIT on the official zai-org model card. The benchmark figures are Z.ai's own published numbers, not benchr tests, and the card evaluates at a 300K context with a 163,840-token generation limit rather than at the full API window. At its promotional rate it was the lowest listed hosted rate in benchr's tracked set on August 28, 2026. The promotion ended on September 9, 2026, and the list rate of $0.15 / $0.50 applies from then."
    },
    {
      "id": "qwen-3-8b",
      "name": "Qwen3-8B",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "Qwen/Qwen3-8B",
      "license": "Apache-2.0",
      "releaseDate": null,
      "context": {
        "windowTokens": 32768,
        "windowTokensExtended": 131072,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null
      },
      "architecture": {
        "type": "Dense",
        "totalParametersB": 8.2,
        "layers": 36,
        "attention": "Grouped-query attention (32 query / 8 KV heads)"
      },
      "benchmarks": {},
      "sources": {
        "api": "https://huggingface.co/Qwen/Qwen3-8B",
        "license": "https://huggingface.co/Qwen/Qwen3-8B",
        "context": "https://huggingface.co/Qwen/Qwen3-8B",
        "architecture": "https://huggingface.co/Qwen/Qwen3-8B"
      },
      "verifiedDate": "2026-08-15",
      "notes": "Official Qwen model card: 8.2B parameters, 32,768 native-token context and YaRN-validated context up to 131,072 tokens. The extended limit requires a compatible runtime configuration; it is not assumed as a default local setting. No provider release date was recorded here because the checked card did not provide one in a directly attributable release field."
    },
    {
      "id": "phi-4-mini-instruct",
      "name": "Phi-4-mini-instruct",
      "provider": "Microsoft",
      "apiModelId": "microsoft/Phi-4-mini-instruct",
      "license": "MIT",
      "releaseDate": null,
      "context": {
        "windowTokens": 128000,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null
      },
      "architecture": {
        "type": "Dense decoder-only Transformer",
        "totalParametersB": 3.8
      },
      "benchmarks": {},
      "sources": {
        "api": "https://huggingface.co/microsoft/Phi-4-mini-instruct",
        "license": "https://huggingface.co/microsoft/Phi-4-mini-instruct",
        "context": "https://huggingface.co/microsoft/Phi-4-mini-instruct",
        "architecture": "https://huggingface.co/microsoft/Phi-4-mini-instruct"
      },
      "verifiedDate": "2026-08-15",
      "notes": "Microsoft's official model card describes a 3.8B-parameter dense decoder-only Transformer with a 128K-token context and MIT license. Its stated intended uses include memory/compute-constrained and latency-bound scenarios; that is a provider positioning statement, not a BenchR throughput measurement."
    },
    {
      "id": "llama-3-2-3b-instruct",
      "name": "Llama 3.2 3B Instruct",
      "provider": "Meta",
      "apiModelId": "meta-llama/Llama-3.2-3B-Instruct",
      "license": "Llama 3.2 Community License",
      "releaseDate": "2024-09-25",
      "context": {
        "windowTokens": 128000,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null
      },
      "architecture": {
        "type": "Dense",
        "totalParametersB": 3.21,
        "attention": "Grouped-query attention"
      },
      "benchmarks": {},
      "sources": {
        "release": "https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct",
        "api": "https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct",
        "license": "https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct",
        "context": "https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct",
        "architecture": "https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct"
      },
      "verifiedDate": "2026-08-15",
      "notes": "Meta's official card lists the text-only 3B model at 3.21B parameters and 128K context. The separate Meta quantized text-only table lists an 8K context; BenchR does not imply that every quantized local artifact preserves the 128K base-model limit."
    },
    {
      "id": "mistral-small-3-1-24b-instruct",
      "name": "Mistral Small 3.1 24B Instruct",
      "provider": "Mistral AI",
      "apiModelId": "mistralai/Mistral-Small-3.1-24B-Instruct-2503",
      "license": "Apache-2.0",
      "releaseDate": null,
      "context": {
        "windowTokens": 128000,
        "maxOutputTokens": null
      },
      "pricing": {
        "selfHost": true,
        "inputPerM": null,
        "outputPerM": null
      },
      "architecture": {
        "type": "Dense",
        "totalParametersB": 24
      },
      "benchmarks": {},
      "sources": {
        "api": "https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503",
        "license": "https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503",
        "context": "https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503",
        "architecture": "https://huggingface.co/mistralai/Mistral-Small-3.1-24B-Instruct-2503"
      },
      "verifiedDate": "2026-08-15",
      "notes": "Mistral's official card identifies this 24B instruction-tuned model as Apache 2.0 with 128K context, text and vision input, and local deployment after quantization. Its note that a quantized build can fit a 32GB RAM MacBook is provider guidance for a particular quantized setup, not a universal performance promise."
    },
    {
      "id": "kimi-k3",
      "name": "Kimi K3",
      "provider": "Moonshot AI",
      "apiModelId": "kimi-k3",
      "license": "Kimi K3 License",
      "releaseDate": "2026-07-16",
      "status": "Available in Kimi products, Kimi Code, and the Kimi API; open weights published by the official Moonshot AI organization.",
      "params": {
        "total": "2.8T",
        "active": "104B",
        "experts": "896 (16 selected/token; 2 shared)"
      },
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": 1048576,
        "defaultMaxOutputTokens": 131072
      },
      "modalities": {
        "input": [
          "text",
          "image",
          "video via hosted API"
        ],
        "output": [
          "text"
        ]
      },
      "reasoning": {
        "supported": true,
        "alwaysEnabled": true,
        "levels": [
          "low",
          "high",
          "max"
        ],
        "default": "max"
      },
      "capabilities": {
        "automaticContextCaching": true,
        "functionCalling": true,
        "toolChoice": true,
        "dynamicToolLoading": true,
        "jsonMode": true,
        "structuredOutputs": true,
        "partialMode": true,
        "batchApi": true,
        "webSearch": "Temporarily not recommended by Moonshot while the tool is being updated",
        "selfHost": true
      },
      "pricing": {
        "inputPerM": 3.0,
        "outputPerM": 15.0,
        "cacheHitInputPerM": 0.3,
        "selfHost": true,
        "note": "Global Kimi API rates. The model uses flat pay-as-you-go pricing with no context-length tier; applicable taxes are extra."
      },
      "benchmarks": {
        "GPQA Diamond": 93.5,
        "DeepSWE": 67.5,
        "ProgramBench": 77.8,
        "Terminal-Bench 2.1": 88.3,
        "FrontierSWE": 81.2,
        "BrowseComp": 91.2,
        "Toolathlon Verified": 76.5,
        "OSWorld Verified": 84.8,
        "MMMU-Pro without tools": 81.6,
        "MMMU-Pro with tools": 83.4
      },
      "sources": {
        "release": "https://www.kimi.ai/blog/kimi-k3",
        "releaseDate": "https://www.kimi.com/code/docs/en/kimi-code/whats-new.html",
        "pricing": "https://platform.kimi.ai/docs/pricing/chat-k3",
        "context": "https://platform.kimi.ai/docs/guide/kimi-k3-quickstart",
        "license": "https://huggingface.co/moonshotai/Kimi-K3",
        "benchmarks": "https://huggingface.co/moonshotai/Kimi-K3"
      },
      "verifiedDate": "2026-08-31",
      "notes": "Moonshot released Kimi K3 on July 16, 2026 and later published its weights. The official model card lists 2.8T total parameters, 104B active parameters, 896 experts with 16 selected per token, 1,048,576 context, native image/text understanding, and provider-run benchmark results. Hosted API documentation additionally supports video files, tools, strict structured output, and up to 1,048,576 completion tokens (131,072 default). Pricing and model card re-verified on the official pricing page and Hugging Face card on 2026-08-31."
    },
    {
      "id": "gemini-3-8-flash",
      "name": "Gemini 3.8 Flash",
      "provider": "Google",
      "apiModelId": "gemini-3.8-flash",
      "license": "proprietary",
      "releaseDate": "2026-09-02",
      "status": "GA (stable) since September 2, 2026.",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": 65536
      },
      "pricing": {
        "inputPerM": 0.75,
        "outputPerM": 3.75,
        "batchInputPerM": 0.375,
        "batchOutputPerM": 1.875,
        "flexInputPerM": 0.375,
        "flexOutputPerM": 1.875,
        "priorityInputPerM": 1.35,
        "priorityOutputPerM": 6.75,
        "cachedInputPerM": 0.075,
        "cacheStoragePerMTokenHour": 0.5,
        "postPromoInputPerM": 1.5,
        "postPromoOutputPerM": 7.5,
        "postPromoCachedInputPerM": 0.15,
        "postPromoStartDate": "2027-01-01",
        "freeApiTier": true,
        "note": "Introductory pricing: $0.75 input / $3.75 output per 1M tokens through December 31, 2026, then $1.50 / $7.50 from January 1, 2027. Every tier doubles on the same date, including batch, flex, priority and cached input."
      },
      "benchmarks": {
        "SWE-bench Verified": null,
        "GPQA Diamond": null,
        "Terminal-Bench 2.1": null,
        "note": null,
        "HLE-Verified": 54.9
      },
      "sources": {
        "release": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/",
        "pricing": "https://ai.google.dev/gemini-api/docs/pricing",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash",
        "deprecations": "https://ai.google.dev/gemini-api/docs/deprecations",
        "benchmarks": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/"
      },
      "verifiedDate": "2026-09-03",
      "notes": "Gemini 3.8 Flash Cyber is a specialised variant of the same foundational model, available only through Google's Fairwind Program for trusted defenders. It is not a separately priced API model and has no entry in the pricing documentation, so benchr does not carry a record for it. Google states 3.7 Flash remains fully supported.",
      "arabic": {
        "name": "جيميني 3.8 فلاش",
        "status": "متاح للجميع منذ 2 سبتمبر 2026.",
        "notes": "نسخة Cyber متغيّر متخصص من النموذج نفسه، متاح فقط عبر برنامج Fairwind، وليس نموذجاً مسعَّراً منفصلاً في التوثيق. وتقول Google إن 3.7 فلاش ما زال مدعوماً بالكامل. السعر التمهيدي 0.75/3.75 دولار لكل مليون رمز حتى 31 ديسمبر 2026، ثم يتضاعف في 1 يناير 2027."
      }
    },
    {
      "id": "qwen-3-8-max",
      "name": "Qwen3.8-Max",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "qwen3.8-max",
      "license": null,
      "releaseDate": "2026-08-03",
      "status": "Available through Alibaba Cloud Model Studio.",
      "params": {
        "total": "2.4T",
        "active": "95B"
      },
      "context": {
        "windowTokens": 1000000,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 2.0,
        "outputPerM": 6.0,
        "note": "Singapore region rate on the Model Studio pricing table, read 2026-09-08. Alibaba publishes region-specific rates; other regions may differ."
      },
      "benchmarks": {
        "note": "Alibaba published arena placements rather than numeric scores at launch — fifth in Text Arena, second in Vision Arena, fourth in Frontend Code Arena. A placement is not a score, so every numeric field stays null."
      },
      "sources": {
        "release": "https://www.alibabacloud.com/blog/alibaba-unveils-qwen3-8-max-its-largest-and-most-capable-flagship-model-to-date_603420",
        "pricing": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "context": "https://www.alibabacloud.com/blog/alibaba-unveils-qwen3-8-max-its-largest-and-most-capable-flagship-model-to-date_603420"
      },
      "verifiedDate": "2026-09-08",
      "notes": "Alibaba announced Qwen3.8-Max on August 3, 2026 as its largest flagship: a sparse mixture-of-experts model with 2.4T total parameters activating 95B per token, multimodal across text and vision, with a context window the announcement describes as 'up to 1 million tokens'. The announcement said weights would follow the next week; benchr has not verified a weights release or a licence on an official page, so license stays null rather than being assumed. No maximum-output figure was published on either the announcement or the pricing table."
    },
    {
      "id": "qwen-3-8-flash",
      "name": "Qwen3.8-Flash",
      "provider": "Alibaba (Qwen)",
      "apiModelId": "qwen3.8-flash",
      "license": null,
      "releaseDate": "2026-08-27",
      "status": "Available through Alibaba Cloud Model Studio.",
      "params": {
        "total": "125B",
        "active": "6B"
      },
      "context": {
        "windowTokens": 262144,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 0.15,
        "outputPerM": 0.47,
        "note": "Singapore region rate on the Model Studio pricing table, read 2026-09-08. Alibaba publishes region-specific rates; other regions may differ."
      },
      "benchmarks": {
        "note": "Alibaba named SWE-bench Pro, CoWorkBench, Toolathlon Verified, MathVision, AndroidWorld and ERQA in the release but published no scores for them, so every numeric field stays null."
      },
      "sources": {
        "release": "https://www.alibabacloud.com/blog/alibaba-releases-qwen3-8-flash-with-innovative-model-architecture-delivering-optimal-price-performance_603503",
        "pricing": "https://www.alibabacloud.com/help/en/model-studio/model-pricing",
        "context": "https://www.alibabacloud.com/blog/alibaba-releases-qwen3-8-flash-with-innovative-model-architecture-delivering-optimal-price-performance_603503"
      },
      "verifiedDate": "2026-09-08",
      "notes": "Released August 27, 2026 as a multimodal mixture-of-experts model: a 125B main model with a further 51B of N-gram embeddings, activating 6B parameters per token. The announcement states a native 262K context extendable to 1 million tokens; benchr records the native figure, because the extended limit is a capability of a configuration rather than the default window. Alibaba's headline calls it open-weight but the announcement names no licence, and benchr has not verified one on an official page, so license stays null. No maximum-output figure was published."
    },
    {
      "id": "gpt-6-astra",
      "name": "GPT-6 Astra",
      "provider": "OpenAI",
      "apiModelId": "gpt-6-astra",
      "license": "proprietary",
      "releaseDate": "2026-09-03",
      "status": "Generally available since September 3, 2026. Listed on the official model directory as OpenAI's most capable model.",
      "context": {
        "windowTokens": 1050000,
        "maxOutputTokens": 128000
      },
      "pricing": {
        "inputPerM": 10.0,
        "outputPerM": 50.0,
        "cachedInputPerM": 1.0,
        "batchInputPerM": 5.0,
        "batchOutputPerM": 25.0,
        "batchCachedInputPerM": 0.5,
        "note": "Read on developers.openai.com/api/docs/pricing on 2026-09-08. Standard $10.00 input / $1.00 cached input / $50.00 output per 1M tokens; Batch is half of each. That is 2.5x GPT-5.6 Sol's post-cut input rate and 2.5x its output rate. No long-context, flex, fast or priority tier is published for this model, unlike Sol."
      },
      "benchmarks": {
        "note": "OpenAI published no benchmark table with the release announcement or on the model directory entry, so every numeric field stays null."
      },
      "sources": {
        "release": "https://developers.openai.com/api/docs/changelog",
        "pricing": "https://developers.openai.com/api/docs/pricing",
        "context": "https://developers.openai.com/api/docs/models/gpt-6-astra"
      },
      "verifiedDate": "2026-09-11",
      "notes": "Announced in OpenAI's API changelog on September 3, 2026 as \"our most capable model, built for the hardest end-to-end work\", aimed at reasoning, coding, computer use, research and document creation. The model directory lists a 1.05M context window and an April 30, 2026 knowledge cutoff; benchr recorded maximum output as null until OpenAI's model page, read on 2026-09-11, listed 128,000 max output tokens and a 922,000-token maximum input alongside the 1,050,000 context window. The same page lists reasoning effort levels low, medium, high, xhigh and max, and support for Chat Completions, Responses and Batch. The changelog records real interface constraints: the model does not support the `none` reasoning-effort level, does not accept custom temperature, top_p or logprobs, requires the Responses API for tool calling rather than Chat Completions, and is subject to misalignment monitoring. Those are migration blockers, not footnotes. benchr previously recorded that Astra was announced but unreleased with no model card, API id or price; that was accurate until this release and is superseded by this record."
    },
    {
      "id": "muse-spark-1-2",
      "name": "Muse Spark 1.2",
      "provider": "Meta",
      "apiModelId": "muse-spark-1.2",
      "license": "proprietary",
      "releaseDate": "2026-08-05",
      "status": "Available through Meta Model API and Muse Code. Meta's model table recommends Muse Spark 1.3 for new work.",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 1.25,
        "outputPerM": 4.25,
        "cachedInputPerM": 0.15,
        "contributorInputPerM": 0.1,
        "contributorOutputPerM": 0.2,
        "contributorCachedInputPerM": 0.002,
        "note": "Read on dev.meta.ai/docs/pricing-rate-limits on 2026-09-19. The contributor rates apply to a separate model ID, muse-spark-1.2-contributor, and buy the discount with permission to train on your prompts and completions. The standard tier is not used to train Meta models."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://developer.meta.com/ai/resources/blog/build-with-muse-code/",
        "pricing": "https://dev.meta.ai/docs/pricing-rate-limits",
        "context": "https://dev.meta.ai/docs/models"
      },
      "verifiedDate": "2026-09-19",
      "notes": "Meta introduced Muse Spark 1.2 alongside Muse Code in a post dated August 5, 2026, which announces the model as available that day. Meta's model table documents text, image, video, audio and PDF input with text output. Meta publishes no maximum output and no numeric benchmark table for it, so both stay empty. benchr added the record on 2026-09-19 while filling a gap in its Meta coverage; the model is not new.",
      "arabic": {
        "name": "ميوز سبارك 1.2",
        "status": "متاح عبر Meta Model API وMuse Code، ويوصي جدول Meta بـMuse Spark 1.3 للعمل الجديد.",
        "notes": "قدّمت Meta نموذج Muse Spark 1.2 مع أداة Muse Code في تدوينة مؤرّخة في 5 أغسطس 2026 تعلن إتاحته في اليوم نفسه. يوثّق جدول Meta إدخال النص والصورة والفيديو والصوت وملفات PDF، وإخراج النص، وسعة سياق 1,048,576 توكنًا. سعره القياسي 1.25 دولار للمدخلات و4.25 للمخرجات لكل مليون توكن، و0.15 للمدخلات المخزّنة، وثمة معرّف منفصل بسعر مخفّض يقايض الخصم بالإذن بالتدريب على مدخلاتك ومخرجاتك. ولا تنشر Meta حدًّا أقصى للإخراج ولا جدول اختبارات رقميًا. أضاف benchr السجل في 19 سبتمبر 2026 سدًّا لثغرة في تغطيته، والنموذج ليس جديدًا."
      }
    },
    {
      "id": "muse-spark-1-3",
      "name": "Muse Spark 1.3",
      "provider": "Meta",
      "apiModelId": "muse-spark-1.3",
      "license": "proprietary",
      "releaseDate": "2026-09-02",
      "status": "Available through Meta Model API and Muse Code since September 2, 2026, and the model Meta's own table recommends for new work.",
      "context": {
        "windowTokens": 1048576,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 1.25,
        "outputPerM": 4.25,
        "cachedInputPerM": 0.15,
        "contributorInputPerM": 0.1,
        "contributorOutputPerM": 0.2,
        "contributorCachedInputPerM": 0.002,
        "note": "Read on dev.meta.ai/docs/pricing-rate-limits on 2026-09-19. Unchanged from Muse Spark 1.2 and 1.1. The contributor rates apply to a separate model ID, muse-spark-1.3-contributor, and buy the discount with permission to train on your prompts and completions."
      },
      "benchmarks": {
        "note": "Meta's announcement shows a scorecard against Muse Spark 1.2, GPT-5.6 Sol (max) and Claude Opus 5 (max) across agent, coding, instruction-following and long-context evaluations, but publishes the numbers in a figure rather than a table benchr could read, so no score is recorded. Meta's stated efficiency claims — about 20% fewer tool calls and about 25% fewer tokens than 1.2 — are Meta's, not benchr's."
      },
      "sources": {
        "release": "https://research.meta.ai/blog/introducing-muse-spark-1-3",
        "pricing": "https://dev.meta.ai/docs/pricing-rate-limits",
        "context": "https://dev.meta.ai/docs/models"
      },
      "verifiedDate": "2026-09-19",
      "notes": "Meta released Muse Spark 1.3 on September 2, 2026, tuned for agentic work — multi-step tool use, browser use and long-horizon tasks — with coding improved over 1.2. Meta's model table documents text, image, video, audio and PDF input with text output. Meta's table notes audio support is limited in 1.3 and that 1.3 alone accepts a \"max\" reasoning effort, on the standard tier only. The price is unchanged from 1.2. Meta documents no maximum output. benchr has not assessed the model editorially, so it carries no ranking and reaches readers through the platform dialog.",
      "arabic": {
        "name": "ميوز سبارك 1.3",
        "status": "متاح عبر Meta Model API وMuse Code منذ 2 سبتمبر 2026، وهو النموذج الذي يوصي به جدول Meta نفسه للعمل الجديد.",
        "notes": "أطلقت Meta نموذج Muse Spark 1.3 في 2 سبتمبر 2026 مضبوطًا لعمل الوكلاء: استخدام الأدوات على خطوات متعددة، وتصفّح الويب، والمهام طويلة الأمد، مع تحسّن في البرمجة عن الإصدار 1.2. يوثّق جدول Meta إدخال النص والصورة والفيديو والصوت وملفات PDF وإخراج النص، ويشير إلى أن دعم الصوت محدود في 1.3، وأنه وحده يقبل مستوى استدلال «أقصى» على الطبقة القياسية. والسعر لم يتغيّر عن 1.2: 1.25 دولار للمدخلات و4.25 للمخرجات لكل مليون توكن. ولا توثّق Meta حدًّا أقصى للإخراج. أما مقارنات الأداء التي نشرتها فهي أرقام Meta في صورة لا جدول، فلم يسجّل benchr منها رقمًا، ولم يقيّم النموذج تحريريًا."
      }
    },
    {
      "id": "fugu-max",
      "name": "Fugu Max",
      "provider": "Sakana AI",
      "apiModelId": "fugu-max-v1.0",
      "license": "proprietary",
      "releaseDate": "2026-09-11",
      "status": "Available through Sakana's OpenAI-compatible API since September 11, 2026. Sakana states the Fugu models are not offered in the EU or the EEA.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 2.0,
        "outputPerM": 6.0,
        "cachedInputPerM": 0.25,
        "webToolCallPerCall": 0.007,
        "note": "Read on sakana.ai/fugu on 2026-09-19. One rate at every context length — Fugu Max has no long-context tier. Sakana bills its built-in web_search and web_fetch tools at $0.007 per call on top of tokens."
      },
      "benchmarks": {
        "note": "Sakana's release post says Fugu Max leads its comparison set on Terminal Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench and SWEFish, but publishes those results as a chart rather than a numeric table, so benchr records no figure. That is Sakana's claim, not a benchr test."
      },
      "sources": {
        "release": "https://sakana.ai/fugu-max-release/",
        "pricing": "https://sakana.ai/fugu/",
        "context": "https://console.sakana.ai/models"
      },
      "verifiedDate": "2026-09-19",
      "notes": "Sakana AI released Fugu Max on September 11, 2026 as the cost tier of its Fugu line. Fugu is an orchestrator rather than a single model: one request to one OpenAI-compatible endpoint, and Fugu routes the work across a pool of other models and assembles the answer. The API accepts fugu-max, which defaults to fugu-max-v1.0. Sakana publishes no context window and no maximum output for the Fugu models, so both stay null. benchr has not assessed the model editorially, so it carries no ranking and reaches readers through the platform dialog.",
      "arabic": {
        "name": "فوغو ماكس",
        "status": "متاح عبر واجهة Sakana المتوافقة مع OpenAI منذ 11 سبتمبر 2026، وتقول الشركة إن نماذج فوغو غير متاحة في الاتحاد الأوروبي والمنطقة الاقتصادية الأوروبية.",
        "notes": "أطلقت Sakana AI نموذج فوغو ماكس في 11 سبتمبر 2026 بوصفه الطبقة الاقتصادية في عائلة فوغو. وفوغو ليس نموذجًا واحدًا بل منسّق: يصل الطلب إلى نقطة واحدة متوافقة مع OpenAI، ثم يوزّع العمل على مجموعة من النماذج الأخرى ويجمع الإجابة. يقبل المعرّف fugu-max ويشير افتراضيًا إلى fugu-max-v1.0. سعره 2 دولار للمدخلات و6 للمخرجات لكل مليون توكن، و0.25 دولار للمدخلات المخزّنة، بسعر واحد مهما طال السياق. ولا تنشر Sakana سعة سياق ولا حدًّا أقصى للإخراج، فبقي الحقلان فارغين. ولم يقيّم benchr النموذج تحريريًا، فلا يحمل ترتيبًا ويصل القارئ إليه عبر نافذة المنصة."
      }
    },
    {
      "id": "fugu-ultra-v2",
      "name": "Fugu Ultra v2",
      "provider": "Sakana AI",
      "apiModelId": "fugu-ultra-v2.0",
      "license": "proprietary",
      "releaseDate": "2026-09-11",
      "status": "Available through Sakana's OpenAI-compatible API since September 11, 2026, and the default behind the fugu-ultra alias. Sakana states the Fugu models are not offered in the EU or the EEA.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": 5.0,
        "outputPerM": 30.0,
        "cachedInputPerM": 0.5,
        "longContextInputPerM": 10.0,
        "longContextOutputPerM": 45.0,
        "longContextCachedInputPerM": 1.0,
        "longContextThresholdTokens": 272000,
        "webToolCallPerCall": 0.007,
        "note": "Read on sakana.ai/fugu on 2026-09-19. The standard rate covers requests up to 272K tokens of context; above that every rate roughly doubles. Sakana bills its built-in web_search and web_fetch tools at $0.007 per call on top of tokens."
      },
      "benchmarks": {
        "Chartography": 48.3,
        "DeepSWE": 74.3,
        "note": "Both figures are Sakana's own published results for Fugu Ultra v2, read on its release post. Sakana reports Fugu Ultra v2 best or joint-best on five of eight benchmarks in that post and gives 27.3 for Claude Opus 5 and 29.5 for Claude Fable 5 on Chartography. benchr has run none of it."
      },
      "sources": {
        "release": "https://sakana.ai/fugu-max-release/",
        "pricing": "https://sakana.ai/fugu/",
        "context": "https://console.sakana.ai/models",
        "benchmarks": "https://sakana.ai/fugu-max-release/"
      },
      "verifiedDate": "2026-09-19",
      "notes": "Sakana AI released Fugu Ultra v2 on September 11, 2026 as the capability tier of its Fugu orchestration line, alongside Fugu Max. The API accepts fugu-ultra, which defaults to fugu-ultra-v2.0; v1.0 and v1.1 remain callable under their own IDs. Sakana gives the training cutoff as August 28, 2026. It publishes a long-context price tier that starts above 272K tokens but states no context window and no maximum output anywhere benchr could read, so both stay null rather than being inferred from the pricing threshold. benchr has not assessed the model editorially, so it carries no ranking and reaches readers through the platform dialog.",
      "arabic": {
        "name": "فوغو ألترا v2",
        "status": "متاح عبر واجهة Sakana المتوافقة مع OpenAI منذ 11 سبتمبر 2026، وهو النموذج الافتراضي خلف الاسم المختصر fugu-ultra. وتقول الشركة إن نماذج فوغو غير متاحة في الاتحاد الأوروبي والمنطقة الاقتصادية الأوروبية.",
        "notes": "أطلقت Sakana AI نموذج فوغو ألترا v2 في 11 سبتمبر 2026 بوصفه الطبقة الأقوى في عائلة فوغو، مع فوغو ماكس. يقبل المعرّف fugu-ultra ويشير افتراضيًا إلى fugu-ultra-v2.0، ويبقى الإصداران v1.0 وv1.1 متاحين بمعرّفيهما. سعره 5 دولارات للمدخلات و30 للمخرجات لكل مليون توكن حتى 272 ألف توكن من السياق، ثم يرتفع إلى 10 و45 فوق ذلك. وتذكر Sakana أن تدريبه ينتهي في 28 أغسطس 2026. ولا تنشر سعة سياق ولا حدًّا أقصى للإخراج، فبقي الحقلان فارغين ولم يُستنتجا من عتبة التسعير. أما أرقام الاختبارات فهي أرقام Sakana المنشورة، لم يجرِ benchr أيًّا منها."
      }
    },
    {
      "id": "gemini-3-8-live",
      "name": "Gemini 3.8 Live",
      "provider": "Google",
      "apiModelId": "gemini-3.8-live",
      "license": "proprietary",
      "releaseDate": "2026-09-15",
      "status": "GA (stable) since September 15, 2026, through the Live API.",
      "context": {
        "windowTokens": 131072,
        "maxOutputTokens": 65536
      },
      "pricing": {
        "textInputPerM": 0.75,
        "textOutputPerM": 4.5,
        "audioInputPerM": 3.0,
        "audioInputPerMinute": 0.005,
        "audioOutputPerM": 12.0,
        "audioOutputPerMinute": 0.018,
        "freeApiTier": true,
        "note": "Read on ai.google.dev/gemini-api/docs/pricing on 2026-09-19. Google prices audio two ways for this model and lists both: per million tokens and per minute. Text and audio are separate billing categories. A free tier is offered with the same features."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "pricing": "https://ai.google.dev/gemini-api/docs/pricing",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live"
      },
      "verifiedDate": "2026-09-19",
      "notes": "Google made Gemini 3.8 Live generally available on September 15, 2026 and calls it the default option for low-latency voice agents, without the delay that reasoning adds. It takes text, images, audio and video in and returns text and audio, and supports audio generation, function calling, search grounding and interleaved thinking. Google documents no caching, no code execution, no image generation, no structured outputs and no Batch API for it. The Live API is the interface. Google publishes no benchmark table for this model, so benchmarks stays empty.",
      "arabic": {
        "name": "جيميني 3.8 لايف",
        "status": "متاح للجميع منذ 15 سبتمبر 2026 عبر واجهة Live.",
        "notes": "أتاحت Google نموذج جيميني 3.8 لايف للجميع في 15 سبتمبر 2026، ووصفته بالخيار الافتراضي لوكلاء الصوت منخفضي الكمون، من دون التأخير الذي يضيفه الاستدلال. يقبل النص والصور والصوت والفيديو، ويخرج نصًا وصوتًا، ويدعم توليد الصوت واستدعاء الأدوات والإسناد إلى البحث والتفكير المتداخل. ولا توثّق Google له تخزينًا مؤقتًا ولا تنفيذ شيفرة ولا توليد صور ولا مخرجات منظمة ولا واجهة الدفعات. سعته 131,072 توكنًا للإدخال و65,536 للإخراج. ويُسعَّر النص بـ0.75 دولار للمدخلات و4.50 للمخرجات لكل مليون توكن، والصوت بـ3 دولارات للمدخلات و12 للمخرجات لكل مليون توكن، أو بالدقيقة. ولم تنشر Google جدول اختبارات لهذا النموذج."
      }
    },
    {
      "id": "gemini-3-8-live-extended-thinking",
      "name": "Gemini 3.8 Live Extended Thinking",
      "provider": "Google",
      "apiModelId": "gemini-3.8-live-extended-thinking",
      "license": "proprietary",
      "releaseDate": "2026-09-15",
      "status": "GA (stable) since September 15, 2026, through the Live API.",
      "context": {
        "windowTokens": 131072,
        "maxOutputTokens": 65536
      },
      "pricing": {
        "textInputPerM": 0.75,
        "textOutputPerM": 4.5,
        "audioInputPerM": 3.0,
        "audioInputPerMinute": 0.005,
        "audioOutputPerM": 12.0,
        "audioOutputPerMinute": 0.018,
        "freeApiTier": true,
        "note": "Read on ai.google.dev/gemini-api/docs/pricing on 2026-09-19. Google's pricing table lists this model at the same rates as Gemini 3.8 Live: the two are priced together, not inferred from one another."
      },
      "benchmarks": {},
      "sources": {
        "release": "https://ai.google.dev/gemini-api/docs/changelog",
        "pricing": "https://ai.google.dev/gemini-api/docs/pricing",
        "context": "https://ai.google.dev/gemini-api/docs/models/gemini-3.8-live-extended-thinking"
      },
      "verifiedDate": "2026-09-19",
      "notes": "Released to general availability alongside Gemini 3.8 Live on September 15, 2026. Google recommends it where a voice interaction needs heavier background reasoning, and documents three differences from the plain Live model: function calling is asynchronous only, and a synchronous blocking call returns an error; the thinking level must be low, medium or high, with minimal unsupported; and proactive audio is permanently on. Google publishes no benchmark table for it.",
      "arabic": {
        "name": "جيميني 3.8 لايف للتفكير الممتد",
        "status": "متاح للجميع منذ 15 سبتمبر 2026 عبر واجهة Live.",
        "notes": "صدر مع جيميني 3.8 لايف في 15 سبتمبر 2026. توصي Google به حين يحتاج الحوار الصوتي إلى استدلال أثقل في الخلفية، وتوثّق ثلاثة فروق عن النموذج العادي: استدعاء الأدوات غير متزامن فقط، والاستدعاء المتزامن يعيد خطأ؛ ومستوى التفكير لا بد أن يكون منخفضًا أو متوسطًا أو مرتفعًا، ولا يقبل الحد الأدنى؛ والصوت الاستباقي مفعَّل دائمًا. وسعره في جدول Google هو سعر جيميني 3.8 لايف نفسه، مذكورًا له لا مستنتجًا منه. ولم تنشر Google جدول اختبارات لهذا النموذج."
      }
    },
    {
      "id": "grok-voice-transcribe-2",
      "name": "Grok Voice Transcribe 2.0",
      "provider": "xAI",
      "apiModelId": "grok-voice-transcribe-2.0",
      "license": "proprietary",
      "releaseDate": "2026-09-18",
      "status": "Available in the xAI Speech-to-Text API since September 18, 2026. grok-voice-transcribe-1.0 is still the default; xAI says 2.0 will become the default soon.",
      "context": {
        "windowTokens": null,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null,
        "batchPerAudioHour": 0.1,
        "streamingPerAudioHour": 0.2,
        "note": "Read on docs.x.ai/docs/models and x.ai/news/grok-voice-transcribe-2 on 2026-09-19. Speech-to-text is billed per hour of audio, not per token: $0.10 an hour over REST and $0.20 an hour streaming. xAI states the price is unchanged from Grok Voice Transcribe 1.0."
      },
      "benchmarks": {
        "note": "xAI says 2.0 is twice as accurate as 1.0 at the same price, that word error rate falls from 20.6% to 6.8% on its short-phrase multilingual sets, and that the model ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. Those are xAI's claims and a third-party leaderboard's placement, not benchr measurements."
      },
      "sources": {
        "release": "https://x.ai/news/grok-voice-transcribe-2",
        "pricing": "https://docs.x.ai/docs/models",
        "context": "https://docs.x.ai/developers/release-notes"
      },
      "verifiedDate": "2026-09-19",
      "notes": "xAI announced Grok Voice Transcribe 2.0 on September 18, 2026 as a speech-to-text model, available beside grok-voice-transcribe-1.0 rather than replacing it. xAI describes support for dozens of languages, automatic language detection, and switching language mid-recording, and says the largest gain over 1.0 is multilingual accuracy, with telephony audio at 8 kHz its strongest result. This is a transcription model, so a context window and a maximum output do not apply and stay null.",
      "arabic": {
        "name": "جروك فويس ترانسكرايب 2.0",
        "status": "متاح في واجهة تحويل الكلام إلى نص لدى xAI منذ 18 سبتمبر 2026، ولا يزال الإصدار 1.0 هو الافتراضي، وتقول الشركة إن 2.0 سيصبح الافتراضي قريبًا.",
        "notes": "أعلنت xAI نموذج جروك فويس ترانسكرايب 2.0 في 18 سبتمبر 2026 لتحويل الكلام إلى نص، وأتاحته إلى جانب الإصدار 1.0 لا بديلًا عنه. تذكر الشركة دعم عشرات اللغات، والتعرّف التلقائي على اللغة، والتبديل بينها أثناء التسجيل، وتقول إن أكبر تحسّن عن 1.0 هو دقة اللغات المتعددة، وإن أقوى نتائجه على صوت الهاتف بتردد 8 كيلوهرتز. الفوترة بساعة الصوت لا بالتوكن: 0.10 دولار للساعة عبر REST و0.20 للبث، وهو سعر الإصدار السابق نفسه. وما تذكره xAI عن مضاعفة الدقة وعن ترتيبه الأول بين 32 نموذج بث هو ادعاء الشركة ولوحة ترتيب خارجية، لا قياس أجراه benchr."
      }
    },
    {
      "id": "atria-dawn-preview",
      "name": "Atria Dawn Preview",
      "provider": "Shanghai AI Laboratory",
      "apiModelId": "internlm/Atria-Dawn-Preview",
      "license": "MIT",
      "releaseDate": null,
      "status": "Open weights published on the Shanghai AI Laboratory's own Hugging Face organisation. A preview release, with no dated vendor announcement and no published price.",
      "params": {
        "total": "753B",
        "active": null
      },
      "context": {
        "windowTokens": 262144,
        "maxOutputTokens": null
      },
      "pricing": {
        "inputPerM": null,
        "outputPerM": null,
        "note": "No price. The model card offers weights under MIT for self-hosting via SGLang or vLLM and mentions hosted endpoints, but states no rate. Read on huggingface.co/internlm/Atria-Dawn-Preview on 2026-09-19."
      },
      "benchmarks": {
        "BrowseComp": 92.5,
        "SWE-bench Pro": 59.6,
        "BFCL v4": 77.0,
        "AutomationBench": 53.8,
        "CyberGym": 86.5,
        "note": "Every figure is the vendor's own, read from the model card. The card also reports DeepSearchQA 96.0, MLE-bench Lite 86.2, Workspace-Bench 65.0 and JobBench 50.3. benchr has run none of them."
      },
      "sources": {
        "release": "https://huggingface.co/internlm/Atria-Dawn-Preview",
        "context": "https://huggingface.co/internlm/Atria-Dawn-Preview",
        "benchmarks": "https://huggingface.co/internlm/Atria-Dawn-Preview"
      },
      "verifiedDate": "2026-09-19",
      "notes": "An agentic model from the Shanghai Artificial Intelligence Laboratory, published as open weights under MIT on the laboratory's internlm Hugging Face organisation. The card gives 753B total parameters in a mixture-of-experts architecture built on a 744B foundation model, a 256K context window, and text-only input. releaseDate is null because the laboratory published no dated announcement: atria-asi.ai carries no post, and the card itself gives no release date. benchr does not date a launch from a repository timestamp or a third-party tracker. The maximum output and the active-parameter count are null for the same reason.",
      "arabic": {
        "name": "أتريا داون بريفيو",
        "status": "أوزان مفتوحة منشورة على حساب مختبر شنغهاي للذكاء الاصطناعي في Hugging Face. إصدار تمهيدي، بلا إعلان مؤرّخ من الجهة المطوّرة وبلا سعر منشور.",
        "notes": "نموذج وكيل من مختبر شنغهاي للذكاء الاصطناعي، نُشرت أوزانه مفتوحةً برخصة MIT على حساب المختبر internlm في Hugging Face. تذكر بطاقة النموذج 753 مليار معامل في بنية خليط خبراء مبنية على نموذج أساس حجمه 744 مليارًا، وسعة سياق 256 ألف توكن، ومدخلات نصية فقط. تُرك تاريخ الإصدار فارغًا لأن المختبر لم ينشر إعلانًا مؤرّخًا: موقع atria-asi.ai خالٍ، والبطاقة نفسها لا تذكر تاريخًا، وbenchr لا يؤرّخ إطلاقًا من طابع زمني لمستودع ولا من متتبّع خارجي. وللسبب نفسه بقي حدّ الإخراج وعدد المعاملات النشطة فارغين. وأرقام الاختبارات كلها أرقام البطاقة، لم يُجرِ benchr أيًّا منها."
      }
    }
  ]
}
