02-347-7730  |  Saeree ERP - Complete ERP System for Thai Businesses Contact Us

Claude vs OpenAI vs Chinese AI, September 2026: How Far Apart Are the Scores? Opus 5.5, GPT-6 Astra, MiMo, GLM, Kimi and DeepSeek on Five Independent Benchmarks

  • Home
  • Articles
  • Claude vs OpenAI vs Chinese AI: Benchmark Scores
Claude vs OpenAI vs Chinese AI, September 2026: How Far Apart Are the Scores? Opus 5.5, GPT-6 Astra, MiMo, GLM, Kimi and DeepSeek on Five Independent Benchmarks
  • 30
  • September

"Claude vs OpenAI vs Chinese AI, September 2026: How Far Apart Are the Scores? Opus 5.5, GPT-6 Astra, MiMo, GLM, Kimi and DeepSeek on Five Independent Benchmarks" — the short answer is as of 30 September 2026, on independent yardsticks, Claude Opus 5.5 leads two of them (Artificial Analysis at 58 and Arena at 1,509), GPT-6 Astra leads the other two (Epoch ECI at 166.6 and ARC-AGI-2 at 95%), and the best Chinese open-weight models, MiMo-V2.6-Pro (46), GLM-5.3 (45) and Kimi K3 (44), trail the top by 12 points on Artificial Analysis, about 9 points on Epoch and 21–30 Elo on Arena, which works out to roughly 4–8 months depending on the yardstick. The gap is almost nil on knowledge exams, very wide on abstract reasoning, and China wins outright on cost per task, by as much as 40 times. This article follows our Claude vs OpenAI, August 2026 comparison and pulls the data from our September China AI update and OpenAI update into one set of tables.

In one line: the two American labs trade the lead depending on the yardstick, China trails by 4–8 months but costs 10–40 times less, and every number this month has a short shelf life.

  • Who leads: Opus 5.5 is #1 on Artificial Analysis (58) and Arena (1,509) · GPT-6 Astra is #1 on Epoch ECI (166.6) and ARC-AGI-2 (95.0%) · Sonnet 5.5 is #1 on Terminal-Bench 4.0 as run by Artificial Analysis (63.6%)
  • Where China sits: MiMo-V2.6-Pro 46 · GLM-5.3 45 · Qwen3.8 Max 45 · Kimi K3 44 · DeepSeek V4.1-Flash 39 on the Artificial Analysis index · Kimi K3 is the highest Chinese model on Epoch (157.68)
  • How far: 12 Artificial Analysis points · 8.9 Epoch points (about 7–8 months at the rate the frontier moves) · 21–30 Arena Elo · but about 34 points on ARC-AGI-2, and no Chinese model has a verified ARC-AGI-3 score yet
  • Cost per task: MiMo-V2.6-Pro at $0.13 per task on the Artificial Analysis index against $5.98 for Opus 5.5 and $7.63 for Fable 5.1
  • Caveats: the Artificial Analysis index changed formula twice in September · SWE-bench Verified has been dropped by OpenAI and Anthropic · there is no 2026 Thai-language benchmark for any of these models

One note for the whole article: facts checked on 30 September 2026 · figures marked "independent" come from outside evaluators (Artificial Analysis, Arena, Epoch AI, ARC Prize, Scale, NIST CAISI); figures marked "vendor" come from the model maker's own announcement · prices are US dollars per 1 million tokens from official pricing pages, not converted to baht · where a model has no score on a yardstick we write "n/a" rather than estimate.

1. Before the Numbers: Five Yardsticks Measure Different Things, and This Month the Yardsticks Themselves Changed

"Which model is better" has no single answer, because each yardstick measures something different. This article leans on yardsticks with an outside evaluator and uses vendor-reported figures only where no independent measurement exists.

YardstickWho measuresWhat it measuresWeakness to know
Artificial Analysis Intelligence Index (v4.3.2)Artificial Analysis, run in-house on 216 modelsA composite of many evaluations covering reasoning, coding, agentic work and knowledge, with private test sets weighted at 45% to deter training to the testThe formula changed twice in September (4 and 19 Sep); scores before and after are not comparable
Arena (formerly LMArena) Text and WebDevReal users vote on paired answers without seeing model namesHuman preference, as an Elo scoreMeasures what people like, not what is correct · new models need votes to accumulate; GPT-6 Astra was added on 11 Sep and is not yet in the Text top 25
Epoch Capabilities Index (ECI)Epoch AIFolds results from many benchmarks into one scale comparable across time; the frontier moves about 14 points a yearModels released late in the month (Opus 5.5, Sonnet 5.5, GPT-6 Sol) have no score yet
ARC-AGI-2 / ARC-AGI-3 (Semi-Private sets)ARC PrizeAbstract reasoning on puzzles the model has never seen, with cost per taskHighly harness-dependent: GPT-6 Astra scores 62.7% or 99.95% on ARC-AGI-3 depending on whether OpenAI's own adapter is used
Terminal-Bench 4.0Artificial Analysis in-house, and vendor self-reportsAgentic tasks in a terminal from start to finish; a new set far harder than the 2.x seriesVendor figures run 6–7 points above independent runs · Chinese models have self-reported figures only

Why this month's scores look lower than last month's when the models did not change

Artificial Analysis rebuilt its index as v4.2 on 4 September and again as v4.3.2 on 19 September, replacing Terminal-Bench 2.1, where every lab scores 86–90 and nothing separates them, with Terminal-Bench 4.0, and replacing τ³-Banking with AutomationBench. The result: Claude Fable 5.1 went from 66 to 53, Kimi K3 from 57 to 44, Qwen3.8 Max from 56 to 45 and GPT-6 Astra from 61.2 to 53 with no change to the models. If the figures in our August article do not match this table, that is why.

Using Claude at work? Get a quote in Thai Baht with a full tax invoice

Our procurement service starts at 5 seats · fewer than that? buy direct from Anthropic · Enterprise: talk to our team

2. The Big Table: 14 Models on Five Independent Yardsticks

Sorted by Artificial Analysis score, except Gemini 3.8 Flash, which sits at the bottom as Google's reference point. Numbers in parentheses are ranks. "Vendor" means the figure was reported by the model's maker.

Model (lab · release)WeightsArtificial AnalysisArena Text (Elo)Epoch ECIARC-AGI-2Terminal-Bench 4.0Price $/1M (input / output)
Claude Opus 5.5 (Anthropic · 22 Sep)Closed58 (#1)1,509 (#1)n/a93.3%59.6% independent · 66.4% vendor$4 / $20
Claude Sonnet 5.5 (Anthropic · 28 Sep)Closed56 (#3)Not yet rankedn/an/a63.6% independent (#1) · 70.6% vendor$2 / $10
Claude Fable 5.1 (Anthropic · 1 Sep)Closed53 (#5)1,501 (#5)165.090.0%52% independent · 55.8% vendor$10 / $50
GPT-6 Astra (OpenAI · 3 Sep)Closed53 (#7)Not in top 25 yet166.6 (#1)95.0% · ARC-AGI-3 62.7%59% independent · 57.7% vendor$10 / $50
GPT-6 Sol (OpenAI · 22 Sep)Closed48 (#20)Not in top 25 yetn/an/an/a$2 / $10
GPT-5.6 Sol (OpenAI · Jul)Closed47 (#21)1,483 (#19)161.9992.5%37.3% vendor$4 / $20
MiMo-V2.6-Pro (Xiaomi · 22 Sep)Open, MIT46 (#1 open-weight)1,480 (#23)n/an/a34.9% vendor$0.43 / $0.87
GLM-5.3 (Zhipu · 14 Aug)Open451,480 (#24)155.56n/an/a$1.40 / $4.40
Qwen3.8 Max (Alibaba · 3 Aug, updated 2 Sep)Closed (open release announced)45 (#27)1,479 (#25)156.69n/an/a$2 / $6
Kimi K3 (Moonshot · 16 Jul)Open (custom license)441,488 (#16)157.68 (highest Chinese)60.4%n/a$3 / $15
DeepSeek V4.1-Flash (DeepSeek · 10 Sep)Open, MIT39Not in top 25 yet155.01n/a31.2% vendor$0.30 / $1.20 peak
DeepSeek V4-Pro (DeepSeek · 13 Aug)Open, MIT36Not in top 25155.3961.3%12.4% vendor$1.32 / $3.96 peak
MiniMax M3 (MiniMax · 1 Jun)Open29Not in top 25147.0n/an/a$0.30 / $1.20
Gemini 3.8 Flash (Google · 2 Sep, reference)Closed41 (#43)1,492 (#10)157.1389.2% · ARC-AGI-3 35.0%19.1% vendor$0.75 / $3.75

Three things the table says at once. First, the two American labs trade the lead: Anthropic leads on the composite score and human preference, OpenAI on Epoch's cross-time index and abstract reasoning. Second, Google currently sits below Kimi K3 and GLM-5.3 on Artificial Analysis because its flagship Gemini 4 is still in training. Third, the four Chinese models cluster at 44–46, only 2 points apart, whereas the American top and second tiers are 5–10 points apart. Model details are in our articles on Opus 5.5, Sonnet 5.5, Fable 5.1, GPT-6 Astra, Kimi and DeepSeek.

3. How Far Apart: The US–China Gap on Each Yardstick, and in Months

The reader's question is "how far apart". This table subtracts the best Chinese model from the leader on each yardstick, and converts to time only where the evaluator publishes a conversion.

YardstickLeaderBest ChineseGapNotes
Artificial Analysis IndexOpus 5.5 = 58MiMo-V2.6-Pro = 4612 pointsBefore Opus 5.5 (22 Sep) the gap was 53 to 45, or 8–9 points, the figure used in our China AI update
Epoch ECIGPT-6 Astra = 166.6Kimi K3 = 157.688.9 pointsEpoch measures the frontier moving about 14 points a year, so this is roughly 7–8 months
Arena TextOpus 5.5 = 1,509Kimi K3 = 1,48821 EloMiMo and GLM-5.3 at 1,480 are 29–30 Elo behind · Stanford HAI measured the US–China Arena gap at 2.7% in March 2026
Arena WebDev (human-voted web code)Opus 5.5 = 1,820Qwen3.8 Max = 1,671149 EloKimi K3 1,659 · GLM-5.3 1,623 · DeepSeek V4.1-Flash 1,620
ARC-AGI-2Astra 95.0% / Opus 5.5 93.3%DeepSeek V4-Pro 61.3% / Kimi K3 60.4%about 34 pointsThe widest gap on any yardstick
ARC-AGI-3Astra 62.7%No Chinese model has a verified score—Opus 5.5 has not been run either · Gemini 3.8 Flash 35.0% · Opus 5 30.2%
Terminal-Bench 4.0 (vendor self-reports)Sonnet 5.5 70.6% / Opus 5.5 66.4%MiMo 34.9% / DeepSeek V4.1-Flash 31.2%about 32–36 pointsVendor-to-vendor comparison only, since there is no independent run for the Chinese models
GPQA Diamond (PhD-level knowledge)Astra 96.0%Kimi K3 93.5%2.5 pointsEvery lab sits at 90–96; this set no longer separates models

The organizations that estimate the gap in time reach similar answers by different routes.

EvaluatorFindingDate
Epoch AIChinese models trail the US frontier by 7 months on average (range 4–14 months)2 Jan 2026
Epoch AIOpen-weight models trail closed models by about 4 months, or 8 ECI points29 May 2026
NIST CAISI (US)DeepSeek V4 Pro lags the frontier by about 8 months but is more cost-efficient on 5 of 7 benchmarks1 May 2026
Mozilla — State of Open Source AI v1.1The best open-weight models trail the closed frontier by 4.4 months · closed models keep a 92-Elo lead on expert knowledge work (GDPval-AA) · "the gap resets every release cycle"15 Sep 2026
InterconnectsChinese open-weight models are 2–5 months behind the closed American frontier, American open-weight models 6–9 months behind · Chinese models take more than 80% of OpenRouter usage21 Sep 2026
Stanford HAI — AI Index 2026US–China gap on Arena 2.7% · closed–open gap 3.3% (from 0.5% in Aug 2024) · the two countries "traded the lead multiple times since early 2025"Data through Mar 2026

4. By Task Type: The Gap Is Not the Same Width Everywhere

The composite hides one fact: the gap is very wide for some kinds of work and nearly absent for others. An organization that knows where its own work sits in this table can choose a model without overpaying.

Task typeMetricUS sideChinese sideReading
Advanced knowledge, academic Q&AGPQA Diamond (vendor)Astra 96.0 · GPT-5.6 Sol 94.6 · Fable 5.1 93.7Kimi K3 93.5 · Qwen3.8 Max 92.6 · MiniMax M3 92.7 (independent) · DeepSeek V4.1-Flash 90.9Nearly equal
Hard multi-discipline examHumanity's Last Exam with tools (vendor)Opus 5.5 67.7 · Fable 5.1 65.0 · Sonnet 5.5 64.5 · Astra 57.2GLM-5.3 62.5 · Qwen3.8 Max 56.2 · Kimi K3 56.0China close, but vendor figures disagree; Scale's independent run gives Astra 54.8 and Fable 5.1 46.5
Human-voted web codingArena WebDev (independent)Opus 5.5 1,820 · Astra 1,792 · Fable 5.1 1,753 · Sonnet 5.5 1,699Qwen3.8 Max 1,671 · Kimi K3 1,659 · Hy4 1,633 · GLM-5.3 1,623150–200 Elo apart
Agent finishing work in a terminalTerminal-Bench 4.0Sonnet 5.5 63.6 · Opus 5.5 59.6 · Astra 59 (independent)MiMo 34.9 · DeepSeek V4.1-Flash 31.2 (vendor)Very wide
Abstract reasoning on unseen puzzlesARC-AGI-2 (independent)Astra 95.0 · Opus 5.5 93.3 · Fable 5.1 90.0DeepSeek V4-Pro 61.3 · Kimi K3 60.4Widest
Expert-level office workGDPval-AA (Elo)Opus 5.5 1,846 · Sonnet 5.5 1,844 (vendor)Mozilla finds closed models 92 Elo aheadModerately wide
Real-world usage volumeTokens on OpenRouter—Chinese models above 80% of tokens (Interconnects) · 7 of the top 10China leads outright, on price

The picture is clear. On work that needs "knowledge", China has caught up. On work that needs "multi-step thinking without a human in the loop", both terminal agents and abstract puzzles, the gap is still 30 points or more. Zhipu's own tables show the same pattern: GLM-5.3 edges Mythos 5 on CyberGym (84.5 to 83.8) but trails badly on ExploitBench (54.4 to 78.0).

5. Price per Unit of Intelligence: The Table China Wins Outright, and Why List Prices Mislead

Artificial Analysis computes "cost per task" from the tokens a model actually used while running the index, not from the list price. That figure tracks a real bill more closely, and in several cases it points the other way from the list price.

ModelArtificial Analysis IndexCost per index task (independent)List price $/1M (input / output)
MiMo-V2.6-Pro46$0.13$0.43 / $0.87
DeepSeek V4.1-Flash39$0.27$0.30 / $1.20 peak
MiniMax M329$0.51$0.30 / $1.20
GPT-6 Sol48$1.05$2 / $10
Gemini 3.8 Flash41$1.24$0.75 / $3.75
GPT-5.6 Sol47$1.99$4 / $20
Kimi K344$2.00$3 / $15
GLM-5.345$2.01$1.40 / $4.40
GPT-6 Astra53$3.26$10 / $50
Qwen3.8 Max45$5.41$2 / $6
Claude Opus 5.558$5.98$4 / $20
Claude Sonnet 5.556$7.60$2 / $10
Claude Fable 5.153$7.63$10 / $50

Three places where the list price misleads. First, Sonnet 5.5 lists at half the price of Opus 5.5 but costs more per index task, because it generated 410 million output tokens during the run against 260 million for Opus 5.5 (Artificial Analysis also notes that Sonnet 5.5 scores lower at max effort than at xhigh). Second, Qwen3.8 Max lists cheaper than Kimi K3 but costs 2.7 times more per task, because it answers at length and slowly. Third, GPT-6 Astra lists at exactly Fable 5.1's price in every column but costs less than half per task, using about 27,000 output tokens per task against 78,000. One more thing to know: Anthropic's pricing page states that models from the 4.7 generation onward use a tokenizer that "produces approximately 30% more tokens for the same text" than Sonnet 4.6 and earlier. Cross-vendor price comparisons should always be made at cost per task.

6. Six Caveats Before Acting on Any of These Numbers

CaveatEvidence
1. SWE-bench Verified is retiredOpenAI stopped reporting it on 23 Feb 2026 because every frontier model "had seen at least some of the problems and solutions during training", and Epoch AI rates the set "Flawed" (3 Sep); Anthropic did not report it for Opus 5.5 · Chinese vendors still headline it (DeepSeek V4 Pro 80.6, MiniMax M3 80.5), and the same model reads 80.6 (vendor), 74 (CAISI) or 96.4 (Vals) depending on checkpoint and harness
2. The harness moves scores more than the model doesAstra scores 62.7% or 99.95% on ARC-AGI-3 depending on the adapter · Fable 5.1 scores 77.9% on OSWorld 2.0 partial but 41.7% strict · vendor Terminal-Bench 4.0 figures (Opus 5.5 66.4 at xhigh, Sonnet 5.5 70.6) sit 6–7 points above Artificial Analysis's runs (59.6, 63.6)
3. Benchmarks saturate fast, so indices keep changing formulaStanford HAI: "evaluations intended to be challenging for years are saturated in months"; HLE alone gained 30 points in a year · GPQA sits at 90–96 for everyone · Terminal-Bench 2.1 sits at 86–90 for everyone, so Artificial Analysis switched to 4.0 and raised private-set weighting to 45% "to prevent gaming"
4. Vendor figures disagree with each other and with independentsGPT-5.6 Sol's HLE is 58.0 in Kimi's table but 64.5 in Zhipu's · Kimi K3's HLE is 56.0 in its own card but 59.8 in Zhipu's · Scale's independent HLE gives Astra 54.8 and Fable 5.1 46.5 against vendor 57.2 and 65.0 · CAISI found DeepSeek's self-reported results "stronger than CAISI's independent testing" on benchmarks outside its technical report
5. The real gap may be wider than measuredEpoch AI warns that open-weight models "tend to optimize on benchmarks more aggressively", so the gap "could be understated" · Interconnects notes that closed labs "sit on their strongest internal systems and release conservatively" · China's cost-per-task advantage, by contrast, is real and measured
6. There is no 2026 Thai-language benchmark for these modelsSEA-HELM (18 Sep 2026) evaluates Thai but lists only mid-size open models, not Opus 5.5, GPT-6 or Kimi K3 · the latest hard Thai-exam numbers for frontier models date from 2024 (Claude 3.5 Sonnet 68.41% vs GPT-4o 66.26%, from the OpenThaiGPT 1.5 paper) · ask for the source before believing any Thai score claimed for a new model

How Should Thai Organizations Use These Numbers?

Five points for anyone choosing a model this quarter

1. Pick the yardstick that matches the work before picking the model — for knowledge Q&A look at GPQA and HLE, where China is close · for agents that run multiple steps on their own look at Terminal-Bench 4.0 and ARC-AGI-2, where the gap is still 30 points or more · for front-end code look at Arena WebDev.

2. Test with 20–50 of your own tasks in Thai — nobody measures these models in Thai for you; use real documents and real spreadsheets, and have someone who knows the work check the output. That is the most trustworthy benchmark your organization will get.

3. Think in cost per task, not list price — measure the tokens actually used per task in the test above and multiply by the price. A cheap-looking model may answer at such length that it costs more (Qwen3.8 Max), and an expensive-looking one may answer briefly enough to cost less (Astra versus Fable 5.1).

4. Read the 4–8 month gap correctly — it means what the top model can do today, an open model at one-tenth the price will do in the first half of next year. Work that is not urgent and has a clear scope can wait; work that needs the best capability now still pays top-tier prices.

5. Design for switching — the top of this table changed three times in one month (Astra on 3 Sep, Opus 5.5 on 22 Sep, Sonnet 5.5 on 28 Sep), and Qwen 4, Haiku 5.5 and Gemini 4 are on the way. Your integration layer must switch models with a single configuration change.

Conclusion

At the end of September 2026 no lab leads on every yardstick. Anthropic leads on the composite score and human preference, OpenAI on the cross-time index and abstract reasoning, and China trails by 4–8 months while leading by a wide margin on cost per task. That gap is very narrow on knowledge work and very wide on agentic work. The organizations that decide well are the ones that know where their own work sits in the table, test with real Thai tasks, and do not tie their systems to this month's winner.

A 12-point gap tells you who leads this month. A cost-per-task difference of more than 40 times tells you which work should go to whom. The two numbers do not contradict each other; they answer different questions.

- Paitoon Butri · Network & Server Security Specialist, Grand Linux Solution Co., Ltd.

References

Want Claude Opus 5.5 or Sonnet 5.5 for your organization, with a Thai tax invoice?

Grand Linux supplies Claude Team (from 5 seats) and Enterprise (from 20 seats) with quotes in Thai baht, tax invoices and purchase-order support, plus integration of Claude into your back-office systems over MCP (Model Context Protocol) with role-based permissions and usage logs.

Get advice / request a quote

Tel 02-347-7730 | sale@grandlinux.com

Saeree ERP Author

About the Author

Paitoon Butri

Network & Server Security Specialist, Grand Linux Solution Co., Ltd.