Skip to content

AGI Companies Ranked by What They Have Actually Shipped

October 3, 2026. Every list of "AGI companies" ranks them by valuation, by headcount or by how often the founder says the word. This one ranks them by the only thing you can check: what each lab has released to the public and how it scores on an independent test nobody controls. The test is ARC-AGI, run by the non-profit ARC Prize, and the measure is each lab's best verified score on ARC-AGI-2, with the interactive ARC-AGI-3 as the tiebreaker, all read from the live leaderboard today (ARC Prize leaderboard, read 3 October 2026). By that measure the order is OpenAI, Anthropic, Google, a small lab called Dots Studio, xAI, Zhipu, DeepSeek, Moonshot, Alibaba and Thinking Machines, with Meta unranked because its current models have not been scored and Safe Superintelligence Inc. at the bottom with nothing to score. The rest of this page explains the method, profiles each lab by what it has shipped, and ends with what the ranking means if you are buying rather than building.

AGI companies ranked by what they have shipped: ARC-AGI-2 scores and releases for OpenAI, Anthropic, Google, xAI, DeepSeek, Meta and SSI
Bar chart of best ARC-AGI-2 scores by lab: OpenAI 95.0, Anthropic 93.3, Google 89.2, Dots Studio 76.8, xAI 67.1, Zhipu 65.8, DeepSeek 61.4, Moonshot 60.4, Alibaba 42.4
ARC Prize leaderboard, read 3 October 2026

Key numbers

ItemNumber
OpenAI GPT-6 Astra (Max), ARC-AGI-2 (September 2, 2026; $1.12 per task)95.0 percent
OpenAI GPT-6.1 Sol (Max), ARC-AGI-2 (September 29, 2026; $0.254 per task)94.2 percent
Anthropic Claude Opus 5.5 (High), ARC-AGI-2 (September 22, 2026; $0.408 per task)93.3 percent
Google Gemini 3.8 Flash (High), ARC-AGI-2 (September 2, 2026; $0.400 per task)89.2 percent
Dots Studio Dots3-Note Preview, ARC-AGI-2 (August 14, 2026; open-weight record; $0.077 per task)76.8 percent
xAI Grok 4.6 (XHigh), ARC-AGI-2 (August 11, 2026; $0.757 per task)67.1 percent
DeepSeek V4 Flash 0731 (Max), ARC-AGI-2 (July 31, 2026; $0.042 per task, cheapest above 60)61.4 percent
Best ARC-AGI-3 standard-harness score (GPT-6 Astra; Anthropic Opus 5 30.2, Gemini 3.8 Flash 10.4, Grok 4.6 2.1)62.7 percent
Meta and SSI models with an ARC-AGI-2 score above zero (Muse not scored; SSI has released nothing)0

ARC Prize leaderboard and results pages, OpenAI, Google, xAI, Meta and ssi.inc, all read 3 October 2026.

How this ranking works

An "AGI company" here is a lab that states artificial general intelligence or superintelligence as its goal: OpenAI's charter commits it to AGI, "highly autonomous systems that outperform humans at most economically valuable work" (OpenAI charter); Anthropic, Google DeepMind, Meta Superintelligence Labs, xAI and SSI say equivalent things. Our explainer on what AGI means covers the definitions. The ranking uses three rules. First, only publicly available or publicly verified systems count; a demo or a blog claim does not. Second, the primary score is ARC-AGI-2, a test of fluid reasoning on which a human panel scores 100 percent, because it is the broadest independent test that most labs have submitted to. Third, ARC-AGI-3, released in 2026, breaks ties, using the standard harness scores because the alternative "provider adapter" harness preserves the model's reasoning state between steps and produces numbers that are not comparable across labs. Costs are ARC Prize's cost per task at list prices. Chat-arena rankings, revenue and valuations are deliberately excluded.

1. OpenAI

Best ARC-AGI-2 score: 95.0 percent with GPT-6 Astra (Max), released September 2, 2026, at $1.12 per task; GPT-6.1 Sol (Max), released September 29, scores 94.2 percent at $0.254 per task. On ARC-AGI-3, Astra holds the top standard-harness score at 62.7 percent, and OpenAI states that under the provider adapter setup Astra "saturates ARC-AGI-3 with a 99.9% score" (OpenAI). What it has shipped: two model generations in a month, both sold through an API at $10 and $50 per million input and output tokens for Astra and $2 and $10 for Sol, plus always-on "dots" agents launched at DevDay on September 29. What it withheld: GPT-6.1 Astra, which OpenAI decided not to release after testing found problems with staying within scope and authorisation; see why GPT-6.1 Astra has no release date. OpenAI leads on every measure this ranking uses, and it is also the lab whose 2026 incident reports are the most detailed; the OpenAI incident report is the companion reading.

2. Anthropic

Best ARC-AGI-2 score: 93.3 percent with Claude Opus 5.5 (High), released September 22, 2026, at $0.408 per task, the cheapest score above 90 on the board apart from GPT-6.1 Sol. Claude Fable 5.1 (Max), released September 1, scores 90.0 percent at $4.49. Anthropic has no published ARC-AGI-3 score for Opus 5.5 or Fable 5.1; Claude Opus 5 (High), from July, scores 30.2 percent under the standard harness, second to OpenAI. What it has shipped: Opus 5.5 at $4 and $20 per million tokens, Fable 5.1 at $10 and $50, and the restricted Mythos tier of the same model for vetted users, plus a system card per release that runs to more than 200 pages; Claude Opus 5.5 explained and what Claude Mythos is cover the lineup. Anthropic is the only lab in this list that publishes model welfare assessments alongside capability results; see is AI conscious.

3. Google DeepMind

Best ARC-AGI-2 score: 89.2 percent with Gemini 3.8 Flash (High), released September 2, 2026, at $0.400 per task (Google). On ARC-AGI-3 it scores 10.4 percent under the standard harness and 35.0 under the provider adapter. Note that Google's best verified entry is a Flash model, its smaller and cheaper tier; the company said on September 23 that Gemini 4 is in early post-training, with an early release intended well before the end of 2026; our note on the Gemini 4 timeline has the quotes. Google ranks third on released capability and first on distribution, since Gemini sits inside Search, Workspace and Android.

4. Dots Studio

Best ARC-AGI-2 score: 76.8 percent with Dots3-Note Preview (Max), released August 14, 2026, at $0.077 per task. ARC Prize notes that as of September 21, 2026 it "sets a new SOTA score for open-weight models" (ARC Prize results). Most readers will not have heard of the lab, which is the point of ranking by output rather than fame: a preview open-weight model from a small team outscores every Chinese major and xAI on this test at a fraction of the cost per task. It is marked "preview", so treat the position as provisional.

5. xAI

Best ARC-AGI-2 score: 67.1 percent with Grok 4.6 (XHigh), released August 11, 2026, at $0.757 per task; on ARC-AGI-3 it scores 2.1 percent under the standard harness (xAI). xAI ships fast and sells through an API and through X, but on this independent test it trails the top three by 22 to 28 points.

6 to 9. Zhipu, DeepSeek, Moonshot, Alibaba

  • Zhipu: GLM-5.3-Flash (Max), August 26, 2026, 65.8 percent at $0.093 per task.
  • DeepSeek: V4 Flash 0731 (Max), July 31, 2026, 61.4 percent at $0.042 per task, the lowest cost per task of any model above 60 on the board; see DeepSeek V4 Flash weights and pricing.
  • Moonshot: Kimi K3 (Max), July 16, 2026, 60.4 percent at $1.59 per task; weights released July 30, covered in Kimi K3 open weights.
  • Alibaba: Qwen3.8-27B (XHigh), August 14, 2026, 42.4 percent at $0.447 per task.

Three of the four are open-weight releases, which is why "what they shipped" cuts in their favour even where the scores do not: anyone can run them. The business implications of that are in how Chinese models are shifting US AI costs.

10. Thinking Machines

Best ARC-AGI-2 score: 40.1 percent with Inkling Small (XHigh), July 30, 2026, at $0.233 per task. The company released Inkling, a 975-billion-parameter open-weight model, on July 15, 2026, positioned as a base for fine-tuning rather than a frontier contender; our note on Inkling has the details. Ranked last among labs with a scored model, first among them in honesty about what it is for.

Unranked: Meta

Meta Superintelligence Labs shipped Muse Spark on April 8, 2026, calling it "the first step on our scaling ladder", Muse Spark 1.1 on July 9, Muse Image on July 7 and the Muse personal agent on September 8 (Meta). None of the Muse models appears on the ARC-AGI leaderboard as of today; Meta's only entries are Llama 4 Maverick and Llama 4 Scout from April 2025, which both score 0 percent on ARC-AGI-2. Meta therefore ranks high on products shipped and cannot be placed on capability until an independent score exists. Meta's personal superintelligence, promised and delivered is the full profile.

Last: Safe Superintelligence Inc.

No model, no paper, no API, no benchmark. SSI's site says it has "one goal and one product: a safe superintelligence" (SSI), and its investors, including Nvidia with a reported $5 billion in July 2026, have priced that promise at $32 billion. On a ranking by what has shipped it is last by construction; Safe Superintelligence Inc: $8B in, zero products out explains why that is deliberate.

What the scores do and do not tell you

ARC-AGI-2 is close to saturated at the top: the gap between first and third is under six points, and the human panel's 100 is within reach of next quarter's releases. That makes it a good test of who has a frontier model and a poor test of who is closest to AGI, which is why ARC Prize built ARC-AGI-3 and why the standard-harness scores on it, 62.7, 30.2, 10.4 and 2.1 for the four labs that have submitted, are the more revealing column. Cost per task matters as much as the score: Anthropic's 93.3 at $0.41 and OpenAI's 94.2 at $0.25 are the numbers a business should look at, not the 95.0 at $1.12. And a benchmark measures a model, not a product. Reliability on your workflow, tool integrations, data handling and price are what you actually buy, which is why the ranking above would look different if sorted by shipped products, where Meta and OpenAI lead, or by open weights, where the Chinese labs and Thinking Machines lead.

What this means if you are buying

Three of the top four labs sell models that score within six points of each other, so the choice between them for a business task should be made on cost, latency and the integrations you need, not on who is "closest to AGI". The practical reading of the board for a company deploying AI this quarter: use OpenAI's GPT-6.1 Sol or Anthropic's Opus 5.5 for tasks that need the strongest reasoning per dollar, a Flash-class model from Google or an open-weight Chinese model for high-volume routine work, and ignore the AGI label entirely when an AI automation agency pitches you. The same logic applies to voice: an AI calling agent for home services runs well on models that were mid-table a year ago, because answering, qualifying and booking a call is a narrow task that the whole top ten can now do. The frontier matters for what becomes possible next year; the ranking of what shipped matters for what you pay this month. For the price side, see our AI model API pricing comparison.

Want the model choice made on cost and reliability for your workflow instead of on AGI headlines?

We design, build, and run it for you, integrated with the tools you already use. Free audit in 24 hours.

Get Your Free Audit

Frequently Asked Questions

By independent test scores, OpenAI: its GPT-6 Astra scores 95.0 percent on ARC-AGI-2 and holds the top standard-harness score of 62.7 percent on the interactive ARC-AGI-3, with Anthropic's Claude Opus 5.5 at 93.3 percent on ARC-AGI-2 and Google's Gemini 3.8 Flash at 89.2. No lab claims to have reached AGI by its own definition.

The labs that state AGI or superintelligence as their goal are OpenAI, Anthropic, Google DeepMind, Meta Superintelligence Labs, xAI and Safe Superintelligence Inc. Chinese labs such as DeepSeek, Moonshot, Zhipu and Alibaba, and newer entrants such as Thinking Machines and Dots Studio, compete on the same benchmarks.

By each lab's best verified score on ARC-AGI-2 from the ARC Prize leaderboard, read October 3, 2026, with standard-harness ARC-AGI-3 scores as the tiebreaker and only publicly available or verified systems counted. Valuations, revenue and chat-arena rankings were excluded.

Meta's current Muse models, released from April 2026 onward, have no score on the ARC-AGI leaderboard as of October 3, 2026; its only entries are Llama 4 Maverick and Scout from April 2025, which both score 0 percent on ARC-AGI-2. Meta ranks high on products shipped but cannot be placed on independent capability until a score exists.

Because it has released nothing: no model, paper, API or benchmark result in 28 months, despite a $32 billion valuation and a reported $5 billion investment from Nvidia in July 2026. Its stated plan is to build a safe superintelligence directly without intermediate products.

Less than price and reliability do. The top three labs score within six points of each other on ARC-AGI-2, so for business tasks the choice should rest on cost per task, latency, integrations and data handling. Narrow tasks such as answering and booking calls run well on models from the whole top ten.

Free Strategy Audit

Ready to put this to work?

Join 200+ businesses already scaling with AI and automation. Get your free audit and a custom roadmap within 48 hours.

Website & marketing performance analysis
AI & automation opportunity mapping
Custom growth roadmap with ROI estimates
Delivered within 48 hours, 100% free
200+
Clients served
48hr
Turnaround
100%
Free, no strings

Get Your Free Audit

Takes 30 seconds. No credit card required.

Prefer to chat?

WhatsApp us