AI-NATIVE STARTUP
← Essays

Pillar · THE RECEIPTS

The 13.86% that built a $26B company

Devin's benchmark score was real. The comparison it was sold against was not like for like. Two years later the company is worth $26 billion, and the most-quoted number about it still has no definition.

2026-07-22 —— [RECORD-LED] —— SOURCE: record/cognition-devin.md

Open source Record →

In March 2024, Cognition came out of stealth with a number. Devin, which it called the first fully autonomous AI software engineer, resolved 13.86% of issues in SWE-bench end to end, unassisted, pass@1 [1][3]. The announcement put that against a prior state of the art of 1.96% [1]. Roughly seven times better than anything before it. That framing travelled further than any other number in agentic software last year.

The score is real. We checked it and it holds [10]. The comparison does not.

Buried in Cognition's own technical report is the caveat that makes the difference [2]. Devin was evaluated unassisted: it had to find the relevant files itself. The models it was compared against were evaluated assisted: they were told exactly which files to edit. Those are not the same test. Read like for like, unassisted against unassisted, the honest comparison is Devin's 13.86% against Claude 2's 4.80% [2].

Still a large win. Roughly 2.9x, not 7x. And Cognition disclosed the caveat itself, in public, in writing. That matters and it should be said plainly: this is not a fabricated number. It is a real number wearing a flattering comparison, and the flattering comparison is what got repeated.

01

What the number bought

Whatever you think of the framing, the trajectory that followed is documented and it is steep.

Cognition's ARR went from $1M in September 2024 to $73M in June 2025 [5], then to $492M by May 2026 [4]. In May 2026 it raised $1B at a $25B pre-money valuation, $26B post, led by Lux Capital, General Catalyst and 8VC [4]. The previous round, eight months earlier in September 2025, valued it at $10.2B post [10]. The valuation more than doubled in eight months. Total funding is above $2.5B. Named customers include Goldman Sachs, Mercedes-Benz, NASA and Santander [4].

At close, $492M ARR against $26B is a multiple of roughly 53x.

Two pieces of fine print belong next to that curve. First, the $492M is combined ARR following the July 2025 acquisition of Windsurf, which contributed roughly $82M at acquisition [10]. Second, the price Cognition paid for Windsurf was never disclosed, while Google separately paid $2.4B for Windsurf's talent and licensing rights in a parallel deal [8]. Cognition took the IP, the technology, the customers, and the employees who turned Google down. What that cost is not public.

02

The number with no definition

The most repeated claim about Cognition today is not the benchmark. It is that Devin writes the company's own code.

Here the two available figures do not agree with each other. CEO Scott Wu said publicly that 90% of Cognition's code is written by its AI [6]. The company's own official announcement, the same month, said 89% of code committed by our engineers is committed by Devin [6].

The one-point gap is not the problem. The problem is that "committed" is never defined [10]. Committed means a commit landed under Devin's name. It does not tell you whether a human wrote the logic and Devin formatted it, whether a human reviewed every line before it merged, or whether Devin generated it and a human accepted it without reading it. Those are three completely different companies, and the same 89% describes all of them. There is no independent audit.

Cognition's other self-reported performance figures carry the same structure: PR merge rate up from 34% to 67% year over year, 4x faster problem solving, 2x more resource efficient [7]. Plausible, directionally consistent with the category, and unreplicated by anyone outside the company.

03

Why this is the pattern, not the exception

Cognition is not being singled out as unusually loose. By the standards of this category it is unusually forthcoming: it published the SWE-bench caveat itself, it puts real ARR figures on its own blog, and it named the acquisition. Most of the companies in our record do less.

The pattern it illustrates is the one that shows up in every audit we run. The numbers that are easy to verify get published. The numbers that would let you check the claim do not. Funding: verified in one search. Valuation: verified. Customer logos: verified. The composition of a 53x revenue multiple, the definition of an autonomy percentage, the price of an acquisition: absent.

There is a useful outside number for calibration here. On Martian's independent Code Review Bench, which scores against roughly 300,000 real pull requests, the best AI code reviewer in the category posts an F1 of 51.2% [9]. That is the measured state of the art for machines checking code, published by a party with nothing to sell. Hold "89% of our code is written by AI" next to "the best available checker catches about half," and the interesting question stops being how much the agent writes. It becomes what reviewed it.

04

What to do with a number like 13.86%

Three checks, usable on any claim in this category:

  1. Ask what the comparison was measured under. Not just the score, the conditions. Assisted against unassisted is the single most common way a real number gets a fake multiple attached.
  2. Ask for the definition before you quote the percentage. "89% committed by Devin" is not a fact until "committed" is defined. If a company publishes a number like that about itself, it should publish the method with it. If it does not, cite it as self-reported and undefined, and say so.
  3. Separate what verifies from what does not, and say which is which. Cognition's funding, valuation, acquisition timing and customer names all verify. Its autonomy percentage, its internal performance multipliers and its acquisition price do not. Both halves are legitimate to write about. Only one is legitimate to state flatly.

The 13.86% was real. Two years later it is also obsolete: multiple agents have passed it on the live leaderboard [10]. That is the ordinary fate of a benchmark. What outlived the score was the framing, and the framing is the part nobody re-checked.

Nothing is asserted here that isn't shown. Every figure above is dated and linked below.

Sources · verify

  1. Cognition, "Introducing Devin" (March 2024): 13.86% SWE-bench, pass@1, unassisted; 1.96% stated as prior state of the art. https://cognition.com/blog/introducing-devin
  2. Cognition, SWE-bench technical report: the 1.96% figure was assisted; like-for-like unassisted, Claude 2 scored 4.80%. https://cognition.ai/blog/swe-bench-technical-report
  3. VentureBeat (March 2024): Cognition emerges from stealth with Devin. https://venturebeat.com/ai/cognition-emerges-from-stealth-to-launch-ai-software-engineer-devin
  4. TechCrunch (27 May 2026): $1B raise at $25B pre-money / $26B post; $492M ARR; customers Goldman Sachs, Mercedes-Benz, NASA, Santander. https://techcrunch.com/2026/05/27/ai-coding-startup-cognition-raises-1b-at-25b-pre-money-valuation/
  5. Cognition blog: ARR $1M (Sept 2024), $73M (June 2025). https://cognition.com/blog/funding-growth-and-the-next-frontier-of-ai-coding-agents
  6. The Next Web (May 2026): CEO states 90%; official announcement states 89% of code committed by engineers is committed by Devin. https://thenextweb.com/news/cognition-just-raised-1-billion-at-a-26-billion-valuation-and-90-of-its-own-code-is-written-by-its-ai
  7. Cognition, "Devin annual performance review 2025": merge rate 34% to 67%; 4x speed; 2x efficiency. Self-reported. https://cognition.com/blog/devin-annual-performance-review-2025
  8. Wikipedia, Cognition AI: founding date, Windsurf acquisition July 2025, Google's parallel $2.4B talent and licensing deal. https://en.wikipedia.org/wiki/Cognition_AI
  9. Martian Code Review Bench: category-leading F1 of 51.2% across ~300,000 real pull requests. https://codereview.withmartian.com/
  10. AI-NATIVE STARTUP verify record (draft, advisory): content/record/cognition-devin.md

More Record-led essays