▮ SUPINT · @SUPINTorg

Superintelligence, measured.

Models · Harnesses · Evidence

Benchmarks and commentary on frontier AI. We give models real engineering work, check every claim they make against the code, and publish the grades: on X first, with the full evidence here.

BENCH 01
Silent Carrier · retired
GRADED
18 models · 525 claims
NEXT
Bench 02 in design
The SUPINT SI-1 bench board An illustrated circuit board. A central SUPINT core is wired by gold traces to six chips, one for each thing the account covers: models, harnesses, claim checks, evidence reruns, rubrics and commentary. Pulses of light carry signals between them. Hover or focus a chip to read about it.
Switch off to pause all motion

// BENCH 01 · SILENT CARRIER · RETIRED

Eighteen models, one firmware bug hunt.

Every model got the same brief on a real radio firmware codebase, in a sealed sandbox, and every claim in its review was checked against the code. This is the final board; the full audit reports stay on bench.innerpulse.net, where they were first published.

“Find why this firmware struggles to decode P25 Phase 1 and Phase 2 voice on the radio’s HR-C6000, and what to do about it.”

DM-1701 · HR-C6000 · AT1846S · STM32F405 · P25 · graded under Clean Room 1.0

Silent Carrier final leaderboard: grade, weighted score out of 100, the six rubric dimension scores, claims that held of those checked, wrong claims, and decode-critical issues found.
#ModelGradeScoreAccuracy30%Coverage25%Root cause15%Fix plan15%Originality10%Clarity5%ClaimsheldWrongCriticalof 4Report
–Claude Opus 5.5unrankedAnthropic · The vocoder wall was a bug.A9395889794969342/440Report ↗
–Claude Opus 5unrankedAnthropic · It measured what everyone argued about.A9294869693959220/210Report ↗
1GPT-6 AstraOpenAI · Reproduced, not asserted.B8992848892908627/290Report ↗
2GPT-5.6 SolOpenAI · Fixes the instrument first.B8491708490848627/280Report ↗
3Grok 4.7xAI · Short, exact, and unreferenced.B8082708686808615/190Report ↗
–Claude Fable 5.1unrankedAnthropic · Counts the calls, misprices them.C+7985628486848815/201Report ↗
4HY4 PreviewTencent · The first one to measure.C+7785568084828824/300Report ↗
5GPT-5.6 LunaOpenAI · Careful, short, and never built.C+7582588080728423/241Report ↗
6UNIONALPHAStealth model · Accurate to the byte. Hesitant about what stops decoding.C7490656070768033/350Report ↗
7Grok 4.6xAI · Finds what erases the signal. Misses what stalls the voice.C7074597075727821/302Report ↗
8Muse Spark 1.3 ContributorMeta · Finds both halves of the filter problem. Points the CPU fix the wrong way.C−6874547271708022/302Report ↗
9DeepSeek V4.1 FlashDeepSeek · Finds the microphone and the filters. Waves de-emphasis through.C−6673496670747428/353Report ↗
10GLM 5.3Z.ai · Proves the real capture bug. Invents a second one.C−6568527468707225/312Report ↗
11Gemini 3.8 FlashGoogle · A buildable plan, on an invented number.D6465467470708423/354Report ↗
12Qwen3.8 MaxAlibaba · Best case yet for the microphone. Never looks at the CPU.D6370397171667821/272Report ↗
13Qwen3.8 FlashAlibaba · Right about the filters. Wrong about the CPU.D5962406870647021/336Report ↗
14GLM 5.3 FlashZ.ai · Audits the memory to the byte. Misreads the receive path.D5664346262586817/284Report ↗
15DeepSeek V4 ProDeepSeek · Right pivot. Wrong arithmetic.D5461336460624817/265Report ↗

Weighted score out of 100 over six dimensions; the best score in each column is in gold. Grades: A ≥ 90 · B 80–89 · C+ 75–79 · C 70–74 · C− 65–69 · D 50–64 · F below 50. Unranked: run through the Claude Code CLI rather than the agent harness every ranked run shared, and without its step cap, so published but not ranked. The auditor is also made by Anthropic; bench discloses this with the grades.

FINDING · 01

No ranked model reached an A

Both A grades (93 and 92) came from Claude runs through a different harness. The top ranked run is GPT-6 Astra at 89. The harness is part of the result.

FINDING · 02

Accurate isn’t the same as useful

UNIONALPHA scored 90 for accuracy, with 33 of 35 claims holding, yet found one of the four issues that actually stop decoding. It finished on a C.

FINDING · 03

Wrong claims pile up at the bottom

The five highest scores made no wrong claims between them. The five lowest made 21.

FINDING · 04

Coverage split the field

Coverage of the decode problems ran from 33 to 88, the widest spread of the six dimensions. Most models could explain what they found; fewer found what mattered.

// COVERAGE

Models, harnesses, and the evidence.

Six things the account measures and writes about. Each maps to a chip on the board above; the SUPINT core in the middle is the rulebook they all run through.

SI-01 · U2

Model benchmarks

Frontier and stealth models on real engineering work: one brief, one codebase, one published rubric, so the grades compare directly.

Details
  • Named, pre-release and stealth models
  • Real codebases, not quiz questions
  • Grades on a fixed rubric
See the first board →
SI-02 · U3

Harness comparisons

The same model through different agent harnesses. The harness is part of the result, so runs through a different one are published but ranked apart.

Details
  • CLI agents and bench harnesses
  • Step caps, tools and retries recorded
  • Ranked only against like runs
Suggest a harness →
SI-03 · U4

Claim checks

Every claim in a model’s answer is checked at the lines it cites. It holds, is qualified, or is wrong, and the count is part of the grade.

Details
  • Checked against source and manuals
  • Holds, qualified or wrong
  • 525 claims on the first board
How claims are checked →
SI-04 · U5

Evidence reruns

Tests, builds and harnesses a model relied on are run again, to separate real findings from artifacts of how they were produced.

Details
  • Suites and builds rerun
  • Missed issues modelled, not just listed
  • Run facts recorded, never graded
Read the method →
SI-05 · U6

Published rubrics

Weights, dimensions and the grade scale are set before the first run and frozen. Change a rule and it becomes a new version, kept apart.

Details
  • Six weighted dimensions
  • A ≥ 90 through F below 50
  • Versions never mixed in one ranking
See Clean Room 1.0 →
SI-06 · U7

Commentary

Release-day takes, leaderboard threads and what the numbers do and don’t show. Posted on X first, with links to the evidence.

Details
  • Threads on new models and harnesses
  • Corrections posted in the open
  • Every claim linked to its evidence
Follow @SUPINTorg →

// CLEAN ROOM 1.0

Comparable because the rules are frozen.

Every Silent Carrier run followed the same seven rules, fixed before the first run. Change a rule and it becomes a new version, kept apart rather than mixed into one ranking. Four of them do most of the work.

F1 · FUSE

Step cap

150 agent steps and no time limit. A run that loops trips the cap instead of running up a bill, and API retries are raised only to survive rate limits.

D1 · DIODE

Sealed sandbox

Tools run with no network and only the repository mounted. The brief goes in; nothing from outside reaches the run.

WDT · WATCHDOG

Every claim checked

Each claim is checked at the lines it cites, against the source and the manuals, and the evidence behind it is rerun.

K1 · RELAY

Withdrawn when tainted

A run that breaks the rules comes off the board. The first P25 runs were withdrawn and replaced by clean-room reruns.

All seven rules

  1. One model per run: no subagents, delegation, skills or memories.
  2. Tools run in a sandbox with no network and only the repository mounted.
  3. A fresh copy of the repository at 46eebda with its own docs and the HR-C6000 manual, and no earlier reviews.
  4. The same brief, word for word, with the same tools and reasoning effort.
  5. No time limit and a 150-step cap; API retries are raised only to survive provider rate limits.
  6. Six fixed dimensions and weights, graded against the verified answer key.
  7. Run facts (time, steps, tool calls, tokens) are recorded but never graded.

// SIGNAL PATH

Bench 02 is in design.

Silent Carrier is closed to new runs. The next bench keeps what worked (a real task, frozen rules, every claim checked) and makes the harness a variable in its own right. Results post to X first.

Choose the task

01 Brief

A real engineering question on a real codebase, with an answer key verified before any model sees it.

Suggest a task →
Run it clean

02 Run

Every model and harness under the same frozen rules, in a sealed sandbox, with run facts recorded but never graded.

Suggest a model or harness →
Check and post

03 Audit

Every claim checked where it points, the evidence rerun, and the grades posted with the full reports behind them.

Follow for results →

// WORK WITH US

Private evaluations

The same method on your code: models, agents or harnesses graded against your own tasks, with every claim checked. Results stay private unless you say otherwise.

Discuss an evaluation

// DIAGNOSTICS

Before you ask.

What is SUPINT?

A benchmarking and commentary account for frontier AI. We grade models and the harnesses that run them on real engineering work, check every claim they make, and write about what the results do and don’t show. Results post first on X at @SUPINTorg; this site keeps the boards and the method.

Why is Silent Carrier retired?

It’s closed to new runs, and its final board stays here as the reference: 18 models graded under one frozen ruleset, 525 claims checked. The full audit reports remain at bench.innerpulse.net, where they were first published. The next bench is in design.

Why are some runs unranked?

Three Claude runs went through the Claude Code CLI instead of the agent harness every ranked run shared, and without its 150-step cap. They’re published because the results are real, but ranking them together would compare harnesses as much as models. The auditor is also made by Anthropic, which bench discloses with the grades.

How is a review graded?

Six weighted dimensions (accuracy 30%, coverage 25%, root cause 15%, fix plan 15%, originality 10%, clarity 5%) give a score out of 100 and a letter: A from 90, B from 80, then C+, C and C− down to 65, D from 50, and F below. Claims are checked where they point and the evidence is rerun before any dimension is scored.

Can I suggest a model or a harness?

Yes: reply on X or email [email protected]. Stealth and pre-release models can be graded without the auditor knowing who made them, as they were on Silent Carrier.

// CONTACT

Tips, corrections, requests.

Spotted a wrong number, want a model on the next board, or know a harness we should run? Follow along on X, or email.

or write to [email protected]