Managers buy decisions, not benchmark points. A polished answer can still carry a recommendation-breaking error.
Can it RunAlfrada?
Models in a real harness
doing real work
This tests the model inside Alfrada — not the model in isolation. Each model used the same production harness to complete two multidisciplinary jobs and ship board-ready artifacts. Seven judges then audited every file.
Markdown copy Full study text with chart image URLs — for sharing, Notion, or LLMs
02What we test
Short answer: a model, inside a fixed production harness, doing a real job. The model supplies judgment and reasoning. Alfrada supplies planning, nineteen live tools, isolated workspaces, code execution, file creation and delivery. Because the harness stayed fixed, the ranking shows how different models operate inside the same system. It is not a ranking of raw model intelligence in isolation.
This is also not a coding benchmark. The job is lateral, multidisciplinary synthesis — research, judgment, computation, communication and delivery. A good harness should narrow the gap between models: checks, tool contracts, workspaces and validation keep work complete, inspectable and on spec while the model still determines the quality of the reasoning.
03The test
Each model became the reasoning engine inside the same production Alfrada session: same tools, same prompts, temperature 0, fresh workspace and strict cold start. No model could read another run or quietly delegate difficult work to a stronger cousin. We then gave that model-plus-harness system two jobs that resemble work a strategy team might actually commission.
Case 1 — The decision package
“Should NorthLine, a $12M-ARR logistics SaaS, enter the UAE in Q3 2026?”
- Market & competitor intelligence with a primary-source evidence standard — staged screenshots must contain the figures they're cited for
- A computed 24-month financial model: TAM/SAM/SOM, CAC, payback, FX sensitivity, an explicitly defined CAC-inclusive breakeven
- A ≥1,500-word decision memo, a ≥10-slide board deck, a 40–50s audio briefing, and scheduled 30/60/90-day review scaffolding
Case 2 — The prediction
“Forecast global semiconductor capex 2026–2030 — as a distribution, not a number.”
- Five quantified drivers with named, dated, findable sources for every prior
- A ≥10,000-trial Monte Carlo with a printed seed and the per-year summary saved as a JSON artifact — quoting numbers the simulation didn't produce is scored as fabrication
- Fan, tornado and calibration charts; a ≥1,500-word memo with two invalidation scenarios and one falsifiable 18-month prediction
Did it ship?
Mechanical checks enforce the minimum: required files, word and slide counts, successful tools, working schedules and zero tool errors.
Does it survive inspection?
Seven judges open the artifacts, replay the calculations, inspect the evidence and score the decision package from 0–10.
The judge panel is Kimi K3, GLM-5.2, Gemini 3.5 Flash, GPT-5.6 Sol, GPT-5.6 Luna, Claude Fable 5 and Claude Opus 4.8. They do not merely read the final answer. They list the workspace, open the memos, re-run Monte Carlos from saved seeds, recompute CAC and payback, and vision-inspect slides and charts. Each of the nine contestants is judged by all seven, producing 126 verdicts. The judge panel was fixed at seven and was not expanded after scoring began, so Qwen and Inkling are contestants but not judges. Inkling's own work also fell below the usability cutoff and did not establish the verification discipline required of an evaluator.
04Results
GPT-5.6 Sol wins on audited quality (8.84 raw; +0.97 calibrated z), and it isn't self-dealing. Its own verdict is excluded from its calibrated rank, as every self-judgment is. Fable and K3 follow. Luna remains the price-performance outlier. Every model passed the mechanical delivery gates; that did not make every delivered job usable.
Manager translation: 6.0 is the practical cutoff, not a percentage grade. Eight models scored 7.05–8.84 and produced usable first passes with varying review burden. Inkling scored 5.60; its package looked complete but failed on recommendation-bearing arithmetic, contradictory source-of-truth files and an irreproducible forecast. Below six, commission substantial rework rather than treating the output as an analyst draft.
Figure 1 — The leaderboard
Rank uses family-balanced judge z-scores; raw composite preserves the familiar 0–10 scale. Click a column to sort or a row for its evidence.
Figure 2 — Who judges whom
Rows are judges, columns are the judged. Click any cell to read the verdict verbatim.
The matrix is where this study earns its keep. Raw means reward generous judges and duplicate architecture families. The calibrated rank instead z-scores each judge within each case, averages Sol/Luna into one GPT-5.6 signal and Fable/Opus into one Claude signal, then gives K3 and GLM one signal each. Self-votes are excluded. Gemini remains fully visible but receives zero ranking weight because its near-uniform marks add almost no discrimination.
Figure 3 — The self-judgment diagonal
Each panel member's score for its own work versus its peers, pooled across both cases. Inkling and Qwen do not appear because they are not judges.
05How to read the scores
Strategy is difficult to score because recommendations, writing and presentation involve taste. Computation does not. A model either reproduced its own Monte Carlo from the saved seed or it did not. CAC arithmetic either reconciled or it did not. A cited screenshot either contained the claimed figure or it did not. Those binary checks give us a hard spine inside an otherwise subjective evaluation.
This is why the scores do not map to SWE-bench or an exam percentage. Coding benchmarks usually ask whether one bounded patch passes tests. Here the coding sandbox is only one tool inside a much wider job. The run must research, model, decide, explain and deliver several mutually consistent artifacts — then survive seven auditors opening the files.
The harness also raises the floor. Schemas, hard gates, file validation and repeatable tools keep output at a steady minimum quality, so model differences appear mostly above that floor. We did not run a “naked model” control, so this study does not quantify the harness uplift causally. What it does show is the difference between shipping and surviving inspection: all nine models cleared every hard gate, while audited quality occupied a 5.60–8.84 band.
Hallucination is not the end of the test. Fabricated facts are penalized, but the higher bar is finding a break in the reasoning chain even when the answer looks polished: evidence that does not support the claim, arithmetic that contradicts the memo, or validation that is circular. That is a broader standard than “did the answer compile?”
Use 6.0 as a go/no-go floor. Below it, the work needs substantial reconstruction. From 6–7, treat it as a supervised draft. From 7–8, it is useful with targeted expert review. Above 8, the package survives more of the audit — but no score removes the need for accountable judgment on consequential decisions.
Defect counts are observations, not rates. “No catalogued load-bearing defect” means none was found in this pinned run; it does not imply a 0% underlying failure probability. The cards below are directional selection hypotheses to validate on your own workload, not procurement guarantees.
Sol · 8.84
The deepest evidence and validation package, with no catalogued load-bearing defect. In this sample, it is the strongest candidate when the cost of a missed issue dominates model spend.
Luna · $0.92
8.04 across both cases in about ten minutes, with honest framing and no catalogued load-bearing defect. The trade is thinner evidence staging and QA.
K3 · 8.40
$3.54, verified core reasoning and no catalogued load-bearing defect. It is slower than Luna, but combines near-frontier quality with a moderate bill.
GLM · 7.95
Strong reproducibility at $4.12. It belongs on the shortlist, with explicit review of the unsupported case-1 breakeven claim found by the panel.
06Load-bearing errors
A load-bearing error sits underneath the recommendation. Remove or correct it and the economics, confidence or proposed action materially changes. A mislabeled column is annoying; a breakeven claim calculated on the wrong profit measure can send a board into the wrong market. This is the most useful distinction for managers because it separates polish problems from decision risk.
By that cut, Sol, Luna and K3 combine zero catalogued load-bearing errors with
evidence the panel could meaningfully inspect. Qwen also shipped clean core arithmetic,
but its evidence pack does not substantiate the load-bearing market claims, so it should not be
treated as equally low-risk. Fable, Opus and GLM each carry one: Fable's "month-26
breakeven" is really the first positive month before two more negative ones; Opus asked the
board to approve a partner-led pilot whose partner economics its model never contains; GLM
told the board month 30–33 when its own JSON records breakeven: null.
Gemini carries two, one in each case. Inkling carries three and fails the cutoff.
Why 5.60 is a failed job, not merely last place. Inkling summed annualized ARR every month, inflating reported 24-month revenue by roughly 12× and claiming breakeven in months one to three. Its JSON, memo and charts disagree on costs and timing. In case 2, eight plausible reconstructions of the stated method failed to reproduce the published forecast. The package shipped; its decision logic did not survive replay.
A second warning: polish can outlive the evidence. Gemini's memo defines breakeven as gross profit exceeding operating expenses — then computes it on revenue, silently skipping the 20% COGS. By its own definition, the UAE branch never breaks even inside 24 months; the memo claims month 21 and recommends GO partly on that number. Four judges found the error independently. In case 2, the same model hand-tuned its backtest parameters with hindsight and presented the sub-1% fit as out-of-sample validation at “92.0% confidence.” In both cases, the reasoning process broke before the prose did.
Directionally, the fleet agrees more than the scores suggest: eight of nine models reached the same case-1 recommendation — a conditional, gated GO — with one principled dissent: Kimi K3 alone concluded conditional no-go, on arithmetic its judges verified. And all six with verified models agree there is no cumulative breakeven within 24 months. Where they genuinely diverge is the forecast:
Figure 4 — Same question, same sources, 2.3× apart
Each model's P50 five-year CAGR for global semiconductor capex, with its own P10–P90 band where reported.
07What this means in practice
The practical result is a controllable quality floor, not guaranteed correctness. The harness made every model ship the required package. Eight produced useful drafts; one demonstrated that completeness checks cannot rescue broken reasoning. Lower-cost entrants can do valuable work inside strong controls, but the controls must preserve enough evidence for an expert to audit the chain.
The accountable expert moves up the stack. Analysts, scientists and reporters spend less time assembling the first package and more time deciding which assumptions matter, replaying the consequential calculations and resolving conflicting evidence. Management shifts from checking whether work exists to checking whether the recommendation is warranted.
For higher-stakes production work, Alfrada's “Beast” workflow adds an integrated critic and judge loop intended to push borderline work toward the high-7s and 8s. That uplift is not claimed as a result of this study because Beast mode was not enabled for the contestant runs.
Route by consequence
Use Sol where a missed issue is expensive, K3 for balanced quality, Luna for fast high-volume work, and GLM where reproducible modelling matters but a targeted review is available.
Audit the spine first
Start with source provenance, definitions, arithmetic, simulation replay and cross-artifact consistency. Line-editing polished prose first is the wrong control.
Keep the working, not just the answer
For analysts, retain source-of-truth models. For scientists, retain code, seeds and calibration. For reporters, retain primary sources and claim-to-evidence links.
AI drafts; a person owns the decision
The benchmark supports model procurement and workflow design. It does not transfer professional responsibility to a composite score.
Figure 5 — Cost against quality
Both cases combined. Bubble area is agent working time; gold marks the efficient frontier.
08The evidence
A benchmark you can't audit is an opinion. Below, per model: every judge's score, the verdicts verbatim, the defect catalogue with its load-bearing tags, the charts each model drew, and a downloadable package — the memos, financial models, simulation outputs and evidence files exactly as each model left them in its workspace, plus the full judge record.
Prefer a portable copy? Download the Markdown export (includes absolute URLs for every chart).
09Method & limitations
Setup. Runs executed July 17–19, 2026 against a production Alfrada build on a single dev machine, one isolated bot user per model run (separate workspaces, budgets, and history), temperature 0, seed 42 in the case config, $15 budget cap per run. Sub-agent workers were pinned server-side to the model under test; the worker requests each model made are recorded and published in the packages. Judge sessions ran under a separate critic user with read-only tools plus code execution for replay.
Scoring. The 0–10 raw composite is the mean of all seven judge scores across both cases and remains visible for interpretability. Rank is calibrated separately: within each case, every judge is z-scored across the nine submissions; Sol/Luna are averaged into one GPT-5.6 family signal, Fable/Opus into one Claude signal, and K3 and GLM each contribute one signal. Self-judgments are excluded. Gemini's verdicts remain published but receive zero calibrated weight because their 9.2–10.0 range provides almost no discrimination. This correction changes the fairness of the aggregation, not the winner or model order.
Usage accounting. The study processed 129,765,534 input/output tokens across 18 contestant sessions and 126 judge sessions. Contestants cost $48.11; judging cost $133.58; total recorded model and tool spend was $181.70. Per-model leaderboard costs intentionally show contestant spend only so procurement comparisons are like-for-like.
Evidence base. The published table is nine attributable entries from a ten-candidate publication cycle; the invalid Qwen swarm result is disclosed below. It sits on hundreds of internal benchmark and development executions accumulated over roughly three months. Because prompts, models and the harness evolved during that period, those runs are not pooled into the displayed scores. They make the broad pattern less surprising; they do not turn this cross-section into a repeated-trials estimate.
Limitations, plainly. Each displayed estimate remains one pinned result per model per case — scores carry roughly ±0.3 of observed run-to-run noise, small gaps should be read as ties, and defect counts must not be read as failure rates. Provenance reconciles as follows: four model entries use July 17 outputs (Fable, Opus, Luna, Gemini), three use July 18 (Sol, GLM, Qwen), and two use July 19 (K3, Inkling). Provider load varies. Judges are themselves models: the panel's biases are measurable (Figure 2), one judge required an automated correction pass after falsely reporting an empty workspace, and rationales — however well-verified — inherit their authors' blind spots. Family averaging reduces duplicate architecture weight but cannot remove correlated blindness: related judges may miss the same error classes together, and z-scoring calibrates scoring severity, not truth. There was no model-without-harness control and no Beast-mode treatment, so neither uplift should be inferred from these scores. Two rows carry disclosures: Kimi K3's result pairs two adjacent runs (its case 1 and case 2 each completed and were fully judged in back-to-back runs on July 19, after we diagnosed why earlier attempts died — a harness bug, not the model: the stream reader ignored keepalive pings during K3's legitimately long tool-history turns and killed live requests as idle; both paired cases scored 8.4 with complete panels). And a swarm-mode Qwen run was excluded after it delegated every worker to GPT-5.6 Sol against explicit instructions, which made its otherwise excellent output unattributable.
Auditable, not externally reproducible. The public evidence packages let outsiders inspect each contestant's memos, models, simulation outputs and full judge record. Outsiders cannot independently rerun the study because the harness is our production system. We therefore claim auditability of the published artifacts, not external reproducibility. Questions or corrections: community@strategize.inc.