Benchmarking · July 24, 2026

Introducing Composite-Bench: the strongest open-weights model isn't Kimi K3

Composite-Bench is the first long-horizon, browser-based computer-use benchmark that has certified optimality and verified compute. It's built on our unique corpus of real professional trajectory and intent data. As far as we can tell, it also provides the first independent long-horizon browser-use results for this month's open-weight releases, and those results disagree with the leaderboard screenshots of Kimi-K3's launch. In our evaluation, GLM-5.2 lands 11 points behind a two-way Claude tie at the top and beats every closed model we measured from OpenAI, Google, and xAI. Kimi-K3 is still an excellent model: it clears OpenAI's flagship and takes second on our hardest pure-search slice. But on certified long-horizon browser use, our results do not match the common verdict of "Kimi for capability, GLM for price."

Introduction

Composite builds products used by hundreds of thousands of professionals: accountants, recruiters, financial-crime analysts, EMS coordinators, engineers, legal staff, and so on. Our products assist that work inside enterprise software, so we capture real browser-use trajectories and, for many users, the intent behind them. That data let us ask a question no public leaderboard answers: which frontier models can do long-horizon enterprise work on a browser or computer, even at full reasoning effort?

We reviewed 104 public web and computer-use benchmarks (via BenchmarkList.com) before building our own. None combined the three properties we needed: resistance to saturation, grounding in real professional work, and genuinely long horizons. Composite-Bench has all three. Success means matching a certified-optimal solution, difficulty is established before any model sees a task, and the compute contract (token budget, reasoning effort, serving route) is pinned, checked at runtime, and published with every row.

We built it for internal use, but the results seemed worth sharing, so we’re publishing the benchmark and the numbers.

The leaderboard

Here is the frontier at pinned maximum reasoning effort. The benchmark has five kinds of tasks (“slices”): pipeline planning, rule extraction, entity deduplication, routing, and information gathering (described below). Each kind has 20 hidden test tasks (“test instances”), so every model faced the same 100 tasks. A model gets one attempt per task (“pass@1”), and it succeeds only if its final answer exactly matches a provably optimal solution. Nothing was thrown out: if a model ran out of budget, hit the step cap, or stopped early, that attempt still counts, scored on whatever it had done. The share of the 100 tasks a model succeeds on is the “macro success” number below.¹

Weights and measures Open weights Closed 0% 20% 40% 60% 80% Claude Fable 5 74% Claude Fable 5 · 74% macro · 95% CI [65, 82] · mean reward 0.965 · $1,166 total eval Claude Opus 4.8 73% Claude Opus 4.8 · 73% macro · 95% CI [64, 81] · mean reward 0.967 · $605 total eval GLM-5.2 63% GLM-5.2 · 63% macro · 95% CI [53, 72] · mean reward 0.936 · $107 total eval Kimi-K3 31% Kimi-K3 · 31% macro · 95% CI [23, 41] · mean reward 0.799 · $187 total eval Gemini 3.5 Flash 27% Gemini 3.5 Flash · 27% macro · 95% CI [19, 36] · mean reward 0.704 · $286 total eval MiniMax-M3 26% MiniMax-M3 · 26% macro · 95% CI [18, 35] · mean reward 0.713 · $40 total eval GPT-5.6 Sol 18% GPT-5.6 Sol · 18% macro · 95% CI [12, 27] · mean reward 0.745 · $671 total eval Grok-4.5 9% Grok-4.5 · 9% macro · 95% CI [5, 16] · mean reward 0.728 · $196 total eval
ModelMacro success95% CIMean rewardNPP ladder*Eval cost²
🥇 Claude Fable 574%[65, 82]0.9650.471$1,166
🥇 Claude Opus 4.873%[64, 81]0.9670.367$605
🥉 GLM-5.263%[53, 72]0.9360.196$107
Kimi-K331%[23, 41]0.7990.412$187
Gemini 3.5 Flash27%[19, 36]0.7040.085$286
MiniMax-M326%[18, 35]0.7130.062$40
GPT-5.6 Sol18%[12, 27]0.7450.366$671
Grok-4.59%[5, 16]0.7280.352$196

*NPP, the sixth kind of task, is scored separately because it is designed to be impossible. The task is to split a set of large numbers into two groups whose sums come out as equal as possible. The statement is simple, but we generate versions where finding the perfect split would take vastly more computation than any model has in a single session. We preregistered a prediction that no model would ever find the perfect answer, and none did in 160 attempts. Since “did you find the perfect answer” is always no, that cannot rank models on success rate. Instead, the ladder score measures how close a model’s best answer gets to the certified optimum on a log scale: 0 means no better than an arbitrary split, 1 means the perfect answer (which never happened), and standard algorithms of increasing quality land at known points in between, so the score reads as which tier of method the model’s answer matches. We’ve excluded this from the macro headline because a task everyone fails at drags every model down by the same amount and tells you nothing about which is better. What it does reveal is how good an answer a model can hold onto when the problem cannot be finished.

Five observations stand out:

  • The launch-week open-weights ranking inverts. Most Kimi-vs-GLM comparisons this week landed on “Kimi for capability, GLM for price.” On long-horizon browser use, the order flips.¹
  • Kimi-K3 is still strong. Kimi clears OpenAI’s flagship on the macro and takes second of all eight models on the NPP search ladder (0.412, behind only Fable).
  • The top is essentially a tie, with different per-slice profiles. Fable sweeps entity-deduplication (20 of 20, every partition exactly optimal); Opus wins spec-extraction and adaptive information-gathering.
  • Grok’s answers look finished but usually aren’t. Its exact success is the board’s lowest (9%), while its mean shaped reward is within 0.02 of Sol’s; 7% of its episodes end at 95%+ solution quality without reaching the optimum. A goal-satisfaction benchmark would rank it several places higher, which is why we score against exact solvers.
  • Price and score don’t track closely. MiniMax’s whole run cost $40 and statistically matches OpenAI’s flagship. GLM hit 85% of the leader’s score for 9% of the cost; Kimi’s run cost 1.7× GLM’s for half the score.²
Worth every penny? Open weights Closed 0% 20% 40% 60% 80% $50 $100 $200 $500 $1,000 Total evaluation cost (USD, log scale) cost-capability frontier Claude Fable 5 Claude Fable 5 · 74% macro · $1,166 total eval · closed Claude Opus 4.8 Claude Opus 4.8 · 73% macro · $605 total eval · closed GLM-5.2 85% of the leader’s score at 9% of the bill GLM-5.2 · 63% macro · $107 total eval · open weights Kimi-K3 Kimi-K3 · 31% macro · $187 total eval · open weights Gemini 3.5 Flash Gemini 3.5 Flash · 27% macro · $286 total eval · closed MiniMax-M3 MiniMax-M3 · 26% macro · $40 total eval · open weights GPT-5.6 Sol GPT-5.6 Sol · 18% macro · $671 total eval · closed Grok-4.5 Grok-4.5 · 9% macro · $196 total eval · closed

For context, Kimi-K3 outranks GLM-5.2 on the intelligence indices, tops Arena’s frontend-code leaderboard (Claude included), and places near the top of a knowledge-work benchmark. Kimi’s second place on our NPP ladder is the same signal.

What the tasks are

Six slices. Each is a deterministic enterprise-web app the agent controls through observe/act tool calls. Every slice has 25 instances: 5 public development instances and 20 held-back test instances, which produce all scored results, so nobody can train on a scored instance. The sixth slice, npp, is excluded from the macro by preregistration (exact success was predicted to be 0% for every model, and it was) and is reported on its own ladder. The five that feed the headline macro:

SliceCapability under test
steward_7dintertemporal resource allocation over a 7-day horizon
rulebook_7dextracting operating rules from a long prose handbook, then planning under them
dedupentity-resolution clustering (partition optimization under transitivity)
orienteeringrouting / subset selection with deep local optima
gauntlet_7gadaptive information-gathering under a spend budget

Per-slice exact successes,³ out of 20 test instances (bold marks each slice’s winner):

Modelsteward_7drulebook_7ddeduporienteeringgauntlet_7g
Claude Fable 51310201318
Claude Opus 4.81113171319
GLM-5.2411171417
Kimi-K3018913
Gemini 3.5 Flash0112212
MiniMax-M3011375
GPT-5.6 Sol003312
Grok-4.500711

As mentioned above, the sixth slice, npp, appears as its own column in the leaderboard. The underlying task is number partitioning: split a set of large numbers into two groups whose sums land as close together as possible. We generate it at the phase transition, the regime where the problem is provably at its hardest and even specialized solvers need orders of magnitude more search than any in-context token budget allows. Across eight models and 160 rows, observed exact success is 0 for 160, matching the preregistered prediction. What the slice measures instead is the quality of the answer a model holds when the problem is impossible to solve optimally. The ladder ordering differs substantially from the macro: Fable leads (0.471); Kimi-K3 is second (0.412); Grok is fifth (0.352, at parity with Opus and Sol). The slice is designed to separate performance on solvable problems from performance on problems that cannot be solved within budget.

Where the tasks come from

Each slice is a distilled form of work we help professionals do every week inside authenticated enterprise tools. The released instances intentionally carry no personal identities, no recorded trajectories, and, crucially, no customer data. Still, the task shapes are as closely representative as possible.

Each slice also carries GDPval sector and occupation labels based on the sessions the tasks were derived from.

  • steward: working a pipeline under budgets and deadlines. Direct ancestors from our corpus include revenue teams working deal pipelines week over week, EMS coordinators building shift and training schedules against daily staffing budgets, accountants sequencing AP/AR against a close date. The shared core: a fixed daily effort budget, actions with different costs and lead times, hard expiry dates, and an irreversible clock. The agent decides what to work on each day so that the week ends with the most value locked in. Models that plan only one day ahead score poorly here; so do people who work that way.
    • Wholesale: Sales Representatives (technical and scientific products)
    • Finance: Financial and Investment Analysts
    • Professional Services: Accountants and Auditors
    • Retail: General and Operations Managers
    • Health Care: Medical and Health Services Managers
  • rulebook: the rules live in a document rather than the interface. In nearly every role we observe, the operating rules of the work live in SOPs, policy wikis, and procedure manuals: tiered client-handling policies, leave-of-absence rules, filing procedures, amendments that supersede base clauses, and provisions that simply don’t apply to the case at hand. Rulebook wraps a steward-class planning problem in a long prose handbook with distractors and plausible misreadings, scored on whether the agent extracts the governing rules and then plans under them.
    • Professional Services: Lawyers; Accountants and Auditors
    • Government: Compliance Officers
    • Health Care: First-Line Supervisors of Office and Administrative Support Workers
  • dedup: entity resolution is everywhere. Financial-crime analysts decide whether entities appearing across alerts are the same actor. CRM administrators merge duplicate accounts. Records clerks match case files. The trap in all of them is transitivity: merging A with B and B with C silently asserts A is C. The slice scores exact partition optimality, so both over-merging and under-merging lose, as they do in a real CRM or case system.
    • Government: Compliance Officers
    • Professional Services: Software Developers
    • Finance: Financial and Investment Analysts
    • Wholesale: Sales Representatives
  • orienteering: opportunity cost of time. Perhaps the most universal pattern in our corpus: a queue of items of unequal value, a limited session, and order-dependent costs. Select ancestors include case officers choosing which matters to advance, account managers choosing which clients to touch this week, coordinators sequencing site visits. The logical core is prize-collecting routing with deep local optima; it penalizes the greedy shortcuts that people and models reach for first.
    • Health Care: Registered Nurses
    • Retail: General and Operations Managers
    • Finance: Financial and Investment Analysts
    • Real Estate: Property and Real Estate Managers
    • Professional Services: Lawyers
  • gauntlet: opportunity cost of retrieving information. Fraud and financial-crime work is adaptive probing: every records pull, log query, or alert expansion costs time or budget, and what you check next should depend on what the last check revealed. Medical-records review and legal evidence review have the same shape. The slice scores whether probing is adaptive and whether the agent commits to an answer before the budget is gone.
    • Government: Compliance Officers
    • Retail: Private Detectives and Investigators
    • Professional Services: Lawyers
    • Health Care: Registered Nurses
    • Finance: Financial and Investment Analysts
  • npp: the control condition, pure search under a binding budget. Real work contains embedded optimization cores: balancing crews across shifts, splitting a book of business evenly across account managers, netting payment batches, assigning balanced cohorts. npp strips away the workflow shell to isolate that core at certified-unsolvable difficulty, so what remains measurable is anytime discipline: whether the model maintains its best answer and improves it monotonically. The failure mode it measures is an agent that continues working and ends with a worse answer than one it held earlier; anytime scoring prices that directly.

These are not loose analogies. Each traces to identified sessions in the corpus the benchmark is calibrated against. steward’s deal-pipeline ancestor is an account executive’s multi-week CRM arcs; its accounting variant is a practitioner sequencing payables and receivables against a close date; its clinical variant is a telehealth provider working dual visit queues, whose replayed work is one of our two measured human baselines. orienteering’s routing structure comes from home-health and field-service visit sequencing. dedup’s seed is an engineer’s recurring account-deduplication queue. gauntlet’s adaptive-probing shape matches financial-crime, medical-records, and legal-evidence review, down to an attorney assembling discovery materials for review. rulebook wraps the steward planning problem in the kind of policy handbook that governs most of these roles. npp is the exception by design: it strips the workflow shell away and so carries no occupational lineage.

Across that calibration corpus the coverage is broad. Its professionals span eight of GDPval’s nine sectors and map to 13 of the 44 occupations; the ninth sector, Government, appears as public-sector legal work, which GDPval files under Professional Services. Just under half of the corpus’s recorded work comes from users in a GDPval-44 occupation, and roughly two thirds from occupational work of any kind; the rest is student, personal, or too sparse to classify. The corpus also holds work GDPval does not cover, such as higher education and telecommunications field service.

Released environments

Though our internal evaluation and training environments are higher fidelity and rebuilt from real behind-login enterprise state, for data security and privacy reasons they cannot be released. The public environments are deliberately schematic: compact mock enterprise applications operated through a plain observe/act tool interface over an accessibility-tree-style screen, with no pixels, no real DOM, and no production UI chrome. Each distills a workflow to its decision core, keeping the decisions professionals actually make and allowing an exact solver to certify outcomes objectively.

Benchmark integrity

The benchmark is built around four integrity properties:

  • Optimality-certified gold. Per-instance exact solvers define success, and a shaped reward in [0,1] measures distance from optimal. A 100% score means every instance was solved to certified optimality.
  • A model-visible, fail-closed compute contract. The episode token budget¹ is announced to the model and tracked live. Reasoning-effort pins are verified before any paid row and fail closed. Providers run under closed, preflighted allowlists with every turn’s serving provider recorded and published.
  • Zero exclusions. Anytime scoring makes every episode terminal and scoreable: no redraws, no dropped rows, zero infrastructure exclusions across every published row. Budget exhaustion is scored as an outcome rather than excluded.
  • Chinese wall. Every instance regenerates bit-identically from (genre, params, seed). Test seeds live in a reserved range excluded from training curricula and stay server-side, so no model can train on (or see) a scored instance.

Grounding in real expert browser-use data

Synthetic environments require calibration against real work. Our products assist hundreds of thousands of professionals daily, which gives us one of the largest corpora of real expert browser-/computer-use data anywhere, including from behind-login enterprise surfaces. As such:

  • Task families come from observed work. Every slice traces to patterns that recur across our corpus.
  • Human expert baselines close the loop. Professionals’ own actual work replays through environments created from that work, scored by the same class of exact verifiers. Across our first two measured cohorts, professionals achieved a median 0.80 and 0.89 of the verifier’s optimum within their own effort envelope. That calibration loop of real experts, replayable environments, and exact scoring allows us to ground the benchmark’s design and difficulty dials.

Running it yourself

Feel free to try this yourself: the harness, generators, and public dev instances are freely available on GitHub. If you have any questions or requests, write to us at research@composite.com.

Work with us

Most of the interesting problems behind this benchmark are still open, and solving them would help us figure out how to create and enable models that let browser-use agents help professionals do their best work. That is the day-to-day ML engineering at Composite. If you’d like to join our team, we’re hiring! composite.com/careers.


Footnotes

¹ Protocol: pinned maximum reasoning effort per provider; announced episode budget of 300,000 output tokens (reasoning included) tracked live and injected into every tool result; per-turn cap min(64k, remaining); 110-step cap; anytime scoring, meaning every episode is terminal and scored as-is; pass@1 over 20 test instances per slice; Wilson 95% CIs. Ties: at n=100 the macro resolves ~10pp, so gaps below that are stated as ties and gaps above it as ranks.

Why 300k: the budget was fixed in the preregistration, before any scored row, at the level that covers roughly the 90th percentile of cumulative output tokens over each model’s clean episodes under the benchmark’s original protocol (per-model p90s: GLM 288k, Sol 287k, Fable 262k, Opus 259k). The original v1.0 protocol had no episode-level budget at all: each turn was capped at 24,000 output tokens, a cap itself raised from an early 6k harness default after reasoning models hit the ceiling mid-thought and emitted zero actions, and truncated turns drew retry nudges. v1.1 replaced per-turn rationing with one announced episode budget so each model plans its own spend; the 64k per-turn ceiling produced zero truncated turns across 201 observed turns of earlier piloting, and nudges were removed entirely.

² Costs are row-recorded native usage × list prices over each model’s full run, prompts included. Budget-exhaustion rates vary widely: Gemini Flash ended 34% of its episodes at budget exhaustion and MiniMax 26% (both disclosed per our >15% prominent-disclosure rule); Grok 0.7%; Kimi 11%, all of it on the npp slice and none on its macro rows.

³ Selection precedes model runs: the oracle scores 1.0, a no-op scores 0, and a strong heuristic baseline scores below 1.0 on every kept instance; hill-climb gap ensembles gate every slice, and npp’s search resistance is certified per seed. Successor benchmarks typically ship after saturation is reached (ScreenSpot→Pro, WebArena→WebChoreArena, OSWorldVerified); Composite-Bench certifies headroom before any model run.

Budget-binding on npp is ~95% for the deep reasoners by construction (disclosed before the campaign) vs ~30% for Sol; budget exhaustion is the designed outcome on this slice, since the instances are certified unsolvable within the budget.

One serving stack does not bound reasoning tokens with max_tokens: 5 of 5,509 Grok turns overran the 64k per-turn request, one reaching 499,996 tokens, producing the only episode in the entire published record to exceed the episode budget’s structural bound (scored as-is under the anytime contract, reward 0). A ninth model, Muse-Spark-1.1, attempted the benchmark but was excluded fail-closed by our serving-maturity gate before any scored row ran; Kimi-K3’s arm worked through two documented serving-layer defects on its first-party lane under the same gate machinery, with its storm history disclosed per-row. Full serving findings are in the release notes.


Appendix A: the compute contract matters (and default-effort leaderboards mislead)

Four of the eight models (Fable, Opus, GLM, Sol) were also measured under the provider-default contract on the benchmark’s original slice set, and the board reorders substantially there: Fable 0.738 ≫ GLM 0.343 ~ Opus 0.294 ~ Sol 0.263 (the mid-table statistically unresolved). Between the default and maximum contracts the ranking itself reorders: Opus +0.45, GLM +0.20, Fable +0.05, Sol −0.28 on matched slices (descriptive deltas; protocol and effort change together). A single-effort leaderboard mis-states frontier capability in both directions, which is why every Composite-Bench table is protocol-versioned and results are never pooled across protocols. (Gemini Flash, MiniMax, Grok, and Kimi were measured at the maximum contract only, against the same never-public test seeds as everyone else, so no default-contract delta exists for them.) Sol’s collapse under the more generous contract is not noise: with a tight per-turn ceiling and truncation nudges, Sol solved 19/20 dedup seeds (on at least one of two rollouts); given a large announced budget and no nudges, it spent the budget refining and failed to lock in answers. The tight ceiling and nudges were what forced it to commit.

Appendix B: reproducibility

Preregistrations (binding from first paid row), amendments disclosed with evidence, per-row artifacts recording the served model and per-turn {provider, finish reason, native usage, generation ID}, and reference transcripts for all eight models on every public dev instance under the current protocol, so any team can calibrate its harness against ours before trusting or submitting results. Public package (GitHub): runnable dev instances for every slice with the full harness and generators; the scored test seeds stay server-side until version retirement.