Enterprise OSS LLM Index · Q2 2026
A quarterly, reproducible benchmark of open-weight models on the document work regulated firms actually run: extraction from degraded scans, retrieval under a fixed judge, and the cost to serve each at scale. Run on our own in-VPC stack and beside the same weights on Amazon Bedrock. It's open to read, with no email gate. The email below is only if you want a note when the next edition ships.
Q2 2026 edition · reproducible · synthetic data · no customer data · full writeup →
Executive summary
Q2 2026 is the first edition. We ran six open-weight models on three measures that matter for regulated document work, on our own stack inside a VPC and beside the same weights on Bedrock and the frontier.
The headline result is the one that lets you stop worrying about who hosts the model: the same weights score the same regardless of who serves them. Extraction F1 on our stack matched Bedrock to the decimal (97.9% against 97.9% on the strongest vision model), and retrieval matched within CI across the set. So where you run an open-weight model is a cost-and-control decision, not a quality one.
On the numbers, no model wins every axis. On retrieval, the repaired metric puts the whole lineup, frontier included, in one 80–86% band with overlapping confidence intervals: no model separates from the pack, so retrieval rank does not pick your model. Qwen3-VL-235B led extraction at 97.9% F1 on degraded scans. Cost separates the field instead: Llama-4-Scout is the cheapest to serve at $1.01 per million output tokens, and DeepSeek-V4-Pro, the newest and largest at 1.6 trillion parameters, is the most expensive at $11.71. The Index is a map of those trade-offs, not a leaderboard with one trophy.
Correction, August 2026: an earlier version of this page ranked retrieval on a defective metric. The judge never saw the retrieved passages, and the question set contained duplicated and ambiguous items. We repaired both, validated the new judge against a human-scored anchor set, and re-ran the lineup; the retrieval numbers below are the repaired ones. Details in the writeup.
Results
| Model | Extract F1 | Retrieval | $/M out |
|---|---|---|---|
| DeepSeek-V4-Pro | — | 80.1% | $11.71 |
| Qwen3-VL-235B | 97.9% | 85.0% | $1.30 |
| Llama-4-Scout | 91.3% | 83.2% | $1.01 |
| GLM-4.7 | — | 83.6% | $1.27 |
| Kimi-K2.5 | — | 82.3% | $2.18 |
| DeepSeek-V3.2 | — | 84.5% | $3.56 |
Methodology
Per-model notes
License terms and per-model deployment notes are in the writeup. Always confirm a model's license against your own use before deploying.
Node economics
Every score in the Index was measured on a dedicated AWS GPU server. We priced those servers on a page of their own. It lists the two best models for each instance type, the committed monthly cost of the box, and the number of tasks one box can serve in a month at the throughput we measured.
The Index is open and free to read. Leave your email and we'll send one note when the Q3 2026 edition is out. One field, no sequence.
Editions: Q2 2026 — current. This is the first edition; the archive grows each quarter.