Enterprise OSS LLM Index · Q2 2026

The open-weight models worth self-hosting, measured.

A quarterly, reproducible benchmark of open-weight models on the document work regulated firms actually run: extraction from degraded scans, retrieval under a fixed judge, and the cost to serve each at scale. Run on our own in-VPC stack and beside the same weights on Amazon Bedrock. It's open to read, with no email gate. The email below is only if you want a note when the next edition ships.

Q2 2026 edition · reproducible · synthetic data · no customer data · full writeup →

Executive summary

The host doesn't change the score. The trade-offs are yours to make.

Q2 2026 is the first edition. We ran six open-weight models on three measures that matter for regulated document work, on our own stack inside a VPC and beside the same weights on Bedrock and the frontier.

The headline result is the one that lets you stop worrying about who hosts the model: the same weights score the same regardless of who serves them. Extraction F1 on our stack matched Bedrock to the decimal (97.9% against 97.9% on the strongest vision model), and retrieval matched within CI across the set. So where you run an open-weight model is a cost-and-control decision, not a quality one.

On the numbers, no model wins every axis. On retrieval, the repaired metric puts the whole lineup, frontier included, in one 80–86% band with overlapping confidence intervals: no model separates from the pack, so retrieval rank does not pick your model. Qwen3-VL-235B led extraction at 97.9% F1 on degraded scans. Cost separates the field instead: Llama-4-Scout is the cheapest to serve at $1.01 per million output tokens, and DeepSeek-V4-Pro, the newest and largest at 1.6 trillion parameters, is the most expensive at $11.71. The Index is a map of those trade-offs, not a leaderboard with one trophy.

Correction, August 2026: an earlier version of this page ranked retrieval on a defective metric. The judge never saw the retrieved passages, and the question set contained duplicated and ambiguous items. We repaired both, validated the new judge against a human-scored anchor set, and re-ran the lineup; the retrieval numbers below are the repaired ones. Details in the writeup.

Results

Q2 2026, per model.

Retrieval quality · self-hosted in-VPC
One fixed context-aware judge · bars to a 100% axis · $/M output at right
Qwen3-VL-235B85.0% · $1.30
DeepSeek-V3.284.5% · $3.56
GLM-4.783.6% · $1.27
Llama-4-Scout83.2% · $1.01
Kimi-K2.582.3% · $2.18
DeepSeek-V4-Pro80.1% · $11.71
# reproducible · synthetic data · intervals overlap: read the band as a tie
Full table · Q2 2026
Model Extract F1 Retrieval $/M out
DeepSeek-V4-Pro80.1%$11.71
Qwen3-VL-235B97.9%85.0%$1.30
Llama-4-Scout91.3%83.2%$1.01
GLM-4.783.6%$1.27
Kimi-K2.582.3%$2.18
DeepSeek-V3.284.5%$3.56
Extraction = F1 on degraded scans, vision models only. Retrieval under one fixed context-aware judge (n=226); the 80–86% spread is inside the confidence intervals, a statistical tie. $/M out = US$ per million output tokens (serving cost), at spot pricing and peak utilization. Same weights on Bedrock score within CI.

Methodology

What we measured, and what we didn't.

Read the full writeup → Methodology & harness on GitHub ↗

Per-model notes

Where each one fits.

License terms and per-model deployment notes are in the writeup. Always confirm a model's license against your own use before deploying.

Node economics

What the hardware costs, and how many tasks it serves.

Every score in the Index was measured on a dedicated AWS GPU server. We priced those servers on a page of their own. It lists the two best models for each instance type, the committed monthly cost of the box, and the number of tasks one box can serve in a month at the throughput we measured.

g6e.xlarge · 1× L40S
$511per month · 3-yr committed
p5.48xlarge · 8× H100
$15,107per month · 3-yr committed
p5en.48xlarge · 8× H200
$17,373per month · 3-yr committed
See the node economics →

Get the next edition when it ships.

The Index is open and free to read. Leave your email and we'll send one note when the Q3 2026 edition is out. One field, no sequence.

Editions: Q2 2026 — current. This is the first edition; the archive grows each quarter.