GPU node economics · companion to the Index
For each AWS GPU server type we benchmarked: the two best models on it, the committed monthly cost of the box, and the tasks per month it supports. Rankings come from the same document extraction and retrieval workloads the Index scores.
us-west-2 · Linux, shared tenancy · pricing pulled 2026-08-05 from the AWS Price List and Savings Plans APIs · throughput measured at 64-concurrent load
1 × L40S
NVIDIA L40S · 48 GB GDDR6
$511/mo dedicated
3-yr all-upfront Savings Plan, $0.700/hr
on-demand: $1,359/mo
Phi-3.5-mini-instruct
Qwen2.5-VL-7B
Dev tier ranked by measured throughput. The RAG quality gold does not cover models this small.
8 × H100
NVIDIA H100 SXM · 80 GB HBM3 each · 640 GB total
$15,107/mo dedicated
3-yr all-upfront Savings Plan, $20.70/hr
on-demand: $40,179/mo
DeepSeek-V3.1
GLM-4.5-Air
The 80–86% retrieval band is a statistical tie; the order here is the point estimate. GLM-4.7 also scores 83.6%; GLM-4.5-Air takes the second slot on throughput, 6,149 tok/s against 3,120.
8 × H200
NVIDIA H200 · 141 GB HBM3e each · 1,128 GB total
$17,373/mo dedicated
3-yr all-upfront Savings Plan, $23.80/hr
on-demand: $46,206/mo
Qwen3-VL-235B
Ornith-1.0-397B
Qwen3-VL and Ornith tie at 85.0% on the point estimate; the order here is throughput. DeepSeek-V4-Pro (80.1%) and Kimi-K2.5 (82.3%) sit lower in the band; V4-Pro also serves at 516 tok/s through a Marlin FP4 fallback on Hopper.
Capacity
One task is one grounded answer, capped at 400 output tokens. The bars show each box at its measured peak over a 730-hour month; multiply by your expected utilization.
Method and caveats
Extraction, retrieval, and serving cost for every model, measured on our stack and beside the same weights on Bedrock.