GPU node economics · companion to the Index

One dedicated box: monthly cost and task capacity.

For each AWS GPU server type we benchmarked: the two best models on it, the committed monthly cost of the box, and the tasks per month it supports. Rankings come from the same document extraction and retrieval workloads the Index scores.

us-west-2 · Linux, shared tenancy · pricing pulled 2026-08-05 from the AWS Price List and Savings Plans APIs · throughput measured at 64-concurrent load

g6e.xlarge

1 × L40S

NVIDIA L40S · 48 GB GDDR6

$511/mo dedicated

3-yr all-upfront Savings Plan, $0.700/hr
on-demand: $1,359/mo

01

Phi-3.5-mini-instruct

BF16 · 3.8B · text

3,494
peak tok/s
23.0M
tasks/mo
$0.02
per 1k tasks
02

Qwen2.5-VL-7B

BF16 · 7B · vision

2,375
peak tok/s
15.6M
tasks/mo
$0.03
per 1k tasks

Dev tier ranked by measured throughput. The RAG quality gold does not cover models this small.

p5.48xlarge

8 × H100

NVIDIA H100 SXM · 80 GB HBM3 each · 640 GB total

$15,107/mo dedicated

3-yr all-upfront Savings Plan, $20.70/hr
on-demand: $40,179/mo

01

DeepSeek-V3.1

AWQ INT4 · retrieval 86.3% · extraction 99.1 (clean)

2,274
peak tok/s
14.9M
tasks/mo
$1.01
per 1k tasks
02

GLM-4.5-Air

BF16 · 106B MoE · retrieval 83.6%

6,149
peak tok/s
40.4M
tasks/mo
$0.37
per 1k tasks

The 80–86% retrieval band is a statistical tie; the order here is the point estimate. GLM-4.7 also scores 83.6%; GLM-4.5-Air takes the second slot on throughput, 6,149 tok/s against 3,120.

p5en.48xlarge

8 × H200

NVIDIA H200 · 141 GB HBM3e each · 1,128 GB total

$17,373/mo dedicated

3-yr all-upfront Savings Plan, $23.80/hr
on-demand: $46,206/mo

01

Qwen3-VL-235B

FP8 · vision · retrieval 85.0% · extraction 97.9

3,838
peak tok/s
25.2M
tasks/mo
$0.69
per 1k tasks
02

Ornith-1.0-397B

BF16 · 397B MoE · vision · retrieval 85.0%

2,710
peak tok/s
17.8M
tasks/mo
$0.98
per 1k tasks

Qwen3-VL and Ornith tie at 85.0% on the point estimate; the order here is throughput. DeepSeek-V4-Pro (80.1%) and Kimi-K2.5 (82.3%) sit lower in the band; V4-Pro also serves at 516 tok/s through a Marlin FP4 fallback on Hopper.

Capacity

Tasks per box, per month.

One task is one grounded answer, capped at 400 output tokens. The bars show each box at its measured peak over a 730-hour month; multiply by your expected utilization.

Task capacity · at measured peak · shared scale
g6e.xlarge · 1× L40S · $511/mo
Phi-3.5-mini23.0M · 3,494 t/s
Qwen2.5-VL-7B15.6M · 2,375 t/s
p5.48xlarge · 8× H100 · $15,107/mo
DeepSeek-V3.114.9M · 2,274 t/s
GLM-4.5-Air40.4M · 6,149 t/s
p5en.48xlarge · 8× H200 · $17,373/mo
Qwen3-VL-235B25.2M · 3,838 t/s
Ornith-1.0-397B17.8M · 2,710 t/s
# tasks/mo = peak tok/s × 2,628,000 s ÷ 400 tokens · t/s = output tokens per second

Method and caveats

The quality scores behind this ranking are in the Index.

Extraction, retrieval, and serving cost for every model, measured on our stack and beside the same weights on Bedrock.