GPU Rental Price per Effective FP8 PFLOP
Newest print Apr 2026Reviewed monthlyCSV — newest observations; full series with a Pro keyAPI
GPU Rental Price per Effective FP8 PFLOP — latest readings
9 observations across 3 series, Q4 2023–Q2 2026.
| Series | Latest ($/hr per PFLOP) | Source |
|---|---|---|
| H100 (on-demand median)Q2 2026 | 1.01 | Our model estimateSilicon Analysts modelApr 2026 |
| B200 (on-demand median)Q2 2026 | 0.78 | Our model estimateSilicon Analysts modelApr 2026 |
| H200 (on-demand median)Q2 2026 | 1.3 | Our model estimateSilicon Analysts modelApr 2026 |
Derived price-per-performance series for rented AI GPUs: public on-demand rental prices divided by NVIDIA peak FP8 (sparse) throughput — H100 SXM5 and H200 at 3,958 TFLOPS, B200 at 9,000 TFLOPS. Tracks whether new silicon generations are priced at a per-performance discount (the neocloud pattern) or at per-PFLOP parity/premium (the hyperscaler scarcity pattern). Companion dataset to the GPU rental premium decomposition analysis.
H200 (on-demand median): $1.3/hr per PFLOP (Q2 2026), Silicon Analysts (derived), Apr 2026
2023–2024: 2 earlier prints — full ledger is Pro
- 1Q1 2025 — AWS H100 -44% (Jun 2025): June 1, 2025: AWS cut H100 P5 on-demand pricing 44% (alongside H200 P5en -25-26% and A100 P4d -33%), resetting the hyperscaler price umbrella over the GPU rental market.
- 2Q1 2026 — AWS H200 Blocks +15% (Jan 2026): January 15, 2026: AWS raised H200 capacity-block pricing ~15%, partially reversing the June 2025 cuts — scarcity pricing returning at the memory-rich end of the ladder.
Source: Silicon Analysts model · Basis: derived ratio · Anonymous view: the newest 1–3 observations per series (Nov 2024–May 2026); 2 earlier prints back to 2023 — full ledger is Pro.
Prints ledger
| Note and copy | |||||||
|---|---|---|---|---|---|---|---|
| Q2 2026 | H100 (on-demand median)Our model estimate | 1.010.68–$1.56/hr per PFLOP | |||||
Source: Silicon Analysts (derived), Apr 2026 · Basis: derived ratio · Estimate Derived: median of neocloud H100 on-demand LIST prices in the Apr 2026 snapshot (RunPod $2.69, Lambda $3.99, CoreWeave $6.16 → median $3.99) ÷ 3.958 PFLOPS FP8 sparse. Basis change vs prior points: three-provider median, not the single CoreWeave observation. Corrected 2026-07-10: CoreWeave observation updated from the retired $4.76 component rate (median unchanged). | |||||||
| Q2 2026 | B200 (on-demand median)Our model estimate | 0.780.61–$0.96/hr per PFLOP | |||||
Source: Silicon Analysts (derived), Apr 2026 · Basis: derived ratio · Estimate Derived: median of B200 on-demand list prices in the Apr 2026 snapshot (Lambda $5.50, CoreWeave $8.60 → median $7.05) ÷ 9.0 PFLOPS FP8 sparse. ~23% below the H100 on-demand median per effective PFLOP; Lambda alone is ~39% below its own H100. Corrected 2026-07-10: CoreWeave B200 observation added. | |||||||
| Q2 2026 | H200 (on-demand median)Our model estimate | 1.31.01–$1.59/hr per PFLOP | |||||
Source: Silicon Analysts (derived), Apr 2026 · Basis: derived ratio · Estimate Derived: median of neocloud H200 on-demand list prices in the Apr 2026 snapshot (RunPod $3.99, CoreWeave $6.31 → median $5.15) ÷ 3.958 PFLOPS FP8 sparse. Corrected 2026-07-10: previously used an unsupported CoreWeave H200 rate of $3.89; CoreWeave’s published price is $50.44/hr per 8-GPU HGX H200 instance ($6.31/GPU-hr). | |||||||
| Q1 2026 | H100 (on-demand median)Our model estimate | 1.56 | |||||
Source: Silicon Analysts (derived), Jan–Mar 2026 · Basis: derived ratio · Estimate Derived: CoreWeave H100 $6.16/GPU-hr (8-GPU bundle $49.24/hr, sole on-demand list observation) ÷ 3.958 PFLOPS FP8 sparse. Basis note: bundle-inclusive (vCPU/RAM/storage), not directly comparable to the 2023–24 GPU-component points. Corrected 2026-07-10: previously derived from a $1.60 effective-rate estimate. | |||||||
| Q1 2026 | B200 (on-demand median)Our model estimate | 0.96 | |||||
Source: Silicon Analysts (derived), Jan–Mar 2026 · Basis: derived ratio · Estimate Derived: CoreWeave HGX B200 $8.60/GPU-hr (8-GPU bundle $68.80/hr, sole on-demand list observation) ÷ 9.0 PFLOPS FP8 sparse. Corrected 2026-07-10: previously derived from a $4.50 SA estimate not supported by CoreWeave’s published list. | |||||||
| Q1 2025 | B200 (on-demand median)Our model estimate | 0.720.61–$0.89/hr per PFLOP | |||||
Source: Silicon Analysts (derived), Mar 2025 · Basis: derived ratio · Estimate Derived: early B200 rental estimate ~$6.50/GPU-hr (range $5.50-8.00, analyst estimate — no public B200 list price existed; CoreWeave’s first Blackwell listing was GB200 NVL72 at $10.50/GPU-hr) ÷ 9.0 PFLOPS FP8 sparse. Even at launch scarcity pricing, Blackwell entered near Hopper per-PFLOP list levels. | |||||||
| Q4 2024 | H100 (on-demand median)Our model estimate | 1.2 | |||||
Source: Silicon Analysts (derived), Dec 2024 · Basis: derived ratio · Estimate Derived: CoreWeave H100 on-demand $4.76/GPU-hr (Dec 2024, list unchanged; sole on-demand list observation) ÷ 3.958 PFLOPS FP8 sparse. Corrected 2026-07-10: previously derived from a $1.89 spot-tracker figure, which is not an on-demand list price. | |||||||
| H100 (on-demand median): 2 earlier prints, Q4 2023–Q2 2024 — full ledger is Pro | |||||||
| Showing the newest 1–3 observations per series (Q4 2024–Q2 2026); 2 earlier observations (Q4 2023–Q2 2024) are in Pro. | |||||||
Provenance
Derived series — provider public list prices ÷ NVIDIA peak FP8 sparse throughput (H100 SXM5 and H200: 3,958 TFLOPS; B200: 9,000 TFLOPS, per lib/constants/chipSpecs.ts). All values are Silicon Analysts estimates; providers do not publish per-performance pricing. Each point's source note names the underlying price observation.
Basis: derived ratio
Derived: provider list price divided by peak FP8 sparse throughput. Silicon Analysts estimate — no provider publishes per-performance pricing.
Our derivations
- Silicon Analysts (derived) (Dec 2023)
- Silicon Analysts (derived) (Jun 2024)
- Silicon Analysts (derived) (Dec 2024)
- Silicon Analysts (derived) (Mar 2025)
- Silicon Analysts (derived) (Jan–Mar 2026)
- Silicon Analysts (derived) (Apr 2026)
Suggested citation
Silicon Analysts, "GPU Rental Price per Effective FP8 PFLOP", as of Apr 2026, https://siliconanalysts.com/market-data/gpu-rental-price-per-tflop, retrieved September 29, 2026.
API: /api/v1/market-data/gpu-rental-price-per-tflopCSVLicence
Showing the newest 1–3 observations per series (Q4 2024–Q2 2026); 2 earlier observations (Q4 2023–Q2 2024) are in Pro.
Related Tools
Related Analysis
Frequently Asked Questions
- How much does it cost to rent a B200 per hour?
- As of the April 2026 snapshot, Lambda lists the B200 on-demand at $5.50 per GPU-hour ($44.00 for an 8-GPU instance), while AWS capacity blocks price the B200 at $9.36 per GPU-hour ($74.88 for the 8-GPU p6.48xlarge). Per unit of FP8 (sparse) compute, Lambda’s B200 works out to roughly $0.61 per effective PFLOP-hour — about 39% below its own H100 — while AWS prices Blackwell at per-PFLOP parity with the H100.
- What does "$/hr per PFLOP" mean for GPU rental pricing?
- It is a derived price-per-performance metric: the public on-demand rental price per GPU-hour divided by the chip’s peak FP8 sparse throughput in PFLOPS (H100 SXM5 and H200: 3.958; B200: 9.0, per NVIDIA spec sheets — dense throughput is roughly half). Providers do not publish per-performance pricing, so all values are Silicon Analysts estimates; each data point cites the underlying price observation.
- Is the B200 cheaper than the H100 per unit of compute?
- On neoclouds, yes: Lambda rents the B200 at 1.38x its H100 price for 2.27x the FP8 (sparse) throughput — about 39% cheaper per effective PFLOP ($0.61 vs $1.01). On AWS capacity blocks the B200 is priced at per-PFLOP parity with the two-year-old H100 (~$1.04 vs ~$0.99), meaning the newest silicon carries zero performance discount — hyperscalers price scarcity, neoclouds price compute.