GPU Rental Price per Effective FP8 PFLOP

Newest print Apr 2026Reviewed monthlyCSV — newest observations; full series with a Pro keyAPI

GPU Rental Price per Effective FP8 PFLOP — latest readings

9 observations across 3 series, Q4 2023–Q2 2026.

Latest freely available reading for each series in GPU Rental Price per Effective FP8 PFLOP, with its basis, source and date
SeriesLatest ($/hr per PFLOP)Source
H100 (on-demand median)Q2 20261.01Our model estimateSilicon Analysts modelApr 2026
B200 (on-demand median)Q2 20260.78Our model estimateSilicon Analysts modelApr 2026
H200 (on-demand median)Q2 20261.3Our model estimateSilicon Analysts modelApr 2026

Derived price-per-performance series for rented AI GPUs: public on-demand rental prices divided by NVIDIA peak FP8 (sparse) throughput — H100 SXM5 and H200 at 3,958 TFLOPS, B200 at 9,000 TFLOPS. Tracks whether new silicon generations are priced at a per-performance discount (the neocloud pattern) or at per-PFLOP parity/premium (the hyperscaler scarcity pattern). Companion dataset to the GPU rental premium decomposition analysis.

H200 (on-demand median): $1.3/hr per PFLOP (Q2 2026), Silicon Analysts (derived), Apr 2026

2023–2024: 2 earlier prints — full ledger is Pro

  1. 1Q1 2025 — AWS H100 -44% (Jun 2025): June 1, 2025: AWS cut H100 P5 on-demand pricing 44% (alongside H200 P5en -25-26% and A100 P4d -33%), resetting the hyperscaler price umbrella over the GPU rental market.
  2. 2Q1 2026 — AWS H200 Blocks +15% (Jan 2026): January 15, 2026: AWS raised H200 capacity-block pricing ~15%, partially reversing the June 2025 cuts — scarcity pricing returning at the memory-rich end of the ladder.
Press / analyst printSilicon Analysts model

Source: Silicon Analysts model · Basis: derived ratio · Anonymous view: the newest 1–3 observations per series (Nov 2024–May 2026); 2 earlier prints back to 2023 — full ledger is Pro.

Prints ledger

Prints ledger: every observation in this dataset with its period, value, basis, type, confidence and source
Note and copy
Q2 2026H100 (on-demand median)Our model estimate1.010.68–$1.56/hr per PFLOP
Q2 2026B200 (on-demand median)Our model estimate0.780.61–$0.96/hr per PFLOP
Q2 2026H200 (on-demand median)Our model estimate1.31.01–$1.59/hr per PFLOP
Q1 2026H100 (on-demand median)Our model estimate1.56
Q1 2026B200 (on-demand median)Our model estimate0.96
Q1 2025B200 (on-demand median)Our model estimate0.720.61–$0.89/hr per PFLOP
Q4 2024H100 (on-demand median)Our model estimate1.2
H100 (on-demand median): 2 earlier prints, Q4 2023–Q2 2024 — full ledger is Pro
Showing the newest 1–3 observations per series (Q4 2024–Q2 2026); 2 earlier observations (Q4 2023–Q2 2024) are in Pro.

Provenance

Derived series — provider public list prices ÷ NVIDIA peak FP8 sparse throughput (H100 SXM5 and H200: 3,958 TFLOPS; B200: 9,000 TFLOPS, per lib/constants/chipSpecs.ts). All values are Silicon Analysts estimates; providers do not publish per-performance pricing. Each point's source note names the underlying price observation.

Basis: derived ratio

Derived: provider list price divided by peak FP8 sparse throughput. Silicon Analysts estimate — no provider publishes per-performance pricing.

Our derivations

  • Silicon Analysts (derived) (Dec 2023)
  • Silicon Analysts (derived) (Jun 2024)
  • Silicon Analysts (derived) (Dec 2024)
  • Silicon Analysts (derived) (Mar 2025)
  • Silicon Analysts (derived) (Jan–Mar 2026)
  • Silicon Analysts (derived) (Apr 2026)

Suggested citation

Silicon Analysts, "GPU Rental Price per Effective FP8 PFLOP", as of Apr 2026, https://siliconanalysts.com/market-data/gpu-rental-price-per-tflop, retrieved September 29, 2026.

API: /api/v1/market-data/gpu-rental-price-per-tflopCSVLicence

Showing the newest 1–3 observations per series (Q4 2024–Q2 2026); 2 earlier observations (Q4 2023–Q2 2024) are in Pro.

Related Tools

Related Analysis

Frequently Asked Questions

How much does it cost to rent a B200 per hour?
As of the April 2026 snapshot, Lambda lists the B200 on-demand at $5.50 per GPU-hour ($44.00 for an 8-GPU instance), while AWS capacity blocks price the B200 at $9.36 per GPU-hour ($74.88 for the 8-GPU p6.48xlarge). Per unit of FP8 (sparse) compute, Lambda’s B200 works out to roughly $0.61 per effective PFLOP-hour — about 39% below its own H100 — while AWS prices Blackwell at per-PFLOP parity with the H100.
What does "$/hr per PFLOP" mean for GPU rental pricing?
It is a derived price-per-performance metric: the public on-demand rental price per GPU-hour divided by the chip’s peak FP8 sparse throughput in PFLOPS (H100 SXM5 and H200: 3.958; B200: 9.0, per NVIDIA spec sheets — dense throughput is roughly half). Providers do not publish per-performance pricing, so all values are Silicon Analysts estimates; each data point cites the underlying price observation.
Is the B200 cheaper than the H100 per unit of compute?
On neoclouds, yes: Lambda rents the B200 at 1.38x its H100 price for 2.27x the FP8 (sparse) throughput — about 39% cheaper per effective PFLOP ($0.61 vs $1.01). On AWS capacity blocks the B200 is priced at per-PFLOP parity with the two-year-old H100 (~$1.04 vs ~$0.99), meaning the newest silicon carries zero performance discount — hyperscalers price scarcity, neoclouds price compute.