AI Accelerators

We Gave Our Data to Two AI Analysts. Both Misread the Same Number.

By Silicon Analysts
7 min read
Market DynamicsMemory & HBM

Why this postmortem exists

When AI agents read your data, prose caveats do not survive the journey — only fields do. Both auditors reconstructed our GB200 NVL72 figure perfectly and classified it wrongly, because the sentence that said "cost, not price" never reached the file they read. The deeper failure was ours alone: our change feed published our own model restatement as a +97.7% market move, and one agent wrote it up as a supply-chain shock that never happened. Data platforms now have two audiences, and the machine audience only reads what is structured.

1Both agents misread $725.8K (a GB200 NVL72 manufacturing BOM) as a rack price and called it "off by 4-5x" against the ~$3.1M street price — while accepting the DGX H100 row one line above, which sits at a lower BOM-to-list ratio. The value was right; the export carried nothing to classify it.
2Perfect labels are not enough: our most speculative row (A16 wafers at $45,000) carried confidence, projection and rumor flags on four separate exported fields — and was still flagged as unmarked speculation. Consumers skip columns; basis must be impossible to miss.
3The error neither agent caught was worse: six of our "what moved this week" cards were curator restatements from one August evening, not market movement. One agent narrated them as "severe supply-chain shocks." We manufactured a market event out of our own bookkeeping.
4The fixes are structural, not editorial - a machine-comparable `basis` field on every exported row, a change engine in which a restatement cannot occupy the market-move array, and every value that steps because our model changed now ships with the correction that explains it.

Last week we ran an experiment we recommend to anyone who publishes numbers: we exported our own dataset to a spreadsheet, handed it to two frontier-model deep-research agents from different vendors, and asked each to audit it against industry consensus.

Both produced long, confident, largely accurate reports. Both made the same headline mistake. And both walked past the one error that actually mattered — because it looked exactly like the market signal our page said it was.

This is the postmortem. We come out of it looking worse than either agent, which is why it is worth publishing.

The number both agents misread

Our AI Server System BOM dataset carries a stacked cost breakdown for the GB200 NVL72 rack: GPU COGS, networking, power and cooling, chassis. The segments summed to $725,800.

Both auditors read that as the price of a rack, benchmarked it against the roughly $3.1 million a hyperscaler pays, and declared it "grossly wrong — off by a factor of 4-5x." One called it a top-tier credibility risk. The other made it a headline finding.

It is a manufacturing bill of materials, not a price. The dataset's methodology note said so in plain English: "All costs are approximate manufacture/procurement costs, not list prices." The largest segment was literally labelled "GPU (COGS)."

Here is the detail that makes this a systems failure rather than an agent failure. One row above the NVL72 sits the DGX H100 — a ~$45,000 BOM against a system that lists for $200,000-$500,000. Both agents accepted that row as "defensible." Our NVL72 BOM is roughly 24% of its street price; the DGX row they waved through sits at 9-23%. They accepted the more aggressive ratio and rejected the more conservative one, because on the row they rejected they had anchored on a famous public number, and nothing in the file pushed back.

Two independent frontier models, same file, same misreading. When the failure rate is 2 for 2, the defect is in the file.

The uncomfortable epilogue: while proving the value was conceptually right, we found it was numerically wrong — in the opposite direction. The GPU segment was still anchored to a pre-August B200 cost that our own HBM re-base had already moved. The corrected roll-up is about $751,000. The auditors said our number was 4x too low as a price; as the BOM it actually is, it was 3.5% too low.

Why the caveat never arrived

The methodology sentence existed. It reached our JSON API. It rendered on the dataset page, below the chart.

It was not a column in the CSV.

The export carried seventeen fields per row — values, units, confidence, projection flags, per-point source notes — everything needed to reconstruct $725,800 and nothing to classify it. The agents did exactly what a careful analyst does with a flat file: summed the stack correctly, then reached for the most famous external anchor to validate it. The anchor was a sell price.

There is a counter-example in the same audit that completes the lesson. Our single most speculative row — A16 wafer pricing at $45,000 — carries confidence: Low, data_type: Projection, is_projection: true, and a source note reading "Rumor-level only," all four in the export. One auditor flagged it as unlabelled speculation anyway. Labels that live in columns a reader can skip are advisory. The classification a number cannot be used without has to be impossible to separate from the number.

The error neither agent caught

While both audits were relitigating a correct number, our "What moved this week" section was showing six large movements: MI325X HBM cost +97.7%, H200 HBM +60%, MI325X manufacturing cost +56.6%, and three more.

Four of the six were not market movements. On the evening of August 16 we had re-based our HBM3e cost model to the August 2026 contract level and normalized a per-stack pricing basis — curator work, done in one sitting, documented in our own correction records as "a methodology correction, not new market data." The next morning's ledger freeze diffed the new values against the old ones and published the steps as the week's market movement.

One of the AI auditors then did what any analyst would do with a movement feed: narrated it. "Severe supply-chain shocks affecting the newer AMD SKUs... the MI325X HBM cost nearly doubled, surging 97.7%."

No such shock occurred. We changed a spreadsheet, our pipeline called it news, and a frontier model wrote the story. That chain — restatement to movement feed to machine-written market narrative — is the single most dangerous failure mode for any data platform whose readers are increasingly agents, because agents do not skim. They cite.

The one card that survived scrutiny, a RunPod H200 rental-price move, was also the only one with no matching entry in our own change history. The feed had buried its one genuine signal under five artifacts of our bookkeeping.

What we changed

Everything below is live, tested, and verifiable against the endpoints this article cites.

Every exported row now carries its basis. Alongside unit and confidence, every CSV and JSON row now ships a machine-comparable basis field — bom_cogs, list_price, contract, spot, revenue_implied, booking_to_delivery, and so on — plus a basis_note stating exclusions in one sentence, and the dataset-level methodology note denormalized onto every row. The NVL72 rows now say, on the row itself: cost stack; excludes vendor gross margin; a comparable system sells for roughly 4x this figure; never benchmark against a rack sell price. Where basis is genuinely unresolved — our HBM market-share series has never settled whether it is revenue share or bit share — the field says unresolved and explains why, because a guessed label is the same defect wearing a fix's clothes.

A restatement can no longer occupy the market-move feed. The change engine now classifies every delta before publishing. A value whose domain is a model — one that moves only when a human edits an input — cannot publish as a market move at all. A delta covered by a published correction releases into a separate, labelled restatements ledger carrying the correction's own explanation. Anything large and unexplained is held for review rather than published. The movement array that feeds our page, API, newsletter and structured data now contains market movements by construction, not by curation discipline.

The steps are visible, not hidden. The August 16 re-base still shows on our changes page — as a restatement, in neutral styling, with the reason attached. Hiding it would repeat the original sin in the other direction: the charts visibly step, and a reader who cannot find out why stops trusting the chart.

What this means beyond our site

Every data publisher now has two audiences. The human one reads your methodology page occasionally. The machine one reads your export, your structured data, and your feed — completely, literally, and without the context in your footer.

Three rules fall out of this experiment:

  1. A classification that lives in prose does not exist for the machine audience. If a number cannot be safely used without knowing what kind of number it is, that knowledge must be a field on the number.

  2. Your change feed is a claim generator. Anything it emits will be narrated by something downstream. If your own revisions can enter it, you are publishing fiction with timestamps.

  3. Audit your data the way it will actually be read. Not on your website, with your tooltips and your methodology section — but from the flat file, by an intelligent reader with no context and a fast anchor. Two of them, ideally, from different vendors. Where they both fail, your export failed.

We got a correct number called wrong by two of the best reading systems ever built, and a wrong feed called interesting by both. The first cost us nothing but pride. The second, uncorrected, would have cost us the only thing a data company sells.

Sources & Methodology

Data Verified PublicAll data sourced from public filings, press releases, and published reports

Methodology

This analysis is based exclusively on publicly available information including quarterly earnings calls, investor presentations, SEC/regulatory filings, published analyst reports, industry conference proceedings, trade publications, and government disclosures. All cost models use cross-validated benchmarks derived from these public sources. No proprietary, classified, or confidential information is used.

The views expressed on this site are my own and do not represent those of my employer. This is a personal research project for educational purposes. All data is sourced exclusively from public filings, press releases, and published industry reports. No proprietary or confidential information is used.

Related Analysis

Silicon Analysts Weekly

The week in AI-chip economics, in your inbox

Pricing signals, HBM & foundry moves, and that week's analysis — sourced and human-reviewed.

One short email a week — only when the data moves. No marketing, ever. One-click unsubscribe.

Explore Our Tools