
Part 3 of "Counting Chips": how boundary-setting authority is built, what it costs to keep, and what would unbundle it
The anomaly sits in a filing. For the quarter ended April 26, 2026, Nvidia reported $75.2 billion in data-center revenue. Compute, the chips, accounted for $60.4 billion of it. But one line item did $14.8 billion on its own, growing 199% year over year. Faster than the GPUs: networking. The switches, the interconnect, the fabric.¹ Nearly a fifth of the data-center revenue of the world's most famous chip company now comes from the equipment between the chips.
That number does less than it might seem to. It does not make the title of this essay literally true; sixty billion dollars of quarterly compute revenue is a chip business by any classification. The networking line is a clue: the visible edge of a change in where Nvidia's differentiation, control, and value capture actually live. That is the sense in which the title is true, and the essay's job is to earn it:
Nvidia's durable advantage is no longer any individual chip but the ability to define the unit of purchase, govern the interfaces inside it, and move that boundary before competitors can stabilize an alternative.
One distinction underwrites everything that follows, so it belongs here at the top. The unit of purchase is the commercial bundle a customer buys. The boundary of authority is the allocation of rights: who specifies a component, who sources it, who qualifies it, who prices it, who can replace it (the instrument Part 2 built). The two are related and not identical. A company can enlarge the bundle it sells without gaining all the rights inside it, and it can open a physical specification while keeping qualification authority. Nvidia's story is the interplay of the two, and every apparent paradox ahead dissolves at exactly this joint.
Nvidia's commercial history is usually told as a graduation: chips, then systems. The chronology is wrong. Nvidia was selling the fully integrated DGX-1 appliance in 2016 and published the HGX board architecture for hyperscalers in 2017. While selling GPUs as components throughout.² The ladder is additive: each rung a new, larger bundle offered on top of the old ones, none retired. The chip, an ingredient in someone else's server. The board, eight GPUs pre-integrated. The rack, arriving in 2024: seventy-two GPUs behaving as a single processor across an NVLink fabric, delivered liquid-cooled and pre-validated. In 2026, the footprint-compatible generation, Vera Rubin, built to occupy the positions its predecessor prepared. And from 2027, the factory reference design: rack, 800-volt power architecture, cooling plant, and a digital twin of the building.
Two clarifications keep the ladder honest. The rungs mix roles that deserve separate names: what Nvidia sells (chips, boards, racks and their contents), what it specifies (reference architectures for power, cooling, facility design), and what it qualifies (the components, partners, and configurations admitted to the platform). The factory rung is mostly the second and third: Nvidia authors the drawings and keeps the admission gate; partners build; customers procure and operate. And there is no single unit of purchase across the market: a hyperscaler buying racks at OEM scale, an enterprise buying a DGX appliance, a sovereign program buying a turnkey deployment, and an OEM building around HGX buy different bundles with different rights attached. The ladder describes the range on offer and its direction of travel. Which is upward, toward bundles where more of the rights sit on Nvidia's side of the line. The ladder is evidence of boundary movement. It is not itself the boundary.
Each new rung changes competition in a specific way: it displaces comparison and raises its cost. A rival accelerator can no longer be evaluated as a chip against a chip;
it must be evaluated as a full deployment: cost per trained model, cost per million tokens at a latency target, time to production, facility adaptation, software migration, and supplier risk.
The comparison is still available. What changed is who can afford to perform it: it now requires owning workloads, software stacks, facilities, and procurement organizations at scale.
Which reveals the real shape of the contest. The buyers most capable of performing the new comparison (Google, Amazon, Microsoft, Meta) are not component designers contemplating an unassembled chip. They own model workloads, networks, facilities, and internal silicon teams. They are integrators. The contest is not Nvidia's integration versus modular custom chips; it is Nvidia's merchant integrated platform versus hyperscaler-controlled integrated platforms, with consortium-governed interfaces forming between them. And every rung Nvidia climbs concentrates more of these buyers' capital with one vendor. Concentration a trillion-dollar customer does not accept but organizes against. The ladder manufactures its own opposition.
The ladder also has a ceiling, and it is visible from here. The nameable rungs above "factory" (energy, selling intelligence directly, sovereign full-stack deployments) are all someone's core business, and increasingly the customers' core business. The climb so far absorbed functions from suppliers and integrators; the rungs above absorb functions from buyers. The ladder ends where climbing means competing with the people who fund the climb, and Nvidia's carefully limited cloud ambitions suggest it knows the wire is there.
The climb has two causes, and it matters for honesty that it would partly have happened without the competitive one.
The technical cause is the subject of this series: AI performance stopped being a property of the accelerator and became a property of the system: memory bandwidth, scale-up communication, power delivery, cooling, deployment speed. When performance is lost across interfaces, someone must co-design across them; optimization migrates above the chip because the physics leaves it no choice. Nvidia would have moved up the ladder to a substantial degree in a world with no custom-silicon threat, because the product it sells, usable intelligence per dollar per watt, is now manufactured at the system level.
The strategic cause is what Nvidia did with that necessity: it converted a genuine engineering requirement into a commercial boundary. Once co-design had to happen somewhere, Nvidia arranged for it to happen inside interfaces it governs: expanding content per system, shifting the basis of competition from benchmarkable chips to deployable platforms, and building its defense against accelerator substitution into the architecture itself. A critic who hears only the second cause will say the essay mistakes engineering for scheming. The accurate claim is that Nvidia took a requirement everyone faced and made it a franchise.
A two-decade-old framework predicts when integration beats assembly.³ Integrated architectures win when the product is not yet good enough: when performance is lost at the interfaces, so whoever controls the interfaces controls the performance. Modular assembly wins when performance overshoots what a buyer needs and competition flips to cost. The framework only works honestly when run by workload rather than on "AI hardware" as one market: frontier training and long-context reasoning remain far from good enough, bottlenecked at exactly the interfaces Part 1 mapped, while narrow, stable, high-volume serving is already good enough on specialized hardware, which is why that is where custom silicon landed first. The not-good-enough regime is real, and it has a border running through the middle of the inference market.
The framework says integration is advantaged where interfaces bind. It does not say Nvidia must be the integrator: Google integrates, Amazon integrates, a consortium can standardize enough interfaces to support several integrated implementations. So a second step is required: what lets Nvidia capture the integration advantage as a merchant, selling to everyone, rather than as a captive integrator serving one owner's workloads? The answer is a compound: the largest installed software ecosystem, the fastest demonstrated path from silicon to deployed capacity, supply-chain scale across foundry and memory and packaging, a scale-up fabric competitors cannot yet buy equivalents of, and the qualification machinery that turns a parts list into a warrantied system. None is individually unassailable (captive integrators match or beat several of them for their own workloads), which is why the merchant/captive line, not the integrated/modular line, is where the industry's real border war runs.
The substitution evidence, calibrated: custom-accelerator programs are growing at roughly 45 percent a year by analyst estimate, and Nvidia's share of accelerator value has drifted from an estimated ~92 percent in 2023 to the low-to-mid 80s.⁴ Inference, the natural beachhead since stable workloads are the ones worth specializing for, is commonly estimated at two-thirds of AI compute, a directional figure.⁵ One analyst comparison credits Google's latest TPU with total cost of ownership roughly 44 percent below a comparable Nvidia system; that is an analyst estimate, not a Google claim, and the distinction matters to its weight.⁶ A frontier lab has committed to roughly a million of those TPUs, which proves the platform is commercially credible, not that the 44 percent figure is right; buyers diversify for supply security, leverage, and resilience as much as for cost.⁷ The calibrated statement: custom silicon has won credible, growing beachheads in the workloads stable enough to specialize for; Nvidia will not retain near-total share; and a falling share of an explosively growing market remains compatible with enormous absolute growth. The chip-level comparison is no longer terrain Nvidia can hold alone, so it expanded the battlefield beyond the chip.
Nvidia's temporal strategy is an annual architecture cadence, and the popular comparison drawn from it needs repair before use. The usual version (custom chips take two to three years, Nvidia ships yearly, therefore every custom chip arrives two generations stale) compares a development duration to a release interval. Both sides pipeline; an established TPU or Trainium program overlaps generations just as Nvidia does. The defensible version concerns realized cadence: Nvidia has demonstrated a roughly annual rhythm from architecture launch to volume deployment across the whole platform (silicon, fabric, software, qualification), and rival merchant platforms have not yet matched that deployed tempo. Where the rhythm holds, the moving-target effect is real: a chip aimed at last year's Nvidia platform meets this year's. Where a rival's pipeline matures, or a workload is stable enough that last year's target is still the right target, the effect fades. An advantage with conditions, not a law.
The cadence also collides with the buyer's clocks: facilities and depreciation schedules measured in years against products refreshed in one. An earlier draft of this essay overclaimed the resolution, so the correction is worth stating. Annual cadence does not strictly require an installed base that absorbs annual change, because a large share of AI capacity growth is greenfield: Nvidia can sell each generation into the next wave while prior generations serve out their useful lives. What upgradeability does is lower the cadence's cost everywhere it touches: preserving facility capital across generations, shortening qualification, shrinking the fear of instant obsolescence that makes buyers hesitate. It is the cadence's accelerant, not its precondition.
Precision about what that upgradeability is, because the vocabulary is doing commercial work. What Nvidia's materials and independent teardowns establish: Vera Rubin preserves the rack footprint, facility interfaces, and cooling architecture of its Blackwell predecessors (footprint-compatible and facility-interface-compatible); internally, cabled connections gave way to a printed-circuit midplane and soldered memory to socketed modules (serviceable and modular); and Nvidia claims a headline installation-time reduction from two days to two hours.⁸ What is not established is field-upgradeable: converting an installed Blackwell rack into a Rubin rack by swapping trays. Same prepared position, yes; facility investment preserved, largely; in-place metamorphosis, not shown. The footprint continuity is the strategically important fact regardless: a customer's building, power, and cooling amortize across Nvidia generations even when racks are replaced whole.
One layer up, attribution needs the same care. The sidecar power racks, prefabricated power rooms, cooling skids, and digital-twin tools that make facilities generation-tolerant are engineered and capitalized mostly by electrical, cooling, and construction partners; Nvidia specifies, coordinates, and certifies more than it builds.⁹ Modular facilities also cut both ways: a standardized sidecar or cooling loop lowers deployment friction for rival racks too. Which exposes the actual pattern: Nvidia standardizes the layers where ubiquity lowers adoption friction (rack mechanics contributed to open standards, facility interfaces, footprints) while keeping control of the layers where differentiation and qualification carry the economics: the scale-up fabric, the platform architecture, the admission gate.¹⁰
This is the essay's center: stabilize the boundary; accelerate what changes inside and above it. Nvidia does not benefit from instability everywhere; a permanently chaotic ecosystem would destroy the software compatibility, qualification investment, and buyer confidence its platform runs on. The strategy has four layers, and the causal modesty of the third is part of the claim. Nvidia stabilizes the platform interfaces it governs. It accelerates its own silicon and system roadmap inside them. Model and workload evolution remains substantially external, set by labs and customers, not by Nvidia. And Nvidia's platform gains value precisely when that external evolution stays rapid, because churn is what keeps buyers needing a programmable, fast-moving platform rather than a frozen appliance. Nvidia profits from architectural churn and amplifies it; it does not control it.
NVLink Fusion is the strategy's furthest extension and its clearest test. The facts: Nvidia licenses third-party CPUs and selected custom accelerators to operate inside its rack architecture and NVLink domain, sharing footprint, power, cooling, networking, and management. With marketing promising that adopters can "decouple data center design and buildout from silicon readiness" and reprovision capacity across silicon mixes.¹¹ AWS has confirmed Trainium4 is designed to support Fusion and the common rack architecture: accepting an Nvidia-governed interface as one deployment route for its flagship silicon, which is interoperability, not exclusivity, and is still no small dependency to accept.¹² Nvidia invested $2 billion in Marvell, the co-design house behind many rival accelerators, to build Fusion-compatible silicon.¹²
Neither clean reading, triumph or retreat, survives the layer-by-layer account. Which is the honest one. Accelerator choice: more flexible; that is the concession, and it is real: GPU exclusivity and some accelerator economics are gone. Scale-up fabric: more Nvidia-dependent; the custom chip now speaks Nvidia's interconnect. Rack architecture: more standardized around Nvidia's design, which is ubiquity rather than loss. Qualification: remains platform-governed. Software: mixed, by workload. Procurement: more choices, inside a governed architecture. Switching costs under Fusion do not rise or fall; they relocate: down at the accelerator layer, up at the fabric and architecture layers.
Whether the trade nets out for Nvidia is an empirical question about monetization, and the components should be kept distinct. The fabric and switch content a Fusion deployment requires; the networking, CPUs, and DPUs it makes likely; the licensing and certification economics it enables. Those attach rates are exactly what the essay's tests must observe; until they are visible, rent collection is a plausible interpretation, not a demonstrated fact. And the redistribution is not two-party. A custom accelerator inside a Fusion rack can strengthen Marvell's or Broadcom's design franchise, the cloud provider's negotiating leverage, and Nvidia's fabric position simultaneously. Authority migrating among five or six parties at once. An essay about where authority goes should resist collapsing that into Nvidia-versus-the-buyer, because the multi-party structure is where Part 4 begins.
Fusion also forces the general question the whole strategy raises: when is opening a concession? Four types are worth separating. Adoption-expanding openness (contributing rack mechanics to open standards so the footprint becomes the industry default) spreads the platform; it looks generous and functions as distribution. Defensive interoperability (admitting third-party silicon so the whole system is not replaced) trades a layer to keep the architecture. Supply-chain underwriting (the Marvell investment, capacity prepayments) is capital spent keeping the ecosystem aligned. Only the fourth, actual rights leakage (a customer taking over specification, qualification, or replacement rights Nvidia previously held), is unambiguously a loss of authority. Most of what looks like Nvidia giving things away belongs to the first three. The signal worth watching, quarter by quarter, is the fourth.
Three parallel costs constrain how high the boundary can move.
Buyer power. Custom silicon programs, procurement probes, and standards coalitions are the organized form of the resistance the ladder manufactures. The direct-memory probe is the sharpest current example (the same teardown that prices a Vera Rubin rack at $7.8 million estimates $6.7 million with buyer-procured memory¹³) And the rights framework is built for exactly this case, so use its gradations rather than a binary. A hyperscaler winning the procurement right while Nvidia retains specification, qualification, firmware, and architecture is procurement unbundling: real, partial, one dimension deep. It is not yet qualification unbundling, and it is far from architectural unbundling. Bundled and unbundled are not all-or-nothing states; the question each quarter is which rights are moving. Consortium standards (UALink, Ultra Ethernet) are the deeper version of the same force: not interfaces nobody administers..... but interfaces administered collectively, converting Nvidia's unilateral governance into shared governance wherever they take hold. The IBM precedent belongs here as conditioned evidence: integrated systems become vulnerable when interfaces standardize, buyers gain bargaining power, and specialized suppliers can meet the newly critical performance dimensions.¹⁴ All three conditions are under active construction by Nvidia's customers. None is complete.
Physical and temporal cost. The cadence and the boundary carry a maintenance bill in physics. Rack power has climbed from roughly 40 kilowatts in the Hopper generation to 120–140 for Blackwell, with Vera Rubin reported in the 190–230 range and the 2027 Kyber generation specified near 600, forcing the 800-volt transition.¹⁵ Footprint continuity holds within a rack family and breaks between families; every few years the upgrade machinery presents a facility-sized bill regardless. The clock is expensive to keep winding — for Nvidia's customers, and through their hesitation, for Nvidia.
Ecosystem cost. Authority over the boundary is financed continuously: prepaid packaging and memory capacity, two billion dollars to a design house that arms competitors, NVLink opened, rack designs donated. Under the four-type taxonomy, most of these are platform strategy rather than rights leakage, but they are not free, and their sum is the honest price tag. A franchise maintained by continuous, selective openness is durable in a way a walled fortress is not. It is also never finished paying.
Tests tied to causal claims, with the terms pinned down enough that two readers could reach the same verdict. Several are diagnostics rather than strict falsifiers (alternative explanations could rescue the thesis), which is why the section is not called "falsifiers."
The boundary-authority claim. Watch rights, not revenue. Against the thesis: hyperscalers winning specification or qualification rights — publishing rack architectures that Nvidia silicon must conform to, qualifying components Nvidia has not blessed, or deploying the same custom silicon at production scale on non-Nvidia fabrics (observable in teardowns and cloud capacity disclosures). For it: Fusion adopters staying inside the architecture across at least two silicon generations, with Nvidia content per third-party rack (fabric, switches, CPUs, DPUs) flat or rising in successive teardown estimates.
The selective-stability claim. Against: Nvidia breaking its own footprint, fabric protocol, or qualification regime in consecutive generations (instability at the boundary), or freezing silicon cadence while defending interfaces (stability inside). For: three consecutive generations of steady interfaces and annual silicon, the pattern's signature corner.
The workload-stability claim. The direct test of the theory, and the hardest to observe, so the admissible evidence should be named in advance: cloud capacity commitments, accelerator-hour and installed-base estimates, capex allocation disclosures, and software-porting timelines. Against, from the left: mature, stable, high-volume serving workloads remaining predominantly on Nvidia systems (call it three-quarters or more of estimated serving capacity) for another two to three years despite credible custom alternatives, which would mean Nvidia's hold rests on something other than the workload border (software gravity, supply, inertia). Against, from the right: frontier training migrating to captive platforms with time-to-production within a quarter or two of Nvidia's, which would mean the not-good-enough advantage is gone.
The cadence claim. Not a single slip; one late generation is execution noise. Against: the gap between Nvidia's volume-deployment dates and rival platforms' compressing toward zero across two full cycles, or buyers demonstrably deferring purchases to wait out generations, visible in utilization and order patterns.
The conclusion the opening clue was pointing at: the equipment between the chips carries a fifth of the revenue because the between is what Nvidia actually sells now: the governed boundary itself: the fabric, the footprint, the qualification gate, and the standing ability to decide where that boundary moves next. The chips remain the largest line on the invoice; the boundary is the reason the invoice comes to Nvidia. That position is held by tempo, financed by concession, and rationally contested by the hyperscalers, the customers rich and technically capable enough to challenge the boundary directly. Who are not buying components in protest but building integrated platforms of their own. Several boundary-authors, colliding, all standing on the same foundries, the same three memory vendors, and the same transformer queue: that is Part 4.
Still here congrats maybe we learned something on this one maybe not. You tell me.
(Pre-publication: each note below is to be converted to a full citation — exact title, publisher, date, link — per the checklist at the end of this section.)
Nvidia Form 8-K and CFO commentary, quarter ended April 26, 2026 (filed May 2026): data-center revenue $75.2B; compute $60.4B; networking $14.8B, +199% y/y. Company-reported. Nvidia has announced reporting-segment changes that may reduce future comparability of the compute/networking split.
Nvidia DGX-1 launch (April 2016) and HGX reference architecture (2017): company launch materials. Basis for presenting the ladder as additive boundary expansion, not product succession.
Clayton M. Christensen, The Innovator's Solution (Harvard Business School Press, 2003), on integration/modularity and the conservation of attractive profits. Framework applied by the author; the workload segmentation and merchant/captive distinction are the author's extensions.
Custom-accelerator growth (~45%/yr) and Nvidia accelerator-value share (low-to-mid 80s percent, from ~92% in 2023): Bloomberg Intelligence estimates, 2025–26 — specific report citation to be attached. Measured in estimated shipment value with internal hyperscaler deployments imputed at estimated prices; no observable market price exists for captive silicon. Directional.
Inference share (~two-thirds of AI compute): analyst estimates with inconsistent denominators (FLOPs, accelerator-hours, spend) across sources. Directional only; "inference" spans stable serving and still-evolving reasoning/agentic workloads.
Ironwood TCO comparison (~44% below comparable Nvidia system): analyst estimate; not present in Google's public Ironwood materials. Originating analyst to be identified and cited before publication.
Anthropic TPU commitment (~1M chips): company announcements, 2025–26. Cited as evidence of commercial credibility and multi-platform strategy, explicitly not as validation of any TCO figure; Anthropic's statements emphasize workload matching and resilience across TPU, Trainium, and Nvidia infrastructure.
Vera Rubin footprint and facility-interface continuity, cableless midplane, socketed SOCAMM memory: Nvidia GTC 2026 materials; SemiAnalysis, "Vera Rubin — Extreme Co-Design" (February 2026). The two-days-to-two-hours figure is a keynote claim whose measured operation (assembly, commissioning, or service swap) is unspecified — verify exact wording against the keynote transcript. Field conversion of installed Blackwell racks to Rubin via tray swap is not established and not claimed.
Sidecar 800-VDC deployment, prefabricated power rooms, cooling skids: Schneider Electric, Delta, Vertiv announcements and industry reporting, 2026 — partner-engineered, Nvidia-specified/coordinated. Siemens quotation from DSX partner materials; vendor statement, verify against original.
Nvidia contribution of MGX/NVL72 electro-mechanical specifications to the Open Compute Project (2024): company and OCP records. OCP frames the objective as common standards and a multi-vendor supply chain.
Nvidia NVLink Fusion product materials, 2026. Vendor marketing; quoted as such; verify quotations against the current page before publication.
AWS statements confirming Trainium4 support for NVLink Fusion and common MGX racks, 2026. Nvidia $2B investment in Marvell: company press releases, March 2026.
Morgan Stanley Vera Rubin rack bill-of-materials analysis, May 2026 ($7.8M rack; $6.7M with direct memory procurement). Analyst estimate, not disclosure; exact report title to be attached.
IBM System/360 unbundling and the plug-compatible era, 1969–1980s. Historical record; conditions-based comparison (standardizable interfaces, buyer power, capable specialized suppliers), explicitly not deterministic; the antitrust dimension has no automatic modern analogue.
Rack power figures: Hopper ~40 kW, Blackwell ~120–142 kW (vendor and industry documentation); Vera Rubin 190–230 kW (analyst-derived reporting; earlier claims near 130 kW conflict — disclosed); Kyber ~600 kW and 800-VDC rationale (Nvidia technical publications; roadmap). Before publication, confirm each figure's measurement basis — compute rack only versus rack plus power sidecar, nameplate versus typical load — so the sequence compares like with like.