Who's Making Trillions Off the AI Boom? | Edited Transcript
A professionally prepared transcript tracing AI infrastructure spending across power, cooling, chips, networking, memory, and storage.
Chapter Timestamps
00:00 $725B in 2026, $7 Trillion by 2030
02:08 Why Spend $725B? Scaling Laws
04:14 What Happens When You Hit Send on an AI App
07:16 Electricity Bottleneck
11:20 Nuclear & Gas Power: Constellation, Vistra, GE Vernova
16:27 The Grid Equipment Layer: Vertiv, Eaton, Caterpillar
18:13 Air-to-Liquid Cooling Bottleneck
23:06 $3 Million per Server Rack: NVIDIA, AMD & Broadcom
30:38 10,000 GPUs Cooperate: Broadcom, Marvell, Astera
36:57 Optics: Coherent, Lumentum & Innolight
41:21 The Next Bottleneck: Co-Packaged Optics
44:17 HBM Memory: SK hynix, Samsung & Micron
46:48 Hard Drives: Seagate & Western Digital
49:21 The Full Journey
Made with: The Transcript Desk Chrome Extension
Full video:
Leo Jiang follows the money through the AI data-center supply chain, tracing how hundreds of billions in capital spending flow into electricity, grid equipment, cooling, GPUs, networking, optics, memory, and storage, and where suppliers do or do not retain pricing power.
Transcript
00:00-02:08 | $725B in 2026, $7 Trillion by 2030
Leo Jiang:
We see it everywhere. AI is changing the world. It’s coming for our jobs, and companies are spending a fortune building it. Just four companies alone are projected to spend around $725 billion this year alone, and even more next year. Almost a trillion dollars in 2027. You pull out your phone and you ask the question on every investor’s mind: who the heck is making trillions of dollars in this AI boom? That’s what I want to do in this video: deep dive into uncovering the truth of this AI boom by following the money.
Who’s getting rich from this? Who’s paying for it? And which companies have pricing power versus those that are just commodity suppliers in this boom cycle? We are in the greatest construction boom in American history. More spending in one year than the inflation-adjusted cost of the entire U.S. interstate highway system. Goldman Sachs projects that AI investments will be around $7.6 trillion from 2026 to 2031, in just five short years. Jensen Huang, the CEO of NVIDIA, stood on stage at Davos in January and called it, in his own words, the greatest infrastructure buildout in human history. We’ll trace this entire journey from the moment you tap enter on your phone’s AI app, through the wires of the internet, to the data centers the size of several football stadiums.
These buildings use as much electricity as a small city, and then we’ll break apart the machines inside, each one costing as much as a house, running so hot that air conditioning no longer works. They have to be liquid-cooled like heavy industrial equipment. So the question this entire video hangs on is simple: where does all this money go? By the end, you’ll be able to answer that. You’ll understand every hand touching this $725 billion ocean of capital: the chip designers, manufacturers, memory companies, fiber optics, power plants, cooling systems, and the software companies building on top.
You’ll know which companies sit at each step, who’s touching the money as it passes, and who’s actually keeping it, who has pricing power, and who’s a cog in the system. I’m Leo. This is for educational purposes only. Please do your own financial diligence before you decide to invest.
02:08-04:14 | Why Spend $725B? Scaling Laws
Leo Jiang:
In late 2022, the ChatGPT moment arrived, and humanity changed. But it was one discovery that predates it, years ago, that is the single most economically important discovery of this decade, and something most people have never even heard of. It’s called the scaling laws. Around 2020, researchers at OpenAI found something almost embarrassingly simple. You take a bigger model, and you give it even more data, and you feed it even more compute, and it gets smarter. Not sometimes, predictably so. And if there’s something corporations love, it’s predictability.
On the chart, it’s nearly a straight line. You spend 10 times more across the board on compute, data, and the model size, and you get reliably a smarter, more capable, and economically potent model. Consider what this means for our corporate CEOs. Since the industrial age, we’ve had to hire our fellow humans, train them, pay them, and coordinate them. And our workers get old, they want rights, and they have their own opinions. They’re unpredictable. The scaling laws have turned intelligence into something that corporations can now train and own forever.
They’ve converted AI from a sci-fi plot into a capital expenditure problem. And here’s the golden thing. American companies may not manufacture much anymore, but if there’s anything America is great at, it’s making a lot of money and spending even more money. Our companies are spending more money on AI than all other countries combined, and it’s not even close. This is why America is so far ahead, and it will continue leading. Remember the scaling laws? Bigger models need bigger computers to run on.
Bigger datasets need more interconnects to move the data. Compute—the lifeblood that makes everything possible—needs data centers, which means electricity, cooling, fiber optics, semiconductors, equipment, and land. And now tech has basically gone from a bunch of office workers working on a laptop to a heavy industrial business, and that’s completely changed who is making money. Now back to the question we sent from our AI app,
04:14-07:16 | What Happens When You Hit Send on an AI App
Leo Jiang:
and follow it through its complete journey. Your question leaves your phone and hits a cell tower, and goes through fiber-optic cables, glass strands carrying data at two-thirds the speed of light, travels hundreds of miles to a data center. The telecom companies built this infrastructure. Think of it as a digital highway. Step two, the front desk. Your question has now traveled hundreds of miles and arrived at the metaphorical front desk. We call it the API gateway. It’s basically a security checkpoint that verifies who you are.
It runs a safety check and determines whether or not you’re allowed to enter the data center. Keep this in mind because it may get more important. As more AI agents roam the internet, the cybersecurity companies are guarding these doors. Cloudflare, Akamai, Palo Alto Networks, CrowdStrike could all be in serious demand. Finally, your request is allowed in and enters the data center. It carries the system instructions, your past conversations, and your current questions. Step three, tokenization. The AI speaks its own language, and this language is grouped into tokens, which are converted into numbers.
Remember this term, tokens. It is the unit of currency that these AIs are measured in. Step four, processing your request. The models process your question in two totally different phases. The first one is called the prefill. The model reads your entire prompt all at once, in parallel, and builds an internal understanding of it. Ever noticed that pause before the first words appear? That is called the prefill. And the data center is reading the work order. Second, prefill produces something called the KV cache.
This is the model’s working memory of your conversation, which holds everything in super fast memory right next to the chip. You don’t need to understand it deeply, but remember the terms prefill and KV cache because these are the steps in the AI factory assembly line that companies are working feverishly to make more efficient. Efficiency here could shift the entire AI infrastructure. One thing to remember here: in the AI era, intelligence works like a factory line. In the industrial age, we were producing T-shirts, cars, and planes.
In the 2020s, the factory of our current age produces tokens. Now the factory line starts running, and the model generates one token at a time. For each token of response, it has to do hundreds of billions of calculations just to produce a single word of response. Then it does the same thing over and over again for every subsequent word, one token at a time. And in 2026, what we saw with the rise of AI agents is that the AI may think entire textbooks worth of text before it ever even responds to you.
This is at its core why AI has exploded in cost. When you watch the answer type itself onto your screen, you’re literally watching the assembly line run.
07:16-11:20 | Electricity Bottleneck
Leo Jiang:
Let’s zoom out and view this journey in layers. This is a map for the rest of the video. Here’s the one-sentence version of the entire story: Electricity enters one side, flows through silicon, and becomes computation and heat. Water carries the heat out, and computation goes through light. What comes out is intelligence, packaged as tokens. That is the factory. So let’s begin this tour at the energy level, the power plant. Because the story of AI in 2026 is a story of bottlenecks. Not just bottlenecks in semiconductor chips, it is a story about electricity.
Here’s the number that frames the problem and explains why the story is only going to get bigger. A traditional server rack, the kind that ran the internet for 20 years, draws about 5 to 10 kilowatts. Each kilowatt is roughly the equivalent of 10 old-fashioned 100-watt light bulbs burning all at once. NVIDIA’s flagship AI rack draws about 120 kilowatts. The next generation of Rubin racks, expected this year, is projected to approach 600 kilowatts per rack. That is a 60 to 100-fold jump in power density in under a decade.
The electricity demand of an entire neighborhood is now concentrated in the size of a phone booth. And the tech companies are betting all their money on building more of these data racks. This power density is the root cause of nearly everything we’ll see in the subsequent two layers. One server rack is basically the size of a phone booth, and this phone booth costs as much as basically all our homes. Now let’s scale up, because the entire data center today can demand a gigawatt or more. A gigawatt is roughly the output of a full-size nuclear reactor, enough electricity for about a million homes.
Individual companies are now planning multiple campuses, each requiring multiple gigawatts. It’s like millions of people suddenly started popping up in the United States, and they all run on electricity. Data centers consumed about 4% to 5% of U.S. electricity before the boom began. Projections show that that’ll increase to 15 percent within the next couple of years. The problem is that the U.S. electric grid was built decades ago for demand growth of 1 to 2 percent a year, which was manageable. But then AI showed up, and now it’s asking for tens of gigawatts almost overnight.
And the grid simply cannot provide that much that quickly. This leads us to these two key bottlenecks. Number one, the interconnection queue. To plug a new facility into the grid, you file a request and wait until the utility company assesses whether the grid can handle the additional load. This queue is currently four to five years long, which makes sense for anyone who’s dealt with utility companies before. The largest companies humanity has ever seen cannot afford to wait that long. According to the Texas Tribune, just in Texas, the massive backlog in the power grid queue has climbed to 475 gigawatts, with data centers making up about 90 percent of total volume requests. In peak demand, that is roughly five times the size of the entire Texas grid waiting in line.
Number two, transformers. Not the deep learning software architecture. The original OG hardware transformers. A large power transformer—a giant gray box that steps voltage up and down—used to take about a year to arrive. Today, delivery takes two to four years, and prices have risen by 80 percent. Even if you secure all the semiconductor chips you need, the grid equipment that powers them may not arrive until 2029. So what do you do if you’re a hyperscaler or a big tech company, and every month of delay collectively costs the industry billions of dollars in revenue?
You don’t wait for the grid, you build around it.
11:20-16:27 | Nuclear & Gas Power: Constellation, Vistra, GE Vernova
Leo Jiang:
The industry calls this behind-the-meter power: generating electricity on-site to bypass the public interconnection queue. That decision, made simultaneously by the largest companies, lit a fire under an entire forgotten corner of the stock market: boring, old-fashioned industrial power companies. Let’s introduce the players, from the Davids to the Goliaths. First, the nuclear resurrection. In 2024, Microsoft signed a deal that would have sounded like satire a decade ago: a 20-year agreement with Constellation Energy to restart the undamaged reactor at Three Mile Island, next to the reactor involved in the 1979 accident. Constellation is spending about $1.6 billion, backed by $1 billion in federal loans, to bring the unit back online in the second half of 2027. Microsoft will buy every single megawatt it produces for the next 20 years. Constellation operates America’s largest nuclear fleet, at about 22 gigawatts. Suddenly, its aging reactors became some of the most valuable energy assets on Earth. Why? Because AI factories run 24 hours a day,
seven days a week, and nuclear is one of the few carbon-free power sources that can do the same. The stock tells the story: Constellation is now roughly a $100 billion company. Its peer Vistra Energy has a 37-gigawatt fleet of nuclear and gas generation and rode the same wave. The moat is beautifully simple: you cannot build a new conventional nuclear plant in America within this decade. This makes the existing reactors nearly irreplaceable. Then there are the small modular reactor upstarts. Most are pre-revenue startups.
Oklo reports a pipeline of more than 14 gigawatts. It is backed by OpenAI CEO Sam Altman. NuScale has a small modular reactor design certified by U.S. regulators. In my view, these are all pre-revenue companies, and they will not deliver electricity until 2030 at the earliest. Both stocks are mostly narrative-driven lottery tickets. But the upside of pre-revenue companies is that they can go as high as the narratives take them. Next, the fastest power in the West. If the grid takes four years and nuclear takes ten, what can you deploy within a year? Fuel cells.
Bloom Energy makes solid-oxide fuel cells: boxes that convert natural gas into electricity chemically, without combustion. You can deploy them quickly behind the meter, right next to a data center. During a single evening in the first half of the year, Bloom announced $7.65 billion in data center contracts. But this July, Hindenburg Research, an activist short-selling firm, published a report claiming that Bloom’s marketed $20 billion backlog was more than 40 times its binding contract obligations, compared with roughly two times for its peers. The report also argued that scaling to meet Bloom’s data center ambitions would require nearly the entire global supply of scandium, a rare-earth metal that China now requires an export license to ship.
Bloom rejects these claims as false and misleading. Still, if the short seller’s comparison between the marketed backlog and Bloom’s SEC filings is accurate, investors have ample reason to be cautious. Finally, there is the elephant in the room: the company best positioned to dominate the next several years, yet perhaps the least glamorous: GE Vernova, General Electric’s power spin-off. It makes giant gas turbines that are realistically the leading near-term power source for AI, because gas is the only generation technology that can be built at scale before 2030.
GE Vernova’s turbine production slots are sold out through the end of the decade. Its backlog is around $163 billion. In the first quarter of 2026 alone, it booked $2.4 billion in data center electrification orders, more than it booked in all of 2025. The stock has risen so much that GE Vernova is now worth $300 billion.
16:27-18:13 | The Grid Equipment Layer: Vertiv, Eaton, Caterpillar
Leo Jiang:
Between the power source and the chips sits a layer of equipment most people never consider: switchgear, busways, and uninterruptible power supplies, essentially giant batteries that catch the load instantly if the grid blinks. Even a half-second outage can ruin a months-long training run, costing hundreds of millions of dollars. These companies dominate this layer, and these three will appear again in the next act: Vertiv, Eaton, and Schneider Electric. Eaton’s electrical backlog grew 48% year over year. Vertiv’s backlog doubled to more than $15 billion, and the company joined the S&P 500 in March.
These companies sell shovels to every miner, regardless of who strikes gold. The first line of defense is rows of backup generators from Caterpillar, diesel engines sitting idle as they wait for the rare moment when the grid fails. Analysts believe that Caterpillar’s data center generators could triple by 2030. Before we move on, we need to confront an uncomfortable reality. Data centers are bidding for power, and they bid against local residents. This is a major political risk. In the regional grid serving 13 states from Illinois to Virginia, data center demand added more than $9 billion to the latest capacity auction.
In parts of Ohio and Maryland, that increase translated to residential electricity bills rising by $16 to $18 a month. And communities are noticing. Officials are proposing moratoriums. Local opposition is becoming a genuine political risk to the buildout. As an investor, you should be wary of this political headwind. So now we’ve talked about the factory getting the power,
18:13-23:06 | Air-to-Liquid Cooling Bottleneck
Leo Jiang:
the power getting to the server racks, but we have a huge issue that arises. As we mentioned, more than 100 kilowatts flow through each rack of chips, and every watt becomes heat, so it’s quickly becoming dangerously hot. A single flagship AI chip today dissipates more than 1,000 watts of heat. You can imagine a chip basically the size of a postcard produces as much heat as a full-size space heater, and then we stack 72 of them into a rack. Add memory and networking, and you have 120 kilowatts of heat, the output of basically 80 space heaters inside a human-sized cabinet that you could wrap your arms around.
So cooling isn’t merely a supporting function, a nice-to-have, in an AI data center. Cooling is an essential part of the job. And for 30 years, the method was to use elaborate air conditioning. Cold air came up through the floor, and hot air was pulled out the back, and giant chillers and cooling towers sat on the roof. Air cooling worked well, up to about 30 to 50 kilowatts per rack, but we have permanently crossed that threshold, and we’re never going back. Air cannot carry heat away from a hundred-something-kilowatt rack fast enough.
You would need almost hurricane-force winds blowing through the servers. So the industry is undergoing the biggest renovation in its history: the shift from air to liquid cooling. And per unit of volume, water can carry heat away roughly 3,000 times faster than air. So here is the technology in kind of one view. The rear-door heat exchangers, the water-cooled radiators bolted to the back of a rack, are transitional solutions. Direct-to-chip cooling has become mainstream in 2026. So these metal plates with liquid channels sit directly on top of each chip, connected to a CDU, or coolant distribution unit.
Think of it as the racks pumping coolant to every chip and carrying the heat into the building’s water loop. NVIDIA’s flagship racks do require liquid cooling. It isn’t optional, by the way. This is a requirement now. At the extreme end is immersion cooling, so literally submerging entire servers in tanks of nonconductive fluid. It’s like staying in a pool on a very, very hot, burning summer day. Investors will encounter two important terms. The first is PUE, power usage effectiveness, a facility’s efficiency score calculated by dividing the total power entering the facility by the power that actually reaches the computers.
A perfect score is 1.0. Older data centers run around 2.0. Modern liquid-cooled facilities can approach 1.0. Multiply that gap by a gigawatt and the price of electricity, and it becomes an increasingly bigger chunk of money. And then there’s the water itself. Many data centers cool their systems by evaporating millions of gallons of water, which is becoming a genuine political risk in drier regions. Closed-loop liquid systems help and are influencing where facilities can be built. So now we’ve explained why this cooling system is exploding in CapEx costs.
Let’s talk about who actually gets paid. It’s largely the same names we saw in the power room, and it’s because the systems are so closely related. Vertiv is the market leader, one of the rare companies selling both power equipment and liquid cooling. It’s a one-stop shop, and it’s grown revenue 28% at 20% margin. And then something remarkable happened. Within months of each other, the two electric giants spent billions to buy their way into liquid cooling. Eaton paid $9.5 billion for Boyd Thermal.
Schneider Electric bought Motivair. Basically, when the industry’s most disciplined buyers are both paying up for the same thing, they’re telling you they believe that every future data center will look like this. And then there are smaller pure plays that focus just on cooling loops and enclosures, but they’re mostly private companies like CoolIT Systems, a specialist whose cold plates ship inside many brand-name servers. The liquid cooling market was worth about $5 billion in 2025. Forecasts put the market between $15 billion and $27 billion by early 2030.
So basically tripling or increasing fivefold within the next couple of years. And it doesn’t matter whether NVIDIA, AMD, or any of the hyperscalers win this race for AI. Heat is still heat, and this will still always be in demand.
23:06-30:38 | $3 Million per Server Rack: NVIDIA, AMD & Broadcom
Leo Jiang:
So the factory has power and the heat is under control. It’s time to walk back and look at the rack and what actually runs everything. And it is the $5 trillion company, the era-defining company of our time: NVIDIA. This is the machine, NVIDIA’s GB200 NVL72. The 72 here in this case means it’s 72 GPUs stacked up together in a rack. GB refers to Blackwell, NVIDIA’s current generation of chips. Basically, these 72 GPUs are wired so tightly together that they behave like one giant computer.
The rack weighs about a ton and a half and draws over 100 kilowatts, as we discussed, and costs roughly $3 million. Let’s look inside and let’s use the kitchen metaphor. The CPU is the head chef. The central processing unit runs the operating system, takes orders, and coordinates everything. It excels at these complex sequential tasks, but there are only a handful of head chefs. For decades, the CPU was the star of computing. It was Intel’s kingdom. But in an AI era, it’s been demoted to management.
The GPUs, graphics processing units, are basically like 10,000 line cooks. They were originally invented mostly for video game graphics and contain thousands of small, simple cores that all perform the same mathematical operation. It turns out the math inside a neural network is exactly that type of work. It’s billions of very similar matrix operations. And one head chef can’t do it all at the same time, but 10,000 line cooks can, each chopping up one operation at a time. One peculiar thing about history is just how things coincidentally play out.
Video game graphics and AI require the same kind of math. And this is the foundation of what really put NVIDIA in such a great position to capitalize and become the main profit center behind this AI boom so far. HBM, or high-bandwidth memory, is like the countertop. SSD is like the pantry. And the NIC is like the network interface card. It’s kind of like the waiter carrying dishes between kitchens. And the power supplies and motherboards are like the plumbing and wiring holding everything together.
Now consider this. NVIDIA finished its most recent fiscal year with $216 billion of revenue, up 65%, of which almost $190 billion came from data centers. It controls roughly 80% to 85% of the AI accelerator market, and its gross margin last quarter was, get this, 75%. Apple, the most admired hardware company in history, operates at around 46%. It became the largest company last October and the world’s first $5 trillion company. Depending on the week, roughly $0.07 of every dollar invested in the S&P 500 index automatically buys NVIDIA stock.
How is that margin possible? Everyone says it’s because NVIDIA obviously has the best chips, and that’s definitely part of the answer. But another part of the answer we can’t overlook at all is its platform, CUDA. For 20 years, every AI developer on Earth has been trained to use it. And for now, the one-line version of this that you have to remember is this: NVIDIA doesn’t just sell chips. It sells the only complete system that the world’s best engineers already know how to operate.
Buying a competitor’s chip can mean retraining an entire workforce. And the bear case: about 40% of NVIDIA’s revenue comes from only four customers, the hyperscalers, and all four of them are building their own chips to replace NVIDIA’s chips. Let’s talk about its closest competitor, AMD. AMD’s GPUs are genuinely competitive for inference, with more memory per chip and, by some estimates, 25% to 40% more tokens per dollar. AMD’s problem has never been the silicon. It’s been its software. Its CUDA alternative now reaches about 90% to 95% of NVIDIA’s performance on standard workloads.
But when a single training run for an AI model costs hundreds of millions of dollars, a product that is only 90% as good and also introduces friction to the workforce is a difficult sell. AMD holds perhaps 5% to 7% of the market. It’s still a distant second. And last and quickly is Intel. Painful as it is to say, Intel is barely in the race. It still sells plenty of CPUs, which have become increasingly important for agentic AI that runs sequentially. But the story of AI so far has been dominated by the GPU, the line cooks, and not the CPU, the head chefs.
And now for the silent elephant in the room: Broadcom. Remember how we said hyperscalers are building their own chips? Well, they can’t do it alone. Designing a frontier AI chip requires years of specialized expertise and intellectual property. Broadcom helped Google, Meta, and now even OpenAI and Anthropic with their custom chips. These custom chips are called ASIC, a term you’ll hear in the news often. Broadcom controls more than 60% of the custom chip market. Its AI revenue rose by 106% last quarter and has a $73 billion backlog.
And its management says it has a line of sight to $100 billion in AI revenue by 2027. In factory terms, NVIDIA sells the entire finished kitchen. Broadcom helps you create your own. It takes a cut either way. Broadcom dominates switching silicon and networking. One company sits at both sides of this. Who creates the racks? And it’s not NVIDIA. NVIDIA designs the core platform. That would be Supermicro. They integrate complete liquid-cooled racks faster than anyone. Its revenue rose by 123% last quarter, with more than 90% coming from AI.
But here’s the key thing: gross margins were only 6% to 10%, depending on the quarter. Dell has taken more than $64 billion in AI server orders and has a $43 billion backlog, an astonishing amount, while server segment margins remain below 9%. And then beneath it all, we have to talk about the equipment manufacturers, the Taiwanese ODMs, or Original Design Manufacturers. Foxconn assembles roughly 40% of the world’s AI racks, along with a couple of other companies. And the same racks pass through many hands.
These companies are mostly Taiwanese, so I’ll skip through them. If you’re interested, we can deep-dive into that in a future video. NVIDIA earns roughly 75% gross margins on its product. But you’ll notice that these companies that are physically assembling these racks together keep just 6% to 10%. Here’s the thing to keep in mind. In hardware, profits accumulate wherever there is scarcity. Chips and software are scarce right now, and assembly is not.
30:38-36:57 | 10,000 GPUs Cooperate: Broadcom, Marvell, Astera
Leo Jiang:
Let’s talk about how these AI models are actually running inside these server racks. Remember what we said about the most important discovery of this decade, scaling laws? AI models are getting more capable with more data, more compute, and larger model sizes. Well, now AI models are humongous. The frontier models have more than a trillion parameters requiring terabytes of ultra-fast memory. But the largest GPUs carry only a few hundred gigabytes, so the model is divided across thousands of chips. And getting these 10,000 line cooks to work as one brain is one of the hardest engineering problems in the building, and it’s where some of the best businesses in the ecosystem are hiding.
So going back to our kitchen metaphor, think of it as a kitchen with 10,000 line cooks preparing one dish. Every cook has only part of the recipe. Every few moments, every cook must send ingredients, measurements, and instructions to each other. If these handoffs are slow, the entire kitchen stops. And these are some of the most expensive line cooks in the world. When the network stalls, the GPUs stall. That means billions of dollars of computer equipment sit idle, not because the chips are too slow, but because the data can’t reach them fast enough.
The network becomes a bottleneck. And there are two bottlenecks that we’re going to look at. Number one, the scale-up network. This is what connects the GPUs inside. So its job is to make dozens of processors behave as if they were one enormous processor. Going back to NVIDIA’s GB200 NVL72, recall that the 72 means there’s 72 Blackwell GPUs. And through NVIDIA’s proprietary software, it can operate as one giant computing machine. And this is NVIDIA’s fortress CUDA. It keeps customers inside their ecosystem.
The second network is called scale-out. So scale-out connects one server rack to thousands of others across data centers. So if scale-up is about making one giant machine, scale-out is about connecting these machines into one giant cluster. This is where Ethernet competes with InfiniBand. So InfiniBand is a specialized networking architecture designed for extremely fast, low-latency computing. NVIDIA gained control of it through an acquisition in 2020. For years, NVIDIA was the default choice for many of the most demanding AI training clusters. They also reinforced customers’ dependency on a single vendor.
This is where Ethernet came on the scene as a challenger. Ethernet is supported by nearly everyone who wants an alternative to NVIDIA’s monopoly. Sorry, I mean vertically integrated network. Ethernet has consequently been gaining ground in new AI clusters. This is the power of open source. And now we need to talk about the traffic controllers that are sending the data around. So we start with the switch chip. Its network switch receives data from many different machines and decides where each packet should go.
Inside the switch is a specialized semiconductor, the traffic controller of the entire network. And one of the strongest businesses here belongs to Broadcom. Broadcom’s Tomahawk chips power many high-performance Ethernet switches. Broadcom does not necessarily build the finished box that appears in the data center. Instead, it sells the silicon inside that box. This is the merchant silicon model. Design a critical chip, sell it to many equipment makers, and let someone else handle the lower-margin assembly. Now recall Broadcom also designs custom chips for the hyperscalers.
So it can earn money from both sides: the chips performing computation and the chips moving data around. Its competitor, Marvell, also competes across many of the same markets. But Marvell is difficult to place in a single box because it participates in almost every part of the connection. Marvell is not merely an optical supplier or a smaller Broadcom. It is a diversified data infrastructure semiconductor company trying to capture value wherever data moves. But selecting which path to traffic data is only part of the problem.
At these speeds, even moving a signal a few inches across a circuit board becomes difficult. Electrical signals weaken and distort as they travel. At lower speeds, that loss may not matter. But when hundreds of billions are moving every second, a small imperfection can corrupt the data. This creates a market for what we call retimers. A retimer receives a degrading electric signal, reconstructs it, and sends it out as a clean version. It’s like placing a translator between two people whose voices become less intelligible each step apart.
This is where Astera Labs first became important. Astera Labs was originally known for its CXL retimers, the small, largely invisible chips that keep accelerators, processors, memory, and network components reliably connected. But describing Astera as merely a retimer company is too narrow. It sells intelligent connectivity products for high-speed cables. Its ambition is to become basically the connectivity platform inside the entire AI rack. The economics showed just how valuable its components can become. In the first quarter of 2026, Astera reported approximately $308 million in revenue, up 93% from the previous year, and with a gross margin of, get this, above 76%.
That doesn’t mean that every retimer carries a 76% margin. It means that the company’s complete product portfolio produced that margin. And the larger lesson is this: when so much money is riding on the AI boom, even the distance between two chips can become a high-margin market. Now let’s talk about the signal carrying this information.
36:57-41:21 | Optics: Coherent, Lumentum & Innolight
Leo Jiang:
Electrical signals work across short distances inside a server rack. Data can travel through copper traces and cables, but copper becomes increasingly inefficient as speed and distance rise. The signal weakens and power consumption increases, and more energy is required to preserve the data. Eventually, the information has to stop traveling as electricity. It travels as light. An optical transceiver performs this conversion. It receives an electrical signal. From a switch, it turns it into pulses of light, sends those pulses through glass fiber, and converts them back into electricity on the other end.
These transceivers sit at the edge of the network like tiny gateways between electronic and optical worlds. And a single transceiver is itself a collection of businesses. It needs a laser to create the light, a modulator to encode information onto that light, and an optical DSP to compensate for distortion. It needs photonic components to guide the signal. It needs fiber to carry it. It needs connectors to join everything together. And then finally, it needs a supplier to manufacture and assemble the finished module with extraordinary precision.
So the optical supply chain is not one homogeneous layer. The first thing to understand about the major players is that they’re not peers. They sit at different links in the chain. Let’s talk about the major players here. Start with the finished product, the pluggable module. The global transceiver market was roughly $24 billion in 2025. The volume leader isn’t American. It’s a Chinese specialist called Innolight that did around $5.3 billion of revenue last year, roughly a fifth of the entire world’s market.
And right behind it is another Chinese specialist called Eoptolink at about $3.5 billion, which in 2025 overtook Coherent for the number two spot. Together, these two companies account for more than a third of all transceivers sold, and they dominate volume because they’re laser-focused on exactly the 800 gigabit and 1.6 terabit modules that hyperscalers are now buying by the millions. Coherent is a very different type of competitor. It’s not the volume leader, but it is the broadest and most vertically integrated photonics company in the West.
It makes its own materials, grows its own laser crystals, builds everything from components to complete modules to the photonic engines for co-packaged optics. Lumentum is more concentrated, high-value lasers and optical components, including the external laser source that the future co-packaged systems will need. It books less revenue than Coherent, but keeps more of each dollar. It has a gross margin of around 44% compared to Coherent’s 38%. One layer above it sits Marvell. It doesn’t sell transceivers at all, but it sells optical DSP.
This is the semiconductor brain inside them and earns semiconductor economics, a margin above 50%, with three-quarters of revenue coming from data centers. At the opposite end of the economics is Fabrinet, the Foxconn of optics. It assembles modules for nearly everyone on the list. Last quarter, it invoiced more revenue than even Lumentum, but it kept a gross margin of about 12%. It touches more money than it keeps. Corning is the arms dealer of the arms dealer. It makes the glass fiber itself, as far upstream as this industry goes.
And Amphenol supplies the connectors and cable assemblies. It’s the knuckles of this entire system. Both are essential, but neither controls the architecture. So all these companies participate in their respective segments of the optical networking. And when we follow the money, it becomes pretty clear: the closer you are to proprietary silicon and lasers, the more money you keep. And the closer you are to assembly and standardized hardware, the more revenue that simply passes through your hand.
41:21-44:17 | The Next Bottleneck: Co-Packaged Optics
Leo Jiang:
The major thing to talk about here is where this industry is headed. And this is one of the hottest trades of this past year. And according to a lot of people, it may get even hotter. And it has to do with the next architectural transition. So what I’m talking about is co-packaged optics, or CPO. In a conventional network switch, the main switch chip sits in the middle of a circuit board. Removable optical transceivers sit at the front of the box. The data begins inside the switch chip, travels electrically across the circuit board, and reaches the transceiver, and only then becomes light.
As switch speeds increase, the electric journey becomes more expensive and power-hungry. CPO shortens this journey. Instead of placing the optical conversion at the front of the switch, the optical engines are placed directly beside the switch, or potentially even beside the chip itself. Electricity travels only a tiny distance, then the data becomes light. Broadcom is the merchant leader. It delivered its first CPO switch, Bailly, at 51.2 terabits, back in 2024, pairing it with its own Tomahawk silicon with eight silicon photonic engines. It’s now on its third generation at 100.2 terabits.
NVIDIA is the full-system player here. It controls the GPU, the networking, and the software, and it’s rolling out photonics versions of its network switches across its data centers. Marvell is the independent challenger: DSP, switch silicon, custom AI chips, and its own CPO platform. And feeding all of them, Coherent, which supplies the lasers, the photonic engines, and the advanced packaging, and Lumentum, specializing in the external laser source that sits outside the package, keeping heat away from the switch chip. Here’s the key point.
CPO is not a new component. It’s a new architecture. And what this means is that it’s basically redrawing the map of the industry. It’s moving the border between the electronics and optics industries, and it threatens the companies that are winning today. For example, Innolight and Eoptolink, those two Chinese firms we have talked about, they dominate the removable module. If the module disappears into the package, some of that value disappears with it, flowing toward Broadcom and NVIDIA, who control the architecture, and toward suppliers like Coherent and Lumentum, whose lasers the new architecture still needs.
No one reports CPO revenue separately, so there isn’t this market share table. But what exists instead is this ranking of control. Who owns the chip? Who owns the system? And who’s merely supplying?
44:17-46:48 | HBM Memory: SK hynix, Samsung & Micron
Leo Jiang:
Let’s talk about memory, the hottest topic in the AI bottleneck story this year. Here’s a secret: the GPU’s massive cores are often not the bottleneck. For each token, the chip has to pull the model’s parameters and working memory, the KV cache we mentioned before, from memory into its cores. The math is fast, but fetching the data is slow, and we call this inference memory bandwidth bound. The line cooks are lightning fast, but the countertop to hold everything can’t feed the ingredients quickly enough, and that makes the countertop memory, in this case, some of the most valuable real estate in technology today.
The industry’s answer is HBM, or high-bandwidth memory. Instead of laying memory chips flat on the board from the processor, HBM stacks them vertically, 8 to 12 stories high, and drills thousands of microscopic elevator shafts through the silicon and places the entire tower directly beside the GPU in the same package. It is a skyscraper of memory downtown instead of a suburb of memory across the highway, and the result is 5 to 6 times the bandwidth of conventional memory at 5 to 6 times the cost.
And NVIDIA happily pays for every single one of them. Memory is now one of the largest cost components inside every AI chip you have heard of, and only three companies can produce it at leading-edge scale. SK hynix owns roughly 60% of the HBM market. It got there by out-executing its giant Korean neighbor, Samsung. It bet on HBM years before it mattered and shipped each generation first, and secured the lion’s share of NVIDIA’s next-generation allocation. Samsung, the largest memory maker overall, was embarrassingly late, and Micron, the American champion, went from an afterthought to selling its entire year’s capacity more than a year in advance.
In May of 2026, all three memory makers crossed a trillion dollars in market value. Less than a decade ago, their value was 16 times less. Memory used to be the most brutal commodity business in technology: boom, bust, bankruptcy, and repeat. HBM is slowly changing this psychology. It is allocated like a scarce resource and priced like a luxury good. The open question is whether these prices will hold once all three giants complete their capacity expansions at the same time.
46:48-49:21 | Hard Drives: Seagate & Western Digital
Leo Jiang:
Now, let’s start wrapping up with the final part: the warehouse. The storage hierarchy can be explained in one line. The closer storage is to the chip, the faster and more expensive it is. Cache sits on the chip itself. HBM sits beside it. Regular DRAM sits on the motherboard. Solid-state drives and flash drives hold hot data. And at the bottom is the technology everyone declared dead 10 years ago: the spinning hard drive. And it’s still unbeatable on cost per terabyte for cold bulk data. AI has become a colossal driver of storage demand because of how large these data sizes are becoming.
And here’s the part no one really predicted: the output. Every conversation, every log, every generated image retained indefinitely. Seagate’s CEO called it the inference inflection. AI does not merely consume more data; it manufactures oceans of data. But the warehouse is really two different businesses. The first is the hard drive, a triopoly. Western Digital and Seagate each hold more than 40% of the market, Toshiba a meaningful third at roughly 17%. Western Digital, now a pure hard drive company after spinning off its flash business, earns 89% of its revenue from its cloud segment and was one of the best performers in the entire S&P 500.
Much of the industry’s high-capacity production is already committed through 2027. And unfortunately, consumer hard drive prices have jumped by half because AI has absorbed the supply. The second business is flash. These are the solid-state drives. And there are six players in the knife fight here. Samsung leads with 29%. SK hynix holds 18%. Kioxia, Micron, SanDisk, and China’s YMTC each around 13%. Flash is fragmented, capital-hungry, and historically the most cyclical market in the segment. When I follow the money, I see that the memory suppliers and the storage suppliers collect very comparable revenue from this AI boom.
But the HBM makers run at roughly 68% operating margins. Storage runs at 33%. The same money is flowing through both, but one side keeps twice as much. In this AI boom, where you sit in this hierarchy matters as much as how much you sell.
49:21-51:27 | The Full Journey
Leo Jiang:
And with that, the tour is over. We’ve seen every layer: power, electricity, cooling, chips, optics, memory, and storage. Let’s watch this journey one final time. Our thumbs hit send on our AI app. The question becomes light in glass fiber, crossing state lines and arriving at a building that draws the power of an entire city. Power from a restarted nuclear power plant, a sold-out gas turbine, or a fuel cell parked behind a meter. Your question is divided into tokens, fed into a $3 million rack assembled in Taiwan, built around chips sold at gross margins that make Apple look like a discount store.
Inside, the metaphorical 10,000 line cooks fetch a trillion parameters from a skyscraper of memory. While liquid coolant carries away the heat of 80 space heaters packed into the size of a cabinet, while interconnects are sending data at light speeds, allowing these chips to think one thought. Software batches your request with a thousand others, and the answer streams back to you token by token. This is already looking like the largest infrastructure buildout humanity has ever attempted. And by the looks of it, we’re not even at the peak yet.
Whether this becomes the most important investment in history, honestly, nobody on Earth knows. But we don’t get to choose the cards we get dealt. We can only choose to follow the money and decide how we play the cards in front of us. If you found this video helpful and there’s any part of the AI infrastructure you’d like to know more about, please leave a comment. This is my first video, so I apologize for the plethora of sloppy editing. This actually took an ungodly amount of time with research and scripting, and I’ve tried to reduce as much as I can.
Not to mention a satanic amount of time editing. Anyways, if you like this video and want to see anything more, please subscribe, like, and I’ll see you all on the other side.
Made with: The Transcript Desk Chrome Extension

