Nebius Co-Founder on AI Infrastructure Bubbles | Edited Transcript
Roman Chernin on why cheaper intelligence expands compute demand, how Nebius moves from megawatts to managed inference, and why consolidation is its largest strategic risk.
Chapter Timestamps
00:00 Intro
01:24 Why AI Infrastructure Is Not a Bubble
04:11 The Real Impact of Open Source on OpenAI & Anthropic
11:03 Jevons Paradox: Why Cheaper AI Creates More Demand
13:06 The Four Layers of AI Infrastructure Explained
18:49 If Nebius Had 10x More Capacity Tomorrow
28:51 The Shift from Training to Inference and Agents
37:18 How Token Factory Cuts AI Costs by 70%
50:34 Sovereign AI, Europe, and the Future of Model Building
53:52 Competing Against Hyperscalers with 10x More Capital
01:08:46 The Biggest Threat to Nebius Isn’t Competition
Made with: The Transcript Desk Chrome Extension
Full video:
Roman Chernin is Co-Founder and Chief Business Officer of Nebius, one of the fastest-growing AI infrastructure companies in the world. Today, Nebius operates some of the largest AI compute clusters globally and serves leading AI labs, enterprises, and developers. Today, Nebius has a market cap of $57BN.
Transcript
00:02-01:47
Roman Chernin: We’re in a capital-intensive game, competing with the best-capitalized companies in the world. Our program this year is $20–25 billion. Our hyperscaler competitors are eight times bigger.
Harry Stebbings: The AI infrastructure race is on. CapEx spending has never been greater. At the center of it is Nebius.
Harry Stebbings: Today, I’m joined by the co-founder of Nebius, a company that has scaled to a $66 billion market cap, going head-to-head with some of the largest hyperscalers in the world.
Roman Chernin: Over the next six months, capital can’t help you. Six months is too short a timeframe. You have what you have, and you need to deliver. The main threat to Nebius as a business is that the world becomes too consolidated.
Harry Stebbings: Today, we uncover the AI infrastructure bubble and so much more.
Roman Chernin: It’s like a shark: you’re alive when you move. So we have to keep moving.
Harry Stebbings: And I’m thrilled to welcome Roman Chernin, who has power against Nvidia. Ready to go.
Harry Stebbings: Roman, I’m so excited for this, dude. I think Nebius is one of the most incredible stories we’ve seen over the past few years. And, holy shit, what exciting years we have ahead. Thank you so much for doing the show.
Roman Chernin: Yeah, thank you for inviting me. Glad to be here.
Harry Stebbings: I’d love to start with a question that’s at the top of a lot of people’s minds: where are we in the AI infrastructure cycle? A lot of people see the capital flowing in and say, “It’s a bubble,” while others say, “It’s just the beginning.” Do you think we’re in an AI infrastructure bubble right now?
01:47-04:11
Roman Chernin: No, I don’t believe it’s a bubble. I mean, it depends how you define a bubble. Do I believe we’ll need to build tens or hundreds of times more infrastructure? I absolutely do.
I’m probably biased. I probably wouldn’t be in this business if I didn’t believe that. But I think we’re only at the beginning of this amazing moment. Jensen calls it “useful AI,” and we’re just at the beginning of real adoption.
Honestly, we may have only one use case that works at scale out of all the possible use cases. That use case is coding. Everybody is talking about coding, and it only really started working a few months ago.
Let’s put that in perspective: we’re only a few months past the point when we got perhaps the first use case that works at scale. We’re starting to see it applied here and there, and I think we’ll see many, many more use cases and much broader adoption.
If you look at almost any company in the world today—perhaps excluding the fastest-moving startups—and examine its AI adoption, you’ll see that it’s using AI for the first percent of its workloads and the first percent of its use cases.
Even large companies that are technologically advanced are just getting started. That tells me we’re only at the beginning. Even if you don’t believe everything Musk says about the future, space, and so on, just looking practically at enterprise adoption, we’re still taking the first steps.
04:11-07:54
Harry Stebbings: We’re completely aligned, but it would be a very boring discussion if I just said I agree with you on everything.
We’ve seen coding work for the past, whatever, six to 12 months. But there’s a view that enterprises will move to locally hosted open-source models because the cost of using frontier models will become too significant. If that happens, wouldn’t it damage providers like OpenAI and Anthropic, as well as Nebius? Why is that view wrong?
Roman Chernin: First of all, I don’t think that’s the future—it’s already happening today.
What we see is that when a customer or product builder reaches scale, they start looking for ways to improve the economics, accelerate growth, and so on. That’s when many of them begin looking at alternative models.
The best way to build today is obviously to start with frontier models from great providers like OpenAI, Anthropic, and Google. They provide the best capabilities in the world.
But once you understand the use case, start seeing adoption, and develop a customer-data loop, you may find a cheaper—or not necessarily cheaper, but higher-quality—way to serve that same use case. You may not need the world’s best universal model. You can create a specialized model that works even better for your specific use case.
That’s when you may want to shift from closed frontier models to open source. The most important thing about these models isn’t just that they’re open source; it’s that they’re tunable and trainable. You can take them, post-train them, and create a specialized model that may perform better in your particular setting.
We see this across many use cases. But why doesn’t it hurt Anthropic and OpenAI? Because they move on to the next frontier.
Going back to the previous point, there are still so many unsolved tasks—and tasks that don’t necessarily have a limited budget attached to solving them. We saw this with DeepSeek a year ago, and we continue to see it now. Every time we find a more efficient way to solve one task, we start solving more complex tasks at the same time.
It’s a continuous journey. You’re always pushing the frontier. There are always more complex problems to figure out. Once you solve them, you can bring the price down or improve quality. But there are so many unsolved tasks that Anthropic, OpenAI, and all the other frontier-model providers still have such a large, unaddressed market. They continue to grow exponentially.
07:54-09:12
Harry Stebbings: Do you buy that? These companies are priced to perfection in many cases, at a trillion dollars. If the value they create is eroded, and they’re constantly playing a game of leapfrogging from one source of value to the next while open source continuously catches up behind them, they have to keep finding new problems to solve. That’s a hard life to live.
Roman Chernin: Actually, most people are concerned about the other side: will we have a strong enough open-source ecosystem, and strong enough specialized models, to build this floor?
I think we’re still at such an early point in adoption, and we have so many unsolved problems. It’s really a question of the total pie. There’s enough room to solve so many tasks in the future. There’s enough pie for frontier capabilities, highly tuned models for specific use cases, and the whole world of open-source and specialized models that we can build on top of them to gain economic and performance advantages once we know what we need.
09:12-11:15
Roman Chernin: You said that every time we get a cheaper model, it hurts the business. My favorite anecdote is from about 15 months ago, during the DeepSeek moment, if you remember.
Nebius stock went down 40% in a week, I think in February or March 2025. But that exact same week, we probably had our best sales week ever. The market was concerned that the market was going down and that infrastructure companies like Nebius would no longer be needed because, if AI became so much cheaper, perhaps it was a bubble.
At the same time, we had our best commercial week in the company’s history. We were still pretty early in our journey, but so many people realized they could run inference in production with DeepSeek and that the economics would work. At the same time, Cursor started growing. I think they were among the first to really benefit from tuning these models for coding and similar use cases.
Every time intelligence gets cheaper, we don’t reduce consumption—we increase it. We can solve more complex tasks with the same budget, or we can finally solve tasks economically that we already knew were possible but couldn’t scale because the economics didn’t work. It’s fascinating to watch those economic improvements play out.
Harry Stebbings: Speaking of Jevons paradox—producing more and seeing that yield even more demand—where are you not moving fast enough today? Where would you like to move faster?
11:15-13:08
Roman Chernin: Everywhere.
When we think about how we build the company, we talk about four dimensions. One is capacity: how many megawatts, gigawatts, and GPUs we deploy. We’re an infrastructure company, so we need to be large. If you’re not large enough, nobody needs you to exist. That’s the physical-world expansion.
The team is doing an amazing job, but it’s never enough. You always want to move as fast as possible, and there are lots of real-world complications that sometimes prevent that. To launch a new data center, you have to work through the entire supply chain, regulations, fires, floods—everything that happens in the real world. That’s one dimension.
Roman Chernin: Another dimension is product. You want to move fast enough to address new types of workloads and new types of customers coming into the market.
Think about it: we started this AI journey with the people building the models—companies like OpenAI, hyperscalers, large labs, and so on. What they need from an infrastructure provider is basically compute: just give them the infrastructure. We see a lot of these large bare-metal deals in the market, and we do them too.
But that’s only the first layer of what we build: scaled physical infrastructure that customers such as Meta or Microsoft can consume in large volumes. That’s the first layer.
13:08-16:27
Roman Chernin: The second layer is what we call multitenant cloud. It still serves research-heavy teams, but now you have hundreds or thousands of teams that don’t want to deal with physical infrastructure. They want managed infrastructure—the classic infrastructure-as-a-service cloud model.
You have storage, compute, networking, virtualization, APIs, observability, security—everything normal teams expect from a cloud. You log in, provision a cluster, and start training or running inference. You manage your application or workflow yourself, while the infrastructure is handled for you.
Roman Chernin: The first layer speaks in megawatts. If you read an announcement that someone has signed a large deal with Meta, Microsoft, or OpenAI, people talk about megawatts. You’re delivering megawatts of compute.
When you talk about managed cloud, people talk about GPU hours, because that’s the key unit you sell: efficient compute hours, alongside storage and complementary services. But you’re still buying managed compute.
The next layer we’re building is managed inference. Customers don’t want to think in GPU hours. They don’t want to figure out whether B200s, H200s, or B300s are better for a particular workload. They don’t want to manage vLLM or SGLang deployments themselves and do all the optimization.
That’s where our product, Nebius Token Factory, comes in. It’s a managed-inference platform. This serves a new type of customer—mostly what we call vertical AI companies and enterprises. The people who actually build products don’t build models; they build products on top of models. This relates to your point about specialized and open-source models: when they need to switch from Anthropic, for example, or diversify the models they use, they need a new primitive—a new kind of platform layer.
Now, we speak in tokens. You don’t pay for GPUs; you consume tokens, and you can build your applications without thinking about the clusters underneath. That’s where we are today.
But I don’t think it’s the final stage. People are now building agentic applications and workflows. When you build an end-to-end agent, you may not even think in terms of a particular model or a particular number of tokens you want to generate. You want the end-to-end task to be executed efficiently and to deliver the expected outcome.
The magic a platform can provide is to think for you: which model is best for this particular call? Do you need to use the smartest model? Or, within the same inference budget, can you ask two lighter models and then use a judge model to select the best result? What context-window size should you use, and so on?
That’s the next layer, where developers may not even think in terms of specific token types. They think in terms of the end-to-end execution of their task.
16:27-18:49
Harry Stebbings: So layer four is a direct competitor to OpenRouter?
Roman Chernin: What we’d love to bring at that level, just as we do in the layers below, is an optimization engine.
You can build your agent using any number of open-source or proprietary tools. But when you need to scale it, you start thinking about economics, reliability, and repeatable execution. At that point, it’s not just a model-choice problem or an outcome problem—it’s a systems problem. You need to make it reliable, repeatable, and economically viable.
That’s probably where Nebius can create value. We don’t tell people how to build their applications. We say: if you need this model to work for you at a given economic profile, we’ll help you optimize it.
It’s the same with agents. If you need an agent to run end-to-end within a particular budget and quality threshold, maybe we can help you optimize that.
To be clear, this is somewhat speculative thinking about what comes next. It’s not something we already have. But this is how we see our customers evolving, and where we think we can create the next layer of the product offering.
18:49-20:46
Harry Stebbings: I love this, and I have all these notes. I want to go through the four pillars you mentioned. Number one: capacity. If you had 10x the capacity today, what would be different? Could you sell it overnight?
Roman Chernin: That’s a good question. Not overnight, but we would definitely have demand for it.
The key question for us isn’t whether we have demand. It’s how we build a portfolio of demand, because there are many different customer types in this market that you can balance across.
Again, going back to the four layers of the product: you can sell bare metal, managed infrastructure, inference, and perhaps new product layers in the future. What we try to do is build a diversified customer portfolio.
We believe that the higher up the stack we move, the more value we can potentially create for customers. And the higher up the stack we move, the larger the customer population we can serve.
At the bare-metal level, there may be only a dozen customers in the world you can work with. In managed infrastructure, there are hundreds. In inference, there are thousands. In agentic applications, there will be tens of thousands of developers building products.
Harry Stebbings: On the customer portfolio, I love that. For capacity, you want customers to be large enough to be meaningful, but not so large that the business relies on them.
Given that tension, what level of revenue concentration with a Meta or a Microsoft are you comfortable with?
20:46-23:02
Roman Chernin: It’s a great question, and I’d say it’s one of the central questions of our business—not just for Nebius, but for the product category as a whole.
We’ve always said publicly, and to our investors and customers, that Nebius’s long-term strategy is to serve as diversified a portfolio as possible. We do our best to work with many customers.
In reality, if you’re serving a dozen companies in the world at the level of Meta or Microsoft—companies that are extremely advanced and have their entire software stack—they literally need only physical infrastructure. They bring everything themselves, deploy it on your infrastructure, and run it.
There’s very little additional value you can provide above the physical infrastructure. Although, by the way, meeting their physical-infrastructure requirements is itself quite a challenge. They’re very demanding, and they need the most scaled infrastructure in the world.
Sometimes people say this is a commodity business, but it’s not really a commodity at that scale. Nothing is commoditized when it comes to operating at truly massive scale.
But, to your point, that’s a relatively small population of customers, and you don’t necessarily need a full software stack to work with them. So, intentionally, from day zero at Nebius, we’ve been building this software stack. We thought it would be much more beneficial for us—and, if I want to be empathetic to the world, to have someone who can support customers not only at the physical infrastructure layer, but beyond that.
Harry Stebbings: For the long-term protection of the business, don’t you have to build the full stack? Otherwise, you become the capacity provider to these mega players, who will make a ton of money, while you’re incredibly concentrated and very vertically focused.
23:02-24:19
Roman Chernin: Yeah, I think so. Again, we don’t know where the world will end up. In a world of infinite demand, you may be able to sustain long-term and midterm bare-metal contracts.
But the more competition you have on the demand side, the more selective you can be about the customers you work with. You can work with customers who appreciate the value of the platform we’ve built.
There are different kinds of customers in the world. Some are more obsessed with price; some are more focused on quality; and some want a much more advanced platform because they want to focus on their own platform or product rather than spend time on the infrastructure.
Harry Stebbings: Before we move to number two, product, just staying on capacity: given the insufficient supply of capacity today, if you doubled pricing, would you see any change in demand?
24:19-25:48
Roman Chernin: It’s a difficult question. We actually raised prices just a couple of months ago, and we still have significant pipeline pressure on supply.
Again, we don’t really know where the balance is. And I’ll tell you why: it’s not just us being greedy and wanting to make as much money as possible because people will still have to pay in a shortage. People need compute in order to build.
But there is a point—especially as we move toward inference. It matters less in training, because training is more of a one-off cost. But inference is the cost of serving the customer. There’s a level at which the economics no longer work.
If our customers’ product economics work, they can grow, and then we can grow with them. It’s not simply a supply-and-demand situation with completely inelastic prices. Prices are elastic to some extent.
We also want to be thoughtful about what our customers need. And, by the way, it’s not only about the GPU-hour cost. It’s about all the optimizations you do and the real total cost of ownership. This is partly why we built the software platform.
25:48-27:46
Roman Chernin: Sorry, I keep coming back to product when you want to talk about capacity. But people are too obsessed with capacity—especially with the nominal price of capacity.
You can price a GPU at $3, $4, or $5, and depending on the use case and the quality of the platform, it can create completely different real costs for the customer. How long does it work? What is the effective uninterrupted amount of time you can run it?
If you’re talking about inference, how many tokens can you extract? We see optimizations that change the price of tokens by orders of magnitude. People talk so much about the cost of a particular GPU, but if you do the right things with the model, you can change the cost dramatically.
All of this needs to work together as a system. If you only provide raw infrastructure, you can only manage the price. But if you build the platform and provide customers with a high level of service, you can unlock much more value than just the infrastructure cost structure.
Harry Stebbings: If we move to that second layer—away from capacity and GPU hours, to the product itself, multi-tenant—what is the main question you ask yourself in that segment? If the question in capacity is revenue concentration, what is the big question in this layer of value?
27:46-31:46
Roman Chernin: What does the customer need? You speak with a lot of product founders, and it’s the same question: what does the customer need at the end of the day? How are customers’ needs evolving? Where is demand moving?
We see the transition from training to inference. We see the transition from simply using models to building agents. And we see a transition from AI labs being the primary consumers of AI compute to enterprises entering the game.
If we want to remain relevant, we need to follow those changes. That’s the main question we ask ourselves on the product side: what do customers need, and what value should Nebius create?
We’re a small company. We can’t build everything, so we need to be very precise about what we can do better than others and where we should focus, given how customers are evolving.
Harry Stebbings: What changes are you seeing in customer needs that aren’t being discussed much publicly?
Roman Chernin: Everybody talks about the shift from training to inference. I think that’s a very high-level, 100,000-foot view. In reality, this shift means people are building specific products, and those products have their own economics and growth trajectories. It’s not simply that the same GPU is being used for another purpose.
I think it creates new requirements. You need to build an inference platform. You need to help customers not only run inference, but also answer: where does the model they’re using for inference come from?
Everybody is taking open-source models and fine-tuning them. How do we help them do that? Then, when they run those models, they generate a lot of data. How do we help customers collect that data once their application and inference workloads are running, and then use it to improve the model or the application they run? People like this flywheel analogy: you run inference, generate data, observe that data, improve the model you’re running, and continue improving the quality of the end product.
I think there are a lot of pieces, both at the systems level and at the AI magic level, if you will. The most fascinating thing for me is that the barrier to building is coming down. We’re seeing more and more builders enter the market who aren’t necessarily AI researchers or inference engineers.
The value companies like Nebius can create is lowering the barrier to building AI-enabled products and applications that actually work—hiding from developers all the complexity of infrastructure, as well as AI complexity, such as how to tune a model or optimize inference. It’s a very research-heavy area. We should let people focus on their customers and use cases, in the same way they do in closed ecosystems like Anthropic and OpenAI.
31:46-34:25
Harry Stebbings: You mentioned differentiation. One theme my partner and I were discussing before this is the importance of building product on top of capacity. When people compare you with other neoclouds—say, CoreWeave—you both run GPUs, both have relationships with NVIDIA, and both have Meta as a customer. What’s the difference?
Roman Chernin: I don’t like comparing ourselves with others. The principle we build around is full-stack integration. You can think of it as full stack down and full stack up.
Full stack down means we go deep into the physical world. We build data centers, racks, servers, and the platform. When you control those things downstream, you can move faster, reduce costs further, and provide more economically viable solutions for customers.
Then, upstream, vertical integration is about product: following customers’ needs and customer segments, and not being limited to the relatively small population of people who just need infrastructure. It’s about serving enterprises and product companies, meeting them where they need us.
I think that’s what makes us different. How that shows up is in lower concentration in the business and a more diversified customer portfolio. Long term, we believe we’re better positioned to go after enterprises, where a lot of demand will eventually come from.
Today, most of our business is with AI-native customers. But there is a huge market of existing enterprises, and someone needs to serve them. They won’t buy raw compute. They’ll need platforms and tools. They’ll need us to respect their legacy environments and work with more complex systems. They aren’t nimble—they have data to migrate and systems to integrate. That’s the big game, and it’s the main direction for us.
34:25-37:18
Harry Stebbings: You mentioned the third layer of the four-layer stack: managed inference. For people who don’t understand it, how do you think about that layer, and how would you explain it?
Roman Chernin: It’s very simple. You’ve built your product on—what would you call it? You write code—
Harry Stebbings: I’m actually an OpenAI-in-code guy.
Roman Chernin: Okay, good enough. You’ve built a great product with OpenAI. You’ve found the use case, started growing, and have amazing traction. The only problem might be that you don’t have enough margin, or you want to use your data more aggressively and tune the behavior of the model—but you can’t do that in a closed ecosystem.
So you go online and read that there are a lot of great open-source models that, on benchmarks, are close to OpenAI. You think, “Great, inference will be 10 times cheaper. I can tune these models, apply my data, make my product better, and accelerate growth.”
So you take the weights from Hugging Face, use an engine to run them—like vLLM or SGLang—and then it doesn’t work. To extract the value you expect, you need to optimize. You need to deploy it properly. It’s not just about generating tokens on one GPU or one host. You have a large product that may run on hundreds or thousands of GPUs. You need orchestration, caching, and observability. Your customers ask you how it works, and so on.
By the way, you had all of that with OpenAI, because it was a production service for you. You don’t think about infrastructure when you work with OpenAI. You subscribe to the plan you need and pay for the outcome.
That’s where a product like Token Factory comes in. Token Factory gives you managed inference using open-source or specialized models. You can run an existing, vanilla open-source model, tune a model, or deploy your own weights. Then we take care of everything else. We apply all the optimization techniques, manage the economics for you, and make it reliable. You don’t need to worry about where you’ll find the next 100 GPUs. It’s a managed service.
37:18-39:31
Harry Stebbings: With Token Factory, you run 60 open-source models, and you said earlier that optimization can cut inference costs by up to 70%.
Harry Stebbings: Through optimization. Can I ask a dumb question? How do you actually make a token cheaper?
Roman Chernin: Yeah, it’s not magic. You take a baseline model and optimize it for the particular scenarios you have.
You can distill the model. You can make a smaller model that delivers the same quality. You can use speculative decoding, optimize caching, and so on.
So, you take the model and build a system around it that, for your particular use case and requirements, delivers optimized economics.
And one reason it’s important for customers to use managed platforms like Token Factory is that models are changing every week, every month. Maybe MiniMax 3 was released today, and another model was just announced. This happens every few weeks.
Each new model may perform better on some benchmarks and worse on others. You want flexibility. You want someone to support you in experimenting with and adopting the best new models for your use case whenever they become available.
Platforms like ours abstract away all the work required to move from one model to another, benchmark them, and so on. You can be confident that you’re always on the frontier. Whenever something new happens, it will be on the platform. You can test it, and if it works better for your use case, you can switch smoothly and transparently.
39:32-41:34
Harry Stebbings: Does the pace of model development sustain? You said every couple of weeks; respectfully, I’d say every couple of days there’s something new. Do we still see that level of iteration in five years?
Roman Chernin: I don’t know. There’s a good chance we’ll continue to see lots of niche models emerge and improve.
I’m a believer that we’re still quite far from the wall, and that we’ll see a lot more model improvement. We’re also seeing many more modalities and specialized models coming into play.
We talk about frontier LLMs, but there’s an entire world of life sciences models, robotics models, world models, video models, and image models. They all have their own use cases. And we’re seeing more and more smaller, highly optimized specialized models for particular use cases.
Just this morning, I spoke with a team here in Israel developing a foundational cyber-defense model—one optimized to build cyber-defense agents. They aren’t starting from scratch. They take an open-source foundational model, then train it for the particular use case and optimize it for the quality and latency required in cyber-defense applications.
I think we’ll continue to see that: lots of specialized, post-trained models that still need optimized inference and optimized infrastructure around them for customers to use.
41:35-45:42
Harry Stebbings: Going back to Token Factory, token costs, and token usage: what are you seeing that people aren’t talking about enough? What has shocked you recently?
Roman Chernin: I think everybody is talking about the same thing: how fast it’s growing. When we see the trajectories of companies like Anthropic, Cursor, and Cognition in coding—and now we’re starting to see this in other verticals too, including healthcare and financial-services use cases—it’s quite amazing.
What’s interesting is seeing how non-AI-native startups are moving. For example, Revolut is a customer of ours. When we started working with them, I think 99% of their inference budget was spent on closed models, primarily OpenAI.
They began to crack some use cases, but some didn’t work economically. They couldn’t practically replace or meaningfully enhance humans in the use cases they wanted to address. So they started moving to open-source models.
But the move wasn’t fast, because they had to spend time building the whole internal engine. First and foremost, they focused on evaluations. I think people underestimate how important it is to build the foundation for improvement—an experimentation engine.
As a company or team, you need to understand what is good for you. You can close a use case and make it work, but then you want to change the model. How do you know you haven’t ruined the quality?
You need metrics. You need a validation mechanism. You need a CI/CD process established for AI development.
What we see with customers like Revolut is that they need to make these foundational investments to understand how to evolve models and safely integrate them into production processes.
But once they solve those foundational problems, they begin growing exponentially. I wouldn’t underestimate how quickly those customers can grow once they build a system that lets them ship fast.
And shipping fast means knowing how to evolve and how to make decisions.
We see this across many customers. They have what you could call foundational investments, or a cold-start problem: how do you start shipping? But once they solve it, they start growing exponentially. They can use different models, build many more products inside the company, and so on. From the outside, people look at these companies and think, “They’re starting small. They’re not growing yet. It takes time.” But if a company has a strong team, it builds the foundation first, and then it starts growing exponentially.
I think we’ll see a lot of explosive growth among enterprises, digital-native companies, and cloud-native companies—companies like Revolut, Shopify, and Booking.com. Once they solve the cold-start problem and build the systems needed to ship, their AI adoption will grow like crazy.
45:42-47:10
Harry Stebbings: How much more do you think Revolut will pay you in three years’ time?
Roman Chernin: I don’t know. I don’t want to comment on that specifically. But I can say that, overall, they’re growing very quickly.
We all see AI companies reporting ARR growth. For enterprises, it isn’t ARR—it’s their AI budget. But I think the most advanced companies are growing their AI budgets at a similar pace. This isn’t some speculative capex race; we see it in their production workloads. Their AI consumption is growing on the same exponential trajectory as AI-native companies’ ARR.
Harry Stebbings: I always push back on people who say open source will be a credible threat to the largest model providers. I say, listen: the biggest enterprises want reliability. They want security. Most of all, they want ease of use. They don’t want to tinker with all the architecture and everything beneath the surface.
What you’re telling me is that you can provide all of that, allowing them to move away from those providers and have a cheaper, better experience because you take away the plumbing. Correct?
47:10-50:35
Roman Chernin: Yes, but again, my point isn’t really closed models versus open-source models, or reliable versus unreliable models. The job of companies like Nebius is to make it possible for customers to use alternative models without having to think about the plumbing.
But fundamentally, it’s about capabilities. Closed-source frontier models are great, and they’ll become even better. They’ll solve many problems that we can’t solve today. At the same time, there is such a diversity of use cases that there will be a market for the smartest models in the world, the fastest models in the world, and the models in between—smart enough, but cheap enough.
As a customer, you’ll be able to choose the right source of tokens for each specific task. And, returning to the agentic-layer point, perhaps it won’t even be the customer’s job to choose which model to call. An engine will understand the capabilities of all the underlying models.
When you use OpenAI for research, you don’t think about how many loops it should run, when it should use an LLM versus search, or which prompt it should call. You give it a task, the reasoning engine determines how to run it, and you get the result.
I think many enterprise use cases and agentic tasks will work the same way. Developers focused on customer needs won’t have to orchestrate all of these tokens and models themselves. We’ll need every kind of model: the smartest models for the most complex intelligence tasks, and fast models for quick iterations.
And we haven’t even discussed all the modalities, or what we’ll need in the physical-AI world. There will be enough room for different models. What we need to do as an infrastructure company is help developers use all these capabilities comfortably. As you rightly said, it’s not only about model capabilities. It’s about removing the plumbing, making the models work, optimizing them, and making them reliable.
50:35-52:58
Harry Stebbings: When we look at the explosion of models, their specialization, how many will be built, and the depth across different use cases, one thing is quite clear: Europe does not have anywhere near the model buildout we’ve seen in the US and China.
How important do you think it is for nations to have their own sovereign models?
Roman Chernin: It looks like the world is becoming divided, whether we like it or not. I think it’s important that sufficiently capable foundational models are available to major parts of the world.
Here in Europe—or at least in this part of the world—we should think about how we ensure enough capabilities are available locally. We’ve had a lot of conversations over the past few years about sovereignty and sovereign AI, but I think the discussion has been too focused on megawatts and power rather than on the builder layer.
The megawatts will come. At Nebius, we’ve always said that we—and companies like us—will build infrastructure if there is demand. And demand comes from the builders.
What we need to care about is having more great companies like Lovable, Black Forest Labs, Mistral, and others. We need enough people investing in research and enough people investing in products. Then they will create enough demand, and we’ll have enough of a flywheel to build good-enough models if we need them. I think that’s something we should care about.
Harry Stebbings: Where is the most interesting area to invest today? I’ll give you four options: infrastructure, horizontal models, vertical models, or the application layer.
52:58-54:21
Roman Chernin: We build infrastructure, so we’re quite happy here. I think it’s a good place to be in the current world.
Even though, to some extent, we’re building the easier part—not that it’s easy; execution is complex—we broadly know what’s needed, and our customers help us understand what’s needed.
I think the most amazing people in this industry are the ones taking the risk to build end-user products. They drive most of the growth. They’re taking the real risk of building something that people may—or may not—need. To me, they’re the heroes of this AI journey.
Harry Stebbings: Speaking of the heroes of AI, before I do a show, I’m fortunate enough to speak to some big people. You mentioned earlier that I’ve interviewed some of them.
One theme that came up was the relationship with Nvidia. Is it really a marriage if one party has more power than the other? How do you think about the power dynamics of a relationship with Nvidia, when they have so much power?
54:21-56:37
Roman Chernin: We look at it very simply: we need to build what we build. We need to build our product, tell our story, and the rest will follow.
What’s fascinating is that Nvidia is still, to a large extent, an engineer-driven company. My view—though they may see it differently—is that the best way to earn Nvidia’s respect is for Nvidia’s engineers to respect your engineers. That gives you the right foundation for a relationship.
I think we’ve proven again and again that we know what we’re building and that we have a strong engineering team. They see that and respect it. We have a lot of engineer-to-engineer relationships: at the physical hardware level, the software layer, and the inference-platform layer.
The more highly Nvidia’s engineers think of you, the better the relationship and partnership you can build. Maybe we’re wrong to think about it this way, but that’s what we believe we can control. We just focus on being reasonable and focused on long-term value.
It sounds fluffy—everyone says it—but at the end of the day, just do your fucking job.
Harry Stebbings: I’m going to title this: “Roman, Just Do Your Fucking Job.”
Roman Chernin: What else can we do? We’re in such a race. We can only do our best to do our work better. That’s it.
Harry Stebbings: “Just do your fucking job.” I know, it’s funny. I like it. But what’s the hardest part of just doing your fucking job today?
56:37-58:36
Roman Chernin: There are four dimensions: building at scale, building the product, working with customers, and capital.
We’ve discussed scale and product. The third is customers. We’re in a field business. We like to say that cloud is a post-sales business: when you sell, you sell a promise, and then you need to fulfill that promise for the customer.
Working with customers, taking care of customers, and having a strong customer-facing engineering team—an FDE team—is the third dimension. Go talk to your customers. Make sure they know you and you know them.
The fourth dimension—the most boring, but also the most exciting—is capital. We’re in a capital-intensive game, competing with the most highly capitalized companies in the world.
Harry Stebbings: If I gave you an unlimited budget, what would you do differently?
Roman Chernin: Build faster. That’s very easy.
Harry Stebbings: Build what faster?
Roman Chernin: Data centers, and fill them with GPUs. Just build faster.
Our capex program this year is €2.5 billion. The hyperscalers have programs that are eight to ten times larger. If I had ten times more capital, I would build more data centers, fill them with GPUs faster, and serve more customers.
That’s where we started: what would I do if I had ten times more supply? I would move faster.
Harry Stebbings: Gavin Baker said, quite intelligently, that permitting, regulation, and the delayed buildout of data centers have actually helped. If I enabled you to build ten times as many data centers today, it would actually create a glut.
58:36-1:00:51
Roman Chernin: It’s a great question. Our investors sometimes ask us what the main bottleneck is. The answer is: everything. But you have to look at it across different time horizons.
Over the next six months, capital can’t really help. Six months is too short; you have what you have, and you need to deliver.
Over the next 12 months, you can accelerate some things, but you’re still constrained by capacity. With capital and execution, we can accelerate certain things.
But over 24 months, you can unlock a huge amount. It’s important to understand that we’re not building one data center; we’re building a portfolio of capacity. The more execution power and capital we have, the more things we can do in parallel.
That’s how we operate: first, we secure power and land. Then we build data centers. Then we fill them with GPUs. Every subsequent stage requires more capital, but we do as much as possible in advance. By the time we reach the next stage, we already have the power secured; once we have enough capital to deploy into GPUs, we’ll have data centers up and running. So it’s a phased investment process.
Again, the bottlenecks differ depending on the time horizon. Obviously, if you have more capital, you can move faster—not within six months, but certainly over 18 to 24 months.
Harry Stebbings: Can I ask: when you think about the data center buildout, we’re seeing more and more public angst toward AI. Eric Schmidt is getting booed off stage—not because of what he’s saying, but because of AI itself.
And we’re seeing public resentment toward data centers. I think 40 out of every 100 projects now aren’t getting built once they go through planning and approvals. How do you think about that internally?
1:00:51-1:02:49
Roman Chernin: This is the environment we need to operate in. There are two sides to it.
First, there’s how we think about it pragmatically as a business. As I said, we view it as a portfolio of projects. We need to make sure we’re oversubscribed, if you will, so that if one data center is delayed, we can still deliver enough capacity to our customers.
Most customers aren’t locked into one physical location. It’s a cloud: we can build in different places and bring workloads to wherever we have capacity. That’s the pragmatic side.
But, obviously, communities and local authorities expect companies like us to work closely with them—to explain what we do, demonstrate the value, understand their concerns, and address them. That’s the reality.
You can compare it to when Uber started growing. In many places there was pushback: “What’s happening? This is something new. It’s moving too fast. We didn’t expect it to move this quickly.” You have to go in, engage, and explain.
It’s simply part of your responsibility to work with the communities that become dependent on you. They have concerns. Sometimes those concerns come from a lack of information; sometimes they’re entirely rational concerns that you need to address. That’s part of the job.
Harry Stebbings: Do you think you’ve done a good job at that so far?
1:02:49-1:05:08
Roman Chernin: We always come from a place of thinking we haven’t done enough. But I think we’ve made quite a lot of progress in the places where we started building.
Historically, we had more experience in Europe. Now, probably 70% to 75% of the new capacity we’re building in the medium term is in the U.S. So we’ve built a significant on-the-ground presence to communicate with local communities in the U.S., and we try to do the best job we can.
We always need to do better, but we’re making progress.
Harry Stebbings: Can you help me on another one? We laughed earlier when we talked about space. Data centers on planet Earth are a—
Roman Chernin: Data centers on Earth are already a very difficult logistical buildout. Data centers in space? I love technology, I’m an optimist, and I hope it happens—but that sounds nuts.
Roman Chernin: No, my view is very simple: so many smart people are now working to make it happen that I may be less pessimistic than I sound. I don’t know whether, in three years, we’ll be building more in space than on Earth. But I’m humble enough to say that so many smart people are trying to solve this problem and bring compute into space—why wouldn’t I believe it could happen?
There are still a lot of challenges and a lot to figure out. But if someone had told us even three years ago that we’d be building multi-gigawatt data centers—large, interconnected compute clusters—would you have believed them? I certainly didn’t think that way. Now we’re here, and it’s routine.
Harry Stebbings: I want to do a quick-fire round with you. I’ll give you a short statement, and you give me your immediate thoughts. What job doesn’t exist today that you think will be very common in five years?
1:05:08-1:07:30
Roman Chernin: One thing that’s obviously happening is that we’re democratizing what people call being a developer. Each of us can become a developer. And by “developer,” I mean someone who can turn an idea into a digital asset.
Again, we have to be optimistic here. I hope this democratization of building—giving every one of us the ability to create—will open up so many opportunities that we can’t even imagine yet. When tens of millions of new people can easily turn their ideas into something that works, we’ll see a lot of new businesses and ideas come to life. They’ll create many new jobs that we don’t even know exist today.
It’s the second-order effect of democratizing the ability to build. But what’s challenging—and what will need to change, and is as much a risk as an opportunity—is education.
When everyone has access to intelligence, what should people learn? You definitely don’t need people to memorize facts. Everything is available; all human knowledge is accessible. So how do you train people to think when they don’t need to think in the same way as before? How do you teach people to adapt continuously when many professions may no longer be stable? How do you help people find themselves in a changing environment, and actually think and learn new concepts constantly?
I think it creates a lot of new opportunities, but also a lot of risks.
Harry Stebbings: You mentioned that you have two teenage daughters. What advice do you give them as they enter the workforce over the next 10 years?
1:07:30-1:09:55
Roman Chernin: What I literally tell them is that I don’t know exactly what will be needed, but I’m sure two things will matter.
First, the ability to communicate with people—with empathy. To understand humans, communicate with them, and be empathetic.
Second, creativity: art, the ability to try new things, and to create. I hope art will continue to exist in some form.
Ten years ago, I thought the most important things they needed to learn were hard skills—math and engineering. Now, I’m much less convinced of that. I’m actually quite happy that they’re much stronger in soft skills than I was as a kid.
If you can help your children develop empathy, the ability to understand and communicate with people, and creativity, I think they’ll be in demand in 10 years.
Harry Stebbings: There’s a question of how you teach creativity, but I completely agree with you.
The big finish: complete this sentence. The biggest threat to Nebius is not competition, but...
Roman Chernin: Consolidation, in general.
I think the main threat to Nebius as a business is a world that becomes too consolidated. As we discussed, we try to stay diversified. We try to solve problems for different customers across different layers.
If we end up in a world where three to five super-models, super-companies, or super-empires control everything, then Nebius, or companies like Nebius, will only be needed to help serve their needs at the physical infrastructure layer.
So, in general, consolidation is our main threat. The more democratized and diversified the world is, the more we’re needed as a business.
Harry Stebbings: Do you think that’s likely? We’re seeing value concentrate among fewer and fewer players. We’re seeing the opposite of diversification.
1:09:55-1:12:17
Roman Chernin: I hope it won’t happen. It’s better for us as a business, but it’s better for us as humans, too—for you and me—that the world remains quite diversified in different ways.
I’m optimistic. There are so many people who want to build things independently. So many people have the need to experiment and create new things. That organically creates pressure for a more diversified world.
So hopefully, it will remain that way.
Harry Stebbings: Penultimate one: Leopold Aschenbrenner is a famous investor right now and has a huge cult following. He recently disclosed a very large position for him—5.3% of the company. I think it’s 15% of his portfolio.
Harry Stebbings: How do you guys view that internally? Are you like, “Yeah, go Leo”?
Roman Chernin: I wouldn’t say we didn’t notice it—obviously, everybody noticed it. The stock jumped, and it was big news all around.
Again, I think we take it as validation of what we do. But when you get that validation, you should tell yourself: these people are giving you credit that you will execute.
I come back again and again to the fact that ours is a post-sale business. Every time we sign a deal, every time someone invests in us, they give us credit and an opportunity to deliver. Then you go back to your job and deliver.
We’re also in a very emotional market, so you have to keep your feet on the ground. Remember that all this growth, all the credit customers give you, is an opportunity to deliver. Go and do your job.
Harry Stebbings: You’re such an Israeli. Americans would be like, “Yeah, go!” You’re like—
1:12:17-1:14:19
Roman Chernin: I think I’m Russian in this way. Russians always know that you need to look at things very pragmatically. You know, Russians always have these faces—always expecting that something will happen, and that you need to be ready.
I think that’s a really important part of what comes from our CEO and founder, Arkady. You wake up, it’s a new customer, it’s a new day, and you need to deliver. Nothing is guaranteed. You need to concentrate on the work.
I know how much effort the team puts into making things work, and how much depends on every day’s work and dedication. The market is moving so fast, and to stay relevant, you need to keep moving at the same pace—or try to move at the same pace—as the market.
On a more romantic note, I would say that we could celebrate a little more. We just don’t take enough opportunities to say kudos to the team. I don’t think we celebrate enough.
I think it’s right that we’re not relaxed, but we could celebrate a little more and give the team more recognition for how much has been done. It wasn’t easy, it still isn’t easy, and it won’t be easy.
But, yeah, never stop. We cannot stop. It’s like a shark: you’re alive when you move. So we have to move.
Harry Stebbings: On that note, I cannot thank you enough for joining me and for putting up with my very meandering questions. You’ve been fantastic, Roman. A really huge thank you.
Roman Chernin: Thank you. You’re too kind to me.
Made with: The Transcript Desk Chrome Extension

