Why the Future of AI Is Not in the Cloud 232

Published 2025-10-31 · Updated 2026-05-23 · 5 min read · AI Hardware and Infrastructure · By Sahin Boydas

After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.

Everyone talks about AI models, but nobody talks about the brutal reality of the hardware that runs them. Here's the unfiltered truth about what it really takes to build and scale AI infrastructure, a truth I learned the hard way over 20 years in Silicon Valley, two exits, and writing checks for over 200 startups, including AI leaders like Anthropic, OpenAI, and Scale AI.

I remember the exact moment the cloud-first AI dream died for me. It was 2019, and at RemoteTeam, we were building a new feature that relied heavily on a series of language models. We started on a major cloud provider, and everything was great. Our initial server bill was maybe $5,000 a month. We felt like geniuses. Then we scaled. Within six months, that bill was pushing $150,000. A month. I’d wake up in the middle of the night sweating, thinking about that number. We were a hostage to our own success, and the cloud provider was the one holding the gun. That experience taught me a lesson that has shaped my entire investment thesis since: for AI, the cloud is a great place to start, but it's a terrible place to finish.

The Great GPU Lie

Right now, building a serious AI company feels like trying to build a city in the middle of a desert with a global sand shortage. That sand is the GPU. Everyone is fighting for the same handful of chips from NVIDIA. We talk about a 'shortage,' but that doesn't capture the reality. It's a strategic bottleneck. If your startup's success depends on your ability to secure a few hundred H100s, you don't have a business plan; you have a prayer.

I’ve seen brilliant teams with groundbreaking models get stuck in neutral for months, not because their tech was flawed, but because they simply couldn't get the hardware. They were waiting in line behind the big tech giants who can pre-order tens of thousands of chips at a time. It’s a rigged game. Relying on the same public clouds and the same chip manufacturer as everyone else is not a strategy for winning; it's a strategy for mediocrity.

This isn't just about cost or availability. It's about control. When you're completely dependent on a third-party cloud, you're subject to their pricing whims, their availability, and their roadmap. You can't innovate faster than they allow you to. You’re building your dream house on rented land, and the landlord can change the terms at any moment.

I have an investment in a small, incredibly bright team that developed a novel compression algorithm for large models. Their tech could reduce inference costs by 80%. But they couldn't get enough GPUs to even run a proper benchmark to prove it. They spent a year in limbo, pitching VCs who loved the tech but saw the hardware dependency as an insurmountable risk. They almost died. That’s not a market failure; it’s a market controlled by a single player.

Data Has Gravity

The other big lie is that data can live anywhere. It can't. Data has gravity. The bigger the dataset, the harder it is to move. The idea that you can just shuttle petabytes of data back and forth to a data center a thousand miles away is a fantasy for most real-world applications.

Think about one of my investments in the autonomous vehicle space. The car generates terabytes of sensor data every single day. You can't upload all of that to the cloud for processing. The latency would be a nightmare, and the cost would be astronomical. The only solution is to process the data right there, in the car. The compute has to go to the data, not the other way around.

This applies everywhere:

  • Manufacturing: Factories with thousands of sensors need real-time analysis to predict equipment failure. You can't wait for a round trip to the cloud. A single millisecond of downtime on a production line can cost tens of thousands of dollars. The decision to shut down a machine has to be made in an instant, not after a leisurely trip to a server farm in Virginia.
  • Healthcare: AI-powered diagnostic tools need to work instantly, on-device, in a doctor's office or a hospital, not on a server in another state. When a surgeon is using an AI-guided robot, the feedback loop has to be instantaneous. Lives are literally on the line.
  • Retail: Analyzing customer behavior in a physical store requires immediate processing at the edge. A retailer wants to know what a customer is looking at right now to send a targeted promotion to their phone, not what they looked at ten minutes ago.
  • Finance: High-frequency trading firms live and die by microseconds. Their algorithms need to be physically located as close to the exchange's servers as possible. The speed of light is a real and unforgiving constraint.

This is the concept of “data gravity.” The compute and the AI inevitably get pulled closer and closer to where the data is generated. That place is almost never a centralized cloud data center.

The Return of the Metal

So what’s the answer? It’s not about abandoning the cloud entirely. It’s about a radical shift in thinking. It's about bringing the compute back in-house, to the edge, and even onto custom silicon.

When we were building MovieLaLa, we hit a wall with our recommendation engine. The off-the-shelf cloud solutions were too generic. We ended up designing a small, specialized hardware accelerator to handle a very specific part of our algorithm. It was a huge pain. It took us nine months and cost a couple of million dollars we barely had. But it cut our processing time by 90% and our costs by more than half. That move was a huge reason Gfycat acquired us.

Today, this idea is becoming mainstream. Look at the most innovative companies:

  • Google has its TPUs (Tensor Processing Units).
  • Apple has the Neural Engine in every iPhone.
  • Tesla is building its own Dojo supercomputer.

They aren't doing this because they like building hardware. They're doing it because they have to. It's the only way to get the performance, efficiency, and cost savings they need to win. They are escaping the GPU bottleneck by building their own tools.

For the next generation of startups, this is the opportunity. You don't need to build a whole data center. You can start with a single rack of servers in a colo. You can design a specific chip for your specific workload. You can focus on building for the edge. This is how you build a real, defensible moat. Your custom hardware and infrastructure become part of your product. It's the difference between being a chef who can only cook with the ingredients the grocery store has in stock, and a chef who owns their own farm.

The New Breed of Engineer

This shift to custom hardware and infrastructure requires a new kind of talent. For the last decade, the best minds in software have been trained to think in terms of abstraction layers. They build on top of AWS, GCP, and Azure. They don't think about the metal. They don't have to.

But to win in this new era, you need engineers who can go deep. People who understand everything from the physics of the silicon to the architecture of the data center. These are the people who can find a 10x performance improvement by rewriting a kernel, or design a cooling system that cuts power consumption in half. This is a throwback to an older generation of engineering, and it's a skillset that's in desperately short supply.

As a founder and an investor, this is what I look for. I look for teams that are obsessed with the full stack, from the application down to the hardware. I look for people who aren't afraid to get their hands dirty, who see the infrastructure not as a commodity, but as a competitive advantage.

My Investment Thesis: The New AI Stack

As an investor, I’m not looking for another company building a slightly better model on top of NVIDIA GPUs in an AWS data center. I’m looking for the rebels. The ones building the picks and shovels for the new era of AI.

My portfolio is full of them. Companies building new types of processors, like quantum-inspired chips that can solve optimization problems that are impossible for classical computers. Startups creating new interconnects that allow for massive, distributed training without the bottlenecks of today's networks. Teams designing AI data centers that are an order of magnitude more power-efficient.

This is where the real value is going to be created in the next decade. It's in the boring, unsexy, and incredibly difficult work of building the physical infrastructure that will power the future of intelligence. The cloud isn't the future of AI. It's the past. The future is happening on the edge, in private data centers, and on custom pieces of silicon designed for a single purpose. The future of AI is not in the cloud; it's on the metal. And that's where I'm placing my bets.

Frequently Asked Questions

How has this view evolved over time?

My thinking on most topics has changed significantly over the years. Early in my career, I held many conventional views that experience proved wrong. I try to update my beliefs when the evidence changes.

What's the most common pushback you get on this?

People often push back by citing exceptions or edge cases. And they're usually right that exceptions exist. But building a strategy around exceptions rather than patterns is a losing game for most founders.

How can I apply this thinking to my own situation?

Start by identifying the core principle behind the opinion, not the specific example. Then ask yourself: does this principle apply to my context? If yes, test it in a small, low-risk way before going all in.

What experience informs this perspective?

This perspective comes from over a decade of building companies in Silicon Valley, two successful exits (RemoteTeam to Gusto, MovieLaLa to Gfycat), and investing in 200+ startups including Anthropic, OpenAI, and Scale AI. I write about what I've lived.

More in AI Hardware and Infrastructure

  • From TPU to Your Own Custom Silicon: A Founder's Journey — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.
  • Surviving the GPU Apocalypse: A Founder's Guide to the Shortage — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.
  • The 6 AI Infrastructure Mistakes That Are Secretly Killing Your Startup — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.
  • Cerebras Systems — Portfolio Company | Angel Investment by Sahin Boydas — Building the world's largest AI chips for training and inference at unprecedented scale.
  • Why the Future of AI Is Not in the Cloud 338 — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.
  • The Counterintuitive Truth About AI Chip Design 869 — After years in the trenches of Silicon Valley, I've seen firsthand how the right AI hardware can make or break a company. I'm sharing the hard-won lessons and contrarian insights I wish I had when I started, from navigating the GPU shortage to building our own custom silicon.

All AI Hardware and Infrastructure articles · Sahin's angel investments · Startups he founded