Ian Buck, vice president of hyperscale and high-performance computing at Nvidia, walked a small group of press through the headquarters of Nvidia in Santa Clara, California.
The HQ — and I suppose its sister building Voyager — is the house that built Vera Rubin, a six-chip platform that is one of the primary reasons that Nvidia has a market valuation of $4.92 trillion. Soon, Nvidia’s Vera Rubin platforms will be shipping to the nations of the world that are building AI factories — I think they would do better to call them cathedrals of science and technology — dedicated to sovereign AI.
I’ve written four very technical stories today about Vera Rubin, with leaders like Buck talking through the strategy and design of their cherished products. But in this story, I’m using some of my gift of flowery writing to tell you why you should care and just how hard this work is.
The buildings themselves at Nvidia are crafted from triangles, the polygonal building blocks of computer graphics. The Nvidia buildings in Silicon Valley are shrines to computing, not unlike the Sagrada Familia cathedral designed by Gaudi in Barcelona, Spain. They’re a construct of the human imagination dedicated to the altar of AI. Only Nvidia’s god will last only two or three years before it’s replaced by the next best god or goddess with a different code name, while Sagrada Familia has been around for 150 years and Gaudi himself never finished designing it.

The Vera CPU, designed for Agentic AI and one of six chips in the Vera Rubin platform, offers two times better performance than the previous Grace Blackwell generation, three times faster core-to-core bandwidth, and 40% better latency, a measure of interaction speeds. The Perplexity AI model can run 1.5 times faster in completing jobs on Vera Rubin compared to the prior generation.
While Vera has 144 billion transistors and Rubin has 336 billion, the six chips of Vera Rubin may be around 695 billion transistors — so says Microsoft Copilot. By comparison, GeForce 256 had 23 million transistors back in 1999, and Intel’s first CPU, the 4004, had 2,300 transistors, the basic switches that are the building blocks of all electronics.
The fairy tale kingdom of Nvidia

The significance of the change in computing power is the difference between arithmetic and intelligence. People expect the high priest of intelligence to be at companies like Anthropic (valued at $965 billion), Open AI, Google DeepMind, xAI, Meta or Microsoft. But maybe the high priest is here inside Nvidia. At least Nvidia’s CEO Jensen Huang is there some of the time. Jensen Huang was away when I visited. He was in Japan, thanking Sega’s leaders like former CEO Shoichiro Irimajiri, who saved Nvidia 30 years ago.
Back then, Nvidia was commissioned to make a graphics chip for Sega’s Dreamcast game console. But the chip fell short of the targeted specs. Sega couldn’t use it. Nvidia had no revenues and no money and the would-be high priest was weeks away from bankruptcy.
But Irimajiri liked Huang, the scrappiest of 80 or so 3D graphics chip makers in Silicon Valley and the rest of the world. Irimajiri approved a $5 million investment and a cash payment for Nvidia, saving the young company from winding up like so many other 3D graphics chip companies — which were akin to this era’s blockchain gaming startups of the 1990s. In 1999, Sega sold off its Nvidia stake for a handy profit of $15 million or so. Had they hung on to that stake, it would have been worth hundreds of billions of dollars today. Now Huang and Sega’s leaders are teaming up again, bringing Sega games to Nvidia’s RTX Spark mini supercomputers.

Only three major graphics chip makers survived the 1990s fray, which gave birth to 3D gaming, which is still the heart of the $200 billion game industry today. And this week I’m heading to an event being held by the rival survivor Advanced Micro Devices. (Intel is the other one, but there’s little inference startups nipping at Nvidia’s heels too). AMD is run by Lisa Su, who is a distant cousin of Huang’s, and they’re both celebrities in Taiwan, which is home to Computex, a computer trade event that drew 111,000 people this year.
Nvidia, started by Huang and two other engineers at a Denny’s restaurant in San Jose, California. Nvidia used the cash from Sega to invest in Riva 128, a next-generation 3D graphics chip, and the subsequent GeForce 256, which kicked off the era of the programmable graphics processing unit (GPU). In turn, with Nvidia’s CUDA software, the GPU became useful for non-graphics computing.
By 2012, GPUs were the heart of parallel computing that had other uses besides gaming. AlexNet, a neural network, used GPUs to do AI computing. AlexNet was an intelligent neural network with 60 million parameters, running on two Nvidia GTX 580 GPUs. Now, just 14 years later, xAI’s Grok5 has six trillion parameters, or 100,000 times more than AlexNet.
This kind of advance in computing is unheard of. If Nvidia had followed Moore’s Law and software followed suit, the hardware should have only been able to deliver a 128 times increase in computing power in 14 years.

That isn’t magic. That’s the work of 42,000 engineers and other employees at Nvidia who have delivered magical results through non-magical scientific mastery of engineering, computer design, architecture, economics, and manufacturing that gives us silicon chips from sand. Yep, if you pound and melt sand enough, it yields glass-like semiconductor crystals that chips are made of. They’re honed into 12-inch-diameter cylinders and then sliced into wafers. On those 12-inch wafers, you could fit maybe 70 or so working Nvidia Blackwell GPUs.
Most of those chips are built in Taiwan at TSMC’s massive factories that cost $25 billion or $35 billion to build. Some of those factories are now being built in the United States again, in part because of a long bipartisan effort to bring tech jobs back from overseas. TSMC itself is building four such factories in Arizona, and Donald Trump is taking the credit. Huang has to behave when it comes to politics, and so he’s been asking the president to let him sell his AI chips in China, even if it helps China win the AI war. Huang’s logic? If you don’t let Nvidia sell its chips in every market, it could lose the AI chip war to a rival, like those in China, that can. And then where would we be? People are going to argue about this for a long time.
But I remember the lesson of short sellers who believed that Nvidia was going to collapse so long ago. Don’t bet against Jensen.
The Terminator syndrome

Trump is very happy to grant that right to Nvidia so long as he brings the jobs back to the U.S. And while those chip factories might be welcome in the U.S., the conundrum is that the data centers that they bring with them are not, as they’re sucking up too much of the electricity and water in America’s small towns. You can tell in some of Huang’s recent interviews that he’s fed up with the analysis that his doomsday chips will bring about Arnold “Ahnold” Scwharzenegger’s Terminator fate. I’ll be back.
But those people in the small towns, they have to get in line to criticize Huang and his company because gamers and game developers are even more upset. The game developers are worried that generative AI in particular could steal the jobs of artists, programmers and even business strategists inside the game companies. They’re also afraid of their human-crafted games getting lost in a sea of AI slop. Gamers fear they copycats and derivatives games will come by the millions, as AI will create games coming in single-sentence prompts coming from the mouths of babes. Even the Roblox kids will probably rue the day when AI arrives and takes over their paradise of player-created games.
The worst crisis in gaming history now has a huge hardware component. Memory chip prices are up seven times for some vendors in the last two years, and Valve, which is making Steam Machines, says it’s only getting worse. The memory chips are used in game consoles and high-end gaming PCs, and they’re in such short supply that game machines are pushing $800 or $1,000, far below the $199 that a Sony PlayStation cost in 1996. Back then, so many people could afford game machines. And now it’s so costly that gaming is becoming an elite thing. And that’s going to be bad news for Sony, Microsoft and Nintendo for years to come.

The prices of game consoles are going up, and that’s never happened before in gaming history, and it’s coming at a time when layoffs are rife among game companies. Gamers might be suspicious of greedy executives at game companies, but they may not realize that this a train — a technological or economic train — that no single CEO or billionaire capitalist on the planet can stop.
I just wish there were more strategy gamers like me out there, as we can see that all of this is a big game. And someday, gaming will get its revenge, as those AI chips will lead to AI games that everybody is going to want to play, and games will once again be most valuable products in the world. Or so I hope. I think that game will be a descendant of Cyberpunk 2077, or perhaps Grand Theft Auto 18.
And all of this context is the backdrop for what’s at stake in Santa Clara and some of the secret places in Silicon Valley where Nvidia is testing its next generation of technology. Buck and is colleague Andrew Bell, senior vice president of hardware engineering at Nvidia, gave the press pack a tour of a data center lab. I felt lucky to be in this tiny group of nerdy journalists. But Nvidia made it clear. It wasn’t a day we could ask questions about geopolitics or finances. That was reserved for Jensen, and he was very far away in Japan.
At the beginning of the briefing, Nvidia’s PR host told the small crowd, “We have a good action-packed day. So thanks again for making the time.”
This was a day where we could sit back and absorb the technology, which is build on such heavy-duty physical infrastructure. It reminded me of the day years ago when I visited Equinix’s secret data centers in San Jose. Outside, they looked like big unmarked warehouses. Inside, they were humming with computers, lights, plumbing and wiring.
Moving to the secret place in Silicon Valley

On the bus road over, longtime tech writer Harry McCracken, global technology writer at Fast Company tells me he has used AI to create his own word processor, spreadsheet and more. I feel behind the times, as I still have to delete stray paragraphs in my WordPress writing and publishing tool. Harry has something that deletes that stuff automatically, and my mouth drops at the miracles of AI. And I feel like I’m years behind Harry.
As our bus rolls up to the non-descript building with no signage, we unload and pile into the foyer, where there are a bunch of data center trays sitting on tables. The cables and cooling tubes are stuff into some of them — only to show that’s the old way of making them. To craft together a rack, all of the cooling tech had to be brought to bear. But over time, Nvidia removed the fans, most of the wires, and tubes. Instead, the Vera Rubin chips and their sea of memory were covered in cases that cooled them and long bars that connected them. The early racks took hours of hand work to build, but robots in AI factories can now assemble the Vera Rubin racks in record short times. The point of showing us this was to say in so many winks that the big launch for Vera Rubin looks like it’s on schedule. (Nobody dared say that, of course).
Before we go into the data center lab, we can already hear the humming. We are given earplugs and protective glasses. It’s a noisy place, and I can barely hear Bell and Buck.
There, we see the bronze racks, yellow cables and steel tubes and wires everywhere. In a big data center, you might see 10,000 racks, Bell said.
There are cables with 800 volts of power that we are warned not to touch. The metrics of the place included 3.8 megawatts of power and 1,100 gallons of water per minute flowing through the tubes of the building to cool the machines down as they run AI workloads.
“We have about 110 kilowatts per shift,” Bell said. There are redundant machines where four racks do the work of three and one is used for backup in case of failure. Each rack has about 230,000 kilowatts. There are a dozen or so partners in Taiwan who can make thousands of racks of computing per day. Even so, AI supercomputers are in short supply.
A lot of the materials are scarce, and Bell cracks a joke. Can Nvidia get its hands on enough materials to make all this stuff? “Every morning I wake up, Jensen asks me that question. We’re doing things at such a scale now that we run out of the world’s XYZ. We have people on our team dedicated to looking at materials for the future.”
Bell and Buck say “we’re good on memory now,” as Nvidia has been in front of that problem. But they have to keep their eye on long-term forecasts. It occurs to me that it’s no wonder that the Iranians targeted data centers in the early days of the Iran War.
Data center economics and ecology

Of course, you can’t help but think about the challenge of using all of the Earth’s resources to make these computing cathedrals. Will we run out of resources before we get to the answers of whether machines will someday outthink us in every way possible? I hope that the brain power going into making these machines is also going into making them sustainably.
Was I in the Belly of the Beast? Or was I in the place where we figure out all of the answers to Life, the Universe and Everything?
Bell was pointing to the different racks that represented Vera Rubin in its different stages of completion. In the earlier days, the air cooling and liquid cooling combination required a lot of fans and copper tubes to dissipate the heat from the rack, which sit on top of each other in vertical rows that are dozens of racks high.
“Infrastructure is physical. It’s really touching. You can see it and it’s what brings AI to life. Here, you’re actually able to touch it,” Buck said.

Later in the day, Gilad Shainer, senior vice president of networking at Nvidia and a veteran of Mellanox in Israel (Nvidia bought Mellanox for $6.9 billion in 2020 and used its networking fabrics to connect a sea of GPUs), passed around the NVLink fabric components that we could touch. “Here’s $1,000 for you,” he said. “And another $1,000 for you. And for you.”
Shainer and others told us that part of the reason why the economics of AI computing still works is the humble chiplet. These big giant racks with multiple Vera CPUs and Rubin GPUs, they’re possible because the designers can disaggregate the parts. Instead of one giant chip, you get different small chips glued together on small boards. Those chips are inexpensive to make compared to the big ones because the yields are better. That is, when you get 70 Blackwell chips on a 12-inch wafer, impurities will zap some of them and they won’t work. But if you can make the components smaller, then they are easier to build, less likely to fail.
And it’s by aggregating chiplets close together to feed the hungry CPUs and GPUs with data that the AI software can run at blazing fast speeds. The old Mellanox chips, now called NVLink networking chips, provide a sea of networking around the chips so they can all communicate with each other and feed data to each other so no single processor goes hungry.
So lots of chiplets can make up a monolithic processor. It gets connected at hyperspeed to another processor and they can work together. A single rack has multiple CPUs or multiple GPus, and they’re surrounded by a sea of memory chips (sorry gamers). Maybe 35 to 45 racks populate a cabinet, and many cabinets exist side by side in a row, and there are many rows in a data center. They’re all connected by tons of cables bringing electricity or bandwidth, and they’re cooled by water that is coming in through maybe eight big rubber tubes, connected to the back of the cabinets like the tubes going into the back of the head of Neo, Keanu Reeves’ character in the Matrix, (sorry Jensen).
The place with humming with fan noise, and mind you this was not a full data center. It was a lab, where Nvidia made sure that Vera Rubin is working perfectly months in advance of its launch this fall. The concept is that all those GPUs in a data center are functioning together like one GPU, a hive mind. The data centers are connected in a cloud that the likes of Amazon and Meta use to run their online stores and social networks and online games.
“This is an area where you know the CPU space is driven by a much broader cloud narrative, and now an agentic AI is really turning it on its head. I think it’s not lost on you that every AI factory is power constrained, and frankly, the most important metric is your delivered performance in a fixed watt data center,” your AI tokens per watt, Buck said.

He said, “All that comes together in your AI packaging. You may rent infrastructure like this for dollars per hour, but you generate revenue with profitable tokens now in the hundreds of dollars per hour.”
He said the revenue generated in those data centers depends on how many AI tokens you are generating. Building the infrastructure happens so fast now, as time is of the essence in the “time to first token.” How quickly does that AI data center start reasoning, thinking and outputting, Buck said.
These things are so valuable that even the old chips are useful. Buck has heard that old GPUs that are over five years old are being rented out for $5 per hour because not enough people can get their hands on new GPUs. That’s why AI factories are so valuable today.
To cram enough computing into a dense set of racks in a data center, Nvidia has to focus on “extreme co-design,” Buck said.
Nvidia designs everything now, “every part of that pod, every rack that goes into it, every tray, every solution as part of the whole end-to-end factory,” Buck said. It takes an ecosystem of hundreds of partners to make AI factories happen, Buck sai.
“And then we’re never done making the AI factory profitable. We’re always optimizing for a lot. We’re always innovating. We’re always improving performance,” Buck said. The racks are being simulated and tested for millions of hours. Nvidia has revealed benchmarking tests for Vera Rubin and the products are fast.
The upshot of this for software engineers? There are perhaps 30 million to 40 million software engineers in the world. They are generating about $3 trillion in salaries, and they’re now more productive than ever before.
“As a result, more engineers are being hired, they’re more productive, and being paid more over time,” Buck said.
Others think that AI is ironically putting software engineers out of work with all of the “vibe coding,” or AI-created software, catching on.
So what does the next generation promise? Buck said it’s likely that we could see another 10 times improvement every year or two. And that is what we’ve come to expect from the post-Moore’s-Law priests and priestesses of technology.