I attended part of the AI Infra Summit this week and heard to fireside chat between Ian Cutress, creator of More than Moore, and David Patterson, a computing pioneer, former UC Berkeley computer science professor, and now a distinguished engineer and fellow at Google.
Their whole conversation was a fascinating look at Patterson’s view of semiconductor chips and the AI era, as interrogated by a technical journalist who has his own doctorate. It was an interesting conversation transcribed in full below.
Patterson is one of the most influential computer architects ever. He is know for co-developing RISC (reduced instruction set computer) that is the foundation of many of today’s processors like Arm-based chips in smartphones, servers and personal computers.
He served as a professor of computer science at UC Berkeley for decades and co-wrote a book with his Stanford rival John Hennessy — Computer Architecture: A Quantitative Approach. In fact, it was so influential I saw an engineer ask Patterson to autograph some books while I was interviewing him once. He was the co-inventor of RAID (redundant array of independent disks) for storage architecture, and he helped get RISC-V, a challenger to Arm, off the ground.
Cutress is a semiconductor industry analyst, technology journalist, and commentator who is known for deep technical coverage of CPUs, GPUs, AI hardware, and chip manufacturing. He is currently the founder and chief analyst of More Than Moore, an independent semiconductor analysis and consulting firm launched in 2022. Before that, he was a senior editor at AnandTech for 11 years, and he earned a PhD from Oxford in computational chemistry.
As expected, their convo drifted from Jensen Huang’s insistence that Moore’s Law — the notion that chips will double in performance every couple of years — is dead to the current memory chip shortage that threatens to stall many industries. They also talked about the AI boom and the future of semiconductor chips and their design.
Here’s an edited transcript of their conversation.

Ian Cutress: We have an interesting history in chip design and David’s been here throughout all of it. I want to get his thoughts. No doubt if you see him around you can give him some of your thoughts as well. The first question I always ask when I bring someone on stage here is one that’s near and dear to my heart. I named my company after Gordon Moore and Moore’s Law. If I speak to people in the sphere, Jensen Huang says Moore’s Law is dead. Intel and Pat Gilsinger will say Moore’s Law is still alive. I once asked Dr. Kevin Zhang of TSMC, a good friend, and he said he doesn’t care. So where do you sit with Moore’s Law?
David Patterson: One thing I do is write textbooks. You have to get the history right. The history is, in 1965 Gordon Moore made one of the most amazing predictions of all time. He plotted a couple of points and said, “I think the number of transistors per chip is going to double every year.” About 10 years later he amended it to maybe double every two years. That was used as a guideline in the semiconductor industry to build fabs for decades.
The electrical engineers of our community–what do I do? I fulfill Moore’s Law. When you point out to them that Gordon Moore–if you plot the points, recently they’re not on that curve. If you say that then Moore’s Law no longer is accurate to what’s going on, that’s fighting words for them. But if you say that Moore’s Law no longer holds, people think that technology has come to a dead end. Even though it doesn’t make grammatical sense to say Moore’s Law is slowing down, those aren’t fighting words. But unquestionably, Jensen is right based on the facts that Gordon Moore originally produced.
Technology is still getting better. It’s not getting as better as quickly as it used to. Different pieces of it are barely improving at all. Some of them continue to improve. That’s my version of the answer. The argument is about packaging. Gordon Moore, when he revised in 1975, he didn’t say anything at all about packaging. That was the great thing in computer design. For decades you could basically ignore packaging and concentrate on what’s on the chip. These days packaging is maybe the major part of what you’re doing.
Cutress: In terms of bottlenecks the industry is experiencing right now, if you talk to most people in the room, they’ll say things like memory. Memory is a limitation in terms of density, in terms of pricing, as is the macroeconomics right now. Others will say it’s the network that allows you to have scale up and scale out domains. Others will say power and power availability. Where do you sit in terms of where the bottlenecks are right now? What do you focus on?
Patterson: I agree about memory as the bottleneck. There’s two pieces of memory. There’s the bandwidth of memory and the capacity of memory. Over the decades, DRAM did this amazing 4X every three years. Now that’s completely gone. DRAM capacity is barely improving. Thanks to this packaging technology of high-bandwidth memory, it delivers phenomenal memory bandwidth. Bandwidth has made great strides. Nevertheless, in computing today a lot of the challenges are not something that memory bandwidth wins.
I would make it memory bandwidth, and then power. We read every day about concerns around building data centers. Why do we have to build a lot of data centers? It’s because the newer chips use a lot more power than the older chips, because of scaling. And then the interconnect, to me, is in third place. I may be biased, because working for Google, we have optical interconnect circuit switching, which is very cool technology. It doesn’t feel as big a limitation as the others, in my view.
Cutress: Are you saying that as an engineer, or as a corporate spokesperson for Google?
Patterson: I think I still have tenure. I was a faculty member at Berkeley for 40 years. I’m still an engineer. I don’t know what the official corporate line is.
Cutress: In that regard, would it be your advice perhaps to people looking to solve some of these problems to focus on memory as one of the key points?
Patterson: I’m wondering whether memory is going to get a big focus of what’s going on in the next five years.
Cutress: In terms of the landscape where we are right now, is there anything you thought would have been solved by now?

Patterson: Kind of? I’m stunned that we’re talking about floating point numbers. That that’s a major focus over the last decade, floating point numbers. In the history of computing, integers–there’s no question how arithmetic works. But floating point is an approximation. One of my colleagues at Berkeley, Bill Kahan, saw the opportunity when microprocessors could start doing floating point, that they could standardize floating point arithmetic. Every company had their own floating point arithmetic. That made it very hard to port scientific programs. He seized the moment and it led to the IEEE floating point standard in 1985. Virtually every company dropped what they were doing and adopted the IEEE floating point standard.
I thought that was it. We had a section in the textbook about floating point arithmetic. I never had to change it. It was very easy. And then 10 years ago people started playing around with floating point numbers, because machine learning, by its statistical nature, didn’t need precision. If you look at a floating point number, if you remember your textbooks, the exponent is the little piece and the fraction is the big piece. Well, they didn’t need that. In fact, at Google, which was one of the early movers here–they were doing it in software, which used too much network bandwidth, so they just whacked off the bottom half of the floating point. Just threw it away. If you talk to a numerical analyst and say that what you did to floating point is just truncate–that’s like nails on the blackboard. They hate this idea.

We went from–real computers did 64-bit floating point. Probably machine learning can do 32-bit. We can just cut off the bottom half to 16, eight, four bits. People are talking about even smaller things than that. I never expected to see innovation in floating point arithmetic again in my lifetime. That was the big surprise for me.
Cutress: I have a friend, a guy named Felix le Clair. He tracks the number of different FP8 formats out there. I think it’s up to 18 now.
Patterson: We need another IEEE floating point standard for little tiny floating point, because not all the companies are doing it the same way.
Cutress: Is this truly innovation, or is this just picking and choosing what’s best?
Patterson: No, it’s real innovation. When you’re going against conventional wisdom–we’re making this all up right now, right? I’d say the history of the microprocessor industry, basically what it did is it looked to the mainframe industry. These older companies. It just imitated the ideas there. Now we’re at this frontier where we’re more invested in special purpose computing than general purpose computing. Nobody knows what the right thing to do is. Especially with this extraordinary rate of the change in the models themselves and building customized hardware. It’s exciting for engineers and scary for leaders of companies, who have to place the investment. We had to add a new chapter to our textbook to cover what’s going on. This is a very exciting time in computer architecture.
Cutress: On that front, in preparation for this I saw a few of your older talks and some of what you’ve presented over the years, both to industry and to students.
Patterson: Let’s see if I’m consistent over time.
Cutress: One of the consistent narratives has always been that we teach students for the general case and then specialize. But right now we’re going into a market of machine learning, which is highly specialized. It doesn’t so much care about the general case. Should we change how we approach this?
Patterson: When my friend and co-author John Hennessy got the Turing Award, about a decade ago, he got a chance to give a lecture and write a paper. They accept the paper no matter what you write. What we said was, it looked to us like–well, what happened in computer architecture history was, in the good old days you could just keep slamming more transistors on and the power didn’t change. That was called Dennard scaling. Bob Dennard noticed that because you dropped the threshold voltage every time you increased the transistors, the power stayed the same. The first two editions of our textbook didn’t even talk about power. It wasn’t something you mentioned, because it was always 20 or 30 watts. That ended in 2005 or so. We still needed more performance. The hardware designers said, “We’re going to multi-core to deliver more performance,” which was a huge upset for the software community. They had to deal with that. They didn’t want to, but there was no other option.
John and I thought, a decade ago, that we’d have to go to domain-specific architectures. With the end of Dennard scaling and the slowing of Moore’s Law, general-purpose computing was only going to improve very slowly after that, and people were still going to want that. If you only have to do one thing, you can throw on a lot of features and throw resources at them. We can make that go faster. But that was the challenge. We predicted that this was going to be a golden age of computer architecture, because you were going to need architects to use those transistors completely differently to accelerate the specific domains. That’s an exciting, but problematic approach. I think students still need to learn how CPUs work, but today, most of the investment is in these domain-specific architectures.
Cutress: We had a couple of conversations just before this talk. I brought up the fact that we do things a bit differently in Europe in our studies. We specialize at the age of 18. We don’t have GE classes.

Patterson: That just seems really dangerous. What would you specialize in now, if you’re 18 years old and you look at the world? What would you specialize in? One of the nice things about teaching at Berkeley is you’re right near Silicon Valley. We would have real engineers come in and give guest lectures in our courses and offer advice. I remember one of them said – this was 30 years ago – be as general as you can. His advice was, when you push a key down on the keyboard, know everything that’s going on in the computer, the physics of what’s going on in the key and the connection and the electronics. The operating system is getting invoked, the interrupts. Learn as much as possible. He said that 30 or 40 years ago.
I just chaired a thing with two rock stars of AI hardware, Bill Dally of Nvidia and Norman Jouppi of Google. I asked them what advice they had. They said to be as general as possible, so you can innovate across all these different interfaces. Don’t keep yourself narrow and rely on what the memory guys say is possible. That was their exact advice. It’s striking to me that in these ages, being as broad as possible is even more important. So, good luck.
Cutress: We’ll get to it in a second, but with the concept of agentic AI now, the ability to request tool use through CPUs without moving back into a more general-purpose orchestration–can we put these two things being specialized together so they can also be generalized?
Patterson: This is another example of a very exciting new area of agentic systems. We don’t know, off the shelf, what the right thing to do is. If you’re invoking a compiler, I don’t understand why you have to have a very close interaction between the CPU and the AI accelerator, GPU or TPU. If it’s taking minutes to compile, how important is that? But there may be other things where the CPU has to be much closer. This will reinvigorate CPU design. If the main use of CPUs is going to be in agentic systems, which is possibly what the future looks like, how will we do CPUs differently for these agentic workloads? Should we redesign our software stack? These tools have been designed for humans to use. Success or failure was the user interface and user experience. With AI engines as your users, how would you do things differently? Very interesting, exciting times going forward, where good ideas could lead to big changes. The industry would welcome big ideas.
Cutress: When I speak to the big players in the industry, whether it’s Jensen or Lisa, they’re not selling you a chip anymore. They’re selling a system, an ecosystem, going into gigawatts of data centers. You already mentioned TPU on the Google side. Jon Masters, over at Google, he thinks the best idea is to keep the chips boring and keep the industry interesting. But we’re at a stage now where engineers like you and I focus a lot on the chip bottom up. We find that interesting, whereas sometimes the real innovation is coming from the top down, from the networking side. Is it better to have boring chips?
Patterson: I don’t see how you could call these chips boring. These nodes are–let’s talk about Nvidia prices. I think they’re going up $10,000 a generation. It’s $50,000 for one node, and kilowatts of power after that. I’d hardly call that boring. It’s an end to end system thing. When I joined Google, the people there realized that there’s two pieces of the problem. There’s the inference piece and the training piece. For inference, they were trying to be economical but fast. Maybe single-chip systems. But for training you wanted a supercomputer. Google has been pushing those supercomputers, which started off 10 years ago at 256 nodes. Now they’re already up to almost 1000 nodes. But the idea that training needs supercomputers, supercomputer interconnect and all that, is still true today. That’s been a wise focus that others are following.
Cutress: Today, can you actually make a chip without thinking about the rack scale solution?
Patterson: It’s all the pieces together. One of the famous guidelines in computer architecture is called Amdahl’s law. Gene Amdahl came up with this about 60 years ago. It’s really the law of diminishing returns. Here’s your pie. If you take this piece, and that’s the only piece you’re working on, you can make that incredibly fast, but the rest of the pie is going to limit how fast you go. Some architects call it Amdahl’s cruel law, because they find this piece that they’re very excited about, and they make it go 10 times faster. Then the rest of it doesn’t go very well. They keep stumbling over it, time and again. Amdahl’s Law again. You focus on this piece and it goes incredibly fast, but from a big system perspective it doesn’t make that much difference, because you didn’t improve the rest of it. That’s always the challenge for big systems. All the pieces have to fit together, which is why we come back to the importance of general knowledge, so you can have a well-balanced system.

Cutress: This is an open-ended question, because you could answer it in 50 million ways. Where do you stand on chiplets?
Patterson: I thought it was a very exciting idea. The way semiconductor manufacturing works, there are flaws. When you build a wafer it’s not 100% yield. Some will have these flaws in them. The idea of chiplets originally was smaller chips, possibly standardized, that you could construct systems out of. That was the original vision. But now chiplets are this full reticle. They call that a chiplet. I don’t know why they call that a chiplet, except I think they’re using some of the interconnect technology. It’s the biggest chip you can possibly make. There’s a paper I’m a co-author of, which I think is the first use of the word chiplet. I don’t see how that can be called a chiplet if it’s full reticle size.
One of these exciting technologies was optical interconnect. People have talked about it for a long time. We use it in data centers, at data center scale, all the time. If we get chip to chip optical interconnect to be economical and efficient, possibly we’ll see chiplets be small, connecting a lot of them together. But right now–whether or not you use the word chiplet, it doesn’t really matter for the systems being deployed right now.
Cutress: I remember IBM having something like oil-filled 80 chiplets in a big package, back in the ‘80s.
Patterson: Yeah, the mainframe people had exotic things like that. So far, the possibility of standardized chiplets, none of that is happening so far.
Cutress: Most of the time we design chips these days – or at least we used to – on this three- or five-year cadence, with pathfinding and then generations. N+1, N+2. But right now we’re in this chasing mode where we keep trying to chase whatever algorithm–today it’s transformers. Now we’re on agentic AI. How can we plan for that future effectively by also being domain-specific?
Patterson: That is the challenge going forward. Companies that have both people who are pushing the state of the art in machine learning and people designing hardware have a potential advantage if those groups interact. They can tell where the puck’s going to go. One of the reasons this domain-specific architecture for machine learning is so exciting is the change in the software stack on top. For CPUs, for decades you were running these legacy programs written in C++. Millions of lines of code. Millions of programs out there. In this AI world, you’re learning from data. The code itself is small. Transformer programs aren’t that big. But they’re looking at billions of pieces of data. The intelligence comes from the data, not from the code.
Because the code is small, it’s not that hard to change. They’re willing to change the model’s code to adapt to your hardware. This is an architect’s dream. Before, you’d say, “I have this great idea! All we have to do is change all the software in the world.” Good luck. It would never happen. But because suddenly this can adapt to your hardware, it can also change very rapidly. It can come up with new ideas every few weeks. It still takes years from the idea before you can have chips in volume, so that’s a scary thing for computer architects. They’ll come up with this great chip and miss some key feature, so they can’t run the most important new software.
Transformers have been around for a long time now. There’s innovation in transformers all the time. I can talk about Google. So far at Google, fingers crossed, it seems to have worked out that the new TPUs run the latest models. But the potential is there for the whole industry, for something to happen that will work well on one guy’s chips but not on all the rest of them. That’s the game we’re playing.
Cutress: Are you saying that the front end of the chip no longer matters, because we can just design the software to it? Or should we still have some standards there?
Patterson: No, it’s the fact that they’re willing to–you can talk to the people developing the models and tell them, “Would you please do this?” And they’ll talk to you. In the old days, software was king. They wouldn’t even talk to you. You could try to bribe them. Please don’t do random memory accesses. You could influence them. But they’re inventing the future here. They’re in stiff competition all over the world. The seven authors of the transformer paper are these demigods in the industry. They would love to create the next transformer. They’re driving hard to do that. The industry, or the marketplace, needs this stuff to get better. There are huge opportunities. They’re trying to innovate as rapidly as they can.

Cutress: We recently had an announcement from OpenAI that they’re developing their own chip. It went from zero RTL to running models in nine months on a chip they taped out. One, that’s really impressive. But two, does that change anything?
Patterson: If the models are changing rapidly, the shorter we can make the time between idea and chips in data centers, that’s a positive. When I saw that–a version of that was, “It took them nine months to design a chip.” That was the press version. A lot of those people came from Google. They were there a year and a half before the clock started. So I think they were probably doing something before that. But still, the RTL to tapeout in nine months is pretty good.
It was great talking to Bill Dally of Nvidia. I didn’t have his perspective. He said that one advantage they have is they have one customer. They have one model they have to run. They don’t have to run models from lots of customers. That’s a potential advantage they have. But it will be interesting to see how that turns out. It was certainly, at the Hot Chips conference a week or so ago, one of the highlights. It looks like they did a good job. I’m anxious to see how it works out.
Cutress: I’m interviewing Richard Ho, their VP of hardware, on Friday. What should I ask him?
Patterson: Oh, boy. Aren’t they using Cerebras 2? How does Jalapeno compare to Cerebras in terms of latency and tokens per second?
Cutress: A lot of your recent work has been about power consumption of data centers. The ability to put the hardware in the right location versus the wrong location. I remember at a previous event, going by the Supermicro booth. It said, “Let’s make 30% of the world’s data centers liquid-cooled by 2030.” When I speak to players in the space, they’re either liquid-cooled or air-cooled. What exactly is right first move here? You guys have to deliver tokens in bulk. Your economics are a bit different to a lot of other people’s.
Patterson: I got involved in the carbon footprint of AI. One of the acronyms we use is the four Ms. You break down the carbon footprint into these four factors. The first M is the model. Get the most efficient models you can. That’s going to have a gigantic impact on the training. The second is the machine. We wrote a paper where we did extremely thorough evaluations of the carbon footprint of the making of a chip and the running of these TPUs. Cradle to grave, mining the materials to manufacturing and all that stuff. We weren’t sure what the bottom line would be, but fortunately the bottom line, the most efficient chip from a carbon footprint perspective is the latest one. It didn’t necessarily have to be that way, but fortunately it’s the latest one. If you get the latest GPU or TPU, it’s probably the most carbon efficient.
The next one was a stretch. We call it mechanization, but it’s really the data centers. Run an efficient data center. Data centers have this widely used metric called power usage effectiveness. If one is the power you need to run the servers, how much more do you need for the power distributions and cooling and all that stuff? Data centers now come down to 1.1, which is about 10% overhead. The model, the machine, and the data center. And then finally it’s the map. Where is it located? How clean is the energy you can get? It’s not only that. That’s the primary factor, the local utility. But companies can, in addition, get carbon-free energy from other sources. Those are the four factors.
Companies that really care about the carbon footprint have to work on all four of those pieces. If you can locate a data center in a place that has lots of carbon-free energy, like near a hydraulic dam or a solar energy farm, that’s going to have a much smaller footprint. What’s exciting about the cloud is that as a customer, you can do this anywhere. Especially for training. It can be anywhere in the world. You can say, “I’m gonna give my business to this data center. It’s running a very efficient model. It’s using efficient machines. It’s in an efficient data center. It has very clean energy.” You can have a lot of impact on that. If customers vote with their feet, then they’ll inspire the cloud companies to try and deliver on that.
Cutress: I recently did some coverage about a new neo-cloud in Japan. They claim to have six megawatts of energy. But their main target is deploying tokens near where they’re needed. They’re going into already existing infrastructure in order to do that. Is that your strategy?

Patterson: The training can be anywhere, but the inference part, the serving, is latency-sensitive. That has to be distributed all over. I would say, in terms of economically–they could make economic sense, ignoring the carbon footprint, going into existing data centers that may be underused and running them more efficiently. It’s not likely that the energy supply in an urban area is very clean energy. If it is, that would be great, but often urban centers don’t have the cleanest energy sources.
Cutress: If we start putting data centers in space, is that clean energy?
Patterson: Yeah. There’s people at Google, people at Musk’s company, talking about data centers in space. I don’t think that’s a right around the corner thing. The argument for space is the sun’s there 24 hours a day. You have free solar energy. It’s also more efficient without the atmosphere. You can discharge heat in space. But what’s the cost to put data center material in a rocket and take it up there? How long does it stay up there? Those are pretty big challenges. I don’t think this is going to happen within a couple of years.
Cutress: Personally I think it’s possible, but it’s the maintenance. If your hardware fails once in five years–
Patterson: You’d have to factor in–in the history of computing, fault-tolerant computing started with the space program. The idea that things break and keep running, that’s fundamental to space technology. That’s the approach you would have to take.
Cutress: Normally, in the enterprise market, we talk about turnover of hardware on a three-year time scale. But now we’re seeing companies amortize their GPUs and XPUs over five, six, and seven years, just because the demand for compute is so high. Do we have to change chip design to deal with that?
Patterson: That was one of the things, when you told me your question–I had an opinion on it and I decided to go check. This is an example where people like–given all the publicity about the new GPUs and how much faster they are, it must be the case that people don’t use the old GPUs. That seems logical. But that’s actually not what the facts are. I went and checked. First of all, you can figure this out yourself. If you go to AWS, Amazon still rents the first GPU that had support for machine learning, the Volta, from 2017. Similarly, at Google, the oldest GPU is the A100, which is six years old.
I looked inside at Google and I was amazed. The TPU V2 and V3, which are from eight and 10 years ago, when used for inference, they’re still being used this many years later. The inference chips get spread all over the world. And then TPU V4, which is the training chip, as far as I can tell, all of those training ones are still in use. Those are six years old. There’s such a desperate need for cycles that people are absolutely keeping them beyond the amortization. I think we’ve amortized our accelerators six and our CPUs eight years. They’re going to last a lot longer. It’s this amazing demand and appetite. People can’t get enough of the latest ones, or its cost-performance is so high, so these other ones are still attractive. And once they’re fully amortized, the way accounting works–we’re going to see these things last a long time. That would be one of the challenges in space. Can you really keep something flying up there for eight or 10 years?
Cutress: Even though you just argued that it’s best to have the best chip–
Patterson: If you have a choice, yeah.
Cutress: But if you have hardware, keep running it.
Patterson: Right. As a consumer, you have a lot of choices. I don’t know what the cost is, but from an environmental perspective, the latest chip is the best one. I assume the pricing will reflect that. But surprisingly, it’s not the case that six-year-old GPUs are so much worse than the current ones that we should rip them all out and replace them. That’s not what’s going on, which is surprising.

Since you asked that question, the head of Google Cloud, Thomas Kurian, just yesterday or so, said they get their money back on GPUs in two years. They make the investment in a GPU server and they can rent it out and it fully pays for itself in two years. It lasts six years. For TPUs it pays for itself in one year. That’s a remarkable money-making situation. Six-year-old TPUs are still pretty exciting technology. It’s remarkable.
Cutress: So the fact that I like to collect weird hardware means I’m going to have a wait a bit longer to find a TPU on eBay.
Patterson: I thought we were going to start throwing away TPU supercomputers. I could figure out good uses for one of those. But that’s not true. People are clinging on to them. It’s fighting words to say you’re going to take them away.
Cutress: I wanted to get your thoughts on the whole memory supercycle. Right now people are demanding HBM in droves, LPDDR in droves. We just saw a report recently about the news in 8-high and 12-high memory stacks. Perhaps it might be better to use 4-high memory stacks. There are other new memory technologies promising to fill in some of the gaps. But everyone I speak to wants to just buy more of the same, more of what’s the latest on the market. Where do you sit with all of this?
Patterson: Is there a change going on? When you look backwards and see it, there’s this chance that we’re entering a memory-centric era. Why are they cutting down HBM stacks? That’s not a good idea, because they need the capacity and the bandwidth. The reason they’re cutting it down is there’s just not enough to go around. They just can’t make enough of them to supply the demand. It takes years to build a fab. This is not going to be solved in the next six months. Then the memory vendors historically have been in these cycles where they overbuild, prices drop, they lose money, and then there’s an upswing. They don’t love a cyclical business. They’d like it to be nice and smooth.
It’s not clear to me that these huge price jumps in the last couple of years are going to disappear momentarily. It could be that memory is becoming a much larger fraction of the so-called barrel of materials. Most of that money is going to memory, and memory bandwidth can be a bottleneck. We may be entering an era where memory-oriented architectures are going to be where you start.

In terms of memory technologies, I’m involved with one that’s called high bandwidth flash. Surprisingly, there’s a way to build flash where it’s not only huge capacity, but instead of having all of those layers being designed to be as cheap as possible, you can transfer data in parallel. It can be high bandwidth and high capacity. Flash is terrible at writing, so you have to find workloads that are mostly read, but that can be another option. SRAM has plateaued. DRAM is barely improving. HBM is this new packaging for DRAM. HBF may be a possible new piece of memory technology that’s practical.
Cutress: We don’t see new memory technology breaking through that often. Things like resistive memory has found a place in embedded, but not at scale.
Patterson: In my career–when I started in computing we had magnetic cores that we successfully shifted to DRAM main memory, and we had flash memory. Those are the two. Holographic memory–
Cutress: I have to ask you. Intel and Micron’s Optane technology. If that had come about today, would it still survive?
Patterson: Well, there were two problems. As a potential technology it would be attractive right now. The way I understand it, Intel rolled it out with restrictions. Only some of their chips supported it. That would still be difficult. It wasn’t universal. To work, you need it to come from multiple vendors, so you don’t get locked in, and you need the software to be able to handle it. You need both of those pieces. That’s why it’s a very sad issue for memory technologies in our field. They almost never work.
Cutress: Is that still true with custom HBM?
Patterson: Both HBM and HBF are not new memory cells. It’s not like resistive RAM or magnetic RAM. They’re standard cells. They’re packaged in a new way. That makes it more stable. The idea of putting processing in your memory is a very promising direction going forward for many of the AI workloads. At least we’ll see.
Cutress: My point with custom HBM is that it no longer becomes a commodity. The custom is vendor-specific.
Patterson: Again, that’s always the memory challenge. Do customers feel willing to get locked in to a single supplier of a memory technology? It’s to their obvious advantage if there are multiple vendors. That’s going to be one of the challenges going forward.
Cutress: A couple of quick fire questions. How are you using AI today personally?
Patterson: First of all, I use multiple ones, since they lie. I try to figure out how I can do things using AI. I don’t write code. All my friends who are programmers are stunned at what these coding agents are doing. I can’t believe any programmer is not playing with these things.
Cutress: We’ve seen a few industry pivot points over the last few years, whether you take into account transformers or the Deepseek moment and so on. Will we always continue to see pivots?
Patterson: I don’t see how that can’t happen. This technology has tremendous upsides and tremendous downsides. The upsides–the science advances, like you mentioned in the video. The guy who was supposed to be here was Jeff Dean. He’s a hero at Google. He left to do a company discovery loop to accelerate the delivery of science. Demis Hassabis at DeepMind, he’s got a company working on that. Advances in science–I’d be surprised if, in the next five or 10 years, we don’t see giant advances in science from AI. That’s one of the great upsides.
Cutress: Do you expect to be answering the same questions in five years, if we were to do this again?
Patterson: No, you’ll come up with much better questions.