Running AI Models Becomes a Memory Challenge

Discussions about AI infrastructure costs often center on Nvidia and GPUs, but memory is becoming a critical piece of the puzzle. With hyperscalers investing billions in new data centers, DRAM chip prices have surged roughly sevenfold in the past year.
Simultaneously, a growing discipline focuses on orchestrating this memory to ensure the right data reaches the right AI agent at the right time. Companies that excel at this can execute the same queries using fewer tokens—a difference that can determine survival in a competitive market.
Semiconductor analyst Dan O’Laughlin offers a compelling perspective on the importance of memory chips in his Substack, featuring a conversation with Val Bercovici, Chief AI Officer at Weka. Both experts come from a semiconductor background, so their discussion leans toward the hardware, though the implications for AI software are equally significant.
One passage stood out, where Bercovici comments on the increasing complexity of Anthropic's prompt-caching documentation:
The clue is on Anthropic's prompt caching pricing page. It began as a simple page six or seven months ago, especially around the launch of Claude Code, essentially saying, "Use caching, it's cheaper." Now, it reads like an encyclopedia of advice on how many cache writes to pre-purchase. You see common industry tiers like 5-minute and 1-hour windows—but nothing longer. That's a telling detail. Then, of course, there are various arbitrage opportunities around cache read pricing based on how many cache writes you've bought in advance.
The core question is how long Claude retains your prompt in cached memory. You can pay for a 5-minute window or a more expensive hour-long window. Accessing data still in the cache is far cheaper, so effective management can lead to substantial savings. However, there's a catch: each new piece of data added to a query might displace something else from the cache.
While the details are complex, the conclusion is straightforward: Memory management within AI models will be a major factor in the future of AI. Companies that master it will gain a significant competitive edge.
There's considerable room for progress in this emerging field. Last October, I wrote about a startup called TensorMesh, which is working on a layer of the stack known as cache optimization.
Techcrunch eventTechCrunch Founder Summit 2026: Tickets Live
On June 23 in Boston, more than 1,100 founders come together at TechCrunch Founder Summit 2026 for a full day focused on growth, execution, and real-world scaling. Learn from founders and investors who have shaped the industry. Connect with peers navigating similar growth stages. Walk away with tactics you can apply immediately
Save up to $300 on your pass or save up to 30% with group tickets for teams of four or more.
TechCrunch Founder Summit: Tickets Live
On June 23 in Boston, more than 1,100 founders come together at TechCrunch Founder Summit 2026 for a full day focused on growth, execution, and real-world scaling. Learn from founders and investors who have shaped the industry. Connect with peers navigating similar growth stages. Walk away with tactics you can apply immediately
Save up to $300 on your pass or save up to 30% with group tickets for teams of four or more.
Boston, MA|June 23, 2026REGISTER NOWOpportunities also exist elsewhere in the stack. Further down, for instance, lies the question of how data centers utilize their different types of memory. (The interview includes an insightful discussion on when to use DRAM chips versus HBM, though it gets quite technical.) Higher up the stack, end-users are learning to structure their model swarms to leverage shared caches effectively.
As companies improve at memory orchestration, they will consume fewer tokens, making inference cheaper. At the same time, models are becoming more efficient at processing each token, driving costs down even further. As server expenses decline, many applications currently deemed unviable will begin to approach profitability.
Related article
Former Infosys Chief’s AI Startup Secures Another $53M
Hang Ten Systems, an AI startup established by former Infosys CEO Vishal Sikka just four months ago, has secured an additional $53 million in seed funding. This latest investment round was finalized merely five weeks after the initial $32 million see
Crypto exchange OKX aims to empower AI agents to hire and pay each other
As AI agents start serving individuals and collaborating with each other, they require mechanisms to locate tasks, compensate for services, and establish credibility. Crypto exchange OKX anticipates this future is arriving sooner than anticipated, in
Thiel-backed startup claims AI can judge journalism, despite risks to whistleblowers
Following his role in the lawsuit that led to Gawker’s bankruptcy, Aron D’Souza identified a critical flaw in the American media landscape: individuals harmed by coverage lacked effective means to respond.His answer is technology. D’Souza’s new ventu
Related Special Topic Recommendations
Comments (1)
0/500

Discussions about AI infrastructure costs often center on Nvidia and GPUs, but memory is becoming a critical piece of the puzzle. With hyperscalers investing billions in new data centers, DRAM chip prices have surged roughly sevenfold in the past year.
Simultaneously, a growing discipline focuses on orchestrating this memory to ensure the right data reaches the right AI agent at the right time. Companies that excel at this can execute the same queries using fewer tokens—a difference that can determine survival in a competitive market.
Semiconductor analyst Dan O’Laughlin offers a compelling perspective on the importance of memory chips in his Substack, featuring a conversation with Val Bercovici, Chief AI Officer at Weka. Both experts come from a semiconductor background, so their discussion leans toward the hardware, though the implications for AI software are equally significant.
One passage stood out, where Bercovici comments on the increasing complexity of Anthropic's prompt-caching documentation:
The clue is on Anthropic's prompt caching pricing page. It began as a simple page six or seven months ago, especially around the launch of Claude Code, essentially saying, "Use caching, it's cheaper." Now, it reads like an encyclopedia of advice on how many cache writes to pre-purchase. You see common industry tiers like 5-minute and 1-hour windows—but nothing longer. That's a telling detail. Then, of course, there are various arbitrage opportunities around cache read pricing based on how many cache writes you've bought in advance.
The core question is how long Claude retains your prompt in cached memory. You can pay for a 5-minute window or a more expensive hour-long window. Accessing data still in the cache is far cheaper, so effective management can lead to substantial savings. However, there's a catch: each new piece of data added to a query might displace something else from the cache.
While the details are complex, the conclusion is straightforward: Memory management within AI models will be a major factor in the future of AI. Companies that master it will gain a significant competitive edge.
There's considerable room for progress in this emerging field. Last October, I wrote about a startup called TensorMesh, which is working on a layer of the stack known as cache optimization.
Techcrunch eventTechCrunch Founder Summit 2026: Tickets Live
On June 23 in Boston, more than 1,100 founders come together at TechCrunch Founder Summit 2026 for a full day focused on growth, execution, and real-world scaling. Learn from founders and investors who have shaped the industry. Connect with peers navigating similar growth stages. Walk away with tactics you can apply immediately
Save up to $300 on your pass or save up to 30% with group tickets for teams of four or more.
TechCrunch Founder Summit: Tickets Live
On June 23 in Boston, more than 1,100 founders come together at TechCrunch Founder Summit 2026 for a full day focused on growth, execution, and real-world scaling. Learn from founders and investors who have shaped the industry. Connect with peers navigating similar growth stages. Walk away with tactics you can apply immediately
Save up to $300 on your pass or save up to 30% with group tickets for teams of four or more.
Boston, MA|June 23, 2026REGISTER NOWOpportunities also exist elsewhere in the stack. Further down, for instance, lies the question of how data centers utilize their different types of memory. (The interview includes an insightful discussion on when to use DRAM chips versus HBM, though it gets quite technical.) Higher up the stack, end-users are learning to structure their model swarms to leverage shared caches effectively.
As companies improve at memory orchestration, they will consume fewer tokens, making inference cheaper. At the same time, models are becoming more efficient at processing each token, driving costs down even further. As server expenses decline, many applications currently deemed unviable will begin to approach profitability.
Former Infosys Chief’s AI Startup Secures Another $53M
Hang Ten Systems, an AI startup established by former Infosys CEO Vishal Sikka just four months ago, has secured an additional $53 million in seed funding. This latest investment round was finalized merely five weeks after the initial $32 million see
Crypto exchange OKX aims to empower AI agents to hire and pay each other
As AI agents start serving individuals and collaborating with each other, they require mechanisms to locate tasks, compensate for services, and establish credibility. Crypto exchange OKX anticipates this future is arriving sooner than anticipated, in
Thiel-backed startup claims AI can judge journalism, despite risks to whistleblowers
Following his role in the lawsuit that led to Gawker’s bankruptcy, Aron D’Souza identified a critical flaw in the American media landscape: individuals harmed by coverage lacked effective means to respond.His answer is technology. D’Souza’s new ventu





Home






