Back to Blog

RAM Prices Are Climbing Through 2028. Here Is What That Means If You Run AI Locally

5 min read

Memory Executives Just Moved the Goalposts

Are you planning to build or upgrade a local LLM rig in the next year? You might want to price it out now instead of waiting.

On October 1, 2026, Micron CEO Sanjay Mehrotra told investors that memory supply-demand will be much tighter in 2027 and 2028 than it is right now. Seventy-five percent of 2027 output is already sold. New clean rooms built for 2028 capacity are ramping slowly. Mehrotra gave no line of sight to supply catching up with demand.

Samsung said something close to identical a few months earlier. On July 31, 2026, the company told investors the shortage would worsen through 2027 and persist until at least 2028, pointing to long-term contracts that have locked in AI demand for years out. SK hynix executives have echoed the same timeline.

This matters for anyone running models locally because DRAM and HBM are the parts you cannot substitute your way around. Server memory and consumer memory draw from the same constrained supply. When AI labs sign multi-year contracts for HBM, consumer RAM gets squeezed right along with it.

None of this is a temporary blip tied to one product cycle. Three separate memory makers are now pointing at the same multi-year window. If your local AI setup depends on adding RAM, the price of that upgrade is set to climb before it levels off.

What You Actually Get Running Models Locally

So what does paying more for RAM actually buy you?

Two things, and they are both real. Zero per-token cost once the hardware is paid off. And your data never leaves the device.

Cloud providers charge anywhere from $0.15 to $60 per million tokens, depending on the model and the provider. That range is enormous because it spans a cheap summarization model and a frontier reasoning model doing agentic work. Local inference does not care which bucket you fall into. Once the box is built, every query is free. You are just paying electricity.

The privacy piece matters just as much for certain workloads. If you handle client data covered by GDPR or HIPAA, or you are running anything where the content cannot touch a third-party server, local inference solves a problem cloud providers cannot solve for you no matter what their terms of service promise. The model runs on your machine. Nothing gets logged, nothing gets sent anywhere, nothing becomes training data for someone else's next release.

What you give up is frontier quality. Models like Llama 4 70B or Qwen running locally handle summarization, coding boilerplate, and routine drafting well. They are not matching GPT-5 or Claude Opus on hard reasoning tasks, and pretending otherwise just sets you up to be disappointed by your own hardware.

Do the Math Before You Buy Hardware

how many queries do you run in a day?

A $2,000 to $5,000 local setup only pencils out if you are using it hard. Analyses from earlier this year put the breakeven for heavy users, meaning 50 or more queries a day or anything running agent workflows around the clock, at 3 to 15 months against equivalent cloud spend. That is a real range depending on what model you are replacing and what you paid for the rig, but it holds up across most of the comparisons.

Sporadic use does not clear that bar. If you are asking a model a handful of questions a day, or using it in bursts for one project and then leaving it idle for weeks, cloud still wins even with RAM prices climbing. You are not running enough tokens through the system to amortize the hardware, and $0.15 per million tokens on the cheap end of cloud pricing is hard to beat with equipment sitting mostly idle.

Agent workflows change the math fastest. An agent calling a model dozens of times to complete one task burns through token volume that would cost real money on a metered cloud plan. That volume is exactly what makes local hardware worth owning.

When Cloud Is Actually the Wrong Default

Is cloud actually cheaper, or does it just feel that way because nobody is tracking the bill?

That question matters more now that RAM is getting expensive on the local side too. Cloud is the right default for most people most of the time. Frontier reasoning, multimodal work, anything sporadic, anything where you are not burning serious token volume day after day, cloud wins on cost and on quality. That has not changed with the shortage.

Where cloud stops being the right default comes down to two specific situations, and both are narrow on purpose. The first is regulated data. GDPR and HIPAA workloads are not a matter of picking a provider with a good privacy policy. The data cannot leave the device, full stop, and no amount of cloud pricing math changes that constraint. If you are handling client records, medical information, or anything else with a compliance obligation attached, local inference is not a cost decision at all.

The second is a runaway token bill. Agent workflows calling a model dozens of times per task, or any setup running 50-plus queries a day, will cross the breakeven point against a $2,000 to $5,000 rig faster than most people expect. If you have never actually added up what you are spending per month on API calls, that is the number to check before assuming cloud is still the cheaper option.

Share:PostShare
RAM Prices Are Climbing Through 2028. Here Is What That Means If You Run AI Locally — PostMimic Blog