Skip to content
support@hexweb.net
Gen. Skobelev 24, Kazanlak, Bulgaria, EU

Local LLM Without a GPU: What Actually Runs on a CPU-Only Laptop

Laptop running a local LLM without a GPU on a desk at night

There is a particular thrill the first time a language model answers you from your own machine. No account. No network. No meter running. You pull the ethernet cable out just to prove it still works — and it does.

If you are anywhere near the start of this, I want to say it up front: running a local LLM without a GPU is worth doing, and it is far more accessible than the guides make it sound. Almost everything written on the subject assumes a gaming rig with a discrete card and a pile of VRAM. I do not have one. My machine is an Intel Core i7-1165G7 — four cores, eight threads, Iris Xe integrated graphics, no discrete GPU at all — with 40GB of RAM. An ordinary work laptop.

So I spent a weekend running fifteen local model configurations on it, with one identical task each time: build a working HTML/CSS/JS mortgage calculator. Then I read every line of code that came out.

What I found genuinely surprised me, and it is not what I expected to write.

Before anything else: why my numbers are slower than the ones you have read

If you have already been reading about CPU inference, you will have seen figures like 12 tokens per second on a 7B model. You are about to see 1.65 here, and I would rather explain that now than have you assume something is broken.

The i7-1165G7 is a four-core, 15-watt ultrabook chip. The processors in most published CPU benchmarks — an i7-12700, say — are twelve-core, 65-watt desktop parts with far more memory bandwidth behind them. A five- to seven-fold gap is exactly what you should expect.

I am not benchmarking the best CPU you can buy. I am benchmarking the laptop most people already own, which is where the question of running a local LLM without a GPU actually comes up.

There is a second reason, and I only found it after most of this testing was already done: this laptop had a vendor power-performance restriction quietly capping the CPU clock, even while plugged into the mains. Every number in this piece was measured before I found and removed it — more on that further down, including by how much. The comparisons between models are unaffected, because every run carried the same handicap. The absolute numbers are not: treat everything below as a floor for this hardware, not a ceiling.

The disappointment: your integrated GPU does nothing

I want to save you the evening I lost to this.

Every inference tool has a GPU offload slider. It looks like the obvious lever. I ran the same model three ways:

  • CPU only — 1.58 tok/s
  • GPU offload enabled — 1.64 tok/s
  • Reduced parallel slots — roughly 1.7 tok/s

That is not an improvement. That is noise. Across every configuration the throughput sat in the same narrow band, and my Iris Xe contributed nothing I could measure.

If your graphics are integrated, the offload slider is not going to rescue you. Skip it and spend the time elsewhere.

The biggest lever wasn’t a model setting

Somewhere in the middle of this I did something a little absurd: I asked gpt-oss-20b how to make gpt-oss-20b run faster.

It answered instantly, confidently, and in a tidy little table. Enable CUDA acceleration. Turn on FP16 mixed precision. Offload more layers to the GPU. All correct advice — for a machine that has a GPU. Mine does not, and the model had no way of knowing that. It cannot see your hardware. It can only pattern-match to the advice that gets written for gaming rigs. This is the Nemotron lesson from earlier, in a different costume: confident output is not informed output.

Animated recording of gpt-oss-20b generating its own speed-optimization answer in LM Studio, sped up 15 times from the real 3.37 tokens per second
The actual answer arriving, sped up 15× from real time so it fits in a few seconds — the genuine run took 110 seconds for 805 tokens, 3.37 tok/sec. That is what “a local LLM without a GPU” costs you in wall-clock time for one moderately long response.

So I went looking for real CPU-only settings instead, and found something more useful than any slider: I had never actually been running at full speed.

This laptop, like a lot of ultrabooks, ships with a vendor power-performance restriction that keeps the CPU clocked down even while plugged into the mains. I had no idea it was active. Removing it took a small Qwen model from crawling to 7 tokens per second, with CPU utilisation finally reaching around 90% for the first time in this whole test. On gpt-oss-20b, the same fix plus a pass of LM Studio settings took it from 3.6 to 5.5 tokens per second — roughly a 50% jump, from a setting that has nothing to do with the model at all.

Here is the diagnostic, in one sentence: if a CPU-bound job is not pushing your CPU close to 100%, the model is not your bottleneck. Check Task Manager before you touch a single slider.

Every timing elsewhere in this post — the table further down included — was measured before I found this, and, like everything else here, from single runs rather than benchmark averages. The comparisons between models still hold, because every run was equally throttled. Treat every absolute number in this piece as a floor. This machine has more in it than the numbers above show.

Once the CPU was actually working, a handful of settings were worth getting right on top of it:

  • GPU offload → 0, explicitly. LM Studio defaulted to offloading 24 layers to a GPU I do not have. Do not just ignore the slider — set it to zero, because zero is not the default.
  • CPU thread count → physical cores, not logical. Four threads, not eight. Hyperthreaded threads do not help CPU-bound matrix maths, and can add scheduling contention instead.
  • Offload KV cache to GPU → off. Same reasoning as the main offload setting.
  • KV cache quantization → on. The one setting here that is a genuine architectural win rather than noise.
  • Keep model in memory → on. Try mmap() → on. Both cut reload overhead for free.
  • Stay on AC power. Battery mode re-throttles most laptops regardless of what the power plan claims.

gpt-oss-20b itself is MXFP4 at 11.28 GB — already quantized to roughly 4 bits. If you were hoping to just quantize it harder, there is not much headroom left; it is already close to as compressed as this architecture gets.

The thing that actually predicts speed

Here is the insight that changed how I choose models, and I wish someone had told me on day one.

On a CPU, generation speed tracks the size of the file on disk. Not the parameter count. Not the model’s reputation. The file size.

Model file Speed
1.89 GB 5 tok/s
4.46 GB 1.65 tok/s
15.7 GB 0.3–0.5 tok/s

Roughly linear, for a wonderfully boring reason: producing each token means streaming the model’s weights through memory. With no GPU, you are limited by memory bandwidth. Bigger file, more bytes per token, slower output. Your CPU is barely the bottleneck — your RAM speed is.

Once that clicked, two results that had confused me suddenly made sense.

Surprise one: the same model got slower when I made it “better”

Quantization is compression — fewer bits per parameter, smaller file, a little quality lost. The standard advice is to take the highest quantization that fits, because higher is more capable.

I ran Qwen2.5-Coder-1.5B at two levels:

  • Q2_K, 752 MB → faster
  • Q8_0, 1.89 GB → measurably slower

Same model. Same 1.5 billion parameters. The only thing that changed was bytes per parameter — and the speed moved with the bytes.

This is where the usual mental model breaks. Everyone talks about model size in billions of parameters, but on a CPU that is the wrong unit entirely. A 1.5B at Q8 and a 1.5B at Q2 are not the same speed, and if you only ever look at the parameter count you will never understand why one of them crawls.

Surprise two: a 22 GB model beat a 4.5 GB one

This should be impossible under the rule I just gave you. It is the exception that proves the mechanism.

  • Qwen2.5-Coder-7B, dense, 4.46 GB → 1.65 tok/s
  • Qwen3-Coder-30B-A3B, Mixture-of-Experts, 22 GB → 3.32 tok/s

Twice the speed at five times the size.

The answer is architecture. A dense model streams every weight for every token. A Mixture-of-Experts model wakes only a fraction of the network per token — so despite being enormous on disk, it moves far fewer bytes to produce each word.

So the rule refines to: speed tracks bytes actually streamed per token. For dense models that equals file size. Mixture-of-Experts breaks the correlation deliberately, and if you are on CPU it is the single most useful thing to look for on a model card.

The catch is real, though. That 22 GB model held 21.6 GB resident and pushed system memory to 97%. It ran, and nothing else was going to run beside it. Sparse compute does not mean sparse memory — you still hold the whole model.

The part I did not expect: almost all of them got it right

I went in assuming that a local LLM without a GPU, running a small model on a slow machine, would produce broken nonsense. I was wrong, and this is the finding I most want to pass on.

Of the twelve local runs that produced a file, eleven produced a mathematically correct mortgage calculator.

Correct means all of: the annual rate converted properly with a division by 100 and by 12, the term converted to months, the real amortization formula, and — mostly — a guard for the 0% interest case. I established every verdict below by opening the generated code and reading it, not by trusting the run log. That matters, because the log was wrong more than once.

Model Verdict
Qwen2.5-Coder-14B Correct — proper conversions, zero-interest guard
Qwen2.5-Coder-7B @ Q4_K_M Correct formula — but no zero-interest guard
Qwen2.5-Coder-1.5B @ Q8_0 Correct — from a 1.9 GB model
Qwen3-Coder-30B-A3B (MoE) Correct — the fast one from the section above
gpt-oss-20b Correct, clean, well guarded
Bonsai-27B (reasoning on) Correct — strongest input validation of the lot
Bonsai-27B (reasoning off) Correct, and much terser
Gemma 4 26B-A4B (MoE, default) Correct — 3 files
Gemma 4 26B-A4B (no reasoning) Correct — 3 files
Gemma 4 E2B (no reasoning) Correct — from a tiny model
Ornith 1.5 35B-A3B (retry) Correct
Nemotron 3 Nano 4B Wrong maths — see below
Ornith 1.5 35B-A3B ×2 No file — though one of them did finish. See below
Phi-4-mini-reasoning No output at all

Look at that list again. Qwen2.5-Coder-1.5B — a 1.9 GB file — wrote a correct mortgage calculator. So did Gemma 4 E2B. These are models that would fit on a phone.

If you are starting out and quietly worried that your laptop is too weak to be useful: on a well-defined, self-contained task, it very probably is not.

The one that got it wrong, and how

Nemotron 3 Nano 4B was the single arithmetic failure, and it is a beautiful example of why you have to read the output.

const interest = parseFloat(document.getElementById('interest').value);
// ...
const r = interest/12; // monthly rate

Spot it? It divides by 12 but never by 100. Enter 5%, and the model computes a monthly rate of 0.4167 instead of 0.004167 — one hundred times too large. The payment figure that comes out is nonsense.

And there, on the same line, sits a comment confidently asserting that this is the monthly rate.

The model was not lying. It had no mechanism to check. Everything about that code looks right — sensible variable names, a tidy comment, a zero-interest guard two lines above that works perfectly. It is wrong in exactly one character’s worth of arithmetic, and no amount of reading the summary would ever have told you.

Read the code. Always read the code.

How long each one actually took

Correctness is only half of it. Here is the same set ordered by wall-clock time on the identical task, timed as each run finished.

Model Time Outcome
Ornith 1.5 35B-A3B (reasoning off) ~3.9 min Finished generating — but the file never got saved. See below
gpt-oss-20b ~7.5 min Correct, clean, 0% case handled first try
Gemma 4 26B-A4B (reasoning off) ~10.9 min Correct — the best-looking result of any local run
Nemotron 3 Nano 4B ~12.5 min Wrong maths
Qwen2.5-Coder-7B @ Q4_K_M ~16 min Correct, no zero-interest guard
Gemma 4 26B-A4B (reasoning on) ~18 min Correct, most polished
Gemma 4 E2B (reasoning off) ~19.6 min Correct — no speed benefit from being tiny
Bonsai-27B (reasoning on) ~22.4 min Correct, strongest input validation
Qwen3-Coder-30B-A3B ~23 min Correct, thousands separators
Ornith 1.5 35B-A3B (retry) ~27 min Correct
Qwen2.5-Coder-14B ~27 min Correct, plain
Phi-4-mini-reasoning >30 min ×2 Killed both times. Zero bytes returned

Two runs are missing from that table on purpose. Qwen2.5-Coder-1.5B has a recorded time, but the run log’s verdict for it disagrees with the file that is actually on my disk — and since I have been trusting files over the log all the way through this piece, I am not going to borrow the log’s stopwatch for the one model where I do not trust its judgement. The first Ornith attempt has no recorded time at all. Better a gap than a guess.

Two things jump out of that middle column.

Small does not mean fast. Gemma 4 E2B is the tiniest model in the test and took 19.6 minutes — nearly three times gpt-oss-20b, which is many times its size. Parameter count told you nothing, again. What mattered was how much of the network woke up per token, and how much of the budget went on reasoning nobody asked for.

Reasoning is the most expensive setting on the page. The same Gemma 4 26B-A4B took 18 minutes with reasoning on and 10.9 with it off, and wrote correct code both times. That is a 40% saving for no loss on this task. It is not a universal switch — suppress Bonsai’s reasoning and it breaks — but it is the first thing to try, and nobody tells you that.

If you only install two, install these

I did not expect to come out of this with favourites. I did.

gpt-oss-20b is the one I would put in front of a beginner. The fastest local run that actually produced a file, at 7.5 minutes, around 3 tokens per second on screen, correct on the first attempt, and it handled the 0% interest case without being asked — which models with far better reputations did not. It is 11.28 GB on disk and Mixture-of-Experts, so only about 3.6B parameters wake per token. That is the whole reason it is quick. It is not a toy you tolerate. It is something you would genuinely leave running.

Gemma 4 26B-A4B is the one I would keep beside it. Also around 3 tokens per second, 10.9 minutes with reasoning switched off, and it produced the best-looking output of any local model here — proper currency formatting via Intl.NumberFormat, which nothing else bothered with. Turn reasoning off and it stays correct while getting a third faster.

Between them they cover most of what you would want from a machine with no GPU: one that is quick and dependable, and one whose output you would not be embarrassed to show someone. Both are Mixture-of-Experts, which by now should not surprise you.

And the models I had the highest hopes for are the ones that let me down. Phi-4-mini-reasoning is small, well reviewed, and appears in an article currently ranking for this exact topic at 12 tokens per second. On my machine it ran past thirty minutes, twice, and returned nothing at all. Specs predict speed badly. Run the thing.

What actually stopped models wasn’t intelligence

Three of fifteen runs produced nothing at all — not bad code, no file. Phi-4-mini-reasoning is the clean case: two attempts, each pushed past thirty minutes, zero bytes returned.

The two Ornith failures are more interesting, and I have to be straight about one of them. It did not fail. It finished generating in under four minutes — the fastest local run in the whole test — and the file never got written because I stopped the run while it was still retrieving the output. That was my tooling, not the model.

Which makes Ornith the strangest result here: 3.9 minutes on one attempt and 27 on another, same model, same task. If you are choosing between local models, that kind of inconsistency deserves more of your attention than any headline speed figure, because you cannot plan around it.

So the barrier to running a local LLM without a GPU for real work is not that the models are stupid. It is that they are slow, that they are erratic, and that some of the time nothing arrives at all.

The trade-off nobody warns you about

Speed and capability pull against each other, hard, and there is no setting that fixes it.

Models fast enough to feel conversational are small enough that you will eventually catch them out. Models capable enough to trust are slow enough that you stop waiting and go and make coffee. That is simply the shape of a local LLM without a GPU, and the sooner you accept it the better your choices get.

So the useful question is not “which local model is best.” It is “what am I actually going to do with it?”

  • Interactive back-and-forth? You need several tokens per second, which caps model size, which caps ambition. Qwen2.5-Coder-1.5B at Q8 was my fast tier at around 5 tok/s — and it still got the calculator right.
  • A real task you will come back to? This is where gpt-oss-20b and Gemma 4 26B-A4B live, at around 3 tokens per second and eight to eleven minutes for a small build. Long enough to make the coffee, short enough that you have not lost the thread. It is the most useful tier on this machine by a distance.
  • Fire-and-forget — generate this, I will look in half an hour? Qwen3-Coder-30B-A3B and Bonsai-27B sit here at 22 to 23 minutes. Both correct, both more polished. Worth it only when the extra quarter of an hour costs you nothing.

I ended up keeping both, for different jobs. That is not a compromise. That is just what the constraint looks like once you stop fighting it.

Would more RAM fix it?

Less than you would hope.

More RAM raises the ceiling — my 40GB let me load a 22 GB model that a 16GB machine simply cannot. That is real, and if you are choosing a machine it matters.

But it does not change throughput, because throughput is bound by memory bandwidth and cores, not capacity. More RAM lets you run bigger models at the same disappointing speed. The upgrade that moves these numbers is a discrete GPU — a different machine, not a bigger stick of memory.

Starting a local LLM without a GPU: what I would do

Start smaller than you think. Qwen2.5-Coder-1.5B at Q8 is about 1.9 GB, runs at a usable speed, and wrote correct code in this test. Get that working in LM Studio before you download anything enormous. The thrill is identical and you get it in ten minutes rather than an hour.

Ignore parameter counts. Look at file size, and check whether it is Mixture-of-Experts. On CPU, MoE is worth actively seeking out.

Set GPU offload to zero, explicitly if your graphics are integrated — don’t just leave the slider alone. Mine defaulted to offloading 24 layers to a GPU that does not exist. Measured contribution of offloading: nothing.

Check whether your laptop is quietly throttling itself before you blame the model. Mine was, on AC power, the whole time. If a CPU-bound run is not pushing your CPU close to 100%, that is where to look first — it was worth more than every other setting on this list combined.

Decide interactive versus batch before you download. It determines everything, and the downloads themselves take tens of minutes.

Turn reasoning off before you conclude a model is slow. It took 40% off Gemma 4’s time here with no loss of correctness. Check the output afterwards, because it breaks some models outright, but try it first.

Expect roughly one run in five to give you nothing — and expect at least one of those to be your own fault, as one of mine was. It is not always the model. Budget for it either way.

And always read the code. Nemotron wrote a comment that said “monthly rate” above a line that was a hundred times wrong. That is the real lesson of the weekend, and it applies well beyond local models.

None of this is what we run for clients, to be clear — that work is ordinary WordPress AI integration, on servers that have proper hardware behind them. This weekend was about finding out where the floor is. If you want the non-technical version of where these models currently stand, we wrote a plain-English guide to the 2026 AI models earlier this year.

Now go and pull the network cable out, and watch it keep answering. It genuinely does not get old.

If nobody's looking after your site, hand it off.

Our Website as a Service model handles upkeep, security, and improvements for a predictable monthly fee — so you never have to think about it.

See how WaaS works
Svetoslav Kodzhamanov
Svetoslav Kodzhamanov

Svetoslav Kodzhamanov is the founder of Hexweb, a WordPress web design and development studio he started in 2018 as a solo venture and has since grown into a small team of designers and developers. He has personally led close to 500 WordPress projects for businesses and agencies across North America and Europe, focused on fast, conversion-driven websites built to last. He works with clients worldwide and speaks English, German, Bulgarian and Russian.

Connect on LinkedIn

One useful email, every other week.

Short, practical web and growth tips for business owners. No spam, unsubscribe anytime.