Local LLM Without a GPU: What Actually Runs on a CPU-Only Laptop
There is a particular thrill the first time a language model answers you from your own machine. No account. No network. No meter running. You pull the ethernet cable out just to prove it still works — and it does.
If you are anywhere near the start of this, I want to say it up front: running a local LLM without a GPU is worth doing, and it is far more accessible than the guides make it sound. Almost everything written on the subject assumes a gaming rig with a discrete card and a pile of VRAM. I do not have one. My machine is an Intel Core i7-1165G7 — four cores, eight threads, Iris Xe integrated graphics, no discrete GPU at all — with 40GB of RAM. An ordinary work laptop.
So I spent a weekend running fifteen local model configurations on it, with one identical task each time: build a working HTML/CSS/JS mortgage calculator. Then I read every line of code that came out.
What I found genuinely surprised me, and it is not what I expected to write.
Before anything else: why my numbers are slower than the ones you have read
If you have already been reading about CPU inference, you will have seen figures like 12 tokens per second on a 7B model. You are about to see 1.65 here, and I would rather explain that now than have you assume something is broken.
The i7-1165G7 is a four-core, 15-watt ultrabook chip. The processors in most published CPU benchmarks — an i7-12700, say — are twelve-core, 65-watt desktop parts with far more memory bandwidth behind them. A five- to seven-fold gap is exactly what you should expect.
I am not benchmarking the best CPU you can buy. I am benchmarking the laptop most people already own, which is where the question of running a local LLM without a GPU actually comes up.
There is a second reason, and I only found it after most of this testing was already done: this laptop had a vendor power-performance restriction quietly capping the CPU clock, even while plugged into the mains. Every number in this piece was measured before I found and removed it — more on that further down, including by how much. The comparisons between models are unaffected, because every run carried the same handicap. The absolute numbers are not: treat everything below as a floor for this hardware, not a ceiling.
The disappointment: your integrated GPU does nothing
I want to save you the evening I lost to this.
Every inference tool has a GPU offload slider. It looks like the obvious lever. I ran the same model three ways:
- CPU only — 1.58 tok/s
- GPU offload enabled — 1.64 tok/s
- Reduced parallel slots — roughly 1.7 tok/s
That is not an improvement. That is noise. Across every configuration the throughput sat in the same narrow band, and my Iris Xe contributed nothing I could measure.
If your graphics are integrated, the offload slider is not going to rescue you. Skip it and spend the time elsewhere.
The biggest lever wasn’t a model setting
Somewhere in the middle of this I did something a little absurd: I asked gpt-oss-20b how to make gpt-oss-20b run faster.
It answered instantly, confidently, and in a tidy little table. Enable CUDA acceleration. Turn on FP16 mixed precision. Offload more layers to the GPU. All correct advice — for a machine that has a GPU. Mine does not, and the model had no way of knowing that. It cannot see your hardware. It can only pattern-match to the advice that gets written for gaming rigs. This is the Nemotron lesson from earlier, in a different costume: confident output is not informed output.

So I went looking for real CPU-only settings instead, and found something more useful than any slider: I had never actually been running at full speed.
This laptop, like a lot of ultrabooks, ships with a vendor power-performance restriction that keeps the CPU clocked down even while plugged into the mains. I had no idea it was active. Removing it took a small Qwen model from crawling to 7 tokens per second, with CPU utilisation finally reaching around 90% for the first time in this whole test. On gpt-oss-20b, the same fix plus a pass of LM Studio settings took it from 3.6 to 5.5 tokens per second — roughly a 50% jump, from a setting that has nothing to do with the model at all.
Here is the diagnostic, in one sentence: if a CPU-bound job is not pushing your CPU close to 100%, the model is not your bottleneck. Check Task Manager before you touch a single slider.
Every timing elsewhere in this post — the table further down included — was measured before I found this, and, like everything else here, from single runs rather than benchmark averages. The comparisons between models still hold, because every run was equally throttled. Treat every absolute number in this piece as a floor. This machine has more in it than the numbers above show.
Once the CPU was actually working, a handful of settings were worth getting right on top of it:
- GPU offload → 0, explicitly. LM Studio defaulted to offloading 24 layers to a GPU I do not have. Do not just ignore the slider — set it to zero, because zero is not the default.
- CPU thread count → physical cores, not logical. Four threads, not eight. Hyperthreaded threads do not help CPU-bound matrix maths, and can add scheduling contention instead.
- Offload KV cache to GPU → off. Same reasoning as the main offload setting.
- KV cache quantization → on. The one setting here that is a genuine architectural win rather than noise.
- Keep model in memory → on. Try mmap() → on. Both cut reload overhead for free.
- Stay on AC power. Battery mode re-throttles most laptops regardless of what the power plan claims.
gpt-oss-20b itself is MXFP4 at 11.28 GB — already quantized to roughly 4 bits. If you were hoping to just quantize it harder, there is not much headroom left; it is already close to as compressed as this architecture gets.
The thing that actually predicts speed
Here is the insight that changed how I choose models, and I wish someone had told me on day one.
On a CPU, generation speed tracks the size of the file on disk. Not the parameter count. Not the model’s reputation. The file size.
| Model file | Speed |
|---|---|
| 1.89 GB | 5 tok/s |
| 4.46 GB | 1.65 tok/s |
| 15.7 GB | 0.3–0.5 tok/s |
Roughly linear, for a wonderfully boring reason: producing each token means streaming the model’s weights through memory. With no GPU, you are limited by memory bandwidth. Bigger file, more bytes per token, slower output. Your CPU is barely the bottleneck — your RAM speed is.
Once that clicked, two results that had confused me suddenly made sense.
Surprise one: the same model got slower when I made it “better”
Quantization is compression — fewer bits per parameter, smaller file, a little quality lost. The standard advice is to take the highest quantization that fits, because higher is more capable.
I ran Qwen2.5-Coder-1.5B at two levels:
- Q2_K, 752 MB → faster
- Q8_0, 1.89 GB → measurably slower
Same model. Same 1.5 billion parameters. The only thing that changed was bytes per parameter — and the speed moved with the bytes.
This is where the usual mental model breaks. Everyone talks about model size in billions of parameters, but on a CPU that is the wrong unit entirely. A 1.5B at Q8 and a 1.5B at Q2 are not the same speed, and if you only ever look at the parameter count you will never understand why one of them crawls.
Surprise two: a 22 GB model beat a 4.5 GB one
This should be impossible under the rule I just gave you. It is the exception that proves the mechanism.
- Qwen2.5-Coder-7B, dense, 4.46 GB → 1.65 tok/s
- Qwen3-Coder-30B-A3B, Mixture-of-Experts, 22 GB → 3.32 tok/s
Twice the speed at five times the size.
The answer is architecture. A dense model streams every weight for every token. A Mixture-of-Experts model wakes only a fraction of the network per token — so despite being enormous on disk, it moves far fewer bytes to produce each word.
So the rule refines to: speed tracks bytes actually streamed per token. For dense models that equals file size. Mixture-of-Experts breaks the correlation deliberately, and if you are on CPU it is the single most useful thing to look for on a model card.
The catch is real, though. That 22 GB model held 21.6 GB resident and pushed system memory to 97%. It ran, and nothing else was going to run beside it. Sparse compute does not mean sparse memory — you still hold the whole model.
The part I did not expect: almost all of them got it right
I went in assuming that a local LLM without a GPU, running a small model on a slow machine, would produce broken nonsense. I was wrong, and this is the finding I most want to pass on.
Of the twelve local runs that produced a file, eleven produced a mathematically correct mortgage calculator.
Correct means all of: the annual rate converted properly with a division by 100 and by 12, the term converted to months, the real amortization formula, and — mostly — a guard for the 0% interest case. I established every verdict below by opening the generated code and reading it, not by trusting the run log. That matters, because the log was wrong more than once.
| Model | Verdict |
|---|---|
| Qwen2.5-Coder-14B | Correct — proper conversions, zero-interest guard |
| Qwen2.5-Coder-7B @ Q4_K_M | Correct formula — but no zero-interest guard |
| Qwen2.5-Coder-1.5B @ Q8_0 | Correct — from a 1.9 GB model |
| Qwen3-Coder-30B-A3B (MoE) | Correct — the fast one from the section above |
| gpt-oss-20b | Correct, clean, well guarded |
| Bonsai-27B (reasoning on) | Correct — strongest input validation of the lot |
| Bonsai-27B (reasoning off) | Correct, and much terser |
| Gemma 4 26B-A4B (MoE, default) | Correct — 3 files |
| Gemma 4 26B-A4B (no reasoning) | Correct — 3 files |
| Gemma 4 E2B (no reasoning) | Correct — from a tiny model |
| Ornith 1.5 35B-A3B (retry) | Correct |
| Nemotron 3 Nano 4B | Wrong maths — see below |
| Ornith 1.5 35B-A3B ×2 | No file — though one of them did finish. See below |
| Phi-4-mini-reasoning | No output at all |
Look at that list again. Qwen2.5-Coder-1.5B — a 1.9 GB file — wrote a correct mortgage calculator. So did Gemma 4 E2B. These are models that would fit on a phone.
If you are starting out and quietly worried that your laptop is too weak to be useful: on a well-defined, self-contained task, it very probably is not.
The one that got it wrong, and how
Nemotron 3 Nano 4B was the single arithmetic failure, and it is a beautiful example of why you have to read the output.
const interest = parseFloat(document.getElementById('interest').value);
// ...
const r = interest/12; // monthly rate
Spot it? It divides by 12 but never by 100. Enter 5%, and the model computes a monthly rate of 0.4167 instead of 0.004167 — one hundred times too large. The payment figure that comes out is nonsense.
And there, on the same line, sits a comment confidently asserting that this is the monthly rate.
The model was not lying. It had no mechanism to check. Everything about that code looks right — sensible variable names, a tidy comment, a zero-interest guard two lines above that works perfectly. It is wrong in exactly one character’s worth of arithmetic, and no amount of reading the summary would ever have told you.
Read the code. Always read the code.
How long each one actually took
Correctness is only half of it. Here is the same set ordered by wall-clock time on the identical task, timed as each run finished.
| Model | Time | Outcome |
|---|---|---|
| Ornith 1.5 35B-A3B (reasoning off) | ~3.9 min | Finished generating — but the file never got saved. See below |
| gpt-oss-20b | ~7.5 min | Correct, clean, 0% case handled first try |
| Gemma 4 26B-A4B (reasoning off) | ~10.9 min | Correct — the best-looking result of any local run |
| Nemotron 3 Nano 4B | ~12.5 min | Wrong maths |
| Qwen2.5-Coder-7B @ Q4_K_M | ~16 min | Correct, no zero-interest guard |
| Gemma 4 26B-A4B (reasoning on) | ~18 min | Correct, most polished |
| Gemma 4 E2B (reasoning off) | ~19.6 min | Correct — no speed benefit from being tiny |
| Bonsai-27B (reasoning on) | ~22.4 min | Correct, strongest input validation |
| Qwen3-Coder-30B-A3B | ~23 min | Correct, thousands separators |
| Ornith 1.5 35B-A3B (retry) | ~27 min | Correct |
| Qwen2.5-Coder-14B | ~27 min | Correct, plain |
| Phi-4-mini-reasoning | >30 min ×2 | Killed both times. Zero bytes returned |
Two runs are missing from that table on purpose. Qwen2.5-Coder-1.5B has a recorded time, but the run log’s verdict for it disagrees with the file that is actually on my disk — and since I have been trusting files over the log all the way through this piece, I am not going to borrow the log’s stopwatch for the one model where I do not trust its judgement. The first Ornith attempt has no recorded time at all. Better a gap than a guess.
Two things jump out of that middle column.
Small does not mean fast. Gemma 4 E2B is the tiniest model in the test and took 19.6 minutes — nearly three times gpt-oss-20b, which is many times its size. Parameter count told you nothing, again. What mattered was how much of the network woke up per token, and how much of the budget went on reasoning nobody asked for.
Reasoning is the most expensive setting on the page. The same Gemma 4 26B-A4B took 18 minutes with reasoning on and 10.9 with it off, and wrote correct code both times. That is a 40% saving for no loss on this task. It is not a universal switch — suppress Bonsai’s reasoning and it breaks — but it is the first thing to try, and nobody tells you that.
If you only install two, install these
I did not expect to come out of this with favourites. I did.
gpt-oss-20b is the one I would put in front of a beginner. The fastest local run that actually produced a file, at 7.5 minutes, around 3 tokens per second on screen, correct on the first attempt, and it handled the 0% interest case without being asked — which models with far better reputations did not. It is 11.28 GB on disk and Mixture-of-Experts, so only about 3.6B parameters wake per token. That is the whole reason it is quick. It is not a toy you tolerate. It is something you would genuinely leave running.
Gemma 4 26B-A4B is the one I would keep beside it. Also around 3 tokens per second, 10.9 minutes with reasoning switched off, and it produced the best-looking output of any local model here — proper currency formatting via Intl.NumberFormat, which nothing else bothered with. Turn reasoning off and it stays correct while getting a third faster.
Between them they cover most of what you would want from a machine with no GPU: one that is quick and dependable, and one whose output you would not be embarrassed to show someone. Both are Mixture-of-Experts, which by now should not surprise you.
And the models I had the highest hopes for are the ones that let me down. Phi-4-mini-reasoning is small, well reviewed, and appears in an article currently ranking for this exact topic at 12 tokens per second. On my machine it ran past thirty minutes, twice, and returned nothing at all. Specs predict speed badly. Run the thing.
What actually stopped models wasn’t intelligence
Three of fifteen runs produced nothing at all — not bad code, no file. Phi-4-mini-reasoning is the clean case: two attempts, each pushed past thirty minutes, zero bytes returned.
The two Ornith failures are more interesting, and I have to be straight about one of them. It did not fail. It finished generating in under four minutes — the fastest local run in the whole test — and the file never got written because I stopped the run while it was still retrieving the output. That was my tooling, not the model.
Which makes Ornith the strangest result here: 3.9 minutes on one attempt and 27 on another, same model, same task. If you are choosing between local models, that kind of inconsistency deserves more of your attention than any headline speed figure, because you cannot plan around it.
So the barrier to running a local LLM without a GPU for real work is not that the models are stupid. It is that they are slow, that they are erratic, and that some of the time nothing arrives at all.
The trade-off nobody warns you about
Speed and capability pull against each other, hard, and there is no setting that fixes it.
Models fast enough to feel conversational are small enough that you will eventually catch them out. Models capable enough to trust are slow enough that you stop waiting and go and make coffee. That is simply the shape of a local LLM without a GPU, and the sooner you accept it the better your choices get.
So the useful question is not “which local model is best.” It is “what am I actually going to do with it?”
- Interactive back-and-forth? You need several tokens per second, which caps model size, which caps ambition. Qwen2.5-Coder-1.5B at Q8 was my fast tier at around 5 tok/s — and it still got the calculator right.
- A real task you will come back to? This is where gpt-oss-20b and Gemma 4 26B-A4B live, at around 3 tokens per second and eight to eleven minutes for a small build. Long enough to make the coffee, short enough that you have not lost the thread. It is the most useful tier on this machine by a distance.
- Fire-and-forget — generate this, I will look in half an hour? Qwen3-Coder-30B-A3B and Bonsai-27B sit here at 22 to 23 minutes. Both correct, both more polished. Worth it only when the extra quarter of an hour costs you nothing.
I ended up keeping both, for different jobs. That is not a compromise. That is just what the constraint looks like once you stop fighting it.
Would more RAM fix it?
Less than you would hope.
More RAM raises the ceiling — my 40GB let me load a 22 GB model that a 16GB machine simply cannot. That is real, and if you are choosing a machine it matters.
But it does not change throughput, because throughput is bound by memory bandwidth and cores, not capacity. More RAM lets you run bigger models at the same disappointing speed. The upgrade that moves these numbers is a discrete GPU — a different machine, not a bigger stick of memory.
Starting a local LLM without a GPU: what I would do
Start smaller than you think. Qwen2.5-Coder-1.5B at Q8 is about 1.9 GB, runs at a usable speed, and wrote correct code in this test. Get that working in LM Studio before you download anything enormous. The thrill is identical and you get it in ten minutes rather than an hour.
Ignore parameter counts. Look at file size, and check whether it is Mixture-of-Experts. On CPU, MoE is worth actively seeking out.
Set GPU offload to zero, explicitly if your graphics are integrated — don’t just leave the slider alone. Mine defaulted to offloading 24 layers to a GPU that does not exist. Measured contribution of offloading: nothing.
Check whether your laptop is quietly throttling itself before you blame the model. Mine was, on AC power, the whole time. If a CPU-bound run is not pushing your CPU close to 100%, that is where to look first — it was worth more than every other setting on this list combined.
Decide interactive versus batch before you download. It determines everything, and the downloads themselves take tens of minutes.
Turn reasoning off before you conclude a model is slow. It took 40% off Gemma 4’s time here with no loss of correctness. Check the output afterwards, because it breaks some models outright, but try it first.
Expect roughly one run in five to give you nothing — and expect at least one of those to be your own fault, as one of mine was. It is not always the model. Budget for it either way.
And always read the code. Nemotron wrote a comment that said “monthly rate” above a line that was a hundred times wrong. That is the real lesson of the weekend, and it applies well beyond local models.
None of this is what we run for clients, to be clear — that work is ordinary WordPress AI integration, on servers that have proper hardware behind them. This weekend was about finding out where the floor is. If you want the non-technical version of where these models currently stand, we wrote a plain-English guide to the 2026 AI models earlier this year.
Now go and pull the network cable out, and watch it keep answering. It genuinely does not get old.
If nobody's looking after your site, hand it off.
Our Website as a Service model handles upkeep, security, and improvements for a predictable monthly fee — so you never have to think about it.
See how WaaS works