Measured · 15 June 2025 · three runs per model
Nine models on the HP ZBook Ultra G1a
dense against mixture of experts
Authors
Machine
HP ZBook Ultra 14 inch G1a
- SoC
- AMD Ryzen AI Max+ PRO 395
- GPU
- Radeon 8060S
- Memory
- 128 GB – dedicated graphics memory set to 512 MB in the BIOS
- OS
- Windows 11 Pro 24H2
Runtime – the same for every model
LM Studio 0.3.16
- Engine
- llama.cpp Vulkan v1.34.1
- Offload
- as many layers on the GPU as the machine allowed
- Context
- 8,192 tokens, identical for every run
- KV-Cache
- on the GPU
Task
Summarise
- Source
- English Wikipedia article on Plato, 27 KB
- Prompt
- a two- to three-sentence introduction plus four bullet points
- Runs
- three per model: one without cache, two with
- Models
- 9 – from 0.6 billion to over 100 billion parameters
Where this report comes from
Article from 2025
- Published
- 15 June 2025
- Evidence
- results table and 47 screenshots from LM Studio
- This edition
- newly typeset, charts computed from the same figures
- Successor
- the measurement of 8 September 2026 on the same machine
What this machine did in 2025
44.0Token/s
Qwen3-0.6b Q8 the fastest model in the field
25.8Token/s
Qwen-30B-A3B Q4_K_M a factor of 5.6 over the equally large QwQ 32B Q8
0.24seconds
wait with cache instead of 119.4 s without, on QwQ 32B Q8
4.3Prozent
widest spread across three runs; every other model less
Speed figures are means over the three runs of each model. The wait without cache comes from the first run, the one with cache is the mean of the two that follow. Every figure appears in the results table of the original article.
01 · The field
Nine models, one machine
Between the fastest and the slowest model lies a factor of 9.5: 44.0 against 4.6 tokens per second. Both ran on the same machine, on the same task.
02 · Build
Why the experts come out ahead
A dense model computes every token using all of its parameters. A mixture-of-experts model stores just as many but computes with only a part. Memory demand follows the first number, speed follows the second.
Two models in this field make the difference immediate, because their names state the same order of magnitude. Qwen-30B-A3B Q4_K_M is a mixture of experts and reaches 25.8 tokens per second. QwQ 32B Q8 is dense, about the same size – and manages 4.6. That is a factor of 5.6 in favour of the build, at practically the same parameter count.
The relationship is not without exception. Llama-4-Scout-17B-16e-Instruct is a mixture of experts as well and stays behind several dense models at 6.1 tokens per second – for a reason that has nothing to do with the build. Chapter 05 shows it.
A year later this became the basis of the next measurement: Qwen3.8-Flash-Next holds 512 experts and activates ten per token. What shows up here as an advantage carries a 176.9-billion-parameter model on the same notebook there.
03 · The cache
What is left the second time
The same prompt, asked a second time: 119.4 seconds of waiting become 0.24. A factor of 487 – and the output speed stays the same throughout.
For actual use this is the difference between workable and unworkable. The first call on a long text makes you wait; no follow-up question on the same text does. Working through a document, you pay for the processing once rather than with every question.
04 · Reproducibility
Asked three times, the same number three times
No model spreads beyond 4.3 percent across its three runs. Seven of nine stay under two.
05 · The limit
Where the machine stopped short in 2025
Eight of the nine models ran entirely on the GPU. The ninth could not: in 2025 the Vulkan driver released at most 64 GB of graphics memory, and Llama-4-Scout needed more. Six of its 48 layers stayed on the CPU. The sensors show what that costs: the GPU clock falls from 2,650 to 846 MHz and utilisation to 72 percent while the CPU joins in at 61 percent.
The 2025 article closed with the assumption that a driver update might lift this limit. The measurement of 8 September 2026 on the same machine answers that, if under a different operating system: on Linux the kernel parameter amdgpu.gttsize releases 110 GiB of system memory to the GPU, and a 176.9-billion-parameter model runs. So the limit sat in the software stack, not in the machine.
06 · Every figure
The complete results table
27 runs, exactly as they appear in the original article. Every chart in this report is computed from this table.
| Model | Run | Tokens/s | Tokens | Wait (s) | Cache |
|---|---|---|---|---|---|
| Qwen3-0.6b Q8 | 1 | 44.47 | 485 | 2.26 | no |
| 2 | 43.71 | 446 | 0.06 | yes | |
| 3 | 43.78 | 429 | 0.06 | yes | |
| Qwen-30B-A3B Q4_K_M | 1 | 25.66 | 668 | 67.12 | no |
| 2 | 25.80 | 947 | 0.08 | yes | |
| 3 | 26.07 | 773 | 0.09 | yes | |
| Qwen-30B-A3B Q8 | 1 | 23.12 | 666 | 62.42 | no |
| 2 | 23.68 | 651 | 0.08 | yes | |
| 3 | 23.51 | 836 | 0.08 | yes | |
| Qwen3-4B Q8 | 1 | 19.61 | 627 | 10.09 | no |
| 2 | 19.85 | 697 | 0.09 | yes | |
| 3 | 19.92 | 598 | 0.09 | yes | |
| Nous-Hermes-2-Mixtral-8x7 Q8 | 1 | 10.65 | 208 | 38.46 | no |
| 2 | 10.75 | 225 | 0.13 | yes | |
| 3 | 10.76 | 266 | 0.13 | yes | |
| Gemma-3 12b | 1 | 7.75 | 324 | 28.61 | no |
| 2 | 7.65 | 289 | 0.17 | yes | |
| 3 | 7.60 | 333 | 0.17 | yes | |
| Magistral-Small Q8 | 1 | 6.60 | 924 | 80.20 | no |
| 2 | 6.56 | 2,561 | 0.19 | yes | |
| 3 | 6.62 | 1,041 | 0.22 | yes | |
| Llama-4-Scout-17B-16e-Instruct | 1 | 5.92 | 301 | 82.54 | no |
| 2 | 6.11 | 356 | 0.23 | yes | |
| 3 | 6.18 | 286 | 0.21 | yes | |
| QwQ 32B Q8 | 1 | 4.65 | 782 | 119.42 | no |
| 2 | 4.64 | 1,125 | 0.24 | yes | |
| 3 | 4.62 | 1,297 | 0.25 | yes |
07 · Limits
What this work does not show
No quality. What was measured is speed, not the quality of the summaries. Whether one model solves the task better than another is not stated here.
One prompt, one length. Every run uses the same text with an 8,192-token context. The measurement says nothing about behaviour on very long contexts – that is what the 2026 report is for.
One stack, one moment. Windows 11, LM Studio 0.3.16, llama.cpp Vulkan v1.34.1, as of June 2025. Drivers and engines have changed since; the figures hold for this stack and this date.
Three runs. Enough to show the spread, too few for a statistical claim about rare outliers.