Measured · 9 and 10 October 2026 · six engines
Qwen3.8-27B as an agent on the Mac Studio M1 Ultra
Rapid-MLX, oMLX, MTPLX, mlx-dspark, llama.cpp, mlx-lm with agent histories up to 98,000 tokens
Authors
Machine
Mac Studio M1 Ultra
- SoC
- Apple M1 Ultra, 20 CPU cores (16 performance, 4 efficiency), 64 GPU cores
- Memory
- 64 GiB unified memory, 56 GiB of it released to the GPU
- OS
- macOS 26.2 (25C56), on mains power
Model
Qwen3.8-27B
- Build
- dense, hybrid: Gated DeltaNet layers with a recurrent state and classic attention
- Weights
- each engine its own 4-bit version, with a draft head where the engine uses one
- Context
- 262,144-token window on every engine
Engines
Version and weights
- Rapid-MLX
0.15.7·Qwen3.8-27B-4bit-MTP-fp16-MLX- oMLX
0.7.0·Qwen3.8-27B-oQ4e-fp16-mtp- MTPLX
2.12.2·Qwen3.8-27B-MTPLX-Optimized-Speed-FP16- mlx-dspark
0.20.3·Qwen3.8-27B-4bit-MTP-fp16-MLX- llama.cpp
0.6.0·Qwen3.8-27B-UD-Q4_K_M.gguf- mlx-lm
0.32.0·Qwen3.8-27B-4bit
Measurement
The way OpenCode works
- API
- OpenAI-compatible, streamed, with tools
- Prompts
- histories of a real coding agent, 2,300 to 98,000 tokens
- Settings
- temperature 0, reasoning effort low, up to 1,024 output tokens
- Procedure
- one engine at a time, pre-check before every start, sensors every 2 s
Rating at a glance
good fair weak
| Engineversion, weights, use | Keeping historyafter side question, second session | Restarthistory afterwards | Readingcold, 13k | Outputagent, 13k | Long contextcorrect, follow-up | Faithful outputwith speculation, no cache |
|---|---|---|---|---|---|---|
| Rapid-MLX0.15.7 · 4bit-MTP-fp16-MLXagent up to 52k | good: 1.3 s | good: 1.2 s | good: 300 tokens/s | good: 31.5 tokens/s | weak: rereads32k: follow-up 116 s | good: identicalin every run |
| oMLX0.7.0 · oQ4e-fp16-mtpagent up to 98k; aborts truncated tool calls | good: 2.3 s | good: 6.6 s | good: 293 tokens/s | fair: 29.5 tokens/s | fair: up to 64k23/24 · 18 s | good: identicalin every run |
| MTPLX2.12.2 · MTPLX-Optimized-Speed-FP16fastest output, but a different one with speculation | good: 1.5 s | good: 5.9 s | good: 285 tokens/s | good: 37.3 tokens/s | weak: up to 32k24/24 · 7.3 s | weak: differsspeed 2.0-fold |
| mlx-dspark0.20.3 · 4bit-MTP-fp16-MLXagent without restarts; cache is not found again | good: 1.3 s | weak: 3.6 min | good: 301 tokens/s | fair: 27.6 tokens/s | fair: up to 64k23/24 · 26 s | fair: longeronly at the token limit |
| llama.cpp0.6.0 · UD-Q4_K_M.ggufdocuments beyond 64k | good: 1.0 s | weak: 4.4 min | fair: 232 tokens/s | weak: 20.2 tokens/s | good: up to 128k24/24 · 1.8 s | good: identicalin every run |
| mlx-lm0.32.0 · 4bitreference; loses the history at every side question | weak: 4.7 min | weak: 4.7 min | weak: 214 tokens/s | weak: 23.4 tokens/s | weak: rereads32k: follow-up 2.6 min | good: identicalin every run |
Levels follow fixed thresholds. Keeping history and restart: first token in the 52k history, good up to 5 or 10 s, fair up to 30 or 60 s. Reading and output: share of the best value, good from 90 or 80 %, fair from 75 or 65 %. Long context: the greatest length with at most one wrong answer and a follow-up within 30 s, good from 128k, fair from 64k. Faithful output: byte-identical to the run without speculation and without cache. All six also solved every tool task.
01 · Setup
Six servers, one model, the same histories
At 4 bits Qwen3.8-27B fits into the 64 GiB of the Mac Studio and leaves room for a long history. On the Mac six servers with an OpenAI-compatible interface run it: mlx-lm as Apple's reference, its offshoots Rapid-MLX, oMLX, MTPLX and mlx-dspark, and llama.cpp with Metal. Every engine ran with the weights it is built for and with its cache switched on the way one would use it day to day.
Four things are measured. First, a session: an agent history of 52,000 tokens growing turn by turn, interrupted by a side question, a second session and a server restart. Second, eight single agent steps from 2,300 to 98,000 tokens, each cold, for plain reading and output speed. Third, follow-up questions about a document of 32,000 to 128,000 tokens with checked answers. Fourth, whether the speculation that four of the engines advertise changes the output.
The agent histories are cut from the trace of a real coding agent (mini-swe-agent); the same set already ran on the HP ZBook Ultra G1a. All times are taken on the measuring script's clock, not from what the servers report. Before every start a script checked that no other server, download or build was running. All six handle tool calls: every engine solved the five tool tasks in three rounds without a formal error.
02 · Session
Whoever keeps the history answers in a second
Every engine reads the first turn cold: 3.5 to 4.7 minutes at 52,000 tokens. After that the next turn starts after 1.0 to 3.0 seconds on all of them. The interruptions make the difference: only Rapid-MLX stays at 1.3 seconds or less in every situation.
The reason lies in the model. Its DeltaNet layers carry a state that sums up the whole history and cannot be cut back to an earlier point. An engine can therefore reuse a stored history only if the new prompt continues it exactly, or if it kept snapshots of the state along the way. Agent steps continue the history; every engine manages that.
A side question or a second session puts another history in between. mlx-lm holds only one entry of this size and then reads for 4.7 min again. Only the engines that write the history to the SSD carry it across a restart: Rapid-MLX (1.2 s), oMLX (6.6 s) and MTPLX (5.9 s). llama.cpp keeps the history in RAM; mlx-dspark does write a disk cache but does not find it again after the restart.
03 · Turns
Reading 52k cold costs three and a half to almost five minutes
Without a cache the M1 Ultra reads a 13,000-token agent history at 214 to 301 tokens per second. At 52,000 tokens the first token takes 3.5 to 4.7 minutes.
Reading sets an agent's wait as soon as the cache loses the history. Output sets the time after that: a step with 1,000 tokens of thinking and a tool call takes about 27 seconds at 37.3 tokens per second and about 50 at 20.2. During the steps the chip drew 95 to 102 watts at 71 to 74 °C; the GPU clock stayed at 1,296 MHz on every engine.
04 · Long context
At 128k only llama.cpp answers at once and without error
Questions about a long document branch: every follow-up starts again right after the document, not after the last answer. For the model's recurrent state this is harder than an agent step.
At 32k all six engines solve all 24 tasks from three versions of the document. At 128k llama.cpp stays error-free, oMLX solves 22 of 24 and then writes only 10.5 tokens per second. MTPLX reads 128k in 19 minutes and then fails on memory errors: 3 of 24. mlx-dspark ran up to 64k; at 128k its memory guard emptied the cache and one follow-up took 16 minutes.
05 · Speculation
Speculation only pays off with MTPLX and changes its output
Speculative decoding lets a small draft head propose several tokens that the model checks in one step. Built correctly it changes no token, only the speed. This is checked per engine with two fresh servers without a cache, once without and once with speculation, the same four prompts twice each.
Re-checked with full text, MTPLX departs in the prose prompt from character 712 of the thinking, in agent task A03 from character 856; there it then writes a different answer and calls the shell with different commands. The result is not wrong, but it is not the model's own without a draft. mlx-dspark writes the same tokens with speculation and only appends accepted draft tokens at the limit of 512 or 1,024 tokens. Rapid-MLX, oMLX, llama.cpp and mlx-lm give byte-identical output in every run.
oMLX switches speculation only in its model settings, not at start; there it ran in both arms. llama.cpp and mlx-lm have none for this model.
06 · Choice
Rapid-MLX for the agent, llama.cpp for long documents
For OpenCode with histories up to 52k, as far as measured: Rapid-MLX. It keeps the history across side questions, a second session in between and restarts, answers within 1.3 seconds, gives the same tokens with and without speculation and writes 31.5 tokens per second at 13k. The client's time-out still has to survive the first turn: a new 52k history takes 3.6 min here too.
For questions about documents beyond 64k: llama.cpp. It is the slowest writer, but the only one that answers at 128k at once and without error. Its cache does not survive a restart.
MTPLX writes fastest and keeps the history too, but changes the output with speculation, sends tool calls in a few large chunks and runs into its memory guard from 52k. oMLX keeps the history and reaches 98k, but aborts a tool call cut off at the token limit with an error. mlx-lm is measured as the reference; for an agent it loses the history too easily.
07 · Every figure
Results per engine
Every chart in this report is computed from these values; the values per request are in docs/m1-ultra-werte.json, the raw data under laeufe/m1-ultra/.
| Step | Rapid-MLX | oMLX | MTPLX | mlx-dspark | llama.cpp | mlx-lm |
|---|---|---|---|---|---|---|
| first turn, cold | 216.1 | 212.3 | 240.1 | 216.3 | 263.4 | 283.5 |
| next turn | 1.3 | 3.0 | 1.1 | 1.3 | 1.0 | 1.2 |
| after side question | 1.3 | 2.3 | 1.5 | 1.3 | 1.0 | 279.3 |
| after second session | 1.3 | 2.2 | 0.8 | 1.3 | 1.0 | 279.6 |
| after restart | 1.2 | 6.6 | 5.9 | 217.3 | 264.9 | 283.7 |
| second session after restart | 0.9 | 1.3 | 2.5 | 45.3 | 57.8 | 67.2 |
| Step | Rapid-MLX | oMLX | MTPLX | mlx-dspark | llama.cpp | mlx-lm |
|---|---|---|---|---|---|---|
| A01 2k | 9 / 36.1 | 12 / 31.5 | 9 / 45.4 | 8 / 32.1 | 10 / 21.6 | 11 / 26.1 |
| A02 5k | 17 / 31.8 | 17 / 31.0 | 19 / 48.2 | 17 / 32.1 | 21 / 21.0 | 23 / 25.3 |
| A03 13k | 44 / 31.5 | 45 / 29.5 | 47 / 37.3 | 43 / 27.6 | 56 / 20.2 | 61 / 23.4 |
| A04 13k | 45 / 30.6 | 46 / 29.5 | 48 / 41.7 | 45 / 27.1 | 58 / 20.2 | 63 / 23.4 |
| A05 52k | 216 / 20.7 | 208 / 23.1 | 239 / – | 215 / 27.7 | 264 / 17.4 | 282 / 18.1 |
| A06 53k | 221 / 24.3 | 213 / 23.0 | 231 / 28.4 | 220 / 25.6 | 269 / 17.4 | 293 / 18.1 |
| A07 97k | – | 453 / 15.0 | – | – | 581 / 15.4 | – |
| A08 98k | – | 464 / – | – | – | 597 / 14.9 | – |
| Engine | Length | correct | Reading (s) | Reading (t/s) | Follow-up (s) | Output (t/s) |
|---|---|---|---|---|---|---|
| Rapid-MLX | 32k | 24 / 24 | 116 | 276 | 116.4 | 27.3 |
| oMLX | 32k | 24 / 24 | 119 | 269 | 16.6 | 26.3 |
| oMLX | 64k | 23 / 24 | 268 | 239 | 18.0 | 19.9 |
| oMLX | 128k | 22 / 24 | 657 | 195 | 14.2 | 10.5 |
| MTPLX | 32k | 24 / 24 | 122 | 264 | 7.3 | 37.1 |
| MTPLX | 64k | 24 / 24 | 340 | 189 | 57.9 | 28.5 |
| MTPLX | 128k | 3 / 24 | 1,142 | 112 | – | 18.4 |
| mlx-dspark | 32k | 24 / 24 | 118 | 273 | 5.1 | 33.3 |
| mlx-dspark | 64k | 23 / 24 | 294 | 218 | 26.5 | 28.5 |
| llama.cpp | 32k | 24 / 24 | 151 | 214 | 1.0 | 18.7 |
| llama.cpp | 64k | 24 / 24 | 342 | 188 | 1.3 | 16.7 |
| llama.cpp | 128k | 24 / 24 | 854 | 150 | 1.8 | 13.7 |
| mlx-lm | 32k | 24 / 24 | 159 | 202 | 157.8 | 21.1 |
| Engine | Prose | Code | A01 | A03 | Chip (W) | °C |
|---|---|---|---|---|---|---|
| Rapid-MLX | identical | identical | identical | identical | 101 | 73.2 |
| oMLX | identical | identical | identical | identical | 100 | 74.2 |
| MTPLX | differs | identical | identical | differs | 102 | 73.8 |
| mlx-dspark | differs | differs | identical | differs | 99 | 73.8 |
| llama.cpp | identical | identical | identical | identical | 95 | 72.2 |
| mlx-lm | identical | identical | identical | identical | 97 | 70.8 |
08 · Limits
What this work does not show
One session, one pass. The session ran once per engine. At temperature 0 the output repeats itself; whether a cache hits may however depend on the state of memory and was not measured repeatedly.
Rapid-MLX in the long-context runs without speculation. Started from a local snapshot, Rapid-MLX turns speculation on only when the draft head is named explicitly, and the long-context runs did not name it. Chapter 05 shows that on this machine it changes neither speed nor output.
Different weights. Every engine ran with the 4-bit version it is built for. Speed and answers therefore compare engine plus weights, not the engine alone.
Capped output. Every agent step ended after at most 1,024 tokens. Some steps hit the limit. That does not affect speed; whether they would have finished is not measured.
One machine, one version. What is measured are the versions named, on 9 and 10 October 2026. The MLX offshoots release almost weekly; a later version may behave differently.