holistechlabsQwen3.8-27B · 6 Engines
DEEN

Measured · 9 and 10 October 2026 · six engines

Qwen3.8-27B as an agent on the Mac Studio M1 Ultra

Rapid-MLX, oMLX, MTPLX, mlx-dspark, llama.cpp, mlx-lm with agent histories up to 98,000 tokens

Authors

Dipl.-Ing. Sören GebbertInstitut für holistische Technologieforschung GmbH Claude Code Opus 5.5Anthropic

Machine

Mac Studio M1 Ultra

SoC
Apple M1 Ultra, 20 CPU cores (16 performance, 4 efficiency), 64 GPU cores
Memory
64 GiB unified memory, 56 GiB of it released to the GPU
OS
macOS 26.2 (25C56), on mains power

Model

Qwen3.8-27B

Build
dense, hybrid: Gated DeltaNet layers with a recurrent state and classic attention
Weights
each engine its own 4-bit version, with a draft head where the engine uses one
Context
262,144-token window on every engine

Engines

Version and weights

Rapid-MLX
0.15.7 · Qwen3.8-27B-4bit-MTP-fp16-MLX
oMLX
0.7.0 · Qwen3.8-27B-oQ4e-fp16-mtp
MTPLX
2.12.2 · Qwen3.8-27B-MTPLX-Optimized-Speed-FP16
mlx-dspark
0.20.3 · Qwen3.8-27B-4bit-MTP-fp16-MLX
llama.cpp
0.6.0 · Qwen3.8-27B-UD-Q4_K_M.gguf
mlx-lm
0.32.0 · Qwen3.8-27B-4bit

Measurement

The way OpenCode works

API
OpenAI-compatible, streamed, with tools
Prompts
histories of a real coding agent, 2,300 to 98,000 tokens
Settings
temperature 0, reasoning effort low, up to 1,024 output tokens
Procedure
one engine at a time, pre-check before every start, sensors every 2 s

Rating at a glance

good fair weak

Engineversion, weights, use Keeping historyafter side question, second session Restarthistory afterwards Readingcold, 13k Outputagent, 13k Long contextcorrect, follow-up Faithful outputwith speculation, no cache
Rapid-MLX0.15.7 · 4bit-MTP-fp16-MLXagent up to 52kgood: 1.3 sgood: 1.2 sgood: 300 tokens/sgood: 31.5 tokens/sweak: rereads32k: follow-up 116 sgood: identicalin every run
oMLX0.7.0 · oQ4e-fp16-mtpagent up to 98k; aborts truncated tool callsgood: 2.3 sgood: 6.6 sgood: 293 tokens/sfair: 29.5 tokens/sfair: up to 64k23/24 · 18 sgood: identicalin every run
MTPLX2.12.2 · MTPLX-Optimized-Speed-FP16fastest output, but a different one with speculationgood: 1.5 sgood: 5.9 sgood: 285 tokens/sgood: 37.3 tokens/sweak: up to 32k24/24 · 7.3 sweak: differsspeed 2.0-fold
mlx-dspark0.20.3 · 4bit-MTP-fp16-MLXagent without restarts; cache is not found againgood: 1.3 sweak: 3.6 mingood: 301 tokens/sfair: 27.6 tokens/sfair: up to 64k23/24 · 26 sfair: longeronly at the token limit
llama.cpp0.6.0 · UD-Q4_K_M.ggufdocuments beyond 64kgood: 1.0 sweak: 4.4 minfair: 232 tokens/sweak: 20.2 tokens/sgood: up to 128k24/24 · 1.8 sgood: identicalin every run
mlx-lm0.32.0 · 4bitreference; loses the history at every side questionweak: 4.7 minweak: 4.7 minweak: 214 tokens/sweak: 23.4 tokens/sweak: rereads32k: follow-up 2.6 mingood: identicalin every run

Levels follow fixed thresholds. Keeping history and restart: first token in the 52k history, good up to 5 or 10 s, fair up to 30 or 60 s. Reading and output: share of the best value, good from 90 or 80 %, fair from 75 or 65 %. Long context: the greatest length with at most one wrong answer and a follow-up within 30 s, good from 128k, fair from 64k. Faithful output: byte-identical to the run without speculation and without cache. All six also solved every tool task.

01 · Setup

Six servers, one model, the same histories

At 4 bits Qwen3.8-27B fits into the 64 GiB of the Mac Studio and leaves room for a long history. On the Mac six servers with an OpenAI-compatible interface run it: mlx-lm as Apple's reference, its offshoots Rapid-MLX, oMLX, MTPLX and mlx-dspark, and llama.cpp with Metal. Every engine ran with the weights it is built for and with its cache switched on the way one would use it day to day.

Four things are measured. First, a session: an agent history of 52,000 tokens growing turn by turn, interrupted by a side question, a second session and a server restart. Second, eight single agent steps from 2,300 to 98,000 tokens, each cold, for plain reading and output speed. Third, follow-up questions about a document of 32,000 to 128,000 tokens with checked answers. Fourth, whether the speculation that four of the engines advertise changes the output.

The agent histories are cut from the trace of a real coding agent (mini-swe-agent); the same set already ran on the HP ZBook Ultra G1a. All times are taken on the measuring script's clock, not from what the servers report. Before every start a script checked that no other server, download or build was running. All six handle tool calls: every engine solved the five tool tasks in three rounds without a formal error.

02 · Session

Whoever keeps the history answers in a second

Every engine reads the first turn cold: 3.5 to 4.7 minutes at 52,000 tokens. After that the next turn starts after 1.0 to 3.0 seconds on all of them. The interruptions make the difference: only Rapid-MLX stays at 1.3 seconds or less in every situation.

Whoever keeps the history answers in a second Time to first token in an agent history of 52,000 tokens, logarithmic. Mac Studio M1 Ultra · 64 GPU cores · 64 GiB · Qwen3.8-27B, each engine with its own 4-bit weights temperature 0 · reasoning effort low · one engine at a time · measured 9 and 10 Oct 2026 Rapid-MLX oMLX MTPLX mlx-dspark llama.cpp mlx-lm 0.5 s 1 s 3 s 10 s 1 min 5 min first token cold next turn 4.7 min after side question 4.7 min after second session 3.6 min 4.4 min 4.7 min after restart
mlx-lm loses the history at a mere side question, llama.cpp and mlx-dspark at a restart. Then the engine reads everything again, as long as on the first turn. Dashed: the median of the cold first turn.

The reason lies in the model. Its DeltaNet layers carry a state that sums up the whole history and cannot be cut back to an earlier point. An engine can therefore reuse a stored history only if the new prompt continues it exactly, or if it kept snapshots of the state along the way. Agent steps continue the history; every engine manages that.

A side question or a second session puts another history in between. mlx-lm holds only one entry of this size and then reads for 4.7 min again. Only the engines that write the history to the SSD carry it across a restart: Rapid-MLX (1.2 s), oMLX (6.6 s) and MTPLX (5.9 s). llama.cpp keeps the history in RAM; mlx-dspark does write a disk cache but does not find it again after the restart.

03 · Turns

Reading 52k cold costs three and a half to almost five minutes

Without a cache the M1 Ultra reads a 13,000-token agent history at 214 to 301 tokens per second. At 52,000 tokens the first token takes 3.5 to 4.7 minutes.

Reading 52k cold costs almost four minutes Time to first token on a new agent history, one line per engine. Mac Studio M1 Ultra · 64 GPU cores · 64 GiB · Qwen3.8-27B, each engine with its own 4-bit weights temperature 0 · reasoning effort low · one engine at a time · measured 9 and 10 Oct 2026 Rapid-MLX oMLX MTPLX mlx-dspark llama.cpp mlx-lm 2k 5k 13k 52k 100k Length of the prompt in tokens (logarithmic) 10 s 30 s 1 min 3 min 10 min first token
The MLX engines read at almost the same speed, llama.cpp and mlx-lm slightly behind. Only oMLX and llama.cpp ran up to 98,000 tokens; there the first step takes over seven minutes.
MTPLX writes fastest; long context slows them all Tokens generated per second on an agent step, read cold. Mac Studio M1 Ultra · 64 GPU cores · 64 GiB · Qwen3.8-27B, each engine with its own 4-bit weights temperature 0 · reasoning effort low · one engine at a time · measured 9 and 10 Oct 2026 Rapid-MLX oMLX MTPLX mlx-dspark llama.cpp mlx-lm 2k 5k 13k 52k 100k Length of the prompt in tokens (logarithmic) 0 10 20 30 40 50 Tokens per second
MTPLX writes fastest at 13k with 37.3 tokens per second, llama.cpp slowest with 20.2. At 52k they all drop; MTPLX's point at the first 52k step is missing because the engine's memory guard stepped in there.

Reading sets an agent's wait as soon as the cache loses the history. Output sets the time after that: a step with 1,000 tokens of thinking and a tool call takes about 27 seconds at 37.3 tokens per second and about 50 at 20.2. During the steps the chip drew 95 to 102 watts at 71 to 74 °C; the GPU clock stayed at 1,296 MHz on every engine.

04 · Long context

At 128k only llama.cpp answers at once and without error

Questions about a long document branch: every follow-up starts again right after the document, not after the last answer. For the model's recurrent state this is harder than an agent step.

At 128k only llama.cpp answers at once and without error Follow-up question about a document, first token; below it tasks solved correctly, three seeds per length. Mac Studio M1 Ultra · 64 GPU cores · 64 GiB · Qwen3.8-27B, each engine with its own 4-bit weights temperature 0 · reasoning effort low · one engine at a time · measured 9 and 10 Oct 2026 Rapid-MLX oMLX MTPLX mlx-dspark llama.cpp mlx-lm 1 s 3 s 10 s 30 s 100 s Follow-up 116 s 24/24 17 s 24/24 7.3 s 24/24 5.1 s 24/24 1.0 s 24/24 2.6 min 24/24 32k – – 18 s 23/24 58 s 24/24 26 s 23/24 1.3 s 24/24 – – 64k – – 14 s 22/24 3/24 – – 1.8 s 24/24 – – 128k correct
llama.cpp keeps snapshots of the state and answers at every length in 1.0 to 1.8 seconds. Rapid-MLX and mlx-lm read every follow-up from scratch and therefore only ran up to 32k. A dash: not measured.

At 32k all six engines solve all 24 tasks from three versions of the document. At 128k llama.cpp stays error-free, oMLX solves 22 of 24 and then writes only 10.5 tokens per second. MTPLX reads 128k in 19 minutes and then fails on memory errors: 3 of 24. mlx-dspark ran up to 64k; at 128k its memory guard emptied the cache and one follow-up took 16 minutes.

05 · Speculation

Speculation only pays off with MTPLX and changes its output

Speculative decoding lets a small draft head propose several tokens that the model checks in one step. Built correctly it changes no token, only the speed. This is checked per engine with two fresh servers without a cache, once without and once with speculation, the same four prompts twice each.

Speculation only pays off with MTPLX and changes its output Tokens generated per second, prose prompt, without and with speculation, both without cache. Mac Studio M1 Ultra · 64 GPU cores · 64 GiB · Qwen3.8-27B, each engine with its own 4-bit weights temperature 0 · reasoning effort low · one engine at a time · measured 9 and 10 Oct 2026 without speculation with speculation 0 10 20 30 40 50 Tokens per second MTPLX 24.4 50.0 output differs Rapid-MLX 33.6 32.9 output identical mlx-dspark 32.8 33.3 only longer oMLX 32.3 cannot be switched output identical llama.cpp 21.5 none output identical mlx-lm 26.3 none output identical
MTPLX doubles the speed from 24.4 to 50.0 tokens per second, but the output is no longer the same. On the M1 Ultra speculation brings no measurable gain with Rapid-MLX and mlx-dspark.

Re-checked with full text, MTPLX departs in the prose prompt from character 712 of the thinking, in agent task A03 from character 856; there it then writes a different answer and calls the shell with different commands. The result is not wrong, but it is not the model's own without a draft. mlx-dspark writes the same tokens with speculation and only appends accepted draft tokens at the limit of 512 or 1,024 tokens. Rapid-MLX, oMLX, llama.cpp and mlx-lm give byte-identical output in every run.

oMLX switches speculation only in its model settings, not at start; there it ran in both arms. llama.cpp and mlx-lm have none for this model.

06 · Choice

Rapid-MLX for the agent, llama.cpp for long documents

For OpenCode with histories up to 52k, as far as measured: Rapid-MLX. It keeps the history across side questions, a second session in between and restarts, answers within 1.3 seconds, gives the same tokens with and without speculation and writes 31.5 tokens per second at 13k. The client's time-out still has to survive the first turn: a new 52k history takes 3.6 min here too.

For questions about documents beyond 64k: llama.cpp. It is the slowest writer, but the only one that answers at 128k at once and without error. Its cache does not survive a restart.

MTPLX writes fastest and keeps the history too, but changes the output with speculation, sends tool calls in a few large chunks and runs into its memory guard from 52k. oMLX keeps the history and reaches 98k, but aborts a tool call cut off at the token limit with an error. mlx-lm is measured as the reference; for an agent it loses the history too easily.

07 · Every figure

Results per engine

Every chart in this report is computed from these values; the values per request are in docs/m1-ultra-werte.json, the raw data under laeufe/m1-ultra/.

Session with 52,000 tokens: first token in seconds
StepRapid-MLXoMLXMTPLXmlx-dsparkllama.cppmlx-lm
first turn, cold216.1212.3240.1216.3263.4283.5
next turn1.33.01.11.31.01.2
after side question1.32.31.51.31.0279.3
after second session1.32.20.81.31.0279.6
after restart1.26.65.9217.3264.9283.7
second session after restart0.91.32.545.357.867.2
Agent steps, cold: first token in seconds / output in tokens per second
StepRapid-MLXoMLXMTPLXmlx-dsparkllama.cppmlx-lm
A01 2k9 / 36.112 / 31.59 / 45.48 / 32.110 / 21.611 / 26.1
A02 5k17 / 31.817 / 31.019 / 48.217 / 32.121 / 21.023 / 25.3
A03 13k44 / 31.545 / 29.547 / 37.343 / 27.656 / 20.261 / 23.4
A04 13k45 / 30.646 / 29.548 / 41.745 / 27.158 / 20.263 / 23.4
A05 52k216 / 20.7208 / 23.1239 / –215 / 27.7264 / 17.4282 / 18.1
A06 53k221 / 24.3213 / 23.0231 / 28.4220 / 25.6269 / 17.4293 / 18.1
A07 97k–453 / 15.0––581 / 15.4–
A08 98k–464 / –––597 / 14.9–
Long context: median over three versions of the document
Engine Length correct Reading (s) Reading (t/s) Follow-up (s) Output (t/s)
Rapid-MLX32k24 / 24116276116.427.3
oMLX32k24 / 2411926916.626.3
oMLX64k23 / 2426823918.019.9
oMLX128k22 / 2465719514.210.5
MTPLX32k24 / 241222647.337.1
MTPLX64k24 / 2434018957.928.5
MTPLX128k3 / 241,142112–18.4
mlx-dspark32k24 / 241182735.133.3
mlx-dspark64k23 / 2429421826.528.5
llama.cpp32k24 / 241512141.018.7
llama.cpp64k24 / 243421881.316.7
llama.cpp128k24 / 248541501.813.7
mlx-lm32k24 / 24159202157.821.1
Output without and with speculation, power draw during the agent steps (median)
EngineProseCodeA01A03 Chip (W)°C
Rapid-MLXidenticalidenticalidenticalidentical10173.2
oMLXidenticalidenticalidenticalidentical10074.2
MTPLXdiffersidenticalidenticaldiffers10273.8
mlx-dsparkdiffersdiffersidenticaldiffers9973.8
llama.cppidenticalidenticalidenticalidentical9572.2
mlx-lmidenticalidenticalidenticalidentical9770.8

08 · Limits

What this work does not show

One session, one pass. The session ran once per engine. At temperature 0 the output repeats itself; whether a cache hits may however depend on the state of memory and was not measured repeatedly.

Rapid-MLX in the long-context runs without speculation. Started from a local snapshot, Rapid-MLX turns speculation on only when the draft head is named explicitly, and the long-context runs did not name it. Chapter 05 shows that on this machine it changes neither speed nor output.

Different weights. Every engine ran with the 4-bit version it is built for. Speed and answers therefore compare engine plus weights, not the engine alone.

Capped output. Every agent step ended after at most 1,024 tokens. Some steps hit the limit. That does not affect speed; whether they would have finished is not measured.

One machine, one version. What is measured are the versions named, on 9 and 10 October 2026. The MLX offshoots release almost weekly; a later version may behave differently.