Measured · 7 to 10 October 2026 · 48 runs each
Qwen3.8-Flash-Next on Apple M1 and Strix Halo
sushi 1.2.1 · MacBook Pro M1 Max against halogen 0.17.0 · HP ZBook Ultra G1a
Authors
Machine and engine
MacBook Pro M1 Max
- SoC
- Apple M1 Max, 10 CPU cores, 32 GPU cores
- Memory
- 64 GiB unified memory, up to 56 GiB for the GPU
(
iogpu.wired_limit_mb) - OS
- macOS 27.0.1, on mains power
- Engine
- sushi 1.2.1, MTP on, KV cache 8 bit (defaults)
- Weights
beamster/Qwen3.8-Flash-Next-Sushi-2.6bpw
Machine and engine
HP ZBook Ultra G1a
- SoC
- Ryzen AI Max+ PRO 395, Radeon 8060S
- Memory
- 128 GiB LPDDR5X-8000 as unified memory
- OS
- Ubuntu 24.04,
platform_profile=performance - Engine
- halogen-flash-server 0.17.0, MTP on (default)
- Weights
v2.hgn, 3.0 bpw
Protocol
Eight questions about a weather log
- Lengths
- 32,000 to 245,000 tokens, native
- Documents
- byte-identical on both machines
- Budget
- 3,072 output tokens, at most 2,048 of them thinking, temperature 0
- Scored
- six tasks; the two arithmetic tasks test sums, not context
337tokens/s
MacBook Pro M1 Max · sushiprompt processing on the first question
12.1minutes
Reading the whole text once, on the first question
27.0seconds
Reading a follow-up the text is already processed
33.0tokens/s
Output when answering follow-up questions
1,212tokens/s
HP ZBook Ultra G1a · halogenprompt processing on the first question
3.4minutes
Reading the whole text once, on the first question
0.5seconds
Reading a follow-up the text is already processed
53.6tokens/s
Output when answering follow-up questions
Medians over all runs at this length. On the first question about a text the machine reads all of it; on every follow-up only the new question is added. Chapter 02 shows the lengths in between.
01 · Setup
The same protocol on two machines
The MacBook Pro M1 Max ran the same 48 runs as the HP ZBook Ultra G1a in the October report, as far as they lie below the native limit of 262,144 tokens: the same documents, regenerated on the Mac and checked byte for byte against the ZBook's raw data, the same eight questions, the same scoring, the same order. Every run got a freshly started server without a disk cache.
Both engines run with their defaults, including their multi-token prediction (MTP). On the Mac it is verified on two requests that sushi produces byte for byte the same output with MTP as without. halogen closes the thinking block once only 1,024 of the 3,072 tokens remain; sushi gets a thinking budget of 2,048 tokens instead.
What is compared are two stacks, not two chips: machine, engine and weights all differ at once. Chapter 06 says what follows from that.
02 · Speed
The ZBook reads faster and answers faster
On the first question about a text the ZBook reads 3.5 to 3.8 times as many tokens per second as the M1 Max. A follow-up takes at most 0.6 seconds on the ZBook and 13 to 27 seconds on the M1 Max.
On the ZBook the engine itself measures the times, on the Mac the measuring script measures them on the stream: reading up to the first token, output after it.
03 · Understanding
In the fixed core both find almost every answer
Across all lengths up to 245k the M1 Max with sushi solves 100 of 108 scorable tasks, the ZBook with halogen 105 of 108.
As in the context report, only the fixed core is scored: the same weather log at every length, padded with an unrelated maintenance journal. Three versions with different random values ran per length.
04 · Memory
At 245k memory gets tight on the M1 Max
Of 5 runs at 245k sushi refused question 6 in one: it would have needed about 8,507 MB of GPU memory, 8,005 MB were free. The repeat of the same run went through completely, and only it is scored. So the limit lies just below this length and depends on the state of the engine, not firmly on the number of tokens.
The M1 Max has 64 GiB, 56 GiB of which are available to the GPU. The ZBook has 128 GiB. The memory figures of the two engines are not set side by side here: macOS and Linux count held memory differently.
05 · Every figure
Results per length step
Speed: medians across all scored runs of the step, reading on the first question. Understanding: the fixed core, six scorable tasks per run. The values per run are in docs/kontext-werte-m1.json and docs/kontext-werte-2026-10.json.
| Length | Runs | Reading M1 (t/s) | Reading ZBook (t/s) | Follow-up M1 (s) | Follow-up ZBook (s) | Output M1 (t/s) | Output ZBook (t/s) |
|---|---|---|---|---|---|---|---|
| 32k | 4 | 363 | 1,398 | 21.7 | 0.5 | 35.5 | 55.3 |
| 64k | 6 | 362 | 1,323 | 20.1 | 0.4 | 37.5 | 55.8 |
| 80k | 3 | 360 | 1,312 | 19.2 | 0.4 | 33.4 | 54.1 |
| 96k | 6 | 359 | 1,308 | 19.0 | 0.5 | 34.7 | 54.7 |
| 112k | 3 | 357 | 1,289 | 18.0 | 0.4 | 34.8 | 54.2 |
| 128k | 6 | 356 | 1,286 | 17.1 | 0.5 | 34.8 | 54.5 |
| 144k | 3 | 355 | 1,262 | 16.2 | 0.5 | 31.7 | 53.2 |
| 160k | 3 | 353 | 1,238 | 14.7 | 0.5 | 32.5 | 53.6 |
| 176k | 3 | 351 | 1,235 | 14.0 | 0.5 | 35.0 | 53.2 |
| 192k | 6 | 350 | 1,231 | 13.1 | 0.6 | 34.3 | 53.4 |
| 229k | 1 | 343 | 1,198 | 27.2 | 0.5 | 34.8 | 53.0 |
| 245k | 4 | 337 | 1,212 | 27.0 | 0.5 | 33.0 | 53.6 |
| Length | Runs | correct M1 | correct ZBook |
|---|---|---|---|
| 32k | 3 | 15 / 18 | 18 / 18 |
| 64k | 3 | 16 / 18 | 18 / 18 |
| 96k | 3 | 18 / 18 | 18 / 18 |
| 128k | 3 | 17 / 18 | 18 / 18 |
| 192k | 3 | 18 / 18 | 18 / 18 |
| 245k | 3 | 16 / 18 | 15 / 18 |
06 · Limits
What this comparison does not show
No comparison of chips. Machine, engine and weights all differ at once. How fast halogen would be on a Mac or sushi on Strix Halo is not measured.
Different weights. halogen computes with v2.hgn (3.0 bpw), sushi with beamster/Qwen3.8-Flash-Next-Sushi-2.6bpw. Differences in understanding may come from the weights.
Only up to the native limit. Nothing beyond 262,144 tokens is measured on the M1 Max; memory already gets tight at 245k.
One kind of document, one session. The tasks ask about a tabular measurement log, one question after another. How both machines fare with source code, several sessions or an agent is not measured here.