Measured · 7 and 8 October 2026 · 54 runs
Long context on the HP ZBook Ultra G1a
halogen-flash-server 0.17.0 from 32,000 to 990,000 tokens
Authors
Machine
HP ZBook Ultra G1a
- SoC
- Ryzen AI Max+ PRO 395
- GPU
- Radeon 8060S, gfx1151
- Memory
- 128 GiB LPDDR5X-8000, soldered – shared by CPU and GPU as unified memory
- SSD
- Samsung 990 PRO, 4 TB, NVMe on PCIe 4.0 x4
- OS
- Ubuntu 24.04,
140 W power supply,
platform_profile=performance
Model
Qwen3.8-Flash-Next
- Size
- about 180B – measured at 176.9 billion parameters
- Build
- “512×56B”, 48 blocks, 512 experts, 10 of them active per token
- Context
- native up to 262,144 tokens, beyond that YaRN with factor 4
Engine
halogen-flash-server
- Version
- image
0.17.0, Podman - Weights
v2.hgnwithout sidecar, 3.0 bpw; the draft head's lookup table sits beside it and is not held in memory- Flags
HALOGEN_CTXper length up to 1,048,576,HALOGEN_ROPE_YARN=4beyond 262k- Budget
- 3,072 output tokens per task, temperature 0, a fresh container per run
Tasks
Eight questions about a weather log
- Text
- 651 lines of weather-station readings, the same at every length
- Length
- padded with an unrelated maintenance journal
- Kind
- look up, compare, count, sum
- Scored
- six; the two arithmetic tasks test sums, not context
1,398tokens/s
Prompt processing on the first question about a text
23seconds
Reading the whole text once, on the first question
0.5seconds
Reading a follow-up the text itself is already processed
55.3tokens/s
Output when answering follow-up questions
86.3GiB
Memory held of 125.1 GiB; left afterwards: 31.3 GiB
947tokens/s
Prompt processing on the first question about a text
17.4minutes
Reading the whole text once, on the first question
0.9seconds
Reading a follow-up the text itself is already processed
50.8tokens/s
Output when answering follow-up questions
99.9GiB
Memory held of 125.1 GiB; left afterwards: 18.3 GiB
Medians over all runs at the given length. On the first question about a text the machine reads all of it. On every follow-up question it is already processed, and only the new question is added. Chapter 02 shows the lengths in between.
01 · Setup
The same task at every length
To find out whether a model understands a long context, you have to change the length and hold the task still. So the questions always concern the same weather log: 651 lines of weather-station readings, about 32,000 tokens, with eight questions that each have one unambiguous answer. An unrelated maintenance journal before and after it brings the text to each length, with 957,999 tokens at 990k. If the result changes, it can only be the length.
Up to 245k three versions of the weather log with different random values ran per step, one beyond: 21 runs for understanding. For speed, all 54 scored runs of those days count, including those with other documents of the same length; how fast the machine reads and answers depends on the number of tokens, not on their content. Discarded are the warm-up run, the repeat of the same document and the comparison runs with YaRN below 262k that chapter 06 shows. Which run is missing and why is in the evaluation in the repository.
02 · Speed
Read once, then it answers in seconds
On the first question about a text the machine reads 1,398 tokens per second at 32k and 947 at 990k. Every follow-up on the same text then takes at most 0.9 seconds, at every length.
For work on a long document this means: you wait once, 3.4 minutes at 245k, and then keep asking in conversation. An agent that extends its history pays for reading only what is new.
03 · Memory
The pool sets the memory, not the prompt
At startup halogen reserves a pool for a fixed number of positions and holds all of it, however long the text turns out to be. Up to 262k every run used the same pool: 86.3 GiB, whether the prompt had 32,000 or 245,000 tokens. For 990k it is 99.9 GiB of 125.1 GiB.
Beyond 262k halogen ran with a generation budget of 16,384 instead of 32,768 tokens. That lowers the working memory from 12.6 to 8.1 GiB, which is why the pool for 396k holds less than the native one.
At 990k it gets tight. After startup the machine had 18.3 to 18.6 GiB left for everything else, and while reserving the pool the kernel already had to compact memory 342 to 3,499 times depending on the start; with the smaller pools not once. Next to a pool this size nothing large should run on the notebook.
04 · Understanding
Flawless up to 192k, slowly slipping beyond
Up to 245k the model solves 105 of 108 scorable tasks.
Up to 192k, length is no obstacle for this model: not a single wrong scorable answer across all runs up to there. The 3 missing answers all lie at 245k, at the top of the native series, spread over 2 of the 3 versions of the weather log.
05 · Beyond 262k
The first thing to fail is the look-up
Beyond the native limit, too, the result stays high. The same weather log solves 6 of 6 scorable tasks at 396k, 5 of 6 at 594k and 4 of 6 at 990k.
The pattern is not an even decline. The model names a real value from the same log, only from the wrong row: at 245k and at 990k a peak wind of 23 km/h instead of 81, at 594k the sensor ID of the neighbouring row. At 396k it reads both values correctly. So it does not lose the text but the exact place in it.
Reading also gets more expensive beyond the limit: at 990k every token costs 28 percent more time than at 245k. Output is almost untouched by it.
06 · YaRN
At equal length, YaRN changes single answers
Beyond 262,144 tokens the model only gets there with YaRN, a stretch of the position encoding. Whether it costs answers itself shows in a comparison at 245k, where both work: the same document once native and once with YaRN.
Behind this are a block of four runs in alternating order on a long weather log of 4,993 lines, whose repeats give the same answers, and one run with the short weather log inside the maintenance journal. In the long log YaRN finds the unique combination that the model reports as absent natively. In the short one YaRN reads the peak wind correctly, which natively comes from the wrong row, and in turn miscounts a subset that is right natively.
Two documents are not enough to credit YaRN with an advantage. They are enough to conclude that YaRN does not cause the model to look up worse beyond 262k: at equal length it looks up at least as well with YaRN as without.
07 · Every figure
Results per length step
Understanding: the weather log inside the maintenance journal, the six scorable tasks across all runs of the step. Speed: medians across all scored runs of the step. Every chart in this report is computed from these values; the values per run are in docs/kontext-werte-2026-10.json.
| Length | Scaling | Runs | correct |
|---|---|---|---|
| 32k | native | 3 | 18 / 18 |
| 64k | native | 3 | 18 / 18 |
| 96k | native | 3 | 18 / 18 |
| 128k | native | 3 | 18 / 18 |
| 192k | native | 3 | 18 / 18 |
| 245k | native | 3 | 15 / 18 |
| 396k | YaRN 4 | 1 | 6 / 6 |
| 594k | YaRN 4 | 1 | 5 / 6 |
| 990k | YaRN 4 | 1 | 4 / 6 |
| Length | Runs | Reading (s) | Reading (t/s) | Output (t/s) | Follow-up (s) |
|---|---|---|---|---|---|
| 32k | 4 | 23.0 | 1,398 | 55.3 | 0.5 |
| 64k | 6 | 48.5 | 1,323 | 55.8 | 0.4 |
| 80k | 3 | 61.1 | 1,312 | 54.1 | 0.4 |
| 96k | 6 | 73.5 | 1,308 | 54.7 | 0.5 |
| 112k | 3 | 87.0 | 1,289 | 54.2 | 0.4 |
| 128k | 6 | 99.7 | 1,286 | 54.5 | 0.5 |
| 144k | 3 | 114.2 | 1,262 | 53.2 | 0.5 |
| 160k | 3 | 129.4 | 1,238 | 53.6 | 0.5 |
| 176k | 3 | 142.7 | 1,235 | 53.2 | 0.5 |
| 192k | 6 | 156.1 | 1,231 | 53.4 | 0.6 |
| 229k | 1 | 191.2 | 1,198 | 53.0 | 0.5 |
| 245k | 4 | 202.4 | 1,212 | 53.6 | 0.5 |
| 396k | 2 | 351.7 | 1,126 | 52.5 | 0.6 |
| 594k | 2 | 573.3 | 1,036 | 51.3 | 0.8 |
| 990k | 2 | 1,046.2 | 947 | 50.8 | 0.9 |
| Pool | Budget | Starts | Weights | KV-Pool | Working | total | left |
|---|---|---|---|---|---|---|---|
| 393,216 | 32,768 | 53 | 62.9 | 10.8 | 12.6 | 86.3 | 31.3–32.0 |
| 425,984 | 16,384 | 2 | 62.9 | 11.7 | 8.1 | 82.8 | 34.9–35.0 |
| 622,592 | 16,384 | 2 | 62.9 | 17.1 | 8.1 | 88.2 | 29.6–29.6 |
| 1,048,576 | 16,384 | 2 | 62.9 | 28.8 | 8.1 | 99.9 | 18.3–18.6 |
08 · Limits
What this work does not show
One version beyond 262k. Above the native limit one version of the weather log ran per length. At temperature 0 the same document repeats the same answer; the spread lies between documents and is not measured there. That 396k solves every scorable task and 245k does not belongs to this. What holds is the course across all steps, not the single point.
One pool for every length up to 262k. Up to 262k every run used the everyday pool of 393,216 positions. How much memory a smaller pool would save for short texts is not measured.
The smaller arena. From 396k halogen works with a generation budget of 16,384 instead of 32,768 tokens. That changes together with the length and is not measured separately.
One kind of document. The tasks ask about a tabular measurement log. For source code, prose or conversations the limit may lie elsewhere.
One engine, one machine. What is measured is halogen 0.17.0 with checkpoint v2 on the HP ZBook Ultra G1a. The speed holds for this stack; whether llama.cpp understands the same at the same lengths is not measured here.