holistechlabshalogen 0.17 · 32k → 990k
DEEN

Measured · 7 and 8 October 2026 · 54 runs

Long context on the HP ZBook Ultra G1a

halogen-flash-server 0.17.0 from 32,000 to 990,000 tokens

Authors

Dipl.-Ing. Sören GebbertInstitut für holistische Technologieforschung GmbH Claude Code Opus 5.5Anthropic

Machine

HP ZBook Ultra G1a

SoC
Ryzen AI Max+ PRO 395
GPU
Radeon 8060S, gfx1151
Memory
128 GiB LPDDR5X-8000, soldered – shared by CPU and GPU as unified memory
SSD
Samsung 990 PRO, 4 TB, NVMe on PCIe 4.0 x4
OS
Ubuntu 24.04, 140 W power supply, platform_profile=performance

Model

Qwen3.8-Flash-Next

Size
about 180B – measured at 176.9 billion parameters
Build
“512×56B”, 48 blocks, 512 experts, 10 of them active per token
Context
native up to 262,144 tokens, beyond that YaRN with factor 4

Engine

halogen-flash-server

Version
image 0.17.0, Podman
Weights
v2.hgn without sidecar, 3.0 bpw; the draft head's lookup table sits beside it and is not held in memory
Flags
HALOGEN_CTX per length up to 1,048,576, HALOGEN_ROPE_YARN=4 beyond 262k
Budget
3,072 output tokens per task, temperature 0, a fresh container per run

Tasks

Eight questions about a weather log

Text
651 lines of weather-station readings, the same at every length
Length
padded with an unrelated maintenance journal
Kind
look up, compare, count, sum
Scored
six; the two arithmetic tasks test sums, not context
32k Context32,000 tokens in the prompt

1,398tokens/s

Prompt processing on the first question about a text

23seconds

Reading the whole text once, on the first question

0.5seconds

Reading a follow-up the text itself is already processed

55.3tokens/s

Output when answering follow-up questions

86.3GiB

Memory held of 125.1 GiB; left afterwards: 31.3 GiB

990k Context990,000 tokens in the prompt

947tokens/s

Prompt processing on the first question about a text

17.4minutes

Reading the whole text once, on the first question

0.9seconds

Reading a follow-up the text itself is already processed

50.8tokens/s

Output when answering follow-up questions

99.9GiB

Memory held of 125.1 GiB; left afterwards: 18.3 GiB

Medians over all runs at the given length. On the first question about a text the machine reads all of it. On every follow-up question it is already processed, and only the new question is added. Chapter 02 shows the lengths in between.

01 · Setup

The same task at every length

To find out whether a model understands a long context, you have to change the length and hold the task still. So the questions always concern the same weather log: 651 lines of weather-station readings, about 32,000 tokens, with eight questions that each have one unambiguous answer. An unrelated maintenance journal before and after it brings the text to each length, with 957,999 tokens at 990k. If the result changes, it can only be the length.

Up to 245k three versions of the weather log with different random values ran per step, one beyond: 21 runs for understanding. For speed, all 54 scored runs of those days count, including those with other documents of the same length; how fast the machine reads and answers depends on the number of tokens, not on their content. Discarded are the warm-up run, the repeat of the same document and the comparison runs with YaRN below 262k that chapter 06 shows. Which run is missing and why is in the evaluation in the repository.

02 · Speed

Read once, then it answers in seconds

On the first question about a text the machine reads 1,398 tokens per second at 32k and 947 at 990k. Every follow-up on the same text then takes at most 0.9 seconds, at every length.

By 990k, reading loses just under a third of its speed Tokens read per second on the first question about a text, one dot per run. HP ZBook Ultra G1a · Ryzen AI Max+ PRO 395 · Qwen3.8-Flash-Next · halogen 0.17.0, checkpoint v2 temperature 0 · eight questions with one clear answer · native up to 262,144 tokens, YaRN 4 beyond · measured 7 and 8 Oct 2026 32k 64k 128k 256k 512k 1M beyond the native limit: YaRN 4 native limit Length of the prompt in tokens (logarithmic) 0 250 500 750 1,000 1,250 Tokens per second 1,398 1,323 1,312 1,308 1,289 1,286 1,262 1,238 1,235 1,231 1,198 1,212 1,126 1,036 947
Reading loses 13 percent of its speed up to the native limit and 32 by 990k. Every dot is one run, the line the median per step.
Wait once, then it answers in seconds Reading time on the first question and on every follow-up, logarithmic. HP ZBook Ultra G1a · Ryzen AI Max+ PRO 395 · Qwen3.8-Flash-Next · halogen 0.17.0, checkpoint v2 temperature 0 · eight questions with one clear answer · native up to 262,144 tokens, YaRN 4 beyond · measured 7 and 8 Oct 2026 follow-up questions first question 32k 64k 128k 256k 512k 1M beyond the native limit: YaRN 4 native limit Length of the prompt in tokens (logarithmic) 1 s 10 s 1 min 5 min 20 min Reading time 23.0 s 48.5 s 1.0 min 1.2 min 1.5 min 1.7 min 1.9 min 2.2 min 2.4 min 2.6 min 3.2 min 3.4 min 5.9 min 9.6 min 17.4 min 0.5 s 0.4 s 0.4 s 0.5 s 0.4 s 0.5 s 0.5 s 0.5 s 0.5 s 0.6 s 0.5 s 0.5 s 0.6 s 0.8 s 0.9 s
The wait comes with the first question. At 990k reading takes 17.4 minutes; after that the text is processed, and every further question costs only its own tokens.
Generation stays almost as fast Tokens generated per second when answering follow-up questions, median per run. HP ZBook Ultra G1a · Ryzen AI Max+ PRO 395 · Qwen3.8-Flash-Next · halogen 0.17.0, checkpoint v2 temperature 0 · eight questions with one clear answer · native up to 262,144 tokens, YaRN 4 beyond · measured 7 and 8 Oct 2026 32k 64k 128k 256k 512k 1M beyond the native limit: YaRN 4 native limit Length of the prompt in tokens (logarithmic) 0 10 20 30 40 50 60 Tokens per second 55.3 55.8 54.1 54.7 54.2 54.5 53.2 53.6 53.2 53.4 53.0 53.6 52.5 51.3 50.8
Output barely depends on the length. From 55.3 to 50.8 tokens per second at thirty times the context. Single answers stalled by the kernel's memory compaction are left out; they change no token.

For work on a long document this means: you wait once, 3.4 minutes at 245k, and then keep asking in conversation. An agent that extends its history pays for reading only what is new.

03 · Memory

The pool sets the memory, not the prompt

At startup halogen reserves a pool for a fixed number of positions and holds all of it, however long the text turns out to be. Up to 262k every run used the same pool: 86.3 GiB, whether the prompt had 32,000 or 245,000 tokens. For 990k it is 99.9 GiB of 125.1 GiB.

The pool sets the memory, not the prompt What halogen holds after starting, per configuration; the engine reports these figures itself at startup. HP ZBook Ultra G1a · Ryzen AI Max+ PRO 395 · Qwen3.8-Flash-Next · halogen 0.17.0, checkpoint v2 temperature 0 · eight questions with one clear answer · native up to 262,144 tokens, YaRN 4 beyond · measured 7 and 8 Oct 2026 working memory KV pool weights 0 20 40 60 80 100 120 gigabytes (GiB) machine: 125.1 GiB 32k to 245k pool 393,216 positions 62.9 10.8 12.6 86.3 GiB left afterwards: 31.3–32.0 GiB 396k pool 425,984 positions 62.9 11.7 8.1 82.8 GiB left afterwards: 34.9–35.0 GiB 594k pool 622,592 positions 62.9 17.1 8.1 88.2 GiB left afterwards: 29.6 GiB 990k pool 1,048,576 positions 62.9 28.8 8.1 99.9 GiB left afterwards: 18.3–18.6 GiB
The weights never change, the KV pool grows with the positions. 10.8 GiB for 393,216 positions, 28.8 GiB for 1,048,576. The engine reports these figures itself at startup, here from 59 starts of dedicated measurement containers.

Beyond 262k halogen ran with a generation budget of 16,384 instead of 32,768 tokens. That lowers the working memory from 12.6 to 8.1 GiB, which is why the pool for 396k holds less than the native one.

At 990k it gets tight. After startup the machine had 18.3 to 18.6 GiB left for everything else, and while reserving the pool the kernel already had to compact memory 342 to 3,499 times depending on the start; with the smaller pools not once. Next to a pool this size nothing large should run on the notebook.

04 · Understanding

Flawless up to 192k, slowly slipping beyond

Up to 245k the model solves 105 of 108 scorable tasks.

Flawless up to 192k, slowly slipping beyond Share of scorable tasks solved correctly; three versions of the weather log up to 245k, one beyond. HP ZBook Ultra G1a · Ryzen AI Max+ PRO 395 · Qwen3.8-Flash-Next · halogen 0.17.0, checkpoint v2 temperature 0 · eight questions with one clear answer · native up to 262,144 tokens, YaRN 4 beyond · measured 7 and 8 Oct 2026 32k 64k 128k 256k 512k 1M beyond the native limit: YaRN 4 native limit Length of the prompt in tokens (logarithmic) 0 % 25 % 50 % 75 % 100 % Share correct 100 % 100 % 100 % 100 % 100 % 83 % 100 % 83 % 67 %
Up to 192k the line stays up. The first errors come at 245k, just below the native limit; beyond, with YaRN, it lies between 67 and 100 percent.

Up to 192k, length is no obstacle for this model: not a single wrong scorable answer across all runs up to there. The 3 missing answers all lie at 245k, at the top of the native series, spread over 2 of the 3 versions of the weather log.

05 · Beyond 262k

The first thing to fail is the look-up

Beyond the native limit, too, the result stays high. The same weather log solves 6 of 6 scorable tasks at 396k, 5 of 6 at 594k and 4 of 6 at 990k.

The first thing to fail is the look-up The same weather log at every length: filled dot correct, open one wrong. HP ZBook Ultra G1a · Ryzen AI Max+ PRO 395 · Qwen3.8-Flash-Next · halogen 0.17.0, checkpoint v2 temperature 0 · eight questions with one clear answer · native up to 262,144 tokens, YaRN 4 beyond · measured 7 and 8 Oct 2026 32k 64k 96k 128k 192k 245k 396k 594k 990k Read a value from a named row (wind) Read a value from a named row (sensor) Distance between two far-apart entries Tell a near-duplicate apart Find a unique combination Find the maximum of a subset Count a subset Sum a subset
The harder tasks hold out longer. Finding a combination, determining a maximum, spotting a near-duplicate, measuring the distance between two entries: correct at every length. Reading the value of a named row fails at 245k, 594k and 990k.

The pattern is not an even decline. The model names a real value from the same log, only from the wrong row: at 245k and at 990k a peak wind of 23 km/h instead of 81, at 594k the sensor ID of the neighbouring row. At 396k it reads both values correctly. So it does not lose the text but the exact place in it.

Reading also gets more expensive beyond the limit: at 990k every token costs 28 percent more time than at 245k. Output is almost untouched by it.

06 · YaRN

At equal length, YaRN changes single answers

Beyond 262,144 tokens the model only gets there with YaRN, a stretch of the position encoding. Whether it costs answers itself shows in a comparison at 245k, where both work: the same document once native and once with YaRN.

At equal length, YaRN changes single answers Each pair the same 245,000-token document, native and with YaRN 4; six scorable tasks. HP ZBook Ultra G1a · Ryzen AI Max+ PRO 395 · Qwen3.8-Flash-Next · halogen 0.17.0, checkpoint v2 temperature 0 · eight questions with one clear answer · native up to 262,144 tokens, YaRN 4 beyond · measured 7 and 8 Oct 2026 YaRN 4 native 0 50 100 150 200 250 300 Reading on the first question, in seconds long weather log, run 1 207 s · 3 / 6 correct 210 s · 4 / 6 correct long weather log, run 2 207 s · 3 / 6 correct 209 s · 4 / 6 correct short log in the maintenance journal 203 s · 5 / 6 correct 204 s · 6 / 6 correct
In every pair YaRN solves one more scorable task. Reading takes at most 1 percent longer with YaRN.

Behind this are a block of four runs in alternating order on a long weather log of 4,993 lines, whose repeats give the same answers, and one run with the short weather log inside the maintenance journal. In the long log YaRN finds the unique combination that the model reports as absent natively. In the short one YaRN reads the peak wind correctly, which natively comes from the wrong row, and in turn miscounts a subset that is right natively.

Two documents are not enough to credit YaRN with an advantage. They are enough to conclude that YaRN does not cause the model to look up worse beyond 262k: at equal length it looks up at least as well with YaRN as without.

07 · Every figure

Results per length step

Understanding: the weather log inside the maintenance journal, the six scorable tasks across all runs of the step. Speed: medians across all scored runs of the step. Every chart in this report is computed from these values; the values per run are in docs/kontext-werte-2026-10.json.

Understanding
Length Scaling Runs correct
32knative318 / 18
64knative318 / 18
96knative318 / 18
128knative318 / 18
192knative318 / 18
245knative315 / 18
396kYaRN 416 / 6
594kYaRN 415 / 6
990kYaRN 414 / 6
Speed
Length Runs Reading (s) Reading (t/s) Output (t/s) Follow-up (s)
32k423.01,39855.30.5
64k648.51,32355.80.4
80k361.11,31254.10.4
96k673.51,30854.70.5
112k387.01,28954.20.4
128k699.71,28654.50.5
144k3114.21,26253.20.5
160k3129.41,23853.60.5
176k3142.71,23553.20.5
192k6156.11,23153.40.6
229k1191.21,19853.00.5
245k4202.41,21253.60.5
396k2351.71,12652.50.6
594k2573.31,03651.30.8
990k21,046.294750.80.9
Memory
Pool Budget Starts Weights KV-Pool Working total left
393,21632,7685362.910.812.686.331.3–32.0
425,98416,384262.911.78.182.834.9–35.0
622,59216,384262.917.18.188.229.6–29.6
1,048,57616,384262.928.88.199.918.3–18.6

08 · Limits

What this work does not show

One version beyond 262k. Above the native limit one version of the weather log ran per length. At temperature 0 the same document repeats the same answer; the spread lies between documents and is not measured there. That 396k solves every scorable task and 245k does not belongs to this. What holds is the course across all steps, not the single point.

One pool for every length up to 262k. Up to 262k every run used the everyday pool of 393,216 positions. How much memory a smaller pool would save for short texts is not measured.

The smaller arena. From 396k halogen works with a generation budget of 16,384 instead of 32,768 tokens. That changes together with the length and is not measured separately.

One kind of document. The tasks ask about a tabular measurement log. For source code, prose or conversations the limit may lie elsewhere.

One engine, one machine. What is measured is halogen 0.17.0 with checkpoint v2 on the HP ZBook Ultra G1a. The speed holds for this stack; whether llama.cpp understands the same at the same lengths is not measured here.