FINDINGS / WHAT THE CAPTURE MEASURED

Eleven things
the capture measured.

One capture series on the instrumented K3 research runtime: ten prompt domains, thirty-two prompts each, 320 prompts in total, with exact expert selections and gate weights recorded for every token at every routed depth. Every number below is a single-model, single-capture result until replicated, and every one is published with its direction, its anchors and its caveat.

Measured resultsPublished 6 September 2026Depth as a fraction of instrumented depthMagnitudes as ratios against stated baselinesOpen the depth explorer ↗
00 / HOW TO READ THESE

The number, its direction,
and its control.

A bare figure is not a result. Each finding below states what the metric measures, which direction would have meant something different, and what it is anchored against: a chance level, a shuffled control, a stated baseline. Where a control exists it is drawn on the chart, not left to a footnote.

01

Depth is a fraction

0.0 is the first captured boundary and 1.0 the last. No absolute layer counts or architecture dimensions are published.

02

Magnitudes are ratios

Every magnitude is a ratio against a stated baseline: the median row, a one-token window, a random set of the same size.

03

One capture, one model

Every finding is single-model and single-capture until replicated. Finding 4’s curve beyond 151 is a fitted extrapolation and is drawn as one.

An observation is not a cause.

None of these findings establish intent, causation or safety. Finding 8 shows where attention landed, not why. What they did change is stated per finding: which engineering lines were closed and which were opened.

01 / ROUTING BREADTHMEASURED

A domain's routing stays broad at every depth

26%–77%of the routed expert pool carries 95% of a layer’s routing mass, depending on depth. Typical: about half.

Across ten prompt domains, the experts carrying 95% of a layer's routing mass are roughly half the pool at most depths; the narrowest layers still use about a quarter.

Share of the routed expert pool needed to carry 95% of routing mass at each captured depth. Line: median across the ten domains. Band: narrowest to broadest domain.
View the data
A domain's routing stays broad at every depth — published values
DepthMedianMinMax
0%77%73%78%
1%69%65%72%
2%53%44%60%
3%50%46%53%
4%62%57%67%
5%74%71%78%
7%55%50%63%
8%59%53%68%
9%56%50%61%
10%60%56%67%
11%59%55%66%
12%39%34%41%
13%70%65%72%
14%66%61%70%
15%63%56%68%
16%59%49%66%
18%58%47%64%
19%55%46%60%
20%55%45%60%
21%59%53%62%
22%64%61%68%
23%62%58%66%
24%55%52%57%
25%68%64%70%
26%68%66%71%
27%63%58%67%
29%64%59%66%
30%68%66%71%
31%59%54%61%
32%59%56%62%
33%58%55%62%
34%64%61%67%
35%63%61%66%
36%60%57%63%
37%56%54%59%
38%59%56%63%
40%48%42%54%
41%46%42%49%
42%49%47%57%
43%49%45%53%
44%48%42%51%
45%37%32%45%
46%41%36%48%
47%42%39%50%
48%38%33%45%
49%34%31%39%
51%39%37%46%
52%48%44%52%
53%36%28%44%
54%39%34%43%
55%44%39%47%
56%34%27%44%
57%26%17%36%
58%30%22%38%
59%31%24%37%
60%26%19%41%
62%29%25%41%
63%33%28%40%
64%28%23%39%
65%45%43%50%
66%55%51%58%
67%54%49%59%
68%52%46%57%
69%47%41%53%
70%49%46%57%
71%51%49%57%
73%45%41%53%
74%50%47%56%
75%51%47%58%
76%55%49%62%
77%54%51%61%
78%48%46%52%
79%36%31%41%
80%47%40%50%
81%53%47%57%
82%35%29%45%
84%37%34%46%
85%41%35%47%
86%38%32%51%
87%47%42%57%
88%42%39%47%
89%50%47%59%
90%48%45%56%
91%38%27%54%
92%40%31%50%
93%34%22%48%
95%38%35%49%
96%37%30%52%
97%38%32%52%
98%39%29%52%
99%44%37%56%
100%48%45%57%
The question
Does a domain concentrate on a small, cacheable set of experts?
The answer
No. Routing is broad by design, not by accident.
What it changed
Set the residency budget and closed the expert-sparsification line before any GPU time was spent.
Caveat
Single capture series, one model. Not replicated.
02 / DOMAIN OVERLAPMEASURED

Domain routing fingerprints overlap, but not fully

0.55mean Jaccard overlap between any two domains’ 95%-mass expert sets (0.45 to 0.67). 1.0 would mean indistinguishable.

The 95%-mass expert sets of any two domains share about half their members: off-diagonal Jaccard mean 0.55, min 0.45, max 0.67.

Jaccard overlap between the 95%-mass expert sets of each pair of domains, averaged over layers. Diagonal cells are the same domain and are left blank.
View the data
Domain routing fingerprints overlap, but not fully — published values
DomainBiology & medicinePython codeGPU kernelsGeneral chatLegal contractsMath proofsFictionRust/C systemsStorage systemsTranslation
Biology & medicine0.550.540.540.600.570.510.540.600.54
Python code0.550.640.530.590.580.530.670.670.55
GPU kernels0.540.640.450.530.530.450.650.670.48
General chat0.540.530.450.560.510.590.460.530.57
Legal contracts0.600.590.530.560.550.550.560.600.58
Math proofs0.570.580.530.510.550.500.540.570.54
Fiction0.510.530.450.590.550.500.470.510.59
Rust/C systems0.540.670.650.460.560.540.470.650.50
Storage systems0.600.670.670.530.600.570.510.650.52
Translation0.540.550.480.570.580.540.590.500.52
The question
Should resident expert sets be built per domain, or is one shared set nearly as good?
The answer
One shared set is nearly as good. The residual differences are what a domain-aware prefetch could use.
What it changed
Removed per-domain residency from the design.
Caveat
Averaged over layers. Single capture series.
03 / EXPERT REUSEMEASURED

Consecutive tokens reuse experts far above chance

29%of the previous token’s experts are re-activated by the next token, against a chance level of 1.8%: a 16.5× lift.

A token re-activates 29.4% of the previous token's experts, against a chance level of 1.8% for random sets of the same size — a 16x lift.

Share of the previous token’s active experts re-activated by the next token at each captured depth. The reference line is the chance level for two random sets of the same size.
View the data
Consecutive tokens reuse experts far above chance — published values
DepthReuseChance
0%10%1.8%
1%23%1.8%
2%30%1.8%
3%37%1.8%
4%20%1.8%
5%8.5%1.8%
7%19%1.8%
8%23%1.8%
9%33%1.8%
10%32%1.8%
11%31%1.8%
12%34%1.8%
13%28%1.8%
14%29%1.8%
15%30%1.8%
16%30%1.8%
18%31%1.8%
19%33%1.8%
20%34%1.8%
21%29%1.8%
22%24%1.8%
23%30%1.8%
24%36%1.8%
25%25%1.8%
26%23%1.8%
27%28%1.8%
29%26%1.8%
30%26%1.8%
31%33%1.8%
32%32%1.8%
33%29%1.8%
34%27%1.8%
35%27%1.8%
36%28%1.8%
37%27%1.8%
38%29%1.8%
40%37%1.8%
41%34%1.8%
42%32%1.8%
43%35%1.8%
44%36%1.8%
45%39%1.8%
46%39%1.8%
47%35%1.8%
48%43%1.8%
49%44%1.8%
51%40%1.8%
52%32%1.8%
53%36%1.8%
54%36%1.8%
55%37%1.8%
56%38%1.8%
57%40%1.8%
58%40%1.8%
59%40%1.8%
60%41%1.8%
62%40%1.8%
63%36%1.8%
64%37%1.8%
65%32%1.8%
66%27%1.8%
67%25%1.8%
68%23%1.8%
69%24%1.8%
70%23%1.8%
71%22%1.8%
73%29%1.8%
74%24%1.8%
75%21%1.8%
76%20%1.8%
77%20%1.8%
78%28%1.8%
79%28%1.8%
80%23%1.8%
81%20%1.8%
82%25%1.8%
84%23%1.8%
85%21%1.8%
86%23%1.8%
87%20%1.8%
88%29%1.8%
89%22%1.8%
90%23%1.8%
91%26%1.8%
92%28%1.8%
93%25%1.8%
95%33%1.8%
96%28%1.8%
97%27%1.8%
98%42%1.8%
99%32%1.8%
100%25%1.8%
The question
How much of the next token's expert traffic is predictable from what was just loaded?
The answer
Enough to prefetch on. Persistence, not prediction, is the usable signal.
What it changed
Made prefetch-by-persistence and a recently-used resident set worth building; the depth profile says where it pays most.
Caveat
Chance level is computed for two random sets of the same size drawn from the same pool.
04 / WINDOW SHARINGMEASURED · EXTRAPOLATION FLAGGED

Relative bytes per token fall as the ingestion window grows

4.6×fewer relative bytes per token at a window of 151, measured. The scrambled-routing control reaches 2.9×.

When a window of tokens shares each layer's expert loads, relative bytes per token fall 4.6x by a window of 151. A control that keeps the same counts but scrambles which experts are chosen reaches only 2.9x, so most of the gain is real reuse rather than the arithmetic of bigger windows.

Relative bytes per token saved when a window of tokens shares each layer’s expert loads, against one-token windows. Solid markers are measurements. The dashed continuation beyond 151 is a fitted power law and is not a measurement.
View the data
Relative bytes per token fall as the ingestion window grows — published values
WindowSaving (measured)Scrambled-routing control
11.00×1.00×
21.16×1.01×
41.37×1.03×
81.66×1.06×
161.99×1.14×
192.11×1.17×
322.41×1.29×
382.64×1.36×
642.96×1.58×
763.44×1.81×
1514.58×2.89×
256 (fitted, not measured)5.11×
512 (fitted, not measured)9.14×
1024 (fitted, not measured)18.29×
The question
How large should an ingestion chunk be before the byte saving flattens out?
The answer
Measured to a window of 151. Beyond that the curve is a fitted power law, not a measurement.
What it changed
Set the prefill chunk size for ingestion and priced the shared-window read path.
Caveat
Windows above 151 are EXTRAPOLATED from a fitted power law and must be labelled as such wherever they appear. Do not present them as measurements.
05 / BACKWARD CONCENTRATIONMEASURED

Backward-pass energy is extremely concentrated

13%of layers hold 90% of gradient energy; 3.0% of sequence positions hold 90% of position energy. Nine objectives, one prompt.

Across nine objectives, 13% of layers hold 90% of the gradient energy and 3% of sequence positions hold 90% of the position energy. Half of the weight-gradient work carries under 2% of the gradient mass.

Cumulative share of gradient energy against the share of layers (with min–max band across nine objectives) and of sequence positions, each sorted by energy. The reference line marks 90%.
View the data
Backward-pass energy is extremely concentrated — published values
Share of layers (sorted)Cumulative energy, meanMinMax
1.1%19%16%27%
2.2%34%31%41%
3.3%47%43%51%
4.3%56%51%60%
5.4%63%57%66%
6.5%69%62%71%
7.6%73%67%75%
8.7%77%71%80%
9.8%81%76%83%
11%85%80%86%
12%87%83%89%
13%90%86%91%
14%92%89%93%
15%94%91%95%
16%95%93%96%
17%96%94%97%
18%97%95%97%
20%97%96%98%
21%98%97%98%
22%98%97%98%
23%98%98%99%
24%99%98%99%
25%99%98%99%
26%99%99%99%
27%99%99%99%
28%99%99%99%
29%99%99%100%
30%100%99%100%
32%100%99%100%
33%100%99%100%
34%100%100%100%
35%100%100%100%
36%100%100%100%
37%100%100%100%
38%100%100%100%
39%100%100%100%
40%100%100%100%
41%100%100%100%
42%100%100%100%
43%100%100%100%
45%100%100%100%
46%100%100%100%
47%100%100%100%
48%100%100%100%
49%100%100%100%
50%100%100%100%
51%100%100%100%
52%100%100%100%
53%100%100%100%
54%100%100%100%
55%100%100%100%
57%100%100%100%
58%100%100%100%
59%100%100%100%
60%100%100%100%
61%100%100%100%
62%100%100%100%
63%100%100%100%
64%100%100%100%
65%100%100%100%
66%100%100%100%
67%100%100%100%
68%100%100%100%
70%100%100%100%
71%100%100%100%
72%100%100%100%
73%100%100%100%
74%100%100%100%
75%100%100%100%
76%100%100%100%
77%100%100%100%
78%100%100%100%
79%100%100%100%
80%100%100%100%
82%100%100%100%
83%100%100%100%
84%100%100%100%
85%100%100%100%
86%100%100%100%
87%100%100%100%
88%100%100%100%
89%100%100%100%
90%100%100%100%
91%100%100%100%
92%100%100%100%
93%100%100%100%
95%100%100%100%
96%100%100%100%
97%100%100%100%
98%100%100%100%
99%100%100%100%
100%100%100%100%
0.7% of positions59%
3.3% of positions91%
5.9% of positions94%
8.6% of positions96%
11% of positions96%
14% of positions97%
16% of positions98%
19% of positions98%
22% of positions98%
24% of positions99%
27% of positions99%
30% of positions99%
32% of positions99%
35% of positions99%
38% of positions99%
40% of positions100%
43% of positions100%
45% of positions100%
48% of positions100%
51% of positions100%
53% of positions100%
56% of positions100%
59% of positions100%
61% of positions100%
64% of positions100%
66% of positions100%
69% of positions100%
72% of positions100%
74% of positions100%
77% of positions100%
80% of positions100%
82% of positions100%
85% of positions100%
88% of positions100%
90% of positions100%
93% of positions100%
95% of positions100%
98% of positions100%
100% of positions100%
The question
When a backward pass is expensive, which weight gradients are worth computing at all?
The answer
A small, stable minority. The cold set is the same across every objective tested.
What it changed
Justified selective backward.
Caveat
One prompt, nine objectives. The objectives are not named publicly.
06 / CONTRIBUTION REDUNDANCYMEASURED

You cannot drop experts: contributions are many and mildly redundant

9.7%of a layer’s expert-write norm is the most any single active expert carries. Keeping the top half recovers 71% (mean) and 62% (10th percentile).

The largest single active expert carries only 9.7% of a layer's expert-write norm. Keeping the top half of active experts still loses about 29% of the write.

Share of a layer’s expert-write norm recovered when only the top share of active experts is kept. Mean and 10th percentile over 600 sampled sites.
View the data
You cannot drop experts: contributions are many and mildly redundant — published values
Share of active experts keptWrite norm recovered, meanWrite norm recovered, p10
6.3%23%12%
13%34%21%
19%42%29%
25%49%36%
31%55%43%
38%61%50%
44%66%56%
50%71%62%
56%75%67%
63%80%73%
69%84%78%
75%87%83%
81%91%87%
88%94%92%
94%97%96%
100%100%100%
The question
Could the model run with fewer active experts per token?
The answer
No. Contributions are spread and only mildly redundant.
What it changed
Closed the sparsification line. Bytes must be cut by residency and I/O, not by skipping computation.
Caveat
600 sampled sites. Mean and 10th percentile both reported; use the p10 when stating a worst case.
07 / ROUTE PREDICTIONMEASURED

The best route predictor is the previous token, not the hidden state

0.28recall from the previous token’s route, against 0.20 for a nearest neighbour on the hidden state, 0.13 for popularity and a shuffled control at 0.07.

Predicting the next layer's active experts one step ahead: the previous token's route recalls 0.28, a nearest-neighbour on the exact pre-router hidden state 0.20, expert popularity 0.13, and a shuffled control 0.07. Accuracy is flat out to four layers ahead, so prefetch lead time is free.

Recall of the next layer’s active experts at a one-token budget, by predictor. The shuffled control is the vacuity floor.
View the data
The best route predictor is the previous token, not the hidden state — published values
PredictorRecall at a one-token budget
previous token0.28
nearest neighbour on exact hidden state0.20
expert popularity0.13
shuffled control0.07
The question
Is a learned route predictor worth building for prefetch, or does the trivial baseline already win?
The answer
The trivial baseline wins. The state carries less routing information than the previous token's choice.
What it changed
Stopped the state-based route-prediction line and moved the value of prediction into retention decisions instead of speculative fetching.
Caveat
Chronological train/test split. The shuffled control is included precisely so the numbers can be read against a floor.
08 / ATTENTION PROVENANCEMEASURED

Attention addresses documents at shallow depth and reads them deep

96%of the answer token’s attention mass sits on one anchor row per document at the shallowest captured depth; 3.0% at the deepest.

With sixteen separately prepared documents attached, shallow attention layers put over 95% of the answer token's mass on one anchor row per document. Deep layers read content instead, and split between the right document and the other fifteen at close to chance.

Where the answer token’s attention mass lands with sixteen separately prepared documents attached: anchor rows, the target document’s content, or the other fifteen documents’ content, by captured depth.
View the data
Attention addresses documents at shallow depth and reads them deep — published values
DepthAnchor rowsTarget document contentOther documents’ content
12%96%0.3%4.1%
25%91%0.8%8.4%
38%74%2.1%24%
51%14%5.5%80%
63%28%5.6%67%
76%28%5.7%66%
89%33%5.5%62%
99%3.0%8.3%89%
The question
Which layers choose between attached documents, and why does retrieval fail when many are attached?
The answer
Addressing happens at the entry layers, reading at the content layers. Crowding happens in the content layers.
What it changed
Located the failure in depth, which is what the fix had to address.
Caveat
A 'document' is a short passage of six invented facts, prepared once alone and attached later without re-reading. Sixteen attached.
09 / RECALL VS DOCUMENTSMEASURED

Attach more documents and recall collapses, confidently

0.92 → 0.00recall with two attached documents versus thirty-two. At sixteen it is 0.08, and the model stays confident throughout.

Recall of a fact falls from 0.92 with two attached documents to 0.08 with sixteen and 0.00 with thirty-two, while the model's first-token margin stays positive throughout. The model retrieves the wrong document with confidence: this is mis-binding, not forgetting.

Left: recall of a fact by number of separately prepared documents attached (twelve questions per group). Right: the mean logit gap between the correct first answer token and its best competitor; positive means confident.
View the data
Attach more documents and recall collapses, confidently — published values
Documents attachedRecallMean first-token margin
20.922.56
40.671.33
80.330.31
160.081.34
320.000.56
The question
How many separately prepared documents can be attached before recall degrades, and does the model know when it is wrong?
The answer
Recall collapses between eight and sixteen. The model does not know it is wrong.
What it changed
Named the failure mode and located it, which set the direction of the fix.
Caveat
Twelve questions per group. The margin is the logit gap between the correct first answer token and the best competitor; positive means confident.
10 / OUTLIER ROWSMEASURED

A handful of residual rows dominate everything near the output

227×the median row norm, reached by a few rows per prompt at depth 98%. Those rows carry 99.9% of the energy.

Residual row norms stay tame through depth, then at the last captured boundary a few rows per prompt reach roughly 200x the median and carry 99.9% of the energy. These rows are position-linked, not content-linked.

Row norms relative to the median row at the same depth, on a logarithmic scale: the 99th percentile and the maximum. Both stay near 1× until the last captured boundary.
View the data
A handful of residual rows dominate everything near the output — published values
Depthp99 ÷ medianMax ÷ medianOutlier rowsEnergy in outliers
38%1.33×1.43×00%
51%1.24×1.31×00%
63%1.22×2.01×00%
76%1.24×1.38×00%
89%1.21×1.27×00%
96%1.23×1.39×00%
98%199.96×226.52×30100%
The question
Is the residual well-behaved enough to compare across runs and fit linear maps on?
The answer
Not at the last boundary. Any comparison there differences these rows first.
What it changed
Invalidated earlier fits that had not excluded these rows: they were measuring position, not memory.
Caveat
Norms are published as ratios to the median at the same depth, never as absolute magnitudes. 22 prompts.
11 / DEPTH REWRITEMEASURED

The residual is rewritten between depth bands

≤ 0.33held-out variance explained by the best linear map between consecutive captured depths, above a constant-predictor floor. Cosine between depths: 0.17–0.62.

Within a single token, consecutive captured depths share under a third of their direction (cosine 0.17-0.31) until the very end. The best linear map from the shallower to the deeper state explains at most 0.33 of held-out variance above a constant-predictor floor, and at one band pair it fails to beat the floor at all.

Within a single token, the cosine between consecutive captured depths and the held-out R² of the best linear map from the shallower to the deeper state, floor-corrected. A negative R² is worse than predicting the mean.
View the data
The residual is rewritten between depth bands — published values
Depth pairCosine between depthsBest linear map, held-out R²
38% -> 51%0.210.13
51% -> 63%0.17-0.04
63% -> 76%0.270.07
76% -> 89%0.290.10
89% -> 96%0.310.16
96% -> 98%0.620.33
The question
Can a token's deeper state be produced from its shallower state without running the layers in between?
The answer
No.
What it changed
Closed the idea of filling a token's deeper cache rows cheaply from its shallower state.
Caveat
Folds grouped by prompt, constant-predictor floor subtracted. A negative R2 means the map is worse than predicting the mean.
12 / PUBLIC DISCLOSURE

Explain the result.
Protect the implementation.

These charts draw from a published pack of series aggregates and matrices. They contain no token-level rows, model weights, index layouts, extraction methods, kernels, thresholds or security controls. The objectives behind finding 5 and the runtime’s dimensions are deliberately not named.

The work was done by a human and an AI working as a pair: the human directs, questions and decides; the AI codes, analyses and drafts; a person reviews before anything is published. The full research record, the development status of Shadow State and the agent-capture boundary are on the research page.

A GOOD QUESTION IS A START

What would you
like to observe?

Tell us about the model, the runtime and the question. Sending delivers the brief to Shadow State Labs by email; you can also keep a local copy.

Sending emails the brief to Shadow State Labs. The website server keeps no copy. Please keep secrets, private captures and credentials out of the form.