Depth is a fraction
0.0 is the first captured boundary and 1.0 the last. No absolute layer counts or architecture dimensions are published.
One capture series on the instrumented K3 research runtime: ten prompt domains, thirty-two prompts each, 320 prompts in total, with exact expert selections and gate weights recorded for every token at every routed depth. Every number below is a single-model, single-capture result until replicated, and every one is published with its direction, its anchors and its caveat.
of the previous token’s experts are re-activated by the next token, against a chance level of 1.8%: a 16.5× lift.
05 / BACKWARD CONCENTRATION13%of layers hold 90% of gradient energy; 3.0% of sequence positions hold 90% of position energy. Nine objectives, one prompt.
09 / RECALL VS DOCUMENTS0.92 → 0.00recall with two attached documents versus thirty-two. At sixteen it is 0.08, and the model stays confident throughout.
07 / ROUTE PREDICTION0.28recall from the previous token’s route, against 0.20 for a nearest neighbour on the hidden state, 0.13 for popularity and a shuffled control at 0.07.
A bare figure is not a result. Each finding below states what the metric measures, which direction would have meant something different, and what it is anchored against: a chance level, a shuffled control, a stated baseline. Where a control exists it is drawn on the chart, not left to a footnote.
0.0 is the first captured boundary and 1.0 the last. No absolute layer counts or architecture dimensions are published.
Every magnitude is a ratio against a stated baseline: the median row, a one-token window, a random set of the same size.
Every finding is single-model and single-capture until replicated. Finding 4’s curve beyond 151 is a fitted extrapolation and is drawn as one.
None of these findings establish intent, causation or safety. Finding 8 shows where attention landed, not why. What they did change is stated per finding: which engineering lines were closed and which were opened.
Across ten prompt domains, the experts carrying 95% of a layer's routing mass are roughly half the pool at most depths; the narrowest layers still use about a quarter.
| Depth | Median | Min | Max |
|---|---|---|---|
| 0% | 77% | 73% | 78% |
| 1% | 69% | 65% | 72% |
| 2% | 53% | 44% | 60% |
| 3% | 50% | 46% | 53% |
| 4% | 62% | 57% | 67% |
| 5% | 74% | 71% | 78% |
| 7% | 55% | 50% | 63% |
| 8% | 59% | 53% | 68% |
| 9% | 56% | 50% | 61% |
| 10% | 60% | 56% | 67% |
| 11% | 59% | 55% | 66% |
| 12% | 39% | 34% | 41% |
| 13% | 70% | 65% | 72% |
| 14% | 66% | 61% | 70% |
| 15% | 63% | 56% | 68% |
| 16% | 59% | 49% | 66% |
| 18% | 58% | 47% | 64% |
| 19% | 55% | 46% | 60% |
| 20% | 55% | 45% | 60% |
| 21% | 59% | 53% | 62% |
| 22% | 64% | 61% | 68% |
| 23% | 62% | 58% | 66% |
| 24% | 55% | 52% | 57% |
| 25% | 68% | 64% | 70% |
| 26% | 68% | 66% | 71% |
| 27% | 63% | 58% | 67% |
| 29% | 64% | 59% | 66% |
| 30% | 68% | 66% | 71% |
| 31% | 59% | 54% | 61% |
| 32% | 59% | 56% | 62% |
| 33% | 58% | 55% | 62% |
| 34% | 64% | 61% | 67% |
| 35% | 63% | 61% | 66% |
| 36% | 60% | 57% | 63% |
| 37% | 56% | 54% | 59% |
| 38% | 59% | 56% | 63% |
| 40% | 48% | 42% | 54% |
| 41% | 46% | 42% | 49% |
| 42% | 49% | 47% | 57% |
| 43% | 49% | 45% | 53% |
| 44% | 48% | 42% | 51% |
| 45% | 37% | 32% | 45% |
| 46% | 41% | 36% | 48% |
| 47% | 42% | 39% | 50% |
| 48% | 38% | 33% | 45% |
| 49% | 34% | 31% | 39% |
| 51% | 39% | 37% | 46% |
| 52% | 48% | 44% | 52% |
| 53% | 36% | 28% | 44% |
| 54% | 39% | 34% | 43% |
| 55% | 44% | 39% | 47% |
| 56% | 34% | 27% | 44% |
| 57% | 26% | 17% | 36% |
| 58% | 30% | 22% | 38% |
| 59% | 31% | 24% | 37% |
| 60% | 26% | 19% | 41% |
| 62% | 29% | 25% | 41% |
| 63% | 33% | 28% | 40% |
| 64% | 28% | 23% | 39% |
| 65% | 45% | 43% | 50% |
| 66% | 55% | 51% | 58% |
| 67% | 54% | 49% | 59% |
| 68% | 52% | 46% | 57% |
| 69% | 47% | 41% | 53% |
| 70% | 49% | 46% | 57% |
| 71% | 51% | 49% | 57% |
| 73% | 45% | 41% | 53% |
| 74% | 50% | 47% | 56% |
| 75% | 51% | 47% | 58% |
| 76% | 55% | 49% | 62% |
| 77% | 54% | 51% | 61% |
| 78% | 48% | 46% | 52% |
| 79% | 36% | 31% | 41% |
| 80% | 47% | 40% | 50% |
| 81% | 53% | 47% | 57% |
| 82% | 35% | 29% | 45% |
| 84% | 37% | 34% | 46% |
| 85% | 41% | 35% | 47% |
| 86% | 38% | 32% | 51% |
| 87% | 47% | 42% | 57% |
| 88% | 42% | 39% | 47% |
| 89% | 50% | 47% | 59% |
| 90% | 48% | 45% | 56% |
| 91% | 38% | 27% | 54% |
| 92% | 40% | 31% | 50% |
| 93% | 34% | 22% | 48% |
| 95% | 38% | 35% | 49% |
| 96% | 37% | 30% | 52% |
| 97% | 38% | 32% | 52% |
| 98% | 39% | 29% | 52% |
| 99% | 44% | 37% | 56% |
| 100% | 48% | 45% | 57% |
The 95%-mass expert sets of any two domains share about half their members: off-diagonal Jaccard mean 0.55, min 0.45, max 0.67.
| Domain | Biology & medicine | Python code | GPU kernels | General chat | Legal contracts | Math proofs | Fiction | Rust/C systems | Storage systems | Translation |
|---|---|---|---|---|---|---|---|---|---|---|
| Biology & medicine | — | 0.55 | 0.54 | 0.54 | 0.60 | 0.57 | 0.51 | 0.54 | 0.60 | 0.54 |
| Python code | 0.55 | — | 0.64 | 0.53 | 0.59 | 0.58 | 0.53 | 0.67 | 0.67 | 0.55 |
| GPU kernels | 0.54 | 0.64 | — | 0.45 | 0.53 | 0.53 | 0.45 | 0.65 | 0.67 | 0.48 |
| General chat | 0.54 | 0.53 | 0.45 | — | 0.56 | 0.51 | 0.59 | 0.46 | 0.53 | 0.57 |
| Legal contracts | 0.60 | 0.59 | 0.53 | 0.56 | — | 0.55 | 0.55 | 0.56 | 0.60 | 0.58 |
| Math proofs | 0.57 | 0.58 | 0.53 | 0.51 | 0.55 | — | 0.50 | 0.54 | 0.57 | 0.54 |
| Fiction | 0.51 | 0.53 | 0.45 | 0.59 | 0.55 | 0.50 | — | 0.47 | 0.51 | 0.59 |
| Rust/C systems | 0.54 | 0.67 | 0.65 | 0.46 | 0.56 | 0.54 | 0.47 | — | 0.65 | 0.50 |
| Storage systems | 0.60 | 0.67 | 0.67 | 0.53 | 0.60 | 0.57 | 0.51 | 0.65 | — | 0.52 |
| Translation | 0.54 | 0.55 | 0.48 | 0.57 | 0.58 | 0.54 | 0.59 | 0.50 | 0.52 | — |
A token re-activates 29.4% of the previous token's experts, against a chance level of 1.8% for random sets of the same size — a 16x lift.
| Depth | Reuse | Chance |
|---|---|---|
| 0% | 10% | 1.8% |
| 1% | 23% | 1.8% |
| 2% | 30% | 1.8% |
| 3% | 37% | 1.8% |
| 4% | 20% | 1.8% |
| 5% | 8.5% | 1.8% |
| 7% | 19% | 1.8% |
| 8% | 23% | 1.8% |
| 9% | 33% | 1.8% |
| 10% | 32% | 1.8% |
| 11% | 31% | 1.8% |
| 12% | 34% | 1.8% |
| 13% | 28% | 1.8% |
| 14% | 29% | 1.8% |
| 15% | 30% | 1.8% |
| 16% | 30% | 1.8% |
| 18% | 31% | 1.8% |
| 19% | 33% | 1.8% |
| 20% | 34% | 1.8% |
| 21% | 29% | 1.8% |
| 22% | 24% | 1.8% |
| 23% | 30% | 1.8% |
| 24% | 36% | 1.8% |
| 25% | 25% | 1.8% |
| 26% | 23% | 1.8% |
| 27% | 28% | 1.8% |
| 29% | 26% | 1.8% |
| 30% | 26% | 1.8% |
| 31% | 33% | 1.8% |
| 32% | 32% | 1.8% |
| 33% | 29% | 1.8% |
| 34% | 27% | 1.8% |
| 35% | 27% | 1.8% |
| 36% | 28% | 1.8% |
| 37% | 27% | 1.8% |
| 38% | 29% | 1.8% |
| 40% | 37% | 1.8% |
| 41% | 34% | 1.8% |
| 42% | 32% | 1.8% |
| 43% | 35% | 1.8% |
| 44% | 36% | 1.8% |
| 45% | 39% | 1.8% |
| 46% | 39% | 1.8% |
| 47% | 35% | 1.8% |
| 48% | 43% | 1.8% |
| 49% | 44% | 1.8% |
| 51% | 40% | 1.8% |
| 52% | 32% | 1.8% |
| 53% | 36% | 1.8% |
| 54% | 36% | 1.8% |
| 55% | 37% | 1.8% |
| 56% | 38% | 1.8% |
| 57% | 40% | 1.8% |
| 58% | 40% | 1.8% |
| 59% | 40% | 1.8% |
| 60% | 41% | 1.8% |
| 62% | 40% | 1.8% |
| 63% | 36% | 1.8% |
| 64% | 37% | 1.8% |
| 65% | 32% | 1.8% |
| 66% | 27% | 1.8% |
| 67% | 25% | 1.8% |
| 68% | 23% | 1.8% |
| 69% | 24% | 1.8% |
| 70% | 23% | 1.8% |
| 71% | 22% | 1.8% |
| 73% | 29% | 1.8% |
| 74% | 24% | 1.8% |
| 75% | 21% | 1.8% |
| 76% | 20% | 1.8% |
| 77% | 20% | 1.8% |
| 78% | 28% | 1.8% |
| 79% | 28% | 1.8% |
| 80% | 23% | 1.8% |
| 81% | 20% | 1.8% |
| 82% | 25% | 1.8% |
| 84% | 23% | 1.8% |
| 85% | 21% | 1.8% |
| 86% | 23% | 1.8% |
| 87% | 20% | 1.8% |
| 88% | 29% | 1.8% |
| 89% | 22% | 1.8% |
| 90% | 23% | 1.8% |
| 91% | 26% | 1.8% |
| 92% | 28% | 1.8% |
| 93% | 25% | 1.8% |
| 95% | 33% | 1.8% |
| 96% | 28% | 1.8% |
| 97% | 27% | 1.8% |
| 98% | 42% | 1.8% |
| 99% | 32% | 1.8% |
| 100% | 25% | 1.8% |
When a window of tokens shares each layer's expert loads, relative bytes per token fall 4.6x by a window of 151. A control that keeps the same counts but scrambles which experts are chosen reaches only 2.9x, so most of the gain is real reuse rather than the arithmetic of bigger windows.
| Window | Saving (measured) | Scrambled-routing control |
|---|---|---|
| 1 | 1.00× | 1.00× |
| 2 | 1.16× | 1.01× |
| 4 | 1.37× | 1.03× |
| 8 | 1.66× | 1.06× |
| 16 | 1.99× | 1.14× |
| 19 | 2.11× | 1.17× |
| 32 | 2.41× | 1.29× |
| 38 | 2.64× | 1.36× |
| 64 | 2.96× | 1.58× |
| 76 | 3.44× | 1.81× |
| 151 | 4.58× | 2.89× |
| 256 (fitted, not measured) | 5.11× | — |
| 512 (fitted, not measured) | 9.14× | — |
| 1024 (fitted, not measured) | 18.29× | — |
Across nine objectives, 13% of layers hold 90% of the gradient energy and 3% of sequence positions hold 90% of the position energy. Half of the weight-gradient work carries under 2% of the gradient mass.
| Share of layers (sorted) | Cumulative energy, mean | Min | Max |
|---|---|---|---|
| 1.1% | 19% | 16% | 27% |
| 2.2% | 34% | 31% | 41% |
| 3.3% | 47% | 43% | 51% |
| 4.3% | 56% | 51% | 60% |
| 5.4% | 63% | 57% | 66% |
| 6.5% | 69% | 62% | 71% |
| 7.6% | 73% | 67% | 75% |
| 8.7% | 77% | 71% | 80% |
| 9.8% | 81% | 76% | 83% |
| 11% | 85% | 80% | 86% |
| 12% | 87% | 83% | 89% |
| 13% | 90% | 86% | 91% |
| 14% | 92% | 89% | 93% |
| 15% | 94% | 91% | 95% |
| 16% | 95% | 93% | 96% |
| 17% | 96% | 94% | 97% |
| 18% | 97% | 95% | 97% |
| 20% | 97% | 96% | 98% |
| 21% | 98% | 97% | 98% |
| 22% | 98% | 97% | 98% |
| 23% | 98% | 98% | 99% |
| 24% | 99% | 98% | 99% |
| 25% | 99% | 98% | 99% |
| 26% | 99% | 99% | 99% |
| 27% | 99% | 99% | 99% |
| 28% | 99% | 99% | 99% |
| 29% | 99% | 99% | 100% |
| 30% | 100% | 99% | 100% |
| 32% | 100% | 99% | 100% |
| 33% | 100% | 99% | 100% |
| 34% | 100% | 100% | 100% |
| 35% | 100% | 100% | 100% |
| 36% | 100% | 100% | 100% |
| 37% | 100% | 100% | 100% |
| 38% | 100% | 100% | 100% |
| 39% | 100% | 100% | 100% |
| 40% | 100% | 100% | 100% |
| 41% | 100% | 100% | 100% |
| 42% | 100% | 100% | 100% |
| 43% | 100% | 100% | 100% |
| 45% | 100% | 100% | 100% |
| 46% | 100% | 100% | 100% |
| 47% | 100% | 100% | 100% |
| 48% | 100% | 100% | 100% |
| 49% | 100% | 100% | 100% |
| 50% | 100% | 100% | 100% |
| 51% | 100% | 100% | 100% |
| 52% | 100% | 100% | 100% |
| 53% | 100% | 100% | 100% |
| 54% | 100% | 100% | 100% |
| 55% | 100% | 100% | 100% |
| 57% | 100% | 100% | 100% |
| 58% | 100% | 100% | 100% |
| 59% | 100% | 100% | 100% |
| 60% | 100% | 100% | 100% |
| 61% | 100% | 100% | 100% |
| 62% | 100% | 100% | 100% |
| 63% | 100% | 100% | 100% |
| 64% | 100% | 100% | 100% |
| 65% | 100% | 100% | 100% |
| 66% | 100% | 100% | 100% |
| 67% | 100% | 100% | 100% |
| 68% | 100% | 100% | 100% |
| 70% | 100% | 100% | 100% |
| 71% | 100% | 100% | 100% |
| 72% | 100% | 100% | 100% |
| 73% | 100% | 100% | 100% |
| 74% | 100% | 100% | 100% |
| 75% | 100% | 100% | 100% |
| 76% | 100% | 100% | 100% |
| 77% | 100% | 100% | 100% |
| 78% | 100% | 100% | 100% |
| 79% | 100% | 100% | 100% |
| 80% | 100% | 100% | 100% |
| 82% | 100% | 100% | 100% |
| 83% | 100% | 100% | 100% |
| 84% | 100% | 100% | 100% |
| 85% | 100% | 100% | 100% |
| 86% | 100% | 100% | 100% |
| 87% | 100% | 100% | 100% |
| 88% | 100% | 100% | 100% |
| 89% | 100% | 100% | 100% |
| 90% | 100% | 100% | 100% |
| 91% | 100% | 100% | 100% |
| 92% | 100% | 100% | 100% |
| 93% | 100% | 100% | 100% |
| 95% | 100% | 100% | 100% |
| 96% | 100% | 100% | 100% |
| 97% | 100% | 100% | 100% |
| 98% | 100% | 100% | 100% |
| 99% | 100% | 100% | 100% |
| 100% | 100% | 100% | 100% |
| 0.7% of positions | 59% | — | — |
| 3.3% of positions | 91% | — | — |
| 5.9% of positions | 94% | — | — |
| 8.6% of positions | 96% | — | — |
| 11% of positions | 96% | — | — |
| 14% of positions | 97% | — | — |
| 16% of positions | 98% | — | — |
| 19% of positions | 98% | — | — |
| 22% of positions | 98% | — | — |
| 24% of positions | 99% | — | — |
| 27% of positions | 99% | — | — |
| 30% of positions | 99% | — | — |
| 32% of positions | 99% | — | — |
| 35% of positions | 99% | — | — |
| 38% of positions | 99% | — | — |
| 40% of positions | 100% | — | — |
| 43% of positions | 100% | — | — |
| 45% of positions | 100% | — | — |
| 48% of positions | 100% | — | — |
| 51% of positions | 100% | — | — |
| 53% of positions | 100% | — | — |
| 56% of positions | 100% | — | — |
| 59% of positions | 100% | — | — |
| 61% of positions | 100% | — | — |
| 64% of positions | 100% | — | — |
| 66% of positions | 100% | — | — |
| 69% of positions | 100% | — | — |
| 72% of positions | 100% | — | — |
| 74% of positions | 100% | — | — |
| 77% of positions | 100% | — | — |
| 80% of positions | 100% | — | — |
| 82% of positions | 100% | — | — |
| 85% of positions | 100% | — | — |
| 88% of positions | 100% | — | — |
| 90% of positions | 100% | — | — |
| 93% of positions | 100% | — | — |
| 95% of positions | 100% | — | — |
| 98% of positions | 100% | — | — |
| 100% of positions | 100% | — | — |
The largest single active expert carries only 9.7% of a layer's expert-write norm. Keeping the top half of active experts still loses about 29% of the write.
| Share of active experts kept | Write norm recovered, mean | Write norm recovered, p10 |
|---|---|---|
| 6.3% | 23% | 12% |
| 13% | 34% | 21% |
| 19% | 42% | 29% |
| 25% | 49% | 36% |
| 31% | 55% | 43% |
| 38% | 61% | 50% |
| 44% | 66% | 56% |
| 50% | 71% | 62% |
| 56% | 75% | 67% |
| 63% | 80% | 73% |
| 69% | 84% | 78% |
| 75% | 87% | 83% |
| 81% | 91% | 87% |
| 88% | 94% | 92% |
| 94% | 97% | 96% |
| 100% | 100% | 100% |
Predicting the next layer's active experts one step ahead: the previous token's route recalls 0.28, a nearest-neighbour on the exact pre-router hidden state 0.20, expert popularity 0.13, and a shuffled control 0.07. Accuracy is flat out to four layers ahead, so prefetch lead time is free.
| Predictor | Recall at a one-token budget |
|---|---|
| previous token | 0.28 |
| nearest neighbour on exact hidden state | 0.20 |
| expert popularity | 0.13 |
| shuffled control | 0.07 |
With sixteen separately prepared documents attached, shallow attention layers put over 95% of the answer token's mass on one anchor row per document. Deep layers read content instead, and split between the right document and the other fifteen at close to chance.
| Depth | Anchor rows | Target document content | Other documents’ content |
|---|---|---|---|
| 12% | 96% | 0.3% | 4.1% |
| 25% | 91% | 0.8% | 8.4% |
| 38% | 74% | 2.1% | 24% |
| 51% | 14% | 5.5% | 80% |
| 63% | 28% | 5.6% | 67% |
| 76% | 28% | 5.7% | 66% |
| 89% | 33% | 5.5% | 62% |
| 99% | 3.0% | 8.3% | 89% |
Recall of a fact falls from 0.92 with two attached documents to 0.08 with sixteen and 0.00 with thirty-two, while the model's first-token margin stays positive throughout. The model retrieves the wrong document with confidence: this is mis-binding, not forgetting.
| Documents attached | Recall | Mean first-token margin |
|---|---|---|
| 2 | 0.92 | 2.56 |
| 4 | 0.67 | 1.33 |
| 8 | 0.33 | 0.31 |
| 16 | 0.08 | 1.34 |
| 32 | 0.00 | 0.56 |
Residual row norms stay tame through depth, then at the last captured boundary a few rows per prompt reach roughly 200x the median and carry 99.9% of the energy. These rows are position-linked, not content-linked.
| Depth | p99 ÷ median | Max ÷ median | Outlier rows | Energy in outliers |
|---|---|---|---|---|
| 38% | 1.33× | 1.43× | 0 | 0% |
| 51% | 1.24× | 1.31× | 0 | 0% |
| 63% | 1.22× | 2.01× | 0 | 0% |
| 76% | 1.24× | 1.38× | 0 | 0% |
| 89% | 1.21× | 1.27× | 0 | 0% |
| 96% | 1.23× | 1.39× | 0 | 0% |
| 98% | 199.96× | 226.52× | 30 | 100% |
Within a single token, consecutive captured depths share under a third of their direction (cosine 0.17-0.31) until the very end. The best linear map from the shallower to the deeper state explains at most 0.33 of held-out variance above a constant-predictor floor, and at one band pair it fails to beat the floor at all.
| Depth pair | Cosine between depths | Best linear map, held-out R² |
|---|---|---|
| 38% -> 51% | 0.21 | 0.13 |
| 51% -> 63% | 0.17 | -0.04 |
| 63% -> 76% | 0.27 | 0.07 |
| 76% -> 89% | 0.29 | 0.10 |
| 89% -> 96% | 0.31 | 0.16 |
| 96% -> 98% | 0.62 | 0.33 |
These charts draw from a published pack of series aggregates and matrices. They contain no token-level rows, model weights, index layouts, extraction methods, kernels, thresholds or security controls. The objectives behind finding 5 and the runtime’s dimensions are deliberately not named.
The work was done by a human and an AI working as a pair: the human directs, questions and decides; the AI codes, analyses and drafts; a person reviews before anything is published. The full research record, the development status of Shadow State and the agent-capture boundary are on the research page.