# Stage 1A final matrix — 2026-10-04

All 13 configurations remain visible. C13 retains the 86 scored requests completed before VS Code opened; request 87 has confirmed overlap and 88–90 have unknown recovery. C10 below is the already completed clean replacement; the earlier partial contaminated attempt is retained separately. No further reruns.

All rows: SEALED_EXECUTION_VALIDATED, E1. That status means evidence completeness, not semantic correctness. PP/TG below are diagnostic native medians in tokens/s. Weight GiB are GGUF file bytes / 2³⁰, not total runtime RAM. P/F/U = semantic PASS/FAIL/UNKNOWN. All-90 pass yield is descriptive only; it mixes domains/conditions and is not a universal score.

| ID | Model · context/KV | Weights GiB | Diagnostic P/F/U | All-90 P/F/U | PASS/all | PP / TG tok/s | Compact / full TTFT s | Pressure / swap GiB | Interaction |
|---|---|---:|---|---|---:|---:|---:|---|---|
| C01 | 14B Q4_K_M · 32K/q4_0 | 8.371 | 17/17/2 | 66/22/2 | 73.3% | 117.02 / 12.21 | 1.844 / 473.025 | normal / 0 | None noted |
| C02 | 14B IQ3_M · 4K/f16 | 6.442 | 18/15/3 | 72/15/3 | 80.0% | 112.55 / 14.70 | 1.773 / 16.718 | normal / 0 | None noted |
| C03 | 32B IQ3_M · 4K/f16 | 13.793 | 24/11/1 | 78/11/1 | 86.7% | 47.52 / 6.65 | 4.572 / 48.470 | normal / 0 | None noted |
| C04 | 14B Q8_0 · 4K/f16 | 14.623 | 17/18/1 | 70/19/1 | 77.8% | 127.79 / 10.96 | 1.556 / 14.376 | normal / 0 | None noted |
| C05 | 7B Q4_K_M · 32K/q8_0 | 4.361 | 18/16/2 | 50/38/2 | 55.6% | 236.73 / 24.48 | 0.879 / 216.882 | normal / 0 | None noted |
| C06 | 3B Q8_0 · 4K/f16 | 3.060 | 11/21/4 | 42/41/7 | 46.7% | 571.77 / 45.36 | 0.339 / 3.203 | normal / 0 | None noted |
| C07 | 7B Q8_0 · 4K/f16 | 7.542 | 18/17/1 | 49/34/7 | 54.4% | 264.55 / 21.55 | 0.756 / 7.718 | normal / 0 | None noted |
| C08 | 32B Q4_K_M · 4K/f16 | 18.488 | 27/8/1 | 81/8/1 | 90.0% | 45.85 / 5.69 | 4.605 / 48.312 | normal / 0 | None noted |
| C09 | 14B Q4_K_M · 4K/f16 | 8.371 | 15/19/2 | 69/19/2 | 76.7% | 112.97 / 12.62 | 1.768 / 16.835 | normal / 0 | None noted |
| C10 | 7B IQ3_M · 4K/f16 | 3.329 | 15/20/1 | 47/41/2 | 52.2% | 223.63 / 27.13 | 0.880 / 8.363 | normal / 0 | Prior partial attempt retained |
| C11 | 14B Q4_K_M · 32K/q8_0 | 8.371 | 15/19/2 | 66/22/2 | 73.3% | 111.30 / 11.76 | 1.811 / 471.615 | normal / 0 | None noted |
| C12 | 7B Q4_K_M · 4K/f16 | 4.361 | 19/15/2 | 58/30/2 | 64.4% | 226.68 / 25.46 | 0.882 / 8.713 | normal / 0 | None noted |
| C13 | 7B Q4_K_M · 32K/q4_0 | 4.361 | 0/1/35 | 0/1/89 | 0.0% | 220.65 / 22.76 | 0.912 / 221.527 | normal / 0 | Partial: see request map |

## Denominators and useful-context controls

Full is 18 tasks; compact and evidence-missing are each 18. Full inputs are 2,028–2,048 native tokens at 4K allocation and 30,689–30,719 at 32K. Compact inputs are 210–225. These are synthetic repository, event and policy tasks. Missing-evidence PASS means the expected abstention/response, not successful retrieval.

| ID | Diagnostic PASS/36 | Diagnostic PASS/scorable | Full P/F/U | Compact P/F/U | Missing P/F/U | Truncated scored outputs |
|---|---:|---:|---|---|---|---:|
| C01 | 47.2% | 17/34 = 50.0% | 14/4/0 | 18/0/0 | 17/1/0 | 0 |
| C02 | 50.0% | 18/33 = 54.5% | 18/0/0 | 18/0/0 | 18/0/0 | 0 |
| C03 | 66.7% | 24/35 = 68.6% | 18/0/0 | 18/0/0 | 18/0/0 | 0 |
| C04 | 47.2% | 17/35 = 48.6% | 18/0/0 | 17/1/0 | 18/0/0 | 0 |
| C05 | 50.0% | 18/34 = 52.9% | 1/17/0 | 13/5/0 | 18/0/0 | 0 |
| C06 | 30.6% | 11/32 = 34.4% | 5/13/0 | 17/1/0 | 9/6/3 | 2 |
| C07 | 50.0% | 18/35 = 51.4% | 1/14/3 | 12/3/3 | 18/0/0 | 0 |
| C08 | 75.0% | 27/35 = 77.1% | 18/0/0 | 18/0/0 | 18/0/0 | 0 |
| C09 | 41.7% | 15/34 = 44.1% | 18/0/0 | 18/0/0 | 18/0/0 | 0 |
| C10 | 41.7% | 15/35 = 42.9% | 4/14/0 | 11/7/0 | 17/0/1 | 0 |
| C11 | 41.7% | 15/34 = 44.1% | 16/2/0 | 18/0/0 | 17/1/0 | 0 |
| C12 | 52.8% | 19/34 = 55.9% | 8/10/0 | 13/5/0 | 18/0/0 | 0 |
| C13 | 0.0% | 0/1 = 0.0% | 0/0/18 | 0/0/18 | 0/0/18 | 74 |

PASS/scorable excludes UNKNOWN and can mislead when very few answers are scorable (C13 has only one). Use counts and PASS/all alongside it. Complete transport responses can still be unscorable. Truncated responses are UNKNOWN under frozen grading, never converted into semantic FAIL.

## Measured memory observations

Per-process footprint is not total unified GPU/model memory. These values cannot establish remaining headroom or exact KV allocation. All configurations: total GPU/model allocation, exact KV bytes and usable headroom = NOT_MEASURED.

| ID | Sampled process physical-footprint peak GiB | Sampled RSS peak GiB | Physical compressor min–max GiB (whole host) | Pressure samples | Swap-in/out counter delta bytes |
|---|---:|---:|---:|---:|---|
| C01 | 1.848 | 10.152 | 0.712–0.838 | 75459 normal | 0 / 0 |
| C02 | 0.849 | 7.254 | 0.679–0.715 | 7302 normal | 0 / 0 |
| C03 | 1.103 | 14.752 | 0.717–0.792 | 16213 normal | 0 / 0 |
| C04 | 0.852 | 15.279 | 0.907–0.928 | 7431 normal | 0 / 0 |
| C05 | 1.080 | 5.436 | 0.875–0.915 | 32236 normal | 0 / 0 |
| C06 | 0.222 | 3.299 | 0.880–0.880 | 2552 normal | 0 / 0 |
| C07 | 0.307 | 7.792 | 0.879–0.880 | 2851 normal | 0 / 0 |
| C08 | 1.107 | 19.650 | 0.862–0.879 | 16415 normal | 0 / 0 |
| C09 | 0.849 | 9.153 | 0.991–0.995 | 7111 normal | 0 / 0 |
| C10 | 0.305 | 3.645 | 0.858–0.861 | 2576 normal | 0 / 0 |
| C11 | 3.350 | 11.653 | 0.840–0.857 | 75070 normal | 0 / 0 |
| C12 | 0.305 | 4.661 | 0.853–0.863 | 2812 normal | 0 / 0 |
| C13 | 0.642 | 4.999 | 0.825–0.840 | 70212 normal | 0 / 0 |

The final row includes the late interaction period; its pre-interaction footprint maximum is 0.642 GiB. Sampling may miss shorter events. The sysctl swap display is rounded; cumulative swap counters also did not increase. Compressor observations show why zero swap should not be paraphrased as zero memory-management activity.

## Diagnostic categories — six tasks per category

Each entry is P/F/U; neither six tasks nor a single session establishes a stable category ranking.

| ID | Coding | Counting | Debugging | Numeric | Ordering | State |
|---|---|---|---|---|---|
| C01 | 4/1/1 | 3/3/0 | 5/1/0 | 1/5/0 | 2/3/1 | 2/4/0 |
| C02 | 4/2/0 | 3/3/0 | 6/0/0 | 2/4/0 | 1/4/1 | 2/2/2 |
| C03 | 5/0/1 | 3/3/0 | 6/0/0 | 1/5/0 | 4/2/0 | 5/1/0 |
| C04 | 4/1/1 | 1/5/0 | 6/0/0 | 1/5/0 | 3/3/0 | 2/4/0 |
| C05 | 4/1/1 | 2/4/0 | 6/0/0 | 2/4/0 | 2/4/0 | 2/3/1 |
| C06 | 3/2/1 | 1/5/0 | 4/2/0 | 1/5/0 | 1/3/2 | 1/4/1 |
| C07 | 5/1/0 | 3/3/0 | 5/1/0 | 2/4/0 | 1/5/0 | 2/3/1 |
| C08 | 5/0/1 | 4/2/0 | 6/0/0 | 2/4/0 | 4/2/0 | 6/0/0 |
| C09 | 3/2/1 | 3/3/0 | 5/1/0 | 1/5/0 | 1/4/1 | 2/4/0 |
| C10 | 5/1/0 | 1/5/0 | 6/0/0 | 1/5/0 | 1/4/1 | 1/5/0 |
| C11 | 3/2/1 | 3/3/0 | 5/1/0 | 1/5/0 | 1/4/1 | 2/4/0 |
| C12 | 4/1/1 | 2/4/0 | 6/0/0 | 2/4/0 | 3/3/0 | 2/3/1 |
| C13 | 0/0/6 | 0/0/6 | 0/0/6 | 0/1/5 | 0/0/6 | 0/0/6 |

## Final-cell temporal sensitivity

| Condition | All requests P/F/U | Pre-interaction P/F/U | Pre-interaction n | All median TTFT s | Pre-interaction median TTFT s |
|---|---|---|---:|---:|---:|
| diagnostic | 0/1/35 | 0/1/35 | 36 | 0.520 | 0.520 |
| full | 0/0/18 | 0/0/17 | 17 | 221.527 | 221.541 |
| compact | 0/0/18 | 0/0/17 | 17 | 0.912 | 0.911 |
| missing | 0/0/18 | 0/0/16 | 16 | 221.597 | 221.657 |

Pre-interaction totals: 0 PASS, 1 FAIL, 85 UNKNOWN across 86 scored requests; 70 truncated. The output anomaly precedes interaction. The four later requests stay in the all-request table; one has proven overlap and three are recovery-unknown. No exact VS Code exit time was recorded, so no uncontaminated post-exit interval is invented.

## Evidence files

- [Exact matrix CSV](matrix.csv) and [structured matrix](matrix.json).
- [1,170 derived per-task records](task-outcomes.csv). Raw prompts, responses, and frozen grader internals are not included in this public bundle.
- Request-level contamination reconstruction is retained in the research repository and is summarized above; it is not included in this public bundle.
- The historical interrupted 7B IQ3 execution remains separate from its completed replacement and is not included in this public bundle.
- [Public export manifest](export-manifest.json) and [figure manifest](charts/manifest.json). Source seal inventories and offline validation records remain in the research repository.

This final interpretation supersedes the blanket cell-13 exclusion in earlier analysis-v2. Old reports and raw evidence remain unchanged. Timing medians include all terminal observations with timings, including truncated outputs; completed-only generation medians are separately available in JSON.
