# Stage 1A clean-frontier protocol summary

**Protocol version:** Stage1A clean-frontier v1  
**Date:** 3 October 2026  
**Evidence level:** E1 exploratory screening

## Scope

The original Stage 1 matrix scheduled 26 cells across two host profiles. Stage 1A selected the 13 LAB_CLEAN cells in their original relative order. It did not add configurations or change the models, model bytes, quantizations, contexts, KV settings, prompts, repetitions, task set, graders, scoring, telemetry, warmup, or thresholds. The other profile, DEV_REFERENCE_V1, was deferred because the synthetic workload orchestration was not sufficiently robust for production screening. No DEV workload conclusions are made from Stage 1A.

The Stage 1A plan contains 1,170 scored requests and 13 warmups. The accepted first clean cell was inherited under its original evidence seal. The remaining 12 cells covered 1,080 scored requests and 12 warmups.

## Execution and admission

Before each cell, the controller checked the pinned model and runtime identities, model-file inventory, independent controller ancestry, AC power, AC Low Power Mode off, unrelated browser/editor/agent/benchmark processes absent, Docker stopped, and the existing baseline pressure and swap-delta conditions. The controller did not start, stop, or configure Docker or any DEV services.

The execution core retained the versioned runner's native adapter, dispatch, clean baseline, grading, token-aware timeout, and single-signal shutdown. The run used the frozen warmup and request rules. An incomplete or failed attempt remained in the audit trail and stopped advancement. A successful prefix could resume without repeating, skipping, overwriting, or resealing prior cells. A zero process exit was insufficient for acceptance: every cell had to pass the evidence validator and Stage 1A identity checks.

## Requests and scoring

Each configuration scheduled 36 synthetic diagnostic tasks and 54 synthetic context requests. Context requests covered full evidence, compact relevant evidence, and missing-evidence controls. PASS, FAIL, and UNKNOWN remain distinct. Execution acceptance does not establish semantic correctness or practical usefulness. No publication usability cutoff or recommendation score was defined.

The study used one screening session per configuration on a single Apple M1 Pro host. Repeated tasks are not independent host or session replications. The resulting evidence is E1. It supports bounded observations about the tested model, runtime, prompt, and host combinations. It does not establish universal quantization rankings, maximum useful context, or general hardware recommendations.

## Deferred study

Stage 1B is a separate future contention study. It may select a few representative configurations from measured Stage 1A results and compare clean, browser/editor, and controlled Docker conditions. Its design and any workload conclusions are outside this Stage 1A protocol.
