# Scope and limitations

- One 32 GiB Apple M1 Pro, pinned Metal llama.cpp build 11146 at source commit `7fe450e19`, Qwen2.5 Instruct models, one screening session per configuration. E1 exploratory evidence.
- The matrix is sparse. It does not test every model, weight format, KV setting or context limit. Results do not transfer automatically to other Macs, AMD, NVIDIA or developer workloads.
- Synthetic tasks: 36 diagnostics and 54 context requests per cell. The task mix is not a universal capability score. PASS/FAIL/UNKNOWN are retained.
- Weight-file bytes are not total runtime RAM. GPU/model allocation, exact KV allocation and available headroom are NOT_MEASURED.
- Sampled pressure was normal and swap was zero. Compressor activity was present. This does not show absence of all memory-management costs.
- Long-context values are TTFT for the tested inputs with zero reused KV tokens. OS disk-cache coldness, total answer time and retrieval quality were not established.
- C13 keeps 86 pre-interaction requests, one confirmed overlap and three recovery-unknown requests. Its initial output anomaly began before VS Code opened. C10's earlier partial attempt is separate from the clean replacement.
- No maximum usable context, energy, thermal behavior, real-world contention or independent uncertainty estimate was measured.
