# Accuracy audit and exhibit verification

Reviewed 11 September 2026. This report covers the attention-microscope revision of **Inside an AI Request**. It is an illustrative release, not a measured inference benchmark.

## What changed

The model view now opens a rotatable attention microscope: three Q/K/V planes, a selectable score/weight matrix, a visible causal mask and connections to the value vectors. Five teaching steps expose projection, matching, masking, normalisation and mixing. Changing one component of the third input recomputes the same mathematical example used by the detailed table. Narrow-screen 3D frames the current operation; native 2D shows projections, matrices and weighted contributions. These are educational values, not captured Qwen activations.

The existing four connected views, six-stop tour, experiments, request clock and atlas integration remain present. The charcoal, ivory, serif and orange portfolio identity is retained, with clearer blue-grey resource geometry.

## Findings and corrections

| Finding | Correction and evidence |
|---|---|
| Resident requests could appear to be computing when another request consumed the prompt budget. | Added explicit scheduler iterations and per-request scheduled work. The load view distinguishes resident slots from actual work. Shared compute activity no longer depends on whether the selected request is busy. |
| A request delivering its final output could return visually to the queue. | Final delivery now stays downstream of compute. Explicit release timestamps avoid subtracting a floating-point delivery delay to reconstruct this boundary. |
| State counts omitted final delivery; the chart interpolated between integer request counts. | Preparing, waiting, resident, final-delivery and complete counts now partition the workload. The chart uses steps rather than fractional transitions. |
| Peak memory sampling could miss the final KV append before release. | Peaks are calculated before releasing completed work. The default request reaches 543 useful positions and 544 block-rounded positions: 6,672,384 and 6,684,672 bytes respectively. Live state after release is zero. |
| The collector parsed the cache-reset response as JSON. | It now accepts the endpoint’s empty HTTP 200. Three isolated Python contract tests cover reset handling and refusal to overwrite evidence. |
| Qwen’s projection explanation omitted biases. | Engineering copy now includes Q/K/V biases and distinguishes the simplified toy from the model. |
| Rounded arithmetic looked like exact equalities. | Approximation symbols and a full-precision calculation note clarify displayed rounding. |
| The concurrency explanation referred to later arrivals despite a simultaneous workload. | It now describes requests farther back in the queue; all arrivals remain simultaneous. |
| Dense mobile 3D, touch handling and rebuilt cell labels reduced usability. | Mobile focuses one operation, one-finger swipes are reserved for page scrolling, and matrix keyboard focus is restored after recalculation. Explicit 2D and camera controls remain available. |

## Source audit

Architecture dimensions were cross-checked against the [official Qwen configuration](https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct/raw/main/config.json): 24 layers, hidden size 896, 14 query heads, two KV heads, MLP width 4,864 and tied embeddings. The checkpoint declares bfloat16; the proposed collector explicitly requests FP16. Both use two bytes, without implying identical numerical behaviour.

The [vLLM Qwen implementation](https://docs.vllm.ai/en/v0.10.2/api/vllm/model_executor/models/qwen2.html) confirms Q/K/V biases, bias-free attention output projection, rotary position encoding and residual/normalisation ordering. Independent parameter accounting reproduces 494,032,768 parameters and 988,065,536 FP16 weight bytes. KV accounting uses KV-head count: 2 × 24 × 2 × 64 × 2 = 12,288 bytes per stored position. These are storage calculations, not process memory measurements.

The [prefix-cache design](https://docs.vllm.ai/en/v0.10.2/design/prefix_caching.html) specifies full-block reuse based on exact preceding tokens and compatible context. The [cache manager](https://docs.vllm.ai/en/v0.10.2/api/vllm/v1/core/kv_cache_manager.html) limits a full prompt hit so the final token is recomputed for logits. The lab’s 16-position blocks and eligibility rule reflect this boundary; its scheduler and timing coefficients are independent teaching assumptions.

The [API server](https://docs.vllm.ai/en/v0.10.2/api/vllm/entrypoints/openai/api_server.html) returns an empty response for cache reset and does not acknowledge the engine’s internal boolean success. The collector therefore does not claim verified cache-hit counts. It waits for preceding requests before requesting reset. [Production metrics](https://docs.vllm.ai/en/v0.10.2/usage/metrics.html) include aggregate histograms; these cannot reconstruct individual internal request traces. Client first-content-chunk time remains explicitly a first-token proxy.

The configured model revision remains `7ae5576`; the collector resolves and records its full immutable SHA before loading. The current official configuration was readable during this audit, but revision-specific Hugging Face endpoints were unavailable through the research tool. Full revision resolution and the serving smoke test have not been re-executed here.

## Verification performed

- **32 JavaScript tests passed**, including attention masking and softmax, future-input independence, storage units and peak boundaries, scheduler termination/determinism across all 36 supported configurations, request-count partitions, prefix eligibility, evidence validation/import and existing portfolio/atlas rendering checks.
- **3 Python collector contract tests passed.** These use fake HTTP responses and temporary files; they are not benchmark recordings.
- **Lab-scoped TypeScript check passed.**
- **Production build and static-export generation passed.** The Sites build validated the Worker fetch handler and hosting manifest. The shared, lazily loaded Three.js/controls chunk still triggers Vite’s 500 kB warning; this is a bundle-size observation, not a measured rendering-performance score.
- Browser: managed desktop Chrome, 1348 × 926 screenshot. The browser had no WebGL, so the actual spatial views ran through Three.js SVGRenderer. Matrix selection, all five calculation steps, rotation controls, input changes and 2D equivalence were exercised.
- The input-change control preserved position 2’s output `[2.340, 0.660]` while changing the future third input. Masking a future cell displayed minus infinity, then exactly zero weight. Keyboard selection retained focus after the diagram rebuilt.
- All six guided stops completed. R8 remained selected and the timeline stayed at 2 simulated ms when moving through GPU, model and request views. Play entered playback; Reduce motion stopped playback and selected 2D.
- All three experiments completed Predict → Run → Inspect → Explain. The longer-prompt cases produced simulated first-token times of 15.4, 46.6 and 171.4 ms; concurrency showed both waiting and aggregate throughput; warm prefix reused 480 illustrative positions while all four cache cases retained 32 outputs. These observations validate the simulator’s UI, not GPU performance.
- Mobile: a 390 × 844 iframe in desktop Chrome, with a 375-pixel content viewport after the scrollbar. Default 2D, optional focused 3D, query/input controls and scroll access were checked. Document width equalled viewport width (375 px), with no horizontal overflow in the tested state. This is responsive verification, not a physical-phone test.
- Screenshots: `screenshots/attention-desktop-audit.jpg` and `screenshots/attention-mobile-audit.jpg`. The mobile image includes the desktop QA frame around the tested viewport.

The browser extension logged metadata-delivery errors and some automation clicks/scrolls timed out. Actions were checked against subsequent DOM/screenshot state, and keyboard interactions were used where necessary. No performance audit score, frame-rate guarantee or clean-browser-console claim is made.

## Remaining limits

`python scripts/ai-request/collect.py --preflight` exits with: “No NVIDIA GPU detected: nvidia-smi is unavailable or failed. No measurements collected.” No compute was purchased or rented. The production recording index is empty and measured mode is disabled.

A measured release still requires a compatible GPU, pinned-model/engine smoke test, genuine recordings for all comparisons, and request-level instrumentation for prefill, queue and cache occupancy where those metrics are required. Client SSE alone cannot provide these. WebGL rendering on a real GPU, physical touchscreen behaviour, OS-level reduced-motion media emulation and screen-reader output were not independently verified in this revision. The software 3D, explicit reduced-motion switch, semantic keyboard controls and 2D paths were exercised.
