# Research execution fieldnote  |  6 October 2026 UTC

Cisco Caceres. AI assistance supported implementation, literature screening, artifact review and writing.

This developer-run report reproduces six historical tool-calling output sets exactly, evaluates a 120-task local verifier/selection pilot with isolated execution, and measures acoustic robustness on frozen speaker-separated recordings. It retains the first verifier contract's failures and the severe-noise ASR and detector failures. A checked attention/optimization reference runs on two local hosts; Echo's recovery fixture exercises actual CLI processes. The sections below identify the inputs, measurements and limits of each result.

These are engineering and exploratory findings, not outside peer review, a novel-method claim or frontier-model performance. Paid model inference, rented compute and training expenditure for this reported work: **$0**. Local electricity, amortized hardware and labor were not measured.

This note separates original freezes from subsequent execution and audit. Original protocols remain historical records; follow-up source changes, decoder warnings and artifact recovery are dated additions, not retroactive preregistration. The dated additions below distinguish the completed verifier engineering pilot and retained negative Echo V1 diagnostic from ongoing Echo V2 qualification and metadata-only acoustic preparation.

## Saved tool-calling artifacts: exact rescore, incomplete model recovery

All six historical prediction files contain 400 unique IDs matching their respective split prefixes. Current evaluation reproduced every stored numerical metric and compact correctness record. This is saved-output reproduction: no new generation or training occurred, and stored latency fields are not new performance measurements.

| Split | Call/no-call items | Prompted macro | Historical SFT macro | Difference |
| --- | ---: | ---: | ---: | ---: |
| Validation | 360/40 | 0.673611 | 0.887500 | +0.213889 |
| Out-of-domain validation | 369/31 | 0.652942 | 0.849987 | +0.197045 |
| No-call canaries | 0/400 | 0.627500 | 0.927500 | +0.300000 |

Macro equally weights call exact-match and correct no-call accuracy when both classes exist. Canaries measure only no-call accuracy. Post-hoc paired item bootstrap and per-class exact McNemar diagnostics are exploratory; prefix sampling, shared tool/template clusters and pretraining exposure limit inference. No production or population superiority is established.

A later read-only Git recovery pass recovered a complete 3,441,185,608-byte safetensors container, SHA256 `db9639dd57242d536149ef49e78735c759748eee19868a3b2e256f9dad2b2b59`. Ancestry and matching metadata identify **abl-sft-names**, a Qwen3-1.7B schema-ablation arm (1,200 steps, learning rate 3e-5, seed 4242), not the original evaluated sft-1.7b checkpoint. At recovery, configuration/tokenizer were absent; the original container alone was not a complete runnable checkpoint. No original selected checkpoint was identified in inspected locations; other archives may exist. Corrupt temporary packs were tested on copies; originals were preserved. Substitution or retraining would constitute a new experiment.

A separately labeled prospective reconstruction now loads the recovered ablation locally. All 310 tensor shapes match the [pinned public Qwen3-1.7B configuration](https://huggingface.co/Qwen/Qwen3-1.7B/blob/70d244cc86ccca08cf5af4e1e306ecf908b1ad5e/config.json), including its tied output head. We used that base snapshot’s cached configuration/tokenizer and the recovered generation configuration; the configuration matched its pinned upstream bytes. PyTorch 2.14.0+cpu and Transformers 5.16.1 loaded the weights, restored the tied head and produced finite logits on a two-token synthetic-input forward pass (CPU BF16, 1.361 seconds including loading). The original recovered files remain unchanged. This establishes local execution viability under declared reconstruction assumptions, not tool-calling accuracy, original training provenance or an exact historical generation reproduction. The original selected SFT remains unidentified.

Artifact: `examples/tool-caller/results/record/reproduction-20261006.json`; methods and guarded commands: `examples/tool-caller/REPRODUCTION.md`. Compact tracked records support limited text-free recomputation. Raw corpora and predictions remain private. [Qwen3-1.7B](https://huggingface.co/Qwen/Qwen3-1.7B) and [ToolACE](https://huggingface.co/datasets/Team-ACE/ToolACE) publish Apache-2.0 metadata; [xLAM irrelevance](https://huggingface.co/datasets/MadeAgents/xlam-irrelevance-7.5k) publishes CC-BY-4.0. Preserve exact revisions, attribution and any gated access requirements rather than inferring source-content rights from repository licensing.

## Verifier development: retain contract failures and saturated success

V2 produced fourteen candidate responses that failed the required contract because they used Markdown. These remain failures; they were not repaired or discarded as warmups. V3 changed the collection contract before its own execution. Twelve candidates across six development tasks then passed both public and held-out hidden tests in an actual disposable networkless VM. Neither version is the proposed confirmation pilot.

The safe V3 VM outcome deposit contains 43 cases: seven qualification cases (one expected acceptance), twelve public candidates, twelve hidden candidates, and six canonical reference cases for each grader. All candidate and reference cases pass. Reports record no NIC, no host mounts and no observed root escape. Guest script hashes match installed source; a separate same-owner review matched five installed files to their manifest and passed six guard tests. The summary records the executed host-runner hash separately from a later strengthened version; the executed snapshot is preserved. The launch template hash was not deposited. Inspection of strengthened current source does not establish those exact bytes were used at historical launch; future runs should deposit all launch hashes.

The saturated candidate set cannot estimate specificity against incorrect candidates, demonstrate benefit from a reliability-aware policy, test H1, measure joint-error advantage, establish natural task-family shift or support frontier transfer. Generated code and tests share an interpreter: accidental-exit guards do not establish grading integrity against malicious candidates. Isolation is engineering evidence, not an independent assessor's attestation.

Methods: `examples/verifier-reliability/docs/PILOT-PROTOCOL.md`, `DEVELOPMENT-PROTOCOL-V3-20261006.md` and `scripts/NATIVE-VM-GRADING.md`. The pinned [HumanEval source](https://github.com/openai/human-eval) is MIT-licensed and useful for engineering qualification; its 2021 tasks cannot establish fresh-data generalization.

The updated research proposal remains an open hypothesis. Its closest-overlap screen includes [ADAP](https://arxiv.org/abs/2605.17609v2), [CAPS](https://arxiv.org/abs/2605.15513), [DAJ](https://arxiv.org/abs/2601.22230), [GRACE](https://arxiv.org/abs/2606.19354), [Inference-Time Pessimism](https://arxiv.org/abs/2503.21878v2) and [Hedging/HedgeTune](https://arxiv.org/html/2506.19248v2). Some entries remain abstract-only screens. Matched information and cost, full algorithm review and applicable baselines are necessary before novelty or advantage claims. No author implementation was copied merely because a repository exists.

## Local verifier pilot: descriptive engineering results


The unfiltered first-candidate comparator passes **109/120 tasks (90.83%)**. Soft best-of-n reaches **91.39%** at the three larger token caps, a **+0.56 percentage-point** three-seed task-mean difference; its paired descriptive 95% interval is **[0, +1.67 percentage points]**. The lowest cap yields no gain. Other unfiltered policies match the first-candidate result; all policies on the fixed 49-task public-qualified subset score **47/49 (95.92%)**. These are exploratory cached-pool results, not confirmatory superiority.

A clearly post-hoc headroom diagnostic finds **437/480 hidden-test candidate passes**, with 108 all-pass pools, ten all-fail pools and only two mixed pools. An oracle choosing any passing candidate could reach 110/120 tasks, leaving just one task (0.83 percentage points) of possible improvement over the first candidate. This small headroom limits what the pilot can reveal about selection methods; it is not a test of the proposed H1.

Collection used pinned `qwen3-coder:30b` and `qwen3.5:9b` through Warden Ollama 0.35.1. Five owned networkless VM batches reconciled all outcomes. All 9,288 public-only decisions were frozen before hidden-label analysis. A separate same-owner agent recomputed the 40 result groups and 10,000-resample task bootstraps from metadata; that check is disclosed assistance, not outside peer review.

The completed collection covers 120 frozen HumanEval tasks, four generated candidates per task (480 candidates), and 480 paired verifier calls. All 960 requests are accounted for. Hidden grading covers all 480 candidates; public grading covers 196 candidates on the original fixed 49-task subset. Invalid source contracts remain failures without execution.

Actual collection used 276,882 prompt tokens and 100,216 returned tokens. Invalid-judge abstentions: 0. Automatic contract-failure grading slots: 4 (slots count hidden and public tests separately). External API spend was $0; local electricity was not measured.

The prospective reconstructed extractor passed 51/57 public reference controls. The original eligibility subset remains 49 tasks; newly passing references were not added. Exact byte parity with the historical inline extractor is not claimed.

Selection was frozen before hidden-label analysis. The table describes cached replay on the same collected candidate pools, at four conservative input-plus-output token caps and three policy seeds (0, 1, 2). Policy seeds and candidates are not independent benchmark samples. Costs below are counterfactual policy consumption, distinct from actual collection cost.

On the 49-task public-qualified subset, single remains an unfiltered first-candidate comparator: it does not invoke or require a passing public test. The other five policies apply the public test filter. Subset membership and public-test filtering are distinct.

| Task subset | Token cap | Policy | Task-mean success | Paired difference vs single [descriptive 95% CI] | Mean consumed tokens | Cap-exhausted seeds |
| --- | ---: | --- | ---: | --- | ---: | ---: |
| 120 unfiltered | 4,608 | best_of_n | 90.83% | +0.00% [+0.00%, +0.00%] | 566.7 | 360/360 |
| 120 unfiltered | 4,608 | bounded_best_of_poisson | 90.83% | +0.00% [+0.00%, +0.00%] | 566.7 | 360/360 |
| 120 unfiltered | 4,608 | single | 90.83% | +0.00% [+0.00%, +0.00%] | 395.0 | 0/360 |
| 120 unfiltered | 4,608 | soft_best_of_n | 90.83% | +0.00% [+0.00%, +0.00%] | 566.7 | 360/360 |
| 120 unfiltered | 9,216 | best_of_n | 90.83% | +0.00% [+0.00%, +0.00%] | 3,109.4 | 9/360 |
| 120 unfiltered | 9,216 | bounded_best_of_poisson | 90.83% | +0.00% [+0.00%, +0.00%] | 3,109.4 | 9/360 |
| 120 unfiltered | 9,216 | single | 90.83% | +0.00% [+0.00%, +0.00%] | 395.0 | 0/360 |
| 120 unfiltered | 9,216 | soft_best_of_n | 91.39% | +0.56% [+0.00%, +1.67%] | 3,109.4 | 9/360 |
| 120 unfiltered | 18,432 | best_of_n | 90.83% | +0.00% [+0.00%, +0.00%] | 3,142.5 | 0/360 |
| 120 unfiltered | 18,432 | bounded_best_of_poisson | 90.83% | +0.00% [+0.00%, +0.00%] | 3,142.5 | 0/360 |
| 120 unfiltered | 18,432 | single | 90.83% | +0.00% [+0.00%, +0.00%] | 395.0 | 0/360 |
| 120 unfiltered | 18,432 | soft_best_of_n | 91.39% | +0.56% [+0.00%, +1.67%] | 3,142.5 | 0/360 |
| 120 unfiltered | 36,864 | best_of_n | 90.83% | +0.00% [+0.00%, +0.00%] | 3,142.5 | 0/360 |
| 120 unfiltered | 36,864 | bounded_best_of_poisson | 90.83% | +0.00% [+0.00%, +0.00%] | 3,142.5 | 0/360 |
| 120 unfiltered | 36,864 | single | 90.83% | +0.00% [+0.00%, +0.00%] | 395.0 | 0/360 |
| 120 unfiltered | 36,864 | soft_best_of_n | 91.39% | +0.56% [+0.00%, +1.67%] | 3,142.5 | 0/360 |
| 49 public-qualified | 4,608 | best_of_n | 95.92% | +0.00% [+0.00%, +0.00%] | 541.9 | 147/147 |
| 49 public-qualified | 4,608 | bounded_best_of_poisson | 95.92% | +0.00% [+0.00%, +0.00%] | 541.9 | 147/147 |
| 49 public-qualified | 4,608 | first_public | 95.92% | +0.00% [+0.00%, +0.00%] | 328.1 | 6/147 |
| 49 public-qualified | 4,608 | single | 95.92% | +0.00% [+0.00%, +0.00%] | 328.1 | 0/147 |
| 49 public-qualified | 4,608 | soft_best_of_n | 95.92% | +0.00% [+0.00%, +0.00%] | 541.9 | 147/147 |
| 49 public-qualified | 4,608 | uniform_public | 95.92% | +0.00% [+0.00%, +0.00%] | 328.1 | 147/147 |
| 49 public-qualified | 9,216 | best_of_n | 95.92% | +0.00% [+0.00%, +0.00%] | 2,540.5 | 0/147 |
| 49 public-qualified | 9,216 | bounded_best_of_poisson | 95.92% | +0.00% [+0.00%, +0.00%] | 2,540.5 | 0/147 |
| 49 public-qualified | 9,216 | first_public | 95.92% | +0.00% [+0.00%, +0.00%] | 372.8 | 0/147 |
| 49 public-qualified | 9,216 | single | 95.92% | +0.00% [+0.00%, +0.00%] | 328.1 | 0/147 |
| 49 public-qualified | 9,216 | soft_best_of_n | 95.92% | +0.00% [+0.00%, +0.00%] | 2,540.5 | 0/147 |
| 49 public-qualified | 9,216 | uniform_public | 95.92% | +0.00% [+0.00%, +0.00%] | 1,304.6 | 0/147 |
| 49 public-qualified | 18,432 | best_of_n | 95.92% | +0.00% [+0.00%, +0.00%] | 2,540.5 | 0/147 |
| 49 public-qualified | 18,432 | bounded_best_of_poisson | 95.92% | +0.00% [+0.00%, +0.00%] | 2,540.5 | 0/147 |
| 49 public-qualified | 18,432 | first_public | 95.92% | +0.00% [+0.00%, +0.00%] | 372.8 | 0/147 |
| 49 public-qualified | 18,432 | single | 95.92% | +0.00% [+0.00%, +0.00%] | 328.1 | 0/147 |
| 49 public-qualified | 18,432 | soft_best_of_n | 95.92% | +0.00% [+0.00%, +0.00%] | 2,540.5 | 0/147 |
| 49 public-qualified | 18,432 | uniform_public | 95.92% | +0.00% [+0.00%, +0.00%] | 1,304.6 | 0/147 |
| 49 public-qualified | 36,864 | best_of_n | 95.92% | +0.00% [+0.00%, +0.00%] | 2,540.5 | 0/147 |
| 49 public-qualified | 36,864 | bounded_best_of_poisson | 95.92% | +0.00% [+0.00%, +0.00%] | 2,540.5 | 0/147 |
| 49 public-qualified | 36,864 | first_public | 95.92% | +0.00% [+0.00%, +0.00%] | 372.8 | 0/147 |
| 49 public-qualified | 36,864 | single | 95.92% | +0.00% [+0.00%, +0.00%] | 328.1 | 0/147 |
| 49 public-qualified | 36,864 | soft_best_of_n | 95.92% | +0.00% [+0.00%, +0.00%] | 2,540.5 | 0/147 |
| 49 public-qualified | 36,864 | uniform_public | 95.92% | +0.00% [+0.00%, +0.00%] | 1,304.6 | 0/147 |

Task means average three seeds; paired bootstrap intervals resample tasks and retain paired seed means. These are conditional descriptive intervals for these fixed tasks and candidate pools, with multiple exploratory comparisons. They do not establish confirmatory superiority.

HumanEval is a familiar public benchmark released in 2021. Training contamination, task familiarity, model-family dependence, fixed candidate pools, and a small public-qualified subset limit generalization. No H1, joint-verifier learning, family-shift robustness, frontier-scale capability, or novel research contribution is established. This result has not received external independent review.

Execution timing represents outer wrapper wall time (15-second bound), with an inner five-second wall cap and three-second CPU cap. CPU usage is unmeasured. Raw responses, source code, hidden tests, and private machine paths are excluded from this summary.

[Download the aggregate pilot results and commitments](https://tuneharness.com/research/2026-10-06/verifier-pilot-summary.json). Source/response/test payloads are excluded. Methods: `examples/verifier-reliability/docs/POSTCOLLECTION-PROTOCOL-V1.md`; analysis SHA256 `4fcd3b98255aa15d7a6f58a303a159fb143f9a411c4bc4426c4d424bc67df3cb`, frozen-decision SHA256 `da9a0cc203fbc44e7511a384cf2257f11becc18c06ab72d8735e44bfdf4d4d25`.

## Acoustic research: frozen public speech, actual failures

Twelve cached LibriSpeech target speakers were separated into development three, calibration three and test six. Six disjoint train-source speakers supplied interference. Thirty-six complete-recording 16 kHz PCM16 WAVs represent clean, 10 dB and 5 dB conditions of twelve targets, not 36 independent speakers. Byte, finite-signal, no-clipping, unique decoded PCM, speaker/source separation and analytic component-SNR checks passed. Source/cache acquisition and earlier exposure lack independent custody; simulated read-English babble is not federal radio, clinical audio or real deployment noise.

[OpenSLR LibriSpeech](https://www.openslr.org/12) and the pinned dataset card declare CC-BY-4.0. Source and transformation attribution are retained. Licensed inputs do not establish model-pretraining nonexposure. No proprietary AMBIE model was evaluated.

Pinned local baseline: openai/whisper-tiny.en revision `87c7102498dcde7456f24cfd30239ca606ed9063`, CPU float32, four threads, greedy beam one, maximum 256 new tokens, explicit pinned English normalizer with no fallback. Data, core weights, decoder, normalizer and execution source were frozen before inference; separate training-source warmup was not counted. Secondary tokenizer-file and full host environment-lock coverage remain incomplete. Model dependencies derive from [Whisper](https://github.com/openai/whisper) and [the model publisher](https://huggingface.co/openai/whisper-tiny.en), not proprietary architecture work.

| Six test speakers | Word errors / reference words | WER | CER | Human speech flagged synthetic by Noise or Voice |
| --- | ---: | ---: | ---: | ---: |
| Clean | 7/155 | 4.52% | 1.40% | 4/6 |
| 10 dB babble | 25/155 | 16.13% | 9.26% | 4/6 |
| 5 dB babble | 274/155 | 176.77% | 159.39% | 5/6 |

WER and CER can exceed 100% through insertions. The 5 dB test condition contains 233 word insertions and one conservative returned-token-length warning at 256. Terminal EOS identifiers were not deposited; this warning does not prove truncation or why generation stopped. Zero ASR processing errors occurred. Retain insertion-heavy failures rather than clipping rates or silently regenerating them. Six test speakers and 155 reference words support a bounded diagnostic, not a population estimate or critical-field/radio claim.

The frozen Noise or Voice complete-recording DSP at threshold 0.5 processed all 36 human recordings without abstentions or errors. Its false alarms across all twelve targets were 5/12, 9/12 and 11/12; test-only counts appear above. There are no synthetic examples, so synthetic miss rate, EER, binary accuracy, speaker authentication and fraud-detection performance are not estimable. Do not tune the threshold on these test observations and relabel the same set blind.

Commitments: acoustic manifest `8636b20bedd6d44b9d58f322cfecf272a59c8f9c2626f401d00becd522ce28fe`; detector source `ca6a48f4f408ec1f83e54546c036dfb8f84c47c0ca9336c7e20ca642cd8b2e51`; detector score deposit `e683068d6ccc134782a29f6dec14be1b8737509d646b00bf8dc91f9373f1876c`. Methods and safe counts are under `examples/acoustic-evaluation/`; waveforms, references and hypotheses are excluded from publication. A later durable-root verifier amendment is an audit convenience, not a change to the original frozen inference protocol.

## Mathematical reference and Echo qualification

An original standard-library educational reference implements one causal attention head, cross-entropy, Q/K/V derivatives, SGD/Adam and incremental KV caching using the established equations of [Attention Is All You Need](https://arxiv.org/abs/1706.03762) and [Adam](https://arxiv.org/abs/1412.6980). Nine tests passed locally and on Feral Python 3.9.6. A 24-coordinate central-difference fixture has maximum gradient error 5.824e-12; full/step cache outputs match exactly. These are mechanism checks, not transformer training, novel architecture or proof of personal frontier research expertise. Learned projections, multihead blocks, positional encoding and full training are absent. Methods: `examples/model-mechanics/README.md`.

Echo's isolated recovery fixture executed nine actual CLI checks and eight selected source-test checks without model inference. It exercises private state persistence, advisory hold/expiry, task ownership and fail-closed recovery; scripted boundaries and logical local hosts do not represent arbitrary autonomous agents or separate physical hosts. Broker checks use owned loopback resources, not the user's broker. This qualification does not compare coding success or demonstrate coordination benefit. The separately qualified V1 original synthetic coding diagnostic is complete: eight tasks, two seeds and three conditions produced 48 retained attempts, with zero accepted pairs (single, plain two-worker and Echo each 0/16). All 138 model requests have known usage: 126,116 prompt tokens and 32,975 returned tokens (159,091 total). External paid inference cost was $0; hardware, electricity and labor were not measured. The frozen analysis binds source/model/VM identities and grades joint acceptance separately from individual features. The negative result exposes tool-contract failures, including missing write payloads; it establishes no coordination benefit, superiority or equivalence. A separately versioned V2 API/schema revision is undergoing new exact-source scripted VM qualification. It reuses the eight exposed development fixtures; no V2 model calls or performance results are reported. V1 failures remain unchanged and are not repaired, replaced or pooled with V2. Methods: Echo `docs/research/recovery-fixture-and-coding-protocol.md`.

## Contribution, access and next evidence

The contribution here is artifact recovery and strict re-scoring, frozen bounded acoustic diagnostics, transparent failure reporting, local collection/isolation infrastructure, checked educational mathematics and coordination qualification. Existing models, datasets, libraries and classic algorithms are credited separately from original integration and audit work. AI-generated assistance is not an external reviewer or independent replication.

Repository/source access is currently private or gated. Local commands and deposited hashes describe how authorized holders can verify artifacts; they do not imply that readers already have a publicly runnable source release. Publicly shareable counts and commitments exclude raw audio, references, hypotheses and generated code. An outside reviewer needs permitted access, exact released revisions and their own execution deposit.

Next verifier evidence needs candidate pools with sufficient incorrect/correct diversity, applicable matched baselines and new task families; the completed cached-pool pilot does not answer the joint-verifier hypothesis. Acoustic preparation has acquired checksum-verified [official ASVspoof 5 metadata](https://zenodo.org/records/14498691) and selected 192 records before scoring: 24 development and 24 evaluation speakers, balanced human/synthetic classes, disjoint speaker IDs and attack labels, and matched uncoded channels. No audio has been acquired or scored; distinct attack labels are not verified generator-family separation, and contents rights/source-utterance linkage still need clarification. The initial Echo V1 diagnostic is complete with zero accepted pairs; separately versioned V2 qualification is underway, with no V2 model results. Original checkpoint recovery remains distinct from an ablation reimplementation. Dated execution addenda must report failures, budget accounting and limitations before any stronger research claims.
