Platform whitepaper / Dated research notes
Research execution fieldnote | 6 October 2026 UTC
Cisco Caceres. AI assistance supported implementation, literature screening, artifact review and writing.
This developer-run report reproduces six historical tool-calling output sets exactly, qualifies local verifier collection and isolated execution, and measures acoustic robustness on frozen speaker-separated recordings. It retains the first verifier contract's failures and the severe-noise ASR and detector failures. A checked attention/optimization reference runs on two local hosts; Echo's recovery fixture exercises actual CLI processes. The sections below identify the inputs, measurements and limits of each result.
These are engineering and exploratory findings, not outside peer review, a novel-method claim or frontier-model performance. Paid model inference, rented compute and training expenditure for this reported work: $0. Local electricity, amortized hardware and labor were not measured.
This note separates original freezes from subsequent execution and audit. Original protocols remain historical records; follow-up source changes, decoder warnings and artifact recovery are dated additions, not retroactive preregistration. Further pilots and VM hardening are in progress and require separate deposits before becoming results.
Saved tool-calling artifacts: exact rescore, incomplete model recovery
All six historical prediction files contain 400 unique IDs matching their respective split prefixes. Current evaluation reproduced every stored numerical metric and compact correctness record. This is saved-output reproduction: no new generation or training occurred, and stored latency fields are not new performance measurements.
| Split | Call/no-call items | Prompted macro | Historical SFT macro | Difference |
|---|---|---|---|---|
| -- - | -- -: | -- -: | -- -: | -- -: |
| Validation | 360/40 | 0.673611 | 0.887500 | +0.213889 |
| Out-of-domain validation | 369/31 | 0.652942 | 0.849987 | +0.197045 |
| No-call canaries | 0/400 | 0.627500 | 0.927500 | +0.300000 |
Macro equally weights call exact-match and correct no-call accuracy when both classes exist. Canaries measure only no-call accuracy. Post-hoc paired item bootstrap and per-class exact McNemar diagnostics are exploratory; prefix sampling, shared tool/template clusters and pretraining exposure limit inference. No production or population superiority is established.
A later read-only Git recovery pass recovered a complete 3,441,185,608-byte safetensors container, SHA256 db9639dd57242d536149ef49e78735c759748eee19868a3b2e256f9dad2b2b59. Ancestry and matching metadata identify abl-sft-names, a Qwen3-1.7B schema-ablation arm (1,200 steps, learning rate 3e-5, seed 4242), not the original evaluated sft-1.7b checkpoint. Missing configuration/tokenizer prevent treating the recovered container as a ready model. No original selected checkpoint was identified in inspected locations; other archives may exist. Corrupt temporary packs were tested on copies; originals were preserved. Substitution or retraining would constitute a new experiment.
A separately labeled prospective reconstruction now loads the recovered ablation locally. All 310 tensor shapes match the pinned public Qwen3-1.7B configuration, including its tied output head. We used that base snapshot’s cached configuration/tokenizer and the recovered generation configuration; the configuration matched its pinned upstream bytes. PyTorch 2.14.0+cpu and Transformers 5.16.1 loaded the weights, restored the tied head and produced finite logits on a two-token synthetic-input forward pass (CPU BF16, 1.361 seconds including loading). The original recovered files remain unchanged. This establishes local execution viability under declared reconstruction assumptions, not tool-calling accuracy, original training provenance or an exact historical generation reproduction. The original selected SFT remains unidentified.
Artifact: examples/tool-caller/results/record/reproduction-20261006.json; methods and guarded commands: examples/tool-caller/REPRODUCTION.md. Compact tracked records support limited text-free recomputation. Raw corpora and predictions remain private. Qwen3-1.7B and ToolACE publish Apache-2.0 metadata; xLAM irrelevance publishes CC-BY-4.0. Preserve exact revisions, attribution and any gated access requirements rather than inferring source-content rights from repository licensing.
Verifier development: retain contract failures and saturated success
V2 produced fourteen candidate responses that failed the required contract because they used Markdown. These remain failures; they were not repaired or discarded as warmups. V3 changed the collection contract before its own execution. Twelve candidates across six development tasks then passed both public and held-out hidden tests in an actual disposable networkless VM. Neither version is the proposed confirmation pilot.
The safe V3 VM outcome deposit contains 43 cases: seven qualification cases (one expected acceptance), twelve public candidates, twelve hidden candidates, and six canonical reference cases for each grader. All candidate and reference cases pass. Reports record no NIC, no host mounts and no observed root escape. Guest script hashes match installed source; a separate same-owner review matched five installed files to their manifest and passed six guard tests. The summary records the executed host-runner hash separately from a later strengthened version; the executed snapshot is preserved. The launch template hash was not deposited. Inspection of strengthened current source does not establish those exact bytes were used at historical launch; future runs should deposit all launch hashes.
The saturated candidate set cannot estimate specificity against incorrect candidates, demonstrate benefit from a reliability-aware policy, test H1, measure joint-error advantage, establish natural task-family shift or support frontier transfer. Generated code and tests share an interpreter: accidental-exit guards do not establish grading integrity against malicious candidates. Isolation is engineering evidence, not an independent assessor's attestation.
Methods: examples/verifier-reliability/docs/PILOT-PROTOCOL.md, DEVELOPMENT-PROTOCOL-V3-20261006.md and scripts/NATIVE-VM-GRADING.md. The pinned HumanEval source is MIT-licensed and useful for engineering qualification; its 2021 tasks cannot establish fresh-data generalization.
The updated research proposal remains an open hypothesis. Its closest-overlap screen includes ADAP, CAPS, DAJ, GRACE, Inference-Time Pessimism and Hedging/HedgeTune. Some entries remain abstract-only screens. Matched information and cost, full algorithm review and applicable baselines are necessary before novelty or advantage claims. No author implementation was copied merely because a repository exists.
Acoustic research: frozen public speech, actual failures
Twelve cached LibriSpeech target speakers were separated into development three, calibration three and test six. Six disjoint train-source speakers supplied interference. Thirty-six complete-recording 16 kHz PCM16 WAVs represent clean, 10 dB and 5 dB conditions of twelve targets, not 36 independent speakers. Byte, finite-signal, no-clipping, unique decoded PCM, speaker/source separation and analytic component-SNR checks passed. Source/cache acquisition and earlier exposure lack independent custody; simulated read-English babble is not federal radio, clinical audio or real deployment noise.
OpenSLR LibriSpeech and the pinned dataset card declare CC-BY-4.0. Source and transformation attribution are retained. Licensed inputs do not establish model-pretraining nonexposure. No proprietary AMBIE model was evaluated.
Pinned local baseline: openai/whisper-tiny.en revision 87c7102498dcde7456f24cfd30239ca606ed9063, CPU float32, four threads, greedy beam one, maximum 256 new tokens, explicit pinned English normalizer with no fallback. Data, core weights, decoder, normalizer and execution source were frozen before inference; separate training-source warmup was not counted. Secondary tokenizer-file and full host environment-lock coverage remain incomplete. Model dependencies derive from Whisper and the model publisher, not proprietary architecture work.
| Six test speakers | Word errors / reference words | WER | CER | Human speech flagged synthetic by Noise or Voice |
|---|---|---|---|---|
| -- - | -- -: | -- -: | -- -: | -- -: |
| Clean | 7/155 | 4.52% | 1.40% | 4/6 |
| 10 dB babble | 25/155 | 16.13% | 9.26% | 4/6 |
| 5 dB babble | 274/155 | 176.77% | 159.39% | 5/6 |
WER and CER can exceed 100% through insertions. The 5 dB test condition contains 233 word insertions and one conservative returned-token-length warning at 256. Terminal EOS identifiers were not deposited; this warning does not prove truncation or why generation stopped. Zero ASR processing errors occurred. Retain insertion-heavy failures rather than clipping rates or silently regenerating them. Six test speakers and 155 reference words support a bounded diagnostic, not a population estimate or critical-field/radio claim.
The frozen Noise or Voice complete-recording DSP at threshold 0.5 processed all 36 human recordings without abstentions or errors. Its false alarms across all twelve targets were 5/12, 9/12 and 11/12; test-only counts appear above. There are no synthetic examples, so synthetic miss rate, EER, binary accuracy, speaker authentication and fraud-detection performance are not estimable. Do not tune the threshold on these test observations and relabel the same set blind.
Commitments: acoustic manifest 8636b20bedd6d44b9d58f322cfecf272a59c8f9c2626f401d00becd522ce28fe; detector source ca6a48f4f408ec1f83e54546c036dfb8f84c47c0ca9336c7e20ca642cd8b2e51; detector score deposit e683068d6ccc134782a29f6dec14be1b8737509d646b00bf8dc91f9373f1876c. Methods and safe counts are under examples/acoustic-evaluation/; waveforms, references and hypotheses are excluded from publication. A later durable-root verifier amendment is an audit convenience, not a change to the original frozen inference protocol.
Mathematical reference and Echo qualification
An original standard-library educational reference implements one causal attention head, cross-entropy, Q/K/V derivatives, SGD/Adam and incremental KV caching using the established equations of Attention Is All You Need and Adam. Nine tests passed locally and on Feral Python 3.9.6. A 24-coordinate central-difference fixture has maximum gradient error 5.824e-12; full/step cache outputs match exactly. These are mechanism checks, not transformer training, novel architecture or proof of personal frontier research expertise. Learned projections, multihead blocks, positional encoding and full training are absent. Methods: examples/model-mechanics/README.md.
Echo's isolated recovery fixture executed nine actual CLI checks and eight selected source-test checks without model inference. It exercises private state persistence, advisory hold/expiry, task ownership and fail-closed recovery; scripted boundaries and logical local hosts do not represent arbitrary autonomous agents or separate physical hosts. Broker checks use owned loopback resources, not the user's broker. This qualification does not compare coding success or demonstrate coordination benefit. The coding comparison remains a designed, unexecuted study. Methods: Echo docs/research/recovery-fixture-and-coding-protocol.md.
Contribution, access and next evidence
The contribution here is artifact recovery and strict re-scoring, frozen bounded acoustic diagnostics, transparent failure reporting, local collection/isolation infrastructure, checked educational mathematics and coordination qualification. Existing models, datasets, libraries and classic algorithms are credited separately from original integration and audit work. AI-generated assistance is not an external reviewer or independent replication.
Repository/source access is currently private or gated. Local commands and deposited hashes describe how authorized holders can verify artifacts; they do not imply that readers already have a publicly runnable source release. Publicly shareable counts and commitments exclude raw audio, references, hypotheses and generated code. An outside reviewer needs permitted access, exact released revisions and their own execution deposit.
Next evidence should be a separately frozen verifier pilot with incorrect-candidate controls and matched baselines; an acoustic follow-up with newly licensed, generator/channel/speaker-separated synthetic and real-noise inputs; and an equal-resource Echo coding comparison. Original checkpoint recovery remains distinct from an ablation reimplementation. Dated execution addenda must report failures, budget accounting and limitations before any stronger research claims.
Download the canonical note (Markdown) · Source checksum and publication metadata