← Phone agents DAS Voice Ultra · September 15, 2026
What our first
voice pilot showed.
We tested eight frozen audio-reasoning questions. This report records the results and the configuration used for that run.
Download the curated scorecard ↓ Post-run engineering rubric 6of 8
Clean completions
75% on this eight-case pilot. This is a custom review metric, not official Big Bench Audio accuracy or a score on the full 1,000 questions.
✓✓✓×✓✓×✓
Two early answers are flagged but counted as clean under this rubric. Keep the denominator in view. Eight selected cases cannot establish a full benchmark score. Our post-run rubric differs from published benchmark grading, so these percentages should not be compared directly with leaderboard or vendor accuracy claims.
The counts, separately
A correct final answer is one part of the result.
All counts below use the same eight selected cases. Final-answer correctness does not remove an earlier wrong answer or repair a failed execution.
- First answer correct
- 7/ 8
- One answer started wrong.
- Final answer correct
- 8/ 8
- Diagnostic only; includes a correction and a failed completion.
- Original protocol completed
- 7/ 8
- One completion-harness failure remains in the result.
- Answers before input ended
- 2/ 8
- Timing concerns are retained, even when the answer was right.
Methodology
Frozen inputs.
A disclosed review.
The source is Artificial Analysis’s Big Bench Audio dataset, which contains 1,000 audio questions across four reasoning categories. We selected two cases per category before execution, using a fixed hash-based selection.
The models received original dataset audio, decoded to 24 kHz mono PCM. They were not given the original question text, case IDs, categories, or answer key.
Artificial Analysis maintains its own speech benchmarking methodology. Our eight-case sample and engineering rubric are different; this report is self-run and has no independent certification.
Configuration tested on September 15, 2026GPT-Live 1 + Cerebras support
GPT-Live 1 used the Marin voice. Cerebras Qwen 3.8 27B ran inside Claude Code as both the delegated reasoning agent and proactive observer. Tools, provider reasoning/thinking, and model fallbacks were disabled. Harness effort was medium.
Each case allowed a 45-second answer window and a three-second settle period. Support was capped at six completions per role per case. This records a test configuration; the DAS Voice package can evolve.
Review methodRecorded audio, reviewed through local ASR
Codex Astra reviewed hash-verified output using cached local small.en speech recognition: a beam-1 pass and a second beam-5 pass, cross-checked with stream captions. Both passes used the same transcription model.
There was no independent human listening or paid judge. Review reused the captured audio without new voice runs. Transcription and interpretation errors remain possible.
Post-run ruleWhat counted as clean
The first and final spoken answers had to match the key, the response could not contain a contradictory answer, and the original execution protocol had to complete. This rubric was defined during review.
Early speech was tracked separately, not made a failure condition. The original automatic scorer returned one correct, one incorrect and six needing review; those grades are preserved below and in the download.
All selected cases
The case-by-case record.
Every selected case remains in the denominator, including the correction and the harness failure. “Clean” refers only to our post-run rubric.
On a narrow screen, scroll the table horizontally to read every column.
What needs attention
Keep the mistakes visible.
Case 883 · wrong, then correctedThe correction does not erase the first answer.
The voice first said Inga tells the truth, then corrected itself. The final answer matched the key; our clean-completion rubric still fails the case. A support note appeared in the same interval, but this run does not establish a causal benefit from the observer.
Case 676 · completion harnessThe right words, a failed execution.
The recorded response said “seven objects.” The original completion predicate did not recognize the number-word answer, so the session reached its duration limit and rejected a later input frame. The original protocol failure remains a failure in this report.
Cases 101 and 295 · early speechCorrect answers can still arrive too soon.
Both answered before the input audio had finished. This matters for conversational behavior even though neither contradicted the answer key. Waiting for the whole question was not required by our post-run rubric, so these cases remain clean with an explicit timing flag.
The observer recorded 25 reviews, three notes, 12 stale results, three rejected results and one error. These counters describe this configuration’s activity; without a paired run with support disabled, they do not measure how much the support system improved the score.
Timing · server audio only
When the final answer finished.
Measured from the end of input playback to the end of the final spoken answer, across all eight cases:
- Minimum
- 0.634 s
- Median
- 1.973 s
- Maximum
- 12.828 s
Offline transcription alignment is approximate. The separately recorded median time to next audio was 1.451 seconds, but that audio can be filler or repetition and the metric omits early speech. These are not PSTN response times, p95 measurements or an SLA.
Cost · modeled inference only
About 23¢ for
this eight-case run.
The estimate combines 238 seconds of provider-reported GPT-Live usage (about 19.8¢) with about 3.0¢ for the support model.
Speech is calculated at OpenAI’s published $0.05 per minute, billed per second. Support uses estimated rates of $1 per million input tokens and $2 per million output tokens.
This is a model-cost estimate, not a reconciled provider invoice or a customer price. It excludes carrier charges, local transcription, hosting and other infrastructure costs.
Reproducibility and limits
Inspect the scorecard.
Keep the scope intact.
The download contains every case result, review rationale, original automatic grade, configuration, timing, cost basis and evidence hash. Hashes identify retained artifacts; raw recordings and execution records are not linked in this release.
Frozen selection and preserved artifact hashes
- Dataset revision
af7bb9c25b015792583ca4da3ee27ec62cb79fe6- Selection seed
das-bba-pilot-v1- Selection method
- sort SHA256(seed:id) independently per category; interleave categories
- Frozen manifest SHA-256
dbe963a384577bc62856937f2644a524d025b9c2fde1f5aab8e7d9b95c5f381a- Original automatic summary SHA-256
076d4296cdfaa28b5a8d46dbdb01a40a3418a29df36c418850179c1f1cc1aa2f- Review decisions SHA-256
2137c41318fb84de294b6f42ec4959aa7dc41d6dd9c6b113eab7b0ac014e0272- Secondary ASR SHA-256
aa529bed79260bfe7b550b0a084dda5d801582341eb7024cefbb4a4cb0a85634
Individual input-audio, output-audio, record and original-grade hashes are included per case in the JSON. A hash alone is not an independently verified result.
What this pilot does not establish
Full-dataset accuracy, independent leaderboard standing, live business-tool success, booking or CRM accuracy, carrier reliability, human-handoff quality, or a causal advantage from subconscious support. Tools were disabled, the sample was eight questions, and the review was not independent human listening.