DavidAgents
← Phone agents

DAS Voice Ultra · September 15, 2026

What our first
voice pilot showed.

We tested eight frozen audio-reasoning questions. This report records the results and the configuration used for that run.

Download the curated scorecard
Post-run engineering rubric
6of 8

Clean completions

75% on this eight-case pilot. This is a custom review metric, not official Big Bench Audio accuracy or a score on the full 1,000 questions.

✓✓✓×✓✓×✓
Two early answers are flagged but counted as clean under this rubric.
Keep the denominator in view. Eight selected cases cannot establish a full benchmark score. Our post-run rubric differs from published benchmark grading, so these percentages should not be compared directly with leaderboard or vendor accuracy claims.

The counts, separately

A correct final answer is one part of the result.

All counts below use the same eight selected cases. Final-answer correctness does not remove an earlier wrong answer or repair a failed execution.

First answer correct
7/ 8
One answer started wrong.
Final answer correct
8/ 8
Diagnostic only; includes a correction and a failed completion.
Original protocol completed
7/ 8
One completion-harness failure remains in the result.
Answers before input ended
2/ 8
Timing concerns are retained, even when the answer was right.

Methodology

Frozen inputs.
A disclosed review.

The source is Artificial Analysis’s Big Bench Audio dataset, which contains 1,000 audio questions across four reasoning categories. We selected two cases per category before execution, using a fixed hash-based selection.

The models received original dataset audio, decoded to 24 kHz mono PCM. They were not given the original question text, case IDs, categories, or answer key.

Artificial Analysis maintains its own speech benchmarking methodology. Our eight-case sample and engineering rubric are different; this report is self-run and has no independent certification.

Configuration tested on September 15, 2026

GPT-Live 1 + Cerebras support

GPT-Live 1 used the Marin voice. Cerebras Qwen 3.8 27B ran inside Claude Code as both the delegated reasoning agent and proactive observer. Tools, provider reasoning/thinking, and model fallbacks were disabled. Harness effort was medium.

Each case allowed a 45-second answer window and a three-second settle period. Support was capped at six completions per role per case. This records a test configuration; the DAS Voice package can evolve.

Review method

Recorded audio, reviewed through local ASR

Codex Astra reviewed hash-verified output using cached local small.en speech recognition: a beam-1 pass and a second beam-5 pass, cross-checked with stream captions. Both passes used the same transcription model.

There was no independent human listening or paid judge. Review reused the captured audio without new voice runs. Transcription and interpretation errors remain possible.

Post-run rule

What counted as clean

The first and final spoken answers had to match the key, the response could not contain a contradictory answer, and the original execution protocol had to complete. This rubric was defined during review.

Early speech was tracked separately, not made a failure condition. The original automatic scorer returned one correct, one incorrect and six needing review; those grades are preserved below and in the download.

All selected cases

The case-by-case record.

Every selected case remains in the denominator, including the correction and the harness failure. “Clean” refers only to our post-run rubric.

On a narrow screen, scroll the table horizontally to read every column.

Eight frozen BBA cases · two per category
Case / categoryAnswer keyFirst → finalOriginal protocolPost-run resultCaveat and original grade
#101Formal fallaciesinvalidinvalidthen invalidCompletedClean · early answer

Answered before the input finished, then repeated the same answer. Early speech is flagged separately and is not a failure under this post-run rubric.

Original automatic grade: needs review (not an unambiguous short answer).
#330NavigationYesyesthen yesCompletedClean

The received answer is Yes, matching the original answer key.

Original automatic grade: correct (strict spoken answer).
#562Object counting99then 9CompletedClean

The received answer says nine fruits, matching 9. Filler and a complete sentence caused the strict initial scorer to defer review.

Original automatic grade: needs review (not an unambiguous short answer).
#883Web of liesNoyesthen noCompletedWrong, then corrected

First said Inga tells the truth, then corrected that answer. An observer note was recorded in the same interval; this does not establish that the observer caused the correction.

Original automatic grade: needs review (not an unambiguous short answer).
#49Formal fallaciesinvalidinvalidthen invalidCompletedClean

After an acknowledgement, the received answer is invalid, matching the original key. Acknowledgement text made the initial exact scorer defer review.

Original automatic grade: needs review (not an unambiguous short answer).
#295NavigationYesyesthen yesCompletedClean · early answer

Answered before the input finished, then repeated the same answer. Early speech is flagged separately and is not a failure under this post-run rubric.

Original automatic grade: needs review (not an unambiguous short answer).
#676Object counting77then 7FailedCompletion failure

Spoken answer was correct. The completion predicate missed it; the session reached its limit and rejected a later input frame. The original failure remains.

Original automatic grade: incorrect (incomplete or error).
#852Web of liesYesyesthen yesCompletedClean

The received answer says Yes, Delbert tells the truth, matching Yes. An uncertain observer note was delivered afterwards but the spoken answer was not changed.

Original automatic grade: needs review (not an unambiguous short answer).

What needs attention

Keep the mistakes visible.

Case 883 · wrong, then corrected

The correction does not erase the first answer.

The voice first said Inga tells the truth, then corrected itself. The final answer matched the key; our clean-completion rubric still fails the case. A support note appeared in the same interval, but this run does not establish a causal benefit from the observer.

Case 676 · completion harness

The right words, a failed execution.

The recorded response said “seven objects.” The original completion predicate did not recognize the number-word answer, so the session reached its duration limit and rejected a later input frame. The original protocol failure remains a failure in this report.

Cases 101 and 295 · early speech

Correct answers can still arrive too soon.

Both answered before the input audio had finished. This matters for conversational behavior even though neither contradicted the answer key. Waiting for the whole question was not required by our post-run rubric, so these cases remain clean with an explicit timing flag.

The observer recorded 25 reviews, three notes, 12 stale results, three rejected results and one error. These counters describe this configuration’s activity; without a paired run with support disabled, they do not measure how much the support system improved the score.

Timing · server audio only

When the final answer finished.

Measured from the end of input playback to the end of the final spoken answer, across all eight cases:

Minimum
0.634 s
Median
1.973 s
Maximum
12.828 s

Offline transcription alignment is approximate. The separately recorded median time to next audio was 1.451 seconds, but that audio can be filler or repetition and the metric omits early speech. These are not PSTN response times, p95 measurements or an SLA.

Cost · modeled inference only

About 23¢ for
this eight-case run.

The estimate combines 238 seconds of provider-reported GPT-Live usage (about 19.8¢) with about 3.0¢ for the support model.

Speech is calculated at OpenAI’s published $0.05 per minute, billed per second. Support uses estimated rates of $1 per million input tokens and $2 per million output tokens.

This is a model-cost estimate, not a reconciled provider invoice or a customer price. It excludes carrier charges, local transcription, hosting and other infrastructure costs.

Reproducibility and limits

Inspect the scorecard.
Keep the scope intact.

The download contains every case result, review rationale, original automatic grade, configuration, timing, cost basis and evidence hash. Hashes identify retained artifacts; raw recordings and execution records are not linked in this release.

Frozen selection and preserved artifact hashes
Dataset revision
af7bb9c25b015792583ca4da3ee27ec62cb79fe6
Selection seed
das-bba-pilot-v1
Selection method
sort SHA256(seed:id) independently per category; interleave categories
Frozen manifest SHA-256
dbe963a384577bc62856937f2644a524d025b9c2fde1f5aab8e7d9b95c5f381a
Original automatic summary SHA-256
076d4296cdfaa28b5a8d46dbdb01a40a3418a29df36c418850179c1f1cc1aa2f
Review decisions SHA-256
2137c41318fb84de294b6f42ec4959aa7dc41d6dd9c6b113eab7b0ac014e0272
Secondary ASR SHA-256
aa529bed79260bfe7b550b0a084dda5d801582341eb7024cefbb4a4cb0a85634

Individual input-audio, output-audio, record and original-grade hashes are included per case in the JSON. A hash alone is not an independently verified result.

What this pilot does not establish

Full-dataset accuracy, independent leaderboard standing, live business-tool success, booking or CRM accuracy, carrier reliability, human-handoff quality, or a causal advantage from subconscious support. Tools were disabled, the sample was eight questions, and the review was not independent human listening.