guides/transcription-adapter-testing

Transcription adapter testing

Status: Current-source coverage reconciled September 16, 2026. The implementation is in isolated worktrees and has not been released. Test totals below describe specific recorded runs, not universal acceptance of the entire design.

Current verification

Latest merge check: the root Hudson package on origin/main 4a60eb2 passed 72 transcription tests and 27 existing HudsonVoice tests with published Vox 9190158 and no local Vox override. Run the root-package commands in the merge validation record. The table below preserves earlier, separately scoped acceptance runs.

CheckRecorded resultEvidence limit
Core, cloud, ElevenLabs current-source harness38 tests across 6 suites passedInjected transports; no vendor audio requests
Direct FluidAudio harness6 tests passedLifecycle/runtime fixtures
Native Parakeet V3File and caller-fed live inference passed on a generated sentenceDoes not prove physical microphone or product insertion behavior
TalkieTranscription host27 reported tests passed; native opt-in skippedIncludes two parameterized completion tests, four cases each; read-only activity inspection; fixed capture routing
TalkieAgent6 tests previously passedRouting fixtures, not physical capture acceptance
Full Talkie validation buildPassed after the latest host completion fixIsolated Termini header packaging; published dependency path remains unverified

The acceptance audit and chronological evidence contain scope and log paths. Earlier counts in the chronological report are historical and are superseded by later runs.

Current executable coverage

SourceAssertions exercised
HudsonTranscriptionTestsDuplicate registration, unknown saved IDs, fingerprints, optional annotations, unsupported/unverified compatibility, accepted IDs, cancellation distinctions, partial replacement, bounded PCM and terminal events
HudsonTranscriptionCloudTestsMAI model/options/normalization/no retry; Gemini upload ownership and cleanup, option rejection, dedicated versus conversational input, ordered PCM, drain timeout and remote uncertainty
ElevenLabsAdapterTestsPublic-contract registration and submission, words/speakers/provenance, unsupported hints before upload
TalkieTranscription WorkspaceTests and LiveWorkspaceTestsConsent, saved selection/run recovery, ordered audio, result identity, contradictory completion payloads, separate meeting-track recovery without duplicate upload

The acceptance matrix below remains a requirements checklist. Presence in that matrix does not assert a passing test for every row. In particular, download cancellation, automatic session rotation, shared cross-adapter resource ownership, and physical UI/capture behavior are not established by the checks above.

Run the recorded checks

The task-owned contract harness symlinks this worktree's current source and tests. It excludes native SDK targets. Run native checks separately.

swift test --package-path "$HOME/Library/Caches/codex-builds/hudson-transcription-contract-harness" \
  --scratch-path "$HOME/Library/Caches/codex-builds/hudson-transcription-contract-check"

For Talkie's host suite, run from the isolated Talkie worktree:

HUDSON_PACKAGE_PATH=/Users/arach/dev/hudson-worktrees/transcription-adapters \
  swift test --package-path apps/macos/TalkieTranscription \
  --scratch-path "$HOME/Library/Caches/codex-builds/talkie-transcription-host"

Native Parakeet acceptance requires a local model folder and an authorized recording. Set HUDSON_PARAKEET_MODEL_DIR and HUDSON_TRANSCRIPTION_ACCEPTANCE_AUDIO. The test does not record a microphone or send audio to a cloud service.

Historical HudsonVoice baseline

These Swift Testing suites exist in this Hudson checkout today. They cover current HudsonVoice behavior that the proposal says to preserve. They are not adapter contract tests.

Verified run, 16 September 2026, assigned origin/main worktree, exit 0. 27 tests in 5 suites passed. This validates existing behavior only, not proposed adapters. See the baseline verification report for scope and limits.

swift test --scratch-path "$HOME/Library/Caches/codex-builds/hudson-transcription-baseline" --filter HudsonVoiceTests
SuitePathWhat it exercisesOut of scope
HudDictation capture start gateHudDictationCaptureStartGateTests.swiftDuplicate begin is rejected, cancel invalidates delayed callbacks, finish is idempotent for the current generationPCM adapters, provider cancellation states
HudPendingUtteranceStoreHudPendingUtterancesTests.swiftHeld dictation audio moves onto durable storage, replays in capture order including sub-second captures, survives store relaunch, discards one utterance, and round-trips capture contextAdapter submit, remote jobs, meeting tracks
HudsonVoicePreferencesHudsonVoicePreferencesTests.swiftHudson preferences save and mirror into embedded Vox preferences, including transcription model idAdapter configuration fingerprints, credential resolvers
HudSpeechHudSpeechTests.swiftSpoken-output provider and credential helpersTranscription adapters
HudSpeechPlaybackHudSpeechPlaybackTests.swiftHost-lent TTS playback through Vox Apple SpeechTranscription adapters

Related current code, not covered by an adapter suite:

SurfacePathNote
HudDictationHudDictation.swiftEmbedded Parakeet plus Apple partials/fallback and parakeetOnly queue. Preserve callers. Existing embedded path remains separate from new adapter tests
HudAudioTranscriberHudAudioTranscriber.swiftApple Speech file facade. No matching test target was found in this checkout
HudVoxLiveSessionHudVoxLiveSession.swiftSends transcribe.startSession. Vox owns mic capture. Not a caller-fed PCM adapter
HudAudioRecorder formattingHudAudioRecorderTests.swiftDuration and file-prefix formatting only

Talkie meeting and engine tests live in the Talkie checkout (EngineService, MeetingAnnotation, AnnotationProviderFactory, MeetingStreamingTranscriptionSession). They were not inventoried here. Preserve those product paths during migration. Do not treat their absence from this matrix as proof they do not exist.

The new HudsonTranscriptionTests target exists alongside this historical baseline.

Required contract acceptance matrix

Use these IDs to track design requirements. Map each claimed pass to a named executable assertion or a recorded native/vendor run. The coverage inventory above identifies implemented checks; remaining rows retain their requirements. Do not infer a pass from a descriptor or a successful build.

Normalization and provenance

IDSetupActionObservable result
C-NORM-1Vendor payload with transcript onlyMap to the shared result typeTranscript present. Timing, speaker, and confidence fields absent, not zeroed
C-NORM-2Vendor payload with word timingsMap to the shared result typeAudio-relative offsets preserved. No wall-clock timestamps invented
C-PROV-1Successful batch jobRead provenanceProvider, model, adapter version, secret-free configuration fingerprint, source digest, run identity, timestamp
C-PROV-2Provider returns a request IDRead provenanceProvider request ID present and handed to the host persistence seam
C-PROV-3Derived speaker labels from a later stepRead provenanceAnnotations marked derived, not native
C-PROV-4Saved model id missing from current discoveryEnumerate models, then load saved configurationIdentifier round-trips. Readiness is unavailable. No silent substitute

Request option conflicts

Evaluate the whole request, including duration and feature combinations, before prepare or upload.

IDSetupActionObservable result
C-OPT-1Gemini file model, speaker labels plus vocabulary hintscompatibilityUnsupported, with a machine-readable reason. No upload
C-OPT-2Gemini file model, speaker labels plus smart formattingcompatibilityUnsupported, with a distinct reason. No upload
C-OPT-3Gemini annotated file longer than 30 minutescompatibilityUnsupported for duration. No automatic chopping
C-OPT-4MAI-Transcribe-2 file with diarization around 15 minutescompatibilityNot represented as unconditionally ready. Unknown or documented risk is visible
C-OPT-5Dedicated gemini-3.5-transcribe-live session past documented ten-minute limitcompatibility or rotation policyReject or rotate with caller-owned offsets. Do not apply this limit to gemini-3.8-live
C-OPT-6Unknown context or session limitcompatibilityUnverified, not treated as unlimited
C-OPT-7Request for MAI-Transcribe-1 or 1.5Model resolveRejected. MAI-Transcribe-2 is required. No implementation fallback
C-OPT-8Request that names gemini-3.5-transcribe, gemini-3.5-transcribe-live, or gemini-3.8-liveModel resolveEach id stays distinct. No silent replacement by a conversational stand-in

Cancellation and unknown remote outcomes

IDSetupActionObservable result
C-CAN-1Local batch in progressCancelWork stops. Resources released. Completion status is cancelled
C-CAN-2Live session after some chunksCancelLate writes rejected. No terminal success. Resources released
C-CAN-3Remote job accepted, cancel supportedRequest cancelStatus is cancellation requested, then cancelled when the vendor confirms
C-CAN-4Remote job accepted, vendor cancel unsupported or unansweredRequest cancel, then time outStatus is remote outcome unknown. Host still has the provider request ID
C-CAN-5Timeout after remote accept, no idempotency keyObserve retry policyAdapter does not resubmit. Coordinator owns retry
C-CAN-6Local model download in progressCancel downloadModel is not marked installed

Partial and final sequencing

IDSetupActionObservable result
C-LIVE-1Live session, several revisions of one utteranceEmit eventsProvisional revisions replace in place. Sequence numbers increase
C-LIVE-2Same session, utterance finalizedEmit eventsOne finalized utterance. Session remains open
C-LIVE-3Finish inputDrain within deadlineExactly one terminal result or error
C-LIVE-4Audio loss mid-sessionFinish or cancelVisible incomplete result, not silent success
C-LIVE-5Session rotationOpen the next sessionCaller-owned source offsets continue. Speaker IDs do not silently continue across sessions
C-LIVE-6Dictation live pathReplaceable partials, then stopOne final utterance, no duplicate insertion
C-LIVE-7Dictation batch-on-releaseStop capture, then transcribeResult labeled as batch, not as live partials

Adding a provider without changing core or Talkie

IDSetupActionObservable result
C-EXT-1Example remote package (ElevenLabs file)Compile against the public contract onlyNo TalkieKit or Hudson core source change
C-EXT-2Example local package (file, batch only)Compile against the public contract onlySame as C-EXT-1. Streaming not published
C-EXT-3Both example packages registered at app compositionRender picker and result viewerNew rows appear from registration. No provider-specific branches
C-EXT-4Conformance fixtures from C-NORM through C-CANRun against each example packageSame fixture set, no core edits
C-EXT-5Future Vox adapter attempt that only starts transcribe.startSessionFeed a Talkie meeting trackMust fail the caller-fed file/PCM requirement. Starting a second microphone is not a pass

Live provider acceptance

Run only with user-approved fixtures and explicit credentials from the host resolver. Do not embed secrets. Do not send diagnostic audio from a readiness check. An explicit sample transcription is a separate action.

Record provider, model, adapter version, API version, source digest, exact options, latency, and vendor usage when supplied. Recheck official catalogs before a live run. If the configured model is unavailable, record unavailable. Do not silently downgrade.

Required named targets:

Provider pathModel idModeStart condition
Microsoft MAI fileMAI-Transcribe-2Batch fileRequired. Do not use 1 or 1.5
Gemini dedicated filegemini-3.5-transcribeBatch fileRequired and distinct
Gemini dedicated livegemini-3.5-transcribe-liveLive PCMRequired and distinct
Gemini 3.8 Live input transcriptiongemini-3.8-liveLive PCM evaluationRequired as its own target. Not a replacement for dedicated Transcribe Live
FluidAudio localthe shipped dependency's modelFile and, if verified, streamDirect adapter, no Vox hop
ElevenLabs referencevendor file modelBatch fileExample package
Talkie ElevenLabs / Deepgramcurrent meeting providersMeeting tracksPreserve existing product paths

MAI Voice Live is a separate implementation path. Leave it unverified until its contract is checked. Do not copy MAI file capabilities onto it.

Gemini 3.8 Live captions-only evaluation must record response behavior, transcript completeness, finalization, latency, and billed usage. Discarding generated audio does not prove generation or its cost is disabled. Speaker labels and word timings stay unverified on this path until measured.

Use-case fixtures

CaseInputExpected behaviorRequired evidence
DictationShort push-to-talk recordingBatch-on-release is valid and labeled honestly. A separate live path shows replaceable partials and one final utteranceCaptured audio, transcript, latency, stop/cancel, no duplicate insertion
Existing recordingSeveral-minute user-approved fixture with names and punctuationCompare MAI-Transcribe-2, Gemini 3.5 file, local FluidAudio, and the ElevenLabs reference adapter on the same fileSame input digest, exact options, output, provenance, measured time, vendor usage where supplied, human-checked errors
MeetingShort multi-speaker fixture plus a 45-minute fixturePreserve mic and system track origins. Reject unsupported long input before upload. No false speaker continuity. Original recording is not lostSpeaker/timing evaluation only on supported paths. 45-minute rejection evidence on constrained models

Remaining product acceptance

The three use-case fixtures above remain in scope. The short local native tests do not replace the several-minute recording comparison or the meeting fixture. The two-track workspace regression verifies ledger recovery, successful-track reuse, and suppression of duplicate uploads after a partial failure. A separate production-helper harness now exercises adapter word mapping, the real merger, track separation, object serialization and revision increments. The database extension exercises the production GRDB publication helper, including concurrent edits and rollback on an injected outbox failure. It does not start the sweep scheduler or sync worker. See the progress report for its command and results.

Account-backed ElevenLabs short-file acceptance passed using the secret-cli meetings key; see the acceptance audit. Gemini acceptance remains pending. MAI via OpenRouter subsequently passed live Swift-adapter acceptance after user authorization. The following earlier checks covered Talkie stores only. Checked stores did not provide readable MAI/Azure or Gemini credentials and MAI endpoint configuration. The ElevenLabs development key lookup returned an authentication failure, which does not prove the key is absent. Resolve credentials through the host before testing; do not print them or add them to fixtures.

For each live path, record the actual model, source digest, options, result, provider request ID, completion behavior and supplied usage. A successful API request is evidence for that request, not a comparative quality ranking.

Settings and failure states to cover

The production mapping/database harness reports six passing tests after adding dictation library publication. It verifies pending insertion, title/notes preservation, completion immunity to a late failure, retained audio references on failure, and no restoration of a deleted row. These are real GRDB tests of the production helper; they do not prove microphone capture, library rendering, or audio playback in the running app.

Native acceptance should include opening an active adapter dictation after more than two minutes and attempting retranscription from another surface. The UI and service now use controller ownership rather than the pending-age heuristic for this case. Also cancel a capture and verify the guard remains in place until the audio writer closes. These interaction checks are not claimed by the database or workspace fixture suites.

Talkie has a native adapter settings view. Source and build checks do not prove its interaction states. Exercise the following states through applicable fixtures and native UI acceptance; do not claim all are verified current UI.

The activity disclosure lists unfinished requests from the app workspace and a read-only snapshot of the recording agent's ledger. It shows source and track identities, configured provider/model, and any provider request ID. An uncertain remote outcome explains the duplicate-upload risk. A persisted submitting status says that a result is pending; it does not claim the process is alive. Refresh reads state only and never retries an upload. Missing ledgers are empty; unreadable ledgers show an error rather than a false empty result. Fixtures cover those distinctions and verify byte-for-byte preservation of inspected ledgers. Native interaction acceptance is still required.

No registered adapters; missing credential; expired credential; model missing; downloading; cancelled download not installed; unavailable platform; unsupported language; incompatible options; unverified capability; offline; rate limited; cancelled; remote outcome unknown; completed; incomplete audio; plugin package unavailable after restart.

Use keyboard-operable native controls and text explanations. Do not encode state in color only.

Deferred or unsupported in this design

  • Runtime installation of arbitrary executable plugins
  • Automatic composition of a transcription engine plus a second diarization engine
  • Automatic audio chopping to bypass duration limits
  • Silent local-to-remote fallback
  • Quality ranking without the recording-fixture measurements
  • MAI live, until its contract is checked and tested
  • Switching existing HudDictation or Vox daemon callers onto the new contract in the same change that introduces the contract

Longer native Parakeet file acceptance

The Talkie host test nativeParakeetLongFileAcceptance accepts HUDSON_PARAKEET_LONG_AUDIO and HUDSON_PARAKEET_MODEL_DIR. Supply at least 180 seconds of generated speech. It checks file compatibility beyond the live limit, completed output, provenance, more than 300 timed words, ordered starts, and a final word within ten seconds of the file duration. It does not send that long file through the live test. Run the host suite with --filter nativeParakeetLongFileAcceptance and the two environment variables.

On 2026-09-16, a 185-second WAV made by repeating the existing generated sentence passed in 53.390 seconds including model preparation. FluidAudio used its native file chunk processor. Log: /tmp/talkie-parakeet-long-file.log. This verifies long-file completion for that synthetic fixture, not accuracy on diverse speech or meeting diarization. No new model installation was needed.

For AI agents