Voice
Overview
Import from hudsonkit/voice to add speech-to-text input, text-to-speech output, or full voice-driven assistant interaction. The subpath is opt-in: it lazily loads @voxd/client (the local Vox Companion) and pulls in voice hooks and reply-shaping utilities. The main hudsonkit package has zero voice dependencies — nothing ships to users who don't import this subpath.
Voice input requires the Vox Companion running locally on 127.0.0.1:43115. You can substitute your own STT provider by passing a transcribe function directly.
useVoiceInput
Captures microphone audio via MediaRecorder, ships it to Vox for transcription, and delivers the transcript via onTranscript.
'use client';
import { useVoiceInput } from 'hudsonkit/voice';
export function VoiceButton() {
const { status, error, start, stop, isSupported } = useVoiceInput({
onTranscript: (text) => console.log('transcript:', text),
});
if (!isSupported) return null;
return status === 'recording'
? <button onClick={stop}>Stop</button>
: <button onClick={start} disabled={status === 'transcribing'}>Speak</button>;
}
UseVoiceInputOptions
| Name | Type | Default | Description |
|---|---|---|---|
onTranscript | (transcript: string) => void | — | Called with the final transcript after recording stops. |
surface | string | "hudson-assistant" | Identifies the calling surface in Vox metadata. |
metadata | Record<string, unknown> | — | Free-form metadata included with every transcribe request. |
language | string | "en" | Spoken language hint. |
transcribe | TranscribeFn | — | Custom STT provider. If omitted, @voxd/client is loaded lazily. |
probe | () => Promise<boolean> | — | Availability probe override. Defaults to the Vox client's probe(). |
UseVoiceInputResult
| Name | Type | Description |
|---|---|---|
status | VoiceStatus | Current state of the input pipeline. |
error | string | null | Human-readable error message, or null. |
lastTranscript | string | null | Most recent transcript — useful for "draft ready" UI before send. |
start | () => Promise<void> | Request mic access and begin recording. |
stop | () => void | Stop recording and trigger transcription. |
isSupported | boolean | true if MediaRecorder and getUserMedia are available. |
useVoiceOutput
Synthesizes speech by posting to your app's /v1/audio/speech route and plays the returned audio.
'use client';
import { useVoiceOutput } from 'hudsonkit/voice';
export function SpeakButton({ text }: { text: string }) {
const { speak, stop, isPlaying } = useVoiceOutput();
return isPlaying
? <button onClick={stop}>Stop</button>
: <button onClick={() => speak(text)}>Play</button>;
}
SpeakOptions
| Name | Type | Default | Description |
|---|---|---|---|
provider | VoiceProvider | — | TTS provider override. |
model | string | — | Model identifier (provider-specific). |
voice | string | — | Voice identifier (provider-specific). |
rate | number | — | Speech rate multiplier. 1.0 = normal. |
metadata | Record<string, unknown> | — | Free-form metadata included with the speech request. |
UseVoiceOutputResult
| Name | Type | Description |
|---|---|---|
status | VoiceStatus | Current state of the output pipeline. |
error | string | null | Human-readable error message, or null. |
speak | (text: string, opts?: SpeakOptions) => Promise<void> | Synthesize and play text. Resolves once playback starts or fails. |
stop | () => void | Stop any in-flight playback and discard pending requests. |
isPlaying | boolean | true while audio is synthesizing or actively playing. |
useAssistantVoice
Combines useVoiceInput and useVoiceOutput into an AssistantVoiceKit — the prop type the built-in <Assistant> component consumes for mic + speaker UI and auto-reply playback.
'use client';
import { useAssistantVoice } from 'hudsonkit/voice';
const voiceKit = useAssistantVoice({
appId: 'my-app',
settings: { speakReplies: true },
});
Heads up:
<AppShell>does not currently forward avoiceKitprop to its built-in Assistant. Today this hook is the building block for wiring voice to a manually mounted<Assistant>or to your own UI. Threading through AppShell will follow.
UseAssistantVoiceOptions
| Name | Type | Default | Description |
|---|---|---|---|
appId | string | — | App ID included in voice metadata. |
settings | Partial<VoiceSettings> | — | Overrides for DEFAULT_VOICE_SETTINGS. |
AssistantVoiceKit shape
interface AssistantVoiceKit {
input: {
status: VoiceKitStatus;
error: string | null;
isSupported: boolean;
start: (onTranscript: (transcript: string) => void) => Promise<void>;
stop: () => void;
};
output: {
status: VoiceKitStatus;
error: string | null;
isPlaying: boolean;
speak: (text: string, opts?: Record<string, unknown>) => Promise<void>;
stop: () => void;
};
settings: {
autoSend: boolean;
speakReplies: boolean;
};
speakReply: (message: Pick<UIMessage, 'parts'>, metadata?: Record<string, unknown>) => void;
}
speakReply extracts text from a UIMessage, shapes it for speech (see Reply shaping), and calls output.speak. Call it from your chat's onFinish to automatically read assistant responses aloud.
Voice settings
VoiceSettings controls both input behavior and how replies are shaped for speech.
import { DEFAULT_VOICE_SETTINGS } from 'hudsonkit/voice';
VoiceSettings
| Name | Type | Default | Description |
|---|---|---|---|
autoSend | boolean | true | Auto-submit the transcript instead of just filling the input. |
speakReplies | boolean | false | Speak assistant replies aloud after streaming completes. |
replyProvider | VoiceProvider | "vox" | TTS provider for spoken replies. |
replyModel | string | "avspeech:system" | TTS model identifier. |
replyVoice | string | "" | Voice identifier. Empty string uses the provider default. |
replyRate | number | 1 | Speech rate multiplier. |
spokenReplyStyle | SpokenReplyStyle | "adaptive" | How much of the reply to speak: "brief", "adaptive", or "full". |
spokenReplyLongResponse | SpokenReplyLongResponse | "invite" | Handling for long replies: "summary", "invite", or "verbatim". |
spokenReplyCodeResponse | SpokenReplyCodeResponse | "summary" | Handling for code-heavy replies: "summary", "mention", or "read". |
spokenReplyMaxChars | number | 720 | Hard character cap for spoken output. |
VoiceProvider
"vox" is the only supported provider. Pass it as replyProvider or in SpeakOptions.provider.
VoiceStatus
| Value | Meaning |
|---|---|
"idle" | No activity. |
"recording" | Mic is active and capturing audio. |
"transcribing" | Audio has been sent; waiting for transcript. |
"ready" | Transcript is available. |
"synthesizing" | Speech request is in flight. |
"speaking" | Audio is playing. |
"unavailable" | Vox Companion is unreachable or warming up. |
"error" | A hard error occurred; see error string for details. |
Reply shaping
These utilities convert raw assistant text into speech-friendly output — stripping markdown, code blocks, and <think> tags, then applying length and style constraints.
import {
createHudsonSpokenReply,
getHudsonMessageDisplayText,
getHudsonVoiceBehaviorPreset,
applyHudsonVoiceBehaviorPreset,
} from 'hudsonkit/voice';
getHudsonMessageDisplayText(message) — Extracts and cleans the text content from a UIMessage. Strips <think> blocks and collapses whitespace.
createHudsonSpokenReply(text, policy) — The main shaping function. Pass the display text and either a VoiceSettings object or a SpokenReplyStyle string. Returns a speech-ready string (or "" if nothing remains after cleaning).
const spoken = createHudsonSpokenReply(displayText, voiceSettings);
if (spoken) voiceOutput.speak(spoken);
getHudsonVoiceBehaviorPreset(settings) — Identifies which named preset ("concise", "balanced", "detailed", or "custom") matches the given settings.
applyHudsonVoiceBehaviorPreset(settings, preset) — Returns a new VoiceSettings with the preset values merged in. Does not mutate the original.
const updated = applyHudsonVoiceBehaviorPreset(settings, 'concise');
Preset mappings:
| Preset | spokenReplyStyle | spokenReplyLongResponse | spokenReplyCodeResponse | spokenReplyMaxChars |
|---|---|---|---|---|
"concise" | "brief" | "invite" | "mention" | 360 |
"balanced" | "adaptive" | "invite" | "summary" | 720 |
"detailed" | "full" | "summary" | "summary" | 960 |
probeVoxAvailability
probeVoxAvailability is exported from the main hudsonkit package (not hudsonkit/voice) so you can check Vox status without pulling in voice dependencies.
import { probeVoxAvailability } from 'hudsonkit';
const availability = await probeVoxAvailability(voxClient);
// → "connected" | "warming" | "unreachable" | "blocked-origin"
Use this to gate voice UI before the user tries to record — for example, showing a "Install Vox" prompt when the result is "unreachable".
Apple (HudsonVoice)
Hudson's Apple SDK ships a Swift counterpart to hudsonkit/voice as the HudsonVoice target inside the HudsonKit Swift package. The product is part of the default package graph, so app code can import it without changing SwiftPM manifest flags.
Build-time code, runtime model
Build Hudson normally:
swift build
# The terminal stack remains optional:
HUDSONKIT_WITH_TERMINAL=1 swift build
HudsonVoice includes the engine integration code, but it does not embed the Parakeet model. Model data is acquired at runtime according to HudVoiceModelDownloadPolicy:
| Policy | Automatic behavior |
|---|---|
.never | Do not automatically download; use an installed model or Apple Speech fallback. |
.onFirstUse | Download and warm when dictation starts for the first time. This is the default. |
.eager | Download and warm when the host calls activate() for its voice surface. |
let dictation = HudDictation(modelDownloadPolicy: .onFirstUse)
// Call when the surface appears. This only acquires the model for `.eager`.
dictation.activate()
// An explicit user action can always request acquisition directly.
dictation.prepare()
Constructing HudDictation, importing HudsonVoice, and running package tests do not download model data.
Model attribution
The downloaded model is not Hudson's. Three layers, three owners:
| Layer | Source | License |
|---|---|---|
| Model | nvidia/parakeet-tdt-0.6b-v3 — 600M-parameter FastConformer encoder + TDT decoder, built with NVIDIA NeMo, trained on the Granary corpus | CC-BY-4.0 |
| Weights | FluidInference/parakeet-tdt-0.6b-v3-coreml — the Core ML conversion Vox fetches (~460 MB), from the FluidAudio project | CC-BY-4.0 |
| Runtime | Vox (VoxEngine) — downloads, caches, and runs the Core ML model on-device; the inference code is Vox's own | see repository |
Attribution is a condition of CC-BY-4.0, and an app shipping HudsonVoice is what causes those weights to land on a user's device — the condition follows the app, not just this repo. HudsonVoiceSettingsView renders the three names on-screen; NOTICE.md is the written form to copy into your own credits.
HudSpeechPlayback — Apple spoken output
HudSpeechPlayback is a thin Hudson facade over one Vox AppleSpeechOutputController per audible surface. It speaks canonical Vox synthesis requests and forwards truthful lifecycle/route events. Product policy — opt-in gating, dedupe, text shaping, fallback copy, credential storage — stays in the host.
Credentials are host-lent and snapshotted at initialization. Unlent providers are configured with explicit blank env keys so Vox cannot pick up ambient process secrets. There is no shared singleton and no pause/seek state. The event callback is not main-actor isolated; hop if you need to update UI.
Canonical system speech is avspeech:system. UI aliases such as system belong at the product boundary, not in this API.
import HudsonVoice
let playback = HudSpeechPlayback(
credentials: [
.openAI: openAIKey, // omit or leave blank to keep the provider fenced
],
onEvent: { event in
// Not MainActor-bound.
print(event.phase, event.requestId, event.modelId, event.provider?.label ?? "Unknown")
}
)
let requestId = await playback.speak(
"Hello from Hudson",
modelId: "avspeech:system"
)
await playback.stop()
await playback.cancel()
The same credential snapshot backs models() and voices(modelId:), so a
surface can populate its picker without constructing a second public facade.
Recreate the surface after a credential change.
HudSpeechSynthesizer remains the generation-only path that returns audio bytes. HudTTS is unchanged.
HudVoicePanel — SwiftUI primitive
HudVoicePanel is a drop-in SwiftUI view that renders the full Vox listen / stop / cancel UI in Hudson's design language (HudCard, HudButton, HudBadge, HudStatusDot). It owns its own HudVoxLiveSession, transcript buffer, and health probe lifecycle.
import SwiftUI
import HudsonVoice
struct VoxScreen: View {
var body: some View {
HudVoicePanel(
options: HudVoxLiveSessionOptions(clientId: "my-app")
)
}
}
Provide a custom endpoint to point at a remote Mac running Vox:
HudVoicePanel(
endpoint: HudVoxEndpoint(host: "macbook.local", port: 42137),
options: HudVoxLiveSessionOptions(clientId: "my-app", language: "en")
)
HudVoxLiveSessionOptions
| Name | Type | Default | Description |
|---|---|---|---|
clientId | String | "HudsonKit" | Identifies the calling surface in Vox metadata. |
modelId | String | "parakeet:v3" | Transcription model identifier. |
language | String? | nil | Spoken language hint. |
mode | HudVoiceMode | .pushToTalk | .pushToTalk or .alwaysOn. |
metadata | [String: String] | [:] | Free-form metadata sent with the session. |
Connection state — iOS pairs with a Mac
Vox is a local daemon that today runs on macOS. On macOS the panel reaches ws://127.0.0.1:42137. On iOS there is typically no Vox daemon on-device — the device pairs with a nearby Mac running Vox, and HudVoxEndpoint should point at that host.
When Vox is unreachable, HudVoicePanel surfaces an OFFLINE badge and a message like "Vox is not reachable at HudVoxProbe.health(...). The Listen button auto-runs the probe before connecting and short-circuits to offline if no health is returned.
Cross-device pairing (discovery, trust, transport) will be handled by the forthcoming HudPairing primitive — see docs/next-up.md. Until then, set the host manually.