Architecture
Voice Interaction Architecture
How on-device speech transcription and a hosted voice model wrap an evidence-grounded assistant: the visitor's voice stays in the browser, the assistant only ever sees text, and validated answers are read aloud on request.
Published October 5, 2026
The Photography Curator and the Engineering Ask on my portfolio can be spoken to and can speak their answers. Neither assistant knows that. Speech recognition turns the visitor's voice into an editable draft in the same text box they could have typed into, and speech synthesis reads the finished, validated answer aloud afterwards. This case study is about keeping those two systems outside the assistant: a small model on the visitor's device for listening, a hosted voice model for speaking, and text as the only thing that crosses into the system that does the reasoning.
The problem
Typing is not always the best way to explore a photo archive. "Show me quiet streets at night" is easier to say than to type, and an answer about a set of photographs is pleasant to hear while you look at them. But voice brings problems that text doesn't have:
- Microphone audio is personal. It can contain things the visitor never meant to send, and sending it to a server is a different promise from sending a typed question.
- Two languages. The site works in English and German. A model that mishears German as English produces a confident, wrong question.
- Browsers differ. Audio playback, microphone access and on-device inference behave differently across Chrome, Firefox and Safari, and across desktop and mobile.
- The assistant is already hard to get right. The Curator is an evidence-grounded system with validators on everything it says. Making it voice-aware would multiply the surface those validators have to cover.
The design constraint
The assistant stays text-native. Speech is an adapter on both sides of it:
- Input: audio → transcription → an editable text draft → the existing submit path.
- Output: the validated text answer → speech synthesis → audio.
The assistant request carries the question and nothing else about how it was produced: there is no voice flag. Nothing the assistant does depends on whether the answer will be read, heard, or both.
Architecture at a glance
INPUT · on the visitor's device
Microphone
│ raw audio, memory only
▼
Audio capture (AudioWorklet)
│ 16 kHz mono samples
▼
Whisper-tiny in a Web Worker
WebGPU, else WebAssembly
│ transcript
▼
Editable draft ◄── Keyboard
│ reviewed, then sent
▼
════════════ TEXT ════════════
│
▼
Existing grounded assistant
retrieval · evidence ·
validation
│
▼
Validated text answer
├─────► shown as text
│
▼ only when asked to speak
Narration text
markup removed, segmented
│ signed answer text
▼
Speech route (my server)
│
▼
Grok Voice, voice "eve",
via OpenRouter
│ streamed PCM audio
▼
Web Audio playback
- Capture and transcription run entirely in the browser. The microphone stream feeds a small audio processor that collects samples until the visitor stops, then hands them to a background worker that runs the speech model.
- The draft is the same text field the visitor types into. The transcript is never submitted automatically, so recognition mistakes are visible and fixable before anything reaches the assistant.
- The assistant is unchanged by voice. It answers a string.
- Narration starts from the exact answer the validators approved. A formatter removes formatting that should not be spoken, such as citation markers and links, and splits the text into short segments so the first sound arrives quickly. Nothing rewrites the answer.
- Synthesis happens on my server's speech route, which calls the hosted voice model and streams the audio back as it is generated.
- Playback uses the browser's Web Audio API, scheduling each segment back to back.
Input: local transcription
The transcription model is Whisper-tiny, multilingual, run with Transformers.js and ONNX Runtime Web in a dedicated Web Worker, so the page stays responsive while it works. It tries WebGPU first and falls back to WebAssembly when WebGPU is missing, fails to initialise, or has already failed earlier in the session.
Where the model comes from. The weights and the runtime are served from my own site, from a fixed list of versioned files, not fetched from a model hub or a third-party CDN at runtime. The browser caches them. The first use downloads about 66 MB of weights; later visits load from cache.
Precision. The first shipped version used full precision throughout, a 154 MB download. Quantising the whole model failed the English quality bar (88.5% intent preservation against a 95% gate). Measuring further showed English quality depends on the encoder's precision and not the decoder's, so the shipped version keeps a full-precision encoder and a quantised decoder: 57% smaller, first-use preparation roughly halved, quality kept.
When it is offered. Voice input appears only where it can work: a secure page, microphone and audio-processing APIs, Web Workers and WebAssembly. It is deliberately disabled on touch-first devices such as phones. Mobile was not measured as viable for on-device inference, so the microphone button doesn't appear there and typing is the path. The model download starts only after the visitor grants microphone permission, so declining costs nothing.
Why transcription runs on the device
The goal was stated before any model was chosen: transcribe English and German entirely on the visitor's device, with no raw audio leaving the browser and no new provider call. The reason is privacy. A spoken question can contain more than the visitor intended, and the safest audio to protect is audio that never leaves the device.
That choice had costs, which I accepted and measured instead of hiding:
- A one-time download of about 66 MB.
- Cold start. First-use preparation took about 7.3 seconds after optimisation, down from about 14.5. Warm, transcribing a short question takes around one to one and a half seconds on WebAssembly and about 0.6 seconds on WebGPU.
- A small model's accuracy. The draft is editable precisely because a tiny model makes mistakes.
I serve the files myself for supply-chain and privacy reasons, not speed: the browser fetches exactly the pinned files I deployed, from my origin, and nothing else. There is deliberately no automatic fallback to a cloud transcription service.
The assistant boundary
Typed and spoken questions converge on one string, in one text field, submitted through one handler. The assistant request has the question and its optional conversation context. It has no field that says "this came from a microphone".
That boundary is the main architectural decision, and it pays off in several ways:
- One reasoning path. Retrieval, evidence selection and the validators are the same for every question. Nothing needed re-validating when voice arrived.
- Testing stays text-based. The assistant's benchmarks and evaluations still measure text in, text out.
- Replaceable parts. The transcription model and the voice model can each change without touching the assistant. The voice is configuration, not code.
- Graceful degradation. If voice input is unavailable, the text box still works. If speech fails, the written answer is already on screen.
- Text stays canonical. What the visitor reads is what was validated; what they hear is a rendering of that, never a separate generation.
The same voice input module serves both assistants, and it is deliberately domain-blind: its contract ends at a transcript string, and it imports nothing from either product.
Output: spoken answers
Narration is off by default and starts only when the visitor asks for it, either per answer or with an "auto" mode that speaks each new answer. It exists only when the answer was actually produced: refusals are never voiced.
The voice model is Grok Voice (grok-voice-tts-1.0) with the voice "eve", reached through OpenRouter's speech endpoint. The server sends text and receives raw 24 kHz PCM audio, streamed as it is generated. Raw PCM means there is no audio codec for the browser to decode, so the format behaves the same everywhere.
Only real answers can be spoken. Each answer carries a signed grant over its exact text and language. The speech route refuses any text it did not itself produce, and is rate-limited. It is not a general-purpose text-to-speech endpoint, so nobody can use it to synthesise arbitrary text at my expense.
Streaming and segmentation. The answer is split into short segments, the first deliberately small. The browser requests segments in a bounded pipeline and starts playback when about a third of a second of audio has arrived. Each segment is scheduled immediately after the previous one, so the answer plays continuously.
Controls and lifecycle. Listen, pause, resume and replay are one control per answer. A new question, closing the panel, or switching to a different photograph stops the current narration and cancels its in-flight requests. Only one narration plays per assistant at a time. Nothing is cached: replay requests the audio again.
Why input and output use different models
The two directions have different jobs, and the model sits where that job is best done.
Listening is optimised for privacy and immediacy. The input is the visitor's own voice; the output is a draft they will read before sending. A small model that runs locally, makes visible mistakes and costs nothing per use is the right trade.
Speaking is optimised for quality and language. The input is text that is already public to the visitor; the output must sound natural in English and German. I measured three hosted voice models on both languages (144 synthesis runs, no failures), scoring each by how well a separate speech recogniser could recover the words:
- Deepgram Aura-2 had the fastest first audio (316 ms median) but the weakest German (word error rate 0.093).
- Microsoft MAI-Voice-2 had the best accuracy, but its first audio took over a second (1,052 ms).
- Grok Voice was the only one with German as intelligible as English (0.043 against 0.047), a first byte around 450 ms, continuous streaming, one recognisable voice in both languages, and half Aura-2's cost per answer.
A human listening pass then accepted it. A small open model was cheaper by an order of magnitude but spoke English only. On-device synthesis was not evaluated as a production option. At about $0.007 per spoken answer, quality was worth more than eliminating the cost.
Language handling
Language flows through the two directions differently, and on purpose:
- Transcription is told the page's language. Whisper's own language detection badly mis-decoded short German questions: intent was preserved in 9.5% of German test utterances with auto-detection, against 95.2% with the language forced. So the site's locale (English or German) decides how the model listens, and there is no second detector.
- The answer's language is detected deterministically from the question text, the same way as for typed questions, and mirrors it.
- The narration language is that detected language, bound into the signed grant together with the answer. The server picks a voice per language from configuration; today both languages use "eve", which speaks both. The provider is not given a language parameter; the voice handles both.
- Anything else is refused: the system supports exactly English and German, end to end.
Privacy boundaries
Four kinds of data travel through this system, and each has its own boundary:
- Raw microphone audio stays in the visitor's browser. It lives in memory between the audio processor and the worker, and is discarded after transcription. No route on my server accepts recorded audio.
- The transcript stays on the device until the visitor chooses to send it. It is then an ordinary typed question.
- The assistant's answer is generated on my server and shown to the visitor. If they ask to hear it, its text is sent to the voice provider through my server.
- Generated audio is streamed back to the browser and played. It is not stored.
So the honest claim is narrow: listening is local; speaking uses a hosted voice model on the answer text. The visitor's voice never leaves their device. The assistant's words do, when the visitor asks to hear them.
Latency
A spoken exchange is a pipeline, and its delay is the sum of very different stages:
- Capture finalisation (local): the visitor stops; the audio is handed to the worker. Starting to listen takes 42–66 ms when warm and about half a second on first use, which is permission and audio setup, not the model.
- Transcription (local, model): about 0.6 s on WebGPU and 1.2–1.4 s on WebAssembly for a short question, warm.
- The visitor's review (human): deliberately not automated.
- Assistant generation (network and model): streamed, and covered in the Photography Assistant case study.
- Speech synthesis (network and model): about 450 ms median to the first audio byte from the provider; around 0.5–0.7 s through my route in a browser.
- Playback start (browser): about 0.35 s of audio is buffered before sound begins, so the voice doesn't stutter.
These numbers come from separate, dated measurements of individual stages. The system has no end-to-end voice latency instrumentation. The client has no telemetry, and the server times only the synthesis call. That is a stated gap, not a hidden one.
Failure isolation
Voice can fail in more places than text, so every voice failure stays a voice failure:
- Microphone denied: the button explains it and typing remains; the model is never downloaded.
- Unsupported browser or a phone: the microphone button doesn't appear.
- Model load or transcription failure, or silence: a short status message, and the text box is untouched.
- Wrong transcript: visible in the draft and editable before sending.
- Voice not configured on the server: no Listen control appears at all.
- Synthesis or network failure: the control shows that the voice is unavailable. The written answer, already on screen, is unaffected.
- Audio blocked by the browser: the voice is not requested at all (so it costs nothing), and a hint asks the visitor to tap Listen.
- New question, closed panel, unmount: the narration stops and its requests are cancelled.
The assistant never fails because of voice, and voice never changes what the assistant said.
Browser constraints
- Voice input is offered on desktop browsers that provide the APIs above. WebGPU is used where it works and WebAssembly elsewhere. My measurements were run in Chromium; real Safari and Firefox runs of on-device transcription are not documented, and phones are excluded by design.
- Spoken answers work in Chrome, Firefox and Safari, after a cross-browser bug I describe below. Browsers require a user gesture before audio can play. Every answer follows a click or a key press, so in practice this only matters for the automatic mode, where a blocked start shows a hint instead of failing silently.
What failed
Spoken answers were silent in Firefox and Safari. Problem: the player checked whether the audio engine was running immediately after asking it to start. Experiment: the same production page in Chromium, Firefox and WebKit, each with a real click. Result: Chromium reported "running" at once. Firefox and WebKit report "suspended" until the engine actually starts, 3–88 ms later, so the player shut the audio down and never requested speech. Autoplay rules and the audio format were both ruled out. Consequence: a bounded wait of up to 500 ms, more than five times the slowest start measured, before deciding audio is blocked.
Making the Curator's camera react to its voice. Problem: I wanted the Curator's camera icon to visibly respond while it spoke. Experiment: an audio analyser on the playback path, a speech-energy envelope, and a shader parameter driven by it, all without React state per frame (about 0.4 µs per sample). Result: the shader parameter did not control the visible silhouette. With every relevant parameter pushed to its limit, the outline could move only 5.2% against the 20–40% needed. Spectrum analysis also showed 95% of speech energy below 390 Hz, so the frequency bands had to be re-placed. Consequence: the shader change was reverted. What shipped is simpler: a soft "membrane" of blurred shapes next to the icon, whose size follows eight speech-frequency bands. One animation loop writes CSS variables to a single element, only while the voice is speaking. A human review accepted it on first look. Under reduced motion it isn't rendered at all.
WebGPU: four attempts. Problem: WebGPU promised faster transcription. Result: the first configuration produced garbage and was 15.6 times slower. Two further attempts were blocked on quality. The fourth shipped WebGPU first, with a proven fallback to WebAssembly: 569 ms against 1,248 ms, warm. Consequence: WebGPU is an acceleration, never a requirement; a WebGPU failure marks it unhealthy for the rest of the session and transcription continues on WebAssembly.
Expressive speech tags. Experiment: adding explicit pause tags to give answers more natural rhythm. Result: the voice already paused about 0.6 s at sentence boundaries; the tags pushed that to over a second, which sounded slower, not better. Consequence: rejected, with no production change.
Leaks in the microphone lifecycle. An audit found four ways the microphone could stay open after the conversation moved on: a hidden panel kept its capture running; a second view could start a second capture alongside it; a permission grant that arrived after the component had gone published a live stream nobody would read; and an error left the microphone on behind the error message. A separate race could show "listening" before capture had connected. All were fixed and checked in a browser. The microphone is now released whenever the session leaves listening, and a late permission grant releases itself if the component is gone.
Trade-offs
- On-device model against download and runtime cost. Privacy and zero per-use cost, paid for with a 66 MB first download, a cold start of several seconds and a small model's accuracy.
- Hosted voice against network dependency. Natural German and English speech, paid for with a provider dependency, about $7 per thousand spoken answers, and a network round trip before the first sound.
- Quality against first-audio time. Grok was not the fastest first byte; Aura-2 was. German intelligibility decided it.
- Voice UX against browser complexity. Audio unlocking, WebGPU health, worker lifecycles and microphone release are all code that exists only because of voice.
- Modality independence against adapter state. Keeping the assistant voice-blind moved the complexity into two adapters, each with its own state machine, cancellation and cleanup. They are tested separately.
- No automatic cloud fallback. A browser that cannot run the model doesn't get voice input. I accepted that rather than silently sending audio elsewhere.
Result
What the architecture provides, and what I can demonstrate:
- A visitor can speak a question in English or German without their voice leaving their device.
- They can hear any validated answer in a natural voice in either language, or read it, or both.
- The assistant, its evidence boundary and its validators are untouched by voice. A spoken question and a typed one are the same request.
- Each voice component can be replaced (the transcription model, the voice, the provider) without touching the assistant.
- When voice fails, it fails alone.
Open items: end-to-end latency measurement, real-device testing beyond Chromium for transcription, and mobile voice input.
Further reading
The Photography Assistant case study describes the grounded assistant that both voice adapters wrap. The Photography page is where the Curator, and its voice, can be tried.