Streaming Voice Agent
Combine live transcription, streamed chat, and streaming speech for low-latency voice turns.
The Voice Interface guide handles one audio file per request. This guide removes most of the waiting by overlapping the three stages: live transcription shows words while the person speaks, a streamed chat reply starts as soon as the transcript is final, and streaming speech turns each finished sentence into audio on one WebSocket connection.
Validate before production
The examples below combine three billed services. Run them end to end against your own account and wallet before you ship, and adapt the error handling to your application. See Streaming for how the individual APIs compare.
How a turn flows
- The browser asks for the microphone and then your server for a Scribe ticket.
- The browser streams 100 ms PCM16 frames and shows
transcript.partialtext. - On
transcript.completed, the browser POSTs the final text to your server. - The server streams the chat reply and cuts it into complete sentences (at
።,?,!or a newline). - Each sentence is spoken in order on one text-to-speech WebSocket connection per user session, with a new request ID per sentence.
- The server forwards the MP3 bytes to the browser in a chunked HTTP response.
- The browser plays them in order.
Before you start
- Keep your API key on the server as
ADDIS_API_KEY. The browser never sees it; it receives only a one-use Scribe ticket. - Scribe live transcription recognizes Amharic only, so this guide uses the Amharic voice
am-hamen. - Ask for microphone permission before you request the ticket. Tickets expire after 60 seconds.
- Keep API keys and tickets out of WebSocket and page URLs.
- Install the SDK on the server.
Build it
Server: create a Scribe ticket
This Express route calls the Scribe sessions endpoint with your API key and returns only the socket URL and ticket. The request ID is generated on the server.
// Example — validate against your account before production
import { randomUUID } from "node:crypto";
import express from "express";
const app = express();
app.use(express.json());
app.post("/api/scribe-ticket", async (req, res) => {
// Authenticate your own user here and enforce your spending policy.
// Store this ID with your user's turn so an interrupted transcript can be
// recovered with GET /api/v1/scribe/requests/{id}.
const requestId = randomUUID();
const response = await fetch("https://api.addisassistant.com/api/v1/scribe/sessions", {
method: "POST",
headers: {
"x-api-key": process.env.ADDIS_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
backend: "standard",
chunk: "320ms",
request_id: requestId,
}),
});
if (!response.ok) return res.status(502).json({ error: "Could not create a session" });
const { data } = await response.json();
res.set("Cache-Control", "no-store"); // Keep tickets out of caches and logs.
res.json({ websocket_url: data.websocket_url, token: data.token });
});Browser: stream the microphone
Mirror this order: microphone first, then the 16 kHz audio context and worklet, then the ticket, then the socket. The worklet posts 100 ms frames of 1,600 samples.
// Example — validate against your account before production
// scribe-pcm-worklet.js (served from your site)
class ScribePCM extends AudioWorkletProcessor {
constructor() {
super();
this.buffer = new Float32Array(1600);
this.offset = 0;
this.port.onmessage = (event) => {
if (!event.data?.flush) return;
if (this.offset) {
const tail = this.buffer.slice(0, this.offset);
this.port.postMessage(tail, [tail.buffer]);
this.offset = 0;
}
this.port.postMessage(null); // Tells the page the last frame was sent.
};
}
process(inputs) {
const input = inputs[0]?.[0];
if (!input) return true;
for (const sample of input) {
this.buffer[this.offset++] = sample;
if (this.offset === 1600) {
this.port.postMessage(this.buffer, [this.buffer.buffer]);
this.buffer = new Float32Array(1600);
this.offset = 0;
}
}
return true;
}
}
registerProcessor("scribe-pcm", ScribePCM);// Example — validate against your account before production
function pcm16(samples) {
const buffer = new ArrayBuffer(samples.length * 2);
const view = new DataView(buffer);
samples.forEach((sample, i) => {
const clamped = Math.max(-1, Math.min(1, sample));
view.setInt16(i * 2, Math.round(clamped * (sample < 0 ? 32768 : 32767)), true);
});
return buffer;
}
async function listen({ onPartial, onFinal }) {
// 1. Microphone permission first, so the 60-second ticket window is usable.
const stream = await navigator.mediaDevices.getUserMedia({
audio: { channelCount: 1, echoCancellation: true, noiseSuppression: true },
video: false,
});
// 2. A 16 kHz context, and a check that the browser honored it.
const context = new AudioContext({ sampleRate: 16000 });
if (context.sampleRate !== 16000) {
throw new Error("This browser cannot capture at 16 kHz.");
}
await context.audioWorklet.addModule("/scribe-pcm-worklet.js");
await context.resume();
// 3. Fetch the ticket from your server.
const response = await fetch("/api/scribe-ticket", { method: "POST" });
if (!response.ok) throw new Error("Could not get a Scribe ticket");
const ticket = await response.json();
// 4. Open the socket and authenticate in the first frame.
const socket = new WebSocket(ticket.websocket_url);
let node;
socket.onopen = () => {
socket.send(JSON.stringify({ type: "session.authenticate", token: ticket.token }));
};
socket.onmessage = ({ data }) => {
const event = JSON.parse(data);
if (event.type === "session.created") {
// 5. Only now start sending audio.
const source = context.createMediaStreamSource(stream);
node = new AudioWorkletNode(context, "scribe-pcm");
const mute = context.createGain();
mute.gain.value = 0;
node.port.onmessage = ({ data: frame }) => {
if (frame && socket.readyState === WebSocket.OPEN) socket.send(pcm16(frame));
};
source.connect(node);
node.connect(mute);
mute.connect(context.destination);
} else if (event.type === "transcript.partial") {
onPartial(event.text); // Provisional; it can change as the person speaks.
} else if (event.type === "transcript.completed") {
onFinal(event.data.text);
socket.close();
} else if (event.type === "error") {
console.error(event.error?.message);
socket.close();
}
};
// Call this when the person stops speaking (a button, or your own VAD).
return function stop() {
node?.port.postMessage({ flush: true });
setTimeout(() => {
if (socket.readyState === WebSocket.OPEN) {
socket.send(JSON.stringify({ type: "audio.finish" }));
}
stream.getTracks().forEach((track) => track.stop());
context.close();
}, 250);
};
}The final text arrives in transcript.completed as data.text, together with data.usage. A session accepts at most 180 seconds of audio, so stop recording before then.
Server: stream the reply as speech
This route streams the chat reply, cuts it into complete sentences, and speaks each sentence on a single text-to-speech WebSocket connection that belongs to the user's session. Audio bytes go straight to the browser as they arrive.
// Example — validate against your account before production
import { randomUUID } from "node:crypto";
import AddisAI from "addisai";
const addis = new AddisAI();
// One text-to-speech connection per user session, with its usage so far.
const sessions = new Map(); // sessionId -> { connection, openedAt, characters, turns }
const queues = new Map(); // sessionId -> promise of the previous turn
// Reuse the connection until it nears a session limit, then open a fresh one.
async function getConnection(sessionId, sentence) {
let s = sessions.get(sessionId);
if (
s &&
(Date.now() - s.openedAt > 9 * 60 * 1000 || // 10-minute session
s.characters + sentence.length > 5000 || // cumulative characters
s.turns >= 100) // request IDs
) {
s.connection.close();
s = undefined;
}
if (!s) {
s = {
connection: await addis.realtime.connect({
voiceId: "am-hamen",
language: "am",
maxTextCharacters: 5000,
}),
openedAt: Date.now(),
characters: 0,
turns: 0,
};
sessions.set(sessionId, s);
}
s.characters += sentence.length;
s.turns += 1;
return s.connection;
}
// Split off complete sentences; keep the unfinished tail in the buffer.
function takeSentences(buffer) {
const sentences = [];
let rest = buffer;
let match;
while ((match = rest.match(/^[\s\S]*?(።|\?|!|\n)/))) {
const sentence = match[0].trim();
if (sentence) sentences.push(sentence);
rest = rest.slice(match[0].length);
}
return { sentences, rest };
}
app.post("/api/voice-turn", async (req, res) => {
// Authenticate your own user and derive a stable session ID here.
const sessionId = req.get("x-session-id") ?? "demo";
const { messages } = req.body; // Prior turns plus the new user text.
res.setHeader("Content-Type", "audio/mpeg");
res.setHeader("Cache-Control", "no-store");
// One generation at a time per connection: wait for this session's previous turn.
const previous = queues.get(sessionId) ?? Promise.resolve();
let finish;
queues.set(sessionId, new Promise((resolve) => (finish = resolve)));
await previous;
try {
let connection;
const speak = async (sentence) => {
connection = await getConnection(sessionId, sentence);
// Store this ID with your turn ID before speaking, so the same turn
// can be retried with the same ID.
const requestId = randomUUID();
for await (const chunk of connection.speak(sentence, requestId)) {
res.write(chunk); // MP3 bytes
}
};
const stream = await addis.chat.completions.create({
system: "Answer naturally and concisely in Amharic.",
messages,
stream: true,
});
let buffer = "";
for await (const chunk of stream) {
buffer += chunk.choices[0].delta.content ?? "";
const { sentences, rest } = takeSentences(buffer);
buffer = rest;
for (const sentence of sentences) await speak(sentence);
}
if (buffer.trim()) await speak(buffer.trim()); // Speak any remaining text.
console.log(connection?.lastCompletion?.usage);
} catch (error) {
console.error(error);
sessions.get(sessionId)?.connection.close();
sessions.delete(sessionId);
} finally {
res.end();
finish(); // Let the next turn for this session start.
}
});A connection lasts 10 minutes and accepts at most 5,000 characters and 100 request IDs. getConnection closes it and opens a new one before any of those limits is reached. Streaming chat cannot be combined with tools, attachments, or audio input.
Browser: play the reply
Send the final transcript to your server and play the response body. Where the browser supports MP3 in MediaSource, append chunks as they arrive. Otherwise, wait for the whole body and play it as one clip.
// Example — validate against your account before production
async function speakReply(messages) {
const response = await fetch("/api/voice-turn", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ messages }),
});
if (!response.ok || !response.body) throw new Error("Voice turn failed");
const audio = new Audio();
if (window.MediaSource && MediaSource.isTypeSupported("audio/mpeg")) {
const mediaSource = new MediaSource();
audio.src = URL.createObjectURL(mediaSource);
await new Promise((resolve) =>
mediaSource.addEventListener("sourceopen", resolve, { once: true }),
);
const buffer = mediaSource.addSourceBuffer("audio/mpeg");
const reader = response.body.getReader();
const append = (bytes) =>
new Promise((resolve) => {
buffer.addEventListener("updateend", resolve, { once: true });
buffer.appendBuffer(bytes); // Append in the order received.
});
audio.play();
for (;;) {
const { done, value } = await reader.read();
if (done) break;
await append(value);
}
mediaSource.endOfStream();
} else {
// Fallback: no streaming playback, but the same audio.
const blob = await response.blob();
audio.src = URL.createObjectURL(blob);
await audio.play();
}
}
// Wire it together.
// const stop = await listen({
// onPartial: (text) => (transcriptEl.textContent = text),
// onFinal: (text) => speakReply([...history, { role: "user", content: text }]),
// });Store the final user text and the assistant's text in your session and send them back as messages on the next turn. Trim or summarize older turns instead of growing the history indefinitely.
Latency, limits, and billing
- Overlap the stages. Partial transcripts appear while the person speaks, chat starts on the final transcript, and the first sentence is spoken before the reply is complete. Shorter sentences reach the speaker sooner.
- Send complete sentences. A single word or a few letters can produce poor speech.
- Partials are provisional. Treat
transcript.partialtext as a preview, and usetranscript.completedas the final text and usage. - Scribe billing. Transcription is billed on final transcribed characters at 3.5 ETB per 1,000 characters. See Speech-to-Text for limits and interrupted requests.
- Voice billing. Speech is billed at 5 ETB per generated minute. Read
connection.lastCompletion?.usagefor the confirmed usage of the latest turn. - One turn at a time. Each text-to-speech connection runs one generation at a time. Open a second connection for parallel speech.
- Session limits. A speech session lasts 10 minutes, with 1–5,000 cumulative characters set at creation and 100 request IDs.
- Capacity errors. A
429withLANGUAGE_CAPACITY_REACHEDincludes aRetry-Afterheader. Retry after that delay with the same request ID. - Disconnects. Work the service has already accepted is still finished and billed if your connection drops or the browser stops reading. See Streaming limits.
- Pricing. See Pricing for current rates.
Production checklist
- Authenticate your own users before creating Scribe tickets, and enforce a per-user spending policy.
- Serve tickets with
Cache-Control: no-store, and keep tickets and API keys out of logs and URLs. - Request microphone permission before the ticket, and handle denial in your interface.
- Save a request ID for every speech turn so you can recover or replay it.
- Reconnect before the 10-minute, 5,000-character, or 100-turn session limits.
- Close each text-to-speech connection when the user session ends.
- Show a fallback (typed input, or a clip played after completion) when the browser lacks 16 kHz capture or MP3
MediaSource. - For a managed, interruption-capable conversation, use the Realtime API instead.