Streaming
Stream Addis Voices 2 speech over HTTP or a persistent WebSocket.
Play speech as it is generated over HTTP or a persistent WebSocket. Choose a voice and enter your API key to try it live.
Addis Voices 2 · Live API
Try streaming speech
Listen as speech is generated. Standard voice rates apply.
29 / 1,000 characters
Check your account, choose a voice, then stream a sentence.
- First audio received
- —
- Audio received
- 0.0 KB
Stopping playback does not cancel a generation that has started; it is still saved and billed once. Recover an interrupted request with no duplicate charge.
Use HTTP for individual requests or WebSockets for successive turns on one connection. For microphone conversations with AI responses, see the Realtime API.
Choose a language and voice
| Language | Code | Example voice | Voice ID |
|---|---|---|---|
| Amharic | am | Hamen | am-hamen |
| Afaan Oromo | om | Bikila | om-bikila |
| Tigrinya | ti | Berhane | ti-berhane |
Fetch the current catalog with voices.list({ language: "ti" }) in Node.js or voices.list(language="ti") in Python. Select an available voice whose language matches the request. The examples below use Berhane (ti-berhane, ti); use Hamen (am-hamen, am) or Bikila (om-bikila, om) in the same calls for the other documented languages.
curl --fail-with-body \
"https://api.addisassistant.com/api/v1/voice/voices?language=ti" \
-H "x-api-key: $ADDIS_API_KEY"Send complete sentences. A single word or a few letters can produce poor speech; let the service divide an utterance into phrases.
Stream with the SDK
Install the SDK and set ADDIS_API_KEY on your server.
Persistent WebSocket
Connect once and reuse the connection for successive turns. Wait for each turn to finish before starting another, and save a unique request ID before sending it.
import { createWriteStream } from "node:fs";
import { randomUUID } from "node:crypto";
import { Readable } from "node:stream";
import { pipeline } from "node:stream/promises";
import AddisAI from "addisai";
const addis = new AddisAI();
const requestId = randomUUID(); // Persist this ID to recover the same turn.
const connection = await addis.realtime.connect({
voiceId: "ti-berhane",
language: "ti",
maxTextCharacters: 1000,
});
try {
await pipeline(
Readable.from(connection.speak(
"ሰላም፣ እንቋዕ ናብ ኣዲስ ኤኣይ ብደሓን መጻእኩም።",
requestId,
)),
createWriteStream("speech.mp3"),
);
console.log(connection.lastCompletion?.usage);
} finally {
connection.close();
}Install the Python SDK with its optional WebSocket dependency:
pip install "addisai[realtime]==0.4.0"from uuid import uuid4
from addisai import AddisAI
request_id = str(uuid4()) # Persist this ID to recover the same turn.
with AddisAI() as addis:
with addis.realtime.connect(
voice_id="ti-berhane", language="ti", max_text_characters=1000,
) as connection:
with open("speech.mp3", "wb") as output:
for chunk in connection.speak(
"ሰላም፣ እንቋዕ ናብ ኣዲስ ኤኣይ ብደሓን መጻእኩም።",
request_id,
):
output.write(chunk)
print(connection.last_completion["usage"])The file examples consume the stream completely. For playback during generation, feed the chunks to a streaming MP3 player in their received order. speak() uses an MP3 session.
HTTP streaming
Use HTTP when each utterance is an independent request. The SDK handles the audio framing and final billing metadata.
import AddisAI from "addisai";
const addis = new AddisAI();
const audio = await addis.voice.stream({
voiceId: "ti-berhane",
language: "ti",
text: "ሰላም፣ እንቋዕ ናብ ኣዲስ ኤኣይ ብደሓን መጻእኩም።",
clientRequestId: "my-persisted-turn-id",
});
await audio.toFile("speech.mp3");
console.log(audio.metadata?.usage);from addisai import AddisAI
with AddisAI() as addis:
audio = addis.voice.stream(
voice_id="ti-berhane", language="ti",
text="ሰላም፣ እንቋዕ ናብ ኣዲስ ኤኣይ ብደሓን መጻእኩም።",
client_request_id="my-persisted-turn-id",
)
audio.to_file("speech.mp3")
print(audio.metadata["usage"])To play audio immediately, iterate audio rather than waiting for the file helper to finish. HTTP streaming accepts mp3_44100; use voice.generate for other formats. Check the format returned in audio events rather than inferring its sample rate from the output format name.
WebSocket protocol and HTTP framing
Connect without an SDK
Create a session on your server
Authenticate your application's user before creating a session, and enforce your own per-user spending policy. Keep the developer API key on your server. Return the session's data object to the authorized client with Cache-Control: no-store; keep tickets out of logs.
curl --fail-with-body \
https://api.addisassistant.com/api/v1/realtime/sessions \
-H "x-api-key: $ADDIS_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"voice_id": "ti-berhane",
"language": "ti",
"audio_format": "mp3",
"max_text_characters": 1000
}'The 201 response contains data.id, data.token, data.websocket_url, data.expires_at, data.session_ttl_seconds, and the requested settings. The ticket expires after 60 seconds and can be used once.
max_text_characters limits cumulative submitted text across the session. It is a text limit, not a guaranteed maximum charge. Billing uses the audio actually generated.
Authenticate the socket
A browser or mobile client uses the scoped ticket. Connect to the returned websocket_url and send session.authenticate as the first message. Wait for session.created before sending text. Do not put a developer API key or ticket in the WebSocket URL.
This browser example assumes /my-server/voice-session is your authenticated backend route and returns the session's data object:
async function streamSpeech(text, requestId, onAudio) {
const response = await fetch("/my-server/voice-session", { method: "POST" });
if (!response.ok) throw new Error("Could not create a voice session");
const session = await response.json();
const socket = new WebSocket(session.websocket_url);
return new Promise((resolve, reject) => {
let completed = false;
socket.onopen = () => {
socket.send(JSON.stringify({
type: "session.authenticate", token: session.token,
}));
};
socket.onmessage = ({ data }) => {
try {
const event = JSON.parse(data);
if (event.type === "session.created") {
socket.send(JSON.stringify({
type: "speech.create", text, request_id: requestId,
}));
} else if (event.type === "audio.delta") {
const bytes = Uint8Array.from(atob(event.audio), c => c.charCodeAt(0));
onAudio(bytes, event.format); // Enqueue chunks in the received order.
} else if (event.type === "speech.completed") {
completed = true;
resolve(event.data); // Saved clip, playback URL, and usage.
socket.close();
} else if (event.type === "error") {
throw new Error(`${event.error.code}: ${event.error.message}`);
}
} catch (error) {
reject(error);
socket.close();
}
};
socket.onerror = () => reject(new Error("Voice connection failed"));
socket.onclose = () => {
if (!completed) reject(new Error("Stream ended before billing confirmation"));
};
});
}Pass a fresh request ID for a new utterance, and keep it with the exact text and settings for recovery. onAudio should enqueue bytes for a streaming player. Avoid decoding each MP3 phrase into a separate independently started player, which can introduce gaps.
Endpoints and events
| Endpoint | Purpose |
|---|---|
GET /api/v1/realtime | Public capabilities |
POST /api/v1/realtime/sessions | Create a ticket using x-api-key |
wss://api.addisassistant.com/api/v1/realtime/voice | Text-to-speech event socket |
POST /api/v1/voice/generations/stream | HTTP voice stream |
GET /api/v1/voice/usage | Voice usage, pricing, and wallet balance |
HTTP endpoints use https://api.addisassistant.com.
The /api/v1/realtime paths and the SDK's addis.realtime.connect serve text-to-speech streaming. They are separate from the Realtime API for microphone conversations.
| Event | Meaning |
|---|---|
session.created | Authenticated socket, session ID, and voice settings |
speech.started | Turn accepted; includes request_id |
audio.delta | request_id, sequence, format, and base64 audio |
speech.completed | Saved clip and usage in data |
speech.cancelled | Clip and usage after delivery was muted |
error | error.code, error.message, and optional request ID/status |
text.appended | Number of characters buffered |
pong | Response to ping |
For generated text, send {"type":"text.append","text":"..."} and then {"type":"text.commit","request_id":"turn-1"} when the utterance is ready. Appending buffers text; committing starts one billed generation. speech.create submits a complete utterance directly.
Choose audio_format: "wav_mp3" when creating a session to receive early WAV pieces followed by MP3 phrases. Each audio.delta.format tells the player which decoder to use. The SDK's raw event iterator supports this mode; its speak() helper expects mp3.
For direct HTTP streaming, send text, language, voice_id, output_format: "mp3_44100", and client_request_id as JSON. The audio protocol is a four-byte big-endian length followed by that many audio bytes. A zero-length frame precedes the final JSON completion. Inspect X-Addis-Audio-Protocol and handle ordinary JSON errors or an idempotent JSON replay before parsing frames. Saving the raw framed HTTP body as an MP3 file will not produce a valid MP3.
Authentication
Use x-api-key: YOUR_API_KEY for HTTP requests. Browser and mobile sockets use the one-use session ticket. Keep API keys on your server and credentials out of socket URLs.
Billing and recovery
Voice streaming uses the same 5 ETB per generated minute pricing and wallet as Addis Voices 2 clips. Read voice.usage() for current pricing and balance. Billing depends on measured audio duration; a text-based estimate can differ from the final charge.
Consume the stream through speech.completed, or the HTTP completion record, to receive the saved clip and confirmed usage. Audio delivery alone does not confirm settlement. Store the returned clip ID and usage with your application's turn ID.
- Keep
request_id/client_request_idstable when retrying the same text, voice, language, and output format. A changed input requires a new ID. - An HTTP replay returns the previously saved clip; the SDK downloads its audio without another synthesis or charge.
- A WebSocket replay returns
speech.completedwithidempotent_replay: trueand the original clip. It does not stream the audio again; fetch the clip to play it. speech.cancelmutes subsequent audio delivery and clears buffered text. A generation already started continues, saves its clip, and is billed. Disconnecting has the same billing behavior.- If a stream ends before completion, keep the original ID and inputs for recovery rather than issuing a new billable generation.
Streaming limits
| Limit | Current value |
|---|---|
| Ticket validity | 60 seconds; one connection |
| Session duration | 10 minutes |
| Cumulative text per session | 1–5,000 characters, set at creation |
| Audio per turn | 60 seconds by default; set max_audio_seconds (1–600) when creating the session |
| Turn concurrency | One generation at a time per connection; open more sessions for parallel speech |
| Turns per session | 100 request IDs |
| Messages per connection | 20 per second |
| Idle receive timeout | 60 seconds; send ping while idle |
There is no fixed cap on sessions or requests per API key. Your wallet balance bounds spending: each session holds credit for one turn and releases what it does not use when it ends. When the voice servers for a language are fully busy, a request returns 429 with LANGUAGE_CAPACITY_REACHED and a Retry-After header (retry_after_seconds on WebSockets). Retry after that delay with the same request ID.
Request IDs accept 1–128 ASCII letters, digits, _, or -. The voice and language are fixed for a session. SDK helpers generate an ID when one is omitted; save your own ID when your application needs to recover a turn after a restart.