Addis AI

Streaming

Stream Addis Voices 2 speech over HTTP or a persistent WebSocket.

Play speech as it is generated over HTTP or a persistent WebSocket. Choose a voice and enter your API key to try it live.

Addis Voices 2 · Live API

Try streaming speech

Listen as speech is generated. Standard voice rates apply.

Your account

Your API key is not saved. Get an API key.

29 / 1,000 characters

Check your account, choose a voice, then stream a sentence.

First audio received
—
Audio received
0.0 KB

Stopping playback does not cancel a generation that has started; it is still saved and billed once. Recover an interrupted request with no duplicate charge.

Use HTTP for individual requests or WebSockets for successive turns on one connection. For microphone conversations with AI responses, see the Realtime API.

Choose a language and voice

LanguageCodeExample voiceVoice ID
AmharicamHamenam-hamen
Afaan OromoomBikilaom-bikila
TigrinyatiBerhaneti-berhane

Fetch the current catalog with voices.list({ language: "ti" }) in Node.js or voices.list(language="ti") in Python. Select an available voice whose language matches the request. The examples below use Berhane (ti-berhane, ti); use Hamen (am-hamen, am) or Bikila (om-bikila, om) in the same calls for the other documented languages.

curl --fail-with-body \
  "https://api.addisassistant.com/api/v1/voice/voices?language=ti" \
  -H "x-api-key: $ADDIS_API_KEY"

Send complete sentences. A single word or a few letters can produce poor speech; let the service divide an utterance into phrases.

Stream with the SDK

Install the SDK and set ADDIS_API_KEY on your server.

Persistent WebSocket

Connect once and reuse the connection for successive turns. Wait for each turn to finish before starting another, and save a unique request ID before sending it.

import { createWriteStream } from "node:fs";
import { randomUUID } from "node:crypto";
import { Readable } from "node:stream";
import { pipeline } from "node:stream/promises";
import AddisAI from "addisai";

const addis = new AddisAI();
const requestId = randomUUID(); // Persist this ID to recover the same turn.
const connection = await addis.realtime.connect({
  voiceId: "ti-berhane",
  language: "ti",
  maxTextCharacters: 1000,
});

try {
  await pipeline(
    Readable.from(connection.speak(
      "ሰላም፣ እንቋዕ ናብ ኣዲስ ኤኣይ ብደሓን መጻእኩም።",
      requestId,
    )),
    createWriteStream("speech.mp3"),
  );
  console.log(connection.lastCompletion?.usage);
} finally {
  connection.close();
}

Install the Python SDK with its optional WebSocket dependency:

pip install "addisai[realtime]==0.4.0"
from uuid import uuid4
from addisai import AddisAI

request_id = str(uuid4())  # Persist this ID to recover the same turn.

with AddisAI() as addis:
    with addis.realtime.connect(
        voice_id="ti-berhane", language="ti", max_text_characters=1000,
    ) as connection:
        with open("speech.mp3", "wb") as output:
            for chunk in connection.speak(
                "ሰላም፣ እንቋዕ ናብ ኣዲስ ኤኣይ ብደሓን መጻእኩም።",
                request_id,
            ):
                output.write(chunk)
        print(connection.last_completion["usage"])

The file examples consume the stream completely. For playback during generation, feed the chunks to a streaming MP3 player in their received order. speak() uses an MP3 session.

HTTP streaming

Use HTTP when each utterance is an independent request. The SDK handles the audio framing and final billing metadata.

import AddisAI from "addisai";

const addis = new AddisAI();
const audio = await addis.voice.stream({
  voiceId: "ti-berhane",
  language: "ti",
  text: "ሰላም፣ እንቋዕ ናብ ኣዲስ ኤኣይ ብደሓን መጻእኩም።",
  clientRequestId: "my-persisted-turn-id",
});

await audio.toFile("speech.mp3");
console.log(audio.metadata?.usage);
from addisai import AddisAI

with AddisAI() as addis:
    audio = addis.voice.stream(
        voice_id="ti-berhane", language="ti",
        text="ሰላም፣ እንቋዕ ናብ ኣዲስ ኤኣይ ብደሓን መጻእኩም።",
        client_request_id="my-persisted-turn-id",
    )

    audio.to_file("speech.mp3")
    print(audio.metadata["usage"])

To play audio immediately, iterate audio rather than waiting for the file helper to finish. HTTP streaming accepts mp3_44100; use voice.generate for other formats. Check the format returned in audio events rather than inferring its sample rate from the output format name.

WebSocket protocol and HTTP framing

Connect without an SDK

Create a session on your server

Authenticate your application's user before creating a session, and enforce your own per-user spending policy. Keep the developer API key on your server. Return the session's data object to the authorized client with Cache-Control: no-store; keep tickets out of logs.

curl --fail-with-body \
  https://api.addisassistant.com/api/v1/realtime/sessions \
  -H "x-api-key: $ADDIS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "voice_id": "ti-berhane",
    "language": "ti",
    "audio_format": "mp3",
    "max_text_characters": 1000
  }'

The 201 response contains data.id, data.token, data.websocket_url, data.expires_at, data.session_ttl_seconds, and the requested settings. The ticket expires after 60 seconds and can be used once.

max_text_characters limits cumulative submitted text across the session. It is a text limit, not a guaranteed maximum charge. Billing uses the audio actually generated.

Authenticate the socket

A browser or mobile client uses the scoped ticket. Connect to the returned websocket_url and send session.authenticate as the first message. Wait for session.created before sending text. Do not put a developer API key or ticket in the WebSocket URL.

This browser example assumes /my-server/voice-session is your authenticated backend route and returns the session's data object:

async function streamSpeech(text, requestId, onAudio) {
  const response = await fetch("/my-server/voice-session", { method: "POST" });
  if (!response.ok) throw new Error("Could not create a voice session");
  const session = await response.json();
  const socket = new WebSocket(session.websocket_url);

  return new Promise((resolve, reject) => {
    let completed = false;
    socket.onopen = () => {
      socket.send(JSON.stringify({
        type: "session.authenticate", token: session.token,
      }));
    };
    socket.onmessage = ({ data }) => {
      try {
        const event = JSON.parse(data);
        if (event.type === "session.created") {
          socket.send(JSON.stringify({
            type: "speech.create", text, request_id: requestId,
          }));
        } else if (event.type === "audio.delta") {
          const bytes = Uint8Array.from(atob(event.audio), c => c.charCodeAt(0));
          onAudio(bytes, event.format); // Enqueue chunks in the received order.
        } else if (event.type === "speech.completed") {
          completed = true;
          resolve(event.data); // Saved clip, playback URL, and usage.
          socket.close();
        } else if (event.type === "error") {
          throw new Error(`${event.error.code}: ${event.error.message}`);
        }
      } catch (error) {
        reject(error);
        socket.close();
      }
    };
    socket.onerror = () => reject(new Error("Voice connection failed"));
    socket.onclose = () => {
      if (!completed) reject(new Error("Stream ended before billing confirmation"));
    };
  });
}

Pass a fresh request ID for a new utterance, and keep it with the exact text and settings for recovery. onAudio should enqueue bytes for a streaming player. Avoid decoding each MP3 phrase into a separate independently started player, which can introduce gaps.

Endpoints and events

EndpointPurpose
GET /api/v1/realtimePublic capabilities
POST /api/v1/realtime/sessionsCreate a ticket using x-api-key
wss://api.addisassistant.com/api/v1/realtime/voiceText-to-speech event socket
POST /api/v1/voice/generations/streamHTTP voice stream
GET /api/v1/voice/usageVoice usage, pricing, and wallet balance

HTTP endpoints use https://api.addisassistant.com.

The /api/v1/realtime paths and the SDK's addis.realtime.connect serve text-to-speech streaming. They are separate from the Realtime API for microphone conversations.

EventMeaning
session.createdAuthenticated socket, session ID, and voice settings
speech.startedTurn accepted; includes request_id
audio.deltarequest_id, sequence, format, and base64 audio
speech.completedSaved clip and usage in data
speech.cancelledClip and usage after delivery was muted
errorerror.code, error.message, and optional request ID/status
text.appendedNumber of characters buffered
pongResponse to ping

For generated text, send {"type":"text.append","text":"..."} and then {"type":"text.commit","request_id":"turn-1"} when the utterance is ready. Appending buffers text; committing starts one billed generation. speech.create submits a complete utterance directly.

Choose audio_format: "wav_mp3" when creating a session to receive early WAV pieces followed by MP3 phrases. Each audio.delta.format tells the player which decoder to use. The SDK's raw event iterator supports this mode; its speak() helper expects mp3.

For direct HTTP streaming, send text, language, voice_id, output_format: "mp3_44100", and client_request_id as JSON. The audio protocol is a four-byte big-endian length followed by that many audio bytes. A zero-length frame precedes the final JSON completion. Inspect X-Addis-Audio-Protocol and handle ordinary JSON errors or an idempotent JSON replay before parsing frames. Saving the raw framed HTTP body as an MP3 file will not produce a valid MP3.

Authentication

Use x-api-key: YOUR_API_KEY for HTTP requests. Browser and mobile sockets use the one-use session ticket. Keep API keys on your server and credentials out of socket URLs.

Billing and recovery

Voice streaming uses the same 5 ETB per generated minute pricing and wallet as Addis Voices 2 clips. Read voice.usage() for current pricing and balance. Billing depends on measured audio duration; a text-based estimate can differ from the final charge.

Consume the stream through speech.completed, or the HTTP completion record, to receive the saved clip and confirmed usage. Audio delivery alone does not confirm settlement. Store the returned clip ID and usage with your application's turn ID.

  • Keep request_id / client_request_id stable when retrying the same text, voice, language, and output format. A changed input requires a new ID.
  • An HTTP replay returns the previously saved clip; the SDK downloads its audio without another synthesis or charge.
  • A WebSocket replay returns speech.completed with idempotent_replay: true and the original clip. It does not stream the audio again; fetch the clip to play it.
  • speech.cancel mutes subsequent audio delivery and clears buffered text. A generation already started continues, saves its clip, and is billed. Disconnecting has the same billing behavior.
  • If a stream ends before completion, keep the original ID and inputs for recovery rather than issuing a new billable generation.

Streaming limits

LimitCurrent value
Ticket validity60 seconds; one connection
Session duration10 minutes
Cumulative text per session1–5,000 characters, set at creation
Audio per turn60 seconds by default; set max_audio_seconds (1–600) when creating the session
Turn concurrencyOne generation at a time per connection; open more sessions for parallel speech
Turns per session100 request IDs
Messages per connection20 per second
Idle receive timeout60 seconds; send ping while idle

There is no fixed cap on sessions or requests per API key. Your wallet balance bounds spending: each session holds credit for one turn and releases what it does not use when it ends. When the voice servers for a language are fully busy, a request returns 429 with LANGUAGE_CAPACITY_REACHED and a Retry-After header (retry_after_seconds on WebSockets). Retry after that delay with the same request ID.

Request IDs accept 1–128 ASCII letters, digits, _, or -. The voice and language are fixed for a session. SDK helpers generate an ID when one is omitted; save your own ID when your application needs to recover a turn after a restart.

On this page