Live transcription
Transcribe microphone audio with Addis Scribe over a WebSocket.
Create a session and connect
For live microphone input, create a one-use ticket with POST /api/v1/scribe/sessions, using your API key in the x-api-key header and this JSON body:
{"backend":"standard","chunk":"320ms","request_id":"my-live-transcript-1"}Connect to the returned data.websocket_url. Send {"type":"session.authenticate","token":"YOUR_ONE_USE_TICKET"} as the first frame. After session.created, send binary 16 kHz mono PCM16 little-endian frames, ideally 100 ms each. Finish with {"type":"audio.finish"}. Tickets expire after 60 seconds and bind the backend, chunk size, and request ID. Keep API keys and tickets out of WebSocket URLs.
Connect live audio with the SDK
Send raw mono PCM16 little-endian at 16 kHz. WAV/MP3 container bytes belong in the file-upload API; strip or decode them before sending PCM. A 100 ms frame is 3,200 bytes. Read events while sending audio, then call finish() to obtain the final transcript and billing result.
import { createReadStream } from "node:fs";
import { setTimeout as delay } from "node:timers/promises";
const live = await addis.scribe.connect({
backend: "standard", chunk: "320ms", requestId: ulid(),
});
try {
await Promise.all([
(async () => {
for await (const event of live) {
if (event.type === "transcript.partial") console.log(event.text);
if (event.type === "transcript.completed") {
console.log(event.data.text, event.data.usage.creditsUsed);
}
}
})(),
(async () => {
for await (const frame of createReadStream("speech.pcm", { highWaterMark: 3200 })) {
live.sendAudio(frame);
await delay(100); // Pace this PCM file like microphone input
}
live.finish();
})(),
]);
} finally {
live.close();
}# pip install "addisai[realtime]"
from concurrent.futures import ThreadPoolExecutor
from time import sleep
def send_pcm(live):
with open("speech.pcm", "rb") as audio:
while frame := audio.read(3200):
live.send_audio(frame)
sleep(0.1) # Pace this raw PCM file like microphone input
live.finish()
with addis.scribe.connect(request_id=ulid(), backend="standard") as live:
with ThreadPoolExecutor(max_workers=1) as executor:
sender = executor.submit(send_pcm, live)
for event in live:
if event["type"] == "transcript.partial":
print(event["text"])
elif event["type"] == "transcript.completed":
print(event["data"]["text"])
print(event["data"]["usage"]["credits_used"])
sender.result()For browser audio, use addis.scribe.createSession({ requestId }) on the server and connectScribe(session) in the browser. Python uses create_session and connect_scribe. The helper authenticates with the ticket in the first frame and waits for admission before audio is sent. A completed replay may immediately return transcript.completed without session.created; it does not accept new audio.
Live speaker labels (preview)
Add "speakers": true to the session body to label who is speaking as you record. Speaker labels need the Turbo backend; with "backend":"standard" the request fails with 422 INVALID_REQUEST.
{"backend":"turbo","chunk":"320ms","request_id":"my-live-transcript-2","speakers":true}session.created then includes "speakers": true. After each pause in speech (0.5 s of quiet after at least 1 s of speech, or every 5 s at most), the session sends a transcript.segment event with the words and caption segments committed since the last one:
{"type":"transcript.segment","request_id":"my-live-transcript-2","words":[{"text":"ሰላም","start":0.32,"end":0.81,"speaker":1},{"text":"እንዴት","start":1.62,"end":2.05,"speaker":2},{"text":"ነህ","start":2.05,"end":2.4,"speaker":2}],"segments":[{"text":"ሰላም","start":0.32,"end":0.81,"speaker":1},{"text":"እንዴት ነህ","start":1.62,"end":2.4,"speaker":2}],"speakers":2}Times are seconds from the start of the session. speaker is numbered from 1 in the order people first talk, not named, and is null when a word cannot be attributed. Labels already sent never change, and later segments keep the same numbering. speakers counts the distinct speakers so far. transcript.partial events still arrive with the unlabelled running text of the current stretch.
transcript.completed includes data.words, data.segments (with speaker) and data.speakers, like an upload with speakers=true, so you can build SRT or VTT captions from it. Its data.text is the joined words.
const lines = [];
socket.onmessage = ({ data }) => {
const event = JSON.parse(data);
if (event.type === "transcript.segment") {
for (const s of event.segments) lines.push(`${s.speaker ? `Speaker ${s.speaker}` : "Unattributed"}: ${s.text}`);
render(lines, ""); // committed, labelled lines
}
if (event.type === "transcript.partial") render(lines, event.text); // unlabelled, may change
if (event.type === "transcript.completed") console.log(event.data.segments, event.data.speakers);
};Live labels are a preview. Each line is usually labelled within about a second of the pause (median 0.3 s for two people, 0.9 s for groups). On two-person conversations they closely match labels from transcribing the whole file. With three or more people, voices are confused more often, so for group recordings upload the file with speakers=true instead. Sessions last up to 180 seconds and label up to 8 speakers.
Limits and billing
Live transcription uses the same limits and billing as the rest of Scribe. See Scribe methods and limits and Billing and interrupted requests.