Addis AI

Speech-to-Text

Transcribe Amharic and Afaan Oromo audio into text.

Our Speech-to-Text (STT) technology is powered by specialized models trained to accurately recognize African languages. Unlike generic transcription services, our models are optimized for the specific phonemes, accents, and dialects of Amharic and Afaan Oromo.

Try Addis Scribe live

Transcribe Amharic live or upload audio. Choose Standard or Turbo · 3× speed and enter your API key. Both modes use the same transcription rate. Uploaded files can also return word timestamps and captions that play in sync with your audio.

Addis Scribe · Live API

Try Amharic transcription

Record live or upload audio. Standard transcription rates apply.

Your account

Your API key is not saved. Get an API key.

Speed

3× speed

Marks who said what as Speaker 1, 2… (not names). Uses Turbo.

Up to 3 minutes.

Marks who said what as Speaker 1, 2… (not names). Uses Turbo.

Timestamps return the full result at once, so partial text is off. Turn timestamps off to see partial text. Live recording returns text only unless Label speakers is on above.

WAV, MP3, M4A, WebM, OGG or FLAC · 25 MB · 3 minutes

Transcript

Record Amharic speech or upload audio.

የእርስዎ ጽሑፍ እዚህ ይታያል።

Disconnected? Recover your transcript with no duplicate charge.

Scribe HTTP and streaming API

The Scribe endpoints are separate from the existing file-transcription API documented below. Upload one audio file, up to 25 MB and 180 seconds, with a unique request_id:

curl 'https://api.addisassistant.com/api/v1/scribe/transcribe?backend=standard&request_id=my-transcript-1' \
  -H "x-api-key: $ADDIS_API_KEY" \
  -F 'audio=@audio.wav'

Use backend=standard for Standard or backend=turbo for Turbo. Add stream=true for newline-delimited transcript.partial events after uploading; transcript.completed includes the final text and confirmed wallet usage. An ordinary HTTP request returns data.text, data.seconds, data.compute_ms, and data.usage as JSON.

For live microphone input, see Live transcription.

transcript.partial text is provisional. transcript.completed confirms the final text and settled charge. Once audio is accepted, disconnecting still finishes transcription and billing. Recover an interrupted response with GET /api/v1/scribe/requests/YOUR_REQUEST_ID and the same account's API key. Replaying a completed request returns its original result without another charge; changed audio or settings require a new ID. Completed results are recoverable for 24 hours; pending settlements remain recoverable. One transcription per wallet is admitted at a time; a busy backend returns an error without switching backends.

GET /api/v1/scribe/usage returns the current STT rate and account balance. Billing uses final transcribed characters, at the existing 3.5 ETB per 1,000 characters rate as of October 4, 2026. Standard and Turbo pricing is the same. Model compute time excludes upload, network, and wallet settlement time.

Scribe SDK quick start

Install the SDK with the commands below. Scribe currently recognizes Amharic; the existing speech.transcribe method below continues to use the separate file-transcription service.

Keep ADDIS_API_KEY on your application server. For browser integrations, create a Scribe session on your server and return the scoped one-use ticket to the browser.

npm install https://github.com/Addis-AI-Org/addisai-js/releases/download/v0.4.0/addisai-0.4.0.tgz
import AddisAI, { fileFromPath, ulid } from "addisai";

const addis = new AddisAI(); // Reads ADDIS_API_KEY
const requestId = ulid();   // Save before sending audio
const result = await addis.scribe.transcribe({
  audio: await fileFromPath("audio.wav"),
  backend: "standard",
  requestId,
});

console.log(result.text);
console.log(result.usage.creditsUsed, result.usage.currency);
console.log(result.seconds, result.computeMs);
// After an interrupted response, recover this ID before retrying:
// const result = await addis.scribe.recover(requestId);
pip install "addisai[realtime]==0.4.0"
from addisai import AddisAI, ulid

request_id = ulid()  # Save before sending audio
with AddisAI() as addis, open("audio.wav", "rb") as audio:
    result = addis.scribe.transcribe(
        audio=audio, backend="standard", request_id=request_id,
    )
    print(result["text"])
    print(result["usage"]["credits_used"], result["usage"]["currency"])
    print(result["seconds"], result["compute_ms"])
    # After an interrupted response, recover this ID before retrying:
    # result = addis.scribe.recover(request_id)

Stream transcript updates after a file upload

scribe.stream uploads the audio first, then yields provisional partials followed by a settled completion. This is useful when you already have a recording. Use a WebSocket session for audio captured as the speaker talks.

const stream = await addis.scribe.stream({
  audio: await fileFromPath("audio.wav"),
  backend: "standard",
  requestId: ulid(),
});
try {
  for await (const event of stream) {
    if (event.type === "transcript.partial") console.log(event.text);
    if (event.type === "transcript.completed") {
      console.log(event.data.text, event.data.usage.creditsUsed);
    }
  }
} finally {
  stream.close();
}
with open("audio.wav", "rb") as audio:
    with addis.scribe.stream(audio=audio, request_id=ulid()) as stream:
        for event in stream:
            if event["type"] == "transcript.partial":
                print(event["text"])
            elif event["type"] == "transcript.completed":
                print(event["data"]["text"])
                print(event["data"]["usage"]["credits_used"])

Connect live audio with the SDK

Stream microphone or PCM audio with scribe.connect. See Connect live audio with the SDK on the Live transcription page.

Scribe methods and limits

OperationNode.jsPython
Read model capabilities and limitsscribe.capabilities()scribe.capabilities()
Check wallet and current character ratescribe.usage()scribe.usage()
Transcribe a filescribe.transcribe(...)scribe.transcribe(...)
File upload with partial updatesscribe.stream(...)scribe.stream(...)
Issue a one-use browser ticketscribe.createSession(...)scribe.create_session(...)
Open a live PCM connectionscribe.connect(...)scribe.connect(...)
Recover an existing requestscribe.recover(requestId)scribe.recover(request_id)
Word timestamps and caption cues (0.5.0)scribe.transcribe({ ..., timestamps: "word" })scribe.transcribe(..., timestamps="word")
Build SRT or VTT captions locally (0.5.0)toSrt(result), toVtt(result)addisai.to_srt(result), addisai.to_vtt(result)
Speaker labels on words and caption cues (0.6.0, Turbo)scribe.transcribe({ ..., backend: "turbo", speakers: true })scribe.transcribe(..., backend="turbo", speakers=True)
SRT or VTT captions with speaker labels (0.6.0)toSrt(result), toVtt(result)addisai.to_srt(result), addisai.to_vtt(result)

HTTP defaults to Standard and 1120ms chunks; live sessions default to Standard and 320ms chunks. Either transport accepts 320ms or 1120ms. File uploads accept at most 25 MiB / 180 seconds. Socket frames accept at most two seconds each and 180 seconds total. Socket tickets are one-use and expire after 60 seconds. An unavailable mode returns an error.

Billing and interrupted requests

Only the final transcript is billable. Partial events are provisional cumulative text and do not add separate charges. Character counts use UTF-16 code units, matching the existing STT wallet ledger. Read the current rate through scribe.usage(); confirmed completion includes the charged characters, rate, credits used, remaining balance, currency, and ledger record ID. Audio duration and model compute time are metrics, not the billing unit.

Save a request ID before sending audio. The SDK disables automatic retries for paid uploads and ticket creation. If the response is interrupted or billing is pending, call scribe.recover with that ID and the same account's key. A completed replay returns its original settled result without another deduction. Use a new ID when changing audio, backend, chunk, or HTTP streaming settings. Accepted audio can still finish and be billed after a socket disconnect or when you stop reading an HTTP stream.

Completed results remain recoverable for 24 hours; pending settlement remains recoverable. Recovery can report IN_PROGRESS, BILLING_PENDING, REQUEST_INTERRUPTED, or NOT_FOUND. Do not display a confirmed charge unless usage.settled is true. A busy account/backend returns HTTP 429, and changed inputs with a reused ID return HTTP 409.

Timestamps and captions New

Scribe can return the time of every word and ready-made caption cues for a completed file upload. Use them to subtitle a video, highlight words while audio plays, or jump to the moment a phrase was said. Ask for JSON with timestamps=word, or ask for a caption file directly with format=srt or format=vtt.

These query parameters are added to POST /api/v1/scribe/transcribe. Everything else about the request stays the same.

Prop

Type

With timestamps=word, the usual data object also contains:

{
  "data": {
    "text": "ሰላም ወዳጆቻችን እንዴት ከረማችሁ ዛሬ እንግዲህ",
    "words": [
      { "text": "ሰላም", "start": 18.9, "end": 19.52 },
      { "text": "ወዳጆቻችን", "start": 19.6, "end": 20.08 }
    ],
    "segments": [
      { "text": "ሰላም ወዳጆቻችን እንዴት ከረማችሁ ዛሬ እንግዲህ", "start": 18.8, "end": 21.8 }
    ]
  }
}
  • Times are seconds from the start of the file, with at most two decimals.
  • words are in spoken order. A word's end is never before its start.
  • segments are caption cues. A cue breaks at a pause of half a second or more, at two lines of 42 characters, or at 7 seconds, and cues never overlap. Segment text is one line; caption files wrap it at 42 characters, two lines at most.

format=srt returns a SubRip file (application/x-subrip). format=vtt returns WebVTT (text/vtt) with a WEBVTT header and 00:00:18.800-style times.

1
00:00:18,800 --> 00:00:21,800
ሰላም ወዳጆቻችን እንዴት ከረማችሁ ዛሬ እንግዲህ እንግዳ አድርጌ
ያቀረኩላችሁ
npm install https://github.com/Addis-AI-Org/addisai-js/releases/download/v0.5.0/addisai-0.5.0.tgz
import { writeFile } from "node:fs/promises";
import AddisAI, { fileFromPath, toSrt, ulid } from "addisai";

const addis = new AddisAI();
const result = await addis.scribe.transcribe({
  audio: await fileFromPath("audio.wav"),
  timestamps: "word",
  requestId: ulid(),
});

for (const word of result.words ?? []) {
  console.log(word.start, word.end, word.text);
}
await writeFile("captions.srt", toSrt(result)); // or toVtt(result)
pip install "addisai[realtime]==0.5.0"
import addisai
from addisai import AddisAI, ulid

with AddisAI() as client, open("audio.wav", "rb") as audio:
    result = client.scribe.transcribe(
        audio=audio, timestamps="word", request_id=ulid(),
    )

for word in result["words"]:
    print(word["start"], word["end"], word["text"])
with open("captions.srt", "w", encoding="utf-8") as captions:
    captions.write(addisai.to_srt(result))  # or addisai.to_vtt(result)
# JSON with words and segments
curl 'https://api.addisassistant.com/api/v1/scribe/transcribe?timestamps=word&request_id=my-captions-1' \
  -H "x-api-key: $ADDIS_API_KEY" \
  -F 'audio=@audio.wav'

# An SRT caption file
curl 'https://api.addisassistant.com/api/v1/scribe/transcribe?format=srt&request_id=my-captions-2' \
  -H "x-api-key: $ADDIS_API_KEY" \
  -F 'audio=@audio.wav' \
  -o captions.srt

toSrt/toVtt and to_srt/to_vtt run locally on the result you already have; they make no API call. To get captions for an earlier request, call GET /api/v1/scribe/requests/YOUR_REQUEST_ID?format=srt (or vtt). This works only for requests made with timestamps.

Limits

  • Timestamps are available for completed file uploads only. They are not available with stream=true or on live WebSocket sessions; combining them with stream=true returns HTTP 422 INVALID_REQUEST.
  • Files are limited to 180 seconds and 25 MiB, as for other uploads.
  • There is no extra charge. Billing stays per transcribed character.
  • Scribe does not add punctuation, so captions break at pauses rather than at the ends of sentences.

Accuracy. Times come from the speech model and are close, not exact. On long Amharic recordings, captions usually appear about 0.15 seconds after the speaker starts and disappear about 0.4 seconds before they finish. That is fine for subtitles and following along, but check by ear if you need frame-exact edits.

Speaker labels New

Scribe can tell you who said what in a recording with more than one voice, such as an interview, a meeting, or a podcast. Every word and every caption cue gets a speaker number. Speakers are numbered Speaker 1, Speaker 2, … in the order they first speak; Scribe does not know or return anyone's name.

Add this query parameter to POST /api/v1/scribe/transcribe. Everything else about the request stays the same.

Prop

Type

With speakers=true, each word and segment has a speaker field, and data.speakers is the number of different speakers found:

{
  "data": {
    "text": "ሰላም ወዳጆቻችን እንዴት ናችሁ",
    "words": [
      { "text": "ሰላም", "start": 0.98, "end": 1.08, "speaker": 1 },
      { "text": "ወዳጆቻችን", "start": 1.12, "end": 1.6, "speaker": 1 },
      { "text": "እንዴት", "start": 1.7, "end": 2.2, "speaker": 2 },
      { "text": "ናችሁ", "start": 2.3, "end": 3.0, "speaker": 2 }
    ],
    "segments": [
      { "text": "ሰላም ወዳጆቻችን", "start": 0.6, "end": 1.6, "speaker": 1 },
      { "text": "እንዴት ናችሁ", "start": 1.7, "end": 3.0, "speaker": 2 }
    ],
    "speakers": 2
  }
}
  • speaker is 1, 2, and so on, numbered by first appearance. It is null when Scribe could not tell who said a word.
  • A caption cue always ends when the speaker changes, so each cue has one speaker. A word with speaker: null stays in the current cue.

Caption files show the speaker on every cue. format=srt puts Speaker N: before the cue text; the label counts toward the 42-character line length:

1
00:00:00,600 --> 00:00:01,600
Speaker 1: ሰላም ወዳጆቻችን

format=vtt uses a WebVTT voice tag, <v Speaker N>, at the start of the cue instead of a text label:

00:00:00.600 --> 00:00:01.600
<v Speaker 1>ሰላም ወዳጆቻችን
npm install https://github.com/Addis-AI-Org/addisai-js/releases/download/v0.6.0/addisai-0.6.0.tgz
import { writeFile } from "node:fs/promises";
import AddisAI, { fileFromPath, toSrt, ulid } from "addisai";

const addis = new AddisAI();
const result = await addis.scribe.transcribe({
  audio: await fileFromPath("interview.wav"),
  backend: "turbo", // Speaker labels need Turbo
  speakers: true,   // Also turns on word timestamps
  requestId: ulid(),
});

console.log(`${result.speakers} speakers`);
for (const segment of result.segments ?? []) {
  console.log(`Speaker ${segment.speaker ?? "?"}:`, segment.text);
}
await writeFile("captions.srt", toSrt(result)); // or toVtt(result)
pip install "addisai[realtime]==0.6.0"
import addisai
from addisai import AddisAI, ulid

with AddisAI() as client, open("interview.wav", "rb") as audio:
    result = client.scribe.transcribe(
        audio=audio,
        backend="turbo",  # Speaker labels need Turbo
        speakers=True,    # Also turns on word timestamps
        request_id=ulid(),
    )

print(result["speakers"], "speakers")
for segment in result["segments"]:
    print(f"Speaker {segment['speaker'] or '?'}:", segment["text"])
with open("captions.srt", "w", encoding="utf-8") as captions:
    captions.write(addisai.to_srt(result))  # or addisai.to_vtt(result)
# JSON with words, segments, and speaker numbers
curl 'https://api.addisassistant.com/api/v1/scribe/transcribe?backend=turbo&speakers=true&request_id=my-speakers-1' \
  -H "x-api-key: $ADDIS_API_KEY" \
  -F 'audio=@interview.wav'

# An SRT caption file with "Speaker N: " labels
curl 'https://api.addisassistant.com/api/v1/scribe/transcribe?backend=turbo&speakers=true&format=srt&request_id=my-speakers-2' \
  -H "x-api-key: $ADDIS_API_KEY" \
  -F 'audio=@interview.wav' \
  -o captions.srt

In SDK 0.6.0, toSrt/toVtt and to_srt/to_vtt add the same speaker labels whenever a segment has a speaker, so local caption files match the API's byte for byte. The SDKs check the request before sending it: speakers without backend: "turbo" (backend="turbo" in Python), or with scribe.stream, is rejected locally.

Limits

  • Speaker labels need backend=turbo. With Standard the API returns HTTP 422 INVALID_REQUEST; with stream=true it also returns HTTP 422. Live WebSocket sessions return text only.
  • scribe.capabilities() lists the backends that support speaker labels and the most speakers Scribe will separate in one file (currently 8).
  • There is no extra charge. Billing stays per transcribed character.
  • A request ID is tied to its settings. Replaying the same ID with a different speakers value returns HTTP 409; use a new ID.
  • Files are limited to 180 seconds and 25 MiB, as for other uploads.

Accuracy. On simulated Amharic conversations, 97.8% of words get the right speaker. It works best with 2 to 4 speakers who take turns. When people talk over each other, words near the overlap are more likely to get the wrong speaker or null. Check the labels before publishing anything where it matters who said what.

Usage Guide

Speech-to-text uploads use multipart/form-data. The SDK handles the transport format when you pass an audio file and language.

Transcribe a File

Upload an audio file to get the text transcription.

import AddisAI, { fileFromPath } from "addisai";

const addis = new AddisAI();
const result = await addis.speech.transcribe({
  audio: await fileFromPath("audio.wav"),
  language: "am",
});

console.log(result.text);
console.log(result.confidence);
from addisai import AddisAI

addis = AddisAI()
with open("audio.wav", "rb") as audio:
    result = addis.speech.transcribe(audio=audio, language="am")

print(result["text"])
print(result.get("confidence"))
curl --location 'https://api.addisassistant.com/api/v2/stt' \\
  --header 'x-api-key: $ADDIS_API_KEY' \\
  --form 'audio=@"audio.wav"' \\
  --form 'request_data="{ \"language_code\": \"am\" }"'

API Reference

Form Data Parameters

These fields are sent as multipart form data.

Prop

Type

Request Data Object

These parameters go inside the request_data JSON string.

Prop

Type

Response Schema

{
  "status": "success",
  "data": {
    "transcription": "ሰላም እንኳን ደህና መጣችሁ",
    "usage_metadata": {
      "totalBilledDuration": "15s",
      "requestId": "69b60667-0000-2a1e-b6d3-d4f547fe6724"
    }
  },
  "confidence": 0.982
}

Supported Formats

We support standard audio containers. For the fastest processing, we recommend WAV.

FormatContent Types (MIME)
WAVaudio/wav, audio/x-wav, audio/wave
MP3audio/mpeg, audio/mp3
M4Aaudio/mp4, audio/x-m4a
WebMaudio/webm

Best Practices

To ensure high accuracy (WER < 10%), follow these recording guidelines.

Audio Specs

Sample Rate: 16kHz or higher is recommended for clarity.

Channel: Mono is preferred. Stereo files are supported but mixed down before processing.

Environment

Noise: Background noise significantly degrades accuracy. Record in quiet environments.

Distance: Keep the speaker 10-30cm from the microphone for optimal volume levels.

Constraints

Max Duration60 Seconds
Max File Size10 MB

Limitations

Speakers: The model is optimized for single-speaker audio. Overlapping voices may result in skipped words.

Technical Terms: Rare technical jargon or code-switching (mixing English heavily) may have lower accuracy.

On this page