Scriptivox API
Transcribe audio and video at $0.20/hour, billed per second. Speaker diarization, word-level timestamps, and automatic language detection across 119 languages — all included.
Powered by Whisper. Pay-as-you-go, no minimums, no subscriptions.
Developer quickstart
Transcribe audio and video files with word-level timestamps, speaker diarization, and language detection. Build transcription into your product in minutes.
import requestsimport timeAPI_KEY = "sk_live_YOUR_KEY"BASE = "https://api.scriptivox.com/v1"# 1. Transcribe from a URLresp = requests.post(f"{BASE}/transcribe",headers={"Authorization": API_KEY},json={"url": "https://example.com/audio.mp3"})job = resp.json()# 2. Get resultwhile True:result = requests.get(f"{BASE}/transcribe/{job['id']}",headers={"Authorization": API_KEY}).json()if result["status"] == "completed":print(result["result"]["full_transcript"])breaktime.sleep(5)
What the API does
Send a recorded audio or video file — by public URL, or by uploading it to a presigned URL we hand you — and get back a structured transcript. Every response carries the full text, an array of utterances with start and end times, and, unless you opt out, a timestamp for every individual word. Speaker labels are added when you ask for diarization. The same transcript can be exported as SRT or WebVTT subtitles, or as plain text, by adding a query parameter to the fetch.
Jobs are asynchronous. POST /v1/transcribe accepts the work and returns immediately with status: "created"; the file is downloaded and validated afterwards. That ordering matters more than anything else on this page: a bad URL, an unreadable file or an unsupported codec surfaces on the poll, as status: "failed", not as an error on the submit call. Clients that only check the response to POST will believe every job succeeded.
This is not a streaming or real-time dictation API. Every job takes a complete file, between 1 second and 10 hours long and at most 5 GB.
Authentication
Every request needs an API key from the dashboard, passed as either Authorization: sk_live_… (a Bearer prefix is accepted) or X-Api-Key: sk_live_…. See Authentication for rotation, the key limit, and what each 401 means.
Endpoints
Base URL: https://api.scriptivox.com/v1
| Method | Path | What it does |
|---|---|---|
| POST | /v1/upload | Get a presigned URL for a file you host yourself. |
| POST | /v1/transcribe | Start a job from a public URL or a completed upload. |
| GET | /v1/transcribe/{id} | Poll status, read the transcript, or export SRT, WebVTT or plain text. |
| DELETE | /v1/transcribe/{id} | Soft-delete a finished transcription. |
| POST | /v1/transcribe/{id}/cancel | Stop an in-flight job and release its reserved balance. |
| GET | /v1/transcriptions | List jobs with filters and cursor pagination. |
| GET | /v1/balance | Check remaining balance and estimated audio hours. |
Every operation, with typed parameters and response schemas, is described in the API reference and in the OpenAPI 3.1 document at /openapi.json (YAML at /openapi.yaml).
Errors
Errors are always JSON, always the same shape, and always carry a stable machine-readable code you can branch on. The message is for humans and may be reworded; the code will not be.
{"error": {"code": "INSUFFICIENT_BALANCE","message": "Add funds to continue transcribing.","docs_url": "https://platform.scriptivox.com/docs/api-reference#error-codes"}}
Two failure paths exist, and they need different handling. Synchronous errors — a bad key, a malformed body, no balance — come back on the call you just made. Asynchronous errors come back later on GET /v1/transcribe/{id} with status: "failed" and a error object describing why. The full code list for both is in the API reference.
Rate limits
Limits are per endpoint and per IP, and every response advertises them, including responses you get before authenticating:
RateLimit-Policy: "endpoint";q=100;w=60, "ip";q=300;w=60RateLimit: "endpoint";r=94;t=41
A 429 carries Retry-After in seconds. Back off on it rather than retrying immediately — the counters are per-minute windows, so a short wait clears them.
Formats and languages
25 container formats are accepted, covering everything common in audio and video: MP3, WAV, FLAC, M4A, AAC, OGG, Opus, MP4, MOV, MKV, WebM and the rest. 119 languages are supported with automatic detection.
Pass language when you know it. Auto-detection usually works, but it can misclassify short clips, code-switched audio, or files that open with music. An explicit ISO code is both faster and more accurate.
Billing
$0.20 per hour of audio, charged per second of actual duration, with no subscription and no minimum. Cost is reserved once the file's duration is known and settled when the job finishes. Failed and cancelled jobs are free — the reservation is released. See Pricing.
Service health
Real-time status, uptime history, and incident reports for every part of the API are published at status.scriptivox.com. Subscribe there to get notified about incidents and scheduled maintenance.
Start building
Quickstart
Get your first transcription running in under 5 minutes
API Reference
Upload, transcribe, and retrieve results via REST endpoints
Webhooks
Receive real-time notifications when transcriptions complete
Pricing
Simple pay-as-you-go at $0.20/hour of audio processed
Capabilities
Fast Transcription
High-accuracy transcription powered by Whisper. Supports 119 languages with automatic detection.
Speaker Diarization
Identify who said what. Detect and label multiple speakers in your audio automatically.
Word-level Timestamps
Precise start and end times for every word, plus confidence scores where the alignment model supports them. On by default — pass align: false to opt out.
Webhook Notifications
Get notified when transcriptions complete. HMAC-signed payloads for security.
URL Transcription
Transcribe from any public URL — Google Drive, Dropbox, OneDrive, or direct file links. No upload step needed.
119 Languages
Automatic language detection across 119 languages. Just send your audio — no configuration needed.
Every page in these docs
- Quickstart — first transcription in five minutes, in Python, JavaScript and curl.
- API reference — every endpoint, parameter, error code, rate limit, format and language.
- Authentication — API keys, the two accepted headers, rotation, and what a 401 means.
- Webhooks — HMAC-signed completion callbacks instead of polling.
- Pricing — pay-as-you-go rates and how reservations settle.
- Versioning — what can change in
/v1without notice, and how deprecations are announced. - Use cases — folder watchers, batch pipelines and other worked examples.
- MCP server — transcription as native tool calls from Claude, ChatGPT and other MCP clients.
- CLI —
npx @scriptivox-api/cli, for shells and scripts.
For AI agents
Every page here has a markdown twin: append .md to its path, or request the same URL with Accept: text/markdown. /llms.txt describes what this service is for and when not to reach for it, and /llms-full.txt is all of this documentation in a single fetch.