Scriptivox logoScriptivox

    Get started

    OverviewQuickstartAuthenticationPricing

    API Reference

    TranscribeFile UploadGet ResultBalanceError CodesRate LimitsFormatsLanguages

    Guides

    WebhooksCLIMCP ServerVersioning

    Use Cases

    Folder Watcher
    Scriptivox logoScriptivoxAPI Documentation

    Scriptivox API

    Transcribe audio and video at $0.20/hour, billed per second. Speaker diarization, word-level timestamps, and automatic language detection across 119 languages — all included.

    Powered by Whisper. Pay-as-you-go, no minimums, no subscriptions.

    Developer quickstart

    Transcribe audio and video files with word-level timestamps, speaker diarization, and language detection. Build transcription into your product in minutes.

    Get started
    import requests
    import time
    API_KEY = "sk_live_YOUR_KEY"
    BASE = "https://api.scriptivox.com/v1"
    # 1. Transcribe from a URL
    resp = requests.post(f"{BASE}/transcribe",
    headers={"Authorization": API_KEY},
    json={"url": "https://example.com/audio.mp3"})
    job = resp.json()
    # 2. Get result
    while True:
    result = requests.get(
    f"{BASE}/transcribe/{job['id']}",
    headers={"Authorization": API_KEY}
    ).json()
    if result["status"] == "completed":
    print(result["result"]["full_transcript"])
    break
    time.sleep(5)

    What the API does

    Send a recorded audio or video file — by public URL, or by uploading it to a presigned URL we hand you — and get back a structured transcript. Every response carries the full text, an array of utterances with start and end times, and, unless you opt out, a timestamp for every individual word. Speaker labels are added when you ask for diarization. The same transcript can be exported as SRT or WebVTT subtitles, or as plain text, by adding a query parameter to the fetch.

    Jobs are asynchronous. POST /v1/transcribe accepts the work and returns immediately with status: "created"; the file is downloaded and validated afterwards. That ordering matters more than anything else on this page: a bad URL, an unreadable file or an unsupported codec surfaces on the poll, as status: "failed", not as an error on the submit call. Clients that only check the response to POST will believe every job succeeded.

    This is not a streaming or real-time dictation API. Every job takes a complete file, between 1 second and 10 hours long and at most 5 GB.

    Authentication

    Every request needs an API key from the dashboard, passed as either Authorization: sk_live_… (a Bearer prefix is accepted) or X-Api-Key: sk_live_…. See Authentication for rotation, the key limit, and what each 401 means.

    Endpoints

    Base URL: https://api.scriptivox.com/v1

    MethodPathWhat it does
    POST/v1/uploadGet a presigned URL for a file you host yourself.
    POST/v1/transcribeStart a job from a public URL or a completed upload.
    GET/v1/transcribe/{id}Poll status, read the transcript, or export SRT, WebVTT or plain text.
    DELETE/v1/transcribe/{id}Soft-delete a finished transcription.
    POST/v1/transcribe/{id}/cancelStop an in-flight job and release its reserved balance.
    GET/v1/transcriptionsList jobs with filters and cursor pagination.
    GET/v1/balanceCheck remaining balance and estimated audio hours.

    Every operation, with typed parameters and response schemas, is described in the API reference and in the OpenAPI 3.1 document at /openapi.json (YAML at /openapi.yaml).

    Errors

    Errors are always JSON, always the same shape, and always carry a stable machine-readable code you can branch on. The message is for humans and may be reworded; the code will not be.

    json
    {
    "error": {
    "code": "INSUFFICIENT_BALANCE",
    "message": "Add funds to continue transcribing.",
    "docs_url": "https://platform.scriptivox.com/docs/api-reference#error-codes"
    }
    }

    Two failure paths exist, and they need different handling. Synchronous errors — a bad key, a malformed body, no balance — come back on the call you just made. Asynchronous errors come back later on GET /v1/transcribe/{id} with status: "failed" and a error object describing why. The full code list for both is in the API reference.

    Rate limits

    Limits are per endpoint and per IP, and every response advertises them, including responses you get before authenticating:

    http
    RateLimit-Policy: "endpoint";q=100;w=60, "ip";q=300;w=60
    RateLimit: "endpoint";r=94;t=41

    A 429 carries Retry-After in seconds. Back off on it rather than retrying immediately — the counters are per-minute windows, so a short wait clears them.

    Formats and languages

    25 container formats are accepted, covering everything common in audio and video: MP3, WAV, FLAC, M4A, AAC, OGG, Opus, MP4, MOV, MKV, WebM and the rest. 119 languages are supported with automatic detection.

    Pass language when you know it. Auto-detection usually works, but it can misclassify short clips, code-switched audio, or files that open with music. An explicit ISO code is both faster and more accurate.

    Billing

    $0.20 per hour of audio, charged per second of actual duration, with no subscription and no minimum. Cost is reserved once the file's duration is known and settled when the job finishes. Failed and cancelled jobs are free — the reservation is released. See Pricing.

    Service health

    Real-time status, uptime history, and incident reports for every part of the API are published at status.scriptivox.com. Subscribe there to get notified about incidents and scheduled maintenance.

    Start building

    Quickstart

    Get your first transcription running in under 5 minutes

    API Reference

    Upload, transcribe, and retrieve results via REST endpoints

    Webhooks

    Receive real-time notifications when transcriptions complete

    Pricing

    Simple pay-as-you-go at $0.20/hour of audio processed

    Capabilities

    Fast Transcription

    High-accuracy transcription powered by Whisper. Supports 119 languages with automatic detection.

    Speaker Diarization

    Identify who said what. Detect and label multiple speakers in your audio automatically.

    Word-level Timestamps

    Precise start and end times for every word, plus confidence scores where the alignment model supports them. On by default — pass align: false to opt out.

    Webhook Notifications

    Get notified when transcriptions complete. HMAC-signed payloads for security.

    URL Transcription

    Transcribe from any public URL — Google Drive, Dropbox, OneDrive, or direct file links. No upload step needed.

    119 Languages

    Automatic language detection across 119 languages. Just send your audio — no configuration needed.

    Every page in these docs

    • Quickstart — first transcription in five minutes, in Python, JavaScript and curl.
    • API reference — every endpoint, parameter, error code, rate limit, format and language.
    • Authentication — API keys, the two accepted headers, rotation, and what a 401 means.
    • Webhooks — HMAC-signed completion callbacks instead of polling.
    • Pricing — pay-as-you-go rates and how reservations settle.
    • Versioning — what can change in /v1 without notice, and how deprecations are announced.
    • Use cases — folder watchers, batch pipelines and other worked examples.
    • MCP server — transcription as native tool calls from Claude, ChatGPT and other MCP clients.
    • CLI — npx @scriptivox-api/cli, for shells and scripts.

    For AI agents

    Every page here has a markdown twin: append .md to its path, or request the same URL with Accept: text/markdown. /llms.txt describes what this service is for and when not to reach for it, and /llms-full.txt is all of this documentation in a single fetch.