Transcribe long audio
in three calls.
Submit an audio file, get a job id back in under a second, then poll until it's done. One path for a 5-second clip and a 3-hour recording alike — nothing times out because nothing is held open.
BASE URLhttps://api.vibevoice-asr.com
01Quickstart
Create a key in the console, then submit a file and poll for the result. Both calls use the same key.
curl -X POST https://api.vibevoice-asr.com/v1/jobs \
-H "Authorization: Bearer $VVASR_KEY" \
-F file=@meeting.mp3
{ "job_id": "8f14e45f-…", "status": "processing", "eta_seconds": 42 }curl https://api.vibevoice-asr.com/v1/jobs/8f14e45f-… \
-H "Authorization: Bearer $VVASR_KEY"
{
"job_id": "8f14e45f-…",
"status": "succeeded",
"eta_seconds": 1,
"result": {
"text": "Let's take the full recording. And keep every speaker attached.",
"segments": [
{ "start": 0.0, "end": 3.4, "speaker": "1", "text": "Let's take the full recording." },
{ "start": 3.4, "end": 6.1, "speaker": "2", "text": "And keep every speaker attached." }
]
}
}Call the API from your server, not from a browser. There are no CORS headers, so a fetch from page JavaScript will be blocked — and it would expose your key to anyone who opens devtools.
02Authentication
Every request needs Authorization: Bearer <key>. Keys look like vvasr_user_… and are shown exactly once, when you create them. Revoking a key takes effect immediately — there is no cache to wait out.
You can hold up to 10 active keys. Jobs belong to your account, not to the key that created them, so any of your keys can read any of your jobs. Someone else's job returns 404, never 403.
03Submit a job
POST /v1/jobs
Send a multipart/form-data form with the audio in it — either the bytes or a link, never both:
# upload the file
curl -X POST https://api.vibevoice-asr.com/v1/jobs -H "Authorization: Bearer $VVASR_KEY" \
-F file=@meeting.mp3
# or let us download it
curl -X POST https://api.vibevoice-asr.com/v1/jobs -H "Authorization: Bearer $VVASR_KEY" \
-F file_url=https://cdn.example.com/meeting.mp3| file | The audio bytes, up to 100 MiB. The format is detected from the bytes themselves, so the file's extension and Content-Type don't matter. |
| file_url | Instead of uploading: a public http(s) URL we download server-side. Same 100 MiB limit, and it must report a Content-Length. |
| callback_url | Optional: an https URL we POST to once when the job finishes. See Completion callbacks. |
| Formats | wav, mp3, flac, ogg, opus, m4a, mp4, webm |
| Video containers | mp4 and webm are accepted; the audio track is extracted server-side. |
Returns 201:
{ "job_id": "8f14e45f-…", "status": "processing", "eta_seconds": 42 }eta_seconds is an estimate of time to completion — use it to pace your first poll. The response also carries an Inference-Id header equal to the job_id.
04Fetch the result
GET /v1/jobs/{job_id}
| status | queued → processing → succeeded | failed | cancelled |
| eta_seconds | Counts down while the job is in flight; floors at 1 |
| result | Present only when succeeded |
| error | Present only when failed, as { code, message } |
Each segment carries start and end in seconds, the text, and speaker when the model identified one. Treat speaker as optional — it is a label, not a number, and it is omitted rather than guessed when the model gives nothing.
How often to poll
First poll after min(eta_seconds × 0.8, 2s), then back off exponentially up to 5s. A job that has not finished within 25 minutes is failed with deadline_exceeded, so that is your upper bound on waiting.
05Completion callbacks
POST {your callback_url} — sent by us, once per job
Leave an https URL in the callback_url form field when you submit, and we POST to it when the job reaches a terminal state, instead of making you poll for the news:
{ "event": "job.completed", "job_id": "8f14e45f-…", "status": "succeeded" }| status | succeeded | failed | cancelled. failed adds error: { code, message }. |
| What's not in it | The transcript. Fetch it with GET /v1/jobs/{job_id} — so the payload contains nothing secret, and a forged callback can at worst make your server do one wasted read. |
| Delivery | One attempt, 5s timeout, no retries. A 2xx from your server means delivered; anything else is logged on our side and dropped. Keep polling as your fallback if you need certainty. |
| Requirements | https only, public host, at most 2048 bytes, no user:pass. Anything else fails the submit with 400 invalid_callback_url. |
| Authenticity | Callbacks aren't signed — the URL is the credential. Put a random token in it (…/hook?t=f3a9…) and check it on arrival. We never follow redirects from your endpoint. |
Jobs that fail synchronously at submit time (4xx/5xx on the POST itself) don't trigger a callback — you already have the error in hand. Jobs we fail on your behalf (deadline_exceeded, dispatch_lost) do: for a client that isn't polling, that callback is the only signal it gets.
06Cancel a job
DELETE /v1/jobs/{job_id}
Stops a job that is still queued or processing and refunds the full reservation. Returns { "job_id": "…", "status": "cancelled" }. A job that already reached a terminal state returns 404 — there is nothing left to cancel.
07What you get charged
$0.0035 per audio minute, computed from the duration the transcription node actually decoded — not your file size, not how long the job took to run. Failed and cancelled jobs cost nothing.
Why your balance dips and then partly comes back
When you submit, we reserve an estimate based on the duration read from the file header, so a long job can't start against an empty balance. When it finishes we settle against the real decoded duration and return the difference. In the console you'll see this as a hold followed by a settle — the second one is usually a partial refund.
08Errors
Every error has the same shape: { "error": { "code": "…", "message": "…" } }. Some carry extra fields, noted below.
| Status | Code | What happened |
|---|---|---|
| 401 | unauthorized | Missing, malformed, unknown, or revoked key. |
| 402 | insufficient_balance | Not enough balance for the reservation. Includes required_nano_usd and balance_nano_usd so you can tell the user how much to add. |
| 413 | payload_too_large | The uploaded file — or the size file_url declares — exceeds 100 MiB. Enforced while the upload streams, so a missing or dishonest Content-Length doesn't get around it. Includes max_bytes. |
| 415 | unsupported_format | The bytes aren't a recognised audio container. Includes supported_formats. |
| 415 | unsupported_media_type | The body isn't multipart/form-data with a valid boundary. Raw bytes and JSON bodies land here too — wrap the file in a form. |
| 400 | invalid_multipart | The form didn't work out: no file and no file_url, both at once, more than one file part, malformed framing, or oversized headers/fields. |
| 400 | invalid_callback_url | callback_url isn't https, points at a private host, embeds credentials, or is too long. |
| 400 | invalid_file_url | file_url isn't http(s), or points at a private or loopback host. |
| 400 | fetch_failed | We couldn't download file_url — unreachable, non-2xx, or no Content-Length. Includes upstream_status when the origin answered. |
| 400 | no_audio_stream | A video container with no audio track. |
| 400 | audio_too_long | Audio runs longer than 60 minutes. Includes max_seconds and audio_seconds, plus estimated: true when we had to infer the duration from the file size rather than read it from the container. |
| 429 | too_many_requests | No capacity right now. Retry with backoff. |
| 503 | no_backend | Capacity existed but no node accepted the job. Retry. |
| 503 | misconfigured | Server-side configuration problem. Not something a retry fixes — tell us. |
| 404 | not_found | Unknown job, or one that isn't yours. We don't distinguish the two on purpose. |
Failures after the job was accepted
These arrive as error.code inside a 200 poll response, with status: "failed": backend_failed, backend_timeout, bad_output, deadline_exceeded, dispatch_lost. None of them are billed — and each also fires your callback, if you left one.
09Limits
| Maximum upload | 100 MiB |
| Maximum audio duration | 60 minutes |
| Maximum time to completion | 25 minutes |
| Uploaded audio retained for | 3 hours |
| Active API keys per account | 10 |
| Callback delivery timeout | 5 s (one attempt) |
| Callback URL length | 2048 bytes |
Maximum time to completion is a ceiling, not a typical duration — the estimate for your own job is the eta_seconds returned when you submit it. Work we predict cannot finish inside that ceiling is declined up front with 429 rather than accepted and timed out.
Transcription parameters (language hints, timestamp granularity) aren't available yet. If you need one, tell us — it helps us prioritise.
Ready?
Create an API key ↗