One hour. One pass.
Preserve global context across continuous audio instead of stitching together disconnected chunks.
A production API for Microsoft's VibeVoice-ASR—turning long-form, multilingual audio into speaker-aware transcripts with precise timestamps.
Let’s take the full recording—no manual slicing.
And keep every speaker and timestamp attached.
That’s exactly what this endpoint is built for.
Built for the recordings that matter: research calls, podcasts, meetings, lectures, and interviews.
Preserve global context across continuous audio instead of stitching together disconnected chunks.
Diarization and timestamps arrive with the transcript—not as a second job.
More than 50 languages, with native support for code-switching.
Sign in with your email, add credit, generate an API key.
POST one continuous recording — raw bytes, up to 100 MiB.
Poll the job and receive who spoke, when, and what they said.
Sign in with your email, create a key, top up a balance.
One queue path for every length. Post your audio file, get a job id back in under a second, then poll until it's done — no timeouts on hour-long files.
curl -X POST \
https://api.vibevoice-asr.com/v1/jobs \
-H "Authorization: Bearer $VVASR_KEY" \
-F file=@meeting.mp3
{ "job_id": "…", "status": "processing", "eta_seconds": 45 }Charged on the audio duration your file actually decodes to — not on upload size, not on wall-clock time. Failed and cancelled jobs cost nothing.
Long-form speech deserves long-form context.