Speech to Text
Two synchronous endpoints turn speech audio into structured text and timing data:
POST /api/v1/audio/transcribe— speech to text with word-level timestamps. Optionally label speaker turns (diarize: true) or get a ready-to-use SRT subtitle track (srt: true). Requestingsrtruns speaker diarization internally (the subtitle exporter needs it), at no extra cost; speaker labels still only appear inwordsand in the SRT when you setdiarize: true.POST /api/v1/audio/align— you already know the text (e.g. the script you sent toPOST /api/v1/tts); force-align it against the spoken audio and get word-level timestamps. This works for audio from any voice engine, so it is the recommended way to build accurate subtitles for TTS output.
Both are in beta and may change. Available to any account with an active paid subscription — no separate allowlisting.
Input
Both endpoints accept the same audio object with exactly one source, the same shape used for
reference media elsewhere in the API:
{
"audio": {"asset_id": "7241059991822401536"}
}{
"audio": {"url": "https://example.com/narration.mp3"}
}{
"audio": {"inline_data": {"mime_type": "audio/mp3", "data": "<base64>"}}
}Input limits: WAV or MP3, up to 15 MB. These are synchronous endpoints best suited to short
clips (a script's worth of narration, an interview segment); very long audio may time out — for
those, split into segments. One-off url / inline_data inputs are not added to your asset
library; upload via POST /api/v1/asset first if you want to reuse the file.
Transcribe
curl -X POST https://openapi.visionstory.ai/api/v1/audio/transcribe \
-H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" \
-d '{"audio": {"url": "https://example.com/interview.mp3"}, "diarize": true, "srt": true}'Response data:
{
"text": "Welcome to the show. Thanks for having me.",
"language": "en",
"words": [
{"text": "Welcome", "start_sec": 0.08, "end_sec": 0.51, "speaker": "speaker_0"},
{"text": "to", "start_sec": 0.55, "end_sec": 0.63, "speaker": "speaker_0"}
],
"srt": "1\n00:00:00,080 --> 00:00:02,110\nWelcome to the show.\n",
"duration_sec": 12.4,
"cost_credit": 1
}Align
curl -X POST https://openapi.visionstory.ai/api/v1/audio/align \
-H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" \
-d '{"audio": {"url": "https://cdn.example.com/tts-output.mp3"}, "text": "Hello from VisionStory."}'Response data carries words (same shape as transcribe, no speaker), duration_sec, and
cost_credit.
Billing and limits
- 1 credit per started 5 minutes of audio, per request; charged only on success.
- Input: WAV or MP3, up to 15 MB. Synchronous — best for short clips; large files may time out (a failed request is not charged).
- A few concurrent requests per key during beta; excess requests are rejected with
429, not queued. - Results are returned inline and not stored; save what you need.