VisionStory OpenAPI
Get API key

Two synchronous endpoints turn speech audio into structured text and timing data:

Both are in beta and may change. Available to any account with an active paid subscription — no separate allowlisting.

Input

Both endpoints accept the same audio object with exactly one source, the same shape used for reference media elsewhere in the API:

{
  "audio": {"asset_id": "7241059991822401536"}
}
{
  "audio": {"url": "https://example.com/narration.mp3"}
}
{
  "audio": {"inline_data": {"mime_type": "audio/mp3", "data": "<base64>"}}
}

Input limits: WAV or MP3, up to 15 MB. These are synchronous endpoints best suited to short clips (a script's worth of narration, an interview segment); very long audio may time out — for those, split into segments. One-off url / inline_data inputs are not added to your asset library; upload via POST /api/v1/asset first if you want to reuse the file.

Transcribe

curl -X POST https://openapi.visionstory.ai/api/v1/audio/transcribe \
  -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" \
  -d '{"audio": {"url": "https://example.com/interview.mp3"}, "diarize": true, "srt": true}'

Response data:

{
  "text": "Welcome to the show. Thanks for having me.",
  "language": "en",
  "words": [
    {"text": "Welcome", "start_sec": 0.08, "end_sec": 0.51, "speaker": "speaker_0"},
    {"text": "to", "start_sec": 0.55, "end_sec": 0.63, "speaker": "speaker_0"}
  ],
  "srt": "1\n00:00:00,080 --> 00:00:02,110\nWelcome to the show.\n",
  "duration_sec": 12.4,
  "cost_credit": 1
}

Align

curl -X POST https://openapi.visionstory.ai/api/v1/audio/align \
  -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" \
  -d '{"audio": {"url": "https://cdn.example.com/tts-output.mp3"}, "text": "Hello from VisionStory."}'

Response data carries words (same shape as transcribe, no speaker), duration_sec, and cost_credit.

Billing and limits