VisionStory OpenAPI
Get API key
Speech to Text/Align text with audio

Align text with audio

POST/api/v1/audio/align

Force-align known text against its spoken audio and get word-level timestamps — e.g. align a script with the MP3 from POST /api/v1/tts to build accurate subtitles for any voice. Billing: 1 credit per started 5 minutes of audio, charged only on success. Synchronous — audio up to 30 minutes. Per-key concurrency is limited during beta; requests beyond the limit are rejected, not queued.

Headers

X-API-KeystringRequired
Your VisionStory API key (sk-vs-...), kept server-side. Create one at OpenApi (Pro plan and up).

Request body

audioMediaRefRequired
Audio to align: a WAV or MP3 up to 15MB (asset_id / url / inline_data).
Show 3 propertiesHide 3 properties
asset_idstring | nullOptional
Asset ID from POST /api/v1/asset; use for materials reused across requests.
urlstring | nullOptional
Publicly accessible media URL for one-off use; not added to your asset library.
inline_dataInlineDataModel | nullOptional
Inline base64 media data for one-off use; not added to your asset library. Images: image/jpeg, image/jpg, image/png, image/webp, image/bmp, image/tiff, image/gif; audio: audio/wav, audio/x-wav, audio/wave, audio/mpeg, audio/mp3; video: video/mp4, video/quicktime, video/mov.
Show 2 propertiesHide 2 properties
mime_typestringRequired
MIME type of the inline data; the gateway uses it to tell image / audio / video apart. The accepted set depends on the endpoint — see the field carrying this object. Talking video accepts audio ['audio/avi', 'audio/mpeg', 'audio/mp3', 'audio/mp4', 'audio/m4a', 'audio/wav'] and images ['image/jpeg', 'image/jpg', 'image/png', 'image/webp', 'image/heic'].
datastringRequired
The file's raw bytes encoded as a base64 string (no data: URI prefix).
textstringRequired
The exact text spoken in the audio (e.g. the script you synthesized with POST /api/v1/tts). Timestamps are aligned against this text.

Response

200Successful Response

Successful calls return a standard envelope: the endpoint payload under data (its fields are documented below), plus a message string ("success") and an ISO 8601 server_time.

Response fields (data)
wordsarray of WordDtoOptional
Word-level timestamps for the given text.
Show 4 propertiesHide 4 properties
textstringRequired
The word as written.
start_secnumberRequired
Word start time in seconds.
end_secnumberRequired
Word end time in seconds.
speakerstring | nullOptional
Speaker label, present only when diarize was requested.
duration_secnumberOptionalDefault 0
Duration of the input audio in seconds.
cost_creditintegerOptionalDefault 0
Credits charged for this request.

Errors

All error responses share one JSON envelope: an error object with a numeric code, a human-readable message, an optional details string, and an optional hint giving an actionable next step (handy for AI agents).

errorErrorDetailRequired
Error payload returned with every non-2xx response. Present only on failure; successful calls use the standard success envelope instead.
Show 4 propertiesHide 4 properties
codeintegerRequired
Machine-readable error code. Mirrors the HTTP status for transport-level failures (e.g. 401, 404, 422, 500) and may carry a business-specific code otherwise.
messagestringRequired
Human-readable explanation of what went wrong. Safe to log or surface to end users; not localized.
detailsstring | nullOptional
Optional structured detail about the failure, e.g. a JSON string of per-field validation errors on a 422. Absent when there is nothing extra to report.
hintstring | nullOptional
Actionable next step for resolving the error, written for both humans and AI agents (e.g. how to fix the request, or where to obtain an API key). May be absent.