VisionStory OpenAPI
Get API key
Speech to Text/Transcribe audio
POST/api/v1/audio/transcribe

Transcribe speech audio to text with word-level timestamps; optionally detect speaker turns (diarize) or render an SRT subtitle track (srt). Billing: 1 credit per started 5 minutes of audio, charged only on success. Synchronous — audio up to 30 minutes. Per-key concurrency is limited during beta; requests beyond the limit are rejected, not queued.

Headers

X-API-KeystringRequired
Your VisionStory API key (sk-vs-...), kept server-side. Create one at OpenApi (Pro plan and up).

Request body

audioMediaRefRequired
Audio to transcribe: a WAV or MP3 up to 15MB (asset_id / url / inline_data). This synchronous endpoint suits short clips; very long audio may time out.
Show 3 propertiesHide 3 properties
asset_idstring | nullOptional
Asset ID from POST /api/v1/asset; use for materials reused across requests.
urlstring | nullOptional
Publicly accessible media URL for one-off use; not added to your asset library.
inline_dataInlineDataModel | nullOptional
Inline base64 media data for one-off use; not added to your asset library. Images: image/jpeg, image/jpg, image/png, image/webp, image/bmp, image/tiff, image/gif; audio: audio/wav, audio/x-wav, audio/wave, audio/mpeg, audio/mp3; video: video/mp4, video/quicktime, video/mov.
Show 2 propertiesHide 2 properties
mime_typestringRequired
MIME type of the inline data; the gateway uses it to tell image / audio / video apart. The accepted set depends on the endpoint — see the field carrying this object. Talking video accepts audio ['audio/avi', 'audio/mpeg', 'audio/mp3', 'audio/mp4', 'audio/m4a', 'audio/wav'] and images ['image/jpeg', 'image/jpg', 'image/png', 'image/webp', 'image/heic'].
datastringRequired
The file's raw bytes encoded as a base64 string (no data: URI prefix).
diarizebooleanOptionalDefault false
Set true to detect speaker turns; each word then carries a speaker label.
srtbooleanOptionalDefault false
Set true to also return an SRT subtitle rendering of the transcript.

Response

200Successful Response

Successful calls return a standard envelope: the endpoint payload under data (its fields are documented below), plus a message string ("success") and an ISO 8601 server_time.

Response fields (data)
textstringRequired
Full transcript text.
languagestringOptionalDefault ""
Detected language code of the audio, e.g. en; may be empty.
wordsarray of WordDtoOptional
Word-level timestamps.
Show 4 propertiesHide 4 properties
textstringRequired
The word as written.
start_secnumberRequired
Word start time in seconds.
end_secnumberRequired
Word end time in seconds.
speakerstring | nullOptional
Speaker label, present only when diarize was requested.
srtstringOptionalDefault ""
SRT subtitle text; empty unless srt: true was requested.
duration_secnumberOptionalDefault 0
Duration of the input audio in seconds.
cost_creditintegerOptionalDefault 0
Credits charged for this request.

Errors

All error responses share one JSON envelope: an error object with a numeric code, a human-readable message, an optional details string, and an optional hint giving an actionable next step (handy for AI agents).

errorErrorDetailRequired
Error payload returned with every non-2xx response. Present only on failure; successful calls use the standard success envelope instead.
Show 4 propertiesHide 4 properties
codeintegerRequired
Machine-readable error code. Mirrors the HTTP status for transport-level failures (e.g. 401, 404, 422, 500) and may carry a business-specific code otherwise.
messagestringRequired
Human-readable explanation of what went wrong. Safe to log or surface to end users; not localized.
detailsstring | nullOptional
Optional structured detail about the failure, e.g. a JSON string of per-field validation errors on a 422. Absent when there is nothing extra to report.
hintstring | nullOptional
Actionable next step for resolving the error, written for both humans and AI agents (e.g. how to fix the request, or where to obtain an API key). May be absent.