VisionStory OpenAPI
Get API key
Voices/List voices
GET/api/v1/voices

List the voices you can synthesize with: the public voice library plus any voices you have cloned. Optionally filter by locale and/or provider. Pass limit (and the returned next_cursor) to page through the filtered public library; without limit the full list is returned in one response.

Query parameters

cursorintegerOptionalDefault 0
Pagination cursor over public_voices, from the previous page's next_cursor. Omit or pass 0 for the first page.
limitinteger | nullOptional
Maximum public_voices per page (1-500). Omit to return the full library unpaginated (my_voices is always returned in full on the first page).
localestring | nullOptional
Filter by BCP 47 locale (case-insensitive). A bare language such as en or zh matches every regional variant; en-GB / zh-TW / zh-HK match that region only. See the locale field of each voice.
providerstring | nullOptional
Filter to voices from this engine (case-insensitive): elevenlabs / seed / gemini / minimax.

Headers

X-API-KeystringRequired
Your VisionStory API key (sk-vs-...), kept server-side. Create one at OpenApi (Pro plan and up).

Response

200Successful Response

Successful calls return a standard envelope: the endpoint payload under data (its fields are documented below), plus a message string ("success") and an ISO 8601 server_time.

Response fields (data)
public_voicesarray of VoiceDtoRequired
Platform-provided voices available for text-to-speech.
Show 11 propertiesHide 11 properties
voice_idstringRequired
Voice identifier. Pass it as voice_id in a text script, or as the target voice when converting uploaded audio.
is_freebooleanRequired
Whether this voice is usable on the free plan; premium voices require a paid plan.
preview_audio_urlstringRequired
URL of a short sample clip demonstrating how this voice sounds.
tagsstring | nullOptional
Free-text descriptors to help pick a voice; superseded by the structured fields below but kept for compatibility. May be empty.
languagestringRequired
Human-readable primary language, e.g. english. Use locale for the machine-readable tag.
localestringOptionalDefault ""
BCP 47 locale of this voice, e.g. en-US / en-GB / zh-TW / zh-HK; a bare language code when no region is known; empty for prompt-driven voices (gemini). Filter with locale, and pass the same value as locale to POST /api/v1/tts to pin pronunciation.
providerstringOptionalDefault ""
Speech engine this voice comes from (e.g. elevenlabs / gemini / minimax), so you can pick voices by engine. Filter with provider. Descriptive, not a guarantee: synthesis may fall back to another engine.
genderstringOptionalDefault ""
Voice gender, e.g. male / female; may be empty.
agestringOptionalDefault ""
Perceived age band, e.g. young / middle-aged / old; may be empty.
accentstringOptionalDefault ""
Accent, e.g. american / british / mandarin; may be empty.
use_casesarray of stringOptional
Suggested use cases, e.g. narration / conversational; may be empty.
my_voicesarray of VoiceDtoRequired
Voices this account has cloned. Included on the first page only when paginating.
Show 11 propertiesHide 11 properties
voice_idstringRequired
Voice identifier. Pass it as voice_id in a text script, or as the target voice when converting uploaded audio.
is_freebooleanRequired
Whether this voice is usable on the free plan; premium voices require a paid plan.
preview_audio_urlstringRequired
URL of a short sample clip demonstrating how this voice sounds.
tagsstring | nullOptional
Free-text descriptors to help pick a voice; superseded by the structured fields below but kept for compatibility. May be empty.
languagestringRequired
Human-readable primary language, e.g. english. Use locale for the machine-readable tag.
localestringOptionalDefault ""
BCP 47 locale of this voice, e.g. en-US / en-GB / zh-TW / zh-HK; a bare language code when no region is known; empty for prompt-driven voices (gemini). Filter with locale, and pass the same value as locale to POST /api/v1/tts to pin pronunciation.
providerstringOptionalDefault ""
Speech engine this voice comes from (e.g. elevenlabs / gemini / minimax), so you can pick voices by engine. Filter with provider. Descriptive, not a guarantee: synthesis may fall back to another engine.
genderstringOptionalDefault ""
Voice gender, e.g. male / female; may be empty.
agestringOptionalDefault ""
Perceived age band, e.g. young / middle-aged / old; may be empty.
accentstringOptionalDefault ""
Accent, e.g. american / british / mandarin; may be empty.
use_casesarray of stringOptional
Suggested use cases, e.g. narration / conversational; may be empty.
next_cursorintegerOptionalDefault 0
Cursor for the next page of public_voices; pass it back as cursor. 0 means no more pages (always 0 when the request did not paginate).

Errors

All error responses share one JSON envelope: an error object with a numeric code, a human-readable message, an optional details string, and an optional hint giving an actionable next step (handy for AI agents).

errorErrorDetailRequired
Error payload returned with every non-2xx response. Present only on failure; successful calls use the standard success envelope instead.
Show 4 propertiesHide 4 properties
codeintegerRequired
Machine-readable error code. Mirrors the HTTP status for transport-level failures (e.g. 401, 404, 422, 500) and may carry a business-specific code otherwise.
messagestringRequired
Human-readable explanation of what went wrong. Safe to log or surface to end users; not localized.
detailsstring | nullOptional
Optional structured detail about the failure, e.g. a JSON string of per-field validation errors on a 422. Absent when there is nothing extra to report.
hintstring | nullOptional
Actionable next step for resolving the error, written for both humans and AI agents (e.g. how to fix the request, or where to obtain an API key). May be absent.