VisionStory API
Create lifelike AI avatar videos with a few lines of code — avatars, voices, talking video, and AI video generation behind one API key.
Capabilities
Everything the API can do. Each capability links to its guide or reference.
Talking Avatar Video
Turn a script (text or audio) into a lip-synced talking avatar video.
Avatars
Create custom avatars from a photo, or pick from the public library.
Voices
Clone a voice from audio samples and use it in any video.
Models
Machine-readable catalog of avatar rendering models.
Billing
Check subscription plan and remaining credits.
AI Video
Generate videos from text or images with frontier AI video models.
Text to Speech
Turn text into natural speech with any public or cloned voice; returns MP3 audio.
Image Generation
Generate or edit images from a text prompt and optional reference images; returns an image URL.
Speech to Text
Transcribe speech to text with word timestamps, or force-align a known script for accurate subtitles.
Media Understanding
Label or extract structured JSON from images, audio, and video with your own schema.
Assets
Upload and manage reusable media assets (image / audio / video).
Models
Talking-avatar rendering models — query GET /api/v1/models for the machine-readable catalog.
vs_character_v4
Character Model v4 provides improved motion quality and stability.
Seedance video models
Frontier AI video generation — text-to-video and image-to-video. Query GET /api/v1/ai_video/models for the live capability sheet.