VisionStory API
Create lifelike AI avatar videos with a few lines of code — avatars, voices, talking video, and AI video generation behind one API key.
Capabilities
Everything the API can do. Each capability links to its guide or reference.
Talking Avatar Video
Turn a script (text or audio) into a lip-synced talking avatar video.
Avatars
Create custom avatars from a photo, or pick from the public library.
Voices
Clone a voice from audio samples and use it in any video.
Models
Machine-readable catalog of avatar rendering models.
Billing
Check subscription plan and remaining credits.
AI Video
Generate videos from text or images with frontier AI video models.
Text to Speech
Turn text into natural speech with any public or cloned voice; returns MP3 audio.
Image Generation
Generate or edit images from a text prompt and optional reference images; returns an image URL.
Assets
Upload and manage reusable media assets (image / audio / video).
Build with agents
Drive the API from an AI agent — three channels, same API key. See the agent guide for setup.
Agent Skill
Installable package — SKILL.md plus a zero-dependency CLI. For coding agents like Claude Code and Codex.
MCP server
Task-oriented tools over the Model Context Protocol. For chat agents like Claude Desktop — no code needed.
llms.txt
Machine-readable docs index; every page is also served as Markdown. For agents that browse and read docs on demand.
Models
Talking-avatar rendering models — query GET /api/v1/models for the machine-readable catalog.
vs_talk_v1
Speech Mode focuses on mouth movements and speech, ideal for clear lip-syncing.
vs_character_v4
Character Model v4 provides improved motion quality and stability.
Seedance video models
Frontier AI video generation — text-to-video and image-to-video. Query GET /api/v1/ai_video/models for the live capability sheet.