AI Video
Generate AI videos with Seedance, Kling and Wan models — text-to-video, image-to-video, and multimodal references — through the same API key and billing you already use for talking avatars.
Beta. These endpoints are in beta and may change. Available to any account with an active paid subscription — no separate allowlisting.
Endpoints
| Method | Path | What it does |
|---|---|---|
GET |
/api/v1/ai_video/models |
Machine-readable capability sheet for every model |
GET |
/api/v1/ai_video/cost |
Exact credit cost of a task before you submit it |
POST |
/api/v1/ai_video |
Submit a generation task |
GET |
/api/v1/ai_video |
Query one task (or up to 20 with video_ids) |
GET |
/api/v1/ai_videos |
List your tasks, newest first |
DELETE |
/api/v1/ai_video |
Delete a task |
POST |
/api/v1/asset |
Upload a reusable media asset |
GET |
/api/v1/assets |
List your assets |
DELETE |
/api/v1/asset |
Delete an asset |
Quick start
Submit a text-to-video task, then poll until it finishes:
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"model_id": "seedance-2.0", "prompt": "A corgi surfing at sunset, cinematic lighting", "duration_sec": 8, "aspect_ratio": "9:16", "resolution": "1080p"}' https://openapi.visionstory.ai/api/v1/ai_video{
"data": {
"video_id": "7241059991822401536",
"status": "queued",
"cost_credit": 64
}
}Poll every 5–10 seconds:
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/ai_video?video_id=7241059991822401536"Models
GET /api/v1/ai_video/models returns allowed values and defaults for each parameter, plus media constraints. Always drive your integration from this endpoint — new models and parameter values appear there without any API change.
| model_id | Best for | Resolution | Duration | Capabilities |
|---|---|---|---|---|
seedance-2.5 |
Latest generation, single clips up to 30 s | 480p / 720p / 1080p | 4–30 s | text-to-video, image-to-video |
seedance-2.0 |
Flagship quality, multimodal references | 480p / 720p / 1080p | 4–15 s | text-to-video, image-to-video |
seedance-2.0-fast |
Lower latency and cost | 480p / 720p | 4–15 s | text-to-video, image-to-video |
seedance-2.0-mini |
Lightweight, most economical | 480p / 720p | 4–15 s | text-to-video, image-to-video |
kling-3.0 |
Highest resolution — the only 4K model | 720p / 1080p / 4k | 5 or 10 s | text-to-video, image-to-video |
kling-3.0-omni |
Reference-guided generation (up to 7 image/video refs) | 720p / 1080p | 3 / 5 / 7 / 10 / 15 s | text-to-video |
wan-3.0 |
All-in-one; the only model with audio references and any duration | 480p / 720p / 1080p | any 2–30 s | text-to-video, image-to-video |
wan-3.0-prime |
Same as wan-3.0 at higher quality, 1.5x the price |
480p / 720p / 1080p | any 2–30 s | text-to-video, image-to-video |
Native audio is on by default at no extra cost everywhere — the price depends only on model, resolution, and duration. Prompt length is 5000 characters on Seedance, 2500 on Kling, 20000 on Wan. These are hard caps, not targets: for Seedance, BytePlus recommends keeping each prompt to about 500 Chinese characters or 1,000 English words — longer prompts are accepted, but the model tends to pick out the main points and drop details. Seedance models accept aspect ratios 16:9 / 9:16 / 4:3 / 3:4 / 1:1. Kling and Wan differ in several ways — read the values from the models endpoint rather than assuming Seedance behaviour:
kling-3.0accepts norefs;kling-3.0-omniaccepts up to 7 image/video references (no audio) at no extra cost. With a video reference Kling cannot produce native audio, sogenerate_audio=truetogether with a video ref is rejected with 400 — leave it off when attaching video.- Kling aspect ratios are
16:9 / 9:16 / 1:1; for image-to-video the aspect ratio followsfirst_frameand an explicitaspect_ratiois rejected with 400. - Kling moderates the prompt and image inputs synchronously at submit (HTTP 403 with
37110/37111, nothing charged); Seedance moderates asynchronously (see Content moderation below). - 480p is available on every Seedance and Wan model and is the cheapest tier —
seedance-2.0-fastandseedance-2.0-miniat 480p are 1 credit/s. Kling has no 480p tier. 4K is available onkling-3.0only, at 8 credits/s — quote it first. kling-3.0-omniis the only model that defaults to1080p; every other model defaults to720p. Opening the cheaper 480p tier did not change any model's default, so it costs nothing to ignore.- Omitting
resolutiononkling-3.0-omnitherefore costs 3 credits/s — passresolutionexplicitly, or quote withGET /api/v1/ai_video/costfirst. - Wan accepts image, video and audio references, and is the only model that takes any integer duration from 2 to 30 s. Per-type reference limits are 10 image / 5 video / 5 audio, with video and audio each capped at 15 s of material in total; the models endpoint returns these as
refs.max_per_kindandrefs.total_duration_sec_max.refs.max(20) is the combined count across all three types, not a per-type limit. - Wan follows the reference material for framing whenever
refsis set, so passingaspect_ratiotogether withrefsis rejected with 400. Its defaultaspect_ratioisadaptive, which also follows the input material; for image-to-video the framing always followsfirst_frame. - Wan reference video and audio longer than 15 s is truncated to the first 15 s rather than rejected.
- Wan moderates synchronously at submit, like Kling.
- The API never substitutes another model: if the model you chose fails, the task resolves to
failedand credits are refunded in full.model_idon the task is always the model that rendered it.
Credits and cost
Generation is billed in credits, charged at submit time and automatically refunded in full if generation fails. Cost depends only on model, resolution, and duration (per second). Native audio and reference materials never change the price. Query the cost endpoint before submitting — it applies exactly the same formula as billing:
# seedance-2.0, 8 s at 1080p: 8 x 8 = 64 credits
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/ai_video/cost?model_id=seedance-2.0&duration_sec=8&resolution=1080p"{
"data": {
"credit": 64
}
}# kling-3.0, 10 s at 4K: 8 x 10 = 80 credits
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/ai_video/cost?model_id=kling-3.0&duration_sec=10&resolution=4k"Check your remaining balance with GET /api/v1/billing/credits.
Media inputs
Every media slot (first_frame, end_frame, refs[]) accepts exactly one of three forms:
| Form | Example | Use when |
|---|---|---|
url |
{"url": "https://your.site/img.jpg"} |
One-off use; fetched by our servers, not added to your asset library |
inline_data |
{"inline_data": {"mime_type": "image/png", "data": "<base64>"}} |
One-off use, no public URL available |
asset_id |
{"asset_id": "7241058823145623552"} |
Reused materials — upload once via the Assets API, reference many times |
Limits differ per model — the exact values for each slot are returned in first_frame and refs.image / refs.video / refs.audio of GET /api/v1/ai_video/models. Summary:
| Kind | Models | Max size | Formats | Constraints |
|---|---|---|---|---|
| Image | Seedance | 30 MB | jpg, jpeg, png, webp, bmp, tiff, gif | 300–6000 px per side, aspect ratio between 1:2.5 and 2.5:1 |
| Image | Kling | 10 MB | jpg, jpeg, png, webp, bmp, tiff, gif | at least 300 px per side, aspect ratio between 1:2.5 and 2.5:1 |
| Video | Seedance | 100 MB | mp4, mov | 2–15 s, 300–6000 px, 24–60 fps, aspect ratio 1:2.5–2.5:1 |
| Video | kling-3.0-omni refs only |
100 MB | mp4, mov | 3–15 s, 720–2160 px per side, aspect ratio 1:2.5–2.5:1 |
| Audio | Seedance | 15 MB | wav, mp3 | 2–15 s; cannot be the only reference |
| Image | Wan | 20 MB | jpg, jpeg, png, bmp, webp | 240–8000 px per side, aspect ratio between 1:8 and 8:1 |
| Video | Wan refs only | 100 MB | mp4, mov | 1–15 s each and 15 s total, 240–4096 px, at least 16 fps |
| Audio | Wan refs only | 15 MB | wav, mp3 | 1–15 s each and 15 s total |
Create a video
POST /api/v1/ai_video has two modes on one endpoint: provide first_frame (optionally end_frame) for image-to-video, provide refs for reference-guided text-to-video (character consistency, style, motion, or voice references), or provide neither for pure text-to-video. refs and first_frame are mutually exclusive. Unknown fields and unsupported values are rejected — nothing is silently ignored.
| Field | Required | Description |
|---|---|---|
model_id |
yes | See Models above |
prompt |
yes | Text prompt, up to 5000 characters (2500 on Kling, 20000 on Wan) |
client_request_id |
no | Idempotency key; resubmitting the same value within 24h returns the original task instead of charging again |
duration_sec |
no | Defaults to the model default |
aspect_ratio |
no | Defaults to the model default; not accepted on kling-3.0 / wan-3.0 image-to-video (follows first_frame), nor on wan-3.0 together with refs (follows the references) |
resolution |
no | Defaults to the model default |
generate_audio |
no | Native audio on the output; defaults to true on every model and never changes the price |
first_frame |
no | Image media object; switches to image-to-video |
end_frame |
no | Image for the last frame; requires first_frame |
refs |
no | Reference media objects, up to refs.max from the models endpoint (9 image/video/audio on Seedance; 7 image/video on kling-3.0-omni; 20 combined on Wan, subject to refs.max_per_kind; not accepted on kling-3.0) |
Image-to-video with first and last frame:
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"model_id": "seedance-2.0", "prompt": "The scene slowly comes alive, gentle camera push-in", "first_frame": {"url": "https://your.site/start.jpg"}, "end_frame": {"url": "https://your.site/end.jpg"}, "duration_sec": 6}' https://openapi.visionstory.ai/api/v1/ai_videoCharacter-consistent generation with a reusable asset:
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"model_id": "seedance-2.0", "prompt": "The same woman walks through a neon-lit street at night", "refs": [{"asset_id": "7241058823145623552"}], "duration_sec": 10}' https://openapi.visionstory.ai/api/v1/ai_videoQuery and poll
Query one task with ?video_id=, or up to 20 at once with ?video_ids=id1,id2,... (batch responses return {"videos": [...]}). Poll every 5–10 seconds.
| status | Meaning |
|---|---|
queued |
Accepted, waiting for a worker |
creating |
Generating |
created |
Done — video_url and cover_url are ready |
failed |
Generation failed — see error; credits were refunded automatically |
Content moderation timing depends on the model. Seedance moderation is asynchronous: a submission that violates content policy is accepted at submit time and later resolves to failed with an explanatory error — always handle the failed state. Kling and Wan models (kling-3.0, kling-3.0-omni, wan-3.0, wan-3.0-prime) check the prompt and any image inputs (first_frame, end_frame, image refs) synchronously at submit and reject with HTTP 403 37110 (prompt) / 37111 (image) before any credits are charged. Failed tasks never consume credits.
List and delete
GET /api/v1/ai_videos lists your tasks, newest first. Pass cursor from the previous page's next_cursor to paginate; next_cursor=0 means no more pages. limit defaults to 20 (max 100). status filters by task status — comma-separated, any of queued, creating, created, failed. DELETE /api/v1/ai_video?video_id= deletes a task — only your own.
status=queued,creating returns exactly the tasks that occupy your per-key concurrency quota, so len(videos) is how many slots you are using. Call it before a burst of submissions to stay under the limit instead of discovering it through 429s:
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/ai_videos?status=queued,creating&limit=100" | jq '.data.videos | length'Assets
An asset is a reusable uploaded material: upload once, reference by asset_id in any number of generation requests — ideal for a recurring character image, brand footage, or a voice sample. Re-uploading identical content returns the existing asset (idempotent):
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"url": "https://your.site/character.jpg"}' https://openapi.visionstory.ai/api/v1/asset{
"data": {
"asset_id": "7241058823145623552",
"kind": "image",
"mime": "image/jpeg",
"width": 1024,
"height": 1536,
"duration_sec": 0,
"created_at": 1754270000
}
}inline_data (base64) is also accepted; size and format limits match Media inputs above. GET /api/v1/assets lists your assets (filter by kind, paginate with cursor / limit). DELETE /api/v1/asset?asset_id= deletes an asset — videos already generated from it are not affected.
Errors
| HTTP | error.code | Meaning |
|---|---|---|
| 401 | 401 | Missing or unrecognised API key — check the X-API-Key header |
| 403 | 42000 | Key is valid, but the account has no active Pro (or above) subscription — renew, the same key resumes working |
| 403 | 403 | Subscription inactive or insufficient — check your plan |
| 403 | 30610 | Active subscription required |
| 403 | 30301 | Insufficient credits |
| 403 | 37101 | Invalid input — parameter or media constraint violation |
| 403 | 37102 | Prompt too long |
| 403 | 37103 | Too many references |
| 403 | 37104 | Reference video or audio exceeds the model's total duration budget (15 s each on Wan) |
| 403 | 37110 | Prompt blocked by content moderation (Kling and Wan models, checked at submit) |
| 403 | 37111 | Image blocked by content moderation (Kling and Wan models, checked at submit) |
| 400 | 400 | Rejected by the gateway (unknown model_id, media too large or unsupported, invalid asset_id) |
| 422 | — | Malformed request body (unknown fields, missing required fields, invalid combinations) |
| 404 | — | Video or asset not found |
| 429 | — | Rate limited |
Asynchronous failures (Seedance moderation, provider errors) never use HTTP errors — the task resolves to status=failed with an error object, and credits are refunded automatically.
Limits and notes
- Concurrency: per-key concurrency is limited during beta; requests beyond the limit are rejected, not queued. The 429 body carries the numbers in
error.details:{"limit": N, "in_flight": M}— readlimitfrom there once and self-throttle, or pollGET /api/v1/ai_videos?status=queued,creatingbefore submitting. - Rate limit: 180 requests per 60 seconds per account. Use batch query (
video_ids) and poll at 5–10 s intervals. - Storage: download
video_urlpromptly if you need long-term storage. Assets stay available until you delete them. - Idempotency: pass a
client_request_idon submit to make retries safe — resubmitting the same value within 24h returns the original task instead of creating (and charging) a new one. Without it, store the returnedvideo_idbefore retrying, since a network-level retry of a successful submit creates (and charges) a new task. - Webhooks are not available yet; polling is the supported integration pattern.
- The beta surface may gain new optional parameters and models over time; existing fields and semantics will not change incompatibly.
Next steps
- Quick start — the talking-avatar flow with the same key.
- Developer tools — SDK, CLI, MCP, and Agent Skill documentation.
- API reference — full schemas for every endpoint above.