VisionStory OpenAPI
Get API key

Generate AI videos with Seedance, Kling and Wan models — text-to-video, image-to-video, and multimodal references — through the same API key and billing you already use for talking avatars.

Beta. These endpoints are in beta and may change. Available to any account with an active paid subscription — no separate allowlisting.

Endpoints

Method Path What it does
GET /api/v1/ai_video/models Machine-readable capability sheet for every model
GET /api/v1/ai_video/cost Exact credit cost of a task before you submit it
POST /api/v1/ai_video Submit a generation task
GET /api/v1/ai_video Query one task (or up to 20 with video_ids)
GET /api/v1/ai_videos List your tasks, newest first
DELETE /api/v1/ai_video Delete a task
POST /api/v1/asset Upload a reusable media asset
GET /api/v1/assets List your assets
DELETE /api/v1/asset Delete an asset

Quick start

Submit a text-to-video task, then poll until it finishes:

curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"model_id": "seedance-2.0", "prompt": "A corgi surfing at sunset, cinematic lighting", "duration_sec": 8, "aspect_ratio": "9:16", "resolution": "1080p"}' https://openapi.visionstory.ai/api/v1/ai_video
{
  "data": {
    "video_id": "7241059991822401536",
    "status": "queued",
    "cost_credit": 64
  }
}

Poll every 5–10 seconds:

curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/ai_video?video_id=7241059991822401536"

Models

GET /api/v1/ai_video/models returns allowed values and defaults for each parameter, plus media constraints. Always drive your integration from this endpoint — new models and parameter values appear there without any API change.

model_id Best for Resolution Duration Capabilities
seedance-2.5 Latest generation, single clips up to 30 s 480p / 720p / 1080p 4–30 s text-to-video, image-to-video
seedance-2.0 Flagship quality, multimodal references 480p / 720p / 1080p 4–15 s text-to-video, image-to-video
seedance-2.0-fast Lower latency and cost 480p / 720p 4–15 s text-to-video, image-to-video
seedance-2.0-mini Lightweight, most economical 480p / 720p 4–15 s text-to-video, image-to-video
kling-3.0 Highest resolution — the only 4K model 720p / 1080p / 4k 5 or 10 s text-to-video, image-to-video
kling-3.0-omni Reference-guided generation (up to 7 image/video refs) 720p / 1080p 3 / 5 / 7 / 10 / 15 s text-to-video
wan-3.0 All-in-one; the only model with audio references and any duration 480p / 720p / 1080p any 2–30 s text-to-video, image-to-video
wan-3.0-prime Same as wan-3.0 at higher quality, 1.5x the price 480p / 720p / 1080p any 2–30 s text-to-video, image-to-video

Native audio is on by default at no extra cost everywhere — the price depends only on model, resolution, and duration. Prompt length is 5000 characters on Seedance, 2500 on Kling, 20000 on Wan. These are hard caps, not targets: for Seedance, BytePlus recommends keeping each prompt to about 500 Chinese characters or 1,000 English words — longer prompts are accepted, but the model tends to pick out the main points and drop details. Seedance models accept aspect ratios 16:9 / 9:16 / 4:3 / 3:4 / 1:1. Kling and Wan differ in several ways — read the values from the models endpoint rather than assuming Seedance behaviour:

Credits and cost

Generation is billed in credits, charged at submit time and automatically refunded in full if generation fails. Cost depends only on model, resolution, and duration (per second). Native audio and reference materials never change the price. Query the cost endpoint before submitting — it applies exactly the same formula as billing:

# seedance-2.0, 8 s at 1080p: 8 x 8 = 64 credits
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/ai_video/cost?model_id=seedance-2.0&duration_sec=8&resolution=1080p"
{
  "data": {
    "credit": 64
  }
}
# kling-3.0, 10 s at 4K: 8 x 10 = 80 credits
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/ai_video/cost?model_id=kling-3.0&duration_sec=10&resolution=4k"

Check your remaining balance with GET /api/v1/billing/credits.

Media inputs

Every media slot (first_frame, end_frame, refs[]) accepts exactly one of three forms:

Form Example Use when
url {"url": "https://your.site/img.jpg"} One-off use; fetched by our servers, not added to your asset library
inline_data {"inline_data": {"mime_type": "image/png", "data": "<base64>"}} One-off use, no public URL available
asset_id {"asset_id": "7241058823145623552"} Reused materials — upload once via the Assets API, reference many times

Limits differ per model — the exact values for each slot are returned in first_frame and refs.image / refs.video / refs.audio of GET /api/v1/ai_video/models. Summary:

Kind Models Max size Formats Constraints
Image Seedance 30 MB jpg, jpeg, png, webp, bmp, tiff, gif 300–6000 px per side, aspect ratio between 1:2.5 and 2.5:1
Image Kling 10 MB jpg, jpeg, png, webp, bmp, tiff, gif at least 300 px per side, aspect ratio between 1:2.5 and 2.5:1
Video Seedance 100 MB mp4, mov 2–15 s, 300–6000 px, 24–60 fps, aspect ratio 1:2.5–2.5:1
Video kling-3.0-omni refs only 100 MB mp4, mov 3–15 s, 720–2160 px per side, aspect ratio 1:2.5–2.5:1
Audio Seedance 15 MB wav, mp3 2–15 s; cannot be the only reference
Image Wan 20 MB jpg, jpeg, png, bmp, webp 240–8000 px per side, aspect ratio between 1:8 and 8:1
Video Wan refs only 100 MB mp4, mov 1–15 s each and 15 s total, 240–4096 px, at least 16 fps
Audio Wan refs only 15 MB wav, mp3 1–15 s each and 15 s total

Create a video

POST /api/v1/ai_video has two modes on one endpoint: provide first_frame (optionally end_frame) for image-to-video, provide refs for reference-guided text-to-video (character consistency, style, motion, or voice references), or provide neither for pure text-to-video. refs and first_frame are mutually exclusive. Unknown fields and unsupported values are rejected — nothing is silently ignored.

Field Required Description
model_id yes See Models above
prompt yes Text prompt, up to 5000 characters (2500 on Kling, 20000 on Wan)
client_request_id no Idempotency key; resubmitting the same value within 24h returns the original task instead of charging again
duration_sec no Defaults to the model default
aspect_ratio no Defaults to the model default; not accepted on kling-3.0 / wan-3.0 image-to-video (follows first_frame), nor on wan-3.0 together with refs (follows the references)
resolution no Defaults to the model default
generate_audio no Native audio on the output; defaults to true on every model and never changes the price
first_frame no Image media object; switches to image-to-video
end_frame no Image for the last frame; requires first_frame
refs no Reference media objects, up to refs.max from the models endpoint (9 image/video/audio on Seedance; 7 image/video on kling-3.0-omni; 20 combined on Wan, subject to refs.max_per_kind; not accepted on kling-3.0)

Image-to-video with first and last frame:

curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"model_id": "seedance-2.0", "prompt": "The scene slowly comes alive, gentle camera push-in", "first_frame": {"url": "https://your.site/start.jpg"}, "end_frame": {"url": "https://your.site/end.jpg"}, "duration_sec": 6}' https://openapi.visionstory.ai/api/v1/ai_video

Character-consistent generation with a reusable asset:

curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"model_id": "seedance-2.0", "prompt": "The same woman walks through a neon-lit street at night", "refs": [{"asset_id": "7241058823145623552"}], "duration_sec": 10}' https://openapi.visionstory.ai/api/v1/ai_video

Query and poll

Query one task with ?video_id=, or up to 20 at once with ?video_ids=id1,id2,... (batch responses return {"videos": [...]}). Poll every 5–10 seconds.

status Meaning
queued Accepted, waiting for a worker
creating Generating
created Done — video_url and cover_url are ready
failed Generation failed — see error; credits were refunded automatically

Content moderation timing depends on the model. Seedance moderation is asynchronous: a submission that violates content policy is accepted at submit time and later resolves to failed with an explanatory error — always handle the failed state. Kling and Wan models (kling-3.0, kling-3.0-omni, wan-3.0, wan-3.0-prime) check the prompt and any image inputs (first_frame, end_frame, image refs) synchronously at submit and reject with HTTP 403 37110 (prompt) / 37111 (image) before any credits are charged. Failed tasks never consume credits.

List and delete

GET /api/v1/ai_videos lists your tasks, newest first. Pass cursor from the previous page's next_cursor to paginate; next_cursor=0 means no more pages. limit defaults to 20 (max 100). status filters by task status — comma-separated, any of queued, creating, created, failed. DELETE /api/v1/ai_video?video_id= deletes a task — only your own.

status=queued,creating returns exactly the tasks that occupy your per-key concurrency quota, so len(videos) is how many slots you are using. Call it before a burst of submissions to stay under the limit instead of discovering it through 429s:

curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/ai_videos?status=queued,creating&limit=100" | jq '.data.videos | length'

Assets

An asset is a reusable uploaded material: upload once, reference by asset_id in any number of generation requests — ideal for a recurring character image, brand footage, or a voice sample. Re-uploading identical content returns the existing asset (idempotent):

curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"url": "https://your.site/character.jpg"}' https://openapi.visionstory.ai/api/v1/asset
{
  "data": {
    "asset_id": "7241058823145623552",
    "kind": "image",
    "mime": "image/jpeg",
    "width": 1024,
    "height": 1536,
    "duration_sec": 0,
    "created_at": 1754270000
  }
}

inline_data (base64) is also accepted; size and format limits match Media inputs above. GET /api/v1/assets lists your assets (filter by kind, paginate with cursor / limit). DELETE /api/v1/asset?asset_id= deletes an asset — videos already generated from it are not affected.

Errors

HTTP error.code Meaning
401 401 Missing or unrecognised API key — check the X-API-Key header
403 42000 Key is valid, but the account has no active Pro (or above) subscription — renew, the same key resumes working
403 403 Subscription inactive or insufficient — check your plan
403 30610 Active subscription required
403 30301 Insufficient credits
403 37101 Invalid input — parameter or media constraint violation
403 37102 Prompt too long
403 37103 Too many references
403 37104 Reference video or audio exceeds the model's total duration budget (15 s each on Wan)
403 37110 Prompt blocked by content moderation (Kling and Wan models, checked at submit)
403 37111 Image blocked by content moderation (Kling and Wan models, checked at submit)
400 400 Rejected by the gateway (unknown model_id, media too large or unsupported, invalid asset_id)
422 — Malformed request body (unknown fields, missing required fields, invalid combinations)
404 — Video or asset not found
429 — Rate limited

Asynchronous failures (Seedance moderation, provider errors) never use HTTP errors — the task resolves to status=failed with an error object, and credits are refunded automatically.

Limits and notes

Next steps