Avatars
An avatar is the on-screen character that speaks your script in a generated video. Use a curated avatar from the public library, or create a custom avatar from a single photo and reuse it across any number of videos.
Endpoints
| Method | Path | What it does |
|---|---|---|
GET |
/api/v1/avatars |
List your custom avatars, or the public library |
POST |
/api/v1/avatar |
Create a custom avatar from an image |
DELETE |
/api/v1/avatar |
Delete one of your custom avatars |
List avatars
GET /api/v1/avatars lists either your own avatars (default) or the public library (is_public=true), newest first, one page at a time:
| Parameter | Default | Meaning |
|---|---|---|
is_public |
false |
false: your own avatars, with their framing. true: the public avatar library |
limit |
20 |
Avatars per page, 1–100 |
cursor |
0 |
Pass the previous page's next_cursor; next_cursor=0 means no more pages |
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/avatars?is_public=true&limit=20"import requests
headers = {"X-API-Key": "sk-vs-xxxxxxxxxxxxxxxxxxx"}
params = {"is_public": "true", "limit": 20}
while True:
resp_data = requests.get("https://openapi.visionstory.ai/api/v1/avatars", params=params, headers=headers,
timeout=10).json()["data"]
for avatar in resp_data["avatars"]:
print(avatar["avatar_id"], avatar["default_voice_id"])
if not resp_data["next_cursor"]:
break
params["cursor"] = resp_data["next_cursor"]Always pick an avatar_id from this endpoint instead of hardcoding one — the public library evolves over time.
Create a custom avatar
Send one portrait image, get back an avatar_id. Provide the image as either a public HTTPS URL or base64 inline_data — exactly one of the two:
curl -s -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"img_url": "https://your.site/portrait.jpg"}' https://openapi.visionstory.ai/api/v1/avatarWith a local file, base64-encode it into inline_data:
import requests
import base64
headers = {"X-API-Key": "sk-vs-xxxxxxxxxxxxxxxxxxx"}
with open("/path/to/image.jpg", "rb") as f:
encoded = base64.b64encode(f.read()).decode("utf-8")
payload = {"inline_data": {"mime_type": "image/jpg", "data": encoded}}
response = requests.post("https://openapi.visionstory.ai/api/v1/avatar", json=payload, headers=headers)
print(response.json()["data"]["avatar_id"])Store the returned avatar_id — you will pass it to every video request that should feature this avatar.
Photo requirements: use a single, clear, front-facing portrait for best results. Supported formats are JPEG, PNG, WEBP, and HEIC, up to 10 MB.
Use an avatar in a video
Pass the avatar_id when creating a video:
{
"model_id": "vs_character_v4",
"avatar_id": "4321918387609092991",
"text_script": { "text": "Hello!", "voice_id": "Alice" }
}See the Quick start for the full generate-poll-download flow.
Framing
Framing decides which part of the avatar's image fills the video frame. Each of your avatars has one framing per aspect ratio (9:16, 16:9, 1:1), shared with the web editor. A new avatar starts with automatic framing around the face; videos always use the current framing, so POST /api/v1/video takes no framing parameters.
You can change the framing of your own avatars only. Public avatars have no framing in API responses and keep their default image; if you reframe a public avatar in the web editor, API videos use that framing too.
Parameters
| Field | Range | Meaning |
|---|---|---|
zoom |
>= 1 |
1 is the largest area of this aspect ratio that fits in the image — the widest framing, e.g. full body. 2 halves the framed width and height, so the person appears twice as large. |
offset_x |
-1 to 1 |
Horizontal position within the room the framed area can move: -1 touches the left edge, 0 is centered, 1 touches the right edge. |
offset_y |
-1 to 1 |
Vertical position: -1 touches the top edge, 0 is centered, 1 touches the bottom edge. |
Any offset in range keeps the framed area inside the image. An offset has no effect in a direction where the framed area already spans the whole image (for example offset_x at zoom: 1 on a tall photo in 9:16).
Read the current framing
POST /api/v1/avatar and GET /api/v1/avatars return framing for each of your own avatars, always three entries in the order 9:16, 16:9, 1:1 (public avatars have framing: null):
{
"avatar_id": "4321918387609092991",
"framing": [
{ "aspect_ratio": "9:16", "zoom": 1.4286, "offset_x": 0.0, "offset_y": -0.62, "image_url": "https://cdn.visionstory.ai/..." },
{ "aspect_ratio": "16:9", "zoom": 1.0, "offset_x": 0.0, "offset_y": -0.5304, "image_url": null },
{ "aspect_ratio": "1:1", "zoom": 1.0, "offset_x": 0.0, "offset_y": -0.7839, "image_url": null }
]
}zoom, offset_x, and offset_y are null when that aspect ratio cannot be reframed — either the avatar predates framing support, or its source image is too small for that ratio (the framed area would fall below the 150px minimum). null means skip it: sending any framing for it returns 400. Other aspect ratios on the same avatar may still be adjustable, so check each entry rather than the avatar as a whole. Create the avatar again from a larger image to make a null ratio adjustable. image_url is the reference image videos in that aspect ratio currently use.
Change the framing
Send all three values for one aspect ratio. To start from the current framing, copy its values from GET /api/v1/avatars and change what you need:
curl -s -X POST https://openapi.visionstory.ai/api/v1/avatar/framing -H "X-API-Key: $VISIONSTORY_API_KEY" -H "Content-Type: application/json" -d '{"avatar_id": "4321918387609092991", "aspect_ratio": "9:16", "zoom": 1, "offset_x": 0, "offset_y": 0}'The response returns the saved zoom, offset_x, offset_y, and the new image_url. Check the image before generating a video. Later videos in that aspect ratio use the new framing, and so does the web editor.
- The request crops the image and checks that a face is visible, which takes a few seconds.
- Values are rounded to four decimals. Sending back the exact values you read keeps the saved framing.
- Public avatars cannot be reframed through the API (
403, code35011). Reframe them in the web editor, or create your own avatar from a photo.
Typical settings:
| Goal | Request values |
|---|---|
| Full body, centered | "zoom": 1, "offset_x": 0, "offset_y": 0 |
| Upper body, head near the top | "zoom": 1.8, "offset_x": 0, "offset_y": -1 |
Errors
| HTTP | error.code | Meaning |
|---|---|---|
| 400 | 400 | zoom is too large: the framed area would be smaller than 150 pixels (details.reason is zoom_too_large, with max_zoom), or the avatar was created before framing was supported (details.reason is framing_unavailable). details is a JSON string |
| 403 | 35011 | The avatar is a public avatar; its framing cannot be changed through the API |
| 403 | 35012 | No face was detected in the framed area — lower zoom or change the offsets |
| 403 | 60006 | The avatar does not exist or is not available to you |
| 422 | 422 | A field is missing, zoom is below 1, or an offset is outside -1 to 1 |
| 503 | 503 | Temporary failure — retry in a few seconds |
Known limitation: the web editor may show framing cached in your browser instead of the framing saved through the API, and generating from that browser can overwrite it.
Delete an avatar
Deleting only affects your own custom avatars; videos already generated with it are unaffected:
curl -s -X DELETE -H "X-API-Key: $VISIONSTORY_API_KEY" "https://openapi.visionstory.ai/api/v1/avatar?avatar_id=YOUR_AVATAR_ID"Next steps
- Voices — pick or clone the voice your avatar speaks with.
- Quick start — generate your first video.
- API reference — full request and response schemas.