Skip to content

Generate

Same three-step pattern per cloud: submit → poll → fetch. The submit call returns immediately with a job handle; polling loops until the job hits a terminal state; fetch grabs the MP4 bytes.

Official docs verified 2026-08-08

Submit → poll → fetch

Azure OpenAI Sora is currently REST-only (no client.videos.* in the openai package during preview). The Python example uses requests against the Azure OpenAI base URL. api-version=preview is the correct API version during the Sora preview period.

import os, time, requests

endpoint = os.environ["AZURE_OPENAI_ENDPOINT"]              # https://<res>.openai.azure.com
headers  = {"api-key": os.environ["AZURE_OPENAI_API_KEY"],
            "Content-Type": "application/json"}
api      = "preview"

# 1. Submit
create = requests.post(
    f"{endpoint}/openai/v1/video/generations/jobs?api-version={api}",
    headers=headers,
    json={
        "model": os.environ.get("AZURE_VIDEO_DEPLOYMENT", "sora-2"),
        "prompt": "A cat playing piano in a jazz bar.",
        "width":  1280,
        "height": 720,
        "n_seconds": 8,
    },
)
create.raise_for_status()
job_id = create.json()["id"]

# 2. Poll
status = None
while status not in ("succeeded", "failed", "cancelled"):
    time.sleep(5)
    st = requests.get(
        f"{endpoint}/openai/v1/video/generations/jobs/{job_id}?api-version={api}",
        headers=headers,
    ).json()
    status = st.get("status")

if status != "succeeded":
    raise RuntimeError(f"Sora job {job_id} ended in status {status!r}")

# 3. Fetch
generation_id = st["generations"][0]["id"]
mp4 = requests.get(
    f"{endpoint}/openai/v1/video/generations/{generation_id}/content/video?api-version={api}",
    headers=headers,
).content

Vertex Veo runs on the google-genai unified client. generate_videos returns an Operation; poll operations.get(op) until op.done. The finished MP4 comes back as a File object that you download and save.

import os, time
from google import genai
from google.genai import types

client = genai.Client(
    vertexai=True,
    project=os.environ["GOOGLE_CLOUD_PROJECT"],
    location=os.environ.get("GOOGLE_CLOUD_LOCATION", "us-central1"),
)

# 1. Submit
op = client.models.generate_videos(
    model=os.environ.get("VERTEX_VIDEO_MODEL", "veo-3.1-generate-preview"),
    prompt="A cat playing piano in a jazz bar.",
    config=types.GenerateVideosConfig(
        aspect_ratio="16:9",   # 16:9 or 9:16
        resolution="720p",     # 720p | 1080p | 4k  (1080p/4k need duration=8)
        duration_seconds="8",  # "4" | "6" | "8"
    ),
)

# 2. Poll
while not op.done:
    time.sleep(10)
    op = client.operations.get(op)

# 3. Fetch
generated = op.response.generated_videos[0]
client.files.download(file=generated.video)
generated.video.save("output.mp4")

Bedrock Nova Reel is fully async. start_async_invoke returns an invocationArn; the finished MP4 lands in the S3 bucket you passed as outputDataConfig.s3OutputDataConfig.s3Uri under output.mp4. Poll with get_async_invoke.

import json, os, time, boto3

br = boto3.client("bedrock-runtime", region_name=os.environ.get("AWS_REGION", "us-east-1"))
s3 = boto3.client("s3",              region_name=os.environ.get("AWS_REGION", "us-east-1"))
bucket = os.environ["CHIRON_VIDEO_BUCKET"]

# 1. Submit
resp = br.start_async_invoke(
    modelId=os.environ.get("BEDROCK_VIDEO_MODEL", "amazon.nova-reel-v1:1"),
    modelInput={
        "taskType": "TEXT_VIDEO",
        "textToVideoParams": {"text": "A cat playing piano in a jazz bar."},
        "videoGenerationConfig": {
            "durationSeconds": 6,          # 6-second increments up to 120
            "fps": 24,
            "dimension": "1280x720",
            "seed": 0,
        },
    },
    outputDataConfig={"s3OutputDataConfig": {"s3Uri": f"s3://{bucket}"}},
)
arn = resp["invocationArn"]

# 2. Poll
while True:
    inv = br.get_async_invoke(invocationArn=arn)
    if inv["status"] in ("Completed", "Failed"):
        break
    time.sleep(30)
if inv["status"] != "Completed":
    raise RuntimeError(f"Nova Reel {arn} failed: {inv.get('failureMessage')}")

# 3. Fetch — output lands at <s3Uri>/output.mp4
out_uri = inv["outputDataConfig"]["s3OutputDataConfig"]["s3Uri"] + "/output.mp4"
key = out_uri.split(f"s3://{bucket}/", 1)[1]
mp4 = s3.get_object(Bucket=bucket, Key=key)["Body"].read()

The service signature

def submit(prompt: str, *, seconds: int = 6) -> str: ...       # returns opaque job_id
def status(job_id: str) -> dict: ...                            # returns {"status": ..., ...}
def fetch(job_id: str) -> bytes: ...                            # returns MP4 bytes once done

The HTTP surface splits these three the same way — POST /video returns { "job_id": ... }, GET /video/{job_id} returns status, GET /video/{job_id}/content returns video/mp4. This keeps the client-side polling contract identical no matter which cloud is behind it.

Run + verify (local, with your own creds)

# Azure — set AZURE_OPENAI_ENDPOINT + AZURE_OPENAI_API_KEY + AZURE_VIDEO_DEPLOYMENT=sora-2
CHIRON_PROVIDER=azure python -c "from video import submit; print(submit('a bear'))"

Same shape on GCP + AWS. Wall-clock is 1–5 minutes per job.