
Multimodality-to-video generation model by ByteDance/Doubao: takes text + image + video + audio references → video (may include audio). 480P/720P, 4-15s, 24fps.
Pricing: 105 creditsper run (≈ $0.53)(base 1 · adjusted by settings)
Used to describe the generated output.
Image for reference
Video for reference
Audio for reference
Chế độ tạo video — đổi bố cục ô ảnh/tham chiếu
Ready to create
Fill the inputs and click Run.
Click an example to load it into the Playground and try it.
README
Multimodality-to-video generation model by ByteDance/Doubao: takes text + image + video + audio references → video (may include audio). 480P/720P, 4-15s, 24fps.
1 credits per run (≈ $0.01). Credits never expire; top up on the Billing page. Some models are priced by parameters (resolution, duration) — the price updates in real time on the Run button before you run.
Sign up for a CinAPI account, create an API key under API Keys, then call the endpoint with a Bearer key. New accounts get 200 free credits.