Task

Text to Video API

Generate a video clip from a written prompt.

Describe the subject, action, setting, camera movement and style, and the model returns a video file. Most models also accept a duration and an aspect ratio, and some can generate a matching audio track.

Text-to-video is the fastest way to prototype a shot. When a specific look or character has to stay consistent, image-to-video or reference-to-video gives more control.

Models
8
Pricing
From about $0.56 per clip
Providers
2

Text to Video models and pricing

Every model below is available with the same API key. Open a model for its full parameter list and request examples.

Text to Video API models
ModelTasksBillingPrice
Dreamina Seedance 2.0
dreamina-seedance-2-0-260128
Provider cost
Provider cost
USD / 1M completion tokens; varies by resolution and video input
Dreamina Seedance 2.0 Fast
dreamina-seedance-2-0-fast-260128
Provider cost
Provider cost
USD / 1M completion tokens; varies by resolution and video input
Dreamina Seedance 2.0 Mini
dreamina-seedance-2-0-mini-260615
Provider cost
Provider cost
USD / 1M completion tokens; varies by resolution and video input
Dreamina Seedance 2.5
dreamina-seedance-2-5-260628
Provider cost
Provider cost
USD / 1M completion tokens; varies by resolution and video input
Kling 3.0 Turbo Pro Text to Video
kwaivgi/kling-video/3.0/turbo/pro/text-to-video
Provider cost
≈ $0.70
per clip, reference price
Kling 3.0 Turbo Standard Text to Video
kwaivgi/kling-video/3.0/turbo/standard/text-to-video
Provider cost
≈ $0.56
per clip, reference price
Kling Video O3 4k Text to Video
kwaivgi/kling-video/o3/4k/text-to-video
Provider cost
≈ $2.10
per clip, reference price
Kling Video V3 4k Text to Video
kwaivgi/kling-video/v3/4k/text-to-video
Provider cost
≈ $2.10
per clip, reference price

Providers for text to video

Frequently asked questions

What is the cheapest Text to Video API?

On ScalingTensor, the lowest-priced text to video model is Kling 3.0 Turbo Standard Text to Video, listed at about $0.56 per clip and billed at the provider's actual cost. Prices across the 8 models are listed in the table above.

Which Text to Video models are available?

ScalingTensor currently offers 8 text to video models: Dreamina Seedance 2.0, Dreamina Seedance 2.0 Fast, Dreamina Seedance 2.0 Mini, Dreamina Seedance 2.5, Kling 3.0 Turbo Pro Text to Video, Kling 3.0 Turbo Standard Text to Video, Kling Video O3 4k Text to Video and Kling Video V3 4k Text to Video.

How do I call a text to video model from my code?

Send POST /api/media/generate with the model ID and its params, authenticated with your apiKey header. The response contains a jobId; poll GET /api/media/result?jobId=... until the status is completed, then download the output from the returned URL.

Start calling Text to Video models

Sign in with Google, top up your wallet from $10 and create an API key. The same key works for every model in the catalog.