How to score a video with music generated to fit it: the request, keeping dialogue audible, changing style across scenes, and what a minute of video costs.
Finding music that fits a video, licensing it and cutting it to length is slow. A video-to-music model does it in one call: it watches the clip and composes a soundtrack that follows its pacing and mood. This guide uses Sonilo Video to Music, which you can call through ScalingTensor with the same API key as every other model.
The simplest request
The only required input is video_url: a public HTTP or HTTPS link to the video. Private or internal addresses are rejected, and the video can be up to 6 minutes long and 300 MB. Uploading the file itself is not supported on the JSON API, so host it first (a signed URL from your storage bucket works).
const response = await fetch("https://restapi.scalingtensor.com/api/media/generate", {
method: "POST",
headers: {
apiKey: process.env.SCALINGTENSOR_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "sonilo/video-to-music",
params: {
video_url: "https://cdn.example.com/launch-teaser.mp4",
},
}),
}).then((r) => r.json());
if (response.code !== 0) throw new Error(response.msg);
console.log(response.data.jobId); // poll /api/media/result with thisLike every media model, the call returns a jobId. Poll GET /api/media/result?jobId=... until the status is completed, then read the generated audio from resultData. The polling guide has a reusable loop.
Steering the result
The optional parameters cover the common production needs:
prompt: a style direction of up to 2,000 characters, such as “warm, acoustic, optimistic”.prompt_influence: 0 to 1, default 0.5. Lower lets the video lead; higher follows the prompt more literally.segments: timed prompts so the music changes at scene boundaries. The first segment starts at 0, and segments must be at least 5 seconds apart.preserve_speech: keep the original dialogue or narration and compose music around it.ducking: also return a track where the music dips under speech, so dialogue stays intelligible.variants_num: 1 to 10 alternative soundtracks in one request, each billed as its own output.stems: split each generated track into drums, bass, vocals and other, for mixing.output_format: m4a by default, or wav or mp3.
A brand film with narration and two moods might send:
{
"model": "sonilo/video-to-music",
"params": {
"video_url": "https://cdn.example.com/brand-film.mp4",
"prompt": "warm, optimistic, acoustic",
"segments": [
{ "start": 0, "prompt": "sparse piano, calm" },
{ "start": 20, "prompt": "full band, building energy" }
],
"preserve_speech": true,
"variants_num": 2
}
}What it costs
Video-to-music is billed by the length of the source video: currently $0.009 per second of source video. A 60-second video costs $0.54. Each extra variant multiplies that. There is no subscription; the charge comes out of your prepaid wallet.
Related audio models
If you need sound effects rather than music, Sonilo Video to Sound Effects adds effects that match what happens on screen. To compose music without a video, use Sonilo Text to Music and set the duration yourself.
| Model | Tasks | Billing | Price |
|---|---|---|---|
| Sonilo Video Dubbing sonilo/dubbing | Per second | $0.0985 per second of source video, multiplied by the number of languages | |
| Sonilo Text to Music sonilo/text-to-music | Per second | $0.00225 per second of output, multiplied by the number of variants_num | |
| Sonilo Text to Sound Effects sonilo/text-to-sfx | Per second | $0.0018 per second of output | |
| Sonilo Video to Music sonilo/video-to-music | Per second | $0.009 per second of source video | |
| Sonilo Video to Sound Effects sonilo/video-to-sfx | Per second | $0.009 per second of source video |