Engineering

Async Media Generation APIs: Polling, Timeouts and Retries Done Right

ScalingTensor Team ·

Video and audio generation jobs take seconds to minutes. How to submit them, poll without hammering the API, survive restarts, and avoid paying twice for the same job.

Generating a video or a soundtrack takes anywhere from a few seconds to several minutes. No HTTP request should stay open that long, so media APIs split the work in two: you submit a job and get an ID back immediately, then you ask for the result until it is ready. ScalingTensor works this way for every media model: POST /api/media/generate returns a jobId, and GET /api/media/result?jobId=... reports its status.

The pattern is simple, but the details decide whether it holds up in production.

Know the states

A job is pending while it queues and renders, then ends as completed or failed. Those two are terminal: stop polling as soon as you see either one. Everything else in your code should treat the job as still running.

Poll with backoff, jitter and a deadline

Polling every few seconds is fine for one job. For hundreds, fixed intervals create bursts of identical requests. Start at 2 to 3 seconds, grow the delay each round, add a little randomness, cap it, and stop at a deadline that matches the model: a short sound effect finishes quickly, a 4K video can take many minutes.

javascript
import { setTimeout as sleep } from "node:timers/promises";

const API = "https://restapi.scalingtensor.com";
const headers = { apiKey: process.env.SCALINGTENSOR_API_KEY };

export async function pollJob(jobId, { timeoutMs = 15 * 60_000, signal } = {}) {
  const deadline = Date.now() + timeoutMs;
  let delay = 2_000;
  while (Date.now() < deadline) {
    let body;
    try {
      const response = await fetch(`${API}/api/media/result?jobId=${jobId}`, { headers, signal });
      body = await response.json();
    } catch (error) {
      if (signal?.aborted) throw error;
      body = null; // network blip: keep the job, try again after the delay
    }
    if (body && body.code !== 0) throw new Error(`jobId ${jobId}: ${body.msg}`);
    if (body?.data.status === "completed") return body.data.resultData;
    if (body?.data.status === "failed") throw new Error(`jobId ${jobId} failed`);
    await sleep(delay + Math.random() * 500, undefined, { signal });
    delay = Math.min(Math.round(delay * 1.5), 20_000);
  }
  throw new Error(`jobId ${jobId} still pending after ${timeoutMs / 1000}s`);
}

Note what the loop does with a network error: it keeps waiting instead of failing the job. The job is still running on the server; losing one status response does not change that.

Persist the jobId before anything else

The job exists, and costs money, from the moment the submit call succeeds. If your process crashes between submitting and polling, the only way to recover the result is the jobId. Write it to durable storage immediately:

sql
create table generation_jobs (
  id           bigserial primary key,
  job_id       text unique,          -- ScalingTensor jobId, set right after submit
  model        text not null,
  params       jsonb not null,
  status       text not null default 'submitting',
  output_url   text,                 -- your copy of the result
  created_at   timestamptz default now()
);

A background worker can then pick up every row that is not finished, poll it, and resume after a deploy or a crash without submitting anything twice.

Retry the right thing

  • Polling fails (timeout, 5xx, connection reset): retry the poll. It is read-only and safe.
  • Submitting fails before you get a response: you cannot tell whether the job was created. Check your usage log in the console before resubmitting, or accept the risk of a duplicate for cheap jobs.
  • Submit returns an error code: read msg. Parameter errors will fail again unchanged, so fix the request instead of retrying it.
  • The job ends as failed: log the jobId and parameters, then decide whether a retry with the same input makes sense.

Copy results to your own storage

When a job completes, resultData holds the provider's result with the output URL. Download the file and put it in your own bucket or CDN before you show it to users. Provider URLs are for retrieval, and their lifetime is not something your product should depend on.

Keep keys off the client

It is tempting to poll from the browser so the user sees progress. Do it through your own backend instead: the browser polls your server, your server polls ScalingTensor with the API key. The key never leaves your infrastructure, and one server-side poll can serve many open tabs.

More guides

Try it with your own key

Sign in with Google, top up your wallet from $10 and create an API key. Every model in these guides works with the same key.