Skip to content
Developer reference

TranscriptLab API

Fetch YouTube captions, transcribe audio, run batch and channel jobs, and search cached transcripts through one REST API.

Base URL: https://api.transcriptlab.dev

01 · Quick start

Make your first request

Create an API key from your account dashboard, keep it in a server-side environment variable, and send it as a Bearer token. X-API-Key is also accepted.

curl -i -X POST \
  -H "Authorization: Bearer $YTGW_KEY" \
  -H "Content-Type: application/json" \
  -d '{"video_id":"dQw4w9WgXcQ","format":"json","lang":"en"}' \
  "https://api.transcriptlab.dev/v1/transcripts"
Show a Node.js example
const response = await fetch('https://api.transcriptlab.dev/v1/transcripts', {
  method: 'POST',
  headers: {
    Authorization: `Bearer ${process.env.YTGW_KEY}`,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({ video_id: 'dQw4w9WgXcQ', format: 'json' }),
});

let result;
if (response.status === 202) {
  // Request the Location URL again until it answers 200; there is no job to poll.
  const location = new URL(response.headers.get('Location'), 'https://api.transcriptlab.dev');
  let poll;
  do {
    await new Promise((resolve) => setTimeout(resolve, 2_000));
    poll = await fetch(location, {
      headers: { Authorization: 'Bearer ' + process.env.YTGW_KEY },
    });
    if (!poll.ok) throw new Error('Transcript read failed: ' + poll.status);
  } while (poll.status === 202);
  result = await poll.json();
} else {
  result = await response.json();
}
console.log(result.transcript.text);
Never embed an API key in browser JavaScript or a public mobile app. Keep it on a server you control.

02 · Authentication

API keys and credentials

Send either Authorization: Bearer ytgw_… or X-API-Key: ytgw_…. If both headers are present, they must contain the same key. Deployments with Clerk configured also accept a valid dashboard session token; API keys are the recommended credential for integrations. Create and manage keys in your account dashboard; the full key is shown only once.

03 · Request flow

Handle both 200 and 202

A YouTube request waits up to deadline_ms. If the transcript is ready, the response is 200. If work continues, the 202 response has a Location header and a body of { video_id, state: "processing" }. Request that URL again every couple of seconds until it answers 200; a 404 means nothing of yours is running, and a 422 means the run failed. There is no job id for a single video.

# A 202 response carries a Location header: the same URL a finished transcript is read from.
# Request it again every couple of seconds. 202 means keep waiting; 200 is the transcript.
curl -sS -i -H "Authorization: Bearer $YTGW_KEY" \
  "https://api.transcriptlab.dev/v1/transcripts/dQw4w9WgXcQ?format=json&lang=en"

Single transcript

A 202 is not a delivered transcript and is not charged as one. The result is billed when delivered.

Batch and channel jobs

These return 202. Poll the job, then use its items_url to inspect each video. These are the only transcripts that have a job.

04 · Transcript requests

Formats and request options

OptionType / defaultDescription
langstring
Default: en
Preferred YouTube caption language (ISO 639-1).
translatestring
Default: —
Translate a YouTube caption track to another ISO 639-1 language. Not supported with engine=stt.
typeany | manual | generated
Default: any
Choose which YouTube caption track to use.
formatjson | text | srt | vtt
Default: json
Choose JSON, plain text, SubRip, or WebVTT output.
metadataboolean
Default: true
Include source metadata in JSON.
alsoarray
Default: —
Embed additional text, SRT, or VTT renderings in a JSON response.
deadline_msinteger
Default: 20,000
How long the request waits for a transcript. Default: 20 seconds. Set 1,000–60,000 ms (1–60 seconds); if it is still processing, the API returns 202.
enginecaptions | stt
Default: captions
captions reads the video's existing YouTube subtitles; stt runs speech-to-text on the audio (billed per minute). If a video has no captions we return an error — we never silently fall back to stt.
diarizeboolean
Default: false
Add speaker labels (Speaker 1, Speaker 2, …) to the transcript. Only works with engine=stt and forces tier=premium, which bills 10 credits per audio minute.
tiernormal | premium | auto
Default: auto
Which speech model runs engine=stt jobs: normal is fast and bills 1 credit per audio minute; premium bills 10 credits per audio minute and is the only tier that can label speakers. auto picks normal unless diarize is set.

YouTube transcript POST bodies accept these fields in JSON. The transcript GET endpoint accepts them as query parameters. For JSON, timestamps are in seconds under transcript.timestamped. Text, SRT, and VTT responses use their corresponding content types.

With engine=captions, language.is_generated distinguishes auto-generated from manual YouTube captions. Speech transcripts instead include metadata.engine="stt" and a detected language. translation is null unless translation was requested, then it reports the target and whether the result was translated. With engine=stt, the speech language is detected from audio; translate is unsupported and type only selects caption tracks.

05 · Endpoint reference

Available API routes

Explore parameters, response schemas, and examples in the interactive Swagger reference. The machine-readable OpenAPI 3.1 document is available too.

YouTube transcripts

Fetch caption tracks, discover languages, search cached transcript text, and browse your transcript library.

POST/v1/transcriptsRequest a transcript. Returns 200 with the transcript or 202 with job and retry URLs.
GET/v1/transcripts/{videoId}Read a published transcript from cache; free once the variant is owned. This route does not start a fetch.
GET/v1/transcripts/{videoId}/discoverList available caption tracks. Discovery is free and plan-quota limited.
POST/v1/transcripts/batchQueue a batch of YouTube video URLs. Accepts an Idempotency-Key header, so a retry never creates a second job.
GET/v1/transcriptsBrowse your delivered transcript library. Pages are consistent: pass next_cursor back as cursor.

Channels and jobs

Queue channel listings and transcription, then poll job state or read item-level results.

POST/v1/channels/transcribeList channel videos or transcribe up to 250 videos per job. Accepts an Idempotency-Key header.
GET/v1/jobsList jobs for the authenticated account. Filter with state=waiting|preparing|processing|done|failed.
GET/v1/jobs/{jobId}Read job status, progress, result URL, and counters.
GET/v1/jobs/{jobId}/itemsRead per-video job results.
DELETE/v1/jobs/{jobId}Cancel a non-terminal job.

Media URLs

Transcribe audio and video files hosted at a public HTTPS URL.

POST/v1/mediaTranscribe a public media-file URL.
GET/v1/media/{mediaId}/transcriptRead a completed media transcript.

06 · Async jobs

Poll, inspect, or cancel work

Job state is the polling field: waiting, preparing, and processing are in progress; done and failed are terminal. The finer-grained status field is also returned. Completed single-result jobs have a result_url; batch and channel jobs expose item results at items_url. DELETE a non-terminal job to cancel it. Each account may have up to 25 open jobs; wait for a job to finish if the API returns a concurrency 429.

07 · Media URLs

Transcribe audio and video from a URL

POST a public HTTPS media-file URL to /v1/media. Speech-to-text accepts audio up to 12 hours and 2,000,000,000 bytes; to transcribe a local file, host it at a URL the API can reach. File uploads are not part of the public API. Direct URLs must resolve to supported public audio/video media; YouTube links use the transcript endpoint instead.

curl -X POST "https://api.transcriptlab.dev/v1/media" \
  -H "Authorization: Bearer $YTGW_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://cdn.example.com/episode.mp3","diarize":true}'

Audio transcription uses explicit STT and supports optional diarization. Translation applies to YouTube captions, not STT. Media transcription is charged at one credit per started audio minute when a transcript is delivered (10 credits per minute for premium tier, which diarize implies).

08 · Webhooks

Verify terminal job events

Batch, channel, and media URL transcription requests accept an HTTPS webhook_url. When the job finishes, TranscriptLab sends job.succeeded, job.partial, job.failed, or job.canceled. Save the webhook_secret from the submission response; it is only returned there.

The X-YTGW-Signature header has the form t=timestamp,v1=signature. Verify the HMAC-SHA256 over timestamp + "." + rawRequestBody using the saved secret before parsing the JSON. Compare in constant time and reject stale timestamps. Return any 2xx to acknowledge; failed deliveries are retried.

// Verify the exact raw request body before parsing it.
const [timestampPart, signaturePart] = request.headers
  .get('X-YTGW-Signature')!.split(',');
const timestamp = timestampPart.slice(2);
const received = signaturePart.slice('v1='.length);
const signedPayload = timestamp + '.' + rawBody;
const expected = createHmac('sha256', webhookSecret)
  .update(signedPayload, 'utf8')
  .digest('hex');
// Compare expected and received with a constant-time comparison;
// reject old timestamps and duplicate event ids.

09 · Errors and pagination

Handle errors and page through lists

Errors use the same JSON shape and include a request ID. Common HTTP statuses include 400 invalid request, 401 unauthorized, 402 insufficient credits, 404 not found, 409 conflict, 413 payload too large, 422 transcript unavailable, and 429 rate limited.

{
  "error": {
    "code": "invalid_request",
    "message": "url or video_id is required",
    "request_id": "req_01JQ8S7X"
  }
}

Responses include X-Request-Id. Transcript reads also report X-Cache, X-Credits-Charged, and usually X-Credits-Remaining. A 429 includes Retry-After. Rate limits depend on the plan; the API does not publish X-RateLimit counters.

List endpoints return has_more. List endpoints also return an opaque next_cursor; pass it back as cursor to get the next page. The transcript library’s cursor keeps the walk on one consistent snapshot. Job and channel submissions accept an Idempotency-Key header: a retry with the same key and body replays the original job.

10 · Credits

Usage and billing

Each transcript variant costs one credit once: the first delivered transcript of a video in a given language and track type is charged and every later delivery of it is free. Audio transcription costs one credit per started minute on the normal tier; tier=premium (and diarize, which implies it) costs ten. Channel listing costs one credit per 100 listed video rows; mode=transcribe also charges for delivered transcript variants not yet owned. Caption discovery and cached transcript queries are free, subject to discovery’s monthly plan quota. A request that returns 202 has not yet delivered a transcript, though channel listing work can be metered as rows are collected. Rate and batch limits depend on the account’s plan. Check the account dashboard for current limits and credit balance.