The YouTube Data API does not return transcripts. captions.download exists, but it only works for videos you own. For anybody else’s video the official answer is: you cannot have it.
The unofficial answer is that YouTube’s own player fetches the caption track over plain HTTP, and for years you could just ask for it yourself. That door is mostly shut now, and the way it is shut is worth understanding before you write any code.
What happens when you roll your own
The classic approach is to pull ytInitialPlayerResponse out of the watch page, read captions.playerCaptionsTracklistRenderer.captionTracks[].baseUrl, and fetch it.
Here is what I measured on 2026-09-21 doing exactly that, from a clean residential IP:
Attempt Resulttimedtext URL taken from the watch page
HTTP 200, Content-Length: 0
timedtext with fmt=json3, fmt=srv3, no fmt
same — 200 and nothing
InnerTube /youtubei/v1/get_transcript (WEB client)
HTTP 400
InnerTube with ANDROID / IOS client contexts
400, or a response with no transcript
Seven variations, no transcript. The 200-with-zero-bytes is the giveaway: this is not an error, it is a refusal. The request is missing a PO Token (Proof of Origin), generated by YouTube’s BotGuard JavaScript. Without it you get a well-formed nothing.
You can run BotGuard in a headless browser to mint tokens. It works, and it costs you a browser per request.
The boring answer that actually works
youtube-transcript-api has solved this already and keeps up with the changes:
from youtube_transcript_api import YouTubeTranscriptApi
api = YouTubeTranscriptApi()
for t in api.list("OrElyY7MFVs"):
print(t.language_code, "generated" if t.is_generated else "human-written")
segments = api.fetch("OrElyY7MFVs")
Enter fullscreen mode Exit fullscreen mode
Two days of my own attempts were worth less than pip install. That is the whole trick, and the rest of this post is about the things that still bite you afterwards.
Bite 1: asking for English
Nearly every transcript tool asks YouTube for English and gives up when there is not one. A Korean video with a perfectly good Korean transcript comes back empty, and you cannot tell whether the video had captions at all.
Prefer a language, do not require one:
def pick_track(tracks, prefer=("en",)):
def key(t):
try:
i = prefer.index(t.language_code)
except ValueError:
i = len(prefer)
# any preferred language first; within that, human-written beats auto
return (0 if i <= 1 else 1, 1 if t.is_generated else 0, i)
return min(tracks, key=key, default=None)
Enter fullscreen mode Exit fullscreen mode
Across 39 long-form videos in a mixed English/Korean sample, preferring instead of requiring took me from a pile of empty rows to 39 transcripts.
Bite 2: the timecode that rounds to sixty
Turning segments into SRT looks trivial:
h, m, s = t // 3600, (t % 3600) // 60, t % 60
ms = round((sec - int(sec)) * 1000)
Enter fullscreen mode Exit fullscreen mode
Feed that 59.9999 and you get 00:00:60,000. That is not a valid timecode, and some players reject the whole file rather than that one cue. Round to milliseconds first, then split:
def ts(sec, vtt=False):
total_ms = int(round(max(0.0, float(sec or 0)) * 1000))
ms, t = total_ms % 1000, total_ms // 1000
return "%02d:%02d:%02d%s%03d" % (
t // 3600, (t % 3600) // 60, t % 60, "." if vtt else ",", ms)
Enter fullscreen mode Exit fullscreen mode
Boundary cases worth a test: 0.9995, 59.9999, 3599.9999, 86399.9999, and a negative.
Bite 3: “translate to 100+ languages”
Several tools advertise this. It comes from YouTube’s own auto-translate, which youtube-transcript-api exposes as Transcript.translate().
I tested it on 2026-09-22, through residential proxies, with three retries and a fresh exit IP on each retry, on four videos:
Video Transcript Translation tozh-Hant
1
ok
blocked
2
ok
blocked
3
ok
target language not offered
4
ok
target language not offered
Transcripts 4/4. Translations 0/4. The listing still advertises translation languages; fetching one gives you nothing. So if a tool promises 100+ languages via YouTube, ask it for one and check what comes back.
Translating it yourself is not hard, and it has a side benefit. Send the segments in batches, keep the order, and the timings never move:
body = "&".join("q=" + quote(s["text"] or ".", safe="") for s in batch)
r = httpx.post(
"https://clients5.google.com/translate_a/t?client=dict-chrome-ex"
"&sl=auto&tl=" + target,
content=body.encode(),
headers={"content-type": "application/x-www-form-urlencoded"})
out = [x[0] for x in r.json()] # one result per q, in order
assert len(out) == len(batch) # ★ never zip mismatched lengths
Enter fullscreen mode Exit fullscreen mode
That assert matters more than it looks. If the API ever returns a different number of results than you sent, a silent zip() shifts every subtitle after that point by one cue — and the file still looks fine until someone watches it.
Because the timings are untouched, you can emit a translated SRT, not just a wall of translated text.
Bite 4: your IP
After about 30 transcript fetches from a home connection, youtube-transcript-api started raising IpBlocked. Not the PO Token wall — ordinary rate limiting. Datacenter IPs are worse: a fresh cloud container is usually blocked on the first call.
Rotate a residential exit per request, and — the part people forget — rotate on retry too. A retry down the same blocked tunnel is just a slower failure.
If you would rather not run any of this
I packaged the above as an Apify Actor: YouTube Transcript Scraper. Give it video URLs, a whole playlist, a channel, or a keyword; it returns full text, timed segments, ready-to-save SRT/VTT, and translation into 55 languages with the timings preserved. transcript_status tells you why when a video has none, and videos whose uploader switched captions off are not charged.
Two related ones, same approach — public pages, no login: YouTube Monitor for new videos on a keyword or channel, and Threads Scraper.
Runnable versions of every snippet here: github.com/XixiSuperMan/threads-youtube-scraper-examples.
All figures above are from runs on 2026-09-21 to 2026-09-23. If YouTube’s auto-translate starts working again, I would genuinely like to know.