ErrLookup › Background articles › ExtractorError in yt-dlp and youtube-dl: what 'said:' messages, login-required errors, and 'Cannot parse data' mean
ExtractorError in yt-dlp and youtube-dl: what 'said:' messages, login-required errors, and 'Cannot parse data' mean
ExtractorError is what yt-dlp and youtube-dl raise when a per-site extractor cannot turn a page or API response into downloadable media: the content needs login, is geo-blocked or removed, the site changed its markup or API since your build was released, or your requests tripped rate limiting. The message often quotes the provider's own text ('YouTube said:', 'Tencent said:', 'This video is only available for registered users', 'Cannot parse data'), which names the real cause. This article covers the mechanisms behind the family's 1137 documented records and the fixes that hold across all of them.
Distilled from 1,137 documented records across 2 repositories.
Background
Both tools are built from per-site modules called extractors, and ExtractorError is the exception class an extractor raises when it cannot produce playable formats or metadata from what the site returned. From the caller's side it is the point where a job stops for one URL: the CLI prints the message and either aborts the extraction (a broken Audiomack album tag kills the whole album run; a broken YouTube comment reply thread aborts the comment download) or skips it when --ignoreerrors is set. The message text comes from two places. Extractor-written strings describe what could not be found: 'Failed to find IDEC id', 'Unable to extract brightcove Video ID', 'No chapters found for course'. Server-forwarded strings quote the provider verbatim: 'LBRY said:', 'Tencent said:', 'NRK said:', YouTube alert text such as 'This video is private', the BBC player's on-page error banner, or the abstract attribute of a thePlatform SMIL fault document.
The class carries an expected flag that changes how the failure should be read. expected=True marks normal site refusals (login walls, geo-restriction, expired rights, unknown ids, risk-control replies), while unmarked raises such as FanCode's missing brightcove id, Facebook's 'Cannot parse data', or UkColumn's 'No embedded video found' usually mean the extractor's assumptions no longer match the live site, which reads as a bug or a staleness problem. Dedicated helpers generate whole sub-families: raise_login_required produces the default 'This video is only available for registered users' plus a login hint, and raise_geo_restricted fires before the generic path on several sites. Several raises also degrade to warnings instead: with ignoreerrors, with --ignore-no-formats-error on metadata-only runs, or when a helper is called with fatal=False as a probe.
Where the failure sits varies by how the site serves data. API-envelope sites (Bilibili, LBRY, Tencent, Polsat Go, NRK, Epicon) answer HTTP 200 with an internal error code and message that the extractor re-raises, so the server's own text reaches you. Scraping extractors (Facebook, LearningOnScreen, CeskaTelevize, Nowness) instead fail with generic 'cannot find/parse' messages when page markup or embedded JSON drifts from what the code traverses. The family spans two repositories: ytdl-org/youtube-dl, the original project, and yt-dlp, its actively maintained fork where extractor fixes are expected to land (Facebook 'markup changes get fixed', Bilibili WBI keys 'change frequently'). That split is itself a cause: an outdated build, or the original youtube-dl where fixes never arrived, reproduces errors current yt-dlp no longer has.
The forwarded messages also tell you whether a failure is retryable. Bilibili risk-control replies gain 'please wait and try later'; YouTube intermittently truncates comment reply-thread continuations; api.lbry.tv has known transient failures. Others are permanent for that URL: NRK programs whose streaming rights lapsed, abandoned LBRY claims, private or deleted YouTube videos, delisted Nintendo Directs. Reading the text before retrying is the core skill this family teaches.
Common causes
- Login-gated content.The content needs an authenticated session and none was supplied. This covers raise_login_required's default message, Instagram's login-gated posts, members-only, subscriber-only, and VIP content, and age-restricted videos. Passing fresh cookies resolves most of these.
- Outdated extractor vs changed site.The site changed its markup, JSON paths, or API contract after your build was released: Facebook's 'Cannot parse data', FanCode's missing brightcove path, CeskaTelevize's moved IDEC fields, Nintendo's drifted GraphQL persisted-query hash, Bilibili's rotated WBI signing keys, Videa's rotated player keys. Updating yt-dlp is the fix.
- Geo-restriction.The provider refuses your region: BBC location notices, NRK's ProgramIsGeoBlocked, LeTV's geo flag, Bilibili region locks, Tencent geo-copyright messages. Retrying from an allowed region via proxy or VPN is the only remedy.
- Removed, private, or expired content.The item is gone: YouTube private or deleted alerts, abandoned LBRY claims, NRK programs whose rights lapsed, delisted Nintendo Directs, deleted videos on Videa, LeTV, and Tencent. These are permanent for that URL; no client-side fix exists.
- Rate limiting and bot detection.Anonymous or aggressive use tripped protection: Bilibili risk-control codes and HTTP 412, YouTube's 'Sign in to confirm you're not a bot' alerts, Instagram's anonymous rate-limit redirect to a login page. Backing off, slowing down, and authenticating usually clear it.
- Wrong or malformed URL or id.The input does not identify real content: a bad Audiomack album tag, a mistyped Nintendo slug, a Safari chapter id used as a course id, a non-canonical Uplynk URL missing its 32-hex id or .m3u8 suffix, a nonexistent Audius handle. Verify the URL opens in a browser and use canonical forms.
- Transient server-side failures.The provider briefly served bad or incomplete data: YouTube's truncated comment reply threads, api.lbry.tv flakiness, occasional Nintendo GraphQL resolver errors. Retrying after a wait succeeds where immediate retries do not.
- Failed login attempt.Credentials were supplied but rejected: Soop/AfreecaTV login codes for wrong passwords, suspended or blocked accounts, a Global account used on Korean Soop, or a required browser-side security check. Completing the check in a browser, then using --cookies-from-browser, bypasses the API login.
What usually fixes it
- Update before debugging: run yt-dlp -U or pip install -U yt-dlp. Extractor fixes for schema and markup changes land fast, and on the original youtube-dl most fixes only exist in the yt-dlp fork, so a current build already resolves a large share of the family.
- Authenticate up front for gated content: --cookies-from-browser BROWSER or a fresh --cookies file, or -u/-p on sites with login support. Re-export cookies on a schedule, because stale sessions reproduce login-required errors silently.
- Read the embedded server text first: 'X said:' messages, YouTube playability reasons, and scraped player banners are the provider's own explanation, and they tell you which remedy applies (sign in, change region, slow down, or accept that the content is gone).
- Verify the URL in a browser, both incognito and logged in, to separate dead or gated content from an extractor bug; prefer canonical URLs copied from the address bar over reconstructed or share links.
- Match the remedy to the failure class: proxy or VPN through an allowed region for geo blocks; back off several minutes and slow requests (--sleep-requests) for risk control; never hammer endpoints whose message says 'please wait and try later'.
- Guard batch jobs: set --ignoreerrors or 'ignoreerrors': True, and classify outcomes as permanent versus retryable so one dead item does not abort the queue.
Go deeper
- Authentication and authorization failures — expired tokens, bad credentials, and missing scopes.
- HTTP status errors: handling 4xx and 5xx responses — how to handle 4xx and 5xx responses properly.
Documented occurrences
- Incomplete data received for comment reply thread. Pass --ignore-errors to ignore and allow rest of comments to download.(yt-dlp/yt-dlp)
- Invalid url for track %d of album url %s(ytdl-org/youtube-dl)
- Instagram sent an empty media response. Check if this post is accessible in your browser without being logged-in. If it is not, then u{self._login_hint()[1:]}. Otherwise, if the post is accessible in browser without being logged-in{bug_reports_message(before=",")}(yt-dlp/yt-dlp)
- GraphQL API error: {errors or "Unknown error"}(yt-dlp/yt-dlp)
- No video found(yt-dlp/yt-dlp)
- This video is only available for registered users(yt-dlp/yt-dlp)
- Unable to extract brightcove Video ID(yt-dlp/yt-dlp)
- reason (dynamic YouTube playability status reason, with optional subreason/captcha suffix)(ytdl-org/youtube-dl)
- Could not find player definition(yt-dlp/yt-dlp)
- Failed to find IDEC id(yt-dlp/yt-dlp)
- Failed to extract content(yt-dlp/yt-dlp)
- Necessary parameters not found in Uplynk URL(yt-dlp/yt-dlp)
- Request failed ({status_code}): {message or "Unknown error"}(yt-dlp/yt-dlp)
- Unable to download video info: {code}: {api_message}(yt-dlp/yt-dlp)
- YouTube said: {errors[-1][1]}(yt-dlp/yt-dlp)
- traverse_obj(smil, (f'{ns}ref/@abstract', ..., any))(yt-dlp/yt-dlp)
- {self.IE_NAME} said: {err.get("code")} - {err.get("message")}(yt-dlp/yt-dlp)
- {self.IE_NAME} said: {message}(yt-dlp/yt-dlp)
- Generic error. flag = %d(yt-dlp/yt-dlp)
- Unable to login: {self.IE_NAME} said: {error}(yt-dlp/yt-dlp)
…and 1,117 more across the corpus — use search.
Honest provenance: generated on 2026-08-22 from AI-assisted analysis of the linked records. See how records are made.