{"record":{"id":"83d6eb7b508482d1","repo":"Mintplex-Labs/anything-llm","slug":"url-could-not-be-scraped-and-no-content-was-found","errorCode":null,"errorMessage":"URL could not be scraped and no content was found.","messagePattern":"URL could not be scraped and no content was found\\.","errorType":"exception","errorClass":"Error","httpStatus":null,"severity":"error","filePath":"server/utils/agentFlows/executors/web-scraping.js","lineNumber":34,"sourceCode":"    `\\x1b[43m[AgentFlowToolExecutor]\\x1b[0m - executing Web Scraping block`\n  );\n\n  if (!url) {\n    throw new Error(\"URL is required for web scraping\");\n  }\n\n  const captureMode = captureAs === \"querySelector\" ? \"html\" : captureAs;\n  introspect(`Scraping the content of ${url} as ${captureAs}`);\n  const { success, content } = await new CollectorApi()\n    .getLinkContent(url, captureMode)\n    .then((res) => {\n      if (captureAs !== \"querySelector\") return res;\n      return parseHTMLwithSelector(res.content, config.querySelector, context);\n    });\n\n  if (!success) {\n    introspect(`Could not scrape ${url}. Cannot use this page's content.`);\n    throw new Error(\"URL could not be scraped and no content was found.\");\n  }\n\n  introspect(`Successfully scraped content from ${url}`);\n  if (!content || content?.length === 0) {\n    introspect(\"There was no content to be collected or read.\");\n    throw new Error(\"There was no content to be collected or read.\");\n  }\n\n  if (!enableSummarization) {\n    logger(`Returning raw content as summarization is disabled`);\n    return content;\n  }\n\n  const tokenCount = new TokenManager(\n    aibitat.defaultProvider.model\n  ).countFromString(content);\n  const contextLimit = Provider.contextLimit(\n    aibitat.defaultProvider.provider,","sourceCodeStart":16,"sourceCodeEnd":52,"githubUrl":"https://github.com/Mintplex-Labs/anything-llm/blob/3aec848f2885144aa8f1e53b9731a04310d5d558/server/utils/agentFlows/executors/web-scraping.js#L16-L52","documentation":"Thrown by executeWebScraping when CollectorApi.getLinkContent(url, captureMode) resolves with success: false — the collector service attempted the fetch and reported that it could not retrieve the page. This is the collector's verdict, not an HTTP exception in the flow itself: unreachable host, blocked/robots-disallowed or blacklisted URL, JS-only page with no static HTML, or an invalid URL scheme.","triggerScenarios":"Scraping a URL the collector cannot fetch: DNS failure or timeout from the collector container, 403/404 from the target, a page that renders entirely client-side, a URL disallowed by the collector's outbound rules, or a malformed URL string.","commonSituations":"Docker deployment where the collector has no route to an internal URL; scraping sites that bot-block non-browser clients; localhost/private addresses unreachable from inside the container; target site down; redirect loops.","solutions":["From the collector's container/host, verify the page is fetchable: curl -A 'Mozilla/5.0' <url>.","If the page is JS-rendered, there is no static HTML to scrape — use the API of the site or a different URL.","Check the collector logs for the specific fetch failure reason.","For internal/private URLs in Docker, use a hostname the collector container can resolve, not 'localhost'."],"exampleFix":null,"handlingStrategy":"fallback","validationCode":"async function urlFetchableFromServer(url) {\n  try {\n    const r = await fetch(url, { method: \"HEAD\", redirect: \"follow\" });\n    return r.ok || r.status === 405; // some servers reject HEAD only\n  } catch {\n    return false;\n  }\n}","typeGuard":null,"tryCatchPattern":"try {\n  result = await AgentFlows.executeFlow(uuid, variables);\n} catch (e) {\n  if (/URL could not be scraped/.test(e.message)) {\n    // fallback: fetch raw HTML yourself or route the step to a different URL/source\n    result = await directFetchFallback(targetUrl);\n  } else throw e;\n}","preventionTips":["Check target reachability from the collector's network (curl from inside the container) before scheduling scrape flows.","Prefer target APIs over page scraping for JS-heavy or bot-protected sites.","Watch collector logs — the flow error intentionally omits the low-level fetch reason."],"tags":["agent-flows","web-scraping","collector","network"],"backgroundTag":"web-scraping-failed","analyzedSha":"3aec848f2885144aa8f1e53b9731a04310d5d558","analyzedAt":"2026-08-18T10:02:21.017Z","contentChangedAt":"2026-08-18T10:02:21.017Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}