Mintplex-Labs/anything-llm · error · Error
URL could not be scraped and no content was found.
Error message
URL could not be scraped and no content was found.
What it means
Thrown by executeWebScraping when CollectorApi.getLinkContent(url, captureMode) resolves with success: false — the collector service attempted the fetch and reported that it could not retrieve the page. This is the collector's verdict, not an HTTP exception in the flow itself: unreachable host, blocked/robots-disallowed or blacklisted URL, JS-only page with no static HTML, or an invalid URL scheme.
Solutions
- From the collector's container/host, verify the page is fetchable: curl -A 'Mozilla/5.0' <url>.
- If the page is JS-rendered, there is no static HTML to scrape — use the API of the site or a different URL.
- Check the collector logs for the specific fetch failure reason.
- For internal/private URLs in Docker, use a hostname the collector container can resolve, not 'localhost'.
Defensive patterns
Strategy: fallback
Validate before calling
async function urlFetchableFromServer(url) {
try {
const r = await fetch(url, { method: "HEAD", redirect: "follow" });
return r.ok || r.status === 405; // some servers reject HEAD only
} catch {
return false;
}
} Try / catch
try {
result = await AgentFlows.executeFlow(uuid, variables);
} catch (e) {
if (/URL could not be scraped/.test(e.message)) {
// fallback: fetch raw HTML yourself or route the step to a different URL/source
result = await directFetchFallback(targetUrl);
} else throw e;
} Prevention
- Check target reachability from the collector's network (curl from inside the container) before scheduling scrape flows.
- Prefer target APIs over page scraping for JS-heavy or bot-protected sites.
- Watch collector logs — the flow error intentionally omits the low-level fetch reason.
When it happens
Trigger: Scraping a URL the collector cannot fetch: DNS failure or timeout from the collector container, 403/404 from the target, a page that renders entirely client-side, a URL disallowed by the collector's outbound rules, or a malformed URL string.
Common situations: Docker deployment where the collector has no route to an internal URL; scraping sites that bot-block non-browser clients; localhost/private addresses unreachable from inside the container; target site down; redirect loops.
Related errors
- Document processing API is not online. Document
- Document processing is unavailable. The collector service…
- There was no content to be collected or read.
- URL is required for web scraping
- An error occurred while deleting the model
AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18).
Data as JSON: /api/errors/83d6eb7b508482d1.
Report an issue: GitHub.
Appendix: source
Thrown at server/utils/agentFlows/executors/web-scraping.js:34
`\x1b[43m[AgentFlowToolExecutor]\x1b[0m - executing Web Scraping block`
);
if (!url) {
throw new Error("URL is required for web scraping");
}
const captureMode = captureAs === "querySelector" ? "html" : captureAs;
introspect(`Scraping the content of ${url} as ${captureAs}`);
const { success, content } = await new CollectorApi()
.getLinkContent(url, captureMode)
.then((res) => {
if (captureAs !== "querySelector") return res;
return parseHTMLwithSelector(res.content, config.querySelector, context);
});
if (!success) {
introspect(`Could not scrape ${url}. Cannot use this page's content.`);
throw new Error("URL could not be scraped and no content was found.");
}
introspect(`Successfully scraped content from ${url}`);
if (!content || content?.length === 0) {
introspect("There was no content to be collected or read.");
throw new Error("There was no content to be collected or read.");
}
if (!enableSummarization) {
logger(`Returning raw content as summarization is disabled`);
return content;
}
const tokenCount = new TokenManager(
aibitat.defaultProvider.model
).countFromString(content);
const contextLimit = Provider.contextLimit(
aibitat.defaultProvider.provider,View on GitHub (pinned to 3aec848f28)