Mintplex-Labs/anything-llm · error · Error
There was no content to be collected or read.
Error message
There was no content to be collected or read.
What it means
Thrown by executeWebScraping when the scrape succeeded (success: true) but `content` is null, undefined, or an empty array/string — the page returned a 200 with nothing extractable in the chosen capture mode. Distinct from the previous error: the fetch worked; the extraction came back empty. In querySelector mode this also happens when the selector matches no elements.
Solutions
- Switch captureAs from querySelector to text to see whether any text exists at all.
- If using a selector, pick one present in the raw HTML (view-source), not one produced by client-side JS.
- Try captureAs=html and inspect what the collector actually received.
- For JS-only pages, target the site's underlying API endpoint instead of the rendered page.
Example fix
// before: web scraping block config
{ "captureAs": "querySelector", "querySelector": ".user-name", "url": "https://example.com/profile" }
// JS-rendered page -> Error: There was no content to be collected or read.
// after
{ "captureAs": "text", "url": "https://example.com/profile" } Defensive patterns
Strategy: fallback
Try / catch
try {
content = await scrape(url, { captureAs: "text" });
} catch (e) {
if (/no content to be collected/.test(e.message)) {
// fallback: try the raw HTML and/or alternate selector before giving up
content = await scrape(url, { captureAs: "html", enableSummarization: false });
if (!content) throw e;
} else throw e;
} Prevention
- Choose CSS selectors that exist in the server-rendered HTML (view-source), not the post-JS DOM.
- Test capture modes (text vs html vs querySelector) once per site template before production use.
- Disable summarization while debugging so you can see exactly what the collector extracted.
When it happens
Trigger: captureAs=text on a page whose body has no text nodes; captureAs=querySelector with a selector that matches nothing (content parsed from HTML is empty); binary or unsupported content-type pages; pages that serve an empty shell to non-JS clients.
Common situations: Selector copied from devtools after JS modified the DOM, so the static HTML the collector sees has no such element; SPA root div with no content; PDF/image URL given to a text scrape; site serves empty body to unknown user agents.
Related errors
- URL could not be scraped and no content was found.
- URL is required for web scraping
- API Call failed
- error.message
- Failed to delete flow
AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18).
Data as JSON: /api/errors/4603ef0c97ae849d.
Report an issue: GitHub.
Appendix: source
Thrown at server/utils/agentFlows/executors/web-scraping.js:40
const captureMode = captureAs === "querySelector" ? "html" : captureAs;
introspect(`Scraping the content of ${url} as ${captureAs}`);
const { success, content } = await new CollectorApi()
.getLinkContent(url, captureMode)
.then((res) => {
if (captureAs !== "querySelector") return res;
return parseHTMLwithSelector(res.content, config.querySelector, context);
});
if (!success) {
introspect(`Could not scrape ${url}. Cannot use this page's content.`);
throw new Error("URL could not be scraped and no content was found.");
}
introspect(`Successfully scraped content from ${url}`);
if (!content || content?.length === 0) {
introspect("There was no content to be collected or read.");
throw new Error("There was no content to be collected or read.");
}
if (!enableSummarization) {
logger(`Returning raw content as summarization is disabled`);
return content;
}
const tokenCount = new TokenManager(
aibitat.defaultProvider.model
).countFromString(content);
const contextLimit = Provider.contextLimit(
aibitat.defaultProvider.provider,
aibitat.defaultProvider.model
);
if (tokenCount < contextLimit) {
logger(
`Content within token limit (${tokenCount}/${contextLimit}). Returning raw content.`View on GitHub (pinned to 3aec848f28)