Mintplex-Labs/anything-llm · error · Error

URL could not be scraped and no content was found.

Error message

URL could not be scraped and no content was found.

What it means

Thrown by executeWebScraping when CollectorApi.getLinkContent(url, captureMode) resolves with success: false — the collector service attempted the fetch and reported that it could not retrieve the page. This is the collector's verdict, not an HTTP exception in the flow itself: unreachable host, blocked/robots-disallowed or blacklisted URL, JS-only page with no static HTML, or an invalid URL scheme.

Solutions

  1. From the collector's container/host, verify the page is fetchable: curl -A 'Mozilla/5.0' <url>.
  2. If the page is JS-rendered, there is no static HTML to scrape — use the API of the site or a different URL.
  3. Check the collector logs for the specific fetch failure reason.
  4. For internal/private URLs in Docker, use a hostname the collector container can resolve, not 'localhost'.
Defensive patterns

Strategy: fallback

Validate before calling

async function urlFetchableFromServer(url) {
  try {
    const r = await fetch(url, { method: "HEAD", redirect: "follow" });
    return r.ok || r.status === 405; // some servers reject HEAD only
  } catch {
    return false;
  }
}

Try / catch

try {
  result = await AgentFlows.executeFlow(uuid, variables);
} catch (e) {
  if (/URL could not be scraped/.test(e.message)) {
    // fallback: fetch raw HTML yourself or route the step to a different URL/source
    result = await directFetchFallback(targetUrl);
  } else throw e;
}

Prevention

When it happens

Trigger: Scraping a URL the collector cannot fetch: DNS failure or timeout from the collector container, 403/404 from the target, a page that renders entirely client-side, a URL disallowed by the collector's outbound rules, or a malformed URL string.

Common situations: Docker deployment where the collector has no route to an internal URL; scraping sites that bot-block non-browser clients; localhost/private addresses unreachable from inside the container; target site down; redirect loops.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18). Data as JSON: /api/errors/83d6eb7b508482d1. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/agentFlows/executors/web-scraping.js:34

    `\x1b[43m[AgentFlowToolExecutor]\x1b[0m - executing Web Scraping block`
  );

  if (!url) {
    throw new Error("URL is required for web scraping");
  }

  const captureMode = captureAs === "querySelector" ? "html" : captureAs;
  introspect(`Scraping the content of ${url} as ${captureAs}`);
  const { success, content } = await new CollectorApi()
    .getLinkContent(url, captureMode)
    .then((res) => {
      if (captureAs !== "querySelector") return res;
      return parseHTMLwithSelector(res.content, config.querySelector, context);
    });

  if (!success) {
    introspect(`Could not scrape ${url}. Cannot use this page's content.`);
    throw new Error("URL could not be scraped and no content was found.");
  }

  introspect(`Successfully scraped content from ${url}`);
  if (!content || content?.length === 0) {
    introspect("There was no content to be collected or read.");
    throw new Error("There was no content to be collected or read.");
  }

  if (!enableSummarization) {
    logger(`Returning raw content as summarization is disabled`);
    return content;
  }

  const tokenCount = new TokenManager(
    aibitat.defaultProvider.model
  ).countFromString(content);
  const contextLimit = Provider.contextLimit(
    aibitat.defaultProvider.provider,

View on GitHub (pinned to 3aec848f28)