Mintplex-Labs/anything-llm · error · Error

Failed to fetch

Error message

Failed to fetch ${url}: ${response.status}

What it means

Identical guard to the Pinecone variant: in QDrant.addDocumentToNamespace's novel-document branch, vectors are only built when the embedding step produced textChunks; an empty chunks array falls to the else branch and throws, intentionally refusing to record a document that contributed no text. The root cause is upstream — the document yielded zero text.

Solutions

  1. Confirm the file contains selectable text; OCR scanned PDFs before uploading.
  2. Verify the file is a supported, non-empty format (pdf, txt, md, docx, html, csv, ...).
  3. Check collector/server logs for the extraction step preceding this throw.
  4. Add a pre-ingestion check that rejects empty extractions with an explicit message instead of letting the vector step fail generically.

Example fix

// before
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });

// after
const text = extractText(doc);
if (!text || text.trim().length === 0) {
  return { success: false, reason: 'No extractable text — scanned PDF or empty file' };
}
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });
Defensive patterns

Strategy: validation

Validate before calling

const text = extractTextFromDocument(doc);
if (!text || text.trim().length === 0) {
  return { success: false, reason: 'No extractable text — scanned PDF or empty file' };
}
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });

Type guard

function hasEmbeddableText(doc) {
  const t = doc?.textContent ?? doc?.text ?? '';
  return typeof t === 'string' && t.trim().length > 0;
}

Try / catch

try {
  await vectorDB.addDocumentToNamespace(namespace, payload);
} catch (e) {
  if (/Could not embed document chunks/.test(e.message)) {
    return markDocumentFailed(doc.id, 'No extractable text — OCR the file or verify it is not empty');
  }
  throw e;
}

Prevention

When it happens

Trigger: Embedding a 0-byte or whitespace-only file; an image-only/scanned PDF with no OCR text layer; a corrupt or unsupported binary that the parser silently reads as empty; a web-crawler page whose text extraction returned nothing.

Common situations: Scanned PDFs uploaded without OCR; blank exports; extraction failures logged earlier in the collector but not surfaced; ingestion scripts that skip extraction-result validation.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18). Data as JSON: /api/errors/f4a63f8c69b0c930. Report an issue: GitHub.

Appendix: source

Thrown at collector/utils/extensions/DrupalWiki/DrupalWiki/index.js:240

  #generateChunkSource(pageId, encryptionWorker) {
    const payload = {
      baseUrl: this.baseUrl,
      pageId: pageId,
      accessToken: this.accessToken,
    };
    return `drupalwiki://${
      this.baseUrl
    }/node/${pageId}?payload=${encryptionWorker.encrypt(
      JSON.stringify(payload)
    )}`;
  }

  async _doFetch(url) {
    const response = await fetch(url, {
      headers: this.#getHeaders(),
    });
    if (!response.ok) {
      throw new Error(`Failed to fetch ${url}: ${response.status}`);
    }
    return response.json();
  }

  #getHeaders() {
    return {
      "Content-Type": "application/json",
      Accept: "application/json",
      Authorization: `Bearer ${this.accessToken}`,
    };
  }

  #prepareStoragePath(baseUrl) {
    const { hostname } = new URL(baseUrl);
    const subFolder = slugify(`drupalwiki-${hostname}`).toLowerCase();
    const outFolder = path.resolve(documentsFolder, subFolder);
    if (!fs.existsSync(outFolder)) fs.mkdirSync(outFolder, { recursive: true });
    return outFolder;

View on GitHub (pinned to 3aec848f28)