Mintplex-Labs/anything-llm · error · Error
Failed to fetch
Error message
Failed to fetch ${url}: ${response.status} What it means
Identical guard to the Pinecone variant: in QDrant.addDocumentToNamespace's novel-document branch, vectors are only built when the embedding step produced textChunks; an empty chunks array falls to the else branch and throws, intentionally refusing to record a document that contributed no text. The root cause is upstream — the document yielded zero text.
Solutions
- Confirm the file contains selectable text; OCR scanned PDFs before uploading.
- Verify the file is a supported, non-empty format (pdf, txt, md, docx, html, csv, ...).
- Check collector/server logs for the extraction step preceding this throw.
- Add a pre-ingestion check that rejects empty extractions with an explicit message instead of letting the vector step fail generically.
Example fix
// before
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });
// after
const text = extractText(doc);
if (!text || text.trim().length === 0) {
return { success: false, reason: 'No extractable text — scanned PDF or empty file' };
}
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata }); Defensive patterns
Strategy: validation
Validate before calling
const text = extractTextFromDocument(doc);
if (!text || text.trim().length === 0) {
return { success: false, reason: 'No extractable text — scanned PDF or empty file' };
}
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata }); Type guard
function hasEmbeddableText(doc) {
const t = doc?.textContent ?? doc?.text ?? '';
return typeof t === 'string' && t.trim().length > 0;
} Try / catch
try {
await vectorDB.addDocumentToNamespace(namespace, payload);
} catch (e) {
if (/Could not embed document chunks/.test(e.message)) {
return markDocumentFailed(doc.id, 'No extractable text — OCR the file or verify it is not empty');
}
throw e;
} Prevention
- Validate extraction output in the collector before the vector step.
- OCR scanned PDFs upstream; reject empty files at upload.
- Surface extraction failures per document so users see which file was skipped.
- Keep ingestion batches resilient: one bad document must not abort the workspace embed.
When it happens
Trigger: Embedding a 0-byte or whitespace-only file; an image-only/scanned PDF with no OCR text layer; a corrupt or unsupported binary that the parser silently reads as empty; a web-crawler page whose text extraction returned nothing.
Common situations: Scanned PDFs uploaded without OCR; blank exports; extraction failures logged earlier in the collector but not surfaced; ingestion scripts that skip extraction-result validation.
Related errors
- Invalid link provided
- Failed to get YouTube video transcription
- FFMPEG binary not found.
- Input file does not exist.
- Could not embed document chunks! This document will not be…
AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18).
Data as JSON: /api/errors/f4a63f8c69b0c930.
Report an issue: GitHub.
Appendix: source
Thrown at collector/utils/extensions/DrupalWiki/DrupalWiki/index.js:240
#generateChunkSource(pageId, encryptionWorker) {
const payload = {
baseUrl: this.baseUrl,
pageId: pageId,
accessToken: this.accessToken,
};
return `drupalwiki://${
this.baseUrl
}/node/${pageId}?payload=${encryptionWorker.encrypt(
JSON.stringify(payload)
)}`;
}
async _doFetch(url) {
const response = await fetch(url, {
headers: this.#getHeaders(),
});
if (!response.ok) {
throw new Error(`Failed to fetch ${url}: ${response.status}`);
}
return response.json();
}
#getHeaders() {
return {
"Content-Type": "application/json",
Accept: "application/json",
Authorization: `Bearer ${this.accessToken}`,
};
}
#prepareStoragePath(baseUrl) {
const { hostname } = new URL(baseUrl);
const subFolder = slugify(`drupalwiki-${hostname}`).toLowerCase();
const outFolder = path.resolve(documentsFolder, subFolder);
if (!fs.existsSync(outFolder)) fs.mkdirSync(outFolder, { recursive: true });
return outFolder;View on GitHub (pinned to 3aec848f28)