Mintplex-Labs/anything-llm · error · Error
Invalid link provided
Error message
Invalid link provided
What it means
In PineconeDB.addDocumentToNamespace, vectors are only built inside a branch that runs when the embedding step produced textChunks. If textChunks is empty the loop never executes and the else branch throws, deliberately skipping the document so no empty/partial record is stored. The real failure happened earlier: the raw document yielded zero text to chunk.
Solutions
- Open the source file and confirm it contains selectable text; re-upload after running OCR if it is a scanned document.
- Check the file is not 0 bytes and is a format the collector actually parses (pdf, txt, md, docx, html, csv, etc.).
- Inspect server/collector logs for the text-extraction step just before this throw — an extraction crash usually logs there.
- Add a pre-flight check that rejects documents with no extractable text before calling addDocumentToNamespace, so the user gets a clear reason instead of this generic error.
Example fix
// before - hand any document straight to the provider
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });
// after - reject empty extractions first
const text = extractText(doc);
if (!text || text.trim().length === 0) {
return { success: false, reason: 'Document has no extractable text (scanned PDF? empty file?)' };
}
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata }); Defensive patterns
Strategy: validation
Validate before calling
const text = extractTextFromDocument(doc); // same extraction the collector uses
if (!text || text.trim().length === 0) {
return { success: false, reason: 'Document has no extractable text (scanned PDF or empty file)' };
}
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata }); Type guard
function hasEmbeddableText(doc) {
const t = doc?.textContent ?? doc?.text ?? '';
return typeof t === 'string' && t.trim().length > 0;
} Try / catch
try {
await vectorDB.addDocumentToNamespace(namespace, payload);
} catch (e) {
if (/Could not embed document chunks/.test(e.message)) {
// mark this single document failed with a user-actionable reason; continue the batch
return markDocumentFailed(doc.id, 'No extractable text — OCR the file or check it is not empty');
}
throw e;
} Prevention
- Validate extraction results in the ingestion pipeline before handing documents to the vector DB.
- OCR scanned/image-only PDFs before upload.
- Reject 0-byte uploads at the UI/API boundary.
- Treat one bad document as a per-item failure so a batch ingest never aborts entirely.
When it happens
Trigger: Embedding a document whose parsed textContent is empty or whitespace-only: a 0-byte file, an image-only/scanned PDF with no OCR layer, a corrupt or unsupported binary, a crawler-fetched page whose text extraction returned nothing.
Common situations: Uploading scanned PDFs without OCR; empty .txt/.md files saved from a blank page; docx/html exports whose text extractor silently failed; ingestion pipelines that don't pre-check extraction results before handing off to the vector DB.
Related errors
- Failed to fetch
- Could not embed document chunks! This document will not be…
- Could not embed document chunks! This document will not be…
- Could not embed document chunks! This document will not be…
- Could not embed document chunks! This document will not be…
AI-assisted analysis of Mintplex-Labs/anything-llm@f92433b4ea (2026-08-18).
Data as JSON: /api/errors/1d9b00d4dcfbc933.
Report an issue: GitHub.
Appendix: source
Thrown at collector/extensions/resync/index.js:23
* value is missing or not a valid URL.
* @param {string|null} baseUrl
* @returns {string} the protocol including its trailing colon, eg "http:"
*/
function protocolOf(baseUrl) {
try {
return new URL(baseUrl).protocol;
} catch {
return "https:";
}
}
/**
* Fetches the content of a raw link. Returns the content as a text string of the link in question.
* @param {object} data - metadata from document (eg: link)
* @param {import("../../middleware/setDataSigner").ResponseWithSigner} response
*/
async function resyncLink({ link }, response) {
if (!link) throw new Error("Invalid link provided");
try {
const { success, content = null, reason } = await getLinkText(link);
if (!success) throw new Error(`Failed to sync link content. ${reason}`);
response.status(200).json({ success, content });
} catch (e) {
console.error(e);
response.status(200).json({
success: false,
content: null,
});
}
}
/**
* Fetches the content of a YouTube link. Returns the content as a text string of the video in question.
* We offer this as there may be some videos where a transcription could be manually edited after initial scraping
* but in general - transcriptions often never change.
* @param {object} data - metadata from document (eg: link)View on GitHub (pinned to f92433b4ea)