Mintplex-Labs/anything-llm · error · Error

Invalid link provided

Error message

Invalid link provided

What it means

In PineconeDB.addDocumentToNamespace, vectors are only built inside a branch that runs when the embedding step produced textChunks. If textChunks is empty the loop never executes and the else branch throws, deliberately skipping the document so no empty/partial record is stored. The real failure happened earlier: the raw document yielded zero text to chunk.

Solutions

  1. Open the source file and confirm it contains selectable text; re-upload after running OCR if it is a scanned document.
  2. Check the file is not 0 bytes and is a format the collector actually parses (pdf, txt, md, docx, html, csv, etc.).
  3. Inspect server/collector logs for the text-extraction step just before this throw — an extraction crash usually logs there.
  4. Add a pre-flight check that rejects documents with no extractable text before calling addDocumentToNamespace, so the user gets a clear reason instead of this generic error.

Example fix

// before - hand any document straight to the provider
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });

// after - reject empty extractions first
const text = extractText(doc);
if (!text || text.trim().length === 0) {
  return { success: false, reason: 'Document has no extractable text (scanned PDF? empty file?)' };
}
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });
Defensive patterns

Strategy: validation

Validate before calling

const text = extractTextFromDocument(doc); // same extraction the collector uses
if (!text || text.trim().length === 0) {
  return { success: false, reason: 'Document has no extractable text (scanned PDF or empty file)' };
}
await vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });

Type guard

function hasEmbeddableText(doc) {
  const t = doc?.textContent ?? doc?.text ?? '';
  return typeof t === 'string' && t.trim().length > 0;
}

Try / catch

try {
  await vectorDB.addDocumentToNamespace(namespace, payload);
} catch (e) {
  if (/Could not embed document chunks/.test(e.message)) {
    // mark this single document failed with a user-actionable reason; continue the batch
    return markDocumentFailed(doc.id, 'No extractable text — OCR the file or check it is not empty');
  }
  throw e;
}

Prevention

When it happens

Trigger: Embedding a document whose parsed textContent is empty or whitespace-only: a 0-byte file, an image-only/scanned PDF with no OCR layer, a corrupt or unsupported binary, a crawler-fetched page whose text extraction returned nothing.

Common situations: Uploading scanned PDFs without OCR; empty .txt/.md files saved from a blank page; docx/html exports whose text extractor silently failed; ingestion pipelines that don't pre-check extraction results before handing off to the vector DB.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@f92433b4ea (2026-08-18). Data as JSON: /api/errors/1d9b00d4dcfbc933. Report an issue: GitHub.

Appendix: source

Thrown at collector/extensions/resync/index.js:23

 * value is missing or not a valid URL.
 * @param {string|null} baseUrl
 * @returns {string} the protocol including its trailing colon, eg "http:"
 */
function protocolOf(baseUrl) {
  try {
    return new URL(baseUrl).protocol;
  } catch {
    return "https:";
  }
}

/**
 * Fetches the content of a raw link. Returns the content as a text string of the link in question.
 * @param {object} data - metadata from document (eg: link)
 * @param {import("../../middleware/setDataSigner").ResponseWithSigner} response
 */
async function resyncLink({ link }, response) {
  if (!link) throw new Error("Invalid link provided");
  try {
    const { success, content = null, reason } = await getLinkText(link);
    if (!success) throw new Error(`Failed to sync link content. ${reason}`);
    response.status(200).json({ success, content });
  } catch (e) {
    console.error(e);
    response.status(200).json({
      success: false,
      content: null,
    });
  }
}

/**
 * Fetches the content of a YouTube link. Returns the content as a text string of the video in question.
 * We offer this as there may be some videos where a transcription could be manually edited after initial scraping
 * but in general - transcriptions often never change.
 * @param {object} data - metadata from document (eg: link)

View on GitHub (pinned to f92433b4ea)