{"record":{"id":"1d9b00d4dcfbc933","repo":"Mintplex-Labs/anything-llm","slug":"invalid-link-provided","errorCode":null,"errorMessage":"Invalid link provided","messagePattern":"Invalid link provided","errorType":"validation","errorClass":"Error","httpStatus":200,"severity":"error","filePath":"collector/extensions/resync/index.js","lineNumber":23,"sourceCode":" * value is missing or not a valid URL.\n * @param {string|null} baseUrl\n * @returns {string} the protocol including its trailing colon, eg \"http:\"\n */\nfunction protocolOf(baseUrl) {\n  try {\n    return new URL(baseUrl).protocol;\n  } catch {\n    return \"https:\";\n  }\n}\n\n/**\n * Fetches the content of a raw link. Returns the content as a text string of the link in question.\n * @param {object} data - metadata from document (eg: link)\n * @param {import(\"../../middleware/setDataSigner\").ResponseWithSigner} response\n */\nasync function resyncLink({ link }, response) {\n  if (!link) throw new Error(\"Invalid link provided\");\n  try {\n    const { success, content = null, reason } = await getLinkText(link);\n    if (!success) throw new Error(`Failed to sync link content. ${reason}`);\n    response.status(200).json({ success, content });\n  } catch (e) {\n    console.error(e);\n    response.status(200).json({\n      success: false,\n      content: null,\n    });\n  }\n}\n\n/**\n * Fetches the content of a YouTube link. Returns the content as a text string of the video in question.\n * We offer this as there may be some videos where a transcription could be manually edited after initial scraping\n * but in general - transcriptions often never change.\n * @param {object} data - metadata from document (eg: link)","sourceCodeStart":5,"sourceCodeEnd":41,"githubUrl":"https://github.com/Mintplex-Labs/anything-llm/blob/f92433b4ea0598492a1e6645ea22addbb4dd1287/collector/extensions/resync/index.js#L5-L41","documentation":"In PineconeDB.addDocumentToNamespace, vectors are only built inside a branch that runs when the embedding step produced textChunks. If textChunks is empty the loop never executes and the else branch throws, deliberately skipping the document so no empty/partial record is stored. The real failure happened earlier: the raw document yielded zero text to chunk.","triggerScenarios":"Embedding a document whose parsed textContent is empty or whitespace-only: a 0-byte file, an image-only/scanned PDF with no OCR layer, a corrupt or unsupported binary, a crawler-fetched page whose text extraction returned nothing.","commonSituations":"Uploading scanned PDFs without OCR; empty .txt/.md files saved from a blank page; docx/html exports whose text extractor silently failed; ingestion pipelines that don't pre-check extraction results before handing off to the vector DB.","solutions":["Open the source file and confirm it contains selectable text; re-upload after running OCR if it is a scanned document.","Check the file is not 0 bytes and is a format the collector actually parses (pdf, txt, md, docx, html, csv, etc.).","Inspect server/collector logs for the text-extraction step just before this throw — an extraction crash usually logs there.","Add a pre-flight check that rejects documents with no extractable text before calling addDocumentToNamespace, so the user gets a clear reason instead of this generic error."],"exampleFix":"// before - hand any document straight to the provider\nawait vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });\n\n// after - reject empty extractions first\nconst text = extractText(doc);\nif (!text || text.trim().length === 0) {\n  return { success: false, reason: 'Document has no extractable text (scanned PDF? empty file?)' };\n}\nawait vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });","handlingStrategy":"validation","validationCode":"const text = extractTextFromDocument(doc); // same extraction the collector uses\nif (!text || text.trim().length === 0) {\n  return { success: false, reason: 'Document has no extractable text (scanned PDF or empty file)' };\n}\nawait vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });","typeGuard":"function hasEmbeddableText(doc) {\n  const t = doc?.textContent ?? doc?.text ?? '';\n  return typeof t === 'string' && t.trim().length > 0;\n}","tryCatchPattern":"try {\n  await vectorDB.addDocumentToNamespace(namespace, payload);\n} catch (e) {\n  if (/Could not embed document chunks/.test(e.message)) {\n    // mark this single document failed with a user-actionable reason; continue the batch\n    return markDocumentFailed(doc.id, 'No extractable text — OCR the file or check it is not empty');\n  }\n  throw e;\n}","preventionTips":["Validate extraction results in the ingestion pipeline before handing documents to the vector DB.","OCR scanned/image-only PDFs before upload.","Reject 0-byte uploads at the UI/API boundary.","Treat one bad document as a per-item failure so a batch ingest never aborts entirely."],"tags":["pinecone","embedding","empty-document","data-ingestion","document-parsing"],"backgroundTag":"empty-embedding-input","analyzedSha":"f92433b4ea0598492a1e6645ea22addbb4dd1287","analyzedAt":"2026-08-18T10:02:21.017Z","contentChangedAt":"2026-08-18T10:02:21.017Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}