{"record":{"id":"f4a63f8c69b0c930","repo":"Mintplex-Labs/anything-llm","slug":"failed-to-fetch-url-response-status","errorCode":null,"errorMessage":"Failed to fetch ${url}: ${response.status}","messagePattern":"Failed to fetch (.+?): (.+?)","errorType":"exception","errorClass":"Error","httpStatus":200,"severity":"error","filePath":"collector/utils/extensions/DrupalWiki/DrupalWiki/index.js","lineNumber":240,"sourceCode":"  #generateChunkSource(pageId, encryptionWorker) {\n    const payload = {\n      baseUrl: this.baseUrl,\n      pageId: pageId,\n      accessToken: this.accessToken,\n    };\n    return `drupalwiki://${\n      this.baseUrl\n    }/node/${pageId}?payload=${encryptionWorker.encrypt(\n      JSON.stringify(payload)\n    )}`;\n  }\n\n  async _doFetch(url) {\n    const response = await fetch(url, {\n      headers: this.#getHeaders(),\n    });\n    if (!response.ok) {\n      throw new Error(`Failed to fetch ${url}: ${response.status}`);\n    }\n    return response.json();\n  }\n\n  #getHeaders() {\n    return {\n      \"Content-Type\": \"application/json\",\n      Accept: \"application/json\",\n      Authorization: `Bearer ${this.accessToken}`,\n    };\n  }\n\n  #prepareStoragePath(baseUrl) {\n    const { hostname } = new URL(baseUrl);\n    const subFolder = slugify(`drupalwiki-${hostname}`).toLowerCase();\n    const outFolder = path.resolve(documentsFolder, subFolder);\n    if (!fs.existsSync(outFolder)) fs.mkdirSync(outFolder, { recursive: true });\n    return outFolder;","sourceCodeStart":222,"sourceCodeEnd":258,"githubUrl":"https://github.com/Mintplex-Labs/anything-llm/blob/3aec848f2885144aa8f1e53b9731a04310d5d558/collector/utils/extensions/DrupalWiki/DrupalWiki/index.js#L222-L258","documentation":"Identical guard to the Pinecone variant: in QDrant.addDocumentToNamespace's novel-document branch, vectors are only built when the embedding step produced textChunks; an empty chunks array falls to the else branch and throws, intentionally refusing to record a document that contributed no text. The root cause is upstream — the document yielded zero text.","triggerScenarios":"Embedding a 0-byte or whitespace-only file; an image-only/scanned PDF with no OCR text layer; a corrupt or unsupported binary that the parser silently reads as empty; a web-crawler page whose text extraction returned nothing.","commonSituations":"Scanned PDFs uploaded without OCR; blank exports; extraction failures logged earlier in the collector but not surfaced; ingestion scripts that skip extraction-result validation.","solutions":["Confirm the file contains selectable text; OCR scanned PDFs before uploading.","Verify the file is a supported, non-empty format (pdf, txt, md, docx, html, csv, ...).","Check collector/server logs for the extraction step preceding this throw.","Add a pre-ingestion check that rejects empty extractions with an explicit message instead of letting the vector step fail generically."],"exampleFix":"// before\nawait vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });\n\n// after\nconst text = extractText(doc);\nif (!text || text.trim().length === 0) {\n  return { success: false, reason: 'No extractable text — scanned PDF or empty file' };\n}\nawait vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });","handlingStrategy":"validation","validationCode":"const text = extractTextFromDocument(doc);\nif (!text || text.trim().length === 0) {\n  return { success: false, reason: 'No extractable text — scanned PDF or empty file' };\n}\nawait vectorDB.addDocumentToNamespace(namespace, { ...doc, metadata });","typeGuard":"function hasEmbeddableText(doc) {\n  const t = doc?.textContent ?? doc?.text ?? '';\n  return typeof t === 'string' && t.trim().length > 0;\n}","tryCatchPattern":"try {\n  await vectorDB.addDocumentToNamespace(namespace, payload);\n} catch (e) {\n  if (/Could not embed document chunks/.test(e.message)) {\n    return markDocumentFailed(doc.id, 'No extractable text — OCR the file or verify it is not empty');\n  }\n  throw e;\n}","preventionTips":["Validate extraction output in the collector before the vector step.","OCR scanned PDFs upstream; reject empty files at upload.","Surface extraction failures per document so users see which file was skipped.","Keep ingestion batches resilient: one bad document must not abort the workspace embed."],"tags":["qdrant","embedding","empty-document","data-ingestion","document-parsing"],"backgroundTag":"empty-embedding-input","analyzedSha":"3aec848f2885144aa8f1e53b9731a04310d5d558","analyzedAt":"2026-08-18T10:02:21.017Z","contentChangedAt":"2026-08-18T10:02:21.017Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}