Mintplex-Labs/anything-llm · error · Error
Failed to fetch documents from Paperless-ngx
Error message
Failed to fetch documents from Paperless-ngx: ${error.message} What it means
The novel-document twin of the cached-path failure: after embedding fresh chunks, the code calls getOrCreateCollection and throws when the result is falsy. Since a genuine 404 from getCollection would reject the promise, a falsy resolve means Qdrant answered with an empty or unexpected payload — collection creation failed server-side (classic cause: not enough free RAM to build a collection) or a client/server version mismatch changed the response shape.
Solutions
- Check collection existence directly: curl $QDRANT_ENDPOINT/collections/<namespace> and read the Qdrant logs for the creation attempt.
- Free or allocate more memory to the Qdrant container (collection creation is memory-hungry), then retry the embed.
- Verify the @qdrant/js-client-rest version is compatible with your Qdrant server series; pin/upgrade so response shapes align.
- If a quota (cloud tier / collection count limit) was hit, raise it or clean unused collections, then re-embed.
Example fix
# before - first embed of a workspace throws 'Failed to create new QDrant collection!'
# after - confirm and pre-create the collection with the engine's dimension
curl -X PUT $QDRANT_ENDPOINT/collections/my-workspace \
-H 'Content-Type: application/json' \
-d '{"vectors":{"size":1536,"distance":"Cosine"}}'
# getOrCreateCollection then takes the exists-branch and skips creation
# also raise qdrant memory limits if creation failed with OOM:
# docker run -m 4g ... qdrant/qdrant Defensive patterns
Strategy: try-catch
Validate before calling
// pre-create the collection with the engine's dimension before first embed
const dim = vectors[0]?.vector?.length;
if (!dim) throw new Error('No vectors to infer dimension from');
const exists = await client.getCollection(namespace).then(() => true).catch(() => false);
if (!exists) {
await client.createCollection(namespace, { vectors: { size: dim, distance: 'Cosine' } });
} Type guard
async function collectionReady(client, namespace) {
const c = await client.getCollection(namespace).catch(() => null);
return c != null && typeof c === 'object';
} Try / catch
try {
await vectorDB.addDocumentToNamespace(namespace, payload);
} catch (e) {
if (/Failed to create new QDrant collection/.test(e.message)) {
// inspect and fix root cause, then retry this single document
const info = await client.getCollection(namespace).catch(() => null);
if (!info) {
await client.createCollection(namespace, { vectors: { size: vectors[0].vector.length, distance: 'Cosine' } });
}
return retryEmbed(doc);
}
throw e;
} Prevention
- Provision collections with a known vector size during workspace/workspace-embedding setup.
- Size Qdrant memory for collection creation (several GB free), especially in containers.
- Pin compatible client/server versions and re-test after upgrading either side.
- Retry first-embed once after fixing the cause — lazy creation failures are usually transient.
When it happens
Trigger: First document embedded into a brand-new workspace collection while Qdrant is memory-constrained; @qdrant/js-client-rest version returning a response the code reads as falsy; race between two concurrent addDocumentToNamespace calls both creating/deleting the same collection; collection creation rejected by a proxy or cloud quota.
Common situations: Qdrant in a small Docker container (default limits) on a busy host; first-ever embed on a fresh deployment failing while later ones succeed; mixing client 1.x with server 1.y across a breaking REST change; Qdrant Cloud free-tier quota exhausted.
Related errors
- Input file does not exist.
- FFMPEG binary not found.
- FFMPEG conversion failed
- Failed to create new AstraDB collection!
- Failed to fetch
AI-assisted analysis of Mintplex-Labs/anything-llm@a145d4d87d (2026-08-18).
Data as JSON: /api/errors/5981a0055b3edf31.
Report an issue: GitHub.
Appendix: source
Thrown at collector/utils/extensions/PaperlessNgx/PaperlessNgxLoader/index.js:82
}
}
console.log(
`Fetched ${documents.length} documents from Paperless-ngx (Pages: ${
page - 1
})`
);
const documentsWithContent = await Promise.all(
documents.map(async (doc) => {
const content = await this.fetchDocumentContent(doc.id);
return { ...doc, content };
})
);
return documentsWithContent.filter((doc) => !!doc.content);
} catch (error) {
throw new Error(
`Failed to fetch documents from Paperless-ngx: ${error.message}`
);
}
}
/**
* Fetches the content of a document from Paperless-ngx
* @param {string} documentId - The ID of the document to fetch
* @returns {Promise<string>} The content of the document
*/
async fetchDocumentContent(documentId) {
try {
const response = await fetch(
`${this.baseUrl}/api/documents/${documentId}/download/`,
{
headers: this.baseHeaders,
}
);View on GitHub (pinned to a145d4d87d)