Mintplex-Labs/anything-llm · error · Error
FFMPEG binary not found.
Error message
FFMPEG binary not found.
What it means
Qdrant collections require an explicit vector size at creation; unlike Pinecone, the server cannot infer it. AnythingLLQ infers the dimension from the first cached chunk (chunks[0][0]?.vector?.length ?? chunks[0][0]?.values?.length); if that yields null — the cached vector-cache file has chunks with no usable vector field — getOrCreateCollection throws rather than creating a collection with a bogus size. The message directs users to GitHub because it indicates an unexpected cache/embedding state.
Solutions
- Delete the document's cached vector file under server/storage (the vector-cache directory) and re-embed the document from scratch so dimensions are re-inferred from fresh embeddings.
- Verify the embedding engine is actually configured and returns vectors (a missing key or broken engine can produce records without vectors).
- If you changed the embedding engine, clear caches for all affected documents — mixed-engine caches are the top cause.
- If it reproduces with a clean cache and a working engine, file the issue on GitHub as the message suggests, including the engine and document type.
Example fix
# before - stale cache with vector-less records triggers the throw # (re-embed keeps reading server/storage/vector-cache/*.json) # after - purge the cache and re-embed rm -rf server/storage/vector-cache/* # then re-upload / re-embed the document in the workspace # getOrCreateCollection now receives a real dimension from the fresh first chunk
Defensive patterns
Strategy: validation
Validate before calling
// before embedding from cache, confirm the cache can yield a dimension
const cached = JSON.parse(await fs.readFile(vectorCachePath, 'utf8')).flat();
const dim = cached[0]?.vector?.length ?? cached[0]?.values?.length;
if (!dim) {
// stale/foreign cache — purge and re-embed instead of letting collection creation fail
await fs.rm(vectorCachePath);
return reEmbedDocument(doc);
}
await qdrant.getOrCreateCollection(client, namespace, dim); Type guard
function cacheYieldsDimension(chunks) {
const first = chunks?.[0]?.[0];
const len = first?.vector?.length ?? first?.values?.length;
return typeof len === 'number' && len > 0;
} Try / catch
try {
await qdrant.addDocumentToNamespace(namespace, payload);
} catch (e) {
if (/Unable to infer vector dimension/i.test(e.message)) {
// cached vectors unusable: clear this document's cache entry and re-embed once
await purgeVectorCache(docId);
return reEmbedDocument(doc);
}
throw e;
} Prevention
- Clear server/storage vector caches whenever the embedding engine (or its model) changes.
- Verify the embedding engine returns vectors with a one-off embedTextInput check before bulk ingest.
- Treat cache files as disposable derived data — delete on version upgrades.
- If a clean cache still fails, capture the engine, model, and doc type and open the GitHub issue the message requests.
When it happens
Trigger: Re-embedding a document whose cached vector-cache file (created by a previous run or a different embedding engine) contains chunk entries without vector/values fields; embedding engine changed (different dimension) so old cache entries no longer match; a vectorization step returned records without vectors; empty chunks array in the cached path.
Common situations: Switching embedding providers (e.g., Azure to OpenAI, or local to OpenAI) then re-embedding without clearing server/storage/document-processor or vector-cache artifacts; interrupted earlier embed runs leaving half-written cache files; upgrading AnythingLLM versions that changed the cache record shape.
Related errors
- Failed to fetch documents from Paperless-ngx
- Failed to fetch
- Input file does not exist.
- Could not embed document chunks! This document will not be…
- Could not embed document chunks! This document will not be…
AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18).
Data as JSON: /api/errors/368731b481dc3af7.
Report an issue: GitHub.
Appendix: source
Thrown at collector/utils/WhisperProviders/ffmpeg/index.js:52
if (this._ffmpegPath) return this._ffmpegPath;
await patchShellEnvironmentPath();
try {
const which = process.platform === "win32" ? "where" : "which";
const result = execSync(`${which} ffmpeg`, { encoding: "utf8" }).trim();
const candidatePath = result?.split("\n")?.[0]?.trim();
if (!candidatePath) throw new Error("FFMPEG candidate path not found.");
if (!this.isValidFFMPEG(candidatePath))
throw new Error("FFMPEG candidate path is not valid ffmpeg binary.");
this.log(`Found FFMPEG binary at ${candidatePath}`);
this._ffmpegPath = candidatePath;
return this._ffmpegPath;
} catch (error) {
this.log(error.message);
}
throw new Error("FFMPEG binary not found.");
}
/**
* Validates that path points to a valid ffmpeg binary.
* Runs ffmpeg -version command.
*
* @param {string} pathToTest - Path of ffmpeg binary
* @returns {boolean}
*/
isValidFFMPEG(pathToTest) {
try {
if (!pathToTest || !fs.existsSync(pathToTest)) return false;
execSync(`"${pathToTest}" -version`, { encoding: "utf8", stdio: "pipe" });
return true;
} catch {
return false;
}
}View on GitHub (pinned to 3aec848f28)