{"record":{"id":"368731b481dc3af7","repo":"Mintplex-Labs/anything-llm","slug":"ffmpeg-binary-not-found","errorCode":null,"errorMessage":"FFMPEG binary not found.","messagePattern":"FFMPEG binary not found\\.","errorType":"exception","errorClass":"Error","httpStatus":200,"severity":"error","filePath":"collector/utils/WhisperProviders/ffmpeg/index.js","lineNumber":52,"sourceCode":"    if (this._ffmpegPath) return this._ffmpegPath;\n    await patchShellEnvironmentPath();\n\n    try {\n      const which = process.platform === \"win32\" ? \"where\" : \"which\";\n      const result = execSync(`${which} ffmpeg`, { encoding: \"utf8\" }).trim();\n      const candidatePath = result?.split(\"\\n\")?.[0]?.trim();\n      if (!candidatePath) throw new Error(\"FFMPEG candidate path not found.\");\n      if (!this.isValidFFMPEG(candidatePath))\n        throw new Error(\"FFMPEG candidate path is not valid ffmpeg binary.\");\n\n      this.log(`Found FFMPEG binary at ${candidatePath}`);\n      this._ffmpegPath = candidatePath;\n      return this._ffmpegPath;\n    } catch (error) {\n      this.log(error.message);\n    }\n\n    throw new Error(\"FFMPEG binary not found.\");\n  }\n\n  /**\n   * Validates that path points to a valid ffmpeg binary.\n   * Runs ffmpeg -version command.\n   *\n   * @param {string} pathToTest - Path of ffmpeg binary\n   * @returns {boolean}\n   */\n  isValidFFMPEG(pathToTest) {\n    try {\n      if (!pathToTest || !fs.existsSync(pathToTest)) return false;\n      execSync(`\"${pathToTest}\" -version`, { encoding: \"utf8\", stdio: \"pipe\" });\n      return true;\n    } catch {\n      return false;\n    }\n  }","sourceCodeStart":34,"sourceCodeEnd":70,"githubUrl":"https://github.com/Mintplex-Labs/anything-llm/blob/3aec848f2885144aa8f1e53b9731a04310d5d558/collector/utils/WhisperProviders/ffmpeg/index.js#L34-L70","documentation":"Qdrant collections require an explicit vector size at creation; unlike Pinecone, the server cannot infer it. AnythingLLQ infers the dimension from the first cached chunk (chunks[0][0]?.vector?.length ?? chunks[0][0]?.values?.length); if that yields null — the cached vector-cache file has chunks with no usable vector field — getOrCreateCollection throws rather than creating a collection with a bogus size. The message directs users to GitHub because it indicates an unexpected cache/embedding state.","triggerScenarios":"Re-embedding a document whose cached vector-cache file (created by a previous run or a different embedding engine) contains chunk entries without vector/values fields; embedding engine changed (different dimension) so old cache entries no longer match; a vectorization step returned records without vectors; empty chunks array in the cached path.","commonSituations":"Switching embedding providers (e.g., Azure to OpenAI, or local to OpenAI) then re-embedding without clearing server/storage/document-processor or vector-cache artifacts; interrupted earlier embed runs leaving half-written cache files; upgrading AnythingLLM versions that changed the cache record shape.","solutions":["Delete the document's cached vector file under server/storage (the vector-cache directory) and re-embed the document from scratch so dimensions are re-inferred from fresh embeddings.","Verify the embedding engine is actually configured and returns vectors (a missing key or broken engine can produce records without vectors).","If you changed the embedding engine, clear caches for all affected documents — mixed-engine caches are the top cause.","If it reproduces with a clean cache and a working engine, file the issue on GitHub as the message suggests, including the engine and document type."],"exampleFix":"# before - stale cache with vector-less records triggers the throw\n# (re-embed keeps reading server/storage/vector-cache/*.json)\n\n# after - purge the cache and re-embed\nrm -rf server/storage/vector-cache/*\n# then re-upload / re-embed the document in the workspace\n# getOrCreateCollection now receives a real dimension from the fresh first chunk","handlingStrategy":"validation","validationCode":"// before embedding from cache, confirm the cache can yield a dimension\nconst cached = JSON.parse(await fs.readFile(vectorCachePath, 'utf8')).flat();\nconst dim = cached[0]?.vector?.length ?? cached[0]?.values?.length;\nif (!dim) {\n  // stale/foreign cache — purge and re-embed instead of letting collection creation fail\n  await fs.rm(vectorCachePath);\n  return reEmbedDocument(doc);\n}\nawait qdrant.getOrCreateCollection(client, namespace, dim);","typeGuard":"function cacheYieldsDimension(chunks) {\n  const first = chunks?.[0]?.[0];\n  const len = first?.vector?.length ?? first?.values?.length;\n  return typeof len === 'number' && len > 0;\n}","tryCatchPattern":"try {\n  await qdrant.addDocumentToNamespace(namespace, payload);\n} catch (e) {\n  if (/Unable to infer vector dimension/i.test(e.message)) {\n    // cached vectors unusable: clear this document's cache entry and re-embed once\n    await purgeVectorCache(docId);\n    return reEmbedDocument(doc);\n  }\n  throw e;\n}","preventionTips":["Clear server/storage vector caches whenever the embedding engine (or its model) changes.","Verify the embedding engine returns vectors with a one-off embedTextInput check before bulk ingest.","Treat cache files as disposable derived data — delete on version upgrades.","If a clean cache still fails, capture the engine, model, and doc type and open the GitHub issue the message requests."],"tags":["qdrant","embedding","vector-dimension","cache-invalidation","collection-creation"],"backgroundTag":"missing-vector-dimension","analyzedSha":"3aec848f2885144aa8f1e53b9731a04310d5d558","analyzedAt":"2026-08-18T10:02:21.017Z","contentChangedAt":"2026-08-18T10:02:21.017Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}