{"record":{"id":"957630bed14d0939","repo":"abhigyanpatwari/GitNexus","slug":"embedding-model-not-initialized-run-embedding-pip","errorCode":null,"errorMessage":"Embedding model not initialized. Run embedding pipeline first.","messagePattern":"Embedding model not initialized\\. Run embedding pipeline first\\.","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"gitnexus/src/core/embeddings/embedding-pipeline.ts","lineNumber":1013,"sourceCode":"      percent: 0,\n      error: errorMessage,\n    });\n\n    throw error;\n  }\n};\n\n/**\n * Perform semantic search using the vector index with chunk deduplication\n */\nexport const semanticSearch = async (\n  executeQuery: (cypher: string) => Promise<any[]>,\n  query: string,\n  k: number = 10,\n  maxDistance: number = getVectorMaxDistance(DEFAULT_VECTOR_MAX_DISTANCE),\n): Promise<SemanticSearchResult[]> => {\n  if (!isEmbedderReady()) {\n    throw new Error('Embedding model not initialized. Run embedding pipeline first.');\n  }\n\n  const queryEmbedding = await embedText(query);\n  const queryVec = embeddingToArray(queryEmbedding);\n  const queryVecStr = `[${queryVec.join(',')}]`;\n\n  let bestChunks = new Map<\n    string,\n    { distance: number; chunkIndex: number; startLine: number; endLine: number }\n  >();\n  // Query/read path: NEVER spawn a network INSTALL on a user query. If the\n  // VECTOR extension was not pre-installed, fall back to exact scan rather than\n  // blocking the query on a download (offline-first; see extension-loader.ts\n  // \"load-only\" — used by all serve/MCP query paths).\n  if (await loadVectorExtension(undefined, { policy: 'load-only' })) {\n    try {\n      bestChunks = await collectBestChunks(k, async (fetchLimit) => {\n        const vectorQuery = `","sourceCodeStart":995,"sourceCodeEnd":1031,"githubUrl":"https://github.com/abhigyanpatwari/GitNexus/blob/aac7515d2a8c50a1f8f923c6fb77218b333560d6/gitnexus/src/core/embeddings/embedding-pipeline.ts#L995-L1031","documentation":"semanticSearch() guards on isEmbedderReady() before embedding the query string. The local transformers.js model is lazily initialized only by the embedding pipeline (or an explicit getEmbedder()); querying before that in this process/state throws immediately rather than triggering an expensive model download on a read path.","triggerScenarios":"Calling semantic search (semanticSearch, or CLI/MCP semantic query paths over the vector index) before `analyze --embeddings` has run against this repo, or in a fresh process where the embedder singleton was never initialized (e.g. querying an index whose embeddings were built by an earlier run).","commonSituations":"Fresh clone or fresh .gitnexus index where embeddings were never generated; embedding step was skipped or failed earlier in the run; querying a different working directory than the one that was embedded.","solutions":["Run the embedding pipeline first: `npx gitnexus analyze --embeddings` from the repo root.","Verify embeddings exist (index status / vector table) before issuing semantic queries.","Guard call sites with isEmbedderReady() and surface a friendly 'run --embeddings first' message.","In HTTP embedding mode, ensure GITNEXUS_EMBEDDING_URL and GITNEXUS_EMBEDDING_MODEL are set for the querying process too."],"exampleFix":"// before\nconst results = await semanticSearch(executeQuery, 'auth flow');\n// after\nimport { isEmbedderReady } from '../core/embeddings/embedder';\nif (!isEmbedderReady()) {\n  throw new Error('Run `npx gitnexus analyze --embeddings` before semantic search.');\n}\nconst results = await semanticSearch(executeQuery, 'auth flow');","handlingStrategy":"type-guard","validationCode":"import { isEmbedderReady } from '../core/embeddings/embedder';\nconst canSearchSemantically = (): boolean => isEmbedderReady();","typeGuard":"import { isEmbedderReady } from '../core/embeddings/embedder';\nif (!isEmbedderReady()) {\n  // tell the user to run `npx gitnexus analyze --embeddings`\n}","tryCatchPattern":"try {\n  results = await semanticSearch(executeQuery, query, k);\n} catch (e) {\n  if (e instanceof Error && e.message.includes('not initialized')) {\n    results = await keywordSearch(executeQuery, query); // graceful fallback\n  } else throw e;\n}","preventionTips":["Always run `analyze --embeddings` before enabling semantic search on a repo.","Check isEmbedderReady() at feature-flag time, not at query time, and disable the semantic entry point in the UI otherwise.","In HTTP mode, export the embedding env vars for every process that queries, not just the one that indexes."],"tags":["embeddings","semantic-search","initialization","lazy-init"],"backgroundTag":"model-not-initialized","analyzedSha":"aac7515d2a8c50a1f8f923c6fb77218b333560d6","analyzedAt":"2026-08-20T23:29:22.980Z","contentChangedAt":"2026-08-20T23:29:22.980Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}