abhigyanpatwari/GitNexus · error

Embedding model not initialized. Run embedding pipeline…

Error message

Embedding model not initialized. Run embedding pipeline first.

What it means

semanticSearch() guards on isEmbedderReady() before embedding the query string. The local transformers.js model is lazily initialized only by the embedding pipeline (or an explicit getEmbedder()); querying before that in this process/state throws immediately rather than triggering an expensive model download on a read path.

Solutions

  1. Run the embedding pipeline first: `npx gitnexus analyze --embeddings` from the repo root.
  2. Verify embeddings exist (index status / vector table) before issuing semantic queries.
  3. Guard call sites with isEmbedderReady() and surface a friendly 'run --embeddings first' message.
  4. In HTTP embedding mode, ensure GITNEXUS_EMBEDDING_URL and GITNEXUS_EMBEDDING_MODEL are set for the querying process too.

Example fix

// before
const results = await semanticSearch(executeQuery, 'auth flow');
// after
import { isEmbedderReady } from '../core/embeddings/embedder';
if (!isEmbedderReady()) {
  throw new Error('Run `npx gitnexus analyze --embeddings` before semantic search.');
}
const results = await semanticSearch(executeQuery, 'auth flow');
Defensive patterns

Strategy: type-guard

Validate before calling

import { isEmbedderReady } from '../core/embeddings/embedder';
const canSearchSemantically = (): boolean => isEmbedderReady();

Type guard

import { isEmbedderReady } from '../core/embeddings/embedder';
if (!isEmbedderReady()) {
  // tell the user to run `npx gitnexus analyze --embeddings`
}

Try / catch

try {
  results = await semanticSearch(executeQuery, query, k);
} catch (e) {
  if (e instanceof Error && e.message.includes('not initialized')) {
    results = await keywordSearch(executeQuery, query); // graceful fallback
  } else throw e;
}

Prevention

When it happens

Trigger: Calling semantic search (semanticSearch, or CLI/MCP semantic query paths over the vector index) before `analyze --embeddings` has run against this repo, or in a fresh process where the embedder singleton was never initialized (e.g. querying an index whose embeddings were built by an earlier run).

Common situations: Fresh clone or fresh .gitnexus index where embeddings were never generated; embedding step was skipped or failed earlier in the run; querying a different working directory than the one that was embedded.

Related errors


AI-assisted analysis of abhigyanpatwari/GitNexus@aac7515d2a (2026-08-20). Data as JSON: /api/errors/957630bed14d0939. Report an issue: GitHub.

Appendix: source

Thrown at gitnexus/src/core/embeddings/embedding-pipeline.ts:1013

      percent: 0,
      error: errorMessage,
    });

    throw error;
  }
};

/**
 * Perform semantic search using the vector index with chunk deduplication
 */
export const semanticSearch = async (
  executeQuery: (cypher: string) => Promise<any[]>,
  query: string,
  k: number = 10,
  maxDistance: number = getVectorMaxDistance(DEFAULT_VECTOR_MAX_DISTANCE),
): Promise<SemanticSearchResult[]> => {
  if (!isEmbedderReady()) {
    throw new Error('Embedding model not initialized. Run embedding pipeline first.');
  }

  const queryEmbedding = await embedText(query);
  const queryVec = embeddingToArray(queryEmbedding);
  const queryVecStr = `[${queryVec.join(',')}]`;

  let bestChunks = new Map<
    string,
    { distance: number; chunkIndex: number; startLine: number; endLine: number }
  >();
  // Query/read path: NEVER spawn a network INSTALL on a user query. If the
  // VECTOR extension was not pre-installed, fall back to exact scan rather than
  // blocking the query on a download (offline-first; see extension-loader.ts
  // "load-only" — used by all serve/MCP query paths).
  if (await loadVectorExtension(undefined, { policy: 'load-only' })) {
    try {
      bestChunks = await collectBestChunks(k, async (fetchLimit) => {
        const vectorQuery = `

View on GitHub (pinned to aac7515d2a)