janhq/jan · error · Error
Failed to count embedding tokens
Error message
Failed to count embedding tokens: ${e instanceof Error ? e.message : String(e)} What it means
splitChunkToFit counts tokens per chunk via llm.countEmbeddingTokens(); if that call throws, the error is rewrapped as 'Failed to count embedding tokens' with the original message. Token counting is required to decide whether a chunk fits the embedding context budget, so a failure here aborts chunking.
Solutions
- Read the appended underlying message to find the actual failure (auth, network, unsupported method).
- Verify the embedding provider is reachable and credentials are valid.
- Retry ingestion after transient network/rate-limit failures resolve.
- If the engine cannot count tokens, switch to an engine/model that implements countEmbeddingTokens or estimate tokens heuristically before ingestion.
Example fix
// before
await ensureChunksFitEmbeddingContext(llm, text) // throws on provider outage
// after
try {
await ensureChunksFitEmbeddingContext(llm, text)
} catch (e) {
if (/Failed to count embedding tokens/.test(String(e))) {
await new Promise(r => setTimeout(r, 2000)) // backoff, then retry
await ensureChunksFitEmbeddingContext(llm, text)
} else throw e
} Defensive patterns
Strategy: retry
Validate before calling
if (typeof llm.countEmbeddingTokens !== 'function') {
throw new Error('embedding engine does not implement countEmbeddingTokens')
} Type guard
function canCountTokens(e: unknown): e is EmbeddingEngine & { countEmbeddingTokens(texts: string[]): Promise<[number]> } {
return typeof (e as any)?.countEmbeddingTokens === 'function'
} Try / catch
try {
const chunks = await ensureChunksFitEmbeddingContext(llm, text)
} catch (e) {
if (/rate limit|timeout|network/i.test(String(e))) {
await backoffRetry(() => ensureChunksFitEmbeddingContext(llm, text), 3)
} else throw e
} Prevention
- Prefer local tokenizers or engines with guaranteed countEmbeddingTokens support.
- Wrap provider calls with retry/backoff for transient network and rate-limit errors.
- Verify credentials and endpoint reachability before batch ingestion of large documents.
When it happens
Trigger: ensureChunksFitEmbeddingContext or a recursive splitChunkToFit call invokes llm.countEmbeddingTokens([text]) and the embedding engine throws — provider unreachable, auth failure, or a broken/unimplemented tokenizer in the engine.
Common situations: Remote embedding API outage or rate limit during ingestion of a large document, invalid provider credentials, or an engine whose countEmbeddingTokens is not supported for the selected model.
Understand the failure class
Background: "API request failed": what wrapped HTTP errors from external APIs mean and how to find the real cause — this error's family across 29 libraries.
Related errors
- Failed to determine embedding context size
- Embedding dimension not available
- Invalid metadata: embedding_length not found or invalid
- llamacpp extension not available
- Tokenize request failed with status
AI-assisted analysis of janhq/jan@7205d770c1 (2026-09-17).
Data as JSON: /api/errors/d15da87a21a2ddec.
Report an issue: GitHub.
Appendix: source
Thrown at extensions/vector-db-extension/src/index.ts:227
return await llm.getEmbeddingContextSize()
} catch (e) {
throw new Error(
`Failed to determine embedding context size: ${e instanceof Error ? e.message : String(e)}`
)
}
}
private async splitChunkToFit(
text: string,
budget: number,
llm: EmbeddingEngine
): Promise<string[]> {
if (!text) return []
let count: number
try {
;[count] = await llm.countEmbeddingTokens([text])
} catch (e) {
throw new Error(
`Failed to count embedding tokens: ${e instanceof Error ? e.message : String(e)}`
)
}
if (count <= budget || text.length <= MIN_CHUNK_SIZE_CHARS) return [text]
const mid = Math.floor(text.length / 2)
return [
...(await this.splitChunkToFit(text.slice(0, mid), budget, llm)),
...(await this.splitChunkToFit(text.slice(mid), budget, llm)),
]
}
async ingestFile(threadId: string, file: VectorDBFileInput, opts: VectorDBIngestOptions): Promise<AttachmentFileInfo> {
// Check for duplicate file (same name + path)
const existingFiles = await vecdb.listAttachments(this.collectionForThread(threadId)).catch(() => [])
const duplicate = existingFiles.find((f: any) => f.name === file.name && f.path === file.path)
if (duplicate) {
throw new Error(`File '${file.name}' has already been attached to this thread`)
}View on GitHub (pinned to 7205d770c1)