janhq/jan · error
Tokenize request failed with status
Error message
Tokenize request failed with status ${res.status} What it means
Thrown when the HTTP tokenize endpoint of the running llama.cpp server returns a non-2xx status. The extension POSTs {content, model} with a Bearer API key per input text and aborts on the first failed response. It surfaces the raw HTTP status so callers can distinguish auth vs server errors.
Solutions
- Check the status code: 401/403 means fix api_key; 404 means server lacks the tokenize endpoint; 5xx means server-side failure
- Verify the llama.cpp server for the session is running and the model is loaded
- Confirm the session's model_id matches a model actually loaded on the server
- Update the llama.cpp backend/server to a version that supports /tokenize
- Retry after restarting the model session if the server was mid-restart
Example fix
// before
const counts = await tokenize(texts)
// after
try {
const counts = await tokenize(texts)
} catch (e) {
logger.warn('Token counting failed, falling back to estimate', e)
const counts = texts.map((t) => Math.ceil(t.length / 4))
} Defensive patterns
Strategy: try-catch
Validate before calling
// Ensure session/server is reachable first
const alive = await fetch(`${baseUrl}/health`).then(r => r.ok).catch(() => false)
if (!alive) throw new Error('llama.cpp server not running; skip tokenize') Try / catch
try {
const counts = await tokenize(texts)
} catch (e) {
const status = /status (\d+)/.exec(String(e))?.[1]
if (status === '401' || status === '403') fixApiKey()
else counts = texts.map(t => Math.ceil(t.length / 4))
} Prevention
- Verify the model is loaded before token counting
- Keep the llama.cpp server version current enough to expose /tokenize
- Validate api_key configuration at startup
- Implement a char-length fallback estimator for non-critical paths
When it happens
Trigger: Calling the token-count/estimate API while the local llama.cpp server is down, the model_id in session info is wrong/unloaded, or the api_key is rejected (401/403); also 404 if the server build lacks the /tokenize route.
Common situations: Counting tokens before a model finished loading, stale session info pointing at a restarted server, misconfigured API key, or an older llama.cpp server without tokenize support.
Understand the failure class
Background: "API error: {status}" and "HTTP 401/403/404/429/5xx" errors: non-2xx HTTP responses explained — this error's family across 27 libraries.
Related errors
- API request failed with status
- Failed to fetch model catalog
- Failed to fetch models from
- Failed to fetch releases
- HTTP request failed
AI-assisted analysis of janhq/jan@7205d770c1 (2026-09-17).
Data as JSON: /api/errors/cb97b956face6bac.
Report an issue: GitHub.
Appendix: source
Thrown at extensions/llamacpp-extension/src/index.ts:2916
* on its session port. Char-based chunking can't reliably predict token
* count (subword tokenizers vary widely by content), so callers that need
* a hard guarantee against exceed_context_size_error should verify with
* this rather than estimating from character length.
*/
async countEmbeddingTokens(texts: string[]): Promise<number[]> {
const sInfo = await this.ensureEmbeddingModelLoaded()
const counts: number[] = []
for (const text of texts) {
const res = await fetch(`http://localhost:${sInfo.port}/tokenize`, {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Authorization': `Bearer ${sInfo.api_key}`,
},
body: JSON.stringify({ content: text, model: sInfo.model_id }),
})
if (!res.ok) {
throw new Error(`Tokenize request failed with status ${res.status}`)
}
const json = (await res.json()) as { tokens?: unknown[] }
counts.push(Array.isArray(json.tokens) ? json.tokens.length : 0)
}
return counts
}
/**
* Token budget for one embedding request.
*
* Deliberately not the engine-wide `ubatch_size`: preset.ts pins every
* embedding model's section to its own ubatch (DEFAULT_EMBEDDING_UBATCH) and
* to `ctx-size = 0`, so the real ceiling is the embedder's own trained
* context -- 512 on MiniLM. llama.cpp rejects a batch wider than either with
* no retry path, so the budget is the smaller of the two.
*/
private async embedBatchBudget(sInfo: SessionInfo): Promise<number> {
let budget = DEFAULT_EMBEDDING_UBATCHView on GitHub (pinned to 7205d770c1)