janhq/jan · error

Tokenize request failed with status

Error message

Tokenize request failed with status ${res.status}

What it means

Thrown when the HTTP tokenize endpoint of the running llama.cpp server returns a non-2xx status. The extension POSTs {content, model} with a Bearer API key per input text and aborts on the first failed response. It surfaces the raw HTTP status so callers can distinguish auth vs server errors.

Solutions

  1. Check the status code: 401/403 means fix api_key; 404 means server lacks the tokenize endpoint; 5xx means server-side failure
  2. Verify the llama.cpp server for the session is running and the model is loaded
  3. Confirm the session's model_id matches a model actually loaded on the server
  4. Update the llama.cpp backend/server to a version that supports /tokenize
  5. Retry after restarting the model session if the server was mid-restart

Example fix

// before
const counts = await tokenize(texts)
// after
try {
  const counts = await tokenize(texts)
} catch (e) {
  logger.warn('Token counting failed, falling back to estimate', e)
  const counts = texts.map((t) => Math.ceil(t.length / 4))
}
Defensive patterns

Strategy: try-catch

Validate before calling

// Ensure session/server is reachable first
const alive = await fetch(`${baseUrl}/health`).then(r => r.ok).catch(() => false)
if (!alive) throw new Error('llama.cpp server not running; skip tokenize')

Try / catch

try {
  const counts = await tokenize(texts)
} catch (e) {
  const status = /status (\d+)/.exec(String(e))?.[1]
  if (status === '401' || status === '403') fixApiKey()
  else counts = texts.map(t => Math.ceil(t.length / 4))
}

Prevention

When it happens

Trigger: Calling the token-count/estimate API while the local llama.cpp server is down, the model_id in session info is wrong/unloaded, or the api_key is rejected (401/403); also 404 if the server build lacks the /tokenize route.

Common situations: Counting tokens before a model finished loading, stale session info pointing at a restarted server, misconfigured API key, or an older llama.cpp server without tokenize support.

Understand the failure class

Background: "API error: {status}" and "HTTP 401/403/404/429/5xx" errors: non-2xx HTTP responses explained — this error's family across 27 libraries.

Related errors


AI-assisted analysis of janhq/jan@7205d770c1 (2026-09-17). Data as JSON: /api/errors/cb97b956face6bac. Report an issue: GitHub.

Appendix: source

Thrown at extensions/llamacpp-extension/src/index.ts:2916

   * on its session port. Char-based chunking can't reliably predict token
   * count (subword tokenizers vary widely by content), so callers that need
   * a hard guarantee against exceed_context_size_error should verify with
   * this rather than estimating from character length.
   */
  async countEmbeddingTokens(texts: string[]): Promise<number[]> {
    const sInfo = await this.ensureEmbeddingModelLoaded()
    const counts: number[] = []
    for (const text of texts) {
      const res = await fetch(`http://localhost:${sInfo.port}/tokenize`, {
        method: 'POST',
        headers: {
          'Content-Type': 'application/json',
          'Authorization': `Bearer ${sInfo.api_key}`,
        },
        body: JSON.stringify({ content: text, model: sInfo.model_id }),
      })
      if (!res.ok) {
        throw new Error(`Tokenize request failed with status ${res.status}`)
      }
      const json = (await res.json()) as { tokens?: unknown[] }
      counts.push(Array.isArray(json.tokens) ? json.tokens.length : 0)
    }
    return counts
  }

  /**
   * Token budget for one embedding request.
   *
   * Deliberately not the engine-wide `ubatch_size`: preset.ts pins every
   * embedding model's section to its own ubatch (DEFAULT_EMBEDDING_UBATCH) and
   * to `ctx-size = 0`, so the real ceiling is the embedder's own trained
   * context -- 512 on MiniLM. llama.cpp rejects a batch wider than either with
   * no retry path, so the budget is the smaller of the two.
   */
  private async embedBatchBudget(sInfo: SessionInfo): Promise<number> {
    let budget = DEFAULT_EMBEDDING_UBATCH

View on GitHub (pinned to 7205d770c1)