{"record":{"id":"cb97b956face6bac","repo":"janhq/jan","slug":"tokenize-request-failed-with-status-res-status","errorCode":null,"errorMessage":"Tokenize request failed with status ${res.status}","messagePattern":"Tokenize request failed with status (.+?)","errorType":"http","errorClass":null,"httpStatus":null,"severity":"error","filePath":"extensions/llamacpp-extension/src/index.ts","lineNumber":2916,"sourceCode":"   * on its session port. Char-based chunking can't reliably predict token\n   * count (subword tokenizers vary widely by content), so callers that need\n   * a hard guarantee against exceed_context_size_error should verify with\n   * this rather than estimating from character length.\n   */\n  async countEmbeddingTokens(texts: string[]): Promise<number[]> {\n    const sInfo = await this.ensureEmbeddingModelLoaded()\n    const counts: number[] = []\n    for (const text of texts) {\n      const res = await fetch(`http://localhost:${sInfo.port}/tokenize`, {\n        method: 'POST',\n        headers: {\n          'Content-Type': 'application/json',\n          'Authorization': `Bearer ${sInfo.api_key}`,\n        },\n        body: JSON.stringify({ content: text, model: sInfo.model_id }),\n      })\n      if (!res.ok) {\n        throw new Error(`Tokenize request failed with status ${res.status}`)\n      }\n      const json = (await res.json()) as { tokens?: unknown[] }\n      counts.push(Array.isArray(json.tokens) ? json.tokens.length : 0)\n    }\n    return counts\n  }\n\n  /**\n   * Token budget for one embedding request.\n   *\n   * Deliberately not the engine-wide `ubatch_size`: preset.ts pins every\n   * embedding model's section to its own ubatch (DEFAULT_EMBEDDING_UBATCH) and\n   * to `ctx-size = 0`, so the real ceiling is the embedder's own trained\n   * context -- 512 on MiniLM. llama.cpp rejects a batch wider than either with\n   * no retry path, so the budget is the smaller of the two.\n   */\n  private async embedBatchBudget(sInfo: SessionInfo): Promise<number> {\n    let budget = DEFAULT_EMBEDDING_UBATCH","sourceCodeStart":2898,"sourceCodeEnd":2934,"githubUrl":"https://github.com/janhq/jan/blob/7205d770c1e097c3daf35a911176410e93bc5564/extensions/llamacpp-extension/src/index.ts#L2898-L2934","documentation":"Thrown when the HTTP tokenize endpoint of the running llama.cpp server returns a non-2xx status. The extension POSTs {content, model} with a Bearer API key per input text and aborts on the first failed response. It surfaces the raw HTTP status so callers can distinguish auth vs server errors.","triggerScenarios":"Calling the token-count/estimate API while the local llama.cpp server is down, the model_id in session info is wrong/unloaded, or the api_key is rejected (401/403); also 404 if the server build lacks the /tokenize route.","commonSituations":"Counting tokens before a model finished loading, stale session info pointing at a restarted server, misconfigured API key, or an older llama.cpp server without tokenize support.","solutions":["Check the status code: 401/403 means fix api_key; 404 means server lacks the tokenize endpoint; 5xx means server-side failure","Verify the llama.cpp server for the session is running and the model is loaded","Confirm the session's model_id matches a model actually loaded on the server","Update the llama.cpp backend/server to a version that supports /tokenize","Retry after restarting the model session if the server was mid-restart"],"exampleFix":"// before\nconst counts = await tokenize(texts)\n// after\ntry {\n  const counts = await tokenize(texts)\n} catch (e) {\n  logger.warn('Token counting failed, falling back to estimate', e)\n  const counts = texts.map((t) => Math.ceil(t.length / 4))\n}","handlingStrategy":"try-catch","validationCode":"// Ensure session/server is reachable first\nconst alive = await fetch(`${baseUrl}/health`).then(r => r.ok).catch(() => false)\nif (!alive) throw new Error('llama.cpp server not running; skip tokenize')","typeGuard":null,"tryCatchPattern":"try {\n  const counts = await tokenize(texts)\n} catch (e) {\n  const status = /status (\\d+)/.exec(String(e))?.[1]\n  if (status === '401' || status === '403') fixApiKey()\n  else counts = texts.map(t => Math.ceil(t.length / 4))\n}","preventionTips":["Verify the model is loaded before token counting","Keep the llama.cpp server version current enough to expose /tokenize","Validate api_key configuration at startup","Implement a char-length fallback estimator for non-critical paths"],"tags":["http","network","tokenization","llamacpp"],"backgroundTag":"http-error-response","analyzedSha":"7205d770c1e097c3daf35a911176410e93bc5564","analyzedAt":"2026-09-17T14:27:30.100Z","contentChangedAt":"2026-09-17T14:27:30.100Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}