{"record":{"id":"c950d57c1b565a83","repo":"janhq/jan","slug":"api-request-failed-with-status-response-status","errorCode":null,"errorMessage":"API request failed with status ${response.status}: ${JSON.stringify(errorData)}","messagePattern":"API request failed with status (.+?): (.+?)","errorType":"http","errorClass":null,"httpStatus":null,"severity":"error","filePath":"extensions/llamacpp-extension/src/index.ts","lineNumber":2422,"sourceCode":"        combinedController.abort(abortController.signal.reason)\n      } else {\n        abortController.signal.addEventListener(\n          'abort',\n          () => combinedController.abort(abortController.signal.reason),\n          { once: true }\n        )\n      }\n    }\n    const response = await fetch(url, {\n      method: 'POST',\n      headers,\n      body,\n      connectTimeout: Number(this.timeout) * 1000, // default 10 minutes\n      signal: combinedController.signal,\n    }).finally(() => clearTimeout(timeoutId))\n    if (!response.ok) {\n      const errorData = await response.json().catch(() => null)\n      throw new Error(\n        `API request failed with status ${response.status}: ${JSON.stringify(\n          errorData\n        )}`\n      )\n    }\n\n    if (!response.body) {\n      throw new Error('Response body is null')\n    }\n\n    const reader = response.body.getReader()\n    const decoder = new TextDecoder('utf-8')\n    let buffer = ''\n    let jsonStr = ''\n    try {\n      while (true) {\n        const { done, value } = await reader.read()\n","sourceCodeStart":2404,"sourceCodeEnd":2440,"githubUrl":"https://github.com/janhq/jan/blob/7205d770c1e097c3daf35a911176410e93bc5564/extensions/llamacpp-extension/src/index.ts#L2404-L2440","documentation":"Thrown by the llamacpp extension's streaming chat-completion request wrapper when the llama.cpp server returns a non-2xx HTTP response. The extension reads the JSON error body (tolerating a missing/invalid body) and embeds both the status code and the raw payload into the message so the caller can see exactly what the inference server rejected.","triggerScenarios":"Any call to the streaming completion path where response.ok is false — e.g. model not loaded (404/400), OOM or server crash (500), invalid request body, or the server port not actually serving llama.cpp.","commonSituations":"Model failed to load or was unloaded while a request was in flight; context length exceeded by prompt; malformed chat template; llama.cpp server returning 500 on internal errors; pointing at the wrong port where another service answers.","solutions":["Check the llama.cpp server logs for the root cause (load failure, OOM, bad request).","Verify the model is loaded before sending requests (loadModel / check session health).","Confirm the base URL/port matches the running llama.cpp server instance.","Reduce prompt/context size or ctx length if the error indicates allocation failure."],"exampleFix":"// before\nconst res = await fetch(url, { body })\nconst reader = res.body.getReader() // crashes or throws generic error\n// after\nconst res = await fetch(url, { body })\nif (!res.ok) {\n  const err = await res.json().catch(() => null)\n  throw new Error(`llamacpp request failed (${res.status}): ${JSON.stringify(err)}`)\n}\nconst reader = res.body.getReader()","handlingStrategy":"try-catch","validationCode":"const ok = await fetch(`${base}/health`).then(r => r.ok).catch(() => false)\nif (!ok) throw new Error('llamacpp server is not reachable')\nif (!(await isModelLoaded(modelId))) await loadModel(modelId)","typeGuard":"function isApiErrorResponse(e: unknown): e is Error & { status?: number } {\n  return e instanceof Error && /API request failed with status \\d+/.test(e.message)\n}","tryCatchPattern":"try {\n  for await (const chunk of stream) { /* ... */ }\n} catch (e) {\n  if (isApiErrorResponse(e)) {\n    const status = Number(e.message.match(/status (\\d+)/)?.[1])\n    if (status >= 500) await restartEngineAndRetry()\n    else showUserError(e.message)\n  } else throw e\n}","preventionTips":["Health-check the llama.cpp server before issuing completions.","Ensure the model is loaded before sending chat requests.","Keep prompt+max_tokens within the configured ctx size.","Pin/verify server version compatibility with the extension."],"tags":["http","api","llamacpp","network"],"backgroundTag":"http-error-response","analyzedSha":"7205d770c1e097c3daf35a911176410e93bc5564","analyzedAt":"2026-09-17T14:27:30.100Z","contentChangedAt":"2026-09-17T14:27:30.100Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}