janhq/jan · error
API request failed with status
Error message
API request failed with status ${response.status}: ${JSON.stringify(errorData)} What it means
Thrown by the llamacpp extension's streaming chat-completion request wrapper when the llama.cpp server returns a non-2xx HTTP response. The extension reads the JSON error body (tolerating a missing/invalid body) and embeds both the status code and the raw payload into the message so the caller can see exactly what the inference server rejected.
Solutions
- Check the llama.cpp server logs for the root cause (load failure, OOM, bad request).
- Verify the model is loaded before sending requests (loadModel / check session health).
- Confirm the base URL/port matches the running llama.cpp server instance.
- Reduce prompt/context size or ctx length if the error indicates allocation failure.
Example fix
// before
const res = await fetch(url, { body })
const reader = res.body.getReader() // crashes or throws generic error
// after
const res = await fetch(url, { body })
if (!res.ok) {
const err = await res.json().catch(() => null)
throw new Error(`llamacpp request failed (${res.status}): ${JSON.stringify(err)}`)
}
const reader = res.body.getReader() Defensive patterns
Strategy: try-catch
Validate before calling
const ok = await fetch(`${base}/health`).then(r => r.ok).catch(() => false)
if (!ok) throw new Error('llamacpp server is not reachable')
if (!(await isModelLoaded(modelId))) await loadModel(modelId) Type guard
function isApiErrorResponse(e: unknown): e is Error & { status?: number } {
return e instanceof Error && /API request failed with status \d+/.test(e.message)
} Try / catch
try {
for await (const chunk of stream) { /* ... */ }
} catch (e) {
if (isApiErrorResponse(e)) {
const status = Number(e.message.match(/status (\d+)/)?.[1])
if (status >= 500) await restartEngineAndRetry()
else showUserError(e.message)
} else throw e
} Prevention
- Health-check the llama.cpp server before issuing completions.
- Ensure the model is loaded before sending chat requests.
- Keep prompt+max_tokens within the configured ctx size.
- Pin/verify server version compatibility with the extension.
When it happens
Trigger: Any call to the streaming completion path where response.ok is false — e.g. model not loaded (404/400), OOM or server crash (500), invalid request body, or the server port not actually serving llama.cpp.
Common situations: Model failed to load or was unloaded while a request was in flight; context length exceeded by prompt; malformed chat template; llama.cpp server returning 500 on internal errors; pointing at the wrong port where another service answers.
Understand the failure class
Background: "API error: {status}" and "HTTP 401/403/404/429/5xx" errors: non-2xx HTTP responses explained — this error's family across 27 libraries.
Related errors
- Failed to fetch models from
- Failed to fetch releases
- Tokenize request failed with status
- Failed to fetch HuggingFace repository
- Failed to fetch model catalog
AI-assisted analysis of janhq/jan@7205d770c1 (2026-09-17).
Data as JSON: /api/errors/c950d57c1b565a83.
Report an issue: GitHub.
Appendix: source
Thrown at extensions/llamacpp-extension/src/index.ts:2422
combinedController.abort(abortController.signal.reason)
} else {
abortController.signal.addEventListener(
'abort',
() => combinedController.abort(abortController.signal.reason),
{ once: true }
)
}
}
const response = await fetch(url, {
method: 'POST',
headers,
body,
connectTimeout: Number(this.timeout) * 1000, // default 10 minutes
signal: combinedController.signal,
}).finally(() => clearTimeout(timeoutId))
if (!response.ok) {
const errorData = await response.json().catch(() => null)
throw new Error(
`API request failed with status ${response.status}: ${JSON.stringify(
errorData
)}`
)
}
if (!response.body) {
throw new Error('Response body is null')
}
const reader = response.body.getReader()
const decoder = new TextDecoder('utf-8')
let buffer = ''
let jsonStr = ''
try {
while (true) {
const { done, value } = await reader.read()
View on GitHub (pinned to 7205d770c1)