janhq/jan · error
error.message
Error message
error.message
What it means
Thrown while parsing the SSE stream when a line begins with 'error: ' — the extension extracts the JSON payload after the prefix, parses it, and re-throws the server-supplied error.message. This propagates llama.cpp's own in-stream error to the caller.
Solutions
- Read error.message for the server's actual reason and address it (usually context overflow or server shutdown).
- Reduce max_tokens/prompt size if the message indicates context limits.
- Check llama.cpp server logs for the corresponding internal error.
- Implement client retry with backoff if the server was restarting under load.
Example fix
// before
} else if (trimmedLine.startsWith('error: ')) {
const error = JSON.parse(trimmedLine.slice(7))
throw new Error(error.message)
}
// after
} else if (trimmedLine.startsWith('error: ')) {
const error = JSON.parse(trimmedLine.slice(7))
const err = new Error(error.message || 'Unknown llamacpp stream error')
err.payload = error
throw err
} Defensive patterns
Strategy: try-catch
Validate before calling
null
Type guard
function isStreamErrorEvent(e: unknown): e is Error & { payload?: { message: string } } {
return e instanceof Error && typeof (e as any).payload === 'object'
} Try / catch
try {
for await (const chunk of stream) yield chunk
} catch (e) {
if (/context|token|OOM|memory/i.test(String((e as Error).message))) {
// surface actionable guidance: reduce context or restart engine
}
throw e
} Prevention
- Cap prompt + max_tokens below the model's context window.
- Monitor engine memory; CUDA OOM mid-stream throws these errors.
- Handle stream errors per-request, not globally, to preserve other sessions.
- Check server logs whenever error.message is opaque.
When it happens
Trigger: The llama.cpp server sends an SSE line 'error: {"message": ...}' mid-stream — e.g. inference aborted, context overflow discovered during generation, or the model crashed mid-completion.
Common situations: Long generations exceeding context after streaming began; server-side abort (client disconnect handling, load shutdown); inference runtime errors (CUDA OOM mid-token).
Related errors
AI-assisted analysis of janhq/jan@7205d770c1 (2026-09-17).
Data as JSON: /api/errors/45537e8f839a00fc.
Report an issue: GitHub.
Appendix: source
Thrown at extensions/llamacpp-extension/src/index.ts:2462
buffer += decoder.decode(value, { stream: true })
// Process complete lines in the buffer
const lines = buffer.split('\n')
buffer = lines.pop() || '' // Keep the last incomplete line in the buffer
for (const line of lines) {
const trimmedLine = line.trim()
if (!trimmedLine || trimmedLine === 'data: [DONE]') {
continue
}
if (trimmedLine.startsWith('data: ')) {
jsonStr = trimmedLine.slice(6)
} else if (trimmedLine.startsWith('error: ')) {
jsonStr = trimmedLine.slice(7)
const error = JSON.parse(jsonStr)
throw new Error(error.message)
} else {
// it should not normally reach here
throw new Error('Malformed chunk')
}
try {
const data = JSON.parse(jsonStr)
const chunk = data as chatCompletionChunk
yield chunk
} catch (e) {
logger.error('Error parsing JSON from stream or server error:', e)
// re‑throw so the async iterator terminates with an error
throw e
}
}
}
} finally {
reader.releaseLock()View on GitHub (pinned to 7205d770c1)