{"record":{"id":"82d7c555460bfd26","repo":"janhq/jan","slug":"the-request-exceeds-the-available-context-size","errorCode":null,"errorMessage":"the request exceeds the available context size.","messagePattern":"the request exceeds the available context size\\.","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"extensions/mlx-extension/src/index.ts","lineNumber":441,"sourceCode":"\n    const response = await fetch(url, {\n      method: 'POST',\n      headers,\n      body,\n      signal: abortController?.signal,\n    })\n\n    if (!response.ok) {\n      const errorData = await response.json().catch(() => null)\n      throw new Error(\n        `MLX API request failed with status ${response.status}: ${JSON.stringify(errorData)}`\n      )\n    }\n\n    const completionResponse = (await response.json()) as chatCompletion\n\n    if (completionResponse.choices?.[0]?.finish_reason === 'length') {\n      throw new Error(OUT_OF_CONTEXT_SIZE)\n    }\n\n    return completionResponse\n  }\n\n  private async *handleStreamingResponse(\n    url: string,\n    headers: HeadersInit,\n    body: string,\n    abortController?: AbortController\n  ): AsyncIterable<chatCompletionChunk> {\n    // AbortSignal.any() is not available in all runtimes (e.g. WebKit/JavaScriptCore),\n    // so we manually combine the timeout and external abort signals.\n    const combinedController = new AbortController()\n    const timeoutId = setTimeout(\n      () => combinedController.abort(new Error('Request timed out')),\n      this.timeout * 1000\n    )","sourceCodeStart":423,"sourceCodeEnd":459,"githubUrl":"https://github.com/janhq/jan/blob/7205d770c1e097c3daf35a911176410e93bc5564/extensions/mlx-extension/src/index.ts#L423-L459","documentation":"After a successful non-streaming completion, the extension inspects finish_reason; when the MLX server reports 'length', the prompt+response exceeded the model's context window and the response was truncated, so the extension throws OUT_OF_CONTEXT_SIZE instead of returning a cut-off answer.","triggerScenarios":"Calling chat() with messages whose total token count plus max_tokens exceeds the model's context size, causing the server to stop generation with finish_reason === 'length'.","commonSituations":"Long conversation histories sent in full every turn, very large system prompts, models loaded with a small context (e.g. 2048) while users paste large documents.","solutions":["Reduce the input: trim or summarize older conversation messages before sending","Lower max_tokens in the request or in app settings","Reload the model with a larger context size (n_ctx / context length setting)","Switch to a model with a larger context window"],"exampleFix":"// before\nawait chat(allMessages, model)\n// after\nconst trimmed = allMessages.slice(-10) // keep recent turns only\nawait chat(trimmed, model)","handlingStrategy":"validation","validationCode":"function estimateTokens(messages) {\n  return messages.reduce((n, m) => n + Math.ceil((m.content?.length ?? 0) / 4), 0)\n}\nconst LIMIT = 2048 // model context size\nif (estimateTokens(messages) + (maxTokens ?? 1024) >= LIMIT) {\n  messages = [messages[0], ...messages.slice(-6)]\n}","typeGuard":"null","tryCatchPattern":"try {\n  return await chat(messages, model)\n} catch (e) {\n  if (String(e.message).includes('context size')) {\n    return chat([messages[0], ...messages.slice(-6)], model)\n  }\n  throw e\n}","preventionTips":["Track running token count across turns and compact history proactively","Set max_tokens below the model context minus the prompt size","Load models with the largest context your RAM allows","Summarize long documents before including them in prompts"],"tags":["mlx","context-length","token-limit","truncation"],"backgroundTag":"context-length-exceeded","analyzedSha":"7205d770c1e097c3daf35a911176410e93bc5564","analyzedAt":"2026-09-17T14:27:30.100Z","contentChangedAt":"2026-09-17T14:27:30.100Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}