Mintplex-Labs/anything-llm · error · Error

Could not load into Foundry Local

Error message

Could not load ${this.model} into Foundry Local: ${error}

What it means

FoundryLLM.assertModelLoaded calls FoundryModels.loadModel(this.model); when that returns { success: false, error }, it rethrows as 'Could not load <model> into Foundry Local: <error>'. The embedded error is either the non-OK HTTP status from the service's /models/load endpoint (see error 149) or 'No Foundry service or model was set.' The class keeps a #loadedModels memo and also forgets evicted models after mid-stream failures so the next message attempts a reload — meaning this error is how load failures surface per chat message.

Solutions

  1. Download the model outside AnythingLLM: `foundry model download <modelId>` (or via Docker Desktop), then retry the message
  2. Confirm the Foundry Local service is running and FOUNDRY_BASE_PATH points at it (curl the /models endpoint)
  3. Check RAM/disk: multi-GB loads abort on the timeout when the machine is starved — free resources or pick a smaller model
  4. Verify the model id matches exactly what `foundry model list` shows (alias vs full id)

Example fix

# before: model chosen in UI but never pulled
# -> Could not load ai/qwen2.5-0.5b into Foundry Local: Foundry could not load ... (HTTP 404)

# after
foundry model download ai/qwen2.5-0.5b-instruct
# then resend the chat message
Defensive patterns

Strategy: retry

Validate before calling

// pre-check that the model is downloaded/loaded before chatting
const loaded = await FoundryModels.loadedModels();
if (!loaded.some((id) => id === model || id.split(":")[0] === model)) {
  const { success, error } = await FoundryModels.loadModel(model);
  if (!success) throw new Error(`Cannot preload ${model}: ${error}`);
}

Try / catch

try {
  await llm.streamGetChatCompletion(messages);
} catch (err) {
  if (/Could not load .* into Foundry Local/.test(err.message)) {
    // evictions/idle timeouts are transient: one reload attempt is reasonable
    if (++attempt === 1) return retryMessage();
    return respond(`Foundry could not load the model: ${err.message}`);
  }
  throw err;
}

Prevention

When it happens

Trigger: getChatCompletion/streamGetChatCompletion on a model that is not currently loaded in Foundry Local, where the subsequent load call fails: the model was never downloaded (HTTP 404 from /models/load), the service is unreachable so the fetch throws, FOUNDRY_BASE_PATH is blank so origin resolution fails, or the load exceeded the multi-GB LOAD_TIMEOUT_MS and aborted.

Common situations: User picked a model id in AnythingLLM that exists in the catalog but was never pulled with `foundry model download`; Foundry Local service stopped or the port changed; machine low on RAM/disk so loading times out; model was loaded earlier but evicted by idle timeout and the reload fails.

Related errors


AI-assisted analysis of Mintplex-Labs/anything-llm@3aec848f28 (2026-08-18). Data as JSON: /api/errors/83dc5f0ff7290b41. Report an issue: GitHub.

Appendix: source

Thrown at server/utils/AiProviders/foundry/index.js:98

  async assertModelLoaded() {
    if (!this.model || FoundryLLM.#loadedModels.has(this.model)) return;
    const FoundryModels = require("./models.js");

    // The service reports fully-qualified variant ids while the preference is
    // usually an alias, so match on either side of the colon-versioned name.
    const loaded = await FoundryModels.loadedModels();
    const isLoaded = loaded.some(
      (id) => id === this.model || id.split(":")[0] === this.model
    );
    if (isLoaded) {
      FoundryLLM.#loadedModels.add(this.model);
      return;
    }

    this.#log(`Loading ${this.model} into Foundry Local...`);
    const { success, error } = await FoundryModels.loadModel(this.model);
    if (!success)
      throw new Error(
        `Could not load ${this.model} into Foundry Local: ${error}`
      );
    FoundryLLM.#loadedModels.add(this.model);
  }

  /**
   * Turn a mid-stream failure into something actionable.
   *
   * A model evicted after we loaded it — by an idle timeout, or from the host —
   * makes the service answer 200 and then drop the socket, which reaches us
   * only as "Premature close". Forget it so the next message reloads it.
   * @param {Error} error
   * @param {string} model
   * @returns {string}
   */
  static explainStreamError(error, model) {
    const isPrematureClose =
      error?.code === "ERR_STREAM_PREMATURE_CLOSE" ||

View on GitHub (pinned to 3aec848f28)