JuliusBrussee/caveman · error · MiddlewareError

recovery_unavailable

recovery_unavailable

Error message

recovery_unavailable

What it means

_run backs a recovery tool bound to the wrapped model via a weakref. If the CavemanLLM/wrapper has been garbage-collected or closed, the recovery binding can no longer execute, so _run raises MiddlewareError with code recovery_unavailable instead of returning data from self._binding.execute.

Solutions

  1. Keep the wrapping model alive for the lifetime of the recovery tool; hold a strong reference while recovery is in use.
  2. Do not close the model until all recovery queries have completed.
  3. Recreate the model (and its recovery binding via runtime.recovery(scope)) and re-run the recovery query.
  4. Handle MiddlewareError code 'recovery_unavailable' by re-establishing the binding instead of retrying _run.

Example fix

// before
tool._run(handle, offset=0)  # model already closed -> MiddlewareError
// after
if not (model is None or model.closed):
    result = tool._run(handle, offset=0)
Defensive patterns

Strategy: try-catch

Validate before calling

if model is None or model.closed:
    # re-bind via runtime.recovery(scope) before calling _run
    ...

Type guard

def recovery_usable(model) -> bool:
    return model is not None and not model.closed

Try / catch

try:
    payload = tool._run(handle, offset=offset)
except MiddlewareError as e:
    if getattr(e, "code", None) == "recovery_unavailable":
        binding = runtime.recovery(scope)  # rebuild and retry once
    else:
        raise

Prevention

When it happens

Trigger: Calling the recovery tool's _run after the wrapping model has been closed (model.closed is True) or after the underlying model object was garbage collected (weakref returns None). The binding itself still exists, but its owner does not.

Common situations: Long-running crews where the LLM wrapper was closed but the recovery tool outlives it; holding a reference to the recovery tool beyond the model's lifetime; agent teardown happening while a background/async recovery query is still pending.

Understand the failure class

Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.

Related errors


AI-assisted analysis of JuliusBrussee/caveman@3ee70a1026 (2026-09-20). Data as JSON: /api/errors/8c1ef5efa5e28481. Report an issue: GitHub.

Appendix: source

Thrown at packages/middleware/python/caveman_middleware/crewai.py:71

class CavemanRecoveryTool(BaseTool):
    """Executed by CrewAI's existing native tool scheduler."""
    name: str = "caveman_retrieve"
    description: str = RECOVERY_DESCRIPTION
    args_schema: type[BaseModel] = _RecoveryInput
    cache_function: Any = _never_cache
    _binding: Any = PrivateAttr()
    _model: Any = PrivateAttr()

    def __init__(self, model):
        super().__init__()
        self._binding = model.runtime.recovery(model.scope)
        self._model = weakref.ref(model)

    def _run(self, handle, offset=0, limit=262144, query=""):
        model = self._model()
        if model is None or model.closed:
            raise MiddlewareError("recovery_unavailable")
        result = self._binding.execute(handle=handle, offset=offset, limit=limit, query=query)
        return json.dumps(result, ensure_ascii=False, separators=(",", ":"))


@dataclass
class _Call:
    model: Any
    attempt: Attempt
    finished: bool = False
    lock: Any = field(default_factory=threading.Lock)

    def finish(self, event, measured=None):
        with self.lock:
            if self.finished:
                return
            self.finished = True
        self.attempt.observe(event, measured)

View on GitHub (pinned to 3ee70a1026)