affaan-m/ECC · error · ValueError

JSON artifact exceeds local size limit

Error message

JSON artifact exceeds local size limit

What it means

When reading a JSON artifact (`parse_json=True`) and no explicit `expected_size` was supplied by the caller beyond the stat, `_read_local` enforces a maximum JSON payload size `_MAX_JSON`. This protects the process from unbounded memory consumption when joining and parsing the file's chunks. A file larger than that cap is refused before any bytes are read.

Solutions

  1. Shrink or split the JSON artifact so it fits under the local size limit, or move bulk data into sidecar files referenced by hash.
  2. If the file is legitimately large, switch the pipeline to a non-JSON (or streamed) artifact path rather than `parse_json=True`.
  3. Check the size before calling: `os.path.getsize(p) <= _MAX_JSON`, and regenerate the artifact if it exceeds it.
  4. Verify the artifact generator isn't duplicating content (e.g. writing debug payloads into the JSON).

Example fix

// before
req = load_application_request("/out/request.json")  # 250MB JSON dump
// after
import os
if os.path.getsize("/out/request.json") > _MAX_JSON:
    slim_request("/out/request.json")  # strip oversized embedded payloads first
req = load_application_request("/out/request.json")
Defensive patterns

Strategy: validation

Validate before calling

import os
def assert_json_fits(path: str, max_json: int) -> None:
    size = os.path.getsize(path)
    if size > max_json:
        raise ValueError(f"JSON artifact {size} bytes exceeds limit {max_json}")

Try / catch

try:
    req = load_application_request(p)
except ValueError as e:
    if str(e) == "JSON artifact exceeds local size limit":
        slim_or_split_artifact(p)   # move bulk data to sidecar files, then retry
        req = load_application_request(p)
    else:
        raise

Prevention

When it happens

Trigger: Calling `load_application_request` (or `_artifact` with `parse_json=True`) on a JSON file whose size exceeds `_MAX_JSON` bytes; an upstream tool accidentally writing a huge log as JSON with a `.json` extension.

Common situations: An application-request JSON that accumulated enormous embedded data (base64 blobs, full dumps); a runaway generator concatenating artifacts; pointing the loader at a data file that was never meant to be a request document.

Understand the failure class

Background: "File too large" / "file size exceeds limit" errors: why libraries cap file sizes and how to fix them — this error's family across 46 libraries.

Related errors


AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16). Data as JSON: /api/errors/01f1a651dfeeb11d. Report an issue: GitHub.

Appendix: source

Thrown at skills/taste-application/scripts/tasteforge/integration.py:127

        raise


def _read_local(raw: str, *, parse_json: bool, expected_size: int | None = None,
                expected_hash: str | None = None) -> Any:
    path = Path(raw)
    if not path.is_absolute() or str(path) != raw or ".." in path.parts:
        raise ValueError("artifact path must be canonical and absolute")
    parent = descriptor = None
    try:
        flags = os.O_RDONLY | os.O_NOFOLLOW | os.O_NONBLOCK
        parent = _parent_fd(path)
        before = os.stat(path.name, dir_fd=parent, follow_symlinks=False)
        if not stat.S_ISREG(before.st_mode) or getattr(before, "st_flags", 0) & 0x40000000:
            raise ValueError("artifact must be a resident regular file")
        if expected_size is None:
            expected_size = before.st_size
        if parse_json and expected_size > _MAX_JSON:
            raise ValueError("JSON artifact exceeds local size limit")
        if before.st_size != expected_size:
            raise ValueError("artifact byte count mismatch")
        descriptor = os.open(path.name, flags, dir_fd=parent)
        if _identity(before) != _identity(os.fstat(descriptor)):
            raise ValueError("artifact changed before reading")
        digest, chunks, count = hashlib.sha256(), [], 0
        while data := os.read(descriptor, 65536):
            count += len(data)
            if count > expected_size:
                raise ValueError("artifact byte count exceeded during reading")
            digest.update(data)
            if parse_json:
                chunks.append(data)
        # Rewalk the named path: a pinned old directory fd can outlive a rename.
        fresh_parent = _parent_fd(path)
        try:
            after = os.stat(path.name, dir_fd=fresh_parent, follow_symlinks=False)
        finally:

View on GitHub (pinned to 8321021c54)