affaan-m/ECC · error · ValueError

artifact byte count exceeded during reading

Error message

artifact byte count exceeded during reading

What it means

While streaming the file in 64KiB chunks, `_read_local` tracks the cumulative byte count and aborts immediately if it exceeds `expected_size`. This catches files that grew after the initial size check or descriptor-level identity checks — a defensive backstop against a file being appended to concurrently while it is being read.

Solutions

  1. Ensure no process appends to the artifact while it is being loaded; wait for the writer to close the file (e.g. check for a `.done` marker).
  2. Regenerate the artifact and its recorded size/hash after all writers finish, then retry the load.
  3. Retry the load once the concurrent appender has stopped — the pre-read checks will then pass.
  4. Change the producer to write atomically (temp file + `os.replace`) instead of appending in place.

Example fix

// before
some_tool >> /out/artifact.json   # still appending during load
// after
wait_for_done_marker("/out/artifact.json.done")
load_application_request("/out/request.json")
Defensive patterns

Strategy: retry

Validate before calling

import os, time
def wait_for_quiet(path: str, quiet_secs: float = 1.0, timeout: float = 30.0) -> None:
    deadline = time.time() + timeout
    last = os.path.getsize(path)
    while time.time() < deadline:
        time.sleep(quiet_secs)
        cur = os.path.getsize(path)
        if cur == last:
            return
        last = cur
    raise TimeoutError(f"file still growing: {path}")

Try / catch

try:
    req = load_application_request(p)
except ValueError as e:
    if str(e) == "artifact byte count exceeded during reading":
        wait_for_quiet(p)
        req = load_application_request(p)   # identity + size checks now pass
    else:
        raise

Prevention

When it happens

Trigger: A writer appends to the artifact file between the `os.open` and the read loop, making the actual byte stream longer than the stat-verified `expected_size`; a misreported or sparse size vs. actual readable bytes on an unusual filesystem.

Common situations: A logging tool still writing to the artifact; `tee`/append redirection continuing into the output file; a producer that appends a footer/trailer after the hash and size were recorded.

Related errors


AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16). Data as JSON: /api/errors/8dbf2fbf7ae07150. Report an issue: GitHub.

Appendix: source

Thrown at skills/taste-application/scripts/tasteforge/integration.py:137

        flags = os.O_RDONLY | os.O_NOFOLLOW | os.O_NONBLOCK
        parent = _parent_fd(path)
        before = os.stat(path.name, dir_fd=parent, follow_symlinks=False)
        if not stat.S_ISREG(before.st_mode) or getattr(before, "st_flags", 0) & 0x40000000:
            raise ValueError("artifact must be a resident regular file")
        if expected_size is None:
            expected_size = before.st_size
        if parse_json and expected_size > _MAX_JSON:
            raise ValueError("JSON artifact exceeds local size limit")
        if before.st_size != expected_size:
            raise ValueError("artifact byte count mismatch")
        descriptor = os.open(path.name, flags, dir_fd=parent)
        if _identity(before) != _identity(os.fstat(descriptor)):
            raise ValueError("artifact changed before reading")
        digest, chunks, count = hashlib.sha256(), [], 0
        while data := os.read(descriptor, 65536):
            count += len(data)
            if count > expected_size:
                raise ValueError("artifact byte count exceeded during reading")
            digest.update(data)
            if parse_json:
                chunks.append(data)
        # Rewalk the named path: a pinned old directory fd can outlive a rename.
        fresh_parent = _parent_fd(path)
        try:
            after = os.stat(path.name, dir_fd=fresh_parent, follow_symlinks=False)
        finally:
            os.close(fresh_parent)
        if (_identity(before) != _identity(os.fstat(descriptor))
                or _identity(before) != _identity(after)):
            raise ValueError("artifact changed during reading")
        if expected_hash is not None and digest.hexdigest() != expected_hash:
            raise ValueError("artifact SHA-256 mismatch")
        return _load_json(b"".join(chunks)) if parse_json else None
    except (OSError, AttributeError) as exc:
        raise ValueError("local artifact unavailable or unsafe") from exc
    finally:

View on GitHub (pinned to 8321021c54)