affaan-m/ECC · error · ValueError

Model failed promotion gates

Error message

Model failed promotion gates: {failures}

What it means

Illustrative gate check from the mle-workflow skill: all required metrics were present, but at least one violated its direction/threshold (e.g. auc below 0.82 or p95 latency above 80ms), so assert_promotion_ready raises listing the failed metrics. This is the deliberate 'do not ship' outcome.

Solutions

  1. Block the release pipeline on gate failure rather than overriding manually
  2. Review failing gates for regressions vs. over-strict thresholds
  3. Keep threshold changes in review like code changes
Defensive patterns

Strategy: validation

When it happens

Trigger: Thrown at skills/mle-workflow/SKILL.md:287 when the library encounters an invalid state.

Common situations: See trigger scenarios.


AI-assisted analysis of affaan-m/ECC@d8409a4b08 (2026-08-26). Data as JSON: /api/errors/a68b2f3eaa2303ef. Report an issue: GitHub.

Appendix: source

Thrown at skills/mle-workflow/SKILL.md:287

    "calibration_error": ("max", 0.04),
    "p95_latency_ms": ("max", 80),
}


def assert_promotion_ready(metrics: dict[str, float]) -> None:
    missing = sorted(name for name in PROMOTION_GATES if name not in metrics)
    if missing:
        raise ValueError(f"Model promotion metrics missing required gates: {missing}")

    failures = {
        name: value
        for name, (direction, threshold) in PROMOTION_GATES.items()
        for value in [metrics[name]]
        if (direction == "min" and value < threshold)
        or (direction == "max" and value > threshold)
    }
    if failures:
        raise ValueError(f"Model failed promotion gates: {failures}")
```

Use offline metrics as gates, not guarantees. When the model changes product behavior, plan shadow evaluation, canary rollout, or A/B testing before full rollout.

### 5. Package for Serving

An ML artifact is production-ready only when the serving contract is testable:

- Model artifact includes version, training data reference, config, and preprocessing
- Input schema rejects invalid, stale, or out-of-range features
- Output schema includes model version and confidence or explanation fields when useful
- Serving path has timeout, batching, resource limits, and fallback behavior
- CPU/GPU requirements are explicit and tested
- Prediction logs avoid PII and include enough identifiers for debugging and label joins
- Integration tests cover missing features, stale features, bad types, empty batches, and fallback path

Never let training-only feature code diverge from serving feature code without a test that proves equivalence.

View on GitHub (pinned to d8409a4b08)