Model updates ship behind a shadow-mode flag first. The new model runs alongside the existing one for a fixed evaluation window; only the existing one affects clinical outputs.
We compare the two on a clinically-curated diff set. If the new model is better on the metrics that matter and no worse on the metrics that don't, we cut over.
We never auto-promote a model. The cutover is a deliberate human decision, with sign-off from clinical leadership.
