Every model is evaluated for performance disparities across demographic groups before it goes to production. Disparities above a clinical-team-defined threshold block the launch.
We also monitor for drift in production. A model that launched fair can become unfair as the data distribution shifts, and that's a real failure mode.
We publish our methodology and the results internally. We expect to publish them externally when the work is mature enough to defend.
