Interactive Explainer AI Governance

The Credit Scoring Drift Monitor

A model is only as fair as the population it was trained on. Move the clock forward — or press play — and watch what happens to accuracy, approval rates and the gap between demographic groups as the world quietly stops matching the training data, without a single line of code changing.

0
Launch 18 months 36 months

Advancing the slider (or pressing play) simulates population drift: the applicant base, macroeconomic conditions and behaviour patterns diverge from the data the model was trained on, while the model itself stays frozen — exactly as it would in a bank that reviews models annually rather than continuously.

0.3%

Reviewers agree with the model almost every time. This is the number a supervisor should ask for.

Model Accuracy

91.4%

baseline

Disparate Impact Ratio

0.94

4/5ths rule threshold: 0.80

Status

Within Tolerance

Based on internal model risk policy

Accuracy & Fairness Over Time

Accuracy Disparate Impact

Traced live as you move the slider. The dashed threshold line marks the regulatory 4/5ths fairness floor.

Fairness floor (0.80) Month 0 Month 36

Approval Rate by Applicant Group

The gap between these two bars is the disparate impact ratio above. Watch it widen as drift accumulates.

68%
Group A
(majority in training data)
64%
Group B
(underrepresented)

Month 0: Launch

The model has just been validated. Accuracy and fairness metrics sit comfortably inside policy. No one is watching for drift yet, because there hasn't been time for any to occur.

What this is actually showing

Data drift ≠ a bug

Nothing breaks. No error is thrown. The model keeps producing confident scores every day; it simply stops being scored against the world it was trained on.

The reviewer can't see this

A human looking at one application at a time cannot detect a population-level shift. Drift is only visible in aggregate, which is exactly why transaction-level "human in the loop" review misses it.

The override rate is the tell

As the model becomes miscalibrated, a genuinely engaged reviewer would start disagreeing with it more often. A flat override rate near zero, even as accuracy degrades, is itself evidence of rubber-stamping.

This explainer accompanies the blog post "The Model That Nobody Told".