Your Fairness Test Has a Shelf Life
Prudential validation measures whether a model still discriminates well and calibrates cleanly. Neither of those degrades when fairness does. A note for validation functions in banks and insurers.
The awkward property of a passed test
Your internal model is revalidated on a cycle. For SSM-significant institutions the cadence and the content of that cycle are shaped by the ECB Guide to Internal Models, whose 2025 release brings the AI Act into the frame of reference for the first time. For insurers, Article 124 of Solvency II requires a regular validation process — backtesting against actual experience, stability analysis, sensitivity testing of key assumptions, data quality assessment. The cycle exists. It is disciplined, documented, and supervised.
Now consider a fairness assessment carried out once, before a model went into production, and signed off.
It has no cycle. It has a date.
That asymmetry is not a compliance oversight. It is a structural feature of how the two disciplines grew up. Prudential validation was built to answer one question — is the model statistically sound, and are the resulting capital requirements adequate — and it answers it well, repeatedly, on schedule. Fairness testing arrived later, from a different literature, and it arrived as an audit: a snapshot, taken once, of a system that does not stand still.
What bias creep is
Static bias is the version most people have in mind: a model that learned discriminatory patterns from historical data at the moment it was built. It can be identified in a one-off audit and, at least in part, corrected.
Bias creep is the slow degradation of fairness metrics across a model’s operating life. It arises from several mechanisms that compound:
→ Data drift — the input distribution shifts. The applicant population in 2026 is not the applicant population the model was trained on.
→ Label shift — the real relationship between groups and outcomes moves away from what the training data encoded.
→ Feedback loops — the model’s outputs shape the data it later learns from. Applicants it declined generate no repayment history; the model converges, quietly, on its own priors.
→ Societal shift — variables that were once neutral proxies become correlated with protected characteristics, or cease to be.
Here is the part that matters for a validation function: through all of this, the model keeps working. Aggregate accuracy holds. Discriminatory power holds — the AUC or Gini statistic you report will not move. Calibration holds, because calibration is measured on the portfolio, and the portfolio-level relationship between predicted and realised default can remain perfectly sound while subgroup outcomes diverge underneath it.
There is no error to find. There is no exception to escalate. The metrics you track are, in the technical sense, fine.
Why this is a mathematical property, not a modelling defect
This is worth stating precisely, because it is the point at which conversations with quantitatively serious people either open up or close down.
A model can exhibit excellent discriminatory power and clean calibration across the whole portfolio and still produce systematically different outcomes for protected groups, with no underlying difference in risk. That is not a hypothesis about badly built models. Once base rates differ between groups, calibration and the common definitions of group fairness cannot all be satisfied at once. Chouldechova established this in 2016, months after the ProPublica investigation into recidivism scoring made the trade-off visible outside the research community. It is a property of the problem, not a flaw in a particular implementation. (We have written about how that literature ended up shaping the text of the AI Act here.)
Which means: no amount of improvement in the metrics your validation already tracks will surface a fairness problem. They are measuring something else. Correctly, rigorously — and something else.
What the AI Act adds, and to whom
Two provisions carry the time dimension.
Article 9(2) frames the risk management system for high-risk AI as a continuous, iterative process across the entire lifecycle, subject to regular systematic review and updating. Not a gate. A cycle.
Article 72 requires providers of high-risk systems to operate a documented post-market monitoring plan: an active, systematic collection and analysis of performance data across the operating life of the system.
Which of the two applies to you depends on a distinction institutions routinely get wrong about themselves. If your institution develops its own credit scoring or life and health pricing models — whether in-house or through a captive model factory serving the group — you are a provider under the Regulation, not merely a deployer. Article 72 is yours, in full.
The relief clause, and its exact limit
If you buy the model in, you are a deployer, and Article 26 governs. Article 26(5) requires deployers to monitor the operation of the system in accordance with the instructions for use and to inform the provider — and, where the system may present a risk, the market surveillance authority, suspending use in the meantime.
And then comes a sentence that is easy to miss and worth reading twice, because the legislator wrote it specifically for you:
“For deployers that are financial institutions subject to requirements regarding their internal governance, arrangements or processes under Union financial services law, the monitoring obligation set out in the first subparagraph shall be deemed to be fulfilled by complying with the rules on internal governance arrangements, processes and mechanisms pursuant to the relevant financial service law.” — Article 26(5), final subparagraph
That is a genuine deeming provision. Comply with your existing internal governance regime, and the deployer monitoring obligation is treated as met. Article 26(6) does the same for logging: keep the logs as part of the documentation you already maintain under financial services law. This is Recital 158’s integration logic made operative, and it is one of the more sensible things in the Regulation.
Three limits, however, and each of them matters:
→ It covers the monitoring obligation of the first subparagraph — not the notification duties. The obligation to inform the provider and the market surveillance authority when a system may present a risk, and to report serious incidents, is unaffected. The deeming clause tells you how to monitor. It does not tell you that you may stay silent about what you find.
→ It applies to you as a deployer. It does not touch Article 72. If you build your own models, you are the provider, and there is no deeming clause on that side of the line. This is precisely the institution type most likely to assume the clause protects it.
→ And it presupposes that your internal governance actually looks at the thing. “Deemed fulfilled by complying with the rules on internal governance” is only as strong as those rules. If your validation cycle tracks discriminatory power and calibration and nothing else, then complying with it means monitoring a system rigorously — for a property that is not the one at issue. The deeming clause is a bridge from one process to another. It is not a bypass around a metric that was never in either.
One further deployer obligation, distinct from all of the above and frequently conflated with Article 86: Article 26(11) requires deployers of Annex III systems that make or assist decisions about natural persons to inform those persons that they are subject to the system. That is notification, not explanation. The explanation right under Article 86 sits on top of it.
ISO/IEC 42001:2023 states the same expectation in the management-system idiom: Clause 9 requires ongoing performance evaluation, not a certificate with a date on it.
The good news, which is structural
None of this calls for a second validation cycle running in parallel to the one you already have. Recital 158 of the AI Act says so directly: for credit institutions under the CRD and for insurers and reinsurers under Solvency II, it is appropriate to integrate the procedural obligations on risk management, post-market monitoring and documentation into existing procedures — with the express aim of avoiding duplication.
You already have the cadence. You already have the second-line function, the documentation standard, the escalation path, the committee that receives the findings.
What you do not have, in most institutions we speak to, is one family of metrics inside that cadence.
Concretely, what has to enter the existing validation cycle:
→ Temporal fairness metrics, measured against a baseline rather than an absolute threshold. The operative question is not “is disparate impact below 0.8 today” but “has the ratio moved since the last cycle, in which subgroups, and under what conditions”. A fairness assessment that reports a single number and no delta is a snapshot wearing the costume of a monitoring process.
→ Slice testing. Model performance broken out across subpopulations that aggregate metrics are designed, by construction, to hide.
→ A documented trade-off decision. Given that calibration and group fairness cannot be jointly satisfied under differing base rates, someone has to choose which definition governs, and write down why. This is not a technical choice. It is a governance choice with a technical component, and it belongs in the validation report where a supervisor — and a court — can read it.
→ Thresholds that trigger action, and a rollback path. A monitoring process that detects degradation and has no defined consequence is a reporting process.
The tooling exists and is unremarkable — AIF360, Fairlearn, Themis-ML integrate into existing data pipelines without ceremony. The hard part was never the tooling.
The question that actually stalls this
In practice, the obstacle we encounter first is rarely methodological. It is this: fairness testing usually means an external provider gains access to sensitive portfolio data — which triggers your outsourcing assessment and your DORA ICT third-party risk review before a single test has run.
That is why waveTest computes client-side, in a Docker container inside your own infrastructure. Only results and reporting leave the house; the portfolio data never does. In governance terms rather than architectural ones: no third-party access to the underlying data, therefore a materially lighter outsourcing assessment, therefore no additional ICT third-party risk to document under DORA.
That is an argument for whoever signs the outsourcing review, not for IT.
On timing
Following the Digital Omnibus, the Annex III high-risk obligations apply from 2 December 2027. We do not think that is a reason for urgency, and we will not argue it that way.
It is, however, roughly one or two validation cycles away. Adding a metric family to an existing cycle is a modest piece of work — if it is added to the cycle. Building a parallel fairness process alongside it, under time pressure, because the cycle was never opened up, is not.
A question, in closing
Article 26(5) lets a financial institution discharge its deployer monitoring obligation by complying with the internal governance rules it already has. That is a real and welcome piece of drafting. It also means the quality of your compliance is now exactly the quality of that internal governance — no better, and no different.
So: in your last validation cycle, was any fairness metric compared against the previous cycle’s baseline — or against an absolute threshold, in isolation, as if the model had just been built?
The answer tends to be diagnostic. And it is usually the start of a more interesting conversation than the one about deadlines.
A more complete delineation — what prudential validation already covers, and the three artefacts the AI Act genuinely adds (Article 10, Article 86, Article 27) — is set out in our note on where the AI Act goes beyond existing model risk management.
Disclaimer
This article is provided for general informational purposes only and does not constitute legal, tax, or financial advice for any individual situation. Its contents reflect our understanding as of the publication date; regulatory frameworks in particular — including the EU AI Act and its national implementation — may have changed since. We make no warranty as to completeness, currency, or accuracy. Reading this article does not create an advisory or client relationship with waveImpact GmbH. For guidance tailored to your specific circumstances, please consult a suitably qualified professional.
Sources
Regulation (EU) 2024/1689 (AI Act) — Articles 9, 10, 26, 72 and Recital 158, consolidated text (EUR-Lex): https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=OJ%3AL_202401689 EUR-Lex does not deep-link individual articles; for day-to-day lookup the unofficial but reliable reference used across the field is more practical: https://artificialintelligenceact.eu/the-act/
Directive 2009/138/EC (Solvency II) — Articles 121, 122, 124 (EUR-Lex): https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32009L0138
ECB Guide to Internal Models (Release 4.0, July 2025): https://www.bankingsupervision.europa.eu/activities/internal_models/html/index.en.html
Regulation (EU) 2022/2554 (DORA) — ICT third-party risk (EUR-Lex): https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32022R2554
ISO/IEC 42001:2023, Clause 9 (performance evaluation): https://www.iso.org/standard/81230.html
Chouldechova, A. (2016), Fair prediction with disparate impact: A study of bias in recidivism prediction instruments: https://arxiv.org/abs/1610.07524