← back to the archiveCover illustration for “Monitor the cohort that justified the model”
POSTday 117·today·by Andy Padia

Monitor the cohort that justified the model

Affirm's largest offline underwriting lift appeared for thin-file users. Keep that cohort visible in production monitoring instead of letting pooled outcomes erase the release case.

Affirm's new underwriting stack reported its largest offline improvement among people with credit-report information but no FICO score. Its first online result, however, was reported across the broader population of eligible new users.

My rule is to keep the cohort that justified a model visible after launch. Give it a named production monitor with its own denominator, outcome window and rollback threshold. Otherwise a healthy portfolio average can erase the evidence that got the model approved.

Affirm's strongest lift came from thin-file users

Affirm's 17 September technical post describes a hybrid system. A transformer learns representations from detailed credit-report records and payment histories. An XGBoost model combines those embeddings with established credit, merchant and user features to produce the risk score.

That hybrid detail matters. Affirm says the transformer-only candidate never beat its best XGBoost model by enough to justify production. The shipped architecture kept the existing decision framework and added a learned representation where it earned measurable value.

Against the same production baseline, the hybrid delivered 1.8 times the FPDQ60 AUC gain of Affirm's next XGBoost iteration on held-out loans. For thin-file consumers, the reported multiple was 2.1 times. Those numbers describe relative gains in risk ranking. They do not mean 2.1 times fewer delinquencies or 2.1 times more approvals.

The online test answered a narrower question

Affirm first used the new score as a second look. It could incrementally approve eligible applications that the existing system had declined, but it could not decline applications the existing system had approved.

Across eligible new users, Affirm reports a 1.2 percentage-point conversion increase, or 3.4% relative, with the additional loans performing better than a comparable expansion under previous models. The launch release presents that as evidence for saying yes to more people at comparable risk.

The public material does not separately disclose online conversion, repayment or adverse-decision outcomes for the thin-file cohort. That is not proof of a problem. It is a boundary on what this launch evidence can establish.

Preserve the cohort from approval to monitoring

In a hypothetical model review, I would attach four fields to the release case: the cohort definition, the offline metric, the online decision policy and the later real-world outcome. If the subgroup cannot be reconstructed from production logs, the launch claim has no durable audit trail.

The monitor should show the cohort's volume as well as its rate. A stable percentage based on a shrinking population can look reassuring while becoming statistically weak. A pooled result can look stable while the launch-defining slice moves in the opposite direction.

This is editorial judgment, not a rule stated by Affirm or a claim that I tested its system. It is consistent with the US banking agencies' April 2026 model-risk guidance, which calls for outcome analysis and ongoing monitoring as products, exposures, clients, data and market conditions change. The guidance does not prescribe this exact cohort dashboard.

I would set the rollback condition before launch: minimum cohort volume, an agreed performance floor, an explanation-quality check and the action taken when any one fails. That makes the subgroup a release control instead of an interesting chart in the research notebook.

What's in it for you

  • Carry the business-case cohort into the production event schema and dashboard.
  • Report its denominator, decision policy and realised outcomes beside the portfolio average.
  • Pre-commit the alert, review and rollback path before the subgroup evidence deteriorates.

If one cohort earns the model its release, that cohort deserves its own production monitor.

#machine-learning#model-monitoring#fintech#evaluation#underwriting
← older drop
System 1 vs System 2 starts with the shape of the work

related drops

explore all 351 drops →
← back to the archiveday 117