← back to the archiveCover illustration for “PRAGMA’s strongest evidence is a benchmark, not a moat”
POSTday 87·2w ago·by Andy Padia

PRAGMA’s strongest evidence is a benchmark, not a moat

Revolut’s model paper supports a shared representation across banking tasks. Durable commercial advantage requires a different comparison.

PRAGMA’s paper reports a 130.2% improvement in credit-scoring PR-AUC over its internal task-specific baseline. That is a defined evaluation result. Rewriting it as 2.3 times better default detection quietly changes the metric.

The Revolut foundation-model paper describes training on 26 million user records and adapting a shared representation across several banking tasks. Its results are interesting precisely because they have an experimental shape: a model, a baseline, tasks and metrics. A claim about a durable competitive moat needs more evidence than that shape supplies.

I would put two separate questions in a review: does the model improve the decision, and can the organisation sustain an advantage after competitors respond?

A useful answer to the first does not settle the second. Nor does working with a common technology partner automatically settle either question against the company. Shared engineering expertise can coexist with distinctive data, execution and distribution. A byline cannot measure how much advantage remains.

Give the metric its original name

PR-AUC summarises a precision-recall curve across thresholds. A relative improvement in that area is not the same as a proportional increase in the number of defaults caught at a particular operational threshold.

That matters because someone must choose the threshold. In a hypothetical lending evaluation, I would want to know the outcomes at an acceptable approval rate and risk level, including how the result varies across relevant applicant groups. A headline about a curve does not choose that operating point for the business.

The paper’s reported improvement remains useful. It tells me where further investigation may be worthwhile. It does not permit me to insert an unmeasured business outcome into a financial plan.

Likewise, customer totals and training cohorts answer different questions. The organisation may serve people whose records were excluded, unavailable or outside the study period. The model assessment should use the cohort actually described by the experiment.

Test what would survive imitation

For a hypothetical enterprise data strategy, my first comparison would be deliberately ordinary. Hold the evaluation cohort and decision target fixed. Compare the existing baseline with the proposed representation and record which additional signals contribute to the improvement.

Then ask what keeps those signals useful. Is the organisation observing an outcome that rivals cannot readily observe? Does it get that outcome earlier? Can it lawfully reuse the data for this purpose? How expensive is maintaining the linkage as products and schemas change?

Those questions locate an advantage in a process that can be assessed. They also expose dependencies. If the gain comes mainly from a one-time cleanup that competitors can repeat, the improvement may still pay for itself. It just needs a business case based on execution rather than permanent exclusivity.

I would keep the research budget and the moat claim separate until both survive their own tests. Research can create valuable shared infrastructure even when the architecture is published. A proprietary dataset can remain expensive inventory if the labels arrive late or the operational decision cannot use them.

The discipline here is to let a good paper be a good paper. Requiring it to prove an entire competitive strategy makes the marketing larger and the evidence less legible.

Preserve the benchmark’s metric, then test the durability of the advantage as a separate claim.

#fintech#models#benchmarks
← older drop
Jalapeño’s benchmark still needs a production-shaped comparison
newer drop →
Redacting the visible transcript does not clear the whole export

related drops

explore all 243 drops →
← back to the archiveday 106