← back to the archiveCover illustration for “Read an AI guarantee through its metric and its remedy”
POSTday 67·5w ago·by Andy Padia

Read an AI guarantee through its metric and its remedy

Cognition’s guarantee uses estimated engineering output and settles shortfalls in usage credits. Ask whether that metric and remedy address the risk your team wants covered.

Ten million dollars is the largest number in Cognition’s productivity guarantee. The words I would read first are “estimate” and “credits.” They explain what the promise actually does.

Cognition’s announcement describes an agent that estimates the useful engineering output of completed Devin sessions. The company says it validated the estimator against engineers’ time assessments. Near the end of an annual contract, estimated hours are converted to value using a standard rate and compared with consumption. A shortfall results in credits, up to the stated cap.

That is a financial commitment with a specific measurement and settlement mechanism. Calling it no commitment at all would be unfair. Treating it as cash reimbursement for any disappointing business outcome would be equally inaccurate.

Ask what the number stands in for

Estimated engineering hours answer a counterfactual question: how long would a human have taken to produce comparable work? That is useful to investigate, but it is not the same as hours removed from payroll or additional customer value delivered.

Cognition makes part of this distinction itself. Its announcement says the estimator does not replace a fuller ROI assessment. It also excludes sessions with unmerged pull requests or otherwise unproductive output from useful output. Those qualifications make the offer more specific than a simple count of generated code.

I would still ask the customer’s engineering owner to inspect a sample of the estimates. Does a difficult-looking task receive a high time estimate even when it was low priority? Does a merged change require substantial review or later rework? How are disagreements handled? A sensible metric needs a way to examine the cases behind the total.

That review does not require assuming the vendor’s estimator is biased. It recognises that the measurement is consequential and that both parties should understand how it maps to the work they value.

Choose a remedy that helps with the actual shortfall

Consider a hypothetical team whose main constraint is senior review capacity. Additional agent usage could be valuable if it produces better accepted work with the same review effort. It could also be difficult to consume if the team already has more proposed changes than it can assess.

In that situation, credits and cash solve different problems. Credits reduce the price of further use. They do not automatically supply the missing reviewer or compensate for every delay. The buyer should decide whether more of the service is a useful remedy under the failure scenario it is most concerned about.

My proposed procurement exercise would define that scenario before celebrating the headline cap. Pick a representative completed task, reconstruct the estimated value and discuss what would happen if the agreed measurement showed a shortfall. Keep the commercial terms and the technical acceptance standard in the same conversation.

I have not audited the estimator, reviewed a customer’s executed agreement or independently measured Devin’s productivity. The public announcement is a description of the offer; the applicable contract determines the actual entitlement and process. I would not infer undisclosed guarantees from the marketing summary.

The important change is that AI suppliers are putting more explicit promises around outcomes. That can improve the discussion, provided customers examine the proxy and the settlement rather than borrowing confidence from the largest number.

An AI guarantee is useful when its measurement reflects work you value and its remedy helps with the shortfall you actually face.

#ai-roi#procurement#coding-agents#measurement
← older drop
Distilled models inherit their teacher's provenance
newer drop →
Price the time it takes to believe an agent is done

related drops

explore all 243 drops →
← back to the archiveday 106