
A tax-agent accuracy claim needs a unit of error
A reported 98 percent accuracy figure cannot size a review process until we know what was scored, which cases were included and when human corrections occurred.
A 98 percent accuracy figure leaves my main operating question unanswered: 98 percent of what?
Before turning a claim like that into an autonomy target, I would ask for the evaluation unit. A field, a calculation, a document and a completed return are different things to score. Their percentages cannot be exchanged without knowing how the result was constructed.
That is not a claim that the reported performance is poor. Faster preparation with reliable review could be a valuable improvement. The decision needs a description of the work that remains, not just the fraction labelled correct.
The denominator determines the review queue
Consider a hypothetical evaluation of 10,000 completed returns. If 98 percent of whole returns met a defined acceptance standard, the remaining 200 would need some form of investigation or correction. That arithmetic describes the hypothetical test, not the reported platform's actual error count.
Now change the unit to fields within returns. The same percentage no longer tells us how many returns contain at least one error. Errors may cluster in a few difficult documents or appear across many. Without that distribution, claiming that “most returns are wrong” would be another unsupported leap.
Severity changes the interpretation too. A minor presentation issue and an incorrect amount can both count as one failed item while requiring very different responses. A useful evaluation keeps the error categories visible rather than compressing every consequence into one average.
I would also ask where human correction sits. Was the score measured on the agent's first output, after review or on the final delivered work? A high final quality score may demonstrate a strong combined process without establishing that the agent can operate unattended.
That combined process can be exactly what the business wants. It should simply be described as such.
Measure the system that will actually run
For a proposed accounting workflow, I would define the eligible work and acceptance standard with qualified practitioners, then test representative cases through the intended review process. Include exceptions and missing information, not only complete, clean examples.
Record preparation time and correction time separately. A fast first draft that requires extensive checking may still help, but the total benefit has to include the checking. Likewise, a slower draft that makes its evidence easy to inspect may reduce the time needed for acceptance.
The operational record should show which errors the process catches, which escape its checks and who handles unresolved cases. That is how an accuracy result becomes a staffing and quality decision rather than a decorative percentage.
This is an evaluation proposal, not tax advice or an estimate of legal liability. The obligations and consequences for actual filings require the relevant professional judgment. A generic accuracy number cannot supply it.
I would be happy to fund a system that makes expert-reviewed work faster and more reliable. I would want the evidence to describe that achievement directly, with its human contribution intact.
Ask an accuracy claim for its unit, error distribution and review stage before using it to decide how much work an agent may perform alone.


