~/
andymental
_
/
search the archive
⌘K
today · day
105
the archive
28 drops
← back to the feed
reading path
all reading paths
Agent engineering
Evaluations and reliability
Security and permissions
AI economics
Enterprise adoption
Search, discovery and media
path link →
Read model claims against tests, failure cases and reproducible evidence.
month
sep
aug
jul
jun
all
filtering by
#evals ✕
#evaluation ✕
#benchmarks ✕
#testing ✕
#reliability ✕
#reproducibility ✕
#quantization ✕
clear all ✕
POST
day 103
A plausible model explanation is not a root cause
2d ago
POST
day 99
The AGI declaration's receipts measure spend, not generality
6d ago
ESSAY
day 94
Quantization damage hides in the flips, not the average
11d ago
POST
day 84
Perfect accuracy cannot reveal the algorithm
3w ago
ESSAY
day 83
ARC-AGI-3's perfect score belongs to the harness
3w ago
ESSAY
day 75
The agent turf war happened in the lab, not the wild
4w ago
ESSAY
day 74
Agent autonomy is a timeout setting, not a capability
4w ago
ESSAY
day 72
QM makes the coding agent a swappable part
4w ago
TIL
day 68
Tool descriptions are not agent guardrails
5w ago
ESSAY
day 65
Guardrail benchmarks are graded before the attacker moves
5w ago
ESSAY
day 59
Model leaderboards can't see the harness
6w ago
POST
day 57
Agent harnesses need maps, not more manuals
6w ago
← prev
page 1 of 3
next →