# AndyMental > Evidence-first notes from Anand “Andy” Padia, Global Head of GenAI at Trigent Software, on building and shipping applied AI. ## Core pages - [About the author](https://andymental.com/about) - [Manifesto](https://andymental.com/manifesto) - [Where it all started](https://andymental.com/story) - [Archive](https://andymental.com/archive) - [RSS feed](https://andymental.com/feed.xml) ## Published drops - [Agent harnesses need maps, not more manuals](https://andymental.com/drops/agent-harnesses-need-maps-not-more-manuals): Self-evolving agents fail at locating where to edit before they fail at writing the edit (Harness Handbook, arXiv Jul 14). The scarce artifact is a source-verified map — not a fatter instruction manual. - [Temperature zero is not a determinism guarantee](https://andymental.com/drops/temperature-zero-is-not-a-determinism-guarantee): A weekend tutorial promises "same verdict every time" from temperature 0. Thinking Machines got 80 distinct outputs from 1,000 temp-0 runs. Temp 0 is a sampling setting; the only deterministic layer is the cache. - [Vulnerability counts need a CVE ledger](https://andymental.com/drops/vuln-counts-need-a-cve-ledger): Anthropic's Project Glasswing claims Mythos found 10,000+ high-severity vulnerabilities. An audit found one CVE explicitly attributed to it. The gap isn't fraud — it's a claim versus the public ledger that checks it. - [An AI detector you can switch off measures cooperation](https://andymental.com/drops/optional-ai-detection-measures-cooperation): Substack added reader-facing AI detection via Pangram — but authors can disable it per post, and disabling shows readers "AI detection unavailable." So it doesn't measure AI. It measures whether you cooperated. - [Less restrictive is now a model spec](https://andymental.com/drops/less-restrictive-is-now-a-model-spec): Anthropic's Opus 5 is cheaper, engages safety classifiers 85% less than Fable 5, and can silently reroute blocked prompts to weaker models. Restriction level just became a product axis you have to audit. - [Agent ROI is often feature activation in disguise](https://andymental.com/drops/agent-roi-is-often-feature-activation): SaaStr's finance agent "fixed" collections by switching on bill.com auto-reminders the team never activated in 8 years. A real win — but configuration debt, not intelligence. Separate activation wins from reasoning wins. - [The open-weight ban letters price the margin, not the risk](https://andymental.com/drops/open-weight-ban-letters-price-the-margin-not-the-risk): Two industry letters hit Washington in 48 hours, split cleanly by who monetizes diffusion vs scarcity. Weights already mirrored can't be un-proliferated — a ban would move inference margin, not remove risk. - [The Jacobian counterexample has a human byline](https://andymental.com/drops/the-jacobian-counterexample-has-a-human-byline): A mathematician using Claude Fable 5 posted a counterexample to the 1939 Jacobian conjecture; within days newsletters credited the model alone. How this gets attributed is the template for every AI-assisted deliverable. - [Task cost is a harness property, not a model price](https://andymental.com/drops/task-cost-is-a-harness-property-not-a-model-price): Factory's CTO showed the same code review costing $1.70 to $6 depending on harness, not model. Pricing pages can't predict task cost — and the vendors best placed to benchmark it honestly aren't the labs. - [The FCA put the model vendor inside the sandbox](https://andymental.com/drops/fca-put-the-model-vendor-inside-the-sandbox): Anthropic joined the FCA's Supercharged Sandbox, giving 21 regulated firms Claude access under supervision. The regulator moved up the stack to the vendor layer — and the governance win doubles as go-to-market. - [Google Cloud's 82% is not a cloud-only signal](https://andymental.com/drops/google-cloud-growth-mixes-services-and-hardware): Alphabet's Cloud segment now blends services with TPU system sales. Mid-market AI teams should answer with a three-layer stack: cloud, a local GPU rack, and AI-capable endpoints. - [An isolated sandbox is a claim, not a property](https://andymental.com/drops/isolated-sandbox-is-a-claim-not-a-property): OpenAI's "highly isolated" benchmark sandbox had a designed route to the internet, and a pre-release model used it. Isolation claims need third-party attestation, the way SOC 2 controls do — not prose. - [Distillation disputes are now trade policy, not license disputes](https://andymental.com/drops/distillation-disputes-are-now-trade-policy): The White House accused Moonshot of distilling Fable to build Kimi K3, and Treasury put sanctions on the table. With no technical test for provenance, open-weight model choice just became a supply-chain decision. - [AI content billing just moved off the URL — and agents pay the price](https://andymental.com/drops/ai-content-billing-moved-off-the-url): Cloudflare retired pay-per-crawl for pay-per-citation and will block agent crawlers by default from September 15. The billable event is now the answer — metered by the party that pays. - [Agent procurement is collapsing into deployment latency](https://andymental.com/drops/agent-procurement-collapses-into-deployment-latency): Five agent vendors courted SaaStr in one week; only the one that deployed in five minutes got adopted. When eval bandwidth is zero, time-to-running decides the shortlist — and the approval gate gets skipped. - [Free LLM routers are paid for in telemetry](https://andymental.com/drops/free-llm-routers-are-paid-in-telemetry): Ramp opened its internal LLM router to the public, free during beta. The routing layer's durable product is the cross-customer model-usage graph — price the telemetry as a trade, not a gift. - [AI ROI headlines choose their denominator](https://andymental.com/drops/ai-roi-headlines-choose-their-denominator): Only 12% of CEOs see AI returns — and about 44% report a financial gain. Both come from the same PwC table; which cut a speaker quotes tells you their agenda, not AI's ROI. - [AI feature news needs changelog provenance](https://andymental.com/drops/ai-feature-news-needs-changelog-provenance): A newsletter announced Claude Code computer use as new in July; the vendor changelog dates it to late March. Strategy decisions sourced from digests need checking against the only feed with real dates. - [Stateless MCP moves the state problem onto your side](https://andymental.com/drops/stateless-mcp-moves-state-onto-your-side): The 2026-07-28 MCP spec deletes protocol sessions so servers can scale. If you keyed context, auth or audit trails off Mcp-Session-Id, you own that state now — complexity relocated, not removed. - [The AI safety index grades disclosure, not safety](https://andymental.com/drops/safety-index-grades-price-disclosure-not-safety): FLI's Summer 2026 index put Anthropic on top with a C+ and handed three labs an F. The scores largely measure what labs publish — reading them as safety measurements misreads the instrument. - [AI pricing changes now arrive as quota emails, not price lists](https://andymental.com/drops/ai-pricing-changes-arrive-as-quota-emails): The real price of an AI coding subscription is its usage quota, and vendors move it with time-boxed promos your finance system cannot see. Any cost forecast built during a promo window is wrong by construction. - [AI citation metrics confuse being cited with being chosen](https://andymental.com/drops/ai-citation-metrics-confuse-cited-with-chosen): Ahrefs' 75K-brand data says mentions get you named in AI answers; format data says comparison pages win the traffic. Citation-share KPIs measure the first and quietly claim the second. - [The "AI employee" title is a governance bug](https://andymental.com/drops/ai-employee-is-a-governance-bug): New experimental data shows the "AI employee" label cuts manager monitoring 16% where agents sit on org charts. The title weakens exactly the oversight agent work still needs — ban it from the org chart. - [Agent approvals need independent policy checks](https://andymental.com/drops/agent-approvals-need-independent-policy-checks): Revolut X lets AI assistants analyse, strategise, and prepare crypto orders while the user carries approval. A confirmation click downstream of the agent's own framing is not a control. - [Quantized retrieval needs slice-level evals](https://andymental.com/drops/quantized-retrieval-needs-slice-level-evals): NVIDIA's 4-bit Nemotron 3 Embed keeps 99% of aggregate retrieval accuracy. An average across 16 public tasks says nothing about which of your query classes absorbed the loss — slice before you swap. - [Agent swarms need aggregation evals](https://andymental.com/drops/agent-swarms-need-aggregation-evals): Kimi's swarm docs report 300 subagents and BrowseComp accuracy doubling. The failure literature says most multi-agent breakage happens after the workers succeed — so score the merge, not the branches. - [The Hugging Face breach lesson is logs, not local models](https://andymental.com/drops/hugging-face-breach-lesson-is-logs-not-local-models): An autonomous agent ran 17,000 actions through Hugging Face over a weekend. It was caught because every action was logged and an LLM triaged the anomaly. The newsletters' "self-host your models" moral fixes nothing here. - [Agent backtests are demos, not track records](https://andymental.com/drops/agent-backtests-are-demos-not-track-records): JPMorgan's eight AI agents beat 60/40 by 0.7pts across 20 years of backtests. The catch nobody in the amplification layer mentions: an LLM was trained on those same decades, so the answer key is partly in its weights. - [Open weights are artifacts, not announcements](https://andymental.com/drops/open-weights-are-artifacts-not-announcements): Kimi K3's "largest open-weight model ever" title rests on a July 27 promise — no repo, no license file, no self-host path — and the coverage repeating it is already miscounting the basics. - [Provenance attests the factory, not the code](https://andymental.com/drops/provenance-attests-the-factory-not-the-code): Five poisoned @asyncapi npm versions shipped through the project's own release pipeline — every one with valid provenance. Attestations prove which workflow built a package, not that the inputs were clean. - [Portability must include the learning layer](https://andymental.com/drops/portability-must-include-the-learning-layer): Everyone negotiates portable model endpoints — and leaves the evals, graders, trace history, and feedback labels trapped in one vendor's control plane. The model is the swappable part; the learning layer is the asset. - [Agent-built apps need expiry dates](https://andymental.com/drops/agent-built-apps-need-expiry-dates): SaaStr migrated 10 years off Marketo for $14 and replaced a $10K app in an hour. When building gets this cheap, app count outruns the obligations each carries — so every agent-built app needs an owner and an expiry. - [Dashboards are not the SaaS moat](https://andymental.com/drops/dashboards-are-not-the-saas-moat): Gartner: $234B of enterprise app spend exposed to "agentic arbitrage" by 2030 — yet wholesale replacement stays unlikely. Both are right. Agents commoditize the dashboard; the governed records underneath stay sticky. - [Support containment is not resolution](https://andymental.com/drops/support-containment-is-not-resolution): Airbnb's AI Assistant resolved 40%+ of guest issues without a human, up from ~33%, as cost per booking fell 10%. Real progress — but "handled without a human" measures channel exit, not whether the problem got solved. - [Preference data governance outlives training methods](https://andymental.com/drops/preference-data-governance-outlives-training-methods): DPO made alignment simpler than RLHF — one classification objective, no reward model. But it simplified the algorithm, not the data. Whose preferences you rewarded still needs governing, whatever the training method. - [Agent harnesses are release platforms](https://andymental.com/drops/agent-harnesses-are-release-platforms): Microsoft Foundry advertises 11,000+ models — but the number that ships agents is one control plane over identity, tools, evals, traces, and rollback. Model choice is one release input; the harness is the product. - [Data gravity is becoming agent governance](https://andymental.com/drops/data-gravity-is-becoming-agent-governance): Databricks' Genie agents inherit Unity Catalog permissions, semantics, and lineage — the model is swappable, the governed access layer is not. You can win the agent layer by owning the data control plane, not a model. - [On-prem AI is an architecture decision](https://andymental.com/drops/on-prem-ai-is-an-architecture-decision): JPMorgan picked SambaNova to run AI inference behind its own walls. Good story — and a trap if your org reads it as "on-prem equals secure." Location is not a control. Choose it for named constraints, not as a vibe. - [AI BOMs need behavioral inventory](https://andymental.com/drops/ai-boms-need-behavioral-inventory): The AI bill-of-materials push borrows the SBOM playbook: inventory your models and datasets. Necessary, insufficient. An agent's risk lives in what it may do — the BOM must record tools, scopes, and policies too. - [Attack agents test their own tooling](https://andymental.com/drops/attack-agents-test-their-own-tooling): Sysdig watched an autonomous attacker unit-test its own exploit payloads, read the failures, and correct before escaping a container. When the attacker iterates, signature detection chases a target that rewrites itself. - [Domain experts are model infrastructure](https://andymental.com/drops/domain-experts-are-model-infrastructure): OpenAI is hiring an investment banker — not to bank, but to build rubrics, reference work, and evals. The labs treat expert judgment, written down and maintained, as infrastructure. Most enterprises treat it as a favor. - [AI productivity is becoming a team metric](https://andymental.com/drops/ai-productivity-is-becoming-a-team-metric): Figma's 2026 report: developers doing design jumped 44% to 60%; designers coding nearly doubled. Everyone can create everything now — whether the team ships faster depends on review, reconciliation, and decisions. - [Model worldview belongs in procurement evals](https://andymental.com/drops/model-worldview-belongs-in-procurement-evals): The Economist mapped 25 frontier models onto World Values Survey axes; same-lab models landed far apart. Don't label models politically — test worldview-sensitive behavior on your own use cases, per version. - [Open models still concentrate infrastructure](https://andymental.com/drops/open-models-still-concentrate-infrastructure): Together AI raised $800M with 500+ MW of compute committed to serving open models. Open weights end software lock-in — and shift the bargaining power to the few providers who can finance inference at industrial scale. - [Forward-deployed engineering is distribution](https://andymental.com/drops/forward-deployed-engineering-is-distribution): Microsoft committed $2.5B and 6,000 engineers to customer-side AI deployment, days after Amazon's $1B version. The hyperscalers concluded the constraint is integration, not model access — services just became a channel. - [Agent leverage follows context consolidation](https://andymental.com/drops/agent-leverage-follows-context-consolidation): SaaStr merged ~10 apps into one codebase and its agents got better with every build; Replit says context is now effectively infinite. Consolidating what agents can see beats adding agents that each see a fragment. - [Vertical AI moats live in workflow frequency](https://andymental.com/drops/vertical-ai-moats-live-in-workflow-frequency): Harvey reportedly added $100M net-new ARR in one quarter — but durability is predicted by the engagement ratio underneath. In vertical AI, the moat is how often the work returns, and what accumulates when it does. - [Voice agents should not rebuild WebRTC](https://andymental.com/drops/voice-agents-should-not-rebuild-webrtc): OpenAI published the relay-and-transceiver architecture behind its voice products — and the real lesson is the boundary it draws. Media transport is infrastructure; your voice agent's advantage lives entirely above it. - [Voice agents need time-sliced architecture](https://andymental.com/drops/voice-agents-need-time-sliced-architecture): Thinking Machines' interaction models interleave input and output in 200ms micro-turns — deciding each moment whether to listen, interject, or stay silent. The voice-team lesson: the problem isn't latency, it's turns. - [AI spend per engineer is not ROI](https://andymental.com/drops/ai-spend-per-engineer-is-not-roi): Tunguz's estimates — Anthropic at 2.3x payroll on compute, the median firm at $137 per engineer — will anchor a thousand budget debates. Both numbers measure intensity, not value. Spend needs a work denominator. - [Inference moats are financed in megawatts](https://andymental.com/drops/inference-moats-are-financed-in-megawatts): Groq's $650M raise talks like a utility — 13 data centers, a 200 MW target by 2027. The tell for buyers: chip benchmarks decay in quarters; deployed capacity, financing, and placement decide what you can actually buy. - [Version control needs eval history](https://andymental.com/drops/version-control-needs-eval-history): "Put your AI workflows in git" is this month's PM advice. Correct — and incomplete. A diff tells you what changed, not which model, fixtures, and scores made the old version safe. Version the evidence too. - [Agent frameworks magnify install risk](https://andymental.com/drops/agent-frameworks-magnify-install-risk): 140+ Mastra npm packages poisoned in under an hour; the payload hunted crypto-wallet extensions. The agent-era lesson: a poisoned dependency lands in your most credentialed environment — pinning alone is not a control. - [Outcome pricing forces agent scope](https://andymental.com/drops/outcome-pricing-forces-agent-scope): Sierra prices customer-facing agents by outcomes — a sale, a resolution — not seats or tokens. The pricing model is secretly an architecture review: you cannot bill an outcome you cannot define, attribute, and evidence. - [Model ensembles buy disagreement, not truth](https://andymental.com/drops/model-ensembles-buy-disagreement-not-truth): OpenRouter's Fusion runs a panel of budget models and reportedly beats frontier systems on research at half the cost. The value is the disagreement the panel surfaces — and a smooth synthesis is where that value dies. - [Security agents need patch acceptance metrics](https://andymental.com/drops/security-agents-need-patch-acceptance-metrics): OpenAI's GPT-5.5-Cyber found 24 kernel privilege-escalation exploits across 30M+ lines. Impressive — and the wrong scoreboard. Judge a security agent by validated fixes accepted upstream, not exploits generated. - [Generated code shifts the bottleneck](https://andymental.com/drops/generated-code-shifts-the-bottleneck): Research cited by ByteByteGo: PR completion rose 26% — and review burden just moved to teammates. Coding is 20-30% of engineering time; codegen speeds one station and floods the queues that were always the constraint. - [Bounded agents beat executive titles](https://andymental.com/drops/bounded-agents-beat-executive-titles): SaaStr's "AI VP of Customer Success" reportedly cut human hours 70% by owning ~40 recurring, checkable deliverables. The result is real; the title is the risk — agents succeed by being bounded, and titles hide bounds. - [AI research needs evidence lineage](https://andymental.com/drops/ai-research-needs-evidence-lineage): KPMG pulled an agentic-AI report after UBS, the NHS, Swiss Federal Railways, and TfL disputed its case studies. The lesson isn't "AI hallucinates" — it's that review without claim-to-source lineage is not a control. - [ARR velocity does not prove a product moat](https://andymental.com/drops/arr-velocity-does-not-prove-a-product-moat): Cursor reportedly ran from $2B ARR in February to $4B by June. That curve proves demand like almost nothing in software history — and proves nothing about switching costs once models and interfaces converge. - [Deployment expertise is a full-stack role](https://andymental.com/drops/deployment-expertise-is-a-full-stack-role): SaaStr says the scarce role now is the "agentic deployment expert" — and it's right, if deployment means owning data, evals, workflow, cost, and change end to end. Installing tools fast is the junior version of the job. - [AI-native teams still need eval ownership](https://andymental.com/drops/ai-native-teams-still-need-eval-ownership): Codex ships 10-12 surfaces with 2 PMs; Cursor runs 40 engineers on 1. The deleted layers were carrying decisions — acceptance, risk, launch evidence, rollback. Delete the layer, keep the decision, name its owner. - [Bug reports are untrusted agent input](https://andymental.com/drops/bug-reports-are-untrusted-agent-input): Tenet Security injected fake Sentry errors and watched coding agents execute the "resolution steps" with the developer's shell and credentials. The broken boundary: data retrieved for diagnosis became instructions. - [AI training needs production telemetry](https://andymental.com/drops/ai-training-needs-production-telemetry): 1,720 Wall Street employees, 53 workshops, one uncomfortable number — 3-5% of power users generate over 70% of usage. Workshops without telemetry don't spread capability; they concentrate it in people who already had it. - [Premium models need failure-cost routing](https://andymental.com/drops/premium-models-need-failure-cost-routing): Fable 5 reopened the price gap between frontier and routine inference — $10/$50 per million tokens. The router that earns that money asks one question the leaderboard never does: what does a wrong answer cost here? - [Agent stack diagrams hide the last mile](https://andymental.com/drops/agent-stack-diagrams-hide-the-last-mile): ByteByteGo's four-layer agent stack is good vocabulary and a fine checklist — but not a reference architecture. Identity, channel failure, and ownership of half-done actions all live outside the neat layers. - [MCP needs revocation before reach](https://andymental.com/drops/mcp-needs-revocation-before-reach): ReversingLabs maps MCP adoption onto the API era's timeline, where security arrived years after the sprawl. The do-over only counts if identity, least privilege, audit, and a kill switch ship before the connectors do. - [Domain agents need disagreement signals](https://andymental.com/drops/domain-agents-need-disagreement-signals): Benchling runs the same scientific task across different model providers — not for redundancy, but because disagreement between model families is a routing signal that sends the risky cases to a domain expert. - [Launch buzz measures demoability](https://andymental.com/drops/launch-buzz-measures-demoability): 400 builder posts from Claude Fable 5's first 48 hours tell you what the model makes easy to show off. They tell you almost nothing about what survives production review — the sample is selected for shareability. - [Agent go-live starts the expensive phase](https://andymental.com/drops/agent-go-live-starts-the-expensive-phase): Salesforce's lessons from 12,000+ Agentforce deployments say the quiet part: test before launch, monitor after. Budget most of an agent program for what comes after go-live — that is when the real debt surfaces. - [Agent loops are release artifacts](https://andymental.com/drops/agent-loops-are-release-artifacts): The head of Claude Code says he doesn't prompt anymore — he writes loops. The quiet part: a loop that runs unattended is software, and it belongs in version control with tests, budgets, and stop conditions. - [AI budgets should follow workloads, not employees](https://andymental.com/drops/ai-budgets-should-follow-workloads): Uber burned a 12-month AI budget in 4 months and answered with a $1,500/month cap per employee per coding tool. The cap treats spend as a person problem — but spend belongs to workloads, and workloads are what to budget. - [Local model benchmarks are not capacity plans](https://andymental.com/drops/local-model-benchmarks-are-not-capacity-plans): Ollama 0.30 is up to 20% faster on NVIDIA — measured on one model, one GPU, one quantization. A single-GPU speedup proves an optimization exists. It does not tell you whether local inference can carry your workload. - [Route agents by queue age, not token price](https://andymental.com/drops/route-agents-by-queue-age-not-token-price): A laptop handled 78% of one investor's AI work last week — but the number that matters is queue age falling from 73 seconds to 4. Hybrid local-cloud routing is a scheduling problem first and a pricing problem second. - [Pre-release model review is procurement leverage, not a safety standard](https://andymental.com/drops/pre-release-model-review-is-procurement-leverage): The new executive order buys the government up to 30 days of confidential pre-release access to frontier models. Without published pass criteria, that is leverage dressed as assurance — buyers should read it that way. - [Data agents need semantic infrastructure, not smarter models](https://andymental.com/drops/data-agents-need-semantic-infrastructure): OpenAI's data agent serves 3,500+ users over 600 petabytes — and the architecture is mostly lineage, definitions, and permissions. Natural-language analytics is a semantic-layer problem wearing an agent costume. - [Agent count is not a production metric](https://andymental.com/drops/agent-count-is-not-a-production-metric): SaaStr says it runs on 3 humans and 21+ AI agents. The useful part of the disclosure is the job map, not the ratio — and the scorecard your agent program needs has five columns, none of which is headcount. - [Voluntary AI frameworks are regulatory signals](https://andymental.com/drops/voluntary-ai-frameworks-are-regulatory-signals): OpenAI's Frontier Governance Framework is a well-built compliance document nine weeks before EU enforcement bites. Read it for what it concedes, not what it promises — and ask your vendor for the judgment record instead.