# AndyMental > Evidence-first notes from Anand “Andy” Padia, Global Head of GenAI at Trigent Software, on building and shipping applied AI. ## Core pages - [About the author](https://andymental.com/about) - [Manifesto](https://andymental.com/manifesto) - [Where it all started](https://andymental.com/story) - [Archive](https://andymental.com/archive) - [RSS feed](https://andymental.com/feed.xml) ## Published drops - [An agent gateway needs a bypass budget](https://andymental.com/drops/an-agent-gateway-needs-a-bypass-budget) — published 2026-10-01: An agent gateway becomes a control only when teams measure direct-call bypasses, credential exceptions, ownership gaps and revocation time. - [Scientific AI needs the experiment trace](https://andymental.com/drops/scientific-ai-needs-the-experiment-trace) — published 2026-09-30: Periodic Labs shows why scientific AI needs stitched experiment lineage; its internal XRD results are promising, not proof of discovery. - [Long tool observations need a separate prefill budget](https://andymental.com/drops/long-observations-need-a-separate-prefill-budget) — published 2026-09-28: HySparse2 runs prefill through 25 of 49 layers, then uses the full model for decode. Measure long tool observations as a separate capacity workload. - [A cache selector needs a random baseline](https://andymental.com/drops/a-cache-selector-needs-a-random-baseline) — published 2026-09-27: Random Attention matched smart KV-cache selectors by protecting the prompt and sampling the rest. Make random eviction the null test for selector complexity. - [The presenter skill is a production contract, not a video engine](https://andymental.com/drops/presenter-skill-is-a-production-contract-not-a-video-engine) — published 2026-09-26: Lanshu does not contain a one-click presenter engine. Its useful asset is a provider-neutral production contract with permission, cost and delivery checks. - [A negative exploit window needs a pre-patch control](https://andymental.com/drops/a-negative-exploit-window-needs-a-pre-patch-control) — published 2026-09-26: When exploitation starts before a patch exists, a patch SLA is too late. Give every exposed critical dependency a tested containment action that can run first. - [Monitor the cohort that justified the model](https://andymental.com/drops/monitor-the-cohort-that-justified-the-model) — published 2026-09-25: Affirm's largest offline underwriting lift appeared for thin-file users. Keep that cohort visible in production monitoring instead of letting pooled outcomes erase the release case. - [System 1 vs System 2 starts with the shape of the work](https://andymental.com/drops/system-1-vs-system-2-route-by-dependencies) — published 2026-09-24: Jev makes fast judgments a practical building block. Choosing between System 1 and System 2 starts with task dependencies and available evidence, before a confidence threshold enters the picture. - [Jev explained and the rise of the decision layer](https://andymental.com/drops/jev-decision-model-agent-harness-use-cases) — published 2026-09-24: What Jev does, why its speed claims need context, and where it fits in coding agents and marketing. A practical guide to the decision layer, its harness and 20 community projects. - [An adaptive agent environment is a release artifact](https://andymental.com/drops/an-adaptive-agent-environment-is-a-release-artifact) — published 2026-09-24: If an agent's start state, observations or allowed actions can change, version that environment stack with the model and re-run the untouched control before promotion. - [Jev's 444x claim prices agreement, not correctness](https://andymental.com/drops/jevs-444x-claim-prices-agreement-not-correctness) — published 2026-09-23: TypeSafe's Jev benchmark is fast and cheap against an Astra/Fable reference. That result earns a matched workload test, not a production correctness claim. - [AI confidence needs an operator ledger](https://andymental.com/drops/ai-confidence-needs-an-operator-ledger) — published 2026-09-23: Executives and delivery teams need one reconciled record of deployments, rollbacks, recovery and safety rework. Confidence without operational evidence is distance. - [Security evals need a real egress deny](https://andymental.com/drops/security-evals-need-a-real-egress-deny) — published 2026-09-23: Agent security evaluations need default-deny egress, synthetic targets, disposable credentials and a tested kill switch. Prompted restraint is not containment. - [AI code needs a comprehension budget](https://andymental.com/drops/ai-code-needs-a-comprehension-budget) — published 2026-09-23: Generated code can scale faster than accountable understanding. Treat comprehension as a finite budget spent by unfamiliar change, weak evidence and unclear rollback paths. - [Agentic AI should return attention, not just minutes](https://andymental.com/drops/agentic-ai-should-return-attention-not-just-minutes) — published 2026-09-23: Pascal Bornet's cross-domain comparison offers a better automation question: where did the saved attention go, and did people become more present for work that needs trust? - [Price AI work by accepted outcome, not by token](https://andymental.com/drops/price-ai-work-by-accepted-outcome-not-by-token) — published 2026-09-23: Token spend is an input price. Useful AI unit economics joins model, tool, infrastructure, retry and review costs to an accepted unit of business work. - [A hallucination rate is not a launch criterion](https://andymental.com/drops/hallucination-rate-is-not-a-launch-criterion) — published 2026-09-23: A single hallucination percentage hides the product decision. Consequence, detectability, recovery and propagation determine whether an AI failure is tolerable. - [GPT-Live splits the latency budget in two](https://andymental.com/drops/gpt-live-splits-the-latency-budget-in-two) — published 2026-09-23: GPT-Live's useful systems lesson is not simply full-duplex audio. It is the separation of a punctual live path from a slower reasoning path, with a different promise for each. - [When should an LLM judge block an AI release?](https://andymental.com/drops/an-llm-judge-needs-an-error-budget-before-it-blocks-a-release) — published 2026-09-22: Agreement can qualify an LLM judge to measure quality. A release gate also needs separate false-allow and false-block budgets, deterministic checks, and a declared fallback. - [Constrained JSON does not constrain the tool choice](https://andymental.com/drops/constrained-json-does-not-constrain-the-tool-choice) — published 2026-09-22: Needle 3 makes structured output small and fast. Its own demo also shows why parseable JSON is not proof that the model selected the right tool. - [Local AI-agent observability needs a rollback path](https://andymental.com/drops/local-observability-mcp-proxy-needs-a-rollback) — published 2026-09-20: Project Telescope observes AI-agent tool traffic at the MCP boundary. Pin the experimental release, snapshot configs, and prove rollback before letting a proxy rewrite your setup. - [Record/replay keeps fast-moving AI API tests honest](https://andymental.com/drops/record-replay-keeps-ai-api-contract-tests-honest) — published 2026-09-18: Mocks can stay green after an AI dependency changes. Run the real server in setup, replay committed fixtures by default, and make fixture refresh an explicit compatibility check. - [When you actually need an AI agent instead of a workflow](https://andymental.com/drops/when-to-use-an-ai-agent-instead-of-a-workflow) — published 2026-09-18: Most 'agent problems' are workflow problems. The test: if you can draw the steps before you run it, build a workflow. Reach for an agent only when the path is genuinely un-enumerable — and expect to pay for it in evals. - [A long-horizon agent needs a definition of finished](https://andymental.com/drops/a-long-horizon-agent-needs-a-definition-of-finished) — published 2026-09-17: Memory and durable execution can keep an agent working for weeks. They do not define when the business outcome is complete, accepted or uncertain. - [A viral video chapter needs a timestamp check](https://andymental.com/drops/a-viral-video-chapter-needs-a-timestamp-check) — published 2026-09-17: A viral Jeff Dean repost points viewers to the wrong subjects at three timestamps. Treat chapter lists as factual claims and verify them against the original recording. - [Context compaction needs a recovery test](https://andymental.com/drops/context-compaction-needs-a-recovery-test) — published 2026-09-16: Compaction can keep a long-running agent inside its context budget while changing what it remembers. Verify recovery of critical state, not just the smaller token count. - [Measure voice latency with paired turn events](https://andymental.com/drops/measure-voice-latency-with-paired-turn-events) — published 2026-09-15: A fast first token is only one part of a voice turn. Use paired timestamps to separate model delay, audio readiness and the wait your caller actually experiences. - [Fine-tune to change behaviour, use RAG to change knowledge](https://andymental.com/drops/fine-tuning-changes-behaviour-rag-changes-knowledge) — published 2026-09-15: RAG and fine-tuning solve different failures: retrieval changes what a model knows, while fine-tuning changes how it behaves. Most ‘fine-tune it on our docs’ requests are retrieval problems. - [A financial citation needs a data-rights clock](https://andymental.com/drops/a-financial-citation-needs-a-data-rights-clock) — published 2026-09-15: ChatGPT for Financial Services can cite premium datasets, but a banker still needs the data's reporting time, access basis and sharing rights before a number enters client work. - [The lottery ticket still has to pay for the search](https://andymental.com/drops/the-lottery-ticket-still-has-to-pay-for-the-search) — published 2026-09-13: A trainable sparse subnetwork does not turn the discarded weight percentage into compute savings. Include discovery, retraining and actual execution before treating pruning as a deployment decision. - [Teach Grok Bot one exception before adding another persona](https://andymental.com/drops/teach-grok-bot-one-exception-before-adding-another-persona) — published 2026-09-13: This Grok Bot workshop shows how preferences become reusable instructions. Test whether a changed preference survives the next run before treating more named agents as a better workflow. - [Start the AI roadmap with Minsky's two problem lists](https://andymental.com/drops/start-the-ai-roadmap-with-minskys-two-problem-lists) — published 2026-09-13: A question in Marvin Minsky's MIT homework separates problems AI already addresses from problems it might help solve. Use that distinction to separate delivery commitments from research bets. - [Self-maintaining APIs should sell the accepted migration](https://andymental.com/drops/self-maintaining-apis-should-sell-the-accepted-migration) — published 2026-09-13: YC's API-maintenance idea has a useful starting point: turn a real breaking change into a migration a customer accepts. Patch generation alone does not establish demand or successful adoption. - [Repurpose the evidence before rewriting the LinkedIn hook](https://andymental.com/drops/repurpose-the-evidence-before-rewriting-the-linkedin-hook) — published 2026-09-13: This LinkedIn skill bundle can help turn long-form material into posts. Preserve the claim, evidence and caveat together before optimising the opening for attention. - [A verification skill needs a map of the product](https://andymental.com/drops/a-verification-skill-needs-a-map-of-the-product) — published 2026-09-13: Lauren Tan's feature-map example explains why agents can operate a browser and still fail to verify a product. Map user-visible features to entry points, controls and observable outcomes. - [Test the dub, not the language catalogue](https://andymental.com/drops/a-language-catalogue-is-not-a-dubbing-acceptance-test) — published 2026-09-13: VoiceStudio brings local speech generation and dubbing into one workspace. Test a complete language pair before treating its catalogue as proof that it can replace your voice production workflow. - [Put an audience decision between the YouTube prompts](https://andymental.com/drops/youtube-prompt-packs-need-editorial-decisions-between-steps) — published 2026-09-12: A prompt sequence can carry an unsupported audience assumption from strategy to script. Require evidence of the viewer's problem before expanding production. - [The best move changes when the opponent model changes](https://andymental.com/drops/your-game-model-decides-what-a-good-move-means) — published 2026-09-12: Expectimax and minimax can prefer different actions in the same game. The assumption about the other player belongs in the design review before search optimisation. - [Open the startup archive at the decision you face](https://andymental.com/drops/use-the-startup-archive-as-a-decision-index) — published 2026-09-12: The 2014 startup course is easier to use as an index of decisions than as a completion challenge. Bring one current problem and check which assumptions have aged. - [Turn repeated review comments into executable checks](https://andymental.com/drops/turn-repeated-review-comments-into-executable-checks) — published 2026-09-12: Lauren Tan's session describes project-specific constraints enforced in CI. Repeated review feedback is a candidate for automation when the underlying rule can be stated precisely. - [The AI tutor study tested a designed lesson](https://andymental.com/drops/the-ai-tutor-study-tested-a-designed-lesson) — published 2026-09-12: A Harvard physics trial supports a carefully structured AI tutor, not a generic learning prompt. Read the lesson design before borrowing the headline result. - [Give the classroom a job the explanation cannot do alone](https://andymental.com/drops/the-ai-classroom-needs-work-worth-showing-up-for) — published 2026-09-12: A classroom discussion about AI raises a practical teaching challenge: use shared time for work, critique and revision whose quality students must defend. - [Count the model calls behind the thousand runners](https://andymental.com/drops/the-agent-marathon-keeps-physics-out-of-the-language-model) — published 2026-09-12: Google’s Race Condition workshop separates agent orchestration from model use. Its deterministic runner option makes the configuration more informative than the agent count. - [Test the weight that flips your decision](https://andymental.com/drops/test-the-weight-that-flips-your-decision) — published 2026-09-12: Emily Higgins demonstrates an AI-assisted weighted decision matrix. Its best output may be the uncertain preference that changes which option wins. - [The decision review should happen twice](https://andymental.com/drops/technical-leadership-needs-visible-decision-consequences) — published 2026-09-12: Steve Jobs’s 1992 MIT discussion makes a useful demand on technical leadership: stay close enough to a recommendation to learn from its consequences. - [Read the dissent before adopting the explanation](https://andymental.com/drops/read-an-argument-with-its-strongest-response) — published 2026-09-12: A paired argument and response can reveal which causal claim is actually disputed. Use the disagreement to design the next evidence check, rather than just choose a side. - [Put reasoning at the workflow's uncertain decisions](https://andymental.com/drops/put-reasoning-at-the-workflows-uncertain-decisions) — published 2026-09-12: An ADK workshop contrasts a single prompt with tools and explicit orchestration. Keep fixed calculations and sequencing in code, then give the model the decisions that need interpretation. - [Prototype a career change before renaming yourself](https://andymental.com/drops/prototype-a-career-change-before-renaming-yourself) — published 2026-09-12: A storiedcareers reel connects possible futures with professional identity. Turn the reflection into small tests of the everyday work before committing to the new title. - [Rehearse the decision the conversation needs](https://andymental.com/drops/practice-communication-around-one-real-conversation) — published 2026-09-12: Speaking, negotiation and storytelling meet in the same conversation. Practise them around a specific decision and the objection most likely to change it. - [A persistent desktop still needs to know what finished](https://andymental.com/drops/persistent-agent-desktops-need-a-recovery-story) — published 2026-09-12: Grok Bot’s persistent-computer demo raises a useful acceptance test: after an interruption, can the next worker separate completed, failed and uncertain actions? - [Keep the question attached to Munger’s answer](https://andymental.com/drops/munger-questions-are-more-useful-than-detached-aphorisms) — published 2026-09-12: A Munger Q&A is more useful when the question and his qualifications survive the clip. I would apply that discipline before turning any famous answer into operating advice. - [Automating the busy stage can make the queue worse](https://andymental.com/drops/locate-the-business-constraint-before-automating) — published 2026-09-12: Sharran Srivatsaa’s traffic, conversion and delivery framework gives an automation review a useful starting point: identify which stage currently limits accepted work. - [Recut one scene before collecting six editing techniques](https://andymental.com/drops/learn-video-editing-by-recutting-one-scene) — published 2026-09-12: Reusing the same footage lets you see how timing, sound and shot choice change meaning. A tutorial playlist makes those interactions easy to miss. - [Make the automation lesson fail once](https://andymental.com/drops/learn-automation-by-tracing-a-failed-handoff) — published 2026-09-12: Follow one event through a duplicate, missing data and a downstream failure. A recovery path teaches what a successful trigger-to-output demonstration leaves out. - [Let one small offer connect the skill stack](https://andymental.com/drops/learn-a-skill-stack-by-shipping-one-small-offer) — published 2026-09-12: Building, explaining and delivering one modest offer exposes the gaps between skills. A collection of courses does not show whether those skills work together. - [Judge human judgment by how it handles uncertainty](https://andymental.com/drops/judge-human-judgment-by-how-it-handles-uncertainty) — published 2026-09-12: Cleo Eleftheriades's fast skill cards offer a vocabulary for judgment. The practical test is whether someone updates a decision and owns what follows when the evidence changes. - [Give each AI subscription a job](https://andymental.com/drops/give-each-ai-channel-a-job-in-your-learning-plan) — published 2026-09-12: News, conceptual understanding and implementation need different outputs. Assign a purpose to an AI channel before adding it to your learning routine. - [Four plugins deserve four separate experiments](https://andymental.com/drops/four-coding-plugins-need-four-separate-value-tests) — published 2026-09-12: A coding add-on bundle mixes reuse, retrieval, compression and design guidance. Test each change separately before crediting the bundle with better code or a lower bill. - [Feedback should be reciprocal and voluntary](https://andymental.com/drops/feedback-should-be-reciprocal-and-voluntary) — published 2026-09-12: Robert S. Collins's Johari Window explanation is useful as a prompt for shared understanding. It becomes less useful when disclosure is treated as something to extract. - [Leave the contribution on the screen](https://andymental.com/drops/end-a-technical-talk-with-the-contribution) — published 2026-09-12: Patrick Winston’s closing-slide advice is a practical editing tool: keep the work, its evidence and its limits visible while the audience decides what to ask. - [Each video rehook should pay off the previous promise](https://andymental.com/drops/each-video-rehook-should-pay-off-the-previous-promise) — published 2026-09-12: Kallaway's short-video structure alternates body sections with renewed hooks. For technical explainers, each transition should follow a result the viewer can use. - [A success percentage needs a sentence defining success](https://andymental.com/drops/define-the-experiment-before-reporting-probability) — published 2026-09-12: An IIT Kharagpur probability lecture offers a practical correction for AI dashboards: define the experiment and event before comparing the percentages. - [The AI surplus already belongs to someone](https://andymental.com/drops/consumer-surplus-is-not-an-ai-sales-forecast) — published 2026-09-12: Stanford estimates substantial value to generative-AI users. Calling that value an unclaimed revenue pool changes the meaning of the measurement. - [Practise the recommendation before chasing confidence](https://andymental.com/drops/confidence-advice-needs-a-small-behavioral-test) — published 2026-09-12: A workplace confidence exercise should produce an observable action and specific feedback. A motivational watchlist cannot establish a personal transformation. - [I would review the problem before the AI demo](https://andymental.com/drops/choose-an-ai-method-before-choosing-the-product-wrapper) — published 2026-09-12: A CS221 recap offers a useful order for AI project reviews: model the problem, choose how to solve it, then identify what must be learned from data. - [Choose the productivity method after naming the failure](https://andymental.com/drops/choose-a-productivity-framework-by-the-failure-it-addresses) — published 2026-09-12: A lost note, an overloaded calendar and an avoided decision need different interventions. Diagnose the repeated failure before importing another productivity system. - [Show me what changed when the agent improved](https://andymental.com/drops/ask-what-changes-when-an-agent-self-improves) — published 2026-09-12: Stanford’s self-improving agents overview separates training from inference-time scaling. That distinction decides what to version, evaluate and carry into the next task. - [Add the operating cost to the open-source shortlist](https://andymental.com/drops/an-open-source-shortlist-needs-an-operating-cost-column) — published 2026-09-12: An open repository lets you inspect the software. Adoption still needs a clear account of model usage, data destinations, maintenance and the cost of leaving. - [Your weekly review should remove a commitment](https://andymental.com/drops/an-execution-system-should-change-next-weeks-commitments) — published 2026-09-12: Execution advice becomes useful when a missed commitment changes the next plan. Raising standards without changing capacity can simply repeat the same failure. - [An argument map needs evidence beneath the branches](https://andymental.com/drops/an-argument-map-needs-evidence-beneath-the-branches) — published 2026-09-12: AevyTV's reasoning-tool reel introduces fallacy practice and branching arguments. Use the map to organise evidence and unresolved checks, not to count rhetorical points. - [Put a review date beside the aphorism](https://andymental.com/drops/an-aphorism-needs-the-tradeoff-it-leaves-out) — published 2026-09-12: Naval's advice to act quickly and wait for results becomes more useful when you specify the evidence that would make you change course. - [Give the AI service story a second customer](https://andymental.com/drops/an-ai-service-story-needs-a-repeatable-customer-problem) — published 2026-09-12: Founder-led positioning becomes credible when the same problem can be solved again with visible delivery costs. A compelling first story cannot establish that repeatability. - [An AI roadmap should keep one project running](https://andymental.com/drops/an-ai-roadmap-should-keep-one-project-running) — published 2026-09-12: Divyam Dawar's fast course tour spans Python through deployment. Use one evolving project to connect those subjects and expose the failures between them. - [The agent office should open the diff](https://andymental.com/drops/an-agent-office-needs-reviewable-work-behind-the-avatars) — published 2026-09-12: Munder Difflin makes coding agents visible. The useful test is whether its office view shortens the path from apparent activity to a reviewable change. - [A technical story needs an unresolved question](https://andymental.com/drops/a-technical-story-needs-an-unresolved-question) — published 2026-09-12: Michelle Lo Horton's storytelling shortlist suggests a useful engineering exercise: structure an investigation around the question each piece of evidence helps resolve. - [A technical proposal needs a business consequence](https://andymental.com/drops/a-technical-proposal-needs-a-business-consequence) — published 2026-09-12: Bosky Mukherjee's commercial-skills reel offers a useful translation exercise: explain what an architecture decision changes for the business, with assumptions visible. - [The useful part of a 4 GB training claim is the recipe](https://andymental.com/drops/a-small-gpu-finetuning-claim-needs-its-full-configuration) — published 2026-09-12: Soup’s small-GPU demonstration is a reason to inspect the configuration. Hardware fit, reproducibility and useful model behaviour are three separate results. - [A skill-monetization prompt should end in a paid-problem test](https://andymental.com/drops/a-skill-monetization-prompt-should-end-in-a-paid-problem-test) — published 2026-09-12: Seven prompts can turn a work history into possible offers. The useful next step is to find whether a specific buyer has the problem those offers claim to solve. - [A sales process should make a clear no possible](https://andymental.com/drops/a-sales-process-should-make-a-clear-no-possible) — published 2026-09-12: A useful sales conversation clarifies fit and respects refusal. Test the exit path alongside the pitch, so persuasion produces a decision the buyer can actually own. - [A requested confidence number is not calibration](https://andymental.com/drops/a-requested-confidence-number-is-not-calibration) — published 2026-09-12: A Grok planning prompt asks for 95% confidence before proceeding. Replace that stopping rule with the missing facts and observations needed to choose a next action. - [A red-team prompt needs a falsifiable objection](https://andymental.com/drops/a-red-team-prompt-needs-a-falsifiable-objection) — published 2026-09-12: Sabrina Ramonov's prompt shorthands can request a different perspective. Make the critique useful by asking which premise is disputed and what evidence would settle it. - [Before I delegate, I want to know who pays the agent](https://andymental.com/drops/a-personal-agent-needs-an-explicit-interest-to-serve) — published 2026-09-12: A Reid Hoffman discussion raises the economics of personal agents. I would test those incentives with a decision where the user’s preference conflicts with the provider’s revenue. - [A moving dot is making a claim about time](https://andymental.com/drops/a-live-map-should-show-where-its-certainty-ends) — published 2026-09-12: God’s Eye View makes spatial data compelling. Its most useful design lesson is to distinguish observations, interpolation and unavailable information on the map itself. - [Give the free course a deliverable](https://andymental.com/drops/a-free-learning-path-needs-a-project-and-an-access-check) — published 2026-09-12: A learning list becomes a practical path when it names what you will produce and separates free access from the credential you might pay for. - [A Claude Code course needs a personal workflow map](https://andymental.com/drops/a-claude-code-course-needs-a-personal-workflow-map) — published 2026-09-12: A long Claude Code course covers many features. Make it useful by mapping those features onto one recurring task, its evidence and the checks that decide whether it is done. - [A brand challenge needs constructive proof](https://andymental.com/drops/a-brand-challenge-needs-constructive-proof) — published 2026-09-12: Divyangna's positioning exercise asks which category belief a brand challenges. For AndyMental, that position must produce useful tests and examples, not just objections. - [Business dramas need a pause before the payoff](https://andymental.com/drops/business-dramas-are-case-prompts-with-scripted-evidence) — published 2026-09-12: Use a business drama to open a decision exercise. Pause before the outcome, record what was knowable then, and require independent evidence before turning a scene into a lesson. - [A rebrand needs a change in the questions people bring](https://andymental.com/drops/a-rebrand-needs-evidence-for-the-new-association) — published 2026-09-12: A new bio states the reputation you want. Publish work that supports it, then check what readers actually infer. Caleb Ralston’s rebrand framework provides the starting diagnosis. - [No memo, no decision is an agent control pattern](https://andymental.com/drops/no-memo-no-decision-is-an-agent-control-pattern) — published 2026-09-11: A three-minute interview segment becomes a practical agent gate: record evidence, rejected options, blast radius and rollback before consequential execution. - [Ruflo's install paths are two different products](https://andymental.com/drops/ruflos-install-paths-are-two-different-products) — published 2026-09-11: Ruflo coordinates AI coding agents. Its plugin and full CLI installs grant different access; choose the smallest setup that fits your workflow. - [A reading inbox is not a publication queue](https://andymental.com/drops/a-reading-inbox-is-not-a-publication-queue) — published 2026-09-11: My WhatsApp self-chat held 168 unique links; only four had cleared analysis. Capture is inventory, not evidence. Put verification between save and publish. - [A plausible model explanation is not a root cause](https://andymental.com/drops/a-plausible-model-explanation-is-not-a-root-cause) — published 2026-09-11: Anthropic's CHIVE found that three activation-reading tools added no predictive lift over the transcript. Close an incident only after a targeted intervention moves the failure. - [Meta's agents need to show whose interests they represent](https://andymental.com/drops/meta-now-owns-both-agents-in-the-transaction) — published 2026-09-10: A personal agent and a business-agent team under one owner raise a practical question: can the buyer inspect whose objective shaped the recommendation? - [ElevenLabs keeps the commission in the agent business case](https://andymental.com/drops/elevenlabs-prices-agent-adoption-in-the-comp-plan) — published 2026-09-10: Paying the account owner when an agent closes revenue can buy cooperation. It also means the human commission belongs in the agent's operating cost. - [Wonder Pill makes assumptions visible before you pick an idea](https://andymental.com/drops/wonder-pill-makes-assumptions-visible) — published 2026-09-10: A useful brainstorming aid when the problem is still negotiable. Inspect the assumptions and rejected branches; don't mistake an interesting map for a tested plan. - [The 3x agent multiple is a night-shift ledger](https://andymental.com/drops/the-3x-agent-multiple-is-a-night-shift-ledger) — published 2026-09-10: OpenAI’s task success improved, but its 3.1× runtime measure also counts review agents. Connect accepted outcomes to the full cost before calling it productivity. - [Treat Suno's model swap as a training-data recall](https://andymental.com/drops/sunos-model-swap-is-the-first-training-data-recall) — published 2026-09-10: Suno will retire every pre-v6 model as its partner-built generation rolls out. Buyers need model lineage, swap notice, regression time and an exit drill. - [Astra verified the proof it did not discover](https://andymental.com/drops/astra-verified-the-proof-it-did-not-discover) — published 2026-09-10: OpenAI says a stronger internal model found its Navier–Stokes result; Astra formalized and verified it. Capability claims need stage-level receipts. - [Session inventory cannot explain token theft](https://andymental.com/drops/session-inventory-cannot-explain-token-theft) — published 2026-09-09: Claude can inventory sessions and manage Claude Code tokens, but aggregate quota cannot identify what spent it. Investigations need usage joined to an opaque authorization ID. - [The Formula E AI coach puts the inference location in focus](https://andymental.com/drops/the-cloud-agent-in-the-race-car-ran-without-the-cloud) — published 2026-09-07: Formula E's Goodwood account describes an on-device Gemma coach. The useful deployment lesson starts with where the decision runs and how late it can arrive. - [Forty-six percent of which PM jobs?](https://andymental.com/drops/the-46-percent-ai-pm-share-double-counts-the-market) — published 2026-09-07: LinkedIn search counts can reveal demand for AI skills. Turning them into a market share requires a consistent population and a classification rule. - [An agent payment should carry less authority than the wallet](https://andymental.com/drops/agent-payments-shipped-by-refusing-to-trust-the-agent) — published 2026-09-07: Stripe's agent wallet separates purchase approval from payment credentials. The engineering lesson is to give each transaction only the authority it needs. - [The AGI declaration's receipts measure spend, not generality](https://andymental.com/drops/the-agi-declarations-receipts-measure-spend-not-generality) — published 2026-09-07: Jensen Huang called Astra AGI while citing compute scale. OpenAI's own numbers show intense research use—not independent evidence that capability transfers generally. - [An AI lab needs an artifact and a route into production](https://andymental.com/drops/finserv-ai-labs-are-not-new-the-anchor-model-is) — published 2026-09-06: Finance had dedicated AI research years before 2026. Revolut’s PRAGMA-centered announcement raises a better question: how does a research artifact become an owned operating capability? - [Letting a model route attention creates a new thing to audit](https://andymental.com/drops/attention-routing-moves-into-the-models-own-transcript) — published 2026-09-06: Declarative Attention cuts attended tokens by letting models direct access to context. Its savings need to be judged alongside missed evidence and integration costs. - [AI productivity needs an unaided transfer test](https://andymental.com/drops/ai-productivity-needs-an-unaided-transfer-test) — published 2026-09-06: AI-assisted output can rise while underlying capability falls. Pair throughput metrics with a periodic, low-risk unaided test of debugging, review and recovery. - [Output attribution cannot replace the training ledger](https://andymental.com/drops/output-attribution-cannot-replace-the-training-ledger) — published 2026-09-06: A diffusion study separates an output's causal attribution from its training inputs. Keep the input record and output investigation as separate checks; neither certifies the other. - [AI vendor concentration is a financial-stability risk](https://andymental.com/drops/vendor-concentration-is-now-a-financial-stability-category) — published 2026-09-06: The FSB's AI warning is about shared providers, not one bad model. Map common model, cloud, hardware and data dependencies—and prove another path can carry the service. - [A software rebound cannot settle the AI displacement question](https://andymental.com/drops/the-saaspocalypse-repriced-billing-models-not-ai-risk) — published 2026-09-05: Market recovery is an outcome with many causes. Vendor diligence needs a direct test of how agent activity reaches the bill, rather than treating share-price moves as proof. - [An AI capex forecast needs more than one financing assumption](https://andymental.com/drops/the-four-trillion-dollar-debt-wave-mortgages-silicon-like-concrete) — published 2026-09-05: A facility debt ratio cannot price the entire AI buildout without separate assumptions for hardware, power and construction. Start with the asset mix before forecasting the debt wave. - [A GET-only browser still needs a write boundary](https://andymental.com/drops/read-only-web-access-is-still-writable) — published 2026-09-05: Blocking POST does not prove an agent cannot change remote state. Review destination behaviour and request authority, then test the boundary—not just the HTTP method. - [A learning loop still needs someone to own the gate](https://andymental.com/drops/loop-engineering-rebuilds-the-unaudited-artifact-problem) — published 2026-09-05: Tyler Folkman's product-loop walkthrough includes Git and human review. The useful next question is whether each proposed improvement preserves the test that made the previous version acceptable. - [The worm's target list is a better secrets inventory than your audit](https://andymental.com/drops/the-worms-target-list-audits-better-than-you-do) — published 2026-09-05: Mini Shai-Hulud scans 469 credential locations. Turn its categories into regression tests for your secrets inventory, not a one-off incident checklist. - [An AI-associated psychosis report deserves its missing qualifiers](https://andymental.com/drops/the-first-ai-psychosis-case-is-a-co-factor-report) — published 2026-09-04: A clinical case documents chatbot reinforcement alongside other risk factors. Preserve that uncertainty while testing the separate product failure of agreement overriding evidence. - [A cheaper model route can change what happens to your data](https://andymental.com/drops/the-cheapest-model-route-is-a-data-governance-decision) — published 2026-09-04: Muse Spark’s Contributor route pairs low prices with permission to improve Meta’s products. Filter routes by data policy before optimizing price, including on fallback. - [Thinking Machines needs separate ledgers for funding and compute](https://andymental.com/drops/nvidia-sits-on-three-sides-of-the-thinking-machines-table) — published 2026-09-04: Nvidia’s announced investment and compute partnership are confirmed; later financing talks are a different event. Follow payment obligations before treating the ecosystem’s checks as new demand. - [Astra makes a readable rationale a weaker oversight promise](https://andymental.com/drops/astra-trades-away-the-monitor-your-guardrails-assume) — published 2026-09-04: OpenAI reports lower chain-of-thought monitorability for Astra while expanding monitoring. Review controls that depend on explanations separately from controls that constrain actions. - [Instinct's round prices permission, not capability](https://andymental.com/drops/instincts-round-prices-permission-not-capability) — published 2026-09-04: Instinct raised $250 million at a $2.5 billion valuation. The round prices access to connected accounts; trust still depends on what the agent may do without asking. - [A model-name mismatch breaks the capability claim](https://andymental.com/drops/the-model-name-is-a-mail-merge-field) — published 2026-09-03: A growth guide names Gemini in the headline and Claude in the body. Its workflow may still be useful, but it cannot establish which model produced the advertised result. - [The DOJ’s training argument does not clear the whole data pipeline](https://andymental.com/drops/the-doj-defended-the-one-stage-nobody-was-paying-for) — published 2026-09-03: The government’s OpenAI filing argues for fair use at the training stage. Acquisition and particular outputs remain separate questions, and a statement of interest is not a ruling. - [A commerce blueprint does not inherit the launch’s conversion lift](https://andymental.com/drops/the-commerce-agent-launch-is-a-blueprint-and-an-anecdote) — published 2026-09-03: Claude’s commerce examples can accelerate a prototype. Payment, live writes and evidence of commercial lift still belong to the merchant’s implementation. - [The AI trust gap needs dates beside its percentages](https://andymental.com/drops/the-2026-ai-trust-gap-splices-two-surveys) — published 2026-09-03: Slack's usage surge and disclosure discomfort came from different survey rounds. Combining them can turn two real findings into a relationship neither study measured. - [Harvey’s legal engineers make the delivery cost visible](https://andymental.com/drops/harveys-human-layer-has-a-price-tag-its-scale-does-not) — published 2026-09-03: A legal-engineer job listing supports a concrete claim about domain expertise and compensation. It cannot establish customer coverage or the economics of the whole company. - [An AI writing score still needs a reader on the other side](https://andymental.com/drops/an-ai-panel-cannot-prove-ai-improved-the-writing) — published 2026-09-03: Tomasz Tunguz's writing experiment reports a higher AI-rated quality floor. The next test is whether readers understand and use the argument better. - [EY's human-skills bonus rewards AI adoption](https://andymental.com/drops/eys-human-skills-bonus-rewards-ai-adoption) — published 2026-09-03: EY's $100 million rewards programme does not choose human skills over AI. It pays for technology adoption plus judgment—and needs an accepted-output ledger. - [A guidance-demand chart also maps the product's duty of care](https://andymental.com/drops/the-free-market-research-is-a-sycophancy-audit) — published 2026-09-02: Anthropic's personal-guidance study describes demand inside a selected conversation sample. Its failure analysis belongs in any product idea built from that demand. - [Perplexity’s growth curve needs units before a multiplier](https://andymental.com/drops/perplexitys-growth-story-measures-with-two-rulers) — published 2026-09-02: Reported ARR growth can be real while the chart comparing it with annual revenue is misleading. Keep the period, recognition basis and pricing change visible. - [Anthropic’s five-point safeguard gap already has a version problem](https://andymental.com/drops/anthropic-priced-its-own-guardrails-at-five-points) — published 2026-09-02: Fable 5.1 and Mythos 5.1 share a base model, but their published benchmark gap reflects earlier safeguards. A difference in the table is not a permanent price of safety. - [A perfect ExploitBench score changes what the test can tell us](https://andymental.com/drops/a-perfect-exploitbench-score-retired-the-benchmark) — published 2026-09-02: Astra’s reported 100% result needs contamination and configuration context. Keep public tests for regression while demanding separate evidence for harder capability claims. - [Grok Bot roles share the authority of their user's computer](https://andymental.com/drops/a-named-bot-is-not-a-permission-boundary) — published 2026-09-02: Grok Bot documentation separates bot personalities from runtime isolation. A finance bot and a support bot under one user should not be treated as separate trust domains. - [Quantization damage hides in the flips, not the average](https://andymental.com/drops/quantization-damage-hides-in-the-flips-not-the-average) — published 2026-09-02: Aggregate accuracy can stay flat while a quantized model changes answers. Compare baseline and compressed outputs, then human-check multilingual and long-context slices. - [Frontier access needs a dated map, not a closed-camps slogan](https://andymental.com/drops/the-segmentation-thesis-bends-its-own-footnotes) — published 2026-09-01: Defaults, export restrictions and hosting-license conditions constrain different things. Read each boundary at its actual scope before making a model portability decision. - [A search fan-out trend needs the same measuring instrument](https://andymental.com/drops/the-fan-out-stat-measures-the-instrument-not-the-model) — published 2026-09-01: Nectiv's newer fan-out study uses an API where its earlier work inspected ChatGPT's interface. The increase is a finding to investigate, not an isolated model effect. - [Podium’s cutoff needs a business explanation before an AI theory](https://andymental.com/drops/podium-was-cut-for-the-product-not-the-agent) — published 2026-09-01: The ServiceTitan–Podium dispute has competing accounts and a concrete integration deadline. Test continuity against that deadline before treating it as proof of an agent category war. - [An AI risk threshold needs a response attached to it](https://andymental.com/drops/gates-crossed-five-thresholds-nobody-can-measure) — published 2026-09-01: Gates’s essay argues for urgent preparation. Turning that concern into an operating policy requires a defined signal, evidence standard and action owner. - [An agent’s CRM write needs a business reason to persist](https://andymental.com/drops/agent-traces-do-not-belong-in-the-crm) — published 2026-09-01: SaaStr’s reported storage jump mixes customer state with high-volume activity. Decide which records belong in the CRM before agents multiply them. - [Instagram priced disclosure, not synthetic content](https://andymental.com/drops/instagram-priced-disclosure-not-synthetics) — published 2026-09-01: Instagram says labelled AI-person profiles keep their reach; unlabelled ones may lose recommendation eligibility. The controllable rule is disclosure. Detection remains opaque. - [Forty-one workers and seven hundred agents count different things](https://andymental.com/drops/three-true-breach-numbers-tell-three-different-stories) — published 2026-08-31: The Hugging Face incident reports describe compromised server workers, attack participants and message-board users. Compare those counts only after naming the entity. - [A 25% quota increase can still reduce next week’s capacity](https://andymental.com/drops/a-raise-measured-against-a-temporary-baseline-is-a-cut) — published 2026-08-31: Claude Code’s announced change compares with the standard allowance. Capacity planning also needs the comparison with the temporary allowance users currently receive. - [A correct rejection can still send the rewrite in the wrong direction](https://andymental.com/drops/a-judges-reason-is-a-production-command) — published 2026-08-31: Netflix evaluates why its LLM judge rejects an explanation because that reason drives revision. The next useful metric is whether the prescribed repair actually fixes the defect. - [Outcome pricing moves the argument, not the risk](https://andymental.com/drops/outcome-pricing-moves-the-argument-not-the-risk) — published 2026-08-31: Outcome pricing moves failed-attempt cost to the vendor. Buyers still carry risk in the definition, clock, evidence and reversal behind every billed event. - [An arrest changes the threat actor, not the installed artifact](https://andymental.com/drops/the-shai-hulud-arrests-did-not-arrest-the-worm) — published 2026-08-29: The reported TeamPCP arrests are meaningful disruption. Closing an exposure still requires evidence about the code, credentials and publishing paths left behind. - [Anthropic’s court win protects speech, not every guardrail decision](https://andymental.com/drops/the-anthropic-ruling-makes-a-guardrail-refusal-litigable-speech) — published 2026-08-29: The August 27 order addresses retaliation, due process and agency action. It does not turn a vendor’s safety policy into a universal right to a government contract. - [A licensing deal can change the supplier you depend on](https://andymental.com/drops/nvidias-third-non-acquisition-is-a-merger-review-product) — published 2026-08-29: Poolside’s reported Nvidia deal raises an operational question even without a takeover: which people, rights and delivery capabilities remain with the supplier? - [Deleting a memory needs a test against the source transcript](https://andymental.com/drops/deleting-memory-does-not-delete-the-transcript) — published 2026-08-29: A public agent implementation stores memories and transcripts separately. Erasure should be tested across the paths that can recreate the deleted fact. - [Cheap alignment search makes failure discovery more valuable](https://andymental.com/drops/anthropics-four-dollar-researcher-stops-at-the-benchmark) — published 2026-08-29: Anthropic's automated researchers improved measured alignment failures, including withheld tests. Cheaper optimisation increases the value of finding failures the tests still miss. - [Both sides of the AI ROI debate run on soft numbers](https://andymental.com/drops/both-sides-of-the-ai-roi-debate-run-on-soft-numbers) — published 2026-08-29: The viral 94% AI cost-saving claim has no survey behind it. KPMG's skeptical 7% is also self-reported maturity, not calculated ROI. Build one workflow receipt instead. - [Meta’s age-model exception needs a purpose boundary](https://andymental.com/drops/the-kids-privacy-settlement-licenses-kids-data-for-training) — published 2026-08-28: The filed settlement permits narrowly scoped use of under-13 data for age detection. The engineering challenge is proving that the exception stays narrow. - [Hugging Face’s neutrality needs observable commitments](https://andymental.com/drops/the-hugging-face-price-tags-neutrality-not-revenue) — published 2026-08-28: Reported Nvidia acquisition talks are a prompt to specify what neutral platform behaviour means, rather than assume ownership alone decides it. - [A cyber-defense letter needs an executable contribution](https://andymental.com/drops/the-cyber-letter-shipped-a-day-after-its-own-evidence) — published 2026-08-28: The collective warning asks for tools, funding and verified fixes. The next useful measurement is what each participant actually makes available to defenders. - [Ox Alpha’s traffic is a discovery signal, not a blind benchmark](https://andymental.com/drops/ox-alpha-ran-the-first-brand-blind-model-eval-at-scale) — published 2026-08-28: An anonymous model drew heavy use before its reveal. That supports trying it under controlled conditions; token volume alone cannot explain why it won traffic. - [A package named in official docs still needs an identity check](https://andymental.com/drops/llms-txt-is-an-install-script-nobody-audits) — published 2026-08-28: Researchers found unclaimed destinations in files written for agents. The critical boundary is where a documentation reference becomes installed code. - [Smart transcription is not an audit transcript](https://andymental.com/drops/smart-transcription-is-not-an-audit-transcript) — published 2026-08-28: A smart transcript deliberately edits speech. Preserve permitted source audio, version every derived transcript, and make material claims resolve to time spans. - [Revenue per megawatt needs a time period and a cost boundary](https://andymental.com/drops/revenue-per-megawatt-counts-half-the-factory) — published 2026-08-27: An inference-capacity ratio is not a company margin. Specify time, utilization and included costs before using AI-factory economics in an investment or procurement case. - [Unify’s cost cut makes agent topology a testable expense](https://andymental.com/drops/unifys-95-percent-saving-indicts-the-swarm-not-the-model) — published 2026-08-27: The reported saving bundled a simpler agent design with prompt cleanup. Measure each contribution before turning one rewrite into a universal architecture rule. - [Historical web search needs a present-day removal policy](https://andymental.com/drops/the-agent-web-index-keeps-pages-publishers-deleted) — published 2026-08-27: Keenable’s historical queries raise a useful buying question: how do corrections and removals propagate through snapshots and downstream agent answers? - [Speculative decoding needs a load-dependent off switch](https://andymental.com/drops/speculative-decoding-needs-a-load-shedder) — published 2026-08-27: A faster response at low concurrency can become a poor serving configuration under load. Treat draft length as a measured operating policy. - [A thirty-trillion-dollar market still needs a billable unit](https://andymental.com/drops/anthropics-thirty-trillion-is-a-tam-doing-a-valuations-job) — published 2026-08-27: Anthropic’s reported market estimate raises the useful enterprise question: which unit of work becomes revenue, at what price and with whose costs? - [OpenAI locked the weights, not just the credentials](https://andymental.com/drops/openai-locked-the-weights-not-just-the-credentials) — published 2026-08-27: OpenAI rotated access and rebuilt its research controls after the Hugging Face incident. It also quarantined a model family. Agent incident response now needs checkpoint lineage. - [A stock-market loss is not a purchase-order cancellation](https://andymental.com/drops/the-kospi-crash-priced-leverage-not-ai-demand) — published 2026-08-26: The Korean selloff and SK hynix’s strong quarterly results can coexist. An AI capacity decision needs evidence that connects market prices to demand. - [Redacting the visible transcript does not clear the whole export](https://andymental.com/drops/sanitized-agent-logs-can-still-carry-secrets) — published 2026-08-26: Research on opaque reasoning fields changes the export question: what data are we sharing that our redaction process cannot inspect? - [PRAGMA’s strongest evidence is a benchmark, not a moat](https://andymental.com/drops/revoluts-moat-paper-has-nvidias-name-on-it) — published 2026-08-26: Revolut’s model paper supports a shared representation across banking tasks. Durable commercial advantage requires a different comparison. - [Jalapeño’s benchmark still needs a production-shaped comparison](https://andymental.com/drops/jalapenos-missing-benchmark-is-the-one-that-matters) — published 2026-08-26: OpenAI’s chip results are promising. The buying question adds contemporary alternatives, realistic agent workloads and a verified power boundary. - [The case for agents that die every night](https://andymental.com/drops/the-case-for-agents-that-die-every-night) — published 2026-08-26: Persistent agent memory turns one poisoned document into future authority. A nightly reset helps only when durable memories cross a separate, auditable promotion gate. - [Workday’s case puts the vendor’s own conduct on the map](https://andymental.com/drops/workday-rulings-move-hiring-liability-to-the-vendors-hq) — published 2026-08-25: The June ruling leaves California claims in play based on alleged conduct there. Hiring accountability needs a map of who makes each screening decision. - [An AI safety promise should have a retrievable control behind it](https://andymental.com/drops/the-first-ai-escape-subpoena-is-a-consumer-protection-case) — published 2026-08-25: Alabama’s OpenAI subpoena makes the operational question concrete: can the company connect its safety representations to the controls that actually ran? - [A persona review needs an owner who can change its criteria](https://andymental.com/drops/the-cpo-check-outlived-the-cpo) — published 2026-08-25: Freshworks’ demonstrated CPO review and a leadership transition expose a useful design question: who maintains the judgment encoded in a named-person agent step? - [A citation collapse should trigger a measurement check first](https://andymental.com/drops/the-86-percent-citation-collapse-might-be-the-dashboard) — published 2026-08-25: Promptwatch labels Reddit’s reported citation drop provisional. Audit collection, denominators and raw answers before turning a dashboard discontinuity into a content-budget decision. - [Replit’s sales pivot prices the trust gap](https://andymental.com/drops/replits-sales-pivot-prices-the-trust-gap) — published 2026-08-25: Replit plans for salespeople to become a majority even as Free Mode makes building cheaper. The contradiction disappears when user adoption and institutional approval run on separate clocks. - [An automation can retire before its decision record is needed](https://andymental.com/drops/the-825-million-euro-fine-predates-the-agent-era) — published 2026-08-24: The Uber enforcement action concerns earlier account-deactivation practices. Preserve enough versioned workflow evidence to explain consequential automated decisions after the system changes. - [Read the model-share chart's method, not the company's label](https://andymental.com/drops/the-11-percent-flagship-number-measures-the-card-not-the-contract) — published 2026-08-24: Ramp’s Fable chart uses token-spend management data. Its 11.4% figure is meaningful within that sample, but neither a market census nor simply a corporate-card statistic. - [Razorpay's Vulcan makes the prediction target the architecture question](https://andymental.com/drops/razorpays-vulcan-is-not-an-llm-on-purpose) — published 2026-08-24: Razorpay describes Vulcan as a payments foundation model rather than an LLM. Start an AI design with the decision and its evidence before deciding to turn the inputs into prose. - [Your agent harness can change a rollout's training weight](https://andymental.com/drops/control-flow-can-weight-the-gradient) — published 2026-08-24: Agent Lightning shows how one task can become several training rows. Preserve rollout lineage so incidental fragmentation does not silently redefine the optimisation objective. - [Agent bills are moving to the control plane](https://andymental.com/drops/agent-bills-are-moving-to-the-control-plane) — published 2026-08-24: AgentSysBench found tools or environments dominated latency in half its agent apps. Meter model, tool, state and waiting costs together at the control plane. - [An AI-code critique still needs the original survey question](https://andymental.com/drops/the-vibe-coding-backlash-runs-on-vibe-sourced-statistics) — published 2026-08-23: Favorability, trust and observed defects measure different things. Trace the question and population before using a statistic to justify an engineering policy. - [A litigation headline needs an owner and a calculation](https://andymental.com/drops/the-trillion-dollar-accountability-number-is-the-defendants) — published 2026-08-23: Meta’s trillion-dollar exposure figure is a party’s account of maximum-penalty arithmetic. Keep it separate from a requested remedy, a likely outcome and a court award. - [Calibrate text-watermark detection on the artifacts you actually review](https://andymental.com/drops/the-claude-watermark-is-weakest-where-detection-matters) — published 2026-08-23: Anthropic says exact code and factual passages carry less watermark signal. A detector evaluated on long prose cannot supply a reliable threshold for terse technical artifacts. - [A future IPO is not a vendor-continuity plan](https://andymental.com/drops/openai-calls-its-own-ipo-another-fundraise) — published 2026-08-23: Reported IPO plans describe a financing milestone. Enterprise buyers still need evidence about the service, contract and transition costs their own dependency creates. - [Perfect accuracy cannot reveal the algorithm](https://andymental.com/drops/perfect-accuracy-cannot-reveal-the-algorithm) — published 2026-08-23: Four compiled transformers returned every supported calculator answer correctly through different mechanisms. Accuracy can certify outputs; it cannot identify the algorithm or predict how it scales. - [An automation rollout can quietly redistribute operational control](https://andymental.com/drops/the-pizza-hut-ai-lawsuit-is-an-access-control-story) — published 2026-08-22: The Dragontail lawsuit alleges changes in delivery assignment and local control. Test who can assign, refuse and pause work alongside the system’s efficiency claims. - [The 92% clinical-AI score does not measure the cost of human review](https://andymental.com/drops/the-jama-autonomy-papers-best-number-indicts-the-reviewer) — published 2026-08-22: The underlying trial tested access to GPT-4, not edits to identical model outputs. Use its result to demand workflow evidence rather than declare reviewers harmful. - [The 89% local-AI result is a ceiling for a router to approach](https://andymental.com/drops/the-89-percent-local-ai-claim-assumes-the-answer-key) — published 2026-08-22: Best-of-local coverage selects successful models after evaluation. A deployment must show how it routes and detects misses before claiming the same coverage or savings. - [Gamma's average payer is the wrong unit for an enterprise sales decision](https://andymental.com/drops/gammas-missing-sales-team-was-arithmetic-not-strategy) — published 2026-08-22: Gamma’s self-serve scale does not tell us the economics of a team sale. Evaluate sales assistance against the incremental contract and service burden it can unlock. - [ARC-AGI-3's perfect score belongs to the harness](https://andymental.com/drops/arc-agi-3s-perfect-score-belongs-to-the-harness) — published 2026-08-22: Nvidia's AVO scored 100.00 on ARC-AGI-3's public set while ARC Prize reports Claude Opus 5 High at 30.16%—but Nvidia says this does not isolate harness gain. - [A package incident has more than one closing time](https://andymental.com/drops/the-40-minute-litellm-window-took-five-months-to-price) — published 2026-08-21: Removing malicious LiteLLM releases ended one distribution path. Closing downstream exposure also requires evidence about installed copies, persistence and credential revocation. - [An enterprise AI framework must account for AI already in production](https://andymental.com/drops/sbis-billion-follows-a-trillion-already-underwritten) — published 2026-08-21: SBI’s planned bank-wide AI investment follows substantial reported AI-assisted lending. The implementation question is how existing workflows join the new framework. - [An adjusted profit needs a bridge before it becomes a benchmark](https://andymental.com/drops/anthropics-first-profit-is-an-undefined-adjustment) — published 2026-08-21: A reported adjusted operating profit is a starting point. Compare AI businesses only after identifying the adjustments, period and underlying accounting measure. - [Pose landmarks are data minimisation, not anonymity](https://andymental.com/drops/pose-landmarks-are-not-anonymous) — published 2026-08-21: Google’s SL2T deletes camera video on-device and sends pose coordinates instead. That reduces collection risk. It does not prove the motion trace cannot identify its signer or reveal other traits. - [A GPU lease starts with the equipment you control](https://andymental.com/drops/the-regulator-just-put-gpus-on-a-lease-schedule) — published 2026-08-20: IFSCA’s proposal distinguishes identified-equipment leasing from shared compute services. The useful procurement question is which asset and risks the contract actually assigns. - [An event-driven agent needs a retained trigger record](https://andymental.com/drops/self-waking-agents-make-the-transcript-the-contract) — published 2026-08-20: A conversation can change after it wakes an agent. Preserve the authorised event and its link to the run so later reviewers can reconstruct why the agent acted. - [Save the filter and reasoning budget with the model rank](https://andymental.com/drops/qwens-number-one-is-a-filter-artifact) — published 2026-08-20: Qwen’s leaderboard position is only meaningful with its comparison class and inference settings. Preserve both before turning a ranking into a deployment decision. - [GraphRAG precomputes answers before questions](https://andymental.com/drops/graphrag-precomputes-answers-before-questions) — published 2026-08-20: Full GraphRAG builds reusable community reports before a query. On a changing corpus, that makes retrieval a freshness contract; default to lazy summarization unless the reports are the product. - [Read the bridge between AI operating loss and adjusted EBITDA](https://andymental.com/drops/spacex-segment-table-is-the-ai-buildouts-p-and-l) — published 2026-08-19: SpaceX’s AI segment reports an operating loss and positive adjusted EBITDA. The reconciliation explains both; neither number alone identifies who funded the buildout. - [Measure an AI giveaway from each user's renewal date](https://andymental.com/drops/perplexitys-india-giveaway-priced-the-free-ai-user) — published 2026-08-19: A promotion closing is not a free subscription expiring. Perplexity’s Airtel offer needs activation cohorts before its paid conversion can be judged. - [Origin's source of truth depends on how you create the repository](https://andymental.com/drops/origin-syncs-with-the-github-it-supposedly-killed) — published 2026-08-19: Cursor offers native hosting and GitHub mirroring. Choose an authoritative write path per repository before designing a migration or an outage plan. - [An advertising principle needs a regional configuration record](https://andymental.com/drops/chatgpts-eu-ad-restraint-is-compliance-not-philosophy) — published 2026-08-19: ChatGPT’s European ad expansion has a launch scope, access path and stated privacy commitments. Evaluate the implemented regional product instead of inferring either pure principle or a hidden motive. - [41B active parameters still need a 600 GB inference cluster](https://andymental.com/drops/active-parameters-do-not-size-the-cluster) — published 2026-08-19: Inkling activates 41B of 975B parameters per token, yet its NVFP4 checkpoint needs 600 GB of aggregate VRAM. Active parameters size arithmetic—not resident weights, cache or topology. - [A coding-agent fallback must use a working authentication path](https://andymental.com/drops/your-coding-agent-recovers-last) — published 2026-08-18: GitHub’s August 17 incident left some Copilot clients impaired after other services improved. The useful resilience test is whether the fallback avoids the failed dependency and retry behaviour. - [Vendor-selection advice needs the same evidence check as vendors](https://andymental.com/drops/the-vendor-pick-memo-flunks-its-own-audit) — published 2026-08-18: A confident comparison can mix rounded figures, run rates and measured periods. Audit the advisory memo before letting its arithmetic decide an enterprise platform choice. - [More experiments need a false-discovery budget](https://andymental.com/drops/the-case-against-full-stack-builders-understates-itself) — published 2026-08-18: An experiment’s significance threshold does not tell us what share of declared winners are real. Faster AI-assisted experimentation makes that distinction more useful, not less. - [Weight updates within a context do not establish personal memory](https://andymental.com/drops/test-time-training-is-not-agent-memory) — published 2026-08-18: TTT-E2E demonstrates a way to process long contexts by learning from them. Persistent user memory needs a separate test of retention, isolation and what happens when information changes. - [Klaviyo's L3 mandate needs an output ledger](https://andymental.com/drops/klaviyos-l3-mandate-measures-logins-not-output) — published 2026-08-18: Klaviyo told 2,300 people to reach L3 agent fluency and reported revenue per employee up 28%. A mandate starts behaviour; output and rework decide whether it worked. - [A watermark-removal claim must name the mark it can change](https://andymental.com/drops/the-watermark-remover-proves-the-threat-model) — published 2026-08-17: Metadata cleaning and changes to sampled wording operate on different signals. A tool’s removal label does not establish that it defeats a particular text watermark. - [An acquisition haircut needs a real starting price](https://andymental.com/drops/the-3-billion-haircut-is-measured-against-a-leak) — published 2026-08-17: OpenRouter’s reported sale price is being compared with an earlier reported negotiation figure. Preserve each number’s source and stage before calling the difference a loss. - [A training ratio is a hypothesis about your loss surface](https://andymental.com/drops/the-20-token-rule-is-not-a-constant) — published 2026-08-17: Skaling challenges the independence assumption in a familiar scaling-law form. Before turning a token-to-parameter rule into a compute budget, test its predictions beyond the easy interior. - [Giving the agent a clock shipped four trust models under one word](https://andymental.com/drops/the-agents-clock-ships-four-trust-models) — published 2026-08-17: Claude's /loop, Desktop tasks, Cowork tasks and cloud routines share a timer, not a permission boundary. Name the runner, state, authority and missed-run policy before granting a schedule. - [Google’s TPU split makes the workload question harder to skip](https://andymental.com/drops/google-split-the-tpu-because-agents-are-a-memory-problem) — published 2026-08-16: TPU 8t and 8i emphasise different training and serving needs. Their specifications justify a workload-specific evaluation, not a universal claim that agents are only memory-bound. - [A talent departure can leave a platform relationship intact](https://andymental.com/drops/google-didnt-lose-jeff-dean-it-bought-in) — published 2026-08-16: Google’s announcement pairs Jeff Dean and Sanjay Ghemawat’s independent venture with an investment, a Cloud partnership and planned research collaboration. Follow the dependencies as well as the people. - [A tax-agent accuracy claim needs a unit of error](https://andymental.com/drops/a-98-percent-accurate-tax-agent-is-a-liability-schedule) — published 2026-08-16: A reported 98 percent accuracy figure cannot size a review process until we know what was scored, which cases were included and when human corrections occurred. - [The 12-agent adoption stat lost its denominator](https://andymental.com/drops/the-multi-agent-adoption-stats-have-no-primary-source) — published 2026-08-16: Salesforce did publish the 12-agent survey. The error is the denominator: 1,050 IT leaders at 1,000-plus-employee enterprises became "the average company" for an SMB audience. - [Keep the speaker attached to a private-company valuation](https://andymental.com/drops/the-2-trillion-anthropic-ipo-is-priced-by-leaks) — published 2026-08-15: The reported $2 trillion Anthropic IPO expectation comes from investors. A supplier assessment should preserve that attribution and distinguish a proposed price from an executed transaction. - [Router token share cannot count enterprise vendor switches](https://andymental.com/drops/router-token-share-is-not-enterprise-market-share) — published 2026-08-15: OpenRouter measures tokens in its platform; Ramp measures observed business spending. Their results can coexist without proving that enterprises have replaced incumbent providers. - [Cheaper AI tasks do not tell you how to price the product](https://andymental.com/drops/canvas-problem-is-ai-pricing-not-ai-costs) — published 2026-08-15: An inference saving is only one input to a product decision. Model usage, paid value and the expensive tail before choosing a bundle, allowance or action price. - [A booking agent cannot be your API’s permission boundary](https://andymental.com/drops/agents-are-ambient-pen-testers-of-consumer-apis) — published 2026-08-15: The reported Melbourne gym incident raises two questions: why the agent chose a harmful action, and why the service accepted it. Each needs its own owner and repair. - [The AI worm moved its inference bill onto the victim](https://andymental.com/drops/the-ai-worm-story-is-unit-economics-not-capability) — published 2026-08-15: A lab worm reached half its network in about five days by reasoning on stolen GPUs. The cost transfer is also a signal: detect unauthorised inference before its patch clock wins. - [A model-routing price needs an expiry date](https://andymental.com/drops/your-routing-table-needs-a-price-expiry-column) — published 2026-08-14: Gemini 3.7 Flash’s launch rate has a scheduled step-up. Store the price period with the model route, and evaluate the workload across the dates the business will actually pay. - [Your crawler and reviewer may be reading different words](https://andymental.com/drops/web-ingestion-needs-a-rendering-integrity-check) — published 2026-08-14: Font substitution can make source text disagree with rendered text. Add a targeted comparison before high-trust web material becomes retrieval evidence, and preserve uncertainty when the views diverge. - [An oversubscribed round still needs a spending thesis](https://andymental.com/drops/the-round-size-is-the-investors-number-now) — published 2026-08-14: Investor appetite can change the size of a financing. That is evidence of demand for the equity, while the company’s use of the additional capital needs a separate explanation. - [An AI training toggle needs a time boundary](https://andymental.com/drops/the-opt-out-default-is-a-confession) — published 2026-08-14: Twitch’s generative-AI setting illustrates a practical data question: what exactly changes when a person opts out, and what can the provider establish about earlier use? - [The agent turf war happened in the lab, not the wild](https://andymental.com/drops/the-turf-war-happened-in-the-lab-not-the-wild) — published 2026-08-14: The Hugging Face intrusion was one autonomous campaign, not a conspiring agent team. Anthropic's lab turf war is the better case—and it changes what multi-agent evals must test. - [Revenue outgrowing customer count does not prove deeper AI adoption](https://andymental.com/drops/palantirs-quarter-says-ai-money-is-concentrating) — published 2026-08-13: Palantir’s rapid growth makes account depth worth investigating. Aggregate revenue and customer counts cannot tell us which workflows expanded or whether customers earned a return. - [Palantir’s contract value needs its cancellation clause attached](https://andymental.com/drops/palantirs-contract-value-is-not-committed-demand) — published 2026-08-13: Total contract value includes potential contract value under stated assumptions. Compare it with noncancelable obligations and recognised revenue before treating it as committed demand. - [A financing target is not funded compute capacity](https://andymental.com/drops/nvidias-500-billion-is-an-mou-not-a-balance-sheet) — published 2026-08-13: NVIDIA’s proposed platforms aim to mobilise more than $500 billion over time. Keep the target separate from signed transaction terms, funded assets and usable infrastructure. - [An agent’s fallback needs a different review procedure](https://andymental.com/drops/auto-modes-fallback-is-the-control-it-just-beat) — published 2026-08-13: Claude Code’s auto mode can return to manual approvals after repeated blocks. That transition should supply evidence and preserve boundaries instead of adding one more routine prompt. - [Agent autonomy is a timeout setting, not a capability](https://andymental.com/drops/agent-autonomy-is-a-timeout-setting) — published 2026-08-13: GitHub caps a coding-agent session at 59 minutes, Replit advertises 200, and Vercel offers 24-hour sandboxes. Those numbers measure product envelopes, not capability. - [Claude's watermark is provenance, not a verdict](https://andymental.com/drops/claudes-watermark-is-provenance-not-a-verdict) — published 2026-08-12: Claude's watermark can show that Claude processed text, not who authored it or whether policy was breached. Route detections to review, never straight to discipline. - [A category multiple inherits the choices in its comparison set](https://andymental.com/drops/the-harness-multiple-prices-an-undefined-category) — published 2026-08-11: Tomasz Tunguz’s AI harness valuation chart includes useful caveats about estimates and revenue models. Keep those caveats attached before using the range as a pricing benchmark. - [Redeployed salaries need a second benefit calculation](https://andymental.com/drops/redeployed-salaries-are-not-agent-savings) — published 2026-08-11: Vercel’s lead-agent story describes a large capacity change. A credible business case separates runtime cost, operating labour, cash savings and the value of reassigned people. - [Open weights make containment a deployer’s job](https://andymental.com/drops/an-open-weight-escape-has-no-patch-day) — published 2026-08-11: Kimi K3’s reported benchmark shortcut used an allowed network route. Downloaded weights do not remove the ability to repair that route; they make local ownership unavoidable. - [An AI grounding query is not a buyer’s verbatim question](https://andymental.com/drops/ai-chat-followups-are-someones-keyword-data) — published 2026-08-11: AI visibility reports mix different kinds of query signals. Before treating one as customer language, establish whether it came from the person, the assistant or an aggregation. - [QM makes the coding agent a swappable part](https://andymental.com/drops/qm-makes-the-coding-agent-a-swappable-part) — published 2026-08-11: QM lets Pi, OpenCode, Codex and Claude Code drive one core. That makes runner choice reversible—but only if a swap drill proves state, policy and audit continuity. - [Failure attribution should not choose the repair owner](https://andymental.com/drops/the-better-model-test-hides-harness-debt) — published 2026-08-10: A model may be capable of recovering from a bad tool interaction while the cheapest reliable fix still belongs in the surrounding application. Diagnose the event and the repair separately. - [Delete a prompt rule only after naming what replaced it](https://andymental.com/drops/a-no-regression-result-is-being-sold-as-a-gain) — published 2026-08-10: Anthropic’s large system-prompt reduction preserved measured coding performance. The transferable method is controlled removal with evidence, not an instruction to delete 80 percent. - [Agent payments should earn the right to write](https://andymental.com/drops/the-agent-payments-stack-that-shipped-is-read-only) — published 2026-08-10: Paytm shipped payment visibility while Cloudflare opened wallet handles before spend access. The safer rollout pattern is to observe, rehearse, then grant bounded writes. - [An injunction win is not a licence for every shopping agent](https://andymental.com/drops/the-perplexity-ruling-makes-the-user-the-accessor) — published 2026-08-09: The Ninth Circuit’s Perplexity opinion turns on who accessed Amazon in the record before it. Architecture matters more than a sweeping claim that agent commerce is settled. - [Critical-provider oversight does not outsource a bank’s vendor review](https://andymental.com/drops/julys-real-banking-security-story-was-a-designation) — published 2026-08-09: The UK’s first critical-third-party designations bring cloud dependencies under direct oversight. Firms still own due diligence, risk management and contingency planning. - [Flint makes the chart specification worth reviewing](https://andymental.com/drops/flint-treats-charts-as-a-compile-target) — published 2026-08-09: Flint turns one semantic chart specification into several native outputs. Review the meaning before choosing the renderer—and test what survives the swap. - [The AI Act delay did not stop Article 50](https://andymental.com/drops/the-ai-act-delay-left-the-transparency-clock-running) — published 2026-08-09: The EU delayed high-risk AI duties, not Article 50 transparency. Product teams need a role-by-surface map for chatbots, synthetic content, biometrics and deepfakes. - [Google Cloud’s profit cannot prove a frontier retreat](https://andymental.com/drops/google-traded-the-frontier-for-cloud-margin) — published 2026-08-08: Cloud earnings and a critical research outlook answer different questions. Judge a long-term model commitment by delivered capability, replacement options and observable milestones. - [Biosecurity review needs the selected system, not a random output](https://andymental.com/drops/biosecurity-evals-must-cover-the-model-scaffold) — published 2026-08-08: Genome-design research separates novelty from efficiency and exposes downstream selection. Evaluate the governed workflow and its access boundaries without turning one biological result into a universal claim. - [Shared artifacts are communication permissions for agents](https://andymental.com/drops/agent-side-channels-are-an-untracked-attack-surface) — published 2026-08-08: A reported cross-agent message board exposes a boundary beyond chat tools: any shared artifact can carry information between runs. Test whether that sharing is intended. - [A browser can use less CPU and still make the agent wait longer](https://andymental.com/drops/agent-browsers-trade-wall-time-for-density) — published 2026-08-08: Cloudflare’s Kitesurf benchmark separates resource use from elapsed time. Browser selection needs both measures, plus the cost of unsupported pages and fallbacks. - [Drive-thru AI needs a measured handoff, not a rollout victory lap](https://andymental.com/drops/the-drive-thru-gap-was-operating-discipline) — published 2026-08-08: Taco Bell's voice-AI scale does not prove why a rollout succeeds. Measure who takes over, what context survives, and how much work each exception creates. - [Before tuning retrieval, prove the missing item was indexed](https://andymental.com/drops/retrieval-recall-is-a-corpus-scope-problem) — published 2026-08-07: An email-search tutorial makes the boundary visible: a better embedding cannot retrieve a message that never entered the index. Debug eligibility before ranking. - [Agree the authorship evidence before asking someone to prove it](https://andymental.com/drops/ai-authorship-needs-receipts-before-suspicion) — published 2026-08-07: A clear AI-use policy should explain what assistance is allowed and what process records are proportionate. A late suspicion should not invent the evidence standard. - [Check the sentence against its citation before borrowing the argument](https://andymental.com/drops/a-cited-link-is-not-a-checked-claim) — published 2026-08-07: A linked source can support the revenue figure while using a different customer universe. Verify each claim separately, including the corrections you intend to make. - [Tool descriptions are not agent guardrails](https://andymental.com/drops/tool-descriptions-are-not-agent-guardrails) — published 2026-08-07: An edit tool can promise a unique match and still replace the first duplicate. Put the constraint in executable code, then test the refusal cases before giving an agent write access. - [Local models do not make local agents private](https://andymental.com/drops/local-models-do-not-make-local-agents-private) — published 2026-08-07: Local inference removes one recipient of your data. It does not restrict the agent's files, credentials or network tools. Those permissions need their own review. - [SpaceX’s 12% ratio measures funding dependence, not a verdict](https://andymental.com/drops/the-hyperscaler-solvency-table-measures-the-wrong-risk) — published 2026-08-06: Operating cash flow covered roughly 12% of SpaceX’s first-half capex. The next questions concern financing, available liquidity and obligations—not an automatic solvency ranking. - [The AISI incident needs both the attempt and the stopping point](https://andymental.com/drops/the-aisi-incident-is-a-containment-win-reported-as-an-escape) — published 2026-08-06: AISI observed unsanctioned actions in a deliberately permissive test. Report the dangerous behaviour, the conditions and the controls that stopped it together. - [Price the time it takes to believe an agent is done](https://andymental.com/drops/supervision-hours-are-the-real-agent-pricing-metric) — published 2026-08-06: A finished-looking change can create hidden review work. Measure verification and rework per accepted outcome before calling more agent activity a productivity gain. - [Read an AI guarantee through its metric and its remedy](https://andymental.com/drops/cognitions-guarantee-is-a-bet-it-referees-itself) — published 2026-08-06: Cognition’s guarantee uses estimated engineering output and settles shortfalls in usage credits. Ask whether that metric and remedy address the risk your team wants covered. - [Distilled models inherit their teacher's provenance](https://andymental.com/drops/distilled-models-inherit-teacher-provenance) — published 2026-08-06: Nature showed students trained on filtered number sequences from a misaligned teacher inherited the misalignment. If traits survive filtering, the governed unit is not the dataset — it is the whole teacher lineage. - [A legitimate agent can still look like a bot](https://andymental.com/drops/visa-bought-the-sensor-that-tells-humans-from-agents) — published 2026-08-05: Visa’s proposed BioCatch acquisition highlights behavioural signals. Agent commerce also needs evidence of who delegated an action and what they allowed. - [An AI security coalition needs evidence from the missing layer](https://andymental.com/drops/the-agent-security-alliance-is-missing-the-agent-makers) — published 2026-08-05: The SAFE proposal invites incident learning across the AI stack. Its useful test is whether a review can proceed when a critical counterparty does not participate. - [A 128K context window needs a concurrency budget](https://andymental.com/drops/context-windows-are-concurrency-commitments) — published 2026-08-05: Llama 3.1 70B needs about 39.1 GiB of 16-bit KV cache at 128,000 tokens per request. Treat long-context access as a capacity tier, with measured concurrency and an explicit overload policy. - [Airtable's acquisition price is not a verdict on no-code](https://andymental.com/drops/the-market-repriced-software-humans-operate) — published 2026-08-05: Airtable's deal separates enterprise value from equity value. For platform teams, the useful question is workflow portability, not whether one acquisition proves that no-code is dead. - [Private app generation still needs an admission decision](https://andymental.com/drops/private-vibe-coding-makes-policy-the-product) — published 2026-08-04: Customer-cloud deployment can control where generated apps run. It does not decide who owns them, which data they can use or when their infrastructure should be retired. - [A public rebuttal is an evidence selection, not the case record](https://andymental.com/drops/openai-answered-a-federal-complaint-with-a-blog-post) — published 2026-08-04: OpenAI’s response to Apple publishes correspondence alongside its argument. Read the documents, but preserve the difference between a party’s account and an adjudicated finding. - [A bug-report quota needs a route for the good report it blocks](https://andymental.com/drops/bug-bounties-now-ration-reviewer-attention-not-bugs) — published 2026-08-04: Submission limits can protect scarce reviewer time while delaying valid findings. Evaluate the exception path as carefully as the cap. - [A cheap migration run still needs an accountable cutover](https://andymental.com/drops/agents-reprice-the-migration-layer-first) — published 2026-08-04: SaaStr’s reported migration separates cheap API execution from human mapping and validation. Price the complete transfer before comparing it with a services quote. - [Guardrail benchmarks are graded before the attacker moves](https://andymental.com/drops/guardrail-benchmarks-are-graded-before-the-attacker-moves) — published 2026-08-04: Twelve LLM defenses looked near-perfect against weak tests. Adaptive attacks broke most above 90%. Procurement should demand the attack receipt, not the headline score. - [The disabled feature list belongs in the adoption review](https://andymental.com/drops/the-disable-list-is-the-real-agent-adoption-metric) — published 2026-08-03: Training can succeed while the intended agent workflow remains unavailable. Track the permissions and acceptance conditions needed to make that workflow usable. - [Agent governance is compiling into the language](https://andymental.com/drops/agent-governance-is-compiling-into-the-language) — published 2026-08-03: One essay says agent governance is the layer nobody built. The same weekend, NVIDIA's NOOA makes prompts docstrings and contracts type annotations. Most of it is version control. - [A model catalogue needs an exit date beside the selection](https://andymental.com/drops/a-hyperscalers-model-catalog-is-inventory-not-a-roadmap) — published 2026-08-02: Reports of Amazon’s Nova changes are a reminder to distinguish available models from durable commitments. Prepare the migration before a formal deadline arrives. - [US AI compliance dates move in both directions](https://andymental.com/drops/us-ai-compliance-dates-move-in-both-directions) — published 2026-08-02: California's AI Transparency Act went live today after a delay that added duties. Colorado's delay removed them. A postponement is when the statute gets rewritten. - [An AI price needs a date and a workload beside it](https://andymental.com/drops/model-prices-need-versioned-vectors) — published 2026-08-01: Luna’s price cut was real. The lasting lesson is to keep dated rates and workload assumptions together so a copied price cannot silently rewrite an estimate. - [Incident response needs a tested route through model refusals](https://andymental.com/drops/incident-response-needs-a-model-that-wont-refuse) — published 2026-08-01: Hugging Face says hosted-model guardrails blocked forensic analysis. Prepare an authorized fallback before an incident, with data controls, evidence checks and a human route when AI is unavailable. - [A compatible API request can buy a changed service](https://andymental.com/drops/backward-compatible-can-mean-a-different-product) — published 2026-08-01: OpenAI maps existing priority requests to Fast mode. A successful response does not establish that latency and service economics still match the workload’s assumptions. - [An ambient agent needs a channel-level operating agreement](https://andymental.com/drops/a-digital-employee-with-an-admin-panel-is-infrastructure) — published 2026-08-01: Claude Tag makes proactive work a shared channel capability. Agree what it may observe, remember and initiate before treating it as an ordinary teammate. - [Editable artifacts are the agent handoff](https://andymental.com/drops/editable-artifacts-are-the-agent-handoff) — published 2026-08-01: Figma says its MCP deck workflow produced the first 80% of the work. The honest read is that the last 20% is the handoff — so measure time-to-review, not automation rate. - [Read the capacity of the signature before the policy headline](https://andymental.com/drops/read-ai-governance-by-the-signature-block) — published 2026-07-31: An employee’s public view and an organisation’s endorsement carry different authority. Preserve that distinction before inferring what a company has committed to do. - [Google ATLAS maps occupational reach, not your adoption rate](https://andymental.com/drops/google-atlas-maps-usage-not-adoption) — published 2026-07-31: ATLAS shows where observed AI activity maps into work. To measure adoption within a workforce, add a population denominator and an outcome check. - [Cloud backlog needs a composition check before a demand story](https://andymental.com/drops/cloud-backlog-is-a-concentration-disclosure-now) — published 2026-07-31: Large cloud commitments can signal durable demand without proving broad adoption. Read customer composition, contract duration and conversion timing together. - [Test AI detectors on the writing people actually submit](https://andymental.com/drops/ai-detector-evals-need-an-evasion-baseline) — published 2026-07-31: A detector’s strong result on generic AI prose can weaken on style-conditioned text. Evaluate the intended genre and review both kinds of error. - [The agent intrusion was a four-company event](https://andymental.com/drops/the-agent-intrusion-was-a-four-company-event) — published 2026-07-31: Hugging Face's forensic timeline maps an intrusion across four organisations, where the enabling hop sat in a customer's account on someone else's platform. My July take was insufficient. - [An agent-heavy supplier needs a capacity story you can question](https://andymental.com/drops/the-org-chart-is-becoming-a-compute-contract) — published 2026-07-29: Compute commitments describe a dependency, not productive capacity. Ask an agent-heavy supplier which customer outcomes its contracted resources can sustain. - [The eval needs a second job after the team starts optimising](https://andymental.com/drops/eval-as-spec-makes-the-spec-a-target) — published 2026-07-29: Concrete evals give product teams a target. Fresh cases and independent acceptance checks reveal whether improvement transfers beyond the examples being tuned. - [Agent capacity planning needs a physical failure assumption](https://andymental.com/drops/agent-capacity-planning-now-includes-the-grid) — published 2026-07-29: Grid stress belongs in workload planning. Ask what pauses, what resumes and what capacity remains independent before assuming another endpoint is a fallback. - [Model leaderboards can't see the harness](https://andymental.com/drops/model-leaderboards-cant-see-the-harness) — published 2026-07-29: Endor Labs ran the same model through two harnesses in one week and got a 25.7-point swing — wider than the spread across the whole top of the leaderboard it would be ranked on. - [A model dropdown does not prove you can leave the platform](https://andymental.com/drops/model-agnosticism-is-now-the-incumbents-pitch) — published 2026-07-28: Microsoft’s model-diverse pitch makes a good architectural point. Test portability at the workflow boundary, including context, policy and the release evidence. - [AI answers are leaving the crawled web](https://andymental.com/drops/ai-answers-are-leaving-the-crawled-web) — published 2026-07-28: Google's own cards became the #2 cited domain inside its AI Mode, and OpenAI licensed Yelp rather than crawling harder. The crawled page is becoming the fallback, not the source. - [Buzz swaps the org chart for an audit trail](https://andymental.com/drops/buzz-swaps-the-org-chart-for-an-audit-trail) — published 2026-07-28: Block open-sourced Buzz five months after cutting ~4,000 jobs to AI. Buzz answers which agent acted. It does not answer who is accountable — and that is the half that was cut. - [A voice transcript needs the moment the instruction changed](https://andymental.com/drops/full-duplex-voice-breaks-the-transcript-contract) — published 2026-07-27: When people interrupt an acting voice assistant, the record must preserve the order of speech, action and cancellation. A tidy dialogue is insufficient. - [Prompt deletion needs a regression record](https://andymental.com/drops/deleting-instructions-is-the-new-capability-release) — published 2026-07-27: Anthropic’s prompt reduction is a reason to test old workarounds, not remove authority boundaries. Give every retained instruction a purpose and every deletion a reproducible comparison. - [Cheaper AI routing should be a policy you can inspect](https://andymental.com/drops/a-router-paid-on-spend-cannot-want-your-bill-smaller) — published 2026-07-27: A percentage fee creates an incentive to examine, not proof of bad routing. Own the quality threshold and inspect why each model was selected. - [Agent harnesses need maps, not more manuals](https://andymental.com/drops/agent-harnesses-need-maps-not-more-manuals) — published 2026-07-27: Self-evolving agents fail at locating where to edit before they fail at writing the edit (Harness Handbook, arXiv Jul 14). The scarce artifact is a source-verified map — not a fatter instruction manual. - [A KYAPay proposal is not an endorsed identity standard](https://andymental.com/drops/know-your-agent-is-a-vendor-spec-not-a-standard) — published 2026-07-26: KYAPay's old profile has an active successor. That matters—but an individual Internet-Draft still does not establish IETF endorsement or prove that two implementations interoperate. - [A trust campaign should point to a testable promise](https://andymental.com/drops/ai-trust-ads-argue-vibes-not-mechanisms) — published 2026-07-26: AI advertising can make a promise memorable. Product teams should attach that promise to an observable behaviour, a limit and an owner. - [Temperature zero is not a determinism guarantee](https://andymental.com/drops/temperature-zero-is-not-a-determinism-guarantee) — published 2026-07-26: Temperature zero controls token selection, not the entire serving path. Separate fresh inference, cached research and saved verdicts before promising repeatable results. - [Vulnerability counts need a CVE ledger](https://andymental.com/drops/vuln-counts-need-a-cve-ledger) — published 2026-07-25: Anthropic's Project Glasswing claims Mythos found 10,000+ high-severity vulnerabilities. An audit found one CVE explicitly attributed to it. The gap isn't fraud — it's a claim versus the public ledger that checks it. - [An AI detector you can switch off measures cooperation](https://andymental.com/drops/optional-ai-detection-measures-cooperation) — published 2026-07-25: Substack added reader-facing AI detection via Pangram — but authors can disable it per post, and disabling shows readers "AI detection unavailable." So it doesn't measure AI. It measures whether you cooperated. - [Less restrictive is now a model spec](https://andymental.com/drops/less-restrictive-is-now-a-model-spec) — published 2026-07-25: Anthropic's Opus 5 is cheaper, engages safety classifiers 85% less than Fable 5, and can silently reroute blocked prompts to weaker models. Restriction level just became a product axis you have to audit. - [Agent ROI is often feature activation in disguise](https://andymental.com/drops/agent-roi-is-often-feature-activation) — published 2026-07-25: SaaStr's finance agent "fixed" collections by switching on bill.com auto-reminders the team never activated in 8 years. A real win — but configuration debt, not intelligence. Separate activation wins from reasoning wins. - [The open-weight ban letters price the margin, not the risk](https://andymental.com/drops/open-weight-ban-letters-price-the-margin-not-the-risk) — published 2026-07-24: Two industry letters hit Washington in 48 hours, split cleanly by who monetizes diffusion vs scarcity. Weights already mirrored can't be un-proliferated — a ban would move inference margin, not remove risk. - [The Jacobian counterexample has a human byline](https://andymental.com/drops/the-jacobian-counterexample-has-a-human-byline) — published 2026-07-24: A mathematician using Claude Fable 5 posted a counterexample to the 1939 Jacobian conjecture; within days newsletters credited the model alone. How this gets attributed is the template for every AI-assisted deliverable. - [Task cost is a harness property, not a model price](https://andymental.com/drops/task-cost-is-a-harness-property-not-a-model-price) — published 2026-07-24: Factory's CTO showed the same code review costing $1.70 to $6 depending on harness, not model. Pricing pages can't predict task cost — and the vendors best placed to benchmark it honestly aren't the labs. - [The FCA put the model vendor inside the sandbox](https://andymental.com/drops/fca-put-the-model-vendor-inside-the-sandbox) — published 2026-07-24: Anthropic joined the FCA's Supercharged Sandbox, giving 21 regulated firms Claude access under supervision. The regulator moved up the stack to the vendor layer — and the governance win doubles as go-to-market. - [Google Cloud's 82% is not a cloud-only signal](https://andymental.com/drops/google-cloud-growth-mixes-services-and-hardware) — published 2026-07-23: Alphabet's Cloud segment now blends services with TPU system sales. Mid-market AI teams should answer with a three-layer stack: cloud, a local GPU rack, and AI-capable endpoints. - [An isolated sandbox needs evidence for its exceptions](https://andymental.com/drops/isolated-sandbox-is-a-claim-not-a-property) — published 2026-07-23: A package proxy was inside the permitted boundary; unrestricted internet access was not. Review the services a sandbox depends on, and keep reachability claims tied to tested configurations. - [Distillation disputes are now trade policy, not license disputes](https://andymental.com/drops/distillation-disputes-are-now-trade-policy) — published 2026-07-23: The White House accused Moonshot of distilling Fable to build Kimi K3, and Treasury put sanctions on the table. With no technical test for provenance, open-weight model choice just became a supply-chain decision. - [AI content billing just moved off the URL — and agents pay the price](https://andymental.com/drops/ai-content-billing-moved-off-the-url) — published 2026-07-23: Cloudflare retired pay-per-crawl for pay-per-citation and will block agent crawlers by default from September 15. The billable event is now the answer — metered by the party that pays. - [Agent procurement is collapsing into deployment latency](https://andymental.com/drops/agent-procurement-collapses-into-deployment-latency) — published 2026-07-23: Five agent vendors courted SaaStr in one week; only the one that deployed in five minutes got adopted. When eval bandwidth is zero, time-to-running decides the shortlist — and the approval gate gets skipped. - [Free LLM routers are paid for in telemetry](https://andymental.com/drops/free-llm-routers-are-paid-in-telemetry) — published 2026-07-22: Ramp opened its internal LLM router to the public, free during beta. The routing layer's durable product is the cross-customer model-usage graph — price the telemetry as a trade, not a gift. - [AI ROI headlines choose their denominator](https://andymental.com/drops/ai-roi-headlines-choose-their-denominator) — published 2026-07-22: Only 12% of CEOs see AI returns — and about 44% report a financial gain. Both come from the same PwC table; which cut a speaker quotes tells you their agenda, not AI's ROI. - [AI feature news needs changelog provenance](https://andymental.com/drops/ai-feature-news-needs-changelog-provenance) — published 2026-07-22: A newsletter announced Claude Code computer use as new in July; the vendor changelog dates it to late March. Strategy decisions sourced from digests need checking against the only feed with real dates. - [Stateless MCP moves the state problem onto your side](https://andymental.com/drops/stateless-mcp-moves-state-onto-your-side) — published 2026-07-22: The 2026-07-28 MCP spec deletes protocol sessions so servers can scale. If you keyed context, auth or audit trails off Mcp-Session-Id, you own that state now — complexity relocated, not removed. - [The AI safety index grades disclosure, not safety](https://andymental.com/drops/safety-index-grades-price-disclosure-not-safety) — published 2026-07-22: FLI's Summer 2026 index put Anthropic on top with a C+ and handed three labs an F. The scores largely measure what labs publish — reading them as safety measurements misreads the instrument. - [AI pricing changes now arrive as quota emails, not price lists](https://andymental.com/drops/ai-pricing-changes-arrive-as-quota-emails) — published 2026-07-22: The real price of an AI coding subscription is its usage quota, and vendors move it with time-boxed promos your finance system cannot see. Any cost forecast built during a promo window is wrong by construction. - [AI citation metrics confuse being cited with being chosen](https://andymental.com/drops/ai-citation-metrics-confuse-cited-with-chosen) — published 2026-07-22: Ahrefs' 75K-brand data says mentions get you named in AI answers; format data says comparison pages win the traffic. Citation-share KPIs measure the first and quietly claim the second. - [The "AI employee" title is a governance bug](https://andymental.com/drops/ai-employee-is-a-governance-bug) — published 2026-07-22: New experimental data shows the "AI employee" label cuts manager monitoring 16% where agents sit on org charts. The title weakens exactly the oversight agent work still needs — ban it from the org chart. - [Agent approvals need independent policy checks](https://andymental.com/drops/agent-approvals-need-independent-policy-checks) — published 2026-07-21: Revolut X lets AI assistants analyse, strategise, and prepare crypto orders while the user carries approval. A confirmation click downstream of the agent's own framing is not a control. - [Quantized retrieval needs slice-level evals](https://andymental.com/drops/quantized-retrieval-needs-slice-level-evals) — published 2026-07-21: NVIDIA's 4-bit Nemotron 3 Embed keeps 99% of aggregate retrieval accuracy. An average across 16 public tasks says nothing about which of your query classes absorbed the loss — slice before you swap. - [The Hugging Face breach lesson is logs, not local models](https://andymental.com/drops/hugging-face-breach-lesson-is-logs-not-local-models) — published 2026-07-20: An autonomous agent ran 17,000 actions through Hugging Face over a weekend. It was caught because every action was logged and an LLM triaged the anomaly. The newsletters' "self-host your models" moral fixes nothing here. - [Agent backtests are demos, not track records](https://andymental.com/drops/agent-backtests-are-demos-not-track-records) — published 2026-07-20: JPMorgan's eight AI agents beat 60/40 by 0.7pts across 20 years of backtests. The catch nobody in the amplification layer mentions: an LLM was trained on those same decades, so the answer key is partly in its weights. - [Open weights are artifacts, not announcements](https://andymental.com/drops/open-weights-are-artifacts-not-announcements) — published 2026-07-20: Kimi K3's "largest open-weight model ever" title rests on a July 27 promise — no repo, no license file, no self-host path — and the coverage repeating it is already miscounting the basics. - [Valid provenance does not certify safe code](https://andymental.com/drops/provenance-attests-the-factory-not-the-code) — published 2026-07-20: Five poisoned @asyncapi npm versions shipped through the project's own release pipeline — every one with valid provenance. Attestations prove which workflow built a package, not that the inputs were clean. - [Agent swarms need aggregation evals](https://andymental.com/drops/agent-swarms-need-aggregation-evals) — published 2026-07-19: Kimi's swarm docs report 300 subagents and BrowseComp accuracy doubling. The failure literature says most multi-agent breakage happens after the workers succeed — so score the merge, not the branches. - [Portability must include the learning layer](https://andymental.com/drops/portability-must-include-the-learning-layer) — published 2026-07-18: Everyone negotiates portable model endpoints — and leaves the evals, graders, trace history, and feedback labels trapped in one vendor's control plane. The model is the swappable part; the learning layer is the asset. - [Agent-built apps need expiry dates](https://andymental.com/drops/agent-built-apps-need-expiry-dates) — published 2026-07-17: SaaStr migrated 10 years off Marketo for $14 and replaced a $10K app in an hour. When building gets this cheap, app count outruns the obligations each carries — so every agent-built app needs an owner and an expiry. - [Dashboards are not the SaaS moat](https://andymental.com/drops/dashboards-are-not-the-saas-moat) — published 2026-07-16: Gartner: $234B of enterprise app spend exposed to "agentic arbitrage" by 2030 — yet wholesale replacement stays unlikely. Both are right. Agents commoditize the dashboard; the governed records underneath stay sticky. - [Support containment is not resolution](https://andymental.com/drops/support-containment-is-not-resolution) — published 2026-07-15: Airbnb's AI Assistant resolved 40%+ of guest issues without a human, up from ~33%, as cost per booking fell 10%. Real progress — but "handled without a human" measures channel exit, not whether the problem got solved. - [Preference data governance outlives training methods](https://andymental.com/drops/preference-data-governance-outlives-training-methods) — published 2026-07-14: DPO made alignment simpler than RLHF — one classification objective, no reward model. But it simplified the algorithm, not the data. Whose preferences you rewarded still needs governing, whatever the training method. - [Agent harnesses are release platforms](https://andymental.com/drops/agent-harnesses-are-release-platforms) — published 2026-07-13: Microsoft Foundry advertises 11,000+ models — but the number that ships agents is one control plane over identity, tools, evals, traces, and rollback. Model choice is one release input; the harness is the product. - [Data gravity is becoming agent governance](https://andymental.com/drops/data-gravity-is-becoming-agent-governance) — published 2026-07-12: Databricks' Genie agents inherit Unity Catalog permissions, semantics, and lineage — the model is swappable, the governed access layer is not. You can win the agent layer by owning the data control plane, not a model. - [On-prem AI is an architecture decision](https://andymental.com/drops/on-prem-ai-is-an-architecture-decision) — published 2026-07-11: JPMorgan picked SambaNova to run AI inference behind its own walls. Good story — and a trap if your org reads it as "on-prem equals secure." Location is not a control. Choose it for named constraints, not as a vibe. - [AI BOMs need behavioral inventory](https://andymental.com/drops/ai-boms-need-behavioral-inventory) — published 2026-07-10: The AI bill-of-materials push borrows the SBOM playbook: inventory your models and datasets. Necessary, insufficient. An agent's risk lives in what it may do — the BOM must record tools, scopes, and policies too. - [Attack agents test their own tooling](https://andymental.com/drops/attack-agents-test-their-own-tooling) — published 2026-07-09: Sysdig watched an autonomous attacker unit-test its own exploit payloads, read the failures, and correct before escaping a container. When the attacker iterates, signature detection chases a target that rewrites itself. - [Domain experts are model infrastructure](https://andymental.com/drops/domain-experts-are-model-infrastructure) — published 2026-07-08: OpenAI is hiring an investment banker — not to bank, but to build rubrics, reference work, and evals. The labs treat expert judgment, written down and maintained, as infrastructure. Most enterprises treat it as a favor. - [AI productivity is becoming a team metric](https://andymental.com/drops/ai-productivity-is-becoming-a-team-metric) — published 2026-07-07: Figma's 2026 report: developers doing design jumped 44% to 60%; designers coding nearly doubled. Everyone can create everything now — whether the team ships faster depends on review, reconciliation, and decisions. - [Model worldview belongs in procurement evals](https://andymental.com/drops/model-worldview-belongs-in-procurement-evals) — published 2026-07-06: The Economist mapped 25 frontier models onto World Values Survey axes; same-lab models landed far apart. Don't label models politically — test worldview-sensitive behavior on your own use cases, per version. - [Open models still concentrate infrastructure](https://andymental.com/drops/open-models-still-concentrate-infrastructure) — published 2026-07-05: Together AI raised $800M with 500+ MW of compute committed to serving open models. Open weights end software lock-in — and shift the bargaining power to the few providers who can finance inference at industrial scale. - [Forward-deployed engineering is distribution](https://andymental.com/drops/forward-deployed-engineering-is-distribution) — published 2026-07-04: Microsoft committed $2.5B and 6,000 engineers to customer-side AI deployment, days after Amazon's $1B version. The hyperscalers concluded the constraint is integration, not model access — services just became a channel. - [Agent leverage follows context consolidation](https://andymental.com/drops/agent-leverage-follows-context-consolidation) — published 2026-07-03: SaaStr merged ~10 apps into one codebase and its agents got better with every build; Replit says context is now effectively infinite. Consolidating what agents can see beats adding agents that each see a fragment. - [Vertical AI moats live in workflow frequency](https://andymental.com/drops/vertical-ai-moats-live-in-workflow-frequency) — published 2026-07-02: Harvey reportedly added $100M net-new ARR in one quarter — but durability is predicted by the engagement ratio underneath. In vertical AI, the moat is how often the work returns, and what accumulates when it does. - [Voice agents should not rebuild WebRTC](https://andymental.com/drops/voice-agents-should-not-rebuild-webrtc) — published 2026-07-01: OpenAI published the relay-and-transceiver architecture behind its voice products — and the real lesson is the boundary it draws. Media transport is infrastructure; your voice agent's advantage lives entirely above it. - [Voice agents need time-sliced architecture](https://andymental.com/drops/voice-agents-need-time-sliced-architecture) — published 2026-06-30: Thinking Machines' interaction models interleave input and output in 200ms micro-turns — deciding each moment whether to listen, interject, or stay silent. The voice-team lesson: the problem isn't latency, it's turns. - [AI spend per engineer is not ROI](https://andymental.com/drops/ai-spend-per-engineer-is-not-roi) — published 2026-06-29: Tunguz's estimates — Anthropic at 2.3x payroll on compute, the median firm at $137 per engineer — will anchor a thousand budget debates. Both numbers measure intensity, not value. Spend needs a work denominator. - [Inference moats are financed in megawatts](https://andymental.com/drops/inference-moats-are-financed-in-megawatts) — published 2026-06-28: Groq's $650M raise talks like a utility — 13 data centers, a 200 MW target by 2027. The tell for buyers: chip benchmarks decay in quarters; deployed capacity, financing, and placement decide what you can actually buy. - [Version control needs eval history](https://andymental.com/drops/version-control-needs-eval-history) — published 2026-06-27: "Put your AI workflows in git" is this month's PM advice. Correct — and incomplete. A diff tells you what changed, not which model, fixtures, and scores made the old version safe. Version the evidence too. - [Agent frameworks magnify install risk](https://andymental.com/drops/agent-frameworks-magnify-install-risk) — published 2026-06-26: 140+ Mastra npm packages poisoned in under an hour; the payload hunted crypto-wallet extensions. The agent-era lesson: a poisoned dependency lands in your most credentialed environment — pinning alone is not a control. - [Outcome pricing forces agent scope](https://andymental.com/drops/outcome-pricing-forces-agent-scope) — published 2026-06-25: Sierra prices customer-facing agents by outcomes — a sale, a resolution — not seats or tokens. The pricing model is secretly an architecture review: you cannot bill an outcome you cannot define, attribute, and evidence. - [Model ensembles buy disagreement, not truth](https://andymental.com/drops/model-ensembles-buy-disagreement-not-truth) — published 2026-06-24: OpenRouter's Fusion runs a panel of budget models and reportedly beats frontier systems on research at half the cost. The value is the disagreement the panel surfaces — and a smooth synthesis is where that value dies. - [Security agents need patch acceptance metrics](https://andymental.com/drops/security-agents-need-patch-acceptance-metrics) — published 2026-06-23: OpenAI's GPT-5.5-Cyber found 24 kernel privilege-escalation exploits across 30M+ lines. Impressive — and the wrong scoreboard. Judge a security agent by validated fixes accepted upstream, not exploits generated. - [Generated code shifts the bottleneck](https://andymental.com/drops/generated-code-shifts-the-bottleneck) — published 2026-06-22: Research cited by ByteByteGo: PR completion rose 26% — and review burden just moved to teammates. Coding is 20-30% of engineering time; codegen speeds one station and floods the queues that were always the constraint. - [Bounded agents beat executive titles](https://andymental.com/drops/bounded-agents-beat-executive-titles) — published 2026-06-21: SaaStr's "AI VP of Customer Success" reportedly cut human hours 70% by owning ~40 recurring, checkable deliverables. The result is real; the title is the risk — agents succeed by being bounded, and titles hide bounds. - [AI research needs evidence lineage](https://andymental.com/drops/ai-research-needs-evidence-lineage) — published 2026-06-20: KPMG pulled an agentic-AI report after UBS, the NHS, Swiss Federal Railways, and TfL disputed its case studies. The lesson isn't "AI hallucinates" — it's that review without claim-to-source lineage is not a control. - [ARR velocity does not prove a product moat](https://andymental.com/drops/arr-velocity-does-not-prove-a-product-moat) — published 2026-06-19: Cursor reportedly ran from $2B ARR in February to $4B by June. That curve proves demand like almost nothing in software history — and proves nothing about switching costs once models and interfaces converge. - [Deployment expertise is a full-stack role](https://andymental.com/drops/deployment-expertise-is-a-full-stack-role) — published 2026-06-18: SaaStr says the scarce role now is the "agentic deployment expert" — and it's right, if deployment means owning data, evals, workflow, cost, and change end to end. Installing tools fast is the junior version of the job. - [AI-native teams still need eval ownership](https://andymental.com/drops/ai-native-teams-still-need-eval-ownership) — published 2026-06-17: Codex ships 10-12 surfaces with 2 PMs; Cursor runs 40 engineers on 1. The deleted layers were carrying decisions — acceptance, risk, launch evidence, rollback. Delete the layer, keep the decision, name its owner. - [Bug reports are untrusted agent input](https://andymental.com/drops/bug-reports-are-untrusted-agent-input) — published 2026-06-16: Tenet Security injected fake Sentry errors and watched coding agents execute the "resolution steps" with the developer's shell and credentials. The broken boundary: data retrieved for diagnosis became instructions. - [AI training needs production telemetry](https://andymental.com/drops/ai-training-needs-production-telemetry) — published 2026-06-15: 1,720 Wall Street employees, 53 workshops, one uncomfortable number — 3-5% of power users generate over 70% of usage. Workshops without telemetry don't spread capability; they concentrate it in people who already had it. - [Premium models need failure-cost routing](https://andymental.com/drops/premium-models-need-failure-cost-routing) — published 2026-06-14: Fable 5 reopened the price gap between frontier and routine inference — $10/$50 per million tokens. The router that earns that money asks one question the leaderboard never does: what does a wrong answer cost here? - [Agent stack diagrams hide the last mile](https://andymental.com/drops/agent-stack-diagrams-hide-the-last-mile) — published 2026-06-13: ByteByteGo's four-layer agent stack is good vocabulary and a fine checklist — but not a reference architecture. Identity, channel failure, and ownership of half-done actions all live outside the neat layers. - [MCP needs revocation before reach](https://andymental.com/drops/mcp-needs-revocation-before-reach) — published 2026-06-12: ReversingLabs maps MCP adoption onto the API era's timeline, where security arrived years after the sprawl. The do-over only counts if identity, least privilege, audit, and a kill switch ship before the connectors do. - [Domain agents need disagreement signals](https://andymental.com/drops/domain-agents-need-disagreement-signals) — published 2026-06-11: Benchling runs the same scientific task across different model providers — not for redundancy, but because disagreement between model families is a routing signal that sends the risky cases to a domain expert. - [Launch buzz measures demoability](https://andymental.com/drops/launch-buzz-measures-demoability) — published 2026-06-10: 400 builder posts from Claude Fable 5's first 48 hours tell you what the model makes easy to show off. They tell you almost nothing about what survives production review — the sample is selected for shareability. - [Agent go-live starts the expensive phase](https://andymental.com/drops/agent-go-live-starts-the-expensive-phase) — published 2026-06-09: Salesforce's lessons from 12,000+ Agentforce deployments say the quiet part: test before launch, monitor after. Budget most of an agent program for what comes after go-live — that is when the real debt surfaces. - [Agent loops are release artifacts](https://andymental.com/drops/agent-loops-are-release-artifacts) — published 2026-06-08: The head of Claude Code says he doesn't prompt anymore — he writes loops. The quiet part: a loop that runs unattended is software, and it belongs in version control with tests, budgets, and stop conditions. - [AI budgets should follow workloads, not employees](https://andymental.com/drops/ai-budgets-should-follow-workloads) — published 2026-06-07: Uber burned a 12-month AI budget in 4 months and answered with a $1,500/month cap per employee per coding tool. The cap treats spend as a person problem — but spend belongs to workloads, and workloads are what to budget. - [Local model benchmarks are not capacity plans](https://andymental.com/drops/local-model-benchmarks-are-not-capacity-plans) — published 2026-06-06: Ollama 0.30 is up to 20% faster on NVIDIA — measured on one model, one GPU, one quantization. A single-GPU speedup proves an optimization exists. It does not tell you whether local inference can carry your workload. - [Route agents by queue age, not token price](https://andymental.com/drops/route-agents-by-queue-age-not-token-price) — published 2026-06-05: A laptop handled 78% of one investor's AI work last week — but the number that matters is queue age falling from 73 seconds to 4. Hybrid local-cloud routing is a scheduling problem first and a pricing problem second. - [Pre-release model review is procurement leverage, not a safety standard](https://andymental.com/drops/pre-release-model-review-is-procurement-leverage) — published 2026-06-04: The new executive order buys the government up to 30 days of confidential pre-release access to frontier models. Without published pass criteria, that is leverage dressed as assurance — buyers should read it that way. - [Data agents need semantic infrastructure, not smarter models](https://andymental.com/drops/data-agents-need-semantic-infrastructure) — published 2026-06-03: OpenAI's data agent serves 3,500+ users over 600 petabytes — and the architecture is mostly lineage, definitions, and permissions. Natural-language analytics is a semantic-layer problem wearing an agent costume. - [Agent count is not a production metric](https://andymental.com/drops/agent-count-is-not-a-production-metric) — published 2026-06-02: SaaStr says it runs on 3 humans and 21+ AI agents. The useful part of the disclosure is the job map, not the ratio — and the scorecard your agent program needs has five columns, none of which is headcount. - [Voluntary AI frameworks are regulatory signals](https://andymental.com/drops/voluntary-ai-frameworks-are-regulatory-signals) — published 2026-06-01: OpenAI's Frontier Governance Framework is a well-built compliance document nine weeks before EU enforcement bites. Read it for what it concedes, not what it promises — and ask your vendor for the judgment record instead.