← back to the archiveCover illustration for “Model worldview belongs in procurement evals”
ESSAYday 36·3 weeks ago·by Andy Padia

Model worldview belongs in procurement evals

The Economist mapped 25 frontier models onto World Values Survey axes; same-lab models landed far apart. Don't label models politically — test worldview-sensitive behavior on your own use cases, per version.

Tomasz Tunguz surfaced an Economist analysis this week that ran 25 frontier models through the World Values Survey — the questionnaire that has mapped moral attitudes across roughly 100 countries since 1981 — and plotted them on its two classic axes: traditional-to-secular values, and survival-to-self-expression. The headline finding is the title's joke: the models cluster overwhelmingly in one quadrant, the secular, self-expression corner populated by rich Western countries. Tunguz adds the structural explanation candidates: Common Crawl is about 46% English, and alignment work happens where the labs are.

Two details deserve more attention than the headline. First: models from the same lab landed far apart. The analysis reportedly places some sibling models as near-strangers on the map while models from rival labs sit as neighbors — which demolishes the lazy heuristic that vendor identity predicts behavioral disposition. Training data and alignment choices, not the logo, set the worldview. Second, the honest caveat: I could not verify the Economist's full methodology, and everything about this genre is sensitive to prompt framing, survey language, and judge choice. A model answering a values questionnaire is performing a task, not confessing a soul.

Which is exactly why the practical conclusion is not "label your models politically." It is narrower and more useful: worldview-sensitive behavior is a product property, it varies between models and between versions of the same model, and almost nobody procures for it.

Where worldview touches revenue

The reflex objection — "we use it for code, who cares" — is fair as far as it goes. Worldview is invisible in a SQL query. But walk the actual enterprise surface area. Policy and advice: a model drafting HR guidance or financial recommendations makes assumptions about authority, risk tolerance, and family structure with every paragraph. Moderation: what counts as offensive, blasphemous, or harmless banter is precisely a values call, made thousands of times a day at the edge of your brand. Customer support: deference versus directness, individual versus family framing — the difference between a reply that lands in Mumbai and one that lands in Munich. Localized products: a recommendation engine that quietly optimizes for self-expression values will feel subtly foreign in markets organized around other priorities.

In each case the failure mode is the same: nothing errors, nothing crashes, and the output is fluently, confidently misaligned with the market it serves. The team discovers it after deployment, from complaints, in the most expensive possible classroom.

rendering diagram…

The procurement move: a behavior suite, not a label

At work, the version of this I now push into every regulated deployment: a locally grounded behavior suite alongside the capability evals. Not a politics quiz — a set of forty-odd scenarios drawn from the product's real surface, written with people from the markets it serves. The support reply to a customer invoking family obligation. The moderation call on the religiously-inflected complaint. The advice draft where risk appetite matters. Each scenario with an acceptance rubric written by someone who actually knows the market, not inferred from a survey quadrant.

Three rules make the suite work. Run it per model version — the same-lab-far-apart finding means every upgrade is a fresh roll of the dice on disposition, so the suite runs on every version bump, exactly like a regression test. Track drift, not ideology — the output is "version N handles authority-framing differently than N-1 in these six cases," a diff a product owner can act on, not a political score nobody can defend in a meeting. Ground it locally — a suite written entirely in English by the head-office team measures the head office's worldview twice. If the market matters enough to serve, it matters enough to hire three hours of a local reviewer's time.

The client conversation that convinced me: a team shipping an assistant into three markets ran precisely one values-adjacent test — the vendor's own safety demo — and discovered in production that the model's handling of deference and formality read as condescending in one market and evasive in another. Same model, same prompts, two different failures. The fix was two weeks of suite-building that should have been procurement week one; the damage was a quarter of brand repair.

Steal this

Before the next model selection or upgrade, write ten scenarios where your product touches authority, religion, family, money, or risk — the WVS's own load-bearing themes — in each market you serve. Get acceptance rubrics from someone local to each. Run them against the incumbent and the candidate, diff the behavior, and file the results with the eval record. One day of work, repeated per version, filed with the eval history so the next upgrade decision starts from evidence. The Economist's map is a conversation starter; your suite is the procurement document.

Every model ships with a worldview whether you tested for it or not — swing the compass before the voyage, and re-swing it every time the needle gets replaced.

#evaluation#procurement#alignment#enterprise-ai#localization
← older drop
Open models still concentrate infrastructure
newer drop →
AI productivity is becoming a team metric

related drops

explore all 78 drops →
← back to the archiveday 59