
Domain experts are model infrastructure
OpenAI is hiring an investment banker — not to bank, but to build rubrics, reference work, and evals. The labs treat expert judgment, written down and maintained, as infrastructure. Most enterprises treat it as a favor.
OpenAI is hiring an investment banker. Not to raise money — to sit inside the Applied AI team and define the quality bar for AI-assisted banking work. The listing asks for at least two years of live transaction experience and pays $185,000 to $205,000 plus equity, three days a week in San Francisco. (The newsletter framing of "$500K to train AI" I could not verify from the listing itself; the equity presumably does the stretching.) Read the actual job description, because it is a template: design realistic hard tasks, create and assess reference work, build grading criteria, diagnose model failures, and map where AI should automate, where it should support, and where humans must stay in the loop — informed by knowing how judgment actually evolves from analyst to director.
Notice what this role is not. It is not an annotator seat — interchangeable, per-task, paid by the label. It is a specification and evaluation role: one expert whose tacit judgment gets converted, deliberately and durably, into assets a model team can reuse — rubrics, reference answers, failure taxonomies. The frontier labs have concluded that domain expertise is infrastructure, and they are staffing it like infrastructure: permanent, senior, expensive, and upstream of the product.
Now compare how the average enterprise treats the same resource. At work, the pattern I see in almost every deployment: the AI team asks a business expert to "take a look at" outputs — informally, in a meeting, as a favor squeezed between their actual job. The expert says "this one's wrong, the tone is off on that one," everyone nods, nothing is written down in reusable form. Six weeks later a model upgrade lands and the entire exercise repeats from scratch, with a different expert, producing different opinions, because nothing was captured the first time. The organization is paying for expert judgment repeatedly and banking none of it.
The difference between those two postures is not budget. It is the conversion step: tacit judgment becomes infrastructure only when it is written down as versioned rubrics (what does "good" mean for this task, in criteria a non-expert could apply), reference sets (worked examples of excellent, acceptable, and subtly-wrong outputs — the subtly-wrong ones are the gold), and a disagreement process (when two experts split, who decides, and does the rubric get amended). Do that once and the asset survives model upgrades, expert attrition, and vendor switches. Skip it and every evaluation is a séance.
The strategic point for enterprises: you need this capability even though you buy the model. OpenAI's banker will define banking quality for OpenAI's products — generic banking, at the industry's center of gravity. Your firm's definition of a good credit memo, your risk appetite, your client conventions live nowhere in that rubric. Buying the model outsources capability; it cannot outsource your acceptance bar. The enterprises that get durable value are building small internal versions of exactly this role — one respected practitioner per critical workflow, given real hours, producing maintained eval assets instead of meeting-room vibes.
Steal the job listing itself: take the OpenAI posting, replace "investment banking" with your critical domain, and use it as the role charter for your best practitioner — even at 20% allocation. First deliverables in month one: ten hard tasks from real work, a one-page rubric, five reference answers including two subtly-wrong ones, and a standing hour to adjudicate disagreements. That is the whole machine. It fits in a repo next to the eval records it feeds.
The labs are paying bankers to write down what good looks like — your experts already know it; the only question is whether it dies in meetings or compounds in a repo.


