August 2026
Every organisation now runs systems whose behaviour depends on what enters a model's context window. Almost none can say whether what enters it is true, sufficient, or worth its cost. That question is a discipline, and it is currently nobody's job.
One role builds the pipe. This one asks whether what is flowing through it is true.
The adjacent roles all exist and all of them are hiring. None is accountable for the empirical question. The gap is structural rather than a staffing oversight, which is why it has stayed open.
None of these is a document about a decision. Each one is the decision itself, in a form somebody else can check.
The value of the role is easiest to see as a list of things that happen when it is vacant. Each of these gets attributed to something else at the time.
No amount of rigour inside a short window extends it. Work drawn from a single week is excellent evidence about a week. Most teams discover this only after shipping.
The window that matters is not a fixed number of days. It is measured against the governing cycle, the period over which the thing being measured actually turns over in that specific deployment.
A single session. No longitudinal observation at all.
Less than one cycle. You have not watched the thing turn over even once, so what you measured was variance.
Approximately one complete cycle.
Multiple complete cycles. Enough to separate a trend from a fluctuation.
Spans a substrate discontinuity, a model rollover or migration, with measurement taken on both sides of it. The rarest and most diagnostic evidence in the field.
A context scientist asks the same four questions everywhere: what counts as ground truth, what cycle governs it, what "sufficient" means here, and what it costs to be wrong. The questions never change. The answers are unrecognisable across domains, which is exactly why the role generalises and the person does not.
| Domain | Ground truth is | Governing cycle | Cost of being wrong |
|---|---|---|---|
| Security operations | The incident outcome, known only afterwards | days | A breach that was visible in the logs |
| Clinical support | Patient outcome, against a guideline that moves | the care episode | Harm to a person |
| Legal practice | Controlling authority, not persuasive text | the matter | Sanction, or malpractice |
| Financial services | Reconciliation and disclosure | the review interval | Enforcement action |
| Industrial and robotics | The physical result, which does not negotiate | the maintenance interval | Downtime, or injury |
| Customer operations | Resolution, not deflection | the relationship | Churn, with the reason never recorded |
Knowing what counts as ground truth in oncology requires oncology. Knowing the review interval in a regulated practice requires having worked under the regulator who sets it. These are earned rather than learned in a quarter, so a context scientist is a specialist in one domain's context and a generalist in method.
Real problems cross these boundaries. A clinical system takes a regulatory feed. A security platform consumes legal determinations. A customer operation touches all four. Each component can be correct against its own notion of sufficient and the composition still fails, because the failures live in the handoff and the handoff belongs to nobody.
The defect nobody owns. Legal determinations refresh once. Threat intelligence refreshes roughly ninety times against them. For all but a few days of the year the legal context feeding that system is stale, and neither specialist can see it, because each is looking at a component that is working correctly. The mismatch is only visible from the seam.
Put several context scientists on a problem spanning their domains and they make four decisions jointly. Each is invisible from inside any single domain, and each is normally made by accident.
The window is finite and every domain believes its material is essential. Somebody has to measure what each contribution is worth at the margin and remove what is not earning its place. This is nearly always adversarial, and should be.
When the clinical answer and the regulatory answer disagree, the system needs a rule, and the rule needs validating rather than merely stating. Resolution that quietly takes the most recent write is a decision made by an implementation detail.
It is the shortest of its inputs, not the average and not the most important. A system is only as current as its fastest-decaying dependency, and this is the single most commonly missed fact about a composed solution.
Each domain prices its own errors. Nobody prices the errors that exist only in combination, which are the ones that reach a customer, a patient or a regulator.
They cannot check each other's facts. A clinical context scientist cannot audit a colleague's threat model and should not pretend to. What they can check is each other's method: the evidence grade behind a claim, the window it was observed over, the confound that was ruled out, the cost that was counted.
That shared method is the whole reason a group of specialists becomes a team rather than a committee. It is also why the discipline needs a name. You cannot staff for a method nobody has named, and you cannot hold people to a standard that has never been written down.
Full assessment is a six-hour protocol. These five questions are what survives when there is only an hour, ordered by how hard they are to fake.
This page describes the role. The rubric behind it is published in full, and the scorer implements it directly so that nothing about how a band is reached is hidden from the person being assessed.
What each of the ten domains measures, with sub-skills and the evidence that demonstrates mastery. Competency anchors, evidence grades, the duration axis, all eight probes, the anti-signals, gates and bands.
The working instrument. Enter anchors and the arithmetic follows: caps shown as they fire with the reason named, gates as pass or fail, and the band on the right scale for the track. Runs entirely in your browser and sends nothing anywhere.