Saluca Labs
Saluca Labs

The Context Scientist

A role definition

August 2026


Every organisation now runs systems whose behaviour depends on what enters a model's context window. Almost none can say whether what enters it is true, sufficient, or worth its cost. That question is a discipline, and it is currently nobody's job.

One role builds the pipe. This one asks whether what is flowing through it is true.

The role, against what it is not

A gap between the people who build the system and the people who use it

The adjacent roles all exist and all of them are hiring. None is accountable for the empirical question. The gap is structural rather than a staffing oversight, which is why it has stayed open.

Prompt engineer
Shapes the instruction. A craft, judged by whether it reads well and works today.
Context engineer
Builds the pipe: retrieval, memory, metadata, tool descriptions. Construction, judged by whether it runs.
ML engineer
Changes the model. Owns the half of the system that is rented and swappable by design.
Data scientist
Interrogates a dataset that holds still. The closest ancestor, and the assumption that breaks.
Context scientist
Asks whether what enters the window is true, sufficient and worth its cost, then supports the answer with evidence that survives an adversarial reviewer.
The separation is the same one that split data engineer from data scientist around 2012. It took roughly three years for the second title to look obvious in hindsight.

What they produce

The role is legible by its artifacts

None of these is a document about a decision. Each one is the decision itself, in a form somebody else can check.

What breaks without one

Failures most teams will already recognise

The value of the role is easiest to see as a list of things that happen when it is vacant. Each of these gets attributed to something else at the time.

The last one is the expensive failure and the reason the role exists. Everything above it is recoverable inside a quarter.

Evidence and span

A claim is bounded by how long you watched

No amount of rigour inside a short window extends it. Work drawn from a single week is excellent evidence about a week. Most teams discover this only after shipping.

The window that matters is not a fixed number of days. It is measured against the governing cycle, the period over which the thing being measured actually turns over in that specific deployment.

T0

A single session. No longitudinal observation at all.

T1

Less than one cycle. You have not watched the thing turn over even once, so what you measured was variance.

T2

Approximately one complete cycle.

T3

Multiple complete cycles. Enough to separate a trend from a fluctuation.

T4

Spans a substrate discontinuity, a model rollover or migration, with measurement taken on both sides of it. The rarest and most diagnostic evidence in the field.

Why the boundary is held strictly: a system that survived a model rollover with no measurement across it is not T4. It is T3 with an anecdote.

Reach across domains

The method is invariant. The context is not.

A context scientist asks the same four questions everywhere: what counts as ground truth, what cycle governs it, what "sufficient" means here, and what it costs to be wrong. The questions never change. The answers are unrecognisable across domains, which is exactly why the role generalises and the person does not.

DomainGround truth isGoverning cycleCost of being wrong
Security operationsThe incident outcome, known only afterwardsdaysA breach that was visible in the logs
Clinical supportPatient outcome, against a guideline that movesthe care episodeHarm to a person
Legal practiceControlling authority, not persuasive textthe matterSanction, or malpractice
Financial servicesReconciliation and disclosurethe review intervalEnforcement action
Industrial and roboticsThe physical result, which does not negotiatethe maintenance intervalDowntime, or injury
Customer operationsResolution, not deflectionthe relationshipChurn, with the reason never recorded

Knowing what counts as ground truth in oncology requires oncology. Knowing the review interval in a regulated practice requires having worked under the regulator who sets it. These are earned rather than learned in a quarter, so a context scientist is a specialist in one domain's context and a generalist in method.

Why the role is plural

Solutions span domains. Cycles do not align.

Real problems cross these boundaries. A clinical system takes a regulatory feed. A security platform consumes legal determinations. A customer operation touches all four. Each component can be correct against its own notion of sufficient and the composition still fails, because the failures live in the handoff and the handoff belongs to nobody.

Threat intelligencecycle: days
Product and engineeringcycle: the quarter
Legal and regulatorycycle: the annual review

The defect nobody owns. Legal determinations refresh once. Threat intelligence refreshes roughly ninety times against them. For all but a few days of the year the legal context feeding that system is stale, and neither specialist can see it, because each is looking at a component that is working correctly. The mismatch is only visible from the seam.

Working together

What a group of them does that none of them can do alone

Put several context scientists on a problem spanning their domains and they make four decisions jointly. Each is invisible from inside any single domain, and each is normally made by accident.

Contribution

What each domain puts in the window, and what gets dropped

The window is finite and every domain believes its material is essential. Somebody has to measure what each contribution is worth at the margin and remove what is not earning its place. This is nearly always adversarial, and should be.

Authority

Whose ground truth wins when two of them conflict

When the clinical answer and the regulatory answer disagree, the system needs a rule, and the rule needs validating rather than merely stating. Resolution that quietly takes the most recent write is a decision made by an implementation detail.

Cadence

What the composed system's governing cycle actually is

It is the shortest of its inputs, not the average and not the most important. A system is only as current as its fastest-decaying dependency, and this is the single most commonly missed fact about a composed solution.

Exposure

What the composite costs, and who pays when it is wrong

Each domain prices its own errors. Nobody prices the errors that exist only in combination, which are the ones that reach a customer, a patient or a regulator.

The shared method is what makes it a team

They cannot check each other's facts. A clinical context scientist cannot audit a colleague's threat model and should not pretend to. What they can check is each other's method: the evidence grade behind a claim, the window it was observed over, the confound that was ruled out, the cost that was counted.

That shared method is the whole reason a group of specialists becomes a team rather than a committee. It is also why the discipline needs a name. You cannot staff for a method nobody has named, and you cannot hold people to a standard that has never been written down.

Recognition

How to tell you are looking at one

Full assessment is a six-hour protocol. These five questions are what survives when there is only an hour, ordered by how hard they are to fake.

Can they show you a negative result?
A portfolio with no failures across a multi-year career has been curated rather than accumulated.
Can they name a claim of their own they now believe is wrong?
This does more discriminating work than anything else on the list, and it cannot be prepared convincingly.
Do they distinguish a control from a measurement?
A threshold that rejects bad writes bounds a problem. It does not measure it. Conflating the two is the sharpest line between building and measuring.
Do they state cost in the same breath as quality?
Tokens, latency, dollars, engineering weeks. Somebody never made to account for the tradeoff has never owned one.
Do they know their governing cycle, and who sets it?
Answering in absolute time rather than in cycles means they have not yet noticed that the window is a property of the deployment.
Expect the profile to be jagged. Nobody has been trained into all of this, because until recently there was nothing to be trained into. The most useful output of an assessment is usually not a score. It is the list of things that turned out to be unevidenced, which tells you what you will have to build regardless of who you hire.

The instruments

Two companion pages carry the assessment itself

This page describes the role. The rubric behind it is published in full, and the scorer implements it directly so that nothing about how a band is reached is hidden from the person being assessed.

The criteria →

What each of the ten domains measures, with sub-skills and the evidence that demonstrates mastery. Competency anchors, evidence grades, the duration axis, all eight probes, the anti-signals, gates and bands.

The scorer →

The working instrument. Enter anchors and the arithmetic follows: caps shown as they fire with the reason named, gates as pass or fail, and the band on the right scale for the track. Runs entirely in your browser and sends nothing anywhere.