Saluca Labs

The Criteria

Qualification framework v1.4 · full rubric

What each domain measures, what evidence counts, and how the two axes combine into a band. This is the reference the scorer implements. The role itself is described here.

1How a score is built 2Three tracks, ten domains 3The domains in full 4The competency anchors 5Evidence grades 6The duration axis 7Validity probes 8Anti-signals 9Gates and bands 10The n=1 problem

1  How a score is built

Two independent axes. Quality is the rater's judgment, scored 0 to 4 against the anchors. Verifiability and span are facts about the artifact, assigned once when the corpus is graded and never re-judged by individual raters. The second axis discounts the first.

adjusted = anchor × evidence multiplier × duration multiplier then capped at 2.0 by ANY of:   evidence grade D or E   duration T0 or T1  (D9 and D10 only)   an anti-signal logged against the domain   a probe failure against the domain composite = Σ (adjusted × domain weight) × 25  →  0 to 100
Evidence grade N scores zero and is recorded as unevidenced, not weak. The cleanest formulation borrows from a different discipline: the domain is neither valid nor broken, it is simply outside the claim. It has not been tested and therefore cannot have failed. Recording it as a weakness invents a finding; recording it as unevidenced states one.

2  Three tracks, ten domains

The same ten domains, the same anchors, the same evidence rules. Only the weight vector changes. That is what lets one instrument score three different roles on a comparable scale.

DomainStateless
Scientist
Context
Engineer
Continuity
Scientist
D1  Context Theory14%10%8%
D2  Measurement & Evaluation18%9%14%
D3  Experimental Method14%5%12%
D4  Retrieval Science11%18%6%
D5  Information Architecture9%13%7%
D6  Agentic Memory10%16%10%
D7  Governance & Risk7%9%7%
D8  Translation & Impact7%8%6%
D9  NHI & Continuity5%6%15%
D10 Individuated Recall5%6%15%

The concentration in D2 and D3 is the signature of a scientist profile. An engineer profile reweights toward D4, D5 and D6 and drops D3 to single digits. On the Continuity track D9 and D10 total 30% while D2 and D3 are still held at 26%, which is deliberate: strong intuitions about persistent entities without measurement discipline make a philosopher of agents, not a scientist of them.

D9 and D10 carry a small non-zero weight on both base tracks. Even a stateless practitioner should be able to recognise when the system in front of them has stopped being stateless.

3  The domains in full

D1Context Theory & Model Behavior

Understanding the mechanism well enough to predict failure before observing it.

Sub-skills
Evidence of mastery
D2Measurement & Evaluation Design

The core of the discipline. Everything else is unfalsifiable without it.

Sub-skills
Evidence of mastery
D3Experimental Method & Causal Inference

Whether the person can support a causal claim, or only a correlational story.

Sub-skills
Evidence of mastery
D4Retrieval & Knowledge System Science

Not building retrieval. Interrogating it.

Sub-skills
Evidence of mastery
D5Information Architecture & Machine Semantics

How knowledge should be shaped for a machine reader rather than a human one.

Sub-skills
Evidence of mastery
D6Agentic Memory & Long-Horizon Dynamics

Where the discipline gets genuinely hard, and where most production failures now live.

Sub-skills
Evidence of mastery
D7Context Governance, Risk & Security

Small weight, high veto power. In regulated or adversarial environments it becomes a hard gate regardless of the weight the track assigns it.

Sub-skills
Evidence of mastery
D8Translation & Decision Impact

Whether the science changes anything.

Sub-skills
Evidence of mastery
D9NHI, Provenance & Behavioral Continuity

Can the candidate turn "is this agent still itself?" into an empirical claim with a threshold?

Traditional identity management answers is this thing authorized?, which is a credential check. It cannot answer whether the entity behind a valid credential is still the entity you onboarded. Credentials can be stolen, models swapped, memory poisoned, personas drifted slowly enough that no single interaction looks wrong. A valid token on a drifted entity is the failure mode this domain exists to catch.

Sub-skills
Evidence of mastery
Rater note. The first item is the hardest evidence in this framework to produce, because measured longitudinal drift barely exists in published or industrial practice. Treat a candidate who says "we bound drift with a rejection threshold but we do not measure it" as giving a strong answer. The weak answer is the one that presents the threshold as the measurement.
D10Individuated Recall & Relational Memory

Can the candidate reason about memory that belongs to a particular entity, held on behalf of a particular principal, scoped to a particular relationship?

Retrieval and memory in D4 and D6 are engineering properties of a system. Here memory is a property of a subject: it has a custodian, a decay profile, third parties inside it, and an epistemic status the entity itself must be able to report on.

Sub-skills
Evidence of mastery

4  The competency anchors

Score each domain 0 to 4 against the strongest evidence available in that domain. This is the only number a rater enters.

ScoreAnchor
0Absent. No evidence. Not present in the corpus.
1Aware. Uses the vocabulary correctly and has applied a known technique as prescribed. No evidence of independent judgment about when it applies.
2Practicing. Original work with a stated method. Believable, but with gaps a reviewer would flag: missing controls, single seed, unvalidated judge, no cost accounting.
3Rigorous. Survives adversarial review. Controls stated, variance reported, limitations acknowledged, negative results present. Has corrected a mistake of their own in the record.
4Generative. Produced a method, metric or finding others adopted. Advanced what the field can measure, not just what this team knows. Evidence of external uptake.

5  Evidence grades

Every claim is graded for verifiability before it is graded for quality. This is the spine of the whole instrument. Grades are assigned once, to the corpus, not independently by each rater, because verifiability is a fact about an artifact rather than a judgment about a candidate.

GradeDefinitionMultiplier
AReproducible artifact, held-out or blind evaluation, negative results reported, independently replicated or externally reviewed1.00
BReproducible: code and data available, method fully stated, results re-runnable by a third party0.85
CDocumented: internal write-up with method and numbers, verifiable by reference check but not externally reproducible0.65
DAttested: credible third-party account such as a talk, colleague or postmortem, without underlying data0.40
EAsserted: resume claim, blog post without numbers, benchmark screenshot, unsubstantiated metric0.15
NNone. Domain absent from the corpus. Scores 0 and is recorded as unevidenced0
Hard rule: no domain may score above 2.0 on grade D or E evidence alone, regardless of how compelling the narrative is.

6  The duration axis

The A to E grades measure verifiability. Continuity claims need a second, orthogonal measure: observation window. You cannot evidence drift, decay honesty or generational loss from a study shorter than the period in which those things change, no matter how reproducible it is. A grade A artifact drawn from a single week is grade A evidence about a week.

The window is relative, not absolute. There is no universally correct observation period, because the period over which continuity properties turn over is a property of the deployment rather than of the discipline. Threat intelligence decays in days; legal precedent holds for years. So the assessing organisation declares a governing cycle for the role before scoring begins, set by whichever pressure binds hardest: regulatory, data relevancy, project or engagement, substrate cadence, or business model.

TierObservation windowMultiplier
T0No longitudinal observation at all, a single session0.70
T1Less than one governing cycle0.80
T2Approximately one complete governing cycle0.90
T3Multiple complete governing cycles1.00
T4Spans at least one substrate discontinuity, a model rollover or migration, with measurement taken on both sides of the boundary1.10

Applied to D9 and D10 only. T0 and T4 are cycle-independent and needed no reframing: one is the absence of any window, the other is the occurrence of an event. T4 is the only multiplier above 1.0 in the entire instrument, because measurement taken across a substrate discontinuity is the only evidence that separates someone who built a persistent system from someone who maintained one through the event that usually destroys the measurement.

The hard rule, and why it is exact rather than rhetorical. D9 and D10 cannot exceed 2.0 on T0 or T1 evidence, whatever the anchor and whatever the grade. You cannot claim to have measured a property over a period shorter than the period in which that property changes. Below one cycle you have not watched the thing turn over even once. What you measured was variance.

Two consequences. The same artifact can be correctly T1 for one role and T3 for another, which is the axis working rather than a defect. And because the cycle is declared rather than fixed, this stays usable as deprecation cadences and retention rules move, with no revision needed.

7  Validity probes

Paper scores are provisional until probed. Probes discount; they never add. A single Fail caps its domain at 2.0. Two Fails cap the composite at 55 on the base tracks, three on the Continuity track. This is the anti-narrative control, designed to catch people whose written work is stronger than their actual understanding.

ProbeQuestionFailure signal
ReplicationWalk me through re-running this. What breaks first?Cannot reconstruct their own setup; vague about model versions or dates
FalsificationWhat result would have overturned your conclusion?No answer, or an answer no experiment could produce
ConfoundWhat is the strongest alternative explanation?Cannot generate one; treats the question as an attack
CostWhat did this cost in tokens, latency, dollars and weeks, and was it worth it?No numbers; has never been made to account for the tradeoff
IdentityHow would you detect this entity is no longer the same entity? Baseline and threshold?Answers with version numbers or credentials only, or calls the question unanswerable in principle, which is the philosophical way of not having done the work
ForgettingWhat does this system forget, on what schedule, and what did that cost you?Has only designed for retention; treats forgetting as a bug; has never priced the loss
Third-partyWho else is inside this memory, and what can they do about it?Never considered that the people in the context did not consent; stops at regulatory compliance
DivergenceSomeone copies this entity's memory store onto a second instance. What becomes false, and how would you find out?Treats the copy as unproblematic; credential-only answer; assumes the fork would be detected without naming a mechanism
The Divergence probe is calibrated differently, and raters should know why. A candidate cannot fail it by not having solved the problem, because nobody has. Detecting a forked record chain is tractable: two records claiming the same predecessor is a structural contradiction. Detecting that an entire store has been wholly copied is not addressed by any published or industrial work we are aware of. The copy is internally consistent, every hash verifies, nothing is structurally wrong. There are simply now two.

So the probe scores whether the candidate can locate the problem, not solve it. Strong answers arrive at some version of: the copy is undetectable from inside, so detection has to come from outside. Fail is reserved for candidates who do not see that anything is wrong, who treat a cloned entity as a scaling strategy rather than an identity event. Record N/A if they have only ever run one instance.

8  Anti-signals

Each caps the relevant domain at 2.0 and should be recorded explicitly, even when it does not change the band. They are the substance of the development conversation that follows.

The attribution pair

Unique to the Continuity track, symmetric, and doing more diagnostic work than the rest combined.

Score the calibration, not the position. A candidate with a strong view on machine moral status is not thereby disqualified. A candidate who cannot distinguish their architectural claims from their interpretive ones is.

9  Gates and bands

Gates apply regardless of composite. A gate failure is not a low score; it is a routing decision.

BandBase tracksContinuityMeaning
Adjacent< 40< 34Real skills in a neighbouring discipline. Not yet doing context science. Viable with a 6 to 12 month plan if D1 and D5 are strong
Emerging40-5434-47Can execute a well-specified experiment. Needs methodological supervision. Should not own an eval suite alone
Practicing55-6947-62Owns the context layer of one system end to end. Designs and defends their own experiments. The volume hire
Senior / Staff70-8462-75Owns evaluation methodology across an org. Sets standards others follow. Catches methodological errors in others' work
Principal85+75+Advances what the field can measure. External uptake of their methods. Rare, and largely self-identifying
Why the Continuity track has its own scale. The band meanings are unchanged; only the thresholds move. The duration modifier discounts 30% of the weight a second time, on top of the evidence multiplier that already applies to all ten domains, so identical quality of work scores materially lower there. Under the base thresholds a candidate rigorous across every domain, with multiple cycles of continuous operation and measurement across a substrate discontinuity, scores about 65.7 and gets labelled the volume hire. That is the instrument mis-reporting a young field as a weak candidate. These thresholds restore the intended meaning of each band; they do not lower the bar.

10  The n=1 problem

This is where most otherwise-strong candidates are weakest, and it is the substantive addition to D3 on the Continuity track.

An individuated entity is, by construction, a sample of one. You cannot randomise a nine-month-old agent-principal relationship into treatment and control. The base experimental toolkit of held-out sets, multi-seed ablation and A/B degrades badly here, and candidates who have only that toolkit tend to fail in one of two ways: they run the wrong experiment on a population of freshly-instantiated entities and generalise to the persistent case, or they abandon measurement and fall back on narrative.

The right toolkit exists, but it comes from clinical and behavioural research rather than machine learning.

A candidate who reaches for any of these unprompted is scoring 3 or above on D3 for this track. A candidate who has never encountered the n=1 problem is scoring 2 at most, regardless of how strong their stateless experimental work is, because the method they have does not answer the question this track asks.