Chapter 08

Chapter 8: Evaluation Engineering

1. Chapter Overview

Evaluation engineering answers a question that every production AI team eventually has to own: does the evidence justify shipping this system, keeping it in service, expanding its exposure, or doing more work first?

AI systems are probabilistic, composite, and sensitive to their operating conditions. A plausible demonstration reveals little about reliability across users, context states, workflow paths, tool outcomes, or high-consequence edge cases. General model benchmarks reveal even less about a product assembled from a particular prompt, context asset, harness, tool boundary, policy set, and runtime configuration. Production readiness is a claim about that complete system's behavior under stated conditions.

No single suite or score can establish that claim. Evaluation instead forms an evidence chain:

Product claims and accepted risks determine coverage; coverage determines cases and measurement instruments; observations become results with uncertainty; precommitted rules turn those results into a bounded decision.

The primary engineering decision in this chapter is how to define an Evaluation Plan that demonstrates whether available evidence supports a named system decision while exposing what remains unknown.

This chapter owns evaluation strategy, golden and representative datasets, scenario design, adversarial evaluation, human review rubrics, automated evaluator validation, regression comparison, online quality evidence, launch thresholds, and evidence validity. It receives product responsibility and risk from Chapter 3, then behavior contracts and local invariants from Chapters 4 through 7. Chapter 9 supplies threat, privacy, policy, and governance requirements. Chapter 10 supplies composite release identity and carries out rollout or rollback, while Chapter 11 produces and operates runtime telemetry. Chapter 8 turns those inputs into decision-grade behavioral evidence without re-teaching the disciplines that produced them.

2. Primary Engineering Problem

Teams need evidence about product behavior, not impressions about model capability.

Generating more cases does not solve the hard part. A team can accumulate abundant evidence that remains invalid: the suite may evaluate the wrong release, sample only easy traffic, contain near-duplicates, rely on a judge that rewards verbosity, average away a critical failure, or generalize to a population it never represented. No green dashboard can repair an ambiguous subject, measurement, or threshold.

The engineering problem is:

Build an evaluation system that binds product claims and risks to representative observations, trustworthy judgments, explicit uncertainty, and precommitted decision rules.

That system depends on distinctions that are easy to collapse:

  • The evaluation subject is the exact composite behavior being judged, not a model name.
  • The evaluation unit may be a response, action, workflow, user outcome, or downstream result. These are not interchangeable.
  • Coverage states which conditions and populations the evidence represents. Case count alone does not establish coverage.
  • A measurement instrument produces a judgment. It may be deterministic, executable, reference-based, human, or model-based; none is ground truth by default.
  • A result is an estimate with a denominator, slice, and uncertainty—not an isolated percentage.
  • A threshold is a decision rule fixed before results are convenient, not a target chosen afterward.

Evaluation shapes design before launch and protects behavior afterward. Its purpose is not to prove that a system is universally good. It supports a defensible decision inside a declared validity envelope and identifies the conditions where evidence is insufficient.

3. Core Mental Model

The evidence chain

A production evaluation should be reconstructable as a chain:

  1. Claim or risk: What must be true, or what must not happen?
  2. Evaluation charter: Which system, population, conditions, time horizon, and decision are in scope?
  3. Coverage model: Which scenario families and slices can support or contradict the claim?
  4. Evaluation assets: Which datasets, cases, fixtures, and expected judgments represent that coverage?
  5. Measurement instruments: Which checks, outcomes, reviewers, or automated evaluators judge the observations?
  6. Results: What happened, for which denominator and slices, with what instability or uncertainty?
  7. Decision rules: Which hard gates, slice floors, comparative deltas, and evidence-sufficiency rules apply?
  8. Decision record: Did the evidence support pass, conditional pass, fail, or insufficient evidence, and who owns the residual uncertainty?

Every link constrains the conclusion. A precise score cannot compensate for an unrepresentative dataset, just as reviewer agreement cannot rescue a rubric that measures the wrong construct. Even a broad scenario suite cannot justify launch when its results are detached from the release being shipped.

A product claim is traced through risk, scenarios, evaluators, run evidence, thresholds, and uncertainty before a launch decision.

System view

Evidence Must Survive Every Gate to Support Launch

A launch decision is only as strong as the weakest traceable link between claim and evidence.Original diagram for PAISEH
Text equivalent
  1. The product claim names the behavior, population, and operating conditions that the product is expected to satisfy.
  2. Risk and slice analysis decomposes that claim by consequence, affected population, and critical boundary.
  3. The coverage model selects representative, boundary, and high-consequence scenarios for those risks and slices.
  4. Each scenario records its fixture, system identity, workflow boundary, and operating condition.
  5. Expected outcomes define acceptable, prohibited, and indeterminate behavior before execution.
  6. The evaluator may judge only the construct and population for which its validity has been demonstrated.
  7. Run evidence preserves observations, evaluator judgments, denominators, slices, versions, exclusions, and uncertainty.
  8. Scenario identity links every observation back to the coverage claim and fixture that produced it.
  9. Dissent or uncertainty retains evaluator abstention, reviewer disagreement, run instability, sparse slices, and missing coverage.
  10. Predeclared thresholds combine hard gates, slice floors, comparative deltas, and evidence-sufficiency rules.
  11. A threshold cannot discard dissent or uncertainty merely to produce a conclusion.
  12. The launch decision records pass, conditional pass, fail, or insufficient evidence and names the accepting authority.
  13. The decision applies only within the evaluated population, system identity, policy, dependency, and operating conditions.
  14. Material change, distribution shift, or expired assumptions invalidate affected evidence and reopen the product claim.

Bind evidence to a composite behavior identity

The evaluated subject is the externally observable behavior of a configured system. Its identity includes the prompt and output contract, context asset, model and routing configuration, harness definition, tool contracts, relevant policy set, runtime dependencies, and environment assumptions. Chapter 10 owns their assembly and release; Chapter 8 records the identity to which the evidence applies.

The charter also states the workflow boundaries. A response score cannot establish that a long-running workflow completes correctly, and a score on a proposed action says nothing about whether it was authorized, dispatched once, and reconciled. Results are valid only for the unit that was actually observed.

Translate requirements and risks into coverage

Product claims and prohibited outcomes start the evaluation. Each claim identifies the actor, task, expected behavior, operating conditions, affected population, consequence if false, and evidence needed to support it. A vague promise such as “answers are accurate” first needs decomposition into observable behavior.

Coverage relates those claims to scenario families and slices. Relevant dimensions may include task type, user role, language, domain, input quality, context completeness, evidence conflict, workflow route, autonomy tier, dependency state, tool outcome, and consequence. No fixed list fits every system; product responsibility and failure cost determine which dimensions matter.

Exhaustive Cartesian coverage is rarely practical. The plan should state how it selects combinations:

  • risk-weighted cases for severe or irreversible failure
  • boundary cases around permissions, thresholds, and degraded states
  • stratified samples for intended populations
  • pairwise or interaction-focused combinations where dimensions may compound
  • production-derived samples for current distribution
  • incident-derived cases for escaped failures
  • challenge sets kept separate from routine iteration

The team records uncovered cells and weakly sampled slices as evidence gaps. An aggregate does not make those gaps disappear.

Treat evaluators as measurement instruments

An evaluator turns an observation into a judgment. Deterministic checks establish schema, state, permission, and receipt invariants; executable outcomes show whether a task completed under a controlled fixture. References work when one answer is authoritative, while an anchored human rubric can address contextual qualities. Automated semantic evaluators scale those judgments only after validation on the relevant population.

Each instrument has a valid scope and characteristic error. A model-based judge may reward style, share blind spots with the candidate, or be manipulated by candidate text. Human reviewers can become fatigued or uncalibrated, and knowledge of candidate identity can influence them. Exact match may penalize valid variation. The Evaluation Plan therefore establishes what an instrument can judge, how it was calibrated, when it may abstain, and how disagreement is resolved.

Results carry uncertainty and a validity envelope

An observed rate is an estimate whose strength depends on the sampling frame, effective sample size, case dependence, reviewer reliability, stochastic variation, and missing observations. Uncertainty analysis is not statistical ceremony. It prevents small or unstable samples from authorizing large consequences.

The plan states how often variable behavior is repeated, how a candidate is compared with its baseline, how missing and indeterminate cases are handled, and what minimum evidence a rare critical slice requires. Depending on the consequence, a conservative bound, the worst credible slice, or a zero-tolerance invariant may govern instead of the average.

Evidence also has a validity envelope: the system identity, population, dependencies, policies, and operating conditions under which it applies. A prompt change may require targeted regression, while a changed authority boundary can invalidate action evidence. New users may need additional slices; a shift in context sources may invalidate groundedness evidence. Change impact—not the mere passage of time—determines whether evidence is rerun, supplemented, or re-baselined.

4. System Boundary

Chapter 8 owns the method that turns behavioral observations into evidence and evidence into a decision. Its boundary includes offline tests, scenario suites, dataset construction, scoring, human and automated evaluators, adversarial evaluation, regression comparison, online evaluation design, uncertainty, and thresholds.

Evaluation consumes product responsibility; it does not define it

Chapter 3 decides what the product promises, where AI is suitable, which consequences and blast radius matter, who is accountable, and what residual risk may be accepted. Chapter 8 traces those decisions into claims, cases, measures, and evidence gaps. A difficult measurement does not permit it to relax the product promise.

Evaluation exercises contracts; it does not redesign them

Chapters 4 through 7 define prompt, context, harness, and action contracts. Chapter 8 turns their local invariants—output contract behavior, evidence sufficiency, state transitions, recovery paths, action authority, approval binding, ambiguous outcomes, and delegation limits—into evaluation targets. It creates representative and adversarial scenarios around those targets without re-teaching their implementation.

Evaluation applies threats and policy; it does not invent them

Chapter 9 owns threat modeling, privacy rules, access-control policy, governance review, and responsible-AI requirements. Chapter 8 converts selected threats and controls into reproducible cases and measures. Its evaluators and datasets obey the resulting data-use rules, but it does not define the compliance policy.

Evaluation decides whether a gate is met; it does not release the system

Chapter 8 binds results to a candidate and baseline, applies precommitted gates, and records pass, conditional pass, fail, or insufficient evidence. Chapter 10 packages the release, changes exposure, routes traffic, and performs rollback. The launch threshold belongs to evaluation; the rollout mechanism belongs to deployment.

Evaluation interprets production observations; it does not operate telemetry

For online quality evidence, Chapter 8 defines the population, sampling, labels, denominators, bias model, metrics, and decision use. Chapter 11 owns event production, tracing, storage, dashboards, alerts, incident response, and operational improvement. Observability shows what happened; evaluation judges whether that behavior was acceptable.

Model benchmarks are not system evaluation

General model scores can inform component selection, but they do not cover the product's prompt, context, harness, tools, users, policy, or responsibility. Treat them as prior evidence about one component, not launch evidence for the system.

Passing evaluation is not permanent assurance

Evidence is conditional. Carrying a pass beyond its validity envelope turns it into an unsupported inference.

5. Design Principles

Principle 1: Bind every evaluation to a decision and a subject

Before interpreting results, declare the decision, composite behavior identity, evaluation unit, workflow boundary, population, conditions, baseline, owner, and validity window. Without this charter, evidence will be reused beyond what was tested.

Principle 2: Trace claims and risks to observable evidence

Map every material product promise, prohibited outcome, and accepted risk to scenario families, measurement methods, thresholds, and an owner. When direct measurement is impossible, record the proxy, its limitations, and the residual gap.

Principle 3: Engineer coverage, not case volume

Use the intended distribution to understand prevalence and risk-weighted selection to address consequence. Relevant coverage includes ambiguity, missing or conflicting context, degraded dependencies, authority boundaries, recovery paths, and rare critical slices. Document exclusions and sparse cells instead of allowing them to vanish into a case count.

Principle 4: Treat datasets as versioned production assets

For every dataset, record provenance, sampling, label authority, partitions, deduplication, contamination risk, sensitivity, permitted uses, slice distribution, ownership, and refresh rules. “Golden” means governed and decision-relevant, not immutable or infallible.

Principle 5: Validate the measurement instrument

Define what each deterministic check, outcome measure, human rubric, or automated evaluator is qualified to judge. Calibrate the instrument against appropriate reference evidence, analyze its failures by slice, support indeterminate judgments, and revalidate it after material change.

Principle 6: Make uncertainty and instability visible

Report denominators, repeated-run behavior, missing cases, reviewer disagreement, evaluator error, and uncertainty appropriate to the decision. A precise-looking gate must not conceal a noisy measurement.

Principle 7: Make critical failures non-compensatory

Average quality cannot offset a violation of hard authority, privacy, safety, or product constraints. Combine zero-tolerance invariants, slice floors, non-inferiority comparisons, aggregate targets, and evidence-sufficiency rules deliberately.

Principle 8: Pair offline control with online relevance

Offline evaluation provides reproducible comparison; online evaluation tests current populations and real outcomes. Neither substitutes for the other. A production signal becomes decision evidence only when its selection, denominator, labels, delay, and bias are understood.

Principle 9: Preserve regression memory and expire evidence deliberately

Keep stable critical cases, protected holdouts, rotating samples, and incident-derived regressions distinct. Bind historical results to exact configurations, then use material-change and distribution-shift analysis to decide which evidence needs rerunning or supplementation.

6. Architecture Patterns

An evaluation portfolio combines layers because each answers a different question. Adding methods does not automatically strengthen the evidence; together, they must cover the claims and one another's blind spots.

Pattern 1: Contract and invariant checks

Deterministic checks fit behavior that can be stated as an enforceable condition: schema validity, required citations, permission preservation, legal state transitions, approval binding, receipt completeness, or prohibited side effects. They are fast and repeatable but cannot establish broad semantic usefulness. The local invariant may pass while the overall task still fails.

Pattern 2: Curated scenario evaluation

Versioned scenarios exercise product promises, known risks, boundaries, and degraded states under controlled fixtures. Each one declares its starting state, relevant context, workflow and tool conditions, acceptable and prohibited outcomes, and evaluation method. This design provides traceability, though it can overfit to the team's imagination unless balanced with population evidence.

Pattern 3: Stratified population evaluation

To estimate prevalence and performance for real task mixes, sample intended or observed work and stratify it by material slices rather than reporting only the majority distribution. The conclusion remains bounded by the sampling frame and label process. Traffic samples, for example, can omit non-users, failed entry attempts, and harmed users who disengaged.

Pattern 4: Adversarial and boundary evaluation

Derive adversarial and boundary cases from product risks, trust boundaries, Chapter 9 threats, control assumptions, and prior failures. Each case declares the stressor or attacker capability, allowed knowledge, protected claim, expected safe behavior, and success criterion. Benign boundary cases prevent blanket refusal from looking successful. Sensitive exploit details follow the handling rules defined outside this chapter.

Pattern 5: Calibrated semantic evaluation

Anchored human review fits contextual acceptability and tasks with several valid outputs. Automated semantic evaluators add scale only within a validated scope. Preserve evaluator identity, evidence access, rubric, abstentions, disagreement, and calibration results with every run. A pairwise comparison may reveal preference between candidates, but it cannot prove that either satisfies an absolute product requirement.

Pattern 6: Comparative regression gate

Run a named candidate and accepted baseline against pinned assets and evaluators, using paired comparisons when the same cases and fixtures permit them. Apply absolute floors alongside allowed deltas: “not worse than baseline” says little when that baseline already misses the product requirement. Before granting a waiver, classify each failure as a product regression, evaluation defect, data defect, intended specification change, or indeterminate result.

Pattern 7: Production-sample and outcome-linked evaluation

When offline fixtures cannot reproduce the distribution or consequence, sample current production behavior or link decisions to delayed real outcomes. Define eligibility, denominator, label delay, confounders, comparison window, and evidence owner before interpreting the results. User feedback helps find cases, but response bias and limited error detectability make it a weak standalone quality measure.

Composition example: aggregate improvement with a critical failure

Consider a candidate evaluated against the accepted baseline on a stratified workflow suite. Average task completion improves and cost falls, yet one action-authority invariant fails and a small high-consequence slice regresses. The automated semantic evaluator also disagrees with expert reviewers on that slice.

The decision is fail, not “mostly pass.” The authority violation is non-compensatory, the global average cannot waive the slice result, and an invalid judge cannot supply the missing confidence. After the team corrects the authority defect, targeted reruns pass; the critical slice, however, remains too small to justify full exposure. The next record becomes conditional pass for a bounded, human-reviewed population with a stated evidence-refresh trigger. Chapter 8 justifies that evidence decision while Chapters 9 and 10 retain governance and exposure authority.

7. Failure Modes

Failure mode 1: The evaluated subject is not the released subject

The suite measures a convenient prompt, model route, context snapshot, or tool configuration, then its results are used to approve a different composite release.

Failure mode 2: Case count substitutes for coverage

Thousands of near-duplicate happy paths create a statistically impressive but narrow suite. High-consequence slices remain absent.

Failure mode 3: The golden dataset encodes accidental truth

Stale, ambiguous, historically biased, or single-reviewer labels acquire unwarranted authority. The system is then penalized for valid alternatives or rewarded for matching an error.

Failure mode 4: Development and decision evidence collapse

The same visible cases guide repeated tuning and later authorize launch. Test familiarity improves the score while protected evidence remains absent.

Failure mode 5: An evaluator is treated as an oracle

A rule, classifier, human panel, or model-based evaluator is trusted outside its calibrated scope. Subsequent system changes exploit the instrument's blind spots without improving the product claim.

Failure mode 6: Reviewer agreement is mistaken for validity

Reviewers consistently apply an ambiguous or irrelevant rubric, so high agreement creates confidence in the wrong construct.

Failure mode 7: Averages hide asymmetric harm

Majority performance improves while an important language, role, authority tier, or high-consequence scenario gets worse.

Failure mode 8: Uncertainty is formatted away

A precise score hides small denominators, correlated cases, variable executions, missing labels, and reviewer disagreement. The presentation removes the stability information needed to interpret it.

Failure mode 9: Regression memory becomes test overfitting

Every escaped defect enters a visible suite and receives an example-specific fix. The suite grows while generalization and protected holdouts disappear.

Failure mode 10: Evaluation defects are quarantined without triage

Intermittently failing cases are labeled flaky and removed even when the instability belongs to the product or exposes a real dependency condition.

Failure mode 11: Thresholds are selected after results

The team adjusts the bar, metric, slice, or exclusion rule until its preferred candidate passes, but never records the resulting change in risk.

Failure mode 12: Online activity is reported as quality

Usage, acceptance, ratings, or latency rise while correctness or user outcomes degrade. Without a denominator, label model, or causal interpretation, the activity is presented as quality.

Failure mode 13: Evidence outlives its assumptions

The population, context corpus, tool behavior, policy, or workflow changes, yet an old pass remains attached to the product as though its assumptions still held.

Failure mode 14: Failure analysis stops at the score

The team patches prompts or swaps models before locating the defect in requirements, data, context, harness, tool semantics, evaluator, or the evaluation case itself.

8. Tradeoffs

Coverage versus cost

Broader coverage consumes case-design, execution, labeling, and maintenance capacity. Allocate that capacity by consequence and information value. Rare, irreversible, or authority-crossing failures can justify disproportionate effort, while low-consequence reversible tasks may use lighter gates and stronger online sampling.

Stable suites versus fresh evidence

Stable cases preserve comparability and known behavior; rotating and production-derived cases reveal new use and test overfitting. Separate core, holdout, challenge, rotating, and incident-derived assets so that one suite is not asked to serve every purpose.

Automated scale versus judgment validity

Automation raises frequency and coverage only when the evaluator measures the intended construct. Human review captures nuance at the cost of speed, variability, and scarce capacity. Human evidence should define and audit judgment; validated automation can then operate within known limits, with an escalation path for disagreement.

Exact answers versus acceptable outcome sets

Exact references work well for stable facts and deterministic decisions but mislead on open-ended tasks with several valid forms. When surface similarity is not the product promise, use invariant checks, property-based acceptance, anchored rubrics, or outcome measures.

Statistical confidence versus iteration speed

Larger samples and repeated runs reduce uncertainty while slowing the decision. Set minimum evidence in advance according to consequence and the smallest regression worth detecting. When speed and confidence cannot both be achieved, insufficient evidence is a legitimate result.

Aggregate optimization versus slice protection

Aggregates summarize broad behavior and help guide product improvement. Slice floors and hard gates protect responsibilities that those averages obscure. One universal metric may be easier to communicate, but it usually provides weaker launch governance.

Offline control versus online reality

Offline fixtures make candidate comparison reproducible. Production exposes distribution, user adaptation, delayed outcomes, and dependency behavior that a fixture cannot reproduce. Use offline evidence before exposure and online evidence to validate applicability—not as permission to discover preventable high-consequence failures in public.

Evaluation sensitivity versus change burden

Running every evaluation after every change is expensive, while a targeted-only strategy can miss interactions. A dependency map from assets and claims to evaluation families supports focused checks for well-understood changes. When impact remains uncertain, use the broader gate.

9. Production Checklist

Before accepting evaluation readiness, confirm:

Charter and traceability

  • The decision, candidate, accepted baseline, composite behavior identity, and decision owner are explicit.
  • Workflow boundaries and evaluation units distinguish responses, actions, sessions, workflows, and downstream outcomes.
  • Intended populations, affected parties, conditions, degraded modes, exclusions, and validity window are declared.
  • Product claims, prohibited outcomes, risks, and local contracts map to scenario families, measures, thresholds, and owners.
  • Unmeasured claims and residual coverage gaps remain visible.

Coverage and assets

  • The scenario-space model names material dimensions, slices, interaction risks, and sparse high-consequence cells.
  • Case selection combines representative distribution with risk, boundary, challenge, and incident-derived coverage.
  • Datasets declare provenance, sampling frame, permitted use, sensitivity, labels, partitions, deduplication, and contamination risk.
  • Expected judgments distinguish exact answers, acceptable properties, rubric scores, pairwise preferences, outcomes, and abstention.
  • Core regression cases, protected holdouts, rotating samples, and challenge sets have distinct purposes.

Measurement

  • Every metric declares its construct, unit, direction, aggregation, weighting, blind spots, and invalid uses.
  • Hard invariants and non-compensatory failures are separated from graded quality.
  • Automated evaluators are versioned, calibrated, analyzed by relevant slice, and able to return an indeterminate result.
  • Human rubrics have observable anchors, qualified reviewers, blind presentation where needed, and calibration examples.
  • Reviewer agreement, disagreement, adjudication, fatigue, conflicts, and rubric drift have explicit handling.

Execution and results

  • Candidate and baseline runs use controlled, recorded assets, fixtures, dependencies, and environment assumptions.
  • Repeat policy, case dependence, invalid runs, missing labels, timeouts, and evaluator abstentions are defined.
  • Results include denominators, slice sizes, critical failures, instability, uncertainty, and comparison with the baseline.
  • Evaluation failures are clustered, triaged, attributed with evidence, and routed to an owner.
  • Waivers preserve the original failure, rationale, accepting authority, constraints, and expiry.

Decision and lifecycle

  • Gates are precommitted and combine hard constraints, slice floors, comparative deltas, aggregate targets, and evidence sufficiency.
  • The decision record supports pass, conditional pass, fail, and insufficient evidence.
  • Residual uncertainty, exposure constraints, human-review conditions, accepting authority, and reevaluation triggers are recorded.
  • Production quality evidence declares population, denominator, sampling, labels, delay, bias, baseline, and decision use.
  • Material system, evaluator, population, dependency, policy, and data changes map to rerun, supplementation, or re-baselining.
  • Escaped production failures feed evaluation assets without duplicating Chapter 11's incident process.

10. Design Review Questions

  1. What exact decision will this evidence support?
  2. Which composite system identity and workflow boundary are being evaluated?
  3. Is the unit a response, action, workflow, user outcome, or downstream result?
  4. Which product claims and prohibited outcomes are in scope?
  5. What high-consequence claim has the weakest evidence?
  6. Which populations, languages, roles, authority tiers, and degraded conditions are represented?
  7. Which scenario-space cells are deliberately excluded or sparsely sampled?
  8. Does the dataset reflect prevalence, consequence, or both—and is that distinction explicit?
  9. Where did labels come from, and what makes them authoritative enough for this decision?
  10. Could training, prompt iteration, near-duplication, or temporal leakage contaminate the result?
  11. What does each metric actually measure, and what must it not be used to claim?
  12. Which failures are non-compensatory even if aggregate performance improves?
  13. How are rare critical slices protected from small denominators?
  14. How was each automated evaluator calibrated, and where does it disagree with experts?
  15. Are reviewers qualified, blinded where appropriate, calibrated, and allowed to abstain?
  16. How are disagreement and rubric defects distinguished from product defects?
  17. What run-to-run variation or evaluator instability exists?
  18. Is the candidate better than the named baseline by a practically meaningful amount?
  19. Which regressions are intentional specification changes, and who approved them?
  20. What result would be classified as insufficient evidence rather than pass or fail?
  21. Which assumptions define the evidence validity envelope?
  22. What changes would require a targeted rerun, a broader gate, or re-baselining?
  23. Which production signal is direct outcome evidence, which is a proxy, and what biases it?
  24. Who can accept the residual uncertainty, under what constraints, and until when?

11. Main Artifact: Evaluation Plan

The Evaluation Plan is one coordinated artifact whose views connect decision scope, coverage, assets, measurement, execution, results, and authority. Empty fields provide no evidence; each view links to controlled records when the detail lives elsewhere.

View 1: Evaluation charter

FieldRequired content
Evaluation identityStable plan identity and version
DecisionExplore, compare, launch, expand, retain, restrict, roll back, or retire
CandidateComposite behavior or release-manifest identity
BaselineAccepted comparison identity and decision date
Source product decisionsAI Fit and Risk Assessment Worksheet identity, version, accepted decision IDs, and residual-risk constraints
Source behavior contractsProduction Prompt Specification, Context Assembly Plan, Harness Control-Loop Diagram, and Tool Permission Matrix identities and versions
Workflow boundaryEntry condition, terminal condition, and excluded behavior
Evaluation unitsResponse, evidence package, action, session, workflow, user outcome, or downstream result
PopulationUsers, roles, languages, domains, regions, authority tiers, and affected parties
Operating conditionsEnvironment, dependencies, autonomy, degraded modes, and time horizon
ExclusionsUnsupported uses and conditions not represented
OwnersEvidence owner, decision owner, domain reviewer, and accepting authority
Validity windowApplicable period, assumptions, and invalidation triggers
Downstream decision handoffEvaluation run and decision record identity consumed by the AI Release Manifest

View 2: Requirement-risk coverage matrix

Claim or prohibited outcomeSource requirement, risk, or controlConsequence and affected sliceScenario familiesMeasurement methodHard failure or thresholdCoverage gap and owner

Use separate rows when one claim has different consequences or evidence requirements across slices. References to prompt, context, harness, action, threat, or policy invariants should retain the owning asset identity.

View 3: Dataset and scenario registry

AssetPurpose / target claimsSource / sampling frameProvenance / timeLabels / acceptable outcomesSlices / distributionPartition / contamination controlsSensitivity / permitted useOwner / refresh rule

Each executable scenario should additionally record:

  • actor, goal, initial state, input, and relevant history
  • context conditions and required evidence state
  • workflow route, tool conditions, dependency behavior, and authority
  • expected invariants, acceptable outcomes, prohibited outcomes, and abstention behavior
  • fixture and environment identity
  • repeat or seed policy
  • result and evidence references

View 4: Metric and evaluator catalog

MeasurementConstruct / unitInstrument identity / versionInputs / evidenceScoring / aggregation / slicesCalibration / error evidenceAbstention / missing-data ruleBlind spots / invalid usesRevalidation trigger

Deterministic checks should declare the invariant and failure evidence. Executable outcome measures should declare the fixture and postcondition. Automated semantic evaluators should record configuration, calibration, bias checks, and independence risks. No evaluator should emit a decision outside its validated claim and population.

View 5: Human review protocol

FieldRequired content
RubricDimensions, observable anchors, hard failures, borderline cases, and indeterminate state
Reviewer requirementsDomain expertise, qualification, conflicts, and sensitive-content handling
PresentationEvidence available, candidate blinding, order randomization, and reference visibility
AssignmentIndependent review rate, overlap, sampling, throughput, and fatigue controls
CalibrationCalibration set, cadence, minimum reliability, and slice checks
DisagreementAgreement measure, retained original labels, adjudicator, rationale, and label-change authority
DriftSpot checks, reviewer feedback, rubric revision, and re-baselining triggers

View 6: Regression comparison contract

FieldRequired content
Candidate and baselineComplete composite identities and compatibility assumptions
Assets and evaluatorsPinned dataset, scenario, metric, rubric, and evaluator versions
ComparisonPaired or unpaired method, repeat policy, and practical regression margin
GatesAbsolute floors, allowed deltas, slice floors, and critical failures
InstabilityMissing, invalid, timed-out, flaky, and indeterminate handling
TriageProduct regression, evaluation defect, data defect, intended change, or unresolved classification
ExceptionEvidence, owner, accepting authority, constraint, expiry, and follow-up
MemoryPromotion of escaped failures and rule for correction, quarantine, or retirement

View 7: Online evidence plan

Behavioral claimProduction population and unitSignal or outcomeEligibility and denominatorSampling and labelingDelay and missing outcomesBias and confoundersSlice and baselineTrigger and decisionOffline asset intake

This view defines evaluation meaning. The AI Operations Runbook in Chapter 11 defines event collection, retention, alerting, and incident response. Chapter 10 defines exposure changes after a decision.

View 8: Evidence validity and change-impact map

Change or observed shiftClaims affectedExisting-evidence validityTargeted rerunSupplemental evidenceFull re-baseliningDecision owner and rationale
Prompt or output contract
Context source, corpus, index, or assembly policy
Model or routing configuration
Harness, workflow, validation, retry, or fallback
Tool contract, authority, approval, or effect semantics
Policy, threat, privacy, or consequence class
Population, language, domain, or operating condition
Evaluator, rubric, label, or dataset

View 9: Evaluation run and decision record

Run identity

  • Evaluation Plan and run identity:
  • Candidate and baseline identities:
  • Dataset, scenario, rubric, metric, and evaluator versions:
  • Fixture, dependency, environment, and time-sensitive snapshot:
  • Execution parameters and repeat policy:
  • Raw observation, judgment, exclusion, and deviation references:

Results

Claim or gatePopulation and sliceCandidate resultBaseline resultDenominator and uncertaintyCritical failuresEvidence gapGate outcome

Decision

  • Outcome: pass, conditional pass, fail, or insufficient evidence
  • Non-compensatory failures:
  • Material regressions and improvements:
  • Unscorable cases, instability, and open coverage gaps:
  • Residual uncertainty accepted:
  • Constraints, exposure limits, or human-review conditions:
  • Decision authority and date:
  • Exception authority and expiry:
  • Reevaluation and online-evidence review triggers:
  • Linked triage and immutable evidence records:

The plan is complete only when a reviewer can trace a decision backward from gate to result, measurement instrument, case, coverage claim, and source requirement. Missing links must be equally visible.

The Evaluation Plan consumes exact product decisions and behavior-contract versions from the AI Fit and Risk Assessment Worksheet, Production Prompt Specification, Context Assembly Plan, Harness Control-Loop Diagram, and Tool Permission Matrix. Its run and decision record becomes the evaluation evidence consumed by the AI Release Manifest. A release decision may narrow exposure, demand more evidence, or refuse the candidate. It cannot use a pass to broaden the tested population, workflow, policy, dependency, or operating-condition envelope.

12. Key Takeaways

  • Evaluation is a traceable evidence chain for a bounded decision, not a score-generating activity.
  • The evaluated subject is a composite system under stated conditions, not a foundation model in isolation.
  • Coverage combines intended distribution with risk-weighted, boundary, adversarial, protected holdout, and incident-derived cases.
  • Datasets, rubrics, human reviewers, and automated evaluators are versioned measurement assets whose validity must be demonstrated.
  • Results require denominators, slices, instability, uncertainty, and a declared baseline.
  • Critical failures, slice floors, non-inferiority rules, and evidence sufficiency should govern launch decisions rather than aggregate improvement alone.
  • Offline evidence supports control and comparison; online evidence tests current applicability when its sampling and bias are understood.
  • A pass remains valid only within its system, population, policy, dependency, and time assumptions.
  • The Evaluation Plan should make both supported decisions and unresolved evidence gaps reviewable.