P0 Scenario Pack · PAISEH-SC-MIA-1.0

Turn production signals into bounded operational decisions.

Use this pack when AI helps summarize monitoring evidence, support incident command, or assemble recovery proof. The assistant organizes authorized evidence; named operators retain incident state, response, communication, and return-to-service authority.

Minimum viable pack

Contract evidence, define meaning, route action, verify recovery.

Missing telemetry is an explicit operating state, not a zero. Generated summaries remain distinguishable from observed facts, operator decisions, controlled actions, effect receipts, and recovery evidence.

Blank working assets

One inspectable path from signal contract to verified recovery.

Preview every asset before use, copy it into the authoritative work location, or download the Markdown. Each runbook names entry conditions, authority, action, verification, stopping, reversal, evidence, and closure.

Start here

Scenario Pack Guide

Defines the advisory boundary, minimum pack, roles, authority, entry and exit criteria, required signal coverage, deeper-record triggers, and reopening rules.

Preview
# Monitoring and Incident Assistant Scenario Pack

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Use this pack when an AI-assisted workflow summarizes production signals, assembles incident evidence, proposes bounded response options, or helps verify recovery. It supports operators; it does not declare an incident, change severity, authorize containment, expose protected telemetry, execute a rollback, or close an incident unless a separately recorded human authority and controlled path permit that action.

The pack turns monitoring and incident work into six inspectable records. It preserves the distinction between observed evidence, generated interpretation, operator decisions, external effects, and verified recovery.

## Smallest Credible Starting Set

For advisory monitoring over a bounded service with no action authority, complete:

1. `ai-observability-contract.md`
2. `signal-objective-catalog.md`
3. `trace-event-schema.md`
4. `alert-specification.md`
5. `incident-evidence-brief.md`
6. `recovery-verification-record.md`

Attach every runbook whose entry condition is plausible. The assistant may summarize and recommend within the recorded evidence boundary. Any containment, configuration change, traffic shift, rollback, tool action, user communication, or incident-state transition requires the applicable authoritative workflow and receipt.

## Users, Owners, and Authority

| Role | Responsibility | Authority boundary |
| --- | --- | --- |
| Service owner | Owns product behavior, service objectives, consequence, and accepted degraded modes | Defines meaning and escalation policy; does not infer recovery from a generated summary |
| Telemetry owner | Owns collection, joins, redaction, retention, freshness, completeness, and evidence health | May repair or invalidate telemetry under policy; cannot redefine product quality |
| Signal owner | Defines one signal's population, denominator, evidence class, threshold, and allowed use | May revise the signal through change control; cannot declare an incident automatically unless policy explicitly binds that rule |
| On-call operator | Receives alerts, validates evidence, and invokes approved response paths | May take only actions granted by the current operations policy |
| Incident commander | Owns incident state, severity, priorities, decisions, and command handoffs | Declares, reclassifies, and closes the incident under policy |
| Domain lead | Interprets system-specific behavior and validates hypotheses | Advises command; cannot broaden action authority |
| Response operator | Executes an authorized containment, rollback, or recovery action | Acts only within the bound target, duration, and reversal conditions |
| Evidence scribe | Maintains the authoritative timeline, sources, hypotheses, decisions, actions, and gaps | Records evidence; does not convert assistant output into fact |
| Recovery verifier | Tests the named technical, behavioral, state, action, user, telemetry, and delayed-outcome claims | Accepts or rejects recovery evidence; does not close the incident unless also named as closure authority |
| Communications owner | Approves affected-user and stakeholder communications | Controls audience and message; generated drafts are not approved notices |
| Decision authority | Accepts residual risk, degraded operation, return to service, or incident closure | Records scope, conditions, expiry, and evidence for the disposition |

One person may hold several roles when policy permits. Record active assignments, handoffs, time windows, and separation-of-duties requirements in the incident system of record.

## Required Inputs and Entry Criteria

- A named service, workflow boundary, consequence class, owning team, on-call route, and incident system of record.
- Current component identities for application code, prompts, model or dependency resolution, context sources or indexes, tool policy, routing, runtime configuration, and evaluation or telemetry logic.
- Defined request, session, workflow, step, attempt, trace, span, action, effect, outcome, release, and incident correlation fields.
- Objectives and signals with explicit population, denominator, window, evidence class, freshness, completeness, delay, missing-data behavior, and owner.
- Externally enforced access, redaction, retention, sampling, and audit rules for telemetry and incident artifacts.
- An alert route with severity, deduplication, suppression, minimum evidence, allowed automation, and a tested response path.
- Approved containment, degraded-mode, rollback, cancellation, reconciliation, communication, and recovery-verification authorities.
- A safe behavior for missing, delayed, contradictory, or unjoinable telemetry.

Do not treat the assistant as ready when component versions cannot be reconstructed, critical telemetry can disappear as zero, protected payloads are broadly exposed, an alert has no credible operator decision, or recovery can be declared from infrastructure health alone.

## Working Sequence

| Stage | Required record | Result |
| --- | --- | --- |
| Contract | AI Observability Contract | System boundary, identities, capture rules, evidence health, views, access, and retention |
| Define | Signal and Objective Catalog | Interpretable technical, behavioral, economic, retrieval, action, safety, drift, and telemetry-health signals |
| Correlate | Trace and Event Schema | Versioned event types, correlation fields, effective component identity, effects, outcomes, and schema health |
| Route | Alert Specification | Actionable detection rules with evidence gates, severity, ownership, deduplication, and safe automation |
| Investigate | Incident Evidence Brief | Authoritative timeline, facts, hypotheses, affected scope, decisions, actions, gaps, and current risk |
| Verify | Recovery Verification Record | Evidence for every recovery dimension, residual risk, observation window, and closure authority |

## Completed Walkthrough

Use `examples/context-index-quality-regression.md` to inspect all six records completed for a citation regression after a context-index release. The example preserves the line between detection, hypothesis, incident authority, release action, recovery recommendation, and closure.

## Assistant Boundary

The assistant may retrieve authorized evidence, identify gaps, group related signals, draft a timeline, label hypotheses, compare observed state with recorded objectives, propose runbook steps, and assemble a recovery evidence packet. Its output remains attributed and reviewable.

Controls outside the model must authenticate the caller, authorize telemetry and payload access, enforce redaction and retention, bind queries to approved scopes, record effective component versions, constrain tools, require action approval, produce effect receipts, and preserve incident-state authority. A confident summary cannot replace missing telemetry, authorize a response, or prove recovery.

## Required Coverage

The records must make these conditions visible:

- request, session, workflow, step, attempt, trace, span, action, outcome, release, and incident identity;
- effective code, prompt, model or dependency, context or index, tool-policy, routing, runtime, evaluator, and telemetry-schema versions;
- end-to-end latency and time to first output, with queue and dependency attribution;
- input, output, cached, and other billable units plus cost per request, workflow, action, tenant, route, and outcome;
- retrieval request, source coverage, authorization filtering, freshness, citation, empty-context, and partial-context states;
- authorization outcome, refusal, fallback, validation failure, cancellation, retry, tool request, approval, effect, and reconciliation;
- sampled quality with eligible population, denominator, evidence class, inclusion probability, label delay, completeness, and Chapter 8 validity reference;
- explicit evidence-health states for missing, late, sampled, dropped, redacted, unjoinable, or contradictory telemetry; and
- a declared operational behavior when critical evidence becomes indeterminate.

## Runbook Selector

| Condition | Runbook |
| --- | --- |
| Valid quality evidence crosses a service boundary or material cohort degrades | `runbooks/quality-regression.md` |
| End-to-end or dependency latency, timeout, saturation, or availability degrades | `runbooks/latency-or-dependency-degradation.md` |
| Unit use or cost departs from its accepted envelope | `runbooks/cost-spike.md` |
| Retrieval returns missing, partial, unauthorized, low-coverage, or unjoinable context | `runbooks/retrieval-failure.md` |
| Admitted sources or context indexes exceed freshness or change expectations | `runbooks/stale-sources.md` |
| A prompt, policy, route, flag, evaluator, or runtime configuration correlates with regression | `runbooks/prompt-or-configuration-regression.md` |
| A tool request, approval, execution, effect, retry, cancellation, or outcome is anomalous | `runbooks/tool-anomaly.md` |
| Critical events are missing, late, invalid, dropped, or no longer join correctly | `runbooks/telemetry-failure.md` |
| Output may be unsafe or protected data may have been exposed | `runbooks/unsafe-output-or-data-exposure.md` |
| Population, behavior, labels, outcomes, dependencies, or evidence applicability drifts | `runbooks/evaluation-drift.md` |
| An authorized prior composite release must be restored | `runbooks/rollback.md` |
| The service must operate temporarily with narrower capability or supervision | `runbooks/degraded-mode.md` |

## Exit and Acceptance Criteria

Close the monitoring task or incident only when:

- current component versions, traffic assignment, population, and consequence are known or explicitly marked unknown;
- all material claims map to retrievable evidence with query version, snapshot time, denominator, freshness, completeness, and limitations;
- telemetry health is verified independently enough to distinguish no failure from no data;
- facts, hypotheses, generated interpretations, operator decisions, actions, effects, and outcomes remain distinguishable;
- containment accounts for new admission, queues, in-flight work, durable state, caches, tool effects, users, and delayed outcomes;
- every response action has authority, scope, expiry, effect, reversal, and reconciliation evidence;
- technical, behavioral, state, action, user, telemetry, and delayed-outcome recovery claims pass their recorded checks;
- open findings, residual uncertainty, accepted degraded conditions, monitoring windows, and reopening triggers have owners;
- the recovery verifier records a recommendation; and
- the decision authority records return-to-service, degraded-operation, residual-risk, or closure disposition.

## Deeper Record Triggers

Attach core toolkit modules when the assistant or incident:

- handles sensitive payloads, personal data, customer content, secrets, or privileged telemetry;
- crosses tenant, region, purpose, or trust boundaries;
- can execute containment, configuration, traffic, release, ticket, restart, or other external actions;
- affects durable state, queued or in-flight work, or an outcome that may be unknown;
- changes component identity, authority, population, autonomy, accepted risk, or consequence;
- relies on sampled, delayed, proxy, or externally labeled quality evidence;
- needs a new evaluation validity decision, threat decision, permission grant, or release candidate; or
- operates during a severe incident with incomplete evidence.

Typical additions are the AI Fit and Risk Assessment Worksheet, Context Assembly Plan, Tool Permission Matrix, Evaluation Plan, AI Threat Model, AI Release Manifest, AI Operations Runbook, and Principal Design Review Checklist and Record.

## Material Change and Reopening Triggers

Reopen affected records when service scope, population, component identity, signal semantics, denominator, threshold, evidence class, sampling, label source, query, event schema, trace join, access policy, redaction, retention, alert route, action authority, degraded mode, rollback target, recovery objective, or consequence changes. Also reopen when a missed incident, false page, telemetry gap, unrecorded effect, contradictory evidence, delayed outcome, or post-incident review invalidates a prior conclusion.

Completed walkthrough

Context-Index Quality Regression

Fills the six operating records for a citation regression after an index release, from evidence-gated detection through conditional recovery and incident closure.

Preview
# Completed Walkthrough: Quality Regression After a Context-Index Release

Toolkit schema: `PAISEH-TK-1.0`

Scenario pack: `PAISEH-SC-MIA-1.0`

Example status: synthetic, completed, and vendor-neutral

This completed scenario shows how operators investigated and recovered from a citation-quality regression after a context-index release. It fills the six pack records at a reviewable level while leaving evaluation validity, release authority, and incident command with their source owners.

## Common Artifact Header

| Field | Completed record |
| --- | --- |
| Artifact ID | `MIA-EX-021`, completed scenario, `1.0` |
| Service and protected outcome | Internal policy assistant; answers must cite an authoritative, current policy source |
| Incident and release | `INC-247`; context index `ctx-2026.07.22.2` replacing `ctx-2026.07.15.4` |
| Population | Internal policy questions in the general route; restricted collections excluded |
| Owners and authority | Service owner: Knowledge Platform; incident commander: Mei; telemetry owner: Arun; recovery verifier: Sal; closure authority: Mei |
| Evidence | `EV-21` through `EV-35`, retained with the incident |
| Known limitation | Sampled answer-quality labels arrive up to 20 minutes late |
| State | Incident closed after conditional recovery and full observation window |
| Reopening triggers | Recurring citation defect, context identity mismatch, telemetry degradation, or recovery-claim failure |

## 1. Completed AI Observability Contract

Every request recorded application, prompt, resolved dependency, context-index, route, runtime configuration, evaluator, and telemetry-query identities. Correlation joined request, retrieval, selected sources, citations, answer validation, sampled review, incident, and recovery records.

| Evidence class | Required operational meaning | Missing behavior |
| --- | --- | --- |
| Retrieval coverage | authorized sources attempted, admitted, selected, empty, partial, and denied states | mark answer evidence `indeterminate`; route telemetry check |
| Source freshness | source version, modified time, index time, and expected refresh | block a source past the accepted lag |
| Citation validity | material answer claims with resolvable admitted-source references | fail answer validation; do not present as supported |
| Effective identity | exact context-index and route used by each request | exclude request from release comparison |
| Sampled quality | eligible population, denominator, evaluator `eval-9`, rubric `rubric-4`, label delay | investigation only until minimum valid sample |
| Evidence health | event admission, identity joins, evaluator lag, and source freshness | show `unknown`, never healthy or zero |

The contract permitted no raw policy text in common events. Protected source and answer references required incident-purpose access. Accepted degraded mode routed questions to the prior index; prohibited mode served answers with an unknown context identity or broken material citations.

## 2. Completed Signal and Objective Catalog

| ID | Definition | Boundary and use |
| --- | --- | --- |
| `OBJ-21` | Eligible answers have a resolvable authoritative citation for every material policy claim | 99.5% over 30 minutes; page only with healthy identity and citation telemetry |
| `SIG-21` | citation-validation failure rate | 10-minute window, at least 200 eligible answers; compare by effective context index |
| `SIG-22` | authoritative-source retrieval coverage | investigation signal; distinguish empty, partial, denied, and failed |
| `SIG-23` | sampled unsupported-claim rate | Chapter 8 validity reference `EP-14`; at least 80 labels; supports quality assessment, not instant recovery |
| `SIG-24` | context identity join completeness | telemetry-health signal; must exceed 99.9% before `SIG-21` can page |
| `SIG-25` | source-to-index freshness lag | block source admission over 60 minutes unless a documented exception applies |

All signals exposed population, denominator, window, release overlay, freshness, and query version. A signal became `indeterminate` when `SIG-24` or its required event stream was unhealthy.

## 3. Completed Trace and Event Schema

Schema `knowledge-events/3.1` required stable request, trace, attempt, release, route, and context identities. Retrieval events recorded authorized source-set identity, source snapshot, index identity, selected-source references, freshness state, result completeness, and terminal outcome. Citation-validation events recorded claim count, valid-reference count, failure class, and protected answer reference—not raw content.

Replay cases passed for a successful cited answer, an empty retrieval, a stale source, a broken citation, a missing context identity, and a delayed revised quality label (`EV-21`). Independent request receipts reconciled 99.97% of eligible events before the incident; the unmatched fraction stayed excluded and visible.

## 4. Completed Alert Specification

Alert `ALT-21 v2` asked: “Did resolvable citation coverage fall for a specific effective context index, and is the evidence healthy enough to act?”

Detection required:

- at least 200 eligible answers in 10 minutes;
- `SIG-21` above 1.0% for two windows;
- `SIG-24` at or above 99.9%;
- fresh citation-validation events and denominators; and
- no approved maintenance suppression.

The alert could open an incident record and attach evidence. It could not declare severity, roll back an index, change routing, or close the incident. Missing component identity opened a telemetry investigation instead of a quality page.

Replay and shadow tests covered known breach, healthy baseline, low volume, missing denominator, telemetry outage, and planned index release. On-call accepted the rule in `ALERT-DEC-21`.

## 5. Completed Incident Evidence Brief

Incident `INC-247` was declared at 09:18 UTC after citation-validation failures reached 8.7% on `ctx-2026.07.22.2`; the prior index remained at 0.3%.

| Time | Type | Fact, decision, or effect | Evidence |
| --- | --- | --- | --- |
| 09:00 | change | new index received 10% of the general route | `EV-22` release and route receipts |
| 09:12 | signal | failures exceeded the alert boundary with healthy telemetry | `EV-23`, `EV-24` |
| 09:18 | declaration | incident opened; severity set by incident commander | `EV-25` |
| 09:24 | finding | 71 of 78 failures referenced a retired policy-document alias | `EV-26` |
| 09:31 | decision | stop new admission to the candidate index; preserve traces and queued work | `DA-21` |
| 09:35 | effect | route returned to prior index; candidate isolated | `EV-27` |
| 10:02 | verification | new-request citation failures returned below boundary | `EV-28` |

Confirmed impact was limited to 412 general-route answers served by the candidate. Restricted collections, tool actions, and durable external state were unaffected. Answer correction for affected requests was assigned to the knowledge owner.

Hypothesis `H-21`—generated alias mappings omitted a retired-to-current document redirect—was supported by index build evidence and rejected against application, prompt, dependency, and evaluator changes. The incident brief did not treat temporal proximity alone as proof.

## 6. Completed Recovery Verification Record

Recovery candidate `REC-247-1` routed all new eligible questions to `ctx-2026.07.15.4`, held the failed index, invalidated affected cached answers, and queued correction notices.

| Claim | Required evidence | Result |
| --- | --- | --- |
| `RC-21` intended prior index is effective for new work | route and per-request context receipts | Pass |
| `RC-22` citation-validation failure is below 0.5% for 60 minutes | healthy `SIG-21` and `SIG-24` | Pass |
| `RC-23` sampled answer quality meets the accepted boundary | 96 valid labels under `EP-14` | Pass after label window |
| `RC-24` affected answers and caches are bounded | request inventory and invalidation receipts | Pass |
| `RC-25` telemetry and alert paths are healthy | independent request reconciliation and alert replay | Pass |

At 10:15 UTC, the verifier recommended `remain degraded` because sampled labels had not reached minimum volume. At 11:20 UTC, after the label window completed, Sal recommended `return to normal` on the prior index. Mei authorized the mode change and later closed the incident under a separate receipt. The failed index remained quarantined pending a new release decision.

## Why a Visual Was Not Added

The engineering value here is exact record content: signal meanings, evidence gates, authority, timeline, and recovery claims. The timeline and claim tables make those relationships directly inspectable; a diagram would repeat them without clarifying a harder boundary.

01 · Contract

AI Observability Contract

Defines composite identity, capture semantics, latency and cost attribution, retrieval and action evidence, evidence health, access, views, and exercises.

Preview
# AI Observability Contract

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../common-artifact-header.md` first. This contract defines what operators can observe, how evidence is protected, and what the service does when evidence is unhealthy. It does not define evaluation validity, grant telemetry access, or authorize incident action.

## 1. Service and Operating Boundary

| Field | Record |
| --- | --- |
| Service, workflow, and owning team | |
| User-visible behavior and protected outcome | |
| Environments, regions, tenants, cohorts, and routes | |
| Consequence and escalation class | |
| Service owner, telemetry owner, on-call route, and incident system | |
| Operations policy and AI Operations Runbook version | |
| Accepted degraded modes and prohibited modes | |
| Contract version, effective date, review date, and supersedes | |

## 2. Composite Identity Contract

Record the identity that was effective for each request, workflow step, action, and outcome. Use resolved values and receipts, not intended aliases alone.

| Component | Required identity | Resolution or receipt source | Change overlay | Missing-identity behavior |
| --- | --- | --- | --- | --- |
| Application and harness code | build / release / commit | | | |
| Prompt and instruction bundle | immutable version | | | |
| Model or external dependency | requested and resolved identity | | | |
| Context source or index | source set, snapshot, index, or policy version | | | |
| Tool and permission policy | tool schema and grant version | | | |
| Routing and fallback | policy and effective route | | | |
| Runtime flags and configuration | versioned snapshot | | | |
| Evaluator and label pipeline | evaluator, rubric, dataset, or label version | | | |
| Telemetry schema and query | event schema and derived-query version | | | |

Minimum correlation identifiers:

- request, session or conversation, workflow, step, and attempt;
- trace and span;
- tool action, approval, effect, reconciliation, and outcome;
- tenant or authorized subject pseudonym where allowed;
- release, route assignment, incident, and recovery verification.

## 3. Capture Contract

| Evidence class | Required fields and units | Capture point | Sampling / inclusion | Freshness / delay | Redaction or tokenization | Retention | Owner |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Request and workflow identity | | | | | | | |
| Effective component versions | | | | | | | |
| Latency and time to first output | | | | | | | |
| Units and cost attribution | | | | | | | |
| Retrieval coverage and freshness | | | | | | | |
| Authorization, refusal, fallback, and validation | | | | | | | |
| Tool request, approval, effect, and outcome | | | | | | | |
| Sampled quality and delayed outcome | | | | | | | |
| Safety or data-handling event | | | | | | | |
| Telemetry-pipeline health | | | | | | | |

Default to allowlisted structured metadata and protected references. Raw prompts, context, outputs, tool payloads, and incident attachments require a named purpose, access rule, redaction treatment, retention, deletion path, and audit evidence.

## 4. Latency and Cost Semantics

| Measure | Start | Stop | Unit | Attribution dimensions | Partial / cancelled handling | Missing behavior |
| --- | --- | --- | --- | --- | --- | --- |
| End-to-end latency | | | | | | |
| Time to first output | | | | | | |
| Queue delay | | | | | | |
| Dependency duration | | | | | | |
| Retrieval duration | | | | | | |
| Tool duration | | | | | | |
| Input and output units | | | | | | |
| Cost per completed workflow | | | | | | |
| Cost per accepted outcome | | | | | | |

Record price-schedule version, currency, cached or discounted unit treatment, shared-cost allocation, credits, retries, failures, and delayed adjustments. A token count or invoice total without a stable denominator is not an operational cost signal.

## 5. Retrieval and Action Semantics

| Condition | Required event and fields | Expected terminal state | Evidence-health dependency |
| --- | --- | --- | --- |
| Retrieval requested | query/context request identity, authorized scope, source set | result / empty / partial / denied / failed / timed out | |
| Source admitted | source identity, snapshot, freshness, authorization decision | admitted / excluded | |
| Context assembled | selected items, coverage, truncation, citations, limitations | complete / partial / empty | |
| Tool requested | action identity, normalized input, permission decision | denied / approval required / admitted | |
| Approval evaluated | approval identity, scope, approver, expiry, bound action | valid / invalid / expired | |
| Effect attempted | attempt, idempotency key, dependency receipt | succeeded / failed / cancelled / unknown | |
| Effect reconciled | authoritative observed state and receipt | confirmed / compensated / unresolved | |

## 6. Behavioral and Safety Evidence

For every sampled quality, refusal, validation, unsafe-output, feedback, behavior, or delayed-outcome signal, reference the corresponding `signal-objective-catalog.md` entry and its Chapter 8 validity decision. Capture does not make a proxy or sampled judgment conclusive.

Record:

- eligible population, denominator, slices, and exclusions;
- evidence class: invariant, proxy, judgment, feedback, behavior, or delayed outcome;
- sampling frame, inclusion probability, minimum volume, and label delay;
- evaluator or label identity, completeness, revision behavior, and blind spots;
- allowed operational use and prohibited interpretation; and
- behavior when the signal is missing, late, contradictory, or inapplicable.

## 7. Evidence-Health Contract

| Pipeline or join | Health objective | Independent evidence | States | Staleness / completeness threshold | Operational behavior | Owner |
| --- | --- | --- | --- | --- | --- | --- |
| Event admission | | | healthy / degraded / failed / unknown | | | |
| Redaction and policy enforcement | | | | | | |
| Collection and transport | | | | | | |
| Trace and outcome joins | | | | | | |
| Derived queries and dashboards | | | | | | |
| Quality evaluator and labels | | | | | | |
| Cost and price data | | | | | | |
| Retention and deletion jobs | | | | | | |

`No data` must render as `unknown` or `indeterminate`, never as zero or healthy. Name the independent infrastructure, action, business, or audit receipts available when the primary telemetry path fails.

## 8. Access, Protection, and Lifecycle

| Evidence class | Allowed roles and purpose | Row / tenant boundary | Payload access | Export rule | Retention / deletion | Access audit |
| --- | --- | --- | --- | --- | --- | --- |
| Operational metadata | | | | | | |
| Protected payload reference | | | | | | |
| Raw content under approved elevation | | | | | | |
| Tool effect and approval receipt | | | | | | |
| Incident artifact and communication | | | | | | |

Incident urgency does not erase access, redaction, legal hold, retention, or deletion boundaries. Record any emergency elevation with scope, authority, start, expiry, access log, and revocation receipt.

## 9. Views and Decision Support

| Audience | Required question | Required fields | Freshness | Query/version display | Known blind spots |
| --- | --- | --- | --- | --- | --- |
| Service owner | Is the protected outcome healthy for the eligible population? | | | | |
| On-call operator | What changed, what is affected, and what action is available now? | | | | |
| Incident commander | What facts, hypotheses, decisions, actions, gaps, and recovery claims are current? | | | | |
| Cost owner | Which route, component, tenant, retry, or outcome explains economic change? | | | | |
| Evaluator owner | Is quality evidence valid, complete, and applicable? | | | | |

Views must expose denominator, window, freshness, sampling, completeness, query version, effective component identity, change overlays, and evidence-health state.

## 10. Validation and Exercise

| Scenario | Expected evidence | Expected missing-data state | Expected route or degraded behavior | Exercise result / date |
| --- | --- | --- | --- | --- |
| Quality regression with healthy infrastructure | | | | |
| Dependency latency and timeout | | | | |
| Cost amplification by retries | | | | |
| Retrieval or source freshness failure | | | | |
| Tool effect with unknown outcome | | | | |
| Telemetry pipeline outage | | | | |
| Unsafe output or protected-data exposure | | | | |
| Rollback with queued and in-flight work | | | | |

## 11. Acceptance and Reopening

- Service owner:
- Telemetry owner:
- Security or privacy reviewer when required:
- Operations reviewer:
- Accepted limitations:
- Effective date and expiry:
- Evidence and decision receipt:

Reopen when the service boundary, component identity, event schema, capture point, sampling, population, denominator, access, redaction, retention, join, query, objective, alert, action authority, degraded behavior, or recovery claim changes, or when an exercise or incident reveals missing or misleading evidence.

02 · Define

Signal and Objective Catalog

Makes technical, behavioral, economic, retrieval, action, safety, drift, and telemetry-health signals interpretable and owned.

Preview
# Signal and Objective Catalog

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../common-artifact-header.md` first. This catalog defines operational meaning and response use. It does not create an evaluation method, prove causality, or authorize action.

## 1. Catalog Basis

| Field | Record |
| --- | --- |
| Service and Observability Contract version | |
| Eligible populations, environments, tenants, routes, and consequences | |
| Signal owner, service owner, telemetry owner, and on-call route | |
| Evaluation Plan and validity references | |
| Release, configuration, and source-change overlays | |
| Review date, expiry, and supersedes | |

## 2. Objective Register

| Objective ID | Protected user or system outcome | Eligible population / denominator | Window | Target / boundary | Consequence of breach | Evidence-health requirement | Owner | Response class |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| OBJ- | | | | | | | | page / investigate / review / report |

An objective is not satisfied when its denominator, freshness, completeness, or applicability is below the stated requirement.

## 3. Signal Register

| Signal ID | Objective ID | Signal class | Subject and measure | Population / denominator / slices | Window and minimum volume | Baseline / threshold / expected direction | Evidence delay and freshness | Missing / late / revision behavior | Owner and allowed use |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| SIG- | | execution / dependency / behavioral / economic / retrieval / authorization / action / safety / drift / telemetry health | | | | | | | |

## 4. Technical and Dependency Signals

For each applicable signal, define:

| Measure | Required semantics |
| --- | --- |
| Availability and terminal outcome | Eligible attempt, success, failure, cancellation, refusal, fallback, unknown outcome, denominator, and retry treatment |
| End-to-end latency | Start, stop, percentiles, queue attribution, partial output, cancellation, and timeout |
| Time to first output | User-observable start, first meaningful output, streaming interruption, and absent-output state |
| Dependency health | Dependency identity, call class, latency, error, throttling, saturation, circuit state, and effect on the workflow |
| Validation failure | Validator identity, stage, rule, severity, blocked effect, and disposition |

## 5. Retrieval and Freshness Signals

| Measure | Required semantics |
| --- | --- |
| Retrieval coverage | Authorized source set, sources attempted, sources responding, selected items, empty and partial states |
| Context freshness | Source snapshot, indexed-at time, source modified time, expected refresh, lag, and stale-source decision |
| Authorization filtering | Candidate, allowed, excluded, and denied items without exposing protected identities |
| Citation or evidence linkage | Material claims with valid source references, missing links, broken links, and unsupported claims |
| Context truncation | Items and tokens admitted, excluded reason, budget, priority, and effect on answer status |

## 6. Authorization, Refusal, Fallback, and Tool Signals

| Measure | Required semantics |
| --- | --- |
| Authorization | Subject or pseudonym, resource class, policy version, decision, denial reason class, and enforcement point |
| Refusal | Requested behavior, refusal class, policy or capability reason, user guidance, and false-refusal review path |
| Fallback | Trigger, prior route, fallback route, capability change, quality or cost implication, and terminal outcome |
| Tool effect | Action class, approval state, attempt, idempotency, effect receipt, cancellation, compensation, reconciliation, and unknown state |

## 7. Behavioral and Quality Signals

| Field | Required record |
| --- | --- |
| Product claim or failure condition | |
| Evidence class | invariant / proxy / sampled judgment / feedback / behavior / delayed outcome |
| Eligible population, denominator, exclusions, and slices | |
| Sampling frame, inclusion probability, minimum volume, and weighting | |
| Evaluator, rubric, label, or outcome identity | |
| Label delay, revision behavior, completeness, and join key | |
| Chapter 8 validity reference and applicable release / population | |
| Baseline, expected direction, uncertainty, and alert use | |
| Blind spots and prohibited interpretations | |

An immediate invariant violation may justify urgent action. A proxy can support investigation or reversible containment under policy. Neither sampled judgment nor delayed feedback should silently become proof of recovery.

## 8. Economic Signals

| Field | Required record |
| --- | --- |
| Units | input / output / cached / tool / retrieval / compute / other |
| Stable denominator | request / completed workflow / accepted outcome / action / tenant / cohort |
| Attribution | component, route, release, retry, fallback, tenant, action, and outcome |
| Price schedule | identity, currency, effective period, discount, credit, and delayed adjustment |
| Budget or objective | boundary, window, owner, alert use, and accepted burst |
| Missing and revision behavior | |

## 9. Drift Signals

| Field | Required record |
| --- | --- |
| Subject and change class | input population / context corpus / dependency / model behavior / route mix / workflow / outcome / policy / label / telemetry |
| Reference population and window | |
| Current population and window | |
| Representation, unit, slices, cadence, seasonality, and expected change | |
| Minimum volume, uncertainty, repeated-alert handling, and freshness | |
| Reference refresh authority and change history | |
| Classification path | expected change / data defect / instrumentation defect / applicability warning / quality-regression candidate / incident |
| Evaluation Plan handoff | |

Chapter 11 detects and routes a change. Chapter 8 determines whether evidence remains valid or a new evaluation is required.

## 10. Evidence-Health Dependencies

| Signal ID | Required event streams, joins, labels, and queries | Freshness / completeness threshold | Independent check | State when unhealthy | Operator behavior |
| --- | --- | --- | --- | --- | --- |
| SIG- | | | | indeterminate / degraded / unavailable | |

## 11. Change, Review, and Retirement

| Signal ID | Change | Reason | Backtest / replay evidence | Alert impact | Owner approval | Effective date | Retirement / migration |
| --- | --- | --- | --- | --- | --- | --- | --- |
| SIG- | | | | | | | |

Reopen a signal when its objective, population, denominator, evidence class, source, evaluator, label, query, join, threshold, baseline, window, expected change, missing-data behavior, operational use, or owner changes. Retire rules that are unactionable, stale, duplicative, or unsupported; preserve their history.

03 · Correlate

Trace and Event Schema

Versions correlation identifiers, effective components, lifecycle events, action effects, terminal states, units, payload boundaries, and schema health.

Preview
# Trace and Event Schema

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../common-artifact-header.md` first. This record defines the event interface needed to reconstruct work, component identity, decisions, actions, effects, outcomes, and evidence health. It does not authorize payload capture or tool execution.

## 1. Schema Identity and Ownership

| Field | Record |
| --- | --- |
| Schema name, semantic version, and effective date | |
| Producer owners and consumer owners | |
| Source Observability Contract | |
| Registry, compatibility policy, and migration path | |
| Event-time source, clock synchronization, and maximum skew | |
| Validation, quarantine, replay, and dead-letter behavior | |
| Retention, deletion, redaction, and audit policies | |

## 2. Common Event Envelope

| Field | Type / format | Required | Semantics and source | Sensitivity | Validation / missing behavior |
| --- | --- | --- | --- | --- | --- |
| `event_id` | stable unique identifier | yes | one admitted event | | |
| `event_type` | versioned enum | yes | lifecycle or evidence event | | |
| `event_schema_version` | semantic version | yes | decoder contract | | |
| `occurred_at` | timestamp | yes | producer event time | | |
| `observed_at` | timestamp | yes | collector admission time | | |
| `producer_id` | stable service identity | yes | emitting component | | |
| `environment` | bounded enum | yes | runtime boundary | | |
| `tenant_or_scope_ref` | authorized pseudonym / scope | conditional | row or tenant boundary | | |
| `request_id` | identifier | conditional | user-visible unit | | |
| `session_id` | identifier | conditional | bounded interaction | | |
| `workflow_id` | identifier | conditional | long-running unit | | |
| `step_id` | identifier | conditional | workflow state | | |
| `attempt_id` | identifier | conditional | one execution attempt | | |
| `trace_id` / `span_id` / `parent_span_id` | identifier | conditional | causal path | | |
| `release_id` | immutable identity | yes | composite release entry point | | |
| `incident_id` | identifier | conditional | incident association | | |
| `data_class` | bounded enum | yes | handling contract | | |
| `payload_ref` | protected reference | conditional | no raw payload by default | | |

## 3. Effective Component Identity

Every start, decision, effect, and outcome event records the effective identities applicable at that point.

| Field | Meaning | Required on | Resolution evidence |
| --- | --- | --- | --- |
| `application_version` | executable code or build | request / workflow / action | |
| `harness_version` | routing, validation, and control logic | decision / transition | |
| `prompt_version` | immutable instruction bundle | model call | |
| `requested_dependency` | configured model or service selector | dependency request | |
| `resolved_dependency` | actual serving model or dependency | dependency result | |
| `context_version` | source set, snapshot, index, or admission policy | retrieval / model call | |
| `tool_policy_version` | tool schema and permission decision logic | action request | |
| `route_version` / `effective_route` | routing policy and observed assignment | route decision | |
| `runtime_config_version` | flags and configuration snapshot | request / change | |
| `evaluator_version` | evaluator, rubric, or label pipeline | quality event | |
| `telemetry_query_version` | derived measure or view | derived signal | |

## 4. Event Type Register

| Event type | Producer | Required correlation | Required body fields | Terminal or paired event | Idempotency / ordering |
| --- | --- | --- | --- | --- | --- |
| request admitted / rejected | | request, release | authorization, route, reason class | request terminal | |
| workflow / step / attempt started | | workflow, step, attempt, trace | input class, deadline, state | corresponding terminal | |
| dependency requested / completed | | attempt, trace, span | dependency identity, units, latency, terminal state | paired | |
| first output emitted | | request, attempt | output class and timestamp | milestone | |
| retrieval requested / completed | | attempt, trace, context | authorized source set, coverage, freshness, selected count, state | paired | |
| context admitted / excluded | | retrieval, source | source identity, policy result, freshness, reason class | per source | |
| validation passed / failed | | step, attempt | validator version, rule, severity, blocked transition | decision | |
| refusal / fallback | | request, step | reason, prior route, new route, capability change | decision | |
| tool requested / approval evaluated | | action, workflow | normalized input hash, permission, approval binding and expiry | paired | |
| tool effect attempted / observed | | action, attempt | idempotency, dependency receipt, effect class, state | paired | |
| action reconciled / compensated | | action | authoritative observed state and receipt | terminal | |
| quality sample / label / outcome | | request or workflow | eligibility, inclusion probability, evaluator, label, delay | may revise | |
| configuration / release / source changed | | release or source | before, after, authority, scope | change overlay | |
| incident declared / changed / closed | | incident | severity, scope, authority, decision receipt | lifecycle | |
| telemetry health changed | | pipeline | freshness, completeness, lag, drop, join state | state change | |

## 5. Result, Error, and Unknown-State Semantics

| Field | Allowed values | Rule |
| --- | --- | --- |
| `terminal_state` | succeeded / failed / refused / cancelled / timed_out / partial / unknown | Never map missing events to success |
| `error_class` | bounded, non-secret taxonomy | Preserve retry and consequence semantics without raw secret-bearing text |
| `retry_disposition` | prohibited / safe_same_key / safe_new_attempt / reconcile_first / manual_only | Determined outside the model |
| `effect_state` | none / proposed / attempted / confirmed / compensated / unknown | Required for external actions |
| `result_completeness` | complete / partial / sampled / truncated / indeterminate | Required where absence could be misleading |

## 6. Latency, Units, and Cost Fields

| Field | Unit | Attribution and calculation rule |
| --- | --- | --- |
| queue duration | milliseconds | |
| dependency duration | milliseconds | |
| time to first output | milliseconds | |
| end-to-end duration | milliseconds | |
| input / output / cached units | provider-neutral unit plus type | |
| retrieval and tool units | named unit | |
| estimated / settled cost | currency and price-schedule version | |
| cost allocation | request, workflow, action, route, tenant, and outcome | |

## 7. Privacy, Redaction, and Payload References

| Data class | Allowed fields | Prohibited fields | Tokenization / redaction | Payload-store access | Retention / deletion |
| --- | --- | --- | --- | --- | --- |
| Public metadata | | | | | |
| Internal operational metadata | | | | | |
| Customer or personal data | | | | | |
| Secret or credential material | | | | | |
| Incident-elevated payload | | | | | |

Do not place raw prompt, context, output, tool payload, database result, credential, or personal identifier in the common envelope. Store an authorized protected reference only when the incident purpose requires it.

## 8. Schema Health and Compatibility

| Check | Threshold | Evidence | Failure state and route |
| --- | --- | --- | --- |
| Required-field validity | | | |
| Unknown event or enum rate | | | |
| Producer schema lag | | | |
| Duplicate and out-of-order rate | | | |
| Orphan request, workflow, action, or outcome rate | | | |
| Event-time skew and late arrival | | | |
| Redaction and policy enforcement | | | |
| End-to-end completeness against independent receipts | | | |

Specify backward, forward, and mixed-version compatibility. Reject, quarantine, or explicitly downgrade incompatible events; never silently coerce them into a healthy signal.

## 9. Test and Replay Cases

| Case | Expected events and order | Expected missing or duplicate behavior | Join assertion | Result |
| --- | --- | --- | --- | --- |
| Successful request with retrieval | | | | |
| Refusal before model call | | | | |
| Fallback after dependency failure | | | | |
| Tool effect succeeds after retry | | | | |
| Tool outcome remains unknown | | | | |
| Workflow cancelled with in-flight step | | | | |
| Telemetry producer fails mid-request | | | | |
| Delayed quality label revises result | | | | |

## 10. Acceptance and Reopening

- Producer-owner acceptance:
- Telemetry-owner acceptance:
- Privacy or security acceptance when required:
- Consumer migration evidence:
- Known compatibility gaps:
- Effective date and decision receipt:

Reopen when an event type, required identity, correlation key, terminal state, unit, cost rule, payload policy, retention rule, producer, consumer, compatibility promise, or independent completeness check changes.

04 · Route

Alert Specification

Connects one protected objective to an actionable evidence gate, severity, route, grouping, suppression, bounded automation, operator payload, and test history.

Preview
# Alert Specification

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../common-artifact-header.md` first. One copy defines one alert or investigation rule. The rule detects and routes evidence; it does not prove cause, authorize unrestricted automation, or close an incident.

## 1. Rule Identity

| Field | Record |
| --- | --- |
| Alert ID, name, version, and state | draft / shadow / active / muted / retired |
| Protected objective and Signal ID(s) | |
| Service, population, environment, tenant, cohort, and route | |
| Signal owner, service owner, telemetry owner, and on-call route | |
| Runbook and incident-policy references | |
| Query, evaluator, schema, and dashboard versions | |
| Effective date, expiry, review date, and supersedes | |

## 2. Decision the Rule Supports

- Operator question:
- Expected immediate decision:
- Allowed response class: `page`, `investigation queue`, `review queue`, or `report only`
- Why the condition is actionable at this urgency:
- Conditions that make it non-actionable:
- Actions the rule never authorizes:

A rule without a credible owner and decision should not page.

## 3. Detection Semantics

| Field | Record |
| --- | --- |
| Eligible population and exclusions | |
| Numerator, denominator, unit, and slices | |
| Window, alignment, lateness, and minimum volume | |
| Baseline, threshold, comparison, and required duration | |
| Evidence class and Chapter 8 validity reference | |
| Change overlays and suspected dependencies | |
| Query or computation | |
| Expected seasonality and planned-change behavior | |

## 4. Evidence Gate and Missing States

| Dependency | Freshness / completeness requirement | Independent check | State when unmet | Routing behavior |
| --- | --- | --- | --- | --- |
| Primary events | | | indeterminate | |
| Denominator | | | | |
| Component identity | | | | |
| Join or outcome data | | | | |
| Evaluator or label pipeline | | | | |
| Cost or price schedule | | | | |

Specify:

- whether missing evidence suppresses the rule, opens a telemetry incident, or activates a degraded mode;
- which invariant violations may bypass minimum volume;
- how late or revised data changes the alert state; and
- how operators see `no failure observed` versus `unable to observe`.

## 5. Severity and Route

| Condition | Severity | Route | Response target | Escalation | Communications trigger |
| --- | --- | --- | --- | --- | --- |
| | | | | | |

Severity considers consequence, affected population, autonomy, external effects, sensitivity, duration, reversibility, evidence confidence, and current containment—not metric magnitude alone.

## 6. Grouping, Deduplication, and Suppression

| Field | Record |
| --- | --- |
| Grouping and fingerprint keys | |
| Related symptom and cause candidates | |
| Deduplication window and ownership | |
| Repeat-notification policy | |
| Maintenance or planned-change suppression | |
| Suppression authority, scope, expiry, and audit | |
| Maximum silent period and independent check | |
| Alert-storm behavior | |

Suppression must expire and remain visible. It cannot convert a breached objective into a healthy state.

## 7. Allowed Automation

| Action | Preconditions | Bound target | Approval | Idempotency / reconciliation | Maximum duration | Reversal | Receipt |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Enrich alert | | | | | | | |
| Open incident record | | | | | | | |
| Apply approved containment | | | | | | | |

Generated summaries, severity proposals, and response recommendations are advisory. External actions use the Tool Permission Matrix and the applicable controlled workflow.

## 8. Operator Payload

The notification must contain:

- objective, signal, condition, population, denominator, window, threshold, and observed value;
- evidence class, freshness, completeness, query version, and evidence-health state;
- effective release, route, prompt, context, tool-policy, evaluator, and telemetry-schema identities;
- first observed, last observed, detection delay, and relevant change overlays;
- affected consequence and current unknowns;
- runbook, dashboard, evidence query, incident record, and service-owner links; and
- the expected decision, allowed actions, and escalation route.

## 9. Test, Shadow, and Review

| Test | Evidence | Expected result | Actual result | Disposition |
| --- | --- | --- | --- | --- |
| Known breach or replay | | | | |
| Healthy baseline | | | | |
| Low volume | | | | |
| Missing denominator | | | | |
| Late or revised labels | | | | |
| Telemetry outage | | | | |
| Planned release or traffic change | | | | |
| Duplicate symptoms and alert storm | | | | |

Record false positives, false negatives, missed incidents, stale-rule findings, operator load, time to credible awareness, time to action, and last exercised date.

## 10. Acceptance and Reopening

- Signal-owner acceptance:
- Service-owner acceptance:
- On-call acceptance:
- Telemetry-owner acceptance:
- Residual limitations:
- Activation evidence and decision receipt:

Reopen when the objective, population, denominator, signal class, validity reference, threshold, evidence gate, query, component identity, severity, route, grouping, suppression, automation, runbook, or owner changes, or when a missed incident or noisy page reveals a semantic defect.

05 · Investigate

Incident Evidence Brief

Separates facts, hypotheses, generated interpretation, decisions, actions, effects, unknowns, scope, component state, and command handoff.

Preview
# Incident Evidence Brief

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../common-artifact-header.md` first. This brief assembles current evidence for incident command. It separates observed facts, hypotheses, assistant-generated interpretation, decisions, actions, and unknowns. It does not declare severity, authorize action, or replace the incident system of record.

## 1. Command Header

| Field | Record |
| --- | --- |
| Incident ID, title, state, severity, and declared time | |
| Affected service, behavior, population, tenants, regions, routes, and consequence | |
| Incident commander, operations lead, domain lead, communications owner, and scribe | |
| Service owner, telemetry owner, recovery verifier, and closure authority | |
| Current command channel and authoritative incident record | |
| Brief version, evidence cutoff, generated-at time, reviewer, and next update | |

## 2. Current Executive Position

- Confirmed impact:
- Potential impact:
- Current containment:
- Service mode: `normal`, `restricted`, `degraded`, `held`, or `stopped`
- Highest-consequence unknown:
- Next decision and decision owner:
- User or stakeholder communication state:

## 3. Evidence Index

| Evidence ID | Observed claim | Authoritative source / query | Snapshot / window | Component identity | Freshness / completeness | Limitation | Reproduction reference | Owner |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| EV- | | | | | | | | |

Evidence may include events, trace snapshots, release and route receipts, approval and effect receipts, configuration snapshots, source freshness records, evaluator results, audit logs, user reports, and independent business outcomes. A dashboard screenshot without query version and denominator is supporting context, not sufficient evidence.

## 4. Authoritative Timeline

| Time and clock source | Type | Observation, decision, action, or effect | Evidence ID / receipt | Actor or authority | Confidence / gap |
| --- | --- | --- | --- | --- | --- |
| | signal / change / report / declaration / decision / action / effect / verification / communication | | | | |

Record clock offsets, late events, corrected entries, and time ranges. Never rewrite history to match the current hypothesis.

## 5. Scope and Impact

| Dimension | Confirmed affected | Confirmed unaffected | Unknown | Evidence ID | Next check |
| --- | --- | --- | --- | --- | --- |
| Requests and workflows | | | | | |
| Cohorts, tenants, regions, and routes | | | | | |
| Queued and in-flight work | | | | | |
| Durable state, caches, and indexes | | | | | |
| Tool actions and external effects | | | | | |
| User-visible outputs and decisions | | | | | |
| Sensitive data or safety exposure | | | | | |
| Cost and resource consumption | | | | | |
| Delayed outcomes | | | | | |

## 6. Effective State and Change Overlay

| Component | Current identity | Prior known-good identity | Change time / receipt | Exposure | Evidence health |
| --- | --- | --- | --- | --- | --- |
| Application and harness | | | | | |
| Prompt and policy | | | | | |
| Model or dependency | | | | | |
| Context source or index | | | | | |
| Tool policy | | | | | |
| Routing and flags | | | | | |
| Evaluator and labels | | | | | |
| Telemetry schema and query | | | | | |

## 7. Hypothesis Register

| Hypothesis ID | Proposed explanation | Supporting evidence | Contradicting evidence | Discriminating test | Risk of test | Owner | State |
| --- | --- | --- | --- | --- | --- | --- | --- |
| H- | | | | | | | proposed / testing / supported / rejected / indeterminate |

Label assistant-generated hypotheses as generated. Do not present correlation with a release, source, prompt, route, or dependency as causation.

## 8. Decisions and Actions

| Decision / Action ID | Decision or action | Authority | Scope and target | Start / expiry | Expected effect | Receipt | Observed effect | Reversal / reconciliation |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| DA- | | | | | | | | |

Unknown outcomes remain open until reconciled. Record rejected options and why they were not chosen when they affected risk.

## 9. Telemetry and Evidence Health

| Evidence dependency | State | Freshness / completeness | Independent check | Effect on conclusions | Degraded behavior | Owner |
| --- | --- | --- | --- | --- | --- | --- |
| Primary event path | healthy / degraded / failed / unknown | | | | | |
| Component identity joins | | | | | | |
| Quality evaluator / labels | | | | | | |
| Retrieval / source freshness | | | | | | |
| Tool effect receipts | | | | | | |
| Cost / price data | | | | | | |

## 10. Findings, Risks, and Open Questions

| Finding ID | Evidence ID(s) | Observation and consequence | Severity | Owner | Required decision or check | Due | State |
| --- | --- | --- | --- | --- | --- | --- | --- |
| F- | | | | | | | |

- Residual uncertainty:
- Access, retention, or disclosure restrictions:
- Decisions blocked by missing evidence:
- Conditions requiring severity or scope reassessment:

## 11. Handoff

- Current command intent:
- Next three bounded actions:
- Actions prohibited or paused:
- Pending approvals:
- Next evidence update:
- Next communications update:
- Handoff accepted by, time, and receipt:

## 12. Reopening and Supersession

Issue a new brief when scope, severity, component identity, evidence cutoff, containment, service mode, hypothesis, telemetry health, user impact, or next decision materially changes. Preserve earlier versions; mark corrected facts and superseded hypotheses rather than deleting them.

06 · Verify

Recovery Verification Record

Tests technical, behavioral, state, action, user, telemetry, and delayed-outcome recovery before an authority changes service or incident state.

Preview
# Recovery Verification Record

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../common-artifact-header.md` first. This record tests explicit recovery claims after containment, rollback, repair, or degraded operation. Green infrastructure alone is not proof of recovery. The verifier recommends a disposition; only the named decision authority may return the service to normal or close the incident.

## 1. Recovery Candidate

| Field | Record |
| --- | --- |
| Incident ID, service, population, and consequence | |
| Candidate service mode | normal / restricted / degraded / held / stopped |
| Recovery change, rollback, repair, or containment identity | |
| Effective release, prompt, dependency, context, tool-policy, route, configuration, evaluator, and telemetry-schema versions | |
| New admission, queue, and in-flight treatment | |
| Recovery verifier and decision authority | |
| Verification start, observation window, delayed-outcome window, and expiry | |

## 2. Recovery Claim Register

| Claim ID | Dimension | Required claim | Population / scope | Success boundary | Evidence-health requirement | Evidence ID | Result |
| --- | --- | --- | --- | --- | --- | --- | --- |
| RC- | technical / behavioral / state / action / user / telemetry / delayed outcome | | | | | | pass / fail / indeterminate |

## 3. Technical Recovery

| Check | Evidence ID | Result | Limitation / next observation |
| --- | --- | --- | --- |
| New requests reach the intended effective components and route | | | |
| Availability, latency, time to first output, saturation, and dependency health meet the accepted boundary | | | |
| Refusal, fallback, validation, and error states are expected | | | |
| Cost per completed workflow and outcome is inside the accepted envelope | | | |
| Retries, cancellations, and timeouts do not amplify load or duplicate work | | | |

## 4. Behavioral and Quality Recovery

| Check | Evidence ID | Result | Limitation / next observation |
| --- | --- | --- | --- |
| Eligible population, denominator, slices, and component identity match the claim | | | |
| Invariants and prohibited-output checks pass | | | |
| Sampled quality evidence meets its minimum volume and validity conditions | | | |
| Retrieval coverage, source freshness, citation, and partial-context behavior are acceptable | | | |
| Proxy, feedback, behavior, and delayed outcomes are labeled and not overclaimed | | | |
| Chapter 8 owner confirms evidence applicability when required | | | |

## 5. State, Queue, and Action Recovery

| Check | Evidence ID | Result | Limitation / next action |
| --- | --- | --- | --- |
| Queued work was drained, cancelled, revalidated, or held under policy | | | |
| In-flight work completed, cancelled, or reconciled | | | |
| Durable state, caches, indexes, and derived artifacts are consistent | | | |
| Every tool action has a terminal effect receipt or an owned unknown outcome | | | |
| Duplicate, partial, or out-of-order effects were prevented or corrected | | | |
| Compensations and reversals reached their verified target state | | | |

## 6. User and Communication Recovery

| Check | Evidence ID | Result | Limitation / next action |
| --- | --- | --- | --- |
| Affected users, decisions, and external effects are identified | | | |
| Unsafe or incorrect outputs were withdrawn, corrected, or bounded | | | |
| Required notices, status updates, remediation, or support paths were delivered | | | |
| Current service capability and degraded limits are accurately communicated | | | |
| Downstream consumers know which receipts or results were superseded | | | |

## 7. Telemetry Recovery

| Check | Evidence ID | Result | Limitation / next observation |
| --- | --- | --- | --- |
| Event admission, transport, joins, queries, and dashboards are healthy | | | |
| Required component identities and correlation fields are complete | | | |
| Evaluator, label, cost, freshness, and deletion pipelines meet their objectives | | | |
| `No data` and indeterminate states render correctly | | | |
| Independent receipts reconcile with primary telemetry | | | |
| Alert rules were re-enabled, exercised, or deliberately retired | | | |

## 8. Delayed Outcomes and Observation

| Outcome or recurrence signal | Earliest credible time | Required window / volume | Owner | Current evidence | Follow-up disposition |
| --- | --- | --- | --- | --- | --- |
| | | | | | |

If delayed evidence is not yet available, record a conditional return-to-service decision, monitoring owner, expiry, and automatic reopening trigger. Do not mark the claim passed.

## 9. Residual Risk and Degraded Conditions

| Condition | Affected scope | Consequence | Compensating control | Owner | Expiry | Reopening trigger | Acceptance authority |
| --- | --- | --- | --- | --- | --- | --- | --- |
| | | | | | | | |

## 10. Verification Outcome

- Claims passed:
- Claims failed:
- Claims indeterminate:
- Open unknown effects:
- Residual uncertainty:
- Recovery verifier recommendation: `return to normal`, `remain degraded`, `continue containment`, `rollback`, or `do not recover`
- Required conditions and observation:
- Decision authority:
- Final disposition, scope, time, expiry, and receipt:
- Incident closure authority and separate closure receipt:

The recovery recommendation does not close the incident unless the verifier is also the recorded closure authority and all closure policy conditions are met.

## 11. Reopening Triggers

Reopen when a recovery claim fails, telemetry becomes indeterminate, the effective component or route differs from the candidate, a queue or action remains unresolved, a delayed outcome contradicts the disposition, the degraded-mode expiry arrives, user impact expands, or the same failure signature recurs.

Behavior runbook

Quality Regression

Validates evidence, bounds affected cohorts, separates correlation from cause, contains consequence, and verifies representative recovery.

Preview
# Runbook: Quality Regression

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when valid production evidence indicates that a protected behavioral outcome, invariant, sampled quality measure, or material cohort has regressed.

## Entry Conditions and Authority

- Objective, Signal ID, population, denominator, window, threshold, and evidence class:
- Evidence-health state, evaluator or label version, minimum volume, and Chapter 8 validity reference:
- Effective release, route, prompt, dependency, context, tool-policy, and configuration identities:
- Signal owner, service owner, incident commander, evaluation owner, and response operator:
- Authority to constrain admission, route traffic, hold actions, enter degraded mode, or invoke rollback:

An assistant summary may open an investigation. It cannot establish evaluation validity, declare severity, or authorize a production change.

## Immediate Action

1. Validate denominator, population, sampling, label delay, evaluator identity, joins, query version, and telemetry completeness.
2. Mark the signal `indeterminate` and route to telemetry failure if evidence validity cannot be established.
3. Preserve affected examples through approved references, current component identities, source or context versions, and change overlays.
4. Bound affected cohorts, routes, time windows, queued work, in-flight actions, and downstream decisions.
5. When consequence requires it, invoke authorized reversible containment without waiting to prove a root cause.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Reproduce | Re-run the versioned signal query and reconcile independent outcome or invariant evidence | Stable regression with known population and completeness | Results disagree or protected evidence is unavailable |
| Correlate | Compare component, route, source, policy, traffic, evaluator, and telemetry changes on one timeline | Candidate change with exposure overlap | Correlation is being treated as cause |
| Slice | Inspect consequence-relevant cohorts without fishing across ungoverned dimensions | Bounded affected and unaffected populations | Slices are underpowered or expose sensitive attributes |
| Test | Use an approved evaluation or replay path to discriminate hypotheses | Evidence that supports or rejects a hypothesis | A new validity decision or risky live experiment is required |
| Contain | Narrow admission, pause effects, require review, use an approved route, or enter degraded mode | Action receipt and reduced exposure | Effect is unknown or containment increases consequence |
| Correct | Route prompt, context, dependency, policy, or release corrections through their owning change process | Approved candidate and fresh evidence | Correction broadens scope or lacks rollback |

## Verification and Stopping Conditions

Recovery requires healthy evidence, correct component identity, sufficient eligible volume, acceptable invariant and quality results for the affected cohorts, reconciled queued and in-flight work, and no concealed user or external-effect impact. Use delayed outcomes when the objective requires them.

Stop and escalate when unsafe output or exposure is possible, action outcomes are unknown, evaluation applicability is disputed, population reach grows, or no approved containment can hold the consequence.

## Reversal

- Revert temporary routing, holds, sampling, capture, or reviewer requirements only under recorded authority.
- Reconcile workflows admitted under both affected and recovery configurations.
- Remove elevated payload access at expiry and verify retention or deletion.
- Supersede incorrect outputs, decisions, dashboards, or receipts.

## Evidence and Closure

Record signal queries, snapshots, denominators, evaluator and component versions, affected examples, hypotheses, decisions, actions, receipts, recovery results, residual uncertainty, and delayed-outcome obligations.

Close only when the recovery verifier accepts all applicable recovery dimensions and the decision authority records the service mode. Reopen when quality, evidence health, or a delayed outcome breaches the accepted condition.

Performance runbook

Latency or Dependency Degradation

Decomposes end-to-end delay, prevents retry amplification, protects active work, and verifies latency without hiding quality or cost regressions.

Preview
# Runbook: Latency or Dependency Degradation

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when end-to-end latency, time to first output, timeout, saturation, throttling, or a dependency objective breaches its accepted boundary.

## Entry Conditions and Authority

- Objective, Signal ID, affected workflow, population, route, and time window:
- End-to-end, queue, retrieval, dependency, tool, and first-output evidence:
- Current dependency identities, limits, timeout policy, fallback, retry, and circuit state:
- On-call operator, service owner, dependency owner, incident commander, and response operator:
- Authority to shed load, pause admission, cancel work, change route, enable fallback, or enter degraded mode:

## Immediate Action

1. Confirm that event admission, clocks, trace joins, and denominators are healthy.
2. Prevent retry storms and new high-consequence admissions when deadlines or dependency capacity are uncertain.
3. Preserve traces, queue depth, saturation, throttling, timeout, cancellation, fallback, and dependency receipts.
4. Bound affected requests, queued and in-flight workflows, tool effects, and user-visible partial output.
5. Invoke approved load shedding, concurrency limits, fallback, or degraded mode when consequence requires containment.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Decompose | Attribute end-to-end time across queue, retrieval, model or dependency, validation, tool, and streaming stages | Dominant stage with trustworthy trace coverage | Trace completeness is insufficient |
| Compare | Overlay releases, routes, configuration, traffic, input size, dependency resolution, and regional changes | Bounded candidate conditions | Investigation would broaden production access |
| Inspect | Check saturation, rate limits, connection pools, circuit state, deadlines, and cancellation propagation | Mechanism consistent with symptoms | Shared platform health is at risk |
| Contain | Shed or queue eligible work, cap concurrency, disable unsafe retries, or route to an approved fallback | Lower exposure and effect receipt | Fallback changes capability without disclosure |
| Recover | Restore capacity or an approved prior configuration through the owning system | Stable latency and terminal outcomes | In-flight work or unknown effects remain |

## Verification and Stopping Conditions

Verify end-to-end and first-output objectives, queue age, dependency health, cancellation, fallback frequency, retry amplification, cost, quality, partial-output handling, and telemetry completeness for the affected population. Confirm that queued and in-flight workflows reached known terminal states.

Stop and escalate when dependency identity is unknown, cancellation does not propagate, retries create load, partial output may mislead users, fallback crosses an authority boundary, or shared capacity continues to degrade.

## Reversal

- Remove temporary admission, concurrency, timeout, routing, or fallback changes through recorded configuration control.
- Drain or cancel overflow queues and reconcile abandoned attempts.
- Restore normal deadlines only after dependency and load evidence stabilizes.
- Notify users when earlier timeouts, partial output, or delayed actions changed expected behavior.

## Evidence and Closure

Record trace and query versions, timing distributions, queue and capacity state, dependency receipts, route changes, action authority, effect, reversal, and recovery observations.

Close when latency and terminal outcomes meet the accepted boundary, retry and fallback behavior is normal, queued and in-flight work is reconciled, and quality and cost remain acceptable. Reopen on recurrence or delayed effect.

Economic runbook

Cost Spike

Reconciles price and usage evidence, normalizes cost by outcome, bounds runaway work, and prevents savings from weakening required controls.

Preview
# Runbook: Cost Spike

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when unit consumption or cost per request, completed workflow, accepted outcome, action, tenant, cohort, or route departs materially from its accepted envelope.

## Entry Conditions and Authority

- Economic Objective and Signal ID, stable denominator, currency, window, threshold, and accepted burst:
- Price-schedule version, settlement delay, credits, cached-unit treatment, and evidence health:
- Affected components, routes, tenants, retries, fallbacks, tools, and outcomes:
- Cost owner, service owner, on-call operator, incident commander, and budget authority:
- Authority to cap admission, disable retries, narrow capability, change route, or accept temporary spend:

## Immediate Action

1. Verify price data, unit events, denominator completeness, currency, shared-cost allocation, and delayed adjustments.
2. Attribute the change before making a quality- or safety-reducing optimization.
3. Bound uncontrolled retries, loops, oversized context, unexpected output, fallback amplification, tool fan-out, and abuse.
4. Apply approved caps or holds when spend rate threatens the accepted budget or signals runaway work.
5. Preserve component identities, unit receipts, cost calculations, traffic mix, and current user impact.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Reconcile | Compare estimated and settled cost against independent usage or billing receipts | Trusted magnitude and attribution | Price or unit evidence is incomplete |
| Normalize | Recalculate per completed workflow and accepted outcome, including failure and retry cost | Stable denominator and cohort view | Completion or outcome joins are unhealthy |
| Localize | Attribute by release, route, dependency, prompt, context size, tool, retry, tenant, and input class | Dominant contributor | A sensitive-tenant inspection is unauthorized |
| Contain | Cap concurrency or budget, stop loops, pause expensive optional work, or use approved degraded behavior | Reduced burn with receipt | Containment breaches quality or safety invariants |
| Correct | Route configuration, prompt, context, policy, or release changes through the owning process | Reviewed correction and cost forecast | A new budget or product decision is required |

## Verification and Stopping Conditions

Verify settled and estimated cost, unit use, retry and fallback rates, completion and accepted-outcome denominators, latency, quality, refusal, retrieval coverage, and affected-user behavior. A lower bill is not recovery if work is silently dropped or quality evidence is missing.

Stop and escalate when spend remains unbounded, attribution is unreliable, abuse or credential compromise is possible, or the only reduction weakens a required control.

## Reversal

- Remove temporary caps, holds, or route changes only after stable observation.
- Reconcile workflows suppressed, delayed, or partially completed during containment.
- Restore optional capability gradually under cost and quality observation.
- Reverse emergency access used for detailed billing or tenant evidence.

## Evidence and Closure

Record unit and billing receipts, price schedules, denominator queries, component and route identities, hypotheses, caps, actions, user impact, and recovery comparisons.

Close when cost stays within the accepted envelope for a representative window, attribution is understood, no control was weakened, suppressed work is reconciled, and the cost owner accepts residual uncertainty. Reopen on late settlement or recurrence.

Context runbook

Retrieval Failure

Distinguishes empty, partial, denied, stale, truncated, and failed retrieval while preserving permissions and visible missing-context behavior.

Preview
# Runbook: Retrieval Failure

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when retrieval is empty, partial, denied unexpectedly, low coverage, unjoinable, mistargeted, or unable to provide required evidence.

## Entry Conditions and Authority

- Retrieval, quality, or evidence-health Signal ID and affected request population:
- Authorized source set, context or index version, query class, freshness contract, and expected coverage:
- Observed empty, partial, denied, truncated, failed, or timed-out state:
- Context owner, source owner, service owner, on-call operator, and incident commander:
- Authority to disable answers, require source review, use an approved fallback, rebuild an index, or narrow population:

## Immediate Action

1. Ensure missing or partial context is explicit in system output and telemetry.
2. Hold high-consequence answers or actions that require unavailable evidence.
3. Verify authorization filtering separately from source availability; do not bypass permissions to restore coverage.
4. Preserve request identity, authorized candidate set, selected items, exclusions, source responses, index version, and truncation evidence.
5. Route stale-source conditions to the stale-sources runbook and telemetry defects to telemetry failure.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Classify | Distinguish true empty result, authorization exclusion, source outage, stale index, query defect, ranking miss, truncation, and telemetry gap | One or more bounded failure classes | Protected source identities would be exposed |
| Scope | Compare affected sources, tenants, cohorts, query classes, regions, and context versions | Affected and unaffected population | Tenant boundaries cannot be verified |
| Reproduce | Replay through an approved non-effecting path with the same authorized scope | Candidate result and stage evidence | Replay changes permissions or source state |
| Contain | Refuse, disclose partiality, require human source inspection, or route to an approved authoritative source | Reduced risk and visible capability change | Fallback lacks equivalent authorization or provenance |
| Correct | Repair source admission, query, index, ranking, budget, or dependency through its owning change path | Fresh context receipt and coverage evidence | The change needs new context or evaluation design |

## Verification and Stopping Conditions

Verify authorized source coverage, freshness, selection, citations, empty and partial behavior, context budget, quality for affected queries, and telemetry completeness. Confirm that no answer or action created during the failure is treated as fully evidenced.

Stop and escalate when permissions are uncertain, source truth conflicts, sensitive content may cross a boundary, or no safe missing-context behavior exists.

## Reversal

- Remove temporary fallbacks, manual-review gates, or disabled query classes through controlled configuration.
- Invalidate caches or generated answers dependent on faulty context.
- Reprocess only eligible work with the original authorization and a new receipt.
- Notify downstream users of superseded claims when required.

## Evidence and Closure

Record source and index identities, authorization decisions, coverage and freshness measures, selected and excluded counts, reproductions, affected answers, actions, corrections, and recovery evidence.

Close when the affected population receives current authorized context or an explicitly accepted safe fallback, earlier outputs are reconciled, and evidence health is stable. Reopen when coverage, authorization, or freshness degrades again.

Freshness runbook

Stale Sources

Invalidates stale context, traces dependent outputs and actions, rebuilds from authoritative sources, and reconciles affected decisions.

Preview
# Runbook: Stale Sources

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when an admitted source, context index, schema, policy, or cached representation exceeds its freshness contract or no longer matches authoritative state.

## Entry Conditions and Authority

- Source, context, or freshness Signal ID; source owner; snapshot or index identity:
- Expected refresh, observed source-modified and indexed-at times, lag, and affected population:
- Known dependent prompts, workflows, answers, actions, and decisions:
- Source owner, context owner, service owner, incident commander, and recovery verifier:
- Authority to invalidate context, disable affected answers, rebuild, reindex, or accept a bounded stale mode:

## Immediate Action

1. Mark affected context and dependent outputs stale or indeterminate.
2. Prevent new high-consequence use when current source truth is required.
3. Preserve authoritative source identity, index or cache receipt, refresh history, failures, and dependency mapping.
4. Bound the stale interval, admitted content, queries, tenants, workflows, and downstream decisions.
5. Do not refresh from an unverified replica or bypass source authorization.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Confirm | Compare authoritative modification evidence with admitted snapshot, index, cache, or policy | Proven freshness breach or semantic mismatch | Authoritative source cannot be identified |
| Localize | Inspect extraction, change detection, queues, transformation, indexing, admission, and cache invalidation | Failed stage and affected versions | Repair would cross a trust boundary |
| Scope | Enumerate dependent requests, answers, citations, actions, and decisions | Bounded affected population and interval | Dependency lineage is unavailable |
| Contain | Disable affected source, disclose staleness, require review, or use an approved current source | Visible bounded behavior | Remaining sources create conflicting truth |
| Restore | Rebuild from the authoritative source and issue a new immutable context identity | Freshness and completeness receipt | Rebuild is partial or permissions differ |

## Verification and Stopping Conditions

Verify source authorization, snapshot time, completeness, transformation, index or cache identity, query coverage, citation linkage, and representative dependent behavior. Reconcile answers and actions produced during the stale interval.

Stop and escalate when source ownership is disputed, conflicting sources lack a resolution policy, a rebuild drops required content, or stale content may have caused an external effect.

## Reversal

- Remove emergency source exclusions or manual gates only after fresh admission evidence.
- Invalidate temporary indexes, caches, and derived artifacts.
- Restore normal refresh cadence and alerting.
- Supersede or retract affected outputs and decision receipts where required.

## Evidence and Closure

Record source and context identities, timestamps, lag, failure stage, affected lineage, containment, rebuild receipt, validation, and user remediation.

Close when the authoritative source is current and complete in the admitted context, dependent behavior passes verification, and affected outputs are reconciled. Reopen on lag recurrence or a contradictory late source event.

Change runbook

Prompt or Configuration Regression

Preserves effective identities, freezes exposure, discriminates candidate changes, and routes versioned corrections through the owning release path.

Preview
# Runbook: Prompt or Configuration Regression

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when a prompt, instruction, policy, route, flag, threshold, evaluator, tool configuration, or runtime setting correlates with a production regression.

## Entry Conditions and Authority

- Signal, symptom, affected population, and first credible time:
- Effective and prior prompt, policy, route, flag, evaluator, tool, and runtime configuration identities:
- Change receipt, exposure assignment, compatibility evidence, and recovery target:
- Configuration owner, release owner, service owner, incident commander, and response operator:
- Authority to freeze rollout, narrow exposure, restore a prior version, or enter degraded mode:

## Immediate Action

1. Freeze further exposure changes and preserve effective—not merely intended—configuration identities.
2. Validate telemetry, denominator, route assignment, and change-overlay completeness.
3. Bound affected cohorts, queued and in-flight workflows, actions, and durable state.
4. Use an approved hold, prior route, or degraded mode when consequence warrants reversible containment.
5. Do not hot-edit an unversioned prompt or setting as an incident shortcut.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Correlate | Compare regression timing and exposure with every relevant component and policy change | Bounded candidate changes | Correlation is being presented as cause |
| Reproduce | Replay representative cases against exact prior and current composite identities | Discriminating outcome evidence | Replay requires protected data or external effects |
| Isolate | Vary one approved candidate while holding other identities stable | Causal support within the test boundary | Mixed-version or dynamic resolution cannot be controlled |
| Contain | Pause rollout, narrow route, require review, or invoke approved prior configuration | Receipt and reduced exposure | Active work cannot retain safe affinity |
| Correct | Create a versioned candidate with owning review, evidence, staged exposure, and rollback | Approved release evidence | A new evaluation or policy decision is required |

## Verification and Stopping Conditions

Verify actual route assignment and component identities, affected quality and safety signals, refusal and fallback, latency, cost, retrieval, tool behavior, queued and in-flight work, and delayed outcomes. A restored configuration is not recovery until exposure and work state are reconciled.

Stop and escalate when the prior version is incompatible, component resolution is dynamic or unknown, rollback would widen another risk, or a tool effect cannot be reconciled.

## Reversal

- Revert temporary holds, route pins, or reviewer gates through the release or configuration system.
- Restore composite version compatibility and remove emergency overrides at expiry.
- Reconcile workflows spanning old and new identities.
- Supersede affected outputs, actions, or decisions.

## Evidence and Closure

Record change and assignment receipts, effective identities, comparisons, tests, decisions, actions, affected work, recovery evidence, and residual uncertainty.

Close when the authorized composite state is stable, affected signals recover with healthy evidence, active work is reconciled, and the decision authority records service mode. Reopen if delayed evidence or route drift contradicts recovery.

Action runbook

Tool Anomaly

Stops unsafe retries, reconstructs approval and permission, reconciles unknown effects, and verifies terminal ownership.

Preview
# Runbook: Tool Anomaly

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when a tool request, permission decision, approval, execution, retry, cancellation, effect, compensation, reconciliation, or terminal outcome differs from the recorded contract.

## Entry Conditions and Authority

- Workflow, action, attempt, approval, idempotency, effect, and reconciliation identities:
- Expected permission, normalized input, target, effect, terminal state, and deadline:
- Observed denial, duplicate, unexpected target, timeout, partial effect, unknown outcome, or missing receipt:
- Tool owner, workflow owner, incident commander, response operator, and reconciliation authority:
- Authority to stop admission, revoke a grant, cancel, compensate, reconcile, or notify an affected owner:

The assistant may group evidence and propose a runbook step. It cannot infer that a timeout means no effect, approve its own action, or retry an unknown outcome.

## Immediate Action

1. Stop automatic retries and new related actions when effect state is unknown or duplication is possible.
2. Preserve normalized input, approval binding and expiry, permission decision, idempotency key, dependency receipts, callbacks, audit records, and observed state.
3. Revoke or quarantine the affected tool path when scope escape, unsafe input, or broad authorization is possible.
4. Bound affected workflows, targets, tenants, durable state, and downstream effects.
5. Assign explicit ownership to every unknown outcome.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Validate | Compare actual request, target, permission, approval, and expiry with the Tool Permission Matrix | Contract match or specific violation | Approval or audit evidence is missing |
| Reconcile | Query the authoritative target using a non-effecting path | Confirmed effect state | Observation itself requires broader authority |
| Classify | Identify denial, validation defect, duplicate, partial effect, dependency failure, cancellation race, or receipt loss | Bounded anomaly class | Multiple targets or tenants may be affected |
| Contain | Disable action class, narrow grant, cancel eligible work, or require manual review | Enforced containment receipt | Cancellation cannot guarantee terminal state |
| Correct | Compensate, complete, or supersede through separately authorized action | New effect and reconciliation receipts | No safe compensation exists |

## Verification and Stopping Conditions

Verify permission enforcement, approval binding, target, idempotency, retry disposition, cancellation, durable state, downstream effects, audit completeness, and terminal ownership. Test a representative safe action after repair without masking unresolved earlier effects.

Stop and escalate when protected data or safety exposure is possible, authority cannot be reconstructed, a destructive effect may have occurred, or unknown outcomes remain unowned.

## Reversal

- Revoke emergency grants, tool credentials, route changes, or manual overrides.
- Execute compensation only through its own approved contract and receipt.
- Reconcile caches, tickets, messages, jobs, databases, and downstream consumers.
- Notify affected owners when an effect cannot be fully reversed.

## Evidence and Closure

Record requests, approvals, permission decisions, attempts, idempotency, dependency evidence, observed effects, compensations, audit gaps, and verification.

Close when every action is confirmed, compensated, or accepted as an explicitly owned residual risk; the tool contract is enforced; and affected workflows are reconciled. Reopen on a late effect or duplicate receipt.

Evidence runbook

Telemetry Failure

Marks conclusions indeterminate, activates independent receipts, localizes evidence loss, and restores signals without replaying effects.

Preview
# Runbook: Telemetry Failure

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when critical events, identities, joins, queries, labels, cost records, redaction results, or dashboards are missing, late, dropped, invalid, contradictory, or stale.

## Entry Conditions and Authority

- Affected event streams, producers, pipelines, joins, queries, views, signals, and time range:
- Evidence-health objective, independent receipt, observed completeness, freshness, and drop or error rate:
- Product decisions, alerts, controls, or recovery claims that depend on the evidence:
- Telemetry owner, service owner, on-call operator, incident commander, and privacy or security owner when required:
- Authority to mark objectives indeterminate, activate independent signals, restrict service, elevate capture, or repair pipelines:

## Immediate Action

1. Mark dependent signals and recovery claims `indeterminate`; never render missing data as zero or healthy.
2. Activate approved independent infrastructure, action, audit, or business receipts.
3. Apply the recorded fail-open, fail-closed, hold, or manual-supervision behavior for critical controls.
4. Preserve producer health, buffers, offsets, validation failures, schema versions, query versions, and access or redaction evidence.
5. Bound the blind interval and every decision made during it.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Localize | Inspect producer, admission, transport, transform, redaction, storage, join, query, evaluator, and view stages | Failed stage and affected interval | The diagnostic path risks protected payloads |
| Quantify | Reconcile events against independent request, release, tool, billing, or outcome receipts | Completeness and population estimate | No independent denominator exists |
| Validate | Check schema compatibility, clock skew, duplicates, late data, orphan rates, and query changes | Specific mechanism and invalidated signals | Evidence corruption extends beyond known scope |
| Contain | Pause dependent automation, page on telemetry health, or enter approved degraded mode | Reduced decision risk | Required control has no independent path |
| Restore | Repair or replay from authorized durable sources without duplicating effects | Rebuilt events and completeness receipt | Replay changes operational state |

## Verification and Stopping Conditions

Verify producer admission, redaction, transport, schema validity, joins, query output, dashboard freshness, alerts, evaluator and label flow, cost attribution, retention, deletion, and independent reconciliation across the full blind interval.

Stop and escalate when redaction failed, data was exposed, events cannot be reconstructed, critical control evidence remains unavailable, or replay could create external effects.

## Reversal

- Remove emergency capture and access elevation at expiry; verify deletion and audit.
- Disable temporary independent queries after normal evidence is trustworthy.
- Re-enable alerts and automation deliberately, with a tested evidence gate.
- Correct or supersede dashboards, incident claims, and decisions made from incomplete data.

## Evidence and Closure

Record blind interval, affected signals, pipeline and schema identities, independent receipts, repair and replay evidence, privacy treatment, decisions made under uncertainty, and restoration tests.

Close when evidence health meets its contract, dependent signals are recomputed, incorrect conclusions are superseded, and service mode is explicitly accepted. Reopen on late gaps, mismatched counts, or failed deletion.

Safety runbook

Unsafe Output or Data Exposure

Stops propagation, routes authority to the owning incident process, protects sensitive evidence, and verifies revocation and remediation.

Preview
# Runbook: Unsafe Output or Data Exposure

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when an output may violate a safety invariant or protected information may have reached an unauthorized prompt, context, log, trace, tool, result, user, export, cache, or incident artifact.

## Entry Conditions and Authority

- Triggering invariant, report, Signal ID, evidence reference, and confidence:
- Potential data class, unsafe behavior, affected population, recipients, systems, and time range:
- Current prompt, context, tool, model or dependency, policy, route, release, and telemetry identities:
- Incident commander, service owner, security or privacy authority, data owner, communications owner, and response operator:
- Authority to stop service, revoke access, quarantine artifacts, preserve evidence, notify users, or invoke legal and policy processes:

The assistant must not retrieve more protected content to “confirm” exposure unless an authorized investigation path requires it.

## Immediate Action

1. Stop propagation and affected high-consequence output or tool paths through approved containment.
2. Invoke the authoritative security, privacy, safety, or data-incident process; that process owns classification and notification.
3. Preserve minimal authorized evidence and access logs without copying the unsafe or sensitive payload into ordinary telemetry or collaboration channels.
4. Revoke exposed credentials, links, grants, caches, exports, or tool routes when applicable.
5. Bound recipients, tenants, outputs, traces, prompts, context sources, actions, and downstream systems.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Classify | Let the owning authority classify the content, data, policy, and consequence | Recorded incident class and handling rule | Classification requires unauthorized access |
| Trace | Follow protected references across ingestion, context, generation, validation, logging, tools, exports, and retention | Bounded propagation map | Cross-tenant or public exposure is possible |
| Identify control | Test authorization, context admission, output validation, redaction, logging, and tool enforcement | Failed or absent control | Test may reproduce exposure to users |
| Contain | Disable source, route, payload capture, export, or action; narrow access and require review | Enforcement and revocation receipts | Containment cannot prevent continued exposure |
| Remediate | Correct through the owning security, context, prompt, tool, telemetry, or release process | Approved fix and adversarial verification | A new policy or risk acceptance is required |

## Verification and Stopping Conditions

Verify revocation, caches, logs, indexes, exports, incident artifacts, downstream stores, affected users, output invariants, authorization, redaction, deletion, and audit coverage. Validate recovery with approved non-sensitive fixtures and targeted evidence.

Stop and escalate when exposure scope is unknown, deletion cannot be verified, credentials or cross-tenant data are involved, unsafe effects occurred, or legal or notification obligations may apply.

## Reversal

- Revoke temporary investigator access and elevated capture.
- Delete or quarantine disallowed copies under the authoritative retention and legal-hold decision.
- Restore sources and features only after access and validation controls pass.
- Correct, retract, or supersede unsafe outputs and reconcile external effects.

## Evidence and Closure

Record protected references, access history, affected scope, containment and revocation receipts, control failure, remediation, deletion or hold evidence, communications, and residual risk. Do not duplicate sensitive payloads in this runbook.

Close only under the owning incident authority after exposure is contained, required remediation and communication are complete, recovery verification passes, and residual obligations have owners. Reopen on new recipients, copies, effects, or failed deletion.

Applicability runbook

Evaluation Drift

Validates population change, preserves reference history, routes evidence applicability to Chapter 8, and prevents silent baseline normalization.

Preview
# Runbook: Evaluation Drift

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when an input population, context corpus, dependency, model behavior, route mix, workflow, outcome, policy, label, or telemetry distribution changes enough to question current evidence applicability.

## Entry Conditions and Authority

- Drift Signal ID, subject, representation, unit, slices, reference and current windows:
- Minimum volume, uncertainty, seasonality, expected change, freshness, and repeated-alert policy:
- Evaluation Plan, applicable release and population, evaluator, label, and telemetry identities:
- Signal owner, evaluation owner, service owner, telemetry owner, and incident commander:
- Authority to refresh a reference, route for evaluation, narrow population, or apply reversible containment:

Detection does not determine whether evidence is invalid. The Chapter 8 evaluation owner makes that decision.

## Immediate Action

1. Validate event, population, label, join, reference, and telemetry health before interpreting the difference.
2. Classify the condition provisionally as expected change, data defect, instrumentation defect, applicability warning, quality-regression candidate, or incident.
3. Preserve current and reference populations, representations, component identities, expected-change records, and release overlays.
4. Bound high-consequence use if existing evidence may no longer apply.
5. Do not refresh the reference merely to silence the alert.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Validate | Recompute with versioned data, representation, slices, and minimum volume | Reproducible difference with uncertainty | Reference or current data is unhealthy |
| Contextualize | Compare planned product, source, traffic, policy, label, dependency, and route changes | Expected or unexplained change | Sensitive slices are not authorized |
| Test applicability | Ask the evaluation owner to assess whether current evidence covers the new population or behavior | Recorded validity disposition | No suitable evaluation evidence exists |
| Contain | Narrow the affected population, require review, or use approved degraded behavior | Reduced consequence and receipt | Containment changes eligibility without authority |
| Update | Create new scenarios, datasets, thresholds, or reference windows through the Evaluation Plan | Versioned evidence and approval | Update would normalize a regression |

## Verification and Stopping Conditions

Verify current and reference data, expected change, evaluation applicability, affected quality and safety measures, population coverage, route and component identities, and telemetry health. Confirm that reference refreshes preserve history and authority.

Stop and escalate when drift coincides with unsafe behavior, affected consequence grows, evidence validity is unknown for a high-risk population, or no owner can accept the applicability decision.

## Reversal

- Revert temporary population restrictions or review gates only after validity and recovery evidence.
- Restore prior references if an unauthorized or defective refresh occurred.
- Supersede dashboards or claims based on invalid comparisons.
- Reconcile decisions made while applicability was uncertain.

## Evidence and Closure

Record subject, representations, windows, volumes, uncertainty, changes, classifications, evaluation-owner decisions, containment, reference history, and recovery evidence.

Close when the condition is classified, evidence applicability is recorded, required evaluation or containment is complete, and operational rules use the accepted reference. Reopen when the population shifts again or delayed quality evidence contradicts the decision.

Release runbook

Composite Rollback

Binds incident authority and work-state treatment to Chapter 10's controlled rollback, then requires operational recovery evidence.

Preview
# Runbook: Composite Rollback

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when incident authority chooses to restore an approved prior composite release or configuration. Chapter 10 owns rollback mechanics; this runbook binds the operational trigger, authority, scope, evidence, and recovery verification.

## Entry Conditions and Authority

- Incident, trigger, affected population, consequence, and rollback decision:
- Current and target release manifests, including code, prompt, dependency, context, tool policy, route, configuration, evaluator, and telemetry compatibility:
- State, schema, cache, index, queue, workflow-affinity, and irreversible-effect constraints:
- Incident commander, release owner, service owner, response operator, and recovery verifier:
- Decision authority, approved target, exposure scope, maximum duration, abort condition, and roll-forward alternative:

## Immediate Action

1. Freeze conflicting releases, configuration changes, context updates, and route changes.
2. Confirm the target is known good for the current population and compatible with current state and dependencies.
3. Bound new admission, queued work, in-flight workflows, long-running affinity, tool effects, and mixed-version operation.
4. Preserve current assignment, change, and effect receipts.
5. Execute only through the authorized release path; do not reconstruct a prior state from memory.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Validate target | Check manifest integrity, artifact availability, compatibility, policy, and prior evidence | Deployable approved target | Prior version has a known current vulnerability or incompatibility |
| Plan state | Classify admission, queues, active workflows, durable state, caches, indexes, and effects | Explicit treatment per state class | Irreversible migration or effect lacks safe handling |
| Stage | Apply target to the approved population and verify actual assignment | Deployment and route receipts | Mixed identity cannot be observed |
| Observe | Check technical, quality, safety, retrieval, action, cost, and telemetry signals | Expected reduction without new breach | Target worsens another protected objective |
| Expand or abort | Follow the decision thresholds in the Release Manifest | Recorded exposure decision | Evidence is indeterminate |

## Verification and Stopping Conditions

Verify actual composite identity, compatibility, new admission, queued and in-flight work, durable state, tool effects, caches and indexes, user behavior, quality, cost, telemetry, and delayed outcomes. The deployment receipt alone is not recovery.

Stop and escalate when the target is unavailable, incompatible, unobservable, unsafe for the present population, or unable to reconcile state and external effects.

## Reversal

- Abort the rollback or roll forward only through a separately recorded release decision.
- Remove emergency route pins, holds, and access at expiry.
- Reconcile mixed-version workflows and supersede stale artifacts.
- Restore paused releases only after incident command and release authority agree.

## Evidence and Closure

Record decision authority, manifests, assignments, state treatment, deployment and route receipts, abort thresholds, actions, effects, verification, and residual risk.

Close the rollback action when the target state and work population are reconciled. Close the incident only after the separate Recovery Verification Record passes and the closure authority acts. Reopen on route drift, compatibility failure, or delayed regression.

Continuity runbook

Degraded Mode

Activates a narrower pre-approved capability with enforced limits, user disclosure, expiry, active-work handling, and verified exit.

Preview
# Runbook: Degraded Mode

Toolkit schema: `PAISEH-TK-1.0`
Scenario pack: `PAISEH-SC-MIA-1.0`

Complete `../../../common-artifact-header.md` first. Use this runbook when the service must temporarily operate with narrower capability, population, authority, automation, freshness, or performance while preserving a protected outcome.

## Entry Conditions and Authority

- Incident, unavailable capability or evidence, affected population, and consequence:
- Approved degraded mode: read-only / draft-only / no-tool / no-retrieval / supervised / limited cohort / queued / other:
- Capabilities preserved, disabled, disclosed, and prohibited:
- Service owner, incident commander, response operator, communications owner, and recovery verifier:
- Authority, activation threshold, maximum duration, expiry, renewal authority, and return-to-normal criteria:

An ad hoc fallback is not an approved degraded mode. The mode must have known semantics, observable enforcement, and a tested exit.

## Immediate Action

1. Select the least-permissive approved mode that preserves the critical user outcome.
2. Enforce the capability boundary outside the model and record effective configuration and route receipts.
3. Account for new admission, queued and in-flight work, actions, durable state, and partial output.
4. Tell affected users what capability, freshness, latency, review, or result guarantees changed.
5. Start the expiry and observation clocks; assign the next renewal or recovery decision.

## Diagnosis and Response

| Step | Action | Expected evidence | Stop / escalate when |
| --- | --- | --- | --- |
| Validate mode | Compare trigger, population, consequence, and unavailable control with the approved mode | Applicable mode and authority | No approved mode covers the condition |
| Activate | Apply route, permission, admission, reviewer, source, or output restrictions | Enforcement and assignment receipt | Actual behavior cannot be observed |
| Reconcile | Classify queued, active, and effecting workflows under the new boundary | Known terminal or held state | Unknown external effects exist |
| Observe | Monitor protected outcome, user impact, quality, safety, cost, capacity, and evidence health | Stable bounded operation | Mode creates a different unacceptable risk |
| Renew or exit | Decide before expiry using current evidence | New receipt or verified return plan | Authority or recovery evidence is missing |

## Verification and Stopping Conditions

Verify that disabled capability cannot be invoked, affected users receive accurate behavior, tool and data boundaries hold, queues and in-flight work are reconciled, and the protected objective remains acceptable. Confirm the mode does not conceal missing telemetry or convert partial results into complete ones.

Stop and escalate when the mode expires, its enforcement is uncertain, affected population grows beyond approval, critical outcome cannot be preserved, or operator capacity becomes inadequate.

## Reversal

- Return to normal through controlled configuration and actual-assignment verification.
- Drain or revalidate work held during the degraded period.
- Remove temporary reviewer roles, grants, route pins, and user notices when authorized.
- Reprocess eligible work only with explicit user and policy treatment.

## Evidence and Closure

Record trigger, mode identity, authority, scope, configuration, assignment, user notice, queues, actions, objective results, renewals, expiry, and exit verification.

Close the degraded-mode action only after normal capability or a newly accepted operating mode passes recovery verification. Reopen when enforcement drifts, expiry arrives, or delayed outcomes reveal harm.