Release Scorecard

Five questions before production launch.

A concise readiness gate for deciding whether an AI system should launch, pilot, narrow scope, or stop.

Problem Fit

Ready:

AI is justified by the workflow, not by novelty or model availability.

Not ready:

The system starts from a demo or capability claim.

Boundaries

Ready:

Inputs, outputs, data access, tool authority, and human approvals are explicit.

Not ready:

The prompt or model is expected to infer product policy.

Grounding

Ready:

Sources are authorized, fresh enough, trusted, and handled honestly when missing.

Not ready:

The model is pressured to answer without required evidence.

Evaluation

Ready:

Release evidence matches the system's risk, autonomy, and user exposure.

Not ready:

Confidence comes from demos, anecdotes, or generic model benchmarks.

Operations

Ready:

The system can be traced, rolled back, monitored, costed, and improved after launch.

Not ready:

Infrastructure health is treated as proof of product correctness.

Decision outcomes

Do not treat readiness as only yes or no.

  1. Proceed when all five areas have evidence and owners.
  2. Pilot when uncertainty remains but blast radius is bounded.
  3. Narrow when the decision surface or autonomy level is too broad.
  4. Block when consequence is high and evidence, control, or accountability is weak.