Case Study / Proofrail

PROOFRAIL.

Provider-neutral evaluation gates with explicit controls, repeated trials, and evidence you can inspect.

Make the release decision inspectable.

Proofrail separates model execution from evaluation. A common JSON envelope carries text, structured outputs, tool calls, traces, artifact metadata, external review scores, cost, and latency. Versioned rubric criteria turn that evidence into a local release review.

From specification to evidence

  • Run the 21-surface preflight to inspect design, controls, trial coverage, evidence policy, isolation declarations, and corpus integrity.
  • Score a known-good Oracle and an empty negative control to check that the rubric can discriminate.
  • Review repeated runs with deterministic evaluators and supplied subjective scores; inspect criterion failures alongside aggregate pass rate.
  • Export Markdown and JSON reports that can be reviewed locally or retained by CI.

Reproduce the synthetic demo

Clone the repository and use Python 3.10 or later. These commands use only the public synthetic LedgerKit fixtures; no provider credentials are required.

python -m unittest discover -s tests -v
python -m proofrail preflight --spec examples/ledgerkit/spec.json --runs examples/ledgerkit/runs.json --corpus examples/regression-corpus/corpus.json --out artifacts/preflight
python -m proofrail review --spec examples/ledgerkit/spec.json --runs examples/ledgerkit/runs.json --corpus examples/corpus.json --out artifacts

Compare the generated reports with the committed text-model evidence and inspect the agent-envelope fixture. A passing sample demonstrates the configured checks on supplied runs, not the quality of a live provider model.

Prototype boundaries

Proofrail is a local evaluation and CI prototype, not an operated enterprise service. It does not execute models, independently calibrate subjective scores, establish statistical confidence, or inspect image/video contents by default. Evidence-policy and isolation flags are declarations, not security enforcement. Run preflight and review together; a release decision applies only to the configured rubric and supplied evidence.

See the enterprise readiness roadmap for remaining execution, governance, security, and statistical work. This public case study uses synthetic examples and contains no confidential evaluation tasks.

← Back to selected work