Agent Behavioral Eval Runner
Scenario
Use after production traces have been normalized and a candidate agent output has been captured in a redacted test run.
Input and output
Input contains normalized cases with constraints and a candidate_output. Output gives per-case PASS, REVIEW, or BLOCKED, failed constraints, missing evidence markers, and an aggregate score.
Execution and recovery
Run the README command. Missing candidates or malformed case arrays return a visible error. A REVIEW result means the evidence or wording needs human inspection; it is not a claim of model quality.
Privacy boundary and acceptance
Use redacted text and synthetic evidence markers only. The sample shows a pass and a review. The runner does not contact an agent, store prompts remotely, or infer business success.