Map the workflow
Identify tools, state, policies, dependencies, permissions and the evidence available when something goes wrong.
Methodology
The goal is not to create a large task count. The goal is to reproduce a failure that matters and determine whether an agent actually avoids it.
Identify tools, state, policies, dependencies, permissions and the evidence available when something goes wrong.
Describe the wrong result a capable agent may produce, why it appears reasonable and what downstream invariant it breaks.
Create deterministic services, representative synthetic or customer-approved data, reset behavior and solver-visible evidence.
Measure resulting behavior and preservation invariants independently from the agent's written answer or self-reported status.
Run oracle, untouched, wrong-target, duplicate-action, shortcut and reward-hacking controls.
Run selected models or the customer's agent, retain trajectories and separate genuine agent failures from evaluation defects.
Delivery evidence
Each engagement can include environment source or hosted access, task and verifier packs, run manifests, machine-readable results, selected trajectories, a failure taxonomy, technical reporting and a next-experiment plan. Access is matched to the customer's security and IP requirements.
Bring us the workflow and the failure you need to detect.