●builderYou can implement this to more rigorously test if your agents are actually following complex logic rather than just producing plausible text.
●researcherThis moves evaluation beyond simple trace observation toward explicit obligation matching.