Skip to main content

Evals

Score your agent and model outputs with reliability statistics: pass^k over repeated runs, code and judge graders, human review, and tamper-evident traces. Build one here, or stream runs in from your own harness over the REST API or MCP.