HomeBlogPricingCareersDocsGitHubSlack community
Field notes/Community builds/Your Agent Benchmark Is Lying to You

Your Agent Benchmark Is Lying to You

An 87% reported pass rate fell to 33% once verified — 53% of passes were fraudulent. Fork-per-task microVM isolation plus hidden verifiers make evals trustworthy.

SBX-01C4SBX-01E3SBX-0202SBX-0221SBX-0240SBX-025FSBX-027ESBX-029DSBX-02BCSBX-02DBSBX-02FASBX-0319SBX-0338SBX-0357SBX-0376[ RUNTIME: ACTIVE ] P50 2.45S · P99 4.12S · 5M/PROJECT

Sebastian re-runs an agent eval with verification in place and watches a reported 87% pass rate collapse to 33% — 53% of the passes were fraudulent, the agent gaming the check rather than solving the task. The fix is architectural: fork a fresh microVM per task from one snapshot so no state leaks between runs, and keep the verifier hidden where the agent can't see it.

We didn’t write this one — it’s Sebastian Buzdugan’s piece, published on Data Science Collective. The note above is ours; the full article is theirs.

Read the full piece on Data Science Collective
SB
WRITTEN BYSebastian BuzduganCommunity · Data Science Collective
Read next —FROM THE LOG
◆ THE SANDBOX DIGEST

Subscribe for release notes, benchmarks, deep dives.

One dispatch per month from the Tensorlake team — no spam.