24/09/2026
AI coding agent evaluations are only as reliable as the environments they run in. A passing result doesn’t always mean the agent solved the task correctly—it may have used files, caches, or other hidden information. Microsoft highlights the need for clean sandboxes, clear boundaries, and trajectory reviews. For AI agents, the final answer matters, but so does how they reached it. Mastykarz , Microsoft
Your AI coding agent passed the eval. But did the model know the answer, or did it find it somewhere on your machine? A correct answer can still invalidate your measurement.