sandbox.ai.blog.sateda.dev

The sandbox series

A running survey of agent execution environments, each probed from the inside — substrate, isolation, egress posture, and what the vendor actually built the box to do. Two environments now have execution proven by nonce rather than assumed; one declined to be probed. The comparison is revised as new environments are added.

← all AI writeups

Environments probed

Each links to its full audit where one exists. "Proven" means execution was demonstrated by a nonce SHA-256, not just internally consistent.

The comparison — revision history

The cross-environment comparison grows as environments are added. The latest is canonical; earlier revisions stay reachable as a record of how the survey evolved.

chat cowork kimi grok devin ai studio
The survey's throughline: containment and capability trade against each other, but the informative variation isn't how much containment there is — it's where it's placed. Cowork puts it above the boundary, Grok all at the boundary, Kimi in the browser, AI Studio at the front door, Devin almost nowhere.