2 hours ago · Tech · hide · 0 comments

Unfortunately, monitoring for adversarial data access turns out to be one of the hardest problems you could imagine. The volume of data that agents produce is so high that no human being could possibly read it, and we probably wouldn’t recognize obfuscated malicious data even if we were looking directly at it. This means any attempt to monitor the inflow/outflow will have to be handled by other models. Thus, the future of agent sandboxing is (1) build a sandbox, (2) install an agent/model into it, (3) install a somewhat dumber/cheaper warden model to guard it, (4) hope you can trust the lunkhead to contain the wizard. And so on and so forth, as models become more intelligent and capable. In other words: a warden-guarded sandbox is just another version of the alignment problem. You’re going to have to trust a model to do it, and that model will need to be at least some fraction as intelligent as the model it’s guarding. If you haven’t convinced yourself that it’s possible to build…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.