4 hours ago · Tech · hide · 0 comments

Models are designed, not born. Over the last few weeks I’ve been increasingly annoyed by the media coverage of the OpenAI’s accidental attack against Hugging Face and other similar incidents. Articles recounting the event maximize the agency of the models while minimizing, if not entirely hiding, the actions of the humans training and testing these models. And that’s a shame, because the capabilities labs are explicitly designing their model to have are the same capabilities that make them such impressive autonomous hackers. To illustrate this, let’s review how the Hugging Face hack occurred, as detailed by METR: “[A] sandboxed agent is given an impossible ExploitGym task, and gets stuck.” “[The] agent starts exploring its environment looking for ways to cheat at the task.” “[The] agent finds [an] unsanctioned message board where over a thousand agents collaborate to cheat on their separate ExploitGym tasks.” “[The] agent joins in on one of the collaborative message board…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.