1 hour ago · 8 min read1599 words · Tech · hide · 0 comments

The highly publicised sandbox agent escapes have certainly made news, and I wrote about the issues with sandboxing agents back in January - though I certainly didn't foresee they would escape the frontier labs. I assumed the real risk was poorly configured sandboxes for end users, so I was surprised to see this happening at the frontier labs. I think it might tell us something about the security philosophy of these organisations. Safety vs security In my mind, AI safety is about "alignment". Will the AI do morally suspect tasks? Will it teach you how to make methamphetamine from household ingredients, encouraging a whole new generation of Jesse Pinkmans? So far, this has really been attempted via two main mechanisms, classifiers (where a separate model checks what the user has been sending, and flags potentially malicious requests and refuses them), and pre/post training safety techniques, where you adjust the weights of the model to itself refuse to obey potentially bad requests.…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.