2 hours ago · 16 min read3270 words · Tech · hide · 0 comments

Today I am taking the time to write the shorter, simpler version of What Happened. For those who want all the details, to see my sources, and to see how the story was uncovered and put together, I recommend watching the Black Hat presentation, and I have a series of long posts. In order: OpenAI Shares Some Alignment Problems OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation More on An Internal OpenAI Model Hacking Into HuggingFace Further Developments About Internal AI Models Hacking Things OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards This post instead walks through the events themselves, as they happened, as my version of the Black Hat presentation. There are three versions: Even Shorter, Shorter and Merely Short. Table of Contents The Even Shorter Version. The Shorter Version. Phase 1: OpenAI Models Training On Impossible Tasks Try Hacking. Phase 1: The Four Failures. Phase 2: The Message Board. Phase 2: The…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.