Superpersuasion will look like bribery 0 ▲ seangoedecke.com 1 hour ago · Culture · hide · 0 comments The idea of “superpersuasion” has been floating around the AI safety community for decades. Now that powerful and difficult-to-control LLMs have appeared, people are again talking about the idea that a sufficiently intelligent AI might be able to persuade people to do whatever it wants. The classic1 version of this idea is an AI persuading some engineer to “let it out of the box”: to grant it access to the internet. The modern version of this idea is the “killswitch”, where if some AI ever does go rogue, the humans in control of its hardware will simply turn it off. Could an AI somehow persuade these humans to not hit the switch? When AI safety people talk about superpersuasion, they often talk about it like rationalist nerds. These people spend their lives trying to follow the most convincing arguments and calibrate their positions as closely towards the truth as possible. In philosophy terms, these are “bullet-biters”: people who are ready to accept a ridiculous-sounding conclusion… No comments yet. Log in to reply on the Fediverse. Comments will appear here.