My new course at UT Austin: AI Alignment Theory 0 ▲ Shtetl-Optimized 3 hours ago · 13 min read2618 words · Culture · hide · 0 comments This semester, I’ve been teaching a brand-new course, entitled CS395T AI Alignment Theory. Here’s the course description: The astounding progress of AI over the past decade has been accompanied by a rising fear: do we really understand how to align and control powerful AI systems—how to get them reliably to do what we wanted, or would want them to do on reflection, rather than merely what we said? If we succeed at building general-purpose superhuman intelligences along the current paradigm, should we expect that development to go well for humanity? Can we modify the design, training, monitoring, or scaffolding of those intelligences to help ensure that it goes well? While there’s been a great deal of recent empirical work touching on these questions, this course will concentrate mainly on theoretical and mathematical foundations. As a warning, the theoretical foundations of AI alignment have not yet gelled into any one coherent body of results accepted as canonical by the field.… No comments yet. Log in to reply on the Fediverse. Comments will appear here.