9 days ago · 9 min read1733 words · Tech · hide · 0 comments

This is a piece I wrote for The Byte, the AI Collective's Tuesday newsletter. You can read it here. There is no generally accepted definition of AGI, but most of us can agree that, in broad strokes, AGI is a label for any LLM that is smarter than most (or all) humans in most (or all) fields. Sycophancy, by contrast, is the tendency of models to agree with, flatter, or defer to users even when doing so requires abandoning correctness.[1][4] In some cases, sycophancy is not a significant problem and may even be useful: flattery can help LLMs be more persuasive and, in some cases, make them easier to talk to and open up to.[9] A bit of sycophancy may be the right tool for achieving specific goals. What we see in today’s models, however, is not intentional behaviour but a symptom of a deeper problem, and it is one of the main roadblocks that may prevent us from reaching true AGI. Reward Hacking Sycophancy is not an inherent feature of large language models. During pre-training, LLMs learn…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.