1 hour ago · Tech · hide · 0 comments

In “Why Sheep Need Pigs in Sheepdog’s Clothing,” Robin Hanson asked three LLMs to score how much status, social skills, and judgment matter for influence across five domains. Judgment came back higher than his argument wanted. So he asked again. The second prompt told the models to focus on “who is selected or is influential in the short run,” and — read this twice — to “consider if they have any concrete basis for seeing judgment as mattering more re intellectuals and innovation.” Judgment dropped. He published both tables and treated the second as the refined estimate. That is a finding. It is not the finding he reported. He wasn’t testing the instrument; he was fixing an answer, and the fix exposed the instrument. Same models, same question, large movement on a rewording that names the direction it wants. A stable instrument returns small output changes for small input changes. This one didn’t. And the movement wasn’t noise. The second prompt doesn’t merely narrow scope — it…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.