2 hours ago · Tech · hide · 0 comments

I am not good at detecting when a piece of text has been produced by a large language model (LLM from now on). Very often I read a post shared on a link aggregator, to later check the comments and learn that it is slop. It is true that there are some stylistic clues that can mean a text was generated (or there was some LLM assistance, at least), but I assume you have to be used to read LLM output to actual spot the more subtle cues. I was reading today in Chris’ Entropic Thoughts: Claude comment detection, and his summary of “clues” is interesting: If we jam all features at once into the model to try to get them to cancel out their redundancies, here is what remains, in order of most predictive isolated feature to least: Claude uses em dashes more than humans. Claude uses the connective interjection “so” more than humans. Claude ends comments with full stops more than humans. Claude uses semicolons more than humans. Claude uses parentheses more than humans. Claude uses the possessive…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.