9 hours ago · 19 min read3750 words · Culture · hide · 0 comments

A long time ago, in what now feels like a galaxy far, far away, large language models were next-word predictors. You gave them an unfinished chunk of text, and they continued from where it left off with plausible-sounding sequences of words. When OpenAI announced in 2019 their ground-breaking GPT-2 model, they boasted that it "generates coherent paragraphs of text" and is "chameleon-like—it adapts to the style and content of the conditioning text". The term "stochastic parrot" would have felt irrefutable to anyone using those models at the time. All an LLM did was trace "word trajectories" through locations in meaning space, i.e. an internalized map of its training dataset. Actually, that is still all an LLM does today. If it feels very different, it's not (only) because it's gotten smarter. Then, at some point, people had the idea of cajoling the LLM into pretending to be in a conversation with someone. It's a very low-tech idea at its core: instead of having it complete an essay or…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.