1 hour ago · 17 min read3419 words · Tech · hide · 0 comments

As I’ve mentioned, many people still think of AI as a text-in text-out language model trained on next-token prediction. But modern AI is much more than that. This post discusses the key AI innovations that improved on "basic" LLMs, the capabilities these innovations unlocked, and what's coming next.See this brief post for the stages of AI before LLMs:Note: I manually wrote my previous Substack posts and just used AI for feedback. For this post, I’m experimenting with using AI to help write significant chunks based on my outline and notes. If you don’t want to read the whole thing, you can just view this picture:Pre-training vs Post-trainingThe first LLMs were literally just built on text prediction, which made it difficult to chat with. For example, if you entered “What is the capital of France”, it might output “What is the capital of Germany”, since that could often appear after similar text. The first LLMs only had “pre-training”, i.e. they were trained on large amounts of text to…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.