Notes on long-running LLM tasks 0 ▲ Nate Meyvis 1 hour ago · Tech · hide · 0 comments I often read about getting LLMs to work for a long time toward a goal. So, for example, Claude has a /goal command, and many of us (directly or indirectly) care a lot about how models do on TerminalBench, which is made up of tasks in which agents get one shot at doing a big chunk of work. More generally, long-running tasks are the sort of thing you hear about on podcasts, read about on Twitter, and so on: how large a bite can an agent chew? There are some obvious, good reasons to care about this: Lots of work fits this paradigm well: certain kinds of optimization, research, and experimentation provide clear examples. It's reasonable to expect that more and more work will be amenable to this kind of structure, so caring about long-running tasks is a way of keeping perspective about the future of the craft. Even if work isn't ideally suited for this structure, it might be better than alternatives (e.g., when no project-relevant human will be available for a while). It's really cool. But… No comments yet. Log in to reply on the Fediverse. Comments will appear here.