AI Software Development – What Does The Data Say? 0 ▲ Codemanship's Blog 1 hour ago · Tech · hide · 0 comments I’m currently pulling together a bunch of sources – that are mostly recent – on the topic of LLMs and their use in software development. Some are peer-reviewed studies. Some are industry studies that haven’t been peer-reviewed. One is statistical physics. Expect more from that angle. Wanna’ know the limits of a technology? Ask a physicist. One is just a blog post, but very useful information about the effect of context size. Most are corroborated by personal experiments and also observations on teams. As time goes on and more data comes in, my picture comes more into focus. Code Evolution & Long-Horizon Agentic Workflows SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integrationhttps://arxiv.org/abs/2603.03823 SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Taskshttps://arxiv.org/html/2603.24755v1 SWE-Milestone: Evaluating AI Agents on Continuous Software Evolutionhttps://arxiv.org/abs/2603.13428 Context Engineering… No comments yet. Log in to reply on the Fediverse. Comments will appear here.