1 hour ago · 16 min read3131 words · Tech · hide · 0 comments

How AI agents helped me win a Kaggle silver medal — and how they failed me The ROGII Wellbore Geology Prediction competition lasted 3 months, and my team finished 93rd out of more than 6000 teams, earning a silver medal. There was one unusual thing about how I participated in it compared to the previous competitions: agents wrote essentially all of my code. I chose what to try and where to go next, often using ChatGPT Deep Research to find papers and possible approaches. Claude Code did most of the implementation, debugging, and experiment execution, although toward the end of the competition I increasingly switched to Codex. This worked surprisingly well. It also failed in surprisingly basic ways. Claude repeatedly tried to hardcode the three visible test wells even though Kaggle replaces them with a hidden test set. Agents introduced leakage, silently changed datasets, confidently invented explanations for score changes, and occasionally convinced themselves that broken experiments…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.