1 hour ago · 25 min read5055 words · Tech · hide · 0 comments

Intro2 months back, I started the Prime Residency with Florian Brand as my peer or shall I say my verifier. My first task was to port the Agents’ Last Exam Linux CLI subset to verifiers. Then one day Florian asked me to look into the data and this quest slowly evolved into finding issues in the benchmark and the making of what we call ALE-Gold. It was more challenging than I expected. I was new to this kind of work, so there was some skill issue on my part. Along the way, I also hit benchmark defects, inference provider failures, and gaps in the freshly released Verifiers v1 (that have been fixed now so you don't need to worry). This post is a work log of my porting work, what the full runs revealed about the benchmark, how that led to ALE-Gold, followed by an analysis of how the models performed. A small odyssey from porting ALE to ALE-Gold, with a few unexpected detours. Table of contents Porting ALE Linux CLI subset to verifiers What is a Rollout? What porting ALE involved and some…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.