1 hour ago · 6 min read1167 words · Tech · hide · 0 comments

These are some things I’ve wandered across on the web this week. 🔖 Elasticsearch isn’t your AI search problem. You are. Why data governance and evaluation matter when building search systems, and especially when doing Retrieval Augnented Generation. llm search 🔖 Can AI agents conduct open-ended AI research? Early evidence from two case studies Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper’s original authors grade its output. We call these shadow evaluations. We ran shadow…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.