54 minutes ago · 6 min read1251 words · Tech · hide · 0 comments

Every frontier lab runs the largest canon-formation project in literary history and calls it a data pipeline. Look at the actual operations. Crawl the written world. Deduplicate it. Filter it for “quality” — a value judgment wearing a lab coat. Weight the mixture. Sequence the curriculum. Then, at the end, write a constitution instructing the finished reader on who to be. Every one of these is an editorial act. It is syllabus design performed by institutions that would be insulted to be called schools, and the reader it produces is the most consequential reader in the world. To be fair to the field, it studies its data obsessively. Contamination checks, toxicity filters, domain weights, influence functions tracing an output back to the documents that caused it, curriculum-ordering experiments, the whole textbooks-are-all-you-need research program. All of it asks one family of questions: what the data does to capability and behavior — what the model can do, will say, might leak. None…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.