1 hour ago · Tech · hide · 0 comments

Source composition and averages. I'm proud to announce Bagaço v3, the third version of the largest European Portuguese pretraining dataset for large language models. Why a new version? After my work on Ginjinha, one thing became pretty apparent: Bagaço was large and diverse, but lacked high-quality data. I believe Bagaço v3 fixes that. This new version builds on top of Bagaço v2 and the great work from the Hugging Face team, and adds documents from Wikipedia (FineWiki) and from PDFs across the web (FinePDFs). But that's not the full story. Let's get into the details. Bagaço v3 sources. To create Bagaço v3, I started by using DataTrove to build a pipeline that merged Bagaço v2 (originally from FineWeb2), FineWiki, and FinePDFs. It filters the last two by European Portuguese score (using my own classifier, which I wrote about here), and deduplicates everything using MinHash deduplication. But we don't want to dump everything into the same place just like that. Just like Bagaço v2, v3…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.