1 hour ago · 44 min read8786 words · Tech · hide · 0 comments

I was nerdsniped! On the HF discussion page for one of the models I created for my previous post, AndrewThompson1233 asked if I'd considered trying out factorised embeddings -- something he's using for his model, Maba, and which was previously used in some other models, like ALBERT. It's a really nifty idea; you use a similar trick to LoRA as a way of reducing the number of parameters used for your embeddings and your output head. And for small models, those can be a disproportionate number of the total. For example, with my 163 million parameter GPT-2 small-style models, the embeddings are about 39 million of them; the output head is the same size, so with those two taken together, that's half the model just getting stuff in and out rather than actually doing the thinking. Even if you use weight tying, like the original GPT-2 models did, you'll wind up spending 30% of your "budget" on the single matrix shared between embeddings and the output head. Similarly, while larger models have…

No comments yet. Log in to reply on the Fediverse. Comments will appear here.