Over the past week or so, I've been working on some models, those released being CodeSoft/sorbet-25m and CodeSoft/sorbet-v2-25m. In general, I'm a little confused because no matter what hyperparameters I change or datasets I add/remove, the benchmarks never move up.
In a recent project, where I attached a TN-gram block to Sorbet-v2-25M, it still stayed the same on benchmarks despite the TN-gram clearly learning (due to the perplexity being lower with the TN-gram attached). When I changed the corpus to favor higher density text (the first paragraphs of Wikipedia articles and synthetic math), the benchmarks either stayed flat or went down.
Does anyone have ideas on what I can do to improve my models? I'd really appreciate any feedback!
Wow, SLM Arena is getting a lot of traffic! Thank you guys for showing your interest!
To handle the growing demand, I’m moving SLM Arena from a CPU Space to a ZeroGPU Space. Hopefully, this will let me add more models to SLM Arena while keeping it running fast.
I've also added a separate arena + leaderboard for base models!
If there are any models or features you’d like to see, let me know in a reply to this post or in a Community post on the Space!