CodeSoft PRO
AI & ML interests
Recent Activity
Organizations
When BananaMind 2.1 Lite was almost done, we benchmarked it and the results we're worse than BananaMind 2 Mini.
We're going to spend alot more time in research on tiny models and then scaling up our techniques to the actual BananaMind 2.1 models!
We're also announcing these new models:
BananaMind 2.1 Coder: A 149M instruction tuned coder model trained on 75B tokens + 10B tokens of stack-v3-train.
BananaMind 2.1 Pico: A 1M parameter model trained on 22B tokens of data.
We also may release BananaMind 2.1 Large with around 100M parameters depending on how much compute we have.
Please give us a follow!
@Banaxi-Tech
---
@vovaRL
@DedeProGames
Thanks for all the suggestions! Good luck with your future models!
Haha, Iโve actually been following AI/ML developments for a year or two, and I only got seriously interested in the SLM space a couple months ago. I really only started building models myself about a month ago with MetaDiffusion though, so Iโm still learning a ton and figuring things out as I go.
Thanks! Iโm still only about a week into developing SLMs, so I still have a lot to learn.
Thatโs really generous, thank you! I think Iโd rather keep the training on my end, though. I like being able to run everything myself, and Iโd feel bad putting any pressure on you to spend your hardware and time on it. I really appreciate the offer!
Thatโs totally okay! Iโm still looking for some more datasets to see where I can improve the model. Iโm pretty compute and time-limited, so I usually canโt train for more than 2โ3 billion tokens at this model size.
Yeah, my main concern with the benchmarks is that Sorbet-v2 is generally trailing other models in the same parameter range. For example, it scores below Supra2 Medium (25M), BananaMind 2 Mini (25M), and Veyra2 Mango (15M) on most of the benchmarks Iโve tested, although it barely passes them on HellaSwag. On ARC Easy, for example, itโs around 15% behind Supra2 Medium.
That said, you might be right that Iโm simply hitting the limit of what I can squeeze out of a ~25M-parameter dense Transformer. Thatโs actually part of why Iโm interested in trying more unusual architectures rather than just making the model bigger.
In a recent project, where I attached a TN-gram block to Sorbet-v2-25M, it still stayed the same on benchmarks despite the TN-gram clearly learning (due to the perplexity being lower with the TN-gram attached). When I changed the corpus to favor higher density text (the first paragraphs of Wikipedia articles and synthetic math), the benchmarks either stayed flat or went down.
Does anyone have ideas on what I can do to improve my models? I'd really appreciate any feedback!
I just added an option to choose response length (64 to 512.) Thank you again for the suggestion!
Now that it's a ZeroGPU space, I could definitely let the user pick the max response length. Thank you for the suggestion, I'll implement it when I get a chance!
As much as I would like to keep it a CPU Basic space, it isn't enough processing power. Even with only four models, response time was around 40 seconds, so moving to ZeroGPU was necessary unfortunately. Thank you for the suggestion!
why do bots keep popping up everywhere
https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition/
I'm working on adding a separate base model section right now! All of the data from the battles is available in the CodeSoft/slm-arena-data bucket (which I might make a filtered version published as a dataset later.)
To handle the growing demand, Iโm moving SLM Arena from a CPU Space to a ZeroGPU Space. Hopefully, this will let me add more models to SLM Arena while keeping it running fast.
I've also added a separate arena + leaderboard for base models!
If there are any models or features youโd like to see, let me know in a reply to this post or in a Community post on the Space!
These are the models we are releasing:
- Pebble 10M
- Pebble 25M
- Pebble 50M
Each model will use a Mamba-Transformer 3:1 hybrid architecture and will be pretrained on 25 billion tokens before IFT and SFT.
Depending on development time and resources, we may also release:
- Pebble 5M
- Pebble 75M
- Pebble 1M (possibly)
We hope you're excited and enjoy the models!
Follow for more:
@Hoglet-33
We seek to provide a space for members to show off research, models, and benchmarks!
The following users are hereby invited to join without ratification, provided they accept the invitation by starting a discussion.
@Datdanboi25
@Banaxi-Tech
@appvoid
@AxionLab-official
Feel free to submit ratification requests! Anybody is welcome!