Boris-2-75M: Trained on 26B tokens -- estimated to start training on August 12th. Boris-2-125M: Trained on 90B tokens -- Estimated to start training on August 20th. Boris-2-250M: Trained on 60B tokens -- Estimated to start training on September 10th.
Why does 125M get more tokens than 250M?
Well, the straight answer is time. It saves time, while still allowing the 250M model to exceed the 125M model.
Furthermore, we are attempting a unique architecture and layering scheme to hopefully end up around the strength of SmolLM2-135M. Fingers crossed!
We're exited to announce BananaMind OS, our OS specically for running BananaMind models! Its able to run BananaMind 2 Nano at 4 bit on only 7-8MB of ram, the 2 bit on 6MB of ram and the 8 bit version on 14MB of RAM! It runs on a 486 or newer! Check out this video and image running BananaMind 2 Nano 4 Bit on 9 MB of RAM and a emulated 486 in QEMU at ~1TPS! We asked it: "What is the first letter of the alphabet?" The response is: "The first letter of the alphabet is: - A. " And if you're asking because of the video, yes I am a arch btw. Comment and like this post for a GitHub link and comment for adding other models!
GPT-X2.5-135M is finally released! 🚀 The new flagship from Axiomic Labs takes 3rd on the open SLM leaderboard trailing only the SmolLMs, check it out and follow us: AxiomicLabs/GPT-X2.5-135M