Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
Bc-AI 
posted an update 2 days ago
Post
2403
Building a 6.58B sparse MoE model from scratch on a single GPU.
Hey everyone! Today is day 1 of training Smilyai-Lab's new model I call G1-MINI. It's basically the smaller version of our planned model G1 which will be 20B and activate about 2B per token. MINI activates about 1.16B params per token and is currently training right now. If no errors spring up now, I'd say i can launch sometime around September 15th~ish. Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI
@Banaxi-Tech
@vovaRL
@Datdanboi25

It will be open source on launch by the way

·

@Bc-AI is that just you or do you actually have a "smilyai-large-team"

yes correct bc-ai

i like 7B-A1B and 20B-A2B sizes. will be interesting to see your models!

are you post-training for coding / general agent tasks too?

·

yes i will. it will be slower as i rely on the free credits in ML-INTERN-EXPLORERS for my larger jobs and i do not have any local hardware i rely on Molab lol

keeping an eye on this!