Post
2403
Building a 6.58B sparse MoE model from scratch on a single GPU.
Hey everyone! Today is day 1 of training Smilyai-Lab's new model I call G1-MINI. It's basically the smaller version of our planned model G1 which will be 20B and activate about 2B per token. MINI activates about 1.16B params per token and is currently training right now. If no errors spring up now, I'd say i can launch sometime around September 15th~ish. Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI
@Banaxi-Tech
@vovaRL
@Datdanboi25
Hey everyone! Today is day 1 of training Smilyai-Lab's new model I call G1-MINI. It's basically the smaller version of our planned model G1 which will be 20B and activate about 2B per token. MINI activates about 1.16B params per token and is currently training right now. If no errors spring up now, I'd say i can launch sometime around September 15th~ish. Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI
@Banaxi-Tech
@vovaRL
@Datdanboi25