@Qozimo I didn't realize using a 2048 vocab would cause this much discourse, so i guess you might not like to hear that i forgot up the context and vocab on Pebble 25M. banaxi is right though chinchilla isn't what we use in the small LM space, BananaMind2 models were trained on 33B tokens and other SLMs go above and beyond chinchilla. I guess i can up the vocab on the next generation of Pebble but if you think Pebble sucks then so be it i just want to see a 10M model from you that beats not only Pebble 10M but also BananaMind2-Nano