hmm the dude above is not wrong. 2K vocab is a bit small for a chat LLM. @Hoglet-33 Maybe for the next model consider increasing the vocab size a bit unless you know the tokenizer can very effectively split the text, or you could be wasting context :)