alkinun
AtAndDev
AI & ML interests
decentralize
Recent Activity
liked a dataset 26 minutes ago
marin-community/token-counts liked a Space 29 minutes ago
marin-community/token-count-viewerOrganizations
reacted to appvoid's post with π₯ about 11 hours ago
reacted to Bc-AI's post with π₯ 6 days ago
Post
2523
Building a 6.58B sparse MoE model from scratch on a single GPU.
Hey everyone! Today is day 1 of training Smilyai-Lab's new model I call G1-MINI. It's basically the smaller version of our planned model G1 which will be 20B and activate about 2B per token. MINI activates about 1.16B params per token and is currently training right now. If no errors spring up now, I'd say i can launch sometime around September 15th~ish. Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI
@Banaxi-Tech
@vovaRL
@Datdanboi25
Hey everyone! Today is day 1 of training Smilyai-Lab's new model I call G1-MINI. It's basically the smaller version of our planned model G1 which will be 20B and activate about 2B per token. MINI activates about 1.16B params per token and is currently training right now. If no errors spring up now, I'd say i can launch sometime around September 15th~ish. Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI
@Banaxi-Tech
@vovaRL
@Datdanboi25
reacted to KlondikeDev's post with π₯ 6 days ago
Post
3437
A Preview of Boris-2!
Hello! Tomorrow, OpenCerebral will be releasing Boris-1.7-D60M-n30M β an experimental architecture. It will be testing a new data mixture, a new tokenizer, and testing Qwen4-like n-gram embeddings.
Following this will be Boris-1.8-D60M-n30M, which will test both the n-gram embeddings AND a new architecture.
Then, Boris-2 will begin training!
Edit: the model is OUT NOW! opencerebral/Boris-1.7-D60M-n30M
Hello! Tomorrow, OpenCerebral will be releasing Boris-1.7-D60M-n30M β an experimental architecture. It will be testing a new data mixture, a new tokenizer, and testing Qwen4-like n-gram embeddings.
Following this will be Boris-1.8-D60M-n30M, which will test both the n-gram embeddings AND a new architecture.
Then, Boris-2 will begin training!
Edit: the model is OUT NOW! opencerebral/Boris-1.7-D60M-n30M
replied to KlondikeDev's post 6 days ago
wow... my eyes..
posted an update 6 days ago
Post
2778
SPECK 2 IS ALREADY OUT: specklabs/Speck2-140M
Pretrained on 4x more tokens than the previous releases (20b vs 5b).
Instruct tuned versions are coming soon.
Very interesting models are coming soon too (hint: super long context).
Thanks for everyone supporting!
Pretrained on 4x more tokens than the previous releases (20b vs 5b).
Instruct tuned versions are coming soon.
Very interesting models are coming soon too (hint: super long context).
Thanks for everyone supporting!
reacted to KlondikeDev's post with π₯ 6 days ago
Post
2382
Boris-1.7-D60M-n30M out NOW!
opencerebral/Boris-1.7-D60M-n30M
Following this will be Boris-1.8-D60M-n30M, which will test both the n-gram embeddings AND a new architecture.
Then, Boris-2 will begin training!
opencerebral/Boris-1.7-D60M-n30M
Following this will be Boris-1.8-D60M-n30M, which will test both the n-gram embeddings AND a new architecture.
Then, Boris-2 will begin training!
reacted to Hoglet-33's post with π₯ 7 days ago
Post
3284
Pebble 10M and Pebble 10M Chat are now released!
Both models use our Mamba/Transformer 3:1 hybrid architecture and were pretrained on 25 billion tokens.
Pebble 10M Chat was additionally fine-tuned on 250 million tokens of Smol-SmolTalk to improve its conversational capabilities.
You can find them here:
- basically-ai/Pebble-10M
- basically-ai/Pebble-10M-Chat
We hope you enjoy using them. The rest of the Pebble family will be released soon.
Follow for more:
@Hoglet-33
basically-ai
Both models use our Mamba/Transformer 3:1 hybrid architecture and were pretrained on 25 billion tokens.
Pebble 10M Chat was additionally fine-tuned on 250 million tokens of Smol-SmolTalk to improve its conversational capabilities.
You can find them here:
- basically-ai/Pebble-10M
- basically-ai/Pebble-10M-Chat
We hope you enjoy using them. The rest of the Pebble family will be released soon.
Follow for more:
@Hoglet-33
reacted to Banaxi-Tech's post with π 8 days ago
Post
3504
We're delaying BananaMind 2.1!
When BananaMind 2.1 Lite was almost done, we benchmarked it and the results we're worse than BananaMind 2 Mini.
We're going to spend alot more time in research on tiny models and then scaling up our techniques to the actual BananaMind 2.1 models!
We're also announcing these new models:
BananaMind 2.1 Coder: A 149M instruction tuned coder model trained on 75B tokens + 10B tokens of stack-v3-train.
BananaMind 2.1 Pico: A 1M parameter model trained on 22B tokens of data.
We also may release BananaMind 2.1 Large with around 100M parameters depending on how much compute we have.
Please give us a follow!
BananaMind
@Banaxi-Tech
---
@vovaRL
@DedeProGames
When BananaMind 2.1 Lite was almost done, we benchmarked it and the results we're worse than BananaMind 2 Mini.
We're going to spend alot more time in research on tiny models and then scaling up our techniques to the actual BananaMind 2.1 models!
We're also announcing these new models:
BananaMind 2.1 Coder: A 149M instruction tuned coder model trained on 75B tokens + 10B tokens of stack-v3-train.
BananaMind 2.1 Pico: A 1M parameter model trained on 22B tokens of data.
We also may release BananaMind 2.1 Large with around 100M parameters depending on how much compute we have.
Please give us a follow!
@Banaxi-Tech
---
@vovaRL
@DedeProGames
reacted to sergiopaniego's post with π₯ 8 days ago
Post
393
catching up on some bookmarked reads from the summer, reading Antidoom from @liquidai
small reasoning models get stuck more easily when the task involves a long thinking trace and a hard problem. It starts repeating the same word over and over again ("Wait", "Alternatively"β¦), each repetition makes the next one likelier, and the generation is spent before it reaches an answer
they measured it, 10.2% of completions for an early LFM2.5-2.6B checkpoint and 22.9% for Qwen3.5-4B at greedy. After training those drop to 1.4% and 1.0%
the fix is FTPO (final token preference optimization). What I like is how narrow it is, it only touches the single token where the loop starts
three ways it differs from DPO:
> trains one token position, mid-generation, instead of whole sequences
> spreads probability across ~20 plausible alternatives instead of swapping one overtrained token for another
> keeps the regularizer in logit space, no softmax, so the rest of the vocabulary stays put
the third one is what makes it usable. If you want to edit one position without disturbing the model, you can't have a loss that reshuffles the other 150k logits on the way
and their explanation abt the result: the training teaches the model nothing new about math or code, it clears the failure mode that was blocking answers the model could already produce
full blog > https://www.liquid.ai/blog/antidoom
FTPO itself comes from Antislop, where it was built to strip overused phrasing. LiquidAI retargeted it to doom loops
and under the hood it's a subclass of TRL's DPOTrainer with compute_loss overridden, around 90 lines of loss and no new trainer
we documented that pattern in TRL's docs
https://huggingface.co/docs/trl/main/en/customization#change-the-training-objective
small reasoning models get stuck more easily when the task involves a long thinking trace and a hard problem. It starts repeating the same word over and over again ("Wait", "Alternatively"β¦), each repetition makes the next one likelier, and the generation is spent before it reaches an answer
they measured it, 10.2% of completions for an early LFM2.5-2.6B checkpoint and 22.9% for Qwen3.5-4B at greedy. After training those drop to 1.4% and 1.0%
the fix is FTPO (final token preference optimization). What I like is how narrow it is, it only touches the single token where the loop starts
three ways it differs from DPO:
> trains one token position, mid-generation, instead of whole sequences
> spreads probability across ~20 plausible alternatives instead of swapping one overtrained token for another
> keeps the regularizer in logit space, no softmax, so the rest of the vocabulary stays put
the third one is what makes it usable. If you want to edit one position without disturbing the model, you can't have a loss that reshuffles the other 150k logits on the way
and their explanation abt the result: the training teaches the model nothing new about math or code, it clears the failure mode that was blocking answers the model could already produce
full blog > https://www.liquid.ai/blog/antidoom
FTPO itself comes from Antislop, where it was built to strip overused phrasing. LiquidAI retargeted it to doom loops
and under the hood it's a subclass of TRL's DPOTrainer with compute_loss overridden, around 90 lines of loss and no new trainer
we documented that pattern in TRL's docs
https://huggingface.co/docs/trl/main/en/customization#change-the-training-objective
posted an update 9 days ago
reacted to OppaAI's post with π₯ 9 days ago
Post
2683
Congrat to HuggingFace Team
https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition/
https://techcrunch.com/2026/08/26/nvidia-closes-in-on-hugging-face-acquisition/
reacted to Banaxi-Tech's post with π₯ 10 days ago
Post
2145
We have updated the BananaMind Base Bench leaderboard!
We now have these benchmark cards, they make it way easier to see which models are actually good!
We've also added the model advisor. It asks you what you want to use the model for and the parameter range and gives you the best model for your task!
Try it out at BananaMind/BananaMindBench-Leaderboard
And please give us a follow to BananaMind!
BananaMind
@Banaxi-Tech
We now have these benchmark cards, they make it way easier to see which models are actually good!
We've also added the model advisor. It asks you what you want to use the model for and the parameter range and gives you the best model for your task!
Try it out at BananaMind/BananaMindBench-Leaderboard
And please give us a follow to BananaMind!
@Banaxi-Tech
reacted to Hoglet-33's post with β€οΈππ₯ 10 days ago
Post
3047
We are announcing the first generation of the Pebble model family!
These are the models we are releasing:
- Pebble 10M
- Pebble 25M
- Pebble 50M
Each model will use a Mamba-Transformer 3:1 hybrid architecture and will be pretrained on 25 billion tokens before IFT and SFT.
Depending on development time and resources, we may also release:
- Pebble 5M
- Pebble 75M
- Pebble 1M (possibly)
We hope you're excited and enjoy the models!
Follow for more:
@Hoglet-33
basically-ai
These are the models we are releasing:
- Pebble 10M
- Pebble 25M
- Pebble 50M
Each model will use a Mamba-Transformer 3:1 hybrid architecture and will be pretrained on 25 billion tokens before IFT and SFT.
Depending on development time and resources, we may also release:
- Pebble 5M
- Pebble 75M
- Pebble 1M (possibly)
We hope you're excited and enjoy the models!
Follow for more:
@Hoglet-33
reacted to CodeSoft's post with π₯β€οΈ 10 days ago
Post
2407
Wow, SLM Arena is getting a lot of traffic! Thank you guys for showing your interest!
To handle the growing demand, Iβm moving SLM Arena from a CPU Space to a ZeroGPU Space. Hopefully, this will let me add more models to SLM Arena while keeping it running fast.
I've also added a separate arena + leaderboard for base models!
If there are any models or features youβd like to see, let me know in a reply to this post or in a Community post on the Space!
To handle the growing demand, Iβm moving SLM Arena from a CPU Space to a ZeroGPU Space. Hopefully, this will let me add more models to SLM Arena while keeping it running fast.
I've also added a separate arena + leaderboard for base models!
If there are any models or features youβd like to see, let me know in a reply to this post or in a Community post on the Space!
reacted to GoktugD's post with β€οΈ 10 days ago
Post
2359
πΉπ· We trained a 1B OCR model specifically for Turkish enterprise documents.
**Werea-DocOCR-1B v2**
The result surprised us:
LightOnOCR-2 base β **64.2% CER**
Werea-DocOCR v1 β **~8.1% CER**
Werea-DocOCR v2 β **0.15% CER** π
Evaluated on a held-out 72-page test set across 12 Turkish document types and 3 different capture conditions.
π 12 Turkish enterprise document types
π§ͺ 12,960 synthetic training pages
π± Digital + scanned + phone photos
π Tables β structured Markdown
βοΈ Full-parameter fine-tuning
π₯οΈ Trained on a single RTX 3090
It handles:
β’ e-Invoices
β’ rental contracts
β’ bank receipts
β’ payroll documents
β’ insurance policies
β’ vehicle documents
β’ official correspondence
β’ trade registry documents
β’ SGK-style tables
β’ and more.
**Model π€**
Werea-co/Werea-DocOCR-1B
**Dataset π**
Werea-co/werea-tr-doc-ocr-enterprise-v2
**Werea πΉπ·**
Werea-co
We're building open AI models from TΓΌrkiye.
This is just the beginning.
#HuggingFace #OCR #DocumentAI #TurkishAI #OpenSourceAI #ComputerVision
**Werea-DocOCR-1B v2**
The result surprised us:
LightOnOCR-2 base β **64.2% CER**
Werea-DocOCR v1 β **~8.1% CER**
Werea-DocOCR v2 β **0.15% CER** π
Evaluated on a held-out 72-page test set across 12 Turkish document types and 3 different capture conditions.
π 12 Turkish enterprise document types
π§ͺ 12,960 synthetic training pages
π± Digital + scanned + phone photos
π Tables β structured Markdown
βοΈ Full-parameter fine-tuning
π₯οΈ Trained on a single RTX 3090
It handles:
β’ e-Invoices
β’ rental contracts
β’ bank receipts
β’ payroll documents
β’ insurance policies
β’ vehicle documents
β’ official correspondence
β’ trade registry documents
β’ SGK-style tables
β’ and more.
**Model π€**
Werea-co/Werea-DocOCR-1B
**Dataset π**
Werea-co/werea-tr-doc-ocr-enterprise-v2
**Werea πΉπ·**
We're building open AI models from TΓΌrkiye.
This is just the beginning.
#HuggingFace #OCR #DocumentAI #TurkishAI #OpenSourceAI #ComputerVision