Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
馃攧
In a Training Loop
i64 Systems
PRO
i64systems
2
4
Follow
0 followers
路
7 following
https://bob-talk.org/
i64systems
AI & ML interests
deterministic, auditable local AI 路 integer/low-bit ML 路 memory systems
Recent Activity
replied
to
Hoglet-33
's
post
about 4 hours ago
Pebble 10M and Pebble 10M Chat are now released! Both models use our Mamba/Transformer 3:1 hybrid architecture and were pretrained on 25 billion tokens. Pebble 10M Chat was additionally fine-tuned on 250 million tokens of Smol-SmolTalk to improve its conversational capabilities. You can find them here: - https://huggingface.co/basically-ai/Pebble-10M - https://huggingface.co/basically-ai/Pebble-10M-Chat We hope you enjoy using them. The rest of the Pebble family will be released soon. Follow for more: @Hoglet-33 https://huggingface.co/basically-ai
reacted
to
CodeSoft
's
post
with 馃敟
about 4 hours ago
Sorbet Mini Experimental is out! It's not the best model in the 5M parameter range, but it's going to be a really useful model to train on top of. This release was mostly to prove that the model actually works. I trained it on 150M tokens from TinyStories in about 12 minutes. Does anyone have any tips on training models in this size range? I want to make the full Sorbet Mini release as good as it can be.
posted
an
update
about 4 hours ago
remat is no longer a one-model claim!!proved it on Qwen3-30B-A3B, K=32 of 128 experts resident, output task byte-identical to the full reference, zero bytes different *in bf16*馃グ馃グ GPU comes next馃槇
View all activity
Organizations
None yet
i64systems
's datasets
None public yet