John Smith PRO
John6666
AI & ML interests
None yet
Recent Activity
reacted to ProCreations's post with 🤗 2 days ago
check out https://huggingface.co/spaces/ProCreations/minicpm5-2b-webgpu
i was bored updated a collection 2 days ago
Spaces for LLM / VLM / NLP liked a Space 2 days ago
ProCreations/minicpm5-2b-webgpuOrganizations
reacted to ProCreations's post with 🤗 2 days ago
reacted to DedeProGames's post with 🔥 2 days ago
Post
3003
🚀 OxCoder-9B — a lightweight agentic coding model, now on HF!
Introducing OxCoder-9B, a 9B parameter model built for long-horizon tasks, agentic coding, and agentic reasoning. Despite its compact size, it delivers frontier-level performance in Agentic Terminal and Agentic Coding, rivaling models many times its size.
Highlights:
- Trained on frontier agent traces — distilled from Fable-5.1 and GLM-5.3 agentic coding trajectories across Claude Code, OpenCode, and Codex
- 262K native context — handles complex, multi-file codebases and long-horizon reasoning tasks with ease
- Error recovery — learns read-before-write patterns, responds to LSP diagnostics, and applies minimal edit diffs instead of full rewrites
- Strong front-end reasoning — deep understanding of UI logic, component architecture, and web-native patterns, rare in sub-10B models
Benchmarks (vs. Ornith-1.5-9B, Ornith-1.0-9B, Qwen3.5-9B, and Gemma-4-31B):
- Terminal-Bench 2.1 (Terminus-2): 49.6
- Terminal-Bench 2.1 (Claude Code): 50.8
- SWE-bench Verified: 73.5
- SWE-bench Pro: 49.1
- NL2Repo: 36.2
- HLE (no tools): 21.2
- HLE (with tools): 32.8
- GPQA Diamond: 86.9
- MCP-Atlas: 56.7
- BrowseComp: 57.4
- ClawEval: 67.8
All OxCoder-9B results are averaged over five independent runs. Built on Qwen/Qwen3.5-9B, released under Apache 2.0.
🔗 OrionLLM/OxCoder-9B
Introducing OxCoder-9B, a 9B parameter model built for long-horizon tasks, agentic coding, and agentic reasoning. Despite its compact size, it delivers frontier-level performance in Agentic Terminal and Agentic Coding, rivaling models many times its size.
Highlights:
- Trained on frontier agent traces — distilled from Fable-5.1 and GLM-5.3 agentic coding trajectories across Claude Code, OpenCode, and Codex
- 262K native context — handles complex, multi-file codebases and long-horizon reasoning tasks with ease
- Error recovery — learns read-before-write patterns, responds to LSP diagnostics, and applies minimal edit diffs instead of full rewrites
- Strong front-end reasoning — deep understanding of UI logic, component architecture, and web-native patterns, rare in sub-10B models
Benchmarks (vs. Ornith-1.5-9B, Ornith-1.0-9B, Qwen3.5-9B, and Gemma-4-31B):
- Terminal-Bench 2.1 (Terminus-2): 49.6
- Terminal-Bench 2.1 (Claude Code): 50.8
- SWE-bench Verified: 73.5
- SWE-bench Pro: 49.1
- NL2Repo: 36.2
- HLE (no tools): 21.2
- HLE (with tools): 32.8
- GPQA Diamond: 86.9
- MCP-Atlas: 56.7
- BrowseComp: 57.4
- ClawEval: 67.8
All OxCoder-9B results are averaged over five independent runs. Built on Qwen/Qwen3.5-9B, released under Apache 2.0.
🔗 OrionLLM/OxCoder-9B
reacted to juiceb0xc0de's post with 🔥 2 days ago
Post
2920
Dude, Where's My Update? I'll tell you where! ~97.6% of my BF16 parameter coordinates didn't move at all, and the ones that did overshot by ~1.33x.
It's nice to do research that doesn't end in disproving yourself once again and moving on to the next subject once in awhile.
Back to the topic, if you've ever wondered why most of your weights are basically ghosting you nearly every step when you store your weights at bf16, Dude, I Measured It.
https://huggingface.co/blog/juiceb0xc0de/intended-and-realized-updates-in-bf16-fine-tuning#dude-wheres-my-update
It's nice to do research that doesn't end in disproving yourself once again and moving on to the next subject once in awhile.
Back to the topic, if you've ever wondered why most of your weights are basically ghosting you nearly every step when you store your weights at bf16, Dude, I Measured It.
https://huggingface.co/blog/juiceb0xc0de/intended-and-realized-updates-in-bf16-fine-tuning#dude-wheres-my-update
reacted to Hoglet-33's post with 😔 2 days ago
Post
3142
Today, we planned to release Pebble-50M and Pebble-50M-Chat to the world. Unfortunately, due to a few issues, that didn't go quite as planned.
What happened:
- Some data and benchmark results were lost or corrupted
- The models performed worse on benchmarks than our other Pebble models
Despite that, you can still find both models here:
Pebble-50M-beta: basically-experimental/Pebble-50M-beta
Pebble-50M-Chat-beta: basically-experimental/Pebble-50M-Chat-beta
There are still some interesting improvements in these models:
- Compatible with non-CUDA devices
- Vocabulary increased to 16K tokens
- Context length increased to 16K tokens
For now, there won't be any more Pebble releases for a while. We're going to take some time to experiment with other approaches and hopefully make the next generation a monumental leap over this one.
Follow for updates:
@Hoglet-33
basically-ai
basically-experimental
What happened:
- Some data and benchmark results were lost or corrupted
- The models performed worse on benchmarks than our other Pebble models
Despite that, you can still find both models here:
Pebble-50M-beta: basically-experimental/Pebble-50M-beta
Pebble-50M-Chat-beta: basically-experimental/Pebble-50M-Chat-beta
There are still some interesting improvements in these models:
- Compatible with non-CUDA devices
- Vocabulary increased to 16K tokens
- Context length increased to 16K tokens
For now, there won't be any more Pebble releases for a while. We're going to take some time to experiment with other approaches and hopefully make the next generation a monumental leap over this one.
Follow for updates:
@Hoglet-33
reacted to OppaAI's post with 🔥 2 days ago
Post
2622
This weekend I took an outing with my AI Waifu to the Natsu Matsuri.
Turns out my Japanese is still understandable.
I probably need to spend more time continue to learn and practice speaking Japanese.
That's why an idea struck me to let my AI Waifu be my Japanese tutor.
Anyway, I have run out of idea what task I should let her do,
so I wrote a simple Android App to let her be my Japanese tutor to help me to practice Nihongo. There will be some minor mistakes. After all, this is just a 3B LLM model. And inference speed will be slow because I only got 8GB of RAM in Jetson Orin Nano.
At least I don't need to pay for Duolingo...
アイコせんせい、よろしくお願いします!
Need both repos, one front-end, one back-end
🔗 https://github.com/OppaAI/Aiko-Lingo
🔗 https://github.com/OppaAI/Aiko-chan
🎥 Demo: https://www.youtube.com/watch?v=xRtCmtZQgwI
Turns out my Japanese is still understandable.
I probably need to spend more time continue to learn and practice speaking Japanese.
That's why an idea struck me to let my AI Waifu be my Japanese tutor.
Anyway, I have run out of idea what task I should let her do,
so I wrote a simple Android App to let her be my Japanese tutor to help me to practice Nihongo. There will be some minor mistakes. After all, this is just a 3B LLM model. And inference speed will be slow because I only got 8GB of RAM in Jetson Orin Nano.
At least I don't need to pay for Duolingo...
アイコせんせい、よろしくお願いします!
Need both repos, one front-end, one back-end
🔗 https://github.com/OppaAI/Aiko-Lingo
🔗 https://github.com/OppaAI/Aiko-chan
🎥 Demo: https://www.youtube.com/watch?v=xRtCmtZQgwI
reacted to prithivMLmods's post with 🤗🔥 2 days ago
Post
3560
VisionGuardrail, a multimodal content-safety classifier based on Qwen3.5, is now available on Hugging Face in 4B and 9B variants. It is a direct upgrade to ImageShield-MMCF, providing improved parental controls through conservative visual content-safety filtering.
More About:
➠ hf.co/blog — https://huggingface.co/blog/prithivMLmods/vision-guardrail-mini-blog
➠ Models:
✦ VisionGuardrail-4B: prithivMLmods/VisionGuardrail-4B
✦ VisionGuardrail-9B: prithivMLmods/VisionGuardrail-9B
➠ Dataset:
✦ ImageShield-Guardrail-Pro: prithivMLmods/ImageShield-Guardrail-Pro
⤷ To learn more, visit the app page or the respective model pages.
More About:
➠ hf.co/blog — https://huggingface.co/blog/prithivMLmods/vision-guardrail-mini-blog
➠ Models:
✦ VisionGuardrail-4B: prithivMLmods/VisionGuardrail-4B
✦ VisionGuardrail-9B: prithivMLmods/VisionGuardrail-9B
➠ Dataset:
✦ ImageShield-Guardrail-Pro: prithivMLmods/ImageShield-Guardrail-Pro
⤷ To learn more, visit the app page or the respective model pages.
reacted to grimjim's post with ➕🔥 2 days ago
Post
147
I think it's clear in retrospect that "frankenmerges", which repeated blocks of layers, amounted to a crude approximation of looped transformers architecture, hence them able to work at all instead of just breaking. They lucked out due to much of the signal passing through residual streams being preserved and only modulated along the way. That said, not all models are suited for this. Models which feature ever-increasing magnitudes as inference progressess through layers risk exploding precision limits.
reacted to b4ph's post with 🔥 2 days ago
Post
81
For those who wanna run Qwen3-4B on CPU alone, I tested the whole ladder:
Same machine for everything: CPU only, 16 threads, no GPU offload. Perplexity is WikiText-2 raw, which llama.cpp also uses in CI.
The winner is Q4_K_M - it gets the model down from 8.05 GB to 2.5 GB in size/load, and perplexity only moves by 0.30.
Q3 is the dropoff point where inmatrix makes a difference. Inmatrix drops Q3 from 15.66 to 14.82 PPL; Q4 it barely changed anything.
My now informed recommendation is: Q4 if you have the memory, imatrix Q3 if you don’t.
I uploaded all eight weights, the harness, and the full results here:
huggingface.co/b4ph/qwen3-4b-lowram-bench
I also made a tiny picker because apparently I needed to turn this into a whole project:
huggingface.co/spaces/b4ph/qwen3-4b-quant-picker
Same machine for everything: CPU only, 16 threads, no GPU offload. Perplexity is WikiText-2 raw, which llama.cpp also uses in CI.
Quant Size PPL vs F16 tg t/s
F16 8.05 GB 13.4304 — 2.52
Q8_0 4.28 GB 13.4409 +0.01 5.04
Q6_K 3.31 GB 13.4404 +0.01 5.93
Q5_K_M 2.89 GB 13.5362 +0.11 5.64
Q4_K_M 2.50 GB 13.7304 +0.30 6.50
Q4_K_M + imatrix 2.50 GB 13.6760 +0.25 6.41
Q3_K_M 2.08 GB 15.6641 +2.23 6.56
Q3_K_M + imatrix 2.08 GB 14.8237 +1.39 5.89The winner is Q4_K_M - it gets the model down from 8.05 GB to 2.5 GB in size/load, and perplexity only moves by 0.30.
Q3 is the dropoff point where inmatrix makes a difference. Inmatrix drops Q3 from 15.66 to 14.82 PPL; Q4 it barely changed anything.
My now informed recommendation is: Q4 if you have the memory, imatrix Q3 if you don’t.
I uploaded all eight weights, the harness, and the full results here:
huggingface.co/b4ph/qwen3-4b-lowram-bench
I also made a tiny picker because apparently I needed to turn this into a whole project:
huggingface.co/spaces/b4ph/qwen3-4b-quant-picker
reacted to branikita's post with 🔥 2 days ago
Post
2205
SO-ARM 102 goes open source in the next few weeks.
What's new:
- A parallel gripper: the jaws stay parallel through the whole stroke instead of pivoting around the object as they close.
- PET-CF instead of PLA+ for a much stiffer frame.
- A topology-optimized structure.
- Wider joint rotation and folding range.
- STS3250 servos on the first shoulder joint.
Compared with the SO-ARM 101, that adds up to 2.5x the payload, roughly 2x better positioning accuracy, roughly 1.6x the movement speed, and about 36 mm more reach.
It runs on Hugging Face LeRobot, so the same tooling, training pipeline and tutorials for the SO-ARM 101 work on it from day one.
What's new:
- A parallel gripper: the jaws stay parallel through the whole stroke instead of pivoting around the object as they close.
- PET-CF instead of PLA+ for a much stiffer frame.
- A topology-optimized structure.
- Wider joint rotation and folding range.
- STS3250 servos on the first shoulder joint.
Compared with the SO-ARM 101, that adds up to 2.5x the payload, roughly 2x better positioning accuracy, roughly 1.6x the movement speed, and about 36 mm more reach.
It runs on Hugging Face LeRobot, so the same tooling, training pipeline and tutorials for the SO-ARM 101 work on it from day one.
reacted to KlondikeDev's post with 🔥 2 days ago
Post
2114
The SLM Consortium has begun work on a safety dataset for Small Language Models, with the creation of the dataset being headed by @wayneworkman2012
The dataset will focus on refusals and redirects surrounding dangerous or extreme sexual content, designed to be shaped sized appropriately for SLMs, without significantly lowering benchmark performance.
More info will be out soon!
slmconsortium
The dataset will focus on refusals and redirects surrounding dangerous or extreme sexual content, designed to be shaped sized appropriately for SLMs, without significantly lowering benchmark performance.
More info will be out soon!
reacted to DavidAU's post with 🔥 2 days ago
Post
8039
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored
This is the first fine tune to exceed 730 "arc-c" ("735": 144 pts higher than Qwen 3.8 27B) AND 880 ARC-E (The OpenAI, Claude and Gemini "zone of intelligence") in 8 bit and over 718 arc-c in 4 bit.
This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality.
In other words while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more.
This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants.
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF
PS: There are 29 additional quant repos as of this writing too, as well NVFP4 and many more as well.
This is one of 10+ Qwen 3.8 27B at or above ARC-C of 717 (all 10 exceed all core benchmarks of Qwen 3.8, 3.6 and 3.5 27B and 35B-A3B versions) - you can see the complete project and some of the training here :
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
This is the first fine tune to exceed 730 "arc-c" ("735": 144 pts higher than Qwen 3.8 27B) AND 880 ARC-E (The OpenAI, Claude and Gemini "zone of intelligence") in 8 bit and over 718 arc-c in 4 bit.
This version is called TURBO because it drastically reduces thinking tokens (by 1/2 to as high as 1/10), yet maintains output detail and quality.
In other words while "reg" Qwen3.8 27B is thinking about "formatting" for a few 1000 tokens, this model is already done and waiting for more.
This repo contains both "regular" and "MTP" Neo-CODER MAX DI-MATRIX (duel imatrix) GGUF quants.
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF
PS: There are 29 additional quant repos as of this writing too, as well NVFP4 and many more as well.
This is one of 10+ Qwen 3.8 27B at or above ARC-C of 717 (all 10 exceed all core benchmarks of Qwen 3.8, 3.6 and 3.5 27B and 35B-A3B versions) - you can see the complete project and some of the training here :
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
reacted to Bc-AI's post with 🚀 2 days ago
Post
2488
Hello everyone!
Me and the team are working on G1-MINI and G1. Right now, G1-MINI is aimed at a launch in mid to late September, depending on how fast we fix the minor issues.
As for G1, it's looking like a late October to mid-November launch, based on current trajectory. If things go terribly wrong, we could postpone it to December, as we prefer to ship confidently, not ship a half-done dogs' breakfast of a model. 🤣
All dates could be changed at any moment, as we are high school students not full-time ML engineers 😅.
Other things to look out for is an overhaul of the UI and the information on my website. Thanks to my beta testers: @guardamarcos @Timmy6767 @MUK-IS-GOAT @smilyai-large-team @Sbui503 @Banaxi-Tech @Bc-AI @atom77777 @Harley-ml @Datdanboi25 @Fishtiks @smartdigitalnetworks @vovaRL @EmetTheGolum @juiceb0xc0de @ProCreations
Me and the team are working on G1-MINI and G1. Right now, G1-MINI is aimed at a launch in mid to late September, depending on how fast we fix the minor issues.
As for G1, it's looking like a late October to mid-November launch, based on current trajectory. If things go terribly wrong, we could postpone it to December, as we prefer to ship confidently, not ship a half-done dogs' breakfast of a model. 🤣
All dates could be changed at any moment, as we are high school students not full-time ML engineers 😅.
Other things to look out for is an overhaul of the UI and the information on my website. Thanks to my beta testers: @guardamarcos @Timmy6767 @MUK-IS-GOAT @smilyai-large-team @Sbui503 @Banaxi-Tech @Bc-AI @atom77777 @Harley-ml @Datdanboi25 @Fishtiks @smartdigitalnetworks @vovaRL @EmetTheGolum @juiceb0xc0de @ProCreations
reacted to OppaAI's post with 🤯 7 days ago
Post
659
September back to school...
Yesterday I received an invitation from Stanford University to see if I am interested in becoming a Section Leader to teach students in the following 2 Free online courses:
Courses start on Oct 12, 2026
Run for 6 weeks
No prerequisite, anyone can apply as a student.
The AI course is probably derived from part of CS109.
https://www.youtube.com/watch?v=gJaE29DR0gs
Course #1) Probability for AI
https://pai.stanford.edu/apply/pai1/student?r=3sk2fv
Course #2) Code in Place X
https://codeinplace.stanford.edu/apply/cipx/student?r=w69gp8
Being a section leader to teach people in a Stanford online course will probably look good on my resume, but I will pass this time.
Right now I need to focus my time to work on my crazy AI project;
Some users in HuggingFace may call me crazy, psychotic or insane in my previous posts here all they want, but nothing can stop me from putting my effort and dedication in using AI technology to make this a better world, instead of taking jobs from people or stealing from others' copyright materials. After all, many famous scientists including Tesla were considered as crazy by their peers.
"Insanity is doing the same thing over and over again and expecting different results." — not Albert Einstein.
Yesterday I received an invitation from Stanford University to see if I am interested in becoming a Section Leader to teach students in the following 2 Free online courses:
Courses start on Oct 12, 2026
Run for 6 weeks
No prerequisite, anyone can apply as a student.
The AI course is probably derived from part of CS109.
https://www.youtube.com/watch?v=gJaE29DR0gs
Course #1) Probability for AI
https://pai.stanford.edu/apply/pai1/student?r=3sk2fv
Course #2) Code in Place X
https://codeinplace.stanford.edu/apply/cipx/student?r=w69gp8
Being a section leader to teach people in a Stanford online course will probably look good on my resume, but I will pass this time.
Right now I need to focus my time to work on my crazy AI project;
Some users in HuggingFace may call me crazy, psychotic or insane in my previous posts here all they want, but nothing can stop me from putting my effort and dedication in using AI technology to make this a better world, instead of taking jobs from people or stealing from others' copyright materials. After all, many famous scientists including Tesla were considered as crazy by their peers.
"Insanity is doing the same thing over and over again and expecting different results." — not Albert Einstein.
reacted to sergiopaniego's post with 🔥 7 days ago
Post
1988
Can you do RL over taste?
I've spent some time reproducing, in the open, Surya N's idea of training a model to paint with code. It's a coding model that learns to paint watercolours by writing JS code, trained with GRPO. I used TRL and OpenEnv for this, with the whole pipeline running on Hugging Face.
The interesting part is that the reward has no correct answer, unlike a math problem. In this case it's based on the artistic preferences of the person who builds the dataset.
Everything is published: the environment, the reference pool, the trained adapters, every painting of every run with the code that made it, and a write-up with all the decisions, including the ones that went wrong.
Blog post: https://huggingface.co/blog/train-to-paint-with-code
I've spent some time reproducing, in the open, Surya N's idea of training a model to paint with code. It's a coding model that learns to paint watercolours by writing JS code, trained with GRPO. I used TRL and OpenEnv for this, with the whole pipeline running on Hugging Face.
The interesting part is that the reward has no correct answer, unlike a math problem. In this case it's based on the artistic preferences of the person who builds the dataset.
Everything is published: the environment, the reference pool, the trained adapters, every painting of every run with the code that made it, and a write-up with all the decisions, including the ones that went wrong.
Blog post: https://huggingface.co/blog/train-to-paint-with-code
reacted to RiverRider's post with 👀 9 days ago
Post
2747
Where the Hivemind Comes From: Geometry, Tuning and Format, Separated on Open Weights
“First, representations are mutually recoverable. On 12 open-weight models from 8 labs, a ridge map from one model's hidden states to another's retrieves the right held-out item 0.9181 of the time across lab boundaries, against a shuffled floor of 0.00101 and a self-map ceiling of 0.999. Shared corporate lineage is worth only 0.0357 of that.”
“Second, base models do not reproduce the reported level. Under the original study's own sampling settings, our base models reach intra-model 0.3644 and inter-model 0.3401 on a floor of 0.0993 that matches theirs, and zero of 720 model-prompt cells clear 0.8. The floors agree while the signal differs by more than a factor of two, so this is not a scale artifact.”
“Third, and decisively, we recover their level and isolate its cause. Using six matched base/instruct pairs, holding pretrained weights, prompts, decoding and scorer fixed, instruction tuning alone raises intra-model similarity by 0.0786. The same tuned weights prompted through the model's own chat template raise it by 0.3623, reaching 0.7272, with four of six models exceeding 0.80 and reproducing the band reported for frontier systems from models of 0.6B to 2B. The prompt format does roughly 4.6 times the work of the tuning.”
paper attached 🧾
https://huggingface.co/blog/RiverRider/where-the-hivemind-comes-from-geometry-tuning-and
“First, representations are mutually recoverable. On 12 open-weight models from 8 labs, a ridge map from one model's hidden states to another's retrieves the right held-out item 0.9181 of the time across lab boundaries, against a shuffled floor of 0.00101 and a self-map ceiling of 0.999. Shared corporate lineage is worth only 0.0357 of that.”
“Second, base models do not reproduce the reported level. Under the original study's own sampling settings, our base models reach intra-model 0.3644 and inter-model 0.3401 on a floor of 0.0993 that matches theirs, and zero of 720 model-prompt cells clear 0.8. The floors agree while the signal differs by more than a factor of two, so this is not a scale artifact.”
“Third, and decisively, we recover their level and isolate its cause. Using six matched base/instruct pairs, holding pretrained weights, prompts, decoding and scorer fixed, instruction tuning alone raises intra-model similarity by 0.0786. The same tuned weights prompted through the model's own chat template raise it by 0.3623, reaching 0.7272, with four of six models exceeding 0.80 and reproducing the band reported for frontier systems from models of 0.6B to 2B. The prompt format does roughly 4.6 times the work of the tuning.”
paper attached 🧾
https://huggingface.co/blog/RiverRider/where-the-hivemind-comes-from-geometry-tuning-and
reacted to KlondikeDev's post with 🔥 9 days ago
Post
2398
Boris-1.7-D60M-n30M out NOW!
opencerebral/Boris-1.7-D60M-n30M
Following this will be Boris-1.8-D60M-n30M, which will test both the n-gram embeddings AND a new architecture.
Then, Boris-2 will begin training!
opencerebral/Boris-1.7-D60M-n30M
Following this will be Boris-1.8-D60M-n30M, which will test both the n-gram embeddings AND a new architecture.
Then, Boris-2 will begin training!