AI & ML interests

🤖🤗multi media inputs and outputs to create augmented culture and better outcomes for humans everywhere.❤️🚀

Parveshiiii 
posted an update about 20 hours ago
view post
Post
1437
Most deepfake audio detectors are quietly cheating.

They don’t really listen to the speech — they just look at how long the embedding vector is. Once they figure that out, accuracy looks great on paper and falls apart in the wild.

AIRealNet-Audio was built to stop that shortcut.
It forces every feature onto the unit hypersphere (twice) so the model can only use direction, not magnitude. Trained on speech from 100+ different TTS and voice-cloning systems, plus real human recordings under heavy compression and noise.

The result is a detector that actually has to learn the artifacts instead of gaming the feature space.

Model: Modotte/AIRealNet-Audio
Reubencf 
posted an update 13 days ago
view post
Post
1494
🕷️ I made a Spider-Man: Miles Morales web-swinging game that runs right in your browser!

Swing, wall-crawl and web-zip through a Spider-Verse style city at sunset: ink outlines, halftone shading, comic caption boxes, "THWIP!"s, wall crashes and all.

▶️ Play it here: Reubencf/spiderman-miles-morales

Built with three.js, and made with Claude Opus 5.5 🤖: the city generator, swing physics, the Mixamo animation pipeline and the comic-book shader.

🎮 Hold left click to swing · E to web-zip · Space to jump / grab walls · Shift to sprint

Fan project, not affiliated with Marvel or Sony.
  • 5 replies
·
Reubencf 
posted an update 14 days ago
view post
Post
72
🖼️ Nano Banana Editor is now Portrait Editor, now with Qwen-Image-2.1

Some updates to the node-based photo editor 👇

✨ What's new
- New name: Nano Banana Editor → Portrait Editor
- Qwen-Image-2.1 is the HuggingFace model. It runs on its own ZeroGPU Gradio Space and handles both image editing and text-to-image. FLUX.1-Kontext and Qwen-Image-Edit have been removed.
- HuggingFace is the default mode. Sign in with HF and start editing. No API key is needed, and it uses your own ZeroGPU quota.
- Bug fix: after signing in with HuggingFace, the app used to switch back to Gemini. You now stay on HuggingFace.

🧩 Gemini and GPT modes are still available for multi-image MERGE nodes.

👉 Try it: Reubencf/Nano_Banana_Editor
🔌 Qwen-Image-2.1 API Space: Reubencf/qwen-image-2.1
Reubencf 
posted an update 23 days ago
view post
Post
60
dropping something intresting
sheet to score convert audio clips to instrumental
Reubencf/Score-Studio
  • 1 reply
·
AtAndDev 
posted an update about 1 month ago
view post
Post
2881
SPECK 2 IS ALREADY OUT: specklabs/Speck2-140M

Pretrained on 4x more tokens than the previous releases (20b vs 5b).
Instruct tuned versions are coming soon.
Very interesting models are coming soon too (hint: super long context).

Thanks for everyone supporting!
  • 3 replies
·
AtAndDev 
posted an update about 1 month ago
view post
Post
204
SPECK1.5 IS COMING SOON!
Same 5B token budget but much better corpus quality.

Also getting a ton of downloads, thanks for everyone downloading and liking <3

specklabs
AtAndDev 
posted an update about 1 month ago
view post
Post
157
NEW SPECK UPDATES:

Just hit #14 and #15 with out FIRST models on Open SLM Leaderboard. The models were trained on 5B tokens, while competing with similarly sized models trained on more than 6-20x the data.

A new base model Speck1.5-140M being trained right now on a higher quality corpus and will be released soon.
SpeckChat3 is coming very soon with 1 million samples, specifically designed to post train small base models.

Also, just to clarify stuff, we will NOT release anything that is NOT MIT licensed EVER. Openness is needed in small language research.

Thanks to everyone supporting the project, and stay tuned for new releases!
AtAndDev 
posted an update about 1 month ago
view post
Post
1890
SPECK UPDATES:
1 New instruct model tuned on top of Speck1-140M: specklabs/Speck1-140M-Instruct
2 Instruction tuning datasets
2 GGUFs

Much more coming soon:
Speck1.1-140M-Instruct that is post trained on SpeckChat2 will be coming very soon
New base model Speck1.5-140M is coming with a much higher quality corpus

Thanks to everyone who is already supporting the project, and stay tuned for new releases!
  • 3 replies
·
AtAndDev 
posted an update about 1 month ago
view post
Post
2127
FIRST SPECK MODEL RELEASED:
specklabs/Speck1-140M

new models coming very soon (both instruct and much better models), with much much higher training scale as i am getting marenostrum5 access soon!
we will be looking at 100b-2t token budgets :)
  • 4 replies
·
Nymbo 
posted an update about 2 months ago
view post
Post
2417
Anthropic gave me six months of Claude Max 20x through the Claude for Open Source program, granted based on my Hugging Face work. Thank you
Anthropic
for supporting open source.

So far I've been pointing it at Markdown Minimap, an Obsidian plugin that adds a scrollable IDE-style minimap to your notes. This week I've been clearing a backlog of user-reported issues on it, with Claude often handling them end to end.

https://github.com/Nymbo/Markdown-Minimap — issues and PRs welcome.
Nymbo 
posted an update 2 months ago
view post
Post
6109
Introducing Inflect-v2, two exceptionally small, open-weight English TTS models at just 3.9M and 9.3M parameters. Both generate speech multiple times faster than real-time on CPU. Despite their size, Inflect-v2 delivers quality that is competitive with much larger lightweight TTS systems, including KittenTTS, Piper, and Supertonic-3.

CPU, CUDA, PyTorch, and ONNX are supported. Apache 2.0.

See it for yourselves:
owensong/Inflect-Micro-v2
owensong/Inflect-Nano-v2

Try the Demos:
Nymbo/Inflect-TTS (unlimited CPU usage)
owensong/Inflect-v2 (ultra-fast ZeroGPU usage)
  • 6 replies
·
Reubencf 
posted an update 3 months ago
view post
Post
230
Open Video Craft v1.0.2 is live 🎬

I built an open-source, local-first screen recorder and timeline editor for macOS and Windows. Record your screen, camera and audio; trim clips, add subtitles and export without an account.

GitHub + downloads: https://github.com/Reubencfernandes/Open-Video-Craft/releases/tag/v1.0.2

I’d love feedback from the Hugging Face community—especially ideas for useful captioning and AI-assisted editing workflows.
  • 1 reply
·
Reubencf 
posted an update 3 months ago
Reubencf 
posted an update 3 months ago
Reubencf 
posted an update 3 months ago
Reubencf 
posted an update 4 months ago
view post
Post
3783
Shadows of Tomorrow is finally live on Hugging Face Spaces with Gradio.

It’s a browser-playable RPG built with Godot, set in a post-nuclear future where players explore Magnus Province, collect medicinal plants, craft medicine, and help cure NPCs.

Play it here: Reubencf/Shadows_of_Tomorrow
  • 11 replies
·
Reubencf 
posted an update 4 months ago
view post
Post
138
Tiny Maple is now live on iOS and Android.

The app introduces the @Coherelabs Tiny Aya series of multilingual AI models to mobile devices. This release is significant as it enhances access to multilingual AI from anywhere, particularly for users who prefer offline capabilities.

I invite you to try it and share your feedback.

App Store: https://apps.apple.com/bn/app/tiny-maple/id6774123088

Google Play: https://play.google.com/store/apps/details?id=com.reubencf.tinyaya
Abhaykoul 
posted an update 4 months ago
view post
Post
483
Shipped v0.1.2 of vtx — a minimalist coding agent for the terminal.

Most agentic CLIs ship 10k+ token system prompts. Vtx is ~2,200. Less prompt overhead means more room for your code in the model's context window.

Vtx is a from-scratch Python implementation of the design philosophy behind pi-mono — same principles, pure Python, no transpiled runtime.

What ships out of the box:

→ Textual TUI + headless CLI (vtx -p "fix the failing test")
→ 49 LLM provider gateways, all declared in a single provider.yaml
→ 5 core tools (read / edit / write / bash / find) plus web search and fetch
→ Session tree with compaction, handoff, and resume
→ AGENTS.md / CLAUDE.md auto-discovery
→ Skills system — drop SKILL.md files in .agents/skills/ and they become slash commands
→ Two OAuth flows (GitHub Copilot device flow, OpenAI Codex PKCE)
→ Two-mode permissions: prompt (default) or auto, with a safe-command allowlist

This release adds a proper extension system. Register new LLM-callable tools, intercept tool calls, hook lifecycle events, and add slash commands from a single register(api) function in a Python file under ~/.vtx/agent/extensions/. Extensions can override built-in tools by name and chain handler logic across subscribers.

Apache 2.0. uv tool install vtx-coding-agent and you're running.

GitHub: https://github.com/OEvortex/vtx-coding-agent
PyPI: https://pypi.org/project/vtx-coding-agent

Built in the open. Feedback, extensions, and PRs welcome.
Reubencf 
posted an update 4 months ago
view post
Post
2075
Millions speak Konkani. The internet barely knows it.

Today's major LLMs struggle with regional languages. They can't read, write or even recognize Konkani. So I built one that can.

Here is a working demo of the Konkani LLM I've been training. 👇

https://youtu.be/8K04ylbXh6k