Peter Szemraj PRO
AI & ML interests
Recent Activity
Organizations
Hey, thanks for following through on that! I see the results on the leaderboard just now.
I need to sift through the different metrics to make sense of things because falconOCR is shockingly bad at least from the main ranking, but maybe it's an out-of-distribution thing, I'm not sure. Or maybe it's a case where it needs its little layout classifier model (or vice versa), which sometimes I find helps, and other times is terrible for quality.
- The being worse than Tesseract is interesting to say the least. Don't really have anything more to tell you yet..
One thing I did want to request, if possible, would be this new nemotron-parse 2.0 from Nvidia, which is very interesting simply from a license and origin perspective in that it is not a fine-tune of some Chinese VLM afaik, but a from-scratch architecture
Meet North Micro Vision: A 2.4B Native-Resolution Vision-Language Model
Full-bandwidth transformer
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
Love this work and great job. Very excited to see a practical OCR task putting models to the test and rankings.
One thing though, I was hoping to see https://hf.co/tiiuae/Falcon-OCR - could it be added and or was there a reason why it was excluded? As it punches above its weight given param count, could be potentially a viable low-cost option