注册并分享邀请链接,可获得视频播放与邀请奖励。

Superman (@thesupermanmx) “China open-sourced a peanut-sized OCR that parses entire 100-page PDFs in one sh” — TopicDigg

Superman 的个人资料封面
Superman 的头像
Superman
@thesupermanmx
AI & Neuroscience
加入 May 2026
0 正在关注    1.8K 粉丝
China open-sourced a peanut-sized OCR that parses entire 100-page PDFs in one shot.. It's called Unlimited-OCR. Only 3B params. Runs locally. Every other OCR tool chops your doc into pages and loses the thread. this one reads the whole thing in a single pass. → One-shot "long-horizon" parsing (32K context window) → Multilingual, out of the box → 93% on the standard parsing benchmark (+6 over baseline) → <0.11 error rate past 40 pages → Runs 100% locally on your own hardware → Works with Transformers, vLLM, SGLang, Docker, Ollama, llama.cpp Traditional cloud OCR (Textract, Google Vision, Azure Doc Intelligence) costs $1.50–$15 per 1,000 pages. This runs on your machine. For free. Forever. Baidu built it explicitly to push DeepSeek-OCR one step further. Already at 1.9M downloads on Hugging Face and most people have no idea it exists yet. 100% open source.
显示更多
0
25
2.8K
288
转发到社区