注册并分享邀请链接,可获得视频播放与邀请奖励。

Superman 的个人资料封面
Superman 的头像

Superman (@thesupermanmx)

@thesupermanmx
AI & Neuroscience
0 正在关注    1.8K 粉丝
China open-sourced a peanut-sized OCR that parses entire 100-page PDFs in one shot.. It's called Unlimited-OCR. Only 3B params. Runs locally. Every other OCR tool chops your doc into pages and loses the thread. this one reads the whole thing in a single pass. → One-shot "long-horizon" parsing (32K context window) → Multilingual, out of the box → 93% on the standard parsing benchmark (+6 over baseline) → <0.11 error rate past 40 pages → Runs 100% locally on your own hardware → Works with Transformers, vLLM, SGLang, Docker, Ollama, llama.cpp Traditional cloud OCR (Textract, Google Vision, Azure Doc Intelligence) costs $1.50–$15 per 1,000 pages. This runs on your machine. For free. Forever. Baidu built it explicitly to push DeepSeek-OCR one step further. Already at 1.9M downloads on Hugging Face and most people have no idea it exists yet. 100% open source.
显示更多
0
25
2.8K
288
转发到社区
someone open-sourced a tool that turns any phone video into a full 3D world in real time. you record a walk around your house, drop the mp4 in, and the entire scene comes back as a walkable 996k-point world at 60fps on your local GPU. no cloud. no subscription. runs 100% locally. 100% open source.
显示更多
0
34
1.9K
176
转发到社区
China open-sourced a model that reconstructs any scene in 3D from a regular video, in real-time. one camera. no LiDAR. 10,000+ frames without falling apart. just walk around with your camera and watch the entire world get rebuilt in 3D at 20 fps. → runs at ~20 FPS on a single GPU → Stable over 10,000+ frames → Beats optimization-based methods on benchmarks → Works on drone footage, driving videos, indoor walkthroughs 100% open source.
显示更多
0
32
2.1K
258
转发到社区