注册并分享邀请链接,可获得视频播放与邀请奖励。

Elastic (@elastic) “You've got 10k slides, scans, and screenshots to search. Your first instinct is” — TopicDigg

Elastic 的个人资料封面
Elastic 的头像
Elastic
@elastic
Where developers learn, build, and share. Your source for hands-on demos, cheat sheets, explainers and more.
加入 October 2009
183 正在关注    65.7K 粉丝
You've got 10k slides, scans, and screenshots to search. Your first instinct is to throw a VLM at it. But a VLM reads one image at a time. Running it across your whole corpus on every query doesn’t scale. Split the pipeline instead. jina-clip embeds every image into a vector once. At query time, the same model embeds your question and does a similarity lookup. Jina-VLM only touches the matched slides: a handful of images, not thousands. Build the index once, then retrieval is instant. The VLM only reasons where it matters.
显示更多
0
8
112
11
转发到社区