注册并分享邀请链接,可获得视频播放与邀请奖励。

与「DeepResearch」相关的搜索结果

DeepResearch 贴吧
一个关键词就是一个贴吧,路径全站唯一。
创建贴吧
用户
未找到
包含 DeepResearch 的内容
传统的 Deep Research 已经卷到头了。 能写出一份漂亮的总结报告 ≠ 能把真实的复杂任务干完。 我拿几十万字的《红楼梦》原著,给 Apodex 1.1 在线工作台,出了个极其变态的任务。 统计 20 个核心人物的 → 出场次数 → 出场回数 → 每个人第一次出场时的原句 最后还要整理成一张可以直接下载的完整表格。 整个完成任务的过程,像是直接「雇佣一个 AI 数据团队」。 上传文件后,它自己开始拆任务。 一个 Agent 负责解析原著,另一个独立分析,多个任务并行推进; 右侧 Task Board 会实时告诉你现在做到哪一步。 更有意思的是,任务跑到一半,我突然改需求: “只分析前 80 回,后面的不要了。” 以前遇到这种情况,AI 很可能重新来一遍。 Apodex 直接保留已经完成的成果,只重规划受影响的部分。 更关键的是,它不是做完就交卷。 交付前,又调起独立核验 Agent,把人物统计、出场回数和 900+ 条出场原句重新检查一遍。 最后交付给我的是: 可下载的结构化表格 + 完整分析结果 + 口径说明。 这可能才是下一代 Deep Research 真正值得关注的变化: 从“帮你生成一份报告”,变成“接管一项复杂任务,并把它做完”。 目前,Apodex 1.1 Web 端已经正式上线。 🎁 注册即送 credits,强烈建议立刻上传个复杂文件自己跑跑看: 🌐 没想到,更炸裂的是,Apodex 居然开源了 开源模型: Apodex 1.1 mini,35B,开放模型权重,支持本地部署。 开源框架: FrontierAgent,Agent 执行 Harness,支持 ReAct 单 Agent + Multi-Agent Team。 本地运行: 支持 macOS / Linux,无需强制依赖 Docker。 组合能力: Apodex 1.1 mini + FrontierAgent,可在本地运行完整 Agent 执行流程。 💻 如果你是开发者,这里有开源 Agent 框架,欢迎顺手点个 ⭐: 👉 🤗 想本地自己跑模型的看这里👇 #Apodex# #DeepResearch#
显示更多
0
20
27
3
转发到社区
llm_wiki v0.4.25 更新发布🎉🎉🎉 主要是集成了 Firecrawl 刚推出的免费搜索,现在 DeepResearch 不再需要付费啦,也不用 API Key,直接启用 Firecrawl就可以了。 另外也修复大量 issues 中提到的缺陷和 Bug
显示更多
用 open-slide 花 4 个小时做了 40+页的 PPT,感觉蛮好的。 先 DeepResearch - 生成课程大纲 - OpenSlide 确定风格 - 生成 slide - 部分图片让 gpt-image-2 生成 - 人工校对和修改。
显示更多
0
11
309
46
转发到社区
闲鱼自动上架这个问题,Claude Opus 4.7 和 ChatGPT Curr 都回答错误。只有 Gemini 搞定了: “ 我在闲鱼上卖虚拟资料。目前卖完一份,资料就自动下架。于是,我要手工再填写资料,再上架。效率非常差。 有什么办法,可以做到卖完资料后,原先的虚拟资料自动上架。调研下现有的做法,按照我是个闲鱼小白,用什么方法最合适 ” Claude 和 ChatGPT ,困在我以往 RPA 问答记录的茧房里,给我的都是写脚本,做插件自动化。 而 Gemini, 平时我几乎不与它切磋编程,只用 DeepResearch, 它少了很多我的历史痕迹,反而给出最朴实的做法,找到下架的商品,再重新上架,搞定 !! 看来在各个 llm 之间同步对话记录,容易污染他们的脑子
显示更多
0
10
43
1
转发到社区
Agent 的动手能力,已经在过去一年经历了显著的跃迁。它不再只是会“聊天”的模型,而是可以真正去动手、去执行复杂任务的智能体。那么现在它能做到什么?已经能解决多复杂的软件工程问题?又该如何在社区里找到最强框架并复用到自己的项目?下面是几条更实用的思路。 要评估一个 Agent 的动手能力,无论它是单一的 LLM,还是 LLM 加上外部工具的工程实现,最终都要回到数据集上。因为数据集定义了“考试题目”,而 benchmark 决定了“评分标准”。目前能全面评估 Agent 工程执行力的两个核心数据集,一个是 OpenAI 的 SWE-bench(software engineering-bench),另一个是 THUDM 提供的 Agent-Bench。前者聚焦真实软件仓库的 bug 修复与功能实现,是“AI 程序员”的试炼场;后者覆盖更广,从软件、操作系统、网络、推理、工具使用到多模态交互,是对 Agent 通用智能和工具操作能力的系统化测评。 什么才算一个好的 Agent,还得回到问题域上看。SWE-bench 的目标是让 Agent 能像程序员一样理解代码、修补缺陷、通过单测;而 Agent-Bench 则像是在考察一个“通才型工程助理”,既要能读懂文档、用命令行、写代码,又要能跨工具协作、执行复杂任务链。前者考工程深度,后者考任务广度。这两个维度,几乎定义了 Agent 的“手工能力边界”。 理解这个边界,还得区分哪些问题是 LLM 本身可以解决的,哪些必须依赖外部工具。从大模型的演进来看,许多原本需要显式工具链配合的能力,正在逐步被“内化”进模型本体。Chain of Thought 已经演化为参数化的推理能力(Reasoning),知识图谱的结构化记忆也被吸收到模型的参数知识(Parametric Knowledge)中。而最近阿里开源的 Tongyi DeepResearch,正是这种趋势的最新代表:它通过强化学习(RL)直接训练模型具备“研究型行为”,主动检索、阅读、摘要、再检索,在真实网络环境中形成自我迭代的探索闭环。 要找到好用的 Agent 框架或最佳实践,最直接的办法就是去看各大数据集的打榜记录,榜单上往往能看到社区最新的开源成果与架构思路。SWE-bench 有一个官方 leaderboard,目前得分最高的方案往往来自一些 AI IDE 工具,比如 TRAE、Augment Code 等,因为 SWE 要解决的软件工程问题,和 AI IDE 的目标几乎完全重叠,它们都想让模型在真实项目里“动手干活”。在这些榜单里,你可以找到大量可以直接复用的开源实现,例如 github@augmentcode/augment-swebench-agent、github@ByteDance-Seed/Seed-Coder 等。 如果你正好在做相关方向的工作,不妨先采取“拿来主义”。SWE-bench 上最好的模型得分已经达到了 78.8 分,意味着这些 Agent 已经能解决绝大多数真实工程问题。要知道,在 2024 年三月,这个榜单的最高分还只有 12.4。短短一年,从“会写代码”到“能维护项目”,AI 的动手能力,已经跨过了一个关键分水岭。
显示更多
0
4
101
17
转发到社区
I was given early access to Grok 3 earlier today, making me I think one of the first few who could run a quick vibe check. Thinking ✅ First, Grok 3 clearly has an around state of the art thinking model ("Think" button) and did great out of the box on my Settler's of Catan question: "Create a board game webpage showing a hex grid, just like in the game Settlers of Catan. Each hex grid is numbered from 1..N, where N is the total number of hex tiles. Make it generic, so one can change the number of "rings" using a slider. For example in Catan the radius is 3 hexes. Single html page please." Few models get this right reliably. The top OpenAI thinking models (e.g. o1-pro, at $200/month) get it too, but all of DeepSeek-R1, Gemini 2.0 Flash Thinking, and Claude do not. ❌ It did not solve my "Emoji mystery" question where I give a smiling face with an attached message hidden inside Unicode variation selectors, even when I give a strong hint on how to decode it in the form of Rust code. The most progress I've seen is from DeepSeek-R1 which once partially decoded the message. ❓ It solved a few tic tac toe boards I gave it with a pretty nice/clean chain of thought (many SOTA models often fail these!). So I upped the difficulty and asked it to generate 3 "tricky" tic tac toe boards, which it failed on (generating nonsense boards / text), but then so did o1 pro. ✅ I uploaded GPT-2 paper. I asked a bunch of simple lookup questions, all worked great. Then asked to estimate the number of training flops it took to train GPT-2, with no searching. This is tricky because the number of tokens is not spelled out so it has to be partially estimated and partially calculated, stressing all of lookup, knowledge, and math. One example is 40GB of text ~= 40B characters ~= 40B bytes (assume ASCII) ~= 10B tokens (assume ~4 bytes/tok), at ~10 epochs ~= 100B token training run, at 1.5B params and with 2+4=6 flops/param/token, this is 100e9 X 1.5e9 X 6 ~= 1e21 FLOPs. Both Grok 3 and 4o fail this task, but Grok 3 with Thinking solves it great, while o1 pro (GPT thinking model) fails. I like that the model *will* attempt to solve the Riemann hypothesis when asked to, similar to DeepSeek-R1 but unlike many other models that give up instantly (o1-pro, Claude, Gemini 2.0 Flash Thinking) and simply say that it is a great unsolved problem. I had to stop it eventually because I felt a bit bad for it, but it showed courage and who knows, maybe one day... The impression overall I got here is that this is somewhere around o1-pro capability, and ahead of DeepSeek-R1, though of course we need actual, real evaluations to look at. DeepSearch Very neat offering that seems to combine something along the lines of what OpenAI / Perplexity call "Deep Research", together with thinking. Except instead of "Deep Research" it is "Deep Search" (sigh). Can produce high quality responses to various researchy / lookupy questions you could imagine have answers in article on the internet, e.g. a few I tried, which I stole from my recent search history on Perplexity, along with how it went: - ✅ "What's up with the upcoming Apple Launch? Any rumors?" - ✅ "Why is Palantir stock surging recently?" - ✅ "White Lotus 3 where was it filmed and is it the same team as Seasons 1 and 2?" - ✅ "What toothpaste does Bryan Johnson use?" - ❌ "Singles Inferno Season 4 cast where are they now?" - ❌ "What speech to text program has Simon Willison mentioned he's using?" ❌ I did find some sharp edges here. E.g. the model doesn't seem to like to reference X as a source by default, though you can explicitly ask it to. A few times I caught it hallucinating URLs that don't exist. A few times it said factual things that I think are incorrect and it didn't provide a citation for it (it probably doesn't exist). E.g. it told me that "Kim Jeong-su is still dating Kim Min-seol" of Singles Inferno Season 4, which surely is totally off, right? And when I asked it to create a report on the major LLM labs and their amount of total funding and estimate of employee count, it listed 12 major labs but not itself (xAI). The impression I get of DeepSearch is that it's approximately around Perplexity DeepResearch offering (which is great!), but not yet at the level of OpenAI's recently released "Deep Research", which still feels more thorough and reliable (though still nowhere perfect, e.g. it, too, quite incorrectly excludes xAI as a "major LLM labs" when I tried with it...). Random LLM "gotcha"s I tried a few more fun / random LLM gotcha queries I like to try now and then. Gotchas are queries that specifically on the easy side for humans but on the hard side for LLMs, so I was curious which of them Grok 3 makes progress on. ✅ Grok 3 knows there are 3 "r" in "strawberry", but then it also told me there are only 3 "L" in LOLLAPALOOZA. Turning on Thinking solves this. ✅ Grok 3 told me 9.11 > 9.9. (common with other LLMs too), but again, turning on Thinking solves it. ✅ Few simple puzzles worked ok even without thinking, e.g. *"Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have?"*. E.g. GPT4o says 2 (incorrectly). ❌ Sadly the model's sense of humor does not appear to be obviously improved. This is a common LLM issue with humor capability and general mode collapse, famously, e.g. 90% of 1,008 outputs asking ChatGPT for joke were repetitions of the same 25 jokes​. Even when prompted in more detail away from simple pun territory (e.g. give me a standup), I'm not sure that it is state of the art humor. Example generated joke: "*Why did the chicken join a band? Because it had the drumsticks and wanted to be a cluck-star!*". In quick testing, thinking did not help, possibly it made it a bit worse. ❌ Model still appears to be just a bit too overly sensitive to "complex ethical issues", e.g. generated a 1 page essay basically refusing to answer whether it might be ethically justifiable to misgender someone if it meant saving 1 million people from dying. ❌ Simon Willison's "*Generate an SVG of a pelican riding a bicycle*". It stresses the LLMs ability to lay out many elements on a 2D grid, which is very difficult because the LLMs can't "see" like people do, so it's arranging things in the dark, in text. Marking as fail because these pelicans are qutie good but, but still a bit broken (see image and comparisons). Claude's are best, but imo I suspect they specifically targeted SVG capability during training. Summary. As far as a quick vibe check over ~2 hours this morning, Grok 3 + Thinking feels somewhere around the state of the art territory of OpenAI's strongest models (o1-pro, $200/month), and slightly better than DeepSeek-R1 and Gemini 2.0 Flash Thinking. Which is quite incredible considering that the team started from scratch ~1 year ago, this timescale to state of the art territory is unprecedented. Do also keep in mind the caveats - the models are stochastic and may give slightly different answers each time, and it is very early, so we'll have to wait for a lot more evaluations over a period of the next few days/weeks. The early LM arena results look quite encouraging indeed. For now, big congrats to the xAI team, they clearly have huge velocity and momentum and I am excited to add Grok 3 to my "LLM council" and hear what it thinks going forward.
显示更多
0
666
16.8K
2.2K
转发到社区
一个让人不舒服的事实。 ChatGPT 每周活跃用户超过 4 亿。绝大多数人用它的方式都一样:打字提问,读回答,关标签页。 他们从没设置 Custom Instructions。从没打开 Memory。从没创建 Project。从没打开 Canvas。从没建 Custom GPT。从没用 Voice Mode。从没上传文件。从没触发 Deep Research。从没安排任务。 他们只用聊天框——15 个功能里的 1 个——然后用零上下文、一次性回答的质量来评判整个平台。 OpenAI 也有责任。他们把 15 个功能塞进每月 20 美元的订阅里,但从不带你过一遍任何一个。没有设置向导在第一天问你"你做什么工作?"没有引导流程帮你建第一个 Custom GPT。没有教程在你第一次通勤时展示 Voice Mode。每个功能都可用。没有一个被解释。 OpenAI 的商业模式是订阅——不是功能采用率。你用 1 个功能还是 15 个,他们都收同样的 20 美元。你的低利用率不会减少他们的收入。它减少的是你的价值。 "ChatGPT 不会给出 C+ 的回答。它给出的回答是根据它拥有的上下文校准的。大多数用户给它零上下文——没有指令、没有记忆、没有项目、没有文件、没有语音、没有历史——然后怪 AI 太泛。AI 不泛。输入泛。配置输入。输出就变了。" 9 个功能。一个晚上。同样的每月 20 美元。 AI 一直都有能力。设置一直缺失。而且没有人——不是 OpenAI、不是 App、不是引导界面——告诉过你打开 Settings。 C+ 的回答从来不是天花板。那是地板。天花板藏在 9 个功能后面,你从订阅那天起就能用——但从没配置过。
显示更多
Grok Bot is now more powerful for sales teams. Connect your Bots to Salesforce, HubSpot, Gong, Clay, Granola, and other GTM tools. Stay on top of accounts, complete follow-ups, and do deep research.
显示更多
0
64
1.5K
111
转发到社区
最近恰逢秋季开学,我也连续推荐了好几门北美顶尖高校新开的 Agent 课程,一个信号越来越明显: Agent Engineering 开始大量被正式写进顶尖高校的课程体系了。 这次是 CMU 新开的 11-768《AI Agents》,看完 syllabus 后感觉内容真的很贴近现在一线在做的东西。 Tool Use、Context、Skills、Memory、Planning,往后还有 Coding Agent、GUI、Deep Research、SFT/RL、Sandboxing 和安全。 作业设计是我更看重的部分。 这个课程要求先自己搭 Agentic Harness,再做 Agent Eval,然后需要用 RL 去训练 Agent,最后做一个研究项目。 这个顺序其实挺能说明 2026 年高校开始怎么理解 Agent 了,会调用模型已经很基础,接下来要学的是怎么把 Agent 搭出来、测明白、训练好,再让它可靠地跑下去。 而且课程资料会公开,视频也刚刚开始上传。 最近想系统补 Agent 的,很推荐跟着这门课学。 课程官网: 课程 Schedule: 课程 Assignments: youtubu 视频:
显示更多
0
46
938
250
转发到社区
ZERO ALPHA Research Preview | Reframing NVDA NVIDIA’s latest earnings report is the trigger event for a new round of deep research. A company already among the largest in the world just delivered 106% year-over-year revenue growth, with Data Center revenue up 117%. What is striking is not simply that NVIDIA beat expectations again, but that its core business has returned to a doubling growth rate from an already enormous base, even as AMD GPUs, hyperscaler-designed chips, and custom AI accelerators continue to enter the market. That prompted us to go back and re-examine NVIDIA’s full growth trajectory since 2023. When revenue growth, earnings growth, stock-price appreciation, and P/E are viewed together, a very different pattern begins to emerge. The first NVIDIA spring was largely top-down. The market recognized the potential of generative AI first, the stock price moved ahead, and earnings later caught up. The second spring now looks increasingly bottom-up. Revenue growth re-accelerated from: 56% → 62% → 73% → 85% → 106% while valuation multiples moved lower rather than higher. In simple terms: First Spring: P led E. Second Spring: E is beginning to lead P. This earnings report therefore may represent more than another earnings beat. It may be a signal that NVDA itself needs to be reframed. It also raises a broader question: What actually defines a true mega-cap growth stock? A high P/E alone does not define growth. The rarest structure may be a company that is already enormous, still grows its core business near 100%, generates earnings faster than its stock price rises, avoids excessive valuation expansion, and continues to create new TAM. Applying this framework to AMD, MU, SNDK, LITE, ALAB, DELL, and the hyperscalers makes the leadership hierarchy increasingly clear. Many of them have strong growth, but each still carries a weakness in valuation, cyclicality, pricing dependence, platform control, or growth durability. NVDA currently presents a more unusual combination. More importantly, at least four additional growth engines are still developing: Pricing Power Supply Efficiency Open Models Inference Specialization If these continue to develop, today’s NVIDIA may not yet represent the peak of this second growth cycle. And NVIDIA’s second spring may not belong to NVIDIA alone. Memory and storage, optical networking, and AI data-center operators could all benefit if another AI infrastructure expansion cycle is now beginning. ZERO ALPHA will therefore use this earnings report — a mega-cap company returning to 100%+ core growth — as the starting point for a six-part NVDA Research Note series: 1/6. NVDA: The Second Spring — From P Leading E to E Leading P 2/6. NVDA: What Defines a True Mega-Cap Growth Stock? 3/6. NVDA: Why It Is Still in Its Prime, Not Near the Peak 4/6. NVDA: Four New Growth Engines — How Far Can the Second Spring Go? 5/6. NVDA: Why Leaders Lose Leadership — Lessons from Intel, Tesla, and AMD 6/6. NVDA: Will the Second Spring Reignite the Entire AI Infrastructure Chain? Each note will focus on one independent question and can be read on its own. ZERO Insight The most important message from this earnings report may not be that NVIDIA beat expectations again. It may be this: When a company already this large returns to 100%+ core growth while trading at a much lower P/E than during its first AI explosion, what needs to be revalued may not be just NVDA’s stock price — but our entire understanding of mega-cap growth.
显示更多