如何解决 AI 写文章的 AI 味儿?
现在的 AI 写数学题或写代码很厉害,但在写学术论文、小说或百科文章时,经常产出空话连篇、表面流畅但缺乏真正深度的文字,这种低质文本常被称为“AI 废话”(AI slop)。
出现这个问题的原因在于,常用的 AI 评分系统很难辨别真正优秀的文字。研究人员发现,如果让目前的先进模型给文章打分,或者让它们自由制定打分标准,模型给出的分数往往偏向 AI 生成的内容,甚至给 AI 文本的分数远高于人类专家的原著。打分裁判本身偏好废话,训练出来的写作模型自然就会写出更多废话。
为了解决这个弊端,META 提出了一套名为 RL-XAR 的新训练方案,核心思路分为三步:
1)准备参照标杆:收集高水平的人类写作范例,比如真实的科研论文、名家小说和维基百科精选词条,截取前半部分作为上下文,分别让人类专家和 AI 模型顺着往下写。
2)训练评分规则:让 AI 裁判反复调整打分标准,直到这些标准能准确把人类高手的文本评为高分,把 AI 生成的文本评为低分,真正拉开两者的差距。
3)针对性强化训练:用这套校准后的高难度评分规则去重新训练写作模型,不断循环这个过程,逼迫模型改掉毛病。
META 在三个领域测试了这种方法:
· 学术论文写作:模型需要补写论文中被拿掉的摘要、引言或结论。经过新方法训练的模型,不再像以前那样把整篇论文翻来覆去复述一遍,而是学会了精简和突出重点。原作者进行盲测评估时,新模型的胜率达到了 89%。
· 文学小说续写:续写普利策奖和诺贝尔奖得主的作品。模型改掉了乱用陈词滥调的比喻这一常见毛病,盲测胜率达到了 95%。
· 维基百科编写:根据标题直接写出正文内容。这项任务难度最高,模型虽然也有进步,但因为现有的 AI 裁判能力有限,提升幅度不如前两项明显。
所以,想让 AI 写出更像人类高手的深刻文字,关键在于先训练出眼光挑剔的评分标准,避免让模型用平庸的标准自我满足。
原文:
Hi,
Tomorrow we are re-opening the Pro $200 subscriptions to new subscribers, but together with it we are also changing how we calculate the usage for it. In effect, if you do the math, it will net out at half the dollar in API spend compared to the old Pro $200 plan.
Now that it's said, let me explain why this is happening and why you will still get more work done than if you were on the Pro $200 subscription one month ago.
(a) We didn't want to compromise in other ways and are committing to not reintroducing the 5h limit, so that you can fully use the weekly usage when you want.
(b) On the subscription, we guarantee that over time you always get more work done and with an increasing level of quality. This means that you will continue to get more value per dollar spent as a result of models getting more efficient and us passing down the improvements in the form of API price reductions.
(c) We don't want to put an incentive on ourselves to artificially inflate the API list prices to make it look like you are getting a lot (and workaround it through discounts, etc). Instead we want to continue to both rapidly reduce prices and increase capabilities of models on the API. This week we introduced GPT-6 Sol and GPT-6 Luna at 50% of their previous price. Over time, we see prices go low enough that it makes sense for most to buy usage as needed without there being a significant gap between what you get in a subscription and what you get in the API for a dollar spent.
(d) Tomorrow, we are adding more things to the subscription that won't draw on the usage, I won't reveal what that is yet.
I wanted to be transparent before all the big announcements tomorrow. Lots of new exciting things are coming to the subscriptions that will make it super compelling, but I wanted to make sure to share this change ahead of time so you can all understand it before we shower you with good news.
Codexingly,
Tibo
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family.
It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
Meta 宣布推出名为“Meta Enterprise Platform”的企业服务平台,正式把服务企业客户作为公司未来的核心业务支柱之一。
Meta 打算把自家的底层技术开放给各类公司和开发者使用。初期开放的技术包括智能助手 Muse、商业助手 Meta Business Agent、编程工具 Muse Code 以及相关的软件开发接口,方便企业把 Meta 的 AI 能力直接接入到自己的业务流程中。
为了负责这项新业务,Meta 挖来了企业软件领域的资深管理者 CJ Desai 出任首席企业平台官,直接向扎克伯格汇报。CJ Desai 之前担任过 MongoDB 的首席执行官,也在 Cloudflare 和 ServiceNow 担任过核心高管。
Meta 表示,后续会把保护数据隐私和系统安全放在首位,帮助不同规模的公司借助 Meta 的 AI 工具提高运转效率、拓展客户群。
官方公告: