Claim: we've solved the AI slop problem (!) 💩🧹✨
Blog post:
🧵1/5
Key idea: take *expert* human writing and learn rubrics that find the gap between experts and models. Train with those rubrics.
We train with RL-XAR (RL with eXpert Aligned Rubrics) & see large performance gains on writing scientific paper sections, Pulitzer prize novel continuations and high quality Wikipedia pages.
显示更多