Claude can now help you build evaluations and hillclimb on them.
In this article, we share guidance on eval design & skills that Claude Code can use to improve your applications.
Testing out some of the mobile offline local knowledge apps that people have been trying to build
(see here )
Definitely getting much better than the one I tried to build myself 2 months ago. But also still much slower and less effective at difficult questions than the models that can run on a laptop.
It's weakest at specialized travel-related queries (eg. my eval is "Tell me the best vegan restaurants in [city I am currently in]", unfortunately none of these performed well on that)
Looking forward to seeing these continue to improve! I hope we can soon get to the point where you can comfortably look up any facts about the world that you care about without needing to access the internet at all.
Remember the time I filmed this super illegal video and someone (me) leaked it and now I see it everywhere? What a time.
Also, swipe for the only other photo I have from the Endgame set. Me and Evans back in 2017.
📸: @chrishemsworth
What is effort really? When do you change it it and why not just use max effort for everything?
I dove deep into this problem, looking into evals and doing my own tests and I was quite surprised by the results.
Super happy to release SmolDataEnvs: 5,000 verifiable RL environment tasks for hill-climbing small models in code and data science by @adithya_s_k
100% open source: environments, evals, training!