For the past few wks since Grok Build launch, we've run SxS testing of complex Codex & Grok Build tasks. Eliminated Claude for needy intern BS. Grok Build wins more trust w/ higher accuracy & less hand-holding. Codex remains solid; each has unique strengths. For heavy math/science/data in aerospace, we're normalizing on Grok Build as of Jul 20. Exciting yet nerve-racking shift after 6mo on Codex. Qs on process/A/B testing welcome.
Congratulations to
@SpaceXAI and the Build TEAM!