Grok 4.6 multimodal is a step change from Grok 4.5. It’s one of the under-discussed improvements, and I’ve been very impressed by it.
My daily work includes reviewing lots of videos and understanding the context; Grok 4.6 improves the productivity of such workloads by at least 10x if not 100x.
Such workflows may not be captured by common VLM benchmarks, but in my use cases it outperforms Gemini 3.5 and Gemma 4, which is considered the SOTA of VLMs in my opinion.
Hats off to the multimodal teams—you did a great job.
显示更多