Daniel's Quest

Read

DataBricks Benchmarks Coding Agents

This is a pretty interesting article that matches my experience. All of my personal projects are going through GLM 5.2 (using Kimi 2.7 Code as an alternate-model plan- and code-reviewer) using the Polytoken coding harness. I use Claude Opus 4.8 (in Claude Code & Cowork) at work. Between the two, I agree with their conclusion that GLM is about as good (Polytoken is way better than Claude Code so that might help a bit). As open models improve, I wonder how thatโ€™ll affect pricing at Anthropic / OpenAI? They may be able to continue to charge a high premium for Fable / Sol models, but how will companies justify paying at least 25% more for the same performance?

Article Link: https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase