Garp Independent AI & technology journalism
Saturday, August 8, 2026 Sign In · Join Subscribe
Latest Naïve raises $28.5M to automate the grunt work of setting up and running a company

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  Databricks makes Chinese open-source model GLM 5.2 its default coding engine after it matched Opus at lower cost

AI News

Databricks makes Chinese open-source model GLM 5.2 its default coding engine after it matched Opus at lower cost

Databricks makes Chinese open-source model GLM 5.2 its default coding engine after…

Databricks benchmarked coding agents on its own multi-million-line codebase and found that the Chinese open-source model GLM 5.2 matched Anthropic’s Opus 4.8 at $1.28 per task versus $1.94. The company plans to roll it out as a daily coding workhorse.

GLM 5.2 hit the top performance cluster at $1.28 per task versus $1.94 for Opus. “The evidence shows it’s time to start deploying these as daily drivers for coding,” write the authors of the blog post, including Databricks co-founder Matei Zaharia. Developer feedback from internal pilots backed up the results, and the company says it’s already working on running GLM at peak performance. Databricks isn’t alone. Coinbase moved to Chinese models including GLM-5.2 and Kimi 2.7, cutting AI spending in half while token usage kept climbing. Lindy ditched Claude entirely for Deepseek v4 and saved millions. Snowflake tested GLM-5.2 against Opus 4.7 and found them nearly tied at a fraction of the cost. On OpenRouter, Chinese models have topped 30 percent of weekly traffic since February 2026, up from 11 percent last year, at 60 to 90 percent lower cost than Western alternatives.Ad No single lab dominates across three performance tiers Overall, the tested models and configs fell into three clusters, according to Databricks. The top group, with an 82 to 90 percent pass rate, includes Opus 4.8, GLM 5.2, and GPT 5.5 in certain configs. A middle group at 71 to 82 percent includes Sonnet 4.6, Sonnet 5, and GPT 5.4, among others. The bottom tier at 51 to 60 percent holds GPT 5.4-mini and Haiku 4.5.AdDEC_D_Incontent-1 An analysis through Unity AI Gateway found that 61 percent of coding tasks from Databricks engineers are medium complexity, about 19 percent low, and only 12 percent high. The most expensive models had been the default. Now the company plans to route more work to cheaper tiers based on task complexity. The Pareto frontier, the best quality-to-cost ratio, is shaped by models from three providers: OpenAI, Anthropic, and open source. Only a mix delivers frontier-level performance, Databricks says.Ad Databricks also points out that token price and actual task cost aren’t the same. Token efficiency matters just as much, like fuel economy in a car, and varies widely by software environment. In one test, the Pi harness sent about three times less context than Claude Code. For Opus 4.8 at “high effort,” Pi was 2.08x cheaper at comparable quality (85 versus 87 percent).