Zhipu Says Its GLM 5.3 Flash Model Served a Week of Record Traffic Entirely on Chinese Chips
The Beijing AI startup Zhipu revealed that its new GLM 5.3 Flash model, which topped third-party leaderboards all week under the anonymous alias Ox Alpha, ran its public test entirely on domestically made chips — a live demonstration that frontier-class AI inference can be served at scale without foreign hardware, timed just before Guiyang's data-industry expo.

For most of last week, one of China's strongest new AI models ran in public under a name that meant nothing: Ox Alpha. On Thursday, the Beijing startup Zhipu AI revealed that the anonymous model — which had become the most-called model of the week on the developer platforms OpenCode and OpenRouter, setting record traffic on both — was in fact GLM 5.3 Flash, and that every request it answered was processed on Chinese-made chips. The company said it had released the model quietly before launch to gather professional feedback, and disclosed that end-to-end serving performance had tripled against its starting baseline on the same hardware.
The scale of the experiment is the point. OpenCode had listed Ox Alpha with a free allowance of 100 trillion tokens a day, and roughly 70 trillion tokens were consumed over the week, according to the Zhihu discussion that greeted the reveal. A day earlier the model had been an open mystery; the Weibo tech commentator Li Nan, one of the first to flag it, called it "the Ox Alpha 'Niu Lai' model... Opus 4.8-level, at one-fortieth the price, 320 billion parameters with 18 billion active" (DS 是谁?牛夫人了 — "So who is DeepSeek? Yesterday's news"). His verdict, posted to more than a hundred likes, captures why the industry took notice: Zhipu claims its new model approaches the capability of the best American frontier models while costing a fraction as much to run — and running, crucially, on domestic silicon.
The timing was no accident. On August 28, the city of Guiyang opens this year's China International Big Data Industry Expo, state media framing the theme as a shift "from green computing to the token economy" (从绿色算力到词元经济) — the idea that the unit of value in the AI era is not raw capacity but the tokens a system can serve cheaply, like kilowatt-hours before it. The exposition is being held in Guizhou, a mountainous, once-impoverished province that rebranded itself as the country's data belt by courting data centers to its cool climate.

The infrastructure behind that pivot is the national East Data West Computing program, which deliberately routes new computing investment to inland provinces where land and power are cheap. CCTV's expo graphics note that more than 70 percent of China's new computing capacity now rises in the west, that over 60 percent of cities can reach a computing cluster within five milliseconds, and that some servers in Guizhou sit in excavated caves, using the surrounding rock for cooling to cut energy use by about 30 percent and save 40 million yuan a year in electricity.

The question Zhihu users were still asking Thursday was where Zhipu found so much spare capacity to give away. The company's own answer — domestic chip clusters, wrung three times more efficient in a week of live traffic — suggests the constraint that once defined China's AI industry is being renegotiated one benchmark at a time.