RC RANDOM CHAOS

Z.ai's GLM-5.3 Helped Build the Inference Stack That Now Runs It

· via Hacker News

Original source

GLM Built Its Own Inference Infrastructure

Hacker News →

Z.ai (the former Zhipu AI) stood up a production inference service for its GLM-5.3-Flash model on a cluster of more than 100,000 Chinese-made AI accelerators — a scale the company claims no one had previously operated on domestic silicon. The notable twist is who did the work: much of the tuning was driven by an ‘Infra Agent’ powered by GLM-5.3 itself, running in a loop where human engineers set objectives, defined system boundaries, and signed off on critical changes while the model handled analysis, hypotheses, and code. The stack pulls together intra-node tensor parallelism, W8A8 and mixed INT8/FP8/BF16 quantization, ReplaySSM, and a disaggregated encode-prefill-decode architecture.

The results Z.ai reports are concrete: from first successful run to production readiness in under two weeks, end-to-end throughput roughly tripling against the initial baseline, and per-token cost and hardware efficiency reaching parity with mainstream Nvidia GPUs. After launch the system processed more than 62 trillion tokens in six days. For a Chinese lab operating under export restrictions, hitting Nvidia-comparable economics on domestic chips is the more strategically important claim than the throughput multiplier.

Z.ai wraps the story in recursive-self-improvement framing — ‘the model optimizes the system; the system runs the model’ — but is careful to walk it back, stating it has not actually reached RSI and that choosing goals, setting limits, and judging risk remain human responsibilities. Read skeptically, this is a vendor-authored account with self-reported numbers and an obvious incentive to lean on the AGI-adjacent narrative, but the underlying engineering result — a large-scale non-Nvidia serving cluster, partly automated by the model it serves — is the part worth watching.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.