1 post
GLM built its own inference infrastructure to serve LLMs cheaply on constrained, mixed hardware. Here's what breaks in generic stacks and what to copy.