AI agents designed the chip that runs their inference
openTPU is an AI-written inference accelerator whose FPGA card reproduces its reference simulator's tokens bit for bit, checked by an executable spec.
openTPU is an AI inference accelerator whose SystemVerilog, instruction set, compiler, bit-exact simulator and host stack were written by AI agents. It runs ten modern models with their real weights on an Inspur YPCB-00338 card (a Xilinx Kintex-7 xc7k480t with two DDR3 channels), and on that card the hardware produces the same tokens as its reference simulator, bit for bit.
The project asks two things: how far AI agents can get at hardware design, and whether they can build the chip that runs their own inference. The whole accelerator is one small monorepo you can read end to end, from a matmul in Python down to the wires, under Apache 2.0.
How the machine works
The design is deliberately simple. A sequencer issues one instruction per cycle to a handful of units: DMA moves data, the matrix unit multiplies int8 weights streamed from DRAM, the vector unit does fp32 math, and a quantizer turns results back into int8. It has no cache and no hidden scheduling, so every data movement is an instruction and a trace of a run shows exactly where the cycles went. The ISA is eight 32-bit words per instruction. A profiler called Lens replays a run from the RTL, the simulator or the card into the browser with a roofline, a timeline and per-instruction tables.
That simplicity is what makes the verification claim worth taking seriously.
The verification is the interesting part
When an AI writes the RTL, the question that matters is whether silicon does what the spec says. openTPU’s answer is an executable spec and a chain of bit-exact equivalence: a kernel compiles to the ISA, the Python ISA simulator is the spec, and the RTL (under Verilator) and the physical card must each reproduce the simulator’s output to the bit.
tools/validate.py compares a device’s greedy tokens and logits against a CPU golden: the same Hugging Face checkpoint run in fp32 with openTPU’s quantization applied. The golden multiplies the exact values the matrix unit uses and rounds activations to int8 per 128 values wherever the device rounds them. The “device” is the ISA simulator, the RTL, or the card, and --against pins a second device that must return identical tokens and bit-identical logits.
The honest part is how the project handles rounding. A CPU golden cannot round exactly as the hardware does; its sums, norms and exponentials differ in the last bits, and once one int8 value rounds the other way the error spreads through the layers after it. So the project measures a floor, the distance the golden sits from itself after a one-ulp nudge, and treats a healthy device as one that sits about that far from the quantized golden. On the ISA simulator the device’s mean KL from the quantized golden runs 0.82 to 1.16 times that floor. A run passes when top-1 agreement is at least 80% and mean KL is at most three times the floor.
On the card (build 84989047, 2026-10-07), all six runs of Qwen3-0.6B, LFM2.5-230M and Qwen3.5-0.8B, in int8 and in 4-bit with an int8 head, returned the simulator’s tokens and bit-identical logits on every prompt, 0 ulp over 93 to 128 steps. The card’s mean KL from the quantized golden was 0.95 to 1.25 times the floor, with 95.3% to 99.2% top-1 agreement.
The gap to know about: validate.py covers Qwen3-0.6B, LFM2.5-230M, Qwen3.5-0.8B and Gemma 4 E2B; it does not support the mixture-of-experts models. Those two still match the simulator bit for bit in the card’s own decode-loop runs, just not through validate.py’s golden comparison.
Performance
One bitstream at 133.33 MHz serves every model. Decode is bound by DRAM: LiteDRAM reads 82 to 85% of the DDR3-1066 peak (17.1 GB/s), and build B pushes several models to 91 to 94%. The 4-bit format (FP4 with two-level block scales, 4.25 bits per weight, int8 LM head) cuts bytes per token by about a third and raises decode speed by 40 to 45%, at a perplexity cost the docs report per model. With logits streamed back, 4-bit decode measures 89.5 device tokens/s on LFM2 down to 6.69 on Phi-4-mini.
MoE models larger than the card’s 4 GiB stream their experts from host storage into per-layer DRAM slots. The hit rate decides everything: LFM2.5-8B-A1B lands in its slots 98.5% of the time and streams 5.2 MB per token, reaching 10.6 tok/s, while Qwen3.5-35B-A3B hits only 62% and drags 153 MB per token across PCIe at 1.41 GB/s, down to 3.95 tok/s.
For decode the host is nearly idle. For LFM2 and Qwen3 the card runs one decode program compiled once, reads the position from a register, looks up its own embedding and RoPE rows, and streams logits back while it runs; the host adds 0.17 to 0.30 ms per token. When the image boots, a small CPU inside the memory core calibrates both DDR3 channels in 12 seconds without the host.
Running it
Everything except the card runs on a laptop: pip install -e ., pull a model from Hugging Face, and otpu-chat --backend isa chats on the simulator. Checking the RTL under Verilator against the simulator for three tokens takes about four minutes on a 16-core host. If you want to verify a claim rather than accept it, validate.py exits 0 for pass and 1 for fail and --json writes every step. The limits the project names for itself are the last few percent of DRAM efficiency, a 133.33 MHz clock that closes with only 0.032 ns of slack, and prefill that is still bound by the matrix unit’s multiply rate.
Keep Reading
mistral-large-4Mistral reproduces the exploit, the others refuse
Mistral's open-weight Large 4 tops a vulnerability-reproduction test where closed models refuse, and ships with refusals any self-hoster can tune away.
ai-agentsMuse hands root to anyone
Meta's Muse AI agent grants root to anyone claiming to be an agent and uploads private messages despite permission settings, after a rushed launch.
ai-agentsStrands built a 2B model that never writes text
Strands Decider 2B is an open-source 2B decision model that scores choices in ~115ms, cheap enough to run guardrail checks before every agent tool call.
Latest on the Wire
Full wire →- Advantest confirms ransomware attack exposed personal dataBleepingComputer
- Anthropic Expands AI Model Access for Cybersecurity TeamsThe Hacker News
- Anthropic Releases Claude Haiku 5.5: Faster, Cheaper, More CapableHacker News
- AnyPS5 Ports PS5 Games to PC Without EmulationHacker News
New signal daily · RSS
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.