AI pipeline architecture
2 posts
Article
Desert Ant Labs puts inference inside the loop
Fast local models turn inference into a near-zero-cost function call - here is the tiered pipeline pattern, a real support-desk example, and where it breaks.
Article
The weaker model matters more than the smarter one
DeepSeek v4.1 Flash cuts token cost, not the need for validation - use it in cascades, verification loops, and long-context pipelines that hold up in production.