RC RANDOM CHAOS

LLM engineering

35 posts

Liquid AI's 8B-A1B drop rewrites inference math
Article

Liquid AI's 8B-A1B drop rewrites inference math

Liquid AI's 8B-A1B MoE trained on 38T tokens shifts LLM inference economics. What it means for engineering pipelines and workforce planning.

One billion fire, eight billion sit in memory
Article

One billion fire, eight billion sit in memory

Liquid AI's 8B-A1B MoE frees compute and latency, not memory. How to match sparse-model architecture to the real constraint in your deployment.

The bottleneck moved past the model
Article

The bottleneck moved past the model

Notes from the Mistral AI Now summit on what the new enterprise stack means for automation pipelines and workforce transformation.

The refund letter addressed to Dear [Name]
Article

The refund letter addressed to Dear [Name]

Why ChatGPT's first output is a draft, not a deliverable, and what production AI systems actually require beyond the prompt.

The smooth line hiding a noisy benchmark
Article

The smooth line hiding a noisy benchmark

The METR AI time horizons graph contains structural errors that mislead teams building agents, automation, and AI workflows. Here is what it actually shows.

Hugging Face revived PapersWithCode in early 2025
Article

Hugging Face revived PapersWithCode in early 2025

Hugging Face's PapersWithCode revival restores the verification substrate LLM engineering teams lost, reshaping pipelines and AI workforce roles.

Sub-JEPA tightens the prediction signal
Article

Sub-JEPA tightens the prediction signal

Sub-JEPA is a small loss-side fix to LeCun's world models that consistently improves performance. Here's how it works, where it fails, and why it matters.

Better AI isn't what separates winning deployments.
Article

Better AI isn't what separates winning deployments.

Stanford studied 51 AI deployments and found a 71 vs 40 productivity gap. The difference was pipeline design, not model choice.

arXiv just raised the bar
Article

arXiv just raised the bar

arXiv's one-year ban on unchecked LLM errors signals a shift: validation pipelines, not better prompts, now define competent AI systems.

Complexity theory never said that
Article

Complexity theory never said that

Complexity theory does not prove human-level ML is impossible. Here is what the theorems actually say and how to design AI systems around real constraints.

Article

AI costs more than humans

Nvidia says AI costs more than human workers. The real issue is architecture, not compute price. Here is how to fix the unit economics.

Managed Agents pricing is an architecture decision
Article

Managed Agents pricing is an architecture decision

Claude Managed Agents pricing isn't a cost center - it's an orchestration lever. Here's how to evaluate it against real total cost of ownership.