MicroLLM Lab runs seven models without touching the network
In-browser tiny LLMs keep prompts off the network, but same-origin scripts, local storage, WebGPU fingerprinting, and weak model safety reshape the risks.
A tiny language model running inside a browser tab does something its cloud cousin cannot: it answers your prompt without a single byte leaving the machine. Open MicroLLM Lab, pick one of its seven models, type a question, and the whole exchange runs on your own GPU. No API call, no server-side log, no vendor holding a transcript of what you asked. That is a genuine privacy property. It is also narrower than it first sounds, and it rearranges the security tradeoffs rather than deleting them.
The technology under these demos is now boringly available. WebGPU shipped in Chrome 113 in May 2023, and runtimes like WebLLM from MLC and Transformers.js from Hugging Face compile models to run against it. That is what makes a seven-model playground possible in a page: a handful of small networks - think Qwen2.5-0.5B, Llama-3.2-1B, TinyLlama-1.1B, SmolLM2, Gemma-2-2B, Phi-3.5-mini - quantized down far enough to load in a tab. Understanding what that actually does to your data requires looking at each layer, not the marketing sentence.
What actually loads when you open the page
When you open one of these demos, the browser downloads a model file. A 4-bit quantized Llama-3.2-1B is roughly 700 to 800 MB. Qwen2.5-0.5B is closer to 400 MB. Phi-3.5-mini runs to several gigabytes even after quantization. That file comes from somewhere - usually the Hugging Face CDN - and gets written to local storage through the Cache API or the Origin Private File System so the second visit is instant.
Two consequences follow. First, the weights are now data your browser trusts and executes against. If the CDN serving them is compromised, or the page is pointed at a poisoned mirror, you run whatever those weights encode. Most in-browser demos do not pin model files with Subresource Integrity hashes, because SRI on a multi-hundred-megabyte artifact split across shards is awkward and rarely implemented. That means the integrity of the thing doing your local inference usually rests on transport security and the CDN’s own controls, not on a hash you can verify. Second, a gigabyte of model now sits in your profile. On a shared or managed machine, that cache is readable evidence of which tool you ran, and it persists until something clears it.
The privacy win is real, and it stops at the tab boundary
“Your data never leaves the device” is true at the network layer and misleading at the JavaScript layer. Your prompt and the model’s output live in the page’s memory, inside the same origin as every other script that page loaded. A local model protects your text from the server operator. It does nothing to protect it from code running in the same tab.
This matters because modern pages are rarely one script. If the site loads a third-party analytics tag, an ad network, a chat widget, or a compromised npm dependency somewhere in its build, that code executes with the same access to the DOM and the JavaScript heap as the chat interface itself. An attacker who lands a cross-site scripting payload on the page can read every prompt you type and every token the model returns, then exfiltrate them to a server - quietly undoing the entire “nothing leaves the device” premise. The local model removed the vendor from the trust equation and left every other script on the page exactly where it was.
So the honest framing is this: in-browser inference shrinks the set of parties who can see your data from “you, the app vendor, and the model host” down to “you and whatever code the page runs.” That is a meaningful reduction for a self-contained, dependency-light tool. It is close to no reduction at all for a page stuffed with third-party tags.
Where the conversation actually gets stored
Local inference does not mean nothing is written to disk. The runtime caches weights in OPFS or the Cache API by design. Separately, the application layer decides what to do with your conversation, and many chat UIs keep history in localStorage or IndexedDB so it survives a refresh. That history is plaintext, scoped to the origin, and readable by any script on that origin - including a future compromised version of the same site.
The practical risk is the ordinary one: shared and borrowed machines. A cloud chatbot ties history to an account you can log out of. A browser tool ties it to the browser profile. If you type something sensitive into a local model on a library terminal or a colleague’s laptop, closing the tab does not clear IndexedDB, and the next person with that profile can read it. “It ran locally” can be worse for privacy than a server session that ended when you signed out, unless the app explicitly clears its own storage.
WebGPU widened the fingerprinting surface
Running a model on the GPU requires exposing the GPU, and that costs you some anonymity. WebGPU hands the page adapter details - vendor, architecture, limits - and precise timing behavior. Those signals add entropy to a browser fingerprint, and they distinguish a machine that has a capable discrete GPU from one that does not. A tracker does not need your prompt to profile you when the mere ability to run a 2B model, plus the shape of your hardware, narrows the field.
GPU-timing side channels are a known research area, not a theatrical one. Shared execution units and memory bandwidth have been used to infer activity across contexts. You do not need to treat this as an emergency, but you should count it: enabling local inference turns on hardware access that is also a discriminator. For most users the tradeoff is fine. For someone whose threat model includes being singled out, “the model is private” and “the session is unlinkable” are two different claims, and only the first one is true here.
Small models make bad gatekeepers
The security failure I expect to see most often has nothing to do with data leaving the device. It is developers wiring a tiny model into a decision it is not fit to make. A 0.5B or 1B model is a convenience feature. Its safety tuning is thin, its refusal behavior is inconsistent, and its hallucination rate is high enough that you cannot treat its output as a fact source without checking. That is fine for drafting text or summarizing a paragraph the user already trusts. It is dangerous the moment it sits in a control path.
Concretely: do not use an in-browser tiny model to classify whether user input is malicious, to sanitize content, to make an authorization call, or to gate access to anything. If the model is processing untrusted text - a pasted email, scraped page content, a document the user did not write - prompt injection applies in full. Instructions hidden in that text can steer a small model easily, and a small model is more suggestible than a frontier one, not less. Treat its output the way you treat any untrusted string: as data to be validated, never as a decision to be trusted. Moving the model into the browser does not change that; it just puts the weak component closer to the user.
What to verify before you ship one
If you are building on top of in-browser inference rather than just playing with it, a short list separates a real privacy tool from a page that only claims to be one.
Check where the weights come from and whether their integrity is verified - pin what you can, and host the model files yourself if the privacy story is the whole point. Audit every third-party script on the page, because each one shares access to the prompts; a genuine local-first tool runs with as few external tags as possible, ideally none. Decide explicitly what the app writes to localStorage, IndexedDB, and OPFS, document it, and give the user a working control to clear conversation history and cached weights. Set a Content Security Policy tight enough that a script injection cannot phone home, since that is the mechanism that would turn a local session into a leak. And keep the model out of any security decision - label it a drafting aid, and validate anything it touches.
The MicroLLM Lab framing - seven tiny models, all in the tab - is a good way to feel what local inference does and does not buy you. What it buys is real: the network operator and the model host drop out of the picture, which is exactly what you want for text you would not send to a server. What it does not buy is protection from the code on the page, unlinkable sessions, safe defaults for storage, or a model you can trust to make calls. Move the model to the edge and you have moved the private part of the pipeline. The rest of the browser is still the browser.
Keep Reading
privacyChatGPT already knows your other browsing.
ChatGPT's ad collector integration joins cross-site tracking data with your prompts under one identity. The exposure, the mechanism, and what stays unconfirmed.
AI modelsMost automation doesn't need a smart model
How 8-29MB automation models like Cactus Needle 3 match large models on narrow tasks, where they fail, and the security tradeoffs of running them locally.
privacyYour Git history is already in the cloud
ZCode transmits local Git history to the cloud with no consent prompt and no notification. A post-incident breakdown of the boundary that was never enforced.
Latest on the Wire
Full wire →- 101 npm packages secretly add devs' WhatsApp to groupsThe Hacker News
- 543,000 Valid Credentials Still Exposed on GitHub Despite SafeguardsBleepingComputer
- 6 Dangerous Browser-Based Attack Techniques in 2026The Hacker News
- AI and the timeless fear of technological replacementHacker News
New signal daily · RSS
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.