Run an LLM Right Inside the User's Browser, No Server, No API Bill

Blog post by Anand: "Run an LLM Right Inside the User's Browser, No Server, No API Bill" (published June 16, 2026; categories: Tech). Run a real LLM fully on the user's device — no server, no API key, no per-message bill. How Transformers.js and web-llm pull it off with WASM and WebGPU, and the honest trade-offs I found.

I'm in the middle of researching my new project. On that, I try to tackle a few problems, simply trying to find a solution. So, in that way, I find one of the pieces of knowledge about making LLMs run on the user's browser. Because that sounds like crazy. But ya, they have possibilities.

Just imagine you have developed an AI product. But you don't like to spend money on an LLM. Then there is a way to run that AI model on the user's WASM/WebGPU. Then use the model in our application.

When I learned about it, I felt, Is it true? Is there an option like that? But ya. They have two different methods which achieve this.

The two methods

After digging more into it, I found there are mainly two methods to pull this off. Both of them run the model on the user side, but they work in their own different way.

Method 1: Transformers.js (package name: @huggingface/transformers)

This one is from Hugging Face, the same people who host almost all the open-source models. Simply a GitHub of open-source LLMs. They made a JavaScript version of their famous Python library, and the package you actually install is called @huggingface/transformers.

Small note. Earlier, it was named @xenova/transformers. Same thing, just Hugging Face officially adopted it and renamed it. So if you see old blogs using @xenova/transformers, that's the same package.

The cool part is that this one runs the model using something called ONNX Runtime, and it works in two ways. WASM uses the normal CPU, and WebGPU uses the user's graphics card. So even if the user device doesn't have a good GPU, it still falls back to the CPU and keeps working. It just becomes slower, but it won't break.

And another good thing, this same package works in the browser and in Node.js also. So you are not locked into one place.

But wait, what is this ONNX, and how does the model "convert" to it?

This part is some more crucial. So, all of them, switch into the serious mode.

I hope you've done it. Ok, I'll explain it simply.

Normally, models are trained in PyTorch (Python world). But a PyTorch model can't run directly in your browser because the browser doesn't understand PyTorch. So Hugging Face uses a middle format called ONNX (Open Neural Network Exchange). Think of ONNX like a PDF for AI models. One common format that many different programs can open and run, no matter where the model came from.

So the model has to be converted from PyTorch into ONNX first. There are two situations here.

First, most of the time you don't convert anything yourself. Hugging Face has already converted thousands of popular models to ONNX and uploaded them under orgs like Xenova and onnx-community. You just point your code to that model name, and it downloads the ready-made ONNX file. Done.

Second, if you want to convert your own model, Hugging Face has a tool called Optimum. With basically one command, you give it a normal model, and it produces the ONNX version that Transformers.js can run in the browser. You can also quantize it at this step to make it smaller and faster.

So the flow is like this. PyTorch model, then Optimum converts it, then you get the ONNX file, then ONNX Runtime runs it in the browser using WASM or WebGPU.

Method 2: web-llm (package name: @mlc-ai/web-llm)

This is the second method, and this one is fully focused on the browser only.

It is built by the MLC team. MLC stands for Machine Learning Compilation, and the project comes out of CMU, which is Carnegie Mellon University, along with the open-source Apache TVM community. Their whole obsession is one thing: make a model run as fast as possible on any device.

How it works

Transformers.js takes a general format (ONNX) and runs it through a general engine. Web-LLM does something different. It compiles the model.

What does "compile" mean here? Instead of using one general engine for all models, MLC uses Apache TVM to take a model like Llama 3 or Phi-3 and turn it into two pieces. One is the quantized model weights, which are the compressed numbers. The other is a small WASM library plus WebGPU shaders, which is basically custom GPU code that is tuned exactly for that model.

A shader is just a tiny program that runs on the graphics card. So at runtime, web-llm sends these shaders straight to the user's GPU through the WebGPU API, and the GPU runs the math at near native speed. The non-GPU parts run in WebAssembly. And the heavy work happens inside a Web Worker, which is a background thread, so your website UI doesn't freeze while the model is thinking.

This is why web-LLM is usually faster than Transformers.js for chat. The model is pre-compiled specifically for the browser GPU, not run through a general engine.

One more nice thing. Their code looks exactly like the OpenAI API. So if you already wrote your app in OpenAI style, you can switch to a web-LLM with almost no change.

But ya, it has one condition. It needs WebGPU. If the user's browser doesn't support WebGPU, it simply won't run. There is no CPU fallback like Transformers.js. And it works only in the browser, not in Node.js.

So, how does this actually work overall?

Let me explain thoroughly, because the first time I also got confused.

A model is basically a big file full of numbers. To use it, some program has to read those numbers and do the math. Normally, that program runs on a big server with a GPU, and that's why companies charge you money: their server is doing the work.

But here, the trick is different. Instead of running on your server, the model file gets downloaded into the user's browser, and then the user's own device does the math. Their CPU or their graphics card. Not yours.

The flow goes like this. The user opens your website. The model file, around 700 MB to 1 GB for a small model, downloads into their browser one time. The browser stores it in the cache, so next time, no need to download again. From there, every question the user asks runs fully on their device. No internet call, no API key, no bill for you.

That's the whole magic. Your hosting, AWS, GCP, Vercel or Netlify or anything, is only sending the website files. The actual AI work happens on the user side.

Comparison: Transformers.js vs web-llm

This is the part many blogs skip, so let me put it clearly in simple words.

The big difference is how they run the model. Transformers.js uses ONNX Runtime, which is a general engine that can run any ONNX model. Web-LLM is different; it compiles the model with TVM into custom WebGPU shaders made just for that model. Because of this, the model format is also different. Transformers.js uses ONNX files; web-llm uses its own pre-compiled format.

Both of them run in the browser; that's the whole point. But Transformers.js also runs in Node.js, while web-llm is browser-only.

Now, the WebGPU part, this is important. Transformers.js doesn't force WebGPU. If the device has no good GPU, it falls back to the CPU using WASM. It becomes slower, but it still works. Web-LLM needs WebGPU, full stop. If the browser doesn't support it, web-llm simply won't run. There is no CPU fallback.

For speed, web-LLM is usually faster for chat because the model is compiled specially for the GPU instead of going through a general engine.

So the simple takeaway. Transformers.js is the all-rounder. It works in more places, it does embeddings, and it has the CPU fallback. Web-LLM is the speed specialist. Faster chat, but browser only, and it needs WebGPU. Each one is doing what it's best at.

But I have to be honest, it's not all perfect

After I got excited, I also tested the reality. And there are some real problems you should know before you build on this.

It uses the user's device power. The model runs on their phone or laptop, so the device can get slow, hot, and the battery drains faster. A strong laptop handles it fine, but a weak phone will struggle.. Incognito mode breaks the cache. In incognito mode, the browser doesn't keep the cache. So the 1 GB model downloads again every time. Same problem if the user clears their cache.. The first load is heavy. That first 700 MB to 1 GB download takes time. A casual visitor might leave before it even finishes.. WebGPU is not everywhere. Around 70 to 75 percent of devices support it. The rest fall back to slow CPU mode, or in web-llm's case, just don't work.. And small models are not geniuses. A 1B or 2B model is good for simple tasks like RAG question answering, todo extraction, and basic chat. But don't expect GPT-4-level reasoning.

So this is not a magic free lunch. It's a trade. You save money, but you give the work to the user's device, and you accept the download problem. For the right type of product, it's worth it. For mass-market mobile users, maybe not.

Final thought

When I first heard about running an LLM inside the browser, I really thought it was impossible. But it's true, and the tools are already there: Transformers.js and web-llm. It's not a perfect solution; the download and device-speed problems are real. But for a small developer who wants to add AI without paying for every single message, this is a genuine option worth knowing.

I'm still researching this for my own project, and I'll share more once I test it more deeply. But ya, the crazy idea is real. The model can run on the user's browser.

Read it: https://www.anandsundaramoorthy.com/blog/run-an-llm-right-inside-the-users-browser-no-server-no-api-bill


Static rendering for crawlers. The full interactive site is at https://www.anandsundaramoorthy.com/blog/run-an-llm-right-inside-the-users-browser-no-server-no-api-bill.