The standard shape of an AI feature is so familiar that most teams never question it. The user does something, the app sends text to a big hosted model, and the model sends text back. It works, and for plenty of problems it's the right call.

Over the last year or so, though, something has shifted. Small language models, with a few billion parameters or fewer, have got good enough at narrow tasks that running them on the user's own device is a real option. Operating systems and browsers now ship with built-in models that apps can call, and open-weight families like Gemma, Phi, Llama and Qwen all come in compact versions built for ordinary hardware. Whether you can run a model locally is settled. The interesting question is when you should.

What small models are good at

Let's be clear about the limits first. A small model won't write your quarterly strategy, work through a complicated contract or answer open-ended questions about the world as reliably as a frontier model. It's good at bounded, well-defined jobs.

It can classify and route: is this support message about billing, a bug or a cancellation, and is this document an invoice or a purchase order? It can pull the date, amount and supplier off a receipt, or turn a free-text address into proper fields. It can condense a long email thread, tidy up dictated notes or change the tone of a draft. It's fine for autocomplete and suggestions, where speed matters more than brilliance and a bad suggestion is easy to ignore. And it can spot and strip personal details before anything leaves the device.

That covers a surprising share of the AI features businesses ship. Plenty of them go to a frontier model out of habit, which is a bit like hiring a senior barrister to sort the mail.

Three reasons to go local

Privacy that doesn't rely on a contract

When the model runs on the device, the data never leaves it. That's a far simpler thing to explain to a privacy officer, a regulator or a nervous customer than "our vendor has agreed not to keep your data". For health, legal, HR and finance workflows it can decide whether a feature gets approved or sits in review for months.

Speed you can design around

A round trip to a hosted model can take hundreds of milliseconds before the first token arrives, and longer when the provider is busy. A local model has no network in the way, so it can respond quickly enough to feel like part of the interface rather than something you wait for. That opens up things that don't work at cloud speed, like live suggestions as you type, instant categorising as items are scanned, and features that keep working in a basement car park with no signal.

Costs that don't grow with usage

Hosted inference is a variable cost that grows with every user and every request, month after month. On-device inference runs on hardware your users already own. For high-volume, low-value jobs like tagging every photo or classifying every incoming message, moving the work onto the device can take a line off the cloud bill entirely.

The catches

Local models have problems of their own, and they bite teams who don't plan for them.

Devices vary. Not everyone has this year's flagship phone, and a model that flies on a new laptop may crawl on a five-year-old Android. You need a plan for devices that can't cope, which usually means falling back to a hosted model or hiding the feature.

Models are big downloads. Even a heavily quantised one can run from hundreds of megabytes to a few gigabytes. The models built into the platform avoid that, but then you don't control exactly which model you get or when it changes.

Testing gets harder. With a hosted model you're testing one thing. On devices, behaviour can shift subtly between hardware, OS versions and quantisation levels, so your evaluation suite has to cover the devices your users really have.

And inference burns battery. A feature that runs constantly in the background will flatten phones and get your app uninstalled.

What works: local first

The design we reach for most isn't local or cloud but local first. A small on-device model handles the quick, frequent, private work of classifying, extracting, redacting and suggesting. When a task is beyond it, say an open-ended question or a long, complicated document, the app hands off to a hosted frontier model, often after the local model has already stripped out the sensitive details.

Done well, users get instant answers most of the time, the inference bill gets smaller and more predictable, and privacy is easier to explain because most data never leaves the device.

The big design decision is where to draw that line, and you settle it with data rather than opinion. Build a test set from real tasks, run it through both the small and the large model, and measure where the small one's quality drops below what you'll accept. The line usually sits further out than people expect.

Got an AI feature that's expensive, slow or stuck in privacy review? It might be a good candidate for running on the device, and we're happy to take a look.