AI Features in Mobile Apps: on-device vs cloud, and what is practical today
A grounded look at adding AI to a mobile app — which tasks belong on-device, which need a server, what it costs, how to handle latency and failure, and the privacy rules that apply.
“Add AI” has become a standard line in product briefs. Most of the resulting features are not hard to build — the difficulty is deciding what belongs on the device, what needs a server, what it costs per user, and what happens when the model is slow, wrong or unavailable.
This is a practical view of what is worth building into a mobile app today.
The first decision: on-device or cloud
| On-device | Cloud | |
|---|---|---|
| Latency | Immediate, no round trip | Network-dependent |
| Offline | Works | Does not |
| Cost | Free after shipping | Per request, forever |
| Privacy | Data never leaves the phone | Data leaves the phone |
| Capability | Narrow, specific tasks | Large, general models |
| Updating | Needs an app release, or model download | Change server-side any time |
| App size | Grows with the model | Unaffected |
The rule that holds up well: perception on-device, generation in the cloud. Recognising, classifying and detecting can usually run locally. Writing, summarising and open-ended reasoning generally cannot.
What works well on-device today
These are mature, fast, free to run, and work offline. In a Flutter app they are typically reached through platform ML kits or a lightweight on-device runtime.
- Text recognition (OCR). Scanning an invoice, meter reading, ID card or price tag. One of the highest-value features per unit of effort.
- Barcode and QR scanning. Effectively solved.
- Face and pose detection for camera framing and liveness cues — detection, not identification.
- Image labelling and object detection for auto-tagging or sorting uploads.
- Language identification and on-device translation for a limited set of languages.
- Speech to textusing the platform's own recognisers.
- Custom small classifiers — a compact model trained for one narrow job, such as sorting product photos into categories.
Newer devices also expose small on-device language models, but capability varies widely by hardware. Treat those as an enhancement on supported devices, with a server fallback, rather than as a baseline you can rely on.
What still belongs on a server
- Anything involving a large language model: summarising, drafting, answering questions.
- Retrieval over your own content, where an index lives server-side anyway.
- Image generation.
- Anything where you need to change the prompt or model without shipping an app update.
Never call a model provider directly from the app. An API key inside an APK can be extracted, and the bill is yours. Route calls through your own backend, which holds the key, authenticates the user, applies rate limits and can swap providers later.
Designing for latency and failure
A cloud model may take several seconds and can fail. Design for that from the start rather than adding a spinner afterwards.
- Stream where possible. Text appearing progressively feels far faster than the same text arriving at once.
- Never block the whole screen. Keep the rest of the app usable.
- Set a timeout with a real fallback.“Couldn't generate a summary — here is the full text” is a good outcome.
- Allow cancelling. Long generations on a slow connection need an exit.
- Cache aggressively. The same input should not be paid for twice.
The cost question, asked properly
Cloud AI is a recurring per-use cost, which is unusual for mobile features. Before building, work out the cost of a single interaction and multiply by realistic usage. A feature costing a fraction of a cent per call is fine at a thousand calls a day and a problem at a million.
Controls that matter:
- Per-user rate limits enforced on your backend, not in the app.
- Caching identical or near-identical requests.
- Choosing a smaller model for simple tasks; not everything needs the largest one.
- Capping input length — users will paste enormous documents.
- Alerting on unusual spend, since abuse shows up as a bill.
Being honest with users
Model output can be confidently wrong. How you present it is a product decision with real consequences.
- Label generated content. Users should know what came from a model.
- Keep the source reachable. A summary should link to the original.
- Avoid advice in regulated areas. Health, legal and financial guidance carries obligations well beyond a disclaimer in settings.
- Offer a way to report a bad result. It is your only feedback signal.
- Do not automate irreversible actions on model output alone. Suggest; let the user confirm.
Privacy and store requirements
- Declare it. If user content is sent to a third-party provider, that belongs in your privacy policy and in the store data declarations.
- Check retention terms. Know whether your provider retains inputs and whether they can be used for training, and configure accordingly.
- Get consent for sensitive data. Health, financial and biometric content needs explicit permission before it leaves the device.
- Watch app size with bundled models, and consider downloading them after install instead.
A sensible way to start
- Pick one task where AI removes real work — scanning a document instead of typing it, summarising a long thread, sorting uploads.
- Check whether an on-device capability already does it. Often one does.
- If it needs a server, route through your own backend and cost a single call.
- Build the failure path before the happy path.
- Ship it to a small group and look at how often the output is actually accepted.
The successful AI features in mobile apps are usually unglamorous: a camera that reads a number so the user does not have to type it. The ambitious ones tend to be expensive, slow and hard to trust.
Read next
Flutter State Management: setState, Provider, BLoC and Riverpod compared
A practical guide to choosing state management in Flutter. What each option is good at, where each breaks down…
Flutter Performance: how to find and fix jank, slow lists and memory problems
Why Flutter apps stutter and what to do about it — rebuild scope, list building, image sizing, expensive widge…