← All articles
Trends10 min read

AI Features in Mobile Apps: on-device vs cloud, and what is practical today

A grounded look at adding AI to a mobile app — which tasks belong on-device, which need a server, what it costs, how to handle latency and failure, and the privacy rules that apply.

AIMobileOn-device MLArchitecture

“Add AI” has become a standard line in product briefs. Most of the resulting features are not hard to build — the difficulty is deciding what belongs on the device, what needs a server, what it costs per user, and what happens when the model is slow, wrong or unavailable.

This is a practical view of what is worth building into a mobile app today.

The first decision: on-device or cloud

On-deviceCloud
LatencyImmediate, no round tripNetwork-dependent
OfflineWorksDoes not
CostFree after shippingPer request, forever
PrivacyData never leaves the phoneData leaves the phone
CapabilityNarrow, specific tasksLarge, general models
UpdatingNeeds an app release, or model downloadChange server-side any time
App sizeGrows with the modelUnaffected

The rule that holds up well: perception on-device, generation in the cloud. Recognising, classifying and detecting can usually run locally. Writing, summarising and open-ended reasoning generally cannot.

What works well on-device today

These are mature, fast, free to run, and work offline. In a Flutter app they are typically reached through platform ML kits or a lightweight on-device runtime.

  • Text recognition (OCR). Scanning an invoice, meter reading, ID card or price tag. One of the highest-value features per unit of effort.
  • Barcode and QR scanning. Effectively solved.
  • Face and pose detection for camera framing and liveness cues — detection, not identification.
  • Image labelling and object detection for auto-tagging or sorting uploads.
  • Language identification and on-device translation for a limited set of languages.
  • Speech to textusing the platform's own recognisers.
  • Custom small classifiers — a compact model trained for one narrow job, such as sorting product photos into categories.

Newer devices also expose small on-device language models, but capability varies widely by hardware. Treat those as an enhancement on supported devices, with a server fallback, rather than as a baseline you can rely on.

What still belongs on a server

  • Anything involving a large language model: summarising, drafting, answering questions.
  • Retrieval over your own content, where an index lives server-side anyway.
  • Image generation.
  • Anything where you need to change the prompt or model without shipping an app update.

Never call a model provider directly from the app. An API key inside an APK can be extracted, and the bill is yours. Route calls through your own backend, which holds the key, authenticates the user, applies rate limits and can swap providers later.

Designing for latency and failure

A cloud model may take several seconds and can fail. Design for that from the start rather than adding a spinner afterwards.

  • Stream where possible. Text appearing progressively feels far faster than the same text arriving at once.
  • Never block the whole screen. Keep the rest of the app usable.
  • Set a timeout with a real fallback.“Couldn't generate a summary — here is the full text” is a good outcome.
  • Allow cancelling. Long generations on a slow connection need an exit.
  • Cache aggressively. The same input should not be paid for twice.

The cost question, asked properly

Cloud AI is a recurring per-use cost, which is unusual for mobile features. Before building, work out the cost of a single interaction and multiply by realistic usage. A feature costing a fraction of a cent per call is fine at a thousand calls a day and a problem at a million.

Controls that matter:

  • Per-user rate limits enforced on your backend, not in the app.
  • Caching identical or near-identical requests.
  • Choosing a smaller model for simple tasks; not everything needs the largest one.
  • Capping input length — users will paste enormous documents.
  • Alerting on unusual spend, since abuse shows up as a bill.

Being honest with users

Model output can be confidently wrong. How you present it is a product decision with real consequences.

  • Label generated content. Users should know what came from a model.
  • Keep the source reachable. A summary should link to the original.
  • Avoid advice in regulated areas. Health, legal and financial guidance carries obligations well beyond a disclaimer in settings.
  • Offer a way to report a bad result. It is your only feedback signal.
  • Do not automate irreversible actions on model output alone. Suggest; let the user confirm.

Privacy and store requirements

  • Declare it. If user content is sent to a third-party provider, that belongs in your privacy policy and in the store data declarations.
  • Check retention terms. Know whether your provider retains inputs and whether they can be used for training, and configure accordingly.
  • Get consent for sensitive data. Health, financial and biometric content needs explicit permission before it leaves the device.
  • Watch app size with bundled models, and consider downloading them after install instead.

A sensible way to start

  1. Pick one task where AI removes real work — scanning a document instead of typing it, summarising a long thread, sorting uploads.
  2. Check whether an on-device capability already does it. Often one does.
  3. If it needs a server, route through your own backend and cost a single call.
  4. Build the failure path before the happy path.
  5. Ship it to a small group and look at how often the output is actually accepted.

The successful AI features in mobile apps are usually unglamorous: a camera that reads a number so the user does not have to type it. The ambitious ones tend to be expensive, slow and hard to trust.

Read next