Decision one: where the model runs
This is the decision every other one depends on, and it is the one the competing articles skip, so the positions below were checked on 27 September 2026.
The model runs on the phone, or it runs on a server you call over the network, and that choice sets your bill, your latency, your privacy posture and half of your review risk. It is also a platform question for any mobile app development team, so if you have not settled native or hybrid mobile app development, settle that first.
On-device AI
On-device AI means that inference happens on the phone rather than on a server. There is no network call, no per-request bill, and the user's data does not leave the device.
What it costs you: app size, because the model or the framework ships with the app; battery and heat, because inference is real work on a real chip; and a device floor, because older phones cannot run it, which shrinks your addressable base.
What it gives you: it works on a train, latency is measured in milliseconds rather than round trips, and the awkward conversation about what you send to a third party does not happen, because you send nothing.
What Apple shipped in September 2025
Apple's newsroom release of 29 September 2025 introduced the Foundation Models framework, which gives developers access to the on-device large language model at the core of Apple Intelligence, roughly three billion parameters, available offline, and, in Apple's words, "all while using AI inference that is free of cost".
Read that last phrase as a founder would rather than as an engineer would. For a text feature that fits a small model, the per-call bill disappears from the business case, so a feature you could not justify at scale becomes free at scale on supported devices.
The limits are real, because it is a small model, so it summarises, classifies and extracts well, and it does not replace a large server model for hard reasoning. And it is iOS-only, which is a cross-platform problem for mobile app development rather than a technical one.
The Android side
Android has the same shape, where Gemini Nano runs on device through ML Kit's generative AI APIs for summarisation, proofreading and similar text work, all without a network call. Custom models run through LiteRT, the runtime formerly called TensorFlow Lite, and the iOS counterpart is Core ML.
The trade is the same in both ecosystems, a device floor in exchange for no inference bill and no user data leaving the phone.
Server inference
Server inference means that your app calls a model over the network each time the feature runs. That is the right answer when the task needs a large model, when the input is long, or when quality matters more than the extra second.
Three of its costs surprise mobile app development teams.
The bill scales with your success. Per-token pricing means a feature that gets popular gets expensive, so model that before you ship rather than after.
You added a network dependency to a mobile product. Phones lose signal in lifts, on trains and in basements, so every server AI feature needs an offline path, and most ship without one.
Someone else's rate limit is now your outage. When the provider throttles or has an incident, your AI feature is down while your support queue is not.
How to choose, in four questions
Does it need to work on a train? If yes, run it on-device, or give it an on-device fallback.
Does the data belong to the user in a way they would care about? Health records, messages, photos and documents all qualify, and if yes, prefer on-device, and read the next section before you decide otherwise.
Does one second of latency break the screen? Typeahead and camera features break under that delay, while a summary on a detail screen does not.
Would the bill scale with usage you cannot price? A free-tier feature running on server inference is a cost you do not control, and it grows fastest in the month your app does well.
Most real apps end up with both, running a small model on the device for the common path and a server model for the hard cases, with a written rule for which one runs.