What is on-device AI?
On-device AI means an AI model runs on the phone, laptop or wearable itself, using that device's own chips and memory. It does not send each request to servers in a remote data center.
Also known as: Local AI, Edge AI, On-device inference
Researched and fact-checked by AI, with no human review. 8 sources listed below. How we verify
Last updated
How it works
IEEE Spectrum reported in November 2025 that most people reach large language models through an online interface, with queries sent to a data center where the model runs. On-device AI performs that processing step, known as inference, on the user's own hardware. Google's Android documentation says its Gemini Nano model executes prompts locally, without server calls or a network connection.
Models that run on phones are far smaller than those in data centers. IEEE Spectrum reports that the label small AI is often applied to language models with at most a few billion parameters. Cutting-edge models can have more than a trillion, it reports. Many are made by pruning parameters from a larger model or by distillation, in which a small model is trained to mimic a big one. Apple said in a 2024 research post that its on-device language model had about 3 billion parameters, compressed to an average of 3.7 bits per weight. It generated 30 tokens per second on an iPhone 15 Pro, Apple said. Microsoft says chips called neural processing units (NPUs) run suitable models faster and extend battery life.
Why it matters
Google's documentation lists privacy, offline use and lower inference costs as benefits and says local processing removes network delay. Apple argues that data kept only on a user's device is not exposed to a central point of attack.
The trade-off is capability. Apple says more sophisticated requests need larger models in the cloud, so Apple Intelligence sends them to servers it calls Private Cloud Compute. Google notes that on-device speed depends on the device's hardware.
Where things stand in 2026
On June 8, 2026, Apple announced a new generation of Apple Foundation Models. It said they were built in collaboration with Google and its Gemini models and run both on device and on Private Cloud Compute servers. Apple said the features would arrive with iOS 27 and its other 2026 operating systems, which began rolling out on September 14, 2026. Google's Android documentation, last updated on September 8, 2026, lists on-device Gemini Nano features including summarization, proofreading, image description and speech recognition.
Not every phone can run such models. The research firm Counterpoint, cited by IEEE Spectrum in July 2026, estimated that slightly more than a third of smartphones shipped worldwide in 2025 could run generative AI. It projected 45 percent by the end of 2026.
Sources
- Gemini Nano, Android Developers (Google)
- Introducing Apple’s On-Device and Server Foundation Models, Apple Machine Learning Research
- Private Cloud Compute: A new frontier for AI privacy in the cloud, Apple Security Research
- Apple Intelligence brings powerful AI capabilities into everyday experiences, Apple Newsroom
- Major updates for Apple’s software platforms are now available, Apple Newsroom
- Small AI Models Gain Traction Around the World, IEEE Spectrum
- Your Laptop Isn’t Ready for LLMs. That’s About to Change, IEEE Spectrum
- Copilot+ PCs developer guide, Microsoft Learn