Skip to content
DopeSwagYolo

Gadgets & Wearables

What is on-device AI?

On-device AI means an AI model runs on the phone, laptop or wearable itself, using that device's own chips and memory. It does not send each request to servers in a remote data center.

Also known as: Local AI, Edge AI, On-device inference

Researched and fact-checked by AI, with no human review. 8 sources listed below. How we verify

Last updated

How it works

IEEE Spectrum reported in November 2025 that most people reach large language models through an online interface, with queries sent to a data center where the model runs. On-device AI performs that processing step, known as inference, on the user's own hardware. Google's Android documentation says its Gemini Nano model executes prompts locally, without server calls or a network connection.

Models that run on phones are far smaller than those in data centers. IEEE Spectrum reports that the label small AI is often applied to language models with at most a few billion parameters. Cutting-edge models can have more than a trillion, it reports. Many are made by pruning parameters from a larger model or by distillation, in which a small model is trained to mimic a big one. Apple said in a 2024 research post that its on-device language model had about 3 billion parameters, compressed to an average of 3.7 bits per weight. It generated 30 tokens per second on an iPhone 15 Pro, Apple said. Microsoft says chips called neural processing units (NPUs) run suitable models faster and extend battery life.

Why it matters

Google's documentation lists privacy, offline use and lower inference costs as benefits and says local processing removes network delay. Apple argues that data kept only on a user's device is not exposed to a central point of attack.

The trade-off is capability. Apple says more sophisticated requests need larger models in the cloud, so Apple Intelligence sends them to servers it calls Private Cloud Compute. Google notes that on-device speed depends on the device's hardware.

Where things stand in 2026

On June 8, 2026, Apple announced a new generation of Apple Foundation Models. It said they were built in collaboration with Google and its Gemini models and run both on device and on Private Cloud Compute servers. Apple said the features would arrive with iOS 27 and its other 2026 operating systems, which began rolling out on September 14, 2026. Google's Android documentation, last updated on September 8, 2026, lists on-device Gemini Nano features including summarization, proofreading, image description and speech recognition.

Not every phone can run such models. The research firm Counterpoint, cited by IEEE Spectrum in July 2026, estimated that slightly more than a third of smartphones shipped worldwide in 2025 could run generative AI. It projected 45 percent by the end of 2026.

Sources

Articles on Gadgets & Wearables