Android Police ran a real local model on a real phone as a daily driver. it's not a GPU replacement — but it covers more than the GPU-shopping conversation usually admits.
Android Police spent real time running an AI model entirely on an Android phone — no cloud, no subscription, no GPU anywhere in the chain — and came away with a blunt verdict: 'There is no need to pay a fee to use AI for these kinds of tasks — a local LLM is good enough,' the outlet wrote in its writeup of the experiment. That's a real finding, not a hypothetical, and it matters for anyone currently eyeing an $800 GPU purchase just to run a small model at home.
what worked
Turning a wall of pasted text (phone reviews, in this case) into a clean pros/cons breakdown
Organizing scattered information into tables
Summarizing long articles or reviews
Pulling out key points from dense text
Explaining a complex topic in plain language
what didn't
Anything needing live web access — no browsing, so content has to be pasted in manually
Factual recall about past events — the model listed IPL cricket champions that have never existed as franchises
Staying concise — small local models tend to over-explain
Anything where being confidently wrong is worse than not answering at all
how fast is 'fast enough'
~40 tok/s
MLC Chat, Qwen3 1.7B, Galaxy S25 Ultra (NPU)
Snapdragon 8 Elite Hexagon NPU
~22 tok/s
MLC Chat, Phi-4 Mini, Galaxy S25 Ultra (NPU)
same device
~16 tok/s
PocketPal, Phi-4 Mini, Galaxy S25 Ultra (Vulkan)
CPU/GPU path, no NPU
$0
Cost per query
after the one-time model download
For context, a budget GPU built for this — something like the RTX 5060 Ti 8GB in our local-AI GPU rundown — pushes well past 100 tokens per second on a similarly sized model and doesn't slow down once you move past a couple of exchanges. The phone isn't competing on raw speed. What it's competing on is that it already exists in your pocket, cost nothing extra to use, and never sends what you type anywhere. We've also covered a phone-based AI agent that went further and started installing its own software, and a model built to run without a GPU at all — the no-GPU thread on this site keeps getting more real, not less.
On-device apps like PocketPal run open models directly on the phone — no data leaves the device. · Unsplash
when you actually need a real GPU
Rough capability ceiling by device
Phone, 1B–2B local modelLight tasks only
8GB GPU (e.g. RTX 5060 Ti 8GB)7B–8B models, comfortably
16GB+ GPU13B+ models, longer context
the honest verdict
Verdict
Our take
A local LLM on your phone is a legitimate free alternative to a cloud subscription for light, private, low-stakes text work — Android Police's own test backs that up. It is not a GPU replacement, and treating it as one will burn you the first time it invents a fact with total confidence. Use the phone for the tasks that don't need accuracy guarantees, and budget for a real card — see best-gpu-for-local-ai-2026 — for everything else.
Best for: Anyone who wants free, private, offline text help and doesn't need it to always be right.
Do I need a GPU to run any local AI at all?
No — small 1B–2B models run fine on a modern phone's CPU or NPU through apps like PocketPal or MLC Chat. You only need a GPU once you want bigger models or more reliable accuracy.
What model did Android Police actually run?
The E2B-sized version of Google's Gemma model, downloaded inside the PocketPal app directly from Hugging Face, with no separate setup required.
Is on-device phone AI private?
Yes — that's the actual selling point. Nothing you type leaves the phone, unlike a cloud subscription.
Can phone-based local AI replace ChatGPT or Gemini entirely?
Not yet, and the people testing it admit that themselves. It's a genuine complement for offline, low-stakes tasks, not a full replacement.
None of this is fixed in place — phone NPUs are getting faster every generation, and the gap between 'pocket AI' and 'desk AI' is narrowing quicker than the GPU-shopping conversation has caught up with. That's a prediction, not a fact, but it's a reasonable one. For now: use the phone for what it's actually good at, and don't feel talked into a GPU purchase for tasks a free app already handles.