how to run an AI that can see — local vision models for image analysis and OCR, on your own PC

A local vision model reads images, describes photos, and pulls text out of screenshots — all offline. Some open ones now beat GPT-4o at OCR. Here's…

Aliteq
Lena Fischer · AI & Local Compute Editor

The short version

Easiest: Ollama — pull a vision model, send an image, get an answer. 100% offline.

The short version

Best OCR (reading text): Qwen3-VL — 8B or 4B; tops document-reading benchmarks. MiniCPM-V beats GPT-4o on OCRBench.

The short version

Best general photo understanding: Llama 3.2 Vision 11B (8-16GB) or the reliable LLaVA 7B.

The short version

Smoothest first run: Gemma 3 4B — also multimodal, easy to start.

The short version

VRAM: 6-8GB runs the small ones (Qwen3-VL-4B/8B, LLaVA 7B); more for the 11B+.

The short version

Uses: describe images, OCR/read screenshots, alt-text, analyze charts & invoices — privately.

Aliteq

Read the full story

how to run an AI that can see — local vision models for image analysis and OCR, on your own PC

Read the full story on Aliteq