
your 8GB graphics card can run real local AI — you're probably picking the wrong model
Phi-4 Mini, Gemma 3, and Llama 3.2 all claim to fit — I checked which one actually leaves room to breathe.
Lena Fischer · Aug 4 · 6 min
Local models, AI rigs and the tools that run them · page 4

Phi-4 Mini, Gemma 3, and Llama 3.2 all claim to fit — I checked which one actually leaves room to breathe.
Lena Fischer · Aug 4 · 6 min

Four real paths to running a 70B-parameter model on your own hardware or by the hour — priced out with actual 2026 numbers.
Lena Fischer · Aug 4 · 8 min

Apple's own memory shortage just raised Mac Studio prices and gutted its top RAM tier — here's whether it's still worth buying over a GPU PC for running models at home.
Lena Fischer · Aug 4 · 8 min

Every model comes in a dozen quant sizes and it's paralyzing. Here's the simple truth: Q4_K_M is the default for almost everyone, and here's exactly when to go higher.
Lena Fischer · Aug 4 · 10 min

You don't need a graphics card to run a local AI. A modern CPU with enough RAM runs small models fine — just slowly. Here's what's realistic, which models to pick, and when you really do need a GPU.
Lena Fischer · Aug 4 · 10 min

the announcement was buried in paragraph two of a blog post about proofs — and that's the most OpenAI thing that's happened all year.
Lena Fischer · Aug 3 · 6 min

A developer wall-calibrated an Apple Silicon power meter against a smart plug and found the real number: fractions of a cent per million tokens.
Lena Fischer · Aug 3 · 6 min

Local models pile up fast and each one is gigabytes. Here's the handful of Ollama commands to list, update, remove, and inspect your models — and keep your disk under control.
Lena Fischer · Aug 3 · 10 min

A '30B' model that runs as fast as a 3B one? That's a Mixture-of-Experts model, and it changes the local-AI math. Here's what MoE means and why it matters for what you can run.
Lena Fischer · Aug 3 · 10 min

For math, logic, and step-by-step problems, 'reasoning' models that think out loud beat regular chatbots. Here's the best one to run locally, from an 8GB card to a 24GB rig.
Lena Fischer · Aug 3 · 10 min

Yes, you can run a real language model on a Raspberry Pi 5. It won't be fast, but for a tiny, private, always-on AI it's genuinely fun and useful. Here's how, and what to expect.
Lena Fischer · Aug 3 · 10 min

A vector database stores the searchable version of your documents. For local RAG, the right one is simpler than the enterprise options everyone benchmarks. Here's what to actually use.
Lena Fischer · Aug 3 · 10 min

Embeddings are what let a local AI search your documents. Pick the wrong one and retrieval is bad; pick the right one and RAG just works. Here's the best local embedding model, by need.
Lena Fischer · Aug 3 · 10 min

Running your local AI stack in Docker keeps it isolated, portable, and easy to reset. It's the tidy way to run Ollama and a web UI together. Here's the simple setup.
Lena Fischer · Aug 3 · 10 min

Those sliders in your local AI app aren't decoration. Temperature and top-p control how creative or focused the model is — and getting them right transforms your results. Explained simply.
Lena Fischer · Aug 3 · 10 min

Local models can be run 'uncensored' — without the refusals and lectures of cloud chatbots. Here's how it works, the legitimate reasons to want it, and how to do it responsibly.
Lena Fischer · Aug 3 · 10 min

Point your editor at a model on your own machine and get autocomplete, chat, and even an agent that edits files — free, unlimited, and with your code never leaving your PC. Here's the setup.
Lena Fischer · Aug 3 · 10 min

A local vision model reads images, describes photos, and pulls text out of screenshots — all offline. Some open ones now beat GPT-4o at OCR. Here's how to run one, in one command.
Lena Fischer · Aug 3 · 10 min