Find the right local LLM

Match models to your hardware and workload.

Basic

Hardware

Computer memory (RAM)

Main use

Pick one

This gets the most weight.

Priority

Balanced
FasterStronger answers
AdvancedTune results8,000 tokens
Also useful for

Secondary uses influence ranking; main use stays primary.

Includes prompt, pasted text and reply.

Power userModel controlsAutomatic model choices

Automatic picks the best-fitting option. Fixed choices can exclude models.

Estimated quantized size before runtime/context overhead.

Leave open for the best overall match.

Replaces the memory budget; speed estimate is disabled.

Matches

Example: 16 GB RAM.

Catalogue Sep 2026

Choose hardware and a main use.

How it works

How the finder works

01

Fit

Model, context and reserve memory are checked against your available memory.

02

Rank

Models are ranked for your selected workload and speed/quality preference.

03

Compare

Known GPUs also get a rough bandwidth-based speed range.

Common questions

Local LLM basics

What is a local LLM?

A language model that runs on your computer. Tools such as Ollama, LM Studio and llama.cpp can run compatible models locally.

How much VRAM do I need?

It depends on model size, quantization and context. Smaller models can use a few GB; larger ones may need 16, 24, 48 GB or more.

Does more RAM or VRAM make a model smarter?

No. More memory lets you run larger models or longer context; it does not improve the same model by itself.

Are speed figures benchmarks?

No. They are rough planning estimates. Runtime, drivers, CPU, settings and prompt length can change real performance.

Before relying on a result

Results are experimental estimates. Check the model publisher, licence and runtime before relying on them or buying hardware.

Methodology · Privacy · Terms & disclaimer