Fit
Model, context and reserve memory are checked against your available memory.
Match models to your hardware and workload.
Example: 16 GB RAM.
Choose hardware and a main use.
How it works
Model, context and reserve memory are checked against your available memory.
Models are ranked for your selected workload and speed/quality preference.
Known GPUs also get a rough bandwidth-based speed range.
Common questions
A language model that runs on your computer. Tools such as Ollama, LM Studio and llama.cpp can run compatible models locally.
It depends on model size, quantization and context. Smaller models can use a few GB; larger ones may need 16, 24, 48 GB or more.
No. More memory lets you run larger models or longer context; it does not improve the same model by itself.
No. They are rough planning estimates. Runtime, drivers, CPU, settings and prompt length can change real performance.
Results are experimental estimates. Check the model publisher, licence and runtime before relying on them or buying hardware.