Recommended models
Names and download steps differ by tool; the table is about size and purpose. Exact pull / filenames follow each tool’s page.
Pick size by hardware
| Rough size | RAM/VRAM rough guide | Best for |
|---|---|---|
| 1B–3B | ~4–8GB | Verify the path first, low-end machines |
| 7B–9B | ~8–16GB | Common sweet spot for roleplay |
| 12B–14B | 16GB+ | Better quality; speed depends on GPU |
| 32B+ | 24GB+ or heavy quant | High-end; phone waits longer too |
Smaller quants (Q4 / Q5 / IQ, etc.) use less resource with a quality trade-off. For GGUF, community favorites are often Q4_K_M / Q5_K_M.
Verify the path (prefer small models)
| Scenario | Ollama example | GGUF / other |
|---|---|---|
| Fastest path check | llama3.2:3b, qwen2.5:3b | Any ~3B instruct GGUF |
| Light Chinese | qwen2.5:3b / qwen2.5:7b | Qwen2.5 Instruct quant |
Get Test succeeding first, then move up in size.
Roleplay tips
- Prefer Instruct / Chat finetunes, not raw base completion models.
- Popular RP weights change over time; use what loads stably in your tool with enough context.
- Too-small context truncates long settings / lorebooks; for ~7B try 4k–8k first, then raise with VRAM.
- The model name in MiniTavern must exactly match the tool list (including tags like
:7b).
Common Ollama pulls
bash
ollama pull llama3.2:3b
ollama pull qwen2.5:7b
ollama pull qwen2.5:14bMore commands: Ollama · Download and pick models.
LM Studio / KoboldCPP / llama.cpp
- Download GGUF from Hugging Face etc. (check license and trust).
- Load the same file in the tool.
- After API / Server is on, copy the model id shown in the UI into MiniTavern.
Next
Tools overview → that tool’s “start server” chapter → Configure in MiniTavern