Skip to content

Recommended models ​

Names and download steps differ by tool; the table is about size and purpose. Exact pull / filenames follow each tool’s page.

Pick size by hardware ​

Rough sizeRAM/VRAM rough guideBest for
1B–3B~4–8GBVerify the path first, low-end machines
7B–9B~8–16GBCommon sweet spot for roleplay
12B–14B16GB+Better quality; speed depends on GPU
32B+24GB+ or heavy quantHigh-end; phone waits longer too

Smaller quants (Q4 / Q5 / IQ, etc.) use less resource with a quality trade-off. For GGUF, community favorites are often Q4_K_M / Q5_K_M.

Verify the path (prefer small models) ​

ScenarioOllama exampleGGUF / other
Fastest path checkllama3.2:3b, qwen2.5:3bAny ~3B instruct GGUF
Light Chineseqwen2.5:3b / qwen2.5:7bQwen2.5 Instruct quant

Get Test succeeding first, then move up in size.

Roleplay tips ​

  • Prefer Instruct / Chat finetunes, not raw base completion models.
  • Popular RP weights change over time; use what loads stably in your tool with enough context.
  • Too-small context truncates long settings / lorebooks; for ~7B try 4k–8k first, then raise with VRAM.
  • The model name in MiniTavern must exactly match the tool list (including tags like :7b).

Common Ollama pulls ​

bash
ollama pull llama3.2:3b
ollama pull qwen2.5:7b
ollama pull qwen2.5:14b

More commands: Ollama · Download and pick models.

LM Studio / KoboldCPP / llama.cpp ​

  1. Download GGUF from Hugging Face etc. (check license and trust).
  2. Load the same file in the tool.
  3. After API / Server is on, copy the model id shown in the UI into MiniTavern.

Next ​

Tools overview → that tool’s “start server” chapter → Configure in MiniTavern