Install and get started (llama.cpp)
llama.cpp provides high-performance GGUF inference; llama-server (or the server binary in your release) can expose an HTTP API for MiniTavern.
Best for: users who already download GGUF and want few dependencies and controllable flags. If you only want click-to-download, start with LM Studio or Ollama.
Get it
- Download a prebuilt package from official Releases, or compile per docs (CUDA / Metal / Vulkan).
- Confirm a server executable is present (
llama-server,server, etc. — follow your build). - Prepare a GGUF model — see Recommended models.
Quick check
Once the model runs locally, enable the HTTP service (next page). First make sure CLI load has no errors.