Skip to content

Install and get started (llama.cpp) ​

llama.cpp provides high-performance GGUF inference; llama-server (or the server binary in your release) can expose an HTTP API for MiniTavern.

Best for: users who already download GGUF and want few dependencies and controllable flags. If you only want click-to-download, start with LM Studio or Ollama.

Get it ​

  1. Download a prebuilt package from official Releases, or compile per docs (CUDA / Metal / Vulkan).
  2. Confirm a server executable is present (llama-server, server, etc. — follow your build).
  3. Prepare a GGUF model — see Recommended models.

Quick check ​

Once the model runs locally, enable the HTTP service (next page). First make sure CLI load has no errors.

Next ​

Start server → Configure in MiniTavern