Free cloud APIs
MiniTavern is BYOK (bring your own key): it does not include cloud credits. You register with a vendor or inference platform and paste your own API Key.
This page summarizes common free / trial quotas and how to fill them into the App. Quotas and policies change — always trust what each vendor’s console shows in real time.
Security note: an API Key is like a password. Store it only in the App’s on-device secure storage. Do not paste it into group chats, public screenshots, or public repos.
Fastest path
| Your situation | Suggestion | App platform |
|---|---|---|
| Mainland China, want to chat ASAP | Zhipu Flash free models or SiliconFlow free models | Custom (OpenAI) |
| One Key for text + image (multimodal) | Agnes AI | Custom (OpenAI) / Images |
One Key for GPT / Gemini *-free + image | AIHubMix | Custom (OpenAI) / Images |
| Coding-oriented gateway, limited Free text | OpenCode Zen | Custom (OpenAI) |
| Official Chinese LLM trial | Qwen / Alibaba Cloud Bailian | Qwen |
| Reasoning / code-oriented | DeepSeek | DeepSeek |
| Overseas account, multimodal text | Google Gemini | Gemini(Google) |
| One Key for many free open-source models | OpenRouter :free | OpenRouter |
| Not sure which free vendors exist | FreeLLM directory (awesome-freellm-apis) | See directory |
| Ultra-fast open-source inference | Groq | Custom (OpenAI) |
| Prefer no cloud | Run local models | Custom (OpenAI) |
General Key steps: Model link. Local deploy: Run local models · 30-minute quickstart.
How to fill in MiniTavern
Path: Settings → Models → Model link → New
- Pick the correct model platform (see “App platform” below).
- Paste the API Key; for custom platforms also fill Base URL.
- Select or type the model ID (must match vendor docs; case-sensitive).
- Tap Test → on success, set as Default.
| Vendor | App platform | Base URL (common) | Notes |
|---|---|---|---|
| Gemini | Gemini(Google) | Preset | Gemini protocol — do not fill as OpenAI |
| OpenRouter | OpenRouter | Preset | Free model IDs often end with :free |
| Qwen / Bailian | Qwen | Preset (intl switchable) | Use a general API Key, not a Coding Plan–only Key |
| DeepSeek | DeepSeek | Preset | |
| Zhipu / SiliconFlow / Agnes / AIHubMix / Zen / Groq / Mistral, etc. | Custom (OpenAI) | See each section | OpenAI-compatible gateways |
China-focused (prefer mainland direct access)
Zhipu AI (GLM)
Best for: long-term free chat (Flash series); friendly for Chinese roleplay.
| Item | Details |
|---|---|
| Site / Key | open.bigmodel.cn → Console → API Keys |
| Free tier | Permanent free models such as GLM-*-Flash; new users often get extra trial credit (see console) |
| Card required | Usually no |
| Base URL | https://open.bigmodel.cn/api/paas/v4 |
| Model examples | glm-4.5-flash, glm-4.7-flash (names follow the model plaza) |
| App | Custom (OpenAI) |
Steps: register → create API Key → in App choose Custom (OpenAI) → Base URL + Key → Flash-series model ID → test.
SiliconFlow
Best for: one Key across many open-source chat models; some image models also have free tiers.
| Item | Details |
|---|---|
| Site / Key | cloud.siliconflow.cn → Account → API keys |
| Free tier | Free models are not billed (name without Pro/ prefix); complete real-name verification for full free catalog |
| Rate limits | Fixed RPM/TPM per model; see official Rate Limits |
| Base URL | https://api.siliconflow.cn/v1 |
| Model examples | In the plaza, filter “Free”, e.g. Qwen/Qwen3-8B (list changes) |
| App | Custom (OpenAI) |
Note: the same name may have a free version and a paid Pro/... version — wrong ID bills you. Free-model usage on the bill should be 0.
Qwen / Alibaba Cloud Bailian
Best for: official Qwen series; MiniTavern has a Qwen preset.
| Item | Details |
|---|---|
| Activate | Bailian console (China North 2 / Beijing) — accept terms to auto-activate |
| Free tier | New-user free quota: most models ~1M tokens / model, ~90 days; quotas do not share across models |
| Key | Bailian → API-Key; use a general Key (Token Plan / Coding Plan–only Keys do not consume free quota) |
| App | Qwen (China preset; international site switchable) |
| Avoid charges | Enable “Stop when free quota is used up” so pay-as-you-go does not start after quota ends |
Official docs: new-user free quota. Calls after quota exhaustion or expiry incur charges.
Image gen (Wanxiang / Qwen-Image, etc.) may use new-user quota — in Settings → Image gen → Image gen link pick Qwen; whether it is free depends on that model’s quota bar in the console.
DeepSeek
Best for: reasoning, long text, high value; cheap after credits, but not permanently free.
| Item | Details |
|---|---|
| Site / Key | platform.deepseek.com → API Keys |
| Free tier | New users usually get signup credit / trial balance (amount and rules change; see console) |
| App | DeepSeek |
| Model examples | deepseek-v4-flash, deepseek-v4-pro (see official docs) |
After credits run out, usage is billed; requests fail with zero balance.
ModelScope API-Inference
Best for: Alibaba Cloud real-name already done; try community open-source inference.
| Item | Details |
|---|---|
| Token | modelscope.cn → Access token |
| Free tier | Registered users get daily API-Inference quotas (totals and per-model caps: official limits) |
| Base URL | https://api-inference.modelscope.cn/v1 |
| App | Custom (OpenAI) |
International / aggregators
Agnes AI
Best for: one free Key for text chat and image gen (both Flash free tier). Video is on the site; the App has no built-in video entry yet.
Full intro, model tables, and Key steps: Agnes AI.
| Item | Details |
|---|---|
| Site / Key | agnes-ai.com · platform.agnes-ai.com |
| Free tier | Free Access: Flash text / image callable for free (RPM limits) |
| Base URL | https://apihub.agnes-ai.com/v1 |
| Text example | agnes-2.5-flash |
| Image example | agnes-image-2.1-flash (image link → Custom (OpenAI Images)) |
| App | Custom (OpenAI) / Custom (OpenAI Images) |
OpenCode Zen
Best for: Free text models on the OpenCode curated gateway (e.g. mimo-v2.5-free). Most flagships are metered — watch auto-recharge.
Full guide: OpenCode Zen.
| Item | Details |
|---|---|
| Site / Key | opencode.ai/auth · docs |
| Free tier | Models marked Free on the pricing table (within caps); others are metered |
| Base URL | https://opencode.ai/zen/v1 |
| Text examples | mimo-v2.5-free, hy3-free, big-pickle |
| App | Custom (OpenAI) (prefer chat/completions-compatible IDs) |
AIHubMix
Best for: one Key for many platform-subsidized *-free models (GPT / Gemini / GLM, etc.) and GPT-Image-2-free image gen; usually no card required.
Full guide: AIHubMix.
| Item | Details |
|---|---|
| Site / Key | aihubmix.com · free models post |
| Free tier | Model IDs ending in -free; per-model RPM / daily caps |
| Base URL | https://aihubmix.com/v1 |
| Text examples | gpt-4o-free, gpt-5.5-free, coding-glm-5.1-free |
| Image example | gpt-image-2-free (Custom (OpenAI Images)) |
| App | Custom (OpenAI) / Custom (OpenAI Images) |
Google Gemini
Best for: multimodal text, long context; text free tier is relatively generous.
| Item | Details |
|---|---|
| Key | aistudio.google.com/app/apikey |
| Free tier | Some Gemini Flash / Lite text models have a Free tier (RPM / RPD as shown in AI Studio) |
| Card required | Text free tier usually no; latest image models mostly have no Free tier |
| Region | EU / UK / Switzerland may lack free tier |
| Privacy | Free-tier prompts may be used to improve products (see Google terms) |
| App | Gemini(Google) |
For chat, pick a text model (e.g. gemini-flash-latest). Do not pick IDs with tts or image: Test sends a text reply request; TTS only outputs AUDIO and errors with response modalities (TEXT) is not supported. Configure voice/image under those features.
If Test/send returns HTTP 404 with no longer available to new users, that model ID is retired or closed to new users (e.g. gemini-2.5-flash-lite). Refresh the model list and pick a live Flash / Lite (e.g. gemini-flash-latest). You do not need the Interactions API.
If send fails with HTTP 429 and generate_content_free_tier_requests (often limit: 20), free-tier RPM is exhausted — wait for Please retry in … then retry. See ai.dev/rate-limit; raise limits via billing in AI Studio. This is not “no money” or a monthly project spend cap.
Image gen: Gemini App web and Developer API quotas differ; API image pricing pages often mark Free tier unavailable. Chat on Gemini + local A1111/ComfyUI for images is a common split — see Image model deploy.
OpenRouter
Best for: one Key across a large open-source / third-party free pool; IDs often end with :free. Full lists, tables, and troubleshooting: OpenRouter.
| Item | Details |
|---|---|
| Key | openrouter.ai/keys |
| Free tier | :free / Free collections; default ~20 RPM, limited daily requests; topping up can raise free daily caps (official rules) |
| App | OpenRouter |
| Model examples | openrouter/free, meta-llama/llama-3.3-70b-instruct:free, openai/gpt-oss-20b:free |
| Collections | Free Models · summary openairouter.net/free-models |
Note: free upstreams may log prompts for training. On HTTP 429: Provider returned error, try another :free, wait, or BYOK.
Groq
Best for: very low latency open-source inference.
| Item | Details |
|---|---|
| Key | console.groq.com/keys |
| Free tier | Permanent Free tier (RPM / RPD / TPM per model; changes over time) |
| Card required | Usually no |
| Base URL | https://api.groq.com/openai/v1 |
| Model examples | llama-3.1-8b-instant, llama-3.3-70b-versatile, openai/gpt-oss-20b |
| App | Custom (OpenAI) |
Mistral AI
| Item | Details |
|---|---|
| Key | console.mistral.ai/api-keys |
| Free tier | Experiment / trial plans (monthly token scale per console); prompts may be used to improve products |
| Base URL | https://api.mistral.ai/v1 |
| App | Custom (OpenAI) |
Cerebras / SambaNova and other fast inference
| Platform | Key entry | Base URL | Notes |
|---|---|---|---|
| Cerebras | cloud.cerebras.ai | https://api.cerebras.ai/v1 | Free tier often has daily token caps; may require a payment method |
| SambaNova | cloud.sambanova.ai | https://api.sambanova.ai/v1 | Permanent free tier + occasional credits; strict rate limits |
App: Custom (OpenAI) for both; model names follow each console.
GitHub Models
| Item | Details |
|---|---|
| Entry | GitHub Models |
| Auth | GitHub Personal Access Token |
| Base URL | https://models.github.ai/inference |
| App | Custom (OpenAI) |
| Limits | Per-model RPM / RPD tiers — light trials |
NVIDIA NIM
| Item | Details |
|---|---|
| Entry | build.nvidia.com |
| Base URL | https://integrate.api.nvidia.com/v1 |
| App | Custom (OpenAI) (or Nvidia preset if restored in App later) |
Trial credits (not permanently free)
These are mostly time-limited credits — fine for evaluation, not as your only long-term backend:
| Vendor | Typical situation | App |
|---|---|---|
| OpenAI | Occasional console trial credit; main API has no stable permanent free tier | ChatGPT(OpenAI) |
| Anthropic Claude | Console starter credits, burn quickly | Claude(Anthropic) |
| DeepSeek | Signup credit above | DeepSeek |
| Alibaba Cloud Bailian | New-user token quota (~90 days) | Qwen |
When credits run out, reduce auto-charge risk (delete Key, enable stop-on-quota-end, or switch to free models / local).
Free image generation
| Path | Cost | Notes |
|---|---|---|
| Local A1111 / ComfyUI | Power + GPU | Unlimited images; phone → PC LAN |
| SiliconFlow free image models | Free models = ¥0 | Image link → Custom (OpenAI Images), same SiliconFlow Base URL …/v1, pick a free image ID |
| Agnes AI image Flash | Free Access | Image link → Custom (OpenAI Images), Base URL https://apihub.agnes-ai.com/v1, e.g. agnes-image-2.1-flash |
| AIHubMix | *-free image models | Image link → Custom (OpenAI Images), Base URL https://aihubmix.com/v1, e.g. gpt-image-2-free |
| Wanxiang / Qwen image | Depends on Bailian free quota for that model | Settings → Image gen → Image gen link → Qwen |
| OpenAI / most Gemini image APIs | Usually paid | Do not treat as “free image” without a stable Free tier |
Chat and image links can differ: e.g. chat on Zhipu Flash, images on home ComfyUI.
Tips and troubleshooting
- Test before chatting: on failure tap “View details” and check Error codes.
- 429 / rate limits: lighter model, shorter memory, off-peak, or a second free platform.
- Unexpected charges: check for
Pro/paid models, Bailian without stop-on-quota-end, or exhausted credits. - Privacy: free tiers often allow vendors to use logs to improve models; for sensitive content prefer local models.
- Quota changes: numbers here were common at research time; trust the vendor site. Broader index: FreeLLM directory (awesome-freellm-apis · freellm.net, unofficial, updated daily).
Related
- Model link — how to fill fields in the App
- Getting started
- FreeLLM directory — more free vendors
- Run local models — local language / image services
- Image model deploy — A1111 / ComfyUI