cheahjs/free-llm-api-resources
A list of free LLM inference resources accessible via API.
Backblaze Generative Media Hackathon
Build the next generation of AI media apps with Genblaze, stored on Backblaze B2. $10,000 in prizes.
Loading star history...
Use Cases & Benefits
- This repository provides a curated list of free and trial-based LLM inference APIs for easy access to various large language models.
- It covers popular LLMs like Llama, Gemini, Mistral, Qwen, and OpenAI GPT variants, with usage limits and access details for each provider.
- Strengths include extensive coverage of legitimate free LLM APIs and clear usage limits; limitations involve dependency on third-party service availability and quota restrictions.
- Organizations can leverage this resource to prototype or integrate LLM capabilities cost-effectively before committing to paid plans in production.
- Ideal for developers and researchers seeking no-cost or low-cost API access to diverse LLMs for experimentation, development, and small-scale applications.
About free-llm-api-resources
Free LLM API resources
This lists various services that provide free access or credits towards API-based LLM usage.
[!NOTE]
Please don't abuse these services, else we might lose them.
[!WARNING]
This list explicitly excludes any services that are not legitimate (eg reverse engineers an existing chatbot)
Free Providers
OpenRouter
Limits:
20 requests/minute
50 requests/day
1000 requests/day with $10 lifetime topup
Models share a common quota.
- DeepCoder 14B Preview
- DeepHermes 3 Llama 3 8B Preview
- DeepSeek R1
- DeepSeek R1 Distill Llama 70B
- DeepSeek R1 Distill Qwen 14B
- DeepSeek V3 0324
- Dolphin 3.0 Mistral 24B
- Dolphin 3.0 R1 Mistral 24B
- Gemma 2 9B Instruct
- Gemma 3 12B Instruct
- Gemma 3 27B Instruct
- Gemma 3 4B Instruct
- Kimi VL A3B Thinking
- Llama 3.1 405B Instruct
- Llama 3.1 Nemotron Ultra 253B v1
- Llama 3.2 3B Instruct
- Llama 3.3 70B Instruct
- Llama 4 Maverick
- Llama 4 Scout
- Mistral 7B Instruct
- Mistral Nemo
- Mistral Small 24B Instruct 2501
- Mistral Small 3.1 24B Instruct
- QwQ 32B ArliAI RpR v1
- Qwen 2.5 72B Instruct
- Qwen 2.5 VL 32B Instruct
- Qwen QwQ 32B
- Qwen2.5 Coder 32B Instruct
- Qwen2.5 VL 72B Instruct
- Reka Flash 3
- Shisa V2 Llama 3.3 70B
- cognitivecomputations/dolphin-mistral-24b-venice-edition:free
- deepseek/deepseek-chat-v3.1:free
- deepseek/deepseek-r1-0528-qwen3-8b:free
- deepseek/deepseek-r1-0528:free
- google/gemma-3n-e2b-it:free
- google/gemma-3n-e4b-it:free
- meta-llama/llama-3.3-8b-instruct:free
- microsoft/mai-ds-r1:free
- mistralai/devstral-small-2505:free
- mistralai/mistral-small-3.2-24b-instruct:free
- moonshotai/kimi-dev-72b:free
- moonshotai/kimi-k2:free
- nvidia/nemotron-nano-9b-v2:free
- openai/gpt-oss-120b:free
- openai/gpt-oss-20b:free
- qwen/qwen3-14b:free
- qwen/qwen3-235b-a22b:free
- qwen/qwen3-30b-a3b:free
- qwen/qwen3-4b:free
- qwen/qwen3-8b:free
- qwen/qwen3-coder:free
- tencent/hunyuan-a13b-instruct:free
- tngtech/deepseek-r1t-chimera:free
- tngtech/deepseek-r1t2-chimera:free
- z-ai/glm-4.5-air:free
Google AI Studio
Data is used for training when used outside of the UK/CH/EEA/EU.
| Model Name | Model Limits |
|---|---|
| Gemini 2.5 Pro | 3,000,000 tokens/day 125,000 tokens/minute 50 requests/day 2 requests/minute |
| Gemini 2.5 Flash | 250,000 tokens/minute 250 requests/day 10 requests/minute |
| Gemini 2.5 Flash-Lite | 250,000 tokens/minute 1,000 requests/day 15 requests/minute |
| Gemini 2.5 Flash Image Preview (Nano Banana) | |
| Gemini 2.0 Flash | 1,000,000 tokens/minute 200 requests/day 15 requests/minute |
| Gemini 2.0 Flash-Lite | 1,000,000 tokens/minute 200 requests/day 30 requests/minute |
| Gemini 2.0 Flash (Experimental) | 250,000 tokens/minute 50 requests/day 10 requests/minute |
| Gemini 1.5 Flash | 250,000 tokens/minute 50 requests/day 15 requests/minute |
| Gemini 1.5 Flash-8B | 250,000 tokens/minute 50 requests/day 15 requests/minute |
| LearnLM 2.0 Flash (Experimental) | 1,500 requests/day 15 requests/minute |
| Gemma 3 27B Instruct | 15,000 tokens/minute 14,400 requests/day 30 requests/minute |
| Gemma 3 12B Instruct | 15,000 tokens/minute 14,400 requests/day 30 requests/minute |
| Gemma 3 4B Instruct | 15,000 tokens/minute 14,400 requests/day 30 requests/minute |
| Gemma 3 1B Instruct | 15,000 tokens/minute 14,400 requests/day 30 requests/minute |
NVIDIA NIM
Phone number verification required. Models tend to be context window limited.
Limits: 40 requests/minute
Mistral (La Plateforme)
- Free tier (Experiment plan) requires opting into data training
- Requires phone number verification.
Limits (per-model): 1 request/second, 500,000 tokens/minute, 1,000,000,000 tokens/month
Mistral (Codestral)
- Currently free to use
- Monthly subscription based
- Requires phone number verification
Limits: 30 requests/minute, 2,000 requests/day
- Codestral
HuggingFace Inference Providers
HuggingFace Serverless Inference limited to models smaller than 10GB. Some popular models are supported even if they exceed 10GB.
Limits: $0.10/month in credits
- Various open models across supported providers
Vercel AI Gateway
Routes to various supported providers.
Limits: $5/month
Cerebras
| Model Name | Model Limits |
|---|---|
| gpt-oss-120b | 30 requests/minute 60,000 tokens/minute 900 requests/hour 1,000,000 tokens/hour 14,400 requests/day 1,000,000 tokens/day |
| Qwen 3 235B A22B Instruct | 30 requests/minute 60,000 tokens/minute 900 requests/hour 1,000,000 tokens/hour 14,400 requests/day 1,000,000 tokens/day |
| Qwen 3 235B A22B Thinking | 30 requests/minute 60,000 tokens/minute 900 requests/hour 1,000,000 tokens/hour 14,400 requests/day 1,000,000 tokens/day |
| Qwen 3 Coder 480B | 10 requests/minute 150,000 tokens/minute 100 requests/hour 1,000,000 tokens/hour 100 requests/day 1,000,000 tokens/day |
| Llama 3.3 70B | 30 requests/minute 64,000 tokens/minute 900 requests/hour 1,000,000 tokens/hour 14,400 requests/day 1,000,000 tokens/day |
| Qwen 3 32B | 30 requests/minute 64,000 tokens/minute 900 requests/hour 1,000,000 tokens/hour 14,400 requests/day 1,000,000 tokens/day |
| Llama 3.1 8B | 30 requests/minute 60,000 tokens/minute 900 requests/hour 1,000,000 tokens/hour 14,400 requests/day 1,000,000 tokens/day |
| Llama 4 Scout | 30 requests/minute 60,000 tokens/minute 900 requests/hour 1,000,000 tokens/hour 14,400 requests/day 1,000,000 tokens/day |
| Llama 4 Maverick | 30 requests/minute 60,000 tokens/minute 900 requests/hour 1,000,000 tokens/hour 14,400 requests/day 1,000,000 tokens/day |
Groq
| Model Name | Model Limits |
|---|---|
| Allam 2 7B | 7,000 requests/day 6,000 tokens/minute |
| DeepSeek R1 Distill Llama 70B | 1,000 requests/day 6,000 tokens/minute |
| Gemma 2 9B Instruct | 14,400 requests/day 15,000 tokens/minute |
| Llama 3.1 8B | 14,400 requests/day 6,000 tokens/minute |
| Llama 3.3 70B | 1,000 requests/day 12,000 tokens/minute |
| Llama 4 Maverick 17B 128E Instruct | 1,000 requests/day 6,000 tokens/minute |
| Llama 4 Scout Instruct | 1,000 requests/day 30,000 tokens/minute |
| Whisper Large v3 | 7,200 audio-seconds/minute 2,000 requests/day |
| Whisper Large v3 Turbo | 7,200 audio-seconds/minute 2,000 requests/day |
| groq/compound | 250 requests/day 70,000 tokens/minute |
| groq/compound-mini | 250 requests/day 70,000 tokens/minute |
| meta-llama/llama-guard-4-12b | 14,400 requests/day 15,000 tokens/minute |
| meta-llama/llama-prompt-guard-2-22m | |
| meta-llama/llama-prompt-guard-2-86m | |
| moonshotai/kimi-k2-instruct | 1,000 requests/day 10,000 tokens/minute |
| moonshotai/kimi-k2-instruct-0905 | 1,000 requests/day 10,000 tokens/minute |
| openai/gpt-oss-120b | 1,000 requests/day 8,000 tokens/minute |
| openai/gpt-oss-20b | 1,000 requests/day 8,000 tokens/minute |
| qwen/qwen3-32b | 1,000 requests/day 6,000 tokens/minute |
Together (Free)
Limits: Up to 60 requests/minute
Cohere
Limits:
20 requests/minute
1,000 requests/month
Models share a common quota.
- Command-A
- Command-R7B
- Command-R+
- Command-R
- Aya Expanse 8B
- Aya Expanse 32B
- Aya Vision 8B
- Aya Vision 32B
GitHub Models
Extremely restrictive input/output token limits.
Limits: Dependent on Copilot subscription tier (Free/Pro/Pro+/Business/Enterprise)
- AI21 Jamba 1.5 Large
- AI21 Jamba 1.5 Mini
- Codestral 25.01
- Cohere Command A
- Cohere Command R 08-2024
- Cohere Command R+ 08-2024
- Cohere Embed v3 English
- Cohere Embed v3 Multilingual
- DeepSeek-R1
- DeepSeek-R1-0528
- DeepSeek-V3-0324
- Grok 3
- Grok 3 Mini
- JAIS 30b Chat
- Llama 4 Maverick 17B 128E Instruct FP8
- Llama 4 Scout 17B 16E Instruct
- Llama-3.2-11B-Vision-Instruct
- Llama-3.2-90B-Vision-Instruct
- Llama-3.3-70B-Instruct
- MAI-DS-R1
- Meta-Llama-3.1-405B-Instruct
- Meta-Llama-3.1-8B-Instruct
- Ministral 3B
- Mistral Large 24.11
- Mistral Medium 3 (25.05)
- Mistral Nemo
- Mistral Small 3.1
- OpenAI GPT-4.1
- OpenAI GPT-4.1-mini
- OpenAI GPT-4.1-nano
- OpenAI GPT-4o
- OpenAI GPT-4o mini
- OpenAI Text Embedding 3 (large)
- OpenAI Text Embedding 3 (small)
- OpenAI gpt-5
- OpenAI gpt-5-chat (preview)
- OpenAI gpt-5-mini
- OpenAI gpt-5-nano
- OpenAI o1
- OpenAI o1-mini
- OpenAI o1-preview
- OpenAI o3
- OpenAI o3-mini
- OpenAI o4-mini
- Phi-4
- Phi-4-mini-instruct
- Phi-4-mini-reasoning
- Phi-4-multimodal-instruct
- Phi-4-reasoning
Cloudflare Workers AI
Limits: 10,000 neurons/day
- @cf/openai/gpt-oss-120b
- @cf/openai/gpt-oss-20b
- DeepSeek R1 Distill Qwen 32B
- Deepseek Coder 6.7B Base (AWQ)
- Deepseek Coder 6.7B Instruct (AWQ)
- Deepseek Math 7B Instruct
- Discolm German 7B v1 (AWQ)
- Falcom 7B Instruct
- Gemma 2B Instruct (LoRA)
- Gemma 3 12B Instruct
- Gemma 7B Instruct
- Gemma 7B Instruct (LoRA)
- Hermes 2 Pro Mistral 7B
- Llama 2 13B Chat (AWQ)
- Llama 2 7B Chat (FP16)
- Llama 2 7B Chat (INT8)
- Llama 2 7B Chat (LoRA)
- Llama 3 8B Instruct
- Llama 3 8B Instruct
- Llama 3 8B Instruct (AWQ)
- Llama 3.1 8B Instruct (AWQ)
- Llama 3.1 8B Instruct (FP8)
- Llama 3.2 11B Vision Instruct
- Llama 3.2 1B Instruct
- Llama 3.2 3B Instruct
- Llama 3.3 70B Instruct (FP8)
- Llama 4 Scout Instruct
- Llama Guard 3 8B
- LlamaGuard 7B (AWQ)
- Mistral 7B Instruct v0.1
- Mistral 7B Instruct v0.1 (AWQ)
- Mistral 7B Instruct v0.2
- Mistral 7B Instruct v0.2 (LoRA)
- Mistral Small 3.1 24B Instruct
- Neural Chat 7B v3.1 (AWQ)
- OpenChat 3.5 0106
- OpenHermes 2.5 Mistral 7B (AWQ)
- Phi-2
- Qwen 1.5 0.5B Chat
- Qwen 1.5 1.8B Chat
- Qwen 1.5 14B Chat (AWQ)
- Qwen 1.5 7B Chat (AWQ)
- Qwen 2.5 Coder 32B Instruct
- Qwen QwQ 32B
- SQLCoder 7B 2
- Starling LM 7B Beta
- TinyLlama 1.1B Chat v1.0
- Una Cybertron 7B v2 (BF16)
- Zephyr 7B Beta (AWQ)
Google Cloud Vertex AI
Very stringent payment verification for Google Cloud.
| Model Name | Model Limits |
|---|---|
| Llama 3.2 90B Vision Instruct | 30 requests/minute Free during preview |
| Llama 3.1 70B Instruct | 60 requests/minute Free during preview |
| Llama 3.1 8B Instruct | 60 requests/minute Free during preview |
Providers with trial credits
Together
Credits: $1 when you add a payment method
Models: Various open models
Fireworks
Credits: $1
Models: Various open models
Baseten
Credits: $30
Models: Any supported model - pay by compute time
Nebius
Credits: $1
Models: Various open models
Novita
Credits: $0.5 for 1 year
Models: Various open models
AI21
Credits: $10 for 3 months
Models: Jamba family of models
Upstage
Credits: $10 for 3 months
Models: Solar Pro/Mini
NLP Cloud
Credits: $15
Requirements: Phone number verification
Models: Various open models
Alibaba Cloud (International) Model Studio
Credits: 1 million tokens/model
Models: Various open and proprietary Qwen models
Modal
Credits: $5/month upon sign up, $30/month with payment method added
Models: Any supported model - pay by compute time
Inference.net
Credits: $1, $25 on responding to email survey
Models: Various open models
nCompass
Credits: $1
Models: Various open models
Hyperbolic
Credits: $1
Models:
- DeepSeek V3
- DeepSeek V3 0324
- Hermes 3 Llama 3.1 70B
- Llama 3 70B Instruct
- Llama 3.1 405B Base
- Llama 3.1 405B Base (FP8)
- Llama 3.1 405B Instruct
- Llama 3.1 70B Instruct
- Llama 3.1 8B Instruct
- Llama 3.2 3B Instruct
- Llama 3.3 70B Instruct
- Pixtral 12B (2409)
- Qwen QwQ 32B
- Qwen QwQ 32B Preview
- Qwen2.5 72B Instruct
- Qwen2.5 Coder 32B Instruct
- Qwen2.5 VL 72B Instruct
- Qwen2.5 VL 7B Instruct
- openai/gpt-oss-120b
- openai/gpt-oss-120b-turbo
- openai/gpt-oss-20b
- qwen/qwen3-235b-a22b-instruct-2507-fp8
- qwen/qwen3-coder-480b-a35b-instruct-fp8
- qwen/qwen3-next-80b-a3b-instruct
- qwen/qwen3-next-80b-a3b-thinking
SambaNova Cloud
Credits: $5 for 3 months
Models:
- E5-Mistral-7B-Instruct
- Llama 3.1 405B
- Llama 3.1 8B
- Llama 3.2 1B
- Llama 3.2 3B
- Llama 3.3 70B
- Llama-4-Maverick-17B-128E-Instruct
- Llama-4-Scout-17B-16E-Instruct
- Llama-Guard-3-8B
- Qwen/QwQ-32B
- Qwen/Qwen2-Audio-7B-Instruct
- Qwen/Qwen3-32B
- deepseek-ai/DeepSeek-R1
- deepseek-ai/DeepSeek-R1-0528
- deepseek-ai/DeepSeek-R1-Distill-Llama-70B
- deepseek-ai/DeepSeek-V3-0324
- deepseek-ai/DeepSeek-V3.1
- test-model
Scaleway Generative APIs
Credits: 1,000,000 free tokens
Models:
- BGE-Multilingual-Gemma2
- DeepSeek R1 Distill Llama 70B
- Gemma 3 27B Instruct
- Llama 3.1 8B Instruct
- Llama 3.3 70B Instruct
- Mistral Nemo 2407
- Mistral Small 3.1 24B Instruct 2503
- Pixtral 12B (2409)
- Qwen2.5 Coder 32B Instruct
- devstral-small-2505
- gpt-oss-120b
- mistral-small-3.2-24b-instruct-2506
- qwen3-235b-a22b-instruct-2507
- qwen3-coder-30b-a3b-instruct
- voxtral-small-24b-2507
Discover Repositories
Search across tracked repositories by name or description