awesome-free-models
A curated list of free AI models, APIs, and tools you can use without paying a cent.
✅ All links verified live on August 1, 2026. 330+ URLs checked. All links working (some sites block automated checks but work in browsers).
Running AI shouldn't require a credit card. This list curates genuinely free models — open-weight models you can self-host, free API tiers from major providers, and tools to run everything locally.
Contents
- 🧠 Open-Weight Models — Downloadable model weights you can run on your own hardware
- 🔌 Free API Providers — Cloud APIs with generous free tiers for model inference
- 🖼️ Image & Video Generation — Open-weight and API-based visual generation models
- 🔀 Free API Routers — Unified gateways routing requests across multiple providers
- 💻 Local Inference Tools — Software to run models locally with full privacy
- 💬 AI Chatbot UIs — Self-hosted web interfaces for chatting with models
- 🎵 Audio & Speech Models — Open-weight TTS, STT, and voice generation models
- 🤖 AI Coding Assistants — IDE extensions and CLI tools for AI-assisted development
- 📝 Code Models — Models specialized for code generation and analysis
- 🧬 Embedding Models — Models for semantic search, RAG, and text representation
- 🔍 RAG & Vector Databases — Vector storage and retrieval for augmented generation
- 🧩 Agentic Frameworks — Frameworks for building autonomous AI agents and multi-agent systems
- 🔧 MCP Servers & Tools — Model Context Protocol servers connecting AI to external tools
- 🎛 Fine-tuning Tools — Tools for adapting models to your specific data
- ✨ Prompt Engineering Tools — Tools for testing, managing, and optimizing prompts
- 📊 LLM Evaluation & Observability — Tracing, evaluation, and monitoring for LLM apps
- 📊 Datasets — Open datasets for training, fine-tuning, and evaluation
- ☁ Model Hosting Platforms — Free cloud platforms for hosting and running models
- 📚 Learning Resources — Free courses, tutorials, and guides for AI engineering
- 🏆 Resources & Leaderboards — Benchmarks, leaderboards, and model discovery tools
- 👥 Communities — Discord servers, subreddits, and forums for discussion
🧠 Open-Weight Models
📅 Last checked: August 1, 2026
Notable open-weight models you can download and run on your own hardware.
- Llama 4 Scout / Maverick — Meta's latest MoE generation. Scout: 109B, 10M context. Maverick: 402B, 1M context. Native multimodal. [License]
- DeepSeek V4 Pro — Apr 2026. 1.6T MoE (49B active). SWE-bench Verified 80.6% (top open-weight). 1M context. MIT license.
- DeepSeek V4 — Core generation with extreme cost-efficiency. 1M context. MIT license.
- DeepSeek-V4-Flash — Apr 2026. Efficiency-focused variant. 284B total (13B active). 1M context. MIT license.
- Gemma 4 31B / 26B MoE / E4B / E2B — Fully permissive Apache 2.0. 256K context, native multimodal. New standard for open-weight.
- Inkling (Thinking Machines Lab) — Jul 2026. 975B MoE (41B active). Leading US open-weight model. Native multimodal (text, image, audio). Apache 2.0. 1M context.
- GLM-5.2 (Zhipu AI) — 744B MoE model optimized for autonomous coding and engineering tasks. 1M-token context. MIT license.
- LongCat-2.0 (ByteDance) — Large-scale open-weight model for heavy agentic coding. MIT license.
- MiniMax M3 — Frontier-tier 1M context, native multimodal + computer use. MSA architecture.
- Trinity (Arcee AI) — 400B parameter enterprise model. Apache 2.0.
- Step 3.7 Flash (StepFun) — May 2026. Apache 2.0. Native multimodal (image+video), strong agentic performance. Efficient enough for high-end local hardware.
- Kimi K3 (Moonshot AI) — Jul 2026. 2.8T-parameter MoE (896 experts, ~50B active). World's largest open-weight model. 1M context, native vision + video. #1 Frontend Code Arena. Modified MIT license. Weights released Jul 27.
- Kimi K2.6 (Moonshot AI) — Apr 2026. 1T-parameter MoE model. Modified MIT license. Exceptional coding (SWE-Bench ~54%) and multi-agent swarm orchestration.
- Qwen 3.6-35B-A3B — Apr 2026. MoE variant with only 3B active parameters. Extremely efficient for consumer hardware. Apache 2.0.
- InternLM 3 (Shanghai AI Lab) — Early 2026. Strong long-context reasoning and agentic performance. Competitive in open-weight benchmarks.
- MiMo-V2.5-Pro (Xiaomi) — Apr 2026. 1.02T-parameter MoE (42B active). Optimized for complex agentic tasks, coding, and long-context.
- Kimi K2.7 Code (Moonshot AI) — Jun 2026. 1T MoE specialized for long-running coding agents. +21.8% over K2.6 on coding benchmarks. Modified MIT.
- Nemotron 3 Super (NVIDIA) — May 2026. 120B total (12B active). 1M context. Published weights, data, recipes, and eval infra. NVIDIA Open Model License.
- Phi-4 14B (Microsoft) — 2025. Compact 14B dense model. Strong reasoning and code. MIT license. Excellent for on-device and small deployments.
- Bonsai 8B (PrismML) — Apr 2026. Groundbreaking 1-bit quantized model. Extremely efficient for edge and consumer hardware (Apple Silicon).
- Aether-7B-5Attn (VIDRAFT) — Jul 2026. 100% open foundation model (weights, data, code, logs). 7B MoE (~3B active) with heterogeneous attention. Apache 2.0.
- Mistral Large 3 (Mistral) — Jun 2026. 675B MoE (41B active). European multilingual flagship. Frontier-class reasoning, native multimodal. Apache 2.0.
- Mistral Small 3.1 (Mistral) — Mar 2025. Versatile 24B multimodal model. Strong text performance with native image understanding and 128K context. Apache 2.0.
- Mistral Small 4 (Mistral) — Mar 2026. Hybrid MoE (6.5B active params) unifying instruction, reasoning, and multimodal capabilities. Efficient frontier-class model. Apache 2.0.
- Command A+ (Cohere) — May 2026. Enterprise multimodal MoE optimized for sovereignty and multilingual RAG across 48 languages. Apache 2.0.
- Apertus 1.5 (ETH Zurich / EPFL) — Jul 2026. Fully open LLM (weights, data, training code). 8B and 70B with image understanding, thinking mode, and tool use. Apache 2.0.
- Hy3 (Tencent) — Jul 2026. 295B MoE (21B active). Strong reasoning and agentic performance. Competes with models 2-5x its size. Apache 2.0.
- Hermes 4 (NousResearch) — Feb 2026. Self-improving agentic model with closed-loop learning. Curates own memory and builds skills from experience. Apache 2.0.
- Snowflake Arctic — Apr 2024. Enterprise MoE model balancing high-quality performance with efficient training costs. Optimized for complex data operations. Apache 2.0.
- Falcon 3 (TII) — Dec 2024. Compact high-performance model with strong reasoning. Designed for efficient deployment on resource-constrained hardware. TII Falcon-LLM License 2.0.
- Apple OpenELM — Apr 2024. Family of efficient on-device SLMs using layer-wise attention scaling. Runs locally on Apple Silicon with full privacy. Apple Sample Code License.
🔌 Free API Providers
📅 Last checked: August 1, 2026
Providers offering free tiers to access models via API — no local hardware required.
- Google AI Studio — Most generous free tier. Access Gemini 2.5 Flash, Gemini 2.0 Flash, and other models. Generous rate limits for prototyping.
- OpenRouter — Aggregates 500+ models. Filter by "Free" to see models available at no cost. Includes experimental and subsidized open-weight models.
- AnyAPI — 400+ models with OpenAI-compatible API. Free tier: 100K tokens/day, unlimited users. Includes free and basic models. No credit card required.
- Groq — Ultra-fast inference. Free tier includes Llama, Gemma, Mixtral, Whisper models with generous daily rate limits.
- Hugging Face Inference Providers — Free tier for thousands of community models. Rate-limited but excellent for testing.
- NVIDIA NIM — Free API access to accelerated versions of Llama, Mistral, Gemma, and more on NVIDIA infrastructure.
- DeepInfra — Serverless inference. Free tier with daily rate limits for popular open-source models.
- Together AI — Free trial credits for new users. Fast inference on open-source models.
- Fireworks AI — Free tier for community models. Optimized for low latency.
- SiliconFlow — Rising platform with free access to many open-source models.
- Cloudflare Workers AI — Free tier for running select open-source models at the edge.
- Black Forest Labs — Free Flux 2 Dev and Flux Kontext Dev image generation via API. Rate-limited, no credit card required.
- Replicate — Free tier with limited credits for running open-source models.
- Poe (Quora) — Free tier with daily credits for GPT-4 mini, Claude instant, and community bots.
- Qwen Studio (Alibaba) — Free access to Qwen 3.6-Plus, Qwen 3.6-Max, and other Qwen models via web chat and API. 1M token context for agentic coding.
- Ollama Cloud — Free tier for running open-source models on Ollama's cloud infrastructure. Light usage with session limits (reset every 5 hours) and weekly limits. 1 concurrent model. Same
ollama runcommand as local. Prompt/response data never logged or trained on. - Mistral AI (La Plateforme) — Free API tier with access to Mistral Large, Mistral Nemo, Codestral and more. 1 req/s, 500k tokens/min. Requires phone verification and data usage opt-in.
- Cohere — Free evaluation API key for Command R, Command R+, Embed, and Rerank models. 20 req/min, 1,000 req/month.
- DeepSeek Platform — Free API credits for new users (5M tokens). Access to DeepSeek V4, DeepSeek-R1, and other models. Generous free allocation.
- GitHub Models — Free tier for GitHub users. Access GPT-4o, Llama 3.3, Mistral, and more with rate-limited playground and API.
- Hyperbolic — Open-access AI cloud with affordable inference. Free compute credits via referral program. Supports Llama, Qwen, DeepSeek, and other open models.
- Novita AI — Free credits for testing 100+ models including Llama, Qwen, DeepSeek, and Mistral. OpenAI-compatible API with competitive pricing beyond the free tier.
- Anakin.ai — 30 daily free credits for accessing multiple AI models. Web chat interface and API access. Supports GPT-4, Claude, and open-weight models.
- Nebius AI — $100 free credits for new users. AI Studio with access to Llama, Qwen, DeepSeek, and other open-weight models. Fast inference on NVIDIA H100 infrastructure.
- Fal.ai — Free starter credits for AI inference. Fast, serverless platform supporting Llama, Flux, and Stable Diffusion models. Pay-as-you-go beyond free tier.
- Vercel AI Gateway — $5/month free credits for the AI Gateway. Proxy and cache requests across multiple LLM providers. SDK is open-source and free.
- AI21 Labs — $10 trial credits for accessing Jamba 1.5, Jamba 1.6, and other AI21 models. Valid for 3 months. Requires account sign-up.
- Amazon Bedrock — $200 AWS credits for new customers. Access to Llama, Mistral, Claude, Titan, and other foundation models via API.
- Microsoft Foundry (Azure) — $200 free trial credits (30 days). Access GPT-4o, Llama, Mistral, Phi, and other models via Azure's unified AI platform.
- RunPod — Free credits for serverless GPU inference. Deploy open-weight models as serverless endpoints. Supports Llama, Qwen, DeepSeek, and more.
- Cerebras — Free tier with ultra-fast inference on Llama, Gemma, and Mistral models. No credit card required. Note: Model availability fluctuates.
- BazaarLink — Free OpenAI-compatible API with
auto:freerouting to zero-cost models. No credit card, no trial expiry. 10 RPM, 130 req/day. - Kimi API (Moonshot) — Free tier for new accounts with access to Kimi K2.5 (128K context). Also free via NVIDIA NIM. OpenAI-compatible.
- Alibaba DashScope — Free tier for Qwen models. 1M tokens/month. OpenAI-compatible API.
- SambaNova Cloud — Free tier with $5 credits (30-day). Fast RDU inference for Llama 3.1 405B, Llama 3.3 70B, DeepSeek V3.1/V3.2, Qwen 2.5, gpt-oss-120b. 20 RPM, 200K tokens/day. No credit card required.
- OVHcloud AI Endpoints — EU-hosted, GDPR-compliant free tier. No registration required for anonymous tier. Models: Qwen, Mistral, Llama, DeepSeek, gpt-oss-120B, embeddings, image generation. 12 RPM. OpenAI-compatible.
- Chutes.ai — Community-powered free GPU inference for open-source models. DeepSeek-R1, Llama 3.1 70B, Qwen 2.5 72B. OpenAI-compatible API. No credit card.
- ModelScope — Chinese platform with 50+ free open models. Qwen, DeepSeek, GLM, and more. No credit card required.
- Z.ai (Zhipu AI) — Free tier for GLM models including GLM-4.5, GLM-4V. No credit card required.
- LongCat AI — Free API for LongCat open-weight models. One-time 10M token grant after signup + KYC. Free cached tokens. OpenAI-compatible. MIT license.
- Coze (ByteDance) — Free bot-building platform with API access to GPT-4o, Gemini 1.5 Pro, and other models. No credit card required. OpenAI-compatible.
- Free.ai — 400+ AI tools via a single OpenAI-compatible API. Free tier: 30K tokens/day, no credit card. Chat, image, video, music, voice, OCR, translation.
- Requesty — Free AI API with 200 requests/day. Works with Claude Code, Cline, Cursor. No credit card. OpenAI-compatible.
- AINative Studio — 84+ models (Llama, DeepSeek, Mistral, Qwen). Free tier: 10M tokens/month. No credit card required. 60 RPM.
- CloudCode.ONE — Free tier for coding agents. Powered by GLM-4.7-Flash. OpenAI and Anthropic-compatible API. No credit card required.
- ZeroLimitAI — Free OpenAI-compatible API with
model: "auto"routing to the best free model. No credit card, no trial expiry. Lifetime free tier available. - Chat Oripe — 2M free tokens/month. OpenAI-compatible API with GPT-4 and Claude access. No credit card required.
- FreeTheAi (da-jb) — Open-source. Free AI API via Discord signup. No daily cap, 30 RPM. OpenAI-compatible, image and video generation.
- OpenCode Zen — Curated AI gateway with 7 free models (DeepSeek V4 Flash Free, MiMo-V2.5 Free, Nemotron 3 Ultra Free, Big Pickle, Qwen 3.6 Plus Free, MiniMax M3 Free, North Mini Code Free). OpenAI-compatible API. No credit card required.
🖼️ Image & Video Generation
📅 Last checked: August 1, 2026
Free, open-weight image and video generation models — run locally or via free APIs.
- FLUX.2-dev (Black Forest Labs) — 2.6K★. 32B rectified flow transformer. SOTA open T2I, single/multi-reference editing, in/out-painting. Updated VAE. FLUX.1-dev Non-Commercial License.
- ERNIE-Image / ERNIE-Image-Turbo (Baidu) — 0.5K★. 8B DiT SOTA among open-weight models. Strong text rendering, layout control. Turbo: 8-step generation. Apache 2.0.
- Z-Image (Tongyi Lab / Alibaba) — 11.8K★. Open-weight T2I with strong GenEval scores. Z-Image-Turbo for 4-step generation. Apache 2.0.
- Pollinations.ai — Free image generation API. No API key or signup needed. Text-to-image, image-to-image. OpenAI-compatible. Integrates with ComfyUI.
- OpenImageGen (Hugging Face) — Free, open-source image generation playground. Supports multiple community models via diffusers. Apache 2.0. Note: Requires Hugging Face login.
- ComfyUI — 123K★. Node-based image and video generation UI. Run FLUX, Stable Diffusion, and more locally. GPL-3.0.
🔀 Free API Routers
📅 Last checked: August 1, 2026
Open-source tools that route requests across multiple AI providers — unified API, automatic failover, and cost optimization.
- 9Router — Open-source gateway connecting 40+ providers with RTK token compression (2-4x reduction). One API key for all services. MIT license. GitHub
- OmniRoute — Full-stack AI gateway with 250+ providers, 90+ free. TypeScript, runs on Web/Desktop/Android. Prompt compression, 3-level proxy for geo restrictions. GitHub
- LiteLLM — Python-based proxy unifying 100+ LLMs behind a single API. Spend tracking, virtual keys, production-ready. MIT license. GitHub
- Portkey AI Gateway — Production guardrails and routing for AI apps. Hybrid open-source (community) and managed (enterprise) tiers. GitHub
💻 Local Inference Tools
📅 Last checked: August 1, 2026
Run models on your own machine — no API keys needed, full privacy.
- Ollama — The easiest way to run local LLMs. One command to download and run any model. macOS, Linux, Windows. GitHub
- LM Studio — Polished desktop GUI. Browse, download, and chat with models. Built-in model browser and local API server.
- llama.cpp — High-performance C++ inference engine. Runs on CPU and GPU. Supports GGUF quantization. Powers most other local tools.
- Jan — Open-source ChatGPT alternative for desktop. Built-in model downloader, local API server. GitHub
- GPT4All — ⚠️ Unmaintained since Feb 2025. Privacy-focused local chatbot. Runs on consumer hardware. Built-in model browser. GitHub
- text-generation-webui (Oobabooga) — Feature-rich web UI. Supports multiple backends (Transformers, llama.cpp, ExLlama, AutoGPTQ).
- LocalAI — Drop-in OpenAI API replacement. Run models locally with an OpenAI-compatible API. GitHub
- KoboldCPP — Single-file executable for running GGUF models. Focused on story generation but general-purpose.
- llamafile (Mozilla) — Distributable single-file executables that run LLMs. No installation needed.
- vLLM — High-throughput production inference engine. Uses PagedAttention for efficient serving.
- SGLang — Fast inference framework with structured generation and RadixAttention.
- TensorRT-LLM (NVIDIA) — NVIDIA's optimized inference engine. Best performance on NVIDIA GPUs.
- ExLlamaV3 — Optimized inference for Llama-family models. Successor to ExLlamaV2. Fastest option for single-GPU inference.
- Aphrodite Engine — High-performance LLM serving engine with advanced quantization support.
- TabbyAPI — Lightweight, fast OpenAI-compatible API server for ExLlamaV2.
- LlamaEdge — Lightweight inference framework for edge devices. OpenAI-compatible API for open-source models. Runs on WasmEdge for portability. GitHub
- MLC LLM — Universal deployment engine by UW/SJTU. Runs LLMs on any hardware — laptops, phones, browsers. OpenAI-compatible API.
- WebLLM — In-browser LLM inference via WebGPU. Runs models directly in your browser with zero setup. No server needed.
- FastChat (LMSYS) — Open platform for training, serving, and evaluating LLMs. Provides OpenAI-compatible API and web UI for local models.
- Hugging Face TGI — 10.9K★. Production-grade serving toolkit for large language models. Note: Archived by Hugging Face (Mar 2026). Consider vLLM, SGLang, or TGI forks for active development.
- DeepSpeed (Microsoft) — Deep learning optimization library with inference acceleration. Enables running larger models on limited hardware through ZeRO optimization.
- AirLLM — Run large models (70B+) on consumer hardware with limited memory. Loads models layer-by-layer for extreme memory efficiency. Actively maintained (last push Jul 2026).
- AI Toolkit for VS Code (Microsoft) — VS Code extension to browse, test, fine-tune, and deploy models locally. Integrates ONNX and llama.cpp.
- Ollama Grid Search — Desktop utility for systematic model evaluation. Test multiple models, prompts, and inference parameters side-by-side via a Rust/React GUI.
- oMLX — 18.2K★. LLM inference server for Apple Silicon with continuous batching, tiered KV caching (hot RAM + cold SSD), and macOS menu bar app. OpenAI and Anthropic compatible. Apache 2.0.
- MTPLX — 1.1K★. Native MTP speculative decoding on Apple Silicon — ~2x faster decode with no external drafter. Mac app + CLI, OpenAI/Anthropic compatible server. Auto-tunes draft depth per machine. Apache 2.0.
💬 AI Chatbot UIs
📅 Last checked: August 1, 2026
Free, open-source web interfaces for chatting with AI models — self-host or use hosted versions.
- Open WebUI — Feature-rich ChatGPT-like interface for Ollama and OpenAI-compatible backends. RAG, image generation, multi-user. GitHub
- LibreChat — Open-source ChatGPT clone supporting 40+ providers, multi-user, plugins, and RAG. Note: Company acquired by ClickHouse; repo actively maintained (last push Aug 2026). GitHub
- AnythingLLM — All-in-one desktop app for chatting with documents and models. Built-in RAG pipeline. GitHub
- Big-AGI — Feature-rich AI chat with personas, multi-model support, voice, and code execution. GitHub
- Lobe Chat — Multi-agent orchestration platform with plugin system and multi-provider support. GitHub
🎵 Audio & Speech Models
📅 Last checked: August 1, 2026
Free, open-weight text-to-speech (TTS), speech-to-text (STT), and voice generation models you can run locally.
- Qwen3-TTS (Alibaba) — 12.5K★. Voice cloning, voice design, 10 languages. Streaming support with 97ms TTFB. 0.6B/1.7B. Apache 2.0.
- Chatterbox (Resemble AI) — 25.5K★. SOTA open-source TTS. Multilingual V3 (23+ languages, 0.5B). Turbo: 350M for low-latency agents. Paralinguistic tags. MIT.
- MOSS-TTS Family (MOSI.AI/OpenMOSS) — 3.9K★. 8B flagship + 100M Nano (CPU). Voice cloning, dialogue generation, sound effects, realtime streaming. Apache 2.0.
- Orpheus-TTS (Canopy Labs) — 6.2K★. Llama-3b backbone, human-like speech, zero-shot voice cloning, emotion tags. ~200ms streaming latency. Apache 2.0.
- NeuTTS (Neuphonic) — 6K★. On-device TTS with instant voice cloning. GGUF quantized for CPU/mobile. 120M Nano and 360M Air variants. Apache 2.0.
- Faster-Whisper — 25K★. CTranslate2-based Whisper for 4x faster transcription. MIT.
🤖 AI Coding Assistants
📅 Last checked: August 1, 2026
Free tools that integrate AI into your development workflow.
- Continue.dev — Company acquired by Cursor; repo actively maintained (last push Aug 2026). Open-source AI code assistant for VS Code and JetBrains. GitHub
- Aider — AI pair programming in the terminal. Edits code in your local git repo. Supports GPT, Claude, and local models. GitHub
- Gemini CLI (Google) — Jul 2026. Open-source terminal agent with generous free Gemini quota. Supports agentic coding workflows.
- Kilo Code — 2026. VS Code/JetBrains agentic coding extension with model-agnostic support and Plan/Act oversight.
- Tabby — Self-hosted AI coding assistant with no dependency on external services. GitHub
- Cody (Sourcegraph) — Free tier for individuals. Chat, autocomplete, and commands with codebase context.
- Llama Coder (Nutlope) — Free AI code generation tool. Generate entire apps from prompts.
- Bolt.new (StackBlitz) — Free tier for AI-powered full-stack web app development in browser.
- Claude Code (Anthropic) — Terminal-based AI coding assistant. Most features require a Claude subscription or API credits. Limited free usage via terminal CLI.
- Cursor 3 — Apr 2026. AI-native code editor with deep model integration and agentic features. Free tier available.
- CodeBuff — CLI-based AI coding assistant that understands entire codebases. Multi-agent architecture, works with any model provider through natural language instructions.
- Pi — Open-source terminal AI coding agent with a unified multi-provider API. Model-agnostic, supports OpenAI, Anthropic, Google, and any OpenAI-compatible endpoint. Extensible plugin architecture. GitHub
- Cline — Popular autonomous VS Code agent. Creates/edits files, runs terminal commands, browses web. Open-source, BYOK (bring your own API key). GitHub
- OpenHands — Autonomous AI software engineer. Navigates file systems, runs shell commands, tests code in browser. Self-hostable. GitHub
- Goose — Open-source CLI agent for complex software engineering tasks. Extensible plugin system. Built by Block/Square. GitHub
- Qwen Code — 2025. Open-source terminal AI coding agent with 26K+ stars. Multi-protocol (OpenAI, Anthropic, Gemini, Qwen). Auto-memory, sub-agents, agent teams, MCP. Apache 2.0.
- CodeWhale — 2026. Terminal coding agent with 40K+ stars. 30+ providers, local models via Ollama/vLLM. TUI, headless mode, web UI. MIT license.
- SideCar — 2026. Free, self-hosted VS Code agent extension. Full agent loop, local Ollama models, inline completions, MCP. Drop-in for Copilot/Claude Code. MIT.
- nanobot (HKUDS) — 2026. Open-source, ultra-lightweight personal AI agent with WebUI, chat channels, MCP, memory, and scheduling. 46K★. MIT.
- MiMoCode (Xiaomi) — Jun 2026. Terminal-native coding agent with persistent memory, subagent orchestration, and goal-driven autonomous loops. 12.4K★. MIT.
📝 Code Models
📅 Last checked: August 1, 2026
Specialized for code generation, completion, and analysis.
- MAI-Code-1-Flash (Microsoft) — Jun 2026. Microsoft's open-weight coding model for lowering infrastructure costs.
- DeepSeek Coder — State-of-the-art open-weight code generation. DeepSeek's coder series leads SWE-bench. MIT license.
- Qwen2.5-Coder (Alibaba) — Highly capable code model series (1.5B–32B). Excellent balance of speed and quality. Apache 2.0.
- Codestral (Mistral) — Mistral's dedicated code generation model — fill-in-the-middle, completion, and instruction.
- CodeGemma (Google) — Google's Gemma architecture fine-tuned for code completion and instruction. Apache 2.0.
- StarCoder2 (BigCode) — Transparently trained code model covering 619 languages. OpenRAIL-M license.
- Yi-Coder (01.AI) — Efficient coding model with strong long-context understanding. Yi License (Apache 2.0 compatible).
- Granite Code (IBM) — IBM's enterprise-grade code model, available in multiple sizes. Apache 2.0.
- Phi-4-mini (Microsoft) — Lightweight model optimized for reasoning and code. Punches above its weight class. MIT license.
- Qwen3-Coder-Next (Alibaba) — Early 2026. Latest generation of Qwen's code series. Strong reasoning and long-context coding capabilities. Apache 2.0.
- CodeLlama (Meta) — Aug 2023. Llama 2-based code generation pioneer. Supports infilling, completion, and instruction. Llama 2 Community License.
- WizardCoder (WizardLM) — 2023. Evol-Instruct fine-tuned for complex coding tasks. Strong general code generation performance. Apache 2.0.
- OpenCodeInterpreter — 2024. Integrates execution feedback to iteratively improve generated code. Bridges generation and execution. Apache 2.0.
- Stable Code 3B (Stability AI) — Aug 2023. Lightweight 3B code model optimized for fill-in-the-middle. Efficient for local autocompletion. StabilityAI license.
- CodeGeeX2 (THUDM) — 2023. Multilingual code model supporting 20+ languages. Strong in both Chinese and English code tasks. Apache 2.0.
- CodeT5+ (Salesforce) — 2023. Encoder-decoder architecture unifying code generation, completion, and understanding. BSD-3 license.
- SantaCoder (BigCode) — 2023. Light 1.1B model specialized for Python, Java, and JavaScript. Fast and efficient for IDE integration.
🧬 Embedding Models
📅 Last checked: August 1, 2026
Free, open-weight embedding and reranker models for semantic search, RAG, and text representation.
- Qwen3-Embedding (Alibaba) — 2K★. #1 on MTEB multilingual leaderboard. Sizes: 0.6B/4B/8B. 32K context, MRL support, instruction-aware. Includes reranker models. Apache 2.0.
- BGE-M3 (BAAI) — Multi-lingual (100+ languages), multi-functionality (dense, sparse, colbert), multi-granularity (8K tokens). MIT.
- FlagEmbedding (BAAI) — 12K★. Framework and model zoo: BGE series, BGE-VL (multimodal), bge-en-icl, bge-multilingual-gemma2 (9B multilingual SOTA). MIT.
- nomic-embed-text-v2 (Nomic AI) — 1.5B MoE embedding model. 8192 context. Matches or exceeds OpenAI text-embedding-3-small. Apache 2.0.
- mxbai-embed-large-v1 — 0.3B lightweight embedding. Top of MTEB among sub-0.5B models. Apache 2.0.
🔍 RAG & Vector Databases
📅 Last checked: August 1, 2026
Free tools for building retrieval-augmented generation pipelines — vector storage, embedding search, and document retrieval.
- Chroma — AI-native open-source embedding database. Runs in-process, no GPU needed. GitHub
- Qdrant — High-performance vector search engine. Free tier on Qdrant Cloud or self-host via Docker. GitHub
- pgvector — Vector similarity search inside PostgreSQL. Free if you already run Postgres.
- LanceDB — Developer-friendly vector database built on Lance columnar format. Runs locally, no server needed. GitHub
- Weaviate — Open-source vector database. Free sandbox tier on Weaviate Cloud. GitHub
- Milvus (Zilliz) — Cloud-native vector database. Free tier on Zilliz Cloud or self-host. GitHub
- txtai — AI-powered semantic search and RAG in a single Python package. GitHub
- R2R (SciPhi) — Production-ready RAG engine with API, user management, and observability. Note: Last release Jun 2025. Consider alternatives like Dify or LangGraph.
- Docling (IBM) — Document understanding and conversion for RAG pipelines. Extracts PDFs, images, and more. GitHub
- Unstructured.io — Preprocessing toolkit for documents (PDF, HTML, Word) for RAG pipelines. Free tier available.
🧩 Agentic Frameworks
📅 Last checked: August 1, 2026
Free, open-source frameworks for building AI agents and multi-agent systems.
- LangGraph (LangChain) — Low-level framework for building stateful, multi-agent applications. GitHub
- CrewAI — Multi-agent framework for orchestrating specialized AI agents to work together. GitHub
- AutoGen (Microsoft) — Extensible framework for building multi-agent conversations. Note: In maintenance mode (last push Apr 2026). GitHub
- Agno (formerly Phidata) — Full-stack AI framework for building multimodal agents with memory, knowledge, and tools. GitHub
- PydanticAI — Agent framework by Pydantic with type-safe outputs and dependency injection. GitHub
- Mastra — TypeScript framework for building AI applications and agent workflows. GitHub
- OpenAI Agents SDK — Lightweight SDK for building single and multi-agent systems. GitHub
- Semantic Kernel (Microsoft) — SDK for orchestrating AI agents with planners, memory, and connectors. GitHub
- Dify — LLM app development platform with visual workflow builder and agent capabilities. GitHub
- Flowise — Low-code visual LLM flow builder with drag-and-drop interface. Note: Company acquired by Workday; repo actively maintained (last push Jul 2026). GitHub
- Fazm — Apr 2026. Open-source local computer-use agent for macOS. Drives apps via accessibility APIs, model-agnostic, faster than screenshot-based agents.
- Smolagents (Hugging Face) — Minimalist agent library where agents "think in code." Lightweight, zero boilerplate. Supports code agents and tool-calling agents.
- Swarms — Enterprise-grade multi-agent orchestration framework. Scalable infrastructure for autonomous agent swarms. Highly modular. Actively maintained (last push Aug 2026).
- Letta (MemGPT) — Framework for long-term agent memory. Virtual memory management that pages data in/out of context like an OS. Persistent agents.
- Griptape — Enterprise agent framework with strictly typed Pipelines, Workflows, and Agents. Structure-first, production-ready.
- Atomic Agents — Framework inspired by Atomic Design. Compose agents from small, reusable, modular components. Testable and scalable.
- PraisonAI — Low-code multi-agent framework. Define agent roles, tasks, and flows via YAML configuration. Wraps underlying agent frameworks.
- Cognee — GraphRAG framework for agent knowledge management. Builds interconnected knowledge graphs from unstructured data.
- MetaGPT — Multi-agent framework simulating a full software team. Assigns Agent, Product Manager, Engineer roles. Implements SOPs for end-to-end code generation. Note: Active development (last push Jan 2026).
- ChatDev (OpenBMB) — Virtual software company driven by multi-agent collaboration. Follows waterfall model through design, coding, testing, and documentation.
- AutoGPT — The original autonomous agent experiment. Sets its own goals, iterates on tasks, and executes without continuous human input. Web browsing and file management.
- Bee Agent Framework (IBM) — Production-ready framework for building reliable AI agents in Python and TypeScript. Modular, with built-in observability and IBM research optimizations.
- Eliza (elizaOS) — Multi-platform agent framework for creating character-driven AI agents. Handles social media interaction, complex decision-making, and autonomous behavior across platforms.
- Qwen-Agent (Alibaba) — Agent framework tightly integrated with the Qwen model family. Optimized for function calling, code execution, RAG, and tool use with Qwen models.
- AGiXT — Extensible modular AI agent automation platform. Plugin system for swapping LLMs, memory backends, and tools. Highly customizable agent workflows.
- Microsoft Agent Framework — 12.3K★. Production-grade multi-agent framework for Python and .NET. Graph-based workflows, streaming, human-in-the-loop. MIT.
- GenericAgent — 13.5K★. Minimal, self-evolving autonomous agent framework. ~3K lines core, 9 atomic tools. Self-crystallizing skill tree from every task. MIT.
- Omnigent — 7.6K★. Open-source meta-harness orchestrating Claude Code, Codex, Cursor, Pi, and custom agents. Real-time collaboration from any device. Apache 2.0.
🔧 MCP Servers & Tools
📅 Last checked: August 1, 2026
Model Context Protocol (MCP) servers that connect AI assistants to external tools, data sources, and APIs.
- GitHub MCP Server — 31.7K★. Official GitHub MCP server by GitHub. Repository management, issue/PR automation, CI/CD intelligence, code analysis. OAuth or PAT auth. MIT license.
- GitMCP — 8.2K★. Free, open-source, remote MCP server for any GitHub project. Zero-setup documentation and code access for AI assistants. Apache 2.0.
- MCP Reference Servers (Anthropic) — 88.9K★. Official reference implementations: Filesystem, Git, Fetch, Memory, Time, Sequential Thinking. Apache 2.0 / MIT.
- MCP Server Toolkit — Semantic code search, docs server, database server (Postgres/MySQL/SQLite), OpenAPI introspection. One-command
npxsetup. MIT. - MCP Depot — Self-hosted MCP server hub with web UI. Connect Jira, GitHub, Confluence, Jenkins, custom APIs. AGPL-3.0.
- free-search-mcp — Free web search via MCP with no API key required. Multiple search backends with automatic failover. Apache 2.0.
🎛 Fine-tuning Tools
📅 Last checked: August 1, 2026
Tools to fine-tune free models on your own data — all free and open-source.
- Unsloth — Fast memory-efficient fine-tuning. 2x faster, 50% less memory. Supports QLoRA, LoRA, full fine-tune.
- Axolotl — Streamlined fine-tuning framework supporting multiple model architectures and quantization methods.
- LLaMA-Factory — Easy-to-use fine-tuning with web UI. Supports 100+ models, multiple training methods.
- Hugging Face TRL — Transformer Reinforcement Learning library. SFT, PPO, DPOTrainer, GRPOTrainer for aligning models.
- XTuner (InternLM) — Efficient fine-tuning toolkit supporting QLoRA, LoRA, and full fine-tune with multiple model architectures.
- Ludwig (Predibase) — Declarative ML framework. Fine-tune models with a simple config file. GitHub
✨ Prompt Engineering Tools
📅 Last checked: August 1, 2026
Free tools for testing, managing, and optimizing prompts.
- Promptfoo — Company acquired by OpenAI; repo actively maintained (last push Aug 2026). Open-source tool for prompt testing and evaluation. Systematic A/B testing of prompts. GitHub
- Fabric (Daniel Miessler) — Open-source framework for augmenting humans with AI. Library of curated prompts (patterns) for common tasks.
- LangFuse — Open-source LLM engineering platform with prompt management, versioning, and evaluation. GitHub
- DSPy (Stanford) — Framework for algorithmically optimizing LM prompts and weights. GitHub
- Agenta — Open-source LLM platform for prompt management, evaluation, and deployment. GitHub
📊 LLM Evaluation & Observability
📅 Last checked: August 1, 2026
Free, open-source tools for tracing, evaluating, and monitoring LLM applications in development and production.
- Langfuse — 31.5K★. Full LLM engineering platform: tracing, evaluations, prompt management, playground, datasets. Self-hostable. MIT (core). GitHub
- Opik (Comet) — 20.7K★. Open-source LLM observability, evaluation, and agent tracing. Datasets, experiments, LLM-as-judge, guardrails, prompt management. Apache 2.0.
- Phoenix (Arize AI) — AI observability platform with OpenTelemetry-based tracing, evals, experiments, and prompt playground. Elastic License 2.0.
- TruLens — 3.4K★. Agent-specific evaluations (7 purpose-built evaluators). OpenTelemetry tracing, MCP support, batch and inline evaluation. MIT.
- OpenLLMetry (Traceloop) — 7K★. OpenTelemetry-based LLM observability. Send traces to any OTLP-compatible backend. Apache 2.0.
📊 Datasets
📅 Last checked: August 1, 2026
Free, open datasets for training, fine-tuning, and evaluating models.
- Hugging Face Datasets — The standard hub for open datasets. 150,000+ datasets across all tasks.
- Common Corpus — Massive open-source dataset for training large language models. (Requires Hugging Face login)
- The Stack v2 (BigCode) — Large-scale code dataset covering 619 programming languages. Permissive license.
- FineWeb (Hugging Face) — High-quality web dataset for LLM pre-training. 15T tokens.
- Dolly (Databricks) — 15k instruction-response pairs for fine-tuning. CC-BY-SA.
- OpenAssistant Conversations — 160k human-generated assistant conversations. Apache 2.0.
- ShareGPT (RyokoAI) — Real user-ChatGPT conversations for fine-tuning.
- UltraChat (Sean C.) — 200k multi-turn conversations synthesized by ChatGPT.
- No Robots (Hugging Face) — 10k high-quality human-written instructions. Apache 2.0.
- MMLU / GSM8K — Standard benchmarks for evaluation.
- CodeAlchemy (IBM) — Jul 2026. ~1T tokens of synthetic code across 15 languages. Includes 1.3M code+execution-trace pairs. Permissive license.
☁ Model Hosting Platforms
📅 Last checked: August 1, 2026
Free platforms that host models — run inference without downloading anything.
- Hugging Face Spaces — Free hosting for ML apps (Gradio, Streamlit). Thousands of community demos.
- Hugging Face Inference Endpoints (Free Tier) — Deploy models with free trial credits.
- Google Colab (Free Tier) — Free GPU (T4, sometimes A100). Perfect for running models and fine-tuning.
- Kaggle Notebooks — Free GPU (T4 x2). 30 hours/week. Good for heavier workloads.
- Lightning AI Studio — Free tier with GPU access for development and prototyping.
- Modal — Free monthly credits for serverless GPU compute.
- Replicate (Free Tier) — Free credits for running community models.
- Deepnote — Free tier with GPU for data science and ML notebooks.
- Zylora — Deploy GPU functions from any language. Free tier: $5 GPU credits/month, no credit card. T4/L4 pool, sub-300ms cold starts.
📚 Learning Resources
📅 Last checked: August 1, 2026
Free courses, books, and tutorials for learning AI and LLMs.
- Fast.ai — Code-first deep learning education. Practical, free courses from fundamentals to advanced.
- Hugging Face NLP Course — Comprehensive free course on transformers, tokenizers, datasets, and deployment.
- DeepLearning.AI Short Courses — Free short courses on LLMs, RAG, LangChain, and AI agents.
- Full Stack Deep Learning — Free course on ML engineering: training, deploying, and maintaining models.
- Andrej Karpathy's Course — From-scratch neural network implementation videos.
- Neural Networks: Zero to Hero — YouTube series building neural networks from scratch.
- Prompt Engineering Guide (DAIR.AI) — Comprehensive free guide on prompt engineering techniques.
- Anthropic Cookbook — Free recipes and patterns for working with Claude.
- OpenAI Cookbook — Free examples and guides for the OpenAI API.
- LearnLLM.dev — Free AI engineering course with 110+ lessons, runnable code, in-browser playground. Covers fundamentals to production agents.
- LLM Zoomcamp (DataTalksClub) — Free 10-week course on building LLM applications with RAG, agents, vector search, and evaluation.
- AI Engineering from Scratch — 2026. 503 lessons across 20 phases. Build LLMs, agents, and MCP servers from scratch. MIT. 41K+ stars.
🏆 Resources & Leaderboards
📅 Last checked: August 1, 2026
- Perplexity — Free AI search and research assistant with real-time answers and source citations.
- BenchLM.ai — New. LLM leaderboard with 281 models compared across 8 categories. Verified benchmark data updated weekly.
- Hugging Face Open LLM Leaderboard — The primary benchmark for open-weight models. Updated regularly.
- LMSYS Chatbot Arena — Human preference rankings of models. Best source for real-world quality comparisons.
- Artificial Analysis — Independent benchmarks for speed, pricing, and quality across providers.
- Hugging Face Models — Search 1M+ models. Filter by license, task, framework.
- OpenRouter Models — Browse models available via API with pricing and free tiers.
- Ollama Library — Browse models available for one-command local setup.
- cheahjs/free-llm-api-resources — Community-maintained list of free LLM API resources.
👥 Communities
📅 Last checked: August 1, 2026
- Hugging Face Discord — Model releases, discussions, and community support.
- r/LocalLLaMA — The largest Reddit community for running local LLMs.
- Ollama Discord — Ollama community for local model enthusiasts.
- LM Studio Discord — LM Studio community.
- Hugging Face Forums — Discussions on models, datasets, and Spaces.
- r/MachineLearning — General ML/AI research and news.
- Discord: AI Agents — Community for AI agent development and agentic frameworks.
License
To the extent possible under law, the author has waived all copyright and related or neighboring rights to this work.
