Mercury
by Inception
Ultra-fast diffusion language models by Inception for coding, reasoning, agents, voice, and real-time AI applications.
Mercury is a family of diffusion-based large language models developed by Inception Labs for high-speed AI applications. Mercury models are designed to deliver frontier-level language intelligence with very low latency, making them suitable for coding, search, agents, voice applications, and real-time workflows. The current flagship Mercury 2.5 supports reasoning, tool use, structured JSON output, and a 260K-token context window. Mercury is available through the Inception API with OpenAI-compatible APIs, as well as platforms including Baseten and OpenRouter.
Provider
Inception
Category
AI Coding
Available On
API, Web Playground, Baseten, OpenRouter, AWS Bedrock, Azure Foundry
Launched
Jun 2025
Use Cases
Coding, Software Development, AI Agents, Search
Best For
Fast coding, agents, search, and real-time AI workflows where low latency and low token cost matter
🧠 Models (4)
Mercury 2.5v2.5
Diffusion Reasoning LLM
Context Window
260K tokens
Input
text
Output
text, structured JSON
Mercury VoicevVoice
Diffusion Voice LLM
Context Window
128K tokens
Input
text, audio
Output
text, audio
Mercury RoutervRouter
AI Model Router
Context Window
N/A
Input
text
Output
model routing decision
Mercury 2v2
Diffusion Language Model
Context Window
Long context
Input
text
Output
text
✨ Features (9)
Diffusion Language Model
Uses a diffusion-based language model architecture designed for high-throughput generation
Ultra-Fast Generation
Mercury models are designed for very high token generation speeds and low-latency AI interactions
Reasoning
Mercury 2.5 provides tunable reasoning for complex tasks
Tool Use
Supports tool calling and parallel tool calls for agentic workflows
Structured JSON
Produces schema-aligned structured JSON for reliable application integration
Long Context
Mercury 2.5 supports a 260K-token context window
OpenAI API Compatible
Mercury models can be integrated using an OpenAI-compatible API interface
Real-Time Voice
Mercury Voice is optimized for low-latency voice-agent applications
Model Routing
Mercury Router can route requests to models based on quality, speed, and cost
💰 Pricing
✓ Free Plan
✕ Free Trial
✓ Paid Plan
✓ Enterprise
Starting Price
$0.04/1M input tokens at launch pricing
Billed monthly
API Pricing
Mercury 2.5: $0.20/1M input tokens and $0.75/1M output tokens; launch price is $0.04/$0.15.
More Info
View pricing page ↗💬 Prompt Library (5)
Generate Code
CodingGenerate fast, production-oriented code
Build a production-ready solution for the following programming task: [TASK]. Explain the approach briefly, then provide clean, efficient, maintainable code with error handling and security considerations.
Code Review
CodingAnalyze and improve existing code
Review the following code for bugs, security issues, performance problems, maintainability, and edge cases. Give specific recommendations and provide an improved version. ```[PASTE CODE HERE]```
Build AI Agent
CodingDesign an agentic AI workflow
Design an AI agent for [TASK]. Define the system workflow, tools it needs, decision logic, error handling, and structured JSON output format. Optimize the design for speed and reliability.
Structured JSON
CodingGenerate reliable structured JSON output
Convert the following information into valid JSON using this exact schema. Do not add extra fields and ensure every value follows the requested data type. Schema: [SCHEMA] Data: [DATA]
Fast Research
ResearchCreate concise structured research output
Analyze the following topic and produce a concise structured response. Separate confirmed facts, assumptions, and recommended next steps. Topic: [TOPIC]
🔗 Integrations (8)
Inception API
api
Official API platform for accessing Mercury models
OpenAI SDK
developer
Mercury provides OpenAI-compatible APIs for developer integrations
Baseten
inference
Mercury models are available through the Baseten inference platform
OpenRouter
api
Access Mercury models through OpenRouter
AWS Bedrock
cloud
Mercury is available through supported AWS enterprise infrastructure
Azure Foundry
cloud
Mercury is available through supported Microsoft Azure AI infrastructure
LangChain
framework
Mercury can be used through supported OpenAI-compatible integrations
LiteLLM
framework
Mercury models are supported through LiteLLM integrations
🎯 Related AI Coding Tools
View all →Tabnine
Enterprise AI coding platform for code completion, chat, code review, testing, agents, and secure software development.
Replit Ghostwriter
AI coding assistant from Replit, originally branded Ghostwriter, for code generation, completion, explanation, transformation, and AI-powered app development.
Windsurf
AI-powered agentic coding IDE by Windsurf with code completion, Cascade agents, codebase context, debugging, and multi-model support.
Cursor
AI-powered code editor and coding agent for building, editing, debugging, reviewing, and shipping software.
GitHub Copilot
AI coding assistant by GitHub for code completion, coding chat, code review, agents, and software development workflows.
Anima
AI design-to-code platform that turns Figma designs, websites, and prompts into production-ready React, HTML, Vue, and CSS code.
⚠️ Limitations
- •Mercury models are primarily developer and API-focused rather than a general consumer chatbot
- •Some models and features are in early access or preview
- •API usage is billed based on token consumption
- •Mercury Voice and Mercury Router availability may require sales or early-access approval
- •Model capabilities and availability can change as Inception releases new versions
- •AI-generated responses may contain inaccurate information
📚 Sources
- https://www.inceptionlabs.ai/models(official)
- https://www.inceptionlabs.ai/blog/introducing-mercury-2-5(official)
- https://www.inceptionlabs.ai/(official)
- https://arxiv.org/abs/2506.17298(research)