Code & Dev
Mercury 2 is a diffusion language model from Inception Labs that fundamentally rethinks how LLMs generate text. Rather than predicting tokens sequentially, it refines entire responses in parallel — delivering over 1,009 tokens per second on NVIDIA Blackwell GPUs while maintaining reasoning-grade output quality. The model features a 128K context window, tunable reasoning depth, native tool use, schema-aligned JSON output, and full OpenAI API compatibility, making it a drop-in replacement anywhere speed is critical. At $0.25 per million input tokens and $0.75 per million output tokens, Mercury 2 targets production AI applications where inference latency compounds across multi-step agent loops. It is accessible via the Inception Labs API platform and a hosted chat interface.