jev-router
Picks the model for each task: Opus 5.5 for hard work, cheaper models for simple turns.
Live catalog — prices update automatically when providers change theirs. Amounts in dinars use today's rate: 1 $ = 250 DA.
Synced automatically with the providers — updated 5 h ago.
jev-router
Picks the model for each task: Opus 5.5 for hard work, cheaper models for simple turns.
claude-opus-5-5
Anthropic's first Claude 5.5 model. Fable 5.1-level performance on most work at 40% lower cost than Opus 5, with a 1M context window; thinking is always on and forced tool_choice is not supported.
claude-fable-5.1
Anthropic's most capable model for demanding reasoning and long-horizon agentic work, with a 1M context window. Successor to Fable 5 with cheaper cache reads; forced tool_choice is not supported.
gpt-6-astra
OpenAI's most capable model, built for the hardest end-to-end work: complex reasoning, coding, computer use, research, and document creation. 1M+ context window, async tool calling, mid-turn steering, and reasoning effort from low through max.
deepseek/deepseek-v4-1-flash
DeepSeek's smallest new-architecture model: natively multimodal with 1M context window, faster and cheaper than V4 Pro while exceeding it on agentic, coding and reasoning tasks. Supports image input, tool calling and thinking mode.
z-ai/glm-5.3
Z-AI's flagship reasoning model on a 1M-token context. Reasoning is always on, with low, high, and max effort levels, and strong coding and tool-calling.
claude-sonnet-5-fast
Anthropic's most agentic Sonnet model, with performance approaching Opus 4.8 at a fraction of the cost.
claude-fable-5-fast
Anthropic's Fable 5 model, with a 1M context window. Shares Opus 4.8's reasoning efforts and modalities.
claude-fable-5
Anthropic's Fable 5 model, with a 1M context window. Shares Opus 4.8's reasoning efforts and modalities.
claude-sonnet-5
Anthropic's most agentic Sonnet model, with performance approaching Opus 4.8 at a fraction of the cost.
claude-opus-4-8
Anthropic's model for complex agentic coding and enterprise work. Near-Fable 5 intelligence at half the price, with a 1M context window.
anthropic/claude-opus-4.8
Anthropic's model for complex agentic coding and enterprise work. Near-Fable 5 intelligence at half the price, with a 1M context window.
claude-opus-5
Anthropic's model for complex agentic coding and enterprise work. Near-Fable 5 intelligence at half the price, with a 1M context window.
gemini-3.6-flash
Google's Flash model with near-Pro intelligence at Flash-tier cost and speed: strong coding, parallel agentic execution, thinking, and search grounding.
gpt-5.6-sol
OpenAI's GPT-5.6 Sol is a frontier model with a 1M+ context window, reasoning, and broad tool support including web search, file search, image generation, code interpreter, hosted shell, computer use, and MCP.
gpt-5.6-terra
OpenAI's cost-optimized variant of GPT-5.6 with a 1M+ context window, reasoning, and broad tool support at a fraction of the flagship cost.
gpt-5.6-luna
OpenAI's most affordable and fastest variant of GPT-5.6 with a 1M+ context window, optimized for high-volume, latency-sensitive applications at minimal cost.
gpt-5.5
OpenAI's GPT-5.5 is a frontier model with a 1M+ context window, reasoning, and broad tool support including web search, file search, image generation, code interpreter, hosted shell, computer use, and MCP. Knowledge cutoff December 1, 2025.
gpt-5.4
OpenAI's GPT-5.4 is OpenAI's latest frontier model with a 1M+ context window, improved reasoning with xhigh effort support, and enhanced capabilities for coding, agentic tasks, and computer use.
gpt-5.4-mini
OpenAI's efficient, cost-optimized variant of GPT-5.4 with strong reasoning and multimodal capabilities at a fraction of the cost.
deepseek/deepseek-v4-flash-0731
DeepSeek's fast, cost-efficient general-purpose model with 1M context window, updated July 2026 with substantially enhanced agentic capabilities. Supports tool calling and thinking mode.
tencent/hy3
Tencent Hunyuan's 295B/21B MoE model for real-world business scenarios: native 256K context, three reasoning modes, strong coding, long-form comprehension, multi-turn dialogue, and agentic task execution.
deepseek/deepseek-v4-flash
DeepSeek's fast, cost-efficient general-purpose model with 1M context window, updated July 2026 with substantially enhanced agentic capabilities. Supports tool calling and thinking mode.
xiaomi/mimo-v2.5
Xiaomi's MiMo-V2.5 native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding within a unified architecture. Built on the MiMo-V2-Flash backbone with dedicated vision and audio encoders.
z-ai/glm-5.2
Z-AI's latest flagship model for long-horizon tasks. Delivers a substantial leap in long-horizon capability over GLM-5.1 on a solid 1M-token context, with stronger coding and flexible thinking effort.
deepseek/deepseek-v4-pro
DeepSeek's flagship reasoning model with 1M context window. Supports tool calling, extended thinking, and high-quality generation.
minimax/minimax-m3
MiniMax's natively multimodal flagship with a 1M-token context window powered by MiniMax Sparse Attention (MSA). Supports image and PDF input, toggleable thinking, strong agentic tool calling, and long-context reasoning.
moonshotai/kimi-k3
Moonshot AI's K3 flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window.
stepfun/step-3.7-flash
StepFun AI's flagship high-efficiency multimodal reasoning model on a sparse Mixture-of-Experts architecture (198B total, ~11B active) with a 256K context window, selectable reasoning effort, native image understanding, and multi-step function calling.
google/gemini-3-flash-preview
Google's latest model with the Flash line's focus on latency, efficiency, and cost, in preview version.
claude-sonnet-4.6
Anthropic's latest Sonnet model with hybrid reasoning, matching near-flagship performance at a fraction of the cost.
google/gemini-2.5-flash-lite
Google's most cost-efficient and lowest-latency multimodal model in the 2.5 family, optimized for high-volume classification, simple data extraction, and low-latency tasks.
openai/gpt-5.6-luna-pro
OpenAI's most affordable and fastest variant of GPT-5.6 with a 1M+ context window, optimized for high-volume, latency-sensitive applications at minimal cost.
google/gemini-2.5-flash
Google's fast and cost-effective model with a 1M token context window. Best for high-volume, low-latency tasks and agentic use cases.
xiaomi/mimo-v2.5-pro
Xiaomi's MiMo-V2.5-Pro is an open-source Mixture-of-Experts (MoE) language model with 1.02T total parameters and 42B active parameters. It utilizes the hybrid attention architecture and 3-layer Multi-Token Prediction (MTP) introduced in MiMo-V2-Flash.
xiaomi/mimo-v2.6-pro
google/gemini-3.1-flash-lite
Google's lightweight and efficient model in the 3.1 family, optimized for low latency and cost, in preview version.
google/gemma-4-31b-it
Dense 30.7B multimodal model from Google DeepMind. Multimodal (text + image input), text output.
openai/gpt-oss-120b
deepseek/deepseek-v3.2
DeepSeek's latest general-purpose model with improved reasoning, coding, and instruction-following capabilities.
google/gemma-4-26b-a4b-it
Efficient MoE variant of Gemma 4 from Google DeepMind. Multimodal (text + image input), text output.
x-ai/grok-4.5
xAI's Grok 4.5 with 500K token context window. Supports configurable reasoning, function calling, structured outputs, and vision.
claude-opus-4.7
Anthropic's most capable generally available model. Step-change improvement in agentic coding over Opus 4.6, with a new tokenizer and 1M context window.
qwen/qwen3.8-max
Alibaba's 2.4-trillion-parameter MoE flagship with a 1M-token context window. Deep-thinking model tuned for coding, professional work, and long-horizon agentic execution, with native visual understanding of documents and video.
claude-haiku-4.5
Anthropic's fastest model with near-frontier intelligence, ideal for high-throughput applications.
openai/gpt-4o-mini
OpenAI's cost-efficient small model. Great for lightweight tasks with fast responses and low cost while maintaining strong capabilities.
minimax/minimax-m2.7
MiniMax's latest flagship model with 200K context window and 131K max output. Advanced agentic capabilities with multi-agent collaboration, strong coding, and tool calling.
google/gemini-3.5-flash-lite
Google's most cost-efficient Gemini model, built for high-throughput and agentic use cases with thinking, tool calling, and search grounding.
google/gemini-3.5-flash
Google's latest Flash-class multimodal model, balancing latency and capability with thinking, function calling, and search grounding.
openai/gpt-5-mini
OpenAI's smaller and faster variant of GPT-5, a more cost-efficient alternative. Great for well-defined tasks and precise prompts.
moonshotai/kimi-k2.6
Moonshot AI's next-gen agentic model built on K2. Long-horizon coding, proactive autonomous execution, and swarm-based task orchestration.
google/gemini-3.1-pro-preview
Google's latest model, in preview version.
google/gemini-3.7-flash
Google's Flash model for agentic and multimodal work: thinking, search grounding, code execution, and computer use over a 1M-token context.
claude-opus-4.6
Anthropic's most recent model. Current leader on agentic coding evaluation, Terminal Bench 2.0, and Humanity's Last Exam.
openai/gpt-oss-20b
qwen/qwen3.7-flash
Alibaba's Qwen3.7 Flash native vision-language model with a 1M-token context window. Hybrid thinking model supporting text and image inputs, tool calling, structured output, and web search, with upgraded multimodal agent and coding capabilities.
qwen/qwen3.6-35b-a3b
Alibaba's Qwen3.6 model with 35B total / 3B active parameters using mixture-of-experts. Strong general-purpose performance.
mistralai/mistral-nemo
moonshotai/kimi-k2.5
Moonshot AI's native multimodal agentic model built on K2. Excels at visual coding, reasoning, and self-directed agent swarm with up to 100 sub-agents.
moonshotai/kimi-k2.7-code
Moonshot AI's K2.7 model specialized for coding. FP4-quantized for fast inference with a 262K context window.
google/gemini-3.1-flash-lite-preview
Google's lightweight and efficient model in the 3.1 family, optimized for low latency and cost, in preview version.
openai/gpt-5.4-nano
OpenAI's most affordable and fastest variant of GPT-5.4, optimized for high-volume, latency-sensitive applications with minimal cost.
qwen/qwen3.7-plus
Alibaba's Qwen3.7 Plus model with a 1M-token context window. Hybrid thinking model supporting text and image inputs, tool calling, structured output, and web search.
openai/gpt-4.1-mini
OpenAI's balanced small model with a massive 1M token context window. Offers great performance at low cost, beating GPT-4o in many benchmarks.
claude-sonnet-4.5
Anthropic's latest Sonnet model featuring exceptional performance on coding, analysis, and instruction following.
z-ai/glm-5.1
Z-AI's next-generation flagship model for agentic engineering. Stronger coding capabilities and state-of-the-art performance on SWE-Bench Pro.
z-ai/glm-5
ZhipuAI's most capable reasoning model with 128K context window. Supports tool calling, web search, and thinking.
qwen/qwen3.7-max
Alibaba's flagship Qwen3.7 agent model with a 1M-token context window. Hybrid thinking model tuned for coding, agentic workflows, and long-horizon autonomous execution. Supports tool calling and reasoning.
x-ai/grok-4.3
xAI's Grok 4.3 with 1M token context window. Supports configurable reasoning, function calling, structured outputs, and vision.
qwen/qwen3-235b-a22b-2507
z-ai/glm-4.7
ZhipuAI's advanced reasoning model with 128K context window. Supports tool calling, web search, and thinking.
amazon/nova-micro-v1
Amazon's fastest text-only model optimized for speed and cost. Ideal for text-based tasks requiring low latency with 128K context window.
openai/gpt-5.6-terra-pro
OpenAI's cost-optimized variant of GPT-5.6 with a 1M+ context window, reasoning, and broad tool support at a fraction of the flagship cost.
nvidia/nemotron-3-ultra-550b-a55b
poolside/laguna-s-2.1
openai/gpt-5.2
OpenAI's latest and greatest, GPT 5.2 is OpenAI's flagship model for coding and agentic tasks across industries.
meta-llama/llama-3.1-8b-instruct
Meta's efficient 8B parameter model optimized for multilingual dialogue. Fast inference with great performance for everyday tasks.
deepseek/deepseek-chat-v3.1
openai/gpt-5-nano
OpenAI's most cost-effective and efficient model in the GPT-5 series, optimized for speed and affordability. Ideal for straightforward tasks and high-volume applications where latency and cost are critical.
inclusionai/ling-3.0-flash
deepseek/deepseek-chat-v3-0324
x-ai/grok-4.20
tencent/hy3-preview
openai/gpt-5.3-codex
OpenAI's GPT-5.3-Codex is the most capable agentic coding model, combining frontier coding performance with reasoning capabilities. Features mid-task steering and 25% faster inference than GPT-5.2-Codex.
minimax/minimax-m2.5
MiniMax's latest flagship model with 200K context window, ~60 tokens/second. Supports function calling and thinking.
claude-opus-5-fast
nvidia/nemotron-3-nano-30b-a3b
qwen/qwen3.6-plus
Alibaba's Qwen3.6 Plus model with a 1M-token context window. Hybrid thinking model supporting text and image inputs, tool calling, structured output, and web search.
qwen/qwen3-30b-a3b-instruct-2507
meta-llama/llama-3.3-70b-instruct
Meta's powerful 70B parameter model quantized to FP8 for fast inference. Excellent for complex reasoning and multilingual tasks.
google/gemma-3-27b-it
Google's largest Gemma 3 model with 27B parameters. Strong performance across reasoning, coding, and multilingual tasks with vision support.
qwen/qwen3.5-9b
Alibaba's Qwen3.5 9B open-source model with a 256K-token context window. Hybrid thinking model supporting text and image inputs, tool calling, and structured output.
openai/gpt-4.1-nano
qwen/qwen3.6-27b
Alibaba's Qwen3.6 27B dense model with reasoning and vision support. Strong general-purpose performance at a small footprint.
z-ai/glm-5v-turbo
deepseek/deepseek-chat
qwen/qwen3-coder-next
Alibaba's latest Qwen3 coding model. Cutting-edge code generation and understanding capabilities.
claude-sonnet-4
deepseek/deepseek-v3.2-exp
meta-llama/llama-4-maverick
Meta's Llama 4 Maverick with 17B active parameters using mixture-of-experts. Excels at coding, reasoning, and multilingual tasks.
qwen/qwen3-32b
Alibaba's Qwen3 dense 32B model. Strong performance across reasoning, math, and coding tasks.
meta-llama/llama-4-scout
Meta's latest Llama 4 model with 17B active parameters using mixture-of-experts architecture.
qwen/qwen3.5-35b-a3b
Alibaba's Qwen3.5 35B-A3B open-source model with a 256K-token context window. Hybrid thinking model supporting text and image inputs, tool calling, structured output, and web search.
nvidia/nemotron-3-super-120b-a12b
deepseek/deepseek-v3.1-terminus
google/gemma-3-12b-it
Google's lightweight open model well-suited for text generation and image understanding. Supports LoRA fine-tuning with 128K context.
mistralai/mistral-small-2603
qwen/qwen3.5-122b-a10b
Alibaba's Qwen3.5 native vision-language model with 122B total / 10B active parameters using a hybrid linear-attention sparse mixture-of-experts architecture.
mistralai/mistral-small-3.2-24b-instruct
Mistral's efficient 24B model optimized for simple tasks with low latency and 128K context. Great for classification, customer support, text generation, and multimodal tasks.
thinkingmachines/inkling
qwen/qwen3-vl-235b-a22b-instruct
Alibaba's Qwen3 vision-language model with 235B total / 22B active parameters. Supports text and image inputs.
openai/gpt-5.1
OpenAI's previous flagship model for coding and agentic tasks with configurable reasoning and non-reasoning effort.
claude-opus-4.5
Anthropic's former flagship model combining maximum intelligence with practical performance.
openai/gpt-5
OpenAI's former advanced reasoning model with enhanced problem-solving capabilities, deeper understanding, and improved accuracy across complex tasks.
z-ai/glm-4.7-flash
ZhipuAI's fast and efficient model variant. Optimized for speed while maintaining strong language capabilities.
qwen/qwen3.5-397b-a17b
Alibaba's Qwen3.5 397B-A17B open-source model with a 256K-token context window. Hybrid thinking model supporting text and image inputs, tool calling, structured output, and web search.
meta-llama/llama-3.1-70b-instruct
Meta's 70B parameter model with 128K context. Strong performance on reasoning, coding, and multilingual tasks.
qwen/qwen3-coder
z-ai/glm-4.6
ZhipuAI's reasoning model with 128K context window. Supports tool calling, web search, and thinking.
openai/gpt-4o
OpenAI's multimodal model optimized for speed and efficiency. Capable of processing text, images, and audio with high accuracy at reduced latency and cost compared to GPT-4 Turbo.
claude-opus-4.8-fast
qwen/qwen3-next-80b-a3b-instruct
Alibaba's latest Qwen3 model with 80B total / 3B active parameters using mixture-of-experts. Strong general-purpose performance.
qwen/qwen3.6-flash
Alibaba's Qwen3.6 Flash model with a 1M-token context window. Hybrid thinking model supporting text and image inputs, tool calling, structured output, and web search.
qwen/qwen3-vl-32b-instruct
qwen/qwen3-vl-8b-instruct
z-ai/glm-4.5-air
openai/gpt-4o-mini-2024-07-18
OpenAI's cost-efficient small model. Great for lightweight tasks with fast responses and low cost while maintaining strong capabilities.
upstage/solar-pro4
deepseek/deepseek-r1-0528
DeepSeek R1 May 2025 update. Reasoning model with strong performance on math, code, and complex reasoning tasks.
mistralai/mistral-medium-3-5
thinkingmachines/inkling-small
moonshotai/kimi-k2-0905
qwen/qwen3.5-27b
Alibaba's Qwen3.5 27B open-source model with a 256K-token context window. Thinking model supporting text and image inputs, tool calling, and structured output.
qwen/qwen3-vl-30b-a3b-instruct
Alibaba's Qwen3 vision-language model with 30B total / 3B active parameters. Supports text and image inputs.
aion-labs/aion-3.0
mistralai/mistral-large-2512
poolside/laguna-xs-2.1
qwen/qwen-2.5-7b-instruct
openai/gpt-oss-safeguard-20b
meta-llama/llama-guard-4-12b
qwen/qwen3.5-plus-02-15
qwen/qwen3-coder-30b-a3b-instruct
Alibaba's Qwen3 coding-focused model with 30B total / 3B active parameters using mixture-of-experts architecture.
meta/muse-glimmer-30b
mistralai/mistral-small-24b-instruct-2501
z-ai/glm-5-turbo
meituan/longcat-2.0
qwen/qwen3-14b
Alibaba's Qwen3 14B model with strong reasoning and instruction-following. A capable mid-sized model for coding and complex tasks.
mistralai/ministral-8b-2512
perceptron/perceptron-mk1
qwen/qwen3-8b
z-ai/glm-4.5
ZhipuAI's efficient reasoning model with 128K context window. Supports tool calling, web search, and thinking.
openai/gpt-5.6-sol-pro
OpenAI's GPT-5.6 Sol is a frontier model with a 1M+ context window, reasoning, and broad tool support including web search, file search, image generation, code interpreter, hosted shell, computer use, and MCP.
qwen/qwen3-235b-a22b-thinking-2507
z-ai/glm-5.3-flash
Z-AI's first natively multimodal GLM-5 model: image and PDF input on a 1M-token context at flash pricing. Reasoning is always on (low/high/max), with function calling and implicit prompt caching.
x-ai/grok-4.6
xAI's frontier model for coding, agentic tasks, and knowledge work, with a 500K token context window. Supports configurable reasoning (including xhigh), function calling, structured outputs, web search, image and PDF input.
claude-opus-4-1
Anthropic's previous flagship model with maximum intelligence. Legacy model - consider using Opus 4.5 for new projects.
mistralai/codestral
Mistral's specialized model for code generation. Optimized for coding tasks including code completion, generation, and explanation.
cohere/command-a
Command A is Cohere's most performant model to date, excelling at tool use, agents, retrieval augmented generation (RAG), and multilingual use cases.
deepseek/deepseek-r1
DeepSeek's reasoning model trained with reinforcement learning. Excels at math, code, and complex reasoning tasks.
deepseek/deepseek-r1-distill-32b
DeepSeek's reasoning model distilled from R1 based on Qwen2.5. Outperforms OpenAI o1-mini across various benchmarks with state-of-the-art dense model results.
deepseek/deepseek-v3-0324
DeepSeek V3 (March 2024 release) — general-purpose text model with strong reasoning, coding, and instruction-following capabilities.
deepseek/deepseek-v3-1
DeepSeek V3.1 hybrid model with optional thinking mode, function calling, and a 128K context window.
deepseek/deepseek-v4-flash-0423
DeepSeek's fast, cost-efficient general-purpose model with 1M context window. Supports tool calling and thinking mode.
deepseek/deepseek-v4-flash-vision-exp
DeepSeek's experimental multimodal model. Matches DeepSeek V4 Flash on text — agentic use, reasoning and world knowledge — and adds image understanding, at V4 Flash pricing. Supports tool calling and thinking mode.
deepseek/deepseek-v4-pro-0813
DeepSeek's flagship reasoning model with 1M context window, updated to the DeepSeek-V4-Pro-0813 build served by DeepSeek's direct API. Supports tool calling, extended thinking, and high-quality generation.
gemini-2.5-pro
Google's most capable model for complex reasoning tasks. Features a 1M token context window and strong performance across benchmarks.
gemini-3.8-flash
Google's Flash model for agentic and multimodal work: thinking, search grounding, code execution, and computer use over a 1M-token context.
gemma-3-4b
Google's lightweight 4B parameter Gemma 3 model. Efficient for simple tasks with vision capabilities.
gemma-4-e4b
Compact per-layer-embedding variant of Gemma 4 from Google DeepMind. Text-only input, text output, with toggleable reasoning.
z-ai/glm-4.5v
ZhipuAI's efficient vision-language model with 128K context window. Supports image inputs and thinking.
z-ai/glm-4.6v
ZhipuAI's vision-language model with 128K context window. Supports image inputs and thinking.
gpt-4.1
OpenAI's improved version of GPT-4, featuring enhanced reasoning capabilities, better contextual understanding, and increased accuracy across a wide range of tasks.
gpt-5.4-pro
OpenAI's most powerful GPT-5.4 variant with a 1M+ context window, designed for the most demanding reasoning, coding, and agentic tasks with maximum capability.
gpt-oss-safeguard-120b
OpenAI's safety-focused open-weight model with 120B parameters. Designed for content moderation and safe AI applications.
grok-4.20-0309-reasoning
xAI's Grok 4.20 reasoning model with 2M token context window. Supports reasoning, function calling, structured outputs, and vision.
grok-4.20-multi-agent-0309
xAI's multi-agent model with 2M token context window. Optimized for multi-agent workflows with reasoning capabilities. No client-side or custom tools: Client-side tools (function calling) and custom tools are not currently supported by the multi-agent model variant.
grok-4.20-non-reasoning
xAI's Grok 4.20 non-reasoning model with 2M token context window. Optimized for speed without reasoning overhead, supports function calling, structured outputs, and vision.
grok-build-0.1
xAI's fast coding model trained specifically for agentic coding. Supports reasoning, function calling, structured outputs, and image input. 256K context window.
ibm-granite/ibm-granite-micro
IBM's ultra-efficient micro model. Small but mighty - perfect for simple tasks requiring minimal latency and cost.
ai21/jamba-1-5-large
AI21's Jamba 1.5 Large model with hybrid SSM-Transformer architecture. Excels at long-context understanding and generation.
ai21/jamba-1-5-mini
AI21's efficient Jamba 1.5 Mini model with hybrid SSM-Transformer architecture. Fast and cost-effective for everyday tasks.
moonshotai/kimi-k2-thinking
Moonshot AI's trillion-parameter MoE reasoning model (32B activated). Excels at multi-step reasoning with 200+ sequential tool calls. Supports function calling and extended thinking.
meta-llama/llama-3-70b-instruct
Meta's original Llama 3 70B model optimized for dialogue. Strong general-purpose performance across a wide range of tasks.
meta-llama/llama-3-8b-instruct
Meta's efficient Llama 3 8B model optimized for dialogue. Fast inference suitable for lightweight tasks.
meta-llama/llama-3.2-1b-instruct
Meta's compact and efficient 1 billion parameter model designed for on-device and edge deployment. Optimized for instruction following with strong performance despite its small size, ideal for resource-constrained environments.
meta-llama/llama-3.2-3b-instruct
Meta's compact 3B parameter model optimized for edge deployment and multilingual dialogue. Great balance of speed and capability.
mistralai/magistral-small-1.2
Mistral's reasoning-focused small model with vision capabilities. Optimized for step-by-step reasoning tasks.
minimax/minimax-m2
MiniMax's flagship agentic language model with 200K context window. Supports function calling and reasoning.
minimax/minimax-m2-1
MiniMax's lightweight MoE model optimized for coding, agentic workflows, and modern application development. 10B activated parameters with strong multilingual code generation.
minimax/minimax-m2-1-highspeed
MiniMax M2.1 optimized for speed at ~100 tokens/second with 200K context window.
minimax/minimax-m2-5-highspeed
MiniMax M2.5 optimized for speed at ~100 tokens/second with 200K context window.
minimax/minimax-m2-7-highspeed
MiniMax M2.7 optimized for speed with 200K context window and 131K max output.
mistralai/ministral-3-14b
Mistral's mid-range 14B parameter model with vision support. Enhanced reasoning over smaller Ministrals.
mistralai/ministral-3-3b
Mistral's ultra-efficient 3B parameter model with vision support. Designed for edge and low-latency applications.
mistralai/ministral-3-8b
Mistral's efficient 8B parameter model with vision support. Good balance of capability and speed for moderate tasks.
mistralai/mistral-large-3
Mistral's flagship 675B parameter model with state-of-the-art reasoning, coding, and multilingual capabilities with vision support.
mistralai/mistral-medium-3
Mistral's frontier-class multimodal model released May 2025.
mistralai/mistral-medium-3.1
Multimodal model from Mistral, released August 2025. Improved tone and performance.
mistralai/mistral-small-3.1
Mistral's efficient 24B model optimized for simple tasks with low latency and 128K context. Great for classification, customer support, text generation, and multimodal tasks.
meta-llama/muse-spark-1.1
Meta's Muse Spark reasoning model with a 1M-token context window. Supports image input, tool calling, web search, and structured output.
meta-llama/muse-spark-1.2
Meta's updated Muse Spark reasoning checkpoint with a 1M-token context window. Supports image input, tool calling, web search, and structured output.
meta-llama/muse-spark-1.2-contributor
Meta's Muse Spark 1.2 checkpoint on the contributor tier: heavily discounted pricing in exchange for permission to train future Meta models on your prompts and completions.
meta-llama/muse-spark-1.3
Meta's latest Muse Spark checkpoint, tuned for agentic multi-step tool workflows and improved coding over 1.2, with a 1M-token context window. Supports image and PDF input, tool calling, web search, and structured output.
meta-llama/muse-spark-1.3-contributor
Meta's Muse Spark 1.3 checkpoint on the contributor tier: heavily discounted pricing in exchange for permission to train future Meta models on your prompts and completions.
nvidia/nemotron-3-120b
NVIDIA's hybrid MoE model with leading accuracy for multi-agent applications and specialized agentic AI systems.
nvidia/nemotron-3-nano-30b
NVIDIA's efficient hybrid MoE model combining Mamba-2 and Transformer layers, with 30B total and 3.5B active parameters. Reasoning can be toggled on or off, and it handles English, code and several European and Asian languages.
nvidia/nemotron-3-nano-omni
NVIDIA's compact multimodal Nemotron 3 Nano with 30B total / 3B active MoE parameters and built-in reasoning. Strong at agentic and analytical tasks.
nvidia/nemotron-3-ultra-nvfp4
NVIDIA's frontier-scale hybrid Latent-MoE model (550B total, 55B active) for the most demanding agentic, reasoning, and long-context workloads across code, math, and science. Trained with an NVFP4 recipe.
nvidia/nemotron-lightning-3.5-30b
NVIDIA's hybrid Mamba-Transformer MoE model (31B total, 3B active) with a multi-token prediction speculative decoding head for low-latency serving. Tuned for high-throughput reasoning and agentic workloads.
amazon/nova-2-lite
Amazon's second-generation lightweight multimodal model. Processes text and images with improved performance over Nova Lite.
amazon/nova-lite
Amazon's multimodal model for processing images, video, and text. Can analyze multiple images with 300K context.
amazon/nova-pro
Amazon's highly capable multimodal model balancing accuracy, speed, and cost. Processes text and images with 300K context.
o1
OpenAI's reasoning model designed to think before responding. Uses chain-of-thought reasoning for complex tasks in science, coding, and math.
o3
OpenAI's reasoning model that excels at math, science, coding, and visual reasoning tasks. Uses chain-of-thought reasoning before responding.
o3-mini
OpenAI's small reasoning model, optimized for cost-efficient chain-of-thought reasoning in coding, math, and science tasks.
o4-mini
OpenAI's latest small o-series reasoning model, optimized for fast, effective reasoning across coding, math, and visual tasks.
writer/palmyra-x4
Writer's enterprise-grade model optimized for business content generation, analysis, and transformation.
writer/palmyra-x5
Writer's most capable enterprise model with enhanced reasoning, analysis, and content generation capabilities.
qwen/qwen-flash
Alibaba's fastest, most cost-efficient model with a 1M-token context window. Hybrid thinking model supporting text and PDF inputs and structured output.
qwen/qwen-plus
Alibaba's balanced mid-tier model with a 1M-token context window. Hybrid thinking model supporting text and PDF inputs and structured output.
qwen/qwen3-235b-a22b-instruct
Alibaba's Qwen3 235B mixture-of-experts model (22B active, 128 experts) with a native 262K context. Operates in non-thinking mode only and emits no reasoning blocks.
qwen/qwen3-30b
Alibaba's Qwen3 model with groundbreaking advancements in reasoning and instruction-following. Excellent for complex tasks and coding.
qwen/qwen3-coder-480b-a35b
Alibaba's open-weights Qwen3 coding model with 480B total / 35B active parameters using mixture-of-experts architecture.
qwen/qwen3-coder-flash
Alibaba's Qwen3 Coder Flash model with a 1M-token context window. Inherits the coding agent capabilities of Qwen3 Coder Plus, tuned for repository-level understanding and multi-turn tool interaction, supporting text and PDF inputs.
qwen/qwen3-coder-plus
Alibaba's Qwen3 Coder Plus model with a 1M-token context window. Hybrid thinking model tuned for coding, supporting text and PDF inputs.
qwen/qwen3-max
Alibaba's flagship Qwen3 Max model with a 256K-token context window. Non-thinking instruct model tuned for agent programming and tool invocation, supporting structured output and web search.
qwen/qwen3-max-thinking
Alibaba's flagship Qwen3 Max with a 256K-token context window, tuned for extended reasoning on hard math, coding and agentic tasks. Text-only, with tool invocation and JSON object output.
qwen/qwen3-vl-flash
Alibaba's small-scale Qwen3 vision-language model with a 256K-token context window. Integrates thinking and non-thinking modes for fast image and document understanding, and supports structured output.
qwen/qwen3-vl-plus
Alibaba's balanced Qwen3 vision-language model with a 256K-token context window. Supports text, image, and PDF inputs and structured output.
qwen/qwen3.5-flash
Alibaba's Qwen3.5 Flash model with a 1M-token context window. Hybrid thinking model supporting text and image inputs, tool calling, structured output, and web search.
qwen/qwen3.5-plus
Alibaba's Qwen3.5 Plus model with a 1M-token context window. Hybrid thinking model supporting text and image inputs, tool calling, structured output, and web search.
qwen/qwen3.8-27b
Alibaba's Qwen3.8 27B dense model with reasoning, vision, and function calling on a 256K context.
qwen/qwq-32b
Alibaba's QwQ reasoning model optimized for analytical tasks. Strong performance on math, coding, and logical reasoning benchmarks.
gpt-6-luna
OpenAI's most affordable and fastest GPT-6 model with a 1M+ context window, optimized for high-volume, latency-sensitive applications at minimal cost.
gpt-6-sol
OpenAI's GPT-6 Sol is the balanced frontier model of the GPT-6 family: strong reasoning, coding, and agentic tool use at mid-tier cost, with a 1M+ context window, reasoning effort from none through max, web search, and native image and PDF input.
grok-4.7
xAI's flagship model for coding, agentic tool calling, and knowledge work, with a 500K token context window and minimal hallucinations. Supports configurable reasoning (low through xhigh), function calling, structured outputs, web search, and image input.
qwen/qwen3.8-flash
Alibaba's Qwen3.8 Flash native multimodal model with a 1M-token context window. Combines powerful reasoning and generation with remarkable speed, shining in coding assistance, agentic workflows, and visual understanding.
gpt-6.1-sol
Create your account, top up from 500 DA and send your first request in minutes.
Create a free account