Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Model catalog

AI Models

Explore 344+ AI models available in MindStudio. From large language models to image generators — no API keys required.

344 models

Anthropic13

Text

Claude 5 Opus

Claude 5 Opus is a text generation model from Anthropic with a 1,000,000-token context window and native tool use.

Text

Claude 5 Sonnet

Claude 5 Sonnet is Anthropic's flagship text and vision model with a 1,000,000-token context window and extended reasoning.

Text

Claude 5 Fable

Claude 5 Fable is a text generation model from Anthropic with a 1,000,000-token context window and built-in reasoning capabilities.

Text

Claude 4.8 Opus

Claude 4.8 Opus is a text generation model from Anthropic with a 1,000,000-token context window and native tool use.

Text

Claude 4.7 Opus

Claude 4.7 Opus is a text generation model from Anthropic with a 1,000,000-token context window and native tool use.

Text

Claude 4.6 Sonnet

Claude 4.6 Sonnet is a text generation model from Anthropic with a 1,000,000-token context window and native tool use.

Text

Claude 4.6 Opus

Claude 4.6 Opus is a text generation model from Anthropic with a 1,000,000-token context window and native tool and MCP support.

Text

Claude 4.5 Opus

Claude 4.5 Opus is a text generation model from Anthropic with a 200,000-token context window and extended reasoning support.

Text

Claude 4.5 Haiku

Claude 4.5 Haiku is a text generation model from Anthropic with a 200,000-token context window and image input support.

Text

Claude 4.5 Sonnet

Claude 4.5 Sonnet is a text generation model from Anthropic with a 200,000-token context window and native tool and MCP support.

Text

Claude 3 SonnetDeprecated

Claude 3 Sonnet is a text generation model from Anthropic with a 200,000 token context window, available via Amazon Bedrock.

Text

Claude 3 HaikuDeprecated

Claude 3 Haiku is a text generation model from Anthropic with a 200,000 token context window, available via Amazon Bedrock.

Text

Claude InstantDeprecated

Claude Instant is a text generation model from Anthropic optimized for speed and efficiency with a 100,000 token context window.

OpenAI47

Text

GPT 5.6 Luna

GPT 5.6 Luna is a text and image-input model from OpenAI with a 1,050,000-token context window.

Text

GPT 5.6 Terra

GPT 5.6 Terra is a text generation model from OpenAI with a 1,050,000-token context window and image input support.

Text

GPT 5.6 Sol

GPT 5.6 Sol is a text generation model from OpenAI with a 1,050,000-token context window and image input support.

Text

GPT-5.5

GPT-5.5 is a multimodal text generation model from OpenAI with a 1,000,000-token context window and adjustable reasoning effort.

Text

GPT 5.5 Pro

GPT 5.5 Pro is a text generation model from OpenAI with a 1,000,000-token context window and reasoning capabilities.

Text

GPT 5.4 Pro

GPT 5.4 Pro is a text generation model from OpenAI with a 1,000,000-token context window and 128,000-token max response size.

Text

GPT 5.4

GPT 5.4 is a text generation model from OpenAI with a 1,000,000 token context window and 128,000 token max response size.

Text

GPT‑5.2 Pro

GPT‑5.2 Pro is a text generation model from OpenAI with a 400,000-token context window and adjustable reasoning effort.

Text

GPT-5.1

GPT-5.1 is a text generation model from OpenAI with a 400,000 token context window and adjustable reasoning effort.

Text

o3-pro

o3-pro is a text generation model from OpenAI that applies more compute to reasoning, with a 200,000-token context window.

Text

GPT-5 nano

GPT-5 nano is a text generation model from OpenAI with a 400,000 token context window and tool-calling support.

Text

GPT-5 mini

GPT-5 mini is a text generation model from OpenAI with a 400,000-token context window and tool-use support.

Text

GPT-5

GPT-5 is a text generation model from OpenAI with a 400,000-token context window and adjustable reasoning effort.

Text

GPT OSS 120B

GPT OSS 120B is an open-source text generation model from OpenAI, served via Groq with a 128,000-token context window.

Text

GPT OSS 20B

GPT OSS 20B is an open-source text generation model from OpenAI with a 128,000-token context window, served via Groq.

Text

o3

o3 is a text generation model from OpenAI released in April 2025, featuring a 200,000 token context window.

Text

o1-pro

o1-pro is a text generation model from OpenAI that applies extended compute to reasoning tasks, with a 200,000-token context window.

Text

o3-mini

o3-mini is a text generation model from OpenAI designed for reasoning-intensive tasks with a 200,000-token context window.

Text

o1

o1 is a text generation model from OpenAI that uses internal chain-of-thought reasoning before producing responses.

Vision

GPT-4o Mini Vision

GPT-4o Mini Vision is a vision-capable model from OpenAI with a 128,000-token context window and low-latency responses.

Vision

GPT-4o Vision

GPT-4o Vision is a multimodal model from OpenAI that processes both images and text within a 128,000-token context window.

Vision

GPT-4 Turbo Vision

GPT-4 Turbo Vision is a multimodal model from OpenAI that processes both images and text with a 128,000-token context window.

Image

GPT Image 2

GPT Image 2 is an image generation model from OpenAI that creates images from text prompts and supports source image inputs.

Image

GPT Image Latest

GPT Image Latest is an image generation model from OpenAI that produces images from text prompts and supports source image inputs.

Image

GPT Image 1.5

GPT Image 1.5 is an image generation model from OpenAI that supports source image inputs and configurable output settings.

Image

GPT Image 1

GPT Image 1 is an image generation model from OpenAI that supports configurable sizes, backgrounds, and source image inputs.

Video

Sora 2 Pro

Sora 2 Pro is a flagship video generation model from OpenAI that accepts text, image, and video inputs to produce video output.

Video

Sora 2

Sora 2 is a video generation model from OpenAI that creates videos from text prompts, images, and character references.

Transcription

GPT Transcribe

GPT Transcribe is a speech-to-text model from OpenAI, released in March 2025, that converts spoken audio into written text.

Transcription

Whisper-1

Whisper-1 is OpenAI's speech-to-text model that transcribes and translates audio at $0.006 per minute.

Transcription

Whisper Large v3 Turbo

Whisper Large v3 Turbo is an open-source speech-to-text model from OpenAI released in October 2024.

Transcription

Whisper Large v3

Whisper Large v3 is an open-source speech-to-text model from OpenAI that transcribes and translates audio across 99 languages.

Text to Speech

GPT-4o-mini TTS

GPT-4o-mini TTS is a text-to-speech model from OpenAI that converts text to audio with adjustable voice, speed, and format.

Text to Speech

TTS HD

TTS HD is a high-definition text-to-speech model from OpenAI that converts text into natural-sounding audio.

Text to Speech

TTS-1

TTS-1 is a text-to-speech model from OpenAI that converts written text into natural-sounding audio across multiple voices.

Embedding

OpenAI Embedding 3 Large

OpenAI Embedding 3 Large is a text embedding model from OpenAI with an 8,191 token context window.

Embedding

OpenAI Embedding 3 Small

OpenAI Embedding 3 Small is a text embedding model from OpenAI with an 8,191 token context window.

Text

GPT-4.5Deprecated

GPT-4.5 is a text generation model from OpenAI with a 128,000-token context window and up to 8,000 tokens of output.

Text

o1-previewDeprecated

o1-preview is a text generation model from OpenAI that uses internal chain-of-thought reasoning to tackle complex problems.

Text

o1-miniDeprecated

o1-mini is a text generation model from OpenAI trained with reinforcement learning to perform complex reasoning across coding, math, and science tasks.

Text

GPT-4o MiniDeprecated

GPT-4o Mini is a low-cost, low-latency text generation model from OpenAI with a 128,000-token context window.

Text

GPT-4oDeprecated

GPT-4o is an omni-modal model from OpenAI that accepts text, audio, and image inputs and generates text, audio, and image outputs.

Text

GPT-3.5Deprecated

GPT-3.5 is a text generation model from OpenAI designed for conversation, content creation, and problem-solving tasks.

Transcription

GPT-4o mini TranscribeDeprecated

GPT-4o mini Transcribe is a speech-to-text model from OpenAI that transcribes audio using GPT-4o mini.

Transcription

GPT-4o TranscribeDeprecated

GPT-4o Transcribe is a speech-to-text model from OpenAI that converts audio input into accurate text transcriptions.

Instruct

GPT-3.5 InstructDeprecated

GPT-3.5 Instruct is an instruction-tuned completion model from OpenAI with a 4,096 token context window.

Instruct

GPT-3Deprecated

GPT-3 (text-davinci-003) is an instruct-tuned language model from OpenAI with a 4,096-token context window.

Google47

Text

Gemini 3.7 Flash

Gemini 3.7 Flash is a text generation model from Google with a 1,048,576-token context window and configurable thinking levels.

Text

Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite is a text generation model from Google designed for low-cost, real-time tasks with a 1M token context window.

Text

Gemini 3.6 Flash

Gemini 3.6 Flash is a text and image understanding model from Google with a 1,048,576-token context window.

Text

Gemini 3.5 Flash

Gemini 3.5 Flash is a text generation model from Google with a 1,048,576-token context window, released in May 2026.

Text

Gemma 4 26B

Gemma 4 26B is a multimodal text generation model from Google with a 262,144 token context window.

Text

Gemma 4 31B

Gemma 4 31B is a multimodal text generation model from Google with a 262,144-token context window.

Text

Gemini 3.1 Pro

Gemini 3.1 Pro is a multimodal text generation model from Google with a 1,048,576-token context window.

Text

Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite is a text generation model from Google with a 1,048,576-token context window and real-time latency.

Text

Gemini 3 Flash

Gemini 3 Flash is a text and image understanding model from Google with a 1,048,576-token context window.

Text

Gemini 2.5 Flash Lite

Gemini 2.5 Flash Lite is a fast text and image understanding model from Google with a 1,000,000-token context window.

Text

Gemma 3.2

Gemma 3.2 is a 27-billion-parameter text generation model from Google with a 128,000-token context window.

Vision

Gemini 2.5 Pro Vision

Gemini 2.5 Pro Vision is Google's multimodal model supporting up to 1,048,576 tokens of context for vision and reasoning tasks.

Vision

Gemini 2.5 Flash Vision

Gemini 2.5 Flash Vision is Google's vision model offering a 1,048,576-token context window with real-time latency.

Vision

Gemini 2.5 Flash

Gemini 2.5 Flash is a multimodal vision model from Google with a 1,000,000-token context window and configurable thinking.

Vision

Gemini 2.5 Pro

Gemini 2.5 Pro is Google's flagship multimodal model with a 1,000,000-token context window and built-in reasoning capabilities.

Image

Gemini 3.1 Flash Lite Image

Gemini 3.1 Flash Lite Image is a Google image generation model supporting source image inputs, aspect ratio control, and custom sizing.

Image

Gemini 3.1 Flash Image

Gemini 3.1 Flash Image is an image generation model from Google that accepts source images and supports configurable aspect ratios and sizes.

Image

Gemini 3 Pro Image

Gemini 3 Pro Image is an image generation model from Google that accepts source images and configurable aspect ratios as inputs.

Image

Gemini 2.5 Flash Image

Gemini 2.5 Flash Image is an image generation model from Google built on the Gemini 2.5 Flash architecture.

Image

Imagen 3

Imagen 3 is a text-to-image generation model from Google DeepMind, available via fal with configurable aspect ratios and seed control.

Image

Imagen 3 Fast

Imagen 3 Fast is a text-to-image generation model from Google designed for rapid image synthesis with configurable aspect ratios.

Image

Imagen 4 Ultra

Imagen 4 Ultra is Google's image generation model offering high-resolution output with support for source image inputs and flexible aspect ratios.

Image

Imagen 4 Fast

Imagen 4 Fast is a Google image generation model optimized for speed, available via fal with a 10,000 token context window.

Video

Gemini Omni Flash

Gemini Omni Flash is a video generation model from Google that supports image and video inputs with multimodal editing capabilities.

Video

Veo 3.1 Lite

Veo 3.1 Lite is a video generation model from Google designed for fast, low-cost video creation with audio support.

Video

Veo 3.1 Fast

Veo 3.1 Fast is a video generation model from Google designed for rapid text-to-video and image-to-video synthesis.

Video

Veo 3.1

Veo 3.1 is a Google video generation model that creates videos from text prompts or images, with optional audio output.

Text to Speech

Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS is a text-to-speech model from Google that converts text into audio with style control.

Music

Lyria 3

Lyria 3 is a music generation model from Google that produces audio from text prompts with a 5000 token context window.

Music

Lyria 3 Pro

Lyria 3 Pro is a music generation model from Google designed to produce full audio tracks from text prompts.

Embedding

Gemini Embedding 2

Gemini Embedding 2 is a text embedding model from Google designed to convert text into dense vector representations.

Embedding

EmbeddingGemma 300M

EmbeddingGemma 300M is a 300-million-parameter text embedding model from Google with a 2048-token context window.

Embedding

Gemini Embedding

Gemini Embedding is a text embedding model from Google that converts text into vector representations with a 2048-token context window.

Document Extraction

Google Document AI

Google Document AI is a cloud-based document extraction service from Google that uses OCR to parse and structure document content.

Text

Gemini 3 ProDeprecated

Gemini 3 Pro is a multimodal text generation model from Google with a 1,000,000-token context window.

Text

Gemini 2.0 Flash ThinkingDeprecated

Gemini 2.0 Flash Thinking is a text generation model from Google that exposes its reasoning process to solve complex problems.

Text

Gemini 2.0 ProDeprecated

Gemini 2.0 Pro is a text generation model from Google with a 1,000,000-token context window built for coding and complex prompts.

Text

Gemini 1.5 FlashDeprecated

Gemini 1.5 Flash is a multimodal text generation model from Google designed for high-volume, cost-effective applications.

Text

Gemini 1.5 ProDeprecated

Gemini 1.5 Pro is a multimodal text generation model from Google with a 2,000,000 token context window.

Text

Gemini 1.0 ProDeprecated

Gemini 1.0 Pro is a text generation model from Google with a 30,720 token context window for content and problem-solving tasks.

Text

PaLM 2Deprecated

PaLM 2 is a text generation model from Google built on the Pathways AI architecture with an 8,000-token context window.

Vision

Gemini 1.5 Flash VisionDeprecated

Gemini 1.5 Flash Vision is Google's multimodal vision model built for high-volume, cost-effective applications with a 1M token context window.

Vision

Gemini 1.5 Pro VisionDeprecated

Gemini 1.5 Pro Vision is a multimodal model from Google that processes images, documents, and text within a 1,000,000-token context window.

Vision

Gemini 1.0 Pro VisionDeprecated

Gemini 1.0 Pro Vision is a Google vision model that accepts both text and image inputs with a 16,384 token context window.

Video

Veo 3 FastDeprecated

Veo 3 Fast is a video generation model from Google that supports image-to-video, audio generation, and flexible aspect ratios.

Video

Veo 3Deprecated

Veo 3 is a video generation model from Google that produces videos with native audio from text or image inputs.

Video

Veo 2Deprecated

Veo 2 is a video generation model from Google that produces videos from text prompts with configurable duration and aspect ratio.

Black Forest Labs13

Image

FLUX.2 [klein] 9B

FLUX.2 [klein] 9B is a 9-billion-parameter image generation model from Black Forest Labs released in January 2026.

Image

FLUX.2 [turbo]

FLUX.2 [turbo] is a text-to-image generation model from Black Forest Labs designed for fast, high-quality image output.

Image

FLUX.1 [dev] Ultra-Fast

FLUX.1 [dev] Ultra-Fast is an image generation model from Black Forest Labs optimized for speed with LoRA and inpainting support.

Image

FLUX.1 [schnell] LoRA

FLUX.1 [schnell] LoRA is an image generation model from Black Forest Labs that supports custom LoRA adapters for stylized output.

Image

FLUX.1 [dev] LoRA

FLUX.1 [dev] LoRA is an image generation model from Black Forest Labs that supports custom LoRA adapters and inpainting.

Image

FLUX.2 [max]

FLUX.2 [max] is a flagship image generation model from Black Forest Labs that accepts source images and custom dimensions as input.

Image

FLUX.2 [dev] LoRA

FLUX.2 [dev] LoRA is an image generation model from Black Forest Labs that supports LoRA adapters and reference image inputs.

Image

FLUX.2 [pro]

FLUX.2 [pro] is an image generation model from Black Forest Labs that accepts source images and configurable dimensions as input.

Image

FLUX.1 Kontext [max]

FLUX.1 Kontext [max] is an image generation model from Black Forest Labs that edits and remixes images using reference inputs and text prompts.

Image

FLUX.1 Kontext [pro]

FLUX.1 Kontext [pro] is an image generation model from Black Forest Labs that edits and remixes images using reference inputs and text prompts.

Image

FLUX 1.1 [pro] Ultra

FLUX 1.1 [pro] Ultra is a high-resolution image generation model from Black Forest Labs supporting outputs up to 4 megapixels.

Image

FLUX 1.1 [pro]

FLUX 1.1 [pro] is a text-to-image generation model from Black Forest Labs, released in October 2024 with prompt upsampling support.

Video

FLUX 3 Video

FLUX 3 Video is a video generation model from Black Forest Labs that supports audio, keyframes, and video continuation.

ByteDance14

Image

Seedream 5.0 Pro

Seedream 5.0 Pro is an image generation model from ByteDance that supports source image input and flexible aspect ratio and resolution controls.

Image

Seedream 4.0

Seedream 4.0 is an image generation model from ByteDance that accepts text prompts and source images to produce generated images.

Image

Seedream 5.0 Lite

Seedream 5.0 Lite is a ByteDance image generation model that accepts source images and custom dimensions at $0.035 per image.

Image

Seedream 4.5

Seedream 4.5 is a ByteDance image generation model that supports reference image inputs and custom output dimensions.

Video

Seedance 2.5 Turbo

Seedance 2.5 Turbo is a video generation model from ByteDance that supports text-to-video, image-to-video, and video editing modes.

Video

Seedance 2.5

Seedance 2.5 is a video generation model from ByteDance that supports text-to-video, image-to-video, video editing, and audio generation.

Video

Seedance 2.0 Mini

Seedance 2.0 Mini is a video generation model from ByteDance that supports text-to-video and image-to-video creation.

Video

Seedance 2.0 Fast Turbo

Seedance 2.0 Fast Turbo is a video generation model from ByteDance that supports text-to-video and image-to-video workflows.

Video

Seedance 2.0 Fast

Seedance 2.0 Fast is a video generation model from ByteDance that supports text-to-video and image-to-video creation.

Video

Seedance 2.0

Seedance 2.0 is a video generation model from ByteDance that supports text-to-video and image-to-video creation.

Video

DreamActor V2

DreamActor V2 is a ByteDance video generation model that animates a source image using motion from a reference video.

Video

Seedance 1.5 Pro

Seedance 1.5 Pro is a video generation model from ByteDance that animates images into videos with configurable resolution and audio.

Lip Sync

LatentSync

LatentSync is a lip sync model from ByteDance that synchronizes video facial movements to a provided audio track.

Lip Sync

Omni Human 1.5

Omni Human 1.5 is a lip sync model from ByteDance that animates a portrait image using an audio input.

Kling10

Image

Kling Image O3

Kling Image O3 is an image generation model from Kling that accepts reference images and supports selectable aspect ratios and resolutions.

Image

Kling Image O1

Kling Image O1 is an image generation model from Kling that accepts reference images and produces outputs at configurable aspect ratios and resolutions.

Video

Kling 3.0 Pro

Kling 3.0 Pro is a video generation model from Kling that supports both text-to-video and image-to-video creation.

Video

Kling 3.0

Kling 3.0 is a video generation model from Kling that supports text-to-video and image-to-video creation.

Video

Kling 2.6

Kling 2.6 is a video generation model from Kling that supports both text-to-video and image-to-video creation.

Video

Kling 3.0 Motion Control

Kling 3.0 Motion Control is a video generation model that transfers motion from a reference video onto a source image.

Video

Kling O3

Kling O3 is a video generation model from Kling that supports image-to-video, video-to-video, and audio generation.

Video

Kling 2.6 Pro Motion Control

Kling 2.6 Pro Motion Control is a video generation model that applies motion from a reference video to a source image.

Video

Kling O1

Kling O1 is a video generation model from Kling that creates videos from source images, videos, and reference frames.

Lip Sync

AI Avatar Standard

AI Avatar Standard is a lip sync model from Kling that animates a portrait image to match a provided audio track.

Mistral19

Text

Ministral 3 3B

Ministral 3 3B is a 3-billion-parameter open-source text generation model from Mistral with a 256,000-token context window.

Text

Ministral 3 8B

Ministral 3 8B is an open-source text generation model from Mistral with a 256,000-token context window.

Text

Ministral 3 14B

Ministral 3 14B is an open-source text generation model from Mistral with a 256,000-token context window.

Text

Mistral Large 3

Mistral Large 3 is an open-source text generation model from Mistral with a 256,000 token context window.

Text

Mistral Medium 3

Mistral Medium 3 is a text generation model from Mistral with a 128,000-token context window and a cost-efficient design.

Text

Mistral Small 3.1 (25.03)

Mistral Small 3.1 (25.03) is a text generation model from Mistral with a 128,000-token context window, released in March 2025.

Text

Mistral Codestral

Mistral Codestral is a code-focused text generation model from Mistral with a 32,000 token context window.

Text

Mistral Nemo

Mistral Nemo is a text generation model from Mistral with a 128,000-token context window, released in July 2024.

Transcription

Voxtral Mini 3B

Voxtral Mini 3B is a 3-billion-parameter open-source speech-to-text model released by Mistral in July 2025.

Document Extraction

Mistral OCR

Mistral OCR is a document extraction model from Mistral designed to recognize and extract text from images and documents.

Text

Mistral 8x7bDeprecated

Mistral 8x7b is a text generation model from Mistral AI using a mixture-of-experts architecture with a 32,768 token context window.

Text

Mixtral 8x7B InstructDeprecated

Mixtral 8x7B Instruct is a text generation model from Mistral using a sparse mixture-of-experts architecture with a 4096 token context window.

Text

Mistral Small 24.02Deprecated

Mistral Small 24.02 is a text generation model from Mistral with a 128,000-token context window, available via Amazon Bedrock.

Text

Mistral Large 24.07Deprecated

Mistral Large 24.07 is a text generation model from Mistral with a 128,000-token context window, available via Amazon Bedrock.

Text

Mistral Large 24.02Deprecated

Mistral Large 24.02 is a text generation model from Mistral with a 128,000-token context window, available via Amazon Bedrock.

Text

Mistral 7B InstructDeprecated

Mistral 7B Instruct is a text generation model from Mistral with a 4,096-token context window, available via Amazon Bedrock.

Text

Mixtral 8x22B InstructDeprecated

Mixtral 8x22B Instruct is a sparse mixture-of-experts text generation model from Mistral with a 64,000 token context window.

Text

Mistral 7B InstructDeprecated

Mistral 7B Instruct is a 7-billion-parameter text generation model from Mistral designed for instruction-following tasks.

Text

Mixtral 8x7B InstructDeprecated

Mixtral 8x7B Instruct is a sparse mixture-of-experts language model from Mistral, licensed under Apache 2.0.

Perplexity11

Text

Sonar Deep Research

Sonar Deep Research is a text generation model from Perplexity that performs multi-step web research and returns cited answers.

Text

Sonar Reasoning Pro

Sonar Reasoning Pro is a text generation model from Perplexity that combines live web search with extended reasoning in a 128K context window.

Text

Sonar

Sonar is a text generation model from Perplexity that combines LLM chat with real-time web search and citation support.

Text

Sonar Pro

Sonar Pro is a text generation model from Perplexity that combines a 200,000-token context window with live web search.

Embedding

Perplexity Embed v1 0.6B

Perplexity Embed v1 0.6B is a quantized embedding model from Perplexity with a 32,000-token context window.

Embedding

Perplexity Embed v1 4B

Perplexity Embed v1 4B is a 4-billion-parameter embedding model from Perplexity with a 32,000-token context window.

Text

Sonar ReasoningDeprecated

Sonar Reasoning is a text generation model from Perplexity combining chain-of-thought reasoning with live internet search and citations.

Text

Sonar Large OnlineDeprecated

Sonar Large Online is a text generation model from Perplexity that retrieves live web data and returns cited answers.

Text

Sonar Large ChatDeprecated

Sonar Large Chat is a text generation model from Perplexity built on Llama 3.1 with a 127,072-token context window.

Text

Sonar Small OnlineDeprecated

Sonar Small Online is a web-connected text generation model from Perplexity built on Llama 3.1 with a 127,072-token context window.

Text

Sonar Small ChatDeprecated

Sonar Small Chat is a text generation model from Perplexity built on Llama 3.1 with a 127,072 token context window.

Qwen23

Text

Qwen3.8 2.4T

Qwen3.8 2.4T is a mixture-of-experts text generation model from Qwen with a 262,144 token context window.

Text

Qwen3.5 Omni Flash

Qwen3.5 Omni Flash is a multimodal text generation model from Qwen with a 256,000 token context window.

Text

Qwen3.6 Flash

Qwen3.6 Flash is a text and image understanding model from Qwen with a 1,000,000-token context window and optional thinking mode.

Text

Qwen3.7 Max

Qwen3.7 Max is a text generation model from Qwen with a 1,000,000-token context window and optional deep thinking mode.

Text

Qwen3.7 Plus

Qwen3.7 Plus is a text generation model from Alibaba's Qwen team with a 1,000,000-token context window and optional deep thinking mode.

Text

Qwen3.6-35B-A3B

Qwen3.6-35B-A3B is a mixture-of-experts text generation model from Qwen with a 262,144 token context window.

Text

Qwen3.5 397B

Qwen3.5 397B is a multimodal text generation model from Qwen with a 262,144-token context window.

Text

Qwen3 235B

Qwen3 235B is a 235-billion-parameter mixture-of-experts text generation model from Qwen with a 262,144-token context window.

Image

Qwen Image 3.0

Qwen Image 3.0 is an image generation model from Qwen that supports text prompts and source image inputs at $0.03 per image.

Image

Qwen Image 3.0 Pro

Qwen Image 3.0 Pro is a flagship image generation model from Qwen supporting text-to-image and image-guided generation.

Image

Qwen Image Edit Plus

Qwen Image Edit Plus is an image generation and editing model from Qwen, priced at $0.02 per image with ControlNet pose support.

Image

Z Image Turbo Controlnet

Z Image Turbo Controlnet is an image generation model from Qwen that uses ControlNet to guide outputs from a reference image.

Image

Qwen 2 Pro

Qwen 2 Pro is an image generation model from Qwen supporting both text-to-image and image-to-image workflows at $0.07 per image.

Image

Qwen Image

Qwen Image is an image generation model from Qwen that supports LoRA adapters and source image inputs for guided generation.

Image

Z Image Turbo

Z Image Turbo is an image generation model from Qwen that supports LoRA adapters and source image input at $0.005 per image.

Transcription

Qwen3 ASR 1.7B

Qwen3 ASR 1.7B is an open-source multilingual speech-to-text model developed by Qwen with 1.7 billion parameters.

Text to Speech

Qwen3 TTS

Qwen3 TTS is an open-source multilingual text-to-speech model from Qwen with a 10,000-token context window.

Embedding

Qwen3 Embedding 0.6B

Qwen3 Embedding 0.6B is a compact embedding model from Qwen with a 32,768-token context window, served via DeepInfra.

Embedding

Qwen3 Embedding 4B

Qwen3 Embedding 4B is a text embedding model from Qwen with a 32,768-token context window, available via DeepInfra.

Embedding

Qwen3 Embedding 8B

Qwen3 Embedding 8B is an 8-billion-parameter text embedding model from Qwen with a 32,768-token context window.

Reranking

Qwen3 Reranker 0.6B

Qwen3 Reranker 0.6B is a lightweight reranking model from Qwen designed to reorder retrieved documents with a 32,768-token context window.

Reranking

Qwen3 Reranker 4B

Qwen3 Reranker 4B is a 4-billion-parameter reranking model from Qwen designed to score and reorder retrieved documents by relevance.

Reranking

Qwen3 Reranker 8B

Qwen3 Reranker 8B is an 8-billion-parameter reranking model from Qwen designed to score and reorder documents for retrieval pipelines.

Wan13

Image

Wan 2.7 Pro

Wan 2.7 Pro is an image generation model from Wan, available via Alibaba, priced at $0.075 per image.

Image

Wan 2.7

Wan 2.7 is an image generation model from Alibaba's Wan that supports source image inputs and custom dimensions at $0.03 per image.

Image

Wan 2.7 Pro

Wan 2.7 Pro is an image generation model from Wan that supports source image input and configurable output dimensions.

Image

Wan 2.7

Wan 2.7 is an image generation model from Wan, available via Wavespeed, priced at $0.03 per image.

Image

Wan 2.6

Wan 2.6 is an image generation model from Wan that accepts source images and text prompts to produce new images at $0.03 per image.

Image

Wan 2.5

Wan 2.5 is an image generation model from Wan that accepts source images and produces new images at $0.03 per image.

Video

Wan 3.0

Wan 3.0 is a video generation model that converts text prompts and images into video clips with optional audio output.

Video

Wan 2.7

Wan 2.7 is a video generation model from Alibaba that supports text-to-video, image-to-video, video editing, and audio generation.

Video

Wan 2.7

Wan 2.7 is a video generation model supporting text-to-video and image-to-video creation with audio and reference image inputs.

Video

Wan 2.6

Wan 2.6 is a video generation model that creates videos from source images and audio with multi-shot control.

Video

Wan 2.5

Wan 2.5 is a video generation model that animates source images with optional audio input, priced at $0.05–$0.15 per second.

Video

Wan 2.2

Wan 2.2 is a video generation model from Wan that converts source images into videos with customizable LoRA fine-tuning.

Image

Wan 2.2Deprecated

Wan 2.2 is an image generation model from Wan, available via Wavespeed, supporting LoRA fine-tuning and configurable resolution.

X.ai22

Text

Grok 4.6

Grok 4.6 is a text generation model from X.ai featuring a 500,000-token context window and configurable reasoning effort.

Text

Grok 4.5

Grok 4.5 is a text generation model from X.ai featuring a 500,000-token context window and configurable reasoning effort.

Text

Grok Build 0.1

Grok Build 0.1 is a text generation model from X.ai featuring a 256,000-token context window for extended conversations.

Text

Grok 4.3

Grok 4.3 is a text generation model from X.ai featuring a 1,000,000-token context window and adjustable reasoning effort.

Text

Grok 4.20 Reasoning

Grok 4.20 Reasoning is a text generation model from X.ai with a 2,000,000-token context window and built-in reasoning.

Text

Grok 4.20

Grok 4.20 is a text generation model from X.ai featuring a 2,000,000 token context window for long-document tasks.

Text

Grok 4.1 Fast Reasoning

Grok 4.1 Fast Reasoning is a text generation model from X.ai designed for rapid reasoning tasks with a 2,000,000 token context window.

Text

Grok 4.1 Fast

Grok 4.1 Fast is a text generation model from X.ai offering a 2,000,000 token context window for large-scale language tasks.

Text

Grok 4 Fast Reasoning

Grok 4 Fast Reasoning is a text generation model from X.ai designed for rapid reasoning tasks with a 2,000,000 token context window.

Text

Grok 4 Fast

Grok 4 Fast is a text and image understanding model from X.ai with a 2,000,000-token context window.

Text

Grok 4

Grok 4 is a text generation model from X.ai with a 256,000-token context window, released in July 2025.

Text

Grok 3 Mini Fast

Grok 3 Mini Fast is a text generation model from X.ai designed for fast inference with a 131,072 token context window.

Text

Grok 3 Mini

Grok 3 Mini is a text generation model from X.ai with a 131,072-token context window and configurable reasoning effort.

Text

Grok 3 Fast

Grok 3 Fast is a text generation model from X.ai designed for speed, with a 131,072 token context window.

Text

Grok 3

Grok 3 is a text generation model from X.ai with a 131,072-token context window, released in February 2025.

Vision

Grok 4.3 Vision

Grok 4.3 Vision is a vision-capable model from X.ai featuring a 2,000,000-token context window for processing large inputs.

Vision

Grok 2 Vision

Grok 2 Vision is a multimodal model from X.ai that processes images and text within a 32,768 token context window.

Image

Grok Imagine 2.0

Grok Imagine 2.0 is an image generation model from X.ai that accepts source images and supports configurable aspect ratio, resolution, and quality.

Image

Grok Imagine Pro

Grok Imagine Pro is an image generation model from X.ai that accepts a source image and aspect ratio as inputs.

Image

Grok Imagine

Grok Imagine is an image generation model from X.ai that accepts a source image and aspect ratio as inputs at $0.02 per image.

Video

Grok Imagine 1.5

Grok Imagine 1.5 is a video generation model from X.ai that accepts text, image, and video inputs to produce generated video output.

Video

Grok Imagine

Grok Imagine is a video generation model from X.ai that creates video from text prompts, source images, or source video clips.

Build with any AI model

No API keys required. Start building AI-powered workflows in minutes.

Get Started Free