AI Models
Explore 344+ AI models available in MindStudio. From large language models to image generators — no API keys required.
Anthropic13
Claude 5 Opus
Claude 5 Opus is a text generation model from Anthropic with a 1,000,000-token context window and native tool use.
TextClaude 5 Sonnet
Claude 5 Sonnet is Anthropic's flagship text and vision model with a 1,000,000-token context window and extended reasoning.
TextClaude 5 Fable
Claude 5 Fable is a text generation model from Anthropic with a 1,000,000-token context window and built-in reasoning capabilities.
TextClaude 4.8 Opus
Claude 4.8 Opus is a text generation model from Anthropic with a 1,000,000-token context window and native tool use.
TextClaude 4.7 Opus
Claude 4.7 Opus is a text generation model from Anthropic with a 1,000,000-token context window and native tool use.
TextClaude 4.6 Sonnet
Claude 4.6 Sonnet is a text generation model from Anthropic with a 1,000,000-token context window and native tool use.
TextClaude 4.6 Opus
Claude 4.6 Opus is a text generation model from Anthropic with a 1,000,000-token context window and native tool and MCP support.
TextClaude 4.5 Opus
Claude 4.5 Opus is a text generation model from Anthropic with a 200,000-token context window and extended reasoning support.
TextClaude 4.5 Haiku
Claude 4.5 Haiku is a text generation model from Anthropic with a 200,000-token context window and image input support.
TextClaude 4.5 Sonnet
Claude 4.5 Sonnet is a text generation model from Anthropic with a 200,000-token context window and native tool and MCP support.
TextClaude 3 SonnetDeprecated
Claude 3 Sonnet is a text generation model from Anthropic with a 200,000 token context window, available via Amazon Bedrock.
TextClaude 3 HaikuDeprecated
Claude 3 Haiku is a text generation model from Anthropic with a 200,000 token context window, available via Amazon Bedrock.
TextClaude InstantDeprecated
Claude Instant is a text generation model from Anthropic optimized for speed and efficiency with a 100,000 token context window.
OpenAI47
GPT 5.6 Luna
GPT 5.6 Luna is a text and image-input model from OpenAI with a 1,050,000-token context window.
TextGPT 5.6 Terra
GPT 5.6 Terra is a text generation model from OpenAI with a 1,050,000-token context window and image input support.
TextGPT 5.6 Sol
GPT 5.6 Sol is a text generation model from OpenAI with a 1,050,000-token context window and image input support.
TextGPT-5.5
GPT-5.5 is a multimodal text generation model from OpenAI with a 1,000,000-token context window and adjustable reasoning effort.
TextGPT 5.5 Pro
GPT 5.5 Pro is a text generation model from OpenAI with a 1,000,000-token context window and reasoning capabilities.
TextGPT 5.4 Pro
GPT 5.4 Pro is a text generation model from OpenAI with a 1,000,000-token context window and 128,000-token max response size.
TextGPT 5.4
GPT 5.4 is a text generation model from OpenAI with a 1,000,000 token context window and 128,000 token max response size.
TextGPT‑5.2 Pro
GPT‑5.2 Pro is a text generation model from OpenAI with a 400,000-token context window and adjustable reasoning effort.
TextGPT-5.1
GPT-5.1 is a text generation model from OpenAI with a 400,000 token context window and adjustable reasoning effort.
Texto3-pro
o3-pro is a text generation model from OpenAI that applies more compute to reasoning, with a 200,000-token context window.
TextGPT-5 nano
GPT-5 nano is a text generation model from OpenAI with a 400,000 token context window and tool-calling support.
TextGPT-5 mini
GPT-5 mini is a text generation model from OpenAI with a 400,000-token context window and tool-use support.
TextGPT-5
GPT-5 is a text generation model from OpenAI with a 400,000-token context window and adjustable reasoning effort.
TextGPT OSS 120B
GPT OSS 120B is an open-source text generation model from OpenAI, served via Groq with a 128,000-token context window.
TextGPT OSS 20B
GPT OSS 20B is an open-source text generation model from OpenAI with a 128,000-token context window, served via Groq.
Texto3
o3 is a text generation model from OpenAI released in April 2025, featuring a 200,000 token context window.
Texto1-pro
o1-pro is a text generation model from OpenAI that applies extended compute to reasoning tasks, with a 200,000-token context window.
Texto3-mini
o3-mini is a text generation model from OpenAI designed for reasoning-intensive tasks with a 200,000-token context window.
Texto1
o1 is a text generation model from OpenAI that uses internal chain-of-thought reasoning before producing responses.
VisionGPT-4o Mini Vision
GPT-4o Mini Vision is a vision-capable model from OpenAI with a 128,000-token context window and low-latency responses.
VisionGPT-4o Vision
GPT-4o Vision is a multimodal model from OpenAI that processes both images and text within a 128,000-token context window.
VisionGPT-4 Turbo Vision
GPT-4 Turbo Vision is a multimodal model from OpenAI that processes both images and text with a 128,000-token context window.
ImageGPT Image 2
GPT Image 2 is an image generation model from OpenAI that creates images from text prompts and supports source image inputs.
ImageGPT Image Latest
GPT Image Latest is an image generation model from OpenAI that produces images from text prompts and supports source image inputs.
ImageGPT Image 1.5
GPT Image 1.5 is an image generation model from OpenAI that supports source image inputs and configurable output settings.
ImageGPT Image 1
GPT Image 1 is an image generation model from OpenAI that supports configurable sizes, backgrounds, and source image inputs.
VideoSora 2 Pro
Sora 2 Pro is a flagship video generation model from OpenAI that accepts text, image, and video inputs to produce video output.
VideoSora 2
Sora 2 is a video generation model from OpenAI that creates videos from text prompts, images, and character references.
TranscriptionGPT Transcribe
GPT Transcribe is a speech-to-text model from OpenAI, released in March 2025, that converts spoken audio into written text.
TranscriptionWhisper-1
Whisper-1 is OpenAI's speech-to-text model that transcribes and translates audio at $0.006 per minute.
TranscriptionWhisper Large v3 Turbo
Whisper Large v3 Turbo is an open-source speech-to-text model from OpenAI released in October 2024.
TranscriptionWhisper Large v3
Whisper Large v3 is an open-source speech-to-text model from OpenAI that transcribes and translates audio across 99 languages.
Text to SpeechGPT-4o-mini TTS
GPT-4o-mini TTS is a text-to-speech model from OpenAI that converts text to audio with adjustable voice, speed, and format.
Text to SpeechTTS HD
TTS HD is a high-definition text-to-speech model from OpenAI that converts text into natural-sounding audio.
Text to SpeechTTS-1
TTS-1 is a text-to-speech model from OpenAI that converts written text into natural-sounding audio across multiple voices.
EmbeddingOpenAI Embedding 3 Large
OpenAI Embedding 3 Large is a text embedding model from OpenAI with an 8,191 token context window.
EmbeddingOpenAI Embedding 3 Small
OpenAI Embedding 3 Small is a text embedding model from OpenAI with an 8,191 token context window.
TextGPT-4.5Deprecated
GPT-4.5 is a text generation model from OpenAI with a 128,000-token context window and up to 8,000 tokens of output.
Texto1-previewDeprecated
o1-preview is a text generation model from OpenAI that uses internal chain-of-thought reasoning to tackle complex problems.
Texto1-miniDeprecated
o1-mini is a text generation model from OpenAI trained with reinforcement learning to perform complex reasoning across coding, math, and science tasks.
TextGPT-4o MiniDeprecated
GPT-4o Mini is a low-cost, low-latency text generation model from OpenAI with a 128,000-token context window.
TextGPT-4oDeprecated
GPT-4o is an omni-modal model from OpenAI that accepts text, audio, and image inputs and generates text, audio, and image outputs.
TextGPT-3.5Deprecated
GPT-3.5 is a text generation model from OpenAI designed for conversation, content creation, and problem-solving tasks.
TranscriptionGPT-4o mini TranscribeDeprecated
GPT-4o mini Transcribe is a speech-to-text model from OpenAI that transcribes audio using GPT-4o mini.
TranscriptionGPT-4o TranscribeDeprecated
GPT-4o Transcribe is a speech-to-text model from OpenAI that converts audio input into accurate text transcriptions.
InstructGPT-3.5 InstructDeprecated
GPT-3.5 Instruct is an instruction-tuned completion model from OpenAI with a 4,096 token context window.
InstructGPT-3Deprecated
GPT-3 (text-davinci-003) is an instruct-tuned language model from OpenAI with a 4,096-token context window.
Google47
Gemini 3.7 Flash
Gemini 3.7 Flash is a text generation model from Google with a 1,048,576-token context window and configurable thinking levels.
TextGemini 3.5 Flash Lite
Gemini 3.5 Flash Lite is a text generation model from Google designed for low-cost, real-time tasks with a 1M token context window.
TextGemini 3.6 Flash
Gemini 3.6 Flash is a text and image understanding model from Google with a 1,048,576-token context window.
TextGemini 3.5 Flash
Gemini 3.5 Flash is a text generation model from Google with a 1,048,576-token context window, released in May 2026.
TextGemma 4 26B
Gemma 4 26B is a multimodal text generation model from Google with a 262,144 token context window.
TextGemma 4 31B
Gemma 4 31B is a multimodal text generation model from Google with a 262,144-token context window.
TextGemini 3.1 Pro
Gemini 3.1 Pro is a multimodal text generation model from Google with a 1,048,576-token context window.
TextGemini 3.1 Flash Lite
Gemini 3.1 Flash Lite is a text generation model from Google with a 1,048,576-token context window and real-time latency.
TextGemini 3 Flash
Gemini 3 Flash is a text and image understanding model from Google with a 1,048,576-token context window.
TextGemini 2.5 Flash Lite
Gemini 2.5 Flash Lite is a fast text and image understanding model from Google with a 1,000,000-token context window.
TextGemma 3.2
Gemma 3.2 is a 27-billion-parameter text generation model from Google with a 128,000-token context window.
VisionGemini 2.5 Pro Vision
Gemini 2.5 Pro Vision is Google's multimodal model supporting up to 1,048,576 tokens of context for vision and reasoning tasks.
VisionGemini 2.5 Flash Vision
Gemini 2.5 Flash Vision is Google's vision model offering a 1,048,576-token context window with real-time latency.
VisionGemini 2.5 Flash
Gemini 2.5 Flash is a multimodal vision model from Google with a 1,000,000-token context window and configurable thinking.
VisionGemini 2.5 Pro
Gemini 2.5 Pro is Google's flagship multimodal model with a 1,000,000-token context window and built-in reasoning capabilities.
ImageGemini 3.1 Flash Lite Image
Gemini 3.1 Flash Lite Image is a Google image generation model supporting source image inputs, aspect ratio control, and custom sizing.
ImageGemini 3.1 Flash Image
Gemini 3.1 Flash Image is an image generation model from Google that accepts source images and supports configurable aspect ratios and sizes.
ImageGemini 3 Pro Image
Gemini 3 Pro Image is an image generation model from Google that accepts source images and configurable aspect ratios as inputs.
ImageGemini 2.5 Flash Image
Gemini 2.5 Flash Image is an image generation model from Google built on the Gemini 2.5 Flash architecture.
ImageImagen 3
Imagen 3 is a text-to-image generation model from Google DeepMind, available via fal with configurable aspect ratios and seed control.
ImageImagen 3 Fast
Imagen 3 Fast is a text-to-image generation model from Google designed for rapid image synthesis with configurable aspect ratios.
ImageImagen 4 Ultra
Imagen 4 Ultra is Google's image generation model offering high-resolution output with support for source image inputs and flexible aspect ratios.
ImageImagen 4 Fast
Imagen 4 Fast is a Google image generation model optimized for speed, available via fal with a 10,000 token context window.
VideoGemini Omni Flash
Gemini Omni Flash is a video generation model from Google that supports image and video inputs with multimodal editing capabilities.
VideoVeo 3.1 Lite
Veo 3.1 Lite is a video generation model from Google designed for fast, low-cost video creation with audio support.
VideoVeo 3.1 Fast
Veo 3.1 Fast is a video generation model from Google designed for rapid text-to-video and image-to-video synthesis.
VideoVeo 3.1
Veo 3.1 is a Google video generation model that creates videos from text prompts or images, with optional audio output.
Text to SpeechGemini 3.1 Flash TTS
Gemini 3.1 Flash TTS is a text-to-speech model from Google that converts text into audio with style control.
MusicLyria 3
Lyria 3 is a music generation model from Google that produces audio from text prompts with a 5000 token context window.
MusicLyria 3 Pro
Lyria 3 Pro is a music generation model from Google designed to produce full audio tracks from text prompts.
EmbeddingGemini Embedding 2
Gemini Embedding 2 is a text embedding model from Google designed to convert text into dense vector representations.
EmbeddingEmbeddingGemma 300M
EmbeddingGemma 300M is a 300-million-parameter text embedding model from Google with a 2048-token context window.
EmbeddingGemini Embedding
Gemini Embedding is a text embedding model from Google that converts text into vector representations with a 2048-token context window.
Document ExtractionGoogle Document AI
Google Document AI is a cloud-based document extraction service from Google that uses OCR to parse and structure document content.
TextGemini 3 ProDeprecated
Gemini 3 Pro is a multimodal text generation model from Google with a 1,000,000-token context window.
TextGemini 2.0 Flash ThinkingDeprecated
Gemini 2.0 Flash Thinking is a text generation model from Google that exposes its reasoning process to solve complex problems.
TextGemini 2.0 ProDeprecated
Gemini 2.0 Pro is a text generation model from Google with a 1,000,000-token context window built for coding and complex prompts.
TextGemini 1.5 FlashDeprecated
Gemini 1.5 Flash is a multimodal text generation model from Google designed for high-volume, cost-effective applications.
TextGemini 1.5 ProDeprecated
Gemini 1.5 Pro is a multimodal text generation model from Google with a 2,000,000 token context window.
TextGemini 1.0 ProDeprecated
Gemini 1.0 Pro is a text generation model from Google with a 30,720 token context window for content and problem-solving tasks.
TextPaLM 2Deprecated
PaLM 2 is a text generation model from Google built on the Pathways AI architecture with an 8,000-token context window.
VisionGemini 1.5 Flash VisionDeprecated
Gemini 1.5 Flash Vision is Google's multimodal vision model built for high-volume, cost-effective applications with a 1M token context window.
VisionGemini 1.5 Pro VisionDeprecated
Gemini 1.5 Pro Vision is a multimodal model from Google that processes images, documents, and text within a 1,000,000-token context window.
VisionGemini 1.0 Pro VisionDeprecated
Gemini 1.0 Pro Vision is a Google vision model that accepts both text and image inputs with a 16,384 token context window.
VideoVeo 3 FastDeprecated
Veo 3 Fast is a video generation model from Google that supports image-to-video, audio generation, and flexible aspect ratios.
VideoVeo 3Deprecated
Veo 3 is a video generation model from Google that produces videos with native audio from text or image inputs.
VideoVeo 2Deprecated
Veo 2 is a video generation model from Google that produces videos from text prompts with configurable duration and aspect ratio.
Amazon4
Amazon Nova 2 Lite
Amazon Nova 2 Lite is a text and image-input model from Amazon with a 1,000,000-token context window.
TextAmazon Nova Pro
Amazon Nova Pro is a text generation model from Amazon with a 300,000-token context window, available via Amazon Bedrock.
TextAmazon Nova Lite
Amazon Nova Lite is a multimodal text generation model from Amazon with a 300,000 token context window.
TextAmazon Nova Micro
Amazon Nova Micro is a text-only generation model from Amazon with a 128,000-token context window and a 5,000-token response limit.
Black Forest Labs13
FLUX.2 [klein] 9B
FLUX.2 [klein] 9B is a 9-billion-parameter image generation model from Black Forest Labs released in January 2026.
ImageFLUX.2 [turbo]
FLUX.2 [turbo] is a text-to-image generation model from Black Forest Labs designed for fast, high-quality image output.
ImageFLUX.1 [dev] Ultra-Fast
FLUX.1 [dev] Ultra-Fast is an image generation model from Black Forest Labs optimized for speed with LoRA and inpainting support.
ImageFLUX.1 [schnell] LoRA
FLUX.1 [schnell] LoRA is an image generation model from Black Forest Labs that supports custom LoRA adapters for stylized output.
ImageFLUX.1 [dev] LoRA
FLUX.1 [dev] LoRA is an image generation model from Black Forest Labs that supports custom LoRA adapters and inpainting.
ImageFLUX.2 [max]
FLUX.2 [max] is a flagship image generation model from Black Forest Labs that accepts source images and custom dimensions as input.
ImageFLUX.2 [dev] LoRA
FLUX.2 [dev] LoRA is an image generation model from Black Forest Labs that supports LoRA adapters and reference image inputs.
ImageFLUX.2 [pro]
FLUX.2 [pro] is an image generation model from Black Forest Labs that accepts source images and configurable dimensions as input.
ImageFLUX.1 Kontext [max]
FLUX.1 Kontext [max] is an image generation model from Black Forest Labs that edits and remixes images using reference inputs and text prompts.
ImageFLUX.1 Kontext [pro]
FLUX.1 Kontext [pro] is an image generation model from Black Forest Labs that edits and remixes images using reference inputs and text prompts.
ImageFLUX 1.1 [pro] Ultra
FLUX 1.1 [pro] Ultra is a high-resolution image generation model from Black Forest Labs supporting outputs up to 4 megapixels.
ImageFLUX 1.1 [pro]
FLUX 1.1 [pro] is a text-to-image generation model from Black Forest Labs, released in October 2024 with prompt upsampling support.
VideoFLUX 3 Video
FLUX 3 Video is a video generation model from Black Forest Labs that supports audio, keyframes, and video continuation.
ByteDance14
Seedream 5.0 Pro
Seedream 5.0 Pro is an image generation model from ByteDance that supports source image input and flexible aspect ratio and resolution controls.
ImageSeedream 4.0
Seedream 4.0 is an image generation model from ByteDance that accepts text prompts and source images to produce generated images.
ImageSeedream 5.0 Lite
Seedream 5.0 Lite is a ByteDance image generation model that accepts source images and custom dimensions at $0.035 per image.
ImageSeedream 4.5
Seedream 4.5 is a ByteDance image generation model that supports reference image inputs and custom output dimensions.
VideoSeedance 2.5 Turbo
Seedance 2.5 Turbo is a video generation model from ByteDance that supports text-to-video, image-to-video, and video editing modes.
VideoSeedance 2.5
Seedance 2.5 is a video generation model from ByteDance that supports text-to-video, image-to-video, video editing, and audio generation.
VideoSeedance 2.0 Mini
Seedance 2.0 Mini is a video generation model from ByteDance that supports text-to-video and image-to-video creation.
VideoSeedance 2.0 Fast Turbo
Seedance 2.0 Fast Turbo is a video generation model from ByteDance that supports text-to-video and image-to-video workflows.
VideoSeedance 2.0 Fast
Seedance 2.0 Fast is a video generation model from ByteDance that supports text-to-video and image-to-video creation.
VideoSeedance 2.0
Seedance 2.0 is a video generation model from ByteDance that supports text-to-video and image-to-video creation.
VideoDreamActor V2
DreamActor V2 is a ByteDance video generation model that animates a source image using motion from a reference video.
VideoSeedance 1.5 Pro
Seedance 1.5 Pro is a video generation model from ByteDance that animates images into videos with configurable resolution and audio.
Lip SyncLatentSync
LatentSync is a lip sync model from ByteDance that synchronizes video facial movements to a provided audio track.
Lip SyncOmni Human 1.5
Omni Human 1.5 is a lip sync model from ByteDance that animates a portrait image using an audio input.
Cohere6
Cohere Embed 4
Cohere Embed 4 is an embedding model from Cohere with a 128,000-token context window, released in April 2025.
RerankingCohere Rerank 4 Fast
Cohere Rerank 4 Fast is a reranking model from Cohere designed to score and reorder search results with a 4096 token context window.
RerankingCohere Rerank 4 Pro
Cohere Rerank 4 Pro is a reranking model from Cohere designed to improve search result relevance with a 4096-token context window.
RerankingCohere Rerank 3.5
Cohere Rerank 3.5 is a reranking model from Cohere that reorders search results by semantic relevance with a 4096 token context window.
TextCommand R+Deprecated
Command R+ is a text generation model from Cohere designed for RAG workflows with a 128,000 token context window.
TextCommand RDeprecated
Command R is a text generation model from Cohere with a 128,000-token context window and multilingual support.
DeepSeek9
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a text generation model from DeepSeek with a 1,048,576-token context window and configurable reasoning effort.
TextDeepSeek V4 Pro
DeepSeek V4 Pro is a text generation model from DeepSeek with a 1,000,000-token context window and adjustable reasoning effort.
TextDeepSeek V4 Flash
DeepSeek V4 Flash is a text generation model from DeepSeek with a 1,000,000-token context window and configurable reasoning effort.
TextDeepSeek V3.2
DeepSeek V3.2 is a text generation model from DeepSeek with a 160,000-token context window, released in December 2025.
TextDeepSeek V3.1
DeepSeek V3.1 is a text generation model from DeepSeek with a 128,000-token context window, released in August 2025.
TextDeepSeek R1 Turbo
DeepSeek R1 Turbo is a reasoning-focused text generation model from DeepSeek with a 128,000 token context window.
TextDeepSeek-R1
DeepSeek-R1 is a reasoning-focused text generation model from DeepSeek with a 64,000-token context window.
TextDeepSeek-V3
DeepSeek-V3 is a text generation model from DeepSeek with a 128,000-token context window, released in December 2024.
TextDeepSeek-V3Deprecated
DeepSeek-V3 is a text generation model from DeepSeek with a 128,000 token context window and up to 8,000 token responses.
ElevenLabs4
Scribe v2
Scribe v2 is a speech-to-text transcription model from ElevenLabs that supports speaker identification.
TranscriptionScribe v1
Scribe v1 is a speech-to-text transcription model from ElevenLabs, available on MindStudio for audio transcription tasks.
Text to SpeechElevenLabs TTS
ElevenLabs TTS is a text-to-speech model from ElevenLabs that converts text into natural-sounding audio speech.
MusicElevenLabs Music
ElevenLabs Music is a music generation model that produces audio tracks from text prompts with a 2000 token context window.
Ideogram8
Ideogram V4
Ideogram V4 is an image generation model from Ideogram that produces images from text prompts with source image input support.
ImageIdeogram V3 Remix
Ideogram V3 Remix is an image generation model from Ideogram that transforms existing images using text prompts and style presets.
ImageIdeogram V3
Ideogram V3 is a flagship image generation model from Ideogram, released in March 2025, with a 10,000-token context window.
ImageIdeogram Upscale
Ideogram Upscale is an image generation model from Ideogram that enhances image resolution with adjustable detail and resemblance controls.
ImageIdeogram V2 Remix
Ideogram V2 Remix is an image generation model from Ideogram that transforms existing images using text prompts and adjustable style controls.
ImageIdeogram V1 Remix
Ideogram V1 Remix is an image generation model from Ideogram that transforms existing images using text prompts and adjustable influence weights.
ImageIdeogram V2
Ideogram V2 is an image generation model from Ideogram that produces images from text prompts, released in August 2024.
VisionIdeogram VisionDeprecated
Ideogram Vision is a multimodal vision model from Ideogram that analyzes and interprets images alongside text prompts.
Kling10
Kling Image O3
Kling Image O3 is an image generation model from Kling that accepts reference images and supports selectable aspect ratios and resolutions.
ImageKling Image O1
Kling Image O1 is an image generation model from Kling that accepts reference images and produces outputs at configurable aspect ratios and resolutions.
VideoKling 3.0 Pro
Kling 3.0 Pro is a video generation model from Kling that supports both text-to-video and image-to-video creation.
VideoKling 3.0
Kling 3.0 is a video generation model from Kling that supports text-to-video and image-to-video creation.
VideoKling 2.6
Kling 2.6 is a video generation model from Kling that supports both text-to-video and image-to-video creation.
VideoKling 3.0 Motion Control
Kling 3.0 Motion Control is a video generation model that transfers motion from a reference video onto a source image.
VideoKling O3
Kling O3 is a video generation model from Kling that supports image-to-video, video-to-video, and audio generation.
VideoKling 2.6 Pro Motion Control
Kling 2.6 Pro Motion Control is a video generation model that applies motion from a reference video to a source image.
VideoKling O1
Kling O1 is a video generation model from Kling that creates videos from source images, videos, and reference frames.
Lip SyncAI Avatar Standard
AI Avatar Standard is a lip sync model from Kling that animates a portrait image to match a provided audio track.
Lightricks3
LTX-2.3 LoRA
LTX-2.3 LoRA is a video generation model from Lightricks that supports text-to-video, image-to-video, and custom LoRA fine-tuning.
VideoLTX-2.3
LTX-2.3 is a video generation model from Lightricks that supports both text-to-video and image-to-video creation.
VideoLTX-2 19b
LTX-2 19b is a 19-billion-parameter video generation model from Lightricks that supports image-to-video synthesis and LoRA fine-tuning.
Luma Labs6
UNI 1.1 Max
UNI 1.1 Max is an image generation model from Luma Labs supporting reference images, style selection, and web search.
ImageUNI 1.1
UNI 1.1 is an image generation model from Luma Labs supporting text prompts, reference images, and multiple aspect ratios.
ImagePhoton 1 Flash
Photon 1 Flash is an image generation model from Luma Labs designed for speed and flexible aspect ratio output.
ImagePhoton 1
Photon 1 is an image generation model from Luma Labs that produces images from text prompts with selectable aspect ratios.
VideoRay 2
Ray 2 is a video generation model from Luma Labs that creates videos from text prompts and reference images.
VideoRay Flash 2
Ray Flash 2 is a video generation model from Luma Labs designed for fast video creation with configurable resolution, aspect ratio, and keyframe inputs.
Meta9
Muse Glimmer 30B
Muse Glimmer 30B is a multimodal text generation model from Meta with a 131,072 token context window.
TextMuse Spark 1.1
Muse Spark 1.1 is a text generation model from Meta supporting images and a 1,048,576-token context window.
TextLlama 4 Scout
Llama 4 Scout is a 17B active parameter mixture-of-experts language model from Meta with a 130,000 token context window.
TextLlama 4 Maverick
Llama 4 Maverick is a 17B active parameter mixture-of-experts text generation model from Meta with a 130,000 token context window.
TextLlama 3 8BDeprecated
Llama 3 8B is a text generation model from Meta with an 8,192-token context window, served via Groq.
TextLlama 3 70BDeprecated
Llama 3 70B is a text generation model from Meta, instruction-tuned for dialogue with an 8,192-token context window.
TextCode LlamaDeprecated
Code Llama is a 34-billion-parameter instruction-tuned model from Meta built for code generation, comprehension, and debugging.
TextLlama-2 13B ChatDeprecated
Llama-2 13B Chat is a conversational text generation model from Meta with a 4096 token context window.
TextLlama-2 70B ChatDeprecated
Llama-2 70B Chat is a text generation model from Meta with 70 billion parameters and a 4096 token context window.
MiniMax6
MiniMax M3
MiniMax M3 is a multimodal text generation model from MiniMax with a 524,288-token context window.
VideoMiniMax H3
MiniMax H3 is a video generation model from MiniMax that supports text-to-video, image-to-video, and reference-based video creation at up to 2K resolution.
VideoHailuo 2.3 Pro
Hailuo 2.3 Pro is a video generation model from MiniMax that supports both text-to-video and image-to-video creation.
Text to SpeechMinimax Speech 2.8 HD
Minimax Speech 2.8 HD is a text-to-speech model from MiniMax supporting emotion, pitch, speed, and multi-format audio output.
MusicMiniMax Music 3.0
MiniMax Music 3.0 is a music generation model from MiniMax that creates songs from lyrics or as instrumentals at $0.15 per song.
MusicMiniMax Music 2.5
MiniMax Music 2.5 is a music generation model that converts lyrics into songs with configurable bitrate and sample rate.
Mistral19
Ministral 3 3B
Ministral 3 3B is a 3-billion-parameter open-source text generation model from Mistral with a 256,000-token context window.
TextMinistral 3 8B
Ministral 3 8B is an open-source text generation model from Mistral with a 256,000-token context window.
TextMinistral 3 14B
Ministral 3 14B is an open-source text generation model from Mistral with a 256,000-token context window.
TextMistral Large 3
Mistral Large 3 is an open-source text generation model from Mistral with a 256,000 token context window.
TextMistral Medium 3
Mistral Medium 3 is a text generation model from Mistral with a 128,000-token context window and a cost-efficient design.
TextMistral Small 3.1 (25.03)
Mistral Small 3.1 (25.03) is a text generation model from Mistral with a 128,000-token context window, released in March 2025.
TextMistral Codestral
Mistral Codestral is a code-focused text generation model from Mistral with a 32,000 token context window.
TextMistral Nemo
Mistral Nemo is a text generation model from Mistral with a 128,000-token context window, released in July 2024.
TranscriptionVoxtral Mini 3B
Voxtral Mini 3B is a 3-billion-parameter open-source speech-to-text model released by Mistral in July 2025.
Document ExtractionMistral OCR
Mistral OCR is a document extraction model from Mistral designed to recognize and extract text from images and documents.
TextMistral 8x7bDeprecated
Mistral 8x7b is a text generation model from Mistral AI using a mixture-of-experts architecture with a 32,768 token context window.
TextMixtral 8x7B InstructDeprecated
Mixtral 8x7B Instruct is a text generation model from Mistral using a sparse mixture-of-experts architecture with a 4096 token context window.
TextMistral Small 24.02Deprecated
Mistral Small 24.02 is a text generation model from Mistral with a 128,000-token context window, available via Amazon Bedrock.
TextMistral Large 24.07Deprecated
Mistral Large 24.07 is a text generation model from Mistral with a 128,000-token context window, available via Amazon Bedrock.
TextMistral Large 24.02Deprecated
Mistral Large 24.02 is a text generation model from Mistral with a 128,000-token context window, available via Amazon Bedrock.
TextMistral 7B InstructDeprecated
Mistral 7B Instruct is a text generation model from Mistral with a 4,096-token context window, available via Amazon Bedrock.
TextMixtral 8x22B InstructDeprecated
Mixtral 8x22B Instruct is a sparse mixture-of-experts text generation model from Mistral with a 64,000 token context window.
TextMistral 7B InstructDeprecated
Mistral 7B Instruct is a 7-billion-parameter text generation model from Mistral designed for instruction-following tasks.
TextMixtral 8x7B InstructDeprecated
Mixtral 8x7B Instruct is a sparse mixture-of-experts language model from Mistral, licensed under Apache 2.0.
Moonshot4
Kimi K3
Kimi K3 is a text generation model from Moonshot with a 1,048,576-token context window and up to 131,072-token responses.
TextKimi K2.7 Code
Kimi K2.7 Code is a text generation model from Moonshot AI built for coding and agentic tasks with a 262,164-token context window.
TextKimi K2.6
Kimi K2.6 is a text generation model from Moonshot AI with a 262,144-token context window, served via Deep Infra.
TextKimi K2.5
Kimi K2.5 is a text generation model from Moonshot AI with a 262,144-token context window and up to 16,384-token responses.
Nvidia5
Nemotron 3.5 Lightning
Nemotron 3.5 Lightning is a text generation model from Nvidia designed for fast reasoning with a 262,144 token context window.
TextNemotron 3 Ultra 550B
Nemotron 3 Ultra 550B is a text generation model from Nvidia with a 262,144-token context window and reasoning support.
TextNemotron 3 Nano 30B
Nemotron 3 Nano 30B is a text generation model from Nvidia with a 1,000,000-token context window and mixture-of-experts architecture.
TextNemotron 3 Super 120B
Nemotron 3 Super 120B is a text generation model from Nvidia with a 1,000,000-token context window and selectable reasoning.
RerankingLlama Nemotron Rerank VL 1B v2
Llama Nemotron Rerank VL 1B v2 is a vision-language reranking model from Nvidia with a 10,240-token context window.
Perplexity11
Sonar Deep Research
Sonar Deep Research is a text generation model from Perplexity that performs multi-step web research and returns cited answers.
TextSonar Reasoning Pro
Sonar Reasoning Pro is a text generation model from Perplexity that combines live web search with extended reasoning in a 128K context window.
TextSonar
Sonar is a text generation model from Perplexity that combines LLM chat with real-time web search and citation support.
TextSonar Pro
Sonar Pro is a text generation model from Perplexity that combines a 200,000-token context window with live web search.
EmbeddingPerplexity Embed v1 0.6B
Perplexity Embed v1 0.6B is a quantized embedding model from Perplexity with a 32,000-token context window.
EmbeddingPerplexity Embed v1 4B
Perplexity Embed v1 4B is a 4-billion-parameter embedding model from Perplexity with a 32,000-token context window.
TextSonar ReasoningDeprecated
Sonar Reasoning is a text generation model from Perplexity combining chain-of-thought reasoning with live internet search and citations.
TextSonar Large OnlineDeprecated
Sonar Large Online is a text generation model from Perplexity that retrieves live web data and returns cited answers.
TextSonar Large ChatDeprecated
Sonar Large Chat is a text generation model from Perplexity built on Llama 3.1 with a 127,072-token context window.
TextSonar Small OnlineDeprecated
Sonar Small Online is a web-connected text generation model from Perplexity built on Llama 3.1 with a 127,072-token context window.
TextSonar Small ChatDeprecated
Sonar Small Chat is a text generation model from Perplexity built on Llama 3.1 with a 127,072 token context window.
Qwen23
Qwen3.8 2.4T
Qwen3.8 2.4T is a mixture-of-experts text generation model from Qwen with a 262,144 token context window.
TextQwen3.5 Omni Flash
Qwen3.5 Omni Flash is a multimodal text generation model from Qwen with a 256,000 token context window.
TextQwen3.6 Flash
Qwen3.6 Flash is a text and image understanding model from Qwen with a 1,000,000-token context window and optional thinking mode.
TextQwen3.7 Max
Qwen3.7 Max is a text generation model from Qwen with a 1,000,000-token context window and optional deep thinking mode.
TextQwen3.7 Plus
Qwen3.7 Plus is a text generation model from Alibaba's Qwen team with a 1,000,000-token context window and optional deep thinking mode.
TextQwen3.6-35B-A3B
Qwen3.6-35B-A3B is a mixture-of-experts text generation model from Qwen with a 262,144 token context window.
TextQwen3.5 397B
Qwen3.5 397B is a multimodal text generation model from Qwen with a 262,144-token context window.
TextQwen3 235B
Qwen3 235B is a 235-billion-parameter mixture-of-experts text generation model from Qwen with a 262,144-token context window.
ImageQwen Image 3.0
Qwen Image 3.0 is an image generation model from Qwen that supports text prompts and source image inputs at $0.03 per image.
ImageQwen Image 3.0 Pro
Qwen Image 3.0 Pro is a flagship image generation model from Qwen supporting text-to-image and image-guided generation.
ImageQwen Image Edit Plus
Qwen Image Edit Plus is an image generation and editing model from Qwen, priced at $0.02 per image with ControlNet pose support.
ImageZ Image Turbo Controlnet
Z Image Turbo Controlnet is an image generation model from Qwen that uses ControlNet to guide outputs from a reference image.
ImageQwen 2 Pro
Qwen 2 Pro is an image generation model from Qwen supporting both text-to-image and image-to-image workflows at $0.07 per image.
ImageQwen Image
Qwen Image is an image generation model from Qwen that supports LoRA adapters and source image inputs for guided generation.
ImageZ Image Turbo
Z Image Turbo is an image generation model from Qwen that supports LoRA adapters and source image input at $0.005 per image.
TranscriptionQwen3 ASR 1.7B
Qwen3 ASR 1.7B is an open-source multilingual speech-to-text model developed by Qwen with 1.7 billion parameters.
Text to SpeechQwen3 TTS
Qwen3 TTS is an open-source multilingual text-to-speech model from Qwen with a 10,000-token context window.
EmbeddingQwen3 Embedding 0.6B
Qwen3 Embedding 0.6B is a compact embedding model from Qwen with a 32,768-token context window, served via DeepInfra.
EmbeddingQwen3 Embedding 4B
Qwen3 Embedding 4B is a text embedding model from Qwen with a 32,768-token context window, available via DeepInfra.
EmbeddingQwen3 Embedding 8B
Qwen3 Embedding 8B is an 8-billion-parameter text embedding model from Qwen with a 32,768-token context window.
RerankingQwen3 Reranker 0.6B
Qwen3 Reranker 0.6B is a lightweight reranking model from Qwen designed to reorder retrieved documents with a 32,768-token context window.
RerankingQwen3 Reranker 4B
Qwen3 Reranker 4B is a 4-billion-parameter reranking model from Qwen designed to score and reorder retrieved documents by relevance.
RerankingQwen3 Reranker 8B
Qwen3 Reranker 8B is an 8-billion-parameter reranking model from Qwen designed to score and reorder documents for retrieval pipelines.
Recraft4
Recraft V4.1
Recraft V4.1 is an image generation model from Recraft that produces images from text prompts at $0.04–$0.08 per image.
ImageRecraft V4.1 Pro
Recraft V4.1 Pro is a flagship image generation model from Recraft offering mode selection and multiple output sizes at $0.25–$0.30 per image.
ImageRecraft 20B
Recraft 20B is an image generation model from Recraft that produces images from text prompts at $0.022 per image.
ImageRecraft Crisp Upscale
Recraft Crisp Upscale is an image upscaling model from Recraft that sharpens and enlarges images at $0.004 per image.
Reka3
Reka EdgeDeprecated
Reka Edge is a 7B parameter text generation model from Reka with a 128,000 token context window.
TextReka FlashDeprecated
Reka Flash is a 21B parameter text generation model from Reka with a 128,000 token context window.
TextReka CoreDeprecated
Reka Core is a text generation model from Reka with a 128,000-token context window designed for complex tasks.
Sentence Transformers1
Stability5
SDXL LoRA
SDXL LoRA is an image generation model from Stability AI that supports custom LoRA adapters for fine-tuned visual styles.
ImageSDXL
SDXL is an open-source image generation model from Stability AI, available at $0.001 per image on MindStudio.
ImageStable Diffusion 3
Stable Diffusion 3 is a text-to-image generation model from Stability AI, released in June 2024.
ImageStable Image Ultra
Stable Image Ultra is a text-to-image generation model from Stability AI, supporting multiple aspect ratios and output formats.
ImageStable Image Core
Stable Image Core is a text-to-image generation model from Stability AI, supporting style presets, aspect ratios, and multiple output formats.
Voyage AI8
Voyage 4
Voyage 4 is an embedding model from Voyage AI designed for semantic search and retrieval with a 32,000-token context window.
EmbeddingVoyage 4 Large
Voyage 4 Large is an embedding model from Voyage AI with a 32,000 token context window for semantic search and retrieval.
EmbeddingVoyage 4 Lite
Voyage 4 Lite is a text embedding model from Voyage AI supporting a 32,000-token context window for semantic search and retrieval tasks.
EmbeddingVoyage Code 4
Voyage Code 4 is a code-specialized embedding model from Voyage AI with a 32,000 token context window.
EmbeddingVoyage Finance 2
Voyage Finance 2 is an embedding model from Voyage AI optimized for financial documents with a 32,000-token context window.
EmbeddingVoyage Law 2
Voyage Law 2 is an embedding model from Voyage AI optimized for legal text retrieval with a 32,000-token context window.
RerankingVoyage Rerank 2.5
Voyage Rerank 2.5 is a reranking model from Voyage AI designed to improve document retrieval relevance with a 32,000-token context window.
RerankingVoyage Rerank 2.5 Lite
Voyage Rerank 2.5 Lite is a reranking model from Voyage AI designed to reorder retrieved documents within a 32,000 token context window.
Wan13
Wan 2.7 Pro
Wan 2.7 Pro is an image generation model from Wan, available via Alibaba, priced at $0.075 per image.
ImageWan 2.7
Wan 2.7 is an image generation model from Alibaba's Wan that supports source image inputs and custom dimensions at $0.03 per image.
ImageWan 2.7 Pro
Wan 2.7 Pro is an image generation model from Wan that supports source image input and configurable output dimensions.
ImageWan 2.7
Wan 2.7 is an image generation model from Wan, available via Wavespeed, priced at $0.03 per image.
ImageWan 2.6
Wan 2.6 is an image generation model from Wan that accepts source images and text prompts to produce new images at $0.03 per image.
ImageWan 2.5
Wan 2.5 is an image generation model from Wan that accepts source images and produces new images at $0.03 per image.
VideoWan 3.0
Wan 3.0 is a video generation model that converts text prompts and images into video clips with optional audio output.
VideoWan 2.7
Wan 2.7 is a video generation model from Alibaba that supports text-to-video, image-to-video, video editing, and audio generation.
VideoWan 2.7
Wan 2.7 is a video generation model supporting text-to-video and image-to-video creation with audio and reference image inputs.
VideoWan 2.6
Wan 2.6 is a video generation model that creates videos from source images and audio with multi-shot control.
VideoWan 2.5
Wan 2.5 is a video generation model that animates source images with optional audio input, priced at $0.05–$0.15 per second.
VideoWan 2.2
Wan 2.2 is a video generation model from Wan that converts source images into videos with customizable LoRA fine-tuning.
ImageWan 2.2Deprecated
Wan 2.2 is an image generation model from Wan, available via Wavespeed, supporting LoRA fine-tuning and configurable resolution.
WaveSpeed8
Chroma
Chroma is an image generation model from WaveSpeed AI, available on MindStudio at $0.015 per image.
3DHunyuan3D V2 Multi-View
Hunyuan3D V2 Multi-View is a 3D generation model from WaveSpeed that converts multi-view images into textured 3D meshes.
3DHunyuan3D v3
Hunyuan3D v3 is a 3D generation model from WaveSpeed that converts images into 3D models with configurable polygon counts and PBR materials.
3DMeshy 6
Meshy 6 is a 3D generation model from WaveSpeed that converts text prompts and images into textured 3D assets.
3DSAM 3D Objects
SAM 3D Objects is a 3D generation model from WaveSpeed that converts images and masked inputs into 3D objects.
3DTripo3D v2.5
Tripo3D v2.5 is a 3D generation model from WaveSpeed that converts source images into 3D models at $0.30 per model.
Lip SyncInfiniteTalk
InfiniteTalk is a lip sync model by WaveSpeed that animates a portrait image to match a provided audio track.
Lip SyncLTX-2 19B Lipsync
LTX-2 19B Lipsync is a lip sync model from WaveSpeed that synchronizes video lip movements to source audio input.
X.ai22
Grok 4.6
Grok 4.6 is a text generation model from X.ai featuring a 500,000-token context window and configurable reasoning effort.
TextGrok 4.5
Grok 4.5 is a text generation model from X.ai featuring a 500,000-token context window and configurable reasoning effort.
TextGrok Build 0.1
Grok Build 0.1 is a text generation model from X.ai featuring a 256,000-token context window for extended conversations.
TextGrok 4.3
Grok 4.3 is a text generation model from X.ai featuring a 1,000,000-token context window and adjustable reasoning effort.
TextGrok 4.20 Reasoning
Grok 4.20 Reasoning is a text generation model from X.ai with a 2,000,000-token context window and built-in reasoning.
TextGrok 4.20
Grok 4.20 is a text generation model from X.ai featuring a 2,000,000 token context window for long-document tasks.
TextGrok 4.1 Fast Reasoning
Grok 4.1 Fast Reasoning is a text generation model from X.ai designed for rapid reasoning tasks with a 2,000,000 token context window.
TextGrok 4.1 Fast
Grok 4.1 Fast is a text generation model from X.ai offering a 2,000,000 token context window for large-scale language tasks.
TextGrok 4 Fast Reasoning
Grok 4 Fast Reasoning is a text generation model from X.ai designed for rapid reasoning tasks with a 2,000,000 token context window.
TextGrok 4 Fast
Grok 4 Fast is a text and image understanding model from X.ai with a 2,000,000-token context window.
TextGrok 4
Grok 4 is a text generation model from X.ai with a 256,000-token context window, released in July 2025.
TextGrok 3 Mini Fast
Grok 3 Mini Fast is a text generation model from X.ai designed for fast inference with a 131,072 token context window.
TextGrok 3 Mini
Grok 3 Mini is a text generation model from X.ai with a 131,072-token context window and configurable reasoning effort.
TextGrok 3 Fast
Grok 3 Fast is a text generation model from X.ai designed for speed, with a 131,072 token context window.
TextGrok 3
Grok 3 is a text generation model from X.ai with a 131,072-token context window, released in February 2025.
VisionGrok 4.3 Vision
Grok 4.3 Vision is a vision-capable model from X.ai featuring a 2,000,000-token context window for processing large inputs.
VisionGrok 2 Vision
Grok 2 Vision is a multimodal model from X.ai that processes images and text within a 32,768 token context window.
ImageGrok Imagine 2.0
Grok Imagine 2.0 is an image generation model from X.ai that accepts source images and supports configurable aspect ratio, resolution, and quality.
ImageGrok Imagine Pro
Grok Imagine Pro is an image generation model from X.ai that accepts a source image and aspect ratio as inputs.
ImageGrok Imagine
Grok Imagine is an image generation model from X.ai that accepts a source image and aspect ratio as inputs at $0.02 per image.
VideoGrok Imagine 1.5
Grok Imagine 1.5 is a video generation model from X.ai that accepts text, image, and video inputs to produce generated video output.
VideoGrok Imagine
Grok Imagine is a video generation model from X.ai that creates video from text prompts, source images, or source video clips.
Z.ai7
GLM-5.2
GLM-5.2 is a text generation model from Z.ai with a 1,048,576-token context window and up to 128,000 token responses.
TextGLM 5.1
GLM 5.1 is a text generation model from Z.ai with a 200,000-token context window and configurable reasoning effort.
TextGLM 5
GLM 5 is a text generation model from Z.ai with a 200,000-token context window and configurable reasoning effort.
TextGLM 4.7
GLM 4.7 is a text generation model from Z.ai with a 131072-token context window and configurable reasoning effort.
TextGLM 4.6V
GLM 4.6V is a multimodal text generation model from Z.ai with a 131,072 token context window and configurable reasoning effort.
TextGLM 4.6
GLM 4.6 is a text generation model from Z.ai with a 200,000-token context window and configurable reasoning effort.
TextGLM 4.7 Flash
GLM 4.7 Flash is a text generation model from Z.ai with a 202,752-token context window and adjustable reasoning effort.
Build with any AI model
No API keys required. Start building AI-powered workflows in minutes.
Get Started Free