Bienvenue dans votre centre d'outils IA. Conçue pour une navigation fluide et un déploiement rapide, cette page réunit l'ensemble de notre intelligence — des puissants grands modèles de langage (LLM) aux générateurs d'images et de vidéos les plus avancés. Évaluez et accédez en un coup d'œil au moteur idéal pour votre flux de travail.

MiniMax H3 Max text-to-video: generate a cinematic video from a text prompt. Supports 480P、768P, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.

MiniMax H3 Max image-to-video: animate a first-frame image (optionally with a last frame) driven by a text prompt. Supports 480P、768P, 5-15s.

MiniMax H3-Developer self-hosted text-to-video: generate a video (with audio) from a text prompt. Supports 768P/1080P, 16:9/9:16/1:1 aspect ratios, tunable seed and inference steps.

MiniMax H3-Developer self-hosted image-to-video: animate a first-frame image (optionally a last frame) driven by a text prompt, with generated audio. Supports 768P/1080P.

MiniMax H3-Developer self-hosted reference-to-video: generate a video that keeps the subject from one or more reference images/videos, driven by a text prompt, with generated audio. Supports 768P/1080P.

ByteDance Seedream 4.7 image editing model with batch generation support. Produce a coherent set of edited images from reference inputs.

ByteDance Seedream 4.7 image editing model. Executes edit instructions precisely while preserving identity, lighting and local structure of the source image.

ByteDance Seedream 4.7 with batch generation support. Generate a set of coherent images in a single request.

ByteDance Seedream 4.7 image generation model. Balanced gains in image quality, aesthetics and instruction following, at the efficiency and cost profile of the 4.0 generation.

All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.

Animate a first frame (optionally with a last frame) into a coherent clip, with native audio and smart-duration up to 30s.

All-in-One Reference: keep subjects consistent from any mix of reference images, videos, and audio; pixel-level identity/voice/space alignment.

All-in-one Wan3.0 renderer: cinematic, hyper-real video from a text prompt, up to 30s with smart-duration and adaptive aspect ratio.

Animate a first frame (optionally with a last frame) into a coherent clip, with native audio and smart-duration up to 30s.

All-in-One Reference: keep subjects consistent from any mix of reference images, videos, and audio; pixel-level identity/voice/space alignment.
DeepSeek V4 Pro is a state-of-the-art large language model combining efficient sparse attention, strong reasoning, and integrated agent capabilities for robust long-context understanding and versatile AI applications.
xAI’s flagship model for advanced reasoning, coding, and agentic tasks.

Multimodal video generation from reference images, videos, and audio. Supports video editing and extension.

Generate videos from a first-frame image (and optional last-frame) with native audio.

Generate videos from text prompts with native audio and optional web search.

MiniMax Music 3.0 is MiniMax's 11.1B-parameter open-weights music model that turns a musical description and optional lyrics into a complete, fully arranged and mixed song of up to five minutes - vocals, instrumentation and production included - in a single generation, with section-tag control over the arrangement and vocal or instrumental output.

MiniMax Lyrics Generation is a dedicated lyric-writing model that turns a one-line theme into a complete, professionally structured set of song lyrics - title, style tags, and sections marked with [Verse]/[Chorus] structure tags - and can also edit, continue, or restructure existing lyrics, with output directly usable as the lyrics input of MiniMax's music models.

xAI Grok Imagine Image 2.0 generates polished visuals from natural-language prompts at 1K or 2K resolution, with 14 aspect ratios and selectable low/medium quality tiers.

xAI Grok Imagine Image 2.0 edits up to three reference images with natural-language instructions at 1K or 2K resolution, with selectable low/medium quality tiers.
Experimental multimodal model optimized for fast visual understanding and reasoning.

Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen strength in complex text rendering and precise prompt adherence

Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes

ByteDance flagship image layer decomposition. Splits a single input image into an editable stack: one base image plus up to 16 transparent PNG layers, each returned with stacking order (z_index), bounding box coordinates, name, and description for downstream drag/scale/recompose editing.
DeepSeek V4 Flash is a state-of-the-art large language model combining efficient sparse attention, strong reasoning, and integrated agent capabilities for robust long-context understanding and versatile AI applications.

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen strength in complex text rendering and precise prompt adherence

Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.

Suno text-to-music via APIMart: inspiration mode (custom=false) turns a description into a song; custom mode (custom=true) uses your own lyrics. Async; returns 2 tracks per generation.
Next-generation flagship model for advanced reasoning, coding, and multimodal AI applications.

MiniMax H3 text-to-video: generate a cinematic video from a text prompt. Supports 2K, 5-15s., and 16:9/9:16/1:1/adaptive aspect ratios.

MiniMax H3 image-to-video: animate a first-frame image (optionally with a last frame) driven by a text prompt. Supports 2K, 5-15s.

MiniMax H3 reference-to-video: generate a video that keeps the subject from a reference image, driven by a text prompt. Supports 2K, 5-15s.
Top-performing open-weight model optimized for frontier reasoning, coding, and enterprise AI applications.
Flagship conversational model built for real-time knowledge exploration, sharp reasoning, and highly engaging AI interactions.

Youchuan V8.2 animates an input image into four 5-second videos at 480p or 720p.

Youchuan automatically removes the background from an input image, returning one transparent-background result.

Youchuan retexture changes the artistic style of an input image while preserving its composition, returning four restyled results.

Youchuan V8.2 blends two to five input images into four fused results, with an optional guiding prompt and native 2K HD.