文生视频、首帧图生视频、首尾帧控制、多模态参考。Text, first-frame, first-and-last-frame and multi-reference.
一个模型覆盖文生视频、首帧图生视频、首尾帧控制和多模态参考,支持原生音频、1080p 与最长 30 秒输出 One model for text, first-frame, first-and-last-frame and multi-reference video with native audio, 1080p and up to 30 seconds
当前 Minuit 视频创作的首选入口。四种生成方式使用同一个公开模型名,便于产品接入和后续扩展。The primary video model on Minuit. Four generation modes share one public model ID for simpler integration and expansion.
从简单提示词到包含图片、视频和音频的参考素材,Wan 3.0 用统一模型完成生成。它适合广告短片、角色内容、产品演示、社交媒体和连续叙事等需要稳定画面与声音的场景。From a simple prompt to image, video and audio references, Wan 3.0 handles generation through one unified model. It fits ads, character content, product stories, social media and narrative work that needs consistent visuals and sound.
文生视频、首帧图生视频、首尾帧控制、多模态参考。Text, first-frame, first-and-last-frame and multi-reference.
比短片段更适合完整动作、镜头调度和连续叙事。More room for complete motion, camera direction and continuous narrative.
可生成对白、环境音与配乐,也可关闭声音获得静音输出。Generate dialogue, ambience and music, or disable audio for silent output.
支持 480p、720p、1080p,并可让模型自动规划画面比例。Supports 480p, 720p and 1080p with adaptive aspect-ratio planning.
以下内容和案例保留用于对比与工作流参考;新项目建议优先选择 Wan 3.0。The material below remains for comparison and workflow reference. Use Wan 3.0 first for new projects.
| 时长Duration | 2秒 - 15秒(灵活)2s - 15s (flexible) |
| 音频Audio | 有声(原生环境音/人声)With audio (native ambient/voice) |
| 分辨率Resolution | 720p / 1080p |
| 成本Cost | $0.13 - $0.20 /sec |
| 输入Input | 图生视频 (I2V)Image-to-Video (I2V) |
| 首尾帧控制First/Last Frame | 暂不支持Not supported yet |
| 时长Duration | 5秒 - 8秒(固定)5s - 8s (fixed) |
| 音频Audio | 无声(需后期配音/BGM)Silent (requires post dubbing/BGM) |
| 分辨率Resolution | 480p / 720p |
| 成本Cost | $0.03 - $0.06 /sec |
| 输入Input | 图生视频 (I2V)Image-to-Video (I2V) |
| 首尾帧控制First/Last Frame | 即将上线Coming Soon |
影视级画质视频内容,直接驱动精品订阅或单次购买。Cinema-grade video content directly driving premium subscriptions or single purchases.
多分支互动故事体验,用户选择决定视频走向。Multi-branch interactive story experience where user choices determine video direction.
为内容创作者提供一键专业视频生成工具。One-click professional video generation tools for content creators.
创作者生成并出售精品 AI 视频内容的交易平台。Marketplace where creators generate and sell premium AI video content.
游戏角色对话场景的微动画素材:呼吸、眨眼、微妙动作。Micro-animation assets for character dialogue: breathing, blinking, subtle movements.
AI 聊天界面中的动态头像,在长期对话中提供视觉反馈。Dynamic avatar in AI chat interface providing visual feedback during conversations.
快节奏短视频,利用 TikTok/Reels/Shorts 上的视觉冲击力进行流量获取。Fast-paced short videos leveraging visual impact on TikTok/Reels/Shorts for traffic acquisition.
用户上传 LoRA 模型,即时将静态照片转换为微动画动态内容。Users upload LoRA model to instantly convert static photos to micro-animation content.
为 Spicy 平台内容卡片生成动态封面缩略图,替代静态图片提升点击率。Generate dynamic cover thumbnails for content cards, replacing static images to boost CTR.
为 VTuber 直播生成动态背景、互动特效和粉丝触发动画。Generate dynamic backgrounds, interactive effects and fan-triggered animations for VTuber streams.
截至 2026-06-05,全部自研能力As of 2026-06-05, all proprietary capabilities
| # | 类别Category | 模型名称Model | 规格Spec | 单位Unit | 价格Price (USD) | 备注Notes |
|---|---|---|---|---|---|---|
| 1 | 图片生成Image Gen | zimage_spicy | Spicy | 张img | $0.013 | |
| 2 | 图片编辑Image Edit | qwen_image_edit_spicy | Spicy | 张img | $0.040 | |
| 3 | 图片换脸Face Swap | faces_swap | Spicy | 张img | $0.013 | |
| 4 | 头部替换Head Swap | head_swap | Spicy | 张img | $0.013 | 脸+头型+发型Face+head+hair |
| 5 | 图生视频I2V | wan22spicy | 480p | 秒sec | $0.040 | 5秒 / 8秒 两档5s / 8s tiers |
| 6 | 图生视频I2V | wan22spicy | 720p | 秒sec | $0.080 | 5秒 / 8秒 两档5s / 8s tiers |
| 7 | 图生视频I2V | wan27spicy | 720p | 秒sec | $0.130 | 2-15秒 / 音频 / Prompt优化2-15s / Audio / Prompt opt |
| 8 | 图生视频I2V | wan27spicy | 1080p | 秒sec | $0.200 | 2-15秒 / 音频 / Prompt优化2-15s / Audio / Prompt opt |
| 9 | 数字人Digital Human | wan_animate | 480p | 秒sec | $0.053 | 5秒 / 8秒 两档5s / 8s tiers |
| 10 | 数字人Digital Human | wan_animate | 720p | 秒sec | $0.110 | 5秒 / 8秒 两档5s / 8s tiers |
| 11 | 视频超分Super-Res | flashvsr | → 2K | 秒sec | $0.032 | |
| 12 | 视频超分Super-Res | flashvsr | → 4K | 秒sec | $0.053 | |
| 13 | Prompt 优化Prompt Opt | zimage_spicy_prompt | 图片生成优化Image gen opt | 次req | $0.001 | |
| 14 | Prompt 优化Prompt Opt | wan22_spicy_prompt | 视频生成优化Video gen opt | 次req | $0.001 |
覆盖从低成本 UGC 引流到高沉浸付费内容的完整转化链路。A complete conversion path from low-cost UGC acquisition to premium immersive content.
查看原始 9 页方案,包含客户画像、成本模型、集成架构与商业化建议。Read the original 9-page brief with customer profiles, cost models, integration architecture and monetization guidance.
咨询 API 与批量价格 →Discuss API & volume pricing → 打开 PDF 方案Open PDF Brief为角色对话生成呼吸、眨眼、轻微晃动等微动态资产。Micro-motion character assets for dialogue, including breathing, blinking and subtle movement.
动态头像与长期对话视觉反馈,支持预生成情感状态视频库。Dynamic avatars and visual feedback for long-running conversations, backed by pre-generated emotional states.
面向 TikTok、Reels、Shorts 的快节奏视觉内容。Fast-paced visual content for TikTok, Reels and Shorts.
上传首图,将静态写真转化为带“呼吸感”的动态内容。Turn a single portrait into a naturally animated digital double.
为内容卡片自动生成 5 秒动态封面,替代静态缩略图。Automatically generate 5-second animated covers for content cards.
循环背景、互动特效与粉丝打赏触发动画。Looping backgrounds, interactive effects and fan-gift-triggered animations.
1080p、原生环境音与 2–15 秒完整动作,适合高价值付费内容。1080p, native ambience and complete 2–15s actions for high-value paid content.
原生衣物摩擦、水声、耳语与画面同步,直接形成沉浸式成片。Synchronized native fabric, water and whisper audio for immersive ready-to-use output.
精心设计首帧与 FlashVSR 高清增强,适合角色营销展示。Carefully designed first frames with FlashVSR enhancement for premium character showcases.
为角色 LoRA 市场生成高清、有声动态预览。Generate high-definition, native-audio previews for character LoRA marketplaces.
用 10–15 秒片段呈现悬念、高潮、反转,并串联为迷你剧。Use 10–15s clips to build suspense, climax and reversals into a mini-series.
通过原生人声和高清画面模拟视频通话式回复。Simulate video-call replies with native voice and high-definition visuals.
优先级:Priority: 生成延迟 > 角色一致性 > 成本 > 画质Latency > consistency > cost > quality
方案:Stack: Wan 2.2 + 换脸 APIFace Swap API + TTS
预生成状态视频缓存 → 实时情感匹配 → 换脸一致性 → TTS 叠加推流。Pre-generated state cache → real-time emotion matching → face consistency → TTS compositing.
优先级:Priority: 总成本 > 批量一致性 > Loop 友好性 > 画质Total cost > consistency > loopability > quality
方案:Stack: Wan 2.2 + 文生图 / 图片编辑image generation / editing
批量生产动态立绘,并与本地语音包精确同步。Mass-produce animated portraits with precise local voice-pack synchronization.
优先级:Priority: 画质 > 音频沉浸感 > 时长灵活性 > 成本Quality > audio immersion > duration flexibility > cost
方案:Stack: Wan 2.7 + FlashVSR
1080p 原生声音可直接形成高价单品,兼顾更新频率与高端形象。Native-audio 1080p output supports premium sales while maintaining cadence and brand quality.
优先级:Priority: 产出效率 > 角色一致性 > 成本 > 画质上限Throughput > consistency > cost > peak quality
方案:Stack: Wan 2.2 + Wan 2.7 + 换脸 APIFace Swap API
2.2 日常量产,2.7 月度精品,换脸保障跨视频 IP 一致性。2.2 for daily volume, 2.7 for monthly hero content, and face swap for cross-video IP consistency.
通过 Minuit API 官网联系我们,获取适合您业务的模型组合。Contact us through the Minuit API website for a solution tailored to your business.
🔊 v2.4 音频全面升级 · 真人视频带声音 · 内置Prompt优化 · 覆盖动漫、亚洲、欧美风格 · 左右滑动浏览🔊 v2.4 Audio Fully Upgraded · Real person video with audio · Built-in prompt optimization · Swipe to browse by style
从图生视频、文字生图到图片编辑、换脸、数字人、视频超分,一站式多模态API能力Image-to-video, text-to-image, image editing, face swap, digital human, super-resolution
配合能力:Prompt优化With capability: Prompt Optimization
配合能力:Prompt优化With capability: Prompt Optimization
Wan2.2 与 2.7 对比分析、Prompt 建议、首图建议Wan2.2 vs 2.7 comparison, prompt tips, first-image guidelines
覆盖视频生成 → 图片生成 → 图片编辑 → 换脸换头全链路,掌握每个环节的 Prompt 技巧Covering Video Gen → Image Gen → Image Edit → Face/Head Swap — master prompt techniques for every stage
图生视频 (I2V) · 5-15 秒 · 带音频 · 逐帧精细控制Image-to-Video (I2V) · 5-15 sec · With Audio · Frame-by-Frame Control
文生视频 (T2V) · 5 秒 · 无声音 · 低成本批量生产Text-to-Video (T2V) · 5 sec · Silent · Low-Cost Batch Production
文生图 · 首帧图制作 · 构图可控Text-to-Image · First Frame Creation · Controllable Composition
精准修改 · 逐步迭代 · 权重调整Precise Modifications · Iterative Refinement · Weight Adjustment
让替换自然无痕 · 角度匹配第一优先Seamless Replacement · Angle Matching Is #1 Priority
点击展开查看完整 Prompt,按能力分类Click to expand, organized by capability
通过多个API能力串联,可以做出更好的视频Chain multiple API capabilities together for superior video output
通过文字描述生成基础图片,作为后续流程的起点Generate a base image from text description as the starting point for the pipeline
对生成的图片进行 Spicy 风格编辑,调整内容方向Apply spicy-style editing to the generated image, adjusting content direction
进一步编辑图片,提升尺度和真实感,使首帧图达到最佳效果Further edit the image to enhance intensity and realism, optimizing the first frame
将图片中的人脸替换为目标人脸,实现角色定制Replace the face in the image with a target face for character customization
替换整个头部(脸型+头型+发型),比 Face Swap 覆盖范围更大,身份一致性更强Replace the entire head (face+head shape+hairstyle), broader coverage than Face Swap for stronger identity consistency
使用 Wan2.7 Spicy 生成带声音的高质量真人视频,效果最佳Use Wan2.7 Spicy for the highest quality real person video with audio
使用 Wan2.2 Spicy 生成无声音视频,成本更低,适合批量生产Use Wan2.2 Spicy for silent video at lower cost, suitable for batch production
将生成的视频超分至 2K 或 4K 分辨率,大幅提升视觉质量,做出更好的视频Upscale generated video to 2K or 4K resolution for significantly enhanced visual quality
产品能力变更与文档更新记录Product capability changes and documentation updates