CRITICAL: When calling the script, you MUST use the exact modelid (second column), NOT the friendly model name. Do NOT infer modelid from the friendly name (e.g., ❌ nano-banana-pro is WRONG; ✅ gemini-3-pro-image is CORRECT).
Quick Reference Table:
| 友好名称 (Friendly Name) | model_id | 说明 (Notes) |
|---|---|---|
| Nano Banana2 | INLINECODE2 | ❌ NOT nano-banana-2, 预算选择 4-13 pts |
| Nano Banana Pro |
gemini-3-pro-image | ❌ NOT nano-banana-pro, 高质量 10-18 pts |
| SeeDream 4.5 | doubao-seedream-4.5 | ✅ Recommended default, 5 pts |
| Midjourney | midjourney | ✅ Same as friendly name, 8-10 pts |
| 友好名称 (Friendly Name) | modelid (t2v) | modelid (i2v) | 说明 (Notes) |
|---|---|---|---|
| Wan 2.6 | INLINECODE6 | INLINECODE7 | ⚠️ Note -t2v/-i2v suffix |
| IMA Video Pro (Sevio 1.0) |
ima-pro | ima-pro | ✅ IMA native quality model |
| IMA Video Pro Fast (Sevio 1.0-Fast) | ima-pro-fast | ima-pro-fast | ✅ IMA native low-latency model |
| Kling O1 | kling-video-o1 | kling-video-o1 | ⚠️ Note video- prefix |
| Kling 2.6 | kling-v2-6 | kling-v2-6 | ⚠️ Note v prefix |
| Hailuo 2.3 | MiniMax-Hailuo-2.3 | MiniMax-Hailuo-2.3 | ⚠️ Note MiniMax- prefix |
| Hailuo 2.0 | MiniMax-Hailuo-02 | MiniMax-Hailuo-02 | ⚠️ Note 02 not 2.0 |
| Google Veo 3.1 | veo-3.1-generate-preview | veo-3.1-generate-preview | ⚠️ Note -generate-preview suffix |
| Sora 2 Pro | sora-2-pro | sora-2-pro | ✅ Straightforward |
| Pixverse | pixverse | pixverse | ✅ Same as friendly name |
| 友好名称 (Friendly Name) | model_id | 说明 (Notes) |
|---|---|---|
| Suno (sonic v4) | INLINECODE26 | ⚠️ Simplified to sonic |
| DouBao BGM |
GenBGM | ❌ NOT doubao-bgm |
| DouBao Song | GenSong | ❌ NOT doubao-song |
| 友好名称 (Friendly Name) | model_id | 说明 (Notes) |
|---|---|---|
| seed-tts-2.0 | INLINECODE29 | ✅ Same as friendly name (default) |
How to get the correct model_id:
--list-models --task-type <type> to query available modelsRuntime truth source:
GET /open/v1/product/list(or--list-models).
Any table in this document is guidance; actual availability depends on current product list.
Example:
# ❌ WRONG: Inferring from friendly name
--model-id nano-banana-pro
# ✅ CORRECT: Using exact model_id from table
--model-id gemini-3-pro-image
This skill is fully runnable as a standalone package.
If ima-knowledge-ai is installed, the agent may read its references for workflow decomposition and consistency guidance.
Recommended optional reads:
ima-knowledge-ai/references/workflow-design.md if:ima-knowledge-ai/references/visual-consistency.md if:ima-knowledge-ai/references/video-modes.md if:ima-knowledge-ai/references/model-selection.md if:Why this matters:
Example multi-media workflow:
CODEBLOCK1
How to check:
CODEBLOCK2
No exceptions — for simple single-media requests, you can proceed directly. For complex multi-media workflows, read the knowledge base first.
Purpose: So that any agent parses user intent consistently, first determine the media type from the user's request, then choose task_type and model.
| User intent / keywords | Media type | task_type examples |
|---|---|---|
| 画 / 生成图 / 图片 / image / 画一张 / 图生图 | image | INLINECODE38 , INLINECODE39 |
| 视频 / 生成视频 / video / 图生视频 / 文生视频 |
text_to_video, image_to_video, first_last_frame_to_video, reference_image_to_video |
| 音乐 / 歌 / BGM / 背景音乐 / music / 作曲 | music | text_to_music |
| 语音 / 朗读 / TTS / 语音合成 / 配音 / speech / read aloud / text-to-speech | speech | text_to_speech |
If the request mixes media (e.g. "宣传片+配乐"), treat as multi-media workflow: read workflow-design.md, then plan image → video → music steps and use the correct task_type for each step.
ima-all-ai:
- Ima Sevio 1.0 → ima-pro
- Ima Sevio 1.0-Fast / Ima Sevio 1.0 Fast → ima-pro-fast
Routing rule:
- Normalize alias first
- Then resolve against runtime product list for the selected task_type
- If model is absent in current category, return available model_ids from --list-models
sonic) vs DouBao BGM/Song — infer from "BGM"/"背景音乐" → BGM; "带歌词"/"人声" → Suno or Song. Use modelid sonic, GenBGM, GenSong per "Recommended Defaults" and "Music Generation" tables below.GET /open/v1/product/list?category=text_to_speech or run script with --task-type text_to_speech --list-models. Map user intent to parameters using product form_config:| User intent / phrasing | Parameter (if in form_config) | Notes |
|------------------------|--------------------------------|--------|
| 女声 / 女声朗读 / female voice | voiceid / voicetype | Use value from form_config options |
| 男声 / 男声朗读 / male voice | voiceid / voicetype | Use value from form_config options |
| 语速快/慢 / speed up/slow | speed | e.g. 0.8–1.2 |
| 音调 / pitch | pitch | If supported |
| 大声/小声 / volume | volume | If supported |
If the user does not specify, use formconfig defaults. Pass extra params via --extra-params '{"speed":1.0}'. Only send parameters present in the product’s creditrules/attributes or form_config (script reflection strips others on retry).
For transparency: This skill uses a bundled Python script (scripts/ima_create.py) to call the IMA Open API. The script:
--user-id only locally as a key for storing your model preferencesWhat gets sent to IMA servers:
What's stored locally:
~/.openclaw/memory/ima_prefs.json - Your model preferences (< 1 KB)| Domain | Owner | Purpose | Data Sent | Privacy |
|---|---|---|---|---|
| INLINECODE67 | IMA Studio | Main API (product list, task creation, task polling) | Prompts, model IDs, generation params, your API key | Standard HTTPS, data processed for AI generation |
| INLINECODE68 |
*.aliyuncs.com, *.esxscloud.com | Alibaba Cloud (OSS) | Image/video storage (file upload, CDN delivery) | Raw image/video bytes (via presigned URL, NO API key) | IMA-managed OSS buckets, presigned URLs expire after 7 days |
Key Points:
text_to_music) and TTS tasks (text_to_speech) only use api.imastudio.com.imapi.liveme.com to obtain presigned URLs for uploading input images.api.imastudio.com and imapi.liveme.com (both owned by IMA Studio).tcpdump -i any -n 'host api.imastudio.com or host imapi.liveme.com'. See this document: 🌐 Network Endpoints Used and ⚠️ Credential Security Notice for full disclosure.Your API key is sent to both IMA-owned domains:
Authorization: Bearer ima_xxx... → api.imastudio.com (main API)appUid=ima_xxx... → imapi.liveme.com (upload service)Security best practices:
https://imastudio.com/dashboard for unauthorized activity.~/.openclaw/logs/ima_skills/ for unexpected API calls.Why two domains? IMA Studio uses a microservices architecture:
api.imastudio.com: Core AI generation APIimapi.liveme.com: Specialized image/video upload service (shared infrastructure)Both domains are operated by IMA Studio. The same API key grants access to both services.
Note for users: You can review the script source at
scripts/ima_create.pyanytime.
The agent uses this script to simplify API calls. Music tasks use onlyapi.imastudio.com, while image/video tasks also callimapi.liveme.comfor file uploads (see "Network Endpoints" above).
Use the bundled script internally for all task types — it ensures correct parameter construction:
CODEBLOCK3
The script outputs JSON with url, model_name, credit — use these values in the UX protocol messages below. The script internals (product list query, parameter construction, polling) are invisible to users.
Call IMA Open API to create AI-generated content. All endpoints require an ima_* API key. The core flow is: query products → create task → poll until done.
This skill is community-maintained and open for inspection.
Full transparency:
scripts/ima_create.py and ima_logger.py anytimeapi.imastudio.com only; image/video tasks also use imapi.liveme.com (see "Network Endpoints" section)~/.openclaw/memory/ima_prefs.json and log filesConfiguration allowed:
export IMA_API_KEY=ima_your_key_hereIMA_API_KEY to agent's environment configurationData control:
rm ~/.openclaw/memory/ima_prefs.json (resets to defaults)rm -rf ~/.openclaw/logs/ima_skills/ (auto-cleanup after 7 days anyway)If you need to modify this skill for your use case:
Note: Modified skills may break API compatibility or introduce security issues. Official support only covers the unmodified version.
Actions that could compromise security:
Why this matters:
What this skill does with your data:
| Data Type | Sent to IMA? | Stored Locally? | User Control |
|---|---|---|---|
| Prompts (image/video/music) | ✅ Yes (required for generation) | ❌ No | None (required) |
| API key |
--user-id value |Privacy recommendations:
--user-id is never sent to IMA servers - it's only used locally as a key for storing preferences in INLINECODE106scripts/ima_create.py to verify network calls (search for create_task function)Get your IMA API key: Visit https://imastudio.com to register and get started.
Version control:
File checksums (optional):
CODEBLOCK4
If users report issues, verify file integrity first.
User preferences have highest priority when they exist. But preferences are only saved when users explicitly express model preferences — not from automatic model selection.
~/.openclaw/memory/ima_prefs.jsonSingle file, shared across all IMA skills:
CODEBLOCK5
Step 1: Get knowledge-ai recommendation (if installed)
CODEBLOCK6
Step 2: Check user preference
CODEBLOCK7
Step 3: Decide which model to use
CODEBLOCK8
Step 4: Check for mismatch (for later hint)
CODEBLOCK9
✅ Save preference when user explicitly specifies a model:
| User says | Action |
|---|---|
INLINECODE110 / 换成XXX / INLINECODE112 | Switch to model XXX + save as preference |
INLINECODE113 / 默认用XXX / INLINECODE115 |
✅ 已记住!以后图片生成默认用 [XXX] |我喜欢XXX / 我更喜欢XXX | Save as preference |
❌ Do NOT save when:
🗑️ Clear preference when user wants automatic selection:
| User says | Action |
|---|---|
INLINECODE119 / 用最合适的 / best / INLINECODE122 | Clear pref + use knowledge-ai recommendation |
INLINECODE123 / 你选一个 / INLINECODE125 |
用默认的 / 用新的 | Clear pref + use knowledge-ai recommendation |试试别的 / 换个试试 (without specific model) | Clear pref + use knowledge-ai recommendation |重新推荐 | Clear pref + use knowledge-ai recommendation |
Implementation:
del prefs[f"user_{user_id}"][task_type]
save_prefs(prefs)
Selection flow:
Important notes:
The defaults below are FALLBACK only. User preferences have highest priority, then knowledge-ai recommendations.
When using user preference for image generation, show a line like:
CODEBLOCK11
When user switches to a different model than their saved preference:
CODEBLOCK12
These are fallback defaults — only used when no user preference exists.
Always default to the newest and most popular model. Do NOT default to the cheapest.
| Task Type | Default Model | modelid | versionid | Cost | Why |
|---|---|---|---|---|---|
| texttoimage | SeeDream 4.5 | INLINECODE131 | INLINECODE132 | 5 pts | Latest doubao flagship, photorealistic 4K |
| texttoimage (budget) |
gemini-3.1-flash-image | gemini-3.1-flash-image | 4 pts | Fastest and cheapest option |
| texttoimage (premium) | Nano Banana Pro | gemini-3-pro-image | gemini-3-pro-image-preview | 10/10/18 pts | Premium quality, 1K/2K/4K options |
| texttoimage (artistic) | Midjourney 🎨 | midjourney | v6 | 8/10 pts | Artist-level aesthetics, creative styles |
| imagetoimage | SeeDream 4.5 | doubao-seedream-4.5 | doubao-seedream-4-5-251128 | 5 pts | Latest, best i2i quality |
| imagetoimage (budget) | Nano Banana2 | gemini-3.1-flash-image | gemini-3.1-flash-image | 4 pts | Cheapest option |
| imagetoimage (premium) | Nano Banana Pro | gemini-3-pro-image | gemini-3-pro-image-preview | 10 pts | Premium quality |
| imagetoimage (artistic) | Midjourney 🎨 | midjourney | v6 | 8/10 pts | Artist-level aesthetics, style transfer |
| texttovideo | Wan 2.6 | wan2.6-t2v | wan2.6-t2v | 25 pts | 🔥 Most popular t2v, balanced cost |
| texttovideo (premium) | Hailuo 2.3 | MiniMax-Hailuo-2.3 | MiniMax-Hailuo-2.3 | 38 pts | Higher quality |
| texttovideo (budget) | Vidu Q2 | viduq2 | viduq2 | 5 pts | Lowest cost t2v |
| imagetovideo | Wan 2.6 | wan2.6-i2v | wan2.6-i2v | 25 pts | 🔥 Most popular i2v, 1080P |
| imagetovideo (premium) | Kling 2.6 | kling-v2-6 | kling-v2-6 | 40-160 pts | Premium Kling i2v |
| firstlastframetovideo | Kling O1 | kling-video-o1 | kling-video-o1 | 48 pts | Newest Kling reasoning model |
| referenceimageto_video | Kling O1 | kling-video-o1 | kling-video-o1 | 48 pts | Best reference fidelity |
| texttomusic | Suno (sonic-v4) | sonic | sonic | 25 pts | Latest Suno engine, best quality |
| texttospeech | (query product list) | — | — | — | Run --task-type text_to_speech --list-models; use first or user-preferred model_id |
Premium options:
Quick selection guide (production as of 2026-02-27, sorted by popularity):
Selection guide by use case:
Image Generation:
Video Generation:
Music Generation:
Speech (TTS) Generation:
text_to_speech. Always query GET /open/v1/product/list?category=text_to_speech (or --list-models) to get current modelid and credit. No fixed default; use first available or user preference. Voice/speed/format parameters: see "Model and parameter parsing" (TTS table) and "Speech (TTS) — textto_speech" in this document.⚠️ Technical Note for Suno:
INLINECODE167 inside
parameters.parameters(e.g.,"sonic-v5") is different from the outermodel_versionfield (which is"sonic"). Always set both correctly when creating Suno tasks.
⚠️ Production Image Models (4 available):
doubao-seedream-4.5) — 5 pts, defaultmidjourney) — 8/10 pts for 480p/720p, artistic stylesgemini-3.1-flash-image) — 4/6/10/13 pts for 512px/1K/2K/4Kgemini-3-pro-image) — 10/10/18 pts for 1K/2K/4KAll other image models mentioned in older documentation are no longer available in production.
🌟 Parameter Support Notes (All Task Types):
🆕 MAJOR UPDATE: Nano Banana series now has NATIVE aspect_ratio support!
aspect_ratio (1:1, 16:9, 9:16, 4:3, 3:4) NATIVELYaspect_ratio (1:1, 16:9, 9:16, 4:3, 3:4) NATIVELYaspect_ratio support details:
attribute_ids, 4-13 pts)attribute_ids, 10-18 pts)attribute_id, 8/10 pts)When user requests unsupported combinations for images:
❌ Midjourney 暂不支持自定义 aspect_ratio(仅支持 1024x1024 方形)
✅ 推荐方案:
1. SeeDream 4.5(支持虚拟参数 aspect_ratio)
• 支持比例:1:1, 16:9, 9:16, 4:3, 3:4, 2:3, 3:2, 21:9
• 成本:5 积分(性价比最佳)
2. Nano Banana Pro/2(原生支持 aspect_ratio)
• 支持比例:1:1, 16:9, 9:16, 4:3, 3:4
• 成本:4-18 积分(按尺寸)
需要我帮你用 SeeDream 4.5 生成吗?
form_config)form_config)Auto-Inference Logic for Pixverse V5.5/V5/V4:
model field in form_config from Product List APImodel parameter (e.g., "v5.5", "v5", "v4")model_name and injects itmodel_name: "Pixverse V5.5" → auto-inject model: "v5.5"model_name: "Pixverse V4" → auto-inject model: "v4"model in form_config (no auto-inference needed)Error Prevention:
Suno sonic-v5 (Full-Featured):
DouBao BGM/Song (Simplified):
🎵 Suno Prompt Writing Guide (for gpt_description_prompt):
When using Suno, structure your prompt with these elements:
"lo-fi hip hop", "orchestral cinematic", "upbeat pop", "dark ambient", "indie folk", INLINECODE203
"80 BPM", "fast tempo", "slow ballad", INLINECODE207
"no vocals" → set make_instrumental=true
- With vocals: "female vocals" → set vocal_gender="female"
- Male vocals: "male vocals" → set vocal_gender="male"
- Mixed: Set INLINECODE214
"happy and energetic", "melancholic", "tense and dramatic", INLINECODE218
negative_tags: "heavy metal, distortion, screaming" to exclude unwanted elements
"60 seconds", "30 second loop", "2 minute track"
- Note: Suno typically generates ~2min, not strictly controllable
Example Suno prompts:
CODEBLOCK14
⚠️ Technical Note for Suno:
INLINECODE224 inside
parameters.parameters(e.g.,"sonic-v5") is different from the outermodel_versionfield (which is"sonic"). Always set both correctly.
CODEBLOCK15
When user requests unsupported combinations:
Note: Image-specific unsupported combinations (Midjourney + aspect_ratio, 8K, non-standard ratios) are documented in the "Image Models" section above.
User preferences have highest priority when they exist. But preferences are only saved when users explicitly express model preferences — not from automatic model selection.
~/.openclaw/memory/ima_prefs.jsonCODEBLOCK16
Step 1: Get knowledge-ai recommendation (if installed)
CODEBLOCK17
Step 2: Check user preference
CODEBLOCK18
Step 3: Decide which model to use
CODEBLOCK19
Step 4: Check for mismatch (for later hint)
CODEBLOCK20
✅ Save preference when user explicitly specifies a model:
| User says | Action |
|---|---|
INLINECODE230 / 换成XXX / INLINECODE232 | Switch to model XXX + save as preference |
INLINECODE233 / 默认用XXX / INLINECODE235 |
✅ 已记住!以后视频生成默认用 [XXX] |我喜欢XXX / 我更喜欢XXX | Save as preference |
❌ Do NOT save when:
🗑️ Clear preference when user wants automatic selection:
| User says | Action |
|---|---|
INLINECODE239 / 用最合适的 / best / INLINECODE242 | Clear pref + use knowledge-ai recommendation |
INLINECODE243 / 你选一个 / INLINECODE245 |
用默认的 / 用新的 | Clear pref + use knowledge-ai recommendation |试试别的 / 换个试试 (without specific model) | Clear pref + use knowledge-ai recommendation |重新推荐 | Clear pref + use knowledge-ai recommendation |
Implementation:
del prefs[f"user_{user_id}"][task_type]
save_prefs(prefs)
Selection flow:
Important notes:
The defaults below are FALLBACK only. User preferences have highest priority, then knowledge-ai recommendations.
v2.0 Updates (aligned with ima-image-ai v1.3):
- - Added Step 0 for correct message ordering (fixes group chat bug)
- Added Step 5 for explicit task completion
- Enhanced Midjourney support with proper timing estimates
- Now 6 steps total (0-5): Acknowledgment → Pre-Gen → Progress → Success/Failure → Done
This skill runs inside IM platforms (Feishu, Discord via OpenClaw).
Generation takes 10 seconds (music) up to 6 minutes (video). Never let users wait in silence.
Always follow all 6 steps below, every single time.
Default to plain-language updates in normal user flows.
If users ask for technical details, provide them transparently (script name, endpoints, and key parameters).
In standard progress messages, prioritize: model name, estimated/actual time, credits consumed, result URL, and natural-language status updates.
| Task Type | Model | Estimated Time | Poll Every | Send Progress Every |
|---|---|---|---|---|
| texttoimage | SeeDream 4.5 | 25~60s | 5s | 20s |
INLINECODE251 = upper bound of the range (e.g. 60 for SeeDream 4.5, 40 for Nano Banana2, 120 for Nano Banana Pro, 90 for Midjourney, 180 for Kling 2.6, 360 for Kling O1).
⚠️ CRITICAL: This step is essential for correct message ordering in IM platforms (Feishu, Discord).
Before doing anything else, reply to the user with a friendly acknowledgment message using your normal reply (not message tool). This reply will automatically appear FIRST in the conversation.
Example acknowledgment messages:
For images:
好的!来帮你画一只萌萌的猫咪 🐱
收到!马上为你生成一张 16:9 的风景照 🏔️
For videos:
好的!来帮你生成一段视频 🎬
For music:
CODEBLOCK27
Rules:
message toolWhy this matters:
After Step 0 reply, use the message tool to push a notification immediately:
CODEBLOCK28
Emoji by content type:
🎬(加注:视频生成需要较长时间,我会定时汇报进度)Cost transparency (new requirement):
Adapt language to match the user (Chinese / English). For video, always add a note that it takes longer. For expensive models, always mention cheaper alternatives unless user explicitly requested premium.
Poll the task detail API every [Poll Every] seconds per the table.
Send a progress update every [Send Progress Every] seconds.
CODEBLOCK29
Progress formula:
CODEBLOCK30
elapsed > estimated_max: freeze at 95%, append INLINECODE263When task status = success:
3.1 Send video player first (IM platforms like Feishu will render inline player):
CODEBLOCK31
Important:
3.2 Then send link as text (for copying/sharing):
CODEBLOCK32
⚠️ Critical for video:
CODEBLOCK33
Important:
Send audio file with player:
CODEBLOCK34
Step 0 — Initial acknowledgment (normal reply)
First reply with a short acknowledgment, e.g.: 好的,正在帮你把这段文字转成语音。 / OK, converting this text to speech.
Step 1 — Pre-generation (message tool)
Push once:
CODEBLOCK35
Step 2 — Progress
Poll every 2–5s. Every 10–15s send: ⏳ 语音合成中… [P]%,已等待 [elapsed]s,预计最长 [max]s. Cap progress at 95% until API returns success.
Step 3 — Success (message tool)
When resource_status == 1 and status != "failed", send media = medias[0].url and caption:
✅ 语音合成成功!
• 模型:[Model Name]
• 耗时:实际 [actual]s
• 消耗积分:[N pts]
🔗 原始链接:[url]
Step 4 — Failure (message tool)
On failure, send user-friendly message. TTS error translation (do not expose raw API errors):
| Technical | ✅ Say (CN) | ✅ Say (EN) |
|---|---|---|
| 401 Unauthorized | 密钥无效或未授权,请至 imaclaw.ai 生成新密钥 | API key invalid; generate at imaclaw.ai |
| 4008 Insufficient points |
Links: API key — https://www.imaclaw.ai/imaclaw/apikey ;Credits — https://www.imaclaw.ai/imaclaw/subscription
Step 5 — Done
After Step 0–4, no further reply needed. Do not send duplicate confirmations.
When task status = failed or any API/network error, send:
CODEBLOCK37
⚠️ CRITICAL: Error Message Translation
NEVER show technical error messages to users. Always translate API errors into natural language.
API key & credits: 密钥与积分管理入口为 imaclaw.ai(与 imastudio.com 同属 IMA 平台)。Key and subscription management: imaclaw.ai (same IMA platform as imastudio.com).
| Technical Error | ❌ Never Say | ✅ Say Instead (Chinese) | ✅ Say Instead (English) |
|---|---|---|---|
| INLINECODE271 🆕 | Invalid API key / 401 Unauthorized | ❌ API密钥无效或未授权<br>💡 生成新密钥: https://www.imaclaw.ai/imaclaw/apikey | ❌ API key is invalid or unauthorized<br>💡 Generate API Key: https://www.imaclaw.ai/imaclaw/apikey |
| INLINECODE272 🆕 |
"Invalid product attribute" / "Insufficient points" | Invalid product attribute | 生成参数配置异常,请稍后重试 | Configuration error, please try again later |Error 6006 (credit mismatch) | Error 6006 | 积分计算异常,系统正在修复 | Points calculation error, system is fixing |Error 6009 (no matching rule) | Error 6009 | 参数组合不匹配,已自动调整 | Parameter mismatch, auto-adjusted |Error 6010 (attribute_id mismatch) | Attribute ID does not match | 模型参数不匹配,请尝试其他模型 | Model parameters incompatible, try another model |error 400 (bad request) | error 400 / Bad request | 请求参数有误,请稍后重试 | Invalid request parameters, please try again |resource_status == 2 | Resource status 2 / Failed | 生成过程遇到问题,建议换个模型试试 | Generation failed, please try another model |status == "failed" (no details) | Task failed | 这次生成没成功,要不换个模型试试? | Generation unsuccessful, try a different model? |timeout | Task timed out / Timeout error | 生成时间过长已超时,建议用更快的模型 | Generation took too long, try a faster model |Generic fallback (when error is unknown):
Best Practices:
auto_lyrics=true)--list-models or shortening text.After sending Step 3 (success) or Step 4 (failure):
Why this step matters:
Exception: If the user explicitly asks "还有别的吗?" or similar, then respond naturally.
The Reflection mechanism (3 automatic retries) now provides specific, actionable suggestions for common errors:
All error handling is automatic and transparent — users receive natural language explanations with next steps.
Failure fallback by task type:
| Task Type | Failed Model | First Alt | Second Alt |
|---|---|---|---|
| texttoimage | SeeDream 4.5 | Nano Banana2 (4pts, fast) | Nano Banana Pro (10-18pts, premium) |
| texttoimage |
--list-models for alternatives | Use another model_id from product list |
Music-specific failure guidance:
TTS-specific failure guidance:
--task-type text_to_speech --list-models and suggest another model_id; or shorten text / simplify content. Use the TTS error translation table in "For TTS Tasks" above for user-facing messages.Source: production
GET /open/v1/product/list(2026-02-27). Model count reduced significantly. Always query product list API at runtime.
| Category | Name | modelid | Cost |
|---|---|---|---|
| texttoimage | SeeDream 4.5 🌟 | INLINECODE290 | 5 pts |
| textto_image |
midjourney | 8/10 pts (480p/720p) |
| texttoimage | Nano Banana2 💚 | gemini-3.1-flash-image | 4/6/10/13 pts |
| texttoimage | Nano Banana Pro | gemini-3-pro-image | 10/10/18 pts |
| imagetoimage | SeeDream 4.5 🌟 | doubao-seedream-4.5 | 5 pts |
| imagetoimage | Midjourney 🎨 | midjourney | 8/10 pts (480p/720p) |
| imagetoimage | Nano Banana2 💚 | gemini-3.1-flash-image | 4/6/10/13 pts |
| imagetoimage | Nano Banana Pro | gemini-3-pro-image | 10 pts |
Midjourney attributeids: 5451/5452 (texttoimage), 5453/5454 (imageto_image)
Nano Banana2 size options: 512px (4pts), 1K (6pts), 2K (10pts), 4K (13pts)
Nano Banana Pro size options: 1K (10pts), 2K (10pts), 4K (18pts for t2i / 10pts for i2i)
⚠️ Critical: Models have varying parameter support. Custom aspect ratios are now supported by multiple models.
| Model | Custom Aspect Ratio | Max Resolution | Size Options | Notes |
|---|---|---|---|---|
| SeeDream 4.5 | ✅ (via virtual params) | 4K (adaptive) | 8 aspect ratios | Supports 1:1, 16:9, 9:16, 4:3, 3:4, 2:3, 3:2, 21:9 (5 pts) |
| Nano Banana2 |
attribute_id |attribute_id |attribute_id | Fixed 1024x1024, artistic style focus |
Key Capabilities:
attribute_ids关键提示: 调用脚本时,必须使用精确的 modelid(第二列),而不是友好的模型名称。请勿从友好名称推断 modelid(例如,❌ nano-banana-pro 是错误的;✅ gemini-3-pro-image 是正确的)。
快速参考表:
| 友好名称 | model_id | 说明 |
|---|---|---|
| Nano Banana2 | gemini-3.1-flash-image | ❌ 不是 nano-banana-2,预算选择 4-13 积分 |
| Nano Banana Pro |
| 友好名称 | modelid (文生视频) | modelid (图生视频) | 说明 |
|---|---|---|---|
| Wan 2.6 | wan2.6-t2v | wan2.6-i2v | ⚠️ 注意 -t2v/-i2v 后缀 |
| IMA Video Pro (Sevio 1.0) |
| 友好名称 | model_id | 说明 |
|---|---|---|
| Suno (sonic v4) | sonic | ⚠️ 简化为 sonic |
| DouBao BGM |
| 友好名称 | model_id | 说明 |
|---|---|---|
| seed-tts-2.0 | seed-tts-2.0 | ✅ 与友好名称相同(默认) |
如何获取正确的 model_id:
运行时真实数据源:GET /open/v1/product/list(或 --list-models)。
本文档中的任何表格仅供参考;实际可用性取决于当前产品列表。
示例:
bash
此技能可作为独立包完整运行。
如果安装了 ima-knowledge-ai,代理可以读取其参考资料以进行工作流分解和一致性指导。
推荐的可选阅读:
为什么这很重要:
多媒体工作流示例:
用户:帮我做个产品宣传MV,有背景音乐,主角是旺财小狗
❌ 错误:
1. 生成狗的图像(随机外观)
2. 生成视频(不同的狗)
3. 生成音乐(不相关)
✅ 正确:
1. 阅读 workflow-design.md + visual-consistency.md
2. 生成主参考:旺财小狗图片
3. 使用图生视频生成视频镜头,以旺财作为第一帧
4. 获取视频时长(例如,15秒)
5. 生成匹配时长和氛围的背景音乐
如何检查:
python
if same subject or series or character:
read(~/.openclaw/skills/ima-knowledge-ai/references/visual-consistency.md)
if video_generation:
read(~/.openclaw/skills/ima-knowledge-ai/references/video-modes.md)
无例外 — 对于简单的单媒体请求,可以直接进行。对于复杂的多媒体工作流,请先阅读知识库。
目的: 使任何代理都能一致地解析用户意图,首先从用户请求中确定媒体类型,然后选择 task_type 和模型。
| 用户意图 / 关键词 | 媒体类型 | tasktype 示例 |
|---|---|---|
| 画 / 生成图 / 图片 / image / 画一张 / 图生图 | 图像 | texttoimage, imagetoimage |
| 视频 / 生成视频 / video / 图生视频 / 文生视频 |
以下为平台配置的接入选项,并非逐项实测通过。能否安装取决于客户端支持、技能来源和运行环境:
帮我安装 SkillHub 和 ima-all-ai-1776112442 技能
设置 SkillHub 为我的优先技能安装源,然后帮我安装 ima-all-ai-1776112442 技能
skillhub install ima-all-ai-1776112442
文件大小: 56.09 KB | 发布时间: 2026-4-14 09:56
