top of page

AI 影片模型怎麼選?Seedance 2.5、MiniMax H3、Wan 3.0 五種情境實測

2天前
讀畢需時 28 分鐘

已更新:5小时前

現在 AI 影片模型越來越多,但實際開始製作之後,很快就會發現一件事:

同一份素材、同一段提示詞,換一個模型,做出來的結果可能完全不一樣。

有些模型比較擅長人物與角色一致性,有些在 Reference 參考影片的跟隨度比較好,也有些模型在細節、運鏡或首尾幀轉換上反而更自然。

所以這次小步透過 ChatArt,挑了目前比較熱門的 3 款 AI 影片模型:

Seedance 2.5、MiniMax H3、Wan 3.0

並且使用相同的素材與提示詞,分別測試 5 種不同的影片類型。

這篇文章除了整理實測結果,也會把每一個案例的製作方式、準備素材,以及影片中使用的提示詞一起整理下來。


實測影片👇



這次測試的 5 種影片類型

測試案例

主要測試能力

小步實測觀察

首尾幀、商品一致性、材質與手部動作

MiniMax H3 表現較自然

角色一致性、表演、鏡頭、台詞與嘴型

Seedance 寫實感較好

圖片+影片的多模態參考能力

Seedance 跟隨度最好

首尾幀轉換、文字形成與畫面細節

MiniMax H3 細節讓人驚喜

多鏡頭理解、角色一致性、運鏡合理性

MiniMax H3 整體較合理

這次測試下來最大的感想其實很簡單:

沒有一個模型適合所有影片。

真正開始製作之前,反而應該先想清楚,這支影片最重要的是角色一致性、Reference 跟隨、畫面細節,還是攝影機運動。


案例 1|電商商品廣告

第一個案例,是最近滿常看到的一種商品影片效果。

商品一開始位在有霧氣的玻璃後方,接著一隻手把玻璃上的霧氣擦開,最後完整露出後方的商品。這類影片看起來不複雜,但如果真的要拿來做商品廣告,其實有幾個地方很重要。

商品本身不能在影片生成過程中突然變形,瓶身、標籤與材質也要盡量維持一致;另外像手部動作、玻璃上的霧氣、水珠,以及擦拭之後留下的痕跡,也都會直接影響畫面的真實感。

實測結果

三個模型放在一起比較之後,小步自己比較喜歡 MiniMax H3 的效果。

主要差異是在霧氣擦掉之後,MiniMax H3 生成的水滴會沿著玻璃緩緩往下滑,質感比另外兩個模型自然一些,整體也比較接近實際拍攝的商品廣告。

製作方式

這個案例需要先準備兩張圖片。

第一張是商品位在霧面玻璃後方的畫面,第二張則是手已經擦開部分霧氣、露出商品的畫面。

如果原本只有一張乾淨的商品圖,也可以先透過 ChatArt 的圖片生成功能,把原始產品圖轉換成需要的兩張素材。

原始商品圖

有霧氣商品圖

手撥開霧氣商品圖


生成有霧氣商品圖提示詞

Create a premium skincare advertisement image based on the same product: VERDANT BOTANICAL SERUM.

Use the same bottle design: a tall cylindrical transparent light green glass bottle, silver pump collar, white pump top, minimalist front label with the text “VERDANT” and “BOTANICAL SERUM”, and a small botanical leaf icon near the bottom.

Place the product behind a foggy glass surface covered with mist and water droplets.

The bottle should still be visible and recognizable, but partially obscured by the frosted moisture on the glass.

The scene should feel elegant, fresh, and mysterious, like a teaser for a luxury skincare ad.

Soft botanical green tones, soft diffused lighting, premium beauty ad aesthetic.

The composition should be centered and vertical.

No hand yet.

Keep the product shape and label consistent with the original product image.

生成手撥開霧氣商品圖提示詞

Create a premium skincare advertisement image based on the same product: VERDANT BOTANICAL SERUM.

Use the same bottle design: a tall cylindrical transparent light green glass bottle, silver pump collar, white pump top, minimalist front label with the text “VERDANT” and “BOTANICAL SERUM”, and a small botanical leaf icon near the bottom.

The product is placed behind a foggy glass surface with mist and water droplets, but now a human hand has wiped away part of the fog, creating a clear transparent area in the center.

Through the wiped-clear area, the product is now sharply visible and revealed.

The hand should remain touching the glass, as if it has just wiped the fog away.

Keep the surrounding edges foggy and covered with moisture so the reveal effect is obvious.

The image should feel fresh, elegant, cinematic, and suitable for a luxury skincare commercial.

Soft botanical green tones, premium beauty ad aesthetic, centered vertical composition.

Keep the product shape, label, and proportions consistent with the original product image.

圖片準備完成之後,再進入圖生影片功能。

這個案例其實直接使用 首尾幀模式 就可以(雖然小步在影片中是使用全能參考模式),將兩張圖片分別放入首幀與尾幀,再加入影片提示詞。

這次小步設定的影片長度是 6 秒,三個模型都使用相同素材、提示詞與長度進行測試。



生成影片提示詞

Create a short premium skincare advertisement based on the provided images. Start from the foggy teaser state and transition naturally into the reveal state.

A human hand gently wipes across the fogged glass surface in one smooth motion, clearing away the mist and moisture. As the glass is wiped, the VERDANT BOTANICAL SERUM bottle is gradually revealed clearly in the center.

Keep the product stationary, centered, and consistent throughout the shot. Preserve the exact bottle shape, proportions, label layout, text, green glass color, silver pump collar, and white pump top. The hand movement should feel natural, elegant, and realistic, like a luxury skincare commercial.

Maintain realistic foggy glass, condensation, water droplets, and wipe marks. The surrounding area should remain slightly foggy and moist, while the wiped area becomes clear and transparent. Use soft, clean, high-end beauty advertisement lighting and a refined premium aesthetic.

Single continuous shot, smooth motion, no cuts, no sudden camera shake, no abrupt scene changes. The result should feel like a polished luxury skincare ad reveal.


案例 2|AI 日系短劇

這類影片不只是要讓人物動起來,而是同時要處理:

  • 角色一致性

  • 人物表演

  • 鏡頭切換

  • 場景與空間連續性

  • 台詞

  • 嘴型

這次我設計的是一段大約 15 秒、9:16 的日系戀愛短劇

情境設定在下雨的夜晚,男女主角在便利商店裡同時看上最後一把透明雨傘。

兩個人原本在爭同一把傘,最後卻決定一起撐傘離開。

故事本身很簡單,但對 AI 影片模型來說,其實同時包含了人物互動、多鏡頭敘事、表情,以及語音對嘴等不同能力。

實測結果

三個模型生成完成之後,如果以 整體寫實感與情境氛圍 來看,我自己會比較喜歡 Seedance 2.5 的結果。

它生成的人物互動和整體畫面,看起來比較接近真人短劇。

這個案例有一個我覺得很有趣的地方。

製作時,我其實只提供了 男女主角兩張角色參考圖

並沒有另外準備便利商店場景、鏡位或每一個畫面的分鏡圖。

所以:

場景長什麼樣子、鏡頭怎麼切、人物怎麼演,其實都是模型根據同一份提示詞自行理解後產生的。

也因此,三個模型最後生成出來的故事雖然相同,但在構圖、鏡頭安排與人物表演上,都可以看到不同模型自己的理解方式。

這類案例也很適合拿來觀察:

模型究竟只是把提示詞裡的內容「做出來」,還是真的能把文字轉換成一段合理的影音敘事。


製作方式

這個案例只需要先準備兩張角色參考圖:

角色素材

用途

女主角參考圖

固定女主角外觀

男主角參考圖

固定男主角外觀

接著進入 ChatArt 的影片生成功能,把兩張人物圖片加入參考素材。

這次影片設定為:

  • 影片比例:9:16

  • 影片長度:約 15 秒

  • 類型:真人日系戀愛短劇

  • 語言:日語

提示詞裡除了描述劇情之外,我也直接把每一段需要出現的台詞寫進去。

這次使用的四句日文台詞是:

「それ、私が先に見つけたんですけど。」「じゃあ、一緒に使います?」「近いです…」「濡れるよりいいでしょ。」

中文大意分別是:

「那個,是我先找到的。」「那就一起用吧?」「太近了……」「總比淋濕好吧。」

這裡其實也是在測模型除了畫面之外,能不能同時處理 人物說話、情緒與嘴型

另外,這支影片的提示詞是用 分段鏡頭 的方式來描述,讓模型知道 15 秒裡每一段大概要發生什麼事情。


點我展開提示詞

生成影片提示詞

@圖片 1:是女主角色參考圖, @圖片 2:是男主角色參考圖。

請根據這兩張角色參考圖,生成一支 15 秒、9:16 直式、日本真人戀愛短劇風格的短片。

整支影片都要穩定維持角色 1(女主)與角色 2(男主)的長相、髮型、服裝、身形比例與整體氣質一致。

【整體風格】

畫面風格像日本真人電視戀愛短劇中的日常橋段。

整體偏寫實、低飽和、自然細膩,有安靜、曖昧、輕微心動的情緒。

人物表演要克制自然,以眼神、停頓、呼吸與細微表情來呈現情緒。

鏡頭語言像真人攝影師在現場拍攝,使用 35mm~50mm 的真人戲劇攝影感,中近景、雙人鏡頭與人物反應特寫為主。

【場景與空間關係】

場景位於東京住宅區街角的一間便利商店門口,時間是夜晚,下著雨。

便利商店門口的傘架上只剩最後一把透明雨傘。

地面潮濕,街道與店面燈光在雨水中形成自然反光。

角色 1(女主)一開始站在便利商店門口靠左的位置,距離傘架約 1 公尺。

角色 2(男主)從右側靠近傘架,兩人初始距離約 1.5 公尺。

鏡頭初始高度約離地 1.6 公尺,距離女主約 2 公尺,位於她左前方約 45 度的位置。

【光影邏輯】

主光源來自便利商店內部與門口招牌的暖白光,主要照亮人物臉部、上半身與騎樓區域。

外部街道以偏冷色的雨夜環境光為主,形成自然的暖冷對比。

人物移動時,臉部受光方向與亮度要依照人物和便利商店之間的實際位置自然改變。

當兩人走出騎樓進入雨中時,便利商店的暖光逐漸減弱,冷色雨夜環境光增加。

透明雨傘可以產生輕微反光與折射,但保持自然,不要過度誇張。

【分鏡與攝影機運動】

Shot 1(0–4 秒)

便利商店門口的傘架上只剩最後一把透明雨傘。

角色 1(女主)與角色 2(男主)幾乎同時伸手握住傘柄。

鏡頭以雙人中景開始,從女主左前方略微向前推近,同時緩慢向右側移動,讓兩人的手與透明雨傘同時進入畫面。

攝影機運動保持自然慣性,不要突然加速或停止。

當兩人的手同時握住傘柄時,鏡頭短暫停留,捕捉兩人抬眼看向對方的反應。

女主帶一點不服氣、但不誇張的語氣說:

「それ、私が先に見つけたんですけど。」

說話時,鏡頭維持在雙人中近景,讓兩人的表情都保持可見。

Shot 2(4–8 秒)

男主沒有立刻放手。

他先側頭看了一眼便利商店外的大雨,再慢慢把視線轉回女主。

鏡頭略微向後退約半步,保留兩人的肩部、手部與透明雨傘,再自然向男主方向偏移,形成男主中近景。

男主以平靜、自然、帶一點溫柔的語氣說:

「じゃあ、一緒に使います?」

男主說完後,鏡頭自然轉向女主。

攝影機可以稍微晚半拍跟上她的反應,捕捉她先愣住、再露出些微驚訝與不好意思的表情。

Shot 3(8–12 秒)

畫面切到兩人一起走出便利商店騎樓,共撐同一把透明雨傘走進雨裡。

因為雨傘不大,兩人的距離縮短到約 30 公分。

鏡頭位於兩人左前側,與兩人保持約 1.5 公尺距離,離地約 1.6 公尺,攝影機一邊側向移動、一邊緩慢向後跟拍。

攝影機運動要有真人跟拍的慣性,允許略微提前或稍微落後人物半拍。

雨傘邊緣與人物肩膀可以短暫遮擋部分畫面,增加自然的現場拍攝感。

女主微微低頭,因為兩人的距離太近而有些害羞,小聲說:

「近いです…」

女主說完後,攝影機可以短暫出現極輕微失焦,再自然重新對焦到兩人的臉部。

Shot 4(12–15 秒)

男主側頭看向女主,露出淡淡的笑意。

鏡頭從雙人中近景緩慢向前推進,先停留在男主說話時的表情,再自然偏移到女主。

男主以輕鬆、自然、帶一點溫柔的語氣說:

「濡れるよりいいでしょ。」

女主聽完後先愣一下,再低下視線,最後忍不住露出有點害羞又心動的淡淡微笑。

最後鏡頭停留在女主近景約 1 秒,讓她的情緒有完整收尾。

【鏡頭呼吸感】

保留真人拍攝的自然呼吸感:

- 攝影機允許略微提前或略微滯後人物動作

- 攝影機移動要有慣性,不要像機械滑軌一樣完全僵硬

- 人物、雨傘或前景物件可以短暫遮擋部分鏡頭

- 快速調整鏡位後,可以出現極短暫的輕微失焦,再自然重新對焦

- 人物說話與反應之間保留自然停頓,不要所有動作完全同步或過度精準

【聲音與表演】

台詞必須由角色自然說出,使用日語發音。

女主台詞:

「それ、私が先に見つけたんですけど。」

「近いです…」

男主台詞:

「じゃあ、一緒に使います?」

「濡れるよりいいでしょ。」

口型盡量配合日語台詞。

環境聲包含:

- 持續但不過大的雨聲

- 遠處微弱車流聲

- 便利商店門口輕微環境聲

人物聲音要清楚,環境音不要蓋過台詞。

【重要要求】

- 男女主角在所有鏡頭中保持一致

- 臉部、髮型、服裝、身形比例保持一致

- 透明雨傘在所有鏡頭中保持一致

- 雙人互動自然

- 手部與握傘動作自然

- 人物表情細膩、克制

- 情緒轉折清楚

- 分鏡之間的空間關係一致

- 光影方向符合實際場景

- 鏡頭運動有明確方向與慣性

- 保留自然的真人攝影感

- 不顯示字幕或畫面文字

- 最終成品要呈現日本真人戀愛短劇的自然敘事感、情緒感與現場拍攝質感


案例 3|Reference 參考影片

這幾天我在 Threads 上看到一種韓國很紅的「寶寶變身魔法少女」影片。

它不是單純用文字告訴 AI:「幫這個角色變身。」

而是同時提供:角色圖片 + 一支已經有完整動作與節奏的參考影片。

讓 AI 參考原影片的:

  • 人物動作

  • 動作節奏

  • 畫面變化

  • 角色走位

  • 鏡頭安排

再把影片裡的人物換成自己的角色。

所以這次我也拿「小小步」來實際測試。


第一輪:30 秒魔法變身影片

第一輪測試,我準備了:

變身前

變身後

生成變身照片提示詞

Transform the character in the reference image into a magical girl version of the same person.


Keep her identity completely consistent with the reference image:

same face, facial features, short wavy reddish-brown hair, black-rimmed glasses, body shape, height, age, and natural body proportions.


Replace only her outfit with an elegant magical girl costume.


The outfit should feature a navy blue, white, and soft pink color palette, with a refined sailor-inspired fantasy design.


Include:

- a white fitted top with a navy sailor-style collar

- a large navy bow on the chest

- subtle pink ribbon accents

- gold star and crescent-moon ornaments

- layered navy and white pleated skirt

- delicate ruffles and decorative ribbons

- white wrist cuffs

- a navy choker with a small star pendant

- white ankle boots with navy ribbons and gold star details


The costume should feel elegant, youthful, magical, and detailed, while still matching the character's original personality.


Keep the same front-facing full-body composition, white studio background, soft lighting, natural pose, and realistic character style.


Do not change her face, hairstyle, glasses, age, body proportions, or identity.

Do not make her taller, thinner, older, or more mature.

She must clearly remain the exact same character as the reference image.


接著讓模型依照 Reference 裡面的動作與節奏,生成小小步自己的魔法變身影片。

不過這一輪其實沒有辦法真正做到三個模型公平比較。

原因並不是影片生成失敗,而是三個模型本身對 Reference 素材長度的限制不同。

這次找到的參考影片大約有 30 秒

  • Seedance 2.5:可以完成測試

  • MiniMax H3:最長生成 15 秒

  • Wan 3.0:Reference 素材不能超過 15 秒

所以 MiniMax H3 與 Wan 3.0 一開始就無法用完全相同的條件測試。

也因為這樣,我後來又另外準備了一組比較短的 Reference。

點我展開提示詞

生成影片提示詞

Use @影片 1:as the REFERENCE VIDEO.

Use @圖片 1:as IMAGE 1.

Use @圖片 2:as IMAGE 2.

Use the provided REFERENCE VIDEO as the primary and fixed reference for the magical transformation sequence, transformation timing, cinematic energy, camera motion, visual progression, transformation choreography, final reveal, ending pose, and overall pacing.

Use the TWO PROVIDED XIAOXIAOBU IMAGES as strict appearance references.

IMAGE 1 = BEFORE TRANSFORMATION REFERENCE

Xiaoxiaobu must start exactly as shown in IMAGE 1.

IMAGE 2 = AFTER TRANSFORMATION REFERENCE

The transformation must finish with Xiaoxiaobu exactly as shown in IMAGE 2.

Do not reinterpret, redesign, simplify, alter, or introduce additional details to either appearance.

Throughout the entire sequence, maintain the exact same Xiaoxiaobu facial identity, apparent age, hairstyle, facial proportions, glasses, body shape, and natural body proportions shown in the provided reference images.

Xiaoxiaobu must NEVER become taller, thinner, older, younger, or significantly different in body proportions at any point during the transformation.

# ==================================================

CHARACTER BODY PROPORTIONS — EXTREMELY IMPORTANT

Maintain the natural youthful female physique and body proportions shown in the provided reference images.

Preserve the original head-to-body ratio, torso proportions, leg-to-body ratio, arm-to-body ratio, shoulder width, hand size, and foot size shown in the provided photographs.

The arms and legs must remain anatomically natural and proportionally consistent throughout the entire transformation.

Do not extend or lengthen the arms or legs for dramatic posing or cinematic movement.

Preserve the original leg-to-body ratio and arm-to-body ratio shown in the provided reference images.

Every body movement and pose must respect Xiaoxiaobu's realistic human anatomy and natural physical proportions.

The hands must remain anatomically natural, with exactly five fingers on each hand.

The feet must remain anatomically natural, with exactly five toes on each foot.

No additional fingers, missing fingers, fused fingers, stretched fingers, extra toes, malformed hands, malformed feet, twisted limbs, elongated limbs, duplicated limbs, or unnatural body proportions.

# ==================================================

TRANSFORMATION SEQUENCE

Xiaoxiaobu starts exactly as shown in IMAGE 1.

Follow the transformation progression, rhythm, motion, and visual energy established by the REFERENCE VIDEO.

A gentle magical glow gradually begins to appear around Xiaoxiaobu's body.

Tiny shimmering particles and delicate points of light begin drifting through the surrounding space.

Glowing ribbons of pale blue, soft white, silver, iridescent pastel light, and subtle cosmic energy gently move around Xiaoxiaobu's body.

The magical energy should flow naturally around Xiaoxiaobu without altering or stretching her anatomy.

The appearance shown in IMAGE 1 gradually transitions into the appearance shown in IMAGE 2.

The transition must feel smooth and continuous rather than resembling a sudden image replacement.

Soft luminous energy moves around Xiaoxiaobu's torso, arms, and legs as the transformation develops.

Elements from IMAGE 1 gently dissolve into sparkling light particles while the corresponding elements from IMAGE 2 naturally emerge through the magical glow.

Do not create intermediate costumes, accessories, hairstyles, or visual elements that are not supported by the provided images.

Avoid sudden replacements or hard cuts between IMAGE 1 and IMAGE 2.

The transformation should feel seamless, elegant, magical, dreamy, and visually satisfying.

During the transformation, use graceful sparkling trails, luminous outlines, translucent magical ribbons, small star-like particles, cosmic dust, soft lens glints, iridescent reflections, and subtle waves of glowing energy.

Keep the magical effects visually beautiful and cinematic rather than explosive or aggressive.

Do not cover or obscure Xiaoxiaobu's face for extended periods.

Xiaoxiaobu's face must remain recognizable and visually consistent whenever it is visible.

# ==================================================

DREAMCORE COSMIC ENVIRONMENT

Create a surreal and dreamy cosmic transformation environment.

The entire scene should take place within an ethereal, dreamlike universe rather than a realistic physical setting.

Use deep midnight blue, indigo, violet, pale blue, and subtle pastel tones to create the cosmic atmosphere.

Surround Xiaoxiaobu with distant stars, drifting particles, soft cosmic dust, dreamy nebulous clouds, translucent waves of light, faint celestial glows, and subtle iridescent gradients.

The background should feel endless, weightless, surreal, nostalgic, whimsical, and dreamcore.

Create the sensation of a magical dream suspended somewhere between outer space and an imaginary fantasy realm.

Soft glowing halos and circular waves of light may appear behind Xiaoxiaobu during important transformation moments.

Small stars and shimmering particles may slowly drift across different depths of the scene to create a sense of dimensionality.

Occasional dreamy light blooms, softly focused celestial shapes, floating luminous particles, and gentle atmospheric haze may appear in the distance.

Keep the environment visually beautiful while maintaining a spacious composition.

Xiaoxiaobu must always remain the strongest visual focus.

Avoid realistic rooms, furniture, buildings, streets, text, logos, crowds, or other distracting elements.

The atmosphere should feel fantastical, magical, cosmic, ethereal, dreamy, and suitable for the character.

# ==================================================

CAMERA

Follow the overall camera language and cinematic progression established by the REFERENCE VIDEO while prioritizing Xiaoxiaobu's anatomy and identity.

Keep Xiaoxiaobu centered during the main transformation moments.

Prefer full-body framing whenever possible so Xiaoxiaobu's natural body proportions remain clearly visible.

Camera movement should feel smooth, floating, magical, and cinematic.

The camera may gently orbit, move closer, pull back, or float around Xiaoxiaobu whenever appropriate to the reference transformation.

Do not use perspective distortion that causes Xiaoxiaobu's legs to appear longer.

Do not use extreme low-angle camera shots.

Do not stretch or distort Xiaoxiaobu's body simply to fill the frame.

Maintain Xiaoxiaobu's original body proportions throughout every camera movement.

# ==================================================

FINAL REVEAL & ENDING POSE — IMPORTANT

Once the transformation is fully complete, clearly reveal Xiaoxiaobu exactly as shown in IMAGE 2.

Do not end the video immediately once the transformation is finished.

Include a distinct FINAL HERO MOMENT after the transformation has been completed.

The magical energy gradually settles and opens around Xiaoxiaobu, clearly revealing the complete transformed appearance shown in IMAGE 2.

Xiaoxiaobu then naturally moves into a cute and confident final ending pose inspired by the final pose and timing of the REFERENCE VIDEO.

The ending pose must remain physically appropriate, natural, and realistic for Xiaoxiaobu.

Keep the pose simple, stable, charming, confident, and achievable with Xiaoxiaobu's natural body proportions.

Do not force Xiaoxiaobu into an exaggerated superhero stance, an unnaturally wide stance, or any anatomically difficult position.

During the final pose, gently pull the camera backward or stabilize it into a clean full-body hero composition.

A soft celestial halo or luminous cosmic glow appears behind Xiaoxiaobu.

Shimmering particles, tiny stars, and magical light trails gradually settle around Xiaoxiaobu.

Clearly hold the completed pose for a brief moment so the viewer can recognize the final IMAGE 2 appearance.

The final frame should feel like a satisfying conclusion to the magical transformation sequence.

Finish with Xiaoxiaobu holding the final pose while the surrounding cosmic particles and dreamlike glow gradually fade away.

# ==================================================

CRITICAL CONSISTENCY RULES

Only one Xiaoxiaobu.

The same Xiaoxiaobu from beginning to end.

IMAGE 1 is strictly the BEFORE transformation reference.

IMAGE 2 is strictly the AFTER transformation reference.

Do not describe, assume, or introduce specific clothing, colors, accessories, or styling beyond what is actually visible in IMAGE 1 and IMAGE 2.

No facial morphing.

No identity changes.

No age progression.

No age regression.

No body growth.

No body shrinking.

No major body proportion changes.

No elongated arms.

No elongated legs.

No oversized hands or feet.

No anatomical deformation.

No additional fingers or toes.

No missing fingers or toes.

No duplicated limbs.

No distorted face.

No gender changes.

No random hairstyle alterations.

No random accessories.

No invented clothing.

No additional characters.

The BEFORE appearance must follow IMAGE 1.

The AFTER appearance must follow IMAGE 2.

The transformation may alter Xiaoxiaobu's appearance from IMAGE 1 into IMAGE 2, but it must NEVER alter Xiaoxiaobu's identity, apparent age, facial structure, glasses, or natural human anatomy.


第二輪:8 秒舞蹈 Reference

第二次我改用最近很紅的韓國「神之八秒」舞蹈。

這次 Reference 影片比較短,因此三個模型都可以用相同條件測試。

我另外生成了三位舞者的人物參考圖,並把老闆放在中間當 C 位。

準備的素材包含:

素材

用途

舞者 1 角色圖

對應左側舞者

老闆真人角色圖

對應中間 C 位

舞者 3 角色圖

對應右側舞者

8 秒舞蹈影片

Reference 動作與節奏

舞者1

老闆

舞者2


這一題真正測的,其實就是模型的 多模態參考能力

不是只看人物長得像不像。

而是:

AI 能不能同時理解「這個角色要長這樣」,以及「這個角色要照影片裡的人這樣動」。

實測結果

我把 Reference 和三個模型的影片放在一起比對之後,這一題我覺得 Seedance 2.5 對參考影片的遵循度最好

尤其是在:

  • 動作

  • 節奏

  • 拍點

幾個地方,都比較接近原本的 Reference。

MiniMax H3 有些動作的節拍會稍微對不上。

Wan 3.0 也有類似的情況。

所以這一題有一個很重要的觀察:

模型支援 Reference,不代表它就能精準跟隨 Reference。

「可以讀取影片」和「真的理解並重現影片裡的動作」,其實是兩件不同的事情。


製作方式

製作 Reference 影片時,除了上傳參考影片之外,角色圖也要準備清楚。

尤其這次有三位人物,所以提示詞不能只寫:

「把影片裡的人換成這三個人。」

而是要描述清楚:

  • 哪一張圖片對應哪一個角色

  • 哪一支影片是動作 Reference

  • 人物在畫面中的位置

  • 整體場景

  • 希望保留哪些原影片特徵

也就是讓模型知道:

圖片負責決定「誰來演」,影片負責決定「怎麼演」。

這也是 Reference 類 AI 影片和一般文生影片最大的差異之一。


點我展開提示詞

生成影片提示詞

以 @影片 1: 作為主要且固定的動作參考,完整跟隨參考影片中的三人舞蹈動作、節奏、站位、肢體幅度與動作順序。

@圖片 1:= 中間主舞者

@圖片 2:= 左側舞者

@圖片 3:= 右側舞者

生成一段約 8 秒的真人舞蹈影片。

三位男性舞者並排站立,中間的老闆為主角,從頭到尾保持在 C 位;左右兩位男舞者依照參考影片中的位置配合演出。

三人完整模仿 @影片 1:中對應角色的舞蹈動作與節奏,包括身體重心轉移、手臂動作、腿部步伐、轉身與同步節拍。三人的動作必須清楚、流暢、有力量感,彼此保持正確距離,不互相穿模或交換位置。

嚴格保持三位角色各自的臉部身份、髮型、服裝、身形比例與外觀一致,不換臉、不交換角色、不改變服裝。

中間老闆必須始終保持主角辨識度,動作自信、自然、有男團主舞者的舞台感。

場景為深色現代練舞室/簡潔舞台,頂部有柔和聚光燈,背景乾淨,不要加入多餘人物或複雜特效。

攝影機以正面全身構圖為主,確保三人的完整舞蹈動作都清楚可見。鏡頭固定或僅做非常輕微的自然跟隨,不做大幅度運鏡,不切鏡。

整段影片的最優先目標是:

1. 準確跟隨 @影片 1:的三人舞蹈動作

2. 保持三位角色身份一致

3. 保持三人站位與隊形穩定

4. 動作自然、同步、符合真人人體結構

不要增加第四個角色。

不要讓三人互換位置或身份。

不要產生多餘手腳、手指或肢體變形。

不要改變臉部、髮型與服裝。

不要加入字幕、文字或 Logo。

案例 4|中文字體動畫

第四個案例表面上看起來是在測 中文字體動畫

但其實我真正想測的,是三個模型對 首尾幀圖片之間的轉換能力

這次我先準備了一張奇幻森林的背景圖,以及完成後帶有「小步奇境」四個繁體中文字的畫面。

我希望 AI 做的不是重新生成另一個標題。

而是讓:「小步奇境」四個字從左到右,依序在畫面裡形成。


這種用法其實很適合:

  • 影片片頭

  • Logo 動畫

  • 標題動畫

  • 活動視覺

  • 已經設計好的海報動態化

因為我們真正需要的,通常不是讓 AI 重新設計整張圖。

而是:

保留原本設計,只把靜態畫面中間的變化補出來。

實測結果

三個模型都有理解「文字由左到右依序出現」這個主要要求。

不過實際生成的風格差異蠻明顯。

製作方式

這個案例需要準備兩張圖片:

圖片

用途

無文字的奇幻森林背景

首幀

完整顯示「小步奇境」的成品圖

尾幀

兩張圖的構圖要盡量一致。

差別主要放在:

「文字還沒有出現」→「文字完整出現」

這樣模型在處理首尾幀時,比較容易理解真正需要生成的是「中間的文字形成動畫」,而不是重新設計整個場景。

這次提示詞的主要要求是:

  • 「小步奇境」由左至右依序形成

  • 文字由金色光粒、螢火蟲、星塵與能量逐漸匯聚

  • 森林薄霧、植物、湖面保留些微動態

  • 文字生成後完整停留

  • 鏡頭保持固定

也就是把影片的動態集中在:

文字生成+環境的細微變化。

而不是另外加入明顯的鏡頭推拉或場景切換。

生成影片提示詞

模擬月夜奇幻森林中的標題顯現動畫,畫面從無文字的夢幻森林背景開始,中央的「小步奇境」四個繁體中文字由左至右逐漸由金色光粒、星光微塵、螢火蟲光點與柔和月光能量聚拢形成,文字筆畫與裝飾圖案依序完整顯現,周圍薄霧輕輕流動,植物輕微晃動,湖面閃爍月光倒影,最後完整停留在「小步奇境」標題畫面,鏡頭固定不動。


案例 5|電影級運鏡

這個案例測的是:如果只有一張圖片,AI 能不能靠提示詞自己完成一段複雜的電影運鏡?

這次我只準備了一張小步的東方奇幻角色圖片,把它當成影片首幀,接著就是透過提示詞,告訴模型:接下來 10 秒鐘,攝影機應該怎麼移動。

這類影片和前面的 Reference 很不一樣。

Reference 是把「怎麼動」直接提供給模型看。

這一題則是完全靠文字描述,測模型能不能理解三維空間中的:

  • 前後景關係

  • 人物遮擋

  • 環繞運鏡

  • 鏡頭重新揭露角色

  • 最後推近人物表情

製作方式

這個案例的素材最簡單:只需要一張首幀角色圖片。


點我展開圖片生成提示詞細節說明

【使用方式】

這份提示詞適合用在:

  • 已經有一張角色參考圖

  • 想做角色設定板 / 角色世界觀概念板

  • 希望後續延伸做影片、系列角色圖、場景圖


使用前,請先準備:

  • 1 張角色參考圖

  • 想好的 角色名稱

  • 想好的 角色定位

  • 想要的 世界觀風格

  • 想放在圖上的 標題文案

【哪些地方要自己改?】

請優先修改這幾個區塊:

  • 【請修改|角色名稱】

  • 【請修改|角色定位】

  • 【請修改|服裝設定】

  • 【請修改|世界觀場景】

  • 【請修改|主標題】

  • 【請修改|副標題】

  • 【請修改|英文標語】

  • 【請修改|品牌句 / 標語】

  • 【請修改|角色文案】

  • 【請修改|色票】

如果暫時沒有靈感,也可以先保留原架構,只改最基本的:

  • 角色名稱

  • 角色定位

  • 世界觀

  • 標題文字

Create a premium character design and worldbuilding concept board for an original character named "【請修改|角色名稱】 — 【請修改|角色定位】".


Use the uploaded reference image as the primary source for the character's appearance.


Keep the character's identity consistent with the reference image, including:

facial features, hairstyle, hair color, eye shape, skin tone, body proportions, age impression, and overall personality.


Preserve the same recognizable character while adapting the styling, costume, and worldbuilding presentation to match this new concept board.

Do not redesign the character into a different person.

Do not significantly change the face shape, identity, or overall appearance.


Overall layout:

Design a polished two-page-style editorial concept sheet with a clean, elegant presentation.

The board should combine a fantasy or themed key visual poster on the left and a structured character design sheet on the right, with additional scene concept panels along the bottom.

The full image should feel like a professional art direction board, suitable for character development and visual worldbuilding.


Character impression:

The character should feel 【請修改|角色氣質,例如:gentle, adventurous, warm, dreamy, intelligent, elegant】.

The overall visual impression should remain appealing, cohesive, and faithful to the uploaded reference character.


Character outfit:

Design a 【請修改|服裝風格,例如:East Asian fantasy traveler / magical girl / cyberpunk explorer / modern urban fashion / medieval adventurer】 costume.

Use a color palette based on 【請修改|主要色系】.

Include 【請修改|服裝重點元素,例如:layered flowing fabrics, embroidered details, fantasy accessories, light armor-inspired accents, ribbons, jewelry, elegant boots】.

The outfit should feel 【請修改|服裝氣質,例如:magical, adventurous, graceful, futuristic, refined, playful】.


Left-side key visual:

Create a cinematic poster featuring the character in the foreground in a visually engaging pose.

Behind the character is 【請修改|世界觀場景,例如:a majestic floating-mountain oriental fantasy world with temples, bridges, waterfalls, clouds, and glowing atmosphere】.

Use beautiful cinematic lighting and strong atmosphere.

The key visual should feel immersive, polished, and emotionally appealing.


Add elegant Chinese typography:

- Large title: "【請修改|主標題】"

- Subtitle: "【請修改|副標題】"

- Small English line: "【請修改|英文標語】"


Also add a handwritten-style phrase on the right area:

"【請修改|品牌句 / 標語】"


At the lower right, add a small poetic line:

"【請修改|補充文案 / 小句子】"


Right-side character sheet:

Show the same character in full-body turnaround views:

- front view

- side view

- back view

- 3/4 view


Also include a heading:

"CHARACTER DESIGN"


Add the Chinese title:

"【請修改|角色名稱】 — 【請修改|角色定位】"


Add a short supporting line:

"【請修改|角色文案】"


Expression variation section:

Include three portrait expressions of the same character:

- cheerful / smiling

- calm / neutral

- surprised / curious


Details section:

Show close-up panels of:

- fabric or clothing detail

- waist accessory / symbolic prop / magical ornament

- shoes or boots design

- accessory detail


Color palette section:

Include 6 circular color swatches based on:

【請修改|色票 1】, 【請修改|色票 2】, 【請修改|色票 3】, 【請修改|色票 4】, 【請修改|色票 5】, 【請修改|色票 6】


Bottom scene concept section:

Add three cinematic environment thumbnails showing:

- 【請修改|場景概念 1】

- 【請修改|場景概念 2】

- 【請修改|場景概念 3】


Label this section:

"SCENE CONCEPT"


and the Chinese subtitle:

"場景概念"


Typography and graphic style:

Use a refined editorial design style with a premium concept-art presentation.

Mix elegant serif English labels with tasteful Chinese text.

Keep the layout clean, airy, balanced, and visually sophisticated.


Overall style:

high-end concept art, polished illustration, soft cinematic lighting, detailed costume design, elegant character sheet presentation, premium moodboard aesthetics, artbook-quality layout.

【小步原始版供參考】 Create a premium character design and worldbuilding concept board for an original character named "小小步 — 東方奇幻旅人".


Use the uploaded reference image as the primary source for the character's appearance.


Keep the character's identity consistent with the reference image, including:

facial features, hairstyle, hair color, eye shape, skin tone, body proportions, age impression, and overall personality.


Preserve the same recognizable character while adapting the styling, costume, and worldbuilding presentation to match this new concept board.

Do not redesign the character into a different person.

Do not significantly change the face shape, identity, or overall appearance.


Overall layout:

Design a polished two-page-style editorial concept sheet with a clean, elegant presentation.

The board should combine a fantasy key visual poster on the left and a structured character design sheet on the right, with additional scene concept panels along the bottom.

The full image should feel like a professional art direction board, suitable for character development and visual worldbuilding.


Character impression:

The character should feel gentle, approachable, warm, slightly dreamy, adventurous, and elegant.

The overall visual impression should remain appealing, cohesive, and faithful to the uploaded reference character.


Character outfit:

Design an East Asian fantasy traveler costume.

Use a color palette based on soft teal, muted blue-green, ivory, cream, and gold.


Include layered flowing fabrics, delicate embroidered details, fantasy accessories, light armor-inspired accents, a decorative waist ornament or magical pendant, and elegant boots.


The outfit should feel magical, adventurous, graceful, refined, and subtly inspired by oriental fantasy aesthetics.


Left-side key visual:

Create a cinematic fantasy poster featuring the character in the foreground, reaching one hand toward the viewer.


Behind the character is a majestic floating-mountain and cliffside oriental fantasy world with temples, bridges, waterfalls, clouds, and glowing atmosphere.


Use beautiful golden-hour lighting, airy mist, soft cinematic light, and an epic yet gentle fantasy atmosphere.


The key visual should feel immersive, polished, and emotionally appealing.


Add elegant Chinese typography:

- Large title: "小步奇境"

- Subtitle: "探索更多可能"

- Small English line: "A SMALL STEP A BIGGER WORLD"


Also add a handwritten-style phrase on the right area:

"用 AI,創造屬於你的精彩旅程"


At the lower right, add a small poetic line:

"從日常出發,走進更大的世界"


Right-side character sheet:

Show the same character in full-body turnaround views:

- front view

- side view

- back view

- 3/4 view


Also include a heading:

"CHARACTER DESIGN"


Add the Chinese title:

"小小步 — 東方奇幻旅人"


Add a short supporting line:

"熟悉的我,走進不一樣的世界"


Expression variation section:

Include three portrait expressions of the same character:

- cheerful / smiling

- calm / neutral

- surprised / curious


Details section:

Show close-up panels of:

- fabric embroidery

- waist accessory / magical ornament

- boots design

- accessory detail


Color palette section:

Include 6 circular color swatches based on:

ivory, pale teal, muted aqua, deep blue-green, warm brown, soft gold


Bottom scene concept section:

Add three cinematic environment thumbnails showing:

- floating fantasy mountains with temples surrounded by clouds

- a balcony or terrace overlooking the fantasy world

- an elegant pavilion or interior passage with sunlight streaming in


Label this section:

"SCENE CONCEPT"


and the Chinese subtitle:

"場景概念"


Typography and graphic style:

Use a refined editorial design style with a premium concept-art presentation.

Mix elegant serif English labels with tasteful Chinese text.

Keep the layout clean, airy, balanced, and visually sophisticated.


Overall style:

high-end fantasy concept art, polished illustration, soft cinematic lighting, airy clouds, detailed costume design, elegant character sheet presentation, premium moodboard aesthetics, artbook-quality layout.


但提示詞反而是五個案例裡相對複雜的一個。

我使用的是 時間軸提示詞

不是只寫一句:「生成電影級環繞運鏡。」

而是把不同時間點需要發生的事情拆開。

大致的運鏡設定是:

0–2 秒 畫面先看到遠處的小步。 一名黑衣劍客從鏡頭下方進入前景並站起來,逐漸遮住小步。

2–6 秒 攝影機靠近黑衣劍客的背部。以低角度向右側做半環繞,大約繞行 120~180 度。

前景黑衣人與遠處小步之間要產生明顯的視差。

6–8 秒 鏡頭繞過黑衣劍客肩膀之後,再次看見遠方的小步。

8–10 秒 鏡頭開始推近小步臉部。小步抬起眼睛,露出輕蔑、不屑,但相對克制的淡笑。


同時要求整支影片保持:

  • 連續鏡頭

  • 真實三維位移

  • 自然前後景視差

  • 小步的臉、紅棕短髮、服裝與身形保持一致

這也是為什麼這一題除了測運鏡之外,其實還同時在考驗 角色一致性


生成影片提示詞

東方奇幻電影場景,昏暗古殿內,暖色燈籠與冷色月光交錯。小步站在大殿遠處中央,保持紅棕色短髮、清秀五官、東方奇幻旅人服裝與臉部特徵完全一致。小步全程不戴眼鏡,雙眼與眉眼細節清楚可見。

0-2 秒:畫面先以小步為遠景主體。突然從鏡頭正下方,一名高大的黑衣劍客緩緩站起進入前景,他的頭頂、肩膀與背部快速佔據畫面中央,短暫遮擋小步,形成強烈前後景層次。

2-6 秒:攝影機立即貼近黑衣劍客的後背,以他為中心進行明顯的低角度半環繞運鏡,從人物背後向右側平順環繞約 120~180 度。鏡頭移動必須具有真實空間位移與明顯視差,古殿石柱、燈籠、石階與背景中的小步依序從畫面邊緣快速滑動,不是單純旋轉或放大畫面。

6-8 秒:當攝影機繞過黑衣劍客的肩膀後,小步重新出現在畫面中心。此時視角已從原本遠距正面改變為側前方低角度近景,她的髮絲、衣襬與飄帶受到微風吹動,神情冷靜而堅定。

8-10 秒:攝影機繼續環繞並略微向前推進,黑衣劍客逐漸滑向畫面側邊,小步的正面完整被揭露。最後鏡頭停在小步的電影主角 Hero Shot,她先抬眼直視鏡頭,接著嘴角微微勾起,露出一個輕蔑、帶點不屑、但非常克制的淡淡笑意,像是早已看穿對手,帶著強烈自信與壓制感。身後古殿、燈火、薄霧與光束完整展開。

整段為一個連續長鏡頭,不切鏡。強調下方入鏡、近距離前景遮擋、低角度環繞、前後景視差與最後人物揭露。攝影機必須真正穿越三維空間,不要只是平移、縮放或讓角色原地旋轉。

保持小步的臉部特徵、紅棕短髮、服裝與身體比例一致。

小步全程不佩戴眼鏡或墨鏡。

不要讓手腳或五官變形。

不要新增其他人物。

五個案例測完之後

一路把五個案例實際測完,可以很明顯看到:

真的沒有哪一個模型,是每一種影片都特別強。

如果今天製作的是商品廣告,我可能更在意材質、手部動作與商品一致性。

製作短劇時,重點會變成人物表演、角色一致性與故事感。

Reference 影片要看的是動作與節奏跟隨。

而電影運鏡又變成空間理解、角色穩定與鏡頭邏輯。

所以與其一直問:「哪一個 AI 影片模型最強?」

更實際的做法應該是先想清楚:「我現在要做的是什麼影片?」

再去判斷這個任務最需要的是:

角色一致性、Reference 跟隨、畫面細節,還是運鏡能力。

最後再選擇比較適合的模型。

這也是我覺得 ChatArt 這類把多個影片模型整合在同一個平台裡,比較方便的地方。

同一份素材可以直接切換不同模型測試,不需要每次重新換平台。

而且模型選對了,也能少掉很多「明明提示詞沒有問題,卻一直重新生成」所浪費的 Token。



 
 
 

留言


正在確認文章閱讀權限……

登入會員後即可閱讀

「AI 影片模型怎麼選?Seedance 2.5、MiniMax H3、Wan 3.0 五種情境實測」為小步學習會員限定內容,免費加入會員並登入後即可閱讀。

​這篇文章你會看到...

如果感覺這篇文章對您有幫助,可以請小步喝咖啡喔~

謝謝您支持小步繼續用心創作。☕ 請小步喝杯咖啡

  • Line
bottom of page