5 Tips to Create Better AI-Generated Videos

Author

Anisa K.

Published on

August 12, 2026

AI tools can be used to instantly create infographic posters, various types of images, and even videos simply by writing the right prompts. Creating videos with AI is relatively easy, as long as you understand the right prompting techniques.

However, you should know that every AI assistant or AI video generator has its own characteristics. That’s why it’s important to understand which prompting techniques work best for each tool.

For example, the prompt structure you use to create a video with Seedance may differ from what you use with Google Invideo, and the same goes for Gemini.

Want to learn some essential tips for creating videos with AI? Read on to find out!

1. Think About the Sequence of Scenes

Writing prompts for video generation is quite different from writing prompts for images, photos, or text. Even for a short video, you have to think about the different scenes you want to include. The same applies when adding audio to support the story.

That’s why it’s important to plan the storyline, scene sequence, video context, and estimated duration of each scene before generating the video. Why does this matter? Some AI tools divide videos into separate scenes to make the editing process easier for users. For example, VEED.IO allows users to add, remove, or replace video clips for free, although the number of edits is limited.

The same principle applies when using Google VEO 3.1, where you need to consider the scene sequence, storyline, video background or location, objects, audio, and many other details.

Therefore, it’s a good idea to outline your scenes and video flow before writing a prompt for AI video generation.

2. Use Layered Prompts

Just like using AI as a research tool often requires layered prompts, the same approach can also be useful for AI-generated videos.

Creating videos with AI requires several key elements in your prompt, including:

  • Who is the subject, character, or talent in the video?
  • What movements or actions should the subject or character perform?
  • What should the mood, time of day, and background look like?
  • What visual style do you want? Realistic, cartoon, dreamy, cinematic, or something else?
  • What text or dialogue should the subject or character say?
  • What speaking style or accent should be used in the video? Should it sound casual, formal, conversational, or use a specific dialect?

Unlike image generation, you can create a layered prompt by describing these elements in complete sentences. You don’t have to use bullet points or a long list, although this approach can also be effective.

Many people use VEO 3.1 to create videos with detailed narrative prompts that describe the scene in depth, even if the prompts may seem quite long.

3. Specify the Video Duration

Some AI tools recommend video durations based on what users need, such as 8 seconds, 15 seconds, 30 seconds, 60 seconds, or even longer. However, some AI tools do not let users specify the video duration directly. In these cases, you can specify and mention the video duration in your prompt.

For example, when using Gemini, you can add instructions to your prompt such as:

Create a 10-second video to keep the output consistent or keep the video to a maximum of 15 seconds so it can be uploaded to a website.

Meanwhile, some other AI tools may not be able to follow specific duration instructions. However, you can work around this by removing unwanted clips that make the final video less effective.

For example, when using VEED.IO, you may not always be able to control the video duration through a prompt. However, you can shorten or extend the video by adding clips and text, as well as manually setting the duration limit through the Settings menu.

4. Use English When Writing Prompts for AI Videos

Using English when writing prompts for AI video generation is highly recommended. This is because AI models are primarily trained on global datasets, especially those used by tools like Gemini, Sora, and VEO 3.1.

So, don’t be surprised if AI-generated videos created with English prompts look more cinematic, include more technical details, and produce results that better match what users have in mind. AI can also interpret English prompts more easily because they often contain common cinematography terms used in filmmaking.

Here are some reasons why English is more recommended when writing prompts for AI video generation:

4.1. English Keywords Work Better with Gemini

It’s worth that Gemini tends to work better with English keywords and prompts, although the AI tool can also generate videos using Indonesian prompts with good results. This is because AI models like Gemini are trained on data from many languages, but English makes up a much larger share of digital content.

English is also more closely associated with terminology commonly used in video production, while Indonesian prompts can sometimes be more ambiguous.

To see the difference, take a look at the following comparison:

  • Example of generating a video in Gemini using an Indonesian prompt

Buat video promosi produk matcha berdurasi 10 detik dengan gaya iklan minuman premium yang realistis dan menarik.

Tampilkan segelas matcha latte berwarna hijau cerah dengan tekstur lembut di atas meja kayu. Pada awal video, kamera melakukan gerakan mendekat secara perlahan ke arah gelas. Bubuk matcha terlihat ditaburkan secara perlahan di atas permukaan minuman. Kemudian, kamera bergerak sedikit ke samping untuk memperlihatkan gelas dari sudut yang lebih menarik. Tambahkan uap tipis dan gerakan cahaya alami agar video terasa hidup dan realistis.

Gunakan pencahayaan alami yang lembut, depth of field dangkal, dan latar belakang yang estetik dengan nuansa hijau dan krem. Pertahankan tampilan matcha yang realistis, tekstur minuman yang natural, dan gerakan kamera yang halus. Jangan mengubah bentuk gelas atau warna utama matcha.

Tambahkan voice-over perempuan dengan suara ramah, hangat, dan persuasif dalam bahasa Indonesia. Ucapkan kalimat promosi berikut dengan jelas dan natural:

“Rasakan creamy-nya matcha premium dalam setiap tegukan. Saatnya nikmati matcha favoritmu!”

Here’s the video generated by AI using the prompt above!

This is the example of an English Prompt for Creating a Video in Gemini:

Create a 10-second promotional video for a matcha brand with a realistic and visually appealing premium beverage advertising style. Show a glass of bright green matcha latte with a smooth texture placed on a wooden table. At the beginning of the video, slowly push the camera toward the glass. Matcha powder is gently sprinkled over the surface of the drink. Then, slightly move the camera sideways to reveal the glass from a more attractive angle. Add subtle steam and natural light movement to make the scene feel alive and realistic.

Use soft natural lighting, shallow depth of field, and an aesthetic green and cream background. Preserve the realistic appearance of the matcha, natural beverage texture, and smooth camera movement. Do not change the shape of the glass or the main color of the matcha.

Add a female voice-over with a warm, friendly, and persuasive tone in Indonesian. Clearly and naturally say the following promotional line:

“Rasakan creamy-nya matcha premium dalam setiap tegukan. Saatnya nikmati matcha favoritmu!”

Add soft background music at a low volume so it does not interfere with the voice-over. Make the entire video feel like a premium matcha product advertisement for social media.

The difference is clear: the result looks much more professional!

Videos created with English prompts actually look more professional. You can see this in the choice of glassware, which resembles what you would typically find in a coffee shop, as well as the camera work. The video does not rely solely on zoom-in and zoom-out shots, but also uses camera movements that feel more intentional and professional.

However, both videos above still need some optimization to avoid visual inconsistencies, where the visuals change between frames or shots and make it obvious to the audience that the video was AI-generated.

The same applies when writing English prompts for creating videos with Google VEO 3.1. Many users who create videos with VEO 3.1 translate their Indonesian prompts into English first to get results that are more relevant, accurate, and aligned with their expectations.

This is because VEO 3.1 can better understand instructions, context, and prompting styles written in English, resulting in video visualizations that more closely match what users have in mind.

Here are some English terms commonly used in video production that AI tools can easily understand:

  • Cinematic lighting
  • Hyperrealistic 4K
  • Documentary style
  • Anime aesthetic
  • Volumetric light
  • Soft focus
  • High detail

You can add these cinematography terms when writing prompts for AI video generation, especially when using VEO 3.1 or Sora.

4.2. VEO 3.1 May Render AI Videos Faster with English Prompts

VEO 3.1 may render videos faster when users submit prompts in English. Why does this happen? Since the AI model is trained on datasets dominated by English-language content, it may be able to recognize and process the prompt structure more efficiently, potentially improving its performance when generating videos.

This can be different when you write prompts in Indonesian, as the AI may take longer to process the language and interpret your instructions. Indonesian has a wide range of vocabulary and expressions, which can make prompts more difficult to interpret, especially when the meaning is ambiguous.

That’s why many VEO 3.1 users write their video scripts in Indonesian first and then translate them into English. This can help make the resulting AI-generated video more accurate and closer to what they have in mind.

If you want to translate your Indonesian video prompts into English, you can use:

  • The Google VEO 3.1 chat feature directly
  • ChatGPT by asking it to create a script or description in English
  • Google Translate

4.3. Try Another AI Tool to Compare the Video Results

Besides Gemini, you can also try other AI tools that can generate videos to compare the results, such as Pippit AI. This AI tool provides a fairly generous number of credits, allowing you to create several videos for free. You can also use it to create product ads, supporting content, and more.

For example, you can use a prompt like this:

“Create a 10-second promotional video for a matcha product in a realistic and engaging premium beverage advertising style. Show a glass of bright green matcha latte with subtle latte art on a wooden table. At the beginning of the video, slowly move the camera closer to the glass. Place a pouch of matcha powder next to the glass. Then, move the camera slightly to the side to show the glass from a more appealing angle.

Use soft, natural lighting, a shallow depth of field, and an aesthetic background with green and cream tones. Keep the matcha looking realistic, with natural beverage texture and smooth camera movements. Do not change the shape of the glass or the matcha’s primary color.

Add a female voice-over with a friendly, warm, and persuasive tone in Indonesian. Clearly and naturally say the following promotional line:

“Rasakan creamy-nya matcha premium dalam setiap tegukan dan nikmati matcha favoritmu!”

 

5. Keep Experimenting

Just like when you’re trying to get relevant and accurate information from AI, you also need to experiment with different prompts to create AI-generated videos that meet your expectations. This process is all about trial and error. Some attempts may not turn out as expected or may not be perfect, and you can see the results directly in the AI-generated video. That’s why it’s important to keep practicing and refining your AI prompts to get better results.

For example, let’s say you’re creating an AI-generated video using Runway with the following prompt:

A young woman barista in a cozy coffee shop, wearing an apron, carefully preparing a matcha latte. She scoops vibrant green matcha powder into a bowl, adds hot water, and whisks it gently using a traditional bamboo whisk. The camera captures close-up shots of the foam forming, the latte art being poured into a ceramic cup. Warm ambient lighting, cinematic style, soft depth of field, calm and peaceful atmosphere.

 

After using the prompt above, you may notice several issues in the generated video, such as:

  • The barista or woman in the video does not appear to be actually looking at the matcha bowl while stirring the matcha.
  • The bamboo whisk seems to move in and out of the bowl as if it is passing through it.
  • Steam appears to be coming from the bowl, suggesting that the water is extremely hot. This is unusual when preparing matcha, which is typically whisked with water at around 80°C.
  • The fingers on her left hand do not look quite right while she is holding the bowl.
  • There are no tools or ingredients for making matcha latte around the bowl or on the table.

Based on these issues, you need to refine the prompt to make the video more consistent with your expectations and create a more realistic-looking activity. In other words, the next step is to improve the prompt. For example:

Halo, please make a video based on this prompt:
A young female barista in a cozy coffee shop, wearing an apron, is carefully whisking matcha using a traditional bamboo whisk. She is fully focused and attentive as she looks at the matcha bowl. There is no visible steam rising from the bowl, as the matcha is prepared using water at a temperature of 70°-80°C. The bamboo whisk does not come out of the bowl, and use the correct matcha stirring techniqueOn the table where she is whisking the matcha, there is a carton of fresh milk, a pouch of matcha powder, a kettle filled with hot water, and an aesthetic glass. The scene looks highly realistic and vivid. Show the barista whisking the matcha naturally. Warm ambient lighting, cinematic style, soft depth of field, and a calm, peaceful atmosphere. Ensure the barista character is complete and flawless, with no missing or distorted fingers, hands, or body parts.”

What Does the AI-Generated Video Look Like After Refining the Prompt Above?

It turns out that the person in the video has also changed. However, if you look closely, Runway has fixed several issues, including:

  • The whisk is now positioned inside the bowl, although the way the bamboo whisk is held is still incorrect.
  • The bamboo whisk now more closely resembles the type of whisk commonly used to prepare matcha.
  • The steam effect from the hot water poured into the bowl of matcha has disappeared.
  • The matcha-making ingredients and tools are placed on the table, making the scene feel more natural and lively.

Despite these improvements, there are still a few minor imperfections, particularly with:

  • The barista’s gaze initially appears focused, but shifts forward a second later, making her look as if she is daydreaming.
  • An odd liquify or wavy distortion appears as the bamboo whisk approaches the rim of the bowl.
  • The fingers on her right hand appear somewhat pointed, making them look different from the fingers on her left hand.

Since the result is still not 100% perfect, you’ll need to experiment repeatedly to get the video to match your expectations.

These are some important tips for creating videos with AI, which often require a few tricks of their own. Be sure to keep experimenting and testing through trial and error to create a video that accurately reflects your prompt and meets your expectations.

Updated on August 31, 2026 by Anisa K.

Mari Bekerja Sama!

Wujudkan situs web impian Anda bersama kami.

Contact Us