Every image to video generation starts with a choice most people never make consciously. You have an image. You feed it to the engine. But where you feed it changes everything about what comes out the other side.

We ran controlled tests in our lab on Seedance 2.5 to isolate exactly what happens when the same image enters as a start frame versus a reference image. The results overturned how we brief every shot at KURACONV, and they will probably overturn yours too.

What a start frame actually does in image to video

The start frame anchors the frame. Not the subject. The frame.

That distinction sounds academic until you watch twenty generations side by side. When an image enters as a start frame, the engine treats it as the world: the light, the composition, the color temperature, the depth of the space. Frame one of your video will be that image, and the engine will honor the atmosphere you gave it.

What the engine will not necessarily honor is your subject. In our tests, the camera routinely drifted away from the hero object, flew over it, or let it deform as motion progressed. The start frame is a contract about where the story begins, not about what the story protects.

So if your priority is world consistency, a specific interior, a precise lighting setup, a composition you fought for in pre-production, the start frame is your instrument. The engine will open on your world and move through it.

What a reference image does in image to video AI

The reference image anchors the object.

Feed the same product shot as a reference instead of a start frame and the behavior inverts. The engine no longer promises to open on your composition. Instead, it carries the identity of the thing across the whole shot: the shape of the bottle, the face of the character, the texture of the fabric. The reference is a contract about what must survive the motion.

In our controlled test, the reference held product identity through camera moves that would have melted a start-frame-only generation. The label stayed the label. The character stayed the character. But the world around it was the engine's invention, loosely guided by prompt, not locked by pixels.

Start frame locks the frame and risks the subject; reference locks the subject and releases the frame.

The engine flies over the subject anyway. Direction corrects that.

Here is the uncomfortable finding: even with a clean start frame, a strong reference, and a disciplined prompt, the engine tends to fly over the subject. It loves camera movement more than it loves your hero. Left alone, it will orbit, crane, and drift, treating your product like scenery.

This is not a bug you prompt away with one magic word. It is a tendency you correct with direction. In practice that means:

  • Naming the camera behavior explicitly, and naming what the camera must not do.
  • Keeping the subject's role in the motion active: the subject turns, the light shifts on it, something happens to it, so the engine has a reason to stay.
  • Reviewing generations as a director reviews dailies, not as a user refreshing a slot machine. When the camera abandons the subject, the fix is usually in the brief, not in the seed.

AI video engines are camera operators with wanderlust. Your job is to be the director standing next to them.

Using both together: the working setup

The strongest results in our lab came from refusing the either-or. Use both, and assign each one its job.

Start frame: the opening world. Build it deliberately, in an image model, with the light and composition you want frame one to have. If the brand needs to be present from the first frame, it lives here.

Reference images: the identities that must survive. Product, character, key prop, and yes, on-screen text when typography matters. Seedance 2.5 accepts multiple references, and each one is a leash on a different element.

Prompt: the motion and the intention. Not a description of the product, the references already carry that. The prompt directs what happens.

When we brief shots this way, the failure rate drops from "generate ten, pray for one" to something that looks like a production schedule.

The decision in one question

Before you generate, ask: what cannot change in this shot?

If the answer is the world, the light, the composition, that image is your start frame. If the answer is the object, the identity, the face, that image is your reference. If the answer is both, and in client work it is almost always both, then build both, feed both, and direct the space between them.

This is the Sentimagem logic we work by, direction over generation: test what the engine actually does instead of trusting the docs, feel where the shot loses its subject, and lock in the setup that holds.

Image to video AI rewards people who test like a lab and decide like a director. The tools are astonishing. They are also indifferent to your brand until you tell them, in the right channel, what matters.

In image to video AI, the start frame anchors the world and the reference image anchors the subject; strong shots feed both and direct the space between them.