There is a rule everyone in AI video repeats like scripture: text does not survive generation. Screens turn to alphabet soup, signage melts into hieroglyphs, and any word longer than "on" becomes a typographic crime scene. So the industry workaround became gospel: never generate text, always composite it in post.
We ran a controlled test in our lab and overturned that rule. Not with a trick prompt. With a change of channel.
Why AI generates wrong text in the first place
Video engines do not write. They paint. When you type "a phone screen displaying the word SUNRISE in bold sans serif," the engine does not typeset the word. It hallucinates the visual statistics of text: strokes, spacing, the rhythm of letterforms. Sometimes it lands close. Usually it produces something that looks like a word until you read it.
Diacritics make it worse. Accented characters, cedillas, tildes, anything outside the plainest ASCII, and the failure rate climbs toward certainty. If your work crosses languages, and ours does, São Paulo alone is a stress test, prompting text is a lottery you lose.
The mistake is not the engine's. The mistake is the channel. A prompt is a description, and typography is not a thing you describe. It is a thing you show.
The reference method for AI video text
Here is what we tested on Seedance 2.5. Instead of describing the text, we built it.
Step one: create the finished text as an image. Real typography, real kerning, real diacritics, set in an image tool or an image model that handles text well, placed in context: lettering on a screen inside the scene, a label on a product, a title card on a monitor.
Step two: feed that image to the video engine as a reference, alongside the start frame that defines the world.
Step three: prompt the motion, not the text. The prompt never mentions what the words say. The reference already carries them.
The result in our controlled runs: the typography survived. The accented characters survived. The engine treated the lettering the way it treats a product or a face fed as reference: as an identity to preserve, not a texture to reinvent. In one test, a reference carrying accented text held the accent through the entire generated shot, something no prompt phrasing had achieved in dozens of attempts.
A reference image carries typography better than any prompt ever will. Text can go through image to video. It just cannot go through the prompt.
Reference anchors the object, and text is an object
This works because of a deeper mechanic we mapped in separate tests. In image to video, a reference image anchors the object: its identity, its details, its surface truth. The start frame anchors the frame: world, light, composition. The prompt directs motion and intent.
Once you see text as an object rather than as language, the method is obvious. A word set in type is a shape with an identity, exactly like a bottle or a logo. Feed it through the channel that protects identity. The engine stops trying to write and starts trying to preserve, and preservation is what engines are actually good at.
Where this changes production
For a studio, this is not a party trick. It restructures the pipeline.
Interface shots. Phone screens, dashboards, kiosks inside a scene: build the screen as a finished still, reference it, and generate the scene around it in motion. No more post-compositing a flat screen onto a moving perspective.
Signage and environments. A storefront name, a poster on a wall, a neon sign: set the type in the still, let the reference hold it while the camera moves.
Multilingual work. Diacritics stop being a risk category. The accent lives in the reference pixels, and pixels do not misspell.
Title moments inside the fiction. When a word needs to exist inside the world of the film rather than on top of it, the reference method is currently the only reliable road.
What it does not replace: supers, subtitles, and legal text still belong in the edit, where you control them frame by frame. The reference method is for text that lives inside the scene, not text that lives on top of the video.
The discipline that makes it hold
Three direction notes from our lab sessions, because the method fails when handled casually.
Finish the type before you generate. The reference must carry the final lettering, final weight, final color. The engine preserves what it sees; it will not improve a sloppy reference.
Keep the text large enough in frame. Tiny type in a wide shot degrades. If the word matters, give it real estate or plan a closer shot for it.
Do not restate the text in the prompt. Describing the words invites the engine to repaint them. Let the reference speak alone and spend your prompt on camera and motion.
This is the Sentimagem method applied to a stubborn problem, direction over generation: test what the engine does with text instead of trusting the received rule, locate where the failure actually lives, in the channel, not the capability, and turn the setup that works into a standard.
Typography survives AI video generation when it enters as a reference image carrying finished lettering, because a reference is a channel of preservation while a prompt is only a description.
