Left is the raw upload. Right is what came back. Nothing was re-shot, re-cut or re-lit — only captions were added. Drag the line and decide for yourself whether that is "subtitles".
Every decision below is made per clip, not per template. Two uploads of the same length can come back completely differently.
The transcript is read before anything is drawn. One word per statement gets the weight; the rest stays flowing text. Filler and auxiliary verbs never become the big word.
Not a fixed band. The frame is measured: face, movement, platform UI. Text dodges the subject, stays inside the safe area, and holds its place instead of jumping every second.
26 animations, picked by meaning. Say "behind me" and the text goes behind you — the person is cut out per frame, so it really passes behind the shoulder.
Push-ins and zooms sit on the strongest moment of the sentence, on the beat when there is music. Sound design lands on the word onset, not on a grid.
3 free minutes every month. No card, no subscription. If the result does not beat what you have now, you have lost four minutes.