Landscape to Vertical, With the Crop Following the Speaker
It runs a face detector four times a second and keeps the biggest box. Knowing that tells you which shots pan like a camera operator and which get mangled.
Premiere calls it Auto Reframe, Resolve Studio calls it Smart Reframe, Final Cut calls it Smart Conform. All three present it as the same gesture: tick a box, and the software finds your subject. The browser version on this site does the same job, and because its workings are readable, it is worth being blunt about them. There is no subject-finding. There is a face detector run four times a second, a smoothing pass with three constants in it, and a single FFmpeg crop expression.
That is not trivia. The mechanism is a prediction tool: once you know it, you know which shots will come out looking like a camera operator followed them and which will come out mangled — before you spend the render.
Your four levers
Take the inventory first, because it is short. Auto Reframe gives you three shape buttons, one "Follow the person in shot" checkbox, and a run button. That is the whole interface. There is no dead-zone slider, no speed control, no manual keyframing, no way to nudge the path, no way to say which face wins, and no preview of the crop before the render — you watch the finished file to find out what it did.
So everything the rest of this page diagnoses resolves to one of four moves:
- Change the target shape. 9:16, 4:5 or 1:1. A wider window forgives a lot.
- Turn follow off. A locked centre crop, on purpose.
- Split the clip first, reframe the pieces separately, rejoin.
- Place the window by hand in Video Crop, which takes four numbers and puts the crop exactly where you say.
Each section below points at which of the four applies.
It finds the biggest face, not the speaker
Auto Reframe samples frames at 4 fps and asks MediaPipe's BlazeFace detector where the faces are. When more than one comes back, it takes the largest box and calls that the subject. It never touches the audio. Nothing in the pipeline knows who is talking, who matters, or that the person on the left just made the point.
For a single presenter that shorthand holds — the only face is the speaker. For an interview it does not. The largest face is normally whoever is nearest the lens, so the crop commits to one person and re-commits every time the other leans forward. Because the path is speed-limited, each switch is a slow slide across the frame rather than a cut, so a lively two-hander reads as a camera drifting between people for no reason. That is the single most common way this ends badly, and lever one — 4:5 instead of 9:16 — fixes more of it than anything else.
It is the same detector Face Blur loads, pointed at a different question: Face Blur samples 30 times a second and acts on every face in the frame, Auto Reframe samples four times a second and throws away all but the biggest. Two limits come with that. Detection runs on a copy of the frame scaled down to at most 640 px wide, so a face filling a tenth of a 1920-wide frame is about 64 pixels across by the time the model sees it — wide masters and stage footage are where detection thins out. And the short-range model variant is the one that loads, so a subject at the far end of a room is already at the edge of what it was built for.
The confidence floor is set low, at 0.35. That buys two good behaviours: marginal detections are kept rather than dropped, and a frame with no detection at all carries the previous position forward instead of snapping back to centre, so a head turning away for a second does not send the crop on a round trip. It also has a cost, which is the first entry in the failure map below.
It lags on purpose, and it also leads
Three things happen to the raw path, and a fourth to the result.
An eleven-sample centred moving average — five each side of the current point, about two and a half seconds of clip at four samples a second. Then a 0.06 dead zone: the subject can wander within six percent of the frame width and the crop does not move at all, which is what stops a talking head making the frame breathe. Then a speed cap of 0.35 frame-widths per second, so on a 1920-wide source the window travels at most 672 px/s. It never has to cross the whole frame: its travel is clamped to what the crop leaves over, which at 9:16 from 1920 is 1312 px, so worst case the window is roughly two seconds behind a subject who walks edge to edge. Finally the smoothed path is reduced to at most 40 keyframes and written as one nested piecewise-linear expression for crop's x, which ramps between keys instead of stepping. That is the whole reason it glides.
Because the average is centred rather than trailing, it reads samples from the future as well as the past, so the crop eases toward where the subject is going shortly before they arrive. Camera operators do the same thing. It also means the tool cannot work live — and it means a hard cut inside your clip is poison for the path. Feed it an edit that cuts from a speaker on the left to a speaker on the right and the crop slides across the join over two or three seconds instead of jumping with it.
Lever three, then: cut first, reframe second. Split at the edits with Video Splitter, reframe each piece, and rejoin with Video Merger — which re-encodes on concat rather than stream-copying, deliberately, because -c copy on the concat demuxer produces broken output. Budget for that extra generation of H.264.
The 40-keyframe ceiling is the other thing to know before you point this at long-form. The thinning runs in two stages: first it drops any point a straight line between its neighbours would already have hit, then, only if more than 40 survive that, it keeps 40 of them spread evenly through the surviving list — evenly by position in the list, not evenly in time. Keyframes cluster wherever the path bends, so busy passages keep their share of the detail and quiet ones lose what little they had. Whether the cap binds at all depends on how much the smoothed path actually moves: a ten-minute lecture from behind a lectern may never reach it, while ten minutes of someone working a room will, and there the pan degrades into a drift with only a loose relation to what anyone is doing. Clips cut out of long-form are what this is for. Whole episodes are not.
The output is rarely 1080×1920
The tool refuses to upscale. It cuts the tallest window of the target shape that fits the source, and it leaves that window at its own size only when the crop already equals the output size exactly. Otherwise a lanczos resample runs — a genuine downscale when the crop is wider than the 1080-wide cap, and a one- or two-pixel rounding correction when it is not, because the even-number rounding on width and the ratio rounding on height rarely agree.
Three real sources, all at 9:16:
- 1920×1080 → crop 608×1080, output 608×1080. No scaling at all.
- 1280×720 → crop 404×720, output 404×718. Lanczos runs, for two pixels of height.
- 3840×2160 → crop 1214×2160, output 1080×1920. A real downscale, and the only common case that lands on the canonical size.
608×1080 is not a broken export. It is every pixel that was recorded of that part of the frame. But "not broken" is not "as sharp as native vertical", and this is the part the reassurance usually skips: the platform displays at 1080 wide, so your file is being upscaled about 1.8× at playback. It shows first on small on-screen text — lower thirds, UI captions, scoreboard digits that were perfectly legible in the 16:9 master are now being drawn from 56 percent of the pixels. From a 1280×720 master the crop is 404 px wide, 37 percent, and it looks it. That is a real argument for sourcing or shooting 4K when the vertical cut is the deliverable.
It is not an argument for upscaling before upload. Video Resizer will hand you a fixed 1080×1920 if a template or a downstream tool demands that container, but the pixels it adds are invented and the platform re-encodes your upload anyway. A fixed 1080×1920 is a container requirement, not a quality one. If you want the full delivery sheet — bitrate, codec, frame rate for TikTok, Reels and Shorts — the export settings guide for football clips has it; read its "1080×1920, do not go higher" line as what the platform wants to receive, not as a verdict on a 608×1080 file that the platform will happily accept.
The video is re-encoded at CRF 21 on the veryfast preset. The audio is stream-copied untouched, which is free quality with one trap attached: the copy goes straight into an MP4 container, so an audio codec MP4 will not carry — PCM from some QuickTime screen recordings, Vorbis from an older WebM — kills the whole render. Worse, the error you get blames the video decoder, so you will chase the wrong thing. If a render dies on a .webm or a .mov and the video codec looks unremarkable, run the file through Video Converter first and try again.
It pans; it never tilts
The crop is taken at y=0 across the full source height, so vertical composition is inherited wholesale from the landscape framing: generous headroom stays generous headroom, a subject sitting low in frame stays low. For 16:9 to 9:16 that is exactly correct, because the window is already as tall as the source and there is nothing to choose.
Where it bites is the case nobody expects. If the target shape is wider than the source — squaring an already-vertical clip, say — the crop has to lose height instead of width, and y=0 means it keeps the top. A 1080×1920 portrait clip cropped to 1:1 gives you the top 1080 pixels, not the middle. For a talking head that is usually lucky. For a full-body shot it removes the legs, and nothing in the interface will move that window down.
Lever four. In Video Crop you type X, Y, width and height against the source dimensions — no drag handles, no presets, but it goes exactly where you put it. For the common job, 1920×1080 to 9:16: W 608, H 1080, Y 0, and X is yours, from 0 at hard left to 1312 at hard right, 656 for dead centre. Same source to 4:5: W 864, H 1080, Y 0, X between 0 and 1056. And for the portrait-to-square case above: W 1080, H 1080, X 0, and Y is precisely the choice the automatic tool took away from you.
A centred crop is usually not a failure
When no face is found in any sampled frame, the tool centres the crop and says so in plain text rather than passing a fixed crop off as a track. That is the correct result, not an error, and re-running will not change it. Turning follow off gives you the same locked centre crop deliberately, which is the right setting for screen recordings, product shots, landscapes and anything else with no people in it.
The other message it shows is less trustworthy. If the source is already as tall as or taller than the shape you have selected, the tool warns that there is nothing to crop away. For 9:16 that is true: a 1080×1920 source gives a 1080-wide window on a 1080-wide frame and the crop has nowhere to travel. For 1:1 and 4:5 it is false. The check only compares aspect ratios, so it fires on that same portrait source while the crop still takes 1080×1080 or 1080×1350 off the top and throws the bottom away — the y=0 case from the section above, announced as a no-op. Trust that warning when the target is 9:16. Ignore it for the other two shapes and go to Video Crop.
What each target shape actually costs
From a 1920-wide source: 9:16 keeps 608 px of width, about 32 percent of the frame. 4:5 keeps 864 px, 45 percent. 1:1 keeps the full 1080, 56 percent.
The gap between 608 and 864 is where two-person footage lives. A 4:5 window is 42 percent wider than a 9:16 one, which is frequently the difference between both heads surviving and the crop having to pick a favourite. And the tracking result is cached — after the first scan the button changes to "Apply this shape" and only re-renders — so trying all three costs one scan and three quick encodes. Do that before assuming 9:16 is the answer.
Split Screen is the other reflex for two speakers, and it is worth knowing its shape before you reach for it: it takes two separate files, stacks them horizontally or vertically, and the output takes the first clip's dimensions. Each input is fitted inside its half and padded rather than stretched, so unless each clip already matches its half's aspect you get letterboxing inside the panels. It will not hand you a filled 1080×1920 two-up. For that you want a desktop NLE.
The rest of the failure map
Faces that are not your subject. This is the price of a 0.35 confidence floor plus "biggest box wins": anything face-shaped counts. A framed portrait on the wall, a face on a second monitor or inside a screen share, a poster, a crew member or an audience head nearer the lens than the speaker — if its box is bigger, it becomes the subject, and the crop can park on a photograph for the length of the clip. For desk, webinar and call recordings this is likelier than any of the motion failures below. Fix it by cropping the offender out first in Video Crop, or by turning follow off and placing the window yourself.
Whip pans and fast handheld moves. The 0.35/s speed cap cannot keep up, by design. The subject leaves frame and the crop arrives late.
Burned-in graphics. Scoreboards, lower thirds, logos and captions sit near the edges of a 16:9 frame, and a 608 px window amputates them. Reframe first and add them afterwards to the vertical cut — Scoreboard Overlay and Video Text work on whatever you feed them, and captions from Auto Captions placed on the vertical file will actually be inside it.
Subjects turned away, in profile, masked or backlit. Detections stop, the last position holds, and the crop parks. Better than thrashing, but it is not tracking any more.
AV1 input. It will scan happily, because your browser can decode it, and then fail at the render: the in-browser FFmpeg build has no AV1 decoder at all, and nothing else on this site can convert it either. That one genuinely is "re-export it as H.264 from wherever it came from, or do this on a desktop."
Before you start the scan
Scanning is not a fast-forward through a decoder. The clip is played back at 4× speed with frames sampled as they are actually presented, and three practical things follow from that.
It takes real time. A six-minute clip needs at least ninety seconds before the encode even begins. And there is no stop button — Reset is disabled while the tool is busy, so once a scan starts on a twenty-minute file your only exit is reloading the tab. Check you picked the right file, at the right length, before you click.
Leave the tab in front while it runs. The sampler only fires when a frame is genuinely presented, so backgrounding the tab during a ninety-second-plus wait is not the free move it looks like. A browser that blocks autoplay outright will fail the scan and say so.
And on the first run in a given browser, the vision runtime and the face model are fetched from public CDNs, so the tracking step needs a network round-trip once. Your video never leaves the tab; the model has to arrive in it.
After that, it is four decisions: which shape, follow or not, split or not, by hand or not. The mechanism only matters because it tells you which one to pick.
Try it yourself — free in your browser
No upload, no signup, no watermark — these tools run on FFmpeg WebAssembly locally.