VFX & MotionAdvanced

You Blurred the Faces. Here's What Still Identifies Them.

Face Blur reports success on the frames it missed. What the detector covers, where the mask lands off-centre, and the identifiers no blur ever touches.

Applicable Software:Premiere ProDaVinci Resolve

Face Blur has no failure state for the thing it gets wrong. When the detector finds nothing in a frame, it does not warn you, log a gap, or slow down. It draws the frame and moves on. The progress bar reaches 100 per cent either way, the preview looks right, and the file you download is the file you send.

So start from the export rather than the tool. Three minutes, phone, 4K, handheld: one person talking to camera, two colleagues crossing a warehouse aisle behind her. Grey ellipses ride on three heads in the preview and the run completes clean. The file is not anonymous, and nothing in the interface will ever tell you so. Face Blur answered a smaller question than the one you asked — "does a face detector find a roughly frontal face in this single image?" You were asking "can anyone who already knows these people recognise them from anything in this file?" The rest of this page is the distance between those two questions.

The same detector, wired two opposite ways

Auto Reframe and Face Blur load the identical model: MediaPipe's blaze_face_short_range, the same WASM build, the same GPU delegate, the same VIDEO running mode. Everything wrapped around it is opposite, and that difference is the whole story.

Auto Reframe samples four times a second on a 640-pixel-wide downscale, accepts detections at 0.35 confidence, keeps only the largest face in frame, carries the last known position through a gap, and smooths the resulting path. It is built to survive a miss.

Face Blur runs the detector once per animation frame at full output resolution, accepts at 0.4, blurs every face it finds, and keeps no state at all between frames. Frame N's mask comes entirely from frame N's detections. When the list comes back empty, the frame is left exactly as it was first drawn: original, sharp, published. One tool forgives a miss. This one ships it.

Two things about that loop. "Once per animation frame" is not once per source frame. The loop is paced by requestAnimationFrame, so on a 60 Hz display over 30 fps source it detects roughly twice per source frame, on a 120 Hz panel roughly four times, and on a loaded machine possibly less than once. The cadence floats with whatever else the browser is doing — which is the practical reason you cannot reason your way to where the dropouts are and have to go look. And the recorder is hardcoded to canvas.captureStream(30). Whatever you feed it, the export is 30 fps: a 60 fps phone clip loses half its frames, 24 fps footage gets resampled. Worth knowing before you plan a delivery, and it is what makes the frame arithmetic below valid.

Every failure mode follows from the no-memory design. Your subject turns towards profile and confidence slides until it crosses 0.4 — where exactly depends on the face, the light and the framing. She passes a loading door and backlights into a silhouette. She turns to point and the pan smears her face across two frames. A colleague walks close to camera facing away, and the back of a head is not a face at any threshold. The two in the aisle are simply too small: the detector works on a fixed-size input, so what counts is a face's size as a fraction of the frame, not how many pixels it occupies. Shooting 4K bought you nothing there.

None of this raises an error. At the fixed 30 fps output, four consecutive missed frames is 133 milliseconds — invisible while scrubbing, entirely sufficient as a screenshot.

Blur is a veil, not a blackout

Before any settings advice: a Gaussian blur is the weakest redaction on this page. It attenuates high-frequency detail but leaves the low-frequency structure of the face intact, which is why blurred and pixelated faces are the standing worked example in obfuscation-defeat research — trained models recover identity from them far above chance. An opaque fill leaves nothing to recover. A blur leaves a reconstruction problem.

That changes how you verify. "Can I read the face?" is the wrong test, because you are not the adversary. If re-identification would have real consequences for the person in frame, the answer is not a larger radius — it is an opaque overlay on a locked-off shot, a crop, or not publishing the shot. Everything below tunes the ordinary case: a colleague who would rather not turn up in a company video. Be honest about which case you are in. The section further down refuses Video Text's 50-per-cent-opacity box as redaction; a 25-pixel blur deserves the same scepticism.

Blur strength is in output pixels, not fractions of a face

The blur is a canvas filter at whatever number the slider says, applied in the coordinate space of the output frame. blur(25px) is a Gaussian with a 25-pixel standard deviation on a 3840×2160 canvas, not 25 per cent of anything.

On the 4K clip a talking-head face box is maybe 700 px wide, so the default σ of 25 is 3.6 per cent of it — enough to look blurred in a thumbnail, not enough to destroy features in a still viewed at 100 per cent. Downscale to 1080p and the same box is around 350 px, so 25 is now 7 per cent and doing twice the work. The slider tops out at 60, which on a tight 4K close-up still cannot reach a comfortable ratio. Run Video Resolution first, down to 1080p, then blur. You lose nothing the detector was using. Rough target for the ordinary case: σ at 10 per cent of face box width or more.

The mask is an ellipse, and near an edge it is in the wrong place

The mask is an ellipse inscribed in the padded detector box, and an inscribed ellipse touches its box at only four midpoints — the corners fall outside it. Work the geometry on the 20 per cent default and the detector's own box corners sit just beyond the ellipse; you need 20.7 per cent before the ellipse fully contains the box it was derived from. And that box hugs the face: no hair, no ears, no underside of the jaw, all recognisable. Treat 35–40 per cent as the working default.

Near a frame edge there is a second and quite different problem, one of placement rather than detection. The code clamps the padded origin into the frame — Math.max(0, box.originX - padX) — without shrinking the padded width to match. A 200-pixel face box sitting 10 pixels from the left edge, at 20 per cent padding, produces a region starting at x = 0 and 280 pixels wide, so the ellipse centre lands at 140 while the face centre is at 110. The mask is 30 pixels off, displaced away from the edge. The 40 pixels of padding you asked for on the exposed side collapse to 10, and the exposed side is exactly where an ellipse is thinnest. Anyone entering or leaving frame gets a mask that is both smaller than you think and in the wrong place — an independent argument for the higher padding number.

What the whole pipeline costs

Face Blur runs on real-time playback, so three minutes of clip costs at least three minutes, and the tab has to stay in the foreground: requestAnimationFrame stops when it is hidden, and the picture freezes while the audio rolls on.

But the blur is the cheap step. The full chain for this clip is Video Resolution (a complete ffmpeg.wasm re-encode of three minutes of 4K, and by a distance the dominant cost), then Face Blur, then Video Converter for WebM to MP4, then Video Crop if you need it, then Metadata Cleaner, which is a stream copy and effectively free. ffmpeg.wasm runs a long way short of native speed and a 4K decode is its worst case. Budget in tens of minutes, not three.

Do not take my estimate — machines differ by more than an order of magnitude. Cut 20 seconds out of the middle with Video Trimmer, run the entire chain on it end to end, and multiply by nine. That is the only number that applies to your laptop.

It also settles the ordering argument. The downscale is the expensive step that makes every step after it cheap, because each later re-encode then runs on a quarter of the pixels — so it pays for itself twice, once in the slider becoming meaningful and once in everything downstream.

One more structural problem: Face Blur has no checkpoint, so an error at 2:40 costs the whole run. Split first. Video Splitter into 30-second parts — tick the precise option if the cuts must land where you asked, since the stream-copy path snaps to the nearest keyframe — then blur each part. A failure now costs one segment, and you can run segments in parallel tabs. Video Merger takes the blurred WebM parts straight back to a single MP4, re-encoding with libx264 and AAC rather than stream-copying, which folds the conversion step in for free. Check the joins: MediaRecorder's timing headers are not the concat demuxer's favourite input.

The QC pass: extract stills, do not scrub

Convert the WebM before inspecting it — partly because most platforms want MP4, partly because MediaRecorder writes WebM with no duration in the header, so browsers report duration as Infinity and anything built on it, scrub sliders included, misbehaves.

Then open the MP4 in Frame Extractor and pull stills where the mechanism predicts trouble:

  • every cut, plus the first and last frame of every shot
  • every turn towards profile, every look down
  • every occlusion: a hand, a mug, a forklift
  • anyone entering or leaving frame — for the off-centre mask as much as for the miss
  • every lighting change, especially anyone in front of a window
  • every background person, at the moment they are largest in frame

The slider steps in 0.01 s, roughly three steps per frame at the fixed 30 fps output, and capture writes a PNG you can open at full size. That is the point: real decoded frames, not a preview the browser may be rendering at reduced rate.

Manual inspection does not scale to a one-hour recording. For long material, desktop FFmpeg: ffmpeg -i blurred.mp4 -vf "select='gt(scene,0.25)'" -vsync vfr qc_%04d.png gives a still at every scene change, and ffmpeg -ss 00:01:12 -to 00:01:16 -i blurred.mp4 qc_%04d.png gives every frame of a suspicious window.

When you find a dropout, re-running changes nothing: the threshold is fixed in the component, so the detector fails the same way every time. You have exactly one lever over a fixed threshold, and it is framing. Detection depends on face size as a fraction of the frame, so cropping in makes everyone still in frame proportionally larger, and a background face that sat below threshold at full width can cross it in a punch-in. Crop the original with Video Crop first, then blur the crop — on 4K source a 1.5× punch-in costs nothing visible at a 1080p delivery.

If framing cannot rescue it, remove it: trim the moment out with Video Trimmer, split it out with Video Splitter, or crop the person out of frame entirely. A dropout inside a shot you cannot lose is where a desktop NLE with a tracked, keyframed mask — Premiere, Resolve — earns its licence fee.

What a face mask never touches

This is the part that burns people. The blur covers one body part; identification does not confine itself to one.

Identifier Why the blur misses it What removes it
Badge, lanyard, embroidered uniform Below the face box Crop, or an opaque overlay if it never moves
Plate, house number, shop sign Not a face Crop the region, or cut the shot
Tattoo, scar, hair, jacket, gait Outside the ellipse Crop or cut; nothing else
Reflection in a monitor or window Small, dim, off-axis — below threshold Crop or cut the shot
Burned-in timestamp or camera OSD Pixels, not metadata Crop the strip off
The voice Audio is never touched Mute, or replace the track
A PA announcement, a doorbell, a shift bell Same Mute or replace that section

Do not use the background box in Video Text as a redaction: it is drawn at black@0.5, a veil rather than a blackout, and lifting exposure on a still brings back what is underneath.

To cover rather than crop, the tool is Logo Watermark with an opaque PNG at opacity 1 — but know its real shape before you commit. It offers five anchor points, not nine: four corners and dead centre. The logo is scaled by width to between 5 and 50 per cent of the frame, the corner margin is fixed at 3 per cent of frame width, and the placement is the same for the entire clip.

That still reaches more than it looks like it does, because the transparent part of a PNG counts. Build the image at the full 50 per cent allowance — transparent everywhere except the opaque rectangle, positioned inside the canvas where you need it — and anchor it to the nearest corner, letting the padding do the aiming. Height is free: the tool scales by width and lets height follow, so a tall canvas reaches as far down the frame as you like. And since 50 per cent plus the 3 per cent margin is more than half the frame, one anchor or the other can reach any horizontal position; run the tool twice for two separate regions. Run it on the MP4 rather than the WebM, since it stream-copies audio into an MP4 container.

The ceiling is real, though: it is one static rectangle for the whole clip. It works on a badge in a locked-off interview. It does nothing for a badge on someone walking. For anything that moves, the honest options are Video Crop or dropping the shot.

Video Crop takes plain X/Y/W/H numbers and runs crop=W:H:X:Y with a re-encode. It will let you type an odd width or height and the H.264 encoder will then refuse the job — keep both even. And a crop tight enough to lose a badge is often tight enough to lose the shot; sometimes the answer is that this shot cannot be published.

The voice is the identifier people rationalise away. The export's audio is the original, pulled off the source element and re-encoded to Opus, otherwise untouched. If the voice is recognisable to the people you are protecting her from — colleagues, a manager, a partner — the blur achieved nothing. Video Mute strips the track with a stream copy on the video, and subtitles or re-recorded narration carry the content. Pitch-shifting is a costume, not a mask: anyone who applies the inverse ratio gets the voice back. It is also not a one-click step here — Audio Speed & Pitch is an audio-only tool and refuses an MP4 outright, so the route is Extract Audio, then the shift, then Audio Replacer to put the track back on the video.

One browser trap: that audio track comes from captureStream() on the source video element, a Chromium API. In Firefox and Safari the call is guarded and returns nothing, so the export is silent with no warning — check before you assume you removed the voice, and before you assume you kept it.

Clean the file last, not first

Run Metadata Cleaner on the file you will actually share — the converted, cropped, trimmed export — not the camera original. The original is the copy you inspect; the export is the copy that leaves the building, and every step in between writes fresh container metadata.

The cleaner scans first and lists what it found, then runs -map_metadata -1 -map_chapters -1 -c copy with bitexact flags. That single argument is what removes the identifying keys: on an iPhone file, com.apple.quicktime.location.iso6709 — the GPS coordinates of the shot — along with device make, model, software version and creation timestamp. Afterwards the tool also deletes the empty moov > udta atom that FFmpeg writes even when you asked for no metadata. Treat that second step as cosmetic: the stub carries nothing, it is only a structural fingerprint, and it is skipped entirely on faststart-layout files, where mdat follows moov and removing bytes would shift every sample offset baked into stco/co64. On a large class of MP4s the stub survives, and that is fine.

On the Face Blur path specifically, the camera atoms were gone before the cleaner ever ran. The export is a fresh MediaRecorder WebM muxed from a canvas capture stream, so nothing from the source container survives the re-render — no GPS, no make, no model. Here the cleaner is cheap insurance against what Converter, Crop and Watermark write fresh. It matters far more on the path that skips Face Blur: a crop-only or trim-only cut of the phone original still carries the whole udta block, GPS included. The cleaner also understands only the MP4 family, which is one more reason to convert the WebM first. For independent confirmation, run exiftool on the download; the Video Metadata viewer shows what the browser exposes, not the atom tree.

Two gaps it cannot close. Filenames accumulate rather than disappear. Face Blur names its download blurred_ plus the original stem, the converter keeps whatever stem it was handed, and the cleaner prefixes cleaned_ — so warehouse_interview_sarah_v3.mp4 arrives at the far end as cleaned_blurred_warehouse_interview_sarah_v3.mp4, carrying her name through every step that was supposed to remove it. Rename it yourself. And stripping a container does nothing about the master in your camera roll, your cloud backup and the email thread you sent it in.

Back to the warehouse clip

Reopen the 4K original, not the blurred export, and answer one question before you touch a tool: does that aisle need to be in shot at all? If it does not, crop it out and you have deleted two detection problems rather than solving them. If it does, punch in so those faces are large enough to be found, and accept the tighter frame as the cost.

Then run it in this order: crop, downscale to 1080p, blur at 35–40 per cent padding, convert to MP4, pull your stills, mute or replace the audio if the voice gives her away, clean the metadata, rename the file. Time the chain on a 20-second test cut before you commit the full three minutes.

And when it finishes, remember what finishing means here. The progress bar reports that the loop ran, not that it worked. The only evidence you will ever get is the frames you go and look at.

Try it yourself — free in your browser

No upload, no signup, no watermark — these tools run on FFmpeg WebAssembly locally.

Tags:face bluranonymizationprivacymetadataquality controlmediapipe