Drop in a finished file — sound, picture, captions — pick where it is going, and find out whether it will be accepted, with the exact timestamp of everything that is wrong. Then, if you want, a corrected copy that is measured again from scratch before it claims to be fixed.
What comes back
Not a dashboard. A list of what the file is, what the target asks for, and the difference between them — in the order that decides whether you can ship today. Checks with nothing to measure get one line between them, not one line each.
✕ Not ready — the failures below would be rejected.
1 failed, 5 warned, 10 passed, 13 not checked
Black Aura (hot master).wav · 03:06 · pcm_s16le 48 kHz stereo · ass 1254 cues sidecar
Loudness over time. The shaded band is what this target asks for; the lighter area is the range each column covers.
Problems occur at 00:00, 00:03, 00:06, 00:07, 00:11, 00:16, 00:20, 00:24, 00:35, 00:37, and 20 more
Thresholds from: YouTube recommended upload encoding settings (published); the loudness figure is measured behaviour, not documented (read 2026-09)
A file with chapter markers also gets a per-chapter table, which makes the chapter recorded eighteen decibels louder than its neighbours impossible to miss.
A list of timestamps tells you to look at 18:07. It does not tell you that 18:07 is one of nine identical spikes and the real problem is a compressor doing something odd.
The chart is bucketed — no display has 36,000 pixels for a ten-hour audiobook — and each bucket keeps its loudest and quietest value rather than an average. Averaging is precisely the operation that hides a spike next to a hole, and spikes next to holes are what this tool is for.
A target band is shaded only when the target states its requirement in the unit the chart is drawn in. ACX asks for an RMS level, which is not LUFS, so on that target no band appears and the chart says so.
Everything else about a caption file measures it against itself. This measures it against the audio, using the silence the audio pass already found — so it costs no decode, and it answers the two questions that otherwise mean scrubbing a two-hour recording: which passages did nobody caption, and is the whole file out of sync.
The gaps reported are the ones within each stretch of sound, not the stretches as a whole — a twelve-second passage with three seconds of caption on the front has nine seconds missing, and pointing at the whole passage would be pointing at the three that are fine. Drift is a median of every cue's distance from the nearest moment sound starts, with a confidence figure, because cues legitimately sit mid-sentence and one of those must not become the answer.
Music and atmosphere are legitimately uncaptioned, so this finds passages to look at rather than faults. What it does is find them in seconds.
The confidence figure is why drift can be trusted at all. Run against thirty-three minutes of speech, 79% of cues sat near a moment sound started and the median said the transcript was 0.19 s early — in sync. Run against three minutes of continuous music, 1% matched, because a song offers almost no onsets to match against; the median of those twelve cues claimed the file was 2.32 seconds out. One of those is a measurement and the other is a number about nothing, and the tool now reports only the first.
A finding carries a timestamp only when the timeline holds the same quantity the rule is about.
Short-term loudness can locate an integrated-loudness failure, because both are loudness. It cannot locate an RMS failure: RMS and LUFS are different measurements, and pointing at a moment measured in one while quoting a threshold in the other would be an invention.
So those findings say measured across the whole file instead of offering a number that looks precise and means nothing.
One engine, two ways in
Drop files or a whole delivery. Media Preflight sorts audio, video and subtitles before asking what each kind is for, estimates the real work, and then runs one visible queue. The verdicts do not change with the mode — only how many decisions are exposed.
Audio, video and subtitle files are identified and grouped. Pick one compatible target for each group, review what will run, then start. A subtitle file is never offered a video target and never enters the picture pipeline.
Give each item its own built-in or saved profile, choose selective or full picture analysis, and inspect the estimate before committing the machine. Compatibility is enforced by the server as well as the dropdown.
Eight WAVs and eight videos are not half-and-half. Progress is weighted by the estimated analysis cost, runs in natural file order, and reports elapsed time and an ETA instead of moving one equal step per file.
A custom target is readable JSON in your own configuration folder: named, validated, editable, diffable and easy to send to a client. Change a published threshold and it is relabelled as your house rule rather than borrowing provenance it no longer deserves.
Switching modes does not re-probe or re-measure anything. The intake stays on this machine under a short-lived local token, and the preference itself stays in the browser.
A folder, not a file
An audiobook is thirty chapters and ACX accepts or rejects the title. Some of what it asks for is a property of the delivery and of no file in it.
Every file passes. The title does not — one chapter is stereo where the rest are mono, which a rule written against one file cannot see. A per-file check would clear all thirty chapters of a real title and let it be rejected on submission.
Not the two extremes. When four chapters agree and a fifth is four decibels up, the quietest of the four is not at fault — it is the reference. Files further from the middle of the set than half its spread are named; when that describes nobody, as in a smooth ramp, both ends are.
A delivery is ordered. A report that lists chapter 10 second is one somebody has to re-sort in their head before they can read it.
chapter-01, -02, -04 is a delivery short one file, and nothing about any file in it is wrong. Only what actually looks like a sequence is checked — three or more files sharing a prefix, a suffix and a digit width — because a folder of unrelated names has no sequence to be missing from, and inventing one would produce a finding about nothing.
A folder checked twice would otherwise start checking the corrected copies it wrote the first time, and a corrected copy of a corrected copy is nobody's delivery.
A file that cannot be read is reported and the rest are checked. One broken file in thirty should not cost you the other twenty-nine.
Correcting each file on its own is exactly what does not fix a set — every chapter above is inside ACX's band. So the delivery decides what its files should agree on: the majority settles the format questions, because a title is almost never wrong in the majority, and the level is settled by the target's band rather than the majority, since bringing four quiet chapters up to meet a loud fifth would satisfy the set rule by making every file wrong.
A stereo chapter downmixed to mono then comes back at a level nothing could have predicted, because how much a downmix costs depends on how alike the two channels were. So the delivery is measured and the files that landed off target are built again from their sources — however many rounds it takes, the number of lossy encodes stays at one.
Where it is going
Each carries the source its thresholds came from and the month they were read. The report prints both, every time.
| Target | What it is | Thresholds |
|---|---|---|
acx | Audiobook, ACX retail delivery | published |
ebu_r128 | Broadcast, EBU R 128 | published |
spotify_podcast | Podcast delivery to Spotify | informal |
youtube | YouTube upload | published |
social_vertical | Instagram / TikTok | informal |
web | Generic web video | informal |
subtitles | Caption readability, checked on its own | informal |
Four of these are marked informal because there is no published specification to point at. Two are platforms that document nothing; the third is a set of subtitling conventions reasonable people disagree about — seventeen characters a second is comfortable for an adult viewer in English and far too fast for a children's programme; and the fourth is Spotify, whose −14 LUFS figure is published for music playback normalisation rather than as a podcast delivery requirement, which Spotify does not publish at all. Their thresholds are offered as a sanity check you should adjust, not as a promise about what anybody requires today. This tool measures your file exactly and compares it against a number you can see and change. It is not, and cannot be, an oracle for somebody else's current ingest rules, and the moment it presents itself as one it is lying to you.
Within one target some numbers are published and some are not. YouTube documents its encoding settings and states no loudness figure anywhere; the −14 LUFS everybody quotes is measured behaviour. Presenting both as equally authoritative would mislead about the more important one, so a finding on a threshold that is not from a specification says so before it says anything else — (measured behaviour, not a published figure.) or (this tool's own threshold, not a rule of the target.)
Every published number here has been read against its source document, and the reading changed three of them. ACX's own page recommends one to five seconds of room tone at both ends — the half-second opening that appears in a great deal of guidance elsewhere is on no ACX page, so that check now warns at a corrected band rather than failing at an invented one. YouTube's guide is explicit that interlaced content must be deinterlaced before upload, so that check fails rather than warns. Spotify asks for true peak below −2 dBTP on masters louder than −14 LUFS, which is now the warning band inside the −1 dBTP limit. EBU R 128 was right as written.
Your own target is a JSON file, not a patch to the tool — a list of rules naming a metric, a band, and what to say when a file falls outside it. Point --target at the path, or drop it in profiles/ and it joins the list in the window.
The corrected copy
Corrections are arithmetic on a signal — a gain change, a limiter, a trim, a re-encode. Nothing here invents audio that was never recorded.
Every correction writes a new file beside the original, refuses to write over the source, and refuses to overwrite anything else unless told twice. There is a test asserting the source's bytes are identical afterwards.
The whole plan is shown in plain English and as the exact ffmpeg command, before anything runs. A correction that cannot be described in a sentence does not belong in the planner.
Raising a quiet recording by eleven decibels raises its peaks and its DC offset by eleven decibels too. The planner works out where those would land and adds a limiter or a high-pass — saying in writing that the source was fine and the gain is what would have broken it.
The chart of a corrected copy carries the earlier reading dashed behind it, both plotted against whichever file is longer — so a trimmed ending shows as the old line outlasting the new one. Scaling them to a common width would hide the very thing worth seeing.
The finished file is analysed again from nothing, and the report you get is that measurement. When it lands off target the correction is recomputed and applied to the original again — never to the copy — so however many attempts it takes, the number of lossy encodes stays at one.
Room tone is copied from a quiet stretch of the file itself wherever there is one, rather than padded with digital zeroes — because a target asking for a quiet ending is asking for the room, not for nothing.
Every other fix re-encodes, and the plan says so in a caveat. Moving an MP4's index in front of its media does not: nothing about the sound or the picture changes, only the order of the boxes in the container. Every stream is copied through — no filter, no encoder — so the corrected copy holds exactly the media the original did, and a test proves it by comparing a hash of the decoded video before and after.
The flashing check counts large frame-to-frame changes in average luminance and flags any second holding three or more, which is where WCAG draws its general flash threshold. It does no spatial analysis, does not measure how much of the screen changed, and knows nothing about the separate red-flash rule. It finds the passages worth looking at with human eyes. Silence from it is not a pass, and every report it appears in says so.
Noise reduction is deliberately absent. A noise floor above the target is reported and left alone, because the answer is to treat the room or re-record — and a tool that quietly ran a denoiser over somebody's audiobook would be doing something its owner did not ask for.
Under it
A second read of a two-hour audiobook costs a minute of somebody's afternoon, and measurements taken in separate passes can disagree about where a moment is.
Metadata goes to stdout and events to stderr, and a filled pipe that nobody is draining is a deadlock — one that looks exactly like a slow file. Draining stderr in a thread is not a detail; it is the difference between a tool that works on a long file and one that appears to hang on it.
ebur128 reports every 100 ms. The timeline holds the loudest momentary and short-term value in each second, because the report quotes seconds — so a ten-hour audiobook costs 36,000 rows instead of 360,000, and nothing is lost that would have survived to the page.
Locating peaks and clipping, and only when the whole-file peak says a peak problem can exist at all. A file whose loudest sample sits below the clipping threshold has no clipped samples — that is arithmetic, not an estimate — so the common case is answered without reading the audio again.
loudnorm's own measurement, because it will not accept another filter's figures; it wants its own threshold and offset.
blackdetect → freezedetect → idet → signalstats, and it runs only when the target actually asks a question about the picture. Decoding a ninety-minute film to count black frames is minutes of somebody's time, and a podcast profile has no reason to spend them.
It also carries only the filters the target's rules actually read, because those four are not equally priced. Timed on a 1080p60 file, the whole chain takes about as long as the file itself — and more than half of that is one filter.
Reading field_order from the container and believing it passes a 3:2-pulldown file that declares itself progressive — which is most telecined material there is. Running ffmpeg's idet and believing that is worse: on progressive material with hard vertical edges and fast motion it calls the majority of frames interlaced.
What separates them is that real interlacing is overwhelmingly one field order and the false positives are mixed, because they are noise rather than field dominance. Both a share test and a dominance test have to pass; material that passes one and fails the other is reported as inconclusive, which is the truthful answer and not a verdict.
Slim distribution builds omit ebur128, and the failure otherwise surfaces halfway through an analysis as a filtergraph error that says nothing about why. The tool looks for the filter, not for a binary with the right name.
Measured on one 8-core laptop, against a three-minute 1080p60 file. The right-hand column is what that filter adds per minute of video at that size; a smaller picture costs proportionally less, because these filters work per pixel.
| Filter | What it answers | Per minute of 1080p60 |
|---|---|---|
blackdetect | Black frames, and where | 2.4 s |
freezedetect | Frozen frames, and where | 4.2 s |
signalstats | Average luminance, for the flash screening | 17 s |
idet | Interlacing and telecine | 27 s |
| decoding | — unavoidable | 3.6 s |
So a ninety-minute feature checked against every picture question is roughly eighty minutes of work, and the interlacing detector alone is half of it. Against a broadcast loudness target, which asks nothing about the picture, it is nothing at all. On the three-minute file above, an Instagram check — which needs black frames and the flash screening, but not fields — took 62 seconds instead of 176.
That choice is offered rather than assumed. The selective pass is the default; --picture full measures everything the picture can be asked whatever the target wants, and the report grows an “also measured in the picture, against nothing” section for the answers no rule read — because “the target does not ask” and “the file is fine” are different sentences, and somebody handing over a master may want both. Either way the tool says what it is about to cost before it starts, and the window puts both times on the dropdown.
Both were tried and both were measured. Filter threading (-filter_threads 8) came back at 28.6 s against a 27.1 s baseline. Hardware decoding (-hwaccel videotoolbox) came back at 27.7 s. Neither is an improvement, and the reason is visible in the table: decoding is 4% of the work. The cost is in the filters, and they run on one core.
The first attempt at measuring them reported 0.03 s, which would have been a fifty-fold speedup and was in fact a command that had failed instantly. A number that good is a bug report.
Reading the picture at half width takes 15.3 s instead of 27.1; at quarter width, 7.9 s. Horizontally only — blending adjacent lines is exactly what idet compares, so vertical scaling is never offered at all.
Across seven test files, black, frozen and flashing findings were identical at full, half and quarter width. Average luminance is the mean of the same pixels either way, so the flash screening is unaffected by construction rather than by luck.
The field checks are the exception, and it is refused rather than degraded. On a near-static picture with one small moving element, full width reports progressive and half width reports that it cannot tell. An inconclusive answer is not a cheaper answer, so asking for both together is an error with a sentence attached, not a quiet loss of confidence.
Running it
Python 3.10 or newer, standard library only. ffmpeg and ffprobe are located at runtime rather than bundled, and their absence is reported as a sentence with the install line for your platform.
# something to double-click: a .app, a .pyz, a .desktop
./build.sh
The build bundles neither Python nor ffmpeg, and that is a decision rather than an omission. Bundling Python would mean a build-time dependency on a tool whose whole claim is that it needs nothing installed; bundling ffmpeg means eighty megabytes and somebody else's licensing decision. So the bundle removes the terminal, not the prerequisites — it looks for both, and when one is missing it says which, and prints the command that installs it.
# the window — a local server and a page in your browser python3 app.py # the command line python3 preflight.py check finished.mp3 --target acx python3 preflight.py batch chapters/ --target acx python3 preflight.py batch chapters/ --target acx --fix python3 preflight.py fix finished.mp3 --target acx --dry-run python3 preflight.py fix finished.mp3 --target acx
check exits 0 when the file passes, 1 when it fails, and 2 when the tool itself could not run — so it drops into a build script without parsing anything.
The window listens on 127.0.0.1 with a token minted at startup. Your file is never uploaded: a browser cannot tell a local program where a dropped file lives, so the page asks the tool to open your system's own file dialog and the path never leaves the machine. That is not a workaround — it is the point.
The wider workshop
Move between guided-meditation authoring, measured speech, lyric video, delivery validation, and the small field tools that support all of them.
Assemble guided sessions and keep published maps beside what was actually found.
Explore Gateway Forge →A transparent Piper workbench for timing, pronunciation, and honest audio export.
Explore Voice Forge →Karaoke and narration video with a vector visor face that mouths every word.
Explore Protoke →Check a finished file against where it is going, and correct it without touching the original.
Explore Media Preflight →Small utilities for audio inspection, corpus work, authoring, and repeatable checks.
Open the toolbox →If a number looks wrong
A threshold that has moved, a target that has changed its rules, a measurement that disagrees with your own meter — all of those are worth an issue. The profiles carry the month they were read precisely so that they can be argued with.