macOS · Windows · Linux · MIT

Media Preflight

Drop in a finished file — sound, picture, captions — pick where it is going, and find out whether it will be accepted, with the exact timestamp of everything that is wrong. Then, if you want, a corrected copy that is measured again from scratch before it claims to be fixed.

7 delivery targets, and your own 1 decode per stream, every measurement 410 checks passing 0 uploads, accounts, credits Local-first. Nothing leaves the machine

What comes back

A verdict, and where to look

Not a dashboard. A list of what the file is, what the target asks for, and the difference between them — in the order that decides whether you can ship today. Checks with nothing to measure get one line between them, not one line each.

✕ Not ready — the failures below would be rejected.

1 failed, 5 warned, 10 passed, 13 not checked

Black Aura (hot master).wav · 03:06 · pcm_s16le 48 kHz stereo · ass 1254 cues sidecar

-20-10clipping — peak 0.00 dBFS at 00:03clipping — peak 0.00 dBFS at 00:06clipping — peak 0.00 dBFS at 00:11clipping — peak 0.00 dBFS at 00:16clipping — peak 0.00 dBFS at 00:20clipping — peak 0.00 dBFS at 00:24clipping — peak 0.00 dBFS at 00:35clipping — peak 0.00 dBFS at 00:37clipping — peak 0.00 dBFS at 00:40clipping — peak 0.00 dBFS at 00:45clipping — peak 0.00 dBFS at 00:58clipping — peak 0.00 dBFS at 01:03clipping — peak 0.00 dBFS at 01:08clipping — peak 0.00 dBFS at 01:16clipping — peak 0.00 dBFS at 01:21clipping — peak 0.00 dBFS at 01:26clipping — peak 0.00 dBFS at 01:30clipping — peak 0.00 dBFS at 01:33clipping — peak 0.00 dBFS at 01:36clipping — peak 0.00 dBFS at 01:40clipping — peak -0.05 dBFS at 01:45clipping — peak 0.00 dBFS at 01:48clipping — peak 0.00 dBFS at 01:52clipping — peak 0.00 dBFS at 01:56clipping — peak 0.00 dBFS at 02:18clipping — peak 0.00 dBFS at 02:20clipping — peak 0.00 dBFS at 02:51clipping — peak 0.00 dBFS at 03:02Integrated loudness at 00:07Integrated loudness at 02:53True peak — first moment the true peak passed the ceiling at 00:00Sound with no caption — 14.1 s of sound with no caption at 02:52Clipped samples — peak 0.00 dBFS at 00:03Clipped samples — peak 0.00 dBFS at 00:06Clipped samples — peak 0.00 dBFS at 00:11Clipped samples — peak 0.00 dBFS at 00:16Clipped samples — peak 0.00 dBFS at 00:20Clipped samples — peak 0.00 dBFS at 00:24Clipped samples — peak 0.00 dBFS at 00:35Clipped samples — peak 0.00 dBFS at 00:37Clipped samples — peak 0.00 dBFS at 00:40Clipped samples — peak 0.00 dBFS at 00:45Clipped samples — peak 0.00 dBFS at 00:58Clipped samples — peak 0.00 dBFS at 01:03Clipped samples — peak 0.00 dBFS at 01:08Clipped samples — peak 0.00 dBFS at 01:16Clipped samples — peak 0.00 dBFS at 01:21Clipped samples — peak 0.00 dBFS at 01:26Clipped samples — peak 0.00 dBFS at 01:30Clipped samples — peak 0.00 dBFS at 01:33Clipped samples — peak 0.00 dBFS at 01:36Clipped samples — peak 0.00 dBFS at 01:40Clipped samples — peak -0.05 dBFS at 01:45Clipped samples — peak 0.00 dBFS at 01:48Clipped samples — peak 0.00 dBFS at 01:52Clipped samples — peak 0.00 dBFS at 01:56Clipped samples — peak 0.00 dBFS at 02:18Clipped samples — peak 0.00 dBFS at 02:20Clipped samples — peak 0.00 dBFS at 02:51Clipped samples — peak 0.00 dBFS at 03:02captioned 00:00–02:52cc00:0001:0002:0003:00

Loudness over time. The shaded band is what this target asks for; the lighter area is the range each column covers.

True peak: 0.2 dBTPRequired: ≤ -1 dBTP at 00:00 (first moment the true peak passed the ceiling) YouTube does not publish a peak ceiling. A decibel of headroom is what stops its AAC encode handing peaks back above full scale.
Integrated loudness: -8.7 LUFSRequired: -17 to -12 LUFS at 00:07, 02:53 (measured behaviour, not a published figure.) YouTube's own encoding guide states no loudness figure; -14 LUFS is what its normalisation is measured to do.
Sound with no caption: 14.13 sRequired: ≤ 0 s at 02:52 (14.1 s of sound with no caption) Music and atmosphere are legitimately uncaptioned, so this finds passages to look at rather than faults — but it finds them in a two-hour recording in seconds.
Clipped samples: 137 sRequired: ≤ 0 s at 00:03, 00:06, 00:11, 00:16, 00:20, 00:24, +20 more
Stereo content: -0.73Required: ≥ -0.5 Sustained negative correlation between the two channels: the sides are out of phase and will cancel in mono.
·13 not checked — nothing in this file to measure them against: fast start, video codec, pixel format, frame height, frame rate mode, interlacing, and 7 more

Problems occur at 00:00, 00:03, 00:06, 00:07, 00:11, 00:16, 00:20, 00:24, 00:35, 00:37, and 20 more

Create corrected copyExport reportExport JSON

Thresholds from: YouTube recommended upload encoding settings (published); the loudness figure is measured behaviour, not documented (read 2026-09)

The window, checking a real three-minute master against YouTube with its word-level lyric subtitles beside it. Two of these findings are real: the fourteen seconds of instrumental outro nobody captioned, and the out-of-phase stereo. The clipping and the peak were planted, by copying the master and pushing it six decibels, because the original has neither.

What it measures

  • Integrated loudness, loudness range, short-term excursions
  • True peak and sample peak, with the moment each ceiling broke
  • Clipped samples, located to the second
  • Noise floor, DC offset, dead channels, channel imbalance
  • Stereo phase — sides that will cancel when someone plays it in mono
  • Room tone at head and tail, mid-file gaps, endings that stop mid-word
  • Black frames and frozen frames, with the seconds they occupy
  • Flashing passages worth a human look, screened at three transitions a second
  • Resolution, pixel format, telecine, and constant versus variable frame rate
  • Interlacing measured from the picture, not read off a header that often lies
  • Caption overlaps, reading speed, line length, cues timed past the last frame
  • Passages of sound nobody captioned, and captions that run out of sync with the audio
  • Subtitle fonts this machine does not have
  • Codec, container, sample rate, channels, bitrate, and whether an MP3 is constant or variable
  • Whether an MP4's index sits in front of its media, so it can start playing before it has finished downloading

A file with chapter markers also gets a per-chapter table, which makes the chapter recorded eighteen decibels louder than its neighbours impossible to miss.

The shape, not just the list

A list of timestamps tells you to look at 18:07. It does not tell you that 18:07 is one of nine identical spikes and the real problem is a compressor doing something odd.

The chart is bucketed — no display has 36,000 pixels for a ten-hour audiobook — and each bucket keeps its loudest and quietest value rather than an average. Averaging is precisely the operation that hides a spike next to a hole, and spikes next to holes are what this tool is for.

A target band is shaded only when the target states its requirement in the unit the chart is drawn in. ACX asks for an RMS level, which is not LUFS, so on that target no band appears and the chart says so.

Do the captions match the programme?

Everything else about a caption file measures it against itself. This measures it against the audio, using the silence the audio pass already found — so it costs no decode, and it answers the two questions that otherwise mean scrubbing a two-hour recording: which passages did nobody caption, and is the whole file out of sync.

The gaps reported are the ones within each stretch of sound, not the stretches as a whole — a twelve-second passage with three seconds of caption on the front has nine seconds missing, and pointing at the whole passage would be pointing at the three that are fine. Drift is a median of every cue's distance from the nearest moment sound starts, with a confidence figure, because cues legitimately sit mid-sentence and one of those must not become the answer.

Music and atmosphere are legitimately uncaptioned, so this finds passages to look at rather than faults. What it does is find them in seconds.

The confidence figure is why drift can be trusted at all. Run against thirty-three minutes of speech, 79% of cues sat near a moment sound started and the median said the transcript was 0.19 s early — in sync. Run against three minutes of continuous music, 1% matched, because a song offers almost no onsets to match against; the median of those twelve cues claimed the file was 2.32 seconds out. One of those is a measurement and the other is a number about nothing, and the tool now reports only the first.

Where the timestamps come from

A finding carries a timestamp only when the timeline holds the same quantity the rule is about.

Short-term loudness can locate an integrated-loudness failure, because both are loudness. It cannot locate an RMS failure: RMS and LUFS are different measurements, and pointing at a moment measured in one while quoting a threshold in the other would be an invention.

So those findings say measured across the whole file instead of offering a number that looks precise and means nothing.

One engine, two ways in

Guided when you want the answer. Professional when you know the handoff.

Drop files or a whole delivery. Media Preflight sorts audio, video and subtitles before asking what each kind is for, estimates the real work, and then runs one visible queue. The verdicts do not change with the mode — only how many decisions are exposed.

Guided / One target per kind

Sort first, decide once

Audio, video and subtitle files are identified and grouped. Pick one compatible target for each group, review what will run, then start. A subtitle file is never offered a video target and never enters the picture pipeline.

Professional / One target per file

Assign the delivery precisely

Give each item its own built-in or saved profile, choose selective or full picture analysis, and inspect the estimate before committing the machine. Compatibility is enforced by the server as well as the dropdown.

Measured progress

The bar follows the expensive work

Eight WAVs and eight videos are not half-and-half. Progress is weighted by the estimated analysis cost, runs in natural file order, and reports elapsed time and an ETA instead of moving one equal step per file.

Portable targets

The wizard writes a real file

A custom target is readable JSON in your own configuration folder: named, validated, editable, diffable and easy to send to a client. Change a published threshold and it is relabelled as your house rule rather than borrowing provenance it no longer deserves.

Switching modes does not re-probe or re-measure anything. The intake stays on this machine under a short-lived local token, and the preference itself stays in the browser.

A folder, not a file

Some faults belong to the set

An audiobook is thirty chapters and ACX accepts or rejects the title. Some of what it asks for is a property of the delivery and of no file in it.

Every file passes. The title does not — one chapter is stereo where the rest are mono, which a rule written against one file cannot see. A per-file check would clear all thirty chapters of a real title and let it be rejected on submission.

It names the file to fix

Not the two extremes. When four chapters agree and a fifth is four decibels up, the quietest of the four is not at fault — it is the reference. Files further from the middle of the set than half its spread are named; when that describes nobody, as in a smooth ramp, both ends are.

Chapter 2 before chapter 10

A delivery is ordered. A report that lists chapter 10 second is one somebody has to re-sort in their head before they can read it.

A missing chapter is visible from the names

chapter-01, -02, -04 is a delivery short one file, and nothing about any file in it is wrong. Only what actually looks like a sequence is checked — three or more files sharing a prefix, a suffix and a digit width — because a folder of unrelated names has no sequence to be missing from, and inventing one would produce a finding about nothing.

Its own output stays out

A folder checked twice would otherwise start checking the corrected copies it wrote the first time, and a corrected copy of a corrected copy is nobody's delivery.

One bad file is not fatal

A file that cannot be read is reported and the rest are checked. One broken file in thirty should not cost you the other twenty-nine.

And it can correct one

Correcting each file on its own is exactly what does not fix a set — every chapter above is inside ACX's band. So the delivery decides what its files should agree on: the majority settles the format questions, because a title is almost never wrong in the majority, and the level is settled by the target's band rather than the majority, since bringing four quiet chapters up to meet a loud fifth would satisfy the set rule by making every file wrong.

A stereo chapter downmixed to mono then comes back at a level nothing could have predicted, because how much a downmix costs depends on how alike the two channels were. So the delivery is measured and the files that landed off target are built again from their sources — however many rounds it takes, the number of lossy encodes stays at one.

Where it is going

Seven targets, and the one you write

Each carries the source its thresholds came from and the month they were read. The report prints both, every time.

TargetWhat it isThresholds
acxAudiobook, ACX retail deliverypublished
ebu_r128Broadcast, EBU R 128published
spotify_podcastPodcast delivery to Spotifyinformal
youtubeYouTube uploadpublished
social_verticalInstagram / TikTokinformal
webGeneric web videoinformal
subtitlesCaption readability, checked on its owninformal

Four of these are marked informal because there is no published specification to point at. Two are platforms that document nothing; the third is a set of subtitling conventions reasonable people disagree about — seventeen characters a second is comfortable for an adult viewer in English and far too fast for a children's programme; and the fourth is Spotify, whose −14 LUFS figure is published for music playback normalisation rather than as a podcast delivery requirement, which Spotify does not publish at all. Their thresholds are offered as a sanity check you should adjust, not as a promise about what anybody requires today. This tool measures your file exactly and compares it against a number you can see and change. It is not, and cannot be, an oracle for somebody else's current ingest rules, and the moment it presents itself as one it is lying to you.

Provenance sits on the rule, not the target

Within one target some numbers are published and some are not. YouTube documents its encoding settings and states no loudness figure anywhere; the −14 LUFS everybody quotes is measured behaviour. Presenting both as equally authoritative would mislead about the more important one, so a finding on a threshold that is not from a specification says so before it says anything else — (measured behaviour, not a published figure.) or (this tool's own threshold, not a rule of the target.)

Every published number here has been read against its source document, and the reading changed three of them. ACX's own page recommends one to five seconds of room tone at both ends — the half-second opening that appears in a great deal of guidance elsewhere is on no ACX page, so that check now warns at a corrected band rather than failing at an invented one. YouTube's guide is explicit that interlaced content must be deinterlaced before upload, so that check fails rather than warns. Spotify asks for true peak below −2 dBTP on masters louder than −14 LUFS, which is now the warning band inside the −1 dBTP limit. EBU R 128 was right as written.

Your own target is a JSON file, not a patch to the tool — a list of rules naming a metric, a band, and what to say when a file falls outside it. Point --target at the path, or drop it in profiles/ and it joins the list in the window.

Writing a target →

The corrected copy

Four promises, kept in code

Corrections are arithmetic on a signal — a gain change, a limiter, a trim, a re-encode. Nothing here invents audio that was never recorded.

01 / Never the source

Your file is the file you still have

Every correction writes a new file beside the original, refuses to write over the source, and refuses to overwrite anything else unless told twice. There is a test asserting the source's bytes are identical afterwards.

02 / Nothing unpreviewed

Sentences first, then the argv

The whole plan is shown in plain English and as the exact ffmpeg command, before anything runs. A correction that cannot be described in a sentence does not belong in the planner.

03 / No new faults

The arithmetic is done in advance

Raising a quiet recording by eleven decibels raises its peaks and its DC offset by eleven decibels too. The planner works out where those would land and adds a limiter or a high-pass — saying in writing that the source was fine and the gain is what would have broken it.

04 / Measured, not assumed

It shows you what changed

The chart of a corrected copy carries the earlier reading dashed behind it, both plotted against whichever file is longer — so a trimmed ending shows as the old line outlasting the new one. Scaling them to a common width would hide the very thing worth seeing.

It reads the copy back

The finished file is analysed again from nothing, and the report you get is that measurement. When it lands off target the correction is recomputed and applied to the original again — never to the copy — so however many attempts it takes, the number of lossy encodes stays at one.

Room tone is copied from a quiet stretch of the file itself wherever there is one, rather than padded with digital zeroes — because a target asking for a quiet ending is asking for the room, not for nothing.

-60-40-20dashed: before the correctionsilence at 00:00silence at 00:24silence at 00:25silence at 00:34silence at 00:36silence at 00:48silence at 01:11silence at 01:12silence at 01:38silence at 01:50silence at 01:52silence at 01:59silence at 02:04silence at 02:05silence at 02:15silence at 02:16silence at 02:18silence at 02:27silence at 02:29silence at 02:30silence at 02:30silence at 02:34silence at 02:42silence at 02:47silence at 02:52silence at 02:55silence at 02:56silence at 02:59silence at 03:02silence at 03:09silence at 03:09silence at 03:12silence at 03:15silence at 03:16silence at 04:54silence at 04:59silence at 05:06silence at 05:12silence at 05:15silence at 05:19silence at 05:56silence at 05:57silence at 06:00silence at 06:03silence at 06:37silence at 07:54silence at 07:57silence at 08:00silence at 08:04silence at 08:10silence at 08:22silence at 08:36silence at 08:40silence at 08:45silence at 09:09silence at 09:21silence at 09:45silence at 09:54silence at 09:57silence at 10:01silence at 10:18silence at 10:39silence at 11:10silence at 11:12silence at 11:12silence at 11:16silence at 11:25silence at 11:42silence at 11:42silence at 11:45silence at 11:45silence at 11:52silence at 12:10silence at 12:41silence at 12:46silence at 12:47silence at 12:50silence at 12:52silence at 12:53silence at 12:54silence at 12:56silence at 14:41silence at 14:45silence at 15:14silence at 15:25silence at 15:31silence at 15:48silence at 15:49silence at 16:15silence at 17:09silence at 17:11silence at 17:24silence at 17:25silence at 17:35silence at 17:36silence at 17:52silence at 18:04silence at 18:05silence at 18:16silence at 18:22silence at 18:37silence at 18:46silence at 18:52silence at 18:57silence at 19:12silence at 19:19silence at 19:20silence at 19:26silence at 19:33silence at 19:37silence at 19:56silence at 19:58silence at 20:00silence at 20:01silence at 20:01silence at 20:11silence at 20:14silence at 20:14silence at 20:20silence at 20:21silence at 20:23silence at 20:24silence at 20:27silence at 20:29silence at 20:39silence at 21:35silence at 21:50silence at 21:52silence at 21:55silence at 21:55silence at 22:12silence at 22:35silence at 22:36silence at 22:57silence at 23:06silence at 23:30silence at 23:30silence at 23:31silence at 23:33silence at 23:43silence at 23:52silence at 24:11silence at 24:13silence at 24:19silence at 24:55silence at 24:56silence at 30:08silence at 30:20silence at 30:30silence at 30:52silence at 30:5500:0010:0020:0030:0033:25
A thirty-three minute spoken-word recording, corrected. The dashed line is the reading taken before: nine decibels quieter, with a peak six decibels above full scale at 30:46 and a second of clipping under it. The solid line is the corrected copy, measured from scratch — and because the plan already had to touch the audio, moving the index to the front of the container came along as one more flag rather than a second pass. Nothing identifying was needed to draw this; a loudness curve says nothing about what was said.

One correction costs nothing at all

Every other fix re-encodes, and the plan says so in a caveat. Moving an MP4's index in front of its media does not: nothing about the sound or the picture changes, only the order of the boxes in the container. Every stream is copied through — no filter, no encoder — so the corrected copy holds exactly the media the original did, and a test proves it by comparing a hash of the decoded video before and after.

What it will not do

The flashing check counts large frame-to-frame changes in average luminance and flags any second holding three or more, which is where WCAG draws its general flash threshold. It does no spatial analysis, does not measure how much of the screen changed, and knows nothing about the separate red-flash rule. It finds the passages worth looking at with human eyes. Silence from it is not a pass, and every report it appears in says so.

Noise reduction is deliberately absent. A noise floor above the target is reported and left alone, because the answer is to treat the room or re-record — and a tool that quietly ran a denoiser over somebody's audiobook would be doing something its owner did not ask for.

Under it

One decode per stream

A second read of a two-hour audiobook costs a minute of somebody's afternoon, and measurements taken in separate passes can disagree about where a moment is.

The audio filter chain: ebur128, aphasemeter, ametadata, silencedetect and astats, run as one pass, with metadata printed to stdout and events to stderr. ebur128 loudness, peak aphasemeter stereo phase ametadata print silencedetect gaps astats levels, channels one pass over the file stdout — metadata every 100 ms, kept as a one-second timeline stderr — silence events and the summaries printed at the end aphasemeter is inserted only for stereo: phase is a statement about the relationship between two channels, and on anything else it reports 1.0 and means nothing by it.

Both pipes are read at once

Metadata goes to stdout and events to stderr, and a filled pipe that nobody is draining is a deadlock — one that looks exactly like a slow file. Draining stderr in a thread is not a detail; it is the difference between a tool that works on a long file and one that appears to hang on it.

Kept to the second

ebur128 reports every 100 ms. The timeline holds the loudest momentary and short-term value in each second, because the report quotes seconds — so a ten-hour audiobook costs 36,000 rows instead of 360,000, and nothing is lost that would have survived to the page.

Two things earn a second pass

Locating peaks and clipping, and only when the whole-file peak says a peak problem can exist at all. A file whose loudest sample sits below the clipping threshold has no clipped samples — that is arithmetic, not an estimate — so the common case is answered without reading the audio again.

loudnorm's own measurement, because it will not accept another filter's figures; it wants its own threshold and offset.

The picture has its own pass, built to order

blackdetect → freezedetect → idet → signalstats, and it runs only when the target actually asks a question about the picture. Decoding a ninety-minute film to count black frames is minutes of somebody's time, and a podcast profile has no reason to spend them.

It also carries only the filters the target's rules actually read, because those four are not equally priced. Timed on a 1080p60 file, the whole chain takes about as long as the file itself — and more than half of that is one filter.

Interlacing is measured, not believed

Reading field_order from the container and believing it passes a 3:2-pulldown file that declares itself progressive — which is most telecined material there is. Running ffmpeg's idet and believing that is worse: on progressive material with hard vertical edges and fast motion it calls the majority of frames interlaced.

What separates them is that real interlacing is overwhelmingly one field order and the false positives are mixed, because they are noise rather than field dominance. Both a share test and a dominance test have to pass; material that passes one and fails the other is reported as inconclusive, which is the truthful answer and not a verdict.

ffmpeg is probed, not trusted

Slim distribution builds omit ebur128, and the failure otherwise surfaces halfway through an analysis as a filtergraph error that says nothing about why. The tool looks for the filter, not for a binary with the right name.

What each filter costs

Measured on one 8-core laptop, against a three-minute 1080p60 file. The right-hand column is what that filter adds per minute of video at that size; a smaller picture costs proportionally less, because these filters work per pixel.

FilterWhat it answersPer minute of 1080p60
blackdetectBlack frames, and where2.4 s
freezedetectFrozen frames, and where4.2 s
signalstatsAverage luminance, for the flash screening17 s
idetInterlacing and telecine27 s
decoding— unavoidable3.6 s

So a ninety-minute feature checked against every picture question is roughly eighty minutes of work, and the interlacing detector alone is half of it. Against a broadcast loudness target, which asks nothing about the picture, it is nothing at all. On the three-minute file above, an Instagram check — which needs black frames and the flash screening, but not fields — took 62 seconds instead of 176.

That choice is offered rather than assumed. The selective pass is the default; --picture full measures everything the picture can be asked whatever the target wants, and the report grows an “also measured in the picture, against nothing” section for the answers no rule read — because “the target does not ask” and “the file is fine” are different sentences, and somebody handing over a master may want both. Either way the tool says what it is about to cost before it starts, and the window puts both times on the dropdown.

Two accelerations that do nothing

Both were tried and both were measured. Filter threading (-filter_threads 8) came back at 28.6 s against a 27.1 s baseline. Hardware decoding (-hwaccel videotoolbox) came back at 27.7 s. Neither is an improvement, and the reason is visible in the table: decoding is 4% of the work. The cost is in the filters, and they run on one core.

The first attempt at measuring them reported 0.03 s, which would have been a fifty-fold speedup and was in fact a command that had failed instantly. A number that good is a bug report.

One that works, and where it stops

Reading the picture at half width takes 15.3 s instead of 27.1; at quarter width, 7.9 s. Horizontally only — blending adjacent lines is exactly what idet compares, so vertical scaling is never offered at all.

Across seven test files, black, frozen and flashing findings were identical at full, half and quarter width. Average luminance is the mean of the same pixels either way, so the flash screening is unaffected by construction rather than by luck.

The field checks are the exception, and it is refused rather than degraded. On a near-static picture with one small moving element, full width reports progressive and half width reports that it cannot tell. An inconclusive answer is not a cheaper answer, so asking for both together is an error with a sentence attached, not a quiet loss of confidence.

Running it

Two commands, no packages

Python 3.10 or newer, standard library only. ffmpeg and ffprobe are located at runtime rather than bundled, and their absence is reported as a sentence with the install line for your platform.

# something to double-click: a .app, a .pyz, a .desktop
./build.sh

The build bundles neither Python nor ffmpeg, and that is a decision rather than an omission. Bundling Python would mean a build-time dependency on a tool whose whole claim is that it needs nothing installed; bundling ffmpeg means eighty megabytes and somebody else's licensing decision. So the bundle removes the terminal, not the prerequisites — it looks for both, and when one is missing it says which, and prints the command that installs it.

# the window — a local server and a page in your browser
python3 app.py

# the command line
python3 preflight.py check finished.mp3 --target acx
python3 preflight.py batch chapters/    --target acx
python3 preflight.py batch chapters/    --target acx --fix
python3 preflight.py fix   finished.mp3 --target acx --dry-run
python3 preflight.py fix   finished.mp3 --target acx

check exits 0 when the file passes, 1 when it fails, and 2 when the tool itself could not run — so it drops into a build script without parsing anything.

The window listens on 127.0.0.1 with a token minted at startup. Your file is never uploaded: a browser cannot tell a local program where a dropped file lives, so the page asks the tool to open your system's own file dialog and the path never leaves the machine. That is not a workaround — it is the point.

The wider workshop

Five tools, one local-first practice.

Move between guided-meditation authoring, measured speech, lyric video, delivery validation, and the small field tools that support all of them.

If a number looks wrong

Say so

A threshold that has moved, a target that has changed its rules, a measurement that disagrees with your own meter — all of those are worth an issue. The profiles carry the month they were read precisely so that they can be argued with.