ContentGuard
A local video player that skips scenes you don't want to watch
Pick a video file. ContentGuard scans it once — a few minutes for a feature-length film — and builds a list of timestamped regions containing explicit content. Hit play, and MPV jumps past them automatically.
Entirely offline. No cloud API, no GPU, no account. That's a design requirement, not a bullet point, and it's the reason most of the rest of the project looks the way it does.
Why it has to run on your machine
There are services that do content filtering for streaming platforms, and none of them help if the file is sitting on your own drive. But the deeper problem with the cloud approach isn't coverage — it's that the only way for a remote service to find explicit content in your video is for your video to be uploaded to it.
Think about what that actually means. To use a hosted scanner, you send your personal media library to someone else's server so their model can look for nudity in it. The content most people would want filtered is precisely the content they'd least want leaving their machine, and the same goes for the fact of which films they're scanning and what they chose to skip. That's a genuinely bad trade, and it isn't fixable with a privacy policy.
So the whole thing runs locally. And once you've committed to local, you inherit the second constraint: it has to run on hardware people actually own. No GPU assumption, no expectation of a recent CPU. That single decision drives almost every technical choice below — you cannot run a heavy vision model across 130,000 frames of a feature film on an old laptop, so the architecture has to be built around not doing that.
Three signals, because one isn't enough
Detection uses an ensemble, and the composition of it is the part I find most interesting, because each member exists specifically to cover a hole in the others.
A part detector (nudenet) localises specific exposed body parts. Precise and reliable when a part is clearly visible in frame — and completely blind to a scene that's obviously explicit without any single part being cleanly in view. It's also measurably weaker on some anatomy than others.
Skin-fraction analysis (classical OpenCV, no neural network involved) measures what proportion of the frame is bare skin and flags anything above a threshold. This is the cheap signal, and it catches the wide shots and near-nudity that a part detector shrugs at, on the simple logic that those frames are mostly skin regardless of what's identifiable in them.
A whole-frame NSFW classifier (an optional Vision Transformer) scores how sexual a frame is as a scene, ignoring the question of which parts are visible. This is what catches sex scenes that the other two miss — where the composition is unmistakable but no individual detection fires.
A frame is flagged if any of the three fires, and each can be disabled independently.
Why OR-gating is the right call here
Requiring agreement between signals would raise precision and would be the wrong design, because the costs of the two error types are wildly lopsided for this specific job.
A false positive means the player skips ten seconds of an innocent scene. Mildly annoying. You notice, you shrug, you keep watching.
A false negative means the tool silently failed at the one thing it exists to do, at the exact moment it mattered most, in front of whoever you were watching with. There is no recovering from that — the content has already been seen. The whole product is worthless if it can't be trusted, and a filter that's right 95% of the time is functionally a filter you have to supervise, which is the same as not having one.
Given that asymmetry, OR-gating three complementary detectors and setting thresholds low is correct even though it demonstrably costs precision. The defaults deliberately over-flag.
Working around the CPU budget
The scan is a one-time cost paid up front and cached, so reopening a file is instant. Everything else is about making that one-time cost tolerable.
Sampling instead of exhaustive analysis. Frames are sampled at an interval rather than every frame decoded and classified. This is the single biggest lever on scan time and it's a direct trade against detection granularity.
Two-stage thresholding. A low sweep threshold flags candidate frames cheaply; a higher confirm threshold decides which candidates actually become written markers. The cheap pass does the volume, the stricter pass does the committing.
A minimum marker duration, so a single anomalous frame doesn't produce a skip.
Configurable everything. Sampling interval, both thresholds, skin trigger fraction, classifier threshold and its normalisation parameters — all in config.ini. The performance/accuracy dial is in the user's hands because the right setting genuinely depends on their hardware and their tolerance, and there is no single defensible default across both.
Degrading gracefully instead of demanding downloads
The app runs with the small bundled detector out of the box. If you drop the larger high-accuracy weights into models/, it auto-detects and uses them. If you add the NSFW classifier, that signal turns on; if the file isn't there, it's skipped silently rather than erroring.
This matters more than it sounds. Nobody wants to download several hundred megabytes of model weights to find out whether an application is worth using. So it works immediately at reduced accuracy, tells you in the console which models it found, and gets better as you feed it more. The classifier config is also deliberately generic — input size, normalisation, and output index are all parameters — so a different NSFW model can be dropped in without touching code. Model weights churn; hardcoding one is a maintenance bill.
The escape hatch is the best decision in the project
Markers can be exported to a plain text file, edited in any text editor, and imported back:
00:18:42 --> 00:19:15 (conf: 0.87)
00:44:03 --> 00:44:51 (conf: 0.92)
Delete a line to remove a false positive. Add a line to catch something the scanner missed. Reimport, and the edited list replaces the generated one and updates the cache, so your corrections persist for every future open. The parser accepts several timestamp formats, ignores comments and blank lines, merges overlapping regions, and reports invalid lines instead of failing on them.
I'd argue this is the most important feature in the application, and it comes from accepting something up front: this detector will be wrong, and no amount of model tuning changes that. Every automated content filter has an error rate. The ones people keep using are the ones where being wrong is a correctable inconvenience instead of a dead end. Exposing the model's output as an editable text file — the most boring, most universally editable format there is — turns "the scan missed a scene" from a reason to uninstall into a thirty-second fix.
It also means the confidence scores are visible, so a user reviewing markers can see which ones the system was unsure about and check those first.
What it doesn't do
Sexual content and nudity only. No violence, no drug use, no language, no other category of mature content. This is a deliberate scope line, not an oversight, but it's worth stating plainly because "content filter" implies more than this delivers.
No audio analysis at all. Explicit dialogue passes through untouched, and a scene that's explicit by sound while visually unremarkable is invisible to every signal in the system. Skipping is driven purely by pixels.
Brief content can fall through the sampling gap. Because frames are sampled at an interval and markers require a minimum duration, content shorter than that window can pass unflagged entirely. This is the sharpest real limitation — it isn't a tuning problem, it's structural to the sampling approach, and closing it means paying much more scan time.
Skin-fraction detection is not equally reliable across skin tones. Classical colour-space skin detection is well known to perform unevenly depending on skin tone and lighting, which means the effective aggressiveness of that signal varies with who is on screen and how the scene is lit. That's a real robustness problem inherited from the technique, and the honest mitigation is that it's one of three signals rather than the deciding one. It also fires on entirely innocent high-skin content — beaches, swimming, contact sport.
There's no measured accuracy figure, and there should be. The thresholds are tuned by observation, not against a labelled test set. I can describe what each signal is good and bad at qualitatively; I can't tell you the precision or recall at the default configuration, and for a tool whose value rests entirely on its false-negative rate, that's the number that matters most. Building a labelled evaluation set and sweeping the thresholds against it is the clear next step.
Stack
Python with a tkinter desktop UI, ONNX Runtime for CPU inference, nudenet for part detection, OpenCV for frame handling and skin analysis, an optional ViT NSFW classifier, FFmpeg for decoding, and libmpv driving playback with a polling skip engine layered over it. Cross-platform across Windows, macOS, and Linux.