Mohammed Mutahar
← Back to Projects

AffectSync

Generative models learn what a joke looks like, never whether whether it made anyone laugh. AffectSync aims to bridge that gap by exporting emotion-aware training data.

Computer VisionEmotion RecognitionMultimodal DataDataset EngineeringReal-Time SyncSpeech-to-TextDeep LearningPythonOpenCVWhisper
AffectSync

AffectSync

Recording what a video makes you feel, frame by frame

AffectSync watches you watch something. While a video plays, it captures your face through the webcam, classifies your emotional expression about ten times a second, transcribes the video's audio, and then stitches all of it together into one annotated JSON file where every emotional reading is attached to the exact line that was being spoken when you had it.

It's a data collection tool.


Why I built it

Generative models for creative content are trained on what content exists, never on what content works.

Feed a model a thousand stand-up specials and it learns the architecture of a joke; the setup, the misdirect, the timing of the turn. What it never learns is whether anyone laughed. The laugh is the only part that actually validates the joke, and it's the part that nobody writes down. Same story with a scene that's supposed to land emotionally, or an ad that's supposed to make you feel something before it asks you for money. The transcript survives. The response evaporates.

I wanted to know what it would take to capture the response and keep it, in a form a model could actually train on. Not sentiment analysis of the script, not a post-hoc survey where someone rates the clip out of ten, a measurement taken at the moment, aligned to the second, so you can point at a specific line and say: this is what it did to a person.

AffectSync is that capture layer. It's deliberately infrastructure rather than a finished research result, because the missing piece in this whole idea isn't the model, it's the dataset.


How it works

Two streams run at once and have to agree about time.

The video stream plays the clip and reports its own playback position. The viewer stream pulls webcam frames, finds the largest face in each one, crops it, and runs it through an emotion classifier that returns seven labels with confidence scores. Both of them read from a single shared monotonic clock, which is the only reason any of the output means anything.

Once playback ends, post-processing takes over: audio is extracted to 16kHz mono, Whisper transcribes it into time-aligned segments, and everything gets merged into a session file.

{
  "timestamp_ms": 1040,
  "emotion": "happy",
  "confidence": 0.87,
  "transcript_segment": "So I was at the DMV the other day...",
  "segment_position": "mid"
}

That's the shape of the thing I was after. A measured response, a timestamp, and the words responsible for it.

Finding the moments that matter

A raw timeline is mostly noise. Ten readings a second across a two-minute clip is over a thousand rows, and the overwhelming majority of them say neutral because most of the time, watching a video, your face is doing nothing at all. Training on that teaches a model that comedy is 94% neutral.

So the pipeline runs a peak detector over the timeline and marks three kinds of events:

  • Onset — a sharp break out of a neutral baseline into something else. This is the reaction moment: the exact frame where the thing landed.
  • Sustained — a non-neutral emotion held for at least half a second. Separates a genuine held reaction from a single-frame classifier twitch.
  • Spike — one frame of unusually high confidence, above 0.80.

Peaks carry their transcript segment with them, which means the useful artifact at the end of a session is a short list of "this line caused this reaction" pairs.


Problems worth talking about

Two clocks are one clock too many. The first thing that breaks in a system like this is time. Video playback position and webcam capture time drift apart immediately if you let each one keep its own accounting, and a hundred milliseconds of drift is enough to attach a laugh to the wrong line, which is worse than no data, because it's confidently wrong. I ended up with a single SessionTimer as the sole source of truth. Nothing in the pipeline is allowed to ask the system what time it is.

Seeking is a trap. My instinct was to seek the video to whatever position the clock reported on every tick. On H.264, seeking is roughly a hundred times slower than a sequential read, because the decoder has to walk back to the nearest keyframe and rebuild forward. The fix was to read frames sequentially and reserve seeking for repositioning after a pause. Obvious in hindsight; not obvious while writing it.

The threading I chose not to write. Playback and inference share one thread. The main loop ticks at 30 FPS for smooth display and emotion inference fires on every third tick, giving 10 FPS of readings. A proper implementation would decouple these into separate threads with a queue between them, and I know that. I skipped it on purpose — the real bottleneck is CPU inference, and adding concurrency to a pipeline whose timing correctness is its entire value proposition seemed like the wrong second problem to take on. It's the clearest deliberate MVP compromise in the codebase.

Testing a thing that needs a face. Every component takes its dependencies as constructor arguments with real instances as defaults, which was a small discipline that paid for itself completely. The whole test suite runs with no webcam attached, no video file present, and no model weights downloaded — camera, video reader, and classifier are all mocked at the seam. I could develop the sync logic on a machine with the lid closed.


What it can't do

Worth being straight about this, because the honest limits are more interesting than the feature list.

A facial expression is not an emotion. It's a proxy, and a leaky one. People suppress smiles. Some people find something hilarious and their face barely moves. Amusement and surprise are genuinely hard to tell apart from geometry alone, and FER-family models inherit every bias in their training distribution. What AffectSync records is visible facial affect, and any downstream use has to treat it as a noisy correlate of the thing we actually care about rather than a direct reading of it. Multi-viewer aggregation is the honest answer here — one person's face is anecdote, thirty faces converging on the same frame is a signal.

It assumes a cooperative environment. Decent lighting, one face, viewer looking roughly at the camera. It is not a field instrument.

Everything else is scope. One viewer per session, clips capped at five minutes, CPU inference downsampled to 10 FPS, and Whisper's base model by default — which is fine on clean single-speaker audio and starts falling apart on overlapping dialogue or strong accents. Moving up to small or medium fixes the transcripts and costs you transcription time.


Where it goes next

Multi-viewer sessions are the priority, for the reason above: aggregating across viewers turns a subjective read into something closer to a consensus reaction curve, and a consensus curve is what you'd actually want to condition a model on. GPU inference would let me stop downsampling and capture reactions at full frame rate, which matters more than it sounds like — comedic timing lives in the 100ms range.

And then the part this was all built for: taking a corpus of these sessions and using it as a preference or reward signal, so a model generating a script has some notion of which version of a line actually worked on a human being.


Stack

Python 3.10+, OpenCV for capture and face detection, DeepFace with FER+ weights on a TensorFlow backend for classification, openai-whisper for local transcription, ffmpeg for audio extraction, pytest for a fully hardware-mocked test suite.

Code: github.com/AsRumi/AffectSync · MIT

Related Projects