The word "bot" makes people picture something clever. On random video chat it is almost always the dullest thing possible: a recording of a real person, played on a loop, to sell you something. That is good news, because a video file cannot do several things a live human does without thinking, and once you know what they are you can tell in about two seconds.
Three things a recording cannot fake
These are worth learning in order, because they get harder to fake as you go down the list.
- It plays the same few seconds twice. Watch for about fifteen seconds and you will often see the same head turn, the same smile, the same glance off camera, in the same order. A person cannot repeat themselves frame for frame. A loop does nothing else.
- Nothing moves at all. The other version is a single photograph propped in front of the lens. It looks like a person until you notice that in thirty seconds not one thing on screen has changed, not a blink, not a shoulder, not the light.
- The caption never moves. Ad bots burn a handle or a link into the video, so it sits on exactly the same pixels while everything behind it moves. Real text held up on real paper drifts, because hands drift.
The two-second check
Ask for something specific and unpredictable, and make it physical rather than conversational. Hold up three fingers. Touch your left ear. Say a number out loud. A recording cannot answer, and the people running them are rarely watching in real time.
The reason this works better than asking a question is that a scripted chat message can answer a question. Nothing in a video file can put three fingers up on request.
Why software is better at this than you are
All three tells are measurable, and none of them needs to know anything about the person on screen. Comparing a frame against the last few seconds of frames catches a loop. Comparing a frame against the previous one catches a still. Watching one region of the picture stay identical while the rest changes catches a burned-in caption.
That is pure pixel arithmetic on a tiny downscaled image. There is no model to download and no face to recognise, which is why this particular check can run for everybody from the very first call, on any machine, at no meaningful cost.
It also means the thing being compared describes video content, not a person. Two strangers who happen to look alike do not collide; the same advertising reel played to a thousand different people does, which is exactly what makes a shared verdict useful.