The pitch is one sentence: take a selfie with your webcam, watch it shatter into a shuffled grid, and put it back together by grabbing tiles with your hand in the air. No download, no signup, no controller. Open a tab and start waving.
That is Snapzule. The interesting part is not that it works — MediaPipe does the hard vision problem for you and does it well. The interesting part is everything between "MediaPipe gives me 21 points" and "this feels like picking something up."
What MediaPipe actually hands you
MediaPipe Hands runs a hand landmark model in the browser and returns, per detected hand, 21 3D landmarks — four points per finger plus a wrist. Coordinates are normalized to the frame, so 0,0 is the top-left corner and 1,1 is the bottom-right, regardless of camera resolution.
That is all it gives you. It does not tell you the hand is making a fist. It does not tell you the user is pointing at a tile. There is no gesture field. Every gesture in Snapzule — palm, fist, peace sign, thumbs up — is something I had to define in terms of where 21 dots are relative to each other.
The question I had going in was whether that meant training a classifier. It did not.
Gestures are geometry, not machine learning
My first instinct was a small neural net on top of the landmarks. Collect samples of each gesture, train a classifier, ship the weights. It is the obvious move and it is the wrong one for this problem, because these gestures have closed-form definitions.
Consider a fist. A finger is curled when its tip is closer to the wrist than its knuckle is. That is it. That is the whole test.
// A finger is extended if the fingertip sits further from the wrist
// than the knuckle does. Curl every finger and you have a fist.
function isExtended(landmarks, tipIdx, pipIdx) {
const wrist = landmarks[0];
return dist(landmarks[tipIdx], wrist) > dist(landmarks[pipIdx], wrist);
}
const FINGERS = [[8, 6], [12, 10], [16, 14], [20, 18]]; // index..pinky
const extended = FINGERS.filter(([tip, pip]) => isExtended(lm, tip, pip)).length;
const isFist = extended === 0;
const isPalm = extended === 4;
const isPeace = extended === 2 && isExtended(lm, 8, 6) && isExtended(lm, 12, 10);
Distance comparisons. No training data, no model file, no inference cost, no accuracy cliff when someone's hand is a different size or skin tone than the training set. It is scale-invariant and rotation-tolerant for free, because it only ever compares two distances against each other, never against a threshold in pixels.
TIP — The rule I took away from this
Reach for a model when the decision boundary is genuinely hard to describe. If you can write the rule down in one sentence — "all fingertips are closer to the wrist than their knuckles" — write the rule down. A model is a way to learn a function you cannot express. This one you can express.
The counter-argument is real: geometric rules are brittle at odd hand angles where a learned model would generalize. In practice, for four coarse gestures held deliberately in front of a webcam, the rules win on every axis that matters — zero payload, zero warm-up, fully explainable when it misfires.
The latency budget is the entire design constraint
Sixty frames per second means every frame has under 16 milliseconds to do everything. In that window Snapzule must:
- Pull a frame from the camera
- Run MediaPipe inference on it
- Classify the gesture from the landmarks
- Smooth the cursor position
- Update puzzle state
- Let React re-render whatever changed
MediaPipe takes the majority of it. Everything else is fighting over what's left, which is why gesture classification being ten distance calculations instead of a forward pass matters more than it looks like it should.
The thing that actually gets you, though, is not the average frame. It is the bad frame.
Two problems nobody warns you about
Raw landmarks jitter. Even with a perfectly still hand, MediaPipe's output wobbles by a few pixels frame to frame. Map that straight to a cursor and the cursor vibrates. Users read a vibrating cursor as "this is broken," not "this is a noisy sensor."
The fix is an exponential moving average — each new position is blended with the last one instead of replacing it:
// alpha near 0 = heavy smoothing, sluggish. Near 1 = responsive, jittery.
smoothed.x = alpha * raw.x + (1 - alpha) * smoothed.x;
smoothed.y = alpha * raw.y + (1 - alpha) * smoothed.y;
One line, one tuning knob, and the knob is not optional — it is the feel of the whole game. Too much smoothing and the cursor lags behind your hand like it is underwater. Too little and it buzzes. There is no correct value you can derive; you sit there moving your hand and adjusting the number until it stops feeling wrong. Every real-time input system has one of these, and it always has to be tuned against an actual human hand rather than reasoned about.
Tracking drops mid-drag. You are holding a tile, you move your hand slightly out of frame or the lighting shifts, and for three frames MediaPipe returns nothing. Handle that naively — no hand, no grip — and the tile drops. It is infuriating, and it is not the user's fault.
Snapzule keeps a 300ms grace window. Lose the hand while dragging and the grip survives; if tracking comes back inside 300ms, nothing happened. Past that, the tile drops. Three hundred milliseconds is long enough to cover essentially every real dropout and short enough that a deliberate release still feels instant.
That single window is responsible for more of the "this feels solid" impression than any other line in the project.
Multiple trackers, one MediaPipe
Snapzule detects gestures in several places — a hook for the palm-hold that triggers the selfie, one for grab-and-drag, one for thumbs-up on the win screen. The naive structure is one MediaPipe instance per hook.
Do not do this. Running several instances against the same video element means several inferences per frame on the same pixels, and the frame budget was already tight with one. Sends have to be serialized through a single tracker with the consumers subscribing to its output. Same landmarks, distributed to whoever cares.
Things I did not expect to matter
Lighting is the real accuracy variable. Not hand size, not skin tone, not distance. Backlit users — a window behind them — are the hard case, because the camera exposes for the window and the hand becomes a silhouette. There is no clever fix from inside the browser; the game just has to fail visibly rather than silently, so the user knows to move.
No audio files. Every sound in Snapzule — the snap, the tile blip, the win arpeggio — is synthesized live with Web Audio oscillators. It started as a way to avoid shipping assets and stayed because a short oscillator burst is genuinely the right tool for a UI blip. Zero bytes over the wire, zero load delay, and pitch becomes a parameter you can vary per tile.
Feedback is what sells the gesture. A tile snapping home fires a synthesized blip, a "+1" pop animation, and a colour shift. Cut those and the exact same tracking code feels unresponsive — users start jabbing at the air because nothing told them the last action registered. Air gestures have no haptics and no physical stop, so every confirmation has to be manufactured.
What it adds up to
- Five grid sizes, 2×2 through 6×6
- Normal, time attack, daily challenge, and upload-your-own-photo modes
- Deterministic scrambles from a seeded PRNG, so a shared URL gives everyone the identical puzzle
- Daily global leaderboard with ghost racing against the current number one
- Server-verified scores — its own post, because trusting the client here was never an option
All of it built on web primitives. React, CSS Grid, the Canvas API, Web Audio, and getUserMedia. No game engine.
Play it at snapzule.outshorts.in — you need a webcam and about forty seconds.