On-device performance

What can your Mac do?
Local logging, un-hyped.

ClipLogger's AI runs 100% on your Mac — which means the speed is bound by your hardware, not our servers. Here's an honest look at how long it takes to log a clip on the machine you already own. No uploads, no cloud required.

Your Mac, per clip

Estimated time to fully log one rich clip with the default on-device model. Pick your Mac to highlight it — or look yours up if it's not in the list.

I've got an
with 32 GB of memory.
Pick your chip to see your estimate.
Not sure which chip? Apple menu → About This Mac.
MacChipUnified memoryBandwidth~ Per clipA 200-clip shoot
Mac StudiodesktopM3 Ultra96–512 GB819 GB/s~2–4 min~8 hrs
Mac Pro · Mac StudiodesktopM2 Ultra64–192 GB800 GB/s~2–4 min~8 hrs
MacBook Pro · Studiolaptop / desktopM4 Max36–128 GB410 GB/s~3–6 min~13 hrs
MacBook ProlaptopM3 Max36–128 GB400 GB/s~4–7 min~16 hrs
MacBook Pro · Studiolaptop / desktopM1 / M2 Max32–96 GB400 GB/s~5–9 min~20 hrs
MacBook Pro · Mac minilaptop / desktopM4 Pro24–64 GB †273 GB/s~6–11 min~25 hrs
MacBook Pro · Mac minilaptop / desktopM2 / M3 Pro18–36 GB †150–200 GB/s~8–14 min~33 hrs
MacBook Air · iMac · minieveryday MacsM1–M4 (base)16–32 GB †100–120 GB/s~14–25 min~2 days
† The default model likes ≥ 32 GB of unified memory to run at full quality. Lighter Macs still log locally — just slower — or hand the job to Rush.
◇ When the clock's against you

This is exactly why Rush exists.

A 200-clip shoot is an afternoon-to-overnight job on your own Mac. Hand it to Rush and it runs parallel across many GPUs — the whole directory logged and named in under 5 minutes, same models, nothing kept. Local for the quiet nights; Rush for the deadline.

The math behind one clip

Where those minutes come from — the real weight of a single clip.

◷ One clip, by the numbers
Frames the model reads~12–20 imgs
Tokens read in (frames + prompt)~40,000
Tokens written out (the log)< 400
Model class~30B VLM · 4-bit

The "robust" default — a ~30-billion-parameter vision model at 4-bit (~21 GB in memory). Frames-per-clip is growing as the models get hungrier, so treat this as today's floor, not its ceiling.

◇ Where the time goes
1

Reading the frames the big one

Pushing ~40,000 image tokens through the model. This is most of the wall-clock — it scales with your GPU's compute.

2

Writing the log

Generating <400 tokens of tags, names and notes. Fast — it rides your memory bandwidth.

Two different bottlenecks — which is why a chip's GPU cores and memory bandwidth both matter, and why more RAM just lets the model fit at all.

About these numbers. These are modeled estimates, not lab benchmarks — real times swing with clip resolution, how many frames the model samples, your quant level, and thermal headroom (laptops throttle under sustained load; desktops don't). They're anchored to published Apple-Silicon inference data and a real Qwen-VL-72B run on an M3 Ultra (~3.5 min for a single full-res image). Treat them as order-of-magnitude, not a stopwatch.