AI proposes, you confirm
Wrong metadata is worse than none, so AI values stay suggestions until a human accepts them and every value says where it came from.
Autopilot tagging is how you make an archive worthless. Not because the models are bad at reading footage, they are decent and improving, but because a wrong value does more damage than a missing one. Missing metadata sends you back to watching clips: slow, annoying, safe. Wrong metadata sends your editor to the wrong clip at 11pm, and worse, it teaches them to stop trusting every other value in the index.
One bad tag poisons the search
Run a model over a 400-clip game day and let it write straight to the record. Say it gets 380 right. That sounds like a win until you ask which twenty are wrong, because nobody knows, including the model. The first time a filter for "touchdown" returns a punt, your editor stops filtering and starts scrubbing, and the whole pass, the 380 correct values included, becomes decoration. Search only works if a hit means something. A log that is 95 percent right and will not tell you which five percent is wrong is worse than an empty one; the empty one at least tells you where you stand.
This is the core reason I do not buy fully automatic logging, and it has nothing to do with model quality. It is an audit problem. Accuracy will keep climbing, and the audit problem will not move, because 99 percent right with no way to find the one percent is still a log you have to re-watch before you can trust it.
The loop: suggest, review, confirm
The fix is not a better model. The fix is a boundary. The model proposes, you confirm, and nothing becomes the record until it crosses that line. ClipLogger is built on exactly this rule: every AI output lands as a visibly pending suggestion, and nothing is written to sidecars, exports, search, or filenames until you accept it. Return accepts the focused suggestion, Delete rejects it, and a rejection is a recorded decision that a later re-run will not resurrect.
The speed argument for autopilot dies here, because review is not slow. Deciding a value from scratch means watching, recalling the vocabulary, typing. Confirming a proposed value is a glance and a keystroke, a second or two per field, and across 400 clips that is the difference between a logged day and a wish. The model does the typing. You keep the judgment.
Every value says where it came from
Confirmation is half the answer. The other half is provenance: every value records which engine produced it, a local model, your own API key, or a Rush cloud run, stamped at the moment it ran. A value from before tracking existed says "provenance unrecorded" instead of guessing.
This sounds like bookkeeping until the day it is not. Two seasons from now, somebody asks why the archive says number 72 scored in a clip where he plainly did not. With provenance you can see the value came from one cloud batch in October, check what else that batch touched, and re-run only those clips. Without it, you get to wonder about every value in the library, which is the poisoned-search problem again, wearing a nicer shirt. The badge sits on the clip itself, not in a settings page you have to remember exists, so the answer is wherever the question comes up.
"The AI tagged it" is not an answer
Picture the call. A client paid for a logged archive and their lawyer wants to know how a name got attached to a person in a sensitive clip. A broadcaster wants to know who verified the metadata on the highlight package before air. "The AI tagged it" is not a sentence you can say out loud in that room. "A model proposed it, a person confirmed it, and here is which engine made the proposal and when" is.
If you bill for logging, the log is your work product, and the confirm is your signature on it. Shipping unreviewed model output under your invoice is subletting your name to a system that will not be on the call when a value turns out wrong. Broadcast and legal both run on chains of accountability, and metadata is part of the deliverable now, whether the invoice says so or not.
Review fast without reviewing carelessly
This part holds in any tool. Review grouped, not clip by clip: fifty proposals for one field or one person move faster as one screen than as fifty visits. Accept in bulk only when the group is uniform, and the moment it is not, slow down for that group alone. Reject decisively; a maybe left pending is a decision you have scheduled for a worse time. And treat repeated wrong suggestions as feedback on your instructions: when the model keeps calling a whip pan an action, the fix is a sharper field definition, not patience.
The rule under all of it: no value enters the record without a human at least glancing at it. That is not distrust of models. It is respect for what the record is for.
A log you can hand to a client is a log where every value had a person say yes.