The count, from my own bench: every clip I've shot or streamed was searchable by exactly one thing. Its filename. Anything silent, anything visual, anything nobody said out loud: gone. Search what was said, lose what was shown Transcript pipelines only index speech. The take where the red jacket walks left, the quiet screen moment, the beat the chat ignored: none of it survives, because nobody said it out loud. This isn't magic, and it isn't cloud. Everything runs on my own hardware, and the first release candidate proves plumbing before it promises understanding. Query without re-seeing The fix is a four-stage spine. Sample the clip on scene cuts and motion, not every frame. Observe it once with a local vision model and write versioned sidecars. Index those sidecars into full-text search plus embeddings. Then query: ranked time spans, no vision re-run, ever. A five-second test clip ingests in 0.38 seconds. A gold harness scores recall on every change, so quality is a number, not a vibe. The release candidate is yours ODVM v0.1.0-rc1 is up for ARCHITECT members: pip install, four commands, your footage stays on your disk. Break it and tell me where. The model is replaceable. The memory you build is the asset. Local first. Query without re-seeing. The discipline beats the tool. Full deep-dive with diagrams on aris-space: https://aris-space.com/documents/tools-and-plugins/odvm-video-memory //A<3