Eyes is an experiment in adaptive video inference for fixed-camera footage. It tests whether a cheap CPU activity scan can identify the most useful parts of a long video before sending footage to the much more expensive Marlin vision model.
Fixed-camera videos often contain long quiet periods and short bursts of useful activity. Processing every second with a GPU model can therefore waste compute. Eyes first scans the complete video on the CPU at 5 FPS and a resolution of 160 × 90. It measures frame difference, foreground area, and persistent scene change, then combines those signals into an activity score.
The resulting schedule divides the source timeline into active, context, and skipped regions. Only selected color clips are sent to Marlin at 2 FPS. The experiment compared three policies on six held-out MEVA videos:
- Eyes: content-adaptive active and context clips.
- Uniform: blind timeline sampling with the same selected duration and Marlin call count as Eyes.
- Full: the complete video timeline.
Four separate development videos were used to calibrate the activity policy. The six test videos were kept out of calibration and were not used to retune the policy after evaluation.
| Measurement | Eyes | Comparison |
|---|---|---|
| CPU scan time for all test videos | 14.87 s | Inexpensive relative to inference |
| Selected source footage | 1,052.1 of 1,801.0 s | 41.58% less than Full |
| Frames sent to Marlin | 2,104 | 43.26% fewer than Full |
| MEVA events covered | 124 of 222 (55.86%) | Uniform covered 119 (53.60%) |
| Marlin calls | 95 | Full required 36 |
| Measured GPU inference | 1,233.15 s | 55.60% higher than Full |
| Accounted end-to-end latency | 2,142.19 s | 13.52% slower than Full |
Eyes placed clips slightly better than Uniform at an exactly matched budget, covering five additional annotated events. However, this advantage was small and inconsistent: Eyes beat Uniform on only two of the six test videos.
The original efficiency hypothesis was only partially supported. Eyes reduced the amount of selected footage, but it did not reduce measured GPU time or end-to-end latency. Its selected windows were fragmented into many short clips, so repeated model generation and RPC overhead outweighed the saving from processing fewer frames. The activity policy also generalized unevenly across cameras.
The experiment's answer is that content-aware selection is useful, but selection and inference packaging must be designed together. The CPU scanner was fast and its signals carried useful information, yet the current MVP needs far fewer or batched Marlin calls and stronger cross-camera calibration. For this experiment, Full remains the faster measured policy, while Eyes remains a promising but unoptimized prototype.
The reported event figures measure temporal coverage of MEVA annotations. They do not establish Marlin's semantic accuracy because the manual semantic review is still pending.
Read the full experimental report for the complete method, measurements, failure analysis, limitations, and reproducibility details.
| Path | Purpose |
|---|---|
src/eyes/ |
Core manifest, scanning, calibration, scheduling, inference, and evaluation modules |
modal_apps/ |
Explicit Modal GPU entrypoints for validation and Marlin inference |
tests/ |
Unit and synthetic-video tests for the local pipeline |
configs/ |
Locked activity-scanner configuration |
data/ |
Local MEVA manifests and videos; excluded from version control |
artifacts/ |
Generated validation, scan, schedule, inference, and evaluation outputs |
docs/ |
Supporting notes, including Marlin model observations |
final_report.md |
Complete experimental results and conclusions |
requirements-local.txt |
Local CPU and Modal client dependencies |
requirements-modal.txt |
Pinned remote Marlin runtime dependencies |
The experiment uses NemoStation/Marlin-2B on Hugging Face.