A CCTV control-room AI operator with zero fine-tuning: define surveillance events in plain language and Llama 4's multimodal vision detects and reports them.
The decisive design call looks like the refusal to fine-tune. Surveillance requirements differ site to site, so anything that assumes you will first collect and label footage never reaches deployment. Letting operators write the events they care about in plain language turns a model development project into a configuration task someone on site can do. It is also the shortest path to demonstrating the host's newly released multimodal model.
Summaries are written by this site. Project pages include demo videos, screenshots and the team's own write-up (videos may autoplay).