The attention problem
Cameras are cheap and can go almost anywhere. What is scarce is the attention needed to look at what they produce — and that, not sensing, is the real constraint on surveillance.
Cameras are cheap and can go almost anywhere. What is scarce is the attention needed to look at what they produce — and that, not sensing, is the real constraint on surveillance.
Surveillance rarely fails for lack of sensors. Cameras are inexpensive and can be placed almost anywhere. What is scarce is not imagery — it is the attention required to look at it.
Watching is a task humans perform poorly over long periods, and the research on this is not ambiguous. Sustained attention on a low-event task degrades measurably within tens of minutes. The irony is exact: good surveillance consists overwhelmingly of uneventful time, which is precisely the condition under which human vigilance is least reliable.
So an organisation faces an unattractive choice. Staff surveillance properly — rotating people frequently enough to maintain attention, which is expensive and consumes the same personnel needed elsewhere. Or staff it improperly, and hold coverage that exists on paper but not in practice.
Adding cameras without adding attention does not increase surveillance. It increases unreviewed footage.
An autonomous system does not become bored. It can hold a pattern for as long as endurance allows, and — more importantly — it can decide what deserves a person's attention rather than requiring a person to find it.
That inverts the workflow. Instead of a human watching a feed hoping to notice something, the system watches continuously and escalates exceptions. The human's attention is spent on judgement, which is what humans are actually good at, rather than on detection, which they are not.
Reporting what changed requires knowing what normal looks like — and normal is specific to a place and a time of day, not something that can be shipped in advance. The system has to build that understanding itself, from observation, and keep updating it as conditions and seasons shift.
This is genuinely hard, and it is where a lot of otherwise capable systems disappoint. A model that flags everything unusual will flag weather, wildlife, shadows and ordinary activity. A model tuned until it stops doing that often stops flagging what matters too.
We would argue that precision matters more than sensitivity in this application, which is not the conventional emphasis. The reasoning is behavioural rather than technical: a system that raises the alarm without cause trains its operators to ignore it, and an ignored system provides no coverage regardless of how good its detection is on paper.
So the metric we care about is not how much a system detects, but how much of what it reports is worth reporting — and whether the confidence it attaches to a report is honest enough to act on.
There is rarely enough communications capacity to send everything, which turns out to be a useful constraint rather than only a limitation. It forces the system to decide what is worth transmitting, which is the same discipline that makes its reports useful to a person.
Sending conclusions rather than raw video is better on every axis: less bandwidth, less exposure, and far less human time consumed at the other end. It also requires the system to be right about what matters, which is the harder engineering but the correct target.
ISR is the largest and most immediate application of autonomy in defence, and the attention economics are why. It is also why our work concentrates on perception quality and honest uncertainty rather than on platform performance figures — the limiting factor was never how long something can stay airborne.
Continue
We publish our positions openly, including the parts we haven't solved.