Connected is not the same as working
Most camera monitoring answers one question: is the device reachable? That question is easy to answer and, on its own, close to useless. The failures that actually cost people footage are the ones where everything looks healthy:
- The stalled stream. The socket is open, the camera answers, and no frames have arrived for hours. Encoder faults, firmware bugs and half-open TCP connections all produce this, and nothing in a reachability check notices.
- The frozen image. Frames keep arriving, but they are identical — the camera is serving a stuck buffer. Detection sees a still life and reports nothing wrong.
- The dead recorder. The camera is fine, the analytics are fine, and the recording process has written zero bytes since Tuesday. Disk full, a crashed encoder subprocess, a permissions change after an update.
- The blind camera. Perfectly online, perfectly recording — a wall, because someone knocked it during a delivery or turned it deliberately.
Each of these fails silently. There is no error, no disconnect, no red light. The system is confidently doing nothing, and it will keep doing nothing until a human happens to look.
The three checks that matter
1. Liveness — when did the last frame arrive?
The single most valuable health signal is a timestamp: when did this camera last deliver a frame? A watchdog that compares that timestamp against a threshold catches the stalled stream that reachability checks miss entirely. It requires no extra network traffic, because the system is already receiving — or not receiving — the frames.
Note that this is a different state from a disconnect. A disconnect is an event the system is told about. A stall is the absence of events, which means it can only be found by something actively looking for silence.
2. Content — is the picture still a picture?
A frame arriving is not a frame worth having. Cheap scene-level checks catch a surprising number of real faults: a view that has gone almost entirely black, one that has become uniform (a covered lens), one whose brightness histogram has shifted radically and stayed there, or one that has not changed at all across a period where change is certain. None of these need a model; they need somebody to have decided they are worth checking.
3. Output — is anything actually being written?
Analytics health and recording health are separate. A system can be perceiving perfectly and recording nothing. The check is blunt and effective: has the recording for this camera grown on disk within the expected window? A recorder process that is alive but producing no bytes is one of the most consistently under-monitored failure modes in video systems, and it is trivially detectable once someone thinks to look at the file rather than the process.
Alerting on health without training people to ignore it
Health alerting has the same failure mode as detection alerting: too many messages and people stop reading them. A few rules keep the channel trustworthy.
- One alert per incident. If two different detectors notice the same outage — say a disconnect and a stall watchdog — they should resolve to one message, not two. Duplicate alerts for one event are how a health channel loses credibility first.
- Debounce before shouting. A grace period absorbs the switch reboots and brief network blips that are not incidents. Alerting instantly on every dropped packet guarantees noise.
- Always send the recovery. A "camera restored" message is what lets someone stop chasing. Systems that alert on failure and stay silent on recovery leave everyone guessing, and the guess is usually wrong.
- Only claim recovery from a failure you actually reported. A restored message for an outage nobody was told about is confusing rather than reassuring.
- Make the current state visible somewhere. Alerts are for transitions; a status view answers "is everything healthy right now", which is the question asked before a holiday weekend rather than during an incident.
Recovery is part of monitoring
Detecting a stall is half the work. A system that notices its recorder has died and does nothing about it has converted a silent failure into a loud one, which is progress but not a solution. The stronger design attempts recovery first — restart the stream, restart the recorder subprocess, reconnect with backoff — and alerts a human when recovery fails or when it has happened often enough to indicate a real fault.
The discipline this belongs to is the same one described in Designing Reliable Perception Systems: every stateful component needs an explicit reset path, and long-run behaviour has to be measured rather than assumed. Soak testing is where these faults surface, because most of them need hours or days to appear at all.
What to monitor, in order of value
- Time since the last frame, per camera.
- Recording file growth, per camera, within an expected window.
- Reconnect frequency — a camera reconnecting hourly is failing, slowly.
- Disk free space and the retention policy's actual behaviour, not its configuration.
- Scene-level sanity: black, uniform, or unchanging views.
- Process memory over days, to catch the leak that kills the system on day nine.
- Time synchronisation — a camera with the wrong clock produces evidence with the wrong timestamp, which is worse than no evidence. See Video Evidence and Chain of Custody.
Where ManasaView fits
ManasaView treats camera liveness as a first-class signal rather than an afterthought. Two independent detectors cover the two shapes of failure — a reconnect cycle for a camera that disconnects, and a no-frames watchdog for one that stays connected and goes silent — and they are deliberately reconciled so a single real outage produces a single message, with a restored message when it comes back. Recording health is checked separately from perception health, because a live system that has quietly stopped writing is the failure that costs you the footage. See it on your own cameras.
Frequently asked questions
Why does a camera show as online when it is not recording?
Because being online and sending video are different things. The network socket can stay open while the camera sends no frames — after a firmware fault, an encoder hang, or a stream the camera has quietly stopped producing. A monitor that only checks reachability will report everything as fine. Only something that checks when the last frame actually arrived can see this failure.
What is the difference between a camera being down and being degraded?
Down means the connection broke and the system knows it — a disconnect event fired. Degraded means the connection is intact but the data is not arriving or is unusable: no frames, a frozen image, a blacked-out view. Degraded is the more dangerous of the two because nothing in the system has failed loudly.
How quickly should a camera outage be alerted?
Fast enough to act on, but debounced enough to survive a brief network blip. A short grace period before alerting — tens of seconds rather than milliseconds — avoids a storm of messages every time a switch reboots, while still surfacing a real outage within a useful window.
Should the system alert once per outage or keep repeating?
Once per real incident, followed by a restored message when it recovers. A repeating alert for an ongoing known outage trains people to mute the channel, which then hides the next genuine alert. The restored message matters as much as the original — without it, nobody knows whether to keep chasing.
Does camera health monitoring detect tampering?
Partly. A cut cable or a powered-down camera looks like an outage and will be caught. A camera that has been covered, sprayed or turned to face a wall stays perfectly online and keeps sending frames, so it needs a different check — scene-level tests for a view that has gone dark, uniform, or abruptly different from its own history.
← Back to Knowledge