Recording and detection want different things

Most IP cameras and recorders produce at least two video streams. The main stream is high resolution and high bitrate, and it is what you record and zoom into after an incident. The sub-stream is smaller and lighter, and it was originally meant for viewing on phones and multi-camera grids.

AI detection normally reads the sub-stream. That surprises people who expect more resolution to mean better detection. In practice a detection model works on a fixed-size input picture anyway. Decoding a 4K stream only to shrink it again costs processing power that is better spent watching more cameras, more reliably. What matters is that the sub-stream is set up properly, and out of the box it very often is not.

The settings we recommend

These are the targets we use when we set up a camera for a CleverCam hub:

Setting Target Why
Codec H.264 The hub decodes H.264 in hardware. H.265 has to be decoded in software, which costs processing power and memory.
Resolution About 704 × 480 Enough detail for people at a useful distance, without wasted pixels.
Shape Widescreen (16:9) preferred Match the shape of the camera's sensor so people are not stretched or squashed.
Frame rate About 8 fps (6 to 10 is fine) Detection does not need 25 frames per second. Extra frames cost decoding and bandwidth.
Bitrate 1 to 2 Mbps Enough to keep night-time pictures clean without flooding the network.
Keyframe interval About the same as the frame rate One full picture roughly every second, so the stream can be picked up quickly and live view starts fast.

Why H.264 still matters

H.265 is the newer, more efficient codec, and camera vendors increasingly ship it as the default. For storage it is a sensible choice. For live AI analysis it is often the wrong one. The CleverCam hub has a hardware decoder for H.264, so an H.264 stream is decoded almost for free. An H.265 stream has to be decoded in software, which uses far more of the hub's processor and memory, and that cost multiplies with every camera.

If your cameras let you choose per stream, a good pattern is H.265 on the main stream for recording and H.264 on the sub-stream for detection. You get efficient storage and efficient analysis.

Resolution is the master setting

The most common problem we find is a sub-stream set to CIF (352 × 288), the old default on many cameras and recorders. At CIF, a person thirty metres away on a wide-angle lens can be just a few pixels tall, which is too small for any detector to find reliably. Many cameras also cap the bitrate at CIF, so the night-time picture turns to mush.

Raising the sub-stream to around 704 × 480 is usually the single biggest improvement you can make to detection on an existing installation.

Watch the shape of the picture

A widescreen sensor feeding a 4:3 sub-stream produces a stretched or squashed picture. People come out taller and thinner, or shorter and wider, than they really are. Detection models are trained on correctly shaped people. Keep the sub-stream the same shape as the main stream where the camera allows it.

Switch off “smart” codec modes on the detection stream

Many cameras offer modes marketed as H.264+, H.265+ or "smart codec". They save storage by stretching the gap between full keyframes to many seconds. That is fine for a recording, but it slows down how quickly a detector (or a person opening live view) can start reading the stream. On the detection stream, keep a keyframe roughly every second.

Older recorders with a CIF-only sub-stream

Some older DVRs and hybrid recorders cannot raise their sub-stream above CIF at all. The workaround is to run detection from the main stream instead, turned down to something like 960 × 576, H.264 at about 8 fps.

Be aware of the trade-off: on a recorder the main stream is also the recording stream, so this lowers the recording resolution too. For many older analogue systems that is a price worth paying, because the recording was never sharper than that in the first place. But it should be a conscious decision, made with the customer.

Bandwidth on the local network

At these settings each camera's detection stream uses around 1 to 2 Mbps on the local network. That is modest, but it adds up on a busy switch or a weak Wi-Fi link. Wire cameras wherever you can. Power over Ethernet (PoE) also puts every camera on the same backed-up power supply, which matters when the lights go out.

How we apply these settings

On Hikvision and Dahua cameras, our team can apply these settings remotely, one camera at a time. We check that detection still works after each change and roll back if it does not, so a camera is never left worse than we found it. For other brands, the settings are applied in the camera's own web page or through the recorder.

If you are an installer, it is worth making these settings part of your commissioning checklist. The ten minutes it takes will save you callbacks for missed detections later. See also how to place cameras for detection.

CleverCam Team

Security insights and product updates from the people building smarter surveillance.

Ready to Secure Your Property?

Upgrade to AI-powered surveillance with CleverCam. Find an accredited installer near you and take the first step toward smarter security.