From licensed black box to owned edge inference
A computer-vision platform that turns camera streams into occupancy events on hardware the organization owns, built rather than licensed, and operated through a real distribution shift.
in production since 2025
System
- IngestRTSP, hardware decode
- InferenceYOLOv8, TensorRT INT8
- TrackingByteTrack, C++ line crossing
- EventsKafka
- Warehousedimensional, zone = INSERT
- Monitorper-camera drift
- Retrainlow-light footage
- Canarycamera subset
- Promote94 to 95.2 mAP
Why INT8
Exporting to ONNX and compiling to TensorRT INT8 bought roughly four times the throughput for under a point of accuracy. That is what makes forty streams fit on two cards instead of eight.
Why a C++ element
Line crossing has to correct for each camera's distortion and tilt, per frame, at frame rate. A Python probe in the pipeline could not hold that budget, so the crossing logic is a native element inside DeepStream.
Why zone = INSERT
Crossings publish to Kafka and land in a dimensional schema. Adding a new zone is a row, not a deploy, so the people who run the site can change what is measured without an engineer.
Constraint
Commercial video-analytics platforms sell per-camera licences for generic detectors. On the geometry actually installed, tilted mounts and fisheye lenses, those detectors lost accuracy that no amount of licence spend would recover. The fee bought a model that could not see the room.
Sending frames to a cloud endpoint was not an option either. Continuous video from dozens of cameras is a bandwidth problem, and the round trip adds latency the count cannot absorb.
Decision
Build the detector. Fine-tuning on footage from the organization's own cameras closed a seven-point accuracy gap, and the weights, the pipeline and the calibration belong to it rather than to a vendor.
This was decided on accuracy, not on price. The licensed option was expensive and wrong.
- what it cost
- roughly twelve months, and about two hundred hours of directed labelling
- what it bought
- no vendor renewal, no per-camera fee, and a model that can be retrained
- what it risked
- an in-house model nobody else can support
Measured result
- 40+
- concurrent streamsat 15 FPS
- ~45 ms
- end to endingest to event
- 94%
- mAP@50 held outagainst 87% stock
- ~4x
- throughputunder a point of accuracy
- 2x
- RTX A4000warm standby
- 12+
- months in productionthrough hardware refreshes
The drift event
Winter lighting moved the input distribution. Per-camera monitoring surfaced it before any person reported a bad number. The model had not changed; the environment had.
Camera health is watched on event-volume drift and in-out imbalance, so a feed that degrades raises an alert before anyone notices it visually. That is what turned a seasonal accuracy loss into a maintenance task rather than an incident.
Retraining on newly labelled low-light footage took held-out mAP@50 from 94% to 95.2%. It canaried on a subset of cameras before going wide. Retraining now runs quarterly on the same path.