← Uday Kumarcase study

From licensed black box to owned edge inference

A computer-vision platform that turns camera streams into occupancy events on hardware the organization owns, built rather than licensed, and operated through a real distribution shift.

in production since 2025

System

  1. IngestRTSP, hardware decode
  2. InferenceYOLOv8, TensorRT INT8
  3. TrackingByteTrack, C++ line crossing
  4. EventsKafka
  5. Warehousedimensional, zone = INSERT
  6. Monitorper-camera drift
  7. Retrainlow-light footage
  8. Canarycamera subset
  9. Promote94 to 95.2 mAP

Why INT8

Exporting to ONNX and compiling to TensorRT INT8 bought roughly four times the throughput for under a point of accuracy. That is what makes forty streams fit on two cards instead of eight.

Why a C++ element

Line crossing has to correct for each camera's distortion and tilt, per frame, at frame rate. A Python probe in the pipeline could not hold that budget, so the crossing logic is a native element inside DeepStream.

Why zone = INSERT

Crossings publish to Kafka and land in a dimensional schema. Adding a new zone is a row, not a deploy, so the people who run the site can change what is measured without an engineer.

Constraint

Commercial video-analytics platforms sell per-camera licences for generic detectors. On the geometry actually installed, tilted mounts and fisheye lenses, those detectors lost accuracy that no amount of licence spend would recover. The fee bought a model that could not see the room.

Sending frames to a cloud endpoint was not an option either. Continuous video from dozens of cameras is a bandwidth problem, and the round trip adds latency the count cannot absorb.

Decision

Build the detector. Fine-tuning on footage from the organization's own cameras closed a seven-point accuracy gap, and the weights, the pipeline and the calibration belong to it rather than to a vendor.

This was decided on accuracy, not on price. The licensed option was expensive and wrong.

what it cost
roughly twelve months, and about two hundred hours of directed labelling
what it bought
no vendor renewal, no per-camera fee, and a model that can be retrained
what it risked
an in-house model nobody else can support

Measured result

40+
concurrent streamsat 15 FPS
~45 ms
end to endingest to event
94%
mAP@50 held outagainst 87% stock
~4x
throughputunder a point of accuracy
2x
RTX A4000warm standby
12+
months in productionthrough hardware refreshes

The drift event

Winter lighting moved the input distribution. Per-camera monitoring surfaced it before any person reported a bad number. The model had not changed; the environment had.

Camera health is watched on event-volume drift and in-out imbalance, so a feed that degrades raises an alert before anyone notices it visually. That is what turned a seasonal accuracy loss into a maintenance task rather than an incident.

Retraining on newly labelled low-light footage took held-out mAP@50 from 94% to 95.2%. It canaried on a subset of cameras before going wide. Retraining now runs quarterly on the same path.