Uday KumarSenior Engineer, Data & AI Platforms · 9 years

Building and operating AI systems in production.

  1. 1raw session
  2. 2sorted by type
  3. 3one denial
  4. 4decision opened

The record behind it and the source that captured it.

one session · 10h 36m · 635 rows recorded

tool calls · 321decisions · 229delegation · 54other · 29failures · 2denied · 1 of 22900:0010h 36m

the refused decision

decision
denied
of
229 recorded decisions
alongside
228 accepted, 2 observed failures
found by
filtering the session
source
native hooks

every row says which of these it is

factthe source recorded it
inferenceit was worked out from what was recorded
unknownnothing covers it, and it says so

Two capture sources: native hooks and OpenTelemetry.

Investigate what your AI agent did. An agent runs unattended, calls tools, delegates, accepts and refuses things, and stops. Afterwards nobody can say what happened. GMAN records the session across two independent sources and rebuilds it, keeping facts, inference and unknowns separate.

session figures published on getmyagentnow.com · developer preview, Claude Code demonstrated today · built for model- and framework-agnostic agent operations

  1. 1citation exists
  2. 2walk the field
  3. 3values compared
  4. 4held for review

The verdict is held for human review.

asserted by the model

severity = P0
cited
ret_4471alerts[3].severity

retrieval exists · an existence check passes here

Existence is not support. The validator walks the field.

evidence drawer · ret_4471

{
  "id": "ret_4471",
  "alerts": [
    { "id": "a-1", "severity": "P3" },
    { "id": "a-2", "severity": "P4" },
    { "id": "a-7", "severity": "P3" },
    { "id": "a-9", "severity": "P2" }
  ]
}
stored P2 ≠ asserted P0needs human

Real citation. Wrong content. Caught.

A real retrieval id pointing at the wrong content is the failure an existence check cannot see. The validator walks the cited field path and compares the stored value to the asserted one, so a citation that resolves but does not support gets caught and routed to a person.

public source · evaluation runs without an API key · 34 numbered architecture decisions, each recording what was rejected

Social agent

  1. 1gathered17rss items in the lookback window
  2. 2assessed17scored against the relevance bar
  3. 3established0reached 3-source convergence
  4. 4briefing1composed and validated

Briefing validated, awaiting an operator, not published.

near misses · 4 of 6 shown

  • agent-orchestrationinsufficient convergence (2 < 3)
  • ai-cost-controlinsufficient convergence (2 < 3)
  • frontier-modelsinsufficient convergence (2 < 3)
  • agent-evalsinsufficient convergence (1 < 3)

publication gate

The scheduled run holds no publication credentials.

VALIDATED
AWAITING OPERATOR
PUBLISHED: FALSE

One real run. Seventeen items gathered, seventeen scored, and not one topic cleared the three-source convergence bar, so the run established nothing and recorded why. The briefing composed, validation passed, and the publication gate stayed shut.

cycle-05 run record and briefing summary · outcome prepared · validationOk true · published false

Edge-inference platform

  1. 1frame
  2. 2detection
  3. 3track
  4. 4event

Count and crossing.

zone_03zone_07zone_11t041t042t043t044t045t046t047

occupancy · derived from crossings

zone_031
zone_072
zone_114

crossings written

  • track_047 │ IN │ zone_03 │ 14:32:04.902
  • track_041 │ IN │ zone_07 │ 14:32:06.118
  • track_042 │ IN │ zone_11 │ 14:32:07.556
  • track_044 │ IN │ zone_11 │ 14:32:08.441

Forty-plus streams at fifteen frames a second, on-premise. Occupancy is derived from crossings.

floor plan and movement are a representative illustration

A build-versus-buy call decided on accuracy, not on price.

From licensed black box to owned edge inference

Commercial video analytics wanted per-camera licensing for generic models that underperformed on the tilted and fisheye geometry actually installed, and cloud inference was a non-starter on bandwidth and latency. I made the call to build. The weights, the pipeline and the calibration belong to the organization.

A YOLOv8 detector fine-tuned on footage from its own cameras holds 94% mAP@50 against 87% stock, compiled to TensorRT INT8 for roughly 4× throughput at under a point of accuracy. RTSP in, hardware decode, batched inference, ByteTrack for identity, then a C++ element I wrote to do calibration-aware line crossing per camera. Crossings publish to Kafka and land in a dimensional schema where adding a zone is an INSERT, not a deploy.

It has been in production over a year, and it has already survived a drift event. Winter lighting moved the input distribution. Per-camera monitoring surfaced it before anyone reported it. The model had not changed; the environment had. Retraining on low-light footage took held-out mAP@50 to 95.2%, canaried on a subset of cameras before it went wide.

Read the full case study

Uday Kumar

udaygkumar33@gmail.com