Edge deployment · video object segmentation
DAVIS priced the speed-up at 3.4 points. On real video it cost the whole track.
A frozen vision transformer that segments an object through a video, moved from a workstation onto a 2019 Jetson Nano with a camera attached, to find out what the edge actually costs. The cheapest way to buy frame rate is to shrink the input, and the benchmark that governs that decision prices it at 3.4 accuracy points. On camera footage with one hand passing over the object, it costs the track: two occlusions, zero recoveries, 66 frames of empty mask. Also here, because it nearly turned this page into a published number instead of a finding, is how an entire fp16 sweep got timed twice before anyone read its output.
The trade the benchmark approves
DAVIS 2017 val says 480×864 → 384×672 costs 3.4 J&F points, 0.767 to 0.733, and buys 2× the frame rate. On that evidence it is the obvious move, and it is the move I would have shipped.
On real video with a real hand occlusion it costs the entire track. The object is covered, the mask empties, the object comes fully back, and at the lower resolution the mask never returns, through a complete reappearance and a second occlusion, 66 frames.
| input | J before | J after | re-acquired | visibility AUC |
|---|---|---|---|---|
| 480×864 | 0.742 | 0.646 | 2 of 2 | 0.845 |
| 384×672 | 0.653 | 0.006 | 0 of 2 | 0.534 |
DAVIS is not wrong about its own question. It is answering a different one: its validation split contains few full occlusions, and it averages over frames rather than asking whether the object was ever found again. Permanent loss of a recoverable object is worth a few points of a frame average and all of the capability. The metric that governs the resolution decision is blind to the thing that decision removes. For anyone deploying this, that is the result that matters.
Ground truth on the camera clip is automatic and exact, which is a property of the scene design rather than the code: make the target trivially segmentable, a saturated object, and occlude it with something that is not. The occluder is never thresholded, so nothing depends on what it is, and the target is free to move.
What the hardware actually delivers
The frame rates the trade is being made against, measured on the board, with accuracy on DAVIS 2017 val. This is the fp32 curve, which is the only one that computes the function; the next section is about why.
| input | patch tokens | J&F | fps | ms/frame |
|---|---|---|---|---|
| 192×336 | 252 | 0.428 | 6.37 | 157 |
| 256×448 | 448 | 0.573 | 3.95 | 253 |
| 320×576 | 720 | 0.666 | 2.28 | 438 |
| 384×672 | 1008 | 0.733 | 1.50 | 666 |
| 480×864 | 1620 | 0.767 | 0.75 | 1337 |
Correctness costs about 1.4× in latency here, because the half-precision build that would have paid for itself does not work on this toolchain.
Two measurements agreed, and both were timing nothing
Five fp16 engines built and ran. trtexec reported 9.04 fps at the smallest
input and 2.06 at the largest, with tight percentiles. A separately written runner
reproduced those timings on real camera frames to within 1%. Two agreeing measurements, so I
stopped checking.
Then 113 MB of saved features gzipped to 153 KB, because they were uniform NaN.
| fp32 engine, first frame | 96,768 / 96,768 finite |
| fp16 engine, first frame | 0 / 96,768 finite |
TensorRT 8.2 has no native LayerNorm, so the exported graph carries the textbook
decomposition, and the variance term overflows fp16 on transformer activations. The engine
is numerically dead from the first block. trtexec feeds random input and never
reads the output it has just timed, which is the correct behaviour for a throughput tool and
is documented. The mistake was mine, and it is the ordinary one: I read a timing number as evidence that
the engine computed the function, then read a second number that agreed. Agreement between two measurements of the same wrong thing is not validation, and nothing
in either output distinguished the two cases.
The same decomposition caused a second failure: at the largest input the fp16 build died outright, every candidate kernel killed by the GPU watchdog. In fp32 that size builds and runs. One root cause, two symptoms, and my first reading of the build failure, that the toolchain rather than throughput set the quality ceiling, was wrong and is withdrawn in the write-up rather than quietly corrected.
Two repairs, and what each one costs
Back to the resolution trap, which is the failure worth repairing. The encoder never lost the object; the voting scheme did. A global match against the object's descriptors still separates it from the background long after the mask is gone: at the failing resolution, 0.95 on target against a 0.67 background percentile. Four patches cannot win a top-5 vote against eight thousand context patches, so the evidence has to be used outside the vote. Re-detection on that basis takes J-after from 0.006 to 0.668 and re-acquisition from 0 of 2 to 2 of 2, in two frames.
The second repair is a silhouette taken from a photograph of the object, segmented once where cost does not matter. At runtime the object is only located, by one dot product, and the known shape is placed there. Borders then come from the reference at full resolution rather than from a 24-pixel patch grid, and occlusion is measured against the object's real shape.
| mask source | J | boundary F | occlusion r |
|---|---|---|---|
| propagation alone | 0.711 | 0.978 | not measured |
| silhouette, reference from the scene | 0.764 | 0.995 | 0.981 |
| silhouette, reference photographed separately | 0.670 | 0.947 | 0.961 |
A photograph shot separately survives a 10.9× scale gap and a change of background, lighting and camera: absolute similarity falls by a third, the margin over background barely moves. It is the only deployable version, since it does not need the object unoccluded and well lit in the first frame of every clip. It is not better than a good in-scene reference.
The condition under which the occlusion number means anything
A second clip was shot to characterise partial coverage and it failed, which turned out to be the more useful outcome. What decides the measurement is the margin: how much more similar the object is to its reference than the most object-looking part of the background.
| first clip, margin 0.252 | occlusion estimate r = 0.96 |
| second clip, margin 0.124 | r = 0.19, at every threshold tried |
The object was larger in the failing clip. Size was not the binding constraint, the background was. So the precondition is not "a big enough object" but a margin above roughly 0.2, and it depends on the scene as much as the subject. That is now a fifteen-second check that runs before anything is recorded, instead of a discovery made after.
What this does not show
One board, one backbone, one object, two usable clips. The occlusion result is one scene: it shows that a frame-averaged benchmark can approve a resolution that destroys re-acquisition, not how often that happens. The NaN is specific to TensorRT 8.2 and this LayerNorm decomposition; the transferable part is that a throughput tool cannot tell you an engine computes the function, and that two agreeing timings do not either. No INT8, pruning or distillation, none of which were tried. No power or thermal measurement: everything ran at maximum clocks with no duty cycle. The occlusion estimate is a detector with a known operating condition, not a calibrated meter. The silhouette fit assumes a roughly rigid object.
Reproduce it
Public CC-BY-4.0 data, every tool in the repository, and the device side installs nothing:
trtexec builds and times the engines, and the inference runner reaches CUDA
through ctypes because the board cannot run a modern pip. The runner aborts on
the first frame if the output is not finite: three lines, and the only reason the fp16 sweep
is a story on this page rather than a row in the table above it.
The write-up also records four conclusions that later measurements overturned, two of them its own earlier drafts. A method that took four wrong turns is worth less than its final numbers suggest, and a reader should be able to see which parts are load-bearing.