Edge video · capture and transport
Hardware encode is 6× cheaper. That gap is the only reason two cameras fit on four cores.
Two cameras on one Jetson Nano, streamed to a browser over WebRTC from a single container. The sensor dictated the design: the onboard IMX219 hands out 10-bit Bayer that FFmpeg cannot debayer, so the two cameras need two unrelated capture stacks, and the only question left is what contract both can satisfy. The measurement that decides whether the answer runs at all is the encoder cost, and the two failures worth writing down were both invisible from the page: a black player reporting Live, and one camera switch running both encoders for sixteen seconds.
What a live camera costs
The board is a 2019 Nano, four A57 cores, measured while a browser was actually watching. Only the camera being watched encodes at all, which is the other half of why two fit.
| idle, nobody watching | load average 0.34 |
| USB camera, software x264, 640×480 with Opus audio | 66.3% of a core |
| CSI camera, hardware NVENC, 1280×720, video only | 10.9% of a core |
| a viewer on each at once | load average 1.07 |
Two software encodes would be most of the board before anything else runs. Two of these, one of each kind, cost about three quarters of one core while both are live.
Why there are two capture stacks at all
The USB webcam is ordinary: v4l2 into software x264 is cheap enough at VGA. The
onboard camera cannot be done that way at all. On JetPack 4.x the IMX219 presents 10-bit
Bayer (RG10) on /dev/video0, which FFmpeg cannot debayer, and
there is no ISP anywhere in the FFmpeg path. It has to go through GStreamer's
nvarguscamerasrc, which reaches the sensor through nvargus-daemon
and the Tegra ISP, and nvv4l2h264enc, which encodes in hardware. The 6× gap in
the table above is not a choice between two options; it is the difference between the path
that exists and the path that does not.
So the design question is not which stack to standardise on, it is what the smallest contract both can satisfy. That contract is three rules and two comment headers, and the project is what is left once it holds.
The player reported Live over a black frame
MediaMTX answers a video-only path with an empty audio m-line that carries no
a=msid. The browser fires ontrack a second time for it, with
e.streams empty. The handler assigned e.streams[0]
unconditionally, so the second event wrote undefined over the video stream it
had bound a moment earlier.
| what the page said | Live, green dot |
| what the connection said | connected, video track flowing |
| what the viewer saw | black |
Nothing in the UI, the connection state or the track list distinguished this from a working stream, and it only happened on the camera with no microphone. Finding it meant reading the SDP that MediaMTX answered with. The fix is a function that returns the stream to bind or null, which refuses an empty event and refuses to rebind the stream already playing.
One viewer switching camera ran both encoders for sixteen seconds
Closing the peer connection does not tell MediaMTX the session is over. It drops the reader once ICE reaches disconnected, about 5 to 6 seconds, and only then does the 10-second on-demand linger start counting. A single viewer switching from one camera to the other therefore left the old camera encoding alongside the new one for roughly 16 seconds, which is precisely the overlap on-demand capture exists to prevent, on a board where one software encode is two thirds of a core.
The player now DELETEs its WHEP session on switch and on page unload, with
keepalive so the request still leaves a page that is closing. The 10-second
linger is kept deliberately: a reload, or a switch away and straight back, reuses the live
pipeline instead of re-opening the device.
The contract, and what it rejects
A camera backend is any executable dropped in cameras.d/. Its filename is the
path name, so front-door.sh is served at :8889/front-door/whep and
appears in the dropdown. Nothing else is edited: not the MediaMTX config, not the
Dockerfile, not the player, because the entrypoint generates the config and the player's
camera list from that directory.
- Publish H264 to
rtsp://127.0.0.1:8554/${MTX_PATH}. - Stay in the foreground. MediaMTX signals the process group on teardown, so a backgrounded pipeline is one that keeps holding the device.
- Exit non-zero, with a message, when the configuration is unusable. The restart policy turns a silent failure into a restart loop.
The rejections are where the time went. CSI_SIZE is validated as two separate
integers, because 1280x and x720 both look numeric once joined and
reach GStreamer as height=, which surfaces as an opaque caps error inside a
restart loop. nvv4l2h264enc rejects 4M and wants bits per second as
an integer. The argus socket is checked before the pipeline is built, with the real causes
in the message, including that restarting nvargus-daemon on the host replaces
the socket's inode and a single-file bind mount cannot follow it.
What is checked, and by what
One command on the target forces every backend to spawn its pipeline and asserts what comes
back. It fails if a backend publishes no video, and also if a backend that declares
camera-audio: no publishes an audio track, which is how the presentation header
stays true rather than decorative.
usb | OK, video and audio |
csi | OK, video |
test | OK, video |
The player has its own assertions, run from the browser console. One of them is that a
stored camera preference which is not a real camera falls back to the default: the cases
that matter are constructor and toString, truthy on any JavaScript
object, which a naive lookup accepts and which would strand the player on a path that 404s
forever.
What this does not show
One board, one USB camera, one sensor, one run of each measurement, read from the host's own process accounting rather than instrumented. The two encoder numbers are not a like-for-like benchmark: different resolutions, and only one of them carries audio. No glass-to-glass latency was measured, so "low latency" here means the transport is WebRTC and the encoders are configured for it, not a number. The container is pinned to JetPack 4.x and L4T R32.7, and its bionic base is a clock already running: this project had to move off Debian bullseye when that pool disappeared. There is no authentication on the streams, so anyone who can reach the port can watch and can cause capture to start, which makes it a LAN or tailnet tool and not an internet-facing one. Audio is USB-only, because the C250 microphone delivers samples only while its own video interface is open. Fast camera switching is verified by hand, not by a test.
Reproduce it
The hardware claims were probed before any of this was built, on the host and then inside
the base image: 60 frames of hardware-encoded 1280×720 CSI capture on the host, 30 frames
repeated in l4t-base:r32.7.1 with the argus socket bind-mounted, plus the
presence of every element and muxer the two pipelines need. That probe is what made the
single-container, on-demand shape viable rather than hoped for.
Deployment is scp over SSH and docker compose on the target: no
registry, no CI, the image built on the Jetson. The repository carries the design document
and the implementation plan it was built from, including the probe tables and the decisions
that were rejected on the way.