Skip to content
FV Francesco Vigni
Francesco Vigni

← Research

Edge video · capture and transport

Hardware encode is 6× cheaper. That gap is the only reason two cameras fit on four cores.

Two cameras on one Jetson Nano, streamed to a browser over WebRTC from a single container. The sensor dictated the design: the onboard IMX219 hands out 10-bit Bayer that FFmpeg cannot debayer, so the two cameras need two unrelated capture stacks, and the only question left is what contract both can satisfy. The measurement that decides whether the answer runs at all is the encoder cost, and the two failures worth writing down were both invisible from the page: a black player reporting Live, and one camera switch running both encoders for sixteen seconds.

What a live camera costs

The board is a 2019 Nano, four A57 cores, measured while a browser was actually watching. Only the camera being watched encodes at all, which is the other half of why two fit.

idle, nobody watchingload average 0.34
USB camera, software x264, 640×480 with Opus audio66.3% of a core
CSI camera, hardware NVENC, 1280×720, video only10.9% of a core
a viewer on each at onceload average 1.07
Two horizontal bars: software x264 at 66.3% of one core, hardware NVENC at 10.9%, against a scale of one core
Not a like-for-like comparison, and it does not need to be: the cheap path is also the one carrying four times the pixels. The expensive path carries the audio.

Two software encodes would be most of the board before anything else runs. Two of these, one of each kind, cost about three quarters of one core while both are live.

Why there are two capture stacks at all

The USB webcam is ordinary: v4l2 into software x264 is cheap enough at VGA. The onboard camera cannot be done that way at all. On JetPack 4.x the IMX219 presents 10-bit Bayer (RG10) on /dev/video0, which FFmpeg cannot debayer, and there is no ISP anywhere in the FFmpeg path. It has to go through GStreamer's nvarguscamerasrc, which reaches the sensor through nvargus-daemon and the Tegra ISP, and nvv4l2h264enc, which encodes in hardware. The 6× gap in the table above is not a choice between two options; it is the difference between the path that exists and the path that does not.

So the design question is not which stack to standardise on, it is what the smallest contract both can satisfy. That contract is three rules and two comment headers, and the project is what is left once it holds.

The player reported Live over a black frame

MediaMTX answers a video-only path with an empty audio m-line that carries no a=msid. The browser fires ontrack a second time for it, with e.streams empty. The handler assigned e.streams[0] unconditionally, so the second event wrote undefined over the video stream it had bound a moment earlier.

what the page saidLive, green dot
what the connection saidconnected, video track flowing
what the viewer sawblack

Nothing in the UI, the connection state or the track list distinguished this from a working stream, and it only happened on the camera with no microphone. Finding it meant reading the SDP that MediaMTX answered with. The fix is a function that returns the stream to bind or null, which refuses an empty event and refuses to rebind the stream already playing.

One viewer switching camera ran both encoders for sixteen seconds

Closing the peer connection does not tell MediaMTX the session is over. It drops the reader once ICE reaches disconnected, about 5 to 6 seconds, and only then does the 10-second on-demand linger start counting. A single viewer switching from one camera to the other therefore left the old camera encoding alongside the new one for roughly 16 seconds, which is precisely the overlap on-demand capture exists to prevent, on a board where one software encode is two thirds of a core.

The player now DELETEs its WHEP session on switch and on page unload, with keepalive so the request still leaves a page that is closing. The 10-second linger is kept deliberately: a reload, or a switch away and straight back, reuses the live pipeline instead of re-opening the device.

The contract, and what it rejects

A camera backend is any executable dropped in cameras.d/. Its filename is the path name, so front-door.sh is served at :8889/front-door/whep and appears in the dropdown. Nothing else is edited: not the MediaMTX config, not the Dockerfile, not the player, because the entrypoint generates the config and the player's camera list from that directory.

  1. Publish H264 to rtsp://127.0.0.1:8554/${MTX_PATH}.
  2. Stay in the foreground. MediaMTX signals the process group on teardown, so a backgrounded pipeline is one that keeps holding the device.
  3. Exit non-zero, with a message, when the configuration is unusable. The restart policy turns a silent failure into a restart loop.

The rejections are where the time went. CSI_SIZE is validated as two separate integers, because 1280x and x720 both look numeric once joined and reach GStreamer as height=, which surfaces as an opaque caps error inside a restart loop. nvv4l2h264enc rejects 4M and wants bits per second as an integer. The argus socket is checked before the pipeline is built, with the real causes in the message, including that restarting nvargus-daemon on the host replaces the socket's inode and a single-file bind mount cannot follow it.

What is checked, and by what

One command on the target forces every backend to spawn its pipeline and asserts what comes back. It fails if a backend publishes no video, and also if a backend that declares camera-audio: no publishes an audio track, which is how the presentation header stays true rather than decorative.

usbOK, video and audio
csiOK, video
testOK, video

The player has its own assertions, run from the browser console. One of them is that a stored camera preference which is not a real camera falls back to the default: the cases that matter are constructor and toString, truthy on any JavaScript object, which a naive lookup accepts and which would strand the player on a path that 404s forever.

What this does not show

One board, one USB camera, one sensor, one run of each measurement, read from the host's own process accounting rather than instrumented. The two encoder numbers are not a like-for-like benchmark: different resolutions, and only one of them carries audio. No glass-to-glass latency was measured, so "low latency" here means the transport is WebRTC and the encoders are configured for it, not a number. The container is pinned to JetPack 4.x and L4T R32.7, and its bionic base is a clock already running: this project had to move off Debian bullseye when that pool disappeared. There is no authentication on the streams, so anyone who can reach the port can watch and can cause capture to start, which makes it a LAN or tailnet tool and not an internet-facing one. Audio is USB-only, because the C250 microphone delivers samples only while its own video interface is open. Fast camera switching is verified by hand, not by a test.

Reproduce it

The hardware claims were probed before any of this was built, on the host and then inside the base image: 60 frames of hardware-encoded 1280×720 CSI capture on the host, 30 frames repeated in l4t-base:r32.7.1 with the argus socket bind-mounted, plus the presence of every element and muxer the two pipelines need. That probe is what made the single-container, on-demand shape viable rather than hoped for.

Deployment is scp over SSH and docker compose on the target: no registry, no CI, the image built on the Jetson. The repository carries the design document and the implementation plan it was built from, including the probe tables and the decisions that were rejected on the way.