Skip to content
FV Francesco Vigni
Francesco Vigni

Research

Medical imaging, and the evidence that says when the model is wrong.

Four recent pieces of work, all on public data, all with the negative results left in. Each card carries one finding; the detail page carries the argument, and the repository stays the source of record.

Endoscopy · acquisition shortcuts

You can tell which hospital a colonoscopy frame came from 96% of the time, without looking at the anatomy, and the standard fix barely helps.

Nine numbers describing the frame the equipment drew (letterboxing, the field-of-view mask, the burned-in date) identify the source dataset out of five, at chance 0.20.

  • 0.961 as acquired
  • 0.859 after the standard crop-and-pad
  • 0.891 from a frozen self-supervised encoder
The same endoscopic frame shown through four preprocessing pipelines, for two source datasets

Four findings, and the one that failed →

Repository

Edge video · capture and transport

Hardware encode costs 10.9% of a core. Software x264 costs 66.3%, which is the only reason two cameras fit on four cores.

The onboard IMX219 hands FFmpeg 10-bit Bayer it cannot debayer, so two cameras on one Jetson Nano need two unrelated capture stacks. Both failures worth writing down were invisible from the page: a black player reporting Live, and one camera switch running both encoders for 16 seconds.

  • 6× cheaper to encode in hardware, at four times the pixels
  • 1.07 load average with a viewer on each camera, idle 0.34
  • 16 s both encoders live after one viewer switched camera
Two bars comparing CPU cost per core, software x264 at 66.3% against hardware NVENC at 10.9%

What the sensor decided, and what the page hid →

Repository

Edge deployment · video object segmentation

DAVIS priced the speed-up at 3.4 accuracy points. On real video it cost the entire track, two occlusions, zero recoveries.

Lowering the input resolution is the cheapest frame rate an edge board sells. The benchmark that approves the trade averages over frames and holds few full occlusions, so it cannot see that the object is never found again: J after a hand passes over it falls to 0.006, against 0.646 one resolution up.

  • 0.006 J after occlusion, at the resolution the benchmark approved
  • 2 of 2 re-acquisitions once the evidence is used outside the vote, from 0
  • 0.75 fps full quality on the board, 33× short of real time
Accuracy against frame rate for five input resolutions on a Jetson Nano, every point far left of the 25 fps line

What the benchmark approved, and what it cost →

Repository

Fetal ultrasound · orientation

Estimating fetal cardiac orientation does not need a trained model: closed-form geometry beats the network by two orders of magnitude.

The useful part is knowing when that shortcut breaks. It does not break where the usual quality score says it should: a mask scoring Dice 0.87 can give a 46° error, one at 0.77 gives 0.22°.

  • 0.28° second-order moments, no training
  • 7.04° trained landmark model, box only
  • ±18° 95% limits of agreement
Six fetal four-chamber ultrasound frames with the annotated and predicted heart axes overlaid, labelled best, median and worst

Why landmarks, and where they break →

Repository

Previously: self-supervised pretraining of a medical-imaging foundation model on a large clinical video corpus under NDA, multi-GPU, with the cross-vendor standardisation and evaluation harness underneath it. Before that, seven years of computer vision and robotics in industry in Germany.

These studies first appeared as a standalone site, still online at portfolio.francescovigni.com. The repositories stay the source of record.

Full background Start a conversation [email protected] GitHub