back to projects

Emberlink

A robot-to-operator link that keeps the useful picture alive when the radio gets bad.

A frozen frame that looks live is the worst thing you can put in front of an operator. So as the link degrades this one sends less rather than falling behind: on a 100 kbps, 10% loss link the oldest thing on screen is 0.75 s old, where TCP is 5.0 s behind and still climbing.

  • C++20
  • UDP
  • Linux
  • netem
  • Node
operator — udp/47001 C++20
$ scripts/link.sh deep      # 100 kbps, 10% loss

stepping down on queueing delay, never on loss
  36 → 24 → 16 → 8 → 4 strips/s

picture age   0.75 s / 1.27 s   p50 / p95
alert         125 ms / 140 ms

Freshness beats completeness

The sender never queues. A frame is four independently decodable 1,200-byte strips, and the pacer sends one strip per token, always from the newest frame. A lost packet costs exactly its own strip — the old pixels stay on screen, visibly ageing, instead of the whole picture stalling.

  • Alerts and telemetry go before pixels. They come off the robot's full-precision frame, before any quantization, and are sent three times 50 ms apart, then retried until the operator acks. An empty send queue is what makes that ordering mean anything.
  • The ladder steps down on queueing delay, never on random loss. 36 down to 4 strips a second, about 363 to 40 kbps. Loss on a radio is noise; a growing queue is the signal that you are sending faster than the link can carry.
  • Silence is treated as failure. If the operator's reports stop for 600 ms the sender drops to the lowest rung rather than keep spending a link that may not be there.
  • Every part of the picture carries its own age. The operator sees which strips are stale rather than a uniform image that might be seconds old.

What it costs when the link is bad

One seeded 80×60 walkthrough at 9 Hz, replayed on all 4 link presets for Emberlink and two baselines, in two network namespaces on one Linux host so both ends read the same clock. Picture age is the age of the oldest part of what the operator is looking at — the number that says how stale the worst of it is.

deep 100 kbps, 10% loss
sender picture age alert latency
Emberlink 0.75 s / 1.27 s 125 ms / 140 ms
TCP 5.0 s / 6.9 s 4.3 s / 5.5 s
naive UDP never complete 768 ms / 890 ms
blackout deep, plus two 3 s outages
sender picture age alert latency
Emberlink 0.75 s / 2.6 s 159 ms / 1.9 s
TCP 5.0 s / 10.2 s 5.6 s / 8.9 s
naive UDP 15.9 s / 30.2 s 761 ms / 2.8 s
wall 500 kbps, 3% loss
sender picture age alert latency
Emberlink 206 ms / 339 ms 39 ms / 42 ms
TCP 340 ms / 514 ms 66 ms / 166 ms
naive UDP 3.0 s / 8.5 s 878 ms / 913 ms
good 10 Mbps, the case it is not for
sender picture age alert latency
Emberlink 85 ms / 261 ms 13 ms / 14 ms
TCP 206 ms / 334 ms 18 ms / 20 ms
naive UDP 82 ms / 129 ms 14 ms / 18 ms

After each of the two three-second outages, Emberlink is showing a current picture again in 2.3 s and 0.6 s. TCP never recovers inside the run.

On the good link it loses, and the last table says so. With 10 Mbps of headroom the right answer is to send everything the moment it is captured, which is what naive UDP does. Emberlink opens at the second-lowest rung and takes about six seconds to climb, and that climb is the whole of its p95. Starting cautiously is what stops it drowning a bad link, and a good link is not the one it is for.

CI reruns the deep and blackout presets on every push and fails the build if Emberlink misses its bounds — which it did on the first real run, because the controller had only ever been tested against a simulator.