ETH-68: Ethernet Audio Interface for Linux

51 points by chabad360 a day ago on hackernews | 26 comments
3dcad

eth68 is an ethernet audio interface for Linux. 

  • Very low latency: 3.620 milliseconds round trip at 48 kHz with 64 sample buffer
  • Scalable: multiple units can be synchronized to increase channel count without increasing round trip audio latency
  • 6 balanced line level inputs on 1/4 inch TRS connectors
  • 8 balanced line level outputs on 1/4 inch TRS connectors
  • 48 kHz and 96 kHz sample rates
  • Burr-Brown PCM3168A codec
  • DIN MIDI connectivity
  • 1U 19-inch rack chassis

Usage with JACK

eth68 runs a custom bare-metal firmware on an STM32H7 microcontroller that emulates a netJACK1 master endpoint. To start exchanging audio with eth68, launch a JACK server with the netone backend:

jackd -d netone

When the first incoming capture packet from eth68 is received, JACK will create IO ports that map to the physical eth68 IOs:

qqq

Start a JACK aware application and connect it to the eth68 IOs:

Usage with PipeWire

It is also possible to plug into a PipeWire session directly using the pw-eth68 client. This PipeWire client implements similar behavior as a JACK server with the netone backend. Start a PipeWire aware application and connect it to the eth68 IOs:

Audio Latency

Audio Latency is the audible latency of the audio signal itself as it traverses the audio hardware and software layers. The audio latency arises from buffering in both the hardware and software, as well as digital filter group delay in the audio codec. The round trip audio latency can be measured with a loopback cable. On systems that suffer from a large audio latency, there is a perceptible delay between pressing a key on a MIDI keyboard and hearing the result from a software synthesizer. Various experiments have established latency perception thresholds ranging from 5 to 20 ms depending on listening contexts.

eth68 achieves 3.620 milliseconds of round trip audio latency at 48 kHz with a buffer size of 64 samples, measured with jack_delay and a loopback cable:

The following table summarizes eth68 round trip latency measurements with jack_delay for various buffer sizes at 48 and 96 kHz1 2:

32 64 96 128 256
48 kHz 2.288 ms 3.620 ms 4.954 ms 6.287 ms 11.621 ms
96 kHz 1.148 ms 1.814 ms 2.480 ms 3.147 ms 5.814 ms

Compared to the RME HDSPe AIO Pro PCIe card 3:

32 64 96 128 256
48 kHz -.--- -- 3.62 ms -.--- -- 6.28 ms 11.62 ms
96 kHz -.--- -- 2.14 ms -.--- -- 3.48 ms 6.14 ms

eth68 matches the latency performance of the RME card at 48 kHz and surpasses it by 0.33 ms at 96 kHz.

Multi-unit Synchronization

Multiple eth68 units can be synchronized to increase the channel count. Synchronization is achieved by:

  1. Sharing a high frequency (HF) clock. One eth68 unit is a clock master, generating an HF clock that is daisy chained to additional eth68 units with BNC cables. Sharing the HF clock eliminates sample drift between units.
  2. Nulling the sample offset between units by sending a broadcast UDP sync command with eth68ctl.py. All units receive the sync broadcast message within ~1 sample time (20.8 us at 48 kHz) and restart their DMA transfers immediately. This results in a sample offset of -1, 0, or +1 samples between units.

Multi-unit operation via the JACK interface does require the use of an intermediate proxy application eth68proxy since netjack isn't designed to merge/split data from/to multiple masters. eth68proxy takes care of merging the capture packets from multiple eth68s into one larger netjack compatible capture packet, as well as splitting up a large playback packet from netjack and delivering the smaller playback packets to each eth68 unit.

With two synchronized eth68s4 and eth68proxy running, netjack automatically detects the doubled channel count and creates the JACK ports:

The following scope shot shows one output from each of two synchronized units, both driven by the same sawtooth oscillator from Reaper:

In the above image, the scope is triggering on the falling edge of the sawtooth, and the green trace has been shifted up by 20 millivolts for clarity. The waveforms are aligned to less than 1 microsecond or 1/20th of a sample at 48 kHz.

When operating with synchronized eth68's, the round trip audio latency is the same as with only a single eth68. This was verified by looping back an output from the first unit to an input on the second synchronized unit and running jack_delay. There is a small penalty on the processing latency due to the extra hop through eth68proxy, but it is not significant enough on my system to measurably increase the xrun frequency.

When using the PipeWire interface, pw-eth68 manages the data streams from/to multiple eth68s, so eth68proxy is not needed.

Processing Latency

Processing Latency is the amount of time it takes to complete all processing of a buffer of audio. The processing latency must be less than the deadline time, otherwise an xrun (overrun or underrun) will occur and result in an audible glitch. In the absence of xruns, processing latency is not audible. The processing latency can be reduced and controlled in several ways including: tuning the host OS for responsiveness, using a higher performance host computer, etc.

eth68 generates a 3.3 V digital signal LATMON for measuring and debugging processing latency. The signal is available on a BNC connector on the rear panel. Here it is on the scope:

LATMON goes high as soon as capture data is available in the capture DMA buffer. LATMON goes low when eth68 is done copying the received playback data into the playback DMA buffer. The LATMON pulse width is therefore an accurate measure of the total processing latency of each cycle and is affected by every element in the data path: processing delay in the microcontroller, network transmission delay, host OS delays, signal processing delay in the DAW, etc.

Histogramming the LATMON pulse widths in real time gives a feel for the processing latency and associated jitter:

In the above plot, the vertical dashed lined is drawn at 1333 microseconds which is the deadline time for 64 sample buffers at 48 kHz. Pulse widths falling on or to the right of the dashed line would result in an xrun. The processing latency for a typical cycle is about 625 microseconds, leaving 708 microseconds (more than 50% margin) to the deadline. The above data was captured while Bitwig was running a project with moderate DSP load.

We can artificially worsen the latency performance by dispatching some hogs with stress -c 6. This makes the host system less responsive by tying it up with busy work, pushing the LATMON pulse width distribution closer to the deadline:

Network Configuration

eth68 sends capture packets to 12.12.12.10:3000 by default. The easiest way to get started is to just set your host IP to 12.12.12.10 and run jackd -d netone. For other configurations, the eth68ctl.py script can be used to adjust the eth68 network settings.

Query the eth68 settings with eth68ctl.py:

python eth68ctl.py 12.12.12.10 query

Output:

--------------------------
header: 2820800604
board_id: 3407900
source_ip: 12.12.12.100  (202116196)
netmask_ip: 255.255.255.0  (4294967040)
gateway_ip: 12.12.12.1  (202116097)
destination_ip: 12.12.12.10  (202116106)
destination_port: 3000
mode: 0
period_size: 64
sample_rate: 48000
capture_channels: 6
playback_channels: 8
footer: 2603798065

From the output, we can see that the eth68 IP address is 12.12.12.100 and it transmits captures packets to 12.12.12.10:3000.

The first argument to ethc68ctl.py is the host IP on the interface that shares a LAN connection with eth68. eth68ctl.py uses broadcast UDP on this interface to control eth68(s). Using broadcast UDP ensures that communication is always possible even if the networking settings are incorrect.

The settings can be changed using the "set" sub-command. For example, to change the destination IP to 12.12.12.50:

python eth68ctl.py 12.12.12.10 --board_id 3407900 set destination_ip 12.12.12.50

eth68ctl.py is also used to change non-network related settings such as sample rate:

python eth68ctl.py 12.12.12.10 --board_id 3407900 set sample_rate 96000

Networking Hardware

It is recommended to use a dedicated LAN for eth68 setups which is isolated from extraneous traffic. Ordinary networking hardware can yield very good results. AVB/TSN ethernet capable hardware is not required. A 1G ethernet switch is recommended for multi-unit setups. eth68 links at 100M full duplex, but having a 1G full duplex link between the host and switch reduces the processing latency. If only a single eth68 is used, it can be connected directly to the host's NIC without a switch.

Results depend significantly on the host's NIC. Excellent results can be had with Intel I210 NICs, which run 15-30$ for a PCIe card in 2026. An important aspect of the NIC is that its interrupt behavior can be modified. NICs often implement some form of interrupt coalescing which can be incompatible with real time applications. Intel I210 apparently implements a reasonable default interrupt coalescing policy since audio with eth68 is smooth with the default settings. However, using the latency debug monitor, it was found that the processing latency can be reduced by a full 175 us (13% of the deadline time) by changing the interrupt coalescing timeout using the following command:

ethtool -C enp9s0 rx-usecs 0

Audio Measurements

I used a Stanford Research DS360 Ultra Low Distortion Function Generator to measure the THD+N performance of the (differential) input:

The -1 dBFS peak input signal was digitally filtered using a steep 20 kHz low-pass filter and a Q = 5 notch filter at the fundamental frequency as per AES17. The PCM3168A datasheet specifies a THD+N of -93 dB typical at 1 kHz for a 48 kHz sample rate. My measurement at 1 kHz is THD+N = -94.8 dBFS (0.0018%).

Note that the sudden improvement of THD+N around 10 kHz is due to the second harmonic shifting into the stop band of the 20 kHz low-pass filter.

Unfortunately, I don't have any test equipment that is sensitive enough to measure the THD+N of the outputs. Instead, I used an eth68 input to analyze the eth68 output performance in a loopback configuration. The measurement is really a bound on the output performance since the input THD+N contributes to the loopback THD+N.

Nonetheless the loopback THD+N is within 0.5 - 1 dBFS of the input THD+N as measured with the DS360.

With the loopback still connected, I did a finer frequency scan to check the passband flatness:

The combined input + output passband flatness is +/- 0.15 dBFS.

Turning off the tone generator and measuring the noise level with an A-weighting filter over 20 kHz of bandwidth, I get:

  • Input only, input grounded: -107.8 dBFS at 48 kHz, -110.5 dBFS at 96 kHz. A-weighted.
  • Input + output (loopback): -106.2 dBFS at 48 kHz, -108.6 dBFS at 96 kHz. A-weighted.

Taking the ratio relative to a full scale -3 dBFS RMS sine wave, the dynamic range is:

  • Input only, input grounded: DR = -104.8 dBFS at 48 kHz, DR = -107.5 dBFS at 96 kHz. A-weighted.
  • Input + output (loopback): DR = -103.2 dBFS at 48 kHz, DR = -105.6 dBFS at 96 kHz. A-weighted.

macOS and Windows

I'm primarily interested in Linux compatibility, but JACK is cross-platform and eth68 is successfully detected by the JACK server on both macOS and Windows. Unfortunately, most audio applications on macOS and Windows are not JACK aware; Ableton Live on macOS does not know how to talk to a JACK server. Even applications that are JACK aware on Linux are not always JACK aware on other OSes; Reaper supports JACK on Linux, but not on macOS. The JACK project aims to solve this problem by bundling a virtual driver "JACK-Router" that bridges between the native audio driver interface (CoreAudio on macOS, ASIO on Windows) and the JACK session.

Current status of JACK-Router:

  • JACK-Router is currently working on Windows. See here. This page also includes a list of JACK aware applications on Windows.
  • JACK-Router on macOS is currently broken due to changes rolled out in macOS 10.15. See here.

Workbench

eth68 Revision B PCB on the bench for loopback testing:

Host System

  • OS: Ubuntu 24.04.3 LTS
  • Kernel: 6.19.10-2-liquorix-amd64
  • CPU: AMD Ryzen 7 5700X
  • Motherboard: ASUS B450-F
  • Switch: TP-LINK TL-SG608
  • NIC: Intel I210