Skip to content

Separate transport, server queue, barrier, and playout latency measurements #80

Description

@christofmuc

Problem

The current latency and jitter observations combine several fundamentally different delays. In particular, an audio packet used for roundtrip measurement may wait in the server ingress queue and at the all-client mix barrier. A client with a faster audio clock then appears to have the worst roundtrip because its packets spend the most time waiting for the slowest client, even when its network path and received audio are healthy.

This makes the measurements difficult to interpret and unsuitable as inputs to automatic server or client buffer controllers.

Goal

Measure transport, server scheduling, and client playout separately so that every reported value answers a specific operational question. No adaptive controller should consume an aggregate latency value when only one component is relevant.

Required measurements

Transport

  • Control-path RTT measured by an immediate server echo that bypasses audio queues and mixing.
  • Uplink packet inter-arrival period and one-sided late residual per source.
  • Downlink packet inter-arrival period and one-sided late residual at each client.
  • Reordered, duplicate, too-late, and missing packet counters.

Transport RTT must use one local clock for both endpoints of the measurement. Do not subtract unsynchronized monotonic timestamps from different machines.

Server ingress and mixing

Per source:

  • current, minimum, maximum, and high-water queue depth;
  • packet residence time from server receive to mix consumption;
  • relative queue-depth/rate slope;
  • barrier-starvation count and duration;
  • number of times the source was the last contributor completing a mix;
  • deliberate drift corrections and hold/flush fast-forwards;
  • packets discarded by each explicit policy.

Per session:

  • time from first available contribution to completion of each mix set;
  • selected adaptive safety target and every target transition;
  • mix processing duration;
  • mix-to-send queue residence and send cadence.

Client playout

  • receive-to-playout queue residence;
  • current, low-water, and high-water queue depth;
  • audio callback consumption cadence;
  • underrun, rebuffer, overflow, and fast-forward counts;
  • selected playout target and every automatic adjustment;
  • concealment type and duration.

User-facing latency

Report aggregate values only with unambiguous names:

  • transport RTT;
  • audio-loop latency;
  • server queue residence;
  • server barrier wait;
  • client playout residence;
  • reported audio-device input/output latency where available.

Document which components are measured, estimated, or unavailable. Do not label audio-loop latency as network RTT.

Implementation constraints

  • Use monotonic clocks for intervals.
  • Keep timestamp domains explicit.
  • Store counters, extrema, rolling histograms, and snapshots without logging or allocating on real-time paths.
  • Format and publish diagnostics from non-real-time threads.
  • Report both milliseconds and frames where meaningful.
  • Use bounded rolling windows and percentiles rather than only averages.
  • Preserve enough packet/mix identifiers to correlate deterministic impairment tests without requiring per-packet production logging.
  • Provide an exportable diagnostic snapshot suitable for issue reports and automated tests.

Acceptance criteria

  • Adding server queue depth raises audio-loop latency but does not change the control-path transport RTT.
  • A simulated fast source clock increases that source's server residence and queue slope without being reported as worse network RTT.
  • Symmetric network delay changes transport RTT by the expected amount.
  • Uplink-only jitter changes server ingress metrics but not client downlink jitter.
  • Downlink-only jitter changes client arrival/playout metrics but not server ingress jitter.
  • A server CPU stall is distinguishable from network delay.
  • A client audio-callback stall is distinguishable from packet lateness.
  • Hold/flush, reorder, burst loss, and deliberate fast-forward tests produce distinct counters.
  • Measurements remain correct across 64/128 network frames, variable device callback sizes, reconnects, and counter rollover.
  • Instrumentation introduces no steady-state allocation, blocking, I/O, or logging on audio callbacks.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions