For decades, the internet has quietly suffered from a paradox: memory became cheap, so network equipment vendors added ever-larger packet buffers to prevent drops, and in doing so they broke the very congestion signaling that TCP depends on to function well. The result is bufferbloat, a condition where seconds of latency accumulate in queues that were meant to smooth microsecond-scale bursts.

Active Queue Management, or AQM, is the discipline devoted to solving this. It sits at the intersection of control theory, protocol design, and empirical measurement, and it has been evolving continuously since Sally Floyd and Van Jacobson introduced Random Early Detection in 1993. Each generation of algorithm has revealed subtleties in how queues, flows, and end-host behaviors interact.

What makes AQM particularly interesting is that it is one of the few areas of networking where deployed practice still lags decisively behind research consensus. RED, CoDel, PIE, FQ-CoDel, and now L4S-adjacent schemes like DualPI2 represent an ongoing refinement of a deceptively simple question: when a router queue starts to fill, what should the router do about it, and how quickly? The answers keep changing because our understanding of congestion itself keeps deepening.

Bufferbloat and the Failure of Drop-Tail

Drop-tail queuing is the default behavior of most network interfaces: accept packets until the buffer is full, then discard whatever arrives next. In an era of small buffers and slow links, this worked adequately because TCP would notice the loss quickly and back off before latency ballooned. The feedback loop between congestion signal and sender response stayed tight.

The introduction of cheap DRAM changed the calculus. Engineers, reasoning that dropped packets were bad and memory was inexpensive, provisioned buffers sized for worst-case bursts on high-bandwidth links. A cable modem with several hundred milliseconds of buffering, paired with a TCP flow trying to fill the pipe, produces a steady-state queue that is nearly full nearly always.

This wrecks interactive traffic. TCP's loss-based congestion control interprets the absence of drops as headroom and keeps pushing. The queue swells, latency climbs into the hundreds of milliseconds or beyond, and every other flow sharing the link, whether a DNS lookup, a VoIP packet, or a game update, waits behind that standing queue. Bandwidth is preserved; responsiveness collapses.

The pathology is subtle because throughput measurements look fine. A speed test reports the expected megabits per second. Only when you measure latency under load, using tools like the RRUL test or flent, does the damage become visible. This measurement gap is part of why bufferbloat persisted for so long as an unnamed problem.

Drop-tail's fundamental limitation is that it provides no early signal. By the time a packet is dropped, the queue is already full and latency is already maximal. AQM inverts this: signal congestion before the buffer saturates, giving TCP time to respond while queues remain shallow.

Takeaway

A buffer that never overflows can be worse than one that occasionally does. Congestion control is a conversation, and silence is not a neutral answer.

From Queue Length to Sojourn Time

Early AQM algorithms like RED tried to preempt buffer saturation by probabilistically dropping packets as average queue length grew beyond a threshold. In principle this worked; in practice, RED required careful tuning of thresholds against link speed and traffic mix, and misconfigured RED often performed worse than plain drop-tail. Operators disabled it, and the algorithm gained a reputation for fragility.

CoDel, introduced by Kathleen Nichols and Van Jacobson in 2012, reframed the problem. Instead of asking how many packets are in the queue, it asks how long each packet has been waiting. This sojourn time is a direct measurement of the delay a queue is imposing on traffic, independent of link speed or packet size.

The algorithm's logic is elegantly minimal. If the minimum sojourn time over a recent interval exceeds a target, typically five milliseconds, CoDel begins dropping packets at a controlled cadence that accelerates until the queue drains below target. It has essentially no tunables that operators need to adjust for their link, which was a deliberate design goal.

PIE, developed at Cisco and standardized for DOCSIS cable systems, reaches a similar conclusion through different mathematics. It estimates current delay from queue length and drain rate, then adjusts a drop probability using a proportional-integral controller. The control-theoretic framing makes PIE's stability properties easier to analyze formally, though the practical behavior converges with CoDel's.

Both algorithms share a deeper insight: queue length is a proxy, and a poor one. A ten-packet queue on a gigabit link is a rounding error; the same queue on a dial-up modem is a catastrophe. Sojourn time normalizes across link speeds automatically because it measures the actual harm being done.

Takeaway

Measure the effect, not the cause. When a metric requires context to interpret, replacing it with one that carries its own context is often the higher-leverage design choice.

Flow Isolation and the FQ-CoDel Synthesis

AQM alone treats a queue as an aggregate, dropping packets without regard to which flow they belong to. This works when all flows are well-behaved TCP, but a single aggressive flow, whether misbehaving software, a non-congestion-controlled UDP stream, or simply a bulk transfer, can starve latency-sensitive traffic sharing the same queue.

Fair queuing addresses this by giving each flow its own logical queue and servicing them in a round-robin or deficit-weighted fashion. Sparse flows, those with little data to send, are naturally prioritized because their queues drain immediately and stay empty. Bulk flows get their fair share of bandwidth but cannot monopolize the link.

FQ-CoDel combines these ideas. It hashes incoming packets into per-flow queues, applies CoDel independently to each queue, and uses a scheduler that gives new or sparse flows preferential access. The result is remarkable: a single link can carry a gigabit bulk download while an interactive SSH session or a video call experiences essentially zero added latency.

The synthesis matters because AQM and fair queuing address different failure modes. AQM handles the vertical dimension of a single queue growing too deep. Fair queuing handles the horizontal dimension of one flow crowding out others. Neither alone is sufficient; together they cover the space of practical misbehavior.

More recent work, particularly the L4S architecture and DualPI2, extends this by separating traffic into two logical queues based on the sender's congestion response, giving scalable congestion controls like DCTCP or Prague sub-millisecond latency while remaining backward compatible with classic TCP. The evolution continues because the ecosystem of endpoints continues to evolve.

Takeaway

Isolation is often cheaper than cooperation. When you cannot trust participants to behave well, structural separation delivers guarantees that policy alone cannot.

The trajectory from RED to CoDel to FQ-CoDel to L4S is not a story of any single algorithm being wrong. It is a story of the network's operating conditions and the endpoints' behaviors coevolving, forcing the queue management layer to keep pace. Each generation solved the previous generation's dominant pathology and revealed the next one.

This is characteristic of mature engineering domains. The problems become subtler, the metrics more refined, and the improvements less spectacular per iteration but no less important in aggregate. Shaving latency from hundreds of milliseconds to tens was dramatic; shaving tens to single digits enables entirely new classes of interactive applications.

The next frontier likely involves AQM at edge and wireless links, where variable capacity complicates every assumption, and tighter integration with congestion control signaling through mechanisms like ECN and L4S. The queue, that most humble of networking primitives, remains a surprisingly deep well of open research.