Skip to content

pool_sv2: the listen backlog is fixed at 128 and not tunable, while the accept queue reaches thousands during ramp #717

Description

@gimballock

An operator cannot raise the pool's accept queue. TcpListener::bind takes no backlog argument, so it inherits tokio's value, which delegates to mio::net::TcpListener::bind — hardcoded to 128 on Linux to match std. Because the effective queue is min(backlog, somaxconn), raising net.core.somaxconn has no effect: the requested backlog is the binding term, and nothing exposes it.

Why 128 is small for this workload

In our scale testing the pool's accept queue peaked at 4,410 entries, at t+197 s of a ramp toward 300,000 connections, measured across 672 samples. It returned to 0 in-plateau, so the depth is a ramp phenomenon — which is precisely when a pool takes on miners after a restart or a network event.

A queue reaching 4,410 against a limit of 128 means the kernel drops connection attempts during exactly the window where a pool most wants to accept them. Miners see a failed connect and retry, which adds further arrivals to an already-saturated queue.

What we changed

We build the listener through socket2 and call listen(8192) explicitly. That is one small change at the bind site plus a direct socket2 dependency, which tokio already pulls in transitively.

The value matters less than making it reachable. Any explicit backlog restores an operator's ability to tune, because somaxconn then becomes the binding term rather than a hardcoded 128.

Two limits on our evidence

We never ran an A/B. Every run carried the explicit backlog and somaxconn=8192, so we measured queue depth with headroom present and inferred that a default-configured pool would truncate it. We did not measure drops at backlog 128.

Our sysctl setup was itself unverified. Our tuning script ended each sysctl in || true and never read the values back, and listen(2) returns success even when the kernel silently clamps to min(backlog, somaxconn). So we cannot prove which effective value applied — though a measured depth of 4,410 implies at least that much was in force, because a clamp to 4096 could not produce it.

Happy to send the socket2 patch, or a configurable version if you would rather expose the backlog in the pool config than hardcode a default.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions