Skip to content

netjoin/netsplit batches should require that clients MUST NOT delay processing #579

Description

@skizzerz

A client that chooses to delay processing of a netjoin or netsplit batch until the end of the batch is received may cause state desyncs vs what is on the network. The language for this batch type should say that the clients MUST NOT delay processing of batch lines, but MAY delay displaying the results of the netjoin/netsplit in the UI until the batch has been fully received.

In these examples, assume that Client A has negotiated the batch capability and is delaying processing of batch messages until end of batch is received, while Client B has not negotiated the batch capability (and thus processes every line as it comes in). <-- indicates a message that client receives and --> indicates a message that client sends. Let's also assume the remote server sending the netjoin is slow enough that the local server A and B are connected to can process intervening messages.

Client A Client B
<-- BATCH +foo netjoin local.server remote.server
<-- @batch=foo :C!user@host JOIN #channel <-- :C!user@host JOIN #channel
--> MODE #channel +o C
<-- :B!user@host MODE #channel +o C <-- :B!user@host MODE #channel +o C
<-- BATCH -foo

End result: Client A processes the fact that B opped C (a nickname it doesn't even know about) before it processes that C joined the channel. It will likely ignore this spurious op, and when it does process that C joined, it will show them in the user list without operator status. However, on the ircd, C is opped, and Client B would show the correct state due to processing everything immediately.

Client A Client B
<-- BATCH +foo netjoin local.server remote.server
<-- @batch=foo :remote.server MODE #channel -o A <-- :remote.server MODE #channel -o A
--> MODE #channel +o A
<-- :B!user@host MODE #channel +o A <-- :B!user@host MODE #channel +o A
<-- @batch=foo :remote.server NOTICE #channel :*** Notice -- TS for #channel changed from 1769042338 to 1769041945 <-- :remote.server NOTICE #channel :*** Notice -- TS for #channel changed from 1769042338 to 1769041945
<-- @batch=foo :C!user@host JOIN #channel <-- :C!user@host JOIN #channel
<-- BATCH -foo

In this example, the channel on the joining server has a lower TS, which causes it to deop A for splitriding. Because A is delaying processing of netjoin messages, they first see B opping them (which is allowed but does nothing as they are already opped). Then they process the netjoin and see the deop. The end result is that A believes they are not opped when server-side they are.

The only sensible way of solving this issue is for clients to not delay processing of netjoin messages. They can delay their display to the user, but not their processing.

netsplits have a similar issue if a split user rejoins on a local server with the same nick before the client processes the QUIT for them in the netsplit batch. The timing window for this is tighter but if you assume a laggy remote server then it is still possible.

Activity

  1. MrIron-no commented on Feb 3, 2026

    @MrIron-no

    The same applies for modes set by the server in a net merge. If the MODE is sent together with the JOIN, such as:
    @Batch=foo :nick!user@host JOIN #channel
    :local.server.org MODE #channel +o nick
    BATCH -foo

    Clients will show the mode being set before the client has joined. It is significantly more complicated to handle a spec at ircd level that has to first flush all JOIN batches before flushing MODEs set, in particular for ircds handling joins and merge modes as part of the same server<->server protocol message.

  2. jwheare commented on Feb 10, 2026

    @jwheare
    Member

    Can we instead ask servers to not interleave state affecting messages when sending these batches?

  3. MrIron-no commented on Feb 10, 2026

    @MrIron-no

    To avoid a state mismatch a server will want to set mode immediately when a client is bursted to the network/channel.

  4. skizzerz commented on Feb 10, 2026

    @skizzerz
    ContributorAuthor

    Can we instead ask servers to not interleave state affecting messages when sending these batches?

    Not really possible for two different reasons/cases:

    1. In the examples I cited in my OP, these interleaved state messages are coming from other connections, because the joining/splitting server is sending messages very slowly or had a network interruption mid-batch and as such it takes a long time, seconds or even tens of seconds, to finish bursting the batch to other servers. Those other servers can't just stop processing all other messages in the meantime, which means unbatched state changes will be interleaved mid-batch.
    2. For the netjoin burst itself doing state changes, refer to what MrIron said: it is not feasible to not interleave these. When a netjoined user is sent to clients, its channel modes, AWAY message, etc. all get sent as well to ensure clients have a complete picture about that new user. Deferring that state info to the end of the batch isn't feasible on the ircd side; for delayed joins (see point 1) that means there's a period of time where the ircd knows the correct state but hasn't shared with clients, and any state changes that clients make need to somehow backflow into what then gets sent out by the batch.

    The only real solution here is for clients to not delay processing of these batches.

  5. jwheare commented on Feb 11, 2026

    @jwheare
    Member

    OK, my next question is who actually implements this spec and is it useful? If the answer is no one, and not really maybe it should just be retired instead of trying to fix it.

  6. skizzerz commented on Feb 11, 2026

    @skizzerz
    ContributorAuthor

    The use, as I see it, is making it easy for clients to collapse netjoins and netsplits in channel buffers to avoid a lot of visual noise when they happen. Without the batch, it's a heuristic-based approach that may fail (either not including enough, or including too much -- e.g. unrelated joins/quits in the middle of the netjoin/netsplit). I personally think that's fairly useful since both netsplits and netjoins can be really noisy. The proposed wording change I gave still allows that benefit to happen, just with the caveat that failing to do the non-display parts of it right away could cause state desyncs between client and server.

    As for who actually has it, it's being deployed to Libera.Chat within the next 1-2 weeks. If it turns out to be a failed experiment, we can roll that particular piece back I suppose, but I think it overall serves a purpose.

  7. jwheare commented on Feb 11, 2026

    @jwheare
    Member

    I think the motivation is sound, but it might just be impractical. This might explain the lack of implementations up till now, and the obvious flaw you've uncovered when trying to implement it.

    It's not useful for a client to delay and group these visually while simultaneously not grouping them for processing. Visual state is an important signal for users on whether someone saw their message. So it would be confusing and unhelpful for the order of messages to be inaccurate if delayed/grouped separately.

    If servers can't guarantee that the full batch can be sent in one go without delaying other state (which is reasonable) I don't think this is really workable as a spec.

  8. jwheare commented on Feb 11, 2026

    @jwheare
    Member

    Worth noting that this spec predates the IRCv3 policy to require implementations before ratifying specs.

  9. progval commented on Feb 11, 2026

    @progval
    Contributor

    There are use-cases for netjoin for bots. Off the top of my head:

    • Not sending welcome messages/notices (aka. entrymsg) or heralding
    • Antispam bots can use netjoin batches to not interpret these joins as spambots.
  10. jwheare commented on Feb 11, 2026

    @jwheare
    Member

    Can we use a message tag for that instead of a batch?

  11. progval commented on Feb 11, 2026

    @progval
    Contributor

    Yes, but it's not significantly different from a batch without delayed processing.

  12. jwheare commented on Feb 11, 2026

    @jwheare
    Member

    A lot simpler and avoids having to special case override the base batch behaviour.

  13. progval commented on Feb 11, 2026

    @progval
    Contributor

    A lot simpler

    indeed

    avoids having to special case override the base batch behaviour.

    I don't know how hard this is for client devs, because in Limnoria I default to not delaying processing.

  14. jwheare commented on Feb 11, 2026

    @jwheare
    Member

    In IRCCloud we only implement it for multiline so it's not an issue we've had to deal with.

  15. jwheare commented on Feb 11, 2026

    @jwheare
    Member

    I suppose netsplit/netjoin as a batch is just a way to effectively apply a message tag to every message within it. The tag could be used to change the visual appearance of messages, but it's probably not useful to group them separately from the linear message order.

    JOIN/PART/QUIT message grouping is a nice feature, but you don't strictly need a batch to be able to do it (and you'd want to do it for non netsplits too). Rather than making servers keep around a potentially long running dangling state, it's probably best to let clients keep track of that themselves. But a message tag would still be a neat lightweight hint.

  16. MrIron-no commented on Feb 12, 2026

    @MrIron-no

    I agree that a message tag would have been just as good. I think the way of looking this as one batch works (for as long as it accepts interleaving MODE, AWAY etc. coming in as part of a net.merge) in a Client -> Server1 -> Server2 perspective, but it gets messy in a several hop topography, with and without lag. As pointed out by skizzers above, state for batch negotiated clients will be detached from network state, and non-burst messages, such as MODEs and even KICKs may reach the client before the remote net merge has completed (although some ircds will require a KICK to be bounced from the remote server to avoid a zombie state).

    This only works if clients (i) update state for each batched message coming in before the batch has ended; (ii) displays and updates state for interleaving events as if the client had already joined before the batch ended. And then its not really a batch; and a tag would be more appropriate.

  17. MrIron-no commented on Feb 12, 2026

    @MrIron-no

    The rule of thumb could be as simple as: BATCH should never touch network state.

  18. skizzerz commented on Feb 12, 2026

    @skizzerz
    ContributorAuthor

    I'd be fine making this a tag instead of a batch, despite having already implemented the batch. It'd be straightforward to move that to a tag instead, only real difference is I no longer need to send the BATCH verbs at the start/end and can remove some edge case code in the implementation (e.g. "need to close out the netjoin batch early if a server splits in the middle of sending a netjoin burst")

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions