Skip to content

A message the writer filters keeps the send-window permit it took #1080

Description

@vharseko

What happens

ServerHandler.take() acquires a permit of the send window before the writer gets to decide
whether the message is sent at all:

  • ServerHandler.take() - getNextMessage(), then acquirePermitInSendWindow()
    (opendj-server-legacy/src/main/java/org/opends/server/replication/server/ServerHandler.java:991,
    :994), with no filter in between;
  • ServerWriter.run() then asks isUpdateMsgFiltered() and drops the message without publishing
    it (ServerWriter.java:102, :110).

The only release is ServerHandler.updateWindow() (ServerHandler.java:1105), which is driven by
the WindowMsg the peer sends for the updates it has received and processed - a directory
server in ReplicationBroker.updateWindowAfterReplay() (ReplicationBroker.java:2702-2713,
:2518), a peer replication server through decAndCheckWindow() / checkWindow()
(ServerHandler.java:311-331). A message the writer dropped never reached the peer, so no credit
for it ever comes back, and the replication server never probes a peer for credit: only a
directory server sends a WindowProbeMsg.

Every filtered message therefore costs the session one permit for the rest of its life.

When the session stalls

Not after windowSize drops but after more than half of them. The peer gives its credit back in
halves of its window: a directory server once halfRcvWindow replays have accumulated, a
replication server once its receive window falls under rcvWindowSizeHalf. With fewer permits
left than that, the replication server cannot send enough for the peer to get there, and the
writer sits in acquirePermitInSendWindow()'s sendWindow.tryAcquire(500, TimeUnit.MILLISECONDS)
loop (ServerHandler.java:1058-1071) until the session is re-established. Nothing is logged. The
same arithmetic was worked out for #1029 in #1034.

windowSize is what the peer advertised in its start message
(ReplicationServerHandler.java:96, DataServerHandler.java:374), and the default of
window-size is 100000 for both a replication server and a replication domain: some 50000 drops.

Which drops

Every return true of ServerWriter.isUpdateMsgFiltered():

ReplicationServerDomain.put() applies the same status and generation id checks before it queues
an update for a peer (isUpdateMsgFiltered(), isDifferentGenerationId()), so the drops do not
come from the queue but from the changelog: the catch-up of a peer which connects behind reads its
whole backlog there, and the writer drops every record of it, one permit each. A status or a
generation id which changes between the queueing and the take adds a few more.

The arms do not weigh the same:

  • the two directory server arms end in a new session: DataServerHandler.changeStatusForResetGenId()
    closes the session of a directory server which leaves BAD_GEN_ID_STATUS, and the end of a
    total update calls broker.reStart(false) (ReplicationDomain.java:2725). A new handler comes
    with a new semaphore, so the lost permits are lost for a session which was receiving nothing
    anyway - which is why [#1029] Send a directory server only the updates it gives send-window credit for #1034 left them alone;
  • the replication server arm does not. ReplicationServerDomain.resetGenerationId() only calls
    rsHandler.setGenerationId() (:2071), and a peer which re-advertises its generation id in a
    TopologyMsg updates the handler in place (ReplicationServerHandler.java:614). A peer which
    connected with another generation id and a backlog of more than half a window is left, once the
    two agree, with a session over which the replication server can send nothing, until something
    breaks it;
  • the protocol version arm drops one ReplicaOfflineMsg per replica which goes offline, over a
    session to an older peer which is never re-established for it either.

This is not new with #1014: at its base the same message was dropped one step later, inside
Session.publish(), after the same permit had been taken. The drop moved; the count did not.

Counted as sent

take() also counts the message as sent - incrementOutCount() and the assured counters
(ServerHandler.java:1017-1021) - before the writer drops it, so the sent-updates,
assured-sr-sent-updates and assured-sd-sent-updates attributes of the monitor entry of the
handler include what the peer was never sent.

Fix

Give the permit back where the message is dropped, once for every arm, and count a message as sent
where it is published:

// ServerWriter.run()
if (isUpdateMsgFiltered(updateMsg))
{
  handler.releasePermitInSendWindow();
  ...
}
else
{
  handler.countSentUpdate(updateMsg);
  ...
}

Taking the permit only once the filter has passed would keep the accounting in one place, but the
filter reads the status of the peer when the message is written, and take() decides on the
assured flag of a message re-read from the changelog after the wait for the permit on purpose, so
both would have to move with it.

Found while reviewing #1019 (@maximthomas).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions