You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
ServerHandler.take() acquires a permit of the send window before the writer gets to decide
whether the message is sent at all:
ServerHandler.take() - getNextMessage(), then acquirePermitInSendWindow()
(opendj-server-legacy/src/main/java/org/opends/server/replication/server/ServerHandler.java:991, :994), with no filter in between;
ServerWriter.run() then asks isUpdateMsgFiltered() and drops the message without publishing
it (ServerWriter.java:102, :110).
The only release is ServerHandler.updateWindow() (ServerHandler.java:1105), which is driven by
the WindowMsg the peer sends for the updates it has received and processed - a directory
server in ReplicationBroker.updateWindowAfterReplay() (ReplicationBroker.java:2702-2713, :2518), a peer replication server through decAndCheckWindow() / checkWindow()
(ServerHandler.java:311-331). A message the writer dropped never reached the peer, so no credit
for it ever comes back, and the replication server never probes a peer for credit: only a
directory server sends a WindowProbeMsg.
Every filtered message therefore costs the session one permit for the rest of its life.
When the session stalls
Not after windowSize drops but after more than half of them. The peer gives its credit back in
halves of its window: a directory server once halfRcvWindow replays have accumulated, a
replication server once its receive window falls under rcvWindowSizeHalf. With fewer permits
left than that, the replication server cannot send enough for the peer to get there, and the
writer sits in acquirePermitInSendWindow()'s sendWindow.tryAcquire(500, TimeUnit.MILLISECONDS)
loop (ServerHandler.java:1058-1071) until the session is re-established. Nothing is logged. The
same arithmetic was worked out for #1029 in #1034.
windowSize is what the peer advertised in its start message
(ReplicationServerHandler.java:96, DataServerHandler.java:374), and the default of window-size is 100000 for both a replication server and a replication domain: some 50000 drops.
Which drops
Every return true of ServerWriter.isUpdateMsgFiltered():
WARN_IGNORING_UPDATE_TO_DS_BADGENID - a directory server in BAD_GEN_ID_STATUS (:241);
WARN_IGNORING_UPDATE_TO_DS_FULLUP - a directory server under a total update (:250);
WARN_IGNORING_UPDATE_TO_RS - a peer replication server whose generation id differs (:266).
ReplicationServerDomain.put() applies the same status and generation id checks before it queues
an update for a peer (isUpdateMsgFiltered(), isDifferentGenerationId()), so the drops do not
come from the queue but from the changelog: the catch-up of a peer which connects behind reads its
whole backlog there, and the writer drops every record of it, one permit each. A status or a
generation id which changes between the queueing and the take adds a few more.
The arms do not weigh the same:
the two directory server arms end in a new session: DataServerHandler.changeStatusForResetGenId()
closes the session of a directory server which leaves BAD_GEN_ID_STATUS, and the end of a
total update calls broker.reStart(false) (ReplicationDomain.java:2725). A new handler comes
with a new semaphore, so the lost permits are lost for a session which was receiving nothing
anyway - which is why [#1029] Send a directory server only the updates it gives send-window credit for #1034 left them alone;
the replication server arm does not. ReplicationServerDomain.resetGenerationId() only calls rsHandler.setGenerationId() (:2071), and a peer which re-advertises its generation id in a
TopologyMsg updates the handler in place (ReplicationServerHandler.java:614). A peer which
connected with another generation id and a backlog of more than half a window is left, once the
two agree, with a session over which the replication server can send nothing, until something
breaks it;
the protocol version arm drops one ReplicaOfflineMsg per replica which goes offline, over a
session to an older peer which is never re-established for it either.
This is not new with #1014: at its base the same message was dropped one step later, inside Session.publish(), after the same permit had been taken. The drop moved; the count did not.
Counted as sent
take() also counts the message as sent - incrementOutCount() and the assured counters
(ServerHandler.java:1017-1021) - before the writer drops it, so the sent-updates, assured-sr-sent-updates and assured-sd-sent-updates attributes of the monitor entry of the
handler include what the peer was never sent.
Fix
Give the permit back where the message is dropped, once for every arm, and count a message as sent
where it is published:
Taking the permit only once the filter has passed would keep the accounting in one place, but the
filter reads the status of the peer when the message is written, and take() decides on the
assured flag of a message re-read from the changelog after the wait for the permit on purpose, so
both would have to move with it.
What happens
ServerHandler.take()acquires a permit of the send window before the writer gets to decidewhether the message is sent at all:
ServerHandler.take()-getNextMessage(), thenacquirePermitInSendWindow()(
opendj-server-legacy/src/main/java/org/opends/server/replication/server/ServerHandler.java:991,:994), with no filter in between;ServerWriter.run()then asksisUpdateMsgFiltered()and drops the message without publishingit (
ServerWriter.java:102,:110).The only release is
ServerHandler.updateWindow()(ServerHandler.java:1105), which is driven bythe
WindowMsgthe peer sends for the updates it has received and processed - a directoryserver in
ReplicationBroker.updateWindowAfterReplay()(ReplicationBroker.java:2702-2713,:2518), a peer replication server throughdecAndCheckWindow()/checkWindow()(
ServerHandler.java:311-331). A message the writer dropped never reached the peer, so no creditfor it ever comes back, and the replication server never probes a peer for credit: only a
directory server sends a
WindowProbeMsg.Every filtered message therefore costs the session one permit for the rest of its life.
When the session stalls
Not after
windowSizedrops but after more than half of them. The peer gives its credit back inhalves of its window: a directory server once
halfRcvWindowreplays have accumulated, areplication server once its receive window falls under
rcvWindowSizeHalf. With fewer permitsleft than that, the replication server cannot send enough for the peer to get there, and the
writer sits in
acquirePermitInSendWindow()'ssendWindow.tryAcquire(500, TimeUnit.MILLISECONDS)loop (
ServerHandler.java:1058-1071) until the session is re-established. Nothing is logged. Thesame arithmetic was worked out for #1029 in #1034.
windowSizeis what the peer advertised in its start message(
ReplicationServerHandler.java:96,DataServerHandler.java:374), and the default ofwindow-sizeis 100000 for both a replication server and a replication domain: some 50000 drops.Which drops
Every
return trueofServerWriter.isUpdateMsgFiltered():WARN_IGNORING_UPDATE_UNSUPPORTED_BY_PEER- the protocol version guard of A ReplicaOfflineMsg which cannot be encoded for a peer is still recorded as forwarded: Session.publish() drops it below protocol V8 #1014 (:220);WARN_IGNORING_UPDATE_TO_DS_BADGENID- a directory server inBAD_GEN_ID_STATUS(:241);WARN_IGNORING_UPDATE_TO_DS_FULLUP- a directory server under a total update (:250);WARN_IGNORING_UPDATE_TO_RS- a peer replication server whose generation id differs (:266).ReplicationServerDomain.put()applies the same status and generation id checks before it queuesan update for a peer (
isUpdateMsgFiltered(),isDifferentGenerationId()), so the drops do notcome from the queue but from the changelog: the catch-up of a peer which connects behind reads its
whole backlog there, and the writer drops every record of it, one permit each. A status or a
generation id which changes between the queueing and the take adds a few more.
The arms do not weigh the same:
DataServerHandler.changeStatusForResetGenId()closes the session of a directory server which leaves
BAD_GEN_ID_STATUS, and the end of atotal update calls
broker.reStart(false)(ReplicationDomain.java:2725). A new handler comeswith a new semaphore, so the lost permits are lost for a session which was receiving nothing
anyway - which is why [#1029] Send a directory server only the updates it gives send-window credit for #1034 left them alone;
ReplicationServerDomain.resetGenerationId()only callsrsHandler.setGenerationId()(:2071), and a peer which re-advertises its generation id in aTopologyMsg updates the handler in place (
ReplicationServerHandler.java:614). A peer whichconnected with another generation id and a backlog of more than half a window is left, once the
two agree, with a session over which the replication server can send nothing, until something
breaks it;
ReplicaOfflineMsgper replica which goes offline, over asession to an older peer which is never re-established for it either.
This is not new with #1014: at its base the same message was dropped one step later, inside
Session.publish(), after the same permit had been taken. The drop moved; the count did not.Counted as sent
take()also counts the message as sent -incrementOutCount()and the assured counters(
ServerHandler.java:1017-1021) - before the writer drops it, so thesent-updates,assured-sr-sent-updatesandassured-sd-sent-updatesattributes of the monitor entry of thehandler include what the peer was never sent.
Fix
Give the permit back where the message is dropped, once for every arm, and count a message as sent
where it is published:
Taking the permit only once the filter has passed would keep the accounting in one place, but the
filter reads the status of the peer when the message is written, and
take()decides on theassured flag of a message re-read from the changelog after the wait for the permit on purpose, so
both would have to move with it.
Found while reviewing #1019 (@maximthomas).