You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
We are investigating established-session traffic failures using Myst 1.39.5-
derived consumers. We have saved, session-fenced traces from the same providers
used concurrently on two environments. We are seeking help correlating these
with provider-side session/payment events, not claiming a proven Myst defect.
Environment and modifications
OKE: one ARM64 Linux worker, 2 OCPUs / 12 GB, host-networked consumer pod.
Mac: ARM64 Linux Docker consumer on a residential network. This is not the
native Mysterium Dark client.
Separate existing consumer identities; one daemon per environment, multiple
proxy-port-keyed sessions, target 150 public US-residential proposals per side.
Myst 1.39.5 source with local atomic consumer payment-accounting correction,
session-reply timeout adjustment and metadata-only diagnostics. OKE also has
a one-line outgoing-TUN-batch allocation correction. These are not stock
binaries, and builds/load/identity differences limit causal comparisons.
Public-provider versions and internal logs are unknown.
Observed failure pattern
In a September 19, 2026 16:00–17:00 UTC comparison, 16 overlapping provider pairs
had one side fail first: eleven OKE-first, five Mac-first. In each case the same
provider's unchanged counterpart later passed HTTPS at least five minutes after
the failed side's removal. None of these paired removals occurred within one
minute of each other. Four counterparts eventually failed later, 6.7–18.8 minutes
apart. This does not look like simultaneous whole-provider outages.
For all 16 pairs we retained two minutes before and five minutes after the first
removal, matched to the exact session, proxy port, WireGuard generation and UDP
binding. Failed-side one-second traces had maximum gaps of 1.041 seconds.
Fifteen failed sessions received no encrypted tunnel packets during the failed
routine request through removal, while continuing to send.
One case received two early packets and completed a fresh WG handshake, but
the request still failed. Please treat it separately from packet silence.
All sixteen recorded DNS operations timing out before HTTP CONNECT completed.
Three independent curl checks through the same proxy also timed out without
CONNECT responses before removal. We did not inject direct-IP traffic into
these specific sessions.
The provider still answered 8–17 P2P keepalive requests after the failed side's
last incoming tunnel packet. Five cases had an isolated keepalive timeout but
continued successful replies.
No recorded endpoint change or increment in invalid TCP checksum counts
occurred in these pre-removal spans. Proxy listeners remained present in the
sampled pre-removal interval.
Example anonymized OKE case P02 (UTC): last incoming WG packet 16:38:53.650;
first failed check starts 16:39:23.330; removal requested 16:40:03.374. Fourteen
keepalive replies followed the last incoming packet. A successful payment-message
reply arrived at 16:39:01.625. The Mac counterpart continued passing traffic.
Payment evidence and limits
All sixteen failed-side windows contain committed consumer payment metadata and
a successful P2P payment-message reply; fourteen replies arrived after the last
incoming tunnel packet. No payment-message error replies or consumer Hermes
errors were recorded in those short windows.
We do not interpret P2P success as invoice/Hermes acceptance. In the inspected
provider source, the handler queues an exchange on paymentEngineChan and
returns before the payment engine's later validation/Hermes work. We cannot see
that later outcome on public providers. No session-destroy message was observed,
but absence does not prove that the provider kept its session open.
Prior work, to avoid repeating broad tests
Similar failures have appeared on OKE, Mac Docker/native tests and GKE; older
comparisons had configuration differences and do not isolate one cloud.
Kernel versus proxy tests did not establish a general kernel-mode remedy.
Broader UDP firewall allowance did not restore reliable operation; historical
flow-log coverage was incomplete, so this does not clear every network layer.
Separate controlled Mac IP-change cases are excluded from this cohort's cause
claim. An owned-provider invoice-registration race was fixed and verified,
but that fix is not deployed to the public providers.
We found and corrected admission/refill and memory-pressure defects locally.
These improve usable capacity, not proof of a tunnel-drop cure.
Specific questions
Does this resemble a regression or related failure from Connection tunnel collapse on failed handshakes #5492, marked fixed
in 1.18.3? That release includes Fix resource allocation in the service #5520, sharing resource allocation across
provider services. Our inspected 1.39.5 source already has the shared allocator;
we have not established the public providers' versions or allocator behavior.
With privately supplied session IDs, provider IDs and UTC timestamps, can you
correlate provider session-close reasons, resource allocation/release,
invoice validation and Hermes outcomes for these pairs?
Can the control channel remain responsive/acknowledge queued payments after
the relevant WG session or forwarding state has failed in current versions?
Is there a supported private channel for the minimal correlation mapping and
redacted metadata? We will not post keys, signatures, payment payloads,
customer payloads or raw wallet data.
We can provide anonymized timelines publicly and session/endpoint mappings
privately after approval. We have no deterministic reproduction of the complete
public-provider failure and are not asking that timing correlation alone be
accepted as proof.
We are investigating established-session traffic failures using Myst 1.39.5-
derived consumers. We have saved, session-fenced traces from the same providers
used concurrently on two environments. We are seeking help correlating these
with provider-side session/payment events, not claiming a proven Myst defect.
Environment and modifications
native Mysterium Dark client.
proxy-port-keyed sessions, target 150 public US-residential proposals per side.
session-reply timeout adjustment and metadata-only diagnostics. OKE also has
a one-line outgoing-TUN-batch allocation correction. These are not stock
binaries, and builds/load/identity differences limit causal comparisons.
c45527af1ea80300ae3d9c92bd37255335b6140d.github.com/mysteriumnetwork/wireguard-goat406b13e8996a(2024-04-16).f29b30f1098165dea0e30386db212f4692116a607e1fe6e2fd35698c770e516c.919f9cde736b51cc5ab387c67ec5c898b4ce2e0e01402b47a4d018d2ed226bff.Observed failure pattern
In a September 19, 2026 16:00–17:00 UTC comparison, 16 overlapping provider pairs
had one side fail first: eleven OKE-first, five Mac-first. In each case the same
provider's unchanged counterpart later passed HTTPS at least five minutes after
the failed side's removal. None of these paired removals occurred within one
minute of each other. Four counterparts eventually failed later, 6.7–18.8 minutes
apart. This does not look like simultaneous whole-provider outages.
For all 16 pairs we retained two minutes before and five minutes after the first
removal, matched to the exact session, proxy port, WireGuard generation and UDP
binding. Failed-side one-second traces had maximum gaps of 1.041 seconds.
routine request through removal, while continuing to send.
the request still failed. Please treat it separately from packet silence.
Three independent curl checks through the same proxy also timed out without
CONNECT responses before removal. We did not inject direct-IP traffic into
these specific sessions.
last incoming tunnel packet. Five cases had an isolated keepalive timeout but
continued successful replies.
occurred in these pre-removal spans. Proxy listeners remained present in the
sampled pre-removal interval.
Example anonymized OKE case P02 (UTC): last incoming WG packet 16:38:53.650;
first failed check starts 16:39:23.330; removal requested 16:40:03.374. Fourteen
keepalive replies followed the last incoming packet. A successful payment-message
reply arrived at 16:39:01.625. The Mac counterpart continued passing traffic.
Payment evidence and limits
All sixteen failed-side windows contain committed consumer payment metadata and
a successful P2P payment-message reply; fourteen replies arrived after the last
incoming tunnel packet. No payment-message error replies or consumer Hermes
errors were recorded in those short windows.
We do not interpret P2P success as invoice/Hermes acceptance. In the inspected
provider source, the handler queues an exchange on
paymentEngineChanandreturns before the payment engine's later validation/Hermes work. We cannot see
that later outcome on public providers. No session-destroy message was observed,
but absence does not prove that the provider kept its session open.
Prior work, to avoid repeating broad tests
comparisons had configuration differences and do not isolate one cloud.
flow-log coverage was incomplete, so this does not clear every network layer.
claim. An owned-provider invoice-registration race was fixed and verified,
but that fix is not deployed to the public providers.
These improve usable capacity, not proof of a tunnel-drop cure.
Specific questions
in 1.18.3? That release includes Fix resource allocation in the service #5520, sharing resource allocation across
provider services. Our inspected 1.39.5 source already has the shared allocator;
we have not established the public providers' versions or allocator behavior.
correlate provider session-close reasons, resource allocation/release,
invoice validation and Hermes outcomes for these pairs?
the relevant WG session or forwarding state has failed in current versions?
redacted metadata? We will not post keys, signatures, payment payloads,
customer payloads or raw wallet data.
We can provide anonymized timelines publicly and session/endpoint mappings
privately after approval. We have no deterministic reproduction of the complete
public-provider failure and are not asking that timing correlation alone be
accepted as proof.