Skip to content

Review: the fork against Evolution API 2.3.7 - #2

Open
AlmogCohen wants to merge 157 commits into
release/2.3.7from
fork/2.3.7
Open

AlmogCohen wants to merge 157 commits into
release/2.3.7from
fork/2.3.7

Conversation

@AlmogCohen

@AlmogCohen AlmogCohen commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

The whole fork as one reviewable diff: fork/2.3.7 against the untouched Evolution API 2.3.7 release (release/2.3.7, at the 2.3.7 tag). This PR is a review view and will not be merged.

Unofficial fork, not endorsed by Evolution Foundation. The license and NOTICE are kept, and the user interface and brand assets are unchanged.

🤖 Generated with Claude Code

AlmogCohen and others added 30 commits September 27, 2026 13:55
Adds FORK.md (why the fork exists, the red-then-green rule, the base) and
NOTICE (Evolution Foundation's attribution from upstream main, plus the
modification notice Apache 2.0 section 4 asks for).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2.3.7 ships 7.0.0-rc.9, which is inside the range of CVE-2026-48063
(message and history-sync spoofing, rc.1 to rc11, GHSA-qvv5-jq5g-4cgg).
rc13 is also the first to learn @lid to phone mappings from the history
sync. Pinned exactly: an RC is not a range.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…Baileys

Evolution has no test suite, and test/ has been in .gitignore since the first
commit. This adds one: vitest 4.1.11 runs the TypeScript source directly, the
real BaileysStartupService is driven through Baileys' real event buffer, and
Prisma and the server module are replaced by in-memory fakes.

- Baileys resolves from the installed package (package.json pins it), and
  deep imports reach the same build. BAILEYS_DIR points the suite at another
  build; BAILEYS_EXPECT fails the run when the resolved version differs.
- Tests start from a production configuration (the minimal profile:
  only the instance stored, local cache, WARN and above). A test that needs
  storage asks for the stored profile.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Typecheck and the vitest suite on the pinned Baileys, on every push to a
fork/ branch and on pull requests. A weekly job runs the same suite
against the newest Baileys release. The upstream workflows are left as they
are: none of them triggers on fork/ branches.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
AGENTS.md gains a first section for this fork: every change starts with a
failing test committed alone, then the fix; how the harness runs; profiles;
what to assert. CLAUDE.md points at it. These are this fork's rules, not
Evolution's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each of the three auth stores (prisma, redis-db, provider-files) saves a
real app-state sync key, is opened again as after a restart, and decodes a
real encrypted contact patch with the key it reads back. Fails today: the
stores revive the key with AppStateSyncKeyData.create(), which leaves
keyData a base64 string, so the patch is skipped and no name comes back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
After a JSON round trip an app-state sync key's keyData is a base64
string. AppStateSyncKeyData.create() keeps it a string, so Baileys derives
the wrong keys and skips every app-state patch after a restart. fromObject()
decodes it back to bytes, as Baileys' own multi-file store does. Applied to
the prisma, redis-db and provider-files stores.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Baileys emits groups.update, which every transport upper-cases to
GROUPS_UPDATE before comparing it with the subscription. /webhook/set (and
every /<transport>/set) accepts only GROUP_UPDATE, and the global env
config is keyed GROUP_UPDATE, so a subscription to group updates never
matches. Through the real WebhookController to a local HTTP destination:
subscribed as GROUP_UPDATE (the spelling /webhook/set accepts) or GROUPS_UPDATE, per instance
and global, the event must arrive, and other events must be unaffected.
The same checks run through the websocket, rabbitmq, nats, sqs, kafka and
pusher controllers with recording clients. Fails today for GROUP_UPDATE,
per instance in all seven transports and global in webhook, rabbitmq,
nats, kafka and pusher.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The event is groups.update (GROUPS_UPDATE once upper-cased), but the
/<transport>/set schemas store GROUP_UPDATE and the global env configs are
keyed GROUP_UPDATE (except SQS's), so group updates reached no subscriber.
A shared isSubscribed() in event.controller.ts accepts both spellings and
replaces the identical comparisons in the webhook, websocket, rabbitmq,
sqs, nats, kafka and pusher controllers, per instance and global. No stored
value is renamed, so existing subscriptions keep working.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Baileys calls the socket's getMessage to answer a retry request for a
message this instance sent once the message has left its recent-message
cache, and relays any truthy answer under the original id. Evolution
answers a database miss with { conversation: '' }. Under the minimal
profile (no messages stored) every lookup misses. The test asks for an
unknown id under both profiles and for a sent message under minimal, and
expects undefined; a message stored under the stored profile must still
come back as stored. Fails today with { conversation: '' }.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A lookup that finds no row used to throw on webMessageInfo[0].message and
be answered by the catch with { conversation: '' }, which Baileys relays
as the retried message under the original id. Return undefined when
nothing is stored, which Baileys reads as "not available" and sends
nothing. A lookup that throws (database error) still gets the old answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
….upsert

Baileys (rc13+) learns which phone number each private @lid belongs to and
hands the mappings over as lidPnMappings on messaging-history.set and as
lid-mapping.update events. Evolution 2.3.7 reads neither, so a consumer never
learns that an @lid chat is a known phone number.

The test drives history through the production path (Baileys' own
processMessage inside ev.createBufferedFunction, a real signal repository)
and asserts the exact CONTACTS_UPSERT items a consumer reads a mapping from:
{ remoteJid: <phone>, pushName: null, lid: <lid>, phoneNumber: <phone>,
instanceId }.

On the unchanged code all seven cases fail with no mapping item emitted at
all (expected [ ...mappings ], received []), because nothing forwards either
source:
- (a) an @lid-keyed conversation, mapped by phoneNumberToLidMappings or its
  own pnJid
- (b) a phone-keyed conversation whose lidJid is set
- (c) a mapping with no conversation of its own
- (d) the mapping in a later, and in an earlier, history batch than its
  conversation
- (e) a live lid-mapping.update from pnForLidChatAction
- the negative control (no mapping for an unmapped @lid, no unrelated pair
  linked), whose positive half expects the two real pairs

Forwarding the event field alone is not enough: Baileys' event buffer drops
lidPnMappings when it flushes a buffered messaging-history.set (rc14
lib/Utils/event-buffer.js, consolidateEvents). Checked against the spike fix
that recovers mappings by asking the signal repository for each @lid chat or
contact id: (a), (e) and the earlier-batch (d) pass, while (b), (c), the
later-batch (d) and the negative control's unconversationed pair still fail.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every @lid to phone mapping Baileys learns now reaches the consumer as one
CONTACTS_UPSERT item: { remoteJid: <phone>, pushName: null, lid: <lid>,
phoneNumber: <phone>, instanceId }. Pairs are oriented, normalized to user
jids and de-duplicated within a batch; anything that is not an @lid and a
phone jid is skipped.

- Live mappings: lid-mapping.update is not a bufferable event, so it reaches
  ev.process intact and is forwarded there.
- History mappings: Baileys puts every mapping of a batch on
  messaging-history.set as lidPnMappings (phoneNumberToLidMappings, plus each
  conversation's pnJid or lidJid, lib/Utils/history.js:45-78), but history is
  processed inside ev.createBufferedFunction and the buffer drops the field
  when it consolidates the batch (lib/Utils/event-buffer.js:543-556). In the
  test's buffered path the handler receives lidPnMappings undefined in every
  batch. So eventHandler() now wraps client.ev.emit once per socket and reads
  lidPnMappings where Baileys emits it, before the buffer sees it, queued on
  eventProcessingQueue like every other event. Plain TypeScript, Baileys is
  not patched, and nothing is looked up.

Chosen because it is the only option that covers every case. Rejected:
- Reading lid / phoneNumber off the contacts Baileys builds from history:
  they survive the buffer, but a contact carries only its own conversation's
  lidJid or pnJid (history.js:56-62). A mapping that exists only in
  phoneNumberToLidMappings (history.js:45-49) has no contact, so an @lid
  conversation mapped that way, a mapping with no conversation, and a
  mapping in a later batch than its conversation are all lost.
- Looking the mapping up in the signal repository: getPNForLID is local
  (cache, then the keys store, lib/Signal/lid-mapping.js:199-270), but
  getLIDForPN falls back to a USYNC query on a miss (lid-mapping.js:155-176),
  so a phone-keyed conversation cannot be resolved without the network, and
  the key store has no way to enumerate its entries, so a mapping with no
  conversation id to ask about, or one arriving in a batch with no
  conversations, cannot be found at all.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… certain

Each contacts.upsert webhook item now carries saved. It is true only for a
name from an app-state contact action (fullName or firstName), and false for
history names (displayName, chat name or username), push names, a message's
pushName, and a contact action name that is just the username. Every existing
field is unchanged, and saved is not written to the stored row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…R,WARN

Drives each path that handles a message, contact or call with a distinctive
text, phone number and push name, under the minimal and stored profiles,
and asserts none of the three is printed. captureOutput now also catches
direct writes to fd 1 and 2, which is how pino (the logger Evolution hands
Baileys) prints.

Paths that leak on 2.3.7:
- ordinary incoming message and placeholder resend: console.log(messageRaw)
  prints text, sender number and push name; the resend also prints the
  whole message after "Message received from phone"
- on-demand history sync: console.log dumps every message in the batch
- message that failed to decrypt ("No session record"): Baileys' error log
  "failed to decrypt message" prints the key with the sender's JID, and
  Evolution's WARN "Message ignored with messageStubParameters" prints the
  whole message with the JID and push name
- Baileys "error in handling message": prints the stanza (from, notify,
  body) through the pino logger Evolution configures
- incoming call: console.log of the CB:call and CB:ack,class:call stanzas
  prints the caller's JID and push name
- group metadata cache lookup: console.log "Cache request for group" prints
  the group JID, which embeds the creator's number in old-format ids
- group participants update when the lookup fails: ERROR prints the group JID
- raw node through baileysSendNode: console.log prints the whole stanza
- stored profile only: "Original message not found for update" WARN prints
  the key, and "Chat insert record ignored" prints the chat JID
- status update under minimal: a bare console.log "CACHE:" bypasses LOG_LEVEL

Guards that already pass: a history batch and an outgoing text message.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A deployment with LOG_LEVEL=ERROR,WARN still printed message text, phone
numbers and push names: bare console.log dumps that ignore LOG_LEVEL, WARN
and ERROR lines that serialised whole messages or keys, and the pino logger
handed to Baileys, which at level error prints message keys, sender JIDs
and whole stanzas.

- Baileys gets makeBaileysLogger (src/utils/log-privacy.ts): the same
  level, but every logged object is reduced to message id, chat type,
  message type, the session-record flag, counts and the error's name and
  message, with anything shaped like a JID or a phone number masked. Used
  for the socket, the signal key store and media downloads.
- Removed console.log(messageRaw); the full payload stays available at
  VERBOSE through the existing logger.verbose.
- On-demand history, "Message received from phone", CB:call and
  CB:ack,class:call, CACHE, group cache lookups and baileysSendNode now go
  through the logger (INFO or VERBOSE) with ids and counts only.
- "Message ignored with messageStubParameters" prints id, chat type and the
  matched decrypt error instead of the whole message.
- "Original message not found for update" prints the message id, not the
  key; "Chat insert record ignored" prints the chat type, not the JID; the
  group participants error drops the group JID.

Webhook payloads and every other behaviour are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A local CDN and local HTTP CONNECT and SOCKS5 proxies on 127.0.0.1, with the
instance's proxy set the way /proxy/set and a connect set it. Downloads via
getBase64FromMediaMessage (the /chat/getBase64FromMediaMessage route and the
S3 storage path) and the messages.upsert webhook base64 path must reach the
CDN through the proxy. On the current code they reach it directly: the socket
gets the proxy, but Baileys downloads with fetch and only a `dispatcher` in
the download options routes it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Baileys 7 downloads media with fetch, which ignores the socket's agent and
fetchAgent and routes only through a `dispatcher` in the download options.
Every downloadMediaMessage and downloadContentFromMessage call in the Baileys
service (getBase64FromMediaMessage and its fallback, and the webhook base64
downloads in messages.upsert and after sending) now passes the instance's
proxy as an undici dispatcher (ProxyAgent for http/https, fetch-socks for
socks4/socks5), built by the existing makeProxyAgentUndici. It is the
socket's own fetchAgent when the socket was created with one, so a
proxyscrape exit is the same exit, and it is rebuilt when /proxy/set changes
the proxy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gh it

Runs Evolution's real connect with makeWASocket replaced by a spy, then runs
Baileys' own getWAUploadToServer with the socket config Evolution built,
against a local HTTPS media host (a test certificate added to the default CAs,
verification on) through local HTTP CONNECT and SOCKS5 proxies. Today the
upload fails on every host: Baileys uploads with http.request under Node and
passes fetchAgent as its agent, and Evolution's fetchAgent is an undici
dispatcher, which http.request refuses with ERR_INVALID_ARG_TYPE.

A connection guard in the helpers refuses any socket not to the loopback
address and records what tried.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Under Node, Baileys uploads media with http.request and passes the socket's
fetchAgent as its agent. Evolution set fetchAgent to an undici dispatcher,
which http.request refuses, so every media send on an instance with a proxy
failed with "Media upload failed on all hosts". fetchAgent is now the same
kind of agent as the websocket's (makeProxyAgent: https-proxy-agent or
socks-proxy-agent), and the undici dispatcher is kept for downloads only.
The proxy is resolved once per connect, so a proxyscrape exit stays one exit
for the socket, uploads and downloads.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Runs Evolution's real connect with a proxy set and makeWASocket replaced by a
spy. The proxy maps web.whatsapp.com and raw.githubusercontent.com to local
HTTPS servers, and the connection guard refuses anything else. Today both the
sw.js request (axios) and the fallback to Baileys' version file (fetch) go
straight out, so the guard records web.whatsapp.com:443 and
raw.githubusercontent.com:443 and the proxy sees nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…roxy

createClient asked web.whatsapp.com for its version (and, failing that,
GitHub for Baileys' version file) before resolving the instance's proxy, so
every account's connect was preceded by a request from the server's own IP.
The proxy is now resolved first, the sw.js request (axios) goes through the
socket's proxy agent with axios' own env proxy off, and the Baileys fallback
(fetch) through the undici dispatcher. Without a proxy nothing changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
connectToWhatsapp starts loadProxy() without waiting for it and goes on to
build the socket. With the Proxy row read taking a database round trip
(50ms here), the socket config carries no proxy agent and the version fetch
goes out directly: the connection guard records web.whatsapp.com and
raw.githubusercontent.com, and a connection opened with the socket's agent
goes straight out too. Covered for the first connect after boot and for a
second connect of the same service (restart, reconnect after close), where
loadProxy first sets enabled to false and so drops a proxy that was already
loaded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
connectToWhatsapp now awaits loadProxy() before createClient, so the socket,
the version fetch and media are built from the Proxy row rather than from
whatever the read had reached. Every connect goes through here: the boot
auto-connect, /instance/create, /instance/connect, /instance/restart, the
reconnect after a closed connection and the Chatwoot reconnect. A failed
Proxy read now fails the connect instead of connecting without the proxy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
connectToWhatsapp starts loadSettings() without waiting for it and builds the
socket from localSettings. With the Setting row read taking a database round
trip (50ms here), the first connect after boot gets syncFullHistory and
markOnlineOnConnect undefined, and shouldIgnoreJid keeps groups that the
instance ignores. Two stored rows, since with full history on Evolution keeps
groups even when groupsIgnore is set. The reconnect cases hold on the current
code (loadSettings does not clear before reading) and guard that path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
connectToWhatsapp now awaits loadSettings() before createClient, so the
socket is built from the instance's Setting row (syncFullHistory,
groupsIgnore, readStatus, alwaysOnline, and the wavoip token read right
after) rather than from defaults when the read is slower than the rest of
the connect.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… session

keyExists and getAuthKey swallowed every database error and returned
false/null, so opening the Prisma auth store during a database blip read
as "no session": it started fresh creds (initAuthCreds) and saved them
over the linked session once the database answered, unlinking the account.

A failed session read now propagates, so the open (and the connect) fails
and can be retried with the stored creds intact. A genuinely absent row
still starts a fresh session. saveKey keeps its existence check inside its
own try, so a failed check skips the save instead of creating a second row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
AlmogCohen and others added 28 commits September 27, 2026 23:00
…that finishes late

- A queued batch checks, when it runs, whether the instance was shut down
  meanwhile, and does nothing if so.
- sendDataWebhook forwards nothing once the instance is shut down (a
  picture lookup, an open's handler or any other late work), except the
  monitor's own remove.instance announcement, sent after the shutdown
  (new test guards it).

Not applied, on purpose: dropping a batch whose socket a reconnect of the
same session replaced while it waited. Its messages were decrypted,
receipted and acknowledged to WhatsApp, which will not send them again,
so dropping them would lose them. A logout already drops batches at run
time (the logout branch), and the queue keeps a logout's close behind the
batches queued before it.

Visible to consumers: no webhook (and no stored row) arrives from an
instance after it was removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cture

contactPictures sends a history batch's pictures on contacts.update once
all its lookups finish, with the name each contact had in the batch; and
profilePicture writes its answer to the cache whatever happened while it
waited.

Red on the current code (2 tests): a contact renamed while its batch's
lookups were running ends, for a consumer applying the webhooks in order,
with the old name; a lookup answered after WhatsApp's 'removed' picture
notification puts the old picture back on the consumer and in the cache.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er outlives a picture notification

- contactPictures sends each contact's newest name: a rename that
  arrives (contacts.update) while the batch's lookups run replaces the
  batch's name for that update and for Chatwoot. The payload keeps its
  shape (remoteJid, pushName, profilePicUrl, instanceId), which the
  recorded live fixture holds it to.
- WhatsApp's picture notification (imgUrl 'changed' or 'removed') bumps a
  per-contact version and drops the kept picture and the lookup under
  way. A lookup started before it neither writes the cache nor is sent on
  the batch's update.

Visible to consumers: a history batch's picture update no longer carries
a stale name or a picture WhatsApp said was removed; otherwise the same.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ress

loadProxy switched the proxy off before reading the Proxy row; a failed
proxyscrape list fetch switched it off before the socket was built; and
a media download for a proxyscrape proxy with no socket exit to share
returned no dispatcher.

Red on the current code (4 tests): a download during a proxy reload goes
directly; after a failed Proxy read every download goes directly; a
connect whose proxyscrape list is down builds a socket with no agent; a
proxyscrape download with no socket exit goes directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ve from the server address

- loadProxy builds the new configuration apart and puts it in force in
  one step after reading the Proxy row; until then the one in force
  stays, and a failed read changes nothing.
- A proxyscrape list that cannot be fetched, or is empty, fails the
  connect (retried by the reconnect) instead of disabling the proxy.
- A media download for a proxyscrape proxy uses the socket's own exit
  even after the configuration was reloaded, and fails (400) when there
  is none yet, instead of going directly.

Visible to consumers: with a proxyscrape proxy whose list is down, the
instance stays disconnected (and keeps retrying) instead of connecting
from the server's address; a download for such an instance before it
has connected answers 400.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…om an error message

An exception's message is arbitrary text: JSON.parse quotes the input it
failed on, and a library error can carry a number as a person typed it,
with spaces or dashes. errorFields masked URLs, JIDs and runs of 6+
digits only.

Red on the current code: a JSON.parse error keeps the quoted text, and
+972 54-111-2233 / 054-111-2233 survive.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eparators

scrub() (used by errorFields and the Baileys logger) now also masks any
quoted part of a message (JSON.parse repeats the input it failed on) and
a run of 7+ digits written with spaces, dashes, dots or brackets, as a
person types a phone number.

Not done here, and why: a name in the unquoted words of an exception
message is not recognisable by pattern, and many inherited catch sites
still log raw errors; the FORK.md claim about logs is narrowed instead
(docs commit). The review's S3 example did not reproduce: a failed
download reaching the S3 upload catch is already the bounded error
getBase64FromMediaMessage throws (checked with a CDN answering 410; the
line logged holds no URL).

Visible to consumers: log lines only (quoted parts and formatted numbers
become [quoted] / [number]).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… only under the @lid

Baileys gives a DM WhatsApp addresses by phone the @lid in remoteJidAlt
(sender_lid) with addressingMode 'pn': the same key the messages.upsert
webhook shows for an @lid-addressed DM after its swap. originalMessageKey
turns every such key into the @lid form, so the phone, which keeps this
message under the phone JID, refuses the re-upload.

Red on the current code: the request names the @lid only, the phone
answers NOT_FOUND, and the download fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r the key as given

A key with the phone JID and its @lid beside it (addressingMode 'pn')
means either an @lid-addressed DM as the webhook shows it, or a DM
WhatsApp addresses by phone as Baileys gave it; nothing in the key tells
them apart. The re-upload asks under the original (@lid) form first, as
before, and when the phone refuses it (a MediaRetryNotification result,
NOT_FOUND and the like, not a timeout) asks once more under the key as
given, within the same 60s budget.

Not done as the review proposed (carrying the original key in the
webhook payload): the payload's key shape is upstream develop's swap,
which consumers already rely on; the retry needs no new field.

Visible to consumers: /chat/getBase64FromMediaMessage now recovers the
media of a phone-addressed DM that failed with reuploadReason NOT_FOUND.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ate answer still acts

Baileys' updateMediaMessage waits for the phone's answer with no timeout
(bindWaitForEvent: a messages.media-update and a connection.update
listener). Evolution's 60s race stops waiting but not the wait.

Red on the current code: after the download fails with no_answer, one
media-update and one connection.update listener remain, and an answer
arriving later still completes the request and emits messages.update.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When the phone does not answer within MEDIA_REUPLOAD_TIMEOUT_MS,
Evolution now answers Baileys' own wait itself: a messages.media-update
carrying an error for that message, on the socket that asked (marked as
Evolution's event in a live recording). Baileys' waiter removes its
listeners and throws, so nothing is left per unanswered request, and an
answer arriving later finds no waiter: the message is not rewritten and
no messages.update follows. A second address (the refusal retry) is not
tried once the request is abandoned.

The test's socket now throws an answer that carries an error, as
Baileys' updateMediaMessage does (it only recorded the answer before);
checked red against the code before this fix with that change.

Visible to consumers: no late messages.update for a media re-upload the
API already answered no_answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e out of the API for good

The monitor's loaders call setInstance without awaiting it, and
setInstance registers the instance only after its connect. A connect
that failed on a read (loadSettings here) rejected unhandled, and the
instance was never registered and never retried.

Red on the current code: the instance is not in waInstances, no socket
is ever built, and the rejection is unhandled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ies it

- setInstance registers the instance in waInstances before its connect,
  and a failed auto-connect (a database read, the network) is logged with
  bounded fields and handed to the service's reconnect backoff
  (retryConnect: 1s, 2s, 4s... up to 60s), instead of rejecting.
- The loaders await each setInstance through loadOne, which logs one
  instance's failure without stopping the others; no rejection is left
  unhandled. (loadInstance runs after the server is listening, so only
  resumeDeletedLogouts now waits for the connects.)

Visible to consumers: an instance whose connect failed at boot answers
in the API (state close, then connecting) and connects when the database
or network returns, instead of 404 until a restart.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ritten, and a failed creds.update save is lost

- saveKey swallowed any error, so saveCreds resolved with the creds only
  in memory.
- A creds row that cannot be parsed read as no session: the store
  started fresh and wrote the fresh creds over it.
- eventHandler's saveCreds on creds.update was neither awaited nor
  retried, and its rejection went unhandled.

Red on the current code (3 tests).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… creds are never replaced

- saveKey lets a database error propagate, so saveCreds rejects instead
  of resolving with the creds only in memory.
- Stored creds that cannot be parsed fail the open (UnreadableCredsError)
  instead of reading as no session, which started fresh creds and wrote
  them over the stored ones.
- A creds.update's save (both the normal and the pending-logout paths) is
  handled: a failure is logged with bounded fields and the save is tried
  again, 1s, 2s, 4s... up to 30s, with the creds as they are by then,
  until one lands. No rejection is left unhandled.

Not done: key files are still written in place (fs.writeFile), not
atomically; no failing test was written for a torn write.

Visible to consumers: an instance whose stored creds are corrupt stays
disconnected (and keeps retrying, logging why) instead of silently
showing a new QR, and so losing the linked device.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… with the socket already built

LiveRecorder.start creates its directory and writes the first manifest
outside its guard, and createClient calls it after makeWASocket and
before eventHandler. Red on the current code: with LIVE_RECORD_DIR under
a file, the connect fails (ENOTDIR) and nothing listens to the socket it
built.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cording, never the connection

- LiveRecorder.start catches a failure to create its directory or write
  the first manifest, warns with the error code only, and returns no
  recorder: createClient goes on to install its handlers.
- attach writes the manifest inside the recorder's guard, so a manifest
  that can no longer be written stops the recording (as any failed write
  already did) instead of throwing out of createClient.

The second test (manifest read-only at attach) was written with this fix
and checked red against the code before it.

Visible to consumers: none unless LIVE_RECORD_DIR is set.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The tape records a batch line for every delivery. A non-bufferable event
emitted while the buffer holds others is delivered at once, as its own
batch, and the buffer keeps holding the rest; the replay flushed the
buffer on every batch line.

Red on the current code: a contacts.upsert and a contacts.update the
live buffer delivered as one batch are replayed as two.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A batch line with a bufferable event is the buffer's flush; one of
non-bufferable events only (a connection.update emitted inside a buffer)
was delivered at once with the buffer still held, so the replay now only
waits for its handling there. The bufferable list is Baileys' own
(BUFFERABLE_EVENT, rc14).

The test's two contacts are now two different people: one contact's
contacts.update is merged into its contacts.upsert by the buffer, which
no live batch line would show as two keys. Checked red against the
previous helper with that change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
FORK.md said more than the tests show (Codex review, item 18):
- the red commits: several fail on the fork's harness and Baileys pin,
  not on 2.3.7 itself;
- saved: the @lid mapping items of contacts.upsert have no saved and a
  null pushName;
- logs: bounded fields and scrubbed exception messages, not a guarantee
  for every line (inherited raw error logs, untested integrations, names
  in unquoted words);
- proxy, sockets, logout, pictures, creds: what this branch's fixes now
  hold, and the limits (transports tested);
- live checks: what the scrubber, gate and guard each check, that a human
  read is still required, and that a replay compares content, not order,
  pictures, decryption or timing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
POST /chat/getBase64FromMediaMessage answers a failed download with the
error's text, Baileys' "Failed to fetch stream from <link>", so the HTTP
answer carries WhatsApp's signed CDN link (oh, oe, the _nc_ parameters and
the directPath). The same object is what the error handler posts to the
errors webhook and what the S3 upload logs at ERROR.

The new file drives the real route against WhatsApp's CDN host, answered by
the undici mock agent: 404 with the phone refusing the re-upload, 410 with
no answer in time, 403 on an expired link, a re-upload whose new link fails
too, and a 403 on a valid link with and without reupload: false. It asserts
the whole answer, its headers and the output meanwhile, the thrown object,
and the S3 upload's error log. All eight fail on the code as it is.

media-skip-reupload.test.ts asserted the old text ("TypeError: fetch failed")
for a fallback that fails without an HTTP status; it now expects the stable
reason text (error kind and network code), and fails until the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
getBase64FromMediaMessage answered a failed download with the error's own
text, which for Baileys' download error is "Failed to fetch stream from"
and the signed CDN link. Its 400 now carries a stable reason instead: the
CDN's HTTP status ("The media could not be downloaded (HTTP 404)"), or,
with no status, the kind of error and its network code. reupload and
reuploadReason are unchanged, and a text Evolution threw itself (it names
no link) stands. The object is built fresh, so no Boom data (data.url)
travels with it: not to the HTTP answer, not to the errors webhook main.ts
posts, and not to the S3 upload's error log.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…concurrent one

The Prisma auth store writes each signal key over its file in place
(fs.writeFile truncates, then writes), and reads a file that does not parse
as no key at all. Four cases, on the real store and a real filesystem:

- Crash: a child process (the real store, bundled with esbuild) writes one
  key over and over with 4 MB values that name their write, and is
  SIGKILLed at random moments, 64 times across 4 folders. After each kill
  the file must hold one whole value. On the code as it is, one run: 32 of
  64 whole, 26 torn, 6 empty.
- A write that fails partway: the child runs under a file size limit
  (ulimit -f), so the second write fails with EFBIG after 1 MiB. The error
  reaches keys.set, and the previous value must survive; it is torn.
- Concurrency: twelve writes of the same key at once must end with the last
  one, whole; they end torn.
- A value written is read back by the store and by a new one after a
  restart (passes today, kept as the guard for the fix).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Prisma auth store now writes a key through writeFileAtomic
(src/utils/atomic-file.ts): the value goes to a temp file in the same
directory (opened exclusively), is flushed with fsync, and is renamed over
the key file, which POSIX makes atomic; the directory is then flushed, so
the rename survives a power loss (renames that land together share one
directory flush). A failed write removes its temp file, leaves the previous
value and rejects, so keys.set fails as saveCreds does. Writes of the same
file run one at a time in call order, so the last one asked for stays.

A killed writer can leave its temp file behind; the store removes those
when it opens its folder (a temp file of a process that no longer runs, or
this process's own that it is not writing), and leaves one of another
running process alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…marker

The marker next to the session's key files is written in place
(fs.writeFile). A write that fails partway (injected here as a full disk:
half the bytes land, then ENOSPC) rejects as it should, but leaves half a
marker over the whole one, which reads as pending with no instance name.
The previous marker must survive, with no temp file left.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
writeLogoutMarker writes through writeFileAtomic, as the key files do: a
temp file in the same directory, fsync, rename over the marker, fsync of the
directory. A write that fails keeps the previous marker, removes its temp
file and rejects, so the logout or delete still answers 500 as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…laced whole

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
NOTICE said the policy was included in this repository; 2.3.7 does not
carry TRADEMARKS.md, so it now links the upstream file. It also states
that the fork does not change the user interface or its brand assets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ket, a pairing-code link)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant