Skip to content

fix: size the connection pool for reuse, not for opening - #14

Merged
bludot merged 1 commit into
mainfrom
fix/connection-pool-reuse
Aug 23, 2026
Merged

bludot merged 1 commit into
mainfrom
fix/connection-pool-reuse

Conversation

@bludot

@bludot bludot commented Aug 23, 2026

Copy link
Copy Markdown
Contributor

The pool was sized to open connections, not to reuse them.

Go retains only MaxIdleConns. Anything opened above that is created for one query and closed the moment it finishes — so a pool of 100 open / 10 idle was never reusing 100 connections. It held 10 and rebuilt up to 90, paying a TCP connect, TLS handshake and Postgres auth per query, against RDS over the internet with sslmode=require.

                  before          after
maxOpenConns      see below         n
maxIdleConns          10            n     ← matched, so the pool is actually a pool
ConnMaxLifetime    5m / 1h         30m
ConnMaxIdleTime    90s / unset     10m

ConnMaxIdleTime at 90s was working against reuse too: on light traffic, connections were reaped between requests and rebuilt on the next one, so the pool never warmed up.

Why the numbers are small

The database is small. db.t4g.micro allows 79 connections in total, and roughly 36 pods share them. The configured ceilings summed to about 935.

  • APIs get 5 — they serve concurrent requests and benefit from parallelism
  • Consumers get 2 — they process one Kafka message at a time, so 25 could never be used, but still counted against the budget

A warm pool of 5 serves more traffic than a churning pool of 25, while claiming a fraction of the budget.

The symptom this explains

Intermittent slow GraphQL responses with no errors. At 48/79 connections the database was not exhausted — requests were waiting on handshakes. auth and user-service were the acute case: each pod was configured for 100, so a single pod could have exhausted the database on its own.

🤖 Generated with Claude Code

Go retains only MaxIdleConns; anything opened above it is closed as soon as the
query finishes. The previous settings kept far fewer idle than they were willing
to open, so most connections were built per query -- TCP, TLS and Postgres auth
each time, against RDS over the internet.

MaxIdleConns now matches MaxOpenConns, which is what makes the pool a pool.

The sizes are small because the database is: db.t4g.micro allows 79 connections
in total and roughly 36 pods share them, against configured ceilings summing to
about 935. APIs take 5, consumers 2 -- a consumer processes one message at a
time, so 25 could never be used but still counted against the budget.

ConnMaxIdleTime moves off 90s, which was reaping connections between requests on
light traffic so the pool never warmed up. Lifetime stays bounded so a failover
or DNS change is still picked up without a restart.
@bludot
bludot merged commit f74b04f into main Aug 23, 2026
2 checks passed
@bludot

bludot commented Aug 23, 2026

Copy link
Copy Markdown
Contributor Author

🎉 This PR is included in version 1.15.1 🎉

The release is available on:

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant