apache / apache/cloudberry

[Bug] Walreciever crashes with SIGABRT for user with scram-sha-256 password

Open
#1,978 1 comment 0 reactions 0 assignees View on GitHub
type: Bug
Dominant language
C
Stars
1.4k
Forks
247
Avg merge
4d 3h
Merged PRs (30d)
39

Description

### Apache Cloudberry version

The main branch 15d75c672ad360a11f1b7204218a35c016479f0c. Older versions should be affected too because the problem exists at greenplum

### What happened

Walreciever crashes with SIGABRT when the connection with primary closes (for ex reconnect) if replication user uses a scram-sha-256 password.

There is a stack trace:
```
(gdb) bt
#0 __pthread_kill_implementation (no_tid=0, signo=6, threadid=) at ./nptl/pthread_kill.c:44
#1 __pthread_kill_internal (signo=6, threadid=) at ./nptl/pthread_kill.c:78
#2 __GI___pthread_kill (threadid=, signo=signo@entry=6) at ./nptl/pthread_kill.c:89
#3 0x000076dc2424527e in __GI_raise (sig=sig@entry=6) at ../sysdeps/posix/raise.c:26
#4 0x000076dc242288ff in __GI_abort () at ./stdlib/abort.c:79
#5 0x000076dc242297b6 in __libc_message_impl (fmt=fmt@entry=0x76dc243ce8d7 "%s\n") at ../sysdeps/posix/libc_fatal.c:134
#6 0x000076dc242a90d5 in malloc_printerr (str=str@entry=0x76dc243d1520 "munmap_chunk(): invalid pointer") at ./malloc/malloc.c:5775
#7 0x000076dc242a955c in munmap_chunk (p=) at ./malloc/malloc.c:3040
#8 0x000076dc242adefa in __GI___libc_free (mem=0x5d2ba9297988) at ./malloc/malloc.c:3388
#9 0x00005d2b6cc4608f in scram_free (opaq=0x5d2ba92d8230) at fe-auth-scram.c:190
#10 0x00005d2b6cc2ffe2 in pqDropConnection (conn=0x5d2ba92d7820, flushInput=false) at fe-connect.c:612
#11 0x00005d2b6cc44b62 in pqReadData (conn=0x5d2ba92d7820) at fe-misc.c:808
#12 0x00005d2b6cc3df28 in PQconsumeInput (conn=0x5d2ba92d7820) at fe-exec.c:2042
#13 0x00005d2b6ce6ae94 in libpqrcv_PQgetResult (streamConn=0x5d2ba92d7820) at libpqwalreceiver/libpqwalreceiver.c:802
#14 0x00005d2b6ce6b070 in libpqrcv_receive (conn=0x5d2ba9297848, buffer=0x7ffc2d8b1b58, wait_fd=0x7ffc2d8b1b28) at libpqwalreceiver/libpqwalreceiver.c:879
#15 0x00005d2b6ce5d572 in WalReceiverMain () at walreceiver.c:480
#16 0x00005d2b6cdeb4f8 in AuxiliaryProcessMain (auxtype=WalReceiverProcess) at auxprocess.c:161
#17 0x00005d2b6cdf8a19 in StartChildProcess (type=WalReceiverProcess) at postmaster.c:6065
#18 0x00005d2b6cdf903a in MaybeStartWalReceiver () at postmaster.c:6315
#19 0x00005d2b6cdf877d in process_pm_pmsignal () at postmaster.c:5857
#20 0x00005d2b6cdf3320 in ServerLoop () at postmaster.c:2098
#21 0x00005d2b6cdf2a3f in PostmasterMain (argc=7, argv=0x5d2ba92961b0) at postmaster.c:1750
#22 0x00005d2b6cc4984d in main (argc=7, argv=0x5d2ba92961b0) at main.c:260
```

This problem occurs due to linkage frontend code into backend, the [scram_free](https://github.com/apache/cloudberry/blob/15d75c672ad360a11f1b7204218a35c016479f0c/src/interfaces/libpq/fe-auth-scram.c#L186-L206) expects that all memory is allocated via malloc and uses free to release it. But the [fe-auth-scram.c](https://github.com/apache/cloudberry/blob/15d75c672ad360a11f1b7204218a35c016479f0c/src/interfaces/libpq/fe-auth-scram.c) is compiled into backend binary where saslprep.c compiled without FRONTEND macro - all [allocation macros uses postgres memory contexts](https://github.com/apache/cloudberry/blob/15d75c672ad360a11f1b7204218a35c016479f0c/src/common/saslprep.c#L38-L41), so all allocation in pg_saslprep include output password is performed with its usage [1](https://github.com/apache/cloudberry/blob/15d75c672ad360a11f1b7204218a35c016479f0c/src/common/saslprep.c#L1067) [2](https://github.com/apache/cloudberry/blob/15d75c672ad360a11f1b7204218a35c016479f0c/src/common/saslprep.c#L1210). Thus free fails to release palloc'd memory.

### What you think should happen instead

No SIGABRT

### How to reproduce

How to reproduce:
- create a cluster with mirrors
```sh
DATADIRS=$HOME/cloudberry-data PORT_BASE=7000 NUM_PRIMARY_MIRROR_PAIRS=1 WITH_MIRRORS=true make create-demo-cluster
```
- set up current user scram-sha-256 password
```sql
set password_encryption to 'scram-sha-256';
alter user gpadmin with password 'password';
```
- for dbfast1 change `pg_hba.conf`- add new record for replication database and `gpadmin` user with auth method to scram-sha-256 (all other disable)
```
host replication gpadmin 127.0.0.1/32 scram-sha-256
```

```sh
sed -i 's/.*replication.*/#&/g' ~/cloudberry-data/dbfast1/demoDataDir0/pg_hba.conf
echo 'host replication gpadmin 127.0.0.1/32 scram-sha-256' >> ~/cloudberry-data/dbfast1/demoDataDir0/pg_hba.conf
```

- create password record:
```sh
echo '*:*:*:gpadmin:password' > $HOME/.pgpass && chmod 0600 $HOME/.pgpass
```
- reload the config
```sh
gpstop -u
```
- reload the mirror - to enable scram connection
```sh
pg_ctl stop -D ~/cloudberry-data/dbfast_mirror1/demoDataDir0/
pg_ctl start -D ~/cloudberry-data/dbfast_mirror1/demoDataDir0/ -o '-c gp_role=execute -p 7003'
```
- stop the primary
```sh
pg_ctl stop -D ~/cloudberry-data/dbfast1/demoDataDir0
```
- the log of `dbfast_mirror1` would contain the info about the problem

### Operating System

Ubuntu 24.04.4

### Anything else

_No response_

### Are you willing to submit PR?

- [x] Yes, I am willing to submit a PR!

### Code of Conduct

- [x] I agree to follow this project's [Code of Conduct](https://github.com/apache/cloudberry/blob/main/CODE_OF_CONDUCT.md).

Contributor guide

Open the contributing guide

Research direction

Start with fe-auth-scram.c at scram_free and compare its allocation assumptions with common/saslprep.c, especially the allocation macros and pg_saslprep call sites. Reproduce the mirror failover scenario using the listed cluster, SCRAM configuration, and pg_ctl commands, then confirm the connection close no longer produces SIGABRT or an invalid-pointer error.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, postgresql
Domain
databases
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
62/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.