apache / apache/cloudberry

[Bug] Expand fails to recognize user-defined tablespaces

Open Beginner friendly
#1,885 0 comments 2 reactions 0 assignees View on GitHub
type: Bug
Dominant language
C
Stars
1.4k
Forks
247
Avg merge
4d 3h
Merged PRs (30d)
39

Description

### Apache Cloudberry version

2.1.0-incubating

### What happened

When user-defined tablespaces exist in the cluster, gpexpand does not create the `newTableSpaceInfo.json` file in the new segment template when running with a config file (non-interactive mode).

This leads to broken symlinks for tablespaces and causes failures during the prepare_schema() stage.

The problem is present in two methods within the same file:
1. `read_tablespace_file()` in `gpMgmt/bin/gpexpand` (lines 1219–1285)
2. `generate_tablespace_inputfile()` in `gpMgmt/bin/gpexpand` (lines 1114–1148)

- `read_tablespace_file()` (line 1244):
```python
tblspc_oids = os.listdir(coordinator_tblspc_dir)
tblspc_oid_names = self.get_tablespace_oid_names()
flag = False
for oid in tblspc_oids:
if oid in tblspc_oid_names: <---
flag = True
if not flag:
return None
```

- `generate_tablespace_inputfile()` (line 1124)
```python
tblspc_oid_names = self.get_tablespace_oid_names()
tblspc_info = {}

for oid in tblspc_oids:
if oid not in tblspc_oid_names: <---
continue
location = os.path.dirname(os.readlink(os.path.join(coordinator_tblspc_dir,
oid)))
tblspc_info[oid] = {"location": location,
"name": tblspc_oid_names[int(oid)]}
```

Root cause: type mismatch between string values returned by `os.listdir()` and integer keys returned by the SQL query `SELECT oid, spcname FROM pg_tablespace`.

When attaching to the process and debugging, the types differ, causing this issue:

- type mismatch

Image

- `newTableSpaceInfo `is `None`

Image

Impact:
- `newTableSpaceInfo `is always None.
- `_handle_tablespace_template()` is never invoked.
- `newTableSpaceInfo.json` is not included in the template tar.
- `gpconfigurenewsegment `on new hosts does not fix tablespace symlinks.
- new segments start with broken symlinks.

### What you think should happen instead

To fix the issue, unify the data types in both methods:
- In `generate_tablespace_inputfile()`:
```python
for oid in tblspc_oids:
if int(oid) not in tblspc_oid_names:
```

- In `read_tablespace_file()`:
```python
tblspc_oids = os.listdir(coordinator_tblspc_dir)
tblspc_oid_names = self.get_tablespace_oid_names()
flag = False
for oid in tblspc_oids:
if int(oid) in tblspc_oid_names:
```

After this change, `newTableSpaceInfo `is no longer `None`:

Image

### How to reproduce

1. Connect to the database:
```sh
gpadmin@cbdb-mdw:~$ psql warehouse
```

2. Create a user-defined tablespace:
```sql
warehouse=# CREATE TABLESPACE ts_stage_logs LOCATION '/tblspc_stage_logs';
```

```
warehouse=# SELECT * FROM pg_tablespace;
oid | spcname | spcowner | spcacl | spcoptions | spcfilehandlersrc | spcfilehandlerbin
-------+---------------+----------+--------+------------+-------------------+-------------------
1663 | pg_default | 10 | | | |
1664 | pg_global | 10 | | | |
17019 | ts_stage_logs | 10 | | | |
(3 rows)
```

3. Create an append-optimized columnar table and populate it with test data:
```sql
warehouse=# CREATE TABLE logs_aot (
id BIGSERIAL,
log_timestamp TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
log_level VARCHAR(10)
)
WITH (
APPENDONLY = TRUE,
ORIENTATION = COLUMN
)
TABLESPACE ts_stage_logs
DISTRIBUTED BY (id);

warehouse=# INSERT INTO logs_aot (
log_timestamp,
log_level
)
SELECT
NOW() - (random() * INTERVAL '90 days'),
(ARRAY['INFO', 'WARN', 'ERROR', 'DEBUG'])[floor(random() * 4 + 1)]
FROM generate_series(1, 8000);
```

4. Prepare the expansion configuration file for adding 2 new hosts to the cluster:
```
sdw3-new|sdw3-new|6000|/primary_data/gpseg4|10|4|p
sdw4-new|sdw4-new|7000|/mirror_data/gpseg4|11|4|m
sdw3-new|sdw3-new|6001|/primary_data/gpseg5|12|5|p
sdw4-new|sdw4-new|7001|/mirror_data/gpseg5|13|5|m
sdw4-new|sdw4-new|6000|/primary_data/gpseg6|14|6|p
sdw3-new|sdw3-new|7000|/mirror_data/gpseg6|15|6|m
sdw4-new|sdw4-new|6001|/primary_data/gpseg7|16|7|p
sdw3-new|sdw3-new|7001|/mirror_data/gpseg7|17|7|m
```

5. Start the expansion process (non-interactive mode):
```sh
gpadmin@cbdb-mdw:~$ gpexpand -i expand.cfg
```

6. The following error appears in the logs (log truncated for readability.)
```
20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Heap checksum setting consistent across cluster
20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Syncing Apache Cloudberry extensions
20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Locking catalog
20260805:18:33:05:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Locked catalog
20260805:18:33:06:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Creating segment template
20260805:18:33:07:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Copying postgresql.conf from existing segment into template
20260805:18:33:08:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Copying pg_hba.conf from existing segment into template
20260805:18:33:09:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Creating schema tar file
...
20260805:18:33:26:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Populating gpexpand.status_detail with data from database postgres
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Populating gpexpand.status_detail with data from database warehouse
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[ERROR]:-gpexpand failed: ERROR: could not open file "pg_tblspc/17019/GPDB_3_302606111/17018/16386": No such file or directory (seg6 203.0.113.5:6000 pid=21449)

Exiting...
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[ERROR]:-gpexpand is past the point of rollback. Any remaining issues must be addressed outside of gpexpand.
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Shutting down gpexpand...
```

Hope this analysis is helpful.

### Operating System

Ubuntu 22.04

### Anything else

_No response_

### Are you willing to submit PR?

- [ ] Yes, I am willing to submit a PR!

### Code of Conduct

- [x] I agree to follow this project's [Code of Conduct](https://github.com/apache/cloudberry/blob/main/CODE_OF_CONDUCT.md).

Contributor guide

Open the contributing guide

Research direction

Start in gpMgmt/bin/gpexpand by reading read_tablespace_file() and generate_tablespace_inputfile(), focusing on how os.listdir() values are compared with get_tablespace_oid_names() results. Run the provided user-defined tablespace and non-interactive gpexpand reproduction, then verify that newTableSpaceInfo.json is generated and tablespace symlinks are fixed on new segments.

Written by the indexing model from the issue text.

Assessment

Tech stack
postgresql, python
Domain
databases
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
76/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.