apache / apache/geaflow

The wcc algorithm output contains a large amount of duplicate data

Open
#761 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Java
Stars
808
Forks
188
Avg merge
3d 22h
Merged PRs (30d)
2

Description

I found that the output file from the WCC algorithm contains duplicate data, and I suspect that intermediate results of the algorithm were also exported.

the wcc sql is:

CREATE GRAPH cc_graph_test (
Vertex nodes (
id bigint ID
),
Edge edges (
srcId bigint SOURCE ID,
targetId bigint DESTINATION ID
)
) WITH (
storeType='memory',
shardCount = 1
);

INSERT INTO cc_graph_test.nodes(id) VALUES
(1),
(2),
(3),
(4),
(5),
(6);

INSERT INTO cc_graph_test.edges VALUES
(1, 2),
(2, 3),
(4, 5),
(5, 6)
;

CREATE TABLE IF NOT EXISTS cc_geaflow_test (
v_id int,
k_value VARCHAR
) WITH (
type='file',
`geaflow.dsl.table.parallelism`= 64,
`geaflow.dsl.source.parallelism` = 64,
`geaflow.file.persistent.config.json` = '{\'*******'}',
`geaflow.dsl.file.path` = '*******',
`geaflow.dsl.column.separator`='\s'
);

USE GRAPH cc_graph_test;
insert into cc_geaflow_test(v_id, k_value)
CALL wcc() YIELD (vid, component)
RETURN vid, component;

output is :
1s1
1s1
2s1
1s1
2s1
3s1
1s1
2s1
3s1
4s4
1s1
2s1
3s1
4s4
5s4
1s1
2s1
3s1
4s4
5s4
6s4

Contributor guide

Open the contributing guide

Research direction

Start by running the supplied CREATE GRAPH, WCC query, and file-output reproduction, then trace how WCC results are emitted and written to cc_geaflow_test. Compare the output with the six input vertices and verify that each vertex/component appears only once when the bug is fixed.

Written by the indexing model from the issue text.

Assessment

Tech stack
java, sql
Domain
data, distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.