Logo image
Parallelizing Maximal Clique Enumeration on GPUs
Conference proceeding

Parallelizing Maximal Clique Enumeration on GPUs

Mohammad Almasri, Yen-Hsiang Chang, Izzat El Hajj, Rakesh Nagi, Jinjun Xiong, Wen-mei Hwu and IEEE
Proceedings of the 32nd International Conference on Parallel Architectures and Compilation Techniques, pp.162-175
ACM Conferences
PACT '23: International Conference on Parallel Architectures and Compilation Techniques
21/10/2023

Abstract

Computing methodologies Computing methodologies -- Computer graphics Computing methodologies -- Computer graphics -- Graphics systems and interfaces Computing methodologies -- Computer graphics -- Graphics systems and interfaces -- Graphics processors Computing methodologies -- Parallel computing methodologies Computing methodologies -- Parallel computing methodologies -- Parallel algorithms Computing methodologies -- Parallel computing methodologies -- Parallel algorithms -- Massively parallel algorithms Computing methodologies -- Parallel computing methodologies -- Parallel programming languages Mathematics of computing Mathematics of computing -- Discrete mathematics Mathematics of computing -- Discrete mathematics -- Graph theory Theory of computation Theory of computation -- Design and analysis of algorithms Theory of computation -- Design and analysis of algorithms -- Graph algorithms analysis Theory of computation -- Design and analysis of algorithms -- Graph algorithms analysis -- Dynamic graph algorithms Theory of computation -- Design and analysis of algorithms -- Parallel algorithms Theory of computation -- Design and analysis of algorithms -- Parallel algorithms -- Massively parallel algorithms
We present a GPU solution for exact maximal clique enumeration (MCE) that performs a search tree traversal following the Bron-Kerbosch algorithm. Prior works on parallelizing MCE on GPUs perform a breadth-first traversal of the tree, which has limited scalability because of the explosion in the number of tree nodes at deep levels. We propose to parallelize MCE on GPUs by performing depth-first traversal of independent subtrees in parallel. Since MCE suffers from high load imbalance and memory capacity requirements, we propose a worker list for dynamic load balancing, as well as partial induced subgraphs and a compact representation of excluded vertex sets to regulate memory consumption. Our evaluation shows that our GPU implementation on a single GPU outperforms the state-of-the-art parallel CPU implementation by a geometric mean of 4.9 × (up to 16.7 ×), and scales efficiently to multiple GPUs. Our code has been open-sourced to enable further research on accelerating MCE.

Metrics

1 Record Views

Details

Logo image