Distributed-memory simulations of turbulent flows on modern GPU systems using an adaptive pencil decomposition library

Romero, Joshua; Costa, P.; Fatica, Massimiliano

Distributed-memory simulations of turbulent flows on modern GPU systems using an adaptive pencil decomposition library

Conference paper (2022)

Authors

Joshua Romero Nvidia Corporation

P. Costa University of Iceland

Massimiliano Fatica Nvidia Corporation

Affiliation

External organisation

Computational fluid dynamics Direct numerical simulation GPU accelerated systems Parallel transpose

To reference this document use:

http://resolver.tudelft.nl/uuid:f433d95a-6d16-4964-ab96-7d492373b2c1

More Info

expand_more

Published Date

2022

Language

English

Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Affiliation

External organisation

Abstract

This paper presents a performance analysis of pencil domain decomposition methodologies for three-dimensional Computational Fluid Dynamics (CFD) codes for turbulence simulations, on several large GPU-accelerated clusters. The performance was assessed for the numerical solution of the Navier-Stokes equations in two codes which require the calculation of Fast-Fourier Transforms (FFT): a tri-periodic pseudo-spectral solver for isotropic turbulence, and a finite-difference solver for canonical turbulent flows, where the FFTs are used in its Poisson solver. Both codes use a newly developed transpose library that automatically determines the optimal domain decomposition and communication backend on each system. We compared the performance across systems with very different node topologies and available network bandwidth, to show how these characteristics impact decomposition selection for best performance. Additionally, we assessed the performance of several communication libraries available on these systems, such as Open-MPI, IBM Spectrum MPI, Cray MPI, the NVIDIA Collective Communication Library (NCCL), and NVSHMEM. Our results show that the optimal combination of communication backend and domain decomposition is highly system-dependent, and that the adaptive decomposition library is key in ensuring efficient resource usage with minimal user effort.

No files available

Metadata only record. There are no files for this conference paper.