[Keyword: SPH] AND [All Categories: Computer Architecture] : Search

In this topic

Advanced Search

SEARCH GUIDE

Results: 1 - 1of1

Follow results:

refine search

Filters

per page:

Sort: Relevance

Context for search term 1Search term 1*

All Dates

LastSelect static range

Custom Range

Select starting monthSelect starting year

Select ending monthSelect ending year

Advanced

Search name	Searched On	Run search
[Keyword: SPH] AND [All Categories: Computer Architecture] (1)	30 Mar 2025	Run
[Keyword: SHG] AND [All Categories: Life Sciences / Biology] (2)	30 Mar 2025	Run
[Keyword: GIS] AND [All Categories: Chemistry] (5)	30 Mar 2025	Run
[Keyword: BIM] AND [All Categories: Computational, Mathematical and Theoretical Phy... (1)	30 Mar 2025	Run
[Keyword: SHG] AND [All Categories: Engineering] (3)	30 Mar 2025	Run

articleNo Access
Efficient parallelization of SPH algorithm on modern multi-core CPUs and massively parallel GPUs
- Pravin Jagtap,
- Rupesh Nasre,
- V. S. Sanapala, and
- B. S. V. Patnaik
International Journal of Modeling, Simulation, and Scientific Computing10 Jul 2021
Preview Abstract
Smoothed Particle Hydrodynamics (SPH) is fast emerging as a practically useful computational simulation tool for a wide variety of engineering problems. SPH is also gaining popularity as the back bone for fast and realistic animations in graphics and video games. The Lagrangian and mesh-free nature of the method facilitates fast and accurate simulation of material deformation, interface capture, etc. Typically, particle-based methods would necessitate particle search and locate algorithms to be implemented efficiently, as continuous creation of neighbor particle lists is a computationally expensive step. Hence, it is advantageous to implement SPH, on modern multi-core platforms with the help of High-Performance Computing (HPC) tools. In this work, the computational performance of an SPH algorithm is assessed on multi-core Central Processing Unit (CPU) as well as massively parallel General Purpose Graphical Processing Units (GP-GPU). Parallelizing SPH faces several challenges such as, scalability of the neighbor search process, force calculations, minimizing thread divergence, achieving coalesced memory access patterns, balancing workload, ensuring optimum use of computational resources, etc. While addressing some of these challenges, detailed analysis of performance metrics such as speedup, global load efficiency, global store efficiency, warp execution efficiency, occupancy, etc. is evaluated. The OpenMP and Compute Unified Device Architecture $(C U D A)$ parallel programming models have been used for parallel computing on Intel Xeon $(R)$ E5- $2630$ multi-core CPU and NVIDIA Quadro M $4000$ and NVIDIA Tesla p $100$ massively parallel GPU architectures. Standard benchmark problems from the Computational Fluid Dynamics (CFD) literature are chosen for the validation. The key concern of how to identify a suitable architecture for mesh-less methods which essentially require heavy workload of neighbor search and evaluation of local force fields from neighbor interactions is addressed.