Improving Parallel System Performance with a NUMA-aware Load Balancer

Multi-core nodes with Non-Uniform Memory Access (NUMA) are now a common architecture for high performance computing. On such NUMA nodes, the shared memory is physically distributed into memory banks connected by a network. Owing to this, memory access costs may vary depending on the distance between the processing unit and the memory bank. Therefore, a key element in improving the performance on these machines is dealing with memory affinity. We propose a NUMA-aware load balancer that combines the information about the NUMA topology with the statistics captured by the Charm++ runtime system. We present speedups of up to 1.8 for synthetic benchmarks running on different NUMA platforms. We also show improvements over existing load balancing strategies both in benchmark performance and in the time for load balancing. In addition, by avoiding unnecessary migrations, our algorithm incurs up to seven times smaller overheads in migration, than the other strategies.

Data and Resources

Additional Info

Field Value
Source https://inria.hal.science/hal-00788813
Author Pilla, Laércio L., Pousa Ribeiro, Christiane, Cordeiro, Daniel, Bhatele, Abhinav, Navaux, Philippe O. A., Mehaut, Jean-François, Kalé, Laxmikant V.
Maintainer CCSD
Last Updated May 14, 2026, 11:25 (UTC)
Created May 14, 2026, 11:25 (UTC)
Identifier hal-00788813
Language en
contributor Instituto de Informática da UFRGS (UFRGS) ; Universidade Federal do Rio Grande do Sul [Porto Alegre – Brasil] = Federal University of Rio Grande do Sul [Porto Alegre – Brazil] (UFRGS)
creator Pilla, Laércio L.
date 2011-05-14T00:00:00
harvest_object_id aa9a03a1-f07e-4776-a29e-3a13a27bbb63
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2025-11-06T00:00:00
set_spec type:REPORT