← Selected research

ML interatomic potentials / HPC

GPU resource trade-offs in MACE molecular dynamics

Measuring how GPU allocation changes the runtime, memory footprint, and compute cost of the same polymer simulation workload.

Engineering pilotSeptember 2026
Two versus four GPUs: runtime falls from 117.3 to 62.7 minutes, maximum sampled per-GPU memory falls from 21.68 to 11.13 GiB, and allocated GPU time rises from 3.91 to 4.18 hours. The three-GPU attempt failed at initialization.
Measured resource use for the same MACE workload. One short engineering run per configuration, on different hosts; scientific density QC was not passed. The three-GPU initialization failure is included in the record.

Research question

How should runtime, device memory, and total resource cost be evaluated together when distributing an ML interatomic potential across GPUs?

Approach

  • A 3,666-atom polymer system with MACE-MH-1, the omol head, and float32 precision; a fixed workload of 5,500 MD steps (1.375 ps).
  • Compared completed two- and four-GPU configurations using the same system, model, precision, and timestep. Wall time excludes queue wait and includes the recorded run stages.
  • Measured sampled per-device memory with NVML and calculated allocated GPU-hours. Retained the three-GPU attempt, which failed during MPI/CUDA initialization before model work.

What the measurements show

  • Observed wall time decreased from 117.3 to 62.7 minutes with four GPUs, a 1.87-fold speedup in this pilot.
  • Maximum sampled memory per device decreased from 21.68 to 11.13 GiB. These are NVML samples, not exact allocator peaks.
  • Allocated GPU time increased from 3.91 to 4.18 GPU-hours, approximately 6.9% more total GPU time.

The useful result is a measured resource trade-off: lower elapsed time does not necessarily mean lower total compute cost.

Scope & limitations

  • Each configuration was run once on different hosts. The observations do not isolate all hardware effects or establish broad scaling performance.
  • The trajectories were short engineering tests and failed scientific density quality checks. They are not qualified density measurements or evidence of long-time equilibration.
  • This is an evaluation of an existing MLIP workflow, not a claim to have introduced a new model-optimization algorithm.

Continue exploring

Preliminary polymer density screening with MACE

Read the case study