Quantum-espresso¶
Support tier: 1
Read information about support tiers.
Installed versions¶
| Resource | Version |
|---|---|
| Arrhenius | 7.6-gcc-2026.03-mpich-eb |
| Dardel-GH/cpe25.03 | 7.5.0 |
| Dardel/cpe26.03 | 7.5.0 |
| Dardel/cpe26.03 | 7.5.0 |
Read information about how to load this software in your environment by searching for Lmod module.
General information¶
Quantum ESPRESSO is an integrated suite of open-source computer codes for electronic-structure calculations and materials modeling at the nanoscale. It is based on density-functional theory, plane waves, and pseudopotentials. For more information, see https://www.quantum-espresso.org.
Running Quantum ESPRESSO on Arrhenius CPU nodes¶
Quantum ESPRESSO is available through the module system. To check which versions are available before loading a module:
To load the 7.6 version of the program:
For supercomputers with a Lustre-based shared file system, such as the one mounted on Arrhenius, keep the main input and output files on the shared file system and store temporary files and wavefunction output on the node-local scratch directory whenever possible. This can be done by adding wfcdir = '/scratch/local/' to the QE input file. Moreover, to avoid unnecessary read/write of large output files, it is advisable to use QE options such as disk_io = 'low' and wf_collect = .false.. For some runs, these settings can substantially reduce excessive file I/O.
Example job scripts¶
Replace naissYYYY-X-XXX-cpu in the examples below with your project account.
Example 1: single node, pure MPI¶
This example uses 256 MPI ranks with one core per rank, filling all physical cores on one CPU node.
#!/bin/bash
#SBATCH --job-name=qe-mpi
#SBATCH --account=naissYYYY-X-XXX-cpu
#SBATCH --partition=cpu
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=256
#SBATCH --cpus-per-task=1
#SBATCH --time=02:00:00
ml QuantumESPRESSO/7.6-gcc-2026.03-mpich-eb
# One thread per MPI rank, including numerical libraries
export OMP_NUM_THREADS=1
export OPENBLAS_NUM_THREADS=1
mpprun pw.x -in scf.in > scf.out
For smaller calculations, reduce --ntasks-per-node, for example to 32 or 64. Using more MPI ranks does not necessarily improve performance. Benchmark a few values to find a suitable setting for your system.
Example 2: single node, hybrid MPI + OpenMP¶
This example uses 64 MPI ranks with four OpenMP threads per rank, using 256 physical cores in total.
#!/bin/bash
#SBATCH --job-name=qe-hybrid
#SBATCH --account=naissYYYY-X-XXX-cpu
#SBATCH --partition=cpu
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=64
#SBATCH --cpus-per-task=4
#SBATCH --time=02:00:00
ml QuantumESPRESSO/7.6-gcc-2026.03-mpich-eb
# OpenMP threads per MPI rank
export OMP_NUM_THREADS="$SLURM_CPUS_PER_TASK"
# Keep BLAS thread counts low to avoid oversubscription
# OpenMP-based numerical routines follow OMP_NUM_THREADS
export OPENBLAS_NUM_THREADS=1
mpprun pw.x -in scf.in > scf.out
Possible full-node configurations to benchmark include:
| MPI ranks per node | OpenMP threads per rank | Cores per node |
|---|---|---|
| 256 | 1 | 256 |
| 128 | 2 | 256 |
| 64 | 4 | 256 |
| 32 | 8 | 256 |
Change --ntasks-per-node and --cpus-per-task together. These settings keep the total number of physical cores used per node constant while changing the MPI/OpenMP balance.
Example 3: two nodes, hybrid MPI + OpenMP¶
This example uses 64 MPI ranks per node and four threads per rank:
- 128 MPI ranks in total
- 256 physical cores per node
- 512 physical cores across two nodes
#!/bin/bash
#SBATCH --job-name=qe-2nodes
#SBATCH --account=naissYYYY-X-XXX-cpu
#SBATCH --partition=cpu
#SBATCH --nodes=2
#SBATCH --ntasks-per-node=64
#SBATCH --cpus-per-task=4
#SBATCH --time=04:00:00
ml QuantumESPRESSO/7.6-gcc-2026.03-mpich-eb
export OMP_NUM_THREADS="$SLURM_CPUS_PER_TASK"
export OPENBLAS_NUM_THREADS=1
mpprun pw.x -in scf.in > scf.out
Benchmark against the single-node case before using this configuration routinely. Communication between nodes can outweigh the benefit of additional cores for smaller systems.
K-point parallelism¶
The examples above leave QE's parallelization parameters at their automatically selected values.
For calculations with sufficient k-points, explicitly distributing them among pools may improve performance. For example, replace the execution line in the two-node script with:
This divides the 128 MPI ranks into eight pools of 16 ranks. Each MPI rank still uses four OpenMP threads.
Choose a pool count that divides the MPI rank count. For balanced work, the number of k-points actually calculated after symmetry reduction should ideally also be divisible by the pool count. For a Gamma-point-only calculation, use one pool (-nk 1) or leave the setting automatic.
Check the output file, here scf.out, for the expected MPI ranks, OpenMP threads, and parallelization layout.
For more advanced parallelization options, see the Quantum ESPRESSO user guide: https://www.quantum-espresso.org/Doc/user_guide/node20.html.
Notes¶
For a first job on Arrhenius, a practical starting configuration is the single-node hybrid case:
- 64 MPI ranks per node
- 4 OpenMP threads per rank
- 1 node
This is a useful baseline for benchmarking before moving to larger jobs or multi-node runs.
If memory usage is too high, reduce the number of MPI ranks or use fewer OpenMP threads per rank. Alternatively, you may use the fat partition.
If performance is poor, benchmark a few combinations of MPI ranks and OpenMP threads instead of assuming that the largest job is always best.
The exact optimal MPI/OpenMP balance depends on the size of the system, the number of k-points, and the memory footprint of the calculation. Start with a conservative configuration, test performance, and scale up only if the workload benefits from it.
How to use Quantum ESPRESSO on Dardel¶
You should always use the option disk_io='low'. With this setting, wave functions are written only at the end of the job instead of after every intermediate step. This substantially reduces the load on the disk system and can make your job run faster.
To load a Quantum ESPRESSO module:
Here is an example job script requesting 128 MPI processes on one node:
#!/bin/bash
#SBATCH -J qejob
#SBATCH -A naissYYYY-X-XX
#SBATCH -p main
#SBATCH -t 01:00:00
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=128
ml PDC/<version>
ml quantum-espresso/7.6.0-cpeGNU-26.03
export OMP_NUM_THREADS=1
srun pw.x -in myjob.in > myjob.out
Since OpenMP is supported by this module, you can also submit a job requesting 64 MPI processes per node and 2 OpenMP threads per MPI process. In this case, you need to specify --cpus-per-task, OMP_NUM_THREADS, and OMP_PLACES.
#!/bin/bash
#SBATCH -J qejob
#SBATCH -A naissYYYY-X-XX
#SBATCH -p main
#SBATCH -t 01:00:00
#SBATCH --nodes=2
#SBATCH --ntasks-per-node=64
#SBATCH --cpus-per-task=2
ml PDC/<version>
ml quantum-espresso/7.6.0-cpeGNU-26.03
export OMP_NUM_THREADS=2
export OMP_PLACES=cores
export SRUN_CPUS_PER_TASK=$SLURM_CPUS_PER_TASK
srun --hint=nomultithread pw.x -in myjob.in > myjob.out
Quantum ESPRESSO is also available as builds for the NVIDIA Grace Hopper nodes. The following example job script runs on two nodes with one task per GPU.
#!/bin/bash
#SBATCH -A pdc.staff
#SBATCH -J qe
#SBATCH -t 02:00:00
#SBATCH -p gpugh
#SBATCH -N 2
#SBATCH -n 8
#SBATCH -c 72
#SBATCH --gpus-per-task 1
# Runtime modules and executable paths
ml PDC/25.03
ml quantum-espresso/7.5.0
# Runtime environment
export MPICH_GPU_SUPPORT_ENABLED=1
export PMPI_GPU_AWARE=1
export OMP_NUM_THREADS=72
export OMP_PLACES=cores
export SRUN_CPUS_PER_TASK=$SLURM_CPUS_PER_TASK
echo "Script initiated at `date` on `hostname`"
srun --hint=nomultithread pw.x -in myjob.in > myjob.out
echo "Script finished at `date` on `hostname`"
How to build Quantum ESPRESSO on Dardel¶
The builds for the AMD CPU nodes were installed using EasyBuild. A build in your local file space can be done with:
See also Installing software using EasyBuild.
A build for the NVIDIA Grace Hopper nodes can be done with:
# Build instructions for Quantum ESPRESSO on Dardel Grace Hopper nodes
# Download and unpack the source code
wget https://gitlab.com/QEF/q-e/-/archive/qe-7.5/q-e-qe-7.5.tar.gz
tar xvf q-e-qe-7.5.tar.gz
cd q-e-qe-7.5
# Load the environment and GNU toolchain
ml PrgEnv-nvidia
ml cudatoolkit/24.11_12.6
ml craype-accel-nvidia90
ml cray-fftw/3.3.10.10
ml cray-hdf5/1.14.3.5
ml cmake/4.1.2
# Configure
mkdir buildNvidiaCuda
cd buildNvidiaCuda
cmake .. -DQE_ENABLE_LIBXC=0 -DQE_ENABLE_OPENMP=1 -DQE_ENABLE_SCALAPACK=1 \
-DQE_ENABLE_WANNIER90=0 -DQE_ENABLE_ELPA=0 -DQE_ENABLE_HDF5=1 \
-DBLAS_LIBRARIES="-L${CRAY_LIBSCI_PREFIX_DIR}/lib -lsci_nvidia_mp" \
-DLAPACK_LIBRARIES="-L${CRAY_LIBSCI_PREFIX_DIR}/lib -lsci_nvidia_mp" \
-DSCALAPACK_LIBRARIES="-L${CRAY_LIBSCI_PREFIX_DIR}/lib -lsci_nvidia_mp" \
-DFFTW3_INCLUDE_DIRS="${FFTW_INC}" \
-DQE_ENABLE_CUDA=1 \
-DQE_ENABLE_MPI_GPU_AWARE=1 \
-DQE_ENABLE_OPENACC=1 \
-DCMAKE_INSTALL_PREFIX=/pdc/software/25.03/other/quantum-espresso/7.5.0 \
> BuildQuantumEspresso_CMakeLog.txt 2>&1
# Build and install
make -j 72 all > BuildQuantumEspresso_make.txt 2>&1
make install
Source repository