Multi-Node Implementation for PyTorch on Olivia
This is part 3 of the PyTorch on Olivia guide. See Single-GPU Implementation for PyTorch on Olivia for single-GPU and Multi-GPU Implementation for PyTorch on Olivia for multi-GPU setup.
Note
The ddp_train.py file does not need any changes when scaling from single node to multiple nodes. The only change required is in the job script.
Multi-node training on Olivia requires a consistent NCCL-enabled module environment and a stable rendezvous endpoint shared by all nodes. The job script below handles both.
Learning Outcomes
By the end of this part, you can:
Launch PyTorch training across multiple nodes with
torchrun.Configure the required module environment for distributed communication.
Set rendezvous parameters correctly for a stable multi-node start.
Job Script for Multi-Node Training
The repo which you cloned earlier already has the job script that uses the NRIS module. Below, you will find the equivalent job scripts for container and EESSI stack.
1#!/bin/bash
2#SBATCH --job-name=pytorch_multinode
3#SBATCH --account=<project_number>
4#SBATCH --output=logs/multinode_%j.out
5#SBATCH --error=logs/multinode_%j.err
6#SBATCH --time=00:30:00
7#SBATCH --partition=accel # GPU partition
8#SBATCH --nodes=2 # Request 2 compute nodes
9#SBATCH --ntasks-per-node=1 # One task (process) on the node
10#SBATCH --cpus-per-task=48 # Right-sized CPU allocation for multi-node ViT DDP
11#SBATCH --mem=192G # Right-sized RAM per node for multi-node ViT DDP
12#SBATCH --gpus-per-node=4 # Number of GPUs per node
13
14module load NRIS/GPU
15module load aws-ofi-nccl/1.19.1-GCCcore-14.3.0-CUDA-13.0.0
16module load libfabric/2.3.1-GCCcore-14.3.0-CUDA-13.0.0
17
18# Get the absolute path to the project directory.
19PROJECT_DIR=$(cd "${SLURM_SUBMIT_DIR}/.." && pwd)
20
21# Path to container and training script
22CONTAINER_PATH="/cluster/work/support/container/pytorch_nvidia_25.05_arm64.sif"
23
24TRAINING_SCRIPT="${PROJECT_DIR}/scripts/train_ddp.py --model vit --dataset tiny-imagenet --epochs 100 --batch-size 2048 --optimizer adamw --base-lr 0.00015 --target-accuracy 0.95 --patience 2 --seed 42 --num-workers 8 --amp"
25
26
27# Host library paths
28HOST_LIBFABRIC_LIB="${EBROOTLIBFABRIC}/lib"
29HOST_AWSOFI_LIB="${EBROOTAWSMINOFIMINNCCL}/lib"
30HOST_CXI_LIB_PATH="/usr/lib64"
31
32# NCCL debug
33#export APPTAINERENV_NCCL_DEBUG=INFO
34#export APPTAINERENV_NCCL_DEBUG_SUBSYS=INIT,NET
35
36
37# Get head node IP
38nodes=( $(scontrol show hostnames $SLURM_JOB_NODELIST) )
39head_node=${nodes[0]}
40export APPTAINERENV_head_node_ip=$(srun --nodes=1 --ntasks=1 -w "$head_node" hostname --ip-address | awk '{print $1}')
41
42echo "Head Node: $head_node"
43echo "Head Node IP: $APPTAINERENV_head_node_ip"
44
45# Pass SLURM variables explicitly
46export APPTAINERENV_SLURM_JOB_NUM_NODES=$SLURM_JOB_NUM_NODES
47export APPTAINERENV_SLURM_GPUS_ON_NODE=$SLURM_GPUS_ON_NODE
48
49# Start GPU utilization monitoring
50GPU_LOG_FILE="${PROJECT_DIR}/jobs/logs/multinode.log"
51echo "Starting GPU utilization monitoring..."
52nvidia-smi --query-gpu=timestamp,index,name,utilization.gpu,utilization.memory,memory.total,memory.used --format=csv -l 5 > $GPU_LOG_FILE &
53NVIDIA_MONITOR_PID=$!
54
55# Run training script with torchrun inside container
56srun apptainer exec --nv \
57 --bind $HOST_LIBFABRIC_LIB:/opt/libfabric/lib \
58 --bind $HOST_AWSOFI_LIB:/opt/aws-ofi-nccl/lib \
59 --bind $HOST_CXI_LIB_PATH:/usr/lib64 \
60 --env head_node_ip=$APPTAINERENV_head_node_ip \
61 --env TRAINING_SCRIPT="$TRAINING_SCRIPT" \
62 --env RDZV_ID=$SLURM_JOB_ID \
63 --env SLURM_JOB_NUM_NODES=$APPTAINERENV_SLURM_JOB_NUM_NODES \
64 --env SLURM_GPUS_ON_NODE=$APPTAINERENV_SLURM_GPUS_ON_NODE \
65 $CONTAINER_PATH \
66 bash -c 'export LD_LIBRARY_PATH=/opt/aws-ofi-nccl/lib:/opt/libfabric/lib:/usr/lib64:$LD_LIBRARY_PATH; \
67 torchrun \
68 --nnodes=$SLURM_JOB_NUM_NODES \
69 --nproc_per_node=$SLURM_GPUS_ON_NODE \
70 --rdzv_id=$RDZV_ID \
71 --rdzv_backend=c10d \
72 --rdzv_endpoint=$head_node_ip:29500 \
73 $TRAINING_SCRIPT'
74
75# Stop GPU utilization monitoring
76echo "Stopping GPU utilization monitoring..."
77kill $NVIDIA_MONITOR_PID 2>/dev/null || true
1#!/bin/bash
2#SBATCH --job-name=pytorch_multinode
3#SBATCH --account=<project_number>
4#SBATCH --output=logs/multinode_%j.out
5#SBATCH --error=logs/multinode_%j.err
6#SBATCH --time=00:30:00
7#SBATCH --partition=accel # GPU partition
8#SBATCH --nodes=2 # Request 2 compute nodes
9#SBATCH --ntasks-per-node=1 # One task (process) on the node
10#SBATCH --cpus-per-task=40 # Reserve 40 CPU cores (Right-sized for multi-node WideResNet)
11#SBATCH --mem=128G # Request 128 GB RAM (Right-sized for multi-node WideResNet)
12#SBATCH --gpus-per-node=4 # Number of GPUs per node
13
14# Activate the EESSI environment
15module load EESSI/2025.06
16module load torchvision/0.27.0-foss-2025b-PyTorch-2.12.0-CUDA-12.9.1
17
18
19# Get the absolute path to the project directory.
20PROJECT_DIR=$(cd "${SLURM_SUBMIT_DIR}/.." && pwd)
21
22# Path to the training script
23TRAINING_SCRIPT="${PROJECT_DIR}/scripts/train_ddp.py --model wideresnet --dataset cifar100 --epochs 100 --batch-size 2048 --base-lr 0.02 --target-accuracy 0.95 --patience 2 --seed 42 --amp"
24
25
26# NCCL Debug
27#export NCCL_DEBUG=INFO
28#export NCCL_DEBUG_SUBSYS=INIT,NET
29
30
31# Get head node IP
32nodes=( $(scontrol show hostnames $SLURM_JOB_NODELIST) )
33head_node=${nodes[0]}
34export head_node_ip=$(srun --nodes=1 --ntasks=1 -w "$head_node" hostname --ip-address | awk '{print $1}')
35
36echo "Head Node: $head_node"
37echo "Head Node IP: $head_node_ip"
38
39
40# Start GPU utilization monitoring
41GPU_LOG_FILE="${PROJECT_DIR}/jobs/logs/multinode.log"
42echo "Starting GPU utilization monitoring..."
43nvidia-smi --query-gpu=timestamp,index,name,utilization.gpu,utilization.memory,memory.total,memory.used --format=csv -l 5 > $GPU_LOG_FILE &
44NVIDIA_MONITOR_PID=$!
45
46# Run training script with torchrun inside container
47srun torchrun \
48 --nnodes=$SLURM_JOB_NUM_NODES \
49 --nproc_per_node=$SLURM_GPUS_ON_NODE \
50 --rdzv_id=$SLURM_JOB_ID \
51 --rdzv_backend=c10d \
52 --rdzv_endpoint=$head_node_ip:29500 \
53 $TRAINING_SCRIPT
54
55# Stop GPU utilization monitoring
56echo "Stopping GPU utilization monitoring..."
57kill $NVIDIA_MONITOR_PID 2>/dev/null || true
Then you can submit and monitor the running job using these commands:
sbatch multinode.sh
squeue -u $USER
tail -f multinode_<jobid>.out
Key Changes from Multi-GPU to Multi-Node
The multi-node-specific additions are:
Change |
Purpose |
|---|---|
|
Requests resources on multiple nodes |
Head-node hostname from |
Defines rendezvous endpoint for all processes |
|
Coordinates multi-node process-group formation |
Note
The key difference from single-node multi-GPU is the rendezvous setup. Single-node uses --standalone, while multi-node requires explicit coordination via --rdzv_backend=c10d and --rdzv_endpoint pointing to the head node.
The output of this job script is shown below:
Epoch 95/100: time=1.271s, train_loss=0.0146, train_acc=0.9994, val_loss=1.4960, val_acc=0.6831, throughput=38657.3 img/s
Epoch 96/100: time=1.271s, train_loss=0.0154, train_acc=0.9990, val_loss=1.5086, val_acc=0.6815, throughput=38674.0 img/s
Epoch 97/100: time=1.290s, train_loss=0.0164, train_acc=0.9987, val_loss=1.4687, val_acc=0.6835, throughput=38103.6 img/s
Epoch 98/100: time=1.282s, train_loss=0.0168, train_acc=0.9991, val_loss=1.4859, val_acc=0.6829, throughput=38335.4 img/s
Epoch 99/100: time=1.315s, train_loss=0.0143, train_acc=0.9994, val_loss=1.4213, val_acc=0.6907, throughput=37383.9 img/s
Epoch 100/100: time=1.270s, train_loss=0.0131, train_acc=0.9994, val_loss=1.3783, val_acc=0.6962, throughput=38694.6 img/s
Training Summary:
Total training time: 131.793 seconds
Throughput: 37294.946 images/second
Total GPUs used: 8
Training completed successfully.
With 8 GPUs across 2 nodes, the throughput increased from ~7367 images/second (single GPU) to ~37294 images/second—a 5x speedup. This sub-linear scaling achieves roughly 63% scaling efficiency across 8 GPUs, where multi-node communication overhead, specifically inter-node network latency during gradient synchronization across nodes is preventing linear scaling. Moreover, the training time dropped from ~667 seconds to just ~131 seconds.
Success criteria for Part 3:
Log shows
Head Nodeand a resolved head-node IPFinal summary reports
Number of nodes: 2andTotal GPUs used: 8Training completes without rendezvous or NCCL startup errors