Distributed LLM Fine-Tuning & Inference on HPC systems, Fall 2026
NRIS Training is organizing a third round of Distributed LLM Fine-Tuning & Inference on HPC systems. This is a two-day, in-person, hands-on course in Bergen. Gain practical, hands-on experience working with single-GPU fine-tuning, multi-GPU scaling on single- and multi-node setups, and optimized LLM inference on a high-performance computing (HPC) system. Attend this course to build applied skills in optimizing large language models in HPC environments.
When: November 18.-19., 2026
Where: Bergen, University Campus
Instructor: Hicham Agueny
HPC System: Olivia
Course program and schedule
Day 1 — Single-GPU Fine-Tuning & HPC Foundations
Theme: Build an efficient single-GPU fine-tuning workflow on an HPC system.
Morning Session (09:30–12:00) — HPC Fundamentals & Fine-Tuning Optimization
HPC Foundations for LLM Workloads
Overview of Olivia Supercomputer
Containerized environments including EESSI
LLM Fine-Tuning Fundamentals
Parameter-efficient fine-tuning with LoRA
Quantized fine-tuning with QLoRA
Afternoon Session (13:00–15:30) — Hands-On: Single-GPU workflow for QA and XSum Tasks
LoRA fine-tuning workflow
Quantized fine-tuning with QLoRA: FP4 vs BF16 comparison
Evaluation of the fine-tuned model
GPU monitoring and memory profiling
Wrap-Up & Discussion (15:30–16:00)
HPC Foundations for LLM Workloads
Overview of Olivia Supercomputer
Containerized environments including EESSI
LLM Fine-Tuning Fundamentals
Parameter-efficient fine-tuning with LoRA
Quantized fine-tuning with QLoRA
LoRA fine-tuning workflow
Quantized fine-tuning with QLoRA: FP4 vs BF16 comparison
Evaluation of the fine-tuned model
GPU monitoring and memory profiling
Wrap-Up & Discussion (15:30–16:00)
Outcome: Participants implement and optimize a complete single-GPU fine-tuning pipeline with performance diagnostics on an HPC system.
Day 2 — Distributed Training & Optimized Inference
Theme: Scale fine-tuning and inference across multiple GPUs while minimizing communication overhead.
Morning Session (09:30–12:00) — Distributed Fine-Tuning
Distributed Training Concepts
Concept of parallelism
DDP vs FSDP
Communication and scaling efficiency
Hands-On: Multi-GPU Fine-Tuning on a single node & acorss nodes for QA and XSum Tasks
Multi-GPU & multi-node LoRA & QLoRA fine-tuning
Evaluation of the fine-tuned model accros multi-GPUs
Profiling distributed workloads
Afternoon Session (13:00–15:30) — Hands-On: Optimized Inference
Introduction to the vLLM inference engine
Single-GPU inference benchmarking
Quantization: torchao, bitsandbytes, GPTQModel
Multi-GPU inference
Wrap-Up & Discussion (15:30–16:00)
Distributed Training Concepts
Concept of parallelism
DDP vs FSDP
Communication and scaling efficiency
Hands-On: Multi-GPU Fine-Tuning on a single node & acorss nodes for QA and XSum Tasks
Multi-GPU & multi-node LoRA & QLoRA fine-tuning
Evaluation of the fine-tuned model accros multi-GPUs
Profiling distributed workloads
Introduction to the vLLM inference engine
Single-GPU inference benchmarking
Quantization: torchao, bitsandbytes, GPTQModel
Multi-GPU inference
Wrap-Up & Discussion (15:30–16:00)
Outcome: Participants scale fine-tuned models and inference across multiple GPUs, interpret performance metrics, and apply optimization strategies suitable for HPC allocations.
Target audience & prerequisites
The course is ideal for researchers, developers, and students with Python experience who want hands-on skills in scalable LLM training and inference on an HPC system.
Registration: Register here
Practical Information
The course is free of charge, but will have a maximum capacity of 25 people. Lunch will be included, and coffee/tea will be served.
Contact us
If there are questions regarding the course or NRIS Training, please contact us at training@nris.no.