AI and Data Science:

Introduction to eXplainable AI for Transformers

What happens between input and prediction? Explore how Explainable AI can help make the decisions of transformer models more transparent—and where explanations reach their limits.

This course is offered by Helmholtz AI in cooperation with HIDA

Helmholtz AI

Helmholtz AI is an application-driven artificial intelligence platform accelerating science across the Helmholtz Association. They enable the development and implementation of AI solutions while promoting collaboration and ensuring access to resources and expertise.

This half-day course provides an introduction to Explainable AI (XAI) for transformer models in natural language processing (NLP). The course focuses on text classification and uses DistilBERT as a running example to introduce explainability methods ranging from attention-based visualizations to architecture-aware attribution techniques.

While the practical sessions focus on text-based transformers, many of the explainability principles discussed in the course are relevant to other transformer architectures and applications.

The course combines alternating theoretical sessions and hands-on exercises. Practical notebooks allow participants to apply different explainability methods to transformer models, while discussions with the instructors help interpret the results and critically assess the strengths and limitations of each approach.

Learning goals

The course aims to provide participants with a practical understanding of explainability methods for transformer-based models.

After a brief introduction to the transformer architecture, participants will explore both model-agnostic and transformer-specific explainability techniques. Starting from attention-based visualizations (raw attention, attention rollout, and attention flow), the course discusses why attention alone is not sufficient as an explanation. It then introduces additional attribution methods, including gradient-based approaches, SHAP, and architecture-aware techniques such as Layer-wise Relevance Propagation (LRP).

Throughout the course, participants will learn how to critically evaluate explanations, understand the assumptions underlying different methods, and recognize their strengths and limitations in practice.

By the end of the course, participants will be able to:

  • understand the main principles of explainability for transformer models;
  • distinguish between model-agnostic and transformer-specific XAI methods;
  • apply standard attribution methods to transformer models;
  • interpret and critically evaluate the explanations produced by different methods.

Course date

Register now: 

Nov 27, 2026

Note: Registration will open Oct 27, 2026, 12 pm.

Prerequisites

  • Basic knowledge of Python
  • Basic understanding of machine learning and deep learning concepts.
  • Familiarity with the fundamentals of transformer models (e.g., encoder architecture, self-attention, token embeddings, and positional encodings) is highly recommended. The course begins with a brief recap of these concepts but does not provide a comprehensive introduction to transformer architectures.
  • Prior familiarity with explainable AI concepts is recommended.Participants are encouraged to have attended the course Introduction to eXplainable AI, or to have equivalent knowledge of model-agnostic explainability methods.

Target group

This course mainly targets Helmholtz PhD students, postdoctoral researchers, and practitioners interested in understanding and interpreting transformer-based models. The course is particularly relevant for researchers applying natural language processing or transformer models in their own work.

This course is free of charge.

 

The Data Science Course Portfolio

This course is part of the Data Science Course Portfolio, curated by the five Helmholtz Information & Data Science Platforms - Helmholtz AI, Helmholtz Imaging, HIDA, HIFIS, HMC. Find out more on the Course Portfolio here

Alternativ-Text

Subscribe newsletter