Data Dynamics – The Fuel of GenAI Models

Is GenAI objective or shaped by data? Students experiment with datasets and discover how data drives outcomes, bias and limitations.

The collage shows multiple repeated tabs with the words "people who liked this also liked...". In the middle, there is a painting of a woman feeding a goat. The goat has been highlighted in a yellow bounding box. The background is a painted nature scene.
Dominika Čupková & Archival Images of AI + AIxDESIGN / https://betterimagesofai.org / https://creativecommons.org/licenses/by/4.0/
  • Group activity
  • Dataset experimentation
  • In class + individual preparation
  • Exploring how data influences GenAI model behavior 
  • All disciplines (especially Social Sciences/Humanities)
  • Intermediate
  • 60-90 min / 1-2 lessons (excluding preparation)
  • approx. 25 students (groups of 3-5)
  • ML tool (e.g. Teachable Machine) or dataset worksheets

Short description

This activity demystifies the “black box” of machine learning by focusing on its primary input: data. Students explore the concept of data through individual research and group analysis. By experimenting with different data samples extracted from a larger dataset, students witness firsthand how data selection influences model behavior, leading to discussions on bias, representation, and explainability.

Competence domain of the Didactic Framework: Foundational AI knowledge

By the end of this activity, students can… 

  • explain “model training” in the machine learning context and describe training procedures. (FLAIR Didactic Framework: LO2) 
  • explain what type of data is used in model training and its role in AI. (FLAIR Didactic Framework: LO3) 
  • define the algorithmic bias and explain how it relates to AI limitations. (FLAIR Didactic Framework: LO5) 
  • examine how different data inputs alter results. 
  • assess the potential for bias in specific datasets. 

Instructions

Before class, students complete a short preparatory task by writing brief definitions of key concepts such as “data”, “types of data”, “personal data”, and “uses of data”. This establishes a baseline understanding and prepares them for in-class discussion. 

In class, students share and compare their definitions. Provide a brief lecture on how machine learning works, focusing on the relationship between training data and model outputs. Emphasise that models do not “understand” the world, but reproduce patterns present – or absent – in the data they are trained on. 

Students form small groups of 3–5 members. Introduce a large dataset and assign different subsets of the data to each group (e.g. data from different regions or demographic groups). Groups use their assigned dataset as input for the machine learning tool/model. The datasets are intentionally varied so that groups can later compare how training data influences model behaviour. 

Example: All groups train a model to classify images with the labels male and female. One group uses mostly images of younger adults, another uses images of older adults, and a third uses a limited and very similar image set.  

After training, groups test their models using the same images across the different models and observe where predictions differ or fail. They document accuracy, errors, and unexpected patterns. 

Students compare results across groups to see how “models” behaved differently based on the individual input. Each group briefly presents their findings to the class. The class collaboratively links these discrepancies to concepts of algorithmic bias and explainability. Students further reflect on how their understanding of data has changed compared to the beginning. 

Assessment 

Formative assessment focuses on the group presentation explaining how the assigned data samples influenced the results and model behavior. Summative assessment may be based on an initial essay on “What is data?” combined with a short post-activity reflection, evaluating how students’ understanding of data has expanded to include its role as training material for GenAI systems. 

Possible challenges

  • Students might struggle to visualize “data training” without a technical background. 
  • Students may find it difficult to interpret differences in results across groups  

How to adress them

  • Use visual analogies (e.g., teaching a child to recognize cats using only pictures of black cats). 
  • Guide students with targeted questions to support interpretation 

Using this resource

This resource is licensed under Creative Commons BY-SA 4.0 license. Suggested citation: Flair Collaboration. (2025). FLAIR Toolkit. Teaching GenAI Competencies.

Creative Commons Licence: Attribution-NonCommercial-ShareAlike 4.0 International