Dataset Personalization Methods based on LLMs for Data Science Education: A Comparative Study of Rescaling and Sampling Approaches

Abstract

This work-in-progress study explores the use of Large Language Models (LLMs) to dynamically personalize datasets used in data science education according to learner interests. The study outlines two dataset personalization methods that transform each variable in the original dataset to a new variable: a scaling method that transforms while preserving the original distributions and correlation structure, and a sampling method that transforms while preserving the original correlation structure but changes distributions to match the new variables. An evaluation with subject matter experts revealed that datasets personalized using the scaling method were not significantly different from the original datasets in terms of the appropriateness of variable names and ranges. Further evaluation using the personalized datasets in the context of instructional materials designed for the original datasets indicated that more inconsistencies were found with the sampling method than the scaling method. These results suggest that dataset personalization can create datasets that serve as drop-in replacements for the original datasets in existing instructional materials.

Publication Title

L@s 2025 Proceedings of the 12th ACM Conference on Learning @ Scale

Share

COinS