Data Quality Aspects of Administrative Data

As researchers and practitioners working with administrative data, we are often given datasets where we do not know the full provenance about how this data set was captured, what kind of processing has been applied to it, and if it has been linked or merged with data from other sources. Complete and up-to-date metadata are not always available.

Not fully understanding the provenance of a data set can lead to assumptions and misconceptions being made about the content and quality of a dataset. This can result in incorrect processing and / or analysis of a dataset which potentially can lead to bad outcomes and decision making.

Course structure 

This in-person course will be a mixture of four hours of interactive presentations (containing small practical exercises) plus two one-hour sessions with group discussions. Additionally, participants will be sent links to watch two introductory/primer videos (approximately 1.5 hours total) prior to attending the workshop.

Course fees

Free - there are no fees for attending this course.

Course dates and location

14 April 2026 - Singleton Park Campus, Swansea University

28 May 2026 - Advanced Research Centre (ARC), Glasgow University

06 October 2026 - UK Research and Innovation (UKRI), London

Course details

The next course will take place:

6 October 2026 - UK Research and Innovation (UKRI), London.

Book your place: Datacise Open Learning

Course audience

This course is aimed both at researchers and practitioners who are working with administrative data, as well as those who are involved in the management of data-centric systems in organisations that act as data custodians, or who are involved in the capture, processing, and linkage of data that potentially will be used for administrative data research. The course requires little technical knowledge and all technical background will be introduced during the course.

Course content

The course will cover data quality dimensions which include technical and social aspects; discuss frameworks that aim to quantify data quality; and provide examples and case studies showing how (the lack of) quality data can lead to bad outcomes for research projects.

The course will not focus on technical aspects of data cleaning, data processing, or data linkage, but rather highlight the issues researchers and practitioners need to be aware of when working with administrative data. It will provide and discuss a set of recommendations, and through interactive sessions participants will be able to share their own experiences of how data quality aspects have led to unexpected outcomes in projects they have worked in.

Lecture topics will cover:

  • The data science workflow
  • Introduction to data wrangling and data analytics/mining
  • Overview of data quality aspects
  • Examples and case studies of what can go wrong in data science.
  • Data quality dimensions, data quality assessments, data quality frameworks
  • How data capturing, data processing, and data linkage can affect data quality
  • Assumptions and misconceptions in data science and how to identify and prevent them
  • Recommendations on how to improve data quality aspects
  • What data scientists should check with their data sources
     

Group discussions will cover:

  • Experiences of data quality issues encountered
  • How to implement recommendations within your own organisation

Taught by

Professor Peter Christen is the Research Lead on the Scottish Historic Population Platform (SHiPP) project, run by the Scottish Centre for Administrative Data Research (SCADR) at the University of Edinburgh. He is also a Professor at the School of Computing at the Australian National University in Canberra.

Peter is a world-leading expert in record linkage with over 20 years’ experience in working with administrative data. He has over 200 publications in the area of data science, including the two books “Data Matching” in 2012 and “Linking Sensitive Data” (co-authored with Thilina Ranbaduge and Rainer Schnell) in 2020. 

As of June 2026, his work has attracted nearly 19,800 citations at Google Scholar.

Course Review

Dr Ana Morales-Gomez, participant in the Glasgow event in May 2026 found this course:

..to be a valuable opportunity to reflect on many of the challenges that arise in practice, particularly around data provenance and linkage, and the assumptions we sometimes make when working with administrative data. The course reinforced the importance of understanding how data are generated and processed before they reach researchers and how these factors can influence research findings. It also prompted us to think more carefully about how these considerations shape the way we interpret and communicate research. Being transparent about the strengths and limitations of administrative data is essential to ensuring that research based on these data is robust, credible, and ultimately delivers meaningful evidence for the public good.