Book page

Open Data, Open Source and Reproducibility of Statistical Algorithms – Introductory Course

Default profile image
Magda CHMIEL @NTTS • 3 March 2026

Course Leader

Ole Mussmann

Target Group

Official statisticians who will start or having an interest to work on Open data - open source and have no specific knowledge on this subject.

Entry Qualifications

Sound command of English. The course targets official statisticians who are planning to begin, or are interested in, working with open data and open-source solutions and who do not yet have specific expertise in this area but familiarity with statistical production processes is recommended.

Objective(s)

Objective is to teach the participants how to use open source and open data by: 

  • Design and implement reproducible workflows for statistical production using collaborative tools 

  • Use and configure open computational environments (e.g. Renkulab, MyBinder,

  • Access and integrate open data sources through recognised standards and vocabularies (SDMX, EU Vocabularies, INSPIRE), 

  • Apply open-source principles in statistical projects, understanding the implications of different licensing models (CC, EUPL, GPL, MIT, etc.).

Contents

The course should focus on the following topics: 

  1. Reproducibility and collaborative environments 

    Reproducible workflows in official statistics. 

  • Hands-on: using GitHub/GitLab for versioning and collaboration. 

  • Jupyter, R Markdown, and computational notebooks for documentation and sharing. 

  • Running environments for open research: renkulab.io, mybinder.org, Docker basics.

  • Demonstration: launching and reproducing a statistical analysis pipeline.

  1. Open Data: access, standards, and interoperability 

    • Principles of open data in official statistics. 

    • Metadata and interoperability standards: SDMX, EU VOC, INSPIRE. 

    • Linking open datasets and APIs; reproducible data acquisition

  2. Open Source and Licensing 

    • What “open source” means for statistical production. 

    • Licensing schemes and their implications: CC-BY, EUPL, GPL, MIT, etc.

    • Best practices for publishing reproducible code and data. 

Expected Outcome

Participants will be able to apply reproducible workflows using collaborative tools, set up open research environments, reproduce statistical analyses, understand open data principles and interoperability standards, and follow best practices for publishing code and data under open-source licenses.

Training Methods
  • Presentations and lectures

  • Exercises


 

Required Reading 

None

Suggested Reading

None

Required Preparation

Basic knowledge of Git and Python are necessary. If not present, the participant should follow the basic courses on git and python on codecademy

Trainer(s)/
Lecturer(s)

Marco Puts (CBS Netherlands)

Ole Mussman (eScience Center) 

Chris Lam (CBS Netherlands)

Darius Keijdener (CBS Netherlands)

 

 

Practical Information

Start date

End Date

Duration

Where

Address

APPLICATION VIA National Contact Point

24 June 2026

26 June 2026

3 days

The Hague, Netherlands

Heerlen (CBS-weg 11 6412 EX Heerlen)

Netherlands

Deadline for application: 08/05/2026