Data Discovery and Data Use Framework - WP5
Work Package 5 (WP5) focuses on enabling effective, privacy‑preserving and scalable reuse of health and genomic data within the PROTECT‑CHILD ecosystem, providing the data discovery, analytical and advanced data use services that allow distributed data to be explored and analysed in practice.
Building on the secure infrastructure and operational services delivered by other work packages, WP5 equips the platform with the capabilities needed to turn data availability into meaningful and compliant data reuse.
From data availability to data use
WP5 addresses the challenge of working with sensitive health and genomic data in a distributed environment. It focuses on enabling data discovery, feasibility assessment and a wide range of privacy‑preserving data use scenarios, ensuring that valuable data can be reused without centralisation and fully in line with ethical and governance constraints.
Rather than focusing on infrastructure, WP5 delivers the services that directly support data understanding, analysis and advanced processing across the PROTECT‑CHILD platform.
What is being developed in WP5
Data discovery and feasibility services
WP5 develops data discovery services based on structured and standardised metadata, allowing users to identify relevant datasets and assess feasibility without accessing raw data.
These services support:
- Metadata‑based dataset discovery
- Privacy‑preserving feasibility assessment
- Exploration of data availability across distributed sites
This approach enables informed decision‑making while ensuring transparency, data protection and compliance with governance rules.
Privacy‑preserving analytical data use
To support analytical data reuse, WP5 delivers privacy‑preserving OLAP functionalities that enable statistical analysis, cohort selection and aggregate queries across distributed datasets.
These capabilities allow meaningful analytical processing while ensuring that sensitive data remain under the control of data holders and are never exposed or centralised.
Federated analytics and federated machine learning
WP5 also enables distributed analytics through the adoption of vantage6 as the framework for federated analytics and federated machine learning.
Through vantage6, analytical code and machine‑learning models are securely executed at local sites, enabling:
- Federated execution of analytics across multiple institutions
- Privacy‑preserving aggregation of results
- Support for federated machine‑learning workflows
- Alignment with distributed data ownership and access control
This ensures that advanced analytics can be performed collaboratively while fully respecting data protection requirements.
Structuring unstructured data and advanced secure computation
Beyond structured data, WP5 addresses the challenge of making unstructured clinical data reusable. It develops AI‑based and multilingual NLP services to extract and structure information from clinical text, improving data quality and usability.
WP5 also explores advanced and quantum‑secure multi‑party computation approaches, strengthening privacy protection for highly sensitive data such as genomics and preparing the platform for future security challenges.
Looking ahead
Over the project lifetime, WP5 will continue to refine and integrate its data discovery, analytical and advanced data use services, supporting pilot studies and real‑world validation. Its focus on privacy‑preserving analytics, federated learning, AI‑based data structuring and future‑proof security ensures that PROTECT‑CHILD can fully exploit distributed health and genomic data in a responsible and scalable way.
WP5 turns distributed data into actionable and trusted knowledge, enabling ethical, secure and impactful reuse of paediatric health data across Europe.
Author: Matteo Gabetta, Biomeris