Barcelona hosted a new meeting of the European project LUCIA, organized by the National Center for Genomic Analysis (CNAG), in which the consortium of which BILBOMÁTICA is a part continued to advance in its lines of work and coordination.
European data spaces
In European health research projects, one of the biggest challenges is not just obtaining data, but ensuring that it can be integrated, analyzed, and used securely, interoperably, and effectively for the scientific community. This challenge becomes even more critical when the objective is to study risk factors related to lung cancer.
In this context, data spaces play a key role. It's not just about gathering information, but about doing so within a framework that guarantees quality, traceability, governance, interoperability, and reuse. This is also the logic driving European data spaces: creating environments where data can be shared and used responsibly to generate knowledge and scientific value.
At BILBOMÁTICA we participate in key tasks related to the LUCIA Health Data Platform (eCRF) and the Virtual Research Environment (VRE). These components function as a common infrastructure for storing data, models and tools, and for facilitating their integration within a secure and operational research environment.
Within the VRE, we have incorporated selection, filtering, and comparative analysis functionalities designed to support research on risk factors associated with lung cancer. The interface allows researchers to work with prospective and retrospective data, apply clinical and sociodemographic criteria, select patients, run inference models, and clearly visualize aggregated results.
The developed functionalities include the comparison of different risk prediction algorithms, such as LCDRAT, BACH, LLPi, PLCOm2012, and LCRAT. The system displays results per patient, risk thresholds, phase distribution, and Venn diagrams that allow for the identification of similarities and differences between models.
This activity directly connects with the increasingly relevant dataspace approach in Europe. Data selection for the VRE is not merely a technical function, but rather part of a broader logic encompassing governance, quality, traceability, interoperability, and reuse. In this sense, LUCIA aligns with the evolution of European dataspaces, where value lies not only in storing information, but also in making it locatable, accessible, interoperable, reusable, and usable under clear security and governance conditions.
In this sense, LUCIA's value lies not in "sending data," but in preparing the technical and methodological ecosystem so that data can be used securely, traceably, and interoperably. The platform facilitates the integration and use of data within the project, while any external sharing would depend on the data owners, applicable permissions, and relevant regulatory requirements.
For this model to work, interoperability, description, and data governance standards are essential. In the EHDS, the European Electronic Health Record Exchange Format (EEHRxF) will be particularly relevant for exchanging electronic health records, as will HealthDCAT-AP for describing health datasets in European catalogs, and standards such as HL7 FHIR, SNOMED CT, LOINC, OMOP CDM, and specific European profiles for structuring, coding, and making data understandable across systems and countries.
In parallel, access for secondary use is not based on a free transfer of data, but on requests, permissions, contracts, authorized purposes, access control, secure processing environments, and traceability obligations. This vision is supported by European infrastructures such as MyHealth@EU, geared towards the primary use of data in cross-border healthcare, and HealthData@EU, focused on secondary use for research and innovation.
Standardization and regulations
This development also aligns with the European framework of rules and standards for trustworthy artificial intelligence, particularly regarding data and source management. Within the context of the AI Act, the harmonized standards promoted by CEN-CENELEC are presented as a way to demonstrate compliance with European regulatory requirements. Applied to VRE, this reinforces the importance of working with clear criteria for data quality, traceability, transparency, representativeness, bias control, governance, security, and risk management.
From this perspective, data selection ceases to be a merely technical matter and becomes an essential element for ensuring that analyses and inference models are based on appropriate, documented, and reusable sources within a controlled scientific environment.
Furthermore, this development is framed within a logic of quality, validation, and operational deployment. The integration of models and functionalities within the VRE allows results to move beyond an isolated experimental phase and be used in a common, secure, and accessible environment for project partners.
Environmental Context
The future validation of risk factor-based models using prospective datasets further reinforces the need for robust data selection criteria, traceability of sources, and evaluation mechanisms that allow for responsible and reproducible interpretation of results.
BILBOMÁTICA's experience in ICT platforms, interoperability, systems integration, and data management is particularly valuable for bringing to the project a vision connected to European data spaces. This vision ensures that data is not only stored but also organized and valued so that it can be securely shared, understood, and reused.
Another component integrated into this ecosystem is the GIS/GeoHealth tool, developed by the Andalusian Health Service (SAS). Its objective is to enrich patient records with socioeconomic and environmental information based on geolocation. In this way, the analysis of risk factors can incorporate not only clinical information but also variables related to the territory, the environment, and environmental exposure.
In this context, Exploratory Spatial Data Analysis (ESDA) provides an additional layer of geospatial analysis by allowing the exploration of territorial patterns, clusters, and potential outliers in environmental, sociodemographic, and clinical data. It incorporates techniques such as Anselin Local Moran’s I (LISA) and Getis-Ord hotspot analysis to study lung cancer rates, the spatial distribution of errors in machine learning models, and territorial clustering. These analyses allow for the extraction of knowledge from the spatial component of the data and its subsequent use as input for geospatial AI models, strengthening the project's capacity to identify risk factors linked to the environment and territory.
This approach reflects one of the great strengths of European R&D projects: combining clinical, technological, analytical, and territorial capabilities to build shared, reusable infrastructures aimed at generating useful knowledge. In LUCIA, the integration of health data, digital tools, data spaces, and geospatial analysis represents an important step toward more contextualized, interoperable, and evidence-based research.
P.S.: Scientific publication: “Understanding lung cancer risk factors and assessing their impact (LUCIA): protocol for a multicenter observational cohort study”
https://bmjopen.bmj.com/content/16/5/e116423
Understanding LUng Cancer risk factors and their Impact Assessment (LUCIA): protocol for multicentre observational cohort study. “… LUCIA is a multicentre, observational, longitudinal cohort study that will recruit approximately 4000 participants across four European regions: Andalusia and the Basque Country (Spain), Liège (Belgium) and Riga (Latvia)…”
El proyecto LUCIA recibe financiación del programa de investigación e innovación Horizon Europe de la Unión Europea en virtud del acuerdo de subvención 101096473.