Platform Architecture
The EpiScientia Data Portal uses a federated architecture enabling secure, privacy-preserving analytics across National Public Health Institutes without centralizing sensitive patient data.
Beneficiary Data Node — Requirements
Server Hardware Requirements
- Processor: Quad Core (AMD or ARM)
- Memory: 32Gb (64Gb recommended)
- Storage: SSD 1Tb or more
- OS: Ubuntu 24.04
Maintenance and General Requirements
- Servers should have an uptime of 99% (follow local KPI guidelines if it's more stringent than our recommended guidelines)
- It is recommended to set up an automated nightly job for regular, nightly backups
- The System Administrator, together with the Network Administrator and/or IT Infra team as needed, should keep the following up to date:
- Security patches
- Software and Docker image versions
- Deployment components
- Server components
cHDP Architecture Overview
EpiScientia NPHIs Data Node: Centralized Health Data Platform
EpiScientia Network Node Server

Components of the network node server: FastAPI (node registration), Prometheus (metrics collection), Grafana (metrics visualization), Keycloak (node authentication) and Superset (data exploration & visualization).
Data Node

Each node runs Ubuntu OS with Docker containers for the ETL, OHDSI tools (Atlas, RStudio, WebAPI, Achilles, ARES, Glue, Traefik) and Vantage6 for federated learning, all around the OMOP CDM database.
Global Infrastructure

End-to-end federated flow: EMR data sources (OpenClinic GA, OpenMRS, DHIS2) feed local data nodes via ETL; NHIC network node servers exchange node health data, registration keys and aggregate results with the project server.
Key Components
Nos de Dados Locais
Health facilities maintain their own data nodes with Hospital EHRs, OpenMRS, and other EMR systems. Each facility uses ETL pipelines to process and standardize data.
Contenorizacao Docker
R, Python, Jupyter, and Vantage environments are containerized with Docker, ensuring consistent execution environments with built-in security firewalls.
Hub Central de Consultas
The Central Data Center receives queries and processes them without storing personal metadata, ensuring privacy-preserving federated analytics.
Analitica por IA
OHDSI/ATLAS integration with R, Python, Machine Learning, and Jupyter notebooks enables advanced analytical capabilities for real-time evidence generation.
How Data Flows
Data Collection
Health facilities collect data through Hospital EHRs, OpenMRS, OpenClinic GA, and other EMR systems.
ETL Processing
Extract-Transform-Load pipelines convert data to OMOP CDM format within secure Docker containers.
Federated Queries
Central hub sends queries to local nodes; only aggregated results return - no patient data leaves the facility.
AI Analytics
Results are processed through OHDSI/ATLAS and ML pipelines for impact analysis and modelling.
Platform Capabilities
- Real-time evidence generation for pandemic preparedness and response
- Impact analysis and epidemiological modelling
- Federated queries across multiple health facilities
- Privacy-preserving data analysis without centralized patient data
- Web portal interface for researchers and analysts
Technology Stack
Docker
Containerization
R
Analytics
Python
Analytics
Jupyter
Notebooks
OHDSI/ATLAS
Standards
Vantage
Data Platform
OpenMRS
EMR
OMOP CDM
Data Model
Machine Learning
AI
SQL
Database
CSV/ETL
Data Pipeline
Firewall
Security
Explore More
Learn more about our data governance, access policies, and available datasets.


