All of Us
Overview
The All of Us dataset is a large, longitudinal U.S. health research resource that integrates clinical, behavioral, genomic, and wearable-device data from a diverse participant population.
Data Types
The dataset includes six major data types collected from participants:
Electronic Health Records (EHRs): Clinical records from participating healthcare organizations (OMOP Common Data Model). It includes diagnoses, medications, procedures, laboratory results, and clinical visits. (393k+ participants)
Participant Surveys: Self-reported data covering demographics, lifestyle, social determinants of health, family and personal health history, health care access and utilization, behavioral health and emotional health history (633k+ participants).
Physical Measurements: Standardized measurements that include height, weight, waist and hip circumference, blood pressure, and heart rate. Collectively from EHR data as well as self-reported through surveys (509k+ participants).
Biosamples: Blood and urine samples from 568,000+ participants, supporting genomic and other laboratory analyses.
Genomics: Whole genome sequences and genotyping arrays from 447,000+ participants, available in the Controlled Tier.
Wearable Devices (Digital Health): Biometric data from Fitbit devices including heart rate, physical activity, sleep patterns, device data. (59k+ participants)
Data Access Tiers
Data are organized into three access tiers based on sensitivity and researcher eligibility:
Public Tier: Aggregate, de-identified data available to everyone via the Data Browser and Data Snapshots. Doesn’t require registration.
Registered Tier: De-identified individual-level data including EHRs, wearables, surveys, and physical measurements. Accessible only to approved researchers on the Researcher Workbench.
Controlled Tier: Contains individual-level genomic data (whole genome sequences and genotyping arrays). Requires additional access approval beyond the Registered Tier.
Methodology and Generation
Data Ecosystem and Pipeline
Data flow through several stages before reaching researchers. Participants contribute data directly via participant portals and biosample donation. Healthcare Provider Organization (HPO) staff submit local EHR data through the HealthPro application. Genomics partners (genetic counseling resources, clinical validation laboratories, and genome centers) process biosamples for sequencing and genotyping. All raw data enters a central Raw Data Repository, then passes through a curation pipeline before being organized into the Curated Data Repository, from which the Public, Registered, and Controlled tiers are made available to researchers through the All of Us Research Hub.
Standardization with OMOP
Electronic health records are standardized using the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM), maintained by the international Observational Health Data Sciences and Informatics (OHDSI) collaborative. The OMOP CDM organizes data around a central Person table linked to standardized tables including Visit Occurrence, Procedure Occurrence, Drug Exposure, Condition Occurrence, Measurement, and Observation. Survey questionnaires and physical measurements are mapped into the Observation and Measurement tables respectively, creating a unified structure across all data sources.
Standardized Healthcare Vocabularies
The dataset uses the OMOP CDM and its associated standardized vocabularies.
SNOMED CT
LOINC
RxNorm
ICD-10-CM / ICD-9-CM
CPT-4 / HCPCS
NDC (National Drug Code)
CVX
PPI (Participant Provided Information): Custom All of Us vocabulary for survey responses not mappable to existing standards.
Data Use Compliance Requirements
Access to Registered and Controlled Tier data requires agreement to the Data User Code of Conduct (DUCC), which includes the following key requirements:
Participants may never be identified or re-identified.
Data may be used for biomedical or health research purposes only.
All of Us data cannot be linked with external datasets at the individual participant level.
Participant-level data cannot be shared, redistributed, or published.
Published statistics must reflect a minimum of 20 participants per cell (aka the ’minimum count rule’). Public Tier applies this same rounding convention.
The program must be notified at least two weeks before any publication or presentation using the data.
Final manuscripts must be submitted to PubMed Central immediately upon publication with no embargo period.
Each Workspace must include an accurate and meaningful description of the research, which becomes publicly viewable in the Research Projects Directory.
References
All of Us Research Program Investigators; Denny JC, Rutter JL, Goldstein DB, Philippakis A, Smoller JW, Jenkins G, Dishman E. (2019). The "All of Us" Research Program. N Engl J Med. 2019 Aug 15;381(7):668-676. doi: 10.1056/NEJMsr1809937. PMID: 31412182; PMCID: PMC8291101.
Mayo KR, Basford MA, Carroll RJ, Dillon M, Fullen H, Leung J, et al. The All of Us Data and Research Center: Creating a Secure, Scalable, and Sustainable Ecosystem for Biomedical Research. Annu Rev Biomed Data Sci. 2023 Aug 10;6:443-464. doi: 10.1146/annurev-biodatasci-122120-104825. PMID: 37561600.
Ginsburg GS, Denny JC, Schully SD. Data-driven science and diversity in the All of Us Research Program. Sci Transl Med. 2023 Dec 13;15(726):eade9214. doi: 10.1126/scitranslmed.ade9214. Epub 2023 Dec 13. PMID: 38091411; PMCID: PMC12374748.
National Institutes of Health. (n.d.). All of Us Research Program.
All of Us Research Hub. (n.d.). Data Snapshots. .
All of Us Research Hub. (n.d.). Data Methods.
Resources
Links
Last updated
