Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
10x Genomics maintains a repository of publicly available datasets generated through three main platforms: Chromium, Visium, and Xenium. The repository includes Chromium single-cell molecular profiles, Visium spatial transcriptomics (ST) data, and Xenium single-cell spatial gene expression measurements. Datasets span a wide range of tissue types, species, and disease states, with each platform offering multiple dataset formats and analysis outputs depending on the assay chosen. As an example, the Human Glioblastoma: Whole Transcriptome Analysis dataset is a Visium dataset containing ST data on fresh frozen human glioblastoma multiforme tissue, and will be referenced throughout this overview to illustrate typical dataset structure and content.
Datasets in the 10x repository are generated using standardized wet-lab and computational protocols specific to each platform. For Visium datasets, tissue is typically obtained from a specialized tissue bank (e.g., BioIVT Asterand, the source for the glioblastoma dataset), encased in a supportive medium, rapidly frozen, and sliced into 10 µm-thick layers. A slice of tissue is placed onto a Visium Gene Expression slide, which is embedded with thousands of microscopic capture spots acting as biological barcodes. The tissue is fixed (typically with methanol), stained with Hematoxylin and Eosin (H&E), and imaged using a microscope such as the Nikon Eclipse Ti2-E. To capture genetic data, a sequencing instrument such as the Illumina NovaSeq 6000 is used to read the RNA in the tissue and build a genomic library accordingly. The resulting data are run through the software Space Ranger to match the genetic code back to the tissue image. In the glioblastoma dataset specifically, this process yielded 3,468 distinct spots occupied by tissue on the slide, with a median of 4,326 genes and 11,596 unique molecular identifier (UMI) counts detected per spot. Figures like these vary by dataset, tissue type, and sequencing depth.
To access a dataset, create a free account through 10x Genomics. Navigate to the Datasets page and use the sidebar to filter by platform, software, species, sample/tissue type, and more. Choose a dataset and look at the "Dataset overview" tab for metadata such as the source, preparation, imaging, sequencing, and key metrics. Look at the "Output and supplemental files" tab for the files to download. For most ST analyses, a minimum of four files are needed: the gene expression matrix, available in either HDF5 (.h5) or Matrix Market (.mtx.gz) format, with .h5 preferred because it is a single file and typically smaller; the spatial coordinates, which map each spot to its location on the slide, available in .csv format; the histology image, required for overlaying expression onto the tissue, available in .png format; and the scale factors, required for aligning spot coordinates with the image, available in .json format. As an example, in the glioblastoma dataset, the display text for the .h5 file is "Feature / barcode matrix HDF5 (filtered)," which downloads a file named "filtered_feature_bc_matrix.h5." For the other three files, the display text "Spatial imaging data" downloads a folder named "spatial" that contains "tissue_positions_list.csv," "tissue_hires_image.png," and "scalefactors_json.json."
In terms of analysis, the software toolkit Seurat is designed to help beginners analyze ST data. Seurat contains user-friendly workflows that streamline complex computational steps to visualize where specific genes are active and identify distinct cell types. Other built-in functions of Seurat include clustering, quality control, and data normalization. While Seurat provides a focused framework for ST analysis, there is also a rapidly expanding ecosystem of ST software. To help navigate these options, Gillespie et al. conducted a meta-review synthesizing benchmarking studies across tissue architecture identification, spatially variable gene detection, cell-cell communication analysis, and deconvolution.This paper can be referenced to help with software selection for ST analysis.
In 10x Genomics gene expression matrices, genes are typically identified using HUGO Gene Nomenclature Committee (HGNC) gene symbols, which are unique abbreviations assigned to each human gene. For example, in the glioblastoma .h5 file, the first three genes are MIR1302-2HG, FAM138A, and OR4F5.
Hao Y, Stuart T, Kowalski MH, Choudhary S, Hoffman P, Hartman A, Srivastava A, Molla G, Madad S, Fernandez-Granda C, Satija R. Nat Biotechnol. 2024 Feb;42(2):293-304.
Gillespie J, Pietrzak M, Song MA, Chung D. Cells. 2025 Jul 10;14(14):1060.
Jaume G, Doucet P, Song AH, Lu MY, Almagro-Pérez C, Wagner SJ, Vaidya AJ, Chen RJ, Williamson DF, Kim A, Mahmood F. Adv Neural Inf Process Syst. 2024 Dec 10;37:53798-833.
Zhang L, Sagan A, Qin B, Wang H, Kim E, Hu B, Osmanbeyoglu HU. Nucleic Acids Res. 2026 Jan 5;54(1):gkaf1473.
Visium-Generated Spatial Transcriptomics Dataset
Row Names (barcodes)
character
unique 16-base pair spatial barcode to identify the physical Visium spot
AAACAGTGTTCCTGGG-1
x
Matrix Object
dgCMatrix
compressed sparse matrix structure that contains all counts
36,601 genes x 3,468 spots
Row Names (features)
numeric
horizontal coordinate on the 2D plane of the image
2361 to 10354 pixels
y
numeric
vertical coordinate on the 2D plane of the image
1442 to 10673 pixels
cell
character
duplicate of spatial barcodes maintained by Seurat
AAACAGTGTTCCTGGG-1
character
biological features in HGNC Gene Symbols
OR4F5
Column Names (cells/spots)
character
unique 16-base pair spatial barcode to identify the physical Visium spot
AAACCGTTCGTCCAGG-1
Matrix Values
numeric
raw UMI transcript counts per gene, per spot
0 to 1,118 UMIs
Table of Contents
10x Genomics maintains a repository of publicly available datasets generated through three main platforms: Chromium, Visium, and Xenium. The repository includes Chromium single-cell molecular profiles, Visium spatial transcriptomics (ST) data, and Xenium single-cell spatial gene expression measurements.
The All of Us dataset is a large, longitudinal U.S. health research resource that integrates clinical, behavioral, genomic, and wearable-device data from a diverse participant population.
The Behavioral Risk Factor Surveillance System (BRFSS) was established in 1984 by the U.S. Centers for Disease Control and Prevention (CDC) to collect prevalence data regarding risk behaviors and preventive health practices among adult residents across all 50 states, the District of Columbia, and participating U.S. territories.
The cBioPortal for Cancer Genomics is an open-access, open-source platform for the interactive exploration of multidimensional cancer genomics datasets with an emphasis on multimodal cancer visualization.
The DailyMed Full SPL Dataset refers to the downloadable full release of all drug labels in the United States.
The Healthcare Cost and Utilization Project (HCUP) is a family of databases, software tools, and related products developed through a Federal-State-Industry partnership and sponsored by the Agency for Healthcare Research and Quality (AHRQ).
The Medical Information Mart for Intensive Care (MIMIC) is a publicly available repository of deidentified electronic health record (EHR) data from patients admitted to the Beth Israel Deaconess Medical Center (BIDMC) in Boston, Massachusetts.
The National Plan and Provider Enumeration System (NPPES) is the national registry, operated by the Centers for Medicare & Medicaid Services (CMS), that assigns and maintains the National Provider Identifier (NPI) — the standard unique identifier for health care providers in the United States.
The Pediatric Psychiatry Emergency Services (PediPES) dataset consists of electronic health record (EHR) data for patients younger than 18 years old evaluated in the emergency department at Hasbro Children's Emergency Department in Providence, Rhode Island for psychiatric concerns.
The Surveillance, Epidemiology, and End Results (SEER) Program provides information on cancer statistics in an effort to reduce the cancer burden among the U.S. population.
SyntheticMass is a Synthea-generated data set that contains realistic but fictional residents of the state of Massachusetts.
The SyntheticRI 2026 datasets consist of patient-centric synthetic health records representing the population of Rhode Island.
The Youth Risk Behavior Surveillance System (YRBSS) was established in 1990 by the U.S. Centers for Disease Control and Prevention (CDC) to monitor priority health risk behaviors and experiences that contribute to death, disability, and social problems among youth and young adults in the United States.
A core component of the URSA Initiative is Brown's Stronghold. Stronghold is a secure computing and storage environment that enables Brown researchers and associates to analyze sensitive data while complying with regulatory or contractual requirements.
Stronghold is maintained by the Brown University's Center for Computation and Visualization (CCV) in the Office of Information Technology (OIT). To learn more about Stronghold, you may refer to CCV's Stronghold Documentation.
URSA Stronghold is one of the research "tenants" within Stronghold. It is collaboratively managed by CCV, the Brown Center for Biomedical Informatics (BCBI), and the Biomedical Informatics, Bioinformatics, and Cyberinfrastructure Enhancement (BIBCE) Core of Advance RI-CTR. URSA Stronghold offers both Linux and Windows computing platforms; database management systems such as Microsoft SQL, PostgreSQL, and MySQL; and, a broad range of data analysis tools including Julia, Python, R, SAS, and Stata.
Researchers may work within the URSA tenant or request their own dedicated Stronghold tenant. Refer to Brown CCV's documentation for more information about available features and how to request a Stronghold tenant.
Oscar (Ocean State Center for Advanced Resources) is Brown University's high performance computing cluster. Oscar is maintained and supported by Brown's Center for Computation and Visualization ().
Brown’s are used to determine data access and storage options. Stronghold is the institutionally-designated environment for Risk Level 3, identified datasets that include protected health information () or personally identifiable information (PII). With approval from the data provider, researchers may leverage Brown's high-performance computing environment or other Brown-managed computers for the storage and analysis of (Brown Risk Level 2) datasets. In addition, Brown’s can be used for storage of de-identified data and other research files. Data analysis is initiated "locally" on a researcher's computer while the data remain secure.
SyntheticMass is a Synthea-generated data set that contains realistic but fictional residents of the state of Massachusetts. The synthetic population aims to statistically mirror the state population in terms of demographics, disease burden, vaccinations, medical visits, and social determinants. [1] Refer to the SyntheticMass website for more information.
There are several data sets available on the SyntheticMass downloads page.
Complete SytheticMass data sets: "SyntheticMass Data Version 2 (24 May, 2017)". This ZIP file is quite large (21GB), so make sure you move the file to a location with enough storage before attempting to unzip.
Sample data sets (<100MB) containing 100 or 1,000 patient records
Specialized data sets that have been generated using Synthea by other study teams. These include COVID-19 data sets, a Childhood Obesity data set and more.
The following versions are available for each data set.
CSV ( describing all CSV tables)
C-CDA (xml files)
FHIR (json files)
MITRE Corporation. (n.d.). . Retrieved May 1, 2024.
Walonoski J, Klaus S, Granger E, Hall D, Gregorowicz A, Neyarapally G, Watson A, Eastman J. . Intell Based Med. 2020 Nov;1:100007. doi: 10.1016/j.ibmed.2020.100007. Epub 2020 Oct 2. PMID: 33043312; PMCID: PMC7531559.
Walonoski J, Kramer M, Nichols J, Quina A, Moesel C, Hall D, Duffett C, Dube K, Gallagher T, McLachlan S. . J Am Med Inform Assoc. 2018 Mar 1;25(3):230-238. doi: 10.1093/jamia/ocx079. Erratum in: J Am Med Inform Assoc. 2018 Jul 1;25(7):921. PMID: 29025144; PMCID: PMC7651916.
NPPES follows a fixed schema, so the full file supports aggregate statistics across the entire registry — summarized below for the most recent monthly release.
1,917,366
(blank — deactivated)
340,360
Active
9,154,078
Deactivated
340,360
F — Female
4,952,833
M — Male
2,246,911
X — Undisclosed
35,500
2005
242,759
2006
1,190,765
2007
793,497
CA
1,145,592
NY
645,005
FL
634,992
106S00000X
Behavior Technician
546,913
1041C0700X
Social Worker, Clinical
U — Undisclosed
1,467
(blank)
1
2008
334,749
2009
242,550
2010
258,206
2011
275,627
2012
286,574
2013
269,626
2014
286,575
2015
297,458
2016
318,321
2017
329,997
2018
391,037
2019
399,994
2020
427,550
2021
463,987
2022
468,470
2023
509,992
2024
548,097
2025
631,472
2026 (partial)
186,775
TX
588,935
OH
395,358
MI
344,935
(blank — deactivated)
340,362
PA
327,945
IL
310,948
NC
267,983
MA
253,121
WA
237,740
NJ
228,549
GA
224,337
MD
207,620
VA
195,085
CO
191,225
AZ
171,484
MN
163,228
TN
156,146
IN
150,919
OR
146,773
MO
145,086
WI
139,254
LA
123,091
341,516
(none — deactivated)
No primary taxonomy on record
340,360
101YM0800X
Counselor, Mental Health
316,829
183500000X
Pharmacist
293,460
225100000X
Physical Therapist
290,096
390200000X
Student in an Organized Health Care Education/Training Program
277,093
363LF0000X
Nurse Practitioner, Family
241,779
235Z00000X
Speech-Language Pathologist
208,312
207Q00000X
Family Medicine (Physician)
206,434
207R00000X
Internal Medicine (Physician)
189,130
1223G0001X
Dentist, General Practice
177,871
363A00000X
Physician Assistant
161,694
171M00000X
Case Manager/Care Coordinator
160,060
101YP2500X
Counselor, Professional
156,214
111N00000X
Chiropractor
152,146
163W00000X
Registered Nurse
143,191
225X00000X
Occupational Therapist
139,532
104100000X
Social Worker
136,288
122300000X
Dentist
122,501
103K00000X
Behavior Analyst
121,669
101Y00000X
Counselor
110,647
106H00000X
Marriage & Family Therapist
107,989
363L00000X
Nurse Practitioner
100,171
101YA0400X
Counselor, Addiction (Substance Use Disorder)
98,615
The National Plan and Provider Enumeration System (NPPES) is the national registry, operated by the Centers for Medicare & Medicaid Services (CMS), that assigns and maintains the National Provider Identifier (NPI) — the standard unique identifier for health care providers in the United States. Established under the Administrative Simplification provisions of HIPAA, NPPES enumerates both individual clinicians and provider organizations, each of which receives a single, permanent 10-digit NPI used to identify them in administrative and financial transactions such as claims, eligibility checks, and remittances. [1,5] The system holds millions of active and deactivated records and functions as core infrastructure for the U.S. health care system, where the NPI serves as a common key for linking providers across enrollment, claims, and quality datasets.
The data are generated primarily through self-reported provider enumeration. A provider becomes part of NPPES by applying for an NPI and submitting the information that forms the core of its record. [2]
Upon a successful application, NPPES generates and assigns a single 10-position, intelligence-free NPI. The identifier is assigned once, is permanent, and is not affected by later changes to name, address, or other fields; nor is it reused or reassigned to a different provider.5 Because the record is built from provider-supplied information, the database depends on providers to maintain their own data and to submit updates when their circumstances change. To keep provider enrollment and enumeration information aligned, NPPES is integrated with CMS's Provider Enrollment, Chain, and Ownership System (PECOS) through an API. Updates made in PECOS can automatically populate the corresponding NPPES record, helping ensure provider information remains accurate and consistent across both systems.
Records can be deactivated as well as created. NPPES carries a deactivation reason code for each deactivated record — death, disbandment, fraud, or other — together with a deactivation date, and these records are retained (and, since 2018, included in the full monthly file rather than dropped). A deactivated NPI is never reissued to another provider and may be reactivated if the original provider returns to practice. CMS does not independently verify most of the self-reported content of a record, which is the methodological reason the data are comprehensive in coverage but variable in currency and accuracy. The publicly available dataset is the FOIA-disclosable subset of these records, extracted and released on the cycles described in Data Access and Dissemination.
CMS releases only the subset of each NPPES record that is disclosable under the Freedom of Information Act; identifying elements such as Social Security Number, ITIN, and date of birth are withheld. Public dissemination of this FOIA-disclosable data began in September 2007 and is offered through three channels:
NPI Registry — a free, real-time web lookup for individual provider records, suited to verifying or retrieving a single NPI.
NPI Registry API — a public programming interface (npiregistry.cms.hhs.gov) that returns individual records as JSON in response to per-record queries. It is built for record-level lookups and does not provide bulk export.
Downloadable files — the route for working with the dataset in bulk. CMS publishes a full replacement monthly NPI file, weekly incremental files (covering NPIs newly assigned, updated, or deactivated during that week), and a monthly deactivation file in Excel format for users who need only deactivated NPIs and their deactivation dates.
The downloadable package is a single ZIP containing a Read Me, the Code Values reference, a Header (file-layout) document, and the FOIA-disclosable data file itself — a comma-separated file carrying both the header row and the data. Since June 18, 2018, the package has also bundled three reference files that externalize repeating attributes so the main file stays manageable: an Other Name reference file (additional names for Type 2 organizations), a Practice Location reference file (non-primary practice locations for Types 1 and 2), and an Endpoint reference file (health information exchange endpoints for Types 1 and 2). Each links back to the core record by NPI.
The weekly file is supplemental and must be used alongside the most recent monthly full-replacement file, which should be re-downloaded each month to stay current. CMS does not retain historical copies of the downloadable file — each release reflects the present state of NPPES — and the full data file contains several million records (individual and organizational, active and deactivated), making it too large to open directly in standard spreadsheet software.
Centers for Medicare & Medicaid Services (US). [Internet]. Baltimore (MD): CMS; 2006 Mar [cited 2026 Jun 12].
Centers for Medicare & Medicaid Services (US). [Internet]. Baltimore (MD): CMS; [cited 2026 Jun 12].
Centers for Medicare & Medicaid Services (US). [Internet]. Baltimore (MD): CMS; [cited 2026 Jun 12].
Bindman AB. Medicare Medicaid Res Rev. 2013;3(3):mmrr.003.03.b03. Published 2013 Jul 30. doi:10.5600/mmrr.003.03.b03. PMID: 24753977; PMCID: PMC3983736.
Department of Health and Human Services (US), Centers for Medicare & Medicaid Services. Fed Regist. 2004 Jan 23;69(15):3434–69.
Centers for Medicare & Medicaid Services (US). [Internet]. Baltimore (MD): CMS; [cited 2026 Jun 16].
Centers for Medicare & Medicaid Services (US). [Internet]. Baltimore (MD): CMS; [cited 2026 Jun 16].
Centers for Medicare & Medicaid Services (US). [Internet]. Baltimore (MD): CMS; 2026 [cited 2026 Jun 16].
Allowed values for coded fields. Source: CMS NPPES Code Values, February 1, 2025.
The Healthcare Cost and Utilization Project (HCUP) is a family of databases, software tools, and related products developed through a Federal-State-Industry partnership and sponsored by the Agency for Healthcare Research and Quality (AHRQ). HCUP databases are derived from administrative data and contain encounter-level, clinical, and nonclinical information including all-listed diagnoses and procedures, discharge status, patient demographics, and charges for all patients, regardless of payer, beginning in 1988. [1]
HCUP offers the following databases:
National (Nationwide) Inpatient Sample (NIS): largest publicly available all-payer hospital inpatient care database in the United States
Kids' Inpatient Database (KID): hospital inpatient stays for children and is specifically designed to allow researchers to study a broad range of conditions and procedures related to children's health
Nationwide Emergency Department Sample (NEDS): emergency department (ED) visits that do not result in an admission as well as ED visits that result in an admission to the same hospital
Nationwide Readmissions Database: designed to support various types of analyses of national readmission rates for all payers and uninsured individuals
State Inpatient Databases (SID): inpatient discharge abstracts from participating States, translated into a uniform format to facilitate multi-State comparisons and analyses
State Ambulatory Surgery and Services Databases (SASD): encounter-level data for ambulatory surgery and other outpatient services from hospital-owned facilities
State Emergency Department Databases (SEDD): discharge information on all emergency department visits that do not result in an admission
Please note: Access to HCUP databases is not free. Database releases must be purchased through the .
Visit for a more information about the databases. Learn more about HCUP on the .
Agency for Healthcare Research and Quality. (n.d.). . Retrieved May 1, 2024.
Below is a list of current Health Data Partners (HDPs) engaged with the URSA Initiative.
To learn more about requesting data from any of these HDPs, please submit the Advance RI-CTR Service Request Form.
Since March 2015, Brown University Health has used LifeChart built on the Epic EHR platform. There are a variety of reporting and analytic tools, such as SlicerDicer and Reporting Workbench, which can be used within Brown University Health. Reports or data extracts can also be requested.
Care New England uses the Epic and Cerner electronic health record (EHR) systems for outpatient and inpatient respectively. Both systems have a suite of reporting and analytic tools for use within the health systems. Reports and data extracts can also be requested.
The Rhode Island Department of Health supports and maintains a range of health datasets and systems that can be used for research.
(RI APCD or HealthFacts RI) is a large-scale database that systematically collects healthcare claims data from a variety of payer sources, including Medicare, Medicaid, and RI’s nine largest commercial payers.
Founded in 2001, the Rhode Island Quality Institute (RIQI) is the state-designated Regional Health Information Organization (RHIO). RIQI manages , Rhode Island’s state-designated Health Information Exchange (HIE). CurrentCare contains EHR and health data from all acute care hospital systems in Rhode Island and from many ambulatory and laboratory facilities across the state. RIQI has provided public health and research data to external partners such as the Rhode Island Department of Health and Brown University. More information about the data in CurrentCare can be found in the .
The SyntheticRI 2026 datasets consist of patient-centric synthetic health records representing the population of Rhode Island. The data was generated using Synthea™, an open-source simulation framework that models the clinical journey of synthetic patients from birth to death. The primary objective of these datasets is to provide a realistic environment for healthcare research, software development, and clinical informatics training.
These datasets contains no real-world patient data. Because all individuals and clinical events are computationally generated from the ground up, the records are entirely free of personally identifiable information (PII). This allows for unrestricted sharing and analysis in compliance with HIPAA and GDPR standards.
To use these datasets for educational or reseach purposes, please contact bcbi@brown.edu
The patient population and their respective health trajectories were simulated using the Synthea™ generator (Walonoski et al., 2018). The simulation utilizes a modular, logic-based approach where clinical outcomes are driven by transition-based state machines.
Demographics: Patient demographics, including age, gender, race, and socio-economic status, are statistically aligned with the United States Census Bureau data for the state of Rhode Island to ensure regional representativeness.
X
Undisclosed
Other
Other reason
4
Former Legal Business Name — Organization
5
Other Name — Both
Mrs.
Mrs.
Dr.
Dr.
Prof.
Prof.
II
II
III
III
IV
IV
V
V
VI
VI
VII
VII
VIII
VIII
IX
IX
X
X
Y
Yes — the individual is a sole proprietor
N
No — not a sole proprietor
X
Not answered
Y
Yes — the organization is a subpart
N
No — not a subpart
X
Not answered
M
Male
F
Female
U
Undisclosed
Death
Provider is deceased
Disbandment
Organization has disbanded
Fraud
NPI was obtained fraudulently
1
Former Name — Individual
2
Professional Name — Individual
3
Doing Business As — Organization
Ms.
Ms.
Mr.
Mr.
Miss
Miss
Jr.
Jr.
Sr.
Sr.
I
I
Y
Yes — primary taxonomy (only one per NPI)
N
No — not the primary taxonomy
X
Not answered
193200000X
Multi-Specialty Group — practitioners with different specializations
193400000X
Single Specialty Group — practitioners with the same specialization
1
Other
5
Medicaid
—
Note: legacy values (02 Medicare UPIN, 04 Medicare ID, 06 BCBS, 07 Medicare PIN) may appear in older records.
Immunology & Allergy
Anemia - Unknown Etiology
Hematology
Appendicitis
Gastroenterology
Asthma
Pulmonology
Atopy
Immunology & Allergy
Atrial Fibrillation
Cardiology
Attention Deficit Disorder
Neurology
Bone Marrow Transplant
Oncology
Breast Cancer
Oncology
Bronchitis
Pulmonology
Cerebral Palsy
Neurology
Chronic Kidney Disease
Nephrology
Colorectal Cancer
Oncology
Congestive Heart Failure
Cardiology
Contraceptive Maintenance
Reproductive Health
Contraceptives
Reproductive Health
COPD
Pulmonology
Covid19
Infectious Disease
Cystic Fibrosis
Pulmonology
Dementia
Neurology
Dental and Oral Examination
Dental
Dentures
Dental
Diabetic Retinopathy Treatment
Ophthalmology
Dialysis
Nephrology
Ear Infections
Infectious Disease
Epilepsy
Neurology
Female Reproduction
Reproductive Health
Fibromyalgia
Rheumatology & Musculoskeletal
Food Allergies
Immunology & Allergy
Gallstones
Gastroenterology
Gout
Rheumatology & Musculoskeletal
HIV Care
Infectious Disease
HIV Diagnosis
Infectious Disease
Home Heatlh Treatment
Health Services & Social Determinants
Home Hospice SNF
Health Services & Social Determinants
Homelessness
Health Services & Social Determinants
Hospice Treatment
Health Services & Social Determinants
Hypertension
Cardiology
Hypothyroidism
Endocrinology & Metabolism
Injuries
Trauma & Emergency
Kidney Transplant
Nephrology
Lung Cancer
Oncology
Lupus
Rheumatology & Musculoskeletal
Veteran Mild TBI
Neurology
Medication Reconciliation
Health Services & Social Determinants
MEND Program
Endocrinology & Metabolism
Metabolic Syndrome Standards of Care
Endocrinology & Metabolism
Metabolic Syndrome Disease Progresion
Endocrinology & Metabolism
Myocardial Infarction
Cardiology
Opiod Addiction
Substance Abuse
Osteoarthritis
Rheumatology & Musculoskeletal
Osteoporosis
Rheumatology & Musculoskeletal
Pregnancy
Reproductive Health
Prescribing Optiods for Chronic Pain and Treatmebt of OUD
Substance Abuse
Rheumatoid Arthritis
Rheumatology & Musculoskeletal
Self Harm
Behavioral Health & Psychiatry
Sepsis
Infectious Disease
Sexual Activity
Reproductive Health
Sinusitis
Pulmonology
Sleep Apnea
Pulmonology
Sore Throat
Infectious Disease
Spina Bifida
Neurology
Stable Ischemic Heart Disease
Cardiology
Stroke
Cardiology
Total Joint Replacement
Rheumatology & Musculoskeletal
Trigger Bone Marrow Transplant
Oncology
Urinary Tract Infections
Infectious Disease
Veteran
Health Services & Social Determinants
Veteran Hyperlipidemia
Cardiology
Veteran Lung Cancer
Oncology
Veteran MDD
Behavioral Health & Psychiatry
Veteran Prostate Cancer
Oncology
Veteran PTSD
Behavioral Health & Psychiatry
Veteran Self Harm
Behavioral Health & Psychiatry
Veteran Substance Abuse Conditions
Substance Abuse
Veteran Substance Abuse Treatment
Substance Abuse
VHD Aortic
Cardiology
VHD Mitral
Cardiology
VHD Pulmonic
Cardiology
VHD Tricuspid
Cardiology
Wellness Encounters
Health Services & Social Determinants
Acute Myeloid Leukemia for PCOR Research
Oncology
Allergic Rhinitis
Immunology & Allergy
Allergies and Treatment
Clinical Logic: Disease progression, treatment pathways, and clinical encounters are dictated by over 80 validated clinical modules. These modules are developed based on peer-reviewed clinical guidelines and public health statistics from sources such as the CDC, NIH, and various specialty medical associations.
The commands used to generate these datasets, including the specific seed, are as follows:
SyntheticRI records include codes from the following vocabularies:
Conditions and Diagnoses: SNOMED-CT (Systematized Nomenclature of Medicine – Clinical Terms)
Observations and Vitals: LOINC (Logical Observation Identifiers Names and Codes)
Medications: RxNorm
Vaccines: CVX (Clinical Vaccine) codes
Walonoski, J., Klaus, M., Granger, E., Hall, D., Gregorowicz, A., Neyarapally, G., Watson, A., & McLachlan, J. (2018). Synthea: An approach, method, and software mechanism for generating synthetic electronic health records. Journal of the American Medical Informatics Association, 25(3), 230–238.
Synthea Documentation & Wiki: Synthetic Health Research. (n.d.). Synthea Wiki.
Source Code: Synthetic Health. (2024). Synthea Patient Population Simulator.
./run_synthea -s 20260508 -p 1000 "Rhode Island"
./run_synthea -s 20260508 -p 300000 "Rhode Island"The Medical Information Mart for Intensive Care (MIMIC) is a publicly available repository of deidentified electronic health record (EHR) data from patients admitted to the Beth Israel Deaconess Medical Center (BIDMC) in Boston, Massachusetts. This resource was created in collaboration with Massachusetts Institute of Technology to lower barriers to reproducible clinical research by making real world hospital data accessible in a deidentified form.
MIMIC is organized into four datasets, which altogether contain five modules. The following info pertain to their latest versions:
MIMIC-IV (version 3.1): This dataset contains info derived from 2008-2022.
The Hosp module contains hospital-level tables related to 546,028 hospitalizations for 223,452 unique individuals.
The ICU module is for for intensive care unit admissions and contains 94,458 ICU stays for 65,366 unique individuals
MIMIC-IV-ED (version 2.2):
The ED module contains tables related to emergency department (ED) visits from 2011 to 2019. There are 425,087 ED visits of which 203,016 led to hospital admissions.
MIMIC-IV-Note (version 2.2):
The Note module contains tables related to 331,794 deidentified discharge summaries from 145,915 patients admitted to the hospital and emergency department.
MIMIC-CXR (latest version 2.0) :
The CXR module contains tables related to 377,110 images corresponding to 227,835 radiographic studies in Digital Imaging and Communications in Medicine (DICOM) format with free text radiology reports made in the emergency department from 2011 to 2016
Other notable modules include the following:
The ECG module contains approximately 800,000 diagnostic electrocardiograms from nearly 160,000 unique patients. These diagnostic ECGs use 12 leads, are 10 seconds in length and are sampled at 500 Hz.
The ECHO module contains 206,488 echocardiogram studies (including 179,928 transthoracic, 16,389 stress, and 10,171 transesophageal echocardiograms) from 91,372 unique patients between 2008 and 2022.
The MIMIC-IV-ECHO module contains structured echocardiographic measurements and DICOM files from echocardiography exams.
Data across the MIMIC-IV family were sourced from Beth Israel Deaconess Medical Center's (BIDMC) systems, with each component drawn from their respective sources. For instance, MIMIC-IV Core drew from the hospital-wide EHR and ICU bedside systems, excluding patients under 18 at first visit or on an enhanced-protection list; MIMIC-IV-Note included only notes within one year of a hospital or ED encounter; MIMIC-IV-ED drew emergency department records in XML format; and MIMIC-CXR separately sourced chest radiograph studies via the local EHR and Radiology Information System (RIS), alongside their associated free-text reports.
Raw data were converted into structured, analysis-ready formats while preserving the integrity of the original clinical record. MIMIC-IV was transformed via custom SQL into a schema organized into the Hosp and ICU modules, designed for backward compatibility with MIMIC-III, with tables exported as CSV files and no data cleaning applied so the dataset reflects real-world clinical data as recorded. Out-of-hospital deaths were linked to Massachusetts state vital records using a custom workflow combining exact and fuzzy matching (name, date of birth, social security number), prioritizing sensitivity over specificity. MIMIC-IV-ED's XML extracts were converted into a denormalized relational database to support analysis. MIMIC-CXR integrated three separately processed streams into a single unified dataset: EHR data, DICOM-format imaging, and extracted radiology reports (stripped of administrative/clinical metadata).
All datasets meet HIPAA Safe Harbor requirements, though each applies methods suited to its data type. Structured data across MIMIC-IV, MIMIC-IV-ED, and MIMIC-CXR had direct identifiers removed and replaced with randomly assigned IDs (subject_id, hadm_id, stay_id) that remain consistent across datasets to enable cross-linkage. Dates were shifted using patient-specific offsets to preserve within-patient chronological relationships. Free-text and unstructured data required additional handling: MIMIC-IV-Note and MIMIC-CXR's radiology reports were deidentified using a combination of rule-based methods, a purpose-built neural network, and manual review; MIMIC-IV-ED's free-text fields were scrubbed and replaced with placeholders; and MIMIC-CXR's DICOM imaging had PHI removed from both metadata and pixel data via a custom algorithm with manual validation. Finally, released datasets were exported in a character-based, comma-delimited format for distribution.
The MIMIC team provides and maintains code in the public mimic-code repository () to generate tables and common variables such as severity scores to make analysis easier.
Access to MIMIC is provided via Physionet and restricted. For Brown University staff and students, access can be provided on OSCAR after meeting the following requirements. Researchers generally need to be credentialed, complete human subjects training via CITI courses and sign a data use agreement before obtaining the full dataset. For more detailed instructions, please refer to your Primary Investigator or Brown faculty.
ICD-9-CM diagnoses
ICD-9-PCS procedures
ICD-10-CM diagnoses
ICD-10-PCS procedures
National Drug Code (NDC)
Logical Observation Identifiers, Names and Codes(LOINC)
DICOM
, Pub. L. No. 104-191.
Johnson AEW, Bulgarelli L, Shen L, Gayles A, Shammout A, Horng S, Pollard TJ, Hao S, Moody B, Gow B, Lehman LH, Celi LA, Mark RG. Sci Data. 2023 Jan 3;10(1):1. doi: 10.1038/s41597-022-01899-x. Erratum in: Sci Data. 2023 Jan 16;10(1):31. doi: 10.1038/s41597-023-01945-2. Erratum in: Sci Data. 2023 Apr 18;10(1):219. doi: 10.1038/s41597-023-02136-9. PMID: 36596836; PMCID: PMC9810617.
Johnson AEW, Pollard TJ, Berkowitz SJ, Greenbaum NR, Lungren MP, Deng CY, Mark RG, Horng S. Sci Data. 2019 Dec 12;6(1):317. doi: 10.1038/s41597-019-0322-0. PMID: 31831740; PMCID: PMC6908718.
Goldberger AL, Amaral LA, Glass L, Hausdorff JM, Ivanov PC, Mark RG, Mietus JE, Moody GB, Peng CK, Stanley HE. Circulation. 2000 Jun 13;101(23):E215-20. doi: 10.1161/01.cir.101.23.e215. PMID: 10851218.
Unless otherwise noted, information in this overview is adapted from the official cBioPortal documentation at docs.cbioportal.org. See references.
The cBioPortal for Cancer Genomics is an open-access, open-source platform for the interactive exploration of multidimensional cancer genomics datasets with an emphasis on multimodal cancer visualization. This resource aims to significantly lower the barriers between complex genomic data and cancer researchers by providing rapid, intuitive, and high-quality access to molecular profiles and clinical attributes from large-scale cancer genomics projects, empowering researchers, clinicians, and translational scientists to extract biological insights and create clinical applications from these rich datasets. cBioPortal provides visualization, analysis, and download from a wide range of datasets predominantly sourced from The Cancer Genome Atlas, with a growing number of institutional and consortium cohorts added over time. Data types span somatic mutations, structural variants and gene fusions, copy number alterations, mRNA and microRNA expression, DNA methylation, and protein/phosphoprotein levels (RPPA), with samples also linked to clinical and treatment attributes. The platform was originally developed at Memorial Sloan Kettering Cancer Center and is now maintained by a multi-institutional team consisting of MSK, the Dana-Farber Cancer Institute, Princess Margaret Cancer Centre, Children's Hospital of Philadelphia, Caris Life Sciences, The Hyve, SE4BIO, and Bilkent University.
cBioPortal was developed to address the data-integration challenges posed by large-scale cancer genomics projects, making raw data from these efforts more easily and directly available to the cancer research community, supporting both basic and translational research. Each study loaded into the portal is built from a standardized set of source files (clinical data tables, mutation files, copy-number segment files, expression matrices, etc.) that are validated and harmonized before import. cBioPortal generally does not generate primary experimental data; it hosts, curates, annotates, harmonizes, and visualizes data generated by external projects, consortia, publications, or institutional pipelines.
The Behavioral Risk Factor Surveillance System (BRFSS) was established in 1984 by the U.S. Centers for Disease Control and Prevention (CDC) to collect prevalence data regarding risk behaviors and preventive health practices among adult residents across all 50 states, the District of Columbia, and participating U.S. territories. The state-based, cross-sectional telephone survey was conducted monthly by state health departments. The survey has consisted of standard core questions on emerging health issues, rotating core questions asked every other year, optional modules, and state-added questions on health priorities specific to the respective state. Collected data was processed through editing, weighting, and analysis to provide critical information on health status. The BRFSS aims to inform the management of health-related policies and priorities at both the state and national levels. Published data, reports, questionnaires, and analysis tools are publicly accessible through the CDC BRFSS website. Data files are provided in ASCII (fixed record length) and SAS Transport (.XTP) formats.
The BRFSS dataset was created using a standardized questionnaire developed in collaboration between the CDC and U.S. state health departments. The survey consisted of a standard core, rotating core, and optional modules. Data was collected by states or private contractors following CDC protocols, which included multiple call attempts and specific schedules to ensure representative participation on weeknights and weekends. The sampling design utilized disproportionate stratified sampling (DSS) for landlines to increase the efficiency of finding residential numbers and a randomly generated frame for cellular phones. The collected data underwent design weighting–which adjusts for the number of adults and telephones in a household–followed by iterative proportional fitting–which aligns the sample with population demographics like age, sex, race, education, and phone ownership to ensure the final results are representative of the population.
HCPCS
N/A
Q2
18-18
What is your sex?
1=Female; 2=Male
N/A
Q3
19-19
In what grade are you?
1=9th; 2=10th; 3=11th; 4=12th; 5=Other
N/A
Q4
20-20
Are you Hispanic or Latino?
1=Yes; 2=No
N/A
Q5
21-28
What is your race?
A=Am. Indian; B=Asian; C=Black; D=Native Hawaiian; E=White
N/A
Variable Name
Data Location
Question Text / Label
Value Codes & Labels
Response of Interest (ROI)
Q6
29-32
Height (meters)
Continuous (from feet/inches)
Variable Name
Data Location
Question Text / Label
Value Codes & Labels
Response of Interest (ROI)
Q8
39-39
Seat belt use
1=Never; 2=Rarely; 3=Sometimes; 4=Most of the time; 5=Always
Variable Name
Data Location
Question Text / Label
Value Codes & Labels
Response of Interest (ROI)
Q12
43-43
Carried weapon (school)
1=0 days; 2=1 day; 3=2-3; 4=4-5; 5=6+
Variable Name
Data Location
Question Text / Label
Value Codes & Labels
Response of Interest (ROI)
Q31
62-62
Ever smoked a cigarette
1=Yes; 2=No
Variable Name
Data Location
Question Text / Label
Value Codes & Labels
Response of Interest (ROI)
Q56
87-87
Ever had sexual intercourse
1=Yes; 2=No
Variable Name
Data Location
Question Text / Label
Value Codes & Labels
Response of Interest (ROI)
QNOBESE
386-386
% with obesity
1=Yes; 2=No
Variable Name
Data Location
Question Text / Label
Value Codes & Labels
Response of Interest (ROI)
WEIGHT
388-397
Weight factor
Continuous decimal
Variable Name
Data Location
Question Text / Label
Value Codes & Labels
Response of Interest (ROI)
Q1
17-17
How old are you?
1=12 or younger; 2=13; 3=14; 4=15; 5=16; 6=17; 7=18 or older
IDATE
Date Interview Completed
Char
MMDDYYYY
Variable Name
Variable Label / Question Text
Type
Values & Labels
SEXVAR
Sex of Respondent
Num
1=Male; 2=Female
Variable Name
Variable Label / Question Text
Type
Values & Labels
_BMI5
Body Mass Index (BMI)
Num
Continuous (4 decimal places)
Variable Name
Variable Label / Question Text
Type
Values & Labels
_STSTR
Sample Design Stratification Variable
Num
Internal Code (used for variance)
Variable Name
Variable Label / Question Text
Type
Values & Labels
_RFHLTH
Adults with good or better health
Num
1=Good/Better; 2=Fair/Poor
Variable Name
Variable Label / Question Text
Type
Values & Labels
_STATE
State FIPS Code
Num
1=Alabama; 2=Alaska... 72=PR
Data in cBioPortal is organized into studies, each of which includes a study description file, clinical meta file, and a clinical data file. Before any study can be loaded into cBioPortal, all files must pass a validation phase to ensure they are following the correct content and formatting guidelines. Studies can be loaded incrementally for certain data types without re-uploading an entire study, and datasets that are subsets or combinations of existing studies can be represented as virtual studies to avoid duplication. For further information on how to load or download study data see the links section below.
cBioPortal does not enforce a standardized clinical ontology across studies for patient and sample attributes. Instead, study authors of the published studies determine the organization of their clinical data. Although cBioPortal does not require a universal clinical ontology across studies, several standardized conventions are used for genomic data to support harmonization and cross-study analyses.
OncoTree (cancer type taxonomy) — used as the hierarchical cancer classification system underlying cBioPortal. Each study specifies a type_of_cancer value, which maps to the portal's cancer type hierarchy; custom cancer type definitions can also be supplied when loading a study.
Reference genome assemblies (hg19/GRCh37, hg38/GRCh38) — genomic datasets specify the reference genome assembly used for sequence alignment and genomic coordinates through study metadata (e.g., reference_genome). Segmented copy-number datasets additionally specify a reference_genome_id field to ensure consistent interpretation of genomic locations.
TCGA Mutation Annotation Format (MAF) — provides the standardized file format and controlled Variant_Classification values (e.g., Missense_Mutation, Nonsense_Mutation, Splice_Site) used for importing somatic mutation data.
HGVS (Human Genome Variation Society) nomenclature — used for standardized reporting of sequence and protein variants, including amino acid changes associated with genomic mutations.
HUGO Gene Nomenclature Committee (HGNC) gene symbols — standardized gene symbols used in genomic datasets through the Hugo_Symbol field.
NCBI Entrez Gene identifiers — stable numeric gene identifiers used through the Entrez_Gene_Id field. Depending on the data type, cBioPortal accepts HGNC symbols, Entrez Gene identifiers, or both to support consistent gene mapping across studies.
For further information see the Links section below.
Demographic information is limited to just AGE and SEX/GENDER with native platform support. Other standard demographic fields are not enforced and vary by study/what the author chooses to include. There are no platform-wide demographic charactersitics published by cBioPortal. At the individual study level: demographic data including race and age IS displayed in Study View, but only when the contributing study authors included it in their clinical data file.
cBioPortal for Cancer Genomics. Overview [Internet]. New York (NY): Memorial Sloan Kettering Cancer Center; 2026 [cited 2026 Jul 15]. Available from: https://docs.cbioportal.org/user-guide/overview/
Cerami E, Gao J, Dogrusoz U, Gross BE, Sumer SO, Aksoy BA, Jacobsen A, Byrne CJ, Heuer ML, Larsson E, Antipin Y, Reva B, Goldberg AP, Sander C, Schultz N. The cBio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data. Cancer Discov. 2012 May;2(5):401-4. doi: 10.1158/2159-8290.CD-12-0095. Erratum in: Cancer Discov. 2012 Oct;2(10):960. PMID: 22588877; PMCID: PMC3956037.
Gao J, Aksoy BA, Dogrusoz U, Dresdner G, Gross B, Sumer SO, Sun Y, Jacobsen A, Sinha R, Larsson E, Cerami E, Sander C, Schultz N. Integrative analysis of complex cancer genomics and clinical profiles using the cBioPortal. Sci Signal. 2013 Apr 2;6(269):pl1. doi: 10.1126/scisignal.2004088. PMID: 23550210; PMCID: PMC4160307.
de Bruijn I, Kundra R, Mastrogiacomo B, Tran TN, Sikina L, Mazor T, Li X, Ochoa A, Zhao G, Lai B, Abeshouse A, Baiceanu D, Ciftci E, Dogrusoz U, Dufilie A, Erkoc Z, Garcia Lara E, Fu Z, Gross B, Haynes C, Heath A, Higgins D, Jagannathan P, Kalletla K, Kumari P, Lindsay J, Lisman A, Leenknegt B, Lukasse P, Madela D, Madupuri R, van Nierop P, Plantalech O, Quach J, Resnick AC, Rodenburg SYA, Satravada BA, Schaeffer F, Sheridan R, Singh J, Sirohi R, Sumer SO, van Hagen S, Wang A, Wilson M, Zhang H, Zhu K, Rusk N, Brown S, Lavery JA, Panageas KS, Rudolph JE, LeNoue-Newton ML, Warner JL, Guo X, Hunter-Zinck H, Yu TV, Pilai S, Nichols C, Gardos SM, Philip J; AACR Project GENIE BPC Core Team, AACR Project GENIE Consortium; Kehl KL, Riely GJ, Schrag D, Lee J, Fiandalo MV, Sweeney SM, Pugh TJ, Sander C, Cerami E, Gao J, Schultz N. Cancer Res. 2023 Dec 1;83(23):3861-3867. doi: 10.1158/0008-5472.CAN-23-0816. PMID: 37668528; PMCID: PMC10690089.
Derived clinical and behavioral metrics, such as Body Mass Index (BMI), which are identifiable in the dataset by a preceding underscore (ie. _BMI).
Specific internal codes including final weight for combined landline/cellular data (_LLCPWT) and stratum weight (_STRWT) used for complex sampling analysis.
Standardized variables for stratification (_STSTR) and clustering (_PSU) required for proper analysis in statistical software.
Standardized result codes (ie. completed interview, refusal, ineligible) established by the American Association of Public Opinion Research (AAPOR) to calculate response and cooperation rates.
Standardized categories for reporting age, race, ethnicity, marital status, education level, and home ownership.
U.S. Centers for Disease Control and Prevention. The BRFSS Data User Guide June 2013 [Internet]. Aug 15, 2013.
U.S. Centers for Disease Control and Prevention. BRFSS Overview 2024 [Internet]. 2025.
U.S. Centers for Disease Control and Prevention. 2024 BRFSS Survey Data and Documentation [Internet]. 2025.
Gundersen DA, ZuWallack RS, Dayton J, Echeverría SE, Delnevo CD. Assessing the Feasibility and Sample Quality of a National Random-digit Dialing Cellular Phone Survey of Young Adults. Am J Epidemiol. 2014 Jan 1;179(1):39–47. doi:10.1093/aje/kwt226 PubMed PMID: 24100957; PubMed Central PMCID: PMC3864711.
Hsia J, Zhao G, Town M, Ren J, Okoro CA, Pierannunzi C, et al. Comparisons of Estimates From the Behavioral Risk Factor Surveillance System and Other National Health Surveys, 2011−2016. American Journal of Preventive Medicine. 2020 Jun 1;58(6):e181–90. doi:10.1016/j.amepre.2020.01.025 PubMed PMID: 32444008.
Hu SS, Pierannunzi C, Balluz L. Integrating a Multimode Design Into a National Random-Digit–Dialed Telephone Survey. Prev Chronic Dis. 2011 Oct 15;8(6):A145. PubMed PMID: 22005638; PubMed Central PMCID: PMC3221584.
Li C, Balluz LS, Ford ES, Okoro CA, Zhao G, Pierannunzi C. . Preventive Medicine. 2012 Jun 1;54(6):381–7. doi:10.1016/j.ypmed.2012.04.003
Pierannunzi C, Hu SS, Balluz L. . BMC Med Res Methodol. 2013 Mar 24;13(1):49. doi:10.1186/1471-2288-13-49
N/A
Q7
33-38
Weight (kilograms)
Continuous (from pounds)
N/A
N/A
QN8
185-185
% did not always wear seat belt
1=Yes; 2=No
1 (Answers A,B,C,D to Q8)
Q9
40-40
Rode with drinking driver
1=0 times; 2=1 time; 3=2-3; 4=4-5; 5=6+
N/A
QN9
186-186
% rode with drinking driver
1=Yes; 2=No
1 (Answers B,C,D,E to Q9)
Q10
41-41
Drove after drinking
1=Did not drive; 2=No; 3=1 time; 4=2-3; 5=4-5; 6=6+
N/A
QN10
187-187
% drove after drinking
1=Yes; 2=No
1 (Answers C,D,E,F to Q10)
N/A
QN12
189-189
% carried weapon (school)
1=Yes; 2=No
1 (Answers B,C,D,E to Q12)
Q26
57-57
Sad or hopeless (2+ weeks)
1=Yes; 2=No
N/A
QN26
203-203
% felt sad or hopeless
1=Yes; 2=No
1 (Answer A to Q26)
Q27
58-58
Seriously considered suicide
1=Yes; 2=No
N/A
QN27
204-204
% considered suicide
1=Yes; 2=No
1 (Answer A to Q27)
N/A
QN31
208-208
% ever smoked
1=Yes; 2=No
1 (Answer A to Q31)
Q35
66-66
Ever used electronic vapor
1=Yes; 2=No
N/A
QN35
212-212
% ever used vapor product
1=Yes; 2=No
1 (Answer A to Q35)
Q43
73-73
Binge drinking (past 30 days)
1=0 days; 2=1; 3=2; 4=3-5; 5=6-9; 6=10-19; 7=20+
N/A
QN43
220-220
% currently binge drinking
1=Yes; 2=No
1 (Answers B,C,D,E,F,G to Q43)
Q46
77-77
Ever used marijuana
1=0 times; 2=1-2; 3=3-9; 4=10-19; 5=20-39; 6=40-99; 7=100+
N/A
QN46
223-223
% ever used marijuana
1=Yes; 2=No
1 (Answers B,C,D,E,F,G to Q46)
N/A
QN56
233-233
% ever had sex
1=Yes; 2=No
1 (Answer A to Q56)
Q61
92-92
Used condom last time
1=Never sex; 2=Yes; 3=No
N/A
QN61
238-238
% used condom last time
1=Yes; 2=No
1 (Answer B to Q61)
1 (>= 95th percentile BMI)
QNOWT
387-387
% overweight
1=Yes; 2=No
1 (>= 85th and < 95th % BMI)
RACEETH
414-415
Standardized Race/Ethnicity
1=Am Indian; 2=Asian; 3=Black; 5=White; 6=Hispanic
N/A
N/A
STRATUM
398-400
Stratum identifier
Continuous
N/A
PSU
401-406
Primary Sampling Unit
Continuous
N/A
_AGEG5TYP
Five-year age group (calculated)
Num
1=18-24; 2=25-29... 13=80+
_RACE
Calculated Race/Ethnicity Category
Num
1=White; 2=Black; 8=Hispanic
_BMI5CAT
BMI Category
Num
1=Underweight; 2=Normal; 3=Overweight; 4=Obese
_PSU
Primary Sampling Unit
Num
Internal Code (used for clustering)
_LLCPWT
Final Weight: Landline and Cell Phone
Num
Continuous (Statistical weight)
_STRWT
Stratum Weight
Num
Continuous (Statistical weight)
DISPCODE
Final Disposition Code
Num
110=Complete; 120=Partial
The DailyMed Full SPL Dataset refers to the downloadable full release of all drug labels in the United States. The dataset is hosted by the National Library of Medicine of the National Institute of Health (NIH). There are five subsets of drug labels available as full SPL releases: human prescription labels, human over-the-counter (OTC) labels, homeopathic labels, animal labels, and remaining labels.
Pharmaceutical companies are required to submit a standard product label (SPL) to the Food and Drug Administration (FDA). SPLs are typically encoded in XML format, standardized by the Health Level Seven International (HL7) schema to ensure the labels are organized and machine readable. This includes the tagging of various sections within the label using Logical Observation Identifiers Names and Codes (LOINC®). After submission to the FDA, labels are stored by the National Library of Medicine, which makes all SPLs available for download via DailyMed.
Accessing the Full SPL release through DailyMed is completely free and does not require an account. Downloading the largest of the datasets requires at least 16.38GB of space on the user’s device. To download, navigate to the DailyMed website, then click “Download Data” under the header NLM SPL Resources. The page will present the option to download full drug labels, indexing and risk evaluation and mitigation strategy (REMS) files, or mapping files that link an SPL’s set id to other relevant information and metadata. To view the full SPL release with labels, click on the first. The full releases will be at the bottom of this page in five parts.
A more common way of leveraging the dataset is through DailyMed’s API client, RESTful. Querying via the RESTful API requires no account, key, or authentication, and returns data in either XML or JSON format. Related endpoints allow lookups by NDC code (/ndcs), drug name (/drugnames), RxNorm concept (/rxcuis), active ingredient (/uniis), and pharmacologic class (/drugclasses).
The main way by which sections of an SPL are standardized is through Logical Observation Identifiers Names and Codes (LOINC®) code tagging. This allows for streamlined indexing into areas such as indications and usage, adverse reactions, dosage and administration, and other common sections.
AMIA Informatics Summit. (2013 Mar 18.). . Retrieved June 25, 2026.
Malec SA, Boyce RD. AMIA Jt Summits Transl Sci Proc. 2020 May 30;2020:403-412. PMID: 32477661; PMCID: PMC7233092.
The Unified Research data Sharing and Access (URSA) Initiative was launched in 2015 with the overall goal of making electronic health record (EHR) and other health data accessible and usable for research purposes across Rhode Island. This initiative is supported by Advance RI-CTR and coordinated by the Advance RI-CTR Biomedical Informatics, Bioinformatics, and Cyberinfrastructure Enhancement () Core with leadership and expertise provided by the Brown Center for Biomedical Informatics ().
The Advance RI-CTR BIBCE Core collaborates with Brown University's Office of Information Technology and the Division of Research, including Research Integrity and Research Agreements & Contracting, and health data partners across Rhode Island to:
Coordinate processes for health data sharing and secure access within and across institutions;
Develop the requisite legal, ethical, and technical infrastructure between Brown and health data sharing partners;
Establish cross-institutional governance, including standard policies, procedures, and protocols for appropriate sharing and use of health data; and,
Provide documentation and training for health data requests, access, and use.
Through the URSA Initiative, the BIBCE Core provides expertise and infrastructure for conducting research using large-scale health datasets. The BIBCE Core is available to help researchers navigate the process of identifying options and solutions for storing, managing, and analyzing data from health data partners or other data sources.
Complete Advance RI-CTR's Service Request Form to schedule a consultation with the BIBCE Core.

Visit other chapters in CODIAC for Health using the or menu in the upper left corner.
One row per NPI. Repeating groups use a _N suffix; positions reflect the CSV column order (330 columns total).
1
NPI
Non-primary practice locations for Type 1 & Type 2 NPIs. Join back to the main file on NPI.
Additional 'other names' for Type 2 (organization) NPIs beyond the one stored in the main file.
Health IT endpoints (Direct messaging, FHIR servers, HIE, etc.) associated with NPIs.
The Youth Risk Behavior Surveillance System (YRBSS) was established in 1990 by the U.S. Centers for Disease Control and Prevention (CDC) to monitor priority health risk behaviors and experiences that contribute to death, disability, and social problems among youth and young adults in the United States. The system monitors critical variables established during childhood and adolescence, including demographics, substance use, sexual behaviors, violence, as well as other student health behaviors and experiences. Data has been collected from over 5 million high school students through more than 2,300 separate surveys at national, state, territorial, tribal, and local levels. The YRBSS aims to determine the prevalence of health risk behaviors and assess trends over time. Published data, reports, questionnaires, and analysis tools are publicly accessible through the CDC YRBSS website. Data files are provided in Access and ASCII formats. All other data must be requested through the YRBSS Data Request Form.
The YRBSS was created using a three-stage cluster sample design to produce a representative sample of 9th through 12th grade students in both public and private high schools. The selection process involved identifying schools systematically with probability proportional to enrollment, followed by systematic equal probability sampling of classes within a required subject or specific period. The survey has been distributed every two years. Once data was collected, responses underwent automated data edits to ensure range validity, logical consistency, and biologically reasonable values for height, weight, and body mass index (BMI). Records that were overly incomplete or with suspicious response patterns were excluded from the final dataset. A weighting factor was applied to each usable record to adjust for nonresponse and oversampling, rendering generalizable data that matches nationwide population projections.
The core alphanumeric identifiers used to categorize each survey question and its corresponding clinical or behavioral concept.
A standardized system of calculated variables where "1" represents a specific "Response of Interest" (ROI) and "2" represents all other valid responses.
The specific reference data and sex-specific percentile limits used to define clinical classifications for BMI, Overweight, and Obesity.
A nominal categorical classification system (values 1-8) that standardizes the reporting of American Indian/Alaskan Native, Asian, Black, Native Hawaiian/Pacific Islander, White, and Hispanic/Latino identities.
Standardized response options used to identify and categorize sexual minority youth (ie. gay, lesbian, or bisexual).
Standardized variables required for statistical software to account for the three-stage cluster sample design, ensuring that prevalence estimates accurately represent the national youth population.
Standardized, shortened versions of response options (under 40 characters) used specifically as value labels in datasets and programs like SAS and SPSS to maintain technical clarity.
U.S. Centers for Disease Control and Prevention. [Internet]. 2024. Site Index.
U.S. Centers for Disease Control and Prevention. [Internet]. 2023 Sep.
Brener ND, Mpofu JJ, Krause KH, Everett Jones S, Thornton JE, Myles Z, et al. . MMWR Suppl. 2024 Oct 10;73(4):1–12. doi:10.15585/mmwr.su7304a1 PubMed PMID: 39378301; PubMed Central PMCID: PMC11559678.
Jorge V. Verlenden P, Ari Fodeman P, Natalie Wilkins P, Sherry Everett Jones P, Shamia Moore MPH, Kelly Cornett MS, et al. . MMWR Suppl. 2024;73. doi:10.15585/mmwr.su7304a9
Swedo EA, Niolon PH, Anderson KN, Li J, Brener N, Mpofu J, et al. . Pediatrics. 2024 Nov 1;154(5):e2024066633. doi:10.1542/peds.2024-066633 PubMed PMID: 39463258; PubMed Central PMCID: PMC11756604.
3
Provider Secondary Practice Location Address – Address Line 2
Street address line 2 (suite, floor, etc.).
4
Provider Secondary Practice Location Address – City Name
City of the secondary location.
5
Provider Secondary Practice Location Address – State Name
State/territory code of the secondary location.
6
Provider Secondary Practice Location Address – Postal Code
ZIP/postal code (may include ZIP+4).
7
Provider Secondary Practice Location Address – Country Code (If outside U.S.)
ISO country code; populated only when outside the U.S.
8
Provider Secondary Practice Location Address – Telephone Number
Phone number at the secondary location.
9
Provider Secondary Practice Location Address – Telephone Extension
Phone extension at the secondary location.
10
Provider Practice Location Address – Fax Number
Fax number at the secondary location.
3
Provider Other Organization Name Type Code
Type of the alternate name (3 = DBA, 4 = Former Legal Business Name, 5 = Other Name).
3
Endpoint Type Description
Full descriptive name of the endpoint type.
4
Endpoint
The actual endpoint address (Direct email, FHIR/web-services URL, etc.).
5
Affiliation
Whether the endpoint is affiliated with another organization (Y/N).
6
Endpoint Description
Free-text description of the endpoint and its purpose.
7
Affiliation Legal Business Name
Legal Business Name of the affiliated organization (when Affiliation = Y).
8
Use Code
Short code for the endpoint's use (e.g., DIRECTMESSAGING).
9
Use Description
Full description of the endpoint's use.
10
Other Use Description
Free-text description when Use Code is 'Other'.
11
Content Type
Short code for content exchanged (e.g., CCD).
12
Content Description
Full description of the content type.
13
Other Content Description
Free-text description when Content Type is 'Other'.
14
Affiliation Address Line One
Street address line 1 of the affiliated organization.
15
Affiliation Address Line Two
Street address line 2 of the affiliated organization.
16
Affiliation Address City
City of the affiliated organization.
17
Affiliation Address State
State of the affiliated organization.
18
Affiliation Address Country
ISO country code of the affiliated organization.
19
Affiliation Address Postal Code
Postal code of the affiliated organization.
Unique 10-digit National Provider Identifier assigned by CMS. Primary key of the file.
—
2
Entity Type Code
1 = Individual, 2 = Organization. Blank only for deactivated records.
Entity Type
3
Replacement NPI
If this NPI was deactivated and replaced, the replacement NPI. Blank otherwise.
—
4
Employer Identification Number (EIN)
IRS EIN of the provider. SUPPRESSED in the public file for privacy.
—
5
Provider Organization Name (Legal Business Name)
Official legal business name of the organization. V.2 expanded this field's length.
—
6
Provider Last Name (Legal Name)
Legal last (family) name of the individual provider.
—
7
Provider First Name
Legal first (given) name. V.2 expanded this field's length.
—
8
Provider Middle Name
Middle name or initial of the individual provider.
—
9
Provider Name Prefix Text
Name prefix (e.g., Dr., Mr., Ms.).
Name Prefix
10
Provider Name Suffix Text
Name suffix (e.g., Jr., Sr., III).
Name Suffix
11
Provider Credential Text
Professional credential(s), e.g., MD, DO, NP, PA.
—
12
Provider Other Organization Name
Single alternate organization name (DBA, former name, etc.).
—
13
Provider Other Organization Name Type Code
Type of the alternate organization name (3, 4, 5).
Other Name Type
14
Provider Other Last Name (Legal Name)
Alternate last name of an individual provider.
—
15
Provider Other First Name
Alternate first name of an individual provider.
—
16
Provider Other Middle Name
Alternate middle name of an individual provider.
—
17
Provider Other Name Prefix Text
Name prefix for the alternate individual name.
Name Prefix
18
Provider Other Name Suffix Text
Name suffix for the alternate individual name.
Name Suffix
19
Provider Other Credential Text
Professional credential for the alternate individual name.
—
20
Provider Other Last Name Type Code
Type of the alternate individual name (1, 2, 5).
Other Name Type
21
Provider First Line Business Mailing Address
Street address line 1 of the mailing address.
—
22
Provider Second Line Business Mailing Address
Street address line 2 (suite, floor, PO box) of the mailing address.
—
23
Provider Business Mailing Address City Name
City of the mailing address.
—
24
Provider Business Mailing Address State Name
State/territory code of the mailing address.
State Codes
25
Provider Business Mailing Address Postal Code
ZIP/postal code (USPS 9-digit; may include ZIP+4).
—
26
Provider Business Mailing Address Country Code (If outside U.S.)
ISO country code; populated only when outside the U.S.
Country Codes
27
Provider Business Mailing Address Telephone Number
Phone number for the mailing address.
—
28
Provider Business Mailing Address Fax Number
Fax number for the mailing address.
—
29
Provider First Line Business Practice Location Address
Street address line 1 of the primary practice location.
—
30
Provider Second Line Business Practice Location Address
Street address line 2 of the primary practice location.
—
31
Provider Business Practice Location Address City Name
City of the primary practice location.
—
32
Provider Business Practice Location Address State Name
State/territory code of the primary practice location.
State Codes
33
Provider Business Practice Location Address Postal Code
ZIP/postal code of the primary practice location.
—
34
Provider Business Practice Location Address Country Code (If outside U.S.)
ISO country code; populated only when outside the U.S.
Country Codes
35
Provider Business Practice Location Address Telephone Number
Phone number at the primary practice location.
—
36
Provider Business Practice Location Address Fax Number
Fax number at the primary practice location.
—
37
Provider Enumeration Date
Date the NPI was first assigned (MM/DD/YYYY).
—
38
Last Update Date
Most recent date any field on the record changed (MM/DD/YYYY).
—
39
NPI Deactivation Reason Code
Reason the NPI was deactivated, if applicable.
Deactivation Reason
40
NPI Deactivation Date
Date the NPI was deactivated. Blank for active providers.
—
41
NPI Reactivation Date
Date the NPI was reactivated after a prior deactivation.
—
42
Provider Sex Code
Sex of an individual provider (formerly 'Provider Gender Code'). Not used for organizations.
Sex
43
Authorized Official Last Name
Last name of the official authorized to act for the organization.
—
44
Authorized Official First Name
First name of the authorized official.
—
45
Authorized Official Middle Name
Middle name or initial of the authorized official.
—
46
Authorized Official Title or Position
Job title/position of the authorized official (e.g., CEO, CFO).
—
47
Authorized Official Telephone Number
Phone number of the authorized official.
—
48–107
Healthcare Provider Taxonomy Code_N
NUCC taxonomy code classifying provider type/specialization (e.g., 207Q00000X). Interleaved per occurrence with the three fields below.
Taxonomy (NUCC)
48–107
Provider License Number_N
State-issued license number for the taxonomy at the same position.
—
48–107
Provider License Number State Code_N
State that issued the license at the same position.
State Codes
48–107
Healthcare Provider Primary Taxonomy Switch_N
Y if this is the provider's primary taxonomy (only one Y per NPI).
Primary Taxonomy Switch
108–307
Other Provider Identifier_N
Legacy/plan-specific identifier (e.g., Medicaid ID). Masked: SSN=$$$$$$$$$, ITIN=*********, EIN=========.
—
108–307
Other Provider Identifier Type Code_N
Type/issuer system of the identifier at the same position.
Other Provider ID Type
108–307
Other Provider Identifier State_N
State code for the identifier (used when issued by a state Medicaid plan).
State Codes
108–307
Other Provider Identifier Issuer_N
Name of the plan/organization that issued the identifier.
—
308
Is Sole Proprietor
Whether the individual operates as a sole proprietor (Y/N/X).
Sole Proprietor
309
Is Organization Subpart
Whether the organization is a subpart of a larger parent (Y/N/X).
Subpart
310
Parent Organization LBN
Legal Business Name of the parent organization. SUPPRESSED in the public file.
—
311
Parent Organization TIN
Tax ID of the parent organization. SUPPRESSED in the public file.
—
312
Authorized Official Name Prefix Text
Name prefix of the authorized official.
Name Prefix
313
Authorized Official Name Suffix Text
Name suffix of the authorized official.
Name Suffix
314
Authorized Official Credential Text
Professional credential of the authorized official.
—
315–329
Healthcare Provider Taxonomy Group_N
Grouping code for the taxonomy at the same position (e.g., 193200000X).
Group Taxonomy
330
Certification Date
Date the provider certified the application information is accurate.
—
1
NPI
NPI of the provider for this secondary location. Join key to the main file.
2
Provider Secondary Practice Location Address – Address Line 1
Street address line 1 of the secondary location.
1
NPI
NPI of the organizational provider this name belongs to. Join key to the main file.
2
Provider Other Organization Name
The alternate organization name (DBA, former name, etc.).
1
NPI
NPI of the provider this endpoint belongs to. Join key to the main file.
2
Endpoint Type
Short code/label for the endpoint type (e.g., DIRECT, FHIR).
90598
26.0
40-59
89353
25.7
60+
106257
30.5
Gender
F
174621
50.2
M
173499
49.8
Race
Asian
12244
3.5
Black
27311
7.8
Hawaiian
4672
1.3
Native
3404
1.0
Other
4602
1.3
White
295887
85.0
Ethnicity
Hispanic
49410
14.2
Nonhispanic
298710
85.8
18-39
318
28.3
40-59
277
24.6
60+
339
30.2
Gender
F
581
51.7
M
543
48.3
Race
Asian
33
2.9
Black
88
7.8
Hawaiian
19
1.7
Native
12
1.1
Other
11
1.0
White
961
85.5
Ethnicity
Hispanic
159
14.1
Nonhispanic
965
85.9
Age
< 18
61912
17.8
18-39
The Pediatric Psychiatry Emergency Services (PediPES) dataset consists of electronic health record (EHR) data for patients younger than 18 years old evaluated in the emergency department at Hasbro Children's Emergency Department in Providence, Rhode Island for psychiatric concerns. The dataset includes visit-level demographic, socioeconomic, geographic, and clinical variables, providing a comprehensive view of the pediatric mental health population served by the hospital. The dataset also captures a diverse range of mental health presentations, including suicidal ideation, suicide attempts, self-injurious behaviors, mood disorders, anxiety disorders, psychotic disorders, and behavioral concerns. This dataset is part of LEAPPES (Leveraging EHRs and AI for Pediatric Psychiatry Emergency Services), a multi-institutional and multi-disciplinary collaboration and offers valuable insight into the mental health needs, care patterns, and outcomes of those under the age of 18 presenting for emergency psychiatric evaluation.
The dataset was generated from electronic health record (EHR) data collected during routine clinical care. Variables were derived from both structured and unstructured data sources. Structured variables were obtained from standardized drop-down menus and other discrete fields within the EHR, while additional clinical information was extracted from free-text documentation using basic natural language processing (NLP) techniques. This approach enabled the inclusion of demographic, socioeconomic, geographic, and clinical variables that may not have been consistently captured in structured fields alone.
Mental health and medical diagnoses are documented using the International Classification of Diseases, Clinical Modification (ICD-10-CM)
Final (Billing) Diagnoses ICD-10 Codes - F Codes: Mental, Behavioral, and Neurodevelopmental Disorders; These are psychiatric and behavioral health diagnoses, which are particularly relevant in a pediatric psychiatric emergency setting.
Final (Billing) Diagnoses ICD-10 Codes - Z Codes: Factors Influencing Health Status and Contact with Health Services; These codes describe social circumstances, healthcare encounters, or reasons for seeking care rather than diseases.
Final (Billing) Diagnoses ICD-10 Codes - Other Codes: All non-F and non-Z ICD-10 codes, representing medical conditions, injuries, poisonings, symptoms, and other health problems.
Suicide risk is assessed using the Columbia-Suicide Severity Rating Scale (C-SSRS): standardized tool that evaluates an individual's risk of suicide by assessing the severity and type of suicidal thoughts and behaviors.
Boarding: A patient remains in the emergency department or another temporary hospital location after being evaluated and determined to need a higher level of psychiatric care, but appropriate placement is not immediately available.
Inpatient: The patient is admitted to a psychiatric hospital or psychiatric unit for 24-hour supervision, stabilization, and intensive treatment.
Outpatient: The patient is discharged from the emergency psychiatric service and receives treatment while continuing to live at home.
Brown KA, Donise KR, Cancilliere MK, Aluthge DP, Chen ES. . AMIA Annu Symp Proc. 2024 Jan 11;2023:864-873. PMID: 38222397; PMCID: PMC10785882.
Kim W, Donise KR, Brown KA, Cancilliere MK, Chen ES. AMIA Jt Summits Transl Sci Proc. 2024 May 31;2024:565-574. PMID: 38827092; PMCID: PMC11141824.
Partial programs: A structured, intensive treatment program that provides several hours of psychiatric care and therapy during the day while allowing the patient to return home each evening. Partial programs serve as an intermediate level of care between inpatient hospitalization and traditional outpatient treatment.
The All of Us dataset is a large, longitudinal U.S. health research resource that integrates clinical, behavioral, genomic, and wearable-device data from a diverse participant population.
The dataset includes six major data types collected from participants:
Electronic Health Records (EHRs): Clinical records from participating healthcare organizations (OMOP Common Data Model). It includes diagnoses, medications, procedures, laboratory results, and clinical visits. (393k+ participants)
Participant Surveys: Self-reported data covering demographics, lifestyle, social determinants of health, family and personal health history, health care access and utilization, behavioral health and emotional health history (633k+ participants).
Physical Measurements: Standardized measurements that include height, weight, waist and hip circumference, blood pressure, and heart rate. Collectively from EHR data as well as self-reported through surveys (509k+ participants).
Biosamples: Blood and urine samples from 568,000+ participants, supporting genomic and other laboratory analyses.
Genomics: Whole genome sequences and genotyping arrays from 447,000+ participants, available in the Controlled Tier.
Wearable Devices (Digital Health): Biometric data from Fitbit devices including heart rate, physical activity, sleep patterns, device data. (59k+ participants)
Data are organized into three access tiers based on sensitivity and researcher eligibility:
Public Tier: Aggregate, de-identified data available to everyone via the Data Browser and Data Snapshots. Doesn’t require registration.
Registered Tier: De-identified individual-level data including EHRs, wearables, surveys, and physical measurements. Accessible only to approved researchers on the Researcher Workbench.
Controlled Tier: Contains individual-level genomic data (whole genome sequences and genotyping arrays). Requires additional access approval beyond the Registered Tier.
Data flow through several stages before reaching researchers. Participants contribute data directly via participant portals and biosample donation. Healthcare Provider Organization (HPO) staff submit local EHR data through the HealthPro application. Genomics partners (genetic counseling resources, clinical validation laboratories, and genome centers) process biosamples for sequencing and genotyping. All raw data enters a central Raw Data Repository, then passes through a curation pipeline before being organized into the Curated Data Repository, from which the Public, Registered, and Controlled tiers are made available to researchers through the All of Us Research Hub.
Electronic health records are standardized using the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM), maintained by the international Observational Health Data Sciences and Informatics (OHDSI) collaborative. The OMOP CDM organizes data around a central Person table linked to standardized tables including Visit Occurrence, Procedure Occurrence, Drug Exposure, Condition Occurrence, Measurement, and Observation. Survey questionnaires and physical measurements are mapped into the Observation and Measurement tables respectively, creating a unified structure across all data sources.
The dataset uses the OMOP CDM and its associated standardized vocabularies.
SNOMED CT
LOINC
RxNorm
ICD-10-CM / ICD-9-CM
Access to Registered and Controlled Tier data requires agreement to the Data User Code of Conduct (DUCC), which includes the following key requirements:
Participants may never be identified or re-identified.
Data may be used for biomedical or health research purposes only.
All of Us data cannot be linked with external datasets at the individual participant level.
All of Us Research Program Investigators; Denny JC, Rutter JL, Goldstein DB, Philippakis A, Smoller JW, Jenkins G, Dishman E. (2019). . N Engl J Med. 2019 Aug 15;381(7):668-676. doi: 10.1056/NEJMsr1809937. PMID: 31412182; PMCID: PMC8291101.
Mayo KR, Basford MA, Carroll RJ, Dillon M, Fullen H, Leung J, et al. . Annu Rev Biomed Data Sci. 2023 Aug 10;6:443-464. doi: 10.1146/annurev-biodatasci-122120-104825. PMID: 37561600.
Ginsburg GS, Denny JC, Schully SD. Sci Transl Med. 2023 Dec 13;15(726):eade9214. doi: 10.1126/scitranslmed.ade9214. Epub 2023 Dec 13. PMID: 38091411; PMCID: PMC12374748.
CPT-4 / HCPCS
NDC (National Drug Code)
CVX
PPI (Participant Provided Information): Custom All of Us vocabulary for survey responses not mappable to existing standards.
Published statistics must reflect a minimum of 20 participants per cell (aka the ’minimum count rule’). Public Tier applies this same rounding convention.
The program must be notified at least two weeks before any publication or presentation using the data.
Final manuscripts must be submitted to PubMed Central immediately upon publication with no embargo period.
Each Workspace must include an accurate and meaningful description of the research, which becomes publicly viewable in the Research Projects Directory.
National Institutes of Health. (n.d.). .
All of Us Research Hub. (n.d.). . .
All of Us Research Hub. (n.d.). .
The Surveillance, Epidemiology, and End Results (SEER) Program provides information on cancer statistics in an effort to reduce the cancer burden among the U.S. population. SEER collects cancer incidence data from population-based cancer registries covering approximately 47.9 percent of the U.S. population. The SEER datasets include data on patient demographics, primary tumor site, tumor morphology, stage at diagnosis, and first course of treatment. [1]
SEER offers the following datasets:
SEER Research Data
Register with any valid email
Excludes geography, month and year of diagnosis, and other demographic fields
SEER Research Plus and NCCR Data
Requires user authentication through eRA Commons or an HHS account.
Includes geography, month, and year of diagnosis, other demographic fields
Visit for a deeper comparison of the datasets. Learn more about SEER and SEER Datasets on the .
National Cancer Institute. (n.d.).. Retrieved May 1, 2024.
2
47961
0.6679
7
7762
0.1081
120101
3576
0.0498
530101
3561
0.0496
370101
2794
0.0389
260101
2453
0.0342
130101
2382
0.0332
180101
2212
0.0308
1
59763
0.8322
4
7735
0.1077
2
4067
0.0566
7
564
0.0079
9
259
0.0036
4
1122412
15.6294
5
1097142
15.2776
3
205142
2.8566
2
104731
1.4584
9
16874
0.235
1
6201
0.0863
8
31
0.0004
9999999999999 9 29 9 9 99999999 99 9
5
0.0001
1
762633
10.6196
0
297044
4.1363
9
6187
0.0862
7
5021
0.0699
4
4027
0.0561
3
1991
0.0277
5
1619
0.0225
21111111112
633
0.0088
3
1411888
19.6604
1
784373
10.9223
4
605305
8.4288
5
231289
3.2207
7
9002
0.1254
9
6852
0.0954
38888 11 522 2132
700
0.0097
28888 11 522 2132
421
0.0059
2
88835
1.237
7
2106
0.0293
9
1509
0.021
1
168134
2.3412
4
7346
0.1023
2
3373
0.047
7
749
0.0104
9
194
0.0027
7
298705
4.1594
6
276241
3.8466
5
215407
2.9995
21
202120
2.8145
4
178366
2.4837
99
168227
2.3425
77
145925
2.032
3
143969
2.0048
2
425445
5.9243
3
234429
3.2644
4
215531
3.0012
7
43925
0.6117
8
30124
0.4195
9
4888
0.0681
1
1818420
25.3213
7
19030
0.265
9
6436
0.0896
9 99999999999999999999 9099999999999999999999 9 29 9 9 99999999
4
0.0001
1107 1 2021012032032031011 3115312188841072502006010172 1 0114210
2
0
2 10120355520220120422 41912588856771351355040171322 11111122031422
2
0
2 1021012032023011021 3243712111841051371255040192 22 11112122551422
2
0
9 99999999999999999999 90999999999999999999990799 1 999999
2
0
264883
3.6885
12
104228
1.4514
72
78653
1.0952
60
43002
0.5988
62
41778
0.5818
65
40228
0.5602
55
39915
0.5558
50
39649
0.5521
58
39535
0.5505
<missing>
5830436
81.1881
1
376807
5.247
<missing>
6705707
93.3761
3
403299
5.6159
<missing>
2787825
38.8201
6
1532501
21.3399
<missing>
4388211
61.1053
2
1098353
15.2944
<missing>
1585951
22.0842
2
1468394
20.4472
<missing>
6303754
87.779
1
785190
10.9337
<missing>
6748789
93.976
3
252809
3.5203
<missing>
4275582
59.5369
8
530471
7.3867
<missing>
3230757
44.9879
1
2996295
41.723
<missing>
2554029
35.5645
2
2411430
33.5789
Unique Patient Count
364,627
Gender
364,627
Female
191,984
Male
172,643
Anchor Age Distribution (nonnormal)
48(29-65)
Living Status
Alive
326,326
Dead
38301
Full hospital admission count
546,028
Unique Patient Count
223,452
Unique Gender Counts
Stays Count
94,458
Unique patients count
65,366
Unique Gender Counts
Full ED visit count
425,087
Unique Patient count
205,504
Led to hospital admissions
203,016
Female
117,736
Male
105,716
Age at admission distribution (non-normal)
61(45-74)
Living Status
Alive
186,570
Dead
36,882
Condensed Race/Ethnicity
White
147,029
Black
28,972
Unknown/Other
23,448
Hispanic/Latino/Portuguese
13,234
Asian American and Pacific Islander
10,496
Multiple race/ethnicity
273
Unique Insurance Type Counts
Medicare
88674
Private
83068
Medicaid
38257
Other
7003
No charge
147
Expanded Race/Ethnicity distribution counts
WHITE
137,946
BLACK/AFRICAN AMERICAN
23,404
UNKNOWN
9552
OTHER
9391
WHITE - OTHER EUROPEAN
6548
ASIAN
4182
HISPANIC/LATINO - PUERTO RICAN
3863
ASIAN - CHINESE
3526
HISPANIC OR LATINO
2618
HISPANIC/LATINO - DOMINICAN
2520
UNABLE TO OBTAIN
2468
BLACK/CAPE VERDEAN
2357
WHITE - RUSSIAN
2299
BLACK/CARIBBEAN ISLAND
1580
BLACK/AFRICAN
1522
PATIENT DECLINED TO ANSWER
998
ASIAN - SOUTH EAST ASIAN
869
ASIAN - ASIAN INDIAN
826
WHITE - EASTERN EUROPEAN
813
HISPANIC/LATINO - GUATEMALAN
780
WHITE - BRAZILIAN
739
PORTUGUESE
738
HISPANIC/LATINO - SALVADORAN
579
AMERICAN INDIAN/ALASKA NATIVE
481
HISPANIC/LATINO - MEXICAN
470
HISPANIC/LATINO - COLUMBIAN
433
SOUTH AMERICAN
369
ASIAN - KOREAN
349
HISPANIC/LATINO - HONDURAN
288
NATIVE HAWAIIAN OR OTHER PACIFIC ISLANDER
257
HISPANIC/LATINO - CUBAN
240
HISPANIC/LATINO - CENTRAL AMERICAN
227
MULTIPLE RACE/ETHNICITY
220
Unique Language Counts
English
202713
Spanish
6783
Chinese
3369
Russian
2415
Kabuverdianu
1750
Portuguese
1452
Haitian
1003
Other
608
Vietnamese
494
Arabic
303
Italian
302
Modern Greek (1453-)
277
Persian
166
Korean
160
American Sign Language
152
Thai
124
Polish
123
Amharic
104
Khmer
103
Hindi
95
Japanese
92
French
82
Bengali
67
Armenian
45
Somali
36
Female
28,646
Male
36,720
Age at ICU admission distribution (nonnormal)
66(55-77)
Length of stay distribution (normal)
1.97(1.09-3.86)
First Care Unit
Medical Intensive Care Unit (MICU)
20703
Medical/Surgical Intensive Care Unit (MICU/SICU)
15449
Cardiac Vascular Intensive Care Unit (CVICU)
14771
Surgical Intensive Care Unit (SICU)
13009
Coronary Care Unit (CCU)
10775
Trauma SICU (TSICU)
10474
Neuro Intermediate
5776
Neuro Surgical Intensive Care Unit (Neuro SICU)
1751
Neuro Stepdown
1421
Surgery/Vascular/Intermediate
145
PACU
122
Intensive Care Unit (ICU)
33
Medicine
16
Surgery/Trauma
10
Medicine/Cardiology Intermediate
1
Med/Surg
1
Neurology
1
Unique Gender Counts
205,504
Female
109,533
Male
95,971
Age at ED visit distribution (normal)
54(35-69)
Arrival Transport
WALK IN
251849
AMBULANCE
155752
UNKNOWN
15352
OTHER
1266
HELICOPTER
868
Disposition
HOME
241632
ADMITTED
158010
TRANSFER
7025
LEFT WITHOUT BEING SEEN
6,155
ELOPED
5,710
OTHER
4,297
LEFT AGAINST MEDICAL ADVICE
1,881
EXPIRED
377
2001
13601
2003
15214
2005
13917
2007
14041
2011
15425
2013
13583
2017
14765
2019
13677
2021
17232
2023
20103
col_222_222
219616
98.64
col_216_216
219558
98.61
col_213_213
219316
98.51
col_364_364
219198
98.45
col_365_365
219153
98.43
col_239_239
219153
98.43
QNSHPARG
219105
98.41
col_238_238
218733
98.24
QNOTHH
218188
98
col_271_271
217150
97.53
col_126_126
213511
95.9
RACEORIG
212726
95.55
QNIUDIMP
212548
95.47
col_282_282
212059
95.25
col_136_136
212059
95.25
col_129_129
211854
95.15
col_138_138
211847
95.15
col_284_284
211847
95.15
col_135_135
211792
95.13
col_281_281
211792
95.13
QNSPDRK3
211754
95.11
col_385_385
211717
95.09
col_283_283
211696
95.08
2007.030569
10.384386
1991
1997
2005
2017
2023
Q1
201865
5.103777
1.243429
1
4
5
6
7
Q7
210300
32.700481
35.218022
1
1
4
63.5
180.99
Q38
211784
1.504944
1.234843
1
1
1
1
7
Q45
203077
2.045323
1.858696
1
1
1
2
8
PSU
220993
227508.118
222704.8756
1
37153
160050
371830
921211
7
30584
3
21452
NaN
20776
1
371
2
287
2
45740
NaN
38121
5
277
3
28507
6
26343
4
14054
7
2385
8
1647
5
986
NaN
12426
3
12027
2
9784
5
9319
1
7803
1.68
5566
B
4964
4
7989
5
7904
1.68
7686
1.63
7486
1.7
7481
1.73
7044
1.6
6954
NaN
9219
54.43
4428
58.97
4389
63.5
3965
68.04
3872
61.24
3613
56.7
3482
5
20049
4
11023
3
10223
6
5848
NaN
40759
4
26813
3
18591
6
833
1991
12272
1993
16296
1995
10904
1997
16263
1999
column
missing_count
missing_percent
SITE
222642
100
col_211_211
222119
year
5
51555
6
51323
4
46293
1
93503
2
91077
NaN
38062
4
46371
3
46302
1
45831
2
67966
1
41242
NaN
39512
E
58300
C
19353
4
15900
1
24558
NaN
22291
2
8056
1
65002
2
23944
3
13398
1
69609
2
58867
NaN
47023
1
47229
5
45914
2
42503
15349
99.77
222642
The table below lists the most frequently occurring ICD-10-CM diagnosis codes recorded in the diagnoses_icd table of the MIMIC-IV Hosp module, ranked by number of unique patients.
1990
81557
1991
87846
1992
96213
1993
102263
1994
Count
1585906
Mean
1464.934
Std
778.4
Column
Missing Count
Missing Percent
BPHIGH4
6748789
93.98
BPHIGH
6723517
YEAR
7181394
2010.506928
2020
2024
NaN
4852521
22
264883
12
104228
NaN
4388211
2
1098353
1
762633
NaN
1585951
2
1468394
3
1411888
NaN
6748789
3
252809
1
168134
NaN
6723517
120101
3576
530101
3561
NaN
5830436
1
376807
2
47961
NaN
2554029
2
2411430
1
1818420
NaN
6705707
3
403299
1
59763
NaN
2787825
6
1532501
4
1122412
NaN
4275582
8
530471
7
298705
NaN
5886236
1
287603
2
43660
I10
Essential (primary) hypertension
190,198
105853
1995
113934
1996
124085
1997
135582
1998
149342
1999
159989
2000
184450
2001
212510
2002
247964
2007
430912
2008
414509
2009
432607
2010
451075
2011
506467
2012
475687
2013
491773
2020
401958
2021
438693
2022
445132
2023
433323
2024
457670
Min
450
25%
550
50%
1696
75%
2366
Max
2366
93.62
DIABETE3
6705707
93.38
DIABETES
6447493
89.78
HLTHPLN1
6303754
87.78
CHECKUP
6229092
86.74
CHOLCHK
5886236
81.97
BLOODCHO
5830436
81.19
RAW_RECORD
5595488
77.92
RAW_RECORD_LENGTH
5595488
77.92
AGE
4852521
67.57
SEX
4388211
61.11
INCOME2
4275582
59.54
CHECKUP1
3230757
44.99
EDUCA
2787825
38.82
72
78653
60
43002
0
297044
9
6187
1
784373
4
605305
4
7346
2
3373
370101
2794
260101
2453
7
7762
120101
3576
7
19030
9
6436
4
7735
2
4067
5
1097142
3
205142
6
276241
5
215407
3
27354
4
14108
91,186
E785
Hyperlipidemia, unspecified
151,863
71,607
K219
Gastro-esophageal reflux disease without esophagitis
104,785
49,543
Z87891
Personal history of nicotine dependence
94,511
46,243
F329
Major depressive disorder, single episode, unspecified
80,240
40,186
N179
Acute kidney failure, unspecified
65,020
38,389
I2510
Atherosclerotic heart disease of native coronary artery without angina pectoris
89,794
37,934
F419
Anxiety disorder, unspecified
61,310
35,075
E119
Type 2 diabetes mellitus without complications
69,343
32,249
D649
Anemia, unspecified
45,532
31,640
I4891
Unspecified atrial fibrillation
58,092
28,276
Z7901
Long term (current) use of anticoagulants
53,592
25,499
N390
Urinary tract infection, site not specified
37,083
25,402
D62
Acute posthemorrhagic anemia
29,544
24,328
E039
Hypothyroidism, unspecified
56,521
24,305
Z20822
Contact with and (suspected) exposure to COVID-19
33,113
23,629
E669
Obesity, unspecified
36,221
22,020
Z66
Do not resuscitate
29,429
21,712
NoDx
37,659
21,685
E871
Hypo-osmolality and hyponatremia
29,893
20,624
E872
Acidosis
25,530
19,921
Y929
Unspecified place or not applicable
24,287
19,660
I129
Hypertensive chronic kidney disease with stage 1 through stage 4 chronic kidney disease, or unspecified chronic kidney disease
41,811
19,507
G4733
Obstructive sleep apnea (adult) (pediatric)
41,489
18,829
Z794
Long term (current) use of insulin
47,242
18,102
I509
Heart failure, unspecified
40,557
17,746
N189
Chronic kidney disease, unspecified
33,252
17,333
J449
Chronic obstructive pulmonary disease, unspecified
36,287
17,278
F17200
Nicotine dependence, unspecified, uncomplicated
29,027
17,260
K5900
Constipation, unspecified
22,403
16,685
J189
Pneumonia, unspecified organism
20,588
16,409
E860
Dehydration
19,470
15,574
D696
Thrombocytopenia, unspecified
20,148
14,902
E780
Pure hypercholesterolemia
22,301
14,817
Z7902
Long term (current) use of antithrombotics/antiplatelets
22,059
14,266
G8929
Other chronic pain
26,416
14,241
F17210
Nicotine dependence, cigarettes, uncomplicated
24,106
14,207
I252
Old myocardial infarction
30,455
14,033
Z370
Single live birth
19,540
13,882
D72829
Elevated white blood cell count, unspecified
16,100
13,638
Z23
Encounter for immunization
14,437
13,392
Z7982
Long term (current) use of aspirin
18,257
13,285
Z8673
Personal history of transient ischemic attack (TIA), and cerebral infarction without residual deficits
25,723
13,211
E876
Hypokalemia
16,738
12,969
R079
Chest pain, unspecified
16,132
12,917
F1010
Alcohol abuse, uncomplicated
20,250
12,780
R001
Bradycardia, unspecified
15,518
12,719
I959
Hypotension, unspecified
14,314
12,467
Y838
Other surgical procedures as the cause of abnormal reaction of the patient, or of later complication, without mention of misadventure at the time of the procedure
14,567
12,425
A419
Sepsis, unspecified organism
15,496
12,000
E875
Hyperkalemia
17,214
11,896
N400
Benign prostatic hyperplasia without lower urinary tract symptoms
22,597
11,873
R197
Diarrhea, unspecified
14,991
11,730
G4700
Insomnia, unspecified
17,324
11,730
J45909
Unspecified asthma, uncomplicated
20,679
11,468
Z515
Encounter for palliative care
12,883
11,378
J9601
Acute respiratory failure with hypoxia
13,339
11,292
D509
Iron deficiency anemia, unspecified
17,448
11,275
M1990
Unspecified osteoarthritis, unspecified site
16,880
10,826
R339
Retention of urine, unspecified
13,112
10,722
M810
Age-related osteoporosis without current pathological fracture
21,191
10,711
R0789
Other chest pain
13,580
10,639
R0902
Hypoxemia
12,387
10,591
Z86718
Personal history of other venous thrombosis and embolism
25,265
10,588
J45998
Other asthma
19,832
10,580
E861
Hypovolemia
11,498
10,150
M109
Gout, unspecified
22,495
10,101
Y848
Other medical procedures as the cause of abnormal reaction of the patient, or of later complication, without mention of misadventure at the time of the procedure
11,442
9,782
Z8249
Family history of ischemic heart disease and other diseases of the circulatory system
11,916
9,767
E1165
Type 2 diabetes mellitus with hyperglycemia
16,359
9,704
R45851
Suicidal ideations
17,489
9,644
Z951
Presence of aortocoronary bypass graft
24,691
9,521
I5032
Chronic diastolic (congestive) heart failure
19,145
9,487
R112
Nausea with vomiting, unspecified
12,016
9,384
I110
Hypertensive heart disease with heart failure
15,711
9,191
N183
Chronic kidney disease, stage 3 (moderate)
16,119
9,026
Y92230
Patient room in hospital as the place of occurrence of the external cause
10,288
9,019
Z781
Physical restraint status
10,069
9,019
R55
Syncope and collapse
10,258
8,983
Z9181
History of falling
11,453
8,959
E8339
Other disorders of phosphorus metabolism
10,410
8,826
R000
Tachycardia, unspecified
10,298
8,822
W19XXXA
Unspecified fall, initial encounter
9,859
8,772
F05
Delirium due to known physiological condition
9,793
8,712
Z006
Encounter for examination for normal comparison and control in clinical research program
9,558
8,593
E870
Hyperosmolality and hypernatremia
10,095
8,588
G92
Toxic encephalopathy
10,272
8,527
I214
Non-ST elevation (NSTEMI) myocardial infarction
10,280
8,388
E6601
Morbid (severe) obesity due to excess calories
14,551
8,087
R6521
Severe sepsis with septic shock
8,969
8,040
Z9861
Coronary angioplasty status
17,023
7,972
J690
Pneumonitis due to inhalation of food and vomit
9,400
7,957
M549
Dorsalgia, unspecified
10,628
7,891
I5033
Acute on chronic diastolic (congestive) heart failure
13,625
7,817
I480
Paroxysmal atrial fibrillation
13,399
7,803
E1122
Type 2 diabetes mellitus with diabetic chronic kidney disease
18,564
7,774
G43909
Migraine, unspecified, not intractable, without status migrainosus
12,782
7,731
I5022
Chronic systolic (congestive) heart failure
14,474
7,710
Y92239
Unspecified place in hospital as the place of occurrence of the external cause
8,253
7,539
Hugo_Symbol
Required*
HUGO gene symbol (*one of Hugo_Symbol or Entrez_Gene_Id required)
Entrez_Gene_Id
. Link does not work from cBioPortal.
Two-dimensional matrix: one row per antibody, one column per sample. Values are log2 protein expression levels or Z-scores.
This format is deprecated. New studies should use the Structural Variant (SV) format above.
Required*
Entrez Gene numeric identifier (*one of Hugo_Symbol or Entrez_Gene_Id required)
[SAMPLE_ID columns]
Required
A sample ID. This field can only contain numbers, letters, points, underscores and hyphens.
cbp_driver
Optional
Custom driver annotation: Putative_Driver, Putative_Passenger, Unknown, NA
cbp_driver_annotation
Optional
Free-text description of driver annotation (max 80 chars). *This field must be present if the cbp_driver is also present in the MAF file
cbp_driver_tiers
Optional
Driver tier label, e.g. 'Highly actionable' (max 20 chars). *This field must be present if the cbp_driver_tiers_annotation is also present in the MAF file
cbp_driver_tiers_annotation
Optional
Description of driver tier (max 80 chars). *This field must be present if the cbp_driver_tiers is also present in the MAF file
Hugo_Symbol
Required*
HUGO gene symbol (*one of Hugo_Symbol or Entrez_Gene_Id required)
Entrez_Gene_Id
Required*
Entrez Gene numeric identifier (*one of Hugo_Symbol or Entrez_Gene_Id required)
ID
Required
Sample identifier
chrom
Required
Index of the chromosome column
Hugo_Symbol
Recommended*
HUGO gene symbol (*one of Hugo_Symbol or Entrez_Gene_Id required)
Entrez_Gene_Id
Recommended*
Entrez Gene numeric identifier (preferred; reduces ambiguity)
Hugo_Symbol
Required
A HUGO gene symbol.
Entrez_Gene_Id
Recommended
A Entrez Gene numeric identifier.
Hugo_Symbol
Required*
HUGO gene symbol (*one of Hugo_Symbol or Entrez_Gene_Id required)
Entrez_Gene_Id
Required*
Entrez Gene numeric identifier (*one of Hugo_Symbol or Entrez_Gene_Id required)
Composite.Element.REF
Required
Antibody identifier encoding gene symbol(s)/Entrez ID(s) and antibody ID, e.g. 'BRAF|B-Raf-M-NA' or 'MAPK1 MAPK3|MAPK_PT202_Y204'
[SAMPLE_ID columns]
Required
One column per sample; real-number protein level per antibody-sample pair
Sample_Id
Required
Sample identifier as defined in the clinical sample file
SV_Status
Required
SOMATIC or GERMLINE
Hugo_Symbol
Required
HUGO gene symbol
Entrez_Gene_Id
Required
Entrez Gene numeric identifier
PATIENT_ID
Required
Patient identifier from the dataset
START_DATE
Required
Days from date of diagnosis (day 0) to event start
chromosome
Required
Chromosome number (without 'chr' prefix)
peak_start
Required
Start coordinate of the region of maximal amplification or deletion
rank
Required
Gene rank by significance
gene
Required
HUGO gene symbol
SAMPLE_ID
Required
Sample identifier
[stable_id columns]
Required
One column per genetic profile (e.g. 'mutations', 'gistic'); value is gene panel stable_id or NA if not profiled
geneset_id
Required
Gene set name (uppercase); must match across score and p-value files
[SAMPLE_ID columns]
Required
Score file: GSVA score between -1.0 and 1.0, or NA. P-value file: p-value for the score
entity_stable_id
Required
Stable identifier for the assay entity (e.g. '1p', '1q', 'SBS1' for mutational signatures)
Name
Required
A column from generic_entity_meta_properties (using the property name as the column header)
[SAMPLE_ID columns]
Required
One column per sample; continuous copy number or log2 value per gene-sample pair
loc.start
Required
Index of the start position column
loc.end
Required
Index of the end position column
num.mark
Required
Index of a probe or description column
seg.mean
Required
No description provided (might be under the data tab on )
[SAMPLE_ID columns]
Required
One column per sample; real-number expression value or NA per gene-sample pair
Center
Optional
The sequencing center.
NCBI_Build
Required
Genome Reference Consortium Build used by variant calling software. Must be GRCh37 or GRCh38 for human, GRCm38 for mouse.
Chromosome
Required
A chromosome number, e.g. '7'.
Start_Position
Recommended
Start position of the event. Required for Cancer Hotspots annotations.
End_Position
Recommended
End position of the event. Required for Cancer Hotspots annotations.
Strand
Optional
Strand of the mutation. Assumed to be reported for the + strand.
Variant_Classification
Required
Translational effect of variant allele, e.g. Missense_Mutation, Nonsense_Mutation, Silent, Splice_Site, Frame_Shift_Del, etc. (TCGA MAF values).
Variant_Type
Optional
Variant type, e.g. SNP, DNP, INS, DEL, etc.
Reference_Allele
Required
The plus strand reference allele at this position.
Tumor_Seq_Allele1
Optional
Primary data genotype (allele 1).
Tumor_Seq_Allele2
Required
Primary data genotype (variant allele).
dbSNP_RS
Optional
Latest dbSNP rs ID for this position.
dbSNP_Val_Status
Optional
dbSNP validation status.
Tumor_Sample_Barcode
Required
Sample ID — either a TCGA barcode (patient ID extracted automatically) or a literal SAMPLE_ID from the clinical data file.
Matched_Norm_Sample_Barcode
Optional
Sample ID for the matched normal sample.
Match_Norm_Seq_Allele1
Optional
Primary data genotype for matched normal (allele 1).
Match_Norm_Seq_Allele2
Optional
Primary data genotype for matched normal (allele 2).
Tumor_Validation_Allele1
Optional
Secondary data from orthogonal technology (tumor allele 1).
Tumor_Validation_Allele2
Optional
Secondary data from orthogonal technology (tumor allele 2).
Match_Norm_Validation_Allele1
Optional
Secondary data from orthogonal technology (normal allele 1).
Match_Norm_Validation_Allele2
Optional
Secondary data from orthogonal technology (normal allele 2).
Verification_Status
Optional
Second pass results from independent attempt using same methods. Values: Verified, Unknown, or NA.
Validation_Status
Optional
Second pass results from orthogonal technology. Values: Valid, Invalid, Untested, Inconclusive, Redacted, Unknown, or NA.
Mutation_Status
Optional
Somatic or Germline are displayed in the Mutations tab. None, LOH, and Wildtype will not be loaded. Other values displayed as text.
Sequencing_Phase
Optional
Indicates current sequencing phase.
Sequence_Source
Optional
Molecular assay type used to produce the analytes used for sequencing.
Validation_Method
Optional
The assay platforms used for the validation call.
Score
Optional
Not used by cBioPortal.
BAM_File
Optional
Not used by cBioPortal.
Sequencer
Optional
Instrument used to produce primary data.
HGVSp_Short
Required
Amino acid change in HGVS notation, e.g. p.V600E.
t_alt_count
Optional
Variant allele count (tumor).
t_ref_count
Optional
Reference allele count (tumor).
n_alt_count
Optional
Variant allele count (normal).
n_ref_count
Optional
Reference allele count (normal).
cbp_driver
Optional
Custom driver annotation: Putative_Driver, Putative_Passenger, Unknown, NA, or empty.
cbp_driver_annotation
Optional
Free-text description of the driver annotation (max 80 chars).
cbp_driver_tiers
Optional
Driver tier label, e.g. 'Highly actionable' (max 20 chars).
cbp_driver_tiers_annotation
Optional
Description of the driver tier value (max 80 chars).
ASCN.ASCN_METHOD
Optional (ASCN)
Method used to obtain allele-specific copy number data, e.g. FACETS.
ASCN.CCF_EXPECTED_COPIES
Optional (ASCN)
Cancer-cell fraction if mutation exists on major allele.
ASCN.CCF_EXPECTED_COPIES_UPPER
Optional (ASCN)
Upper error bound for cancer-cell fraction estimate.
ASCN.EXPECTED_ALT_COPIES
Optional (ASCN)
Estimated number of copies harboring the mutant allele.
ASCN.CLONAL
Optional (ASCN)
Clonal status: Clonal, Subclonal, or Indeterminate.
ASCN.TOTAL_COPY_NUMBER
Optional (ASCN)
Total copy number of the gene.
ASCN.MINOR_COPY_NUMBER
Optional (ASCN)
Copy number of the minor allele.
ASCN.ASCN_INTEGER_COPY_NUMBER
Optional (ASCN)
Absolute integer copy-number estimate.
Site2_Region
Recommended
Region type: 5_Prime_UTR, 3_Prime_UTR, Promoter, Exon, or Intron
Site2_Chromosome
Recommended
Chromosome of gene 2
Site2_Position
Recommended
Genomic position of breakpoint at gene 2
NCBI_Build
Optional
Genome reference build: GRCh37 or GRCh38
Class
Optional
Deletion, Duplication, Insertion, Inversion, or Translocation
Event_Info
Optional
Free-text description of the event, e.g. 'TMPRSS2-ERG fusion'
Annotation
Optional
Free-text description of the gene/transcript rearrangement
Site1_Ensembl_Transcript_Id
Optional
Ensembl transcript ID of gene 1 (required for SV tab visualization)
Site2_Ensembl_Transcript_Id
Optional
Ensembl transcript ID of gene 2 (required for SV tab visualization)
DNA_Support
Optional
Yes or No — fusion detected from DNA data
RNA_Support
Optional
Yes or No — fusion detected from RNA data
SV_Length
Optional
Length of the structural variant in bases
Tumor_Split_Read_Count
Optional
Number of split reads supporting the call in tumor
Tumor_Paired_End_Read_Count
Optional
Number of paired-end reads supporting the call in tumor
Comments
Optional
Any free-text comments
[SAMPLE_ID columns]
Required
One column per sample; methylation beta-value per gene-sample pair
Site1_Hugo_Symbol
Recommended
HUGO gene symbol of gene 1 (left/3' site)
Site1_Ensembl_Transcript_Id
Optional
Ensembl transcript ID of gene 1 (required for SV tab visualization)
Site1_Entrez_Gene_Id
Recommended
Entrez Gene identifier of gene 1
Site1_Region_Number
Recommended
Region number of Site 1, e.g. exon 2
Site1_Region
Recommended
Region type: 5_Prime_UTR, 3_Prime_UTR, Promoter, Exon, or Intron
Site1_Chromosome
Recommended
Chromosome of gene 1
Site1_Contig
Optional
The contig of Site 1
Site1_Position
Recommended
Genomic position of breakpoint at gene 1
Site1_Description
Optional
Description of this event at site 2. This could be the location of the 2nd breakpoint in case of a fusion event.
Site2_Hugo_Symbol
Recommended
HUGO gene symbol of gene 2 (right/5' site)
Site2_Ensembl_Transcript_Id
Optional
Ensembl transcript ID of gene 2 (required for SV tab visualization)
Site2_Entrez_Gene_Id
Recommended
Entrez Gene identifier of gene 2
Site2_Region_Number
Recommended
Region number of Site 2
Site2_Region
Recommended
Region type: 5_Prime_UTR, 3_Prime_UTR, Promoter, Exon, or Intron
Site2_Chromosome
Recommended
Chromosome of gene 2
Site2_Contig
Optional
The contig of Site 2
Site2_Position
Recommended
Genomic position of breakpoint at gene 2
Site2_Description
Optional
Description of this event at site 1. This could be the location of the 1st breakpoint in case of a fusion event.
Site2_Effect_On_Frame
Optional
The effect of frame reading in gene 2. Frame_shift or InFrame (free text)
NCBI_Build
Optional
Genome reference build: GRCh37 or GRCh38
Class
Optional
Deletion, Duplication, Insertion, Inversion, or Translocation
Tumor_Split_Read_Count
Optional
Number of split reads supporting the call in tumor
Tumor_Paired_End_Read_Count
Optional
Number of paired-end reads supporting the call in tumor
Event_Info
Optional
Free-text description of the event, e.g. 'TMPRSS2-ERG fusion'
Connection_Type
Optional
Which direction the connection is made
Breakpoint_Type
Optional
PRECISE or IMPRECISE which explain the resolution. Fill in PRECISE if the breakpoint resolution is known down to the base pair
Annotation
Optional
Free-text description of the gene/transcript rearrangement
DNA_Support
Optional
Yes or No — fusion detected from DNA data
RNA_Support
Optional
Yes or No — fusion detected from RNA data
SV_Length
Optional
Length of the structural variant in bases
Normal_Read_Count
Optional
The total number of reads of the normal tissue.
Tumor_Read_Count
Optional
The total number of reads of the tumor tissue.
Normal_Variant_Count
Optional
The number of reads of the normal tissue that have the variant/allele.
Tumor_Variant_Count
Optional
The number of reads of the tumor tissue that have the variant/allele.
Normal_Paired_End_Read_Count
Optional
The number of paired-end reads of the normal tissue that support the call.
Normal_Split_Read_Count
Optional
The number of split reads of the normal tissue that support the call.
Comments
Optional
Any free-text comments
Center
Required
Sequencing center
Tumor_Sample_Barcode
Required
Sample ID
Fusion
Required
Description of the fusion, e.g. 'TMPRSS2-ERG fusion'
DNA_support
Required
Fusion detected from DNA: yes or no
RNA_support
Required
Fusion detected from RNA: yes or no
Method
Required
Algorithm/tool used for fusion detection
Frame
Required
in-frame or frameshift
Fusion_Status
Optional
Assesment of mutation type: SOMATIC, GERMLINE, UNKNOWN, or empty.
STOP_DATE
Required
Days from date of diagnosis to event end (blank if point-in-time event)
EVENT_TYPE
Required
Category of event: TREATMENT, LAB_TEST, IMAGING, STATUS, SPECIMEN, or any custom type
TREATMENT_TYPE
Optional
For TREATMENT events: Medical Therapy or Radiation Therapy
SUBTYPE
Optional
For TREATMENT events: Chemotherapy, Hormone Therapy, Targeted Therapy, etc.
AGENT
Optional
For TREATMENT events: agent name with number of cycles if applicable
TEST
Optional
For LAB_TEST events: type of test performed
RESULT
Optional
For LAB_TEST events: corresponding test result value
DIAGNOSTIC_TYPE
Optional
For IMAGING events: diagnostic tool used (MRI, CT scan, etc.)
STATUS
Optional
For STATUS events: best response or disease progression stage
SPECIMEN_SITE
Optional
For SPECIMEN events: site from which specimen was collected
SPECIMEN_TYPE
Optional
For SPECIMEN events: tissue or blood
STYLE_SHAPE
Optional
Render shape for this event: circle, square, triangle, diamond, star, or camera
STYLE_COLOR
Optional
Hexadecimal color value for rendering this event, e.g. #ffffff
Agent_Class
Optional
Suggested for TREATMENT events to classify agents into groups
Diagnostic_Type_Detailed
Optional
Suggested for IMAGING events as a detailed description
Source
Optional
Appears as a suggested column for IMAGING, STATUS, and SPECIMEN events
peak_end
Required
End coordinate of the region of maximal amplification or deletion
genes_in_region
Required
Comma-separated list of HUGO gene symbols in the wide peak
amp
Required
1 for amplification, 0 for deletion
cytoband
Required
Cytogenetic band specification including chromosome (Giemsa stain)
q_value
Required
Q-value (FDR-corrected p-value) of the peak region
N (or Nnon)
Required
Number of bases covered
n (or nnon)
Required
Number of mutations observed
p
Required
P-value: probability mutations are due to background processes
q
Required
Q-value: p-value corrected for multiple testing
Description
Required
Another generic_entity_meta_properties column
URL
Required
Another generic_entity_meta_properties column
[SAMPLE_ID columns]
Required
One column per sample; numeric value per entity-sample pair
How old were you when you had sexual intercourse for
Q49
1
148809
66.8378
How old were you when you tried marijuana for the first time?
Q46
1
133816
60.1037
How old were you when you tried marijuana for the first time?
Q36
1
133363
59.9002
How old were you when you smoked a whole cigarette
Q25
2
104615
46.988
How old were you when you first started smoking
Q27
1
97632
43.8516
How old were you when you had your first drink of alcohol other than
Q40
1
95370
42.8356
How old were you when you had sexual intercourse for the first time?
Q58
1
90869
40.814
How old were you when you smoked a whole cigarette
Q25
1
69859
31.3773
How old were you when you first started smoking
Q27
<missing>
62654
28.1411
How old were you when you had sexual intercourse for the first time?
Q58
2
57567
25.8563
How old were you when you had sexual intercourse for
Q49
<missing>
54754
24.5928
How old are you?
Q1
5
51555
23.156
How old are you?
Q1
6
51323
23.0518
How old were you when you tried marijuana for the
Q36
<missing>
47463
21.3181
How old are you?
Q1
4
46293
20.7926
How old were you when you first started smoking
Q27
2
42753
19.2026
How old were you when you smoked a whole cigarette
Q25
<missing>
40160
18.0379
How old were you when you tried marijuana for the first time?
Q46
<missing>
40087
18.0051
How old were you when you had sexual intercourse for the first time?
Q58
<missing>
33931
15.2402
How old were you when you had your first drink of alcohol other than
Q40
<missing>
31177
14.0032
How old are you?
Q1
7
30584
13.7369
How old were you when you had your first drink of alcohol other than
Q40
5
23180
10.4113
How old are you?
Q1
3
21452
9.6352
How old are you?
Q1
<missing>
20776
9.3316
How old were you when you had your first drink of alcohol other than
Q40
6
19383
8.7059
How old were you when you had your first drink of alcohol other than
Q40
2
18478
8.2994
How old were you when you had your first drink of alcohol other than
Q40
4
14115
6.3398
How old were you when you had your first drink of alcohol other than
Q40
3
13939
6.2607
How old were you when you tried marijuana for the first time?
Q46
6
12007
5.393
How old were you when you tried marijuana for the first time?
Q46
5
10911
4.9007
How old were you when you tried marijuana for the
Q36
5
10211
4.5863
How old were you when you tried marijuana for the
Q36
2
9800
4.4017
How old were you when you tried marijuana for the first time?
Q46
2
8566
3.8474
How old were you when you had sexual intercourse for the first time?
Q58
7
8301
3.7284
How old were you when you had sexual intercourse for the first time?
Q58
4
8072
3.6256
How old were you when you had sexual intercourse for the first time?
Q58
3
7255
3.2586
How old were you when you had sexual intercourse for the first time?
Q58
5
7133
3.2038
How old were you when you had sexual intercourse for the first time?
Q58
6
7130
3.2025
How old were you when you had your first drink of alcohol other than
Q40
7
7000
3.1441
How old were you when you tried marijuana for the
Q36
3
6873
3.087
How old were you when you had sexual intercourse for
Q49
2
6673
2.9972
How old were you when you tried marijuana for the
Q36
6
6497
2.9181
How old were you when you tried marijuana for the first time?
Q46
3
6397
2.8732
How old were you when you first started smoking
Q27
5
5736
2.5763
How old were you when you tried marijuana for the
Q36
4
5677
2.5498
How old were you when you tried marijuana for the first time?
Q46
4
5453
2.4492
How old were you when you first started smoking
Q27
3
5311
2.3854
How old were you when you tried marijuana for the first time?
Q46
7
4585
2.0594
How old were you when you first started smoking
Q27
6
3954
1.7759
How old were you when you first started smoking
Q27
4
3657
1.6425
How old were you when you had sexual intercourse for
Q49
3
3524
1.5828
How old were you when you had sexual intercourse for
Q49
6
2884
1.2954
How old were you when you smoked a whole cigarette
Q25
3
2684
1.2055
How old were you when you had sexual intercourse for
Q49
4
2324
1.0438
How old were you when you smoked a whole cigarette
Q25
5
2229
1.0012
How old were you when you tried marijuana for the
Q36
7
2114
0.9495
How old were you when you had sexual intercourse for
Q49
5
2027
0.9104
How old were you when you smoked a whole cigarette
Q25
4
1449
0.6508
How old were you when you smoked a whole cigarette
Q25
6
1300
0.5839
How old were you when you had sexual intercourse for
Q49
7
1057
0.4748
How old were you when you first started smoking
Q27
7
945
0.4244
How old are you?
Q1
1
371
0.1666
How old were you when you smoked a whole cigarette
Q25
7
346
0.1554
How old are you?
Q1
2
287
0.1289
the past 7 days, how many times did you drink a can, bottle, orQ95 glass of a sports drink such as Gatorade or PowerAde?
During
<missing>
210733
94.6511
During your life, how many times have you used heroin (also called
Q52
1
185111
83.1429
Among students who had sexual intercourse during the past
QNBCNONE
<missing>
208535
93.6638
How tall are you without your shoes on? (Note: Data are in meters.)
Q6
1
24558
11.0303
During your life, how many times have you used synthetic marijuana?
Q48
1
150532
67.6117
Percentage of students who actually attempted suicide (one or more
QN28
<missing>
129206
58.0331
Race as originally scanned
RACEORIG
<missing>
212726
95.5462
During the past 30 days, on how many days did you smoke cigars,
Q38
1
170175
76.4344
Weight_2
Weight_2
<missing>
208725
93.7492
During the past 30 days, how did you usually get the alcohol you
Q44
1
142858
64.1649
During the past 30 days, how did you usually get the alcohol you
Q45
1
136813
61.4498
Percentage of students who had their first drink of alcohol
QN40
<missing>
134826
60.5573
Percentage of students who had at least one drink of alcohol
QN39
<missing>
126782
56.9443
Percentage of students who drank alcohol or used drugs before last
QN61
<missing>
123381
55.4168
During the past 7 days, how many times did you drink a can, bottle, or
Q90
<missing>
118346
53.1553
Percentage of students who had five or more drinks of
QN42
2
95886
43.0673
Percentage of students who had five or more drinks of
QN42
<missing>
93935
42.1911
Percentage of students who currently drank alcohol (at least one drink
QN41
<missing>
89242
40.0832
Percentage of students who currently drank alcohol (at least one drink
QN41
2
87479
39.2913
Percentage of students who drank alcohol or used drugs before last
QN61
2
66438
29.8407
Percentage of students who had their first drink of alcohol
QN40
2
59933
26.919
Percentage of students who had at least one drink of alcohol
QN39
2
54900
24.6584
During the past 7 days, how many times did you drink a can, bottle, or
Q90
1
52140
23.4188
Percentage of students who currently drank alcohol (at least one drink
QN41
1
42467
19.0741
During the past 30 days, how did you usually get the alcohol you
Q44
<missing>
37002
16.6195
Percentage of students who had at least one drink of alcohol
QN39
1
35427
15.9121
Percentage of students who had five or more drinks of
QN42
1
30084
13.5123
During the past 7 days, how many times did you drink a can, bottle, or
Q90
2
27199
12.2165
Percentage of students who had their first drink of alcohol
QN40
1
24703
11.0954
During the past 30 days, how did you usually get the alcohol you
Q45
2
21835
9.8072
During the past 30 days, how did you usually get the alcohol you
Q45
<missing>
19565
8.7877
Percentage of students who drank alcohol or used drugs before last
QN61
1
19535
8.7742
Percentage of students who drank alcohol or used drugs before last
QN61
.
13288
5.9683
During the past 30 days, how did you usually get the alcohol you
Q45
6
11459
5.1468
During the past 30 days, how did you usually get the alcohol you
Q45
5
11315
5.0821
During the past 30 days, how did you usually get the alcohol you
Q44
2
10119
4.545
During the past 7 days, how many times did you drink a can, bottle, or
Q90
3
8254
3.7073
During the past 30 days, how did you usually get the alcohol you
Q45
7
7292
3.2752
During the past 30 days, how did you usually get the alcohol you
Q44
6
7058
3.1701
During the past 30 days, how did you usually get the alcohol you
Q44
3
7026
3.1557
During the past 30 days, how did you usually get the alcohol you
Q44
7
6823
3.0646
During the past 30 days, how did you usually get the alcohol you
Q45
3
6370
2.8611
During the past 30 days, how did you usually get the alcohol you
Q45
4
6138
2.7569
Percentage of students who had at least one drink of alcohol
QN39
.
5533
2.4852
the past 7 days, how many times did you drink a can, bottle, orQ95 glass of a sports drink such as Gatorade or PowerAde?
During
1
5359
2.407
During the past 30 days, how did you usually get the alcohol you
Q44
5
5290
2.376
During the past 7 days, how many times did you drink a can, bottle, or
Q90
4
4978
2.2359
During the past 30 days, how did you usually get the alcohol you
Q44
4
4244
1.9062
During the past 7 days, how many times did you drink a can, bottle, or
Q90
5
3845
1.727
the past 7 days, how many times did you drink a can, bottle, orQ95 glass of a sports drink such as Gatorade or PowerAde?
During
2
3754
1.6861
Percentage of students who currently drank alcohol (at least one drink
QN41
.
3454
1.5514
During the past 7 days, how many times did you drink a can, bottle, or
Q90
6
3364
1.5109
Percentage of students who had their first drink of alcohol
QN40
.
3180
1.4283
Percentage of students who had five or more drinks of
QN42
.
2737
1.2293
During the past 7 days, how many times did you drink a can, bottle, or
Q90
8
2561
1.1503
the past 7 days, how many times did you drink a can, bottle, orQ95 glass of a sports drink such as Gatorade or PowerAde?
During
3
1306
0.5866
the past 7 days, how many times did you drink a can, bottle, orQ95 glass of a sports drink such as Gatorade or PowerAde?
During
4
642
0.2884
the past 7 days, how many times did you drink a can, bottle, orQ95 glass of a sports drink such as Gatorade or PowerAde?
During
5
358
0.1608
the past 7 days, how many times did you drink a can, bottle, orQ95 glass of a sports drink such as Gatorade or PowerAde?
During
7
301
0.1352
the past 7 days, how many times did you drink a can, bottle, orQ95 glass of a sports drink such as Gatorade or PowerAde?
During
6
189
0.0849
During your life, how many times have you used ecstasy (also called
Q54
1
151349
67.9786
During your life, how many times have you used methamphetamines
Q53
1
150263
67.4909
Percenta ge of students who used ecstasy one or more times
QN54
<missing>
123874
55.6382
Percent age of students who used heroin one or more times
QN52
<missing>
123701
55.5605
Percentage of students who were offered, sold, or given an illegal drug
QN56
<missing>
119726
53.7751
Percentage of students who took steroid pills or shots
QN55
<missing>
114377
51.3726
During your life, how many times have you used hallucinogenic drugs,
Q89
<missing>
110390
49.5818
"speed," "crystal meth," "crank," "ice," or "meth," one or more times
QN53
2
109473
49.17
"speed," "crystal meth," "crank," "ice," or "meth," one or more times
QN53
<missing>
106830
47.9829
Percent age of students who used any form of cocaine,
QN49
<missing>
106382
47.7816
Percent age of students who used any form of cocaine,
QN49
2
103125
46.3188
Percentage of students who took steroid pills or shots
QN55
2
102686
46.1216
Percent age of students who used heroin one or more times
QN52
2
93379
41.9413
Percenta ge of students who used ecstasy one or more times
QN54
2
91802
41.233
Percentage of students who were offered, sold, or given an illegal drug
QN56
2
90106
40.4712
During your life, how many times have you used hallucinogenic drugs,
Q89
1
60279
27.0744
During your life, how many times have you used ecstasy (also called
Q54
<missing>
41978
18.8545
During your life, how many times have you used methamphetamines
Q53
2
40252
18.0792
During your life, how many times have you used hallucinogenic drugs,
Q89
2
25622
11.5082
During your life, how many times have you used methamphetamines
Q53
<missing>
25328
11.3761
During your life, how many times have you used heroin (also called
Q52
<missing>
24662
11.077
During your life, how many times have you used ecstasy (also called
Q54
2
19334
8.6839
Percentage of students who were offered, sold, or given an illegal drug
QN56
1
11677
5.2447
Percent age of students who used any form of cocaine,
QN49
1
11322
5.0853
During your life, how many times have you used hallucinogenic drugs,
Q89
3
10445
4.6914
During your life, how many times have you used hallucinogenic drugs,
Q89
4
8857
3.9781
During your life, how many times have you used heroin (also called
Q52
3
6656
2.9896
Percenta ge of students who used ecstasy one or more times
QN54
1
5814
2.6114
During your life, how many times have you used hallucinogenic drugs,
Q89
5
5595
2.513
During your life, how many times have you used methamphetamines
Q53
3
4570
2.0526
"speed," "crystal meth," "crank," "ice," or "meth," one or more times
QN53
1
4541
2.0396
Percent age of students who used heroin one or more times
QN52
1
4246
1.9071
During your life, how many times have you used heroin (also called
Q52
2
4081
1.833
Percentage of students who took steroid pills or shots
QN55
1
3937
1.7683
During your life, how many times have you used ecstasy (also called
Q54
3
3816
1.714
During your life, how many times have you used ecstasy (also called
Q54
4
3386
1.5208
Percent age of students who used any form of cocaine,
QN49
.
1813
0.8143
"speed," "crystal meth," "crank," "ice," or "meth," one or more times
QN53
.
1798
0.8076
Percentage of students who took steroid pills or shots
QN55
.
1642
0.7375
During your life, how many times have you used ecstasy (also called
Q54
5
1448
0.6504
Percent age of students who used heroin one or more times
QN52
.
1316
0.5911
During your life, how many times have you used ecstasy (also called
Q54
6
1160
0.521
During your life, how many times have you used methamphetamines
Q53
6
1153
0.5179
Percenta ge of students who used ecstasy one or more times
QN54
.
1152
0.5174
Percentage of students who were offered, sold, or given an illegal drug
QN56
.
1133
0.5089
During your life, how many times have you used heroin (also called
Q52
6
1080
0.4851
During your life, how many times have you used hallucinogenic drugs,
Q89
7
892
0.4006
During your life, how many times have you used methamphetamines
Q53
4
673
0.3023
During your life, how many times have you used heroin (also called
Q52
4
661
0.2969
During your life, how many times have you used methamphetamines
Q53
5
403
0.181
During your life, how many times have you used heroin (also called
Q52
5
391
0.1756
During your life, how many times have you used hallucinogenic drugs,
Q89
6
342
0.1536
During your life, how many times have you used ecstasy (also called
Q54
7
171
0.0768
Among students who had sexual intercourse during the past
QNDUALBC
<missing>
208489
93.6432
Percentage of students who have never had sex, haven’t had
QNRESPSX
<missing>
207428
93.1666
Percentile for body mass index, by age and sex
BMIPct
<missing>
194684
87.4426
Among students who had sexual intercourse during the past
QN63
<missing>
164705
73.9775
Have you ever been physically forced to have sexual intercourse when
Q19
1
134269
60.3071
Of students who had sexual
QN62
<missing>
131640
59.1263
Percentage of st udents who had sexual intercourse with four
QN60
<missing>
125984
56.5859
Percentage of st udents who had sexual intercourse for the
QN59
<missing>
125909
56.5522
Percentage of students who ever had sexual intercourse
QN58
<missing>
125137
56.2055
touching, or being physically forced to have sexual intercourse] that
QN20
<missing>
110558
49.6573
The last time you had sexual intercourse with an opposite-sex partner,
Q63
1
106099
47.6545
touching, or being physically forced to have sexual intercourse] that
QN20
2
98106
44.0645
What is your sex?
Q2
1
93503
41.997
What is your sex?
Q2
2
91077
40.9074
The last time you had sexual intercourse with an opposite-sex partner,
Q62
1
89345
40.1294
Percentage of st udents who had sexual intercourse with four
QN60
2
71907
32.2971
Percentage of st udents who had sexual intercourse for the
QN59
2
71381
32.0609
Percentage of students who ever had sexual intercourse
QN58
2
66170
29.7204
Have you ever been physically forced to have sexual intercourse when
Q19
2
52734
23.6856
The last time you had sexual intercourse with an opposite-sex partner,
Q62
<missing>
47794
21.4667
The last time you had sexual intercourse with an opposite-sex partner,
Q63
<missing>
47060
21.1371
Of students who had sexual
QN62
2
43271
19.4352
The last time you had sexual intercourse with an opposite-sex partner,
Q62
3
40539
18.2082
What is your sex?
Q2
<missing>
38062
17.0956
Of students who had sexual
QN62
.
29442
13.2239
The last time you had sexual intercourse with an opposite-sex partner,
Q62
2
29138
13.0874
Have you ever been physically forced to have sexual intercourse when
Q19
<missing>
29104
13.0721
The last time you had sexual intercourse with an opposite-sex partner,
Q63
2
25941
11.6514
Percentage of students who ever had sexual intercourse
QN58
1
25482
11.4453
Among students who had sexual intercourse during the past
QN63
2
25387
11.4026
The last time you had sexual intercourse with an opposite-sex partner,
Q63
3
23425
10.5214
Percentage of st udents who had sexual intercourse for the
QN59
1
20511
9.2125
Among students who had sexual intercourse during the past
QN63
.
19574
8.7917
Percentage of st udents who had sexual intercourse with four
QN60
1
19071
8.5658
Of students who had sexual
QN62
1
18289
8.2145
The last time you had sexual intercourse with an opposite-sex partner,
Q63
4
14978
6.7274
Among students who had sexual intercourse during the past
QN63
1
12976
5.8282
Among students who had sexual intercourse during the past
QNDUALBC
2
12891
5.79
touching, or being physically forced to have sexual intercourse] that
QN20
1
12857
5.7747
Among students who had sexual intercourse during the past
QNBCNONE
2
12113
5.4406
Percentage of students who have never had sex, haven’t had
QNRESPSX
1
11200
5.0305
The last time you had sexual intercourse with an opposite-sex partner,
Q62
4
10853
4.8746
Percentage of students who ever had sexual intercourse
QN58
.
5853
2.6289
Percentage of st udents who had sexual intercourse with four
QN60
.
5680
2.5512
Percentage of st udents who had sexual intercourse for the
QN59
.
4841
2.1743
Have you ever been physically forced to have sexual intercourse when
Q19
3
3726
1.6735
The last time you had sexual intercourse with an opposite-sex partner,
Q62
5
2628
1.1804
Percentage of students who have never had sex, haven’t had
QNRESPSX
.
2147
0.9643
Among students who had sexual intercourse during the past
QNBCNONE
1
1994
0.8956
The last time you had sexual intercourse with an opposite-sex partner,
Q63
5
1923
0.8637
Percentage of students who have never had sex, haven’t had
QNRESPSX
2
1867
0.8386
Percentile for body mass index, by age and sex
BMIPct
.
1777
0.7981
The last time you had sexual intercourse with an opposite-sex partner,
Q63
6
1671
0.7505
Among students who had sexual intercourse during the past
QNDUALBC
1
1262
0.5668
touching, or being physically forced to have sexual intercourse] that
QN20
.
1121
0.5035
Have you ever been physically forced to have sexual intercourse when
Q19
4
1029
0.4622
Percentile for body mass index, by age and sex
BMIPct
99
1003
0.4505
The last time you had sexual intercourse with an opposite-sex partner,
Q63
7
927
0.4164
The last time you had sexual intercourse with an opposite-sex partner,
Q62
6
922
0.4141
Percentile for body mass index, by age and sex
BMIPct
98
824
0.3701
The last time you had sexual intercourse with an opposite-sex partner,
Q62
7
810
0.3638
Have you ever been physically forced to have sexual intercourse when
Q19
5
765
0.3436
Percentile for body mass index, by age and sex
BMIPct
97
694
0.3117
Have you ever been physically forced to have sexual intercourse when
Q19
8
689
0.3095
Percentile for body mass index, by age and sex
BMIPct
96
615
0.2762
Percentile for body mass index, by age and sex
BMIPct
95
603
0.2708
Percentile for body mass index, by age and sex
BMIPct
94
558
0.2506
Have you ever been physically forced to have sexual intercourse when
Q19
6
200
0.0898
In what grade are you?
Q3
4
46371
20.8276
In what grade are you?
Q3
3
46302
20.7966
In what grade are you?
Q3
1
45831
20.5851
In what grade are you?
Q3
2
45740
20.5442
In what grade are you?
Q3
<missing>
38121
17.1221
In what grade are you?
Q3
5
277
0.1244
How tall are you without your shoes on? (Note: Data are in meters.)
Q6
<missing>
22291
10.012
How tall are you without your shoes on? (Note: Data are in meters.)
Q6
2
8056
3.6184
How tall are you without your shoes on? (Note: Data are in meters.)
Q6
4
7989
3.5883
How tall are you without your shoes on? (Note: Data are in meters.)
Q6
5
7904
3.5501
How tall are you without your shoes on? (Note: Data are in meters.)
Q6
1.68
7686
3.4522
How tall are you without your shoes on? (Note: Data are in meters.)
Q6
1.63
7486
3.3623
How tall are you without your shoes on? (Note: Data are in meters.)
Q6
1.7
7481
3.3601
During the past 30 days, how many times did you use marijuana?
Q47
1
146777
65.9251
Percentage of students who used marijuana one or more
QN44
<missing>
132172
59.3653
Percentage of students who tried marijuana for the first time before age
QN46
<missing>
128714
57.8121
Percentage of students who ever used marijuana (one or more times
QN45
<missing>
115997
52.1002
Percentage of students who ever used synthetic marijuana (one or more
QN48
<missing>
114057
51.2289
Percentage of students who currently used marijuana (one or more
QN47
<missing>
106261
47.7273
Percentage of students who currently used marijuana (one or more
QN47
2
96868
43.5084
Percentage of students who ever used synthetic marijuana (one or more
QN48
2
94794
42.5769
Percentage of students who ever used marijuana (one or more times
QN45
2
86953
39.0551
Percentage of students who tried marijuana for the first time before age
QN46
2
70888
31.8395
Percentage of students who used marijuana one or more
QN44
2
57964
26.0346
During your life, how many times have you used synthetic marijuana?
Q48
<missing>
46881
21.0567
During the past 30 days, how many times did you use marijuana?
Q47
<missing>
38532
17.3067
Percentage of students who tried marijuana for the first time before age
QN46
1
21499
9.6563
Percentage of students who used marijuana one or more
QN44
1
21055
9.4569
Percentage of students who ever used marijuana (one or more times
QN45
1
18156
8.1548
Percentage of students who currently used marijuana (one or more
QN47
1
18080
8.1207
Percentage of students who ever used synthetic marijuana (one or more
QN48
1
12466
5.5991
Percentage of students who used marijuana one or more
QN44
.
11451
5.1432
During your life, how many times have you used synthetic marijuana?
Q48
2
10943
4.9151
During the past 30 days, how many times did you use marijuana?
Q47
2
10854
4.8751
During the past 30 days, how many times did you use marijuana?
Q47
6
7781
3.4948
During the past 30 days, how many times did you use marijuana?
Q47
5
6428
2.8871
During the past 30 days, how many times did you use marijuana?
Q47
3
5749
2.5822
During your life, how many times have you used synthetic marijuana?
Q48
6
4511
2.0261
During the past 30 days, how many times did you use marijuana?
Q47
4
3880
1.7427
During your life, how many times have you used synthetic marijuana?
Q48
3
3540
1.59
During your life, how many times have you used synthetic marijuana?
Q48
5
3290
1.4777
During the past 30 days, how many times did you use marijuana?
Q47
7
2641
1.1862
During your life, how many times have you used synthetic marijuana?
Q48
4
2484
1.1157
Percentage of students who tried marijuana for the first time before age
QN46
.
1541
0.6921
Percentage of students who ever used marijuana (one or more times
QN45
.
1536
0.6899
Percentage of students who currently used marijuana (one or more
QN47
.
1433
0.6436
Percentage of students who ever used synthetic marijuana (one or more
QN48
.
1325
0.5951
During your life, how many times have you used synthetic marijuana?
Q48
7
461
0.2071
Percentage of students who seriously considered attempting suicide
QN26
<missing>
119070
53.4805
Percentage of students who seriously considered attempting suicide
QN26
2
82799
37.1893
Percentage of students who actually attempted suicide (one or more
QN28
2
62797
28.2054
Percentage of students who actually attempted suicide (one or more
QN28
1
28975
13.0142
Percentage of students who seriously considered attempting suicide
QN26
1
15515
6.9686
Percentage of students who seriously considered attempting suicide
QN26
.
5258
2.3616
Percentage of students who actually attempted suicide (one or more
QN28
.
1664
0.7474
Race/ethnicity as originally scanned
Q4ORIG
<missing>
208956
93.8529
Race/Ethnicity
RaceEth
<missing>
208849
93.8049
RACEETH
RACEETH
<missing>
130067
58.4198
What is your race?
Q5
E
58300
26.1855
RACEETH
RACEETH
5
43400
19.4932
What is your race?
Q5
C
19353
8.6924
What is your race?
Q5
4
15900
7.1415
RACEETH
RACEETH
3
14709
6.6066
RACEETH
RACEETH
7
13011
5.8439
What is your race?
Q5
<missing>
12426
5.5812
What is your race?
Q5
3
12027
5.4019
What is your race?
Q5
2
9784
4.3945
What is your race?
Q5
5
9319
4.1856
RACEETH
RACEETH
6
8934
4.0127
What is your race?
Q5
1
7803
3.5047
Race/ethnicity as originally scanned
Q4ORIG
F
6122
2.7497
Race as originally scanned
RACEORIG
E
6096
2.738
Race/Ethnicity
RaceEth
5
5775
2.5939
RACEETH
RACEETH
8
5630
2.5287
RACEETH
RACEETH
2
4078
1.8316
Race/ethnicity as originally scanned
Q4ORIG
C
3347
1.5033
Race/Ethnicity
RaceEth
3
2931
1.3165
Race as originally scanned
RACEORIG
C
2662
1.1956
RACEETH
RACEETH
1
2175
0.9769
Race/ethnicity as originally scanned
Q4ORIG
D
2069
0.9293
Race/Ethnicity
RaceEth
6
2008
0.9019
Race/Ethnicity
RaceEth
7
1868
0.839
Race/ethnicity as originally scanned
Q4ORIG
D F
651
0.2924
Race/Ethnicity
RaceEth
2
428
0.1922
Race/Ethnicity
RaceEth
8
383
0.172
Race/ethnicity as originally scanned
Q4ORIG
B
368
0.1653
Race as originally scanned
RACEORIG
B
323
0.1451
Race as originally scanned
RACEORIG
A
314
0.141
Race/Ethnicity
RaceEth
1
297
0.1334
Race/ethnicity as originally scanned
Q4ORIG
A D
188
0.0844
Race/ethnicity as originally scanned
Q4ORIG
A
147
0.066
Race as originally scanned
RACEORIG
D
125
0.0561
Race as originally scanned
RACEORIG
A E
122
0.0548
Race as originally scanned
RACEORIG
C E
64
0.0287
Percentage of students who used any tobacco during the past
QNANYTOB
<missing>
151221
67.9211
the days they smoked during the 30 days before the survey, among
QN33
<missing>
149093
66.9654
During the past 30 days, on how many days did you smoke cigarettes?
Q32
1
129641
58.2285
Percentage of students who used chewing tobacco, snuff, or
QN36
<missing>
128103
57.5377
Percent age of students who usually got their own cigarettes
QN32
<missing>
124426
55.8861
Percentage of students who smoked cigarettes on 20 or more
QNFRCIG
<missing>
124299
55.8291
Percent age of students who smoked a whole cigarette for the
QN29
<missing>
120611
54.1726
(including e-cigarettes, vapes, vape pens, e-cigars, e-hookahs, hookah
QN34
<missing>
117912
52.9604
Of st udents who are current smokers, the percentage who
QN35
<missing>
117286
52.6792
Percent age of students who smoked cigarettes on one or
QN30
<missing>
116792
52.4573
During the past 30 days, on the days you smoked, how many cigarettes
Q33
1
112645
50.5947
Percentage of students who first tried cigarette smoking before age 13
QN31
<missing>
110944
49.8307
Percentage of students who currently smoked cigars (cigars, cigarillos,
QN38
2
109352
49.1156
During the past 12 months, did you ever try to quit using all tobacco
Q39
1
106254
47.7241
Percentage of students who currently smoked cigars (cigars, cigarillos,
QN38
<missing>
100108
44.9637
Have you ever smoked cigarettes regularly, that is,
Q26
1
95556
42.9191
Percentage of students who smoked cigarettes on 20 or more
QNFRCIG
2
90160
40.4955
Percent age of students who smoked cigarettes on one or
QN30
2
84462
37.9362
Percent age of students who smoked a whole cigarette for the
QN29
2
83702
37.5949
Have you ever smoked cigarettes regularly, that is,
Q26
2
79194
35.5701
(including e-cigarettes, vapes, vape pens, e-cigars, e-hookahs, hookah
QN34
2
77394
34.7616
Percentage of students who first tried cigarette smoking before age 13
QN31
2
73716
33.1097
Percentage of students who used chewing tobacco, snuff, or
QN36
2
71816
32.2563
During the past 30 days, on the days you smoked, how many cigarettes
Q33
<missing>
69887
31.3899
During the past 30 days, on how many days did you smoke cigarettes?
Q32
<missing>
57597
25.8698
Percent age of students who usually got their own cigarettes
QN32
2
55187
24.7873
the days they smoked during the 30 days before the survey, among
QN33
2
53302
23.9407
Of st udents who are current smokers, the percentage who
QN35
2
50395
22.635
Percentage of students who used any tobacco during the past
QNANYTOB
2
50196
22.5456
Have you ever smoked cigarettes regularly, that is,
Q26
<missing>
44789
20.117
During the past 12 months, did you ever try to quit using all tobacco
Q39
<missing>
42567
19.119
Percent age of students who usually got their own cigarettes
QN32
.
38022
17.0776
Of st udents who are current smokers, the percentage who
QN35
.
36165
16.2436
(including e-cigarettes, vapes, vape pens, e-cigars, e-hookahs, hookah
QN34
1
24429
10.9723
Percentage of students who first tried cigarette smoking before age 13
QN31
.
24388
10.9539
During the past 30 days, on the days you smoked, how many cigarettes
Q33
2
22963
10.3139
Of st udents who are current smokers, the percentage who
QN35
1
18796
8.4423
During the past 12 months, did you ever try to quit using all tobacco
Q39
2
18788
8.4387
During the past 12 months, did you ever try to quit using all tobacco
Q39
3
18727
8.4113
Percent age of students who smoked cigarettes on one or
QN30
1
18578
8.3443
Percentage of students who used any tobacco during the past
QNANYTOB
1
16010
7.1909
the days they smoked during the 30 days before the survey, among
QN33
.
15290
6.8675
During the past 30 days, on how many days did you smoke cigars,
Q38
2
14963
6.7207
Percentage of students who used chewing tobacco, snuff, or
QN36
.
14769
6.6335
Percent age of students who smoked a whole cigarette for the
QN29
1
14643
6.5769
Percentage of students who first tried cigarette smoking before age 13
QN31
1
13594
6.1058
During the past 30 days, on how many days did you smoke cigars,
Q38
<missing>
10858
4.8769
Percentage of students who currently smoked cigars (cigars, cigarillos,
QN38
1
10347
4.6474
During the past 12 months, did you ever try to quit using all tobacco
Q39
4
10071
4.5234
During the past 12 months, did you ever try to quit using all tobacco
Q39
5
9261
4.1596
During the past 12 months, did you ever try to quit using all tobacco
Q39
7
9134
4.1026
During the past 30 days, on how many days did you smoke cigarettes?
Q32
2
8895
3.9952
During the past 30 days, on how many days did you smoke cigars,
Q38
3
8567
3.8479
Percentage of students who used chewing tobacco, snuff, or
QN36
1
7954
3.5726
During the past 12 months, did you ever try to quit using all tobacco
Q39
6
7840
3.5213
During the past 30 days, on how many days did you smoke cigarettes?
Q32
5
7573
3.4014
During the past 30 days, on how many days did you smoke cigars,
Q38
5
6518
2.9276
During the past 30 days, on how many days did you smoke cigars,
Q38
4
5629
2.5283
Percentage of students who smoked cigarettes on 20 or more
QNFRCIG
1
5271
2.3675
Percentage of students who used any tobacco during the past
QNANYTOB
.
5215
2.3423
Percent age of students who usually got their own cigarettes
QN32
1
5007
2.2489
During the past 30 days, on how many days did you smoke cigarettes?
Q32
4
4998
2.2449
During the past 30 days, on how many days did you smoke cigarettes?
Q32
6
4968
2.2314
the days they smoked during the 30 days before the survey, among
QN33
1
4957
2.2264
During the past 30 days, on the days you smoked, how many cigarettes
Q33
3
4709
2.1151
During the past 30 days, on the days you smoked, how many cigarettes
Q33
4
4553
2.045
During the past 30 days, on how many days did you smoke cigarettes?
Q32
3
4486
2.0149
During the past 30 days, on how many days did you smoke cigars,
Q38
6
3709
1.6659
Percent age of students who smoked a whole cigarette for the
QN29
.
3686
1.6556
During the past 30 days, on how many days did you smoke cigarettes?
Q32
7
3595
1.6147
During the past 30 days, on the days you smoked, how many cigarettes
Q33
5
2938
1.3196
Percentage of students who smoked cigarettes on 20 or more
QNFRCIG
.
2912
1.3079
(including e-cigarettes, vapes, vape pens, e-cigars, e-hookahs, hookah
QN34
.
2907
1.3057
Percentage of students who currently smoked cigars (cigars, cigarillos,
QN38
.
2835
1.2733
Percent age of students who smoked cigarettes on one or
QN30
.
2810
1.2621
During the past 30 days, on the days you smoked, how many cigarettes
Q33
7
2582
1.1597
Have you ever smoked cigarettes regularly, that is,
Q26
3
2304
1.0348
During the past 30 days, on how many days did you smoke cigars,
Q38
7
2223
0.9985
During the past 30 days, on the days you smoked, how many cigarettes
Q33
6
2037
0.9149
Have you ever smoked cigarettes regularly, that is,
Q26
5
481
0.216
Have you ever smoked cigarettes regularly, that is,
Q26
4
318
0.1428
Weight_3
Weight_3
<missing>
208601
93.6935
Percentage of students who were overweight
QNOWT
<missing>
181492
81.5174
Percentage of students who are overweight
QNOVWGT
<missing>
179470
80.6092
Percentage of students who were trying to lose weight
QN65
<missing>
171268
76.9253
Percentage of students who were trying to lose weight
QN67
<missing>
129497
58.1638
Weight variable*
WEIGHT
<missing>
92780
41.6723
How much do you weigh without your shoes on?
Q7
1
65002
29.1957
Percentage of students who were trying to lose weight
QN67
2
52925
23.7713
Percentage of students who were trying to lose weight
QN67
1
39153
17.5856
Percentage of students who are overweight
QNOVWGT
2
34558
15.5218
Percentage of students who were overweight
QNOWT
2
33539
15.0641
Percentage of students who were trying to lose weight
QN65
2
25631
11.5122
How much do you weigh without your shoes on?
Q7
2
23944
10.7545
Percentage of students who were trying to lose weight
QN65
1
14925
6.7036
How much do you weigh without your shoes on?
Q7
3
13398
6.0177
Percentage of students who were trying to lose weight
QN65
.
10818
4.8589
How much do you weigh without your shoes on?
Q7
<missing>
9219
4.1407
Percentage of students who were overweight
QNOWT
1
6471
2.9065
Percentage of students who are overweight
QNOVWGT
1
5499
2.4699
How much do you weigh without your shoes on?
Q7
54.43
4428
1.9888
How much do you weigh without your shoes on?
Q7
58.97
4389
1.9713
How much do you weigh without your shoes on?
Q7
63.5
3965
1.7809
How much do you weigh without your shoes on?
Q7
68.04
3872
1.7391
Percentage of students who are overweight
QNOVWGT
.
3115
1.3991
Percentage of students who were overweight
QNOWT
.
1140
0.512
Percentage of students who were trying to lose weight
QN67
.
1067
0.4792
Weight variable*
WEIGHT
5.8559
196
0.088
Weight variable*
WEIGHT
5.5349
140
0.0629
Weight variable*
WEIGHT
5.4651
128
0.0575
Weight variable*
WEIGHT
5.586273339
117
0.0526
Weight variable*
WEIGHT
0.069
114
0.0512
Weight variable*
WEIGHT
0.0231
113
0.0508
Weight variable*
WEIGHT
5.2648
108
0.0485
Weight_2
Weight_2
0.1884
67
0.0301
Weight_2
Weight_2
0.1955
60
0.0269
Weight_2
Weight_2
0.1195
54
0.0243
Weight_3
Weight_3
0.3813
52
0.0234
Weight_3
Weight_3
0.3042
51
0.0229
Weight_2
Weight_2
3.0334
51
0.0229
Weight_2
Weight_2
1.8008
50
0.0225
Weight_2
Weight_2
1.8685
50
0.0225
Weight_2
Weight_2
1.8942
48
0.0216
Weight_3
Weight_3
2.055
48
0.0216
Weight_3
Weight_3
0.4907
46
0.0207
Weight_3
Weight_3
2.135
44
0.0198
Weight_3
Weight_3
0.2068
41
0.0184
Weight_3
Weight_3
0.2747
40
0.018