For the complete documentation index, see llms.txt. This page is also available as Markdown.

cBioPortal for Cancer Genomics

Unless otherwise noted, information in this overview is adapted from the official cBioPortal documentation at docs.cbioportal.org. See references.

Overview

The cBioPortal for Cancer Genomics is an open-access, open-source platform for the interactive exploration of multidimensional cancer genomics datasets with an emphasis on multimodal cancer visualization. This resource aims to significantly lower the barriers between complex genomic data and cancer researchers by providing rapid, intuitive, and high-quality access to molecular profiles and clinical attributes from large-scale cancer genomics projects, empowering researchers, clinicians, and translational scientists to extract biological insights and create clinical applications from these rich datasets. cBioPortal provides visualization, analysis, and download from a wide range of datasets predominantly sourced from The Cancer Genome Atlas, with a growing number of institutional and consortium cohorts added over time. Data types span somatic mutations, structural variants and gene fusions, copy number alterations, mRNA and microRNA expression, DNA methylation, and protein/phosphoprotein levels (RPPA), with samples also linked to clinical and treatment attributes. The platform was originally developed at Memorial Sloan Kettering Cancer Center and is now maintained by a multi-institutional team consisting of MSK, the Dana-Farber Cancer Institute, Princess Margaret Cancer Centre, Children's Hospital of Philadelphia, Caris Life Sciences, The Hyve, SE4BIO, and Bilkent University.

Methodology and Generation

cBioPortal was developed to address the data-integration challenges posed by large-scale cancer genomics projects, making raw data from these efforts more easily and directly available to the cancer research community, supporting both basic and translational research. Each study loaded into the portal is built from a standardized set of source files (clinical data tables, mutation files, copy-number segment files, expression matrices, etc.) that are validated and harmonized before import. cBioPortal generally does not generate primary experimental data; it hosts, curates, annotates, harmonizes, and visualizes data generated by external projects, consortia, publications, or institutional pipelines.

Data in cBioPortal is organized into studies, each of which includes a study description file, clinical meta file, and a clinical data file. Before any study can be loaded into cBioPortal, all files must pass a validation phase to ensure they are following the correct content and formatting guidelines. Studies can be loaded incrementally for certain data types without re-uploading an entire study, and datasets that are subsets or combinations of existing studies can be represented as virtual studies to avoid duplication. For further information on how to load or download study data see the links section below.

Standardized Healthcare Vocabularies

cBioPortal does not enforce a standardized clinical ontology across studies for patient and sample attributes. Instead, study authors of the published studies determine the organization of their clinical data. Although cBioPortal does not require a universal clinical ontology across studies, several standardized conventions are used for genomic data to support harmonization and cross-study analyses.

  • OncoTree (cancer type taxonomy) — used as the hierarchical cancer classification system underlying cBioPortal. Each study specifies a type_of_cancer value, which maps to the portal's cancer type hierarchy; custom cancer type definitions can also be supplied when loading a study.

  • Reference genome assemblies (hg19/GRCh37, hg38/GRCh38) — genomic datasets specify the reference genome assembly used for sequence alignment and genomic coordinates through study metadata (e.g., reference_genome). Segmented copy-number datasets additionally specify a reference_genome_id field to ensure consistent interpretation of genomic locations.

  • TCGA Mutation Annotation Format (MAF) — provides the standardized file format and controlled Variant_Classification values (e.g., Missense_Mutation, Nonsense_Mutation, Splice_Site) used for importing somatic mutation data.

  • HGVS (Human Genome Variation Society) nomenclature — used for standardized reporting of sequence and protein variants, including amino acid changes associated with genomic mutations.

  • HUGO Gene Nomenclature Committee (HGNC) gene symbols — standardized gene symbols used in genomic datasets through the Hugo_Symbol field.

  • NCBI Entrez Gene identifiers — stable numeric gene identifiers used through the Entrez_Gene_Id field. Depending on the data type, cBioPortal accepts HGNC symbols, Entrez Gene identifiers, or both to support consistent gene mapping across studies.

For further information see the Links section below.

Demographics

Demographic information is limited to just AGE and SEX/GENDER with native platform support. Other standard demographic fields are not enforced and vary by study/what the author chooses to include. There are no platform-wide demographic charactersitics published by cBioPortal. At the individual study level: demographic data including race and age IS displayed in Study View, but only when the contributing study authors included it in their clinical data file.

References

cBioPortal for Cancer Genomics. Overview [Internet]. New York (NY): Memorial Sloan Kettering Cancer Center; 2026 [cited 2026 Jul 15]. Available from: https://docs.cbioportal.org/user-guide/overview/

Resources

Articles

  • Cerami E, Gao J, Dogrusoz U, Gross BE, Sumer SO, Aksoy BA, Jacobsen A, Byrne CJ, Heuer ML, Larsson E, Antipin Y, Reva B, Goldberg AP, Sander C, Schultz N. The cBio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data. Cancer Discov. 2012 May;2(5):401-4. doi: 10.1158/2159-8290.CD-12-0095. Erratum in: Cancer Discov. 2012 Oct;2(10):960. PMID: 22588877; PMCID: PMC3956037.

  • Gao J, Aksoy BA, Dogrusoz U, Dresdner G, Gross B, Sumer SO, Sun Y, Jacobsen A, Sinha R, Larsson E, Cerami E, Sander C, Schultz N. Integrative analysis of complex cancer genomics and clinical profiles using the cBioPortal. Sci Signal. 2013 Apr 2;6(269):pl1. doi: 10.1126/scisignal.2004088. PMID: 23550210; PMCID: PMC4160307.

  • de Bruijn I, Kundra R, Mastrogiacomo B, Tran TN, Sikina L, Mazor T, Li X, Ochoa A, Zhao G, Lai B, Abeshouse A, Baiceanu D, Ciftci E, Dogrusoz U, Dufilie A, Erkoc Z, Garcia Lara E, Fu Z, Gross B, Haynes C, Heath A, Higgins D, Jagannathan P, Kalletla K, Kumari P, Lindsay J, Lisman A, Leenknegt B, Lukasse P, Madela D, Madupuri R, van Nierop P, Plantalech O, Quach J, Resnick AC, Rodenburg SYA, Satravada BA, Schaeffer F, Sheridan R, Singh J, Sirohi R, Sumer SO, van Hagen S, Wang A, Wilson M, Zhang H, Zhu K, Rusk N, Brown S, Lavery JA, Panageas KS, Rudolph JE, LeNoue-Newton ML, Warner JL, Guo X, Hunter-Zinck H, Yu TV, Pilai S, Nichols C, Gardos SM, Philip J; AACR Project GENIE BPC Core Team, AACR Project GENIE Consortium; Kehl KL, Riely GJ, Schrag D, Lee J, Fiandalo MV, Sweeney SM, Pugh TJ, Sander C, Cerami E, Gao J, Schultz N. Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal. Cancer Res. 2023 Dec 1;83(23):3861-3867. doi: 10.1158/0008-5472.CAN-23-0816. PMID: 37668528; PMCID: PMC10690089.

Last updated