> For the complete documentation index, see [llms.txt](https://docs.bcbi.brown.edu/codiac-for-health/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.bcbi.brown.edu/codiac-for-health/ursa-ri/datasets/cbioportal-for-cancer-genomics.md).

# cBioPortal for Cancer Genomics

## Overview

The cBioPortal for Cancer Genomics is an open-access, open-source platform for the interactive exploration of multidimensional cancer genomics datasets with an emphasis on multimodal cancer visualization. This resource aims to significantly lower the barriers between complex genomic data and cancer researchers by providing rapid, intuitive, and high-quality access to molecular profiles and clinical attributes from large-scale cancer genomics projects, empowering researchers, clinicians, and translational scientists to extract biological insights and create clinical applications from these rich datasets. cBioPortal provides visualization, analysis, and download from a wide range of datasets predominantly sourced from The Cancer Genome Atlas, with a growing number of institutional and consortium cohorts added over time. Data types span somatic mutations, structural variants and gene fusions, copy number alterations, mRNA and microRNA expression, DNA methylation, and protein/phosphoprotein levels (RPPA), with samples also linked to clinical and treatment attributes. The platform was originally developed at Memorial Sloan Kettering Cancer Center and is now maintained by a multi-institutional team consisting of MSK, the Dana-Farber Cancer Institute, Princess Margaret Cancer Centre, Children's Hospital of Philadelphia, Caris Life Sciences, The Hyve, SE4BIO, and Bilkent University.

## Methodology and Generation

cBioPortal was developed to address the data-integration challenges posed by large-scale cancer genomics projects, making raw data from these efforts more easily and directly available to the cancer research community, supporting both basic and translational research. Each study loaded into the portal is built from a standardized set of source files (clinical data tables, mutation files, copy-number segment files, expression matrices, etc.) that are validated and harmonized before import. cBioPortal generally does not generate primary experimental data; it hosts, curates, annotates, harmonizes, and visualizes data generated by external projects, consortia, publications, or institutional pipelines.&#x20;

Data in cBioPortal is organized into studies, each of which includes a study description file, clinical meta file, and a clinical data file. Before any study can be loaded into cBioPortal, all files must pass a validation phase to ensure they are following the correct content and formatting guidelines. Studies can be loaded incrementally for certain data types without re-uploading an entire study, and datasets that are subsets or combinations of existing studies can be represented as virtual studies to avoid duplication. For further information on how to load or download study data see the links section below.

## Standardized Healthcare Vocabularies

cBioPortal does not enforce a standardized clinical ontology across studies for patient and sample attributes. Instead, study authors of the published studies determine the organization of their clinical data. Although cBioPortal does not require a universal clinical ontology across studies, several standardized conventions are used for genomic data to support harmonization and cross-study analyses.

* **OncoTree (cancer type taxonomy)** — used as the hierarchical cancer classification system underlying cBioPortal. Each study specifies a `type_of_cancer` value, which maps to the portal's cancer type hierarchy; custom cancer type definitions can also be supplied when loading a study.
* **Reference genome assemblies (hg19/GRCh37, hg38/GRCh38)** — genomic datasets specify the reference genome assembly used for sequence alignment and genomic coordinates through study metadata (e.g., `reference_genome`). Segmented copy-number datasets additionally specify a `reference_genome_id` field to ensure consistent interpretation of genomic locations.
* **TCGA Mutation Annotation Format (MAF)** — provides the standardized file format and controlled `Variant_Classification` values (e.g., `Missense_Mutation`, `Nonsense_Mutation`, `Splice_Site`) used for importing somatic mutation data.
* **HGVS (Human Genome Variation Society) nomenclature** — used for standardized reporting of sequence and protein variants, including amino acid changes associated with genomic mutations.
* **HUGO Gene Nomenclature Committee (HGNC) gene symbols** — standardized gene symbols used in genomic datasets through the `Hugo_Symbol` field.
* **NCBI Entrez Gene identifiers** — stable numeric gene identifiers used through the `Entrez_Gene_Id` field. Depending on the data type, cBioPortal accepts HGNC symbols, Entrez Gene identifiers, or both to support consistent gene mapping across studies.

&#x20;For further information see the Links section below.

## Demographics

Demographic information is limited to just AGE and SEX/GENDER with native platform support. Other standard demographic fields are not enforced and vary by study/what the author chooses to include. There are no platform-wide demographic charactersitics published by cBioPortal. At the individual study level: demographic data including race and age IS displayed in Study View, but only when the contributing study authors included it in their clinical data file.

## References

cBioPortal for Cancer Genomics. Overview \[Internet]. New York (NY): Memorial Sloan Kettering Cancer Center; 2026 \[cited 2026 Jul 15]. Available from: <https://docs.cbioportal.org/user-guide/overview/>

## Resources

### Articles

* Cerami E, Gao J, Dogrusoz U, Gross BE, Sumer SO, Aksoy BA, Jacobsen A, Byrne CJ, Heuer ML, Larsson E, Antipin Y, Reva B, Goldberg AP, Sander C, Schultz N. [The cBio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data.](https://pubmed.ncbi.nlm.nih.gov/22588877/) Cancer Discov. 2012 May;2(5):401-4. doi: 10.1158/2159-8290.CD-12-0095. Erratum in: Cancer Discov. 2012 Oct;2(10):960. PMID: 22588877; PMCID: PMC3956037.
* Gao J, Aksoy BA, Dogrusoz U, Dresdner G, Gross B, Sumer SO, Sun Y, Jacobsen A, Sinha R, Larsson E, Cerami E, Sander C, Schultz N. [Integrative analysis of complex cancer genomics and clinical profiles using the cBioPortal.](https://pubmed.ncbi.nlm.nih.gov/23550210/) Sci Signal. 2013 Apr 2;6(269):pl1. doi: 10.1126/scisignal.2004088. PMID: 23550210; PMCID: PMC4160307.
* de Bruijn I, Kundra R, Mastrogiacomo B, Tran TN, Sikina L, Mazor T, Li X, Ochoa A, Zhao G, Lai B, Abeshouse A, Baiceanu D, Ciftci E, Dogrusoz U, Dufilie A, Erkoc Z, Garcia Lara E, Fu Z, Gross B, Haynes C, Heath A, Higgins D, Jagannathan P, Kalletla K, Kumari P, Lindsay J, Lisman A, Leenknegt B, Lukasse P, Madela D, Madupuri R, van Nierop P, Plantalech O, Quach J, Resnick AC, Rodenburg SYA, Satravada BA, Schaeffer F, Sheridan R, Singh J, Sirohi R, Sumer SO, van Hagen S, Wang A, Wilson M, Zhang H, Zhu K, Rusk N, Brown S, Lavery JA, Panageas KS, Rudolph JE, LeNoue-Newton ML, Warner JL, Guo X, Hunter-Zinck H, Yu TV, Pilai S, Nichols C, Gardos SM, Philip J; AACR Project GENIE BPC Core Team, AACR Project GENIE Consortium; Kehl KL, Riely GJ, Schrag D, Lee J, Fiandalo MV, Sweeney SM, Pugh TJ, Sander C, Cerami E, Gao J, Schultz N. [Analysis and Visualization of Longitudinal Genomic and Clinical Data from the AACR Project GENIE Biopharma Collaborative in cBioPortal.](https://pubmed.ncbi.nlm.nih.gov/37668528/) Cancer Res. 2023 Dec 1;83(23):3861-3867. doi: 10.1158/0008-5472.CAN-23-0816. PMID: 37668528; PMCID: PMC10690089.

### Links

* [cBioPortal Tutorial Slides](https://docs.cbioportal.org/user-guide/overview/)&#x20;
* [cBioPortal Data Loading Documentation](https://docs.cbioportal.org/data-loading/)&#x20;
* [cBioPortal File Formats Documentation](https://docs.cbioportal.org/file-formats/#cancer-study)

<br>

<br>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.bcbi.brown.edu/codiac-for-health/ursa-ri/datasets/cbioportal-for-cancer-genomics.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
