Data Dictionary
1. Discrete Copy Number Data
Hugo_Symbol
Required*
HUGO gene symbol (*one of Hugo_Symbol or Entrez_Gene_Id required)
Entrez_Gene_Id
Required*
Entrez Gene numeric identifier (*one of Hugo_Symbol or Entrez_Gene_Id required)
[SAMPLE_ID columns]
Required
A sample ID. This field can only contain numbers, letters, points, underscores and hyphens.
cbp_driver
Optional
Custom driver annotation: Putative_Driver, Putative_Passenger, Unknown, NA
cbp_driver_annotation
Optional
Free-text description of driver annotation (max 80 chars). *This field must be present if the cbp_driver is also present in the MAF file
cbp_driver_tiers
Optional
Driver tier label, e.g. 'Highly actionable' (max 20 chars). *This field must be present if the cbp_driver_tiers_annotation is also present in the MAF file
cbp_driver_tiers_annotation
Optional
Description of driver tier (max 80 chars). *This field must be present if the cbp_driver_tiers is also present in the MAF file
2. Continuous Copy Number Data
Hugo_Symbol
Required*
HUGO gene symbol (*one of Hugo_Symbol or Entrez_Gene_Id required)
Entrez_Gene_Id
Required*
Entrez Gene numeric identifier (*one of Hugo_Symbol or Entrez_Gene_Id required)
[SAMPLE_ID columns]
Required
One column per sample; continuous copy number or log2 value per gene-sample pair
3. Segmented Data (SEG)
Should follow this format. Link does not work from cBioPortal.
ID
Required
Sample identifier
chrom
Required
Index of the chromosome column
loc.start
Required
Index of the start position column
loc.end
Required
Index of the end position column
num.mark
Required
Index of a probe or description column
4. Expression Data (mRNA / microRNA)
Hugo_Symbol
Recommended*
HUGO gene symbol (*one of Hugo_Symbol or Entrez_Gene_Id required)
Entrez_Gene_Id
Recommended*
Entrez Gene numeric identifier (preferred; reduces ambiguity)
[SAMPLE_ID columns]
Required
One column per sample; real-number expression value or NA per gene-sample pair
5. Mutation Data (MAF)
Hugo_Symbol
Required
A HUGO gene symbol.
Entrez_Gene_Id
Recommended
A Entrez Gene numeric identifier.
Center
Optional
The sequencing center.
NCBI_Build
Required
Genome Reference Consortium Build used by variant calling software. Must be GRCh37 or GRCh38 for human, GRCm38 for mouse.
Chromosome
Required
A chromosome number, e.g. '7'.
Start_Position
Recommended
Start position of the event. Required for Cancer Hotspots annotations.
End_Position
Recommended
End position of the event. Required for Cancer Hotspots annotations.
Strand
Optional
Strand of the mutation. Assumed to be reported for the + strand.
Variant_Classification
Required
Translational effect of variant allele, e.g. Missense_Mutation, Nonsense_Mutation, Silent, Splice_Site, Frame_Shift_Del, etc. (TCGA MAF values).
Variant_Type
Optional
Variant type, e.g. SNP, DNP, INS, DEL, etc.
Reference_Allele
Required
The plus strand reference allele at this position.
Tumor_Seq_Allele1
Optional
Primary data genotype (allele 1).
Tumor_Seq_Allele2
Required
Primary data genotype (variant allele).
dbSNP_RS
Optional
Latest dbSNP rs ID for this position.
dbSNP_Val_Status
Optional
dbSNP validation status.
Tumor_Sample_Barcode
Required
Sample ID — either a TCGA barcode (patient ID extracted automatically) or a literal SAMPLE_ID from the clinical data file.
Matched_Norm_Sample_Barcode
Optional
Sample ID for the matched normal sample.
Match_Norm_Seq_Allele1
Optional
Primary data genotype for matched normal (allele 1).
Match_Norm_Seq_Allele2
Optional
Primary data genotype for matched normal (allele 2).
Tumor_Validation_Allele1
Optional
Secondary data from orthogonal technology (tumor allele 1).
Tumor_Validation_Allele2
Optional
Secondary data from orthogonal technology (tumor allele 2).
Match_Norm_Validation_Allele1
Optional
Secondary data from orthogonal technology (normal allele 1).
Match_Norm_Validation_Allele2
Optional
Secondary data from orthogonal technology (normal allele 2).
Verification_Status
Optional
Second pass results from independent attempt using same methods. Values: Verified, Unknown, or NA.
Validation_Status
Optional
Second pass results from orthogonal technology. Values: Valid, Invalid, Untested, Inconclusive, Redacted, Unknown, or NA.
Mutation_Status
Optional
Somatic or Germline are displayed in the Mutations tab. None, LOH, and Wildtype will not be loaded. Other values displayed as text.
Sequencing_Phase
Optional
Indicates current sequencing phase.
Sequence_Source
Optional
Molecular assay type used to produce the analytes used for sequencing.
Validation_Method
Optional
The assay platforms used for the validation call.
Score
Optional
Not used by cBioPortal.
BAM_File
Optional
Not used by cBioPortal.
Sequencer
Optional
Instrument used to produce primary data.
HGVSp_Short
Required
Amino acid change in HGVS notation, e.g. p.V600E.
t_alt_count
Optional
Variant allele count (tumor).
t_ref_count
Optional
Reference allele count (tumor).
n_alt_count
Optional
Variant allele count (normal).
n_ref_count
Optional
Reference allele count (normal).
cbp_driver
Optional
Custom driver annotation: Putative_Driver, Putative_Passenger, Unknown, NA, or empty.
cbp_driver_annotation
Optional
Free-text description of the driver annotation (max 80 chars).
cbp_driver_tiers
Optional
Driver tier label, e.g. 'Highly actionable' (max 20 chars).
cbp_driver_tiers_annotation
Optional
Description of the driver tier value (max 80 chars).
ASCN.ASCN_METHOD
Optional (ASCN)
Method used to obtain allele-specific copy number data, e.g. FACETS.
ASCN.CCF_EXPECTED_COPIES
Optional (ASCN)
Cancer-cell fraction if mutation exists on major allele.
ASCN.CCF_EXPECTED_COPIES_UPPER
Optional (ASCN)
Upper error bound for cancer-cell fraction estimate.
ASCN.EXPECTED_ALT_COPIES
Optional (ASCN)
Estimated number of copies harboring the mutant allele.
ASCN.CLONAL
Optional (ASCN)
Clonal status: Clonal, Subclonal, or Indeterminate.
ASCN.TOTAL_COPY_NUMBER
Optional (ASCN)
Total copy number of the gene.
ASCN.MINOR_COPY_NUMBER
Optional (ASCN)
Copy number of the minor allele.
ASCN.ASCN_INTEGER_COPY_NUMBER
Optional (ASCN)
Absolute integer copy-number estimate.
Site2_Region
Recommended
Region type: 5_Prime_UTR, 3_Prime_UTR, Promoter, Exon, or Intron
Site2_Chromosome
Recommended
Chromosome of gene 2
Site2_Position
Recommended
Genomic position of breakpoint at gene 2
NCBI_Build
Optional
Genome reference build: GRCh37 or GRCh38
Class
Optional
Deletion, Duplication, Insertion, Inversion, or Translocation
Event_Info
Optional
Free-text description of the event, e.g. 'TMPRSS2-ERG fusion'
Annotation
Optional
Free-text description of the gene/transcript rearrangement
Site1_Ensembl_Transcript_Id
Optional
Ensembl transcript ID of gene 1 (required for SV tab visualization)
Site2_Ensembl_Transcript_Id
Optional
Ensembl transcript ID of gene 2 (required for SV tab visualization)
DNA_Support
Optional
Yes or No — fusion detected from DNA data
RNA_Support
Optional
Yes or No — fusion detected from RNA data
SV_Length
Optional
Length of the structural variant in bases
Tumor_Split_Read_Count
Optional
Number of split reads supporting the call in tumor
Tumor_Paired_End_Read_Count
Optional
Number of paired-end reads supporting the call in tumor
Comments
Optional
Any free-text comments
6. Methylation Data
Hugo_Symbol
Required*
HUGO gene symbol (*one of Hugo_Symbol or Entrez_Gene_Id required)
Entrez_Gene_Id
Required*
Entrez Gene numeric identifier (*one of Hugo_Symbol or Entrez_Gene_Id required)
[SAMPLE_ID columns]
Required
One column per sample; methylation beta-value per gene-sample pair
7. Protein Level Data (RPPA / Mass Spectrometry)
Two-dimensional matrix: one row per antibody, one column per sample. Values are log2 protein expression levels or Z-scores.
Composite.Element.REF
Required
Antibody identifier encoding gene symbol(s)/Entrez ID(s) and antibody ID, e.g. 'BRAF|B-Raf-M-NA' or 'MAPK1 MAPK3|MAPK_PT202_Y204'
[SAMPLE_ID columns]
Required
One column per sample; real-number protein level per antibody-sample pair
8. Structural Variant Data (SV)
Sample_Id
Required
Sample identifier as defined in the clinical sample file
SV_Status
Required
SOMATIC or GERMLINE
Site1_Hugo_Symbol
Recommended
HUGO gene symbol of gene 1 (left/3' site)
Site1_Ensembl_Transcript_Id
Optional
Ensembl transcript ID of gene 1 (required for SV tab visualization)
Site1_Entrez_Gene_Id
Recommended
Entrez Gene identifier of gene 1
Site1_Region_Number
Recommended
Region number of Site 1, e.g. exon 2
Site1_Region
Recommended
Region type: 5_Prime_UTR, 3_Prime_UTR, Promoter, Exon, or Intron
Site1_Chromosome
Recommended
Chromosome of gene 1
Site1_Contig
Optional
The contig of Site 1
Site1_Position
Recommended
Genomic position of breakpoint at gene 1
Site1_Description
Optional
Description of this event at site 2. This could be the location of the 2nd breakpoint in case of a fusion event.
Site2_Hugo_Symbol
Recommended
HUGO gene symbol of gene 2 (right/5' site)
Site2_Ensembl_Transcript_Id
Optional
Ensembl transcript ID of gene 2 (required for SV tab visualization)
Site2_Entrez_Gene_Id
Recommended
Entrez Gene identifier of gene 2
Site2_Region_Number
Recommended
Region number of Site 2
Site2_Region
Recommended
Region type: 5_Prime_UTR, 3_Prime_UTR, Promoter, Exon, or Intron
Site2_Chromosome
Recommended
Chromosome of gene 2
Site2_Contig
Optional
The contig of Site 2
Site2_Position
Recommended
Genomic position of breakpoint at gene 2
Site2_Description
Optional
Description of this event at site 1. This could be the location of the 1st breakpoint in case of a fusion event.
Site2_Effect_On_Frame
Optional
The effect of frame reading in gene 2. Frame_shift or InFrame (free text)
NCBI_Build
Optional
Genome reference build: GRCh37 or GRCh38
Class
Optional
Deletion, Duplication, Insertion, Inversion, or Translocation
Tumor_Split_Read_Count
Optional
Number of split reads supporting the call in tumor
Tumor_Paired_End_Read_Count
Optional
Number of paired-end reads supporting the call in tumor
Event_Info
Optional
Free-text description of the event, e.g. 'TMPRSS2-ERG fusion'
Connection_Type
Optional
Which direction the connection is made
Breakpoint_Type
Optional
PRECISE or IMPRECISE which explain the resolution. Fill in PRECISE if the breakpoint resolution is known down to the base pair
Annotation
Optional
Free-text description of the gene/transcript rearrangement
DNA_Support
Optional
Yes or No — fusion detected from DNA data
RNA_Support
Optional
Yes or No — fusion detected from RNA data
SV_Length
Optional
Length of the structural variant in bases
Normal_Read_Count
Optional
The total number of reads of the normal tissue.
Tumor_Read_Count
Optional
The total number of reads of the tumor tissue.
Normal_Variant_Count
Optional
The number of reads of the normal tissue that have the variant/allele.
Tumor_Variant_Count
Optional
The number of reads of the tumor tissue that have the variant/allele.
Normal_Paired_End_Read_Count
Optional
The number of paired-end reads of the normal tissue that support the call.
Normal_Split_Read_Count
Optional
The number of split reads of the normal tissue that support the call.
Comments
Optional
Any free-text comments
9. Fusion Data (DEPRECATED — use Structural Variant Data instead)
This format is deprecated. New studies should use the Structural Variant (SV) format above.
Hugo_Symbol
Required
HUGO gene symbol
Entrez_Gene_Id
Required
Entrez Gene numeric identifier
Center
Required
Sequencing center
Tumor_Sample_Barcode
Required
Sample ID
Fusion
Required
Description of the fusion, e.g. 'TMPRSS2-ERG fusion'
DNA_support
Required
Fusion detected from DNA: yes or no
RNA_support
Required
Fusion detected from RNA: yes or no
Method
Required
Algorithm/tool used for fusion detection
Frame
Required
in-frame or frameshift
Fusion_Status
Optional
Assesment of mutation type: SOMATIC, GERMLINE, UNKNOWN, or empty.
10. Timeline Data
PATIENT_ID
Required
Patient identifier from the dataset
START_DATE
Required
Days from date of diagnosis (day 0) to event start
STOP_DATE
Required
Days from date of diagnosis to event end (blank if point-in-time event)
EVENT_TYPE
Required
Category of event: TREATMENT, LAB_TEST, IMAGING, STATUS, SPECIMEN, or any custom type
TREATMENT_TYPE
Optional
For TREATMENT events: Medical Therapy or Radiation Therapy
SUBTYPE
Optional
For TREATMENT events: Chemotherapy, Hormone Therapy, Targeted Therapy, etc.
AGENT
Optional
For TREATMENT events: agent name with number of cycles if applicable
TEST
Optional
For LAB_TEST events: type of test performed
RESULT
Optional
For LAB_TEST events: corresponding test result value
DIAGNOSTIC_TYPE
Optional
For IMAGING events: diagnostic tool used (MRI, CT scan, etc.)
STATUS
Optional
For STATUS events: best response or disease progression stage
SPECIMEN_SITE
Optional
For SPECIMEN events: site from which specimen was collected
SPECIMEN_TYPE
Optional
For SPECIMEN events: tissue or blood
STYLE_SHAPE
Optional
Render shape for this event: circle, square, triangle, diamond, star, or camera
STYLE_COLOR
Optional
Hexadecimal color value for rendering this event, e.g. #ffffff
Agent_Class
Optional
Suggested for TREATMENT events to classify agents into groups
Diagnostic_Type_Detailed
Optional
Suggested for IMAGING events as a detailed description
Source
Optional
Appears as a suggested column for IMAGING, STATUS, and SPECIMEN events
11. GISTIC 2.0 Data
chromosome
Required
Chromosome number (without 'chr' prefix)
peak_start
Required
Start coordinate of the region of maximal amplification or deletion
peak_end
Required
End coordinate of the region of maximal amplification or deletion
genes_in_region
Required
Comma-separated list of HUGO gene symbols in the wide peak
amp
Required
1 for amplification, 0 for deletion
cytoband
Required
Cytogenetic band specification including chromosome (Giemsa stain)
q_value
Required
Q-value (FDR-corrected p-value) of the peak region
12. MutSig Data
rank
Required
Gene rank by significance
gene
Required
HUGO gene symbol
N (or Nnon)
Required
Number of bases covered
n (or nnon)
Required
Number of mutations observed
p
Required
P-value: probability mutations are due to background processes
q
Required
Q-value: p-value corrected for multiple testing
13. Gene Panel Data (Gene Panel Matrix)
SAMPLE_ID
Required
Sample identifier
[stable_id columns]
Required
One column per genetic profile (e.g. 'mutations', 'gistic'); value is gene panel stable_id or NA if not profiled
14. Gene Set Data (GSVA Scores and P-values)
geneset_id
Required
Gene set name (uppercase); must match across score and p-value files
[SAMPLE_ID columns]
Required
Score file: GSVA score between -1.0 and 1.0, or NA. P-value file: p-value for the score
15. Generic Assay Data (incl. Arm-Level CNA, Mutational Signatures)
entity_stable_id
Required
Stable identifier for the assay entity (e.g. '1p', '1q', 'SBS1' for mutational signatures)
Name
Required
A column from generic_entity_meta_properties (using the property name as the column header)
Description
Required
Another generic_entity_meta_properties column
URL
Required
Another generic_entity_meta_properties column
[SAMPLE_ID columns]
Required
One column per sample; numeric value per entity-sample pair
Last updated
