DriverDBv5: A database for human cancer driver gene research



What is DriverDB?

DriverDB is an integrative cancer omics database that combines somatic mutation, RNA expression, miRNA expression, protein expression, methylation, copy number variation (CNV), and clinical data with curated annotations and published bioinformatics algorithms for driver gene and driver event identification. Featured in the 2014, 2016, 2020, and 2024 Nucleic Acids Research Database Issues, DriverDB applies state-of-the-art computational methods to characterize cancer drivers across molecular layers.

DriverDB provides three major analytical modules:
  • Cancer – Summarizes driver gene predictions for a selected cancer type across multiple omics layers using published driver identification tools.
  • Gene – Visualizes multi-omics features of a user-selected gene, including differential expression, mutation, CNV, methylation, survival, miRNA regulation, protein expression, and integrated multi-omics evidence.
  • Customized Analysis – Allows users to perform subgroup comparisons, survival analyses, multi-omics driver exploration, prognostic signature construction, and multivariate Cox modeling based on user-defined clinical or molecular criteria.

1. Cancer

1.1 Cancer Module Overview

The Cancer module summarizes driver gene and driver event predictions for a user-selected cancer type by integrating multi-omics data — including somatic mutations, copy number variation (CNV), methylation, RNA expression, miRNA expression, and clinical information — through published bioinformatics algorithms and curated annotation sources. This module provides a cancer-centric overview of dysregulated molecular features and highlights candidate driver genes, their regulatory mechanisms, and their functional significance across molecular layers.

For mutation, CNV, and methylation, and RNA, a Survival Relevance tab evaluates whether identified driver genes are associated with patient survival using multiple analysis methods, including Cox regression, cure model, and machine learning-based approaches. For multi-omics analysis, additional machine learning results are provided, including prognostic signature identification, Kaplan–Meier survival plots, predictive performance plots, and a gene-level summary of survival associations across omics types, endpoints, and algorithms.


1.2 Dataset Selection: Browse by Cancer Type

The DriverDBv5 Cancer interface currently provides 104 cancer projects for selection. Available projects are selected based on the availability of sufficient analysis results for visualization and interpretation.

A. Tissue Type (Optional)

Filter available datasets by tissue origin to quickly locate cancers related to a specific anatomical site.
For example, selecting Lung narrows the list to datasets such as:
  • Lung Adenocarcinoma (TCGA-US)
  • Lung Squamous Cell Carcinoma (TCGA-US)
  • Lung Cancer – KR (ICGC-KR)

B. Related Dataset

Select the specific cancer dataset you wish to analyze. Each dataset label includes its data source (e.g., TCGA-US, ICGC-KR), allowing users to choose cohorts most relevant to their research.

C. Submit

After making your selections, click Submit to load driver gene summaries and molecular features for the chosen cancer type. All downstream tabs, including Mutation, CNV, Methylation, Survival, miRNA, and Multi-Omics, will display results based on the selected dataset.


1.3 Overview of Result Tabs

The Cancer module contains several result tabs, each summarizing driver evidence derived from a different omics layer:
  • Summary – integrates dysfunction and dysregulation evidence across omics layers to highlight candidate driver genes and miRNA drivers for the selected cancer type, visualized through an interactive network.
  • Mutation – identifies mutation-based driver genes using multiple detection tools, and evaluates their association with patient survival through the Survival Relevance tab.
  • CNV – visualizes driver genes with significant copy number gain or loss, including CNV–expression relationships, and evaluates their association with patient survival through the Survival Relevance tab.
  • Methylation – highlights hypermethylation and hypomethylation driver genes and locus enrichment distributions, and evaluates their association with patient survival through the Survival Relevance tab.
  • RNA – identifies expression-based candidate driver genes based on differential expression and evaluates their association with patient survival through the Survival Relevance tab.
  • miRNA – identifies expression-based candidate driver genes based on differential expression and evaluates their association with patient survival through the Survival Relevance tab.
  • Multi-Omics – integrates multiple omics layers to identify cross-omics driver genes and functional patterns, and provides machine learning-based prognostic signature identification, Kaplan–Meier survival plots, predictive performance plots, and a gene-level summary of survival associations across omics types, endpoints, and algorithms.

1.4 Cancer Summary

1.4.1 Overview

The Cancer Summary tab provides an integrated overview of cancer driver genes and miRNA-mediated regulation in the selected cancer type by combining driver evidence across multiple omics levels, including mutation, CNV, methylation, RNA expression, and miRNA regulation.

This section contains two main components:
  1. Summary Network
  2. Driver Summary Table

The Summary Network visualizes driver genes and their molecular interactions, while the Driver Summary Table summarizes driver and dysregulation evidence for individual genes.

1.4.2 Summary Network

Purpose

The Summary Network provides an integrated visualization of driver genes and their molecular relationships in the selected cancer type. It combines driver and dysregulation evidence across RNA expression, mutation, CNV, and methylation, together with miRNA-mediated regulation and gene–gene interactions.

Nodes

Driver gene nodes are displayed as circular nodes divided into four quadrants, each corresponding to an omics feature:
  • Upper left: RNA expression status — upregulated or downregulated
  • Upper right: Mutation driver status
  • Lower left: Methylation driver status — hypermethylated or hypomethylated
  • Lower right: CNV driver status — gain or loss

Each quadrant is colored when the corresponding molecular alteration is identified; otherwise, it remains white. The overall appearance of the node therefore reflects the combination of molecular alterations identified for the gene. A red star indicates a gene identified by multi-omics tools. Detailed color definitions are provided in the Gene Node Legend.

miRNAs are represented as yellow diamond-shaped nodes and can be included through the Regulatory Layer filter.


Edges

Edges between gene nodes represent protein–protein interactions (PPIs) or cross-omics synergistic effects. Edge colors indicate the corresponding omics combination for synergistic effects, as shown in the Edge Legend.

miRNA–gene edges represent regulatory relationships between miRNAs and their associated genes.


Network Node Selection

When the number of eligible nodes is large, genes are prioritized using the Weighted Evidence Score (WES). The WES integrates evidence from multiple molecular layers, including mutation, CNV, RNA expression, methylation, miRNA regulation, and multi-omics evidence. Genes receive higher scores when they are supported by more mutation detection tools, associated with more miRNAs, or supported by evidence across multiple omics layers, thereby prioritizing genes with stronger and more diverse driver evidence.

Starting from higher-scoring genes, connected neighboring nodes are iteratively included to construct a subnetwork of approximately 50 nodes. Because connected neighbors are added together during network expansion, the final number of displayed nodes may slightly exceed 50.

Interaction Guide

The Summary Network is interactive:

Selecting and Highlighting
  • Click a node to highlight its connected nodes and edges.
  • Click blank space to clear the selection and restore the full network view.
  • Use the dropdown menu to select a specific gene.
  • Rearrange recalculates the network layout to reorganize node positions without changing the displayed nodes or interactions.
Filtering Options
  • Gene Source: Select CGC, NCG, or all genes.
  • Gene Driver Evidence:
  • Select mutation drivers, CNV drivers, methylation drivers, and/or RNA dysregulation.
  • Regulatory Layer:
  • Control whether the miRNA network is included.
  • Gene–Gene Interaction Type:
  • Display PPIs, synergistic effects, or both.

1.4.3 Driver Summary Table

The Driver Summary Table provides an integrated summary of driver and dysregulation evidence for genes identified in the selected cancer type. It combines cancer gene annotations with evidence from mutation, CNV, methylation, RNA expression, miRNA regulation, and multi-omics analyses.

Columns:
  • Cancer Project: Cancer project in which the gene is identified.
  • ENSG: Ensembl gene identifier.
  • GENE: Official gene symbol.
  • CGC: Indicates whether the gene is included in the Cancer Gene Census (CGC).
  • NCG: Indicates whether the gene is included in the Network of Cancer Genes (NCG).
  • Mutation: Number of mutation driver detection tools supporting the gene.
  • CNV: CNV driver status, shown as gain or loss.
  • Methylation: Methylation driver status, shown as hyper or hypo.
  • RNA: RNA dysregulation status, shown as Upregulated or Downregulated.
  • Multi-omics: Indicates whether the gene is supported by multi-omics integration analysis.
  • miRNA: miRNAs associated with the gene through miRNA-mediated regulation; multiple miRNAs are listed when applicable.

1.5 Cancer Mutation

1.5.1 Overview

The Cancer Mutation section identifies and visualizes mutation-based driver genes and their survival relevance in the selected cancer type. Results are organized into two tabs: Driver Genes and Survival Relevance.

The Driver Genes tab focuses on mutation driver genes identified by multiple published computational tools. Mutation driver genes are defined as genes supported by at least three tools. The degree of consensus across tools provides a measure of confidence in each gene's driver role.
This tab contains two components:
  1. Mutation Driver Summary by Tools — summarizes how many genes are identified by varying numbers of mutation driver-detection tools and lists tool support counts for each driver gene.
  2. Mutation Profiles of Top 30 Driver Genes — visualizes mutation patterns, impact levels, and tool support for the top 30 mutation driver genes ranked by tool support across the patient cohort.
The Survival Relevance tab evaluates the prognostic relevance of mutation driver genes in the selected cancer type.
  1. Survival Gene Distribution Summary — bar charts and Venn diagrams summarizing the number and overlap of survival-related driver genes across four survival endpoints and four survival analysis methods.
  2. Survival Gene Summary Table — lists survival-related driver genes with their survival associations across endpoints and analysis methods, including log-transformed hazard ratios (log HRs) and machine learning identification status.
  3. Trans-Omics Synergistic Effect — evaluates whether pairs of genes or molecular features show combined survival effects, identifying cross-omics interactions where the combined hazard ratio exceeds that of either individual feature alone.

Together, the Driver Genes and Survival Relevance tabs help users identify which genes are supported as mutation drivers by computational tools, which driver genes are associated with patient survival, and which show synergistic survival effects in combination with other molecular features.

1.5.2 Driver genes

Mutation Driver Summary by Tools

Purpose

This panel summarizes how many genes are identified by varying numbers of mutation driver–detection tools.
Stronger consensus across tools indicates stronger evidence supporting a gene’s driver role.

Components

Distribution of Mutation Driver Genes by Tool Support (Left Plot)
  • Displays a bar plot showing the number of mutation driver genes supported by 3 or more tools.
  • Each bar represents the number of driver genes supported by a specific number of tools.
  • Higher bars at larger tool counts indicate stronger multi-tool agreement.

Mutation Summary Table (Right Table)

  • Located to the right of the plot.
  • Lists the tool support count for each mutation driver gene.
    Details of the mutation driver-detection tools are provided in FAQ4.
  • The plot on the left is derived from this table.

Mutation Profiles of Top 30 Driver Genes

Purpose

This section visualizes mutation patterns for the top 30 mutation driver genes ranked by tool support, helping users examine:
  • Mutation burden per gene
  • Mutation impact distribution
  • How mutations are distributed across samples
  • Multi-tool support for each top gene

It contains two interactive components.

Components

Mutation Impact Distribution of Top 30 Driver Genes (Left Plot)
The plot displays mutation data across the top 30 driver genes, with each row representing a different driver gene and each column representing an individual sample. Each cell within the plot indicates whether that particular sample carries a mutation in the corresponding gene, and if so, the predicted impact level of that mutation—categorized as High, Moderate, Low, or Modifier..
Additional Elements:
  • Left panel (A): total mutation percentage per gene.
  • Top bar chart (B): total mutation count per sample.
  • Right bar chart (C): total mutation count per gene

Tool Support for Top 30 Driver Genes (Right Plot)

The plot displays a bar chart where each bar represents a gene, with the height of the bar indicating the number of mutation tools that identified that gene as a mutation driver. Genes that are supported by a greater number of tools suggest higher-confidence driver roles, as consensus across multiple computational methods provides stronger evidence for their functional importance in cancer development.

1.6 Cancer CNV

1.6.1 Overview

The Cancer CNV section identifies and visualizes copy number variation (CNV) driver genes and their survival relevance in the selected cancer type. Results are organized into two tabs: Driver Genes and Survival Relevance.

The Driver Genes tab focuses on identifying genes with significant copy number gain or loss based on statistical significance, sample proportion, and copy number fold change. This tab contains three components:
  1. Visualization of Top 30 CNV Driver Genes — visualizes CNV gain and loss patterns of the top 30 CNV driver genes across the patient cohort.
  2. Locus Enrichment — summarizes the chromosomal distribution and enriched chromosomal loci of CNV driver genes.
  3. CNV Driver Gene Summary Table — lists CNV driver genes with detailed statistics for copy number gain and loss, sample proportions, fold changes, and CNV–expression correlations.
The Survival Relevance tab evaluates the prognostic relevance of CNV driver genes across multiple survival endpoints and analysis methods. This tab contains three components:
  1. Survival Gene Distribution Summary — bar charts and Venn diagrams summarizing the number and overlap of survival-related driver genes across four survival endpoints and four survival analysis methods.
  2. Survival Gene Summary Table — lists survival-related driver genes and their survival associations across endpoints and analysis methods.
  3. Trans-Omics Synergistic Effect — evaluates whether pairs of genes or molecular features show combined survival effects, identifying cross-omics interactions where the combined hazard ratio exceeds that of either individual feature alone.

Together, the Driver Genes and Survival Relevance tabs help users explore CNV driver genes, their copy number alteration patterns and chromosomal distribution, and their associations with patient survival.

1.6.2 Driver genes

Visualization of Top 30 CNV Driver Genes

This panel presents CNV gain, loss, and neutral patterns for the top 30 CNV driver genes in the selected cancer type.

CNV Gain and Loss Distribution of Top 30 Genes (Top Chart)

The bar chart summarizes the percentage of samples exhibiting CNV gain, CNV loss, or no CNV alteration for each of the top 30 CNV driver genes. Each bar is color-coded to indicate CNV gain (red), CNV loss (blue), and no CNV alteration (gray). Hover over each bar segment to view the exact percentage for each CNV state.

CNV Patterns of Top 30 Genes Across Cancer Samples (Bottom Heatmap)

The heatmap visualizes CNV status across the top 30 CNV driver genes and cancer samples. Rows represent genes and columns represent samples, with each cell indicating CNV gain (red), CNV loss (blue), or no CNV alteration (gray). The left panel (A) shows the total percentage of CNV occurrences for each gene, the top bar chart (B) shows the total number of CNV occurrences per sample, and the right bar chart (C) shows the total number of CNV occurrences per gene.

Locus Enrichment

This section shows the chromosomal distribution of CNV driver genes and identifies significantly enriched chromosomal loci.

Chromosomal Locus Enrichment of CNV Driver Genes (Left Plot)

The plot displays the chromosomal positions of CNV driver genes with available genomic position information. Each red dot represents a gene and its corresponding genomic location. Hovering over a dot reveals the chromosome, genomic position, correlation value, and gene symbol. The correlation value reflects the relationship between CNV and RNA expression: positive values indicate that higher copy number is associated with higher expression, whereas negative values indicate an inverse relationship.

Locus Enrichment Summary Table (Right Table)

The table lists significantly enriched chromosomal loci among the CNV driver genes. The −log10(p-value) indicates the statistical significance of each enriched locus, and the associated genes are listed for each locus.

CNV Driver Gene Summary Table

This table provides detailed statistics for each CNV driver gene, including significance measures for CNV gain and loss, sample proportions, copy number fold changes, and CNV–expression correlations.

Columns:
  • Gene symbol / ENSG: Gene symbol and Ensembl gene identifier.
  • Gain/Loss p-value: Statistical significance of copy number gain or loss.
  • Gain/Loss FDR: FDR-adjusted p-value for copy number gain or loss.
  • Gain/Loss sample proportion: Proportion of samples exhibiting copy number gain or loss.
  • Normal sample proportion: Proportion of samples without copy number gain or loss.
  • Gain/Loss log2FC: Log2 fold change associated with copy number gain or loss.
  • DIGGIT Spearman p-value: Statistical significance of the CNV–expression association evaluated by DIGGIT.
  • Spearman correlation coefficient / p-value: Spearman correlation coefficient between CNV and gene expression and its corresponding p-value.
  • Max correlation coefficient / p-value: Maximum CNV–expression correlation coefficient and its corresponding p-value.

1.7 Cancer Methylation

1.7.1 Overview

The Cancer Methylation section visualizes methylation driver genes in the selected cancer type and evaluates whether gene-level methylation status is associated with patient survival. Results are organized into two tabs: Driver Genes and Survival Relevance.

The Driver Genes tab provides an overview of methylation driver genes across patient samples and chromosomal locations, helping users explore methylation patterns, methylation–expression relationships, and the genomic distribution of methylation alterations. Methylation driver genes are defined by MethylMix, while available ELMER results provide additional probe-level, statistical, and gene-expression correlation information. This tab contains three components:
  1. Visualization of Top 30 Methylation Driver Genes — displays hypermethylation and hypomethylation patterns for the top 30 methylation driver genes across the patient cohort.
  2. Locus Enrichment — summarizes the chromosomal distribution and enriched chromosomal loci of methylation driver genes.
  3. Methylation Driver Gene Summary Table — lists MethylMix-defined driver genes with methylation proportions and available ELMER probe information, statistical significance, and gene-expression correlation information.
The Survival Relevance tab evaluates the prognostic relevance of methylation driver genes across multiple survival endpoints and analysis methods. This tab contains three components:
  1. Survival Gene Distribution Summary — bar charts and Venn diagrams summarizing the number and overlap of survival-related driver genes across four survival endpoints and four survival analysis methods.
  2. Survival Gene Summary Table — lists survival-related driver genes and their survival associations across endpoints and analysis methods.
  3. Trans-Omics Synergistic Effect — evaluates whether pairs of genes or molecular features show combined survival effects, identifying cross-omics interactions where the combined hazard ratio exceeds that of either individual feature alone.

Together, the Driver Genes and Survival Relevance tabs help users explore methylation driver genes, their methylation patterns and chromosomal distribution, and their associations with patient survival.

1.7.2 Driver genes

Visualization of Top 30 Methylation Driver Genes

This panel summarizes the methylation status of the top 30 methylation driver genes and shows how hypermethylation and hypomethylation patterns are distributed across cancer samples.

Methylation Status of Top 30 Genes (Top Bar Chart)

The bar chart summarizes the proportion of samples showing hypermethylation, hypomethylation, or no methylation alteration for each of the top 30 methylation driver genes. Red represents hypermethylation, blue represents hypomethylation, and gray represents no methylation alteration. Hover over each bar to view the exact methylation percentages for each gene.

Methylation Patterns Across Cancer Samples (Bottom Heatmap)

The heatmap displays methylation profiles of the top 30 methylation driver genes across cancer samples. Rows represent genes and columns represent samples, with each cell indicating hypermethylation (red), hypomethylation (blue), or no methylation alteration (gray). The left panel (A) shows the total methylation percentage for each gene, the top bar chart (B) shows the total number of methylation events per sample, and the right bar chart (C) shows the total number of methylation events per gene.

Locus Enrichment

This section maps methylation driver genes to their chromosomal positions and summarizes significantly enriched chromosomal loci.

Chromosomal Locus Enrichment of Methylation Driver Genes (Left Plot)

The plot displays the chromosomal positions of methylation driver genes with available genomic position information. Each red dot represents a gene and its corresponding genomic location. Hover over a dot to view the chromosome, genomic position, gene symbol, and correlation value. The correlation value represents the relationship between methylation level and RNA expression: positive values indicate that higher methylation is associated with higher expression, whereas negative values indicate an inverse relationship.

Locus Enrichment Summary Table (Right Table)

The table lists significantly enriched chromosomal loci among the methylation driver genes.

Methylation Driver Gene Summary Table

This table lists methylation-related statistics for each MethylMix-defined driver gene in the selected cancer type. It summarizes methylation proportions together with available ELMER probe information, statistical significance, and correlations with gene expression.

Columns:
  • Gene symbol / ENSG: Gene symbol and Ensembl gene identifier.
  • Hyper / Hypo / None percent: Proportion of samples classified by MethylMix as hypermethylated, hypomethylated, or without a methylation alteration.
  • Probe: ELMER-associated methylation probe identifier.
  • Distance: Genomic distance between the ELMER-associated probe and the gene.
  • Adjust p-value: Adjusted significance value for the methylation event.
  • Methylation type: Classification of the methylation pattern (e.g., hyper or hypo) based on the available ELMER result.
  • Spearman correlation coefficient / Spearman p-value: Spearman correlation between methylation level and gene expression and its corresponding p-value.

1.8 Cancer RNA

1.8.1 Overview

The Cancer RNA section provides a comprehensive analysis of gene expression alterations in a selected cancer type. It identifies RNA driver genes based on differential expression and integrates evidence from other omics data types to characterize their potential roles in cancer. The section also evaluates the prognostic relevance of RNA driver genes through downstream survival analyses. Results are organized into two tabs: Driver Genes and Survival Relevance.

The Driver Genes tab identifies and visualizes RNA driver genes in the selected cancer type. Genes with significant differential expression are identified based on expression level, fold change, and statistical significance. Evidence from CNV, methylation, and miRNA analyses is integrated to indicate whether each RNA driver gene is also supported by other omics data types. This tab contains two components:
  1. Visualization of RNA Driver Genes — displays RNA driver genes using MA and volcano plots and indicates whether each gene is also supported by CNV, methylation, or miRNA evidence.
  2. RNA Driver Gene Summary Table — lists RNA driver genes with differential-expression statistics and integrated multi-omics driver evidence.
The Survival Relevance tab evaluates the prognostic relevance of RNA driver genes across multiple survival endpoints and analysis methods. This tab contains three components:
  1. Survival Gene Distribution Summary — bar charts and Venn diagrams summarizing the number and overlap of survival-related driver genes across four survival endpoints and four survival analysis methods.
  2. Survival Gene Summary Table — lists survival-related driver genes and their survival associations across endpoints and analysis methods.
  3. Trans-Omics Synergistic Effect — evaluates whether pairs of genes or molecular features show combined survival effects, identifying cross-omics interactions where the combined hazard ratio exceeds that of either individual feature alone.

Together, the Driver Genes and Survival Relevance tabs help users explore RNA driver genes, their differential-expression patterns and multi-omics support, and their associations with patient survival.

1.8.2 Driver genes

Visualization of RNA Driver Genes

This panel visualizes RNA driver genes based on their expression levels, fold changes, and statistical significance, while indicating whether each RNA driver is also supported by other omics evidence.

RNA Driver MA Plot

The MA plot displays RNA driver genes according to their mean expression level and differential expression. The x-axis represents log10 mean expression, calculated using the RNA expression measure defined for the selected dataset, and the y-axis represents log2 fold change. The dashed horizontal lines at log2 fold change = −1 and 1 indicate the fold-change thresholds used for RNA driver identification.

RNA Driver Volcano Plot

The volcano plot displays RNA driver genes according to their differential expression and statistical significance. The x-axis represents log2 fold change, and the y-axis represents −log10 adjusted p-value. The dashed vertical lines at log2 fold change = −1 and 1 indicate the fold-change thresholds used for RNA driver identification.

In both plots, each point represents an RNA driver gene. Colors indicate whether the RNA driver is also supported by other omics evidence, including CNV, methylation, and miRNA. Hover over a point to view detailed information for the corresponding gene. Click a legend item to show or hide the corresponding driver omics group.


RNA Driver Gene Summary Table

The RNA Driver Gene Summary Table lists differential-expression statistics and multi-omics evidence for RNA driver genes in the selected cancer type.

Columns:
  • Project: Cancer dataset.
  • Gene: Gene symbol.
  • Driver omics: Omics data types supporting the gene as a driver, including RNA and available CNV, methylation, or miRNA evidence.
  • Reference group / Comparison group: Sample groups used for differential-expression analysis.
  • Detection tool: Method used for differential-expression analysis.
  • log10 Mean Expression: Log10-transformed mean gene expression, calculated using the RNA expression measure defined for the selected dataset.
  • log10 Expression Cutoff: Expression threshold used for RNA driver identification.
  • log2 Fold Change: Log2 fold change between the comparison and reference groups.
  • P-value: P-value from the differential-expression analysis.
  • Adjusted P-value: Multiple-testing-adjusted p-value from the differential-expression analysis.

1.9 Cancer miRNA

1.9.1 Overview

The Cancer miRNA section analyzes regulatory relationships between differentially expressed (DE) miRNAs and their target genes in the selected cancer type. It integrates experimentally validated and computationally predicted miRNA–gene interactions with expression correlation information to characterize potential miRNA-mediated regulatory relationships in cancer.

The section consists of three main components:
  1. miRNA–Gene Interaction Network
  2. Heatmap of Differentially Expressed Genes and miRNAs
  3. Gene–miRNA Correlation Summary Table

1.9.2 miRNA-Gene Interaction Network

Purpose

The miRNA–Gene Interaction Network displays validated and predicted interactions between miRNAs and their target genes in the selected cancer type. The network enables users to explore miRNA–gene regulatory relationships and identify interactions supported by experimental validation, computational prediction, or both.

Nodes

  • Green nodes represent genes.
  • Yellow diamond-shaped nodes represent miRNAs.
  • Node size reflects network connectivity, with larger nodes representing features connected to more interaction partners.

Edges

Two types of miRNA–gene interactions are shown:
  1. Validated interactions
    • Experimentally supported miRNA–target interactions recorded in miRTarBase
    • Displayed as solid lines.
  2. Predicted interactions
    • Computationally predicted miRNA–target interactions supported by miRNA target prediction tools.
    • Displayed as dotted lines.
    • Users can specify the minimum prediction support, including any predicted interaction(>0), ≥6, ≥8, or ≥10 tools.

When both Predicted interactions and Validated interactions are selected, interactions satisfying either criterion are included.

Interaction Guide

  • Click a node to highlight its connected interaction partners.
  • Use the gene dropdown and Search button to locate and highlight a specific gene.
  • Rearrange recalculates the network layout without changing the selected interaction set.

Analysis Filters

Users can refine the displayed interaction set using:
  • Gene Source: CGC, NCG, or All
  • Interaction Type: Predicted interactions, Validated interactions, or both.
  • Prediction Support: Any predicted interaction, ≥6, ≥8, or ≥10 prediction tools.

Clicking Apply Filters updates the interaction network, heatmap, and correlation summary table using the selected criteria.


1.9.3 Heatmap of Differentially Expressed Genes and miRNAs

Purpose

The Heatmap of Differentially Expressed Genes and miRNAs visualizes expression profiles of DE genes and miRNAs across tumor (TP) and normal (NT) samples in the selected cancer type. The genes and miRNAs included in the heatmap are based on the interaction set defined by the Analysis Filters.

Gene expression values are based on the RNA expression type defined for the selected dataset, while miRNA expression values are based on RPKM for TCGA datasets and TPM for other datasets. Expression values are standardized for heatmap visualization to show relative expression patterns across samples.

Heatmap Display

  • Columns represent individual samples.
  • Rows represent DE genes and/or DE miRNAs included in the selected interaction set.
  • Red indicates higher relative expression, whereas blue indicates lower relative expression.
  • TP represents tumor samples and NT represents normal samples.
  • Dendrograms indicate similarity among samples and expression features based on their expression profiles.

Visualization Modes

The Visualization by panel allows users to switch among:
  • DE miRNA: displays DE miRNAs.
  • DE gene: displays DE genes.
  • DE miRNA + DE gene: displays both simultaneously.

These visualization modes change which features are displayed in the heatmap without changing the interaction criteria selected in the Analysis Filters.

Interpretation

The heatmap can be used to examine expression patterns of interacting miRNAs and genes across tumor and normal samples. Opposing expression patterns between a miRNA and its target gene may be consistent with a potential repressive relationship; however, expression patterns alone do not establish a direct regulatory mechanism. Correlation statistics and interaction evidence can be examined in the Gene–miRNA Correlation Summary Table.

1.9.4. Gene–miRNA Correlation Summary Table

The Gene–miRNA Correlation Summary Table summarizes gene–miRNA interactions that meet the criteria defined by the Analysis Filters. It provides validation evidence, prediction-tool support, and expression correlation statistics for each interaction.

The table includes the following columns:
  • Project: Cancer dataset.
  • miRNA / Gene / ENSG: Identifiers for the miRNA and its target gene.
  • Validated: In-house categorical score derived from miRTarBase miRNA–target evidence; higher scores indicate stronger support for the miRNA–gene interaction. https://mirtarbase.cuhk.edu.cn
  • Number of tools: Number of prediction tools supporting the interaction. See FAQ4 for descriptions of the prediction algorithms/tools.
  • Pearson / Spearman / Kendall correlation coefficients and p-values: Expression correlation between the miRNA and gene across samples and the corresponding statistical significance.

Negative correlations indicate that higher miRNA expression is associated with lower expression of the target gene, which is consistent with a potential repressive regulatory relationship. Correlation, however, does not by itself establish direct miRNA-mediated regulation and should be interpreted together with validation and prediction evidence.


1.10 Cancer Multi-omics

1.10.1 Overview

The Cancer Multi-Omics section visualizes driver evidence identified through multi-omics integration tools and evaluates survival-related results for the selected cancer type. By integrating results across mutation, CNV, methylation, mRNA expression, and miRNA regulation, this section summarizes driver evidence across molecular layers and provides views of associated biological functions, tool support, omics distributions, and prognostic signatures.

Users may filter results by gene set:
  • All: includes all identified multi-omics driver genes
  • CGC: includes only genes listed in the Cancer Gene Census
  • NCG 6.0: includes only genes listed in the Network of Cancer Genes

The section contains six components:
  1. Multi-Layer Relationship Diagram of Multi-Omics Drivers and Biological Functions — visualizes the relationships between multi-omics driver genes and their associated biological functions across molecular layers.
  2. Distribution of Multi-Omics Drivers Across Omics Layers — summarizes how driver genes are distributed across mutation, CNV, methylation, mRNA expression, and miRNA regulation layers.
  3. Cross-Tool Comparison of Multi-Omics Driver Detection — compares the distribution of driver gene identification across multiple integration tools and omics types.
  4. Machine Learning Result Table — summarizes significant prognostic signatures identified by LASSO, Random Forest, and I-Boost across omics data types and survival endpoints, with hazard ratios, confidence intervals, and patient group sizes.
  5. Signature Results — displays the prognostic signature, Kaplan–Meier survival plot, and predictive performance plot for a user-selected algorithm and survival endpoint combination.
  6. Multi-Omics Survival Gene Summary — provides an overview of survival-related genes identified across omics types, endpoints, and algorithms, including bar charts summarizing gene distributions and a detailed gene table.

1.10.2 Multi-Layer Relationship Diagram of Multi-Omics Drivers and Biological Functions

This section presents a diagram illustrating the hierarchical relationships from the selected cancer project → omics layers → multi-omics driver genes → Gene Ontology Biological Process (GO BP) terms. The diagram shows how driver evidence is distributed across omics types, genes, and their associated GO Biological Process terms. A summary table below the diagram provides the corresponding gene-, omics-, tool-, and GO-related information.

1.10.3 Distribution of Multi-Omics Drivers Across Omics Layers

Purpose

This section summarizes the distribution of tool-supported driver evidence across genes and omics types.

It consists of two complementary plots:
  1. Left Heatmap – Tool Support per Gene and Omics Layer
  2. Right Bar Chart – Top Genes by Multi-Omics Tool Support

Tool Support Across Omics Layers (Left Heatmap)

This heatmap displays genes on the y-axis and omics types on the x-axis. Each cell represents the number of tools supporting the corresponding gene within that omics type. Hovering over a cell displays the gene, omics type, and exact number of supporting tools. Darker cells indicate a larger number of supporting tools for that gene–omics combination.

Top Genes by Tool Evidence (Right Bar Chart)

This stacked bar chart displays prioritized genes according to accumulated tool evidence across omics types. The x-axis represents the number of supporting tool records, and the y-axis lists gene symbols. Colored segments indicate the contribution from each omics type, allowing users to compare how tool evidence is distributed across molecular layers for each gene.


1.10.4 Cross-Tool Comparison of Multi-Omics Driver Detection

Purpose

This section compares how multi-omics identification tools distribute their detected genes across omics types and summarizes the number of genes supported by different numbers of tools.

It contains:
  1. Left Heatmap – Tool vs. Omics Layer Coverage
  2. Right Bar Chart – Gene Counts by Tool Support Level

Proportion of Genes Identified by Each Tool (Left Heatmap)

This heatmap displays multi-omics identification tools on the y-axis and omics types on the x-axis. Each cell represents the proportion of genes identified by that tool that belong to the corresponding omics type. Proportions are calculated within each tool across all represented omics types. Hovering over a cell displays the exact proportion.

Gene Counts by Tool Support Level (Right Bar Chart)

This bar chart shows how many genes are supported by a given number of distinct multi-omics tools. The x-axis represents the number of genes, and the y-axis represents the number of supporting tools. Hovering over a bar displays the exact number of genes at each tool-support level.


1.10.5 Machine Learning Results

The machine learning result table summarizes significant prognostic signatures identified by machine learning algorithms for the selected cancer type. Each row represents a significant result for a specific survival endpoint and algorithm combination.

The table includes the following columns:
  • Endpoint: the survival endpoint evaluated, including overall survival (OS), disease-specific survival (DSS), disease-free interval (DFI), or progression-free interval (PFI).
  • Algorithm: the machine learning algorithm that identified the signature — LASSO, Random Forest, or I-Boost.
  • HR: the hazard ratio comparing survival outcomes between the high- and low-risk groups defined by the composite signature score. Values greater than 1 indicate higher risk in the high-risk group; values less than 1 indicate lower risk.
  • L95 / U95: the lower and upper bounds of the 95% confidence interval for the hazard ratio, reflecting the precision of the risk estimate.
  • Log-rank p-value: the p-value from the log-rank test evaluating whether the survival difference between the high- and low-risk groups is statistically significant.
  • Patients in high-risk: the number of patients assigned to the high-risk group based on the composite signature score.
  • Patients in low-risk: the number of patients assigned to the low-risk group based on the composite signature score.

Users can reorder the table by clicking on any column name. Selecting a row displays the corresponding Kaplan–Meier survival plot and predictive performance plot in the Signature Results panel below.

1.10.6 Signature Results

The Signature Results panel displays the prognostic signature identified by the selected machine learning algorithm and survival endpoint for the selected cancer type. Users can select a machine learning algorithm — LASSO, Random Forest, or I-Boost — and a survival endpoint from the left menu. The corresponding signature gene table, Kaplan–Meier survival plot, and predictive performance plot are displayed on the right. For detailed algorithm descriptions and reference links, please refer to FAQ4.

  1. Signature gene table
    The signature gene table lists all molecular features included in the selected prognostic signature. The table always includes the following columns:
    • Omic: the omics data type from which the feature is derived — RNA expression, mutation, CNV, or methylation. When the signature includes features from multiple omics layers, each feature's source is identified here.
    • Gene: the gene symbol of the molecular feature.
    Additional columns depend on the selected algorithm:
    • LASSO: the Coefficient column shows the weight assigned to each feature. A positive coefficient indicates that a higher feature value is associated with worse survival (higher risk), shown in red; a negative coefficient indicates that a higher feature value is associated with better survival (lower risk), shown in blue. The interpretation of feature value depends on the omics type — for example, expression level in RNA data, mutation presence versus wild-type in mutation data, copy number level in CNV data, or methylation level in methylation data.
    • Random Forest: the Depth column indicates how early a feature appears in the decision trees, with shallower depth reflecting stronger discriminative power. The Relative Frequency column reflects how consistently the feature is used as a splitting variable across all trees, expressed as a proportion. The relative frequency column is color-coded according to its value.
    • I-Boost: the Coefficient column is interpreted the same way as in LASSO — positive values in red indicate higher risk and negative values in blue indicate lower risk.
    Users can reorder the table by clicking on any column name.
  2. Kaplan–Meier survival plot
    The Kaplan–Meier plot displays survival differences between patient groups stratified by their composite signature score, computed from the combined weighted contributions of all features in the signature regardless of omics type. Patients are divided into two groups based on the median signature score of the cohort:
    • High: patients with a signature score above the median, indicating higher overall risk
    • Low: patients with a signature score below the median, indicating lower overall risk
    The survival statistics are shown above the plot. The x-axis represents survival time from initial cancer diagnosis and the y-axis represents survival probability. Users can hover over the curves to view detailed survival information, and curves can be shown or hidden by clicking the corresponding legend labels.
  3. Predictive performance plot

    For LASSO and Random Forest, ROC curves evaluate the predictive performance of the signature at different survival time points. The x-axis represents the false-positive rate and the y-axis represents the true-positive rate. Users can hover over the curves to view the false-positive rate, true-positive rate, and cutoff value at each point. ROC curves for different survival times can be shown or hidden by clicking the corresponding legend labels.

    For I-Boost, a cumulative hazard plot is displayed instead of ROC curves. This plot shows the cumulative hazard over time for the high- and low-risk groups. Higher cumulative hazard values indicate a greater accumulated risk of the survival event occurring up to that time point. The x-axis represents survival time from initial cancer diagnosis and the y-axis represents cumulative hazard. Users can hover over the curves to view detailed information, and curves can be shown or hidden by clicking the corresponding legend labels.

1.10.7 Multi-Omics Survival Gene Summary

The Multi-Omics Survival Gene Summary panel provides an overview of survival-related genes identified by LASSO, Random Forest, and I-Boost across omics data types and survival endpoints for the selected cancer type.


Bar charts
The three bar charts at the top summarize the distribution of significant survival-related genes from different perspectives. Hover over each bar to view the exact gene count.
(A) Significant genes by omics type: Shows the number of significant survival-related genes identified from each omics data type — RNA expression, mutation (MUT), CNV, and methylation (MET). This chart helps users assess which omics layer contributes the most survival-related features in the selected cancer type.
(B) Significant genes by survival endpoint: Shows the number of significant survival-related genes associated with each survival endpoint — OS, PFI, DSS, and DFI. This chart helps users compare the breadth of survival associations across endpoints.
(C) Significant genes by algorithm: Shows the number of significant survival-related genes identified by each machine learning algorithm — LASSO, Random Forest, and I-Boost. This chart helps users assess whether results are consistent across algorithms or driven predominantly by one method.

Survival gene table (D)
The table below the bar charts lists all survival-related genes identified across omics types, endpoints, and algorithms. Each row represents a unique omics-gene combination, so a gene identified across multiple omics types appears as separate rows. The table includes the following columns:
Molecular: the omics data type from which the feature is derived — RNA expression, mutation (MUT), CNV, or methylation (MET).
Gene: the gene symbol of the molecular feature.
Count: the total number of + marks in that row, reflecting how many endpoint-algorithm combinations identified the gene as survival-related. Higher counts indicate more consistent survival relevance across endpoints and algorithms.
Endpoint columns: each survival endpoint — OS, PFI, DSS, and DFI — is represented as a column group with three sub-columns corresponding to LASSO, Random Forest, and I-Boost. A + indicates that the gene was identified as survival-related by that algorithm under that endpoint. A blank cell indicates the gene was not identified under that combination.

Users can reorder the table by clicking on any column name. For algorithm descriptions and references, please refer to FAQ4.

1. Cancer Survival Relevance

The Survival Relevance tab is available for RNA, mutation, copy number variation (CNV), and methylation in the Cancer section. It evaluates the prognostic relevance of driver genes from the selected omics data type across multiple survival endpoints and analysis methods.

Overall summary

The bar charts and Venn diagrams summarize the number and overlap of survival-related driver genes from the selected omics data type across four survival endpoints and four survival analysis methods.

The four survival endpoints include overall survival (OS), progression-free interval (PFI), disease-specific survival (DSS), and disease-free interval (DFI). The four survival analysis methods include Cox univariate regression, Cox multivariate regression adjusted for clinical covariates, cure model analysis, and machine learning (ML). The ML analysis includes LASSO, Random Forest, and I-Boost; a gene is considered supported by ML if it is identified by at least one of these three algorithms. For more information about these methods, please refer to FAQ4.

  1. Number of Survival-Related Driver Genes by Analysis Method
    This bar chart shows the number of survival-related driver genes identified by each analysis method. The x-axis represents the analysis method, and the y-axis represents the number of genes. Hover over each bar to view the exact count.
  2. Number of Survival-Related Driver Genes by Survival Endpoint
    This bar chart shows the number of survival-related driver genes identified for each survival endpoint. The x-axis represents the survival endpoint, and the y-axis represents the number of genes. Hover over each bar to view the exact count.
  3. Overlap of Survival-Related Driver Genes among Analysis Methods
    This Venn diagram shows the overlap of survival-related driver genes identified across the four analysis methods. Genes appearing in overlapping regions are supported by multiple methods, indicating more consistent evidence of survival relevance. Hover over each region to view the number and percentage of genes in that subset.
  4. Overlap of Survival-Related Driver Genes among Survival Endpoints
    This Venn diagram shows the overlap of survival-related driver genes across the four survival endpoints. Genes appearing in overlapping regions are associated with multiple endpoints, indicating broader prognostic relevance across clinical outcomes. Hover over each region to view the number and percentage of genes in that subset.


Survival gene summary table

The Survival Gene Summary table lists survival-related driver genes in the selected cancer type for the selected omics data type. Results are organized into four tabs corresponding to overall survival (OS), progression-free interval (PFI), disease-specific survival (DSS), and disease-free interval (DFI).

Each endpoint-specific table includes the following columns: Gene Symbol, Cox Uni, Cox Multi (Clinical), Cure Model, Machine Learning, and Number of Algorithms.
Genes identified by at least one of the four analysis methods (Number of Algorithms ≥ 1) are included in the summary table.

For Cox Uni, Cox Multi (Clinical), and Cure Model, values represent log-transformed hazard ratios (log HRs), centered around zero to facilitate comparison of risk directions. Positive values are shown in red and indicate higher risk, whereas negative values are shown in blue and indicate lower risk. Greater color intensity represents a larger absolute log HR. Blank cells indicate that the gene did not meet the significance criterion or was not evaluated by that method.

For Machine Learning, genes identified by at least one of the three machine learning algorithms—LASSO, Random Forest, or I-Boost—are marked with “+”. A blank cell indicates that the gene was not identified by any of the three algorithms.

The Number of Algorithms column indicates how many of the four analysis methods — Cox Univariate, Cox Multivariate (Clinical), Cure Model, and Machine Learning—support the gene as survival-related, on a scale of 1 to 4. Higher values indicate more consistent evidence of survival relevance across methods. Users can reorder the table by clicking any column name. See FAQ4 for algorithm descriptions and reference links.


Trans-Omics Synergistic Effect

The trans-omics synergistic effect evaluates whether pairs of molecular features from different omics layers show a combined survival effect within the selected cancer type. The analysis considers cross-omics interactions among RNA expression, mutation, copy number variation (CNV), and methylation, and is currently available for overall survival (OS) only.

For each interaction, the combined survival effect of the paired features is compared with the survival effect of the individual features. The resulting synergistic interactions are summarized in a table, and selecting an interaction displays the corresponding Kaplan–Meier survival plots.

Within each omics tab, results are restricted to interactions in which the gene associated with the current omics layer is identified as a driver. Users can further filter the results by gene set (All, CGC, or NCG 6.0) and hazard ratio direction (All, HR > 1, or HR < 1).

Result Table

The result table summarizes each synergistic survival interaction using the following columns:

  • cancer: Cancer type in which the interaction was identified.
  • trans_omics: Pair of omics layers involved in the interaction.
  • gene1 / gene2: Gene symbols of the two molecular features forming the interaction.
  • omic1 / omic2: Omics types corresponding to gene1 and gene2, respectively.
  • HR.FC: Synergy fold-change score comparing the combined effect with the stronger single-feature effect.
  • Hazard_ratio: Hazard ratio representing the combined effect of the paired features.
  • log2HR.dir: Direction of the survival effect associated with the first omics feature.
  • adjust_pval: Benjamini–Hochberg adjusted log-rank p-value for the combined interaction.

The table can be sorted by clicking any column header.

Kaplan–Meier Survival Plots

Selecting a row in the result table displays the corresponding Kaplan–Meier survival plots below.

The left plot shows the unadjusted Kaplan–Meier survival curves, whereas the right plot shows covariate-adjusted survival curves when sufficient clinical covariate data are available. Survival curves are displayed for the first 5 years of follow-up.

The x-axis represents survival time from the initial cancer diagnosis, and the y-axis represents survival probability. Users can hover over the curves to inspect survival information and click the legend to show or hide individual groups.

Synergy Score

The synergistic effect is quantified using HR.FC, which compares the hazard ratio of the combined feature pair with the stronger hazard ratio of the individual features.

The combined effect is represented by HR.pair, while the individual feature effects are represented by HR.single1 and HR.single2. The value HR.single corresponds to the stronger single-feature effect after accounting for the direction of the survival effect.

An HR.FC > 1 indicates that the combined feature pair shows a stronger survival effect than the stronger individual feature alone, after accounting for the direction of the survival effect.

Patient Group Stratification

Patient groups are defined according to the combined molecular states of the two paired features.

  • RNA expression: Patients are divided into high- and low-expression groups using the median expression cutoff.
  • Mutation: Patients are classified as mutated or wild-type.
  • CNV: Patients are classified as gain, loss, or neutral according to iGC.
  • Methylation: Patients are divided into high- and low-methylation groups using the median beta value.

Group Label Definitions

The labels shown in the Kaplan–Meier plots describe the combined molecular states of the two features. The order of gene1 and gene2 follows the interaction type shown in the result table.

RNA–Mutation
  • high_mut: High RNA expression in gene1 and mutation in gene2.
  • high_wt: High RNA expression in gene1 and wild-type gene2.
  • low_mut: Low RNA expression in gene1 and mutation in gene2.
  • low_wt: Low RNA expression in gene1 and wild-type gene2.
RNA–CNV
  • high_gain: High RNA expression in gene1 and CNV gain in gene2.
  • high_loss: High RNA expression in gene1 and CNV loss in gene2.
  • high_none: High RNA expression in gene1 and neutral CNV in gene2.
  • low_gain: Low RNA expression in gene1 and CNV gain in gene2.
  • low_loss: Low RNA expression in gene1 and CNV loss in gene2.
  • low_none: Low RNA expression in gene1 and neutral CNV in gene2.
RNA–Methylation
  • high_meth: High RNA expression in gene1 and high methylation in gene2.
  • high_unmeth: High RNA expression in gene1 and low methylation in gene2.
  • low_meth: Low RNA expression in gene1 and high methylation in gene2.
  • low_unmeth: Low RNA expression in gene1 and low methylation in gene2.
Mutation–CNV
  • mut_gain: Mutated gene1 and CNV gain in gene2.
  • mut_loss: Mutated gene1 and CNV loss in gene2.
  • mut_none: Mutated gene1 and neutral CNV in gene2.
  • wt_gain: Wild-type gene1 and CNV gain in gene2.
  • wt_loss: Wild-type gene1 and CNV loss in gene2.
  • wt_none: Wild-type gene1 and neutral CNV in gene2.
Mutation–Methylation
  • mut_meth: Mutated gene1 and high methylation in gene2.
  • mut_unmeth: Mutated gene1 and low methylation in gene2.
  • wt_meth: Wild-type gene1 and high methylation in gene2.
  • wt_unmeth: Wild-type gene1 and low methylation in gene2.
CNV–Methylation
  • gain_meth: CNV gain in gene1 and high methylation in gene2.
  • gain_unmeth: CNV gain in gene1 and low methylation in gene2.
  • none_meth: Neutral CNV in gene1 and high methylation in gene2.
  • none_unmeth: Neutral CNV in gene1 and low methylation in gene2.
  • loss_meth: CNV loss in gene1 and high methylation in gene2.
  • loss_unmeth: CNV loss in gene1 and low methylation in gene2.

Interpretation

A synergistic interaction indicates that the combined molecular states of two cross-omics features are associated with a stronger survival effect than the stronger individual feature alone. Such interactions may highlight complementary molecular alterations that are jointly associated with patient prognosis within the selected cancer type.

2. Gene

2.1 Gene Module Overview

The Gene module provides a comprehensive, multi-omics overview of a user-selected gene across multiple cancer types. By integrating expression, mutation, CNV, methylation, miRNA regulation, protein expression, multi-omics driver evidence, and survival analyses, this module helps users understand how a gene behaves across the cancer landscape and how its molecular alterations may relate to patient outcomes.


2.2 Input Selection

Search Mode

To begin, choose how you want to search for the gene:
  • Gene Name – Enter a gene symbol or alias (e.g., EGFR or ERK).
  • Ensembl ID – Enter the Ensembl gene identifier (e.g., ENSG00000141510)

After entering the query, click Submit. Matching genes with significant results available for downstream analyses will be displayed in the result table below.

Search Result Selection

The result table provides the corresponding Ensembl ID, gene symbol, official symbol, and aliases for each matched record. If multiple records are returned, use this information to identify the desired gene.

Select the radio button next to the desired gene to open its corresponding Gene page and view the available downstream analyses.



2.3 Overview of Result Tabs

The Gene module contains several result tabs, each summarizing multi-omics evidence and functional insights for a selected gene across different cancer types:
  • Summary – Provides an overview of multi-omics evidence for the selected gene across projects, cohorts, and tissues. Bar plots and boxplots summarize global cross-cohort results, including RNA, CNV, methylation, mutation, and miRNA findings. The heatmap displays tissue-specific project-level results based on the tissue or organ selected from the body diagram.
  • RNA – Displays gene expression patterns across cancer projects and cohorts, with results organized by sample type or tumor stage. Users can explore organ- or tissue-specific expression patterns across projects as well as detailed cancer-specific results. A separate survival analysis evaluates the association between expression of the selected gene and patient survival.
  • Mutation – Provides mutation-focused results for the selected gene through three tabs: mutation rate, mutation percentage, and exon distribution. The mutation rate and mutation percentage tabs include heatmaps showing hotspot mutation regions (HMRs) across multiple cancer types, where mutation frequency is calculated based on mutation count relative to sample count. Cancer project-specific mutation rate bar charts and survival analysis results are also provided.
  • CNV – Visualizes copy-number alterations affecting the gene, including amplification and deletion frequencies across cancer types, CNV–expression correlations, and survival analysis based on copy-number status for the user-selected gene.
  • Methylation – Highlights methylation status at the gene locus, including hyper- and hypomethylation patterns, methylation–expression relationships, and survival analysis based on methylation status for the user-selected gene.
  • miRNA – Shows regulatory interactions between the gene and miRNAs, integrating predicted interactions, experimentally validated evidence, and cancer-specific expression correlations.
  • Protein – Displays protein-level expression data, post-translational modifications, and protein–protein interactions relevant to the gene product.
  • Multi-Omics – Integrates multi-omics evidence for the selected gene across RNA, mutation, CNV, and methylation analyses. Visualizations summarize the distribution of supporting tool evidence across cancer projects and omics types, while a Sankey diagram illustrates the relationships between the selected gene, supported omics types, and cancer projects.

2.4 Gene Summary

Gene Overview

The Gene Overview page provides a high-level overview of multi-omics evidence for the selected gene across projects, cohorts, and tissues. These visualizations help users quickly identify where the gene shows qualifying molecular alterations, how consistently these patterns appear across cohorts, and whether tissue-specific patterns are present.

The summary bar plots and boxplot provide global cross-cohort summaries for the selected gene, while the heatmap shows project-level results for the selected tissue or organ. The bar plots report the proportion of cohorts showing each class of alteration in RNA expression, somatic mutation, copy-number variation (CNV), and DNA methylation, while the boxplot summarizes the number of associated miRNAs across projects.
Detailed information on the computational algorithms and tools used in DriverDBv5 can be found in FAQ4.

Color definitions, statistical cutoffs, and asterisk criteria are described in the next Visualization Color & Asterisk Reference section.

Select a tissue or organ from the body diagram to update the heatmap and view tissue-specific results across RNA, mutation, CNV, methylation, and miRNA.


Visualization Color & Asterisk Reference

  • RNA expression: Red indicates upregulation, blue indicates downregulation, white indicates no significant alteration, and gray indicates no available data.
  • Somatic mutation: Orange indicates a mutation driver supported by at least three distinct mutation detection tools, white indicates support by fewer than three tools, and gray indicates no available data.
  • CNV: Red indicates copy-number gain, blue indicates copy-number loss, the combined color indicates both gain and loss, white indicates no significant alteration, and gray indicates no available data.
  • DNA methylation: Red indicates hypermethylation, blue indicates hypomethylation, the combined color indicates both hyper- and hypomethylation, white indicates no significant alteration, and gray indicates no available data.
  • miRNA: The orange color gradient represents the number of associated miRNAs, with darker colors indicating larger numbers.

Asterisks (*) indicate significant survival associations where applicable. Detailed criteria for each category are provided below.


2.5 Gene RNA

2.5.1 Overview

The RNA panel visualizes expression patterns and survival associations for the selected gene across multiple cancer types.

Users can explore RNA expression results using two grouping options: sample type and tumor stage. Each grouping option includes an Organ-specific Project View and a Cancer-Specific View, allowing users to compare expression patterns across projects associated with a selected tissue or organ or examine expression details within a specific cancer project.

In the Organ-specific Project View, users can select or click a tissue or organ on the body diagram to display relevant cancer projects. RNA expression is shown as z-scores, enabling comparison of expression distributions across projects with different expression metrics. Results are grouped by either sample type or tumor stage according to the selected view.

In the Cancer-Specific View, expression distributions are shown as violin plots with embedded boxplots. The expression metric is defined by the selected project and may therefore vary across datasets. Pairwise comparisons between available groups are provided in a summary table below the plot.

A separate Survival Map and Survival Analysis section evaluates whether RNA expression of the selected gene is associated with patient survival. Patients are grouped into high- and low-expression groups, and the survival results reflect the overall expression level of the gene rather than the sample type or tumor stage groupings shown in the expression panels. Survival results are available for TCGA cohorts only.

The RNA section is organized into the following components:
  1. Expression by Sample Type
  2. Expression by Tumor Stage
  3. Survival Map
  4. Survival Analysis

2.5.2 Expression by Sample Type

This tab displays expression across all available sample types (e.g., NT, TP, TM, TRBM, TBM).

Organ-specific Project Expression by Sample Type

This panel displays z-score expression distributions of the selected gene across all projects associated with the selected tissue or organ, grouped by sample type.
Users can:
  • Select or click a tissue or organ to view relevant cancer projects for the selected gene
  • Toggle specific sample-type groups using the legend.
  • Hover over individual dots to view sample-level details.
  • Hover near box areas to view distribution summary statistics.
Sample Type Abbreviations:
  • NT — Solid Tissue Normal
  • NB - Blood Derived Normal
  • TAP — Additional New Primary
  • TB — Primary Blood-Derived Cancer
  • TBM — Metastatic Blood-Derived Cancer
  • TM — Metastatic
  • TP — Primary Solid Tumor
  • TR — Recurrent Solid Tumor
  • TRBM — Recurrent Blood-Derived Metastatic

Cancer-Specific View: Expression by Sample Type

Shows gene expression distributions of the selected gene within a specific cancer project, grouped by sample type.

Users first select an organ or tissue from the left panel (A), which filters the available cancer projects to those associated with the selected organ or tissue. Users then select a specific cancer project from the filtered list (B). Once a project is selected, the corresponding project description (C) is displayed at the top, followed by violin plots with embedded boxplots (D) showing the expression distribution of the selected gene across available sample types within that project. A summary table (E) is displayed below the plots, reporting pairwise comparisons between sample types, including p-values that indicate whether gene expression differs significantly between the compared groups.

Expression values are shown using the expression metric defined for the selected project, which may vary across datasets.

Use the sample type controls to show or hide specific groups. Hover over individual dots to view sample-level details, or hover over the violin or boxplot areas to view distribution statistics, including the maximum, upper fence, Q3, median, Q1, lower fence, and minimum.


2.5.3 Expression by Tumor Stage

This tab examines expression variation across tumor stages (Stage I–IV).

Organ-specific Project Expression by Tumor Stage

This panel displays z-score expression distributions of the selected gene across all projects associated with the selected tissue or organ, grouped by tumor stage (Stage I–IV).
Users can:
  • Select or click a tissue or organ to view relevant cancer projects for the selected gene.
  • Toggle specific tumor stages using the legend.
  • Hover over individual dots to view sample-level details, including cancer project, tumor stage, and z-score.
  • Hover near box areas to view distribution summary statistics.

Cancer-Specific View: Expression by Tumor Stage

Shows gene expression distributions of the selected gene within a specific cancer project, grouped by tumor stage.

Users first select an organ or tissue from the left panel (A), which filters the available cancer projects to those associated with the selected organ or tissue. Users then select a specific cancer project from the filtered list (B). Once a project is selected, the corresponding project description (C) is displayed at the top, followed by violin plots with embedded boxplots (D) showing the expression distribution of the selected gene across available tumor stages within that project. A summary table (E) is displayed below the plots, reporting pairwise comparisons between tumor stages, including p-values that indicate whether gene expression differs significantly between the compared groups.

Expression values are shown using the expression metric defined for the selected project, which may vary across datasets.

Use the stage controls to show or hide specific tumor stages. Hover over individual dots to view sample-level details, or hover over the violin or boxplot areas to view distribution statistics, including the maximum, upper fence, Q3, median, Q1, lower fence, and minimum.


2.5.4 Survival Map & Survival Analysis

The Survival Map and Survival Analysis sections evaluate the prognostic relevance of RNA expression of the selected gene across cancer cohorts and survival endpoints using multiple analysis methods. Survival results are available for TCGA cohorts only.

Detailed information on interpreting the Survival Map and Survival Analysis panels is provided in Gene-Survival.


2.6 Gene Mutation

2.6.1 Overview

The Mutation interface visualizes mutation patterns and statistics of the selected gene across multiple cancer types, with mutations mapped along the protein sequence and aligned with functional protein domains.

Users can explore three mutation-level summaries: Mutation Rate, Mutation Percent, and Exon Distribution. Mutation Rate reflects the frequency of mutations per sample, while Mutation Percent reflects the proportion of samples carrying at least one mutation in the selected gene. Exon Distribution summarizes how mutations are distributed across the exonic regions of the gene. Each summary includes both a Pan-Cancer View and a Cancer-Specific View, allowing users to compare mutation patterns across cancer types or examine mutation details within a selected cancer type. Mutation hotspots — positions where mutations cluster more frequently than expected — can be identified by examining the distribution of mutations along the protein coordinates.

Each of the three mutation summaries includes a dedicated Survival Map and Survival Analysis section. These evaluate whether mutation status of the selected gene is associated with patient survival using the same patient grouping — mutated versus wild-type — regardless of which mutation summary is selected. As a result, the survival results are consistent across the three summaries and reflect the overall mutation status of the gene rather than the specific mutation-level metric displayed above. Survival results are available for TCGA cohorts only.

Each mutation summary section is organized as follows:
  1. Pan-Cancer View — displays the selected mutation metric across all available cancer types, mapped along protein coordinates and aligned with functional protein domains.
  2. Cancer-Specific View — displays the selected mutation metric within a selected cancer type, allowing detailed examination of mutation patterns and hotspots.
  3. Survival Map — summarizes survival associations of the selected gene's mutation status across cancer types and survival endpoints.
  4. Survival Analysis — provides detailed survival analyses evaluating the association between mutation status and patient prognosis using multiple analysis methods.

This structure is consistent across all three mutation summaries: Mutation Rate, Mutation Percent, and Exon Distribution.

2.6.2 Mutation Rate

Pan-Cancer View: Mutation Rate Heatmap

This view integrates multiple coordinated panels to show where mutations occur along the protein and how frequently they appear across cancer types, using mutation rate as the metric.

Components
A. Pan-Cancer Mutation Hotspot Heatmap

This heatmap displays projects as rows and protein positions as columns, with cell color indicating the mutation rate at each specific position. Users can hover over cells to view the cancer type, protein position, and mutation rate, allowing them to identify protein regions with recurrent mutation hotspots across multiple cancer types.

B. Protein Region Impact Bar Plot

This plot aggregates mutation rates per protein region, with bars stacked by impact level (High, Moderate, or Low) to show the relative contribution of different mutation severities. Hovering over bar segments reveals the region name, impact category, and mutation rate, demonstrating which functional regions of the protein accumulate the highest mutation load.

C. Dataset-Level Mutation Burden Bar Plot

This bar chart displays the overall mutation rate per dataset or cancer type, with bars stacked by mutation impact level to show the distribution of mutation severities. Users can hover to view the dataset, tissue type, impact level, and mutation rate, providing a quick comparison of which cancers have the heaviest mutation burden for the selected gene.

D. Dataset & Tissue Legend

This companion panel lists the tissue type, project ID, and cancer type for each dataset included in the analysis, with each color corresponding to a specific tissue type to help users interpret the color-coding used throughout the visualization.

E. Protein Domain Annotation Track

This track shows annotated protein domains from Pfam or InterPro databases, displaying the domain name, protein coordinate range, and functional description (accessible via hover). This annotation aligns functional domains with mutation hotspots, helping users understand whether mutations cluster in functionally important regions of the protein.

F. Exon Annotation Track

This track displays exon boundaries aligned to protein coordinates, with each exon shown as a distinct colored block to illustrate the genomic structure underlying the protein sequence and how mutations map to specific exons.

G&H. Legends

Two legends accompany the visualization: a mutation rate legend providing a continuous color scale for heatmap intensity, and an impact legend showing categorical colors for High, Moderate, and Low mutation impacts to help users interpret the color-coding throughout all components.



Cancer-Specific View: Mutation Rate Bar Chart

Displays the mutation rate of the selected gene across protein positions within a selected cancer project, with mutations categorized by predicted impact level.

Users first select an organ or tissue from the left panel (A), which filters the available cancer projects to those associated with the selected organ or tissue. Users then select a specific cancer project from the filtered list (B). Once a project is selected, the corresponding project description (C) is displayed at the top, followed by the bar chart (D), where each bar represents a protein position and the height reflects the proportion of samples carrying a mutation at that position, expressed as a rate.

Mutation impacts are stacked within each bar into three categories — High, Moderate, and Low — to illustrate the impact composition at each protein position. This allows users to identify not only mutation-enriched regions along the protein sequence but also whether mutations at those positions are predominantly high-impact or low-impact.

Use the legend (E) to show or hide specific impact categories for focused comparisons. Hover over individual bars to view the protein position, mutation rate, impact category, and cancer project.



2.6.3 Mutation Percent

Pan-Cancer View: Mutation Percent Heatmap

This visualization is structurally identical to the Mutation Rate view but uses mutation percentage—the proportion of mutated samples in each dataset—rather than mutation rate.

Components

A. Pan-Cancer Mutation Hotspot Heatmap

This heatmap displays projects as rows and protein positions as columns, with cell color indicating the mutation percentage at each specific position. Users can hover over cells to view the cancer type, protein position, and mutation percentage, revealing which protein positions are most frequently mutated across patient samples in different cancer types.

B. Protein Region Impact Bar Plot

This plot shows mutation percentage per protein region, with bars stacked by mutation impact level to display the relative contribution of High, Moderate, and Low impact mutations. Hovering over bar segments reveals the region name, impact level, and mutation percentage, highlighting protein regions with high prevalence of mutations among patients and indicating which functional domains are most commonly affected.

C. Dataset-Level Mutation Burden Bar Plot

This bar chart displays mutation percentage per dataset or cancer type, with bars stacked by impact level to show the distribution of mutation severities. Users can hover to view the dataset, tissue type, impact level, and mutation percentage, identifying cancer types where mutations in the gene are widespread across patient populations.

D. Dataset & Tissue Legend

Color-coded tissue and dataset identifiers for interpreting the heatmap rows.

E. Protein Domain Annotation Track

This track displays Pfam and InterPro protein domains aligned to protein coordinates, showing the domain name, coordinate range, and functional details accessible through hovering, allowing users to determine whether mutations cluster within functionally important protein domains.

F. Exon Annotation Track

This track displays exon boundaries aligned to protein structure, with each exon shown as a distinct block to illustrate how the genomic organization corresponds to the protein sequence and mutation positions.

G&H. Legends

Two legends accompany the visualization: a mutation percent legend providing a color scale for mutation proportions, and an impact legend showing colors for High, Moderate, and Low mutation impact categories to help users interpret the color-coding throughout all components.


Cancer-Specific View: Mutation Percent Bar Chart

Displays the mutation percentage of the selected gene across protein positions within a selected cancer project, with mutations categorized by predicted impact level.

Users first select an organ or tissue from the left panel (A), which filters the available cancer projects to those associated with the selected organ or tissue. Users then select a specific cancer project from the filtered list (B). Once a project is selected, the corresponding project description (C) is displayed at the top, followed by the bar chart (D), where each bar represents a protein position and the height reflects the proportion of samples carrying a mutation at that position, expressed as a percentage.

Mutation impacts are stacked within each bar into three categories — High, Moderate, and Low — to illustrate the impact composition at each protein position. This allows users to identify not only mutation-enriched regions along the protein sequence but also whether mutations at those positions are predominantly high-impact or low-impact.

Use the legend (E) to show or hide specific impact categories for focused comparisons. Hover over individual bars to view the protein position, mutation percentage, impact category, and cancer project.

2.6.4 Exon Distribution

Pan-Cancer View: Exon Mutation Distribution

Components

A. Mutation Count by Exon

Shows the number of mutations per exon across all cancer types. X-axis = exon number; Y-axis = mutation count. Bars are stacked by mutation impact (High, Moderate, Low, Modifier). Hover to view exon number, impact, and mutation count.

B. Mutation Percentage by Exon

Displays the proportion of mutated samples per exon. X-axis = exon number; Y-axis = mutation percentage. Hover to view exon number, impact, and mutation percentage.

C. Protein Domain Panel

Annotated functional domains with Pfam ID, InterPro ID, position, and description. Hover for details on each domain.

D. Exon Annotation Track

Each colored block represents an exon aligned to the protein coordinate axis.

E. Impact Legend


Cancer-Specific View: Exon Mutation Distribution

Displays the exon-level mutation distribution of the selected gene within a selected cancer project, with mutations categorized by predicted impact level.

Users first select the visualization metric — mutation count or mutation percentage — from panel (A) to determine whether the bar chart displays the absolute number of mutations or the proportion of samples carrying a mutation per exon. Users then select an organ or tissue from panel (B), which filters the available cancer projects to those associated with the selected organ or tissue, and select a specific cancer project from the filtered list (C).

Once selections are made, the corresponding project description (D) is displayed at the top, followed by the bar chart (E) showing the exon-level mutation distribution for the selected cancer project and metric. Each bar represents an exon, and mutations are stacked within each bar into three impact categories — High, Moderate, and Low — to illustrate the impact composition at each exon. This allows users to identify which exons harbor the highest mutation load or frequency and whether mutations within those exons are predominantly high-impact or low-impact.

Use the legend (F) to show or hide specific impact categories for focused comparisons. Hover over individual bars to view the exon number, impact category, and mutation count or percentage.


2.6.5 Survival Map & Survival Analysis

The Survival Map and Survival Analysis sections evaluate the prognostic relevance of mutations in the selected gene across cancer cohorts and survival endpoints using multiple analysis methods. Survival results are available for TCGA cohorts only.

Detailed information on interpreting the Survival Map and Survival Analysis panels is provided in Gene-Survival.


2.7 Gene CNV

2.7.1 Overview

The Copy Number Variation interface visualizes CNV patterns of the selected gene across multiple cancer types and explores how copy number changes relate to gene expression and patient survival.

This interface integrates results from two complementary CNV analysis tools that operate at different levels of analysis. iGC identifies significant copy number gains and losses for the selected gene across cancer types, providing a gene-level summary of CNV status. DIGGIT identifies genes whose copy number alterations are significantly correlated with downstream gene expression changes, inferring potential CNV driver genes — that is, genes whose copy number changes may confer a functional advantage by altering the expression of downstream targets. Together, these tools allow users to explore both the CNV status of the selected gene and its potential functional consequences.

Users can explore CNV gain or loss significance across cancer types, examine how copy number changes correlate with gene expression levels, and identify cancers where the selected gene may act as a CNV driver.

The Survival section evaluates whether copy number variation of the selected gene is associated with patient survival. Survival results include a Survival Map summarizing associations across cancer types and endpoints, and detailed survival analyses based on multiple analysis methods. Survival results are available for TCGA cohorts only.

This panel includes five sections:
  1. Pan-Cancer View: Copy Number Variation Overview — summarizes CNV gain and loss status of the selected gene across all available cancer types.
  2. Cancer-Specific View: CNV Distribution and Correlation — displays CNV distributions and the relationship between copy number status and gene expression for a selected cancer type.
  3. CNV Summary Table — provides a structured summary of CNV details across cancer types, including iGC and DIGGIT results.
  4. Survival Map — summarizes survival associations of the selected gene's CNV status across cancer types and survival endpoints.
  5. Survival Analysis — provides detailed survival analyses evaluating the association between CNV status and patient prognosis using multiple analysis methods.

2.7.2 Pan-Cancer View: Copy Number Variation Overview

This visualization summarizes CNV gain, loss, and no-change states for the selected gene across available cancer projects. Only projects in which the selected gene meets the iGC CNV driver criteria are included. The display format depends on the number of available projects: when more than five projects are available, results are shown as a bar chart; when five or fewer projects are available, results are shown as pie charts.

For results displayed as a bar chart, the CNV driver panel at the top indicates the level of CNV driver support for the selected gene in each cancer project:
  • 1 tool (light grey): the gene meets the iGC CNV driver criteria.
  • 2 tools (dark grey): the gene meets the iGC CNV driver criteria and is additionally supported by DIGGIT-related evidence together with a significant positive CNV–expression correlation.
CNV Driver Criteria:
  • iGC: CNV gain or loss with FDR < 0.001, sample proportion > 0.15, and absolute log2 fold change > 1.
  • Additional DIGGIT-related support: DIGGIT Spearman p-value < 0.001, CNV–expression Spearman correlation > 0.3, and correlation p-value < 0.05.

The main panel displays the proportions of samples classified by iGC as CNV gain, CNV loss, or no CNV change. Each bar represents one cancer project, and the height of each segment represents the proportion of samples in the corresponding CNV state. Hover over the chart to view CNV status and sample-proportion information.

When five or fewer projects are available, the same sample-proportion information is displayed as pie charts. Each pie chart represents one cancer project, and each segment represents the proportion of samples classified as CNV gain, CNV loss, or no CNV change.

Together, these views allow users to compare CNV alteration patterns across cancer projects and determine whether the selected gene is supported as a CNV driver by iGC alone or by iGC together with additional DIGGIT-related evidence.


2.7.3 Cancer-Specific View: CNV Distribution and Correlation

This visualization shows the relationship between copy number variation and gene expression within a selected cancer project, combining sample-level CNV segment mean and expression values with comparisons across CNV status groups.

Users first select an organ or tissue from panel (A), which filters the available cancer projects to those associated with the selected organ or tissue. Users then select a specific cancer project from the filtered list (B). Once a project is selected, the corresponding project description (C) is displayed at the top, followed by a 2 × 2 grid of visualization panels.

D. CNV–Expression Correlation Scatter Plot
Displays the relationship between CNV segment mean (x-axis) and gene expression (y-axis) for individual samples. Each point represents one sample and is colored according to CNV status. Expression values are shown using the expression metric defined for the selected project, which may vary across datasets. Hover over a point to view its segment mean and expression value.

E. Expression by CNV Status Boxplot
Summarizes gene expression across gain, loss, no-change, and normal groups, allowing comparison of expression distributions among CNV states. Hover over the boxplot areas to view summary statistics, including the maximum, upper fence, Q3, median, Q1, lower fence, and minimum.

F. Segment Mean by CNV Status Boxplot
Summarizes the distribution of CNV segment mean values across gain, loss, no-change, and normal groups. Hover over the boxplot areas to view the corresponding distribution statistics.

G. Correlation Summary Panel
Displays the Spearman correlation coefficient and corresponding p-value for the association between CNV segment mean and gene expression. A positive coefficient indicates that higher CNV segment mean values tend to be associated with higher gene expression, whereas a negative coefficient indicates an inverse relationship.

H. Legend
The legend displays the CNV status color coding: red for gain, blue for loss, light grey for no change, and dark grey for normal samples. Click legend entries to show or hide specific groups across the visualization.


2.7.4 CNV Summary Table

The CNV Summary Table provides detailed CNV statistics for the selected gene across available cancer projects, including results from iGC, DIGGIT-related evidence, and CNV–expression correlation analysis.

Column Descriptions:
  • Cancer type abbreviation: Abbreviated cancer type name (e.g., LUAD, BRCA)
  • Gene symbol: The official gene symbol of the user-selected gene (e.g., EGFR).
  • ENSG: The Ensembl gene identifier corresponding to the selected gene.
  • iGC Gain p-value: p-value from the iGC tool assessing the statistical significance of copy-number gain in the selected gene.
  • iGC Gain FDR: False Discovery Rate–adjusted p-value for CNV gain from iGC, used to correct for multiple testing.
  • iGC Loss p-value: p-value from iGC evaluating the significance of copy-number loss in the selected gene.
  • iGC Loss FDR: FDR-adjusted p-value for CNV loss from iGC.
  • iGC Gain sample proportion: The proportion of samples classified as copy-number gain by iGC within the cancer dataset.
  • iGC Normal sample proportion: The proportion of samples classified as normal (diploid) by iGC.
  • iGC Loss sample proportion: The proportion of samples classified as copy-number loss by iGC.
  • iGC Gain log₂FC: Log₂ fold change in expression between gain and normal CNV groups, as estimated by iGC.
  • iGC Loss log₂FC: Log₂ fold change in expression between loss and normal CNV groups, as estimated by iGC.
  • DIGGIT Spearman p-value: Spearman-based p-value from DIGGIT used as additional evidence for CNV driver support.
  • Segment mean vs. expression Spearman correlation coefficient: Spearman correlation coefficient (ρ) measuring the association between segment mean values and gene expression (TPM).
  • Segment mean vs. expression Spearman p-value: p-value testing the statistical significance of the Spearman correlation between CNV segment mean and gene expression.

Together, these results allow users to compare CNV alteration patterns and their associations with gene expression across cancer projects.

2.7.5 Survival Map & Survival Analysis

The Survival Map and Survival Analysis sections evaluate the prognostic relevance of copy number variation of the selected gene across cancer cohorts and survival endpoints using multiple analysis methods. Survival results are available for TCGA cohorts only.

Detailed information on interpreting the Survival Map and Survival Analysis panels is provided in Gene-Survival.


2.8 Gene Methylation

2.8.1 Overview

The Methylation interface visualizes DNA methylation patterns of the selected gene across multiple cancer projects and explores the relationship between methylation, gene expression, and patient survival.

Methylation driver evidence is identified primarily using MethylMix, which identifies aberrant methylation patterns of the selected gene. ELMER provides additional probe-level evidence for MethylMix-identified methylation drivers. Together, these results allow users to explore the methylation status of the selected gene and its relationship with gene expression.

The Survival section evaluates whether methylation status of the selected gene is associated with patient survival. Survival results include a Survival Map summarizing associations across cancer projects and survival endpoints, and detailed survival analyses based on multiple analysis methods. Survival results are available for TCGA cohorts only.

This panel includes five sections:
  1. Pan-Cancer View: Methylation Status Overview — summarizes methylation status and methylation driver support for the selected gene across available cancer projects.
  2. Cancer-Specific View: Methylation Distribution and Correlation — displays methylation distributions and the relationship between beta value and gene expression within a selected cancer project.
  3. Methylation Summary Table — provides detailed methylation statistics across cancer projects, including MethylMix results, ELMER evidence, and methylation–expression correlation results.
  4. Survival Map — summarizes survival associations of the selected gene's methylation status across cancer projects and survival endpoints.
  5. Survival Analysis — provides detailed survival analyses evaluating the association between methylation status and patient prognosis using multiple analysis methods.

2.8.2 Pan-Cancer View: Methylation Status Overview

This visualization summarizes DNA methylation results for the selected gene across available cancer projects. Only projects in which the selected gene is identified as a methylation driver by MethylMix are included.

The methylation driver panel at the top summarizes methylation driver support for the selected gene across cancer projects. Light grey indicates that the gene is identified as a methylation driver by MethylMix, whereas dark grey indicates that the MethylMix prediction is additionally supported by ELMER.

Below the driver panel, sample proportions classified by MethylMix as hypermethylated, hypomethylated, or showing no methylation change are displayed for each cancer project. The display format depends on the number of available projects: when more than five projects are available, results are shown as a bar chart; when five or fewer projects are available, results are shown as pie charts.

Users can hover over the visualization to view methylation status and sample-proportion information. Together, these views allow users to compare methylation patterns across cancer projects and determine whether the selected gene is supported as a methylation driver by MethylMix alone or by MethylMix together with additional ELMER evidence.

2.8.3 Cancer-Specific View: Methylation Distribution and Correlation

This visualization helps users assess the relationship between DNA methylation and gene expression within a selected cancer project, displaying the correlation between beta value and gene expression alongside group-level comparisons across methylation status categories.

Users first select an organ or tissue from panel (A), which filters the available cancer projects to those associated with the selected organ or tissue. Users then select a specific cancer project from the filtered list (B). Once a project is selected, the corresponding project description (C) is displayed at the top, followed by a 2 × 2 grid of visualization panels.

D. Methylation–Expression Correlation Scatter Plot (upper right)
Displays the relationship between beta value (x-axis) and gene expression (y-axis) for individual samples, where each point represents one sample. Expression values are shown using the expression metric defined for the selected project, which may vary across datasets. Points are colored according to methylation status: red for hypermethylated, blue for hypomethylated, light grey for no methylation change, and dark grey for normal samples. Hover over a point to view the corresponding beta value and expression value.

E. Expression by Methylation Status Boxplot (upper left)
Summarizes gene expression across hypermethylated, hypomethylated, no methylation change, and normal groups, allowing users to compare expression distributions among methylation states. Expression values are shown using the expression metric defined for the selected project. Hover over the boxplot areas to view summary statistics, including the maximum, upper fence, Q3, median, Q1, lower fence, and minimum.

F. Beta Value by Methylation Status Boxplot (bottom right)
Summarizes the distribution of beta values across hypermethylated, hypomethylated, no methylation change, and normal groups, allowing users to compare methylation levels among groups. Hover over the boxplot areas to view the corresponding distribution statistics.

G. Correlation Summary Panel (bottom left)
Displays the Spearman correlation coefficient and corresponding p-value for the association between beta value and gene expression. A negative coefficient indicates that higher beta values tend to be associated with lower gene expression, whereas a positive coefficient indicates that higher beta values tend to be associated with higher gene expression.

H. Legend
The legend displays the methylation status color coding: red for hypermethylated, blue for hypomethylated, light grey for no methylation change, and dark grey for normal samples. Normal represents normal-tissue (NT) samples, whereas no methylation change represents tumor samples without a methylation alteration identified by MethylMix. Click the legend entries to show or hide specific methylation status groups for focused comparison.



2.8.4 Methylation Summary Table

The Methylation Summary Table provides detailed methylation and expression-correlation statistics for the selected gene across available cancer projects, including results from MethylMix and additional evidence from ELMER.

Column Descriptions:
  • Cancer Project: The abbreviated name of the cancer project (e.g., LUAD-TCGA).
  • Gene symbol: The official symbol of the selected gene (e.g., EGFR).
  • ENSG: The Ensembl gene identifier corresponding to the selected gene.
  • MethylMix Hyper percent: The proportion of samples classified as hypermethylated by MethylMix.
  • MethylMix Hypo percent: The proportion of samples classified as hypomethylated by MethylMix.
  • MethylMix None percent: The proportion of samples classified as having no methylation change by MethylMix.
  • ELMER Probe: The probe associated with the selected gene in the ELMER result.
  • ELMER Distance: The genomic distance between the ELMER probe and the selected gene.
  • ELMER Adjusted p-value: The adjusted p-value reported for the ELMER result.
  • ELMER Methylation type: Indicates whether the ELMER-associated methylation alteration is classified as hypermethylation or hypomethylation.
  • Beta-value vs. expression Spearman correlation coefficient: Spearman correlation coefficient (ρ) measuring the association between beta value and gene expression. Negative values indicate that higher beta values are associated with lower expression, whereas positive values indicate that higher beta values are associated with higher expression.
  • Beta-value vs. expression Spearman p-value: The p-value assessing the statistical significance of the Spearman correlation between beta value and gene expression.

The Spearman correlation coefficient (ρ) indicates the direction and strength of the association between beta value and gene expression, while the corresponding p-value indicates its statistical significance.

2.8.5 Survival Map & Survival Analysis

The Survival Map and Survival Analysis sections evaluate the prognostic relevance of DNA methylation of the selected gene across cancer cohorts and survival endpoints using multiple analysis methods. Survival results are available for TCGA cohorts only.

Detailed information on interpreting the Survival Map and Survival Analysis panels is provided in Gene-Survival.


2.9 Gene miRNA

2.9.1 Overview

The Gene miRNA module visualizes regulatory relationships between the selected gene and its associated miRNAs across cancer types, integrating predicted interactions, experimentally validated evidence, and expression-correlation information.

Gene–miRNA interaction predictions are integrated from 12 prediction tools, while experimentally validated miRNA–target interactions are obtained from miRTarBase. Detailed information on the prediction tools, scoring logic, data sources, and methodology is provided in FAQ4.

This module contains two result sections:
  1. Gene–miRNA Interaction Network — provides an interactive visualization of predicted and experimentally validated gene–miRNA interactions with flexible filtering by gene source, interaction type, and prediction support.
  2. Gene–miRNA Correlation Table — provides cancer-specific interaction evidence and expression-correlation statistics for individual gene–miRNA pairs.

2.9.2 Gene-miRNA Interaction Network

This interactive network displays regulatory relationships between the selected gene and associated miRNAs, integrating predicted and experimentally validated interaction evidence across cancer types.

Data Sources

Predicted gene–miRNA interactions are integrated from 12 prediction tools, while experimentally validated miRNA–target interactions are obtained from miRTarBase. Expression correlations between gene–miRNA pairs are additionally provided in the Gene–miRNA Correlation Table. Detailed information on the prediction tools, scoring logic, references, and expression-correlation methodology is provided in FAQ4.

Network Representation

The network represents genes and miRNAs as nodes and their interactions as edges. Green nodes represent genes, whereas yellow diamond-shaped nodes represent miRNAs.

Edge style indicates validation evidence: solid lines indicate interactions with validated evidence, whereas dashed lines indicate predicted interactions without validated evidence.

Edge color indicates how many cancer types support an interaction: light grey represents one cancer type, dark grey represents two cancer types, and black represents three or more cancer types.

Filtering Options

Users can refine the network using the following filters:
  1. Gene Source: Filters genes according to CGC, NCG, or All.
  2. Interaction Type: Displays predicted interactions, validated interactions, or both. When both are selected, interactions meeting either criterion are displayed.
  3. Prediction Support: Filters predicted interactions according to the number of supporting prediction tools. Users can display any predicted interaction or require support from at least 6, 8, or 10 tools. The selected prediction-support threshold does not exclude interactions with validated evidence when validated interactions are selected.

Network Interaction

Users can select a gene or miRNA from the search field to locate and highlight its interaction neighborhood. Clicking a node highlights the selected node and its directly connected interactions, while clicking empty space restores the complete network. The Re-arrange function recalculates the network layout to provide an alternative arrangement of the displayed nodes and edges.

Interpretation

The network allows users to distinguish predicted from experimentally validated gene–miRNA interactions, examine the level of computational prediction support, and assess how broadly individual interactions are supported across cancer types. Together, these features provide an interactive overview of computational and experimental evidence for potential miRNA-mediated regulation of the selected gene.



2.9.3 Gene-miRNA Correlation Table

The Gene–miRNA Correlation Table provides cancer-specific interaction evidence and expression-correlation statistics for individual gene–miRNA pairs.

Column Descriptions:
  • Cancer type abbreviation: The abbreviated cancer type name (e.g., LUAD, BRCA).
  • miRNA: The miRNA associated with the selected gene.
  • Gene symbol: The official symbol of the selected gene (e.g., EGFR).
  • ENSG: The Ensembl gene identifier corresponding to the selected gene.
  • Validated: An in-house categorical score derived from miRTarBase miRNA–target evidence; higher scores indicate stronger experimental support for the miRNA–gene interaction.
    See miRTarBase for details.
  • Number of tool: The number of prediction tools supporting the miRNA–gene interaction.
  • Pearson correlation coefficient: Pearson correlation coefficient measuring the expression correlation between the gene and miRNA.
  • Pearson p-value: The p-value assessing the statistical significance of the Pearson correlation.
  • Spearman correlation coefficient: Spearman correlation coefficient measuring the expression correlation between the gene and miRNA.
  • Spearman p-value: The p-value assessing the statistical significance of the Spearman correlation.
  • Kendall correlation coefficient: Kendall correlation coefficient measuring the expression correlation between the gene and miRNA.
  • Kendall p-value: The p-value assessing the statistical significance of the Kendall correlation.

Negative correlation coefficients indicate that higher miRNA expression tends to be associated with lower expression of the selected gene, whereas positive coefficients indicate that gene and miRNA expression tend to vary in the same direction. The corresponding p-values indicate the statistical significance of each correlation.

Note: The correlation results shown in this table are based on TCGA cohorts.


2.10 Gene Protein

2.10.1 Overview

The Gene Protein module visualizes protein-level variation of the selected gene across cancers and examines how protein abundance relates to mRNA expression and post-translational modifications (PTMs).

Analyses are organized into three tabs:
  1. Clinical Stages – grouped by clinical tumor stages
  2. Mutation Classes – grouped by mutation impact levels
  3. PTM Sites – grouped by specific phosphorylation sites (e.g., pY1068, pY1173)

All analyses support interactive exploration, including sample-level tooltips, togglable groups, and mRNA–protein scatter plots.

2.10.2 Clinical Stages

This tab evaluates how protein expression and mRNA–protein associations vary across clinical tumor stages.

Protein Expression by Clinical Stage (Pan-Cancer)

Purpose:

Visualizes protein expression levels of the selected gene across all TCGA cancer types, grouped by stage.

Plot Features:
  • Boxplots display protein abundance for each cancer type (x-axis), grouped by stage (colors).
  • Legend toggling: Show or hide specific stages (e.g., Stage I, Stage II, Stage III, Stage IV).
  • Hover interactions:
    • Hover over dots → sample-level details (sample ID, expression, tissue).
    • Hover near box areas → summary statistics (median, Q1/Q3, upper/lower fences, min/max).
  • Optional PTM selection: Choose to display None, pY1068, or pY1173 to inspect PTM-specific protein patterns.
Interpretation:
Differences across stages may indicate stage-dependent dysregulation of protein abundance.

mRNA–Protein Correlation by Clinical Stage (Pan-Cancer)

Purpose:

Assesses whether mRNA abundance explains protein expression patterns across cancers within each stage group.

Plot Features:
  • Bar chart showing Spearman correlation coefficients (ρ) between mRNA (FPKM-UQ) and protein expression across cancer types.
  • Bars are grouped by clinical stage.
  • Toggle individual stages via the legend.
  • Hover interactions: Cancer type, stage, Spearman ρ, p-value.
  • Click bar → opens a scatter plot (mRNA vs. protein), including the correlation and p-value.
Interpretation:
  • ρ > 0: mRNA and protein increase together → transcriptionally consistent regulation.
  • ρ < 0: expression moves in opposite directions → post-transcriptional regulation or translational inhibition.


Cancer-Specific: Stage-Specific Protein Expression

This visualization shows protein expression patterns within a selected cancer type grouped by tumor stage, displaying a violin plot with stage groups (I–IV) on the x-axis and protein expression levels on the y-axis, where users can toggle stages using the legend and hover over dots to view sample-level information or hover over violin areas to see statistical summaries including median, quartiles, and fences. An accompanying statistical table compares pairs of stages with columns showing Group 1, Group 2, p-value, significance level, and sample counts, indicating whether stage-specific differences are statistically significant (p < 0.05) and helping users determine if protein expression changes progressively across disease stages or shows distinct patterns at specific stages of cancer development.

2.10.3 Mutation Classes

This tab evaluates how protein expression varies across mutation impact categories and how mutation classes influence mRNA–protein correlations.

Mutation impact groups:
  • High
  • Moderate
  • Low
  • Modifier
  • Normal tissue
  • Tumors without mutation

Protein Expression by Mutation Class (Pan-Cancer)

Purpose:

Visualizes protein expression across cancers grouped by mutation impact level.

Plot Features:
  • Boxplots grouped by impact class,, one set per cancer type
  • Legend toggling for impact classes
  • Hover for sample-level details and boxplot summary statistics
  • Optional PTM filtering (None, pY1068, pY1173)
Interpretation:

Allows users to assess whether specific impact classes (e.g., high-impact mutations) correspond to altered protein levels.



mRNA–Protein Correlation by Mutation Class (Pan-Cancer)

Purpose:

Examines mRNA–protein concordance across mutation-defined sample groups.

Features:
  • Bar chart of Spearman ρ for each cancer type, grouped by mutation impact
  • Hover for cancer type, impact class, ρ, p-value
  • Click bar → opens the corresponding mRNA–protein scatter plot
Interpretation:
  • Positive ρ: protein expression tracks mRNA → transcriptionally driven response
  • Negative ρ: mutation-class–specific post-transcriptional or PTM-dependent regulation


Cancer-Specific: Mutation-Impact–Specific Protein Expression

This visualization shows protein expression differences within a selected cancer type grouped by mutation class, displaying a violin plot with mutation impact class on the x-axis and protein expression levels on the y-axis, where users can hover for statistical summaries and sample information and toggle impact classes using the legend. An accompanying statistical table provides pairwise comparisons between impact classes with columns showing Group 1, Group 2, p-value, significance level, and sample counts. This analysis helps identify whether high-impact mutation carriers show altered protein levels relative to other mutation groups, revealing whether mutations influence not only gene expression at the transcript level but also at the protein level, which may have more direct functional consequences for cancer phenotypes.

2.10.4 PTM Sites

This tab evaluates how post-translational modifications (PTMs)—specifically phosphorylation sites—modify the relationship between mRNA and protein expression.

mRNA-Protein Correlation by PTM Site (Pan-Cancer)

Purpose:

Assesses how PTMs (e.g., pY1068, pY1173) influence mRNA–protein coupling across cancers.

Plot Features:
  • Bar chart of Spearman correlation coefficients (ρ) across cancer types
  • Groups correspond to PTM sites:
    • None (total protein)
    • pY1068
    • pY1173
  • Toggle PTM sites using the legend
  • Hover for Cancer type, PTM site, ρ, p-value
  • Click bar → opens a PTM-specific mRNA–protein scatter plot
Interpretation of ρ:
  • Positive ρ (> 0):PTM-site–specific protein levels track mRNA → transcriptionally driven regulation
  • Negative ρ (< 0):Protein/PTM levels diverge from mRNA → post-transcriptional or PTM-dependent modulation
    (e.g., phosphorylation buffering, kinase pathway activation independent of transcript levels)
This analysis reveals whether phosphorylation alters mRNA–protein consistency across cancers.

2.10.5 Survival

Purpose

This analysis evaluates whether total protein abundance or site-specific post-translational modification abundance is associated with patient survival across selected cancer cohorts.

When phosphorylation data are available, results are presented separately for:
  • Total protein, representing overall protein abundance.
  • PTM sites, representing the abundance of individual phosphorylation sites, such as pY1068 or pY1173.

Comparing total-protein and PTM-site results can reveal site-specific prognostic associations that may not be apparent from overall protein abundance.

Analysis Workflow

Users first select one of three survival-analysis methods:
  1. Cox Uni — univariate Cox proportional hazards analysis.
  2. Cox Multi (Clinical) — multivariable Cox proportional hazards analysis with clinical covariate adjustment.
  3. Cure Model — survival analysis designed to capture both short-term risk and long-term survival patterns.
For Cox Uni and Cox Multi, users select:
  • Cancer type: the patient cohort to nanalyze.
  • Survival type: the clinical endpoint.
  • Survival time: either the full available follow-up period or follow-up restricted to 5 years.
  • Stratification method: median or best cutpoint.

For the Cure Model, users currently select only the cancer type. The survival endpoint is overall survival (OS), follow-up includes all available years, and abundance stratification is based on the median abundance of the analyzed protein feature.

Survival Endpoint

The available survival endpoints are:
  • Overall Survival (OS): time from diagnosis or study entry to death from any cause.
  • Progression-Free Interval (PFI): time to disease progression, recurrence, a new primary tumor, or death, according to the endpoint definition used in the source cohort.
  • Disease-Free Interval (DFI): time from completion of initial treatment or achievement of disease-free status to recurrence or a new disease event.
  • Disease-Specific Survival (DSS): time to death attributed to the cancer under study; deaths from other causes are generally censored.

Endpoint availability may vary among cancer cohorts.

Follow-up Time

Users may analyze:
  • All time: all available follow-up data are included.
  • 5 years: follow-up is restricted to the first 60 months.

For a 5-year analysis, patients who remain event-free beyond 60 months should normally be administratively censored at 60 months rather than excluded. The plot x-axis and risk estimates then represent outcomes during the first five years of follow-up.

Sample Stratification

For each total-protein or PTM-site feature, patients with valid abundance and survival data are divided into abundance groups before plotting the survival curves.

Median Stratification

The median abundance value among the eligible patients is used as the cutpoint.

  • Patients above the median are assigned to the High group.
  • Patients below the median are assigned to the Low group.

Median stratification is simple, reproducible, and generally produces groups of similar size. However, it may miss an association when the biologically relevant threshold is not close to the median.

Best-Cutpoint Stratification

A set of candidate abundance thresholds is evaluated, and the cutpoint that produces the strongest separation between the survival groups is selected.

The optimal threshold is commonly chosen by maximizing a survival-separation statistic, such as the standardized log-rank statistic, or equivalently by minimizing the corresponding log-rank p-value within an allowed range of cutpoints.

Patients are then classified as:
  • High: abundance above the selected cutpoint.
  • Low: abundance at or below the selected cutpoint, or according to the exact boundary rule used by the application.

Candidate cutpoints should be restricted so that neither group becomes too small. Because the same dataset is used to select and test the threshold, best-cutpoint results can overestimate effect size and statistical significance. These results should therefore be interpreted as exploratory and ideally validated in an independent cohort.

Cox Uni

Cox Uni evaluates one molecular feature at a time without adjustment for clinical characteristics. The feature may be total protein abundance or the abundance of an individual PTM site.

The Cox model estimates the relative hazard for the High group compared with the Low group. A hazard ratio greater than 1 indicates a higher event rate in the High group, whereas a hazard ratio below 1 indicates a lower event rate.

Kaplan-Meier Plot

The Kaplan–Meier plot displays the observed survival probability over time for the High and Low abundance groups.

Greater separation between the curves indicates a larger difference in survival experience. The log-rank p-value tests whether the survival distributions differ between the groups.

The Kaplan–Meier plot is an unadjusted comparison and does not account for clinical covariates.

Cumulative Hazard Plot

The cumulative hazard plot displays the accumulated event hazard over time for the same High and Low groups.

A curve that rises more rapidly indicates faster accumulation of risk. Separation between the curves suggests different event rates between the abundance groups.

The cumulative hazard plot complements the Kaplan–Meier plot but should not be interpreted as the instantaneous hazard at a particular time point.

Interpretation

A significant result suggests that the selected protein or PTM-site abundance is associated with the selected survival endpoint in an unadjusted analysis. It does not establish that the feature is independent of tumor stage, age, or other clinical factors.

Cox Multi (clinical)

Cox Multi fits a multivariable Cox proportional hazards model to evaluate the association between a protein or PTM feature and survival after accounting for available clinical covariates.

Depending on the cohort and data availability, covariates may include variables such as age, sex, tumor stage, grade, or other relevant clinical characteristics.

Kaplan-Meier Plot

The Kaplan–Meier plot displays the observed, unadjusted survival experience of the High and Low abundance groups.

Although it is shown alongside the multivariable analysis, the Kaplan–Meier curve itself does not adjust for clinical covariates. Clinical adjustment is provided by the Cox model and summarized in the forest plot.

When the plot is generated from a multivariable Cox model, it is labelled adjusted survival curve rather than Kaplan–Meier plot.

Forest Plot The forest plot summarizes the hazard ratio and 95% confidence interval for:
  • The High-versus-Low protein or PTM abundance group.
  • Each clinical covariate included in the model.
Interpretation of the hazard ratio is as follows:
  • HR > 1: higher hazard, corresponding to a worse outcome for the modeled comparison.
  • HR < 1: lower hazard, corresponding to a better outcome.
  • HR = 1: no estimated difference in hazard.

The dashed vertical reference line at HR = 1 represents no association. A confidence interval that crosses 1 indicates that the effect is not statistically distinguishable from no association at the corresponding confidence level.

Interpretation

If the protein or PTM abundance group remains statistically significant after clinical adjustment, the result supports an association with survival that is not explained by the clinical covariates included in the model.

This should be described as an independent association within the fitted model, rather than proof that the feature is biologically or causally independent. Residual confounding may remain, and results depend on the quality and availability of the clinical variables.

Cure Model

The cure model is intended for survival settings in which a proportion of patients may experience sustained long-term survival and no longer show the same event risk as the susceptible patient population.

Unlike a standard Cox model, a cure model can separately characterize:
  • Long-term survival or cure fraction: the estimated probability of belonging to the long-term event-free group.
  • Short-term survival among susceptible patients: the timing or risk of events among patients who remain at risk.

The term “cure” is statistical and does not necessarily indicate confirmed clinical eradication of disease.

Survival Plots

Cure-model results are displayed for total protein and, when available, separately for each PTM site.

The plots may show:
  • Long-term survival: differences in the estimated long-term-surviving or cured fraction between abundance groups.
  • Short-term survival: differences in survival among patients considered susceptible to the event.

Separation between the High and Low curves suggests group-specific survival behavior. A plateau in the late portion of a survival curve may be consistent with a long-term-surviving fraction, although a plateau can also result from limited follow-up or few patients remaining at risk.

Interpretation

Differences in the long-term component suggest that abundance may be associated with the estimated long-term-surviving fraction. Differences in the short-term component suggest an association with event timing among susceptible patients.

Comparisons between total protein and individual PTM sites can identify site-specific long-term or short-term survival patterns that are not reflected by total protein abundance


2.11 Gene Multi-omics

2.11.1 Overview

The Gene Multi-Omics interface summarizes multi-omics driver evidence for the user-selected gene across cancer projects, integrating results from RNA, mutation, CNV, and methylation analyses. Results are organized by omics type, cancer project, and supporting integration tools to show the distribution and extent of support for the selected gene across different cancer contexts.

This interface includes three result sections:
  • Integrated Multi-Omics Overview
  • Omics Connectivity Network
  • Multi-Omics Driver Event Table

Together, these sections provide complementary views of the distribution of multi-omics driver evidence across omics types, cancer projects, and supporting integration tools.

2.11.2 Integrated Multi-Omics Overview

The Integrated Multi-Omics Overview summarizes multi-omics driver evidence for the selected gene across cancer projects and omics types. The visualization consists of two bar charts and a combination matrix.

Top Bar Chart — Tool Evidence Across Cancer Projects

The top bar chart shows the accumulated tool evidence for each cancer project across omics types. Each tool-supported event contributes to the count; therefore, the bar height reflects the overall level of supporting evidence rather than the number of unique tools.

Left Bar Chart — Tool Evidence Across Omics Types

The left bar chart shows the accumulated tool evidence for each omics type across cancer projects. Each tool-supported event contributes to the count; therefore, the bar length reflects the overall level of supporting evidence rather than the number of unique tools.

Combination Matrix — Omics Evidence Across Cancer Projects

The combination matrix indicates which omics types contribute evidence in each cancer project. Solid dots indicate omics types with supporting evidence in the corresponding cancer project, while connected dots indicate evidence from multiple omics types within the same project.



2.11.3 Omics Connectivity Network

The Omics Connectivity Network illustrates how multi-omics evidence for the selected gene is distributed across omics types and cancer projects.

Nodes represent the selected gene, omics types, and cancer projects. Links connect the selected gene to supported omics types and the corresponding cancer projects, providing a Gene → Omics → Cancer Project view of the evidence. Link width reflects the number of represented gene–omic–project relationships.

Cancer projects belonging to the same cancer type are displayed using the same node color to facilitate comparison across related datasets.



2.11.4 Multi-Omics Driver Event Table

The Multi-Omics Driver Event Table provides detailed multi-omics driver evidence for the selected gene across omics types and cancer projects.

Column Descriptions:
  • Gene: Official gene symbol.
  • Omic: Omics type associated with the evidence.
  • Cancer Project: Cancer project or cohort in which the evidence was identified.
  • CGC / NCG: Indicates whether the gene is annotated in CGC or NCG.
  • tool: Integration tool(s) supporting the event.
  • nTool: Number of distinct supporting tools for that gene–omic–project combination.

The table allows users to examine which integration tools contribute evidence for the selected gene within individual omics types and cancer projects. Unlike the accumulated Tool Evidence shown in the Integrated Multi-Omics Overview, nTool represents the number of distinct supporting tools within each gene–omic–project combination.

2. Gene Survival

Survival Map

The Survival Map displays the survival impact of the selected gene across multiple cancer types and four survival endpoints: overall survival (OS), progression-free interval (PFI), disease-free interval (DFI), and disease-specific survival (DSS). The map supports multiple omics data types, including RNA expression, mutation, copy number variation (CNV), and methylation.

  1. Cancer type abbreviations
    Cancer types are shown using abbreviations. In the visualizations, the -TCGA suffix is omitted from cancer type labels; therefore, when looking up full cancer type names, please add the -TCGA suffix to the displayed cancer type abbreviation.
  2. Survival significance and hazard ratio

    Each heatmap cell represents a combination of cancer type and survival analysis method. For Cox univariate, Cox multivariate, and cure model analyses, a colored cell indicates that the selected molecular feature is significantly associated with survival based on the hazard ratio and p-value. The selected molecular feature may represent RNA expression level, mutation status, CNV status, or methylation level, depending on the selected omics type.

    The color gradient represents the hazard ratio: red indicates a hazard ratio greater than 1, suggesting higher risk, while blue indicates a hazard ratio less than 1, suggesting lower risk. The direction of risk is interpreted relative to the omics-specific reference group used in the analysis — for example, high versus low RNA expression, mutated versus wild-type status, or copy number gain or loss relative to neutral status. For methylation data, the reference group depends on the grouping method: high versus low methylation level when using beta-value median stratification, or hypermethylated versus hypomethylated status when using MethylMix-based classification.

    For machine learning–based results, a colored cell indicates that the selected gene was identified as survival-related by at least one of the machine learning algorithms — LASSO, Random Forest, or I-Boost. To determine which specific algorithm identified the gene, users can refer to the detailed results in the Survival Analysis section.

    Hover over a colored cell to view detailed survival information, including cancer type, survival endpoint, analysis method, omics type, grouping or stratification method, hazard ratio, log-rank p-value, and cutoff or grouping value when available.

    Grouping methods depend on the selected omics type. RNA expression is stratified using the best cutoff, which identifies the threshold that maximizes survival difference between groups, or the median cutoff. Mutation data are grouped by mutated versus wild-type status. CNV data are grouped using iGC or GISTIC-based copy number calls into gain, loss, or neutral categories. For methylation data, patients are grouped using beta-value median stratification, beta-value best cutoff stratification, or MethylMix-based classification.

    Note that p-values for Cox univariate, Cox univariate 5-year, and machine learning analyses are calculated using the log-rank test, while p-values for Cox multivariate and Cox multivariate 5-year analyses are calculated using the Cox proportional hazards model.

  3. Survival analysis methods

    The map includes multiple survival analysis approaches: Cox univariate regression, Cox multivariate regression adjusted for clinical covariates, cure model analysis, and machine learning–based survival analysis using LASSO, Random Forest, and I-Boost.

    Available survival analysis methods vary by survival endpoint. For the OS endpoint, available results include Cox univariate regression, Cox univariate regression 5-year, Cox multivariate regression adjusted for clinical covariates, Cox multivariate regression 5-year adjusted for clinical covariates, cure model short-term effect, cure model long-term effect, and machine learning–based results. For the PFI, DFI, and DSS endpoints, available results include Cox univariate regression, Cox univariate regression 5-year, Cox multivariate regression adjusted for clinical covariates, Cox multivariate regression 5-year adjusted for clinical covariates, and machine learning–based results. Cure model results are available only for the OS endpoint.

References of all survival analysis methods, including the machine learning–based approaches, are available in FAQ4. Click a colored cell to open the corresponding Kaplan–Meier plot. If the selected cell represents a machine learning result, the detailed output opens in a new tab. For figure and table manipulation, please refer to FAQ3.

Survival Analysis

The Survival Analysis panel evaluates whether the selected gene's molecular features are associated with patient prognosis. Users first select a survival analysis type:
  • Cox Univariate (Cox Uni) – Evaluates the association between the selected molecular feature and patient survival without adjusting for additional clinical variables.
  • Cox Multivariate (clinical) (Cox Multi) – Evaluates the association between the selected molecular feature and patient survival while adjusting for available clinical covariates, such as age, gender, stage, or other cohort-specific clinical variables.
  • Cure Model – Models survival patterns while accounting for the possibility that a subset of patients may experience long-term survival or reduced risk over time. This framework supports evaluation of both short-term and long-term survival effects.
  • Machine Learning – Uses supervised learning-based approaches to identify molecular features or signatures associated with prognosis. Available methods include LASSO, Random Forest, and I-Boost.
  • Trans-Omics Synergistic Effect – Evaluates whether two molecular features, potentially from different omics layers, have a combined or interaction-based prognostic effect on patient survival.

References of all survival analysis methods, including the machine learning–based approaches, are available in FAQ4.

After an analysis type is selected, the available filters and result views update accordingly.

Depending on the selected analysis framework, users can choose relevant options such as cancer type, survival endpoint, survival time, stratification method, or machine learning algorithm. The resulting plots and tables help compare survival patterns, estimate risk differences, and identify molecular features or interactions associated with patient outcomes.

Cox Univariate

The Cox Univariate results section evaluates whether the selected molecular feature of the user-selected gene is associated with patient survival in a selected cancer type. The analysis is performed using univariate Cox proportional hazards regression and Kaplan–Meier survival analysis, without adjusting for any clinical covariates.

Users can select a cancer type, survival endpoint, survival time, and grouping method from the dropdown menus. The available survival endpoints include overall survival (OS), progression-free interval (PFI), disease-free interval (DFI), and disease-specific survival (DSS). The survival time options include all-time and 5-year analyses.

Patient grouping depends on the selected omics type. For RNA expression data, patients are stratified into high- and low-expression groups using either the median cutoff or the best cutoff, which identifies the expression threshold that maximizes the log-rank test statistic across all possible cutpoints to produce the most statistically significant survival separation between groups. For mutation data, patients are grouped by mutation status — mutated or wild-type. For CNV data, patients are grouped into copy number gain, loss, or neutral categories based on copy number status defined by either iGC or GISTIC. For methylation data, patients are grouped using beta-value median stratification, beta-value best cutoff stratification, or MethylMix-based classification. For beta-value stratification methods, patients are divided into high and low methylation level groups; for MethylMix-based classification, patients are divided into hypermethylated and hypomethylated groups.

After the selections are made, two Kaplan–Meier survival plots are displayed. The left plot shows the unadjusted survival curves based solely on the selected omics-specific patient grouping, reflecting the univariate association between the molecular feature and survival. The right plot shows covariate-adjusted survival curves generated from a separate Cox model that accounts for available clinical covariates such as age, gender, stage, or other cohort-specific variables; this plot is displayed only when sufficient clinical covariate data are available for the selected cancer type and endpoint.

Each plot displays survival analysis results above the figure, such as hazard ratio and p-value. The x-axis represents survival time starting from the initial cancer diagnosis, and the y-axis represents survival probability.

Patient groups are defined according to the selected omics type:
  • RNA expression: high versus low expression of the selected gene.
  • Mutation: mutated versus wild-type status of the selected gene.
  • CNV: copy number gain, loss, or no copy number variation of the selected gene, depending on the available grouping.
  • Methylation: high versus low methylation level of the selected gene (when using beta-value median stratification), or hypermethylated versus hypomethylated status (when using MethylMix-based classification).

Users can hover over the curves to view detailed survival information at specific time points. Curves can also be shown or hidden by clicking the corresponding labels in the legend. For figure and table manipulation, please refer to FAQ3. For algorithm descriptions and references, please refer to FAQ4.

Cox Multivariate (clinical)

The Cox Multivariate results section evaluates whether the selected molecular feature of the user-selected gene is independently associated with patient survival after adjusting for available clinical covariates. The analysis is performed using multivariate Cox proportional hazards regression across multiple cancer types.

Users can select a cancer type, survival endpoint, survival time, and grouping method from the dropdown menus. The available survival endpoints include overall survival (OS), progression-free interval (PFI), disease-free interval (DFI), and disease-specific survival (DSS). The survival time options include all-time and 5-year analyses.

Patient grouping depends on the selected omics type. For RNA expression data, patients are stratified into high- and low-expression groups using either the median cutoff or the best cutoff, which identifies the expression threshold that maximizes the log-rank test statistic across all possible cutpoints to produce the most statistically significant survival separation between groups. For mutation data, patients are grouped by mutation status — mutated or wild-type. For CNV data, patients are grouped into copy number gain, loss, or neutral categories based on copy number status defined by either iGC or GISTIC. For methylation data, patients are grouped using beta-value median stratification, beta-value best cutoff stratification, or MethylMix-based classification. For beta-value stratification methods, patients are divided into high and low methylation level groups; for MethylMix-based classification, patients are divided into hypermethylated and hypomethylated groups.

After the selections are made, the results display covariate-adjusted survival curves and a forest plot. The adjusted survival curves show model-estimated survival differences among patient groups defined by the selected omics-specific grouping method, after accounting for available clinical covariates such as age, gender, stage, or other cohort-specific variables. The x-axis represents survival time from initial cancer diagnosis, and the y-axis represents survival probability.

Patient groups are defined according to the selected omics type:
  • RNA expression: high versus low expression of the selected gene.
  • Mutation: mutated versus wild-type status of the selected gene.
  • CNV: copy number gain, loss, or no copy number variation of the selected gene, depending on the available grouping.
  • Methylation: high versus low methylation level of the selected gene (when using beta-value median stratification), or hypermethylated versus hypomethylated status (when using MethylMix-based classification).

Users can hover over the survival curves to view detailed survival information at specific time points. Curves can also be shown or hidden by clicking the corresponding labels in the legend.

The forest plot summarizes the hazard ratios and 95% confidence intervals for the selected molecular feature and all clinical covariates included in the multivariate Cox model. Each row represents one variable, with the point estimate indicating the hazard ratio and the horizontal line indicating the confidence interval. Values greater than 1 indicate higher risk and values less than 1 indicate lower risk relative to the reference group. The forest plot helps users compare the relative association of each variable with survival after mutual adjustment. Clicking on the forest plot opens a full-sized version for closer inspection.

For figure and table manipulation, please refer to FAQ3. For algorithm descriptions and references, please refer to FAQ4.

Cure Model

The Cure Model results section evaluates whether the selected molecular feature of the user-selected gene is associated with long-term and short-term survival outcomes. The cure model estimates two types of survival effects: short-term and long-term effects. The short-term effect reflects the association between the selected molecular feature and survival time among patients who remain at risk of the event. The long-term effect reflects the association between the selected molecular feature and the probability of long-term survival, or the estimated cured fraction. Short-term and long-term p-values are reported to indicate whether the selected molecular feature is significantly associated with each component of the cure model. The short-term and long-term p-values are displayed above the plot alongside the other survival statistics.

Cure model results are available only for overall survival (OS) using all-time survival data. Patient grouping depends on the selected omics type. For RNA expression data, patients are separated into high- and low-expression groups using the median cutoff. For mutation data, patients are separated into mutated and wild-type groups. For CNV data, patients are grouped into copy number gain, loss, or neutral categories based on copy number status defined by either iGC or GISTIC. For methylation data, patients are grouped using either beta-value median stratification or MethylMix-based methylation states.

Users can select a cancer type from the dropdown menu to view the corresponding cure model results.

After a cancer type is selected, the estimated survival curves are displayed. The values calculated by the survival analysis are shown above the plot. The x-axis represents survival time starting from the initial cancer diagnosis, and the y-axis represents survival probability.

Patient groups are defined according to the selected omics type:
  • RNA expression: high versus low expression of the selected gene.
  • Mutation: mutated versus wild-type status of the selected gene.
  • CNV: copy number gain, loss, or no copy number variation of the selected gene, depending on the available grouping.
  • Methylation: high versus low methylation level of the selected gene (when using beta-value median stratification), or hypermethylated versus hypomethylated status (when using MethylMix-based classification).

Users can hover over the curves to view detailed survival information at specific time points. Curves can also be shown or hidden by clicking the corresponding labels in the legend. For figure and table manipulation, please refer to FAQ3. For algorithm descriptions and references, please refer to FAQ4.

Machine Learning

Machine Learning–based survival analysis identifies molecular features associated with patient survival using three algorithms: LASSO, Random Forest, and I-Boost. Each method builds a multi-feature prognostic signature, and patients are stratified into high- and low-risk groups based on their composite signature score. Results are presented as Kaplan–Meier survival curves and ROC curves evaluating the predictive performance of each signature.

LASSO (Least Absolute Shrinkage and Selection Operator)

The LASSO results section displays significant survival signatures involving the user-selected gene across 33 cancer types. LASSO is a regression-based method that selects the most relevant survival-related molecular features by shrinking less informative gene coefficients toward zero. Features with non-zero coefficients are retained to build a prognostic signature.

In the visualizations, the -TCGA suffix is omitted from cancer type labels; therefore, when looking up full cancer type names, please add the -TCGA suffix to the displayed cancer type abbreviation.

  1. Significant LASSO results are shown in the signature selection table. The table includes cancer type, survival endpoint, significance, and the number of genes included in the signature. Selecting a specific result from the table displays the corresponding selected gene table, Kaplan–Meier survival plot, and ROC curves.
  2. The selected gene table lists the genes included in the LASSO signature and their coefficients. A positive coefficient indicates that higher expression of that gene is associated with worse survival (higher risk), shown in red. A negative coefficient indicates that higher expression is associated with better survival (lower risk), shown in blue. The interpretation of feature value depends on the selected data type — for example, expression level in RNA data, mutation presence versus wild-type in mutation data, copy number level in CNV data, or methylation level in methylation data. Users can reorder the table by clicking on any column name.
  3. The Kaplan–Meier plot displays survival differences between patient groups stratified by their composite LASSO signature score, which is computed as a weighted sum of all selected feature values using their LASSO coefficients. Patients are divided into two groups based on the median signature score of the cohort:
    • High: patients with a signature score above the median, indicating higher overall risk
    • Low: patients with a signature score below the median, indicating lower overall riskLow: patients with a signature score below the median, indicating lower overall risk

    Note that this grouping reflects the combined behavior of all features in the signature, not the value of any single feature. The survival statistics are shown above the plot. The x-axis represents survival time from initial cancer diagnosis, and the y-axis represents survival probability. Users can hover over the curves to view detailed survival information, and curves can be shown or hidden by clicking the corresponding legend labels.

  4. The ROC curves evaluate the predictive performance of the LASSO signature at different survival times. The x-axis represents the false-positive rate, and the y-axis represents the true-positive rate. Users can hover over the curves to view false-positive rate, true-positive rate, and cutoff value information. ROC curves for different survival times can be shown or hidden by clicking the corresponding labels in the legend.

For figure and table manipulation, please refer to FAQ3. For detailed algorithm descriptions and reference links, please refer to FAQ4.

Random Forest

The Random Forest results section displays significant survival signatures involving the user-selected gene across 33 cancer types. Random Forest is a machine learning method that uses many decision trees to identify molecular features that help distinguish different survival outcomes. Features are ranked based on their contribution to prediction performance.

In the visualizations, the -TCGA suffix is omitted from cancer type labels; therefore, when looking up full cancer type names, please add the -TCGA suffix to the displayed cancer type abbreviation.

  1. Significant Random Forest results are shown in the signature selection table. The table includes cancer type, survival endpoint, significance, and the number of genes included in the signature. Selecting a specific result from the table displays the corresponding selected gene table, Kaplan–Meier survival plot, and ROC curves.
  2. The selected gene table lists the molecular features included in the Random Forest signature, along with their depth and relative frequency within the forest:
    • Depth refers to how early a gene tends to appear in the decision trees. Gene at shallower depths (lower values) are used earlier in the trees, indicating stronger discriminative power for survival outcomes.
    • Relative frequency reflects how often a gene is used as a splitting variable across all trees in the forest, expressed as a proportion. Higher values indicate that the gene contributes more consistently to survival prediction across the model.
    The interpretation of feature values depends on the selected data type — for example, expression level in RNA data, mutation presence versus wild-type in mutation data, copy number level in CNV data, or methylation level in methylation data. The relative frequency column is color-coded according to its value. Users can reorder the table by clicking on any column name.
  3. The Kaplan–Meier plot displays survival differences between patient groups stratified by their composite Random Forest signature score, derived from the combined predictive output of all selected genes in the model. Patients are divided into two groups based on the median signature score of the cohort:
    • High: patients with a signature score above the median, indicating higher overall risk
    • Low: patients with a signature score below the median, indicating lower overall risk
    Note that this grouping reflects the combined behavior of all genes in the signature, not the expression level of any single gene. The survival statistics are shown above the plot. The x-axis represents survival time from initial cancer diagnosis, and the y-axis represents survival probability. Users can hover over the curves to view detailed survival information, and curves can be shown or hidden by clicking the corresponding legend labels.
  4. The ROC curves evaluate the predictive performance of the Random Forest signature at different survival time points. The x-axis represents the false-positive rate and the y-axis represents the true-positive rate. Users can hover over the curves to view the false-positive rate, true-positive rate, and cutoff value at each point. ROC curves for different survival times can be shown or hidden by clicking the corresponding legend labels.

For figure and table manipulation, please refer to FAQ3. For detailed algorithm descriptions and reference links, please refer to FAQ4.

I-Boost

The I-Boost results section displays significant survival signatures involving the user-selected gene across 33 cancer types. I-Boost is a boosting-based machine learning method that builds a survival prediction model by combining multiple weak predictors into a stronger signature. Features are selected and assigned coefficients based on their cumulative contribution to survival prediction boosting iterations.

In the visualizations, the -TCGA suffix is omitted from cancer type labels; therefore, when looking up full cancer type names, please add the -TCGA suffix to the displayed cancer type abbreviation.

  1. Significant I-Boost results are shown in the signature selection table. The table includes cancer type, survival endpoint, significance, and the number of genes included in the signature. Selecting a specific result from the table displays the corresponding selected gene table, Kaplan–Meier survival plot, and cumulative hazard plot.
  2. The selected feature table lists the molecular features included in the I-Boost signature along with their coefficients. A positive coefficient indicates that a higher feature value is associated with worse survival (higher risk), shown in red; a negative coefficient indicates that a higher feature value is associated with better survival (lower risk), shown in blue. The interpretation of feature value depends on the selected data type — for example, expression level in RNA data, mutation presence versus wild-type in mutation data, copy number level in CNV data, or methylation level in methylation data. Users can reorder the table by clicking on any column name.
  3. The Kaplan–Meier plot displays survival differences between patient groups stratified by their composite I-Boost signature score, computed from the combined weighted contributions of all selected features in the model. Patients are divided into two groups based on the median signature score of the cohort:
    • High: patients with a signature score above the median, indicating higher overall risk
    • Low: patients with a signature score below the median, indicating lower overall risk

    Note that this grouping reflects the combined behavior of all features in the signature, not the value of any single feature. The survival statistics are shown above the plot. The x-axis represents survival time from initial cancer diagnosis, and the y-axis represents survival probability. Users can hover over the curves to view detailed survival information, and curves can be shown or hidden by clicking the corresponding legend labels.

  4. The cumulative hazard plot displays the cumulative hazard over time for patients in the high and low signature score groups. Higher cumulative hazard values indicate a greater accumulated risk of the survival event occurring up to that time point. The survival statistics are shown above the plot. The x-axis represents survival time from initial cancer diagnosis, and the y-axis represents cumulative hazard. Users can hover over the curves to view detailed information, and curves can be shown or hidden by clicking the corresponding legend labels.

For figure and table manipulation, please refer to FAQ3. For detailed algorithm descriptions and reference links, please refer to FAQ4.

Trans-Omics Synergistic Effect

The trans-omics synergistic effect evaluates whether the selected gene shows a combined survival effect with molecular features from different omics layers within a cancer type. The analysis considers cross-omics interactions among RNA expression, mutation, copy number variation (CNV), and methylation, and is currently available for overall survival (OS) only.

For each interaction, the combined survival effect of the selected gene and its paired feature is compared with the survival effects of the individual features. The resulting synergistic interactions are summarized in a table, and selecting an interaction displays the corresponding Kaplan–Meier survival plots.

Result Table

The result table summarizes significant synergistic survival interactions involving the selected gene using the following columns:

  • cancer: Cancer type in which the interaction was identified.
  • interaction: Pair of omics layers involved in the interaction.
  • omic1 / omic2: Omics types corresponding to gene1 and gene2, respectively.
  • gene1 / gene2: Gene symbols of the two molecular features forming the interaction.
  • variation: CNV state (gain or loss) evaluated for interactions involving CNV; not applicable to interactions without CNV.
  • HR.FC: Synergy fold-change score comparing the combined effect with the stronger single-feature effect.
  • pval.logrank.pair: Log-rank p-value comparing survival among the combined-omics patient groups.
  • adjust.pval.pair: Benjamini–Hochberg adjusted log-rank p-value for the combined interaction.
  • HR.pair: Hazard ratio representing the combined effect of the paired features.

The table can be sorted by clicking any column header.

Kaplan–Meier Survival Plots

Selecting a row in the result table displays the corresponding Kaplan–Meier survival plots below.

The left plot shows the unadjusted Kaplan–Meier survival curves, whereas the right plot shows covariate-adjusted survival curves when sufficient clinical covariate data are available. Survival curves are displayed for the first 5 years of follow-up.

The x-axis represents survival time from the initial cancer diagnosis, and the y-axis represents survival probability. Users can hover over the curves to inspect survival information and click the legend to show or hide individual groups.

Synergy Score

The synergistic effect is quantified using HR.FC, which compares the hazard ratio of the combined feature pair with the stronger hazard ratio of the individual features.

The combined effect is represented by HR.pair, while the individual feature effects are represented by HR.single1 and HR.single2. The value HR.single corresponds to the stronger single-feature effect after accounting for the direction of the survival effect.

An HR.FC > 1 indicates that the combined feature pair shows a stronger survival effect than the stronger individual feature alone, after accounting for the direction of the survival effect.

Patient Group Stratification

Patient groups are defined according to the combined molecular states of the two paired features.

  • RNA expression: Patients are divided into high- and low-expression groups using the median expression cutoff.
  • Mutation: Patients are classified as mutated or wild-type.
  • CNV: Patients are classified as gain, loss, or neutral according to iGC.
  • Methylation: Patients are divided into high- and low-methylation groups using the median beta value.

Group Label Definitions

The labels shown in the Kaplan–Meier plots describe the combined molecular states of the two features. The order of gene1 and gene2 follows the interaction type shown in the result table.

RNA–Mutation
  • high_mut: High RNA expression in gene1 and mutation in gene2.
  • high_wt: High RNA expression in gene1 and wild-type gene2.
  • low_mut: Low RNA expression in gene1 and mutation in gene2.
  • low_wt: Low RNA expression in gene1 and wild-type gene2.
RNA–CNV
  • high_gain: High RNA expression in gene1 and CNV gain in gene2.
  • high_loss: High RNA expression in gene1 and CNV loss in gene2.
  • high_none: High RNA expression in gene1 and neutral CNV in gene2.
  • low_gain: Low RNA expression in gene1 and CNV gain in gene2.
  • low_loss: Low RNA expression in gene1 and CNV loss in gene2.
  • low_none: Low RNA expression in gene1 and neutral CNV in gene2.
RNA–Methylation
  • high_meth: High RNA expression in gene1 and high methylation in gene2.
  • high_unmeth: High RNA expression in gene1 and low methylation in gene2.
  • low_meth: Low RNA expression in gene1 and high methylation in gene2.
  • low_unmeth: Low RNA expression in gene1 and low methylation in gene2.
Mutation–CNV
  • mut_gain: Mutated gene1 and CNV gain in gene2.
  • mut_loss: Mutated gene1 and CNV loss in gene2.
  • mut_none: Mutated gene1 and neutral CNV in gene2.
  • wt_gain: Wild-type gene1 and CNV gain in gene2.
  • wt_loss: Wild-type gene1 and CNV loss in gene2.
  • wt_none: Wild-type gene1 and neutral CNV in gene2.
Mutation–Methylation
  • mut_meth: Mutated gene1 and high methylation in gene2.
  • mut_unmeth: Mutated gene1 and low methylation in gene2.
  • wt_meth: Wild-type gene1 and high methylation in gene2.
  • wt_unmeth: Wild-type gene1 and low methylation in gene2.
CNV–Methylation
  • gain_meth: CNV gain in gene1 and high methylation in gene2.
  • gain_unmeth: CNV gain in gene1 and low methylation in gene2.
  • none_meth: Neutral CNV in gene1 and high methylation in gene2.
  • none_unmeth: Neutral CNV in gene1 and low methylation in gene2.
  • loss_meth: CNV loss in gene1 and high methylation in gene2.
  • loss_unmeth: CNV loss in gene1 and low methylation in gene2.

Interpretation

A synergistic interaction indicates that the combined molecular states of two cross-omics features are associated with a stronger survival effect than the stronger individual feature alone. Such interactions may highlight complementary molecular alterations that are jointly associated with patient prognosis within the selected cancer type.

3. Customized Analysis

3.1 Overview

The Customized Analysis module enables researchers to perform user-defined analyses using clinical subgroups, gene features, and survival outcomes.
Unlike the Cancer and Gene modules—which summarize fixed results—Customized Analysis allows flexible, interactive, and hypothesis-driven exploration.

This module includes five major analytical categories:
  1. Subgroup Comparison Analyses
  2. Survival Analyses
  3. Multi-Omics Driver Analysis
  4. Prognostic Signature Identification
  5. Clinical Relevance Analysis (Multivariate Survival Analysis)

Below is detailed help for the Subgroup Comparison Analyses section.


3.2 Subgroup Comparison Analyses

Subgroup Comparison Analyses evaluate how gene-level molecular features differ across clinically defined patient subgroups.
Users define subgroups using any combination of clinical parameters (e.g., stage, grade, receptor status), and analyses are performed per selected gene and dataset.

Available comparison types:
  • Expression
  • Mutation
  • CNV
  • Methylation

Each analysis helps uncover biology associated with disease progression, risk groups, treatment response, or other clinically important factors.

3.2.1 Expression Subgroup Comparison

This analysis assesses whether gene expression varies across clinically defined patient subgroups within a cancer dataset.

Workflow

  1. Select a gene of interest.
  2. Choose a dataset.
  3. Define the analysis cohort by selecting one or more clinical criteria (e.g., Stage I+II, ER−, Grade 3).
    • Users may select multiple criteria simultaneously; sample counts update automatically.
  4. Choose a subgroup factor (e.g., stage, grade, receptor status), which determines the x-axis grouping in plots.

Output: Expression Comparison Across Clinical Subgroups

Violin Plot (log₁₀ TPM)
  • Displays expression distributions after log transformation.
  • Log scale reduces extreme variance and highlights differences between subgroup distributions.
  • Hover for summary statistics (median, quartiles, fences) or sample-level data.
Violin Plot (TPM)
  • Shows raw expression values without transformation.
  • Useful for interpreting absolute expression magnitude.
  • Hover to view sample-level details.
Statistical Comparison Table
Pairwise comparisons are automatically generated across subgroup levels.
Columns include:
  • Group1 / Group2
  • p-value
  • Significance
    • ns (≥0.05)
    • * (<0.05)
    • ** (<0.01)
    • *** (<0.001)
    • **** (<0.0001, optional)
  • Sample counts

Interpretation

Together, the violin plots and comparison table help determine whether gene expression differs meaningfully across clinical categories such as:
  • Tumor stage
  • Histologic grade
  • Receptor/HER2 status
  • Molecular subtype
  • Treatment response groups
Use this analysis to explore potential biomarkers or subgroup-specific molecular patterns.

3.2.2 Mutation Subgroup Comparison

This analysis evaluates whether mutation frequency of the selected gene differs between two clinically defined patient groups.

Workflow

  1. Select a gene and dataset.
  2. Define Group 1 and Group 2 using one or more clinical criteria (e.g., Stage I vs Stage III–IV, ER+ vs ER−).
    • Each group must contain ≥20 samples to ensure statistical validity.
  3. Run the analysis to generate contingency and statistical results.

Output

Mutation Contingency Table
Displays mutation counts for each group:
  • Group 1 Mutated / Wild-type
  • Group 2 Mutated / Wild-type
  • Total sample counts

This table summarizes how mutation events are distributed across the two subpopulations.

Fisher’s Exact Test Statistics
Computed to determine whether mutation frequencies differ significantly.
Outputs include:
  • Odds Ratio (OR)
    • OR > 1 → mutations more common in Group 1
    • OR < 1 → mutations more common in Group 2
  • 95% Confidence Interval (CI)
  • p-value (Fisher’s Exact Test)
Interpretation
Use this analysis to determine whether clinical subgroups differ in mutation burden for the selected gene.
Biological questions supported include:
  • Are late-stage tumors more mutated?
  • Do ER− patients have higher mutation frequency?
  • Are responders and non-responders genetically distinct?
This interpretation is essential for biomarker validation and subgroup-specific mutation profiling.

3.2.3 CNV Subgroup Comparison

The CNV Subgroup Comparison evaluates whether copy number variation (CNV) patterns differ between two clinically defined patient groups.
Events are categorized as Gain, Loss, or None (neutral).

Workflow

  1. Select a gene and dataset.
  2. Define Group 1 and Group 2 using one or more clinical criteria (e.g., Stage I vs Stage III–IV).
    • Each group must include ≥20 samples.
  3. Run the analysis to generate CNV contingency tables and Fisher’s Exact Test results.

Output

CNV Contingency Tables (Four Comparisons)
  1. All CNV Categories Combined
    • Shows counts of Gain, Loss, and None in both groups.
    • Provides an overall view of CNV distribution.
  2. Gain vs. Loss
    • Excludes neutral samples.
    • Tests whether amplification vs. deletion trends differ.
  3. Gain vs. None
    • Compares Gain against neutral CNV states.
  4. Loss vs. None
    • Compares Loss against neutral states.
Each 2×2 table includes:
  1. Odds Ratio (OR):
    • OR > 1 → CNV event more common in Group 1
    • OR < 1 → CNV event more common in Group 2
  2. 95% CI
  3. p-value (Fisher’s Exact Test)

Interpretation

This analysis reveals whether the selected gene exhibits different CNV profiles across clinical subpopulations—for example:
  • Amplification enriched in late-stage cancers
  • Deletion enriched in a specific molecular subtype
  • Neutral copy number predominance in certain risk groups
These insights help characterize subgroup-specific genomic alterations.

3.2.4 Methylation Subgroup Comparison

The Methylation Subgroup Comparison evaluates whether DNA methylation levels (β-values) differ across clinically defined patient subgroups.

Workflow

  1. Choose a gene and dataset.
  2. Filter the cohort by selecting clinical criteria.
  3. Select a subgroup factor (e.g., stage, grade, receptor status).
    • This determines the x-axis grouping.
  4. Run the analysis to generate violin plots and statistical comparisons.

Output

Methylation Violin Plot (β-values)
  1. Y-axis: β-value (0–1)
    • 0 = unmethylated
    • 1 = fully methylated
  2. X-axis: subgroup factor categories
  3. Each violin shows the methylation distribution within each subgroup.
  4. Hover for summary statistics (median, quartiles, fences) or sample-level details.
Statistical Comparison Table
Includes pairwise subgroup comparisons:
  • Group 1 / Group 2
  • p-value
  • Significance (ns, *, **, ***, ****)
  • Sample counts per subgroup

Interpretation

This analysis helps determine whether epigenetic regulation of the gene differs across patient groups—for example:
  • Hypermethylation enriched in high-grade tumors
  • Hypomethylation associated with specific receptor status
  • Stage-dependent methylation differences
These patterns can reveal clinically relevant epigenetic dysregulation.


3.3 Survival Analyses

Survival Analyses evaluate how mutations, gene expression levels, or miRNA expression levels influence patient outcomes within a user-defined subpopulation. Users define the patient cohort via clinical criteria and select stratification methods. This section begins with Mutation-Based Survival Analysis, followed by Expression-Based Survival Analysis and miRNA-Based Survival Analysis.

3.3.1 Mutation-Based Survival Analysis

The Mutation-Based Survival Analysis assesses whether mutations in the selected gene list are associated with survival differences in a clinically defined subpopulation.

Workflow

  1. Input a gene list.
  2. Select a dataset.
  3. Use clinical criteria to filter the patient cohort (e.g., Stage II only, ER+ only).
    • This determines which patients will be evaluated.
  4. The results tab then allows you to dynamically select:
    • Stratification method (Mutation vs. Wild type or Number of mutated genes)
    • Time interval (All follow-up or 5-year)

These two options control both the survival table and the Kaplan–Meier plots below.

Output

A. Mutation Oncoprint

This visualization provides a visual summary of mutation patterns across patients, with rows representing genes from the input list, columns representing individual patients, and cells color-coded by mutation impact where red indicates high impact, blue indicates moderate impact, and additional colors are used as applicable. Side panels provide complementary information: the left panel displays the percentage of mutated samples for each gene, while the top panel shows mutation burden or impact summary, often displayed as a combination impact score (e.g., 0–2). This visualization quickly shows which genes are frequently mutated and how mutation profiles vary across patients within the selected cohort.

B. Survival Control Panel

Located directly above the survival table and Kaplan–Meier plots, this panel includes dropdown menus for selecting the stratification method (Mutation vs. Wild type or By number of mutated genes) and time interval (All follow-up or 5-year survival). Changing these settings immediately updates the KM curves to reflect the selected analysis parameters.

C. Survival Statistics Table

This comprehensive table summarizes survival analysis results for each cancer type, gene or gene set, and survival endpoint combination. Key information includes the cancer type abbreviation, the gene(s) evaluated under the selected stratification, the outcome analyzed (OS for Overall Survival, PFI for Progression-Free Interval, DFI for Disease-Free Interval, DSS for Disease-Specific Survival), the stratification method used, which groups serve as the comparison factor versus reference, log-rank and Cox p-values for both all follow-up time and the 5-year interval, hazard ratios and their log2 transforms for both time periods, and sample counts for mutated and wild-type groups. HR values greater than 1 indicate the mutated group has worse prognosis, HR values less than 1 indicate the mutated group has better prognosis, and p-values less than 0.05 indicate significant survival differences between groups.

D. Kaplan–Meier (KM) Survival Plots

For every analysis, four Kaplan–Meier plots are generated—one for each survival endpoint: OS (Overall Survival), PFI (Progression-Free Interval), DFI (Disease-Free Interval), and DSS (Disease-Specific Survival). The KM curves automatically update based on the selected stratification method (Mutation vs. Wild type or Number of mutated genes) and time interval (All follow-up or 5-year survival). Each KM plot includes color-coded survival curves for the selected groups, survival probability over time in months, log-rank test p-value, and hover interaction to view timepoint-specific survival values. These four KM plots allow users to visually compare survival differences across mutation-defined groups for all major survival outcomes, with the multi-endpoint output being particularly useful for identifying consistent trends or endpoint-specific associations across different measures of patient prognosis.


3.3.2 Expression-based Survival Analysis

The Expression-Based Survival Analysis evaluates whether gene expression levels are associated with survival outcomes (OS, PFI, DSS, DFI) in a user-defined patient subpopulation, allowing users to stratify patients based on gene expression and examine survival differences across clinical groups.

Workflow

  1. Select gene(s)
    Enter one or multiple genes whose expression will be used for group stratification.
  2. Select dataset
    Choose the cancer dataset on which survival analysis will be performed.
  3. Define patient subpopulation
    Use clinical criteria (e.g., stage, grade, ER/PR/HER2 status, molecular subtype) to filter the cohort.
    These filters determine which patients will be included in the analysis.
  4. Submit
    The system loads the analysis interface.

Output

A. Stratification Control Panel

B. Survival Table

This table summarizes survival statistics for each gene, cancer type, and survival endpoint. Key metrics include hazard ratios (HR) and p-values for both all follow-up and 5-year intervals, expression cutoff thresholds, stratification methods, and sample counts for high and low-expression groups. HR > 1 indicates high expression is associated with worse prognosis, HR < 1 indicates better prognosis, and p < 0.05 indicates significant survival differences.

C. Kaplan-Meier (KM) Survival Plots

For each survival outcome, KM curves visualize survival differences between expression-defined groups, with one curve per group (High vs Low, All-high vs Others, etc.), log-rank p-value and HR displayed on the plot, and curves automatically updating based on cutoff method, grouping method, and time interval (5-year vs all follow-up). These plots show whether expression differences translate into clinically meaningful survival divergence and help users assess the prognostic value of the selected gene(s) in the filtered patient cohort.

D. Boxplots of Gene Expression

These boxplots summarize the expression distributions of the selected gene(s) within the filtered cohort, displaying TPM or log10(TPM) values with hover functionality to view sample-level values and summary statistics including median, Q1, Q3, fences, minimum, and maximum. This visualization helps confirm that expression-defined patient groups are meaningfully different in terms of gene expression levels before analyzing survival outcomes, ensuring that stratification produces biologically distinct groups for comparison.


3.3.3 miRNA-based Survival Analysis

Workflow

  1. Input a miRNA (e.g., hsa-miR-21-5p) to define the expression feature.
  2. Select an analysis framework: Cox Uni, Cox Multi, or Cure Model.
  3. After choosing the framework, configure the analysis using dropdown menus:
    • Cancer type – choose the cohort to analyze
    • Survival endpoint – select the outcome type
    • Survival time – follow-up time scale on the x-axis (months)
    • Stratification method – split patients into groups based on miRNA expression (e.g., High vs Low)
  4. View results as survival/hazard curves and (when applicable) risk estimates.

Output

A. Cox Uni Results (Univariate Cox Analysis)

This analysis displays two complementary visualizations when Cox Uni is selected and dropdown menus are configured. The Kaplan–Meier (KM) survival plot shows survival probability (y-axis) over time in months (x-axis) for miRNA-defined groups (e.g., High vs Low), allowing users to compare curve separation between groups to assess outcome differences and use the reported log-rank p-value to evaluate whether the group difference is statistically supported. The cumulative hazard plot shows cumulative hazard (accumulated risk of the event) over months for the same stratified groups, where steeper curves indicate risk accumulating more quickly and can be used alongside the KM plot to view group differences in terms of risk accumulation rather than survival probability. Clear separation between High versus Low groups suggests the miRNA is associated with prognosis without clinical adjustment, indicating a univariate association between miRNA expression and patient outcomes.

B. Cox Multi (Multivariate Cox Analysis with Clinical Adjustment)

This analysis displays two complementary visualizations when Cox Multi is selected and dropdown menus are configured. The adjusted survival curve shows model-predicted survival probability (y-axis) over months (x-axis) for miRNA-defined groups after adjusting for available clinical covariates in that cohort, allowing users to assess whether group separation persists after clinical adjustment and provides evidence that the miRNA offers prognostic information beyond standard clinical variables within the available covariates. The forest plot displays hazard ratios (HRs) with 95% confidence intervals for the miRNA group term (e.g., High vs Low) and each included clinical covariate that was automatically selected based on cohort availability, where HR > 1 indicates higher risk (worse outcome), HR < 1 indicates lower risk (better outcome), a dashed vertical line at HR = 1 indicates no effect, and wider confidence intervals indicate greater uncertainty. A significant miRNA-group HR after adjustment suggests the miRNA is an independent prognostic factor given the included covariates, providing evidence that miRNA expression contributes prognostic value beyond traditional clinical variables.

C. Cure Model

This analysis displays a cure-model survival curve when Cure Model is selected and dropdown menus are configured, showing cure-model–estimated survival probability (y-axis) over months (x-axis) for miRNA-defined groups. Users can compare curves to evaluate group-specific outcome differences under a model designed to capture long-term survival patterns and look for late-time plateaus that can reflect sustained survival behavior. Group separation indicates prognostic differences in a framework designed for long-term survival dynamics, which can be particularly informative when standard proportional hazards assumptions may not fully reflect the data and when a subset of patients may experience extended disease-free survival.


3.4 Multi-omics Driver Analysis

Overview

The Multi-omics Driver Analysis identifies driver events that differ between two clinically defined patient groups, integrating:
  • Gene expression
  • Mutation
  • Copy number variation (CNV)
  • Methylation

It highlights which genes and pathways are most likely driving group differences (e.g., responders vs non-responders, early vs late stage) and how consistently they are supported across omics types and tools.

Use this analysis when you want to understand which genomic and epigenomic alterations underlie clinical subgroup differences.

Workflow

  1. Select dataset
    Choose the cancer dataset to analyze.
  2. Define Group 1 and Group 2
    Use clinical criteria (e.g., stage, grade, response status, receptor subtype) to define two patient groups.
    • Each group should contain ≥ 20 samples for robust statistical analysis.
  3. Choose gene set option (Gene Dataset)
    • All – analyze all eligible genes in the dataset
    • CGC – restrict to genes listed in Cancer Gene Census (CGC)
    • NCG – restrict to genes listed in Network of Cancer Genes (NCG 6.0)
  4. Run analysis
    The results page will display driver events and multiple integrated visualizations.

Output

A. Multi-layer Driver-Function Relationship Diagram & Driver Summary Table
Multi-layer Driver-Function Relationship Diagram

This network-like diagram connects the selected cancer dataset, omics layers (mRNA, Mutation, CNV, Methylation), driver genes, and functional/pathway terms such as GO terms to illustrate the relationships between molecular alterations and biological functions. The structure flows hierarchically: the cancer node connects to each omics node (mRNA, Mutation, CNV, Methylation), each omics node connects to driver genes identified in that layer, and driver genes connect to GO term or function nodes representing enriched pathways or processes. Users can trace paths from clinical groups through omics alterations to driver genes and finally to biological functions, identifying which omics layers contribute most to observed group differences, which driver genes are shared across omics layers, and which biological functions and pathways are most impacted by these alterations.

Driver Summary Table

This table lists all identified drivers with their cancer type/dataset, omics layer (mRNA, Mutation, CNV, Methylation), driver gene symbol, Cancer Gene Census (CGC) status, Network of Cancer Genes (NCG) status, integration method that detected the driver, number of tools supporting this driver event (nTools), and associated Gene Ontology terms indicating pathways or functions. Higher nTools values indicate stronger cross-tool evidence for a driver event, CGC/NCG = Yes provides additional external support as a known cancer gene, and GO_term entries reveal potential biological roles and affected pathways. Together, the diagram and table summarize how multi-omics drivers are identified and what functional roles they may play in cancer biology.


B. Distribution of Drivers Across Omics Types and Tools

This section helps evaluate how strongly each driver is supported across omics layers and computational methods through two complementary visualizations.

Omics-by-Gene Tool Support Heatmap (Left)

This heatmap displays genes as rows and omics types (mRNA, Mutation, CNV, Methylation) as columns, with each cell value representing the number of tools that identified that gene as a driver in that specific omic layer. Users can hover on cells to see the exact tool count for each gene-omic combination, enabling identification of robust multi-omics drivers including genes with high support across multiple omics layers and genes supported by many tools in at least one omic category. This visualization helps prioritize genes based on the breadth and depth of computational evidence supporting their driver status.

Tools per Gene by Omics Bar Plot (Right)

This bar plot displays each driver gene with bar height representing the total number of tools that support that gene as a driver, with color coding indicating the contributions from different omics layers (mRNA, Mutation, CNV, Methylation). Users can compare tool support between genes to identify the most robustly detected drivers and use the legend to toggle specific omics types on or off, allowing focused examination of particular molecular layers. Genes with high tool support across several omics types are high-confidence multi-omics drivers, as convergent evidence from multiple computational methods and molecular mechanisms strengthens the reliability of their identification as functionally important cancer genes.


C. Coverage and Consistency of Multi-omics Identification Tools

This section focuses on tools rather than genes, evaluating how comprehensively tools cover omics layers and how consistent their driver identifications are across methods.

Omics-by-Tool Coverage Heatmap (Left)

This heatmap displays tools as rows and omics types as columns, with each cell showing the proportion of drivers in each omic layer detected by each specific tool. Users can examine which tools have broad coverage across multiple omics types versus those with more selective detection patterns focused on particular molecular layers, and identify tools that contribute most substantially to driver identification within a specific omics type. This visualization reveals the complementary nature of different computational approaches and helps users understand which tools are most effective for detecting drivers in each molecular context.

Tool Overlap Distribution Plot (Right)

This bar chart displays the number of tools (x-axis) versus the number of genes detected by that many tools (y-axis), revealing the degree of consensus among computational methods in driver identification. Genes detected by multiple tools are generally more reliable as convergent evidence from independent methods strengthens confidence in their driver status, while a right-shifted distribution with more genes supported by many tools suggests strong cross-tool consistency in the analytical pipeline. This plot helps users assess the overall reproducibility of driver detection and identify which genes have the most robust computational support across the integrated multi-omics framework.



3.5 Prognostic Signature Identification

3.5.1 Overview

The Prognostic Signature Identification analysis constructs a survival-predictive gene signature from a user-provided gene list. Using LASSO (Least Absolute Shrinkage and Selection Operator) and Random Forest models, the analysis identifies survival-associated genes, builds a multigene risk-score model, and evaluates its predictive performance through survival statistics, ROC curves, risk stratification plots, and feature-selection diagnostics.

Use this analysis if you already have a candidate gene list and want to determine:
  • Which genes are most predictive of survival
  • How these genes can be combined into a prognostic signature
  • How well the signature stratifies patients into risk groups
  • What biological functions are enriched among signature genes

3.5.2 Workflow

1. Input a Gene List

Users may provide a list of candidate genes using either method:
  • Type or paste genes directly into the text box
  • Upload a .txt file containing one gene symbol per line

These genes will be used to identify survival-related markers and construct the prognostic signature.

2. Select Dataset Settings

After submitting the gene list, users configure all analysis settings on the results page.

Select a Tissue
Choose a broad tissue category (e.g., Breast, Lung, Colon).
This filters the available cancer datasets.
Select Cancer Type

From the filtered list, select a TCGA (or other) cancer dataset for training the prognostic signature model.

3. Select Data Type(s) for Model Construction

Users may choose one or multiple omics types used for signature construction:

  • RNA expression
  • Copy Number Variation (CNV)
  • Mutation
  • Methylation

Selected data types define which molecular features contribute to the LASSO and Random Forest survival models.

4. Select a Survival Endpoint

Choose the patient outcome to be modeled:
  • Overall Survival (OS)
  • Progression-Free Interval (PFI)
  • Disease-Free Interval (DFI)
  • Disease-Specific Survival (DSS)

The endpoint determines how prognostic performance is evaluated.

5. Define Patient Subpopulation (Clinical Criteria Filter)

Filter the dataset to analyze a specific patient subpopulation by applying one or more clinical criteria.
Each criterion includes multiple groups with sample counts.
Users may combine multiple criteria to define a precise analysis cohort.

3.5.3 Results Overview

  1. Statistical Summary Table
  2. Kaplan-Meier Plot
  3. Time-Dependent ROC Curves
  4. Prognostic Risk Score Model
  5. Risk Heatmap
  6. Lambda Screening Plot
  7. Functional Annotation Barplots
  8. Shrinkage Gene List

3.6 Clinical Relevance Analysis

3.6.1 Overview

In the Clinical Relevance Analysis, over one hundred clinical factors are available for selection to construct a comprehensive prognostic model. If you are interested in comparing specific candidate gene(s) with well-known clinical prognostic biomarkers, this analysis will construct a multivariate model with customized clinical factors in the CoxPH framework. The generated report includes corresponding statistical results, Kaplan–Meier plots, and point-estimated values of all factors displayed in a forest plot.

3.6.2 Workflow

1. Input a Gene List or Signature list

Users choose the input type—either gene name or signature—and provide their list of candidate genes by typing or pasting gene symbols directly into the text box. These genes will be used to identify survival-related markers and construct the prognostic signature for the selected cancer cohort.

2. Select Dataset Settings

After submitting the gene list, users configure all analysis settings on the results page.

Select a Tissue

Choose a broad tissue category (e.g., Breast, Lung, Colon).
This filters the available cancer datasets.

Select Cancer Type

From the filtered list, select a TCGA (or other) cancer dataset for training the prognostic signature model.

3. Select Confounding Factors

This section determines how clinical variables are handled in the survival model.
Users:
  1. Select clinical factors (e.g., Age, Stage, Gender, Grade) to adjust for.
  2. Optionally enable: “Confounding factors selected by LASSO”
Behavior:
  • Unchecked:
    All selected confounding factors are forced into the Cox model—always included.
  • Checked:
    Selected confounders are also subjected to LASSO.
    LASSO selects only the clinical factors that meaningfully contribute to prognosis, shrinking others to zero.
This allows users to choose between:
  • Full adjustment (all selected confounders included)
  • Sparse adjustment (LASSO optimizes both genes and clinical variables)

4. Select Data Type(s) for Model Construction

Users may choose one or multiple omics types used for signature construction:
  • RNA expression
  • Copy Number Variation (CNV)
  • Mutation
  • Methylation

Selected data types define which molecular features contribute to the LASSO and Random Forest survival models.

5. Select a Survival Endpoint

Choose the patient outcome to be modeled:
  • Overall Survival (OS)
  • Progression-Free Interval (PFI)
  • Disease-Free Interval (DFI)
  • Disease-Specific Survival (DSS)

The endpoint determines how prognostic performance is evaluated.

6. Define Patient Subpopulation (Clinical Criteria Filter)

Filter the dataset to analyze a specific patient subpopulation by applying one or more clinical criteria.
Each criterion includes multiple groups with sample counts.
Users may combine multiple criteria to define a precise analysis cohort.

3.6.3 Results Overview

The analysis generates comprehensive results for each selected data type, including a summary table with statistical metrics, Kaplan–Meier plots showing survival curves for risk-stratified groups, and forest plots displaying hazard ratios with confidence intervals for all genes and clinical factors included in the final multivariate model.

Download

DriverDBv5 provides pan-cancer driver gene sets across five omics categories for download. RNA driver genes are identified based on differential expression criteria. Mutation driver genes are defined by at least three mutation detection tools across various cancers. CNV driver genes are identified based on significant copy-number gain or loss by iGC. Methylation driver genes are identified by MethylMix. Multi-omics driver genes are identified by multi-omics integration methods across multiple molecular data types. All datasets are available in TXT or CSV format.



Copyright© 2010-2025. All Rights Reserved. ©版權所有. Ver. 1.00.003未經允許請勿任意轉載、複製或做商業用途