No terms match your search.

Enrichment Analysis

Pathway Enrichment
Genes Matched
The number of genes from your submitted list that are also known members of that pathway, the numerator used to calculate the Overlap Ratio for that row.
Overlap Ratio
The fraction of a disease or pathway's total known genes that are also in your submitted list (matched genes ÷ total genes). A higher ratio means more of that disease or pathway's genes are represented in your input.
P-value (Hypergeometric Probability)
How likely it is to see at least this much overlap between your gene list and the pathway's genes purely by chance (a one-sided hypergeometric test). Lower means more significant. Only your input genes that appear in GeDiPNet's pathway data count towards the test. Because many pathways are tested at once, use the Adjusted P-value (q-value) to judge significance. By default GeDiPNet shows results with q below 0.05.
Adjusted P-value (q-value, BH FDR)
The p-value corrected for testing many pathways at once, using the Benjamini-Hochberg false discovery rate (FDR). If you keep every result with q below 0.05, about 5% of them are expected to be false positives. GeDiPNet filters and colours results by q < 0.05 by default; switch "Significance measure" to Raw p-value to filter on p instead.
Fold Enrichment
How many times more often a pathway's genes appear in your list than expected by chance: (genes matched ÷ your input genes) ÷ (pathway size ÷ all background genes). 1 means no enrichment; 5 means five times more than expected. Read it alongside the q-value: a large fold from only one or two genes can still be weak evidence.
Legacy P-value (deprecated)
The p-value this page reported under its previous method (the probability of exactly the observed number of matches, rather than that many or more), reproduced only so earlier or published results can be traced. It is deprecated: don't use it to judge significance; use the P-value and Adjusted P-value (q-value) columns instead. Shown in Version 1 only, whose data is frozen so the old values can be recreated exactly.
Disease Enrichment (Genes)
Genes Matched
The number of genes from your submitted list that are also known to be associated with that disease, the numerator used to calculate the Overlap Ratio for that row.
Overlap Ratio
The fraction of a disease's total known genes that are also in your submitted list (matched genes / total genes). A higher ratio means more of that disease's genes are represented in your input.
P-value (Hypergeometric Probability)
How likely it is to see at least this much overlap between your gene list and a disease's known genes purely by chance (a one-sided hypergeometric test). Lower means more significant. Only your input genes that appear in GeDiPNet's disease-gene data count towards the test, and each matched gene is counted once. Because many diseases are tested at once, use the Adjusted P-value (q-value) to judge significance. By default GeDiPNet shows results with q below 0.05.
Adjusted P-value (q-value, BH FDR)
The p-value corrected for testing many diseases at once, using the Benjamini-Hochberg false discovery rate (FDR). If you keep every result with q below 0.05, about 5% of them are expected to be false positives. GeDiPNet filters and colours results by q < 0.05 by default; switch "Significance measure" to Raw p-value to filter on p instead.
Fold Enrichment
How many times more often a disease's genes appear in your list than expected by chance: (genes matched ÷ your input genes) ÷ (disease gene-set size ÷ all background genes). 1 means no enrichment; 5 means five times more than expected. Read it alongside the q-value: a large fold from only one or two genes can still be weak evidence.
Legacy P-value (deprecated)
The p-value this page reported under its previous method (the probability of exactly the observed number of matches, with the old way of counting matched genes, rather than that many or more), reproduced only so earlier or published results can be traced. It is deprecated: don't use it to judge significance; use the P-value and Adjusted P-value (q-value) columns instead. Shown in Version 1 only, whose data is frozen so the old values can be recreated exactly.
Domain Enrichment
Genes Matched
The number of genes from your submitted list that are also known to carry that protein domain (Pfam), the numerator used to calculate the Overlap Ratio for that row.
Overlap Ratio
The fraction of a domain's total known genes that are also in your submitted list (matched genes / total genes). A higher ratio means more of that domain's genes are represented in your input.
P-value (Hypergeometric Probability)
How likely it is to see at least this much overlap between your gene list and a protein domain's known genes purely by chance (a one-sided hypergeometric test). Lower means more significant. Only your input genes that carry a described Pfam domain count towards the test. Because many domains are tested at once, use the Adjusted P-value (q-value) to judge significance. By default GeDiPNet shows results with q below 0.05.
Adjusted P-value (q-value, BH FDR)
The p-value corrected for testing many protein domains at once, using the Benjamini-Hochberg false discovery rate (FDR). If you keep every result with q below 0.05, about 5% of them are expected to be false positives. GeDiPNet filters and colours results by q < 0.05 by default; switch "Significance measure" to Raw p-value to filter on p instead.
Fold Enrichment
How many times more often a domain's genes appear in your list than expected by chance: (genes matched ÷ your input genes) ÷ (genes carrying the domain ÷ all background genes). 1 means no enrichment; 5 means five times more than expected. Read it alongside the q-value: a large fold from only one or two genes can still be weak evidence.
Legacy P-value (deprecated)
The p-value this page reported under its previous method (the probability of exactly the observed number of matches, rather than that many or more), reproduced only so earlier or published results can be traced. It is deprecated: don't use it to judge significance; use the P-value and Adjusted P-value (q-value) columns instead. Shown in Version 1 only, whose data is frozen so the old values can be recreated exactly.
Disease Enrichment (Gene Ontologies)
P-value (Hypergeometric Probability)
How likely it is to see at least this much overlap between your gene list's GO Biological Process terms and a disease's GO terms purely by chance (a one-sided hypergeometric test over human GO annotations). Because GO terms are nested and correlated rather than independent, treat this as a similarity score for ranking diseases rather than a strict significance level. By default GeDiPNet shows results with an Adjusted P-value (q-value) below 0.05.
Overlap Ratio
The fraction of a disease's total associated GO Biological Process terms that are also found among your submitted gene list's GO terms (matched GO terms / total GO terms for that disease). This compares shared biological function, not shared genes directly.
GO Terms Matched
The number of GO Biological Process terms shared between your submitted gene list and that disease's known genes. This analysis matches on biological function annotations rather than on the genes themselves.
GO Term (Biological Process)
A standardized Gene Ontology label describing a biological process a gene product takes part in. It is one of GO's three branches (alongside Molecular Function and Cellular Component). This analysis groups genes and diseases by shared Biological Process terms rather than by shared genes.
Disease Name (Merged) vs (Original)
Disease Name (Merged) is GeDiPNet's standardized name across spelling/wording variants (disease_merge); Disease Name (Original) is the exact wording as it appeared in the source record before merging. Both are shown since one merged name can group several original disease-term variants.
GO Term (Gene Ontology)
A standardized label from the Gene Ontology describing a gene product's function, the biological process it takes part in, or where in the cell it acts. GeDiPNet uses these to compare genes or diseases by shared function rather than just shared gene identity.
Adjusted P-value (q-value, BH FDR)
The p-value corrected for testing many diseases at once, using the Benjamini-Hochberg false discovery rate (FDR). GeDiPNet filters and colours results by q < 0.05 by default; switch "Significance measure" to Raw p-value to filter on p instead. On this page the underlying p-value is a similarity score over GO terms, so use q to rank diseases rather than as a strict significance level.
Fold Enrichment
How many times more often a disease's GO terms appear among your genes' GO terms than expected by chance: (GO terms matched ÷ your genes' GO terms) ÷ (disease's GO terms ÷ all GO Biological Process terms). 1 means no enrichment.
Tissue-Specific Pathway Enrichment
P-value (Hypergeometric Probability)
How likely it is to see at least this much overlap between your tissue-expressed genes and a pathway's genes purely by chance (a one-sided hypergeometric test), before the Benjamini-Hochberg multiple-testing correction. The background is the KEGG genes expressed in the selected tissue at or above the chosen nTPM cutoff. See Adjusted P-value (q-value) for the corrected figure, which GeDiPNet filters on by default.
Overlap Ratio
The fraction of a pathway's total known genes that are also in your (tissue-filtered) submitted list (matched genes / total genes).
Genes Matched
The number of genes from your tissue-filtered list that are also known members of that pathway, the numerator used to calculate the Overlap Ratio for that row.
Overlapping Genes
The actual gene symbols from your tissue-filtered list that are also known members of that pathway, listed out. Genes Matched is just the count of this same list.
Adjusted P-value (q-value, BH FDR)
The p-value corrected for testing many pathways at once, using the Benjamini-Hochberg false discovery rate (FDR). If you keep every result with q below 0.05, about 5% of them are expected to be false positives. GeDiPNet filters and colours results by q < 0.05 by default; switch "Significance measure" to Raw p-value to filter on p instead.
nTPM (normalized Transcripts Per Million)
Measures how strongly a gene is expressed in the tissue you selected. Only your genes expressed at or above the nTPM cutoff are tested, and the same cutoff defines the background: just the KEGG genes expressed at or above it in that tissue are counted, so your list is compared like with like. Raising the cutoff gives a stricter, more tissue-specific analysis.
Fold Enrichment
How many times more often a pathway's genes appear in your tissue-expressed genes than expected by chance: (genes matched ÷ your input genes in the background) ÷ (pathway size ÷ all background genes), all counted at the selected nTPM cutoff. 1 means no enrichment; 5 means five times more than expected.
Legacy P-value (deprecated)
The p-value this page reported under its previous method (the probability of exactly the observed number of matches, measured against every gene listed for the tissue), reproduced only so earlier or published results can be traced. It is deprecated: don't use it to judge significance; use the p-value and Adjusted P-value (q-value) columns instead. Shown in Version 1 only, whose data is frozen so the old values can be recreated exactly.
Gene Network Analysis
Legend (Hub / Bottleneck / Regular Genes)
Color-codes every gene in the network diagram by its structural role: Hub genes (highly connected), Bottleneck genes (key connectors bridging otherwise separate parts of the network), and Regular genes (everything else). Use the checkboxes above the diagram to show or hide each type.
Hub Gene
The most highly connected genes in the network: strong candidates for polypharmacological drug targets, since affecting them influences many other genes at once.
Bottleneck Gene
Connects otherwise separate parts (modules) of the network. They don't need the most connections, but removing them would disconnect large parts of the network from each other.
Regular Gene
A gene that's neither a hub nor a bottleneck: the majority of genes in the network, not specially flagged as highly connected or as a bridge between modules.
Degree
How strongly a gene is connected to others in this network, accounting for the strength of each connection rather than just counting them. It is the same underlying measure as Weighted Degree elsewhere on GeDiPNet, shown here simply as Degree.
Closeness Centrality
How quickly a gene can reach every other gene in the network, based on shortest paths between them. A higher value means it sits closer, on average, to everything else in the network.
Betweenness Centrality
How often a gene lies on the shortest path between two other genes. A high value means it acts as a bridge controlling the flow of interactions between different parts of the network.
View Tissue Specific Interaction
Opens a popup showing that gene's known protein-protein interactions, letting you check which tissues its interaction partners are also expressed in.

Polypharmacological Analysis

Network Analysis
Weighted Degree
How strongly a gene is connected to others in this network. It sums the strength of all of a gene's connections, not just how many it has, so a higher value means heavier overall interaction with the rest of the network.
Closeness Centrality
How quickly a gene can reach every other gene in the network, based on shortest paths between them. A higher value means it sits closer, on average, to everything else in the network.
Betweenness Centrality
How often a gene lies on the shortest path between two other genes. A high value means it acts as a bridge controlling the flow of interactions between different parts of the network.
Hub Gene
The most highly connected genes in the network: strong candidates for polypharmacological drug targets, since affecting them influences many other genes at once.
Bottleneck Gene
Connects otherwise separate parts (modules) of the network. They don't need the most connections, but removing them would disconnect large parts of the network from each other.
Polypharmacology
The idea that a single drug (or gene target) can act on multiple targets or influence multiple diseases at once, rather than one drug/one target. GeDiPNet's polypharmacological analysis looks for genes connected across multiple diseases or pathways as candidates for this kind of multi-target effect.
Drug Target (Known / Predicted)
A "known" drug target is a gene with an existing drug already recorded as acting on it (from DGIdb/CTD); a "predicted" target is a gene GeDiPNet flags as similar to a known target (by sequence, function, or network position) but without a confirmed drug on file yet: a hypothesis for further validation, not a confirmed result.
Type
Classifies each gene as Hub (highly connected), Bottleneck (a key connector bridging otherwise separate parts of the network), or left unlabeled if it meets neither threshold.
Enriched Pathways
KEGG Pathway
A biological pathway from the KEGG database that came up significantly enriched among the genes in your query.
Associated Genes
Which of your query's genes belong to that KEGG pathway.
P-value
How likely it is to see at least this many of your query genes in the same KEGG pathway purely by chance (a one-sided hypergeometric test, counting only your query genes found in KEGG). Because many pathways are tested at once, the tab lists pathways whose Adjusted P-value (q-value) is below 0.05.
Adjusted P-value (q-value, BH FDR)
The p-value corrected for testing many pathways at once, using the Benjamini-Hochberg false discovery rate (FDR). If you keep every result with q below 0.05, about 5% of them are expected to be false positives. Only pathways with q < 0.05 are listed, and they are the ones carried into the network and summary tabs.
Fold Enrichment
How many times more often a pathway's genes appear among your query genes than expected by chance: (genes matched ÷ your query genes found in KEGG) ÷ (pathway size ÷ all KEGG genes). 1 means no enrichment. Read it alongside the q-value: a large fold from only one or two genes can still be weak evidence.
Legacy P-value (deprecated)
The p-value this tab reported under its previous method (the probability of exactly the observed number of matches, rather than that many or more), reproduced only so earlier or published results can be traced. It is deprecated: don't use it to judge significance; use the P-value and Adjusted P-value (q-value) columns instead. Shown in Version 1 only, whose data is frozen so the old values can be recreated exactly.
Summary
Polypharmacological Target (Gene)
A gene from your critical (Hub/Bottleneck) list on the Network Analysis tab, shown here with its pathways and any known drug already targeting it. Together these form the shortlist most relevant to a multi-target drug strategy.
Associated Pathways
The KEGG pathways this gene belongs to, carried over from the Enriched Pathways tab.
Known Drug
An existing drug on record (from DGIdb/CTD) already known to act on this gene.
Interaction Type
How a known drug interacts with this gene/target (e.g. inhibitor, agonist, antagonist), as recorded in the underlying drug-gene interaction data (DGIdb/CTD).
Node Type
Whether this gene was flagged as a Hub or Bottleneck on the Network Analysis tab. Only critical (hub/bottleneck) genes appear in this Summary table.
Drug Targets
Drug Name
A drug already on record (via DGIdb/CTD) as interacting with at least one of the critical (Hub/Bottleneck) genes in your query.
Known Targets / Predicted Targets
A "known" target is a gene with an existing drug already recorded as acting on it (from DGIdb/CTD); a "predicted" target is a gene GeDiPNet flags as similar to a known target (by sequence, function, or network position) but without a confirmed drug on file yet: a hypothesis for further validation, not a confirmed result.
Deep-Dive
Opens the Drug Deep-Dive view for this specific drug, pulling together 5 things in one place so you can assess it in detail: Clinical Trial Status, Tissue Expression (GTEx), Transcriptomic Cross-reference, Side Effects, and ADMET Screening.
Drug Deep-Dive
SMILES
Simplified Molecular Input Line Entry System: a compact text notation for a molecule's chemical structure, using letters and symbols to represent atoms and bonds. For example, Metformin's SMILES is CN(C)C(=N)NC(=N)N, and this single string is what gets fed into ADMET prediction to compute a drug's properties directly from its chemical structure.
ADMET
Absorption, Distribution, Metabolism, Excretion, and Toxicity: five properties that determine whether a drug candidate behaves safely and effectively in the body, beyond just hitting its intended target. For example, a drug might bind its target perfectly in a lab test but still fail as a medicine if it's poorly absorbed or toxic to the liver. ADMET screening flags issues like these early, before expensive clinical testing.
Blood-Brain Barrier (BBB) Penetration
A prediction of whether a drug molecule can cross from the bloodstream into the brain. For example, a drug meant to treat a brain condition needs high BBB penetration to reach its target, while a drug for a condition outside the brain (like diabetes) usually doesn't, and low BBB penetration can even be desirable there, to avoid unwanted central nervous system side effects.
TPM (Transcripts Per Million)
A standardized unit for measuring how much a gene is expressed (how actively it's being read and used) in a tissue, normalized so different samples and tissues can be fairly compared. For example, if a gene shows 90 TPM in Spleen but only 6 TPM in Brain, that gene is far more active in Spleen. GeDiPNet's Tissue Expression panel uses TPM values from GTEx to check whether a drug's predicted target gene is actually expressed in a tissue relevant to the disease.
mRNA Expression Z-score
A statistical score showing how much a gene's expression in one sample differs from the average, measured in standard deviations: 0 means average, +2 means notably higher than average, -2 means notably lower. For example, a Z-score of 3 for a gene in a tumor sample means that gene is expressed unusually highly in that tumor. GeDiPNet's Transcriptomic Cross-reference panel uses this to flag genes with strongly altered expression in cancer datasets from TCGA.
PRR (Proportional Reporting Ratio)
A statistical measure used in drug safety monitoring to flag side effects reported unusually often for a specific drug, compared to how often that side effect is reported for all other drugs combined. A PRR of 1 means no unusual association; higher values mean a stronger signal. For example, a PRR of 10 for "jaundice" means that side effect is reported about 10 times more often for that drug than expected by chance. GeDiPNet's Side Effects panel uses PRR values from the OFFSIDES database to surface these off-label signals.
QED (Quantitative Estimate of Drug-likeness)
A single score from 0 to 1 summarizing how similar a molecule's overall properties are to those of successful, approved oral drugs. The closer to 1, the more "drug-like" it looks structurally. For example, a QED of 0.8 suggests favorable size, solubility, and other properties typical of real medicines, while a very low QED (like 0.1) flags a molecule that may be hard to develop into an actual pill, even if it binds its target well.
ChEMBL ID
A unique identifier (e.g. CHEMBL1431) assigned by the ChEMBL database to a specific drug or bioactive molecule, used to reliably look up its exact chemical structure regardless of what name it goes by. For example, GeDiPNet uses a drug's ChEMBL ID, when available, to fetch its precise chemical structure for ADMET screening, since drug names alone can be ambiguous (brand, generic, or chemical names) while a ChEMBL ID always points to one exact molecule.
Clinical Trial Phase
The stage of testing a drug has reached in human trials, numbered 1 through 4. Phase 1 tests safety in a small healthy group; Phase 2 tests effectiveness and side effects in a larger patient group; Phase 3 confirms effectiveness in an even larger group, often required before approval; Phase 4 monitors a drug after it's already approved and in public use. For example, a drug listed as "Phase 3" in GeDiPNet's Clinical Trial Status panel is in late-stage testing, much closer to potential approval than a "Phase 1" compound.
Clinical Trial Status: Not yet recruiting
The study has been registered but has not started recruiting participants yet. For example, a drug shown as "Not yet recruiting" for a disease on GeDiPNet's Clinical Trial Status panel is a very early, unconfirmed repurposing lead, years away from any result.
Clinical Trial Status: Recruiting
The study is actively enrolling and currently looking for participants who meet its eligibility criteria. This is a live, in-progress trial: the strongest "this is genuinely being investigated right now" signal short of a completed result.
Clinical Trial Status: Enrolling by invitation
The study is selecting participants from a specific pre-decided population and directly inviting them, rather than being open to anyone who meets the eligibility criteria. For example, a trial studying a drug in people who already have a specific rare genetic marker would enroll this way instead of accepting open sign-ups.
Clinical Trial Status: Active, not recruiting
The study is still ongoing and participants already enrolled are receiving the intervention or being examined, but no new participants are being recruited. Enrollment is closed, but the trial hasn't reported a result yet.
Clinical Trial Status: Suspended
The study has stopped early, but might resume later. This is a caution flag: something (often a safety concern, funding issue, or administrative problem) paused the trial before it reached a normal conclusion.
Clinical Trial Status: Terminated
The study stopped early and will not resume; participants are no longer being examined or treated. Worth investigating why (often listed in the trial's own record on ClinicalTrials.gov) before treating a drug as a promising repurposing candidate, since termination can mean anything from a safety failure to a simple funding cut.
Clinical Trial Status: Completed
The study ended normally: every participant's last visit has already happened. This is the status most likely to have an actual published result attached, making it the most useful status to check for an existing answer on a drug's effectiveness or safety.
Clinical Trial Status: Withdrawn
The study stopped before enrolling even its first participant. Functionally, no data exists from this trial at all; it was registered but never actually ran.
Clinical Trial Status: Unknown
The trial's last known status was "recruiting," "not yet recruiting," or "active, not recruiting," but it has now passed its expected completion date without being re-verified in the past 2 years. In practice this usually means the trial record is stale and nobody has updated ClinicalTrials.gov on what actually happened to it.
ClinicalTrials.gov
The U.S. National Library of Medicine's public registry of clinical trials worldwide. GeDiPNet's Clinical Trial Status panel queries it live to check whether a predicted drug is already approved, in trial, or completely untested for a given disease.
GTEx (Genotype-Tissue Expression)
A large reference project measuring normal gene expression levels across dozens of human tissues from donated samples. GeDiPNet's Tissue Expression panel uses GTEx to check whether a predicted drug target gene is actually expressed in a tissue relevant to the disease being studied. A target with near-zero expression there is a weak candidate no matter how good its other evidence looks.
Expression Atlas
EMBL-EBI's public database of gene expression experiments across many diseases and conditions, both healthy and diseased. GeDiPNet's Transcriptomic Cross-reference panel searches it for studies relevant to the disease being investigated, linking out to the actual experiments.
cBioPortal
A public cancer genomics portal providing real patient-level data from large cancer studies, including TCGA (The Cancer Genome Atlas). GeDiPNet's Transcriptomic Cross-reference panel uses it specifically for cancer-related diseases, to show real mRNA expression Z-scores for a target gene across an actual patient cohort.
SIDER (Side Effect Resource)
A database of side effects extracted from the official printed labels/package inserts of marketed drugs. GeDiPNet's Side Effects panel uses SIDER for the "known, on-label" side effects of a candidate drug.
OFFSIDES (nSIDES)
A database of statistically significant drug side effect signals mined from real-world FDA adverse event reports, specifically the ones NOT listed on a drug's official label. GeDiPNet's Side Effects panel uses OFFSIDES for the "off-label, unexpected" side effect signals, scored by PRR.
ADMETlab 3.0
A free online tool that predicts a molecule's ADMET properties (absorption, distribution, metabolism, excretion, toxicity) directly from its chemical structure, using machine learning models. GeDiPNet's ADMET Screening panel submits a drug's SMILES structure to ADMETlab 3.0 to generate these predictions.
ChEMBL
EMBL-EBI's manually curated database of bioactive, drug-like molecules and their properties. GeDiPNet uses a drug's ChEMBL ID, when available, to reliably fetch its exact chemical structure (SMILES) for ADMET screening.
PubChem
The NIH's large, open database of chemical structures and properties. GeDiPNet uses PubChem as a fallback to resolve a drug's chemical structure by name whenever a ChEMBL ID isn't available.

Comorbidity Analysis

Comorbidity Analysis (Gene Result)
Shared Diseases Score
score = (diseases shared by a gene pair / the larger gene's total disease count) x 100. It measures how related two genes are based on the number of diseases they're both linked to. Higher score means more overlap, shown as a color-coded heatmap (light = low, dark = high).
Disease Uniqueness (weighting)
Weights shared diseases by how rare they are across GeDiPNet's genes: sharing a disease linked to very few other genes counts for more than sharing a broadly-linked one.
Heatmap
A grid coloring every compared pair by its score (light for low scores, dark for high ones), so the strongest relationships in a large comparison are visible at a glance without reading every number individually.
Shared Phenotype (Comorbidity)
A tab on the Gene Comorbidity Analysis results page that scores how related two genes are based on the clinical signs and symptoms (HPO, or Human Phenotype Ontology, terms) they are both linked to. The score is the number of shared HPO terms divided by the larger of the two genes' total term counts, shown as a percentage.
Shared Phenotype Score
score = (HPO terms shared by a gene pair divided by the larger gene's total HPO term count) x 100. This is the number behind the Shared Phenotype tab's heatmap color, same normalization convention as every other comorbidity score (higher = more overlap).
Shared Symptoms/Signs
The specific HPO (Human Phenotype Ontology) terms two genes are both linked to, listed in the Shared Phenotype tab's results table; each links out to its HPO browser page. Sourced from hpo_gene_phenotype (HPO's genes_to_phenotype.txt), with inheritance-pattern and age-of-onset annotations filtered out, since those describe the record rather than something a patient actually presents with.
Tissue-Specific (Comorbidity)
A tab on the Gene Comorbidity Analysis results page that scores how related two genes are based on overlap in the tissues where they are both actively expressed (nTPM >= 1). Reveals whether two genes' relatedness looks localized to specific tissues (e.g. both active in gut and brain) or systemic (active almost everywhere).
Tissue Specificity Score
score = (tissues shared by a gene pair, each gene's own full expression profile at nTPM >= 1, divided by the larger gene's total tissue count) x 100. This is the number behind the Tissue Specificity tab's heatmap color.
Shared Tissues
The specific tissues (Human Protein Atlas expression data) two genes are both actively expressed in (nTPM >= 1), listed in the Tissue Specificity tab's results table alongside each gene's own nTPM value there.
nTPM
Normalized Transcripts Per Million: the Human Protein Atlas's measure of a gene's expression level in a given tissue, normalized so values are comparable across tissues and samples. GeDiPNet treats nTPM >= 1 as "actively expressed" throughout its tissue-based features (Tissue-Specific Pathway Enrichment, Tissue Specificity comorbidity).
Comorbidity Analysis (Disease Result)
Comorbidity Score
score = (genes shared by a disease/gene pair ÷ the larger of the two items' total gene counts) × 100. A higher score means the pair's genes overlap more heavily, shown as a color-coded heatmap (light = low risk, dark = high risk).
Gene Uniqueness (weighting)
Weights shared genes by how rare they are across GeDiPNet's diseases: sharing an unusual gene counts for more than sharing a very common one.
Shared Ontologies
Evaluates comorbidity risk between diseases by their shared Gene Ontology (GO) terms, rather than by shared genes directly. Two diseases can score highly here even with little gene overlap, if their genes converge on the same biological processes.
Heatmap
A grid coloring every compared pair by its comorbidity score (light for low scores, dark for high ones), so the strongest relationships in a large comparison are visible at a glance without reading every number individually.
Shared Phenotype (Comorbidity)
A tab on the Disease Comorbidity Analysis results page that scores how related two diseases are based on the clinical signs and symptoms (HPO terms) their associated genes are linked to. It goes from each disease to its genes to those genes' HPO phenotypes, so diseases are never matched to HPO by name directly.
Shared Phenotype Score
score = (HPO terms shared by a disease pair's gene sets divided by the larger disease's term count) x 100. Each disease's terms are its top 15 HPO phenotypes, ranked by how many of the disease's own genes carry that term, to keep a disease with 100+ genes from saturating toward nearly every HPO term that exists.
Shared Phenotypes
The specific HPO terms two diseases' gene sets are both linked to, listed in the Shared Phenotype tab's results table; each links out to its HPO browser page. A disease whose genes come mostly from rare/Mendelian-disease data (HPO's main source) gets a more reliable signature here than one whose genes are linked mainly through common-disease evidence (GWAS, DisGeNET, etc.); the tab flags this when it applies.
Tissue-Specific (Comorbidity)
A tab on the Disease Comorbidity Analysis results page that scores how related two diseases are based on overlap in the tissues where their associated genes are actively expressed (nTPM >= 1). It goes from each disease to its genes to those genes' tissue expression.
Tissue Specificity Score
score = (tissues shared by a disease pair's gene sets divided by the larger disease's tissue count) x 100. Each disease's tissues are its top 10 by mean expression across the disease's genes, to keep a disease with 100+ genes from saturating toward nearly every tissue in the body.
Shared Tissues
The specific tissues two diseases' gene sets are both active in, listed in the Tissue Specificity tab's results table alongside each disease's average nTPM there.
nTPM
Normalized Transcripts Per Million: the Human Protein Atlas's measure of a gene's expression level in a given tissue, normalized so values are comparable across tissues and samples. GeDiPNet treats nTPM >= 1 as "actively expressed" throughout its tissue-based features (Tissue-Specific Pathway Enrichment, Tissue Specificity comorbidity).
Disease Clusters
Disease Cluster
A group of diseases that share a large number of curated genes with each other, computed via label propagation over a shared-gene similarity graph, so diseases land in the same cluster by transitive gene overlap, not just direct pairwise similarity.
Cluster Gene Count
Total distinct genes across every disease in a cluster: the "n" used as the background size in that cluster's pathway/GO enrichment significance test.
Overlap Genes (x / y)
x = genes shared between a cluster and a given pathway/GO term; y = that pathway/GO term's total gene count. A higher x relative to y (and to the cluster's own size) means a tighter biological match.
P-value / FDR q-value (Disease Clusters)
Is a pathway/GO term's overlap with the cluster more than chance? Upper-tail hypergeometric test, Benjamini-Hochberg corrected across every pathway/term tested for that cluster. Prefer the q-value, since it accounts for testing many at once.
Similarity Score (Pairs Within Cluster)
Jaccard-based gene overlap between two specific diseases inside the same cluster, using the same metric as the Shared-Gene Disease Pairs page, shown here per pair so you can see which pairs are driving the cluster.
Shared-Gene Disease Pairs
Shared Genes (Disease Pairs)
Curated genes linked to both diseases in a pair. A second figure shows how many of those are "corroborated": backed by 2+ independent database sources combined across both diseases, not resting on a single source's say-so.
Similarity Score (Disease Pairs)
Jaccard-based: shared genes ÷ the union of both diseases' entire gene sets, with a small boost from users who bookmarked both diseases. Treats both diseases symmetrically.
Overlap Coefficient
Shared genes ÷ the smaller disease's own total gene count. Complements the similarity score for asymmetric pairs. For example, a rare disease almost entirely "contained" in a common disease's much larger gene set scores low on Jaccard but high here.
P-value / FDR q-value (Disease Pairs)
Is a pair's gene overlap more than chance? Upper-tail hypergeometric test, Benjamini-Hochberg corrected across every pair tested. Prefer the q-value, since it accounts for testing thousands of pairs at once.
Shared Cluster
Links to a multi-disease cluster (see Disease Clusters) if both diseases of a pair were independently grouped together by that separate analysis.
Pathway
KEGG vs Reactome
KEGG and Reactome are the two independent pathway databases GeDiPNet curates from, and a pathway can appear from either (or both). They sometimes name overlapping biology differently, so it's normal to see both KEGG and Reactome entries covering related pathways. You can tell them apart by ID format: KEGG IDs look like "hsa#####", Reactome IDs contain "R-HSA-".
Genes in Pathway / Diseases Associated
Genes in Pathway counts the distinct genes shown in this pathway's tree diagram; Diseases Associated counts the distinct diseases linked to any of those genes. Both are computed directly from the diagram, not separate database lookups.
SNP/Variant
Clinical Significance
ClinVar's own classification of how strongly a variant is linked to disease. Common values are Pathogenic, Likely pathogenic, Uncertain significance, Likely benign, and Benign.
dbSNP ID
The unique identifier NCBI's dbSNP database assigns to a specific genetic variant (e.g. rs123456). Click it to open that variant's full NCBI record.
Variant Type
Describes the kind of DNA change, e.g. single nucleotide variant, deletion, insertion, duplication, or indel.
RCV Accession ID
ClinVar's record identifier for one specific variant-condition pairing. Click it to see ClinVar's full submission and evidence for that classification.
Phenotype/Disease
The condition ClinVar associated with that variant when it was submitted. Click it to open that disease's page on GeDiPNet.
Consequence
The functional effect of this variant on the gene (e.g. missense, nonsense, synonymous), as classified in dbSNP.
Visualize Variation
Opens a separate tool showing where this variant sits within the gene/protein structure, rather than an inline chart on this page.
SNP Type (Coding / Non-coding)
Whether this SNP falls within a gene's coding sequence or not. This is a different classification from Variant Type (which describes the kind of DNA change, e.g. substitution vs deletion), shown in GeDiPNet's SNP Analyzer tool.
QTL Type
The kind of quantitative trait locus association found for this SNP (e.g. an expression QTL/eQTL-style result), sourced from QTLbase, alongside the Tissue column showing which tissue/dataset that association was observed in.
Protein
PDB ID
Points to an experimentally solved 3D structure of the protein in the RCSB Protein Data Bank. Click it to view that structure. Not every protein has one; an empty PDB column just means no solved structure is on file yet.
UniProt ID (Entry)
That protein's unique identifier in the UniProt database. Click it to open UniProt's full curated record for the protein.
Sequence Length
The number of amino acids in the protein's full sequence, shown alongside the sequence itself in the table.
Protein Domain (Pfam)
A distinct structural or functional region within a protein, as classified by the Pfam database. Proteins sharing a domain often share a related function or evolutionary origin, even if the rest of the protein differs.
Position in Sequence / Type (Pfam)
Where on the protein's amino acid sequence a Pfam domain sits (start position to end position), plus its Type (e.g. domain, family, or repeat) as classified by Pfam.
View Interactions
Opens a popup showing this protein's known protein-protein interactions from the STRING database.
Gene Names (Protein search results)
A single UniProt protein entry can be linked to more than one gene symbol; this column lists all of them, each clickable to that gene's own page.
Gene
Chromosomal Location
The gene's cytogenetic map position (e.g. 17q21.31): chromosome number, arm (p/q), and band.
Entrez / NCBI Gene ID
The numeric identifier NCBI's Gene database assigns to a gene (distinct from its text symbol, e.g. BRCA1). It's a stable reference that doesn't change even if a gene's preferred symbol or name is later renamed.
Gene Synonyms / Aliases
Alternative names or symbols a gene has been known by (older nomenclature, common lab names, etc.), beyond its current official HGNC symbol. Searching any synonym will still find the gene under its current symbol.
Match Type (Direct / Via Disease Link)
Direct match means the gene's own record (name, symbol, synonym, or chromosomal location) contains your search term. Via disease link (indirect match) means the gene itself doesn't mention your term, but it's associated with a disease that does.
Gene Ontology (GO) Annotation
Each row pairs a GO branch (Biological Process, Molecular Function, or Cellular Component) with one specific GO term describing that gene's function, process, or location. It is not a dictionary-style definition, just the GO term's own name.
GO Evidence Code
A short code (e.g. IEA, TAS, IDA) from the Gene Ontology's standard vocabulary indicating how that GO annotation was determined. For example, IEA means inferred electronically (not manually reviewed), while IDA means inferred from a direct experimental assay.
Regulation (Transcription Factors)
How the listed transcription factor regulates this gene (e.g. activates or represses its expression), sourced from the TRRUST v2 database.
Experiments (miRNA)
The experimental method(s) (e.g. reporter assay, western blot) used to validate that this miRNA targets this gene, sourced from miRTarBase.
Other IDs (MIM / HGNC / Ensembl)
Cross-references to this gene's record in three other databases: OMIM (Mendelian Inheritance in Man), HGNC (the official gene nomenclature committee), and Ensembl; each links out to that database's own page for the gene.
Disease
Evidence Source / References
Links out to the original evidence source for a gene-disease association, e.g. MESH, HPO, OMIM, MedGen, MONDO, or Orphanet, depending on where GeDiPNet curated it from.
Disease Name vs Disease Term
GeDiPNet merges related disease spellings/variants under one standardized "Disease Name" (disease_merge); the "Disease Term" is the original, more specific wording as it appeared in the source record before merging. Searching either will surface the same underlying association.
PMID (PubMed ID)
The unique identifier PubMed assigns to a published article. GeDiPNet attaches PMIDs to gene-disease associations where a supporting publication is on record, linking directly to that article's PubMed page.
Evidence Score (Star Rating)
A 1-to-5 star rating GeDiPNet computes for how strongly a gene-disease association is supported: 1 star = found only via text mining; 2 stars = text mining plus an Unknown/Other-Associations source; 3 stars = reported in Unknown/Other Associations by 2 or more independent sources; 4 stars = a ClinVar Pathogenic/Likely Pathogenic variant (fewer than 5 such variants); 5 stars = a ClinVar Pathogenic/Likely Pathogenic variant with 5 or more such variants on record.
Causal vs Unknown/Other Associations
Causal means the association is backed by a Pathogenic or Likely Pathogenic variant in ClinVar. Unknown/Other means the association instead comes from ClinVar entries with uncertain or conflicting evidence, or from other databases (OMIM, Orphanet, GWAS Catalog, GenCC, ClinGen, HPO, DisGeNET, CTD) where the gene isn't established as a direct cause.
Relationship Type (Text Mining)
The type of relationship a text-mining pipeline extracted between a gene and disease from published literature (associate, stimulate, or inhibit). It is not a clinical or causal classification, just what the source text described.
ClinGen Report (References column)
When a gene-disease association's evidence comes from ClinGen, the References column links to ClinGen's own evidence report instead of a PubMed ID; not every entry in that column is a PMID.
Venn Analysis
Venn Diagram
A diagram of overlapping circles, one per set you compared (genes, pathways, GO terms, or protein domains). Where circles overlap, the items in that region are shared by every set whose circle covers it; a region covered by only one circle holds items unique to that set alone.