Shared-Gene Disease Pairs?
Disease pairs ranked by curated gene overlap — a data-driven way to spot diseases that aren't normally considered related but share a large number of underlying genes. Looking for groups of more than two? See Disease Clusters.
What do these columns mean?
- Shared genes
- Curated genes linked to both diseases (disease_gdp). The line below it shows how many of those are "corroborated" -- backed by 2+ independent database sources combined across both diseases, not resting on a single source's say-so.
- Similarity score
- Jaccard-based: shared genes ÷ the union of both diseases' entire gene sets, with a small boost from users who bookmarked both diseases. Treats both diseases symmetrically.
- Overlap coefficient
- Shared genes ÷ the smaller disease's own total gene count. Complements the similarity score above for asymmetric pairs -- e.g. a rare disease almost entirely "contained" in a common disease's much larger gene set scores low on Jaccard but high here.
- P-value / FDR q-value
- Is this gene overlap more than chance? An upper-tail hypergeometric test, Benjamini-Hochberg corrected across every tested pair (the q-value is the one that accounts for testing thousands of pairs at once -- prefer it over the raw p-value).
- Shared cluster
- Links to a multi-disease cluster (see Disease Clusters) if both diseases of this pair were independently grouped together by that separate analysis.
0 selected
·
Loading clusters...
Showing 13 of 22038 pairs, sorted by significance (ascending). Click a column header to sort.