Dataset Versions
Compare the live Version 2 dataset against the frozen Version 1 snapshot: what's in each, and when to use which.
GeDiPNet keeps a frozen Version 1 snapshot alongside the live Version 2 dataset, so the data an analysis runs against in Version 1 never changes, even as Version 2 keeps growing. The statistical methods applied to it are shared by both versions and may be corrected over time; any such change is noted below. Every browse and analysis page on the site lets you choose which one to query.
Updated on an ongoing basis with new curation, corrections, and analysis features.
What's New
- RAG chatbot for grounded, database-only Q&A
- Domain (Pfam) enrichment analysis
- Drug Deep-Dive: clinical trials, ADMET, side effects, tissue expression, transcriptomic cross-reference
- Comorbidity scoring by gene uniqueness and shared HPO phenotype
- Gene-based polypharmacological target prediction
Locked at its initial release. Its data never changes, so analyses run against it stay reproducible. Statistical methods may still be corrected (see the note below).
What Version 1 covers
- Disease- and gene-based comorbidity analysis
- KEGG/Reactome pathway, disease, and GO enrichment analysis
- Disease-based polypharmacological target prediction
- Venn comparison across genes, pathways, domains, and GO terms
Choosing a version: use Version 2 for the most current data. Use Version 1 when you need to work from exactly the data a previous analysis or publication used, since its data never changes. Results computed from it can still differ if a statistical method has since been corrected (see below).
Methods update (enrichment analysis, both versions): enrichment p-values now use the one-sided hypergeometric test (probability of k or more matches; previously the probability of exactly k), with Benjamini-Hochberg FDR-adjusted q-values and fold enrichment added and the default filter now q < 0.05. Input genes missing from an analysis's annotation data no longer count towards the test, disease enrichment (genes) now counts each matched gene once, tissue-specific enrichment uses KEGG genes expressed at the chosen nTPM cutoff as its background, and GO-based disease enrichment uses human Gene Ontology annotations only. The underlying data in both versions is unchanged. To trace a result reported under the previous method, run it against Version 1: there, the pathway, disease (genes), domain and tissue-specific enrichment pages include a “Legacy p-value (deprecated)” column that reproduces the old value; GO-based disease enrichment results from the previous method cannot be reproduced. The Enriched Pathways tab of Polypharmacological Target Prediction uses the same updated method, so it now lists pathways with q < 0.05, and the network and summary tabs, which are built from that list, can change accordingly. In Version 1 that tab has a “Show previous-method list” switch that displays the list exactly as it appeared under the previous method.