ClinVar is a freely accessible, public archive of reports of the relationships among human variations and phenotypes, with supporting evidence. ClinVar thus facilitates access to and communication about the relationships asserted between human variation and observed health status, and the history of that interpretation. ClinVar processes submissions reporting variants found in patient samples, assertions made regarding their clinical significance, information about the submitter, and other supporting data. The alleles described in submissions are mapped to reference sequences.
The level of confidence in the accuracy of variation calls and assertions of clinical significance depends in large part on the supporting evidence, so this information, when available, is collected and visible to users. Because the availability of supporting evidence may vary, particularly in regard to retrospective data aggregated from published literature, the archive accepts submissions from multiple groups, and aggregates related information, to reflect transparently both consensus and conflicting assertions of clinical significance. A review status is also assigned to any assertion, to support communication about the trustworthiness of any assertion. ClinVar archives submitted information, and adds identifiers and other data that may be available about a variant or condition from other public resources. However, ClinVar neither curates content nor modifies interpretations independent of an explicit submission.
ClinVar is one of the most powerful tools available to us. However, it is important to understand that any pathogenicity prediction tools, especially with public variant submission, can have discrepancies. ClinVar occasionally will have misclassified variants, and a large number (>70%) will be classified as Variants of Uncertain Significance (VUS).
https://www.ncbi.nlm.nih.gov/clinvar/
COSMIC – the Catalogue of Somatic Mutations in Cancer – is the world’s largest source of manually curated somatic mutation information relating to human cancers. Data is outlined in terms of structure, content and scope making it easier for you to evaluate what you will find in COSMIC, and how best to access it to fulfill your research needs.
The COSMIC database combines two main types of data:
High Precision Data, Manually Curated by Experts:
- Targeted gene-screening panels
- Over 25,000 peer reviewed papers
- Metadata (environmental factors and patient history)
- Focused on known and suspected cancer genes and mutations
- Objective frequency data as a result of mutation negative samples
- Full details of the curation process and data captured
Genome-wide Screen Data:
- Over 32,000 genomes, consisting of:
- Provides unbiased, genome-level profiling of diseases
- Objective frequency data, by interpreting non-mutant genes across each genome
- Can be used to discover novel driver genes
Together, this compilation of data provides extensive coverage of the cancer genomic landscape from a somatic perspective. New and potentially significant data are continually captured and made available through quarterly updates to COSMIC each year.
COSMIC is the go-to database when interested in cancer genetics. It can be used along with gnomAD & ClinVar to gain a better understanding of the genetic mechanisms that are driving disease.
https://cancer.sanger.ac.uk/cosmic
The Genome Aggregation Database (gnomAD), is a database of sequenced individuals in an attempt to understand genome and exome sequencing data from a variety of large-scale sequencing projects, and to make summary data available for the wider scientific community. In its first release, which contained exclusively exome data, it was known as the Exome Aggregation Consortium (ExAC). In the last few years large amounts of full genome data have been added.
The data set provided on this website spans 123,136 exomes and 15,496 genomes from unrelated individuals sequenced as part of various disease-specific and population genetic studies. Individuals known to be affected by severe pediatric disease have been removed from the data set, as well as their first-degree relatives, so this data set should serve as a useful reference set of allele frequencies for severe disease studies – however, note that some individuals with severe disease may still be included in the data set, albeit likely at a frequency equivalent to or lower than that seen in the general population.
The gnomAD data set contains individuals sequenced using multiple exome capture methods and sequencing chemistries, so coverage varies between individuals and across sites. This variation in coverage is incorporated into the variant frequency calculations for each variant.
gnomAD is yet another useful tool in the discovery and analysis of non-embryonic lethal missense and loss-of-function variants. It should be used alongside ClinVar and any other variant databases that are used.
http://gnomad.broadinstitute.org/
The UniProt Knowledgebase (UniProtKB) is the central hub for the collection of functional information on proteins, with accurate, consistent and rich annotation. In addition to capturing the core data mandatory for each UniProtKB entry (mainly, the amino acid sequence, protein name or description, taxonomic data and citation information), as much annotation information as possible is added. This includes widely accepted biological ontologies, classifications and cross-references, and clear indications of the quality of annotation in the form of evidence attribution of experimental and computational data.
The UniProt Knowledgebase consists of two sections: a section containing manually-annotated records with information extracted from literature and curator-evaluated computational analysis, and a section with computationally analyzed records that await full manual annotation. More than 95 % of the protein sequences provided by UniProtKB are derived from the translation of the coding sequences (CDS). In order to have minimal redundancy and to improve sequence reliability, all protein sequences encoded by a same gene are merged into a single UniProtKB/Swiss-Prot entry.
Oftentimes UniProt will be the first database accessed to gain an understanding of a gene. Another utility of UniProt is using the FASTA sequence to model the protein in YASARA.