Comparative Genomics: Methods, Tools, Significance

Human beings share about 99% of DNA with other people, 98.5% with chimpanzees, 85% with mice, and a surprising 60% with bananas. But what is the portion that varies in these organisms? The study of these differences in the genome is known as comparative genomics.

Comparative Genomics
Comparative Genomics

What is Comparative Genomics?

Comparative genomics is a bioinformatics tool in which the genomes of different organisms, such as humans, bats, mice, insects, and others, are compared to understand the molecular and evolutionary differences among them.

  • The underlying genomic features of the organism help us understand and differentiate varying life forms from each other at a molecular level. 
  • It is a powerful tool to study the similarities and differences between genomes, to elucidate conserved and divergent genes among species with their unique phenotypical function. 
  • To date, over 1,000 prokaryotic genomes and 1,300 species have been entirely sequenced, with the number steadily growing each year. 
  • Among them, one of the most notable achievements is the sequencing of the human genome through an international collaborative effort, called the Human Genome Project (HGP). The project sequenced almost all of the human genome, providing a comprehensive blueprint of human biology. Similarly, other important model organisms such as chimpanzee, mouse, puffer fish, rat, fruit fly, roundworms, yeast, E. coli, and others have also been completely sequenced. 

What is a Genome, and What to Compare?

A genome can be defined as the entire DNA sequence of an organism. It contains all the instructions needed for an individual to develop and function. The human genome is distributed to 23 pairs of chromosomes inside the nucleus and a tiny chromosome in the mitochondria

The genome contains coding sequences (exons) and non-coding sequences (introns) of DNA. The coding region is responsible for producing proteins, while the non-coding sequences encode regulatory RNAs and repetitive sequences such as tandem repeats and interspersed repeats. In addition to this, genomic sequences include transposable elements, retrotransposons in eukaryotes, long terminal repeats (LTRs), and non-long terminal repeats (non-LTRs). 

The genomes of organisms can be compared based on: 

  • Size of the genome: Genomes in living beings are not equal. Tiny microorganisms such as bacteria and viruses have a small genomic size, whereas the genomes of plants can be much larger. To elaborate, the bacterium Escherichia coli contains 4.6 million base pairs, in contrast to the 150 billion base pairs of the genome of Paris japonica. Genome comparison typically begins with simpler features such as genome size, gene count, and chromosome number. However, a larger genome size, higher gene number, or more chromosomes does not indicate that the organism is more complex or functional. These features vary widely among species and are not always correlated with biological complexity.
  • Genome organization: Prokaryotes have circular genomes with single-stranded (ss) or double-stranded DNA (ds), whereas the genomes of eukaryotes are packaged into chromosomes. 
  • Coding region and non-coding regions: As the complexity of the organism increases, the non-coding regions increase. Microorganisms have little to no non-coding regions; therefore, they have a polycistronic expression system. In contrast to this, the eukaryotic genomes are more complex with non-coding regions interspersed between the coding regions. These regions have roles in RNA splicing during transcription. It follows a monocistronic expression system. 
  • Gene structure: A gene can be defined as a sequence of DNA that has a phenotypic function. A gene can be a culmination of exons with a series of introns in between. Genomic studies help us map the number of exons and introns, their lengths, and sequence similarity. 
  • Gene characteristics: The characteristics of the genes, such as splice sites for transcription, codon usage in translation, and a conserved block of sequences, can be compared within the studied genomes. 

Other comparisons include open-reading frames (ORFs), their average length, repetitive DNA  such as telomeres, junk DNA, etc. 

OrganismSize of the genomeNumber of chromosome pairsEstimated gene number
Human3.1 billion4625,000
Arabidopsis thaliana157 million1025,000
Fruit fly (Drosophila melanogaster)165 million813,000
Yeast (Saccharomyces cerevisiae)12 million326,000
Round worm (Caenorrhabditis elegans)97 million1219,000
Escherichia coli4.6 million13,200

Methods for Comparative Genomics

Comparative genomics uses a wide range of computational tools to process the large amount of data encoded within the genome. They are described as follows:

Genome Assembly and Genome Annotation

  • Genome assembly is the process of reconstructing a genome from millions or billions of small reads from DNA sequencing technologies. 
  • Genome assembly can be de novo (reconstruction without a reference) or reference-guided (reconstruction with a closely related reference genome). The prior is a lot more challenging and is done only for newly sequenced species, whereas the latter is relatively easier and more precise. 
  • Assembly uses bioinformatics pipelines such as FASTQC, trimmomatic, SAMtools, and others to read, clean, and process sequencing data for generating a high-quality genome sequence.
  • Genome annotation is the process of identifying genes, coding regions, non-coding regions, functional elements, and their regulatory features in the assembled genome.
  • It involves three key steps:

Repeat masking to identify and mask repetitive sequences in the genome and reduce its noise. 

Gene prediction involves locating the important constituents of the genome (exons, introns, transposons, etc.).

Functional annotation gives functions to the predicted genes. 

UCSC Genome Browser on Human (GRCh38/hg38)
UCSC Genome Browser on Human (GRCh38/hg38). Source: Human hg38 chr17:7,637,623-7,706,022 UCSC Genome Browser v482

Sequence alignment

  • Sequence alignment is the method of arranging DNA sequences to identify regions of similarity and differences in a species. As the genome size is larger, faster alignment methods utilizing global pairwise alignment are often used. 
  • FASTA and BLAST are among the most common methods for local and global pairwise alignment. However, specific programs such as MUMmer, VISTA, BLASTZ, DIAMOND, SyRI, etc., are tailored towards global alignment for comparative genomics. 

Identification of Homologous Regions

  • Homologs are the basis of evolutionary studies. They are mainly of two types: orthologs and paralogs. Orthologs are genes that have the same function in different species, whereas paralogs are genes that have undergone a whole-genome duplication event. Paralogs may disappear or evolve into new or specialized functions.
  • OrthoFinder is used to identify and extract orthologs and paralogs for sequences from multiple species.

Genome Mapping

  • Synteny is defined as the order of genes or markers on chromosomes across related species descended from a common ancestor. 
  • Programs like MCScanX build conserved syntenic blocks between genomes.  
  • For example, in the figure below, the colors in mouse chromosomes represent the homologous bands of the same color in humans. Homologous sequences in mouse chromosome 1 are mainly distributed in human chromosomes 1 and 2. 
Synteny of Human and Mouse
Synteny of Human and Mouse. Source: Sinha & Meller, 2007

Phylogenetic Analysis

  • Phylogenetic and evolutionary analysis is used to determine evolutionary relationships between common ancestors. 
  • Phylogeny can also be used to predict the time at which the mutation or a gene/allele was distributed in the population.
  • Trees are built by using pipelines such as IQTree, RaxML, etc.

Commonly Used Tools for Comparative Genomics

Because of the massive nature of genomic data, it can be extremely difficult to annotate and analyze data. The field utilizes a variety of computational tools to analyze large genomic data. A few of the most significant tools used in comparative genomics are attached below:

UCSC Browser: The UCSC browser contains all the reference sequences for a large collection of genomes. It can also track multiple sequence alignments and annotate a segment of DNA. 

Ensembl: It contains a large repertoire of genomic data for vertebrates and other eukaryotic organisms. 

OrthoFinder: The OrthoFinder tool uses orthogroups of species to determine an entire set of orthologous genes and their gene duplication events from genomic sequences. 

SyRI: SyRI or Synteny and Rearrangement Identifier for analyzing and visualizing synteny. It also predicts genomic differences between related genomes. 

MCScanX: MCScanX is a tool that can be used to scan multiple genomes to identify homologous regions of chromosomes, establish synteny, and align them by using genes as anchors. 

SynVisio: SynVisio can be used to analyze and visualize syntenic blocks and detect collinearity from the results of MCScanX. 

VISTA and PipMaker: Visualization Tool for Alignment (VISTA) and Percent Identity Plot Maker (PipMaker) are one of the most common visualization tools in comparative genomics. They use global alignment to turn raw sequence data from multiple species into graphical plots for interpretation.  

IGV: Integrative Genomics Viewer (IGV) is a widely used tool for analyzing and visualizing genomic data. It supports alignments, variants, and annotation across multiple genomes. 

MUMmer: MUMmer is a collection of tools for whole-genome alignment and comparison, often used for the identification of similarities, differences, and evolutionary events among varying genomic scales. 

Significance of Comparative Genomics

Phylogenetic studies

While genomes have an integral biological role within a species, when they are compared to those of another species, a more detailed picture of their evolution can be revealed. Such comparative studies help to track divergence, speciation events, whole genome duplications (WGD), and construct phylogenetic trees. These evolutionary studies have contributed to hypotheses, such as the 2R hypothesis, which proposes two rounds of WGD events in the early vertebrate evolution.

Phylogenetic Tree Showing the 2R Hypothesis
Phylogenetic Tree Showing the 2R Hypothesis. Source: Kasahara, 2007

Gene function prediction

Several genes still have an unknown function. Comparative genomics can help assign functions to such unannotated genes by comparison to known genes in reference organisms. This includes comparing a virulent strain of a pathogen (virus, bacteria) with non-virulent strains to identify the genes and their respective proteins.  

Identifying non-coding elements

Non-coding sequences are one of the most mysterious parts of the genome. They contain enhancers, promoters, or other regulatory elements. Comparing such sequences across species helps in identifying cis-regulatory motifs in gene regulation. Such studies can lead to the discovery of potential biomarkers in diseases such as cancers, in which the regulatory elements are disrupted. 

Study of Genome Organization and Synteny

Comparative genomics helps in the investigation of the evolution of chromosomes and their mutations, such as inversions, translocations, and duplications (ploidy). Syntenic blocks might help reconstruct ancestral genomes of extinct species. 

References

  1. Bornstein, K., Gryan, G., Chang, E. S., Marchler-Bauer, A., & Schneider, V. A. (2023). The NIH Comparative Genomics Resource: Addressing the promises and challenges of comparative genomics on human health. BMC Genomics, 24(1), 575. https://doi.org/10.1186/s12864-023-09643-4
  2. Genereux, D. P., Serres, A., Armstrong, J., Johnson, J., Marinescu, V. D., Murén, E., Juan, D., Bejerano, G., Casewell, N. R., Chemnick, L. G., Damas, J., Di Palma, F., Diekhans, M., Fiddes, I. T., Garber, M., Gladyshev, V. N., Goodman, L., Haerty, W., Houck, M. L., … Zoonomia Consortium. (2020). A comparative genomics multitool for scientific discovery and conservation. Nature, 587(7833), 240–245. https://doi.org/10.1038/s41586-020-2876-6
  3. Genetic maps and the use of synteny—PubMed. (n.d.). Retrieved June 17, 2025, from https://pubmed.ncbi.nlm.nih.gov/19347649/
  4. Human hg38 chr17:7,637,623-7,706,022 UCSC Genome Browser v482. (n.d.). Retrieved June 17, 2025, from https://genome.ucsc.edu/cgi-bin/hgTracks?db=hg38&lastVirtModeType=default&lastVirtModeExtraState=&virtModeType=default&virtMode=0&nonVirtPosition=&position=chr17%3A7637623%2D7706022&hgsid=2685246404_wrQ1r3I2ZjRzHtFjaS5TfaNAeQBZ
  5. Kasahara, M. (2007). The 2R hypothesis: An update. Current Opinion in Immunology, 19(5), 547–552. https://doi.org/10.1016/j.coi.2007.07.009
  6. OrthoFinder: Phylogenetic orthology inference for comparative genomics | Genome Biology | Full Text. (n.d.). Retrieved June 17, 2025, from https://genomebiology.biomedcentral.com/articles/10.1186/s13059-019-1832-y
  7. Pennacchio, L. A., & Rubin, E. M. (2003). Comparative genomic tools and databases: Providing insights into the human genome. Journal of Clinical Investigation, 111(8), 1099–1106. https://doi.org/10.1172/JCI17842
  8. Ph.D, D. S. P. (2019, May 22). What is Comparative Genomics? News-Medical. https://www.news-medical.net/life-sciences/What-is-Comparative-Genomics.aspx
  9. Sinha, A. U., & Meller, J. (2007). Cinteny: Flexible analysis and visualization of synteny and genome rearrangements in multiple organisms. BMC Bioinformatics, 8, 82. https://doi.org/10.1186/1471-2105-8-82
  10. Wang, Y., Tang, H., DeBarry, J. D., Tan, X., Li, J., Wang, X., Lee, T., Jin, H., Marler, B., Guo, H., Kissinger, J. C., & Paterson, A. H. (2012). MCScanX: A toolkit for detection and evolutionary analysis of gene synteny and collinearity. Nucleic Acids Research, 40(7), e49. https://doi.org/10.1093/nar/gkr1293

About Author

Photo of author

Rashal Shakya

Rashal Shakya has a bachelor’s degree (B.Tech.) in Biotechnology from Kathmandu University. He has actively contributed to multiple academic and research projects. His notable work includes the isolation and characterization of endophytic microbiomes in Paris polyphylla Sm., published in the Nepal Journal of Biotechnology. Rashal has gained hands-on experience through internships at leading research institutes, Kathmandu Research Institute for Biological Sciences (KRIBS) and Research Institute for Bioscience and Biotechnology (RIBB). With a growing interest in the intricacies of molecular biology and cellular machineries, he aims to contribute meaningfully to applied biosciences and translational research.

Leave a Comment