Eukaryotic Genome: Features, Organization, Genes, Regions

The eukaryotic genome refers to the entire genetic material in eukaryotes, ranging from single-celled fungi to complex multicellular organisms.

Eukaryotic Genome
Eukaryotic Genome

Eukaryotes have a wide range of DNA inside the nucleus, mitochondria, and plastids (in plants). The arrangement of the genome is more organized, efficient, and contains functional and non-functional regions.

General Features of the Eukaryotic Genome

  • In contrast to the prokaryotic genome, eukaryotic genomes are larger, about 10 Mb to 100,000 Mb. Because of this, the number of functional genes is also enormous.
  • The DNA is packaged into chromosomes, which are all linear and are made up of nucleosomes. A chromosome comprises euchromatin and heterochromatin regions for varying degrees of expression. 
  • The genome is normally diploid and possesses two copies within a cell, except for sex-related cells. This, however, varies with the organism. 
  • Apart from the genomic DNA inside the nucleus, eukaryotic cells have mitochondrial (mt) DNA in mitochondria, and chloroplast DNA (cpDNA) in chloroplasts. 
  • The eukaryotic genome consists of coding regions, non-coding regions, transposons, intergenic segments, repeated sequences,  and others. 
  • All of the eukaryotic genomes have introns and exons. The introns are spliced in transcription, whereas the exons are utilized in the chromosome. 
  • Compared to prokaryotic genomes, horizontal gene transfer is quite rare in eukaryotes; however, genome/gene duplication events are frequent. 

Size of the Eukaryotic Genome

  • The size of the eukaryotic genome is enormous and can range from less than 10 Mb to well over 100,000 Mb. The size depends on the organism; plants are known for their high genomic size, but some organisms have a 3 MB size (the smallest eukaryotic genome: 0.38 Mb). 
  • In this regard, the genome size depends on the complexity of the organisms; however, similar organisms in the same clade can have vastly different genome sizes. For example, humans have a genome size of about 4200 Mb, but lizards, which are less complicated than us have a genome size of 1570 Mb. This is known as the C-value paradox.
  • The complexity depends on the gene density rather than the genomic size. Gene density is a simple measure of the average number of genes per megabase pair (Mb) of DNA. To elaborate, if an organism has a genomic size of about 50 Mb, its gene density is 100 genes/Mb. 
  • A rough correlation between the complexity and gene density can be elucidated, i.e., the less complex the organism, the higher its gene density.  
  • The simplest unicellular eukaryote, Saccharomyces cerevisiae, has a greater gene density (about 500/Mb than that of prokaryotes; however, the gene density of Homo sapiens is even less (about 50-fold decrease).
OrganismGenome Size (Mb)Number of GenesNumber of Chromosomes
Protists
Entamoeba histolytica20 8,20014
Plasmodium falciparum23.8~5,300
Tetrahymena thermophila>120~25,0005
Fungi
Sachharomyces cerevisiae12.26,00016
Aspergillus niger34.8512,000
Amanita muscaria40.7018,000
Plants
Arabidopsis thaliana12033,4105
Oryza sativa43067,39312
Zea mays2,50053,76410
Animals
Homo sapiens320020,000-25,00023
Caenorhabditis elegans10020,0006
Drosophila melanogaster18013,6004

Structural Components of the Eukaryotic Genome

Genome Organization
Genome Organization

Chromosomes

  • Since the genomes of the eukaryotes are substantially larger, the DNA needs to be condensed and packaged to fit inside the tiny cell. Therefore, the DNA is wrapped around a set of histone proteins, called the nucleosome. 
  • Each nucleosome contains over 100 million bp of DNA, meaning this association shortens the length of DNA by about 7-fold. To elaborate, a DNA of length 1 m will be condensed to a string-of-beads chromatin fiber of 14 cm. 
  • However, this shortening is still large and cannot possibly fit inside the 10-20 micrometer nucleus. So, the chromatin is further coiled into a much thicker and shorter fiber about the size of 30 nm in diameter, a structure known as the 30-nanometer fiber (30 nm).
  • This fiber forms loops of 300 nm in length, which are compressed and folded to produce a 250 nm-wide fiber. Finally, the tight coiling of the 250-nm fiber produces the chromatid of a chromosome. 
  • The chromosome is divided into heterochromatin and euchromatin regions. The prior is tightly associated with the nucleosome, whereas the latter is loosely arranged. This controls the expression of the genes.
Chromosome
Chromosome

Telomeres

  • Telomeres are repetitive sequences of DNA located at the ends of a linear chromosome in eukaryotes. 
  • The repeats generally comprise short TG-rich sequences. These sequences, however, vary with organisms. To elaborate, human telomeres have a repeating sequence of 5’-TTAGGG-3’.  
  • The replication of telomeres is carried out by a special enzyme known as telomerase.
  • Telomeres protect the ends of the chromosomes during replication. 
Telomeres protect the ends of the Chromosomes
Telomeres protect the ends of the Chromosomes

Centromere

  • The centromere is the region of the eukaryotic chromosome where segregation takes place during cell division. 
  • The replicated copies of chromosomes, the sister chromatids, are separated into two daughter cells by the action of a protein complex known as the kinetochore, which assembles at the centromeres. 
  • The kinetochore assembles at each centromeric DNA and binds to protein filaments called microtubules that eventually pull and segregate the sister chromatids. 
  • When this region is absent, the separation process is random. The presence of more than one centromere is also disastrous, as the association of two centromeres pulling in opposite directions can cause chromosomal breakage. 
Condition of Separation by Centromeres. A. Normal condition; only one centromere. B. No centromeres. C. Two centromeres
Condition of Separation by Centromeres. A. Normal condition; only one centromere. B. No centromeres. C. Two centromeres . Source: Pederson, 2015

Origin of Replication in the Eukaryotic Genome

  • The origin of replication (ORI) is the site of DNA at which the DNA replication machinery is assembled, and replication is initiated. 
  • In a eukaryotic cell, ORIs are found 30-40 kb apart throughout the length of each eukaryotic chromosome. They are situated in the non-coding regions.
  • Unlike prokaryotes, ORIs are present in multiple copies in a single chromosome.

Genes and Related Sequences

Genes

  • Genes are the basic and functional units of heredity. They are protein-coding regions of the genome. Humans consist of about 42 Mb of genes.  
  • Exons encode amino acids, which are interspaced by non-coding regions of introns. Most of the genes contain about 10 introns. 
  • The variation in these regions can have a detrimental impact that can alter protein function or expression, influencing physical traits, disease onset, and drug response. 
  • The reshuffling of genes during meiosis creates variation and diversity in the offspring. 

Related sequences

Non-coding parts of genes include introns, UTRs, pseudogenes, and gene fragments. 

Introns

  • Introns are segments within the genome that do not encode proteins or structural RNAs. They are spliced out or removed during transcription. 
  • About 60% of the human genome is composed of introns, and many of these portions have no known function. The human genome constitutes 95% introns and only 5% exons. 
  • Introns are responsible for alternative splicing, regulation of gene expression, and are also a source of regulatory non-coding RNAs. They promote variation, allowing proteins to gain, lose, or shuffle functional properties. 
Central Dogma in Eukaryotes
Central Dogma in Eukaryotes

Gene Fragments, Pseudogenes, and Reverse Transcriptase

  • Because of the growing complexity of organisms, the segments that are needed to regulate sequences also grow in complexity and in size. 
  • The unique introns of the genome comprise non-functional mutant DNA, gene fragments, and pseudogenes. 
  • The pseudogenes are a result of the enzyme known as reverse transcriptase, which synthesizes double-stranded DNA (dsDNA) from RNA. This forms the copy DNA or cDNA. 
  • Reverse transcriptase is only expressed in a certain type of viruses that require this enzyme to reproduce. However, in some rare instances, cellular mRNAs are copied into host DNA, resulting in DNA fragments that reintegrate into the genome. 
  • These copies do not have any function because introns are rapidly removed from the copy of the gene from which they are derived. They lack the promoter and other regulatory sequences to direct expression. 
 Functional Gene and Non-functional Pseudogene
 Functional Gene and Non-functional Pseudogene

Promoters and Enhancers

  • In eukaryotes, promoters are regions of DNA located upstream of the gene where transcriptional machinery, such as RNA polymerase, transcription factors, and regulatory complexes, bind to initiate RNA synthesis.  
  • A promoter contains specific sequences like the TATA box (consensus sequence: TATAAA), which help recruit TFIIB and RNA polymerase II for transcription. 
  • On the other hand, enhancers are transcriptional start sites located several kilobases away from the promoters that promote the transcriptional activity of associated genes. They have short DNA motifs of about 6 to 10 bp, and are recognized by transcriptional activators such as Oct4, Sox2, etc.
  • Enhancers can loop to their target promoters, even at vast distances (tens of thousands of nucleotides away), to bring protein complexes to proximity for transcription. 
Role of Promoters and Enhancers in Transcription of the Eukaryotic Genome
Role of Promoters and Enhancers in Transcription of the Eukaryotic Genome

MicroRNA and lincRNAs

  • The vastness of the unique intergenic regions might have functional roles that have not yet been understood. 
  • MicroRNA (miRNA) is a tiny structural RNA that regulates the expression of other genes by altering either the stability of the mRNA product or its ability to be translated. 
  • It has been estimated that human cells contain more than 500 miRNA genes. 
  • Similarly, longer segments known as long-intervening non-coding RNAs (lincRNAs) also occupy a substantial amount of the genome. These RNAs do not encode for any proteins, but they act to regulate gene expression in a positive and negative manner, which has not been determined. 

Noncoding and Repetitive DNA

Intergenic Regions

Intergenic regions, unlike introns, are vast segments of DNA located between genes. They are composed of non-functional “junk DNA” (pseudogenes, repeats, transposons), and functional regulatory sequences (promoters, enhancers, origin of replication). They can be unique and repeated. One of the major contributors to the increase in these intergenic regions is direct transcription regulation, called regulatory sequences.

Microsatellites

  • Microsatellites are short repetitive DNA motifs, about the size of 1-6 base pairs, that occur thousands of times in the eukaryotic genome. Almost half of the human genome is composed of these short repeated sequences.
  • Microsatellites make up about 3% of the human genome and are found in non-coding as well as in coding regions. They are abundantly found near telomeres and centromeres, in varying frequencies across the species.  
  • The most common microsatellites are dinucleotide repeats such as CACACACA, ATATATATA, etc. They arise from difficulties in replicating the DNA. 

Genome-wide Repeats

  • Genome-wide repeats are a large chunk of repetitive DNA, much larger than the microsatellites. Each unit is greater than 100 bp, and many are above 1000 bp. 
  • These repeats can be found as single copies dispersed throughout the genome, or even as closely spaced clusters. 
  • The most common feature of genome-wide repeats is that they are “jumping DNA” or transposons.

Transposable elements

  • Transposable elements or transposons are segments of DNA that can move from one place to another in the genome. Through transposition, these sequences can multiply and accumulate throughout the genome. 
  • Compared to prokaryotes, the movement of transposons is relatively rare in human cells. However, throughout evolution, these elements have been successful at propagating throughout the genome (more than 45%). 

Organelle Genomes

In addition to the nuclear genome, eukaryotic cells also contain separate genomes inside the mitochondria and plastids. These organelles are derived from ancient prokaryotic organisms that were engulfed by primitive eukaryotes via a process called endosymbiosis. Because of this reason, the organelle genome follows the prokaryotic polycistronic gene expression, in which multiple genes are transcribed into a single mRNA transcript. 

Mitochondrial genome

  • Mitochondrial DNA (mtDNA) is a circular genome about the size of 16,569 bp. It lacks introns and is mostly compact with coding regions.  
  • In humans, the mt genome comprises 37 genes that encode for 13 proteins, 22 transfer RNAs (tRNAs), and 2 ribosomal RNAs (rRNAs). These genes are responsible for the synthesis of proteins related to oxidative phosphorylation.
  • The mitochondrial genome is maternally inherited, meaning that it is passed down from the mother to the offspring.  
Mitochondrial Genome (mtDNA) of Humans
Mitochondrial Genome (mtDNA) of Humans. Source: Butenko et al., 2024

Plastids genome

  • Plastids, such as chloroplasts, are found only in plants, algae, and a few protists. These organelles have their own genome as they were once a single organism that got integrated through endosymbiosis. 
  • The chloroplast genome (cpDNA) is about 110-150 kb in length and encodes about 150 genes. The size and number of genes within this genome vary with species. In the case of Arabidopsis thaliana, the genome encodes 117 proteins. 
  • Unlike the nucleus, chloroplasts can be present in multiple copies per plant cell (up to 300 copies). Since each of these organelles has its own DNA, there is an overall abundance of cpDNA in the cell, leading to more chloroplast-associated gene expression. 
  • The chloroplast (plastid) DNA is maternally inherited, because of which there is a lower risk of inheriting a transgenic plant. 
Chloroplast genome (cpDNA) of Arabidopsis thaliana
Chloroplast genome (cpDNA) of Arabidopsis thaliana.

References

  1. Butenko, A., Lukeš, J., Speijer, D., & Wideman, J. G. (2024). Mitochondrial genomes revisited: Why do different lineages retain different genes? BMC Biology, 22(1), 15. https://doi.org/10.1186/s12915-024-01824-1
  2. Chauhan, D. T. (2018, August 22). What is DNA packaging in eukaryotes? Genetic Education. https://geneticeducation.co.in/what-is-dna-packaging-in-eukaryotes/
  3. Cooper, G. M. (2000). The Complexity of Eukaryotic Genomes. In The Cell: A Molecular Approach. 2nd edition. Sinauer Associates. https://www.ncbi.nlm.nih.gov/books/NBK9846/
  4. Elliott, T. A., & Gregory, T. R. (2015). What’s in a genome? The C-value enigma and the evolution of eukaryotic genome content. Philosophical Transactions of the Royal Society B: Biological Sciences, 370(1678), 20140331. https://doi.org/10.1098/rstb.2014.0331
  5. Eukaryotic Genome Complexity | Learn Science at Scitable. (n.d.). Retrieved July 14, 2025, from http://www.nature.com/scitable/topicpage/eukaryotic-genome-complexity-437
  6. Gene Expression | Learn Science at Scitable. (n.d.). Retrieved July 14, 2025, from https://www.nature.com/scitable/topicpage/gene-expression-14121669/
  7. Mitochondrial DNA in human identification: A review [PeerJ]. (n.d.). Retrieved July 14, 2025, from https://peerj.com/articles/7314/
  8. Pederson, T. (2015). Molecular Biology of the Gene: By James D. Watson: W. A. Benjamin (1965): New York, New York. FASEB Journal: Official Publication of the Federation of American Societies for Experimental Biology, 29(11), 4399–4401. https://doi.org/10.1096/fj.15-1101ufm
  9. Promoter. (n.d.). Retrieved July 14, 2025, from https://www.genome.gov/genetics-glossary/Promoter
  10. Saitou, N. (2013). Eukaryote Genomes. Introduction to Evolutionary Genomics, 17, 193–222. https://doi.org/10.1007/978-1-4471-5304-7_8
  11. Tutar, Y. (2012). Pseudogenes. Comparative and Functional Genomics, 2012, 424526. https://doi.org/10.1155/2012/424526

About Author

Photo of author

Rashal Shakya

Rashal Shakya has a bachelor’s degree (B.Tech.) in Biotechnology from Kathmandu University. He has actively contributed to multiple academic and research projects. His notable work includes the isolation and characterization of endophytic microbiomes in Paris polyphylla Sm., published in the Nepal Journal of Biotechnology. Rashal has gained hands-on experience through internships at leading research institutes, Kathmandu Research Institute for Biological Sciences (KRIBS) and Research Institute for Bioscience and Biotechnology (RIBB). With a growing interest in the intricacies of molecular biology and cellular machineries, he aims to contribute meaningfully to applied biosciences and translational research.

Leave a Comment