First, a point of clarification:
coding and
encoding are two different things. "Encoding" is a CS term that doesn't come up much in biology. (You could say that DNA is an
encoding of a base-4 number sequence.) "Coding" is a bioinformatics term — DNA base-pairs
code for particular RNA sequences; and then, only if they're in a
coding region. (Analogy: flux patterns on a floppy disk
code for particular bytes, but only if they're in a
sector.)
Second: what are genes? A "gene" is an abstraction.
Let's first define a "genome". A genome is the complete dump of raw data of "what the DNA base-pair sequence string says" — before any coding or expression occurs. Whenever we sequence DNA, it's [parts of] this raw base-pair string that we get back — not the codons, not the expressed RNA sequences. And this is why we care — sequencing is the tool we have, so it's the lens by which we look at DNA. And that lens shows us the raw data.
When we talk about "genes" (or "SNPs" / other bioinformatics terms), we're basically talking about the ways in which that raw data we can extract through sequencing, relates to observed phenotypic changes. A "gene" is a particular part of a "genome" (DNA base-pair sequence) that can be uniquely identified as being the cause of interesting phenotypic consequences when changes are made to it. (Which in turn is how we figure out what genes are "responsible for" doing what.)
Note how these concepts, "genome" and "gene", both completely ignore epigenetics / gene expression. That's because these terms were invented before epigenetics was invented, and these terms are part of a model that uses a lens (DNA sequencing) that itself ignores epigenetics. Gene sequencing shows you the world of DNA as if epigenetics didn't exist.
Now let's talk about DNA.
If you think that epigenetics is about changing how DNA encodes information, then you might be thinking of "DNA" via the lie-to-children model, of it just being a long sequence of nucleotides.
But consider: why do chromosomes look the way they do — little X shapes? A long string of base pairs, on its own, has no reason to assemble into that macroscopic shape.
This is because "DNA" — which is really a shorthand for "the DNA complex" (i.e. a complex of multiple weakly-bound molecules that together form the chromatin of a chromosome) — is not just nucleotide base-pairs. The nucleotide base-pairs form one molecule (the deoxyribonucleic acid string itself); but then you've got other stuff. You've got histones — little tape-spools the DNA string is wrapped onto. You've got methyl groups — little markers hanging off particular places on the DNA string. You've got other stuff I don't personally understand/know about.
When lay-people talk about "DNA", they're really referring to the whole complex of molecules that makes up each one of your chromosomes.
"Genes" are just the bioinformatic data representation of the deoxyribonucleic-acid-base-pair-string part of that complex of molecules.
But gene coding, and gene expression, are both a result of what the other molecules in that complex are doing.
(Think about it: if gene coding + expression were encoded in-band by the DNA itself, then that information would be appear in DNA sequencing — and so we'd be unable to differentiate that information from changes to the DNA itself — and so we wouldn't even have a concept of "epigenetics", because it would all just look like "genetics"!)
Here's a picture: https://www.genome.gov/sites/default/files/media/images/2022...
Histones — those little tape spools — attract or repel each-other due to chemical modifications to the histones themselves. Two histones that "snap together", prevent the region of DNA "tape" between them from being physically accessed by the RNA polymerase enzymes that "read the tape" to produce RNA.
Methylation markers hang off of the initial "landing sites" for RNA polymerase enzymes (CpG dinucleotides — think "floppy-disk sector header"), and repel them.
When you sequence DNA, you ignore these transcription-silencing signals. You can picture DNA sequencing as "unrolling" the deoxyribonucleic acid off of its histone carriers, rendering them irrelevant; and then using molecular tools to read the sequence — tools that, unlike RNA polymerase, are not repelled by methyl-groups.
If you use more-modern techniques to capture this additional information about where these silenced regions are, and by what mechanism they've been silenced, then you get what we call an "epigenome" — which isn't a post-translated version of DNA, but rather sort of a "metadata track" that runs parallel to the DNA "data track."
(And you can, in theory, combine the two to calculate an "expressed genome." I don't think we've ever done that yet — partly because "whole-epigenome sequencing" doesn't yet exist in the way that "whole-genome sequencing" does; and partly because epigenetic metadata is probabilistic — with some modifications decreasing the probability of a region coding for something, rather than turning it off altogether — and so current approaches, that derive the epigenome "by reaction", would observe something like "weak bits" on a floppy disk — regions that read differently each time, requiring many passes to calculate a "flux strength" for each region and to find each silenced region's true borders.)