← blog
May 2026

Pathogenic synonymous variants: the signal that REVEL can't see

Synonymous variants don't change the protein, but they do alter the mRNA thermodynamic profile. First CNN predictor with pure thermodynamic signal: AUC 0.683. Conceptual framework Λ = S ⊗ Φ. Paper 1 of a trilogy.

bio-ia paper trilogy · paper 1

Genomics' blind spot

A synonymous variant is a DNA mutation that does not change the encoded amino acid. The codon changes, but the resulting protein is identical. For decades, synonymous variants were considered "silent" — with no functional effect.

We know that is false. Synonymous variants cause cystic fibrosis, cancer, and dozens of diseases. But the most powerful genomics predictors — REVEL, AlphaMissense, CADD — are fundamentally blind to them. They operate at the protein level: if the protein doesn't change, they have no signal to measure.

This is where QMetrika's two research lines converge.

If the protein doesn't change but the patient becomes ill, the pathology is in the mRNA, not in the protein. And the mRNA has its own physics that a synonymous mutation can alter.

The framework: Λ = S ⊗ Φ

Λ = S ⊗ Φ
Λ — biological viability of the transcript
S — genetic code (protein sequence) · semantic identity
Φ — biophysical constraints (mRNA thermodynamic form factor)
⊗ — systemic dependency between both axes

Every mRNA transcript has two layers of information. S is the digital message: the amino acid sequence it encodes. Φ is the analog channel: the mRNA thermodynamics that determine how efficiently that message is translated — local stability, translation speed, cotranslational folding.

In a missense variant, both axes change: ΔS ≠ 0, ΔΦ ≠ 0. In a synonymous variant, the argument is elegant: ΔS = 0, ΔΦ ≠ 0. The protein doesn't change, but the mRNA form factor does. If that causes disease, the pathology is purely thermodynamic.

This framework is presented in Paper 1 as a qualitative conceptual framework — not as formal algebra. Paper 2 makes it concrete in the opposite, stronger direction: the signal decomposes into two interpretable physical observables (signed stacking ΔΔG + G>A) that reproduce the deep classifier.

Pathogenic mechanisms

How can a variant that doesn't change the protein cause disease? Four mechanisms that the ΔG profile can capture:

mRNA destabilization

Alteration of local stacking that reduces thermodynamic stability → accelerated mRNA degradation → less protein produced.

Aberrant splicing

Local secondary structure change that alters splice sites or disrupts exonic enhancers/silencers (ESE/ESS).

Translation speed

Rare codons or stacking changes that slow the ribosome → cotranslational misfolding of the protein.

Cis regulation

Disruption of regulatory elements that depend on the nucleotide sequence, not the protein sequence.

Results

0.683
Dedicated AUC
0.625
LOGO AUC
2.60×
G>A enrichment

Dedicated CNN: AUC 0.683 using only mRNA thermodynamic signal (ΔG channels). No evolutionary information, no conservation, no 3D structure. Pure nucleotide physics.

LOGO (Leave-One-Gene-Out): AUC 0.625 — the signal generalizes across genes; it is not an artifact of a particular gene.

Mechanistic finding — G>A enrichment: Pathogenic synonymous variants are 2.60× enriched in G>A transitions. This has a direct thermodynamic explanation: G>A reduces stacking (guanine stacks more strongly than adenine), and this effect is captured by the ΔG profile.

CDS length effect: Genes with longer CDS have stronger thermodynamic signal. This is consistent with the model: longer mRNAs have more opportunities for a local perturbation to propagate.

What this implies

An AUC of 0.683 does not compete with the best missense predictors (which reach 0.99). But it doesn't have to. The question is not "Is it better than AlphaMissense?" but "Does a thermodynamic signal exist in synonymous variants?" And the answer is yes.

This opens a niche that no published predictor covers. REVEL, AlphaMissense, CADD — all are blind to synonymous variants because they operate at the protein level. EnergyFingerprint operates at the nucleotide level. The biophysical channel already works for synonymous variants without modification.

EnergyFingerprint is not competing with AlphaMissense on its turf. It opens new ground where nobody can operate — because nobody else looks at the physics of the mRNA.

The trilogy

This paper is the first of three that progressively unfold the Λ = S ⊗ Φ framework:

Three papers, one escalating question

1

"The signal exists" (this paper) — A dedicated CNN demonstrates that the thermodynamic signal discriminates pathogenic synonymous variants from benign ones. The Λ = S ⊗ Φ framework is presented as a qualitative conceptual framework.

2

"The signal is interpretable" (horizon 2027) — The dedicated CNN reduces to two physical observables: the signed nearest-neighbor stacking ΔΔG (Turner parameters) plus the G>A mutational bias. A two-feature model matches the deep network and generalizes across genes — full interpretability of the σ operator, decisive for clinical adoption.

3

"The predicted physics occurs in the cell" (horizon 2027-2028) — Direct experimental validation. DMS-MaPseq measures in vivo accessibility. Reporter assays measure expression. Does predicted ΔΦ = actual structural change?

L1 × L2 convergence

This work is the natural convergence of QMetrika's two research lines. L1 (EnergyFingerprint) provides the CNN and the transferability gradient. L2 (mRNA Design) demonstrates that the ΔG profile is a useful and orthogonal dimension.

Together they predict exactly this result: if ΔG is informative for missense and orthogonal to MFE, it should also be informative for synonymous variants — where it is the only available signal, because the protein doesn't change.

The natural extension is the evolutionary channel. ESM-1v operates at the protein level — for synonymous variants it does not directly apply. But a masked marginal at the DNA level (not protein) could capture evolutionary pressure on the nucleotide sequence. That is the subject of Paper 2.

Read it in the open

Preprint with dedicated CNN, LOGO analysis, mechanistic findings and the Λ = S ⊗ Φ framework.