top of page

How DNA Language Models Are Changing Codon Optimisation

  • Aug 12
  • 5 min read

Updated: 5 days ago

A promising new drug candidate may have the right target, binding profile, and functional activity, yet still become difficult to progress because its coding sequence does not support efficient protein production. When poor expression is encountered, the usual approach is to redesign the construct, synthesise the new sequence, test again, and repeat. 


Suboptimal codon usage is one potential culprit that can lead to poor protein expression due to reduced titres and increased protein misfolding. Codon optimisation is meant to solve this problem. However, conventional codon optimisation approaches often reduce the process to a familiar, but overly simplistic rule: replace problematic codons with those preferred by the expression host.


This approach may improve basic host compatibility, but, more often than not, it fails to capture the full sequence-design challenge.


While synonymous DNA sequences encode the same protein, they do not necessarily translate to equal expression levels. Codon choice can affect mRNA structure, translation dynamics, regulatory motifs, and co-translational folding1. These properties emerge across the transcript, not one codon at a time.


Instead of optimising each position in isolation, the EMLy™ codon optimisation pipeline takes a fundamentally different approach. Using a DNA large language model (DNA-LLM), EMLy models the coding sequence as a whole and generates multiple context-aware designs for experimental testing. Keep reading to see how this approach captures what conventional codon-optimisation methods may miss.



The problem with isolated codon optimisation


Traditional codon-optimisation tools are built around host-specific codon-usage tables. These tables describe which synonymous codons are commonly found in a target organism, often using highly expressed genes as a reference2.


Unfortunately, the most frequently used codon is not automatically the best choice in every sequence context.


Changing one codon can alter more than host adaptation. Depending on where it sits in the transcript, a synonymous substitution may affect:

  • mRNA secondary structure

  • local and regional GC content

  • translation kinetics

  • codon-pair patterns

  • cryptic splice or regulatory motifs

  • co-translational protein folding


A design that scores well on codon adaptation may therefore still contain sequence features that limit expression.


The challenge becomes more complex as the number of interacting constraints increases. Improve one metric, and another may get worse. Remove one motif and a new structural feature may appear elsewhere. Apply the same codon preferences repeatedly and the resulting sequences may cluster around a narrow set of similar designs.


These interacting constraints highlight why codon optimisation should not be thought of as a simple substitution problem, but as a global optimisation problem.


Flowchart illustrating how codon optimisation
Flowchart illustrating how codon optimisation alters transcript-level properties like host codon compatibility, mRNA secondary structure, translation kinetics, and co-translational folding, impacting protein expression levels, despite identical protein sequences.

From static rules to learned sequence context


Rule-based tools can screen for known sequence liabilities. They may remove restriction sites, maintain a target GC range, avoid repetitive elements, or reduce predicted RNA structures near the start codon.


But every additional rule increases the complexity of the design space.

Machine learning approaches began to address this by learning patterns directly from biological sequences. Recurrent neural networks, for example, can incorporate neighbouring sequence information rather than relying exclusively on fixed codon frequencies.


Transformer-based language models extend this further.


Originally developed for natural language, transformer architectures use attention mechanisms to model relationships among positions across a sequence. Applied to DNA, this makes it possible to evaluate codon choices in the context of the wider transcript rather than only the position immediately being optimised.


That is the shift DNA language models bring to codon optimisation: from static preferences to learned, context-dependent sequence generation.


DNA Large Language Model (LLM) for codon optimisation
Diagram illustrating the training process of a DNA Large Language Model (LLM) for codon optimisation, highlighting the consideration of global sequence context and the application of region-specific mutations.

 

What a DNA language model sees differently


A DNA language model does not understand a coding sequence in the human sense. It learns statistical relationships from large collections of biological sequences and uses those patterns to generate or evaluate new designs.


For codon optimisation, that enables several important capabilities.


Full-transcript context

Each codon choice can be evaluated in relation to the surrounding sequence and more distant positions across the transcript.


Multiple candidate designs

Instead of producing one deterministic “optimised” sequence, the model can generate several synonymous alternatives for comparison and testing.


Learned sequence patterns

The model can identify recurring patterns in natural coding sequences that individual codon-frequency scores may not capture.

CodonTransformer is one example of this approach. The multispecies model was trained on more than one million DNA–protein pairs from 164 organisms and combines organism, amino acid, and codon information to generate host-specific coding sequences.


The practical difference is significant. Codon optimisation no longer has to ask only:

"What is the preferred codon at this position?"

It can ask:

"What coding sequence is most appropriate in the context of this entire transcript?"



How EMLy applies DNA language models to codon optimisation


The codon optimisation pipeline in EMLy™ is purpose-built to focus on global sequence design context.


The workflow uses a host-conditioned DNA language model to generate multiple synonymous coding sequences. For antibody and TCR programmes, Etcembly further fine-tunes the model on immune-receptor sequence data to reflect the unique structural and compositional characteristics of these molecules.


Each candidate is passed through Etcembly’s BioSmartness refinement workflow, which evaluates predicted 5′ RNA structure, GC and CpG content, restriction sites, cryptic splice sites, internal poly-A/T tracts, and other sequence-level constraints.


The output is not one sequence. It is a diverse, host-aware candidate panel with each sequence scored and ranked to give the experimental programme a stronger starting point.


In internal and partnered studies, more than 96% of tested EMLy-optimised sequences improved antibody yield relative to matched controls across IgG, VHH, and multispecific formats.


What this means for biologics development


For a biotech or pharma team working with a low-expressing antibody, TCR, bispecific, or other recombinant protein, the candidate's coding sequence can determine how much time and resources are spent before a usable construct is identified.


DNA language models change the economics of that search. They move more of the sequence exploration in silico, account for broader sequence context, and give teams multiple candidates to validate rather than one design to hope for.


The lab remains essential. But it is used to confirm the most promising sequences, not to discover them through repeated trial and error.


Keep an eye out on Etcembly's white paper (coming out soon) to explore the EMLy codon-optimisation workflow, BioSmartness refinement, and validation across multiple antibody formats.




If you have a TCR or antibody engineering challenge, we'd like to hear about it. Get in touch at hello@etcembly.io

 

References

Mauro VP & Chappell SA. A critical analysis of codon optimization in human therapeutics. Trends Mol Med. 20: 604–13 (2014). doi: 10.1016/j.molmed.2014.09.003

Athey J, et al. A new and updated resource for codon usage tables. BMC Bioinformatics. 18: 391 (2017). doi: 10.1186/s12859-017-1793-7

 
 
 

Comments


Commenting on this post isn't available anymore. Contact the site owner for more info.
bottom of page