Artificial Intelligence: Language models that see every letter
en-GBde-DEes-ESfr-FR

Artificial Intelligence: Language models that see every letter


Most AI language models never see the individual letters of a word directly. A new method changes that by retrofitting existing models at a fraction of the usual training cost.

Before large language models (LLMs) can process text, they split it into chunks such as words or word fragments. The LLM behind ChatGPT, for example, splits “LMU München” into the three chunks “LM”, “U” and “München”, so it never directly sees the individual letters in “München”. This step, which is known as subword tokenization, makes LLMs efficient but causes a range of problems, such as limited character-level understanding. Models that instead read text byte by byte (the basic units in which computers store text, roughly one per letter) avoid these problems. In practice, however, they have not yet become a viable alternative to LLMs based on subword tokenization.

In a new study published in the journal Nature, researchers from LMU, the Allen Institute for AI, the University of Cambridge, the University of Washington and Imperial College London have developed a method that converts existing LLMs into byte-level models without retraining them from scratch. This “byteification” requires less than one percent of the training typically needed for such a model and largely preserves the performance of the original LLM while adding the benefits of byte-level processing. The resulting models outperform all previously published byte-level LLMs of comparable size.

“We believe that byteification, and byte-level LLMs more generally, have the potential to overcome some long-standing shortcomings of LLMs,” says Valentin Hofmann, Junior Professor for Information and Language Processing Using AI Methods at LMU Munich and last author of the study. “Traditional LLMs struggle with tasks that require character-level capabilities, such as spelling a word backwards. Our byteified models are substantively better at this.” These capabilities are more than an academic curiosity: “A good representation of the low-level structure of text is critical in many areas of science, for example when working with code or biological sequences,” Hofmann explains.

The authors hope that byteification will make byte-level models a practical alternative to today’s LLMs and open up new research directions. The models, code, and training data are publicly available.

Minixhofer, B., Murray, T., Limisiewicz, T. et al. Retrofitting language models to operate over bytes. Nature (2026).
DOI: https://doi.org/10.1038/s41586-026-11111-4
Regions: Europe, Germany
Keywords: Applied science, Artificial Intelligence

Disclaimer: AlphaGalileo is not responsible for the accuracy of content posted to AlphaGalileo by contributing institutions or for the use of any information through the AlphaGalileo system.

Testimonios

We have used AlphaGalileo since its foundation but frankly we need it more than ever now to ensure our research news is heard across Europe, Asia and North America. As one of the UK’s leading research universities we want to continue to work with other outstanding researchers in Europe. AlphaGalileo helps us to continue to bring our research story to them and the rest of the world.
Peter Dunn, Director of Press and Media Relations at the University of Warwick
AlphaGalileo has helped us more than double our reach at SciDev.Net. The service has enabled our journalists around the world to reach the mainstream media with articles about the impact of science on people in low- and middle-income countries, leading to big increases in the number of SciDev.Net articles that have been republished.
Ben Deighton, SciDevNet
AlphaGalileo is a great source of global research news. I use it regularly.
Robert Lee Hotz, LA Times

Trabajamos en estrecha colaboración con...


  • The Research Council of Norway
  • SciDevNet
  • Swiss National Science Foundation
  • iesResearch
Copyright 2026 by DNN Corp Terms Of Use Privacy Statement