Add BrightSurf on Google Email

Artificial Intelligence: Language models that see every letter

10.09.26 | Ludwig-Maximilians-Universität München

Most AI language models never see the individual letters of a word directly. A new method changes that by retrofitting existing models at a fraction of the usual training cost.

Before large language models (LLMs) can process text, they split it into chunks such as words or word fragments. The LLM behind ChatGPT, for example, splits “LMU München” into the three chunks “LM”, “U” and “München”, so it never directly sees the individual letters in “München”. This step, which is known as subword tokenization, makes LLMs efficient but causes a range of problems, such as limited character-level understanding. Models that instead read text byte by byte (the basic units in which computers store text, roughly one per letter) avoid these problems. In practice, however, they have not yet become a viable alternative to LLMs based on subword tokenization.

In a new study published in the journal Nature , researchers from LMU, the Allen Institute for AI, the University of Cambridge, the University of Washington and Imperial College London have developed a method that converts existing LLMs into byte-level models without retraining them from scratch. This “byteification” requires less than one percent of the training typically needed for such a model and largely preserves the performance of the original LLM while adding the benefits of byte-level processing. The resulting models outperform all previously published byte-level LLMs of comparable size.

“We believe that byteification, and byte-level LLMs more generally, have the potential to overcome some long-standing shortcomings of LLMs,” says Valentin Hofmann , Junior Professor for Information and Language Processing Using AI Methods at LMU Munich and last author of the study. “Traditional LLMs struggle with tasks that require character-level capabilities, such as spelling a word backwards. Our byteified models are substantively better at this.” These capabilities are more than an academic curiosity: “A good representation of the low-level structure of text is critical in many areas of science, for example when working with code or biological sequences,” Hofmann explains.

The authors hope that byteification will make byte-level models a practical alternative to today’s LLMs and open up new research directions. The models, code, and training data are publicly available.

Nature

10.1038/s41586-026-11111-4

Retrofitting language models to operate over bytes

7-Oct-2026

Keywords

Article Information

Contact Information

Dominic Anders
Ludwig-Maximilians-Universität München
Dominic.Anders@lmu.de

Source

This article is based on a news release from Ludwig-Maximilians-Universität München. BrightSurf curates and republishes science news from research institutions worldwide; the original release is linked below.

How to Cite This Article

APA:
Ludwig-Maximilians-Universität München. (2026, October 9). Artificial Intelligence: Language models that see every letter. Brightsurf News. https://www.brightsurf.com/news/LPE494O8/artificial-intelligence-language-models-that-see-every-letter.html
MLA:
"Artificial Intelligence: Language models that see every letter." Brightsurf News, Oct. 9 2026, https://www.brightsurf.com/news/LPE494O8/artificial-intelligence-language-models-that-see-every-letter.html.