Researchers have proposed a way for a language model to keep its internal working state between words instead of rebuilding it every time. The paper, posted to arXiv on 29 September 2026, names the design LIFT, short for Latent Information Feedback Transformer.
Today’s models write one token at a time, and the chosen token is the only thing that travels from one step to the next. LIFT adds a second channel, a compact summary of the model’s internal state that rides along with each token.
Training stays parallel because the learning-time states come from an existing pretrained model. At generation time the model uses its own predicted states. The authors say this adds a small cost that shrinks as models grow.
All results are the authors’ own, on models from 135 million to 1 billion parameters. They report gains over standard transformers on language modeling and reasoning when training text is held equal. On one state-tracking task, a tiny LIFT beat same-size models trained on eight times more data.
No frontier-scale model was tested. Any lab weighing this for its next pretraining run will want a replication well above 1 billion parameters first.
Reported from the paper posted on arXiv by its authors, submitted 29 September 2026.