Yandex Research scientists have proposed a method to accelerate text generation using diffusion language models. The new method helps such neural networks perform more computations in parallel without losing accuracy in their responses. In experiments with models having 16 billion parameters, the quality on some tasks even improved.

Conventional language models, including most popular chatbots, generate text sequentially — first selecting one token, then the next, and so on. Diffusion models are structured differently. They start with an incomplete or noisy sequence and gradually refine it, working with several text positions at once in a single step. This allows them to generate text in parallel.

However, achieving good quality with this generation method is not easy. For this, the model is additionally trained after the main training. Such methods can usually require complex computations or additional models.

The Yandex Research team proposed a simpler approach. It compares texts generated by the neural network with real examples from the training data. The Maximum Mean Discrepancy (MMD) mathematical method is used for comparison. It assesses how much two datasets differ.

The comparison is not made directly by words, but by internal features extracted by an already trained language model. Its parameters remain unchanged. This allows it to be used as a kind of measuring tool, without training another large model specifically to evaluate the results.

The authors tested the approach on two types of diffusion models — those working with discrete tokens and those with continuous data representations.

What the experiments showed

On texts from the OpenWebText dataset, the new method helped the model better reproduce patterns of human speech while maintaining the diversity of responses. On GSM8K mathematical tasks, the model began to find a more successful balance between the accuracy of answers and the number of computational steps. The fewer steps needed to get the correct answer, the faster the model can potentially cope with the task.

The method was also tested on models with 16 billion parameters. They were additionally trained on eight NVIDIA H100 accelerators, and, according to researchers, this took about 1.3–1.9 minutes. In four mathematical tests, the models processed 10–16% more parts of the text in one computational step, maintaining the same accuracy or slightly improving it. In programming tasks, the accuracy of answers increased by 2.4–3.8 percentage points.

Diffusion neural networks can change the way text is generated. However, they are currently less common than neural networks that sequentially predict the next token.

Read more on the topic:

Never miss our newsAdd this site to your preferred sources to see us more oftenAdd on Google

Сейчас на главной