
Yandex proposes a new method to accelerate language models without expensive fine-tuning
The method allows processing more operations simultaneously while maintaining the quality of generated text

The method allows processing more operations simultaneously while maintaining the quality of generated text