SyncAI.news, a Varaisys broadcasting
Introducing RWKV - An RNN with the advantages of a transformer
HF

Hugging Face Blog

· 1 min read

AI LabsHugging Face Blog

Introducing RWKV - An RNN with the advantages of a transformer

ChatGPT and chatbot-powered applications have captured significant attention in the Natural Language Processing (NLP) domain. The community is constantly seeking strong, reliable and open-source models for their applications and use cases. The rise of these powerful models stems from the democratization and widespread adoption of transformer-based models, first introduced by Vaswani et al. in 2017. These models significantly outperformed previous SoTA NLP models based on Recurrent Neural Networks (RNNs), which were considered dead after that paper. Through this blogpost, we will introduce the integration of a new architecture, RWKV, that combines the advantages of both RNNs and transformers, and that has been recently integrated into the Hugging Face transformers library.

Overview of the RWKV project

The RWKV project was kicked off and is being led by Bo Peng, who is actively contributing and maintaining the project. The community, organized in the official discord channel, is constantly enhancing the project’s artifacts on various topics such as performance (RWKV.cpp, quantization, etc.), scalability (dataset processing & scrapping) and research (chat-fine tuning, multi-modal finetuning, etc.). The GPUs for training RWKV models are donated by Stability AI.

You can get involved by joining the official discord channel and learn more about the general ideas behind RWKV in these two blogposts: https://johanwind.github.io/2023/03/23/rwkv_overview.html / https://johanwind.github.io/2023/03/23/rwkv_details.html

Transformer Architecture vs RNNs

Overview of possible configurations of using RNNs. Source: Andrej Karpathy's blogpost
Formulation of attention scores in transformer models. Source: Jay Alammar's blogpost
Formulation of attention scores in RWKV models. Source: RWKV blogpost

The RWKV architecture

RWKV as a combination of RNNs and transformers

The major drawbacks of traditional RNN models and how RWKV is different:

RWKV attention formulation

Existing checkpoints

| |

Original source

This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on huggingface.co

Similar News