Reinforcement Learning (RL) has become one of the core ingredients of modern Large Language Model (LLM) post-training. While pre-training teaches a model to predict the next token from massive text corpora, …
Large Language Models (LLMs) are everywhere and are substantially changing the way we perform many daily tasks. Recently, I noticed a growing number of positions related to the post-training of such models. Interestingly, two …