Skip to content

TheLLM Brief

← All stories

Research

SFT vs RL Is the Wrong Question in Post-Training

A new arXiv paper reframes the debate around 'accessible support', not training method.

Sourced from arXiv.org

A paper posted to arXiv reframes how the field should think about post-training. The central claim: supervised fine-tuning and reinforcement learning both reweight a pretrained model's reference distribution. The real question is whether that reweighting stays inside behaviors the model could already reach, or expands that reachable set.

The authors introduce the concept of 'accessible support', the set of behaviors a model can practically produce under finite compute budgets. Training that shifts probability mass within that support is capability elicitation. Training that changes the support itself is capability creation. The free-energy framing treats demonstration signals and reward signals as two versions of the same mechanism, not fundamentally different operations.

This matters for anyone building on top of foundation models. Labs, operators, and evaluators have used SFT versus RL as a proxy for 'imitation versus discovery.' That proxy is now contested. The paper argues search, interaction, tool use, and new information are the real levers for capability creation. Watch how safety evaluators and post-training teams update their benchmarks in response.

Analysis

Elicitation versus creation is the contrast that now governs post-training investment decisions. Labs that conflate the two will misread what their fine-tuning actually buys.

Research this with your AI

Copy the research prompt into your AI assistant to see how this story affects you.

Then paste it into ChatGPT, Claude, Gemini, Grok and others.
Runs in your own assistant with your own context. Nothing is sent to us.
Show the prompt
I just read this AI news story and want to understand it in my own context.

Title: SFT vs RL Is the Wrong Question in Post-Training
Summary: Researchers argue post-training debates misplace the key distinction. What matters is whether training expands a model's reachable behaviors, not whether the method is SFT or RL.
Category: Research
Source: arXiv.org, https://arxiv.org/abs/2605.08368

Using my own history and context, help me understand:
1. What is the core development and why does it matter?
2. Who are the major players involved and what are their motivations?
3. How does this fit into the broader AI landscape right now?
4. How does this apply to my own work, and what should I do or watch next?

Be specific and plain spoken.

Newsletter

The day's AI stories, with the editor's take, in one email.

Free. Unsubscribe in one click.