Experiential Reinforcement Learning (2026)
The basic reinforcement learning loop is simple: try something, receive a reward, and repeat. Poor behavior is gradually corrected as reward signals accumulate across trials.
This paper proposes Experiential Reinforcement Learning (ERL), which adds a reflection step. When an attempt receives a poor reward, the model reflects on what went wrong and produces an adjustment, denoted by delta, to inform a second attempt. That second attempt is then evaluated through the usual reward mechanism.
The authors call this an experience–reflection–consolidation loop. Here is a more intuitive illustration of the process.
The key question is how to perform reflection and obtain that adjustment. Since this paper focuses on reinforcement learning for LLMs, the approach is to ask the LLM itself to reflect.
In the figure below, “Feedback” refers to textual feedback from the environment. This setup therefore assumes an environment that can provide such feedback—for example, compiler messages in a code-generation task.
The formulation is shown below. Here, m represents reflection memory. Reflections generated during training are retained and supplied alongside subsequent inputs, allowing the model to reuse insights accumulated from earlier attempts.
A reflection is added to memory only when the associated reward exceeds a specified threshold.
Another important goal is to preserve single-pass inference. Following the full training procedure at inference time would require an additional reflection step and a second attempt.
To avoid this, the authors use selective distillation: outputs from the second attempt are selectively distilled, guided by their rewards, into a model that can answer directly on the first attempt.
The approach makes intuitive sense to me, and the reported experimental results suggest that it is effective.
Could this approach work outside LLM training? I think the central question is whether we can construct a useful reflection policy. For tasks where such a policy can be designed, this seems like an approach worth exploring.
One further question occurred to me: if reflection improves the second reward relative to the first, it increases the proportion of successful attempts. Could that have side effects? Does learning need a balance of good and bad outcomes, and could reflection shift that balance too far?
My tentative interpretation is that this is less concerning in settings where positive rewards are sparse to begin with. The method also computes separate losses for the first and second attempts, so the original failures are not simply replaced by successful retries. A higher success rate alone does not necessarily imply a harmful imbalance; what matters is how those experiences contribute to learning.
