<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Nick One-Drop]]></title><description><![CDATA[Nick One-Drop]]></description><link>https://nickonedrop.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6a9304f15a81010d6b32b818/3291cd81-e728-4efe-865f-2ba2e233922a.png</url><title>Nick One-Drop</title><link>https://nickonedrop.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Wed, 30 Sep 2026 08:59:32 GMT</lastBuildDate><atom:link href="https://nickonedrop.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Experiential Reinforcement Learning (2026)]]></title><description><![CDATA[The basic reinforcement learning loop is simple: try something, receive a reward, and repeat. Poor behavior is gradually corrected as reward signals accumulate across trials.
This paper proposes Exper]]></description><link>https://nickonedrop.hashnode.dev/experiential-reinforcement-learning-2026</link><guid isPermaLink="true">https://nickonedrop.hashnode.dev/experiential-reinforcement-learning-2026</guid><category><![CDATA[Machine Learning]]></category><category><![CDATA[DeepLearning]]></category><category><![CDATA[Reinforcement Learning]]></category><dc:creator><![CDATA[Nick One-Drop]]></dc:creator><pubDate>Sun, 27 Sep 2026 15:27:35 GMT</pubDate><content:encoded><![CDATA[<p>The basic reinforcement learning loop is simple: try something, receive a reward, and repeat. Poor behavior is gradually corrected as reward signals accumulate across trials.</p>
<p>This paper proposes <strong>Experiential Reinforcement Learning (ERL)</strong>, which adds a reflection step. When an attempt receives a poor reward, the model reflects on what went wrong and produces an adjustment, denoted by delta, to inform a second attempt. That second attempt is then evaluated through the usual reward mechanism.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/dd6a5006-a2fe-46e2-884d-5c0c7a306232.png" alt="" style="display:block;margin:0 auto" />

<p>The authors call this an <strong>experience–reflection–consolidation loop</strong>. Here is a more intuitive illustration of the process.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/7078d456-df63-47f8-a04a-9bdde82da74e.png" alt="" style="display:block;margin:0 auto" />

<p>The key question is how to perform reflection and obtain that adjustment. Since this paper focuses on reinforcement learning for LLMs, the approach is to ask the LLM itself to reflect.</p>
<p>In the figure below, “Feedback” refers to textual feedback from the environment. This setup therefore assumes an environment that can provide such feedback—for example, compiler messages in a code-generation task.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/ed327e39-2882-4225-ba60-b997ac8f72c6.png" alt="" style="display:block;margin:0 auto" />

<p>The formulation is shown below. Here, <strong>m</strong> represents reflection memory. Reflections generated during training are retained and supplied alongside subsequent inputs, allowing the model to reuse insights accumulated from earlier attempts.</p>
<p>A reflection is added to memory only when the associated reward exceeds a specified threshold.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/3c00f797-acf8-45bb-b6c3-5b39e4b7a4fd.png" alt="" style="display:block;margin:0 auto" />

<p>Another important goal is to preserve <strong>single-pass inference</strong>. Following the full training procedure at inference time would require an additional reflection step and a second attempt.</p>
<p>To avoid this, the authors use <strong>selective distillation</strong>: outputs from the second attempt are selectively distilled, guided by their rewards, into a model that can answer directly on the first attempt.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/05c87c98-4a3c-45d9-82c5-47f67edaa28a.png" alt="" style="display:block;margin:0 auto" />

<p>The approach makes intuitive sense to me, and the reported experimental results suggest that it is effective.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/84261e1f-65fa-4a1e-b860-698ceecb2473.png" alt="" style="display:block;margin:0 auto" />

<p>Could this approach work outside LLM training? I think the central question is whether we can construct a useful <strong>reflection policy</strong>. For tasks where such a policy can be designed, this seems like an approach worth exploring.</p>
<p>One further question occurred to me: if reflection improves the second reward relative to the first, it increases the proportion of successful attempts. Could that have side effects? Does learning need a balance of good and bad outcomes, and could reflection shift that balance too far?</p>
<p>My tentative interpretation is that this is less concerning in settings where positive rewards are sparse to begin with. The method also computes separate losses for the first and second attempts, so the original failures are not simply replaced by successful retries. A higher success rate alone does not necessarily imply a harmful imbalance; what matters is how those experiences contribute to learning.</p>
<hr />
<p><em>Paper: <a href="https://arxiv.org/abs/2602.13949">https://arxiv.org/abs/2602.13949</a></em></p>
]]></content:encoded></item><item><title><![CDATA[GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance (2026)]]></title><description><![CDATA[This paper explores using LLMs for game quality assurance. The evaluation uses 30 games that were themselves created by LLM agents. Together, these games form the GBQA benchmark.


The focus is on eva]]></description><link>https://nickonedrop.hashnode.dev/gbqa-a-game-benchmark-for-evaluating-llms-as-quality-assurance-2026</link><guid isPermaLink="true">https://nickonedrop.hashnode.dev/gbqa-a-game-benchmark-for-evaluating-llms-as-quality-assurance-2026</guid><category><![CDATA[Machine Learning]]></category><category><![CDATA[DeepLearning]]></category><category><![CDATA[Game Development]]></category><dc:creator><![CDATA[Nick One-Drop]]></dc:creator><pubDate>Fri, 25 Sep 2026 16:20:00 GMT</pubDate><content:encoded><![CDATA[<p>This paper explores using LLMs for game quality assurance. The evaluation uses 30 games that were themselves created by LLM agents. Together, these games form the <strong>GBQA</strong> benchmark.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/1dc03333-377f-4df2-9c54-05258ded0aeb.png" alt="" style="display:block;margin:0 auto" />

<p>The focus is on evaluating LLMs in the role of a <strong>QA agent</strong>.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/9d308278-1149-4160-9f6a-30130ab6de5d.png" alt="" style="display:block;margin:0 auto" />

<p>This does not involve agents watching gameplay through a vision-language model and interacting with what they see on screen. Instead, they work with structured game information, such as states and actions, alongside data about bug locations.</p>
<p>I would describe this as QA for game-state management and the consistency of game logic.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/0faf4e52-8572-4c55-8721-6e448b271bb6.png" alt="" style="display:block;margin:0 auto" />

<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/7a55f01e-e7f6-4369-962f-1a96624dcffb.png" alt="" style="display:block;margin:0 auto" />

<p>Visual elements such as game graphics are therefore outside the scope of this benchmark. Extending the evaluation to cover those aspects seems like a possible direction for future work.</p>
<p>The results for different LLMs are shown below. Without a human baseline, however, it is difficult to judge how these scores compare with human QA performance.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/a665a5e3-313f-42b4-98b0-bcb2ed3e771d.png" alt="" style="display:block;margin:0 auto" />

<p>My reading of the paper’s main message is that there is a substantial gap between writing code and finding bugs: the reported coding performance is around 80%, while bug-finding performance remains below 50%.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/54a124d0-e861-476d-91f4-de4f9987b287.png" alt="" style="display:block;margin:0 auto" />

<hr />
<p><em>Paper: <a href="https://arxiv.org/abs/2604.02648">https://arxiv.org/abs/2604.02648</a></em></p>
]]></content:encoded></item><item><title><![CDATA[Training an AlphaZero AI for a Unity Game: Separating Self-Play from Runtime]]></title><description><![CDATA[I recently released K-CHESS OMEGA, a Korean chess game whose AI was trained using an AlphaZero-style approach.
While building it, I ran into an interesting engineering problem: the environment used fo]]></description><link>https://nickonedrop.hashnode.dev/training-an-alphazero-ai-for-a-unity-game-separating-self-play-from-runtime</link><guid isPermaLink="true">https://nickonedrop.hashnode.dev/training-an-alphazero-ai-for-a-unity-game-separating-self-play-from-runtime</guid><category><![CDATA[DeepLearning]]></category><category><![CDATA[unity]]></category><category><![CDATA[Reinforcement Learning]]></category><dc:creator><![CDATA[Nick One-Drop]]></dc:creator><pubDate>Sun, 06 Sep 2026 18:32:16 GMT</pubDate><content:encoded><![CDATA[<p>I recently released <strong>K-CHESS OMEGA</strong>, a Korean chess game whose AI was trained using an AlphaZero-style approach.</p>
<p>While building it, I ran into an interesting engineering problem: the environment used for large-scale self-play and the environment used in the actual Unity game had very different requirements.</p>
<p>Solving that problem eventually led me to maintain the same game and search logic in both C# and Python—something I would normally have tried to avoid.</p>
<p>Coding AI made that trade-off much more practical than I expected.</p>
<h2>Starting with Unity and ML-Agents</h2>
<p>The game state, actions, rules, and MCTS logic were initially written in C#, since they ultimately had to run inside the Unity game.</p>
<p>For training, self-play ran in Python in a distributed environment, with the Unity environment connected through ML-Agents.</p>
<p>Conceptually, the initial setup looked like this:</p>
<p><strong>Python self-play / MCTS ↔ ML-Agents ↔ Unity / C# game environment</strong></p>
<p>Functionally, this worked well.</p>
<p>The problem was speed.</p>
<p>AlphaZero-style MCTS performs hundreds of simulations to decide a single move. During those simulations, the search repeatedly interacts with the game environment to apply actions, inspect states, and continue the search.</p>
<p>With Unity and the Python self-play process repeatedly communicating through ML-Agents, the communication overhead accumulated quickly.</p>
<p>For the amount of self-play I wanted to generate, I needed a different approach.</p>
<h2>Moving the self-play environment to Python</h2>
<p>I started to understand why some AlphaZero implementations go as far as implementing the game environment and MCTS in C++.</p>
<p>I didn't go that far.</p>
<p>Instead, I ported both the C# game environment and the relevant MCTS logic to Python.</p>
<p>The training side then became fully Python-based:</p>
<p><strong>Python game environment ↔ Python MCTS / self-play → PyTorch training → ONNX</strong></p>
<p>Unity was no longer involved in the inner self-play loop.</p>
<p>This removed the Unity/ML-Agents communication overhead from self-play while keeping the production game itself in Unity.</p>
<img src="https://cdn.hashnode.com/uploads/covers/6a9304f15a81010d6b32b818/3d93fc79-58ea-4965-a894-d74b05c1ca97.jpg" alt="" style="display:block;margin:0 auto" />

<h2>The downside: two implementations of the same logic</h2>
<p>There was, of course, an obvious downside.</p>
<p>I now had the same game rules and search behavior implemented in two languages.</p>
<p><strong>C#:</strong></p>
<ul>
<li><p>Production game environment in Unity</p>
</li>
<li><p>Runtime MCTS</p>
</li>
</ul>
<p><strong>Python:</strong></p>
<ul>
<li><p>Training game environment</p>
</li>
<li><p>Self-play MCTS</p>
</li>
</ul>
<p>Normally, I would be reluctant to choose this architecture.</p>
<p>Every rule change has to be reflected correctly in both implementations. Even a subtle difference between the Python training environment and the C# production environment could mean that the model is trained on a game that behaves differently from the one the player actually plays.</p>
<p>The same concern applies to the MCTS implementation.</p>
<p>This kind of duplicate implementation creates a real maintenance cost and a risk of the two versions diverging.</p>
<h2>Coding AI changed the trade-off</h2>
<p>This was where coding AI turned out to be particularly useful.</p>
<p>Rather than rewriting everything manually, I used coding AI to port the existing C# implementation to Python.</p>
<p>More importantly, as the C# implementation evolved, I used it to keep the corresponding Python implementation in sync.</p>
<p>For this kind of task—translating already-existing logic between languages while preserving behavior—current coding models worked extremely well for me.</p>
<p>That changed the trade-off.</p>
<p>Maintaining two implementations was still something I needed to be careful about, but the cost was low enough that an architecture I might previously have avoided became a practical option.</p>
<p>This was one of the more interesting lessons from the project for me: coding AI wasn't only helping me write code faster. It affected which software architecture I was willing to choose.</p>
<h2>Bringing the trained model back into Unity</h2>
<p>Once trained, the PyTorch model was exported to ONNX.</p>
<p>At runtime, Unity Sentis runs the ONNX model for policy and value inference, while the C# implementation of MCTS performs the actual search.</p>
<p>So the final architecture became:</p>
<p><strong>Training</strong></p>
<p>Python game environment<br />↔ Python MCTS / self-play<br />→ PyTorch<br />→ ONNX</p>
<p><strong>Runtime</strong></p>
<p>Unity / C# game environment<br />↔ C# MCTS<br />↔ Unity Sentis / ONNX inference</p>
<p>The expensive self-play process could now run entirely outside Unity, while the trained model could still be integrated directly into the production game.</p>
<h2>Final thoughts</h2>
<p>What started as a straightforward question—how to train an AlphaZero-style AI for a Unity game—eventually became a question of where the boundary between training and runtime should be.</p>
<p>For this project, separating them worked well: Python provided a self-contained environment for self-play and training, while Unity remained the production environment where the trained model and MCTS ultimately had to run.</p>
<p>The unexpected part was how practical it became to maintain the corresponding C# and Python implementations with the help of coding AI.</p>
<p>K-CHESS OMEGA ended up combining Unity, Sentis, PyTorch, ONNX, and coding AI in one system.</p>
<p>It was a fun engineering experience, and I thought the architecture and the trade-offs behind it might be useful to other developers working with game AI or reinforcement learning in Unity.</p>
]]></content:encoded></item><item><title><![CDATA[Does Socialization Emerge in AI Agent Society?]]></title><description><![CDATA[I recently read Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook.
The paper asks a simple question: if millions of AI agents continuously interact, will something like human soc]]></description><link>https://nickonedrop.hashnode.dev/does-socialization-emerge-in-ai-agent-society</link><guid isPermaLink="true">https://nickonedrop.hashnode.dev/does-socialization-emerge-in-ai-agent-society</guid><dc:creator><![CDATA[Nick One-Drop]]></dc:creator><pubDate>Sat, 29 Aug 2026 16:24:21 GMT</pubDate><content:encoded><![CDATA[<p>I recently read <em>Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook</em>.</p>
<p>The paper asks a simple question: if millions of AI agents continuously interact, will something like human society naturally emerge?</p>
<p>Its answer is: <strong>not necessarily.</strong></p>
<h2>What the study found</h2>
<p>The study finds some convergence in vocabulary and global semantic patterns.</p>
<p>But individual agents do not meaningfully converge, adapt to one another, or form lasting influence hierarchies.</p>
<p>One thing I wondered about is whether global averages are the right way to measure this. Socialization may happen within smaller communities without causing much change in the population-wide average.</p>
<h2>A systems question</h2>
<p>There is also a systems question.</p>
<p>Why would agents socialize without persistent social memory, repeated relationships, or mechanisms that let interactions shape future behavior?</p>
<p>Perhaps scale and communication are not enough.</p>
<p>Society may require memory, attachment, incentives—and possibly something analogous to emotion.</p>
<hr />
<p>Paper:<br /><em>Does Socialization Emerge in AI Agent Society? A Case Study of Moltbook</em><br /><a href="https://arxiv.org/abs/2602.14299">https://arxiv.org/abs/2602.14299</a></p>
]]></content:encoded></item></channel></rss>