Discussion about this post

User's avatar
AutomationLabs's avatar

The critic-based PPO section is the part I keep coming back to — everyone adopted GRPO for the memory savings, but uniform trajectory-level advantage really does fall apart once compaction splits a trajectory into uneven sub-traces. The compaction-aware detail (training the critic on the lossy summaries agents actually see in production, not pristine histories) is exactly the train/serve alignment most teams skip. Good to see the systems-engineering-beats-single-techniques point made with real numbers rather than vibes.

No posts

Ready for more?