THE SKYRL BLOG
Reinforcement learning for the open model stack.
Engineering notes, research updates, and practical guides from the team building SkyRL — a modular full-stack RL library for LLMs.

FEATURED · CASE STUDY
Training frontier knowledge work agents: A 397B RL training guide with SkyRL
A step-by-step guide to post-training Qwen3.5-397B-A17B with SkyRL, including the infrastructure, systems tuning, and lessons behind the hero run.
LATEST ARTICLES
18 storiesSkyRL v0: Train real-world long-horizon agents via reinforcement learning
Introducing SkyRL v0, an open framework for training capable long-horizon agents with reinforcement learning.
SkyRL SQL: Reinforcement learning with structured data
Training agents to reason over databases, structured data, and reliable tool interactions.
SkyRL v0.1: Building a stronger foundation for agent RL
The next step in making full-stack reinforcement learning research more modular and reproducible.
Search-R1: Training reasoning models to search
Exploring reinforcement learning for agents that learn when and how to search for information.
SkyRL DeepResearch: Long-horizon research agents
A look at the systems and training recipes behind capable research-oriented agents.
SkyRL TX: Scaling training for tool-using agents
New infrastructure for efficient reinforcement learning with tool use and external environments.
SkyRL TX v0.0.2: Iterating on agent training infrastructure
An incremental release focused on making agent training workflows easier to run and extend.
SkyRL TX 0.0.3: More reliable experiments
Updates to the TX stack for reproducible, production-minded reinforcement learning research.
SkyRL TX v0.1.0: A new milestone for tool-use RL
The first major TX release brings a more complete foundation for training tool-using agents.
SkyRL TX v0.2: Advancing scalable agent training
New capabilities for scaling experiments across environments, tools, and model families.
SkyRL TX Release — February 2026
A release update covering the latest progress in SkyRL TX and tool-using agent research.
SkyRL TX v0.3.0: Toward dependable agent RL
A new TX release with improvements for dependable, scalable training workflows.
SkyRL TX v0.2.1: Refining the training loop
A focused update that improves the day-to-day experience of running TX experiments.
SkyRL Tinker: Faster iteration for RL researchers
Making it easier to prototype, debug, and iterate on reinforcement learning ideas.
On-policy distillation for capable language agents
A practical look at distilling strong behaviors into models through on-policy training.
SkyRL Harbor: A home for agent environments
An environment layer for bringing richer tasks and real-world interaction into agent RL.
Improving GPU power utilization on agentic RL with async training
A case study on using asynchronous training to improve utilization in agentic reinforcement learning.
SkyRL v0.3: The open stack keeps moving
The latest SkyRL milestone and what comes next for open, full-stack reinforcement learning.
