THE SKYRL BLOG

Reinforcement learning for the open model stack.

Engineering notes, research updates, and practical guides from the team building SkyRL — a modular full-stack RL library for LLMs.

SkyRL architecture diagram showing the VeRL backend, tools, environment runtime, and sandbox server

FEATURED · CASE STUDY

A step-by-step guide to post-training Qwen3.5-397B-A17B with SkyRL, including the infrastructure, systems tuning, and lessons behind the hero run.

·Mercor + SkyRL

LATEST ARTICLES

18 stories
RESEARCH

SkyRL v0: Train real-world long-horizon agents via reinforcement learning

Introducing SkyRL v0, an open framework for training capable long-horizon agents with reinforcement learning.

RESEARCH

SkyRL SQL: Reinforcement learning with structured data

Training agents to reason over databases, structured data, and reliable tool interactions.

RELEASE

SkyRL v0.1: Building a stronger foundation for agent RL

The next step in making full-stack reinforcement learning research more modular and reproducible.

RESEARCH

Search-R1: Training reasoning models to search

Exploring reinforcement learning for agents that learn when and how to search for information.

RESEARCH

SkyRL DeepResearch: Long-horizon research agents

A look at the systems and training recipes behind capable research-oriented agents.

ENGINEERING

SkyRL TX: Scaling training for tool-using agents

New infrastructure for efficient reinforcement learning with tool use and external environments.

RELEASE

SkyRL TX v0.0.2: Iterating on agent training infrastructure

An incremental release focused on making agent training workflows easier to run and extend.

RELEASE

SkyRL TX 0.0.3: More reliable experiments

Updates to the TX stack for reproducible, production-minded reinforcement learning research.

RELEASE

SkyRL TX v0.1.0: A new milestone for tool-use RL

The first major TX release brings a more complete foundation for training tool-using agents.

RELEASE

SkyRL TX v0.2: Advancing scalable agent training

New capabilities for scaling experiments across environments, tools, and model families.

RELEASE

SkyRL TX Release — February 2026

A release update covering the latest progress in SkyRL TX and tool-using agent research.

RELEASE

SkyRL TX v0.3.0: Toward dependable agent RL

A new TX release with improvements for dependable, scalable training workflows.

RELEASE

SkyRL TX v0.2.1: Refining the training loop

A focused update that improves the day-to-day experience of running TX experiments.

ENGINEERING

SkyRL Tinker: Faster iteration for RL researchers

Making it easier to prototype, debug, and iterate on reinforcement learning ideas.

RESEARCH

On-policy distillation for capable language agents

A practical look at distilling strong behaviors into models through on-policy training.

ENGINEERING

SkyRL Harbor: A home for agent environments

An environment layer for bringing richer tasks and real-world interaction into agent RL.

CASE STUDY

Improving GPU power utilization on agentic RL with async training

A case study on using asynchronous training to improve utilization in agentic reinforcement learning.

RELEASE

SkyRL v0.3: The open stack keeps moving

The latest SkyRL milestone and what comes next for open, full-stack reinforcement learning.