RELEASESkyRL v0.4.0 · 2026

SkyRL v0.4.0 Release

Overview

In this SkyRL release we focused on improving SkyRL for large scale post-training, with the following key new features:

  • Model Support: Support for the following new models:
  • GLM 5.3 Flash
  • GLM 5.2/5.3
  • Kimi K2.6/2.7
  • Qwen3.8
  • Nemotron 3.5-Lightning        
  • RL stability
  • Top-P Sampler Replay
  • Optimized R3 data transfer
  • Full FP8/MXFP8 RL
  • Weight Syncing:
  • Delta Weight Sync
  • Sharded Weight Transfer with Ray Direct Transfer (RDT) + NIXL
  • Scalability
  • Large Scale SFT
  • Tinker Server Scalability

New Model Support

GLM 5.3-Flash and GLM 5.3

SkyRL now supports GLM 5.3 Flash and GLM 5.3 training. We show stable LoRA and full finetuning training with low logprob diff on DAPO with rollout router replay (R3). This enablement effort was in collaboration with Trajectory - more details can be found in their blog post: https://www.trajectory.ai/field-notes/enabling-frontier-scale-training-for-the-glm-5-3-family.

PRs:

Kimi K2.6/2.7

SkyRL now has experimental support for training Kimi 2.6/2.7. Training large 1T+ parameter models is often expensive, requiring 10TB+ of GPU memory even for simple single turn RL tasks. SkyRL now supports efficient training for Kimi 2.6/2.7 models with LoRA, allowing users to train a 1T parameter model on just 2 B300 nodes (see example script here)! We use INT4 serving for inference and apply fake-quantization on the trainer to minimize training inference mismatch.

Thanks to @casper_hansen for the contribution!

PRs:

Faster R3 + Top-P Sampler Replay

This release features major improvements for stable RL training with SkyRL: support for top-p sampler replay, as well as throughput improvements for rollout router replay (R3) for large MoE models.

These features were contributed by Dominic Yurk from Lila Sciences. With both features turned on for training Nemotron 3 Super on Lila Sciences’ tasks, the logprob gap stays under 0.1 and reward is still climbing at step 160. Without them, the gap climbs past 0.8 and reward collapses.

Top-P Sampler Replay

SkyRL now supports top-p and top-k sampling during RL, using the Keep Sampling Mask approach from DeepSeek-V3.2. With generator.inference_engine.enable_return_sample_support_set=true, vLLM returns the candidates that survived filtering for each generated token. With trainer.algorithm.enable_sample_support_replay=true, the trainer normalizes each token's logprob over that same set. The support travels through the same packed path as the routes, and the trainer scores it without materializing full-vocabulary logits. It works on both Megatron and FSDP.

Bounded sampling helps keep entropy from collapsing over long runs. On GLM-4.7-Flash with Nemotron-Terminal-Synthetic-Tasks, reward is unchanged with sampler replay, entropy settles around 0.14 instead of 0.05, and Pass@8 improves from 0.631 to 0.641. Example scripts are here.

PRs: #2237, #2238, #2239, #2240, #2241

Rollout Routing Replay (R3)

SkyRL already supported rollout routing replay (R3), but the routing data gets large at scale: one expert ID per token, per MoE layer, per top-k slot. For Nemotron 3 Super (120B), that's nearly 900× the size of the tokens themselves, and the original implementation made each trainer step more than 6× slower.

Most of that cost was CPU time spent serializing, parsing, and copying routes, as well as spillover to disk from the ray object store. In v0.4.0, routes stay in compact packed arrays from vLLM to the trainer, large objects are transferred through ray object store via zero-copy views, and multi-turn generators only send the routes for each new turn. With these improvements, throughput for training Nemotron 3 Super with R3 goes from 111 to 483 tokens/s/GPU. 

This release also fixes several multi-turn correctness bugs, including one where replay could silently use stale routes. To enable R3 (Megatron only), set generator.inference_engine.enable_return_routed_experts=true and trainer.policy.megatron_config.moe_enable_routing_replay=true.

PRs: #2232, #2233, #2234, #2235, #2236

Full FP8 + MXFP8 Training

SkyRL now supports Blockwise FP8 and MXFP8 accelerated training and rollout with on-policy weight sync. In long-run RL experiments, the FP8 configurations closely track BF16 convergence while reducing end-to-end step time by up to 23%. More details can be found in our blog post from @jinghanyao1-hub.

PRs:

  • Full FP8 weight sync: #1898
  • MXFP8 weight sync: #2072

Weight Syncing

Weight sync was reworked substantially in v0.4.0. SkyRL now uses vLLM’s new abstractions that decouple trainer send logic and inference receive logic. The result is an easier path for customizing weight transfer for RL frameworks built on top of vLLM.

Delta Weight Syncing

RL training on disaggregated clusters allows users to better utilize available resources in disparate regions , different cloud providers or different hardware. A key bottleneck for disaggregated RL is weight transfer between inference and trainer nodes. During RL training, most of the weights are typically not updated in a given training step (only ~1-3% are updated). Delta weight sync enables efficient disaggregated RL over ethernet. The trainer computes byte-exact XOR deltas against the previously published BF16 checkpoint-format weights, compresses the deltas, publishes a manifest, and asks inference workers to fetch/process the next version before pausing generation. During the weight update, inference workers load the prepared weights from disk.

The implementation supports disk, S3, and GCS-style transfer paths and overlaps receiver-side fetch with generation to reduce idle time in fully async RL setups.

PR: Add Delta Weight Sync Support in SkyRL

Sharded RDT (NIXL)

SkyRL now includes an efficient sharded weight transfer backend that leverages NIXL via Ray Direct Transport (RDT). This is useful for very large MoE runs where simple NCCL broadcast-based transfer becomes a bottleneck. With NCCL, trainer rank 0 forms a collective communication group with all the inference ranks and transfers full weights via broadcast. With the sharded weight transfer engine, we utilize all trainer ranks in the transfer and only send the shard that is needed. The transfer is further optimized to avoid gathering weights across PP ranks, and skips gathering expert layers. The sharded RDT backend is able to achieve weight transfer for the Kimi K2 model in BF16 in 7.53 seconds on 48 8xH100 nodes (32 nodes for the trainer, 16 for inference). We share more details in our blog post.

PR:
Sharded RDT weight sync

LoRA In-Memory Weight Syncing

SkyRL now supports in-memory weight syncing with LoRA training. Typically, weight syncing with LoRA in vLLM is implemented with disk based transfer: trainer saves updated LoRA adapters to disk, and the inference ranks re-load the updated adapters into the adapter cache from disk. We implemented simple in-memory weight syncing for LoRA-based training where LoRA weights are transferred via NCCL/CUDA IPC to the inference ranks, which update the active adapter directly on GPU memory. We will be upstreaming our vLLM-side changes in the coming weeks.

PR: In-memory LoRA adapter sync for Megatron

Large Scale SFT

The native SFTTrainer now supports supervised finetraining on very large (10B-1T+ tokens) datasets. We support loading directly from pretokenized datasets on disk for users who run large-scale, complex data processing pipelines. Datasets are loaded in a memory-mapped fashion for memory-efficient data loading. In v0.4.0, SFT training also supports custom sampling schemes, and mixing multiple datasets, allowing researchers to easily experiment with different data mixtures.

Tinker Server Improvements

In SkyRL v0.3.0, the Tinker API server struggled to support over 1,024 concurrent sample requests. Every sample went through a small SQLite connection pool shared with training futures and heartbeats, and the pool ran out quickly. In v0.4.0, the same server handles 131,072 concurrent trajectories with 212k-token outputs. The throughput improvements came from some small and some major fixes. The first fix was keeping forwarded sample futures and their results in memory, so sampling never touches SQLite (thanks to Trajectory AI!). We then fixed the API server’s bottlenecks one by one: replacing httpx with aiohttp, keeping vLLM completion response outputs in protobuf, and reducing unwanted logging were some of the improvements implemented.

PRs: #2097, #2160, #2161, #2162, #2163, #2164

Github Release

Read the full SkyRL v0.4.0 release notes and access the latest code on GitHub.

View SkyRL v0.4.0 on GitHub