Kimi K2.x: 1T-param RL with INT4 rollouts + LoRA
Requires the megatron backend. The reference run uses two 8-GPU nodes
(verified on 2 x 8 B300) at the full 262,144-token context.
This example trains moonshotai/Kimi-K2.7-Code — a 1T-parameter
DeepSeek-V3-architecture MoE (61 layers, 384 routed experts, MLA) released as
a KimiK25ForConditionalGeneration checkpoint whose expert weights are
INT4 QAT-quantized (compressed-tensors pack-quantized, group_size=32,
Kimi convention scale_divisor=7.0 / q_min=-7) — with GRPO (DAPO-style),
LoRA, and bit-exact fake-INT4 quantization-aware training.
How the pieces fit together
- vLLM serves the release INT4 checkpoint as-is (compressed-tensors
W4A16, ~595 GB), one engine per node, with the LoRA adapter hot-loaded from
disk after every optimizer step (
merge_lora=false). The INT4 base weights in the engine never change. - The Megatron trainer holds BF16 "master" weights dequantized from the
INT4 release, with LoRA adapters on
linear_proj(attention out) andlinear_fc1/linear_fc2(dense MLP, shared expert, and all routed experts — the routed experts carry ~96% of the parameters). - Fake-INT4 QAT (
trainer.policy.model.fake_int4_qat) fake-quantizes the frozen expert GEMMs onto the same INT4 grid in the trainer's forward pass (straight-through backward). For Kimi QAT checkpoints the dequantized masters are a fixed point of this quantizer, so the trainer's experts are bit-exact with what vLLM serves; TIS corrects the residual (LoRA-off-policy) mismatch. - Text-only training (
language_model_only=trueon both the trainer and the inference engine): the vision tower stays frozen inside vLLM and is never built on the Megatron side. - Capacity-normalized expert LoRA
(
trainer.policy.megatron_config.lora_config.normalize_moe_lora=true): the 384 routed experts get rankrank // moe_router_topkadapters (non-expert layers keep the full rank) so the per-step PEFT export stays small enough to write and hot-load on every engine.
One-time setup
Every node needs the repo checkout, the project venv, the HF cache and the BF16 masters at identical paths.
0) Install the megatron stack into the project venv. The root uv.lock
is CUDA-13 native (torch 2.13 cu130, vLLM 0.28, Transformer Engine 2.16; every
CUDA library, cuDNN included, comes from pip), which is what Blackwell-Ultra
(B300 / sm103) requires — NVIDIA driver R580 or newer:
uv sync --extra megatron1) Dequantize the INT4 release to BF16 masters on every node (or a shared filesystem). This also verifies, tensor by tensor, that re-quantizing the masters reproduces the released INT4 grid bit-for-bit:
uv run --isolated examples/train/megatron/dequantize_compressed_tensors_int4.py \
--input <path-to-Kimi-K2.7-Code-snapshot> \
--output /data/skyrl/models/Kimi-K2.7-Code-BF16(~595 GB in, ~2.1 TB out. For RTN-quantized checkpoints — scale_divisor=7.5
— use the original BF16 release as masters instead; a dequantized RTN grid is
not its own fixed point and the script will fail loudly if you try.)
2) Download the math data on the head node (any dataset/environment works; the reference script reuses the standard DAPO prep):
bash examples/train/algorithms/dapo/prepare_dapo_data.sh3) Start Ray on both nodes from the checkout, then run the training script
on the head node only. The environment block at the bottom of the run script
(LD_LIBRARY_PATH / CUDNN_PATH pointing into the venv's pip CUDA libraries,
SKYRL_LD_LIBRARY_PATH_EXPORT=1, UV_PROJECT_ENVIRONMENT) plus any
site-specific cross-node NCCL settings (NCCL_SOCKET_IFNAME, NCCL_IB_HCA,
GLOO_SOCKET_IFNAME) must be exported in the shell that runs ray start on
every node: Ray workers inherit the raylet's environment, and only
LD_LIBRARY_PATH and the SKYRL_* / UV_* / HF_* variables are re-exported
through the Ray runtime env.
export RAY_RUNTIME_ENV_HOOK=ray._private.runtime_env.uv_runtime_env_hook.hook
# head
uv run --no-sync --extra megatron ray start --head --port=6379 --num-gpus=8
# worker
uv run --no-sync --extra megatron ray start --address=<head-ip>:6379 --num-gpus=8Running
bash examples/train/megatron/run_megatron_dapo_kimi_k2.7_code_lora_int4_qat.shThe key configuration (2 nodes x 8 GPUs, colocated):
# vLLM: one engine per node serving the INT4 release, TP=8
NUM_INFERENCE_ENGINES=2
INFERENCE_ENGINE_TENSOR_PARALLEL_SIZE=8
# Megatron: TP=2 x CP=8 (DP=1), experts sharded EP=16 (24 experts/GPU/layer,
# ~127 GB of frozen BF16 masters + ~12 GB non-expert per GPU)
MEGATRON_TP=2 MEGATRON_CP=8 MEGATRON_EP=16 MEGATRON_ETP=1 MEGATRON_PP=1
# Full 262,144-token context
MAX_PROMPT_LENGTH=$((1024 * 8))
MAX_RESPONSE_LENGTH=$((262144 - MAX_PROMPT_LENGTH))
# INT4-served actor: trainer loads BF16 masters, fake-quantizes experts
trainer.policy.model.path="moonshotai/Kimi-K2.7-Code"
trainer.policy.model.fake_int4_qat.enabled=true
trainer.policy.model.fake_int4_qat.scale_divisor=7.0 # Kimi QAT convention
trainer.policy.model.fake_int4_qat.q_min=-7
trainer.policy.model.fake_int4_qat.bf16_base_path=/data/skyrl/models/Kimi-K2.7-Code-BF16
trainer.policy.megatron_config.lora_config.merge_lora=false
trainer.policy.megatron_config.lora_config.normalize_moe_lora=true
trainer.policy.language_model_only=true
generator.inference_engine.language_model_only=true
trainer.flash_attn=false # MLA + sample packing needs TE cuDNN fused attentionThe script reads its scale knobs from the environment, so a quick end-to-end check of a fresh cluster is a matter of shrinking them (this is how the reference setup is smoke-tested before a real run):
LOGGER=console MAX_RESPONSE_LENGTH=2048 TRAIN_BATCH_SIZE=8 MINI_BATCH_SIZE=8 \
N_SAMPLES_PER_PROMPT=4 \
bash examples/train/megatron/run_megatron_dapo_kimi_k2.7_code_lora_int4_qat.sh \
trainer.max_training_steps=2QAT_MODE=off bash ... runs the ablation: the INT4 engine stays, but the
trainer's fake-quant and TIS are disabled, exposing the uncorrected
train/infer mismatch (same toggle as the Qwen3.6 INT4 study).
Notes
- Multi-node adapter sync: with
merge_lora=falsethe adapter files are written once per node (each vLLM worker readslora_sync_pathfrom its local filesystem), so no shared filesystem is required. - Frozen-master offload: under
colocate_allthe trainer's frozen BF16 masters (~1.1 TB per node) are offloaded while the engines generate. They go to file-backed mmaps underSKYRL_FROZEN_OFFLOAD_DIR(default/data/skyrl/frozen-offload; clean, evictable page cache), so point it at a fast local disk on every node — or set it to0to keep the offload in pinned host RAM. - Blackwell-Ultra (B300 / sm103) runs on the default CUDA-13 lock — no
separate dependency set is needed anymore. On hosts that also carry a
system CUDA/cuDNN, the venv's pip libraries must come first on
LD_LIBRARY_PATHandCUDNN_PATHmust point at the pip cuDNN (the run script exports both), or Transformer Engine mixes the systemlibcudnncore with the pip sublibraries (CUDNN_STATUS_SUBLIBRARY_LOADING_FAILED). A system cuDNN 9 of a different 9.x than the pip one is still a problem even then: cuDNN's graph library probes optional sublibraries (libcudnn_engines_tensor_ir) that the pip wheel pinned by torch may not ship, the dynamic loader falls through to the system copy, and TE's fused attention fails withCUDNN_STATUS_SUBLIBRARY_VERSION_MISMATCH. Remove the system cuDNN or install a pipnvidia-cudnn-cu13of the same version. - Engine bring-up time: the 595 GB INT4 load plus
FULL_DECODE_ONLYCUDA-graph capture can exceed SkyRL's default 600 s engine health deadline on a cold page cache; the script raisesSKYRL_WAIT_UNTIL_INFERENCE_SERVER_HEALTHY_TIMEOUT_SandVLLM_EXECUTE_MODEL_TIMEOUT_SECONDSaccordingly. - MLA prefill backend: the script pins vLLM's
attention_config.mla_prefill_backend=FLASHINFER; the default FA4 CuTe-DSL prefill is incompatible with the lockednvidia-cutlass-dsl. - The LoRA targets deliberately exclude the MLA q/kv projections: mcore names
them differently (
linear_q_down_proj, ...) and vLLM's MLA weight absorption does not apply adapters tokv_b_proj. - Checkpoints save the LoRA adapters + optimizer state only (the frozen base
is reloaded from
bf16_base_pathon resume); exported adapters reference the INT4 release asbase_model_name_or_path, so they can be served directly against the public checkpoint.