SkyRL
Tinker API

Configuration

This page describes how to configure the SkyRL Tinker backend, including GPU allocation, training parameters, and inference settings.

When spinning up the Tinker server, the --backend-config flag accepts a JSON dictionary of dot-notation overrides that are applied to the underlying SkyRL-Train configuration. For example:

uv run --extra tinker --extra fsdp -m skyrl.tinker.api \
    --base-model "Qwen/Qwen3-4B-Instruct-2507" --backend fsdp \
    --backend-config '{"trainer.placement.policy_num_gpus_per_node": 4, "generator.inference_engine.num_engines": 4}'

Any field in the SkyRL-Train config can be overridden this way (see SkyRLTrainConfig in skyrl/train/config/config.py for all available keys and defaults). The most commonly used options are listed below.

JSON types are preserved exactly, so use JSON literals rather than strings: null for an unset field, true/false for booleans, and 4 rather than "4" for numbers.

GPU and Parallelism

KeyDefaultDescription
trainer.placement.policy_num_gpus_per_node1Number of GPUs for training
trainer.placement.policy_num_nodes1Number of nodes for training
generator.inference_engine.num_engines1Number of vLLM inference engines for sampling
generator.inference_engine.tensor_parallel_size1Tensor parallel size per inference engine
trainer.micro_forward_batch_size_per_gpu1Micro-batch size per GPU (for forward pass)
trainer.micro_train_batch_size_per_gpu1Micro-batch size per GPU (for gradient accumulation)
generator.inference_engine.gpu_memory_utilization0.8Fraction of GPU memory for vLLM KV cache

When running a small model on multiple GPUs, you typically want to set policy_num_gpus_per_node and generator.inference_engine.num_engines to the same value. For example, on a 4-GPU node:

--backend-config '{"trainer.placement.policy_num_gpus_per_node": 4, "generator.inference_engine.num_engines": 4}'

For large models that don't fit on a single GPU for inference, increase generator.inference_engine.tensor_parallel_size and decrease generator.inference_engine.num_engines accordingly. For example, on 4 GPUs with TP=2:

--backend-config '{"trainer.placement.policy_num_gpus_per_node": 4, "generator.inference_engine.num_engines": 2, "generator.inference_engine.tensor_parallel_size": 2}'

LoRA

LoRA is configured from the client side, not the server. When creating a model via the Tinker SDK, pass a lora_config with the desired rank. For example, in tinker-cookbook recipes:

# LoRA training (default in most recipes)
python -m tinker_cookbook.recipes.sl_loop ... lora_rank=32

# Full-parameter fine-tuning
python -m tinker_cookbook.recipes.sl_loop ... lora_rank=0

No server-side configuration is needed to switch between single-tenant LoRA and full-parameter fine-tuning.

Multi-tenant LoRA

Hosting multiple LoRA tenants concurrently against one server does require server-side configuration on the Megatron backend. At minimum:

{
    "trainer.placement.colocate_all": false,
    "trainer.policy.megatron_config.lora_config.merge_lora": false,
    "trainer.policy.model.lora.max_loras": <max concurrent adapters in a single batch>,
    "trainer.policy.model.lora.max_cpu_loras": <total adapter capacity>
}

merge_lora: false is required so vLLM serves each tenant's adapter by name (with merge_lora: true vLLM only sees the merged base and per-tenant sampling returns the wrong weights). max_cpu_loras must be sized to the peak number of concurrent tenants — there is no on-demand reload, and if vLLM evicts an adapter the next sample() against it 404s. All adapters on one server must share the same (rank, alpha, target_modules) signature; mismatched signatures are hard-rejected at create_model.

See Multi-tenancy for the full operator contract and SFT/RL quickstarts.

torch.profiler

Traces of policy-worker training steps, driven over plain HTTP on a running server rather than through the Tinker SDK. Start the server with --torch-profiler to enable it; without the flag the endpoints return 404.

uv run --extra tinker --extra fsdp -m skyrl.tinker.api \
    --base-model "Qwen/Qwen3-0.6B" --backend fsdp \
    --torch-profiler '{"export_dir": "/tmp/skyrl_traces", "ranks": [0]}'
FieldDefaultMeaning
export_dirrequiredWhere traces are written. An absolute local path, or a cloud URI (s3://, gs://, gcs://)
ranks[0]Global ranks to profile
max_session_duration_sec7200Releases the profiling slot if a client never calls /stop_profiling

Then bracket the steps you care about:

curl -X POST localhost:8000/start_profiling -H 'Content-Type: application/json' -d '{
    "model_id": "<your model id>", "global_step": 120, "export_path_extra": "before-fix",
    "schedule_options": {"skip_first": 0, "wait": 0, "warmup": 1, "active": 5, "repeat": 1}
}'
# ... run training ...
curl -X POST localhost:8000/stop_profiling -H 'Content-Type: application/json' -d '{"model_id": "<your model id>"}'
curl localhost:8000/profiling_status

Traces land in {export_dir}/{global_step}_{export_path_extra}. schedule_options is passed to torch.profiler.schedule and profile_options to torch.profiler.profile. Both are expressed in optim steps: the profiler advances once per optim_step for the requesting model_id, and a trace is written when a window closes, on the transition out of the last active step. /stop_profiling flushes an open window, and for a cloud export_dir blocks until the upload completes.

Only one profiling session runs at a time server-wide — a second /start_profiling gets a 409, and only the owning model_id may stop it. A non-empty target directory is also a 409 unless overwrite: true is passed.

On the FSDP backend with colocate_all: true, profiling requires trainer.policy.fsdp_config.cpu_offload: true. The policy is offloaded to CPU between requests, and the default manual offload path moves parameters with torch.utils.swap_tensors, which fails while the profiler holds references to them.

Only policy workers are profiled; inference engines are not. Setting trainer.policy.torch_profiler_config.* in backend_config is rejected at startup — static config would compete with these endpoints for the same profiler.

See examples/tinker/torch_profiling/ for a runnable client.

Full Config Reference

For the complete list of configuration options, see the SkyRL-Train configuration API reference and the SkyRLTrainConfig dataclass.

On this page