v27

latestOpenAPI 3.1.0raw.githubusercontent.com2026-04-22216424765.0 KB
gateway.openapi_Gateway

Resume Rlor Trainer Job

post/v1/accounts/{account_id}/rlorTrainerJobs/{rlor_trainer_job_id}:resume

Path parameters

account_idstring required

The Account Id

rlor_trainer_job_idstring required

The Rlor Trainer Job Id

Request body

GatewayResumeRlorTrainerJobBody required

Response

A successful response.

namestring
displayNamestring
createTimestring date-time
completedTimestring date-time
datasetstring

The name of the dataset used for training.

evaluationDatasetstring

The name of a separate dataset to use for evaluation.

evalAutoCarveoutboolean

Whether to auto-carve the dataset for eval.

state'JOB_STATE_UNSPECIFIED' | 'JOB_STATE_CREATING' | 'JOB_STATE_RUNNING' | 'JOB_STATE_COMPLETED' | 'JOB_STATE_FAILED' | 'JOB_STATE_CANCELLED' | 'JOB_STATE_DELETING' | 'JOB_STATE_WRITING_RESULTS' | 'JOB_STATE_VALIDATING' | 'JOB_STATE_DELETING_CLEANING_UP' | 'JOB_STATE_PENDING' | 'JOB_STATE_EXPIRED' | 'JOB_STATE_RE_QUEUEING' | 'JOB_STATE_CREATING_INPUT_DATASET' | 'JOB_STATE_IDLE' | 'JOB_STATE_CANCELLING' | 'JOB_STATE_EARLY_STOPPED' | 'JOB_STATE_PAUSED'

JobState represents the state an asynchronous job can be in.

  • JOB_STATE_PAUSED: Job is paused, typically due to account suspension or manual intervention.
createdBystring

The email address of the user who initiated this fine-tuning job.

rewardWeightsstring[]

A list of reward metrics to use for training in format of "<reward_name>=<weight>".

keepAliveboolean
rolloutDeploymentNamestring

Rollout deployment name associated with this RLOR trainer job. This is optional. If not set, trainer will not trigger weight sync to rollout engine.

nodeCountinteger

The number of nodes to use for the fine-tuning job. If not specified, the default is 1.

acceleratorSecondsobject

Accelerator seconds used by the job, keyed by accelerator type (e.g., "NVIDIA_H100_80GB"). Updated periodically.

serviceModeboolean

Service-mode RLOR trainers currently support full-parameter tuning only. When enabled, trainingConfig.loraRank must be 0 (loraRank>0 is rejected).

directRouteHandlestring
hotLoadDeploymentIdstring

The deployment ID used for hot loading. When set, checkpoints are saved to this deployment's hot load bucket, enabling weight swaps on inference. Only valid for service-mode or keep-alive jobs.

usePurposestring

Use dedicated resources for the job. The only supported value currently is "pilot". Defaults to empty.

forwardOnlyboolean

When true, run the trainer in forward-only mode (no backward/optimizer). Used for reference models in GRPO that only need forward passes.