latestOpenAPI 3.1.02026-08-19270462702.4 KB

b92f75fd3b61

felix

Create Training Job

Create a new training job.

The provider is resolved from the registry based on (base_model, training_type).

Args: request: Incoming FastAPI request (required by SlowAPI key function). training_request: Training configuration with dataset references and base_model. auth: Authenticated user context.

post/felix/training-jobs

Request body

model_namestring required

User-friendly name for the trained model

base_modelstring required

HuggingFace model identifier (e.g. 'fastino/gliner2-base-v1', 'deepseek-ai/DeepSeek-V4-Flash', or 'nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16').

training_type'full' | 'lora'

Training type: 'full' or 'lora'

validation_data_percentagenumber

Fraction of data held out for validation.

nr_epochsinteger

Maximum training epochs. With early stopping enabled, training typically terminates well before this ceiling.

learning_ratenumber nullable

Peak learning rate for AdamW. When omitted, the trainer's default applies — the per-model catalog rate (2e-4) for decoder LoRA on Modal, 2e-5 elsewhere. Pinning the old 2e-5 default on a decoder LoRA run under-trains the adapter by an order of magnitude.

batch_sizeinteger

Training batch size. Prefer omitting this field so the training service applies the catalog default for base_model. Explicit values equal to this Field default (4) that exceed the model's safe maximum are treated as legacy unset clients and clamped to the catalog default at the training service boundary; any other oversize value is rejected.

seedinteger nullable

Optional reproducibility seed for Modal decoder training. Requests must pin provider_name to 'modal'; Fireworks and non-decoder training reject this field. When omitted, the established Modal worker default (3407) is used.

save_stepsinteger

Save checkpoint every N steps

profile_trainingboolean

Enable structured training profiling for this run and persist a training_profile.json artifact.

wandb_api_keystring nullable

Optional W&B API key for logging

project_idstring nullable

Project ID to associate with this training job. When omitted, the job is anchored to the caller's auto-managed "Default" project so it is always deployable and fleet-eligible.

lora_rinteger nullable

LoRA rank. When omitted, the trainer's default applies: the per-model catalog rank for decoder LoRA on Modal (including 32 for the qualified Nemotron 3.5 Lightning profile), or 16 elsewhere.

lora_alphainteger nullable

LoRA alpha. When omitted, the trainer's default applies: the per-model catalog alpha for decoder LoRA on Modal, or 32 elsewhere.

lora_dropoutnumber nullable

LoRA dropout. When omitted, the trainer's default applies: the per-model catalog dropout for decoder LoRA on Modal, or 0.1 elsewhere.

mask_historyboolean

Decoder SFT loss masking knob. Dense decoder LoRA currently rejects true until assistant-only loss masking is supported by the active trainer. Defaults false to preserve the stock recipe.

warmup_rationumber nullable

Fraction of total training steps for linear LR warmup. Ignored when warmup_steps is set. When omitted, the trainer's default applies — the per-model catalog warmup (0.03) for decoder LoRA on Modal, none elsewhere.

warmup_stepsinteger nullable

Absolute number of linear LR warmup steps. When set, takes priority over warmup_ratio.

lr_scheduler_typestring

LR decay schedule after warmup: 'constant', 'linear', or 'cosine'.

weight_decaynumber

AdamW weight decay (L2 penalty). 0 disables weight decay. Honoured by Modal decoder/dense-LoRA training paths. Not forwarded to GLiNER2 Modal training, which uses its own container default.

early_stopping_patienceinteger

Validation epochs without improvement before stopping. Requires validation_data_percentage > 0. Set to 0 to disable.

early_stopping_min_deltanumber

Minimum validation loss improvement to count as progress. Prevents early stopping from triggering on noise.

provider_namestring nullable

Pin training to a specific provider (e.g. 'modal'). Bypasses automatic provider selection.

system_promptstring nullable

Canonical system prompt written to every decoder training row. When populated, it is persisted on training_jobs.system_prompt and re-injected by inference providers for API-direct callers that omit the system message (train/serve alignment), and prefills the inference-page system-prompt editor. Leave null for PAFT datasets, mixed-prompt uploads, or any case where no single prompt should be pinned at serve time. Ignored for non-decoder tasks.

encoder_learning_ratenumber nullable

GLiNER only: learning rate applied to encoder parameters. When omitted, falls back to learning_rate.

task_learning_ratenumber nullable

GLiNER only: learning rate applied to task-head parameters. When omitted, falls back to learning_rate.

gradient_accumulation_stepsinteger nullable

Accumulate gradients over N mini-batches before each optimizer step. Effective batch size = batch_size * N. Honoured by GLiNER and dense decoder LoRA training; when omitted, dense LoRA applies the per-model catalog default (8 for the H200 Nemotron 3.5 profiles, whose per-device batch is pinned to 1).

auto_data_sizingboolean nullable

GLiNER only. Opt-in: when true, downsample each training dataset to min(max_samples_per_dataset, max(min_samples_per_dataset, samples_per_label * num_labels)). When omitted or false, the full provided dataset is used (no silent downsampling). Defaults to false in the Modal container.

min_samples_per_datasetinteger nullable

GLiNER only: lower bound for auto-sized dataset cap.

max_samples_per_datasetinteger nullable

GLiNER only: upper bound for auto-sized dataset cap.

samples_per_labelinteger nullable

GLiNER only: scaling factor used when computing the auto-sized per-dataset cap.

min_training_stepsinteger nullable

GLiNER only: minimum number of optimizer steps; raises epoch count if the provided nr_epochs would yield fewer steps.

training_algorithm'sft' | 'grpo' | 'dpo'

Training algorithm: 'sft' (default), 'grpo', or 'dpo'. GRPO and DPO are dispatched to the Modal RL entrypoint. GRPO requires rl_config.reward_type from the built-in menu; DPO requires {prompt, chosen, rejected} columns and optional rl_config.dpo_beta / loss_type.

rl_configobject nullable

Algorithm-specific hyperparameters for RL training. Supported keys (all optional unless noted, TRL-aligned defaults applied container-side): max_steps, kl_beta, group_size, sampling_temperature, max_completion_length, reward_type (GRPO; required, one of the built-in reward function names — see rl_training.BUILTIN_REWARDS); dpo_beta, loss_type (DPO); logging_steps (both; defaults to 25, lower for short smoke runs). When reward_type == 'llm_as_judge' (GRPO only) the judge call is routed through brain's '/v1/chat/completions' API authenticated with a per-run pio_sk* key minted by ModalTrainingHandler._launch_and_monitor immediately before spawning the Modal function (the user never supplies the key — minted in the workqueue handler so the cleartext value never enters the SQS message body, injected into the Modal payload at spawn time, revoked from the post-training cleanup hook on terminal status). Additional knobs: llm_judge_model (HuggingFace model id, default 'meta-llama/Llama-3.1-8B-Instruct' — must resolve to a brain catalog entry via resolve_catalog_model_id), llm_judge_rubric (template string with {prompt}/{completion}/{answer} placeholders; falls back to a generic faithfulness/quality rubric scored 1-10 when absent), llm_judge_score_scale (raw max score for normalisation to [0,1], default 10), llm_judge_timeout_s (HTTP timeout per judge call, default 30), llm_judge_max_concurrent (parallelism cap on judge HTTP calls, default 8), llm_judge_max_retries (per-row retry budget on transient HTTP errors, default 1), llm_judge_retry_backoff_s (sleep between retries, default 2.0).

Response

Successful Response

idstring required
user_idstring required
project_idstring nullable

Project ID this training job is associated with

model_namestring nullable
base_modelstring required
validation_data_percentagenumber required
nr_epochsinteger required
learning_ratenumber required
batch_sizeinteger required
seedinteger nullable

Effective reproducibility seed for Modal decoder training. Null for Fireworks, unknown, and other providers that cannot honor this contract.

trained_model_pathstring nullable
job_referencestring nullable
instance_typestring nullable
statusstring required
normalized_statusstring nullable

Canonical status alias for compatibility handling (requested, running, complete, deployed, failed, cancelled)

is_terminal_statusboolean nullable

Whether this status is terminal for polling loops

error_messagestring nullable
created_atstring required
updated_atstring required
started_atstring nullable
completed_atstring nullable
model_auto_selectedboolean nullable
model_selection_reasonstring nullable
task_typestring nullable

Task type derived from training datasets: 'ner', 'classification', 'custom', or 'decoder'

training_typestring nullable

Raw training method as persisted: 'lora', 'qlora', or 'full'.

model_kind'lora' | 'full' nullable

Normalized fine-tune kind: 'lora' for adapters (lora/qlora) or 'full' for merged weights. Null when the persisted training type is unrecognised -- clients must not claim a kind in that case.

artifact_readyboolean nullable

Whether an artifact location is recorded, so there is something to serve.

provider_readyboolean nullable

Whether a provider is already serving this artifact. False is not a deployment blocker: promotion provisions or re-warms a provider.

is_deployableboolean nullable

Whether this job passes server-side deployability validation for its own project. Authoritative -- the same check the deployment endpoints enforce.

deployability_reasonstring nullable

Why the job is not deployable (e.g. 'job_incomplete', 'missing_artifact', 'provider_incompatible'). Null when deployable.

labelsstring[] nullable

Merged labels from training datasets (entity types for NER, class labels for classification)

examplestring nullable

Sample text to pre-load into inference input

metricsobject nullable

Training and evaluation metrics dictionary. Contains final_training_loss, final_validation_loss, best_validation_loss from training logs, and optional evaluation metrics (f1_score, precision_score, recall_score, accuracy) if an evaluation has been run.

version_numberstring nullable

Version number for this training job (e.g., '1', '2', '3')

root_job_idstring nullable

ID of the original/root training job this version derives from

provider_deploymentsobject nullable

Provider-specific deployment metadata written by the training monitor, keyed by provider: {"modal": {...}}.

provider_namestring nullable

Training provider that handled this job (e.g. 'modal'). Jobs predating a provider removal carry an 'archived_<provider>' label.

progress_percentinteger nullable

Overall training completion percentage (0-100). Updated live during training.

current_epochinteger nullable

Epoch currently in progress (1-indexed). Updated live during training.

deployment_statusstring nullable required

Deprecated. Always returns None -- deployment_status no longer exists.

Kept for backward compat with clients that read this field.