Hosted RL

Train RL policies on cloud GPUs

Pick an environment — gymnasium classic control, Atari, MuJoCo, LLM RLHF, or robotics sim — a baseline repo, and a GPU. Metrics stream live. When the reward curve flattens, deploy the policy as a scale-to-zero cog or serverless endpoint.

CLI

app api POST /api/rl/jobs '{
  "envId": "cartpole-v1",
  "algoId": "ppo",
  "hardware": "gpu-rtx4090"
}'
app api GET  /api/rl/jobs/<job-id>          # poll metrics
app api POST /api/rl/jobs/<job-id>/deploy \
  '{"target": "serverless"}'                # policy endpoint

Per-second GPU billing from the shared credit balance; jobs scale to zero when training finishes.

1 · environment

Pick an environment

2 · algorithm repo

Pick a baseline

3 · gpu tier

Pick a GPU

Per-second billing with the platform margin included, the same catalog as Cog Studio and pricing.

live training metrics

Reward, loss, entropy, and throughput

Launch a run above to stream reward, loss, entropy, and throughput here, polled from /api/rl/jobs/{id}. Signed-out visitors get a simulated run.

Train the policy, then serve it

Train on per-second billed GPUs, deploy the policy as a scale-to-zero endpoint, and call it over HTTP.