GPU workloads — NuPaaS Docs
Infrastructure & servers

GPU workloads

GPU workloads are batch jobs — a container image, a GPU count, and a timeout — scheduled onto nodes in your own fleet that have NVIDIA GPUs available. This is a job runner, not a long-lived inference service.

What this is

You submit a container image; NuPaaS creates a Kubernetes job for it on a node in your organization's compute pool that has free GPUs. The job runs to completion, and its exit code, start time and completion time are recorded against the workload. Submit and monitor from /orgs/<org>/platform/gpu-workloads.

Because the work runs on your servers rather than on shared capacity, the ceiling is whatever GPUs you have brought. There is no pool to burst into: a workload asking for more GPUs than any single node can offer will not be placed.

Before you start

Three preconditions, all of which are enforced rather than advisory:

  • The feature must be enabled for your organization. If GPU workloads are not available to you, the page redirects to your projects list with the reason attached rather than rendering an unusable form.
  • You must be an owner or an admin. Members and viewers are redirected. There is no read-only view of the workload list for lower roles.
  • You need at least one server with a usable GPU. GPU inventory is discovered by the node agent and recorded on the node record. A node with no discovered GPUs is not a candidate.

Submitting a workload

Give the form an image, optionally a tag, the number of GPUs, a timeout, and any environment variables the job needs. The image is required; submitting with it blank is rejected before any request is made.

A minimal submission
image:        nvidia/cuda:12.2.0-base-ubuntu22.04
tag:          latest
gpuCount:     1
timeout:      3600
env:          MODEL_PATH=/data/model  BATCH_SIZE=32

On success the platform returns the workload's identifier and the name of the Kubernetes job it created, so you can correlate the row in the panel with what is actually running.

Submission fields

ParameterTypeDescription
imagestringContainer image to run. Required.
tagstringImage tag. Optional; omitted means the image reference is used as given.
gpuCountnumberHow many GPUs the job requests. Must be satisfiable by a single server.
timeoutSecondsnumberWall-clock limit. The panel defaults to 3600. Exceeding it moves the workload to timeout.
envVarsmapKey/value environment variables injected into the container.
projectIdstringOptional. Associates the workload with one of your projects for attribution.

Status reference

ParameterTypeDescription
pendingstatusAccepted and waiting for a server with free GPUs.
runningstatusThe container is executing.
completedstatusFinished. Check the recorded exit code — completed means the job ended, not that it succeeded.
failedstatusThe job ended in error.
cancelledstatusYou cancelled it.
timeoutstatusThe job exceeded its timeout and was stopped.

Cancelling a workload

A pending or running workload can be cancelled from its row in the panel. Cancellation stops the underlying job; it does not roll back anything the job already wrote to storage or to a database. If the job has side effects, make them idempotent before you rely on being able to stop it midway.

Limitations

  • Workloads are batch jobs. There is no autoscaling, no request routing, and no persistent endpoint attached to one.
  • A workload runs on one node. Multi-node distributed training is not modelled here.
  • Scheduling is best-effort against your own capacity. A workload can sit in pending indefinitely if nothing in your fleet can satisfy its GPU request.
  • Logs and traces from the job flow into the same observability surface as everything else — see Logs and telemetry.