Skip to main content

lium clusters

Rent several whole machines on one InfiniBand or RoCE fabric as a single cluster, inspect it, and remove it. Multi-node clusters explains what a cluster is and how to run a job across it; this page is the CLI.

lium clusters [list|up|ps|show|rm] [OPTIONS]

lium clusters on its own is lium clusters list. Every subcommand takes --format [table|json] (default table; --json is the same as --format json).

lium clusters list​

Fabrics with free nodes that can be rented as one cluster: fabric id, type (InfiniBand or RoCE), node configuration, free and total nodes, link rate, whether the link was measured, price per node-hour and location.

lium clusters
lium clusters --format json

The listing is saved, so lium clusters up 1 can name a fabric by its row number. In JSON each fabric has fabric_id, fabric_type, link_rate, fabric_measured, node_count, free_count, gpu_type, gpus_per_node, price_per_node_hour and nodes.

lium clusters up​

Rent N whole nodes of one fabric as a single all-or-nothing cluster.

lium clusters up FABRIC --nodes N --name NAME [OPTIONS]

FABRIC is a row number from the last lium clusters, a fabric id, or a unique prefix of one. The cheapest free nodes are taken. Every member runs the cluster template, shares the pod name, and gets a private overlay address (10.42.0.<rank+1>); node 0 is the master.

FlagEffect
-N, --nodes NHow many whole nodes to rent, at least 2 (required)
-n, --name NAMEPod name given to every member (required)
-t, --template_id IDTemplate ID (default: the newest daturaai/lium-cluster template)
-p, --ports NPorts to expose per node
--ttl DURATIONAuto-terminate every member after a duration (6h, 45m, 2d)
--wait / --no-waitWait until every member is RUNNING with SSH (default: wait)
--timeout SECONDSHow long --wait waits (default 900)
-y, --yesSkip the confirmation prompt, which shows the hourly price of the whole cluster

The --ttl schedule is set on every member right after the rent, before any waiting, so a wait that times out does not leave nodes billing with nothing scheduled.

Once the rent has gone through, every later failure (a member that fails to start, a wait that times out, a TTL that could not be scheduled) exits 3 with a hint not to run lium clusters up again: the nodes are rented and billing, and a second run rents a second cluster. Run lium clusters ps and remove the cluster with lium clusters rm instead.

lium clusters ps​

Your cluster rentals: id, name, status, number of nodes, node configuration, master and price per hour.

lium clusters ps

lium clusters show​

Members of a cluster: rank, pod, status, overlay IP, GPUs and SSH command, which is everything a launcher needs.

lium clusters show CLUSTER [OPTIONS]

CLUSTER is a cluster id, a unique prefix of one, or the pod name its members share.

FlagEffect
--hostfilePrint only an mpirun/DeepSpeed hostfile: one <overlay ip> slots=<gpus> line per node
--torchrun RANKPrint only the torchrun rendezvous flags for the node with this rank: --nnodes N --node_rank R --master_addr A --master_port 29500

lium clusters rm​

Remove every member pod of a cluster in one API call.

lium clusters rm CLUSTER [-y]

The API answers one line per member. A member that failed is printed with its error, the others are still removed, and the command exits 3; run it again to retry the ones that failed, or remove one with lium rm <huid>.

Examples​

lium clusters # Fabrics with free nodes
lium clusters up 1 --nodes 2 -n job # Rent 2 nodes of fabric #1 as one cluster
lium clusters up 1 -N 4 -n job --ttl 12h -y --format json
lium clusters ps # Your clusters
lium clusters show job --hostfile # mpirun / DeepSpeed hostfile
lium clusters show job --torchrun 0 # torchrun flags for rank 0
lium clusters rm job -y # Remove every member

See also​