lium clusters
Rent several whole machines on one InfiniBand or RoCE fabric as a single cluster, inspect it, and remove it. Multi-node clusters explains what a cluster is and how to run a job across it; this page is the CLI.
lium clusters [list|up|ps|show|rm] [OPTIONS]
lium clusters on its own is lium clusters list. Every subcommand takes --format [table|json] (default table; --json is the same as --format json).
lium clusters list
Fabrics with free nodes that can be rented as one cluster: fabric id, type (InfiniBand or RoCE), node configuration, free and total nodes, link rate, whether the link was measured, price per node-hour and location.
lium clusters
lium clusters --format json
The listing is saved, so lium clusters up 1 can name a fabric by its row number. In JSON each fabric has fabric_id, fabric_type, link_rate, fabric_measured, node_count, free_count, gpu_type, gpus_per_node, price_per_node_hour and nodes.
lium clusters up
Rent N whole nodes of one fabric as a single all-or-nothing cluster.
lium clusters up FABRIC --nodes N --name NAME [OPTIONS]
FABRIC is a row number from the last lium clusters, a fabric id, or a unique prefix of one. The cheapest free nodes are taken. Every member runs the cluster template, shares the pod name, and gets a private overlay address (10.42.0.<rank+1>); node 0 is the master.
| Flag | Effect |
|---|---|
-N, --nodes N | How many whole nodes to rent, at least 2 (required) |
-n, --name NAME | Pod name given to every member (required) |
-t, --template_id ID | Template ID (default: the newest daturaai/lium-cluster template) |
-p, --ports N | Ports to expose per node |
--ttl DURATION | Auto-terminate every member after a duration (6h, 45m, 2d) |
--wait / --no-wait | Wait until every member is RUNNING with SSH (default: wait) |
--timeout SECONDS | How long --wait waits (default 900) |
-y, --yes | Skip the confirmation prompt, which shows the hourly price of the whole cluster |
The --ttl schedule is set on every member right after the rent, before any waiting, so a wait that times out does not leave nodes billing with nothing scheduled.
Once the rent has gone through, every later failure (a member that fails to start, a wait that times out, a TTL that could not be scheduled) exits 3 with a hint not to run lium clusters up again: the nodes are rented and billing, and a second run rents a second cluster. Run lium clusters ps and remove the cluster with lium clusters rm instead.
lium clusters ps
Your cluster rentals: id, name, status, number of nodes, node configuration, master and price per hour.
lium clusters ps
lium clusters show
Members of a cluster: rank, pod, status, overlay IP, GPUs and SSH command, which is everything a launcher needs.
lium clusters show CLUSTER [OPTIONS]
CLUSTER is a cluster id, a unique prefix of one, or the pod name its members share.
| Flag | Effect |
|---|---|
--hostfile | Print only an mpirun/DeepSpeed hostfile: one <overlay ip> slots=<gpus> line per node |
--torchrun RANK | Print only the torchrun rendezvous flags for the node with this rank: --nnodes N --node_rank R --master_addr A --master_port 29500 |
lium clusters rm
Remove every member pod of a cluster in one API call.
lium clusters rm CLUSTER [-y]
The API answers one line per member. A member that failed is printed with its error, the others are still removed, and the command exits 3; run it again to retry the ones that failed, or remove one with lium rm <huid>.
Examples
lium clusters # Fabrics with free nodes
lium clusters up 1 --nodes 2 -n job # Rent 2 nodes of fabric #1 as one cluster
lium clusters up 1 -N 4 -n job --ttl 12h -y --format json
lium clusters ps # Your clusters
lium clusters show job --hostfile # mpirun / DeepSpeed hostfile
lium clusters show job --torchrun 0 # torchrun flags for rank 0
lium clusters rm job -y # Remove every member
See also
- Multi-node clusters — what you get and how to run a job across it
lium up— rent a single podlium rm— remove one member by its huid