# Welcome to Lium # Welcome to Lium Lium is a decentralized GPU rental marketplace built on **Bittensor Subnet 51**. Providers contribute GPU resources to a global pool; renters access those resources on demand; validators keep everyone honest by scoring GPU performance and distributing TAO rewards fairly. ## Pick your audience Choose the section that matches what you're here to do: | Audience | I want to… | Start here | |----------|-----------|------------| | **Providers** | Contribute GPU resources and earn TAO | [Providers →](/providers) | | **Validators** | Run validation, score providers, maintain integrity | [Validators →](/validators) | | **Renters** | Rent GPU compute for ML / data workloads on lium.io | [Renters →](/pod-users) | | **Developers** | Build with the Lium API or CLI | [Developers →](/developers) | | **AI agents** | Go from no account to a running GPU pod, headless | [AI Agents →](/developers/agents) | ## Agent access This site is agent-native: - **If you are an AI agent, start at [AI Agents](/developers/agents)** — the shortest path from no account to a running pod, on one page. - Every page is reachable as `.md` (e.g. [`/providers/quickstart.md`](/providers/quickstart.md)). - Send `Accept: text/markdown` to any page URL to get raw markdown. - AI agents can use the `/mcp` endpoint (`search` + `read_page` tools). - `/llms.txt` and `/llms-full.txt` are at the site root per the [llmstxt.org](https://llmstxt.org) spec. ## What is Bittensor Subnet 51? The Lium compute subnet on Bittensor is a decentralized network where providers contribute GPU resources to a global pool. Users rent these resources for computational tasks — machine learning, data analysis, inference, training — and the system ensures fair compensation for providers based on the quality and performance of their GPUs through the Bittensor validation mechanism. - **Providers**: Provide GPU resources, evaluated and scored by validators. - **Validators**: Securely connect to provider machines to verify hardware specs and performance. - **Renters**: Rent computational resources through the [lium.io](https://lium.io) platform. - **Bittensor**: The decentralized blockchain where compensation is managed and paid out in $TAO. ## Contact and support - **Discord**: [Dedicated channel](https://discord.com/channels/799672011265015819/1291754566957928469) - **Email**: [support@lium.io](mailto:support@lium.io) - **GitHub**: [Datura-ai/lium-io](https://github.com/Datura-ai/lium-io) --- # Official Support and Scam Prevention # Official Support and Scam Prevention Scam and impersonation attempts against Lium users are becoming more common. This page explains how those scams usually work, what information you must never share, and the safest way to contact Lium support. :::warning Keep support conversations in the official ticket The safest support path is the official Lium Discord server → [`#contact-lium-support`](https://discord.com/channels/1350984082733142026/1367618603079307275) → Ticket Tool form → created support ticket. ::: ## How Scams Usually Look Scammers often try to intercept users right after they ask for help. Common patterns include: - You ask a question in a public Discord channel, then receive an automated-looking reply or message telling you to move somewhere else. - Someone invites or adds you to another channel or Discord server that looks like a Lium support ticket. - A person contacts you with the same name, avatar, or wording as a real Lium admin. - You are sent a link for "RPC setup", "wallet connect", "refund", "compensation", "penalty appeal", or "support verification". - You are asked to enter a seed phrase, private key, coldkey, hotkey, password, recovery phrase, or wallet file. A scammer can make a profile look identical or nearly identical to a real Lium team member. Do not rely on display name, avatar, or channel name as proof. ## Information You Must Never Share Never share: - Passwords - Seed phrases - Recovery phrases - Private keys - Coldkeys - Hotkeys - Wallet files - Sensitive credentials - API keys or access tokens Lium support does not need these secrets to help you. Any request for this information is a scam attempt. ## The Safest Way to Contact Lium Support The safest way to contact Lium support is through the official Lium Discord support flow. 1. Open the official Lium Discord server. 2. Go to [`#contact-lium-support`](https://discord.com/channels/1350984082733142026/1367618603079307275). 3. Use the Ticket Tool form. 4. Choose whether your request is about `miner` or `rental`. 5. Fill out the form. 6. Continue the conversation inside the created ticket. ![Official Lium Discord support flow](./assets/contact-lium-support-ticket-tool.png) In the screenshot: 1. **Official Lium Discord server** — make sure you are in the official `lium.io` server. 2. **`#contact-lium-support`** — use this channel to start the support flow. 3. **Ticket Tool form** — fill out the form to create the official support ticket. Lium representatives who respond inside a ticket created through this official channel are the trusted support path. ## Stay Inside the Ticket The safest place to discuss a support issue is inside the official ticket you created. In some cases, a Lium team member may ask to continue in DMs. This can be legitimate only if the move is first confirmed inside your official ticket by a team member who is already participating in that ticket. Before continuing in DMs: - Make sure the request was made inside your official ticket. - Open the Lium team member's profile from inside the official ticket and initiate the DM yourself. - Do not trust a separate DM from someone who only looks like a Lium admin. - Never share seed phrases, private keys, coldkeys, hotkeys, passwords, recovery phrases, wallet files, or sensitive credentials. If someone asks you to move to DMs, another server, another channel, or an external website without confirmation inside your official ticket, treat it as suspicious. ## Red Flags Stop and verify through the official ticket if: - You were manually invited or added to a new "support" channel. - Someone sends you an unsolicited DM after you asked a question in public. - Someone DMs you while impersonating a Lium team member, including using the same display picture or a very similar profile photo. - Someone sends you to another Discord server or another support channel. - Someone sends a wallet connection, RPC, refund, compensation, penalty appeal, or support verification link. - A website asks for your seed phrase, private key, coldkey, hotkey, password, recovery phrase, or wallet file. - Someone pressures you to act quickly outside the official ticket. ## If You Are Not Sure If anything feels suspicious: 1. Stop. 2. Do not click links. 3. Do not share secrets. 4. Open a fresh ticket through [`#contact-lium-support`](https://discord.com/channels/1350984082733142026/1367618603079307275). 5. Ask the Lium team inside that ticket to verify the request. When in doubt, stay inside the official ticket. ## Report Scam Attempts If you identify a scam attempt, report the user to Discord. If you can do so safely, also notify the Lium team inside your official support ticket or in a public channel in the official Lium Discord server. --- # Providers # Providers Contribute GPU resources to the Lium network on Bittensor Subnet 51 and earn TAO rewards. As a provider, you operate a **self-hosted provider** (a lightweight CPU coordinator) that connects one or more GPU nodes — each node earns you rewards proportional to its performance and validator scores. You can run a self-hosted provider yourself or opt into the **Lium.io Central Provider Server** — a hosted instance Lium operates so you don't have to. Pick your path on the [Provider Configuration page](/providers/provider-configuration). ## Quick links - [Get started in 5 min](/providers/quickstart) — register, configure, and launch your first provider - [Rewards](/providers/rewards) — understand earnings, fees, and payouts - [Architecture](/providers/architecture) — mental model of how the coordinator, nodes, validators, and Provider Portal fit together - [Provider Configuration](/providers/provider-configuration) — pick between self-hosting the coordinator or using the Lium.io Central Provider Server - [Self-hosted provider setup](/providers/self-hosted-provider) — full install guide for the do-it-yourself coordinator - [Nodes](/providers/nodes) — bring GPU nodes online (CLI quickstart, manual setup, Sysbox, GPU splitting, CVM) - [Provider Portal](/providers/portal/overview) — manage nodes and pricing via the web UI --- # Provider Quickstart # Provider Quickstart ## Get started in 5 min Three steps to a live provider on Bittensor Subnet 51. Want to know how much you'll earn? See [Rewards](rewards/index.mdx). :::warning Security Notice Run `btcli` on your **secure local machine**, not on the provider server. Only hotkeys belong on provider infrastructure — never your coldkey. ::: ### What GPU models can I bring to this subnet? The validator scores a fixed allow-list of NVIDIA models — Hopper / Blackwell datacenter (H100, H200, H800, B200, B300), Ampere & inference datacenter (A100, A10, T4, V100, P100, P40, M40), workstation (RTX 6000 / 5880 / 5000 Ada, RTX PRO 4000/5000/6000 Blackwell, RTX A-series, L4 / L40 / L40S, Quadro & TITAN), and consumer (full RTX 50/40/30/20-series and GTX 10/16-series). The complete list with per-model reward weight, the unrented-pool eligibility rule, and CVM-eligible GPUs lives in [Architecture → Supported GPUs](./architecture.md#supported-gpus). The source-of-truth is [`GPU_MODEL_RATES` in `lium-io`](https://github.com/Datura-ai/lium-io/blob/main/neurons/validators/src/services/const.py). ### Before you start — wallet & TAO You need a Bittensor wallet (one coldkey + at least one hotkey) and enough TAO on the coldkey to cover the dynamic subnet 51 registration fee. - Don't have a wallet yet? Install `btcli` (`pip install bittensor`) and create one — see the [Bittensor wallet docs](https://docs.bittensor.com/getting-started/wallets) for the `wallet new_coldkey` / `new_hotkey` flow. - Need TAO? Acquire it on a supported exchange and transfer it to your coldkey's SS58 address. The registration fee is dynamic — check the live cost on the [TaoMarketCap Subnet 51 registration page](https://taomarketcap.com/subnets/51/registration) before submitting (or run `btcli subnet list --netuid 51`). ### Step 1 — Register your provider on Subnet 51 ```bash btcli subnet register --netuid 51 --wallet.name --wallet.hotkey ``` Registration burns the dynamic subnet 51 fee from your coldkey balance. The transaction either confirms (your hotkey now has a UID on subnet 51) or fails (insufficient TAO, or the burn cost spiked between command issue and chain inclusion). ### Step 2 — Sign in to the Provider Portal 1. Add your provider hotkey to a supported wallet extension. Any SS58-compatible browser extension works — common choices are the [Polkadot.js extension](https://chromewebstore.google.com/detail/polkadot%7Bjs%7D-extension/mopnmbcafieddcagagdcbnhejhlodfdd) and the [Bittensor Wallet extension](https://chromewebstore.google.com/detail/bittensor-wallet/bdgmdoedahdcjmpmifafdhnffjinddgc). 2. Sign in at [provider.lium.io/signature-login](https://provider.lium.io/signature-login) with that hotkey. 3. Open [Profile Settings](https://provider.lium.io/settings) → **Central Provider** → select **Lium.io Central Provider Server**. 4. In the same settings page, [connect Discord](./portal/discord.md) to keep extra subnet incentives enabled. This opt-in lets Lium operate the CPU coordinator for you, so you can skip the self-hosted provider setup. Prefer to run your own? See [Provider Configuration](./provider-configuration.md) for the self-hosted path. ### Step 3 — Set up your first node A provider with no GPU nodes earns nothing. Bring one online with the [Node Quickstart](./nodes/quickstart.md) — three steps: install [Sysbox](./nodes/sysbox.md) (**required** — validators reject any node without the `sysbox-runc` runtime), run the one-line `mine.sh` installer, and register the node back in the Provider Portal. You're live 🎉 :::info "Live" vs "earning" Being live means the provider is registered and the node is reachable. Earning starts after validators score your node — the first validator probe cycle takes roughly 15 minutes, and meaningful rental revenue depends on actual customer rentals matching your GPU type and price. Track scoring in the [Provider Portal monitoring view](./portal/monitoring.md) and emission externally on [TaoMarketCap](https://taomarketcap.com/subnets/51/miners). ::: ## Next steps - [Node Quickstart](./nodes/quickstart.md) — bring your first GPU online - [Connect Discord](./portal/discord.md) — provider Discord setup - [Provider Configuration](./provider-configuration.md) — self-hosted vs Central Provider Server trade-offs - [Self-hosted provider setup](./self-hosted-provider.md) — full install if you opted out of Central Provider Server - [Architecture](./architecture.md) — mental model of the system - [Provider Portal overview](./portal/overview.md) — day-2 management UI --- # Architecture # Architecture How the pieces of a Lium provider setup fit together. Read this once to build a mental model — then proceed to [Quickstart](./quickstart.md), [Self-hosted provider setup](./self-hosted-provider.md), or [Node Quickstart](./nodes/quickstart.md) depending on the path you've chosen. ## The two-component model A Lium provider setup has two distinct components: 1. **Provider** — a lightweight CPU coordinator (no GPU). It signs with your registered subnet 51 hotkey, talks to validators, and pushes/pulls node configuration. 2. **Nodes** — one or more GPU machines that perform actual rental workloads. Validators probe each node and renters connect to it. ```mermaid graph TD A[Provider
CPU coordinator] --> B[Node 1
GPU] A --> C[Node 2
GPU] A --> D[Node N
GPU] V[Validators] -.scoring.-> B V -.scoring.-> C V -.scoring.-> D P[Provider Portal
provider.lium.io] --> A R[Renters] -.rentals.-> B R -.rentals.-> C R -.rentals.-> D style A fill:#1A75FF,stroke:#0066FF,stroke-width:2px,color:#fff style B fill:#161B22,stroke:#30363D,stroke-width:2px,color:#E6EDF3 style C fill:#161B22,stroke:#30363D,stroke-width:2px,color:#E6EDF3 style D fill:#161B22,stroke:#30363D,stroke-width:2px,color:#E6EDF3 ``` A provider with no nodes earns nothing. Each node you bring online adds to your scoring surface and rental revenue. ## Where the provider runs — your choice The CPU coordinator can run in two places: | Option | Who runs it | Setup | |---|---|---| | **Self-hosted provider** | You, on a Linux box you provide | [Self-hosted provider setup](./self-hosted-provider.md) | | **Lium.io Central Provider Server** | Lium operates it for you | Toggle on in your [Provider Portal profile](https://provider.lium.io/profile) | The choice is a per-account toggle in the [Provider Portal profile page](https://provider.lium.io/profile). You can flip back and forth — see [Provider Configuration](./provider-configuration.md) for the trade-offs and the toggle flow. The nodes are always run by you, regardless of which path you pick for the coordinator. ## How validators interact with your nodes Validators on subnet 51 do two things continuously: - **Probe each node** for the required `sysbox-runc` runtime, hardware specs, and synthetic workload performance. Nodes missing Sysbox are rejected — they earn no emission and cannot be rented. Setup is in [Sysbox](./nodes/sysbox.md). - **Score each provider** based on those probes plus rental activity. Scores translate into TAO emission via Bittensor's standard subnet mechanism. Rental income is separate from emission — renters pay you (in the platform's billing system) for actual GPU-hour usage. See [Provider Portal → Payments](./portal/payments.md) for the payout view.
**What goes on each component** (component breakdown) **On the provider host (CPU coordinator)** - Bittensor `btcli` for hotkey operations - The provider Docker container (from `lium-io/neurons/miners/`) - The registered subnet 51 hotkey (only the hotkey — never the coldkey) **On each node host (GPU)** - NVIDIA driver + Container Toolkit - [Sysbox runtime](./nodes/sysbox.md) (required) - The node Docker container (from `lium-io/neurons/executor/`) - Optionally: [XFS-backed Docker storage](./nodes/docker-storage.md) for [GPU splitting](./nodes/gpu-splitting.md), or [TDX-isolated CVM](./nodes/cvm.md) for confidential workloads **Off-host (managed for you)** - The [Provider Portal](./portal/overview.md) — node registration, pricing, monitoring, payouts - The Bittensor network — subnet 51 emission, hotkey state - Optionally, the Lium.io Central Provider Server if you opted in
## Supported GPUs ### What GPU models can I bring to this subnet? The validator scores a fixed allow-list of NVIDIA models — anything outside this list is rejected at machine-spec time. Within the list, only some models are admitted to the **unrented pool** (so they earn the unrented-emission share); the rest are *operationally supported* (you can register and rent them out) but receive `0` from the unrented-pool reward. The live list — with each model's reference rental price — is available from the API: ```bash curl https://lium.io/api/machines ``` Price limits apply to the reference price returned by this endpoint: the portal enforces a **0.5× floor** and a **3× ceiling** (HTTP 400 outside that range). The static list below is a snapshot; the API is authoritative. The validator source also documents each model's emission weight — see [`neurons/validators/src/services/const.py` → `GPU_MODEL_RATES`](https://github.com/Datura-ai/lium-io/blob/main/neurons/validators/src/services/const.py). #### Datacenter — Hopper / Blackwell `B300 SXM6 AC`, `B200`, `H200`, `H200 NVL`, `H100 80GB HBM3`, `H100 NVL`, `H100 PCIe`, `H800 80GB HBM3`, `H800 NVL`, `H800 PCIe`. #### Datacenter — Ampere & inference `A100 80GB PCIe`, `A100-SXM4-80GB`, `A800 80GB PCIe`, `A10 Tensor Core`, `CMP 170HX`, `T4 Tensor Core`, `Tesla V100 Tensor Core`, `Tesla P100`, `Tesla P40`, `Tesla M40`. #### Workstation — Ada, Ampere, Blackwell `RTX 6000 Ada`, `RTX 5880 Ada`, `RTX 5000 Ada`, `RTX 4500 Ada`, `RTX PRO 6000 Blackwell` (Server / Workstation Editions), `RTX PRO 6000D Blackwell` (China SKU), `RTX PRO 5000 Blackwell`, `RTX PRO 4500 Blackwell`, `RTX PRO 4000 Blackwell`, `RTX PRO 2000 Blackwell`, `RTX A6000`, `RTX A5000`, `RTX A4500`, `RTX A4000`, `RTX A2000`, `L40S`, `L40`, `L4`. #### Workstation — older Quadro / TITAN `Quadro RTX 8000`, `Quadro RTX 6000`, `Quadro RTX 5000`, `Quadro P4000`, `TITAN RTX`, `TITAN V`, `TITAN Xp`. #### Consumer — RTX 50-series `RTX 5090`, `RTX 5080`, `RTX 5070 Ti`, `RTX 5070`, `RTX 5060 Ti`, `RTX 5060`. #### Consumer — RTX 40-series `RTX 4090`, `RTX 4090 D`, `RTX 4080 SUPER`, `RTX 4080`, `RTX 4070 Ti SUPER`, `RTX 4070 Ti`, `RTX 4070 SUPER`, `RTX 4070`, `RTX 4060 Ti`, `RTX 4060`. #### Consumer — RTX 30-series `RTX 3090 Ti`, `RTX 3090`, `RTX 3080 Ti`, `RTX 3080`, `RTX 3070 Ti`, `RTX 3070`, `RTX 3060 Ti`, `RTX 3060`, `RTX 3060 Laptop`, `RTX 3050`. #### Consumer — RTX 20-series & GTX 10/16-series `RTX 2080 Ti`, `RTX 2080 SUPER`, `RTX 2070 SUPER`, `RTX 2060 SUPER`, `RTX 2060`, `GTX 1660 Ti`, `GTX 1660 SUPER`, `GTX 1660`, `GTX 1080 Ti`, `GTX 1080`, `GTX 1070 Ti`, `GTX 1070`, `GTX 1060`. :::info Rented vs unrented pool eligibility Every model on the list is eligible for **rental revenue** (the customer-paid stream). Only a subset earns the **unrented-emission** share — concretely, the models with a non-zero entry in `GPU_MODEL_RATES`. Models marked `0.0` in that dict (e.g. RTX 5080, RTX 4080, A10, T4, RTX A2000, every model older than the RTX 30-series) still mine and rent normally; they just don't accrue unrented-pool weight when idle. For live per-model reference prices, query `https://lium.io/api/machines`. ::: :::tip CVM-eligible GPUs For [Confidential VM (CVM)](./nodes/cvm.md) nodes, only Hopper (H100, H200) and Blackwell (B200, GB200) are supported — consumer and workstation cards do not implement NVIDIA Confidential Computing. ::: ## Where to go next - New provider: [Quickstart](./quickstart.md) (5 min, recommended) or [Self-hosted provider setup](./self-hosted-provider.md) for the do-it-yourself path - Adding GPU nodes: [Node Quickstart](./nodes/quickstart.md) - Day-2 operations: [Provider Portal](./portal/overview.md) --- # Provider Configuration # Provider Configuration Choose how your provider is run — self-host it, or let Lium operate one for you. ## What is the provider? The **provider** is the lightweight CPU coordinator every Subnet 51 provider operator needs. It connects your GPU nodes to validators and the Provider Portal. It is *not* a GPU machine — it just orchestrates them. See [Architecture](./architecture.md) for the full picture. There are two ways to run one, and you choose between them on this page: | Option | Who runs it | When to pick | |---|---|---| | **Self-hosted provider** | You | You already have a Linux box, or you want full control over the process and logs | | **Lium.io Central Provider Server** | Lium | You'd rather skip the CPU server setup and have Lium operate the coordinator for you | The rest of this page is about the second option — the **Lium.io Central Provider Server** opt-in. For self-hosted setup, see [Self-hosted provider setup](./self-hosted-provider.md). ## Key Benefits of the Lium.io Central Provider Server ### Simplified Infrastructure - **No separate CPU machine required**: Skip the dedicated CPU host - **Reduced setup time**: Get earning faster with minimal configuration - **Unified management**: Lium.io's hosted instance acts as your provider for all operations ### Problem Solved Previously, providers faced complex multi-node setups requiring both a CPU host *and* GPU nodes, manual environment configuration, and terminal-based operations that created barriers for newcomers. The Lium.io Central Provider Server opt-in removes the CPU host step. ## Configuration Options You set your preference on the [profile page](https://provider.lium.io/profile). The portal offers two options: ### Lium.io Central Provider Server (hosted) **Best for**: Providers who want the simplest setup with minimal infrastructure management. - **Hosted service**: Lium.io operates the provider for you - **Zero setup**: No CPU host to configure or maintain - **Automatic updates**: Server maintenance and upgrades are handled by Lium - **Reliable infrastructure**: Backed by Lium.io's managed uptime - **Cost**: Free — there is no fee or rental-fee cut for using the Central Provider Server. The trade-off is operational, not financial: because Lium operates the coordinator, your scoring is bounded by its uptime and performance. If the central instance underperforms, your rewards are affected the same way they would be for a degraded self-hosted provider. - **Validator scoring**: Identical to self-hosted. Validators score the nodes themselves; they do not distinguish whether the coordinator is Lium-operated or self-operated. There is no scoring or visibility advantage either way. ### Self-hosted provider **Best for**: Providers who want full control and customization over their infrastructure. - **Full control**: Operate your own provider end-to-end - **Self-managed**: You handle the host, the updates, and the logs - **Direct visibility**: Direct access to server logs and monitoring on your machine ### Switching from self-hosted to Central Provider Server If you already have a self-hosted provider running and decide to opt into the Central Provider Server: 1. Toggle on **Lium.io Central Provider Server** on your [profile page](https://provider.lium.io/profile). 2. From the [Provider Portal Nodes page](https://provider.lium.io/executors), click **Sync From Provider Server** to copy your current node configurations into the central instance. Once the sync completes the Central Provider Server is the source of truth for that hotkey. You can shut down the self-hosted instance whenever it's convenient — there's no automatic migration of the running container, but leaving it on briefly does no harm because validators only score one coordinator per hotkey. ## Getting Started
### Prerequisites Before configuring your provider preference: 1. **Registered Provider**: You need to have a Bittensor coldkey and hotkey registered on subnet 51 2. **Portal Access**: Be logged into the [Provider Portal](https://provider.lium.io) 3. **Profile Access**: Navigate to your profile page ### Configuration Steps 1. **Access Profile Settings**: - Log into the [Provider Portal](https://provider.lium.io) - Navigate to the [profile page](https://provider.lium.io/profile) 2. **Select Your Preference**: - Find the **Central Provider Server** section on the profile page - Toggle on **Lium.io Central Provider Server** to use the hosted instance, or leave it off to keep using your own self-hosted provider 3. **Save Your Selection**: - Confirm your choice - Your preference will be saved and applied immediately :::info Configuration Change You can change your provider preference at any time from your profile page. Changes take effect immediately and will be used for all new job requests from validators. ::: ## Understanding the Choice ### When to use the Lium.io Central Provider Server Choose Lium.io's hosted service if you: - Want the fastest and simplest setup process - Prefer not to manage a separate CPU host - Are new to providing GPUs and want to minimize complexity - Want automatic updates and maintenance handled for you ### When to self-host your provider Run your own if you: - Already have infrastructure and want to use it - Need specific customizations or configurations - Want complete control over your provider infrastructure - Have experience managing server infrastructure ## Next Steps After configuring your provider preference: 1. **Add Your Nodes**: Register your GPU nodes to start earning 2. **Monitor Performance**: Track your nodes and earnings through the portal 3. **Manage Operations**: Handle maintenance, pricing, and customer requests :::info Note If you self-host, make sure your provider is properly configured and running. Setup instructions are in [Self-hosted provider setup](./self-hosted-provider.md). ::: --- # Self-Hosted Provider Setup # Self-Hosted Provider Setup This page covers installing and operating the **self-hosted provider** — the lightweight CPU coordinator that connects your GPU nodes to the Lium network. Skip this entire page if you opted into the [Lium.io Central Provider Server](./provider-configuration.md), which runs the coordinator for you. For the conceptual picture of how the coordinator, nodes, validators, and Provider Portal fit together, see [Architecture](./architecture.md). :::warning Coldkey safety Never place your coldkey on a provider server. Coldkeys belong on a secure, offline machine. Provider infrastructure only needs the registered hotkey. ::: ## Prerequisites Before installing the self-hosted provider, you should have: - A registered Bittensor wallet (coldkey + hotkey) — see the [Bittensor wallet docs](https://docs.bittensor.com/getting-started/wallets) if you haven't created one. - Your hotkey already registered on subnet 51 (`btcli subnet register --netuid 51 ...`) — see [Quickstart Step 1](./quickstart.md#step-1--register-your-provider-on-subnet-51). - Basic Linux administration and Docker familiarity. ## System Requirements The self-hosted provider is a lightweight CPU coordinator — it does **not** need a GPU. | Component | Minimum | Recommended | |-----------|---------|-------------| | CPU | 4 cores | 8 cores | | RAM | 8 GB | 16 GB | | Storage | 50 GB | 100 GB SSD | | Network | 100 Mbps | 1 Gbps | | OS | Ubuntu 20.04+ | Ubuntu 22.04 | GPU nodes have their own separate requirements — see [Node Quickstart → Requirements](./nodes/quickstart.md#requirements). ## Installation ### Step 1 — Clone the repo ```bash git clone https://github.com/Datura-ai/lium-io.git cd lium-io ``` ### Step 2 — Run the installer ```bash chmod +x scripts/install_miner_on_ubuntu.sh ./scripts/install_miner_on_ubuntu.sh ``` Verify the tools are present: ```bash btcli --version docker --version ``` If either is missing, install them directly: - **Bittensor**: [official installation guide](https://github.com/opentensor/bittensor/blob/master/README.md#install-bittensor-sdk) - **Docker**: [Docker's documentation](https://docs.docker.com/engine/install/) ### Step 3 — Configure the `.env` Create the config from the template: ```bash cp neurons/miners/.env.template neurons/miners/.env ``` Set these values: ```bash # Your Bittensor wallet name (check with: btcli wallet list) BITTENSOR_WALLET_NAME=your_wallet_name # Your hotkey name (must already be registered on subnet 51) BITTENSOR_WALLET_HOTKEY_NAME=your_hotkey_name # Public IP of this provider host EXTERNAL_IP_ADDRESS=your.server.ip.address # Path to the Bittensor wallet directory HOST_WALLET_DIR=/home/your_user/.bittensor/wallets # Network ports (EXTERNAL_PORT must be open in the firewall) INTERNAL_PORT=8000 EXTERNAL_PORT=8000 ``` :::caution Network configuration `EXTERNAL_PORT` must be: - open in your firewall / cloud security group - not blocked by your ISP - forwarded correctly if the host is behind NAT ::: ### Step 4 — Start the provider ```bash cd neurons/miners docker compose up -d ``` Verify it's running: ```bash docker compose ps docker compose logs -f ``` ## Adding GPU nodes Once the self-hosted provider is up, you have two ways to attach GPU nodes: 1. **Via the Provider Portal** (recommended) — register each node from the [portal UI](./portal/managing-nodes.md#adding-new-nodes); the portal pushes config to the self-hosted provider automatically. 2. **Via the provider CLI** — run `add-executor` directly inside the provider container (described below). Useful for scripted / batch operations. ### Add a node from the CLI ```bash docker exec pdm run src/cli.py add-executor \ --address \ --port \ --validator ``` | Argument | Meaning | |----------|---------| | `` | Container ID/name (find with `docker ps`) | | `` | Public IP of the node host | | `` | Node's listening port (default `8001`) | | `` | Validator SS58 to associate this node with. **Optional** — see note below. | :::tip Picking a validator Today there is exactly **one active validator sending jobs** on Subnet 51: `5F7X5UpKSr26KU3jKfpLmT8kuKtBNyHhEnfS8xtxPCqCb13p`. You can pass that hotkey explicitly, but the argument is fully optional — if you omit it, the provider falls back to the main Lium validator automatically. There is no benefit to spreading or rotating across multiple validators on this subnet at the moment. ::: ### Remove a node ```bash docker exec pdm run src/cli.py remove-executor \ --address \ --port ``` ### Switch a node's validator ```bash docker exec -it pdm run src/cli.py switch-validator \ --address \ --port \ --validator ``` ### List attached nodes ```bash docker exec pdm run src/cli.py list-executors ``` ## Logs and basic monitoring ```bash # Recent logs docker compose logs --tail=100 # Follow logs docker compose logs -f # List currently attached nodes docker exec pdm run src/cli.py list-executors ``` For node-level performance and validator scoring, use the [Provider Portal monitoring view](./portal/monitoring.md) or [TaoMarketCap Subnet 51](https://taomarketcap.com/subnets/51/miners). To check your wallet status from the host: ```bash btcli wallet overview --netuid 51 ``` ## Operational security The self-hosted provider is exposed to the internet and signs with your hotkey. Treat the host accordingly. 1. **Key management** - Coldkeys stay offline. Never copy them to a provider host. - Only the registered hotkey is needed on the provider. - Rotate hotkeys promptly if a host is compromised. 2. **Network** - Restrict `EXTERNAL_PORT` and SSH to known sources where possible. - Use SSH key authentication, disable password login. - Keep the host patched. 3. **Operations** - Watch for node disconnections — they translate directly into lost emission. - Back up `neurons/miners/.env` and the wallet directory. ## Troubleshooting | Symptom | Likely cause | |---------|--------------| | Node not connecting | Firewall / NAT blocking `EXTERNAL_PORT`, wrong ``, or node host down | | Provider container crashes on start | Wallet path missing or unreadable; check `HOST_WALLET_DIR` mount | | Low or zero rewards | Node not scored yet (wait one validator cycle ~15 min), validator assignment issue, or node failing rental checks — see [Monitoring](./portal/monitoring.md) | | Registration command failed | Insufficient TAO balance for the dynamic registration fee | Pre-support checklist: - [ ] Hotkey registered on subnet 51 (`btcli wallet overview --netuid 51`) - [ ] Provider container running (`docker compose ps`) - [ ] `EXTERNAL_PORT` reachable from the public internet - [ ] Each node's `MINER_HOTKEY_SS58_ADDRESS` matches your registered hotkey - [ ] Wallet directory mounted into the container correctly ## Next - [Bring a GPU node online](./nodes/quickstart.md) - [Provider Portal overview](./portal/overview.md) — manage nodes and pricing in the UI - [Architecture](./architecture.md) — conceptual picture of the system --- # How providers earn on Subnet 51 import {sn51, protocol, links} from "@site/src/data/rewards-numbers"; import Pct from "@site/src/components/Pct"; # How providers earn on Subnet 51 Providers on Subnet 51 earn in two parallel streams: rental fees (paid by renters for using your GPUs) and subnet emission (distributed by Bittensor validators). Both are denominated in alpha, but they land in different places: rental fees are transferred to your **coldkey** once per day (with a delay), while subnet emission accrues as alpha **stake on your hotkey** every tempo (~72 minutes). For live numbers, see the rewards calculator. ## The 30-second answer - **Rental fees**: You earn a share of every USD that renters pay to use your GPUs. A renter pays hourly; the Lium billing system transfers your share to your **coldkey** once per day. - **Subnet emission**: Bittensor emits {protocol.provider_emission_share_pct}% of every epoch's emission to Subnet 51 providers. Validators set your weight via Yuma Consensus, and the emission accrues as alpha **stake on your hotkey** every tempo (~72 minutes) — it does not flow through Lium's billing system. - **Two destinations**: Rental fees arrive as a discrete daily transfer to your coldkey; emission shows up as continuous stake growth on your hotkey each epoch. To see emission, look at your hotkey's **Stake** on the Subnet 51 page — it is not a wallet "Transfers" entry. ## Worked example Suppose you operate 8×H100 GPUs priced at $1.26/GPU/hour with typical rental utilization: - **Rental income** (typical daily usage): ~$193 daily in USD → you receive your share (via the calculator for current fee split) - **Emission income** (estimated): ~45–90 α per day (varies with validator score and total subnet activity) - **Total daily**: Roughly $140–$250 equivalent in alpha — rental fees to your coldkey (with a delay), emission as stake on your hotkey (each tempo) **Note:** This is illustrative. For current numbers, use the Rewards Calculator, which queries live Subnet 51 scores and rental rates. ## Two-stream diagram ```mermaid graph LR Renter["Renter"] Backend["Lium
Backend"] ProviderColdkey["Your
Coldkey
(Alpha)"] ProviderHotkey["Your
Hotkey
(Alpha Stake)"] Emission["Bittensor
Emission
(per epoch)"] Validator["Subnet 51
Validators
(Yuma)"] Weights["Your Validator
Weight
(Score)"] Renter -->|"Pays USD
per hour"| Backend Backend -->|"Fee split,
daily transfer"| ProviderColdkey Emission -->|"Provider share
per epoch"| Validator Validator -->|"Weights"| Weights Weights -->|"Alpha stake
per tempo"| ProviderHotkey style ProviderColdkey fill:#90EE90 style ProviderHotkey fill:#90EE90 style Emission fill:#87CEEB style Validator fill:#87CEEB ``` ## Next steps - **[Rental Fees](./rental-fees.mdx)** — How your share is calculated, where the money comes from, and when it's withheld - **[Subnet Emission](./emission.mdx)** — How Bittensor's per-tempo emission flows to providers, and how the burn-vs-rental split works - **[Rewards Calculator](./calculator.mdx)** — Interactive tool with every input explained - **[Getting Paid](./payouts.mdx)** — When you receive alpha, how to unstake, and current feature-flag state - **[FAQ](./faq.mdx)** — The most-asked questions answered with source links --- # Rental fees import {sn51, protocol, links} from "@site/src/data/rewards-numbers"; import Pct from "@site/src/components/Pct"; # Rental fees When a renter books your GPUs on the Lium platform, they pay an hourly USD price. That payment is split: you receive your share, and Lium retains the platform fee. Earnings are billed daily and paid to your coldkey in alpha tokens. ## What is a rental fee? A rental fee is the USD amount a renter pays per hour (or per epoch) for access to your GPUs. The price is set by you (the provider) and displayed on the Lium portal. When a renter books GPUs for 10 hours at $1.26/GPU/hour for 2 GPUs, the total rental fee is $25.20. This fee is split between you and Lium daily. You receive your share; Lium covers infrastructure costs (billing, on-chain transfers, validator operations, API hosting). ## Your share You earn of every rental payment. This value is configured in production and cannot be changed per-provider; it is the same for all providers on Subnet 51. **Example:** If a renter pays $100 for your GPUs in a day, you receive × $100 in alpha-token equivalent, paid to your coldkey. ## How the price is set You set the hourly price for your GPUs using the **Provider Portal** → **Node Management** interface or the CLI. The CLI command runs **inside the self-hosted provider Docker container** — replace `` with your container name (e.g. from `docker ps`): ```bash docker exec pdm run src/cli.py update-executor-price \ --address --port --price ``` The `update-executor-price` command sets the hourly rate on the node. This rate is advertised to renters and used to calculate daily billing. For self-hosted provider setup, see [Self-hosted provider](../self-hosted-provider.md). ### Price limits **You cannot freely set any price** — two constraints apply: - **Portal floor (0.5×)**: The portal rejects prices below 0.5× the GPU model's reference price with an HTTP 400 error. - **Portal hard cap (4×)**: The portal rejects prices above 4× the GPU model's reference price with an HTTP 400 error. Current limits are 0.5× (floor) and 4.0× (ceiling) of the GPU model's reference price. Fetch the live reference prices from: ```bash curl https://lium.io/api/machines ``` For the full explanation of both caps, see [Price limits](../portal/managing-nodes.md#price-limits) in Managing Nodes. ## Where to see your history Log in to the **Provider Portal** at [provider.lium.io/dashboard](https://provider.lium.io/dashboard) and navigate to **Payments**. You will see: - Daily rental income (in USD equivalent and alpha received) - Rental utilization by node - Historical trends - Unpaid balance (if any) For details on payout mechanics and the current feature-flag status, see [Getting Paid](./payouts.mdx). ## When fees are withheld If a rental fails because of a problem on your node, that rental's fees for the day it fails and the day before are withheld and refunded to the renter. See [Penalties](./penalties.mdx) for the triggers, the exemptions, and how to dispute one. --- # Subnet emission import {sn51, protocol, links} from "@site/src/data/rewards-numbers"; import Pct from "@site/src/components/Pct"; # Subnet emission Subnet 51 receives a share of Bittensor's per-tempo (~72 minutes) TAO emission. Validators split that share across **three pools** — rented, unrented, and burn — and pay the result to providers as alpha.
**Bittensor protocol layer** (general context) Every {protocol.tempo_minutes_approx} minutes (~{protocol.tempo_blocks} blocks), Bittensor emits new TAO to the network: - **{protocol.provider_emission_share_pct}%** to providers (the protocol's "miner" role) - **{protocol.validator_staker_emission_share_pct}%** to validators and stakers - **{protocol.subnet_owner_emission_share_pct}%** to subnet owners Subnet 51 competes for the provider share alongside other subnets, then distributes its slice via **Yuma Consensus**. See Bittensor Emissions and Yuma Consensus.
## Subnet 51 layer Each epoch the validator splits the SN51 provider emission into three pools. ### The three pools | Pool | Share | Who earns | Live source of truth | |------|-------|-----------|----------------------| | **Rented** | — fixed | Nodes whose GPU is currently rented | Grafana — GPU model rate (per-model `gpu_portion`) | | **Unrented** | 0 – — dynamic | Unrented nodes of eligible GPU models | Lium rewards calculator (per-model USD/epoch) | | **Burn** | minus the unrented share | Designated burner UIDs only | — | The unrented and burn pools share a ceiling of . When unrented activity grows, burn shrinks toward zero; when unrented activity is sparse, burn absorbs the slack. "Fixed" here means the rented pool does not move with network activity the way the unrented/burn split does — not that the number can never change. The ceiling is a platform setting the validator reads at the start of every cycle, so Lium can retune it without a subnet release (it moved from to in mid-2026). The current value is always published at lium.io/api/v1/shared-config as `total_burn_emission`; the rented pool is whatever is left over. ### Rented pool Distributed across rented nodes in proportion to a live per-model `gpu_portion`, which the validator updates from rental revenue every time the Lium platform reports a new figure. The current values are on the Grafana dashboard — treat that as the live truth. ### Unrented pool Distributed across unrented nodes of eligible GPU models. The eligible models, the per-`(model, gpu_count)` USD/hr prices, and the per-bucket supply caps all change over time as Lium tunes the validator's incentive parameters — so use the rewards calculator rather than memorizing a list. See [Rewards calculator](./calculator.mdx) for a walkthrough. Eligibility is four things: your GPU is **not currently rented**, it is **not running your own [Default Job](../portal/default-jobs.md)** (running your own Default Job forfeits the unrented incentive for that node), its base model has a **positive bucket cap** (a per-model limit on how many unrented GPUs can earn from this pool), and your validator-assigned **score is > 0**. Flagship nodes have one extra condition: an 8× H200, B200, or B300 node must also offer [GPU splitting](../nodes/gpu-splitting.md) (with a minimum below the node's total), [GPU Profiling (ncu)](../nodes/gpu-profiling.md), or run inside a [confidential VM](../nodes/cvm.md) whose TDX quote the validator verifies. With none of the three, the node stays active and rentable but forfeits this pool while idle. ### GPU-count tiers and capacity caps The unrented pool is not one pot per GPU model. Each model is divided into **GPU-count tiers**, and every tier has its own hourly rate and its own capacity cap (the "bucket cap"). Today the priced tiers are **1 GPU** and **8 GPUs**; Lium retunes this over time, so confirm against the rewards calculator rather than memorizing it. Your node is placed in the tier matching its GPU count. Two consequences follow: - **A GPU count with no priced tier earns nothing from this pool.** A 4-GPU node sits in neither the 1-GPU nor the 8-GPU tier, so it has no rate and no capacity, and its unrented incentive is `0`. It still earns [rental fees](./rental-fees.mdx) normally when someone rents it — only the idle-time incentive is lost. [GPU splitting](../nodes/gpu-splitting.md) is the fix; see below. - **A tier over its cap dilutes everyone in it.** If a tier's cap is 64 GPUs and 128 unrented GPUs are sitting in it, every node there is paid at `64 / 128 = 0.5` of its rate. Nobody is thrown out — the tier is shared proportionally. Your node's tier, its cap, and the multiplier it was paid at all appear on every scoring run in the [Grafana job logs](../troubleshooting.md#1-pull-error-details-from-grafana-job-logs). ### When your GPU-count tier is full Turning on [GPU splitting](../nodes/gpu-splitting.md) also sets a **minimum GPU count per rental**. That minimum names a second tier the validator can rate your idle node against. It gets used in two situations: - **Your GPU count has no priced tier at all** — a 4-GPU node, for example. It is rated against your minimum-split tier from the start, which is the difference between earning and earning `0`. - **Your GPU count has a tier, but that tier is over its cap** — the node may be moved into your minimum-split tier for that cycle, under the rules below. The move in the second case happens only when **all** of these are true: 1. Your node is **idle** — a rented node earns from the rented pool instead. 2. Your own GPU-count tier is **over its cap** this cycle. 3. Your minimum-split tier is a **priced tier with a cap**. 4. Your node's **whole GPU count fits** in that tier's remaining capacity. A node is never split across two tiers. A node that moves is paid at the split tier's rate with **no dilution**, because rule 4 only admits it when it fits under the cap. Nodes already in the split tier are never diluted by an arriving node. **Example.** The 8-GPU H200 tier holds 128 unrented GPUs against a cap of 64, so it pays every node there at `0.5`. Your 8×H200 node has splitting enabled with a minimum of 1 GPU, and the 1-GPU H200 tier holds 2 GPUs against a cap of 10 — 8 free slots, exactly enough for your node. It is rated against the 1-GPU tier at `1.0` instead of `0.5`: twice the effective rate for that cycle. When it happens, the node's job log says so: > Unrented incentive: the 8x NVIDIA H200 tier was over its capacity when this executor was placed, so it is rated against its 1x split tier, which had free capacity and pays a better rate. This is re-evaluated every scoring cycle. The free capacity is shared with every other provider, so a node can be moved one cycle and not the next. If more nodes qualify than there is room for, they are admitted in a fixed internal order until the tier is full — there is no queue position you can buy or influence.
**Burn pool** (ordinary providers do not earn from this pool) The residual: burn share = total burn ceiling − unrented share. Distributed equally across designated burner UIDs.
### Sysbox is required Sysbox is a hard precondition. Without `sysbox-runc` running on your node, your `sysbox_multiplier` is `0` and you earn nothing — from any pool. The legacy partial penalty (a 0.8× multiplier for non-sysbox rented nodes) only applied to rentals created **before** the cutoff at **2026-04-03 12:00 UTC**, which has now passed. See [Sysbox setup](../nodes/sysbox.md). ### Discord is required for extra incentives Providers must keep Discord connected in the Provider Portal to remain eligible for extra subnet incentives. Nodes from providers without connected Discord are treated like spot nodes for incentives: they do not earn from the rented or unrented incentive pools until Discord is connected again. This does not change rental fees paid by renters for active GPU usage. See [Connect Discord](../portal/discord.md#discord-and-extra-incentives).
**What about my score?** (math details) The math depends on which pool your node is earning from. The formulas below assume `sysbox-runc` is running on your node (without it, every score is `0`) and the provider account meets the current incentive eligibility requirements. **Rented pool** (your GPU is currently rented): ``` mining_score = score × gpu_portion × gpu_count / total_gpu_count_of_model incentive = mining_share × mining_score / total_mining_score ``` `gpu_portion` is the live per-model value on the Grafana dashboard. `total_gpu_count_of_model` is the total count of your GPU model across all rented nodes in the epoch — so your slice shrinks as more competitors of the same model come online. **Unrented pool** (your GPU is unrented and its base model is eligible): ``` effective_rate = hourly_rate × unrented_cap_multiplier incentive = rental_share × gpu_count × effective_rate / total_rental_cost ``` `unrented_cap_multiplier = min(unrented_count_in_bucket, max_cap) / unrented_count_in_bucket` — cap dilution per `(base_model, gpu_count_bucket)`. If your bucket already has more unrented GPUs than its cap, every node in it scales down proportionally. `hourly_rate` is the rate of the bucket your node is rated against. For a splitting-capable node that is the better of its own GPU-count rate and its minimum-split rate, and the bucket itself may be the minimum-split bucket — see [When your GPU-count tier is full](#when-your-gpu-count-tier-is-full).
### How emissions arrive Subnet 51 emission accrues as **alpha stake on your hotkey**, added every tempo (~72 minutes) when the subnet's validators set weights via Yuma Consensus. It is **not** processed by Lium's billing system and is **not** transferred to your coldkey — it lands directly on the hotkey you registered on Subnet 51, as stake in the SN51 pool. **Why you may not see it on your hotkey.** Emission is *stake growth*, not a transfer, so it does not appear in a wallet's "Transfers" or extrinsics tab — check the **Stake** view for your hotkey on the Subnet 51 page (e.g. taostats / TAO Market Cap). Two things can also make a hotkey look empty even though emission is arriving: it accrues continuously in small per-tempo increments rather than as one visible payout, and your coldkey may periodically re-stake or move the accrued alpha to a delegate. Rental fees are the separate stream that lands on your coldkey — see [Getting Paid](./payouts.mdx). ### Where to monitor - Grafana — GPU model rate (live rented-pool portion per model) - Lium rewards calculator (live unrented-pool USD/epoch per model) - TAO Market Cap — Subnet 51 - Subnet Alpha — Lium - TAO Stats — Subnets On-chain weights update at the end of each tempo (~72 minutes). The Grafana dashboard updates continuously as the validator processes new rental-revenue messages. --- # Rewards calculator import {sn51, protocol, links} from "@site/src/data/rewards-numbers"; import Pct from "@site/src/components/Pct"; # Rewards calculator The [Lium rewards calculator](https://provider.lium.io/rewards-calculator) is your fastest way to estimate earnings. It runs against live Subnet 51 data, accounting for your GPU model, rental utilization, and node configuration. This page walks you through every input and shows the math underneath. ## Try the calculator Start here: **https://provider.lium.io/rewards-calculator** Enter your GPU model and count, set your rental rate (or use the platform average), then toggle your node features. The calculator returns hourly, daily, and monthly projections. ## Understanding each input ### GPU model & count Select your GPU model from the dropdown (H100, H200, A100, L40S, etc.) and how many you plan to rent. The calculator queries the live SN51 subnet for reward rates specific to that model and quantity. If no estimate is available (e.g. you set GPU splitting or request a multi-GPU configuration the validator has not yet scored), the calculator displays "Estimate unavailable" — this is normal during network ramp-up. Cached fallback estimates use the most recent validator snapshot. ### Base price (hourly rental rate) This is what you set in the provider portal: the USD per hour per GPU you want to charge renters. Your actual earnings depend on this price and the platform's rental fee split. **Example:** If you set $0.25/hr and your rental pool has a moderate-to-high utilization, the expected rental income combines: - Income from rented hours: pool earnings at your price - Emission from unrented hours: subnet emission allocated to unrented GPUs ### Rental rate (utilization) The fraction of your GPUs expected to be rented in a given time period. If you don't know, the calculator defaults to the platform average utilization. The slider ranges from idle to full utilization and directly affects the balance between rental income and emission rewards. ### Node configuration toggles #### Sysbox runtime Enable if you run sysbox (container-level virtualization). **Sysbox is required for any score on Subnet 51 today** — the post-cutoff `sysbox_multiplier` is binary (`0` or `1`), so leaving it off zeroes both the unrented- and rented-pool rewards. The toggle exists in the calculator only to model the legacy partial penalty (a 0.8× multiplier) for rented nodes created before the cutoff (2026-04-03 12:00 UTC); for any node running today, leave it on. See [Subnet emission → Sysbox runtime](./emission.mdx#sysbox-is-required). #### TDX attestation passed Enable if your hardware supports Intel Trust Domain Extensions and you have completed attestation. TDX-enabled nodes may score higher under future SN51 incentive phases. #### GPU splitting Enable if you run [GPU splitting](../nodes/gpu-splitting.md) on the node. Splitting configurations always need a live validator query and are never served from cache, so you will see "Estimate unavailable" whenever no validator is connected. The toggle is on/off only. The calculator always assumes a **minimum split count of 1** — the value that reaches a priced GPU-count tier under today's caps. If the minimum you actually set on the node is higher, treat the estimate as optimistic: the validator rates your node against the tier your real minimum names, and if that tier has no cap the node earns nothing from the unrented pool. See [When your GPU-count tier is full](./emission.mdx#when-your-gpu-count-tier-is-full). ## The formula The calculator uses two main components: ### Rental income per hour ``` Rental income = basePrice × (1 − PLATFORM_FEE) × gpuCount ``` where `PLATFORM_FEE` is Lium's share of each rental dollar. Providers receive of every rental payment. ### Expected earnings The calculator blends rental income with subnet emission based on your utilization rate: ``` Expected per hour = rentedPerHour × rentalRate + unrentedPerHour × (1 − rentalRate) ``` - **rentedPerHour**: rental income + emission earned from rented GPUs - **unrentedPerHour**: emission earned from unrented GPUs - **rentalRate**: your estimated utilization (0.0–1.0) Monthly estimates multiply hourly by 24 × 30 (730 hours, a calendar approximation).
**How the calculator gets its numbers** (advanced) When you click **Calculate**, the calculator calls the SN51 validator via the `/machines/estimate` API endpoint. The validator returns: - **usd_per_epoch**: reward value (emission or rental) per SN51 epoch - **effective_rate**: your provider's unrented-pool rate after the cap-dilution factor is applied (`hourly_rate × unrented_cap_multiplier`); `sysbox_multiplier` is a binary gate that zeros the rate if sysbox isn't running, not a tunable factor - **mining_score**: your competitive score in the rented pool - **gpu_portion**: the live per-GPU-model portion the validator publishes on the Grafana GPU model rate dashboard - **rental_share / burn_share**: the current dynamic split of the ceiling between the unrented pool and the burn pool - **sysbox_multiplier**: binary `0` or `1` post-cutoff (see [emission docs](./emission.mdx#sysbox-is-required)) If the validator is unreachable, the calculator falls back to a cached set of single-GPU estimates (stored in Redis). Multi-GPU and gpu-splitting estimates always require a live validator connection.
## Monthly = hourly × 24 × 30 The calculator uses a 30-day calendar month (730 hours) for monthly projections. Your actual monthly payout varies based on: - Actual utilization (may drift from your estimate) - Validator scoring factors that actually move the score today: sysbox (binary), `unrented_cap_multiplier` (per-bucket cap dilution), and misbehavior windows - Network emission changes (if the dynamic split between the unrented pool and the burn pool shifts) Use the calculator as a planning tool, not a guarantee. Check the [FAQ](./faq.mdx) for reasons your actual earnings might differ. ## Caveats & limits - **Validator unavailable?** The calculator caches fallback estimates; you'll see "Estimate unavailable" for multi-GPU or splitting configs until a validator reconnects. - **Emission volatility:** The subnet's split between the unrented pool and the burn pool shifts as rental demand changes (the rented pool itself holds at regardless of demand). Your emission earnings are not fixed. - **Live per-model rate moves:** The rented pool's per-model `gpu_portion` is an EMA of rental revenue and can shift between epochs. Watch the Grafana dashboard. - **Cap dilution:** If many providers post unrented GPUs in your bucket, `unrented_cap_multiplier` shrinks below 1 and your unrented-pool reward drops proportionally. With GPU splitting on, an idle node in an over-cap bucket may be rated against its minimum-split bucket instead — see [GPU-count tiers and capacity caps](./emission.mdx#gpu-count-tiers-and-capacity-caps). - **Misbehavior windows:** Failed verifications and SSH/connectivity incidents zero out your score for the affected day. - **Price competition:** If many providers undercut your base price, your rental rate (utilization) may fall below your estimate. For the most up-to-date numbers and live SN51 stats, visit [TaoMarketCap Subnet 51](https://taomarketcap.com/subnets/51/miners) or [TaoStats Subnets](https://taostats.io/subnets). --- # Payouts import {sn51, protocol, links} from "@site/src/data/rewards-numbers"; import Pct from "@site/src/components/Pct"; # Payouts This page describes how provider earnings are paid out on Subnet 51. ### Destination The two earning streams land in **different places**: - **Rental fees** are transferred to your **provider coldkey** — the same coldkey you registered with when you joined Subnet 51 — as a daily on-chain transfer. You do not need to provide any additional wallet address. - **Subnet emission** accrues as alpha **stake on your hotkey**, added every tempo (~72 minutes) via Yuma Consensus. It is not routed through Lium's billing and is not transferred to your coldkey. See [Subnet emission → How emissions arrive](./emission.mdx#how-emissions-arrive). ### Currency & amount Both streams are denominated in **alpha tokens**. One alpha token represents a stake in the SN51 dynamic-TAO pool. Your earnings break down as: - **Rental fees earned** (if you rented GPUs): your share of renter USD, converted to alpha and transferred to your coldkey at the daily settlement (with the delay below) - **Subnet emission** (allocated to all eligible providers): your share of SN51 emission, added as alpha stake on your hotkey each tempo (~72 minutes) ### Schedule & delay **Rental fees** are processed **daily**: your earnings from a given day are computed and a transfer to your coldkey is initiated, arriving with a delay of from the earning date. **Subnet emission** does not follow this daily schedule — it accrues continuously as hotkey stake each tempo (~72 minutes), with no billing delay. There is no expedited payout path. If you need faster liquidity, you can unstake your alpha once it arrives and convert it to TAO or another token (see "[How to receive as TAO](#how-to-receive-as-tao)" below). ### Misbehavior decline If a rental fails because of your node, a penalty declines that rental's fees for the day it fires and the day before, and the renter is refunded. Payouts already settled are never clawed back, and subnet emission is unaffected. See [Penalties](./penalties.mdx) for the triggers, the exemptions, and how to dispute one. ### Why don't I see my emission on my hotkey? Emission is **stake**, not a transfer, so it will not show up in your wallet's "Transfers" or extrinsics history. Look at the **Stake** view for your hotkey on the Subnet 51 page (taostats / TAO Market Cap) instead. It also arrives in small per-tempo increments rather than as a single daily payout, and your coldkey may periodically re-stake or move the accrued alpha to a delegate — either of which can make the hotkey look empty even though emission is landing. This is separate from rental fees, which arrive as a discrete daily transfer on your coldkey. ## How to receive as TAO Alpha tokens live in the SN51 dynamic-TAO pool. To convert them to standard TAO or swap for other tokens: 1. **Unstake** your alpha from the SN51 pool (returns it to your coldkey as standard TAO). See the [Bittensor dynamic-TAO unstaking guide](https://docs.learnbittensor.org/subnets/understanding-subnets#unstaking-from-dynamic-subnets). 2. **Swap** TAO for the stablecoin or token of your choice on any decentralized exchange (e.g. MEXC, Coinbase, Uniswap). Unstaking has a default unbonding period of 7 days on the Bittensor network. Plan accordingly if you need the tokens sooner. --- # Penalties import {sn51} from "@site/src/data/rewards-numbers"; import Pct from "@site/src/components/Pct"; # Penalties When a rental fails because of a problem on your node, Lium withholds that rental's fees for the day it fails and the day before, and refunds the renter the same amount. This page describes what triggers a penalty, what exactly is withheld, and what to do if you believe a penalty is wrong. Penalties apply to **rental fees only**. [Subnet emission](./emission.mdx) — the alpha stake that accrues on your hotkey each tempo — is never touched by a penalty. ## What a penalty withholds A penalty applies to the **single rental** that failed, not to your whole fleet, and it covers a window of : the day the penalty is applied (day D) and the day before it (D-1). | What | Effect | |---|---| | Rental fees for that rental, days D and D-1 | Withheld — not paid out to you | | The renter | Refunded the same amount | | Payouts already settled (D-2 and older) | Untouched — never clawed back | | Subnet emission on your hotkey | Untouched | | Your other rentals and other nodes | Untouched | Payouts are settled with a delay of (see [Payouts](./payouts.mdx)). That delay is what makes this window possible: when a penalty is applied, the affected days have not been paid out yet, so no money has to be taken back. ## What triggers a penalty Three events apply a penalty automatically: | Reason code | What happened | |---|---| | `POD_UNDEPLOY_FAILED` | The validator could not undeploy or clean up the pod on your node, and ran out of retries. | | `EXECUTOR_INACTIVE_MID_RENTAL` | Your node stopped sending updates for more than 1 hour while a rental was active. | | `BROKEN_BY_PROVIDER` | You marked a stuck rented pod as broken in the Provider Portal, which force-closes the rental. | A fourth code, `MANUAL`, marks a penalty applied by the Lium team by hand. It is not an automatic trigger. :::note Force-closing a genuinely stuck pod is still the right action — it frees your node and ends a rental the renter cannot use. The penalty is what refunds the renter for that failed rental. ::: ## When a penalty is recorded but not charged Some events are logged with `dry_run = true` and no money is withheld: - **Spot-tier rentals.** A rental that started while the node was [Spot](../portal/node-tier.md) is never charged a penalty. The tier is frozen when the rental starts, so a rental that started on a Secure node can still be penalized even if the node becomes Spot later. - **Scheduled notice periods.** If you set a notice period at least 24 hours ahead and then reclaim the node, the undeploy-failure and inactivity penalties are waived. The waiver extends a short time past the end of the period, because these events are detected with a delay. Two further cases skip the penalty entirely, with nothing recorded at all: the rental produced no billable usage, or the affected days have already been paid out. ## Where to see your penalties The [Penalty Events dashboard](https://grafana.lium.io/d/b7b972f3-a083-415f-886d-ae5b12e2371b/penalty-events?orgId=2) lists every penalty; narrow it down with the `miner_hotkey` and `executor_id` variables at the top of the dashboard. The columns that matter: | Column | Meaning | |---|---| | `reason_code` | Which event caused the penalty | | `amount_withheld` | How much was withheld for that rental | | `dry_run` | `true` = recorded only, `false` = actually charged | | `action` | `apply`, or `revert` if the penalty was undone | See [Grafana Dashboards](../grafana.md#penalty-events) for the rest of the dashboard catalogue. ## Effect on your node tier Penalties also feed the reliability rule that decides your tier. **Penalty coverage** is the total amount withheld by penalties over the trailing , as a share of your billed rental revenue in the same window. If it rises above , all of your nodes move from Secure to Spot and stop earning subnet incentive until your coverage drops back below . It is calculated across all of your nodes together, and reverted penalties do not count toward it. See [Node Tier](../portal/node-tier.md) for the full rule. :::note Penalties recorded with `dry_run = true` withhold no money, but their recorded amount still counts toward penalty coverage. ::: ## If you think a penalty is wrong A penalty can be reverted. When it is, the withheld amount is paid to you and the penalty no longer counts toward your penalty coverage. Open a support ticket through the [official Discord support flow](../../security/official-support-and-scams.md#the-safest-way-to-contact-lium-support) with: - the node id (or hotkey) and the pod id, - the time the penalty was applied, from the dashboard, - why you believe it is wrong — for example the node was reachable the whole time, or the pod had already been removed. The team checks the event against the validator and backend records, and reverts the penalty if it was applied in error. ## Reducing penalties - Keep the node sending updates: most `EXECUTOR_INACTIVE_MID_RENTAL` events come from the agent or the host going silent, not from a GPU problem. See [Troubleshooting](../troubleshooting.md). - Schedule a notice period before planned maintenance instead of pulling a rented node offline. - Switch nodes you cannot keep stable to [Spot](../portal/node-tier.md). They earn no subnet incentive, but an interrupted rental on them costs you nothing. --- # FAQ import {sn51, protocol, links} from "@site/src/data/rewards-numbers"; import Pct from "@site/src/components/Pct"; # Rewards FAQ ## 1. How much will I earn with N×H100/H200/A100 on Subnet 51? **Start with the rewards calculator.** Enter your GPU model, count, and rental rate, and you'll get hourly, daily, and monthly projections based on live SN51 data. Example: 8×H100 at $0.20/hr at typical SN51 utilization might earn $3,000–$4,000 per month from rental fees plus subnet emission, depending on current network conditions. But the actual number depends on: - Current SN51 demand for H100 capacity (drives both your rental utilization and the live `gpu_portion` for the rented pool — see the Grafana GPU model rate dashboard) - Your competitive pricing - Whether sysbox is running on your node (mandatory; without it your score is `0` post-cutoff — see the [emission docs](./emission.mdx#sysbox-is-required)) - The bucket cap dilution for your `(GPU model, count)` pairing (see [the unrented-pool eligibility rules](./emission.mdx#unrented-pool)) - Validator scoring at the time The calculator accounts for all of these. Use it as your primary planning tool, and revisit it weekly as network conditions change. ## 2. Why are there two reward streams instead of one? Subnet 51 combines two independent reward sources: 1. **Rental fees** (platform side) — when a renter books your GPU, Lium collects USD and distributes your provider share daily as alpha. 2. **Subnet emission** (Bittensor protocol side) — the SN51 subnet mints new alpha tokens every ~72 minutes (one tempo) and distributes them to eligible providers via Yuma Consensus. The two streams coexist because rental demand is not always saturated. SN51 emission itself is split into **three pools** (rented, unrented, burn — see [Subnet emission](./emission.mdx#the-three-pools)): when your GPU is unrented and its model is eligible, you earn from the unrented pool; when your GPU is rented, you earn rental income *plus* a slice of the rented pool. See the [Rewards index](./index.mdx) for a diagram showing both streams. ## 3. Why did I receive less than the calculator predicted? Common reasons: - **Actual utilization fell short** — your real rental utilization tracked below the figure you used in the calculator (for example, you assumed steady-high utilization but averaged moderate). Lower utilization means lower rental income. - **Sysbox dropped out** — sysbox is mandatory. Any window where `sysbox-runc` was not running on your node zeros that period's score. - **Live `gpu_portion` shifted** — the rented-pool portion for your GPU model is an EMA of platform-wide rental revenue and moves between epochs. Watch the Grafana GPU model rate dashboard. - **Cap dilution kicked in** — if more unrented GPUs landed in your `(model, count)` bucket than the per-model cap allows, every node in that bucket is scaled down by `cap / unrented_count`. See [GPU-count tiers and capacity caps](./emission.mdx#gpu-count-tiers-and-capacity-caps). - **Your GPU count has no priced tier** — the unrented pool pays per `(model, GPU count)` tier, and only some counts are priced (today 1 and 8). A 4-GPU node earns `0` from that pool unless you enable [GPU splitting](../nodes/gpu-splitting.md) with a minimum count that does have a tier. - **Misbehavior windows** — nodes with SSH failures, crashes, or policy violations forfeit that day's payout. Even one misbehavior incident can reduce your monthly total. - **Emission allocation shifted** — the dynamic split between the unrented pool and the burn pool moved; if rental activity dropped, more emission went to the burn pool (paid to designated burner UIDs only) and less to you. The calculator shows a best-estimate snapshot. Your actual earnings are determined by live network conditions and your node's real-time score. Check the [Payouts](./payouts.mdx) page for details on misbehavior policy. ## 4. When do I get paid? Can I be paid faster? Payouts are processed **daily**. Your earnings from a given day are computed at the end of the day, and a payout is initiated. You receive your alpha with a delay of . **There is no expedited payout path.** If you need faster liquidity, unstake your alpha once it arrives (7-day unbonding period on Bittensor) and swap for a stablecoin on your exchange of choice. ## 5. Why do I have alpha and not TAO? How do I cash out? Subnet 51 uses **dynamic TAO (alpha)**, a subnet-specific token backed by stake in the SN51 pool. When you earn rewards, you receive alpha in your coldkey. To convert to standard TAO or another token: 1. **Unstake** your alpha from the SN51 pool. This returns it to your wallet as standard TAO. See the [Bittensor dynamic-TAO guide](https://docs.learnbittensor.org/subnets/understanding-subnets). 2. **Swap or sell** your TAO on a centralized (Coinbase, MEXC) or decentralized (Uniswap) exchange. Unstaking has a 7-day unbonding period on the Bittensor network. Plan accordingly if you need the tokens urgently. ## 6. What is the burn pool and how big is it? The burn pool is **dynamic**, not a fixed slice. SN51 emission splits into three pools each epoch: - **Rented pool** — . It does not move with network activity, though Lium can retune the figure itself. - **Unrented + burn** — share the remaining ceiling. The split between them is recomputed every epoch by the validator from real rental activity. When unrented activity is low, most of the ceiling falls into the burn pool — paid to designated burner UIDs rather than concentrated on a small set of active providers. As unrented activity ramps up, the burn pool shrinks toward zero. In mature operation with saturated unrented activity, burn can reach zero and the entire flows to active unrented eligible nodes. For the full three-pool table, see [Subnet emission](./emission.mdx#the-three-pools). ## 7. How do I increase my score and earn more? Today, the score depends on a small set of factors that actually move the needle: - **Sysbox running (mandatory)** — without `sysbox-runc` on your node, your score is `0`. See [Sysbox setup](../nodes/sysbox.md). - **Run an eligible GPU model** — only base models with a positive bucket cap earn from the unrented pool. The eligible/ineligible split is tuned over time; use the rewards calculator for the live snapshot rather than relying on a static list. - **Pick a (model, count) pairing that isn't capped out** — if your bucket already has more unrented GPUs than its cap, every node in that bucket is diluted by `cap / unrented_count`. Watch supply. - **Turn on GPU splitting on multi-GPU nodes** — your minimum split count gives the validator a second GPU-count tier to rate the idle node against. It rescues a node whose GPU count has no priced tier at all, and it lets a node escape its own tier when that tier is over its cap. See [GPU splitting](../nodes/gpu-splitting.md#when-splitting-actually-helps). - **Competitive pricing** — set your rental price competitively. Higher utilization shifts you from the unrented pool into rental income + the rented pool. The rented pool's per-model `gpu_portion` is updated live from rental revenue (visible on the Grafana dashboard), so models that get rented heavily earn more from the rented pool too. - **Mind your Default Job** — running your own [Default Job](../portal/default-jobs.md) on an idle node forfeits the unrented incentive for that node. Leave the node in **Lium Default Job** to keep the incentive (Lium runs its own Default Job there instead). - **No misbehavior** — failed verifications and SSH/connectivity incidents zero out your score for the affected day. To increase your earnings: 1. Make sure `sysbox-runc` is installed and running on every node. 2. Run an eligible GPU model and check whether your `(model, count)` bucket is already saturated. 3. Adjust your price to match demand; too high and you won't rent, too low and you leave money on the table. 4. Use the rewards calculator and the Grafana GPU model rate dashboard to plan against live numbers — the static tables in this doc and in the validator config change over time. For detailed operational guidance, see the [Provider portal](../portal/managing-nodes.md) documentation (if available) or the [Getting started guide](../quickstart.md). ## 8. My idle node's rate changed and I didn't change anything. Why? The unrented pool is divided into `(GPU model, GPU count)` tiers with capacity caps, and those caps are shared with every other provider on the subnet. Two things move on their own: - **Other providers arrived or left.** When a tier goes over its cap, every node in it is diluted by `cap / unrented_count`. When those GPUs get rented or go offline, the dilution lifts. - **Your node changed tier.** If you run [GPU splitting](../nodes/gpu-splitting.md), an idle node in an over-cap tier is rated against your minimum-split tier whenever the node fits in that tier's free capacity. That free capacity is shared, so a node can be moved one cycle and not the next. Neither is a penalty, and nothing on your host changed. Both are visible on every scoring run in the node's [Grafana job logs](../troubleshooting.md#1-pull-error-details-from-grafana-job-logs) — check the tier, its cap, and the multiplier. The full rules are in [When your GPU-count tier is full](./emission.mdx#when-your-gpu-count-tier-is-full). The way to stop depending on shared idle capacity is to get the node rented: price it competitively so it earns rental fees plus the rented pool, neither of which is capped this way. --- # Nodes # Nodes A **node** is a GPU machine you connect to your provider so it can be rented out by Lium customers. A provider alone earns nothing — your score, emission, and rental revenue all come from the nodes you bring online. ## Pick your setup path | Path | Best for | Time | |------|----------|------| | [**Quickstart with the `lium` CLI**](./quickstart.md) | Most providers — single command sets up the node end-to-end | ~5 min | | [Confidential VM (CVM)](./cvm.md) | TDX-isolated nodes for sensitive customer workloads | Advanced | ## Required follow-ups After the node is running, complete the setup: - [**Docker storage setup**](./docker-storage.md) — required for GPU splitting today, and rolling out to every node soon. Migrates `/var/lib/docker` to an `overlay2` + `xfs` (`pquota`, `ftype=1`) backing. - [**GPU splitting**](./gpu-splitting.md) — optional, lets a single node serve multiple customers at once. - [**GPU Profiling (ncu)**](./gpu-profiling.md) — optional, opens GPU performance counters so renters can run Nsight Compute. Makes the node whole-host-only while enabled. :::info 8× H200 / B200 / B300: offer splitting or profiling A full 8× node of a flagship model (H200, B200, B300) earns the [unrented incentive](../rewards/emission.mdx#unrented-pool) only if at least one of [GPU splitting](./gpu-splitting.md) or [GPU Profiling (ncu)](./gpu-profiling.md) is enabled. With neither, the node stays active and rentable but earns no idle payout. Enabling GPU Profiling suspends GPU splitting — open counters make the node whole-host-only — so pick the one that fits your fleet. ::: ## Operate your nodes Day-to-day tasks (registering in the portal, updating prices, notice periods, deletion, troubleshooting) live under the [Provider Portal → Managing Nodes](../portal/managing-nodes.md) page. --- # Node Quickstart # Node Quickstart ## Get started in 5 min Three commands and a portal click bring a GPU machine online as a Lium node. ### Requirements - **OS**: Ubuntu 22.04 (recommended) - **GPU**: NVIDIA driver installed (`nvidia-smi` works) - **Docker**: installed and running - **Network**: 50 Mbps download, public IP, ports open (default `8080` HTTP and `2200` SSH) - **Hardware**: ≥ 8 GB RAM, ≥ 100 GB free storage - **A provider already registered on Subnet 51** — see [Provider Quickstart](../quickstart). You'll need its hotkey SS58 address. ### Optional (strongly recommended) - **[Docker Storage Setup](./docker-storage.md)** — provision XFS-backed Docker storage with project quotas so each rented pod is capped at the disk size advertised in the portal. :::danger Why this matters Without XFS + `pquota`, a renter can write across the **entire host disk** instead of the slice they paid for. The validator detects this as a misbehaving node and your **score drops to 0**, which means **no emission and no rentals** until you reprovision storage. ::: ### Step 1 — Install Sysbox (required) [Sysbox](./sysbox.md) is **required**. Validators reject any node missing the `sysbox-runc` runtime — without it the node earns no emission and cannot be rented. Install it before anything else. The `lium-io` repo ships an installer that also handles the NVIDIA Container Toolkit: ```bash curl -fsSL https://raw.githubusercontent.com/Datura-ai/lium-io/main/neurons/executor/nvidia_docker_sysbox_setup.sh | sudo bash ``` Verify with the same command our validator uses: ```bash docker run --rm --runtime=sysbox-runc --gpus all daturaai/compute-subnet-executor:latest nvidia-smi ``` If `nvidia-smi` output appears — you're good. Running Docker ≥ 29.2.0, or hitting a `permission denied` error? See the [Sysbox page](./sysbox.md#docker--2920-cdi-compatibility-fix) for the CDI fix. ### Step 2 — Run the node in one script ```bash curl -fsSL https://lium.io/mine.sh | bash -s -- -k ``` This installs the [`lium`](https://github.com/Datura-ai/lium) CLI, then runs `lium mine`, which: | Step | Action | |------|--------| | 1 | Clone or update the `lium-io` repo into `./compute-subnet` | | 2 | Run `scripts/install_executor_on_ubuntu.sh` (Docker + NVIDIA Container Toolkit checks) | | 3 | Render `neurons/executor/.env` from the template (ports + hotkey) | | 4 | `docker compose up -d` and wait for the container to report `healthy` | | 5 | Run the validator's own check via `daturaai/lium-validator:latest` | You'll be prompted for ports interactively. Defaults work for most providers: - **Service port** — node HTTP API (default `8080`) - **Node SSH port** — used by validators to SSH into the container (default `2200`) - **Public SSH port** — only if behind NAT and forwarding a different port - **Renting port range** — optional, e.g. `2000-2005` To skip prompts and accept all defaults, append `--auto`: ```bash curl -fsSL https://lium.io/mine.sh | bash -s -- -k --auto ``` When the script finishes, note the printed **GPU type, IP address, port, and GPU count** — you'll paste them into the portal next. ### Step 3 — Register the node in the Provider Portal 1. Sign in at [provider.lium.io](https://provider.lium.io) — if you haven't yet, follow [Provider Quickstart → Step 2](../quickstart#step-2--sign-in-to-the-provider-portal). 2. In the sidebar, click **Add Node**. 3. Fill the form with the values printed at the end of Step 2: - **GPU Type** (e.g. `RTX 4090`) - **GPU Count** - **IP Address** - **Port** (default `8080`) - **Price per GPU (USD/hr)** 4. Click **Add Node** to submit. After submitting, the node will appear in your portal — confirm it shows as active. Detailed walk-through: [Managing Nodes → Adding New Nodes](../portal/managing-nodes#adding-new-nodes). ## What's next 1. [Docker Storage Setup](./docker-storage.md) — provision XFS-backed Docker storage for stable rentals. 2. [GPU splitting](./gpu-splitting.md) — let multiple customers rent partial GPUs from one node. 3. [CVM (Confidential VM)](./cvm.md) — TDX-isolated nodes for sensitive workloads. 4. [Manage prices and notice periods](../portal/managing-nodes.md) in the Provider Portal. ## Other useful `lium` commands ```bash lium gpu-splitting check # inspect host before enabling GPU splitting lium gpu-splitting setup # provision XFS-backed Docker storage non-interactively lium gpu-splitting verify # confirm the host meets the splitting requirements ``` See [GPU Splitting](./gpu-splitting.md) for the full workflow. --- # Sysbox # Sysbox [Sysbox](https://github.com/nestybox/sysbox) is a container runtime that lets customers run Docker-in-Docker securely inside pods — without `--privileged` mode. Many workloads (custom image builds, CI/CD pipelines, system-level tooling) require it. :::danger Sysbox is required Every validator probes every node for the `sysbox-runc` runtime. Nodes missing it are rejected — they earn **no emission and cannot be rented**. Sysbox is a hard requirement for every Lium node, not an optional optimization. Setup takes ~5 minutes. ::: ## Install The `lium-io` repo ships an installer that takes care of NVIDIA Container Toolkit + Sysbox in one go. From any Ubuntu host: ```bash curl -fsSL https://raw.githubusercontent.com/Datura-ai/lium-io/main/neurons/executor/nvidia_docker_sysbox_setup.sh | sudo bash ``` Or, if you already have the [`lium-io` repo](https://github.com/Datura-ai/lium-io) cloned locally: ```bash cd lium-io/neurons/executor chmod +x nvidia_docker_sysbox_setup.sh sudo ./nvidia_docker_sysbox_setup.sh ``` Confirm `/etc/docker/daemon.json` includes the sysbox runtime: ```json { "runtimes": { "sysbox-runc": { "path": "/usr/bin/sysbox-runc" } } } ``` Restart Docker: ```bash sudo systemctl restart docker ``` ## Verify Run the same command our validator uses: ```bash docker run --rm --runtime=sysbox-runc --gpus all daturaai/compute-subnet-executor:latest nvidia-smi ``` If you see `nvidia-smi` output — you're good. ## Docker ≥ 29.2.0: CDI compatibility fix Running **Docker 29.2.0 or later**? Sysbox + GPU will fail with a `permission denied` error: ``` OCI runtime create failed: ... failed to open OCI spec file: ... permission denied ``` This is because Docker 29.2.x [enables CDI (Container Device Interface) by default](https://docs.docker.com/reference/cli/dockerd/#disable-cdi-devices), routing `--gpus` through CDI — incompatible with sysbox's user namespace. Disable CDI by adding `"features": {"cdi": false}` to `/etc/docker/daemon.json`: ```json { "features": { "cdi": false } } ``` Restart Docker and re-run the verify command: ```bash sudo systemctl restart docker ``` This change is safe and reversible — remove the `features` block and restart Docker to re-enable CDI. ## Troubleshooting - **`sysbox-runc not found`** — the installer didn't finish. Re-run `nvidia_docker_sysbox_setup.sh` and check its output. - **GPU not visible inside the container** — confirm NVIDIA Container Toolkit is installed (`nvidia-container-cli --version`) and the Docker daemon was restarted after the install. - **Validator still reports Sysbox missing** — wait one validation cycle (~15 min) and re-check from the Provider Portal. --- # Docker Storage Setup # Docker Storage Setup Lium expects every node's Docker daemon to use the **`overlay2`** storage driver backed by an **`xfs`** filesystem mounted with **`pquota`** and `ftype=1`. The [`lium`](https://github.com/Datura-ai/lium) CLI inspects the host, formats the target device, and migrates `/var/lib/docker` for you. :::info When you need this - **Required today** if you want to enable [GPU splitting](./gpu-splitting.md) — preflight blocks the portal toggle until the storage driver and backing filesystem match. - **Required for every node soon** — Lium will extend this requirement to all nodes regardless of GPU splitting, so completing this setup now future-proofs the host. ::: The rest of this page covers the host-side setup and verification. Once the host is ready, GPU-splitting providers finish the portal-side step from [Managing Nodes → Setting Minimum GPU Count for Rental](../portal/managing-nodes.md#setting-minimum-gpu-count-for-rental). ## Before You Run the CLI The `--device` argument is the storage target Lium will use for Docker storage. You do **not** need to: - pre-format it as XFS - mount it at `/var/lib/docker` - manually configure `overlay2` Lium handles partition creation, XFS formatting, mount setup, and Docker storage migration after the target passes safety checks. ### Safe Device States Lium will use the device when it is one of these states: - a blank whole disk - a whole disk with one unambiguous free-space region - an unused partition with no filesystem - an empty XFS partition with `ftype=1` ### Device States That Will Be Rejected Lium will skip or reject the device when it is: - the root disk - the current Docker backing device - mounted or otherwise in use - a partition with an existing non-XFS filesystem - a non-empty XFS partition - a disk with an ambiguous layout - a loop or device-mapper target ## Prepare the Device ### Recommended Beginner Path The safest path is to attach a fresh extra SSD/NVMe and leave it unused and unmounted before you run the CLI. 1. Attach a new extra disk. 2. Make sure it is not your root disk. 3. Make sure it is not the device currently backing `/var/lib/docker`. 4. Pass that whole-disk path to the CLI. ### Quick Checks Use these commands before running setup: ```bash # See the candidate device, filesystem, and mount state lsblk -o NAME,PATH,TYPE,FSTYPE,SIZE,MOUNTPOINTS # See which device backs / findmnt / # See which device currently backs Docker storage findmnt /var/lib/docker ``` If you want to reuse an existing XFS partition, confirm it is XFS with `ftype=1`: ```bash sudo xfs_info /dev/ ``` ### Common Remediation - If the device is mounted, unmount it first and confirm nothing important is using it. - If the device has old filesystems or old data, the safest option is to use a different empty disk. - If you want to reuse the same disk and it is safe to erase, reset it to a clean disk or unused partition before rerunning the CLI. :::warning Single-disk hosts If you only have one drive, there is **no supported single-disk path**. The `lium gpu-splitting` CLI explicitly rejects the root disk and the disk currently backing `/var/lib/docker`, and this page does not document a manual workaround. Reusing the only drive on the host would mean repartitioning the live system — that is risky, unsupported, and not covered by the CLI safety checks. The only supported options today are: attach a second disk, or reinstall the host onto an XFS root before bringing the node back online. ::: Example: reset a dedicated extra disk so Lium can partition and format it itself: ```bash # Warning: destructive. Replace /dev/nvme1n1 with the extra disk you want to erase. sudo umount /dev/nvme1n1p1 2>/dev/null || true sudo parted -s /dev/nvme1n1 mklabel gpt ``` After that, rerun: ```bash lium gpu-splitting check --device /dev/nvme1n1 ``` ## Run the Storage Setup Use the CLI in this order: ```bash # Inspect the host and confirm the device is eligible lium gpu-splitting check --device /dev/ # Apply the storage migration sudo lium gpu-splitting setup --device /dev/ # Verify the final state lium gpu-splitting verify ``` What each command does: - `check` inspects the host and prints the exact storage plan without changing anything - `setup` performs the Docker storage migration - `verify` confirms the final Docker storage state :::note CLI naming The commands live under `lium gpu-splitting` because GPU splitting is the first feature that depends on this storage layout. The migration itself is a generic Docker storage setup and the same commands are used regardless of whether you plan to enable GPU splitting. ::: :::warning Destructive operation `setup` stops Docker, formats the target partition as XFS, and remounts `/var/lib/docker`. It backs up existing Docker data first, but you should have an independent backup before running it on a host with live containers. ::: ## Verify the Host After setup, confirm these requirements are satisfied: - Docker storage driver is `overlay2` - Docker backing filesystem is `xfs` - `Supports d_type` is `true` - `/var/lib/docker` is mounted with `pquota` or `prjquota` Quick verification: ```bash docker info findmnt /var/lib/docker lium gpu-splitting verify ``` ## Manual setup If you can't or don't want to use the CLI, the manual sequence is: stop Docker → format the target partition as XFS with `ftype=1` → mount it at `/var/lib/docker` with `pquota` → restore the previous contents → restart Docker. ```bash # Stop Docker sudo systemctl stop docker docker.socket # Make sure /etc/docker/daemon.json sets the overlay2 storage driver # { # "storage-driver": "overlay2" # } # Format the new partition (replace $new_partition) new_partition= sudo rsync -aXS /var/lib/docker/ /tmp/docker-backup/ sudo mkfs.xfs -n ftype=1 "$new_partition" -f sudo mount -t xfs -o defaults,inode64,pquota "$new_partition" /var/lib/docker sudo rsync -aXS /tmp/docker-backup/ /var/lib/docker/ # Make the mount permanent UUID=$(sudo blkid -s UUID -o value "$new_partition") echo "UUID=${UUID} /var/lib/docker xfs defaults,inode64,pquota 0 2" | sudo tee -a /etc/fstab # Start Docker sudo systemctl start docker ``` Then run `docker info` and confirm `Storage Driver: overlay2`, `Backing Filesystem: xfs`, `Supports d_type: true`. ## Next step (GPU splitting only) If you're enabling GPU splitting, finish the portal-side configuration: 1. Open the [Managing Nodes guide](../portal/managing-nodes.md#setting-minimum-gpu-count-for-rental). 2. Follow the **Setting Minimum GPU Count for Rental** steps. For nodes that aren't using GPU splitting, no further action is needed — the host is ready and stays ready when the requirement extends to all providers. ## Troubleshooting If the CLI does not use the device you passed, check these first: - it is not the root or current Docker device - it is not mounted - it does not contain an existing non-XFS filesystem - it is not a non-empty XFS partition - the disk layout is simple enough for Lium to classify safely If you are unsure, the most reliable fix is to use a fresh extra disk and rerun `lium gpu-splitting check --device /dev/`. --- # GPU Splitting # GPU Splitting GPU splitting lets a single node serve multiple customers at once — each renting an integer subset of GPUs. It increases utilization and lets you price flexibly. ## Host prerequisite: Docker storage GPU splitting requires Docker on `overlay2` + XFS with `pquota` and `ftype=1`. Run the CLI workflow in [Docker Storage Setup](./docker-storage.md) before enabling the feature in the portal — the `Edit` button only unlocks once preflight passes. (This same storage layout will be required for every node in the near future, so the work is not GPU-splitting specific.) ## Enable GPU splitting in the portal After the host meets the prerequisites: 1. Open the node's details page in the [Provider Portal](https://provider.lium.io). 2. Find the **GPU Splitting** panel — an **Edit** button appears once preflight passes. 3. Set the minimum GPU count per rental — at least `1`, and no more than the node's total GPU count. On flagship nodes it must be **below** the total to count for the idle payout (see below). Customers can then rent any integer count between your minimum and the node's total. To disable splitting, clear the minimum — only allowed when no pod is currently using a partial allocation. ## When splitting actually helps **More renters can match your node.** A meaningful share of customers want a single GPU rather than a full 8× node, and an 8-GPU node without splitting is invisible to that segment. Widening the pool of renters that can match your node usually raises utilization on multi-GPU hosts. **On flagship nodes it protects the idle payout.** An 8× H200, B200, or B300 node must offer GPU splitting, [GPU Profiling (ncu)](./gpu-profiling.md), or run inside an attested [confidential VM](./cvm.md) to earn the [unrented incentive](../rewards/emission.mdx#unrented-pool) while idle. Splitting counts only when the minimum GPU count per rental is set **below** the node's total — a minimum equal to the full node size is a whole-host rental in practice and does not qualify. **Your minimum split count also gives the validator a second GPU-count tier for your idle node.** The unrented incentive is paid per `(GPU model, GPU count)` tier, and each tier has its own rate and its own capacity cap. Your minimum names a second tier the validator can rate the node against while it is idle: - **If your GPU count has no priced tier, splitting is what makes the node eligible at all.** Only some GPU counts are priced — today 1 and 8. A 4-GPU node earns `0` from the idle pool; the same node with a minimum split count of 1 is rated in the 1-GPU tier and earns. - **If your GPU-count tier is over its capacity cap, the node can be rated against your minimum-split tier instead** — but only when the node's whole GPU count fits in that tier's free capacity. Being moved means no cap dilution for that cycle. Both incentive effects apply only while the node is **idle**. Once it is rented it earns rental fees plus the rented pool, and neither depends on splitting. Splitting never changes your rental price or your rental fees, and on single-GPU nodes it has no effect. For the exact tier rules, an example, and the log line the validator writes when it moves a node, see [Subnet emission → When your GPU-count tier is full](../rewards/emission.mdx#when-your-gpu-count-tier-is-full). --- # Confidential Virtual Machine (CVM) # Confidential Virtual Machine (CVM)
## Key features - **Flexible management**: full control over your machines and availability - **Pause new rentals**: stop a rented node from accepting future rentals without interrupting the current renter - **Maintenance scheduling**: plan downtime without penalties - **Network-wide visibility**: see every node on the platform alongside your own - **Grafana integration**: detailed performance dashboards on each node's detail page ## Getting started Before you begin, ensure you have: 1. **Registered provider** — Bittensor coldkey and hotkey registered on subnet 51 2. **Provider running** — either a [self-hosted provider](../self-hosted-provider.md) or the [Lium.io Central Provider Server](../provider-configuration.md) opt-in 3. **Wallet extension** — any SS58-compatible browser extension (e.g. [Polkadot.js extension](https://chromewebstore.google.com/detail/polkadot%7Bjs%7D-extension/mopnmbcafieddcagagdcbnhejhlodfdd) or [Bittensor Wallet extension](https://chromewebstore.google.com/detail/bittensor-wallet/bdgmdoedahdcjmpmifafdhnffjinddgc)) with your registered provider hotkey imported, used to sign in to the portal 4. **Nodes** — at least one GPU node running. See [Node Quickstart](../nodes/quickstart.md). ## Navigation The Provider Portal is organized into several sections: - **Nodes** — manage your registered GPU machines (status, pricing, notice periods, notifications) and view all nodes across the platform - **Add Node** — register a new GPU node - **Provider Payouts** — earnings from rental fees, broken down by day, validator, and node - **Machine Requests** — pending rental inquiries from renters that you can match to your nodes ## Next steps 1. [Sign in with your provider hotkey](./authentication.md) 2. [Add your first node](./managing-nodes.md) 3. [Connect Discord](./discord.md) so renters can request private pod support channels 4. [Monitor performance](./monitoring.md) 5. [Track payouts](./payments.md) --- # Authentication # Authentication This covers the signup and signin process for the [Provider Portal](https://provider.lium.io), including how to authenticate using your Bittensor provider hotkey. A provider account is created on first successful wallet-hotkey login after the hotkey is registered on subnet 51. ## Prerequisites Before you can access the Provider portal, you'll need: - A valid Bittensor provider hotkey registered on subnet 51 - A provider hotkey you can sign with, either through a browser wallet extension or the Lium CLI on a machine with the Bittensor wallet available - A Discord account if you plan to enable extra subnet incentives ## Login options For first login, use wallet-hotkey authentication. For subsequent logins, you can use either: 1. Login with Password 2. Login with Wallet Hotkey ## Login with Provider Hotkey 1. **Add your Provider Hotkey into a wallet extension**: import your provider hotkey into any SS58-compatible wallet extension. Common options include the [Polkadot.js extension](https://chromewebstore.google.com/detail/polkadot%7Bjs%7D-extension/mopnmbcafieddcagagdcbnhejhlodfdd) and the [Bittensor Wallet extension](https://chromewebstore.google.com/detail/bittensor-wallet/bdgmdoedahdcjmpmifafdhnffjinddgc). 2. **Login with Wallet Hotkey**: Go to [signature login page](https://provider.lium.io/signature-login) and login with wallet hotkey. 3. If Discord is not connected yet, the portal shows a Discord onboarding prompt. Follow it, or continue and connect later from **Profile Settings**. ## Browserless login with the CLI If you do not want to use a browser extension, the CLI can create and authenticate the provider account with the local Bittensor hotkey: ```bash lium config set provider.coldkey lium config set provider.hotkey lium provider portal login ``` For automation, add `--json`: ```bash lium provider portal login --json ``` The command signs the current Unix timestamp with the hotkey, calls the Provider Portal API, and stores the provider JWT locally under `~/.lium/provider/`. After login, connect Discord from the CLI: ```bash lium provider config connect-discord ``` On headless machines or agent flows: ```bash lium provider config connect-discord --json --no-wait ``` The JSON output includes the Discord authorization URL. Open that URL as the Discord account owner to approve the connection. ## Login with password 1. After wallet-hotkey login, go to [Profile Settings](https://provider.lium.io/settings) and set your password. 2. Your password will be linked with your hotkey. 3. Once you set your password, you can use password login for later sessions. You can also set the password from the CLI after `lium provider portal login`: ```bash lium provider config set-password ``` :::info Note The Profile page password form asks for your current password. If no password is set yet, leave the current password field empty. The CLI `set-password` command uses hotkey signature authentication instead. ::: ## Connect Discord after login After you can access the Provider Portal, connect Discord from **Profile Settings**. See [Connect Discord](./discord.md) for the incentive eligibility and renter support details. ## Troubleshooting Authentication ### Common Issues #### "Provider Hotkey not registered in Bittensor subnet 51" Error - Confirm your hotkey is properly registered on the Bittensor network - Confirm your provider is running properly — for self-hosted setups see the [self-hosted provider setup](../self-hosted-provider.md); if you opted into the [Lium.io Central Provider Server](../provider-configuration.md), check that the toggle is still enabled in your profile --- # Connect Discord # Connect Discord Connect Discord in the Provider Portal so renters can invite you into private pod support channels. Without a connected provider Discord account, renters see **Provider has not connected Discord** on pods running on your machines and cannot create the shared channel from the pod detail page. Keeping Discord connected also keeps your provider account eligible for extra subnet incentives. ## Discord and extra incentives Providers must keep Discord connected to remain eligible for extra subnet incentives. If Discord is not connected, your nodes do not earn the incentive portion until Discord is connected again. This requirement applies to subnet incentives only. Rental fees paid by renters for active GPU usage are separate. ## Connect your provider account 1. Sign in to [provider.lium.io](https://provider.lium.io). 2. Open **Profile Settings**. 3. Find the **Discord** section. 4. Click **Connect Discord**. 5. Approve the Discord OAuth prompt. 6. Return to the Provider Portal and confirm the Discord section shows **Connected**. The Provider Portal stores your Discord user ID on your provider profile. The support channel flow uses that ID to grant you access to channels for pods rented from your machines. ## Connect from the CLI After logging in with the provider CLI: ```bash lium provider portal login lium provider config connect-discord ``` The CLI opens the Discord authorization URL when possible. On a headless machine or in an agent workflow, use: ```bash lium provider config connect-discord --json --no-wait ``` Open the returned `authorization_url` as the Discord account owner and approve the OAuth prompt. Check readiness with: ```bash lium provider status --json ``` Look for `discord_connected: true` and `extra_incentive_eligible: true`. ## Disconnect Discord Use **Disconnect Discord** from the same **Discord** section if you need to unlink the account. After disconnecting, new renter support channels cannot include you until you connect Discord again. ## How renter support channels work When a renter clicks **Request Discord Support** on a pod detail page: 1. Lium checks that the renter has connected Discord. 2. Lium checks that the provider assigned to that pod has connected Discord. 3. Lium creates a private Discord text channel under the configured Lium support category, or returns the existing pod channel if one already exists. 4. Lium grants channel access to the renter, the provider, and the configured support role. 5. Lium posts a pod context message with the pod ID, machine details, Docker image, ports, SSH command, hardware, price, and location. The channel name is generated from the pod name and pod ID. The channel is stored on the pod record, so repeated support requests reopen the same channel instead of creating duplicates. When the pod is removed, Lium attempts to delete the stored Discord channel during pod cleanup. ## Troubleshooting | Issue | What to check | | --- | --- | | **Connect Discord** fails immediately | The Provider Portal backend Discord OAuth configuration may be missing. Contact Lium support. | | OAuth returns **Discord Connection Failed** | Retry the connection. If it repeats, contact Lium support with the error shown in the notification. | | The portal says Discord is connected but renters still cannot request support | Confirm the pod is running on a machine owned by the same provider hotkey that has Discord connected. | | You cannot see a support channel after a renter says they requested one | Confirm you are signed in to Discord with the connected account. If you recently disconnected/reconnected, ask the renter to retry from the pod detail page. | --- # Managing Nodes # Managing Nodes This guide covers all aspects of node management, from adding new machines to configuring settings and handling customer requests. :::tip Browserless / agent-driven node management Every action in this page is also available from the [`lium provider`](../../developers/cli/reference/provider.md) CLI — same surface as the portal frontend, scriptable, with `--json` envelopes for AI agents and CI. For example: `lium provider node add --gpu-type "NVIDIA H200 NVL" --gpu-count 8 --ip --port 8080 --price 1.85 --yes`. ::: ## Adding New Nodes ### Prerequisites Before adding a node, ensure: 1. **Node Software Running**: Your GPU machine has the node program running 2. **Network Access**: Machine is accessible from the internet 3. **Provider connected**: Your provider (self-hosted, or the [Lium.io Central Provider Server](../provider-configuration.md)) is running and connected to the portal ### Registration Process 1. **Access Add Node Modal**: - Navigate to the Nodes page - Click the **"Add Node"** button - A popup modal will appear 2. **Fill Required Information**: - **GPU Type**: Select from available options (RTX 4090, A100, etc.) - **GPU Count**: Enter the number of GPUs in the machine - **Machine IP Address**: Public IP address of your node - **Port**: Port number where the node program is running - **GPU Price**: Hourly price per GPU (in USD) 3. **Submit Registration**: - Review all information for accuracy - Click **"Add Node"** to register ## Synchronization The sync controls below apply when you run a **self-hosted provider**. If you've opted into the [Lium.io Central Provider Server](../provider-configuration.md), Lium handles sync for you and these buttons are not needed. ### Sync From the self-hosted provider Sync nodes from your self-hosted provider to the portal: - **Purpose**: Import node configurations from your self-hosted provider - **When to Use**: After setting up nodes directly on your server - **Process**: Click **"Sync From Provider Server"** button - **Result**: Nodes configured on your server appear in the portal :::info Note When you first login on portal, the sync process will be done automatically for you. ::: ### Sync Into the self-hosted provider Sync node configurations from the portal to your self-hosted provider: - **Purpose**: Export portal configurations to your self-hosted provider - **When to Use**: After making changes in the portal that need to be reflected on your server if your server lost connection at that moment. - **Process**: Click **"Sync Into Provider Server"** button - **Result**: Portal configurations are applied to your self-hosted provider :::info Note The syncing process is done automatically as long as your self-hosted provider is connected to the portal backend and you barely need to do manual sync at all. ::: :::warning Synchronization Always ensure your self-hosted provider is running and accessible before attempting synchronization. ::: ### Post-Registration Steps After successful registration: 1. **Monitor Status**: Check that the node appears as active 2. **Configure Settings**: Set notice periods and other parameters ## Price Management ### Updating Prices 1. **Access Price Settings**: - Navigate to your node details page - Click the **"Update Price"** button - Enter the new hourly rate per GPU (in USD) in the **GPU Price** field - Note: The **Machine Price** field is deprecated and will be ignored 2. **Price Considerations**: - **Minimum Viable Price**: Set prices that cover your operational costs - **Competitive Pricing**: Research market rates to stay competitive - **GPU Count Impact**: Total rental cost = GPU Price × Number of GPUs rented :::info Price Updates Price changes take effect immediately and apply to all new rental requests. Existing active rentals continue at their original rate until they end. ::: ## Pausing New Rentals Use **Pause New Rentals** when a node is currently rented and you want it to stop accepting future rentals after the current renter is done. This does not stop, delete, or interrupt the active pod. ### When to use it - You plan to take the machine offline after the current rental ends - You want to drain a node before maintenance or hardware changes - You want the current renter to finish normally, but do not want another renter to immediately take the node :::info Rented nodes stay active While the pause is pending, the current rental remains active and the node is still checked by validators. The node is removed from new-rental availability immediately, including partial GPU rentals on GPU-splitting nodes. ::: ### Request a pause 1. Open the node details page in the Provider Portal. 2. Click **Pause New Rentals**. 3. Confirm the action. After the request is saved: - The node no longer appears as available for new rentals. - Existing pods on the node continue to run. - The node shows a pending pause state until the last active pod on that node ends. ### After the current rental ends When the final active pod on the node is deleted, the pause becomes active. At that point: - The node stays hidden from new-rental availability. - The node is not eligible for unrented availability rewards. - The node can be brought back with **Resume New Rentals**. ### Cancel or resume - Use **Cancel New Rentals Pause** while the pause is still pending and the node is still rented. - Use **Resume New Rentals** after the pause has fully applied and you want the node to accept rentals again. ### Price limits There are two hard constraints on the price you can set: | Constraint | Effect | |---|---| | Price < **0.5×** the GPU model's reference price | Portal rejects with HTTP 400 — the price is not saved | | Price > **4×** the GPU model's reference price | Portal rejects with HTTP 400 — the price is not saved | Reference prices change as the platform adds GPU models or rebalances rates. Always fetch the current values from the machines endpoint: ```bash curl https://lium.io/api/machines ``` The response lists each supported GPU model with its reference price in USD/hr. Multiply your GPU model's reference price by 0.5 (floor) and 4.0 (ceiling) to get the allowed range. ## GPU Splitting GPU splitting allows you to rent individual GPUs from a single node to multiple customers simultaneously, enabling more efficient resource utilization and flexible pricing. ### Prerequisites GPU splitting requires Docker to run on `overlay2` with an XFS backing filesystem (`pquota`, `ftype=1`). Complete the host-side setup from [Docker Storage Setup](../nodes/docker-storage.md) before continuing — the **Edit** button on the GPU Splitting panel stays disabled until preflight reports the host as ready. ### Setting Minimum GPU Count for Rental After configuring the prerequisites, you must set a minimum GPU count to enable GPU splitting on your node. 1. **Access GPU Splitting Settings**: - Navigate to your node details page in the portal - Locate the **"GPU Splitting"** section in the details panel - If your node meets all prerequisites, you will see an **"Edit"** button 2. **Configure Minimum GPU Count**: - Click the **"Edit"** button - Set the minimum GPU count (must be greater than 1) - Save your changes :::info Minimum GPU Count GPU splitting requires a minimum GPU count to be configured. If the minimum GPU count is not set, the node will not support GPU splitting, and customers will only be able to rent all GPUs on the node. Once configured, customers can rent any number of GPUs from the minimum up to the total available GPUs on the node. **Disabling GPU Splitting:** You can remove the minimum GPU count to disable GPU splitting. However, you can only remove it when there are no pods currently renting a partial number of GPUs. ::: ## Notice Period Management ### Understanding Notice Periods Notice periods allow you to schedule maintenance for rented machines while keeping customers informed: - **Requirement** - Notice period should be scheduled 24 hours ahead. - Maximum period is 60 mins - **Purpose**: - Give customers advance warning of planned downtime - Give providers a safe maintenance window for the machine ### Setting Notice Periods - Go to node management page - Find the **"Notice Period"** section - Select the start time. (minimum 24 hours later from now) - Enter desired notice period (maximum 60 mins) :::info Note Once the notice period is scheduled, it will send an email to the customer who rented the machine and let them know. ::: ## Notify Machine Request from Customer Respond to customer requests efficiently: 1. **Access Notify Feature**: - Go to your node details page - Click the **"Notify"** button in the top-right corner - A modal window displays available machine requests 2. **Select and Notify**: - Review available machine requests - Select appropriate requests based on your node's capabilities - Click **"Notify"** to send notification to the renter ## Force-closing a pod (Mark broken) ### Why this exists & when to use it Sometimes a pod's container disappears or dies on your host machine, but Lium still shows the pod as **rented**. When that happens the node gets stuck: the rental isn't actually running, so the node keeps scoring **zero incentive**, and you have no way to clear the dangling pod on your own — leaving the machine unusable even after a penalty has been applied. Force-closing is the **recovery path** for exactly this situation. By marking the dead pod broken you accept the provider penalty and refund the renter what they paid; in return, the pod is cleared and your node returns to the rentable pool so it can be rented and **earn its score again**. Use it when: - A rented pod's container is **gone or dead on the host** but still shows as rented/running in Lium, **and** - You have no other way to recover the node. It's a recovery tool, not a routine action — it costs you a penalty and disrupts the renter, so only reach for it when the node is genuinely stuck. ### How to do it 1. Open the node's details page in the Provider Portal. 2. In the **pod status table**, find the pod and click **Force close / Mark broken**. 3. Read the confirmation, then click **Mark as broken**. ### What happens - A **provider penalty** is applied: the rental earnings for that pod are declined from you and **credited back to the renter** for the disruption. - The pod's GPUs are **released back to the rentable pool**, so the node can be rented again and recover its score. - For the renter, the pod turns to **BROKEN** — they're notified they were credited and the pod can only be deleted. :::warning The container is NOT torn down for you Marking a pod broken only updates billing and state. **The platform does not stop the container** — you must kill it yourself on the host machine. Otherwise it keeps running and consuming resources. ::: :::info Tier affects the penalty On **Secure** nodes a force-close applies the penalty above. On **Spot** nodes there is **no penalty** (renters already accept that Spot pods may be interrupted). See [Node Tier: Secure vs Spot](./node-tier.md). ::: ## Node Deletion ### When to Delete a Node Consider deleting a node when: - **Hardware Issues**: Machine is no longer functional - **Upgrade Plans**: Replacing with better hardware - **Cost Optimization**: Reducing operational costs - **Relocation**: Moving to a different location ### Deletion Process 1. **Delete Node**: - Delete from the portal ## Troubleshooting ### Common Issues #### Node Not Appearing - **Check Registration**: Verify all required fields were filled correctly - **Network Connectivity**: Ensure machine is accessible from the internet - **Provider connected**: Confirm your provider is running and connected (self-hosted, or [Lium.io Central Provider Server](../provider-configuration.md) toggle on) - **Validation**: Wait for validation process (minimum 15 minutes) to complete --- # Node Tier: Secure vs Spot # Node Tier: Secure vs Spot By default every node is **Secure** — it earns [subnet incentive](../rewards/emission.mdx) and is subject to penalties if a rental is interrupted. Providers can switch any non-rented node to **Spot**: | | Secure (default) | Spot | |--------------------------------|------------------|------| | Rental income | yes | yes | | Subnet incentive | yes | no | | Penalty on rental interruption | yes | no | | "Spot" badge shown to renters | no | yes | ## Requirements - The node must not be currently rented ## Changing the tier 1. Open the node details page in the Provider Portal. 2. Click **Change Tier** in the top toolbar (between *Add Notice Period* and *Remove*). 3. Select **Secure** or **Spot** and confirm. ![Change Tier modal](./assets/spot-tier-modal.png) The current tier is shown in the **General** section as the `TIER` badge. :::info Renter visibility Renters see a **Spot** badge on these nodes in the rental UI and are aware the pod may be interrupted without compensation. ::: --- # Monitoring Nodes # Monitoring Nodes This guide covers monitoring your nodes, including status tracking and performance metrics. ## Overview Monitor your nodes to ensure optimal performance and maximize earnings: - **Status Monitoring**: Track node health and availability - **Performance Metrics**: Monitor key performance indicators - **Grafana Integration**: Access detailed logs and dashboards ## Node Status Monitoring ![Node Status Monitoring](../assets/executor-status-monitoring.png) ### Primary Status Indicators Monitor these key status indicators to ensure your nodes are operating correctly: #### 1. Rental Check Status The rental check status indicates the verification state of your node: - **RENTAL_CHECK_SUCCESS**: ✅ Rental verification completed successfully - Node is properly configured and ready for rentals - All system checks have passed - **RENTAL_CHECK_FAILED**: ❌ Rental verification failed due to issues - Configuration problems detected - Network connectivity issues - Hardware or software problems - **RENTAL_CHECK_PENDING**: ⏳ Rental verification is queued for processing - Verification request submitted - Waiting for system processing - **RENTAL_CHECK_IN_PROGRESS**: 🔄 Rental verification is currently being performed - System is actively testing the node - Verification process in progress - Wait for completion before taking action #### 2. Active Status Indicates whether the machine is currently operational: - **Yes**: ✅ Machine is running and available for rental - **No**: ❌ Machine is offline or not responding #### 3. Rent Status Shows the rented status: - **Yes**: 🟢 Machine is already rented by customer - **No**: 🔴 Machine is not rented yet ## Performance Monitoring ### Key Performance Indicators (KPIs) Track these metrics to optimize your node performance: #### 1. Uptime Metrics - **Active Time**: Time the node has been active in the system - **Availability**: Percentage of time the node is online and operational #### 2. Rental Price - **Current Rate**: Your current rental price per hour - **Market Position**: How your pricing compares to similar nodes ### Real-time emission and validator scores The Provider Portal does not display sub-day TAO emission. For live emission, check the external [TaoMarketCap → Subnet 51](https://taomarketcap.com/subnets/51/miners) for your hotkey. For real-time validator scoring (the upstream input to emission), Lium exposes a [Grafana dashboard suite](../grafana.md) that providers can read directly. The same dashboards are embedded inside the Provider Portal on each node's detail page (the **Grafana** tab) — that's where you'll see the synthetic-job and scoring telemetry without leaving the portal. Use TaoMarketCap for the resulting TAO emission and the embedded Grafana for the validator-side score signals that produce it. ## Monitoring Tools ### 1. Node Details Page Access comprehensive node information: 1. **Navigate to Node**: Go to `https://provider.lium.io/executors/{executor-id}` 2. **Review Status**: Check all status indicators 3. **Monitor Performance**: Track key metrics ### 2. Grafana Integration Access detailed monitoring dashboards: 1. **Access Grafana**: Click the **"Grafana"** tab on your node details page 2. **View Dashboards**: Access comprehensive monitoring dashboards 3. **Analyze Metrics**: Review detailed performance data ![Node Grafana Monitoring](../assets/executor-grafana-monitoring.png) :::info Grafana Access The Grafana integration provides real-time monitoring of all synthetic jobs. ::: ## Troubleshooting ### Common Issues #### Rental Check Failed - **Check Network**: Ensure stable internet connection - **Verify Software**: Confirm node software is running - **Review Configuration**: Check all settings are correct #### Inactive Status - **Check Services**: Verify node software is running - **Network Issues**: Test connectivity and firewall settings - **Hardware**: Ensure all hardware components are functioning #### Low Rental Activity - **Pricing**: Review and adjust your rental rates - **Availability**: Ensure node is consistently online - **Performance**: Monitor and improve execution quality --- # Default Jobs # Default Jobs Default Jobs are reusable fallback workload profiles for provider nodes. Assign a profile to a node, and Lium can run that workload while the node is idle and unrented. :::info Runtime priority Default Jobs are best-effort idle workloads. Customer rentals always have priority, and a running Default Job can be stopped when the node is needed for a rental or when the assignment is disabled, cleared, changed, or deleted. ::: ![Default Jobs overview](./assets/default-jobs-overview.png) ## What Default Jobs are Use Default Jobs to: - Save a reusable workload profile in the Provider Portal. - Assign that profile to one or more owned nodes. - Bulk-assign a profile across a larger fleet. - Leave a node in the **Lium Default Job** state to let Lium run its own idle workload there and keep the unrented incentive. Default Jobs do not create customer rental records, billing records, or customer pods. ## Default Jobs and the unrented incentive On an idle node, running your own Default Job and earning the [unrented incentive](../rewards/emission.mdx#unrented-pool) are **mutually exclusive**. Per node you choose one: - **Lium Default Job** — Lium runs its own Default Job on the idle node; the node keeps the **unrented incentive**. See [Lium Default Jobs on idle nodes](#lium-default-jobs-on-idle-nodes). - **Miner: your job** — you assign your own Custom Docker Default Job; it earns for you, but the node forfeits the **unrented incentive** while it's assigned. - **Miner Empty Job** — a tiny no-op container holds the node so Lium's idle jobs can't start; the node earns **no unrented incentive** but stays fully rentable. See [Create a Miner Empty Job](#create-a-miner-empty-job). This trade-off applies only while the node is idle. When a customer rents the node, it earns rental income and the rented-pool share exactly as before, regardless of any Default Job assignment. ## Lium Default Jobs on idle nodes When a node is idle and has no Default Job of your own assigned, Lium automatically runs its own Default Job on it. This turns unused capacity into revenue that backs the [unrented incentive](../rewards/emission.mdx#unrented-pool), while the node keeps earning that incentive as usual. How Lium Default Jobs behave: - **Automatic.** Nothing to configure on your side; the rollout is gradual across the fleet. - **Customer rentals always win.** When a customer rents the node, the Lium Default Job stops within seconds and the rental starts as usual. Rentability and rental earnings are not affected. When the rental ends and the node is idle again, a Lium Default Job can start again automatically. - **Your own Default Job takes priority.** If you assign your own Default Job to the node, the Lium Default Job yields and your job runs instead — with the [unrented-incentive trade-off](#default-jobs-and-the-unrented-incentive) described above. - **Visible in Node Activity.** Every Lium Default Job start and stop appears in **Node Activity** in the [Provider Portal](https://provider.lium.io), alongside your rentals and your own Default Job runs. ### Which Lium Default Job runs on your node The Lium Default Job is one of two workloads, chosen automatically per node: - **Dolphin** — a paid AI inference worker for the dphn.ai network, run on capable idle nodes. It is a light workload and runs at full GPU power. - **Pearl** — the fallback Lium runs on nodes that are not eligible for Dolphin. It holds the GPU at sustained full load, so while it runs Lium automatically lowers the GPU power limit to protect the card; the limit is restored as soon as the job stops, and a rented node always runs at full power. Either way, Lium earns from the workload; your node keeps its [unrented incentive](../rewards/emission.mdx#unrented-pool) and stays fully rentable, exactly as described above. There is nothing to enable — eligible nodes are picked up automatically. A node is eligible for Dolphin when it has: - An **Ada-generation or newer GPU** (RTX 40-series/Ada, Hopper, Blackwell — RTX PRO 6000 included). Ampere and older (A100, RTX 30-series, V100, T4) are not eligible. - **70 GB or more of GPU memory** in total across the node's GPUs (a single smaller card does not qualify; several cards can add up). - **NVIDIA driver 580 or newer** (CUDA 13). Older drivers are skipped. - **At least ~130 GB of free disk** and an **x86 (amd64) CPU**. Nodes that do not meet these requirements run the Pearl fallback while idle and still keep the unrented incentive — Dolphin simply does not run on them. To make an eligible card qualify for Dolphin, keep its NVIDIA driver on 580 or newer. ## Before you start Before creating a Default Job, make sure: 1. You are signed in to the [Provider Portal](https://provider.lium.io) with your provider hotkey. 2. Your provider is connected, either self-hosted or through the [Lium.io Central Provider Server](../provider-configuration.md). 3. You have at least one registered node. See [Managing Nodes](./managing-nodes.md). 4. For Custom Docker, you have a public image and a command that keeps the container running. ## Create a Custom Docker profile Use Custom Docker when you want to run your own long-running container on idle nodes. 1. Open **Default Jobs** in the Provider Portal. 2. Click **Create Job**. 3. Select **Custom Docker**. 4. Fill the required fields: - **Job Name**: a readable name for the profile. - **Image** and **Image Tag**: for example, `alpine` and `3.20`. - **Command / Args**: a command that keeps the container alive, such as `sleep infinity`. - **Disk Limit GB**: required, from 1 to 100 GB. 5. Add optional fields if your image needs them: - **Environment Variables** as `KEY=value` lines. - **Volumes** as comma-separated mount paths. - **Internal Ports** as comma-separated container ports. - **Minimum VRAM MB** and **Minimum GPU Count** for node-fit requirements. 6. Click **Save Default Job**. :::warning Keep the process running If the Custom Docker command exits, the Default Job run ends. Use a long-running worker process, service command, or an explicit command such as `sleep infinity` for smoke testing. ::: ![Custom Docker Default Job modal](./assets/default-jobs-custom-docker-modal.png) ## Create a Miner Empty Job Use **Miner Empty Job** to keep both your own and Lium's default jobs off a node while it is idle — it runs a tiny no-op container that holds the node without doing any work, uses no GPU, and has no fields to configure. The simplest way is to assign it straight from the node's dropdown; the portal creates the preset for you: 1. Open **Default Jobs** in the Provider Portal. 2. In the **Node Assignment** table, open the node's **Default Job** dropdown (or tick several nodes to use the bulk dropdown). 3. Select **Miner Empty Job**. For a bulk assignment, then click **Apply to selected**. You can also create it as a saved profile from **Create Job** — it appears alongside Custom Docker — but that is optional, since selecting it in the dropdown creates and assigns it in one step. While a node is assigned Miner Empty Job: - **Lium does not run its own Default Job** on that node. - **The node earns no unrented incentive** while it sits idle — the same trade-off as running your own Default Job, because Miner Empty Job *is* your own job. - **The node stays fully rentable.** Customer rentals and rental earnings are unaffected. To let Lium's Default Jobs and the unrented incentive resume, set the node back to **Lium Default Job**. ## Assign profiles to nodes After saving a profile, assign it in the **Node Assignment** table. ### Assign one node 1. Find the node row. 2. Open the **Default Job** dropdown. 3. Select the profile. ### Assign multiple nodes 1. Select the checkboxes for the target nodes. 2. Choose a profile from the bulk assignment dropdown. 3. Click **Apply to selected**. To remove an assignment, choose **Lium Default Job**. Each node can have only one active Default Job assignment. ![Default Job node assignment](./assets/default-jobs-assignment.png) ## Monitor runs and activity The Default Jobs page shows runtime signals in two places: - **Node Assignment**: current assignment state, runtime status, last event time, runtime today, and failure reason when available. - **Node Activity**: recent customer rentals and Default Job runs across your owned nodes, including [Lium Default Job](#lium-default-jobs-on-idle-nodes) runs started automatically on your idle nodes. Assignment statuses include **Lium Default Job**, **Assigned**, and **Disabled**. Runtime statuses can include **Idle**, **Starting**, **Running**, **Stopping**, **Failed**, or **Blocked**. ![Default Job activity table](./assets/default-jobs-activity.png) ## Disable, clear, or delete Use the smallest action that matches what you want: | Action | Where | Result | |---|---|---| | Disable a profile | Edit the Default Job and turn the **Enabled** switch off | The profile is kept, but the runtime will not start it. Existing assignments remain visible. | | Clear an assignment | Set a node to **Lium Default Job** | The node no longer has an active fallback workload assignment. | | Delete a profile | Open the profile details and click **Delete** | The profile is removed and its active assignments are cleared. | When a running Default Job becomes inactive because the profile or assignment changed, the runtime stops it on a later runner cycle. ## Limits and behavior - Default Jobs run only on active, idle nodes with an enabled profile assignment. - Running your own Default Job forfeits the node's unrented incentive while it is assigned; leaving a node in **Lium Default Job** keeps the incentive. See [Default Jobs and the unrented incentive](#default-jobs-and-the-unrented-incentive). - Customer rentals have priority over Default Jobs. - Heavy Default Jobs that hold the GPU at full load, such as the Lium default Pearl job, run with a reduced GPU power limit to protect the hardware; the limit is restored when the job stops. Lighter workloads such as Dolphin are not power-limited. - A node can have only one active Default Job assignment. - Custom Docker disk limits must be between 1 and 100 GB. - The runtime checks node fit before launch, including free GPUs, optional minimum GPU count, optional minimum VRAM, disk space, and required internal ports. - Runs are best-effort and can be skipped or stopped by runtime conditions. --- # Provider Payouts # Provider Payouts > For total-rewards math, see the [Rewards](../rewards/index.mdx) section. Track your earnings from GPU rental fees with detailed payout history, amounts, validator information, and node specifications. ![Provider Payments](./assets/provider-payments.png) ## What you actually earn Total provider earnings on Subnet 51 are the **sum of two streams**: - **Subnet emission** — standard Bittensor emission, paid in **alpha** and scaled by your validator score. It accrues as **stake on your hotkey** each tempo (~72 minutes), not through this payout page. Track it on [TaoMarketCap → Subnet 51](https://taomarketcap.com/subnets/51/miners). - **Rental fees** — what renters pay for actual GPU-hour usage on your nodes. See [Rental fees](../rewards/rental-fees.mdx) for the current provider share. Lium retains the remainder. The Provider Payouts page on this site shows the **rental-fee** stream. Subnet emission flows through the standard Bittensor mechanism and is not displayed here. ## Overview The Provider Payouts section displays your revenue from GPU rentals. The backend automatically calculates and distributes your portion of rental fees based on rental activity. :::info Payout Timing Rental fee distribution is delayed to allow validation, quality assurance, and the misbehavior penalty window to clear. See [Payouts](../rewards/payouts.mdx) for the current delay value. ::: ## Daily Payout Display The Provider Payouts page shows your earnings organized by day, providing a clear view of your daily revenue from GPU rentals. Each day's entry consolidates all payouts from your nodes for that specific date. ## Payout Information Each daily payout record includes: - **Date**: The specific day for which earnings are displayed - **Total Daily Amount**: Your combined earnings from all nodes for that day - **Payout Status**: Current status of the daily payout (processed, pending, declined) - **Node Breakdown**: Individual earnings from each node for that day - **Validator Information**: Which validators were responsible for scoring your work - **Rental Activity**: Summary of rental periods and GPU utilization for the day ## Understanding Your Earnings Your payouts are calculated based on: - **Rental Duration**: How long your nodes were rented - **GPU Performance**: The performance and availability of your nodes - **Validator Scoring**: Scores assigned by validators for your node performance - **Market Rates**: Current rental rates for your GPU types ## Where payouts arrive This section describes the **rental-fee** payout shown on this page. Subnet emission is a separate stream that accrues as alpha stake on your **hotkey** each tempo (~72 minutes) — see [Payouts → Destination](../rewards/payouts.mdx#destination). - **Destination**: Your **coldkey** — the same coldkey that registered the hotkey on Subnet 51. Rental-fee payouts are delivered as a daily stake transfer through the Subtensor `stake_transfer` API. - **Currency**: **alpha** (the SN51 dynamic-TAO token). Rental rates are listed in USD, but distributions settle in alpha at the prevailing rate. - **Mechanics**: **Automatic** — there is no claim button and no minimum threshold. The backend pushes each day's rental earnings to your coldkey on the standard schedule. - **Schedule**: A given day's rental earnings arrive on your coldkey after the rental day, once the validation and misbehavior windows have cleared (see [Payouts](../rewards/payouts.mdx) for the current delay value). If misbehavior is recorded for that node on that day, the day's payout for that node is declined (see below). :::warning Misbehavior Policy If misbehavior is detected on a node during any day, the entire day's payment for that node will be declined. This ensures network quality and protects customers from poor service. ::: The main misbehavior trigger is **bringing a node down while it has an active rental** without using the supported maintenance flow. To avoid this, schedule maintenance through the **Notice Period** feature ([Managing Nodes → Notice Period Management](./managing-nodes.md#notice-period-management)) — it notifies the renter in advance and protects the day's payout. --- # Machine Requests # Machine Requests Machine requests are rental inquiries from renters seeking specific hardware configurations. ## What are Machine Requests? Machine requests are detailed specifications submitted by renters who need GPU compute resources. ![Machine Requests Page](./assets/machine-requests-2.png) ## How It Works ### 1. View Machine Requests Providers can check all available machine requests from renters in the Machine Requests page. Each request includes detailed specifications of the hardware and configuration the renter needs. ### 2. Respond to Requests When you have a node that matches a renter's requirements, you can respond to the request directly from the Nodes page: ![Notify Renter](./assets/machine-requests-1.png) ### 3. Notification Flow When you select a machine request for your node: - The renter receives an email notification about your available machine - The renter can see your machine in their dashboard - The renter can review your node's specifications and decide whether to proceed with the rental --- # Grafana Dashboards # Grafana Dashboards Lium runs a Grafana instance at **[grafana.lium.io](https://grafana.lium.io/dashboards/f/cfh8umv50c9oga/subnet-51)** that exposes the validator-side telemetry behind subnet 51 — the same data that drives your score, your emission, and your rental activity. Use it for live debugging when the [Provider Portal](./portal/overview.md) doesn't go deep enough. ## Access - **Direct:** open [grafana.lium.io](https://grafana.lium.io/dashboards/f/cfh8umv50c9oga/subnet-51) and sign in. All dashboards covered here live in the **Subnet 51** folder. - **Embedded:** each node's detail page in the [Provider Portal](./portal/monitoring.md) has a **Grafana** tab that surfaces the relevant panels for that one node — no separate login. Most per-provider views accept `miner_hotkey` and `executor_id` template variables in the top toolbar — set them once and every panel filters down to your own machines. ## Dashboards | Dashboard | What it answers | |---|---| | [**Dashboard**](https://grafana.lium.io/d/bejrh6ldw5lhcc/dashboard) | Subnet-wide health: total GPUs, revenue ($/hr), rented-vs-idle ratio, GPUs per provider / validator / node. Use to compare the network against your own footprint. | | [**GPU Demand Analytics**](https://grafana.lium.io/d/gpu-demand-analytics-v1/gpu-demand-analytics) | Supply, demand, and utilization broken down by GPU model. Use before adding hardware to see which models are oversaturated. | | [**GPU model rate**](https://grafana.lium.io/d/eenyghq29eqrke/gpu-model-rate) | Score-portion share by GPU model over time — how much of total subnet score each model is earning. | | [**Job Logs**](https://grafana.lium.io/d/aejriu31349hcb/job-logs) | Raw validator scoring logs per node. Use to find out *why* a specific node scored what it scored. | | [**Penalty Events**](https://grafana.lium.io/d/b7b972f3-a083-415f-886d-ae5b12e2371b/penalty-events) | Every penalty applied to a provider: reason code, action, amount withheld, and the deploy/undeploy failure that triggered it. | | [**Weights**](https://grafana.lium.io/d/cem1ol3fw0740e/weights) | Per-validator weight bars to each provider. Use to confirm a specific validator is actually weighting your hotkey. | ## Panels in each dashboard ### Dashboard Network-wide overview. No template variables. - **Total GPUs**, **Revenue ($/hr)**, **R / I** (rented/idle), **Celium Revenue ($ / hr)** — top-line stats and trends. - **GPUs per Type** — count per GPU model over time. - **GPUs Per Provider**, **Gpus Per Validator**, **Nodes Per Provider**, **Nodes Per Validator** — GPU and node counts broken down per provider and per validator. - **GPUs Per Node (GPU not changed)** and **GPUs Per Node (GPU changed)** — split nodes by whether their GPU set has been swapped between runs. - **Nodes without GPU** — table of nodes reporting zero GPUs. ### GPU Demand Analytics Supply/demand by GPU model. Template variable: `gpu_type` (multi-select). - **Total GPUs in Network**, **Currently Rented GPUs**, **Available GPUs**, **Overall Utilization** — current snapshot. - **Rented GPUs**, **GPU Utilization Rate**, **Supply / Demand** — per-model time series. - **Current Utilization by GPU Type** — table sortable by demand. - **Current Demand** — bar chart of active demand per model. ### GPU model rate - **Score Portion** — time series of each GPU model's share of total subnet score. Useful for spotting which model the validators are currently rewarding most. ### Job Logs Three Loki-style log feeds from `prod_executors`, all keyed by provider hotkey and node ID: - **Job Logs** — every scoring run, including the score and uptime. - **Job Error Logs (Score = 0 AND GPU > 0)** — runs where the node was alive but failed scoring (e.g., synthetic-job error, sysbox missing). - **Machine scrape Error Logs (Score = 0 AND GPU = 0)** — runs where the validator couldn't even read the machine spec (unreachable, auth failure, no GPU detected). If your node is earning nothing, look here first: the second panel tells you the node is up but failing tests; the third tells you it isn't reachable at all. ### Penalty Events Template variables: `miner_hotkey`, `executor_id`. Single table with one row per penalty. For what a penalty withholds and how to dispute one, see [Penalties](./rewards/penalties.mdx): | Column | Meaning | |---|---| | `event_time` | When the penalty fired | | `executor_id`, `miner_hotkey` | Who got penalized | | `reason_code` | Machine-readable reason | | `action` | `apply` for a penalty, `revert` when one was undone | | `dry_run` | `true` = recorded but not enforced | | `amount_withheld` | Reward withheld for this event | | `failure_type` | Linked `deploy_failed` / `undeploy_failed` event from the past 24 h, if any | | `comment` | Free-text comment + extracted error message | ### Weights Template variable: `validator`. Single bar chart of the chosen validator's weight vector across all provider hotkeys. Use it to verify a given validator is actually weighting your hotkey, and at what magnitude relative to peers. ## Common questions → which dashboard | Question | Open | |---|---| | "Why is my node scoring zero?" | **Job Logs** — filter by your hotkey, check the two `Score = 0` panels | | "Were any rewards withheld from me?" | **Penalty Events** — filter by your hotkey | | "Is this validator giving me weight?" | **Weights** — pick the validator | | "Should I add more H100s or 4090s?" | **GPU Demand Analytics** + **GPU model rate** | | "How does my fleet compare to the network?" | **Dashboard** — see GPUs/Nodes per provider | For the resulting TAO emission (which is downstream of these score signals), the [TaoMarketCap subnet 51 page](https://taomarketcap.com/subnets/51/miners) is the canonical source. The Grafana dashboards above are the *inputs* that produce that emission. --- # Troubleshooting # Troubleshooting When a node is online but earning nothing — or a freshly-installed node is failing validator verification — start here. This page covers the two diagnostic surfaces providers reach for first: 1. **Grafana Job Logs** — the canonical place to read raw validator-side error messages for a specific node. 2. **The `verifyx-test` benchmark** — runs the same network, storage, and memory checks the validator runs, on demand. If neither narrows down the problem, the [Provider Portal monitoring view](./portal/monitoring.md) and the [full Grafana suite](./grafana.md) cover the rest of the day-2 telemetry. ## 1. Pull error details from Grafana Job Logs The [**Job Logs** dashboard](https://grafana.lium.io/d/aejriu31349hcb/job-logs) is the per-node scoring log. Every validator run for your node — successful, failed, or unreachable — lands here with the exact error message the validator produced. ### Open the dashboard with your filters 1. Go to [grafana.lium.io → Job Logs](https://grafana.lium.io/d/aejriu31349hcb/job-logs). 2. In the top toolbar, set the two template variables: - **`miner_hotkey`** — your provider SS58 hotkey (the one registered on subnet 51). - **`executor_id`** — the UUID of the specific node you're debugging. You can copy this from the [Provider Portal node detail page](./portal/managing-nodes.md) or from `provider.lium.io/executors/{executor-id}`. 3. Set the time range (top-right) to cover the period you care about. Default is the last 1 hour; widen it if your node has been silent. You can also reach the same view without leaving the portal — each node's detail page has a **Grafana** tab that pre-fills both variables for that one node (see [Monitoring Nodes → Grafana Integration](./portal/monitoring.md#2-grafana-integration)). ### Read the three log feeds The dashboard shows three separate panels, all keyed by the same `miner_hotkey` + `executor_id`: | Panel | When it has rows | What to look for | |---|---|---| | **Job Logs** | Every scoring run | The score the validator gave you, the uptime, and any non-fatal warnings. | | **Job Error Logs (Score = 0 AND GPU > 0)** | The node was reachable but failed scoring | Synthetic-job errors, missing `sysbox-runc` runtime, GPU verification failures, network-metric failures. | | **Machine scrape Error Logs (Score = 0 AND GPU = 0)** | The validator could not even read your machine | SSH / auth failures, unreachable host, no GPU detected, port closed. | **Triage rule of thumb:** - Rows in panel 3 → it is a connectivity or configuration problem (network, firewall, SSH, machine offline). Fix the host before doing anything else. - Rows in panel 2 → the host is up but a specific check is failing. Read the error message; for verifyx network or storage failures, jump to [§2](#2-run-the-verifyx-benchmark-to-reproduce-validator-checks) and reproduce locally. - Rows in panel 1 with non-zero scores → your node is being scored normally; revenue concerns belong in [Rewards](./rewards/index.mdx) and the [rewards FAQ](./rewards/faq.mdx). ## 2. Run the verifyx benchmark to reproduce validator checks When the Job Logs panel 2 reports a network or storage failure, the fastest way to confirm and iterate on a fix is to run the validator's own checks locally. We ship them as a Docker image, `daturaai/verifyx-test`, that runs on the node host itself. Run it before registering a new node, and re-run it any time validator-side network or storage scoring drops. ### Quick start (network only) ```bash docker run --gpus all --rm daturaai/verifyx-test:latest ``` Runs a 10-iteration network benchmark and prints download and upload speeds. Finishes in roughly 2–3 minutes. ### Full benchmark (network + storage + memory) ```bash docker run --gpus all --rm daturaai/verifyx-test:latest python3 benchmark.py --full ``` Runs all three test suites. Expect 5–10 minutes due to the 5 GB storage I/O test. ### What is tested | Test | What it measures | Minimum requirement | |------|-----------------|---------------------| | Network download | PyPI package download speed | **50 Mbps** | | Network upload | Cloudflare speedtest upload | — | | Storage write | Sequential write throughput (5 GB) | **100 GB free space** | | Storage read | Sequential read throughput (5 GB) | — | | Memory allocation | Allocates 75 % of available RAM (8–128 GB range) | **8 GB RAM** | ### Interpreting results The benchmark runs **10 iterations** and reports an Exponential Moving Average (EMA) with `alpha = 0.3` (decay = 0.7), so recent runs carry more weight. The final printed values are the EMA-smoothed numbers — a single slow run will not dominate the result. The last section of the output also shows your hardware stats (total RAM, free disk space, utilization) from the most recent successful run. ### Prerequisites Your host must have the **NVIDIA Container Toolkit** installed so Docker can pass GPU access through the `--gpus all` flag. If you followed the [Node setup guide](./nodes/quickstart.md), this is already configured. ```bash # Verify the toolkit is installed docker run --gpus all --rm nvidia/cuda:12.8.1-base-ubuntu22.04 nvidia-smi ``` ## 3. Common issues | Symptom (Job Logs message or benchmark output) | Likely cause | Fix | |---|---|---| | Download speed below 50 Mbps | ISP throttling or congested uplink | Contact your hosting provider; consider a dedicated port. | | Storage test fails — not enough free space | Less than 100 GB free on the node disk | Free space or resize the volume. See [Docker storage](./nodes/docker-storage.md). | | Memory allocation fails | Less than 8 GB RAM available to Docker | Reduce other processes consuming RAM; upgrade RAM if needed. | | `--gpus all` flag not recognised | NVIDIA Container Toolkit not installed | Follow the [Node setup guide](./nodes/quickstart.md). | | Job Logs panel 3 has rows; SSH / auth errors | Validator cannot reach the node | Check firewall, the node's public port, SSH service, and the credentials registered in the Provider Portal. | | Job Logs panel 2 mentions `sysbox-runc` | Sysbox runtime missing or not the default | Install Sysbox — see [Sysbox setup](./nodes/sysbox.md). It is required; validators reject any node without it. | ## Next steps - [Grafana Dashboards](./grafana.md) — full dashboard catalogue beyond Job Logs (Penalty Events, Weights, GPU Demand Analytics). - [Provider Portal — Monitoring](./portal/monitoring.md) — portal-side status indicators and embedded Grafana tab. - [Node Quickstart](./nodes/quickstart.md) — full setup if you suspect a misconfiguration at the host level. --- # Validators # Validators Lium runs a single validator on Bittensor Subnet 51, operated by the Lium team. ## Lium team validator | Hotkey | |--------| | `5F7X5UpKSr26KU3jKfpLmT8kuKtBNyHhEnfS8xtxPCqCb13p` | ## Running your own validator Other neurons who want to validate Subnet 51 should use the community auditor project: [github.com/Datura-ai/sn51-auditor](https://github.com/Datura-ai/sn51-auditor) --- # Renters # Renters Rent GPUs on [lium.io](https://lium.io) — browse the marketplace, deploy a containerized pod in minutes, pay only for what you use. This section is **UI-first**: every page walks through the dashboard at [lium.io](https://lium.io) and shows where to click. The matching API and CLI live in the [Developers](/developers) section and are referenced from each page in a collapsed "For agents and automation" block. ## Get going | Step | What you do | Page | |------|-------------|------| | 1 | Sign up, add a payment method, paste an SSH key | [Quickstart](/pod-users/quickstart) | | 2 | Pick a machine, configure a pod, click Deploy | [Create a pod](/pod-users/create-pod) | | 3 | Pick the right Docker image (PyTorch, custom, …) | [Templates](/pod-users/templates) | | 4 | Keep your data across pod sessions | [Volumes](/pod-users/volumes) | | 5 | Schedule automatic backups of your work | [Backups](/pod-users/backups) | | 6 | Restore a backup into a new or existing pod | [Restores](/pod-users/restores) | | 7 | Auto-terminate pods so they stop charging | [Scheduled termination](/pod-users/scheduled-termination) | | 8 | Open a private Discord channel with the provider for a rented pod | [Discord support](/pod-users/discord-support) | ## Reference (don't read top-to-bottom) - [Broken pods](/pod-users/broken-pods) — what a **BROKEN** status means (provider force-closed the pod) and how to clear it. - [Pod overview & security model](/pod-users/overview) — what a pod actually is and what the GPU provider can/can't see. - [API Keys](/pod-users/api-keys) — generate keys for the CLI, SDK, MCP, or your own AI agent. - [Pod security](/pod-users/security) — what to never put on a non-CVM pod. - [GPU profiling with ncu](/pod-users/gpu-profiling) — rent a whole-host machine where Nsight Compute works. --- # Quickstart # Quickstart ## Get started in 5 min Goal: a running GPU pod you can SSH into. All steps are in the dashboard at [lium.io](https://lium.io). ### 1. Create an account Sign up at [lium.io/register](https://lium.io/register). Click **Billing** in the sidebar and add a payment method — your account balance shows in the top-left of the dashboard. ### 2. Add an SSH key Click **Access** in the sidebar (key icon, bottom group) → **SSH Keys** tab → **ADD NEW +**. Paste your public key (e.g. `~/.ssh/id_ed25519.pub`) and give it a name. ```bash # Generate a key if you don't have one ssh-keygen -t ed25519 -C "you@example.com" cat ~/.ssh/id_ed25519.pub ``` ![SSH Keys tab in Access page](./assets/access-ssh-keys.png) ### 3. Pick a machine Click **Browse Pods** in the sidebar. Use the right-hand filter rail to narrow by GPU type (RTX 4090, A100, H100, L40, …), GPU count, location, **Confidential Computing** for CVM-isolated nodes, or **GPU Profiling (ncu)** for machines where [Nsight Compute profiling](./gpu-profiling) works. ![Browse Pods list with filters](./assets/browse-pods.png) :::tip Sensitive workload? Toggle the **Confidential Computing** filter on the right. CVM nodes prevent the GPU provider from reading your pod's memory or filesystem. See [Pod security](./security). ::: ### 4. Configure and deploy Click **RENT NOW** on the row you want. On the **Create Pod** page: 1. Tweak the auto-generated **Pod Name** if you want. 2. Confirm the **Template** — `Pytorch (Cuda) - daturaai/pytorch` is selected by default. 3. Confirm your **SSH Key** is in the chip list. 4. (Optional) attach a **Volume** for persistent storage at `/mnt`. 5. (Optional) set **Auto-Termination** so the pod stops billing after N hours. 6. Review the **Total cost** in the right-hand summary. 7. Click **Deploy**. ![Create Pod page with summary panel on the right](./assets/create-pod.png) ### 5. Connect via SSH Open **Your Pods**. Once the card flips to **RUNNING**, click **SEE DETAILS**. Copy the SSH command from the **SSH CONNECTION** field at the top: ```bash ssh root@ -p ``` ![Pod detail page with SSH connection string](./assets/pod-detail.png) You're in. Your home directory is `/root` and is mapped to the pod's local volume. Anything outside `/root` (and `/mnt` if you attached a volume) is **not** persisted across pod restarts. ## Next steps - Save your work between sessions → [Volumes](./volumes) - Avoid losing checkpoints if a pod dies → [Backups](./backups) - Spin up the same environment from one image → [Templates](./templates) - Cap costs with auto-termination → [Scheduled termination](./scheduled-termination) - Drive everything from the CLI / an agent → [API Keys](./api-keys) --- # Create a pod # Create a pod Walk-through of the **Create Pod** page at [lium.io](https://lium.io). You land here when you click **RENT NOW** on a row in **Browse Pods**. ![Create Pod page top — pod name, machine card, template, SSH key](./assets/create-pod-top.png) :::warning Don't put secrets on a non-CVM pod The page itself reminds you: *"Do not upload private keys, seed phrases, API keys, or other secret files to a rented pod. The GPU provider may have access to the pod environment."* If you must run sensitive workloads, filter for **Confidential Computing** machines on Browse Pods. Full model: [Pod security](./security) and [CVM guide](../providers/nodes/cvm.md). ::: ## The fields, top to bottom ### Pod Name Auto-generated like `Proud Machine Cloud`. Click and rename to anything you'll recognize in the **Your Pods** list later. ### Template The Docker image your container is built from. The default is the most-used template for the chosen GPU (`Pytorch (Cuda) - daturaai/pytorch` for most NVIDIA cards), pre-cached on the host so deploys finish in ~1 minute. - Click the **×** on the template card to swap. The picker drawer marks the host's canonical image with a blue **Fast deploy** badge and shows a green **Previously used** badge on templates you've deployed before. - After picking, watch the right-hand summary's **Est. Deploy Time** row. A small **cached** badge there means the image is on this specific node right now (deploy in seconds); no badge means a pull is needed (1–10 min). - Need something specific (vLLM, SD-WebUI, your own image)? See [Templates](./templates) for picking and creating one. ### SSH Key The chips show which keys will be installed in `~/.ssh/authorized_keys` on the pod. The dropdown picks from existing keys; **Add an SSH Key** opens an inline form to paste a new public key — same as **Access → SSH Keys**. You **need at least one SSH key** to reach the pod. ### Volume Optional. Mount an external [Volume](./volumes) at `/mnt` so data survives pod termination. One volume per pod at a time. If you don't attach a volume, anything outside the pod's local mount (`/root`) is gone when the pod is deleted. ### Encrypt local volume This option appears for official templates that support it. Encryption is on by default and stores the pod's local volume (`/root`) as ciphertext on the provider's disk. Turn it off for IO-bound workloads where disk throughput matters more than at-rest protection. Full details: [Encrypted local volumes](./encrypted-volumes). ### Initial Port Count How many TCP ports to forward from the host. Each pod gets a port range; one is used for SSH, the rest are free for your services (Jupyter, vLLM, etc.). The right side shows the host's max — click **Max** to grab them all. ### Auto-Termination (Optional) Type a number of hours and the pod will be deleted automatically when it elapses. Cheap insurance against forgetting. You can also set this later — see [Scheduled termination](./scheduled-termination). ### Install Jupyter Tick the box and the pod's startup script installs JupyterLab and exposes it on a forwarded port. The pod detail page will show the URL. ### Restore (optional) Bring a [backup](./backups) into the new pod's filesystem on first boot. - **Volume path** is fixed to `/root` — the local volume mount where your backup will be extracted. - Click **Select a Backup** to pick one of your existing archives, or tick **Enter backup ID directly** to paste a UUID. Full details and path rules: [Restores](./restores). ### # of GPUs (when GPU splitting is on) If the host supports **GPU splitting**, a count selector appears. The provider sets a minimum (e.g. 2 of 8); you can pick any count between that minimum and the host's total. CPU, memory, and storage are sliced proportionally to the GPU count you take: ``` cpu = total_cpu × rented_gpu_count / total_gpu_count memory_gb = (total_mem_gb - 2) × rented_gpu_count / total_gpu_count disk = free_disk × rented_gpu_count / total_gpu_count volume_limit = int(disk × 2 / 3) storage_limit = int(disk × 1 / 3) ``` On a machine with the **GPU Profiling** badge the selector never appears: profiling machines are rented whole-host only, all GPUs included. See [GPU profiling with ncu](./gpu-profiling). ### Advanced — Skip agent SSH key The platform installs a small **agent key** alongside yours. It powers the web terminal and pod-health checks. Untick **Skip agent SSH key** to remove it — you'll lose the web terminal and we can't tell the GPU provider if SSH breaks. Default: leave it on. ## Right-hand summary The summary panel mirrors your selections — CPU, memory, hard disk, network, estimated deploy time, location, max CUDA driver, GPU cost per hour, GPU count, **Total cost / hour**. Only the **Total cost** number includes splitting; verify before you click **Deploy**. ![Right-hand summary panel with Deploy button](./assets/create-pod-summary.png) ## After deploying Click **Deploy**, and the pod appears in **Your Pods** with status **PROVISIONING** then **RUNNING**. From there: - Copy the SSH command from the **SSH CONNECTION** strip on the pod detail page. - Connect Discord and request a private provider support channel if you need help with the running pod — see [Discord support](./discord-support.md). - Set up automated [Backups](./backups). - [Schedule termination](./scheduled-termination) if you forgot during creation. If a pod ever shows a red **BROKEN** status, the provider force-closed it — you're credited and it can only be deleted. See [Broken pods](./broken-pods). --- # Templates # Templates A **template** is the recipe Lium uses to start your pod's container. It bundles a Docker image, an image tag, environment variables, port exposures, a startup command, and the volume mount path — everything that decides "what software is on this GPU when it boots." Pick a good template and your pod is ready to run training/inference 30 seconds after **Deploy**. Pick a custom one and your environment is reproducible across teammates and runs. ## Why templates matter | Benefit | What it means in practice | |---|---| | **Faster cold boots** | The picker shows a **Fast deploy** badge on the canonical image for the chosen GPU class, and the right-hand summary shows a **cached** badge next to **Est. Deploy Time** when the exact image is already on the chosen node — your container starts in seconds instead of minutes. | | **Reproducible environments** | Pin a CUDA / PyTorch / driver combo once, reuse it on every pod for the project. | | **Shareable** | Make a template public and your team (or the whole community) can deploy with one click. | | **Agent-friendly** | The CLI, SDK, and MCP all accept a template ID — your AI agent can spin up identical pods on demand. | ## Browsing & picking a template Two places to choose: 1. **Templates page** — Click **Templates** in the sidebar. The page splits into **My templates** (yours) and **Browse templates** (everyone's). Search by name, filter by category. 2. **Inline on Create Pod** — When you click **RENT NOW**, the **Template** card on the Create Pod page is pre-filled with the most-used template for that GPU. Click the **×** to swap. ![Templates list page with My templates and Browse templates](./assets/templates-list.png) Each card shows: - Image name (e.g. `daturaai/pytorch`) and tag. - The template's **Category** (PyTorch, TensorFlow, Custom, …). - Labels (only render in the **Create Pod** template picker — the standalone `/templates` page doesn't pass the host context the badges need): - **Previously used** — green chip on a template you've deployed before. - **Fast deploy** — blue chip on the template whose `docker_image:tag` matches the canonical default image for the chosen host's GPU + driver. Picking it means the host is most likely to have the image cached and you skip the long pull. - An **EDIT** button on templates you own. :::tip Two different speed signals **Fast deploy** is per-GPU-class: "this is the canonical image for this card." **cached** (the small badge that shows up in the right-hand summary next to **Est. Deploy Time**) is per-node: "this exact image is already on this specific machine right now." Both pointing at the template and the host means a near-instant deploy. Pulling a non-cached image adds 1–10 minutes. ::: ## Creating a custom template When the official templates don't fit (custom CUDA stack, your team's image, vLLM/SGLang launcher, …) make your own. ### Steps in the UI 1. **Templates** sidebar → **CREATE A NEW TEMPLATE** (top-right blue button). 2. Give it a **Template name**. 3. Set the visibility toggle: **Private** (only you can see it) or **Public** (everyone can use it). 4. Fill in the **Configure** tab: - **Category** — picks the icon and sorts the template into the matching tab. PyTorch / TensorFlow / Custom / … - **Docker Credential** — only required if your image is in a private registry. Click **+** to add credentials. Lium passes them to the host at pull time only. - **Container Image** — required, e.g. `pytorch/pytorch`. - **Container Image Digest** — optional. Pin to a specific `sha256:…` so the image you tested is the image that runs. - **Container Image Tag** — required, e.g. `2.2.0-cuda12.1-cudnn8-runtime`. - **Container Start Command** — runs as the container's CMD. Leave blank to use the image's default; or set something like `bash -c "pip install -r /root/requirements.txt && sleep infinity"`. - **Entrypoint** — overrides the image's ENTRYPOINT. Most users leave this blank. - **Volume** — the mount path for the local persistent volume **inside** the container. Default `/root`. **Most users should leave this as `/root`** — backups, restores, and the SSH home all assume this path. - **Internal Ports** — ports your software listens on (e.g. `22` for SSH, `8888` for Jupyter, `8000` for vLLM). Lium maps each to a unique external port on the host. - **Environment Variables** — `KEY=value` pairs injected into the container. 5. (Optional) The **Read me** tab takes Markdown that other users see when they pick your template. Document the ports, env vars, and how to run it. 6. **SAVE**. The template appears in **My templates** and is selectable on Create Pod. ![Create new template form: Configure tab fields](./assets/templates-create.png) ### Verification status Templates carry a `status` field — `CREATED`, `UPDATED`, `VERIFY_PENDING`, `VERIFY_FAILED`, or `VERIFY_SUCCESS` — and an associated `verification_logs` text field. **Today, automatic verification on user-created templates is disabled.** When you save a template (private or public), the backend marks it `VERIFY_SUCCESS` straight away and the template is immediately usable. There is no test-deploy / SSH-check pipeline gating your template's availability. That means **the burden is on you to make sure the template actually boots cleanly** before relying on it for important work. Sanity-check it locally: ```bash docker run --rm -d : docker ps # the container should be in the list, not exited ``` …or just deploy a cheap pod (RTX 3090, 1× GPU) with the template and confirm SSH lights up. If automatic verification is re-enabled in the future, the same `status` / `verification_logs` fields will carry the result. ### Common pitfalls - **Container exits immediately.** Your start command must keep PID 1 alive. End it with `sleep infinity` (or run a daemon in the foreground) or the pod will be marked failed. - **Image not public, no credentials.** The pull will fail. Add Docker credentials via **Access → Docker Credentials**, then attach them in the template's Docker Credential field. - **Volume mount changed away from `/root`.** Backups and Jupyter still assume `/root`. Only change this if you really know what you're doing. - **No port for SSH.** Add `22` to **Internal Ports**, otherwise the agent can't open a connection and your **SSH CONNECTION** field on the pod page will be empty. - **Docker-in-Docker template + external Volume.** These work together — attach an external [Volume](./volumes#attach-a-volume-to-a-pod) to a DinD template and nested Docker keeps running. A few nodes are still unable to serve both; when that happens the deploy log says Docker-in-Docker was turned off, and moving the pod to another node fixes it. ## Editing or deleting a template In **Templates → My templates**, click **EDIT** on your card. Same form as creation. Changes apply to **future** pod deployments only — pods already running keep their original template. Public templates that other users have rented from cannot be deleted while pods are still using them. ## What gets baked into a deployed pod When you deploy from a template, the platform records *everything* in the template at that moment onto the pod. Editing the template later does not retroactively change running pods. To re-deploy the latest version of a template into a fresh pod, just hit **RENT NOW** again.
For agents and automation: API + CLI Templates are first-class objects in the API. Use these when you're driving Lium from a script, the CLI, or an LLM agent. You'll need an [API key](./api-keys). ```bash # List all templates (yours + public) curl https://lium.io/api/templates \ -H "X-API-Key: $LIUM_API_KEY" # Filter by GPU model + driver — the response's `compatibility` field tells you # whether the template's CUDA requirements line up with the target GPU. curl "https://lium.io/api/templates?gpu_model=NVIDIA+H100+80GB+HBM3&driver_version=535.86.10" \ -H "X-API-Key: $LIUM_API_KEY" # Get one curl https://lium.io/api/templates/ \ -H "X-API-Key: $LIUM_API_KEY" # Create one curl -X POST https://lium.io/api/templates \ -H "X-API-Key: $LIUM_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "name": "My PyTorch Template", "docker_image": "pytorch/pytorch", "docker_image_tag": "2.2.0-cuda12.1-cudnn8-runtime", "volumes": ["/root"], "environment": {"PYTHONUNBUFFERED": "1"}, "internal_ports": [22, 8888], "startup_commands": "pip install -r /root/requirements.txt && sleep infinity" }' ``` The CLI wraps the same calls (`lium templates list`, `lium templates create …`). See the [`lium templates` reference](/developers/cli/reference/templates).
--- # Volumes # Volumes A **Volume** is persistent storage that lives outside any single pod. Attach one to a pod and the data survives pod termination, can be moved to another pod later, and never has to fit on the host's local disk. When attached, the volume is mounted at **`/mnt`** inside your container. Use it for datasets, model weights, training checkpoints, evaluation outputs — anything you don't want to lose when a pod is deleted. > Volumes are backed by S3-compatible object storage. Throughput is fine for streaming/loading, but sustained random I/O is slower than the pod's local disk. For tight inner loops, copy the working set into `/root` and write results back to `/mnt`. ## Create a volume In the dashboard: 1. Click **Volumes** in the sidebar. 2. Click **ADD NEW +** (top-right). 3. Enter a **Name** (letters, digits, hyphens, underscores, dots, spaces, and `,()!?/` — max 100 chars). Description is optional. 4. **Create**. ![Volumes page with rate calculator and existing volume](./assets/volumes-list.png) The page shows your volumes with their **Size**, **Files**, and **Date created**. The **Current Rate** widget at the top is a calculator — set Storage Amount × Time Period to estimate hourly cost. ## Attach a volume to a pod When deploying a pod (Browse Pods → **RENT NOW**): 1. Find the **Volume** field on the Create Pod form. 2. Open **Choose Volume** and select an existing volume — or click **Add a Volume** to create one inline. 3. **Deploy**. The volume mounts at `/mnt` automatically. Forgot to pick one at deploy time? Open the pod detail page and click **Volume** in the action bar. It works only while the pod has no volume yet. :::warning A pod's volume is permanent Pick the right volume before you attach it. Once a pod has one, the **Volume** action is greyed out and the API rejects the change, so it can't be swapped for a different volume or detached while the pod runs. To stop using a volume on a pod, terminate the pod. The data on the volume survives that. ::: :::info One volume, several pods The reverse is not restricted: the same volume can be attached to more than one pod at a time, including pods on different nodes. Each pod mounts the same bucket, so they all see the same files. Nothing coordinates those pods. Two pods writing the same file will not merge their changes, the last write wins. And the check that stops you deleting a volume that is in use only looks at one pod, so with several pods attached it can let the deletion through. Deleting a volume destroys the data for every pod still using it. ::: :::info Volumes work with Docker-in-Docker You can attach a volume to a Docker-in-Docker template — nested Docker keeps working and `/mnt` stays writable. Files on the volume belong to `root` inside the pod, and permissions you set on them are stored with the objects. On a small number of nodes the volume still turns Docker-in-Docker off. The deploy log says so when it happens; move the pod to another node if you need both. ::: ## Use it from inside the pod ```bash # SSH into the pod ssh root@ -p # Volume is at /mnt ls /mnt df -h /mnt # confirm it's mounted # Suggested layout /mnt/ ├── datasets/ # input data ├── models/ # trained weights you want to keep ├── checkpoints/ # in-progress training state └── outputs/ # results, evaluation reports ``` Tip: keep the working set on `/root` (the local mount, fast disk) during training, then `rsync /root/checkpoints/ /mnt/checkpoints/` between epochs. ## Billing - Charged **per-GB / per-hour** based on the actual data stored, not the allocated capacity. - Continues to accrue while the volume is **detached** — your data stays available. - Stops only when you **delete** the volume (irreversible). - Pay-as-you-use AWS pass-through — no markup. The Volumes page calculator at the top of the list shows the current `$/GB/hour` rate. ## Delete a volume In the **Volumes** list, click the trash icon on the row. **This permanently destroys the data** — there is no undo and no archive. If you might want the data back, take a [Backup](./backups) of the relevant pod first, then delete the volume. ## Limits & rules | Rule | Detail | |---|---| | Mount point | Always `/mnt` | | Volumes per pod | 1, fixed once attached | | Pods per volume | No limit — they all mount the same bucket, with no locking | | Size cap | None — volumes scale automatically | | Name pattern | `[a-zA-Z0-9\s\-_.,()!?/]+`, 1–100 chars |
For agents and automation: API + CLI ```bash # List your volumes curl https://lium.io/api/volumes \ -H "X-API-Key: $LIUM_API_KEY" # Get one curl https://lium.io/api/volumes/ \ -H "X-API-Key: $LIUM_API_KEY" # Create one curl -X POST https://lium.io/api/volumes \ -H "X-API-Key: $LIUM_API_KEY" \ -H "Content-Type: application/json" \ -d '{"name": "my-dataset", "description": "Training data v2"}' # Delete one (data is gone — no recovery) curl -X DELETE https://lium.io/api/volumes/ \ -H "X-API-Key: $LIUM_API_KEY" ``` CLI: `lium volumes list / new / rm` ([reference](/developers/cli/reference/volumes)). Get an [API key](./api-keys) first.
--- # Encrypted local volumes # Encrypted local volumes Lium can store the pod's local volume at `/root` as ciphertext on the provider's disk. Inside the running pod, it looks like a normal filesystem. ## What it protects Encryption covers data at rest. A leftover volume, a disk image, or a drive that leaves the datacentre is unreadable by itself. It does not protect a running pod. A provider with host access can read `/root` while your pod is up, inspect its memory, or capture the key. If the provider must never see your workload, rent a [CVM-enabled machine](../providers/nodes/cvm.md) instead. Full model: [Pod security](./security). ## Using it On the Create Pod page, **Encrypt local volume** appears for official templates that support it. Encryption is on by default. The API, SDK, and CLI request encryption by default too. To turn it off: - API: `"enable_volume_encryption": false` in the rent request. - SDK: `Lium.up(..., enable_volume_encryption=False)`. - CLI: `lium up --no-volume-encryption`. The template decides whether the UI shows the option. At deploy time, Lium checks the template's Docker image. If the image does not support encryption, the pod starts with a plain volume instead of failing the rental. ## Tradeoffs - Disk is slower. On an A6000 node with an EPYC 7763, a mixed 70/30 random read/write test went from 1008 to 625 MiB/s read and 435 to 271 MiB/s write. That is roughly 38% less throughput. Your numbers depend on the host and workload. - Backups need a running pod. [Backups](./backups) and [restores](./restores) of an encrypted volume only work while the pod is up. ## Status on the pod page After deploy, **Volume encryption** shows what actually happened: - **Enabled**: `/root` is encrypted. - **Unavailable**: the image did not support encryption, so the pod got a plain volume. - **Disabled**: encryption was turned off in the request. - **Setup failed**: the image supported encryption, but the encrypted mount did not come up.
For agents and automation The volume uses [gocryptfs](https://nuetzlich.net/gocryptfs/). The Docker volume holds ciphertext, and the decrypted view is mounted at `/root`. An image qualifies by shipping gocryptfs and carrying the `lium.volume_encryption.enable=1` Docker label. The validator derives a passphrase per pod and hands it to the container at mount time. The passphrase is never written to the volume.
--- # Backups # Backups Lium runs a scheduled backup job for any pod you configure: it ZIPs a directory inside the pod, uploads the archive to AWS S3 in a bucket scoped to your account, and prunes archives older than the retention window. Use it to protect training checkpoints, model weights, and any work-in-progress that would be expensive to redo if the pod went away. ## How it works 1. You pick a **Backup Path** inside the pod and a **Frequency** + **Retention**. 2. The platform installs a cron job on the pod. 3. On every tick, the job ZIPs the directory and uploads the archive to S3. 4. Archives older than the retention window are auto-deleted. 5. Backup status (started / completed / failed / deleted / expired) shows up on the pod's **Backups** tab and on the global **Backups** page. ## Backup path: the rule that matters **The backup path must be a subfolder of the pod's local volume mount.** The local mount is `/root` for every standard template — that's also your SSH home directory and where templates land your code. The Backup Configuration modal calls this out at the top: *"Volume Path: `/root`. The backup path should be a subfolder of the volume path to ensure data persistence."* ### ✅ Valid backup paths ``` /root # entire local volume /root/models # one subdirectory /root/checkpoints # another subdirectory /root/project/data # nested subdirectory ``` ### ❌ Invalid backup paths ``` /home/user/documents # not under the local volume mount /tmp/backup # tmpfs, not persistent — erased on pod restart /var/log # system directory outside the local volume /mnt # external Volume mount — outside the local volume /mnt/datasets # same — external Volumes are NOT in the backup scope .. # path traversal, rejected by the API ``` :::info The /mnt distinction **External [Volumes](./volumes)** mounted at `/mnt` are *not* backed up by this system — they have their own durability via S3 already, and including them would double-bill your storage. Backup is for the **local** pod volume at `/root`. ::: If a template uses a non-default volume mount (rare — some custom templates change `Volume` in the template definition), the rule becomes "subfolder of *that* path." The modal's banner reflects whatever the pod's mount actually is. ## Configure a backup in the UI From the pod detail page: 1. Click **Backup** in the top-right action bar. 2. The **Backup Configuration** modal opens. 3. **Backup Path** — defaults to `/root`. Narrow it (e.g. `/root/checkpoints`) if the rest of `/root` is large and you don't need it backed up. 4. **Backup Frequency** (hours, 1–168). 5. **Retention** (days, 1–365). 6. **Save**. ![Backup Configuration modal on the pod detail page](./assets/backups-config-modal.png) Tweak it later by clicking **Backup** again — the modal pre-fills with the existing config. ## See your backups - **Per pod** — pod detail page → **Backups** tab. Shows recent runs and a **Restore From Backup** section listing restore jobs. - **Globally** — sidebar → **Backups**. Filter by Type (Manual / Automatic), Status (Completed / Failed / Deleted / Expired), Pod ID, or search by name. Use **Delete All COMPLETED** to clean up. ![Global Backups page with filters and history](./assets/backups-list.png) ## Frequency cheatsheet | You're doing | `backup_frequency_hours` | |---|---| | Active training run, costly to lose | `6` | | Regular daily workload | `12`–`24` | | Stable / inference-only project | `24`–`72` | | Long-term archival | `168` (weekly) | Pair with **Retention**: longer retention = bigger S3 bill. 7–30 days is typical for active work. ## What gets stored where The archive is `.zip`, uploaded to a per-account S3 prefix the platform manages. You don't get raw S3 credentials — pull archives back through the [Restores](./restores) flow or via the API. ## Billing Backups are billed transparently with no markup: - **Storage** — by the actual compressed archive size, charged hourly. - **Transfer** — pay AWS rates for the upload. - **Retention runs** — free. - **Pod stopped or deleted** — your archives stay live in S3 for the full retention window. Storage cost continues. To stop charges entirely: delete the backup configuration **and** delete the existing archives (use **Delete All COMPLETED** on the Backups page, or delete by ID via the API). ## Triggering a manual backup From the pod detail page **Backups** tab, the next scheduled run is governed by the cron job. To kick off a one-off backup right now, use the API call below — set `is_manual: true`. (A "Run now" button is on the roadmap.) ## Update or delete a configuration In the dashboard, click **Backup** on the pod detail page and change the fields, or delete the configuration entirely from there.
For agents and automation: API + CLI ```bash # Create a backup configuration curl -X POST https://lium.io/api/backup-configs \ -H "X-API-Key: $LIUM_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "pod_id": "", "backup_path": "/root/checkpoints", "backup_frequency_hours": 6, "retention_days": 30 }' # List configurations curl https://lium.io/api/backup-configs \ -H "X-API-Key: $LIUM_API_KEY" # Get the config for one pod curl https://lium.io/api/backup-configs/pod/ \ -H "X-API-Key: $LIUM_API_KEY" # Update curl -X PUT https://lium.io/api/backup-configs/ \ -H "X-API-Key: $LIUM_API_KEY" \ -H "Content-Type: application/json" \ -d '{"backup_frequency_hours": 12, "retention_days": 90, "backup_path": "/root"}' # Delete (archives in S3 are kept for the retention window) curl -X DELETE https://lium.io/api/backup-configs/ \ -H "X-API-Key: $LIUM_API_KEY" ``` ### Configuration schema | Field | Type | Range | Notes | |---|---|---|---| | `pod_id` | UUID | — | Pod whose volume gets archived | | `backup_path` | string | ≤ 500 chars | Must be under the pod's local volume mount (default `/root`); `..` rejected | | `backup_frequency_hours` | int | 1–168 | 168 = weekly | | `retention_days` | int | 1–365 | Older archives auto-deleted | CLI shortcuts: `lium bk now / set / show / restore / rm / logs` ([reference](/developers/cli/reference/bk)). Get an [API key](./api-keys) first.
--- # Restores # Restores A **restore** pulls a [backup](./backups) archive out of S3 and extracts it into a directory inside a target pod. Use it to: - Recover from a pod that crashed or got terminated. - Migrate work from one pod to a different (often cheaper or different-GPU) machine. - Boot a new pod that already has yesterday's checkpoints, datasets, or weights ready to go. ## Two ways to restore ### Option 1 — Restore into a running pod Use this when your current pod is fine and you want yesterday's data back, or when you're moving data into a pod that's already running. 1. Open the pod's detail page. 2. Click the **Backups** tab. 3. Find the backup row you want and click **Restore**. 4. The modal asks for: - **Target pod** — the pod where the archive will be extracted. - **Restore Path** — a subfolder of the target pod's local volume mount (`/root` by default). 5. **Confirm**. The archive downloads from S3, extracts into `restore_path`, and progress shows up on the pod's **Backups** tab → **Restore From Backup** section. ![Pod detail Backups tab showing restore section](./assets/backups-tab-pod.png) ### Option 2 — Restore on a new pod's first boot Use this when you're starting a fresh pod and want it to come up *with* the backed-up data already in place. 1. **Browse Pods** → click **RENT NOW** on a row. 2. On the Create Pod page, scroll to **Restore** near the bottom. 3. **Volume path** is fixed to `/root` (or whatever the chosen template's local volume mount is). 4. Click **Select a Backup** to pick from your archives, or tick **Enter backup ID directly** to paste a UUID. 5. Click **Deploy**. The new pod boots, the agent extracts the archive into the volume path, and SSH lights up once everything is in place. ![Restore section on the Create Pod page](./assets/create-pod-restore.png) ## Restore path: the rule that matters **The restore path must be a subfolder of the target pod's local volume mount** — the same rule as backups. The default mount is `/root`. ### ✅ Valid restore paths ``` /root # extract over the whole local volume /root/models # subdirectory /root/checkpoints # subdirectory /root/project/data # nested subdirectory ``` ### ❌ Invalid restore paths ``` /home/user/documents # not under the local volume mount /tmp # tmpfs — wiped on pod restart, the restore wouldn't survive /var/log # system directory /mnt # external Volume — restore is for the local volume only /mnt/datasets # same as above ``` :::tip Match the layout you backed up If the backup was taken at `/root` then the archive's top-level paths look like `models/…`, `checkpoints/…`. Extracting that archive into `/root` puts everything back where it was. Extracting it into `/root/from-backup` nests it one level deeper — useful when you want to compare or merge. ::: ## Watch progress Both flows write a **restore log** that you can read on: - The pod detail page → **Backups** tab → **Restore From Backup** section. The log includes status (`pending`, `running`, `completed`, `failed`), timestamps, and a `progress` field that ticks from `0.0` to `1.0` as the archive extracts. ## What restore does (and doesn't do) - It **overwrites** files in the restore path that share names with the archive's contents. - It does **not** delete files in the restore path that aren't in the archive — so partial backups won't wipe your other data. - It does **not** restart the container — your processes keep running. Restart the daemon yourself if you need it to re-read files. - It does **not** restore environment variables, ports, or template settings — only filesystem contents under `restore_path`. ## When restores fail The restore log surfaces an error message. Most common reasons: - **Backup expired or deleted.** Check the backup row's Status on the global Backups page. - **Restore path outside the volume mount.** Same rule as backups — see above. - **Disk full on the target pod.** Check `df -h /root` inside the pod. - **Pod has no agent SSH key.** The agent needs SSH to drive the extraction; see the **Skip agent SSH key** option on [Create Pod](./create-pod). ## First, you need a backup Restores require an archive that already exists. If you haven't set one up yet, see [Backups](./backups) — it takes about 30 seconds to configure.
For agents and automation: API + CLI ```bash # Trigger a restore into a running pod curl -X POST https://lium.io/api/restores \ -H "X-API-Key: $LIUM_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "backup_log_id": "", "target_pod_id": "", "restore_path": "/root" }' # Read restore status curl https://lium.io/api/restore-logs/ \ -H "X-API-Key: $LIUM_API_KEY" # List restores for a pod curl https://lium.io/api/restore-logs?pod_id= \ -H "X-API-Key: $LIUM_API_KEY" ``` The agent on the pod posts progress to `PUT /restore-logs/{id}/progress` so the UI can show a percentage. CLI: [`lium bk restore --id `](/developers/cli/reference/bk) (use `lium bk logs` to see job status). Get an [API key](./api-keys) first.
--- # Scheduled termination # Scheduled termination Schedule a pod to be deleted automatically at a future time. Cheap insurance against forgetting a `kill` overnight — billing stops the moment the pod is terminated. ## Set it when deploying On the **Create Pod** page, the **Auto-Termination (Optional)** field takes a number of hours. The pod will be deleted that many hours after deploy. > Pod will automatically terminate after specified hours. This sets `termination_hours` on the pod. Leave the field blank if you want the pod to keep running until you delete it manually. ## Set it on a running pod From the pod detail page: 1. Click **DELETE** (top-right action bar). 2. In the menu, choose **Schedule pod deletion**. 3. Pick a preset (1 h, 2 h, 3 h, 6 h, 12 h, 1 day) or pick a **custom date and time**. The pod detail page now shows a "Scheduled to terminate at …" line, and a **Remove Scheduled Terminate** action lets you cancel. ![Schedule deletion menu on the pod detail page](./assets/scheduled-termination.png) ## Cancel a schedule Pod detail page → **Remove Scheduled Terminate**. The pod returns to "until deleted" mode. ## Notes - Billing **stops** at the scheduled termination time — you don't pay for the partial hour after it terminates. - You can't bulk-schedule multiple pods at once from the UI; do it per pod, or use the API in a loop. - The schedule survives pod restarts — rebooting the pod from inside doesn't cancel the timer.
For agents and automation: API + CLI ```bash # Schedule removal at a specific UTC datetime curl -X POST https://lium.io/api/pods//schedule-removal \ -H "X-API-Key: $LIUM_API_KEY" \ -H "Content-Type: application/json" \ -d '{"removal_scheduled_at": "2026-05-01T12:00:00Z"}' # Cancel a schedule curl -X DELETE https://lium.io/api/pods//schedule-removal \ -H "X-API-Key: $LIUM_API_KEY" ``` The scheduled time lives on the pod object as `removal_scheduled_at` (ISO 8601 UTC). When you create a pod via API, set `termination_hours` on the create request to bake the schedule in at deploy time. CLI: bake the schedule in at create time with [`lium up --ttl`](/developers/cli/reference/up) or `--until`; manage existing schedules with [`lium schedules list / rm`](/developers/cli/reference/schedules). Get an [API key](./api-keys) first.
--- # Broken pods # Broken pods Sometimes a pod's status turns red and reads **BROKEN**. This means the **GPU provider force-closed your pod** — the container and your rental are already gone. Hover the status for the in-app note: > This pod was force-closed by the provider. You have been credited. You can delete it. ## What it means for you - **Billing has stopped** and you've been **credited** for the interrupted rental — you aren't charged for a broken pod. - The pod is a dead shell. Anything that wasn't on a [Volume](./volumes) or in a [Backup](./backups) is gone. ## What you can do A broken pod can **only be deleted**. Every other action is disabled, because there's no live container to act on: - ❌ Reboot, edit, switch template, install Jupyter, add/remove SSH keys - ❌ Back up or restore - ✅ **Delete** — removes the pod from **Your Pods** Open the pod and click **DELETE**. ## Renting the same node again To rent the same machine again, **delete the broken pod first**. If you try to rent it while the broken pod is still there, the rental is blocked with a message that names the pod to remove — for example: > This node still has a broken pod (``) that must be removed before you can rent it again. Please delete the broken pod first. Delete the broken pod, then deploy as usual from [Create a pod](./create-pod).
For agents and automation: API ```bash # Delete a broken pod curl -X DELETE https://lium.io/api/pods/ \ -H "X-API-Key: $LIUM_API_KEY" ``` A pod's `status` is `BROKEN` once the provider force-closes it. Any mutating call on it (reboot, edit, switch-template, install-jupyter, SSH keys, backup, restore) returns **400** — only delete succeeds. Renting the node again while your broken pod still exists also returns **400**, with the broken pod's id in the message, until you delete it. Get an [API key](./api-keys) first.
--- # Discord support # Discord support Rented pods can open a private Discord support channel for the renter, the GPU provider, and the Lium support role when that role is configured. Use it when a running pod needs coordination with the provider: SSH access trouble, unexpected restarts, networking questions, performance issues, or container debugging. ## What must be connected first Both sides need a linked Discord account: - **Renter** — connect Discord from your Lium account settings at [lium.io](https://lium.io). - **Provider** — connect Discord from the Provider Portal profile/settings page at [provider.lium.io](https://provider.lium.io). See [Provider Discord setup](../providers/portal/discord.md). If either side has not connected Discord, the pod detail page will explain what is missing before it lets you request a channel. ## Connect Discord as a renter 1. Sign in to [lium.io](https://lium.io). 2. Open **Settings** or **Profile**. 3. In the **Discord** section, click **Connect Discord**. 4. Approve the Discord OAuth prompt. 5. Return to Lium and confirm the Discord section shows **Connected**. The OAuth connection stores your Discord user ID on your Lium account. Validator accounts cannot use the renter Discord connection flow. You can disconnect the account later from the same Discord section. Disconnecting prevents new support channels from including you until you connect again. ## Request a pod support channel 1. Open **Your Pods**. 2. Select the pod that needs support. 3. Find the **Discord Support** section near the top of the pod detail page. 4. Click **Request Discord Support**. When the request succeeds, Lium opens the Discord channel in a new browser tab. If the pod already has a support channel, the button changes to **Open Discord channel** and returns the existing channel instead of creating a duplicate. ## What Lium sends to Discord The support request seeds the channel with pod context so the provider and support team can identify the machine quickly. The message includes details such as: - pod ID, pod name, and pod status - provider and validator hotkeys - node ID, machine name, IP address, and port - container name - Docker image, tag, digest, and template - forwarded ports and SSH command - GPU, CPU, RAM, price, and location Do not paste secrets, private keys, seed phrases, API keys, or proprietary data into the support channel unless you are comfortable sharing them with the channel participants. ## Channel visibility The channel is private by default. Lium creates Discord permission overwrites for: - the renter Discord user - the provider Discord user - the configured Lium support role, when enabled The channel is tied to the pod record. When Lium removes the pod during normal pod cleanup, it also attempts to delete the stored Discord channel. ## Troubleshooting | Message | What it means | What to do | | --- | --- | --- | | **Connect Discord first** | Your renter account does not have a Discord ID linked. | Open **Settings** and connect Discord, then retry from the pod detail page. | | **Provider has not connected Discord** | The provider for this pod has not linked Discord in the Provider Portal. | Wait for the provider to connect Discord, or contact Lium support through your normal support path. | | **Discord support channel could not be created** | Lium could not create the Discord channel. This can happen if Discord integration is unavailable or misconfigured. | Retry once. If it still fails, contact Lium support with the pod ID. | | Browser does not open Discord | The backend returned success but your browser blocked the new tab. | Allow popups for `lium.io`, then click **Open Discord channel** again. | --- # Overview # Overview Pods in Lium represent individual GPU rental units that users can lease for their computational needs. Each pod is a containerized environment that provides secure, isolated access to GPU resources through the Bittensor subnet infrastructure. Security depends on the underlying machine you choose. CVM-enabled machines provide stronger isolation for sensitive workloads, while non-CVM machines may allow the GPU provider to access the rented pod. Review the [Create Pod guide](./create-pod.md) before deployment, and read the [Confidential Virtual Machine (CVM) guide](../providers/nodes/cvm.md) if you need the full security model. --- # Pod Security # Pod Security :::warning Keep your secrets off rented pods Do not upload private keys, seed phrases, API keys, or other secret files to a rented pod. The GPU provider may have access to the pod environment. ::: ## What GPU Providers Can Do on Non-CVM Machines When you rent a pod on a machine that is **not** CVM-enabled, the GPU provider retains full host-level access to the underlying server. This means they can: - **Inspect container contents** — read any file inside your pod, including those you copied in or generated at runtime - **Read environment variables** — any secrets passed via `ENV`, `--env`, or `.env` files are visible - **Read process memory** — secrets held in memory (loaded keys, decrypted tokens) can be extracted - **Intercept network traffic** — plaintext traffic within the pod can be observed - **Replace or modify the container** — the node image is not integrity-protected on non-CVM hosts This is not a bug — it is a fundamental property of how Linux containers work. Containers share the host kernel; a privileged host process can always inspect container memory and filesystems. ## What You Must Never Place in a Non-CVM Pod | Secret type | Examples | |---|---| | Blockchain private keys | Bittensor coldkeys/hotkeys, Ethereum/Solana wallets | | Seed phrases / mnemonics | BIP39 recovery phrases for any wallet | | SSH private keys | `~/.ssh/id_rsa`, `~/.ssh/id_ed25519` | | API keys | OpenAI, Anthropic, HuggingFace, AWS, GCP, Azure tokens | | Database credentials | PostgreSQL, MySQL, Redis passwords and connection URIs | | OAuth/app secrets | Client secrets, JWT signing keys, webhook secrets | | Model weights with IP value | Proprietary fine-tuned models, trade-secret weights | ## What to Do Instead ### Option 1 — Use a CVM-enabled machine (recommended for secrets) CVM machines use Intel TDX hardware isolation. The GPU provider runs the VM but **cannot read its memory or filesystem contents**. See the [CVM guide](../providers/nodes/cvm.md) for hardware requirements and how to identify CVM-enabled nodes. ### Option 2 — Keep the local disk encrypted (on by default) [Encrypted local volumes](./encrypted-volumes) store `/root` as ciphertext on the provider's disk. Leftover volumes and disk images are unreadable after your pod is gone. This does not change the table above. Those are attacks on a *running* pod, where the provider's host access wins regardless. Encryption protects data at rest. A CVM protects it in use. ### Option 3 — Fetch secrets at runtime from an external secrets manager If you must use a non-CVM machine, avoid placing secrets in the pod image or container filesystem. Instead: - Pull secrets from an external vault (HashiCorp Vault, AWS Secrets Manager, GCP Secret Manager) at runtime - Rotate secrets immediately after the pod session ends - Use short-lived tokens scoped to the minimum required permissions - Treat the pod as fully compromised — rotate any secrets used in it after termination ### Option 4 — Design workloads to not require secrets - Use pre-signed URLs or token-based object storage access that expires - Run inference against an external API rather than loading weights locally - Separate the credential-holding component from the GPU-intensive component ## How to Identify CVM-Enabled Machines When browsing available nodes in the Lium UI or CLI, CVM-capable machines are labeled accordingly. The [CVM guide](../providers/nodes/cvm.md) describes the required hardware (Intel TDX + NVIDIA Confidential Computing) and how attestation works. If a machine is not labeled as CVM-enabled, assume the GPU provider can access your pod. --- # API Keys # API Keys Most renters never need this page. Use it when: - You want to drive Lium from the **CLI** (`lium ...`). - You want an **AI agent** (Claude, Cursor, your own) to deploy and manage pods on your behalf via the API or [MCP server](/developers/mcp). - You're integrating Lium into a CI pipeline, a Slack bot, or your own tool. If you're just here to rent a GPU and SSH in, the dashboard at [lium.io](https://lium.io) is the friendlier path — see [Quickstart](./quickstart). ## What an API key does An API key authenticates HTTP requests to `https://lium.io/api/...` as **you**. Anyone with the key can: - Deploy pods, attach volumes, start backups — and **be billed to your account**. - Read your pod list, backup history, and account balance. - Delete things. The API has no separate "read-only" scope yet. Treat the key like a password. Never paste it into a public repo, an issue tracker, a screenshot, or a non-CVM pod. ## Generate a key 1. Sidebar → **Access** (key icon). 2. Click the **API Keys** tab. 3. **ADD NEW +**, give the key a **Name** (something like `cli-laptop`, `ci-deploy-bot`, `claude-agent`), **ADD**. 4. The full key shows **once** in a copy-to-clipboard chip. Copy it now — Lium only stores a hashed version, so you can't view it again. ![Access page → API Keys tab](./assets/api-keys-list.png) The list view shows each key's **Name**, masked **Key** prefix, **Active** flag, **Last used**, **Expires at**, and **Date created**. ## Use the key Pass the key in the `X-API-Key` header on every request: ```bash export LIUM_API_KEY=sk_yjujqGP...full-key curl https://lium.io/api/pods \ -H "X-API-Key: $LIUM_API_KEY" ``` For the CLI: ```bash pip install lium.io # or: curl https://lium.io/install.sh | sh lium init # guided login (or set LIUM_API_KEY in your env) lium ps # list your pods ``` For the [SDK](/developers/sdk) and [MCP server](/developers/mcp), set `LIUM_API_KEY` in the environment. ## Quick "API Key" button (top-left) The **API Key** button at the very top of the sidebar (under your balance) **copies your most recent active key to the clipboard** and shows a toast. Handy for quick re-pastes; not a substitute for managing keys on the Access page. ## Rotate, deactivate, delete On the **Access → API Keys** row: - ✏️ **Edit** — rename or set an expiry date. - 🗑️ **Delete** — revokes the key immediately. Anything using it (CLI session, running agent) will start getting 401s. There's no separate "deactivate but keep" toggle yet — delete and recreate when you need to rotate. ## What this unlocks for AI agents The whole Renters surface (pods, templates, volumes, backups, restores, scheduled termination) is exposed via the same REST API and via the [MCP server](/developers/mcp). With one API key, an agent can: - Watch the marketplace and deploy a cheap pod when an A100 drops below your price ceiling. - Spin up a fresh pod for each job, run training, take a final backup, terminate. - Pull a backup down for inspection in your laptop, then fan out restores into many pods. - Manage templates: keep your team's images, tags, and entrypoints in sync from CI. Schemas live in the [OpenAPI spec](/developers/openapi). The Renters pages above each have a "For agents and automation" block at the bottom showing the curl/CLI equivalents of every UI flow. ## Key hygiene - One key per integration. If your laptop's key leaks, you only revoke that one. - Set an **Expires at** for short-lived agents (CI runners, hackathon scripts). - Don't bake keys into Docker images or templates. Set them as env vars at runtime, or use the agent itself to deploy with a short-lived key. - Never put a key in a non-CVM pod's filesystem — see [Pod security](./security). --- # Developers # Developers Build on top of Lium programmatically — whether you're a human writing a service or an AI agent driving the platform on a user's behalf. Everything Lium exposes (REST API, CLI, Python SDK, MCP, llms.txt) is documented here. The external URLs in the lists below are stable and safe to bookmark. ## For AI agents Start at **[AI Agents](/developers/agents)** — a single page that wires up all four agent surfaces: - [`https://docs.lium.io/mcp`](/developers/mcp) — docs MCP server (search + read pages) - [`https://docs.lium.io/llms-full.txt`](/developers/llms-txt) — full docs as one markdown bundle - [`https://lium.io/api/openapi.json`](/developers/openapi) — live OpenAPI 3.1 spec for the platform API - [`curl -fsSL https://lium.io/install.sh | bash`](/developers/cli/overview) — shell-driven pod management ## For developers - [Get started in 5 min](/developers/quickstart) — your first authenticated API call - [OpenAPI spec](/developers/openapi) — every REST endpoint, payload, and response - [SDK](/developers/sdk) — typed Python client (`@lium.machine`, `Lium()`) - [CLI overview](/developers/cli/overview) — the `lium` command-line tool - [Installation](/developers/cli/installation) · [Quickstart](/developers/cli/quickstart) · [Reference](/developers/cli/reference) - [MCP endpoint](/developers/mcp) — Model Context Protocol docs server - [llms.txt](/developers/llms-txt) — `/llms.txt` and `/llms-full.txt` for LLM context The live Swagger UI is hosted at [lium.io/documents](https://lium.io/documents); the raw spec lives at [lium.io/api/openapi.json](https://lium.io/api/openapi.json). --- # AI Agents # AI Agents This page is the entry point for AI agents (Claude, ChatGPT, Cursor, in-house copilots, …) integrating with Lium. Everything else in the Developers section also applies to agents — this page just stitches the pieces together. Prefer to work from a single file? [`lium.io/llms-full.txt`](https://lium.io/llms-full.txt) is the whole agent reference in one fetch — install, signup, the [agent skill](https://github.com/Datura-ai/lium-skill), and every CLI and SDK command — enough to go from no account to a rented GPU without reading anything else. ## Get started in 5 min The CLI is the preferred surface for an agent: it covers signup, funding, renting and SSH, so there is no web form to fill and no HTTP to hand-roll. Drop to the REST API only where the CLI cannot go — no shell in your sandbox, or an endpoint it does not wrap. 1. **Install the CLI and get an account:** ```bash curl -fsSL https://lium.io/install.sh | bash lium signup --email you@example.com # no account yet lium init --no-browser # account exists — prints an approval URL ``` 2. **Rent a GPU and connect:** ```bash lium ls --format json # browse machines lium up --yes # rent lium ssh # connect ``` 3. **Ground yourself in the docs** — add the MCP server for targeted search: ```json { "mcpServers": { "lium-docs": { "url": "https://docs.lium.io/mcp" } } } ``` Or pull the whole corpus in one fetch: `curl -sSL https://docs.lium.io/llms-full.txt`. 4. **Call the REST API** for anything the CLI does not cover: ```bash curl -sSL https://lium.io/api/openapi.json ``` ## Pick the right surface | Goal | Use this | Why | |------|----------|-----| | Answer questions about Lium docs | [MCP endpoint](./mcp) — `https://docs.lium.io/mcp` | Targeted search + page reads, no token bloat | | Ground a long conversation in the full docs | [llms.txt / llms-full.txt](./llms-txt) | One fetch, every page concatenated | | Read a single page as raw markdown | Append `.md` to any URL — e.g. `/providers/quickstart.md` | No HTML to parse | | Call an endpoint the CLI does not wrap | [OpenAPI spec](./openapi) — `https://lium.io/api/openapi.json` | Source of truth for every REST endpoint | | Drive pods from a shell agent | [CLI](./cli/overview) | Stable subcommands, JSON output, no auth headers to manage | | Take a user from no account to an API key | [`lium signup`](./cli/reference/signup.md) | Creates the account and mints its key from the terminal — no web form | | Teach your agent Lium's commands once | [Agent skill](https://github.com/Datura-ai/lium-skill) — `npx skills add Datura-ai/lium-skill --skill lium` | Ships with the agent, so no doc fetch mid-task | | Automate the provider portal (Subnet 51) | [`lium provider`](./cli/reference/provider.md) | Same surface as [lium.io/portal](/providers/portal/overview): node lifecycle, central-miner-server config, sync, billing, machine requests — all `--json`-able | | Write code against typed Python clients | [SDK](./sdk) | `lium.machine` decorator + `Lium()` client | ## Stable agent-facing URLs These URLs are stable and safe to hard-code in agent configurations: ```text https://docs.lium.io/mcp # docs MCP JSON-RPC endpoint https://docs.lium.io/llms.txt # docs index + summaries https://docs.lium.io/llms-full.txt # all docs concatenated https://docs.lium.io/.md # raw markdown for any page https://lium.io/api/openapi.json # live platform OpenAPI 3.1 spec https://lium.io/documents # Swagger UI for the same spec https://lium.io/llms.txt # platform agent index: CLI + skill install https://lium.io/llms-full.txt # full CLI/SDK reference, self-contained ``` ## Start from no account An agent creates the account itself — no web form. [`lium signup`](./cli/reference/signup.md) creates it and stores the API key it mints. New accounts usually get a **$5 credit**, enough to rent straight away, but it is granted once per signup IP — read `signup_credit_granted` from the `--json` output rather than assuming it landed. ```bash lium signup --email you@example.com --json ``` Ask for the user's real email: the confirmation link goes there, and clicking it is the one step an agent cannot do for them — renting fails with `403 User is not verified` until they do. ## Authentication The platform API authenticates with an API key, sent in the **`X-API-Key`** header on every request; the CLI stores the key and sends it for you. The docs surfaces (MCP, llms.txt, `.md` URLs) require no auth. [`lium signup`](./cli/reference/signup.md) mints and stores the key for a new account. For an account that already exists, [`lium init --no-browser`](./cli/reference/init.md) prints an approval URL **and** a session ID; once the user has approved it, `lium init --session ` saves the key. Plain `lium init` opens a browser and is not usable by an agent. A key copied from **Settings → API Keys** at [lium.io](https://lium.io) works too. ```bash export LIUM_API_KEY=your_api_key_here curl https://lium.io/api/pods -H "X-API-Key: $LIUM_API_KEY" ``` ## Self-funding (top up your own balance) An agent with its own crypto wallet can keep itself funded without a human. Use [`lium topup`](./cli/reference/topup.md): create a stablecoin invoice, read the deposit address, send the funds from your wallet, then poll the balance. Every subcommand supports `--json`. ```bash # 1. Create a $20 USDT (TRON) invoice; capture the address + exact amount INVOICE=$(lium topup create -a 20 -c USDT -n tron --json) ADDRESS=$(echo "$INVOICE" | jq -r .deposit_address) AMOUNT=$(echo "$INVOICE" | jq -r .crypto_amount) # 2. Send exactly $AMOUNT USDT on TRON to $ADDRESS from your own wallet. # Lium never moves your funds — only you can sign this transfer. # 3. Poll until the credit lands lium balance --json ``` Lium issues the invoice and credits the balance once the payment is confirmed on-chain; the transfer itself stays in your wallet. See [`lium topup`](./cli/reference/topup.md) for currency discovery and the full invoice response. ## Recommended agent loop 1. Install the CLI and authenticate — [`lium signup`](./cli/reference/signup.md) for a new account, [`lium init --no-browser`](./cli/reference/init.md) for an existing one. 2. Drive the platform with `lium` commands, reading results as JSON. 3. Drop to the REST API only where the CLI has no equivalent: resolve the endpoint from the OpenAPI spec and send `X-API-Key`. 4. Pull context from MCP `search` — or feed `llms-full.txt` once — when you need to explain or troubleshoot rather than act. ## Related - [MCP endpoint](./mcp) — JSON-RPC contract and tool definitions - [llms.txt](./llms-txt) — llmstxt.org-format docs bundle - [OpenAPI spec](./openapi) — code-gen and tool-use examples - [CLI overview](./cli/overview) — full command surface - [SDK](./sdk) — typed Python client --- # Developer Quickstart # Developer Quickstart ## Get started in 5 min This guide takes you from zero to your first authenticated API call against the Lium platform API. ### Step 1 — Get an API key 1. Log in to [lium.io](https://lium.io). 2. Navigate to **Settings → API Keys**. 3. Click **Create API Key**, copy the key, and store it securely. ### Step 2 — Make your first request List available templates (public endpoint, no auth required): ```bash curl https://lium.io/api/templates ``` List your pods (requires API key): ```bash export LIUM_API_KEY=your_api_key_here curl https://lium.io/api/pods \ -H "X-API-Key: $LIUM_API_KEY" ``` ### Step 3 — Create a pod ```bash curl -X POST https://lium.io/api/pods \ -H "X-API-Key: $LIUM_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "machine_id": "", "template_id": "", "gpu_count": 1, "ssh_key_ids": [""] }' ``` The response includes a `pod_id`. Poll `GET /pods/{pod_id}` until status is `running`. **End state**: You have a running pod created via the API. The live OpenAPI reference is available at [lium.io/documents](https://lium.io/documents). ### Step 4 — Explore the CLI Install the CLI for scripted workflows: ```bash curl -fsSL https://lium.io/install.sh | bash # or `pip install lium.io` lium --help ``` See the [CLI installation guide](./cli/installation) for platform notes, version pinning, and the Python package alternative. ## Next steps - [OpenAPI spec](./openapi) — every endpoint, payload, and response - [SDK](./sdk) — typed Python client built on the same API - [CLI installation and quickstart](./cli/quickstart) - [AI Agents](./agents) — MCP, llms.txt, OpenAPI, and CLI in one map --- # OpenAPI Spec # OpenAPI Spec The Lium platform API publishes a live [OpenAPI 3.1](https://spec.openapis.org/oas/v3.1.0) schema. Point any code-generator, agent runtime, or HTTP client at it to discover every endpoint, payload, and response. | Resource | URL | |----------|-----| | Raw JSON spec | [`https://lium.io/api/openapi.json`](https://lium.io/api/openapi.json) | | Swagger UI | [`https://lium.io/documents`](https://lium.io/documents) | The schema is generated directly from the FastAPI app at request time, so it always matches the live deployment. ## Fetch the spec ```bash curl -sSL https://lium.io/api/openapi.json -o lium-openapi.json ``` ```python import httpx spec = httpx.get("https://lium.io/api/openapi.json").json() print(spec["info"]["title"], spec["info"]["version"]) ``` ## Use it from an AI agent ### Ground Claude with the spec Pass the spec to Claude as cached context, then let the model reason about which endpoint to call. The `cache_control` block keeps you from re-billing input tokens on every request — the spec is large but stable. ```python import json, httpx, anthropic spec_text = httpx.get("https://lium.io/api/openapi.json").text client = anthropic.Anthropic() resp = client.messages.create( model="claude-sonnet-4-5", max_tokens=1024, system=[ {"type": "text", "text": "You manage Lium GPU pods. Use the OpenAPI spec to choose the right endpoint."}, {"type": "text", "text": f"\n{spec_text}\n", "cache_control": {"type": "ephemeral"}}, ], messages=[{"role": "user", "content": "List my running pods."}], ) print(resp.content[0].text) ``` For real tool calling (model invokes endpoints directly), walk `spec["paths"]` and emit one Anthropic [tool definition](https://docs.anthropic.com/en/docs/build-with-claude/tool-use) per operation — `name = operationId`, `input_schema = requestBody.content["application/json"].schema`. ### MCP-aware agents If your agent already speaks [MCP](./mcp), point it at the docs MCP endpoint for prose questions and at this OpenAPI URL for direct API calls. Both are stable and require no extra configuration. ### Code generation Generate a typed client in any language with the standard OpenAPI tooling: ```bash # TypeScript / fetch npx openapi-typescript https://lium.io/api/openapi.json -o lium.d.ts # Python (httpx) pip install openapi-python-client openapi-python-client generate --url https://lium.io/api/openapi.json ``` ## Authentication Authenticated endpoints accept your API key in the **`X-API-Key`** request header. Get a key from **Settings → API Keys** at [lium.io](https://lium.io). See the [Developer Quickstart](./quickstart) for a complete first-call walkthrough. ```bash export LIUM_API_KEY=your_api_key_here curl https://lium.io/api/pods -H "X-API-Key: $LIUM_API_KEY" ``` The OpenAPI spec also advertises a `JwtAccessBearer` security scheme — that one is for JWT session tokens used by the lium.io frontend. For server-to-server calls with an API key, use `X-API-Key`. ## Related - [Developer Quickstart](./quickstart) — first authenticated request in 5 minutes - [SDK](./sdk) — typed Python client built on top of the same API - [MCP endpoint](./mcp) — search and read the docs over JSON-RPC - [llms.txt](./llms-txt) — bulk markdown bundle for LLM context --- # SDK # Lium SDK The `lium.io` package ships both the [CLI](./cli/overview) and a **Python SDK** for managing GPU pods programmatically. Install it once and use whichever interface fits the job. :::info SDK Reference The generated Python SDK reference covers the public client, data models, exceptions, configuration, and decorators. **[Open SDK Reference](./sdk/reference)** ::: ## Installation ```bash pip install lium.io ``` ## Authentication For local development, authenticate once with the CLI: ```bash lium init ``` This saves your API key to `~/.lium/config.ini`, which is shared by both the CLI and SDK. After that, `Lium()` can authenticate automatically: ```python from lium.sdk import Lium lium = Lium() ``` For CI, scripts, or temporary overrides, set `LIUM_API_KEY` instead: ```bash export LIUM_API_KEY="sk_..." ``` `LIUM_API_KEY` takes precedence over the saved config file. ## Two Entry Points The SDK exposes two ways to run work on Lium GPUs: - **`@lium.machine` decorator** — annotate a Python function and offload it to a GPU pod. Best for quickly running isolated workloads. - **`Lium()` client** — a direct client for long-lived orchestration code that manages pod lifecycles. ### `@lium.machine` decorator Annotate a function with the machine type and dependencies, then call it like a normal Python function. The SDK handles provisioning, code upload, execution, and teardown. ```python import lium @lium.machine(machine="A100", requirements=["torch", "transformers", "accelerate"]) def infer(prompt: str) -> str: from transformers import AutoTokenizer, AutoModelForCausalLM model_id = "HuggingFaceTB/SmolLM2-135M-Instruct" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, device_map="cuda") tokens = tokenizer.apply_chat_template( [{"role": "user", "content": prompt}], return_tensors="pt", add_generation_prompt=True, ).to("cuda") out = model.generate( tokens, max_new_tokens=64, do_sample=False, pad_token_id=tokenizer.eos_token_id, ) return tokenizer.decode(out[0][tokens.shape[-1]:], skip_special_tokens=True).strip() print(infer("Who discovered penicillin?")) ``` For a complete remote inference workflow, see [Remote Inference with `@lium.machine`](./sdk/examples/machine-inference). ### `Lium()` client The client mirrors the CLI's pod lifecycle — list nodes, bring a pod up, wait until it's ready, execute commands, and tear it down. ```python from lium.sdk import Lium lium = Lium() ready = None try: executor = lium.ls(gpu_type="A100", gpu_count=1)[0] pod = lium.up(executor_id=executor.id, name="demo") ready = lium.wait_ready(pod, timeout=600) if ready is None: raise RuntimeError("Pod did not become ready before the timeout") print(lium.exec(ready, command="nvidia-smi")["stdout"]) finally: if ready is not None: lium.down(ready) ``` For the complete CLI-equivalent workflow, see [Pod Lifecycle with `Lium()`](./sdk/examples/pod-lifecycle). ## Next Steps Ready to go deeper? Open the **[generated SDK reference](./sdk/reference)** for signatures, arguments, return types, data models, and exceptions, or follow the **[SDK examples](./sdk/examples)** for complete workflows. ## Related - [SDK Examples](./sdk/examples) — complete SDK workflows - [CLI Installation](./cli/installation) — install the `lium.io` package - [CLI Reference](./cli/reference) — per-command reference, grouped by category - [CLI Quickstart](./cli/quickstart) — get started with the CLI --- # SDK Examples # SDK Examples These examples show runnable SDK workflows. ## Available Examples | Example | What it shows | | --- | --- | | [Pod lifecycle with `Lium()`](/developers/sdk/examples/pod-lifecycle) | Create a pod, wait for readiness, run commands, transfer files, and clean up. | | [Train a model on a pod](/developers/sdk/examples/training-workflow) | Sync a training project, stream the training run, download a checkpoint, and stop the pod. | | [Custom Docker template](/developers/sdk/examples/custom-template) | Create a private template with a Docker image, startup command, ports, environment, and volume paths. | | [Remote inference with `@lium.machine`](/developers/sdk/examples/machine-inference) | Offload a Python function to a GPU pod and run a small instruct model. | | [Run a vLLM server](/developers/sdk/examples/vllm) | Create a vLLM template, start a GPU pod, and expose the OpenAI-compatible API. | Use the [SDK reference](/developers/sdk/reference) when you need exact signatures, return types, and model fields. --- # Pod Lifecycle with Lium() # Pod Lifecycle with `Lium()` This example shows the core pod workflow in Python: choose a GPU, create a pod, wait for SSH readiness, run a command, move a small file, and stop the pod. Use this pattern when you want a script, notebook, CI job, or agent to manage pod lifecycle directly instead of shelling out to the CLI. ```python #!/usr/bin/env python3 """Create a Lium pod, run a command, transfer a file, and clean up.""" from pathlib import Path from lium.sdk import Lium GPU_TYPE = "A100" GPU_COUNT = 1 POD_NAME = "sdk-lifecycle-demo" LOCAL_NOTE = Path("lium-sdk-note.txt") DOWNLOADED_NOTE = Path("lium-sdk-note.remote.txt") REMOTE_NOTE = "/root/lium-sdk-note.txt" def require_success(result: dict, label: str) -> None: """Raise a useful local error when a remote command fails.""" if result["success"]: return stderr = result.get("stderr", "").strip() stdout = result.get("stdout", "").strip() details = stderr or stdout or f"exit code {result.get('exit_code')}" raise RuntimeError(f"{label} failed: {details}") lium = Lium() created_pod_id = None ready_pod = None try: # 1. Find a matching GPU executor. executors = lium.ls(gpu_type=GPU_TYPE, gpu_count=GPU_COUNT) if not executors: raise RuntimeError(f"No available {GPU_COUNT}x {GPU_TYPE} executors") # Pick the lowest listed hourly price for this small demo. executor = min(executors, key=lambda item: item.price_per_hour) print( f"Using {executor.machine_name} at ${executor.price_per_hour:.2f}/hr " f"({executor.huid})" ) # 2. Create the pod. pod = lium.up( executor_id=executor.id, name=POD_NAME, ports=1, ) created_pod_id = pod["id"] print(f"Created pod {pod.get('name', POD_NAME)} ({created_pod_id})") # 3. Wait until the pod is running and has SSH metadata. ready_pod = lium.wait_ready(pod, timeout=600) if ready_pod is None: raise RuntimeError("Pod did not become ready before the timeout") print(f"Ready: {ready_pod.name} ({ready_pod.huid})") print(lium.ssh(ready_pod)) # 4. Inspect active pods, like `lium ps`. print("Active pods:") for active in lium.ps(): print(f"- {active.name}: {active.status} ({active.huid})") # 5. Run a command over SSH, like `lium exec`. gpu_info = lium.exec(ready_pod, command="nvidia-smi") require_success(gpu_info, "nvidia-smi") print(gpu_info["stdout"]) # 6. Upload and download a file, like `lium scp`. LOCAL_NOTE.write_text("hello from the local machine\n", encoding="utf-8") lium.upload(ready_pod, local=str(LOCAL_NOTE), remote=REMOTE_NOTE) cat_note = lium.exec(ready_pod, command=f"cat {REMOTE_NOTE}") require_success(cat_note, "cat uploaded note") print(cat_note["stdout"].strip()) lium.download(ready_pod, remote=REMOTE_NOTE, local=str(DOWNLOADED_NOTE)) print(f"Downloaded: {DOWNLOADED_NOTE.read_text(encoding='utf-8').strip()}") finally: # Always stop temporary pods so they do not keep accruing charges. pod_to_stop = ready_pod if pod_to_stop is None and created_pod_id: pod_to_stop = next((p for p in lium.ps() if p.id == created_pod_id), None) if pod_to_stop is not None: lium.down(pod_to_stop) print(f"Stopped pod {pod_to_stop.name} ({pod_to_stop.huid})") ``` ## Notes - Call `wait_ready()` before `exec`, `upload`, `download`, `rsync`, or `ssh`; those operations need SSH connection metadata. - Use `try` / `finally` around temporary pods so failures do not leave a pod running. - Use `stream_exec()` instead of `exec()` for long-running jobs where you want incremental output. - Use `rsync()` for directory syncs only when your container image has `rsync` installed. The [training workflow](./training-workflow) shows how to check for it and install it when missing. --- # Train a Model on a Pod # Train a Model on a Pod This example runs a local training project on a Lium GPU pod. It selects a GPU, starts a pod with a PyTorch template, syncs code and data, installs dependencies, streams the training output, downloads the best checkpoint, and stops the pod. The script assumes this local project layout: ```text . |-- requirements.txt |-- src/ | `-- train.py `-- data/ ``` `src/train.py` should accept `--data-dir` and `--output-dir`, then write the checkpoint to `checkpoints/best_model.pt`. Run the orchestration script from the project root: ```python #!/usr/bin/env python3 """Run a local training project on a Lium GPU pod.""" from pathlib import Path from lium.sdk import Lium GPU_TYPE = "A100" GPU_COUNT = 1 MAX_PRICE_PER_HOUR = 1.50 MIN_CUDA_VERSION = 12.4 POD_NAME = "sdk-training-demo" PROJECT_ROOT = Path.cwd() LOCAL_REQUIREMENTS = PROJECT_ROOT / "requirements.txt" LOCAL_SRC = PROJECT_ROOT / "src" LOCAL_DATA = PROJECT_ROOT / "data" LOCAL_CHECKPOINT = PROJECT_ROOT / "checkpoints" / "best_model.pt" REMOTE_WORKSPACE = "/workspace" REMOTE_REQUIREMENTS = f"{REMOTE_WORKSPACE}/requirements.txt" REMOTE_SRC = f"{REMOTE_WORKSPACE}/src" REMOTE_DATA = f"{REMOTE_WORKSPACE}/data" REMOTE_CHECKPOINTS = f"{REMOTE_WORKSPACE}/checkpoints" REMOTE_BEST_MODEL = f"{REMOTE_CHECKPOINTS}/best_model.pt" def require_local_inputs() -> None: missing = [ str(path) for path in (LOCAL_REQUIREMENTS, LOCAL_SRC, LOCAL_DATA) if not path.exists() ] if missing: raise RuntimeError(f"Missing local training inputs: {', '.join(missing)}") def require_success(result: dict, label: str) -> None: if result["success"]: return stderr = result.get("stderr", "").strip() stdout = result.get("stdout", "").strip() details = stderr or stdout or f"exit code {result.get('exit_code')}" raise RuntimeError(f"{label} failed: {details}") def ensure_rsync(lium: Lium, pod) -> None: check = lium.exec(pod, command="which rsync") if check["success"]: return install = lium.exec( pod, command="apt-get update -qq && apt-get install -y rsync -qq", ) require_success(install, "install rsync") def select_executor(lium: Lium): executors = lium.ls( gpu_type=GPU_TYPE, gpu_count=GPU_COUNT, min_cuda_version=MIN_CUDA_VERSION, ) executors = [ executor for executor in executors if executor.price_per_hour <= MAX_PRICE_PER_HOUR ] if not executors: raise RuntimeError( f"No {GPU_COUNT}x {GPU_TYPE} executor under " f"${MAX_PRICE_PER_HOUR:.2f}/hr with CUDA >= {MIN_CUDA_VERSION}" ) return min(executors, key=lambda executor: executor.price_per_hour) def select_pytorch_template(lium: Lium): templates = lium.templates(filter="pytorch") verified = [ template for template in templates if template.status.upper() == "VERIFY_SUCCESS" ] if not verified: raise RuntimeError("No verified PyTorch template is available") return verified[0] require_local_inputs() lium = Lium() created_pod_id = None ready_pod = None try: executor = select_executor(lium) template = select_pytorch_template(lium) print( f"Using {executor.machine_name} at ${executor.price_per_hour:.2f}/hr " f"with template {template.name}" ) pod = lium.up( executor_id=executor.id, name=POD_NAME, template_id=template.id, ) created_pod_id = pod["id"] ready_pod = lium.wait_ready(pod, timeout=600) if ready_pod is None: raise RuntimeError("Pod did not become ready before the timeout") mkdir = lium.exec( ready_pod, command=f"mkdir -p {REMOTE_SRC} {REMOTE_DATA} {REMOTE_CHECKPOINTS}", ) require_success(mkdir, "create remote workspace") lium.upload( ready_pod, local=str(LOCAL_REQUIREMENTS), remote=REMOTE_REQUIREMENTS, ) ensure_rsync(lium, ready_pod) lium.rsync(ready_pod, local=f"{LOCAL_SRC}/", remote=f"{REMOTE_SRC}/") lium.rsync(ready_pod, local=f"{LOCAL_DATA}/", remote=f"{REMOTE_DATA}/") install = lium.exec( ready_pod, command=f"cd {REMOTE_WORKSPACE} && python -m pip install -r requirements.txt", ) require_success(install, "install dependencies") train_command = ( f"cd {REMOTE_WORKSPACE} && " "PYTHONUNBUFFERED=1 python src/train.py " "--data-dir data " "--output-dir checkpoints" ) for chunk in lium.stream_exec(ready_pod, command=train_command): print(chunk["data"], end="") checkpoint = lium.exec(ready_pod, command=f"test -f {REMOTE_BEST_MODEL}") require_success(checkpoint, "check training artifact") LOCAL_CHECKPOINT.parent.mkdir(parents=True, exist_ok=True) lium.download( ready_pod, remote=REMOTE_BEST_MODEL, local=str(LOCAL_CHECKPOINT), ) print(f"Downloaded checkpoint to {LOCAL_CHECKPOINT}") finally: pod_to_stop = ready_pod if pod_to_stop is None and created_pod_id: pod_to_stop = next((p for p in lium.ps() if p.id == created_pod_id), None) if pod_to_stop is not None: lium.down(pod_to_stop) print(f"Stopped pod {pod_to_stop.name} ({pod_to_stop.huid})") ``` ## Adapting the Workflow - Change `GPU_TYPE`, `GPU_COUNT`, `MAX_PRICE_PER_HOUR`, and `MIN_CUDA_VERSION` to match the hardware your job needs. - Update `train_command` if your training script uses different flags or writes artifacts to a different path. - Use `lium.exec()` instead of `stream_exec()` when you want a final `stdout`, `stderr`, and `exit_code` result dictionary instead of live output. - `ensure_rsync()` mirrors the CLI behavior: it checks whether the pod has `rsync`, then installs it with `apt-get` when missing. - For multiple output files, write a small archive on the pod with `tar` and download that single archive with `download()`. ## Minimal `train.py` If you do not already have a training script, use this minimal PyTorch example to verify the workflow. It trains a tiny linear model, uses CUDA when available, streams progress, and writes `checkpoints/best_model.pt`. With the PyTorch template above, `requirements.txt` can be empty. ```python title="src/train.py" #!/usr/bin/env python3 import argparse from pathlib import Path import torch parser = argparse.ArgumentParser() parser.add_argument("--data-dir", required=True) parser.add_argument("--output-dir", required=True) args = parser.parse_args() device = torch.device("cuda" if torch.cuda.is_available() else "cpu") data_dir = Path(args.data_dir) output_dir = Path(args.output_dir) records = sorted(path.name for path in data_dir.iterdir()) print(f"input files: {records}", flush=True) print(f"torch: {torch.__version__}", flush=True) print(f"device: {device}", flush=True) torch.manual_seed(7) x = torch.linspace(-1, 1, 256, device=device).unsqueeze(1) y = 2.0 * x + 0.3 model = torch.nn.Linear(1, 1).to(device) optimizer = torch.optim.SGD(model.parameters(), lr=0.2) loss_fn = torch.nn.MSELoss() for step in range(1, 31): optimizer.zero_grad() loss = loss_fn(model(x), y) loss.backward() optimizer.step() if step % 10 == 0: print(f"step {step:02d} loss={loss.item():.6f}", flush=True) output_dir.mkdir(parents=True, exist_ok=True) torch.save( { "state_dict": model.state_dict(), "final_loss": loss.item(), "device": str(device), "input_files": records, }, output_dir / "best_model.pt", ) print("checkpoint written", flush=True) ``` --- # Custom Docker Template # Custom Docker Template Use a template when you want pods to start from a specific Docker image with predefined ports, environment variables, volume paths, and a startup command. This example creates a private template from a known PyTorch image, starts a Python HTTP server from the template's startup command, prints the mapped public URL, checks an environment variable, and stops the pod. ```python #!/usr/bin/env python3 """Create a custom template, launch a pod from it, and inspect runtime settings.""" from datetime import datetime, timezone from lium.sdk import Lium lium = Lium() ready_pod = None created_pod_id = None template_name = "sdk-custom-template-" + datetime.now(timezone.utc).strftime("%Y%m%d%H%M%S") try: template = lium.create_template( name=template_name, docker_image="daturaai/pytorch", docker_image_tag="2.4.0-py3.11-cuda12.4.1-devel-ubuntu22.04", ports=[22, 8000], start_command="python -m http.server 8000 --bind 0.0.0.0 --directory /workspace", environment={ "LIUM_EXAMPLE_MODE": "custom-template", }, volumes=["/workspace"], description="Private SDK example template", is_private=True, ) template = lium.wait_template_ready(template.id, timeout=600) if template is None: raise RuntimeError("Template did not verify before the timeout") executors = lium.ls( gpu_type="A100", gpu_count=1, min_cuda_version=12.4, ) if not executors: raise RuntimeError("No matching executor is currently available") executor = min(executors, key=lambda item: item.price_per_hour) pod = lium.up( executor_id=executor.id, template_id=template.id, name="sdk-custom-template-demo", ports=2, ) created_pod_id = pod["id"] ready_pod = lium.wait_ready(pod, timeout=600) if ready_pod is None: raise RuntimeError("Pod did not become ready before the timeout") details = lium.pod(ready_pod.id) ports_mapping = details["ports_mapping"] app_port = ports_mapping.get("8000") if app_port is None: raise RuntimeError(f"Port 8000 was not exposed: {ports_mapping}") result = lium.exec(ready_pod, command="printenv LIUM_EXAMPLE_MODE") if not result["success"]: raise RuntimeError(result["stderr"] or result["stdout"]) print(result["stdout"].strip()) print(f"Open http://{ready_pod.host}:{app_port}/") finally: pod_to_stop = ready_pod if pod_to_stop is None and created_pod_id: pod_to_stop = next((p for p in lium.ps() if p.id == created_pod_id), None) if pod_to_stop is not None: lium.down(pod_to_stop) print(f"Stopped pod {pod_to_stop.name} ({pod_to_stop.huid})") ``` ## What Each Field Does - `docker_image` and `docker_image_tag` choose the container image for the template. - `ports` declares internal container ports that Lium should expose. Include `22` for SSH access, plus application ports such as `8000` for an API server. - `start_command` is the command run when the container starts. This example starts Python's built-in HTTP server on `0.0.0.0:8000` so the service can be reached through Lium's public port mapping. - `environment` sets container environment variables. - `volumes` declares mount paths inside the container. - `is_private=True` keeps the template scoped to your account. When you create a pod, request enough public ports for the SSH port and the application port. Fetch the pod with `lium.pod(...)` and read `ports_mapping` to see the assigned host ports. --- # Remote Inference with @lium.machine # Remote Inference with `@lium.machine` This example runs a local Python function on a GPU pod. The decorator provisions the pod, uploads the function, installs Python dependencies, runs inference, returns the result, and tears the pod down. Authenticate before running the script: ```bash lium init ``` For CI, scripts, or temporary overrides, set `LIUM_API_KEY` instead. ```python #!/usr/bin/env python3 """Run a small instruct model on a remote Lium GPU.""" import lium @lium.machine(machine="A100", requirements=["torch", "transformers", "accelerate"]) def infer(prompt: str) -> str: from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "HuggingFaceTB/SmolLM2-135M-Instruct" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained( model_id, torch_dtype="auto", device_map="cuda", ) messages = [{"role": "user", "content": prompt}] text = tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, ) inputs = tokenizer(text, return_tensors="pt").to(model.device) outputs = model.generate( **inputs, max_new_tokens=64, do_sample=False, pad_token_id=tokenizer.eos_token_id, ) response_tokens = outputs[0][inputs["input_ids"].shape[-1]:] return tokenizer.decode(response_tokens, skip_special_tokens=True).strip() print(infer("In one sentence, who discovered penicillin?")) ``` The exact generated text can vary by model version, but it should identify Alexander Fleming. --- # Run a vLLM Server # Run a vLLM Server This example creates a reusable vLLM template, starts a GPU pod, and prints the OpenAI-compatible `/v1/models` endpoint. Authenticate before running the script: ```bash lium init ``` For CI, scripts, or temporary overrides, set `LIUM_API_KEY` instead. For gated Hugging Face models, also set `HF_TOKEN`. ```python #!/usr/bin/env python3 """Run vLLM on a Lium GPU pod.""" import os from lium.sdk import Lium lium = Lium() GPU_TYPE = "A100" # Volume paths cache Hugging Face models between pods. template = lium.create_template( name="vllm-smollm", docker_image="vllm/vllm-openai", ports=[22, 8000], environment={ "HF_HOME": "/root/.cache/huggingface", "HF_TOKEN": os.environ.get("HF_TOKEN", ""), }, # The vllm/vllm-openai image expects API-server args, not a shell command. start_command="--model HuggingFaceTB/SmolLM-135M --host 0.0.0.0 --port 8000", volumes=["/workspace", "/root/.cache/huggingface"], ) executors = lium.ls(gpu_type=GPU_TYPE) if not executors: raise RuntimeError(f"No {GPU_TYPE} executors are currently available") first = executors[0] pod = lium.up(executor_id=first.id, template_id=template.id) ready_pod = lium.wait_ready(pod["id"], timeout=600) if not ready_pod: raise RuntimeError("Pod did not become ready before the timeout") pod = lium.pod(pod_id=pod["id"]) if "8000" in pod["ports_mapping"]: port = pod["ports_mapping"]["8000"] host = pod["executor"]["executor_ip_address"] print("vLLM is loading. Check the model endpoint:") print(f"http://{host}:{port}/v1/models") ``` Once vLLM finishes loading, check the endpoint printed by the script: ```bash curl http://:/v1/models ``` You can later change the model name: ```python lium.edit( pod["id"], startup_commands="--model meta-llama/Llama-3.2-1B-Instruct --host 0.0.0.0 --port 8000", environment={ "HF_HOME": "/root/.cache/huggingface", "HF_TOKEN": os.environ.get("HF_TOKEN", ""), }, ) ``` Stop it when you are done: ```python lium.down(ready_pod) ``` --- # SDK Reference # SDK Reference Reference documentation for the public `lium.sdk` Python API. Install the SDK with: ```bash pip install lium.io ``` ## Client - [Lium](/developers/sdk/reference/client/lium) - Clean Unix-style SDK for Lium. - [AlphaQuote](/developers/sdk/reference/client/alpha-quote) - USD -> alpha quote from GET /balance/convert/alpha. ## Configuration - [Config](/developers/sdk/reference/config) ## Models - [ExecutorInfo](/developers/sdk/reference/models/executor-info) - [PodInfo](/developers/sdk/reference/models/pod-info) - [Template](/developers/sdk/reference/models/template) - Template information. - [BackupConfig](/developers/sdk/reference/models/backup-config) - Backup configuration information. - [BackupLog](/developers/sdk/reference/models/backup-log) - Backup log information. - [RestoreLog](/developers/sdk/reference/models/restore-log) - Restore log information. - [VolumeInfo](/developers/sdk/reference/models/volume-info) - Volume information. - [SSHKey](/developers/sdk/reference/models/ssh-key) - Public SSH key registered for the current user. ## Exceptions - [LiumError](/developers/sdk/reference/exceptions/lium-error) - Base exception for Lium SDK. - [LiumAuthError](/developers/sdk/reference/exceptions/lium-auth-error) - Authentication error. - [LiumRateLimitError](/developers/sdk/reference/exceptions/lium-rate-limit-error) - Rate limit exceeded. - [LiumServerError](/developers/sdk/reference/exceptions/lium-server-error) - Server error. - [LiumNotFoundError](/developers/sdk/reference/exceptions/lium-not-found-error) - Resource not found (404). ## Decorators - [machine](/developers/sdk/reference/decorators/machine) - Decorator to execute a function on a remote Lium machine. --- # CLI Overview # CLI Overview The Lium CLI is a powerful command-line interface that allows you to manage GPU pods directly from your terminal. It provides comprehensive control over pod creation, management, and monitoring without needing to use the web interface. ## Key Features ### Resource Management - List available GPU nodes with real-time pricing - Create and manage pods with customizable configurations - Monitor pod status and resource usage - Manage multiple pods simultaneously ### File Operations - Copy files to and from pods - Synchronize directories with rsync - Execute commands remotely on pods ### Development Workflow - SSH access to pods for interactive development - Template management for common configurations - Execute commands remotely without SSH ### Provider-side Automation - Browserless parity with the [provider portal](/providers/portal/overview): portal session, node lifecycle, central-miner-server config, sync, billing, and machine-request queries - Scripted with `--json` envelopes — see [`lium provider`](./reference/provider.md) ## How It Works The Lium CLI interacts with the Lium platform through: - **API Layer**: RESTful API for pod management - **SDK Integration**: Built on lium-sdk for robust functionality - **SSH Protocol**: Secure shell access to pods - **Template System**: Pre-configured Docker templates for common workloads ## Getting Started To get started with the Lium CLI: 1. **[Installation](./installation)** — Install and configure the CLI 2. **[Quick Start](./quickstart)** — Step-by-step guide for your first pod 3. **[Reference](./reference/)** — Complete reference for every CLI command, grouped by category ## Related - **[OpenAPI spec](../openapi)** — every REST endpoint the CLI calls under the hood - **[SDK](../sdk)** — typed Python client packaged with `lium.io` - **[AI Agents](../agents)** — wire the CLI into a shell-driven agent loop --- # Installation # Installation This guide covers the installation and initial setup of the Lium CLI. ## Prerequisites - **Operating system**: Linux or macOS (`amd64` or `arm64`). Windows users should run the binary install inside WSL, or use the Python package. - **Network**: Internet connection for API access and binary download. - **SSH client**: Required for pod access (`ssh`, `scp`, `rsync`). - **Python**: Only required if you install via `pip` — version 3.10 or higher, below 3.15. ## Installation methods ### Binary install (recommended) The fastest way to get a working `lium` command — no Python required. The installer downloads a precompiled binary for your platform, verifies its checksum, and adds it to your `PATH`: ```bash curl -fsSL https://lium.io/install.sh | bash ``` Supported platforms: `darwin-amd64`, `darwin-arm64`, `linux-amd64`, `linux-arm64`. The installer: 1. Downloads the latest release of `lium` from [GitHub releases](https://github.com/Datura-ai/lium/releases) into `~/.lium/versions//lium`. 2. Verifies the SHA-256 checksum against the published `checksums.txt`. 3. Creates a managed symlink at `~/.lium/bin/lium` pointing to the versioned binary. 4. Adds `~/.lium/bin` to `PATH` in your `~/.bashrc`, `~/.zshrc`, or `~/.config/fish/config.fish`. After the install finishes, restart your shell (or `source` your rc file) so the updated `PATH` takes effect: ```bash exec $SHELL -l # or: source ~/.bashrc / ~/.zshrc lium --version ``` #### Customizing the binary install The installer respects a few environment variables for non-default setups: | Variable | Purpose | |----------|---------| | `LIUM_INSTALL_DIR` | Override the symlink directory (default `~/.lium/bin`). | | `LIUM_VERSION` | Pin a specific release (e.g. `LIUM_VERSION=0.0.7`). | Example — pin a specific version: ```bash curl -fsSL https://lium.io/install.sh | LIUM_VERSION=0.0.7 bash ``` #### Upgrading Re-run the same command to upgrade to the latest release. The installer will download the newer versioned binary and re-point the `~/.lium/bin/lium` symlink at it; older versions remain under `~/.lium/versions/` until you remove them. ```bash curl -fsSL https://lium.io/install.sh | bash ``` #### Uninstalling ```bash rm -rf ~/.lium/bin ~/.lium/versions # Optional: remove the PATH line the installer added to your shell rc ``` ### Python package Install the published package from PyPI when you also want the Python SDK alongside the CLI, or when you can't run a precompiled binary: ```bash pip install lium.io ``` To upgrade: ```bash pip install --upgrade lium.io ``` It's recommended to install into a virtual environment to avoid conflicts: ```bash python -m venv lium-env # Linux/macOS: source lium-env/bin/activate # Windows: lium-env\Scripts\activate pip install lium.io ``` ### From source For development or to track unreleased changes: ```bash git clone https://github.com/Datura-ai/lium.git cd lium pip install -e . ``` ## Initial configuration ### No account yet Create one from the terminal — see [`lium signup`](./reference/signup): ```bash lium signup --email you@example.com ``` It creates the account, saves the API key and sets up a local SSH key, so nothing below is needed afterwards. Click the confirmation link in the email before you rent. ### Running the setup wizard If you already have an account, run: ```bash lium init ``` This interactive wizard helps you configure: 1. **API key** — your Lium platform API key. 2. **SSH keys** — generate a new pair or point at an existing one. 3. **Default settings** — preferences such as theme. ### Manual configuration You can also edit `~/.lium/config.ini` directly: ```ini [api] api_key = your-api-key-here [ssh] key_path = /path/to/your/ssh/key ``` ### Environment variables Any config key can be supplied through the environment — `LIUM_` plus the dotted key in upper case, with dots as underscores. Useful for CI: ```bash export LIUM_API_KEY=your-api-key-here # api.api_key export LIUM_SSH_KEY_PATH=/path/to/ssh/key # ssh.key_path ``` Note that the SSH connection itself does not read `ssh.key_path`: the client picks the first of `~/.ssh/id_ed25519`, `~/.ssh/id_rsa`, `~/.ssh/id_ecdsa` that exists. Keep the key you want used at one of those paths. ## Verification ```bash lium --version # confirm install lium config show # confirm config lium ls # list available nodes ``` ## Getting your API key 1. Register at [lium.io](https://lium.io/register). 2. Open your account settings and find the API section. 3. Generate an API key. 4. Paste it when `lium init` prompts for it (or export `LIUM_API_KEY`). ## SSH key setup ### Generating new SSH keys If you don't have SSH keys, the CLI can generate them — choose **Generate new SSH key pair** in `lium init`. ### Using existing keys 1. Locate your key pair (typically `~/.ssh/`). 2. During `lium init`, provide the path to your private key. 3. The CLI uploads the matching public key to Lium. ### Key permissions ```bash chmod 600 ~/.ssh/id_ed25519 chmod 644 ~/.ssh/id_ed25519.pub ``` ## Next steps - [Quick Start Guide](./quickstart) — start using the CLI - [CLI Reference](./reference/) — every command, grouped by category --- # Quick Start Guide # Quick Start Guide This guide will walk you through your first pod creation and management workflow with the Lium CLI. ## Prerequisites Before starting, make sure you have installed and configured the Lium CLI. See the [Installation Guide](./installation) for detailed instructions. ## Step 1: Browse Available GPUs List available GPU nodes: ```bash lium ls ``` This shows a table with: - Node index numbers - GPU types (H100, A100, RTX 4090, etc.) - Pricing per hour - Available memory and storage - Pareto-optimal choices marked with ★ Filter by GPU type: ```bash lium ls --gpu H100 # Show only H100 GPUs lium ls --gpu A100 # Show only A100 GPUs ``` ## Step 2: Create Your First Pod Create a pod using a node from the list: ```bash lium up 1 # Uses node #1 from the list ``` You'll be prompted to select a template from the available options. You can filter the templates during selection. ### Filtering Nodes You can filter nodes automatically using various criteria: ```bash # Filter by GPU type lium up --gpu H200 # Filter by GPU type and count lium up --gpu A6000 -c 2 # Filter by country lium up --country US # Combine filters lium up --gpu H200 --country FR # Ensure minimum available ports (and allocate them) lium up --ports 5 ``` ## Step 3: Check Pod Status View your active pods: ```bash lium ps ``` This displays: - Pod names and IDs - Status (running, stopped, etc.) - Uptime and costs - SSH connection details ## Step 4: Connect to Your Pod ### SSH Access Connect via SSH: ```bash lium ssh my-pod ``` Or use the pod number from `lium ps`: ```bash lium ssh 1 ``` ### Execute Commands Run commands without SSH: ```bash lium exec my-pod "nvidia-smi" lium exec my-pod "python --version" ``` ## Step 5: Transfer Files ### Copy Files to Pod Copy a single file: ```bash lium scp my-pod ./script.py ``` Copy to specific location: ```bash lium scp my-pod ./data.csv /root/datasets/ ``` Copy to multiple pods: ```bash lium scp 1,2,3 ./model.py lium scp all ./config.json # All pods ``` ### Sync Directories Synchronize entire directories: ```bash lium rsync my-pod ./project lium rsync my-pod ./data /root/datasets/ ``` ## Step 6: Stop Your Pod Remove a pod when done: ```bash lium rm my-pod ``` Remove multiple pods: ```bash lium rm pod1 pod2 pod3 lium rm all # Remove all pods ``` ## Complete Example Workflow Here's a complete machine learning training workflow: ```bash # 1. List available GPUs and choose one lium ls --gpu A100 # 2. Create a pod with filters (auto-selects best matching node) lium up --gpu A100 --country US --name ml-training # Or create with a specific node lium up 1 --name ml-training # 3. Copy your training code lium scp ml-training ./train.py lium rsync ml-training ./data /root/datasets/ # 4. Install dependencies lium exec ml-training "pip install -r requirements.txt" # 5. Start training lium exec ml-training "python train.py --epochs 100" # 6. Monitor progress lium ssh ml-training # Inside pod: tail -f training.log # 7. Copy results back (from local machine) scp root@:/root/models/best_model.pt ./ # 8. Clean up lium rm ml-training ``` ## Using Templates List available templates: ```bash lium templates ``` Search for specific templates: ```bash lium templates pytorch lium templates tensorflow ``` Create pod and select template: ```bash lium up 1 # You'll be prompted to select from available templates ``` ## Cost Management ### Monitor Spending Check current costs: ```bash lium ps # Shows hourly rate and total spent ``` ### Fund Your Account Add funds from a Bittensor wallet (TAO): ```bash lium fund # Interactive mode lium fund -w my-wallet -a 10.0 # Fund 10 TAO ``` Or top up with a stablecoin — `lium topup` returns a deposit address to send to (ideal for agents that self-fund): ```bash lium topup create -a 20 -c USDT -n tron # $20 invoice; prints deposit address ``` See [`lium topup`](./reference/topup.md) for the full self-funding loop. ## Tips and Best Practices ### 1. Use Pareto-Optimal Nodes Look for ★ symbols in `lium ls` - these offer the best price/performance. ### 2. Name Your Pods Use descriptive names for easier management: ```bash lium up 1 --name experiment-bert-v2 ``` ### 3. Batch Operations Copy files to multiple pods efficiently: ```bash lium scp all ./updated_config.json ``` ### 4. Monitor Resources Check GPU utilization: ```bash lium exec my-pod "nvidia-smi" ``` ### 5. Clean Up Always remove pods when done to avoid charges: ```bash lium rm all ``` ## Next Steps - [CLI Reference](./reference/) — Per-command documentation grouped by category - [Installation Guide](./installation) — Installation and setup details --- # CLI Reference # CLI Reference Complete reference for every `lium` command, grouped by what you're trying to do. Each command has its own page with arguments, options, and runnable examples. If you're new, start at the [Quickstart](../quickstart.md) and circle back here once you know which command you need. ## Global Options Available on every command: | Flag | Effect | |------|--------| | `--help` | Show help for the command | | `--version` | Print CLI version (on the root `lium` command) | ## Pods Day-to-day pod lifecycle, access, and observation. | Command | Purpose | |---------|---------| | [`lium ls`](./ls.md) | List available GPU nodes | | [`lium up`](./up.md) | Create a new pod | | [`lium ps`](./ps.md) | List your active pods | | [`lium ssh`](./ssh.md) | Open an SSH session | | [`lium exec`](./exec.md) | Run a command on a pod | | [`lium scp`](./scp.md) | Copy files to/from a pod | | [`lium rsync`](./rsync.md) | Sync directories to a pod | | [`lium rm`](./rm.md) | Stop / remove pods | | [`lium reboot`](./reboot.md) | Reboot one or more pods | | [`lium logs`](./logs.md) | Stream or tail pod logs | | [`lium update`](./update.md) | Update a running pod (e.g. install Jupyter) | | [`lium port-forward`](./port-forward.md) | Forward a pod port to your local machine | | [`lium schedules`](./schedules.md) | List / cancel auto-termination schedules | ## Storage Persistent volumes and pod backups. | Command | Purpose | |---------|---------| | [`lium volumes`](./volumes.md) | Manage persistent volumes (`list` / `new` / `rm`) | | [`lium bk`](./bk.md) | Pod backups: create, schedule, restore, list, delete | ## Templates | Command | Purpose | |---------|---------| | [`lium templates`](./templates.md) | List and search Docker templates and images | ## Account & Config | Command | Purpose | |---------|---------| | [`lium signup`](./signup.md) | Create an account from the terminal and store its API key | | [`lium init`](./init.md) | First-time setup for an existing account (API key + SSH key) | | [`lium config`](./config.md) | Read / write CLI configuration | | [`lium theme`](./theme.md) | Change CLI color theme | | [`lium balance`](./balance.md) | Show your current account balance | | [`lium fund`](./fund.md) | Fund your account with TAO or free Subnet 51 alpha from a Bittensor wallet | | [`lium topup`](./topup.md) | Top up your balance with a stablecoin (agent self-funding) | | [`lium ssh-keys`](./ssh-keys.md) | List / sync SSH keys on your account | ## Providers Commands for users running provider nodes on Subnet 51. See also the [Providers section](/providers/) for setup and architecture. | Command | Purpose | |---------|---------| | [`lium mine`](./mine.md) | Bootstrap a provider node end-to-end | | [`lium gpu-splitting`](./gpu-splitting.md) | Prepare Docker storage for multi-tenant rental | | [`lium provider`](./provider.md) | Full provider-portal automation: portal session, node lifecycle, central-miner-server config, sync, billing, and machine-request queries | ## Patterns That Work Across Commands ### Pod selectors Most pod-targeting commands (`ssh`, `exec`, `scp`, `rsync`, `rm`, `reboot`) accept any of: - A pod name: `lium ssh my-pod` - An index from `lium ps`: `lium ssh 1` - A comma-separated list: `lium exec 1,2,3 "uptime"` - The keyword `all`: `lium scp all ./config.json` ### Output formats Listing commands accept `--format` (`table` or `json`): ```bash lium ls --format json | jq '.[] | .gpu_type' lium ps --format json > pods.json ``` ### Environment variables Override config without editing files: ```bash LIUM_API_KEY=xyz lium ls ``` For provider commands, `LIUM_PROVIDER_COLDKEY` and `LIUM_PROVIDER_HOTKEY` set the wallet identity (or persist them via `lium config set provider.coldkey ...` / `lium config set provider.hotkey ...`). `LIUM_PORTAL_URL` overrides the portal base URL, and `LIUM_PROVIDER_ACK=1` auto-confirms the persona gate on spend-affecting `lium provider …` subcommands. See [`lium provider`](./provider.md) for the full surface. ## Exit codes | Code | Meaning | |------|---------| | `0` | Success | | `1` | General error | | `2` | Configuration error — bad arguments, unreadable script, missing configuration | | `4` | SSH error | | `5` | Pod not found | Two exceptions: - `lium exec` exits with the **remote command's** exit code, so `lium exec my-pod "cmd" && next-step` chains the way you expect. - `lium provider …` has [its own exit-code map](./provider.md#exit-codes--errors) — the same numbers carry different meanings there, and it emits `3`, `6` and `7`, which this table does not list. ## See also - [Quick Start](../quickstart.md) — first pod in 5 minutes - [Installation](../installation.md) — install and authenticate - [SDK](../../sdk.md) — Python client used under the hood - [OpenAPI](../../openapi.md) — REST endpoints the CLI calls --- # `lium ls` # `lium ls` List available GPU nodes. ```bash lium ls [OPTIONS] ``` ## Options | Flag | Effect | |------|--------| | `--gpu TYPE` | Filter by GPU type (`H100`, `A100`, `RTX4090`, …) | | `--count N` | Exact GPU count to match (e.g. `1`, `8`) | | `--min-cuda VERSION` | Minimum CUDA version supported by the node (e.g. `12.4`) | | `--lat LAT` | Latitude for distance filtering | | `--lon LON` | Longitude for distance filtering | | `--max-distance MILES` | Maximum distance in miles from `--lat`/`--lon` | | `--sort FIELD` | Sort by `download` (default), `upload`, `price_gpu`, `price_total`, `loc`, `id`, or `gpu` | | `--limit N` | Limit the number of rows shown | | `--format FORMAT` | Output format: `table` (default) or `json` | ## Examples ```bash lium ls # List all nodes lium ls --gpu H100 # Only H100 GPUs lium ls --min-cuda 12.4 # Only nodes supporting CUDA 12.4 or higher lium ls --sort price_gpu # Cheapest $/GPU·h first lium ls --limit 10 # Show only the first 10 rows lium ls --format json # JSON output ``` Pareto-optimal choices in the table are marked with ★. ## Tier column Each node shows its **Tier**, so you can see the interruption risk before renting: - **secure** — the default. If the rental is interrupted, the provider is penalized (their rental fee is withheld), so these nodes are the stable choice. - **spot** — the provider can reclaim the node at any time with no penalty, and you are not compensated if the pod is interrupted. Use spot for cheaper, interruptible, fault-tolerant workloads. With `--format json`, the same value is exposed as the `tier` field on each node (`"spot"`, `"secure"`, or `null` when unknown). See [Node Tier: Secure vs Spot](/providers/portal/node-tier) for the full comparison. ## See also - [`lium up`](./up) — create a pod on a listed node - [`lium ps`](./ps) — list pods you already own --- # `lium up` # `lium up` Create a new pod. ```bash lium up [NODE_ID] [OPTIONS] ``` ## Arguments - `NODE_ID` — Node UUID, HUID, or index from the last [`lium ls`](./ls.md). Optional: when you omit it, the filter flags below narrow the list and the best node is picked **automatically**, with no menu — the only prompt is the confirmation before renting, which `--yes` skips. ## Options ### Selection & naming | Flag | Effect | |------|--------| | `--name, -n NAME` | Custom pod name (auto-generated if not specified) | | `--template_id, -t ID` | Specify template ID directly (skips template prompt) | | `--gpu TYPE` | Filter nodes by GPU type (e.g. `H200`, `A6000`) | | `--count, -c NUM` | Filter by number of GPUs per pod | | `--country CODE` | Filter nodes by ISO country code (`US`, `FR`, …) | | `--ports, -p NUM` | Only consider nodes with at least NUM available ports, and request NUM ports on the pod | | `--yes, -y` | Skip confirmation prompt | | `--no-ssh` | Create the pod and return, instead of opening an SSH session | By default a successful `lium up` ends by opening an interactive SSH session to the new pod (in `--image` mode it streams the container logs instead). Pass `--no-ssh` when you want the command to return — in a script, or when you plan to connect later with [`lium ssh`](./ssh.md). ### Auto-termination | Flag | Effect | |------|--------| | `--ttl DURATION` | Auto-terminate after duration (e.g. `6h`, `45m`, `2d`) | | `--until WHEN` | Auto-terminate at local time (`today 23:00`, `tomorrow 01:00`, `2026-05-08 15:30`) | Use [`lium schedules`](./schedules.md) to view and cancel an existing schedule. ### Container | Flag | Effect | |------|--------| | `--image IMAGE` | Docker image to run (e.g. `pytorch/pytorch:2.0`, `nvidia/cuda:12.0`) | | `--dockerfile PATH` | Build the pod image from a local Dockerfile (custom build). Mutually exclusive with `--image` and `--template_id`. | | `--internal-ports LIST` | Internal ports to expose (comma-separated, e.g. `22,8000,8080`) | | `-e, --env KEY=VAL` | Environment variable; can be repeated | | `--entrypoint CMD` | Override container entrypoint | | `--cmd CMD` | Override container command | | `--jupyter` | Install Jupyter Notebook (auto-selects an available port) | | `--ssh-name NAME` | Name for a freshly registered SSH key (default: `cli-@`) | **Custom Dockerfile builds** (`--dockerfile`): the image is built **on the node** from your Dockerfile — no build context is uploaded, so the Dockerfile must be self-contained. Network access is disabled during the build (`--network=none`) and `ADD ` / `ADD ${var}` directives are rejected. Keep the file under 64 KiB. The pod is SSH-accessible once the build finishes, just like a template-based pod. The `--env`, `--entrypoint`, `--cmd`, and `--internal-ports` flags don't apply here — define those in the Dockerfile itself. ### Volumes | Flag | Effect | |------|--------| | `--volume, -v SPEC` | Volume spec — see below | | `--volume-encryption` / `--no-volume-encryption` | Encrypt the local volume where the node supports it. Enabled by default. | Volume spec forms: - `id:` — attach an existing volume by HUID - `new:name=[,desc=]` — create and attach a new volume ## Examples ```bash # Basic lium up # Auto-picks the best node, asks to confirm lium up cosmic-hawk-f2 # Specific node by HUID lium up 1 # Node #1 from last ls lium up --name dev-pod # Custom name # Filtering nodes lium up --gpu H200 lium up --gpu A6000 -c 2 lium up --gpu H200 --country FR lium up --ports 5 # Auto-terminate lium up --gpu H200 --ttl 6h # Stop after 6h lium up --gpu H200 --until "tomorrow 09:00" # Stop tomorrow morning # Container customization lium up 1 --image pytorch/pytorch:2.0 lium up 1 --image my-image -e API_KEY=xyz -e DEBUG=1 lium up 1 --internal-ports 22,8000,8080 lium up 1 --jupyter # Spin up with Jupyter # Custom Dockerfile build (image built remotely from your Dockerfile) lium up --gpu A6000 --dockerfile ./Dockerfile lium up 1 --dockerfile ./Dockerfile --name my-build # Volumes lium up 1 --volume id:brave-fox-3a lium up 1 --volume new:name=my-data,desc="Training data" # Direct template + skip prompt lium up 1 --template_id abc123 lium up --gpu H200 --yes ``` ## See also - [`lium ls`](./ls.md) — find a node first - [`lium templates`](./templates.md) — pick a template - [`lium volumes`](./volumes.md) — pre-create a volume - [`lium schedules`](./schedules.md) — manage `--ttl` / `--until` schedules --- # `lium ps` # `lium ps` List your active pods. ```bash lium ps [POD_ID] [OPTIONS] ``` ## Arguments - `POD_ID` — Optional pod name, index, or id. When given, shows just that pod. ## Options | Flag | Effect | |------|--------| | `--format FORMAT` | Output format: `table` (default) or `json` | ## Examples ```bash lium ps # All your pods lium ps my-pod # A single pod by name/index/id lium ps --format json # JSON output ``` The numeric index in the first column is reused by [`lium ssh`](./ssh), [`lium exec`](./exec), [`lium scp`](./scp), [`lium rsync`](./rsync), and [`lium rm`](./rm). ## See also - [`lium up`](./up) — create a pod - [`lium rm`](./rm) — stop a pod --- # `lium ssh` # `lium ssh` SSH into a pod. ```bash lium ssh TARGET ``` ## Arguments - `TARGET` — Pod name or index from [`lium ps`](./ps). ## Examples ```bash lium ssh my-pod # Interactive shell lium ssh 1 # SSH to pod #1 ``` To run a one-shot command without holding a shell, use [`lium exec`](./exec) instead. ## See also - [`lium exec`](./exec) — run a command without holding a TTY - [`lium scp`](./scp) / [`lium rsync`](./rsync) — move files --- # `lium exec` # `lium exec` Execute a command on one or more pods. ```bash lium exec TARGETS [COMMAND] [OPTIONS] ``` ## Arguments - `TARGETS` — Pod name, index, comma-separated list, or `all`. - `COMMAND` — The command to execute. Optional when you pass `--script`. ## Options | Flag | Effect | |------|--------| | `--script, -s PATH` | Execute a local script file on the pod(s) instead of an inline command | | `--env, -e KEY=VALUE` | Set an environment variable for the command (repeatable) | | `--json` | Print machine-readable JSON instead of raw output | `lium exec` exits with the remote command's own exit code, so `lium exec my-pod "cmd" && next-step` behaves the way you expect. The `--json` payload is always an envelope with an array, even for a single pod: ```json { "ok": true, "results": [ {"pod": "eager-wolf-aa", "stdout": "…", "stderr": "", "exit_code": 0, "error": null} ] } ``` `ok` is true only when every pod succeeded. Read a single result with `jq -r '.results[0].stdout'` — there is no top-level `stdout`. ## Examples ```bash lium exec my-pod "python train.py" lium exec 1 "nvidia-smi" lium exec 1,2,3 "apt update" # Multiple pods lium exec all "pip install numpy" # Every running pod lium exec my-pod --script ./setup.sh # Run a local script lium exec my-pod -e WANDB_MODE=offline "python train.py" lium exec my-pod --json "python train.py" # Parse the result programmatically ``` ## See also - [`lium ssh`](./ssh) — interactive shell - [`lium scp`](./scp) — push code first, then `exec` --- # `lium scp` # `lium scp` Copy files between your machine and pods. Upload is the default; add `--download` to pull files back. ```bash lium scp TARGETS SOURCE_PATH [DESTINATION_PATH] [OPTIONS] ``` ## Arguments - `TARGETS` — Pod name, index, comma-separated list, or `all`. - `SOURCE_PATH` — On upload, a local file. On download (`-d`), a remote path. - `DESTINATION_PATH` — Optional. On upload, the remote path (default: `/root/`). On download, the local path. ## Options | Flag | Effect | |------|--------| | `--download, -d` | Download from the pod to your local machine instead of uploading | ## Examples ```bash lium scp my-pod ./script.py # Upload → /root/script.py lium scp 1 ./data.csv /root/datasets/ # Upload to a specific dir lium scp all ./config.json # Every running pod lium scp 1,2,3 ./model.py /root/models/ # Multiple pods lium scp 2 /root/output.log ./outputs -d # Download from pod #2 ``` ## See also - [`lium rsync`](./rsync) — for repeated syncs and large directories --- # `lium rsync` # `lium rsync` Synchronize directories to pods. ```bash lium rsync TARGETS LOCAL_PATH [REMOTE_PATH] ``` ## Arguments - `TARGETS` — Pod name, index, list, or `all`. - `LOCAL_PATH` — Local directory to sync (must exist). - `REMOTE_PATH` — Optional destination path on the pod. ## Examples ```bash lium rsync my-pod ./project # Sync project lium rsync 1 ./data ~/datasets/ # Specific path lium rsync all ./configs # Sync to every running pod ``` ## See also - [`lium scp`](./scp) — single-shot copy --- # `lium rm` # `lium rm` Remove / stop pods. ```bash lium rm [TARGETS] [OPTIONS] ``` ## Arguments - `TARGETS` — Pod name, index, comma-separated list, or `all`. ## Options | Flag | Effect | |------|--------| | `--all, -a` | Remove all active pods | | `--yes, -y` | Skip the confirmation prompt on `--all` | | `--in DURATION` | Schedule removal after a duration (e.g. `6h`, `45m`, `2d`) | | `--at TIME` | Schedule removal at a local time (e.g. `"tomorrow 09:00"`) | Removal is irreversible. The command exits non-zero when nothing matched `TARGETS`, so a typo cannot look like a successful teardown. Only `lium rm --all` asks for confirmation, and only on a terminal — that is the one place `--yes` changes anything. Removing a named pod, an index or a list never prompts, with or without `-y`. ## Examples ```bash lium rm my-pod # Remove a single pod lium rm 1,2,3 # Remove several lium rm all # Remove every active pod (positional) lium rm --all # Same, via the flag lium rm --all -y # Remove everything without the confirmation lium rm my-pod --in 6h # Schedule removal in 6 hours lium rm my-pod --at "tomorrow 09:00" ``` ## See also - [`lium schedules`](./schedules.md) — view / cancel scheduled removals (created via `lium up --ttl`) - [`lium volumes`](./volumes.md) — manage detached volumes --- # `lium reboot` # `lium reboot` Reboot one or more pods. ```bash lium reboot [TARGETS] [OPTIONS] ``` ## Arguments - `TARGETS` — Pod name, index, or comma-separated list. Optional when `--all` is set. ## Options | Flag | Effect | |------|--------| | `-a, --all` | Reboot every active pod | | `--volume-id HUID` | Volume ID to attach when rebooting | ## Examples ```bash lium reboot my-pod lium reboot 1,2,3 lium reboot --all lium reboot my-pod --volume-id brave-fox-3a # Attach a volume on reboot ``` ## See also - [`lium ps`](./ps.md) — find the pod to reboot - [`lium rm`](./rm.md) — fully stop a pod instead of rebooting --- # `lium logs` # `lium logs` Stream or tail pod logs. ```bash lium logs POD_ID [OPTIONS] ``` ## Arguments - `POD_ID` — Pod name or index from [`lium ps`](./ps.md). ## Options | Flag | Effect | |------|--------| | `-n, --tail N` | Number of lines to show from the end (default: `100`) | | `-f, --follow` | Follow log output (like `tail -f`) | ## Examples ```bash lium logs my-pod # Last 100 lines lium logs 1 --tail 500 # Last 500 lines lium logs my-pod -f # Stream live lium logs my-pod -n 50 -f # Tail and follow ``` ## See also - [`lium exec`](./exec.md) — run an ad-hoc command on a pod - [`lium ssh`](./ssh.md) — open an interactive shell --- # `lium update` # `lium update` Update a running pod in place. Currently handles installing Jupyter on a chosen internal port. ```bash lium update TARGET [OPTIONS] ``` ## Arguments - `TARGET` — Pod name or index. ## Options | Flag | Effect | |------|--------| | `--jupyter PORT` | Install Jupyter Notebook on the specified internal port | ## Examples ```bash lium update my-pod --jupyter 8888 lium update 1 --jupyter 8000 ``` ## See also - [`lium up`](./up.md) — `--jupyter` flag installs Jupyter at create time instead - [`lium port-forward`](./port-forward.md) — expose the internal Jupyter port locally --- # `lium port-forward` # `lium port-forward` Forward a port from a pod to your local machine over the SSH tunnel. ```bash lium port-forward TARGET PORT [OPTIONS] ``` ## Arguments - `TARGET` — Pod name or index. - `PORT` — Internal pod port to forward. ## Options | Flag | Effect | |------|--------| | `-l, --local-port PORT` | Local port to bind (default: same as internal port) | ## Examples ```bash lium port-forward my-pod 8888 # Forward 8888:8888 lium port-forward my-pod 8888 -l 9000 # Forward 8888 → local 9000 lium port-forward 1 22 -l 2222 # SSH on a custom local port ``` The forward stays open until you `Ctrl-C`. ## See also - [`lium update --jupyter`](./update.md) — open a Jupyter port on a running pod, then forward it - [`lium ssh`](./ssh.md) — interactive shell over the same tunnel --- # `lium schedules` # `lium schedules` List and cancel scheduled pod removals. Schedules are **created** via [`lium up --ttl`](./up.md) or `--until`; this command manages the existing ones. ```bash lium schedules SUBCOMMAND [ARGS] ``` Running `lium schedules` with no subcommand defaults to `list`. ## Subcommands ### `lium schedules list` List all scheduled pod removals on your account. ```bash lium schedules list ``` The output assigns each schedule a numeric index, used by `rm`. ### `lium schedules rm INDICES` Cancel one or more scheduled removals by index. ```bash lium schedules rm 1 lium schedules rm 1,2,5 # Multiple ``` Indices must be numbers (single or comma-separated); there is no `all` keyword. ## Examples ```bash # Schedule a pod to auto-stop in 6 hours (creates the schedule) lium up --gpu H200 --ttl 6h # Or at a specific local time lium up --gpu H200 --until "tomorrow 09:00" # List active schedules lium schedules list # Cancel one lium schedules rm 1 ``` ## See also - [Scheduled termination](/pod-users/scheduled-termination) — concept guide - [`lium up`](./up.md) — `--ttl` and `--until` flags - [`lium rm`](./rm.md) — stop a pod immediately instead --- # `lium volumes` # `lium volumes` Manage persistent volumes that can be attached to pods. ```bash lium volumes SUBCOMMAND [OPTIONS] ``` Running `lium volumes` with no subcommand defaults to `list`. ## Subcommands | Subcommand | Purpose | |------------|---------| | `list` | List your volumes | | `new NAME [--desc DESC]` | Create a new volume | | `rm INDICES [-y]` | Delete volumes by index from `list` | ### `lium volumes list` ```bash lium volumes list ``` The output assigns each volume a numeric index, used by `rm`. The HUID column is what you pass to [`lium up --volume id:`](./up.md). ### `lium volumes new NAME` | Flag | Effect | |------|--------| | `-d, --desc TEXT` | Volume description | ```bash lium volumes new training-data lium volumes new datasets --desc "ImageNet 2012" ``` ### `lium volumes rm INDICES` | Flag | Effect | |------|--------| | `-y, --yes` | Skip the confirmation prompt | ```bash lium volumes rm 1 # Delete volume #1 lium volumes rm 1,2,5 # Multiple lium volumes rm 1 -y # No confirmation ``` A volume must be detached before it can be deleted. ## Attaching to a pod ```bash lium up 1 --volume id:brave-fox-3a # Existing volume lium up 1 --volume new:name=training,desc="ML data" # Create + attach ``` ## See also - [Volumes](/pod-users/volumes) — concept guide - [`lium up`](./up.md) — attach a volume at pod creation --- # `lium bk` # `lium bk` Pod backups and restores. Each subcommand operates on a single pod (by HUID or index from `lium ps`). ```bash lium bk SUBCOMMAND POD_ID [OPTIONS] ``` ## Subcommands | Subcommand | Purpose | |------------|---------| | `now` | Trigger an immediate one-off backup | | `set` | Configure a recurring backup schedule for a pod | | `show` | Show the backup config and recent backups for a pod | | `restore` | Restore a backup into a pod | | `rm` | Delete the backup config for a pod | | `logs` | Show backup job logs | | `restore-logs` | Show restore job logs | ### `lium bk now POD_ID` Trigger a backup right now. | Flag | Effect | |------|--------| | `-n, --name NAME` | Backup name (e.g. `pre-release`) | | `-d, --description TEXT` | Backup description | ```bash lium bk now my-pod -n pre-release -d "snapshot before deploy" ``` ### `lium bk set POD_ID` Configure a recurring backup schedule. | Flag | Effect | |------|--------| | `--path PATH` | Backup path (default: `/root`) | | `--every DURATION` | Frequency (e.g. `1h`, `6h`, `24h`) | | `--keep DURATION` | Retention (e.g. `1d`, `7d`, `30d`) | | `-y, --yes` | Skip the confirmation prompt | ```bash lium bk set my-pod --every 6h --keep 7d lium bk set my-pod --path /workspace --every 24h --keep 30d -y ``` ### `lium bk show POD_ID` Show the backup config and history for a pod. ```bash lium bk show my-pod ``` ### `lium bk restore POD_ID --id BACKUP_ID` Restore a backup into a pod. | Flag | Effect | |------|--------| | `--id BACKUP_ID` | **Required.** Backup ID to restore | | `--to PATH` | Restore path (default: `/root`) | | `-y, --yes` | Skip confirmation | ```bash lium bk restore my-pod --id bkp-abc123 lium bk restore my-pod --id bkp-abc123 --to /workspace -y ``` ### `lium bk rm POD_ID` Delete the backup configuration for a pod. ```bash lium bk rm my-pod -y ``` ### `lium bk logs [POD_ID]` Show backup job logs. Pod ID optional; with `--id`, shows details for one backup. ```bash lium bk logs # Recent jobs across all pods lium bk logs my-pod # Jobs for one pod lium bk logs --id bkp-abc123 # Details for one backup ``` ### `lium bk restore-logs [POD_ID]` Show restore job logs for one pod. With `--id`, shows details for one restore. Specify either `POD_ID` or `--id`. | Flag | Effect | |------|--------| | `--id RESTORE_ID` | Restore log ID to inspect | ```bash lium bk restore-logs my-pod # Restore jobs for one pod lium bk restore-logs --id rst-abc123 # Details for one restore ``` ## See also - [Backups](/pod-users/backups) — concept guide and schedule policies - [Restores](/pod-users/restores) — concept guide - [`lium volumes`](./volumes.md) — for detached storage instead of pod backups --- # `lium templates` # `lium templates` List available Docker templates and images used to create pods. ```bash lium templates [SEARCH] ``` ## Arguments - `SEARCH` — Optional filter string. Templates and images whose name matches are shown; omit it to list everything. ## Examples ```bash lium templates # List all templates and images lium templates pytorch # Show entries matching "pytorch" ``` To create a pod from a template, pass its ID to [`lium up --template_id`](./up). ## See also - [Templates](/pod-users/templates) — concept guide - [`lium up`](./up) — pick a template at pod creation --- # `lium signup` # `lium signup` Create a Lium account from the terminal and store the API key it mints. Use this when you have no account yet. [`lium init`](./init) authenticates an account that already exists — it cannot create one. ```bash lium signup --email you@example.com ``` ## Options | Flag | Effect | |------|--------| | `--email ADDRESS` | Your real email address — the confirmation link is sent there. **Required.** | | `--name NAME` | Display name (defaults to the local part of the email) | | `--password PASSWORD` | Account password. One is generated when you omit it. | | `--json` | Print machine-readable JSON | Prefer the `LIUM_SIGNUP_PASSWORD` environment variable over `--password`: a flag value is left behind in your shell history and in `ps` output. ## What it does 1. Creates the account. 2. Writes the minted API key to `~/.lium/config.ini` under `api.api_key`. 3. Sets up an SSH key **locally**: picks the first of `~/.ssh/id_ed25519`, `id_rsa`, `id_ecdsa` that exists, generates `id_ed25519` when none does, and records the path in the config. Nothing is uploaded yet — the public key reaches your Lium account on the first [`lium up`](./up). After it finishes, [`lium ls`](./ls) and [`lium up`](./up) work with no further setup. The command prints the password — it is your login at [lium.io](https://lium.io) and is not stored anywhere else. ## Confirm your email before renting Renting stays blocked until you click the link in the **"Please confirm your email"** message. The separate "Welcome to Celium!" message carries no link. ## When an account already exists locally `lium signup` refuses to run while an API key is already in effect, rather than creating a second account you cannot reach. Where the key comes from decides how to clear it — the error message names the right one: ```bash # Key stored in ~/.lium/config.ini lium config unset api.api_key # Key exported in the environment — `lium config unset` will not help here unset LIUM_API_KEY lium signup --email you@example.com ``` ## Examples ```bash lium signup --email ada@example.com lium signup --email ada@example.com --name Ada lium signup --email ada@example.com --json LIUM_SIGNUP_PASSWORD=... lium signup --email ada@example.com ``` `--json` output: ```json { "api_key": "sk_...", "email": "ada@example.com", "next_steps": ["...", "...", "..."], "password": "generated-or-supplied", "signup_credit_granted": true, "ssh_key_configured": true } ``` `signup_credit_granted` answers whether the signup credit landed: `true` granted, `false` not granted, `null` not reported by the backend — read [`lium balance --json`](./balance) in that case. If the command fails *after* the account was created, the error still reports the email and password, so you can log in at [lium.io](https://lium.io) and copy an API key from the dashboard. ## See also - [`lium init`](./init) — authenticate an existing account - [`lium balance`](./balance) — check the balance - [`lium ls`](./ls) — find a node to rent --- # `lium init` # `lium init` Run first-time setup: authenticate in the browser, then save your API key and configure your SSH key locally. ```bash lium init [OPTIONS] ``` By default `lium init` opens your browser to authenticate. For headless machines, use the two-step flow with `--no-browser` and `--session`. ## Options | Flag | Effect | |------|--------| | `--no-browser` | Print the auth URL and a session ID instead of opening a browser (step 1 of headless auth) | | `--session ID` | Verify a pending auth session and save the API key (step 2 of headless auth) | ## Examples ```bash lium init # Opens the browser to authenticate lium init --no-browser # Prints the auth URL + session ID lium init --session # Verifies that session and saves the API key ``` ## See also - [Installation](../installation) — install the CLI first - [`lium config`](./config) — read/write individual config values later --- # `lium config` # `lium config` Read and write CLI configuration values. ```bash lium config SUBCOMMAND [OPTIONS] ``` ## Subcommands | Subcommand | Purpose | |------------|---------| | `show` | Display the entire current configuration | | `get KEY` | Read a single value by dotted key | | `set KEY [VALUE]` | Write a single value (interactive if `VALUE` is omitted) | | `unset KEY` | Remove a single value | | `reset` | Reset configuration to defaults (`--confirm` skips the prompt) | | `path` | Print the path to the config file | | `edit` | Open the config file in your editor | ## Examples ```bash lium config show # Print full config lium config get api.api_key # Read a value lium config set api.api_key # Set interactively (value hidden) lium config unset api.api_key # Remove a value lium config path # Where the config file lives lium config reset --confirm # Reset to defaults, no prompt lium config edit # Open in $EDITOR ``` ## Environment variable override Override the API key for a single invocation: ```bash LIUM_API_KEY=xyz lium ls ``` ## See also - [`lium init`](./init) — first-time setup - [`lium theme`](./theme) — change the CLI color theme --- # `lium theme` # `lium theme` Change the CLI color theme. ```bash lium theme THEME_NAME ``` ## Arguments - `THEME_NAME` — Required. The theme to apply: `dark` or `light`. ## Examples ```bash lium theme dark # Apply the dark theme lium theme light # Apply the light theme ``` ## See also - [`lium config`](./config) — view / edit the persisted theme value --- # `lium balance` # `lium balance` Show your current Lium account balance. ```bash lium balance [OPTIONS] ``` ## Options | Flag | Effect | |------|--------| | `--json` | Print machine-readable JSON instead of the formatted view | ## Examples ```bash lium balance # Human-readable balance lium balance --json # Machine-readable output ``` ## See also - [`lium fund`](./fund) — add TAO or alpha funds - [`lium topup`](./topup) — top up with stablecoins --- # `lium fund` # `lium fund` Fund your account with TAO — or Subnet 51 **alpha** — from a Bittensor wallet. ```bash lium fund [OPTIONS] ``` Run without flags for an interactive prompt that walks you through wallet, hotkey, and amount. :::warning Alpha funding moves only your *free* alpha With `--alpha`, the transfer uses your **free** alpha only — your stake minus any locked amount. If you ask to move more alpha than you have free, the transfer is **not attributed and your balance is not credited at all**. Keep the amount at or below your free alpha. ::: ## Options | Flag | Effect | |------|--------| | `--wallet, -w NAME` | Wallet name | | `--hotkey, -k NAME` | Hotkey the funds (or alpha) come from. Required with `--alpha`. | | `--amount, -a AMOUNT` | Amount to fund. TAO by default; **USD** when `--alpha` is set.[^usd] | | `--alpha` | Fund with free Subnet 51 alpha stake (via `transfer_stake`) instead of TAO. | | `--json` | Print machine-readable JSON (non-interactive; pass `--yes`). | | `--yes, -y` | Skip confirmation | [^usd]: Under `--alpha` the amount is denominated in USD. Lium quotes the alpha to move and the destination subnet at transfer time, then credits the actual alpha moved revalued at on-chain inclusion time, so the credited USD can differ slightly from the quote. ## Examples ```bash lium fund # Interactive (TAO) lium fund -w default -a 10.0 # Fund 10 TAO lium fund -w mywallet -k hotkey1 -a 5.0 -y # No prompts (TAO) # Alpha funding — -a is USD, -k (origin hotkey) is required lium fund --alpha -k -a 25 # Fund ~$25 of free alpha lium fund --alpha -w default -k myhotkey -a 25 # -k may be a wallet hotkey name lium fund --alpha -k -a 25 -y --json # Non-interactive, JSON output ``` ## See also - [`lium ps`](./ps) — see hourly burn rate after funding --- # `lium topup` # `lium topup` Top up your Lium balance with a **stablecoin** (e.g. USDT). Unlike [`lium fund`](./fund.md) — which transfers TAO from a Bittensor wallet — `topup` creates a payment invoice and hands you a **deposit address** plus the exact crypto amount to send. You make the transfer from your own wallet; the balance is credited automatically once the network confirms it. This is the recommended path for **AI agents that self-fund**: every subcommand supports `--json`, so an agent can create an invoice, read the deposit address, send the funds, then poll `lium balance` until the credit lands — no human in the loop. ```bash lium topup COMMAND [OPTIONS] ``` ## Subcommands | Command | Purpose | |---------|---------| | `lium topup currencies` | List supported stablecoins and networks | | `lium topup create` | Create an invoice and print the deposit address + amount | ## `lium topup currencies` List the stablecoins and networks you can pay with. ```bash lium topup currencies [--refresh] [--json] ``` | Flag | Effect | |------|--------| | `--refresh` | Bypass the cache and re-fetch from the provider | | `--json` | Print machine-readable JSON | ```bash lium topup currencies lium topup currencies --json ``` ## `lium topup create` Create a top-up invoice and print the deposit address plus the exact amount to send. ```bash lium topup create -a AMOUNT -c CURRENCY -n NETWORK [--json] ``` | Flag | Effect | |------|--------| | `--amount, -a FLOAT` | Top-up amount in **USD** (required) | | `--currency, -c CODE` | Stablecoin code, e.g. `USDT` (required) | | `--network, -n NAME` | Network the coin is sent on, e.g. `tron` (required) | | `--json` | Print machine-readable JSON | ```bash lium topup create -a 20 -c USDT -n tron lium topup create -a 20 -c USDT -n tron --json ``` The response includes everything needed to pay: | Field | Meaning | |-------|---------| | `invoice_id` | Provider invoice identifier | | `deposit_address` | **Send the funds here** | | `crypto_amount` | Exact amount of `crypto_currency` to send | | `crypto_currency` / `crypto_network` | Coin and network to send on | | `fiat_amount` / `fiat_currency` | USD value being credited | | `expires_at` | Invoice expiry — send before this (invoices live ~30 min) | | `hosted_invoice_url` | Browser-payable page for the same invoice | :::warning Send the exact amount, on the right network, before expiry Underpaying credits only what is received. Sending on the wrong network, or after `expires_at`, can lose the funds. Always read `crypto_amount` and `crypto_network` back from the invoice rather than assuming them. ::: ## Agent self-funding loop A shell agent with wallet access can top itself up end-to-end: ```bash # 1. Discover a supported coin/network lium topup currencies --json # 2. Create a $20 invoice and capture the deposit address + amount INVOICE=$(lium topup create -a 20 -c USDT -n tron --json) ADDRESS=$(echo "$INVOICE" | jq -r .deposit_address) AMOUNT=$(echo "$INVOICE" | jq -r .crypto_amount) # 3. Send exactly $AMOUNT USDT (TRON) to $ADDRESS from your own wallet. # This transfer happens in your wallet — Lium does not move your funds. # 4. Poll until the balance is credited lium balance --json ``` The transfer in step 3 is yours to make — Lium only issues the invoice and credits the balance once the provider confirms the on-chain payment. ## See also - [`lium fund`](./fund.md) — fund with TAO from a Bittensor wallet instead - [`lium balance`](./index.md) — check the credited balance - [AI Agents](../../agents.md) — full agent integration guide --- # `lium ssh-keys` # `lium ssh-keys` Manage the SSH keys registered on your Lium account. ```bash lium ssh-keys SUBCOMMAND ``` Running `lium ssh-keys` with no subcommand defaults to `list`. ## Subcommands ### `lium ssh-keys list` List the SSH keys currently registered to your account. ```bash lium ssh-keys list ``` ### `lium ssh-keys sync` Sync SSH keys between your local machine and the Lium account, registering any local public keys that aren't already on the account. ```bash lium ssh-keys sync ``` ## See also - [`lium init`](./init.md) — register your first SSH key during setup - [`lium up --ssh-name`](./up.md) — name a fresh key registered on pod creation --- # `lium mine` # `lium mine` Bootstrap a Lium provider node end-to-end: install prereqs, clone `compute-subnet`, set up the executor environment, and bring it online. ```bash lium mine [OPTIONS] ``` This is the same workflow that `https://lium.io/mine.sh` uses under the hood. Run it directly if you've already installed the CLI. ## Options | Flag | Effect | |------|--------| | `--hotkey, -k SS58` | Miner hotkey SS58 address | | `--dir, -d PATH` | Target directory (default: `compute-subnet`) | | `--branch, -b NAME` | Branch to clone (default: `main`) | | `--auto, -a` | Non-interactive — auto-detect GPU/IP and use defaults | | `--verbose, -v` | Show the plan banner before executing | ## Examples ```bash # One-shot installer (this is what mine.sh runs) curl -fsSL https://lium.io/mine.sh | bash -s -- -k # Direct CLI invocation lium mine -k lium mine -k --auto # No prompts lium mine -k -b staging # Use the staging branch lium mine -k -d /opt/compute-subnet # Custom install path ``` ## See also - [Provider node quickstart](/providers/nodes/quickstart) — full setup guide - [`lium gpu-splitting`](./gpu-splitting.md) — enable multi-tenant rental on the same node - [`lium provider`](./provider.md) — provider portal authentication and status --- # `lium gpu-splitting` # `lium gpu-splitting` Prepare Docker storage on a provider node so the host can be split across multiple renters. This is a **provider-side** command, run on the host. ```bash lium gpu-splitting SUBCOMMAND [OPTIONS] ``` GPU splitting requires an XFS-backed Docker storage driver on a non-root disk. These subcommands inspect the host, verify it meets the requirements, and (with `setup`) reconfigure Docker on a target device. ## Subcommands ### `lium gpu-splitting check` Inspect the host and print the plan **without making any changes**. ```bash lium gpu-splitting check [--device /dev/] ``` | Flag | Effect | |------|--------| | `--device PATH` | Inspect a specific block device instead of auto-selecting | Outputs the current Docker state, candidate target devices, and what `setup` would do. ### `lium gpu-splitting verify` Verify the host already matches the GPU-splitting requirements. ```bash lium gpu-splitting verify ``` Exits non-zero if any requirement is not satisfied. Use this after `setup` to confirm. ### `lium gpu-splitting setup` Perform end-to-end Docker storage setup for GPU splitting. This is the only destructive subcommand. ```bash sudo lium gpu-splitting setup [--device /dev/] [--yes] ``` | Flag | Effect | |------|--------| | `--device PATH` | Target a specific block device (else auto-selected) | | `--yes` | Skip the interactive confirmation | `setup` rejects the root disk and the disk currently backing `/var/lib/docker` — it requires a separate, non-root device. There is no supported single-disk path; attach a second disk or reinstall onto an XFS root first. ## Examples ```bash # Inspect before doing anything lium gpu-splitting check lium gpu-splitting check --device /dev/nvme1n1 # Apply on a specific device sudo lium gpu-splitting setup --device /dev/nvme1n1 sudo lium gpu-splitting setup --device /dev/nvme1n1 --yes # Confirm requirements after setup lium gpu-splitting verify ``` ## See also - [GPU splitting](/providers/nodes/gpu-splitting) — concept guide - [Docker storage](/providers/nodes/docker-storage) — XFS / device requirements - [`lium mine`](./mine.md) — bring the node online first --- # `lium provider` # `lium provider` Provider-side CLI for Bittensor Subnet 51. The `lium provider` namespace covers everything the provider portal frontend at [lium.io/portal](https://lium.io/portal) does — portal authentication, node lifecycle, central-miner-server configuration, batch sync, and read-only billing / machine-request queries — so an AI agent (or any automation) can run a Lium provider account end-to-end without a browser. Distinct from [`lium mine`](./mine.md), which is the on-host bootstrap: it clones the miner repo, installs the executor tooling and starts it on the GPU machine itself. Both are provider-side — the CLI's own `lium provider --help` blurb calls `lium mine` a renter workflow, which is wrong. Hotkey registration on SN51 itself is performed directly with `btcli subnet register`; everything after that is `lium provider …`. ```bash lium provider [OPTIONS] SUBCOMMAND [SUBCOMMAND-OPTIONS] ``` ## Group-level options Set once on the `lium provider` group; every subcommand inherits them. CLI flags beat env vars beat `~/.lium/config.ini`. | Flag | Env var | Config key | Effect | |------|---------|------------|--------| | `-w, --coldkey NAME` | `LIUM_PROVIDER_COLDKEY` | `provider.coldkey` | Bittensor coldkey (wallet) name | | `-k, --hotkey NAME` | `LIUM_PROVIDER_HOTKEY` | `provider.hotkey` | Bittensor hotkey on the coldkey | | `--portal-url URL` | `LIUM_PORTAL_URL` | `provider.portal_url` | Override portal base URL (default: production) | | `--json` | — | — | Emit one machine-readable JSON envelope per command | | `--debug` | — | — | Verbose logging; include error context on stderr | | `-y, --yes` | `LIUM_PROVIDER_ACK=1` | — | Auto-confirm the persona gate for spend-affecting subcommands | | `--dry-run` | — | — | Skip irreversible subprocess calls (e.g. `ssh`) and report intent only | Persist hotkey/coldkey across sessions instead of repeating flags: ```bash lium config set provider.coldkey miner-prod lium config set provider.hotkey miner-1 ``` After that, every `lium provider …` invocation auto-resolves the wallet identity from `~/.lium/config.ini`. ## Subcommand map | Group | Purpose | Spend-affecting? | |-------|---------|------------------| | [`portal`](#lium-provider-portal) | JWT session against the provider portal | no | | [`status`](#lium-provider-status) | Aggregated provider snapshot | no | | [`node`](#lium-provider-node) | GPU node lifecycle on the portal | most subcommands yes | | [`config`](#lium-provider-config) | Portal-account, Discord, password & central-miner-server settings | some mutating subcommands | | [`sync`](#lium-provider-sync) | Batch sync between portal and central miner server | yes | | [`billing`](#lium-provider-billing) | Paginated billing history | no | | [`machine-request`](#lium-provider-machine-request) | Pending tenant machine requests | no | | [`machine`](#lium-provider-machine) | GPU machine catalogue + reward estimates | no | Spend-affecting subcommands run a persona gate that prompts unless `--yes` is set or `LIUM_PROVIDER_ACK=1` is exported. ## `lium provider portal` Manage the cached JWT used to authenticate against the provider portal. The token is keyed by hotkey ss58 and stored under `~/.lium/provider/`. ```bash lium provider portal login [--force] # Exchange a hotkey signature for a JWT lium provider portal logout # Drop the cached JWT for this hotkey lium provider portal whoami # Call /auth/me with the cached token ``` | Flag (on `login`) | Effect | |-------------------|--------| | `--force` | Bypass the local token cache and re-authenticate | `login` requires `--hotkey` (or `LIUM_PROVIDER_HOTKEY`, or `provider.hotkey` in config). `logout` and `whoami` do too. In JSON mode, `login` includes Discord connection and extra incentive eligibility when the portal profile is available: ```json { "discord_connected": false, "extra_incentive_eligible": false, "hotkey": "5...", "provider_id": "...", "token_present": true } ``` In human mode, if Discord is not connected, `login` also shows the next command to run: ```bash lium provider config connect-discord ``` ## `lium provider status` One-shot health snapshot composed from the subtensor metagraph, portal `/auth/me`, the node list, extra incentive eligibility, and validator weights. Sources that fail are skipped and surfaced as `warnings` in the JSON envelope. ```bash lium provider status [--netuid 51] ``` | Flag | Effect | |------|--------| | `--netuid INT` | Subnet to query (default: `51`) | When the portal profile is available, `status` includes `discord_connected` and `extra_incentive_eligible`. If Discord is not connected, human output shows the same extra incentive next step as `portal login`. ## `lium provider node` Node lifecycle on the portal. Mutating subcommands (`add`, `rm`, `update-*`, `min-gpu set/unset`, `notice-period set/unset`, `notify-added`) run the persona gate. ```bash lium provider node list [--miner-hotkey HK] [--page N] [--limit N] lium provider node get lium provider node add --gpu-type TYPE --ip IP [--port 8080] [--price USD] --gpu-count N lium provider node rm lium provider node update-price --price USD lium provider node update-gpu --gpu-type TYPE --gpu-count N lium provider node min-gpu set lium provider node min-gpu unset lium provider node pods lium provider node machine-requests lium provider node notice-period set lium provider node notice-period unset lium provider node notify-added --request-id REQ ``` | Subcommand | Args / flags | Purpose | |------------|--------------|---------| | `list` | `--miner-hotkey HK`, `--page N`, `--limit N` | Paginated node listing for the active provider | | `get` | `NODE_ID` | Fetch a single node record | | `add` | `--gpu-type`, `--ip`, `--port` (default `8080`), `--price` (optional; auto-filled from public [shared-config](#default-prices--shared-config) when omitted), `--gpu-count` (default `1`) | Queue a new node addition (`POST /executors`). Persona-gated | | `rm` | `NODE_ID` | Delete the node. Persona-gated | | `update-price` | `NODE_ID --price USD` | Set price-per-GPU/hour. Persona-gated | | `update-gpu` | `NODE_ID --gpu-type TYPE --gpu-count N` | Change GPU type/count. Persona-gated | | `min-gpu set` | `NODE_ID COUNT` | Minimum GPU count required for rental matchmaking. Persona-gated | | `min-gpu unset` | `NODE_ID` | Clear the minimum-GPU rule. Persona-gated | | `pods` | `NODE_ID` | List pods currently rented on this node | | `machine-requests` | `NODE_ID` | List pending tenant requests targeting this node | | `notice-period set` | `NODE_ID` | Open a notice period before maintenance / decommission. Persona-gated | | `notice-period unset` | `NODE_ID` | Cancel an open notice period. Persona-gated | | `notify-added` | `NODE_ID --request-id REQ` | Mark a tenant machine request fulfilled (`POST /machine-added`). Persona-gated | The portal still exposes these routes under `/executors/...`; the CLI verb is `node` because that is the user-facing terminology after the 2026 Provider rename. ### Default prices & shared-config `lium provider node add` accepts `--price` as **optional**. When omitted, the CLI fetches the public, unauthenticated `GET /v1/shared-config` snapshot from `https://lium.io/api` (the same data the provider portal frontend uses to populate its Add-Node modal) and looks up the base USD/GPU/hour for `--gpu-type`. The chosen price is printed to stderr so you can confirm before the call hits the portal. If the GPU type isn't in the public price table, the CLI exits `1` with `ARG_INVALID` and lists known types — pass `--price` explicitly or correct the spelling. To browse all known GPU types interactively, use [`lium provider machine list`](#lium-provider-machine). Override the snapshot URL with the `LIUM_SHARED_CONFIG_URL` environment variable (useful when pointing at staging or a self-hosted backend). ## `lium provider config` Portal-*account* state held server-side — distinct from `lium config` (CLI-side `~/.lium/config.ini`) and `lium provider portal` (JWT/session). Spend-affecting and existing account-mutating subcommands note when they are persona-gated. ```bash lium provider config show lium provider config opt-in # use lium.io's central miner server lium provider config opt-out # run your own central miner server lium provider config set-email lium provider config set-password [--password PASSWORD] lium provider config connect-discord [--no-wait] [--timeout N] [--poll-interval S] lium provider config set-subscriptions [--gpu TYPE]... # repeat --gpu, or pass none to clear ``` | Subcommand | Effect | |------------|--------| | `show` | Full `GET /auth/me` profile, including `discord_connected` and `extra_incentive_eligible` | | `opt-in` / `opt-out` | Toggle the lium.io central miner server. Persona-gated | | `set-email` | Update the contact email. Persona-gated | | `set-password` | Set the Provider Portal password using fresh hotkey signature authentication. Prompts when `--password` is omitted outside JSON mode; automation can use `--password` or `LIUM_PROVIDER_NEW_PASSWORD` | | `connect-discord` | Start Discord OAuth linking. Best-effort opens a browser; always prints/returns the authorization URL | | `set-subscriptions` | Set machine-request notification subscriptions by GPU type (repeat `--gpu`); pass none to clear. Persona-gated | By default, `connect-discord` opens the Discord authorization URL when a browser is available, then waits up to 120 seconds for the portal profile to show Discord as connected. Use `--no-wait` to print/return the URL immediately. `connect-discord --json --no-wait` is the recommended agent/headless form. It returns: ```json { "authorization_url": "https://discord.com/oauth2/authorize?...", "browser_opened": false, "discord_connected": false, "extra_incentive_eligible": false, "next_action": "open_authorization_url_and_complete_discord_oauth" } ``` Open `authorization_url` as the provider's Discord account owner, then rerun `lium provider status --json` to confirm `discord_connected` and `extra_incentive_eligible` are true. Without Discord connected, extra subnet incentives are disabled. ## `lium provider sync` Batch sync between the portal and the central miner server — the same buttons the portal frontend exposes as "Sync From Miner Server" and "Sync Into Miner Server". Both subcommands are persona-gated. ```bash lium provider sync from-miner-server # pull node state from the central miner server lium provider sync to-miner-server # push node state to the central miner server ``` ## `lium provider billing` ```bash lium provider billing list [--miner-hotkey HK] [--page N] [--limit N] ``` Paginated billing-history entries for the active provider, filtered by hotkey if provided. ## `lium provider machine-request` ```bash lium provider machine-request list # all pending tenant requests lium provider machine-request get # single tenant machine request ``` ## `lium provider machine` ```bash lium provider machine list lium provider machine estimate --gpu-type TYPE --gpu-count N [--gpu-price USD] ``` | Subcommand | Args | Purpose | |------------|------|---------| | `list` | — | Available GPU machine catalogue | | `estimate` | `--gpu-type`, `--gpu-count`, optional `--gpu-price` | Estimated rewards for a given GPU configuration | ## End-to-end examples ```bash # One-time setup: persist wallet identity so flags aren't needed on every call lium config set provider.coldkey miner-prod lium config set provider.hotkey miner-1 # Authenticate against the provider portal lium provider portal login lium provider portal whoami # Connect Discord for extra incentive eligibility lium provider config connect-discord # Browserless/headless form for agents lium provider config connect-discord --json --no-wait # Health snapshot (registration, portal session, node count, validator weights) lium provider status lium provider --json status # JSON envelope for scripts / agents # Provision: opt in to the lium.io central miner server, then add a node lium provider config opt-in --yes # Explicit price (override): lium provider node add \ --gpu-type "NVIDIA H200 NVL" \ --gpu-count 8 \ --ip 203.0.113.42 \ --port 8080 \ --price 1.85 \ --yes # Or omit --price to auto-fill from the public shared-config baseline: lium provider node add \ --gpu-type "NVIDIA H200 NVL" \ --gpu-count 8 \ --ip 203.0.113.42 \ --yes # Inspect & manage the node lium provider node list --limit 50 lium provider node get lium provider node pods lium provider node update-price --price 2.10 --yes # Reward forecasting and tenant requests lium provider machine estimate --gpu-type "NVIDIA H200 NVL" --gpu-count 8 lium provider machine-request list lium provider node machine-requests lium provider node notify-added --request-id --yes # Schedule maintenance, then close out lium provider node notice-period set --yes # ... maintenance happens ... lium provider node notice-period unset --yes # Sync state between portal and central miner server lium provider sync from-miner-server --yes lium provider sync to-miner-server --yes # Billing & history lium provider billing list --page 1 --limit 50 # Drop the token when rotating hotkeys lium provider portal logout ``` Add `--json` to any command for a machine-readable envelope (scripted automation, CI runs, AI-agent loops). Pair with `--yes` (or `LIUM_PROVIDER_ACK=1`) to skip the persona prompt on spend-affecting calls. ## Exit codes & errors `lium provider …` does **not** use the exit codes in the [CLI reference index](./index.md#exit-codes) — it has its own map, and the same number means something different here: | Exit code | Meaning here | Error codes | |-----------|--------------|-------------| | `0` | Success | — | | `1` | User / argument error, and the fallback for anything unmapped | `ARG_INVALID`, `PORTS_INVALID`, `HOTKEY_NOT_REGISTERED` | | `2` | Authentication error | `PORTAL_AUTH_INVALID`, `PORTAL_AUTH_EXPIRED`, `WALLET_NOT_FOUND` | | `3` | Portal error (server-side, not auth) | `PORTAL_SERVER_ERROR`, `PORTAL_NOT_FOUND`, `PORTAL_CONTRACT_DRIFT`, `PORTAL_RATE_LIMIT` | | `5` | SSH error | `SSH_UNREACHABLE`, `SSH_AUTH_FAILED` | | `6` | Config error | `CONFIG_MISSING` | | `7` | Token-cache contention — another `lium provider` process is refreshing the JWT; retry | `PORTAL_AUTH_REFRESH_RACE` | The error code is the stable thing to branch on. Under `--json` it arrives as `{"ok": false, "error": {"code": "...", "message": "...", "hint": "...", "context": {}}}`: | Error code | Meaning | |-----------|---------| | `ARG_INVALID` | Required flag or argument missing (e.g. `--hotkey`) | | `PORTAL_AUTH_INVALID` | JWT rejected — run `lium provider portal login` | | `PORTAL_AUTH_EXPIRED` | JWT expired — run `lium provider portal login` | | `PORTAL_CONTRACT_DRIFT` | Portal returned an unexpected envelope shape (422) | | `PORTAL_SERVER_ERROR` | Portal 5xx — retry with `--debug` for context | | `PORTAL_NOT_FOUND` | Portal returned 404 — wrong UUID, or already removed | | `PORTAL_RATE_LIMIT` | Backing off; retry shortly | | `WALLET_NOT_FOUND` | The named coldkey/hotkey does not exist locally | | `HOTKEY_NOT_REGISTERED` | Register on SN51 first with `btcli subnet register` | | `SSH_UNREACHABLE` / `SSH_AUTH_FAILED` | Host unreachable, or key/user rejected | | `INSTALLER_PARTIAL_FAIL` | `mine.sh` did not finish — see `/tmp/lium-mine.log` on the host | | `CONFIG_MISSING` | A required config value is unset | ## See also - [`lium mine`](./mine.md) — install and start the node binary on the host - [`lium gpu-splitting`](./gpu-splitting.md) — prepare Docker storage for multi-tenant rental - [Provider Portal — Overview](/providers/portal/overview) - [Provider Portal — Managing Nodes](/providers/portal/managing-nodes) - [Provider Portal — Machine Requests](/providers/portal/machine-requests) - [AI Agents integration guide](../../agents.md) --- # MCP Endpoint # MCP Endpoint The Lium docs site exposes a [Model Context Protocol](https://modelcontextprotocol.io) endpoint at `/mcp`. AI agents and MCP-compatible clients can use this to search the documentation and read individual pages. :::info For agents This endpoint is for the **docs**. To call the **platform** API directly (create pods, manage volumes, …), use the live [OpenAPI spec](./openapi) at `https://lium.io/api/openapi.json`. The [AI Agents](./agents) page shows when to reach for which surface. ::: ## Endpoint ``` POST https://docs.lium.io/mcp ``` The endpoint uses the MCP HTTP transport (JSON-RPC 2.0 over HTTP POST). ## Available tools ### `search` Full-text search across all documentation pages, with optional audience filtering. **Parameters:** | Parameter | Type | Required | Description | |-----------|------|----------|-------------| | `query` | string | yes | Search query | | `audience` | string | no | Filter to one audience: `providers`, `validators`, `pod-users`, `developers` | | `limit` | number | no | Max results (default: 10) | **Example:** ```json { "jsonrpc": "2.0", "id": 1, "method": "tools/call", "params": { "name": "search", "arguments": { "query": "register a node", "audience": "providers", "limit": 5 } } } ``` ### `read_page` Read a page's source markdown by URL or slug. Resolution order: full URL > absolute slug (`/providers/quickstart`) > basename (`quickstart`), with tie-breaking by audience nav order. **Parameters:** | Parameter | Type | Required | Description | |-----------|------|----------|-------------| | `url_or_slug` | string | yes | Full URL, absolute slug, or page basename | **Example:** ```json { "jsonrpc": "2.0", "id": 2, "method": "tools/call", "params": { "name": "read_page", "arguments": { "url_or_slug": "/providers/quickstart" } } } ``` ## Using with Claude To connect Claude to the Lium docs MCP server, add it to your MCP configuration: ```json { "mcpServers": { "lium-docs": { "url": "https://docs.lium.io/mcp" } } } ``` ## `.md` URLs Every page on the docs site is also available as a raw markdown URL by appending `.md`: ```bash curl https://docs.lium.io/providers/quickstart.md ``` You can also use `Accept: text/markdown` on any page URL: ```bash curl -H 'Accept: text/markdown' https://docs.lium.io/providers/quickstart ``` --- # llms.txt # llms.txt The Lium docs site publishes `llms.txt` and `llms-full.txt` at the site root, conforming to the [llmstxt.org specification](https://llmstxt.org). ## Endpoints | URL | Description | |-----|-------------| | `https://docs.lium.io/llms.txt` | Index: page titles + 1-line summaries + `.md` URLs | | `https://docs.lium.io/llms-full.txt` | Full bundle: all pages concatenated with H1 separators | ## Format `llms.txt` is an H1 title, a brief description, and then an audience-grouped link list: ``` # Lium Documentation > Decentralized GPU rental marketplace on Bittensor Subnet 51. ## Providers - [Provider Quickstart](/providers/quickstart.md): Register and launch your first provider node in 5 minutes. - [Architecture](/providers/architecture.md): Self-hosted provider + node architecture. ... ``` `llms-full.txt` concatenates every page's source markdown, separated by H1 headings: ``` # Provider Quickstart --- # Architecture ... ``` If `llms-full.txt` exceeds 5 MB, it is split into `llms-full.part-1.txt`, `llms-full.part-2.txt`, etc., referenced from the main `llms-full.txt`. ## Usage with LLMs Pass the full bundle to an LLM as context to enable zero-shot Q&A over the entire Lium documentation: ```python import anthropic import httpx llms_full = httpx.get("https://docs.lium.io/llms-full.txt").text client = anthropic.Anthropic() response = client.messages.create( model="claude-sonnet-4-5", max_tokens=1024, messages=[{ "role": "user", "content": f"\n{llms_full}\n\n\nHow do I register a provider node?" }] ) print(response.content[0].text) ``` For programmatic access with tool use, prefer the [MCP endpoint](./mcp). To call the platform API itself (not the docs), use the live [OpenAPI spec](./openapi). --- # AlphaQuote # AlphaQuote ```python from lium.sdk import AlphaQuote ``` Defined in `lium.sdk.client`. USD -> alpha quote from ``GET /balance/convert/alpha``. ``netuid`` is the subnet the alpha must be transferred on (the same subnet the pay-tao-api-v2 listener credits), so it — not a hardcoded constant — drives the on-chain ``transfer_stake``. ```python AlphaQuote( usd: decimal.Decimal, alpha_amount: decimal.Decimal, rate: decimal.Decimal, netuid: int ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `usd` | decimal.Decimal | | | `alpha_amount` | decimal.Decimal | | | `rate` | decimal.Decimal | | | `netuid` | int | | --- # Lium # Lium ```python from lium.sdk import Lium ``` Defined in `lium.sdk.client`. Clean Unix-style SDK for Lium. ```python Lium(config: Optional[lium.sdk.config.Config] = None, source: str = 'sdk') ``` ## Attributes | Name | Type | Description | | --- | --- | --- | | `config` | | | | `headers` | | | ## Methods | Method | Description | | --- | --- | | [`list_ssh_keys`](#list_ssh_keys) | Return SSH keys registered for the current user. | | [`register_ssh_key`](#register_ssh_key) | Register a new SSH public key under the current user. | | [`default_ssh_key_name`](#default_ssh_key_name) | cli-<user>@<host> sanitised to A-Za-z0-9.@-. | | [`up`](#up) | Start a new pod on a specific node. | | [`pod`](#pod) | Retrieve detailed information about a specific pod. | | [`logs`](#logs) | Stream logs from a pod. | | [`edit`](#edit) | Edit a pod's template configuration. | | [`ls`](#ls) | List available nodes. | | [`ps`](#ps) | List active pods. | | [`down`](#down) | Stop a pod. | | [`rm`](#rm) | Remove pod (alias for down()). | | [`reboot`](#reboot) | Reboot a pod. | | [`get_default_images`](#get_default_images) | Get default images for GPU type and driver version. | | [`default_docker_template`](#default_docker_template) | Resolve the best default template for a node ID. | | [`templates`](#templates) | List available templates. | | [`get_executor`](#get_executor) | Resolve a node by ID. | | [`gpu_types`](#gpu_types) | Get list of available GPU types. | | [`get_template`](#get_template) | Fetch a template by ID/HUID/name. | | [`get_template_by_image_name`](#get_template_by_image_name) | Fetch a template by its Docker image + tag. | | [`ssh_connection`](#ssh_connection) | SSH connection context manager. | | [`exec`](#exec) | Execute a shell command on a pod over SSH. | | [`stream_exec`](#stream_exec) | Execute a shell command and stream incremental output. | | [`exec_all`](#exec_all) | Execute a shell command on multiple pods in parallel. | | [`wait_ready`](#wait_ready) | Poll until a pod reports RUNNING + SSH metadata. | | [`scp`](#scp) | Upload a local file to a pod via SFTP. | | [`download`](#download) | Download a file from a pod via SFTP. | | [`upload`](#upload) | Upload a file to a pod via SFTP. | | [`ssh`](#ssh) | Get SSH command string for connecting to a pod. | | [`rsync`](#rsync) | Sync directories with rsync. | | [`switch_template`](#switch_template) | Switch the template of a running pod. | | [`create_template`](#create_template) | Create a new template. | | [`wait_template_ready`](#wait_template_ready) | Wait for template verification to complete. | | [`get_my_user_id`](#get_my_user_id) | Get the current user's ID. | | [`update_template`](#update_template) | Update an existing template owned by the caller. | | [`wallets`](#wallets) | Get the caller's configured funding wallets. | | [`add_wallet`](#add_wallet) | Link a Bittensor wallet with the user account. | | [`convert_alpha`](#convert_alpha) | Quote usd (USD) -> alpha via GET /balance/convert/alpha. | | [`company_wallet`](#company_wallet) | Resolve the Lium destination coldkey via GET /wallet/company/?appid=. | | [`backup_create`](#backup_create) | Create or replace a backup configuration for a pod. | | [`backup_now`](#backup_now) | Trigger an immediate backup for a pod. | | [`backup_config`](#backup_config) | Return the backup configuration for a pod if one exists. | | [`backup_list`](#backup_list) | List all backup configurations across all pods. | | [`backup_logs`](#backup_logs) | Get recent backup logs for a pod. | | [`backup_delete`](#backup_delete) | Delete a backup configuration by ID. | | [`restore`](#restore) | Restore a backup to a pod. | | [`restore_logs`](#restore_logs) | Get recent restore logs for a pod. | | [`get_deployment_estimate`](#get_deployment_estimate) | Estimate deployment time for a template on a node. | | [`balance`](#balance) | Get current account balance. | | [`topup_currencies`](#topup_currencies) | List stablecoin currencies/networks supported for self-serve top-ups. | | [`topup_create_invoice`](#topup_create_invoice) | Create a stablecoin top-up invoice for the current account. | | [`volumes`](#volumes) | List all volumes for the current user. | | [`volume`](#volume) | Get a specific volume by ID. | | [`volume_create`](#volume_create) | Create a new volume. | | [`volume_update`](#volume_update) | Update a volume's metadata. | | [`volume_delete`](#volume_delete) | Delete a volume. | | [`schedule_termination`](#schedule_termination) | Schedule a pod for automatic termination at a future date and time. | | [`cancel_scheduled_termination`](#cancel_scheduled_termination) | Cancel a scheduled termination for a pod. | | [`install_jupyter`](#install_jupyter) | Install Jupyter Notebook on a pod. | ### list_ssh_keys ```python def list_ssh_keys() -> List[lium.sdk.models.SSHKey]: ``` Return SSH keys registered for the current user. ### register_ssh_key ```python def register_ssh_key(*, name: str, public_key: str) -> lium.sdk.models.SSHKey: ``` Register a new SSH public key under the current user. ### default_ssh_key_name ```python def default_ssh_key_name() -> str: ``` ``cli-@`` sanitised to ``[A-Za-z0-9._@-]``. ### up ```python def up( *, executor_id: str, name: str = 'Your Pod', template_id: Optional[str] = None, dockerfile_content: Optional[str] = None, volume_id: Optional[str] = None, ports: Optional[int] = None, ssh_keys: Optional[List[str]] = None, ssh_name: Optional[str] = None, enable_volume_encryption: bool | None = True ) -> Dict[str, Any]: ``` Start a new pod on a specific node. **Arguments:** - **executor_id:** Target node ID string. - **name:** Human-friendly pod name (defaults to ``"Your Pod"``). - **template_id:** Template ID. Defaults to the node's default template. Mutually exclusive with ``dockerfile_content``. - **dockerfile_content:** Raw Dockerfile text to build the pod image from on the node (custom build). Mutually exclusive with ``template_id`` — pass exactly one. The image is built remotely with no network access, so the Dockerfile must be self-contained (no ``ADD `` or ``ADD $\{var\}`` directives). - **volume_id:** Optional volume ID to attach on spawn. - **ports:** Number of exposed ports to request. - **ssh_keys:** SSH public keys to authorize. Defaults to the keys discovered by the Config. - **ssh_name:** Optional name to use when registering a new SSH key with the backend. Defaults to ``cli-@``. Only applied to keys that are not already registered server-side. - **enable_volume_encryption:** Whether to request encryption for the local pod volume. Enabled by default. The image must support Lium volume encryption. **Returns:** Pod metadata as returned by the rent API (id, name, status, ssh command, etc.). ### pod ```python def pod(pod_id: str) -> Dict[str, Any]: ``` Retrieve detailed information about a specific pod. **Arguments:** - **pod_id:** The unique identifier of the pod to retrieve. **Returns:** Raw pod data dictionary including template, node, status, and connection info. ### logs ```python def logs( pod_id: str, *, tail: int = 100, follow: bool = False ) -> Generator[bytes, NoneType, NoneType]: ``` Stream logs from a pod. **Arguments:** - **pod_id:** The unique identifier of the pod. - **tail:** Number of lines to retrieve from the end of the logs (default: 100). - **follow:** If True, stream logs continuously (default: False). **Yields:** Log lines as bytes. ### edit ```python def edit(pod_id: str, **kwargs) -> Dict[str, Any]: ``` Edit a pod's template configuration. Updates the template associated with a pod by merging the provided keyword arguments with the existing template settings. **Arguments:** - **pod_id:** The unique identifier of the pod whose template to edit. - `**kwargs`: Template fields to update. Common fields include: - docker_image (str): Docker image repository. - docker_image_tag (str): Docker image tag. - startup_commands (str): Commands to run on container start. - internal_ports (List[int]): Ports to expose. - environment (Dict[str, str]): Environment variables. - volumes (List[str]): Volume mount paths. **Returns:** Updated template data dictionary from the API. **Example:** ```python lium.edit(pod_id, startup_commands="python main.py", environment={"DEBUG": "1"}) ``` ### ls ```python def ls( *, gpu_type: Optional[str] = None, gpu_count: Optional[int] = None, lat: Optional[float] = None, lon: Optional[float] = None, max_distance_miles: Optional[int] = None, min_cuda_version: Optional[float] = None ) -> List[lium.sdk.models.ExecutorInfo]: ``` List available nodes. **Arguments:** - **gpu_type:** Optional GPU filter such as ``"A100"`` or ``"H200"``. - **gpu_count:** Exact GPU count to match (defaults to 8, pass ``None`` to disable). - **lat:** Optional latitude for geospatial filtering. Must be used together with ``lon`` and ``max_distance_miles``. - **lon:** Optional longitude for geospatial filtering. Must be used together with ``lat`` and ``max_distance_miles``. - **max_distance_miles:** Optional radius (in miles) for geospatial filtering. Must be used together with ``lat`` and ``lon``. - **min_cuda_version:** Optional minimum CUDA version to require (e.g. ``12.4``). Nodes whose ``max_cuda_version`` is ``None`` or below this threshold are excluded. NVIDIA drivers are backward compatible, so a node with a higher driver CUDA version satisfies the requirement. **Returns:** A list of [`ExecutorInfo`](/developers/sdk/reference/models/executor-info) objects that satisfy the filters. ### ps ```python def ps() -> List[lium.sdk.models.PodInfo]: ``` List active pods. **Returns:** List of [`PodInfo`](/developers/sdk/reference/models/pod-info) objects representing the caller's running pods. ### down ```python def down(pod: lium.sdk.models.PodInfo) -> Dict[str, Any]: ``` Stop a pod. **Arguments:** - **pod:** Pod to terminate. **Returns:** API response payload from the delete call. ### rm ```python def rm(pod: lium.sdk.models.PodInfo) -> Dict[str, Any]: ``` Remove pod (alias for `down()`). **Arguments:** - **pod:** Pod to terminate. **Returns:** API response payload from the delete call. ### reboot ```python def reboot( pod: lium.sdk.models.PodInfo, volume_id: Optional[str] = None ) -> Dict[str, Any]: ``` Reboot a pod. **Arguments:** - **pod:** Pod to reboot. - **volume_id:** Optional volume ID to attach for the reboot request. **Returns:** Pod data from the API response after issuing the reboot. ### get_default_images ```python def get_default_images(gpu_model: Optional[str], driver_version: Optional[str]) -> list[dict]: ``` Get default images for GPU type and driver version. ### default_docker_template ```python def default_docker_template(executor_id: str) -> lium.sdk.models.Template: ``` Resolve the best default template for a node ID. **Arguments:** - **executor_id:** Node identifier returned by `ls()`. **Returns:** [`Template`](/developers/sdk/reference/models/template) best suited for the node. **Raises:** - **ValueError:** If no matching node or template exists. ### templates ```python def templates( filter: Optional[str] = None, only_my: bool = False ) -> List[lium.sdk.models.Template]: ``` List available templates. **Arguments:** - **filter:** Optional substring to filter by image or name. - **only_my:** When ``True`` return only templates owned by the caller. **Returns:** List of [`Template`](/developers/sdk/reference/models/template). ### get_executor ```python def get_executor(executor: str) -> Optional[lium.sdk.models.ExecutorInfo]: ``` Resolve a node by ID. **Arguments:** - **executor:** Node ID string. **Returns:** Matching [`ExecutorInfo`](/developers/sdk/reference/models/executor-info) or ``None`` if not found. ### gpu_types ```python def gpu_types() -> set[str]: ``` Get list of available GPU types. **Returns:** Set of GPU type strings advertised by the API. ### get_template ```python def get_template(template_id: str) -> Optional[lium.sdk.models.Template]: ``` Fetch a template by ID/HUID/name. **Arguments:** - **template_id:** Template ID, HUID, or name to match. **Returns:** Matching [`Template`](/developers/sdk/reference/models/template) or ``None`` if not found. ### get_template_by_image_name ```python def get_template_by_image_name( image_name: Optional[str] = None, image_tag: Optional[str] = None ) -> Optional[lium.sdk.models.Template]: ``` Fetch a template by its Docker image + tag. **Arguments:** - **image_name:** Repository/image name. - **image_tag:** Tag to match. **Returns:** Matching [`Template`](/developers/sdk/reference/models/template) or ``None`` if not found. ### ssh_connection ```python def ssh_connection(pod: lium.sdk.models.PodInfo, timeout: int = 30): ``` SSH connection context manager. **Arguments:** - **pod:** Pod whose SSH metadata is used. - **timeout:** Connection timeout in seconds. **Yields:** An active ``paramiko.SSHClient``. ### exec ```python def exec( pod: lium.sdk.models.PodInfo, *, command: str, env: Optional[Dict[str, str]] = None ) -> Dict[str, Any]: ``` Execute a shell command on a pod over SSH. **Arguments:** - **pod:** Pod to target. - **command:** Shell command to run remotely. - **env:** Optional environment variables exported before the command runs. **Returns:** Dict containing stdout, stderr, exit_code, and success flag. ### stream_exec ```python def stream_exec( pod: lium.sdk.models.PodInfo, *, command: str, env: Optional[Dict[str, str]] = None ) -> Generator[Dict[str, str], NoneType, NoneType]: ``` Execute a shell command and stream incremental output. **Arguments:** - **pod:** Pod to target. - **command:** Shell command to run remotely. - **env:** Optional environment variables exported before the command runs. **Yields:** Streaming output chunks as ``\{"type": "stdout"|"stderr", "data": str\}``. ### exec_all ```python def exec_all( pods: List[lium.sdk.models.PodInfo], *, command: str, env: Optional[Dict[str, str]] = None, max_workers: int = 10 ) -> List[Dict]: ``` Execute a shell command on multiple pods in parallel. **Arguments:** - **pods:** List of pods to target. - **command:** Shell command to run on each pod. - **env:** Optional environment variables exported before each command. - **max_workers:** Maximum number of SSH workers to spawn. **Returns:** List of result dictionaries mirroring `exec()`. ### wait_ready ```python def wait_ready( pod: Union[str, lium.sdk.models.PodInfo, Dict], *, timeout: int = 300, poll_interval: int = 10 ) -> Optional[lium.sdk.models.PodInfo]: ``` Poll until a pod reports RUNNING + SSH metadata. **Arguments:** - **pod:** Pod identifier, PodInfo, or dict with an ``id`` field. - **timeout:** Maximum number of seconds to wait. - **poll_interval:** Interval between successive ``ps`` calls. **Returns:** PodInfo when the pod is ready, otherwise ``None`` if timeout expires. ### scp ```python def scp(pod: lium.sdk.models.PodInfo, *, local: str, remote: str) -> None: ``` Upload a local file to a pod via SFTP. ### download ```python def download(pod: lium.sdk.models.PodInfo, *, remote: str, local: str) -> None: ``` Download a file from a pod via SFTP. **Arguments:** - **pod:** The pod to download from. - **remote:** Remote file path on the pod. - **local:** Local destination path. **Raises:** - **ValueError:** If SSH is not configured for the pod. ### upload ```python def upload(pod: lium.sdk.models.PodInfo, *, local: str, remote: str) -> None: ``` Upload a file to a pod via SFTP. This is an alias for `scp()` for parity with the CLI. **Arguments:** - **pod:** The pod to upload to. - **local:** Local file path to upload. - **remote:** Remote destination path on the pod. **Raises:** - **ValueError:** If SSH is not configured for the pod. ### ssh ```python def ssh(pod: lium.sdk.models.PodInfo) -> str: ``` Get SSH command string for connecting to a pod. **Arguments:** - **pod:** The pod to generate SSH command for. **Returns:** SSH command string with the configured SSH key path. **Raises:** - **ValueError:** If SSH is not configured for the pod or no SSH key path is set. ### rsync ```python def rsync(pod: lium.sdk.models.PodInfo, *, local: str, remote: str) -> None: ``` Sync directories with rsync. **Arguments:** - **pod:** Pod to sync. - **local:** Local path or directory (rsync source). - **remote:** Remote path on the pod. **Raises:** - **RuntimeError:** If the rsync command fails. ### switch_template ```python def switch_template( pod: lium.sdk.models.PodInfo, *, template_id: str ) -> lium.sdk.models.PodInfo: ``` Switch the template of a running pod. **Arguments:** - **pod:** Pod to update. - **template_id:** ID of the template to switch to. **Returns:** PodInfo object with updated pod information. ### create_template ```python def create_template( name: str, docker_image: str, docker_image_digest: str = '', docker_image_tag: str = 'latest', ports: Optional[List[int]] = None, start_command: Optional[str] = None, **kwargs ) -> lium.sdk.models.Template: ``` Create a new template. **Arguments:** - **name:** Friendly template name. - **docker_image:** Image repository (e.g., ``"daturaai/pytorch"``). - **docker_image_digest:** Digest string for pinning (defaults to empty string). - **docker_image_tag:** Image tag (defaults to ``"latest"``). - **ports:** Internal ports to expose (defaults to ``[22, 8000]``). - **start_command:** Optional command executed on container start. - `**kwargs`: Additional template fields: - category (str): Template category (defaults to ``"UBUNTU"``). - is_private (bool): Whether template is private (defaults to ``True``). - volumes (List[str]): Volume mount paths (defaults to ``["/workspace"]``). - description (str): Template description. - environment (Dict[str, str]): Environment variables. - entrypoint (str): Container entrypoint. - one_time_template (bool): Whether to delete template after pod removal (defaults to ``False``). **Returns:** Newly created [`Template`](/developers/sdk/reference/models/template). ### wait_template_ready ```python def wait_template_ready( template_id: str, timeout: int = 300 ) -> Optional[lium.sdk.models.Template]: ``` Wait for template verification to complete. **Arguments:** - **template_id:** Template identifier. - **timeout:** Maximum seconds to wait. **Returns:** Template when verification succeeds, otherwise ``None`` if the timeout expires. **Raises:** - **LiumError:** If template verification fails. ### get_my_user_id ```python def get_my_user_id() -> str: ``` Get the current user's ID. **Returns:** The ID returned by ``/users/me``. ### update_template ```python def update_template( template_id: str, name: str, docker_image: str, docker_image_digest: str, docker_image_tag: str = 'latest', ports: Optional[List[int]] = None, start_command: Optional[str] = None, **kwargs ) -> lium.sdk.models.Template: ``` Update an existing template owned by the caller. **Arguments:** - **template_id:** Template identifier. - **name:** Friendly name. - **docker_image:** Image repository. - **docker_image_digest:** Optional digest. - **docker_image_tag:** Image tag. - **ports:** Internal ports to expose. - **start_command:** Startup command. - `**kwargs`: Additional override fields. **Returns:** Updated [`Template`](/developers/sdk/reference/models/template). **Raises:** - **ValueError:** If the template is missing or not owned by the caller. ### wallets ```python def wallets() -> List[Dict[str, Any]]: ``` Get the caller's configured funding wallets. **Returns:** Raw wallet records returned by the pay API. ### add_wallet ```python def add_wallet(bt_wallet: Any) -> tuple[str, str]: ``` Link a Bittensor wallet with the user account. **Arguments:** - **bt_wallet:** Wallet object exposing ``coldkey``/``coldkeypub`` for signing. **Returns:** ``(app_id, customer_id)`` parsed from the ``/tao/create-transfer`` redirect — surfaced so the alpha funding flow can reuse the same single round-trip for company-wallet lookup instead of issuing a second POST. **Raises:** - **LiumError:** If verification or wallet polling fails. ### convert_alpha ```python def convert_alpha(usd: Any) -> lium.sdk.client.AlphaQuote: ``` Quote ``usd`` (USD) -> alpha via ``GET /balance/convert/alpha``. The response carries both the alpha amount to transfer (``converted``) and the subnet ``netuid`` the transfer must happen on. Hard-fails (no fallback) on a pay-API error: ``_request`` maps 503 -> `[`LiumServerError`](/developers/sdk/reference/exceptions/lium-server-error)` (a `[`LiumError`](/developers/sdk/reference/exceptions/lium-error)`), so a down subtensor / unavailable alpha price aborts the fund before any on-chain call. ### company_wallet ```python def company_wallet(app_id: str) -> str: ``` Resolve the Lium destination coldkey via ``GET /wallet/company/?app_id=``. Returns the company ``wallet_hash`` (the SS58 the pay-tao-api-v2 listener credits). Hard-fails (no fallback): a 404 (app has no wallet) maps to `[`LiumNotFoundError`](/developers/sdk/reference/exceptions/lium-not-found-error)` (a `[`LiumError`](/developers/sdk/reference/exceptions/lium-error)`), aborting before any on-chain call. ### backup_create ```python def backup_create( pod: lium.sdk.models.PodInfo, *, path: str = '/home', frequency_hours: int = 6, retention_days: int = 7 ) -> lium.sdk.models.BackupConfig: ``` Create or replace a backup configuration for a pod. **Arguments:** - **pod:** Pod to configure. - **path:** Filesystem path to back up. - **frequency_hours:** Backup interval in hours. - **retention_days:** Retention period in days. **Returns:** Created [`BackupConfig`](/developers/sdk/reference/models/backup-config). ### backup_now ```python def backup_now( pod: lium.sdk.models.PodInfo, *, name: str, description: str = '' ) -> Dict[str, Any]: ``` Trigger an immediate backup for a pod. **Arguments:** - **pod:** Pod to back up. - **name:** Backup name. - **description:** Optional description. **Returns:** API response payload from the run-now endpoint. ### backup_config ```python def backup_config(pod: lium.sdk.models.PodInfo) -> Optional[lium.sdk.models.BackupConfig]: ``` Return the backup configuration for a pod if one exists. **Arguments:** - **pod:** Pod to inspect. **Returns:** [`BackupConfig`](/developers/sdk/reference/models/backup-config) if present, otherwise ``None``. ### backup_list ```python def backup_list() -> List[lium.sdk.models.BackupConfig]: ``` List all backup configurations across all pods. **Returns:** List of [`BackupConfig`](/developers/sdk/reference/models/backup-config). ### backup_logs ```python def backup_logs(pod: lium.sdk.models.PodInfo) -> List[lium.sdk.models.BackupLog]: ``` Get recent backup logs for a pod. **Arguments:** - **pod:** Pod to inspect. **Returns:** List of [`BackupLog`](/developers/sdk/reference/models/backup-log) entries (possibly empty). ### backup_delete ```python def backup_delete(config_id: str) -> Dict[str, Any]: ``` Delete a backup configuration by ID. **Arguments:** - **config_id:** Backup configuration identifier. **Returns:** API response payload. ### restore ```python def restore( pod: lium.sdk.models.PodInfo, *, backup_id: str, restore_path: str = '/root' ) -> Dict[str, Any]: ``` Restore a backup to a pod. **Arguments:** - **pod:** Pod to restore to. - **backup_id:** ID of the backup to restore. - **restore_path:** Path where to restore the backup (default: /root). **Returns:** Response from the restore API. ### restore_logs ```python def restore_logs(pod: lium.sdk.models.PodInfo) -> List[lium.sdk.models.RestoreLog]: ``` Get recent restore logs for a pod. **Arguments:** - **pod:** Pod to inspect. **Returns:** List of [`RestoreLog`](/developers/sdk/reference/models/restore-log) entries (possibly empty). ### get_deployment_estimate ```python def get_deployment_estimate(executor_id: str, template_id: str) -> dict: ``` Estimate deployment time for a template on a node. **Arguments:** - **executor_id:** Node UUID. - **template_id:** Template UUID. **Returns:** Dict with ``estimated_seconds``, ``is_slow_machine``, ``warning_message``, ``is_cached_template``, and ``docker_image_size`` (image size in bytes, or ``None`` if unknown). ### balance ```python def balance() -> float: ``` Get current account balance. **Returns:** Floating-point balance value reported by ``/users/me``. ### topup_currencies ```python def topup_currencies(refresh: bool = False) -> List[Dict[str, Any]]: ``` List stablecoin currencies/networks supported for self-serve top-ups. **Arguments:** - **refresh:** Bypass the server-side cache and re-fetch from the provider. **Returns:** List of ``\{"code", "network", "decimals", "display_decimals"\}`` dicts. ### topup_create_invoice ```python def topup_create_invoice( amount: float, crypto_currency: str, crypto_network: str ) -> Dict[str, Any]: ``` Create a stablecoin top-up invoice for the current account. The returned ``deposit_address`` is where the exact ``crypto_amount`` of ``crypto_currency`` (on ``crypto_network``) must be sent. Once the provider confirms the transfer, the account balance is credited automatically. **Arguments:** - **amount:** Top-up amount in USD. - **crypto_currency:** Stablecoin code (e.g. ``"USDT"``), see `topup_currencies()`. - **crypto_network:** Network the stablecoin is sent on (e.g. ``"tron"``). **Returns:** Invoice dict including ``invoice_id``, ``deposit_address``, ``crypto_amount``, ``crypto_currency``, ``crypto_network``, ``exchange_rate`` and ``expires_at``. ### volumes ```python def volumes() -> List[lium.sdk.models.VolumeInfo]: ``` List all volumes for the current user. **Returns:** List of [`VolumeInfo`](/developers/sdk/reference/models/volume-info). ### volume ```python def volume(volume_id: str) -> lium.sdk.models.VolumeInfo: ``` Get a specific volume by ID. **Arguments:** - **volume_id:** Volume identifier. **Returns:** [`VolumeInfo`](/developers/sdk/reference/models/volume-info) for the requested volume. ### volume_create ```python def volume_create(name: str, *, description: str = '') -> lium.sdk.models.VolumeInfo: ``` Create a new volume. **Arguments:** - **name:** Volume name. - **description:** Optional description. **Returns:** Created [`VolumeInfo`](/developers/sdk/reference/models/volume-info). ### volume_update ```python def volume_update( volume_id: str, *, name: Optional[str] = None, description: Optional[str] = None ) -> lium.sdk.models.VolumeInfo: ``` Update a volume's metadata. **Arguments:** - **volume_id:** Volume identifier. - **name:** Optional new name. - **description:** Optional description. **Returns:** Updated [`VolumeInfo`](/developers/sdk/reference/models/volume-info). **Raises:** - **ValueError:** If neither ``name`` nor ``description`` is provided. ### volume_delete ```python def volume_delete(volume_id: str) -> Dict[str, Any]: ``` Delete a volume. **Arguments:** - **volume_id:** Volume identifier. **Returns:** API response payload from the delete request. ### schedule_termination ```python def schedule_termination(pod: lium.sdk.models.PodInfo, *, termination_time: str) -> Dict[str, Any]: ``` Schedule a pod for automatic termination at a future date and time. **Arguments:** - **pod:** Pod to schedule - **termination_time:** ISO 8601 formatted datetime string (e.g., "2025-10-17T15:30:00Z") **Returns:** Response from the schedule termination API ### cancel_scheduled_termination ```python def cancel_scheduled_termination(pod: lium.sdk.models.PodInfo) -> Dict[str, Any]: ``` Cancel a scheduled termination for a pod. **Arguments:** - **pod:** Pod to cancel the schedule for **Returns:** Response from the cancel scheduled termination API ### install_jupyter ```python def install_jupyter( pod: lium.sdk.models.PodInfo, *, jupyter_internal_port: int ) -> Dict[str, Any]: ``` Install Jupyter Notebook on a pod. **Arguments:** - **pod:** Pod to install Jupyter on - **jupyter_internal_port:** Internal port for Jupyter Notebook **Returns:** Response from the install Jupyter API --- # Config # Config ```python from lium.sdk import Config ``` Defined in `lium.sdk.config`. ```python Config( api_key: str, base_url: str = 'https://lium.io/api', base_pay_url: str = 'https://pay-api.lium.io', ssh_key_path: Optional[pathlib.Path] = None ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `api_key` | str | | | `base_url` | str | `'https://lium.io/api'` | | `base_pay_url` | str | `'https://pay-api.lium.io'` | | `ssh_key_path` | Optional[pathlib.Path] | `None` | ## Properties | Name | Type | Description | | --- | --- | --- | | `ssh_public_keys` | List[str] | Get SSH public keys. | ## Methods | Method | Description | | --- | --- | | [`load`](#load) | Load config from env/file with smart defaults. | ### load ```python def load(cls) -> lium.sdk.config.Config: ``` Load config from env/file with smart defaults. --- # machine # machine ```python from lium.sdk import machine ``` Defined in `lium.sdk.decorators`. ```python def machine( machine: str, template_id: Optional[str] = None, cleanup: bool = True, requirements: Optional[Sequence[str]] = None ): ``` Decorator to execute a function on a remote Lium machine. Creates a new pod, sends function source code and executes it remotely, returns the result, and optionally cleans up the pod. **Arguments:** - **machine:** Machine type (e.g., "1xH200", "1xA100") - **template_id:** Docker template ID (optional, uses default if not specified) - **cleanup:** Whether to delete the pod after execution (default: True) - **requirements:** Optional iterable of pip-installable packages to install on the pod --- # LiumAuthError # LiumAuthError ```python from lium.sdk import LiumAuthError ``` Defined in `lium.sdk.exceptions`. Authentication error. ```python LiumAuthError(*args, **kwargs) ``` --- # LiumError # LiumError ```python from lium.sdk import LiumError ``` Defined in `lium.sdk.exceptions`. Base exception for Lium SDK. ```python LiumError(*args, **kwargs) ``` --- # LiumNotFoundError # LiumNotFoundError ```python from lium.sdk import LiumNotFoundError ``` Defined in `lium.sdk.exceptions`. Resource not found (404). ```python LiumNotFoundError(*args, **kwargs) ``` --- # LiumRateLimitError # LiumRateLimitError ```python from lium.sdk import LiumRateLimitError ``` Defined in `lium.sdk.exceptions`. Rate limit exceeded. ```python LiumRateLimitError(*args, **kwargs) ``` --- # LiumServerError # LiumServerError ```python from lium.sdk import LiumServerError ``` Defined in `lium.sdk.exceptions`. Server error. ```python LiumServerError(*args, **kwargs) ``` --- # BackupConfig # BackupConfig ```python from lium.sdk import BackupConfig ``` Defined in `lium.sdk.models`. Backup configuration information. ```python BackupConfig( id: str, huid: str, pod_executor_id: str, backup_frequency_hours: int, retention_days: int, backup_path: str, is_active: bool, created_at: str, updated_at: Optional[str] = None ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `id` | str | | | `huid` | str | | | `pod_executor_id` | str | | | `backup_frequency_hours` | int | | | `retention_days` | int | | | `backup_path` | str | | | `is_active` | bool | | | `created_at` | str | | | `updated_at` | Optional[str] | `None` | --- # BackupLog # BackupLog ```python from lium.sdk import BackupLog ``` Defined in `lium.sdk.models`. Backup log information. ```python BackupLog( id: str, huid: str, backup_config_id: str, status: str, started_at: str, completed_at: Optional[str] = None, error_message: Optional[str] = None, progress: Optional[float] = None, backup_volume_id: Optional[str] = None, created_at: Optional[str] = None ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `id` | str | | | `huid` | str | | | `backup_config_id` | str | | | `status` | str | | | `started_at` | str | | | `completed_at` | Optional[str] | `None` | | `error_message` | Optional[str] | `None` | | `progress` | Optional[float] | `None` | | `backup_volume_id` | Optional[str] | `None` | | `created_at` | Optional[str] | `None` | --- # ExecutorInfo # ExecutorInfo ```python from lium.sdk import ExecutorInfo ``` Defined in `lium.sdk.models`. ```python ExecutorInfo( id: str, huid: str, machine_name: str, gpu_type: str, gpu_count: int, price_per_hour: float, price_per_gpu: float, location: Dict, specs: Dict, status: str, docker_in_docker: bool, ip: str, available_port_count: Optional[int] = None, effective_upload_speed_mbps: Optional[float] = None, effective_download_speed_mbps: Optional[float] = None, max_cuda_version: Optional[float] = None, tier: Optional[str] = None ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `id` | str | | | `huid` | str | | | `machine_name` | str | | | `gpu_type` | str | | | `gpu_count` | int | | | `price_per_hour` | float | | | `price_per_gpu` | float | | | `location` | Dict | | | `specs` | Dict | | | `status` | str | | | `docker_in_docker` | bool | | | `ip` | str | | | `available_port_count` | Optional[int] | `None` | | `effective_upload_speed_mbps` | Optional[float] | `None` | | `effective_download_speed_mbps` | Optional[float] | `None` | | `max_cuda_version` | Optional[float] | `None` | | `tier` | Optional[str] | `None` | ## Properties | Name | Type | Description | | --- | --- | --- | | `driver_version` | str | Extract GPU driver version from specs. | | `gpu_model` | str | Extract GPU model name from specs. | | `download_speed` | float | Effective download speed in Mbps (backend-authoritative; 0.0 if unknown). | | `upload_speed` | float | Effective upload speed in Mbps (backend-authoritative; 0.0 if unknown). | --- # PodInfo # PodInfo ```python from lium.sdk import PodInfo ``` Defined in `lium.sdk.models`. ```python PodInfo( id: str, name: str, status: str, huid: str, ssh_cmd: Optional[str], ports: Dict, created_at: str, updated_at: str, executor: Optional[lium.sdk.models.ExecutorInfo], template: Dict, removal_scheduled_at: Optional[str], jupyter_installation_status: Optional[str], jupyter_url: Optional[str], enable_volume_encryption: bool | None = None, volume_encryption_status: str | None = None ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `id` | str | | | `name` | str | | | `status` | str | | | `huid` | str | | | `ssh_cmd` | Optional[str] | | | `ports` | Dict | | | `created_at` | str | | | `updated_at` | str | | | `executor` | Optional[[ExecutorInfo](/developers/sdk/reference/models/executor-info)] | | | `template` | Dict | | | `removal_scheduled_at` | Optional[str] | | | `jupyter_installation_status` | Optional[str] | | | `jupyter_url` | Optional[str] | | | `enable_volume_encryption` | bool | None | `None` | | `volume_encryption_status` | str | None | `None` | ## Properties | Name | Type | Description | | --- | --- | --- | | `host` | Optional[str] | | | `username` | Optional[str] | | | `ssh_port` | int | Extract SSH port from command. | --- # RestoreLog # RestoreLog ```python from lium.sdk import RestoreLog ``` Defined in `lium.sdk.models`. Restore log information. ```python RestoreLog( id: str, huid: str, backup_id: str, pod_id: str, status: str, progress: float, created_at: str, started_at: Optional[str] = None, completed_at: Optional[str] = None, error_message: Optional[str] = None, logs: Optional[List[str]] = None, restore_path: Optional[str] = None ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `id` | str | | | `huid` | str | | | `backup_id` | str | | | `pod_id` | str | | | `status` | str | | | `progress` | float | | | `created_at` | str | | | `started_at` | Optional[str] | `None` | | `completed_at` | Optional[str] | `None` | | `error_message` | Optional[str] | `None` | | `logs` | Optional[List[str]] | `None` | | `restore_path` | Optional[str] | `None` | --- # SSHKey # SSHKey ```python from lium.sdk import SSHKey ``` Defined in `lium.sdk.models`. Public SSH key registered for the current user. ```python SSHKey( id: str, name: str, public_key: str, created_at: Optional[str] = None ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `id` | str | | | `name` | str | | | `public_key` | str | | | `created_at` | Optional[str] | `None` | --- # Template # Template ```python from lium.sdk import Template ``` Defined in `lium.sdk.models`. Template information. ```python Template( id: str, name: str, huid: str, docker_image: str, docker_image_tag: str, category: str, status: str ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `id` | str | | | `name` | str | | | `huid` | str | | | `docker_image` | str | | | `docker_image_tag` | str | | | `category` | str | | | `status` | str | | --- # VolumeInfo # VolumeInfo ```python from lium.sdk import VolumeInfo ``` Defined in `lium.sdk.models`. Volume information. ```python VolumeInfo( id: str, huid: str, name: str, description: str, created_at: str, updated_at: Optional[str] = None, current_size_bytes: int = 0, current_file_count: int = 0, current_size_gb: float = 0.0, current_size_mb: float = 0.0, last_metrics_update: Optional[str] = None ) ``` ## Fields | Name | Type | Default | | --- | --- | --- | | `id` | str | | | `huid` | str | | | `name` | str | | | `description` | str | | | `created_at` | str | | | `updated_at` | Optional[str] | `None` | | `current_size_bytes` | int | `0` | | `current_file_count` | int | `0` | | `current_size_gb` | float | `0.0` | | `current_size_mb` | float | `0.0` | | `last_metrics_update` | Optional[str] | `None` | --- # Screenshots for the Renters section # Screenshots for the Renters section The Renters pages reference these images. Capture each on a clean, signed-in [lium.io](https://lium.io) browser at viewport ≥ 1440×900 with the dark theme (default) and save as PNG into this directory. | File | URL to capture | What's in frame | |------|----------------|-----------------| | `browse-pods.png` | `https://lium.io/` | Browse Pods list with the right-hand filter rail visible. | | `access-ssh-keys.png` | `https://lium.io/ssh-keys` | Access page with the **SSH Keys** tab selected. | | `api-keys-list.png` | `https://lium.io/api-keys` | Access page with the **API Keys** tab selected (one+ row visible). | | `templates-list.png` | `https://lium.io/templates` | Templates page showing **My templates** cards and the start of **Browse templates**. | | `templates-create.png` | `https://lium.io/templates/create` | Create new template form, scrolled so Container Image, Tag, Volume, Internal Ports are all visible. | | `volumes-list.png` | `https://lium.io/volumes` | Volumes page with the rate calculator at the top and at least one existing volume row. | | `backups-list.png` | `https://lium.io/backups` | Global Backups page with filters and a few rows of history. | | `create-pod-top.png` | Click RENT NOW on any row from `/` | Top of the Create Pod page — Pod Name, machine card, Template card, SSH Key chip. | | `create-pod-summary.png` | Same as above, scrolled down | Right-hand summary panel and the **Deploy** button. | | `create-pod.png` | Same as above | Whole page (full-page screenshot if your tool supports it; otherwise the upper half). | | `create-pod-restore.png` | Same as above, scrolled to bottom | The **Restore** subsection on the Create Pod form. | | `pod-detail.png` | Click SEE DETAILS on a running pod | Pod detail header with **SSH CONNECTION**, INFO panel, TEMPLATE panel. | | `backups-tab-pod.png` | Pod detail page → **Backups** tab | Backups + Restore From Backup sections. | | `backups-config-modal.png` | Pod detail page → click **Backup** | The Backup Configuration modal open over the page. | | `scheduled-termination.png` | Pod detail page → click **DELETE** | The "Schedule pod deletion" menu / preset chooser. | ## Quick capture recipe The Chrome extension lets us grab page-level screenshots, but it can't write directly to the WSL filesystem in this build. Easiest manual flow: 1. Open the URL above in a logged-in browser at 1440×900. 2. Use your OS screenshot tool (or a full-page extension like GoFullPage) to save PNG. 3. Drop the file into this `docs/pod-users/assets/` folder using the exact name from the table. 4. `npm run build` from the lium-docs root and confirm the images render. ## Why these images matter The Renters pages are now UI-first — they walk a renter through the dashboard step by step. Without the images, every workflow still works (the prose is self-contained), but with them, a new user can match every "click X" to the actual button on screen and get to a running pod in 5 minutes flat. This file is generated by hand; update it when you add or rename a screenshot reference in any `pod-users/*.md` file. --- # GPU profiling with ncu # GPU profiling with ncu Profile CUDA kernels with [NVIDIA Nsight Compute](https://developer.nvidia.com/nsight-compute) (`ncu`) inside a Lium pod. ## Why a special machine `ncu` reads GPU performance counters, and on a default driver configuration those counters are restricted to admin users on the host. Inside an ordinary pod, any profiling attempt fails with: ``` ==ERROR== ERR_NVGPUCTRPERM - The user does not have permission to access NVIDIA GPU Performance Counters on the target device ``` Some providers open the counters on their machines (see the [provider guide](../providers/nodes/gpu-profiling.md)). Those machines carry a teal **GPU Profiling** badge in **Browse Pods**, and on them `ncu` reads the counters from an ordinary pod — no special container flags, just the CUDA toolkit (below). ## Find one 1. Open **Browse Pods** on [lium.io](https://lium.io). 2. Toggle **GPU Profiling (ncu)** in the filter rail. It shows machines where the counters are open **and** the whole machine is currently free. 3. Rent as usual. Because the counters are host-wide, a profiling machine is always rented **as a whole host**: the GPU-count selector is absent, every GPU goes to your pod, and nobody shares the machine while you hold it. The hourly price covers all GPUs — budget accordingly. The exclusivity is also your protection: open counters would let a co-tenant observe your GPU activity, so Lium never co-schedules anyone with you on these machines. ## Run ncu in the pod `ncu` ships with the CUDA toolkit. The default PyTorch template includes only the CUDA runtime, so install the toolkit first (or pick a `devel`-flavored template that bundles it). Pick the toolkit version the machine's driver supports — the **Max CUDA driver** row in the create-pod summary: ```bash apt-get update && apt-get install -y cuda-toolkit-12-6 ``` Then profile as you would on your own machine: ```bash ncu --version # sanity check ncu -o profile ./my_kernel_binary # profile a binary, write profile.ncu-rep ncu --set full -o profile python train.py ``` Copy the `.ncu-rep` report to your laptop (`scp`/`rsync` over the pod's SSH port) and open it in the Nsight Compute UI. If `ncu` still prints `ERR_NVGPUCTRPERM`, you are on a machine without open counters — look for the **GPU Profiling** badge and move to a machine that has it. --- # GPU Profiling (ncu) # GPU Profiling (ncu) GPU Profiling is an opt-in node capability: open the GPU performance counters on the host, and renters can run [NVIDIA Nsight Compute](https://developer.nvidia.com/nsight-compute) (`ncu`) inside the pods they rent from you. The switch is a single host driver flag — there is no portal toggle. Once the platform detects the flag, the node gets a teal **GPU Profiling** badge in the marketplace and starts matching the renter-side **GPU Profiling (ncu)** filter. ## Why enable it On a default driver configuration, NVIDIA restricts GPU performance counters to admin users, so `ncu` inside a rented pod fails with `ERR_NVGPUCTRPERM`. Machines with open counters are scarce on every GPU cloud, and renters do ask for them — kernel and inference-engine developers need `ncu` to tune CUDA code, and the demand is strongest on the latest silicon (B200 / B300), where kernels are still being optimized for the new architecture. The upside is demand-side — your node becomes visible to a renter segment that ordinary nodes cannot serve at all. On flagship nodes it is also one of the three ways to keep the idle payout: an 8× H200, B200, or B300 node must offer GPU Profiling, [GPU splitting](./gpu-splitting.md), or run inside an attested [confidential VM](./cvm.md) to earn the [unrented incentive](../rewards/emission.mdx#unrented-pool); with none of the three, the node stays active and rentable but forfeits that incentive while idle. ## The trade-off: whole-host rentals only GPU performance counters are **host-wide**. With the flag on, any workload on the machine can read the counters — including a co-tenant on a split rental, who could watch a neighbor's GPU activity through them. To keep tenants isolated, Lium leases a profiling-enabled node **only as a whole host**: - GPU splitting is suspended while the flag is on — every rental takes all GPUs; - renters see the node in the **GPU Profiling (ncu)** filter only while it has no active rental. Idle-time [default jobs](../portal/default-jobs.md) keep running while the node has no renter; the platform clears them before a renter's pod starts, as on any node. Enable the flag if whole-node rentals fit your fleet. A node that mostly earns from 1×/2× split rentals loses that segment while the counters are open. ## Enable the counters :::warning Only change the flag on an idle node Opening the counters while rented pods are running exposes those renters to the side-channel above and trips a platform alert. Wait until the node has no active rentals (or drain it first). ::: The commands below are for the Ubuntu host image Lium nodes ship with; on another distribution use its own initramfs tool. 1. Set the driver module option on the host: ```bash echo 'options nvidia NVreg_RestrictProfilingToAdminUsers=0' | sudo tee /etc/modprobe.d/nvidia-profiling.conf ``` 2. Rebuild the initramfs and reboot: ```bash sudo update-initramfs -u sudo reboot ``` 3. Verify the loaded state after reboot: ```bash grep RmProfilingAdminOnly /proc/driver/nvidia/params ``` Expected output: ``` RmProfilingAdminOnly: 0 ``` The validator reads the loaded driver state on its next hardware scan of the node. Expect the **GPU Profiling** badge on [lium.io](https://lium.io) within an hour, not instantly. Detection is fail-closed: if the platform cannot read the state (no NVIDIA driver, unreadable `/proc/driver/nvidia/params`), the node is treated as *not* profiling-capable. ## Disable the counters Remove the option and reboot — again, only while the node has no active rentals: ```bash sudo rm /etc/modprobe.d/nvidia-profiling.conf sudo update-initramfs -u sudo reboot ``` After the next hardware scan the badge disappears and the node returns to normal rental rules, including GPU splitting if configured.