User Tools

Site Tools


tech:linux

This is an old revision of the document!


AI server Build

Ubunut 24.04

Install all drivers needed for using V100 on Ubuntu 24.04 LTS

Setting up an enterprise data center GPU like the NVIDIA V100 (Volta architecture) on Ubuntu 24.04 LTS requires the proprietary driver and the CUDA Toolkit. Since the V100 is typically used in a headless server environment for compute, machine learning, and AI workloads, a compute-only/headless driver installation is ideal. This avoids dragging in unnecessary graphical desktop packages.

Here is the clean, streamlined method using the official NVIDIA network repositories for Ubuntu 24.04.

Step 1: Clean Up Existing Drivers

To prevent conflicts with generic open-source or older drivers, remove any current installations:

sudo apt-get purge 'nvidia\*' 'cuda\*' -y  
sudo apt-get autoremove -y  

Step 2: Install Linux Kernel Headers

Ensure you have the correct kernel headers installed so the NVIDIA kernel modules can compile properly:

sudo apt-get update  
sudo apt-get install linux-headers-$(uname -r) build-essential -y  

Step 3: Setup the NVIDIA & CUDA Network Repository

NVIDIA provides a repository pinned to Ubuntu 24.04 (noble). Fetch the repository configurations and GPG keys so apt can safely pull the latest stable packages:

# Download the repository pin file to set priority
 
wget <https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86>\_64/cuda-ubuntu2404.pin  
sudo mv cuda-ubuntu2404.pin /etc/apt/preferences.d/cuda-repository-pin-600
 
# Download and install the repository package
 
wget <https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2404/x86>\_64/cuda-keyring\_1.1-1\_all.deb  
sudo dpkg -i cuda-keyring\_1.1-1\_all.deb
 
# Update apt package listings
 
sudo apt-get update  

Step 4: Install the Drivers and CUDA Toolkit

For enterprise cards like the V100, the stable branch (such as the nvidia-driver-550-server or nvidia-driver-565-server depending on repository syncs) is highly recommended.

To fetch the headless, data-center optimal drivers alongside the CUDA toolkit, run:

sudo apt-get -y install cuda-toolkit-12-8 nvidia-headless-server-550 nvidia-utils-550  

*Note: If you plan on deploying Docker containers that utilize the V100, you should install the container runtime package as well:* `sudo apt-get install -y nvidia-container-toolkit`

Step 5: Update Environment Variables

To ensure your system paths find the CUDA binaries, append them to your shell configuration (\~/.bashrc):

echo 'export PATH=/usr/local/cuda/bin${PATH:+:${PATH}}' \>\> \~/.bashrc  
echo 'export LD\_LIBRARY\_PATH=/usr/local/cuda/lib64${LD\_LIBRARY\_PATH:+:${LD\_LIBRARY\_PATH}}' \>\> \~/.bashrc  
source \~/.bashrc  

Step 6: Reboot and Verify

Restart the machine to initialize the newly compiled kernel modules:

sudo reboot  

Once the system is back up, verify that the OS recognizes the V100 and that the drivers are operational:

nvidia-smi  

You should see an output matrix detailing your V100, its current temperature, power consumption, and the driver/CUDA version currently running.

GeForce 210 Driver Override Guide

To configure a Driver Override for your GeForce 210, follow these steps to isolate the card from the proprietary NVIDIA driver and allow the open-source nouveau driver to manage the desktop display.

Step 1: Identify Device IDs

Use lspci to find the exact Hardware IDs for your GeForce 210 at PCI address 0000:01:00.0.

lspci -nn | grep -i “GeForce 210” Note the hexadecimal IDs (e.g., 10de:0a65).

root@test1:~# lspci -nn | grep "210"
01:00.0 VGA compatible controller [0300]: NVIDIA Corporation GT218 [GeForce 210] [10de:0a65] (rev a2)

Step 2: Isolate the Card with pci-stub

Bind the device to pci-stub at boot time so the proprietary driver cannot claim it.

* Open your GRUB configuration:

sudo nano /etc/default/grub * Append pci-stub.ids= with your hardware IDs to GRUB_CMDLINE_LINUX_DEFAULT:

GRUB_CMDLINE_LINUX_DEFAULT=“quiet splash pci-stub.ids=10de:0a65” * Update GRUB:

root@test1:~# grep GRUB_CMDLINE_LINUX_DEFAULT  /etc/default/grub
GRUB_CMDLINE_LINUX_DEFAULT="pci-stub.ids=10de:0a65"
sudo update-grub

Step 3: Ensure nouveau is Loaded

Ensure the open-source driver is not blacklisted and is forced to load at boot.

* Remove any blacklist nouveau lines found in files within /etc/modprobe.d/. * Add nouveau to your modules list:

echo “nouveau” | sudo tee -a /etc/modules-load.d/modules.conf

Step 4: Manual Driver Override (Optional)

If the card does not bind correctly after rebooting, create a manual override script to forcibly unbind any driver and attach nouveau.

#!/bin/bash
## Unbind from any existing driver
echo "0000:01:00.0" > /sys/bus/pci/devices/0000:01:00.0/driver/unbind
## Force bind to nouveau
echo "nouveau" > /sys/bus/pci/devices/0000:01:00.0/driver_override
echo "0000:01:00.0" > /sys/bus/pci/drivers/nouveau/bind

Step 5: Verify Configuration

Reboot your system and verify the kernel driver in use for each card.

lspci -nnk | grep -A 3 "VGA|3D"

The GeForce 210 should show Kernel driver in use: nouveau, while your V100s should remain under the nvidia driver.

Installing Docker on Ubuntu 24.04

To install Docker Engine on Ubuntu 24.04, the official and recommended method is to set up Docker's official apt repository. This ensures you get the latest stable version and automatic future updates.

Step 1: Remove Conflicting Packages

Before starting, remove any default or unofficial Docker packages to prevent conflicts:

sudo apt remove docker.io docker-doc docker-compose docker-compose-v2 podman-docker containerd runc

Step 2: Set Up the Docker Repository

Update your package index and install the prerequisite security certificates:

sudo apt update
sudo apt install -y ca-certificates curl gnupg

Next, add Docker's official GPG key to verify package authenticity:

sudo install -m 0755 -d /etc/apt/keyrings
sudo curl -fsSL https://docker.com -o /etc/apt/keyrings/docker.asc
sudo chmod a+r /etc/apt/keyrings/docker.asc

Add the stable repository to your system's apt sources:

echo \
  "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.asc] https://docker.com \
  $(. /etc/os-release && echo "$VERSION_CODENAME") stable" | \
  sudo tee /etc/apt/sources.list.d/docker.list > /dev/null

Step 3: Install Docker Engine & Plugins

Update your repository index to include Docker's packages, then install Docker Engine, the CLI, and Docker Compose V2:

sudo apt update
sudo apt install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin

Step 4: Verify the Installation

Check if the service is up and running:

sudo systemctl status docker

Run the default test container to verify it pulls and executes properly:

sudo docker run hello-world

Optional: Run Docker Without Sudo

By default, Docker requires root privileges. To run commands as a regular user, add yourself to the docker group:

  • Create the group (usually exists already):
sudo groupadd docker
  • Add your user to the group:
sudo usermod -aG docker $USER
  • Apply the changes immediately without logging out:
newgrp docker
  • Verify by running without sudo:
docker run hello-world

For detailed troubleshooting or specific alternative guides, you can refer to the Official Docker Engine Installation Docs.

Enabling NVIDIA AI and GPU Acceleration on Docker

To enable NVIDIA AI and GPU acceleration on Docker, you must install and configure the NVIDIA Container Toolkit. This acts as the bridge that allows Docker containers to access your host machine's physical GPU.

Here is how to set it up on a Linux host (such as Ubuntu).

1. Install Prerequisites

Ensure you have the official NVIDIA drivers installed on your host system, along with Docker Engine. Validate your driver installation by running:

nvidia-smi

(You should see a table displaying your GPU information.)

2. Install the NVIDIA Container Toolkit

Add the official package repositories and install the toolkit:

# Add the NVIDIA package repository
curl -fsSL https://github.io | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg

curl -s -L https://github.io | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list

# Update package list and install
sudo apt-get update
sudo apt-get install -y nvidia-container-toolkit

3. Configure the Docker Runtime

Configure the Docker daemon to automatically recognize the NVIDIA container runtime, then restart the Docker service:

# Configure Docker to use the NVIDIA runtime
sudo nvidia-ctk runtime configure --runtime=docker

# Restart the Docker service to apply changes
sudo systemctl restart docker

4. Verify GPU Access in Docker

Test the configuration by running a lightweight, official CUDA container. Passing the –gpus all flag tells Docker to expose your graphics cards to the container:

sudo docker run --rm --gpus all ubuntu nvidia-smi

If successful, the container will print out your host GPU details exactly like the native nvidia-smi command did.


Alternative: Docker Desktop (Windows / Mac)

If you are developing locally via Docker Desktop, you do not need to manually install the toolkit. Instead, you can leverage native AI features:

  • WSL 2 (Windows): Ensure WSL integration is turned on under Settings > Resources > WSL Integration in Docker Desktop.
  • Docker Model Runner (DMR): Go to Settings > AI, check Enable Docker Model Runner, and select Enable GPU-backed inference to pull and manage local AI models natively using commands like docker model run.

If you plan to deploy enterprise-grade AI workloads, you can also authenticate your Docker client to the NVIDIA NGC Container Registry using an API key. This gives you direct access to fine-tuned, GPU-optimized AI models and microservices.

tech/linux.1781481543.txt.gz · Last modified: by glong

Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki