Ultralytics YOLO27:

AMD GPU Training and Inference with Ultralytics YOLO, ROCm and MIGraphX#

Linux x86_64 only

This guide targets Linux x86_64 hosts with the AMD GPU kernel driver (amdgpu) installed, and all commands below are for a Linux shell. The MIGraphX execution provider is published only as Linux x86_64 wheels for Python 3.11 or newer, so native Windows is not supported, and Windows Subsystem for Linux (WSL2) has not been validated yet.

Ultralytics YOLO runs on AMD Instinct and supported Radeon GPUs through ROCm, AMD's open GPU compute stack. With a ROCm build of PyTorch, you train, validate, and predict with .pt models using the same device=0 argument as on any other GPU, with no code changes.

For deployment, export the trained model to ONNX and run it through MIGraphX, AMD's graph-optimization and inference engine. The Ultralytics ONNX backend detects ROCm and selects the ONNX Runtime MIGraphXExecutionProvider automatically, so the same predict call that runs on NVIDIA GPUs runs on AMD GPUs. Scheduled Ultralytics continuous integration runs ONNX export and MIGraphX inference for every task on AMD GPU hardware.

Installation#

The Python stack installs entirely through pip from AMD's ROCm wheel indexes, so no apt packages or root access are needed for PyTorch, the plugin, or the ROCm runtime libraries, provided the host already has the AMD GPU kernel driver (amdgpu / /dev/kfd) in place.

Installation

Detect your GPU architecture (for example gfx1151 on Strix Halo) and install only the PyTorch kernels it needs:

# Read the GPU architecture from the amdgpu driver, e.g. device-gfx1151
GPU_ARCH=$(grep -rhs '^gfx_target_version' /sys/class/kfd/kfd/topology/nodes | awk '$2 {printf "device-gfx%d%d%x\n", int($2/10000), int($2/100)%100, $2%100}' | sort -u | paste -sd, -)

# ROCm PyTorch + torchvision for the detected architecture
pip install "torch[${GPU_ARCH:?No AMD GPU found}]" "torchvision[$GPU_ARCH]" --index-url https://stable.repo.amd.com/rocm/whl-next/

# Ultralytics
pip install ultralytics

Ultralytics installs onnx automatically on the first ONNX export. On the first ONNX inference on a ROCm system, it also installs the MIGraphX execution provider plugin (onnxruntime-ep-migraphx) and the MIGraphX runtime libraries it loads (migraphx-libs), pinned to versions tested together. The plugin is published for Python 3.11 or newer on Linux x86_64 only; on other Python versions or architectures, or with YOLO_AUTOINSTALL=false and no plugin installed, ONNX inference falls back to the CPU with a warning. For detailed instructions and best practices, check our YOLO26 Installation guide; if you encounter any difficulties, consult our Common Issues guide.

Verify that PyTorch sees the AMD GPU before training or exporting:

import torch

print(torch.cuda.is_available())  # True on a working ROCm install
print(torch.cuda.get_device_name(0))  # AMD GPU name, e.g. AMD Radeon 8060S
print(torch.version.hip)  # HIP version of the ROCm PyTorch build, None on CUDA or CPU builds

Train on AMD GPUs with PyTorch ROCm#

Native training, validation, and prediction on .pt models run on AMD GPUs through PyTorch ROCm, independent of the MIGraphX inference path below. Install a ROCm build of PyTorch as shown in Installation and use the same device arguments as any other Ultralytics Train run:

Why AMD GPUs report as CUDA in Ultralytics

The ROCm build of PyTorch uses HIP internally but deliberately reuses the torch.cuda interfaces, so torch.cuda.is_available() returns True and AMD GPUs are addressed with standard CUDA-style IDs. Use device=0 or device=cuda:0 for AMD GPUs; rocm is not a PyTorch device type. See the PyTorch HIP semantics for details.

ROCm Training
from ultralytics import YOLO

model = YOLO("yolo26n.pt")

# Train on the first AMD GPU exposed by PyTorch ROCm
results = model.train(data="coco8.yaml", epochs=100, imgsz=640, device=0)

# Train across two AMD GPUs
results = model.train(data="coco8.yaml", epochs=100, imgsz=640, device=[0, 1])

Validation and prediction of .pt models use the same device selection:

yolo detect val model=path/to/best.pt data=coco8.yaml device=0
yolo predict model=path/to/best.pt source=path/to/image.jpg device=0
Mixed precision on ROCm

Ultralytics enables AMP by default and disables it automatically if a pre-training check finds that mixed-precision results diverge from full precision. ROCm AMP behavior can change with the PyTorch and ROCm versions, so if a run still produces NaN losses or zero mAP, train with amp=False:

yolo detect train data=coco8.yaml model=yolo26n.pt device=0 amp=False

Export and Inference with MIGraphX#

ONNX Runtime is a cross-platform inference engine that runs a single ONNX model on many hardware backends through pluggable execution providers (EPs). Each EP maps the ONNX graph onto a specific accelerator: CUDA for NVIDIA GPUs, CoreML for Apple silicon, and MIGraphX for AMD GPUs on ROCm, which ships as the onnxruntime-ep-migraphx plugin and compiles the graph into a tuned program for your GPU.

MIGraphX inference uses the standard ONNX export. Export the model to ONNX (format="onnx") and Ultralytics inference will automatically use the MIGraphXExecutionProvider backend to run the ONNX model on your AMD GPU.

Key Features of MIGraphX Inference#

  • Automatic provider selection: On a ROCm host the ONNX backend registers the plugin and selects MIGraphXExecutionProvider with no code changes. If the plugin is missing or fails to load, inference falls back to the CPU with a warning.
  • Graph optimization: MIGraphX applies operator fusion, memory planning, and kernel selection tuned for AMD GPU architectures.
  • Zero-copy IO binding: Inputs and outputs are bound directly to GPU tensors through the DLPack protocol, avoiding host round-trips during inference.
  • Precision options: Run FP32 or export an FP16 ONNX model for reduced-precision inference.
  • Portable artifact: A single .onnx file runs on CPUs, NVIDIA GPUs, and AMD GPUs, letting you target multiple platforms from one export.
  • Reproducible deployment: The full stack (ROCm PyTorch, the MIGraphX plugin, and its libraries) installs through pip from AMD's ROCm wheel indexes.

Supported Tasks#

MIGraphX inference supports all seven Ultralytics tasks. Semantic segmentation and depth estimation are available only with YOLO26, the only family that ships those heads.

TaskYOLOv8YOLO11YOLO26
Detect✅✅✅
Segment✅✅✅
Semantic❌❌✅
Depth❌❌✅
Classify✅✅✅
Pose✅✅✅
OBB✅✅✅

Usage#

Before diving into the usage instructions, be sure to check out the range of YOLO26 models offered by Ultralytics. This will help you choose the most appropriate model for your project requirements.

The ONNX format supports the Export, Predict, and Validate modes. Inference and validation on an AMD GPU require a ROCm system with the MIGraphX plugin installed. Export your model, then load the exported model to run inference or validate its accuracy on device=0.

Conflict with onnxruntime-gpu

onnxruntime-gpu and the standard onnxruntime package install into the same onnxruntime Python module, so whichever is installed last overwrites the other. Uninstalling only one of them leaves the module broken, and older onnxruntime-gpu releases can crash MIGraphX inference. If an earlier setup installed onnxruntime-gpu, run pip uninstall -y onnxruntime-gpu onnxruntime, and Ultralytics reinstalls the standard onnxruntime on the next ONNX inference.

Export
from ultralytics import YOLO

# Load a YOLO26 model
model = YOLO("yolo26n.pt")

# Export the model to ONNX format
model.export(format="onnx")  # creates 'yolo26n.onnx'
Predict
from ultralytics import YOLO

# Load the exported ONNX model and run inference on an AMD GPU
model = YOLO("yolo26n.onnx")

# The MIGraphX execution provider is selected automatically on ROCm
results = model.predict("https://ultralytics.com/images/bus.jpg", device=0)
Validate
from ultralytics import YOLO

# Load the exported ONNX model
model = YOLO("yolo26n.onnx")

# Validate accuracy on the COCO8 dataset on an AMD GPU
metrics = model.val(data="coco8.yaml", device=0)

Export Arguments#

MIGraphX inference reuses the ONNX export arguments. The most relevant options for AMD GPU deployment are:

ArgumentTypeDefaultDescription
formatstr'onnx'Target format for the exported model. Use onnx for MIGraphX EP inference.
imgszint or tuple640Desired image size for the model input. Can be an integer for square images or a tuple (height, width).
quantizeint or strNonePrecision of the exported ONNX model: 16 (FP16) for reduced-precision inference; 32/unset is FP32.
dynamicboolFalseAllows dynamic input sizes. Static shapes let MIGraphX compile a specialized program and enable zero-copy IO binding.
simplifyboolTrueSimplifies the model graph with onnxslim, potentially improving performance and compatibility.
opsetintNoneONNX opset version for compatibility with different runtimes. If not set, uses the latest supported version.
nmsbool, optionalNoneSelect raw output (None, default), embedded NMS (True), or the NMS-free head (False). Embedded NMS (True) is not yet supported by the MIGraphX EP (ROCm/AMDMIGraphX#5246).
batchint1Export batch size, or the max number of images the exported model processes concurrently in predict mode.
devicestrNoneDevice for exporting: GPU (device=0), CPU (device=cpu).

For the full list of export arguments, see the ONNX integration and the Ultralytics documentation page on exporting.

Deploying on AMD GPUs with MIGraphX#

Compiled-program cache

The MIGraphX EP compiles the graph on the first session, which dominates initial load time. Ultralytics caches the compiled program per model under the Ultralytics config directory so later loads of the same model skip recompilation. The cache keeps the 8 most recently used models and evicts older ones, so it cannot grow without bound. Set ORT_MIGRAPHX_CACHE_DIR to override the location, and delete the cache directory to force recompilation after upgrading ROCm or MIGraphX.

Compile time

Ultralytics disables MIGraphX Winograd convolution kernels by default (MIGRAPHX_DISABLE_WINOGRAD=1) to cut cold-compile time on YOLO graphs with no measurable inference change (ROCm/AMDMIGraphX#5234); set MIGRAPHX_DISABLE_WINOGRAD=0 to re-enable them.

For a ready-to-run environment, Dockerfile-amd builds the ultralytics/ultralytics:latest-amd image with ROCm PyTorch and the MIGraphX EP preinstalled. See Using GPUs in the Docker Quickstart for the docker run flags that expose AMD GPUs to the container.

Benchmarks#

YOLO26 benchmarks below were run by the Ultralytics team on an AMD Radeon 8060S GPU (Ryzen AI Max+ PRO 395), comparing speed and accuracy between PyTorch and ONNX on the MIGraphX execution provider. Both formats run at FP32 precision on the GPU, PyTorch through ROCm and ONNX through MIGraphX.

Performance
ModelFormatStatusSize (MB)metrics/mAP50-95(B)Inference time (ms/im)
YOLO26nPyTorch✅5.30.40894.1
YOLO26nONNX (MIGraphX)✅9.50.40892.1
YOLO26sPyTorch✅19.50.48507.2
YOLO26sONNX (MIGraphX)✅36.50.48504.7
YOLO26mPyTorch✅42.20.532314.6
YOLO26mONNX (MIGraphX)✅78.20.532310.9
YOLO26lPyTorch✅50.70.548918.5
YOLO26lONNX (MIGraphX)✅95.00.548915.0
YOLO26xPyTorch✅113.20.574837.4
YOLO26xONNX (MIGraphX)✅212.90.574827.1

Benchmarked with Ultralytics 8.4.165 (depth estimation with 8.4.172)

Note

Validation for the above benchmarks was done at batch size 1 on the full validation sets, at 640 for detection, segmentation and pose estimation, 2048 for semantic segmentation, 768 for depth estimation, 224 for classification and 1024 for OBB. Inference time does not include pre/post-processing or the one-time MIGraphX compile.

Support at a Glance#

Support for one AMD product or runtime does not imply support for every AMD accelerator. This table summarizes the current status in the Ultralytics Python package.

AMD product or runtimeSupportUsage or status
AMD Instinct and supported Radeon GPUs with ROCm✅Train, validate, and run native PyTorch models with device=0 or device=cuda:0.
MIGraphX inference✅Run exported ONNX models on AMD GPUs through the MIGraphX EP. All YOLO26 tasks are supported.
Multi-GPU ROCm✅Use device=0,1 or device=[0, 1]; distributed execution follows the installed PyTorch ROCm stack.
ROCm Automatic Mixed Precision (AMP)⚠️Available when the installed PyTorch and ROCm versions pass Ultralytics AMP checks; use amp=False if incompatible.
AMD Docker image✅latest-amd ships ROCm PyTorch with the MIGraphX EP preinstalled.
Native MIGraphX export❌Export to ONNX with format="onnx" and run it on the MIGraphX EP for AMD GPU inference.
Windows DirectML❌No DirectML training or prediction backend in the Python package.
Ryzen AI NPU❌No native NPU integration; external ONNX/Vitis AI workflows are community-managed.
AMD Xilinx Versal AI Edge Gen 2 NPUs✅Export with format="xilinx"; see the AMD Xilinx guide.
AMD Xilinx Zynq, Kria, Versal AI Edge Gen 1❌No native export; use AMD Vitis AI flows in the AMD Xilinx guide.
AMD CPUs✅ CPUUse device=cpu; standard CPU execution, not an AMD-specific acceleration backend.
Check AMD and PyTorch compatibility first

ROCm availability depends on the exact GPU, operating system, ROCm version, and PyTorch build. Confirm your hardware in AMD's ROCm compatibility matrix before installing.

Summary#

In this guide, you learned how to train Ultralytics YOLO26 models on AMD GPUs with PyTorch ROCm using the standard device=0 argument, then export them to ONNX for accelerated inference through the ONNX Runtime MIGraphX execution provider. The ONNX backend selects MIGraphXExecutionProvider automatically, caches the compiled program for fast subsequent loads, and supports all YOLO26 tasks with no code changes.

For other deployment targets, browse the integration guide page, and compare export formats with Benchmark mode.

FAQ#

  • Yes. Install a ROCm build of PyTorch as shown in Installation, then train .pt models with device=0, or device=0,1 for multiple GPUs. Training runs through PyTorch ROCm and does not use MIGraphX or the ONNX plugin. See Train on AMD GPUs with PyTorch ROCm for examples.

  • Export your model to ONNX, then run it with device=0 on a ROCm system. Ultralytics installs the MIGraphX plugin on the first ONNX inference:

    Usage
    from ultralytics import YOLO
    
    # Export to ONNX
    model = YOLO("yolo26n.pt")
    model.export(format="onnx")  # creates 'yolo26n.onnx'
    
    # Run inference on the AMD GPU (MIGraphX EP selected automatically)
    onnx_model = YOLO("yolo26n.onnx")
    results = onnx_model.predict("https://ultralytics.com/images/bus.jpg", device=0)
  • No. On a ROCm system, the Ultralytics ONNX backend detects HIP, installs and registers the onnxruntime-ep-migraphx plugin, and selects MIGraphXExecutionProvider automatically. If the plugin cannot be installed or loaded, for example on Python older than 3.11 or on ARM64, inference falls back to the CPU with a warning.

  • MIGraphX compiles the ONNX graph into an optimized program on the first session, which dominates initial load time. Ultralytics caches the compiled program, so later loads of the same model skip recompilation. See the compiled-program cache note for its location and size limit.

  • The installed PyTorch package is probably a CPU or CUDA build rather than a ROCm build, or the GPU is not supported by the active ROCm stack. Reinstall PyTorch from AMD's ROCm wheel index as shown in Installation, check that torch.version.hip is not None, and confirm the GPU in AMD's ROCm compatibility matrix. Also make sure the user can access /dev/kfd and /dev/dri, typically by joining the video and render groups.

  • This is expected. PyTorch ROCm intentionally reuses the torch.cuda API and CUDA-style device strings for Python compatibility. Use device=0 or device=cuda:0; the model still executes through HIP and ROCm on the AMD GPU.

  • Not through the Python package. DirectML has no training or prediction backend, and Ryzen AI NPUs are not exposed through PyTorch ROCm. Community workflows may export to ONNX and run with AMD's external Ryzen AI or Vitis AI tools, but those runtimes are outside the supported Ultralytics execution path. For embedded AMD Xilinx Versal AI Edge Series Gen 2 NPUs, export with format="xilinx"; for Zynq, Kria and first-generation Versal AI Edge devices, see the AMD Xilinx guide.

  • Pass the GPU index directly, for example device=1 for the second GPU, for both PyTorch and MIGraphX EP inference. To restrict a process or container to specific GPUs, set HIP_VISIBLE_DEVICES (or ROCR_VISIBLE_DEVICES for containers), for example HIP_VISIBLE_DEVICES=2; the visible GPUs are then renumbered from device=0.

Contributors

Comments