1. 报错信息分析

在使用GPU训练YOLOV9的时候,出现报错:

torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 3.35 GiB (GPU 0; 23.65 GiB total capacity; 17.72 GiB already allocated; 1.29 GiB free; 21.27 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF

问题诊断:

  • 当 reserved 远大于 allocated ,说明‌显存碎片化严重‌

2. 解决方案:

2.1 方案一:优化Pytorch内存分配策略(亲测可行)

import os
# 调整内存碎片整理阈值,
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:128"

max_split_size_mb:限制 PyTorch 内存块分割最大值,推荐值:

  • 默认值:不设限制
  • 推荐值范围:32/64/128(需根据报错信息调整)
  • 减少显存碎片化,避免小块内存分配失败

2.2 方案二:降低BatchSize

  • 降低训练batchsize
  • 配合 num_workers 减少数据加载进程数

2.3 方案三:启用混合精度训练

  • 训练YOLO时,将amp设置为True

更多推荐