{
  "id": 551482,
  "title": "How to enable Callback in YOLO",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/551482",
  "author_name": "yoshio13",
  "post_date": "2024-12-13T13:10:17.714000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I tried to manage this <a href=\"https://www.kaggle.com/code/itsuki9180/czii-yolo11-submission-baseline\" target=\"_blank\">'CZII YOLO11 Training Baseline'</a>  with wandb and set up the callback, but the callback was being ignored. </p>\n<p>It seems that the cause of this issue is related to DDP. (<a href=\"https://github.com/ultralytics/ultralytics/issues/6168\" target=\"_blank\">https://github.com/ultralytics/ultralytics/issues/6168</a>)</p>\n<p>The solution is to save the code you want to run as a .py file and execute it with '! python -m torch.distributed.run --nproc_per_node=2 ./train.py'.\"</p>\n<pre><code>script = \nfrom ultralytics import YOLO\n\n# YOLOモデルのロード\nmodel = YOLO()  # 小さなYOLOv8モデルを使用\n\n# コールバック関数を定義\ndef on_train_batch_end(trainer):\n    ()\n    ()\n\n# コールバックを登録\nmodel.add_callback(, on_train_batch_end)\n\nresults = model.train(\n    data=,\n    epochs=,\n    warmup_epochs=,\n    optimizer=,\n    cos_lr=True,\n    lr0=-,\n    lrf=,\n    imgsz=,\n    device=, # &lt;-  one device(device=), you can  directly\n    weight_decay=,\n    batch=,\n    scale=,\n    flipud=,\n    fliplr=,\n    degrees=,\n    shear=,\n    mixup=,\n    copy_paste=,\n    seed=, # (｡•◡•｡)\n)\n\n\nwith (, )  :\n    .(script)\n</code></pre>\n<p>and run</p>\n<pre><code>! python -m torch.distributed. =2 ./train.py\n</code></pre>",
  "messages": [
    {
      "id": 3071181,
      "postDate": "2024-12-13T13:10:17.713Z",
      "content": "<p>I tried to manage this <a href=\"https://www.kaggle.com/code/itsuki9180/czii-yolo11-submission-baseline\" target=\"_blank\">'CZII YOLO11 Training Baseline'</a>  with wandb and set up the callback, but the callback was being ignored. </p>\n<p>It seems that the cause of this issue is related to DDP. (<a href=\"https://github.com/ultralytics/ultralytics/issues/6168\" target=\"_blank\">https://github.com/ultralytics/ultralytics/issues/6168</a>)</p>\n<p>The solution is to save the code you want to run as a .py file and execute it with '! python -m torch.distributed.run --nproc_per_node=2 ./train.py'.\"</p>\n<pre><code>script = \nfrom ultralytics import YOLO\n\n# YOLOモデルのロード\nmodel = YOLO()  # 小さなYOLOv8モデルを使用\n\n# コールバック関数を定義\ndef on_train_batch_end(trainer):\n    ()\n    ()\n\n# コールバックを登録\nmodel.add_callback(, on_train_batch_end)\n\nresults = model.train(\n    data=,\n    epochs=,\n    warmup_epochs=,\n    optimizer=,\n    cos_lr=True,\n    lr0=-,\n    lrf=,\n    imgsz=,\n    device=, # &lt;-  one device(device=), you can  directly\n    weight_decay=,\n    batch=,\n    scale=,\n    flipud=,\n    fliplr=,\n    degrees=,\n    shear=,\n    mixup=,\n    copy_paste=,\n    seed=, # (｡•◡•｡)\n)\n\n\nwith (, )  :\n    .(script)\n</code></pre>\n<p>and run</p>\n<pre><code>! python -m torch.distributed. =2 ./train.py\n</code></pre>",
      "rawMarkdown": "I tried to manage this ['CZII YOLO11 Training Baseline'](https://www.kaggle.com/code/itsuki9180/czii-yolo11-submission-baseline)  with wandb and set up the callback, but the callback was being ignored. \n\nIt seems that the cause of this issue is related to DDP. (https://github.com/ultralytics/ultralytics/issues/6168)\n\nThe solution is to save the code you want to run as a .py file and execute it with '! python -m torch.distributed.run --nproc_per_node=2 ./train.py'.\"\n\n```\nscript = \"\"\"\nfrom ultralytics import YOLO\n\n# YOLOモデルのロード\nmodel = YOLO('yolov8n.pt')  # 小さなYOLOv8モデルを使用\n\n# コールバック関数を定義\ndef on_train_batch_end(trainer):\n    print(f\"Batch {trainer.batch_size} ended.\")\n    print(f\"Loss: {trainer.loss.item()}\")\n\n# コールバックを登録\nmodel.add_callback(\"on_train_batch_end\", on_train_batch_end)\n\nresults = model.train(\n    data=\"/kaggle/input/czii-yolo-datasets/czii_conf.yaml\",\n    epochs=1,\n    warmup_epochs=10,\n    optimizer='AdamW',\n    cos_lr=True,\n    lr0=3e-4,\n    lrf=0.03,\n    imgsz=640,\n    device=\"0,1\", # <- if one device(device=\"0\"), you can execute directly\n    weight_decay=0.005,\n    batch=32,\n    scale=0,\n    flipud=0.5,\n    fliplr=0.5,\n    degrees=45,\n    shear=5,\n    mixup=0.2,\n    copy_paste=0.25,\n    seed=8620, # (｡•◡•｡)\n)\n\"\"\"\n\nwith open(\"train.py\", \"w\") as f:\n    f.write(script)\n```\n\nand run\n\n```\n! python -m torch.distributed.run --nproc_per_node=2 ./train.py\n```",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3071181": "I tried to manage this ['CZII YOLO11 Training Baseline'](https://www.kaggle.com/code/itsuki9180/czii-yolo11-submission-baseline)  with wandb and set up the callback, but the callback was being ignored. \n\nIt seems that the cause of this issue is related to DDP. (https://github.com/ultralytics/ultralytics/issues/6168)\n\nThe solution is to save the code you want to run as a .py file and execute it with '! python -m torch.distributed.run --nproc_per_node=2 ./train.py'.\"\n\n```\nscript = \"\"\"\nfrom ultralytics import YOLO\n\n# YOLOモデルのロード\nmodel = YOLO('yolov8n.pt')  # 小さなYOLOv8モデルを使用\n\n# コールバック関数を定義\ndef on_train_batch_end(trainer):\n    print(f\"Batch {trainer.batch_size} ended.\")\n    print(f\"Loss: {trainer.loss.item()}\")\n\n# コールバックを登録\nmodel.add_callback(\"on_train_batch_end\", on_train_batch_end)\n\nresults = model.train(\n    data=\"/kaggle/input/czii-yolo-datasets/czii_conf.yaml\",\n    epochs=1,\n    warmup_epochs=10,\n    optimizer='AdamW',\n    cos_lr=True,\n    lr0=3e-4,\n    lrf=0.03,\n    imgsz=640,\n    device=\"0,1\", # <- if one device(device=\"0\"), you can execute directly\n    weight_decay=0.005,\n    batch=32,\n    scale=0,\n    flipud=0.5,\n    fliplr=0.5,\n    degrees=45,\n    shear=5,\n    mixup=0.2,\n    copy_paste=0.25,\n    seed=8620, # (｡•◡•｡)\n)\n\"\"\"\n\nwith open(\"train.py\", \"w\") as f:\n    f.write(script)\n```\n\nand run\n\n```\n! python -m torch.distributed.run --nproc_per_node=2 ./train.py\n```"
  }
}