{
  "id": 64734,
  "title": "19th place solution overview [and trained weights]",
  "url": "/competitions/google-ai-open-images-object-detection-track/writeups/radek-19th-place-solution-overview-and-trained-wei",
  "author_name": "",
  "post_date": "2018-09-01T05:51:06.133201500Z",
  "votes": 24,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I trained <a href=\"https://pjreddie.com/darknet/yolo/\">yolov3</a> with SPP using <a href=\"https://pjreddie.com/media/files/darknet53.conv.74\">darknet53.conv.74</a> weights pretrained on Imagenet.</p>\n\n<p>I followed the training regime and used the same architecture that the creators of yolo used for training on COCO. Most likely better results could be achieved on the Open Images dataset should the capacity of the model be increased.</p>\n\n<p>I emulated the conditions of training on COCO and created two custom datasets. I initially (up to ~ 250k batches IIRC) trained on ~320k files, increasing the dataset to ~430k files onwards. I picked the files to train on to ensure as balanced a representation of each class as possible (picked all images that contained rare labels - limited number of images to include for more popular labels to arrive at the target file count).</p>\n\n<p>I trained the model thx to the GCP coupon using a preemptible instance / premptible GPUs. I went with 16 vCPUs and 4 preemptible P100s. I didn't measure this precisely but I think in 24 hrs the model trained on ~50k batches, maybe a bit more. The full training schedule is just over 500k batches.</p>\n\n<p>I believe that appropriate makeup of the trainset is key to achieving a good result. This is an observation that I have not verified though.</p>\n\n<p>I have not modified the train set labels in any way. I also don't believe I have expanded the labels before training to include parent categories but might be wrong on that one. Doing something reasonable with the label hierarchy should grant further improvement in performance.</p>\n\n<p>My private LB score is slightly different as I tried training on a subset of labels / files for trainingtrain set selected in a specific way to help me in the VRD track but I didn't get very far (partially due to this likely being the wrong approach and partially due to running out of time).</p>\n\n<p>I am sharing the final weights and the config files in this github <a href=\"https://github.com/radekosmulski/yolo_open_images\">repository</a>.</p>\n\n<p>Here is the citation of the paper that introduces yolov3:</p>\n\n<p>&gt; @article{yolov3,   title={YOLOv3: An Incremental Improvement}, <br>\n&gt; author={Redmon, Joseph and Farhadi, Ali},   journal = {arXiv}, <br>\n&gt; year={2018} }</p>\n\n<p>And <a href=\"https://arxiv.org/abs/1804.02767\">here</a> is where you can find the paper on arxiv.</p>",
  "messages": [
    {
      "id": "379875",
      "postDate": "09/01/2018 05:51:06",
      "content": "<p>I trained <a href=\"https://pjreddie.com/darknet/yolo/\">yolov3</a> with SPP using <a href=\"https://pjreddie.com/media/files/darknet53.conv.74\">darknet53.conv.74</a> weights pretrained on Imagenet.</p>\n\n<p>I followed the training regime and used the same architecture that the creators of yolo used for training on COCO. Most likely better results could be achieved on the Open Images dataset should the capacity of the model be increased.</p>\n\n<p>I emulated the conditions of training on COCO and created two custom datasets. I initially (up to ~ 250k batches IIRC) trained on ~320k files, increasing the dataset to ~430k files onwards. I picked the files to train on to ensure as balanced a representation of each class as possible (picked all images that contained rare labels - limited number of images to include for more popular labels to arrive at the target file count).</p>\n\n<p>I trained the model thx to the GCP coupon using a preemptible instance / premptible GPUs. I went with 16 vCPUs and 4 preemptible P100s. I didn't measure this precisely but I think in 24 hrs the model trained on ~50k batches, maybe a bit more. The full training schedule is just over 500k batches.</p>\n\n<p>I believe that appropriate makeup of the trainset is key to achieving a good result. This is an observation that I have not verified though.</p>\n\n<p>I have not modified the train set labels in any way. I also don't believe I have expanded the labels before training to include parent categories but might be wrong on that one. Doing something reasonable with the label hierarchy should grant further improvement in performance.</p>\n\n<p>My private LB score is slightly different as I tried training on a subset of labels / files for trainingtrain set selected in a specific way to help me in the VRD track but I didn't get very far (partially due to this likely being the wrong approach and partially due to running out of time).</p>\n\n<p>I am sharing the final weights and the config files in this github <a href=\"https://github.com/radekosmulski/yolo_open_images\">repository</a>.</p>\n\n<p>Here is the citation of the paper that introduces yolov3:</p>\n\n<p>&gt; @article{yolov3,   title={YOLOv3: An Incremental Improvement}, <br>\n&gt; author={Redmon, Joseph and Farhadi, Ali},   journal = {arXiv}, <br>\n&gt; year={2018} }</p>\n\n<p>And <a href=\"https://arxiv.org/abs/1804.02767\">here</a> is where you can find the paper on arxiv.</p>",
      "rawMarkdown": "I trained [yolov3][1] with SPP using [darknet53.conv.74](https://pjreddie.com/media/files/darknet53.conv.74) weights pretrained on Imagenet.\n\nI followed the training regime and used the same architecture that the creators of yolo used for training on COCO. Most likely better results could be achieved on the Open Images dataset should the capacity of the model be increased.\n\nI emulated the conditions of training on COCO and created two custom datasets. I initially (up to ~ 250k batches IIRC) trained on ~320k files, increasing the dataset to ~430k files onwards. I picked the files to train on to ensure as balanced a representation of each class as possible (picked all images that contained rare labels - limited number of images to include for more popular labels to arrive at the target file count).\n\nI trained the model thx to the GCP coupon using a preemptible instance / premptible GPUs. I went with 16 vCPUs and 4 preemptible P100s. I didn't measure this precisely but I think in 24 hrs the model trained on ~50k batches, maybe a bit more. The full training schedule is just over 500k batches.\n\nI believe that appropriate makeup of the trainset is key to achieving a good result. This is an observation that I have not verified though.\n\nI have not modified the train set labels in any way. I also don't believe I have expanded the labels before training to include parent categories but might be wrong on that one. Doing something reasonable with the label hierarchy should grant further improvement in performance.\n\nMy private LB score is slightly different as I tried training on a subset of labels / files for trainingtrain set selected in a specific way to help me in the VRD track but I didn't get very far (partially due to this likely being the wrong approach and partially due to running out of time).\n\nI am sharing the final weights and the config files in this github [repository](https://github.com/radekosmulski/yolo_open_images).\n\nHere is the citation of the paper that introduces yolov3:\n\n&gt; @article{yolov3,   title={YOLOv3: An Incremental Improvement},  \n&gt; author={Redmon, Joseph and Farhadi, Ali},   journal = {arXiv},  \n&gt; year={2018} }\n\nAnd [here][2] is where you can find the paper on arxiv.\n\n  [1]: https://pjreddie.com/darknet/yolo/\n  [2]: https://arxiv.org/abs/1804.02767",
      "votes": null
    },
    {
      "id": "379977",
      "postDate": "09/01/2018 11:19:19",
      "content": "<p>Congrats Radek! Amazing achievement!</p>",
      "rawMarkdown": "Congrats Radek! Amazing achievement!",
      "votes": null
    },
    {
      "id": "380697",
      "postDate": "09/03/2018 08:54:24",
      "content": "<p>Congratulations on the silver medal!</p>\n\n<p>Can you share what input resolution did you use in your model?</p>\n\n<p>EDIT: Found it on the link :)</p>",
      "rawMarkdown": "Congratulations on the silver medal!\n\nCan you share what input resolution did you use in your model?\n\nEDIT: Found it on the link :)",
      "votes": null
    },
    {
      "id": "380705",
      "postDate": "09/03/2018 09:22:37",
      "content": "<p>Thank you! Darknet does the rescaling on the fly for you to a couple of different resolutions to train on objects of various sizes. I trained with a max size of 608x608.</p>\n\n<p>The data I downloaded and used for training came from a storage bucket whose address I believe I have found on github in the object detection API repository, though might have been somewhere else. I don't recall now (and don't have access to my rig at the moment) but I believe the images were rescaled to either 1024 for the longer edge or were slightly bigger.</p>",
      "rawMarkdown": "Thank you! Darknet does the rescaling on the fly for you to a couple of different resolutions to train on objects of various sizes. I trained with a max size of 608x608.\n\nThe data I downloaded and used for training came from a storage bucket whose address I believe I have found on github in the object detection API repository, though might have been somewhere else. I don't recall now (and don't have access to my rig at the moment) but I believe the images were rescaled to either 1024 for the longer edge or were slightly bigger.",
      "votes": null
    },
    {
      "id": "381109",
      "postDate": "09/04/2018 04:25:30",
      "content": "<p>Congrats Radek!! Tks for sharing.</p>",
      "rawMarkdown": "Congrats Radek!! Tks for sharing.",
      "votes": null
    },
    {
      "id": "388814",
      "postDate": "09/17/2018 16:26:08",
      "content": "<p>Congrats!\nHow much GPU mem is required to run inference?\nI have OOM with 8GB RAM.\nIs 12GB enough to run it? (I am trying to run webcam demo)</p>",
      "rawMarkdown": "Congrats!\nHow much GPU mem is required to run inference?\nI have OOM with 8GB RAM.\nIs 12GB enough to run it? (I am trying to run webcam demo)",
      "votes": null
    },
    {
      "id": "389079",
      "postDate": "09/18/2018 05:41:29",
      "content": "<p>I have not tried running the webcam demo, quite surprised 8GB is not enough. I performed inference on a  1080TI which has 11GB of RAM.</p>\n\n<p>I pushed some changes to the config - not sure if this will help but worth a try.</p>",
      "rawMarkdown": "I have not tried running the webcam demo, quite surprised 8GB is not enough. I performed inference on a  1080TI which has 11GB of RAM.\n\nI pushed some changes to the config - not sure if this will help but worth a try.",
      "votes": null
    },
    {
      "id": "389130",
      "postDate": "09/18/2018 07:55:10",
      "content": "<p>Now I'm able to run it on 980M with 6fps.\nHowever, it is not able to find any objects at all.</p>\n\n<p>Command:</p>\n\n<p><code>\n../darknet detector demo yolo.data yolov3-spp.cfg yolov3-spp_final.weights\n</code>\nI downloaded your repo to darknet.\nDarknet is built with CUDA 9.0, CUDNN 7 and OpenCV 3 (needed to run demo)</p>\n\n<p>2 weeks ago I was trying to train YOLOv3 on OIDv4 by myself but had some problems...\nI've never trained darknet, but had successes using TF OD API</p>",
      "rawMarkdown": "Now I'm able to run it on 980M with 6fps.\nHowever, it is not able to find any objects at all.\n\nCommand:\n\n```\n../darknet detector demo yolo.data yolov3-spp.cfg yolov3-spp_final.weights\n```\nI downloaded your repo to darknet.\nDarknet is built with CUDA 9.0, CUDNN 7 and OpenCV 3 (needed to run demo)\n\n2 weeks ago I was trying to train YOLOv3 on OIDv4 by myself but had some problems...\nI've never trained darknet, but had successes using TF OD API",
      "votes": null
    },
    {
      "id": "398204",
      "postDate": "10/03/2018 16:55:10",
      "content": "<p>Hi. I am unable to download the weights, it says:\n\"Fetching master\nGit LFS: (0 of 1 files) 0 B / 249.15 MB\nbatch response: This repository is over its data quota. Purchase more data packs to restore access.\nerror: failed to fetch some objects from '<a href=\"https://github.com/radekosmulski/yolo_open_images.git/info/lfs\">https://github.com/radekosmulski/yolo_open_images.git/info/lfs</a>\"</p>\n\n<p>Have you also uploaded these weights somewhere else like google drive? If not, can you please do and share the link here? Thanks!</p>",
      "rawMarkdown": "Hi. I am unable to download the weights, it says:\n\"Fetching master\nGit LFS: (0 of 1 files) 0 B / 249.15 MB\nbatch response: This repository is over its data quota. Purchase more data packs to restore access.\nerror: failed to fetch some objects from 'https://github.com/radekosmulski/yolo_open_images.git/info/lfs\"\n\nHave you also uploaded these weights somewhere else like google drive? If not, can you please do and share the link here? Thanks!",
      "votes": null
    },
    {
      "id": "398465",
      "postDate": "10/04/2018 06:36:40",
      "content": "<p>Could you please try pulling the repository from <a href=\"https://bitbucket.org/dreamwalkerrr/yolo_open_images\">bitbucket</a>? I think this should work</p>",
      "rawMarkdown": "Could you please try pulling the repository from [bitbucket][1]? I think this should work\n\n  [1]: https://bitbucket.org/dreamwalkerrr/yolo_open_images",
      "votes": null
    },
    {
      "id": "527937",
      "postDate": "05/06/2019 16:35:35",
      "content": "<p>yolov3-spp is better than yolov3?</p>",
      "rawMarkdown": "yolov3-spp is better than yolov3?",
      "votes": null
    },
    {
      "id": "625559",
      "postDate": "09/13/2019 07:27:29",
      "content": "<p>tks for sharing!</p>",
      "rawMarkdown": "tks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 379977,
      "author_name": "sermakarevich",
      "author_url": "",
      "post_date": "09/01/2018 11:19:19",
      "content": "<p>Congrats Radek! Amazing achievement!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 380697,
      "author_name": "barty777",
      "author_url": "",
      "post_date": "09/03/2018 08:54:24",
      "content": "<p>Congratulations on the silver medal!</p>\n\n<p>Can you share what input resolution did you use in your model?</p>\n\n<p>EDIT: Found it on the link :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 380705,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "09/03/2018 09:22:37",
          "content": "<p>Thank you! Darknet does the rescaling on the fly for you to a couple of different resolutions to train on objects of various sizes. I trained with a max size of 608x608.</p>\n\n<p>The data I downloaded and used for training came from a storage bucket whose address I believe I have found on github in the object detection API repository, though might have been somewhere else. I don't recall now (and don't have access to my rig at the moment) but I believe the images were rescaled to either 1024 for the longer edge or were slightly bigger.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 381109,
      "author_name": "haimd1801",
      "author_url": "",
      "post_date": "09/04/2018 04:25:30",
      "content": "<p>Congrats Radek!! Tks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 388814,
      "author_name": "karolmajek",
      "author_url": "",
      "post_date": "09/17/2018 16:26:08",
      "content": "<p>Congrats!\nHow much GPU mem is required to run inference?\nI have OOM with 8GB RAM.\nIs 12GB enough to run it? (I am trying to run webcam demo)</p>",
      "votes": null,
      "replies": [
        {
          "id": 389079,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "09/18/2018 05:41:29",
          "content": "<p>I have not tried running the webcam demo, quite surprised 8GB is not enough. I performed inference on a  1080TI which has 11GB of RAM.</p>\n\n<p>I pushed some changes to the config - not sure if this will help but worth a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 389130,
          "author_name": "karolmajek",
          "author_url": "",
          "post_date": "09/18/2018 07:55:10",
          "content": "<p>Now I'm able to run it on 980M with 6fps.\nHowever, it is not able to find any objects at all.</p>\n\n<p>Command:</p>\n\n<p><code>\n../darknet detector demo yolo.data yolov3-spp.cfg yolov3-spp_final.weights\n</code>\nI downloaded your repo to darknet.\nDarknet is built with CUDA 9.0, CUDNN 7 and OpenCV 3 (needed to run demo)</p>\n\n<p>2 weeks ago I was trying to train YOLOv3 on OIDv4 by myself but had some problems...\nI've never trained darknet, but had successes using TF OD API</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 398204,
      "author_name": "jigsawcoder",
      "author_url": "",
      "post_date": "10/03/2018 16:55:10",
      "content": "<p>Hi. I am unable to download the weights, it says:\n\"Fetching master\nGit LFS: (0 of 1 files) 0 B / 249.15 MB\nbatch response: This repository is over its data quota. Purchase more data packs to restore access.\nerror: failed to fetch some objects from '<a href=\"https://github.com/radekosmulski/yolo_open_images.git/info/lfs\">https://github.com/radekosmulski/yolo_open_images.git/info/lfs</a>\"</p>\n\n<p>Have you also uploaded these weights somewhere else like google drive? If not, can you please do and share the link here? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 398465,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "10/04/2018 06:36:40",
          "content": "<p>Could you please try pulling the repository from <a href=\"https://bitbucket.org/dreamwalkerrr/yolo_open_images\">bitbucket</a>? I think this should work</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 527937,
      "author_name": "hqwsky",
      "author_url": "",
      "post_date": "05/06/2019 16:35:35",
      "content": "<p>yolov3-spp is better than yolov3?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 625559,
      "author_name": "fyq1999",
      "author_url": "",
      "post_date": "09/13/2019 07:27:29",
      "content": "<p>tks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "379875": "I trained [yolov3][1] with SPP using [darknet53.conv.74](https://pjreddie.com/media/files/darknet53.conv.74) weights pretrained on Imagenet.\n\nI followed the training regime and used the same architecture that the creators of yolo used for training on COCO. Most likely better results could be achieved on the Open Images dataset should the capacity of the model be increased.\n\nI emulated the conditions of training on COCO and created two custom datasets. I initially (up to ~ 250k batches IIRC) trained on ~320k files, increasing the dataset to ~430k files onwards. I picked the files to train on to ensure as balanced a representation of each class as possible (picked all images that contained rare labels - limited number of images to include for more popular labels to arrive at the target file count).\n\nI trained the model thx to the GCP coupon using a preemptible instance / premptible GPUs. I went with 16 vCPUs and 4 preemptible P100s. I didn't measure this precisely but I think in 24 hrs the model trained on ~50k batches, maybe a bit more. The full training schedule is just over 500k batches.\n\nI believe that appropriate makeup of the trainset is key to achieving a good result. This is an observation that I have not verified though.\n\nI have not modified the train set labels in any way. I also don't believe I have expanded the labels before training to include parent categories but might be wrong on that one. Doing something reasonable with the label hierarchy should grant further improvement in performance.\n\nMy private LB score is slightly different as I tried training on a subset of labels / files for trainingtrain set selected in a specific way to help me in the VRD track but I didn't get very far (partially due to this likely being the wrong approach and partially due to running out of time).\n\nI am sharing the final weights and the config files in this github [repository](https://github.com/radekosmulski/yolo_open_images).\n\nHere is the citation of the paper that introduces yolov3:\n\n&gt; @article{yolov3,   title={YOLOv3: An Incremental Improvement},  \n&gt; author={Redmon, Joseph and Farhadi, Ali},   journal = {arXiv},  \n&gt; year={2018} }\n\nAnd [here][2] is where you can find the paper on arxiv.\n\n  [1]: https://pjreddie.com/darknet/yolo/\n  [2]: https://arxiv.org/abs/1804.02767",
    "379977": "Congrats Radek! Amazing achievement!",
    "380697": "Congratulations on the silver medal!\n\nCan you share what input resolution did you use in your model?\n\nEDIT: Found it on the link :)",
    "380705": "Thank you! Darknet does the rescaling on the fly for you to a couple of different resolutions to train on objects of various sizes. I trained with a max size of 608x608.\n\nThe data I downloaded and used for training came from a storage bucket whose address I believe I have found on github in the object detection API repository, though might have been somewhere else. I don't recall now (and don't have access to my rig at the moment) but I believe the images were rescaled to either 1024 for the longer edge or were slightly bigger.",
    "381109": "Congrats Radek!! Tks for sharing.",
    "388814": "Congrats!\nHow much GPU mem is required to run inference?\nI have OOM with 8GB RAM.\nIs 12GB enough to run it? (I am trying to run webcam demo)",
    "389079": "I have not tried running the webcam demo, quite surprised 8GB is not enough. I performed inference on a  1080TI which has 11GB of RAM.\n\nI pushed some changes to the config - not sure if this will help but worth a try.",
    "389130": "Now I'm able to run it on 980M with 6fps.\nHowever, it is not able to find any objects at all.\n\nCommand:\n\n```\n../darknet detector demo yolo.data yolov3-spp.cfg yolov3-spp_final.weights\n```\nI downloaded your repo to darknet.\nDarknet is built with CUDA 9.0, CUDNN 7 and OpenCV 3 (needed to run demo)\n\n2 weeks ago I was trying to train YOLOv3 on OIDv4 by myself but had some problems...\nI've never trained darknet, but had successes using TF OD API",
    "398204": "Hi. I am unable to download the weights, it says:\n\"Fetching master\nGit LFS: (0 of 1 files) 0 B / 249.15 MB\nbatch response: This repository is over its data quota. Purchase more data packs to restore access.\nerror: failed to fetch some objects from 'https://github.com/radekosmulski/yolo_open_images.git/info/lfs\"\n\nHave you also uploaded these weights somewhere else like google drive? If not, can you please do and share the link here? Thanks!",
    "398465": "Could you please try pulling the repository from [bitbucket][1]? I think this should work\n\n  [1]: https://bitbucket.org/dreamwalkerrr/yolo_open_images",
    "527937": "yolov3-spp is better than yolov3?",
    "625559": "tks for sharing!"
  },
  "source": "meta"
}