{
  "id": 63547,
  "title": "Train Setup & Time & Result",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/63547",
  "author_name": "",
  "post_date": "2018-08-18T02:54:16.920584100Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>This is my first competition and I struggled to preprocess data and acquire right setup initially, but I finally got everything squared away. I'm now just waiting my model to train using Yolo V3 on Darknet on Amazon p3.2xlarge instance. It seems like I will need days of training to achieve good results....</p>\n\n<p>Since there are so many different possible setup, I was curious if anyone would be interested in sharing the following : 1) Hardware 2) Model 3) Training Time 4) Achieved mAP.\nIf I had alot of time, I would have tried myself, but wasted a lot of time to even get to training my model properly.</p>\n\n<p>Thank you for anyone who's sharing here! :) </p>",
  "messages": [
    {
      "id": "371993",
      "postDate": "08/18/2018 02:54:16",
      "content": "<p>This is my first competition and I struggled to preprocess data and acquire right setup initially, but I finally got everything squared away. I'm now just waiting my model to train using Yolo V3 on Darknet on Amazon p3.2xlarge instance. It seems like I will need days of training to achieve good results....</p>\n\n<p>Since there are so many different possible setup, I was curious if anyone would be interested in sharing the following : 1) Hardware 2) Model 3) Training Time 4) Achieved mAP.\nIf I had alot of time, I would have tried myself, but wasted a lot of time to even get to training my model properly.</p>\n\n<p>Thank you for anyone who's sharing here! :) </p>",
      "rawMarkdown": "This is my first competition and I struggled to preprocess data and acquire right setup initially, but I finally got everything squared away. I'm now just waiting my model to train using Yolo V3 on Darknet on Amazon p3.2xlarge instance. It seems like I will need days of training to achieve good results....\n\nSince there are so many different possible setup, I was curious if anyone would be interested in sharing the following : 1) Hardware 2) Model 3) Training Time 4) Achieved mAP.\nIf I had alot of time, I would have tried myself, but wasted a lot of time to even get to training my model properly.\n\nThank you for anyone who's sharing here! :)",
      "votes": null
    },
    {
      "id": "371996",
      "postDate": "08/18/2018 03:09:44",
      "content": "<p>My first one as well.... Yolov3 did not converge for me. I have trained it for two days. At first the model converged nicely but after about a day loss was fluctuating around 1.6 if memory serves and testing the checkpoints showed unremarkable results, to say the least.\nNaturally, it might have been my fault completely (took me a while to get it going as well... Its not a very user friendly environment), so, ymmv... </p>",
      "rawMarkdown": "My first one as well.... Yolov3 did not converge for me. I have trained it for two days. At first the model converged nicely but after about a day loss was fluctuating around 1.6 if memory serves and testing the checkpoints showed unremarkable results, to say the least.\nNaturally, it might have been my fault completely (took me a while to get it going as well... Its not a very user friendly environment), so, ymmv...",
      "votes": null
    },
    {
      "id": "371997",
      "postDate": "08/18/2018 03:19:25",
      "content": "<p>I wasted about a week wondering why it's not improving, then I figured out I got y_min wrong.... Yolo V3's zero coordinate is left top, not left bottom. Since I corrected this issue, this is my mAP&amp;IoU vs iterations looks like.</p>\n\n<p><a href=\"https://www.flickr.com/photos/143579366@N02/44103321291/in/dateposted-public/\">https://www.flickr.com/photos/143579366@N02/44103321291/in/dateposted-public/</a></p>\n\n<p>It looks like I still have long way to train it...</p>",
      "rawMarkdown": "I wasted about a week wondering why it's not improving, then I figured out I got y_min wrong.... Yolo V3's zero coordinate is left top, not left bottom. Since I corrected this issue, this is my mAP&amp;IoU vs iterations looks like.\n\nhttps://www.flickr.com/photos/143579366@N02/44103321291/in/dateposted-public/\n\nIt looks like I still have long way to train it...",
      "votes": null
    },
    {
      "id": "372000",
      "postDate": "08/18/2018 03:31:39",
      "content": "<p>That might be it! Good luck! Its more art then science with darknet. I found the alex fork a bit more approachable but wasted a few good days and Google credit on it</p>",
      "rawMarkdown": "That might be it! Good luck! Its more art then science with darknet. I found the alex fork a bit more approachable but wasted a few good days and Google credit on it",
      "votes": null
    },
    {
      "id": "372001",
      "postDate": "08/18/2018 03:33:04",
      "content": "<p>Btw, there is a single tweet by the author of yolo about training on oid going well, and silence ever since</p>",
      "rawMarkdown": "Btw, there is a single tweet by the author of yolo about training on oid going well, and silence ever since",
      "votes": null
    },
    {
      "id": "372003",
      "postDate": "08/18/2018 03:43:45",
      "content": "<p>Thanks Moshel for the input, I will check his tweet.\nAnd grats on your current standing in the leaderboard! Do you mind me asking how long it took you to train your model to achieve 0.39491 mAP?</p>",
      "rawMarkdown": "Thanks Moshel for the input, I will check his tweet.\nAnd grats on your current standing in the leaderboard! Do you mind me asking how long it took you to train your model to achieve 0.39491 mAP?",
      "votes": null
    },
    {
      "id": "372023",
      "postDate": "08/18/2018 05:02:36",
      "content": "<p>It's actually an ensamble of different models. I have trained retinanet, it took a few days to converge on a massive machine... (Thanks Google for the credits!). By itself it achieved about 0.27 iirc. It is very difficult to train properly on this set because of its size and the number of classes. </p>",
      "rawMarkdown": "It's actually an ensamble of different models. I have trained retinanet, it took a few days to converge on a massive machine... (Thanks Google for the credits!). By itself it achieved about 0.27 iirc. It is very difficult to train properly on this set because of its size and the number of classes.",
      "votes": null
    },
    {
      "id": "372025",
      "postDate": "08/18/2018 05:10:49",
      "content": "<p>Thank you for sharing your experience! Good luck :)</p>",
      "rawMarkdown": "Thank you for sharing your experience! Good luck :)",
      "votes": null
    },
    {
      "id": "372372",
      "postDate": "08/19/2018 06:20:41",
      "content": "<p>Hi. I struggle to train this size of dataset. I train with SSD Mxnet on google cloud 2 P100 GPUs, for a subset of 50,000 images. Yet it takes me a week to achieve some 60% accuracy.  Please tell me your hardware configuration. Thank you.</p>",
      "rawMarkdown": "Hi. I struggle to train this size of dataset. I train with SSD Mxnet on google cloud 2 P100 GPUs, for a subset of 50,000 images. Yet it takes me a week to achieve some 60% accuracy.  Please tell me your hardware configuration. Thank you.",
      "votes": null
    },
    {
      "id": "372398",
      "postDate": "08/19/2018 08:35:19",
      "content": "<p>60% is amazing but i am afraid you are over fitting. 50000 images is very small for 500 classes even with aug.\nMy machine was 1 p100 with 48gb of memory iirc. The drive was ssd. It's important with so many images. </p>",
      "rawMarkdown": "60% is amazing but i am afraid you are over fitting. 50000 images is very small for 500 classes even with aug.\nMy machine was 1 p100 with 48gb of memory iirc. The drive was ssd. It's important with so many images.",
      "votes": null
    },
    {
      "id": "372408",
      "postDate": "08/19/2018 09:16:30",
      "content": "<p>Thank you very much. I use Standard Persistent disk. Perhaps that's why I cannot train that fast?  I even use 2 p100 , but my RAM only 12 GB. Is it because of RAM?</p>\n\n<p>Let me clarify my 60% accuracy a bit. I am only using 4 classes of 50,000 images. My MAP of these 4 classes is 50%. The best class got 60%. I am doing this small because, you know, the training is slow.</p>",
      "rawMarkdown": "Thank you very much. I use Standard Persistent disk. Perhaps that's why I cannot train that fast?  I even use 2 p100 , but my RAM only 12 GB. Is it because of RAM?\n\nLet me clarify my 60% accuracy a bit. I am only using 4 classes of 50,000 images. My MAP of these 4 classes is 50%. The best class got 60%. I am doing this small because, you know, the training is slow.",
      "votes": null
    },
    {
      "id": "372425",
      "postDate": "08/19/2018 10:33:32",
      "content": "<p>I am not an expert in this, but i would say that 2 p100 are probably mostly idle if you cant feed them the data fast enough. So, i would go with ssd and much more memory. Deep learning uses up a lot of memory! </p>",
      "rawMarkdown": "I am not an expert in this, but i would say that 2 p100 are probably mostly idle if you cant feed them the data fast enough. So, i would go with ssd and much more memory. Deep learning uses up a lot of memory!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 371996,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "08/18/2018 03:09:44",
      "content": "<p>My first one as well.... Yolov3 did not converge for me. I have trained it for two days. At first the model converged nicely but after about a day loss was fluctuating around 1.6 if memory serves and testing the checkpoints showed unremarkable results, to say the least.\nNaturally, it might have been my fault completely (took me a while to get it going as well... Its not a very user friendly environment), so, ymmv... </p>",
      "votes": null,
      "replies": [
        {
          "id": 371997,
          "author_name": "silvernine",
          "author_url": "",
          "post_date": "08/18/2018 03:19:25",
          "content": "<p>I wasted about a week wondering why it's not improving, then I figured out I got y_min wrong.... Yolo V3's zero coordinate is left top, not left bottom. Since I corrected this issue, this is my mAP&amp;IoU vs iterations looks like.</p>\n\n<p><a href=\"https://www.flickr.com/photos/143579366@N02/44103321291/in/dateposted-public/\">https://www.flickr.com/photos/143579366@N02/44103321291/in/dateposted-public/</a></p>\n\n<p>It looks like I still have long way to train it...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372000,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "08/18/2018 03:31:39",
          "content": "<p>That might be it! Good luck! Its more art then science with darknet. I found the alex fork a bit more approachable but wasted a few good days and Google credit on it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372001,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "08/18/2018 03:33:04",
          "content": "<p>Btw, there is a single tweet by the author of yolo about training on oid going well, and silence ever since</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372003,
          "author_name": "silvernine",
          "author_url": "",
          "post_date": "08/18/2018 03:43:45",
          "content": "<p>Thanks Moshel for the input, I will check his tweet.\nAnd grats on your current standing in the leaderboard! Do you mind me asking how long it took you to train your model to achieve 0.39491 mAP?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372023,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "08/18/2018 05:02:36",
          "content": "<p>It's actually an ensamble of different models. I have trained retinanet, it took a few days to converge on a massive machine... (Thanks Google for the credits!). By itself it achieved about 0.27 iirc. It is very difficult to train properly on this set because of its size and the number of classes. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372025,
          "author_name": "silvernine",
          "author_url": "",
          "post_date": "08/18/2018 05:10:49",
          "content": "<p>Thank you for sharing your experience! Good luck :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 372372,
      "author_name": "peerajak",
      "author_url": "",
      "post_date": "08/19/2018 06:20:41",
      "content": "<p>Hi. I struggle to train this size of dataset. I train with SSD Mxnet on google cloud 2 P100 GPUs, for a subset of 50,000 images. Yet it takes me a week to achieve some 60% accuracy.  Please tell me your hardware configuration. Thank you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 372398,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "08/19/2018 08:35:19",
          "content": "<p>60% is amazing but i am afraid you are over fitting. 50000 images is very small for 500 classes even with aug.\nMy machine was 1 p100 with 48gb of memory iirc. The drive was ssd. It's important with so many images. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372408,
          "author_name": "peerajak",
          "author_url": "",
          "post_date": "08/19/2018 09:16:30",
          "content": "<p>Thank you very much. I use Standard Persistent disk. Perhaps that's why I cannot train that fast?  I even use 2 p100 , but my RAM only 12 GB. Is it because of RAM?</p>\n\n<p>Let me clarify my 60% accuracy a bit. I am only using 4 classes of 50,000 images. My MAP of these 4 classes is 50%. The best class got 60%. I am doing this small because, you know, the training is slow.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 372425,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "08/19/2018 10:33:32",
          "content": "<p>I am not an expert in this, but i would say that 2 p100 are probably mostly idle if you cant feed them the data fast enough. So, i would go with ssd and much more memory. Deep learning uses up a lot of memory! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "371993": "This is my first competition and I struggled to preprocess data and acquire right setup initially, but I finally got everything squared away. I'm now just waiting my model to train using Yolo V3 on Darknet on Amazon p3.2xlarge instance. It seems like I will need days of training to achieve good results....\n\nSince there are so many different possible setup, I was curious if anyone would be interested in sharing the following : 1) Hardware 2) Model 3) Training Time 4) Achieved mAP.\nIf I had alot of time, I would have tried myself, but wasted a lot of time to even get to training my model properly.\n\nThank you for anyone who's sharing here! :)",
    "371996": "My first one as well.... Yolov3 did not converge for me. I have trained it for two days. At first the model converged nicely but after about a day loss was fluctuating around 1.6 if memory serves and testing the checkpoints showed unremarkable results, to say the least.\nNaturally, it might have been my fault completely (took me a while to get it going as well... Its not a very user friendly environment), so, ymmv...",
    "371997": "I wasted about a week wondering why it's not improving, then I figured out I got y_min wrong.... Yolo V3's zero coordinate is left top, not left bottom. Since I corrected this issue, this is my mAP&amp;IoU vs iterations looks like.\n\nhttps://www.flickr.com/photos/143579366@N02/44103321291/in/dateposted-public/\n\nIt looks like I still have long way to train it...",
    "372000": "That might be it! Good luck! Its more art then science with darknet. I found the alex fork a bit more approachable but wasted a few good days and Google credit on it",
    "372001": "Btw, there is a single tweet by the author of yolo about training on oid going well, and silence ever since",
    "372003": "Thanks Moshel for the input, I will check his tweet.\nAnd grats on your current standing in the leaderboard! Do you mind me asking how long it took you to train your model to achieve 0.39491 mAP?",
    "372023": "It's actually an ensamble of different models. I have trained retinanet, it took a few days to converge on a massive machine... (Thanks Google for the credits!). By itself it achieved about 0.27 iirc. It is very difficult to train properly on this set because of its size and the number of classes.",
    "372025": "Thank you for sharing your experience! Good luck :)",
    "372372": "Hi. I struggle to train this size of dataset. I train with SSD Mxnet on google cloud 2 P100 GPUs, for a subset of 50,000 images. Yet it takes me a week to achieve some 60% accuracy.  Please tell me your hardware configuration. Thank you.",
    "372398": "60% is amazing but i am afraid you are over fitting. 50000 images is very small for 500 classes even with aug.\nMy machine was 1 p100 with 48gb of memory iirc. The drive was ssd. It's important with so many images.",
    "372408": "Thank you very much. I use Standard Persistent disk. Perhaps that's why I cannot train that fast?  I even use 2 p100 , but my RAM only 12 GB. Is it because of RAM?\n\nLet me clarify my 60% accuracy a bit. I am only using 4 classes of 50,000 images. My MAP of these 4 classes is 50%. The best class got 60%. I am doing this small because, you know, the training is slow.",
    "372425": "I am not an expert in this, but i would say that 2 p100 are probably mostly idle if you cant feed them the data fast enough. So, i would go with ssd and much more memory. Deep learning uses up a lot of memory!"
  },
  "source": "meta"
}