{
  "id": 35150,
  "title": "huge submission set and computational times",
  "url": "/competitions/noaa-fisheries-steller-sea-lion-population-count/discussion/35150",
  "author_name": "",
  "post_date": "2017-06-22T22:16:01.311367600Z",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I understand that the competition sponsors don't want participants cheating on the test images by say, manually counting sea lions.  Therefore, the test set is large.  However, I am experiencing over 24 hours of compute time to generate one submission.csv.</p>\n\n<p>There are plenty of submissions on the leader board, so I was wondering if anyone else is experiencing this issue?</p>",
  "messages": [
    {
      "id": "195181",
      "postDate": "06/22/2017 22:16:01",
      "content": "<p>I understand that the competition sponsors don't want participants cheating on the test images by say, manually counting sea lions.  Therefore, the test set is large.  However, I am experiencing over 24 hours of compute time to generate one submission.csv.</p>\n\n<p>There are plenty of submissions on the leader board, so I was wondering if anyone else is experiencing this issue?</p>",
      "rawMarkdown": "I understand that the competition sponsors don't want participants cheating on the test images by say, manually counting sea lions.  Therefore, the test set is large.  However, I am experiencing over 24 hours of compute time to generate one submission.csv.\n\nThere are plenty of submissions on the leader board, so I was wondering if anyone else is experiencing this issue?",
      "votes": null
    },
    {
      "id": "195185",
      "postDate": "06/22/2017 22:31:08",
      "content": "<p>Yeah, huge submission time is a major issue for me too. Right now it takes about 12 hours to generate the full submission using one GPU. And I could definitely improve it a bit by taking say 4x longer to predict :)</p>\n\n<p>The way to ~~waste~~ make many submissions is to have a multi-stage pipeline where you tweak just the last step which is faster. Say you predict bounding boxes or segmentation masks or heatmaps, and then as the last step predict lion counts from those. So the last step takes less time, and you can make more submissions tweaking the last step.</p>",
      "rawMarkdown": "Yeah, huge submission time is a major issue for me too. Right now it takes about 12 hours to generate the full submission using one GPU. And I could definitely improve it a bit by taking say 4x longer to predict :)\n\nThe way to ~~waste~~ make many submissions is to have a multi-stage pipeline where you tweak just the last step which is faster. Say you predict bounding boxes or segmentation masks or heatmaps, and then as the last step predict lion counts from those. So the last step takes less time, and you can make more submissions tweaking the last step.",
      "votes": null
    },
    {
      "id": "195189",
      "postDate": "06/22/2017 22:48:56",
      "content": "<p>That's a good suggestion. I can see it working if the outputs from each step in the pipeline can be stored.  Thanks for the tip!    </p>",
      "rawMarkdown": "That's a good suggestion. I can see it working if the outputs from each step in the pipeline can be stored.  Thanks for the tip!",
      "votes": null
    },
    {
      "id": "195191",
      "postDate": "06/22/2017 22:53:04",
      "content": "<p>The same for me. I am predicting on 1024x1024 tiles and thus, it takes around 6-7 seconds per image. Overall 30-35 hours are spent on an AWS spot instance. With 3 instances I have to wait around 11 hours. Still waiting for my first and final result :) I will do this just once - It takes too much time.</p>",
      "rawMarkdown": "The same for me. I am predicting on 1024x1024 tiles and thus, it takes around 6-7 seconds per image. Overall 30-35 hours are spent on an AWS spot instance. With 3 instances I have to wait around 11 hours. Still waiting for my first and final result :) I will do this just once - It takes too much time.",
      "votes": null
    },
    {
      "id": "195211",
      "postDate": "06/23/2017 00:01:24",
      "content": "<p>As pointed out by @toshi_k in <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/34671\">this thread</a> if you are using tile level based prediction you can use some binary classifier with high confidence(tune it to have no false negatives) to filter out areas with water, ground, etc. it will reduce computations of 'counting' classifier.</p>",
      "rawMarkdown": "As pointed out by @toshi_k in [this thread][1] if you are using tile level based prediction you can use some binary classifier with high confidence(tune it to have no false negatives) to filter out areas with water, ground, etc. it will reduce computations of 'counting' classifier.\n\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/34671",
      "votes": null
    },
    {
      "id": "195351",
      "postDate": "06/23/2017 10:49:26",
      "content": "<p>I would have liked to do that, but unfortunately I didn't have time for it. I've just fine-tuned a VGG and trained it on 1024x1024 tiles that included at least 3 lions. And finally, test-set prediction has finished :)</p>",
      "rawMarkdown": "I would have liked to do that, but unfortunately I didn't have time for it. I've just fine-tuned a VGG and trained it on 1024x1024 tiles that included at least 3 lions. And finally, test-set prediction has finished :)",
      "votes": null
    },
    {
      "id": "195436",
      "postDate": "06/23/2017 15:48:55",
      "content": "<p>I've tried AWS in the past, but I didn't think the GPU instances were very good because each GPU has only 4GB memory.  Is this one of the reasons you are having to break-up images down to 1024x1024 tiles?</p>",
      "rawMarkdown": "I've tried AWS in the past, but I didn't think the GPU instances were very good because each GPU has only 4GB memory.  Is this one of the reasons you are having to break-up images down to 1024x1024 tiles?",
      "votes": null
    },
    {
      "id": "195532",
      "postDate": "06/23/2017 22:14:39",
      "content": "<p>Anyone else is experiencing this issue</p>",
      "rawMarkdown": "Anyone else is experiencing this issue",
      "votes": null
    },
    {
      "id": "195533",
      "postDate": "06/23/2017 22:29:35",
      "content": "<p>This is by design to avoid cheating, but there are tricks to iterate faster. Keep in mind the following:</p>\n\n<ul>\n<li><p>over 80% of the total area do <strong>not</strong> contain any animals. If you ran a less accurate model over all the test dataset before you already have some idea of which areas have higher density or not. You can use this information to your advantage.</p>\n\n<ul><li>given the metric, crowded areas contributes more to RMSE</li></ul></li>\n</ul>",
      "rawMarkdown": "This is by design to avoid cheating, but there are tricks to iterate faster. Keep in mind the following:\n\n - over 80% of the total area do **not** contain any animals. If you ran a less accurate model over all the test dataset before you already have some idea of which areas have higher density or not. You can use this information to your advantage.\n\n- given the metric, crowded areas contributes more to RMSE",
      "votes": null
    },
    {
      "id": "195538",
      "postDate": "06/23/2017 23:19:09",
      "content": "<p>For GPU computing you can choose one of three instances:</p>\n\n<ul>\n<li>p2.xlarge: 4 vCPUs, 61 GB RAM, 1 GPU</li>\n<li>p2.8xlarge: 32 vCPUs, 488 GB RAM, 8 GPUs</li>\n<li>p2.16xlarge: 64 vCPUs,732 GB RAM, 16 GPUs</li>\n</ul>\n\n<p>The GPU is an NVIDIA K80 with 12 GB Memory. The amount of used memory is not only depended on the image input size but also on the complexity of the model (VGG16/U-net in my case). I ended up using 1024x1024 for p2.xlarge. I was getting out-of-memory for bigger tiles.</p>",
      "rawMarkdown": "For GPU computing you can choose one of three instances:\n\n - p2.xlarge: 4 vCPUs, 61 GB RAM, 1 GPU\n - p2.8xlarge: 32 vCPUs, 488 GB RAM, 8 GPUs\n - p2.16xlarge: 64 vCPUs,732 GB RAM, 16 GPUs\n\nThe GPU is an NVIDIA K80 with 12 GB Memory. The amount of used memory is not only depended on the image input size but also on the complexity of the model (VGG16/U-net in my case). I ended up using 1024x1024 for p2.xlarge. I was getting out-of-memory for bigger tiles.",
      "votes": null
    },
    {
      "id": "195556",
      "postDate": "06/24/2017 01:10:53",
      "content": "<p>This is good advice. I am doing a 4-5 day sprint on this to see if I can get top50ish, and every decision has to take compute time into account. These two ideas are good ones. </p>",
      "rawMarkdown": "This is good advice. I am doing a 4-5 day sprint on this to see if I can get top50ish, and every decision has to take compute time into account. These two ideas are good ones.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 195185,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "06/22/2017 22:31:08",
      "content": "<p>Yeah, huge submission time is a major issue for me too. Right now it takes about 12 hours to generate the full submission using one GPU. And I could definitely improve it a bit by taking say 4x longer to predict :)</p>\n\n<p>The way to ~~waste~~ make many submissions is to have a multi-stage pipeline where you tweak just the last step which is faster. Say you predict bounding boxes or segmentation masks or heatmaps, and then as the last step predict lion counts from those. So the last step takes less time, and you can make more submissions tweaking the last step.</p>",
      "votes": null,
      "replies": [
        {
          "id": 195189,
          "author_name": "jeffalltogether",
          "author_url": "",
          "post_date": "06/22/2017 22:48:56",
          "content": "<p>That's a good suggestion. I can see it working if the outputs from each step in the pipeline can be stored.  Thanks for the tip!    </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 195191,
      "author_name": "firolino",
      "author_url": "",
      "post_date": "06/22/2017 22:53:04",
      "content": "<p>The same for me. I am predicting on 1024x1024 tiles and thus, it takes around 6-7 seconds per image. Overall 30-35 hours are spent on an AWS spot instance. With 3 instances I have to wait around 11 hours. Still waiting for my first and final result :) I will do this just once - It takes too much time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 195436,
          "author_name": "jeffalltogether",
          "author_url": "",
          "post_date": "06/23/2017 15:48:55",
          "content": "<p>I've tried AWS in the past, but I didn't think the GPU instances were very good because each GPU has only 4GB memory.  Is this one of the reasons you are having to break-up images down to 1024x1024 tiles?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195538,
          "author_name": "firolino",
          "author_url": "",
          "post_date": "06/23/2017 23:19:09",
          "content": "<p>For GPU computing you can choose one of three instances:</p>\n\n<ul>\n<li>p2.xlarge: 4 vCPUs, 61 GB RAM, 1 GPU</li>\n<li>p2.8xlarge: 32 vCPUs, 488 GB RAM, 8 GPUs</li>\n<li>p2.16xlarge: 64 vCPUs,732 GB RAM, 16 GPUs</li>\n</ul>\n\n<p>The GPU is an NVIDIA K80 with 12 GB Memory. The amount of used memory is not only depended on the image input size but also on the complexity of the model (VGG16/U-net in my case). I ended up using 1024x1024 for p2.xlarge. I was getting out-of-memory for bigger tiles.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 195211,
      "author_name": "mrgloom",
      "author_url": "",
      "post_date": "06/23/2017 00:01:24",
      "content": "<p>As pointed out by @toshi_k in <a href=\"https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/34671\">this thread</a> if you are using tile level based prediction you can use some binary classifier with high confidence(tune it to have no false negatives) to filter out areas with water, ground, etc. it will reduce computations of 'counting' classifier.</p>",
      "votes": null,
      "replies": [
        {
          "id": 195351,
          "author_name": "firolino",
          "author_url": "",
          "post_date": "06/23/2017 10:49:26",
          "content": "<p>I would have liked to do that, but unfortunately I didn't have time for it. I've just fine-tuned a VGG and trained it on 1024x1024 tiles that included at least 3 lions. And finally, test-set prediction has finished :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 195532,
      "author_name": "sparamonov",
      "author_url": "",
      "post_date": "06/23/2017 22:14:39",
      "content": "<p>Anyone else is experiencing this issue</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 195533,
      "author_name": "aamaia",
      "author_url": "",
      "post_date": "06/23/2017 22:29:35",
      "content": "<p>This is by design to avoid cheating, but there are tricks to iterate faster. Keep in mind the following:</p>\n\n<ul>\n<li><p>over 80% of the total area do <strong>not</strong> contain any animals. If you ran a less accurate model over all the test dataset before you already have some idea of which areas have higher density or not. You can use this information to your advantage.</p>\n\n<ul><li>given the metric, crowded areas contributes more to RMSE</li></ul></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 195556,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "06/24/2017 01:10:53",
          "content": "<p>This is good advice. I am doing a 4-5 day sprint on this to see if I can get top50ish, and every decision has to take compute time into account. These two ideas are good ones. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "195181": "I understand that the competition sponsors don't want participants cheating on the test images by say, manually counting sea lions.  Therefore, the test set is large.  However, I am experiencing over 24 hours of compute time to generate one submission.csv.\n\nThere are plenty of submissions on the leader board, so I was wondering if anyone else is experiencing this issue?",
    "195185": "Yeah, huge submission time is a major issue for me too. Right now it takes about 12 hours to generate the full submission using one GPU. And I could definitely improve it a bit by taking say 4x longer to predict :)\n\nThe way to ~~waste~~ make many submissions is to have a multi-stage pipeline where you tweak just the last step which is faster. Say you predict bounding boxes or segmentation masks or heatmaps, and then as the last step predict lion counts from those. So the last step takes less time, and you can make more submissions tweaking the last step.",
    "195189": "That's a good suggestion. I can see it working if the outputs from each step in the pipeline can be stored.  Thanks for the tip!",
    "195191": "The same for me. I am predicting on 1024x1024 tiles and thus, it takes around 6-7 seconds per image. Overall 30-35 hours are spent on an AWS spot instance. With 3 instances I have to wait around 11 hours. Still waiting for my first and final result :) I will do this just once - It takes too much time.",
    "195211": "As pointed out by @toshi_k in [this thread][1] if you are using tile level based prediction you can use some binary classifier with high confidence(tune it to have no false negatives) to filter out areas with water, ground, etc. it will reduce computations of 'counting' classifier.\n\n\n  [1]: https://www.kaggle.com/c/noaa-fisheries-steller-sea-lion-population-count/discussion/34671",
    "195351": "I would have liked to do that, but unfortunately I didn't have time for it. I've just fine-tuned a VGG and trained it on 1024x1024 tiles that included at least 3 lions. And finally, test-set prediction has finished :)",
    "195436": "I've tried AWS in the past, but I didn't think the GPU instances were very good because each GPU has only 4GB memory.  Is this one of the reasons you are having to break-up images down to 1024x1024 tiles?",
    "195532": "Anyone else is experiencing this issue",
    "195533": "This is by design to avoid cheating, but there are tricks to iterate faster. Keep in mind the following:\n\n - over 80% of the total area do **not** contain any animals. If you ran a less accurate model over all the test dataset before you already have some idea of which areas have higher density or not. You can use this information to your advantage.\n\n- given the metric, crowded areas contributes more to RMSE",
    "195538": "For GPU computing you can choose one of three instances:\n\n - p2.xlarge: 4 vCPUs, 61 GB RAM, 1 GPU\n - p2.8xlarge: 32 vCPUs, 488 GB RAM, 8 GPUs\n - p2.16xlarge: 64 vCPUs,732 GB RAM, 16 GPUs\n\nThe GPU is an NVIDIA K80 with 12 GB Memory. The amount of used memory is not only depended on the image input size but also on the complexity of the model (VGG16/U-net in my case). I ended up using 1024x1024 for p2.xlarge. I was getting out-of-memory for bigger tiles.",
    "195556": "This is good advice. I am doing a 4-5 day sprint on this to see if I can get top50ish, and every decision has to take compute time into account. These two ideas are good ones."
  },
  "source": "meta"
}