{
  "id": 74443,
  "title": "1st place solution for Algorithm Speed Prize",
  "url": "/competitions/airbus-ship-detection/writeups/ods-ai-alex-sologub-1st-place-solution-for-algorit",
  "author_name": "",
  "post_date": "2018-12-12T10:10:40.583Z",
  "votes": 41,
  "comment_count": 22,
  "views": 0,
  "content": "<p>Hi there!\nHere's a quick breakdown of the 1st place solution:</p>\n\n<ol>\n<li>PyTorch.</li>\n<li>SE-ResNet50 as a classifier: took around 30 seconds to infer all images.</li>\n<li>LinkNet as a segmentation network: took around 140 seconds to infer all positively classified images.</li>\n<li>Extracting ship instances from binary mask via scipy.ndimage. Ignore instances with small area (less than 80px): took around 40 seconds.</li>\n<li>Didn't use TTA.</li>\n</ol>\n\n<p>Initial profiling has shown that I/O is the slowest part of the pipeline, so I've spent at least a week trying to optimize it. However, after a some rule clarifications, it was clear that I/O should be ignored completely, so I switched to optimizing the networks. After a bunch of experiments, I ended up with a slightly modified LinkNet, which was at least x10 faster than my model from the first stage of the competition.</p>\n\n<h1>Key insights:</h1>\n\n<ol>\n<li>Adapt to the kernel. GPU kernel has 2 very slow CPU cores. I won 1 minute of inference time just by transferring image normalization to the GPU.</li>\n<li>Classifier speed and accuracy is critical: it saves a lot of time, since segmentation network is much slower.</li>\n<li>Predicting ship borders and/or using watershed requires too much post-processing.</li>\n</ol>\n\n<h1>Training the classifier:</h1>\n\n<ol>\n<li>Train and predict on resized 224x224 images.</li>\n<li>Nesterov SGD with LR 0.001, batch size 16, weight decay 1e-3, momentum 0.9.</li>\n<li>BCE loss.</li>\n<li>Augmentations: rotate90, flip, random brightness, gamma, bunch of blurs.</li>\n<li>Use location-based stratification and split on 5 folds.</li>\n<li>Balance the dataset (roughly 40/60 ratio between images with ships/without ships)</li>\n</ol>\n\n<h1>Training the segmentation network:</h1>\n\n<ol>\n<li>Train in 3 stages: on 256x256 crops containing ships, then finetune on 384x384, and finally on 512x512. Inference on full-sized 768x768 images.</li>\n<li>Augmentations: rotate90 and flip. LinkNet had a trouble converging with heavy augmentations.</li>\n<li>Lovasz hinge loss with ELU + 1 trick.</li>\n<li>Drop last residual connection.</li>\n<li>Use location-based stratification and split on 5 folds.</li>\n<li>Balance the dataset (roughly 40/60 ratio between images with ships/without ships)</li>\n</ol>\n\n<h1>Some of the failed experiments:</h1>\n\n<ol>\n<li>FP16. K80 GPUs do not seem to support half-precision very well.</li>\n<li>Separable convolution. Got marginal speed improvement on LinkNet, but the model couldn't reach good F2 score.</li>\n<li>Simplifying larger network (U-Net based FPN with ResNet34). I used this model during the first stage of the competition. Couldn't get inference speed below 15 minutes.</li>\n<li>Simpler classifier networks. I tried smaller ResNets, but couldn't get them to the same level of accuracy. What they gave in terms of speed, was then taken by false positives during segmentation.</li>\n<li>Resizing images on GPU. Resizing itself was quicker on GPU, but transfer of large images from RAM and GPU was super-slow.</li>\n<li>Replacing MaxPool layers with a strided Conv. Again, marginal speed improvement, but pretty low F2 score.</li>\n<li>More aggressive pooling inside the network and upsampling at the final layer.</li>\n</ol>\n\n<h1>If I had more time I would:</h1>\n\n<ol>\n<li>Prune the network.</li>\n<li>Try to compete on CPU kernel, but with quantization and stuff. It would require using OpenVINO or PyTorch Glow.</li>\n<li>Reimplement scipy.ndimage on GPU. It took around 40 seconds of overall time.</li>\n</ol>",
  "messages": [
    {
      "id": "437659",
      "postDate": "12/12/2018 09:36:00",
      "content": "<p>Hi there!\nHere's a quick breakdown of the 1st place solution:</p>\n\n<ol>\n<li>PyTorch.</li>\n<li>SE-ResNet50 as a classifier: took around 30 seconds to infer all images.</li>\n<li>LinkNet as a segmentation network: took around 140 seconds to infer all positively classified images.</li>\n<li>Extracting ship instances from binary mask via scipy.ndimage. Ignore instances with small area (less than 80px): took around 40 seconds.</li>\n<li>Didn't use TTA.</li>\n</ol>\n\n<p>Initial profiling has shown that I/O is the slowest part of the pipeline, so I've spent at least a week trying to optimize it. However, after a some rule clarifications, it was clear that I/O should be ignored completely, so I switched to optimizing the networks. After a bunch of experiments, I ended up with a slightly modified LinkNet, which was at least x10 faster than my model from the first stage of the competition.</p>\n\n<h1>Key insights:</h1>\n\n<ol>\n<li>Adapt to the kernel. GPU kernel has 2 very slow CPU cores. I won 1 minute of inference time just by transferring image normalization to the GPU.</li>\n<li>Classifier speed and accuracy is critical: it saves a lot of time, since segmentation network is much slower.</li>\n<li>Predicting ship borders and/or using watershed requires too much post-processing.</li>\n</ol>\n\n<h1>Training the classifier:</h1>\n\n<ol>\n<li>Train and predict on resized 224x224 images.</li>\n<li>Nesterov SGD with LR 0.001, batch size 16, weight decay 1e-3, momentum 0.9.</li>\n<li>BCE loss.</li>\n<li>Augmentations: rotate90, flip, random brightness, gamma, bunch of blurs.</li>\n<li>Use location-based stratification and split on 5 folds.</li>\n<li>Balance the dataset (roughly 40/60 ratio between images with ships/without ships)</li>\n</ol>\n\n<h1>Training the segmentation network:</h1>\n\n<ol>\n<li>Train in 3 stages: on 256x256 crops containing ships, then finetune on 384x384, and finally on 512x512. Inference on full-sized 768x768 images.</li>\n<li>Augmentations: rotate90 and flip. LinkNet had a trouble converging with heavy augmentations.</li>\n<li>Lovasz hinge loss with ELU + 1 trick.</li>\n<li>Drop last residual connection.</li>\n<li>Use location-based stratification and split on 5 folds.</li>\n<li>Balance the dataset (roughly 40/60 ratio between images with ships/without ships)</li>\n</ol>\n\n<h1>Some of the failed experiments:</h1>\n\n<ol>\n<li>FP16. K80 GPUs do not seem to support half-precision very well.</li>\n<li>Separable convolution. Got marginal speed improvement on LinkNet, but the model couldn't reach good F2 score.</li>\n<li>Simplifying larger network (U-Net based FPN with ResNet34). I used this model during the first stage of the competition. Couldn't get inference speed below 15 minutes.</li>\n<li>Simpler classifier networks. I tried smaller ResNets, but couldn't get them to the same level of accuracy. What they gave in terms of speed, was then taken by false positives during segmentation.</li>\n<li>Resizing images on GPU. Resizing itself was quicker on GPU, but transfer of large images from RAM and GPU was super-slow.</li>\n<li>Replacing MaxPool layers with a strided Conv. Again, marginal speed improvement, but pretty low F2 score.</li>\n<li>More aggressive pooling inside the network and upsampling at the final layer.</li>\n</ol>\n\n<h1>If I had more time I would:</h1>\n\n<ol>\n<li>Prune the network.</li>\n<li>Try to compete on CPU kernel, but with quantization and stuff. It would require using OpenVINO or PyTorch Glow.</li>\n<li>Reimplement scipy.ndimage on GPU. It took around 40 seconds of overall time.</li>\n</ol>",
      "rawMarkdown": "Hi there!\nHere's a quick breakdown of the 1st place solution:\n\n0. PyTorch.\n1. SE-ResNet50 as a classifier: took around 30 seconds to infer all images.\n2. LinkNet as a segmentation network: took around 140 seconds to infer all positively classified images.\n3. Extracting ship instances from binary mask via scipy.ndimage. Ignore instances with small area (less than 80px): took around 40 seconds.\n4. Didn't use TTA.\n\nInitial profiling has shown that I/O is the slowest part of the pipeline, so I've spent at least a week trying to optimize it. However, after a some rule clarifications, it was clear that I/O should be ignored completely, so I switched to optimizing the networks. After a bunch of experiments, I ended up with a slightly modified LinkNet, which was at least x10 faster than my model from the first stage of the competition.\n\n#Key insights:\n\n1. Adapt to the kernel. GPU kernel has 2 very slow CPU cores. I won 1 minute of inference time just by transferring image normalization to the GPU.\n2. Classifier speed and accuracy is critical: it saves a lot of time, since segmentation network is much slower.\n3. Predicting ship borders and/or using watershed requires too much post-processing.\n\n#Training the classifier:\n\n1. Train and predict on resized 224x224 images.\n2. Nesterov SGD with LR 0.001, batch size 16, weight decay 1e-3, momentum 0.9.\n3. BCE loss.\n4. Augmentations: rotate90, flip, random brightness, gamma, bunch of blurs.\n5. Use location-based stratification and split on 5 folds.\n6. Balance the dataset (roughly 40/60 ratio between images with ships/without ships)\n\n#Training the segmentation network:\n\n1. Train in 3 stages: on 256x256 crops containing ships, then finetune on 384x384, and finally on 512x512. Inference on full-sized 768x768 images.\n2. Augmentations: rotate90 and flip. LinkNet had a trouble converging with heavy augmentations.\n3. Lovasz hinge loss with ELU + 1 trick.\n4. Drop last residual connection.\n5. Use location-based stratification and split on 5 folds.\n6. Balance the dataset (roughly 40/60 ratio between images with ships/without ships)\n\n#Some of the failed experiments:\n\n1. FP16. K80 GPUs do not seem to support half-precision very well.\n2. Separable convolution. Got marginal speed improvement on LinkNet, but the model couldn't reach good F2 score.\n3. Simplifying larger network (U-Net based FPN with ResNet34). I used this model during the first stage of the competition. Couldn't get inference speed below 15 minutes.\n4. Simpler classifier networks. I tried smaller ResNets, but couldn't get them to the same level of accuracy. What they gave in terms of speed, was then taken by false positives during segmentation.\n5. Resizing images on GPU. Resizing itself was quicker on GPU, but transfer of large images from RAM and GPU was super-slow.\n6. Replacing MaxPool layers with a strided Conv. Again, marginal speed improvement, but pretty low F2 score.\n7. More aggressive pooling inside the network and upsampling at the final layer.\n\n#If I had more time I would:\n\n 1. Prune the network.\n 2. Try to compete on CPU kernel, but with quantization and stuff. It would require using OpenVINO or PyTorch Glow.\n 3. Reimplement scipy.ndimage on GPU. It took around 40 seconds of overall time.",
      "votes": null
    },
    {
      "id": "437671",
      "postDate": "12/12/2018 09:49:22",
      "content": "<blockquote>\n  <p>However, after a some rule clarifications, it was clear that I/O should be ignored completely, so I switched to optimizing the networks.</p>\n</blockquote>\n\n<p>What do you mean by this? Did you include the time to copy the data to the GPU?</p>",
      "rawMarkdown": "&gt; However, after a some rule clarifications, it was clear that I/O should be ignored completely, so I switched to optimizing the networks.\n\nWhat do you mean by this? Did you include the time to copy the data to the GPU?",
      "votes": null
    },
    {
      "id": "437676",
      "postDate": "12/12/2018 09:55:10",
      "content": "<p>Yes, I included the time to copy tensors to from CPU to GPU.\nHowever, reading images from the disk was ignored.</p>",
      "rawMarkdown": "Yes, I included the time to copy tensors to from CPU to GPU.\nHowever, reading images from the disk was ignored.",
      "votes": null
    },
    {
      "id": "437686",
      "postDate": "12/12/2018 10:09:07",
      "content": "<blockquote>\n  <p>SE-ResNet50 as a classifier: took around 30 seconds to infer all images.</p>\n</blockquote>\n\n<p>So this is (30 / 15606) * 1000 = 1.92ms per image for the classifier?</p>",
      "rawMarkdown": "&gt; SE-ResNet50 as a classifier: took around 30 seconds to infer all images.\n\nSo this is (30 / 15606) * 1000 = 1.92ms per image for the classifier?",
      "votes": null
    },
    {
      "id": "437700",
      "postDate": "12/12/2018 10:28:46",
      "content": "<p>This doesn't make sense to me. This is about as fast as a RTX2080TI. 16 (batch-size) * 1.92ms = 30.72ms (K80) vs 28.4ms (RTX2080TI).  This is SE-ResNet50 vs VGG-16 at 224 resolution. I would expect the SE-ResNet50 to be much slower and K80 as well. I got the numbers for the RTX2080TI from here: <a href=\"https://nikolasent.github.io/hardware,/deeplearning/2018/11/06/Benchmarking-RTX-2080-Ti-vs-Pascal-GPUs-with-DL-tasks.html\">https://nikolasent.github.io/hardware,/deeplearning/2018/11/06/Benchmarking-RTX-2080-Ti-vs-Pascal-GPUs-with-DL-tasks.html</a></p>\n\n<p>I think your timings are wrong. The classifier <strong>has</strong> to take much longer...</p>",
      "rawMarkdown": "This doesn't make sense to me. This is about as fast as a RTX2080TI. 16 (batch-size) * 1.92ms = 30.72ms (K80) vs 28.4ms (RTX2080TI).  This is SE-ResNet50 vs VGG-16 at 224 resolution. I would expect the SE-ResNet50 to be much slower and K80 as well. I got the numbers for the RTX2080TI from here: https://nikolasent.github.io/hardware,/deeplearning/2018/11/06/Benchmarking-RTX-2080-Ti-vs-Pascal-GPUs-with-DL-tasks.html\n\nI think your timings are wrong. The classifier **has** to take much longer...",
      "votes": null
    },
    {
      "id": "437701",
      "postDate": "12/12/2018 10:36:36",
      "content": "<p>But according to this benchmark: <a href=\"https://arxiv.org/pdf/1810.00736.pdf\">https://arxiv.org/pdf/1810.00736.pdf</a>\nSE-ResNet50 is much faster than VGG16. It also correlates with my experiments.</p>\n\n<p>You can also try to reproduce the results. I used the implementation from here:\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/senet.py\">https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/senet.py</a></p>",
      "rawMarkdown": "But according to this benchmark: https://arxiv.org/pdf/1810.00736.pdf\nSE-ResNet50 is much faster than VGG16. It also correlates with my experiments.\n\nYou can also try to reproduce the results. I used the implementation from here:\nhttps://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/senet.py",
      "votes": null
    },
    {
      "id": "437702",
      "postDate": "12/12/2018 10:37:27",
      "content": "<p>This also sounds a bit confusing to me... We used resnet18 based classifier on 384x384 images with very aggressive pooling and inference time was about 1.5 minutes :/</p>",
      "rawMarkdown": "This also sounds a bit confusing to me... We used resnet18 based classifier on 384x384 images with very aggressive pooling and inference time was about 1.5 minutes :/",
      "votes": null
    },
    {
      "id": "437717",
      "postDate": "12/12/2018 11:07:12",
      "content": "<p>I just tried running SE-ResNet50 in a kernel and got approximately 0.4 seconds for 32 224x224 batch which gives roughly 180 seconds for all 15k images. I wonder, what you did to make it run so fast?</p>",
      "rawMarkdown": "I just tried running SE-ResNet50 in a kernel and got approximately 0.4 seconds for 32 224x224 batch which gives roughly 180 seconds for all 15k images. I wonder, what you did to make it run so fast?",
      "votes": null
    },
    {
      "id": "437721",
      "postDate": "12/12/2018 11:11:27",
      "content": "<blockquote>\n  <p>SE-ResNet50 is much faster than VGG16. It also correlates with my experiments.</p>\n</blockquote>\n\n<p>Thanks for posting the paper. I just skimmed through it. Were did you find that SE-ResNet50 is faster? I see 200 FPS(VGG-16) vs 150 FPS(SE-ResNet50). So actually VGG <strong>is</strong> faster.</p>\n\n<p>Update: It is SE-ResNet50 not SE-ResNext50\n<img src=\"https://i.imgur.com/mDScLYa.png\" alt=\"FPS\"></p>",
      "rawMarkdown": "&gt;  SE-ResNet50 is much faster than VGG16. It also correlates with my experiments.\n\nThanks for posting the paper. I just skimmed through it. Were did you find that SE-ResNet50 is faster? I see 200 FPS(VGG-16) vs 150 FPS(SE-ResNet50). So actually VGG **is** faster.\n\nUpdate: It is SE-ResNet50 not SE-ResNext50\n![FPS][1]\n\n\n  [1]: https://i.imgur.com/mDScLYa.png",
      "votes": null
    },
    {
      "id": "437828",
      "postDate": "12/12/2018 15:21:45",
      "content": "<p><a href=\"/marvelousninja\">@marvelousninja</a> <a href=\"/inversion\">@inversion</a> <a href=\"/jeffaudi\">@jeffaudi</a></p>\n\n<p>I created a kernel: <a href=\"https://www.kaggle.com/seesee/se-resnet50-vs-vgg-16\">https://www.kaggle.com/seesee/se-resnet50-vs-vgg-16</a> . I can't reproduce  anything claimed in this post. Neither the 30 seconds, nor SE-ResNet-50 being \"much faster\" than VGG-16. Please tell me where I made the mistake. Thanks!</p>",
      "rawMarkdown": "marvelousninja @inversion @jeffaudi\n\nI created a kernel: https://www.kaggle.com/seesee/se-resnet50-vs-vgg-16 . I can't reproduce  anything claimed in this post. Neither the 30 seconds, nor SE-ResNet-50 being \"much faster\" than VGG-16. Please tell me where I made the mistake. Thanks!",
      "votes": null
    },
    {
      "id": "437941",
      "postDate": "12/12/2018 19:56:30",
      "content": "<p>I think it would be nice to make all kernels (or at least high-scoring ones) submitted for the SpeedPrize public. I suppose it would eliminate all misunderstandings and suspicions.</p>\n\n<p><a href=\"/inversion\">@inversion</a> <a href=\"/jeffaudi\">@jeffaudi</a></p>",
      "rawMarkdown": "I think it would be nice to make all kernels (or at least high-scoring ones) submitted for the SpeedPrize public. I suppose it would eliminate all misunderstandings and suspicions.\n\n@inversion @jeffaudi",
      "votes": null
    },
    {
      "id": "438391",
      "postDate": "12/13/2018 15:28:21",
      "content": "<p><a href=\"/seesee\">@seesee</a> Thank you for sharing this kernel. We will wait for <a href=\"/marvelousninja\">@marvelousninja</a> feedback. \n<a href=\"/ddanevskyi\">@ddanevskyi</a> We will discuss this with <a href=\"/inversion\">@inversion</a> and the Kaggle team.</p>",
      "rawMarkdown": "seesee Thank you for sharing this kernel. We will wait for @marvelousninja feedback. \n@ddanevskyi We will discuss this with @inversion and the Kaggle team.",
      "votes": null
    },
    {
      "id": "438729",
      "postDate": "12/14/2018 04:56:31",
      "content": "<p>I want to express thanks for the efforts and debate. There are quite a number of things that make a Speed Test complicated and messy. The use of Kernels was an attempt to balance some practical considerations while also try to get a reasonable signal. This, unfortunately, means that there will be plenty to disagree with and to dispute. </p>\n\n<p>As this is a special prize, added above and beyond the normal competition prizes, we give latitude to the competition sponsors to make their best, good faith, effort to properly assess the entries, and provide the prize to those they deem to have best meet the guidelines as provided. We will though, as always, take all the feedback into consideration if a special prize is offered in future competitions.</p>",
      "rawMarkdown": "I want to express thanks for the efforts and debate. There are quite a number of things that make a Speed Test complicated and messy. The use of Kernels was an attempt to balance some practical considerations while also try to get a reasonable signal. This, unfortunately, means that there will be plenty to disagree with and to dispute. \n\nAs this is a special prize, added above and beyond the normal competition prizes, we give latitude to the competition sponsors to make their best, good faith, effort to properly assess the entries, and provide the prize to those they deem to have best meet the guidelines as provided. We will though, as always, take all the feedback into consideration if a special prize is offered in future competitions.",
      "votes": null
    },
    {
      "id": "438794",
      "postDate": "12/14/2018 07:35:19",
      "content": "<p><a href=\"/inversion\">@inversion</a> Does this mean you are not going to double check our submissions?</p>",
      "rawMarkdown": "inversion Does this mean you are not going to double check our submissions?",
      "votes": null
    },
    {
      "id": "438804",
      "postDate": "12/14/2018 08:05:52",
      "content": "<p>It was not my intend to dispute, rather clarify. The results look just too good to be true (from my point of view). Though, I will wait for <a href=\"/marvelousninja\">@marvelousninja</a> feedback. Let's see, maybe it is possible.</p>\n\n<p>As you noted, benchmarking code is indeed not easy. I was therefore really surprised when you changed the rules/template and relied on the users to correctly measure their code. If we just sticked with this (it is still in the speed prize tab):</p>\n\n<pre><code>import time\ninference_start = time.time()\n#### inference code ####\ninference_end = time.time()\nprint('Inference Time: %0.2f Minutes'%((inference_end - inference_start)/60))\n</code></pre>\n\n<p>we would not be in the current situation.</p>\n\n<p>I agree with <a href=\"/ddanevskyi\">@ddanevskyi</a> that, at this point, the best would be to make the kernels public.</p>",
      "rawMarkdown": "It was not my intend to dispute, rather clarify. The results look just too good to be true (from my point of view). Though, I will wait for @marvelousninja feedback. Let's see, maybe it is possible.\n\nAs you noted, benchmarking code is indeed not easy. I was therefore really surprised when you changed the rules/template and relied on the users to correctly measure their code. If we just sticked with this (it is still in the speed prize tab):\n\n\n    import time\n    inference_start = time.time()\n    #### inference code ####\n    inference_end = time.time()\n    print('Inference Time: %0.2f Minutes'%((inference_end - inference_start)/60))\n\nwe would not be in the current situation.\n\nI agree with @ddanevskyi that, at this point, the best would be to make the kernels public.",
      "votes": null
    },
    {
      "id": "439000",
      "postDate": "12/14/2018 14:54:56",
      "content": "<p><a href=\"/seesee\">@seesee</a> Concerning the measurement of the time, I was the one to clarify/update the rules probably without full understanding of the consequences. We run our machine learning predictors on the cloud as micro-services (using docker containers) and they got the imagery sent to them inside so I am really concerned about the speed of the prediction and not the speed of I/O or startup up of the containers. This is why I proposed to exclude some parts to fit more precisely our use case. I hope that this will not cause too much trouble and that everyone will feel comfortable with the outcome of the speed prize challenge. </p>",
      "rawMarkdown": "seesee Concerning the measurement of the time, I was the one to clarify/update the rules probably without full understanding of the consequences. We run our machine learning predictors on the cloud as micro-services (using docker containers) and they got the imagery sent to them inside so I am really concerned about the speed of the prediction and not the speed of I/O or startup up of the containers. This is why I proposed to exclude some parts to fit more precisely our use case. I hope that this will not cause too much trouble and that everyone will feel comfortable with the outcome of the speed prize challenge.",
      "votes": null
    },
    {
      "id": "439029",
      "postDate": "12/14/2018 16:06:28",
      "content": "<p>Hi everyone!\nI've spent a couple of nights checking the timings.</p>\n\n<p><strong>TL; DR</strong> - We measured time incorrectly. Wrapping the code with .time() calls is not enough.\nI have assembled a kernel that should explain what happened:</p>\n\n<p><a href=\"https://www.kaggle.com/marvelousninja/pytorch-synchronization-issue-1-0\">https://www.kaggle.com/marvelousninja/pytorch-synchronization-issue-1-0</a></p>\n\n<p>For the kernel with SE-ResNet50, it resulted in approximately 120 seconds gain.\nWith synchronization enabled, it brings the time to 5.5 minutes.</p>\n\n<p>Considering the nature of the bug, I suggested to use a different kernel as my best submission.\nWith synchronization enabled, it has a time of 3.96 minutes, which is still enough to win.\nIt uses ResNet-18 as a classifier on 224x224 images and LinkNet for segmentation on 768x768 images.</p>\n\n<p>As a show of good sportsmanship, I'm ready to publish this kernel.\nOf course, if <a href=\"/inversion\">@inversion</a> and <a href=\"/jeffaudi\">@jeffaudi</a> approve this.</p>",
      "rawMarkdown": "Hi everyone!\nI've spent a couple of nights checking the timings.\n\n**TL; DR** - We measured time incorrectly. Wrapping the code with .time() calls is not enough.\nI have assembled a kernel that should explain what happened:\n\nhttps://www.kaggle.com/marvelousninja/pytorch-synchronization-issue-1-0\n\nFor the kernel with SE-ResNet50, it resulted in approximately 120 seconds gain.\nWith synchronization enabled, it brings the time to 5.5 minutes.\n\nConsidering the nature of the bug, I suggested to use a different kernel as my best submission.\nWith synchronization enabled, it has a time of 3.96 minutes, which is still enough to win.\nIt uses ResNet-18 as a classifier on 224x224 images and LinkNet for segmentation on 768x768 images.\n\nAs a show of good sportsmanship, I'm ready to publish this kernel.\nOf course, if @inversion and @jeffaudi approve this.",
      "votes": null
    },
    {
      "id": "439041",
      "postDate": "12/14/2018 16:31:02",
      "content": "<p><a href=\"/marvelousninja\">@marvelousninja</a> No problem from my side and <a href=\"/inversion\">@inversion</a> to publish your kernel. </p>",
      "rawMarkdown": "marvelousninja No problem from my side and @inversion to publish your kernel.",
      "votes": null
    },
    {
      "id": "439076",
      "postDate": "12/14/2018 17:32:57",
      "content": "<p>Ok, here is the kernel!\n<a href=\"https://www.kaggle.com/marvelousninja/kaggle-airbus-3-96-minutes\">https://www.kaggle.com/marvelousninja/kaggle-airbus-3-96-minutes</a></p>",
      "rawMarkdown": "Ok, here is the kernel!\nhttps://www.kaggle.com/marvelousninja/kaggle-airbus-3-96-minutes",
      "votes": null
    },
    {
      "id": "439159",
      "postDate": "12/14/2018 19:59:22",
      "content": "<p>Hello, I just read this topic here, <a href=\"/marvelousninja\">@marvelousninja</a>, and as I understood you did not submit the correct kernel in time. \nSo \"as a show of good sportsmanship\" after two weeks after this competition a submission of a new kernel with the correct timing is NOT \"enough to win\", unfortunately.  </p>\n\n<p>The rules are clear: November 30, 2018 - Final submission deadline for speed prize kernel submission.</p>\n\n<p>Please, make rules clear and fair for all of the applicants, Kaggle platform is for that, <a href=\"/inversion\">@inversion</a>!  And also, take into a consideration just kernels that were submitted before the deadline. <a href=\"/jeffaudi\">@jeffaudi</a></p>",
      "rawMarkdown": "Hello, I just read this topic here, @marvelousninja, and as I understood you did not submit the correct kernel in time. \nSo \"as a show of good sportsmanship\" after two weeks after this competition a submission of a new kernel with the correct timing is NOT \"enough to win\", unfortunately.  \n\nThe rules are clear: November 30, 2018 - Final submission deadline for speed prize kernel submission.\n\nPlease, make rules clear and fair for all of the applicants, Kaggle platform is for that, @inversion!  And also, take into a consideration just kernels that were submitted before the deadline. @jeffaudi",
      "votes": null
    },
    {
      "id": "439377",
      "postDate": "12/15/2018 09:37:35",
      "content": "<p><a href=\"/jeffaudi\">@jeffaudi</a> <a href=\"/inversion\">@inversion</a></p>\n\n<p>FYI, our initial submission (4.69) takes only 2.90 to run. We did <strong>NOT</strong> change anything in the model, only added <code>torch.backends.cudnn.benchmark = True</code> and warmup rounds before the predictions loops.</p>",
      "rawMarkdown": "jeffaudi @inversion\n\nFYI, our initial submission (4.69) takes only 2.90 to run. We did **NOT** change anything in the model, only added `torch.backends.cudnn.benchmark = True` and warmup rounds before the predictions loops.",
      "votes": null
    },
    {
      "id": "439406",
      "postDate": "12/15/2018 11:31:14",
      "content": "<blockquote>\n  <p>With synchronization enabled, it brings the time to 5.5 minutes.</p>\n</blockquote>\n\n<p>It is great that you fixed the bug. Nevertheless, this time is only the 3rd place (assuming correct timings). Team [attention heads] and my kernel would be faster.</p>\n\n<blockquote>\n  <p>We did NOT change anything in the model only added torch.backends.cudnn.benchmark = True and warmup rounds before the predictions loops.</p>\n</blockquote>\n\n<p>How is this \"not change anything\" :) ?. To me this sounds like a couple of lines of code...</p>\n\n<p>Anyway, I don't see this is leading to an outcome \"everyone will feel comfortable with\". I'd be fine to call it a day and split the prize. I see 3 great solutions to the problem of detecting ships as fast as possible. The hosts can choose what they want to put into production. Please let me know what you think.</p>",
      "rawMarkdown": "&gt; With synchronization enabled, it brings the time to 5.5 minutes.\n\nIt is great that you fixed the bug. Nevertheless, this time is only the 3rd place (assuming correct timings). Team [attention heads] and my kernel would be faster.\n \n&gt; We did NOT change anything in the model only added torch.backends.cudnn.benchmark = True and warmup rounds before the predictions loops.\n\nHow is this \"not change anything\" :) ?. To me this sounds like a couple of lines of code...\n\nAnyway, I don't see this is leading to an outcome \"everyone will feel comfortable with\". I'd be fine to call it a day and split the prize. I see 3 great solutions to the problem of detecting ships as fast as possible. The hosts can choose what they want to put into production. Please let me know what you think.",
      "votes": null
    },
    {
      "id": "456794",
      "postDate": "01/16/2019 14:57:48",
      "content": "<p>can you share your code?</p>",
      "rawMarkdown": "can you share your code?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 437671,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "12/12/2018 09:49:22",
      "content": "<blockquote>\n  <p>However, after a some rule clarifications, it was clear that I/O should be ignored completely, so I switched to optimizing the networks.</p>\n</blockquote>\n\n<p>What do you mean by this? Did you include the time to copy the data to the GPU?</p>",
      "votes": null,
      "replies": [
        {
          "id": 437676,
          "author_name": "marvelousninja",
          "author_url": "",
          "post_date": "12/12/2018 09:55:10",
          "content": "<p>Yes, I included the time to copy tensors to from CPU to GPU.\nHowever, reading images from the disk was ignored.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437686,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "12/12/2018 10:09:07",
          "content": "<blockquote>\n  <p>SE-ResNet50 as a classifier: took around 30 seconds to infer all images.</p>\n</blockquote>\n\n<p>So this is (30 / 15606) * 1000 = 1.92ms per image for the classifier?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437700,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "12/12/2018 10:28:46",
          "content": "<p>This doesn't make sense to me. This is about as fast as a RTX2080TI. 16 (batch-size) * 1.92ms = 30.72ms (K80) vs 28.4ms (RTX2080TI).  This is SE-ResNet50 vs VGG-16 at 224 resolution. I would expect the SE-ResNet50 to be much slower and K80 as well. I got the numbers for the RTX2080TI from here: <a href=\"https://nikolasent.github.io/hardware,/deeplearning/2018/11/06/Benchmarking-RTX-2080-Ti-vs-Pascal-GPUs-with-DL-tasks.html\">https://nikolasent.github.io/hardware,/deeplearning/2018/11/06/Benchmarking-RTX-2080-Ti-vs-Pascal-GPUs-with-DL-tasks.html</a></p>\n\n<p>I think your timings are wrong. The classifier <strong>has</strong> to take much longer...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437701,
          "author_name": "marvelousninja",
          "author_url": "",
          "post_date": "12/12/2018 10:36:36",
          "content": "<p>But according to this benchmark: <a href=\"https://arxiv.org/pdf/1810.00736.pdf\">https://arxiv.org/pdf/1810.00736.pdf</a>\nSE-ResNet50 is much faster than VGG16. It also correlates with my experiments.</p>\n\n<p>You can also try to reproduce the results. I used the implementation from here:\n<a href=\"https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/senet.py\">https://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/senet.py</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437702,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "12/12/2018 10:37:27",
          "content": "<p>This also sounds a bit confusing to me... We used resnet18 based classifier on 384x384 images with very aggressive pooling and inference time was about 1.5 minutes :/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437717,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "12/12/2018 11:07:12",
          "content": "<p>I just tried running SE-ResNet50 in a kernel and got approximately 0.4 seconds for 32 224x224 batch which gives roughly 180 seconds for all 15k images. I wonder, what you did to make it run so fast?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437721,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "12/12/2018 11:11:27",
          "content": "<blockquote>\n  <p>SE-ResNet50 is much faster than VGG16. It also correlates with my experiments.</p>\n</blockquote>\n\n<p>Thanks for posting the paper. I just skimmed through it. Were did you find that SE-ResNet50 is faster? I see 200 FPS(VGG-16) vs 150 FPS(SE-ResNet50). So actually VGG <strong>is</strong> faster.</p>\n\n<p>Update: It is SE-ResNet50 not SE-ResNext50\n<img src=\"https://i.imgur.com/mDScLYa.png\" alt=\"FPS\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437828,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "12/12/2018 15:21:45",
          "content": "<p><a href=\"/marvelousninja\">@marvelousninja</a> <a href=\"/inversion\">@inversion</a> <a href=\"/jeffaudi\">@jeffaudi</a></p>\n\n<p>I created a kernel: <a href=\"https://www.kaggle.com/seesee/se-resnet50-vs-vgg-16\">https://www.kaggle.com/seesee/se-resnet50-vs-vgg-16</a> . I can't reproduce  anything claimed in this post. Neither the 30 seconds, nor SE-ResNet-50 being \"much faster\" than VGG-16. Please tell me where I made the mistake. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437941,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "12/12/2018 19:56:30",
          "content": "<p>I think it would be nice to make all kernels (or at least high-scoring ones) submitted for the SpeedPrize public. I suppose it would eliminate all misunderstandings and suspicions.</p>\n\n<p><a href=\"/inversion\">@inversion</a> <a href=\"/jeffaudi\">@jeffaudi</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438391,
          "author_name": "jeffaudi",
          "author_url": "",
          "post_date": "12/13/2018 15:28:21",
          "content": "<p><a href=\"/seesee\">@seesee</a> Thank you for sharing this kernel. We will wait for <a href=\"/marvelousninja\">@marvelousninja</a> feedback. \n<a href=\"/ddanevskyi\">@ddanevskyi</a> We will discuss this with <a href=\"/inversion\">@inversion</a> and the Kaggle team.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438729,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "12/14/2018 04:56:31",
          "content": "<p>I want to express thanks for the efforts and debate. There are quite a number of things that make a Speed Test complicated and messy. The use of Kernels was an attempt to balance some practical considerations while also try to get a reasonable signal. This, unfortunately, means that there will be plenty to disagree with and to dispute. </p>\n\n<p>As this is a special prize, added above and beyond the normal competition prizes, we give latitude to the competition sponsors to make their best, good faith, effort to properly assess the entries, and provide the prize to those they deem to have best meet the guidelines as provided. We will though, as always, take all the feedback into consideration if a special prize is offered in future competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438794,
          "author_name": "yaroshevskiy",
          "author_url": "",
          "post_date": "12/14/2018 07:35:19",
          "content": "<p><a href=\"/inversion\">@inversion</a> Does this mean you are not going to double check our submissions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438804,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "12/14/2018 08:05:52",
          "content": "<p>It was not my intend to dispute, rather clarify. The results look just too good to be true (from my point of view). Though, I will wait for <a href=\"/marvelousninja\">@marvelousninja</a> feedback. Let's see, maybe it is possible.</p>\n\n<p>As you noted, benchmarking code is indeed not easy. I was therefore really surprised when you changed the rules/template and relied on the users to correctly measure their code. If we just sticked with this (it is still in the speed prize tab):</p>\n\n<pre><code>import time\ninference_start = time.time()\n#### inference code ####\ninference_end = time.time()\nprint('Inference Time: %0.2f Minutes'%((inference_end - inference_start)/60))\n</code></pre>\n\n<p>we would not be in the current situation.</p>\n\n<p>I agree with <a href=\"/ddanevskyi\">@ddanevskyi</a> that, at this point, the best would be to make the kernels public.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439000,
          "author_name": "jeffaudi",
          "author_url": "",
          "post_date": "12/14/2018 14:54:56",
          "content": "<p><a href=\"/seesee\">@seesee</a> Concerning the measurement of the time, I was the one to clarify/update the rules probably without full understanding of the consequences. We run our machine learning predictors on the cloud as micro-services (using docker containers) and they got the imagery sent to them inside so I am really concerned about the speed of the prediction and not the speed of I/O or startup up of the containers. This is why I proposed to exclude some parts to fit more precisely our use case. I hope that this will not cause too much trouble and that everyone will feel comfortable with the outcome of the speed prize challenge. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439029,
          "author_name": "marvelousninja",
          "author_url": "",
          "post_date": "12/14/2018 16:06:28",
          "content": "<p>Hi everyone!\nI've spent a couple of nights checking the timings.</p>\n\n<p><strong>TL; DR</strong> - We measured time incorrectly. Wrapping the code with .time() calls is not enough.\nI have assembled a kernel that should explain what happened:</p>\n\n<p><a href=\"https://www.kaggle.com/marvelousninja/pytorch-synchronization-issue-1-0\">https://www.kaggle.com/marvelousninja/pytorch-synchronization-issue-1-0</a></p>\n\n<p>For the kernel with SE-ResNet50, it resulted in approximately 120 seconds gain.\nWith synchronization enabled, it brings the time to 5.5 minutes.</p>\n\n<p>Considering the nature of the bug, I suggested to use a different kernel as my best submission.\nWith synchronization enabled, it has a time of 3.96 minutes, which is still enough to win.\nIt uses ResNet-18 as a classifier on 224x224 images and LinkNet for segmentation on 768x768 images.</p>\n\n<p>As a show of good sportsmanship, I'm ready to publish this kernel.\nOf course, if <a href=\"/inversion\">@inversion</a> and <a href=\"/jeffaudi\">@jeffaudi</a> approve this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439041,
          "author_name": "jeffaudi",
          "author_url": "",
          "post_date": "12/14/2018 16:31:02",
          "content": "<p><a href=\"/marvelousninja\">@marvelousninja</a> No problem from my side and <a href=\"/inversion\">@inversion</a> to publish your kernel. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439076,
          "author_name": "marvelousninja",
          "author_url": "",
          "post_date": "12/14/2018 17:32:57",
          "content": "<p>Ok, here is the kernel!\n<a href=\"https://www.kaggle.com/marvelousninja/kaggle-airbus-3-96-minutes\">https://www.kaggle.com/marvelousninja/kaggle-airbus-3-96-minutes</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439377,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "12/15/2018 09:37:35",
          "content": "<p><a href=\"/jeffaudi\">@jeffaudi</a> <a href=\"/inversion\">@inversion</a></p>\n\n<p>FYI, our initial submission (4.69) takes only 2.90 to run. We did <strong>NOT</strong> change anything in the model, only added <code>torch.backends.cudnn.benchmark = True</code> and warmup rounds before the predictions loops.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439406,
          "author_name": "seesee",
          "author_url": "",
          "post_date": "12/15/2018 11:31:14",
          "content": "<blockquote>\n  <p>With synchronization enabled, it brings the time to 5.5 minutes.</p>\n</blockquote>\n\n<p>It is great that you fixed the bug. Nevertheless, this time is only the 3rd place (assuming correct timings). Team [attention heads] and my kernel would be faster.</p>\n\n<blockquote>\n  <p>We did NOT change anything in the model only added torch.backends.cudnn.benchmark = True and warmup rounds before the predictions loops.</p>\n</blockquote>\n\n<p>How is this \"not change anything\" :) ?. To me this sounds like a couple of lines of code...</p>\n\n<p>Anyway, I don't see this is leading to an outcome \"everyone will feel comfortable with\". I'd be fine to call it a day and split the prize. I see 3 great solutions to the problem of detecting ships as fast as possible. The hosts can choose what they want to put into production. Please let me know what you think.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 439159,
      "author_name": "orgunova",
      "author_url": "",
      "post_date": "12/14/2018 19:59:22",
      "content": "<p>Hello, I just read this topic here, <a href=\"/marvelousninja\">@marvelousninja</a>, and as I understood you did not submit the correct kernel in time. \nSo \"as a show of good sportsmanship\" after two weeks after this competition a submission of a new kernel with the correct timing is NOT \"enough to win\", unfortunately.  </p>\n\n<p>The rules are clear: November 30, 2018 - Final submission deadline for speed prize kernel submission.</p>\n\n<p>Please, make rules clear and fair for all of the applicants, Kaggle platform is for that, <a href=\"/inversion\">@inversion</a>!  And also, take into a consideration just kernels that were submitted before the deadline. <a href=\"/jeffaudi\">@jeffaudi</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 456794,
      "author_name": "xiaojidan",
      "author_url": "",
      "post_date": "01/16/2019 14:57:48",
      "content": "<p>can you share your code?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "437659": "Hi there!\nHere's a quick breakdown of the 1st place solution:\n\n0. PyTorch.\n1. SE-ResNet50 as a classifier: took around 30 seconds to infer all images.\n2. LinkNet as a segmentation network: took around 140 seconds to infer all positively classified images.\n3. Extracting ship instances from binary mask via scipy.ndimage. Ignore instances with small area (less than 80px): took around 40 seconds.\n4. Didn't use TTA.\n\nInitial profiling has shown that I/O is the slowest part of the pipeline, so I've spent at least a week trying to optimize it. However, after a some rule clarifications, it was clear that I/O should be ignored completely, so I switched to optimizing the networks. After a bunch of experiments, I ended up with a slightly modified LinkNet, which was at least x10 faster than my model from the first stage of the competition.\n\n#Key insights:\n\n1. Adapt to the kernel. GPU kernel has 2 very slow CPU cores. I won 1 minute of inference time just by transferring image normalization to the GPU.\n2. Classifier speed and accuracy is critical: it saves a lot of time, since segmentation network is much slower.\n3. Predicting ship borders and/or using watershed requires too much post-processing.\n\n#Training the classifier:\n\n1. Train and predict on resized 224x224 images.\n2. Nesterov SGD with LR 0.001, batch size 16, weight decay 1e-3, momentum 0.9.\n3. BCE loss.\n4. Augmentations: rotate90, flip, random brightness, gamma, bunch of blurs.\n5. Use location-based stratification and split on 5 folds.\n6. Balance the dataset (roughly 40/60 ratio between images with ships/without ships)\n\n#Training the segmentation network:\n\n1. Train in 3 stages: on 256x256 crops containing ships, then finetune on 384x384, and finally on 512x512. Inference on full-sized 768x768 images.\n2. Augmentations: rotate90 and flip. LinkNet had a trouble converging with heavy augmentations.\n3. Lovasz hinge loss with ELU + 1 trick.\n4. Drop last residual connection.\n5. Use location-based stratification and split on 5 folds.\n6. Balance the dataset (roughly 40/60 ratio between images with ships/without ships)\n\n#Some of the failed experiments:\n\n1. FP16. K80 GPUs do not seem to support half-precision very well.\n2. Separable convolution. Got marginal speed improvement on LinkNet, but the model couldn't reach good F2 score.\n3. Simplifying larger network (U-Net based FPN with ResNet34). I used this model during the first stage of the competition. Couldn't get inference speed below 15 minutes.\n4. Simpler classifier networks. I tried smaller ResNets, but couldn't get them to the same level of accuracy. What they gave in terms of speed, was then taken by false positives during segmentation.\n5. Resizing images on GPU. Resizing itself was quicker on GPU, but transfer of large images from RAM and GPU was super-slow.\n6. Replacing MaxPool layers with a strided Conv. Again, marginal speed improvement, but pretty low F2 score.\n7. More aggressive pooling inside the network and upsampling at the final layer.\n\n#If I had more time I would:\n\n 1. Prune the network.\n 2. Try to compete on CPU kernel, but with quantization and stuff. It would require using OpenVINO or PyTorch Glow.\n 3. Reimplement scipy.ndimage on GPU. It took around 40 seconds of overall time.",
    "437671": "&gt; However, after a some rule clarifications, it was clear that I/O should be ignored completely, so I switched to optimizing the networks.\n\nWhat do you mean by this? Did you include the time to copy the data to the GPU?",
    "437676": "Yes, I included the time to copy tensors to from CPU to GPU.\nHowever, reading images from the disk was ignored.",
    "437686": "&gt; SE-ResNet50 as a classifier: took around 30 seconds to infer all images.\n\nSo this is (30 / 15606) * 1000 = 1.92ms per image for the classifier?",
    "437700": "This doesn't make sense to me. This is about as fast as a RTX2080TI. 16 (batch-size) * 1.92ms = 30.72ms (K80) vs 28.4ms (RTX2080TI).  This is SE-ResNet50 vs VGG-16 at 224 resolution. I would expect the SE-ResNet50 to be much slower and K80 as well. I got the numbers for the RTX2080TI from here: https://nikolasent.github.io/hardware,/deeplearning/2018/11/06/Benchmarking-RTX-2080-Ti-vs-Pascal-GPUs-with-DL-tasks.html\n\nI think your timings are wrong. The classifier **has** to take much longer...",
    "437701": "But according to this benchmark: https://arxiv.org/pdf/1810.00736.pdf\nSE-ResNet50 is much faster than VGG16. It also correlates with my experiments.\n\nYou can also try to reproduce the results. I used the implementation from here:\nhttps://github.com/Cadene/pretrained-models.pytorch/blob/master/pretrainedmodels/models/senet.py",
    "437702": "This also sounds a bit confusing to me... We used resnet18 based classifier on 384x384 images with very aggressive pooling and inference time was about 1.5 minutes :/",
    "437717": "I just tried running SE-ResNet50 in a kernel and got approximately 0.4 seconds for 32 224x224 batch which gives roughly 180 seconds for all 15k images. I wonder, what you did to make it run so fast?",
    "437721": "&gt;  SE-ResNet50 is much faster than VGG16. It also correlates with my experiments.\n\nThanks for posting the paper. I just skimmed through it. Were did you find that SE-ResNet50 is faster? I see 200 FPS(VGG-16) vs 150 FPS(SE-ResNet50). So actually VGG **is** faster.\n\nUpdate: It is SE-ResNet50 not SE-ResNext50\n![FPS][1]\n\n\n  [1]: https://i.imgur.com/mDScLYa.png",
    "437828": "marvelousninja @inversion @jeffaudi\n\nI created a kernel: https://www.kaggle.com/seesee/se-resnet50-vs-vgg-16 . I can't reproduce  anything claimed in this post. Neither the 30 seconds, nor SE-ResNet-50 being \"much faster\" than VGG-16. Please tell me where I made the mistake. Thanks!",
    "437941": "I think it would be nice to make all kernels (or at least high-scoring ones) submitted for the SpeedPrize public. I suppose it would eliminate all misunderstandings and suspicions.\n\n@inversion @jeffaudi",
    "438391": "seesee Thank you for sharing this kernel. We will wait for @marvelousninja feedback. \n@ddanevskyi We will discuss this with @inversion and the Kaggle team.",
    "438729": "I want to express thanks for the efforts and debate. There are quite a number of things that make a Speed Test complicated and messy. The use of Kernels was an attempt to balance some practical considerations while also try to get a reasonable signal. This, unfortunately, means that there will be plenty to disagree with and to dispute. \n\nAs this is a special prize, added above and beyond the normal competition prizes, we give latitude to the competition sponsors to make their best, good faith, effort to properly assess the entries, and provide the prize to those they deem to have best meet the guidelines as provided. We will though, as always, take all the feedback into consideration if a special prize is offered in future competitions.",
    "438794": "inversion Does this mean you are not going to double check our submissions?",
    "438804": "It was not my intend to dispute, rather clarify. The results look just too good to be true (from my point of view). Though, I will wait for @marvelousninja feedback. Let's see, maybe it is possible.\n\nAs you noted, benchmarking code is indeed not easy. I was therefore really surprised when you changed the rules/template and relied on the users to correctly measure their code. If we just sticked with this (it is still in the speed prize tab):\n\n\n    import time\n    inference_start = time.time()\n    #### inference code ####\n    inference_end = time.time()\n    print('Inference Time: %0.2f Minutes'%((inference_end - inference_start)/60))\n\nwe would not be in the current situation.\n\nI agree with @ddanevskyi that, at this point, the best would be to make the kernels public.",
    "439000": "seesee Concerning the measurement of the time, I was the one to clarify/update the rules probably without full understanding of the consequences. We run our machine learning predictors on the cloud as micro-services (using docker containers) and they got the imagery sent to them inside so I am really concerned about the speed of the prediction and not the speed of I/O or startup up of the containers. This is why I proposed to exclude some parts to fit more precisely our use case. I hope that this will not cause too much trouble and that everyone will feel comfortable with the outcome of the speed prize challenge.",
    "439029": "Hi everyone!\nI've spent a couple of nights checking the timings.\n\n**TL; DR** - We measured time incorrectly. Wrapping the code with .time() calls is not enough.\nI have assembled a kernel that should explain what happened:\n\nhttps://www.kaggle.com/marvelousninja/pytorch-synchronization-issue-1-0\n\nFor the kernel with SE-ResNet50, it resulted in approximately 120 seconds gain.\nWith synchronization enabled, it brings the time to 5.5 minutes.\n\nConsidering the nature of the bug, I suggested to use a different kernel as my best submission.\nWith synchronization enabled, it has a time of 3.96 minutes, which is still enough to win.\nIt uses ResNet-18 as a classifier on 224x224 images and LinkNet for segmentation on 768x768 images.\n\nAs a show of good sportsmanship, I'm ready to publish this kernel.\nOf course, if @inversion and @jeffaudi approve this.",
    "439041": "marvelousninja No problem from my side and @inversion to publish your kernel.",
    "439076": "Ok, here is the kernel!\nhttps://www.kaggle.com/marvelousninja/kaggle-airbus-3-96-minutes",
    "439159": "Hello, I just read this topic here, @marvelousninja, and as I understood you did not submit the correct kernel in time. \nSo \"as a show of good sportsmanship\" after two weeks after this competition a submission of a new kernel with the correct timing is NOT \"enough to win\", unfortunately.  \n\nThe rules are clear: November 30, 2018 - Final submission deadline for speed prize kernel submission.\n\nPlease, make rules clear and fair for all of the applicants, Kaggle platform is for that, @inversion!  And also, take into a consideration just kernels that were submitted before the deadline. @jeffaudi",
    "439377": "jeffaudi @inversion\n\nFYI, our initial submission (4.69) takes only 2.90 to run. We did **NOT** change anything in the model, only added `torch.backends.cudnn.benchmark = True` and warmup rounds before the predictions loops.",
    "439406": "&gt; With synchronization enabled, it brings the time to 5.5 minutes.\n\nIt is great that you fixed the bug. Nevertheless, this time is only the 3rd place (assuming correct timings). Team [attention heads] and my kernel would be faster.\n \n&gt; We did NOT change anything in the model only added torch.backends.cudnn.benchmark = True and warmup rounds before the predictions loops.\n\nHow is this \"not change anything\" :) ?. To me this sounds like a couple of lines of code...\n\nAnyway, I don't see this is leading to an outcome \"everyone will feel comfortable with\". I'd be fine to call it a day and split the prize. I see 3 great solutions to the problem of detecting ships as fast as possible. The hosts can choose what they want to put into production. Please let me know what you think.",
    "456794": "can you share your code?"
  },
  "source": "meta"
}