{
  "id": 71664,
  "title": "14th place solution: data and metric comprehension",
  "url": "/competitions/airbus-ship-detection/discussion/71664",
  "author_name": "robga",
  "post_date": "2018-11-15T12:30:55.258000",
  "votes": 35,
  "comment_count": 10,
  "views": 0,
  "content": "<p>My team and I are happy to achieve a 14th place. Congratulations to everyone else in medal positions! And many thanks to my team mates <a href=\"https://www.kaggle.com/radek1\">radek</a> and <a href=\"https://www.kaggle.com/keremt\">kerem</a> for inspirational collaboration.</p>\n\n<p>This is my first attempt at a Kaggle competition with tiers/points. 12 months ago I had not heard of kaggle or pytorch, deep learning was a mystery, and python a snake, so it feels gratifying to do well. For that, I must thank both <em>kaggle</em> and <em>fast.ai</em> for creating communities and tools that allow autodidacts such as myself with a single GPU to progress so rapidly, and sponsors like Airbus for being so brave as to be open with data. I give back by explaining our solution ...</p>\n\n<p><strong>The TL;DR; Solution</strong></p>\n\n<p>The solution is based on the idea of</p>\n\n<ul>\n<li>careful data preparation (read: exploit data issues)</li>\n<li>understanding the balance between false positive rate and segmentation score (read: exploit the metric)</li>\n<li>local validation (read: trust local results not public scores)</li>\n<li>likelihood of the public leaderboard and private leaderboard differing in score-effecting ways (read: exploit LB carving)</li>\n<li>using dead simple out of the box models until the above had been exhausted</li>\n</ul>\n\n<p><strong>Data Prep</strong></p>\n\n<p>I was glad to learn the train data wasn't being re-released as the overlaps and leaks within it created opportunity. I created mosaics slightly differently to others, by using the train-masks rather than the train-images, which possibly made it much simpler and perhaps even better as these smaller ship-dense mosaics, separated by sea, could be stratified more homogeneously. I couldn't find tools to do this, I just used python hash values of the masks and re-assembled tiles by position. A 768px window was slid over mosaics in 256px steps to create a 50k image train set, striped into 5 folds.</p>\n\n<p><strong>Binary Classifier</strong></p>\n\n<p>A simple 256 resnet34 classifier, best used at the end not the start of a prediction pipeline.</p>\n\n<p><strong>Segmentation</strong></p>\n\n<p>A rudimentary resnet34 encoded unet, direct from <a href=\"https://www.fast.ai/2018/10/02/fastai-ai/\">fast.ai v1</a>, which uses Leslie Smith's   <a href=\"https://arxiv.org/abs/1803.09820\">1cycle policy</a> for learning. Training directly on 768px for 4 frozen then 32+ unfrozen epochs, small batch size of 4.  Dihedral augmentation with some limited brightness/contrast augmentation in training.</p>\n\n<p>I used a single GPU with only 8GB memory use, albeit on a 2080 ti with mixed precision training (thanks again to fast.ai v1) and 16 hour training runs. With this and the good data, single folds of 0.849 were achieved. Evidently a 5-fold CV didn't help any on the private LB despite big public LB differences (.729-.741).</p>\n\n<p>Our team ensembled this with resnet18 and resnext50-se models for a 0.001-2 boost. Our one regret is we didn't explore these ensembles further, due to public LB disappointment, and lack of time for local validation examination, but I am confident higher scores could have been achieved. Our best ensemble submission would have scored 11th if selected.</p>\n\n<p>We used dice loss, focal loss, BCE loss, and mixed versions, and do not consider the choice of loss function to have been significant.</p>\n\n<p><strong>Post processing</strong></p>\n\n<p>The intuition here was that the severe penalty (0 score) for labelling a ship where there is none meant it was necessary to sieve the results based on ship size, f2 score for that size, number of ships, circularity, whether the blob was on an edge, and classifier confidence for the image so as to avoid the penalty. </p>\n\n<p>And that the difference in the ship-laden public LB and the ship-barren private LB created an opportunity, albeit a simple algebraic one yet one that seems to have passed many people by.</p>\n\n<p>Ships &lt;100px had f2 scores of just .2-.3 in validation meaning 1 false positive image in every 4-5 would wipe out any gain. Whereas ships &gt; 5000px had an f2 of .7 and are unlikely to trigger a false positive. So in local validation we took the c2900 ship images that the segmenter said existed and used a crude heuristic to sieve out c500 ships: small ships if not very confident (.99+) with the binary classifier, and medium ships if not somewhat confident (.95+). We only submitted predictions for 2400 ie 2/3 of the 3700 images with ships (from the .765 private LB).</p>\n\n<p>While the curious data situation and puzzling LB split created an opportunity and helped us achieve a good result in this case, I plea for well provisioned data and more representative LBs going forward, if only to preserve sanity.</p>\n\n<p><strong>Attempted but didn't seem to help</strong></p>\n\n<ul>\n<li>More complex models didn't seem to help as much as understanding the data, the metric, and local validation results did.</li>\n<li>We trained a 'coast and structures' classifier </li>\n<li>We trained a segmentation model for 20-100px ships</li>\n<li>A <a href=\"http://www.itfind.or.kr/Report01/200302/IITA/IITA-2015-037/IITA-2015-037.pdf\">Circle Frequency Filter</a> gave cool results on images, but didn't have time to include in a model and to refine ship widths in post processing </li>\n<li>Fitting BB's, going so far as to try IOU measured BBs over blobs.</li>\n</ul>\n\n<p><strong>Things we didn't have time to explore enough that may have helped</strong></p>\n\n<ul>\n<li>Varying the pixel selection threshold away from 0.5 for different sized ships</li>\n<li>Ensembling strategies</li>\n<li>Stacking ship characteristics with segmentation predictions</li>\n<li>Other mask strategies e.g. trimaps</li>\n</ul>\n\n<p><strong>What wasn't attempted</strong></p>\n\n<ul>\n<li>More test time augmentation</li>\n<li>Breaking apart conjoined ships</li>\n<li>A 'wake detector' that may have helped binary classification of very small sub 20px ships.</li>\n<li>Probing to see a ship-size makeup of the public LB</li>\n</ul>\n\n<p><strong>fast.ai</strong></p>\n\n<p>As an addendum, I can't sing the praises of fast.ai enough. It really does make deep learning 'uncool' and accelerates experimentation around problem solving, not model making. No need to be a programming whiz. No need to subscribe to paid DL services. Indeed I don't need to make any code public, as the training needed to get 14th place was as simple as:</p>\n\n<pre><code>#prep data\n#define unet learner\nlearn = get_learner(data)\nlr=2e-3\nlearn.freeze_to(1)\nlearn.fit_one_cycle(4, lr, div_factor=100, pct_start=.3)\nlearn.unfreeze()\nlearn.fit_one_cycle(32, [lr/64,lr/8,lr], div_factor=25, pct_start=.3)\n</code></pre>",
  "messages": [
    {
      "id": 421792,
      "postDate": "2018-11-15T12:30:55.257Z",
      "content": "<p>My team and I are happy to achieve a 14th place. Congratulations to everyone else in medal positions! And many thanks to my team mates <a href=\"https://www.kaggle.com/radek1\">radek</a> and <a href=\"https://www.kaggle.com/keremt\">kerem</a> for inspirational collaboration.</p>\n\n<p>This is my first attempt at a Kaggle competition with tiers/points. 12 months ago I had not heard of kaggle or pytorch, deep learning was a mystery, and python a snake, so it feels gratifying to do well. For that, I must thank both <em>kaggle</em> and <em>fast.ai</em> for creating communities and tools that allow autodidacts such as myself with a single GPU to progress so rapidly, and sponsors like Airbus for being so brave as to be open with data. I give back by explaining our solution ...</p>\n\n<p><strong>The TL;DR; Solution</strong></p>\n\n<p>The solution is based on the idea of</p>\n\n<ul>\n<li>careful data preparation (read: exploit data issues)</li>\n<li>understanding the balance between false positive rate and segmentation score (read: exploit the metric)</li>\n<li>local validation (read: trust local results not public scores)</li>\n<li>likelihood of the public leaderboard and private leaderboard differing in score-effecting ways (read: exploit LB carving)</li>\n<li>using dead simple out of the box models until the above had been exhausted</li>\n</ul>\n\n<p><strong>Data Prep</strong></p>\n\n<p>I was glad to learn the train data wasn't being re-released as the overlaps and leaks within it created opportunity. I created mosaics slightly differently to others, by using the train-masks rather than the train-images, which possibly made it much simpler and perhaps even better as these smaller ship-dense mosaics, separated by sea, could be stratified more homogeneously. I couldn't find tools to do this, I just used python hash values of the masks and re-assembled tiles by position. A 768px window was slid over mosaics in 256px steps to create a 50k image train set, striped into 5 folds.</p>\n\n<p><strong>Binary Classifier</strong></p>\n\n<p>A simple 256 resnet34 classifier, best used at the end not the start of a prediction pipeline.</p>\n\n<p><strong>Segmentation</strong></p>\n\n<p>A rudimentary resnet34 encoded unet, direct from <a href=\"https://www.fast.ai/2018/10/02/fastai-ai/\">fast.ai v1</a>, which uses Leslie Smith's   <a href=\"https://arxiv.org/abs/1803.09820\">1cycle policy</a> for learning. Training directly on 768px for 4 frozen then 32+ unfrozen epochs, small batch size of 4.  Dihedral augmentation with some limited brightness/contrast augmentation in training.</p>\n\n<p>I used a single GPU with only 8GB memory use, albeit on a 2080 ti with mixed precision training (thanks again to fast.ai v1) and 16 hour training runs. With this and the good data, single folds of 0.849 were achieved. Evidently a 5-fold CV didn't help any on the private LB despite big public LB differences (.729-.741).</p>\n\n<p>Our team ensembled this with resnet18 and resnext50-se models for a 0.001-2 boost. Our one regret is we didn't explore these ensembles further, due to public LB disappointment, and lack of time for local validation examination, but I am confident higher scores could have been achieved. Our best ensemble submission would have scored 11th if selected.</p>\n\n<p>We used dice loss, focal loss, BCE loss, and mixed versions, and do not consider the choice of loss function to have been significant.</p>\n\n<p><strong>Post processing</strong></p>\n\n<p>The intuition here was that the severe penalty (0 score) for labelling a ship where there is none meant it was necessary to sieve the results based on ship size, f2 score for that size, number of ships, circularity, whether the blob was on an edge, and classifier confidence for the image so as to avoid the penalty. </p>\n\n<p>And that the difference in the ship-laden public LB and the ship-barren private LB created an opportunity, albeit a simple algebraic one yet one that seems to have passed many people by.</p>\n\n<p>Ships &lt;100px had f2 scores of just .2-.3 in validation meaning 1 false positive image in every 4-5 would wipe out any gain. Whereas ships &gt; 5000px had an f2 of .7 and are unlikely to trigger a false positive. So in local validation we took the c2900 ship images that the segmenter said existed and used a crude heuristic to sieve out c500 ships: small ships if not very confident (.99+) with the binary classifier, and medium ships if not somewhat confident (.95+). We only submitted predictions for 2400 ie 2/3 of the 3700 images with ships (from the .765 private LB).</p>\n\n<p>While the curious data situation and puzzling LB split created an opportunity and helped us achieve a good result in this case, I plea for well provisioned data and more representative LBs going forward, if only to preserve sanity.</p>\n\n<p><strong>Attempted but didn't seem to help</strong></p>\n\n<ul>\n<li>More complex models didn't seem to help as much as understanding the data, the metric, and local validation results did.</li>\n<li>We trained a 'coast and structures' classifier </li>\n<li>We trained a segmentation model for 20-100px ships</li>\n<li>A <a href=\"http://www.itfind.or.kr/Report01/200302/IITA/IITA-2015-037/IITA-2015-037.pdf\">Circle Frequency Filter</a> gave cool results on images, but didn't have time to include in a model and to refine ship widths in post processing </li>\n<li>Fitting BB's, going so far as to try IOU measured BBs over blobs.</li>\n</ul>\n\n<p><strong>Things we didn't have time to explore enough that may have helped</strong></p>\n\n<ul>\n<li>Varying the pixel selection threshold away from 0.5 for different sized ships</li>\n<li>Ensembling strategies</li>\n<li>Stacking ship characteristics with segmentation predictions</li>\n<li>Other mask strategies e.g. trimaps</li>\n</ul>\n\n<p><strong>What wasn't attempted</strong></p>\n\n<ul>\n<li>More test time augmentation</li>\n<li>Breaking apart conjoined ships</li>\n<li>A 'wake detector' that may have helped binary classification of very small sub 20px ships.</li>\n<li>Probing to see a ship-size makeup of the public LB</li>\n</ul>\n\n<p><strong>fast.ai</strong></p>\n\n<p>As an addendum, I can't sing the praises of fast.ai enough. It really does make deep learning 'uncool' and accelerates experimentation around problem solving, not model making. No need to be a programming whiz. No need to subscribe to paid DL services. Indeed I don't need to make any code public, as the training needed to get 14th place was as simple as:</p>\n\n<pre><code>#prep data\n#define unet learner\nlearn = get_learner(data)\nlr=2e-3\nlearn.freeze_to(1)\nlearn.fit_one_cycle(4, lr, div_factor=100, pct_start=.3)\nlearn.unfreeze()\nlearn.fit_one_cycle(32, [lr/64,lr/8,lr], div_factor=25, pct_start=.3)\n</code></pre>",
      "rawMarkdown": "My team and I are happy to achieve a 14th place. Congratulations to everyone else in medal positions! And many thanks to my team mates [radek][3] and [kerem][4] for inspirational collaboration.\n\nThis is my first attempt at a Kaggle competition with tiers/points. 12 months ago I had not heard of kaggle or pytorch, deep learning was a mystery, and python a snake, so it feels gratifying to do well. For that, I must thank both *kaggle* and *fast.ai* for creating communities and tools that allow autodidacts such as myself with a single GPU to progress so rapidly, and sponsors like Airbus for being so brave as to be open with data. I give back by explaining our solution ...\n\n**The TL;DR; Solution**\n\nThe solution is based on the idea of\n\n - careful data preparation (read: exploit data issues)\n - understanding the balance between false positive rate and segmentation score (read: exploit the metric)\n - local validation (read: trust local results not public scores)\n - likelihood of the public leaderboard and private leaderboard differing in score-effecting ways (read: exploit LB carving)\n - using dead simple out of the box models until the above had been exhausted\n\n**Data Prep**\n\nI was glad to learn the train data wasn't being re-released as the overlaps and leaks within it created opportunity. I created mosaics slightly differently to others, by using the train-masks rather than the train-images, which possibly made it much simpler and perhaps even better as these smaller ship-dense mosaics, separated by sea, could be stratified more homogeneously. I couldn't find tools to do this, I just used python hash values of the masks and re-assembled tiles by position. A 768px window was slid over mosaics in 256px steps to create a 50k image train set, striped into 5 folds.\n\n**Binary Classifier**\n\nA simple 256 resnet34 classifier, best used at the end not the start of a prediction pipeline.\n\n**Segmentation**\n\nA rudimentary resnet34 encoded unet, direct from [fast.ai v1][1], which uses Leslie Smith's   [1cycle policy][2] for learning. Training directly on 768px for 4 frozen then 32+ unfrozen epochs, small batch size of 4.  Dihedral augmentation with some limited brightness/contrast augmentation in training.\n\nI used a single GPU with only 8GB memory use, albeit on a 2080 ti with mixed precision training (thanks again to fast.ai v1) and 16 hour training runs. With this and the good data, single folds of 0.849 were achieved. Evidently a 5-fold CV didn't help any on the private LB despite big public LB differences (.729-.741).\n\nOur team ensembled this with resnet18 and resnext50-se models for a 0.001-2 boost. Our one regret is we didn't explore these ensembles further, due to public LB disappointment, and lack of time for local validation examination, but I am confident higher scores could have been achieved. Our best ensemble submission would have scored 11th if selected.\n\nWe used dice loss, focal loss, BCE loss, and mixed versions, and do not consider the choice of loss function to have been significant.\n\n**Post processing**\n\nThe intuition here was that the severe penalty (0 score) for labelling a ship where there is none meant it was necessary to sieve the results based on ship size, f2 score for that size, number of ships, circularity, whether the blob was on an edge, and classifier confidence for the image so as to avoid the penalty. \n\nAnd that the difference in the ship-laden public LB and the ship-barren private LB created an opportunity, albeit a simple algebraic one yet one that seems to have passed many people by.\n\nShips &lt;100px had f2 scores of just .2-.3 in validation meaning 1 false positive image in every 4-5 would wipe out any gain. Whereas ships &gt; 5000px had an f2 of .7 and are unlikely to trigger a false positive. So in local validation we took the c2900 ship images that the segmenter said existed and used a crude heuristic to sieve out c500 ships: small ships if not very confident (.99+) with the binary classifier, and medium ships if not somewhat confident (.95+). We only submitted predictions for 2400 ie 2/3 of the 3700 images with ships (from the .765 private LB).\n\nWhile the curious data situation and puzzling LB split created an opportunity and helped us achieve a good result in this case, I plea for well provisioned data and more representative LBs going forward, if only to preserve sanity.\n\n**Attempted but didn't seem to help**\n\n- More complex models didn't seem to help as much as understanding the data, the metric, and local validation results did.\n- We trained a 'coast and structures' classifier \n- We trained a segmentation model for 20-100px ships\n- A [Circle Frequency Filter][5] gave cool results on images, but didn't have time to include in a model and to refine ship widths in post processing \n- Fitting BB's, going so far as to try IOU measured BBs over blobs.\n\n**Things we didn't have time to explore enough that may have helped**\n\n- Varying the pixel selection threshold away from 0.5 for different sized ships\n- Ensembling strategies\n- Stacking ship characteristics with segmentation predictions\n- Other mask strategies e.g. trimaps\n\n**What wasn't attempted**\n\n- More test time augmentation\n- Breaking apart conjoined ships\n- A 'wake detector' that may have helped binary classification of very small sub 20px ships.\n- Probing to see a ship-size makeup of the public LB\n\n**fast.ai**\n\nAs an addendum, I can't sing the praises of fast.ai enough. It really does make deep learning 'uncool' and accelerates experimentation around problem solving, not model making. No need to be a programming whiz. No need to subscribe to paid DL services. Indeed I don't need to make any code public, as the training needed to get 14th place was as simple as:\n\n    #prep data\n    #define unet learner\n    learn = get_learner(data)\n    lr=2e-3\n    learn.freeze_to(1)\n    learn.fit_one_cycle(4, lr, div_factor=100, pct_start=.3)\n    learn.unfreeze()\n    learn.fit_one_cycle(32, [lr/64,lr/8,lr], div_factor=25, pct_start=.3)\n\n  [1]: https://www.fast.ai/2018/10/02/fastai-ai/\n  [2]: https://arxiv.org/abs/1803.09820\n  [3]: https://www.kaggle.com/radek1\n  [4]: https://www.kaggle.com/keremt\n  [5]: http://www.itfind.or.kr/Report01/200302/IITA/IITA-2015-037/IITA-2015-037.pdf",
      "votes": 35
    },
    {
      "id": 423787,
      "postDate": "2018-11-19T03:15:58.693Z",
      "content": "<p>Thank you for your detailed explanation and Congratulations to your team. would you like to open-source your code implementation of this solution? Thanks!</p>",
      "rawMarkdown": "Thank you for your detailed explanation and Congratulations to your team. would you like to open-source your code implementation of this solution? Thanks!",
      "votes": 12
    },
    {
      "id": 423626,
      "postDate": "2018-11-18T18:30:02.997Z",
      "content": "<p>Why are the div factors set to the numbers they are set at i.e 100,25?</p>",
      "rawMarkdown": "Why are the div factors set to the numbers they are set at i.e 100,25?",
      "votes": 1,
      "replies": [
        {
          "id": 423662,
          "postDate": "2018-11-18T20:10:26.113Z",
          "content": "<p>The div factor represents the min to max learning rate through a training cycle. With a max_lr of 2e-3 a div factor of 100 means a lr moving through 2e-5 to 2e-3 while 25 is 8e-5 through 2e-3. A small difference and frankly my starting point divs for frozen and unfrozen training. That's the beauty of 1cycle in my experience, making the lr hyperparameter choice less fragile than it may otherwise be.</p>",
          "rawMarkdown": "The div factor represents the min to max learning rate through a training cycle. With a max_lr of 2e-3 a div factor of 100 means a lr moving through 2e-5 to 2e-3 while 25 is 8e-5 through 2e-3. A small difference and frankly my starting point divs for frozen and unfrozen training. That's the beauty of 1cycle in my experience, making the lr hyperparameter choice less fragile than it may otherwise be.",
          "votes": 1
        }
      ]
    },
    {
      "id": 427689,
      "postDate": "2018-11-26T00:35:59.700Z",
      "content": "<p>Congrats <a href=\"/robga\">@robga</a>,  and thanks for sharing.</p>",
      "rawMarkdown": "Congrats @robga,  and thanks for sharing."
    },
    {
      "id": 423237,
      "postDate": "2018-11-17T18:39:44.500Z",
      "content": "<p>Great writeup! Love seeing fastai being used for great results. I see that in your training code you use different div factors in fit one cycle. How did you decide on these, and did they give you improvements over the default?</p>",
      "rawMarkdown": "Great writeup! Love seeing fastai being used for great results. I see that in your training code you use different div factors in fit one cycle. How did you decide on these, and did they give you improvements over the default?",
      "replies": [
        {
          "id": 423581,
          "postDate": "2018-11-18T15:49:23.033Z",
          "content": "<p>What would that mean div factor</p>",
          "rawMarkdown": "What would that mean div factor"
        }
      ]
    },
    {
      "id": 422045,
      "postDate": "2018-11-15T18:05:58.460Z",
      "content": "<p>hi nice solution..i  too used simple FAI only...\nCould you try to help  more understand \nOur team ensembled this with resnet18 and resnext50-se models for a 0.001-2 boost\nHow this can be achieved... i read about this word ensembling in many solutions.. m curious to know how we achieve this..</p>",
      "rawMarkdown": "hi nice solution..i  too used simple FAI only...\nCould you try to help  more understand \nOur team ensembled this with resnet18 and resnext50-se models for a 0.001-2 boost\nHow this can be achieved... i read about this word ensembling in many solutions.. m curious to know how we achieve this..\n",
      "replies": [
        {
          "id": 422079,
          "postDate": "2018-11-15T18:58:32.060Z",
          "content": "<p>Ensembling simply means combining multiple predictions, for instance those made with different models. You can take the average of many predictions, weight averages, take the median, or employ various voting strategies. Different ensembling techniques work in different circumstances. </p>",
          "rawMarkdown": "Ensembling simply means combining multiple predictions, for instance those made with different models. You can take the average of many predictions, weight averages, take the median, or employ various voting strategies. Different ensembling techniques work in different circumstances. ",
          "votes": 1
        },
        {
          "id": 423337,
          "postDate": "2018-11-18T00:22:14.197Z",
          "content": "<p>First Congrats for the results. But how Can we use Ensembling for a segmentation problem.</p>",
          "rawMarkdown": "First Congrats for the results. But how Can we use Ensembling for a segmentation problem.",
          "votes": 1
        }
      ]
    },
    {
      "id": 423223,
      "postDate": "2018-11-17T18:14:26.593Z",
      "rawMarkdown": "",
      "votes": 7,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 423787,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-19T03:15:58.693000",
      "content": "<p>Thank you for your detailed explanation and Congratulations to your team. would you like to open-source your code implementation of this solution? Thanks!</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 423626,
      "author_name": "Rahul Deora",
      "author_url": "",
      "post_date": "2018-11-18T18:30:02.997000",
      "content": "<p>Why are the div factors set to the numbers they are set at i.e 100,25?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 423662,
          "author_name": "robga",
          "author_url": "",
          "post_date": "2018-11-18T20:10:26.113000",
          "content": "<p>The div factor represents the min to max learning rate through a training cycle. With a max_lr of 2e-3 a div factor of 100 means a lr moving through 2e-5 to 2e-3 while 25 is 8e-5 through 2e-3. A small difference and frankly my starting point divs for frozen and unfrozen training. That's the beauty of 1cycle in my experience, making the lr hyperparameter choice less fragile than it may otherwise be.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 427689,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2018-11-26T00:35:59.700000",
      "content": "<p>Congrats <a href=\"/robga\">@robga</a>,  and thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 423237,
      "author_name": "William Horton",
      "author_url": "",
      "post_date": "2018-11-17T18:39:44.500000",
      "content": "<p>Great writeup! Love seeing fastai being used for great results. I see that in your training code you use different div factors in fit one cycle. How did you decide on these, and did they give you improvements over the default?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 423581,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2018-11-18T15:49:23.033000",
          "content": "<p>What would that mean div factor</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 422045,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2018-11-15T18:05:58.460000",
      "content": "<p>hi nice solution..i  too used simple FAI only...\nCould you try to help  more understand \nOur team ensembled this with resnet18 and resnext50-se models for a 0.001-2 boost\nHow this can be achieved... i read about this word ensembling in many solutions.. m curious to know how we achieve this..</p>",
      "votes": 0,
      "replies": [
        {
          "id": 422079,
          "author_name": "robga",
          "author_url": "",
          "post_date": "2018-11-15T18:58:32.060000",
          "content": "<p>Ensembling simply means combining multiple predictions, for instance those made with different models. You can take the average of many predictions, weight averages, take the median, or employ various voting strategies. Different ensembling techniques work in different circumstances. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 423337,
          "author_name": "JM100",
          "author_url": "",
          "post_date": "2018-11-18T00:22:14.197000",
          "content": "<p>First Congrats for the results. But how Can we use Ensembling for a segmentation problem.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 423223,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-17T18:14:26.593000",
      "content": "",
      "votes": 7,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "421792": "My team and I are happy to achieve a 14th place. Congratulations to everyone else in medal positions! And many thanks to my team mates [radek][3] and [kerem][4] for inspirational collaboration.\n\nThis is my first attempt at a Kaggle competition with tiers/points. 12 months ago I had not heard of kaggle or pytorch, deep learning was a mystery, and python a snake, so it feels gratifying to do well. For that, I must thank both *kaggle* and *fast.ai* for creating communities and tools that allow autodidacts such as myself with a single GPU to progress so rapidly, and sponsors like Airbus for being so brave as to be open with data. I give back by explaining our solution ...\n\n**The TL;DR; Solution**\n\nThe solution is based on the idea of\n\n - careful data preparation (read: exploit data issues)\n - understanding the balance between false positive rate and segmentation score (read: exploit the metric)\n - local validation (read: trust local results not public scores)\n - likelihood of the public leaderboard and private leaderboard differing in score-effecting ways (read: exploit LB carving)\n - using dead simple out of the box models until the above had been exhausted\n\n**Data Prep**\n\nI was glad to learn the train data wasn't being re-released as the overlaps and leaks within it created opportunity. I created mosaics slightly differently to others, by using the train-masks rather than the train-images, which possibly made it much simpler and perhaps even better as these smaller ship-dense mosaics, separated by sea, could be stratified more homogeneously. I couldn't find tools to do this, I just used python hash values of the masks and re-assembled tiles by position. A 768px window was slid over mosaics in 256px steps to create a 50k image train set, striped into 5 folds.\n\n**Binary Classifier**\n\nA simple 256 resnet34 classifier, best used at the end not the start of a prediction pipeline.\n\n**Segmentation**\n\nA rudimentary resnet34 encoded unet, direct from [fast.ai v1][1], which uses Leslie Smith's   [1cycle policy][2] for learning. Training directly on 768px for 4 frozen then 32+ unfrozen epochs, small batch size of 4.  Dihedral augmentation with some limited brightness/contrast augmentation in training.\n\nI used a single GPU with only 8GB memory use, albeit on a 2080 ti with mixed precision training (thanks again to fast.ai v1) and 16 hour training runs. With this and the good data, single folds of 0.849 were achieved. Evidently a 5-fold CV didn't help any on the private LB despite big public LB differences (.729-.741).\n\nOur team ensembled this with resnet18 and resnext50-se models for a 0.001-2 boost. Our one regret is we didn't explore these ensembles further, due to public LB disappointment, and lack of time for local validation examination, but I am confident higher scores could have been achieved. Our best ensemble submission would have scored 11th if selected.\n\nWe used dice loss, focal loss, BCE loss, and mixed versions, and do not consider the choice of loss function to have been significant.\n\n**Post processing**\n\nThe intuition here was that the severe penalty (0 score) for labelling a ship where there is none meant it was necessary to sieve the results based on ship size, f2 score for that size, number of ships, circularity, whether the blob was on an edge, and classifier confidence for the image so as to avoid the penalty. \n\nAnd that the difference in the ship-laden public LB and the ship-barren private LB created an opportunity, albeit a simple algebraic one yet one that seems to have passed many people by.\n\nShips &lt;100px had f2 scores of just .2-.3 in validation meaning 1 false positive image in every 4-5 would wipe out any gain. Whereas ships &gt; 5000px had an f2 of .7 and are unlikely to trigger a false positive. So in local validation we took the c2900 ship images that the segmenter said existed and used a crude heuristic to sieve out c500 ships: small ships if not very confident (.99+) with the binary classifier, and medium ships if not somewhat confident (.95+). We only submitted predictions for 2400 ie 2/3 of the 3700 images with ships (from the .765 private LB).\n\nWhile the curious data situation and puzzling LB split created an opportunity and helped us achieve a good result in this case, I plea for well provisioned data and more representative LBs going forward, if only to preserve sanity.\n\n**Attempted but didn't seem to help**\n\n- More complex models didn't seem to help as much as understanding the data, the metric, and local validation results did.\n- We trained a 'coast and structures' classifier \n- We trained a segmentation model for 20-100px ships\n- A [Circle Frequency Filter][5] gave cool results on images, but didn't have time to include in a model and to refine ship widths in post processing \n- Fitting BB's, going so far as to try IOU measured BBs over blobs.\n\n**Things we didn't have time to explore enough that may have helped**\n\n- Varying the pixel selection threshold away from 0.5 for different sized ships\n- Ensembling strategies\n- Stacking ship characteristics with segmentation predictions\n- Other mask strategies e.g. trimaps\n\n**What wasn't attempted**\n\n- More test time augmentation\n- Breaking apart conjoined ships\n- A 'wake detector' that may have helped binary classification of very small sub 20px ships.\n- Probing to see a ship-size makeup of the public LB\n\n**fast.ai**\n\nAs an addendum, I can't sing the praises of fast.ai enough. It really does make deep learning 'uncool' and accelerates experimentation around problem solving, not model making. No need to be a programming whiz. No need to subscribe to paid DL services. Indeed I don't need to make any code public, as the training needed to get 14th place was as simple as:\n\n    #prep data\n    #define unet learner\n    learn = get_learner(data)\n    lr=2e-3\n    learn.freeze_to(1)\n    learn.fit_one_cycle(4, lr, div_factor=100, pct_start=.3)\n    learn.unfreeze()\n    learn.fit_one_cycle(32, [lr/64,lr/8,lr], div_factor=25, pct_start=.3)\n\n  [1]: https://www.fast.ai/2018/10/02/fastai-ai/\n  [2]: https://arxiv.org/abs/1803.09820\n  [3]: https://www.kaggle.com/radek1\n  [4]: https://www.kaggle.com/keremt\n  [5]: http://www.itfind.or.kr/Report01/200302/IITA/IITA-2015-037/IITA-2015-037.pdf",
    "423787": "Thank you for your detailed explanation and Congratulations to your team. would you like to open-source your code implementation of this solution? Thanks!",
    "423626": "Why are the div factors set to the numbers they are set at i.e 100,25?",
    "427689": "Congrats @robga,  and thanks for sharing.",
    "423237": "Great writeup! Love seeing fastai being used for great results. I see that in your training code you use different div factors in fit one cycle. How did you decide on these, and did they give you improvements over the default?",
    "422045": "hi nice solution..i  too used simple FAI only...\nCould you try to help  more understand \nOur team ensembled this with resnet18 and resnext50-se models for a 0.001-2 boost\nHow this can be achieved... i read about this word ensembling in many solutions.. m curious to know how we achieve this..\n",
    "423223": ""
  }
}