{
  "id": 82994,
  "title": "87 place solution - thanks Martin",
  "url": "/competitions/humpback-whale-identification/writeups/ods-ai-borys-tymchenko-87-place-solution-thanks-ma",
  "author_name": "",
  "post_date": "2019-03-05T21:01:51.914126700Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Congratulations to all the winners!\nThaks all for the meaningful discussions!</p>\n\n<p>At first, I tried to use Martin's awesome kernel as is just to replicate the results but did not succeded to outperform the pretrained model.\nAfter it, I experimented with triplet loss and semi-hard negative mining. This gave me a score ~0.5 along with very long training times (Keras), so I abandoned this idea.\nLater on, I was using Martin's kernel as an approach backbone and modified it.</p>\n\n<p>To reduce training time, I trained encoder as a classification model first. \nEmbedding layer was L2-normalized and then scaled with learnable parameter. To make embeddings more compact, I added L1L2 regularization to scale parameter.\nHead was taken from Martin's kernel but with bigger internal dimension.\nIn Martin's kernel I modified schedule of random values to linear decreasing from 0.75 to 0.15 over 50 epochs.</p>\n\n<p>All models were trained without local validation :D</p>\n\n<h2>Things that worked</h2>\n\n<ul>\n<li>Ensembling</li>\n<li>BCE loss</li>\n<li>SGD with warm restarts</li>\n<li>Progressive resizing while training classification model</li>\n<li>Heavy augmentations (almost like in imgaug demo)</li>\n<li>SEResNet34/50</li>\n<li>Weight decay</li>\n<li>Label smoothing while finetuning</li>\n<li>Threshold selection based on 30% new whales fraction</li>\n<li>Linear layer after GAP</li>\n</ul>\n\n<h2>Things that did not work</h2>\n\n<ul>\n<li>DenseNets (were too memory-hungry)</li>\n<li>Focal loss</li>\n<li>Global Max Pooling (-0.05 LB)</li>\n<li>Convolutional block attention module</li>\n<li>Triplet loss (I was not trying that hard)</li>\n<li>Mixup (again)</li>\n<li>Light augmentations led to underfitting (but why?!)</li>\n<li>Half-fluke crops</li>\n<li>Images larger than 512x512</li>\n</ul>\n\n<h2>Top solution</h2>\n\n<p>Top performing solution scored 0.922 on private.\nTop individual model scored 0.904 on private.\nIt was an ensemble of 4 the best models: 2xSEResNet34 with 512x512 resolution and 2xSEResNet50 with 384x384 resolution.</p>\n\n<h2>Hardware</h2>\n\n<p>We used primarily the server with one 1080Ti and 64GB of RAM.\nTo reduce disk bottleneck, we put all training data into ramdisk with plenty of swap (~3x load speedup).</p>",
  "messages": [
    {
      "id": "484336",
      "postDate": "03/05/2019 21:01:51",
      "content": "<p>Congratulations to all the winners!\nThaks all for the meaningful discussions!</p>\n\n<p>At first, I tried to use Martin's awesome kernel as is just to replicate the results but did not succeded to outperform the pretrained model.\nAfter it, I experimented with triplet loss and semi-hard negative mining. This gave me a score ~0.5 along with very long training times (Keras), so I abandoned this idea.\nLater on, I was using Martin's kernel as an approach backbone and modified it.</p>\n\n<p>To reduce training time, I trained encoder as a classification model first. \nEmbedding layer was L2-normalized and then scaled with learnable parameter. To make embeddings more compact, I added L1L2 regularization to scale parameter.\nHead was taken from Martin's kernel but with bigger internal dimension.\nIn Martin's kernel I modified schedule of random values to linear decreasing from 0.75 to 0.15 over 50 epochs.</p>\n\n<p>All models were trained without local validation :D</p>\n\n<h2>Things that worked</h2>\n\n<ul>\n<li>Ensembling</li>\n<li>BCE loss</li>\n<li>SGD with warm restarts</li>\n<li>Progressive resizing while training classification model</li>\n<li>Heavy augmentations (almost like in imgaug demo)</li>\n<li>SEResNet34/50</li>\n<li>Weight decay</li>\n<li>Label smoothing while finetuning</li>\n<li>Threshold selection based on 30% new whales fraction</li>\n<li>Linear layer after GAP</li>\n</ul>\n\n<h2>Things that did not work</h2>\n\n<ul>\n<li>DenseNets (were too memory-hungry)</li>\n<li>Focal loss</li>\n<li>Global Max Pooling (-0.05 LB)</li>\n<li>Convolutional block attention module</li>\n<li>Triplet loss (I was not trying that hard)</li>\n<li>Mixup (again)</li>\n<li>Light augmentations led to underfitting (but why?!)</li>\n<li>Half-fluke crops</li>\n<li>Images larger than 512x512</li>\n</ul>\n\n<h2>Top solution</h2>\n\n<p>Top performing solution scored 0.922 on private.\nTop individual model scored 0.904 on private.\nIt was an ensemble of 4 the best models: 2xSEResNet34 with 512x512 resolution and 2xSEResNet50 with 384x384 resolution.</p>\n\n<h2>Hardware</h2>\n\n<p>We used primarily the server with one 1080Ti and 64GB of RAM.\nTo reduce disk bottleneck, we put all training data into ramdisk with plenty of swap (~3x load speedup).</p>",
      "rawMarkdown": "Congratulations to all the winners!\nThaks all for the meaningful discussions!\n\nAt first, I tried to use Martin's awesome kernel as is just to replicate the results but did not succeded to outperform the pretrained model.\nAfter it, I experimented with triplet loss and semi-hard negative mining. This gave me a score ~0.5 along with very long training times (Keras), so I abandoned this idea.\nLater on, I was using Martin's kernel as an approach backbone and modified it.\n\nTo reduce training time, I trained encoder as a classification model first. \nEmbedding layer was L2-normalized and then scaled with learnable parameter. To make embeddings more compact, I added L1L2 regularization to scale parameter.\nHead was taken from Martin's kernel but with bigger internal dimension.\nIn Martin's kernel I modified schedule of random values to linear decreasing from 0.75 to 0.15 over 50 epochs.\n\nAll models were trained without local validation :D\n\n## Things that worked ##\n- Ensembling\n- BCE loss\n- SGD with warm restarts\n- Progressive resizing while training classification model\n- Heavy augmentations (almost like in imgaug demo)\n- SEResNet34/50\n- Weight decay\n- Label smoothing while finetuning\n- Threshold selection based on 30% new whales fraction\n- Linear layer after GAP\n\n## Things that did not work ##\n- DenseNets (were too memory-hungry)\n- Focal loss\n- Global Max Pooling (-0.05 LB)\n- Convolutional block attention module\n- Triplet loss (I was not trying that hard)\n- Mixup (again)\n- Light augmentations led to underfitting (but why?!)\n- Half-fluke crops\n- Images larger than 512x512\n\n## Top solution ##\nTop performing solution scored 0.922 on private.\nTop individual model scored 0.904 on private.\nIt was an ensemble of 4 the best models: 2xSEResNet34 with 512x512 resolution and 2xSEResNet50 with 384x384 resolution.\n\n## Hardware ##\nWe used primarily the server with one 1080Ti and 64GB of RAM.\nTo reduce disk bottleneck, we put all training data into ramdisk with plenty of swap (~3x load speedup).",
      "votes": null
    },
    {
      "id": "484364",
      "postDate": "03/05/2019 22:12:16",
      "content": "<p>What's hard for me to grasp is: Why did the results from copying Martin's awesome kernel give such a wide variation in results?</p>",
      "rawMarkdown": "What's hard for me to grasp is: Why did the results from copying Martin's awesome kernel give such a wide variation in results?",
      "votes": null
    },
    {
      "id": "484560",
      "postDate": "03/06/2019 07:13:46",
      "content": "<p>Martin's kernel is really awesome idea-wise. Unfortunately, its code is not that good at all. \nAlso, its original version takes a lot of time to train from scratch and (probably) highly depends or random seed. My top result of just running it as-is was something like ~0.81 on public LB in 500 epochs (4 days on 1080 Ti!).</p>\n\n<p>Some people, probably took substitutions for quick LAP from discussions. And these substitutions are not the complete alternative to a original solution, so proper modifications to epoch schedule are also needed...</p>\n\n<p>The devil is in the detail.</p>",
      "rawMarkdown": "Martin's kernel is really awesome idea-wise. Unfortunately, its code is not that good at all. \nAlso, its original version takes a lot of time to train from scratch and (probably) highly depends or random seed. My top result of just running it as-is was something like ~0.81 on public LB in 500 epochs (4 days on 1080 Ti!).\n\nSome people, probably took substitutions for quick LAP from discussions. And these substitutions are not the complete alternative to a original solution, so proper modifications to epoch schedule are also needed...\n\nThe devil is in the detail.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 484364,
      "author_name": "chrisfreiling",
      "author_url": "",
      "post_date": "03/05/2019 22:12:16",
      "content": "<p>What's hard for me to grasp is: Why did the results from copying Martin's awesome kernel give such a wide variation in results?</p>",
      "votes": null,
      "replies": [
        {
          "id": 484560,
          "author_name": "spsancti",
          "author_url": "",
          "post_date": "03/06/2019 07:13:46",
          "content": "<p>Martin's kernel is really awesome idea-wise. Unfortunately, its code is not that good at all. \nAlso, its original version takes a lot of time to train from scratch and (probably) highly depends or random seed. My top result of just running it as-is was something like ~0.81 on public LB in 500 epochs (4 days on 1080 Ti!).</p>\n\n<p>Some people, probably took substitutions for quick LAP from discussions. And these substitutions are not the complete alternative to a original solution, so proper modifications to epoch schedule are also needed...</p>\n\n<p>The devil is in the detail.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "484336": "Congratulations to all the winners!\nThaks all for the meaningful discussions!\n\nAt first, I tried to use Martin's awesome kernel as is just to replicate the results but did not succeded to outperform the pretrained model.\nAfter it, I experimented with triplet loss and semi-hard negative mining. This gave me a score ~0.5 along with very long training times (Keras), so I abandoned this idea.\nLater on, I was using Martin's kernel as an approach backbone and modified it.\n\nTo reduce training time, I trained encoder as a classification model first. \nEmbedding layer was L2-normalized and then scaled with learnable parameter. To make embeddings more compact, I added L1L2 regularization to scale parameter.\nHead was taken from Martin's kernel but with bigger internal dimension.\nIn Martin's kernel I modified schedule of random values to linear decreasing from 0.75 to 0.15 over 50 epochs.\n\nAll models were trained without local validation :D\n\n## Things that worked ##\n- Ensembling\n- BCE loss\n- SGD with warm restarts\n- Progressive resizing while training classification model\n- Heavy augmentations (almost like in imgaug demo)\n- SEResNet34/50\n- Weight decay\n- Label smoothing while finetuning\n- Threshold selection based on 30% new whales fraction\n- Linear layer after GAP\n\n## Things that did not work ##\n- DenseNets (were too memory-hungry)\n- Focal loss\n- Global Max Pooling (-0.05 LB)\n- Convolutional block attention module\n- Triplet loss (I was not trying that hard)\n- Mixup (again)\n- Light augmentations led to underfitting (but why?!)\n- Half-fluke crops\n- Images larger than 512x512\n\n## Top solution ##\nTop performing solution scored 0.922 on private.\nTop individual model scored 0.904 on private.\nIt was an ensemble of 4 the best models: 2xSEResNet34 with 512x512 resolution and 2xSEResNet50 with 384x384 resolution.\n\n## Hardware ##\nWe used primarily the server with one 1080Ti and 64GB of RAM.\nTo reduce disk bottleneck, we put all training data into ramdisk with plenty of swap (~3x load speedup).",
    "484364": "What's hard for me to grasp is: Why did the results from copying Martin's awesome kernel give such a wide variation in results?",
    "484560": "Martin's kernel is really awesome idea-wise. Unfortunately, its code is not that good at all. \nAlso, its original version takes a lot of time to train from scratch and (probably) highly depends or random seed. My top result of just running it as-is was something like ~0.81 on public LB in 500 epochs (4 days on 1080 Ti!).\n\nSome people, probably took substitutions for quick LAP from discussions. And these substitutions are not the complete alternative to a original solution, so proper modifications to epoch schedule are also needed...\n\nThe devil is in the detail."
  },
  "source": "meta"
}