{
  "id": 94817,
  "title": "Solution | Private 4th",
  "url": "/competitions/imet-2019-fgvc6/writeups/pudae-solution-private-4th",
  "author_name": "",
  "post_date": "2019-06-11T16:51:19.293Z",
  "votes": 46,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Congrats to all the winners.\nThanks to Kaggle and hosting team for an interesting competition.</p>\n\n<p>Here is my solution summary.</p>\n\n<h3>Pure Classification Model</h3>\n\n<p>Because LB scores of our classification models are similar or worse with many other competitors, I briefly describe the configurations.</p>\n\n<ul>\n<li><strong>Dataset</strong>: 10 folds CV with Multilabel Iterative Stratification</li>\n<li><strong>Models</strong>: se-resnext101, senet154. After last convolution layer, one more fc layer is inserted.</li>\n<li><strong>Augmentation</strong>: random resized crop, horizontal flip.</li>\n<li><strong>Loss</strong>: class balanced focal loss. But it didn’t improve score compared to the pure focal loss.</li>\n<li><strong>Optimization</strong>: AdamW, \n<ul><li>learning rate 0.00025, weight decay 0.01 for se-resnext101.</li>\n<li>learning rate 0.0001, weight decay 0.05 for senet154</li></ul></li>\n<li><strong>Learning rate schedule</strong>: CosineAnnealingLR. T_max 15, eta_min 0.00001</li>\n<li>Input size 320x320. Batch size 24</li>\n</ul>\n\n<p>With above configurations, the LB score of single fold se-resnext101 with hflip TTA was 0.618.</p>\n\n<h3>Tag Relevance Prediction</h3>\n\n<p>We’ve tried several methods, but all were failed. So we’ve decided to improve prediction methods. \nWe’ve found the method based on the <a href=\"http://class.inrialpes.fr/pub/guillaumin-iccv09b.pdf\">paper</a> improves LB score significantly. </p>\n\n<p>The detailed are the following:</p>\n\n<p>For the training set, the presence probability of label <em>L</em> is defined as <em>1 – epsilon</em> if the example has label L, otherwise <em>epsilon</em>. The label presence probabilities of test example are the weighted sum over the nearest K train examples.</p>\n\n<p>The weights are determined based on the distances between test and train examples. As a distance metric, 1 – cosine similarity between train and test example is used. The output of the last convolution layer is used as embedding features.</p>\n\n<p>The weights of training example j for an test example i are defined as:</p>\n\n<p>&gt; \\(\\pi_{i,j}=exp(-d * distance(i,j))/\\sum exp(-d * distance(i,j')))\\)</p>\n\n<p>The probability of class <em>w</em> for test image <em>i</em> is:</p>\n\n<p>&gt; \\( p(y_{i,w}=1)=\\sum \\pi_{i,j} * p(y_{j,w})\\)</p>\n\n<p>The parameter d and K has been chosen based on the validation score.</p>\n\n<p>With this method, the LB score of single fold se-resnext101 with hflip TTA was 0.634.\nThe average of two predictions is LB 0.650.</p>\n\n<p><strong>Minor improvement</strong></p>\n\n<p>We’ve found the images that have no close images have a low recall. So, we choose thresholds according to the distance from the nearest train image. If it has a higher distance, a lower threshold is assigned. Using this method, LB was improved +0.001~0.002.</p>\n\n<p>Our final submission is an ensemble of 4 folds se-resnext101 and 6 folds senet154. </p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "546998",
      "postDate": "06/07/2019 05:42:41",
      "content": "<p>Congrats to all the winners.\nThanks to Kaggle and hosting team for an interesting competition.</p>\n\n<p>Here is my solution summary.</p>\n\n<h3>Pure Classification Model</h3>\n\n<p>Because LB scores of our classification models are similar or worse with many other competitors, I briefly describe the configurations.</p>\n\n<ul>\n<li><strong>Dataset</strong>: 10 folds CV with Multilabel Iterative Stratification</li>\n<li><strong>Models</strong>: se-resnext101, senet154. After last convolution layer, one more fc layer is inserted.</li>\n<li><strong>Augmentation</strong>: random resized crop, horizontal flip.</li>\n<li><strong>Loss</strong>: class balanced focal loss. But it didn’t improve score compared to the pure focal loss.</li>\n<li><strong>Optimization</strong>: AdamW, \n<ul><li>learning rate 0.00025, weight decay 0.01 for se-resnext101.</li>\n<li>learning rate 0.0001, weight decay 0.05 for senet154</li></ul></li>\n<li><strong>Learning rate schedule</strong>: CosineAnnealingLR. T_max 15, eta_min 0.00001</li>\n<li>Input size 320x320. Batch size 24</li>\n</ul>\n\n<p>With above configurations, the LB score of single fold se-resnext101 with hflip TTA was 0.618.</p>\n\n<h3>Tag Relevance Prediction</h3>\n\n<p>We’ve tried several methods, but all were failed. So we’ve decided to improve prediction methods. \nWe’ve found the method based on the <a href=\"http://class.inrialpes.fr/pub/guillaumin-iccv09b.pdf\">paper</a> improves LB score significantly. </p>\n\n<p>The detailed are the following:</p>\n\n<p>For the training set, the presence probability of label <em>L</em> is defined as <em>1 – epsilon</em> if the example has label L, otherwise <em>epsilon</em>. The label presence probabilities of test example are the weighted sum over the nearest K train examples.</p>\n\n<p>The weights are determined based on the distances between test and train examples. As a distance metric, 1 – cosine similarity between train and test example is used. The output of the last convolution layer is used as embedding features.</p>\n\n<p>The weights of training example j for an test example i are defined as:</p>\n\n<p>&gt; \\(\\pi_{i,j}=exp(-d * distance(i,j))/\\sum exp(-d * distance(i,j')))\\)</p>\n\n<p>The probability of class <em>w</em> for test image <em>i</em> is:</p>\n\n<p>&gt; \\( p(y_{i,w}=1)=\\sum \\pi_{i,j} * p(y_{j,w})\\)</p>\n\n<p>The parameter d and K has been chosen based on the validation score.</p>\n\n<p>With this method, the LB score of single fold se-resnext101 with hflip TTA was 0.634.\nThe average of two predictions is LB 0.650.</p>\n\n<p><strong>Minor improvement</strong></p>\n\n<p>We’ve found the images that have no close images have a low recall. So, we choose thresholds according to the distance from the nearest train image. If it has a higher distance, a lower threshold is assigned. Using this method, LB was improved +0.001~0.002.</p>\n\n<p>Our final submission is an ensemble of 4 folds se-resnext101 and 6 folds senet154. </p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Congrats to all the winners.\nThanks to Kaggle and hosting team for an interesting competition.\n\nHere is my solution summary.\n\n### Pure Classification Model\nBecause LB scores of our classification models are similar or worse with many other competitors, I briefly describe the configurations.\n\n-\t**Dataset**: 10 folds CV with Multilabel Iterative Stratification\n-\t**Models**: se-resnext101, senet154. After last convolution layer, one more fc layer is inserted.\n-\t**Augmentation**: random resized crop, horizontal flip.\n-\t**Loss**: class balanced focal loss. But it didn’t improve score compared to the pure focal loss.\n-\t**Optimization**: AdamW, \n   - learning rate 0.00025, weight decay 0.01 for se-resnext101.\n   - learning rate 0.0001, weight decay 0.05 for senet154\n-\t**Learning rate schedule**: CosineAnnealingLR. T_max 15, eta_min 0.00001\n-\tInput size 320x320. Batch size 24\n\nWith above configurations, the LB score of single fold se-resnext101 with hflip TTA was 0.618.\n\n\n### Tag Relevance Prediction\n\nWe’ve tried several methods, but all were failed. So we’ve decided to improve prediction methods. \nWe’ve found the method based on the [paper](http://class.inrialpes.fr/pub/guillaumin-iccv09b.pdf) improves LB score significantly. \n\nThe detailed are the following:\n\nFor the training set, the presence probability of label *L* is defined as *1 – epsilon* if the example has label L, otherwise *epsilon*. The label presence probabilities of test example are the weighted sum over the nearest K train examples.\n\n\nThe weights are determined based on the distances between test and train examples. As a distance metric, 1 – cosine similarity between train and test example is used. The output of the last convolution layer is used as embedding features.\n\nThe weights of training example j for an test example i are defined as:\n\n&gt; \\\\(\\pi_{i,j}=exp(-d * distance(i,j))/\\sum exp(-d * distance(i,j')))\\\\)\n\nThe probability of class *w* for test image *i* is:\n\n&gt; \\\\( p(y_{i,w}=1)=\\sum \\pi_{i,j} * p(y_{j,w})\\\\)\n\nThe parameter d and K has been chosen based on the validation score.\n\nWith this method, the LB score of single fold se-resnext101 with hflip TTA was 0.634.\nThe average of two predictions is LB 0.650.\n\n**Minor improvement**\n\nWe’ve found the images that have no close images have a low recall. So, we choose thresholds according to the distance from the nearest train image. If it has a higher distance, a lower threshold is assigned. Using this method, LB was improved +0.001~0.002.\n\nOur final submission is an ensemble of 4 folds se-resnext101 and 6 folds senet154. \n\nThanks.",
      "votes": null
    },
    {
      "id": "547002",
      "postDate": "06/07/2019 05:54:04",
      "content": "<p>Well done and congratulations!! It really didn't occur to me that metric learning would apply to this competition.</p>",
      "rawMarkdown": "Well done and congratulations!! It really didn't occur to me that metric learning would apply to this competition.",
      "votes": null
    },
    {
      "id": "547111",
      "postDate": "06/07/2019 09:44:34",
      "content": "<p>Congrats ! Thanks for sharing.</p>",
      "rawMarkdown": "Congrats ! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "547560",
      "postDate": "06/07/2019 21:34:19",
      "content": "<p>Congrats and thanks for sharing!! Really impressive prediction method!!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!! Really impressive prediction method!!",
      "votes": null
    },
    {
      "id": "548397",
      "postDate": "06/09/2019 10:33:00",
      "content": "<p>Congrats! Do you think additional FC layer was helpful? My guess would be, SENet already has enough layers to discriminate everything...</p>",
      "rawMarkdown": "Congrats! Do you think additional FC layer was helpful? My guess would be, SENet already has enough layers to discriminate everything...",
      "votes": null
    },
    {
      "id": "548978",
      "postDate": "06/10/2019 06:50:32",
      "content": "<p>Thanks for sharing .When use focal loss, performances of different thresholds is almost same, but when use BCEloss, there is a big gap between  performances of different thresholds. is that right? by the way, what is the best threshold of your model?</p>",
      "rawMarkdown": "Thanks for sharing .When use focal loss, performances of different thresholds is almost same, but when use BCEloss, there is a big gap between  performances of different thresholds. is that right? by the way, what is the best threshold of your model?",
      "votes": null
    },
    {
      "id": "549229",
      "postDate": "06/10/2019 12:43:48",
      "content": "<p>In my case, additional FC layer with dropout 0.5 was helpful. I replaced last linear with linear -&gt; batchnorm -&gt; relu -&gt; dropout -&gt; linear.</p>",
      "rawMarkdown": "In my case, additional FC layer with dropout 0.5 was helpful. I replaced last linear with linear -&gt; batchnorm -&gt; relu -&gt; dropout -&gt; linear.",
      "votes": null
    },
    {
      "id": "549238",
      "postDate": "06/10/2019 12:54:17",
      "content": "<p>In my case, the score was different depending on the threshold. \nThe best threshold is 0.32 for classification model, 0.15 for tag relevant prediction and 0.22 for the ensemble of them. </p>",
      "rawMarkdown": "In my case, the score was different depending on the threshold. \nThe best threshold is 0.32 for classification model, 0.15 for tag relevant prediction and 0.22 for the ensemble of them.",
      "votes": null
    },
    {
      "id": "550074",
      "postDate": "06/11/2019 09:16:19",
      "content": "<p>thanks a lot，i will try it again</p>",
      "rawMarkdown": "thanks a lot，i will try it again",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 547002,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "06/07/2019 05:54:04",
      "content": "<p>Well done and congratulations!! It really didn't occur to me that metric learning would apply to this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 547111,
      "author_name": "harshthaker",
      "author_url": "",
      "post_date": "06/07/2019 09:44:34",
      "content": "<p>Congrats ! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 547560,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "06/07/2019 21:34:19",
      "content": "<p>Congrats and thanks for sharing!! Really impressive prediction method!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 548397,
      "author_name": "artyomp",
      "author_url": "",
      "post_date": "06/09/2019 10:33:00",
      "content": "<p>Congrats! Do you think additional FC layer was helpful? My guess would be, SENet already has enough layers to discriminate everything...</p>",
      "votes": null,
      "replies": [
        {
          "id": 549229,
          "author_name": "pudae81",
          "author_url": "",
          "post_date": "06/10/2019 12:43:48",
          "content": "<p>In my case, additional FC layer with dropout 0.5 was helpful. I replaced last linear with linear -&gt; batchnorm -&gt; relu -&gt; dropout -&gt; linear.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 548978,
      "author_name": "valencebond",
      "author_url": "",
      "post_date": "06/10/2019 06:50:32",
      "content": "<p>Thanks for sharing .When use focal loss, performances of different thresholds is almost same, but when use BCEloss, there is a big gap between  performances of different thresholds. is that right? by the way, what is the best threshold of your model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 549238,
          "author_name": "pudae81",
          "author_url": "",
          "post_date": "06/10/2019 12:54:17",
          "content": "<p>In my case, the score was different depending on the threshold. \nThe best threshold is 0.32 for classification model, 0.15 for tag relevant prediction and 0.22 for the ensemble of them. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 550074,
          "author_name": "valencebond",
          "author_url": "",
          "post_date": "06/11/2019 09:16:19",
          "content": "<p>thanks a lot，i will try it again</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "546998": "Congrats to all the winners.\nThanks to Kaggle and hosting team for an interesting competition.\n\nHere is my solution summary.\n\n### Pure Classification Model\nBecause LB scores of our classification models are similar or worse with many other competitors, I briefly describe the configurations.\n\n-\t**Dataset**: 10 folds CV with Multilabel Iterative Stratification\n-\t**Models**: se-resnext101, senet154. After last convolution layer, one more fc layer is inserted.\n-\t**Augmentation**: random resized crop, horizontal flip.\n-\t**Loss**: class balanced focal loss. But it didn’t improve score compared to the pure focal loss.\n-\t**Optimization**: AdamW, \n   - learning rate 0.00025, weight decay 0.01 for se-resnext101.\n   - learning rate 0.0001, weight decay 0.05 for senet154\n-\t**Learning rate schedule**: CosineAnnealingLR. T_max 15, eta_min 0.00001\n-\tInput size 320x320. Batch size 24\n\nWith above configurations, the LB score of single fold se-resnext101 with hflip TTA was 0.618.\n\n\n### Tag Relevance Prediction\n\nWe’ve tried several methods, but all were failed. So we’ve decided to improve prediction methods. \nWe’ve found the method based on the [paper](http://class.inrialpes.fr/pub/guillaumin-iccv09b.pdf) improves LB score significantly. \n\nThe detailed are the following:\n\nFor the training set, the presence probability of label *L* is defined as *1 – epsilon* if the example has label L, otherwise *epsilon*. The label presence probabilities of test example are the weighted sum over the nearest K train examples.\n\n\nThe weights are determined based on the distances between test and train examples. As a distance metric, 1 – cosine similarity between train and test example is used. The output of the last convolution layer is used as embedding features.\n\nThe weights of training example j for an test example i are defined as:\n\n&gt; \\\\(\\pi_{i,j}=exp(-d * distance(i,j))/\\sum exp(-d * distance(i,j')))\\\\)\n\nThe probability of class *w* for test image *i* is:\n\n&gt; \\\\( p(y_{i,w}=1)=\\sum \\pi_{i,j} * p(y_{j,w})\\\\)\n\nThe parameter d and K has been chosen based on the validation score.\n\nWith this method, the LB score of single fold se-resnext101 with hflip TTA was 0.634.\nThe average of two predictions is LB 0.650.\n\n**Minor improvement**\n\nWe’ve found the images that have no close images have a low recall. So, we choose thresholds according to the distance from the nearest train image. If it has a higher distance, a lower threshold is assigned. Using this method, LB was improved +0.001~0.002.\n\nOur final submission is an ensemble of 4 folds se-resnext101 and 6 folds senet154. \n\nThanks.",
    "547002": "Well done and congratulations!! It really didn't occur to me that metric learning would apply to this competition.",
    "547111": "Congrats ! Thanks for sharing.",
    "547560": "Congrats and thanks for sharing!! Really impressive prediction method!!",
    "548397": "Congrats! Do you think additional FC layer was helpful? My guess would be, SENet already has enough layers to discriminate everything...",
    "548978": "Thanks for sharing .When use focal loss, performances of different thresholds is almost same, but when use BCEloss, there is a big gap between  performances of different thresholds. is that right? by the way, what is the best threshold of your model?",
    "549229": "In my case, additional FC layer with dropout 0.5 was helpful. I replaced last linear with linear -&gt; batchnorm -&gt; relu -&gt; dropout -&gt; linear.",
    "549238": "In my case, the score was different depending on the threshold. \nThe best threshold is 0.32 for classification model, 0.15 for tag relevant prediction and 0.22 for the ensemble of them.",
    "550074": "thanks a lot，i will try it again"
  },
  "source": "meta"
}