{
  "id": 94783,
  "title": "[Solution] 22 Place, LB 0.650",
  "url": "/competitions/imet-2019-fgvc6/writeups/ods-ai-yury-dzerin-solution-22-place-lb-0-650",
  "author_name": "",
  "post_date": "2019-06-11T08:09:02.660Z",
  "votes": 19,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Here are my findings and tricks, which helped me to get LB 0.650 by combining 3 models: se_resnext101_32x4d, pnasnet5large, senet154.</p>\n\n<p><strong>Validation</strong>: I used CV 5. Folds were made by the <a href=\"https://github.com/trent-b/iterative-stratification\">iterative stratification package</a>. The package is extremely useful for unbalanced dataset and multiclass classification.</p>\n\n<p><strong>Scheduler</strong>: is one of the most crucial part for the fast convergence. I used \nCosineAnnealingWarmRestarts increasing the frequency by factor of two after each restart (see the sketch). Using this strategy 15 Epochs(3 restarts) were enough to converge. And I was able to test different ideas much faster.\n<img src=\"https://i.ibb.co/ggyVkRJ/sgdr.jpg\" alt=\"Learning rate\"></p>\n\n<p><strong>Augmentation:</strong> everything is very standard: resize + random crop, small color jitter, hflip and small affine:\n<code>\n    transforms.Resize(img_size),\n    transforms.RandomCrop((img_size, img_size), padding=0, pad_if_needed=True),\n    transforms.ColorJitter(brightness=0.1,  contrast=0.1, saturation=0.1, hue=0.1),\n    transforms.RandomHorizontalFlip(p=0.5),\n    transforms.RandomAffine(degrees = 10, translate= (0.0, 0.2), scale=(0.8, 1.2), shear=10)\n</code>\n<strong>Mixup</strong>: is a trick from the paper: <a href=\"https://arxiv.org/abs/1812.01187\">Bag of Tricks for Image Classification with Convolutional Neural Networks</a>. The parameter of beta distribution was reduced by the same scheduler as the learning rate. It was reduced from 1 (very aggressive mixup) to 0 (no mixup at all). Mixup produced better validation results, but made the effect of TTA smaller. At the end it added around 0.002 per model. Adding this trick also made the convergence slower.</p>\n\n<p><strong>Dropout</strong>: The finding for the se_resnext101_32x4d: the standard dropout is 0.3, after increasing the dropout to 0.5 the CV and LB increased by ~0.005.</p>\n\n<p><strong>Loss</strong>: BCEWithLogitsLoss. The experiments with using Focal Loss or combining Focal Loss with BCE were not successful.</p>\n\n<p><strong>Batch Size</strong>: gradient accumulation didn’t help me much, so I just fit the batch size to occupy the whole GPU memory.</p>\n\n<p><strong>Threshold fitting</strong>: I guess it helped a lot to get nice results and made it easier to compare different experiments based on LB. \nFirst, I used validation to get the optimal threshold value, then I used this threshold to calculate average amount of labels per instance for the validation. It was equal to 5.2. During the inference I dynamically calculated the threshold such that I will approximately have 5.2 labels per image. Here is the code:</p>\n\n<p><code>\ndesired_mean = 5.2\nfor th in np.arange(1000)/10000. + 0.01:\n    pred = (np.array(res) &amp;gt; th).astype(np.float)\n    if np.abs(pred.sum()/len(pred) - desired_mean) &amp;lt; closest:\n        closest = np.abs(pred.sum()/len(pred) - desired_mean)\n        fix_th = th\n</code></p>\n\n<p>I also have another strategy, where I try to have a certain distribution of the predictions. It performed just a bit better (&lt;0.001) than the previous strategy, but it is quite hard to explain (I can share the code if someone will be interested).</p>\n\n<p><strong>TTA</strong>: All augmentations from training, parameters of RandomAffine and ColorJitter are smaller.</p>\n\n<p><strong>Results</strong>:</p>\n\n<p>| model | description | CV | LB |\n| --- | --- | --- | --- |\n| se_resnext101_32x4d | size: 300x300; + TTA4 | 0.607 | 0.642 |\n| pnasnet5large | size: 331x331; + mixup + TTA4 | 0.609 | 0.641 |\n| senet154 | size: 224x224; + mixup + TTA4 | 0.605 | 0.639 |</p>\n\n<p>Ensemble of this 3 Networks with TTA2: 0.650.</p>",
  "messages": [
    {
      "id": "546737",
      "postDate": "06/06/2019 21:13:03",
      "content": "<p>Here are my findings and tricks, which helped me to get LB 0.650 by combining 3 models: se_resnext101_32x4d, pnasnet5large, senet154.</p>\n\n<p><strong>Validation</strong>: I used CV 5. Folds were made by the <a href=\"https://github.com/trent-b/iterative-stratification\">iterative stratification package</a>. The package is extremely useful for unbalanced dataset and multiclass classification.</p>\n\n<p><strong>Scheduler</strong>: is one of the most crucial part for the fast convergence. I used \nCosineAnnealingWarmRestarts increasing the frequency by factor of two after each restart (see the sketch). Using this strategy 15 Epochs(3 restarts) were enough to converge. And I was able to test different ideas much faster.\n<img src=\"https://i.ibb.co/ggyVkRJ/sgdr.jpg\" alt=\"Learning rate\"></p>\n\n<p><strong>Augmentation:</strong> everything is very standard: resize + random crop, small color jitter, hflip and small affine:\n<code>\n    transforms.Resize(img_size),\n    transforms.RandomCrop((img_size, img_size), padding=0, pad_if_needed=True),\n    transforms.ColorJitter(brightness=0.1,  contrast=0.1, saturation=0.1, hue=0.1),\n    transforms.RandomHorizontalFlip(p=0.5),\n    transforms.RandomAffine(degrees = 10, translate= (0.0, 0.2), scale=(0.8, 1.2), shear=10)\n</code>\n<strong>Mixup</strong>: is a trick from the paper: <a href=\"https://arxiv.org/abs/1812.01187\">Bag of Tricks for Image Classification with Convolutional Neural Networks</a>. The parameter of beta distribution was reduced by the same scheduler as the learning rate. It was reduced from 1 (very aggressive mixup) to 0 (no mixup at all). Mixup produced better validation results, but made the effect of TTA smaller. At the end it added around 0.002 per model. Adding this trick also made the convergence slower.</p>\n\n<p><strong>Dropout</strong>: The finding for the se_resnext101_32x4d: the standard dropout is 0.3, after increasing the dropout to 0.5 the CV and LB increased by ~0.005.</p>\n\n<p><strong>Loss</strong>: BCEWithLogitsLoss. The experiments with using Focal Loss or combining Focal Loss with BCE were not successful.</p>\n\n<p><strong>Batch Size</strong>: gradient accumulation didn’t help me much, so I just fit the batch size to occupy the whole GPU memory.</p>\n\n<p><strong>Threshold fitting</strong>: I guess it helped a lot to get nice results and made it easier to compare different experiments based on LB. \nFirst, I used validation to get the optimal threshold value, then I used this threshold to calculate average amount of labels per instance for the validation. It was equal to 5.2. During the inference I dynamically calculated the threshold such that I will approximately have 5.2 labels per image. Here is the code:</p>\n\n<p><code>\ndesired_mean = 5.2\nfor th in np.arange(1000)/10000. + 0.01:\n    pred = (np.array(res) &amp;gt; th).astype(np.float)\n    if np.abs(pred.sum()/len(pred) - desired_mean) &amp;lt; closest:\n        closest = np.abs(pred.sum()/len(pred) - desired_mean)\n        fix_th = th\n</code></p>\n\n<p>I also have another strategy, where I try to have a certain distribution of the predictions. It performed just a bit better (&lt;0.001) than the previous strategy, but it is quite hard to explain (I can share the code if someone will be interested).</p>\n\n<p><strong>TTA</strong>: All augmentations from training, parameters of RandomAffine and ColorJitter are smaller.</p>\n\n<p><strong>Results</strong>:</p>\n\n<p>| model | description | CV | LB |\n| --- | --- | --- | --- |\n| se_resnext101_32x4d | size: 300x300; + TTA4 | 0.607 | 0.642 |\n| pnasnet5large | size: 331x331; + mixup + TTA4 | 0.609 | 0.641 |\n| senet154 | size: 224x224; + mixup + TTA4 | 0.605 | 0.639 |</p>\n\n<p>Ensemble of this 3 Networks with TTA2: 0.650.</p>",
      "rawMarkdown": "Here are my findings and tricks, which helped me to get LB 0.650 by combining 3 models: se\\_resnext101\\_32x4d, pnasnet5large, senet154.\n\n**Validation**: I used CV 5. Folds were made by the [iterative stratification package](https://github.com/trent-b/iterative-stratification). The package is extremely useful for unbalanced dataset and multiclass classification.\n\n**Scheduler**: is one of the most crucial part for the fast convergence. I used \nCosineAnnealingWarmRestarts increasing the frequency by factor of two after each restart (see the sketch). Using this strategy 15 Epochs(3 restarts) were enough to converge. And I was able to test different ideas much faster.\n![Learning rate](https://i.ibb.co/ggyVkRJ/sgdr.jpg)\n\n**Augmentation:** everything is very standard: resize + random crop, small color jitter, hflip and small affine:\n```\n    transforms.Resize(img_size),\n    transforms.RandomCrop((img_size, img_size), padding=0, pad_if_needed=True),\n    transforms.ColorJitter(brightness=0.1,  contrast=0.1, saturation=0.1, hue=0.1),\n    transforms.RandomHorizontalFlip(p=0.5),\n    transforms.RandomAffine(degrees = 10, translate= (0.0, 0.2), scale=(0.8, 1.2), shear=10)\n```\n**Mixup**: is a trick from the paper: [Bag of Tricks for Image Classification with Convolutional Neural Networks](https://arxiv.org/abs/1812.01187). The parameter of beta distribution was reduced by the same scheduler as the learning rate. It was reduced from 1 (very aggressive mixup) to 0 (no mixup at all). Mixup produced better validation results, but made the effect of TTA smaller. At the end it added around 0.002 per model. Adding this trick also made the convergence slower.\n\n**Dropout**: The finding for the se\\_resnext101\\_32x4d: the standard dropout is 0.3, after increasing the dropout to 0.5 the CV and LB increased by ~0.005.\n\n**Loss**: BCEWithLogitsLoss. The experiments with using Focal Loss or combining Focal Loss with BCE were not successful.\n\n**Batch Size**: gradient accumulation didn’t help me much, so I just fit the batch size to occupy the whole GPU memory.\n\n**Threshold fitting**: I guess it helped a lot to get nice results and made it easier to compare different experiments based on LB. \nFirst, I used validation to get the optimal threshold value, then I used this threshold to calculate average amount of labels per instance for the validation. It was equal to 5.2. During the inference I dynamically calculated the threshold such that I will approximately have 5.2 labels per image. Here is the code:\n\n```\ndesired_mean = 5.2\nfor th in np.arange(1000)/10000. + 0.01:\n    pred = (np.array(res) &gt; th).astype(np.float)\n    if np.abs(pred.sum()/len(pred) - desired_mean) &lt; closest:\n        closest = np.abs(pred.sum()/len(pred) - desired_mean)\n        fix_th = th\n```\n\n\nI also have another strategy, where I try to have a certain distribution of the predictions. It performed just a bit better (&lt;0.001) than the previous strategy, but it is quite hard to explain (I can share the code if someone will be interested).\n\n\n**TTA**: All augmentations from training, parameters of RandomAffine and ColorJitter are smaller.\n\n**Results**:\n\n| model | description | CV | LB |\n| --- | --- | --- | --- |\n| se\\_resnext101\\_32x4d | size: 300x300; + TTA4 | 0.607 | 0.642 |\n| pnasnet5large | size: 331x331; + mixup + TTA4 | 0.609 | 0.641 |\n| senet154 | size: 224x224; + mixup + TTA4 | 0.605 | 0.639 |\n\n\nEnsemble of this 3 Networks with TTA2: 0.650.",
      "votes": null
    },
    {
      "id": "547761",
      "postDate": "06/08/2019 08:21:35",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "547886",
      "postDate": "06/08/2019 13:23:24",
      "content": "<p>Nice strategy, thanks for sharing.</p>",
      "rawMarkdown": "Nice strategy, thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 547761,
      "author_name": "dhaqui",
      "author_url": "",
      "post_date": "06/08/2019 08:21:35",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 547886,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "06/08/2019 13:23:24",
      "content": "<p>Nice strategy, thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "546737": "Here are my findings and tricks, which helped me to get LB 0.650 by combining 3 models: se\\_resnext101\\_32x4d, pnasnet5large, senet154.\n\n**Validation**: I used CV 5. Folds were made by the [iterative stratification package](https://github.com/trent-b/iterative-stratification). The package is extremely useful for unbalanced dataset and multiclass classification.\n\n**Scheduler**: is one of the most crucial part for the fast convergence. I used \nCosineAnnealingWarmRestarts increasing the frequency by factor of two after each restart (see the sketch). Using this strategy 15 Epochs(3 restarts) were enough to converge. And I was able to test different ideas much faster.\n![Learning rate](https://i.ibb.co/ggyVkRJ/sgdr.jpg)\n\n**Augmentation:** everything is very standard: resize + random crop, small color jitter, hflip and small affine:\n```\n    transforms.Resize(img_size),\n    transforms.RandomCrop((img_size, img_size), padding=0, pad_if_needed=True),\n    transforms.ColorJitter(brightness=0.1,  contrast=0.1, saturation=0.1, hue=0.1),\n    transforms.RandomHorizontalFlip(p=0.5),\n    transforms.RandomAffine(degrees = 10, translate= (0.0, 0.2), scale=(0.8, 1.2), shear=10)\n```\n**Mixup**: is a trick from the paper: [Bag of Tricks for Image Classification with Convolutional Neural Networks](https://arxiv.org/abs/1812.01187). The parameter of beta distribution was reduced by the same scheduler as the learning rate. It was reduced from 1 (very aggressive mixup) to 0 (no mixup at all). Mixup produced better validation results, but made the effect of TTA smaller. At the end it added around 0.002 per model. Adding this trick also made the convergence slower.\n\n**Dropout**: The finding for the se\\_resnext101\\_32x4d: the standard dropout is 0.3, after increasing the dropout to 0.5 the CV and LB increased by ~0.005.\n\n**Loss**: BCEWithLogitsLoss. The experiments with using Focal Loss or combining Focal Loss with BCE were not successful.\n\n**Batch Size**: gradient accumulation didn’t help me much, so I just fit the batch size to occupy the whole GPU memory.\n\n**Threshold fitting**: I guess it helped a lot to get nice results and made it easier to compare different experiments based on LB. \nFirst, I used validation to get the optimal threshold value, then I used this threshold to calculate average amount of labels per instance for the validation. It was equal to 5.2. During the inference I dynamically calculated the threshold such that I will approximately have 5.2 labels per image. Here is the code:\n\n```\ndesired_mean = 5.2\nfor th in np.arange(1000)/10000. + 0.01:\n    pred = (np.array(res) &gt; th).astype(np.float)\n    if np.abs(pred.sum()/len(pred) - desired_mean) &lt; closest:\n        closest = np.abs(pred.sum()/len(pred) - desired_mean)\n        fix_th = th\n```\n\n\nI also have another strategy, where I try to have a certain distribution of the predictions. It performed just a bit better (&lt;0.001) than the previous strategy, but it is quite hard to explain (I can share the code if someone will be interested).\n\n\n**TTA**: All augmentations from training, parameters of RandomAffine and ColorJitter are smaller.\n\n**Results**:\n\n| model | description | CV | LB |\n| --- | --- | --- | --- |\n| se\\_resnext101\\_32x4d | size: 300x300; + TTA4 | 0.607 | 0.642 |\n| pnasnet5large | size: 331x331; + mixup + TTA4 | 0.609 | 0.641 |\n| senet154 | size: 224x224; + mixup + TTA4 | 0.605 | 0.639 |\n\n\nEnsemble of this 3 Networks with TTA2: 0.650.",
    "547761": "Thanks for sharing!",
    "547886": "Nice strategy, thanks for sharing."
  },
  "source": "meta"
}