{
  "id": 95289,
  "title": "OSError on stage2, code for ~20th place shared",
  "url": "/competitions/imet-2019-fgvc6/discussion/95289",
  "author_name": "Artyom Palvelev",
  "post_date": "2019-06-11T09:02:39.677000",
  "votes": 21,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Ok, I wasn't expecting this kind of final, but it turned out that writing all predicts to the disk is not ok. Kaggle kernels don't have enough disk space for 35 models. Here is my solution anyway. I'm pretty confident it would get 0.647 on Private LB. The code is here: <a href=\"https://github.com/artyompal/imet\">https://github.com/artyompal/imet</a></p>\n\n<p>I decided that a big number of TTAs means essentially a blend of the same model, so I chose to use many different models. I even ran out of 20 Gb space of private datasets, so I added encrypted models in public datasets (they are not being used in the best ensemble, though).</p>\n\n<p>My best ensemble is 7 models with TTA x2. It was supposed to finish Stage 2 prediction in 8 hrs 40 mins:\n* SE-ResNext50 at 288x288;\n* SE-ResNext101 at 288x288;\n* CBAM-ResNet50 at 288x288;\n* PNASNet5 Large at 288x288;\n* another SE-ResNext101 at 288x288 with fewer augmentations and dropout 0.3;\n* two more SE-ResNext101 at 352x352 with different augmentations.</p>\n\n<p>The batch size was 16. This worked better than 32 or 64, and batch accumulation 4x or 8x didn't improve the score.</p>\n\n<p>I started with simple LR scheduler ReduceLrOnPlateau but then moved to Cosine LR.</p>\n\n<p>I used cross-entropy loss. Focal loss and F2 loss weren't better. After I realized there's a lot of missing labels, I came up with this loss:\n<code>\nclass ForgivingLoss(nn.Module):\n    def __init__(self, weight: float) -&amp;gt; None:\n        super().__init__()\n        self.bce_loss = binary_cross_entropy()\n        self.weight = weight\n    def forward(self, logits: torch.Tensor, labels: torch.Tensor) -&amp;gt; torch.Tensor:\n        return self.bce_loss(logits, labels) + self.bce_loss(logits * labels, labels) * self.weight\n</code></p>\n\n<p>Interestingly, it worked a little better with Inception-like models, namely Xception and InceptionResNetV2, but not with ResNext-like models.</p>\n\n<p>I trained 5 folds of each model. Then I used <code>scipy.optimize.minimize</code> to find the best blend coefficients.</p>\n\n<p>I used Kostia's method of deployment, which packs everything into a single py-file. Also, I wrote a code which automatically searches for all available models in <code>../input/</code> and decrypts them.</p>\n\n<p>What didn't work: I tried pseudo-labeling, it improved a single fold score from 611 to 622, but worked worse on LB. Probably, I generated too many labels. I would be possible to achieve gold otherwise.</p>\n\n<p>\"If we compete only for the win, we may lose, but if we compete for learning and providing a useful solution to the host, there's nothing to lose\" (c) bestfitting</p>",
  "messages": [
    {
      "id": 550055,
      "postDate": "2019-06-11T09:02:39.677Z",
      "content": "<p>Ok, I wasn't expecting this kind of final, but it turned out that writing all predicts to the disk is not ok. Kaggle kernels don't have enough disk space for 35 models. Here is my solution anyway. I'm pretty confident it would get 0.647 on Private LB. The code is here: <a href=\"https://github.com/artyompal/imet\">https://github.com/artyompal/imet</a></p>\n\n<p>I decided that a big number of TTAs means essentially a blend of the same model, so I chose to use many different models. I even ran out of 20 Gb space of private datasets, so I added encrypted models in public datasets (they are not being used in the best ensemble, though).</p>\n\n<p>My best ensemble is 7 models with TTA x2. It was supposed to finish Stage 2 prediction in 8 hrs 40 mins:\n* SE-ResNext50 at 288x288;\n* SE-ResNext101 at 288x288;\n* CBAM-ResNet50 at 288x288;\n* PNASNet5 Large at 288x288;\n* another SE-ResNext101 at 288x288 with fewer augmentations and dropout 0.3;\n* two more SE-ResNext101 at 352x352 with different augmentations.</p>\n\n<p>The batch size was 16. This worked better than 32 or 64, and batch accumulation 4x or 8x didn't improve the score.</p>\n\n<p>I started with simple LR scheduler ReduceLrOnPlateau but then moved to Cosine LR.</p>\n\n<p>I used cross-entropy loss. Focal loss and F2 loss weren't better. After I realized there's a lot of missing labels, I came up with this loss:\n<code>\nclass ForgivingLoss(nn.Module):\n    def __init__(self, weight: float) -&amp;gt; None:\n        super().__init__()\n        self.bce_loss = binary_cross_entropy()\n        self.weight = weight\n    def forward(self, logits: torch.Tensor, labels: torch.Tensor) -&amp;gt; torch.Tensor:\n        return self.bce_loss(logits, labels) + self.bce_loss(logits * labels, labels) * self.weight\n</code></p>\n\n<p>Interestingly, it worked a little better with Inception-like models, namely Xception and InceptionResNetV2, but not with ResNext-like models.</p>\n\n<p>I trained 5 folds of each model. Then I used <code>scipy.optimize.minimize</code> to find the best blend coefficients.</p>\n\n<p>I used Kostia's method of deployment, which packs everything into a single py-file. Also, I wrote a code which automatically searches for all available models in <code>../input/</code> and decrypts them.</p>\n\n<p>What didn't work: I tried pseudo-labeling, it improved a single fold score from 611 to 622, but worked worse on LB. Probably, I generated too many labels. I would be possible to achieve gold otherwise.</p>\n\n<p>\"If we compete only for the win, we may lose, but if we compete for learning and providing a useful solution to the host, there's nothing to lose\" (c) bestfitting</p>",
      "rawMarkdown": "Ok, I wasn't expecting this kind of final, but it turned out that writing all predicts to the disk is not ok. Kaggle kernels don't have enough disk space for 35 models. Here is my solution anyway. I'm pretty confident it would get 0.647 on Private LB. The code is here: https://github.com/artyompal/imet\n\nI decided that a big number of TTAs means essentially a blend of the same model, so I chose to use many different models. I even ran out of 20 Gb space of private datasets, so I added encrypted models in public datasets (they are not being used in the best ensemble, though).\n\nMy best ensemble is 7 models with TTA x2. It was supposed to finish Stage 2 prediction in 8 hrs 40 mins:\n* SE-ResNext50 at 288x288;\n* SE-ResNext101 at 288x288;\n* CBAM-ResNet50 at 288x288;\n* PNASNet5 Large at 288x288;\n* another SE-ResNext101 at 288x288 with fewer augmentations and dropout 0.3;\n* two more SE-ResNext101 at 352x352 with different augmentations.\n\nThe batch size was 16. This worked better than 32 or 64, and batch accumulation 4x or 8x didn't improve the score.\n\nI started with simple LR scheduler ReduceLrOnPlateau but then moved to Cosine LR.\n\nI used cross-entropy loss. Focal loss and F2 loss weren't better. After I realized there's a lot of missing labels, I came up with this loss:\n```\nclass ForgivingLoss(nn.Module):\n    def __init__(self, weight: float) -&gt; None:\n        super().__init__()\n        self.bce_loss = binary_cross_entropy()\n        self.weight = weight\n    def forward(self, logits: torch.Tensor, labels: torch.Tensor) -&gt; torch.Tensor:\n        return self.bce_loss(logits, labels) + self.bce_loss(logits * labels, labels) * self.weight\n```\n\nInterestingly, it worked a little better with Inception-like models, namely Xception and InceptionResNetV2, but not with ResNext-like models.\n\nI trained 5 folds of each model. Then I used `scipy.optimize.minimize` to find the best blend coefficients.\n\nI used Kostia's method of deployment, which packs everything into a single py-file. Also, I wrote a code which automatically searches for all available models in `../input/` and decrypts them.\n\nWhat didn't work: I tried pseudo-labeling, it improved a single fold score from 611 to 622, but worked worse on LB. Probably, I generated too many labels. I would be possible to achieve gold otherwise.\n\n\"If we compete only for the win, we may lose, but if we compete for learning and providing a useful solution to the host, there's nothing to lose\" (c) bestfitting",
      "votes": 21
    },
    {
      "id": 550298,
      "postDate": "2019-06-11T13:24:28.923Z",
      "content": "<p>You did an amazing work, thanks for sharing your solution 👍 \n<br>\n<img src=\"https://i.kym-cdn.com/photos/images/masonry/000/248/081/7d1.jpg\" alt=\"\"></p>",
      "rawMarkdown": "You did an amazing work, thanks for sharing your solution 👍 \n<br>\n![](https://i.kym-cdn.com/photos/images/masonry/000/248/081/7d1.jpg)",
      "votes": 1,
      "replies": [
        {
          "id": 550302,
          "postDate": "2019-06-11T13:29:49.483Z",
          "content": "<p>Bad things do happen sometimes :)</p>",
          "rawMarkdown": "Bad things do happen sometimes :)"
        }
      ]
    },
    {
      "id": 550204,
      "postDate": "2019-06-11T12:01:20.913Z",
      "content": "<p>Sorry to hear that. Good luck with your next competition! </p>",
      "rawMarkdown": "Sorry to hear that. Good luck with your next competition! ",
      "votes": 1,
      "replies": [
        {
          "id": 550257,
          "postDate": "2019-06-11T12:50:12.517Z",
          "content": "<p>Thanks, and congrats with your medal!</p>",
          "rawMarkdown": "Thanks, and congrats with your medal!",
          "votes": 1
        }
      ]
    },
    {
      "id": 550192,
      "postDate": "2019-06-11T11:50:32.313Z",
      "content": "<p>Very sad. I remember your progress on Leaderboard during competition.\nMaybe this is not the key for your problem, but when I met with such situation, I uploaded new models with the same names (-&gt; rewriting). This step prevented me from disk space error.</p>",
      "rawMarkdown": "Very sad. I remember your progress on Leaderboard during competition.\nMaybe this is not the key for your problem, but when I met with such situation, I uploaded new models with the same names (-&gt; rewriting). This step prevented me from disk space error.",
      "votes": 1
    },
    {
      "id": 550161,
      "postDate": "2019-06-11T11:15:24.923Z",
      "content": "<p>hi, how do you know that it is due to OSError in stage2? For me, the submission page look like this, it seems that my kernel has not been rerun?\n<img src=\"http://chuantu.xyz/t6/702/1560253170x1954578459.png\"></p>",
      "rawMarkdown": "hi, how do you know that it is due to OSError in stage2? For me, the submission page look like this, it seems that my kernel has not been rerun?\n<img src=\"http://chuantu.xyz/t6/702/1560253170x1954578459.png\">",
      "replies": [
        {
          "id": 550189,
          "postDate": "2019-06-11T11:46:25.330Z",
          "content": "<p>Try and run your selected kernel. When it finishes, download the log and look at the end.</p>",
          "rawMarkdown": "Try and run your selected kernel. When it finishes, download the log and look at the end.",
          "votes": 1
        },
        {
          "id": 550259,
          "postDate": "2019-06-11T12:51:09.610Z",
          "content": "<p>thanks, I will try.</p>",
          "rawMarkdown": "thanks, I will try."
        },
        {
          "id": 550831,
          "postDate": "2019-06-12T04:30:15.643Z",
          "content": "<p>ouch, my kernel runs into memory problem, sad ...</p>",
          "rawMarkdown": "ouch, my kernel runs into memory problem, sad ..."
        }
      ]
    },
    {
      "id": 550070,
      "postDate": "2019-06-11T09:13:00.780Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 550298,
      "author_name": "Nanashi",
      "author_url": "",
      "post_date": "2019-06-11T13:24:28.923000",
      "content": "<p>You did an amazing work, thanks for sharing your solution 👍 \n<br>\n<img src=\"https://i.kym-cdn.com/photos/images/masonry/000/248/081/7d1.jpg\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 550302,
          "author_name": "Artyom Palvelev",
          "author_url": "",
          "post_date": "2019-06-11T13:29:49.483000",
          "content": "<p>Bad things do happen sometimes :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 550204,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2019-06-11T12:01:20.913000",
      "content": "<p>Sorry to hear that. Good luck with your next competition! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 550257,
          "author_name": "Artyom Palvelev",
          "author_url": "",
          "post_date": "2019-06-11T12:50:12.517000",
          "content": "<p>Thanks, and congrats with your medal!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 550192,
      "author_name": "Alexander Kireev",
      "author_url": "",
      "post_date": "2019-06-11T11:50:32.313000",
      "content": "<p>Very sad. I remember your progress on Leaderboard during competition.\nMaybe this is not the key for your problem, but when I met with such situation, I uploaded new models with the same names (-&gt; rewriting). This step prevented me from disk space error.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 550161,
      "author_name": "good good study",
      "author_url": "",
      "post_date": "2019-06-11T11:15:24.923000",
      "content": "<p>hi, how do you know that it is due to OSError in stage2? For me, the submission page look like this, it seems that my kernel has not been rerun?\n<img src=\"http://chuantu.xyz/t6/702/1560253170x1954578459.png\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 550189,
          "author_name": "Artyom Palvelev",
          "author_url": "",
          "post_date": "2019-06-11T11:46:25.330000",
          "content": "<p>Try and run your selected kernel. When it finishes, download the log and look at the end.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 550259,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2019-06-11T12:51:09.610000",
          "content": "<p>thanks, I will try.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 550831,
          "author_name": "good good study",
          "author_url": "",
          "post_date": "2019-06-12T04:30:15.643000",
          "content": "<p>ouch, my kernel runs into memory problem, sad ...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 550070,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-11T09:13:00.780000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "550055": "Ok, I wasn't expecting this kind of final, but it turned out that writing all predicts to the disk is not ok. Kaggle kernels don't have enough disk space for 35 models. Here is my solution anyway. I'm pretty confident it would get 0.647 on Private LB. The code is here: https://github.com/artyompal/imet\n\nI decided that a big number of TTAs means essentially a blend of the same model, so I chose to use many different models. I even ran out of 20 Gb space of private datasets, so I added encrypted models in public datasets (they are not being used in the best ensemble, though).\n\nMy best ensemble is 7 models with TTA x2. It was supposed to finish Stage 2 prediction in 8 hrs 40 mins:\n* SE-ResNext50 at 288x288;\n* SE-ResNext101 at 288x288;\n* CBAM-ResNet50 at 288x288;\n* PNASNet5 Large at 288x288;\n* another SE-ResNext101 at 288x288 with fewer augmentations and dropout 0.3;\n* two more SE-ResNext101 at 352x352 with different augmentations.\n\nThe batch size was 16. This worked better than 32 or 64, and batch accumulation 4x or 8x didn't improve the score.\n\nI started with simple LR scheduler ReduceLrOnPlateau but then moved to Cosine LR.\n\nI used cross-entropy loss. Focal loss and F2 loss weren't better. After I realized there's a lot of missing labels, I came up with this loss:\n```\nclass ForgivingLoss(nn.Module):\n    def __init__(self, weight: float) -&gt; None:\n        super().__init__()\n        self.bce_loss = binary_cross_entropy()\n        self.weight = weight\n    def forward(self, logits: torch.Tensor, labels: torch.Tensor) -&gt; torch.Tensor:\n        return self.bce_loss(logits, labels) + self.bce_loss(logits * labels, labels) * self.weight\n```\n\nInterestingly, it worked a little better with Inception-like models, namely Xception and InceptionResNetV2, but not with ResNext-like models.\n\nI trained 5 folds of each model. Then I used `scipy.optimize.minimize` to find the best blend coefficients.\n\nI used Kostia's method of deployment, which packs everything into a single py-file. Also, I wrote a code which automatically searches for all available models in `../input/` and decrypts them.\n\nWhat didn't work: I tried pseudo-labeling, it improved a single fold score from 611 to 622, but worked worse on LB. Probably, I generated too many labels. I would be possible to achieve gold otherwise.\n\n\"If we compete only for the win, we may lose, but if we compete for learning and providing a useful solution to the host, there's nothing to lose\" (c) bestfitting",
    "550298": "You did an amazing work, thanks for sharing your solution 👍 \n<br>\n![](https://i.kym-cdn.com/photos/images/masonry/000/248/081/7d1.jpg)",
    "550204": "Sorry to hear that. Good luck with your next competition! ",
    "550192": "Very sad. I remember your progress on Leaderboard during competition.\nMaybe this is not the key for your problem, but when I met with such situation, I uploaded new models with the same names (-&gt; rewriting). This step prevented me from disk space error.",
    "550161": "hi, how do you know that it is due to OSError in stage2? For me, the submission page look like this, it seems that my kernel has not been rerun?\n<img src=\"http://chuantu.xyz/t6/702/1560253170x1954578459.png\">",
    "550070": ""
  }
}