{
  "id": 110390,
  "title": "14th place solution, [0.989] Colab and Kaggle kernels",
  "url": "/competitions/recursion-cellular-image-classification/discussion/110390",
  "author_name": "Maksym Pyrozhok",
  "post_date": "2019-09-27T09:24:14.538000",
  "votes": 17,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I want to thank Recursion Pharmaceuticals and Kaggle for organizing this competition. I also want to thank everyone who participated and especially <a href=\"https://www.kaggle.com/zaharch\">nosound</a> for explaining and providing sample code for leak exploitation.</p>\n\n<p>My solution was developed using only freely available computational resources. I used Colab at the start of the competition, then briefly switched to Kaggle kernels up to the point when Kaggle enforced GPU time limitations and then Colab again. My final submission was a single densenet201-based model.</p>\n\n<p><strong>Tools</strong>\n- PyTorch with Apex and a bit of Ignite.</p>\n\n<p><strong>Data</strong>\n- 6x512x512. First I've tried to use smaller resolutions, but understood fairly quickly that full-size images are necessary for maximum performance.\n- For augmentation I've used rotations, horizontal flip, cutout and random brightness/contrast. I've also used 4 images rotated by 90 degrees for TTA.</p>\n\n<p><strong>Model</strong>\n- densenet201 with ArcFace loss.\n- separate 1108 classes for every cell type.</p>\n\n<p><strong>Training process</strong>\n- AdamW and cyclical learning rate with restarts, triangular schedule.\n- validating on 10% of train dataset to find a point where model begins to overfit. Training on the whole dataset up to the cycle that showed improvement on validation.\n- using the last checkpoint in a cycle for prediction.\n- averaging distances for two sites.</p>\n\n<p><strong>Post-processing</strong>\n- using leak. Getting maximum prediction combination using lapjv for every plate.</p>\n\n<p>Pseudo-labeling and blending different models' predictions would have likely boosted my score, but unfortunately there was not enough time and computational resources.</p>\n\n<p><strong>Congratulations to the winners!</strong></p>",
  "messages": [
    {
      "id": 635229,
      "postDate": "2019-09-27T09:24:14.537Z",
      "content": "<p>I want to thank Recursion Pharmaceuticals and Kaggle for organizing this competition. I also want to thank everyone who participated and especially <a href=\"https://www.kaggle.com/zaharch\">nosound</a> for explaining and providing sample code for leak exploitation.</p>\n\n<p>My solution was developed using only freely available computational resources. I used Colab at the start of the competition, then briefly switched to Kaggle kernels up to the point when Kaggle enforced GPU time limitations and then Colab again. My final submission was a single densenet201-based model.</p>\n\n<p><strong>Tools</strong>\n- PyTorch with Apex and a bit of Ignite.</p>\n\n<p><strong>Data</strong>\n- 6x512x512. First I've tried to use smaller resolutions, but understood fairly quickly that full-size images are necessary for maximum performance.\n- For augmentation I've used rotations, horizontal flip, cutout and random brightness/contrast. I've also used 4 images rotated by 90 degrees for TTA.</p>\n\n<p><strong>Model</strong>\n- densenet201 with ArcFace loss.\n- separate 1108 classes for every cell type.</p>\n\n<p><strong>Training process</strong>\n- AdamW and cyclical learning rate with restarts, triangular schedule.\n- validating on 10% of train dataset to find a point where model begins to overfit. Training on the whole dataset up to the cycle that showed improvement on validation.\n- using the last checkpoint in a cycle for prediction.\n- averaging distances for two sites.</p>\n\n<p><strong>Post-processing</strong>\n- using leak. Getting maximum prediction combination using lapjv for every plate.</p>\n\n<p>Pseudo-labeling and blending different models' predictions would have likely boosted my score, but unfortunately there was not enough time and computational resources.</p>\n\n<p><strong>Congratulations to the winners!</strong></p>",
      "rawMarkdown": "I want to thank Recursion Pharmaceuticals and Kaggle for organizing this competition. I also want to thank everyone who participated and especially [nosound](https://www.kaggle.com/zaharch) for explaining and providing sample code for leak exploitation.\n\nMy solution was developed using only freely available computational resources. I used Colab at the start of the competition, then briefly switched to Kaggle kernels up to the point when Kaggle enforced GPU time limitations and then Colab again. My final submission was a single densenet201-based model.\n\n**Tools**\n- PyTorch with Apex and a bit of Ignite.\n\n**Data**\n- 6x512x512. First I've tried to use smaller resolutions, but understood fairly quickly that full-size images are necessary for maximum performance.\n- For augmentation I've used rotations, horizontal flip, cutout and random brightness/contrast. I've also used 4 images rotated by 90 degrees for TTA.\n\n**Model**\n- densenet201 with ArcFace loss.\n- separate 1108 classes for every cell type.\n\n**Training process**\n- AdamW and cyclical learning rate with restarts, triangular schedule.\n- validating on 10% of train dataset to find a point where model begins to overfit. Training on the whole dataset up to the cycle that showed improvement on validation.\n- using the last checkpoint in a cycle for prediction.\n- averaging distances for two sites.\n\n**Post-processing**\n- using leak. Getting maximum prediction combination using lapjv for every plate.\n\nPseudo-labeling and blending different models' predictions would have likely boosted my score, but unfortunately there was not enough time and computational resources.\n\n**Congratulations to the winners!**",
      "votes": 17
    },
    {
      "id": 635239,
      "postDate": "2019-09-27T09:34:54.473Z",
      "content": "<p>Yeah, agree, I would say in this competitions one could get a good score even with limited resources. I spend lots of time doing some fancy stuff, but my best scores were from DenseNets with plain cross-entropy, a bit of data augmentation and well-known methods like 1-cycle, cosine annealing, AdamW, Lookahead, etc. So if I would train a bigger model and for a longer period of time, I believe the score would be much higher :)</p>\n\n<p>So a good lesson to learn is to try bigger networks (the competitions become more and more involved), and use long training.</p>",
      "rawMarkdown": "Yeah, agree, I would say in this competitions one could get a good score even with limited resources. I spend lots of time doing some fancy stuff, but my best scores were from DenseNets with plain cross-entropy, a bit of data augmentation and well-known methods like 1-cycle, cosine annealing, AdamW, Lookahead, etc. So if I would train a bigger model and for a longer period of time, I believe the score would be much higher :)\n\nSo a good lesson to learn is to try bigger networks (the competitions become more and more involved), and use long training.",
      "votes": 1,
      "replies": [
        {
          "id": 635439,
          "postDate": "2019-09-27T14:26:13.633Z",
          "content": "<p>The training itself did not take prohibitively long time actually. About 30 to 40 epochs were enough and it took about 3 days on T4. But actually making decisions about pipeline/model/loss was not a straightforward process at all. I've experimented with less and more heavy architectures, different normalization schemes, different resolutions, losses, validation schemes etc. The bottleneck was colab instability, its overall performance and that it does not allow to train for an acceptable amount of time. This hinders iterating over ideas significantly and is actually a painful process. Unfortunately my kaggle account was locked during competition for almost two months and I was not able to use Kaggle kernels at all. After it was unlocked there was not that much time left and then GPU limits came. </p>",
          "rawMarkdown": "The training itself did not take prohibitively long time actually. About 30 to 40 epochs were enough and it took about 3 days on T4. But actually making decisions about pipeline/model/loss was not a straightforward process at all. I've experimented with less and more heavy architectures, different normalization schemes, different resolutions, losses, validation schemes etc. The bottleneck was colab instability, its overall performance and that it does not allow to train for an acceptable amount of time. This hinders iterating over ideas significantly and is actually a painful process. Unfortunately my kaggle account was locked during competition for almost two months and I was not able to use Kaggle kernels at all. After it was unlocked there was not that much time left and then GPU limits came. ",
          "votes": -1
        }
      ]
    },
    {
      "id": 637887,
      "postDate": "2019-10-01T10:44:32.790Z",
      "content": "<p>Hi <a href=\"/mpyrozhok\">@mpyrozhok</a>,\nDo you mind sharing how to make kaggle submission from google colab? I have been struggling over it for the past few days. </p>\n\n<p>Best 🙂</p>",
      "rawMarkdown": "Hi @mpyrozhok,\nDo you mind sharing how to make kaggle submission from google colab? I have been struggling over it for the past few days. \n\nBest 🙂",
      "replies": [
        {
          "id": 638102,
          "postDate": "2019-10-01T14:34:33.513Z",
          "content": "<p>You can refer to this <a href=\"https://colab.research.google.com/drive/17MsV4f8Bap8y2MU8BcKBNXRdTNE2ftP3\">notebook</a>.</p>",
          "rawMarkdown": "You can refer to this [notebook](https://colab.research.google.com/drive/17MsV4f8Bap8y2MU8BcKBNXRdTNE2ftP3)."
        }
      ]
    },
    {
      "id": 635330,
      "postDate": "2019-09-27T11:40:18.243Z",
      "content": "<p>Will it be possible for you to share your code.I would really like to play around with arcface and learn the details(couldn't make it work).\nThanks</p>",
      "rawMarkdown": "Will it be possible for you to share your code.I would really like to play around with arcface and learn the details(couldn't make it work).\nThanks\n",
      "replies": [
        {
          "id": 635413,
          "postDate": "2019-09-27T13:48:13.423Z",
          "content": "<p>Unfortunately I currently don't have time to bring my codebase to a level acceptable for public release. My implementation of ArcFace does not deviate from its description in the paper. You can refer to <a href=\"https://github.com/ronghuaiyang/arcface-pytorch/blob/master/models/metrics.py\">this</a> for details. My implementation is even less numerically stable.</p>",
          "rawMarkdown": "Unfortunately I currently don't have time to bring my codebase to a level acceptable for public release. My implementation of ArcFace does not deviate from its description in the paper. You can refer to [this](https://github.com/ronghuaiyang/arcface-pytorch/blob/master/models/metrics.py) for details. My implementation is even less numerically stable.",
          "votes": 2
        }
      ]
    },
    {
      "id": 635236,
      "postDate": "2019-09-27T09:32:49.347Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 635239,
      "author_name": "Ilia Zaitsev",
      "author_url": "",
      "post_date": "2019-09-27T09:34:54.473000",
      "content": "<p>Yeah, agree, I would say in this competitions one could get a good score even with limited resources. I spend lots of time doing some fancy stuff, but my best scores were from DenseNets with plain cross-entropy, a bit of data augmentation and well-known methods like 1-cycle, cosine annealing, AdamW, Lookahead, etc. So if I would train a bigger model and for a longer period of time, I believe the score would be much higher :)</p>\n\n<p>So a good lesson to learn is to try bigger networks (the competitions become more and more involved), and use long training.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 635439,
          "author_name": "Maksym Pyrozhok",
          "author_url": "",
          "post_date": "2019-09-27T14:26:13.633000",
          "content": "<p>The training itself did not take prohibitively long time actually. About 30 to 40 epochs were enough and it took about 3 days on T4. But actually making decisions about pipeline/model/loss was not a straightforward process at all. I've experimented with less and more heavy architectures, different normalization schemes, different resolutions, losses, validation schemes etc. The bottleneck was colab instability, its overall performance and that it does not allow to train for an acceptable amount of time. This hinders iterating over ideas significantly and is actually a painful process. Unfortunately my kaggle account was locked during competition for almost two months and I was not able to use Kaggle kernels at all. After it was unlocked there was not that much time left and then GPU limits came. </p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 637887,
      "author_name": "Callistus Ezenwaka",
      "author_url": "",
      "post_date": "2019-10-01T10:44:32.790000",
      "content": "<p>Hi <a href=\"/mpyrozhok\">@mpyrozhok</a>,\nDo you mind sharing how to make kaggle submission from google colab? I have been struggling over it for the past few days. </p>\n\n<p>Best 🙂</p>",
      "votes": 0,
      "replies": [
        {
          "id": 638102,
          "author_name": "Maksym Pyrozhok",
          "author_url": "",
          "post_date": "2019-10-01T14:34:33.513000",
          "content": "<p>You can refer to this <a href=\"https://colab.research.google.com/drive/17MsV4f8Bap8y2MU8BcKBNXRdTNE2ftP3\">notebook</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 635330,
      "author_name": "Deepshad",
      "author_url": "",
      "post_date": "2019-09-27T11:40:18.243000",
      "content": "<p>Will it be possible for you to share your code.I would really like to play around with arcface and learn the details(couldn't make it work).\nThanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 635413,
          "author_name": "Maksym Pyrozhok",
          "author_url": "",
          "post_date": "2019-09-27T13:48:13.423000",
          "content": "<p>Unfortunately I currently don't have time to bring my codebase to a level acceptable for public release. My implementation of ArcFace does not deviate from its description in the paper. You can refer to <a href=\"https://github.com/ronghuaiyang/arcface-pytorch/blob/master/models/metrics.py\">this</a> for details. My implementation is even less numerically stable.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 635236,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-27T09:32:49.347000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "635229": "I want to thank Recursion Pharmaceuticals and Kaggle for organizing this competition. I also want to thank everyone who participated and especially [nosound](https://www.kaggle.com/zaharch) for explaining and providing sample code for leak exploitation.\n\nMy solution was developed using only freely available computational resources. I used Colab at the start of the competition, then briefly switched to Kaggle kernels up to the point when Kaggle enforced GPU time limitations and then Colab again. My final submission was a single densenet201-based model.\n\n**Tools**\n- PyTorch with Apex and a bit of Ignite.\n\n**Data**\n- 6x512x512. First I've tried to use smaller resolutions, but understood fairly quickly that full-size images are necessary for maximum performance.\n- For augmentation I've used rotations, horizontal flip, cutout and random brightness/contrast. I've also used 4 images rotated by 90 degrees for TTA.\n\n**Model**\n- densenet201 with ArcFace loss.\n- separate 1108 classes for every cell type.\n\n**Training process**\n- AdamW and cyclical learning rate with restarts, triangular schedule.\n- validating on 10% of train dataset to find a point where model begins to overfit. Training on the whole dataset up to the cycle that showed improvement on validation.\n- using the last checkpoint in a cycle for prediction.\n- averaging distances for two sites.\n\n**Post-processing**\n- using leak. Getting maximum prediction combination using lapjv for every plate.\n\nPseudo-labeling and blending different models' predictions would have likely boosted my score, but unfortunately there was not enough time and computational resources.\n\n**Congratulations to the winners!**",
    "635239": "Yeah, agree, I would say in this competitions one could get a good score even with limited resources. I spend lots of time doing some fancy stuff, but my best scores were from DenseNets with plain cross-entropy, a bit of data augmentation and well-known methods like 1-cycle, cosine annealing, AdamW, Lookahead, etc. So if I would train a bigger model and for a longer period of time, I believe the score would be much higher :)\n\nSo a good lesson to learn is to try bigger networks (the competitions become more and more involved), and use long training.",
    "637887": "Hi @mpyrozhok,\nDo you mind sharing how to make kaggle submission from google colab? I have been struggling over it for the past few days. \n\nBest 🙂",
    "635330": "Will it be possible for you to share your code.I would really like to play around with arcface and learn the details(couldn't make it work).\nThanks\n",
    "635236": ""
  }
}