{
  "id": 110366,
  "title": "Our solution: 9th place",
  "url": "/competitions/recursion-cellular-image-classification/discussion/110366",
  "author_name": "Ivan Sosin",
  "post_date": "2019-09-27T06:34:51.839000",
  "votes": 39,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Dear all, thank you for excellent spirit of competition and astonishing results! Without such favorable pressure we wouldn’t be able to achieve such scores.</p>\n\n<p>I will split description into two main parts: pipeline and prediction with post-processing.</p>\n\n<p><strong>Pipeline:</strong>\n- <strong>6-channel images.</strong>\n- <strong>Super Resolution.</strong> During exploratory data analysis we suddenly realize that the size of individual cells is rather small and convolutional layer might have a hard time learning all patterns. So we use simple bicubic superresolution technique. That really boost our scores. We checked up to 960 * 960. We have tried to use GAN’s (ESRGAN in particular) for this task but its performance was worse compared to vanilla bicubic interpolation.\n- <strong>3Fold by experiments.</strong>\n- <strong>Add all controls to train.</strong> We tried to add controls only from train experiments, but it didn’t work.\n- <strong>Mixed Precision:</strong> Due to the large size of images we decided to use only Mixed Precision learning because of inevitable graphics card’s memory shortage.\n- <strong>Gradual layer unfreezing:</strong> We found out that gradual layer unfreezing was crucial condition for our models not to diverge. First 5 epochs we unfreeze 20% of layers from head each epoch. Presumably it was because of Mixed Precision - some sort of gradient explosion or something similar.\n- <strong>Metric learning - CosFace loss:</strong> We took my favorite metric learning loss. We also tried ArcLoss and AdaCos but they fail to converge or achieve worse results in Mixed precision mode.\n- <strong>Cyclic Linear Lr Annealing.</strong>\n- <strong>Cyclic Linear Scale annealing:</strong> This idea is inspired by AdaCos.  <a href=\"https://arxiv.org/abs/1905.00292\">https://arxiv.org/abs/1905.00292</a> where authors decrease scale as model achieve high scores. S_max = 64, S_min = 24.\n- <strong>Classes for distinct cell types are distinct as well:</strong> So we predict vector of 1108 * 4 + control classes. It  improved convergence drastically.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F362676%2F318492d44eeea2d6e21b5ec53a34ffac%2FScreenshot%20from%202019-09-24%2020-27-00.png?generation=1569525109617780&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li><p><strong>Unsupervised Domain Adaptation by Backpropagation:</strong> <a href=\"https://arxiv.org/abs/1409.7495\">https://arxiv.org/abs/1409.7495</a>. We took plate and cell type as domain label and schedule gradient reversal layer coefficient as lr and Scale for CosFace. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F362676%2F81d59df47d3bf1faf271f11d74245223%2FScreenshot%20from%202019-09-24%2020-55-48.png?generation=1569525156026046&amp;alt=media\" alt=\"\"></p></li>\n<li><p><strong>Backbones:</strong> seresnet50 (img_size up to 960), senet154 (img_size up to 672), xception (img_size up to 860), inceptionresnetv2 (img_size up to 512), densenet161 (img_size up to 840), densenet201 (img_size up to 720). </p></li>\n<li>SGD optimizer with momentum. We tried Adam, Lookahead and Ranger but they didn’t improve validation score.</li>\n<li>Gradient normalization.</li>\n<li>Batch size = 16 with accumulation up to 32.</li>\n<li>Focal loss: gamma = 32.</li>\n</ul>\n\n<p><strong>Prediction:</strong>\n- <strong>Prediction balancing.</strong> This technique can be used when classes are equally balanced. Excellent implementation: <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270</a>.\n- <strong>Progressive prediction.</strong> Idea is simple: for every experiment each class can be presented only once. So if we predict some class with high probability there is a little chance that this class can be predicted further with lower probability. We take 50% most probable classes, remove them from consideration and predict the next sample and so on.\n- <strong>Leak.</strong></p>\n\n<p><strong>What didn’t work:</strong>\n- Mix-up\n- ArcFace\n- AdaCos\n- Deformable ConvNet - <a href=\"https://arxiv.org/pdf/1811.11168.pdf\">https://arxiv.org/pdf/1811.11168.pdf</a>\n- Deep Supervision\n- Weight Standardisation\n- Concatenation of two sites together to make 12-channel image</p>",
  "messages": [
    {
      "id": 635088,
      "postDate": "2019-09-27T06:34:51.840Z",
      "content": "<p>Dear all, thank you for excellent spirit of competition and astonishing results! Without such favorable pressure we wouldn’t be able to achieve such scores.</p>\n\n<p>I will split description into two main parts: pipeline and prediction with post-processing.</p>\n\n<p><strong>Pipeline:</strong>\n- <strong>6-channel images.</strong>\n- <strong>Super Resolution.</strong> During exploratory data analysis we suddenly realize that the size of individual cells is rather small and convolutional layer might have a hard time learning all patterns. So we use simple bicubic superresolution technique. That really boost our scores. We checked up to 960 * 960. We have tried to use GAN’s (ESRGAN in particular) for this task but its performance was worse compared to vanilla bicubic interpolation.\n- <strong>3Fold by experiments.</strong>\n- <strong>Add all controls to train.</strong> We tried to add controls only from train experiments, but it didn’t work.\n- <strong>Mixed Precision:</strong> Due to the large size of images we decided to use only Mixed Precision learning because of inevitable graphics card’s memory shortage.\n- <strong>Gradual layer unfreezing:</strong> We found out that gradual layer unfreezing was crucial condition for our models not to diverge. First 5 epochs we unfreeze 20% of layers from head each epoch. Presumably it was because of Mixed Precision - some sort of gradient explosion or something similar.\n- <strong>Metric learning - CosFace loss:</strong> We took my favorite metric learning loss. We also tried ArcLoss and AdaCos but they fail to converge or achieve worse results in Mixed precision mode.\n- <strong>Cyclic Linear Lr Annealing.</strong>\n- <strong>Cyclic Linear Scale annealing:</strong> This idea is inspired by AdaCos.  <a href=\"https://arxiv.org/abs/1905.00292\">https://arxiv.org/abs/1905.00292</a> where authors decrease scale as model achieve high scores. S_max = 64, S_min = 24.\n- <strong>Classes for distinct cell types are distinct as well:</strong> So we predict vector of 1108 * 4 + control classes. It  improved convergence drastically.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F362676%2F318492d44eeea2d6e21b5ec53a34ffac%2FScreenshot%20from%202019-09-24%2020-27-00.png?generation=1569525109617780&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li><p><strong>Unsupervised Domain Adaptation by Backpropagation:</strong> <a href=\"https://arxiv.org/abs/1409.7495\">https://arxiv.org/abs/1409.7495</a>. We took plate and cell type as domain label and schedule gradient reversal layer coefficient as lr and Scale for CosFace. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F362676%2F81d59df47d3bf1faf271f11d74245223%2FScreenshot%20from%202019-09-24%2020-55-48.png?generation=1569525156026046&amp;alt=media\" alt=\"\"></p></li>\n<li><p><strong>Backbones:</strong> seresnet50 (img_size up to 960), senet154 (img_size up to 672), xception (img_size up to 860), inceptionresnetv2 (img_size up to 512), densenet161 (img_size up to 840), densenet201 (img_size up to 720). </p></li>\n<li>SGD optimizer with momentum. We tried Adam, Lookahead and Ranger but they didn’t improve validation score.</li>\n<li>Gradient normalization.</li>\n<li>Batch size = 16 with accumulation up to 32.</li>\n<li>Focal loss: gamma = 32.</li>\n</ul>\n\n<p><strong>Prediction:</strong>\n- <strong>Prediction balancing.</strong> This technique can be used when classes are equally balanced. Excellent implementation: <a href=\"https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270\">https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270</a>.\n- <strong>Progressive prediction.</strong> Idea is simple: for every experiment each class can be presented only once. So if we predict some class with high probability there is a little chance that this class can be predicted further with lower probability. We take 50% most probable classes, remove them from consideration and predict the next sample and so on.\n- <strong>Leak.</strong></p>\n\n<p><strong>What didn’t work:</strong>\n- Mix-up\n- ArcFace\n- AdaCos\n- Deformable ConvNet - <a href=\"https://arxiv.org/pdf/1811.11168.pdf\">https://arxiv.org/pdf/1811.11168.pdf</a>\n- Deep Supervision\n- Weight Standardisation\n- Concatenation of two sites together to make 12-channel image</p>",
      "rawMarkdown": "Dear all, thank you for excellent spirit of competition and astonishing results! Without such favorable pressure we wouldn’t be able to achieve such scores.\n\nI will split description into two main parts: pipeline and prediction with post-processing.\n\n**Pipeline:**\n- **6-channel images.**\n- **Super Resolution.** During exploratory data analysis we suddenly realize that the size of individual cells is rather small and convolutional layer might have a hard time learning all patterns. So we use simple bicubic superresolution technique. That really boost our scores. We checked up to 960 * 960. We have tried to use GAN’s (ESRGAN in particular) for this task but its performance was worse compared to vanilla bicubic interpolation.\n- **3Fold by experiments.**\n- **Add all controls to train.** We tried to add controls only from train experiments, but it didn’t work.\n- **Mixed Precision:** Due to the large size of images we decided to use only Mixed Precision learning because of inevitable graphics card’s memory shortage.\n- **Gradual layer unfreezing:** We found out that gradual layer unfreezing was crucial condition for our models not to diverge. First 5 epochs we unfreeze 20% of layers from head each epoch. Presumably it was because of Mixed Precision - some sort of gradient explosion or something similar.\n- **Metric learning - CosFace loss:** We took my favorite metric learning loss. We also tried ArcLoss and AdaCos but they fail to converge or achieve worse results in Mixed precision mode.\n- **Cyclic Linear Lr Annealing.**\n- **Cyclic Linear Scale annealing:** This idea is inspired by AdaCos.  https://arxiv.org/abs/1905.00292 where authors decrease scale as model achieve high scores. S_max = 64, S_min = 24.\n- **Classes for distinct cell types are distinct as well:** So we predict vector of 1108 * 4 + control classes. It  improved convergence drastically.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F362676%2F318492d44eeea2d6e21b5ec53a34ffac%2FScreenshot%20from%202019-09-24%2020-27-00.png?generation=1569525109617780&amp;alt=media)\n\n- **Unsupervised Domain Adaptation by Backpropagation:** https://arxiv.org/abs/1409.7495. We took plate and cell type as domain label and schedule gradient reversal layer coefficient as lr and Scale for CosFace. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F362676%2F81d59df47d3bf1faf271f11d74245223%2FScreenshot%20from%202019-09-24%2020-55-48.png?generation=1569525156026046&amp;alt=media)\n\n- **Backbones:** seresnet50 (img_size up to 960), senet154 (img_size up to 672), xception (img_size up to 860), inceptionresnetv2 (img_size up to 512), densenet161 (img_size up to 840), densenet201 (img_size up to 720). \n- SGD optimizer with momentum. We tried Adam, Lookahead and Ranger but they didn’t improve validation score.\n- Gradient normalization.\n- Batch size = 16 with accumulation up to 32.\n- Focal loss: gamma = 32.\n\n\n**Prediction:**\n- **Prediction balancing.** This technique can be used when classes are equally balanced. Excellent implementation: https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270.\n- **Progressive prediction.** Idea is simple: for every experiment each class can be presented only once. So if we predict some class with high probability there is a little chance that this class can be predicted further with lower probability. We take 50% most probable classes, remove them from consideration and predict the next sample and so on.\n- **Leak.**\n\n\n**What didn’t work:**\n- Mix-up\n- ArcFace\n- AdaCos\n- Deformable ConvNet - https://arxiv.org/pdf/1811.11168.pdf\n- Deep Supervision\n- Weight Standardisation\n- Concatenation of two sites together to make 12-channel image\n\n",
      "votes": 39
    },
    {
      "id": 635102,
      "postDate": "2019-09-27T06:55:10.760Z",
      "content": "<p>Thank you so much for your writeup. \nI havent worked with metric learning much. \nSorry if it is a stupid question. How can you have softmax probability from metric learning model so that you can apply balance technique?</p>",
      "rawMarkdown": "Thank you so much for your writeup. \nI havent worked with metric learning much. \nSorry if it is a stupid question. How can you have softmax probability from metric learning model so that you can apply balance technique?",
      "votes": 1,
      "replies": [
        {
          "id": 635108,
          "postDate": "2019-09-27T07:02:35.673Z",
          "content": "<p><a href=\"/backaggle\">@backaggle</a>, [Name]Face block outputs probabilities that further goes to cross-entropy loss. So you can use embeddings or probabilities from [Name]Face block. </p>",
          "rawMarkdown": "@backaggle, [Name]Face block outputs probabilities that further goes to cross-entropy loss. So you can use embeddings or probabilities from [Name]Face block. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 635123,
      "postDate": "2019-09-27T07:19:12.960Z",
      "content": "<p>Congratulations to your team on winning gold medal position and being well stable in LB. Is it possible for you to open-source the code solution, even the baseline structure? I would like to recreate and re-run it. many thanks again!</p>",
      "rawMarkdown": "Congratulations to your team on winning gold medal position and being well stable in LB. Is it possible for you to open-source the code solution, even the baseline structure? I would like to recreate and re-run it. many thanks again!"
    },
    {
      "id": 635147,
      "postDate": "2019-09-27T07:45:41.667Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 635145,
      "postDate": "2019-09-27T07:44:39.697Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 635102,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2019-09-27T06:55:10.760000",
      "content": "<p>Thank you so much for your writeup. \nI havent worked with metric learning much. \nSorry if it is a stupid question. How can you have softmax probability from metric learning model so that you can apply balance technique?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 635108,
          "author_name": "Ivan Sosin",
          "author_url": "",
          "post_date": "2019-09-27T07:02:35.673000",
          "content": "<p><a href=\"/backaggle\">@backaggle</a>, [Name]Face block outputs probabilities that further goes to cross-entropy loss. So you can use embeddings or probabilities from [Name]Face block. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 635123,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2019-09-27T07:19:12.960000",
      "content": "<p>Congratulations to your team on winning gold medal position and being well stable in LB. Is it possible for you to open-source the code solution, even the baseline structure? I would like to recreate and re-run it. many thanks again!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635147,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-27T07:45:41.667000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 635145,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-27T07:44:39.697000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "635088": "Dear all, thank you for excellent spirit of competition and astonishing results! Without such favorable pressure we wouldn’t be able to achieve such scores.\n\nI will split description into two main parts: pipeline and prediction with post-processing.\n\n**Pipeline:**\n- **6-channel images.**\n- **Super Resolution.** During exploratory data analysis we suddenly realize that the size of individual cells is rather small and convolutional layer might have a hard time learning all patterns. So we use simple bicubic superresolution technique. That really boost our scores. We checked up to 960 * 960. We have tried to use GAN’s (ESRGAN in particular) for this task but its performance was worse compared to vanilla bicubic interpolation.\n- **3Fold by experiments.**\n- **Add all controls to train.** We tried to add controls only from train experiments, but it didn’t work.\n- **Mixed Precision:** Due to the large size of images we decided to use only Mixed Precision learning because of inevitable graphics card’s memory shortage.\n- **Gradual layer unfreezing:** We found out that gradual layer unfreezing was crucial condition for our models not to diverge. First 5 epochs we unfreeze 20% of layers from head each epoch. Presumably it was because of Mixed Precision - some sort of gradient explosion or something similar.\n- **Metric learning - CosFace loss:** We took my favorite metric learning loss. We also tried ArcLoss and AdaCos but they fail to converge or achieve worse results in Mixed precision mode.\n- **Cyclic Linear Lr Annealing.**\n- **Cyclic Linear Scale annealing:** This idea is inspired by AdaCos.  https://arxiv.org/abs/1905.00292 where authors decrease scale as model achieve high scores. S_max = 64, S_min = 24.\n- **Classes for distinct cell types are distinct as well:** So we predict vector of 1108 * 4 + control classes. It  improved convergence drastically.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F362676%2F318492d44eeea2d6e21b5ec53a34ffac%2FScreenshot%20from%202019-09-24%2020-27-00.png?generation=1569525109617780&amp;alt=media)\n\n- **Unsupervised Domain Adaptation by Backpropagation:** https://arxiv.org/abs/1409.7495. We took plate and cell type as domain label and schedule gradient reversal layer coefficient as lr and Scale for CosFace. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F362676%2F81d59df47d3bf1faf271f11d74245223%2FScreenshot%20from%202019-09-24%2020-55-48.png?generation=1569525156026046&amp;alt=media)\n\n- **Backbones:** seresnet50 (img_size up to 960), senet154 (img_size up to 672), xception (img_size up to 860), inceptionresnetv2 (img_size up to 512), densenet161 (img_size up to 840), densenet201 (img_size up to 720). \n- SGD optimizer with momentum. We tried Adam, Lookahead and Ranger but they didn’t improve validation score.\n- Gradient normalization.\n- Batch size = 16 with accumulation up to 32.\n- Focal loss: gamma = 32.\n\n\n**Prediction:**\n- **Prediction balancing.** This technique can be used when classes are equally balanced. Excellent implementation: https://www.kaggle.com/c/quickdraw-doodle-recognition/discussion/73803#latest-438270.\n- **Progressive prediction.** Idea is simple: for every experiment each class can be presented only once. So if we predict some class with high probability there is a little chance that this class can be predicted further with lower probability. We take 50% most probable classes, remove them from consideration and predict the next sample and so on.\n- **Leak.**\n\n\n**What didn’t work:**\n- Mix-up\n- ArcFace\n- AdaCos\n- Deformable ConvNet - https://arxiv.org/pdf/1811.11168.pdf\n- Deep Supervision\n- Weight Standardisation\n- Concatenation of two sites together to make 12-channel image\n\n",
    "635102": "Thank you so much for your writeup. \nI havent worked with metric learning much. \nSorry if it is a stupid question. How can you have softmax probability from metric learning model so that you can apply balance technique?",
    "635123": "Congratulations to your team on winning gold medal position and being well stable in LB. Is it possible for you to open-source the code solution, even the baseline structure? I would like to recreate and re-run it. many thanks again!",
    "635147": "",
    "635145": ""
  }
}