{
  "id": 110337,
  "title": "4th place solution",
  "url": "/competitions/recursion-cellular-image-classification/discussion/110337",
  "author_name": "yu4u",
  "post_date": "2019-09-27T01:02:06.428000",
  "votes": 44,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Congrats to all prize and medal winners!\nOur brief solution summary:</p>\n\n<p>1) heavy model ensemble\n  - 1108-way classification\n  - seresnext50, 101, densenet, efficientnet x 6C6, 5C6 4C6 channel selection x cross validation\n  - 512x512 input, 90rot+clip aug</p>\n\n<p>2) solve linear sum assignment problem to make the most of 'the groups of 277 per plate restriction'\n<a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-624588\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-624588</a></p>\n\n<p><code>\nfrom scipy.optimize import linear_sum_assignment\nplate_prob = plate_prob / plate_prob.sum(axis=0, keepdims=True)\nrow_ind, col_ind = linear_sum_assignment(1 - plate_prob)\n</code></p>\n\n<p>3) make psuedo label and go back to 1 (with a few selected models like efficientnet)</p>\n\n<p>The code can be found in <a href=\"https://github.com/ngxbac/Kaggle-Recursion-Cellular\">https://github.com/ngxbac/Kaggle-Recursion-Cellular</a>\n(mainly the 1) part)</p>",
  "messages": [
    {
      "id": 634933,
      "postDate": "2019-09-27T01:02:06.430Z",
      "content": "<p>Congrats to all prize and medal winners!\nOur brief solution summary:</p>\n\n<p>1) heavy model ensemble\n  - 1108-way classification\n  - seresnext50, 101, densenet, efficientnet x 6C6, 5C6 4C6 channel selection x cross validation\n  - 512x512 input, 90rot+clip aug</p>\n\n<p>2) solve linear sum assignment problem to make the most of 'the groups of 277 per plate restriction'\n<a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-624588\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-624588</a></p>\n\n<p><code>\nfrom scipy.optimize import linear_sum_assignment\nplate_prob = plate_prob / plate_prob.sum(axis=0, keepdims=True)\nrow_ind, col_ind = linear_sum_assignment(1 - plate_prob)\n</code></p>\n\n<p>3) make psuedo label and go back to 1 (with a few selected models like efficientnet)</p>\n\n<p>The code can be found in <a href=\"https://github.com/ngxbac/Kaggle-Recursion-Cellular\">https://github.com/ngxbac/Kaggle-Recursion-Cellular</a>\n(mainly the 1) part)</p>",
      "rawMarkdown": "Congrats to all prize and medal winners!\nOur brief solution summary:\n\n1) heavy model ensemble\n  - 1108-way classification\n  - seresnext50, 101, densenet, efficientnet x 6C6, 5C6 4C6 channel selection x cross validation\n  - 512x512 input, 90rot+clip aug\n\n2) solve linear sum assignment problem to make the most of 'the groups of 277 per plate restriction'\nhttps://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-624588\n\n```\nfrom scipy.optimize import linear_sum_assignment\nplate_prob = plate_prob / plate_prob.sum(axis=0, keepdims=True)\nrow_ind, col_ind = linear_sum_assignment(1 - plate_prob)\n```\n\n3) make psuedo label and go back to 1 (with a few selected models like efficientnet)\n\nThe code can be found in https://github.com/ngxbac/Kaggle-Recursion-Cellular\n(mainly the 1) part)",
      "votes": 44
    },
    {
      "id": 634969,
      "postDate": "2019-09-27T02:21:35.213Z",
      "content": "<p>Thank yu4u for your brief writeup. <br>\nI would love to share my insights in this competition. </p>\n\n<h1>How many channels?</h1>\n\n<p>I try different combination of channels: [1,2,3,4,5], [1,2,3,4,6], ... and recognize that different combinations give different performances on different siRNAs and they are really diversity. <br>\nI am not the domain expert, so I really appriciate if someone can confirm this: My hypothesis is, each siRNA should need different information from different channels. <br>\nEx: siRNA = 1 needs [1,2,3] for recognize while siRNA = 2 needs [4,5,6] and so on. <br>\nThis obsevervation brings me to train multiple combinations of channel and ensemble them.</p>\n\n<h1>Loss function</h1>\n\n<p>I use <code>LabelSmoothingCrossEntropy</code> loss which is better than CrossEntropy. I was surprised and thought that: <br>\n<code>May label be wrong?</code>. </p>\n\n<p>But no, I think label is not wrong. It is a <code>weak label</code> because of the way we use channels. As the example above, if we use [4,5,6] to train for siRNA = 1, it does not make sense. So, one again, this loss function also prove my hypothesis. </p>\n\n<h1>Data augmentations</h1>\n\n<p>I use \n<code>\nRandomRotate90(),\nFlip(),\nTranspose(),\nIAAAffine(shear=(-10, 10)),\n</code></p>\n\n<p>Especially, <code>ChannelDropout</code> increases the performance 2%. We random drop the channel during training. It provides diversity and removes useless channel of combination. This point also prove for the hypothesis above. </p>\n\n<h1>Learning rate, Optimizer and scheduler</h1>\n\n<p>We train two stages: \n- First 3 epochs with Nadam, LR=0.001 for FC layers only. \n- 40 epochs with Nadam for all network. We use <code>OneCycleLR</code> as follows\n<code>\n      scheduler: OneCycleLR\n      num_steps: &amp;amp;num_epochs 40\n      lr_range: [0.0005, 0.00001]\n      # lr_range: [0.0015, 0.00003]\n      warmup_steps: 5\n      momentum_range: [0.85, 0.95]\n</code></p>\n\n<h1>How to use control images</h1>\n\n<p>Using control images boosts us 2%. \nFirst, we train the model with control images only (31 classes classifier). Second, we take the pretrained weight and continue training with 1108 classes. It should convex faster. </p>\n\n<h1>How many sites ?</h1>\n\n<p>We random select one site for training and average 2 sites for testing. It helps much. </p>",
      "rawMarkdown": "Thank yu4u for your brief writeup.  \nI would love to share my insights in this competition. \n\n# How many channels? \nI try different combination of channels: [1,2,3,4,5], [1,2,3,4,6], ... and recognize that different combinations give different performances on different siRNAs and they are really diversity.  \nI am not the domain expert, so I really appriciate if someone can confirm this: My hypothesis is, each siRNA should need different information from different channels.  \nEx: siRNA = 1 needs [1,2,3] for recognize while siRNA = 2 needs [4,5,6] and so on.  \nThis obsevervation brings me to train multiple combinations of channel and ensemble them.\n\n# Loss function \nI use `LabelSmoothingCrossEntropy` loss which is better than CrossEntropy. I was surprised and thought that:  \n`May label be wrong?`. \n\nBut no, I think label is not wrong. It is a `weak label` because of the way we use channels. As the example above, if we use [4,5,6] to train for siRNA = 1, it does not make sense. So, one again, this loss function also prove my hypothesis. \n\n# Data augmentations \nI use \n```\nRandomRotate90(),\nFlip(),\nTranspose(),\nIAAAffine(shear=(-10, 10)),\n```\n\nEspecially, `ChannelDropout` increases the performance 2%. We random drop the channel during training. It provides diversity and removes useless channel of combination. This point also prove for the hypothesis above. \n\n# Learning rate, Optimizer and scheduler \nWe train two stages: \n- First 3 epochs with Nadam, LR=0.001 for FC layers only. \n- 40 epochs with Nadam for all network. We use `OneCycleLR` as follows\n```\n      scheduler: OneCycleLR\n      num_steps: &amp;num_epochs 40\n      lr_range: [0.0005, 0.00001]\n      # lr_range: [0.0015, 0.00003]\n      warmup_steps: 5\n      momentum_range: [0.85, 0.95]\n```\n\n# How to use control images \nUsing control images boosts us 2%. \nFirst, we train the model with control images only (31 classes classifier). Second, we take the pretrained weight and continue training with 1108 classes. It should convex faster. \n\n# How many sites ? \nWe random select one site for training and average 2 sites for testing. It helps much. ",
      "votes": 11,
      "replies": [
        {
          "id": 635029,
          "postDate": "2019-09-27T04:51:44.327Z",
          "content": "<p>Congratulations...\nThank you for Sharing your Approach &amp; Insights...!! <a href=\"/backaggle\">@backaggle</a> <a href=\"/ren4yu\">@ren4yu</a> </p>",
          "rawMarkdown": "Congratulations...\nThank you for Sharing your Approach &amp; Insights...!! @backaggle @ren4yu ",
          "votes": 1
        },
        {
          "id": 635051,
          "postDate": "2019-09-27T05:23:33.740Z",
          "content": "<p>Congratulations! Thank you very much for sharing.</p>",
          "rawMarkdown": "Congratulations! Thank you very much for sharing.",
          "votes": 1
        },
        {
          "id": 635132,
          "postDate": "2019-09-27T07:28:08.037Z",
          "content": "<p>Hi, thanks for your insights!\nDid you try all possible combinations of channels (e.g. there are 15 combinations just for subset of 4 channels!) or reduced the number of it somehow?</p>",
          "rawMarkdown": "Hi, thanks for your insights!\nDid you try all possible combinations of channels (e.g. there are 15 combinations just for subset of 4 channels!) or reduced the number of it somehow?"
        },
        {
          "id": 635296,
          "postDate": "2019-09-27T10:44:28.750Z",
          "content": "<p>I tried all 6C4 and 6C5 :D. The training time was too long. seresnext101 with kfold 6C4 took 32days </p>",
          "rawMarkdown": "I tried all 6C4 and 6C5 :D. The training time was too long. seresnext101 with kfold 6C4 took 32days ",
          "votes": 1
        }
      ]
    },
    {
      "id": 634988,
      "postDate": "2019-09-27T02:59:33.493Z",
      "content": "<p>many thanks for sharing your code solutions! Congratulations on winning the gold medal! :-)</p>",
      "rawMarkdown": "many thanks for sharing your code solutions! Congratulations on winning the gold medal! :-)",
      "votes": 1
    },
    {
      "id": 635139,
      "postDate": "2019-09-27T07:36:43.497Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 634969,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2019-09-27T02:21:35.213000",
      "content": "<p>Thank yu4u for your brief writeup. <br>\nI would love to share my insights in this competition. </p>\n\n<h1>How many channels?</h1>\n\n<p>I try different combination of channels: [1,2,3,4,5], [1,2,3,4,6], ... and recognize that different combinations give different performances on different siRNAs and they are really diversity. <br>\nI am not the domain expert, so I really appriciate if someone can confirm this: My hypothesis is, each siRNA should need different information from different channels. <br>\nEx: siRNA = 1 needs [1,2,3] for recognize while siRNA = 2 needs [4,5,6] and so on. <br>\nThis obsevervation brings me to train multiple combinations of channel and ensemble them.</p>\n\n<h1>Loss function</h1>\n\n<p>I use <code>LabelSmoothingCrossEntropy</code> loss which is better than CrossEntropy. I was surprised and thought that: <br>\n<code>May label be wrong?</code>. </p>\n\n<p>But no, I think label is not wrong. It is a <code>weak label</code> because of the way we use channels. As the example above, if we use [4,5,6] to train for siRNA = 1, it does not make sense. So, one again, this loss function also prove my hypothesis. </p>\n\n<h1>Data augmentations</h1>\n\n<p>I use \n<code>\nRandomRotate90(),\nFlip(),\nTranspose(),\nIAAAffine(shear=(-10, 10)),\n</code></p>\n\n<p>Especially, <code>ChannelDropout</code> increases the performance 2%. We random drop the channel during training. It provides diversity and removes useless channel of combination. This point also prove for the hypothesis above. </p>\n\n<h1>Learning rate, Optimizer and scheduler</h1>\n\n<p>We train two stages: \n- First 3 epochs with Nadam, LR=0.001 for FC layers only. \n- 40 epochs with Nadam for all network. We use <code>OneCycleLR</code> as follows\n<code>\n      scheduler: OneCycleLR\n      num_steps: &amp;amp;num_epochs 40\n      lr_range: [0.0005, 0.00001]\n      # lr_range: [0.0015, 0.00003]\n      warmup_steps: 5\n      momentum_range: [0.85, 0.95]\n</code></p>\n\n<h1>How to use control images</h1>\n\n<p>Using control images boosts us 2%. \nFirst, we train the model with control images only (31 classes classifier). Second, we take the pretrained weight and continue training with 1108 classes. It should convex faster. </p>\n\n<h1>How many sites ?</h1>\n\n<p>We random select one site for training and average 2 sites for testing. It helps much. </p>",
      "votes": 11,
      "replies": [
        {
          "id": 635029,
          "author_name": "Ailurophile",
          "author_url": "",
          "post_date": "2019-09-27T04:51:44.327000",
          "content": "<p>Congratulations...\nThank you for Sharing your Approach &amp; Insights...!! <a href=\"/backaggle\">@backaggle</a> <a href=\"/ren4yu\">@ren4yu</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 635051,
          "author_name": "Putalay",
          "author_url": "",
          "post_date": "2019-09-27T05:23:33.740000",
          "content": "<p>Congratulations! Thank you very much for sharing.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 635132,
          "author_name": "Vladislav Myrov",
          "author_url": "",
          "post_date": "2019-09-27T07:28:08.037000",
          "content": "<p>Hi, thanks for your insights!\nDid you try all possible combinations of channels (e.g. there are 15 combinations just for subset of 4 channels!) or reduced the number of it somehow?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 635296,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-09-27T10:44:28.750000",
          "content": "<p>I tried all 6C4 and 6C5 :D. The training time was too long. seresnext101 with kfold 6C4 took 32days </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 634988,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2019-09-27T02:59:33.493000",
      "content": "<p>many thanks for sharing your code solutions! Congratulations on winning the gold medal! :-)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 635139,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-27T07:36:43.497000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "634933": "Congrats to all prize and medal winners!\nOur brief solution summary:\n\n1) heavy model ensemble\n  - 1108-way classification\n  - seresnext50, 101, densenet, efficientnet x 6C6, 5C6 4C6 channel selection x cross validation\n  - 512x512 input, 90rot+clip aug\n\n2) solve linear sum assignment problem to make the most of 'the groups of 277 per plate restriction'\nhttps://www.kaggle.com/c/recursion-cellular-image-classification/discussion/102905#latest-624588\n\n```\nfrom scipy.optimize import linear_sum_assignment\nplate_prob = plate_prob / plate_prob.sum(axis=0, keepdims=True)\nrow_ind, col_ind = linear_sum_assignment(1 - plate_prob)\n```\n\n3) make psuedo label and go back to 1 (with a few selected models like efficientnet)\n\nThe code can be found in https://github.com/ngxbac/Kaggle-Recursion-Cellular\n(mainly the 1) part)",
    "634969": "Thank yu4u for your brief writeup.  \nI would love to share my insights in this competition. \n\n# How many channels? \nI try different combination of channels: [1,2,3,4,5], [1,2,3,4,6], ... and recognize that different combinations give different performances on different siRNAs and they are really diversity.  \nI am not the domain expert, so I really appriciate if someone can confirm this: My hypothesis is, each siRNA should need different information from different channels.  \nEx: siRNA = 1 needs [1,2,3] for recognize while siRNA = 2 needs [4,5,6] and so on.  \nThis obsevervation brings me to train multiple combinations of channel and ensemble them.\n\n# Loss function \nI use `LabelSmoothingCrossEntropy` loss which is better than CrossEntropy. I was surprised and thought that:  \n`May label be wrong?`. \n\nBut no, I think label is not wrong. It is a `weak label` because of the way we use channels. As the example above, if we use [4,5,6] to train for siRNA = 1, it does not make sense. So, one again, this loss function also prove my hypothesis. \n\n# Data augmentations \nI use \n```\nRandomRotate90(),\nFlip(),\nTranspose(),\nIAAAffine(shear=(-10, 10)),\n```\n\nEspecially, `ChannelDropout` increases the performance 2%. We random drop the channel during training. It provides diversity and removes useless channel of combination. This point also prove for the hypothesis above. \n\n# Learning rate, Optimizer and scheduler \nWe train two stages: \n- First 3 epochs with Nadam, LR=0.001 for FC layers only. \n- 40 epochs with Nadam for all network. We use `OneCycleLR` as follows\n```\n      scheduler: OneCycleLR\n      num_steps: &amp;num_epochs 40\n      lr_range: [0.0005, 0.00001]\n      # lr_range: [0.0015, 0.00003]\n      warmup_steps: 5\n      momentum_range: [0.85, 0.95]\n```\n\n# How to use control images \nUsing control images boosts us 2%. \nFirst, we train the model with control images only (31 classes classifier). Second, we take the pretrained weight and continue training with 1108 classes. It should convex faster. \n\n# How many sites ? \nWe random select one site for training and average 2 sites for testing. It helps much. ",
    "634988": "many thanks for sharing your code solutions! Congratulations on winning the gold medal! :-)",
    "635139": ""
  }
}