{
  "id": 110391,
  "title": "6th place solution",
  "url": "/competitions/recursion-cellular-image-classification/writeups/looking-for-a-job-6th-place-solution",
  "author_name": "",
  "post_date": "2019-09-27T10:59:31.210Z",
  "votes": 25,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi kagglers! Nice one, congratulations to all the winners and all who've learned something new. Congratulations and many thanks my team Andriy ( <a href=\"/ayaroshevskiy\">@ayaroshevskiy</a> ), Sasha ( <a href=\"/pajari\">@pajari</a> ),  Vitaliy ( <a href=\"/ladler0320\">@ladler0320</a> ). Short summary of our very simple and straightforward solution.</p>\n\n<p>About \"<strong>leak</strong>\". I won't argue about the definition. But actually, that was simply target class co-occurance and I recommend to check this in every EDA of every multiclassification challenge same as you check class imbalance. EDA is must.\nFor example:\n<code>df.groupby([\"experiment\", \"plate\"]).sirna.apply(lambda x: x.sort_values().unique())</code>\nSo why do they co-occur that way? Why experiments are shuffled only within the plates? I guess its by design hard to shuffle them every single experiment for 1100+ classes but anyway that knowledge reduced an error significantly.</p>\n\n<p>Also, as you already know - crucial was to solve <strong>assignment problem</strong> having raw/softmax predicts from neural network. So we've used hungarian algorithm as well. </p>\n\n<ul>\n<li><p>Validation. I've used 30/33 experiments as a validation splitted by <strong>3 folds</strong>. (3 other experiments went to all the folds). Iteratively, I've manually swapped experiments to align their scores. HUVEC-18 went to the train too. We've found that public LB contains U2OS-04 and it's hard to achieve good score for that particular one. I removed it from most of our submissions to see score for other tree cells. But also that gave us a chance that someone will overfit to public LB. Private U2OS-05 was alright and looked much more like U2OS-03/U2OS-01.</p></li>\n<li><p>Image <strong>normalization</strong>. Organizers gave us pixel stats so we've used it in exact way - scaled all experiments by their mean/std. By the end of competition I've found that scaling U2OS-04 with mean/std of U2OS-02 increased its score by 10%. But who cares about public LB experiments ¯\\_(ツ)_/¯</p></li>\n<li><p>Networks. We've trained our first classification model with softmax - efficientnet-b0 in kaggle kernels before the gpus quotas.  Then in my 2x1080ti devbox. I didn't have much time for modeling because of lack of gpus but I've found EfficientNets work comparably better than densenets and resnets/se-resnets. One fold of b0 takes about 9 hours but the CPU was a bottleneck.</p></li>\n<li><p>Hard <strong>augmentations</strong>. All flips/rotates. Hard Brightness and Contrast with prob=0.8. Gamma, ChannelDropout. One thing I've found - uint8 augmentations work worse than float because of distribution discretization. </p></li>\n<li><p>Training process. I train everything with <strong>Adam</strong>, ReduceOnPlateau and few warmup iterations. I've spent so much time trying to understand why people declare SGD works better but no chance - Adam performed better as usual.</p></li>\n<li><p><strong>Controls</strong>. We've simply added train and test (!) controls as additional classes 1108:1139. That boosted score significantly. </p></li>\n</ul>\n\n<p><strong>First iteration</strong>. Pretrained <strong>EfficientNets</strong> b0, b1, b2 achieve almost the same score but blending their raw predictions and than solving assignment problem boosted score nicely. Averaging this 3 models scores 0.744/0.750 (3 cells, U2OS-04 reduced) public and 0.990 private LB (13th).</p>\n\n<p>I saw few strategies of how to split test experiments for <strong>pseudo-labeling</strong>. And actually I've tried around 6 of them with one or two iterations of PL. By the end I've came with very careful strategy - soft labels and balanced experiments split between folds:\n<code>\npl_all_folds = [\"HEPG2-10\", \"HUVEC-19\", \"RPE-10\", \"U2OS-05\"]\npl_folds = [[\"HEPG2-08\", \"HUVEC-17\", \"HUVEC-23\", \"RPE-08\"],\n        [\"HEPG2-09\", \"HUVEC-20\", \"HUVEC-21\", \"RPE-09\"],\n        [\"HEPG2-11\", \"HUVEC-22\", \"HUVEC-24\", \"RPE-11\"]]\n</code>\nAs you see U2OS-04 was reduced.\n<strong>Soft</strong> pseudo-labelling seemed to work better than hard. Now I understand that the best strategy might be to add them iteratively based on some threshold after solving assignment problem.</p>\n\n<p><strong>Second iteration</strong>. Same b0, b1, b2 now with pseudo-labels. People don't usual train same architectures in PL stage so I've added b4 and b6 as well. EfficientNet-b4 scored best validation and LB score and unfortunately EfficientNet-b6 didn't converge well in time. For big models we used <strong>gradient accumulations</strong> up to x12. TTA: 2 sites with 8 combinations of flip/rotate each. </p>\n\n<p>What didn't work:\n- SGD with CLRs (triangles, cosine etc)\n- Lookahead and Ranger\n- ArcFace (achieved almost the same score as Softmax)\n- My favourite se-resnext50</p>",
  "messages": [
    {
      "id": "635237",
      "postDate": "09/27/2019 09:33:16",
      "content": "<p>Hi kagglers! Nice one, congratulations to all the winners and all who've learned something new. Congratulations and many thanks my team Andriy ( <a href=\"/ayaroshevskiy\">@ayaroshevskiy</a> ), Sasha ( <a href=\"/pajari\">@pajari</a> ),  Vitaliy ( <a href=\"/ladler0320\">@ladler0320</a> ). Short summary of our very simple and straightforward solution.</p>\n\n<p>About \"<strong>leak</strong>\". I won't argue about the definition. But actually, that was simply target class co-occurance and I recommend to check this in every EDA of every multiclassification challenge same as you check class imbalance. EDA is must.\nFor example:\n<code>df.groupby([\"experiment\", \"plate\"]).sirna.apply(lambda x: x.sort_values().unique())</code>\nSo why do they co-occur that way? Why experiments are shuffled only within the plates? I guess its by design hard to shuffle them every single experiment for 1100+ classes but anyway that knowledge reduced an error significantly.</p>\n\n<p>Also, as you already know - crucial was to solve <strong>assignment problem</strong> having raw/softmax predicts from neural network. So we've used hungarian algorithm as well. </p>\n\n<ul>\n<li><p>Validation. I've used 30/33 experiments as a validation splitted by <strong>3 folds</strong>. (3 other experiments went to all the folds). Iteratively, I've manually swapped experiments to align their scores. HUVEC-18 went to the train too. We've found that public LB contains U2OS-04 and it's hard to achieve good score for that particular one. I removed it from most of our submissions to see score for other tree cells. But also that gave us a chance that someone will overfit to public LB. Private U2OS-05 was alright and looked much more like U2OS-03/U2OS-01.</p></li>\n<li><p>Image <strong>normalization</strong>. Organizers gave us pixel stats so we've used it in exact way - scaled all experiments by their mean/std. By the end of competition I've found that scaling U2OS-04 with mean/std of U2OS-02 increased its score by 10%. But who cares about public LB experiments ¯\\_(ツ)_/¯</p></li>\n<li><p>Networks. We've trained our first classification model with softmax - efficientnet-b0 in kaggle kernels before the gpus quotas.  Then in my 2x1080ti devbox. I didn't have much time for modeling because of lack of gpus but I've found EfficientNets work comparably better than densenets and resnets/se-resnets. One fold of b0 takes about 9 hours but the CPU was a bottleneck.</p></li>\n<li><p>Hard <strong>augmentations</strong>. All flips/rotates. Hard Brightness and Contrast with prob=0.8. Gamma, ChannelDropout. One thing I've found - uint8 augmentations work worse than float because of distribution discretization. </p></li>\n<li><p>Training process. I train everything with <strong>Adam</strong>, ReduceOnPlateau and few warmup iterations. I've spent so much time trying to understand why people declare SGD works better but no chance - Adam performed better as usual.</p></li>\n<li><p><strong>Controls</strong>. We've simply added train and test (!) controls as additional classes 1108:1139. That boosted score significantly. </p></li>\n</ul>\n\n<p><strong>First iteration</strong>. Pretrained <strong>EfficientNets</strong> b0, b1, b2 achieve almost the same score but blending their raw predictions and than solving assignment problem boosted score nicely. Averaging this 3 models scores 0.744/0.750 (3 cells, U2OS-04 reduced) public and 0.990 private LB (13th).</p>\n\n<p>I saw few strategies of how to split test experiments for <strong>pseudo-labeling</strong>. And actually I've tried around 6 of them with one or two iterations of PL. By the end I've came with very careful strategy - soft labels and balanced experiments split between folds:\n<code>\npl_all_folds = [\"HEPG2-10\", \"HUVEC-19\", \"RPE-10\", \"U2OS-05\"]\npl_folds = [[\"HEPG2-08\", \"HUVEC-17\", \"HUVEC-23\", \"RPE-08\"],\n        [\"HEPG2-09\", \"HUVEC-20\", \"HUVEC-21\", \"RPE-09\"],\n        [\"HEPG2-11\", \"HUVEC-22\", \"HUVEC-24\", \"RPE-11\"]]\n</code>\nAs you see U2OS-04 was reduced.\n<strong>Soft</strong> pseudo-labelling seemed to work better than hard. Now I understand that the best strategy might be to add them iteratively based on some threshold after solving assignment problem.</p>\n\n<p><strong>Second iteration</strong>. Same b0, b1, b2 now with pseudo-labels. People don't usual train same architectures in PL stage so I've added b4 and b6 as well. EfficientNet-b4 scored best validation and LB score and unfortunately EfficientNet-b6 didn't converge well in time. For big models we used <strong>gradient accumulations</strong> up to x12. TTA: 2 sites with 8 combinations of flip/rotate each. </p>\n\n<p>What didn't work:\n- SGD with CLRs (triangles, cosine etc)\n- Lookahead and Ranger\n- ArcFace (achieved almost the same score as Softmax)\n- My favourite se-resnext50</p>",
      "rawMarkdown": "Hi kagglers! Nice one, congratulations to all the winners and all who've learned something new. Congratulations and many thanks my team Andriy ( @ayaroshevskiy ), Sasha ( @pajari ),  Vitaliy ( @ladler0320 ). Short summary of our very simple and straightforward solution.\n\nAbout \"**leak**\". I won't argue about the definition. But actually, that was simply target class co-occurance and I recommend to check this in every EDA of every multiclassification challenge same as you check class imbalance. EDA is must.\nFor example:\n`df.groupby([\"experiment\", \"plate\"]).sirna.apply(lambda x: x.sort_values().unique())`\nSo why do they co-occur that way? Why experiments are shuffled only within the plates? I guess its by design hard to shuffle them every single experiment for 1100+ classes but anyway that knowledge reduced an error significantly.\n\nAlso, as you already know - crucial was to solve **assignment problem** having raw/softmax predicts from neural network. So we've used hungarian algorithm as well. \n\n- Validation. I've used 30/33 experiments as a validation splitted by **3 folds**. (3 other experiments went to all the folds). Iteratively, I've manually swapped experiments to align their scores. HUVEC-18 went to the train too. We've found that public LB contains U2OS-04 and it's hard to achieve good score for that particular one. I removed it from most of our submissions to see score for other tree cells. But also that gave us a chance that someone will overfit to public LB. Private U2OS-05 was alright and looked much more like U2OS-03/U2OS-01.\n\n- Image **normalization**. Organizers gave us pixel stats so we've used it in exact way - scaled all experiments by their mean/std. By the end of competition I've found that scaling U2OS-04 with mean/std of U2OS-02 increased its score by 10%. But who cares about public LB experiments ¯\\\\_(ツ)_/¯\n\n- Networks. We've trained our first classification model with softmax - efficientnet-b0 in kaggle kernels before the gpus quotas.  Then in my 2x1080ti devbox. I didn't have much time for modeling because of lack of gpus but I've found EfficientNets work comparably better than densenets and resnets/se-resnets. One fold of b0 takes about 9 hours but the CPU was a bottleneck.\n\n- Hard **augmentations**. All flips/rotates. Hard Brightness and Contrast with prob=0.8. Gamma, ChannelDropout. One thing I've found - uint8 augmentations work worse than float because of distribution discretization. \n\n- Training process. I train everything with **Adam**, ReduceOnPlateau and few warmup iterations. I've spent so much time trying to understand why people declare SGD works better but no chance - Adam performed better as usual.\n\n- **Controls**. We've simply added train and test (!) controls as additional classes 1108:1139. That boosted score significantly. \n\n**First iteration**. Pretrained **EfficientNets** b0, b1, b2 achieve almost the same score but blending their raw predictions and than solving assignment problem boosted score nicely. Averaging this 3 models scores 0.744/0.750 (3 cells, U2OS-04 reduced) public and 0.990 private LB (13th).\n\nI saw few strategies of how to split test experiments for **pseudo-labeling**. And actually I've tried around 6 of them with one or two iterations of PL. By the end I've came with very careful strategy - soft labels and balanced experiments split between folds:\n```\npl_all_folds = [\"HEPG2-10\", \"HUVEC-19\", \"RPE-10\", \"U2OS-05\"]\npl_folds = [[\"HEPG2-08\", \"HUVEC-17\", \"HUVEC-23\", \"RPE-08\"],\n        [\"HEPG2-09\", \"HUVEC-20\", \"HUVEC-21\", \"RPE-09\"],\n        [\"HEPG2-11\", \"HUVEC-22\", \"HUVEC-24\", \"RPE-11\"]]\n```\nAs you see U2OS-04 was reduced.\n**Soft** pseudo-labelling seemed to work better than hard. Now I understand that the best strategy might be to add them iteratively based on some threshold after solving assignment problem.\n\n**Second iteration**. Same b0, b1, b2 now with pseudo-labels. People don't usual train same architectures in PL stage so I've added b4 and b6 as well. EfficientNet-b4 scored best validation and LB score and unfortunately EfficientNet-b6 didn't converge well in time. For big models we used **gradient accumulations** up to x12. TTA: 2 sites with 8 combinations of flip/rotate each. \n\nWhat didn't work:\n- SGD with CLRs (triangles, cosine etc)\n- Lookahead and Ranger\n- ArcFace (achieved almost the same score as Softmax)\n- My favourite se-resnext50",
      "votes": null
    },
    {
      "id": "635268",
      "postDate": "09/27/2019 10:23:31",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights...!! <a href=\"/yaroshevskiy\">@yaroshevskiy</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights...!! @yaroshevskiy",
      "votes": null
    },
    {
      "id": "635272",
      "postDate": "09/27/2019 10:28:11",
      "content": "<p>My favourite backbone is also se-resnext50 : )\nCongrats!</p>\n\n<blockquote>\n  <p>Now I understand that the best strategy might be to add them iteratively based on some threshold after solving assignment problem.</p>\n</blockquote>\n\n<p>Lesson learned, thanks!</p>",
      "rawMarkdown": "My favourite backbone is also se-resnext50 : )\nCongrats!\n&gt; Now I understand that the best strategy might be to add them iteratively based on some threshold after solving assignment problem.\n\nLesson learned, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 635268,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "09/27/2019 10:23:31",
      "content": "<p>Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights...!! <a href=\"/yaroshevskiy\">@yaroshevskiy</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 635272,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "09/27/2019 10:28:11",
      "content": "<p>My favourite backbone is also se-resnext50 : )\nCongrats!</p>\n\n<blockquote>\n  <p>Now I understand that the best strategy might be to add them iteratively based on some threshold after solving assignment problem.</p>\n</blockquote>\n\n<p>Lesson learned, thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "635237": "Hi kagglers! Nice one, congratulations to all the winners and all who've learned something new. Congratulations and many thanks my team Andriy ( @ayaroshevskiy ), Sasha ( @pajari ),  Vitaliy ( @ladler0320 ). Short summary of our very simple and straightforward solution.\n\nAbout \"**leak**\". I won't argue about the definition. But actually, that was simply target class co-occurance and I recommend to check this in every EDA of every multiclassification challenge same as you check class imbalance. EDA is must.\nFor example:\n`df.groupby([\"experiment\", \"plate\"]).sirna.apply(lambda x: x.sort_values().unique())`\nSo why do they co-occur that way? Why experiments are shuffled only within the plates? I guess its by design hard to shuffle them every single experiment for 1100+ classes but anyway that knowledge reduced an error significantly.\n\nAlso, as you already know - crucial was to solve **assignment problem** having raw/softmax predicts from neural network. So we've used hungarian algorithm as well. \n\n- Validation. I've used 30/33 experiments as a validation splitted by **3 folds**. (3 other experiments went to all the folds). Iteratively, I've manually swapped experiments to align their scores. HUVEC-18 went to the train too. We've found that public LB contains U2OS-04 and it's hard to achieve good score for that particular one. I removed it from most of our submissions to see score for other tree cells. But also that gave us a chance that someone will overfit to public LB. Private U2OS-05 was alright and looked much more like U2OS-03/U2OS-01.\n\n- Image **normalization**. Organizers gave us pixel stats so we've used it in exact way - scaled all experiments by their mean/std. By the end of competition I've found that scaling U2OS-04 with mean/std of U2OS-02 increased its score by 10%. But who cares about public LB experiments ¯\\\\_(ツ)_/¯\n\n- Networks. We've trained our first classification model with softmax - efficientnet-b0 in kaggle kernels before the gpus quotas.  Then in my 2x1080ti devbox. I didn't have much time for modeling because of lack of gpus but I've found EfficientNets work comparably better than densenets and resnets/se-resnets. One fold of b0 takes about 9 hours but the CPU was a bottleneck.\n\n- Hard **augmentations**. All flips/rotates. Hard Brightness and Contrast with prob=0.8. Gamma, ChannelDropout. One thing I've found - uint8 augmentations work worse than float because of distribution discretization. \n\n- Training process. I train everything with **Adam**, ReduceOnPlateau and few warmup iterations. I've spent so much time trying to understand why people declare SGD works better but no chance - Adam performed better as usual.\n\n- **Controls**. We've simply added train and test (!) controls as additional classes 1108:1139. That boosted score significantly. \n\n**First iteration**. Pretrained **EfficientNets** b0, b1, b2 achieve almost the same score but blending their raw predictions and than solving assignment problem boosted score nicely. Averaging this 3 models scores 0.744/0.750 (3 cells, U2OS-04 reduced) public and 0.990 private LB (13th).\n\nI saw few strategies of how to split test experiments for **pseudo-labeling**. And actually I've tried around 6 of them with one or two iterations of PL. By the end I've came with very careful strategy - soft labels and balanced experiments split between folds:\n```\npl_all_folds = [\"HEPG2-10\", \"HUVEC-19\", \"RPE-10\", \"U2OS-05\"]\npl_folds = [[\"HEPG2-08\", \"HUVEC-17\", \"HUVEC-23\", \"RPE-08\"],\n        [\"HEPG2-09\", \"HUVEC-20\", \"HUVEC-21\", \"RPE-09\"],\n        [\"HEPG2-11\", \"HUVEC-22\", \"HUVEC-24\", \"RPE-11\"]]\n```\nAs you see U2OS-04 was reduced.\n**Soft** pseudo-labelling seemed to work better than hard. Now I understand that the best strategy might be to add them iteratively based on some threshold after solving assignment problem.\n\n**Second iteration**. Same b0, b1, b2 now with pseudo-labels. People don't usual train same architectures in PL stage so I've added b4 and b6 as well. EfficientNet-b4 scored best validation and LB score and unfortunately EfficientNet-b6 didn't converge well in time. For big models we used **gradient accumulations** up to x12. TTA: 2 sites with 8 combinations of flip/rotate each. \n\nWhat didn't work:\n- SGD with CLRs (triangles, cosine etc)\n- Lookahead and Ranger\n- ArcFace (achieved almost the same score as Softmax)\n- My favourite se-resnext50",
    "635268": "Congratulations\nGreat Write-Up\nThanks for Sharing your Approach &amp; Insights...!! @yaroshevskiy",
    "635272": "My favourite backbone is also se-resnext50 : )\nCongrats!\n&gt; Now I understand that the best strategy might be to add them iteratively based on some threshold after solving assignment problem.\n\nLesson learned, thanks!"
  },
  "source": "meta"
}