{
  "id": 220803,
  "title": "8th Place solution ",
  "url": "/competitions/rfcx-species-audio-detection/writeups/shino-8th-place-solution",
  "author_name": "",
  "post_date": "2021-03-02T08:25:49.667Z",
  "votes": 12,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thank you for opening this competition<br>\nAlso the notebooks and discussions have helped me a lot, thank you all!</p>\n<p>My solution was similar to Beluga &amp; Peter in 7th Place.<br>\n<a href=\"https://www.kaggle.com/https\" target=\"_blank\">@https</a>://<a href=\"http://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443\" target=\"_blank\">www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443</a>.</p>\n<ul>\n<li>Multi-class multi-label problem</li>\n<li>Data cleaning train_data with hand labels</li>\n<li>SED model</li>\n</ul>\n<p></p>\n<p><strong>[hand labels]</strong><br>\n While observing the data from t_min and t_max in the given train__tp.csv, I found that there are many kinds of bird calls mixed together.<br>\n So I decided to treat it as a multi-class multi-label problem.<br>\n It was also mentioned in the discussion that the test labels were eventually carefully labeled by humans.<br>\n The TPs in the given t_min, t_max range are all easy to understand, but there are many TPs in the 60s clip that are difficult to understand and not labeled.<br>\n I thought it would be better to label them carefully by myself to make the condition as close to test as possible in case such incomprehensible calls are also labeled in test.<br>\nAnd I was thinking of doing Pseudo labeling after the accuracy of the model improves.</p>\n<p>I trimmed the 5s~ around t_min and t_max in train__tp.csv.<br>\n Hand labels took about a week.<br>\n As a result, a total of 2428 clips and 5s chunks were used as train_data.<br>\n The distribution of the train_data classes looks like this<br>\n(I couldn't upload the image, so I'll post it later)</p>\n<p>class nb<br>\ns3 1257<br>\ns12 520<br>\ns18 512<br>\n:<br>\ns16 100<br>\ns17 100<br>\ns6 97</p>\n<p>I can see that there is a label imbalance, especially for s3, s12, and s18, because their labels co-occur among the other classes of clips.</p>\n<p>In particular, s3 is dominant, so it tends to output high probability, while a few classes output low probability, so i thought this is a bad problem for this evaluation index.<br>\nTherefore, in order to achieve a more balanced distribution, I oversampled the minority classes and undersampled the majority classes.<br>\n However, the LB became worse.<br>\n Looking back, I didn't think of approaching the test distribution, as Chris pointed out.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389</a></p>\n<p>Eventually, i ensembled 15 models with various losses(FocalLoss,FbetaLoss+FocalLoss,ResampleLoss).</p>\n<p><strong>[trainning]</strong><br>\nexample single model <br>\nPANNsDense161 (public_LB 0.95548, private_LB 0.96300)</p>\n<p>I also tried EfficientNet_b0, Dense121, etc., but Dense161 worked well.</p>\n<p>train_data(sr=48000,5s)<br>\n  window_size=2048,hop_size=512,mel_bins=256<br>\n  MultilabelStratifiedKFold 5fold<br>\n  BCEFocalLoss(α=0.25,β=2)<br>\n  GradualWarmupScheduler,CosineAnnealingLR(lr = 0.001,multiplier=10,epo35)</p>\n<p><strong>Augmentation</strong><br>\n  GaussianNoise(p=0.5)<br>\n  GaussianSNR(p=0.5)<br>\n  FrequencyMask(min_frequency_band=0.0, max_frequency_band=0.2, p=0.3)<br>\n  TimeMask(min_band_part=0.0, max_band_part=0.2, p=0.8)<br>\n  PitchShift(min_semitones=-0.5, max_semitones=0.5, p=0.1)<br>\n  Shift(p=0.1)<br>\n  Gain(p=0.2)</p>\n<p><strong>[inference]</strong><br>\n  stride=1<br>\n  framewise_output max<br>\n  No TTA (I used it in the final ensemble model)</p>\n<p>Finally, I've uploaded the train_data_wav (sr=48000) and csv that I used. <br>\n<a href=\"https://www.kaggle.com/shinoda18/rainforest-data\" target=\"_blank\">https://www.kaggle.com/shinoda18/rainforest-data</a></p>",
  "messages": [
    {
      "id": "1210666",
      "postDate": "02/19/2021 15:37:50",
      "content": "<p>Thank you for opening this competition<br>\nAlso the notebooks and discussions have helped me a lot, thank you all!</p>\n<p>My solution was similar to Beluga &amp; Peter in 7th Place.<br>\n<a href=\"https://www.kaggle.com/https\" target=\"_blank\">@https</a>://<a href=\"http://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443\" target=\"_blank\">www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443</a>.</p>\n<ul>\n<li>Multi-class multi-label problem</li>\n<li>Data cleaning train_data with hand labels</li>\n<li>SED model</li>\n</ul>\n<p></p>\n<p><strong>[hand labels]</strong><br>\n While observing the data from t_min and t_max in the given train__tp.csv, I found that there are many kinds of bird calls mixed together.<br>\n So I decided to treat it as a multi-class multi-label problem.<br>\n It was also mentioned in the discussion that the test labels were eventually carefully labeled by humans.<br>\n The TPs in the given t_min, t_max range are all easy to understand, but there are many TPs in the 60s clip that are difficult to understand and not labeled.<br>\n I thought it would be better to label them carefully by myself to make the condition as close to test as possible in case such incomprehensible calls are also labeled in test.<br>\nAnd I was thinking of doing Pseudo labeling after the accuracy of the model improves.</p>\n<p>I trimmed the 5s~ around t_min and t_max in train__tp.csv.<br>\n Hand labels took about a week.<br>\n As a result, a total of 2428 clips and 5s chunks were used as train_data.<br>\n The distribution of the train_data classes looks like this<br>\n(I couldn't upload the image, so I'll post it later)</p>\n<p>class nb<br>\ns3 1257<br>\ns12 520<br>\ns18 512<br>\n:<br>\ns16 100<br>\ns17 100<br>\ns6 97</p>\n<p>I can see that there is a label imbalance, especially for s3, s12, and s18, because their labels co-occur among the other classes of clips.</p>\n<p>In particular, s3 is dominant, so it tends to output high probability, while a few classes output low probability, so i thought this is a bad problem for this evaluation index.<br>\nTherefore, in order to achieve a more balanced distribution, I oversampled the minority classes and undersampled the majority classes.<br>\n However, the LB became worse.<br>\n Looking back, I didn't think of approaching the test distribution, as Chris pointed out.<br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389</a></p>\n<p>Eventually, i ensembled 15 models with various losses(FocalLoss,FbetaLoss+FocalLoss,ResampleLoss).</p>\n<p><strong>[trainning]</strong><br>\nexample single model <br>\nPANNsDense161 (public_LB 0.95548, private_LB 0.96300)</p>\n<p>I also tried EfficientNet_b0, Dense121, etc., but Dense161 worked well.</p>\n<p>train_data(sr=48000,5s)<br>\n  window_size=2048,hop_size=512,mel_bins=256<br>\n  MultilabelStratifiedKFold 5fold<br>\n  BCEFocalLoss(α=0.25,β=2)<br>\n  GradualWarmupScheduler,CosineAnnealingLR(lr = 0.001,multiplier=10,epo35)</p>\n<p><strong>Augmentation</strong><br>\n  GaussianNoise(p=0.5)<br>\n  GaussianSNR(p=0.5)<br>\n  FrequencyMask(min_frequency_band=0.0, max_frequency_band=0.2, p=0.3)<br>\n  TimeMask(min_band_part=0.0, max_band_part=0.2, p=0.8)<br>\n  PitchShift(min_semitones=-0.5, max_semitones=0.5, p=0.1)<br>\n  Shift(p=0.1)<br>\n  Gain(p=0.2)</p>\n<p><strong>[inference]</strong><br>\n  stride=1<br>\n  framewise_output max<br>\n  No TTA (I used it in the final ensemble model)</p>\n<p>Finally, I've uploaded the train_data_wav (sr=48000) and csv that I used. <br>\n<a href=\"https://www.kaggle.com/shinoda18/rainforest-data\" target=\"_blank\">https://www.kaggle.com/shinoda18/rainforest-data</a></p>",
      "rawMarkdown": "Thank you for opening this competition\nAlso the notebooks and discussions have helped me a lot, thank you all!\n\nMy solution was similar to Beluga & Peter in 7th Place.\n@https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443.\n- Multi-class multi-label problem\n- Data cleaning train_data with hand labels\n- SED model\n\n~~In my case, Pseudo labeling did not work well, so I did not use it.~~\n\n**[hand labels]**\n While observing the data from t_min and t_max in the given train__tp.csv, I found that there are many kinds of bird calls mixed together.\n So I decided to treat it as a multi-class multi-label problem.\n It was also mentioned in the discussion that the test labels were eventually carefully labeled by humans.\n The TPs in the given t_min, t_max range are all easy to understand, but there are many TPs in the 60s clip that are difficult to understand and not labeled.\n I thought it would be better to label them carefully by myself to make the condition as close to test as possible in case such incomprehensible calls are also labeled in test.\nAnd I was thinking of doing Pseudo labeling after the accuracy of the model improves.\n\n I trimmed the 5s~ around t_min and t_max in train__tp.csv.\n Hand labels took about a week.\n As a result, a total of 2428 clips and 5s chunks were used as train_data.\n The distribution of the train_data classes looks like this\n(I couldn't upload the image, so I'll post it later)\n\nclass nb\ns3 1257\ns12 520\ns18 512\n:\ns16 100\ns17 100\ns6 97\n\n I can see that there is a label imbalance, especially for s3, s12, and s18, because their labels co-occur among the other classes of clips.\n\n In particular, s3 is dominant, so it tends to output high probability, while a few classes output low probability, so i thought this is a bad problem for this evaluation index.\nTherefore, in order to achieve a more balanced distribution, I oversampled the minority classes and undersampled the majority classes.\n However, the LB became worse.\n Looking back, I didn't think of approaching the test distribution, as Chris pointed out.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\n\nEventually, i ensembled 15 models with various losses(FocalLoss,FbetaLoss+FocalLoss,ResampleLoss).\n\n**[trainning]**\nexample single model \nPANNsDense161 (public_LB 0.95548, private_LB 0.96300)\n\nI also tried EfficientNet_b0, Dense121, etc., but Dense161 worked well.\n\n  train_data(sr=48000,5s)\n  window_size=2048,hop_size=512,mel_bins=256\n  MultilabelStratifiedKFold 5fold\n  BCEFocalLoss(α=0.25,β=2)\n  GradualWarmupScheduler,CosineAnnealingLR(lr = 0.001,multiplier=10,epo35)\n\n**Augmentation**\n  GaussianNoise(p=0.5)\n  GaussianSNR(p=0.5)\n  FrequencyMask(min_frequency_band=0.0, max_frequency_band=0.2, p=0.3)\n  TimeMask(min_band_part=0.0, max_band_part=0.2, p=0.8)\n  PitchShift(min_semitones=-0.5, max_semitones=0.5, p=0.1)\n  Shift(p=0.1)\n  Gain(p=0.2)\n\n**[inference]**\n  stride=1\n  framewise_output max\n  No TTA (I used it in the final ensemble model)\n\nFinally, I've uploaded the train_data_wav (sr=48000) and csv that I used. \nhttps://www.kaggle.com/shinoda18/rainforest-data",
      "votes": null
    },
    {
      "id": "1211942",
      "postDate": "02/20/2021 17:23:22",
      "content": "<p>a strong effort! a solo Gold… I think u and Chris were the only solos on the Top 10 list. Brilliant!</p>",
      "rawMarkdown": "a strong effort! a solo Gold... I think u and Chris were the only solos on the Top 10 list. Brilliant!",
      "votes": null
    },
    {
      "id": "1212172",
      "postDate": "02/21/2021 00:44:53",
      "content": "<p>Thank you!<br>\nI saw your discussion.<br>\nGreat insights and hard work!<br>\nI'll see you at another convention somewhere!</p>",
      "rawMarkdown": "Thank you!\nI saw your discussion.\nGreat insights and hard work!\nI'll see you at another convention somewhere!",
      "votes": null
    },
    {
      "id": "1212175",
      "postDate": "02/21/2021 00:51:46",
      "content": "<p>\"In my case, Pseudo labeling did not work well, so I did not use it.\"</p>\n<p>since you have both hand label and pseudo label, I suggest you can do this:</p>\n<ol>\n<li>measure label purity of  pseudo label, based on your hand label</li>\n<li>train with a subset of the purer pseudo label and see if results improve.</li>\n</ol>\n<p>if you can find out the reasons, why it doesn't work and then make pseudo label works, I think it is useful for future competition.</p>\n<p>the issue here is I think pseudo label is not as good as your hand label. But we want to ask the question:</p>\n<ul>\n<li>how much difference is there between the 2 set?</li>\n<li>can I modify so that the 2 distribution are more similar?</li>\n</ul>",
      "rawMarkdown": "\"In my case, Pseudo labeling did not work well, so I did not use it.\"\n\nsince you have both hand label and pseudo label, I suggest you can do this:\n1. measure label purity of  pseudo label, based on your hand label\n2. train with a subset of the purer pseudo label and see if results improve.\n\nif you can find out the reasons, why it doesn't work and then make pseudo label works, I think it is useful for future competition.\n\nthe issue here is I think pseudo label is not as good as your hand label. But we want to ask the question:\n- how much difference is there between the 2 set?\n- can I modify so that the 2 distribution are more similar?",
      "votes": null
    },
    {
      "id": "1212201",
      "postDate": "02/21/2021 02:03:35",
      "content": "<p>I am always amazed at your discussions and solutions!<br>\nYou are one of the kaggler I personally respect!</p>\n<p>Sorry, there is a mistake in my text!<br>\nI tried the hand label=&gt; Pseudo label<br>\nI did not try Pseudo label from the beginning.</p>\n<blockquote>\n  <p>how much difference is there between the 2 set?</p>\n</blockquote>\n<p>So I don't have a set of pseudo labels with me.<br>\nI was not able to compare them.<br>\nI can't say that the pseudo labels did not work. I will fix it.</p>\n<blockquote>\n  <p>can I modify so that the 2 distribution are more similar?</p>\n</blockquote>\n<p>I've looked at all the train_data.<br>\ns3,s7,s11,s12,s18 seem to be particularly frequent.<br>\nIn Chris's notebook.<br>\n<a href=\"https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\" target=\"_blank\">https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970</a><br>\n[113,204,44,923,53,41,3,213,44,23,26,149,255,14,123,222,46,6,474,4,17,18,23,72].<br>\nThe distribution is similar to</p>\n<p>If we do a Pseudo label using all the train_data, we may get a distribution exactly like test_data.</p>\n<p>The hand label=&gt; Pseudo label I tried mainly upsampled the tail class to eliminate the class imbalance, so it did not work well because it was far from the test_data distribution.</p>",
      "rawMarkdown": "I am always amazed at your discussions and solutions!\nYou are one of the kaggler I personally respect!\n\nSorry, there is a mistake in my text!\nI tried the hand label=> Pseudo label\nI did not try Pseudo label from the beginning.\n\n> how much difference is there between the 2 set?\n\nSo I don't have a set of pseudo labels with me.\nI was not able to compare them.\nI can't say that the pseudo labels did not work. I will fix it.\n\n> can I modify so that the 2 distribution are more similar?\n\nI've looked at all the train_data.\ns3,s7,s11,s12,s18 seem to be particularly frequent.\nIn Chris's notebook.\nhttps://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\n[113,204,44,923,53,41,3,213,44,23,26,149,255,14,123,222,46,6,474,4,17,18,23,72].\nThe distribution is similar to\n\nIf we do a Pseudo label using all the train_data, we may get a distribution exactly like test_data.\n\n\nThe hand label=> Pseudo label I tried mainly upsampled the tail class to eliminate the class imbalance, so it did not work well because it was far from the test_data distribution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1211942,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "02/20/2021 17:23:22",
      "content": "<p>a strong effort! a solo Gold… I think u and Chris were the only solos on the Top 10 list. Brilliant!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1212172,
          "author_name": "shinoda18",
          "author_url": "",
          "post_date": "02/21/2021 00:44:53",
          "content": "<p>Thank you!<br>\nI saw your discussion.<br>\nGreat insights and hard work!<br>\nI'll see you at another convention somewhere!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212175,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/21/2021 00:51:46",
      "content": "<p>\"In my case, Pseudo labeling did not work well, so I did not use it.\"</p>\n<p>since you have both hand label and pseudo label, I suggest you can do this:</p>\n<ol>\n<li>measure label purity of  pseudo label, based on your hand label</li>\n<li>train with a subset of the purer pseudo label and see if results improve.</li>\n</ol>\n<p>if you can find out the reasons, why it doesn't work and then make pseudo label works, I think it is useful for future competition.</p>\n<p>the issue here is I think pseudo label is not as good as your hand label. But we want to ask the question:</p>\n<ul>\n<li>how much difference is there between the 2 set?</li>\n<li>can I modify so that the 2 distribution are more similar?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1212201,
          "author_name": "shinoda18",
          "author_url": "",
          "post_date": "02/21/2021 02:03:35",
          "content": "<p>I am always amazed at your discussions and solutions!<br>\nYou are one of the kaggler I personally respect!</p>\n<p>Sorry, there is a mistake in my text!<br>\nI tried the hand label=&gt; Pseudo label<br>\nI did not try Pseudo label from the beginning.</p>\n<blockquote>\n  <p>how much difference is there between the 2 set?</p>\n</blockquote>\n<p>So I don't have a set of pseudo labels with me.<br>\nI was not able to compare them.<br>\nI can't say that the pseudo labels did not work. I will fix it.</p>\n<blockquote>\n  <p>can I modify so that the 2 distribution are more similar?</p>\n</blockquote>\n<p>I've looked at all the train_data.<br>\ns3,s7,s11,s12,s18 seem to be particularly frequent.<br>\nIn Chris's notebook.<br>\n<a href=\"https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\" target=\"_blank\">https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970</a><br>\n[113,204,44,923,53,41,3,213,44,23,26,149,255,14,123,222,46,6,474,4,17,18,23,72].<br>\nThe distribution is similar to</p>\n<p>If we do a Pseudo label using all the train_data, we may get a distribution exactly like test_data.</p>\n<p>The hand label=&gt; Pseudo label I tried mainly upsampled the tail class to eliminate the class imbalance, so it did not work well because it was far from the test_data distribution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1210666": "Thank you for opening this competition\nAlso the notebooks and discussions have helped me a lot, thank you all!\n\nMy solution was similar to Beluga & Peter in 7th Place.\n@https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443.\n- Multi-class multi-label problem\n- Data cleaning train_data with hand labels\n- SED model\n\n~~In my case, Pseudo labeling did not work well, so I did not use it.~~\n\n**[hand labels]**\n While observing the data from t_min and t_max in the given train__tp.csv, I found that there are many kinds of bird calls mixed together.\n So I decided to treat it as a multi-class multi-label problem.\n It was also mentioned in the discussion that the test labels were eventually carefully labeled by humans.\n The TPs in the given t_min, t_max range are all easy to understand, but there are many TPs in the 60s clip that are difficult to understand and not labeled.\n I thought it would be better to label them carefully by myself to make the condition as close to test as possible in case such incomprehensible calls are also labeled in test.\nAnd I was thinking of doing Pseudo labeling after the accuracy of the model improves.\n\n I trimmed the 5s~ around t_min and t_max in train__tp.csv.\n Hand labels took about a week.\n As a result, a total of 2428 clips and 5s chunks were used as train_data.\n The distribution of the train_data classes looks like this\n(I couldn't upload the image, so I'll post it later)\n\nclass nb\ns3 1257\ns12 520\ns18 512\n:\ns16 100\ns17 100\ns6 97\n\n I can see that there is a label imbalance, especially for s3, s12, and s18, because their labels co-occur among the other classes of clips.\n\n In particular, s3 is dominant, so it tends to output high probability, while a few classes output low probability, so i thought this is a bad problem for this evaluation index.\nTherefore, in order to achieve a more balanced distribution, I oversampled the minority classes and undersampled the majority classes.\n However, the LB became worse.\n Looking back, I didn't think of approaching the test distribution, as Chris pointed out.\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\n\nEventually, i ensembled 15 models with various losses(FocalLoss,FbetaLoss+FocalLoss,ResampleLoss).\n\n**[trainning]**\nexample single model \nPANNsDense161 (public_LB 0.95548, private_LB 0.96300)\n\nI also tried EfficientNet_b0, Dense121, etc., but Dense161 worked well.\n\n  train_data(sr=48000,5s)\n  window_size=2048,hop_size=512,mel_bins=256\n  MultilabelStratifiedKFold 5fold\n  BCEFocalLoss(α=0.25,β=2)\n  GradualWarmupScheduler,CosineAnnealingLR(lr = 0.001,multiplier=10,epo35)\n\n**Augmentation**\n  GaussianNoise(p=0.5)\n  GaussianSNR(p=0.5)\n  FrequencyMask(min_frequency_band=0.0, max_frequency_band=0.2, p=0.3)\n  TimeMask(min_band_part=0.0, max_band_part=0.2, p=0.8)\n  PitchShift(min_semitones=-0.5, max_semitones=0.5, p=0.1)\n  Shift(p=0.1)\n  Gain(p=0.2)\n\n**[inference]**\n  stride=1\n  framewise_output max\n  No TTA (I used it in the final ensemble model)\n\nFinally, I've uploaded the train_data_wav (sr=48000) and csv that I used. \nhttps://www.kaggle.com/shinoda18/rainforest-data",
    "1211942": "a strong effort! a solo Gold... I think u and Chris were the only solos on the Top 10 list. Brilliant!",
    "1212172": "Thank you!\nI saw your discussion.\nGreat insights and hard work!\nI'll see you at another convention somewhere!",
    "1212175": "\"In my case, Pseudo labeling did not work well, so I did not use it.\"\n\nsince you have both hand label and pseudo label, I suggest you can do this:\n1. measure label purity of  pseudo label, based on your hand label\n2. train with a subset of the purer pseudo label and see if results improve.\n\nif you can find out the reasons, why it doesn't work and then make pseudo label works, I think it is useful for future competition.\n\nthe issue here is I think pseudo label is not as good as your hand label. But we want to ask the question:\n- how much difference is there between the 2 set?\n- can I modify so that the 2 distribution are more similar?",
    "1212201": "I am always amazed at your discussions and solutions!\nYou are one of the kaggler I personally respect!\n\nSorry, there is a mistake in my text!\nI tried the hand label=> Pseudo label\nI did not try Pseudo label from the beginning.\n\n> how much difference is there between the 2 set?\n\nSo I don't have a set of pseudo labels with me.\nI was not able to compare them.\nI can't say that the pseudo labels did not work. I will fix it.\n\n> can I modify so that the 2 distribution are more similar?\n\nI've looked at all the train_data.\ns3,s7,s11,s12,s18 seem to be particularly frequent.\nIn Chris's notebook.\nhttps://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\n[113,204,44,923,53,41,3,213,44,23,26,149,255,14,123,222,46,6,474,4,17,18,23,72].\nThe distribution is similar to\n\nIf we do a Pseudo label using all the train_data, we may get a distribution exactly like test_data.\n\n\nThe hand label=> Pseudo label I tried mainly upsampled the tail class to eliminate the class imbalance, so it did not work well because it was far from the test_data distribution."
  },
  "source": "meta"
}