{
  "id": 220339,
  "title": "Top Solutions and Approaches: Rainforest Audio Detection",
  "url": "/competitions/rfcx-species-audio-detection/discussion/220339",
  "author_name": "",
  "post_date": "2021-02-18T03:32:29.998485500Z",
  "votes": 12,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>Congratulations to all the winners!!</p>\n<p>I am consolidating the top 50 results, as they become available. The discussion as well as the code in GitHub/Kaggle. Hoping this becomes  easy to refer to different top solutions to learn from!<br>\nIn case, I missed your discussion/code, please comment and I will update! Thanks for sharing!</p>\n<p>Note: Names below are for the one who posted, many are part of a larger team</p>\n<p>Rank 1 <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563</a></p>\n<p>Rank 2 <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220760\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220760</a></p>\n<p>Rank 3 <a href=\"https://www.kaggle.com/dicksonchin93\" target=\"_blank\">@dicksonchin93</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522</a><br>\nRank 3  <a href=\"https://www.kaggle.com/meaninglesslives\" target=\"_blank\">@meaninglesslives</a><br>\nCode: <a href=\"https://www.kaggle.com/meaninglesslives/rfcx-minimal\" target=\"_blank\">https://www.kaggle.com/meaninglesslives/rfcx-minimal</a></p>\n<p>Rank 4 <a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342</a></p>\n<p>Rank 5 <a href=\"https://www.kaggle.com/takamichitoda\" target=\"_blank\">@takamichitoda</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220432\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220432</a></p>\n<p>Rank 6  <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220446\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220446</a><br>\nRank 6  <a href=\"https://www.kaggle.com/amezet\" target=\"_blank\">@amezet</a><br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220981\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220981</a></p>\n<p>Rank 7 <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443</a></p>\n<p>Rank 8 <a href=\"https://www.kaggle.com/shinoda18\" target=\"_blank\">@shinoda18</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220803\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220803</a></p>\n<p>Rank 9 <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389</a><br>\nCode: <a href=\"https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\" target=\"_blank\">https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970</a></p>\n<p>===========</p>\n<p>Rank 10 <a href=\"https://www.kaggle.com/bigironsphere\" target=\"_blank\">@bigironsphere</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220319\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220319</a></p>\n<p>Rank 11 <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304</a></p>\n<p>Rank 13 <a href=\"https://www.kaggle.com/reppic\" target=\"_blank\">@reppic</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220308\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220308</a></p>\n<p>Rank 14 <a href=\"https://www.kaggle.com/prateekagnihotri\" target=\"_blank\">@prateekagnihotri</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220725\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220725</a></p>\n<p>Rank 16 <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220450\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220450</a></p>\n<p>Rank 18 <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309</a></p>\n<p>===============</p>\n<p>Rank 21 <a href=\"https://www.kaggle.com/xmpgek\" target=\"_blank\">@xmpgek</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220436\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220436</a></p>\n<p>Rank 23 <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220972\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220972</a></p>\n<p>Rank 25 <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220314\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220314</a></p>\n<p>Rank 27 <a href=\"https://www.kaggle.com/mnpinto\" target=\"_blank\">@mnpinto</a><br>\nDiscussion:<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220306\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220306</a></p>\n<p>Rank 32 <a href=\"https://www.kaggle.com/harangdev\" target=\"_blank\">@harangdev</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220322\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220322</a></p>\n<p>Rank 33 <a href=\"https://www.kaggle.com/yururoi\" target=\"_blank\">@yururoi</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220753\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220753</a></p>\n<p>Rank 38 : <a href=\"https://www.kaggle.com/vzaguskin\" target=\"_blank\">@vzaguskin</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220386\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220386</a></p>\n<p>Rank 43 : <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220335\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220335</a></p>\n<p>Congratulations to everyone! and thanks for sharing!!</p>",
  "messages": [
    {
      "id": "1207797",
      "postDate": "02/18/2021 03:32:30",
      "content": "<p>Hi All,</p>\n<p>Congratulations to all the winners!!</p>\n<p>I am consolidating the top 50 results, as they become available. The discussion as well as the code in GitHub/Kaggle. Hoping this becomes  easy to refer to different top solutions to learn from!<br>\nIn case, I missed your discussion/code, please comment and I will update! Thanks for sharing!</p>\n<p>Note: Names below are for the one who posted, many are part of a larger team</p>\n<p>Rank 1 <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563</a></p>\n<p>Rank 2 <a href=\"https://www.kaggle.com/selimsef\" target=\"_blank\">@selimsef</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220760\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220760</a></p>\n<p>Rank 3 <a href=\"https://www.kaggle.com/dicksonchin93\" target=\"_blank\">@dicksonchin93</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522</a><br>\nRank 3  <a href=\"https://www.kaggle.com/meaninglesslives\" target=\"_blank\">@meaninglesslives</a><br>\nCode: <a href=\"https://www.kaggle.com/meaninglesslives/rfcx-minimal\" target=\"_blank\">https://www.kaggle.com/meaninglesslives/rfcx-minimal</a></p>\n<p>Rank 4 <a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342</a></p>\n<p>Rank 5 <a href=\"https://www.kaggle.com/takamichitoda\" target=\"_blank\">@takamichitoda</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220432\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220432</a></p>\n<p>Rank 6  <a href=\"https://www.kaggle.com/antorsae\" target=\"_blank\">@antorsae</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220446\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220446</a><br>\nRank 6  <a href=\"https://www.kaggle.com/amezet\" target=\"_blank\">@amezet</a><br>\n<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220981\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220981</a></p>\n<p>Rank 7 <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443</a></p>\n<p>Rank 8 <a href=\"https://www.kaggle.com/shinoda18\" target=\"_blank\">@shinoda18</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220803\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220803</a></p>\n<p>Rank 9 <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389</a><br>\nCode: <a href=\"https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\" target=\"_blank\">https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970</a></p>\n<p>===========</p>\n<p>Rank 10 <a href=\"https://www.kaggle.com/bigironsphere\" target=\"_blank\">@bigironsphere</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220319\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220319</a></p>\n<p>Rank 11 <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304</a></p>\n<p>Rank 13 <a href=\"https://www.kaggle.com/reppic\" target=\"_blank\">@reppic</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220308\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220308</a></p>\n<p>Rank 14 <a href=\"https://www.kaggle.com/prateekagnihotri\" target=\"_blank\">@prateekagnihotri</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220725\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220725</a></p>\n<p>Rank 16 <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220450\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220450</a></p>\n<p>Rank 18 <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309</a></p>\n<p>===============</p>\n<p>Rank 21 <a href=\"https://www.kaggle.com/xmpgek\" target=\"_blank\">@xmpgek</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220436\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220436</a></p>\n<p>Rank 23 <a href=\"https://www.kaggle.com/dathudeptrai\" target=\"_blank\">@dathudeptrai</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220972\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220972</a></p>\n<p>Rank 25 <a href=\"https://www.kaggle.com/rytisva88\" target=\"_blank\">@rytisva88</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220314\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220314</a></p>\n<p>Rank 27 <a href=\"https://www.kaggle.com/mnpinto\" target=\"_blank\">@mnpinto</a><br>\nDiscussion:<a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220306\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220306</a></p>\n<p>Rank 32 <a href=\"https://www.kaggle.com/harangdev\" target=\"_blank\">@harangdev</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220322\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220322</a></p>\n<p>Rank 33 <a href=\"https://www.kaggle.com/yururoi\" target=\"_blank\">@yururoi</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220753\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220753</a></p>\n<p>Rank 38 : <a href=\"https://www.kaggle.com/vzaguskin\" target=\"_blank\">@vzaguskin</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220386\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220386</a></p>\n<p>Rank 43 : <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a><br>\nDiscussion: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220335\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220335</a></p>\n<p>Congratulations to everyone! and thanks for sharing!!</p>",
      "rawMarkdown": "Hi All,\n\nCongratulations to all the winners!!\n\nI am consolidating the top 50 results, as they become available. The discussion as well as the code in GitHub/Kaggle. Hoping this becomes  easy to refer to different top solutions to learn from!\nIn case, I missed your discussion/code, please comment and I will update! Thanks for sharing!\n\nNote: Names below are for the one who posted, many are part of a larger team\n\nRank 1 @philippsinger\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563\n\n\nRank 2 @selimsef\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220760\n\n\nRank 3 @dicksonchin93\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522\nRank 3  @meaninglesslives\nCode: https://www.kaggle.com/meaninglesslives/rfcx-minimal\n\nRank 4 @kupchanski\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342\n\nRank 5 @takamichitoda\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220432\n\nRank 6  @antorsae\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220446\nRank 6  @amezet\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220981\n\nRank 7 @pestipeti\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443\n\nRank 8 @shinoda18\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220803\n\n\nRank 9 @cdeotte\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\nCode: https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\n\n===========\n\nRank 10 @bigironsphere\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220319\n\nRank 11 @cpmpml\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\n\nRank 13 @reppic\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220308\n\nRank 14 @prateekagnihotri\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220725\n\nRank 16 @hidehisaarai1213\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220450\n\nRank 18 @hengck23\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309\n\n===============\n\nRank 21 @xmpgek\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220436\n\nRank 23 @dathudeptrai\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220972\n\nRank 25 @rytisva88\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220314\n\nRank 27 @mnpinto\nDiscussion:https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220306\n\nRank 32 @harangdev\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220322\n\n\nRank 33 @yururoi\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220753\n\nRank 38 : @vzaguskin\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220386\n\nRank 43 : @shinmurashinmura\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220335\n\n\nCongratulations to everyone! and thanks for sharing!!",
      "votes": null
    },
    {
      "id": "1207915",
      "postDate": "02/18/2021 04:57:50",
      "content": "<p>when will top3 GitHub links be revealed?</p>",
      "rawMarkdown": "when will top3 GitHub links be revealed?",
      "votes": null
    },
    {
      "id": "1207955",
      "postDate": "02/18/2021 05:08:38",
      "content": "<p>Typically in 2-3 days, most clean up their code before posting<br>\nHowever, some may choose to just put the approach rather than the code</p>",
      "rawMarkdown": "Typically in 2-3 days, most clean up their code before posting\nHowever, some may choose to just put the approach rather than the code",
      "votes": null
    },
    {
      "id": "1207983",
      "postDate": "02/18/2021 05:19:47",
      "content": "<p>I think you mentioned me twiсe for 4 and 14 places)</p>",
      "rawMarkdown": "I think you mentioned me twiсe for 4 and 14 places)",
      "votes": null
    },
    {
      "id": "1207998",
      "postDate": "02/18/2021 05:35:21",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a>, corrected!!</p>",
      "rawMarkdown": "Thanks @kupchanski, corrected!!",
      "votes": null
    },
    {
      "id": "1209500",
      "postDate": "02/18/2021 23:47:18",
      "content": "<p><img src=\"https://i.ibb.co/1z07JZv/Selection-081.png\" alt=\"\"></p>\n<p>I also want to mention that serious kaggling is hard work:</p>\n<p>top-1: trained 120 models<br>\ntop-3: numerous hand-labeling</p>\n<p>one may want to remember this formula in your future competition development cycle:<br>\nbetter modeling --&gt; better data --&gt; better biasing</p>",
      "rawMarkdown": "![](https://i.ibb.co/1z07JZv/Selection-081.png)\n\nI also want to mention that serious kaggling is hard work:\n\ntop-1: trained 120 models\ntop-3: numerous hand-labeling\n\n\none may want to remember this formula in your future competition development cycle:\nbetter modeling --> better data --> better biasing",
      "votes": null
    },
    {
      "id": "1209707",
      "postDate": "02/19/2021 02:41:00",
      "content": "<p>Couldn't agree more <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !!</p>",
      "rawMarkdown": "Couldn't agree more @hengck23 !!",
      "votes": null
    },
    {
      "id": "1209737",
      "postDate": "02/19/2021 02:57:29",
      "content": "<p>Nice plot and suggestions Heng.</p>",
      "rawMarkdown": "Nice plot and suggestions Heng.",
      "votes": null
    },
    {
      "id": "1211574",
      "postDate": "02/20/2021 10:36:17",
      "content": "<p>Great analysis.</p>\n<p>I never do better biasing because it is a recipe for overfitting in real life ML.  But in Kaggle, when we fully know the test data it is very effective.  Indeed it becomes a matter of overfitting to private test data.  And when public/private split is random, it then becomes a matter of overftting to public test data.</p>\n<p>The downside is that when public/private split is not random then the above results in a major fall in shakedown.  Given I don't overfit to public LB I almost always move up in shakeup.</p>\n<p>Now the question for me is: should i stick to sound ML practice, or should I enter the overfitting game?  I don't like to have to make a choice here.</p>\n<p>From the host perspective, I wonder what is the value of solutions that overfit to test data.  </p>\n<p>I would recommend that public/private test split is never random so that generalization power of models is tested.</p>\n<p>Edit.  Better biasing based on training data is sound in real life ML.  Doing it whenever applicable is great.</p>",
      "rawMarkdown": "Great analysis.\n\nI never do better biasing because it is a recipe for overfitting in real life ML.  But in Kaggle, when we fully know the test data it is very effective.  Indeed it becomes a matter of overfitting to private test data.  And when public/private split is random, it then becomes a matter of overftting to public test data.\n\nThe downside is that when public/private split is not random then the above results in a major fall in shakedown.  Given I don't overfit to public LB I almost always move up in shakeup.\n\nNow the question for me is: should i stick to sound ML practice, or should I enter the overfitting game?  I don't like to have to make a choice here.\n\nFrom the host perspective, I wonder what is the value of solutions that overfit to test data.  \n\nI would recommend that public/private test split is never random so that generalization power of models is tested.\n\nEdit.  Better biasing based on training data is sound in real life ML.  Doing it whenever applicable is great.",
      "votes": null
    },
    {
      "id": "1211598",
      "postDate": "02/20/2021 10:51:13",
      "content": "<p>You should update ranks now that LB is final.</p>",
      "rawMarkdown": "You should update ranks now that LB is final.",
      "votes": null
    },
    {
      "id": "1211638",
      "postDate": "02/20/2021 12:00:23",
      "content": "<p>Thanks for the reminder!! Updated the ranks and some new discussions  as well</p>",
      "rawMarkdown": "Thanks for the reminder!! Updated the ranks and some new discussions  as well",
      "votes": null
    },
    {
      "id": "1211795",
      "postDate": "02/20/2021 15:10:07",
      "content": "<p>It looks like that most of the top sharings in this competition are idea-based, no github links… A bit disappointed…</p>",
      "rawMarkdown": "It looks like that most of the top sharings in this competition are idea-based, no github links... A bit disappointed...",
      "votes": null
    },
    {
      "id": "1211938",
      "postDate": "02/20/2021 17:20:46",
      "content": "<p>Yes, I agree that all top solutions are overfitted to the test data. Having 3 and 18 species in majority of samples is a leak in some sense, especially for LWRAP metric. <br>\nThough with pseudolabeling on early stages performance is really amazing judging from the logloss on OOF for TP and FP (later stages have worse logloss on FP samples). So with better validation  they can choose models that suit them (for example by using test data for validation and employing other metrics)</p>",
      "rawMarkdown": "Yes, I agree that all top solutions are overfitted to the test data. Having 3 and 18 species in majority of samples is a leak in some sense, especially for LWRAP metric. \nThough with pseudolabeling on early stages performance is really amazing judging from the logloss on OOF for TP and FP (later stages have worse logloss on FP samples). So with better validation  they can choose models that suit them (for example by using test data for validation and employing other metrics)",
      "votes": null
    },
    {
      "id": "1211984",
      "postDate": "02/20/2021 18:08:17",
      "content": "<p>I wonder why they did not use a metric per species, like roc-auc, or f1 score.  </p>\n<blockquote>\n  <p>(later stages have worse logloss on FP samples). </p>\n</blockquote>\n<p>Interesting, it shows that the bias in FP sampling is detrimental at some point.</p>",
      "rawMarkdown": "I wonder why they did not use a metric per species, like roc-auc, or f1 score.  \n\n> (later stages have worse logloss on FP samples). \n\nInteresting, it shows that the bias in FP sampling is detrimental at some point.",
      "votes": null
    },
    {
      "id": "1212779",
      "postDate": "02/21/2021 15:29:34",
      "content": "<p><a href=\"https://www.kaggle.com/meaninglesslives\" target=\"_blank\">@meaninglesslives</a> has already shared a gold class kernel - <a href=\"https://www.kaggle.com/meaninglesslives/rfcx-minimal\" target=\"_blank\">https://www.kaggle.com/meaninglesslives/rfcx-minimal</a>. Please go thru' and upvote that one which is already there. We can wait till other Github links are published..</p>",
      "rawMarkdown": "meaninglesslives has already shared a gold class kernel - https://www.kaggle.com/meaninglesslives/rfcx-minimal. Please go thru' and upvote that one which is already there. We can wait till other Github links are published..",
      "votes": null
    },
    {
      "id": "1213195",
      "postDate": "02/22/2021 00:16:01",
      "content": "<p>from a competition point of view, it is ok to overfit test data. you need very good skills to overfit test data. The private test data is not completely visible, it is a test of skill to see you extract information to guess how the private test data looks like.</p>\n<p>if you know how to overfit test data, you would also know how to generalize, interpret results, etc</p>\n<p>from industry or application point of view, we never overfit test data. instead, I would rather spend more time collect data.</p>\n<hr>\n<p>\"I would recommend that public/private test split is never random so that generalization power of models is tested.\"</p>\n<p>if kaggle want industrial-grade solution, the testing should be \"pure black box\". then I afraid, there will be many shakeup.</p>\n<p>i often regard kaggle solution as only proof-of-concept if we are talking about a commercial application development pipeline. It is just the minimum viable model stage (MVM). some team will take over from the kaggle solution to map up a development milestone and scale up the model for real industrial applications.</p>",
      "rawMarkdown": "from a competition point of view, it is ok to overfit test data. you need very good skills to overfit test data. The private test data is not completely visible, it is a test of skill to see you extract information to guess how the private test data looks like.\n\nif you know how to overfit test data, you would also know how to generalize, interpret results, etc\n\nfrom industry or application point of view, we never overfit test data. instead, I would rather spend more time collect data.\n\n---\n\n\"I would recommend that public/private test split is never random so that generalization power of models is tested.\"\n\nif kaggle want industrial-grade solution, the testing should be \"pure black box\". then I afraid, there will be many shakeup.\n\ni often regard kaggle solution as only proof-of-concept if we are talking about a commercial application development pipeline. It is just the minimum viable model stage (MVM). some team will take over from the kaggle solution to map up a development milestone and scale up the model for real industrial applications.",
      "votes": null
    },
    {
      "id": "1213197",
      "postDate": "02/22/2021 00:17:56",
      "content": "<blockquote>\n  <p>all top solutions are overfitted to the test data</p>\n</blockquote>\n<p>because of the addition of extra hand-labels in both training and validation, our lwlrap CV score correlates to the increase in lb, I wouldn't say it's overfitting to test set in our case.</p>",
      "rawMarkdown": "> all top solutions are overfitted to the test data\n\nbecause of the addition of extra hand-labels in both training and validation, our lwlrap CV score correlates to the increase in lb, I wouldn't say it's overfitting to test set in our case.",
      "votes": null
    },
    {
      "id": "1213261",
      "postDate": "02/22/2021 02:28:42",
      "content": "<blockquote>\n  <p>if kaggle want industrial-grade solution, the testing should be \"pure black box\". then I afraid, there will be many shakeup.</p>\n</blockquote>\n<p>Yes, there will be shakeup and robust models will survive.  We see it happening each time public test distribution is not close to private test distribution.</p>\n<p>Is it a bad thing?</p>",
      "rawMarkdown": "> if kaggle want industrial-grade solution, the testing should be \"pure black box\". then I afraid, there will be many shakeup.\n\nYes, there will be shakeup and robust models will survive.  We see it happening each time public test distribution is not close to private test distribution.\n\nIs it a bad thing?",
      "votes": null
    },
    {
      "id": "1213328",
      "postDate": "02/22/2021 04:06:39",
      "content": "<p>\"Is it a bad thing?\"</p>\n<p>it is a  good thing, from an algorithm point of view. <br>\nbut if you know there will be shakeup, then one may be less interested in taking part.<br>\n(e.g. realtime stock prediction)</p>\n<p>I propose something like:</p>\n<ol>\n<li><p>like now: visible public/private test (i.e. similar distribution), this test your algorithm development skill: some prize money</p></li>\n<li><p>new: robustness test (black box test set) this test if your algorithm is really robust:<br>\nadditional prize money</p></li>\n</ol>\n<p>this is what I did in the past for commercial algorithm development. you won't get paid fully unless you pass the black box test.</p>",
      "rawMarkdown": "\"Is it a bad thing?\"\n\nit is a  good thing, from an algorithm point of view. \nbut if you know there will be shakeup, then one may be less interested in taking part.\n(e.g. realtime stock prediction)\n\n\nI propose something like:\n1. like now: visible public/private test (i.e. similar distribution), this test your algorithm development skill: some prize money\n\n2. new: robustness test (black box test set) this test if your algorithm is really robust:\nadditional prize money\n\nthis is what I did in the past for commercial algorithm development. you won't get paid fully unless you pass the black box test.",
      "votes": null
    },
    {
      "id": "1213655",
      "postDate": "02/22/2021 08:58:57",
      "content": "<p><a href=\"https://www.kaggle.com/wubinbai\" target=\"_blank\">@wubinbai</a> How about you share your github link? Refactoring a solution into a github repository which is clean enough to be useful to others is serious effort. It makes me angry to see people taking it for granted or even demanding it. Its courtesy of top teams to share their solutions with the community, knowing that it most likely will work against them in the next competition. Thats even more true with shared code.</p>",
      "rawMarkdown": "wubinbai How about you share your github link? Refactoring a solution into a github repository which is clean enough to be useful to others is serious effort. It makes me angry to see people taking it for granted or even demanding it. Its courtesy of top teams to share their solutions with the community, knowing that it most likely will work against them in the next competition. Thats even more true with shared code.",
      "votes": null
    },
    {
      "id": "1213837",
      "postDate": "02/22/2021 11:24:41",
      "content": "<p>That would be great indeed, but it means more work from host, hence I doubt it will happen.  </p>",
      "rawMarkdown": "That would be great indeed, but it means more work from host, hence I doubt it will happen.",
      "votes": null
    },
    {
      "id": "1213858",
      "postDate": "02/22/2021 11:35:30",
      "content": "<p>Oops… Thanks, I can totally understand the downside of sharing code to winners. I didn't think take it for granted, but it looks to me that the 2019 kaggle freesound competition's top three teams shared their code(previously I was naively thinking that it may due to the policy(which may/ may not be the case)), and a lot of top teams did, unlike this competition, so I just really don't know why. Thanks for the effort, though……</p>",
      "rawMarkdown": "Oops... Thanks, I can totally understand the downside of sharing code to winners. I didn't think take it for granted, but it looks to me that the 2019 kaggle freesound competition's top three teams shared their code(previously I was naively thinking that it may due to the policy(which may/ may not be the case)), and a lot of top teams did, unlike this competition, so I just really don't know why. Thanks for the effort, though......",
      "votes": null
    },
    {
      "id": "1213883",
      "postDate": "02/22/2021 12:00:20",
      "content": "<p>Airbus applied blackbox testing after the competition and selected 2nd team (ours) solution for their system. Unfortunately there were no additional prizes😄 They adapted the codebase, retrained on new data and use it in production <a href=\"https://blog.usejournal.com/important-things-you-should-know-before-organizing-a-kaggle-competition-3911b71701fb\" target=\"_blank\">https://blog.usejournal.com/important-things-you-should-know-before-organizing-a-kaggle-competition-3911b71701fb</a><br>\nThere are just a few known production use cases of Kaggle solutions. I guess if hosts shared their experience after 6-12 months that would be helpful.</p>",
      "rawMarkdown": "Airbus applied blackbox testing after the competition and selected 2nd team (ours) solution for their system. Unfortunately there were no additional prizes😄 They adapted the codebase, retrained on new data and use it in production https://blog.usejournal.com/important-things-you-should-know-before-organizing-a-kaggle-competition-3911b71701fb\nThere are just a few known production use cases of Kaggle solutions. I guess if hosts shared their experience after 6-12 months that would be helpful.",
      "votes": null
    },
    {
      "id": "1213926",
      "postDate": "02/22/2021 12:56:24",
      "content": "<blockquote>\n  <p>I would recommend that public/private test split is never random so that generalization power of models is tested.</p>\n</blockquote>\n<p>I don't think this is a good idea for competitions, as it can make the solutions random. If it is not a random split, then there is some inherent logic in the split, making it even more prone to the top solutions not being robust. In an industry data science project I would always encourage to have a validation set that resembles the data you want to apply it on. In best case you would have validation data including different scenarios and maybe also data shifts, so that you can get a fuller picture. But data collection is unfortunately complicated in many cases.</p>\n<p>In this competition, the training data was obviously very different than the data you want to apply it on, so public LB was your validation set. And that also means to me that hosts expect their future data to come from the same distribution as public and private data. So those solutions that perform better, at the metric of interest, on public LB, performed better on private LB and most likely will also perform better on future data coming from the same distribution.</p>\n<p>In worst case the private LB would suddenly have let's say S1 as the most common one. Then still apparent robust solutions would not perform well, but probably some methods way down the LB that had some lucky scalings on those columns. This has happened before on Kaggle competitions. I think a bit more biased public/private splits can only work for competitions if training data is reasonably representative of the whole test population, which was not the case at all in this competition.</p>\n<p>Overall, it is a fair question to discuss what robust means though. I am all for introducing multiple metrics that can test different aspects of the model more thoroughly. I think there is potential to score solutions on more than one metric. And also, I think that metrics that are not so prone to scalings should be preferred.</p>",
      "rawMarkdown": "> I would recommend that public/private test split is never random so that generalization power of models is tested.\n\nI don't think this is a good idea for competitions, as it can make the solutions random. If it is not a random split, then there is some inherent logic in the split, making it even more prone to the top solutions not being robust. In an industry data science project I would always encourage to have a validation set that resembles the data you want to apply it on. In best case you would have validation data including different scenarios and maybe also data shifts, so that you can get a fuller picture. But data collection is unfortunately complicated in many cases.\n\nIn this competition, the training data was obviously very different than the data you want to apply it on, so public LB was your validation set. And that also means to me that hosts expect their future data to come from the same distribution as public and private data. So those solutions that perform better, at the metric of interest, on public LB, performed better on private LB and most likely will also perform better on future data coming from the same distribution.\n\nIn worst case the private LB would suddenly have let's say S1 as the most common one. Then still apparent robust solutions would not perform well, but probably some methods way down the LB that had some lucky scalings on those columns. This has happened before on Kaggle competitions. I think a bit more biased public/private splits can only work for competitions if training data is reasonably representative of the whole test population, which was not the case at all in this competition.\n\nOverall, it is a fair question to discuss what robust means though. I am all for introducing multiple metrics that can test different aspects of the model more thoroughly. I think there is potential to score solutions on more than one metric. And also, I think that metrics that are not so prone to scalings should be preferred.",
      "votes": null
    },
    {
      "id": "1213956",
      "postDate": "02/22/2021 13:21:45",
      "content": "<p>IMHO, what matters is what provides most value to host.  It would be interesting to get host feedback as Selim suggested.  It maybe that some are looking for code they could put in production, in which case black box private test maybe best.  in other cases they may just want to know what's the best achievable, in which case fitting to test may be fine.  In other cases they just want a review of applicable models and techniques, and getmaterial for a scientific publication, etc.  </p>\n<p>For the rest, we have an opinion vs opinion discussion.  It is very a interesting discussion because opinions are well motivated and backed by solid experience.  Having host tell us what maters would help settle the discussion.</p>\n<p>To conclude, I'll express a variant of my opinion. I like that  public/private split is often not random and time based split in in forecasting competitions Kaggle comps.  And yes, these often have larger shakeup than other competitions.</p>",
      "rawMarkdown": "IMHO, what matters is what provides most value to host.  It would be interesting to get host feedback as Selim suggested.  It maybe that some are looking for code they could put in production, in which case black box private test maybe best.  in other cases they may just want to know what's the best achievable, in which case fitting to test may be fine.  In other cases they just want a review of applicable models and techniques, and getmaterial for a scientific publication, etc.  \n\nFor the rest, we have an opinion vs opinion discussion.  It is very a interesting discussion because opinions are well motivated and backed by solid experience.  Having host tell us what maters would help settle the discussion.\n\nTo conclude, I'll express a variant of my opinion. I like that  public/private split is often not random and time based split in in forecasting competitions Kaggle comps.  And yes, these often have larger shakeup than other competitions.",
      "votes": null
    },
    {
      "id": "1213962",
      "postDate": "02/22/2021 13:30:10",
      "content": "<p>back tot this specific competition, it really looks like data (both train and test recordings) was sampled randomly from the overall distribution described in the host paper, hence any \"overfit\" is not an overfit.  What is at odd is the TP/FP distribution,  I totally missed that and I learned a lot from those who used PP or PL successfully to fit to the data distribution.  Please don't take any of what I write as a way to undermine what I missed.</p>",
      "rawMarkdown": "back tot this specific competition, it really looks like data (both train and test recordings) was sampled randomly from the overall distribution described in the host paper, hence any \"overfit\" is not an overfit.  What is at odd is the TP/FP distribution,  I totally missed that and I learned a lot from those who used PP or PL successfully to fit to the data distribution.  Please don't take any of what I write as a way to undermine what I missed.",
      "votes": null
    },
    {
      "id": "1217563",
      "postDate": "02/25/2021 06:34:11",
      "content": "<p>Thanks for sharing!<br>\nFor those of you who are VERY INTERESTED:<br>\nFYI, I am going to convert all these discussions(may be with codes) into a single txt/pdf file and upload to github in two weeks.(before March 10th) Please upvote this thread if you like. Thanks!</p>",
      "rawMarkdown": "Thanks for sharing!\nFor those of you who are VERY INTERESTED:\nFYI, I am going to convert all these discussions(may be with codes) into a single txt/pdf file and upload to github in two weeks.(before March 10th) Please upvote this thread if you like. Thanks!",
      "votes": null
    },
    {
      "id": "1232070",
      "postDate": "03/09/2021 13:16:57",
      "content": "<p>My first version of all these materials is done! A total of 209 pages pdf!!! Please check my github below, thanks! Any comments or suggestions are welcome!!!<br>\n<strong><em># Let me see this upvote and the github starred, :D</em></strong><br>\n<a href=\"https://github.com/wubinbai/kaggle-rainforest-top-solutions\" target=\"_blank\">https://github.com/wubinbai/kaggle-rainforest-top-solutions</a></p>",
      "rawMarkdown": "My first version of all these materials is done! A total of 209 pages pdf!!! Please check my github below, thanks! Any comments or suggestions are welcome!!!\n***# Let me see this upvote and the github starred, :D***\nhttps://github.com/wubinbai/kaggle-rainforest-top-solutions",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207915,
      "author_name": "wubinbai",
      "author_url": "",
      "post_date": "02/18/2021 04:57:50",
      "content": "<p>when will top3 GitHub links be revealed?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207955,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/18/2021 05:08:38",
          "content": "<p>Typically in 2-3 days, most clean up their code before posting<br>\nHowever, some may choose to just put the approach rather than the code</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211795,
          "author_name": "wubinbai",
          "author_url": "",
          "post_date": "02/20/2021 15:10:07",
          "content": "<p>It looks like that most of the top sharings in this competition are idea-based, no github links… A bit disappointed…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1212779,
          "author_name": "allohvk",
          "author_url": "",
          "post_date": "02/21/2021 15:29:34",
          "content": "<p><a href=\"https://www.kaggle.com/meaninglesslives\" target=\"_blank\">@meaninglesslives</a> has already shared a gold class kernel - <a href=\"https://www.kaggle.com/meaninglesslives/rfcx-minimal\" target=\"_blank\">https://www.kaggle.com/meaninglesslives/rfcx-minimal</a>. Please go thru' and upvote that one which is already there. We can wait till other Github links are published..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213655,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "02/22/2021 08:58:57",
          "content": "<p><a href=\"https://www.kaggle.com/wubinbai\" target=\"_blank\">@wubinbai</a> How about you share your github link? Refactoring a solution into a github repository which is clean enough to be useful to others is serious effort. It makes me angry to see people taking it for granted or even demanding it. Its courtesy of top teams to share their solutions with the community, knowing that it most likely will work against them in the next competition. Thats even more true with shared code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213858,
          "author_name": "wubinbai",
          "author_url": "",
          "post_date": "02/22/2021 11:35:30",
          "content": "<p>Oops… Thanks, I can totally understand the downside of sharing code to winners. I didn't think take it for granted, but it looks to me that the 2019 kaggle freesound competition's top three teams shared their code(previously I was naively thinking that it may due to the policy(which may/ may not be the case)), and a lot of top teams did, unlike this competition, so I just really don't know why. Thanks for the effort, though……</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207983,
      "author_name": "kupchanski",
      "author_url": "",
      "post_date": "02/18/2021 05:19:47",
      "content": "<p>I think you mentioned me twiсe for 4 and 14 places)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207998,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/18/2021 05:35:21",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/kupchanski\" target=\"_blank\">@kupchanski</a>, corrected!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209500,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/18/2021 23:47:18",
      "content": "<p><img src=\"https://i.ibb.co/1z07JZv/Selection-081.png\" alt=\"\"></p>\n<p>I also want to mention that serious kaggling is hard work:</p>\n<p>top-1: trained 120 models<br>\ntop-3: numerous hand-labeling</p>\n<p>one may want to remember this formula in your future competition development cycle:<br>\nbetter modeling --&gt; better data --&gt; better biasing</p>",
      "votes": null,
      "replies": [
        {
          "id": 1209707,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/19/2021 02:41:00",
          "content": "<p>Couldn't agree more <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> !!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1209737,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/19/2021 02:57:29",
          "content": "<p>Nice plot and suggestions Heng.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211574,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/20/2021 10:36:17",
          "content": "<p>Great analysis.</p>\n<p>I never do better biasing because it is a recipe for overfitting in real life ML.  But in Kaggle, when we fully know the test data it is very effective.  Indeed it becomes a matter of overfitting to private test data.  And when public/private split is random, it then becomes a matter of overftting to public test data.</p>\n<p>The downside is that when public/private split is not random then the above results in a major fall in shakedown.  Given I don't overfit to public LB I almost always move up in shakeup.</p>\n<p>Now the question for me is: should i stick to sound ML practice, or should I enter the overfitting game?  I don't like to have to make a choice here.</p>\n<p>From the host perspective, I wonder what is the value of solutions that overfit to test data.  </p>\n<p>I would recommend that public/private test split is never random so that generalization power of models is tested.</p>\n<p>Edit.  Better biasing based on training data is sound in real life ML.  Doing it whenever applicable is great.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211938,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/20/2021 17:20:46",
          "content": "<p>Yes, I agree that all top solutions are overfitted to the test data. Having 3 and 18 species in majority of samples is a leak in some sense, especially for LWRAP metric. <br>\nThough with pseudolabeling on early stages performance is really amazing judging from the logloss on OOF for TP and FP (later stages have worse logloss on FP samples). So with better validation  they can choose models that suit them (for example by using test data for validation and employing other metrics)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211984,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/20/2021 18:08:17",
          "content": "<p>I wonder why they did not use a metric per species, like roc-auc, or f1 score.  </p>\n<blockquote>\n  <p>(later stages have worse logloss on FP samples). </p>\n</blockquote>\n<p>Interesting, it shows that the bias in FP sampling is detrimental at some point.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213195,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/22/2021 00:16:01",
          "content": "<p>from a competition point of view, it is ok to overfit test data. you need very good skills to overfit test data. The private test data is not completely visible, it is a test of skill to see you extract information to guess how the private test data looks like.</p>\n<p>if you know how to overfit test data, you would also know how to generalize, interpret results, etc</p>\n<p>from industry or application point of view, we never overfit test data. instead, I would rather spend more time collect data.</p>\n<hr>\n<p>\"I would recommend that public/private test split is never random so that generalization power of models is tested.\"</p>\n<p>if kaggle want industrial-grade solution, the testing should be \"pure black box\". then I afraid, there will be many shakeup.</p>\n<p>i often regard kaggle solution as only proof-of-concept if we are talking about a commercial application development pipeline. It is just the minimum viable model stage (MVM). some team will take over from the kaggle solution to map up a development milestone and scale up the model for real industrial applications.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213197,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "02/22/2021 00:17:56",
          "content": "<blockquote>\n  <p>all top solutions are overfitted to the test data</p>\n</blockquote>\n<p>because of the addition of extra hand-labels in both training and validation, our lwlrap CV score correlates to the increase in lb, I wouldn't say it's overfitting to test set in our case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213261,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/22/2021 02:28:42",
          "content": "<blockquote>\n  <p>if kaggle want industrial-grade solution, the testing should be \"pure black box\". then I afraid, there will be many shakeup.</p>\n</blockquote>\n<p>Yes, there will be shakeup and robust models will survive.  We see it happening each time public test distribution is not close to private test distribution.</p>\n<p>Is it a bad thing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213328,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/22/2021 04:06:39",
          "content": "<p>\"Is it a bad thing?\"</p>\n<p>it is a  good thing, from an algorithm point of view. <br>\nbut if you know there will be shakeup, then one may be less interested in taking part.<br>\n(e.g. realtime stock prediction)</p>\n<p>I propose something like:</p>\n<ol>\n<li><p>like now: visible public/private test (i.e. similar distribution), this test your algorithm development skill: some prize money</p></li>\n<li><p>new: robustness test (black box test set) this test if your algorithm is really robust:<br>\nadditional prize money</p></li>\n</ol>\n<p>this is what I did in the past for commercial algorithm development. you won't get paid fully unless you pass the black box test.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213837,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/22/2021 11:24:41",
          "content": "<p>That would be great indeed, but it means more work from host, hence I doubt it will happen.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213883,
          "author_name": "selimsef",
          "author_url": "",
          "post_date": "02/22/2021 12:00:20",
          "content": "<p>Airbus applied blackbox testing after the competition and selected 2nd team (ours) solution for their system. Unfortunately there were no additional prizes😄 They adapted the codebase, retrained on new data and use it in production <a href=\"https://blog.usejournal.com/important-things-you-should-know-before-organizing-a-kaggle-competition-3911b71701fb\" target=\"_blank\">https://blog.usejournal.com/important-things-you-should-know-before-organizing-a-kaggle-competition-3911b71701fb</a><br>\nThere are just a few known production use cases of Kaggle solutions. I guess if hosts shared their experience after 6-12 months that would be helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213926,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/22/2021 12:56:24",
          "content": "<blockquote>\n  <p>I would recommend that public/private test split is never random so that generalization power of models is tested.</p>\n</blockquote>\n<p>I don't think this is a good idea for competitions, as it can make the solutions random. If it is not a random split, then there is some inherent logic in the split, making it even more prone to the top solutions not being robust. In an industry data science project I would always encourage to have a validation set that resembles the data you want to apply it on. In best case you would have validation data including different scenarios and maybe also data shifts, so that you can get a fuller picture. But data collection is unfortunately complicated in many cases.</p>\n<p>In this competition, the training data was obviously very different than the data you want to apply it on, so public LB was your validation set. And that also means to me that hosts expect their future data to come from the same distribution as public and private data. So those solutions that perform better, at the metric of interest, on public LB, performed better on private LB and most likely will also perform better on future data coming from the same distribution.</p>\n<p>In worst case the private LB would suddenly have let's say S1 as the most common one. Then still apparent robust solutions would not perform well, but probably some methods way down the LB that had some lucky scalings on those columns. This has happened before on Kaggle competitions. I think a bit more biased public/private splits can only work for competitions if training data is reasonably representative of the whole test population, which was not the case at all in this competition.</p>\n<p>Overall, it is a fair question to discuss what robust means though. I am all for introducing multiple metrics that can test different aspects of the model more thoroughly. I think there is potential to score solutions on more than one metric. And also, I think that metrics that are not so prone to scalings should be preferred.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213956,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/22/2021 13:21:45",
          "content": "<p>IMHO, what matters is what provides most value to host.  It would be interesting to get host feedback as Selim suggested.  It maybe that some are looking for code they could put in production, in which case black box private test maybe best.  in other cases they may just want to know what's the best achievable, in which case fitting to test may be fine.  In other cases they just want a review of applicable models and techniques, and getmaterial for a scientific publication, etc.  </p>\n<p>For the rest, we have an opinion vs opinion discussion.  It is very a interesting discussion because opinions are well motivated and backed by solid experience.  Having host tell us what maters would help settle the discussion.</p>\n<p>To conclude, I'll express a variant of my opinion. I like that  public/private split is often not random and time based split in in forecasting competitions Kaggle comps.  And yes, these often have larger shakeup than other competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1213962,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/22/2021 13:30:10",
          "content": "<p>back tot this specific competition, it really looks like data (both train and test recordings) was sampled randomly from the overall distribution described in the host paper, hence any \"overfit\" is not an overfit.  What is at odd is the TP/FP distribution,  I totally missed that and I learned a lot from those who used PP or PL successfully to fit to the data distribution.  Please don't take any of what I write as a way to undermine what I missed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1211598,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/20/2021 10:51:13",
      "content": "<p>You should update ranks now that LB is final.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1211638,
          "author_name": "kmldas",
          "author_url": "",
          "post_date": "02/20/2021 12:00:23",
          "content": "<p>Thanks for the reminder!! Updated the ranks and some new discussions  as well</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1217563,
      "author_name": "wubinbai",
      "author_url": "",
      "post_date": "02/25/2021 06:34:11",
      "content": "<p>Thanks for sharing!<br>\nFor those of you who are VERY INTERESTED:<br>\nFYI, I am going to convert all these discussions(may be with codes) into a single txt/pdf file and upload to github in two weeks.(before March 10th) Please upvote this thread if you like. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1232070,
          "author_name": "wubinbai",
          "author_url": "",
          "post_date": "03/09/2021 13:16:57",
          "content": "<p>My first version of all these materials is done! A total of 209 pages pdf!!! Please check my github below, thanks! Any comments or suggestions are welcome!!!<br>\n<strong><em># Let me see this upvote and the github starred, :D</em></strong><br>\n<a href=\"https://github.com/wubinbai/kaggle-rainforest-top-solutions\" target=\"_blank\">https://github.com/wubinbai/kaggle-rainforest-top-solutions</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1207797": "Hi All,\n\nCongratulations to all the winners!!\n\nI am consolidating the top 50 results, as they become available. The discussion as well as the code in GitHub/Kaggle. Hoping this becomes  easy to refer to different top solutions to learn from!\nIn case, I missed your discussion/code, please comment and I will update! Thanks for sharing!\n\nNote: Names below are for the one who posted, many are part of a larger team\n\nRank 1 @philippsinger\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220563\n\n\nRank 2 @selimsef\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220760\n\n\nRank 3 @dicksonchin93\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220522\nRank 3  @meaninglesslives\nCode: https://www.kaggle.com/meaninglesslives/rfcx-minimal\n\nRank 4 @kupchanski\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220342\n\nRank 5 @takamichitoda\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220432\n\nRank 6  @antorsae\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220446\nRank 6  @amezet\nhttps://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220981\n\nRank 7 @pestipeti\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220443\n\nRank 8 @shinoda18\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220803\n\n\nRank 9 @cdeotte\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220389\nCode: https://www.kaggle.com/cdeotte/rainforest-post-process-lb-0-970\n\n===========\n\nRank 10 @bigironsphere\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220319\n\nRank 11 @cpmpml\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220304\n\nRank 13 @reppic\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220308\n\nRank 14 @prateekagnihotri\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220725\n\nRank 16 @hidehisaarai1213\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220450\n\nRank 18 @hengck23\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220309\n\n===============\n\nRank 21 @xmpgek\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220436\n\nRank 23 @dathudeptrai\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220972\n\nRank 25 @rytisva88\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220314\n\nRank 27 @mnpinto\nDiscussion:https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220306\n\nRank 32 @harangdev\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220322\n\n\nRank 33 @yururoi\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220753\n\nRank 38 : @vzaguskin\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220386\n\nRank 43 : @shinmurashinmura\nDiscussion: https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220335\n\n\nCongratulations to everyone! and thanks for sharing!!",
    "1207915": "when will top3 GitHub links be revealed?",
    "1207955": "Typically in 2-3 days, most clean up their code before posting\nHowever, some may choose to just put the approach rather than the code",
    "1207983": "I think you mentioned me twiсe for 4 and 14 places)",
    "1207998": "Thanks @kupchanski, corrected!!",
    "1209500": "![](https://i.ibb.co/1z07JZv/Selection-081.png)\n\nI also want to mention that serious kaggling is hard work:\n\ntop-1: trained 120 models\ntop-3: numerous hand-labeling\n\n\none may want to remember this formula in your future competition development cycle:\nbetter modeling --> better data --> better biasing",
    "1209707": "Couldn't agree more @hengck23 !!",
    "1209737": "Nice plot and suggestions Heng.",
    "1211574": "Great analysis.\n\nI never do better biasing because it is a recipe for overfitting in real life ML.  But in Kaggle, when we fully know the test data it is very effective.  Indeed it becomes a matter of overfitting to private test data.  And when public/private split is random, it then becomes a matter of overftting to public test data.\n\nThe downside is that when public/private split is not random then the above results in a major fall in shakedown.  Given I don't overfit to public LB I almost always move up in shakeup.\n\nNow the question for me is: should i stick to sound ML practice, or should I enter the overfitting game?  I don't like to have to make a choice here.\n\nFrom the host perspective, I wonder what is the value of solutions that overfit to test data.  \n\nI would recommend that public/private test split is never random so that generalization power of models is tested.\n\nEdit.  Better biasing based on training data is sound in real life ML.  Doing it whenever applicable is great.",
    "1211598": "You should update ranks now that LB is final.",
    "1211638": "Thanks for the reminder!! Updated the ranks and some new discussions  as well",
    "1211795": "It looks like that most of the top sharings in this competition are idea-based, no github links... A bit disappointed...",
    "1211938": "Yes, I agree that all top solutions are overfitted to the test data. Having 3 and 18 species in majority of samples is a leak in some sense, especially for LWRAP metric. \nThough with pseudolabeling on early stages performance is really amazing judging from the logloss on OOF for TP and FP (later stages have worse logloss on FP samples). So with better validation  they can choose models that suit them (for example by using test data for validation and employing other metrics)",
    "1211984": "I wonder why they did not use a metric per species, like roc-auc, or f1 score.  \n\n> (later stages have worse logloss on FP samples). \n\nInteresting, it shows that the bias in FP sampling is detrimental at some point.",
    "1212779": "meaninglesslives has already shared a gold class kernel - https://www.kaggle.com/meaninglesslives/rfcx-minimal. Please go thru' and upvote that one which is already there. We can wait till other Github links are published..",
    "1213195": "from a competition point of view, it is ok to overfit test data. you need very good skills to overfit test data. The private test data is not completely visible, it is a test of skill to see you extract information to guess how the private test data looks like.\n\nif you know how to overfit test data, you would also know how to generalize, interpret results, etc\n\nfrom industry or application point of view, we never overfit test data. instead, I would rather spend more time collect data.\n\n---\n\n\"I would recommend that public/private test split is never random so that generalization power of models is tested.\"\n\nif kaggle want industrial-grade solution, the testing should be \"pure black box\". then I afraid, there will be many shakeup.\n\ni often regard kaggle solution as only proof-of-concept if we are talking about a commercial application development pipeline. It is just the minimum viable model stage (MVM). some team will take over from the kaggle solution to map up a development milestone and scale up the model for real industrial applications.",
    "1213197": "> all top solutions are overfitted to the test data\n\nbecause of the addition of extra hand-labels in both training and validation, our lwlrap CV score correlates to the increase in lb, I wouldn't say it's overfitting to test set in our case.",
    "1213261": "> if kaggle want industrial-grade solution, the testing should be \"pure black box\". then I afraid, there will be many shakeup.\n\nYes, there will be shakeup and robust models will survive.  We see it happening each time public test distribution is not close to private test distribution.\n\nIs it a bad thing?",
    "1213328": "\"Is it a bad thing?\"\n\nit is a  good thing, from an algorithm point of view. \nbut if you know there will be shakeup, then one may be less interested in taking part.\n(e.g. realtime stock prediction)\n\n\nI propose something like:\n1. like now: visible public/private test (i.e. similar distribution), this test your algorithm development skill: some prize money\n\n2. new: robustness test (black box test set) this test if your algorithm is really robust:\nadditional prize money\n\nthis is what I did in the past for commercial algorithm development. you won't get paid fully unless you pass the black box test.",
    "1213655": "wubinbai How about you share your github link? Refactoring a solution into a github repository which is clean enough to be useful to others is serious effort. It makes me angry to see people taking it for granted or even demanding it. Its courtesy of top teams to share their solutions with the community, knowing that it most likely will work against them in the next competition. Thats even more true with shared code.",
    "1213837": "That would be great indeed, but it means more work from host, hence I doubt it will happen.",
    "1213858": "Oops... Thanks, I can totally understand the downside of sharing code to winners. I didn't think take it for granted, but it looks to me that the 2019 kaggle freesound competition's top three teams shared their code(previously I was naively thinking that it may due to the policy(which may/ may not be the case)), and a lot of top teams did, unlike this competition, so I just really don't know why. Thanks for the effort, though......",
    "1213883": "Airbus applied blackbox testing after the competition and selected 2nd team (ours) solution for their system. Unfortunately there were no additional prizes😄 They adapted the codebase, retrained on new data and use it in production https://blog.usejournal.com/important-things-you-should-know-before-organizing-a-kaggle-competition-3911b71701fb\nThere are just a few known production use cases of Kaggle solutions. I guess if hosts shared their experience after 6-12 months that would be helpful.",
    "1213926": "> I would recommend that public/private test split is never random so that generalization power of models is tested.\n\nI don't think this is a good idea for competitions, as it can make the solutions random. If it is not a random split, then there is some inherent logic in the split, making it even more prone to the top solutions not being robust. In an industry data science project I would always encourage to have a validation set that resembles the data you want to apply it on. In best case you would have validation data including different scenarios and maybe also data shifts, so that you can get a fuller picture. But data collection is unfortunately complicated in many cases.\n\nIn this competition, the training data was obviously very different than the data you want to apply it on, so public LB was your validation set. And that also means to me that hosts expect their future data to come from the same distribution as public and private data. So those solutions that perform better, at the metric of interest, on public LB, performed better on private LB and most likely will also perform better on future data coming from the same distribution.\n\nIn worst case the private LB would suddenly have let's say S1 as the most common one. Then still apparent robust solutions would not perform well, but probably some methods way down the LB that had some lucky scalings on those columns. This has happened before on Kaggle competitions. I think a bit more biased public/private splits can only work for competitions if training data is reasonably representative of the whole test population, which was not the case at all in this competition.\n\nOverall, it is a fair question to discuss what robust means though. I am all for introducing multiple metrics that can test different aspects of the model more thoroughly. I think there is potential to score solutions on more than one metric. And also, I think that metrics that are not so prone to scalings should be preferred.",
    "1213956": "IMHO, what matters is what provides most value to host.  It would be interesting to get host feedback as Selim suggested.  It maybe that some are looking for code they could put in production, in which case black box private test maybe best.  in other cases they may just want to know what's the best achievable, in which case fitting to test may be fine.  In other cases they just want a review of applicable models and techniques, and getmaterial for a scientific publication, etc.  \n\nFor the rest, we have an opinion vs opinion discussion.  It is very a interesting discussion because opinions are well motivated and backed by solid experience.  Having host tell us what maters would help settle the discussion.\n\nTo conclude, I'll express a variant of my opinion. I like that  public/private split is often not random and time based split in in forecasting competitions Kaggle comps.  And yes, these often have larger shakeup than other competitions.",
    "1213962": "back tot this specific competition, it really looks like data (both train and test recordings) was sampled randomly from the overall distribution described in the host paper, hence any \"overfit\" is not an overfit.  What is at odd is the TP/FP distribution,  I totally missed that and I learned a lot from those who used PP or PL successfully to fit to the data distribution.  Please don't take any of what I write as a way to undermine what I missed.",
    "1217563": "Thanks for sharing!\nFor those of you who are VERY INTERESTED:\nFYI, I am going to convert all these discussions(may be with codes) into a single txt/pdf file and upload to github in two weeks.(before March 10th) Please upvote this thread if you like. Thanks!",
    "1232070": "My first version of all these materials is done! A total of 209 pages pdf!!! Please check my github below, thanks! Any comments or suggestions are welcome!!!\n***# Let me see this upvote and the github starred, :D***\nhttps://github.com/wubinbai/kaggle-rainforest-top-solutions"
  },
  "source": "meta"
}