{
  "id": 220949,
  "title": "What were the top learnings from this competition?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/220949",
  "author_name": "",
  "post_date": "2021-02-20T08:26:14.512038600Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I certainly learnt a lot by following the discussion and public notebooks. Thanks to the community for sharing their best ideas and encouraging everyone. Some of my learnings are as below</p>\n<ul>\n<li>Optimize on single model and then create ensemble</li>\n<li>Trust your cross validation score and not the leaderboard score</li>\n<li>Bagging of different models gives better results</li>\n</ul>\n<p>Would be great if we get opinion on other Kaggle experts as well here.</p>",
  "messages": [
    {
      "id": "1211447",
      "postDate": "02/20/2021 08:26:14",
      "content": "<p>I certainly learnt a lot by following the discussion and public notebooks. Thanks to the community for sharing their best ideas and encouraging everyone. Some of my learnings are as below</p>\n<ul>\n<li>Optimize on single model and then create ensemble</li>\n<li>Trust your cross validation score and not the leaderboard score</li>\n<li>Bagging of different models gives better results</li>\n</ul>\n<p>Would be great if we get opinion on other Kaggle experts as well here.</p>",
      "rawMarkdown": "I certainly learnt a lot by following the discussion and public notebooks. Thanks to the community for sharing their best ideas and encouraging everyone. Some of my learnings are as below\n\n- Optimize on single model and then create ensemble\n- Trust your cross validation score and not the leaderboard score\n- Bagging of different models gives better results\n\nWould be great if we get opinion on other Kaggle experts as well here.",
      "votes": null
    },
    {
      "id": "1211980",
      "postDate": "02/20/2021 18:05:47",
      "content": "<p>A good thread form <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339#1209500\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339#1209500</a></p>",
      "rawMarkdown": "A good thread form @hengck23 https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339#1209500",
      "votes": null
    },
    {
      "id": "1212287",
      "postDate": "02/21/2021 04:32:44",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for sharing, will follow that.</p>",
      "rawMarkdown": "Thanks @cpmpml for sharing, will follow that.",
      "votes": null
    },
    {
      "id": "1212782",
      "postDate": "02/21/2021 15:38:19",
      "content": "<p>I dont think option 2 is applicable here is it? One of the rare comps where most top folks relied on LB?</p>\n<p>Another thing missing is PP. I am sure you would have checked out Chri's great post on Post-processing. </p>\n<p>pseudo labelling seemed to have worked for most. For me I was trying to see how I could get more TPs as my (binary) model was learning to tilt more towards classifying things as FP. I thought of plucking some data from 'test' but quickly banished my thoughts as it amounted to leakage. Later I learnt that pseudo-labelling is a technique to be used. I assume folks would have used the train data (the 1000's of unused seconds that we typically discard when cropping) to get more TPs…though I did read an external blog where someone uses test data also.. Will try to rebuild some of my models using some of the learnings and will try to post sometime in the future…</p>",
      "rawMarkdown": "I dont think option 2 is applicable here is it? One of the rare comps where most top folks relied on LB?\n\nAnother thing missing is PP. I am sure you would have checked out Chri's great post on Post-processing. \n\npseudo labelling seemed to have worked for most. For me I was trying to see how I could get more TPs as my (binary) model was learning to tilt more towards classifying things as FP. I thought of plucking some data from 'test' but quickly banished my thoughts as it amounted to leakage. Later I learnt that pseudo-labelling is a technique to be used. I assume folks would have used the train data (the 1000's of unused seconds that we typically discard when cropping) to get more TPs...though I did read an external blog where someone uses test data also.. Will try to rebuild some of my models using some of the learnings and will try to post sometime in the future...",
      "votes": null
    },
    {
      "id": "1212931",
      "postDate": "02/21/2021 18:21:41",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> !</p>",
      "rawMarkdown": "Thanks @allohvk !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1211980,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/20/2021 18:05:47",
      "content": "<p>A good thread form <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339#1209500\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339#1209500</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1212287,
          "author_name": "saurabhbagchi",
          "author_url": "",
          "post_date": "02/21/2021 04:32:44",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for sharing, will follow that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212782,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "02/21/2021 15:38:19",
      "content": "<p>I dont think option 2 is applicable here is it? One of the rare comps where most top folks relied on LB?</p>\n<p>Another thing missing is PP. I am sure you would have checked out Chri's great post on Post-processing. </p>\n<p>pseudo labelling seemed to have worked for most. For me I was trying to see how I could get more TPs as my (binary) model was learning to tilt more towards classifying things as FP. I thought of plucking some data from 'test' but quickly banished my thoughts as it amounted to leakage. Later I learnt that pseudo-labelling is a technique to be used. I assume folks would have used the train data (the 1000's of unused seconds that we typically discard when cropping) to get more TPs…though I did read an external blog where someone uses test data also.. Will try to rebuild some of my models using some of the learnings and will try to post sometime in the future…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1212931,
          "author_name": "saurabhbagchi",
          "author_url": "",
          "post_date": "02/21/2021 18:21:41",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/allohvk\" target=\"_blank\">@allohvk</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1211447": "I certainly learnt a lot by following the discussion and public notebooks. Thanks to the community for sharing their best ideas and encouraging everyone. Some of my learnings are as below\n\n- Optimize on single model and then create ensemble\n- Trust your cross validation score and not the leaderboard score\n- Bagging of different models gives better results\n\nWould be great if we get opinion on other Kaggle experts as well here.",
    "1211980": "A good thread form @hengck23 https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/220339#1209500",
    "1212287": "Thanks @cpmpml for sharing, will follow that.",
    "1212782": "I dont think option 2 is applicable here is it? One of the rare comps where most top folks relied on LB?\n\nAnother thing missing is PP. I am sure you would have checked out Chri's great post on Post-processing. \n\npseudo labelling seemed to have worked for most. For me I was trying to see how I could get more TPs as my (binary) model was learning to tilt more towards classifying things as FP. I thought of plucking some data from 'test' but quickly banished my thoughts as it amounted to leakage. Later I learnt that pseudo-labelling is a technique to be used. I assume folks would have used the train data (the 1000's of unused seconds that we typically discard when cropping) to get more TPs...though I did read an external blog where someone uses test data also.. Will try to rebuild some of my models using some of the learnings and will try to post sometime in the future...",
    "1212931": "Thanks @allohvk !"
  },
  "source": "meta"
}