{
  "id": 206138,
  "title": "What is your validation scheme?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/206138",
  "author_name": "Serk0",
  "post_date": "2020-12-23T11:42:14.292000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I split train_tp with stratified kfold where's y is species_id + song_type.<br>\nTrain CNN on cropped 1s audio of species song.</p>\n<p>When do inference, use sliding window of 1 sec with hop_size 0.5 sec on 1min audio. Agregate final probabilities with mean fuction.</p>\n<p>I found that this validation scheme has big difference with LB. For example, Validation lalwrap 0.39, LB 0.32 (I have CNN with 2 conv layers only)</p>\n<p>Think that it could be because of train/test difference,   for example, train has only 65+ audios where 2 or more species has represented. </p>\n<p>What kind of validation scheme do you use? Thanks.</p>",
  "messages": [
    {
      "id": 1123635,
      "postDate": "2020-12-23T11:42:14.293Z",
      "content": "<p>I split train_tp with stratified kfold where's y is species_id + song_type.<br>\nTrain CNN on cropped 1s audio of species song.</p>\n<p>When do inference, use sliding window of 1 sec with hop_size 0.5 sec on 1min audio. Agregate final probabilities with mean fuction.</p>\n<p>I found that this validation scheme has big difference with LB. For example, Validation lalwrap 0.39, LB 0.32 (I have CNN with 2 conv layers only)</p>\n<p>Think that it could be because of train/test difference,   for example, train has only 65+ audios where 2 or more species has represented. </p>\n<p>What kind of validation scheme do you use? Thanks.</p>",
      "rawMarkdown": "I split train_tp with stratified kfold where's y is species_id + song_type.\nTrain CNN on cropped 1s audio of species song.\n\nWhen do inference, use sliding window of 1 sec with hop_size 0.5 sec on 1min audio. Agregate final probabilities with mean fuction.\n\nI found that this validation scheme has big difference with LB. For example, Validation lalwrap 0.39, LB 0.32 (I have CNN with 2 conv layers only)\n\nThink that it could be because of train/test difference,   for example, train has only 65+ audios where 2 or more species has represented. \n\nWhat kind of validation scheme do you use? Thanks."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1123635": "I split train_tp with stratified kfold where's y is species_id + song_type.\nTrain CNN on cropped 1s audio of species song.\n\nWhen do inference, use sliding window of 1 sec with hop_size 0.5 sec on 1min audio. Agregate final probabilities with mean fuction.\n\nI found that this validation scheme has big difference with LB. For example, Validation lalwrap 0.39, LB 0.32 (I have CNN with 2 conv layers only)\n\nThink that it could be because of train/test difference,   for example, train has only 65+ audios where 2 or more species has represented. \n\nWhat kind of validation scheme do you use? Thanks."
  }
}