{
  "id": 215233,
  "title": "Improper validation procedure (and how to fix it)",
  "url": "/competitions/rfcx-species-audio-detection/discussion/215233",
  "author_name": "",
  "post_date": "2021-01-29T06:58:57.616859300Z",
  "votes": 26,
  "comment_count": 8,
  "views": 0,
  "content": "<p>One of the first things I do in most competitions is check out what people have already done and what has seemed to be successful. I have noticed and wanted to point out a few things I have researched that might be useful to others and help mitigate some of the gap between train, validation and leaderboard.</p>\n<p>The first thing I have seen is that on the validation step people are still using the cropping trick to focus on the regions that were labeled. This is problematic because the procedure is not matched on the test set, you cannot crop in on the regions of interest on the test set. You must make predictions for all of the audio. For this reason you must mimic that procedure during validation, predict on all 5 second chunks and then aggregate and evaluate the performance. How you perform on a crop right near the audio of interest is not a perfect analog to how your model will perform on multiple segments aggregated.</p>\n<p><img src=\"https://i.imgur.com/ddNJ8jL.png\" alt=\"\"></p>\n<p>The above is what validation currently looks like. A random crop is placed around where a label is found and then the model is evaluated just for that single region. This misses the model's ability to reject any misleading signal that may have occurred anywhere else in the audio. On top of that you can see variation with a static model, one that has not had any training applied to it will yield a different result. This is problematic if you are applying early stopping because you might see a model appear to have better performance when really it just got a lucky crop upon validation. An additional thing that this is missing is when an important sound event might be split across multiple splits on the test set. </p>\n<p><img src=\"https://i.imgur.com/awvPync.png\" alt=\"\"><br>\nAn alternate more appropriate validation scheme is to apply the same procedure that will be used on the test set. That is, use the same inference procedure and see how it performs given the full audio clip and using the aggregation method you settle. It is important to validate like this because a significant number of things can happen with those extra audio segments you are passing the model. Slightly terrifyingly this model was not only validating against a different task than it will see at inference time, it is mimicking the training procedure that is making the model blind to these undocumented regions of audio. Someone could theoretically have a model that achieves perfect performance on the cropped regions but shows poorly on the leaderboard because those untrained and unvalidated regions are causing problems with predictions. </p>\n<p>Using the original method I see the kernel achieving ~.85 on validation but ~.74 on lb. Correcting this validation error alone makes it so validation is .8 and lb is .776. Close to within the variation I see between epochs and probably attributable to .8 being the peak number. Interestingly you would think without altering the training procedure at all there would be no score change, but actually selecting a better early stopping checkpoint against a more rigorous validation procedure can be the explanation for this. </p>",
  "messages": [
    {
      "id": "1175436",
      "postDate": "01/29/2021 06:58:57",
      "content": "<p>One of the first things I do in most competitions is check out what people have already done and what has seemed to be successful. I have noticed and wanted to point out a few things I have researched that might be useful to others and help mitigate some of the gap between train, validation and leaderboard.</p>\n<p>The first thing I have seen is that on the validation step people are still using the cropping trick to focus on the regions that were labeled. This is problematic because the procedure is not matched on the test set, you cannot crop in on the regions of interest on the test set. You must make predictions for all of the audio. For this reason you must mimic that procedure during validation, predict on all 5 second chunks and then aggregate and evaluate the performance. How you perform on a crop right near the audio of interest is not a perfect analog to how your model will perform on multiple segments aggregated.</p>\n<p><img src=\"https://i.imgur.com/ddNJ8jL.png\" alt=\"\"></p>\n<p>The above is what validation currently looks like. A random crop is placed around where a label is found and then the model is evaluated just for that single region. This misses the model's ability to reject any misleading signal that may have occurred anywhere else in the audio. On top of that you can see variation with a static model, one that has not had any training applied to it will yield a different result. This is problematic if you are applying early stopping because you might see a model appear to have better performance when really it just got a lucky crop upon validation. An additional thing that this is missing is when an important sound event might be split across multiple splits on the test set. </p>\n<p><img src=\"https://i.imgur.com/awvPync.png\" alt=\"\"><br>\nAn alternate more appropriate validation scheme is to apply the same procedure that will be used on the test set. That is, use the same inference procedure and see how it performs given the full audio clip and using the aggregation method you settle. It is important to validate like this because a significant number of things can happen with those extra audio segments you are passing the model. Slightly terrifyingly this model was not only validating against a different task than it will see at inference time, it is mimicking the training procedure that is making the model blind to these undocumented regions of audio. Someone could theoretically have a model that achieves perfect performance on the cropped regions but shows poorly on the leaderboard because those untrained and unvalidated regions are causing problems with predictions. </p>\n<p>Using the original method I see the kernel achieving ~.85 on validation but ~.74 on lb. Correcting this validation error alone makes it so validation is .8 and lb is .776. Close to within the variation I see between epochs and probably attributable to .8 being the peak number. Interestingly you would think without altering the training procedure at all there would be no score change, but actually selecting a better early stopping checkpoint against a more rigorous validation procedure can be the explanation for this. </p>",
      "rawMarkdown": "One of the first things I do in most competitions is check out what people have already done and what has seemed to be successful. I have noticed and wanted to point out a few things I have researched that might be useful to others and help mitigate some of the gap between train, validation and leaderboard.\n\nThe first thing I have seen is that on the validation step people are still using the cropping trick to focus on the regions that were labeled. This is problematic because the procedure is not matched on the test set, you cannot crop in on the regions of interest on the test set. You must make predictions for all of the audio. For this reason you must mimic that procedure during validation, predict on all 5 second chunks and then aggregate and evaluate the performance. How you perform on a crop right near the audio of interest is not a perfect analog to how your model will perform on multiple segments aggregated.\n\n![](https://i.imgur.com/ddNJ8jL.png)\n\nThe above is what validation currently looks like. A random crop is placed around where a label is found and then the model is evaluated just for that single region. This misses the model's ability to reject any misleading signal that may have occurred anywhere else in the audio. On top of that you can see variation with a static model, one that has not had any training applied to it will yield a different result. This is problematic if you are applying early stopping because you might see a model appear to have better performance when really it just got a lucky crop upon validation. An additional thing that this is missing is when an important sound event might be split across multiple splits on the test set. \n\n![](https://i.imgur.com/awvPync.png)\nAn alternate more appropriate validation scheme is to apply the same procedure that will be used on the test set. That is, use the same inference procedure and see how it performs given the full audio clip and using the aggregation method you settle. It is important to validate like this because a significant number of things can happen with those extra audio segments you are passing the model. Slightly terrifyingly this model was not only validating against a different task than it will see at inference time, it is mimicking the training procedure that is making the model blind to these undocumented regions of audio. Someone could theoretically have a model that achieves perfect performance on the cropped regions but shows poorly on the leaderboard because those untrained and unvalidated regions are causing problems with predictions. \n\nUsing the original method I see the kernel achieving ~.85 on validation but ~.74 on lb. Correcting this validation error alone makes it so validation is .8 and lb is .776. Close to within the variation I see between epochs and probably attributable to .8 being the peak number. Interestingly you would think without altering the training procedure at all there would be no score change, but actually selecting a better early stopping checkpoint against a more rigorous validation procedure can be the explanation for this.",
      "votes": null
    },
    {
      "id": "1175443",
      "postDate": "01/29/2021 07:08:23",
      "content": "<p>Down that avenue of thinking, I wanted to close the gap between training and validation/test so I played with the idea of moving away from the varied cropping technique entirely. Instead of using a single crop and its label as a training sample I tried training on the fill collection of overlapping segments and the full label for the audio and then applied max pooling across the segments and then used BCE to train with that reduction. Surprisingly that seemed to perform pretty close to the cropping technique. </p>\n<p>I would think since the training samples were very similar every pass forward it would memorize the results pretty quickly, but I guess between specaugment and the other noise augmentations it managed to perform reasonably well. </p>\n<p>This is obviously much slower and more compute intensive because it is the full 1 minute of audio all at once when really only a small region is really worth learning about, but figured it was worth exploring and was surprised to find the gap between the two was not very large.</p>",
      "rawMarkdown": "Down that avenue of thinking, I wanted to close the gap between training and validation/test so I played with the idea of moving away from the varied cropping technique entirely. Instead of using a single crop and its label as a training sample I tried training on the fill collection of overlapping segments and the full label for the audio and then applied max pooling across the segments and then used BCE to train with that reduction. Surprisingly that seemed to perform pretty close to the cropping technique. \n\n I would think since the training samples were very similar every pass forward it would memorize the results pretty quickly, but I guess between specaugment and the other noise augmentations it managed to perform reasonably well. \n\nThis is obviously much slower and more compute intensive because it is the full 1 minute of audio all at once when really only a small region is really worth learning about, but figured it was worth exploring and was surprised to find the gap between the two was not very large.",
      "votes": null
    },
    {
      "id": "1175450",
      "postDate": "01/29/2021 07:14:18",
      "content": "<p>An additional thing I tried was applying a second loss to the model with slightly stronger labeling. So on top of training on the predictions from the max pooling of the segments I also trained the model on the segment level labels. The 60s of audio is split into 11 overlapping 10s pieces with 5s stride and I could associate labels with each segment. </p>\n<p>If there was no segment then I would zero out the loss so the model was not penalized in regions that were simply unlabeled. So then the model is trying not only to get the correct prediction at the macro level it is also trying to make sure that the segment where the audio actually occurred was the one that lit up the max pooling layer. </p>\n<p>This performed almost exactly on par as the single loss model, but converged much more quickly. Not sure if I can attribute this to it truly helping or just the doubling of the loss value with both of them active. </p>\n<p>This technique could be further extended to include the false positives as well because we know those regions of the audio are 0's so we can persuade our model to output zeros rather than just nullifying the loss. </p>",
      "rawMarkdown": "An additional thing I tried was applying a second loss to the model with slightly stronger labeling. So on top of training on the predictions from the max pooling of the segments I also trained the model on the segment level labels. The 60s of audio is split into 11 overlapping 10s pieces with 5s stride and I could associate labels with each segment. \n\nIf there was no segment then I would zero out the loss so the model was not penalized in regions that were simply unlabeled. So then the model is trying not only to get the correct prediction at the macro level it is also trying to make sure that the segment where the audio actually occurred was the one that lit up the max pooling layer. \n\nThis performed almost exactly on par as the single loss model, but converged much more quickly. Not sure if I can attribute this to it truly helping or just the doubling of the loss value with both of them active. \n\nThis technique could be further extended to include the false positives as well because we know those regions of the audio are 0's so we can persuade our model to output zeros rather than just nullifying the loss.",
      "votes": null
    },
    {
      "id": "1175743",
      "postDate": "01/29/2021 10:10:41",
      "content": "<p>Thanks. It's vary valuable to see how others approach problems. Truly appreciate it.</p>",
      "rawMarkdown": "Thanks. It's vary valuable to see how others approach problems. Truly appreciate it.",
      "votes": null
    },
    {
      "id": "1176455",
      "postDate": "01/29/2021 16:24:41",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> </p>\n<p>Thanks for nice write up on the validation topic.</p>\n<p>I also did a migration from random cropping to full clip validation ( exactly the same as discussed in your write-up). Although the LB and CV gap narrows as a result, i find that full clip validation score is still very noisy.  Training a single fold out of 5 fold, i found that slight improvement on lwlrap_framewise_aggregate score and validation loss have no correlation with the LB performance.  A slight improvement in val loss and lwlrap could lead to 1-3% drop in LB.</p>\n<p>Trying to establish a better correlation between local validation and LB, I am switch to relying on computing validation metric on the 5 fold concatenated oof predictions. I have not been able to conclude if this resolve the lack of CV &amp; LB correlation issue. But so far, I have a <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215389\" target=\"_blank\">counter-example</a>.  </p>\n<p>1 explanation for people's failures to establish good correlation is the poor quality of training labels ( many labels missing). Therefore, slight improvement or worsening of valid_loss and metric really don't mean that muhc ?? Along this line, I am also trying to find different metric ( other than the comp eval metric) that is more robust to imperfect label and thus better correlates CV and LB score.</p>\n<p>BTW, I have experimented with using different stride in inference. And I found that inferring on overlapping segments ( stride &lt; period ) of audio produces slightly worse result ( on the exact same model) compared to non-overlapping segments. It seems to be pretty counter-intuitive to me. </p>\n<p>Hope your exploration could lead to good CV scheme. Right now, this lack fo CV &amp; LB correlation is driving me crazy. Every submission feels like a lottery to me</p>",
      "rawMarkdown": "ryches \n\nThanks for nice write up on the validation topic.\n\nI also did a migration from random cropping to full clip validation ( exactly the same as discussed in your write-up). Although the LB and CV gap narrows as a result, i find that full clip validation score is still very noisy.  Training a single fold out of 5 fold, i found that slight improvement on lwlrap_framewise_aggregate score and validation loss have no correlation with the LB performance.  A slight improvement in val loss and lwlrap could lead to 1-3% drop in LB.\n\nTrying to establish a better correlation between local validation and LB, I am switch to relying on computing validation metric on the 5 fold concatenated oof predictions. I have not been able to conclude if this resolve the lack of CV & LB correlation issue. But so far, I have a [counter-example](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215389).  \n\n1 explanation for people's failures to establish good correlation is the poor quality of training labels ( many labels missing). Therefore, slight improvement or worsening of valid_loss and metric really don't mean that muhc ?? Along this line, I am also trying to find different metric ( other than the comp eval metric) that is more robust to imperfect label and thus better correlates CV and LB score.\n\nBTW, I have experimented with using different stride in inference. And I found that inferring on overlapping segments ( stride < period ) of audio produces slightly worse result ( on the exact same model) compared to non-overlapping segments. It seems to be pretty counter-intuitive to me. \n\nHope your exploration could lead to good CV scheme. Right now, this lack fo CV & LB correlation is driving me crazy. Every submission feels like a lottery to me",
      "votes": null
    },
    {
      "id": "1176830",
      "postDate": "01/29/2021 20:26:56",
      "content": "<p>Yeah there is some expected movement between validation and leaderboard. I think one thing that might lead to decreased variation between folds but increased variation against the leaderboard is some people are doing stratified folds. Doing this makes it so each training and validation set are balanced, but that assumption doesn't necessarily always hold true against a small test set so you training and validation are an ideal scenario while test can be anywhere on the spectrum of sampling</p>\n<p>The quality of labels definitely plays a factor but I presume the quality is similar between validation and test</p>",
      "rawMarkdown": "Yeah there is some expected movement between validation and leaderboard. I think one thing that might lead to decreased variation between folds but increased variation against the leaderboard is some people are doing stratified folds. Doing this makes it so each training and validation set are balanced, but that assumption doesn't necessarily always hold true against a small test set so you training and validation are an ideal scenario while test can be anywhere on the spectrum of sampling\n\nThe quality of labels definitely plays a factor but I presume the quality is similar between validation and test",
      "votes": null
    },
    {
      "id": "1176831",
      "postDate": "01/29/2021 20:28:18",
      "content": "<p>Interesting that you have found that relationship with the stride. I was thinking of playing with that. Wasn't sure if smaller strides meant more chances of getting the correct detection given multiple views of the same audio or greater chance of a false positive, probably a balance. </p>\n<p>Potentially fragile if using small segments and max pooling</p>",
      "rawMarkdown": "Interesting that you have found that relationship with the stride. I was thinking of playing with that. Wasn't sure if smaller strides meant more chances of getting the correct detection given multiple views of the same audio or greater chance of a false positive, probably a balance. \n\nPotentially fragile if using small segments and max pooling",
      "votes": null
    },
    {
      "id": "1177184",
      "postDate": "01/30/2021 05:52:53",
      "content": "<p><a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782#1101468\" target=\"_blank\">Discussion</a> here by the competition host seems imply the test labels are higher quality.</p>",
      "rawMarkdown": "[Discussion](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782#1101468) here by the competition host seems imply the test labels are higher quality.",
      "votes": null
    },
    {
      "id": "1182069",
      "postDate": "02/02/2021 09:36:35",
      "content": "<p>Thanks for sharing your great ideas!</p>\n<p>I think there are two things that make validation difficult:</p>\n<ol>\n<li>Incomplete labels as mentioned by others making the full image/recording LWAP estimation noisy</li>\n<li>I think (this is just a hunch) that most of the TP labels that have been given to us are exemplar or very clean (i.e. easy to predict). This means using the TP labels in their current form results in great validation performance but poor test performance on noisier labels. Perhaps adding noise/augmentation to the validation set might help?</li>\n</ol>",
      "rawMarkdown": "Thanks for sharing your great ideas!\n\nI think there are two things that make validation difficult:\n1. Incomplete labels as mentioned by others making the full image/recording LWAP estimation noisy\n2. I think (this is just a hunch) that most of the TP labels that have been given to us are exemplar or very clean (i.e. easy to predict). This means using the TP labels in their current form results in great validation performance but poor test performance on noisier labels. Perhaps adding noise/augmentation to the validation set might help?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1175443,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "01/29/2021 07:08:23",
      "content": "<p>Down that avenue of thinking, I wanted to close the gap between training and validation/test so I played with the idea of moving away from the varied cropping technique entirely. Instead of using a single crop and its label as a training sample I tried training on the fill collection of overlapping segments and the full label for the audio and then applied max pooling across the segments and then used BCE to train with that reduction. Surprisingly that seemed to perform pretty close to the cropping technique. </p>\n<p>I would think since the training samples were very similar every pass forward it would memorize the results pretty quickly, but I guess between specaugment and the other noise augmentations it managed to perform reasonably well. </p>\n<p>This is obviously much slower and more compute intensive because it is the full 1 minute of audio all at once when really only a small region is really worth learning about, but figured it was worth exploring and was surprised to find the gap between the two was not very large.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1175450,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "01/29/2021 07:14:18",
          "content": "<p>An additional thing I tried was applying a second loss to the model with slightly stronger labeling. So on top of training on the predictions from the max pooling of the segments I also trained the model on the segment level labels. The 60s of audio is split into 11 overlapping 10s pieces with 5s stride and I could associate labels with each segment. </p>\n<p>If there was no segment then I would zero out the loss so the model was not penalized in regions that were simply unlabeled. So then the model is trying not only to get the correct prediction at the macro level it is also trying to make sure that the segment where the audio actually occurred was the one that lit up the max pooling layer. </p>\n<p>This performed almost exactly on par as the single loss model, but converged much more quickly. Not sure if I can attribute this to it truly helping or just the doubling of the loss value with both of them active. </p>\n<p>This technique could be further extended to include the false positives as well because we know those regions of the audio are 0's so we can persuade our model to output zeros rather than just nullifying the loss. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1175743,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "01/29/2021 10:10:41",
          "content": "<p>Thanks. It's vary valuable to see how others approach problems. Truly appreciate it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1176455,
      "author_name": "garfieldchh",
      "author_url": "",
      "post_date": "01/29/2021 16:24:41",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> </p>\n<p>Thanks for nice write up on the validation topic.</p>\n<p>I also did a migration from random cropping to full clip validation ( exactly the same as discussed in your write-up). Although the LB and CV gap narrows as a result, i find that full clip validation score is still very noisy.  Training a single fold out of 5 fold, i found that slight improvement on lwlrap_framewise_aggregate score and validation loss have no correlation with the LB performance.  A slight improvement in val loss and lwlrap could lead to 1-3% drop in LB.</p>\n<p>Trying to establish a better correlation between local validation and LB, I am switch to relying on computing validation metric on the 5 fold concatenated oof predictions. I have not been able to conclude if this resolve the lack of CV &amp; LB correlation issue. But so far, I have a <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215389\" target=\"_blank\">counter-example</a>.  </p>\n<p>1 explanation for people's failures to establish good correlation is the poor quality of training labels ( many labels missing). Therefore, slight improvement or worsening of valid_loss and metric really don't mean that muhc ?? Along this line, I am also trying to find different metric ( other than the comp eval metric) that is more robust to imperfect label and thus better correlates CV and LB score.</p>\n<p>BTW, I have experimented with using different stride in inference. And I found that inferring on overlapping segments ( stride &lt; period ) of audio produces slightly worse result ( on the exact same model) compared to non-overlapping segments. It seems to be pretty counter-intuitive to me. </p>\n<p>Hope your exploration could lead to good CV scheme. Right now, this lack fo CV &amp; LB correlation is driving me crazy. Every submission feels like a lottery to me</p>",
      "votes": null,
      "replies": [
        {
          "id": 1176830,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "01/29/2021 20:26:56",
          "content": "<p>Yeah there is some expected movement between validation and leaderboard. I think one thing that might lead to decreased variation between folds but increased variation against the leaderboard is some people are doing stratified folds. Doing this makes it so each training and validation set are balanced, but that assumption doesn't necessarily always hold true against a small test set so you training and validation are an ideal scenario while test can be anywhere on the spectrum of sampling</p>\n<p>The quality of labels definitely plays a factor but I presume the quality is similar between validation and test</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1176831,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "01/29/2021 20:28:18",
          "content": "<p>Interesting that you have found that relationship with the stride. I was thinking of playing with that. Wasn't sure if smaller strides meant more chances of getting the correct detection given multiple views of the same audio or greater chance of a false positive, probably a balance. </p>\n<p>Potentially fragile if using small segments and max pooling</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1177184,
          "author_name": "garfieldchh",
          "author_url": "",
          "post_date": "01/30/2021 05:52:53",
          "content": "<p><a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782#1101468\" target=\"_blank\">Discussion</a> here by the competition host seems imply the test labels are higher quality.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1182069,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "02/02/2021 09:36:35",
      "content": "<p>Thanks for sharing your great ideas!</p>\n<p>I think there are two things that make validation difficult:</p>\n<ol>\n<li>Incomplete labels as mentioned by others making the full image/recording LWAP estimation noisy</li>\n<li>I think (this is just a hunch) that most of the TP labels that have been given to us are exemplar or very clean (i.e. easy to predict). This means using the TP labels in their current form results in great validation performance but poor test performance on noisier labels. Perhaps adding noise/augmentation to the validation set might help?</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1175436": "One of the first things I do in most competitions is check out what people have already done and what has seemed to be successful. I have noticed and wanted to point out a few things I have researched that might be useful to others and help mitigate some of the gap between train, validation and leaderboard.\n\nThe first thing I have seen is that on the validation step people are still using the cropping trick to focus on the regions that were labeled. This is problematic because the procedure is not matched on the test set, you cannot crop in on the regions of interest on the test set. You must make predictions for all of the audio. For this reason you must mimic that procedure during validation, predict on all 5 second chunks and then aggregate and evaluate the performance. How you perform on a crop right near the audio of interest is not a perfect analog to how your model will perform on multiple segments aggregated.\n\n![](https://i.imgur.com/ddNJ8jL.png)\n\nThe above is what validation currently looks like. A random crop is placed around where a label is found and then the model is evaluated just for that single region. This misses the model's ability to reject any misleading signal that may have occurred anywhere else in the audio. On top of that you can see variation with a static model, one that has not had any training applied to it will yield a different result. This is problematic if you are applying early stopping because you might see a model appear to have better performance when really it just got a lucky crop upon validation. An additional thing that this is missing is when an important sound event might be split across multiple splits on the test set. \n\n![](https://i.imgur.com/awvPync.png)\nAn alternate more appropriate validation scheme is to apply the same procedure that will be used on the test set. That is, use the same inference procedure and see how it performs given the full audio clip and using the aggregation method you settle. It is important to validate like this because a significant number of things can happen with those extra audio segments you are passing the model. Slightly terrifyingly this model was not only validating against a different task than it will see at inference time, it is mimicking the training procedure that is making the model blind to these undocumented regions of audio. Someone could theoretically have a model that achieves perfect performance on the cropped regions but shows poorly on the leaderboard because those untrained and unvalidated regions are causing problems with predictions. \n\nUsing the original method I see the kernel achieving ~.85 on validation but ~.74 on lb. Correcting this validation error alone makes it so validation is .8 and lb is .776. Close to within the variation I see between epochs and probably attributable to .8 being the peak number. Interestingly you would think without altering the training procedure at all there would be no score change, but actually selecting a better early stopping checkpoint against a more rigorous validation procedure can be the explanation for this.",
    "1175443": "Down that avenue of thinking, I wanted to close the gap between training and validation/test so I played with the idea of moving away from the varied cropping technique entirely. Instead of using a single crop and its label as a training sample I tried training on the fill collection of overlapping segments and the full label for the audio and then applied max pooling across the segments and then used BCE to train with that reduction. Surprisingly that seemed to perform pretty close to the cropping technique. \n\n I would think since the training samples were very similar every pass forward it would memorize the results pretty quickly, but I guess between specaugment and the other noise augmentations it managed to perform reasonably well. \n\nThis is obviously much slower and more compute intensive because it is the full 1 minute of audio all at once when really only a small region is really worth learning about, but figured it was worth exploring and was surprised to find the gap between the two was not very large.",
    "1175450": "An additional thing I tried was applying a second loss to the model with slightly stronger labeling. So on top of training on the predictions from the max pooling of the segments I also trained the model on the segment level labels. The 60s of audio is split into 11 overlapping 10s pieces with 5s stride and I could associate labels with each segment. \n\nIf there was no segment then I would zero out the loss so the model was not penalized in regions that were simply unlabeled. So then the model is trying not only to get the correct prediction at the macro level it is also trying to make sure that the segment where the audio actually occurred was the one that lit up the max pooling layer. \n\nThis performed almost exactly on par as the single loss model, but converged much more quickly. Not sure if I can attribute this to it truly helping or just the doubling of the loss value with both of them active. \n\nThis technique could be further extended to include the false positives as well because we know those regions of the audio are 0's so we can persuade our model to output zeros rather than just nullifying the loss.",
    "1175743": "Thanks. It's vary valuable to see how others approach problems. Truly appreciate it.",
    "1176455": "ryches \n\nThanks for nice write up on the validation topic.\n\nI also did a migration from random cropping to full clip validation ( exactly the same as discussed in your write-up). Although the LB and CV gap narrows as a result, i find that full clip validation score is still very noisy.  Training a single fold out of 5 fold, i found that slight improvement on lwlrap_framewise_aggregate score and validation loss have no correlation with the LB performance.  A slight improvement in val loss and lwlrap could lead to 1-3% drop in LB.\n\nTrying to establish a better correlation between local validation and LB, I am switch to relying on computing validation metric on the 5 fold concatenated oof predictions. I have not been able to conclude if this resolve the lack of CV & LB correlation issue. But so far, I have a [counter-example](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/215389).  \n\n1 explanation for people's failures to establish good correlation is the poor quality of training labels ( many labels missing). Therefore, slight improvement or worsening of valid_loss and metric really don't mean that muhc ?? Along this line, I am also trying to find different metric ( other than the comp eval metric) that is more robust to imperfect label and thus better correlates CV and LB score.\n\nBTW, I have experimented with using different stride in inference. And I found that inferring on overlapping segments ( stride < period ) of audio produces slightly worse result ( on the exact same model) compared to non-overlapping segments. It seems to be pretty counter-intuitive to me. \n\nHope your exploration could lead to good CV scheme. Right now, this lack fo CV & LB correlation is driving me crazy. Every submission feels like a lottery to me",
    "1176830": "Yeah there is some expected movement between validation and leaderboard. I think one thing that might lead to decreased variation between folds but increased variation against the leaderboard is some people are doing stratified folds. Doing this makes it so each training and validation set are balanced, but that assumption doesn't necessarily always hold true against a small test set so you training and validation are an ideal scenario while test can be anywhere on the spectrum of sampling\n\nThe quality of labels definitely plays a factor but I presume the quality is similar between validation and test",
    "1176831": "Interesting that you have found that relationship with the stride. I was thinking of playing with that. Wasn't sure if smaller strides meant more chances of getting the correct detection given multiple views of the same audio or greater chance of a false positive, probably a balance. \n\nPotentially fragile if using small segments and max pooling",
    "1177184": "[Discussion](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782#1101468) here by the competition host seems imply the test labels are higher quality.",
    "1182069": "Thanks for sharing your great ideas!\n\nI think there are two things that make validation difficult:\n1. Incomplete labels as mentioned by others making the full image/recording LWAP estimation noisy\n2. I think (this is just a hunch) that most of the TP labels that have been given to us are exemplar or very clean (i.e. easy to predict). This means using the TP labels in their current form results in great validation performance but poor test performance on noisier labels. Perhaps adding noise/augmentation to the validation set might help?"
  },
  "source": "meta"
}