{
  "id": 212906,
  "title": "How to Combine 5 fold CV model for this comp ?  I am getting worse result after combining folds",
  "url": "/competitions/rfcx-species-audio-detection/discussion/212906",
  "author_name": "NakedKoala",
  "post_date": "2021-01-20T17:05:56.442000",
  "votes": 7,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I am using SED method  discussed <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830\" target=\"_blank\">here</a>. My best 1 fold model achieves 0.80 on LB.</p>\n<p>I was expecting that combining 5 models from 5 folds would improve my result, but my 5 folds model performs worse than a single fold model. </p>\n<p>Here is what I did:</p>\n<p>Standard Stratified 5 fold Splits </p>\n<p>I trained each fold model for 30 epochs and save the model that achieves the highest LWLRAP on the validation set for that fold.   Then I  try to methods to combine the test prediction: 1) averaging 2) approach detailed <a href=\"https://www.kaggle.com/kneroma/rfcx-bagging\" target=\"_blank\">here</a>. </p>\n<h3>Questions:</h3>\n<ol>\n<li>Are there any obvious mistake in my approach ? </li>\n<li>For people who had success in improving score after combining 5 folds, what did you do differently ? </li>\n</ol>\n<p>Appreciate your input.  Thanks ! 🙏</p>",
  "messages": [
    {
      "id": 1161648,
      "postDate": "2021-01-20T17:05:56.443Z",
      "content": "<p>I am using SED method  discussed <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830\" target=\"_blank\">here</a>. My best 1 fold model achieves 0.80 on LB.</p>\n<p>I was expecting that combining 5 models from 5 folds would improve my result, but my 5 folds model performs worse than a single fold model. </p>\n<p>Here is what I did:</p>\n<p>Standard Stratified 5 fold Splits </p>\n<p>I trained each fold model for 30 epochs and save the model that achieves the highest LWLRAP on the validation set for that fold.   Then I  try to methods to combine the test prediction: 1) averaging 2) approach detailed <a href=\"https://www.kaggle.com/kneroma/rfcx-bagging\" target=\"_blank\">here</a>. </p>\n<h3>Questions:</h3>\n<ol>\n<li>Are there any obvious mistake in my approach ? </li>\n<li>For people who had success in improving score after combining 5 folds, what did you do differently ? </li>\n</ol>\n<p>Appreciate your input.  Thanks ! 🙏</p>",
      "rawMarkdown": "I am using SED method  discussed [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830). My best 1 fold model achieves 0.80 on LB.\n\nI was expecting that combining 5 models from 5 folds would improve my result, but my 5 folds model performs worse than a single fold model. \n\nHere is what I did:\n\nStandard Stratified 5 fold Splits \n\nI trained each fold model for 30 epochs and save the model that achieves the highest LWLRAP on the validation set for that fold.   Then I  try to methods to combine the test prediction: 1) averaging 2) approach detailed [here](https://www.kaggle.com/kneroma/rfcx-bagging). \n\n### Questions:\n\n1. Are there any obvious mistake in my approach ? \n2. For people who had success in improving score after combining 5 folds, what did you do differently ? \n\n\nAppreciate your input.  Thanks ! 🙏",
      "votes": 7
    },
    {
      "id": 1162198,
      "postDate": "2021-01-21T03:31:31.500Z",
      "content": "<ol>\n<li><p>I think you are right.</p></li>\n<li><p>I used EfficientNetB3 with BCE. And a single model's score was 0.86-0.89. Finally, these 5-folds average ensemble score was 0.913. </p></li>\n</ol>\n<blockquote>\n  <p>I trained each fold model for 30 epochs and save the model that achieves the highest LWLRAP on the validation set for that fold. </p>\n</blockquote>\n<p>Your local LWLRAP may not be correlation with LB. In my environment, it is more correlated with Loss value than LWLRAP.</p>",
      "rawMarkdown": "1. I think you are right.\n\n2. I used EfficientNetB3 with BCE. And a single model's score was 0.86-0.89. Finally, these 5-folds average ensemble score was 0.913. \n\n> I trained each fold model for 30 epochs and save the model that achieves the highest LWLRAP on the validation set for that fold. \n\nYour local LWLRAP may not be correlation with LB. In my environment, it is more correlated with Loss value than LWLRAP.",
      "votes": 1,
      "replies": [
        {
          "id": 1162441,
          "postDate": "2021-01-21T06:37:27.650Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a> <br>\nBTW,<br>\nWhat is your experience with using larger backbone ? </p>\n<p>One of your earlier post mentioned that you are doing EfficientNetB0.  Did B3 outperform B0 in your setup ?</p>\n<p>I saw another person mentioned he uses B5 ( same PANNs arch).</p>\n<p>I didn't try larger backbone because I saw other discussion about how larger model leads to worse result. But if larger model does help, then I will go ahead and try it out</p>",
          "rawMarkdown": "@shinmura0 \nBTW,\nWhat is your experience with using larger backbone ? \n\nOne of your earlier post mentioned that you are doing EfficientNetB0.  Did B3 outperform B0 in your setup ?\n\nI saw another person mentioned he uses B5 ( same PANNs arch).\n\nI didn't try larger backbone because I saw other discussion about how larger model leads to worse result. But if larger model does help, then I will go ahead and try it out"
        },
        {
          "id": 1162561,
          "postDate": "2021-01-21T08:15:26.857Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1162560,
          "postDate": "2021-01-21T08:15:26.857Z",
          "content": "<blockquote>\n  <p>Did B3 outperform B0 in your setup ?</p>\n</blockquote>\n<p>Yes. The following result is 5-folds average ensemble.</p>\n<ul>\n<li>EfficientNetB0:0.887</li>\n<li>EfficientNetB3:0.913</li>\n</ul>",
          "rawMarkdown": "> Did B3 outperform B0 in your setup ?\n\nYes. The following result is 5-folds average ensemble.\n+ EfficientNetB0:0.887\n+ EfficientNetB3:0.913\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1161678,
      "postDate": "2021-01-20T17:20:46.830Z",
      "content": "<ol>\n<li>I dont think that you are doing any mistake </li>\n<li>I didnt did anything special thing :) my sed 6 fold score of a single model is 0.86x in lb and combining with my other models i am able to achieve 0.89 . <br>\ni still think there is a lot of improvements needed for SED as my cv for SED is 0.89x  </li>\n</ol>",
      "rawMarkdown": "1.  I dont think that you are doing any mistake \n2. I didnt did anything special thing :) my sed 6 fold score of a single model is 0.86x in lb and combining with my other models i am able to achieve 0.89 . \ni still think there is a lot of improvements needed for SED as my cv for SED is 0.89x  \n ",
      "votes": 1,
      "replies": [
        {
          "id": 1161804,
          "postDate": "2021-01-20T18:54:58.843Z",
          "content": "<p><a href=\"https://www.kaggle.com/trooperog\" target=\"_blank\">@trooperog</a> <br>\nThanks for sharing your model score. </p>\n<p>Here is a summary my SED approach. Any hint on how may i improve toward your single model result ? </p>\n<p>Architecture = PANNs SED architecture<br>\nBackbone = EfficientNetB0 <br>\nAugmentation = [MixUp, RandomCrop of tp audio] <br>\nOptimizer = Adam with Cosine Annealing <br>\nData:  log mel spectrogram<br>\nLoss: BCE</p>",
          "rawMarkdown": "@trooperog \nThanks for sharing your model score. \n\nHere is a summary my SED approach. Any hint on how may i improve toward your single model result ? \n\nArchitecture = PANNs SED architecture\nBackbone = EfficientNetB0 \nAugmentation = [MixUp, RandomCrop of tp audio] \nOptimizer = Adam with Cosine Annealing \nData:  log mel spectrogram\nLoss: BCE\n"
        },
        {
          "id": 1161819,
          "postDate": "2021-01-20T19:07:11.447Z",
          "content": "<p>my SED arch is almost the same as yours , i did some changes in Panns official architecture and some more augmentations and a custom loss and also The backbone i used was efficientnet_b5_ns from timm github repo.</p>",
          "rawMarkdown": "my SED arch is almost the same as yours , i did some changes in Panns official architecture and some more augmentations and a custom loss and also The backbone i used was efficientnet_b5_ns from timm github repo.",
          "votes": 1
        },
        {
          "id": 1161869,
          "postDate": "2021-01-20T19:30:28.330Z",
          "content": "<p><a href=\"https://www.kaggle.com/trooperog\" target=\"_blank\">@trooperog</a> </p>\n<p>Any hint for where I should shop for custom loss ? Like what keyword to search in literature / medium. </p>\n<p>Does your loss specifically optimize for ranking objective ? </p>",
          "rawMarkdown": "@trooperog \n\nAny hint for where I should shop for custom loss ? Like what keyword to search in literature / medium. \n\nDoes your loss specifically optimize for ranking objective ? "
        },
        {
          "id": 1161884,
          "postDate": "2021-01-20T19:45:09.967Z",
          "content": "<p>in tensorflow there is a loss called SigmoidFocalCrossEntropy my custom loss is a pytorch implementation of this with some changes in its hyper parameters .</p>",
          "rawMarkdown": "in tensorflow there is a loss called SigmoidFocalCrossEntropy my custom loss is a pytorch implementation of this with some changes in its hyper parameters .",
          "votes": 2
        },
        {
          "id": 1161951,
          "postDate": "2021-01-20T20:58:16.930Z",
          "content": "<p>I will look into it.  Many thanks !</p>",
          "rawMarkdown": "I will look into it.  Many thanks !"
        }
      ]
    },
    {
      "id": 1162076,
      "postDate": "2021-01-21T00:36:11.163Z",
      "content": "<p>It's an interesting question here. </p>\n<p>Perhaps I just didn't get what other people are doing. As far as I see, it's pretty hard to get a proper local estimation of your model's LWLRAP for 60s audio. I don't have it at all.</p>\n<p>The given <code>train_tp</code> and <code>train_fp</code> only contains some information of the training audios, namely, in specific windows there is or is not a particular species. It doesn't say anything about other species in the specified windows, nor anything outside of the windows. So, we don't have accurate full set of species that exists in one 60s audio, which is necessary to compute a trustworthy LWLRAP.</p>\n<p>What do you think?</p>",
      "rawMarkdown": "It's an interesting question here. \n\nPerhaps I just didn't get what other people are doing. As far as I see, it's pretty hard to get a proper local estimation of your model's LWLRAP for 60s audio. I don't have it at all.\n\nThe given `train_tp` and `train_fp` only contains some information of the training audios, namely, in specific windows there is or is not a particular species. It doesn't say anything about other species in the specified windows, nor anything outside of the windows. So, we don't have accurate full set of species that exists in one 60s audio, which is necessary to compute a trustworthy LWLRAP.\n\nWhat do you think?",
      "replies": [
        {
          "id": 1162086,
          "postDate": "2021-01-21T01:17:29.050Z",
          "content": "<p>You are saying that the tp labels do not provide sufficient info for us to compute clipwise_lwrap. Therefore, my 5 fold best model based on clipwise_lwrap are probably sub-optimal / random ?</p>\n<p>What kind of validation scheme are you doing ?   I used to validate on random 10 secs crop of audio that contains the annotated event. But I am getting overly optimistic performance estimate( val lwrap = 0.87 and LB = 0.80).  After I did clipwise_lwrap, the over estimation issue is much reduced.</p>",
          "rawMarkdown": "You are saying that the tp labels do not provide sufficient info for us to compute clipwise_lwrap. Therefore, my 5 fold best model based on clipwise_lwrap are probably sub-optimal / random ?\n\nWhat kind of validation scheme are you doing ?   I used to validate on random 10 secs crop of audio that contains the annotated event. But I am getting overly optimistic performance estimate( val lwrap = 0.87 and LB = 0.80).  After I did clipwise_lwrap, the over estimation issue is much reduced.\n\n",
          "votes": 1
        },
        {
          "id": 1162094,
          "postDate": "2021-01-21T01:43:04.023Z",
          "content": "<blockquote>\n  <p>`You are saying that the tp labels do not provide sufficient info for us to compute clipwise_lwrap. Therefore, my 5 fold best model based on clipwise_lwrap are probably sub-optimal / random ?</p>\n</blockquote>\n<p>Yes, I think the raw <code>train_tp</code> does not have enough information for either 60s audios or the small windows it points to. In some sense, I didn't validate, I just train for a fixed number of epochs.</p>\n<p>Perhaps one-hot type of training labels made simply from <code>train_tp</code> is also an inaccurate loss. Bird calls together, but we have only one non-zero label for each sample.</p>",
          "rawMarkdown": "> `You are saying that the tp labels do not provide sufficient info for us to compute clipwise_lwrap. Therefore, my 5 fold best model based on clipwise_lwrap are probably sub-optimal / random ?\n\nYes, I think the raw `train_tp` does not have enough information for either 60s audios or the small windows it points to. In some sense, I didn't validate, I just train for a fixed number of epochs.\n\nPerhaps one-hot type of training labels made simply from `train_tp` is also an inaccurate loss. Bird calls together, but we have only one non-zero label for each sample."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1162198,
      "author_name": "shinmura0",
      "author_url": "",
      "post_date": "2021-01-21T03:31:31.500000",
      "content": "<ol>\n<li><p>I think you are right.</p></li>\n<li><p>I used EfficientNetB3 with BCE. And a single model's score was 0.86-0.89. Finally, these 5-folds average ensemble score was 0.913. </p></li>\n</ol>\n<blockquote>\n  <p>I trained each fold model for 30 epochs and save the model that achieves the highest LWLRAP on the validation set for that fold. </p>\n</blockquote>\n<p>Your local LWLRAP may not be correlation with LB. In my environment, it is more correlated with Loss value than LWLRAP.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1162441,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-21T06:37:27.650000",
          "content": "<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a> <br>\nBTW,<br>\nWhat is your experience with using larger backbone ? </p>\n<p>One of your earlier post mentioned that you are doing EfficientNetB0.  Did B3 outperform B0 in your setup ?</p>\n<p>I saw another person mentioned he uses B5 ( same PANNs arch).</p>\n<p>I didn't try larger backbone because I saw other discussion about how larger model leads to worse result. But if larger model does help, then I will go ahead and try it out</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1162561,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-21T08:15:26.857000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1162560,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2021-01-21T08:15:26.857000",
          "content": "<blockquote>\n  <p>Did B3 outperform B0 in your setup ?</p>\n</blockquote>\n<p>Yes. The following result is 5-folds average ensemble.</p>\n<ul>\n<li>EfficientNetB0:0.887</li>\n<li>EfficientNetB3:0.913</li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1161678,
      "author_name": "Shubham Thapa",
      "author_url": "",
      "post_date": "2021-01-20T17:20:46.830000",
      "content": "<ol>\n<li>I dont think that you are doing any mistake </li>\n<li>I didnt did anything special thing :) my sed 6 fold score of a single model is 0.86x in lb and combining with my other models i am able to achieve 0.89 . <br>\ni still think there is a lot of improvements needed for SED as my cv for SED is 0.89x  </li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 1161804,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-20T18:54:58.843000",
          "content": "<p><a href=\"https://www.kaggle.com/trooperog\" target=\"_blank\">@trooperog</a> <br>\nThanks for sharing your model score. </p>\n<p>Here is a summary my SED approach. Any hint on how may i improve toward your single model result ? </p>\n<p>Architecture = PANNs SED architecture<br>\nBackbone = EfficientNetB0 <br>\nAugmentation = [MixUp, RandomCrop of tp audio] <br>\nOptimizer = Adam with Cosine Annealing <br>\nData:  log mel spectrogram<br>\nLoss: BCE</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161819,
          "author_name": "Shubham Thapa",
          "author_url": "",
          "post_date": "2021-01-20T19:07:11.447000",
          "content": "<p>my SED arch is almost the same as yours , i did some changes in Panns official architecture and some more augmentations and a custom loss and also The backbone i used was efficientnet_b5_ns from timm github repo.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1161869,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-20T19:30:28.330000",
          "content": "<p><a href=\"https://www.kaggle.com/trooperog\" target=\"_blank\">@trooperog</a> </p>\n<p>Any hint for where I should shop for custom loss ? Like what keyword to search in literature / medium. </p>\n<p>Does your loss specifically optimize for ranking objective ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1161884,
          "author_name": "Shubham Thapa",
          "author_url": "",
          "post_date": "2021-01-20T19:45:09.967000",
          "content": "<p>in tensorflow there is a loss called SigmoidFocalCrossEntropy my custom loss is a pytorch implementation of this with some changes in its hyper parameters .</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1161951,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-20T20:58:16.930000",
          "content": "<p>I will look into it.  Many thanks !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1162076,
      "author_name": "Buffalo Spdwy",
      "author_url": "",
      "post_date": "2021-01-21T00:36:11.163000",
      "content": "<p>It's an interesting question here. </p>\n<p>Perhaps I just didn't get what other people are doing. As far as I see, it's pretty hard to get a proper local estimation of your model's LWLRAP for 60s audio. I don't have it at all.</p>\n<p>The given <code>train_tp</code> and <code>train_fp</code> only contains some information of the training audios, namely, in specific windows there is or is not a particular species. It doesn't say anything about other species in the specified windows, nor anything outside of the windows. So, we don't have accurate full set of species that exists in one 60s audio, which is necessary to compute a trustworthy LWLRAP.</p>\n<p>What do you think?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1162086,
          "author_name": "NakedKoala",
          "author_url": "",
          "post_date": "2021-01-21T01:17:29.050000",
          "content": "<p>You are saying that the tp labels do not provide sufficient info for us to compute clipwise_lwrap. Therefore, my 5 fold best model based on clipwise_lwrap are probably sub-optimal / random ?</p>\n<p>What kind of validation scheme are you doing ?   I used to validate on random 10 secs crop of audio that contains the annotated event. But I am getting overly optimistic performance estimate( val lwrap = 0.87 and LB = 0.80).  After I did clipwise_lwrap, the over estimation issue is much reduced.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1162094,
          "author_name": "Buffalo Spdwy",
          "author_url": "",
          "post_date": "2021-01-21T01:43:04.023000",
          "content": "<blockquote>\n  <p>`You are saying that the tp labels do not provide sufficient info for us to compute clipwise_lwrap. Therefore, my 5 fold best model based on clipwise_lwrap are probably sub-optimal / random ?</p>\n</blockquote>\n<p>Yes, I think the raw <code>train_tp</code> does not have enough information for either 60s audios or the small windows it points to. In some sense, I didn't validate, I just train for a fixed number of epochs.</p>\n<p>Perhaps one-hot type of training labels made simply from <code>train_tp</code> is also an inaccurate loss. Bird calls together, but we have only one non-zero label for each sample.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1161648": "I am using SED method  discussed [here](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830). My best 1 fold model achieves 0.80 on LB.\n\nI was expecting that combining 5 models from 5 folds would improve my result, but my 5 folds model performs worse than a single fold model. \n\nHere is what I did:\n\nStandard Stratified 5 fold Splits \n\nI trained each fold model for 30 epochs and save the model that achieves the highest LWLRAP on the validation set for that fold.   Then I  try to methods to combine the test prediction: 1) averaging 2) approach detailed [here](https://www.kaggle.com/kneroma/rfcx-bagging). \n\n### Questions:\n\n1. Are there any obvious mistake in my approach ? \n2. For people who had success in improving score after combining 5 folds, what did you do differently ? \n\n\nAppreciate your input.  Thanks ! 🙏",
    "1162198": "1. I think you are right.\n\n2. I used EfficientNetB3 with BCE. And a single model's score was 0.86-0.89. Finally, these 5-folds average ensemble score was 0.913. \n\n> I trained each fold model for 30 epochs and save the model that achieves the highest LWLRAP on the validation set for that fold. \n\nYour local LWLRAP may not be correlation with LB. In my environment, it is more correlated with Loss value than LWLRAP.",
    "1161678": "1.  I dont think that you are doing any mistake \n2. I didnt did anything special thing :) my sed 6 fold score of a single model is 0.86x in lb and combining with my other models i am able to achieve 0.89 . \ni still think there is a lot of improvements needed for SED as my cv for SED is 0.89x  \n ",
    "1162076": "It's an interesting question here. \n\nPerhaps I just didn't get what other people are doing. As far as I see, it's pretty hard to get a proper local estimation of your model's LWLRAP for 60s audio. I don't have it at all.\n\nThe given `train_tp` and `train_fp` only contains some information of the training audios, namely, in specific windows there is or is not a particular species. It doesn't say anything about other species in the specified windows, nor anything outside of the windows. So, we don't have accurate full set of species that exists in one 60s audio, which is necessary to compute a trustworthy LWLRAP.\n\nWhat do you think?"
  }
}