{
  "id": 206551,
  "title": "Need help, big gap between local and LB",
  "url": "/competitions/rfcx-species-audio-detection/discussion/206551",
  "author_name": "",
  "post_date": "2020-12-25T07:09:23.922084200Z",
  "votes": 6,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I got single models 5 folds CV around 0.920 metric from here <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198418\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198418</a> while no subs got high than 0.836. acc around 0.87x. <br>\nAfter two weeks struggle this gap still exists. 😧 </p>",
  "messages": [
    {
      "id": "1125912",
      "postDate": "12/25/2020 07:09:23",
      "content": "<p>I got single models 5 folds CV around 0.920 metric from here <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198418\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198418</a> while no subs got high than 0.836. acc around 0.87x. <br>\nAfter two weeks struggle this gap still exists. 😧 </p>",
      "rawMarkdown": "I got single models 5 folds CV around 0.920 metric from here https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198418 while no subs got high than 0.836. acc around 0.87x. \nAfter two weeks struggle this gap still exists. 😧",
      "votes": null
    },
    {
      "id": "1126554",
      "postDate": "12/25/2020 17:34:31",
      "content": "<p>Need more details :)</p>\n<p>Are you early stopping on validation set?</p>\n<p>I am also from Chengdu.<br>\nMerry Christmas. 🎉</p>",
      "rawMarkdown": "Need more details :)\n\nAre you early stopping on validation set?\n\nI am also from Chengdu.\nMerry Christmas. 🎉",
      "votes": null
    },
    {
      "id": "1126568",
      "postDate": "12/25/2020 17:43:59",
      "content": "<p>No I’m training 40 epochs and saving best valid loss checkpoint. Thanks:) </p>",
      "rawMarkdown": "No I’m training 40 epochs and saving best valid loss checkpoint. Thanks:)",
      "votes": null
    },
    {
      "id": "1126657",
      "postDate": "12/25/2020 19:00:34",
      "content": "<p>I think what I say below is general common sense. Please don't feel offended if it seems too basic :)</p>\n<p>OK,  then, validation set is used to select model and compute the 0.920 CV metric. So, validation set is used to fit model in a general sense.</p>\n<p>I doubt it is the reason, but since the given positive samples for each species is very few, perhaps choosing a best validation loss checkpoint may have over fitting effect more severely than other cases where data is more abundant.</p>\n<p>If you fix the number of epochs for your 5 CV models, does it still look the same?</p>",
      "rawMarkdown": "I think what I say below is general common sense. Please don't feel offended if it seems too basic :)\n\nOK,  then, validation set is used to select model and compute the 0.920 CV metric. So, validation set is used to fit model in a general sense.\n\nI doubt it is the reason, but since the given positive samples for each species is very few, perhaps choosing a best validation loss checkpoint may have over fitting effect more severely than other cases where data is more abundant.\n\nIf you fix the number of epochs for your 5 CV models, does it still look the same?",
      "votes": null
    },
    {
      "id": "1126673",
      "postDate": "12/25/2020 19:47:56",
      "content": "<p>The learning rate is really small at last epochs, so fix epoch and choose the last epoch to predict won't make big difference. </p>",
      "rawMarkdown": "The learning rate is really small at last epochs, so fix epoch and choose the last epoch to predict won't make big difference.",
      "votes": null
    },
    {
      "id": "1126766",
      "postDate": "12/25/2020 23:00:19",
      "content": "<p>So, does your validation loss does ever go up ? Do you use validation set to determine when to shrink step size? Those are places validation-test difference comes in, right?</p>",
      "rawMarkdown": "So, does your validation loss does ever go up ? Do you use validation set to determine when to shrink step size? Those are places validation-test difference comes in, right?",
      "votes": null
    },
    {
      "id": "1126768",
      "postDate": "12/25/2020 23:04:18",
      "content": "<p>Hi MaChaogong,</p>\n<p>you can for example use this method (the class activation map method) to understand what the model is focusing on in order to make the prediction <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206393\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206393</a></p>\n<p>I think your model is memorizing the training data (overfitting), and you are not seeing the overfitting trend (high train acc and low validation acc) since the validation is <em>very</em> similar to the training set.</p>\n<p>Best,</p>\n<p>Guglielmo</p>",
      "rawMarkdown": "Hi MaChaogong,\n\nyou can for example use this method (the class activation map method) to understand what the model is focusing on in order to make the prediction [https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206393](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206393)\n\nI think your model is memorizing the training data (overfitting), and you are not seeing the overfitting trend (high train acc and low validation acc) since the validation is *very* similar to the training set.\n\nBest,\n\nGuglielmo",
      "votes": null
    },
    {
      "id": "1127874",
      "postDate": "12/27/2020 01:17:48",
      "content": "<p>Not sure how you're splitting your data but if you're training on n second chunks and not splitting train/val by file then there is a chance your train and val sets are overlapping for some samples. </p>",
      "rawMarkdown": "Not sure how you're splitting your data but if you're training on n second chunks and not splitting train/val by file then there is a chance your train and val sets are overlapping for some samples.",
      "votes": null
    },
    {
      "id": "1128065",
      "postDate": "12/27/2020 06:38:21",
      "content": "<p>thanks for your reply. yes I'm overfitting for sure while a less overfitting checkpoint has much worse LB performance. I got almost 100% acc for training set, while valid set is around 0.87x acc</p>",
      "rawMarkdown": "thanks for your reply. yes I'm overfitting for sure while a less overfitting checkpoint has much worse LB performance. I got almost 100% acc for training set, while valid set is around 0.87x acc",
      "votes": null
    },
    {
      "id": "1128068",
      "postDate": "12/27/2020 06:41:32",
      "content": "<p>just 5 folds. yeah I'm training on n seconds chunks by file. <br>\nmaybe training preprocess is not same as testing. I'm predicting multi fix seconds slides for test data, and take the max for each class. </p>",
      "rawMarkdown": "just 5 folds. yeah I'm training on n seconds chunks by file. \nmaybe training preprocess is not same as testing. I'm predicting multi fix seconds slides for test data, and take the max for each class.",
      "votes": null
    },
    {
      "id": "1128099",
      "postDate": "12/27/2020 07:08:00",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1538601%2F0a2ecfa7abcfe83ab63194e456b46453%2F(1).png?generation=1609052865978139&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1538601%2Ffb83ddc82edeed557bff62292d5d3a20%2F.png?generation=1609052875095521&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1538601%2F0a2ecfa7abcfe83ab63194e456b46453%2F(1).png?generation=1609052865978139&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1538601%2Ffb83ddc82edeed557bff62292d5d3a20%2F.png?generation=1609052875095521&alt=media)",
      "votes": null
    },
    {
      "id": "1128107",
      "postDate": "12/27/2020 07:15:37",
      "content": "<p>Maybe the wav length I'm using is too small. seeing from your image and mine</p>",
      "rawMarkdown": "Maybe the wav length I'm using is too small. seeing from your image and mine",
      "votes": null
    },
    {
      "id": "1135082",
      "postDate": "01/01/2021 21:43:18",
      "content": "<p>Are you only considering lwlrap at the chunk level (e.g. during training)? Or are you computing the metric at the file level after training is over?  For me the chunk level lwlrap tends to be higher and less reliable vs lb. </p>",
      "rawMarkdown": "Are you only considering lwlrap at the chunk level (e.g. during training)? Or are you computing the metric at the file level after training is over?  For me the chunk level lwlrap tends to be higher and less reliable vs lb.",
      "votes": null
    },
    {
      "id": "1135095",
      "postDate": "01/01/2021 22:11:00",
      "content": "<p>I’m considering lwlrap and the loss. Both from epoch level. The gap still exists while smaller now:)</p>",
      "rawMarkdown": "I’m considering lwlrap and the loss. Both from epoch level. The gap still exists while smaller now:)",
      "votes": null
    },
    {
      "id": "1139902",
      "postDate": "01/05/2021 17:31:30",
      "content": "<p>I found that I lose 5 points on public lb when I use mixup, even though my CV goes up a bit. I concluded that it might be because the labelling method is different for the training data and the test data.</p>",
      "rawMarkdown": "I found that I lose 5 points on public lb when I use mixup, even though my CV goes up a bit. I concluded that it might be because the labelling method is different for the training data and the test data.",
      "votes": null
    },
    {
      "id": "1140423",
      "postDate": "01/06/2021 02:13:19",
      "content": "<p>Hi Alexander,</p>\n<p>In my case, mixup helps a lot! I think it depends on:</p>\n<ul>\n<li>the implementation:<ul>\n<li>mixup all the elements of batch?</li>\n<li>use the same lambda for all the samples in the batch?</li></ul></li>\n<li>the \"strength\" of mixup:<ul>\n<li>the alpha parameter that select the lambda that combines the samples</li></ul></li>\n</ul>\n<p>Best,</p>\n<p>Guglielmo</p>",
      "rawMarkdown": "Hi Alexander,\n\nIn my case, mixup helps a lot! I think it depends on:\n\n- the implementation:\n  - mixup all the elements of batch?\n  - use the same lambda for all the samples in the batch?\n- the \"strength\" of mixup:\n  - the alpha parameter that select the lambda that combines the samples\n\nBest,\n\nGuglielmo",
      "votes": null
    },
    {
      "id": "1140431",
      "postDate": "01/06/2021 02:21:19",
      "content": "<p>Hi Alex, for your LB rank is not that good now. you shall first build a good single baseline model then try other tricks. </p>",
      "rawMarkdown": "Hi Alex, for your LB rank is not that good now. you shall first build a good single baseline model then try other tricks.",
      "votes": null
    },
    {
      "id": "1142659",
      "postDate": "01/07/2021 14:23:14",
      "content": "<p><a href=\"https://www.kaggle.com/guglielmo\" target=\"_blank\">@guglielmo</a> I took a slightly different approach. I overlaid spectra with pixel wise max and set labels to 1 for any TP present. CV improved</p>",
      "rawMarkdown": "guglielmo I took a slightly different approach. I overlaid spectra with pixel wise max and set labels to 1 for any TP present. CV improved",
      "votes": null
    },
    {
      "id": "1142997",
      "postDate": "01/07/2021 18:00:19",
      "content": "<p>1.yes. validation loss first go down then go up(for learning rate is small, the loss won't go up too high).<br>\n2.yes. I'm sometimes using valid loss to dicide when to shrink step size.<br>\n3.yes they are.<br>\nI will sub a fixed epochs model and a less overfitting  model to see how LB changes. Thanks for your nice suggestions. I'm struck and rethinking everything again:))</p>",
      "rawMarkdown": "1.yes. validation loss first go down then go up(for learning rate is small, the loss won't go up too high).\n2.yes. I'm sometimes using valid loss to dicide when to shrink step size.\n3.yes they are.\nI will sub a fixed epochs model and a less overfitting  model to see how LB changes. Thanks for your nice suggestions. I'm struck and rethinking everything again:))",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1126554,
      "author_name": "barnwellguy",
      "author_url": "",
      "post_date": "12/25/2020 17:34:31",
      "content": "<p>Need more details :)</p>\n<p>Are you early stopping on validation set?</p>\n<p>I am also from Chengdu.<br>\nMerry Christmas. 🎉</p>",
      "votes": null,
      "replies": [
        {
          "id": 1126568,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "12/25/2020 17:43:59",
          "content": "<p>No I’m training 40 epochs and saving best valid loss checkpoint. Thanks:) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126657,
          "author_name": "barnwellguy",
          "author_url": "",
          "post_date": "12/25/2020 19:00:34",
          "content": "<p>I think what I say below is general common sense. Please don't feel offended if it seems too basic :)</p>\n<p>OK,  then, validation set is used to select model and compute the 0.920 CV metric. So, validation set is used to fit model in a general sense.</p>\n<p>I doubt it is the reason, but since the given positive samples for each species is very few, perhaps choosing a best validation loss checkpoint may have over fitting effect more severely than other cases where data is more abundant.</p>\n<p>If you fix the number of epochs for your 5 CV models, does it still look the same?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126673,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "12/25/2020 19:47:56",
          "content": "<p>The learning rate is really small at last epochs, so fix epoch and choose the last epoch to predict won't make big difference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126766,
          "author_name": "barnwellguy",
          "author_url": "",
          "post_date": "12/25/2020 23:00:19",
          "content": "<p>So, does your validation loss does ever go up ? Do you use validation set to determine when to shrink step size? Those are places validation-test difference comes in, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142997,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "01/07/2021 18:00:19",
          "content": "<p>1.yes. validation loss first go down then go up(for learning rate is small, the loss won't go up too high).<br>\n2.yes. I'm sometimes using valid loss to dicide when to shrink step size.<br>\n3.yes they are.<br>\nI will sub a fixed epochs model and a less overfitting  model to see how LB changes. Thanks for your nice suggestions. I'm struck and rethinking everything again:))</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1126768,
      "author_name": "guglielmocamporese",
      "author_url": "",
      "post_date": "12/25/2020 23:04:18",
      "content": "<p>Hi MaChaogong,</p>\n<p>you can for example use this method (the class activation map method) to understand what the model is focusing on in order to make the prediction <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206393\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206393</a></p>\n<p>I think your model is memorizing the training data (overfitting), and you are not seeing the overfitting trend (high train acc and low validation acc) since the validation is <em>very</em> similar to the training set.</p>\n<p>Best,</p>\n<p>Guglielmo</p>",
      "votes": null,
      "replies": [
        {
          "id": 1128065,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "12/27/2020 06:38:21",
          "content": "<p>thanks for your reply. yes I'm overfitting for sure while a less overfitting checkpoint has much worse LB performance. I got almost 100% acc for training set, while valid set is around 0.87x acc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128099,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "12/27/2020 07:08:00",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1538601%2F0a2ecfa7abcfe83ab63194e456b46453%2F(1).png?generation=1609052865978139&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1538601%2Ffb83ddc82edeed557bff62292d5d3a20%2F.png?generation=1609052875095521&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128107,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "12/27/2020 07:15:37",
          "content": "<p>Maybe the wav length I'm using is too small. seeing from your image and mine</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1127874,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "12/27/2020 01:17:48",
      "content": "<p>Not sure how you're splitting your data but if you're training on n second chunks and not splitting train/val by file then there is a chance your train and val sets are overlapping for some samples. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1128068,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "12/27/2020 06:41:32",
          "content": "<p>just 5 folds. yeah I'm training on n seconds chunks by file. <br>\nmaybe training preprocess is not same as testing. I'm predicting multi fix seconds slides for test data, and take the max for each class. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1135082,
      "author_name": "reppic",
      "author_url": "",
      "post_date": "01/01/2021 21:43:18",
      "content": "<p>Are you only considering lwlrap at the chunk level (e.g. during training)? Or are you computing the metric at the file level after training is over?  For me the chunk level lwlrap tends to be higher and less reliable vs lb. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1135095,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "01/01/2021 22:11:00",
          "content": "<p>I’m considering lwlrap and the loss. Both from epoch level. The gap still exists while smaller now:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1139902,
      "author_name": "alexandersoare",
      "author_url": "",
      "post_date": "01/05/2021 17:31:30",
      "content": "<p>I found that I lose 5 points on public lb when I use mixup, even though my CV goes up a bit. I concluded that it might be because the labelling method is different for the training data and the test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1140423,
          "author_name": "guglielmocamporese",
          "author_url": "",
          "post_date": "01/06/2021 02:13:19",
          "content": "<p>Hi Alexander,</p>\n<p>In my case, mixup helps a lot! I think it depends on:</p>\n<ul>\n<li>the implementation:<ul>\n<li>mixup all the elements of batch?</li>\n<li>use the same lambda for all the samples in the batch?</li></ul></li>\n<li>the \"strength\" of mixup:<ul>\n<li>the alpha parameter that select the lambda that combines the samples</li></ul></li>\n</ul>\n<p>Best,</p>\n<p>Guglielmo</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140431,
          "author_name": "mcggood",
          "author_url": "",
          "post_date": "01/06/2021 02:21:19",
          "content": "<p>Hi Alex, for your LB rank is not that good now. you shall first build a good single baseline model then try other tricks. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142659,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "01/07/2021 14:23:14",
          "content": "<p><a href=\"https://www.kaggle.com/guglielmo\" target=\"_blank\">@guglielmo</a> I took a slightly different approach. I overlaid spectra with pixel wise max and set labels to 1 for any TP present. CV improved</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1125912": "I got single models 5 folds CV around 0.920 metric from here https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198418 while no subs got high than 0.836. acc around 0.87x. \nAfter two weeks struggle this gap still exists. 😧",
    "1126554": "Need more details :)\n\nAre you early stopping on validation set?\n\nI am also from Chengdu.\nMerry Christmas. 🎉",
    "1126568": "No I’m training 40 epochs and saving best valid loss checkpoint. Thanks:)",
    "1126657": "I think what I say below is general common sense. Please don't feel offended if it seems too basic :)\n\nOK,  then, validation set is used to select model and compute the 0.920 CV metric. So, validation set is used to fit model in a general sense.\n\nI doubt it is the reason, but since the given positive samples for each species is very few, perhaps choosing a best validation loss checkpoint may have over fitting effect more severely than other cases where data is more abundant.\n\nIf you fix the number of epochs for your 5 CV models, does it still look the same?",
    "1126673": "The learning rate is really small at last epochs, so fix epoch and choose the last epoch to predict won't make big difference.",
    "1126766": "So, does your validation loss does ever go up ? Do you use validation set to determine when to shrink step size? Those are places validation-test difference comes in, right?",
    "1126768": "Hi MaChaogong,\n\nyou can for example use this method (the class activation map method) to understand what the model is focusing on in order to make the prediction [https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206393](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/206393)\n\nI think your model is memorizing the training data (overfitting), and you are not seeing the overfitting trend (high train acc and low validation acc) since the validation is *very* similar to the training set.\n\nBest,\n\nGuglielmo",
    "1127874": "Not sure how you're splitting your data but if you're training on n second chunks and not splitting train/val by file then there is a chance your train and val sets are overlapping for some samples.",
    "1128065": "thanks for your reply. yes I'm overfitting for sure while a less overfitting checkpoint has much worse LB performance. I got almost 100% acc for training set, while valid set is around 0.87x acc",
    "1128068": "just 5 folds. yeah I'm training on n seconds chunks by file. \nmaybe training preprocess is not same as testing. I'm predicting multi fix seconds slides for test data, and take the max for each class.",
    "1128099": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1538601%2F0a2ecfa7abcfe83ab63194e456b46453%2F(1).png?generation=1609052865978139&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1538601%2Ffb83ddc82edeed557bff62292d5d3a20%2F.png?generation=1609052875095521&alt=media)",
    "1128107": "Maybe the wav length I'm using is too small. seeing from your image and mine",
    "1135082": "Are you only considering lwlrap at the chunk level (e.g. during training)? Or are you computing the metric at the file level after training is over?  For me the chunk level lwlrap tends to be higher and less reliable vs lb.",
    "1135095": "I’m considering lwlrap and the loss. Both from epoch level. The gap still exists while smaller now:)",
    "1139902": "I found that I lose 5 points on public lb when I use mixup, even though my CV goes up a bit. I concluded that it might be because the labelling method is different for the training data and the test data.",
    "1140423": "Hi Alexander,\n\nIn my case, mixup helps a lot! I think it depends on:\n\n- the implementation:\n  - mixup all the elements of batch?\n  - use the same lambda for all the samples in the batch?\n- the \"strength\" of mixup:\n  - the alpha parameter that select the lambda that combines the samples\n\nBest,\n\nGuglielmo",
    "1140431": "Hi Alex, for your LB rank is not that good now. you shall first build a good single baseline model then try other tricks.",
    "1142659": "guglielmo I took a slightly different approach. I overlaid spectra with pixel wise max and set labels to 1 for any TP present. CV improved",
    "1142997": "1.yes. validation loss first go down then go up(for learning rate is small, the loss won't go up too high).\n2.yes. I'm sometimes using valid loss to dicide when to shrink step size.\n3.yes they are.\nI will sub a fixed epochs model and a less overfitting  model to see how LB changes. Thanks for your nice suggestions. I'm struck and rethinking everything again:))"
  },
  "source": "meta"
}