{
  "id": 313286,
  "title": "[CV vs LB]",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/313286",
  "author_name": "",
  "post_date": "2022-03-16T11:13:45.488857200Z",
  "votes": 11,
  "comment_count": 32,
  "views": 0,
  "content": "<p>lets start the ancient kaggle thread!</p>\n<p>Model : effnet_b1 (no pretrained weights, trained from scratch)<br>\nFolds: 5<br>\nCV: 0.39<br>\nLB: 0.51<br>\n(Stratified k folds based on genre)<br>\n(used melspectrograms)</p>",
  "messages": [
    {
      "id": "1724592",
      "postDate": "03/16/2022 11:13:45",
      "content": "<p>lets start the ancient kaggle thread!</p>\n<p>Model : effnet_b1 (no pretrained weights, trained from scratch)<br>\nFolds: 5<br>\nCV: 0.39<br>\nLB: 0.51<br>\n(Stratified k folds based on genre)<br>\n(used melspectrograms)</p>",
      "rawMarkdown": "lets start the ancient kaggle thread!\n\nModel : effnet_b1 (no pretrained weights, trained from scratch)\nFolds: 5\nCV: 0.39\nLB: 0.51\n(Stratified k folds based on genre)\n(used melspectrograms)",
      "votes": null
    },
    {
      "id": "1724594",
      "postDate": "03/16/2022 11:15:13",
      "content": "<p>Quite a large gap between CV/LB, anybody have the idea why?</p>",
      "rawMarkdown": "Quite a large gap between CV/LB, anybody have the idea why?",
      "votes": null
    },
    {
      "id": "1724784",
      "postDate": "03/16/2022 14:04:44",
      "content": "<p>The metric is F1 but I remember seeing somewhere that it is 'macro' or 'micro'?</p>",
      "rawMarkdown": "The metric is F1 but I remember seeing somewhere that it is 'macro' or 'micro'?",
      "votes": null
    },
    {
      "id": "1724790",
      "postDate": "03/16/2022 14:11:38",
      "content": "<p>My score is very close to the accuracy and not f1 :/</p>",
      "rawMarkdown": "My score is very close to the accuracy and not f1 :/",
      "votes": null
    },
    {
      "id": "1724805",
      "postDate": "03/16/2022 14:24:57",
      "content": "<p><a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> - it should be <code>micro</code> for the first day of the competition I had it as <code>macro</code> but then switched.</p>",
      "rawMarkdown": "dienhoa - it should be `micro` for the first day of the competition I had it as `macro` but then switched.",
      "votes": null
    },
    {
      "id": "1725321",
      "postDate": "03/17/2022 02:03:16",
      "content": "<p>Ohh I think thats why theres the gap,I have been measuring the macro F1 in my local setup</p>",
      "rawMarkdown": "Ohh I think thats why theres the gap,I have been measuring the macro F1 in my local setup",
      "votes": null
    },
    {
      "id": "1729182",
      "postDate": "03/19/2022 18:10:31",
      "content": "<p>In my experiment, the F1 micro score is always equal to accuracy, and I think it's true after reading this: <a href=\"https://stackoverflow.com/questions/37358496/is-f1-micro-the-same-as-accuracy#:~:text=4%20Answers&amp;text=is%20not%20useful-,Show%20activity%20on%20this%20post.,case%20in%20multi-label%20classification\" target=\"_blank\">https://stackoverflow.com/questions/37358496/is-f1-micro-the-same-as-accuracy#:~:text=4%20Answers&amp;text=is%20not%20useful-,Show%20activity%20on%20this%20post.,case%20in%20multi-label%20classification</a>.</p>\n<p>Sometimes I have a big gap between CV and LB (CV~0.55 but LB~0.50). Did anyone experience the same?</p>\n<p>Thanks</p>",
      "rawMarkdown": "In my experiment, the F1 micro score is always equal to accuracy, and I think it's true after reading this: https://stackoverflow.com/questions/37358496/is-f1-micro-the-same-as-accuracy#:~:text=4%20Answers&text=is%20not%20useful-,Show%20activity%20on%20this%20post.,case%20in%20multi-label%20classification.\n\nSometimes I have a big gap between CV and LB (CV~0.55 but LB~0.50). Did anyone experience the same?\n\nThanks",
      "votes": null
    },
    {
      "id": "1729453",
      "postDate": "03/20/2022 05:16:23",
      "content": "<p>How are you calculating CV? When I take k fold mode then cv lb gap is less, however, when I take avg probs and then argmax, cv lb gap is significant. </p>",
      "rawMarkdown": "How are you calculating CV? When I take k fold mode then cv lb gap is less, however, when I take avg probs and then argmax, cv lb gap is significant.",
      "votes": null
    },
    {
      "id": "1729523",
      "postDate": "03/20/2022 07:02:42",
      "content": "<p>Because my model takes quite a long time to learn so now I am just based on one fold :D, I will try to average 5 folds to see what happens.</p>\n<p>If I remember correctly, LB has 2500 samples so do you think it's big enough to generalize the result?</p>\n<p>Maybe my model is not generalized enough :/  </p>",
      "rawMarkdown": "Because my model takes quite a long time to learn so now I am just based on one fold :D, I will try to average 5 folds to see what happens.\n\nIf I remember correctly, LB has 2500 samples so do you think it's big enough to generalize the result?\n\nMaybe my model is not generalized enough :/",
      "votes": null
    },
    {
      "id": "1729530",
      "postDate": "03/20/2022 07:16:58",
      "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> I've just submitted an average of 5 folds with each fold having a score &gt; 0.56 and my LB is 0.53093 :D. I have no idea why </p>",
      "rawMarkdown": "pheadrus I've just submitted an average of 5 folds with each fold having a score > 0.56 and my LB is 0.53093 :D. I have no idea why",
      "votes": null
    },
    {
      "id": "1729533",
      "postDate": "03/20/2022 07:32:13",
      "content": "<p>Yes, I think averaging probs across folds might not be a very good idea, as the class thresholds might change for different models. Did you try generating class labels for each fold and then np.mode?   </p>",
      "rawMarkdown": "Yes, I think averaging probs across folds might not be a very good idea, as the class thresholds might change for different models. Did you try generating class labels for each fold and then np.mode?",
      "votes": null
    },
    {
      "id": "1729537",
      "postDate": "03/20/2022 07:35:25",
      "content": "<p>I think someone should do an adv validation to see if the train and test come from same distributions or not. If they don't come from same dist then there are implications about how much you can trust local CV.</p>",
      "rawMarkdown": "I think someone should do an adv validation to see if the train and test come from same distributions or not. If they don't come from same dist then there are implications about how much you can trust local CV.",
      "votes": null
    },
    {
      "id": "1729885",
      "postDate": "03/20/2022 16:53:10",
      "content": "<p>I witnessed the same and I thought there is some bug in my code :D Then I saw the confusion matrix of every fold and understood what's wrong. </p>\n<p>Agree with <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> hard voting maybe a better idea for this problem.</p>",
      "rawMarkdown": "I witnessed the same and I thought there is some bug in my code :D Then I saw the confusion matrix of every fold and understood what's wrong. \n\nAgree with @pheadrus hard voting maybe a better idea for this problem.",
      "votes": null
    },
    {
      "id": "1729896",
      "postDate": "03/20/2022 17:05:44",
      "content": "<p>Thanks a lot, <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> , <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> . I have tried hard voting which I think makes more sense that every model is treated equally. Unfortunately, the result is not better.</p>\n<p>Do you think Oversampling can help? How can I check if the test set and training set come from the same distribution? Maby comparing the histogram of the validation prediction and the test prediction? </p>",
      "rawMarkdown": "Thanks a lot, @pheadrus , @harveenchadha . I have tried hard voting which I think makes more sense that every model is treated equally. Unfortunately, the result is not better.\n\nDo you think Oversampling can help? How can I check if the test set and training set come from the same distribution? Maby comparing the histogram of the validation prediction and the test prediction?",
      "votes": null
    },
    {
      "id": "1729920",
      "postDate": "03/20/2022 17:27:08",
      "content": "<p>I will try oversampling with audio augmentations in my next experiments. Something like this.</p>\n<pre><code>def get_transforms(*, data):\n    if data == 'train':\n        return tA.Compose(\n                transforms=[\n                     tA.ShuffleChannels(p=0.1,mode=\"per_example\",p_mode=\"per_example\", sample_rate= 16000),\n                    tA.PeakNormalization(p=0.5,apply_to='only_too_loud_sounds'),\n                     tA.AddColoredNoise(p=0.1,mode=\"per_channel\",p_mode=\"per_channel\", sample_rate=16000,max_snr_in_db = 15),\n                     tA.Shift(p=0.1,mode=\"per_example\",p_mode=\"per_example\", sample_rate=16000,\n                              max_shift=0.025, min_shift=-0.025),\n                ])\n\n    elif data == 'valid':\n        return tA.Compose([\n        ])\n</code></pre>",
      "rawMarkdown": "I will try oversampling with audio augmentations in my next experiments. Something like this.\n\n\n```\ndef get_transforms(*, data):\n    if data == 'train':\n        return tA.Compose(\n                transforms=[\n                     tA.ShuffleChannels(p=0.1,mode=\"per_example\",p_mode=\"per_example\", sample_rate= 16000),\n                    tA.PeakNormalization(p=0.5,apply_to='only_too_loud_sounds'),\n                     tA.AddColoredNoise(p=0.1,mode=\"per_channel\",p_mode=\"per_channel\", sample_rate=16000,max_snr_in_db = 15),\n                     tA.Shift(p=0.1,mode=\"per_example\",p_mode=\"per_example\", sample_rate=16000,\n                              max_shift=0.025, min_shift=-0.025),\n                ])\n\n    elif data == 'valid':\n        return tA.Compose([\n        ])\n\n```",
      "votes": null
    },
    {
      "id": "1729922",
      "postDate": "03/20/2022 17:29:42",
      "content": "<p>To check if train, test come from same distribution. Make new labels 1 for train, 0 for test. And do binary classification. If AUC ~ 0.5 then they both come from same distribution, else, no. No need to split data in kfolds, use entire datasets. </p>",
      "rawMarkdown": "To check if train, test come from same distribution. Make new labels 1 for train, 0 for test. And do binary classification. If AUC ~ 0.5 then they both come from same distribution, else, no. No need to split data in kfolds, use entire datasets.",
      "votes": null
    },
    {
      "id": "1738972",
      "postDate": "03/29/2022 16:45:43",
      "content": "<p>CV 0.5635 (stratified 5 folds) - LB 0.56766</p>",
      "rawMarkdown": "CV 0.5635 (stratified 5 folds) - LB 0.56766",
      "votes": null
    },
    {
      "id": "1739225",
      "postDate": "03/29/2022 20:54:35",
      "content": "<p>That's impressive !! I'm still struggling that my CV and LB is not consistent :/ </p>",
      "rawMarkdown": "That's impressive !! I'm still struggling that my CV and LB is not consistent :/",
      "votes": null
    },
    {
      "id": "1739464",
      "postDate": "03/30/2022 04:24:49",
      "content": "<p><a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> how much difference are you seeing in CV vs LB? I am seeing consistent gap of ~2% (5 fold mode ensemble LB vs OOF CV)</p>",
      "rawMarkdown": "dienhoa how much difference are you seeing in CV vs LB? I am seeing consistent gap of ~2% (5 fold mode ensemble LB vs OOF CV)",
      "votes": null
    },
    {
      "id": "1739674",
      "postDate": "03/30/2022 08:45:22",
      "content": "<p>Maybe it is because of myself that one fold takes a lot of time to train (~2hours/fold) so I just push the results of 1 -&gt; 3 folds, The gap usually is &gt; 3%. </p>\n<p>My latest submission is a lot better that I don't have a big gap ( I'm really happy that I have improved after 2 weeks ! ) I hope it is not just pure luck :D</p>\n<p>How much time does it take for your to train the whole 5 epochs? Do you have any strategy to experiment fastly different ideas (Train with less genre, train with less data, …) ?</p>\n<p>Thanks </p>",
      "rawMarkdown": "Maybe it is because of myself that one fold takes a lot of time to train (~2hours/fold) so I just push the results of 1 -> 3 folds, The gap usually is > 3%. \n\nMy latest submission is a lot better that I don't have a big gap ( I'm really happy that I have improved after 2 weeks ! ) I hope it is not just pure luck :D\n\nHow much time does it take for your to train the whole 5 epochs? Do you have any strategy to experiment fastly different ideas (Train with less genre, train with less data, ...) ?\n\nThanks",
      "votes": null
    },
    {
      "id": "1739680",
      "postDate": "03/30/2022 08:52:33",
      "content": "<p>It takes me less than 30 minutes per fold.</p>\n<p>I use precomputed spectrograms though, the cost is that I can't do augmentations on waveforms but it's much faster than loading the waveforms and then computing the spectrograms on the fly :)</p>",
      "rawMarkdown": "It takes me less than 30 minutes per fold.\n\nI use precomputed spectrograms though, the cost is that I can't do augmentations on waveforms but it's much faster than loading the waveforms and then computing the spectrograms on the fly :)",
      "votes": null
    },
    {
      "id": "1739710",
      "postDate": "03/30/2022 09:25:29",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> . I have reduced lots of time loading the waveform by using the .wav instead of .ogg, but definitely calculating the spectrogram takes a lot of time !</p>",
      "rawMarkdown": "Thanks a lot @theoviel . I have reduced lots of time loading the waveform by using the .wav instead of .ogg, but definitely calculating the spectrogram takes a lot of time !",
      "votes": null
    },
    {
      "id": "1739733",
      "postDate": "03/30/2022 09:51:05",
      "content": "<p>For me each epoch is 10-15 mins, so around 1.5 hours for each fold (7 epochs). <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> are you experimenting with spec augs such as mixup? I experimented a little bit with raw audio augs, but the results were not encouraging, maybe my aug were not appropriate. </p>\n<pre><code>tA.Shift(p=0.5,mode=\"per_example\",p_mode=\"per_example\", sample_rate=16000,\n                              max_shift=0.025, min_shift=-0.025)\n</code></pre>",
      "rawMarkdown": "For me each epoch is 10-15 mins, so around 1.5 hours for each fold (7 epochs). @theoviel are you experimenting with spec augs such as mixup? I experimented a little bit with raw audio augs, but the results were not encouraging, maybe my aug were not appropriate. \n\n```\ntA.Shift(p=0.5,mode=\"per_example\",p_mode=\"per_example\", sample_rate=16000,\n                              max_shift=0.025, min_shift=-0.025)\n```",
      "votes": null
    },
    {
      "id": "1739774",
      "postDate": "03/30/2022 10:18:18",
      "content": "<p>I tried specaugment (time masking &amp; frequency masking), but it didn't really help.<br>\nMixup is on my todo :) </p>",
      "rawMarkdown": "I tried specaugment (time masking & frequency masking), but it didn't really help.\nMixup is on my todo :)",
      "votes": null
    },
    {
      "id": "1739802",
      "postDate": "03/30/2022 10:49:26",
      "content": "<p>I used mixup and it really helps, without it, the model will easily overfit, with mixup I can train for a long time and the training and validation score converge </p>",
      "rawMarkdown": "I used mixup and it really helps, without it, the model will easily overfit, with mixup I can train for a long time and the training and validation score converge",
      "votes": null
    },
    {
      "id": "1739809",
      "postDate": "03/30/2022 11:00:08",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> </p>",
      "rawMarkdown": "Thanks @dienhoa",
      "votes": null
    },
    {
      "id": "1743562",
      "postDate": "04/03/2022 05:18:17",
      "content": "<p>Densenet121D CV 55.01 LB 58. Don't know whats going on.</p>",
      "rawMarkdown": "Densenet121D CV 55.01 LB 58. Don't know whats going on.",
      "votes": null
    },
    {
      "id": "1743616",
      "postDate": "04/03/2022 06:31:12",
      "content": "<p>Maybe we can ask <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> does he still has the consistent LB vs CV?</p>",
      "rawMarkdown": "Maybe we can ask @theoviel does he still has the consistent LB vs CV?",
      "votes": null
    },
    {
      "id": "1743670",
      "postDate": "04/03/2022 07:40:45",
      "content": "<p>Yes please. <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> </p>",
      "rawMarkdown": "Yes please. @theoviel",
      "votes": null
    },
    {
      "id": "1743767",
      "postDate": "04/03/2022 09:33:27",
      "content": "<p>CV vs LB is not consistent at all for me either :(</p>\n<p>I think my first submissions were just lucky on that matter.</p>",
      "rawMarkdown": "CV vs LB is not consistent at all for me either :(\n\nI think my first submissions were just lucky on that matter.",
      "votes": null
    },
    {
      "id": "1743895",
      "postDate": "04/03/2022 11:59:59",
      "content": "<p>Makes me think if <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> has done some kind of sorcery with test dataset. 👀👀</p>",
      "rawMarkdown": "Makes me think if @robikscube has done some kind of sorcery with test dataset. 👀👀",
      "votes": null
    },
    {
      "id": "1744769",
      "postDate": "04/04/2022 09:44:35",
      "content": "<p>In my last version, I reach a CV ~ 0.605 and has a LB ~ 0.55. Do you all usually have a big gap like this (which is always the case for me).</p>\n<p>My approach is using a CNN model on spectrogram of ~5 seconds of music, not many data augmentations at all, just cropping the 5s chunk at difference places, changing the brightness of the spectrogram. I was thinking that may be 5s chunk is not big enough to generalize the classification and have tried 15s chunk but not succeeded.</p>\n<p>I think maybe the 30s of clip in the training set is sometimes from the same audios so that we might overfit the training set but not well generalize at all.</p>\n<p>Or maybe CNN spectrogram is not good at generalizing audio classification. Do you think Time Series Classification can work better?</p>\n<p>I'm running out of ideas :)) Can someone shed me some light ?</p>\n<p>Thanks</p>",
      "rawMarkdown": "In my last version, I reach a CV ~ 0.605 and has a LB ~ 0.55. Do you all usually have a big gap like this (which is always the case for me).\n \n My approach is using a CNN model on spectrogram of ~5 seconds of music, not many data augmentations at all, just cropping the 5s chunk at difference places, changing the brightness of the spectrogram. I was thinking that may be 5s chunk is not big enough to generalize the classification and have tried 15s chunk but not succeeded.\n \n I think maybe the 30s of clip in the training set is sometimes from the same audios so that we might overfit the training set but not well generalize at all.\n \n Or maybe CNN spectrogram is not good at generalizing audio classification. Do you think Time Series Classification can work better?\n \n I'm running out of ideas :)) Can someone shed me some light ?\n \n Thanks",
      "votes": null
    },
    {
      "id": "1744787",
      "postDate": "04/04/2022 10:13:18",
      "content": "<p>I have similar results (~0.6 CV / ~0.55 LB). </p>\n<p>My hypothesis is that the unstability comes from the metric : some classes are hard to differentiate and two (similar) models do not necessarily agree on the label to give for a lot of samples (taking the argmax is quite brutal if you have two classes with similar scores), which results in heavy f1-score fluctuations. </p>",
      "rawMarkdown": "I have similar results (~0.6 CV / ~0.55 LB). \n\nMy hypothesis is that the unstability comes from the metric : some classes are hard to differentiate and two (similar) models do not necessarily agree on the label to give for a lot of samples (taking the argmax is quite brutal if you have two classes with similar scores), which results in heavy f1-score fluctuations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1724594,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "03/16/2022 11:15:13",
      "content": "<p>Quite a large gap between CV/LB, anybody have the idea why?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1724790,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/16/2022 14:11:38",
          "content": "<p>My score is very close to the accuracy and not f1 :/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729182,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/19/2022 18:10:31",
          "content": "<p>In my experiment, the F1 micro score is always equal to accuracy, and I think it's true after reading this: <a href=\"https://stackoverflow.com/questions/37358496/is-f1-micro-the-same-as-accuracy#:~:text=4%20Answers&amp;text=is%20not%20useful-,Show%20activity%20on%20this%20post.,case%20in%20multi-label%20classification\" target=\"_blank\">https://stackoverflow.com/questions/37358496/is-f1-micro-the-same-as-accuracy#:~:text=4%20Answers&amp;text=is%20not%20useful-,Show%20activity%20on%20this%20post.,case%20in%20multi-label%20classification</a>.</p>\n<p>Sometimes I have a big gap between CV and LB (CV~0.55 but LB~0.50). Did anyone experience the same?</p>\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729453,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/20/2022 05:16:23",
          "content": "<p>How are you calculating CV? When I take k fold mode then cv lb gap is less, however, when I take avg probs and then argmax, cv lb gap is significant. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729523,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/20/2022 07:02:42",
          "content": "<p>Because my model takes quite a long time to learn so now I am just based on one fold :D, I will try to average 5 folds to see what happens.</p>\n<p>If I remember correctly, LB has 2500 samples so do you think it's big enough to generalize the result?</p>\n<p>Maybe my model is not generalized enough :/  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729530,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/20/2022 07:16:58",
          "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> I've just submitted an average of 5 folds with each fold having a score &gt; 0.56 and my LB is 0.53093 :D. I have no idea why </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729533,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/20/2022 07:32:13",
          "content": "<p>Yes, I think averaging probs across folds might not be a very good idea, as the class thresholds might change for different models. Did you try generating class labels for each fold and then np.mode?   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729537,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/20/2022 07:35:25",
          "content": "<p>I think someone should do an adv validation to see if the train and test come from same distributions or not. If they don't come from same dist then there are implications about how much you can trust local CV.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729885,
          "author_name": "harveenchadha",
          "author_url": "",
          "post_date": "03/20/2022 16:53:10",
          "content": "<p>I witnessed the same and I thought there is some bug in my code :D Then I saw the confusion matrix of every fold and understood what's wrong. </p>\n<p>Agree with <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> hard voting maybe a better idea for this problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729896,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/20/2022 17:05:44",
          "content": "<p>Thanks a lot, <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> , <a href=\"https://www.kaggle.com/harveenchadha\" target=\"_blank\">@harveenchadha</a> . I have tried hard voting which I think makes more sense that every model is treated equally. Unfortunately, the result is not better.</p>\n<p>Do you think Oversampling can help? How can I check if the test set and training set come from the same distribution? Maby comparing the histogram of the validation prediction and the test prediction? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729920,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/20/2022 17:27:08",
          "content": "<p>I will try oversampling with audio augmentations in my next experiments. Something like this.</p>\n<pre><code>def get_transforms(*, data):\n    if data == 'train':\n        return tA.Compose(\n                transforms=[\n                     tA.ShuffleChannels(p=0.1,mode=\"per_example\",p_mode=\"per_example\", sample_rate= 16000),\n                    tA.PeakNormalization(p=0.5,apply_to='only_too_loud_sounds'),\n                     tA.AddColoredNoise(p=0.1,mode=\"per_channel\",p_mode=\"per_channel\", sample_rate=16000,max_snr_in_db = 15),\n                     tA.Shift(p=0.1,mode=\"per_example\",p_mode=\"per_example\", sample_rate=16000,\n                              max_shift=0.025, min_shift=-0.025),\n                ])\n\n    elif data == 'valid':\n        return tA.Compose([\n        ])\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1729922,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/20/2022 17:29:42",
          "content": "<p>To check if train, test come from same distribution. Make new labels 1 for train, 0 for test. And do binary classification. If AUC ~ 0.5 then they both come from same distribution, else, no. No need to split data in kfolds, use entire datasets. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1724784,
      "author_name": "dienhoa",
      "author_url": "",
      "post_date": "03/16/2022 14:04:44",
      "content": "<p>The metric is F1 but I remember seeing somewhere that it is 'macro' or 'micro'?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1724805,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/16/2022 14:24:57",
          "content": "<p><a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> - it should be <code>micro</code> for the first day of the competition I had it as <code>macro</code> but then switched.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1725321,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "03/17/2022 02:03:16",
          "content": "<p>Ohh I think thats why theres the gap,I have been measuring the macro F1 in my local setup</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1738972,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "03/29/2022 16:45:43",
      "content": "<p>CV 0.5635 (stratified 5 folds) - LB 0.56766</p>",
      "votes": null,
      "replies": [
        {
          "id": 1739225,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/29/2022 20:54:35",
          "content": "<p>That's impressive !! I'm still struggling that my CV and LB is not consistent :/ </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1739464,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/30/2022 04:24:49",
          "content": "<p><a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> how much difference are you seeing in CV vs LB? I am seeing consistent gap of ~2% (5 fold mode ensemble LB vs OOF CV)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1739674,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/30/2022 08:45:22",
          "content": "<p>Maybe it is because of myself that one fold takes a lot of time to train (~2hours/fold) so I just push the results of 1 -&gt; 3 folds, The gap usually is &gt; 3%. </p>\n<p>My latest submission is a lot better that I don't have a big gap ( I'm really happy that I have improved after 2 weeks ! ) I hope it is not just pure luck :D</p>\n<p>How much time does it take for your to train the whole 5 epochs? Do you have any strategy to experiment fastly different ideas (Train with less genre, train with less data, …) ?</p>\n<p>Thanks </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1739680,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/30/2022 08:52:33",
          "content": "<p>It takes me less than 30 minutes per fold.</p>\n<p>I use precomputed spectrograms though, the cost is that I can't do augmentations on waveforms but it's much faster than loading the waveforms and then computing the spectrograms on the fly :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1739710,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/30/2022 09:25:29",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> . I have reduced lots of time loading the waveform by using the .wav instead of .ogg, but definitely calculating the spectrogram takes a lot of time !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1739733,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/30/2022 09:51:05",
          "content": "<p>For me each epoch is 10-15 mins, so around 1.5 hours for each fold (7 epochs). <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> are you experimenting with spec augs such as mixup? I experimented a little bit with raw audio augs, but the results were not encouraging, maybe my aug were not appropriate. </p>\n<pre><code>tA.Shift(p=0.5,mode=\"per_example\",p_mode=\"per_example\", sample_rate=16000,\n                              max_shift=0.025, min_shift=-0.025)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1739774,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/30/2022 10:18:18",
          "content": "<p>I tried specaugment (time masking &amp; frequency masking), but it didn't really help.<br>\nMixup is on my todo :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1739802,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "03/30/2022 10:49:26",
          "content": "<p>I used mixup and it really helps, without it, the model will easily overfit, with mixup I can train for a long time and the training and validation score converge </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1739809,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "03/30/2022 11:00:08",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/dienhoa\" target=\"_blank\">@dienhoa</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1743562,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "04/03/2022 05:18:17",
      "content": "<p>Densenet121D CV 55.01 LB 58. Don't know whats going on.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1743616,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/03/2022 06:31:12",
          "content": "<p>Maybe we can ask <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> does he still has the consistent LB vs CV?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1743670,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "04/03/2022 07:40:45",
          "content": "<p>Yes please. <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1743767,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "04/03/2022 09:33:27",
          "content": "<p>CV vs LB is not consistent at all for me either :(</p>\n<p>I think my first submissions were just lucky on that matter.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1743895,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "04/03/2022 11:59:59",
          "content": "<p>Makes me think if <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> has done some kind of sorcery with test dataset. 👀👀</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1744769,
          "author_name": "dienhoa",
          "author_url": "",
          "post_date": "04/04/2022 09:44:35",
          "content": "<p>In my last version, I reach a CV ~ 0.605 and has a LB ~ 0.55. Do you all usually have a big gap like this (which is always the case for me).</p>\n<p>My approach is using a CNN model on spectrogram of ~5 seconds of music, not many data augmentations at all, just cropping the 5s chunk at difference places, changing the brightness of the spectrogram. I was thinking that may be 5s chunk is not big enough to generalize the classification and have tried 15s chunk but not succeeded.</p>\n<p>I think maybe the 30s of clip in the training set is sometimes from the same audios so that we might overfit the training set but not well generalize at all.</p>\n<p>Or maybe CNN spectrogram is not good at generalizing audio classification. Do you think Time Series Classification can work better?</p>\n<p>I'm running out of ideas :)) Can someone shed me some light ?</p>\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1744787,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "04/04/2022 10:13:18",
          "content": "<p>I have similar results (~0.6 CV / ~0.55 LB). </p>\n<p>My hypothesis is that the unstability comes from the metric : some classes are hard to differentiate and two (similar) models do not necessarily agree on the label to give for a lot of samples (taking the argmax is quite brutal if you have two classes with similar scores), which results in heavy f1-score fluctuations. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1724592": "lets start the ancient kaggle thread!\n\nModel : effnet_b1 (no pretrained weights, trained from scratch)\nFolds: 5\nCV: 0.39\nLB: 0.51\n(Stratified k folds based on genre)\n(used melspectrograms)",
    "1724594": "Quite a large gap between CV/LB, anybody have the idea why?",
    "1724784": "The metric is F1 but I remember seeing somewhere that it is 'macro' or 'micro'?",
    "1724790": "My score is very close to the accuracy and not f1 :/",
    "1724805": "dienhoa - it should be `micro` for the first day of the competition I had it as `macro` but then switched.",
    "1725321": "Ohh I think thats why theres the gap,I have been measuring the macro F1 in my local setup",
    "1729182": "In my experiment, the F1 micro score is always equal to accuracy, and I think it's true after reading this: https://stackoverflow.com/questions/37358496/is-f1-micro-the-same-as-accuracy#:~:text=4%20Answers&text=is%20not%20useful-,Show%20activity%20on%20this%20post.,case%20in%20multi-label%20classification.\n\nSometimes I have a big gap between CV and LB (CV~0.55 but LB~0.50). Did anyone experience the same?\n\nThanks",
    "1729453": "How are you calculating CV? When I take k fold mode then cv lb gap is less, however, when I take avg probs and then argmax, cv lb gap is significant.",
    "1729523": "Because my model takes quite a long time to learn so now I am just based on one fold :D, I will try to average 5 folds to see what happens.\n\nIf I remember correctly, LB has 2500 samples so do you think it's big enough to generalize the result?\n\nMaybe my model is not generalized enough :/",
    "1729530": "pheadrus I've just submitted an average of 5 folds with each fold having a score > 0.56 and my LB is 0.53093 :D. I have no idea why",
    "1729533": "Yes, I think averaging probs across folds might not be a very good idea, as the class thresholds might change for different models. Did you try generating class labels for each fold and then np.mode?",
    "1729537": "I think someone should do an adv validation to see if the train and test come from same distributions or not. If they don't come from same dist then there are implications about how much you can trust local CV.",
    "1729885": "I witnessed the same and I thought there is some bug in my code :D Then I saw the confusion matrix of every fold and understood what's wrong. \n\nAgree with @pheadrus hard voting maybe a better idea for this problem.",
    "1729896": "Thanks a lot, @pheadrus , @harveenchadha . I have tried hard voting which I think makes more sense that every model is treated equally. Unfortunately, the result is not better.\n\nDo you think Oversampling can help? How can I check if the test set and training set come from the same distribution? Maby comparing the histogram of the validation prediction and the test prediction?",
    "1729920": "I will try oversampling with audio augmentations in my next experiments. Something like this.\n\n\n```\ndef get_transforms(*, data):\n    if data == 'train':\n        return tA.Compose(\n                transforms=[\n                     tA.ShuffleChannels(p=0.1,mode=\"per_example\",p_mode=\"per_example\", sample_rate= 16000),\n                    tA.PeakNormalization(p=0.5,apply_to='only_too_loud_sounds'),\n                     tA.AddColoredNoise(p=0.1,mode=\"per_channel\",p_mode=\"per_channel\", sample_rate=16000,max_snr_in_db = 15),\n                     tA.Shift(p=0.1,mode=\"per_example\",p_mode=\"per_example\", sample_rate=16000,\n                              max_shift=0.025, min_shift=-0.025),\n                ])\n\n    elif data == 'valid':\n        return tA.Compose([\n        ])\n\n```",
    "1729922": "To check if train, test come from same distribution. Make new labels 1 for train, 0 for test. And do binary classification. If AUC ~ 0.5 then they both come from same distribution, else, no. No need to split data in kfolds, use entire datasets.",
    "1738972": "CV 0.5635 (stratified 5 folds) - LB 0.56766",
    "1739225": "That's impressive !! I'm still struggling that my CV and LB is not consistent :/",
    "1739464": "dienhoa how much difference are you seeing in CV vs LB? I am seeing consistent gap of ~2% (5 fold mode ensemble LB vs OOF CV)",
    "1739674": "Maybe it is because of myself that one fold takes a lot of time to train (~2hours/fold) so I just push the results of 1 -> 3 folds, The gap usually is > 3%. \n\nMy latest submission is a lot better that I don't have a big gap ( I'm really happy that I have improved after 2 weeks ! ) I hope it is not just pure luck :D\n\nHow much time does it take for your to train the whole 5 epochs? Do you have any strategy to experiment fastly different ideas (Train with less genre, train with less data, ...) ?\n\nThanks",
    "1739680": "It takes me less than 30 minutes per fold.\n\nI use precomputed spectrograms though, the cost is that I can't do augmentations on waveforms but it's much faster than loading the waveforms and then computing the spectrograms on the fly :)",
    "1739710": "Thanks a lot @theoviel . I have reduced lots of time loading the waveform by using the .wav instead of .ogg, but definitely calculating the spectrogram takes a lot of time !",
    "1739733": "For me each epoch is 10-15 mins, so around 1.5 hours for each fold (7 epochs). @theoviel are you experimenting with spec augs such as mixup? I experimented a little bit with raw audio augs, but the results were not encouraging, maybe my aug were not appropriate. \n\n```\ntA.Shift(p=0.5,mode=\"per_example\",p_mode=\"per_example\", sample_rate=16000,\n                              max_shift=0.025, min_shift=-0.025)\n```",
    "1739774": "I tried specaugment (time masking & frequency masking), but it didn't really help.\nMixup is on my todo :)",
    "1739802": "I used mixup and it really helps, without it, the model will easily overfit, with mixup I can train for a long time and the training and validation score converge",
    "1739809": "Thanks @dienhoa",
    "1743562": "Densenet121D CV 55.01 LB 58. Don't know whats going on.",
    "1743616": "Maybe we can ask @theoviel does he still has the consistent LB vs CV?",
    "1743670": "Yes please. @theoviel",
    "1743767": "CV vs LB is not consistent at all for me either :(\n\nI think my first submissions were just lucky on that matter.",
    "1743895": "Makes me think if @robikscube has done some kind of sorcery with test dataset. 👀👀",
    "1744769": "In my last version, I reach a CV ~ 0.605 and has a LB ~ 0.55. Do you all usually have a big gap like this (which is always the case for me).\n \n My approach is using a CNN model on spectrogram of ~5 seconds of music, not many data augmentations at all, just cropping the 5s chunk at difference places, changing the brightness of the spectrogram. I was thinking that may be 5s chunk is not big enough to generalize the classification and have tried 15s chunk but not succeeded.\n \n I think maybe the 30s of clip in the training set is sometimes from the same audios so that we might overfit the training set but not well generalize at all.\n \n Or maybe CNN spectrogram is not good at generalizing audio classification. Do you think Time Series Classification can work better?\n \n I'm running out of ideas :)) Can someone shed me some light ?\n \n Thanks",
    "1744787": "I have similar results (~0.6 CV / ~0.55 LB). \n\nMy hypothesis is that the unstability comes from the metric : some classes are hard to differentiate and two (similar) models do not necessarily agree on the label to give for a lot of samples (taking the argmax is quite brutal if you have two classes with similar scores), which results in heavy f1-score fluctuations."
  },
  "source": "meta"
}