{
  "id": 93664,
  "title": "Hitting a dead end",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/93664",
  "author_name": "",
  "post_date": "2019-05-29T04:20:50.994764300Z",
  "votes": 10,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I am doing ok in this compeptition with a single 2D CNN model with LB 0.652, and LB 0.684 with 5-fold CV. This uses mixup and TTA. However, I have had significant challenges improving this score. I have tried:\n- Changing to ResNet50 and Inceptionv4 \n- Increasing the number of rounds of TTA and adding more augmentations to TTA\n- Increasing number of folds!\n- Increasing number of epochs (originally at 200 epochs, increased to 400)</p>\n\n<p>All of these have shown no improvement in score (based on analysis of local CV and public LB). \nI was interested in trying unsupervised techniques but I hear it does not help very much due to the different distribution of curated and noisy labels.</p>\n\n<p>Any suggestions or tips? Are you getting similar results or is there something wrong on my end?</p>",
  "messages": [
    {
      "id": "538751",
      "postDate": "05/29/2019 04:20:50",
      "content": "<p>I am doing ok in this compeptition with a single 2D CNN model with LB 0.652, and LB 0.684 with 5-fold CV. This uses mixup and TTA. However, I have had significant challenges improving this score. I have tried:\n- Changing to ResNet50 and Inceptionv4 \n- Increasing the number of rounds of TTA and adding more augmentations to TTA\n- Increasing number of folds!\n- Increasing number of epochs (originally at 200 epochs, increased to 400)</p>\n\n<p>All of these have shown no improvement in score (based on analysis of local CV and public LB). \nI was interested in trying unsupervised techniques but I hear it does not help very much due to the different distribution of curated and noisy labels.</p>\n\n<p>Any suggestions or tips? Are you getting similar results or is there something wrong on my end?</p>",
      "rawMarkdown": "I am doing ok in this compeptition with a single 2D CNN model with LB 0.652, and LB 0.684 with 5-fold CV. This uses mixup and TTA. However, I have had significant challenges improving this score. I have tried:\n- Changing to ResNet50 and Inceptionv4 \n- Increasing the number of rounds of TTA and adding more augmentations to TTA\n- Increasing number of folds!\n- Increasing number of epochs (originally at 200 epochs, increased to 400)\n\nAll of these have shown no improvement in score (based on analysis of local CV and public LB). \nI was interested in trying unsupervised techniques but I hear it does not help very much due to the different distribution of curated and noisy labels.\n\nAny suggestions or tips? Are you getting similar results or is there something wrong on my end?",
      "votes": null
    },
    {
      "id": "538798",
      "postDate": "05/29/2019 06:12:43",
      "content": "<p>Meanwhile, I cannot figure out how to use KFold in fast.ai :/ </p>",
      "rawMarkdown": "Meanwhile, I cannot figure out how to use KFold in fast.ai :/",
      "votes": null
    },
    {
      "id": "538870",
      "postDate": "05/29/2019 08:06:15",
      "content": "<p>I feel you man. My first submission scored 0.659, 2nd was 5fold and scored 0.684 and nothing worked after that... Tried everything from your list,... my CV lwl goes to 0.86+ (avg for 5kfold) but no success on the LB</p>",
      "rawMarkdown": "I feel you man. My first submission scored 0.659, 2nd was 5fold and scored 0.684 and nothing worked after that... Tried everything from your list,... my CV lwl goes to 0.86+ (avg for 5kfold) but no success on the LB",
      "votes": null
    },
    {
      "id": "538883",
      "postDate": "05/29/2019 08:37:45",
      "content": "<p>I'm stuck at similar levels. I somehow managed to get to get 0.7 LB using a complicated several stage training process with noisy data, but I'm not sure if that's a consistent improvement or just luck and overfit to LB.\nWhat frustrates me most is that I'm able to up local CV to 0.9+ using noisy data(CV is still measured on curated only), but that doesn't lead to any meaningful LB improvements.</p>\n\n<p>I use all the same things as you(2D CNN, mixup, TTA, k-fold averaging)</p>",
      "rawMarkdown": "I'm stuck at similar levels. I somehow managed to get to get 0.7 LB using a complicated several stage training process with noisy data, but I'm not sure if that's a consistent improvement or just luck and overfit to LB.\nWhat frustrates me most is that I'm able to up local CV to 0.9+ using noisy data(CV is still measured on curated only), but that doesn't lead to any meaningful LB improvements.\n\nI use all the same things as you(2D CNN, mixup, TTA, k-fold averaging)",
      "votes": null
    },
    {
      "id": "538895",
      "postDate": "05/29/2019 09:08:38",
      "content": "<p>Wow, a 0.9 local CV sounds incredible! I never managed to get higher than 0.87 locally. However, contrary to your experience, I consistently observe that LB score increases as my local score improves. </p>",
      "rawMarkdown": "Wow, a 0.9 local CV sounds incredible! I never managed to get higher than 0.87 locally. However, contrary to your experience, I consistently observe that LB score increases as my local score improves.",
      "votes": null
    },
    {
      "id": "538911",
      "postDate": "05/29/2019 09:33:41",
      "content": "<p>I generally have a correlation between CV and LB, but the variance is big. CV around 0.83 can lead to LB from 0.62 to 0.67(single model) </p>",
      "rawMarkdown": "I generally have a correlation between CV and LB, but the variance is big. CV around 0.83 can lead to LB from 0.62 to 0.67(single model)",
      "votes": null
    },
    {
      "id": "539072",
      "postDate": "05/29/2019 13:29:14",
      "content": "<p>I trained a 12 layers ResNet (not predefined model) with only curated data and achieved 0.68 LB without k-fold cv (using mixup, cutout, TTA). I spent a long time searching good parameters for making mel-spectrogram images. I expect that both data handling and model training are important in this competition.</p>",
      "rawMarkdown": "I trained a 12 layers ResNet (not predefined model) with only curated data and achieved 0.68 LB without k-fold cv (using mixup, cutout, TTA). I spent a long time searching good parameters for making mel-spectrogram images. I expect that both data handling and model training are important in this competition.",
      "votes": null
    },
    {
      "id": "539289",
      "postDate": "05/29/2019 20:42:15",
      "content": "<p>You are using different mel-spectrogram parameters apart from the ones used by <a href=\"/daisukelab\">@daisukelab</a> in those kernels?</p>",
      "rawMarkdown": "You are using different mel-spectrogram parameters apart from the ones used by @daisukelab in those kernels?",
      "votes": null
    },
    {
      "id": "539366",
      "postDate": "05/30/2019 01:47:37",
      "content": "<p>Yes. My program for making mel-spectrogram images is based on <a href=\"/daisukelab\">@daisukelab</a> kernel (it helps me a lot). But <code>sampling_rate</code>, <code>duration</code>, <code>hop_length</code> were changed. In addition, procedures and parameters for silence trimming were also changed.</p>\n\n<p>These tuning improved my LB by about 0.02.</p>",
      "rawMarkdown": "Yes. My program for making mel-spectrogram images is based on @daisukelab kernel (it helps me a lot). But `sampling_rate`, `duration`, `hop_length` were changed. In addition, procedures and parameters for silence trimming were also changed.\n\nThese tuning improved my LB by about 0.02.",
      "votes": null
    },
    {
      "id": "539939",
      "postDate": "05/30/2019 18:26:22",
      "content": "<p>You can used KFolds to generate a list of indexes for your validation set. Then you can use fastai's <code>.split_by_idx(val_index)</code> to use these for a validation set. </p>\n\n<p>Here is a useful starting point for KFolds: <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/80809\">https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/80809</a></p>",
      "rawMarkdown": "You can used KFolds to generate a list of indexes for your validation set. Then you can use fastai's `.split_by_idx(val_index)` to use these for a validation set. \n\nHere is a useful starting point for KFolds: https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/80809",
      "votes": null
    },
    {
      "id": "540649",
      "postDate": "05/31/2019 19:47:50",
      "content": "<p>That’s interesting <a href=\"/takedarts\">@takedarts</a> ! Hearing from you, I ‘ll try to explore more on mel-spectrogram parameters (I have tried, but failed to improve) ... In any cases, <em>after the end of competition</em> , hope that you would share your optimal parameters and how you trim the data!</p>\n\n<p>By the way <a href=\"/takedarts\">@takedarts</a> , what is your local CV for the .68LB ?</p>",
      "rawMarkdown": "That’s interesting @takedarts ! Hearing from you, I ‘ll try to explore more on mel-spectrogram parameters (I have tried, but failed to improve) ... In any cases, *after the end of competition* , hope that you would share your optimal parameters and how you trim the data!\n\nBy the way @takedarts , what is your local CV for the .68LB ?",
      "votes": null
    },
    {
      "id": "540794",
      "postDate": "06/01/2019 05:29:50",
      "content": "<p>There is also that: <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/92858#latest-539290\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/92858#latest-539290</a></p>\n\n<p>Hope it helps!</p>",
      "rawMarkdown": "There is also that: https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/92858#latest-539290\n\nHope it helps!",
      "votes": null
    },
    {
      "id": "540800",
      "postDate": "06/01/2019 05:42:26",
      "content": "<p>thank you <a href=\"/toldo171\">@toldo171</a> . Appreciate it :)</p>",
      "rawMarkdown": "thank you @toldo171 . Appreciate it :)",
      "votes": null
    },
    {
      "id": "540801",
      "postDate": "06/01/2019 05:42:39",
      "content": "<p>thank you <a href=\"/joshvarty\">@joshvarty</a> :)</p>",
      "rawMarkdown": "thank you @joshvarty :)",
      "votes": null
    },
    {
      "id": "541004",
      "postDate": "06/01/2019 14:27:06",
      "content": "<p>My local CV was 0.85+ when my model achieved 0.68 LB.\nNow, my local CV is 0.87+, but LB score of the model is worse than 0.68. I feel it is too difficult to understand what the local CV scores indicate.</p>",
      "rawMarkdown": "My local CV was 0.85+ when my model achieved 0.68 LB.\nNow, my local CV is 0.87+, but LB score of the model is worse than 0.68. I feel it is too difficult to understand what the local CV scores indicate.",
      "votes": null
    },
    {
      "id": "541297",
      "postDate": "06/02/2019 07:43:39",
      "content": "<p><a href=\"/takedarts\">@takedarts</a> My single fold score (no K-fold) is similar to yours (0.677LB, 0.85x CV ; the two scores have been consistently up along together), but I took almost extreme opposite path; my score is achieved by super complicated training with noisy data. But I could not at all improve the preprocessed data; it would be interesting to see our K-folds score, and also interesting to see whether our methods can be compatible / combined.</p>\n\n<p>By the way does your preprocess will take a lot of time in the 2nd stage inference ??</p>",
      "rawMarkdown": "takedarts My single fold score (no K-fold) is similar to yours (0.677LB, 0.85x CV ; the two scores have been consistently up along together), but I took almost extreme opposite path; my score is achieved by super complicated training with noisy data. But I could not at all improve the preprocessed data; it would be interesting to see our K-folds score, and also interesting to see whether our methods can be compatible / combined.\n\nBy the way does your preprocess will take a lot of time in the 2nd stage inference ??",
      "votes": null
    },
    {
      "id": "541299",
      "postDate": "06/02/2019 07:49:03",
      "content": "<p><a href=\"/vzaguskin\">@vzaguskin</a> Do you achieve 0.9+ CV score to all K folds ??, or that just measure on one validation data split ?</p>",
      "rawMarkdown": "vzaguskin Do you achieve 0.9+ CV score to all K folds ??, or that just measure on one validation data split ?",
      "votes": null
    },
    {
      "id": "541496",
      "postDate": "06/02/2019 14:50:11",
      "content": "<p>All 5 folds.</p>",
      "rawMarkdown": "All 5 folds.",
      "votes": null
    },
    {
      "id": "541514",
      "postDate": "06/02/2019 15:26:59",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> In kernel environment, my preprocess takes 10 min in 1st stage, so I expect that it will take about 30 min in 2nd stage. As talked above, my preprocess is close to <a href=\"/daisukelab\">@daisukelab</a> kernel except a silent trimming procedure which is not so complex.\nCurrently, my best model is a shallow CNN with a simple training schedule. But, I guess it is pretty different from others because my training uses only Momentum SGD (many published kernels use Adam).</p>",
      "rawMarkdown": "ratthachat In kernel environment, my preprocess takes 10 min in 1st stage, so I expect that it will take about 30 min in 2nd stage. As talked above, my preprocess is close to @daisukelab kernel except a silent trimming procedure which is not so complex.\nCurrently, my best model is a shallow CNN with a simple training schedule. But, I guess it is pretty different from others because my training uses only Momentum SGD (many published kernels use Adam).",
      "votes": null
    },
    {
      "id": "541668",
      "postDate": "06/02/2019 21:21:33",
      "content": "<p>I've done a small research about popular optimizers and Adam seems to be recommended as \"way to go\" default choise.\nBut the problem with Adam is that it doesn't support momentums so if you want to use cyclic lr schedulers, then you are looking for SGD or RmsProp.</p>",
      "rawMarkdown": "I've done a small research about popular optimizers and Adam seems to be recommended as \"way to go\" default choise.\nBut the problem with Adam is that it doesn't support momentums so if you want to use cyclic lr schedulers, then you are looking for SGD or RmsProp.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 538798,
      "author_name": "tanmaypandey",
      "author_url": "",
      "post_date": "05/29/2019 06:12:43",
      "content": "<p>Meanwhile, I cannot figure out how to use KFold in fast.ai :/ </p>",
      "votes": null,
      "replies": [
        {
          "id": 539939,
          "author_name": "joshvarty",
          "author_url": "",
          "post_date": "05/30/2019 18:26:22",
          "content": "<p>You can used KFolds to generate a list of indexes for your validation set. Then you can use fastai's <code>.split_by_idx(val_index)</code> to use these for a validation set. </p>\n\n<p>Here is a useful starting point for KFolds: <a href=\"https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/80809\">https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/80809</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540794,
          "author_name": "toldo171",
          "author_url": "",
          "post_date": "06/01/2019 05:29:50",
          "content": "<p>There is also that: <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/92858#latest-539290\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/92858#latest-539290</a></p>\n\n<p>Hope it helps!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540800,
          "author_name": "tanmaypandey",
          "author_url": "",
          "post_date": "06/01/2019 05:42:26",
          "content": "<p>thank you <a href=\"/toldo171\">@toldo171</a> . Appreciate it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540801,
          "author_name": "tanmaypandey",
          "author_url": "",
          "post_date": "06/01/2019 05:42:39",
          "content": "<p>thank you <a href=\"/joshvarty\">@joshvarty</a> :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 538870,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "05/29/2019 08:06:15",
      "content": "<p>I feel you man. My first submission scored 0.659, 2nd was 5fold and scored 0.684 and nothing worked after that... Tried everything from your list,... my CV lwl goes to 0.86+ (avg for 5kfold) but no success on the LB</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 538883,
      "author_name": "vzaguskin",
      "author_url": "",
      "post_date": "05/29/2019 08:37:45",
      "content": "<p>I'm stuck at similar levels. I somehow managed to get to get 0.7 LB using a complicated several stage training process with noisy data, but I'm not sure if that's a consistent improvement or just luck and overfit to LB.\nWhat frustrates me most is that I'm able to up local CV to 0.9+ using noisy data(CV is still measured on curated only), but that doesn't lead to any meaningful LB improvements.</p>\n\n<p>I use all the same things as you(2D CNN, mixup, TTA, k-fold averaging)</p>",
      "votes": null,
      "replies": [
        {
          "id": 538895,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "05/29/2019 09:08:38",
          "content": "<p>Wow, a 0.9 local CV sounds incredible! I never managed to get higher than 0.87 locally. However, contrary to your experience, I consistently observe that LB score increases as my local score improves. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538911,
          "author_name": "vzaguskin",
          "author_url": "",
          "post_date": "05/29/2019 09:33:41",
          "content": "<p>I generally have a correlation between CV and LB, but the variance is big. CV around 0.83 can lead to LB from 0.62 to 0.67(single model) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 541299,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "06/02/2019 07:49:03",
          "content": "<p><a href=\"/vzaguskin\">@vzaguskin</a> Do you achieve 0.9+ CV score to all K folds ??, or that just measure on one validation data split ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 541496,
          "author_name": "vzaguskin",
          "author_url": "",
          "post_date": "06/02/2019 14:50:11",
          "content": "<p>All 5 folds.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 539072,
      "author_name": "takedarts",
      "author_url": "",
      "post_date": "05/29/2019 13:29:14",
      "content": "<p>I trained a 12 layers ResNet (not predefined model) with only curated data and achieved 0.68 LB without k-fold cv (using mixup, cutout, TTA). I spent a long time searching good parameters for making mel-spectrogram images. I expect that both data handling and model training are important in this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 539289,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "05/29/2019 20:42:15",
          "content": "<p>You are using different mel-spectrogram parameters apart from the ones used by <a href=\"/daisukelab\">@daisukelab</a> in those kernels?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539366,
          "author_name": "takedarts",
          "author_url": "",
          "post_date": "05/30/2019 01:47:37",
          "content": "<p>Yes. My program for making mel-spectrogram images is based on <a href=\"/daisukelab\">@daisukelab</a> kernel (it helps me a lot). But <code>sampling_rate</code>, <code>duration</code>, <code>hop_length</code> were changed. In addition, procedures and parameters for silence trimming were also changed.</p>\n\n<p>These tuning improved my LB by about 0.02.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 540649,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "05/31/2019 19:47:50",
          "content": "<p>That’s interesting <a href=\"/takedarts\">@takedarts</a> ! Hearing from you, I ‘ll try to explore more on mel-spectrogram parameters (I have tried, but failed to improve) ... In any cases, <em>after the end of competition</em> , hope that you would share your optimal parameters and how you trim the data!</p>\n\n<p>By the way <a href=\"/takedarts\">@takedarts</a> , what is your local CV for the .68LB ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 541004,
          "author_name": "takedarts",
          "author_url": "",
          "post_date": "06/01/2019 14:27:06",
          "content": "<p>My local CV was 0.85+ when my model achieved 0.68 LB.\nNow, my local CV is 0.87+, but LB score of the model is worse than 0.68. I feel it is too difficult to understand what the local CV scores indicate.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 541297,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "06/02/2019 07:43:39",
          "content": "<p><a href=\"/takedarts\">@takedarts</a> My single fold score (no K-fold) is similar to yours (0.677LB, 0.85x CV ; the two scores have been consistently up along together), but I took almost extreme opposite path; my score is achieved by super complicated training with noisy data. But I could not at all improve the preprocessed data; it would be interesting to see our K-folds score, and also interesting to see whether our methods can be compatible / combined.</p>\n\n<p>By the way does your preprocess will take a lot of time in the 2nd stage inference ??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 541514,
          "author_name": "takedarts",
          "author_url": "",
          "post_date": "06/02/2019 15:26:59",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> In kernel environment, my preprocess takes 10 min in 1st stage, so I expect that it will take about 30 min in 2nd stage. As talked above, my preprocess is close to <a href=\"/daisukelab\">@daisukelab</a> kernel except a silent trimming procedure which is not so complex.\nCurrently, my best model is a shallow CNN with a simple training schedule. But, I guess it is pretty different from others because my training uses only Momentum SGD (many published kernels use Adam).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 541668,
          "author_name": "vandalko",
          "author_url": "",
          "post_date": "06/02/2019 21:21:33",
          "content": "<p>I've done a small research about popular optimizers and Adam seems to be recommended as \"way to go\" default choise.\nBut the problem with Adam is that it doesn't support momentums so if you want to use cyclic lr schedulers, then you are looking for SGD or RmsProp.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "538751": "I am doing ok in this compeptition with a single 2D CNN model with LB 0.652, and LB 0.684 with 5-fold CV. This uses mixup and TTA. However, I have had significant challenges improving this score. I have tried:\n- Changing to ResNet50 and Inceptionv4 \n- Increasing the number of rounds of TTA and adding more augmentations to TTA\n- Increasing number of folds!\n- Increasing number of epochs (originally at 200 epochs, increased to 400)\n\nAll of these have shown no improvement in score (based on analysis of local CV and public LB). \nI was interested in trying unsupervised techniques but I hear it does not help very much due to the different distribution of curated and noisy labels.\n\nAny suggestions or tips? Are you getting similar results or is there something wrong on my end?",
    "538798": "Meanwhile, I cannot figure out how to use KFold in fast.ai :/",
    "538870": "I feel you man. My first submission scored 0.659, 2nd was 5fold and scored 0.684 and nothing worked after that... Tried everything from your list,... my CV lwl goes to 0.86+ (avg for 5kfold) but no success on the LB",
    "538883": "I'm stuck at similar levels. I somehow managed to get to get 0.7 LB using a complicated several stage training process with noisy data, but I'm not sure if that's a consistent improvement or just luck and overfit to LB.\nWhat frustrates me most is that I'm able to up local CV to 0.9+ using noisy data(CV is still measured on curated only), but that doesn't lead to any meaningful LB improvements.\n\nI use all the same things as you(2D CNN, mixup, TTA, k-fold averaging)",
    "538895": "Wow, a 0.9 local CV sounds incredible! I never managed to get higher than 0.87 locally. However, contrary to your experience, I consistently observe that LB score increases as my local score improves.",
    "538911": "I generally have a correlation between CV and LB, but the variance is big. CV around 0.83 can lead to LB from 0.62 to 0.67(single model)",
    "539072": "I trained a 12 layers ResNet (not predefined model) with only curated data and achieved 0.68 LB without k-fold cv (using mixup, cutout, TTA). I spent a long time searching good parameters for making mel-spectrogram images. I expect that both data handling and model training are important in this competition.",
    "539289": "You are using different mel-spectrogram parameters apart from the ones used by @daisukelab in those kernels?",
    "539366": "Yes. My program for making mel-spectrogram images is based on @daisukelab kernel (it helps me a lot). But `sampling_rate`, `duration`, `hop_length` were changed. In addition, procedures and parameters for silence trimming were also changed.\n\nThese tuning improved my LB by about 0.02.",
    "539939": "You can used KFolds to generate a list of indexes for your validation set. Then you can use fastai's `.split_by_idx(val_index)` to use these for a validation set. \n\nHere is a useful starting point for KFolds: https://www.kaggle.com/c/santander-customer-transaction-prediction/discussion/80809",
    "540649": "That’s interesting @takedarts ! Hearing from you, I ‘ll try to explore more on mel-spectrogram parameters (I have tried, but failed to improve) ... In any cases, *after the end of competition* , hope that you would share your optimal parameters and how you trim the data!\n\nBy the way @takedarts , what is your local CV for the .68LB ?",
    "540794": "There is also that: https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/92858#latest-539290\n\nHope it helps!",
    "540800": "thank you @toldo171 . Appreciate it :)",
    "540801": "thank you @joshvarty :)",
    "541004": "My local CV was 0.85+ when my model achieved 0.68 LB.\nNow, my local CV is 0.87+, but LB score of the model is worse than 0.68. I feel it is too difficult to understand what the local CV scores indicate.",
    "541297": "takedarts My single fold score (no K-fold) is similar to yours (0.677LB, 0.85x CV ; the two scores have been consistently up along together), but I took almost extreme opposite path; my score is achieved by super complicated training with noisy data. But I could not at all improve the preprocessed data; it would be interesting to see our K-folds score, and also interesting to see whether our methods can be compatible / combined.\n\nBy the way does your preprocess will take a lot of time in the 2nd stage inference ??",
    "541299": "vzaguskin Do you achieve 0.9+ CV score to all K folds ??, or that just measure on one validation data split ?",
    "541496": "All 5 folds.",
    "541514": "ratthachat In kernel environment, my preprocess takes 10 min in 1st stage, so I expect that it will take about 30 min in 2nd stage. As talked above, my preprocess is close to @daisukelab kernel except a silent trimming procedure which is not so complex.\nCurrently, my best model is a shallow CNN with a simple training schedule. But, I guess it is pretty different from others because my training uses only Momentum SGD (many published kernels use Adam).",
    "541668": "I've done a small research about popular optimizers and Adam seems to be recommended as \"way to go\" default choise.\nBut the problem with Adam is that it doesn't support momentums so if you want to use cyclic lr schedulers, then you are looking for SGD or RmsProp."
  },
  "source": "meta"
}