{
  "id": 93216,
  "title": "How to break 0.66 for k-fold?",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/93216",
  "author_name": "",
  "post_date": "2019-05-24T11:31:43.195959Z",
  "votes": 3,
  "comment_count": 11,
  "views": 0,
  "content": "<p>As posted in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-535917\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-535917</a> , many kagglers are able to get 0.68+ with only k-fold and curated dataset. But It has been a hard time for me to break 0.66 using k-fold. I have tried:</p>\n\n<ol>\n<li>different kind of augmentation</li>\n<li>different image size to train CNN</li>\n<li>different use of noisy data</li>\n<li>different CNN architecture</li>\n<li>mixup (mix up two image with different weight and combine their label)</li>\n<li>tta</li>\n</ol>\n\n<p>some of them help, but still can not break 0.66, there might be something that I had been missed.... What's your trick to break 0.66? any idea will be appreciated.</p>",
  "messages": [
    {
      "id": "536379",
      "postDate": "05/24/2019 11:31:43",
      "content": "<p>As posted in <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-535917\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-535917</a> , many kagglers are able to get 0.68+ with only k-fold and curated dataset. But It has been a hard time for me to break 0.66 using k-fold. I have tried:</p>\n\n<ol>\n<li>different kind of augmentation</li>\n<li>different image size to train CNN</li>\n<li>different use of noisy data</li>\n<li>different CNN architecture</li>\n<li>mixup (mix up two image with different weight and combine their label)</li>\n<li>tta</li>\n</ol>\n\n<p>some of them help, but still can not break 0.66, there might be something that I had been missed.... What's your trick to break 0.66? any idea will be appreciated.</p>",
      "rawMarkdown": "As posted in https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-535917 , many kagglers are able to get 0.68+ with only k-fold and curated dataset. But It has been a hard time for me to break 0.66 using k-fold. I have tried:\n\n1. different kind of augmentation\n2. different image size to train CNN\n3. different use of noisy data\n4. different CNN architecture\n5. mixup (mix up two image with different weight and combine their label)\n6. tta\n\nsome of them help, but still can not break 0.66, there might be something that I had been missed.... What's your trick to break 0.66? any idea will be appreciated.",
      "votes": null
    },
    {
      "id": "536388",
      "postDate": "05/24/2019 11:45:19",
      "content": "<p>mixup &amp; tta</p>",
      "rawMarkdown": "mixup &amp; tta",
      "votes": null
    },
    {
      "id": "536403",
      "postDate": "05/24/2019 11:58:46",
      "content": "<p>Thanks for your reply, actually, I have tried both (forget to mention), </p>\n\n<p>by saying 'mixup', do you mean 'mix up two image with different weight and combine their label'? this is what I am currently doing, and it does not give improvement for local cv and public lb...</p>",
      "rawMarkdown": "Thanks for your reply, actually, I have tried both (forget to mention), \n\nby saying 'mixup', do you mean 'mix up two image with different weight and combine their label'? this is what I am currently doing, and it does not give improvement for local cv and public lb...",
      "votes": null
    },
    {
      "id": "536458",
      "postDate": "05/24/2019 14:08:21",
      "content": "<p>Regarding mixup, did you see train &amp; valid loss relationship drastically change? And it would give 0.1~0.2 improvement I guess.\nIf not, you can check out working implementations. Here's links:\n- My first call, fastai implementation: <a href=\"https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py\">https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py</a>\n- For Keras, this really helped me last year: <a href=\"https://github.com/yu4u/mixup-generator\">https://github.com/yu4u/mixup-generator</a>\n- Facebook research also has: <a href=\"https://github.com/facebookresearch/mixup-cifar10\">https://github.com/facebookresearch/mixup-cifar10</a></p>\n\n<p>I think ... we need to re-re-re-confirm standing on solid ground. Let's check our implementation, is it working correctly as expected?</p>",
      "rawMarkdown": "Regarding mixup, did you see train &amp; valid loss relationship drastically change? And it would give 0.1~0.2 improvement I guess.\nIf not, you can check out working implementations. Here's links:\n- My first call, fastai implementation: https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py\n- For Keras, this really helped me last year: https://github.com/yu4u/mixup-generator\n- Facebook research also has: https://github.com/facebookresearch/mixup-cifar10\n\nI think ... we need to re-re-re-confirm standing on solid ground. Let's check our implementation, is it working correctly as expected?",
      "votes": null
    },
    {
      "id": "536482",
      "postDate": "05/24/2019 14:36:49",
      "content": "<p>And regarding TTA, did you check out how the augmented test images? If not, please take your time for debugging it. It also give us +0.1~0.2 for sure.\nMy public kernel uses fastai's implementation as it is, but it doesn't work in that case. It is for normal images.\nWe have to check how the augmentation is working inside the TTA implementation.\nI mean, visualize what's the exact sample of TTA augmentation one by one. Then we would see what to be done next naturally.</p>\n\n<p>Again, debugging is really important I think.</p>",
      "rawMarkdown": "And regarding TTA, did you check out how the augmented test images? If not, please take your time for debugging it. It also give us +0.1~0.2 for sure.\nMy public kernel uses fastai's implementation as it is, but it doesn't work in that case. It is for normal images.\nWe have to check how the augmentation is working inside the TTA implementation.\nI mean, visualize what's the exact sample of TTA augmentation one by one. Then we would see what to be done next naturally.\n\nAgain, debugging is really important I think.",
      "votes": null
    },
    {
      "id": "536546",
      "postDate": "05/24/2019 16:35:25",
      "content": "<p>Agree with daisukelab,  mixup should have a significant positive impact on your results. So if its not helping, there is probably some issue with your implementation. Same with TTA, this should also help a lot. </p>",
      "rawMarkdown": "Agree with daisukelab,  mixup should have a significant positive impact on your results. So if its not helping, there is probably some issue with your implementation. Same with TTA, this should also help a lot.",
      "votes": null
    },
    {
      "id": "536578",
      "postDate": "05/24/2019 17:57:13",
      "content": "<p>I'm interested in how you guys make inference using TTA. Are you taking N random clips of each sample, doing different augmentations on different clips, feeding them into model, and taking average of the N predictions?</p>",
      "rawMarkdown": "I'm interested in how you guys make inference using TTA. Are you taking N random clips of each sample, doing different augmentations on different clips, feeding them into model, and taking average of the N predictions?",
      "votes": null
    },
    {
      "id": "536635",
      "postDate": "05/24/2019 21:48:16",
      "content": "<p>I'm currently splitting every samples into N pieces (with overlaps if needed), then apply very little augmentations.\n- Just splitting into some N, and taking mean of N predictions will give better result.\n- Augmentation has to be modest and realistic, or would worsen your result.\n- As reported in other thread, random splitting could be better. (Though I'm doing equal split, for better reproducibility)</p>",
      "rawMarkdown": "I'm currently splitting every samples into N pieces (with overlaps if needed), then apply very little augmentations.\n- Just splitting into some N, and taking mean of N predictions will give better result.\n- Augmentation has to be modest and realistic, or would worsen your result.\n- As reported in other thread, random splitting could be better. (Though I'm doing equal split, for better reproducibility)",
      "votes": null
    },
    {
      "id": "536642",
      "postDate": "05/24/2019 22:13:37",
      "content": "<p>Yeah I saw that post as well. Thanks for the info! Let me do both and see which one is better</p>",
      "rawMarkdown": "Yeah I saw that post as well. Thanks for the info! Let me do both and see which one is better",
      "votes": null
    },
    {
      "id": "536662",
      "postDate": "05/25/2019 00:03:09",
      "content": "<p>I found that \"A note about using Mix~up\" in this thread from last competition clearly shows the important know-how.\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/64262#latest-523914\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/64262#latest-523914</a>\nI think it is the same with what's implemented in fastai's.</p>",
      "rawMarkdown": "I found that \"A note about using Mix~up\" in this thread from last competition clearly shows the important know-how.\nhttps://www.kaggle.com/c/freesound-audio-tagging/discussion/64262#latest-523914\nI think it is the same with what's implemented in fastai's.",
      "votes": null
    },
    {
      "id": "536696",
      "postDate": "05/25/2019 03:17:28",
      "content": "<p>thanks for your reply, yes, there must be something wrong with my implementation, since most people get improvement using this tricks.</p>",
      "rawMarkdown": "thanks for your reply, yes, there must be something wrong with my implementation, since most people get improvement using this tricks.",
      "votes": null
    },
    {
      "id": "536718",
      "postDate": "05/25/2019 04:36:57",
      "content": "<p>pretrained model . Even it is not allowed. </p>",
      "rawMarkdown": "pretrained model . Even it is not allowed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 536388,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "05/24/2019 11:45:19",
      "content": "<p>mixup &amp; tta</p>",
      "votes": null,
      "replies": [
        {
          "id": 536403,
          "author_name": "zjucor",
          "author_url": "",
          "post_date": "05/24/2019 11:58:46",
          "content": "<p>Thanks for your reply, actually, I have tried both (forget to mention), </p>\n\n<p>by saying 'mixup', do you mean 'mix up two image with different weight and combine their label'? this is what I am currently doing, and it does not give improvement for local cv and public lb...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536458,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/24/2019 14:08:21",
          "content": "<p>Regarding mixup, did you see train &amp; valid loss relationship drastically change? And it would give 0.1~0.2 improvement I guess.\nIf not, you can check out working implementations. Here's links:\n- My first call, fastai implementation: <a href=\"https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py\">https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py</a>\n- For Keras, this really helped me last year: <a href=\"https://github.com/yu4u/mixup-generator\">https://github.com/yu4u/mixup-generator</a>\n- Facebook research also has: <a href=\"https://github.com/facebookresearch/mixup-cifar10\">https://github.com/facebookresearch/mixup-cifar10</a></p>\n\n<p>I think ... we need to re-re-re-confirm standing on solid ground. Let's check our implementation, is it working correctly as expected?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536482,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/24/2019 14:36:49",
          "content": "<p>And regarding TTA, did you check out how the augmented test images? If not, please take your time for debugging it. It also give us +0.1~0.2 for sure.\nMy public kernel uses fastai's implementation as it is, but it doesn't work in that case. It is for normal images.\nWe have to check how the augmentation is working inside the TTA implementation.\nI mean, visualize what's the exact sample of TTA augmentation one by one. Then we would see what to be done next naturally.</p>\n\n<p>Again, debugging is really important I think.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536546,
          "author_name": "jamesrequa",
          "author_url": "",
          "post_date": "05/24/2019 16:35:25",
          "content": "<p>Agree with daisukelab,  mixup should have a significant positive impact on your results. So if its not helping, there is probably some issue with your implementation. Same with TTA, this should also help a lot. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536578,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "05/24/2019 17:57:13",
          "content": "<p>I'm interested in how you guys make inference using TTA. Are you taking N random clips of each sample, doing different augmentations on different clips, feeding them into model, and taking average of the N predictions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536635,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/24/2019 21:48:16",
          "content": "<p>I'm currently splitting every samples into N pieces (with overlaps if needed), then apply very little augmentations.\n- Just splitting into some N, and taking mean of N predictions will give better result.\n- Augmentation has to be modest and realistic, or would worsen your result.\n- As reported in other thread, random splitting could be better. (Though I'm doing equal split, for better reproducibility)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536642,
          "author_name": "jihangz",
          "author_url": "",
          "post_date": "05/24/2019 22:13:37",
          "content": "<p>Yeah I saw that post as well. Thanks for the info! Let me do both and see which one is better</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536662,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/25/2019 00:03:09",
          "content": "<p>I found that \"A note about using Mix~up\" in this thread from last competition clearly shows the important know-how.\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging/discussion/64262#latest-523914\">https://www.kaggle.com/c/freesound-audio-tagging/discussion/64262#latest-523914</a>\nI think it is the same with what's implemented in fastai's.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 536696,
          "author_name": "zjucor",
          "author_url": "",
          "post_date": "05/25/2019 03:17:28",
          "content": "<p>thanks for your reply, yes, there must be something wrong with my implementation, since most people get improvement using this tricks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 536718,
      "author_name": "baomengjiao",
      "author_url": "",
      "post_date": "05/25/2019 04:36:57",
      "content": "<p>pretrained model . Even it is not allowed. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "536379": "As posted in https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91881#latest-535917 , many kagglers are able to get 0.68+ with only k-fold and curated dataset. But It has been a hard time for me to break 0.66 using k-fold. I have tried:\n\n1. different kind of augmentation\n2. different image size to train CNN\n3. different use of noisy data\n4. different CNN architecture\n5. mixup (mix up two image with different weight and combine their label)\n6. tta\n\nsome of them help, but still can not break 0.66, there might be something that I had been missed.... What's your trick to break 0.66? any idea will be appreciated.",
    "536388": "mixup &amp; tta",
    "536403": "Thanks for your reply, actually, I have tried both (forget to mention), \n\nby saying 'mixup', do you mean 'mix up two image with different weight and combine their label'? this is what I am currently doing, and it does not give improvement for local cv and public lb...",
    "536458": "Regarding mixup, did you see train &amp; valid loss relationship drastically change? And it would give 0.1~0.2 improvement I guess.\nIf not, you can check out working implementations. Here's links:\n- My first call, fastai implementation: https://github.com/fastai/fastai/blob/master/fastai/callbacks/mixup.py\n- For Keras, this really helped me last year: https://github.com/yu4u/mixup-generator\n- Facebook research also has: https://github.com/facebookresearch/mixup-cifar10\n\nI think ... we need to re-re-re-confirm standing on solid ground. Let's check our implementation, is it working correctly as expected?",
    "536482": "And regarding TTA, did you check out how the augmented test images? If not, please take your time for debugging it. It also give us +0.1~0.2 for sure.\nMy public kernel uses fastai's implementation as it is, but it doesn't work in that case. It is for normal images.\nWe have to check how the augmentation is working inside the TTA implementation.\nI mean, visualize what's the exact sample of TTA augmentation one by one. Then we would see what to be done next naturally.\n\nAgain, debugging is really important I think.",
    "536546": "Agree with daisukelab,  mixup should have a significant positive impact on your results. So if its not helping, there is probably some issue with your implementation. Same with TTA, this should also help a lot.",
    "536578": "I'm interested in how you guys make inference using TTA. Are you taking N random clips of each sample, doing different augmentations on different clips, feeding them into model, and taking average of the N predictions?",
    "536635": "I'm currently splitting every samples into N pieces (with overlaps if needed), then apply very little augmentations.\n- Just splitting into some N, and taking mean of N predictions will give better result.\n- Augmentation has to be modest and realistic, or would worsen your result.\n- As reported in other thread, random splitting could be better. (Though I'm doing equal split, for better reproducibility)",
    "536642": "Yeah I saw that post as well. Thanks for the info! Let me do both and see which one is better",
    "536662": "I found that \"A note about using Mix~up\" in this thread from last competition clearly shows the important know-how.\nhttps://www.kaggle.com/c/freesound-audio-tagging/discussion/64262#latest-523914\nI think it is the same with what's implemented in fastai's.",
    "536696": "thanks for your reply, yes, there must be something wrong with my implementation, since most people get improvement using this tricks.",
    "536718": "pretrained model . Even it is not allowed."
  },
  "source": "meta"
}