{
  "id": 95382,
  "title": "[8th place solution] : SpecMix and warm-up pipeline",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/eric-bouteillon-8th-place-solution-specmix-and-war",
  "author_name": "",
  "post_date": "2019-06-28T19:27:47.777Z",
  "votes": 42,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>I released on a <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">github repository</a> my solution. It is highly inspired from <a href=\"/daisukelab\">@daisukelab</a> notebooks for preprocessing and <a href=\"/mhiro2\">@mhiro2</a> for his simple CNN model.</p>\n\n<p>Keys points of this solution, 2 techniques I imagined (maybe they exist under another name 😄) \n- <strong>warm-up pipeline</strong> : try to use noisy dataset as pretrain model and as semi-supervised to increase diversity in generating a set of models. In another words, first use noisy set to warmup model training, then fine tune model with curated set and then use noisy set again in a semi-supervised way.\n- <strong>SpecMix</strong> : my new data-augmentation technique which takes what I consider the best from SpecAugment and mixup. It generates new samples by applying frequency replacement and time replacement on inputs and compute a weighted average on targets. More details in the github <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">README</a></p>\n\n<p>Inference is done using an ensemble of 2 models (<a href=\"/mhiro2\">@mhiro2</a> simple CNN model + VGG16) on 10 folds CV.</p>\n\n<p>More details are available in the github <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">README</a>, as well code and weights.</p>",
  "messages": [
    {
      "id": "550652",
      "postDate": "06/11/2019 21:59:21",
      "content": "<p>Hello,</p>\n\n<p>I released on a <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">github repository</a> my solution. It is highly inspired from <a href=\"/daisukelab\">@daisukelab</a> notebooks for preprocessing and <a href=\"/mhiro2\">@mhiro2</a> for his simple CNN model.</p>\n\n<p>Keys points of this solution, 2 techniques I imagined (maybe they exist under another name 😄) \n- <strong>warm-up pipeline</strong> : try to use noisy dataset as pretrain model and as semi-supervised to increase diversity in generating a set of models. In another words, first use noisy set to warmup model training, then fine tune model with curated set and then use noisy set again in a semi-supervised way.\n- <strong>SpecMix</strong> : my new data-augmentation technique which takes what I consider the best from SpecAugment and mixup. It generates new samples by applying frequency replacement and time replacement on inputs and compute a weighted average on targets. More details in the github <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">README</a></p>\n\n<p>Inference is done using an ensemble of 2 models (<a href=\"/mhiro2\">@mhiro2</a> simple CNN model + VGG16) on 10 folds CV.</p>\n\n<p>More details are available in the github <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">README</a>, as well code and weights.</p>",
      "rawMarkdown": "Hello,\n\nI released on a [github repository](https://github.com/ebouteillon/freesound-audio-tagging-2019) my solution. It is highly inspired from @daisukelab notebooks for preprocessing and @mhiro2 for his simple CNN model.\n\nKeys points of this solution, 2 techniques I imagined (maybe they exist under another name 😄) \n- **warm-up pipeline** : try to use noisy dataset as pretrain model and as semi-supervised to increase diversity in generating a set of models. In another words, first use noisy set to warmup model training, then fine tune model with curated set and then use noisy set again in a semi-supervised way.\n- **SpecMix** : my new data-augmentation technique which takes what I consider the best from SpecAugment and mixup. It generates new samples by applying frequency replacement and time replacement on inputs and compute a weighted average on targets. More details in the github [README](https://github.com/ebouteillon/freesound-audio-tagging-2019)\n\nInference is done using an ensemble of 2 models (@mhiro2 simple CNN model + VGG16) on 10 folds CV.\n\nMore details are available in the github [README](https://github.com/ebouteillon/freesound-audio-tagging-2019), as well code and weights.",
      "votes": null
    },
    {
      "id": "550767",
      "postDate": "06/12/2019 02:26:31",
      "content": "<p>Thank for sharing a solid solution!</p>",
      "rawMarkdown": "Thank for sharing a solid solution!",
      "votes": null
    },
    {
      "id": "550789",
      "postDate": "06/12/2019 03:14:01",
      "content": "<p>Thank you for sharing! Your solution is quite simple enough and strong.</p>",
      "rawMarkdown": "Thank you for sharing! Your solution is quite simple enough and strong.",
      "votes": null
    },
    {
      "id": "550806",
      "postDate": "06/12/2019 03:36:28",
      "content": "<p>Thank you for using many of my codes, it's my pleasure, really!  # including it's easy for me to read your code ;)\nAnd your two ideas are really nice and making good performance actually, it's much better than straightforward mine.\nThis is exactly what I was expecting to see. Wishing your good luck in the 2nd stage!</p>",
      "rawMarkdown": "Thank you for using many of my codes, it's my pleasure, really!  # including it's easy for me to read your code ;)\nAnd your two ideas are really nice and making good performance actually, it's much better than straightforward mine.\nThis is exactly what I was expecting to see. Wishing your good luck in the 2nd stage!",
      "votes": null
    },
    {
      "id": "551146",
      "postDate": "06/12/2019 11:43:55",
      "content": "<p>Thank you for your nice words.\nWishing you all the best for the second stage!</p>",
      "rawMarkdown": "Thank you for your nice words.\nWishing you all the best for the second stage!",
      "votes": null
    },
    {
      "id": "551148",
      "postDate": "06/12/2019 11:46:52",
      "content": "<p>It should be me that is thanking you for the model that I borrowed, it helped me a lot :smile:\nWishing you all the best for the second stage!</p>",
      "rawMarkdown": "It should be me that is thanking you for the model that I borrowed, it helped me a lot :smile:\nWishing you all the best for the second stage!",
      "votes": null
    },
    {
      "id": "551149",
      "postDate": "06/12/2019 11:48:33",
      "content": "<p>My code is a bit crappy, I should refactor it and comment it. Thanks for your kernels and nice comment. :smile:\nWishing you all the best for the second stage!</p>",
      "rawMarkdown": "My code is a bit crappy, I should refactor it and comment it. Thanks for your kernels and nice comment. :smile:\nWishing you all the best for the second stage!",
      "votes": null
    },
    {
      "id": "551176",
      "postDate": "06/12/2019 12:29:45",
      "content": "<p>thanks for sharing <a href=\"/ebouteillon\">@ebouteillon</a> .</p>",
      "rawMarkdown": "thanks for sharing @ebouteillon .",
      "votes": null
    },
    {
      "id": "551364",
      "postDate": "06/12/2019 16:21:00",
      "content": "<p>You are welcomed. Good luck!</p>",
      "rawMarkdown": "You are welcomed. Good luck!",
      "votes": null
    },
    {
      "id": "551504",
      "postDate": "06/12/2019 19:46:49",
      "content": "<p>This is fantastic and really well documented. Learned a lot, thanks for sharing. </p>",
      "rawMarkdown": "This is fantastic and really well documented. Learned a lot, thanks for sharing.",
      "votes": null
    },
    {
      "id": "551613",
      "postDate": "06/12/2019 23:46:20",
      "content": "<p>Your kind words are appreciated</p>",
      "rawMarkdown": "Your kind words are appreciated",
      "votes": null
    },
    {
      "id": "551641",
      "postDate": "06/13/2019 00:56:27",
      "content": "<p>Thanks for sharing <a href=\"/ebouteillon\">@ebouteillon</a> , warm-up pipeline really make sense.</p>",
      "rawMarkdown": "Thanks for sharing @ebouteillon , warm-up pipeline really make sense.",
      "votes": null
    },
    {
      "id": "551845",
      "postDate": "06/13/2019 07:04:40",
      "content": "<p>Thank you for your positive feedback. An important point to take care of is not leaking data between folds and stages during the warm-up.\nGood luck for the second stage. </p>",
      "rawMarkdown": "Thank you for your positive feedback. An important point to take care of is not leaking data between folds and stages during the warm-up.\nGood luck for the second stage.",
      "votes": null
    },
    {
      "id": "553546",
      "postDate": "06/15/2019 22:14:21",
      "content": "<p>Hi my Kaggling friends,\nJust to let you know that I updated the <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">README</a> in the repository. I tried to clarify some parts of the description, added instructions to reproduce my results, acknowledgments where due, and of course useless badges... 😄 \nFeedback are of course welcomed.</p>",
      "rawMarkdown": "Hi my Kaggling friends,\nJust to let you know that I updated the [README](https://github.com/ebouteillon/freesound-audio-tagging-2019) in the repository. I tried to clarify some parts of the description, added instructions to reproduce my results, acknowledgments where due, and of course useless badges... 😄 \nFeedback are of course welcomed.",
      "votes": null
    },
    {
      "id": "557622",
      "postDate": "06/21/2019 10:24:38",
      "content": "<p>“Dad, you are the best and you will be at the very top”  . THIS is real motivation !!! Congrats :)</p>",
      "rawMarkdown": "“Dad, you are the best and you will be at the very top”  . THIS is real motivation !!! Congrats :)",
      "votes": null
    },
    {
      "id": "557659",
      "postDate": "06/21/2019 11:18:04",
      "content": "<p>Thank for noticing it. All credits goes to him. 😉</p>",
      "rawMarkdown": "Thank for noticing it. All credits goes to him. 😉",
      "votes": null
    },
    {
      "id": "563922",
      "postDate": "06/28/2019 19:28:59",
      "content": "<p>Note: Updated title and happy to become master 🥳</p>",
      "rawMarkdown": "Note: Updated title and happy to become master 🥳",
      "votes": null
    },
    {
      "id": "573149",
      "postDate": "07/11/2019 21:10:35",
      "content": "<p>thank you so much for sharing</p>",
      "rawMarkdown": "thank you so much for sharing",
      "votes": null
    },
    {
      "id": "573861",
      "postDate": "07/12/2019 21:07:14",
      "content": "<p>You’re welcomed. </p>",
      "rawMarkdown": "You’re welcomed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 550767,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "06/12/2019 02:26:31",
      "content": "<p>Thank for sharing a solid solution!</p>",
      "votes": null,
      "replies": [
        {
          "id": 551146,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "06/12/2019 11:43:55",
          "content": "<p>Thank you for your nice words.\nWishing you all the best for the second stage!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550789,
      "author_name": "mhiro2",
      "author_url": "",
      "post_date": "06/12/2019 03:14:01",
      "content": "<p>Thank you for sharing! Your solution is quite simple enough and strong.</p>",
      "votes": null,
      "replies": [
        {
          "id": 551148,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "06/12/2019 11:46:52",
          "content": "<p>It should be me that is thanking you for the model that I borrowed, it helped me a lot :smile:\nWishing you all the best for the second stage!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550806,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "06/12/2019 03:36:28",
      "content": "<p>Thank you for using many of my codes, it's my pleasure, really!  # including it's easy for me to read your code ;)\nAnd your two ideas are really nice and making good performance actually, it's much better than straightforward mine.\nThis is exactly what I was expecting to see. Wishing your good luck in the 2nd stage!</p>",
      "votes": null,
      "replies": [
        {
          "id": 551149,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "06/12/2019 11:48:33",
          "content": "<p>My code is a bit crappy, I should refactor it and comment it. Thanks for your kernels and nice comment. :smile:\nWishing you all the best for the second stage!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551176,
      "author_name": "karanjakhar",
      "author_url": "",
      "post_date": "06/12/2019 12:29:45",
      "content": "<p>thanks for sharing <a href=\"/ebouteillon\">@ebouteillon</a> .</p>",
      "votes": null,
      "replies": [
        {
          "id": 551364,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "06/12/2019 16:21:00",
          "content": "<p>You are welcomed. Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551504,
      "author_name": "madeupmasters",
      "author_url": "",
      "post_date": "06/12/2019 19:46:49",
      "content": "<p>This is fantastic and really well documented. Learned a lot, thanks for sharing. </p>",
      "votes": null,
      "replies": [
        {
          "id": 551613,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "06/12/2019 23:46:20",
          "content": "<p>Your kind words are appreciated</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551641,
      "author_name": "sailorwei",
      "author_url": "",
      "post_date": "06/13/2019 00:56:27",
      "content": "<p>Thanks for sharing <a href=\"/ebouteillon\">@ebouteillon</a> , warm-up pipeline really make sense.</p>",
      "votes": null,
      "replies": [
        {
          "id": 551845,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "06/13/2019 07:04:40",
          "content": "<p>Thank you for your positive feedback. An important point to take care of is not leaking data between folds and stages during the warm-up.\nGood luck for the second stage. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 553546,
      "author_name": "ebouteillon",
      "author_url": "",
      "post_date": "06/15/2019 22:14:21",
      "content": "<p>Hi my Kaggling friends,\nJust to let you know that I updated the <a href=\"https://github.com/ebouteillon/freesound-audio-tagging-2019\">README</a> in the repository. I tried to clarify some parts of the description, added instructions to reproduce my results, acknowledgments where due, and of course useless badges... 😄 \nFeedback are of course welcomed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 557622,
      "author_name": "nikosdim",
      "author_url": "",
      "post_date": "06/21/2019 10:24:38",
      "content": "<p>“Dad, you are the best and you will be at the very top”  . THIS is real motivation !!! Congrats :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 557659,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "06/21/2019 11:18:04",
          "content": "<p>Thank for noticing it. All credits goes to him. 😉</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 563922,
      "author_name": "ebouteillon",
      "author_url": "",
      "post_date": "06/28/2019 19:28:59",
      "content": "<p>Note: Updated title and happy to become master 🥳</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 573149,
      "author_name": "lingzhikang",
      "author_url": "",
      "post_date": "07/11/2019 21:10:35",
      "content": "<p>thank you so much for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 573861,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "07/12/2019 21:07:14",
          "content": "<p>You’re welcomed. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "550652": "Hello,\n\nI released on a [github repository](https://github.com/ebouteillon/freesound-audio-tagging-2019) my solution. It is highly inspired from @daisukelab notebooks for preprocessing and @mhiro2 for his simple CNN model.\n\nKeys points of this solution, 2 techniques I imagined (maybe they exist under another name 😄) \n- **warm-up pipeline** : try to use noisy dataset as pretrain model and as semi-supervised to increase diversity in generating a set of models. In another words, first use noisy set to warmup model training, then fine tune model with curated set and then use noisy set again in a semi-supervised way.\n- **SpecMix** : my new data-augmentation technique which takes what I consider the best from SpecAugment and mixup. It generates new samples by applying frequency replacement and time replacement on inputs and compute a weighted average on targets. More details in the github [README](https://github.com/ebouteillon/freesound-audio-tagging-2019)\n\nInference is done using an ensemble of 2 models (@mhiro2 simple CNN model + VGG16) on 10 folds CV.\n\nMore details are available in the github [README](https://github.com/ebouteillon/freesound-audio-tagging-2019), as well code and weights.",
    "550767": "Thank for sharing a solid solution!",
    "550789": "Thank you for sharing! Your solution is quite simple enough and strong.",
    "550806": "Thank you for using many of my codes, it's my pleasure, really!  # including it's easy for me to read your code ;)\nAnd your two ideas are really nice and making good performance actually, it's much better than straightforward mine.\nThis is exactly what I was expecting to see. Wishing your good luck in the 2nd stage!",
    "551146": "Thank you for your nice words.\nWishing you all the best for the second stage!",
    "551148": "It should be me that is thanking you for the model that I borrowed, it helped me a lot :smile:\nWishing you all the best for the second stage!",
    "551149": "My code is a bit crappy, I should refactor it and comment it. Thanks for your kernels and nice comment. :smile:\nWishing you all the best for the second stage!",
    "551176": "thanks for sharing @ebouteillon .",
    "551364": "You are welcomed. Good luck!",
    "551504": "This is fantastic and really well documented. Learned a lot, thanks for sharing.",
    "551613": "Your kind words are appreciated",
    "551641": "Thanks for sharing @ebouteillon , warm-up pipeline really make sense.",
    "551845": "Thank you for your positive feedback. An important point to take care of is not leaking data between folds and stages during the warm-up.\nGood luck for the second stage.",
    "553546": "Hi my Kaggling friends,\nJust to let you know that I updated the [README](https://github.com/ebouteillon/freesound-audio-tagging-2019) in the repository. I tried to clarify some parts of the description, added instructions to reproduce my results, acknowledgments where due, and of course useless badges... 😄 \nFeedback are of course welcomed.",
    "557622": "“Dad, you are the best and you will be at the very top”  . THIS is real motivation !!! Congrats :)",
    "557659": "Thank for noticing it. All credits goes to him. 😉",
    "563922": "Note: Updated title and happy to become master 🥳",
    "573149": "thank you so much for sharing",
    "573861": "You’re welcomed."
  },
  "source": "meta"
}