{
  "id": 96680,
  "title": "6th place solution fastai",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/miguel-pinto-6th-place-solution-fastai",
  "author_name": "",
  "post_date": "2019-06-28T16:13:45.120Z",
  "votes": 27,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Here are the write-up and code for my solution!</p>\n\n<p><strong>Blog post:</strong> <a href=\"https://link.medium.com/Kv5kyHjcIX\">https://link.medium.com/Kv5kyHjcIX</a>\n<strong>Code:</strong> <a href=\"https://github.com/mnpinto/audiotagging2019\">https://github.com/mnpinto/audiotagging2019</a></p>\n\n<p><strong>Summary:</strong>\n* Models: xresnets\n* Image size: 256x256\n* Mixup sampling from a uniform distribution\n* Horizontal and Vertical Flip as new labels (total 320 labels)\n* Compute loss only for samples with F2 score (with a threshold of 0.2) less than 1.\n* Noisy data: ~3500 \"good noisy samples\" used the same way as curated data\n* TTA: Slice clips each 128px in the time axis (no overlap), generate predictions for each slice and compute the <code>max</code> for each class.\n* Final submissions: 1) Average of 2 models: public LB 0.742, private LB 0.74620; 2) Average of 6 models: public LB 0.742, private LB 0.75421.</p>\n\n<p><strong>Additional observations:</strong>\n* I found better results when using random crops of 128x128 and rescale to 256x256, comparing to random crops of 128x256 and rescale to 256x256, I was expecting the opposite.\n* I wonder why max_zoom=1.5 works; I would not expect so.</p>\n\n<p><strong>Acknowledgements:</strong>\n<a href=\"/daisukelab\">@daisukelab</a> thanks for the code to generate the mel spectrograms! Thanks to everyone that contributed in the discussions or with kernels. And finally, thanks to the organizers for this great competition!</p>",
  "messages": [
    {
      "id": "558161",
      "postDate": "06/21/2019 22:46:35",
      "content": "<p>Here are the write-up and code for my solution!</p>\n\n<p><strong>Blog post:</strong> <a href=\"https://link.medium.com/Kv5kyHjcIX\">https://link.medium.com/Kv5kyHjcIX</a>\n<strong>Code:</strong> <a href=\"https://github.com/mnpinto/audiotagging2019\">https://github.com/mnpinto/audiotagging2019</a></p>\n\n<p><strong>Summary:</strong>\n* Models: xresnets\n* Image size: 256x256\n* Mixup sampling from a uniform distribution\n* Horizontal and Vertical Flip as new labels (total 320 labels)\n* Compute loss only for samples with F2 score (with a threshold of 0.2) less than 1.\n* Noisy data: ~3500 \"good noisy samples\" used the same way as curated data\n* TTA: Slice clips each 128px in the time axis (no overlap), generate predictions for each slice and compute the <code>max</code> for each class.\n* Final submissions: 1) Average of 2 models: public LB 0.742, private LB 0.74620; 2) Average of 6 models: public LB 0.742, private LB 0.75421.</p>\n\n<p><strong>Additional observations:</strong>\n* I found better results when using random crops of 128x128 and rescale to 256x256, comparing to random crops of 128x256 and rescale to 256x256, I was expecting the opposite.\n* I wonder why max_zoom=1.5 works; I would not expect so.</p>\n\n<p><strong>Acknowledgements:</strong>\n<a href=\"/daisukelab\">@daisukelab</a> thanks for the code to generate the mel spectrograms! Thanks to everyone that contributed in the discussions or with kernels. And finally, thanks to the organizers for this great competition!</p>",
      "rawMarkdown": "Here are the write-up and code for my solution!\n\n**Blog post:** https://link.medium.com/Kv5kyHjcIX\n**Code:** [https://github.com/mnpinto/audiotagging2019](https://github.com/mnpinto/audiotagging2019)\n\n**Summary:**\n* Models: xresnets\n* Image size: 256x256\n* Mixup sampling from a uniform distribution\n* Horizontal and Vertical Flip as new labels (total 320 labels)\n* Compute loss only for samples with F2 score (with a threshold of 0.2) less than 1.\n* Noisy data: ~3500 \"good noisy samples\" used the same way as curated data\n* TTA: Slice clips each 128px in the time axis (no overlap), generate predictions for each slice and compute the `max` for each class.\n* Final submissions: 1) Average of 2 models: public LB 0.742, private LB 0.74620; 2) Average of 6 models: public LB 0.742, private LB 0.75421.\n\n**Additional observations:**\n* I found better results when using random crops of 128x128 and rescale to 256x256, comparing to random crops of 128x256 and rescale to 256x256, I was expecting the opposite.\n* I wonder why max_zoom=1.5 works; I would not expect so.\n\n**Acknowledgements:**\n@daisukelab thanks for the code to generate the mel spectrograms! Thanks to everyone that contributed in the discussions or with kernels. And finally, thanks to the organizers for this great competition!",
      "votes": null
    },
    {
      "id": "558172",
      "postDate": "06/21/2019 23:14:02",
      "content": "<p>Hello <a href=\"/mnpinto\">@mnpinto</a> ,\nit seems attention layer gave you a boost like <a href=\"/romul0212\">@romul0212</a> . I will probably try it in next competition. \nNice write up on medium. 👍\nThanks</p>",
      "rawMarkdown": "Hello @mnpinto ,\nit seems attention layer gave you a boost like @romul0212 . I will probably try it in next competition. \nNice write up on medium. 👍\nThanks",
      "votes": null
    },
    {
      "id": "558346",
      "postDate": "06/22/2019 07:24:32",
      "content": "<p>How much did your score improve with \"Horizontal and Vertical Flip as new labels (total 320 labels)\"?</p>",
      "rawMarkdown": "How much did your score improve with \"Horizontal and Vertical Flip as new labels (total 320 labels)\"?",
      "votes": null
    },
    {
      "id": "558403",
      "postDate": "06/22/2019 09:09:10",
      "content": "<p>Thanks, <a href=\"/ebouteillon\">@ebouteillon</a>! And congrats for your great results also! The attention layer makes the training slower though. I will try to understand the performance of different models better after the competition is finalised.</p>",
      "rawMarkdown": "Thanks, @ebouteillon! And congrats for your great results also! The attention layer makes the training slower though. I will try to understand the performance of different models better after the competition is finalised.",
      "votes": null
    },
    {
      "id": "558411",
      "postDate": "06/22/2019 09:17:09",
      "content": "<p><a href=\"/action\">@action</a>, I'm not sure how much it did improve, but after the competition is finalised, I will check that and give some feedback. Perhaps what I also need to check is how does this compare with ensembling four models, one for each \"flip type\".</p>",
      "rawMarkdown": "action, I'm not sure how much it did improve, but after the competition is finalised, I will check that and give some feedback. Perhaps what I also need to check is how does this compare with ensembling four models, one for each \"flip type\".",
      "votes": null
    },
    {
      "id": "591545",
      "postDate": "08/03/2019 22:29:06",
      "content": "<p>Congratulations <a href=\"/mnpinto\">@mnpinto</a> and thanks for sharing your solution with nice medium write-up!\n(And you are welcome, thanks for using my code.)\nI'm impressed with your approaches, and especially:\n- Adding new labels for flipped clips, great idea.\n- Curriculum learning.</p>",
      "rawMarkdown": "Congratulations @mnpinto and thanks for sharing your solution with nice medium write-up!\n(And you are welcome, thanks for using my code.)\nI'm impressed with your approaches, and especially:\n- Adding new labels for flipped clips, great idea.\n- Curriculum learning.",
      "votes": null
    },
    {
      "id": "623756",
      "postDate": "09/11/2019 09:21:12",
      "content": "<p>Thanks for publishing this, it's helped me a lot with a personal project that I have been working on :)</p>\n\n<p>I was hoping you could explain a little bit how you decided on your AudioMixup implementation? What made you go with a uniform distribution for lambda rather than beta distribution that is used in the stock MixUp callback in the Fast AI library (which I am assuming you modified from).</p>\n\n<p>I'm experimenting between the two and don't want to reinvent the wheel :)</p>\n\n<p>Thank you very much!</p>",
      "rawMarkdown": "Thanks for publishing this, it's helped me a lot with a personal project that I have been working on :)\n\nI was hoping you could explain a little bit how you decided on your AudioMixup implementation? What made you go with a uniform distribution for lambda rather than beta distribution that is used in the stock MixUp callback in the Fast AI library (which I am assuming you modified from).\n\nI'm experimenting between the two and don't want to reinvent the wheel :)\n\nThank you very much!",
      "votes": null
    },
    {
      "id": "623792",
      "postDate": "09/11/2019 10:04:42",
      "content": "<p><a href=\"/daisukelab\">@daisukelab</a> Thanks! Your code was a great starting point for many people in this competition.</p>",
      "rawMarkdown": "daisukelab Thanks! Your code was a great starting point for many people in this competition.",
      "votes": null
    },
    {
      "id": "623799",
      "postDate": "09/11/2019 10:13:59",
      "content": "<p><a href=\"/zacheism\">@zacheism</a> You are welcome, I'm glad my solution did help on your project. I choose a uniform distribution because for audio spectrograms combining clips seems more natural than combining regular pictures. But I think I didn't compare the two, at least in my final models. I would be interested to know if you find differences between the two approaches in your experiments.</p>",
      "rawMarkdown": "zacheism You are welcome, I'm glad my solution did help on your project. I choose a uniform distribution because for audio spectrograms combining clips seems more natural than combining regular pictures. But I think I didn't compare the two, at least in my final models. I would be interested to know if you find differences between the two approaches in your experiments.",
      "votes": null
    },
    {
      "id": "623999",
      "postDate": "09/11/2019 13:47:51",
      "content": "<p>Appreciate the quick response! Okay yea I suppose it makes sense intuitively -- I'll test them both anyway and let ya know how it goes. Thanks again!</p>",
      "rawMarkdown": "Appreciate the quick response! Okay yea I suppose it makes sense intuitively -- I'll test them both anyway and let ya know how it goes. Thanks again!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 558172,
      "author_name": "ebouteillon",
      "author_url": "",
      "post_date": "06/21/2019 23:14:02",
      "content": "<p>Hello <a href=\"/mnpinto\">@mnpinto</a> ,\nit seems attention layer gave you a boost like <a href=\"/romul0212\">@romul0212</a> . I will probably try it in next competition. \nNice write up on medium. 👍\nThanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 558403,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "06/22/2019 09:09:10",
          "content": "<p>Thanks, <a href=\"/ebouteillon\">@ebouteillon</a>! And congrats for your great results also! The attention layer makes the training slower though. I will try to understand the performance of different models better after the competition is finalised.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 558346,
      "author_name": "action",
      "author_url": "",
      "post_date": "06/22/2019 07:24:32",
      "content": "<p>How much did your score improve with \"Horizontal and Vertical Flip as new labels (total 320 labels)\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 558411,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "06/22/2019 09:17:09",
          "content": "<p><a href=\"/action\">@action</a>, I'm not sure how much it did improve, but after the competition is finalised, I will check that and give some feedback. Perhaps what I also need to check is how does this compare with ensembling four models, one for each \"flip type\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 591545,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "08/03/2019 22:29:06",
      "content": "<p>Congratulations <a href=\"/mnpinto\">@mnpinto</a> and thanks for sharing your solution with nice medium write-up!\n(And you are welcome, thanks for using my code.)\nI'm impressed with your approaches, and especially:\n- Adding new labels for flipped clips, great idea.\n- Curriculum learning.</p>",
      "votes": null,
      "replies": [
        {
          "id": 623792,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "09/11/2019 10:04:42",
          "content": "<p><a href=\"/daisukelab\">@daisukelab</a> Thanks! Your code was a great starting point for many people in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 623756,
      "author_name": "zacheism",
      "author_url": "",
      "post_date": "09/11/2019 09:21:12",
      "content": "<p>Thanks for publishing this, it's helped me a lot with a personal project that I have been working on :)</p>\n\n<p>I was hoping you could explain a little bit how you decided on your AudioMixup implementation? What made you go with a uniform distribution for lambda rather than beta distribution that is used in the stock MixUp callback in the Fast AI library (which I am assuming you modified from).</p>\n\n<p>I'm experimenting between the two and don't want to reinvent the wheel :)</p>\n\n<p>Thank you very much!</p>",
      "votes": null,
      "replies": [
        {
          "id": 623799,
          "author_name": "mnpinto",
          "author_url": "",
          "post_date": "09/11/2019 10:13:59",
          "content": "<p><a href=\"/zacheism\">@zacheism</a> You are welcome, I'm glad my solution did help on your project. I choose a uniform distribution because for audio spectrograms combining clips seems more natural than combining regular pictures. But I think I didn't compare the two, at least in my final models. I would be interested to know if you find differences between the two approaches in your experiments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 623999,
          "author_name": "zacheism",
          "author_url": "",
          "post_date": "09/11/2019 13:47:51",
          "content": "<p>Appreciate the quick response! Okay yea I suppose it makes sense intuitively -- I'll test them both anyway and let ya know how it goes. Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "558161": "Here are the write-up and code for my solution!\n\n**Blog post:** https://link.medium.com/Kv5kyHjcIX\n**Code:** [https://github.com/mnpinto/audiotagging2019](https://github.com/mnpinto/audiotagging2019)\n\n**Summary:**\n* Models: xresnets\n* Image size: 256x256\n* Mixup sampling from a uniform distribution\n* Horizontal and Vertical Flip as new labels (total 320 labels)\n* Compute loss only for samples with F2 score (with a threshold of 0.2) less than 1.\n* Noisy data: ~3500 \"good noisy samples\" used the same way as curated data\n* TTA: Slice clips each 128px in the time axis (no overlap), generate predictions for each slice and compute the `max` for each class.\n* Final submissions: 1) Average of 2 models: public LB 0.742, private LB 0.74620; 2) Average of 6 models: public LB 0.742, private LB 0.75421.\n\n**Additional observations:**\n* I found better results when using random crops of 128x128 and rescale to 256x256, comparing to random crops of 128x256 and rescale to 256x256, I was expecting the opposite.\n* I wonder why max_zoom=1.5 works; I would not expect so.\n\n**Acknowledgements:**\n@daisukelab thanks for the code to generate the mel spectrograms! Thanks to everyone that contributed in the discussions or with kernels. And finally, thanks to the organizers for this great competition!",
    "558172": "Hello @mnpinto ,\nit seems attention layer gave you a boost like @romul0212 . I will probably try it in next competition. \nNice write up on medium. 👍\nThanks",
    "558346": "How much did your score improve with \"Horizontal and Vertical Flip as new labels (total 320 labels)\"?",
    "558403": "Thanks, @ebouteillon! And congrats for your great results also! The attention layer makes the training slower though. I will try to understand the performance of different models better after the competition is finalised.",
    "558411": "action, I'm not sure how much it did improve, but after the competition is finalised, I will check that and give some feedback. Perhaps what I also need to check is how does this compare with ensembling four models, one for each \"flip type\".",
    "591545": "Congratulations @mnpinto and thanks for sharing your solution with nice medium write-up!\n(And you are welcome, thanks for using my code.)\nI'm impressed with your approaches, and especially:\n- Adding new labels for flipped clips, great idea.\n- Curriculum learning.",
    "623756": "Thanks for publishing this, it's helped me a lot with a personal project that I have been working on :)\n\nI was hoping you could explain a little bit how you decided on your AudioMixup implementation? What made you go with a uniform distribution for lambda rather than beta distribution that is used in the stock MixUp callback in the Fast AI library (which I am assuming you modified from).\n\nI'm experimenting between the two and don't want to reinvent the wheel :)\n\nThank you very much!",
    "623792": "daisukelab Thanks! Your code was a great starting point for many people in this competition.",
    "623799": "zacheism You are welcome, I'm glad my solution did help on your project. I choose a uniform distribution because for audio spectrograms combining clips seems more natural than combining regular pictures. But I think I didn't compare the two, at least in my final models. I would be interested to know if you find differences between the two approaches in your experiments.",
    "623999": "Appreciate the quick response! Okay yea I suppose it makes sense intuitively -- I'll test them both anyway and let ya know how it goes. Thanks again!"
  },
  "source": "meta"
}