{
  "id": 233453,
  "title": "Making a bachelor's thesis out of BirdClef21",
  "url": "/competitions/birdclef-2021/discussion/233453",
  "author_name": "",
  "post_date": "2021-04-19T11:50:21.488414500Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm studying a Computer Science Bsc. in Denmark and I'm basing my bachelor's thesis on the BirdClef21 competition. I've had two courses on Machine Learning, but they were very basic. Other than that I did a summer project on genre classification of music, but it ended up being mostly a literature review. Half up my thesis time is up and I haven't done much yet. I've written notes for the CEUR Proceedings overviews of the different BirdClef years, but that's it. Right now I'm really overwhelmed by all the code being shared here. </p>\n<p>Me and my supervisor have talked about, that I should probably get a baseline working and from that baseline I could try out some of the things that have been recommended by the community, in papers, etc. to do small improvements over the baseline. And then I'd write my thesis around that. Nothing groundbreaking, no attempts to become top10 or anything.</p>\n<p>This is what I've learned from the CEUR Proceedings Overviews of the latest years:</p>\n<ol>\n<li>CNN is the way to go. </li>\n<li>Very deep networks performed best.</li>\n<li>Sophisticated augmentation strategies are of particular importance since they provide the needed variance to the input data distribution which prevents overfitting, also providing improvements in temporal and frequency domain.<ul>\n<li>addition of background noise</li>\n<li>sounds from other files belonging to the same bird species withrandom intensity, in order to stimulate artificially numerous con-texts.</li>\n<li>including a randomized background noise addition phase makes itmore robot to the diversity of noise encountered in the test data.</li></ul></li>\n<li>use of validation data to fine-tune the pre-trained networks has asignificant effect on the overall scores.</li>\n<li>Inception-based CNNs provide the best performance.</li>\n<li>Ensembles of networks improves performance.</li>\n<li>Smaller FFT window might give a small edge on soundscape, pos-sibly because it’s more likely to catch only one bird singing,  thanseveral, overlapping ones. (But the normals benefit from larger FFTwindow!)</li>\n</ol>\n<p>I guess I would just like to show with my thesis that the addition of some data augmentation or specific approach, boosted the performance over the baseline. But I'm having a bit of anxiety in relation to what I should choose as my baseline. There's the <a href=\"https://github.com/kahst/BirdCLEF-Baseline\" target=\"_blank\">BirdClef 2018 Baseline</a> that was released for participants to improve upon. I guess this would be a nice option to go with, because of simplicity and because it's well documented and structured. However, it's a Lasagne/Theano implementation, and I would really prefer PyTorch (or Tensorflow/Keras) as my courses used this. But I guess I could try and re-write the code. Otherwise, there's e.g. the <a href=\"https://www.kaggle.com/hidehisaarai1213/pytorch-inference-birdclef2021-starter\" target=\"_blank\">BirdClef21 starter provided by hidehisaarai1213</a>. </p>\n<p>Hidehisaarai's starter already does noise augmentation as far as I can see. So I was wondering, if anyone has a suggestion, for what could be an interesting thing to persue in the thesis, that is also something a beginner could do. </p>\n<p>Thanks in advance, sincerely Tobias. </p>",
  "messages": [
    {
      "id": "1277941",
      "postDate": "04/19/2021 11:50:21",
      "content": "<p>Hi everyone,</p>\n<p>I'm studying a Computer Science Bsc. in Denmark and I'm basing my bachelor's thesis on the BirdClef21 competition. I've had two courses on Machine Learning, but they were very basic. Other than that I did a summer project on genre classification of music, but it ended up being mostly a literature review. Half up my thesis time is up and I haven't done much yet. I've written notes for the CEUR Proceedings overviews of the different BirdClef years, but that's it. Right now I'm really overwhelmed by all the code being shared here. </p>\n<p>Me and my supervisor have talked about, that I should probably get a baseline working and from that baseline I could try out some of the things that have been recommended by the community, in papers, etc. to do small improvements over the baseline. And then I'd write my thesis around that. Nothing groundbreaking, no attempts to become top10 or anything.</p>\n<p>This is what I've learned from the CEUR Proceedings Overviews of the latest years:</p>\n<ol>\n<li>CNN is the way to go. </li>\n<li>Very deep networks performed best.</li>\n<li>Sophisticated augmentation strategies are of particular importance since they provide the needed variance to the input data distribution which prevents overfitting, also providing improvements in temporal and frequency domain.<ul>\n<li>addition of background noise</li>\n<li>sounds from other files belonging to the same bird species withrandom intensity, in order to stimulate artificially numerous con-texts.</li>\n<li>including a randomized background noise addition phase makes itmore robot to the diversity of noise encountered in the test data.</li></ul></li>\n<li>use of validation data to fine-tune the pre-trained networks has asignificant effect on the overall scores.</li>\n<li>Inception-based CNNs provide the best performance.</li>\n<li>Ensembles of networks improves performance.</li>\n<li>Smaller FFT window might give a small edge on soundscape, pos-sibly because it’s more likely to catch only one bird singing,  thanseveral, overlapping ones. (But the normals benefit from larger FFTwindow!)</li>\n</ol>\n<p>I guess I would just like to show with my thesis that the addition of some data augmentation or specific approach, boosted the performance over the baseline. But I'm having a bit of anxiety in relation to what I should choose as my baseline. There's the <a href=\"https://github.com/kahst/BirdCLEF-Baseline\" target=\"_blank\">BirdClef 2018 Baseline</a> that was released for participants to improve upon. I guess this would be a nice option to go with, because of simplicity and because it's well documented and structured. However, it's a Lasagne/Theano implementation, and I would really prefer PyTorch (or Tensorflow/Keras) as my courses used this. But I guess I could try and re-write the code. Otherwise, there's e.g. the <a href=\"https://www.kaggle.com/hidehisaarai1213/pytorch-inference-birdclef2021-starter\" target=\"_blank\">BirdClef21 starter provided by hidehisaarai1213</a>. </p>\n<p>Hidehisaarai's starter already does noise augmentation as far as I can see. So I was wondering, if anyone has a suggestion, for what could be an interesting thing to persue in the thesis, that is also something a beginner could do. </p>\n<p>Thanks in advance, sincerely Tobias. </p>",
      "rawMarkdown": "Hi everyone,\n\nI'm studying a Computer Science Bsc. in Denmark and I'm basing my bachelor's thesis on the BirdClef21 competition. I've had two courses on Machine Learning, but they were very basic. Other than that I did a summer project on genre classification of music, but it ended up being mostly a literature review. Half up my thesis time is up and I haven't done much yet. I've written notes for the CEUR Proceedings overviews of the different BirdClef years, but that's it. Right now I'm really overwhelmed by all the code being shared here. \n\nMe and my supervisor have talked about, that I should probably get a baseline working and from that baseline I could try out some of the things that have been recommended by the community, in papers, etc. to do small improvements over the baseline. And then I'd write my thesis around that. Nothing groundbreaking, no attempts to become top10 or anything.\n\nThis is what I've learned from the CEUR Proceedings Overviews of the latest years:\n1. CNN is the way to go. \n2. Very deep networks performed best.\n3. Sophisticated augmentation strategies are of particular importance since they provide the needed variance to the input data distribution which prevents overfitting, also providing improvements in temporal and frequency domain.\n    - addition of background noise\n    - sounds from other files belonging to the same bird species withrandom intensity, in order to stimulate artificially numerous con-texts.\n    - including a randomized background noise addition phase makes itmore robot to the diversity of noise encountered in the test data.\n4. use of validation data to fine-tune the pre-trained networks has asignificant effect on the overall scores.\n5. Inception-based CNNs provide the best performance.\n6. Ensembles of networks improves performance.\n7. Smaller FFT window might give a small edge on soundscape, pos-sibly because it’s more likely to catch only one bird singing,  thanseveral, overlapping ones. (But the normals benefit from larger FFTwindow!)\n\nI guess I would just like to show with my thesis that the addition of some data augmentation or specific approach, boosted the performance over the baseline. But I'm having a bit of anxiety in relation to what I should choose as my baseline. There's the [BirdClef 2018 Baseline](https://github.com/kahst/BirdCLEF-Baseline) that was released for participants to improve upon. I guess this would be a nice option to go with, because of simplicity and because it's well documented and structured. However, it's a Lasagne/Theano implementation, and I would really prefer PyTorch (or Tensorflow/Keras) as my courses used this. But I guess I could try and re-write the code. Otherwise, there's e.g. the [BirdClef21 starter provided by hidehisaarai1213](https://www.kaggle.com/hidehisaarai1213/pytorch-inference-birdclef2021-starter). \n\nHidehisaarai's starter already does noise augmentation as far as I can see. So I was wondering, if anyone has a suggestion, for what could be an interesting thing to persue in the thesis, that is also something a beginner could do. \n\nThanks in advance, sincerely Tobias.",
      "votes": null
    },
    {
      "id": "1278224",
      "postDate": "04/19/2021 17:00:02",
      "content": "<p><a href=\"https://www.kaggle.com/tobiasbonnesen\" target=\"_blank\">@tobiasbonnesen</a> as part of your thesis are you submitting to the competition or just working with this data in github? Just curious if you are required to show a working model. </p>",
      "rawMarkdown": "tobiasbonnesen as part of your thesis are you submitting to the competition or just working with this data in github? Just curious if you are required to show a working model.",
      "votes": null
    },
    {
      "id": "1278259",
      "postDate": "04/19/2021 17:35:38",
      "content": "<p>Thanks for replying Charlie! I'm expected to be able to make a prediction using my model, and report how well I did. The big joker here is of course, that the test dataset is not available outside of Kaggle submissions. So that sort of forces me to participate. Does that answer the question? :-) </p>",
      "rawMarkdown": "Thanks for replying Charlie! I'm expected to be able to make a prediction using my model, and report how well I did. The big joker here is of course, that the test dataset is not available outside of Kaggle submissions. So that sort of forces me to participate. Does that answer the question? :-)",
      "votes": null
    },
    {
      "id": "1278263",
      "postDate": "04/19/2021 17:41:06",
      "content": "<p>I suggest you explore what top teams shared in previous competition.  I'd start with kaggle forum and follow links to additional material whenever possible: <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition</a></p>\n<p>You can also look at a somewhat related competition, Rainforest: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection</a></p>",
      "rawMarkdown": "I suggest you explore what top teams shared in previous competition.  I'd start with kaggle forum and follow links to additional material whenever possible: https://www.kaggle.com/c/birdsong-recognition\n\nYou can also look at a somewhat related competition, Rainforest: https://www.kaggle.com/c/rfcx-species-audio-detection",
      "votes": null
    },
    {
      "id": "1278349",
      "postDate": "04/19/2021 19:27:13",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> That's a great recommendation for this and anyone new to Kaggle 😊</p>",
      "rawMarkdown": "cpmpml That's a great recommendation for this and anyone new to Kaggle 😊",
      "votes": null
    },
    {
      "id": "1278352",
      "postDate": "04/19/2021 19:29:05",
      "content": "<p>Well, there is a test set. Do you mean the holdback of the final leaderboard public score set? </p>",
      "rawMarkdown": "Well, there is a test set. Do you mean the holdback of the final leaderboard public score set?",
      "votes": null
    },
    {
      "id": "1278991",
      "postDate": "04/20/2021 13:38:22",
      "content": "<p>Yes there is, but we agree that it's hidden unless you make a submission, right? 🙂<br>\nFrom <a href=\"https://www.kaggle.com/stefankahl/birdclef2021-sample-submission\" target=\"_blank\">Stefans sample submission notebook here</a>:</p>\n<blockquote>\n  <p>The hidden test set will only appear if you submit the notebook </p>\n</blockquote>",
      "rawMarkdown": "Yes there is, but we agree that it's hidden unless you make a submission, right? 🙂\nFrom [Stefans sample submission notebook here](https://www.kaggle.com/stefankahl/birdclef2021-sample-submission ):\n> The hidden test set will only appear if you submit the notebook",
      "votes": null
    },
    {
      "id": "1279307",
      "postDate": "04/20/2021 19:16:40",
      "content": "<p>If that is a problem for the thesis you could set aside 2 or 3 of the soundscapes in the training data for your own validation set. I am pretty sure this is common practice anyways for papers and would probably give you enough information to write the thesis!</p>",
      "rawMarkdown": "If that is a problem for the thesis you could set aside 2 or 3 of the soundscapes in the training data for your own validation set. I am pretty sure this is common practice anyways for papers and would probably give you enough information to write the thesis!",
      "votes": null
    },
    {
      "id": "1298952",
      "postDate": "05/09/2021 11:03:41",
      "content": "<p>Would be nice if you could try out non-CNN models. Because CNN is guaranteed to be the winner, there is little exploration in other areas. But CNN has its limitations. Sounds aren't images. You can't move the patches around in the images without completely ruining the (associated) sounds. </p>\n<p>This is the line what I would have chosen if I had time off work.. But I will try to work on this even after the comp has ended and see if I can contribute something.</p>",
      "rawMarkdown": "Would be nice if you could try out non-CNN models. Because CNN is guaranteed to be the winner, there is little exploration in other areas. But CNN has its limitations. Sounds aren't images. You can't move the patches around in the images without completely ruining the (associated) sounds. \n\nThis is the line what I would have chosen if I had time off work.. But I will try to work on this even after the comp has ended and see if I can contribute something.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1278224,
      "author_name": "crained",
      "author_url": "",
      "post_date": "04/19/2021 17:00:02",
      "content": "<p><a href=\"https://www.kaggle.com/tobiasbonnesen\" target=\"_blank\">@tobiasbonnesen</a> as part of your thesis are you submitting to the competition or just working with this data in github? Just curious if you are required to show a working model. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1278259,
          "author_name": "tobiasbonnesen",
          "author_url": "",
          "post_date": "04/19/2021 17:35:38",
          "content": "<p>Thanks for replying Charlie! I'm expected to be able to make a prediction using my model, and report how well I did. The big joker here is of course, that the test dataset is not available outside of Kaggle submissions. So that sort of forces me to participate. Does that answer the question? :-) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278352,
          "author_name": "crained",
          "author_url": "",
          "post_date": "04/19/2021 19:29:05",
          "content": "<p>Well, there is a test set. Do you mean the holdback of the final leaderboard public score set? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1278991,
          "author_name": "tobiasbonnesen",
          "author_url": "",
          "post_date": "04/20/2021 13:38:22",
          "content": "<p>Yes there is, but we agree that it's hidden unless you make a submission, right? 🙂<br>\nFrom <a href=\"https://www.kaggle.com/stefankahl/birdclef2021-sample-submission\" target=\"_blank\">Stefans sample submission notebook here</a>:</p>\n<blockquote>\n  <p>The hidden test set will only appear if you submit the notebook </p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279307,
          "author_name": "aarons0",
          "author_url": "",
          "post_date": "04/20/2021 19:16:40",
          "content": "<p>If that is a problem for the thesis you could set aside 2 or 3 of the soundscapes in the training data for your own validation set. I am pretty sure this is common practice anyways for papers and would probably give you enough information to write the thesis!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1278263,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/19/2021 17:41:06",
      "content": "<p>I suggest you explore what top teams shared in previous competition.  I'd start with kaggle forum and follow links to additional material whenever possible: <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition</a></p>\n<p>You can also look at a somewhat related competition, Rainforest: <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1278349,
          "author_name": "crained",
          "author_url": "",
          "post_date": "04/19/2021 19:27:13",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> That's a great recommendation for this and anyone new to Kaggle 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1298952,
      "author_name": "allohvk",
      "author_url": "",
      "post_date": "05/09/2021 11:03:41",
      "content": "<p>Would be nice if you could try out non-CNN models. Because CNN is guaranteed to be the winner, there is little exploration in other areas. But CNN has its limitations. Sounds aren't images. You can't move the patches around in the images without completely ruining the (associated) sounds. </p>\n<p>This is the line what I would have chosen if I had time off work.. But I will try to work on this even after the comp has ended and see if I can contribute something.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1277941": "Hi everyone,\n\nI'm studying a Computer Science Bsc. in Denmark and I'm basing my bachelor's thesis on the BirdClef21 competition. I've had two courses on Machine Learning, but they were very basic. Other than that I did a summer project on genre classification of music, but it ended up being mostly a literature review. Half up my thesis time is up and I haven't done much yet. I've written notes for the CEUR Proceedings overviews of the different BirdClef years, but that's it. Right now I'm really overwhelmed by all the code being shared here. \n\nMe and my supervisor have talked about, that I should probably get a baseline working and from that baseline I could try out some of the things that have been recommended by the community, in papers, etc. to do small improvements over the baseline. And then I'd write my thesis around that. Nothing groundbreaking, no attempts to become top10 or anything.\n\nThis is what I've learned from the CEUR Proceedings Overviews of the latest years:\n1. CNN is the way to go. \n2. Very deep networks performed best.\n3. Sophisticated augmentation strategies are of particular importance since they provide the needed variance to the input data distribution which prevents overfitting, also providing improvements in temporal and frequency domain.\n    - addition of background noise\n    - sounds from other files belonging to the same bird species withrandom intensity, in order to stimulate artificially numerous con-texts.\n    - including a randomized background noise addition phase makes itmore robot to the diversity of noise encountered in the test data.\n4. use of validation data to fine-tune the pre-trained networks has asignificant effect on the overall scores.\n5. Inception-based CNNs provide the best performance.\n6. Ensembles of networks improves performance.\n7. Smaller FFT window might give a small edge on soundscape, pos-sibly because it’s more likely to catch only one bird singing,  thanseveral, overlapping ones. (But the normals benefit from larger FFTwindow!)\n\nI guess I would just like to show with my thesis that the addition of some data augmentation or specific approach, boosted the performance over the baseline. But I'm having a bit of anxiety in relation to what I should choose as my baseline. There's the [BirdClef 2018 Baseline](https://github.com/kahst/BirdCLEF-Baseline) that was released for participants to improve upon. I guess this would be a nice option to go with, because of simplicity and because it's well documented and structured. However, it's a Lasagne/Theano implementation, and I would really prefer PyTorch (or Tensorflow/Keras) as my courses used this. But I guess I could try and re-write the code. Otherwise, there's e.g. the [BirdClef21 starter provided by hidehisaarai1213](https://www.kaggle.com/hidehisaarai1213/pytorch-inference-birdclef2021-starter). \n\nHidehisaarai's starter already does noise augmentation as far as I can see. So I was wondering, if anyone has a suggestion, for what could be an interesting thing to persue in the thesis, that is also something a beginner could do. \n\nThanks in advance, sincerely Tobias.",
    "1278224": "tobiasbonnesen as part of your thesis are you submitting to the competition or just working with this data in github? Just curious if you are required to show a working model.",
    "1278259": "Thanks for replying Charlie! I'm expected to be able to make a prediction using my model, and report how well I did. The big joker here is of course, that the test dataset is not available outside of Kaggle submissions. So that sort of forces me to participate. Does that answer the question? :-)",
    "1278263": "I suggest you explore what top teams shared in previous competition.  I'd start with kaggle forum and follow links to additional material whenever possible: https://www.kaggle.com/c/birdsong-recognition\n\nYou can also look at a somewhat related competition, Rainforest: https://www.kaggle.com/c/rfcx-species-audio-detection",
    "1278349": "cpmpml That's a great recommendation for this and anyone new to Kaggle 😊",
    "1278352": "Well, there is a test set. Do you mean the holdback of the final leaderboard public score set?",
    "1278991": "Yes there is, but we agree that it's hidden unless you make a submission, right? 🙂\nFrom [Stefans sample submission notebook here](https://www.kaggle.com/stefankahl/birdclef2021-sample-submission ):\n> The hidden test set will only appear if you submit the notebook",
    "1279307": "If that is a problem for the thesis you could set aside 2 or 3 of the soundscapes in the training data for your own validation set. I am pretty sure this is common practice anyways for papers and would probably give you enough information to write the thesis!",
    "1298952": "Would be nice if you could try out non-CNN models. Because CNN is guaranteed to be the winner, there is little exploration in other areas. But CNN has its limitations. Sounds aren't images. You can't move the patches around in the images without completely ruining the (associated) sounds. \n\nThis is the line what I would have chosen if I had time off work.. But I will try to work on this even after the comp has ended and see if I can contribute something."
  },
  "source": "meta"
}