{
  "id": 511419,
  "title": "What I learned in BirdCLEF2024",
  "url": "/competitions/birdclef-2024/discussion/511419",
  "author_name": "Cody_Null",
  "post_date": "2024-06-10T16:15:42.029000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi all, as we bring this comp to a close I thought I would share some of what I learned. Hopefully we can use this as an opportunity to focus on growth and distract ourselves from what will no doubt be a shaky leaderboard. Sadly this competition had quite a difference from training to testing data, unlike previous years, to my knowledge. Therefore this one could be interesting as many have indicated they think they have massively overfit to the public LB. </p>\n<ol>\n<li><p>More Pytorch familiarity - I always like to try new things in competition and this is fortunately my smoothest notebook yet. So I am very pleased to be ale to deliver a custom pipeline from scratch that hopefully shakes up. </p></li>\n<li><p>Inference optimization - I spent a ton of time trying to optimize original scripts for speed but sadly did not have a ton of luck. I did however make a ton of progress when rebuilding my pipeline to be compatible with ONNX, which I will admit was not super intuitive and I am very excited to hear more about how everyone did this.</p></li>\n<li><p>Audio processing - Due to moving over to ONNX I did need to redo much of my audio processing as my original  method of processing didnt work. I also tried many methods in order to be able to extract more signal from the data but it was not significant use past what the more common methods are. </p></li>\n<li><p>Data loaders - I have not had to work with a ton of image data many times outside Kaggle so this is not something I was very good at before but now I think I am quite comfortable doing data loading and processing from scratch with this :)</p></li>\n<li><p>Lots of patience - This competition was quite annoying because of the lack of correlation between the CV and the LB. I tried many many things that did not work significantly because of this so it was a growing point I guess? I will say I am a little confused on why the noisy data in test  is not also just noisy in train. It has been referenced a few times that there is much more going on in the background of the test data. In my eyes it would be more useful for both us and hosts if the training data was purely representative of the test data. </p></li>\n</ol>\n<p>I am sure there is more to highlight here but I dont want to give away too much before the end of the competition. I do have some cool insight on trying different models and how I was able to get some non standard models to work. Good luck in the private LB!</p>",
  "messages": [
    {
      "id": 2865260,
      "postDate": "2024-06-10T16:15:42.030Z",
      "content": "<p>Hi all, as we bring this comp to a close I thought I would share some of what I learned. Hopefully we can use this as an opportunity to focus on growth and distract ourselves from what will no doubt be a shaky leaderboard. Sadly this competition had quite a difference from training to testing data, unlike previous years, to my knowledge. Therefore this one could be interesting as many have indicated they think they have massively overfit to the public LB. </p>\n<ol>\n<li><p>More Pytorch familiarity - I always like to try new things in competition and this is fortunately my smoothest notebook yet. So I am very pleased to be ale to deliver a custom pipeline from scratch that hopefully shakes up. </p></li>\n<li><p>Inference optimization - I spent a ton of time trying to optimize original scripts for speed but sadly did not have a ton of luck. I did however make a ton of progress when rebuilding my pipeline to be compatible with ONNX, which I will admit was not super intuitive and I am very excited to hear more about how everyone did this.</p></li>\n<li><p>Audio processing - Due to moving over to ONNX I did need to redo much of my audio processing as my original  method of processing didnt work. I also tried many methods in order to be able to extract more signal from the data but it was not significant use past what the more common methods are. </p></li>\n<li><p>Data loaders - I have not had to work with a ton of image data many times outside Kaggle so this is not something I was very good at before but now I think I am quite comfortable doing data loading and processing from scratch with this :)</p></li>\n<li><p>Lots of patience - This competition was quite annoying because of the lack of correlation between the CV and the LB. I tried many many things that did not work significantly because of this so it was a growing point I guess? I will say I am a little confused on why the noisy data in test  is not also just noisy in train. It has been referenced a few times that there is much more going on in the background of the test data. In my eyes it would be more useful for both us and hosts if the training data was purely representative of the test data. </p></li>\n</ol>\n<p>I am sure there is more to highlight here but I dont want to give away too much before the end of the competition. I do have some cool insight on trying different models and how I was able to get some non standard models to work. Good luck in the private LB!</p>",
      "rawMarkdown": "Hi all, as we bring this comp to a close I thought I would share some of what I learned. Hopefully we can use this as an opportunity to focus on growth and distract ourselves from what will no doubt be a shaky leaderboard. Sadly this competition had quite a difference from training to testing data, unlike previous years, to my knowledge. Therefore this one could be interesting as many have indicated they think they have massively overfit to the public LB. \n\n1. More Pytorch familiarity - I always like to try new things in competition and this is fortunately my smoothest notebook yet. So I am very pleased to be ale to deliver a custom pipeline from scratch that hopefully shakes up. \n\n2. Inference optimization - I spent a ton of time trying to optimize original scripts for speed but sadly did not have a ton of luck. I did however make a ton of progress when rebuilding my pipeline to be compatible with ONNX, which I will admit was not super intuitive and I am very excited to hear more about how everyone did this.\n\n3. Audio processing - Due to moving over to ONNX I did need to redo much of my audio processing as my original  method of processing didnt work. I also tried many methods in order to be able to extract more signal from the data but it was not significant use past what the more common methods are. \n\n4. Data loaders - I have not had to work with a ton of image data many times outside Kaggle so this is not something I was very good at before but now I think I am quite comfortable doing data loading and processing from scratch with this :)\n\n5. Lots of patience - This competition was quite annoying because of the lack of correlation between the CV and the LB. I tried many many things that did not work significantly because of this so it was a growing point I guess? I will say I am a little confused on why the noisy data in test  is not also just noisy in train. It has been referenced a few times that there is much more going on in the background of the test data. In my eyes it would be more useful for both us and hosts if the training data was purely representative of the test data. \n\nI am sure there is more to highlight here but I dont want to give away too much before the end of the competition. I do have some cool insight on trying different models and how I was able to get some non standard models to work. Good luck in the private LB!\n\n",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2865260": "Hi all, as we bring this comp to a close I thought I would share some of what I learned. Hopefully we can use this as an opportunity to focus on growth and distract ourselves from what will no doubt be a shaky leaderboard. Sadly this competition had quite a difference from training to testing data, unlike previous years, to my knowledge. Therefore this one could be interesting as many have indicated they think they have massively overfit to the public LB. \n\n1. More Pytorch familiarity - I always like to try new things in competition and this is fortunately my smoothest notebook yet. So I am very pleased to be ale to deliver a custom pipeline from scratch that hopefully shakes up. \n\n2. Inference optimization - I spent a ton of time trying to optimize original scripts for speed but sadly did not have a ton of luck. I did however make a ton of progress when rebuilding my pipeline to be compatible with ONNX, which I will admit was not super intuitive and I am very excited to hear more about how everyone did this.\n\n3. Audio processing - Due to moving over to ONNX I did need to redo much of my audio processing as my original  method of processing didnt work. I also tried many methods in order to be able to extract more signal from the data but it was not significant use past what the more common methods are. \n\n4. Data loaders - I have not had to work with a ton of image data many times outside Kaggle so this is not something I was very good at before but now I think I am quite comfortable doing data loading and processing from scratch with this :)\n\n5. Lots of patience - This competition was quite annoying because of the lack of correlation between the CV and the LB. I tried many many things that did not work significantly because of this so it was a growing point I guess? I will say I am a little confused on why the noisy data in test  is not also just noisy in train. It has been referenced a few times that there is much more going on in the background of the test data. In my eyes it would be more useful for both us and hosts if the training data was purely representative of the test data. \n\nI am sure there is more to highlight here but I dont want to give away too much before the end of the competition. I do have some cool insight on trying different models and how I was able to get some non standard models to work. Good luck in the private LB!\n\n"
  }
}