{
  "id": 243927,
  "title": "[1st Place] Detailed Solution",
  "url": "/competitions/birdclef-2021/discussion/243927",
  "author_name": "",
  "post_date": "2021-06-04T13:49:39.019746Z",
  "votes": 77,
  "comment_count": 16,
  "views": 0,
  "content": "<h1>tl;dr</h1>\n<p>My teammate <a href=\"https://www.kaggle.com/kami634\" target=\"_blank\">@kami634</a> has already posted a quick version of our solution. Our solution pipeline is composed of three stages and you can overview the stages info here.<br>\n<a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243304\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/243304</a><br>\nAlso, this is an image of our pipeline. (only inference part, so some of the steps including the first stage are omitted.)</p>\n<p><img src=\"https://drive.google.com/uc?id=1teMrPuBoKeCEf5fFpxoNpzD4fi1IgyxX\" alt=\"img_name\"></p>\n<p>The complete code used in this competition has been uploaded to the following github. All of the ipynb files there have been confirmed to work properly in the current kaggle notebook environment (2021/6/4).<br>\n<a href=\"https://github.com/namakemono/kaggle-birdclef-2021\" target=\"_blank\">https://github.com/namakemono/kaggle-birdclef-2021</a></p>\n<p><br></p>\n<h1>Detailed Solution</h1>\n<p>The first thing we did when we joined this competition was to find out the percentage of nocalls in train_soundscapes and test_soundscapes. In train_soundscapes, it is easy to find out, and in test_soundscapes (Public LB), it can be found by submitting all lines as nocall. The results were 0.637 for train_soundscapes and 0.54 for Public LB. Similarly, in the 2020 competition, we submitted all lines as nocall as late submission, and the results were 0.577 for Public LB and 0.544 for Private LB. In all cases, the majority of the targets were nocalls, and we speculated that the difference between birdcalls by bird species and the difference between birds singing and nocall might be qualitatively different. That’s why we considered nocall detection and bird identification as separate tasks. This was the origin of the binary nocall detector from melspectrograms(1). Initially, we came up with the flow to use the nocall detector to eliminate nocalls for sure, and then predict some birds with nocall labels disabled for the rest, but this did not work. The next idea was to modify the weak labels when building a multilabel classifier, in other words, multiplying the call probabilities obtained from the nocall detector by the labels of train_short_audio. This worked well.</p>\n<p><br></p>\n<p>Next, we built a multilabel classifier (backbone was resnest) for the melspectrogram as the second stage, but there were following four problems at this stage.</p>\n<p><br></p>\n<p>A. The labels of train_short_audio were weak, so it was not clear whether primary or secondary labeled birds were sounding in each frame.<br>\nB. It was possible that the labels itself were wrong.<br>\nC. The output was 397-dimensional probability vectors, and the measure of whether some bird was singing or not was unclear.<br>\nD. Information before and after the current frame might be meaningful.</p>\n<p><br></p>\n<p>To deal with these problems, we created another table competition by ourselves, using the results of the nocall detector, train_metadata, and time series information before and after the current frame. The procedure for creating this table competition data was as follows:</p>\n<p><br></p>\n<p>Ⅰ. Extract the top N candidates in each frame as rows from the results of the multilabel classifier.<br>\nⅡ. Among the N candidates, assigned label 1 to the common set with primary or secondary label, and 0 to the others. However, if the output of the nocall detector indicated that the possibilites of birds singing were low, all the candidates were set to 0. This would be the target variable.<br>\nⅢ. For each candidate, the meta data (time and location information) of the frame and the probabilities that the candidate bird was singing before and after the frame (output of multilabel classifier) were assigned. In addition, more features were added by feature engineering (2), and the table was completed. Then It was analyzed by lightGBM.</p>\n<p><br></p>\n<p>The above steps were expected to have the following effects on the issues A to D mentioned above.</p>\n<p><br></p>\n<p>A. Even if the given data has only a weak label, it is possible to reassign the label for each frame.<br>\nB. The noisy labels can be somewhat cancelled.<br>\nC. The output of the nocall detector can be taken into account.<br>\nD. Information before and after the current frame can be incorporated.</p>\n<p><br></p>\n<p>In addition, by incorporating metadata into the table, it was not necessary to create multiple melspectrogram multilabel classifiers for each region or season, which saved time and computational resources. In fact, in our first-place submission, we used only 15 minutes out of the 3-hour time limit. The fact that we were able to compete based on colab pro might be due to the fact that we succeeded in making this audio competition come down to the table competition.</p>\n<p><br></p>\n<p>The output of our table competition was one probability value per candidate (5 candidates x number of frames in total). The next task was to find an appropriate threshold value for these. In this case, we used train_soundscapes as a reference and searched for the value that maximizes the F1 score by the ternary search. In this process, we hypothesized that the threshold would be different depending on the percentage of nocalls in train_soundscapes, and experimented by reducing the percentage of nocalls, but the conclusion was that the threshold did not change much.</p>\n<p><br></p>\n<p>Finally, we would like to propose a technique we call “nocall injection.” In the current algorithm, rows predicted as “nocall” does not intersected with rows in which some birds were predicted, and vice versa. However, due to the nature of the F1 score, if you predict only one out of the two correct answers, you get 2/3 o<a href=\"url\" target=\"_blank\"></a>f the score, and if you miss the only one label in some row, you get 0. Both are the same in that only one label is wrong, but the size of the penalty is different. Therefore, in order to make sure that nocalls are caught, we added “nocall” to test frames with a high possibility of being nocalls (That would be determined from the output of lightgbm), even if they had been already labeled as some birds. We called this nocall injection. As a result, the nocall and the bird appeared in the same row, and that row could not get the full F1 score, but the benefit of capturing the nocall outweighed it, so the score increased.</p>\n<p><br></p>\n<p>Sorry for writing such a long thread but that's our 3-week activities. It was our first time to compete in a sound recognition competition and we didn't know anything about it at first, but <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>, <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and many others helped us a lot with their comments and code. I would like to take this opportunity to thank them again.</p>\n<p><br></p>\n<hr>\n<h1>Note</h1>\n<p>(1) The nocall detector took the external data (freefield1010) as input because the data from train_soundscapes was not enough. The accuracy was also improved by using data augmentation for images.</p>\n<p><br></p>\n<p>(2) The features we created were as follows.<br>\n  \"year\": The year when the audio was recorded.<br>\n  \"month\": The month in which the audio was recorded.<br>\n  \"sum_prob\": The sum of the probability values of the 397 dimensions for the frame.<br>\n  \"mean_prob\": Mean of the 397 dimensional probability values for the frame.<br>\n  \"max_prob\": Maximum 397-dimensional probability values for the frame.<br>\n  \"prev_prob\" ~ \"prev6_prob\": Probability values for the candidate birds in the previous n frames.<br>\n  \"prob\": Probability value for the candidate bird in the target frame.<br>\n  \"next_prob\" ~ \"next6_prob\": Probability values of the candidate bird in the next n frames.<br>\n  \"rank\": The probability that the bird of interest has the highest probability value out of 397 in the frame.<br>\n  \"latitude\":<br>\n  \"longitude\":<br>\n  \"bird_id\": Which of the 397 birds is the target.<br>\n  \"seconds\": The number of seconds in the frame.<br>\n  \"num_appear\": The number of files where the bird is the primary label in train_short_audio<br>\n  \"site_num_appear\": how many times the bird in train_short_audio appeared in the region<br>\n  \"site_appear_ratio\": site_num_appear/num_appear<br>\n  \"prob_diff\": Difference between prob and average of prob values for 3 frames </p>",
  "messages": [
    {
      "id": "1335887",
      "postDate": "06/04/2021 13:49:39",
      "content": "<h1>tl;dr</h1>\n<p>My teammate <a href=\"https://www.kaggle.com/kami634\" target=\"_blank\">@kami634</a> has already posted a quick version of our solution. Our solution pipeline is composed of three stages and you can overview the stages info here.<br>\n<a href=\"https://www.kaggle.com/c/birdclef-2021/discussion/243304\" target=\"_blank\">https://www.kaggle.com/c/birdclef-2021/discussion/243304</a><br>\nAlso, this is an image of our pipeline. (only inference part, so some of the steps including the first stage are omitted.)</p>\n<p><img src=\"https://drive.google.com/uc?id=1teMrPuBoKeCEf5fFpxoNpzD4fi1IgyxX\" alt=\"img_name\"></p>\n<p>The complete code used in this competition has been uploaded to the following github. All of the ipynb files there have been confirmed to work properly in the current kaggle notebook environment (2021/6/4).<br>\n<a href=\"https://github.com/namakemono/kaggle-birdclef-2021\" target=\"_blank\">https://github.com/namakemono/kaggle-birdclef-2021</a></p>\n<p><br></p>\n<h1>Detailed Solution</h1>\n<p>The first thing we did when we joined this competition was to find out the percentage of nocalls in train_soundscapes and test_soundscapes. In train_soundscapes, it is easy to find out, and in test_soundscapes (Public LB), it can be found by submitting all lines as nocall. The results were 0.637 for train_soundscapes and 0.54 for Public LB. Similarly, in the 2020 competition, we submitted all lines as nocall as late submission, and the results were 0.577 for Public LB and 0.544 for Private LB. In all cases, the majority of the targets were nocalls, and we speculated that the difference between birdcalls by bird species and the difference between birds singing and nocall might be qualitatively different. That’s why we considered nocall detection and bird identification as separate tasks. This was the origin of the binary nocall detector from melspectrograms(1). Initially, we came up with the flow to use the nocall detector to eliminate nocalls for sure, and then predict some birds with nocall labels disabled for the rest, but this did not work. The next idea was to modify the weak labels when building a multilabel classifier, in other words, multiplying the call probabilities obtained from the nocall detector by the labels of train_short_audio. This worked well.</p>\n<p><br></p>\n<p>Next, we built a multilabel classifier (backbone was resnest) for the melspectrogram as the second stage, but there were following four problems at this stage.</p>\n<p><br></p>\n<p>A. The labels of train_short_audio were weak, so it was not clear whether primary or secondary labeled birds were sounding in each frame.<br>\nB. It was possible that the labels itself were wrong.<br>\nC. The output was 397-dimensional probability vectors, and the measure of whether some bird was singing or not was unclear.<br>\nD. Information before and after the current frame might be meaningful.</p>\n<p><br></p>\n<p>To deal with these problems, we created another table competition by ourselves, using the results of the nocall detector, train_metadata, and time series information before and after the current frame. The procedure for creating this table competition data was as follows:</p>\n<p><br></p>\n<p>Ⅰ. Extract the top N candidates in each frame as rows from the results of the multilabel classifier.<br>\nⅡ. Among the N candidates, assigned label 1 to the common set with primary or secondary label, and 0 to the others. However, if the output of the nocall detector indicated that the possibilites of birds singing were low, all the candidates were set to 0. This would be the target variable.<br>\nⅢ. For each candidate, the meta data (time and location information) of the frame and the probabilities that the candidate bird was singing before and after the frame (output of multilabel classifier) were assigned. In addition, more features were added by feature engineering (2), and the table was completed. Then It was analyzed by lightGBM.</p>\n<p><br></p>\n<p>The above steps were expected to have the following effects on the issues A to D mentioned above.</p>\n<p><br></p>\n<p>A. Even if the given data has only a weak label, it is possible to reassign the label for each frame.<br>\nB. The noisy labels can be somewhat cancelled.<br>\nC. The output of the nocall detector can be taken into account.<br>\nD. Information before and after the current frame can be incorporated.</p>\n<p><br></p>\n<p>In addition, by incorporating metadata into the table, it was not necessary to create multiple melspectrogram multilabel classifiers for each region or season, which saved time and computational resources. In fact, in our first-place submission, we used only 15 minutes out of the 3-hour time limit. The fact that we were able to compete based on colab pro might be due to the fact that we succeeded in making this audio competition come down to the table competition.</p>\n<p><br></p>\n<p>The output of our table competition was one probability value per candidate (5 candidates x number of frames in total). The next task was to find an appropriate threshold value for these. In this case, we used train_soundscapes as a reference and searched for the value that maximizes the F1 score by the ternary search. In this process, we hypothesized that the threshold would be different depending on the percentage of nocalls in train_soundscapes, and experimented by reducing the percentage of nocalls, but the conclusion was that the threshold did not change much.</p>\n<p><br></p>\n<p>Finally, we would like to propose a technique we call “nocall injection.” In the current algorithm, rows predicted as “nocall” does not intersected with rows in which some birds were predicted, and vice versa. However, due to the nature of the F1 score, if you predict only one out of the two correct answers, you get 2/3 o<a href=\"url\" target=\"_blank\"></a>f the score, and if you miss the only one label in some row, you get 0. Both are the same in that only one label is wrong, but the size of the penalty is different. Therefore, in order to make sure that nocalls are caught, we added “nocall” to test frames with a high possibility of being nocalls (That would be determined from the output of lightgbm), even if they had been already labeled as some birds. We called this nocall injection. As a result, the nocall and the bird appeared in the same row, and that row could not get the full F1 score, but the benefit of capturing the nocall outweighed it, so the score increased.</p>\n<p><br></p>\n<p>Sorry for writing such a long thread but that's our 3-week activities. It was our first time to compete in a sound recognition competition and we didn't know anything about it at first, but <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>, <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a>, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> and many others helped us a lot with their comments and code. I would like to take this opportunity to thank them again.</p>\n<p><br></p>\n<hr>\n<h1>Note</h1>\n<p>(1) The nocall detector took the external data (freefield1010) as input because the data from train_soundscapes was not enough. The accuracy was also improved by using data augmentation for images.</p>\n<p><br></p>\n<p>(2) The features we created were as follows.<br>\n  \"year\": The year when the audio was recorded.<br>\n  \"month\": The month in which the audio was recorded.<br>\n  \"sum_prob\": The sum of the probability values of the 397 dimensions for the frame.<br>\n  \"mean_prob\": Mean of the 397 dimensional probability values for the frame.<br>\n  \"max_prob\": Maximum 397-dimensional probability values for the frame.<br>\n  \"prev_prob\" ~ \"prev6_prob\": Probability values for the candidate birds in the previous n frames.<br>\n  \"prob\": Probability value for the candidate bird in the target frame.<br>\n  \"next_prob\" ~ \"next6_prob\": Probability values of the candidate bird in the next n frames.<br>\n  \"rank\": The probability that the bird of interest has the highest probability value out of 397 in the frame.<br>\n  \"latitude\":<br>\n  \"longitude\":<br>\n  \"bird_id\": Which of the 397 birds is the target.<br>\n  \"seconds\": The number of seconds in the frame.<br>\n  \"num_appear\": The number of files where the bird is the primary label in train_short_audio<br>\n  \"site_num_appear\": how many times the bird in train_short_audio appeared in the region<br>\n  \"site_appear_ratio\": site_num_appear/num_appear<br>\n  \"prob_diff\": Difference between prob and average of prob values for 3 frames </p>",
      "rawMarkdown": "# tl;dr\nMy teammate @kami634 has already posted a quick version of our solution. Our solution pipeline is composed of three stages and you can overview the stages info here.\nhttps://www.kaggle.com/c/birdclef-2021/discussion/243304\nAlso, this is an image of our pipeline. (only inference part, so some of the steps including the first stage are omitted.)\n\n![img_name](https://drive.google.com/uc?id=1teMrPuBoKeCEf5fFpxoNpzD4fi1IgyxX)\n\nThe complete code used in this competition has been uploaded to the following github. All of the ipynb files there have been confirmed to work properly in the current kaggle notebook environment (2021/6/4).\nhttps://github.com/namakemono/kaggle-birdclef-2021\n\n<br>\n\n# Detailed Solution\n\nThe first thing we did when we joined this competition was to find out the percentage of nocalls in train_soundscapes and test_soundscapes. In train_soundscapes, it is easy to find out, and in test_soundscapes (Public LB), it can be found by submitting all lines as nocall. The results were 0.637 for train_soundscapes and 0.54 for Public LB. Similarly, in the 2020 competition, we submitted all lines as nocall as late submission, and the results were 0.577 for Public LB and 0.544 for Private LB. In all cases, the majority of the targets were nocalls, and we speculated that the difference between birdcalls by bird species and the difference between birds singing and nocall might be qualitatively different. That’s why we considered nocall detection and bird identification as separate tasks. This was the origin of the binary nocall detector from melspectrograms(1). Initially, we came up with the flow to use the nocall detector to eliminate nocalls for sure, and then predict some birds with nocall labels disabled for the rest, but this did not work. The next idea was to modify the weak labels when building a multilabel classifier, in other words, multiplying the call probabilities obtained from the nocall detector by the labels of train_short_audio. This worked well.\n\n<br>\n\nNext, we built a multilabel classifier (backbone was resnest) for the melspectrogram as the second stage, but there were following four problems at this stage.\n\n<br>\n\nA. The labels of train_short_audio were weak, so it was not clear whether primary or secondary labeled birds were sounding in each frame.\nB. It was possible that the labels itself were wrong.\nC. The output was 397-dimensional probability vectors, and the measure of whether some bird was singing or not was unclear.\nD. Information before and after the current frame might be meaningful.\n\n<br>\n\nTo deal with these problems, we created another table competition by ourselves, using the results of the nocall detector, train_metadata, and time series information before and after the current frame. The procedure for creating this table competition data was as follows:\n\n<br>\n\nⅠ. Extract the top N candidates in each frame as rows from the results of the multilabel classifier.\nⅡ. Among the N candidates, assigned label 1 to the common set with primary or secondary label, and 0 to the others. However, if the output of the nocall detector indicated that the possibilites of birds singing were low, all the candidates were set to 0. This would be the target variable.\nⅢ. For each candidate, the meta data (time and location information) of the frame and the probabilities that the candidate bird was singing before and after the frame (output of multilabel classifier) were assigned. In addition, more features were added by feature engineering (2), and the table was completed. Then It was analyzed by lightGBM.\n\n<br>\n\nThe above steps were expected to have the following effects on the issues A to D mentioned above.\n\n<br>\n\nA. Even if the given data has only a weak label, it is possible to reassign the label for each frame.\nB. The noisy labels can be somewhat cancelled.\nC. The output of the nocall detector can be taken into account.\nD. Information before and after the current frame can be incorporated.\n\n<br>\n\nIn addition, by incorporating metadata into the table, it was not necessary to create multiple melspectrogram multilabel classifiers for each region or season, which saved time and computational resources. In fact, in our first-place submission, we used only 15 minutes out of the 3-hour time limit. The fact that we were able to compete based on colab pro might be due to the fact that we succeeded in making this audio competition come down to the table competition.\n\n<br>\n\nThe output of our table competition was one probability value per candidate (5 candidates x number of frames in total). The next task was to find an appropriate threshold value for these. In this case, we used train_soundscapes as a reference and searched for the value that maximizes the F1 score by the ternary search. In this process, we hypothesized that the threshold would be different depending on the percentage of nocalls in train_soundscapes, and experimented by reducing the percentage of nocalls, but the conclusion was that the threshold did not change much.\n\n<br>\n\nFinally, we would like to propose a technique we call “nocall injection.” In the current algorithm, rows predicted as “nocall” does not intersected with rows in which some birds were predicted, and vice versa. However, due to the nature of the F1 score, if you predict only one out of the two correct answers, you get 2/3 o[](url)f the score, and if you miss the only one label in some row, you get 0. Both are the same in that only one label is wrong, but the size of the penalty is different. Therefore, in order to make sure that nocalls are caught, we added “nocall” to test frames with a high possibility of being nocalls (That would be determined from the output of lightgbm), even if they had been already labeled as some birds. We called this nocall injection. As a result, the nocall and the bird appeared in the same row, and that row could not get the full F1 score, but the benefit of capturing the nocall outweighed it, so the score increased.\n\n<br>\n\nSorry for writing such a long thread but that's our 3-week activities. It was our first time to compete in a sound recognition competition and we didn't know anything about it at first, but @kneroma, @hidehisaarai1213, @cpmpml and many others helped us a lot with their comments and code. I would like to take this opportunity to thank them again.\n\n<br>\n\n<hr>\n\n# Note\n\n(1) The nocall detector took the external data (freefield1010) as input because the data from train_soundscapes was not enough. The accuracy was also improved by using data augmentation for images.\n\n<br>\n\n(2) The features we created were as follows.\n  \"year\": The year when the audio was recorded.\n  \"month\": The month in which the audio was recorded.\n  \"sum_prob\": The sum of the probability values of the 397 dimensions for the frame.\n  \"mean_prob\": Mean of the 397 dimensional probability values for the frame.\n  \"max_prob\": Maximum 397-dimensional probability values for the frame.\n  \"prev_prob\" ~ \"prev6_prob\": Probability values for the candidate birds in the previous n frames.\n  \"prob\": Probability value for the candidate bird in the target frame.\n  \"next_prob\" ~ \"next6_prob\": Probability values of the candidate bird in the next n frames.\n  \"rank\": The probability that the bird of interest has the highest probability value out of 397 in the frame.\n  \"latitude\":\n  \"longitude\":\n  \"bird_id\": Which of the 397 birds is the target.\n  \"seconds\": The number of seconds in the frame.\n  \"num_appear\": The number of files where the bird is the primary label in train_short_audio\n  \"site_num_appear\": how many times the bird in train_short_audio appeared in the region\n  \"site_appear_ratio\": site_num_appear/num_appear\n  \"prob_diff\": Difference between prob and average of prob values for 3 frames",
      "votes": null
    },
    {
      "id": "1335890",
      "postDate": "06/04/2021 13:50:39",
      "content": "<h2>More about the machine…</h2>\n<p>We mainly used the Colab Pro as the machine, and were lucky to be assigned a V100 occasionally, but usually a P100.<br>\nA lot of computational resources were needed for the second stage of training. We had converted the images into melspectrograms in advance, so the cost during training was kept low.<br>\nIt took a few minutes per epoch, so a total of 2-5 hours per model to train.</p>\n<p>We had multiple notebooks open at the same time for training and inference(for validation), so even ‘Colab Pro’ sometimes stopped if we used too much.<br>\nIf you get suspended for using the regular ‘Colab’ too much, Google will recommend that you subscribe to ‘Colab Pro’.<br>\nIf you use ‘Colab Pro’ too much and get suspended, Google will recommend that you subscribe to ‘Colab Pro’! If only they offered ‘<strong>Colab Pro Pro</strong>’, I would subscribe! </p>",
      "rawMarkdown": "## More about the machine...\n\nWe mainly used the Colab Pro as the machine, and were lucky to be assigned a V100 occasionally, but usually a P100.\nA lot of computational resources were needed for the second stage of training. We had converted the images into melspectrograms in advance, so the cost during training was kept low.\nIt took a few minutes per epoch, so a total of 2-5 hours per model to train.\n\nWe had multiple notebooks open at the same time for training and inference(for validation), so even ‘Colab Pro’ sometimes stopped if we used too much.\nIf you get suspended for using the regular ‘Colab’ too much, Google will recommend that you subscribe to ‘Colab Pro’.\nIf you use ‘Colab Pro’ too much and get suspended, Google will recommend that you subscribe to ‘Colab Pro’! If only they offered ‘**Colab Pro Pro**’, I would subscribe!",
      "votes": null
    },
    {
      "id": "1335891",
      "postDate": "06/04/2021 13:51:32",
      "content": "<p>Congrats.  I became convinced that a stacking model like yours could be useful, only too late.  It is good to see it worked well.</p>",
      "rawMarkdown": "Congrats.  I became convinced that a stacking model like yours could be useful, only too late.  It is good to see it worked well.",
      "votes": null
    },
    {
      "id": "1335917",
      "postDate": "06/04/2021 14:06:41",
      "content": "<p>XD haha, <strong>Colab Pro Pro</strong></p>",
      "rawMarkdown": "XD haha, **Colab Pro Pro**",
      "votes": null
    },
    {
      "id": "1335973",
      "postDate": "06/04/2021 14:59:17",
      "content": "<p>Thank you. <br>\nDuring the competition, we suffered from the feeling that your score was endlessly far away, actually.</p>",
      "rawMarkdown": "Thank you. \nDuring the competition, we suffered from the feeling that your score was endlessly far away, actually.",
      "votes": null
    },
    {
      "id": "1336128",
      "postDate": "06/04/2021 16:34:12",
      "content": "<p>Congrats once again!! Seems that the limitation on resources make you explore other ways to use the data and pays you back at the end - Well done!! </p>\n<p>ps: I have some questions but I'll look in more detail to your code first… thanks for sharing</p>",
      "rawMarkdown": "Congrats once again!! Seems that the limitation on resources make you explore other ways to use the data and pays you back at the end - Well done!! \n\nps: I have some questions but I'll look in more detail to your code first... thanks for sharing",
      "votes": null
    },
    {
      "id": "1336203",
      "postDate": "06/04/2021 17:59:40",
      "content": "<p>Thank you for the detailed write-up, and congrats on your win. I think your solution was very innovative overall, but I especially liked the idea of the no-call injection, which I think might have been the difference-maker in the end.</p>",
      "rawMarkdown": "Thank you for the detailed write-up, and congrats on your win. I think your solution was very innovative overall, but I especially liked the idea of the no-call injection, which I think might have been the difference-maker in the end.",
      "votes": null
    },
    {
      "id": "1337714",
      "postDate": "06/05/2021 20:16:59",
      "content": "<p>Thanks for sharing your solution and congratulations on the first place. As for uploading images, sadly it has been disabled for now. One workaround, is to upload to google drive, share the file with everyone and then use the following URL template: </p>\n<p><code>![img_name](https://drive.google.com/uc?id=img_id)</code></p>\n<p>where the <code>img_id</code> is the id of your image: you can find it by clicking on the image.</p>\n<p>I hope this helps!</p>",
      "rawMarkdown": "Thanks for sharing your solution and congratulations on the first place. As for uploading images, sadly it has been disabled for now. One workaround, is to upload to google drive, share the file with everyone and then use the following URL template: \n\n`![img_name](https://drive.google.com/uc?id=img_id)`\n\nwhere the `img_id` is the id of your image: you can find it by clicking on the image.\n\nI hope this helps!",
      "votes": null
    },
    {
      "id": "1338514",
      "postDate": "06/06/2021 14:04:04",
      "content": "<p>Is it possible to download Academic Torrents data like original freefield1010 without torrent software? Torrent software is blocked  by my security software.</p>",
      "rawMarkdown": "Is it possible to download Academic Torrents data like original freefield1010 without torrent software? Torrent software is blocked  by my security software.",
      "votes": null
    },
    {
      "id": "1339003",
      "postDate": "06/07/2021 00:04:52",
      "content": "<p>Thank you so much for your help!<br>\nI didn't know that way. I succeeded in uploading our pipeline image.</p>",
      "rawMarkdown": "Thank you so much for your help!\nI didn't know that way. I succeeded in uploading our pipeline image.",
      "votes": null
    },
    {
      "id": "1341174",
      "postDate": "06/08/2021 13:33:49",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shigemitsutomizawa\" target=\"_blank\">@shigemitsutomizawa</a>, sorry for the late reply.</p>\n<p>Actually, we did not download the original freefield1010 audio data, but used the one that was already uploaded as a kaggle dataset. Please check the below dataset.<br>\n<a href=\"https://www.kaggle.com/rlmwang/ff1010bird\" target=\"_blank\">https://www.kaggle.com/rlmwang/ff1010bird</a></p>\n<p>We apologize for not being able to meet your expectations.</p>",
      "rawMarkdown": "Hi @shigemitsutomizawa, sorry for the late reply.\n\nActually, we did not download the original freefield1010 audio data, but used the one that was already uploaded as a kaggle dataset. Please check the below dataset.\nhttps://www.kaggle.com/rlmwang/ff1010bird\n\nWe apologize for not being able to meet your expectations.",
      "votes": null
    },
    {
      "id": "1341295",
      "postDate": "06/08/2021 15:02:15",
      "content": "<p>Awesome! The graph looks neat. 👌</p>",
      "rawMarkdown": "Awesome! The graph looks neat. 👌",
      "votes": null
    },
    {
      "id": "1345189",
      "postDate": "06/11/2021 11:43:12",
      "content": "<p>Very usefull solution, thank's for sharing. Congrats once again !</p>",
      "rawMarkdown": "Very usefull solution, thank's for sharing. Congrats once again !",
      "votes": null
    },
    {
      "id": "1375829",
      "postDate": "07/04/2021 13:59:26",
      "content": "<p>Thanks for sharing your detailed solution and congratulations on the first place. <br>\nI would like to ask you two questions, even though it has been a long time since the competition ended.</p>\n<ol>\n<li>Since the y-axis of a spectrogram is a frequency, it is not always a good idea to flip it (especially vertically) in the same way as an image. (As mentioned in this notebook (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english</a>).) On the other hand your team is using them when making the nocall detector. Did you use them because you saw an improvement in CV or something?</li>\n<li>Regarding the training of lightGBM, in train short audio, some data is about 5 seconds long, how did you calculate features such as prev_prob?  If the clip does not exist, did you replace it with 0?  Also, I thought that if it is 0, it does not mean that there is no clip, but it may train that the bird is not calling. Did you consider anything about this?</li>\n</ol>",
      "rawMarkdown": "Thanks for sharing your detailed solution and congratulations on the first place. \nI would like to ask you two questions, even though it has been a long time since the competition ended.\n1. Since the y-axis of a spectrogram is a frequency, it is not always a good idea to flip it (especially vertically) in the same way as an image. (As mentioned in this notebook ([https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english](url)).) On the other hand your team is using them when making the nocall detector. Did you use them because you saw an improvement in CV or something?\n2. Regarding the training of lightGBM, in train short audio, some data is about 5 seconds long, how did you calculate features such as prev_prob?  If the clip does not exist, did you replace it with 0?  Also, I thought that if it is 0, it does not mean that there is no clip, but it may train that the bird is not calling. Did you consider anything about this?",
      "votes": null
    },
    {
      "id": "1376437",
      "postDate": "07/05/2021 06:25:16",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>, thank you for your questions.</p>\n<ol>\n<li><p>As you pointed out, the y-axis of the mel spectrogram does indeed refer to frequency, and the vertical flip seems to be nonsense. However, I remember that we adopted it because it actually increased the CV score. This is just a hypothesis, but I think that the CV score increased because the vertical flip could preserve the shape of signals (or can we call it \"waveform\" here?) that appear on the mel spectrogram, even though it twisted the sound information itself.</p></li>\n<li><p>LightGBM can accept NaN as input.</p></li>\n</ol>\n<p>I hope these will help…</p>",
      "rawMarkdown": "Hi @naoism, thank you for your questions.\n\n1. As you pointed out, the y-axis of the mel spectrogram does indeed refer to frequency, and the vertical flip seems to be nonsense. However, I remember that we adopted it because it actually increased the CV score. This is just a hypothesis, but I think that the CV score increased because the vertical flip could preserve the shape of signals (or can we call it \"waveform\" here?) that appear on the mel spectrogram, even though it twisted the sound information itself.\n\n2. LightGBM can accept NaN as input.\n\nI hope these will help...",
      "votes": null
    },
    {
      "id": "1378353",
      "postDate": "07/06/2021 13:26:14",
      "content": "<p>Thank you for your quick reply.<br>\nIf the CV got better, then the new information appeared by flipping it. It is also possible that by flipping the spectrogram, one bird's call became a different bird's call (with a different main frequency). On the other hand, I think this is only better because of the nocall detector. Of coarse, that's just my hypothesis.</p>",
      "rawMarkdown": "Thank you for your quick reply.\nIf the CV got better, then the new information appeared by flipping it. It is also possible that by flipping the spectrogram, one bird's call became a different bird's call (with a different main frequency). On the other hand, I think this is only better because of the nocall detector. Of coarse, that's just my hypothesis.",
      "votes": null
    },
    {
      "id": "3197208",
      "postDate": "05/07/2025 22:14:42",
      "content": "<p>Thanks for your explanations. I couldn't figure out the repository description for producing same results. You are mentioning about imitating the directory structure like Kaggle, but if I want to use Kaggle itself, can I create such a directory? Or it is expected to create this directory on my personal sources? I am confused..</p>",
      "rawMarkdown": "Thanks for your explanations. I couldn't figure out the repository description for producing same results. You are mentioning about imitating the directory structure like Kaggle, but if I want to use Kaggle itself, can I create such a directory? Or it is expected to create this directory on my personal sources? I am confused..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1335890,
      "author_name": "kami634",
      "author_url": "",
      "post_date": "06/04/2021 13:50:39",
      "content": "<h2>More about the machine…</h2>\n<p>We mainly used the Colab Pro as the machine, and were lucky to be assigned a V100 occasionally, but usually a P100.<br>\nA lot of computational resources were needed for the second stage of training. We had converted the images into melspectrograms in advance, so the cost during training was kept low.<br>\nIt took a few minutes per epoch, so a total of 2-5 hours per model to train.</p>\n<p>We had multiple notebooks open at the same time for training and inference(for validation), so even ‘Colab Pro’ sometimes stopped if we used too much.<br>\nIf you get suspended for using the regular ‘Colab’ too much, Google will recommend that you subscribe to ‘Colab Pro’.<br>\nIf you use ‘Colab Pro’ too much and get suspended, Google will recommend that you subscribe to ‘Colab Pro’! If only they offered ‘<strong>Colab Pro Pro</strong>’, I would subscribe! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1335917,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "06/04/2021 14:06:41",
          "content": "<p>XD haha, <strong>Colab Pro Pro</strong></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1335891,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/04/2021 13:51:32",
      "content": "<p>Congrats.  I became convinced that a stacking model like yours could be useful, only too late.  It is good to see it worked well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1335973,
          "author_name": "startjapan",
          "author_url": "",
          "post_date": "06/04/2021 14:59:17",
          "content": "<p>Thank you. <br>\nDuring the competition, we suffered from the feeling that your score was endlessly far away, actually.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1336128,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "06/04/2021 16:34:12",
      "content": "<p>Congrats once again!! Seems that the limitation on resources make you explore other ways to use the data and pays you back at the end - Well done!! </p>\n<p>ps: I have some questions but I'll look in more detail to your code first… thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1336203,
      "author_name": "agneev",
      "author_url": "",
      "post_date": "06/04/2021 17:59:40",
      "content": "<p>Thank you for the detailed write-up, and congrats on your win. I think your solution was very innovative overall, but I especially liked the idea of the no-call injection, which I think might have been the difference-maker in the end.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1337714,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "06/05/2021 20:16:59",
      "content": "<p>Thanks for sharing your solution and congratulations on the first place. As for uploading images, sadly it has been disabled for now. One workaround, is to upload to google drive, share the file with everyone and then use the following URL template: </p>\n<p><code>![img_name](https://drive.google.com/uc?id=img_id)</code></p>\n<p>where the <code>img_id</code> is the id of your image: you can find it by clicking on the image.</p>\n<p>I hope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1339003,
          "author_name": "startjapan",
          "author_url": "",
          "post_date": "06/07/2021 00:04:52",
          "content": "<p>Thank you so much for your help!<br>\nI didn't know that way. I succeeded in uploading our pipeline image.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1341295,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "06/08/2021 15:02:15",
          "content": "<p>Awesome! The graph looks neat. 👌</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1338514,
      "author_name": "shigemitsutomizawa",
      "author_url": "",
      "post_date": "06/06/2021 14:04:04",
      "content": "<p>Is it possible to download Academic Torrents data like original freefield1010 without torrent software? Torrent software is blocked  by my security software.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1341174,
          "author_name": "startjapan",
          "author_url": "",
          "post_date": "06/08/2021 13:33:49",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/shigemitsutomizawa\" target=\"_blank\">@shigemitsutomizawa</a>, sorry for the late reply.</p>\n<p>Actually, we did not download the original freefield1010 audio data, but used the one that was already uploaded as a kaggle dataset. Please check the below dataset.<br>\n<a href=\"https://www.kaggle.com/rlmwang/ff1010bird\" target=\"_blank\">https://www.kaggle.com/rlmwang/ff1010bird</a></p>\n<p>We apologize for not being able to meet your expectations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1345189,
      "author_name": "raphaelbordeau",
      "author_url": "",
      "post_date": "06/11/2021 11:43:12",
      "content": "<p>Very usefull solution, thank's for sharing. Congrats once again !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1375829,
      "author_name": "naoism",
      "author_url": "",
      "post_date": "07/04/2021 13:59:26",
      "content": "<p>Thanks for sharing your detailed solution and congratulations on the first place. <br>\nI would like to ask you two questions, even though it has been a long time since the competition ended.</p>\n<ol>\n<li>Since the y-axis of a spectrogram is a frequency, it is not always a good idea to flip it (especially vertically) in the same way as an image. (As mentioned in this notebook (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english</a>).) On the other hand your team is using them when making the nocall detector. Did you use them because you saw an improvement in CV or something?</li>\n<li>Regarding the training of lightGBM, in train short audio, some data is about 5 seconds long, how did you calculate features such as prev_prob?  If the clip does not exist, did you replace it with 0?  Also, I thought that if it is 0, it does not mean that there is no clip, but it may train that the bird is not calling. Did you consider anything about this?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1376437,
          "author_name": "startjapan",
          "author_url": "",
          "post_date": "07/05/2021 06:25:16",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/naoism\" target=\"_blank\">@naoism</a>, thank you for your questions.</p>\n<ol>\n<li><p>As you pointed out, the y-axis of the mel spectrogram does indeed refer to frequency, and the vertical flip seems to be nonsense. However, I remember that we adopted it because it actually increased the CV score. This is just a hypothesis, but I think that the CV score increased because the vertical flip could preserve the shape of signals (or can we call it \"waveform\" here?) that appear on the mel spectrogram, even though it twisted the sound information itself.</p></li>\n<li><p>LightGBM can accept NaN as input.</p></li>\n</ol>\n<p>I hope these will help…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1378353,
          "author_name": "naoism",
          "author_url": "",
          "post_date": "07/06/2021 13:26:14",
          "content": "<p>Thank you for your quick reply.<br>\nIf the CV got better, then the new information appeared by flipping it. It is also possible that by flipping the spectrogram, one bird's call became a different bird's call (with a different main frequency). On the other hand, I think this is only better because of the nocall detector. Of coarse, that's just my hypothesis.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3197208,
      "author_name": "ayagler",
      "author_url": "",
      "post_date": "05/07/2025 22:14:42",
      "content": "<p>Thanks for your explanations. I couldn't figure out the repository description for producing same results. You are mentioning about imitating the directory structure like Kaggle, but if I want to use Kaggle itself, can I create such a directory? Or it is expected to create this directory on my personal sources? I am confused..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1335887": "# tl;dr\nMy teammate @kami634 has already posted a quick version of our solution. Our solution pipeline is composed of three stages and you can overview the stages info here.\nhttps://www.kaggle.com/c/birdclef-2021/discussion/243304\nAlso, this is an image of our pipeline. (only inference part, so some of the steps including the first stage are omitted.)\n\n![img_name](https://drive.google.com/uc?id=1teMrPuBoKeCEf5fFpxoNpzD4fi1IgyxX)\n\nThe complete code used in this competition has been uploaded to the following github. All of the ipynb files there have been confirmed to work properly in the current kaggle notebook environment (2021/6/4).\nhttps://github.com/namakemono/kaggle-birdclef-2021\n\n<br>\n\n# Detailed Solution\n\nThe first thing we did when we joined this competition was to find out the percentage of nocalls in train_soundscapes and test_soundscapes. In train_soundscapes, it is easy to find out, and in test_soundscapes (Public LB), it can be found by submitting all lines as nocall. The results were 0.637 for train_soundscapes and 0.54 for Public LB. Similarly, in the 2020 competition, we submitted all lines as nocall as late submission, and the results were 0.577 for Public LB and 0.544 for Private LB. In all cases, the majority of the targets were nocalls, and we speculated that the difference between birdcalls by bird species and the difference between birds singing and nocall might be qualitatively different. That’s why we considered nocall detection and bird identification as separate tasks. This was the origin of the binary nocall detector from melspectrograms(1). Initially, we came up with the flow to use the nocall detector to eliminate nocalls for sure, and then predict some birds with nocall labels disabled for the rest, but this did not work. The next idea was to modify the weak labels when building a multilabel classifier, in other words, multiplying the call probabilities obtained from the nocall detector by the labels of train_short_audio. This worked well.\n\n<br>\n\nNext, we built a multilabel classifier (backbone was resnest) for the melspectrogram as the second stage, but there were following four problems at this stage.\n\n<br>\n\nA. The labels of train_short_audio were weak, so it was not clear whether primary or secondary labeled birds were sounding in each frame.\nB. It was possible that the labels itself were wrong.\nC. The output was 397-dimensional probability vectors, and the measure of whether some bird was singing or not was unclear.\nD. Information before and after the current frame might be meaningful.\n\n<br>\n\nTo deal with these problems, we created another table competition by ourselves, using the results of the nocall detector, train_metadata, and time series information before and after the current frame. The procedure for creating this table competition data was as follows:\n\n<br>\n\nⅠ. Extract the top N candidates in each frame as rows from the results of the multilabel classifier.\nⅡ. Among the N candidates, assigned label 1 to the common set with primary or secondary label, and 0 to the others. However, if the output of the nocall detector indicated that the possibilites of birds singing were low, all the candidates were set to 0. This would be the target variable.\nⅢ. For each candidate, the meta data (time and location information) of the frame and the probabilities that the candidate bird was singing before and after the frame (output of multilabel classifier) were assigned. In addition, more features were added by feature engineering (2), and the table was completed. Then It was analyzed by lightGBM.\n\n<br>\n\nThe above steps were expected to have the following effects on the issues A to D mentioned above.\n\n<br>\n\nA. Even if the given data has only a weak label, it is possible to reassign the label for each frame.\nB. The noisy labels can be somewhat cancelled.\nC. The output of the nocall detector can be taken into account.\nD. Information before and after the current frame can be incorporated.\n\n<br>\n\nIn addition, by incorporating metadata into the table, it was not necessary to create multiple melspectrogram multilabel classifiers for each region or season, which saved time and computational resources. In fact, in our first-place submission, we used only 15 minutes out of the 3-hour time limit. The fact that we were able to compete based on colab pro might be due to the fact that we succeeded in making this audio competition come down to the table competition.\n\n<br>\n\nThe output of our table competition was one probability value per candidate (5 candidates x number of frames in total). The next task was to find an appropriate threshold value for these. In this case, we used train_soundscapes as a reference and searched for the value that maximizes the F1 score by the ternary search. In this process, we hypothesized that the threshold would be different depending on the percentage of nocalls in train_soundscapes, and experimented by reducing the percentage of nocalls, but the conclusion was that the threshold did not change much.\n\n<br>\n\nFinally, we would like to propose a technique we call “nocall injection.” In the current algorithm, rows predicted as “nocall” does not intersected with rows in which some birds were predicted, and vice versa. However, due to the nature of the F1 score, if you predict only one out of the two correct answers, you get 2/3 o[](url)f the score, and if you miss the only one label in some row, you get 0. Both are the same in that only one label is wrong, but the size of the penalty is different. Therefore, in order to make sure that nocalls are caught, we added “nocall” to test frames with a high possibility of being nocalls (That would be determined from the output of lightgbm), even if they had been already labeled as some birds. We called this nocall injection. As a result, the nocall and the bird appeared in the same row, and that row could not get the full F1 score, but the benefit of capturing the nocall outweighed it, so the score increased.\n\n<br>\n\nSorry for writing such a long thread but that's our 3-week activities. It was our first time to compete in a sound recognition competition and we didn't know anything about it at first, but @kneroma, @hidehisaarai1213, @cpmpml and many others helped us a lot with their comments and code. I would like to take this opportunity to thank them again.\n\n<br>\n\n<hr>\n\n# Note\n\n(1) The nocall detector took the external data (freefield1010) as input because the data from train_soundscapes was not enough. The accuracy was also improved by using data augmentation for images.\n\n<br>\n\n(2) The features we created were as follows.\n  \"year\": The year when the audio was recorded.\n  \"month\": The month in which the audio was recorded.\n  \"sum_prob\": The sum of the probability values of the 397 dimensions for the frame.\n  \"mean_prob\": Mean of the 397 dimensional probability values for the frame.\n  \"max_prob\": Maximum 397-dimensional probability values for the frame.\n  \"prev_prob\" ~ \"prev6_prob\": Probability values for the candidate birds in the previous n frames.\n  \"prob\": Probability value for the candidate bird in the target frame.\n  \"next_prob\" ~ \"next6_prob\": Probability values of the candidate bird in the next n frames.\n  \"rank\": The probability that the bird of interest has the highest probability value out of 397 in the frame.\n  \"latitude\":\n  \"longitude\":\n  \"bird_id\": Which of the 397 birds is the target.\n  \"seconds\": The number of seconds in the frame.\n  \"num_appear\": The number of files where the bird is the primary label in train_short_audio\n  \"site_num_appear\": how many times the bird in train_short_audio appeared in the region\n  \"site_appear_ratio\": site_num_appear/num_appear\n  \"prob_diff\": Difference between prob and average of prob values for 3 frames",
    "1335890": "## More about the machine...\n\nWe mainly used the Colab Pro as the machine, and were lucky to be assigned a V100 occasionally, but usually a P100.\nA lot of computational resources were needed for the second stage of training. We had converted the images into melspectrograms in advance, so the cost during training was kept low.\nIt took a few minutes per epoch, so a total of 2-5 hours per model to train.\n\nWe had multiple notebooks open at the same time for training and inference(for validation), so even ‘Colab Pro’ sometimes stopped if we used too much.\nIf you get suspended for using the regular ‘Colab’ too much, Google will recommend that you subscribe to ‘Colab Pro’.\nIf you use ‘Colab Pro’ too much and get suspended, Google will recommend that you subscribe to ‘Colab Pro’! If only they offered ‘**Colab Pro Pro**’, I would subscribe!",
    "1335891": "Congrats.  I became convinced that a stacking model like yours could be useful, only too late.  It is good to see it worked well.",
    "1335917": "XD haha, **Colab Pro Pro**",
    "1335973": "Thank you. \nDuring the competition, we suffered from the feeling that your score was endlessly far away, actually.",
    "1336128": "Congrats once again!! Seems that the limitation on resources make you explore other ways to use the data and pays you back at the end - Well done!! \n\nps: I have some questions but I'll look in more detail to your code first... thanks for sharing",
    "1336203": "Thank you for the detailed write-up, and congrats on your win. I think your solution was very innovative overall, but I especially liked the idea of the no-call injection, which I think might have been the difference-maker in the end.",
    "1337714": "Thanks for sharing your solution and congratulations on the first place. As for uploading images, sadly it has been disabled for now. One workaround, is to upload to google drive, share the file with everyone and then use the following URL template: \n\n`![img_name](https://drive.google.com/uc?id=img_id)`\n\nwhere the `img_id` is the id of your image: you can find it by clicking on the image.\n\nI hope this helps!",
    "1338514": "Is it possible to download Academic Torrents data like original freefield1010 without torrent software? Torrent software is blocked  by my security software.",
    "1339003": "Thank you so much for your help!\nI didn't know that way. I succeeded in uploading our pipeline image.",
    "1341174": "Hi @shigemitsutomizawa, sorry for the late reply.\n\nActually, we did not download the original freefield1010 audio data, but used the one that was already uploaded as a kaggle dataset. Please check the below dataset.\nhttps://www.kaggle.com/rlmwang/ff1010bird\n\nWe apologize for not being able to meet your expectations.",
    "1341295": "Awesome! The graph looks neat. 👌",
    "1345189": "Very usefull solution, thank's for sharing. Congrats once again !",
    "1375829": "Thanks for sharing your detailed solution and congratulations on the first place. \nI would like to ask you two questions, even though it has been a long time since the competition ended.\n1. Since the y-axis of a spectrogram is a frequency, it is not always a good idea to flip it (especially vertically) in the same way as an image. (As mentioned in this notebook ([https://www.kaggle.com/hidehisaarai1213/rfcx-audio-data-augmentation-japanese-english](url)).) On the other hand your team is using them when making the nocall detector. Did you use them because you saw an improvement in CV or something?\n2. Regarding the training of lightGBM, in train short audio, some data is about 5 seconds long, how did you calculate features such as prev_prob?  If the clip does not exist, did you replace it with 0?  Also, I thought that if it is 0, it does not mean that there is no clip, but it may train that the bird is not calling. Did you consider anything about this?",
    "1376437": "Hi @naoism, thank you for your questions.\n\n1. As you pointed out, the y-axis of the mel spectrogram does indeed refer to frequency, and the vertical flip seems to be nonsense. However, I remember that we adopted it because it actually increased the CV score. This is just a hypothesis, but I think that the CV score increased because the vertical flip could preserve the shape of signals (or can we call it \"waveform\" here?) that appear on the mel spectrogram, even though it twisted the sound information itself.\n\n2. LightGBM can accept NaN as input.\n\nI hope these will help...",
    "1378353": "Thank you for your quick reply.\nIf the CV got better, then the new information appeared by flipping it. It is also possible that by flipping the spectrogram, one bird's call became a different bird's call (with a different main frequency). On the other hand, I think this is only better because of the nocall detector. Of coarse, that's just my hypothesis.",
    "3197208": "Thanks for your explanations. I couldn't figure out the repository description for producing same results. You are mentioning about imitating the directory structure like Kaggle, but if I want to use Kaggle itself, can I create such a directory? Or it is expected to create this directory on my personal sources? I am confused.."
  },
  "source": "meta"
}