{
  "id": 412871,
  "title": "8th Place Solution: Implementing Multimodal Data Augmentation Methods (last updated 05/28)",
  "url": "/competitions/birdclef-2023/writeups/furu-nag-8th-place-solution-implementing-multimoda",
  "author_name": "",
  "post_date": "2023-05-28T09:28:22.960Z",
  "votes": 17,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Congratulations to all the winners!<br>\nOur gratitude extends to Kaggle and the Cornell Lab of Ornithology for their exceptional efforts in organizing this intriguing competition. In particular, I was able to contribute to the improvement of accuracy with highly original methods, so I was able to enjoy the competition very much.<br>\nAllow me to give a succinct overview of my solution. I’ll provide a more detailed update on the solution in the following days.</p>\n<h1><strong>Summary</strong></h1>\n<p>In my solution, only one type of backbone is used, and multimodal data augmentation techniques such as object detection, recommendation systems, and NLP are used to increase diversity and improve model accuracy.</p>\n<h2>Preprocess Pipeline Summary</h2>\n<p>We built a preprocessing pipeline that prevents overfitting by devising preprocessing so that the combinations that are actually possible while maintaining the diversity of the data. In particular, in preprocessing for Sound Object Detection, weighting the signal density of the spectrum made it easier to pick out the areas where the contribution of secondary labels is included. Thanks to that, the training accuracy has improved even when the value of the secondary labels is large.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2Ffc17c767400faa897953a3034352d16e%2F.drawio%20(1).png?generation=1685242284411201&amp;alt=media\" alt=\"\"></p>\n<h2>Training Pipeline</h2>\n<p>In order to grasp the features of bird calls with various periodicities, build a pipeline that first learns with a long spectrum input, then lowers the time axis direction for each epoch, and finally learns local features. Did. Preprocessing by SOD prevents noise-like offsets from being selected even for local spectra, so training is stable.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2F090bc781738d7cfe844801e0bcd23f6e%2F8thplacesolution-2.drawio.png?generation=1685260705814564&amp;alt=media\" alt=\"\"></p>\n<p>code detail: <a href=\"https://www.kaggle.com/code/kunihikofurugori/8th-solution-down-scaling-detail/notebook?scriptVersionId=131297676\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/8th-solution-down-scaling-detail/notebook?scriptVersionId=131297676</a></p>\n<h1><strong>MultiModal Data Augments Part(in Detailed)</strong></h1>\n<p>Code detail: <a href=\"https://www.kaggle.com/code/kunihikofurugori/8th-solution-notebook/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/8th-solution-notebook/notebook</a></p>\n<h3>・Time Segment DownScaling Augmentation (Global to Local)</h3>\n<p>Regarding the mini-batch method represented by the 2nd place solution of birdclef2021, we trained to learn local features while pre-learning global features by training to evolve the number of mini-batches from 15 to 1 at each epoch.</p>\n<p>By first learning the global features and then learning the local features, it becomes possible to predict the local features after understanding the global features.</p>\n<h3>・Weak to Strong Preprocessing by YOLOv8 (Sound Object Detection)</h3>\n<p>I use the data pointed out by the competition host to train the yolov8 object detection model, and select an offset that will increase the number of object detections as much as possible.<br>\nAlso, by minimizing the bounding box and reliability-weighted IoU for the frequency masking, Frequency Masking effectively hides noise.</p>\n<h4>Sample Inference</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2F058db543cb9d88e54a699785bd32c3b2%2F__results___3_2.png?generation=1685058390267527&amp;alt=media\" alt=\"Inferece Sample Sound Object Detection\"></p>\n<p>Reference Notebook<br>\n[1]<a href=\"https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-prepare-annot-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-prepare-annot-data/notebook</a><br>\n[2]<a href=\"https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectiontrain/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectiontrain/notebook</a><br>\n[3]<a href=\"https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-batch-predict/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-batch-predict/notebook</a><br>\n[4]<a href=\"https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectionanalysis/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectionanalysis/notebook</a></p>\n<h2>・Various Mixup Strategy</h2>\n<p>By using various mix-up methods, we tried to make it impossible to choose unnatural pairs. Details of the method are described below.</p>\n<h3>・Geometrical Mixup</h3>\n<p>Assuming that geographically close samples should have similar background noise, we created a model that mixes up geographically close samples without learning the background noise.</p>\n<p>If the distance matrix is properly constructed, the amount of calculation will be O(n^2), so the code is written so that the parts that match up to the second decimal place of the longitude and latitude are treated as the same group.</p>\n<h3>・Resonance Mixup (Inspired by collaborative filtering)</h3>\n<p>Choosing a pair that resonates easily as a mixup destination is important for creating a more natural sound. Implemented a mixup that preferentially selects bird samples that are likely to resonate based on the collaborative filtering idea of the recommender system.</p>\n<h3>・Inner mixup for geometrical clustering space</h3>\n<p>Groups that are geographically close and have the same label do not change the label characteristics even if they share inputs and mixup. Therefore, we clustered such groups and made it possible to perform Inputmixup freely.</p>\n<h1><strong>Models</strong></h1>\n<p>My models are based on the BirdCLEF 2021 2nd place solution using backbone, eca_nfnet_l0.</p>\n<h1><strong>Dataset</strong></h1>\n<ol>\n<li>Bird CLEF 2023</li>\n<li>Bird CLEF 2022/2021/2020 (Pretraining)</li>\n<li>ff1010bird_nocall</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358\" target=\"_blank\">Zendo data (YOLOv8) </a></li>\n</ol>\n<h1><strong>CV setup</strong></h1>\n<p>I thought it would be difficult to create a multifold model and check the LB due to resource limitations, so I did a simple test-train split instead of out-of-fold.</p>\n<p>For the test data, we built validation by splitting 30 seconds into 5 second intervals and bootstrap sampling to the same size as the LB dataset size (8400 samples).</p>\n<p>Furthermore, we found that using only the data with the primary label showed behavior similar to LB, so we excluded the data with the secondary label from the validation and calculated.</p>\n<p>Inference Kernel: <a href=\"https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel</a><br>\nGithub link: <a href=\"https://github.com/furu-kaggle/birdclef2023\" target=\"_blank\">https://github.com/furu-kaggle/birdclef2023</a></p>\n<h1>Result</h1>\n<table>\n<thead>\n<tr>\n<th>commit_id</th>\n<th>single inference kernel</th>\n<th>LB</th>\n<th>PB</th>\n<th>tag</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>4a57a984d1f4101b9f470095e6a6fae0a6cc5af2</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=130154067\" target=\"_blank\">link</a></td>\n<td>0.82628</td>\n<td><strong>0.73777</strong></td>\n<td>ecanfnet_nmel128_best1</td>\n</tr>\n<tr>\n<td>55557aea59976c5befc11cbcd077f22c1d8289c0</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129843959\" target=\"_blank\">link</a></td>\n<td>0.82475</td>\n<td>0.73301</td>\n<td>ecanfnet_nmel128_best3</td>\n</tr>\n<tr>\n<td>ccc435294a9f5f9989f5865def93da0817c3b030</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129843959\" target=\"_blank\">link</a></td>\n<td><strong>0.82786</strong></td>\n<td>0.73341</td>\n<td>ecanfnet_nmel128_best2</td>\n</tr>\n<tr>\n<td>c9e077c172368c31c2d1974e0288554ebd014b2f</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=130288082\" target=\"_blank\">link</a></td>\n<td>0.82279</td>\n<td>0.72918</td>\n<td>ecanfnet_nmel64_best2</td>\n</tr>\n<tr>\n<td>a3791c561544adcc37507d5c27ec98e9d1c603a7</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129584055\" target=\"_blank\">link</a></td>\n<td>0.82172</td>\n<td>0.73223</td>\n<td>ecanfnet_nmel64_best1</td>\n</tr>\n<tr>\n<td>ea70ecc95de1625212465b8e0148bc996e0237f9</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129564172\" target=\"_blank\">link</a></td>\n<td>0.82149</td>\n<td>0.72850</td>\n<td>ecanfnet_nmel64_best3</td>\n</tr>\n<tr>\n<td>ecanfnet_nmel128_best1~3(300 stride pred) + ecanfnet_nmel64_best1~3 ensemble</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel?scriptVersionId=130834386\" target=\"_blank\">link</a></td>\n<td><strong>0.83616</strong></td>\n<td><strong>0.75285</strong></td>\n<td>best_submission</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2274037",
      "postDate": "05/25/2023 15:14:24",
      "content": "<p>Congratulations to all the winners!<br>\nOur gratitude extends to Kaggle and the Cornell Lab of Ornithology for their exceptional efforts in organizing this intriguing competition. In particular, I was able to contribute to the improvement of accuracy with highly original methods, so I was able to enjoy the competition very much.<br>\nAllow me to give a succinct overview of my solution. I’ll provide a more detailed update on the solution in the following days.</p>\n<h1><strong>Summary</strong></h1>\n<p>In my solution, only one type of backbone is used, and multimodal data augmentation techniques such as object detection, recommendation systems, and NLP are used to increase diversity and improve model accuracy.</p>\n<h2>Preprocess Pipeline Summary</h2>\n<p>We built a preprocessing pipeline that prevents overfitting by devising preprocessing so that the combinations that are actually possible while maintaining the diversity of the data. In particular, in preprocessing for Sound Object Detection, weighting the signal density of the spectrum made it easier to pick out the areas where the contribution of secondary labels is included. Thanks to that, the training accuracy has improved even when the value of the secondary labels is large.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2Ffc17c767400faa897953a3034352d16e%2F.drawio%20(1).png?generation=1685242284411201&amp;alt=media\" alt=\"\"></p>\n<h2>Training Pipeline</h2>\n<p>In order to grasp the features of bird calls with various periodicities, build a pipeline that first learns with a long spectrum input, then lowers the time axis direction for each epoch, and finally learns local features. Did. Preprocessing by SOD prevents noise-like offsets from being selected even for local spectra, so training is stable.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2F090bc781738d7cfe844801e0bcd23f6e%2F8thplacesolution-2.drawio.png?generation=1685260705814564&amp;alt=media\" alt=\"\"></p>\n<p>code detail: <a href=\"https://www.kaggle.com/code/kunihikofurugori/8th-solution-down-scaling-detail/notebook?scriptVersionId=131297676\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/8th-solution-down-scaling-detail/notebook?scriptVersionId=131297676</a></p>\n<h1><strong>MultiModal Data Augments Part(in Detailed)</strong></h1>\n<p>Code detail: <a href=\"https://www.kaggle.com/code/kunihikofurugori/8th-solution-notebook/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/8th-solution-notebook/notebook</a></p>\n<h3>・Time Segment DownScaling Augmentation (Global to Local)</h3>\n<p>Regarding the mini-batch method represented by the 2nd place solution of birdclef2021, we trained to learn local features while pre-learning global features by training to evolve the number of mini-batches from 15 to 1 at each epoch.</p>\n<p>By first learning the global features and then learning the local features, it becomes possible to predict the local features after understanding the global features.</p>\n<h3>・Weak to Strong Preprocessing by YOLOv8 (Sound Object Detection)</h3>\n<p>I use the data pointed out by the competition host to train the yolov8 object detection model, and select an offset that will increase the number of object detections as much as possible.<br>\nAlso, by minimizing the bounding box and reliability-weighted IoU for the frequency masking, Frequency Masking effectively hides noise.</p>\n<h4>Sample Inference</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2F058db543cb9d88e54a699785bd32c3b2%2F__results___3_2.png?generation=1685058390267527&amp;alt=media\" alt=\"Inferece Sample Sound Object Detection\"></p>\n<p>Reference Notebook<br>\n[1]<a href=\"https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-prepare-annot-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-prepare-annot-data/notebook</a><br>\n[2]<a href=\"https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectiontrain/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectiontrain/notebook</a><br>\n[3]<a href=\"https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-batch-predict/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-batch-predict/notebook</a><br>\n[4]<a href=\"https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectionanalysis/notebook\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectionanalysis/notebook</a></p>\n<h2>・Various Mixup Strategy</h2>\n<p>By using various mix-up methods, we tried to make it impossible to choose unnatural pairs. Details of the method are described below.</p>\n<h3>・Geometrical Mixup</h3>\n<p>Assuming that geographically close samples should have similar background noise, we created a model that mixes up geographically close samples without learning the background noise.</p>\n<p>If the distance matrix is properly constructed, the amount of calculation will be O(n^2), so the code is written so that the parts that match up to the second decimal place of the longitude and latitude are treated as the same group.</p>\n<h3>・Resonance Mixup (Inspired by collaborative filtering)</h3>\n<p>Choosing a pair that resonates easily as a mixup destination is important for creating a more natural sound. Implemented a mixup that preferentially selects bird samples that are likely to resonate based on the collaborative filtering idea of the recommender system.</p>\n<h3>・Inner mixup for geometrical clustering space</h3>\n<p>Groups that are geographically close and have the same label do not change the label characteristics even if they share inputs and mixup. Therefore, we clustered such groups and made it possible to perform Inputmixup freely.</p>\n<h1><strong>Models</strong></h1>\n<p>My models are based on the BirdCLEF 2021 2nd place solution using backbone, eca_nfnet_l0.</p>\n<h1><strong>Dataset</strong></h1>\n<ol>\n<li>Bird CLEF 2023</li>\n<li>Bird CLEF 2022/2021/2020 (Pretraining)</li>\n<li>ff1010bird_nocall</li>\n<li><a href=\"https://www.kaggle.com/competitions/birdclef-2023/discussion/394358\" target=\"_blank\">Zendo data (YOLOv8) </a></li>\n</ol>\n<h1><strong>CV setup</strong></h1>\n<p>I thought it would be difficult to create a multifold model and check the LB due to resource limitations, so I did a simple test-train split instead of out-of-fold.</p>\n<p>For the test data, we built validation by splitting 30 seconds into 5 second intervals and bootstrap sampling to the same size as the LB dataset size (8400 samples).</p>\n<p>Furthermore, we found that using only the data with the primary label showed behavior similar to LB, so we excluded the data with the secondary label from the validation and calculated.</p>\n<p>Inference Kernel: <a href=\"https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel\" target=\"_blank\">https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel</a><br>\nGithub link: <a href=\"https://github.com/furu-kaggle/birdclef2023\" target=\"_blank\">https://github.com/furu-kaggle/birdclef2023</a></p>\n<h1>Result</h1>\n<table>\n<thead>\n<tr>\n<th>commit_id</th>\n<th>single inference kernel</th>\n<th>LB</th>\n<th>PB</th>\n<th>tag</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>4a57a984d1f4101b9f470095e6a6fae0a6cc5af2</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=130154067\" target=\"_blank\">link</a></td>\n<td>0.82628</td>\n<td><strong>0.73777</strong></td>\n<td>ecanfnet_nmel128_best1</td>\n</tr>\n<tr>\n<td>55557aea59976c5befc11cbcd077f22c1d8289c0</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129843959\" target=\"_blank\">link</a></td>\n<td>0.82475</td>\n<td>0.73301</td>\n<td>ecanfnet_nmel128_best3</td>\n</tr>\n<tr>\n<td>ccc435294a9f5f9989f5865def93da0817c3b030</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129843959\" target=\"_blank\">link</a></td>\n<td><strong>0.82786</strong></td>\n<td>0.73341</td>\n<td>ecanfnet_nmel128_best2</td>\n</tr>\n<tr>\n<td>c9e077c172368c31c2d1974e0288554ebd014b2f</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=130288082\" target=\"_blank\">link</a></td>\n<td>0.82279</td>\n<td>0.72918</td>\n<td>ecanfnet_nmel64_best2</td>\n</tr>\n<tr>\n<td>a3791c561544adcc37507d5c27ec98e9d1c603a7</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129584055\" target=\"_blank\">link</a></td>\n<td>0.82172</td>\n<td>0.73223</td>\n<td>ecanfnet_nmel64_best1</td>\n</tr>\n<tr>\n<td>ea70ecc95de1625212465b8e0148bc996e0237f9</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129564172\" target=\"_blank\">link</a></td>\n<td>0.82149</td>\n<td>0.72850</td>\n<td>ecanfnet_nmel64_best3</td>\n</tr>\n<tr>\n<td>ecanfnet_nmel128_best1~3(300 stride pred) + ecanfnet_nmel64_best1~3 ensemble</td>\n<td><a href=\"https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel?scriptVersionId=130834386\" target=\"_blank\">link</a></td>\n<td><strong>0.83616</strong></td>\n<td><strong>0.75285</strong></td>\n<td>best_submission</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Congratulations to all the winners!\nOur gratitude extends to Kaggle and the Cornell Lab of Ornithology for their exceptional efforts in organizing this intriguing competition. In particular, I was able to contribute to the improvement of accuracy with highly original methods, so I was able to enjoy the competition very much.\nAllow me to give a succinct overview of my solution. I’ll provide a more detailed update on the solution in the following days.\n\n# **Summary**\nIn my solution, only one type of backbone is used, and multimodal data augmentation techniques such as object detection, recommendation systems, and NLP are used to increase diversity and improve model accuracy.\n\n## Preprocess Pipeline Summary\nWe built a preprocessing pipeline that prevents overfitting by devising preprocessing so that the combinations that are actually possible while maintaining the diversity of the data. In particular, in preprocessing for Sound Object Detection, weighting the signal density of the spectrum made it easier to pick out the areas where the contribution of secondary labels is included. Thanks to that, the training accuracy has improved even when the value of the secondary labels is large.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2Ffc17c767400faa897953a3034352d16e%2F.drawio%20(1).png?generation=1685242284411201&alt=media)\n\n## Training Pipeline\nIn order to grasp the features of bird calls with various periodicities, build a pipeline that first learns with a long spectrum input, then lowers the time axis direction for each epoch, and finally learns local features. Did. Preprocessing by SOD prevents noise-like offsets from being selected even for local spectra, so training is stable.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2F090bc781738d7cfe844801e0bcd23f6e%2F8thplacesolution-2.drawio.png?generation=1685260705814564&alt=media)\n\ncode detail: https://www.kaggle.com/code/kunihikofurugori/8th-solution-down-scaling-detail/notebook?scriptVersionId=131297676\n\n# **MultiModal Data Augments Part(in Detailed)**\nCode detail: https://www.kaggle.com/code/kunihikofurugori/8th-solution-notebook/notebook\n### ・Time Segment DownScaling Augmentation (Global to Local)\nRegarding the mini-batch method represented by the 2nd place solution of birdclef2021, we trained to learn local features while pre-learning global features by training to evolve the number of mini-batches from 15 to 1 at each epoch.\n\nBy first learning the global features and then learning the local features, it becomes possible to predict the local features after understanding the global features.\n\n### ・Weak to Strong Preprocessing by YOLOv8 (Sound Object Detection)\nI use the data pointed out by the competition host to train the yolov8 object detection model, and select an offset that will increase the number of object detections as much as possible.\nAlso, by minimizing the bounding box and reliability-weighted IoU for the frequency masking, Frequency Masking effectively hides noise.\n\n#### Sample Inference\n![Inferece Sample Sound Object Detection](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2F058db543cb9d88e54a699785bd32c3b2%2F__results___3_2.png?generation=1685058390267527&alt=media)\n\nReference Notebook\n[1]https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-prepare-annot-data/notebook\n[2]https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectiontrain/notebook\n[3]https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-batch-predict/notebook\n[4]https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectionanalysis/notebook\n\n## ・Various Mixup Strategy\nBy using various mix-up methods, we tried to make it impossible to choose unnatural pairs. Details of the method are described below.\n\n### ・Geometrical Mixup\nAssuming that geographically close samples should have similar background noise, we created a model that mixes up geographically close samples without learning the background noise.\n\nIf the distance matrix is properly constructed, the amount of calculation will be O(n^2), so the code is written so that the parts that match up to the second decimal place of the longitude and latitude are treated as the same group.\n\n### ・Resonance Mixup (Inspired by collaborative filtering)\nChoosing a pair that resonates easily as a mixup destination is important for creating a more natural sound. Implemented a mixup that preferentially selects bird samples that are likely to resonate based on the collaborative filtering idea of the recommender system.\n\n\n### ・Inner mixup for geometrical clustering space\nGroups that are geographically close and have the same label do not change the label characteristics even if they share inputs and mixup. Therefore, we clustered such groups and made it possible to perform Inputmixup freely.\n\n# **Models**\nMy models are based on the BirdCLEF 2021 2nd place solution using backbone, eca_nfnet_l0.\n\n# **Dataset**\n1. Bird CLEF 2023\n2. Bird CLEF 2022/2021/2020 (Pretraining)\n3. ff1010bird_nocall\n4. [Zendo data (YOLOv8) ](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358)\n\n# **CV setup**\nI thought it would be difficult to create a multifold model and check the LB due to resource limitations, so I did a simple test-train split instead of out-of-fold.\n\nFor the test data, we built validation by splitting 30 seconds into 5 second intervals and bootstrap sampling to the same size as the LB dataset size (8400 samples).\n\nFurthermore, we found that using only the data with the primary label showed behavior similar to LB, so we excluded the data with the secondary label from the validation and calculated.\n\nInference Kernel: https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel\nGithub link: https://github.com/furu-kaggle/birdclef2023\n\n# Result\n|commit_id | single inference kernel| LB | PB |tag|\n| --- | --- | --- | --- |--- |\n|4a57a984d1f4101b9f470095e6a6fae0a6cc5af2|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=130154067) | 0.82628 | **0.73777** |ecanfnet_nmel128_best1|\n|55557aea59976c5befc11cbcd077f22c1d8289c0|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129843959) | 0.82475 | 0.73301 |ecanfnet_nmel128_best3|\n|ccc435294a9f5f9989f5865def93da0817c3b030|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129843959) | **0.82786** | 0.73341 |ecanfnet_nmel128_best2|\n|c9e077c172368c31c2d1974e0288554ebd014b2f|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=130288082) | 0.82279 | 0.72918 |ecanfnet_nmel64_best2|\n|a3791c561544adcc37507d5c27ec98e9d1c603a7|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129584055) | 0.82172 | 0.73223 |ecanfnet_nmel64_best1|\n|ea70ecc95de1625212465b8e0148bc996e0237f9|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129564172) | 0.82149 | 0.72850 |ecanfnet_nmel64_best3|\n|ecanfnet_nmel128_best1~3(300 stride pred) + ecanfnet_nmel64_best1~3 ensemble|[link](https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel?scriptVersionId=130834386) | **0.83616** | **0.75285** | best_submission|",
      "votes": null
    },
    {
      "id": "2279369",
      "postDate": "05/29/2023 10:11:56",
      "content": "<p>Congratulations on winning the eighth place in the competition and winning a gold medal! Your plan has benefited me a lot. By the way, in your Github, what do the 6 folders represent? I would like to reproduce it according to your report. Let me show you the structure of the experiment so that I can understand your thinking more deeply.</p>",
      "rawMarkdown": "Congratulations on winning the eighth place in the competition and winning a gold medal! Your plan has benefited me a lot. By the way, in your Github, what do the 6 folders represent? I would like to reproduce it according to your report. Let me show you the structure of the experiment so that I can understand your thinking more deeply.",
      "votes": null
    },
    {
      "id": "2279402",
      "postDate": "05/29/2023 10:47:44",
      "content": "<p><a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> <br>\nThank you for your appreciation. github<br>\n6 folders are linked to commit_id of Result. For example, 4a57a98…script represents the training script for 4a57a984d1f4101b9f470095e6a6fae0a6cc5af2. Please ask again if you have any questions.</p>",
      "rawMarkdown": "gentlezdh \nThank you for your appreciation. github\n6 folders are linked to commit_id of Result. For example, 4a57a98...script represents the training script for 4a57a984d1f4101b9f470095e6a6fae0a6cc5af2. Please ask again if you have any questions.",
      "votes": null
    },
    {
      "id": "2279685",
      "postDate": "05/29/2023 14:21:51",
      "content": "<p>Thank you for your answer. I am looking forward to discussing your plan with you by email or just asking here. This is really helpful for an undergraduate.</p>",
      "rawMarkdown": "Thank you for your answer. I am looking forward to discussing your plan with you by email or just asking here. This is really helpful for an undergraduate.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2279369,
      "author_name": "gentlezdh",
      "author_url": "",
      "post_date": "05/29/2023 10:11:56",
      "content": "<p>Congratulations on winning the eighth place in the competition and winning a gold medal! Your plan has benefited me a lot. By the way, in your Github, what do the 6 folders represent? I would like to reproduce it according to your report. Let me show you the structure of the experiment so that I can understand your thinking more deeply.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2279402,
          "author_name": "kunihikofurugori",
          "author_url": "",
          "post_date": "05/29/2023 10:47:44",
          "content": "<p><a href=\"https://www.kaggle.com/gentlezdh\" target=\"_blank\">@gentlezdh</a> <br>\nThank you for your appreciation. github<br>\n6 folders are linked to commit_id of Result. For example, 4a57a98…script represents the training script for 4a57a984d1f4101b9f470095e6a6fae0a6cc5af2. Please ask again if you have any questions.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2279685,
              "author_name": "gentlezdh",
              "author_url": "",
              "post_date": "05/29/2023 14:21:51",
              "content": "<p>Thank you for your answer. I am looking forward to discussing your plan with you by email or just asking here. This is really helpful for an undergraduate.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2274037": "Congratulations to all the winners!\nOur gratitude extends to Kaggle and the Cornell Lab of Ornithology for their exceptional efforts in organizing this intriguing competition. In particular, I was able to contribute to the improvement of accuracy with highly original methods, so I was able to enjoy the competition very much.\nAllow me to give a succinct overview of my solution. I’ll provide a more detailed update on the solution in the following days.\n\n# **Summary**\nIn my solution, only one type of backbone is used, and multimodal data augmentation techniques such as object detection, recommendation systems, and NLP are used to increase diversity and improve model accuracy.\n\n## Preprocess Pipeline Summary\nWe built a preprocessing pipeline that prevents overfitting by devising preprocessing so that the combinations that are actually possible while maintaining the diversity of the data. In particular, in preprocessing for Sound Object Detection, weighting the signal density of the spectrum made it easier to pick out the areas where the contribution of secondary labels is included. Thanks to that, the training accuracy has improved even when the value of the secondary labels is large.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2Ffc17c767400faa897953a3034352d16e%2F.drawio%20(1).png?generation=1685242284411201&alt=media)\n\n## Training Pipeline\nIn order to grasp the features of bird calls with various periodicities, build a pipeline that first learns with a long spectrum input, then lowers the time axis direction for each epoch, and finally learns local features. Did. Preprocessing by SOD prevents noise-like offsets from being selected even for local spectra, so training is stable.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2F090bc781738d7cfe844801e0bcd23f6e%2F8thplacesolution-2.drawio.png?generation=1685260705814564&alt=media)\n\ncode detail: https://www.kaggle.com/code/kunihikofurugori/8th-solution-down-scaling-detail/notebook?scriptVersionId=131297676\n\n# **MultiModal Data Augments Part(in Detailed)**\nCode detail: https://www.kaggle.com/code/kunihikofurugori/8th-solution-notebook/notebook\n### ・Time Segment DownScaling Augmentation (Global to Local)\nRegarding the mini-batch method represented by the 2nd place solution of birdclef2021, we trained to learn local features while pre-learning global features by training to evolve the number of mini-batches from 15 to 1 at each epoch.\n\nBy first learning the global features and then learning the local features, it becomes possible to predict the local features after understanding the global features.\n\n### ・Weak to Strong Preprocessing by YOLOv8 (Sound Object Detection)\nI use the data pointed out by the competition host to train the yolov8 object detection model, and select an offset that will increase the number of object detections as much as possible.\nAlso, by minimizing the bounding box and reliability-weighted IoU for the frequency masking, Frequency Masking effectively hides noise.\n\n#### Sample Inference\n![Inferece Sample Sound Object Detection](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6220676%2F058db543cb9d88e54a699785bd32c3b2%2F__results___3_2.png?generation=1685058390267527&alt=media)\n\nReference Notebook\n[1]https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-prepare-annot-data/notebook\n[2]https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectiontrain/notebook\n[3]https://www.kaggle.com/code/kunihikofurugori/sound-object-detection-batch-predict/notebook\n[4]https://www.kaggle.com/code/kunihikofurugori/birdclef2023-soundobjectdetectionanalysis/notebook\n\n## ・Various Mixup Strategy\nBy using various mix-up methods, we tried to make it impossible to choose unnatural pairs. Details of the method are described below.\n\n### ・Geometrical Mixup\nAssuming that geographically close samples should have similar background noise, we created a model that mixes up geographically close samples without learning the background noise.\n\nIf the distance matrix is properly constructed, the amount of calculation will be O(n^2), so the code is written so that the parts that match up to the second decimal place of the longitude and latitude are treated as the same group.\n\n### ・Resonance Mixup (Inspired by collaborative filtering)\nChoosing a pair that resonates easily as a mixup destination is important for creating a more natural sound. Implemented a mixup that preferentially selects bird samples that are likely to resonate based on the collaborative filtering idea of the recommender system.\n\n\n### ・Inner mixup for geometrical clustering space\nGroups that are geographically close and have the same label do not change the label characteristics even if they share inputs and mixup. Therefore, we clustered such groups and made it possible to perform Inputmixup freely.\n\n# **Models**\nMy models are based on the BirdCLEF 2021 2nd place solution using backbone, eca_nfnet_l0.\n\n# **Dataset**\n1. Bird CLEF 2023\n2. Bird CLEF 2022/2021/2020 (Pretraining)\n3. ff1010bird_nocall\n4. [Zendo data (YOLOv8) ](https://www.kaggle.com/competitions/birdclef-2023/discussion/394358)\n\n# **CV setup**\nI thought it would be difficult to create a multifold model and check the LB due to resource limitations, so I did a simple test-train split instead of out-of-fold.\n\nFor the test data, we built validation by splitting 30 seconds into 5 second intervals and bootstrap sampling to the same size as the LB dataset size (8400 samples).\n\nFurthermore, we found that using only the data with the primary label showed behavior similar to LB, so we excluded the data with the secondary label from the validation and calculated.\n\nInference Kernel: https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel\nGithub link: https://github.com/furu-kaggle/birdclef2023\n\n# Result\n|commit_id | single inference kernel| LB | PB |tag|\n| --- | --- | --- | --- |--- |\n|4a57a984d1f4101b9f470095e6a6fae0a6cc5af2|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=130154067) | 0.82628 | **0.73777** |ecanfnet_nmel128_best1|\n|55557aea59976c5befc11cbcd077f22c1d8289c0|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129843959) | 0.82475 | 0.73301 |ecanfnet_nmel128_best3|\n|ccc435294a9f5f9989f5865def93da0817c3b030|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129843959) | **0.82786** | 0.73341 |ecanfnet_nmel128_best2|\n|c9e077c172368c31c2d1974e0288554ebd014b2f|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=130288082) | 0.82279 | 0.72918 |ecanfnet_nmel64_best2|\n|a3791c561544adcc37507d5c27ec98e9d1c603a7|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129584055) | 0.82172 | 0.73223 |ecanfnet_nmel64_best1|\n|ea70ecc95de1625212465b8e0148bc996e0237f9|[link](https://www.kaggle.com/code/kunihikofurugori/best-single-model-lb-0-826-pb-0-73777?scriptVersionId=129564172) | 0.82149 | 0.72850 |ecanfnet_nmel64_best3|\n|ecanfnet_nmel128_best1~3(300 stride pred) + ecanfnet_nmel64_best1~3 ensemble|[link](https://www.kaggle.com/code/kunihikofurugori/8th-place-solution-inference-kernel?scriptVersionId=130834386) | **0.83616** | **0.75285** | best_submission|",
    "2279369": "Congratulations on winning the eighth place in the competition and winning a gold medal! Your plan has benefited me a lot. By the way, in your Github, what do the 6 folders represent? I would like to reproduce it according to your report. Let me show you the structure of the experiment so that I can understand your thinking more deeply.",
    "2279402": "gentlezdh \nThank you for your appreciation. github\n6 folders are linked to commit_id of Result. For example, 4a57a98...script represents the training script for 4a57a984d1f4101b9f470095e6a6fae0a6cc5af2. Please ask again if you have any questions.",
    "2279685": "Thank you for your answer. I am looking forward to discussing your plan with you by email or just asking here. This is really helpful for an undergraduate."
  },
  "source": "meta"
}