{
  "id": 245708,
  "title": "3rd place submission",
  "url": "/competitions/birdclef-2021/writeups/shiro-3rd-place-submission",
  "author_name": "",
  "post_date": "2021-06-12T01:05:42.957Z",
  "votes": 31,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, thank you to the Host, the Kaggle Team for having organised such a good competition. I am very happy to have participated in this one.</p>\n<p>In the following, I want to explain roughtly what was the approach I used.</p>\n<p><strong>Summary</strong>:</p>\n<p>The final solution was an ensembling of 18 checkpoints trained on different CNNS and fold. During training each models are trained on a clip of 20 seconds. A post processing is also applied durgin inference based on the prediction at the \"clip level\" (20 seconds) and the \"segment level\" (5 seconds)</p>\n<p><strong>Explanation:</strong></p>\n<p>The first observation we can make is that :</p>\n<ul>\n<li>For training, we have weak label. So we can use a supervised learning method based on stochastic gradient descent to minimise the loss function.</li>\n<li>For testing, we are looking to predicting every 5 seconds. Therefore, there is a gap, between the label given to the model during training (sample level) and during inference (5-second segment-level) .</li>\n</ul>\n<p>To overcome this issue, I implement  an architecture which allows getting prediction for a clip (20 seconds or more depending of your hardware) and also make prediction every 5 seconds in the clip. </p>\n<p><em>Why training on a clip of 20 seconds and not 5 seconds ?</em> <br>\nThe hypothesis I have made is this one : *training on small clip will introduce noise to our training. *<br>\nIndeed, we don’t know where and which birds are singing in the 5-second clip , so it is possible to select a clip with nocall for instance. Increasing the length of the clip, allow to reduce this noise.<br>\nThen, these clips are divided into segments of 5 secondes. For each of these segment, I create a mel-spectrogram based on the torchlibrosa package. Therefore, for a clip of 20 seconds, 4 spectrograms will be generated. We can use these spectrogram to feed a Deep Learning model.</p>\n<p><strong>Model :</strong></p>\n<p>The architecture of the model is composed of :</p>\n<ul>\n<li>A backbone based on SED model (CNN) which can use different architectures such as seresnet50, EfficientNetB2, EfficientNetB3, etc with the model coming from this github : <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a><br>\nthe backbone has two outputs : the first one is for classifying each segment (one output for every 5 seconds), the second one is for classifying each timestep of the segment. I use the second output  to decide if a bird is present in the 5 seconds or not for inference on test (I took the max probability across the time-step axis for each class). I used the first one for feeding the attention block of my \"sequential model\" used during training.</li>\n<li>An attention blocks in order to compute the final output (clip-level prediction) which will be needed for the training. The Attention block is applied on the first output of the backbone (the dimension should be (BS, 4, num_class) )<br>\nDuring the training, we are trying to minimise the output of the attention block with a cross entropy loss. We are trying here to predict the birds present in the 20 seconds clip. I use label smoothing of 0.05 as the label was still noisy. That could help to reduce overfitting. The metric used to select best checkpoint is the PR auc, as F1 score is based on precision and recall, I thought it would be a good idea to use PR AUC for each class then average it in order to have a robust estimation of the model.<br>\nDuring inference, I am using the outputs of each timesteps of the segments. The dimension should be (BS, 4, timesteps, num_classes). I take the max probability across the timesteps dimension and got a vector of (BS, 4, num_classes). These vector corresponds to the prediction for each 5-second segment inside the 20 seconds clip. Then a post-processing is applied, using the final output (prediction at 20-second clip-level).</li>\n</ul>\n<p><a href=\"https://ibb.co/hFfvZ6n\"><img src=\"https://i.ibb.co/nMCSnGV/Processus.jpg\" alt=\"Processus\"></a></p>\n<p>Different CNNs have been trained such as SeResNet50,  EfficientNetB2, EfficientNetB3, EfficientNetB4, EfficientNetB5, EfficientNetB6, EfficientNetB7. My initial goal was to train each of these model with a 5-folds strategy but because of hardware issues(too long to train), I did not train them on all folds. After that, I just ensemble these models by CNN type. </p>\n<p><strong>Post processing Inference :</strong></p>\n<p>To do that I create two ensemblings : one based on the final output (20 sec clip) and the second one based on the max of the prediction of the timesteps segments. Then, to decide if a bird appeared, I need to look first if the bird appeared in the ensembling's prediction of the max timesteps segments (based on a threshold t1) and if the bird also appeared in the ensembling's final prediction (based on a threshold t2)</p>\n<p><a href=\"https://ibb.co/mF6B9Mx\"><img src=\"https://i.ibb.co/JF5dvMh/ensembling.jpg\" alt=\"ensembling\"></a></p>\n<p>Data augmentation used : mixup (signal and spectrogram worked)</p>\n<p>Hardware : GTX 1080TI + Colab </p>",
  "messages": [
    {
      "id": "1345936",
      "postDate": "06/12/2021 01:02:39",
      "content": "<p>First of all, thank you to the Host, the Kaggle Team for having organised such a good competition. I am very happy to have participated in this one.</p>\n<p>In the following, I want to explain roughtly what was the approach I used.</p>\n<p><strong>Summary</strong>:</p>\n<p>The final solution was an ensembling of 18 checkpoints trained on different CNNS and fold. During training each models are trained on a clip of 20 seconds. A post processing is also applied durgin inference based on the prediction at the \"clip level\" (20 seconds) and the \"segment level\" (5 seconds)</p>\n<p><strong>Explanation:</strong></p>\n<p>The first observation we can make is that :</p>\n<ul>\n<li>For training, we have weak label. So we can use a supervised learning method based on stochastic gradient descent to minimise the loss function.</li>\n<li>For testing, we are looking to predicting every 5 seconds. Therefore, there is a gap, between the label given to the model during training (sample level) and during inference (5-second segment-level) .</li>\n</ul>\n<p>To overcome this issue, I implement  an architecture which allows getting prediction for a clip (20 seconds or more depending of your hardware) and also make prediction every 5 seconds in the clip. </p>\n<p><em>Why training on a clip of 20 seconds and not 5 seconds ?</em> <br>\nThe hypothesis I have made is this one : *training on small clip will introduce noise to our training. *<br>\nIndeed, we don’t know where and which birds are singing in the 5-second clip , so it is possible to select a clip with nocall for instance. Increasing the length of the clip, allow to reduce this noise.<br>\nThen, these clips are divided into segments of 5 secondes. For each of these segment, I create a mel-spectrogram based on the torchlibrosa package. Therefore, for a clip of 20 seconds, 4 spectrograms will be generated. We can use these spectrogram to feed a Deep Learning model.</p>\n<p><strong>Model :</strong></p>\n<p>The architecture of the model is composed of :</p>\n<ul>\n<li>A backbone based on SED model (CNN) which can use different architectures such as seresnet50, EfficientNetB2, EfficientNetB3, etc with the model coming from this github : <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a><br>\nthe backbone has two outputs : the first one is for classifying each segment (one output for every 5 seconds), the second one is for classifying each timestep of the segment. I use the second output  to decide if a bird is present in the 5 seconds or not for inference on test (I took the max probability across the time-step axis for each class). I used the first one for feeding the attention block of my \"sequential model\" used during training.</li>\n<li>An attention blocks in order to compute the final output (clip-level prediction) which will be needed for the training. The Attention block is applied on the first output of the backbone (the dimension should be (BS, 4, num_class) )<br>\nDuring the training, we are trying to minimise the output of the attention block with a cross entropy loss. We are trying here to predict the birds present in the 20 seconds clip. I use label smoothing of 0.05 as the label was still noisy. That could help to reduce overfitting. The metric used to select best checkpoint is the PR auc, as F1 score is based on precision and recall, I thought it would be a good idea to use PR AUC for each class then average it in order to have a robust estimation of the model.<br>\nDuring inference, I am using the outputs of each timesteps of the segments. The dimension should be (BS, 4, timesteps, num_classes). I take the max probability across the timesteps dimension and got a vector of (BS, 4, num_classes). These vector corresponds to the prediction for each 5-second segment inside the 20 seconds clip. Then a post-processing is applied, using the final output (prediction at 20-second clip-level).</li>\n</ul>\n<p><a href=\"https://ibb.co/hFfvZ6n\"><img src=\"https://i.ibb.co/nMCSnGV/Processus.jpg\" alt=\"Processus\"></a></p>\n<p>Different CNNs have been trained such as SeResNet50,  EfficientNetB2, EfficientNetB3, EfficientNetB4, EfficientNetB5, EfficientNetB6, EfficientNetB7. My initial goal was to train each of these model with a 5-folds strategy but because of hardware issues(too long to train), I did not train them on all folds. After that, I just ensemble these models by CNN type. </p>\n<p><strong>Post processing Inference :</strong></p>\n<p>To do that I create two ensemblings : one based on the final output (20 sec clip) and the second one based on the max of the prediction of the timesteps segments. Then, to decide if a bird appeared, I need to look first if the bird appeared in the ensembling's prediction of the max timesteps segments (based on a threshold t1) and if the bird also appeared in the ensembling's final prediction (based on a threshold t2)</p>\n<p><a href=\"https://ibb.co/mF6B9Mx\"><img src=\"https://i.ibb.co/JF5dvMh/ensembling.jpg\" alt=\"ensembling\"></a></p>\n<p>Data augmentation used : mixup (signal and spectrogram worked)</p>\n<p>Hardware : GTX 1080TI + Colab </p>",
      "rawMarkdown": "First of all, thank you to the Host, the Kaggle Team for having organised such a good competition. I am very happy to have participated in this one.\n\nIn the following, I want to explain roughtly what was the approach I used.\n\n**Summary**:\n\nThe final solution was an ensembling of 18 checkpoints trained on different CNNS and fold. During training each models are trained on a clip of 20 seconds. A post processing is also applied durgin inference based on the prediction at the \"clip level\" (20 seconds) and the \"segment level\" (5 seconds)\n\n**Explanation:**\n\n\nThe first observation we can make is that :\n-\tFor training, we have weak label. So we can use a supervised learning method based on stochastic gradient descent to minimise the loss function.\n-\tFor testing, we are looking to predicting every 5 seconds. Therefore, there is a gap, between the label given to the model during training (sample level) and during inference (5-second segment-level) .\n\nTo overcome this issue, I implement  an architecture which allows getting prediction for a clip (20 seconds or more depending of your hardware) and also make prediction every 5 seconds in the clip. \n\n*Why training on a clip of 20 seconds and not 5 seconds ?* \nThe hypothesis I have made is this one : *training on small clip will introduce noise to our training. *\nIndeed, we don’t know where and which birds are singing in the 5-second clip , so it is possible to select a clip with nocall for instance. Increasing the length of the clip, allow to reduce this noise.\nThen, these clips are divided into segments of 5 secondes. For each of these segment, I create a mel-spectrogram based on the torchlibrosa package. Therefore, for a clip of 20 seconds, 4 spectrograms will be generated. We can use these spectrogram to feed a Deep Learning model.\n\n**Model :**\n\nThe architecture of the model is composed of :\n-\tA backbone based on SED model (CNN) which can use different architectures such as seresnet50, EfficientNetB2, EfficientNetB3, etc with the model coming from this github : https://github.com/rwightman/pytorch-image-models\nthe backbone has two outputs : the first one is for classifying each segment (one output for every 5 seconds), the second one is for classifying each timestep of the segment. I use the second output  to decide if a bird is present in the 5 seconds or not for inference on test (I took the max probability across the time-step axis for each class). I used the first one for feeding the attention block of my \"sequential model\" used during training.\n-\tAn attention blocks in order to compute the final output (clip-level prediction) which will be needed for the training. The Attention block is applied on the first output of the backbone (the dimension should be (BS, 4, num_class) )\nDuring the training, we are trying to minimise the output of the attention block with a cross entropy loss. We are trying here to predict the birds present in the 20 seconds clip. I use label smoothing of 0.05 as the label was still noisy. That could help to reduce overfitting. The metric used to select best checkpoint is the PR auc, as F1 score is based on precision and recall, I thought it would be a good idea to use PR AUC for each class then average it in order to have a robust estimation of the model.\nDuring inference, I am using the outputs of each timesteps of the segments. The dimension should be (BS, 4, timesteps, num_classes). I take the max probability across the timesteps dimension and got a vector of (BS, 4, num_classes). These vector corresponds to the prediction for each 5-second segment inside the 20 seconds clip. Then a post-processing is applied, using the final output (prediction at 20-second clip-level).\n \n<a href=\"https://ibb.co/hFfvZ6n\"><img src=\"https://i.ibb.co/nMCSnGV/Processus.jpg\" alt=\"Processus\" border=\"0\"></a>\n\nDifferent CNNs have been trained such as SeResNet50,  EfficientNetB2, EfficientNetB3, EfficientNetB4, EfficientNetB5, EfficientNetB6, EfficientNetB7. My initial goal was to train each of these model with a 5-folds strategy but because of hardware issues(too long to train), I did not train them on all folds. After that, I just ensemble these models by CNN type. \n\n**Post processing Inference :**\n\nTo do that I create two ensemblings : one based on the final output (20 sec clip) and the second one based on the max of the prediction of the timesteps segments. Then, to decide if a bird appeared, I need to look first if the bird appeared in the ensembling's prediction of the max timesteps segments (based on a threshold t1) and if the bird also appeared in the ensembling's final prediction (based on a threshold t2)\n \n\n<a href=\"https://ibb.co/mF6B9Mx\"><img src=\"https://i.ibb.co/JF5dvMh/ensembling.jpg\" alt=\"ensembling\" border=\"0\"></a>\n\nData augmentation used : mixup (signal and spectrogram worked)\n\nHardware : GTX 1080TI + Colab",
      "votes": null
    },
    {
      "id": "1355967",
      "postDate": "06/18/2021 16:40:49",
      "content": "<p>Congratulations on a fantastic score, and thank you for the write-up, very interesting.<br>\nAny chance of sharing your code? </p>",
      "rawMarkdown": "Congratulations on a fantastic score, and thank you for the write-up, very interesting.\nAny chance of sharing your code?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1355967,
      "author_name": "botkop",
      "author_url": "",
      "post_date": "06/18/2021 16:40:49",
      "content": "<p>Congratulations on a fantastic score, and thank you for the write-up, very interesting.<br>\nAny chance of sharing your code? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1345936": "First of all, thank you to the Host, the Kaggle Team for having organised such a good competition. I am very happy to have participated in this one.\n\nIn the following, I want to explain roughtly what was the approach I used.\n\n**Summary**:\n\nThe final solution was an ensembling of 18 checkpoints trained on different CNNS and fold. During training each models are trained on a clip of 20 seconds. A post processing is also applied durgin inference based on the prediction at the \"clip level\" (20 seconds) and the \"segment level\" (5 seconds)\n\n**Explanation:**\n\n\nThe first observation we can make is that :\n-\tFor training, we have weak label. So we can use a supervised learning method based on stochastic gradient descent to minimise the loss function.\n-\tFor testing, we are looking to predicting every 5 seconds. Therefore, there is a gap, between the label given to the model during training (sample level) and during inference (5-second segment-level) .\n\nTo overcome this issue, I implement  an architecture which allows getting prediction for a clip (20 seconds or more depending of your hardware) and also make prediction every 5 seconds in the clip. \n\n*Why training on a clip of 20 seconds and not 5 seconds ?* \nThe hypothesis I have made is this one : *training on small clip will introduce noise to our training. *\nIndeed, we don’t know where and which birds are singing in the 5-second clip , so it is possible to select a clip with nocall for instance. Increasing the length of the clip, allow to reduce this noise.\nThen, these clips are divided into segments of 5 secondes. For each of these segment, I create a mel-spectrogram based on the torchlibrosa package. Therefore, for a clip of 20 seconds, 4 spectrograms will be generated. We can use these spectrogram to feed a Deep Learning model.\n\n**Model :**\n\nThe architecture of the model is composed of :\n-\tA backbone based on SED model (CNN) which can use different architectures such as seresnet50, EfficientNetB2, EfficientNetB3, etc with the model coming from this github : https://github.com/rwightman/pytorch-image-models\nthe backbone has two outputs : the first one is for classifying each segment (one output for every 5 seconds), the second one is for classifying each timestep of the segment. I use the second output  to decide if a bird is present in the 5 seconds or not for inference on test (I took the max probability across the time-step axis for each class). I used the first one for feeding the attention block of my \"sequential model\" used during training.\n-\tAn attention blocks in order to compute the final output (clip-level prediction) which will be needed for the training. The Attention block is applied on the first output of the backbone (the dimension should be (BS, 4, num_class) )\nDuring the training, we are trying to minimise the output of the attention block with a cross entropy loss. We are trying here to predict the birds present in the 20 seconds clip. I use label smoothing of 0.05 as the label was still noisy. That could help to reduce overfitting. The metric used to select best checkpoint is the PR auc, as F1 score is based on precision and recall, I thought it would be a good idea to use PR AUC for each class then average it in order to have a robust estimation of the model.\nDuring inference, I am using the outputs of each timesteps of the segments. The dimension should be (BS, 4, timesteps, num_classes). I take the max probability across the timesteps dimension and got a vector of (BS, 4, num_classes). These vector corresponds to the prediction for each 5-second segment inside the 20 seconds clip. Then a post-processing is applied, using the final output (prediction at 20-second clip-level).\n \n<a href=\"https://ibb.co/hFfvZ6n\"><img src=\"https://i.ibb.co/nMCSnGV/Processus.jpg\" alt=\"Processus\" border=\"0\"></a>\n\nDifferent CNNs have been trained such as SeResNet50,  EfficientNetB2, EfficientNetB3, EfficientNetB4, EfficientNetB5, EfficientNetB6, EfficientNetB7. My initial goal was to train each of these model with a 5-folds strategy but because of hardware issues(too long to train), I did not train them on all folds. After that, I just ensemble these models by CNN type. \n\n**Post processing Inference :**\n\nTo do that I create two ensemblings : one based on the final output (20 sec clip) and the second one based on the max of the prediction of the timesteps segments. Then, to decide if a bird appeared, I need to look first if the bird appeared in the ensembling's prediction of the max timesteps segments (based on a threshold t1) and if the bird also appeared in the ensembling's final prediction (based on a threshold t2)\n \n\n<a href=\"https://ibb.co/mF6B9Mx\"><img src=\"https://i.ibb.co/JF5dvMh/ensembling.jpg\" alt=\"ensembling\" border=\"0\"></a>\n\nData augmentation used : mixup (signal and spectrogram worked)\n\nHardware : GTX 1080TI + Colab",
    "1355967": "Congratulations on a fantastic score, and thank you for the write-up, very interesting.\nAny chance of sharing your code?"
  },
  "source": "meta"
}