{
  "id": 183222,
  "title": "36th place solution",
  "url": "/competitions/birdsong-recognition/writeups/oriental-honey-buzzard-36th-place-solution",
  "author_name": "",
  "post_date": "2020-09-16T02:12:27.378006500Z",
  "votes": 19,
  "comment_count": 6,
  "views": 0,
  "content": "<p>At first, congratulations to the winners and all participants who finished this competition. And also thanks to <a href=\"https://www.kaggle.com/shonenkov\" target=\"_blank\">@shonenkov</a> for kindly posting the kernel for submission.</p>\n<p>In the early stage of this competition, I was worried about the shake in private LB. But I continued with this competition, changed my thought as it is a relatively stable. I think this task was relevant to real issues and very interesiting. I would like to thank <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>,  <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a> and Kaggle.</p>\n<p>So let me shamelessly share a simple solution.<br>\nThe figure below is the overview of my model pipeline. As I'm sure there are already some great solutions out there, and will be more to come, so I'd like to give just a few points of my own.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F479538%2Fd51398618f29983e9f7c73c327f42112%2Fbirdcall_model-pipeline_01.png?generation=1600220214358611&amp;alt=media\" alt=\"\"></p>\n<p>1) Event aware extraction<br>\nAs the participants noticed, not all the time of an audio had bird voices. Therefore, if we randomly extracted some parts of an audio (e.g. 5 sec or so), there would be no birdsong at all, resulting in a form of mislabeling for training. To mitigate that, I used a naive algorithm to extract the signal parts of the audio as shown below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F479538%2Ff84b1e09ab70cd92ce87b957d69b38f5%2Fbirdcall_model-pipeline_02.png?generation=1600220869680424&amp;alt=media\" alt=\"\"></p>\n<p>2) LogMel Mixup<br>\nWhen I mixed 2 LogMels, once I convert it back into the power domain and then did a logarithmic transformation again. This is because LogMels are logarithmic values, and just linear summing is not the sum of the powers.<br>\na * log(X1) + b * log(X2) != log(a * X1 + b * X2)<br>\nAnd labels were not scaled with the mixup coefficients used in features, but an union of them.</p>\n<p>3) Binary Model<br>\nI made a binary model (ResNet18) to classify call/nocall audio chunks. This slightly improved my score in publicLB, but as a result, it was a big improvement in privateLB.</p>\n<p>4) Multi Label Model (Multi Task Learning)<br>\nAlthough the primary label was provided in the data for this competition, it was clear that there were actually other bird calls in the background that fell under the competition's predictive label. So, in training the model, I did multitasking learning by splitting the model top into two parts, one for the primary label and the other for the background. And I trained the models as the learning strategy below.</p>\n<p>Step1: primary only<br>\nStep2: primary + background<br>\nStep3: psuedo soft labeling using background predictions</p>\n<p>Last step3 improved my score by 0.004 in public and 0.006 in private, respectively.</p>\n<p>Now, that's what I'm going to share with you. I'll see you at the next competition somewhere else. <br>\nUntil next time!</p>",
  "messages": [
    {
      "id": "1012264",
      "postDate": "09/16/2020 02:12:27",
      "content": "<p>At first, congratulations to the winners and all participants who finished this competition. And also thanks to <a href=\"https://www.kaggle.com/shonenkov\" target=\"_blank\">@shonenkov</a> for kindly posting the kernel for submission.</p>\n<p>In the early stage of this competition, I was worried about the shake in private LB. But I continued with this competition, changed my thought as it is a relatively stable. I think this task was relevant to real issues and very interesiting. I would like to thank <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a>,  <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>, <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a> and Kaggle.</p>\n<p>So let me shamelessly share a simple solution.<br>\nThe figure below is the overview of my model pipeline. As I'm sure there are already some great solutions out there, and will be more to come, so I'd like to give just a few points of my own.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F479538%2Fd51398618f29983e9f7c73c327f42112%2Fbirdcall_model-pipeline_01.png?generation=1600220214358611&amp;alt=media\" alt=\"\"></p>\n<p>1) Event aware extraction<br>\nAs the participants noticed, not all the time of an audio had bird voices. Therefore, if we randomly extracted some parts of an audio (e.g. 5 sec or so), there would be no birdsong at all, resulting in a form of mislabeling for training. To mitigate that, I used a naive algorithm to extract the signal parts of the audio as shown below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F479538%2Ff84b1e09ab70cd92ce87b957d69b38f5%2Fbirdcall_model-pipeline_02.png?generation=1600220869680424&amp;alt=media\" alt=\"\"></p>\n<p>2) LogMel Mixup<br>\nWhen I mixed 2 LogMels, once I convert it back into the power domain and then did a logarithmic transformation again. This is because LogMels are logarithmic values, and just linear summing is not the sum of the powers.<br>\na * log(X1) + b * log(X2) != log(a * X1 + b * X2)<br>\nAnd labels were not scaled with the mixup coefficients used in features, but an union of them.</p>\n<p>3) Binary Model<br>\nI made a binary model (ResNet18) to classify call/nocall audio chunks. This slightly improved my score in publicLB, but as a result, it was a big improvement in privateLB.</p>\n<p>4) Multi Label Model (Multi Task Learning)<br>\nAlthough the primary label was provided in the data for this competition, it was clear that there were actually other bird calls in the background that fell under the competition's predictive label. So, in training the model, I did multitasking learning by splitting the model top into two parts, one for the primary label and the other for the background. And I trained the models as the learning strategy below.</p>\n<p>Step1: primary only<br>\nStep2: primary + background<br>\nStep3: psuedo soft labeling using background predictions</p>\n<p>Last step3 improved my score by 0.004 in public and 0.006 in private, respectively.</p>\n<p>Now, that's what I'm going to share with you. I'll see you at the next competition somewhere else. <br>\nUntil next time!</p>",
      "rawMarkdown": "At first, congratulations to the winners and all participants who finished this competition. And also thanks to @shonenkov for kindly posting the kernel for submission.\n\nIn the early stage of this competition, I was worried about the shake in private LB. But I continued with this competition, changed my thought as it is a relatively stable. I think this task was relevant to real issues and very interesiting. I would like to thank @stefankahl,  @tomdenton, @holgerklinck and Kaggle.\n\n  \nSo let me shamelessly share a simple solution.\nThe figure below is the overview of my model pipeline. As I'm sure there are already some great solutions out there, and will be more to come, so I'd like to give just a few points of my own.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F479538%2Fd51398618f29983e9f7c73c327f42112%2Fbirdcall_model-pipeline_01.png?generation=1600220214358611&alt=media)\n\n  \n1) Event aware extraction\nAs the participants noticed, not all the time of an audio had bird voices. Therefore, if we randomly extracted some parts of an audio (e.g. 5 sec or so), there would be no birdsong at all, resulting in a form of mislabeling for training. To mitigate that, I used a naive algorithm to extract the signal parts of the audio as shown below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F479538%2Ff84b1e09ab70cd92ce87b957d69b38f5%2Fbirdcall_model-pipeline_02.png?generation=1600220869680424&alt=media)\n\n\n2) LogMel Mixup\nWhen I mixed 2 LogMels, once I convert it back into the power domain and then did a logarithmic transformation again. This is because LogMels are logarithmic values, and just linear summing is not the sum of the powers.\na * log(X1) + b * log(X2) != log(a * X1 + b * X2)\nAnd labels were not scaled with the mixup coefficients used in features, but an union of them.\n\n3) Binary Model\nI made a binary model (ResNet18) to classify call/nocall audio chunks. This slightly improved my score in publicLB, but as a result, it was a big improvement in privateLB.\n\n4) Multi Label Model (Multi Task Learning)\nAlthough the primary label was provided in the data for this competition, it was clear that there were actually other bird calls in the background that fell under the competition's predictive label. So, in training the model, I did multitasking learning by splitting the model top into two parts, one for the primary label and the other for the background. And I trained the models as the learning strategy below.\n\nStep1: primary only\nStep2: primary + background\nStep3: psuedo soft labeling using background predictions\n\nLast step3 improved my score by 0.004 in public and 0.006 in private, respectively.\n\n\nNow, that's what I'm going to share with you. I'll see you at the next competition somewhere else. \nUntil next time!",
      "votes": null
    },
    {
      "id": "1012351",
      "postDate": "09/16/2020 03:32:06",
      "content": "<p>Your model definitely is not a simple solution. It is well thought out, wish you could have seen the results if you used bigger models like resnest50 or Efficientnet. Sorry for your tropical room, but at least it is a sauna for free =D Congratulations on medal!</p>",
      "rawMarkdown": "Your model definitely is not a simple solution. It is well thought out, wish you could have seen the results if you used bigger models like resnest50 or Efficientnet. Sorry for your tropical room, but at least it is a sauna for free =D Congratulations on medal!",
      "votes": null
    },
    {
      "id": "1012367",
      "postDate": "09/16/2020 03:45:51",
      "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> <br>\nThank you for your warm comment :)<br>\nCongratulations, too!</p>\n<p>In Japan, it was very very hot this summer. This made me decide to try this competition with a light model. But I’m also curious same as you, how my solution can be improved with more complicated models, like resnext or others. Even if my room become a perfect sauna:-) I will check it later with a slightly deeper model (I’m sorry this is the best I can do in this season) which I did not use, although had implemented.</p>",
      "rawMarkdown": "returnofsputnik \nThank you for your warm comment :)\nCongratulations, too!\n\nIn Japan, it was very very hot this summer. This made me decide to try this competition with a light model. But I’m also curious same as you, how my solution can be improved with more complicated models, like resnext or others. Even if my room become a perfect sauna:-) I will check it later with a slightly deeper model (I’m sorry this is the best I can do in this season) which I did not use, although had implemented.",
      "votes": null
    },
    {
      "id": "1012465",
      "postDate": "09/16/2020 05:37:02",
      "content": "<p><a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a>, thank you for describing your approach. I'm curious how much LogMel Mixup got improvement against simple MixUp? I tried to experiment with it a while back, but in my initial tests it didn't give any improvement, while it must be a right way of doing MixUp for mel spetrograms.</p>",
      "rawMarkdown": "maxwell110, thank you for describing your approach. I'm curious how much LogMel Mixup got improvement against simple MixUp? I tried to experiment with it a while back, but in my initial tests it didn't give any improvement, while it must be a right way of doing MixUp for mel spetrograms.",
      "votes": null
    },
    {
      "id": "1012477",
      "postDate": "09/16/2020 05:45:01",
      "content": "<p>I tried it. Just MixUp gave the best rating. !<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3577701%2Ff41fcbbcd4f025ad66e74a6598a3d8a5%2Fs.png?generation=1600235161280755&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I tried it. Just MixUp gave the best rating. !![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3577701%2Ff41fcbbcd4f025ad66e74a6598a3d8a5%2Fs.png?generation=1600235161280755&alt=media)",
      "votes": null
    },
    {
      "id": "1012849",
      "postDate": "09/16/2020 10:50:30",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> <br>\nThank you for comments. Let me answer as much as I can understand.</p>\n<blockquote>\n  <p>I'm curious how much LogMel Mixup got improvement against simple MixUp?</p>\n</blockquote>\n<p>I experimented this during the early stages of tiral and error. So do not know the exact degree of improvement. But I did the same thing <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> did above, and I found if one of the main signals was small, the signals became extremely small after mixup and became invisible.<br>\nI suppose this is probably a matter of normalizing after mixup and setting of hyper parameters used in mixup coefficients: parameters of beta distribution for mixup coefficients.</p>",
      "rawMarkdown": "iafoss @vlomme \nThank you for comments. Let me answer as much as I can understand.\n\n> I'm curious how much LogMel Mixup got improvement against simple MixUp?\n\nI experimented this during the early stages of tiral and error. So do not know the exact degree of improvement. But I did the same thing @vlomme did above, and I found if one of the main signals was small, the signals became extremely small after mixup and became invisible.\nI suppose this is probably a matter of normalizing after mixup and setting of hyper parameters used in mixup coefficients: parameters of beta distribution for mixup coefficients.",
      "votes": null
    },
    {
      "id": "1013326",
      "postDate": "09/16/2020 16:33:37",
      "content": "<p>The meaning of the original mixup on mel spectrograms is convolution of two signals: <code>s = a*log(s1) + (1-a)*log(s2) = log(s1^a*s2^(1-a)) = log(FFT(conv(x1^a,x2^(1-a))))</code>, where x1,x2 are the original waves. It's different from <code>s = log(FFT(a*x1 + (1-a)*x2))</code>, which is more natural, and it is one u have used. </p>",
      "rawMarkdown": "The meaning of the original mixup on mel spectrograms is convolution of two signals: `s = a*log(s1) + (1-a)*log(s2) = log(s1^a*s2^(1-a)) = log(FFT(conv(x1^a,x2^(1-a))))`, where x1,x2 are the original waves. It's different from `s = log(FFT(a*x1 + (1-a)*x2))`, which is more natural, and it is one u have used.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012351,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "09/16/2020 03:32:06",
      "content": "<p>Your model definitely is not a simple solution. It is well thought out, wish you could have seen the results if you used bigger models like resnest50 or Efficientnet. Sorry for your tropical room, but at least it is a sauna for free =D Congratulations on medal!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012367,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "09/16/2020 03:45:51",
          "content": "<p><a href=\"https://www.kaggle.com/returnofsputnik\" target=\"_blank\">@returnofsputnik</a> <br>\nThank you for your warm comment :)<br>\nCongratulations, too!</p>\n<p>In Japan, it was very very hot this summer. This made me decide to try this competition with a light model. But I’m also curious same as you, how my solution can be improved with more complicated models, like resnext or others. Even if my room become a perfect sauna:-) I will check it later with a slightly deeper model (I’m sorry this is the best I can do in this season) which I did not use, although had implemented.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012465,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "09/16/2020 05:37:02",
      "content": "<p><a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a>, thank you for describing your approach. I'm curious how much LogMel Mixup got improvement against simple MixUp? I tried to experiment with it a while back, but in my initial tests it didn't give any improvement, while it must be a right way of doing MixUp for mel spetrograms.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012477,
          "author_name": "vlomme",
          "author_url": "",
          "post_date": "09/16/2020 05:45:01",
          "content": "<p>I tried it. Just MixUp gave the best rating. !<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3577701%2Ff41fcbbcd4f025ad66e74a6598a3d8a5%2Fs.png?generation=1600235161280755&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012849,
          "author_name": "maxwell110",
          "author_url": "",
          "post_date": "09/16/2020 10:50:30",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> <br>\nThank you for comments. Let me answer as much as I can understand.</p>\n<blockquote>\n  <p>I'm curious how much LogMel Mixup got improvement against simple MixUp?</p>\n</blockquote>\n<p>I experimented this during the early stages of tiral and error. So do not know the exact degree of improvement. But I did the same thing <a href=\"https://www.kaggle.com/vlomme\" target=\"_blank\">@vlomme</a> did above, and I found if one of the main signals was small, the signals became extremely small after mixup and became invisible.<br>\nI suppose this is probably a matter of normalizing after mixup and setting of hyper parameters used in mixup coefficients: parameters of beta distribution for mixup coefficients.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1013326,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "09/16/2020 16:33:37",
          "content": "<p>The meaning of the original mixup on mel spectrograms is convolution of two signals: <code>s = a*log(s1) + (1-a)*log(s2) = log(s1^a*s2^(1-a)) = log(FFT(conv(x1^a,x2^(1-a))))</code>, where x1,x2 are the original waves. It's different from <code>s = log(FFT(a*x1 + (1-a)*x2))</code>, which is more natural, and it is one u have used. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1012264": "At first, congratulations to the winners and all participants who finished this competition. And also thanks to @shonenkov for kindly posting the kernel for submission.\n\nIn the early stage of this competition, I was worried about the shake in private LB. But I continued with this competition, changed my thought as it is a relatively stable. I think this task was relevant to real issues and very interesiting. I would like to thank @stefankahl,  @tomdenton, @holgerklinck and Kaggle.\n\n  \nSo let me shamelessly share a simple solution.\nThe figure below is the overview of my model pipeline. As I'm sure there are already some great solutions out there, and will be more to come, so I'd like to give just a few points of my own.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F479538%2Fd51398618f29983e9f7c73c327f42112%2Fbirdcall_model-pipeline_01.png?generation=1600220214358611&alt=media)\n\n  \n1) Event aware extraction\nAs the participants noticed, not all the time of an audio had bird voices. Therefore, if we randomly extracted some parts of an audio (e.g. 5 sec or so), there would be no birdsong at all, resulting in a form of mislabeling for training. To mitigate that, I used a naive algorithm to extract the signal parts of the audio as shown below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F479538%2Ff84b1e09ab70cd92ce87b957d69b38f5%2Fbirdcall_model-pipeline_02.png?generation=1600220869680424&alt=media)\n\n\n2) LogMel Mixup\nWhen I mixed 2 LogMels, once I convert it back into the power domain and then did a logarithmic transformation again. This is because LogMels are logarithmic values, and just linear summing is not the sum of the powers.\na * log(X1) + b * log(X2) != log(a * X1 + b * X2)\nAnd labels were not scaled with the mixup coefficients used in features, but an union of them.\n\n3) Binary Model\nI made a binary model (ResNet18) to classify call/nocall audio chunks. This slightly improved my score in publicLB, but as a result, it was a big improvement in privateLB.\n\n4) Multi Label Model (Multi Task Learning)\nAlthough the primary label was provided in the data for this competition, it was clear that there were actually other bird calls in the background that fell under the competition's predictive label. So, in training the model, I did multitasking learning by splitting the model top into two parts, one for the primary label and the other for the background. And I trained the models as the learning strategy below.\n\nStep1: primary only\nStep2: primary + background\nStep3: psuedo soft labeling using background predictions\n\nLast step3 improved my score by 0.004 in public and 0.006 in private, respectively.\n\n\nNow, that's what I'm going to share with you. I'll see you at the next competition somewhere else. \nUntil next time!",
    "1012351": "Your model definitely is not a simple solution. It is well thought out, wish you could have seen the results if you used bigger models like resnest50 or Efficientnet. Sorry for your tropical room, but at least it is a sauna for free =D Congratulations on medal!",
    "1012367": "returnofsputnik \nThank you for your warm comment :)\nCongratulations, too!\n\nIn Japan, it was very very hot this summer. This made me decide to try this competition with a light model. But I’m also curious same as you, how my solution can be improved with more complicated models, like resnext or others. Even if my room become a perfect sauna:-) I will check it later with a slightly deeper model (I’m sorry this is the best I can do in this season) which I did not use, although had implemented.",
    "1012465": "maxwell110, thank you for describing your approach. I'm curious how much LogMel Mixup got improvement against simple MixUp? I tried to experiment with it a while back, but in my initial tests it didn't give any improvement, while it must be a right way of doing MixUp for mel spetrograms.",
    "1012477": "I tried it. Just MixUp gave the best rating. !![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3577701%2Ff41fcbbcd4f025ad66e74a6598a3d8a5%2Fs.png?generation=1600235161280755&alt=media)",
    "1012849": "iafoss @vlomme \nThank you for comments. Let me answer as much as I can understand.\n\n> I'm curious how much LogMel Mixup got improvement against simple MixUp?\n\nI experimented this during the early stages of tiral and error. So do not know the exact degree of improvement. But I did the same thing @vlomme did above, and I found if one of the main signals was small, the signals became extremely small after mixup and became invisible.\nI suppose this is probably a matter of normalizing after mixup and setting of hyper parameters used in mixup coefficients: parameters of beta distribution for mixup coefficients.",
    "1013326": "The meaning of the original mixup on mel spectrograms is convolution of two signals: `s = a*log(s1) + (1-a)*log(s2) = log(s1^a*s2^(1-a)) = log(FFT(conv(x1^a,x2^(1-a))))`, where x1,x2 are the original waves. It's different from `s = log(FFT(a*x1 + (1-a)*x2))`, which is more natural, and it is one u have used."
  },
  "source": "meta"
}