{
  "id": 220725,
  "title": "14th Place Solution - Binary Classification on cropped frequency x time",
  "url": "/competitions/rfcx-species-audio-detection/discussion/220725",
  "author_name": "Prateek",
  "post_date": "2021-02-19T09:57:12.280000",
  "votes": 19,
  "comment_count": 8,
  "views": 0,
  "content": "<p>First of all thanks to Kaggle and Hosts for organizing this competition.</p>\n<p>We ( me &amp; <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> ) were initially working in Cassava Leaf Disease Classification but thanks to this <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/212610\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> that gave us the direction that something different needs to be tried then what is going on in the public. On closer inspection, we found something similar to him. The frequency-time crops in the audios for a specie_id are almost constant (i.e. say for specie_id 23 - in most of the recordings audio frequency lies between 6459 and 11628, and its duration lasts for about 16 seconds). This gave us the idea of cropping out all the potential regions from the spectrogram and perform binary classification on them.</p>\n<p>Our approach can be summarized as -</p>\n<ol>\n<li>Crop images from spectrogram with frequency ranging between max-min frequency observed for a specie_id and with time duration 2 times the max duration observed for a specie_id </li>\n<li>Pre-Process: resize crops to size 128 x 256, scale between 0 and 1, and perform augmentation</li>\n<li>Train B0 binary classifier detecting the presence of specie (a single binary classifier - here we tracked back using the frequency information of the crop that which ID we are asking classifier to detect for)</li>\n<li>Generate Pseudo-labels</li>\n<li>Retrain</li>\n<li>Perform inference on the test, and take the mean of max n(in our case it was 3) probabilities observed for a specie_id in a recording as the probability of that specie_id</li>\n</ol>\n<p>Note: On the very first submission, a single model with the above approach gave us 0.921 as public LB (has private LB 0.927), then pseudo labeling and a little bit of blending took private LB to 0.948</p>\n<h2>Cropping</h2>\n<p>From each spectrogram, for each specie_id x songtype, we cropped out image sequences with the frequency range between min and max frequency observed for that specie_id x songtype, and then created image sequences with duration 2 times the max time interval the audio lasted for that specie_id x songtype in the train.<br>\n<img src=\"https://github.com/PRATEEKKUMARAGNIHOTRI/CMS-trigger/blob/master/images/RFCX1.png\" alt=\"Img\"><br>\n<a href=\"https://github.com/PRATEEKKUMARAGNIHOTRI/CMS-trigger/blob/master/images/RFCX1.png\" target=\"_blank\">Here</a> - In case img not visible</p>\n<h2>Augmentation</h2>\n<p>Along with adding random noise, we took a false positive sample of the same specie_id and added that to the audio sample. After this augmentation, the label of the recording id x specie id remained the same(i.e. a false-negative remained the false negative and a true positive remained the true positive)</p>",
  "messages": [
    {
      "id": 1210265,
      "postDate": "2021-02-19T09:57:12.280Z",
      "content": "<p>First of all thanks to Kaggle and Hosts for organizing this competition.</p>\n<p>We ( me &amp; <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> ) were initially working in Cassava Leaf Disease Classification but thanks to this <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/212610\" target=\"_blank\">post</a> by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> that gave us the direction that something different needs to be tried then what is going on in the public. On closer inspection, we found something similar to him. The frequency-time crops in the audios for a specie_id are almost constant (i.e. say for specie_id 23 - in most of the recordings audio frequency lies between 6459 and 11628, and its duration lasts for about 16 seconds). This gave us the idea of cropping out all the potential regions from the spectrogram and perform binary classification on them.</p>\n<p>Our approach can be summarized as -</p>\n<ol>\n<li>Crop images from spectrogram with frequency ranging between max-min frequency observed for a specie_id and with time duration 2 times the max duration observed for a specie_id </li>\n<li>Pre-Process: resize crops to size 128 x 256, scale between 0 and 1, and perform augmentation</li>\n<li>Train B0 binary classifier detecting the presence of specie (a single binary classifier - here we tracked back using the frequency information of the crop that which ID we are asking classifier to detect for)</li>\n<li>Generate Pseudo-labels</li>\n<li>Retrain</li>\n<li>Perform inference on the test, and take the mean of max n(in our case it was 3) probabilities observed for a specie_id in a recording as the probability of that specie_id</li>\n</ol>\n<p>Note: On the very first submission, a single model with the above approach gave us 0.921 as public LB (has private LB 0.927), then pseudo labeling and a little bit of blending took private LB to 0.948</p>\n<h2>Cropping</h2>\n<p>From each spectrogram, for each specie_id x songtype, we cropped out image sequences with the frequency range between min and max frequency observed for that specie_id x songtype, and then created image sequences with duration 2 times the max time interval the audio lasted for that specie_id x songtype in the train.<br>\n<img src=\"https://github.com/PRATEEKKUMARAGNIHOTRI/CMS-trigger/blob/master/images/RFCX1.png\" alt=\"Img\"><br>\n<a href=\"https://github.com/PRATEEKKUMARAGNIHOTRI/CMS-trigger/blob/master/images/RFCX1.png\" target=\"_blank\">Here</a> - In case img not visible</p>\n<h2>Augmentation</h2>\n<p>Along with adding random noise, we took a false positive sample of the same specie_id and added that to the audio sample. After this augmentation, the label of the recording id x specie id remained the same(i.e. a false-negative remained the false negative and a true positive remained the true positive)</p>",
      "rawMarkdown": "First of all thanks to Kaggle and Hosts for organizing this competition.\n\nWe ( me & @ks2019 ) were initially working in Cassava Leaf Disease Classification but thanks to this [post](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/212610) by @cpmpml that gave us the direction that something different needs to be tried then what is going on in the public. On closer inspection, we found something similar to him. The frequency-time crops in the audios for a specie_id are almost constant (i.e. say for specie_id 23 - in most of the recordings audio frequency lies between 6459 and 11628, and its duration lasts for about 16 seconds). This gave us the idea of cropping out all the potential regions from the spectrogram and perform binary classification on them.\n\nOur approach can be summarized as -\n1. Crop images from spectrogram with frequency ranging between max-min frequency observed for a specie_id and with time duration 2 times the max duration observed for a specie_id \n2. Pre-Process: resize crops to size 128 x 256, scale between 0 and 1, and perform augmentation\n3. Train B0 binary classifier detecting the presence of specie (a single binary classifier - here we tracked back using the frequency information of the crop that which ID we are asking classifier to detect for)\n4. Generate Pseudo-labels\n5. Retrain\n4. Perform inference on the test, and take the mean of max n(in our case it was 3) probabilities observed for a specie_id in a recording as the probability of that specie_id\n\nNote: On the very first submission, a single model with the above approach gave us 0.921 as public LB (has private LB 0.927), then pseudo labeling and a little bit of blending took private LB to 0.948\n\n## Cropping\nFrom each spectrogram, for each specie_id x songtype, we cropped out image sequences with the frequency range between min and max frequency observed for that specie_id x songtype, and then created image sequences with duration 2 times the max time interval the audio lasted for that specie_id x songtype in the train.\n![Img](https://github.com/PRATEEKKUMARAGNIHOTRI/CMS-trigger/blob/master/images/RFCX1.png)\n[Here](https://github.com/PRATEEKKUMARAGNIHOTRI/CMS-trigger/blob/master/images/RFCX1.png) - In case img not visible\n\n## Augmentation\nAlong with adding random noise, we took a false positive sample of the same specie_id and added that to the audio sample. After this augmentation, the label of the recording id x specie id remained the same(i.e. a false-negative remained the false negative and a true positive remained the true positive)",
      "votes": 19
    },
    {
      "id": 1322459,
      "postDate": "2021-05-25T13:11:39.943Z",
      "content": "<p>Thanks for the write-up and congrat for silver medal! <br>\nThis solution is pretty clever that nicely exploits the time x frequency properties of each species. <br>\nI have a few questions:</p>\n<ol>\n<li>If I understand your approach correctly, your model is a single binary classifier. And your binary classifier predicting 1 means \"The crop spectrogram contains at least one of the 24 species\". Is it correct?</li>\n<li>I am wondering how you would do inference with your binary classifier. Is it you divide a single recording into 24 spectrogram sequences (lets say T frames per sequence) and then you run prediction on each sequence (in total 24*T prediction calls)? While each series of sequence is bounded by f_max, f_min specific to the target species?</li>\n</ol>\n<p>I hope its not too late for me to write a comment and question here. I would really appreciate for your answers! </p>",
      "rawMarkdown": "Thanks for the write-up and congrat for silver medal! \nThis solution is pretty clever that nicely exploits the time x frequency properties of each species. \nI have a few questions:\n1. If I understand your approach correctly, your model is a single binary classifier. And your binary classifier predicting 1 means \"The crop spectrogram contains at least one of the 24 species\". Is it correct?\n2. I am wondering how you would do inference with your binary classifier. Is it you divide a single recording into 24 spectrogram sequences (lets say T frames per sequence) and then you run prediction on each sequence (in total 24*T prediction calls)? While each series of sequence is bounded by f_max, f_min specific to the target species?\n\nI hope its not too late for me to write a comment and question here. I would really appreciate for your answers! ",
      "votes": 1,
      "replies": [
        {
          "id": 1324257,
          "postDate": "2021-05-26T18:34:31.150Z",
          "content": "<p>Yes, your understanding is correct!<br>\nOnly instead of T frames per sequence, we cropped along the time axis too with <code>duration 2 times the max time interval the audio lasted for that specie_id x songtype in the train</code>. And the hop length was half of the sequence cropping length.</p>",
          "rawMarkdown": "Yes, your understanding is correct!\nOnly instead of T frames per sequence, we cropped along the time axis too with `duration 2 times the max time interval the audio lasted for that specie_id x songtype in the train`. And the hop length was half of the sequence cropping length.",
          "votes": 1
        },
        {
          "id": 1324497,
          "postDate": "2021-05-27T02:54:46.493Z",
          "content": "<p><a href=\"https://www.kaggle.com/prateekagnihotri\" target=\"_blank\">@prateekagnihotri</a> <br>\nThanks for the reply! Understood what you mean. <br>\nI had a doubt whether the classifier could learn well when all characteristics of all species x songtype are packed into 1 single class, but surprisingly results that you got is really impressive. <br>\nI would say its a simpler approach with more computationally cost on inference part. (i.e. you needa run prediction on a recording for (# species x # songtype x # cropped frame along time-axis) times.<br>\nProbably I will try to reproduce your approach (the basic part) in sometime!</p>",
          "rawMarkdown": "@prateekagnihotri \nThanks for the reply! Understood what you mean. \nI had a doubt whether the classifier could learn well when all characteristics of all species x songtype are packed into 1 single class, but surprisingly results that you got is really impressive. \nI would say its a simpler approach with more computationally cost on inference part. (i.e. you needa run prediction on a recording for (# species x # songtype x # cropped frame along time-axis) times.\nProbably I will try to reproduce your approach (the basic part) in sometime!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1212631,
      "postDate": "2021-02-21T12:21:20.260Z",
      "content": "<p>Thanks for sharing.<br>\nI have a question.<br>\nIn your binary classification('target data' vs 'other data'), dose 'other data' means fp data or other speicies data?</p>",
      "rawMarkdown": "Thanks for sharing.\nI have a question.\nIn your binary classification('target data' vs 'other data'), dose 'other data' means fp data or other speicies data?",
      "votes": 1,
      "replies": [
        {
          "id": 1212670,
          "postDate": "2021-02-21T13:18:40.053Z",
          "content": "<p>There is only one classifier and for that classifier, all tp-samples are one and all fp-sample are zero.</p>\n<p>During inference, if the model says one, then we track back using the frequency region info of the crop, for which class it is saying one.</p>",
          "rawMarkdown": "There is only one classifier and for that classifier, all tp-samples are one and all fp-sample are zero.\n\nDuring inference, if the model says one, then we track back using the frequency region info of the crop, for which class it is saying one.",
          "votes": 1
        },
        {
          "id": 1212714,
          "postDate": "2021-02-21T14:02:44.763Z",
          "content": "<p>Oh,I missunderstood you used 24 binary model like me. <br>\nGreat solution.</p>",
          "rawMarkdown": "Oh,I missunderstood you used 24 binary model like me. \nGreat solution.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1210374,
      "postDate": "2021-02-19T11:36:48.123Z",
      "content": "<p>Thanks for sharing and congrats on the result.  I also appreciate the acknowledgement very much!</p>",
      "rawMarkdown": "Thanks for sharing and congrats on the result.  I also appreciate the acknowledgement very much!\n",
      "votes": 1
    },
    {
      "id": 1324794,
      "postDate": "2021-05-27T08:16:36.890Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1322459,
      "author_name": "Alex Lau",
      "author_url": "",
      "post_date": "2021-05-25T13:11:39.943000",
      "content": "<p>Thanks for the write-up and congrat for silver medal! <br>\nThis solution is pretty clever that nicely exploits the time x frequency properties of each species. <br>\nI have a few questions:</p>\n<ol>\n<li>If I understand your approach correctly, your model is a single binary classifier. And your binary classifier predicting 1 means \"The crop spectrogram contains at least one of the 24 species\". Is it correct?</li>\n<li>I am wondering how you would do inference with your binary classifier. Is it you divide a single recording into 24 spectrogram sequences (lets say T frames per sequence) and then you run prediction on each sequence (in total 24*T prediction calls)? While each series of sequence is bounded by f_max, f_min specific to the target species?</li>\n</ol>\n<p>I hope its not too late for me to write a comment and question here. I would really appreciate for your answers! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1324257,
          "author_name": "Prateek",
          "author_url": "",
          "post_date": "2021-05-26T18:34:31.150000",
          "content": "<p>Yes, your understanding is correct!<br>\nOnly instead of T frames per sequence, we cropped along the time axis too with <code>duration 2 times the max time interval the audio lasted for that specie_id x songtype in the train</code>. And the hop length was half of the sequence cropping length.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1324497,
          "author_name": "Alex Lau",
          "author_url": "",
          "post_date": "2021-05-27T02:54:46.493000",
          "content": "<p><a href=\"https://www.kaggle.com/prateekagnihotri\" target=\"_blank\">@prateekagnihotri</a> <br>\nThanks for the reply! Understood what you mean. <br>\nI had a doubt whether the classifier could learn well when all characteristics of all species x songtype are packed into 1 single class, but surprisingly results that you got is really impressive. <br>\nI would say its a simpler approach with more computationally cost on inference part. (i.e. you needa run prediction on a recording for (# species x # songtype x # cropped frame along time-axis) times.<br>\nProbably I will try to reproduce your approach (the basic part) in sometime!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1212631,
      "author_name": "UEMU",
      "author_url": "",
      "post_date": "2021-02-21T12:21:20.260000",
      "content": "<p>Thanks for sharing.<br>\nI have a question.<br>\nIn your binary classification('target data' vs 'other data'), dose 'other data' means fp data or other speicies data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1212670,
          "author_name": "Prateek",
          "author_url": "",
          "post_date": "2021-02-21T13:18:40.053000",
          "content": "<p>There is only one classifier and for that classifier, all tp-samples are one and all fp-sample are zero.</p>\n<p>During inference, if the model says one, then we track back using the frequency region info of the crop, for which class it is saying one.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1212714,
          "author_name": "UEMU",
          "author_url": "",
          "post_date": "2021-02-21T14:02:44.763000",
          "content": "<p>Oh,I missunderstood you used 24 binary model like me. <br>\nGreat solution.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1210374,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-02-19T11:36:48.123000",
      "content": "<p>Thanks for sharing and congrats on the result.  I also appreciate the acknowledgement very much!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1324794,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-27T08:16:36.890000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1210265": "First of all thanks to Kaggle and Hosts for organizing this competition.\n\nWe ( me & @ks2019 ) were initially working in Cassava Leaf Disease Classification but thanks to this [post](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/212610) by @cpmpml that gave us the direction that something different needs to be tried then what is going on in the public. On closer inspection, we found something similar to him. The frequency-time crops in the audios for a specie_id are almost constant (i.e. say for specie_id 23 - in most of the recordings audio frequency lies between 6459 and 11628, and its duration lasts for about 16 seconds). This gave us the idea of cropping out all the potential regions from the spectrogram and perform binary classification on them.\n\nOur approach can be summarized as -\n1. Crop images from spectrogram with frequency ranging between max-min frequency observed for a specie_id and with time duration 2 times the max duration observed for a specie_id \n2. Pre-Process: resize crops to size 128 x 256, scale between 0 and 1, and perform augmentation\n3. Train B0 binary classifier detecting the presence of specie (a single binary classifier - here we tracked back using the frequency information of the crop that which ID we are asking classifier to detect for)\n4. Generate Pseudo-labels\n5. Retrain\n4. Perform inference on the test, and take the mean of max n(in our case it was 3) probabilities observed for a specie_id in a recording as the probability of that specie_id\n\nNote: On the very first submission, a single model with the above approach gave us 0.921 as public LB (has private LB 0.927), then pseudo labeling and a little bit of blending took private LB to 0.948\n\n## Cropping\nFrom each spectrogram, for each specie_id x songtype, we cropped out image sequences with the frequency range between min and max frequency observed for that specie_id x songtype, and then created image sequences with duration 2 times the max time interval the audio lasted for that specie_id x songtype in the train.\n![Img](https://github.com/PRATEEKKUMARAGNIHOTRI/CMS-trigger/blob/master/images/RFCX1.png)\n[Here](https://github.com/PRATEEKKUMARAGNIHOTRI/CMS-trigger/blob/master/images/RFCX1.png) - In case img not visible\n\n## Augmentation\nAlong with adding random noise, we took a false positive sample of the same specie_id and added that to the audio sample. After this augmentation, the label of the recording id x specie id remained the same(i.e. a false-negative remained the false negative and a true positive remained the true positive)",
    "1322459": "Thanks for the write-up and congrat for silver medal! \nThis solution is pretty clever that nicely exploits the time x frequency properties of each species. \nI have a few questions:\n1. If I understand your approach correctly, your model is a single binary classifier. And your binary classifier predicting 1 means \"The crop spectrogram contains at least one of the 24 species\". Is it correct?\n2. I am wondering how you would do inference with your binary classifier. Is it you divide a single recording into 24 spectrogram sequences (lets say T frames per sequence) and then you run prediction on each sequence (in total 24*T prediction calls)? While each series of sequence is bounded by f_max, f_min specific to the target species?\n\nI hope its not too late for me to write a comment and question here. I would really appreciate for your answers! ",
    "1212631": "Thanks for sharing.\nI have a question.\nIn your binary classification('target data' vs 'other data'), dose 'other data' means fp data or other speicies data?",
    "1210374": "Thanks for sharing and congrats on the result.  I also appreciate the acknowledgement very much!\n",
    "1324794": ""
  }
}