{
  "id": 394162,
  "title": "Pytorch Lightning Baseline - LB 0.78 CV 0.831",
  "url": "/competitions/birdclef-2023/discussion/394162",
  "author_name": "",
  "post_date": "2023-03-12T12:43:49.019752100Z",
  "votes": 63,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Glad to see another BirdClef competition, also thankful to hosts for a much stable threshold free metric :))</p>\n<p>I am currently making my baseline public which is written in pytorch lightning, currently it scores 0.77 on the leaderboard which is also my current 5th-place submission. </p>\n<p><strong>Generate Spectrograms &amp; Make split:</strong> <a href=\"https://www.kaggle.com/code/nischaydnk/split-creating-melspecs-stage-1\" target=\"_blank\">notebook link</a><br>\n<strong>Training notebook:</strong> <a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-training-w-cmap\" target=\"_blank\">notebook link</a><br>\n<strong>Submission notebook:</strong>  <a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook\" target=\"_blank\">notebook link</a></p>\n<p>Train and validation split csv: <a href=\"https://www.kaggle.com/datasets/nischaydnk/bc2023-train-val-df\" target=\"_blank\">https://www.kaggle.com/datasets/nischaydnk/bc2023-train-val-df</a></p>\n<p>It took around <strong>75</strong> minutes to generate spectrograms, just <strong>12</strong> minutes for training &amp; <strong>29</strong> minutes for submission.</p>\n<p><strong>Some additional things to add:</strong></p>\n<ul>\n<li><p>Backbone: efficientnet b0 </p></li>\n<li><p>Split: <strong>Stratified train test split</strong> based on primary birds [80% - 20%] , as there is not enough submission time for multiple folds.</p></li>\n<li><p>There are few birds with just 1 sample, so I only put them in training data so that model could learn.</p></li>\n<li><p>Generated spectrograms with a duration of 5 seconds.</p></li>\n<li><p><strong>CV :</strong><br>\nC-MAP score padding = 5 ---&gt; 0.81<br>\nC-MAP score padding = 3 ---&gt; 0.755<br>\nAP score -----&gt; 0.784</p></li>\n</ul>\n<p>Please feel free to reach out incase you find any part of code hard to understand, I usually forget to add markdown comments in notebooks :))<br>\nA few of the code snippets are based on <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> sharing in BirdClef 2021.</p>\n<p>happy kaggling!!</p>",
  "messages": [
    {
      "id": "2178440",
      "postDate": "03/12/2023 12:43:49",
      "content": "<p>Glad to see another BirdClef competition, also thankful to hosts for a much stable threshold free metric :))</p>\n<p>I am currently making my baseline public which is written in pytorch lightning, currently it scores 0.77 on the leaderboard which is also my current 5th-place submission. </p>\n<p><strong>Generate Spectrograms &amp; Make split:</strong> <a href=\"https://www.kaggle.com/code/nischaydnk/split-creating-melspecs-stage-1\" target=\"_blank\">notebook link</a><br>\n<strong>Training notebook:</strong> <a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-training-w-cmap\" target=\"_blank\">notebook link</a><br>\n<strong>Submission notebook:</strong>  <a href=\"https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook\" target=\"_blank\">notebook link</a></p>\n<p>Train and validation split csv: <a href=\"https://www.kaggle.com/datasets/nischaydnk/bc2023-train-val-df\" target=\"_blank\">https://www.kaggle.com/datasets/nischaydnk/bc2023-train-val-df</a></p>\n<p>It took around <strong>75</strong> minutes to generate spectrograms, just <strong>12</strong> minutes for training &amp; <strong>29</strong> minutes for submission.</p>\n<p><strong>Some additional things to add:</strong></p>\n<ul>\n<li><p>Backbone: efficientnet b0 </p></li>\n<li><p>Split: <strong>Stratified train test split</strong> based on primary birds [80% - 20%] , as there is not enough submission time for multiple folds.</p></li>\n<li><p>There are few birds with just 1 sample, so I only put them in training data so that model could learn.</p></li>\n<li><p>Generated spectrograms with a duration of 5 seconds.</p></li>\n<li><p><strong>CV :</strong><br>\nC-MAP score padding = 5 ---&gt; 0.81<br>\nC-MAP score padding = 3 ---&gt; 0.755<br>\nAP score -----&gt; 0.784</p></li>\n</ul>\n<p>Please feel free to reach out incase you find any part of code hard to understand, I usually forget to add markdown comments in notebooks :))<br>\nA few of the code snippets are based on <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a> sharing in BirdClef 2021.</p>\n<p>happy kaggling!!</p>",
      "rawMarkdown": "Glad to see another BirdClef competition, also thankful to hosts for a much stable threshold free metric :))\n\nI am currently making my baseline public which is written in pytorch lightning, currently it scores 0.77 on the leaderboard which is also my current 5th-place submission. \n\n\n**Generate Spectrograms & Make split:** [notebook link](https://www.kaggle.com/code/nischaydnk/split-creating-melspecs-stage-1)\n**Training notebook:** [notebook link](https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-training-w-cmap)\n**Submission notebook:**  [notebook link](https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook)\n\nTrain and validation split csv: https://www.kaggle.com/datasets/nischaydnk/bc2023-train-val-df\n\n\nIt took around **75** minutes to generate spectrograms, just **12** minutes for training & **29** minutes for submission.\n\n**Some additional things to add:**\n - Backbone: efficientnet b0 \n - Split: **Stratified train test split** based on primary birds [80% - 20%] , as there is not enough submission time for multiple folds.\n - There are few birds with just 1 sample, so I only put them in training data so that model could learn.\n - Generated spectrograms with a duration of 5 seconds.\n \n - **CV :**\nC-MAP score padding = 5 ---> 0.81\nC-MAP score padding = 3 ---> 0.755\nAP score -----> 0.784\n\nPlease feel free to reach out incase you find any part of code hard to understand, I usually forget to add markdown comments in notebooks :))\nA few of the code snippets are based on @kneroma sharing in BirdClef 2021.\n\nhappy kaggling!!",
      "votes": null
    },
    {
      "id": "2178725",
      "postDate": "03/12/2023 16:57:02",
      "content": "<p>nice to see pytorch baseline,  many thanks for sharing, seems we need a single strong model ,  ensemble not work here.</p>",
      "rawMarkdown": "nice to see pytorch baseline,  many thanks for sharing, seems we need a single strong model ,  ensemble not work here.",
      "votes": null
    },
    {
      "id": "2178790",
      "postDate": "03/12/2023 17:57:32",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a> , yeah I think we can't fit more than 4-5 smaller backbone models this time. </p>",
      "rawMarkdown": "Thanks @evilpsycho42 , yeah I think we can't fit more than 4-5 smaller backbone models this time.",
      "votes": null
    },
    {
      "id": "2179738",
      "postDate": "03/13/2023 11:19:32",
      "content": "<p><strong>Update 1:</strong><br>\nDid few minor changes and replaced model with efn b1, good improvement in score. I think there is going to be a tradeoff between using denser architecture vs multiple models ensemble in this competition. </p>\n<p>Architecture: tf_effcientnet_b0_ns --&gt; tf_efficientnet_b1_ns<br>\nCV: 0.81 --&gt; 0.83<br>\nLB: 0.77 --&gt; 0.78<br>\nSubmission time: 42 min</p>",
      "rawMarkdown": "**Update 1:**\nDid few minor changes and replaced model with efn b1, good improvement in score. I think there is going to be a tradeoff between using denser architecture vs multiple models ensemble in this competition. \n\nArchitecture: tf_effcientnet_b0_ns --> tf_efficientnet_b1_ns\nCV: 0.81 --> 0.83\nLB: 0.77 --> 0.78\nSubmission time: 42 min",
      "votes": null
    },
    {
      "id": "2181672",
      "postDate": "03/14/2023 17:18:02",
      "content": "<p>Hi KKY, thanks for your great baseline!<br>\nBased on your work, I tried to use only pytorch for more flexibility. But both training loss and validation loss(BCEWithLogit) quickly decreased to around 0.02 during training, which means the predictions were all near zero, but the metrics is not good(around 0.5). Considering the class number, I think it is reasonable that loss decreases quickly because there are so many zeros in the label.</p>\n<p>I wonder how you keep the loss around 0.4 duiring the training, or do you have any insight about the phenomenon I described above?</p>",
      "rawMarkdown": "Hi KKY, thanks for your great baseline!\nBased on your work, I tried to use only pytorch for more flexibility. But both training loss and validation loss(BCEWithLogit) quickly decreased to around 0.02 during training, which means the predictions were all near zero, but the metrics is not good(around 0.5). Considering the class number, I think it is reasonable that loss decreases quickly because there are so many zeros in the label.\n\nI wonder how you keep the loss around 0.4 duiring the training, or do you have any insight about the phenomenon I described above?",
      "votes": null
    },
    {
      "id": "2181837",
      "postDate": "03/14/2023 19:20:59",
      "content": "<p>Hi, the loss which is being calculated and used in my pipeline is actually CrossEntropyLoss along with Mixup, which is why you could see higher value. Mistakenly, I kept validation loss as BCELogitsLoss, which gives misleading and unstable valid loss, you can change that to crossentropy as well.</p>",
      "rawMarkdown": "Hi, the loss which is being calculated and used in my pipeline is actually CrossEntropyLoss along with Mixup, which is why you could see higher value. Mistakenly, I kept validation loss as BCELogitsLoss, which gives misleading and unstable valid loss, you can change that to crossentropy as well.",
      "votes": null
    },
    {
      "id": "2182386",
      "postDate": "03/15/2023 04:56:19",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> , thanks a lot for this great kernel. I have a question - why are you not resizing the mel spectograms to be of same size and how it an image of different sizes giving output after being sent in the model?</p>",
      "rawMarkdown": "Hey @nischaydnk , thanks a lot for this great kernel. I have a question - why are you not resizing the mel spectograms to be of same size and how it an image of different sizes giving output after being sent in the model?",
      "votes": null
    },
    {
      "id": "2183150",
      "postDate": "03/15/2023 13:47:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> , Can I ask why in the <code>get_audios_as_images</code> function in the <a target=\"_blank\">Split &amp; Creating MelSpecs [Stage 1]</a> you chose <code>step=Config.duration*0.666*Config.sampling_rate</code>. why did you choose <code>5*0.666*32000</code>? Does <code>step</code> matter much when you convert audio into images? Thank you</p>",
      "rawMarkdown": "Hi @nischaydnk , Can I ask why in the `get_audios_as_images` function in the [Split & Creating MelSpecs [Stage 1]](uhttps://www.kaggle.com/code/nischaydnk/split-creating-melspecs-stage-1#Let's-Think-of-a-Better-Split-Methodrl) you chose `step=Config.duration*0.666*Config.sampling_rate`. why did you choose `5*0.666*32000`? Does `step` matter much when you convert audio into images? Thank you",
      "votes": null
    },
    {
      "id": "2183217",
      "postDate": "03/15/2023 14:27:58",
      "content": "<p>I have answered in the notebook's comment of yours. Also sharing here if someone else want to know.</p>\n<p>timm models are compatible with any image size except for few transformer based models. So, given any image size: final feature map comes out to be in the format of [Batch size, Num features,  Final H, Final W] , then we apply adaptive average pooling layer which converts the latter two dimension --&gt; [1,1]. Hence, for different image size or aspect ratios, it will generate different feature map but output after pooling will be of same shape. </p>\n<p>Coming onto your question, It is upto experiments, you may resize the image to maintain aspect ratio 1:1 , or tune the hop length to increase/decrease spectrogram dimensions, and can compare results. Our goal is to minimise the loss of information about the signals present in the data, either we try to manually do preprocessing or we just leave it up to the CNNs to do the work. </p>",
      "rawMarkdown": "I have answered in the notebook's comment of yours. Also sharing here if someone else want to know.\n\n timm models are compatible with any image size except for few transformer based models. So, given any image size: final feature map comes out to be in the format of [Batch size, Num features,  Final H, Final W] , then we apply adaptive average pooling layer which converts the latter two dimension --> [1,1]. Hence, for different image size or aspect ratios, it will generate different feature map but output after pooling will be of same shape. \n\nComing onto your question, It is upto experiments, you may resize the image to maintain aspect ratio 1:1 , or tune the hop length to increase/decrease spectrogram dimensions, and can compare results. Our goal is to minimise the loss of information about the signals present in the data, either we try to manually do preprocessing or we just leave it up to the CNNs to do the work.",
      "votes": null
    },
    {
      "id": "2183241",
      "postDate": "03/15/2023 14:44:57",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/locbaop\" target=\"_blank\">@locbaop</a> , if I remember correctly steps are used to determine the overlapping ratio of regions while splitting the audios in several windows. So, 0.66 basically means 2̶/̶3̶r̶d̶  1/3rd (thanks <a href=\"https://www.kaggle.com/andtem2000\" target=\"_blank\">@andtem2000</a>) part of the previous segment is in the next segment. Idea is to extract better temporal information, you can still tune the parameter and find out results, I used it as this was commonly used in previous competitions. </p>",
      "rawMarkdown": "Hey @locbaop , if I remember correctly steps are used to determine the overlapping ratio of regions while splitting the audios in several windows. So, 0.66 basically means 2̶/̶3̶r̶d̶  1/3rd (thanks @andtem2000) part of the previous segment is in the next segment. Idea is to extract better temporal information, you can still tune the parameter and find out results, I used it as this was commonly used in previous competitions.",
      "votes": null
    },
    {
      "id": "2183252",
      "postDate": "03/15/2023 14:51:00",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    },
    {
      "id": "2184368",
      "postDate": "03/16/2023 10:17:15",
      "content": "<p>\"0.66 basically means 2/3rd part of the previous segment is in the next segment.\"<br>\n1/3</p>",
      "rawMarkdown": "\"0.66 basically means 2/3rd part of the previous segment is in the next segment.\"\n1/3",
      "votes": null
    },
    {
      "id": "2186761",
      "postDate": "03/18/2023 04:57:13",
      "content": "<p>Oh, I see the loss. Thanks! I managed to reproduce your result!</p>",
      "rawMarkdown": "Oh, I see the loss. Thanks! I managed to reproduce your result!",
      "votes": null
    },
    {
      "id": "2229936",
      "postDate": "04/21/2023 20:35:55",
      "content": "<p>When you say \"CV 0.831\" does \"CV\" mean Cross Validation? If I understand correctly, Cross Validation is a technique for partitioning the dataset and then running multiple training experiments with multiple models to determine the best hyperparams, before retraining a new model on the entire training data using those hyperparams. But here (and in most posts/notebooks that mention \"CV X\" on kaggle) it sounds like it just means \"test score\"--how well a single model does on a single held out test set. Or am I misunderstanding something?</p>",
      "rawMarkdown": "When you say \"CV 0.831\" does \"CV\" mean Cross Validation? If I understand correctly, Cross Validation is a technique for partitioning the dataset and then running multiple training experiments with multiple models to determine the best hyperparams, before retraining a new model on the entire training data using those hyperparams. But here (and in most posts/notebooks that mention \"CV X\" on kaggle) it sounds like it just means \"test score\"--how well a single model does on a single held out test set. Or am I misunderstanding something?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2178725,
      "author_name": "evilpsycho42",
      "author_url": "",
      "post_date": "03/12/2023 16:57:02",
      "content": "<p>nice to see pytorch baseline,  many thanks for sharing, seems we need a single strong model ,  ensemble not work here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2178790,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "03/12/2023 17:57:32",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a> , yeah I think we can't fit more than 4-5 smaller backbone models this time. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2179738,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "03/13/2023 11:19:32",
      "content": "<p><strong>Update 1:</strong><br>\nDid few minor changes and replaced model with efn b1, good improvement in score. I think there is going to be a tradeoff between using denser architecture vs multiple models ensemble in this competition. </p>\n<p>Architecture: tf_effcientnet_b0_ns --&gt; tf_efficientnet_b1_ns<br>\nCV: 0.81 --&gt; 0.83<br>\nLB: 0.77 --&gt; 0.78<br>\nSubmission time: 42 min</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2181672,
      "author_name": "honglihang",
      "author_url": "",
      "post_date": "03/14/2023 17:18:02",
      "content": "<p>Hi KKY, thanks for your great baseline!<br>\nBased on your work, I tried to use only pytorch for more flexibility. But both training loss and validation loss(BCEWithLogit) quickly decreased to around 0.02 during training, which means the predictions were all near zero, but the metrics is not good(around 0.5). Considering the class number, I think it is reasonable that loss decreases quickly because there are so many zeros in the label.</p>\n<p>I wonder how you keep the loss around 0.4 duiring the training, or do you have any insight about the phenomenon I described above?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2181837,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "03/14/2023 19:20:59",
          "content": "<p>Hi, the loss which is being calculated and used in my pipeline is actually CrossEntropyLoss along with Mixup, which is why you could see higher value. Mistakenly, I kept validation loss as BCELogitsLoss, which gives misleading and unstable valid loss, you can change that to crossentropy as well.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2186761,
              "author_name": "honglihang",
              "author_url": "",
              "post_date": "03/18/2023 04:57:13",
              "content": "<p>Oh, I see the loss. Thanks! I managed to reproduce your result!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2182386,
      "author_name": "hackingpirate",
      "author_url": "",
      "post_date": "03/15/2023 04:56:19",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> , thanks a lot for this great kernel. I have a question - why are you not resizing the mel spectograms to be of same size and how it an image of different sizes giving output after being sent in the model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2183217,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "03/15/2023 14:27:58",
          "content": "<p>I have answered in the notebook's comment of yours. Also sharing here if someone else want to know.</p>\n<p>timm models are compatible with any image size except for few transformer based models. So, given any image size: final feature map comes out to be in the format of [Batch size, Num features,  Final H, Final W] , then we apply adaptive average pooling layer which converts the latter two dimension --&gt; [1,1]. Hence, for different image size or aspect ratios, it will generate different feature map but output after pooling will be of same shape. </p>\n<p>Coming onto your question, It is upto experiments, you may resize the image to maintain aspect ratio 1:1 , or tune the hop length to increase/decrease spectrogram dimensions, and can compare results. Our goal is to minimise the loss of information about the signals present in the data, either we try to manually do preprocessing or we just leave it up to the CNNs to do the work. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2183252,
              "author_name": "hackingpirate",
              "author_url": "",
              "post_date": "03/15/2023 14:51:00",
              "content": "<p>Thanks a lot!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2183150,
      "author_name": "locbaop",
      "author_url": "",
      "post_date": "03/15/2023 13:47:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> , Can I ask why in the <code>get_audios_as_images</code> function in the <a target=\"_blank\">Split &amp; Creating MelSpecs [Stage 1]</a> you chose <code>step=Config.duration*0.666*Config.sampling_rate</code>. why did you choose <code>5*0.666*32000</code>? Does <code>step</code> matter much when you convert audio into images? Thank you</p>",
      "votes": null,
      "replies": [
        {
          "id": 2183241,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "03/15/2023 14:44:57",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/locbaop\" target=\"_blank\">@locbaop</a> , if I remember correctly steps are used to determine the overlapping ratio of regions while splitting the audios in several windows. So, 0.66 basically means 2̶/̶3̶r̶d̶  1/3rd (thanks <a href=\"https://www.kaggle.com/andtem2000\" target=\"_blank\">@andtem2000</a>) part of the previous segment is in the next segment. Idea is to extract better temporal information, you can still tune the parameter and find out results, I used it as this was commonly used in previous competitions. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2184368,
              "author_name": "andtem2000",
              "author_url": "",
              "post_date": "03/16/2023 10:17:15",
              "content": "<p>\"0.66 basically means 2/3rd part of the previous segment is in the next segment.\"<br>\n1/3</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2229936,
      "author_name": "robbynevels",
      "author_url": "",
      "post_date": "04/21/2023 20:35:55",
      "content": "<p>When you say \"CV 0.831\" does \"CV\" mean Cross Validation? If I understand correctly, Cross Validation is a technique for partitioning the dataset and then running multiple training experiments with multiple models to determine the best hyperparams, before retraining a new model on the entire training data using those hyperparams. But here (and in most posts/notebooks that mention \"CV X\" on kaggle) it sounds like it just means \"test score\"--how well a single model does on a single held out test set. Or am I misunderstanding something?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2178440": "Glad to see another BirdClef competition, also thankful to hosts for a much stable threshold free metric :))\n\nI am currently making my baseline public which is written in pytorch lightning, currently it scores 0.77 on the leaderboard which is also my current 5th-place submission. \n\n\n**Generate Spectrograms & Make split:** [notebook link](https://www.kaggle.com/code/nischaydnk/split-creating-melspecs-stage-1)\n**Training notebook:** [notebook link](https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-training-w-cmap)\n**Submission notebook:**  [notebook link](https://www.kaggle.com/code/nischaydnk/birdclef-2023-pytorch-lightning-inference/notebook)\n\nTrain and validation split csv: https://www.kaggle.com/datasets/nischaydnk/bc2023-train-val-df\n\n\nIt took around **75** minutes to generate spectrograms, just **12** minutes for training & **29** minutes for submission.\n\n**Some additional things to add:**\n - Backbone: efficientnet b0 \n - Split: **Stratified train test split** based on primary birds [80% - 20%] , as there is not enough submission time for multiple folds.\n - There are few birds with just 1 sample, so I only put them in training data so that model could learn.\n - Generated spectrograms with a duration of 5 seconds.\n \n - **CV :**\nC-MAP score padding = 5 ---> 0.81\nC-MAP score padding = 3 ---> 0.755\nAP score -----> 0.784\n\nPlease feel free to reach out incase you find any part of code hard to understand, I usually forget to add markdown comments in notebooks :))\nA few of the code snippets are based on @kneroma sharing in BirdClef 2021.\n\nhappy kaggling!!",
    "2178725": "nice to see pytorch baseline,  many thanks for sharing, seems we need a single strong model ,  ensemble not work here.",
    "2178790": "Thanks @evilpsycho42 , yeah I think we can't fit more than 4-5 smaller backbone models this time.",
    "2179738": "**Update 1:**\nDid few minor changes and replaced model with efn b1, good improvement in score. I think there is going to be a tradeoff between using denser architecture vs multiple models ensemble in this competition. \n\nArchitecture: tf_effcientnet_b0_ns --> tf_efficientnet_b1_ns\nCV: 0.81 --> 0.83\nLB: 0.77 --> 0.78\nSubmission time: 42 min",
    "2181672": "Hi KKY, thanks for your great baseline!\nBased on your work, I tried to use only pytorch for more flexibility. But both training loss and validation loss(BCEWithLogit) quickly decreased to around 0.02 during training, which means the predictions were all near zero, but the metrics is not good(around 0.5). Considering the class number, I think it is reasonable that loss decreases quickly because there are so many zeros in the label.\n\nI wonder how you keep the loss around 0.4 duiring the training, or do you have any insight about the phenomenon I described above?",
    "2181837": "Hi, the loss which is being calculated and used in my pipeline is actually CrossEntropyLoss along with Mixup, which is why you could see higher value. Mistakenly, I kept validation loss as BCELogitsLoss, which gives misleading and unstable valid loss, you can change that to crossentropy as well.",
    "2182386": "Hey @nischaydnk , thanks a lot for this great kernel. I have a question - why are you not resizing the mel spectograms to be of same size and how it an image of different sizes giving output after being sent in the model?",
    "2183150": "Hi @nischaydnk , Can I ask why in the `get_audios_as_images` function in the [Split & Creating MelSpecs [Stage 1]](uhttps://www.kaggle.com/code/nischaydnk/split-creating-melspecs-stage-1#Let's-Think-of-a-Better-Split-Methodrl) you chose `step=Config.duration*0.666*Config.sampling_rate`. why did you choose `5*0.666*32000`? Does `step` matter much when you convert audio into images? Thank you",
    "2183217": "I have answered in the notebook's comment of yours. Also sharing here if someone else want to know.\n\n timm models are compatible with any image size except for few transformer based models. So, given any image size: final feature map comes out to be in the format of [Batch size, Num features,  Final H, Final W] , then we apply adaptive average pooling layer which converts the latter two dimension --> [1,1]. Hence, for different image size or aspect ratios, it will generate different feature map but output after pooling will be of same shape. \n\nComing onto your question, It is upto experiments, you may resize the image to maintain aspect ratio 1:1 , or tune the hop length to increase/decrease spectrogram dimensions, and can compare results. Our goal is to minimise the loss of information about the signals present in the data, either we try to manually do preprocessing or we just leave it up to the CNNs to do the work.",
    "2183241": "Hey @locbaop , if I remember correctly steps are used to determine the overlapping ratio of regions while splitting the audios in several windows. So, 0.66 basically means 2̶/̶3̶r̶d̶  1/3rd (thanks @andtem2000) part of the previous segment is in the next segment. Idea is to extract better temporal information, you can still tune the parameter and find out results, I used it as this was commonly used in previous competitions.",
    "2183252": "Thanks a lot!",
    "2184368": "\"0.66 basically means 2/3rd part of the previous segment is in the next segment.\"\n1/3",
    "2186761": "Oh, I see the loss. Thanks! I managed to reproduce your result!",
    "2229936": "When you say \"CV 0.831\" does \"CV\" mean Cross Validation? If I understand correctly, Cross Validation is a technique for partitioning the dataset and then running multiple training experiments with multiple models to determine the best hyperparams, before retraining a new model on the entire training data using those hyperparams. But here (and in most posts/notebooks that mention \"CV X\" on kaggle) it sounds like it just means \"test score\"--how well a single model does on a single held out test set. Or am I misunderstanding something?"
  },
  "source": "meta"
}