{
  "id": 410876,
  "title": "How picky can we be about the training samples?",
  "url": "/competitions/birdclef-2023/discussion/410876",
  "author_name": "",
  "post_date": "2023-05-16T20:45:00.644368700Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I've based a lot of my hopes in this competition on producing a really clean, and somewhat better balanced dataset.  So I built a separate detector, trained on a fraction of the training data, including a no-call class, with extra data for that.  Then I ran that through all the training data, built myself a completely new dataset based on 8 second chunks that were very likely to contain the bird of interest.  And less likely to contain the secondary bird, or a no-call, chose fewer samples from the most common classes.</p>\n<p>The new dataset trains brilliantly,  I get private CMAP5 and LRAP scores into the .90's with lots of augmentation and no pre-training.</p>\n<p>But…..    It scores horribly on submission, 0.76-0.78   No better than the original detector.   Either there is a bug in my inference code, that I've tried quite hard to find already.   Or this strategy is actually wrong.   Have I just picked out all the samples that were easy to train on, and ignored the hard ones that my network really needed to learn on?  Or maybe I have moved any bias learned by the first detector onto the second?  </p>\n<p>I'm running out of ideas at this point.  Just doing an experiment where I remove any selection based on the primary bird probability score, and down-prioritize samples that score high on no-call, or secondary bird probabilities.  </p>",
  "messages": [
    {
      "id": "2262332",
      "postDate": "05/16/2023 20:45:00",
      "content": "<p>I've based a lot of my hopes in this competition on producing a really clean, and somewhat better balanced dataset.  So I built a separate detector, trained on a fraction of the training data, including a no-call class, with extra data for that.  Then I ran that through all the training data, built myself a completely new dataset based on 8 second chunks that were very likely to contain the bird of interest.  And less likely to contain the secondary bird, or a no-call, chose fewer samples from the most common classes.</p>\n<p>The new dataset trains brilliantly,  I get private CMAP5 and LRAP scores into the .90's with lots of augmentation and no pre-training.</p>\n<p>But…..    It scores horribly on submission, 0.76-0.78   No better than the original detector.   Either there is a bug in my inference code, that I've tried quite hard to find already.   Or this strategy is actually wrong.   Have I just picked out all the samples that were easy to train on, and ignored the hard ones that my network really needed to learn on?  Or maybe I have moved any bias learned by the first detector onto the second?  </p>\n<p>I'm running out of ideas at this point.  Just doing an experiment where I remove any selection based on the primary bird probability score, and down-prioritize samples that score high on no-call, or secondary bird probabilities.  </p>",
      "rawMarkdown": "I've based a lot of my hopes in this competition on producing a really clean, and somewhat better balanced dataset.  So I built a separate detector, trained on a fraction of the training data, including a no-call class, with extra data for that.  Then I ran that through all the training data, built myself a completely new dataset based on 8 second chunks that were very likely to contain the bird of interest.  And less likely to contain the secondary bird, or a no-call, chose fewer samples from the most common classes.\n\nThe new dataset trains brilliantly,  I get private CMAP5 and LRAP scores into the .90's with lots of augmentation and no pre-training.\n\nBut.....    It scores horribly on submission, 0.76-0.78   No better than the original detector.   Either there is a bug in my inference code, that I've tried quite hard to find already.   Or this strategy is actually wrong.   Have I just picked out all the samples that were easy to train on, and ignored the hard ones that my network really needed to learn on?  Or maybe I have moved any bias learned by the first detector onto the second?  \n\nI'm running out of ideas at this point.  Just doing an experiment where I remove any selection based on the primary bird probability score, and down-prioritize samples that score high on no-call, or secondary bird probabilities.",
      "votes": null
    },
    {
      "id": "2262516",
      "postDate": "05/17/2023 02:03:39",
      "content": "<p>Thats very interesting, thx for sharing, might have saved me a lot of time since I thought that implementing this would really help as well.<br>\nI can't bring up any explaination to why that didn't work tho, it's the same for me, any intuition based experience I try doesn't work at all..</p>",
      "rawMarkdown": "Thats very interesting, thx for sharing, might have saved me a lot of time since I thought that implementing this would really help as well.\nI can't bring up any explaination to why that didn't work tho, it's the same for me, any intuition based experience I try doesn't work at all..",
      "votes": null
    },
    {
      "id": "2262550",
      "postDate": "05/17/2023 02:27:46",
      "content": "<p>Cleaning data, even manually has worked for others in past competitions.  It may be just something to do with my implementations, so I wouldn't necessarily want to discourage you from trying something like this.  It's a little bit like pseudo labeling, which has also worked for others in the past.   </p>\n<p>Maybe it would be better just to build a simple bird-nobird classifier, so you can just remove the chunks of sound from training data with false positives.</p>",
      "rawMarkdown": "Cleaning data, even manually has worked for others in past competitions.  It may be just something to do with my implementations, so I wouldn't necessarily want to discourage you from trying something like this.  It's a little bit like pseudo labeling, which has also worked for others in the past.   \n\nMaybe it would be better just to build a simple bird-nobird classifier, so you can just remove the chunks of sound from training data with false positives.",
      "votes": null
    },
    {
      "id": "2263760",
      "postDate": "05/17/2023 21:58:42",
      "content": "<p>Did you do any error analysis? What kind of errors does your model do, are there any specific patterns? Did you try to visualize the inputs that your model fail to detect. And also you can compare them with the same class examples but correctly classified. Let your model guide your next step if you run out of ideas. </p>",
      "rawMarkdown": "Did you do any error analysis? What kind of errors does your model do, are there any specific patterns? Did you try to visualize the inputs that your model fail to detect. And also you can compare them with the same class examples but correctly classified. Let your model guide your next step if you run out of ideas.",
      "votes": null
    },
    {
      "id": "2263876",
      "postDate": "05/18/2023 01:56:57",
      "content": "<p>Great suggestion, thanks, but no I don't see any patterns.  The scores are pretty consistent across the classes.  Just four or five birds that don't do so well.  </p>\n<p>If I take a step back, and ask what I'm doing differently, I'm training on way less data than everybody else, getting sky-high internal scores, but not generalising well to the the test dataset.  I think this is steering me to a lack of diversity in my training samples.  I did get a measurable improvement last night by cranking up my augmentation strategy, which also points towards this I think.   Next step is going to be to throw back more training samples, until my internal scores start falling.  I'll see if that helps!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5224109%2F61eefa44cbc5eb91778ec7bdcd04e448%2Fvalid_cmap.png?generation=1684374775954997&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Great suggestion, thanks, but no I don't see any patterns.  The scores are pretty consistent across the classes.  Just four or five birds that don't do so well.  \n\nIf I take a step back, and ask what I'm doing differently, I'm training on way less data than everybody else, getting sky-high internal scores, but not generalising well to the the test dataset.  I think this is steering me to a lack of diversity in my training samples.  I did get a measurable improvement last night by cranking up my augmentation strategy, which also points towards this I think.   Next step is going to be to throw back more training samples, until my internal scores start falling.  I'll see if that helps!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5224109%2F61eefa44cbc5eb91778ec7bdcd04e448%2Fvalid_cmap.png?generation=1684374775954997&alt=media)",
      "votes": null
    },
    {
      "id": "2264404",
      "postDate": "05/18/2023 12:03:12",
      "content": "<p>You can get a lot of ideas on data augmentation from this notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/vlomme/surfin-bird-2nd-place</a>, which might help during the training. Additionally, I think ensemble models are also key to score improvement. E.g. you could train two other versions of your model with different train-validate splits, and take average (or max) of the three results in the submission_df. These are in my plans as next trials, if I can manage before the end.. </p>",
      "rawMarkdown": "You can get a lot of ideas on data augmentation from this notebook [https://www.kaggle.com/code/vlomme/surfin-bird-2nd-place](url), which might help during the training. Additionally, I think ensemble models are also key to score improvement. E.g. you could train two other versions of your model with different train-validate splits, and take average (or max) of the three results in the submission_df. These are in my plans as next trials, if I can manage before the end..",
      "votes": null
    },
    {
      "id": "2264415",
      "postDate": "05/18/2023 12:13:11",
      "content": "<p>Are these only the classes in which you train your models? If so, there are 264 classes in total (if I'm not mistaken). How do you handle the rest of the classes?</p>",
      "rawMarkdown": "Are these only the classes in which you train your models? If so, there are 264 classes in total (if I'm not mistaken). How do you handle the rest of the classes?",
      "votes": null
    },
    {
      "id": "2265116",
      "postDate": "05/19/2023 01:58:36",
      "content": "<p>It's just the way the graph is displayed.  I'm using 265 classes. One extra for no-call.  You just can't see them all on the y axis labels.</p>\n<p>I've also tried pretraining the backbone on the 360 or so classes from 21, then replacing the classifier head, but it hasn't helped yet.</p>",
      "rawMarkdown": "It's just the way the graph is displayed.  I'm using 265 classes. One extra for no-call.  You just can't see them all on the y axis labels.\n\nI've also tried pretraining the backbone on the 360 or so classes from 21, then replacing the classifier head, but it hasn't helped yet.",
      "votes": null
    },
    {
      "id": "2265122",
      "postDate": "05/19/2023 02:12:56",
      "content": "<p>Ooh, I see. Then, makes sense of course. 👍</p>",
      "rawMarkdown": "Ooh, I see. Then, makes sense of course. 👍",
      "votes": null
    },
    {
      "id": "2266403",
      "postDate": "05/20/2023 03:06:22",
      "content": "<p>Thanks Hakan,   we'll that link has gone down, but anyway I think my Augmentation is pretty good now, I just had some settings dialed down too low.  The problem for me lies elsewhere.   I think I know where now, but still not quite confirmed.</p>\n<p>I'll try and ensemble a couple of models, if the submission time allows.  Or at least a couple of different classifier heads on a shared backbone.   Good luck with your trials!</p>",
      "rawMarkdown": "Thanks Hakan,   we'll that link has gone down, but anyway I think my Augmentation is pretty good now, I just had some settings dialed down too low.  The problem for me lies elsewhere.   I think I know where now, but still not quite confirmed.\n\nI'll try and ensemble a couple of models, if the submission time allows.  Or at least a couple of different classifier heads on a shared backbone.   Good luck with your trials!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2262516,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "05/17/2023 02:03:39",
      "content": "<p>Thats very interesting, thx for sharing, might have saved me a lot of time since I thought that implementing this would really help as well.<br>\nI can't bring up any explaination to why that didn't work tho, it's the same for me, any intuition based experience I try doesn't work at all..</p>",
      "votes": null,
      "replies": [
        {
          "id": 2262550,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "05/17/2023 02:27:46",
          "content": "<p>Cleaning data, even manually has worked for others in past competitions.  It may be just something to do with my implementations, so I wouldn't necessarily want to discourage you from trying something like this.  It's a little bit like pseudo labeling, which has also worked for others in the past.   </p>\n<p>Maybe it would be better just to build a simple bird-nobird classifier, so you can just remove the chunks of sound from training data with false positives.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2263760,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "05/17/2023 21:58:42",
      "content": "<p>Did you do any error analysis? What kind of errors does your model do, are there any specific patterns? Did you try to visualize the inputs that your model fail to detect. And also you can compare them with the same class examples but correctly classified. Let your model guide your next step if you run out of ideas. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2263876,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "05/18/2023 01:56:57",
          "content": "<p>Great suggestion, thanks, but no I don't see any patterns.  The scores are pretty consistent across the classes.  Just four or five birds that don't do so well.  </p>\n<p>If I take a step back, and ask what I'm doing differently, I'm training on way less data than everybody else, getting sky-high internal scores, but not generalising well to the the test dataset.  I think this is steering me to a lack of diversity in my training samples.  I did get a measurable improvement last night by cranking up my augmentation strategy, which also points towards this I think.   Next step is going to be to throw back more training samples, until my internal scores start falling.  I'll see if that helps!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5224109%2F61eefa44cbc5eb91778ec7bdcd04e448%2Fvalid_cmap.png?generation=1684374775954997&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 2264415,
              "author_name": "snnclsr",
              "author_url": "",
              "post_date": "05/18/2023 12:13:11",
              "content": "<p>Are these only the classes in which you train your models? If so, there are 264 classes in total (if I'm not mistaken). How do you handle the rest of the classes?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2265116,
                  "author_name": "ollypowell",
                  "author_url": "",
                  "post_date": "05/19/2023 01:58:36",
                  "content": "<p>It's just the way the graph is displayed.  I'm using 265 classes. One extra for no-call.  You just can't see them all on the y axis labels.</p>\n<p>I've also tried pretraining the backbone on the 360 or so classes from 21, then replacing the classifier head, but it hasn't helped yet.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2265122,
                      "author_name": "snnclsr",
                      "author_url": "",
                      "post_date": "05/19/2023 02:12:56",
                      "content": "<p>Ooh, I see. Then, makes sense of course. 👍</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2264404,
      "author_name": "hakandogan",
      "author_url": "",
      "post_date": "05/18/2023 12:03:12",
      "content": "<p>You can get a lot of ideas on data augmentation from this notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/vlomme/surfin-bird-2nd-place</a>, which might help during the training. Additionally, I think ensemble models are also key to score improvement. E.g. you could train two other versions of your model with different train-validate splits, and take average (or max) of the three results in the submission_df. These are in my plans as next trials, if I can manage before the end.. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2266403,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "05/20/2023 03:06:22",
          "content": "<p>Thanks Hakan,   we'll that link has gone down, but anyway I think my Augmentation is pretty good now, I just had some settings dialed down too low.  The problem for me lies elsewhere.   I think I know where now, but still not quite confirmed.</p>\n<p>I'll try and ensemble a couple of models, if the submission time allows.  Or at least a couple of different classifier heads on a shared backbone.   Good luck with your trials!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2262332": "I've based a lot of my hopes in this competition on producing a really clean, and somewhat better balanced dataset.  So I built a separate detector, trained on a fraction of the training data, including a no-call class, with extra data for that.  Then I ran that through all the training data, built myself a completely new dataset based on 8 second chunks that were very likely to contain the bird of interest.  And less likely to contain the secondary bird, or a no-call, chose fewer samples from the most common classes.\n\nThe new dataset trains brilliantly,  I get private CMAP5 and LRAP scores into the .90's with lots of augmentation and no pre-training.\n\nBut.....    It scores horribly on submission, 0.76-0.78   No better than the original detector.   Either there is a bug in my inference code, that I've tried quite hard to find already.   Or this strategy is actually wrong.   Have I just picked out all the samples that were easy to train on, and ignored the hard ones that my network really needed to learn on?  Or maybe I have moved any bias learned by the first detector onto the second?  \n\nI'm running out of ideas at this point.  Just doing an experiment where I remove any selection based on the primary bird probability score, and down-prioritize samples that score high on no-call, or secondary bird probabilities.",
    "2262516": "Thats very interesting, thx for sharing, might have saved me a lot of time since I thought that implementing this would really help as well.\nI can't bring up any explaination to why that didn't work tho, it's the same for me, any intuition based experience I try doesn't work at all..",
    "2262550": "Cleaning data, even manually has worked for others in past competitions.  It may be just something to do with my implementations, so I wouldn't necessarily want to discourage you from trying something like this.  It's a little bit like pseudo labeling, which has also worked for others in the past.   \n\nMaybe it would be better just to build a simple bird-nobird classifier, so you can just remove the chunks of sound from training data with false positives.",
    "2263760": "Did you do any error analysis? What kind of errors does your model do, are there any specific patterns? Did you try to visualize the inputs that your model fail to detect. And also you can compare them with the same class examples but correctly classified. Let your model guide your next step if you run out of ideas.",
    "2263876": "Great suggestion, thanks, but no I don't see any patterns.  The scores are pretty consistent across the classes.  Just four or five birds that don't do so well.  \n\nIf I take a step back, and ask what I'm doing differently, I'm training on way less data than everybody else, getting sky-high internal scores, but not generalising well to the the test dataset.  I think this is steering me to a lack of diversity in my training samples.  I did get a measurable improvement last night by cranking up my augmentation strategy, which also points towards this I think.   Next step is going to be to throw back more training samples, until my internal scores start falling.  I'll see if that helps!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5224109%2F61eefa44cbc5eb91778ec7bdcd04e448%2Fvalid_cmap.png?generation=1684374775954997&alt=media)",
    "2264404": "You can get a lot of ideas on data augmentation from this notebook [https://www.kaggle.com/code/vlomme/surfin-bird-2nd-place](url), which might help during the training. Additionally, I think ensemble models are also key to score improvement. E.g. you could train two other versions of your model with different train-validate splits, and take average (or max) of the three results in the submission_df. These are in my plans as next trials, if I can manage before the end..",
    "2264415": "Are these only the classes in which you train your models? If so, there are 264 classes in total (if I'm not mistaken). How do you handle the rest of the classes?",
    "2265116": "It's just the way the graph is displayed.  I'm using 265 classes. One extra for no-call.  You just can't see them all on the y axis labels.\n\nI've also tried pretraining the backbone on the 360 or so classes from 21, then replacing the classifier head, but it hasn't helped yet.",
    "2265122": "Ooh, I see. Then, makes sense of course. 👍",
    "2266403": "Thanks Hakan,   we'll that link has gone down, but anyway I think my Augmentation is pretty good now, I just had some settings dialed down too low.  The problem for me lies elsewhere.   I think I know where now, but still not quite confirmed.\n\nI'll try and ensemble a couple of models, if the submission time allows.  Or at least a couple of different classifier heads on a shared backbone.   Good luck with your trials!"
  },
  "source": "meta"
}