{
  "id": 326972,
  "title": "How to determine what \"worked\" and what didn't?",
  "url": "/competitions/birdclef-2022/discussion/326972",
  "author_name": "",
  "post_date": "2022-05-25T05:21:34.411545900Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, so our ensemble scores 0.78 public set / 0.70 private set - meaning that we are among those who failed the hardest; so we would like to review what we did and maybe find out what worked, what didn't, and why our ensemble failed so miserably, so that we could perhaps take this opportunity to actually learn something, gain insights, maybe write writeups, etc. In fact, we would be more than happy to analyze our experimentations and share the analysis with the community. However, it is not immediately clear how the analysis can even be done.</p>\n<p>I noticed that many of you who already wrote writeups seemed quite confident about what worked and what didn't. Could you teach me how to figure this out? Or, more like…could you define \"worked\"? How much improvement is \"worked\"?</p>\n<p>OK, maybe the question needs further clarification. I'll explain why I think \"just look at the leaderboard, Jayeon!\"  is just not enough for an answer. Well, can the private leaderboard score be trusted? Sure, it has more samples than the public leaderboard, but just like how a model successful in the public set can fail miserably in the private set, a model successful in the private set can also fail hard in other real-life test-sets, if there were any. In short, the private leaderboard tells us less about how a specific tweak contributes to a robustness of a model.</p>\n<p>Furthermore, since the decimal points given as the score are so limited, it is very hard to see what augmentations led to a increase in performance, and which ones do not. Combine this with the robustness issue I mentioned above, and everything becomes even more opaque…statistically.</p>\n<p><strong>Q. How do you determine what \"worked\"? How do you determine what didn't?</strong></p>",
  "messages": [
    {
      "id": "1800627",
      "postDate": "05/25/2022 05:21:34",
      "content": "<p>Hi, so our ensemble scores 0.78 public set / 0.70 private set - meaning that we are among those who failed the hardest; so we would like to review what we did and maybe find out what worked, what didn't, and why our ensemble failed so miserably, so that we could perhaps take this opportunity to actually learn something, gain insights, maybe write writeups, etc. In fact, we would be more than happy to analyze our experimentations and share the analysis with the community. However, it is not immediately clear how the analysis can even be done.</p>\n<p>I noticed that many of you who already wrote writeups seemed quite confident about what worked and what didn't. Could you teach me how to figure this out? Or, more like…could you define \"worked\"? How much improvement is \"worked\"?</p>\n<p>OK, maybe the question needs further clarification. I'll explain why I think \"just look at the leaderboard, Jayeon!\"  is just not enough for an answer. Well, can the private leaderboard score be trusted? Sure, it has more samples than the public leaderboard, but just like how a model successful in the public set can fail miserably in the private set, a model successful in the private set can also fail hard in other real-life test-sets, if there were any. In short, the private leaderboard tells us less about how a specific tweak contributes to a robustness of a model.</p>\n<p>Furthermore, since the decimal points given as the score are so limited, it is very hard to see what augmentations led to a increase in performance, and which ones do not. Combine this with the robustness issue I mentioned above, and everything becomes even more opaque…statistically.</p>\n<p><strong>Q. How do you determine what \"worked\"? How do you determine what didn't?</strong></p>",
      "rawMarkdown": "Hi, so our ensemble scores 0.78 public set / 0.70 private set - meaning that we are among those who failed the hardest; so we would like to review what we did and maybe find out what worked, what didn't, and why our ensemble failed so miserably, so that we could perhaps take this opportunity to actually learn something, gain insights, maybe write writeups, etc. In fact, we would be more than happy to analyze our experimentations and share the analysis with the community. However, it is not immediately clear how the analysis can even be done.\n\nI noticed that many of you who already wrote writeups seemed quite confident about what worked and what didn't. Could you teach me how to figure this out? Or, more like...could you define \"worked\"? How much improvement is \"worked\"?\n\nOK, maybe the question needs further clarification. I'll explain why I think \"just look at the leaderboard, Jayeon!\"  is just not enough for an answer. Well, can the private leaderboard score be trusted? Sure, it has more samples than the public leaderboard, but just like how a model successful in the public set can fail miserably in the private set, a model successful in the private set can also fail hard in other real-life test-sets, if there were any. In short, the private leaderboard tells us less about how a specific tweak contributes to a robustness of a model.\n\nFurthermore, since the decimal points given as the score are so limited, it is very hard to see what augmentations led to a increase in performance, and which ones do not. Combine this with the robustness issue I mentioned above, and everything becomes even more opaque...statistically.\n\n**Q. How do you determine what \"worked\"? How do you determine what didn't?**",
      "votes": null
    },
    {
      "id": "1800733",
      "postDate": "05/25/2022 07:12:40",
      "content": "<p>I think,<br>\nfor example, When cv or lb goes up by applying some way like augmentation. then it is called \"what worked\"<br>\notherwise what didn't work</p>",
      "rawMarkdown": "I think,\nfor example, When cv or lb goes up by applying some way like augmentation. then it is called \"what worked\"\notherwise what didn't work",
      "votes": null
    },
    {
      "id": "1800781",
      "postDate": "05/25/2022 08:08:57",
      "content": "<p>goes up…by how much? Do you have a rule-of-thumb?</p>",
      "rawMarkdown": "goes up...by how much? Do you have a rule-of-thumb?",
      "votes": null
    },
    {
      "id": "1800790",
      "postDate": "05/25/2022 08:14:55",
      "content": "<p>In this competition, difficult point is…</p>\n<ul>\n<li>Domain shift</li>\n<li>Weak label</li>\n</ul>\n<p>These point is contained BirdClef2021. I recommend using BirdClef2021.<br>\nIn BirdClef2021, you can use train_soundscape. <br>\nTrain_soundscape is recorded with same style of test data.<br>\nAnd you can listen this and you evaluate model with this data.</p>\n<p>In fact, in this competition, I experiment with BirdClef2021. And I evaluated.<br>\nOnly BirdClef2021 improvement methods were used in this competition.</p>",
      "rawMarkdown": "In this competition, difficult point is...\n- Domain shift\n- Weak label\n\nThese point is contained BirdClef2021. I recommend using BirdClef2021.\nIn BirdClef2021, you can use train_soundscape. \nTrain_soundscape is recorded with same style of test data.\nAnd you can listen this and you evaluate model with this data.\n\nIn fact, in this competition, I experiment with BirdClef2021. And I evaluated.\nOnly BirdClef2021 improvement methods were used in this competition.",
      "votes": null
    },
    {
      "id": "1801253",
      "postDate": "05/25/2022 14:57:28",
      "content": "<p>It depends on the competition. If one method was applied in this competition, even 1% of the scores would have contributed to the improvement of the score, the ranking would have risen significantly. then it becomes \"what worked\".</p>",
      "rawMarkdown": "It depends on the competition. If one method was applied in this competition, even 1% of the scores would have contributed to the improvement of the score, the ranking would have risen significantly. then it becomes \"what worked\".",
      "votes": null
    },
    {
      "id": "1801319",
      "postDate": "05/25/2022 16:02:06",
      "content": "<p>Ahh, makes sense. Thank you!</p>",
      "rawMarkdown": "Ahh, makes sense. Thank you!",
      "votes": null
    },
    {
      "id": "1801335",
      "postDate": "05/25/2022 16:11:26",
      "content": "<p>Congratulations on the gold! I think  evaluating on BirdCLEF2021 is actually a very resourceful move, and I am in awe.</p>\n<p>This is a bit funny for us - our team actually applied almost all of the augmentation tricks that were mentioned in BirdCLEF 2021/2020 high-place writeups, including adding freefield1010 noise, adding pink noise and/or white noise, using mixup, using cutmix, training on longer segments instead of shorter ones to avoid false positives in label, spectrogram flipping, cutout, etc etc etc, we trained multiple models with/without these and funnily enough it turns out that none of them <em>actually</em> \"worked\", (as defined by <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> !) Seems like our lack of validation mechanism got us bad ;D </p>\n<p>Again, congratulations! By the way do you have plans to write writeups/papers? I'm curious as to what kinds of tricks worked with a SED, because we used SEDs, too…🤔</p>",
      "rawMarkdown": "Congratulations on the gold! I think  evaluating on BirdCLEF2021 is actually a very resourceful move, and I am in awe.\n\nThis is a bit funny for us - our team actually applied almost all of the augmentation tricks that were mentioned in BirdCLEF 2021/2020 high-place writeups, including adding freefield1010 noise, adding pink noise and/or white noise, using mixup, using cutmix, training on longer segments instead of shorter ones to avoid false positives in label, spectrogram flipping, cutout, etc etc etc, we trained multiple models with/without these and funnily enough it turns out that none of them *actually* \"worked\", (as defined by @deepkim !) Seems like our lack of validation mechanism got us bad ;D \n\nAgain, congratulations! By the way do you have plans to write writeups/papers? I'm curious as to what kinds of tricks worked with a SED, because we used SEDs, too...🤔",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1800733,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "05/25/2022 07:12:40",
      "content": "<p>I think,<br>\nfor example, When cv or lb goes up by applying some way like augmentation. then it is called \"what worked\"<br>\notherwise what didn't work</p>",
      "votes": null,
      "replies": [
        {
          "id": 1800781,
          "author_name": "jayeonyi",
          "author_url": "",
          "post_date": "05/25/2022 08:08:57",
          "content": "<p>goes up…by how much? Do you have a rule-of-thumb?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801253,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "05/25/2022 14:57:28",
          "content": "<p>It depends on the competition. If one method was applied in this competition, even 1% of the scores would have contributed to the improvement of the score, the ranking would have risen significantly. then it becomes \"what worked\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1801319,
          "author_name": "jayeonyi",
          "author_url": "",
          "post_date": "05/25/2022 16:02:06",
          "content": "<p>Ahh, makes sense. Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1800790,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "05/25/2022 08:14:55",
      "content": "<p>In this competition, difficult point is…</p>\n<ul>\n<li>Domain shift</li>\n<li>Weak label</li>\n</ul>\n<p>These point is contained BirdClef2021. I recommend using BirdClef2021.<br>\nIn BirdClef2021, you can use train_soundscape. <br>\nTrain_soundscape is recorded with same style of test data.<br>\nAnd you can listen this and you evaluate model with this data.</p>\n<p>In fact, in this competition, I experiment with BirdClef2021. And I evaluated.<br>\nOnly BirdClef2021 improvement methods were used in this competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1801335,
          "author_name": "jayeonyi",
          "author_url": "",
          "post_date": "05/25/2022 16:11:26",
          "content": "<p>Congratulations on the gold! I think  evaluating on BirdCLEF2021 is actually a very resourceful move, and I am in awe.</p>\n<p>This is a bit funny for us - our team actually applied almost all of the augmentation tricks that were mentioned in BirdCLEF 2021/2020 high-place writeups, including adding freefield1010 noise, adding pink noise and/or white noise, using mixup, using cutmix, training on longer segments instead of shorter ones to avoid false positives in label, spectrogram flipping, cutout, etc etc etc, we trained multiple models with/without these and funnily enough it turns out that none of them <em>actually</em> \"worked\", (as defined by <a href=\"https://www.kaggle.com/deepkim\" target=\"_blank\">@deepkim</a> !) Seems like our lack of validation mechanism got us bad ;D </p>\n<p>Again, congratulations! By the way do you have plans to write writeups/papers? I'm curious as to what kinds of tricks worked with a SED, because we used SEDs, too…🤔</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1800627": "Hi, so our ensemble scores 0.78 public set / 0.70 private set - meaning that we are among those who failed the hardest; so we would like to review what we did and maybe find out what worked, what didn't, and why our ensemble failed so miserably, so that we could perhaps take this opportunity to actually learn something, gain insights, maybe write writeups, etc. In fact, we would be more than happy to analyze our experimentations and share the analysis with the community. However, it is not immediately clear how the analysis can even be done.\n\nI noticed that many of you who already wrote writeups seemed quite confident about what worked and what didn't. Could you teach me how to figure this out? Or, more like...could you define \"worked\"? How much improvement is \"worked\"?\n\nOK, maybe the question needs further clarification. I'll explain why I think \"just look at the leaderboard, Jayeon!\"  is just not enough for an answer. Well, can the private leaderboard score be trusted? Sure, it has more samples than the public leaderboard, but just like how a model successful in the public set can fail miserably in the private set, a model successful in the private set can also fail hard in other real-life test-sets, if there were any. In short, the private leaderboard tells us less about how a specific tweak contributes to a robustness of a model.\n\nFurthermore, since the decimal points given as the score are so limited, it is very hard to see what augmentations led to a increase in performance, and which ones do not. Combine this with the robustness issue I mentioned above, and everything becomes even more opaque...statistically.\n\n**Q. How do you determine what \"worked\"? How do you determine what didn't?**",
    "1800733": "I think,\nfor example, When cv or lb goes up by applying some way like augmentation. then it is called \"what worked\"\notherwise what didn't work",
    "1800781": "goes up...by how much? Do you have a rule-of-thumb?",
    "1800790": "In this competition, difficult point is...\n- Domain shift\n- Weak label\n\nThese point is contained BirdClef2021. I recommend using BirdClef2021.\nIn BirdClef2021, you can use train_soundscape. \nTrain_soundscape is recorded with same style of test data.\nAnd you can listen this and you evaluate model with this data.\n\nIn fact, in this competition, I experiment with BirdClef2021. And I evaluated.\nOnly BirdClef2021 improvement methods were used in this competition.",
    "1801253": "It depends on the competition. If one method was applied in this competition, even 1% of the scores would have contributed to the improvement of the score, the ranking would have risen significantly. then it becomes \"what worked\".",
    "1801319": "Ahh, makes sense. Thank you!",
    "1801335": "Congratulations on the gold! I think  evaluating on BirdCLEF2021 is actually a very resourceful move, and I am in awe.\n\nThis is a bit funny for us - our team actually applied almost all of the augmentation tricks that were mentioned in BirdCLEF 2021/2020 high-place writeups, including adding freefield1010 noise, adding pink noise and/or white noise, using mixup, using cutmix, training on longer segments instead of shorter ones to avoid false positives in label, spectrogram flipping, cutout, etc etc etc, we trained multiple models with/without these and funnily enough it turns out that none of them *actually* \"worked\", (as defined by @deepkim !) Seems like our lack of validation mechanism got us bad ;D \n\nAgain, congratulations! By the way do you have plans to write writeups/papers? I'm curious as to what kinds of tricks worked with a SED, because we used SEDs, too...🤔"
  },
  "source": "meta"
}