{
  "id": 321744,
  "title": "Question about comparisons with previous competitions",
  "url": "/competitions/birdclef-2022/discussion/321744",
  "author_name": "",
  "post_date": "2022-04-28T12:41:35.328419100Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi folks,</p>\n<p>This is my first audio competition. And I am experimenting everyday, learning from past solutions.</p>\n<p>I am sure that some of the people who participate in this competition have participated in similar competitions(BirdCLEF 2021 - Birdcall Identification, Cornell Birdcall Identification etc.) in the past.</p>\n<p>So I have one question. <br>\nHow is the noisiness of the data and labels and the magnitude of the domain shift different compared to past competitions?</p>\n<p>I am waiting to hear your experiences.<br>\nThank you!</p>",
  "messages": [
    {
      "id": "1770624",
      "postDate": "04/28/2022 12:41:35",
      "content": "<p>Hi folks,</p>\n<p>This is my first audio competition. And I am experimenting everyday, learning from past solutions.</p>\n<p>I am sure that some of the people who participate in this competition have participated in similar competitions(BirdCLEF 2021 - Birdcall Identification, Cornell Birdcall Identification etc.) in the past.</p>\n<p>So I have one question. <br>\nHow is the noisiness of the data and labels and the magnitude of the domain shift different compared to past competitions?</p>\n<p>I am waiting to hear your experiences.<br>\nThank you!</p>",
      "rawMarkdown": "Hi folks,\n\nThis is my first audio competition. And I am experimenting everyday, learning from past solutions.\n\nI am sure that some of the people who participate in this competition have participated in similar competitions(BirdCLEF 2021 - Birdcall Identification, Cornell Birdcall Identification etc.) in the past.\n\nSo I have one question. \nHow is the noisiness of the data and labels and the magnitude of the domain shift different compared to past competitions?\n\nI am waiting to hear your experiences.\nThank you!",
      "votes": null
    },
    {
      "id": "1771339",
      "postDate": "04/29/2022 06:14:59",
      "content": "<p>I participated in BirdCLEF2021 last year.<br>\nThe degree of noise in the training data is probably the same as last year since data source (xeno-canto) does not change. It seems that effective augmentation methods are similar. <br>\nTarget labels are more imbalanced and metric can be more sensitive to imbalance labels. Secondary labels were not very effective for my models in last year's competition, but were quite effective this year. <br>\nI have no idea as for domain shift because only one test_soundscape is opened. In that respect, Cornell Birdcall Identification is closer than BirdCLEF2021. In addition to this point, metric is unclear making it very difficult to build a reliable validation.</p>",
      "rawMarkdown": "I participated in BirdCLEF2021 last year.\nThe degree of noise in the training data is probably the same as last year since data source (xeno-canto) does not change. It seems that effective augmentation methods are similar. \nTarget labels are more imbalanced and metric can be more sensitive to imbalance labels. Secondary labels were not very effective for my models in last year's competition, but were quite effective this year. \nI have no idea as for domain shift because only one test_soundscape is opened. In that respect, Cornell Birdcall Identification is closer than BirdCLEF2021. In addition to this point, metric is unclear making it very difficult to build a reliable validation.",
      "votes": null
    },
    {
      "id": "1771515",
      "postDate": "04/29/2022 09:55:42",
      "content": "<p>Thanks for helpful informations!</p>\n<ul>\n<li>Label imbalance</li>\n<li>Weak label</li>\n<li>Small sample(scored_birds)</li>\n<li>Domain shift(different thresholds between train and test)</li>\n</ul>\n<p>These things bother me.</p>\n<p>I am considering different validation strategies, but to be honest, I have not yet found a clear way to do so.</p>\n<p>It's still going to be a long journey😂</p>",
      "rawMarkdown": "Thanks for helpful informations!\n\n- Label imbalance\n- Weak label\n- Small sample(scored_birds)\n- Domain shift(different thresholds between train and test)\n\nThese things bother me.\n\nI am considering different validation strategies, but to be honest, I have not yet found a clear way to do so.\n\nIt's still going to be a long journey😂",
      "votes": null
    },
    {
      "id": "1773084",
      "postDate": "04/30/2022 21:41:17",
      "content": "<p>The data is definitely noisier than what has been used in past competitions. The magnitude of the domain shift is also much larger - making it a more difficult challenge.</p>",
      "rawMarkdown": "The data is definitely noisier than what has been used in past competitions. The magnitude of the domain shift is also much larger - making it a more difficult challenge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1771339,
      "author_name": "naka2ka",
      "author_url": "",
      "post_date": "04/29/2022 06:14:59",
      "content": "<p>I participated in BirdCLEF2021 last year.<br>\nThe degree of noise in the training data is probably the same as last year since data source (xeno-canto) does not change. It seems that effective augmentation methods are similar. <br>\nTarget labels are more imbalanced and metric can be more sensitive to imbalance labels. Secondary labels were not very effective for my models in last year's competition, but were quite effective this year. <br>\nI have no idea as for domain shift because only one test_soundscape is opened. In that respect, Cornell Birdcall Identification is closer than BirdCLEF2021. In addition to this point, metric is unclear making it very difficult to build a reliable validation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1771515,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "04/29/2022 09:55:42",
          "content": "<p>Thanks for helpful informations!</p>\n<ul>\n<li>Label imbalance</li>\n<li>Weak label</li>\n<li>Small sample(scored_birds)</li>\n<li>Domain shift(different thresholds between train and test)</li>\n</ul>\n<p>These things bother me.</p>\n<p>I am considering different validation strategies, but to be honest, I have not yet found a clear way to do so.</p>\n<p>It's still going to be a long journey😂</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1773084,
      "author_name": "",
      "author_url": "",
      "post_date": "04/30/2022 21:41:17",
      "content": "<p>The data is definitely noisier than what has been used in past competitions. The magnitude of the domain shift is also much larger - making it a more difficult challenge.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1770624": "Hi folks,\n\nThis is my first audio competition. And I am experimenting everyday, learning from past solutions.\n\nI am sure that some of the people who participate in this competition have participated in similar competitions(BirdCLEF 2021 - Birdcall Identification, Cornell Birdcall Identification etc.) in the past.\n\nSo I have one question. \nHow is the noisiness of the data and labels and the magnitude of the domain shift different compared to past competitions?\n\nI am waiting to hear your experiences.\nThank you!",
    "1771339": "I participated in BirdCLEF2021 last year.\nThe degree of noise in the training data is probably the same as last year since data source (xeno-canto) does not change. It seems that effective augmentation methods are similar. \nTarget labels are more imbalanced and metric can be more sensitive to imbalance labels. Secondary labels were not very effective for my models in last year's competition, but were quite effective this year. \nI have no idea as for domain shift because only one test_soundscape is opened. In that respect, Cornell Birdcall Identification is closer than BirdCLEF2021. In addition to this point, metric is unclear making it very difficult to build a reliable validation.",
    "1771515": "Thanks for helpful informations!\n\n- Label imbalance\n- Weak label\n- Small sample(scored_birds)\n- Domain shift(different thresholds between train and test)\n\nThese things bother me.\n\nI am considering different validation strategies, but to be honest, I have not yet found a clear way to do so.\n\nIt's still going to be a long journey😂",
    "1773084": "The data is definitely noisier than what has been used in past competitions. The magnitude of the domain shift is also much larger - making it a more difficult challenge."
  },
  "source": "meta"
}