{
  "id": 230639,
  "title": "The key is \"label quality\"",
  "url": "/competitions/birdclef-2021/discussion/230639",
  "author_name": "",
  "post_date": "2021-04-05T03:29:43.586896400Z",
  "votes": 34,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all</p>\n<p>I joined in <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">last year's competition</a> and the <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">previous one</a>. The key of these competitions was \"label quality\". I think \"label quality\" will be the key in this competition as well.</p>\n<p>In this competition, I think \"weak label\" will have the most influence on label quality. Why? Here is an example.</p>\n<ul>\n<li>XC147969.ogg(about 22 sec sound file)<br>\nprimary label(weak label) : <em>andsol1</em></li>\n</ul>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/ea1634ef-4489-4cfe-66e4-47b8407e0529.png\" alt=\"image.png\"></p>\n<p>Half of this sound file has no_sound (by my ear). And the end of sound file has human voice. If I make 5sec random clip in this sound file, I may make \"only no_sound\" clip( or \"only human voice\" clip). These sound don't have primary label sound. I don't hope this.</p>\n<p>The simplest solution of this issue is \"the length of time\" of sound file.<br>\nShort time sound file has less no_sound duration than long one.</p>\n<ul>\n<li>XC121287.ogg(about 10 sec sound file)<br>\nprimary label(weak label) : <em>andsol1</em></li>\n</ul>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/80c01e85-018b-cb91-1326-bf8791df5d49.png\" alt=\"image.png\"></p>\n<p>Using this sound file, if I get 5sec random clip, I can get primary label(andsol1) sound. I hope this.</p>\n<p>I think that the short time sound file was probably edited from the long one by the recorder. Therefore, in the short time sound file, there is less no_sound duration and maybe there is less secondary label than long one. I can get <strong>clean primary label sound</strong> by using short time sound file.</p>",
  "messages": [
    {
      "id": "1263037",
      "postDate": "04/05/2021 03:29:43",
      "content": "<p>Hi all</p>\n<p>I joined in <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">last year's competition</a> and the <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection\" target=\"_blank\">previous one</a>. The key of these competitions was \"label quality\". I think \"label quality\" will be the key in this competition as well.</p>\n<p>In this competition, I think \"weak label\" will have the most influence on label quality. Why? Here is an example.</p>\n<ul>\n<li>XC147969.ogg(about 22 sec sound file)<br>\nprimary label(weak label) : <em>andsol1</em></li>\n</ul>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/ea1634ef-4489-4cfe-66e4-47b8407e0529.png\" alt=\"image.png\"></p>\n<p>Half of this sound file has no_sound (by my ear). And the end of sound file has human voice. If I make 5sec random clip in this sound file, I may make \"only no_sound\" clip( or \"only human voice\" clip). These sound don't have primary label sound. I don't hope this.</p>\n<p>The simplest solution of this issue is \"the length of time\" of sound file.<br>\nShort time sound file has less no_sound duration than long one.</p>\n<ul>\n<li>XC121287.ogg(about 10 sec sound file)<br>\nprimary label(weak label) : <em>andsol1</em></li>\n</ul>\n<p><img src=\"https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/80c01e85-018b-cb91-1326-bf8791df5d49.png\" alt=\"image.png\"></p>\n<p>Using this sound file, if I get 5sec random clip, I can get primary label(andsol1) sound. I hope this.</p>\n<p>I think that the short time sound file was probably edited from the long one by the recorder. Therefore, in the short time sound file, there is less no_sound duration and maybe there is less secondary label than long one. I can get <strong>clean primary label sound</strong> by using short time sound file.</p>",
      "rawMarkdown": "Hi all\n\nI joined in [last year's competition](https://www.kaggle.com/c/birdsong-recognition) and the [previous one](https://www.kaggle.com/c/rfcx-species-audio-detection). The key of these competitions was \"label quality\". I think \"label quality\" will be the key in this competition as well.\n\nIn this competition, I think \"weak label\" will have the most influence on label quality. Why? Here is an example.\n\n+ XC147969.ogg(about 22 sec sound file)\nprimary label(weak label) : *andsol1*\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/ea1634ef-4489-4cfe-66e4-47b8407e0529.png)\n\nHalf of this sound file has no_sound (by my ear). And the end of sound file has human voice. If I make 5sec random clip in this sound file, I may make \"only no_sound\" clip( or \"only human voice\" clip). These sound don't have primary label sound. I don't hope this.\n\nThe simplest solution of this issue is \"the length of time\" of sound file.\nShort time sound file has less no_sound duration than long one.\n\n+ XC121287.ogg(about 10 sec sound file)\nprimary label(weak label) : *andsol1*\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/80c01e85-018b-cb91-1326-bf8791df5d49.png)\n\nUsing this sound file, if I get 5sec random clip, I can get primary label(andsol1) sound. I hope this.\n\nI think that the short time sound file was probably edited from the long one by the recorder. Therefore, in the short time sound file, there is less no_sound duration and maybe there is less secondary label than long one. I can get **clean primary label sound** by using short time sound file.",
      "votes": null
    },
    {
      "id": "1282895",
      "postDate": "04/24/2021 11:25:53",
      "content": "<p>Thanks for this interesting point. Do we know what things are considered for the label quality feature and how it was generated? For example, if there is much room in a recording without calls while still having good quality overall, will this be lower rated? Same with a human voice and such.</p>",
      "rawMarkdown": "Thanks for this interesting point. Do we know what things are considered for the label quality feature and how it was generated? For example, if there is much room in a recording without calls while still having good quality overall, will this be lower rated? Same with a human voice and such.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1282895,
      "author_name": "wspinkaggle",
      "author_url": "",
      "post_date": "04/24/2021 11:25:53",
      "content": "<p>Thanks for this interesting point. Do we know what things are considered for the label quality feature and how it was generated? For example, if there is much room in a recording without calls while still having good quality overall, will this be lower rated? Same with a human voice and such.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1263037": "Hi all\n\nI joined in [last year's competition](https://www.kaggle.com/c/birdsong-recognition) and the [previous one](https://www.kaggle.com/c/rfcx-species-audio-detection). The key of these competitions was \"label quality\". I think \"label quality\" will be the key in this competition as well.\n\nIn this competition, I think \"weak label\" will have the most influence on label quality. Why? Here is an example.\n\n+ XC147969.ogg(about 22 sec sound file)\nprimary label(weak label) : *andsol1*\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/ea1634ef-4489-4cfe-66e4-47b8407e0529.png)\n\nHalf of this sound file has no_sound (by my ear). And the end of sound file has human voice. If I make 5sec random clip in this sound file, I may make \"only no_sound\" clip( or \"only human voice\" clip). These sound don't have primary label sound. I don't hope this.\n\nThe simplest solution of this issue is \"the length of time\" of sound file.\nShort time sound file has less no_sound duration than long one.\n\n+ XC121287.ogg(about 10 sec sound file)\nprimary label(weak label) : *andsol1*\n\n![image.png](https://qiita-image-store.s3.ap-northeast-1.amazonaws.com/0/264781/80c01e85-018b-cb91-1326-bf8791df5d49.png)\n\nUsing this sound file, if I get 5sec random clip, I can get primary label(andsol1) sound. I hope this.\n\nI think that the short time sound file was probably edited from the long one by the recorder. Therefore, in the short time sound file, there is less no_sound duration and maybe there is less secondary label than long one. I can get **clean primary label sound** by using short time sound file.",
    "1282895": "Thanks for this interesting point. Do we know what things are considered for the label quality feature and how it was generated? For example, if there is much room in a recording without calls while still having good quality overall, will this be lower rated? Same with a human voice and such."
  },
  "source": "meta"
}