{
  "id": 175864,
  "title": "Question to the organizers about train data markup",
  "url": "/competitions/birdsong-recognition/discussion/175864",
  "author_name": "",
  "post_date": "2020-08-19T16:59:56.334164200Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Dear competition organizers, <strong>is additional marking of training data allowed?</strong></p>\n<p>As noted in other topics, the training data mixes several different classes in some files. (or the required class may be missing, it seems to me).</p>\n<p>Or is the purpose of the competition to create a model on such inaccurate markings?</p>\n<p>I am not yet competent enough to understand this. That's why this question arose.</p>\n<p>For example, such markup might look like:<br>\n<strong>label – file – start - duration</strong><br>\naldfly - XC16967.mp3 – 1.26 – 0.39<br>\naldfly - XC16967.mp3 – 6.70 – 0.43<br>\n…</p>\n<p>[Of course, we are only talking about training data in the \"train_audio\" folder]</p>",
  "messages": [
    {
      "id": "977715",
      "postDate": "08/19/2020 16:59:56",
      "content": "<p>Dear competition organizers, <strong>is additional marking of training data allowed?</strong></p>\n<p>As noted in other topics, the training data mixes several different classes in some files. (or the required class may be missing, it seems to me).</p>\n<p>Or is the purpose of the competition to create a model on such inaccurate markings?</p>\n<p>I am not yet competent enough to understand this. That's why this question arose.</p>\n<p>For example, such markup might look like:<br>\n<strong>label – file – start - duration</strong><br>\naldfly - XC16967.mp3 – 1.26 – 0.39<br>\naldfly - XC16967.mp3 – 6.70 – 0.43<br>\n…</p>\n<p>[Of course, we are only talking about training data in the \"train_audio\" folder]</p>",
      "rawMarkdown": "Dear competition organizers, **is additional marking of training data allowed?**\n\nAs noted in other topics, the training data mixes several different classes in some files. (or the required class may be missing, it seems to me).\n\nOr is the purpose of the competition to create a model on such inaccurate markings?\n\nI am not yet competent enough to understand this. That's why this question arose.\n\nFor example, such markup might look like:\n**label – file – start - duration**\naldfly - XC16967.mp3 – 1.26 – 0.39\naldfly - XC16967.mp3 – 6.70 – 0.43\n…\n\n[Of course, we are only talking about training data in the \"train_audio\" folder]",
      "votes": null
    },
    {
      "id": "978098",
      "postDate": "08/20/2020 00:07:08",
      "content": "<p>you can use the train data in anyway you like. you can use your own annotation (e.g. hand-mark intervals)</p>\n<p>\"Or is the purpose of the competition to create a model on such inaccurate markings?\"</p>\n<ul>\n<li>its is difficult to obtain strong labels (interval labels) for audio problems. if you start to do hand marking, you will know why.</li>\n<li>most audio problems are weak-supervised problem (only clip labels) are given.</li>\n</ul>\n<p>if problems can be solved by human annotation, then we would like to use that. But not all problems can be solved by human annotation … sometimes we don't know how to annotate (especially if it is non-visual data and something we don't see)</p>\n<p>\" (or the required class may be missing, it seems to me).\"<br>\ni check most of the data. the class is present.</p>",
      "rawMarkdown": "you can use the train data in anyway you like. you can use your own annotation (e.g. hand-mark intervals)\n\n\"Or is the purpose of the competition to create a model on such inaccurate markings?\"\n- its is difficult to obtain strong labels (interval labels) for audio problems. if you start to do hand marking, you will know why.\n- most audio problems are weak-supervised problem (only clip labels) are given.\n\nif problems can be solved by human annotation, then we would like to use that. But not all problems can be solved by human annotation ... sometimes we don't know how to annotate (especially if it is non-visual data and something we don't see)\n\n\" (or the required class may be missing, it seems to me).\"\ni check most of the data. the class is present.",
      "votes": null
    },
    {
      "id": "978361",
      "postDate": "08/20/2020 06:10:56",
      "content": "<p>Thank you, I found your old <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/169538\" target=\"_blank\">topic </a> with a discussion of this task</p>",
      "rawMarkdown": "Thank you, I found your old [topic ](https://www.kaggle.com/c/birdsong-recognition/discussion/169538) with a discussion of this task",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 978098,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/20/2020 00:07:08",
      "content": "<p>you can use the train data in anyway you like. you can use your own annotation (e.g. hand-mark intervals)</p>\n<p>\"Or is the purpose of the competition to create a model on such inaccurate markings?\"</p>\n<ul>\n<li>its is difficult to obtain strong labels (interval labels) for audio problems. if you start to do hand marking, you will know why.</li>\n<li>most audio problems are weak-supervised problem (only clip labels) are given.</li>\n</ul>\n<p>if problems can be solved by human annotation, then we would like to use that. But not all problems can be solved by human annotation … sometimes we don't know how to annotate (especially if it is non-visual data and something we don't see)</p>\n<p>\" (or the required class may be missing, it seems to me).\"<br>\ni check most of the data. the class is present.</p>",
      "votes": null,
      "replies": [
        {
          "id": 978361,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "08/20/2020 06:10:56",
          "content": "<p>Thank you, I found your old <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/169538\" target=\"_blank\">topic </a> with a discussion of this task</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "977715": "Dear competition organizers, **is additional marking of training data allowed?**\n\nAs noted in other topics, the training data mixes several different classes in some files. (or the required class may be missing, it seems to me).\n\nOr is the purpose of the competition to create a model on such inaccurate markings?\n\nI am not yet competent enough to understand this. That's why this question arose.\n\nFor example, such markup might look like:\n**label – file – start - duration**\naldfly - XC16967.mp3 – 1.26 – 0.39\naldfly - XC16967.mp3 – 6.70 – 0.43\n…\n\n[Of course, we are only talking about training data in the \"train_audio\" folder]",
    "978098": "you can use the train data in anyway you like. you can use your own annotation (e.g. hand-mark intervals)\n\n\"Or is the purpose of the competition to create a model on such inaccurate markings?\"\n- its is difficult to obtain strong labels (interval labels) for audio problems. if you start to do hand marking, you will know why.\n- most audio problems are weak-supervised problem (only clip labels) are given.\n\nif problems can be solved by human annotation, then we would like to use that. But not all problems can be solved by human annotation ... sometimes we don't know how to annotate (especially if it is non-visual data and something we don't see)\n\n\" (or the required class may be missing, it seems to me).\"\ni check most of the data. the class is present.",
    "978361": "Thank you, I found your old [topic ](https://www.kaggle.com/c/birdsong-recognition/discussion/169538) with a discussion of this task"
  },
  "source": "meta"
}