{
  "id": 164741,
  "title": "Secondary Labels with no time stamp?",
  "url": "/competitions/birdsong-recognition/discussion/164741",
  "author_name": "",
  "post_date": "2020-07-07T10:43:18.929340600Z",
  "votes": 6,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello Everyone, I had gone through the data and in train.csv file there is a <code>secondary_labels</code> column. But there is no time stamp given for the same. Is this column have no use. Please correct me if I am wrong.</p>",
  "messages": [
    {
      "id": "918592",
      "postDate": "07/07/2020 10:43:18",
      "content": "<p>Hello Everyone, I had gone through the data and in train.csv file there is a <code>secondary_labels</code> column. But there is no time stamp given for the same. Is this column have no use. Please correct me if I am wrong.</p>",
      "rawMarkdown": "Hello Everyone, I had gone through the data and in train.csv file there is a `secondary_labels` column. But there is no time stamp given for the same. Is this column have no use. Please correct me if I am wrong.",
      "votes": null
    },
    {
      "id": "920972",
      "postDate": "07/09/2020 01:49:23",
      "content": "<p>What do you mean by timestamp? do you mean when the bird was heard in the audio file? Even for primary labels we don't even have those labels right? My understanding of the secondary label is that it is other birds that were also heard in the recording. If you're using the entire audio file to train, you can probably extract the <code>secondary_labels</code> and add them to your training labels.\n<code>python\nimport ast\nmeta = pd.read_csv('train.csv')\nebird_codes = {code: index for code, index in enumerate(meta.ebird_code.unique())}\nprimary_labels = {}\nfor i in range(len(meta)):\n    if meta.iloc[i].primary_label not in primary_labels:\n        primary_labels[meta.iloc[i].primary_label] = meta.iloc[i].ebird_code\nmeta.secondary_labels = meta.secondary_labels.apply(ast.literal_eval)\nsearch_secondary = lambda x: [primary_labels[v] for v in x]\nmeta.secondary_labels = meta.secondary_labels.apply(search_secondary)\n</code>\nI hope this helps!</p>",
      "rawMarkdown": "What do you mean by timestamp? do you mean when the bird was heard in the audio file? Even for primary labels we don't even have those labels right? My understanding of the secondary label is that it is other birds that were also heard in the recording. If you're using the entire audio file to train, you can probably extract the `secondary_labels` and add them to your training labels.\n```python\nimport ast\nmeta = pd.read_csv('train.csv')\nebird_codes = {code: index for code, index in enumerate(meta.ebird_code.unique())}\nprimary_labels = {}\nfor i in range(len(meta)):\n    if meta.iloc[i].primary_label not in primary_labels:\n        primary_labels[meta.iloc[i].primary_label] = meta.iloc[i].ebird_code\nmeta.secondary_labels = meta.secondary_labels.apply(ast.literal_eval)\nsearch_secondary = lambda x: [primary_labels[v] for v in x]\nmeta.secondary_labels = meta.secondary_labels.apply(search_secondary)\n```\nI hope this helps!",
      "votes": null
    },
    {
      "id": "921043",
      "postDate": "07/09/2020 03:41:39",
      "content": "<p><code>do you mean when the bird was heard in the audio file?</code>\nYes this is what i mean. Got it. We can train by using primary and secondary labels combined for entire audio file and this problem become multi-label classification.</p>",
      "rawMarkdown": "`do you mean when the bird was heard in the audio file? `\nYes this is what i mean. Got it. We can train by using primary and secondary labels combined for entire audio file and this problem become multi-label classification.",
      "votes": null
    },
    {
      "id": "921060",
      "postDate": "07/09/2020 03:54:06",
      "content": "<p>I think you can use the method that I talked about before to do multi-label classification. Correct me if I'm wrong, but I don't think time stamps exist even for primary labels? (There are no time stamps that state which sections of the audio are ambient while others are the bird call).</p>",
      "rawMarkdown": "I think you can use the method that I talked about before to do multi-label classification. Correct me if I'm wrong, but I don't think time stamps exist even for primary labels? (There are no time stamps that state which sections of the audio are ambient while others are the bird call).",
      "votes": null
    },
    {
      "id": "921090",
      "postDate": "07/09/2020 04:26:52",
      "content": "<p><code>I think you can use the method that I talked about before to do multi-label classification.</code>\nYes sure.\n<code>Correct me if I'm wrong, but I don't think time stamps exist even for primary labels?</code>\nYes you are right</p>",
      "rawMarkdown": "`I think you can use the method that I talked about before to do multi-label classification.`\nYes sure.\n`Correct me if I'm wrong, but I don't think time stamps exist even for primary labels? `\nYes you are right",
      "votes": null
    },
    {
      "id": "1000890",
      "postDate": "09/06/2020 21:52:20",
      "content": "<p><a href=\"https://www.kaggle.com/rishabhdhiman\" target=\"_blank\">@rishabhdhiman</a> <a href=\"https://www.kaggle.com/alansun17904\" target=\"_blank\">@alansun17904</a> Did you achieve any better results with using <code>secondary_labels</code>? Using them did not increase my score on LB and I heard some people saying that it just causes bad results and does not make training better.</p>",
      "rawMarkdown": "rishabhdhiman @alansun17904 Did you achieve any better results with using `secondary_labels`? Using them did not increase my score on LB and I heard some people saying that it just causes bad results and does not make training better.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 920972,
      "author_name": "alansun17904",
      "author_url": "",
      "post_date": "07/09/2020 01:49:23",
      "content": "<p>What do you mean by timestamp? do you mean when the bird was heard in the audio file? Even for primary labels we don't even have those labels right? My understanding of the secondary label is that it is other birds that were also heard in the recording. If you're using the entire audio file to train, you can probably extract the <code>secondary_labels</code> and add them to your training labels.\n<code>python\nimport ast\nmeta = pd.read_csv('train.csv')\nebird_codes = {code: index for code, index in enumerate(meta.ebird_code.unique())}\nprimary_labels = {}\nfor i in range(len(meta)):\n    if meta.iloc[i].primary_label not in primary_labels:\n        primary_labels[meta.iloc[i].primary_label] = meta.iloc[i].ebird_code\nmeta.secondary_labels = meta.secondary_labels.apply(ast.literal_eval)\nsearch_secondary = lambda x: [primary_labels[v] for v in x]\nmeta.secondary_labels = meta.secondary_labels.apply(search_secondary)\n</code>\nI hope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 921043,
          "author_name": "rishabhdhiman",
          "author_url": "",
          "post_date": "07/09/2020 03:41:39",
          "content": "<p><code>do you mean when the bird was heard in the audio file?</code>\nYes this is what i mean. Got it. We can train by using primary and secondary labels combined for entire audio file and this problem become multi-label classification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 921060,
          "author_name": "alansun17904",
          "author_url": "",
          "post_date": "07/09/2020 03:54:06",
          "content": "<p>I think you can use the method that I talked about before to do multi-label classification. Correct me if I'm wrong, but I don't think time stamps exist even for primary labels? (There are no time stamps that state which sections of the audio are ambient while others are the bird call).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 921090,
          "author_name": "rishabhdhiman",
          "author_url": "",
          "post_date": "07/09/2020 04:26:52",
          "content": "<p><code>I think you can use the method that I talked about before to do multi-label classification.</code>\nYes sure.\n<code>Correct me if I'm wrong, but I don't think time stamps exist even for primary labels?</code>\nYes you are right</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1000890,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "09/06/2020 21:52:20",
          "content": "<p><a href=\"https://www.kaggle.com/rishabhdhiman\" target=\"_blank\">@rishabhdhiman</a> <a href=\"https://www.kaggle.com/alansun17904\" target=\"_blank\">@alansun17904</a> Did you achieve any better results with using <code>secondary_labels</code>? Using them did not increase my score on LB and I heard some people saying that it just causes bad results and does not make training better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "918592": "Hello Everyone, I had gone through the data and in train.csv file there is a `secondary_labels` column. But there is no time stamp given for the same. Is this column have no use. Please correct me if I am wrong.",
    "920972": "What do you mean by timestamp? do you mean when the bird was heard in the audio file? Even for primary labels we don't even have those labels right? My understanding of the secondary label is that it is other birds that were also heard in the recording. If you're using the entire audio file to train, you can probably extract the `secondary_labels` and add them to your training labels.\n```python\nimport ast\nmeta = pd.read_csv('train.csv')\nebird_codes = {code: index for code, index in enumerate(meta.ebird_code.unique())}\nprimary_labels = {}\nfor i in range(len(meta)):\n    if meta.iloc[i].primary_label not in primary_labels:\n        primary_labels[meta.iloc[i].primary_label] = meta.iloc[i].ebird_code\nmeta.secondary_labels = meta.secondary_labels.apply(ast.literal_eval)\nsearch_secondary = lambda x: [primary_labels[v] for v in x]\nmeta.secondary_labels = meta.secondary_labels.apply(search_secondary)\n```\nI hope this helps!",
    "921043": "`do you mean when the bird was heard in the audio file? `\nYes this is what i mean. Got it. We can train by using primary and secondary labels combined for entire audio file and this problem become multi-label classification.",
    "921060": "I think you can use the method that I talked about before to do multi-label classification. Correct me if I'm wrong, but I don't think time stamps exist even for primary labels? (There are no time stamps that state which sections of the audio are ambient while others are the bird call).",
    "921090": "`I think you can use the method that I talked about before to do multi-label classification.`\nYes sure.\n`Correct me if I'm wrong, but I don't think time stamps exist even for primary labels? `\nYes you are right",
    "1000890": "rishabhdhiman @alansun17904 Did you achieve any better results with using `secondary_labels`? Using them did not increase my score on LB and I heard some people saying that it just causes bad results and does not make training better."
  },
  "source": "meta"
}