{
  "id": 311331,
  "title": "Task understanding",
  "url": "/competitions/birdclef-2022/discussion/311331",
  "author_name": "",
  "post_date": "2022-03-06T11:06:52.878998200Z",
  "votes": 8,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi there,</p>\n<p>I've been on kaggle for years but this is the first competition I participate in - so beware of noob questions.</p>\n<p>I looked at the data and sample submissions but I'm not a 100% sure I got the task right. So, there's &gt;150 species in the train dataset but only 21 are relevant/scored, i.e. appear in the hidden test set. The submission has two columns: <code>row_id</code> is <code>&lt;filename&gt;_&lt;scpecies&gt;_&lt;second_end&gt;</code> and <code>target</code> is boolean, whether the species in <code>&lt;species&gt;</code> is present in the sample <code>&lt;filename&gt;</code> within seconds <code>&lt;second_end&gt; - 5</code> until <code>&lt;second_end&gt;</code>, right?<br>\nSo are we supposed to make that prediction for each scored species? Because the filename doesn't have the species. That means this is a multilabel task (but why that format for the submission - why isn't there a column per species then or why don't we just print either the true or the false case?).<br>\nWith 5.5k test samples, 1 minute each and 21 species this would result in 5500x12x21 = about 1.4 million rows. Is that correct?</p>\n<p>Then, if I got this right, we have to tell if the species is present in a 5-sec chunk. But we don't have that information in the training data - it's only telling us <em>that</em> the species (and seconday ones) are present *somewhere * in the <em>whole</em> sample. So we can't just add a \"silence\" label and train it. We probably have to throw some bird/silence classifier or heuristic at the 5 seconds chunk first, right?</p>\n<p>Another question regarding the notebook: Are the nine hours for inference only or for the \"whole thing\", i.e. including training? Do I train my model(s) and they get saved and submitted anlongside the notebook?</p>\n<p>Thanks in advance and happy training!</p>",
  "messages": [
    {
      "id": "1713762",
      "postDate": "03/06/2022 11:06:52",
      "content": "<p>Hi there,</p>\n<p>I've been on kaggle for years but this is the first competition I participate in - so beware of noob questions.</p>\n<p>I looked at the data and sample submissions but I'm not a 100% sure I got the task right. So, there's &gt;150 species in the train dataset but only 21 are relevant/scored, i.e. appear in the hidden test set. The submission has two columns: <code>row_id</code> is <code>&lt;filename&gt;_&lt;scpecies&gt;_&lt;second_end&gt;</code> and <code>target</code> is boolean, whether the species in <code>&lt;species&gt;</code> is present in the sample <code>&lt;filename&gt;</code> within seconds <code>&lt;second_end&gt; - 5</code> until <code>&lt;second_end&gt;</code>, right?<br>\nSo are we supposed to make that prediction for each scored species? Because the filename doesn't have the species. That means this is a multilabel task (but why that format for the submission - why isn't there a column per species then or why don't we just print either the true or the false case?).<br>\nWith 5.5k test samples, 1 minute each and 21 species this would result in 5500x12x21 = about 1.4 million rows. Is that correct?</p>\n<p>Then, if I got this right, we have to tell if the species is present in a 5-sec chunk. But we don't have that information in the training data - it's only telling us <em>that</em> the species (and seconday ones) are present *somewhere * in the <em>whole</em> sample. So we can't just add a \"silence\" label and train it. We probably have to throw some bird/silence classifier or heuristic at the 5 seconds chunk first, right?</p>\n<p>Another question regarding the notebook: Are the nine hours for inference only or for the \"whole thing\", i.e. including training? Do I train my model(s) and they get saved and submitted anlongside the notebook?</p>\n<p>Thanks in advance and happy training!</p>",
      "rawMarkdown": "Hi there,\n\nI've been on kaggle for years but this is the first competition I participate in - so beware of noob questions.\n\nI looked at the data and sample submissions but I'm not a 100% sure I got the task right. So, there's >150 species in the train dataset but only 21 are relevant/scored, i.e. appear in the hidden test set. The submission has two columns: `row_id` is `<filename>_<scpecies>_<second_end>` and `target` is boolean, whether the species in `<species>` is present in the sample `<filename>` within seconds `<second_end> - 5` until `<second_end>`, right?\nSo are we supposed to make that prediction for each scored species? Because the filename doesn't have the species. That means this is a multilabel task (but why that format for the submission - why isn't there a column per species then or why don't we just print either the true or the false case?).\nWith 5.5k test samples, 1 minute each and 21 species this would result in 5500x12x21 = about 1.4 million rows. Is that correct?\n\nThen, if I got this right, we have to tell if the species is present in a 5-sec chunk. But we don't have that information in the training data - it's only telling us *that* the species (and seconday ones) are present *somewhere * in the *whole* sample. So we can't just add a \"silence\" label and train it. We probably have to throw some bird/silence classifier or heuristic at the 5 seconds chunk first, right?\n\nAnother question regarding the notebook: Are the nine hours for inference only or for the \"whole thing\", i.e. including training? Do I train my model(s) and they get saved and submitted anlongside the notebook?\n\nThanks in advance and happy training!",
      "votes": null
    },
    {
      "id": "1714379",
      "postDate": "03/06/2022 22:53:55",
      "content": "<p><code>Are the nine hours for inference only or for the \"whole thing\", i.e. including training? Do I train my model(s) and they get saved and submitted anlongside the notebook?</code></p>\n<p>9h is just for inference. The submission/inference notebook does not have to do training. You can train your models and folds on another notebook and save the weights. Then you add these weights as a dataset or as 'Notebook Output Files' to your inference notebook, load them and perform inference.</p>",
      "rawMarkdown": "`Are the nine hours for inference only or for the \"whole thing\", i.e. including training? Do I train my model(s) and they get saved and submitted anlongside the notebook?`\n\n9h is just for inference. The submission/inference notebook does not have to do training. You can train your models and folds on another notebook and save the weights. Then you add these weights as a dataset or as 'Notebook Output Files' to your inference notebook, load them and perform inference.",
      "votes": null
    },
    {
      "id": "1715272",
      "postDate": "03/07/2022 19:34:41",
      "content": "<p>Thanks. So that's \"kernel chaining\", eh? I can just use the outputs of one notebook (aka kernel) as input/dataset for another?</p>",
      "rawMarkdown": "Thanks. So that's \"kernel chaining\", eh? I can just use the outputs of one notebook (aka kernel) as input/dataset for another?",
      "votes": null
    },
    {
      "id": "1715408",
      "postDate": "03/08/2022 00:54:10",
      "content": "<p>Yep <a href=\"https://www.kaggle.com/m02ph3u5\" target=\"_blank\">@m02ph3u5</a> I think that's the only way to be competitive. I am spending ~6h to train just 1 fold from 1 model architecture on kaggle GPU. So if you want to use several folds and architectures (ensemble) to make your final predictions you will have to do it this way.</p>\n<p>On the other hand, inference is very fast, compared to training. On my models I can make inference in a couple of minutes (and then it runs for  ~2h before the score is available).</p>",
      "rawMarkdown": "Yep @m02ph3u5 I think that's the only way to be competitive. I am spending ~6h to train just 1 fold from 1 model architecture on kaggle GPU. So if you want to use several folds and architectures (ensemble) to make your final predictions you will have to do it this way.\n\nOn the other hand, inference is very fast, compared to training. On my models I can make inference in a couple of minutes (and then it runs for  ~2h before the score is available).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1714379,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "03/06/2022 22:53:55",
      "content": "<p><code>Are the nine hours for inference only or for the \"whole thing\", i.e. including training? Do I train my model(s) and they get saved and submitted anlongside the notebook?</code></p>\n<p>9h is just for inference. The submission/inference notebook does not have to do training. You can train your models and folds on another notebook and save the weights. Then you add these weights as a dataset or as 'Notebook Output Files' to your inference notebook, load them and perform inference.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1715272,
          "author_name": "m02ph3u5",
          "author_url": "",
          "post_date": "03/07/2022 19:34:41",
          "content": "<p>Thanks. So that's \"kernel chaining\", eh? I can just use the outputs of one notebook (aka kernel) as input/dataset for another?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1715408,
          "author_name": "hinepo",
          "author_url": "",
          "post_date": "03/08/2022 00:54:10",
          "content": "<p>Yep <a href=\"https://www.kaggle.com/m02ph3u5\" target=\"_blank\">@m02ph3u5</a> I think that's the only way to be competitive. I am spending ~6h to train just 1 fold from 1 model architecture on kaggle GPU. So if you want to use several folds and architectures (ensemble) to make your final predictions you will have to do it this way.</p>\n<p>On the other hand, inference is very fast, compared to training. On my models I can make inference in a couple of minutes (and then it runs for  ~2h before the score is available).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1713762": "Hi there,\n\nI've been on kaggle for years but this is the first competition I participate in - so beware of noob questions.\n\nI looked at the data and sample submissions but I'm not a 100% sure I got the task right. So, there's >150 species in the train dataset but only 21 are relevant/scored, i.e. appear in the hidden test set. The submission has two columns: `row_id` is `<filename>_<scpecies>_<second_end>` and `target` is boolean, whether the species in `<species>` is present in the sample `<filename>` within seconds `<second_end> - 5` until `<second_end>`, right?\nSo are we supposed to make that prediction for each scored species? Because the filename doesn't have the species. That means this is a multilabel task (but why that format for the submission - why isn't there a column per species then or why don't we just print either the true or the false case?).\nWith 5.5k test samples, 1 minute each and 21 species this would result in 5500x12x21 = about 1.4 million rows. Is that correct?\n\nThen, if I got this right, we have to tell if the species is present in a 5-sec chunk. But we don't have that information in the training data - it's only telling us *that* the species (and seconday ones) are present *somewhere * in the *whole* sample. So we can't just add a \"silence\" label and train it. We probably have to throw some bird/silence classifier or heuristic at the 5 seconds chunk first, right?\n\nAnother question regarding the notebook: Are the nine hours for inference only or for the \"whole thing\", i.e. including training? Do I train my model(s) and they get saved and submitted anlongside the notebook?\n\nThanks in advance and happy training!",
    "1714379": "`Are the nine hours for inference only or for the \"whole thing\", i.e. including training? Do I train my model(s) and they get saved and submitted anlongside the notebook?`\n\n9h is just for inference. The submission/inference notebook does not have to do training. You can train your models and folds on another notebook and save the weights. Then you add these weights as a dataset or as 'Notebook Output Files' to your inference notebook, load them and perform inference.",
    "1715272": "Thanks. So that's \"kernel chaining\", eh? I can just use the outputs of one notebook (aka kernel) as input/dataset for another?",
    "1715408": "Yep @m02ph3u5 I think that's the only way to be competitive. I am spending ~6h to train just 1 fold from 1 model architecture on kaggle GPU. So if you want to use several folds and architectures (ensemble) to make your final predictions you will have to do it this way.\n\nOn the other hand, inference is very fast, compared to training. On my models I can make inference in a couple of minutes (and then it runs for  ~2h before the score is available)."
  },
  "source": "meta"
}