{
  "id": 312554,
  "title": "Competition Dos and Don'ts",
  "url": "/competitions/kaggle-pog-series-s01e02/discussion/312554",
  "author_name": "",
  "post_date": "2022-03-12T17:51:57.893604300Z",
  "votes": 16,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I'm so excited about this competition, it should be a lot of fun. I know everyone who takes part will learn something new. All of the information below can be found in the competition pages, but I wanted to highlight these so there is no confusion. Please keep these in mind when making your submissions:</p>\n<ul>\n<li><strong>Do</strong>: Research about machine learning techniques for audio data. Read papers, learn about state-of-the-art model archetectures, and try to implement them. If you like, you can also share what you've learned on the forums or in notebooks (up until the week before the deadline).</li>\n<li><strong>Don't</strong>: Use external data sources or pretrained models. Don't attempt to search for additional information about the songs (public or private set), and obviously don't hand label the test. This will disqualify you from the competition and prizes.</li>\n<li><strong>Do</strong>: Read the competition description and <a href=\"https://www.kaggle.com/c/kaggle-pog-series-s01e02/overview/prizes\" target=\"_blank\">prize elidgeability requirements</a>. Follow all the steps to ensure you are elidgeable to win the GPU and DLI vouchers.</li>\n<li><strong>Don't</strong>: Discuss the competition with people outside of your team, as it's considered private sharing and will disqualify you from prize elidgeability. We are here to learn together but not to cheat!</li>\n<li><strong>Do</strong>: Make sure your submission inference is done in a kaggle notebook.  Even though you can upload a CSV for submission, CSV only submissions will not be cosidered for the prize. Training can be done offline and uploaded as a dataset, however you may be required to show the training code as proof that you did not use exernal data or pretrained models.</li>\n<li><strong>Don't</strong>: Overthink it. This competition is intended to be fun so feel free to ask questions if you are unsure. If you learn something new then you've already won.</li>\n<li><strong>Do</strong>: Have fun!!!!!!!</li>\n</ul>\n<p>That is all. I am happy to clarify anything, just ask below.</p>",
  "messages": [
    {
      "id": "1720354",
      "postDate": "03/12/2022 17:51:57",
      "content": "<p>I'm so excited about this competition, it should be a lot of fun. I know everyone who takes part will learn something new. All of the information below can be found in the competition pages, but I wanted to highlight these so there is no confusion. Please keep these in mind when making your submissions:</p>\n<ul>\n<li><strong>Do</strong>: Research about machine learning techniques for audio data. Read papers, learn about state-of-the-art model archetectures, and try to implement them. If you like, you can also share what you've learned on the forums or in notebooks (up until the week before the deadline).</li>\n<li><strong>Don't</strong>: Use external data sources or pretrained models. Don't attempt to search for additional information about the songs (public or private set), and obviously don't hand label the test. This will disqualify you from the competition and prizes.</li>\n<li><strong>Do</strong>: Read the competition description and <a href=\"https://www.kaggle.com/c/kaggle-pog-series-s01e02/overview/prizes\" target=\"_blank\">prize elidgeability requirements</a>. Follow all the steps to ensure you are elidgeable to win the GPU and DLI vouchers.</li>\n<li><strong>Don't</strong>: Discuss the competition with people outside of your team, as it's considered private sharing and will disqualify you from prize elidgeability. We are here to learn together but not to cheat!</li>\n<li><strong>Do</strong>: Make sure your submission inference is done in a kaggle notebook.  Even though you can upload a CSV for submission, CSV only submissions will not be cosidered for the prize. Training can be done offline and uploaded as a dataset, however you may be required to show the training code as proof that you did not use exernal data or pretrained models.</li>\n<li><strong>Don't</strong>: Overthink it. This competition is intended to be fun so feel free to ask questions if you are unsure. If you learn something new then you've already won.</li>\n<li><strong>Do</strong>: Have fun!!!!!!!</li>\n</ul>\n<p>That is all. I am happy to clarify anything, just ask below.</p>",
      "rawMarkdown": "I'm so excited about this competition, it should be a lot of fun. I know everyone who takes part will learn something new. All of the information below can be found in the competition pages, but I wanted to highlight these so there is no confusion. Please keep these in mind when making your submissions:\n\n- **Do**: Research about machine learning techniques for audio data. Read papers, learn about state-of-the-art model archetectures, and try to implement them. If you like, you can also share what you've learned on the forums or in notebooks (up until the week before the deadline).\n- **Don't**: Use external data sources or pretrained models. Don't attempt to search for additional information about the songs (public or private set), and obviously don't hand label the test. This will disqualify you from the competition and prizes.\n- **Do**: Read the competition description and [prize elidgeability requirements](https://www.kaggle.com/c/kaggle-pog-series-s01e02/overview/prizes). Follow all the steps to ensure you are elidgeable to win the GPU and DLI vouchers.\n- **Don't**: Discuss the competition with people outside of your team, as it's considered private sharing and will disqualify you from prize elidgeability. We are here to learn together but not to cheat!\n- **Do**: Make sure your submission inference is done in a kaggle notebook.  Even though you can upload a CSV for submission, CSV only submissions will not be cosidered for the prize. Training can be done offline and uploaded as a dataset, however you may be required to show the training code as proof that you did not use exernal data or pretrained models.\n- **Don't**: Overthink it. This competition is intended to be fun so feel free to ask questions if you are unsure. If you learn something new then you've already won.\n- **Do**: Have fun!!!!!!!\n\nThat is all. I am happy to clarify anything, just ask below.",
      "votes": null
    },
    {
      "id": "1720705",
      "postDate": "03/13/2022 04:00:21",
      "content": "<p><code>Don't: Use external data sources or pretrained models.</code><br>\nCant we use the timm library for transfer learning?</p>",
      "rawMarkdown": "`Don't: Use external data sources or pretrained models.`\nCant we use the timm library for transfer learning?",
      "votes": null
    },
    {
      "id": "1720790",
      "postDate": "03/13/2022 06:29:09",
      "content": "<p>Qouting the rules tab, it seems this is not permitted:</p>\n<blockquote>\n  <p>No pretrained models. Everything needs to be created from scratch. </p>\n</blockquote>\n<p>I think this would make the approaches more interesting</p>",
      "rawMarkdown": "Qouting the rules tab, it seems this is not permitted:\n\n> No pretrained models. Everything needs to be created from scratch. \n\nI think this would make the approaches more interesting",
      "votes": null
    },
    {
      "id": "1721111",
      "postDate": "03/13/2022 12:24:18",
      "content": "<p><a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> is correct, pretrained models are not allowed. You are, however allowed to use <strong>architectures</strong> from timm, huggingface, etc. You just need to make sure to use <code>pretrained=False</code>. I don't forsee this being a big deal because a model pretrained on imagenet labels won't be a big help for a model trying to classifiy melspectorgrams. But either way it's not allowed to used pretrained weights.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "init27 is correct, pretrained models are not allowed. You are, however allowed to use **architectures** from timm, huggingface, etc. You just need to make sure to use `pretrained=False`. I don't forsee this being a big deal because a model pretrained on imagenet labels won't be a big help for a model trying to classifiy melspectorgrams. But either way it's not allowed to used pretrained weights.\n\nGood luck!",
      "votes": null
    },
    {
      "id": "1751341",
      "postDate": "04/10/2022 16:27:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> quick question. </p>\n<p>Is it okay to upload the predictions (npy/csv) generated on local system and use them directly in inference notebooks? Or do we have to restrict to generating prediction in notebooks via model.predict functions? </p>\n<p>The reason I ask is I am seeing a big difference in performance from locally generated predictions vs ones generated on kernels. I am still investigating the reasons. </p>",
      "rawMarkdown": "Hi @robikscube quick question. \n\nIs it okay to upload the predictions (npy/csv) generated on local system and use them directly in inference notebooks? Or do we have to restrict to generating prediction in notebooks via model.predict functions? \n\nThe reason I ask is I am seeing a big difference in performance from locally generated predictions vs ones generated on kernels. I am still investigating the reasons.",
      "votes": null
    },
    {
      "id": "1751367",
      "postDate": "04/10/2022 17:07:40",
      "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> - in order to be eligible for the prizes your inference must be made through the notebook. Uploading CSV would not be a valid submission for the prize.</p>",
      "rawMarkdown": "pheadrus - in order to be eligible for the prizes your inference must be made through the notebook. Uploading CSV would not be a valid submission for the prize.",
      "votes": null
    },
    {
      "id": "1751382",
      "postDate": "04/10/2022 17:38:43",
      "content": "<p>Got it, thanks. </p>",
      "rawMarkdown": "Got it, thanks.",
      "votes": null
    },
    {
      "id": "1762422",
      "postDate": "04/20/2022 17:07:13",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> one question.</p>\n<p>Is it okay to use the duplicate files but convert them into a .wav? </p>\n<p>That files don't have extra data only the same files but in .wav</p>",
      "rawMarkdown": "Hi, @robikscube one question.\n\nIs it okay to use the duplicate files but convert them into a .wav? \n\nThat files don't have extra data only the same files but in .wav",
      "votes": null
    },
    {
      "id": "1762626",
      "postDate": "04/20/2022 20:46:31",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/kaggle-pog-series-s01e02/discussion/313271#1724815\" target=\"_blank\">https://www.kaggle.com/competitions/kaggle-pog-series-s01e02/discussion/313271#1724815</a></p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/kaggle-pog-series-s01e02/discussion/313271#1724815",
      "votes": null
    },
    {
      "id": "1762669",
      "postDate": "04/20/2022 22:25:11",
      "content": "<p><a href=\"https://www.kaggle.com/asisheriberto\" target=\"_blank\">@asisheriberto</a> - if it's just the original dataset in a different format then that is OK to use.</p>",
      "rawMarkdown": "asisheriberto - if it's just the original dataset in a different format then that is OK to use.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1720705,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "03/13/2022 04:00:21",
      "content": "<p><code>Don't: Use external data sources or pretrained models.</code><br>\nCant we use the timm library for transfer learning?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1720790,
          "author_name": "init27",
          "author_url": "",
          "post_date": "03/13/2022 06:29:09",
          "content": "<p>Qouting the rules tab, it seems this is not permitted:</p>\n<blockquote>\n  <p>No pretrained models. Everything needs to be created from scratch. </p>\n</blockquote>\n<p>I think this would make the approaches more interesting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1721111,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "03/13/2022 12:24:18",
          "content": "<p><a href=\"https://www.kaggle.com/init27\" target=\"_blank\">@init27</a> is correct, pretrained models are not allowed. You are, however allowed to use <strong>architectures</strong> from timm, huggingface, etc. You just need to make sure to use <code>pretrained=False</code>. I don't forsee this being a big deal because a model pretrained on imagenet labels won't be a big help for a model trying to classifiy melspectorgrams. But either way it's not allowed to used pretrained weights.</p>\n<p>Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1751341,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "04/10/2022 16:27:13",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> quick question. </p>\n<p>Is it okay to upload the predictions (npy/csv) generated on local system and use them directly in inference notebooks? Or do we have to restrict to generating prediction in notebooks via model.predict functions? </p>\n<p>The reason I ask is I am seeing a big difference in performance from locally generated predictions vs ones generated on kernels. I am still investigating the reasons. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1751367,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "04/10/2022 17:07:40",
          "content": "<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> - in order to be eligible for the prizes your inference must be made through the notebook. Uploading CSV would not be a valid submission for the prize.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1751382,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "04/10/2022 17:38:43",
          "content": "<p>Got it, thanks. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1762422,
      "author_name": "asisheriberto",
      "author_url": "",
      "post_date": "04/20/2022 17:07:13",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> one question.</p>\n<p>Is it okay to use the duplicate files but convert them into a .wav? </p>\n<p>That files don't have extra data only the same files but in .wav</p>",
      "votes": null,
      "replies": [
        {
          "id": 1762626,
          "author_name": "dsx001",
          "author_url": "",
          "post_date": "04/20/2022 20:46:31",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/kaggle-pog-series-s01e02/discussion/313271#1724815\" target=\"_blank\">https://www.kaggle.com/competitions/kaggle-pog-series-s01e02/discussion/313271#1724815</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1762669,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "04/20/2022 22:25:11",
          "content": "<p><a href=\"https://www.kaggle.com/asisheriberto\" target=\"_blank\">@asisheriberto</a> - if it's just the original dataset in a different format then that is OK to use.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1720354": "I'm so excited about this competition, it should be a lot of fun. I know everyone who takes part will learn something new. All of the information below can be found in the competition pages, but I wanted to highlight these so there is no confusion. Please keep these in mind when making your submissions:\n\n- **Do**: Research about machine learning techniques for audio data. Read papers, learn about state-of-the-art model archetectures, and try to implement them. If you like, you can also share what you've learned on the forums or in notebooks (up until the week before the deadline).\n- **Don't**: Use external data sources or pretrained models. Don't attempt to search for additional information about the songs (public or private set), and obviously don't hand label the test. This will disqualify you from the competition and prizes.\n- **Do**: Read the competition description and [prize elidgeability requirements](https://www.kaggle.com/c/kaggle-pog-series-s01e02/overview/prizes). Follow all the steps to ensure you are elidgeable to win the GPU and DLI vouchers.\n- **Don't**: Discuss the competition with people outside of your team, as it's considered private sharing and will disqualify you from prize elidgeability. We are here to learn together but not to cheat!\n- **Do**: Make sure your submission inference is done in a kaggle notebook.  Even though you can upload a CSV for submission, CSV only submissions will not be cosidered for the prize. Training can be done offline and uploaded as a dataset, however you may be required to show the training code as proof that you did not use exernal data or pretrained models.\n- **Don't**: Overthink it. This competition is intended to be fun so feel free to ask questions if you are unsure. If you learn something new then you've already won.\n- **Do**: Have fun!!!!!!!\n\nThat is all. I am happy to clarify anything, just ask below.",
    "1720705": "`Don't: Use external data sources or pretrained models.`\nCant we use the timm library for transfer learning?",
    "1720790": "Qouting the rules tab, it seems this is not permitted:\n\n> No pretrained models. Everything needs to be created from scratch. \n\nI think this would make the approaches more interesting",
    "1721111": "init27 is correct, pretrained models are not allowed. You are, however allowed to use **architectures** from timm, huggingface, etc. You just need to make sure to use `pretrained=False`. I don't forsee this being a big deal because a model pretrained on imagenet labels won't be a big help for a model trying to classifiy melspectorgrams. But either way it's not allowed to used pretrained weights.\n\nGood luck!",
    "1751341": "Hi @robikscube quick question. \n\nIs it okay to upload the predictions (npy/csv) generated on local system and use them directly in inference notebooks? Or do we have to restrict to generating prediction in notebooks via model.predict functions? \n\nThe reason I ask is I am seeing a big difference in performance from locally generated predictions vs ones generated on kernels. I am still investigating the reasons.",
    "1751367": "pheadrus - in order to be eligible for the prizes your inference must be made through the notebook. Uploading CSV would not be a valid submission for the prize.",
    "1751382": "Got it, thanks.",
    "1762422": "Hi, @robikscube one question.\n\nIs it okay to use the duplicate files but convert them into a .wav? \n\nThat files don't have extra data only the same files but in .wav",
    "1762626": "https://www.kaggle.com/competitions/kaggle-pog-series-s01e02/discussion/313271#1724815",
    "1762669": "asisheriberto - if it's just the original dataset in a different format then that is OK to use."
  },
  "source": "meta"
}