{
  "id": 88064,
  "title": "Competition Rules discussion thread",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/88064",
  "author_name": "Frederic Font",
  "post_date": "2019-04-05T14:03:23.190000",
  "votes": 17,
  "comment_count": 106,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>Make sure that you check the <strong>competition rules</strong> before participating. The complete rules are explained in the rules section:  <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/rules\">https://www.kaggle.com/c/freesound-audio-tagging-2019/rules</a>\nThis is a summary of the most important rules:</p>\n\n<ul>\n<li>Unlike last year's edition of this task, participants are <strong>not allowed to use external data</strong> for system development. This also excludes the use of pre-trained models.</li>\n<li>Participants are <strong>not allowed to make subjective judgement​s</strong> of the test data, nor to annotate it (this includes the use of statistics about the evaluation dataset in the decision making). The test set cannot be used to train the submitted system.</li>\n<li>The winning teams are <strong>required to publish their systems under an open-source license</strong> in order to be considered winners. See <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/93054#latest-535382\">this thread</a> for more info.​</li>\n</ul>\n\n<p>Please post any questions and doubts you might have about competition rules in this discussion thread.</p>",
  "messages": [
    {
      "id": 508005,
      "postDate": "2019-04-05T14:03:23.190Z",
      "content": "<p>Hi,</p>\n\n<p>Make sure that you check the <strong>competition rules</strong> before participating. The complete rules are explained in the rules section:  <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/rules\">https://www.kaggle.com/c/freesound-audio-tagging-2019/rules</a>\nThis is a summary of the most important rules:</p>\n\n<ul>\n<li>Unlike last year's edition of this task, participants are <strong>not allowed to use external data</strong> for system development. This also excludes the use of pre-trained models.</li>\n<li>Participants are <strong>not allowed to make subjective judgement​s</strong> of the test data, nor to annotate it (this includes the use of statistics about the evaluation dataset in the decision making). The test set cannot be used to train the submitted system.</li>\n<li>The winning teams are <strong>required to publish their systems under an open-source license</strong> in order to be considered winners. See <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/93054#latest-535382\">this thread</a> for more info.​</li>\n</ul>\n\n<p>Please post any questions and doubts you might have about competition rules in this discussion thread.</p>",
      "rawMarkdown": "Hi,\n\nMake sure that you check the **competition rules** before participating. The complete rules are explained in the rules section:  https://www.kaggle.com/c/freesound-audio-tagging-2019/rules\nThis is a summary of the most important rules:\n\n- Unlike last year's edition of this task, participants are **not allowed to use external data** for system development. This also excludes the use of pre-trained models.\n- Participants are **not allowed to make subjective judgement​s** of the test data, nor to annotate it (this includes the use of statistics about the evaluation dataset in the decision making). The test set cannot be used to train the submitted system.\n- The winning teams are **required to publish their systems under an open-source license** in order to be considered winners. See [this thread](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/93054#latest-535382) for more info.​\n\nPlease post any questions and doubts you might have about competition rules in this discussion thread.",
      "votes": 17
    },
    {
      "id": 508290,
      "postDate": "2019-04-05T23:37:19.103Z",
      "content": "<p>I have the same question as <a href=\"/daisukelab\">@daisukelab</a>, the rules seem to be a bit bizarre to me.</p>\n\n<p>As I understand it, I can upload a pretrained model for inference, however, that model must NOT be pretrained on any other dataset before trained on Freesound data. So the question is, how do you know? How can you be sure that any teams outside the prize range didn't cheat?</p>\n\n<p>Another question is why did you choose to prohibit ImageNet models? I changed <code>pretrained=False</code> to <code>True</code> in FastAI and my validation score almost doubled up, I know there must be reasons behind it but to be honest it seems a bit silly to me that such a simple (and amazing) thing is not allowed.</p>",
      "rawMarkdown": "I have the same question as @daisukelab, the rules seem to be a bit bizarre to me.\n\nAs I understand it, I can upload a pretrained model for inference, however, that model must NOT be pretrained on any other dataset before trained on Freesound data. So the question is, how do you know? How can you be sure that any teams outside the prize range didn't cheat?\n\nAnother question is why did you choose to prohibit ImageNet models? I changed ```pretrained=False``` to ```True``` in FastAI and my validation score almost doubled up, I know there must be reasons behind it but to be honest it seems a bit silly to me that such a simple (and amazing) thing is not allowed.",
      "votes": 7,
      "replies": [
        {
          "id": 508361,
          "postDate": "2019-04-06T03:51:44Z",
          "content": "<p>Yeah it sounds more logical to either allow the use of pre-trained models or make it strictly kernel-based (including pre-processing, training, predicting). </p>",
          "rawMarkdown": "Yeah it sounds more logical to either allow the use of pre-trained models or make it strictly kernel-based (including pre-processing, training, predicting). ",
          "votes": 4
        },
        {
          "id": 509727,
          "postDate": "2019-04-08T07:47:55.610Z",
          "content": "<p>We want participants to have the freedom to train and do their experiments offline, outside Kaggle kernels. At the same time we want to encourage not to use external data or pre-trained models because we want participants to focus on the research challenges which are characteristic to our dataset. We believe in participant's honour to avoid cheating. For the winning submissions we'll revise them manually so we'll be able to detect that (code needs to be open sourced).</p>",
          "rawMarkdown": "We want participants to have the freedom to train and do their experiments offline, outside Kaggle kernels. At the same time we want to encourage not to use external data or pre-trained models because we want participants to focus on the research challenges which are characteristic to our dataset. We believe in participant's honour to avoid cheating. For the winning submissions we'll revise them manually so we'll be able to detect that (code needs to be open sourced).",
          "votes": 4
        },
        {
          "id": 509747,
          "postDate": "2019-04-08T08:27:04.857Z",
          "content": "<p>Thank you for answering my question <a href=\"/fredericfont\">@fredericfont</a>. </p>\n\n<p>Just to be clear, are we allowed to upload multiple models for ensembling, as long as the total inference time does not break the 1 hour limit?</p>",
          "rawMarkdown": "Thank you for answering my question @fredericfont. \n\nJust to be clear, are we allowed to upload multiple models for ensembling, as long as the total inference time does not break the 1 hour limit?"
        },
        {
          "id": 509758,
          "postDate": "2019-04-08T08:38:52.717Z",
          "content": "<p>yes</p>",
          "rawMarkdown": "yes",
          "votes": 1
        },
        {
          "id": 510430,
          "postDate": "2019-04-09T04:28:35.427Z",
          "content": "<p>How do you verify that the weight uploaded by the competition team derived from the pre-trained models?</p>",
          "rawMarkdown": "How do you verify that the weight uploaded by the competition team derived from the pre-trained models?",
          "votes": 4
        },
        {
          "id": 514592,
          "postDate": "2019-04-11T19:11:12.477Z",
          "content": "<p>We use a combination of the honor system and code review. Winning submissions will need to give us the code they used to train their models. This worked well for us last year where we placed restrictions on external datasets.</p>",
          "rawMarkdown": "We use a combination of the honor system and code review. Winning submissions will need to give us the code they used to train their models. This worked well for us last year where we placed restrictions on external datasets.",
          "votes": 1
        },
        {
          "id": 515865,
          "postDate": "2019-04-13T08:25:11.270Z",
          "content": "<p>I think this is a very bad decision. I would either allow usage of all possible data or totally forbid. Since only top3 will be checked, you can be 4th and get a solo gold by violating the rules. Solo gold + high kaggle points is most probably more valuable than money prize of 3rd:) <a href=\"/inversion\">@inversion</a> what do you think?</p>",
          "rawMarkdown": "I think this is a very bad decision. I would either allow usage of all possible data or totally forbid. Since only top3 will be checked, you can be 4th and get a solo gold by violating the rules. Solo gold + high kaggle points is most probably more valuable than money prize of 3rd:) @inversion what do you think?",
          "votes": 9
        },
        {
          "id": 524479,
          "postDate": "2019-04-28T21:23:00.307Z",
          "content": "<p>I've only been actively using Kaggle for a month or two and I already have seen people who cheat by making multiple accounts to test more submissions.  If people would go through all the trouble of making more accounts just for more submissions, they probably already have happily gone through with exactly what Ahmet has described above.  It's really a shame.</p>",
          "rawMarkdown": "I've only been actively using Kaggle for a month or two and I already have seen people who cheat by making multiple accounts to test more submissions.  If people would go through all the trouble of making more accounts just for more submissions, they probably already have happily gone through with exactly what Ahmet has described above.  It's really a shame.",
          "votes": 1
        },
        {
          "id": 535999,
          "postDate": "2019-05-23T19:09:42.770Z",
          "content": "<p><a href=\"/divrikwicky\">@divrikwicky</a> I am on your side. Seems only top3 will be checked if they used pre-trained external data. Teams are not in the top3 will not be checked, this is unfair for everyone.</p>",
          "rawMarkdown": "@divrikwicky I am on your side. Seems only top3 will be checked if they used pre-trained external data. Teams are not in the top3 will not be checked, this is unfair for everyone.",
          "votes": 1
        },
        {
          "id": 536010,
          "postDate": "2019-05-23T19:46:33.807Z",
          "content": "<p>Please see <a href=\"/addisonhoward\">@addisonhoward</a> 's message earlier in this thread\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064525061\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064525061</a></p>\n\n<p>The organizers and Kaggle reserve the right to check the submissions of any teams, not just the final winners, and the penalties could include banning of your account from Kaggle entirely. Also note that the final winners are determined by ranking on the private test set, and that will be determined only after the competition ends and may not be the top 3 from the public leaderboard.</p>\n\n<p>So if you try to cheat, you are gambling your Kaggle account on the chance that you will not end up as a winner on the private test set, and that you will not be among the teams that we choose to verify.  I'm not sure if losing your account is worth a few points from one challenge.</p>",
          "rawMarkdown": "Please see @addisonhoward 's message earlier in this thread\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064525061\n\nThe organizers and Kaggle reserve the right to check the submissions of any teams, not just the final winners, and the penalties could include banning of your account from Kaggle entirely. Also note that the final winners are determined by ranking on the private test set, and that will be determined only after the competition ends and may not be the top 3 from the public leaderboard.\n\nSo if you try to cheat, you are gambling your Kaggle account on the chance that you will not end up as a winner on the private test set, and that you will not be among the teams that we choose to verify.  I'm not sure if losing your account is worth a few points from one challenge.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 508071,
      "postDate": "2019-04-05T15:14:15.970Z",
      "content": "<p>Hi, let me confirm about pre-trained models.\nI'd like to work with ImageNet pre-trained CNN model provided by PyTorch by default.\nIt is NOT ALLOWED, correct?</p>\n\n<p>But <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements\">Kernels Requirements here</a> also shows that:\n- \"External data and pre-trained models are not allowed. If you train your model offline, you may upload your model as a Kaggle dataset for use in creating your inference Kernel.\"</p>\n\n<p>It sounds like we can use pre-trained models from Kaggle datasets, and we have PyTorch models here:\n- <a href=\"https://www.kaggle.com/pvlima/pretrained-pytorch-models\">https://www.kaggle.com/pvlima/pretrained-pytorch-models</a></p>\n\n<p>Can I use <a href=\"https://www.kaggle.com/pvlima/pretrained-pytorch-models\">https://www.kaggle.com/pvlima/pretrained-pytorch-models</a>?</p>\n\n<p>Thank you.</p>",
      "rawMarkdown": "Hi, let me confirm about pre-trained models.\nI'd like to work with ImageNet pre-trained CNN model provided by PyTorch by default.\nIt is NOT ALLOWED, correct?\n\nBut [Kernels Requirements here](https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements) also shows that:\n- \"External data and pre-trained models are not allowed. If you train your model offline, you may upload your model as a Kaggle dataset for use in creating your inference Kernel.\"\n\nIt sounds like we can use pre-trained models from Kaggle datasets, and we have PyTorch models here:\n- https://www.kaggle.com/pvlima/pretrained-pytorch-models\n\nCan I use https://www.kaggle.com/pvlima/pretrained-pytorch-models?\n\nThank you.",
      "votes": 6,
      "replies": [
        {
          "id": 509731,
          "postDate": "2019-04-08T07:55:29.787Z",
          "content": "<p>You can use models that you trained yourself with the data we provide and use them for inference, this is so that you can also work offline and not all training needs to happen in Kernel runtime. Use of models pre-trained with external data is forbidden.</p>",
          "rawMarkdown": "You can use models that you trained yourself with the data we provide and use them for inference, this is so that you can also work offline and not all training needs to happen in Kernel runtime. Use of models pre-trained with external data is forbidden.",
          "votes": 7
        },
        {
          "id": 509778,
          "postDate": "2019-04-08T09:08:22.173Z",
          "content": "<p>Thank you, it's clear now :)</p>",
          "rawMarkdown": "Thank you, it's clear now :)"
        }
      ]
    },
    {
      "id": 525061,
      "postDate": "2019-04-30T05:58:40.053Z",
      "content": "<p>Hi All,</p>\n\n<p>Just a reminder that Kaggle and the Freesound team reserve the right to review all kernels, not just those that win the competition, for adherence to the competition rules. Should it be determined that a participant intentioned tried to \"fly under the radar\" by finishing high enough for strong Kaggle points/medals, but hoping to not have their code reviewed, we reserve the right to not only disqualify that team from the competition, but also to ban the account entirely.</p>\n\n<p>These rules have been the same for all competitions which prohibit external data, and the integrity of the Kaggle community has resulted in only a few violations to this effect.</p>",
      "rawMarkdown": "Hi All,\n\nJust a reminder that Kaggle and the Freesound team reserve the right to review all kernels, not just those that win the competition, for adherence to the competition rules. Should it be determined that a participant intentioned tried to \"fly under the radar\" by finishing high enough for strong Kaggle points/medals, but hoping to not have their code reviewed, we reserve the right to not only disqualify that team from the competition, but also to ban the account entirely.\n\nThese rules have been the same for all competitions which prohibit external data, and the integrity of the Kaggle community has resulted in only a few violations to this effect.",
      "votes": 4,
      "replies": [
        {
          "id": 544763,
          "postDate": "2019-06-05T22:11:46.870Z",
          "content": "<p>Hi <a href=\"/addisonhoward\">@addisonhoward</a> , is it true that a team using pre-trained models will be banned from Kaggle? I've seem banned from competition, but from Kaggle I believe its the first time. </p>\n\n<p>Our solution uses only non-pretrained models, but we checked and using pre-trained models scores much higher in LB. I just want to be sure that all teams using pre-trained will be dropped from the LB. </p>",
          "rawMarkdown": "Hi @addisonhoward , is it true that a team using pre-trained models will be banned from Kaggle? I've seem banned from competition, but from Kaggle I believe its the first time. \n\nOur solution uses only non-pretrained models, but we checked and using pre-trained models scores much higher in LB. I just want to be sure that all teams using pre-trained will be dropped from the LB. ",
          "votes": 3
        },
        {
          "id": 544842,
          "postDate": "2019-06-06T00:25:46.877Z",
          "content": "<p>If we determine that a team has violated the rules in such a way as to obtain a high ranking without having to disclose not high enough to have to disclose their model (thus breaking the rules but hoping not to get caught), we reserve the right to ban them from the site at our discretion</p>",
          "rawMarkdown": "If we determine that a team has violated the rules in such a way as to obtain a high ranking without having to disclose not high enough to have to disclose their model (thus breaking the rules but hoping not to get caught), we reserve the right to ban them from the site at our discretion",
          "votes": 2
        },
        {
          "id": 546347,
          "postDate": "2019-06-06T13:59:51.177Z",
          "content": "<p>They'll create another account ). Kaggle is addictive, and we all know that, so the ban will not stop people from kaggling, imho ) (I am not in this competition, just saw it occasionally)</p>",
          "rawMarkdown": "They'll create another account ). Kaggle is addictive, and we all know that, so the ban will not stop people from kaggling, imho ) (I am not in this competition, just saw it occasionally)",
          "votes": 1
        },
        {
          "id": 546358,
          "postDate": "2019-06-06T14:05:33.600Z",
          "content": "<p>Very happy to see rules enforcement beyond prize winners.  We all can think of cases where there was a high suspicion of private sharing or external data use, but they flew under the radar.  If kernel only help in that matter then I'm way more interested in kernel only competition than before!</p>",
          "rawMarkdown": "Very happy to see rules enforcement beyond prize winners.  We all can think of cases where there was a high suspicion of private sharing or external data use, but they flew under the radar.  If kernel only help in that matter then I'm way more interested in kernel only competition than before!",
          "votes": 1
        },
        {
          "id": 546399,
          "postDate": "2019-06-06T14:41:16.937Z",
          "content": "<p>Thank you <a href=\"/addisonhoward\">@addisonhoward</a> . Hope this encourage teams not to use pretrained models, testset or any other external data :)</p>",
          "rawMarkdown": "Thank you @addisonhoward . Hope this encourage teams not to use pretrained models, testset or any other external data :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 527084,
      "postDate": "2019-05-04T14:48:12.070Z",
      "content": "<p>Hi, is it possible to clarify if we can use  “Utility Script On” or not which is introduced in this discussion?\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91241#latest-526152\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91241#latest-526152</a>\nThanks!</p>",
      "rawMarkdown": "Hi, is it possible to clarify if we can use  “Utility Script On” or not which is introduced in this discussion?\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91241#latest-526152\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 527545,
          "postDate": "2019-05-05T19:13:01.197Z",
          "content": "<p>I believe it should be fine to import any code you like. <a href=\"/addisonhoward\">@addisonhoward</a> </p>",
          "rawMarkdown": "I believe it should be fine to import any code you like. @addisonhoward ",
          "votes": 1
        },
        {
          "id": 527617,
          "postDate": "2019-05-06T00:23:02.997Z",
          "content": "<p>You can import code to train, but remember that the inference kernel must run standalone, and not be a part of linking to multiple other Kernels to try and circumnavigate the single inference kernel for scoring.</p>",
          "rawMarkdown": "You can import code to train, but remember that the inference kernel must run standalone, and not be a part of linking to multiple other Kernels to try and circumnavigate the single inference kernel for scoring.",
          "votes": 1
        },
        {
          "id": 527619,
          "postDate": "2019-05-06T00:29:59.500Z",
          "content": "<p>Thank you for clarification, I understood that final submission kernel CANNOT import code.\nI will keep final kernel to be standalone.\n(I was hoping not to copy&amp;paste basic library code)\nThanks again, it's clear.</p>",
          "rawMarkdown": "Thank you for clarification, I understood that final submission kernel CANNOT import code.\nI will keep final kernel to be standalone.\n(I was hoping not to copy&amp;paste basic library code)\nThanks again, it's clear."
        },
        {
          "id": 527646,
          "postDate": "2019-05-06T02:27:48.040Z",
          "content": "<p>I think you can import code just fine. Addison was pointing out that you should not try and link kernels in such a way that other kernels also run when we run the inference kernel.</p>\n\n<p>Note that in addition to using the new Kaggle feature of utility scripts, you can also import code by simply including the code in a Dataset (along with your pre-trained model) and attaching the Dataset to the kernel.</p>",
          "rawMarkdown": "I think you can import code just fine. Addison was pointing out that you should not try and link kernels in such a way that other kernels also run when we run the inference kernel.\n\nNote that in addition to using the new Kaggle feature of utility scripts, you can also import code by simply including the code in a Dataset (along with your pre-trained model) and attaching the Dataset to the kernel.",
          "votes": 5
        },
        {
          "id": 527647,
          "postDate": "2019-05-06T02:32:13.843Z",
          "content": "<p>Thanks Manoj - that is the correct interpretation</p>",
          "rawMarkdown": "Thanks Manoj - that is the correct interpretation"
        },
        {
          "id": 527671,
          "postDate": "2019-05-06T04:00:07.770Z",
          "content": "<p>Thank you <a href=\"/plakal\">@plakal</a> <a href=\"/addisonhoward\">@addisonhoward</a> for correction, I appreciate your help!</p>",
          "rawMarkdown": "Thank you @plakal @addisonhoward for correction, I appreciate your help!"
        }
      ]
    },
    {
      "id": 524089,
      "postDate": "2019-04-27T22:20:03.193Z",
      "content": "<p>Could you please be more specific on <em>The test set cannot be used to train the submitted system.</em>? Will unsupervised methods such as domain adaptation methods be considered as the rules violation?</p>",
      "rawMarkdown": "Could you please be more specific on *The test set cannot be used to train the submitted system.*? Will unsupervised methods such as domain adaptation methods be considered as the rules violation?",
      "votes": 1,
      "replies": [
        {
          "id": 524949,
          "postDate": "2019-04-29T20:47:33.687Z",
          "content": "<p>I think <a href=\"/fredericfont\">@fredericfont</a> was fairly clear in his post: you cannot use the test set in any way as an input in the training of your submission.</p>",
          "rawMarkdown": "I think @fredericfont was fairly clear in his post: you cannot use the test set in any way as an input in the training of your submission.",
          "votes": 1
        },
        {
          "id": 524951,
          "postDate": "2019-04-29T20:52:53.593Z",
          "content": "<p>Thank you for clarification :)</p>",
          "rawMarkdown": "Thank you for clarification :)"
        },
        {
          "id": 546359,
          "postDate": "2019-06-06T14:05:58.937Z",
          "content": "<p>I wish that was made clear in LANL competition as well!</p>",
          "rawMarkdown": "I wish that was made clear in LANL competition as well!"
        }
      ]
    },
    {
      "id": 535353,
      "postDate": "2019-05-22T18:32:17.640Z",
      "content": "<p>Hi,  I have one question about a kernel.  </p>\n\n<p>As everyone confirmed, we are allowed to upload our models as private external data. When we attached our external data to the inference scripts, directory environment will change. More specifically, in the default environment, sample_submission.csv and other related csv files will be just under the <code>input</code> folder. But after we added external data, the environment will be changed, and csv files will be under <code>input/freesound-audio-tagging-2019</code> folder.  </p>\n\n<p>I am now writing my kernel script to refer the changed folder location when reading csv and other audio files. <br>\nIs it safe to use such a hard-coded directory name (e.g. <code>input/freesound-audio-tagging-2019</code>) in 2nd stage ?  </p>\n\n<p>Thanks in advance.  </p>",
      "rawMarkdown": "Hi,  I have one question about a kernel.  \n  \nAs everyone confirmed, we are allowed to upload our models as private external data. When we attached our external data to the inference scripts, directory environment will change. More specifically, in the default environment, sample_submission.csv and other related csv files will be just under the `input` folder. But after we added external data, the environment will be changed, and csv files will be under `input/freesound-audio-tagging-2019` folder.  \n  \nI am now writing my kernel script to refer the changed folder location when reading csv and other audio files.  \nIs it safe to use such a hard-coded directory name (e.g. `input/freesound-audio-tagging-2019`) in 2nd stage ?  \n  \nThanks in advance.  ",
      "votes": 2,
      "replies": [
        {
          "id": 535415,
          "postDate": "2019-05-22T21:22:33.823Z",
          "content": "<p>You could get into trouble if you import a dataset into your script that later gets deleted. If you are sure you won't have that issue, you could hard code paths.</p>\n\n<p>A safer alternative might be something like:</p>\n\n<p><code>\ntry:\n &lt;read long path&gt;\nexcept:\n &lt;read short path&gt;\n</code></p>",
          "rawMarkdown": "You could get into trouble if you import a dataset into your script that later gets deleted. If you are sure you won't have that issue, you could hard code paths.\n\nA safer alternative might be something like:\n\n```\ntry:\n ",
          "votes": 3
        },
        {
          "id": 535427,
          "postDate": "2019-05-22T22:36:21.857Z",
          "content": "<p><a href=\"/inversion\">@inversion</a> \nThank you for your reply. <br>\nI understand. Actually I faced that situation in another kernel, it wired me...  </p>\n\n<p>So as long as I do NOT delete an external dataset attached to the inference kernel, hard coding the paths will not be a trouble, right ?</p>",
          "rawMarkdown": "@inversion \nThank you for your reply.  \nI understand. Actually I faced that situation in another kernel, it wired me...  \n  \nSo as long as I do NOT delete an external dataset attached to the inference kernel, hard coding the paths will not be a trouble, right ?",
          "votes": 1
        }
      ]
    },
    {
      "id": 533357,
      "postDate": "2019-05-19T04:18:20.780Z",
      "content": "<p>Hi, let me make sure the rule about pre-trained model.\nWe are allowed to train a model beforehand with the data set in this competition, but not with external data.\nAnd we are also not allowed to use pre-trained model even if it's included in some frameworks.</p>\n\n<p>I want to clarify what will happen after finishing the competition.\nThe codes of prize winners will be checked their reproducibility, but how about other medal winners?\nWill all of medal winner's codes be also reviewed whether they keep the rule or not?</p>\n\n<p>I think it should be so but it seems difficult to check over 100 teams' code.\nBut if not checked, that means someone against the rule can win the competition secretly.</p>\n\n<p>I just want to make sure so that every one can be fair and enjoy the competition!</p>",
      "rawMarkdown": "Hi, let me make sure the rule about pre-trained model.\nWe are allowed to train a model beforehand with the data set in this competition, but not with external data.\nAnd we are also not allowed to use pre-trained model even if it's included in some frameworks.\n\nI want to clarify what will happen after finishing the competition.\nThe codes of prize winners will be checked their reproducibility, but how about other medal winners?\nWill all of medal winner's codes be also reviewed whether they keep the rule or not?\n\nI think it should be so but it seems difficult to check over 100 teams' code.\nBut if not checked, that means someone against the rule can win the competition secretly.\n\nI just want to make sure so that every one can be fair and enjoy the competition!",
      "votes": 2,
      "replies": [
        {
          "id": 533733,
          "postDate": "2019-05-19T20:38:18.197Z",
          "content": "<p>Hi, you are correct: external data is not allowed (including models pre-trained on external data).</p>\n\n<p>I understand your concern. I will refer to this post by <a href=\"/addisonhoward\">@addisonhoward</a> \n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#525061\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#525061</a></p>\n\n<p>I will add that we like to trust participants' honor code. Besides, infringing the rules comes accompanied with the risk of being disqualified from the competition, and even being banned from Kaggle. It seems a high price to pay for a few Kaggle points/medals.</p>",
          "rawMarkdown": "Hi, you are correct: external data is not allowed (including models pre-trained on external data).\n\nI understand your concern. I will refer to this post by @addisonhoward \n[https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#525061](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#525061)\n\nI will add that we like to trust participants' honor code. Besides, infringing the rules comes accompanied with the risk of being disqualified from the competition, and even being banned from Kaggle. It seems a high price to pay for a few Kaggle points/medals.",
          "votes": 1
        },
        {
          "id": 533773,
          "postDate": "2019-05-20T02:09:09.857Z",
          "content": "<p>OK, thanks! I was relieved to hear that!</p>",
          "rawMarkdown": "OK, thanks! I was relieved to hear that!"
        }
      ]
    },
    {
      "id": 530806,
      "postDate": "2019-05-13T18:03:30.990Z",
      "content": "<p>Will the sample_submission.csv update in stage 2? I'd like to know it because I use it as a reference for file path to test data.</p>",
      "rawMarkdown": "Will the sample_submission.csv update in stage 2? I'd like to know it because I use it as a reference for file path to test data.",
      "votes": 2,
      "replies": [
        {
          "id": 530815,
          "postDate": "2019-05-13T18:28:41.297Z",
          "content": "<p>Correct - we'll be updating the sample submission file for stage 2</p>",
          "rawMarkdown": "Correct - we'll be updating the sample submission file for stage 2",
          "votes": 2
        },
        {
          "id": 531343,
          "postDate": "2019-05-14T17:27:27.353Z",
          "content": "<p>Thank you for your clarification!</p>",
          "rawMarkdown": "Thank you for your clarification!"
        }
      ]
    },
    {
      "id": 549520,
      "postDate": "2019-06-10T18:33:39.490Z",
      "content": "<p>I have a question about submitting our kernels. I have a kernel that I have been working on and have different scores on as I've made changes. On the submissions page I can select two scores/submissions that will be used as my final leaderboard score.</p>\n\n<p>If I select two scores from the same kernel (but different versions with different public scores) will everything work properly? As in: Does it automatically use the latest version of a kernel, or will it use the version and environment that produced the selected public LB score?</p>",
      "rawMarkdown": "I have a question about submitting our kernels. I have a kernel that I have been working on and have different scores on as I've made changes. On the submissions page I can select two scores/submissions that will be used as my final leaderboard score.\n\nIf I select two scores from the same kernel (but different versions with different public scores) will everything work properly? As in: Does it automatically use the latest version of a kernel, or will it use the version and environment that produced the selected public LB score?",
      "replies": [
        {
          "id": 549599,
          "postDate": "2019-06-10T21:04:18.840Z",
          "content": "<p>You can select a submission based on a specific kernel version - it will not automatically use the latest version of the kernel.</p>",
          "rawMarkdown": "You can select a submission based on a specific kernel version - it will not automatically use the latest version of the kernel."
        }
      ]
    },
    {
      "id": 548931,
      "postDate": "2019-06-10T05:41:30.723Z",
      "content": "<p>Hi, I have one question about a 2nd stage data .</p>\n\n<p>Is the fname of sample_submission.csv updated in 2nd stages guaranteed to be unique?</p>",
      "rawMarkdown": "Hi, I have one question about a 2nd stage data .\n\nIs the fname of sample_submission.csv updated in 2nd stages guaranteed to be unique?"
    },
    {
      "id": 546344,
      "postDate": "2019-06-06T13:52:42.583Z",
      "content": "<p>The second part of this statement about the kernel requirements is confusing. \"External data and <strong>pre-trained models</strong> are not allowed. If you train your model offline, you may upload your model as a Kaggle dataset for use in creating your inference Kernel.\" What does \"upload your model as a Kaggle dataset\" mean? What's the difference between a \"model\" and a \"<strong>pre-trained</strong> model\"? Does it mean you can upload a model description, like a json file? That seems rather useless. Does it mean we can upload pre-processed data? That would save tremendously on pre-training time, since it takes 1 hour to read in 1000 training files, 90 minutes for 1120 test files. From reading through the comments here, it doesn't seem that anybody really understands what it means. Please clarify.</p>",
      "rawMarkdown": "The second part of this statement about the kernel requirements is confusing. \"External data and **pre-trained models** are not allowed. If you train your model offline, you may upload your model as a Kaggle dataset for use in creating your inference Kernel.\" What does \"upload your model as a Kaggle dataset\" mean? What's the difference between a \"model\" and a \"**pre-trained** model\"? Does it mean you can upload a model description, like a json file? That seems rather useless. Does it mean we can upload pre-processed data? That would save tremendously on pre-training time, since it takes 1 hour to read in 1000 training files, 90 minutes for 1120 test files. From reading through the comments here, it doesn't seem that anybody really understands what it means. Please clarify.",
      "replies": [
        {
          "id": 546492,
          "postDate": "2019-06-06T16:10:04.293Z",
          "content": "<p>Pre-trained models in this case specifically refer to models that were trained by someone else. For example, in image classification, it is very common to use ImageNet as a pretrained model and to perform additional tuning from there. In this case, no pre-trained models are allowed. You can, however train your own model and upload it to a Kernel.</p>",
          "rawMarkdown": "Pre-trained models in this case specifically refer to models that were trained by someone else. For example, in image classification, it is very common to use ImageNet as a pretrained model and to perform additional tuning from there. In this case, no pre-trained models are allowed. You can, however train your own model and upload it to a Kernel."
        },
        {
          "id": 546556,
          "postDate": "2019-06-06T17:19:50.423Z",
          "content": "<p>Reallly!!? Cool. So, how do you do that? I think I attempted to attach a dataset to a kernel for this competition and the kernel failed to produce a submission.csv file. I figured it was because I was reaching for \"outside\" data.</p>",
          "rawMarkdown": "Reallly!!? Cool. So, how do you do that? I think I attempted to attach a dataset to a kernel for this competition and the kernel failed to produce a submission.csv file. I figured it was because I was reaching for \"outside\" data."
        }
      ]
    },
    {
      "id": 543034,
      "postDate": "2019-06-04T09:36:45.737Z",
      "content": "<p>Hi, I have a question about the size of the second-stage test set.\nDoes it mean the \"total duration\" of the second-stage test set is approximately three times the first-stage?\nOr it means the \"total number of clips\" of the second-stage test set is approximately three times the first-stage?</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Hi, I have a question about the size of the second-stage test set.\nDoes it mean the \"total duration\" of the second-stage test set is approximately three times the first-stage?\nOr it means the \"total number of clips\" of the second-stage test set is approximately three times the first-stage?\n\nThank you",
      "replies": [
        {
          "id": 543662,
          "postDate": "2019-06-04T17:18:08.950Z",
          "content": "<p>Hi Kai,</p>\n\n<p>The second-stage test set is based on the number of samples. The distributions of length and size should be approximately the same.</p>",
          "rawMarkdown": "Hi Kai,\n\nThe second-stage test set is based on the number of samples. The distributions of length and size should be approximately the same.",
          "votes": 1
        }
      ]
    },
    {
      "id": 542753,
      "postDate": "2019-06-04T04:33:23.020Z",
      "content": "<p>I have a question regarding to this rule </p>\n\n<blockquote>\n  <p>Participants are not allowed to make subjective judgement​s of the test data, nor to annotate it (this includes the use of statistics about the evaluation dataset in the decision making). The test set cannot be used to train the submitted system. </p>\n</blockquote>\n\n<p>I also got a similar answer from you bellow: </p>\n\n<blockquote>\n  <p>Hi, we mean that using a <strong>priori stats of the public test set to make decisions</strong> over the private test set is not allowed. Specifically, answering your question: weighting factors to tune the distribution of classes in the answer set are not allowed. The class label distribution in the test set ensures a minimum number of labels per class, but the distribution is a bit imbalanced.</p>\n</blockquote>\n\n<p>My addition question is: \nIf I use a postprocessing method to make the distribution of the final submission as same as <strong>train set</strong>. Is it allowed ? </p>\n\n<p>Thank you </p>",
      "rawMarkdown": "I have a question regarding to this rule \n\n&gt; Participants are not allowed to make subjective judgement​s of the test data, nor to annotate it (this includes the use of statistics about the evaluation dataset in the decision making). The test set cannot be used to train the submitted system. \n\nI also got a similar answer from you bellow: \n&gt; Hi, we mean that using a **priori stats of the public test set to make decisions** over the private test set is not allowed. Specifically, answering your question: weighting factors to tune the distribution of classes in the answer set are not allowed. The class label distribution in the test set ensures a minimum number of labels per class, but the distribution is a bit imbalanced.\n\nMy addition question is: \nIf I use a postprocessing method to make the distribution of the final submission as same as **train set**. Is it allowed ? \n\nThank you ",
      "replies": [
        {
          "id": 543976,
          "postDate": "2019-06-05T02:11:55.817Z",
          "content": "<p>If it is about the train set, I think it is allowed. But, assuming you mean distribution of class labels, note that the train set is quasi balanced and the private test set might not be that way...</p>",
          "rawMarkdown": "If it is about the train set, I think it is allowed. But, assuming you mean distribution of class labels, note that the train set is quasi balanced and the private test set might not be that way...",
          "votes": 2
        }
      ]
    },
    {
      "id": 540135,
      "postDate": "2019-05-31T04:45:47.310Z",
      "content": "<p>Which of these situations would count as \"training on the test set\"?\n1)  I decide on an architecture because it does well on the test set.\n2)  I average my two best (public score) models together . <br>\n3) Two teams merge and average their individual best (public score) models.  </p>",
      "rawMarkdown": "Which of these situations would count as \"training on the test set\"?\n1)  I decide on an architecture because it does well on the test set.\n2)  I average my two best (public score) models together .  \n3) Two teams merge and average their individual best (public score) models.  ",
      "replies": [
        {
          "id": 540153,
          "postDate": "2019-05-31T05:28:04.293Z",
          "content": "<p>None</p>",
          "rawMarkdown": "None"
        },
        {
          "id": 540540,
          "postDate": "2019-05-31T16:02:15.867Z",
          "content": "<p>You are free to use your public leaderboard scores to make decisions about your models, including architecture, hyperparameters, etc.</p>\n\n<p>What we disallow is directly feeding the test data into the training pipeline of your model, whether as unlabeled data (for unsupervised/semi-supervised techniques, or just for computing statistics), or as labeled data (where you label it yourself).</p>\n\n<p>At the end of the contest, it should be possible to train your submitted model from scratch using the training data we have given you, and no other data.</p>",
          "rawMarkdown": "You are free to use your public leaderboard scores to make decisions about your models, including architecture, hyperparameters, etc.\n\nWhat we disallow is directly feeding the test data into the training pipeline of your model, whether as unlabeled data (for unsupervised/semi-supervised techniques, or just for computing statistics), or as labeled data (where you label it yourself).\n\nAt the end of the contest, it should be possible to train your submitted model from scratch using the training data we have given you, and no other data."
        }
      ]
    },
    {
      "id": 534681,
      "postDate": "2019-05-21T17:13:35.887Z",
      "content": "<p>Will it be possible to make a submission to be evaluated (but, of course, not rated) after the end of the competition? </p>",
      "rawMarkdown": "Will it be possible to make a submission to be evaluated (but, of course, not rated) after the end of the competition? ",
      "replies": [
        {
          "id": 534683,
          "postDate": "2019-05-21T17:16:30.720Z",
          "content": "<p>Yes - you will be able to make a submission after the competition has concluded, in which you will receive a score, but will not be shown on the leaderboard.</p>",
          "rawMarkdown": "Yes - you will be able to make a submission after the competition has concluded, in which you will receive a score, but will not be shown on the leaderboard."
        }
      ]
    },
    {
      "id": 533426,
      "postDate": "2019-05-19T07:54:41.793Z",
      "content": "<p>Hi, I have two questions.</p>\n\n<p>In the second stage, it is you to rerun our kernel? Not ourselves? So how we upload the p retrained weights as a dataset so that you could find it? If we upload them as a private dataset, add them in the kernel, will you be able to access them?</p>\n\n<p>Also, the commit button in kaggle takes me a lot of time usually. Before seeing the output file, it will stay in that page for a long time…. will that time be counted as run time in the second stage?</p>",
      "rawMarkdown": "Hi, I have two questions.\n\nIn the second stage, it is you to rerun our kernel? Not ourselves? So how we upload the p retrained weights as a dataset so that you could find it? If we upload them as a private dataset, add them in the kernel, will you be able to access them?\n\nAlso, the commit button in kaggle takes me a lot of time usually. Before seeing the output file, it will stay in that page for a long time…. will that time be counted as run time in the second stage?"
    },
    {
      "id": 531326,
      "postDate": "2019-05-14T16:45:45.473Z",
      "content": "<p>Can we use the pre-trained melspectrogram script for our final submission? </p>",
      "rawMarkdown": "Can we use the pre-trained melspectrogram script for our final submission? ",
      "replies": [
        {
          "id": 533717,
          "postDate": "2019-05-19T19:56:18.263Z",
          "content": "<p>Hi, not sure I understand what you mean by <em>pre-trained melspectrogram script</em>. Could you please clarify?</p>",
          "rawMarkdown": "Hi, not sure I understand what you mean by *pre-trained melspectrogram script*. Could you please clarify?"
        }
      ]
    },
    {
      "id": 528306,
      "postDate": "2019-05-07T13:05:34.100Z",
      "content": "<p>Hello all,\non the competition rules, it is stated that the locally or kernel trained models should be uploaded as a kaggle dataset. In case that i want to fit a model on a kernel, can i just import the models as a kernel output, on the final submission kernel, or i will have to download them locally and then upload them as a kaggle dataset?</p>",
      "rawMarkdown": "Hello all,\non the competition rules, it is stated that the locally or kernel trained models should be uploaded as a kaggle dataset. In case that i want to fit a model on a kernel, can i just import the models as a kernel output, on the final submission kernel, or i will have to download them locally and then upload them as a kaggle dataset?",
      "replies": [
        {
          "id": 529426,
          "postDate": "2019-05-09T21:47:50.293Z",
          "content": "<p>Hi, \nbased on this response from <a href=\"/addisonhoward\">@addisonhoward</a> I think you must run the inference kernel standalone, and not link it to other Kernels.</p>\n\n<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#527617\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#527617</a></p>",
          "rawMarkdown": "Hi, \nbased on this response from @addisonhoward I think you must run the inference kernel standalone, and not link it to other Kernels.\n\n[https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#527617](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#527617)"
        },
        {
          "id": 529433,
          "postDate": "2019-05-09T22:09:42.470Z",
          "content": "<p>Thank you for your response\nSo, in order to be sure that any error will not ocure, i will have to make a kernel with my uploaded weights as a kaggle dataset and i will have to read the test data from the directory:\n'../input/freesound-audio-tagging-2019/test/xxxxxx.wav'</p>",
          "rawMarkdown": "Thank you for your response\nSo, in order to be sure that any error will not ocure, i will have to make a kernel with my uploaded weights as a kaggle dataset and i will have to read the test data from the directory:\n'../input/freesound-audio-tagging-2019/test/xxxxxx.wav'"
        }
      ]
    },
    {
      "id": 527362,
      "postDate": "2019-05-05T08:20:32.800Z",
      "content": "<p>Hi,\nWill the maximum kernel run time of the second stage of the competition change? As the size of the test data increases, the time taken to process these data will increase and I am not sure that it will be easy to keep up with 60 minutes on GPU...</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi,\nWill the maximum kernel run time of the second stage of the competition change? As the size of the test data increases, the time taken to process these data will increase and I am not sure that it will be easy to keep up with 60 minutes on GPU...\n\nThanks!",
      "replies": [
        {
          "id": 527546,
          "postDate": "2019-05-05T19:14:23.597Z",
          "content": "<p>The run-time limits will not change. You will need to make sure that your submission kernel can run within the limits when we pass in a private test set that is around 3 times the size of the public test set.</p>",
          "rawMarkdown": "The run-time limits will not change. You will need to make sure that your submission kernel can run within the limits when we pass in a private test set that is around 3 times the size of the public test set.",
          "votes": 1
        }
      ]
    },
    {
      "id": 525396,
      "postDate": "2019-04-30T22:06:18.953Z",
      "content": "<p>I have a problem. I use kernel named <em>submission</em> that uses the output from kernel <em>training</em>. Training kernel uses the output from <em>features</em>. In this pipeline, I'm not able to use <em>submission</em> kernel because of the error:</p>\n\n<blockquote>\n  <p>your input kernel [training] cannot use kernel outputs as a data source for this competition</p>\n</blockquote>\n\n<p>I think it's very disturbing. The architecture of the pipeline I've chosen has to be rewritten. Did anyone have the same problem? Is it necessary to block such data flow? Was it mentioned anywhere here?</p>\n\n<p>I can apparently add only one kernel depth to submission kernel.</p>",
      "rawMarkdown": "I have a problem. I use kernel named *submission* that uses the output from kernel *training*. Training kernel uses the output from *features*. In this pipeline, I'm not able to use *submission* kernel because of the error:\n\n&gt; your input kernel [training] cannot use kernel outputs as a data source for this competition\n\nI think it's very disturbing. The architecture of the pipeline I've chosen has to be rewritten. Did anyone have the same problem? Is it necessary to block such data flow? Was it mentioned anywhere here?\n\nI can apparently add only one kernel depth to submission kernel.",
      "replies": [
        {
          "id": 525438,
          "postDate": "2019-05-01T01:47:41.303Z",
          "content": "<p>One alternative is to use the outputs from your features &amp; training kernels as a Dataset (it can be private) and then just attach the dataset to the kernel that you want to use those files in. </p>",
          "rawMarkdown": "One alternative is to use the outputs from your features &amp; training kernels as a Dataset (it can be private) and then just attach the dataset to the kernel that you want to use those files in. ",
          "votes": 1
        },
        {
          "id": 525456,
          "postDate": "2019-05-01T04:03:47.527Z",
          "content": "<p>I also recommend to do as James wrote. We could minimize dependency to the complex Kaggle system. System can be down anytime...</p>",
          "rawMarkdown": "I also recommend to do as James wrote. We could minimize dependency to the complex Kaggle system. System can be down anytime...",
          "votes": 2
        }
      ]
    },
    {
      "id": 522813,
      "postDate": "2019-04-25T04:31:43.830Z",
      "content": "<p>Hi Frederic,\nAs you said \"the size of the second-stage test set is approximately three times the size of the first\", maybe 3.3 or 2.7.\nBecause of the limited time in second-stage, how many samples in the second-stage test set?\nThanks</p>",
      "rawMarkdown": "Hi Frederic,\nAs you said \"the size of the second-stage test set is approximately three times the size of the first\", maybe 3.3 or 2.7.\nBecause of the limited time in second-stage, how many samples in the second-stage test set?\nThanks",
      "replies": [
        {
          "id": 522948,
          "postDate": "2019-04-25T09:16:26.663Z",
          "content": "<p>Hi, I can't tell you the exact number, the 3x guideline should be enough as is it quite realistic.</p>",
          "rawMarkdown": "Hi, I can't tell you the exact number, the 3x guideline should be enough as is it quite realistic."
        },
        {
          "id": 522958,
          "postDate": "2019-04-25T09:34:04.477Z",
          "content": "<p>Thank you</p>",
          "rawMarkdown": "Thank you"
        }
      ]
    },
    {
      "id": 522102,
      "postDate": "2019-04-23T21:31:07.253Z",
      "content": "<p>Hi,\nThe competition rules state that <code>No custom packages enabled in kernels</code>.\nI keep all of my training/preprocessing/inference code in a GitHub repository, does this mean that my only option is to manually copy-paste all my codebase each time I make a submission? </p>\n\n<p>Also, just out of curiosity, why aren't we allowed to use custom packages? (:</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi,\nThe competition rules state that `No custom packages enabled in kernels`.\nI keep all of my training/preprocessing/inference code in a GitHub repository, does this mean that my only option is to manually copy-paste all my codebase each time I make a submission? \n\nAlso, just out of curiosity, why aren't we allowed to use custom packages? (:\n\nThanks",
      "replies": [
        {
          "id": 522248,
          "postDate": "2019-04-24T05:09:06.363Z",
          "content": "<p>You should be able to upload all your (Python) code as well as your trained model checkpoint into a Kaggle Dataset which you can attach to your kernel for submission.</p>\n\n<p><a href=\"/addisonhoward\">@addisonhoward</a> for the motivation behind the custom packages restriction.</p>",
          "rawMarkdown": "You should be able to upload all your (Python) code as well as your trained model checkpoint into a Kaggle Dataset which you can attach to your kernel for submission.\n\n@addisonhoward for the motivation behind the custom packages restriction."
        },
        {
          "id": 522655,
          "postDate": "2019-04-24T19:14:36.847Z",
          "content": "<p>Great! Thanks.</p>",
          "rawMarkdown": "Great! Thanks."
        }
      ]
    },
    {
      "id": 519847,
      "postDate": "2019-04-19T18:18:58.183Z",
      "content": "<p>Hello to all and thank the organizers for an interesting competition...</p>\n\n<p>Regarding the second phase of this competition : \n1. Is the second phase inference going to be automatic like the Petfinder competition or are we going to manually rerun our best performing kernels? What happens if we exceed the 1hour runtime. Are we automatically off or we will be given a second chance to use a \"lighter\" version of the kernel?</p>\n\n<ol>\n<li>In the second case, are we allowed to preprocess eg. extract spectrogram features from the test set before the final submission? This way we may save a couple of minutes from the runtime. On the other hand test audio preprocessing is considered part of the inference.. Just want to make sure I got it right..</li>\n</ol>\n\n<p>I hope I made my self clear....</p>",
      "rawMarkdown": "Hello to all and thank the organizers for an interesting competition...\n\nRegarding the second phase of this competition : \n1. Is the second phase inference going to be automatic like the Petfinder competition or are we going to manually rerun our best performing kernels? What happens if we exceed the 1hour runtime. Are we automatically off or we will be given a second chance to use a \"lighter\" version of the kernel?\n\n2. In the second case, are we allowed to preprocess eg. extract spectrogram features from the test set before the final submission? This way we may save a couple of minutes from the runtime. On the other hand test audio preprocessing is considered part of the inference.. Just want to make sure I got it right..\n\nI hope I made my self clear....\n",
      "replies": [
        {
          "id": 521351,
          "postDate": "2019-04-22T20:18:41.677Z",
          "content": "<ol>\n<li><p>We will run all submitted kernels on the private test set after the competition ends on June 10. Kernels that hit the limits described at <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements\">https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements</a> will be disqualified.  There will be no second chances, so you should submit a kernel that can scale to a 3x larger  data set.</p></li>\n<li><p>I'm not sure I follow because we will be running your submitted kernel on an unseen private test set so no amount of preprocessing using the public test set is going to help you there. If your question is about separating preprocessing time from inference time in general, then we do not support any such separation. The kernel time limits are for all computation that you need to perform on the test set, including preprocessing and inference.</p></li>\n</ol>",
          "rawMarkdown": "1. We will run all submitted kernels on the private test set after the competition ends on June 10. Kernels that hit the limits described at https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements will be disqualified.  There will be no second chances, so you should submit a kernel that can scale to a 3x larger  data set.\n\n2. I'm not sure I follow because we will be running your submitted kernel on an unseen private test set so no amount of preprocessing using the public test set is going to help you there. If your question is about separating preprocessing time from inference time in general, then we do not support any such separation. The kernel time limits are for all computation that you need to perform on the test set, including preprocessing and inference."
        },
        {
          "id": 521391,
          "postDate": "2019-04-22T21:13:21.987Z",
          "content": "<p>Perfect.. thank you very much.</p>",
          "rawMarkdown": "Perfect.. thank you very much."
        }
      ]
    },
    {
      "id": 516225,
      "postDate": "2019-04-13T19:57:44.380Z",
      "content": "<p>Hi,\nThis would be my first competition, and I'm wondering if someone wouldn't mind answering my newbie questions.\n1. What does it mean to use kernels only for inference?\n2. Can I preprocess the files and save features like mfcc in advance and upload to the Kernel to use later during the evaluation?\n3. Can I train the model in advance using the provided files, and upload the model weights to the kernel to use to predict?\n4. When my model gets evaluated with the private test set, how do I make sure my kernel can process the private test set? Would the files in test folder and sample_submission.csv be automatically swapped with the private data set?\nThanks for your help!</p>",
      "rawMarkdown": "Hi,\nThis would be my first competition, and I'm wondering if someone wouldn't mind answering my newbie questions.\n1. What does it mean to use kernels only for inference?\n2. Can I preprocess the files and save features like mfcc in advance and upload to the Kernel to use later during the evaluation?\n3. Can I train the model in advance using the provided files, and upload the model weights to the kernel to use to predict?\n4. When my model gets evaluated with the private test set, how do I make sure my kernel can process the private test set? Would the files in test folder and sample_submission.csv be automatically swapped with the private data set?\nThanks for your help!",
      "replies": [
        {
          "id": 518145,
          "postDate": "2019-04-16T22:59:34.503Z",
          "content": "<p>Hi James,</p>\n\n<p>1) Inference only means that your kernel only needs to make predictions, it does not need to also train the data. If you win the competition, you will also need to provide your full solution, but to submit to the competition, you only need the inference kernel.\n2-3) you can train/preprocess in another kernel or on your local machine. You can then upload those weights or other features to your inference kernel.\n4) The second-stage test set is approximately <strong>three times the size of the first</strong>. You should plan your kernel's memory, disk, and runtime footprint accordingly.</p>",
          "rawMarkdown": "Hi James,\n\n1) Inference only means that your kernel only needs to make predictions, it does not need to also train the data. If you win the competition, you will also need to provide your full solution, but to submit to the competition, you only need the inference kernel.\n2-3) you can train/preprocess in another kernel or on your local machine. You can then upload those weights or other features to your inference kernel.\n4) The second-stage test set is approximately **three times the size of the first**. You should plan your kernel's memory, disk, and runtime footprint accordingly."
        }
      ]
    },
    {
      "id": 515536,
      "postDate": "2019-04-12T18:30:47.730Z",
      "content": "<p>Hello, </p>\n\n<p>Are we allowed to manually label train data?</p>",
      "rawMarkdown": "Hello, \n\nAre we allowed to manually label train data?",
      "replies": [
        {
          "id": 517734,
          "postDate": "2019-04-16T12:48:09.333Z",
          "content": "<p>Yes you can experiment with labelling <em>train set data</em>, but never the test set. In any case remember that the goal of the task is precisely how to solve the label noise problem through <strong>automatic</strong> methods. Manually replacing noisy labels with better ones is therefore not a solution to the problem.</p>",
          "rawMarkdown": "Yes you can experiment with labelling *train set data*, but never the test set. In any case remember that the goal of the task is precisely how to solve the label noise problem through **automatic** methods. Manually replacing noisy labels with better ones is therefore not a solution to the problem.",
          "votes": 1
        },
        {
          "id": 553900,
          "postDate": "2019-06-16T14:57:16.653Z",
          "content": "<p><a href=\"/fredericfont\">@fredericfont</a> Our team has never thought manual relabeling. Also, we didn't remove the six corrupted files because the robustness to noisy data is the aim of this competition. But now, it revealed manual relabeling is a very effective strategy in this competition. It's not matched to the concept of this competition that asking training with limited data. Relabeling should have been considered as external data, I think. Even if a competition that permits external data, hand relabeling is considered as external data and required to share in public. This is a very strange rule. I'm sorry for I didn't notice this before this competition ends.</p>",
          "rawMarkdown": "@fredericfont Our team has never thought manual relabeling. Also, we didn't remove the six corrupted files because the robustness to noisy data is the aim of this competition. But now, it revealed manual relabeling is a very effective strategy in this competition. It's not matched to the concept of this competition that asking training with limited data. Relabeling should have been considered as external data, I think. Even if a competition that permits external data, hand relabeling is considered as external data and required to share in public. This is a very strange rule. I'm sorry for I didn't notice this before this competition ends.",
          "votes": 4
        },
        {
          "id": 554101,
          "postDate": "2019-06-17T01:34:03.940Z",
          "content": "<p><a href=\"/fredericfont\">@fredericfont</a> I agree with OsciiArt. As we all know, the concept of this competition is to create a model from small curated data and large noisy data. The essence of this concept is to create a robust model without making a lot of curated data. I think allowing relabeling is against this concept.</p>",
          "rawMarkdown": "@fredericfont I agree with OsciiArt. As we all know, the concept of this competition is to create a model from small curated data and large noisy data. The essence of this concept is to create a robust model without making a lot of curated data. I think allowing relabeling is against this concept."
        },
        {
          "id": 554586,
          "postDate": "2019-06-17T17:36:34.547Z",
          "content": "<p>Can you elaborate on this revelation that manual relabeling is a very effective strategy in this competition? Are you talking about a specific team? Also I'm curious about how you know the strategy to be effective before the private leaderboard has been released.</p>",
          "rawMarkdown": "Can you elaborate on this revelation that manual relabeling is a very effective strategy in this competition? Are you talking about a specific team? Also I'm curious about how you know the strategy to be effective before the private leaderboard has been released.",
          "votes": 2
        },
        {
          "id": 554602,
          "postDate": "2019-06-17T18:07:40.607Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924</a>\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95785\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95785</a></p>",
          "rawMarkdown": "https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95785"
        },
        {
          "id": 554608,
          "postDate": "2019-06-17T18:20:31.650Z",
          "content": "<p>My suggestion is to wait for the private leaderboard :)</p>",
          "rawMarkdown": "My suggestion is to wait for the private leaderboard :)",
          "votes": 1
        },
        {
          "id": 564175,
          "postDate": "2019-06-29T04:48:02.467Z",
          "content": "<p>I saw the final result. I strongly hope the organizer considers changing the rule next year.</p>",
          "rawMarkdown": "I saw the final result. I strongly hope the organizer considers changing the rule next year.",
          "votes": 1
        },
        {
          "id": 564263,
          "postDate": "2019-06-29T07:59:11.337Z",
          "content": "<p>Agreed it defeats the purpose of the competition to manually relabel the training set</p>",
          "rawMarkdown": "Agreed it defeats the purpose of the competition to manually relabel the training set"
        }
      ]
    },
    {
      "id": 514284,
      "postDate": "2019-04-11T14:51:19.570Z",
      "content": "<p>Hi Frederic. \nI see there are 1120 samples in the current test set. So, how many test samples in the stage 2? Would all the test sample of stage 2  also come from Freesound Dataset (FSD) source?</p>",
      "rawMarkdown": "Hi Frederic. \nI see there are 1120 samples in the current test set. So, how many test samples in the stage 2? Would all the test sample of stage 2  also come from Freesound Dataset (FSD) source?",
      "replies": [
        {
          "id": 514316,
          "postDate": "2019-04-11T15:07:59.817Z",
          "content": "<p>Hi, as explained in the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">data section</a> of the competition page, the contents of the test set consist of manually-labeled data from FSD, and the size of the second-stage test set is approximately three times the size of the first.</p>",
          "rawMarkdown": "Hi, as explained in the [data section](https://www.kaggle.com/c/freesound-audio-tagging-2019/data) of the competition page, the contents of the test set consist of manually-labeled data from FSD, and the size of the second-stage test set is approximately three times the size of the first.",
          "votes": 3
        },
        {
          "id": 514323,
          "postDate": "2019-04-11T15:12:04.623Z",
          "content": "<p>Got it. Thank you!</p>",
          "rawMarkdown": "Got it. Thank you!"
        }
      ]
    },
    {
      "id": 511862,
      "postDate": "2019-04-10T11:48:44.167Z",
      "content": "<p>Just in case, I'd like to confirm whether the trained weight of my model should be <strong>public</strong> on Kaggle dataset, if I decide to train my model on local machine? Or it can be kept <strong>private</strong>?</p>",
      "rawMarkdown": "Just in case, I'd like to confirm whether the trained weight of my model should be **public** on Kaggle dataset, if I decide to train my model on local machine? Or it can be kept **private**?",
      "replies": [
        {
          "id": 512222,
          "postDate": "2019-04-10T16:31:28.070Z",
          "content": "<p>Hi Kenmatsu,</p>\n\n<p>You do not need to make your own trained weights public. But you can not use external data to train your weights :)</p>",
          "rawMarkdown": "Hi Kenmatsu,\n\nYou do not need to make your own trained weights public. But you can not use external data to train your weights :)"
        }
      ]
    },
    {
      "id": 1402079,
      "postDate": "2021-07-27T20:32:39.173Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 543714,
      "postDate": "2019-06-04T18:03:21.217Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 543744,
          "postDate": "2019-06-04T18:54:37.293Z",
          "content": "<p>Sorry - the rules acceptance deadline has passed. Once the competition concludes, you can make a \"late submission\" where you will receive a score, but the leaderboard will not update.</p>",
          "rawMarkdown": "Sorry - the rules acceptance deadline has passed. Once the competition concludes, you can make a \"late submission\" where you will receive a score, but the leaderboard will not update."
        }
      ]
    },
    {
      "id": 543698,
      "postDate": "2019-06-04T17:53:55.637Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 542955,
      "postDate": "2019-06-04T08:21:21.833Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 537074,
      "postDate": "2019-05-26T05:59:27.847Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 532719,
      "postDate": "2019-05-17T15:21:09.043Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 533728,
          "postDate": "2019-05-19T20:21:43.057Z",
          "content": "<p>Hi, we mean that using a priori stats of the public test set to make decisions over the private test set is not allowed. Specifically, answering your question: weighting factors to tune the distribution of classes in the answer set are not allowed.</p>\n\n<p>The class label distribution in the test set ensures a minimum number of labels per class, but the distribution is a bit imbalanced.</p>",
          "rawMarkdown": "Hi, we mean that using a priori stats of the public test set to make decisions over the private test set is not allowed. Specifically, answering your question: weighting factors to tune the distribution of classes in the answer set are not allowed.\n\nThe class label distribution in the test set ensures a minimum number of labels per class, but the distribution is a bit imbalanced.",
          "votes": 1
        }
      ]
    },
    {
      "id": 509649,
      "postDate": "2019-04-08T05:00:09.740Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 510209,
          "postDate": "2019-04-08T20:22:15.970Z",
          "content": "<p>Are you talking about using kernels for training or for inference? There are different limits for training (the default Kaggle-wide limits on kernels <a href=\"https://www.kaggle.com/docs/kernels#technical-specifications\">https://www.kaggle.com/docs/kernels#technical-specifications</a>) and inference (the challenge-specific limits listed at <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements\">https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements</a>). </p>\n\n<p>You can do whatever you like if you are just training in a kernel, subject to the Kaggle-wide kernel limits. If those kernel limits are too restrictive for your particular training, then you are welcome to train on your local machine without using kernels, and use kernels only for inference when making your submission.</p>\n\n<p>We expect the inference kernel limits (4 hr on CPU or 1 hr on GPU) should be more than enough for inference on the test set (public and private).</p>",
          "rawMarkdown": "Are you talking about using kernels for training or for inference? There are different limits for training (the default Kaggle-wide limits on kernels https://www.kaggle.com/docs/kernels#technical-specifications) and inference (the challenge-specific limits listed at https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements). \n\nYou can do whatever you like if you are just training in a kernel, subject to the Kaggle-wide kernel limits. If those kernel limits are too restrictive for your particular training, then you are welcome to train on your local machine without using kernels, and use kernels only for inference when making your submission.\n\nWe expect the inference kernel limits (4 hr on CPU or 1 hr on GPU) should be more than enough for inference on the test set (public and private)."
        },
        {
          "id": 512245,
          "postDate": "2019-04-10T16:38:55.927Z",
          "content": "<p>If I understand it correctly, we can make some pre-procession in some kernel (or locally), and than use the output pre-processed sample to train our models within the introduced limits for Competition Kernels, right? </p>",
          "rawMarkdown": "If I understand it correctly, we can make some pre-procession in some kernel (or locally), and than use the output pre-processed sample to train our models within the introduced limits for Competition Kernels, right? "
        },
        {
          "id": 512263,
          "postDate": "2019-04-10T16:51:57.220Z",
          "content": "<p>You are free to make pre-processed versions of the data for your training or inference in kernels.</p>\n\n<p>Again, as I said in my previous message, this challenge only imposes limits on kernels where you run inference to make submissions. We have not placed any limits on kernels where you do training, where you can do what you want, subject to the general limits placed by Kaggle on all kernels (I think that is curently around 9 hours of run time, please see  <a href=\"https://www.kaggle.com/docs/kernels#technical-specifications\">https://www.kaggle.com/docs/kernels#technical-specifications</a>).</p>",
          "rawMarkdown": "You are free to make pre-processed versions of the data for your training or inference in kernels.\n\nAgain, as I said in my previous message, this challenge only imposes limits on kernels where you run inference to make submissions. We have not placed any limits on kernels where you do training, where you can do what you want, subject to the general limits placed by Kaggle on all kernels (I think that is curently around 9 hours of run time, please see  https://www.kaggle.com/docs/kernels#technical-specifications)."
        },
        {
          "id": 512318,
          "postDate": "2019-04-10T17:25:28.753Z",
          "content": "<p>Thank you for clarification! </p>",
          "rawMarkdown": "Thank you for clarification! "
        },
        {
          "id": 544773,
          "postDate": "2019-06-05T22:35:11.857Z",
          "content": "<p>Okay, I'm still confused? Are there instructions on how to \"upload your [not pretrained] model as a Kaggle dataset\"? </p>",
          "rawMarkdown": "Okay, I'm still confused? Are there instructions on how to \"upload your [not pretrained] model as a Kaggle dataset\"? "
        }
      ]
    },
    {
      "id": 524227,
      "postDate": "2019-04-28T08:49:58.023Z",
      "content": "<p>Thank you for clarification</p>",
      "rawMarkdown": "Thank you for clarification"
    }
  ],
  "comments": [
    {
      "id": 508290,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2019-04-05T23:37:19.103000",
      "content": "<p>I have the same question as <a href=\"/daisukelab\">@daisukelab</a>, the rules seem to be a bit bizarre to me.</p>\n\n<p>As I understand it, I can upload a pretrained model for inference, however, that model must NOT be pretrained on any other dataset before trained on Freesound data. So the question is, how do you know? How can you be sure that any teams outside the prize range didn't cheat?</p>\n\n<p>Another question is why did you choose to prohibit ImageNet models? I changed <code>pretrained=False</code> to <code>True</code> in FastAI and my validation score almost doubled up, I know there must be reasons behind it but to be honest it seems a bit silly to me that such a simple (and amazing) thing is not allowed.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 508361,
          "author_name": "ChewZY",
          "author_url": "",
          "post_date": "2019-04-06T03:51:44",
          "content": "<p>Yeah it sounds more logical to either allow the use of pre-trained models or make it strictly kernel-based (including pre-processing, training, predicting). </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 509727,
          "author_name": "Frederic Font",
          "author_url": "",
          "post_date": "2019-04-08T07:47:55.610000",
          "content": "<p>We want participants to have the freedom to train and do their experiments offline, outside Kaggle kernels. At the same time we want to encourage not to use external data or pre-trained models because we want participants to focus on the research challenges which are characteristic to our dataset. We believe in participant's honour to avoid cheating. For the winning submissions we'll revise them manually so we'll be able to detect that (code needs to be open sourced).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 509747,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2019-04-08T08:27:04.857000",
          "content": "<p>Thank you for answering my question <a href=\"/fredericfont\">@fredericfont</a>. </p>\n\n<p>Just to be clear, are we allowed to upload multiple models for ensembling, as long as the total inference time does not break the 1 hour limit?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 509758,
          "author_name": "Frederic Font",
          "author_url": "",
          "post_date": "2019-04-08T08:38:52.717000",
          "content": "<p>yes</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 510430,
          "author_name": "qrfaction",
          "author_url": "",
          "post_date": "2019-04-09T04:28:35.427000",
          "content": "<p>How do you verify that the weight uploaded by the competition team derived from the pre-trained models?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 514592,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-04-11T19:11:12.477000",
          "content": "<p>We use a combination of the honor system and code review. Winning submissions will need to give us the code they used to train their models. This worked well for us last year where we placed restrictions on external datasets.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 515865,
          "author_name": "Ahmet Erdem",
          "author_url": "",
          "post_date": "2019-04-13T08:25:11.270000",
          "content": "<p>I think this is a very bad decision. I would either allow usage of all possible data or totally forbid. Since only top3 will be checked, you can be 4th and get a solo gold by violating the rules. Solo gold + high kaggle points is most probably more valuable than money prize of 3rd:) <a href=\"/inversion\">@inversion</a> what do you think?</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 524479,
          "author_name": "James Conley",
          "author_url": "",
          "post_date": "2019-04-28T21:23:00.307000",
          "content": "<p>I've only been actively using Kaggle for a month or two and I already have seen people who cheat by making multiple accounts to test more submissions.  If people would go through all the trouble of making more accounts just for more submissions, they probably already have happily gone through with exactly what Ahmet has described above.  It's really a shame.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 535999,
          "author_name": "Jie Wu",
          "author_url": "",
          "post_date": "2019-05-23T19:09:42.770000",
          "content": "<p><a href=\"/divrikwicky\">@divrikwicky</a> I am on your side. Seems only top3 will be checked if they used pre-trained external data. Teams are not in the top3 will not be checked, this is unfair for everyone.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 536010,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-05-23T19:46:33.807000",
          "content": "<p>Please see <a href=\"/addisonhoward\">@addisonhoward</a> 's message earlier in this thread\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064525061\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064525061</a></p>\n\n<p>The organizers and Kaggle reserve the right to check the submissions of any teams, not just the final winners, and the penalties could include banning of your account from Kaggle entirely. Also note that the final winners are determined by ranking on the private test set, and that will be determined only after the competition ends and may not be the top 3 from the public leaderboard.</p>\n\n<p>So if you try to cheat, you are gambling your Kaggle account on the chance that you will not end up as a winner on the private test set, and that you will not be among the teams that we choose to verify.  I'm not sure if losing your account is worth a few points from one challenge.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 508071,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "2019-04-05T15:14:15.970000",
      "content": "<p>Hi, let me confirm about pre-trained models.\nI'd like to work with ImageNet pre-trained CNN model provided by PyTorch by default.\nIt is NOT ALLOWED, correct?</p>\n\n<p>But <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements\">Kernels Requirements here</a> also shows that:\n- \"External data and pre-trained models are not allowed. If you train your model offline, you may upload your model as a Kaggle dataset for use in creating your inference Kernel.\"</p>\n\n<p>It sounds like we can use pre-trained models from Kaggle datasets, and we have PyTorch models here:\n- <a href=\"https://www.kaggle.com/pvlima/pretrained-pytorch-models\">https://www.kaggle.com/pvlima/pretrained-pytorch-models</a></p>\n\n<p>Can I use <a href=\"https://www.kaggle.com/pvlima/pretrained-pytorch-models\">https://www.kaggle.com/pvlima/pretrained-pytorch-models</a>?</p>\n\n<p>Thank you.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 509731,
          "author_name": "Frederic Font",
          "author_url": "",
          "post_date": "2019-04-08T07:55:29.787000",
          "content": "<p>You can use models that you trained yourself with the data we provide and use them for inference, this is so that you can also work offline and not all training needs to happen in Kernel runtime. Use of models pre-trained with external data is forbidden.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 509778,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-04-08T09:08:22.173000",
          "content": "<p>Thank you, it's clear now :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 525061,
      "author_name": "Addison Howard",
      "author_url": "",
      "post_date": "2019-04-30T05:58:40.053000",
      "content": "<p>Hi All,</p>\n\n<p>Just a reminder that Kaggle and the Freesound team reserve the right to review all kernels, not just those that win the competition, for adherence to the competition rules. Should it be determined that a participant intentioned tried to \"fly under the radar\" by finishing high enough for strong Kaggle points/medals, but hoping to not have their code reviewed, we reserve the right to not only disqualify that team from the competition, but also to ban the account entirely.</p>\n\n<p>These rules have been the same for all competitions which prohibit external data, and the integrity of the Kaggle community has resulted in only a few violations to this effect.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 544763,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-06-05T22:11:46.870000",
          "content": "<p>Hi <a href=\"/addisonhoward\">@addisonhoward</a> , is it true that a team using pre-trained models will be banned from Kaggle? I've seem banned from competition, but from Kaggle I believe its the first time. </p>\n\n<p>Our solution uses only non-pretrained models, but we checked and using pre-trained models scores much higher in LB. I just want to be sure that all teams using pre-trained will be dropped from the LB. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 544842,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-06-06T00:25:46.877000",
          "content": "<p>If we determine that a team has violated the rules in such a way as to obtain a high ranking without having to disclose not high enough to have to disclose their model (thus breaking the rules but hoping not to get caught), we reserve the right to ban them from the site at our discretion</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 546347,
          "author_name": "Blonde",
          "author_url": "",
          "post_date": "2019-06-06T13:59:51.177000",
          "content": "<p>They'll create another account ). Kaggle is addictive, and we all know that, so the ban will not stop people from kaggling, imho ) (I am not in this competition, just saw it occasionally)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 546358,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-06T14:05:33.600000",
          "content": "<p>Very happy to see rules enforcement beyond prize winners.  We all can think of cases where there was a high suspicion of private sharing or external data use, but they flew under the radar.  If kernel only help in that matter then I'm way more interested in kernel only competition than before!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 546399,
          "author_name": "Giba",
          "author_url": "",
          "post_date": "2019-06-06T14:41:16.937000",
          "content": "<p>Thank you <a href=\"/addisonhoward\">@addisonhoward</a> . Hope this encourage teams not to use pretrained models, testset or any other external data :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 527084,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "2019-05-04T14:48:12.070000",
      "content": "<p>Hi, is it possible to clarify if we can use  “Utility Script On” or not which is introduced in this discussion?\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91241#latest-526152\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91241#latest-526152</a>\nThanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 527545,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-05-05T19:13:01.197000",
          "content": "<p>I believe it should be fine to import any code you like. <a href=\"/addisonhoward\">@addisonhoward</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527617,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-05-06T00:23:02.997000",
          "content": "<p>You can import code to train, but remember that the inference kernel must run standalone, and not be a part of linking to multiple other Kernels to try and circumnavigate the single inference kernel for scoring.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527619,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-05-06T00:29:59.500000",
          "content": "<p>Thank you for clarification, I understood that final submission kernel CANNOT import code.\nI will keep final kernel to be standalone.\n(I was hoping not to copy&amp;paste basic library code)\nThanks again, it's clear.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527646,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-05-06T02:27:48.040000",
          "content": "<p>I think you can import code just fine. Addison was pointing out that you should not try and link kernels in such a way that other kernels also run when we run the inference kernel.</p>\n\n<p>Note that in addition to using the new Kaggle feature of utility scripts, you can also import code by simply including the code in a Dataset (along with your pre-trained model) and attaching the Dataset to the kernel.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 527647,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-05-06T02:32:13.843000",
          "content": "<p>Thanks Manoj - that is the correct interpretation</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527671,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-05-06T04:00:07.770000",
          "content": "<p>Thank you <a href=\"/plakal\">@plakal</a> <a href=\"/addisonhoward\">@addisonhoward</a> for correction, I appreciate your help!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 524089,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-04-27T22:20:03.193000",
      "content": "<p>Could you please be more specific on <em>The test set cannot be used to train the submitted system.</em>? Will unsupervised methods such as domain adaptation methods be considered as the rules violation?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 524949,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-04-29T20:47:33.687000",
          "content": "<p>I think <a href=\"/fredericfont\">@fredericfont</a> was fairly clear in his post: you cannot use the test set in any way as an input in the training of your submission.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 524951,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-04-29T20:52:53.593000",
          "content": "<p>Thank you for clarification :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 546359,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-06-06T14:05:58.937000",
          "content": "<p>I wish that was made clear in LANL competition as well!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 535353,
      "author_name": "Maxwell",
      "author_url": "",
      "post_date": "2019-05-22T18:32:17.640000",
      "content": "<p>Hi,  I have one question about a kernel.  </p>\n\n<p>As everyone confirmed, we are allowed to upload our models as private external data. When we attached our external data to the inference scripts, directory environment will change. More specifically, in the default environment, sample_submission.csv and other related csv files will be just under the <code>input</code> folder. But after we added external data, the environment will be changed, and csv files will be under <code>input/freesound-audio-tagging-2019</code> folder.  </p>\n\n<p>I am now writing my kernel script to refer the changed folder location when reading csv and other audio files. <br>\nIs it safe to use such a hard-coded directory name (e.g. <code>input/freesound-audio-tagging-2019</code>) in 2nd stage ?  </p>\n\n<p>Thanks in advance.  </p>",
      "votes": 2,
      "replies": [
        {
          "id": 535415,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "2019-05-22T21:22:33.823000",
          "content": "<p>You could get into trouble if you import a dataset into your script that later gets deleted. If you are sure you won't have that issue, you could hard code paths.</p>\n\n<p>A safer alternative might be something like:</p>\n\n<p><code>\ntry:\n &lt;read long path&gt;\nexcept:\n &lt;read short path&gt;\n</code></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 535427,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-05-22T22:36:21.857000",
          "content": "<p><a href=\"/inversion\">@inversion</a> \nThank you for your reply. <br>\nI understand. Actually I faced that situation in another kernel, it wired me...  </p>\n\n<p>So as long as I do NOT delete an external dataset attached to the inference kernel, hard coding the paths will not be a trouble, right ?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 533357,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2019-05-19T04:18:20.780000",
      "content": "<p>Hi, let me make sure the rule about pre-trained model.\nWe are allowed to train a model beforehand with the data set in this competition, but not with external data.\nAnd we are also not allowed to use pre-trained model even if it's included in some frameworks.</p>\n\n<p>I want to clarify what will happen after finishing the competition.\nThe codes of prize winners will be checked their reproducibility, but how about other medal winners?\nWill all of medal winner's codes be also reviewed whether they keep the rule or not?</p>\n\n<p>I think it should be so but it seems difficult to check over 100 teams' code.\nBut if not checked, that means someone against the rule can win the competition secretly.</p>\n\n<p>I just want to make sure so that every one can be fair and enjoy the competition!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 533733,
          "author_name": "Eduardo Fonseca",
          "author_url": "",
          "post_date": "2019-05-19T20:38:18.197000",
          "content": "<p>Hi, you are correct: external data is not allowed (including models pre-trained on external data).</p>\n\n<p>I understand your concern. I will refer to this post by <a href=\"/addisonhoward\">@addisonhoward</a> \n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#525061\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#525061</a></p>\n\n<p>I will add that we like to trust participants' honor code. Besides, infringing the rules comes accompanied with the risk of being disqualified from the competition, and even being banned from Kaggle. It seems a high price to pay for a few Kaggle points/medals.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 533773,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2019-05-20T02:09:09.857000",
          "content": "<p>OK, thanks! I was relieved to hear that!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 530806,
      "author_name": "OsciiArt",
      "author_url": "",
      "post_date": "2019-05-13T18:03:30.990000",
      "content": "<p>Will the sample_submission.csv update in stage 2? I'd like to know it because I use it as a reference for file path to test data.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 530815,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-05-13T18:28:41.297000",
          "content": "<p>Correct - we'll be updating the sample submission file for stage 2</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 531343,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2019-05-14T17:27:27.353000",
          "content": "<p>Thank you for your clarification!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 549520,
      "author_name": "Josh Varty",
      "author_url": "",
      "post_date": "2019-06-10T18:33:39.490000",
      "content": "<p>I have a question about submitting our kernels. I have a kernel that I have been working on and have different scores on as I've made changes. On the submissions page I can select two scores/submissions that will be used as my final leaderboard score.</p>\n\n<p>If I select two scores from the same kernel (but different versions with different public scores) will everything work properly? As in: Does it automatically use the latest version of a kernel, or will it use the version and environment that produced the selected public LB score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 549599,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-06-10T21:04:18.840000",
          "content": "<p>You can select a submission based on a specific kernel version - it will not automatically use the latest version of the kernel.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 548931,
      "author_name": "uratatsu",
      "author_url": "",
      "post_date": "2019-06-10T05:41:30.723000",
      "content": "<p>Hi, I have one question about a 2nd stage data .</p>\n\n<p>Is the fname of sample_submission.csv updated in 2nd stages guaranteed to be unique?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 546344,
      "author_name": "Phil",
      "author_url": "",
      "post_date": "2019-06-06T13:52:42.583000",
      "content": "<p>The second part of this statement about the kernel requirements is confusing. \"External data and <strong>pre-trained models</strong> are not allowed. If you train your model offline, you may upload your model as a Kaggle dataset for use in creating your inference Kernel.\" What does \"upload your model as a Kaggle dataset\" mean? What's the difference between a \"model\" and a \"<strong>pre-trained</strong> model\"? Does it mean you can upload a model description, like a json file? That seems rather useless. Does it mean we can upload pre-processed data? That would save tremendously on pre-training time, since it takes 1 hour to read in 1000 training files, 90 minutes for 1120 test files. From reading through the comments here, it doesn't seem that anybody really understands what it means. Please clarify.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 546492,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-06-06T16:10:04.293000",
          "content": "<p>Pre-trained models in this case specifically refer to models that were trained by someone else. For example, in image classification, it is very common to use ImageNet as a pretrained model and to perform additional tuning from there. In this case, no pre-trained models are allowed. You can, however train your own model and upload it to a Kernel.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 546556,
          "author_name": "Phil",
          "author_url": "",
          "post_date": "2019-06-06T17:19:50.423000",
          "content": "<p>Reallly!!? Cool. So, how do you do that? I think I attempted to attach a dataset to a kernel for this competition and the kernel failed to produce a submission.csv file. I figured it was because I was reaching for \"outside\" data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543034,
      "author_name": "Kany",
      "author_url": "",
      "post_date": "2019-06-04T09:36:45.737000",
      "content": "<p>Hi, I have a question about the size of the second-stage test set.\nDoes it mean the \"total duration\" of the second-stage test set is approximately three times the first-stage?\nOr it means the \"total number of clips\" of the second-stage test set is approximately three times the first-stage?</p>\n\n<p>Thank you</p>",
      "votes": 0,
      "replies": [
        {
          "id": 543662,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-06-04T17:18:08.950000",
          "content": "<p>Hi Kai,</p>\n\n<p>The second-stage test set is based on the number of samples. The distributions of length and size should be approximately the same.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 542753,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2019-06-04T04:33:23.020000",
      "content": "<p>I have a question regarding to this rule </p>\n\n<blockquote>\n  <p>Participants are not allowed to make subjective judgement​s of the test data, nor to annotate it (this includes the use of statistics about the evaluation dataset in the decision making). The test set cannot be used to train the submitted system. </p>\n</blockquote>\n\n<p>I also got a similar answer from you bellow: </p>\n\n<blockquote>\n  <p>Hi, we mean that using a <strong>priori stats of the public test set to make decisions</strong> over the private test set is not allowed. Specifically, answering your question: weighting factors to tune the distribution of classes in the answer set are not allowed. The class label distribution in the test set ensures a minimum number of labels per class, but the distribution is a bit imbalanced.</p>\n</blockquote>\n\n<p>My addition question is: \nIf I use a postprocessing method to make the distribution of the final submission as same as <strong>train set</strong>. Is it allowed ? </p>\n\n<p>Thank you </p>",
      "votes": 0,
      "replies": [
        {
          "id": 543976,
          "author_name": "Eduardo Fonseca",
          "author_url": "",
          "post_date": "2019-06-05T02:11:55.817000",
          "content": "<p>If it is about the train set, I think it is allowed. But, assuming you mean distribution of class labels, note that the train set is quasi balanced and the private test set might not be that way...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 540135,
      "author_name": "Granddad",
      "author_url": "",
      "post_date": "2019-05-31T04:45:47.310000",
      "content": "<p>Which of these situations would count as \"training on the test set\"?\n1)  I decide on an architecture because it does well on the test set.\n2)  I average my two best (public score) models together . <br>\n3) Two teams merge and average their individual best (public score) models.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 540153,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2019-05-31T05:28:04.293000",
          "content": "<p>None</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 540540,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-05-31T16:02:15.867000",
          "content": "<p>You are free to use your public leaderboard scores to make decisions about your models, including architecture, hyperparameters, etc.</p>\n\n<p>What we disallow is directly feeding the test data into the training pipeline of your model, whether as unlabeled data (for unsupervised/semi-supervised techniques, or just for computing statistics), or as labeled data (where you label it yourself).</p>\n\n<p>At the end of the contest, it should be possible to train your submitted model from scratch using the training data we have given you, and no other data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 534681,
      "author_name": "Pavel Mazaev",
      "author_url": "",
      "post_date": "2019-05-21T17:13:35.887000",
      "content": "<p>Will it be possible to make a submission to be evaluated (but, of course, not rated) after the end of the competition? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 534683,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-05-21T17:16:30.720000",
          "content": "<p>Yes - you will be able to make a submission after the competition has concluded, in which you will receive a score, but will not be shown on the leaderboard.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 533426,
      "author_name": "Gavin Bao",
      "author_url": "",
      "post_date": "2019-05-19T07:54:41.793000",
      "content": "<p>Hi, I have two questions.</p>\n\n<p>In the second stage, it is you to rerun our kernel? Not ourselves? So how we upload the p retrained weights as a dataset so that you could find it? If we upload them as a private dataset, add them in the kernel, will you be able to access them?</p>\n\n<p>Also, the commit button in kaggle takes me a lot of time usually. Before seeing the output file, it will stay in that page for a long time…. will that time be counted as run time in the second stage?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 531326,
      "author_name": "duvallwh",
      "author_url": "",
      "post_date": "2019-05-14T16:45:45.473000",
      "content": "<p>Can we use the pre-trained melspectrogram script for our final submission? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 533717,
          "author_name": "Eduardo Fonseca",
          "author_url": "",
          "post_date": "2019-05-19T19:56:18.263000",
          "content": "<p>Hi, not sure I understand what you mean by <em>pre-trained melspectrogram script</em>. Could you please clarify?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 528306,
      "author_name": "Dimosthenis Karaflos",
      "author_url": "",
      "post_date": "2019-05-07T13:05:34.100000",
      "content": "<p>Hello all,\non the competition rules, it is stated that the locally or kernel trained models should be uploaded as a kaggle dataset. In case that i want to fit a model on a kernel, can i just import the models as a kernel output, on the final submission kernel, or i will have to download them locally and then upload them as a kaggle dataset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 529426,
          "author_name": "Eduardo Fonseca",
          "author_url": "",
          "post_date": "2019-05-09T21:47:50.293000",
          "content": "<p>Hi, \nbased on this response from <a href=\"/addisonhoward\">@addisonhoward</a> I think you must run the inference kernel standalone, and not link it to other Kernels.</p>\n\n<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#527617\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/88064#527617</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 529433,
          "author_name": "Dimosthenis Karaflos",
          "author_url": "",
          "post_date": "2019-05-09T22:09:42.470000",
          "content": "<p>Thank you for your response\nSo, in order to be sure that any error will not ocure, i will have to make a kernel with my uploaded weights as a kaggle dataset and i will have to read the test data from the directory:\n'../input/freesound-audio-tagging-2019/test/xxxxxx.wav'</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 527362,
      "author_name": "Datasciensyash",
      "author_url": "",
      "post_date": "2019-05-05T08:20:32.800000",
      "content": "<p>Hi,\nWill the maximum kernel run time of the second stage of the competition change? As the size of the test data increases, the time taken to process these data will increase and I am not sure that it will be easy to keep up with 60 minutes on GPU...</p>\n\n<p>Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 527546,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-05-05T19:14:23.597000",
          "content": "<p>The run-time limits will not change. You will need to make sure that your submission kernel can run within the limits when we pass in a private test set that is around 3 times the size of the public test set.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 525396,
      "author_name": "DavidS",
      "author_url": "",
      "post_date": "2019-04-30T22:06:18.953000",
      "content": "<p>I have a problem. I use kernel named <em>submission</em> that uses the output from kernel <em>training</em>. Training kernel uses the output from <em>features</em>. In this pipeline, I'm not able to use <em>submission</em> kernel because of the error:</p>\n\n<blockquote>\n  <p>your input kernel [training] cannot use kernel outputs as a data source for this competition</p>\n</blockquote>\n\n<p>I think it's very disturbing. The architecture of the pipeline I've chosen has to be rewritten. Did anyone have the same problem? Is it necessary to block such data flow? Was it mentioned anywhere here?</p>\n\n<p>I can apparently add only one kernel depth to submission kernel.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 525438,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2019-05-01T01:47:41.303000",
          "content": "<p>One alternative is to use the outputs from your features &amp; training kernels as a Dataset (it can be private) and then just attach the dataset to the kernel that you want to use those files in. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 525456,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-05-01T04:03:47.527000",
          "content": "<p>I also recommend to do as James wrote. We could minimize dependency to the complex Kaggle system. System can be down anytime...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 522813,
      "author_name": "DungNB",
      "author_url": "",
      "post_date": "2019-04-25T04:31:43.830000",
      "content": "<p>Hi Frederic,\nAs you said \"the size of the second-stage test set is approximately three times the size of the first\", maybe 3.3 or 2.7.\nBecause of the limited time in second-stage, how many samples in the second-stage test set?\nThanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 522948,
          "author_name": "Frederic Font",
          "author_url": "",
          "post_date": "2019-04-25T09:16:26.663000",
          "content": "<p>Hi, I can't tell you the exact number, the 3x guideline should be enough as is it quite realistic.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522958,
          "author_name": "DungNB",
          "author_url": "",
          "post_date": "2019-04-25T09:34:04.477000",
          "content": "<p>Thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 522102,
      "author_name": "Max",
      "author_url": "",
      "post_date": "2019-04-23T21:31:07.253000",
      "content": "<p>Hi,\nThe competition rules state that <code>No custom packages enabled in kernels</code>.\nI keep all of my training/preprocessing/inference code in a GitHub repository, does this mean that my only option is to manually copy-paste all my codebase each time I make a submission? </p>\n\n<p>Also, just out of curiosity, why aren't we allowed to use custom packages? (:</p>\n\n<p>Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 522248,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-04-24T05:09:06.363000",
          "content": "<p>You should be able to upload all your (Python) code as well as your trained model checkpoint into a Kaggle Dataset which you can attach to your kernel for submission.</p>\n\n<p><a href=\"/addisonhoward\">@addisonhoward</a> for the motivation behind the custom packages restriction.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 522655,
          "author_name": "Max",
          "author_url": "",
          "post_date": "2019-04-24T19:14:36.847000",
          "content": "<p>Great! Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 519847,
      "author_name": "Costas Voglis",
      "author_url": "",
      "post_date": "2019-04-19T18:18:58.183000",
      "content": "<p>Hello to all and thank the organizers for an interesting competition...</p>\n\n<p>Regarding the second phase of this competition : \n1. Is the second phase inference going to be automatic like the Petfinder competition or are we going to manually rerun our best performing kernels? What happens if we exceed the 1hour runtime. Are we automatically off or we will be given a second chance to use a \"lighter\" version of the kernel?</p>\n\n<ol>\n<li>In the second case, are we allowed to preprocess eg. extract spectrogram features from the test set before the final submission? This way we may save a couple of minutes from the runtime. On the other hand test audio preprocessing is considered part of the inference.. Just want to make sure I got it right..</li>\n</ol>\n\n<p>I hope I made my self clear....</p>",
      "votes": 0,
      "replies": [
        {
          "id": 521351,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-04-22T20:18:41.677000",
          "content": "<ol>\n<li><p>We will run all submitted kernels on the private test set after the competition ends on June 10. Kernels that hit the limits described at <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements\">https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements</a> will be disqualified.  There will be no second chances, so you should submit a kernel that can scale to a 3x larger  data set.</p></li>\n<li><p>I'm not sure I follow because we will be running your submitted kernel on an unseen private test set so no amount of preprocessing using the public test set is going to help you there. If your question is about separating preprocessing time from inference time in general, then we do not support any such separation. The kernel time limits are for all computation that you need to perform on the test set, including preprocessing and inference.</p></li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 521391,
          "author_name": "Costas Voglis",
          "author_url": "",
          "post_date": "2019-04-22T21:13:21.987000",
          "content": "<p>Perfect.. thank you very much.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 516225,
      "author_name": "James Lee",
      "author_url": "",
      "post_date": "2019-04-13T19:57:44.380000",
      "content": "<p>Hi,\nThis would be my first competition, and I'm wondering if someone wouldn't mind answering my newbie questions.\n1. What does it mean to use kernels only for inference?\n2. Can I preprocess the files and save features like mfcc in advance and upload to the Kernel to use later during the evaluation?\n3. Can I train the model in advance using the provided files, and upload the model weights to the kernel to use to predict?\n4. When my model gets evaluated with the private test set, how do I make sure my kernel can process the private test set? Would the files in test folder and sample_submission.csv be automatically swapped with the private data set?\nThanks for your help!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 518145,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-04-16T22:59:34.503000",
          "content": "<p>Hi James,</p>\n\n<p>1) Inference only means that your kernel only needs to make predictions, it does not need to also train the data. If you win the competition, you will also need to provide your full solution, but to submit to the competition, you only need the inference kernel.\n2-3) you can train/preprocess in another kernel or on your local machine. You can then upload those weights or other features to your inference kernel.\n4) The second-stage test set is approximately <strong>three times the size of the first</strong>. You should plan your kernel's memory, disk, and runtime footprint accordingly.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 515536,
      "author_name": "Dmytro Danevskyi",
      "author_url": "",
      "post_date": "2019-04-12T18:30:47.730000",
      "content": "<p>Hello, </p>\n\n<p>Are we allowed to manually label train data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 517734,
          "author_name": "Frederic Font",
          "author_url": "",
          "post_date": "2019-04-16T12:48:09.333000",
          "content": "<p>Yes you can experiment with labelling <em>train set data</em>, but never the test set. In any case remember that the goal of the task is precisely how to solve the label noise problem through <strong>automatic</strong> methods. Manually replacing noisy labels with better ones is therefore not a solution to the problem.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 553900,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2019-06-16T14:57:16.653000",
          "content": "<p><a href=\"/fredericfont\">@fredericfont</a> Our team has never thought manual relabeling. Also, we didn't remove the six corrupted files because the robustness to noisy data is the aim of this competition. But now, it revealed manual relabeling is a very effective strategy in this competition. It's not matched to the concept of this competition that asking training with limited data. Relabeling should have been considered as external data, I think. Even if a competition that permits external data, hand relabeling is considered as external data and required to share in public. This is a very strange rule. I'm sorry for I didn't notice this before this competition ends.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 554101,
          "author_name": "iwashi",
          "author_url": "",
          "post_date": "2019-06-17T01:34:03.940000",
          "content": "<p><a href=\"/fredericfont\">@fredericfont</a> I agree with OsciiArt. As we all know, the concept of this competition is to create a model from small curated data and large noisy data. The essence of this concept is to create a robust model without making a lot of curated data. I think allowing relabeling is against this concept.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 554586,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-06-17T17:36:34.547000",
          "content": "<p>Can you elaborate on this revelation that manual relabeling is a very effective strategy in this competition? Are you talking about a specific team? Also I'm curious about how you know the strategy to be effective before the private leaderboard has been released.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 554602,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2019-06-17T18:07:40.607000",
          "content": "<p><a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95924</a>\n<a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95785\">https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/95785</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 554608,
          "author_name": "Manoj Plakal",
          "author_url": "",
          "post_date": "2019-06-17T18:20:31.650000",
          "content": "<p>My suggestion is to wait for the private leaderboard :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 564175,
          "author_name": "OsciiArt",
          "author_url": "",
          "post_date": "2019-06-29T04:48:02.467000",
          "content": "<p>I saw the final result. I strongly hope the organizer considers changing the rule next year.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 564263,
          "author_name": "James Requa",
          "author_url": "",
          "post_date": "2019-06-29T07:59:11.337000",
          "content": "<p>Agreed it defeats the purpose of the competition to manually relabel the training set</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 514284,
      "author_name": "Cocoxili",
      "author_url": "",
      "post_date": "2019-04-11T14:51:19.570000",
      "content": "<p>Hi Frederic. \nI see there are 1120 samples in the current test set. So, how many test samples in the stage 2? Would all the test sample of stage 2  also come from Freesound Dataset (FSD) source?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 514316,
          "author_name": "Frederic Font",
          "author_url": "",
          "post_date": "2019-04-11T15:07:59.817000",
          "content": "<p>Hi, as explained in the <a href=\"https://www.kaggle.com/c/freesound-audio-tagging-2019/data\">data section</a> of the competition page, the contents of the test set consist of manually-labeled data from FSD, and the size of the second-stage test set is approximately three times the size of the first.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 514323,
          "author_name": "Cocoxili",
          "author_url": "",
          "post_date": "2019-04-11T15:12:04.623000",
          "content": "<p>Got it. Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 511862,
      "author_name": "kenmatsu4",
      "author_url": "",
      "post_date": "2019-04-10T11:48:44.167000",
      "content": "<p>Just in case, I'd like to confirm whether the trained weight of my model should be <strong>public</strong> on Kaggle dataset, if I decide to train my model on local machine? Or it can be kept <strong>private</strong>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 512222,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-04-10T16:31:28.070000",
          "content": "<p>Hi Kenmatsu,</p>\n\n<p>You do not need to make your own trained weights public. But you can not use external data to train your weights :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1402079,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-27T20:32:39.173000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 543714,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T18:03:21.217000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 543744,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2019-06-04T18:54:37.293000",
          "content": "<p>Sorry - the rules acceptance deadline has passed. Once the competition concludes, you can make a \"late submission\" where you will receive a score, but the leaderboard will not update.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 543698,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T17:53:55.637000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 542955,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-04T08:21:21.833000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 537074,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-26T05:59:27.847000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 532719,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-17T15:21:09.043000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 533728,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-19T20:21:43.057000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 509649,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-08T05:00:09.740000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 510209,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-08T20:22:15.970000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 512245,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-10T16:38:55.927000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 512263,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-10T16:51:57.220000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 512318,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-10T17:25:28.753000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 544773,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-06-05T22:35:11.857000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 524227,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-28T08:49:58.023000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "508005": "Hi,\n\nMake sure that you check the **competition rules** before participating. The complete rules are explained in the rules section:  https://www.kaggle.com/c/freesound-audio-tagging-2019/rules\nThis is a summary of the most important rules:\n\n- Unlike last year's edition of this task, participants are **not allowed to use external data** for system development. This also excludes the use of pre-trained models.\n- Participants are **not allowed to make subjective judgement​s** of the test data, nor to annotate it (this includes the use of statistics about the evaluation dataset in the decision making). The test set cannot be used to train the submitted system.\n- The winning teams are **required to publish their systems under an open-source license** in order to be considered winners. See [this thread](https://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/93054#latest-535382) for more info.​\n\nPlease post any questions and doubts you might have about competition rules in this discussion thread.",
    "508290": "I have the same question as @daisukelab, the rules seem to be a bit bizarre to me.\n\nAs I understand it, I can upload a pretrained model for inference, however, that model must NOT be pretrained on any other dataset before trained on Freesound data. So the question is, how do you know? How can you be sure that any teams outside the prize range didn't cheat?\n\nAnother question is why did you choose to prohibit ImageNet models? I changed ```pretrained=False``` to ```True``` in FastAI and my validation score almost doubled up, I know there must be reasons behind it but to be honest it seems a bit silly to me that such a simple (and amazing) thing is not allowed.",
    "508071": "Hi, let me confirm about pre-trained models.\nI'd like to work with ImageNet pre-trained CNN model provided by PyTorch by default.\nIt is NOT ALLOWED, correct?\n\nBut [Kernels Requirements here](https://www.kaggle.com/c/freesound-audio-tagging-2019/overview/kernels-requirements) also shows that:\n- \"External data and pre-trained models are not allowed. If you train your model offline, you may upload your model as a Kaggle dataset for use in creating your inference Kernel.\"\n\nIt sounds like we can use pre-trained models from Kaggle datasets, and we have PyTorch models here:\n- https://www.kaggle.com/pvlima/pretrained-pytorch-models\n\nCan I use https://www.kaggle.com/pvlima/pretrained-pytorch-models?\n\nThank you.",
    "525061": "Hi All,\n\nJust a reminder that Kaggle and the Freesound team reserve the right to review all kernels, not just those that win the competition, for adherence to the competition rules. Should it be determined that a participant intentioned tried to \"fly under the radar\" by finishing high enough for strong Kaggle points/medals, but hoping to not have their code reviewed, we reserve the right to not only disqualify that team from the competition, but also to ban the account entirely.\n\nThese rules have been the same for all competitions which prohibit external data, and the integrity of the Kaggle community has resulted in only a few violations to this effect.",
    "527084": "Hi, is it possible to clarify if we can use  “Utility Script On” or not which is introduced in this discussion?\nhttps://www.kaggle.com/c/freesound-audio-tagging-2019/discussion/91241#latest-526152\nThanks!",
    "524089": "Could you please be more specific on *The test set cannot be used to train the submitted system.*? Will unsupervised methods such as domain adaptation methods be considered as the rules violation?",
    "535353": "Hi,  I have one question about a kernel.  \n  \nAs everyone confirmed, we are allowed to upload our models as private external data. When we attached our external data to the inference scripts, directory environment will change. More specifically, in the default environment, sample_submission.csv and other related csv files will be just under the `input` folder. But after we added external data, the environment will be changed, and csv files will be under `input/freesound-audio-tagging-2019` folder.  \n  \nI am now writing my kernel script to refer the changed folder location when reading csv and other audio files.  \nIs it safe to use such a hard-coded directory name (e.g. `input/freesound-audio-tagging-2019`) in 2nd stage ?  \n  \nThanks in advance.  ",
    "533357": "Hi, let me make sure the rule about pre-trained model.\nWe are allowed to train a model beforehand with the data set in this competition, but not with external data.\nAnd we are also not allowed to use pre-trained model even if it's included in some frameworks.\n\nI want to clarify what will happen after finishing the competition.\nThe codes of prize winners will be checked their reproducibility, but how about other medal winners?\nWill all of medal winner's codes be also reviewed whether they keep the rule or not?\n\nI think it should be so but it seems difficult to check over 100 teams' code.\nBut if not checked, that means someone against the rule can win the competition secretly.\n\nI just want to make sure so that every one can be fair and enjoy the competition!",
    "530806": "Will the sample_submission.csv update in stage 2? I'd like to know it because I use it as a reference for file path to test data.",
    "549520": "I have a question about submitting our kernels. I have a kernel that I have been working on and have different scores on as I've made changes. On the submissions page I can select two scores/submissions that will be used as my final leaderboard score.\n\nIf I select two scores from the same kernel (but different versions with different public scores) will everything work properly? As in: Does it automatically use the latest version of a kernel, or will it use the version and environment that produced the selected public LB score?",
    "548931": "Hi, I have one question about a 2nd stage data .\n\nIs the fname of sample_submission.csv updated in 2nd stages guaranteed to be unique?",
    "546344": "The second part of this statement about the kernel requirements is confusing. \"External data and **pre-trained models** are not allowed. If you train your model offline, you may upload your model as a Kaggle dataset for use in creating your inference Kernel.\" What does \"upload your model as a Kaggle dataset\" mean? What's the difference between a \"model\" and a \"**pre-trained** model\"? Does it mean you can upload a model description, like a json file? That seems rather useless. Does it mean we can upload pre-processed data? That would save tremendously on pre-training time, since it takes 1 hour to read in 1000 training files, 90 minutes for 1120 test files. From reading through the comments here, it doesn't seem that anybody really understands what it means. Please clarify.",
    "543034": "Hi, I have a question about the size of the second-stage test set.\nDoes it mean the \"total duration\" of the second-stage test set is approximately three times the first-stage?\nOr it means the \"total number of clips\" of the second-stage test set is approximately three times the first-stage?\n\nThank you",
    "542753": "I have a question regarding to this rule \n\n&gt; Participants are not allowed to make subjective judgement​s of the test data, nor to annotate it (this includes the use of statistics about the evaluation dataset in the decision making). The test set cannot be used to train the submitted system. \n\nI also got a similar answer from you bellow: \n&gt; Hi, we mean that using a **priori stats of the public test set to make decisions** over the private test set is not allowed. Specifically, answering your question: weighting factors to tune the distribution of classes in the answer set are not allowed. The class label distribution in the test set ensures a minimum number of labels per class, but the distribution is a bit imbalanced.\n\nMy addition question is: \nIf I use a postprocessing method to make the distribution of the final submission as same as **train set**. Is it allowed ? \n\nThank you ",
    "540135": "Which of these situations would count as \"training on the test set\"?\n1)  I decide on an architecture because it does well on the test set.\n2)  I average my two best (public score) models together .  \n3) Two teams merge and average their individual best (public score) models.  ",
    "534681": "Will it be possible to make a submission to be evaluated (but, of course, not rated) after the end of the competition? ",
    "533426": "Hi, I have two questions.\n\nIn the second stage, it is you to rerun our kernel? Not ourselves? So how we upload the p retrained weights as a dataset so that you could find it? If we upload them as a private dataset, add them in the kernel, will you be able to access them?\n\nAlso, the commit button in kaggle takes me a lot of time usually. Before seeing the output file, it will stay in that page for a long time…. will that time be counted as run time in the second stage?",
    "531326": "Can we use the pre-trained melspectrogram script for our final submission? ",
    "528306": "Hello all,\non the competition rules, it is stated that the locally or kernel trained models should be uploaded as a kaggle dataset. In case that i want to fit a model on a kernel, can i just import the models as a kernel output, on the final submission kernel, or i will have to download them locally and then upload them as a kaggle dataset?",
    "527362": "Hi,\nWill the maximum kernel run time of the second stage of the competition change? As the size of the test data increases, the time taken to process these data will increase and I am not sure that it will be easy to keep up with 60 minutes on GPU...\n\nThanks!",
    "525396": "I have a problem. I use kernel named *submission* that uses the output from kernel *training*. Training kernel uses the output from *features*. In this pipeline, I'm not able to use *submission* kernel because of the error:\n\n&gt; your input kernel [training] cannot use kernel outputs as a data source for this competition\n\nI think it's very disturbing. The architecture of the pipeline I've chosen has to be rewritten. Did anyone have the same problem? Is it necessary to block such data flow? Was it mentioned anywhere here?\n\nI can apparently add only one kernel depth to submission kernel.",
    "522813": "Hi Frederic,\nAs you said \"the size of the second-stage test set is approximately three times the size of the first\", maybe 3.3 or 2.7.\nBecause of the limited time in second-stage, how many samples in the second-stage test set?\nThanks",
    "522102": "Hi,\nThe competition rules state that `No custom packages enabled in kernels`.\nI keep all of my training/preprocessing/inference code in a GitHub repository, does this mean that my only option is to manually copy-paste all my codebase each time I make a submission? \n\nAlso, just out of curiosity, why aren't we allowed to use custom packages? (:\n\nThanks",
    "519847": "Hello to all and thank the organizers for an interesting competition...\n\nRegarding the second phase of this competition : \n1. Is the second phase inference going to be automatic like the Petfinder competition or are we going to manually rerun our best performing kernels? What happens if we exceed the 1hour runtime. Are we automatically off or we will be given a second chance to use a \"lighter\" version of the kernel?\n\n2. In the second case, are we allowed to preprocess eg. extract spectrogram features from the test set before the final submission? This way we may save a couple of minutes from the runtime. On the other hand test audio preprocessing is considered part of the inference.. Just want to make sure I got it right..\n\nI hope I made my self clear....\n",
    "516225": "Hi,\nThis would be my first competition, and I'm wondering if someone wouldn't mind answering my newbie questions.\n1. What does it mean to use kernels only for inference?\n2. Can I preprocess the files and save features like mfcc in advance and upload to the Kernel to use later during the evaluation?\n3. Can I train the model in advance using the provided files, and upload the model weights to the kernel to use to predict?\n4. When my model gets evaluated with the private test set, how do I make sure my kernel can process the private test set? Would the files in test folder and sample_submission.csv be automatically swapped with the private data set?\nThanks for your help!",
    "515536": "Hello, \n\nAre we allowed to manually label train data?",
    "514284": "Hi Frederic. \nI see there are 1120 samples in the current test set. So, how many test samples in the stage 2? Would all the test sample of stage 2  also come from Freesound Dataset (FSD) source?",
    "511862": "Just in case, I'd like to confirm whether the trained weight of my model should be **public** on Kaggle dataset, if I decide to train my model on local machine? Or it can be kept **private**?",
    "1402079": "",
    "543714": "",
    "543698": "",
    "542955": "",
    "537074": "",
    "532719": "",
    "509649": "",
    "524227": "Thank you for clarification"
  }
}