{
  "id": 141554,
  "title": "Are we required to do training/inference inside the kernel?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/141554",
  "author_name": "",
  "post_date": "2020-04-06T13:46:59.635544900Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Both the public and private test sets are visible, so one can just do training and even inference elsewhere, or across multiple kernels. Is this allowed?</p>",
  "messages": [
    {
      "id": "799500",
      "postDate": "04/06/2020 13:46:59",
      "content": "<p>Both the public and private test sets are visible, so one can just do training and even inference elsewhere, or across multiple kernels. Is this allowed?</p>",
      "rawMarkdown": "Both the public and private test sets are visible, so one can just do training and even inference elsewhere, or across multiple kernels. Is this allowed?",
      "votes": null
    },
    {
      "id": "799770",
      "postDate": "04/06/2020 18:32:35",
      "content": "<p>No you're not required to do both training &amp; inference in the Kaggle notebook. You could do training offline and load into your inference notebook for submission. The requirement is that the submission.csv be generated as an output from a notebook/kernel, so inference on Kaggle would be required.</p>\n\n<p>Now, if you're interested in the TPU Star prize, then you'll want to use TPUs throughout. You could opt to do this piecemeal with multiple notebooks to take advantage of our integration.</p>",
      "rawMarkdown": "No you're not required to do both training &amp; inference in the Kaggle notebook. You could do training offline and load into your inference notebook for submission. The requirement is that the submission.csv be generated as an output from a notebook/kernel, so inference on Kaggle would be required.\n\nNow, if you're interested in the TPU Star prize, then you'll want to use TPUs throughout. You could opt to do this piecemeal with multiple notebooks to take advantage of our integration.",
      "votes": null
    },
    {
      "id": "800230",
      "postDate": "04/07/2020 07:33:14",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> As this is a code competition with full test data available it is technically possible to just upload submission.csv to a new kernel and sub that to also overcome time limits. I assume this is not allowed? How are you going to check that?</p>",
      "rawMarkdown": "juliaelliott As this is a code competition with full test data available it is technically possible to just upload submission.csv to a new kernel and sub that to also overcome time limits. I assume this is not allowed? How are you going to check that?",
      "votes": null
    },
    {
      "id": "800703",
      "postDate": "04/07/2020 16:56:10",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Good clarification question. In this competition, we haven't set it up as a strict code competition (i.e. we're not re-running code), and as such, what you describe (where you write simple code to load a CSV of predictions and submit from that output) is entirely permitted, particularly if working entirely locally. The submission from a notebook requirement is intended to ease submission out of a notebook, making use of the TPU integration more seamless, rather than serving an enforcement of code function.</p>\n\n<p>However, what is still (and always) <em>not</em> permitted is that these predictions cannot be the result of or incorporate any hand-labeling of the test set. Those in winning standing will have to hand over their code (whether this is code run within Kaggle or locally in any form). And if it's discovered that a violation is committed, the team will be disqualified.</p>",
      "rawMarkdown": "philippsinger Good clarification question. In this competition, we haven't set it up as a strict code competition (i.e. we're not re-running code), and as such, what you describe (where you write simple code to load a CSV of predictions and submit from that output) is entirely permitted, particularly if working entirely locally. The submission from a notebook requirement is intended to ease submission out of a notebook, making use of the TPU integration more seamless, rather than serving an enforcement of code function.\n\nHowever, what is still (and always) *not* permitted is that these predictions cannot be the result of or incorporate any hand-labeling of the test set. Those in winning standing will have to hand over their code (whether this is code run within Kaggle or locally in any form). And if it's discovered that a violation is committed, the team will be disqualified.",
      "votes": null
    },
    {
      "id": "800724",
      "postDate": "04/07/2020 17:18:51",
      "content": "<p>Ok thanks for clarification, so fitting locally and uploading the csv only is allowed.</p>",
      "rawMarkdown": "Ok thanks for clarification, so fitting locally and uploading the csv only is allowed.",
      "votes": null
    },
    {
      "id": "804741",
      "postDate": "04/11/2020 21:52:07",
      "content": "<p>Thanks for the clarification!  Is this information in the 'Overview' section somewhere and I just missed it?  If not, feels like it should be.</p>",
      "rawMarkdown": "Thanks for the clarification!  Is this information in the 'Overview' section somewhere and I just missed it?  If not, feels like it should be.",
      "votes": null
    },
    {
      "id": "804901",
      "postDate": "04/12/2020 05:38:56",
      "content": "<p>Everything is allowed except what is disallowed in the rules. The only relevant rule to this discussion is section #5 of the rules:</p>\n\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>",
      "rawMarkdown": "Everything is allowed except what is disallowed in the rules. The only relevant rule to this discussion is section #5 of the rules:\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 799770,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "04/06/2020 18:32:35",
      "content": "<p>No you're not required to do both training &amp; inference in the Kaggle notebook. You could do training offline and load into your inference notebook for submission. The requirement is that the submission.csv be generated as an output from a notebook/kernel, so inference on Kaggle would be required.</p>\n\n<p>Now, if you're interested in the TPU Star prize, then you'll want to use TPUs throughout. You could opt to do this piecemeal with multiple notebooks to take advantage of our integration.</p>",
      "votes": null,
      "replies": [
        {
          "id": 800230,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/07/2020 07:33:14",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> As this is a code competition with full test data available it is technically possible to just upload submission.csv to a new kernel and sub that to also overcome time limits. I assume this is not allowed? How are you going to check that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800703,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "04/07/2020 16:56:10",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Good clarification question. In this competition, we haven't set it up as a strict code competition (i.e. we're not re-running code), and as such, what you describe (where you write simple code to load a CSV of predictions and submit from that output) is entirely permitted, particularly if working entirely locally. The submission from a notebook requirement is intended to ease submission out of a notebook, making use of the TPU integration more seamless, rather than serving an enforcement of code function.</p>\n\n<p>However, what is still (and always) <em>not</em> permitted is that these predictions cannot be the result of or incorporate any hand-labeling of the test set. Those in winning standing will have to hand over their code (whether this is code run within Kaggle or locally in any form). And if it's discovered that a violation is committed, the team will be disqualified.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800724,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/07/2020 17:18:51",
          "content": "<p>Ok thanks for clarification, so fitting locally and uploading the csv only is allowed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804741,
          "author_name": "yeayates21",
          "author_url": "",
          "post_date": "04/11/2020 21:52:07",
          "content": "<p>Thanks for the clarification!  Is this information in the 'Overview' section somewhere and I just missed it?  If not, feels like it should be.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 804901,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "04/12/2020 05:38:56",
          "content": "<p>Everything is allowed except what is disallowed in the rules. The only relevant rule to this discussion is section #5 of the rules:</p>\n\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "799500": "Both the public and private test sets are visible, so one can just do training and even inference elsewhere, or across multiple kernels. Is this allowed?",
    "799770": "No you're not required to do both training &amp; inference in the Kaggle notebook. You could do training offline and load into your inference notebook for submission. The requirement is that the submission.csv be generated as an output from a notebook/kernel, so inference on Kaggle would be required.\n\nNow, if you're interested in the TPU Star prize, then you'll want to use TPUs throughout. You could opt to do this piecemeal with multiple notebooks to take advantage of our integration.",
    "800230": "juliaelliott As this is a code competition with full test data available it is technically possible to just upload submission.csv to a new kernel and sub that to also overcome time limits. I assume this is not allowed? How are you going to check that?",
    "800703": "philippsinger Good clarification question. In this competition, we haven't set it up as a strict code competition (i.e. we're not re-running code), and as such, what you describe (where you write simple code to load a CSV of predictions and submit from that output) is entirely permitted, particularly if working entirely locally. The submission from a notebook requirement is intended to ease submission out of a notebook, making use of the TPU integration more seamless, rather than serving an enforcement of code function.\n\nHowever, what is still (and always) *not* permitted is that these predictions cannot be the result of or incorporate any hand-labeling of the test set. Those in winning standing will have to hand over their code (whether this is code run within Kaggle or locally in any form). And if it's discovered that a violation is committed, the team will be disqualified.",
    "800724": "Ok thanks for clarification, so fitting locally and uploading the csv only is allowed.",
    "804741": "Thanks for the clarification!  Is this information in the 'Overview' section somewhere and I just missed it?  If not, feels like it should be.",
    "804901": "Everything is allowed except what is disallowed in the rules. The only relevant rule to this discussion is section #5 of the rules:\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records."
  },
  "source": "meta"
}