{
  "id": 147331,
  "title": "Submission Code Kernel Requirements",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/147331",
  "author_name": "",
  "post_date": "2020-04-30T09:12:41.897668900Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hello everyone after recent release of some kernels, I am a bit confused about the competition rules. Does the kernel to be submitted needs to perform training in the code itself? Or we can upload our outputs to a kernel and then submit? which in my opinion would not be fair to everyone?</p>\n\n<p>Since the kernels will not be re - run, my only concern is after the competition ends, will the kernels be checked to see if they are training using the given data, which again seems inplausible. Anyone could please clarify the rules here.</p>",
  "messages": [
    {
      "id": "827398",
      "postDate": "04/30/2020 09:12:41",
      "content": "<p>Hello everyone after recent release of some kernels, I am a bit confused about the competition rules. Does the kernel to be submitted needs to perform training in the code itself? Or we can upload our outputs to a kernel and then submit? which in my opinion would not be fair to everyone?</p>\n\n<p>Since the kernels will not be re - run, my only concern is after the competition ends, will the kernels be checked to see if they are training using the given data, which again seems inplausible. Anyone could please clarify the rules here.</p>",
      "rawMarkdown": "Hello everyone after recent release of some kernels, I am a bit confused about the competition rules. Does the kernel to be submitted needs to perform training in the code itself? Or we can upload our outputs to a kernel and then submit? which in my opinion would not be fair to everyone?\n\nSince the kernels will not be re - run, my only concern is after the competition ends, will the kernels be checked to see if they are training using the given data, which again seems inplausible. Anyone could please clarify the rules here.",
      "votes": null
    },
    {
      "id": "827762",
      "postDate": "04/30/2020 14:24:32",
      "content": "<p>As its a  code competition, the output/submission file must be generated from your Kaggle kernel. That means you can't just upload predictions for submission.</p>\n\n<blockquote>\n  <p>Since the kernels will not be re - run, my only concern is after the competition ends, will the kernels be checked to see if they are training using the given data</p>\n</blockquote>\n\n<p>Anyone from kaggle should clarify on this.</p>",
      "rawMarkdown": "As its a  code competition, the output/submission file must be generated from your Kaggle kernel. That means you can't just upload predictions for submission.\n\n&gt; Since the kernels will not be re - run, my only concern is after the competition ends, will the kernels be checked to see if they are training using the given data\n\nAnyone from kaggle should clarify on this.",
      "votes": null
    },
    {
      "id": "827841",
      "postDate": "04/30/2020 15:17:55",
      "content": "<p>This is wrong, you can just upload csv to kernel and submit from kernel.</p>\n\n<p>Confirmation: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141554#800703\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141554#800703</a></p>",
      "rawMarkdown": "This is wrong, you can just upload csv to kernel and submit from kernel.\n\nConfirmation: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141554#800703",
      "votes": null
    },
    {
      "id": "827850",
      "postDate": "04/30/2020 15:23:44",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a>  Ok then thanks for clarifying, I was going by the standard rules of code competitions.</p>",
      "rawMarkdown": "philippsinger  Ok then thanks for clarifying, I was going by the standard rules of code competitions.",
      "votes": null
    },
    {
      "id": "827872",
      "postDate": "04/30/2020 15:46:43",
      "content": "<p>Well, you are technically correct, but the output is just from reading and dumping the same csv. It is unfortunate that the whole test set can be seen in this competition.</p>",
      "rawMarkdown": "Well, you are technically correct, but the output is just from reading and dumping the same csv. It is unfortunate that the whole test set can be seen in this competition.",
      "votes": null
    },
    {
      "id": "827895",
      "postDate": "04/30/2020 16:04:44",
      "content": "<p>Thanks for helping out <a href=\"/philippsinger\">@philippsinger</a> , I understood the rules, however this is quite dangerous. People aiming for a good rank or in the gold zone, apart from those winning prizes, can still do handlabelling and its not possible for Kaggle to check each and every submission.</p>",
      "rawMarkdown": "Thanks for helping out @philippsinger , I understood the rules, however this is quite dangerous. People aiming for a good rank or in the gold zone, apart from those winning prizes, can still do handlabelling and its not possible for Kaggle to check each and every submission.",
      "votes": null
    },
    {
      "id": "828833",
      "postDate": "05/01/2020 10:22:02",
      "content": "<p><a href=\"/nikhilmishradev\">@nikhilmishradev</a> 100% agree, unfortunately people don't always play fair. The majority does, but black sheep exist. Giving them an easy way to do it is unfortunate.</p>",
      "rawMarkdown": "nikhilmishradev 100% agree, unfortunately people don't always play fair. The majority does, but black sheep exist. Giving them an easy way to do it is unfortunate.",
      "votes": null
    },
    {
      "id": "830448",
      "postDate": "05/02/2020 15:42:05",
      "content": "<p>In general Kaggle prevents   handlabelling in CV competitions by using Two stages data and forcing everyone to upload his code and model at the end of the first stage.  The second stage data is made available only one week before the end.</p>\n\n<p>I don't know why the same thing is not done here ?  (And that would still be compatible with internet use for TPU)\nAnyway, All gold solutions  should be rigorously checked IMHO. \nAnd some random checks in silver/Top 50 may be done too. </p>\n\n<p>Fortunately for this competition,  the test dataset is relatively huge and mixes many languages. So handlabelling may take some efforts ^^</p>",
      "rawMarkdown": "In general Kaggle prevents   handlabelling in CV competitions by using Two stages data and forcing everyone to upload his code and model at the end of the first stage.  The second stage data is made available only one week before the end.\n\n I don't know why the same thing is not done here ?  (And that would still be compatible with internet use for TPU)\nAnyway, All gold solutions  should be rigorously checked IMHO. \nAnd some random checks in silver/Top 50 may be done too. \n\n\nFortunately for this competition,  the test dataset is relatively huge and mixes many languages. So handlabelling may take some efforts ^^",
      "votes": null
    },
    {
      "id": "830465",
      "postDate": "05/02/2020 15:55:44",
      "content": "<blockquote>\n  <p>Fortunately for this competition, the test dataset is relatively huge and mixes many languages. So handlabelling may take some efforts ^^</p>\n</blockquote>\n\n<p>One can simply translate them to common language :)</p>",
      "rawMarkdown": "&gt;Fortunately for this competition, the test dataset is relatively huge and mixes many languages. So handlabelling may take some efforts ^^\n\nOne can simply translate them to common language :)",
      "votes": null
    },
    {
      "id": "830469",
      "postDate": "05/02/2020 16:03:50",
      "content": "<p>You're right. </p>\n\n<p>But these translations may \"detoxify\" the language..and, hopefully, handlabelling will end up with mislabelling on private LB ^^</p>\n\n<p>Anyway , All top solutions need to be checked. </p>",
      "rawMarkdown": "You're right. \n\nBut these translations may \"detoxify\" the language..and, hopefully, handlabelling will end up with mislabelling on private LB ^^\n\nAnyway , All top solutions need to be checked.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 827762,
      "author_name": "stephenmugisha",
      "author_url": "",
      "post_date": "04/30/2020 14:24:32",
      "content": "<p>As its a  code competition, the output/submission file must be generated from your Kaggle kernel. That means you can't just upload predictions for submission.</p>\n\n<blockquote>\n  <p>Since the kernels will not be re - run, my only concern is after the competition ends, will the kernels be checked to see if they are training using the given data</p>\n</blockquote>\n\n<p>Anyone from kaggle should clarify on this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 827841,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/30/2020 15:17:55",
          "content": "<p>This is wrong, you can just upload csv to kernel and submit from kernel.</p>\n\n<p>Confirmation: <a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141554#800703\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141554#800703</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 827850,
          "author_name": "stephenmugisha",
          "author_url": "",
          "post_date": "04/30/2020 15:23:44",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>  Ok then thanks for clarifying, I was going by the standard rules of code competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 827872,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "04/30/2020 15:46:43",
          "content": "<p>Well, you are technically correct, but the output is just from reading and dumping the same csv. It is unfortunate that the whole test set can be seen in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 827895,
          "author_name": "nikhilmishradev",
          "author_url": "",
          "post_date": "04/30/2020 16:04:44",
          "content": "<p>Thanks for helping out <a href=\"/philippsinger\">@philippsinger</a> , I understood the rules, however this is quite dangerous. People aiming for a good rank or in the gold zone, apart from those winning prizes, can still do handlabelling and its not possible for Kaggle to check each and every submission.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 828833,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "05/01/2020 10:22:02",
          "content": "<p><a href=\"/nikhilmishradev\">@nikhilmishradev</a> 100% agree, unfortunately people don't always play fair. The majority does, but black sheep exist. Giving them an easy way to do it is unfortunate.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 830448,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/02/2020 15:42:05",
          "content": "<p>In general Kaggle prevents   handlabelling in CV competitions by using Two stages data and forcing everyone to upload his code and model at the end of the first stage.  The second stage data is made available only one week before the end.</p>\n\n<p>I don't know why the same thing is not done here ?  (And that would still be compatible with internet use for TPU)\nAnyway, All gold solutions  should be rigorously checked IMHO. \nAnd some random checks in silver/Top 50 may be done too. </p>\n\n<p>Fortunately for this competition,  the test dataset is relatively huge and mixes many languages. So handlabelling may take some efforts ^^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 830465,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "05/02/2020 15:55:44",
          "content": "<blockquote>\n  <p>Fortunately for this competition, the test dataset is relatively huge and mixes many languages. So handlabelling may take some efforts ^^</p>\n</blockquote>\n\n<p>One can simply translate them to common language :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 830469,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/02/2020 16:03:50",
          "content": "<p>You're right. </p>\n\n<p>But these translations may \"detoxify\" the language..and, hopefully, handlabelling will end up with mislabelling on private LB ^^</p>\n\n<p>Anyway , All top solutions need to be checked. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "827398": "Hello everyone after recent release of some kernels, I am a bit confused about the competition rules. Does the kernel to be submitted needs to perform training in the code itself? Or we can upload our outputs to a kernel and then submit? which in my opinion would not be fair to everyone?\n\nSince the kernels will not be re - run, my only concern is after the competition ends, will the kernels be checked to see if they are training using the given data, which again seems inplausible. Anyone could please clarify the rules here.",
    "827762": "As its a  code competition, the output/submission file must be generated from your Kaggle kernel. That means you can't just upload predictions for submission.\n\n&gt; Since the kernels will not be re - run, my only concern is after the competition ends, will the kernels be checked to see if they are training using the given data\n\nAnyone from kaggle should clarify on this.",
    "827841": "This is wrong, you can just upload csv to kernel and submit from kernel.\n\nConfirmation: https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141554#800703",
    "827850": "philippsinger  Ok then thanks for clarifying, I was going by the standard rules of code competitions.",
    "827872": "Well, you are technically correct, but the output is just from reading and dumping the same csv. It is unfortunate that the whole test set can be seen in this competition.",
    "827895": "Thanks for helping out @philippsinger , I understood the rules, however this is quite dangerous. People aiming for a good rank or in the gold zone, apart from those winning prizes, can still do handlabelling and its not possible for Kaggle to check each and every submission.",
    "828833": "nikhilmishradev 100% agree, unfortunately people don't always play fair. The majority does, but black sheep exist. Giving them an easy way to do it is unfortunate.",
    "830448": "In general Kaggle prevents   handlabelling in CV competitions by using Two stages data and forcing everyone to upload his code and model at the end of the first stage.  The second stage data is made available only one week before the end.\n\n I don't know why the same thing is not done here ?  (And that would still be compatible with internet use for TPU)\nAnyway, All gold solutions  should be rigorously checked IMHO. \nAnd some random checks in silver/Top 50 may be done too. \n\n\nFortunately for this competition,  the test dataset is relatively huge and mixes many languages. So handlabelling may take some efforts ^^",
    "830465": "&gt;Fortunately for this competition, the test dataset is relatively huge and mixes many languages. So handlabelling may take some efforts ^^\n\nOne can simply translate them to common language :)",
    "830469": "You're right. \n\nBut these translations may \"detoxify\" the language..and, hopefully, handlabelling will end up with mislabelling on private LB ^^\n\nAnyway , All top solutions need to be checked."
  },
  "source": "meta"
}