{
  "id": 77919,
  "title": "Loophole in the Rules",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77919",
  "author_name": "Theo Viel",
  "post_date": "2019-01-17T16:15:53.088000",
  "votes": 9,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Recently, an upsetting kernel has been published : \n<a href=\"https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string\">https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string</a> (EDIT)</p>\n\n<p>It basically allows us to reuse previous submissions into a string. Which can lead to a significant boost in performance.\nAs this is not recognized as external data but still kinda is, can anyone (from kaggle?) confirm that such practices are forbidden (or not but I would be surprised) ?</p>\n\n<p>Thanks in advance</p>",
  "messages": [
    {
      "id": 457753,
      "postDate": "2019-01-18T03:14:08.927Z",
      "content": "<p>The final score will be tested on a new private test set, this should not be a big problem. BTW, this is cheating, indeed</p>",
      "rawMarkdown": "The final score will be tested on a new private test set, this should not be a big problem. BTW, this is cheating, indeed",
      "votes": 10,
      "replies": [
        {
          "id": 457868,
          "postDate": "2019-01-18T08:34:03.147Z",
          "content": "<p>Still, having more labelled samples can be useful for pseudo-labelling and can improve a model.</p>",
          "rawMarkdown": "Still, having more labelled samples can be useful for pseudo-labelling and can improve a model."
        }
      ]
    },
    {
      "id": 457527,
      "postDate": "2019-01-17T16:15:53.090Z",
      "content": "<p>Recently, an upsetting kernel has been published : \n<a href=\"https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string\">https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string</a> (EDIT)</p>\n\n<p>It basically allows us to reuse previous submissions into a string. Which can lead to a significant boost in performance.\nAs this is not recognized as external data but still kinda is, can anyone (from kaggle?) confirm that such practices are forbidden (or not but I would be surprised) ?</p>\n\n<p>Thanks in advance</p>",
      "rawMarkdown": "Recently, an upsetting kernel has been published : \nhttps://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string (EDIT)\n\nIt basically allows us to reuse previous submissions into a string. Which can lead to a significant boost in performance.\nAs this is not recognized as external data but still kinda is, can anyone (from kaggle?) confirm that such practices are forbidden (or not but I would be surprised) ?\n\nThanks in advance",
      "votes": 9
    },
    {
      "id": 457878,
      "postDate": "2019-01-18T08:47:22.967Z",
      "content": "<p>This is an insincere kernel :)</p>",
      "rawMarkdown": "This is an insincere kernel :)",
      "votes": 5
    },
    {
      "id": 457904,
      "postDate": "2019-01-18T10:08:50.963Z",
      "content": "<p><a href=\"/theoviel\">@theoviel</a>, don't worry about it as this is clearly against the rule.</p>\n\n<p>The whole idea of a kernel only competition is that you run your model end-to-end within the kernel environment with time constraint and all.  I just finished competing in the <a href=\"https://www.kaggle.com/c/traveling-santa-2018-prime-paths\">\"Travelling Santa 2018 - prime paths\"</a> contest and a lot of similar questions were asked about the kernel prize part. The Kaggle team made it very clear that you have to do all your pre-processing plus modelling within the kernel environment.</p>",
      "rawMarkdown": "@theoviel, don't worry about it as this is clearly against the rule.\n\nThe whole idea of a kernel only competition is that you run your model end-to-end within the kernel environment with time constraint and all.  I just finished competing in the [\"Travelling Santa 2018 - prime paths\"][1] contest and a lot of similar questions were asked about the kernel prize part. The Kaggle team made it very clear that you have to do all your pre-processing plus modelling within the kernel environment.\n\n\n  [1]: https://www.kaggle.com/c/traveling-santa-2018-prime-paths",
      "votes": 2,
      "replies": [
        {
          "id": 458300,
          "postDate": "2019-01-19T09:43:09.143Z",
          "content": "<p>I think there are several problems if someone wants to cheat.  </p>\n\n<p>Suppose that I did fixing misspelling words out side of Kaggle kernel environment. Let say I have a dictionary including 10,000 fixed words. If I directly copy my dictionary into kernel, I bet the kernel will be crashed. So, using this trick I think I can have an advantage for cleaning data. Of course, this data does not depend on testset. It should be no problem on phase 2. <br>\nIt is a simple case. Let think about if someone wants to bring their \"pretrained weight\" into their models ;).  </p>\n\n<p>I wonder if Kaggle team can review all the kernel codes to know if they are cheating or not</p>",
          "rawMarkdown": "I think there are several problems if someone wants to cheat.  \n\nSuppose that I did fixing misspelling words out side of Kaggle kernel environment. Let say I have a dictionary including 10,000 fixed words. If I directly copy my dictionary into kernel, I bet the kernel will be crashed. So, using this trick I think I can have an advantage for cleaning data. Of course, this data does not depend on testset. It should be no problem on phase 2.   \nIt is a simple case. Let think about if someone wants to bring their \"pretrained weight\" into their models ;).  \n\nI wonder if Kaggle team can review all the kernel codes to know if they are cheating or not",
          "votes": 1
        },
        {
          "id": 458361,
          "postDate": "2019-01-19T12:49:36.810Z",
          "content": "<p>Your pre-trained weights will be considered an external data hence against the rule.</p>\n\n<p>On your other point, Kaggle go through cheaters removal at the end of every competition before finalizing the results. I remember in my 1st Kaggle competition, I jumped 2 percentage points i.e. from top 15% to 13% a few days after the closing. So food for thought.</p>",
          "rawMarkdown": "Your pre-trained weights will be considered an external data hence against the rule.\n\nOn your other point, Kaggle go through cheaters removal at the end of every competition before finalizing the results. I remember in my 1st Kaggle competition, I jumped 2 percentage points i.e. from top 15% to 13% a few days after the closing. So food for thought."
        }
      ]
    },
    {
      "id": 457537,
      "postDate": "2019-01-17T16:48:58.243Z",
      "content": "<p>I think you mean this kernel: <a href=\"https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string\">https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string</a></p>",
      "rawMarkdown": "I think you mean this kernel: https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string",
      "votes": 2,
      "replies": [
        {
          "id": 457662,
          "postDate": "2019-01-17T22:05:49.593Z",
          "content": "<p>Oops yeah, messed up my c/c !</p>",
          "rawMarkdown": "Oops yeah, messed up my c/c !"
        }
      ]
    },
    {
      "id": 457563,
      "postDate": "2019-01-17T18:09:50.083Z",
      "content": "<p>Well it is pretty obvious that this is cheating so I would not worry too much about it.</p>",
      "rawMarkdown": "Well it is pretty obvious that this is cheating so I would not worry too much about it.",
      "votes": -1
    },
    {
      "id": 457898,
      "postDate": "2019-01-18T09:55:27.677Z",
      "content": "<p>Training your model locally and submitting only indexes saves a lot of time, so it is not cheating. Using external information in the final kernel, such as pseudo labeled public test data is cheating.</p>",
      "rawMarkdown": "Training your model locally and submitting only indexes saves a lot of time, so it is not cheating. Using external information in the final kernel, such as pseudo labeled public test data is cheating.",
      "votes": -5
    },
    {
      "id": 457538,
      "postDate": "2019-01-17T16:49:39.847Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 457753,
      "author_name": "good good study",
      "author_url": "",
      "post_date": "2019-01-18T03:14:08.927000",
      "content": "<p>The final score will be tested on a new private test set, this should not be a big problem. BTW, this is cheating, indeed</p>",
      "votes": 10,
      "replies": [
        {
          "id": 457868,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2019-01-18T08:34:03.147000",
          "content": "<p>Still, having more labelled samples can be useful for pseudo-labelling and can improve a model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 457878,
      "author_name": "Juju",
      "author_url": "",
      "post_date": "2019-01-18T08:47:22.967000",
      "content": "<p>This is an insincere kernel :)</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 457904,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-01-18T10:08:50.963000",
      "content": "<p><a href=\"/theoviel\">@theoviel</a>, don't worry about it as this is clearly against the rule.</p>\n\n<p>The whole idea of a kernel only competition is that you run your model end-to-end within the kernel environment with time constraint and all.  I just finished competing in the <a href=\"https://www.kaggle.com/c/traveling-santa-2018-prime-paths\">\"Travelling Santa 2018 - prime paths\"</a> contest and a lot of similar questions were asked about the kernel prize part. The Kaggle team made it very clear that you have to do all your pre-processing plus modelling within the kernel environment.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 458300,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2019-01-19T09:43:09.143000",
          "content": "<p>I think there are several problems if someone wants to cheat.  </p>\n\n<p>Suppose that I did fixing misspelling words out side of Kaggle kernel environment. Let say I have a dictionary including 10,000 fixed words. If I directly copy my dictionary into kernel, I bet the kernel will be crashed. So, using this trick I think I can have an advantage for cleaning data. Of course, this data does not depend on testset. It should be no problem on phase 2. <br>\nIt is a simple case. Let think about if someone wants to bring their \"pretrained weight\" into their models ;).  </p>\n\n<p>I wonder if Kaggle team can review all the kernel codes to know if they are cheating or not</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 458361,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-01-19T12:49:36.810000",
          "content": "<p>Your pre-trained weights will be considered an external data hence against the rule.</p>\n\n<p>On your other point, Kaggle go through cheaters removal at the end of every competition before finalizing the results. I remember in my 1st Kaggle competition, I jumped 2 percentage points i.e. from top 15% to 13% a few days after the closing. So food for thought.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 457537,
      "author_name": "Mark Worrall",
      "author_url": "",
      "post_date": "2019-01-17T16:48:58.243000",
      "content": "<p>I think you mean this kernel: <a href=\"https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string\">https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 457662,
          "author_name": "Theo Viel",
          "author_url": "",
          "post_date": "2019-01-17T22:05:49.593000",
          "content": "<p>Oops yeah, messed up my c/c !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 457563,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2019-01-17T18:09:50.083000",
      "content": "<p>Well it is pretty obvious that this is cheating so I would not worry too much about it.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 457898,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-18T09:55:27.677000",
      "content": "<p>Training your model locally and submitting only indexes saves a lot of time, so it is not cheating. Using external information in the final kernel, such as pseudo labeled public test data is cheating.</p>",
      "votes": -5,
      "replies": []
    },
    {
      "id": 457538,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-17T16:49:39.847000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "457753": "The final score will be tested on a new private test set, this should not be a big problem. BTW, this is cheating, indeed",
    "457527": "Recently, an upsetting kernel has been published : \nhttps://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string (EDIT)\n\nIt basically allows us to reuse previous submissions into a string. Which can lead to a significant boost in performance.\nAs this is not recognized as external data but still kinda is, can anyone (from kaggle?) confirm that such practices are forbidden (or not but I would be surprised) ?\n\nThanks in advance",
    "457878": "This is an insincere kernel :)",
    "457904": "@theoviel, don't worry about it as this is clearly against the rule.\n\nThe whole idea of a kernel only competition is that you run your model end-to-end within the kernel environment with time constraint and all.  I just finished competing in the [\"Travelling Santa 2018 - prime paths\"][1] contest and a lot of similar questions were asked about the kernel prize part. The Kaggle team made it very clear that you have to do all your pre-processing plus modelling within the kernel environment.\n\n\n  [1]: https://www.kaggle.com/c/traveling-santa-2018-prime-paths",
    "457537": "I think you mean this kernel: https://www.kaggle.com/paulorzp/compressing-a-binary-submission-within-a-string",
    "457563": "Well it is pretty obvious that this is cheating so I would not worry too much about it.",
    "457898": "Training your model locally and submitting only indexes saves a lot of time, so it is not cheating. Using external information in the final kernel, such as pseudo labeled public test data is cheating.",
    "457538": ""
  }
}