{
  "id": 147234,
  "title": "There is no completely hidden test set!?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/147234",
  "author_name": "",
  "post_date": "2020-04-30T00:46:14.122276800Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I noticed this in the code requirements section. I believe full test set is what we have already available in data section. People without a fair play spirit may indeed find an ensemble of native annotators for each language to hard code a submission. Is this a correct understanding of the situation or is there still a catch that will avoid this kind of behavior that I am unable to understand? Thanks!</p>",
  "messages": [
    {
      "id": "826934",
      "postDate": "04/30/2020 00:46:14",
      "content": "<p>I noticed this in the code requirements section. I believe full test set is what we have already available in data section. People without a fair play spirit may indeed find an ensemble of native annotators for each language to hard code a submission. Is this a correct understanding of the situation or is there still a catch that will avoid this kind of behavior that I am unable to understand? Thanks!</p>",
      "rawMarkdown": "I noticed this in the code requirements section. I believe full test set is what we have already available in data section. People without a fair play spirit may indeed find an ensemble of native annotators for each language to hard code a submission. Is this a correct understanding of the situation or is there still a catch that will avoid this kind of behavior that I am unable to understand? Thanks!",
      "votes": null
    },
    {
      "id": "829106",
      "postDate": "05/01/2020 14:01:08",
      "content": "<p>Manual annotation is against the rule.</p>\n\n<p>Rule5\n<code>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</code></p>\n\n<p>Don't worry Kaggle team is very strict, they have disqualified countless (both intended and unintended) cheaters even grandmasters in the past</p>",
      "rawMarkdown": "Manual annotation is against the rule.\n\nRule5\n`Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.`\n\nDon't worry Kaggle team is very strict, they have disqualified countless (both intended and unintended) cheaters even grandmasters in the past",
      "votes": null
    },
    {
      "id": "829122",
      "postDate": "05/01/2020 14:23:46",
      "content": "<p>Thanks for the reply!</p>",
      "rawMarkdown": "Thanks for the reply!",
      "votes": null
    },
    {
      "id": "841114",
      "postDate": "05/10/2020 15:52:51",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a> Do you think the same disqualification will apply to people who will just use csv files for submissions as well, e.g. there are forked public kernels that ensemble other inference kernel outputs?</p>",
      "rawMarkdown": "ratthachat Do you think the same disqualification will apply to people who will just use csv files for submissions as well, e.g. there are forked public kernels that ensemble other inference kernel outputs?",
      "votes": null
    },
    {
      "id": "841757",
      "postDate": "05/11/2020 02:22:03",
      "content": "<p>Hi <a href=\"/keremt\">@keremt</a> ! , as long as the submitted csv files can be reproduced, I think it is not aganst the competition rule (note that in “Code competition” like tweeter the rules are different) , e.g. either csv come from public kernel or your own kernel ...</p>\n\n<p>However, say, if we use some csv from public kernels which “do not” share Complete private data, and we could not reproduce. In this case, it may be against the rule.</p>",
      "rawMarkdown": "Hi @keremt ! , as long as the submitted csv files can be reproduced, I think it is not aganst the competition rule (note that in “Code competition” like tweeter the rules are different) , e.g. either csv come from public kernel or your own kernel ...\n\nHowever, say, if we use some csv from public kernels which “do not” share Complete private data, and we could not reproduce. In this case, it may be against the rule.",
      "votes": null
    },
    {
      "id": "842592",
      "postDate": "05/11/2020 13:45:30",
      "content": "<p>This might be a problem. Is submission.csv file allowed to be uploaded to the notebook instead of uploading a model -&gt; making predictions -&gt; producing submission.csv from within the notebook ?</p>\n\n<p>According to the Rules only public external data or pretrained models are allowed. The submission.csv file is neither public nor a model. ???</p>",
      "rawMarkdown": "This might be a problem. Is submission.csv file allowed to be uploaded to the notebook instead of uploading a model -&gt; making predictions -&gt; producing submission.csv from within the notebook ?\n\nAccording to the Rules only public external data or pretrained models are allowed. The submission.csv file is neither public nor a model. ???",
      "votes": null
    },
    {
      "id": "842750",
      "postDate": "05/11/2020 16:06:21",
      "content": "<p>Yeah, the rules for this competition seems a bit vague. I am also worried about LB probing as there is no additional private test set. I am able to consistently get better results just by simple rank averaging trying out my different submissions and it wouldn't be a problem to programmatically submit a variant of multiple models every day with the max allowed submission. There will be no shake up  and anyone can just brute force their way up.</p>",
      "rawMarkdown": "Yeah, the rules for this competition seems a bit vague. I am also worried about LB probing as there is no additional private test set. I am able to consistently get better results just by simple rank averaging trying out my different submissions and it wouldn't be a problem to programmatically submit a variant of multiple models every day with the max allowed submission. There will be no shake up  and anyone can just brute force their way up.",
      "votes": null
    },
    {
      "id": "843551",
      "postDate": "05/12/2020 05:34:53",
      "content": "<p><a href=\"/isakev\">@isakev</a> <a href=\"/keremt\">@keremt</a> It's a nature of this competition type to allow us to make an extreme ensemble. However, it will not be easy to avoid shake up as there's only 30% in public test set. </p>\n\n<p>And it seems to me that the labels are very noisy as they are subjective to both languages and labellers ... So I think huge shake up is coming for this competition.</p>\n\n<p>In fact, I participated two similar competitions e.g. VSB Powerline prediction and Cloud Detection where I and many participants were severely shaken down in both of them :) </p>\n\n<p>Therefore, the key in this competition to me is to find the solid validation pipeline, which is really very difficult. In my opinion, To make our models robust, we should assumesthat private test set can come from different distributions from public test set.</p>",
      "rawMarkdown": "isakev @keremt It's a nature of this competition type to allow us to make an extreme ensemble. However, it will not be easy to avoid shake up as there's only 30% in public test set. \n\nAnd it seems to me that the labels are very noisy as they are subjective to both languages and labellers ... So I think huge shake up is coming for this competition.\n\nIn fact, I participated two similar competitions e.g. VSB Powerline prediction and Cloud Detection where I and many participants were severely shaken down in both of them :) \n\nTherefore, the key in this competition to me is to find the solid validation pipeline, which is really very difficult. In my opinion, To make our models robust, we should assumesthat private test set can come from different distributions from public test set.",
      "votes": null
    },
    {
      "id": "854308",
      "postDate": "05/20/2020 00:25:20",
      "content": "<p><a href=\"/keremt\">@keremt</a> I thought it was worth chiming in here that our current infrastructure that facilitates the TPU integration is currently blocking our ability to design this competition as a code competition with a fully hidden test set. This is something engineering is actively tackling, but is not currently available to us. So we definitely concur, that it would be far more ideal to have a fully hidden test set, and we are working towards this possibility. In the meantime, the rules prohibitions mentioned by other users will be our mechanism for disqualifying winners who are found to have committed hand-labeling violations.</p>",
      "rawMarkdown": "keremt I thought it was worth chiming in here that our current infrastructure that facilitates the TPU integration is currently blocking our ability to design this competition as a code competition with a fully hidden test set. This is something engineering is actively tackling, but is not currently available to us. So we definitely concur, that it would be far more ideal to have a fully hidden test set, and we are working towards this possibility. In the meantime, the rules prohibitions mentioned by other users will be our mechanism for disqualifying winners who are found to have committed hand-labeling violations.",
      "votes": null
    },
    {
      "id": "854486",
      "postDate": "05/20/2020 04:01:43",
      "content": "<p>Thank you <a href=\"/juliaelliott\">@juliaelliott</a> for clarification. Definitely, being able to use the speed of TPU outweighs the drawbacks of openly disclosed test set, at least for those of us who enjoy less than 8 GPUs in our arsenal.</p>\n\n<p>BUT, PLEASE CLARIFY, do we need to upload all the employed models into the notebook to produce submission.csv or is it enough and legit to upload submissiin 'csv' files (they are not public) without uploading models and running predictions ?</p>",
      "rawMarkdown": "Thank you @juliaelliott for clarification. Definitely, being able to use the speed of TPU outweighs the drawbacks of openly disclosed test set, at least for those of us who enjoy less than 8 GPUs in our arsenal.\n\nBUT, PLEASE CLARIFY, do we need to upload all the employed models into the notebook to produce submission.csv or is it enough and legit to upload submissiin 'csv' files (they are not public) without uploading models and running predictions ?",
      "votes": null
    },
    {
      "id": "855453",
      "postDate": "05/21/2020 00:03:16",
      "content": "<p><a href=\"/isakev\">@isakev</a> Since this is not a true code re-run competition, the only restriction is that the <code>submission.csv</code> must be generated out of a Kaggle notebook. But we do not say how that submission should be generated. Therefore, yes, you could make a submission that doesn't involve uploaded models/code and simply loads the <code>submission.csv</code> with predictions for the known test set id's. Winners' code will then be inspected (to ensure hand-labeling was not used) to earn a prize.</p>",
      "rawMarkdown": "isakev Since this is not a true code re-run competition, the only restriction is that the `submission.csv` must be generated out of a Kaggle notebook. But we do not say how that submission should be generated. Therefore, yes, you could make a submission that doesn't involve uploaded models/code and simply loads the `submission.csv` with predictions for the known test set id's. Winners' code will then be inspected (to ensure hand-labeling was not used) to earn a prize.",
      "votes": null
    },
    {
      "id": "855469",
      "postDate": "05/21/2020 00:39:53",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>As I said in another topic,  a better idea would be a two stages Competition, as done in many computer vision competitions, with the 2nd stage test set available only a week before the end. </p>\n\n<p>That would be compatible with internet enabled for TPU</p>",
      "rawMarkdown": "juliaelliott \n\nAs I said in another topic,  a better idea would be a two stages Competition, as done in many computer vision competitions, with the 2nd stage test set available only a week before the end. \n\nThat would be compatible with internet enabled for TPU",
      "votes": null
    },
    {
      "id": "858471",
      "postDate": "05/23/2020 14:14:47",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Does it mean only money prize winners' code will be inspected and not everyone in the medal zone? Wouldn't this allow a lot people to exploit the current way of testing except for the people who chases the prize?</p>",
      "rawMarkdown": "juliaelliott Does it mean only money prize winners' code will be inspected and not everyone in the medal zone? Wouldn't this allow a lot people to exploit the current way of testing except for the people who chases the prize?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 829106,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "05/01/2020 14:01:08",
      "content": "<p>Manual annotation is against the rule.</p>\n\n<p>Rule5\n<code>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</code></p>\n\n<p>Don't worry Kaggle team is very strict, they have disqualified countless (both intended and unintended) cheaters even grandmasters in the past</p>",
      "votes": null,
      "replies": [
        {
          "id": 829122,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "05/01/2020 14:23:46",
          "content": "<p>Thanks for the reply!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 841114,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "05/10/2020 15:52:51",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a> Do you think the same disqualification will apply to people who will just use csv files for submissions as well, e.g. there are forked public kernels that ensemble other inference kernel outputs?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 841757,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "05/11/2020 02:22:03",
          "content": "<p>Hi <a href=\"/keremt\">@keremt</a> ! , as long as the submitted csv files can be reproduced, I think it is not aganst the competition rule (note that in “Code competition” like tweeter the rules are different) , e.g. either csv come from public kernel or your own kernel ...</p>\n\n<p>However, say, if we use some csv from public kernels which “do not” share Complete private data, and we could not reproduce. In this case, it may be against the rule.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 842592,
          "author_name": "isakev",
          "author_url": "",
          "post_date": "05/11/2020 13:45:30",
          "content": "<p>This might be a problem. Is submission.csv file allowed to be uploaded to the notebook instead of uploading a model -&gt; making predictions -&gt; producing submission.csv from within the notebook ?</p>\n\n<p>According to the Rules only public external data or pretrained models are allowed. The submission.csv file is neither public nor a model. ???</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 842750,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "05/11/2020 16:06:21",
          "content": "<p>Yeah, the rules for this competition seems a bit vague. I am also worried about LB probing as there is no additional private test set. I am able to consistently get better results just by simple rank averaging trying out my different submissions and it wouldn't be a problem to programmatically submit a variant of multiple models every day with the max allowed submission. There will be no shake up  and anyone can just brute force their way up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 843551,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "05/12/2020 05:34:53",
          "content": "<p><a href=\"/isakev\">@isakev</a> <a href=\"/keremt\">@keremt</a> It's a nature of this competition type to allow us to make an extreme ensemble. However, it will not be easy to avoid shake up as there's only 30% in public test set. </p>\n\n<p>And it seems to me that the labels are very noisy as they are subjective to both languages and labellers ... So I think huge shake up is coming for this competition.</p>\n\n<p>In fact, I participated two similar competitions e.g. VSB Powerline prediction and Cloud Detection where I and many participants were severely shaken down in both of them :) </p>\n\n<p>Therefore, the key in this competition to me is to find the solid validation pipeline, which is really very difficult. In my opinion, To make our models robust, we should assumesthat private test set can come from different distributions from public test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 854308,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "05/20/2020 00:25:20",
      "content": "<p><a href=\"/keremt\">@keremt</a> I thought it was worth chiming in here that our current infrastructure that facilitates the TPU integration is currently blocking our ability to design this competition as a code competition with a fully hidden test set. This is something engineering is actively tackling, but is not currently available to us. So we definitely concur, that it would be far more ideal to have a fully hidden test set, and we are working towards this possibility. In the meantime, the rules prohibitions mentioned by other users will be our mechanism for disqualifying winners who are found to have committed hand-labeling violations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 854486,
          "author_name": "isakev",
          "author_url": "",
          "post_date": "05/20/2020 04:01:43",
          "content": "<p>Thank you <a href=\"/juliaelliott\">@juliaelliott</a> for clarification. Definitely, being able to use the speed of TPU outweighs the drawbacks of openly disclosed test set, at least for those of us who enjoy less than 8 GPUs in our arsenal.</p>\n\n<p>BUT, PLEASE CLARIFY, do we need to upload all the employed models into the notebook to produce submission.csv or is it enough and legit to upload submissiin 'csv' files (they are not public) without uploading models and running predictions ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 855453,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "05/21/2020 00:03:16",
          "content": "<p><a href=\"/isakev\">@isakev</a> Since this is not a true code re-run competition, the only restriction is that the <code>submission.csv</code> must be generated out of a Kaggle notebook. But we do not say how that submission should be generated. Therefore, yes, you could make a submission that doesn't involve uploaded models/code and simply loads the <code>submission.csv</code> with predictions for the known test set id's. Winners' code will then be inspected (to ensure hand-labeling was not used) to earn a prize.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 855469,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "05/21/2020 00:39:53",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>As I said in another topic,  a better idea would be a two stages Competition, as done in many computer vision competitions, with the 2nd stage test set available only a week before the end. </p>\n\n<p>That would be compatible with internet enabled for TPU</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 858471,
          "author_name": "keremt",
          "author_url": "",
          "post_date": "05/23/2020 14:14:47",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Does it mean only money prize winners' code will be inspected and not everyone in the medal zone? Wouldn't this allow a lot people to exploit the current way of testing except for the people who chases the prize?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "826934": "I noticed this in the code requirements section. I believe full test set is what we have already available in data section. People without a fair play spirit may indeed find an ensemble of native annotators for each language to hard code a submission. Is this a correct understanding of the situation or is there still a catch that will avoid this kind of behavior that I am unable to understand? Thanks!",
    "829106": "Manual annotation is against the rule.\n\nRule5\n`Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.`\n\nDon't worry Kaggle team is very strict, they have disqualified countless (both intended and unintended) cheaters even grandmasters in the past",
    "829122": "Thanks for the reply!",
    "841114": "ratthachat Do you think the same disqualification will apply to people who will just use csv files for submissions as well, e.g. there are forked public kernels that ensemble other inference kernel outputs?",
    "841757": "Hi @keremt ! , as long as the submitted csv files can be reproduced, I think it is not aganst the competition rule (note that in “Code competition” like tweeter the rules are different) , e.g. either csv come from public kernel or your own kernel ...\n\nHowever, say, if we use some csv from public kernels which “do not” share Complete private data, and we could not reproduce. In this case, it may be against the rule.",
    "842592": "This might be a problem. Is submission.csv file allowed to be uploaded to the notebook instead of uploading a model -&gt; making predictions -&gt; producing submission.csv from within the notebook ?\n\nAccording to the Rules only public external data or pretrained models are allowed. The submission.csv file is neither public nor a model. ???",
    "842750": "Yeah, the rules for this competition seems a bit vague. I am also worried about LB probing as there is no additional private test set. I am able to consistently get better results just by simple rank averaging trying out my different submissions and it wouldn't be a problem to programmatically submit a variant of multiple models every day with the max allowed submission. There will be no shake up  and anyone can just brute force their way up.",
    "843551": "isakev @keremt It's a nature of this competition type to allow us to make an extreme ensemble. However, it will not be easy to avoid shake up as there's only 30% in public test set. \n\nAnd it seems to me that the labels are very noisy as they are subjective to both languages and labellers ... So I think huge shake up is coming for this competition.\n\nIn fact, I participated two similar competitions e.g. VSB Powerline prediction and Cloud Detection where I and many participants were severely shaken down in both of them :) \n\nTherefore, the key in this competition to me is to find the solid validation pipeline, which is really very difficult. In my opinion, To make our models robust, we should assumesthat private test set can come from different distributions from public test set.",
    "854308": "keremt I thought it was worth chiming in here that our current infrastructure that facilitates the TPU integration is currently blocking our ability to design this competition as a code competition with a fully hidden test set. This is something engineering is actively tackling, but is not currently available to us. So we definitely concur, that it would be far more ideal to have a fully hidden test set, and we are working towards this possibility. In the meantime, the rules prohibitions mentioned by other users will be our mechanism for disqualifying winners who are found to have committed hand-labeling violations.",
    "854486": "Thank you @juliaelliott for clarification. Definitely, being able to use the speed of TPU outweighs the drawbacks of openly disclosed test set, at least for those of us who enjoy less than 8 GPUs in our arsenal.\n\nBUT, PLEASE CLARIFY, do we need to upload all the employed models into the notebook to produce submission.csv or is it enough and legit to upload submissiin 'csv' files (they are not public) without uploading models and running predictions ?",
    "855453": "isakev Since this is not a true code re-run competition, the only restriction is that the `submission.csv` must be generated out of a Kaggle notebook. But we do not say how that submission should be generated. Therefore, yes, you could make a submission that doesn't involve uploaded models/code and simply loads the `submission.csv` with predictions for the known test set id's. Winners' code will then be inspected (to ensure hand-labeling was not used) to earn a prize.",
    "855469": "juliaelliott \n\nAs I said in another topic,  a better idea would be a two stages Competition, as done in many computer vision competitions, with the 2nd stage test set available only a week before the end. \n\nThat would be compatible with internet enabled for TPU",
    "858471": "juliaelliott Does it mean only money prize winners' code will be inspected and not everyone in the medal zone? Wouldn't this allow a lot people to exploit the current way of testing except for the people who chases the prize?"
  },
  "source": "meta"
}