{
  "id": 97652,
  "title": "Any synchronous kernel competition has a data leakage from the private test set?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/97652",
  "author_name": "",
  "post_date": "2019-06-28T07:39:00.413312800Z",
  "votes": 11,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Overview page tells us:\n<code>\nIn synchronous KO competitions, note that a submission that results in an error--either within the kernel or within the process of scoring--will count against your daily submission limit and will not return the specific error message. This is necessary to prevent probing the private test set.\n</code>\nSo, we can conditionally prepare submissions that would fail on the private test set. By doing this, we can get 1 bit of information about the private test set. We still have two months, 5 submits per day, so at least we can get some useful stats about private test set, its classes distribution, etc. Of course, you can get more information if you're in an X5 team.</p>",
  "messages": [
    {
      "id": "563361",
      "postDate": "06/28/2019 07:39:00",
      "content": "<p>Overview page tells us:\n<code>\nIn synchronous KO competitions, note that a submission that results in an error--either within the kernel or within the process of scoring--will count against your daily submission limit and will not return the specific error message. This is necessary to prevent probing the private test set.\n</code>\nSo, we can conditionally prepare submissions that would fail on the private test set. By doing this, we can get 1 bit of information about the private test set. We still have two months, 5 submits per day, so at least we can get some useful stats about private test set, its classes distribution, etc. Of course, you can get more information if you're in an X5 team.</p>",
      "rawMarkdown": "Overview page tells us:\n```\nIn synchronous KO competitions, note that a submission that results in an error--either within the kernel or within the process of scoring--will count against your daily submission limit and will not return the specific error message. This is necessary to prevent probing the private test set.\n```\nSo, we can conditionally prepare submissions that would fail on the private test set. By doing this, we can get 1 bit of information about the private test set. We still have two months, 5 submits per day, so at least we can get some useful stats about private test set, its classes distribution, etc. Of course, you can get more information if you're in an X5 team.",
      "votes": null
    },
    {
      "id": "563671",
      "postDate": "06/28/2019 14:51:39",
      "content": "<p>Think so, you can check <a href=\"https://www.kaggle.com/c/instant-gratification\">Instant Gratification</a> and how chris tested private LB <a href=\"https://www.kaggle.com/cdeotte/private-lb-probing-0-950\">Private LB Probing</a>.\nbtw, if you want to create a X5 team ... I'm in :D</p>",
      "rawMarkdown": "Think so, you can check [Instant Gratification](https://www.kaggle.com/c/instant-gratification) and how chris tested private LB [Private LB Probing](https://www.kaggle.com/cdeotte/private-lb-probing-0-950).\nbtw, if you want to create a X5 team ... I'm in :D",
      "votes": null
    },
    {
      "id": "563831",
      "postDate": "06/28/2019 18:12:16",
      "content": "<p>As I suggested before, there should be a firewall in the form of a synthetic Private set to help catch major errors while avoiding probing. The actual Private set should only be accessible to probing shortly before the deadline (to iron out the remaining errors).</p>",
      "rawMarkdown": "As I suggested before, there should be a firewall in the form of a synthetic Private set to help catch major errors while avoiding probing. The actual Private set should only be accessible to probing shortly before the deadline (to iron out the remaining errors).",
      "votes": null
    },
    {
      "id": "563873",
      "postDate": "06/28/2019 19:00:55",
      "content": "<p>Yes, that would help</p>",
      "rawMarkdown": "Yes, that would help",
      "votes": null
    },
    {
      "id": "565715",
      "postDate": "07/01/2019 09:50:07",
      "content": "<p>Don't forget to count me in, guys 👍 </p>",
      "rawMarkdown": "Don't forget to count me in, guys 👍",
      "votes": null
    },
    {
      "id": "571948",
      "postDate": "07/10/2019 09:02:46",
      "content": "<p>This post is misleading. You only get that information for the data itself, not for the label. Have fun reconstructing the private test images one bit at a time! This is no leakage.</p>",
      "rawMarkdown": "This post is misleading. You only get that information for the data itself, not for the label. Have fun reconstructing the private test images one bit at a time! This is no leakage.",
      "votes": null
    },
    {
      "id": "595678",
      "postDate": "08/09/2019 14:34:39",
      "content": "<p>As we know (submission.csv, LB score) pair for each submission, we can get log2(N)-bit information from each submission by changing submission file according to desired test dataset information, where N is the number of known (submission.csv, LB score) pairs (~= the number of total submissions).\nIs this correct?</p>",
      "rawMarkdown": "As we know (submission.csv, LB score) pair for each submission, we can get log2(N)-bit information from each submission by changing submission file according to desired test dataset information, where N is the number of known (submission.csv, LB score) pairs (~= the number of total submissions).\nIs this correct?",
      "votes": null
    },
    {
      "id": "596188",
      "postDate": "08/10/2019 09:33:11",
      "content": "<p>Well, if it's the best usage you can possibly think of, then yes.</p>",
      "rawMarkdown": "Well, if it's the best usage you can possibly think of, then yes.",
      "votes": null
    },
    {
      "id": "596190",
      "postDate": "08/10/2019 09:33:43",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "596913",
      "postDate": "08/11/2019 13:57:42",
      "content": "<p>Assuming we have N submissions that we know each LB scores (and all scores are different), we can encode pattern n (1&lt;= n &lt;= N) by selecting n-th submission in submitted kernel according to information about test private dateset. We can know which pattern was selected by seeing LB score.</p>",
      "rawMarkdown": "Assuming we have N submissions that we know each LB scores (and all scores are different), we can encode pattern n (1&lt;= n &lt;= N) by selecting n-th submission in submitted kernel according to information about test private dateset. We can know which pattern was selected by seeing LB score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 563671,
      "author_name": "jesucristo",
      "author_url": "",
      "post_date": "06/28/2019 14:51:39",
      "content": "<p>Think so, you can check <a href=\"https://www.kaggle.com/c/instant-gratification\">Instant Gratification</a> and how chris tested private LB <a href=\"https://www.kaggle.com/cdeotte/private-lb-probing-0-950\">Private LB Probing</a>.\nbtw, if you want to create a X5 team ... I'm in :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 565715,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "07/01/2019 09:50:07",
          "content": "<p>Don't forget to count me in, guys 👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 563831,
      "author_name": "stocks",
      "author_url": "",
      "post_date": "06/28/2019 18:12:16",
      "content": "<p>As I suggested before, there should be a firewall in the form of a synthetic Private set to help catch major errors while avoiding probing. The actual Private set should only be accessible to probing shortly before the deadline (to iron out the remaining errors).</p>",
      "votes": null,
      "replies": [
        {
          "id": 563873,
          "author_name": "artyomp",
          "author_url": "",
          "post_date": "06/28/2019 19:00:55",
          "content": "<p>Yes, that would help</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 571948,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "07/10/2019 09:02:46",
      "content": "<p>This post is misleading. You only get that information for the data itself, not for the label. Have fun reconstructing the private test images one bit at a time! This is no leakage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 596188,
          "author_name": "artyomp",
          "author_url": "",
          "post_date": "08/10/2019 09:33:11",
          "content": "<p>Well, if it's the best usage you can possibly think of, then yes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 595678,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "08/09/2019 14:34:39",
      "content": "<p>As we know (submission.csv, LB score) pair for each submission, we can get log2(N)-bit information from each submission by changing submission file according to desired test dataset information, where N is the number of known (submission.csv, LB score) pairs (~= the number of total submissions).\nIs this correct?</p>",
      "votes": null,
      "replies": [
        {
          "id": 596190,
          "author_name": "artyomp",
          "author_url": "",
          "post_date": "08/10/2019 09:33:43",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 596913,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "08/11/2019 13:57:42",
          "content": "<p>Assuming we have N submissions that we know each LB scores (and all scores are different), we can encode pattern n (1&lt;= n &lt;= N) by selecting n-th submission in submitted kernel according to information about test private dateset. We can know which pattern was selected by seeing LB score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "563361": "Overview page tells us:\n```\nIn synchronous KO competitions, note that a submission that results in an error--either within the kernel or within the process of scoring--will count against your daily submission limit and will not return the specific error message. This is necessary to prevent probing the private test set.\n```\nSo, we can conditionally prepare submissions that would fail on the private test set. By doing this, we can get 1 bit of information about the private test set. We still have two months, 5 submits per day, so at least we can get some useful stats about private test set, its classes distribution, etc. Of course, you can get more information if you're in an X5 team.",
    "563671": "Think so, you can check [Instant Gratification](https://www.kaggle.com/c/instant-gratification) and how chris tested private LB [Private LB Probing](https://www.kaggle.com/cdeotte/private-lb-probing-0-950).\nbtw, if you want to create a X5 team ... I'm in :D",
    "563831": "As I suggested before, there should be a firewall in the form of a synthetic Private set to help catch major errors while avoiding probing. The actual Private set should only be accessible to probing shortly before the deadline (to iron out the remaining errors).",
    "563873": "Yes, that would help",
    "565715": "Don't forget to count me in, guys 👍",
    "571948": "This post is misleading. You only get that information for the data itself, not for the label. Have fun reconstructing the private test images one bit at a time! This is no leakage.",
    "595678": "As we know (submission.csv, LB score) pair for each submission, we can get log2(N)-bit information from each submission by changing submission file according to desired test dataset information, where N is the number of known (submission.csv, LB score) pairs (~= the number of total submissions).\nIs this correct?",
    "596188": "Well, if it's the best usage you can possibly think of, then yes.",
    "596190": "",
    "596913": "Assuming we have N submissions that we know each LB scores (and all scores are different), we can encode pattern n (1&lt;= n &lt;= N) by selecting n-th submission in submitted kernel according to information about test private dateset. We can know which pattern was selected by seeing LB score."
  },
  "source": "meta"
}