{
  "id": 117473,
  "title": "LB probing ",
  "url": "/competitions/tensorflow2-question-answering/discussion/117473",
  "author_name": "Yury Kashnitsky",
  "post_date": "2019-11-15T20:02:30.459000",
  "votes": 14,
  "comment_count": 9,
  "views": 0,
  "content": "<p>With this competition format it's relatively easy to know more about the (hidden) test dataset.\nIt's explained in <a href=\"https://www.kaggle.com/kashnitsky/one-bit-at-a-time-lb-probing\">this Notebook</a>. </p>\n\n<p>```\nWhen you press \"Commit\", you execute the code with public test dataset. But when you hit \"Submit\", in the background, the same code is run against private test dataset.</p>\n\n<p>Thus we can find out more about the private (hidden) test set. Which actually defines the prizes.</p>\n\n<p>The approach is straightforward:</p>\n\n<ol>\n<li>ask a binary question about the test dataset (eg. whether it's longer than 1000)</li>\n<li>if the condition doesn't hold - you submit the sample submission file and get zero on Public LB</li>\n<li>if the condition holds - you make the script fail (raise an error). Thus after submitting you'll see an error and conclude that that binary condition holds for the hidden test dataset.\n```</li>\n</ol>\n\n<p>I don't know whether it's cheating... \"gray zone\" as they like to say here on Kaggle. But formally not forbidden. </p>\n\n<p>Feel free to share your insights about the private test set. And don't run out of submissions too early (btw you might not be able with other teams in case of too many submissions).</p>",
  "messages": [
    {
      "id": 674023,
      "postDate": "2019-11-15T20:02:30.460Z",
      "content": "<p>With this competition format it's relatively easy to know more about the (hidden) test dataset.\nIt's explained in <a href=\"https://www.kaggle.com/kashnitsky/one-bit-at-a-time-lb-probing\">this Notebook</a>. </p>\n\n<p>```\nWhen you press \"Commit\", you execute the code with public test dataset. But when you hit \"Submit\", in the background, the same code is run against private test dataset.</p>\n\n<p>Thus we can find out more about the private (hidden) test set. Which actually defines the prizes.</p>\n\n<p>The approach is straightforward:</p>\n\n<ol>\n<li>ask a binary question about the test dataset (eg. whether it's longer than 1000)</li>\n<li>if the condition doesn't hold - you submit the sample submission file and get zero on Public LB</li>\n<li>if the condition holds - you make the script fail (raise an error). Thus after submitting you'll see an error and conclude that that binary condition holds for the hidden test dataset.\n```</li>\n</ol>\n\n<p>I don't know whether it's cheating... \"gray zone\" as they like to say here on Kaggle. But formally not forbidden. </p>\n\n<p>Feel free to share your insights about the private test set. And don't run out of submissions too early (btw you might not be able with other teams in case of too many submissions).</p>",
      "rawMarkdown": "With this competition format it's relatively easy to know more about the (hidden) test dataset.\nIt's explained in [this Notebook](https://www.kaggle.com/kashnitsky/one-bit-at-a-time-lb-probing). \n\n```\nWhen you press \"Commit\", you execute the code with public test dataset. But when you hit \"Submit\", in the background, the same code is run against private test dataset.\n\nThus we can find out more about the private (hidden) test set. Which actually defines the prizes.\n\nThe approach is straightforward:\n\n1. ask a binary question about the test dataset (eg. whether it's longer than 1000)\n2. if the condition doesn't hold - you submit the sample submission file and get zero on Public LB\n3. if the condition holds - you make the script fail (raise an error). Thus after submitting you'll see an error and conclude that that binary condition holds for the hidden test dataset.\n```\n\nI don't know whether it's cheating... \"gray zone\" as they like to say here on Kaggle. But formally not forbidden. \n\nFeel free to share your insights about the private test set. And don't run out of submissions too early (btw you might not be able with other teams in case of too many submissions).",
      "votes": 14
    },
    {
      "id": 674475,
      "postDate": "2019-11-16T14:49:55.523Z",
      "content": "<p>You can extract much more information (multiple bits) than answering a binary question (1 bit) per probe run by exploiting time and binary encoding answers to the questions.</p>",
      "rawMarkdown": "You can extract much more information (multiple bits) than answering a binary question (1 bit) per probe run by exploiting time and binary encoding answers to the questions.",
      "votes": 3,
      "replies": [
        {
          "id": 674507,
          "postDate": "2019-11-16T15:48:54.970Z",
          "content": "<p>Interesting. Would be cool to see an example </p>",
          "rawMarkdown": "Interesting. Would be cool to see an example "
        }
      ]
    },
    {
      "id": 674033,
      "postDate": "2019-11-15T20:24:31.730Z",
      "content": "<p>This makes next 2 months, exploring hidden test set new EDA erra!!. Nice tricks you always surprise me always. Thanks <a href=\"/kashnitsky\">@kashnitsky</a></p>",
      "rawMarkdown": "This makes next 2 months, exploring hidden test set new EDA erra!!. Nice tricks you always surprise me always. Thanks @kashnitsky",
      "votes": 1,
      "replies": [
        {
          "id": 674303,
          "postDate": "2019-11-16T08:09:22.793Z",
          "content": "<p>Oh yes, Probing EDA :) </p>\n\n<p>btw. this probing has nothing to do with that 0.73 Public LB score. In this context, probing only helps to know the private test set better. </p>",
          "rawMarkdown": "Oh yes, Probing EDA :) \n\nbtw. this probing has nothing to do with that 0.73 Public LB score. In this context, probing only helps to know the private test set better. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 675209,
      "postDate": "2019-11-17T18:40:05.010Z",
      "content": "<p>There was a discussion on this in the <a href=\"https://www.kaggle.com/c/instant-gratification/discussion/93080#535910\">Instant gratification competition</a>.</p>\n\n<p>The limited number of submissions (5 per day) helps to  limit its impact. But i'm sure people can do some useful things with it. For example, I think it would be  particularly relevant in time-series analysis one could if assumptions about the data hold in the private test set (e.g., you could test if it was stationary).</p>",
      "rawMarkdown": "There was a discussion on this in the [Instant gratification competition](https://www.kaggle.com/c/instant-gratification/discussion/93080#535910).\n\nThe limited number of submissions (5 per day) helps to  limit its impact. But i'm sure people can do some useful things with it. For example, I think it would be  particularly relevant in time-series analysis one could if assumptions about the data hold in the private test set (e.g., you could test if it was stationary).",
      "votes": 2
    },
    {
      "id": 675620,
      "postDate": "2019-11-18T10:30:50.090Z",
      "content": "<p>wow secrets exposed!</p>",
      "rawMarkdown": "wow secrets exposed!"
    },
    {
      "id": 683230,
      "postDate": "2019-11-28T08:20:27.943Z",
      "content": "<p>Thanks to Yury.  Private dataset size in interval (2500; 3500] examples</p>",
      "rawMarkdown": "Thanks to Yury.  Private dataset size in interval (2500; 3500] examples",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 683342,
          "postDate": "2019-11-28T10:21:23.170Z",
          "content": "<p>Binary search :)</p>",
          "rawMarkdown": "Binary search :)"
        },
        {
          "id": 692434,
          "postDate": "2019-12-11T08:54:07.707Z",
          "content": "<p>Private dataset size in interval [3000;3500] examples\n;)</p>",
          "rawMarkdown": "Private dataset size in interval [3000;3500] examples\n;)",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 674475,
      "author_name": "Yusaku Sako",
      "author_url": "",
      "post_date": "2019-11-16T14:49:55.523000",
      "content": "<p>You can extract much more information (multiple bits) than answering a binary question (1 bit) per probe run by exploiting time and binary encoding answers to the questions.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 674507,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-11-16T15:48:54.970000",
          "content": "<p>Interesting. Would be cool to see an example </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 674033,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2019-11-15T20:24:31.730000",
      "content": "<p>This makes next 2 months, exploring hidden test set new EDA erra!!. Nice tricks you always surprise me always. Thanks <a href=\"/kashnitsky\">@kashnitsky</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 674303,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-11-16T08:09:22.793000",
          "content": "<p>Oh yes, Probing EDA :) </p>\n\n<p>btw. this probing has nothing to do with that 0.73 Public LB score. In this context, probing only helps to know the private test set better. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 675209,
      "author_name": "FChmiel",
      "author_url": "",
      "post_date": "2019-11-17T18:40:05.010000",
      "content": "<p>There was a discussion on this in the <a href=\"https://www.kaggle.com/c/instant-gratification/discussion/93080#535910\">Instant gratification competition</a>.</p>\n\n<p>The limited number of submissions (5 per day) helps to  limit its impact. But i'm sure people can do some useful things with it. For example, I think it would be  particularly relevant in time-series analysis one could if assumptions about the data hold in the private test set (e.g., you could test if it was stationary).</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 675620,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-18T10:30:50.090000",
      "content": "<p>wow secrets exposed!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 683230,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-28T08:20:27.943000",
      "content": "<p>Thanks to Yury.  Private dataset size in interval (2500; 3500] examples</p>",
      "votes": 3,
      "replies": [
        {
          "id": 683342,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-11-28T10:21:23.170000",
          "content": "<p>Binary search :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 692434,
          "author_name": "Rohit Agarwal",
          "author_url": "",
          "post_date": "2019-12-11T08:54:07.707000",
          "content": "<p>Private dataset size in interval [3000;3500] examples\n;)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "674023": "With this competition format it's relatively easy to know more about the (hidden) test dataset.\nIt's explained in [this Notebook](https://www.kaggle.com/kashnitsky/one-bit-at-a-time-lb-probing). \n\n```\nWhen you press \"Commit\", you execute the code with public test dataset. But when you hit \"Submit\", in the background, the same code is run against private test dataset.\n\nThus we can find out more about the private (hidden) test set. Which actually defines the prizes.\n\nThe approach is straightforward:\n\n1. ask a binary question about the test dataset (eg. whether it's longer than 1000)\n2. if the condition doesn't hold - you submit the sample submission file and get zero on Public LB\n3. if the condition holds - you make the script fail (raise an error). Thus after submitting you'll see an error and conclude that that binary condition holds for the hidden test dataset.\n```\n\nI don't know whether it's cheating... \"gray zone\" as they like to say here on Kaggle. But formally not forbidden. \n\nFeel free to share your insights about the private test set. And don't run out of submissions too early (btw you might not be able with other teams in case of too many submissions).",
    "674475": "You can extract much more information (multiple bits) than answering a binary question (1 bit) per probe run by exploiting time and binary encoding answers to the questions.",
    "674033": "This makes next 2 months, exploring hidden test set new EDA erra!!. Nice tricks you always surprise me always. Thanks @kashnitsky",
    "675209": "There was a discussion on this in the [Instant gratification competition](https://www.kaggle.com/c/instant-gratification/discussion/93080#535910).\n\nThe limited number of submissions (5 per day) helps to  limit its impact. But i'm sure people can do some useful things with it. For example, I think it would be  particularly relevant in time-series analysis one could if assumptions about the data hold in the private test set (e.g., you could test if it was stationary).",
    "675620": "wow secrets exposed!",
    "683230": "Thanks to Yury.  Private dataset size in interval (2500; 3500] examples"
  }
}