{
  "id": 105763,
  "title": "Private LB probing - there might be no shakeup",
  "url": "/competitions/aptos2019-blindness-detection/discussion/105763",
  "author_name": "",
  "post_date": "2019-08-26T08:32:18.994813700Z",
  "votes": 25,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I probed the private LB to determine that the private test set is quite similar public test set! Here's how I did it (kernel available <a href=\"https://www.kaggle.com/tanlikesmath/private-lb-probing-adversarial-validation\">here!</a>):</p>\n\n<h1>Principles of probing private dataset</h1>\n\n<h2>for synchronous KO competitions</h2>\n\n<p>Private dataset probing of synchronous KO competitions has been done before. Chris Deotte has nicely outlined how to do this in this forum post over <a href=\"https://www.kaggle.com/c/instant-gratification/discussion/93080\">here</a>\nThe idea is that information can be returned via:\n1. altering submission.csv and consequently altering public LB score\n2. altering execution time\n3. choosing to return an error\nNow, option 1 can be done by either submitting a null submission (LB=0) or submitting an actual submission (LB != 0). This provides a single level of information. Option 3 also provides another level of information. If the code in your kernel runs pretty fast, option 2 can provide a lot of information. In case there is variability, I thought about having a resolution of every 10 minutes for 90 minutes, so there's a whopping 90 levels. In total, there's approximately 92  levels of information! And this is just a low value, as there are 9 hours, not 90 minutes available for runtime. However, I won't need such high level of detail for now.</p>\n\n<h1>Adversarial validation</h1>\n\n<p>With this knowledge, I wanted to gather information about the private dataset. In particular, it would be helpful to know how similar the private test set is to the public test set. This will indicate if there could be a significant shake-up in the competition or not. That is exactly what adversarial validation does.\nThe idea behind adversarial validation is that we can train a classifer to classify between our training and testing set (or in this case both testing sets) and if the AUC is close to 0.5, which indicates the classifier is performing no better than random, then the datasets probably come from identical distributions.\n<a href=\"/konradb\">@konradb</a> provided a great kernel for adversarial validation which I will adapt for my case. Therefore, my goal of this experiment is:\n<strong>Goal: To perform adversarial validation between the public and private testing set</strong>\nI wanted to obtain an approximation for AUC score. The training of a model to distinguish between private and public test sets takes 2 min to train. Therefore, we can have levels at every 10 min with my kernel running a maximum of 90 minutes. I could have higher resolution, but this close to the end of the competition, I wanted to focus on my actual submissions.</p>\n\n<h1>Results</h1>\n\n<p>Running the kernel on the private dataset (by submitting it), it took 51 min to run. Based on how I set up the code, this corresponds to an AUC between 0.5 and 0.6. This means that the private dataset is quite similar to the public dataset. This indicates the LB  score is actually a good representation of the private LB score.\n<strong>Conclusion: The private LB is probably going to be similar to the public LB</strong></p>\n\n<h2>What's next</h2>\n\n<p>First,  I would appreciate if the community checks <a href=\"https://www.kaggle.com/tanlikesmath/private-lb-probing-adversarial-validation\">my code</a> and verify my conclusion. Also, if you find any flaw in any of my reasoning, please don't hesitate to let me know! </p>\n\n<p>Second, there are so many ways to probe the LB. However, I am not entirely sure what other things to probe. The question of public vs. private seemed to be the most important. So I hope maybe you guys could find some other useful information about the private dataset, and share with us!</p>",
  "messages": [
    {
      "id": "608034",
      "postDate": "08/26/2019 08:32:18",
      "content": "<p>I probed the private LB to determine that the private test set is quite similar public test set! Here's how I did it (kernel available <a href=\"https://www.kaggle.com/tanlikesmath/private-lb-probing-adversarial-validation\">here!</a>):</p>\n\n<h1>Principles of probing private dataset</h1>\n\n<h2>for synchronous KO competitions</h2>\n\n<p>Private dataset probing of synchronous KO competitions has been done before. Chris Deotte has nicely outlined how to do this in this forum post over <a href=\"https://www.kaggle.com/c/instant-gratification/discussion/93080\">here</a>\nThe idea is that information can be returned via:\n1. altering submission.csv and consequently altering public LB score\n2. altering execution time\n3. choosing to return an error\nNow, option 1 can be done by either submitting a null submission (LB=0) or submitting an actual submission (LB != 0). This provides a single level of information. Option 3 also provides another level of information. If the code in your kernel runs pretty fast, option 2 can provide a lot of information. In case there is variability, I thought about having a resolution of every 10 minutes for 90 minutes, so there's a whopping 90 levels. In total, there's approximately 92  levels of information! And this is just a low value, as there are 9 hours, not 90 minutes available for runtime. However, I won't need such high level of detail for now.</p>\n\n<h1>Adversarial validation</h1>\n\n<p>With this knowledge, I wanted to gather information about the private dataset. In particular, it would be helpful to know how similar the private test set is to the public test set. This will indicate if there could be a significant shake-up in the competition or not. That is exactly what adversarial validation does.\nThe idea behind adversarial validation is that we can train a classifer to classify between our training and testing set (or in this case both testing sets) and if the AUC is close to 0.5, which indicates the classifier is performing no better than random, then the datasets probably come from identical distributions.\n<a href=\"/konradb\">@konradb</a> provided a great kernel for adversarial validation which I will adapt for my case. Therefore, my goal of this experiment is:\n<strong>Goal: To perform adversarial validation between the public and private testing set</strong>\nI wanted to obtain an approximation for AUC score. The training of a model to distinguish between private and public test sets takes 2 min to train. Therefore, we can have levels at every 10 min with my kernel running a maximum of 90 minutes. I could have higher resolution, but this close to the end of the competition, I wanted to focus on my actual submissions.</p>\n\n<h1>Results</h1>\n\n<p>Running the kernel on the private dataset (by submitting it), it took 51 min to run. Based on how I set up the code, this corresponds to an AUC between 0.5 and 0.6. This means that the private dataset is quite similar to the public dataset. This indicates the LB  score is actually a good representation of the private LB score.\n<strong>Conclusion: The private LB is probably going to be similar to the public LB</strong></p>\n\n<h2>What's next</h2>\n\n<p>First,  I would appreciate if the community checks <a href=\"https://www.kaggle.com/tanlikesmath/private-lb-probing-adversarial-validation\">my code</a> and verify my conclusion. Also, if you find any flaw in any of my reasoning, please don't hesitate to let me know! </p>\n\n<p>Second, there are so many ways to probe the LB. However, I am not entirely sure what other things to probe. The question of public vs. private seemed to be the most important. So I hope maybe you guys could find some other useful information about the private dataset, and share with us!</p>",
      "rawMarkdown": "I probed the private LB to determine that the private test set is quite similar public test set! Here's how I did it (kernel available [here!](https://www.kaggle.com/tanlikesmath/private-lb-probing-adversarial-validation)):\n# Principles of probing private dataset\n## for synchronous KO competitions\nPrivate dataset probing of synchronous KO competitions has been done before. Chris Deotte has nicely outlined how to do this in this forum post over [here](https://www.kaggle.com/c/instant-gratification/discussion/93080)\nThe idea is that information can be returned via:\n1. altering submission.csv and consequently altering public LB score\n2. altering execution time\n3. choosing to return an error\nNow, option 1 can be done by either submitting a null submission (LB=0) or submitting an actual submission (LB != 0). This provides a single level of information. Option 3 also provides another level of information. If the code in your kernel runs pretty fast, option 2 can provide a lot of information. In case there is variability, I thought about having a resolution of every 10 minutes for 90 minutes, so there's a whopping 90 levels. In total, there's approximately 92  levels of information! And this is just a low value, as there are 9 hours, not 90 minutes available for runtime. However, I won't need such high level of detail for now.\n# Adversarial validation\nWith this knowledge, I wanted to gather information about the private dataset. In particular, it would be helpful to know how similar the private test set is to the public test set. This will indicate if there could be a significant shake-up in the competition or not. That is exactly what adversarial validation does.\nThe idea behind adversarial validation is that we can train a classifer to classify between our training and testing set (or in this case both testing sets) and if the AUC is close to 0.5, which indicates the classifier is performing no better than random, then the datasets probably come from identical distributions.\n@konradb provided a great kernel for adversarial validation which I will adapt for my case. Therefore, my goal of this experiment is:\n**Goal: To perform adversarial validation between the public and private testing set**\nI wanted to obtain an approximation for AUC score. The training of a model to distinguish between private and public test sets takes 2 min to train. Therefore, we can have levels at every 10 min with my kernel running a maximum of 90 minutes. I could have higher resolution, but this close to the end of the competition, I wanted to focus on my actual submissions.\n# Results\nRunning the kernel on the private dataset (by submitting it), it took 51 min to run. Based on how I set up the code, this corresponds to an AUC between 0.5 and 0.6. This means that the private dataset is quite similar to the public dataset. This indicates the LB  score is actually a good representation of the private LB score.\n**Conclusion: The private LB is probably going to be similar to the public LB**\n\n## What's next\nFirst,  I would appreciate if the community checks [my code](https://www.kaggle.com/tanlikesmath/private-lb-probing-adversarial-validation) and verify my conclusion. Also, if you find any flaw in any of my reasoning, please don't hesitate to let me know! \n\nSecond, there are so many ways to probe the LB. However, I am not entirely sure what other things to probe. The question of public vs. private seemed to be the most important. So I hope maybe you guys could find some other useful information about the private dataset, and share with us!",
      "votes": null
    },
    {
      "id": "608116",
      "postDate": "08/26/2019 11:06:38",
      "content": "<p>This is interesting. I am wondering how can this classifier be trained within just 2 minutes? (or perhaps I misunderstood your method)</p>",
      "rawMarkdown": "This is interesting. I am wondering how can this classifier be trained within just 2 minutes? (or perhaps I misunderstood your method)",
      "votes": null
    },
    {
      "id": "608117",
      "postDate": "08/26/2019 11:08:15",
      "content": "<p>I used Konrad's kernel but just did a single fold rather than 5-folds. Otherwise it takes ~10 minutes with 5 folds.</p>",
      "rawMarkdown": "I used Konrad's kernel but just did a single fold rather than 5-folds. Otherwise it takes ~10 minutes with 5 folds.",
      "votes": null
    },
    {
      "id": "608119",
      "postDate": "08/26/2019 11:11:07",
      "content": "<p>Here's the <a href=\"https://www.kaggle.com/konradb/adversarial-validation-quick-fast-ai-approach\">kernel</a></p>",
      "rawMarkdown": "Here's the [kernel](https://www.kaggle.com/konradb/adversarial-validation-quick-fast-ai-approach)",
      "votes": null
    },
    {
      "id": "608123",
      "postDate": "08/26/2019 11:18:33",
      "content": "<p>Great work!!\nI think we also need to know the class histogram of private test dataset in order to know how much SHAKE will be occurred.\nThe difference of class histogram also affects QWK score.</p>",
      "rawMarkdown": "Great work!!\nI think we also need to know the class histogram of private test dataset in order to know how much SHAKE will be occurred.\nThe difference of class histogram also affects QWK score.",
      "votes": null
    },
    {
      "id": "608191",
      "postDate": "08/26/2019 13:22:29",
      "content": "<p>Nice and i hope you'll make it. \nBut there is sth i don't understand. How can you perform a model to distinguish between private and public test sets ?.  But it's a kernel competition and we can only use the public test. You can do it with the training and the test sets if there's discrepancy between CV and LB public scores.</p>",
      "rawMarkdown": "Nice and i hope you'll make it. \nBut there is sth i don't understand. How can you perform a model to distinguish between private and public test sets ?.  But it's a kernel competition and we can only use the public test. You can do it with the training and the test sets if there's discrepancy between CV and LB public scores.",
      "votes": null
    },
    {
      "id": "608241",
      "postDate": "08/26/2019 14:32:32",
      "content": "<p>Nice probing. Thanks!</p>",
      "rawMarkdown": "Nice probing. Thanks!",
      "votes": null
    },
    {
      "id": "608424",
      "postDate": "08/26/2019 18:36:58",
      "content": "<p>When you submit a kernel, it also runs on the private test set. This is a feature of the synchronous KO competition. </p>",
      "rawMarkdown": "When you submit a kernel, it also runs on the private test set. This is a feature of the synchronous KO competition.",
      "votes": null
    },
    {
      "id": "608543",
      "postDate": "08/26/2019 22:54:04",
      "content": "<p>Thanks for such an interesting work! It's my first time seeing LB probing being used for something more worth while  :)</p>",
      "rawMarkdown": "Thanks for such an interesting work! It's my first time seeing LB probing being used for something more worth while  :)",
      "votes": null
    },
    {
      "id": "608959",
      "postDate": "08/27/2019 09:40:30",
      "content": "<p>I don't quite sure whether I have got your point, but I should point out that, when you click <code>submit</code> bottom and wait for your LB score, the kernel is just running on exactly the same images as your public test set. The difference of the running time is just because Kaggle have duplicated those images to around 6x~ when evaluating LB score for ensuring your 9h time limitation.</p>\n\n<p>So I don't think you can see any information of private test set before the game end.</p>",
      "rawMarkdown": "I don't quite sure whether I have got your point, but I should point out that, when you click `submit` bottom and wait for your LB score, the kernel is just running on exactly the same images as your public test set. The difference of the running time is just because Kaggle have duplicated those images to around 6x~ when evaluating LB score for ensuring your 9h time limitation.\n\nSo I don't think you can see any information of private test set before the game end.",
      "votes": null
    },
    {
      "id": "608976",
      "postDate": "08/27/2019 10:06:08",
      "content": "<p>From <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/overview/kernels-requirements\">this page</a>:</p>\n\n<blockquote>\n  <p>Your kernel will re-run automatically against an unseen test set</p>\n</blockquote>\n\n<p>The synchronous KO competitions run the kernel also on the private set, so the results can be released immediately after the comp. end. instead of rerunning the kernels. So unless there is a change of rules I don't know about, the longer submission times are definitely due to running on the private set.</p>",
      "rawMarkdown": "From [this page](https://www.kaggle.com/c/aptos2019-blindness-detection/overview/kernels-requirements):\n\n&gt; Your kernel will re-run automatically against an unseen test set\n\nThe synchronous KO competitions run the kernel also on the private set, so the results can be released immediately after the comp. end. instead of rerunning the kernels. So unless there is a change of rules I don't know about, the longer submission times are definitely due to running on the private set.",
      "votes": null
    },
    {
      "id": "608978",
      "postDate": "08/27/2019 10:07:16",
      "content": "<p>Unfortunately this cannot be done, as there is no way of obtaining information regarding the labels of the private test set AFAIK.</p>",
      "rawMarkdown": "Unfortunately this cannot be done, as there is no way of obtaining information regarding the labels of the private test set AFAIK.",
      "votes": null
    },
    {
      "id": "608990",
      "postDate": "08/27/2019 10:17:14",
      "content": "<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/102718\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/102718</a></p>\n\n<p>You should try this, to find out why I said that they are just duplicating the public test set</p>",
      "rawMarkdown": "https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/102718\n\nYou should try this, to find out why I said that they are just duplicating the public test set",
      "votes": null
    },
    {
      "id": "609053",
      "postDate": "08/27/2019 11:32:11",
      "content": "<p>I think the inference result of a well-trained model(over 0.8LB score) is similar to a real class histogram.</p>",
      "rawMarkdown": "I think the inference result of a well-trained model(over 0.8LB score) is similar to a real class histogram.",
      "votes": null
    },
    {
      "id": "609084",
      "postDate": "08/27/2019 12:00:34",
      "content": "<p>I'm afraid that the kernel tried to distinguish \"public test set\" and \"public+private test set\" because, in my understanding, the test.csv in the second execution includes both public and private test set.</p>",
      "rawMarkdown": "I'm afraid that the kernel tried to distinguish \"public test set\" and \"public+private test set\" because, in my understanding, the test.csv in the second execution includes both public and private test set.",
      "votes": null
    },
    {
      "id": "609593",
      "postDate": "08/27/2019 22:20:17",
      "content": "<p>I thought private LB is only on the private data? Not combined...</p>\n\n<blockquote>\n  <p>The final results will be based on the other 85%, so the final standings may be different.\n  -from the leaderboard description.</p>\n</blockquote>",
      "rawMarkdown": "I thought private LB is only on the private data? Not combined...\n&gt; The final results will be based on the other 85%, so the final standings may be different.\n-from the leaderboard description.",
      "votes": null
    },
    {
      "id": "609618",
      "postDate": "08/27/2019 23:37:21",
      "content": "<p>thanks for that</p>",
      "rawMarkdown": "thanks for that",
      "votes": null
    },
    {
      "id": "609621",
      "postDate": "08/27/2019 23:44:57",
      "content": "<p>Interesting... well I don't have such a model yet lol but when I do I will look into this...</p>",
      "rawMarkdown": "Interesting... well I don't have such a model yet lol but when I do I will look into this...",
      "votes": null
    },
    {
      "id": "622852",
      "postDate": "09/10/2019 06:54:27",
      "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a>\nIt finally turns out that submitting kernels are running on the private set. Sorry for that.\nBut it also turns out that the information of private test set is invisible before the competition end ;)</p>\n\n<p>And congrats to your silver medal!</p>",
      "rawMarkdown": "tanlikesmath\nIt finally turns out that submitting kernels are running on the private set. Sorry for that.\nBut it also turns out that the information of private test set is invisible before the competition end ;)\n\nAnd congrats to your silver medal!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 608116,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "08/26/2019 11:06:38",
      "content": "<p>This is interesting. I am wondering how can this classifier be trained within just 2 minutes? (or perhaps I misunderstood your method)</p>",
      "votes": null,
      "replies": [
        {
          "id": 608117,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/26/2019 11:08:15",
          "content": "<p>I used Konrad's kernel but just did a single fold rather than 5-folds. Otherwise it takes ~10 minutes with 5 folds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608119,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/26/2019 11:11:07",
          "content": "<p>Here's the <a href=\"https://www.kaggle.com/konradb/adversarial-validation-quick-fast-ai-approach\">kernel</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609618,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/27/2019 23:37:21",
          "content": "<p>thanks for that</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 608123,
      "author_name": "hirune924",
      "author_url": "",
      "post_date": "08/26/2019 11:18:33",
      "content": "<p>Great work!!\nI think we also need to know the class histogram of private test dataset in order to know how much SHAKE will be occurred.\nThe difference of class histogram also affects QWK score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 608978,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/27/2019 10:07:16",
          "content": "<p>Unfortunately this cannot be done, as there is no way of obtaining information regarding the labels of the private test set AFAIK.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609053,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "08/27/2019 11:32:11",
          "content": "<p>I think the inference result of a well-trained model(over 0.8LB score) is similar to a real class histogram.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609621,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/27/2019 23:44:57",
          "content": "<p>Interesting... well I don't have such a model yet lol but when I do I will look into this...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 608191,
      "author_name": "seif95",
      "author_url": "",
      "post_date": "08/26/2019 13:22:29",
      "content": "<p>Nice and i hope you'll make it. \nBut there is sth i don't understand. How can you perform a model to distinguish between private and public test sets ?.  But it's a kernel competition and we can only use the public test. You can do it with the training and the test sets if there's discrepancy between CV and LB public scores.</p>",
      "votes": null,
      "replies": [
        {
          "id": 608424,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/26/2019 18:36:58",
          "content": "<p>When you submit a kernel, it also runs on the private test set. This is a feature of the synchronous KO competition. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 608241,
      "author_name": "songwonho",
      "author_url": "",
      "post_date": "08/26/2019 14:32:32",
      "content": "<p>Nice probing. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 608543,
      "author_name": "joonl04",
      "author_url": "",
      "post_date": "08/26/2019 22:54:04",
      "content": "<p>Thanks for such an interesting work! It's my first time seeing LB probing being used for something more worth while  :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 608959,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "08/27/2019 09:40:30",
      "content": "<p>I don't quite sure whether I have got your point, but I should point out that, when you click <code>submit</code> bottom and wait for your LB score, the kernel is just running on exactly the same images as your public test set. The difference of the running time is just because Kaggle have duplicated those images to around 6x~ when evaluating LB score for ensuring your 9h time limitation.</p>\n\n<p>So I don't think you can see any information of private test set before the game end.</p>",
      "votes": null,
      "replies": [
        {
          "id": 608976,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/27/2019 10:06:08",
          "content": "<p>From <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/overview/kernels-requirements\">this page</a>:</p>\n\n<blockquote>\n  <p>Your kernel will re-run automatically against an unseen test set</p>\n</blockquote>\n\n<p>The synchronous KO competitions run the kernel also on the private set, so the results can be released immediately after the comp. end. instead of rerunning the kernels. So unless there is a change of rules I don't know about, the longer submission times are definitely due to running on the private set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608990,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "08/27/2019 10:17:14",
          "content": "<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/102718\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/102718</a></p>\n\n<p>You should try this, to find out why I said that they are just duplicating the public test set</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 622852,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "09/10/2019 06:54:27",
          "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a>\nIt finally turns out that submitting kernels are running on the private set. Sorry for that.\nBut it also turns out that the information of private test set is invisible before the competition end ;)</p>\n\n<p>And congrats to your silver medal!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 609084,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "08/27/2019 12:00:34",
      "content": "<p>I'm afraid that the kernel tried to distinguish \"public test set\" and \"public+private test set\" because, in my understanding, the test.csv in the second execution includes both public and private test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 609593,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "08/27/2019 22:20:17",
          "content": "<p>I thought private LB is only on the private data? Not combined...</p>\n\n<blockquote>\n  <p>The final results will be based on the other 85%, so the final standings may be different.\n  -from the leaderboard description.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "608034": "I probed the private LB to determine that the private test set is quite similar public test set! Here's how I did it (kernel available [here!](https://www.kaggle.com/tanlikesmath/private-lb-probing-adversarial-validation)):\n# Principles of probing private dataset\n## for synchronous KO competitions\nPrivate dataset probing of synchronous KO competitions has been done before. Chris Deotte has nicely outlined how to do this in this forum post over [here](https://www.kaggle.com/c/instant-gratification/discussion/93080)\nThe idea is that information can be returned via:\n1. altering submission.csv and consequently altering public LB score\n2. altering execution time\n3. choosing to return an error\nNow, option 1 can be done by either submitting a null submission (LB=0) or submitting an actual submission (LB != 0). This provides a single level of information. Option 3 also provides another level of information. If the code in your kernel runs pretty fast, option 2 can provide a lot of information. In case there is variability, I thought about having a resolution of every 10 minutes for 90 minutes, so there's a whopping 90 levels. In total, there's approximately 92  levels of information! And this is just a low value, as there are 9 hours, not 90 minutes available for runtime. However, I won't need such high level of detail for now.\n# Adversarial validation\nWith this knowledge, I wanted to gather information about the private dataset. In particular, it would be helpful to know how similar the private test set is to the public test set. This will indicate if there could be a significant shake-up in the competition or not. That is exactly what adversarial validation does.\nThe idea behind adversarial validation is that we can train a classifer to classify between our training and testing set (or in this case both testing sets) and if the AUC is close to 0.5, which indicates the classifier is performing no better than random, then the datasets probably come from identical distributions.\n@konradb provided a great kernel for adversarial validation which I will adapt for my case. Therefore, my goal of this experiment is:\n**Goal: To perform adversarial validation between the public and private testing set**\nI wanted to obtain an approximation for AUC score. The training of a model to distinguish between private and public test sets takes 2 min to train. Therefore, we can have levels at every 10 min with my kernel running a maximum of 90 minutes. I could have higher resolution, but this close to the end of the competition, I wanted to focus on my actual submissions.\n# Results\nRunning the kernel on the private dataset (by submitting it), it took 51 min to run. Based on how I set up the code, this corresponds to an AUC between 0.5 and 0.6. This means that the private dataset is quite similar to the public dataset. This indicates the LB  score is actually a good representation of the private LB score.\n**Conclusion: The private LB is probably going to be similar to the public LB**\n\n## What's next\nFirst,  I would appreciate if the community checks [my code](https://www.kaggle.com/tanlikesmath/private-lb-probing-adversarial-validation) and verify my conclusion. Also, if you find any flaw in any of my reasoning, please don't hesitate to let me know! \n\nSecond, there are so many ways to probe the LB. However, I am not entirely sure what other things to probe. The question of public vs. private seemed to be the most important. So I hope maybe you guys could find some other useful information about the private dataset, and share with us!",
    "608116": "This is interesting. I am wondering how can this classifier be trained within just 2 minutes? (or perhaps I misunderstood your method)",
    "608117": "I used Konrad's kernel but just did a single fold rather than 5-folds. Otherwise it takes ~10 minutes with 5 folds.",
    "608119": "Here's the [kernel](https://www.kaggle.com/konradb/adversarial-validation-quick-fast-ai-approach)",
    "608123": "Great work!!\nI think we also need to know the class histogram of private test dataset in order to know how much SHAKE will be occurred.\nThe difference of class histogram also affects QWK score.",
    "608191": "Nice and i hope you'll make it. \nBut there is sth i don't understand. How can you perform a model to distinguish between private and public test sets ?.  But it's a kernel competition and we can only use the public test. You can do it with the training and the test sets if there's discrepancy between CV and LB public scores.",
    "608241": "Nice probing. Thanks!",
    "608424": "When you submit a kernel, it also runs on the private test set. This is a feature of the synchronous KO competition.",
    "608543": "Thanks for such an interesting work! It's my first time seeing LB probing being used for something more worth while  :)",
    "608959": "I don't quite sure whether I have got your point, but I should point out that, when you click `submit` bottom and wait for your LB score, the kernel is just running on exactly the same images as your public test set. The difference of the running time is just because Kaggle have duplicated those images to around 6x~ when evaluating LB score for ensuring your 9h time limitation.\n\nSo I don't think you can see any information of private test set before the game end.",
    "608976": "From [this page](https://www.kaggle.com/c/aptos2019-blindness-detection/overview/kernels-requirements):\n\n&gt; Your kernel will re-run automatically against an unseen test set\n\nThe synchronous KO competitions run the kernel also on the private set, so the results can be released immediately after the comp. end. instead of rerunning the kernels. So unless there is a change of rules I don't know about, the longer submission times are definitely due to running on the private set.",
    "608978": "Unfortunately this cannot be done, as there is no way of obtaining information regarding the labels of the private test set AFAIK.",
    "608990": "https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/102718\n\nYou should try this, to find out why I said that they are just duplicating the public test set",
    "609053": "I think the inference result of a well-trained model(over 0.8LB score) is similar to a real class histogram.",
    "609084": "I'm afraid that the kernel tried to distinguish \"public test set\" and \"public+private test set\" because, in my understanding, the test.csv in the second execution includes both public and private test set.",
    "609593": "I thought private LB is only on the private data? Not combined...\n&gt; The final results will be based on the other 85%, so the final standings may be different.\n-from the leaderboard description.",
    "609618": "thanks for that",
    "609621": "Interesting... well I don't have such a model yet lol but when I do I will look into this...",
    "622852": "tanlikesmath\nIt finally turns out that submitting kernels are running on the private set. Sorry for that.\nBut it also turns out that the information of private test set is invisible before the competition end ;)\n\nAnd congrats to your silver medal!"
  },
  "source": "meta"
}