{
  "id": 183079,
  "title": "Leaderboard Probing ",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/183079",
  "author_name": "",
  "post_date": "2020-09-15T12:20:17.226611700Z",
  "votes": 18,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Leaderboard probing revealed the following.</p>\n<ol>\n<li>there are 188 unique patient id's in the test data.</li>\n<li>There is no common patient id for train data and test data.  </li>\n</ol>\n<p>This notebook is an example of a Leaderboard Probing<br>\n<a href=\"https://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data\" target=\"_blank\">https://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data</a></p>\n<p>Is there anything else that leaderboard proofing can tell us?</p>\n<p>This information from Leaderboard Probing can be used to estimate the amount of time available for image processing and inference and to select a validation strategy.</p>",
  "messages": [
    {
      "id": "1011375",
      "postDate": "09/15/2020 12:20:17",
      "content": "<p>Leaderboard probing revealed the following.</p>\n<ol>\n<li>there are 188 unique patient id's in the test data.</li>\n<li>There is no common patient id for train data and test data.  </li>\n</ol>\n<p>This notebook is an example of a Leaderboard Probing<br>\n<a href=\"https://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data\" target=\"_blank\">https://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data</a></p>\n<p>Is there anything else that leaderboard proofing can tell us?</p>\n<p>This information from Leaderboard Probing can be used to estimate the amount of time available for image processing and inference and to select a validation strategy.</p>",
      "rawMarkdown": "Leaderboard probing revealed the following.\n\n1. there are 188 unique patient id's in the test data.\n2. There is no common patient id for train data and test data.  \n\nThis notebook is an example of a Leaderboard Probing\nhttps://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data\n\nIs there anything else that leaderboard proofing can tell us?\n\nThis information from Leaderboard Probing can be used to estimate the amount of time available for image processing and inference and to select a validation strategy.",
      "votes": null
    },
    {
      "id": "1011393",
      "postDate": "09/15/2020 12:38:32",
      "content": "<p>I guess the obvious question I was asking myself were whether the sample is stratified (e.g. by smoking status and sex - continous factors like age &amp; baseline FVC might be harder to probe for, but could also have been deliberately balanced, or not). Knowing that would indeed help for validation strategy.</p>",
      "rawMarkdown": "I guess the obvious question I was asking myself were whether the sample is stratified (e.g. by smoking status and sex - continous factors like age & baseline FVC might be harder to probe for, but could also have been deliberately balanced, or not). Knowing that would indeed help for validation strategy.",
      "votes": null
    },
    {
      "id": "1012651",
      "postDate": "09/16/2020 07:55:57",
      "content": "<p>Since the number of patients in the test data is 188 and since the last three per patient are evaluated, the total number of data to be evaluated is 564.</p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 15% of the test data. The final results will be based on the other 85%, so the final standings may be different.</p>\n</blockquote>\n<p>And The number of data in PublicLB and PrivateLB is as follows.</p>\n<ul>\n<li>Public:  564 * 15% = 84.6   </li>\n<li>private:   564 * 85% = 479.4</li>\n</ul>\n<p>I don't think it's clear whether PublicLB and PrivateLB data are included in only one of them per Patient or not. </p>",
      "rawMarkdown": "Since the number of patients in the test data is 188 and since the last three per patient are evaluated, the total number of data to be evaluated is 564.\n\n> This leaderboard is calculated with approximately 15% of the test data. The final results will be based on the other 85%, so the final standings may be different.\n\nAnd The number of data in PublicLB and PrivateLB is as follows.\n- Public: ~~846 * 15% = 126.9~~ 564 * 15% = 84.6   \n- private: ~~846 * 85%= 719.1~~  564 * 85% = 479.4\n\nI don't think it's clear whether PublicLB and PrivateLB data are included in only one of them per Patient or not.",
      "votes": null
    },
    {
      "id": "1012848",
      "postDate": "09/16/2020 10:49:53",
      "content": "<p>Dont you mean this?:</p>\n<ul>\n<li>Public: 564*15% </li>\n<li>private: 564*85%</li>\n</ul>",
      "rawMarkdown": "Dont you mean this?:\n\n* Public: 564*15% \n* private: 564*85%",
      "votes": null
    },
    {
      "id": "1012881",
      "postDate": "09/16/2020 11:15:59",
      "content": "<p><a href=\"https://www.kaggle.com/htopper\" target=\"_blank\">@htopper</a> Thanks. I fixed it.</p>",
      "rawMarkdown": "htopper Thanks. I fixed it.",
      "votes": null
    },
    {
      "id": "1013112",
      "postDate": "09/16/2020 14:06:39",
      "content": "<p>It will be interesting to see some samples from public test distribution (features &amp; target), but it seems impossible. Test set is completely hidden, so we have to make probing not only for target, but for features too. The last one requires lots of submissions to figure out the hidden values.</p>",
      "rawMarkdown": "It will be interesting to see some samples from public test distribution (features & target), but it seems impossible. Test set is completely hidden, so we have to make probing not only for target, but for features too. The last one requires lots of submissions to figure out the hidden values.",
      "votes": null
    },
    {
      "id": "1013354",
      "postDate": "09/16/2020 16:52:20",
      "content": "<p>I agree and we can do adversarial validation, like this <a href=\"https://www.kaggle.com/supreethmanyam/adversarial-moa-private-test-included\" target=\"_blank\">notebook</a>.</p>",
      "rawMarkdown": "I agree and we can do adversarial validation, like this [notebook](https://www.kaggle.com/supreethmanyam/adversarial-moa-private-test-included).",
      "votes": null
    },
    {
      "id": "1018735",
      "postDate": "09/19/2020 22:47:41",
      "content": "<p>Perhaps, the number of Female at test.csv is 40.<br>\nBut, It's hard to keep checking…</p>",
      "rawMarkdown": "Perhaps, the number of Female at test.csv is 40.\nBut, It's hard to keep checking...",
      "votes": null
    },
    {
      "id": "1019290",
      "postDate": "09/20/2020 10:31:50",
      "content": "<p><a href=\"https://www.kaggle.com/currypurin\" target=\"_blank\">@currypurin</a>  <a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a> Can you please explain a bit more about <strong>Leaderboard Probing</strong>? Sorry, If its a very basic question. I am New to Kaggle competitions and still learning. </p>",
      "rawMarkdown": "currypurin  @koza4ukdmitrij Can you please explain a bit more about **Leaderboard Probing**? Sorry, If its a very basic question. I am New to Kaggle competitions and still learning.",
      "votes": null
    },
    {
      "id": "1019316",
      "postDate": "09/20/2020 10:54:41",
      "content": "<p>Leaderboard probing is about figuring out what a Public test set labels would be. For instance, if we have regression competition with MAE metric we can submit two submissions:</p>\n<ul>\n<li>the first with zero prediction</li>\n<li>the second with zero prediction except one sample <code>sample0</code> with 1 prediction</li>\n</ul>\n<p>Then you will get two Public LB scores: MAE1 and MAE2. Difference MAE1 - MAE2 consists only information about <code>sample0</code> and with basic transformation you can extract true label for <code>sample0</code>. </p>\n<p>After that, you can extend your training set with such samples <code>sample0</code>. It may be useful cause of:</p>\n<ul>\n<li>you extend training data with samples from new distribution (hope that closer to private test set)</li>\n<li>you get more training examples for better fitting</li>\n</ul>\n<p>That scheme works if you know features for test set, but in this competition you don't so leaderboard probing have to be more sophisticating: you have to find not only true label for <code>sample0</code>, but all features too. It's really hard and in my opinion chances are it doesn't work.</p>\n<p>And, you're welcome to the Kaggle competitions ;)</p>",
      "rawMarkdown": "Leaderboard probing is about figuring out what a Public test set labels would be. For instance, if we have regression competition with MAE metric we can submit two submissions:\n- the first with zero prediction\n- the second with zero prediction except one sample `sample0` with 1 prediction\n\nThen you will get two Public LB scores: MAE1 and MAE2. Difference MAE1 - MAE2 consists only information about `sample0` and with basic transformation you can extract true label for `sample0`. \n\nAfter that, you can extend your training set with such samples `sample0`. It may be useful cause of:\n- you extend training data with samples from new distribution (hope that closer to private test set)\n- you get more training examples for better fitting\n\nThat scheme works if you know features for test set, but in this competition you don't so leaderboard probing have to be more sophisticating: you have to find not only true label for `sample0`, but all features too. It's really hard and in my opinion chances are it doesn't work.\n\nAnd, you're welcome to the Kaggle competitions ;)",
      "votes": null
    },
    {
      "id": "1019984",
      "postDate": "09/20/2020 19:29:59",
      "content": "<p>So we can extract true labels for test data, using leaderboard probing. <br>\nIs this a preferred way of training? or Wouldnt we just overfit our model for test set?  </p>",
      "rawMarkdown": "So we can extract true labels for test data, using leaderboard probing. \nIs this a preferred way of training? or Wouldnt we just overfit our model for test set?",
      "votes": null
    },
    {
      "id": "1031153",
      "postDate": "09/29/2020 09:07:41",
      "content": "<p>good job ! good</p>",
      "rawMarkdown": "good job ! good",
      "votes": null
    },
    {
      "id": "1919632",
      "postDate": "08/30/2022 15:07:00",
      "content": "<p>Very Nice!</p>",
      "rawMarkdown": "Very Nice!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1011393,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "09/15/2020 12:38:32",
      "content": "<p>I guess the obvious question I was asking myself were whether the sample is stratified (e.g. by smoking status and sex - continous factors like age &amp; baseline FVC might be harder to probe for, but could also have been deliberately balanced, or not). Knowing that would indeed help for validation strategy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1018735,
          "author_name": "betweens",
          "author_url": "",
          "post_date": "09/19/2020 22:47:41",
          "content": "<p>Perhaps, the number of Female at test.csv is 40.<br>\nBut, It's hard to keep checking…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1012651,
      "author_name": "currypurin",
      "author_url": "",
      "post_date": "09/16/2020 07:55:57",
      "content": "<p>Since the number of patients in the test data is 188 and since the last three per patient are evaluated, the total number of data to be evaluated is 564.</p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 15% of the test data. The final results will be based on the other 85%, so the final standings may be different.</p>\n</blockquote>\n<p>And The number of data in PublicLB and PrivateLB is as follows.</p>\n<ul>\n<li>Public:  564 * 15% = 84.6   </li>\n<li>private:   564 * 85% = 479.4</li>\n</ul>\n<p>I don't think it's clear whether PublicLB and PrivateLB data are included in only one of them per Patient or not. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1012848,
          "author_name": "htopper",
          "author_url": "",
          "post_date": "09/16/2020 10:49:53",
          "content": "<p>Dont you mean this?:</p>\n<ul>\n<li>Public: 564*15% </li>\n<li>private: 564*85%</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012881,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "09/16/2020 11:15:59",
          "content": "<p><a href=\"https://www.kaggle.com/htopper\" target=\"_blank\">@htopper</a> Thanks. I fixed it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1013112,
      "author_name": "koza4ukdmitrij",
      "author_url": "",
      "post_date": "09/16/2020 14:06:39",
      "content": "<p>It will be interesting to see some samples from public test distribution (features &amp; target), but it seems impossible. Test set is completely hidden, so we have to make probing not only for target, but for features too. The last one requires lots of submissions to figure out the hidden values.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1013354,
          "author_name": "currypurin",
          "author_url": "",
          "post_date": "09/16/2020 16:52:20",
          "content": "<p>I agree and we can do adversarial validation, like this <a href=\"https://www.kaggle.com/supreethmanyam/adversarial-moa-private-test-included\" target=\"_blank\">notebook</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1019290,
      "author_name": "amritpal333",
      "author_url": "",
      "post_date": "09/20/2020 10:31:50",
      "content": "<p><a href=\"https://www.kaggle.com/currypurin\" target=\"_blank\">@currypurin</a>  <a href=\"https://www.kaggle.com/koza4ukdmitrij\" target=\"_blank\">@koza4ukdmitrij</a> Can you please explain a bit more about <strong>Leaderboard Probing</strong>? Sorry, If its a very basic question. I am New to Kaggle competitions and still learning. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1019316,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "09/20/2020 10:54:41",
          "content": "<p>Leaderboard probing is about figuring out what a Public test set labels would be. For instance, if we have regression competition with MAE metric we can submit two submissions:</p>\n<ul>\n<li>the first with zero prediction</li>\n<li>the second with zero prediction except one sample <code>sample0</code> with 1 prediction</li>\n</ul>\n<p>Then you will get two Public LB scores: MAE1 and MAE2. Difference MAE1 - MAE2 consists only information about <code>sample0</code> and with basic transformation you can extract true label for <code>sample0</code>. </p>\n<p>After that, you can extend your training set with such samples <code>sample0</code>. It may be useful cause of:</p>\n<ul>\n<li>you extend training data with samples from new distribution (hope that closer to private test set)</li>\n<li>you get more training examples for better fitting</li>\n</ul>\n<p>That scheme works if you know features for test set, but in this competition you don't so leaderboard probing have to be more sophisticating: you have to find not only true label for <code>sample0</code>, but all features too. It's really hard and in my opinion chances are it doesn't work.</p>\n<p>And, you're welcome to the Kaggle competitions ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1019984,
          "author_name": "amritpal333",
          "author_url": "",
          "post_date": "09/20/2020 19:29:59",
          "content": "<p>So we can extract true labels for test data, using leaderboard probing. <br>\nIs this a preferred way of training? or Wouldnt we just overfit our model for test set?  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1919632,
      "author_name": "torquecoffee",
      "author_url": "",
      "post_date": "08/30/2022 15:07:00",
      "content": "<p>Very Nice!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1031153,
      "author_name": "naim99",
      "author_url": "",
      "post_date": "09/29/2020 09:07:41",
      "content": "<p>good job ! good</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1011375": "Leaderboard probing revealed the following.\n\n1. there are 188 unique patient id's in the test data.\n2. There is no common patient id for train data and test data.  \n\nThis notebook is an example of a Leaderboard Probing\nhttps://www.kaggle.com/currypurin/osic-lb-probing-number-of-patients-in-test-data\n\nIs there anything else that leaderboard proofing can tell us?\n\nThis information from Leaderboard Probing can be used to estimate the amount of time available for image processing and inference and to select a validation strategy.",
    "1011393": "I guess the obvious question I was asking myself were whether the sample is stratified (e.g. by smoking status and sex - continous factors like age & baseline FVC might be harder to probe for, but could also have been deliberately balanced, or not). Knowing that would indeed help for validation strategy.",
    "1012651": "Since the number of patients in the test data is 188 and since the last three per patient are evaluated, the total number of data to be evaluated is 564.\n\n> This leaderboard is calculated with approximately 15% of the test data. The final results will be based on the other 85%, so the final standings may be different.\n\nAnd The number of data in PublicLB and PrivateLB is as follows.\n- Public: ~~846 * 15% = 126.9~~ 564 * 15% = 84.6   \n- private: ~~846 * 85%= 719.1~~  564 * 85% = 479.4\n\nI don't think it's clear whether PublicLB and PrivateLB data are included in only one of them per Patient or not.",
    "1012848": "Dont you mean this?:\n\n* Public: 564*15% \n* private: 564*85%",
    "1012881": "htopper Thanks. I fixed it.",
    "1013112": "It will be interesting to see some samples from public test distribution (features & target), but it seems impossible. Test set is completely hidden, so we have to make probing not only for target, but for features too. The last one requires lots of submissions to figure out the hidden values.",
    "1013354": "I agree and we can do adversarial validation, like this [notebook](https://www.kaggle.com/supreethmanyam/adversarial-moa-private-test-included).",
    "1018735": "Perhaps, the number of Female at test.csv is 40.\nBut, It's hard to keep checking...",
    "1019290": "currypurin  @koza4ukdmitrij Can you please explain a bit more about **Leaderboard Probing**? Sorry, If its a very basic question. I am New to Kaggle competitions and still learning.",
    "1019316": "Leaderboard probing is about figuring out what a Public test set labels would be. For instance, if we have regression competition with MAE metric we can submit two submissions:\n- the first with zero prediction\n- the second with zero prediction except one sample `sample0` with 1 prediction\n\nThen you will get two Public LB scores: MAE1 and MAE2. Difference MAE1 - MAE2 consists only information about `sample0` and with basic transformation you can extract true label for `sample0`. \n\nAfter that, you can extend your training set with such samples `sample0`. It may be useful cause of:\n- you extend training data with samples from new distribution (hope that closer to private test set)\n- you get more training examples for better fitting\n\nThat scheme works if you know features for test set, but in this competition you don't so leaderboard probing have to be more sophisticating: you have to find not only true label for `sample0`, but all features too. It's really hard and in my opinion chances are it doesn't work.\n\nAnd, you're welcome to the Kaggle competitions ;)",
    "1019984": "So we can extract true labels for test data, using leaderboard probing. \nIs this a preferred way of training? or Wouldnt we just overfit our model for test set?",
    "1031153": "good job ! good",
    "1919632": "Very Nice!"
  },
  "source": "meta"
}