{
  "id": 69801,
  "title": "Is Public/Private split random?",
  "url": "/competitions/PLAsTiCC-2018/discussion/69801",
  "author_name": "",
  "post_date": "2018-10-27T13:22:06.051726100Z",
  "votes": 14,
  "comment_count": 10,
  "views": 0,
  "content": "<p>If I remember correctly, the way test data is split was documented in past competitions.  Either random, or time based, or ...</p>\n\n<p>Here we have no information.  Could it be that the split is non random, so that public is not representative of private, as train is not representative of public?</p>",
  "messages": [
    {
      "id": "411123",
      "postDate": "10/27/2018 13:22:06",
      "content": "<p>If I remember correctly, the way test data is split was documented in past competitions.  Either random, or time based, or ...</p>\n\n<p>Here we have no information.  Could it be that the split is non random, so that public is not representative of private, as train is not representative of public?</p>",
      "rawMarkdown": "If I remember correctly, the way test data is split was documented in past competitions.  Either random, or time based, or ...\n\nHere we have no information.  Could it be that the split is non random, so that public is not representative of private, as train is not representative of public?",
      "votes": null
    },
    {
      "id": "411124",
      "postDate": "10/27/2018 13:25:17",
      "content": "<p>Summary of your question: can I probe LB for class_99? :)</p>",
      "rawMarkdown": "Summary of your question: can I probe LB for class_99? :)",
      "votes": null
    },
    {
      "id": "411130",
      "postDate": "10/27/2018 13:32:19",
      "content": "<p>I don't want to experiment what others experimented in TGS: work for a month on a promising idea (mosaic) to see that it was useless on private ;)</p>\n\n<p>I hope that the split is not random.</p>",
      "rawMarkdown": "I don't want to experiment what others experimented in TGS: work for a month on a promising idea (mosaic) to see that it was useless on private ;)\n\nI hope that the split is not random.",
      "votes": null
    },
    {
      "id": "411141",
      "postDate": "10/27/2018 13:41:04",
      "content": "<p>This information is very important from ML view point. My team in TGS competition lost many LB spots just by believing Private/Public was splited randomly and it wasn't.</p>\n\n<p>By the way, my LB probing experiments, up to now, told me that distribution of class_99 Galactic is around 10x smaller than class_99 ExtraGalactic.</p>",
      "rawMarkdown": "This information is very important from ML view point. My team in TGS competition lost many LB spots just by believing Private/Public was splited randomly and it wasn't.\n\nBy the way, my LB probing experiments, up to now, told me that distribution of class_99 Galactic is around 10x smaller than class_99 ExtraGalactic.",
      "votes": null
    },
    {
      "id": "411238",
      "postDate": "10/27/2018 17:01:46",
      "content": "<p>@Giba your finding is supported by probabilities used in this kernel : <a href=\"https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081\">https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081</a></p>",
      "rawMarkdown": "Giba your finding is supported by probabilities used in this kernel : https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081",
      "votes": null
    },
    {
      "id": "412371",
      "postDate": "10/30/2018 05:11:46",
      "content": "<p>Does it really matter if the data is split based on time in this case though? Even ignoring the fact that the data is synthetic, as long as the same object does not appear twice in the split sets, we can safely assume that the property of the universe will not change in a few years and time-based is as good as random...</p>\n\n<p>I think the only major difference between train and test should be that there are no class 99 in train, so the class distribution can be slightly different.</p>",
      "rawMarkdown": "Does it really matter if the data is split based on time in this case though? Even ignoring the fact that the data is synthetic, as long as the same object does not appear twice in the split sets, we can safely assume that the property of the universe will not change in a few years and time-based is as good as random...\n\nI think the only major difference between train and test should be that there are no class 99 in train, so the class distribution can be slightly different.",
      "votes": null
    },
    {
      "id": "431857",
      "postDate": "12/03/2018 02:50:15",
      "content": "<p>Is there any confirmation on how the public and private LB has been done please?</p>",
      "rawMarkdown": "Is there any confirmation on how the public and private LB has been done please?",
      "votes": null
    },
    {
      "id": "432034",
      "postDate": "12/03/2018 09:42:17",
      "content": "<p>I don't think they will give us this information but if I had to guess I would say it would be random because they have nothing to gain from making it non-random. The train data is not representative of public test in order to simulate the real LSST conditions, where much more and fainter/further/unknown object data will be obtained than what we currently have. I don't think we expect to get yet another set of very different data after LSST has been running for a while. At the worst it would be time-based but (as Mithrillion mentioned) it shouldn't matter.</p>",
      "rawMarkdown": "I don't think they will give us this information but if I had to guess I would say it would be random because they have nothing to gain from making it non-random. The train data is not representative of public test in order to simulate the real LSST conditions, where much more and fainter/further/unknown object data will be obtained than what we currently have. I don't think we expect to get yet another set of very different data after LSST has been running for a while. At the worst it would be time-based but (as Mithrillion mentioned) it shouldn't matter.",
      "votes": null
    },
    {
      "id": "432059",
      "postDate": "12/03/2018 10:41:34",
      "content": "<p>hello, thanks for your inputs. I just wanted to know, if public and private test sets have similar distribution for classes.  </p>",
      "rawMarkdown": "hello, thanks for your inputs. I just wanted to know, if public and private test sets have similar distribution for classes.",
      "votes": null
    },
    {
      "id": "436067",
      "postDate": "12/09/2018 13:58:03",
      "content": "<p>Good question.</p>",
      "rawMarkdown": "Good question.",
      "votes": null
    },
    {
      "id": "439270",
      "postDate": "12/15/2018 04:07:12",
      "content": "<p>I need to raise this question again. Could we get some response from the host? @Sohier Dane</p>",
      "rawMarkdown": "I need to raise this question again. Could we get some response from the host? @Sohier Dane",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 411124,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "10/27/2018 13:25:17",
      "content": "<p>Summary of your question: can I probe LB for class_99? :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 411130,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "10/27/2018 13:32:19",
          "content": "<p>I don't want to experiment what others experimented in TGS: work for a month on a promising idea (mosaic) to see that it was useless on private ;)</p>\n\n<p>I hope that the split is not random.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 411141,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "10/27/2018 13:41:04",
      "content": "<p>This information is very important from ML view point. My team in TGS competition lost many LB spots just by believing Private/Public was splited randomly and it wasn't.</p>\n\n<p>By the way, my LB probing experiments, up to now, told me that distribution of class_99 Galactic is around 10x smaller than class_99 ExtraGalactic.</p>",
      "votes": null,
      "replies": [
        {
          "id": 411238,
          "author_name": "ogrellier",
          "author_url": "",
          "post_date": "10/27/2018 17:01:46",
          "content": "<p>@Giba your finding is supported by probabilities used in this kernel : <a href=\"https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081\">https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 412371,
      "author_name": "mithrillion",
      "author_url": "",
      "post_date": "10/30/2018 05:11:46",
      "content": "<p>Does it really matter if the data is split based on time in this case though? Even ignoring the fact that the data is synthetic, as long as the same object does not appear twice in the split sets, we can safely assume that the property of the universe will not change in a few years and time-based is as good as random...</p>\n\n<p>I think the only major difference between train and test should be that there are no class 99 in train, so the class distribution can be slightly different.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 431857,
      "author_name": "chinta",
      "author_url": "",
      "post_date": "12/03/2018 02:50:15",
      "content": "<p>Is there any confirmation on how the public and private LB has been done please?</p>",
      "votes": null,
      "replies": [
        {
          "id": 432034,
          "author_name": "taniaj",
          "author_url": "",
          "post_date": "12/03/2018 09:42:17",
          "content": "<p>I don't think they will give us this information but if I had to guess I would say it would be random because they have nothing to gain from making it non-random. The train data is not representative of public test in order to simulate the real LSST conditions, where much more and fainter/further/unknown object data will be obtained than what we currently have. I don't think we expect to get yet another set of very different data after LSST has been running for a while. At the worst it would be time-based but (as Mithrillion mentioned) it shouldn't matter.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 432059,
          "author_name": "chinta",
          "author_url": "",
          "post_date": "12/03/2018 10:41:34",
          "content": "<p>hello, thanks for your inputs. I just wanted to know, if public and private test sets have similar distribution for classes.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 436067,
      "author_name": "longyin2",
      "author_url": "",
      "post_date": "12/09/2018 13:58:03",
      "content": "<p>Good question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 439270,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "12/15/2018 04:07:12",
      "content": "<p>I need to raise this question again. Could we get some response from the host? @Sohier Dane</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "411123": "If I remember correctly, the way test data is split was documented in past competitions.  Either random, or time based, or ...\n\nHere we have no information.  Could it be that the split is non random, so that public is not representative of private, as train is not representative of public?",
    "411124": "Summary of your question: can I probe LB for class_99? :)",
    "411130": "I don't want to experiment what others experimented in TGS: work for a month on a promising idea (mosaic) to see that it was useless on private ;)\n\nI hope that the split is not random.",
    "411141": "This information is very important from ML view point. My team in TGS competition lost many LB spots just by believing Private/Public was splited randomly and it wasn't.\n\nBy the way, my LB probing experiments, up to now, told me that distribution of class_99 Galactic is around 10x smaller than class_99 ExtraGalactic.",
    "411238": "Giba your finding is supported by probabilities used in this kernel : https://www.kaggle.com/darbin/weighted-naive-benchmark-lb-2-081",
    "412371": "Does it really matter if the data is split based on time in this case though? Even ignoring the fact that the data is synthetic, as long as the same object does not appear twice in the split sets, we can safely assume that the property of the universe will not change in a few years and time-based is as good as random...\n\nI think the only major difference between train and test should be that there are no class 99 in train, so the class distribution can be slightly different.",
    "431857": "Is there any confirmation on how the public and private LB has been done please?",
    "432034": "I don't think they will give us this information but if I had to guess I would say it would be random because they have nothing to gain from making it non-random. The train data is not representative of public test in order to simulate the real LSST conditions, where much more and fainter/further/unknown object data will be obtained than what we currently have. I don't think we expect to get yet another set of very different data after LSST has been running for a while. At the worst it would be time-based but (as Mithrillion mentioned) it shouldn't matter.",
    "432059": "hello, thanks for your inputs. I just wanted to know, if public and private test sets have similar distribution for classes.",
    "436067": "Good question.",
    "439270": "I need to raise this question again. Could we get some response from the host? @Sohier Dane"
  },
  "source": "meta"
}