{
  "id": 378151,
  "title": "The way to split test dataset into public/private",
  "url": "/competitions/otto-recommender-system/discussion/378151",
  "author_name": "",
  "post_date": "2023-01-14T14:43:40.743781300Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thankfully, competition host shared the code to split train/test dataset.<br>\n<a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py</a></p>\n<p>But, the way  to split test dataset into public/private is not shared.<br>\nIf it is splited by the specific order  (e.g. session-ID order, time-series order), the risk of overfitting with public testset would become larger. </p>\n<p>How do you think about this issue ?<br>\nDoes anyone verify the way to split by using public LB ?</p>",
  "messages": [
    {
      "id": "2099595",
      "postDate": "01/14/2023 14:43:40",
      "content": "<p>Thankfully, competition host shared the code to split train/test dataset.<br>\n<a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py</a></p>\n<p>But, the way  to split test dataset into public/private is not shared.<br>\nIf it is splited by the specific order  (e.g. session-ID order, time-series order), the risk of overfitting with public testset would become larger. </p>\n<p>How do you think about this issue ?<br>\nDoes anyone verify the way to split by using public LB ?</p>",
      "rawMarkdown": "Thankfully, competition host shared the code to split train/test dataset.\nhttps://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\n\nBut, the way  to split test dataset into public/private is not shared.\nIf it is splited by the specific order  (e.g. session-ID order, time-series order), the risk of overfitting with public testset would become larger. \n\nHow do you think about this issue ?\nDoes anyone verify the way to split by using public LB ?",
      "votes": null
    },
    {
      "id": "2099787",
      "postDate": "01/14/2023 17:38:20",
      "content": "<p>I haven't seen any verification probably because everyone assumes it is split by time. It could also be different like you said.</p>",
      "rawMarkdown": "I haven't seen any verification probably because everyone assumes it is split by time. It could also be different like you said.",
      "votes": null
    },
    {
      "id": "2100119",
      "postDate": "01/15/2023 00:13:25",
      "content": "<p>I believe that they are simply split randomly by session.<br>\nI can’t think of any reasonable reason for host by splitting with time or any specific orders. If they want to see if the model’s generalization over data-shift, splitting without leakage (no-overlapping sessions) like the split between train/test datasets is more reasonable. <br>\nAlthough this could be wrong.</p>",
      "rawMarkdown": "I believe that they are simply split randomly by session.\nI can’t think of any reasonable reason for host by splitting with time or any specific orders. If they want to see if the model’s generalization over data-shift, splitting without leakage (no-overlapping sessions) like the split between train/test datasets is more reasonable. \nAlthough this could be wrong.",
      "votes": null
    },
    {
      "id": "2100132",
      "postDate": "01/15/2023 00:26:44",
      "content": "<p>If you want to validate the assumption that public/private sets are split with specific time orders, you can verify this by the following process.</p>\n<p>First split the test dataset into two parts so that they have the same size (let’s call them A and B), whereas part A has its mean ‘ts’ &lt; threshold, and part B has its mean ‘ts’ &gt;= threshold. (Note that they are group by sessions.)<br>\nAfter that, make the following submissions:</p>\n<ol>\n<li>mask your model’s prediction for part A</li>\n<li>mask your model’s prediction for part B</li>\n</ol>\n<p>And finally, compare these scores.<br>\nIf they differs significantly, there could be a chance that the data are split in the specific time order.<br>\n(E.g., if the data B contains more private samples, the score drop of submission 2 would be less compared to submission 1.)</p>",
      "rawMarkdown": "If you want to validate the assumption that public/private sets are split with specific time orders, you can verify this by the following process.\n\nFirst split the test dataset into two parts so that they have the same size (let’s call them A and B), whereas part A has its mean ‘ts’ < threshold, and part B has its mean ‘ts’ >= threshold. (Note that they are group by sessions.)\nAfter that, make the following submissions:\n\n1. mask your model’s prediction for part A\n2. mask your model’s prediction for part B\n\nAnd finally, compare these scores.\nIf they differs significantly, there could be a chance that the data are split in the specific time order.\n(E.g., if the data B contains more private samples, the score drop of submission 2 would be less compared to submission 1.)",
      "votes": null
    },
    {
      "id": "2100670",
      "postDate": "01/15/2023 10:46:12",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a>, thanks for your question! I can confirm that <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>’s assumption is correct, and we indeed used a random split for dividing our sessions into public and private test sets.</p>",
      "rawMarkdown": "Hi @toshik, thanks for your question! I can confirm that @tatamikenn’s assumption is correct, and we indeed used a random split for dividing our sessions into public and private test sets.",
      "votes": null
    },
    {
      "id": "2100680",
      "postDate": "01/15/2023 10:56:18",
      "content": "<p>Thank you for your answer ! Your information make it possible to work on essential parts for us.</p>",
      "rawMarkdown": "Thank you for your answer ! Your information make it possible to work on essential parts for us.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2099787,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "01/14/2023 17:38:20",
      "content": "<p>I haven't seen any verification probably because everyone assumes it is split by time. It could also be different like you said.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2100119,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "01/15/2023 00:13:25",
      "content": "<p>I believe that they are simply split randomly by session.<br>\nI can’t think of any reasonable reason for host by splitting with time or any specific orders. If they want to see if the model’s generalization over data-shift, splitting without leakage (no-overlapping sessions) like the split between train/test datasets is more reasonable. <br>\nAlthough this could be wrong.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2100132,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/15/2023 00:26:44",
          "content": "<p>If you want to validate the assumption that public/private sets are split with specific time orders, you can verify this by the following process.</p>\n<p>First split the test dataset into two parts so that they have the same size (let’s call them A and B), whereas part A has its mean ‘ts’ &lt; threshold, and part B has its mean ‘ts’ &gt;= threshold. (Note that they are group by sessions.)<br>\nAfter that, make the following submissions:</p>\n<ol>\n<li>mask your model’s prediction for part A</li>\n<li>mask your model’s prediction for part B</li>\n</ol>\n<p>And finally, compare these scores.<br>\nIf they differs significantly, there could be a chance that the data are split in the specific time order.<br>\n(E.g., if the data B contains more private samples, the score drop of submission 2 would be less compared to submission 1.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2100670,
      "author_name": "pnormann",
      "author_url": "",
      "post_date": "01/15/2023 10:46:12",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/toshik\" target=\"_blank\">@toshik</a>, thanks for your question! I can confirm that <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>’s assumption is correct, and we indeed used a random split for dividing our sessions into public and private test sets.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2100680,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "01/15/2023 10:56:18",
          "content": "<p>Thank you for your answer ! Your information make it possible to work on essential parts for us.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2099595": "Thankfully, competition host shared the code to split train/test dataset.\nhttps://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\n\nBut, the way  to split test dataset into public/private is not shared.\nIf it is splited by the specific order  (e.g. session-ID order, time-series order), the risk of overfitting with public testset would become larger. \n\nHow do you think about this issue ?\nDoes anyone verify the way to split by using public LB ?",
    "2099787": "I haven't seen any verification probably because everyone assumes it is split by time. It could also be different like you said.",
    "2100119": "I believe that they are simply split randomly by session.\nI can’t think of any reasonable reason for host by splitting with time or any specific orders. If they want to see if the model’s generalization over data-shift, splitting without leakage (no-overlapping sessions) like the split between train/test datasets is more reasonable. \nAlthough this could be wrong.",
    "2100132": "If you want to validate the assumption that public/private sets are split with specific time orders, you can verify this by the following process.\n\nFirst split the test dataset into two parts so that they have the same size (let’s call them A and B), whereas part A has its mean ‘ts’ < threshold, and part B has its mean ‘ts’ >= threshold. (Note that they are group by sessions.)\nAfter that, make the following submissions:\n\n1. mask your model’s prediction for part A\n2. mask your model’s prediction for part B\n\nAnd finally, compare these scores.\nIf they differs significantly, there could be a chance that the data are split in the specific time order.\n(E.g., if the data B contains more private samples, the score drop of submission 2 would be less compared to submission 1.)",
    "2100670": "Hi @toshik, thanks for your question! I can confirm that @tatamikenn’s assumption is correct, and we indeed used a random split for dividing our sessions into public and private test sets.",
    "2100680": "Thank you for your answer ! Your information make it possible to work on essential parts for us."
  },
  "source": "meta"
}