{
  "id": 329275,
  "title": "Be aware of the counts per time! ",
  "url": "/competitions/amex-default-prediction/discussion/329275",
  "author_name": "Laura Fink",
  "post_date": "2022-06-05T19:25:37.508000",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hey Kagglers,</p>\n<p>Have you seen that there are differences between the test and the train data in the value counts of the temporal feature S_2? I'm not sure yet what it means. In the test data it seems that we have more customer IDs between mid of October 2018 until May 2019 or at least more entries…</p>\n<p>Should we expect that something like this will also be present in the private test data?</p>\n<p><img src=\"https://i.imgur.com/VtABUMS.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1812422,
      "postDate": "2022-06-05T19:25:37.510Z",
      "content": "<p>Hey Kagglers,</p>\n<p>Have you seen that there are differences between the test and the train data in the value counts of the temporal feature S_2? I'm not sure yet what it means. In the test data it seems that we have more customer IDs between mid of October 2018 until May 2019 or at least more entries…</p>\n<p>Should we expect that something like this will also be present in the private test data?</p>\n<p><img src=\"https://i.imgur.com/VtABUMS.png\" alt=\"\"></p>",
      "rawMarkdown": "Hey Kagglers,\n\nHave you seen that there are differences between the test and the train data in the value counts of the temporal feature S_2? I'm not sure yet what it means. In the test data it seems that we have more customer IDs between mid of October 2018 until May 2019 or at least more entries...\n\nShould we expect that something like this will also be present in the private test data?\n\n\n![](https://i.imgur.com/VtABUMS.png)",
      "votes": 16
    },
    {
      "id": 1812500,
      "postDate": "2022-06-05T23:33:57.187Z",
      "content": "<p>Hi Laura, nice observation. For each customer we have 1 year of data. All the train data is one window (1 year) of time but the test data is two windows of time (2 overlaping 1 year intervals)</p>\n<p>The test data from Apr 1 2018 thru April 2019 is the public test and the data from October 2018 thru October 2019 is the private test. Half the test dataset belongs to each, so we see double the data where they overlap. </p>",
      "rawMarkdown": "Hi Laura, nice observation. For each customer we have 1 year of data. All the train data is one window (1 year) of time but the test data is two windows of time (2 overlaping 1 year intervals)\n\nThe test data from Apr 1 2018 thru April 2019 is the public test and the data from October 2018 thru October 2019 is the private test. Half the test dataset belongs to each, so we see double the data where they overlap. ",
      "votes": 9,
      "replies": [
        {
          "id": 1812517,
          "postDate": "2022-06-06T00:26:56.577Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 1812522,
          "postDate": "2022-06-06T00:58:53.713Z",
          "content": "<p>yes. Ryota explains it <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "yes. Ryota explains it [here][1]\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926",
          "votes": 3
        },
        {
          "id": 1812691,
          "postDate": "2022-06-06T05:54:02.210Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>\n<p>Thank you for your explanation and the link. I wasn't aware of it that's possible to have a glimpse on the overall test distribution (public an private) when looking at the histograms of the data description page ;-). Crazy… In my opinion some kind of leakage. At least one can now search for features with extreme outliers in public or private (in comparison with EDA on the public test data) to understand LB differences. </p>",
          "rawMarkdown": "Hi @cdeotte !\n\nThank you for your explanation and the link. I wasn't aware of it that's possible to have a glimpse on the overall test distribution (public an private) when looking at the histograms of the data description page ;-). Crazy... In my opinion some kind of leakage. At least one can now search for features with extreme outliers in public or private (in comparison with EDA on the public test data) to understand LB differences. ",
          "votes": 2
        },
        {
          "id": 1813315,
          "postDate": "2022-06-06T17:33:21.437Z",
          "content": "<p><a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a>  </p>\n<p>I usually get to know data in a tool I am comfortable and is intuitive (like a Tableau or Excel).</p>\n<p>I created a calculated field that finds the max date per customer in Tableau … I then trended that to help me get to know the data better.<br>\nBelow, you can see 1 chunk for train (Mar 2018) and 2 chunks for test (april 2019, oct 2019).</p>\n<p><img src=\"https://i.imgur.com/NeeYGZ8.png\"></p>",
          "rawMarkdown": "@allunia  \n\nI usually get to know data in a tool I am comfortable and is intuitive (like a Tableau or Excel).\n\nI created a calculated field that finds the max date per customer in Tableau … I then trended that to help me get to know the data better.\nBelow, you can see 1 chunk for train (Mar 2018) and 2 chunks for test (april 2019, oct 2019).\n\n<img src=\"https://i.imgur.com/NeeYGZ8.png\">"
        }
      ]
    },
    {
      "id": 1812463,
      "postDate": "2022-06-05T20:57:26.863Z",
      "content": "<p>Very likely, the majority didn't see any difference between the test and the train data yet.<br>\nI hope some experienced professional (or any one) can answer and contribute with your topic, clarifying that doubt.</p>",
      "rawMarkdown": "Very likely, the majority didn't see any difference between the test and the train data yet.\nI hope some experienced professional (or any one) can answer and contribute with your topic, clarifying that doubt.",
      "votes": 1,
      "replies": [
        {
          "id": 1813455,
          "postDate": "2022-06-06T20:39:34.130Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's answer is the best for the OP's specific question. But for train vs test deltas, so far there's one topic that I'm aware of, related to B_29:</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756</a>  </p>",
          "rawMarkdown": "@cdeotte 's answer is the best for the OP's specific question. But for train vs test deltas, so far there's one topic that I'm aware of, related to B_29:\n\nhttps://www.kaggle.com/competitions/amex-default-prediction/discussion/328756  "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1812500,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-06-05T23:33:57.187000",
      "content": "<p>Hi Laura, nice observation. For each customer we have 1 year of data. All the train data is one window (1 year) of time but the test data is two windows of time (2 overlaping 1 year intervals)</p>\n<p>The test data from Apr 1 2018 thru April 2019 is the public test and the data from October 2018 thru October 2019 is the private test. Half the test dataset belongs to each, so we see double the data where they overlap. </p>",
      "votes": 9,
      "replies": [
        {
          "id": 1812517,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-06T00:26:56.577000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1812522,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-06T00:58:53.713000",
          "content": "<p>yes. Ryota explains it <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327926\" target=\"_blank\">here</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1812691,
          "author_name": "Laura Fink",
          "author_url": "",
          "post_date": "2022-06-06T05:54:02.210000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>\n<p>Thank you for your explanation and the link. I wasn't aware of it that's possible to have a glimpse on the overall test distribution (public an private) when looking at the histograms of the data description page ;-). Crazy… In my opinion some kind of leakage. At least one can now search for features with extreme outliers in public or private (in comparison with EDA on the public test data) to understand LB differences. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1813315,
          "author_name": "David Dirring",
          "author_url": "",
          "post_date": "2022-06-06T17:33:21.437000",
          "content": "<p><a href=\"https://www.kaggle.com/allunia\" target=\"_blank\">@allunia</a>  </p>\n<p>I usually get to know data in a tool I am comfortable and is intuitive (like a Tableau or Excel).</p>\n<p>I created a calculated field that finds the max date per customer in Tableau … I then trended that to help me get to know the data better.<br>\nBelow, you can see 1 chunk for train (Mar 2018) and 2 chunks for test (april 2019, oct 2019).</p>\n<p><img src=\"https://i.imgur.com/NeeYGZ8.png\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1812463,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-06-05T20:57:26.863000",
      "content": "<p>Very likely, the majority didn't see any difference between the test and the train data yet.<br>\nI hope some experienced professional (or any one) can answer and contribute with your topic, clarifying that doubt.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1813455,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-06-06T20:39:34.130000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's answer is the best for the OP's specific question. But for train vs test deltas, so far there's one topic that I'm aware of, related to B_29:</p>\n<p><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/328756</a>  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1812422": "Hey Kagglers,\n\nHave you seen that there are differences between the test and the train data in the value counts of the temporal feature S_2? I'm not sure yet what it means. In the test data it seems that we have more customer IDs between mid of October 2018 until May 2019 or at least more entries...\n\nShould we expect that something like this will also be present in the private test data?\n\n\n![](https://i.imgur.com/VtABUMS.png)",
    "1812500": "Hi Laura, nice observation. For each customer we have 1 year of data. All the train data is one window (1 year) of time but the test data is two windows of time (2 overlaping 1 year intervals)\n\nThe test data from Apr 1 2018 thru April 2019 is the public test and the data from October 2018 thru October 2019 is the private test. Half the test dataset belongs to each, so we see double the data where they overlap. ",
    "1812463": "Very likely, the majority didn't see any difference between the test and the train data yet.\nI hope some experienced professional (or any one) can answer and contribute with your topic, clarifying that doubt."
  }
}