{
  "id": 328756,
  "title": "The distribution of missing values over time",
  "url": "/competitions/amex-default-prediction/discussion/328756",
  "author_name": "AmbrosM",
  "post_date": "2022-06-02T21:02:58.864000",
  "votes": 100,
  "comment_count": 9,
  "views": 0,
  "content": "<p>If we count the number of rows which our three datasets (train, public test and private test) have per month, we see that the three datasets have equal size: Every dataset covers 13 months, starts at 400000 customers and slowly grows with the months:</p>\n<p><img src=\"https://i.imgur.com/oiMWKbo.png\" alt=\"all.png\"></p>\n<p>Now compare this diagram to the next one, which shows the distribution of non-null values over time. B_29 is an interesting feature. Given that each of the three datasets contains almost half a million customers, we see that until May of 2019 the B_29 value is available for fewer than a tenth of the customers. The other nine tenths are missing. In June of 2019, American Express started collecting B_29 data for all customers:</p>\n<p><img src=\"https://i.imgur.com/ze8pr7w.png\" alt=\"B_29.png\"></p>\n<p>What insight can we get from these diagrams? The distribution of the missing B_29 differs between the train and test datasets. Whereas in the training and public leaderboard data &gt;90 % are missing, during the last five months of private leaderboard, we have B_29 data for almost every customer. If we use this feature in our models, we should be prepared for surprises in the private leaderboard. Could dropping the feature be a better solution?</p>\n<p>Source code is <a href=\"https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": 1809605,
      "postDate": "2022-06-02T21:02:58.863Z",
      "content": "<p>If we count the number of rows which our three datasets (train, public test and private test) have per month, we see that the three datasets have equal size: Every dataset covers 13 months, starts at 400000 customers and slowly grows with the months:</p>\n<p><img src=\"https://i.imgur.com/oiMWKbo.png\" alt=\"all.png\"></p>\n<p>Now compare this diagram to the next one, which shows the distribution of non-null values over time. B_29 is an interesting feature. Given that each of the three datasets contains almost half a million customers, we see that until May of 2019 the B_29 value is available for fewer than a tenth of the customers. The other nine tenths are missing. In June of 2019, American Express started collecting B_29 data for all customers:</p>\n<p><img src=\"https://i.imgur.com/ze8pr7w.png\" alt=\"B_29.png\"></p>\n<p>What insight can we get from these diagrams? The distribution of the missing B_29 differs between the train and test datasets. Whereas in the training and public leaderboard data &gt;90 % are missing, during the last five months of private leaderboard, we have B_29 data for almost every customer. If we use this feature in our models, we should be prepared for surprises in the private leaderboard. Could dropping the feature be a better solution?</p>\n<p>Source code is <a href=\"https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "If we count the number of rows which our three datasets (train, public test and private test) have per month, we see that the three datasets have equal size: Every dataset covers 13 months, starts at 400000 customers and slowly grows with the months:\n\n![all.png](https://i.imgur.com/oiMWKbo.png)\n\nNow compare this diagram to the next one, which shows the distribution of non-null values over time. B_29 is an interesting feature. Given that each of the three datasets contains almost half a million customers, we see that until May of 2019 the B_29 value is available for fewer than a tenth of the customers. The other nine tenths are missing. In June of 2019, American Express started collecting B_29 data for all customers:\n\n![B_29.png](https://i.imgur.com/ze8pr7w.png)\n\nWhat insight can we get from these diagrams? The distribution of the missing B_29 differs between the train and test datasets. Whereas in the training and public leaderboard data >90 % are missing, during the last five months of private leaderboard, we have B_29 data for almost every customer. If we use this feature in our models, we should be prepared for surprises in the private leaderboard. Could dropping the feature be a better solution?\n\nSource code is [here](https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense).\n",
      "votes": 97
    },
    {
      "id": 1809608,
      "postDate": "2022-06-02T21:08:28.007Z",
      "content": "<p>Great insight. Analyzing the differences between private and train+public will be important to avoid shakedown.</p>",
      "rawMarkdown": "Great insight. Analyzing the differences between private and train+public will be important to avoid shakedown.",
      "votes": 10
    },
    {
      "id": 1813883,
      "postDate": "2022-06-07T10:10:36.463Z",
      "content": "<p>Great insights. This could be something that has happened due to the data collection process that was present prior to June 2019. I have seen this many times for financial datasets. A feature can take on greater importance and only begins to be collected more efficiently moving forward. Many times the data will not be available in the past as it was not requested from the customer or it was something that a data collection system could only process for certain customers. Thanks for sharing 😊</p>",
      "rawMarkdown": "Great insights. This could be something that has happened due to the data collection process that was present prior to June 2019. I have seen this many times for financial datasets. A feature can take on greater importance and only begins to be collected more efficiently moving forward. Many times the data will not be available in the past as it was not requested from the customer or it was something that a data collection system could only process for certain customers. Thanks for sharing 😊",
      "votes": 5
    },
    {
      "id": 1811774,
      "postDate": "2022-06-05T05:59:47.823Z",
      "content": "<p>This is good stuff <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>!</p>\n<p>Because <em>B</em> variables are assigned to <em>Balance</em> features, maybe this particular attribute was a new tracking metric implemented later.</p>\n<p>Or dare I say clerical error/bug at AMEX?</p>",
      "rawMarkdown": "This is good stuff @ambrosm!\n\nBecause _B_ variables are assigned to _Balance_ features, maybe this particular attribute was a new tracking metric implemented later.\n\nOr dare I say clerical error/bug at AMEX?",
      "votes": 3
    },
    {
      "id": 1811013,
      "postDate": "2022-06-04T07:09:30.623Z",
      "content": "<p>Well spotted, <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>!</p>",
      "rawMarkdown": "Well spotted, @ambrosm!",
      "votes": 3
    },
    {
      "id": 1810154,
      "postDate": "2022-06-03T10:03:00.817Z",
      "content": "<p>Thanks for your sharing Very good insight!!👍</p>",
      "rawMarkdown": "Thanks for your sharing Very good insight!!👍",
      "votes": 4
    },
    {
      "id": 1883715,
      "postDate": "2022-08-04T03:35:54.390Z",
      "content": "<p>Thanks for your sharing Very good insight!! But we should be more careful about whether we need to delete whole missing values</p>",
      "rawMarkdown": "Thanks for your sharing Very good insight!! But we should be more careful about whether we need to delete whole missing values",
      "votes": 1
    },
    {
      "id": 1812823,
      "postDate": "2022-06-06T08:57:32.713Z",
      "content": "<p>Very nice find. In my opinion, I don't think that we should drop this kind of features necessarily. But we should definitely be careful how we pass  the NaN information to our models.</p>",
      "rawMarkdown": "Very nice find. In my opinion, I don't think that we should drop this kind of features necessarily. But we should definitely be careful how we pass  the NaN information to our models.",
      "votes": 1
    },
    {
      "id": 1811769,
      "postDate": "2022-06-05T05:49:25.360Z",
      "content": "<p>Great point.</p>",
      "rawMarkdown": "Great point.",
      "votes": 1
    },
    {
      "id": 1812702,
      "postDate": "2022-06-06T06:14:24.763Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1809608,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-06-02T21:08:28.007000",
      "content": "<p>Great insight. Analyzing the differences between private and train+public will be important to avoid shakedown.</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 1813883,
      "author_name": "James McNeill",
      "author_url": "",
      "post_date": "2022-06-07T10:10:36.463000",
      "content": "<p>Great insights. This could be something that has happened due to the data collection process that was present prior to June 2019. I have seen this many times for financial datasets. A feature can take on greater importance and only begins to be collected more efficiently moving forward. Many times the data will not be available in the past as it was not requested from the customer or it was something that a data collection system could only process for certain customers. Thanks for sharing 😊</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1811774,
      "author_name": "Kris Smith",
      "author_url": "",
      "post_date": "2022-06-05T05:59:47.823000",
      "content": "<p>This is good stuff <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>!</p>\n<p>Because <em>B</em> variables are assigned to <em>Balance</em> features, maybe this particular attribute was a new tracking metric implemented later.</p>\n<p>Or dare I say clerical error/bug at AMEX?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1811013,
      "author_name": "Ruchi Bhatia",
      "author_url": "",
      "post_date": "2022-06-04T07:09:30.623000",
      "content": "<p>Well spotted, <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1810154,
      "author_name": "ds.wook",
      "author_url": "",
      "post_date": "2022-06-03T10:03:00.817000",
      "content": "<p>Thanks for your sharing Very good insight!!👍</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1883715,
      "author_name": "Pengfei",
      "author_url": "",
      "post_date": "2022-08-04T03:35:54.390000",
      "content": "<p>Thanks for your sharing Very good insight!! But we should be more careful about whether we need to delete whole missing values</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1812823,
      "author_name": "Nikola Bacic",
      "author_url": "",
      "post_date": "2022-06-06T08:57:32.713000",
      "content": "<p>Very nice find. In my opinion, I don't think that we should drop this kind of features necessarily. But we should definitely be careful how we pass  the NaN information to our models.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1811769,
      "author_name": "SgangX",
      "author_url": "",
      "post_date": "2022-06-05T05:49:25.360000",
      "content": "<p>Great point.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1812702,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-06T06:14:24.763000",
      "content": "",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1809605": "If we count the number of rows which our three datasets (train, public test and private test) have per month, we see that the three datasets have equal size: Every dataset covers 13 months, starts at 400000 customers and slowly grows with the months:\n\n![all.png](https://i.imgur.com/oiMWKbo.png)\n\nNow compare this diagram to the next one, which shows the distribution of non-null values over time. B_29 is an interesting feature. Given that each of the three datasets contains almost half a million customers, we see that until May of 2019 the B_29 value is available for fewer than a tenth of the customers. The other nine tenths are missing. In June of 2019, American Express started collecting B_29 data for all customers:\n\n![B_29.png](https://i.imgur.com/ze8pr7w.png)\n\nWhat insight can we get from these diagrams? The distribution of the missing B_29 differs between the train and test datasets. Whereas in the training and public leaderboard data >90 % are missing, during the last five months of private leaderboard, we have B_29 data for almost every customer. If we use this feature in our models, we should be prepared for surprises in the private leaderboard. Could dropping the feature be a better solution?\n\nSource code is [here](https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense).\n",
    "1809608": "Great insight. Analyzing the differences between private and train+public will be important to avoid shakedown.",
    "1813883": "Great insights. This could be something that has happened due to the data collection process that was present prior to June 2019. I have seen this many times for financial datasets. A feature can take on greater importance and only begins to be collected more efficiently moving forward. Many times the data will not be available in the past as it was not requested from the customer or it was something that a data collection system could only process for certain customers. Thanks for sharing 😊",
    "1811774": "This is good stuff @ambrosm!\n\nBecause _B_ variables are assigned to _Balance_ features, maybe this particular attribute was a new tracking metric implemented later.\n\nOr dare I say clerical error/bug at AMEX?",
    "1811013": "Well spotted, @ambrosm!",
    "1810154": "Thanks for your sharing Very good insight!!👍",
    "1883715": "Thanks for your sharing Very good insight!! But we should be more careful about whether we need to delete whole missing values",
    "1812823": "Very nice find. In my opinion, I don't think that we should drop this kind of features necessarily. But we should definitely be careful how we pass  the NaN information to our models.",
    "1811769": "Great point.",
    "1812702": ""
  }
}