{
  "id": 327651,
  "title": "Strange Histograms",
  "url": "/competitions/amex-default-prediction/discussion/327651",
  "author_name": "Chris Deotte",
  "post_date": "2022-05-28T11:47:51.486000",
  "votes": 73,
  "comment_count": 14,
  "views": 0,
  "content": "<h1>Strange Histograms</h1>\n<p>I noticed that many features have strange histograms. For example, at first glance, it appears that column <code>B_8</code> is a binary feature that has values 0 and 1.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b1.png\" alt=\"\"></p>\n<p>However if we zoom in, we see that there are many values between 0 and 0.01. And many values between 1.0 and 1.01</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b2.png\" alt=\"\"></p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b3.png\" alt=\"\"></p>\n<p>Specifically in the train data with <code>5_531_451</code> rows, there are <code>3_053_857</code> values between 0.0 and 0.1. And there are <code>2_455_326</code> values between 1.0 and 1.1.</p>\n<h1>Questions</h1>\n<p>This prompts questions. </p>\n<ul>\n<li>If we convert this column into two values, namely 0 and 1, do we lose information?</li>\n<li>Note when we convert this column from <code>float32</code> to <code>float16</code> we decrease the number of unique values from <code>2_948_222</code> to <code>8_501</code>. Does this lose important information? </li>\n<li>Why are so many values uniformly distributed between 0.0 and 0.01? Uniform distributions seldom happen in the real world. Did AMEX just add uniform noise to the 0's and noise to the 1's?</li>\n<li>If we use an NN, do we need to process this column differently than standard scaler for the NN to access all these tightly packed values?</li>\n</ul>\n<h1>UPDATE</h1>\n<p>At the same time I posted this discussion, Raddar posted a similar topic. Raddar hypothesizes that random uniform noise has been injected into the data. His discussion is <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\" target=\"_blank\">here</a> and notebook <a href=\"https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\" target=\"_blank\">here</a>. He points out that columns related to money such as <code>B_8</code> (which is a <code>balance variable</code>) should have less unique values than  <code>2_948_222</code>. The existence of many unique values supports the argument that random noise has been added.</p>\n<h2>UPDATE2 - XGB Starter Notebook - LB 794 - without Noise</h2>\n<p>Raddar has converted the entire train and test dataset to remove this noise <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>. His Kaggle dataset shrinks the original data from 50GB to 5GB! He posted a discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a>. </p>\n<p>Afterward, I published an XGBoost starter notebook using his Kaggle dataset <a href=\"https://tinyurl.com/yb769r9c\" target=\"_blank\">here</a> which achieves CV 0.792 and LB 0.794. This is the same accuracy as other public notebooks using similar features. Therefore it appears that this noise does not contain signal.</p>",
  "messages": [
    {
      "id": 1803973,
      "postDate": "2022-05-28T11:47:51.487Z",
      "content": "<h1>Strange Histograms</h1>\n<p>I noticed that many features have strange histograms. For example, at first glance, it appears that column <code>B_8</code> is a binary feature that has values 0 and 1.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b1.png\" alt=\"\"></p>\n<p>However if we zoom in, we see that there are many values between 0 and 0.01. And many values between 1.0 and 1.01</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b2.png\" alt=\"\"></p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b3.png\" alt=\"\"></p>\n<p>Specifically in the train data with <code>5_531_451</code> rows, there are <code>3_053_857</code> values between 0.0 and 0.1. And there are <code>2_455_326</code> values between 1.0 and 1.1.</p>\n<h1>Questions</h1>\n<p>This prompts questions. </p>\n<ul>\n<li>If we convert this column into two values, namely 0 and 1, do we lose information?</li>\n<li>Note when we convert this column from <code>float32</code> to <code>float16</code> we decrease the number of unique values from <code>2_948_222</code> to <code>8_501</code>. Does this lose important information? </li>\n<li>Why are so many values uniformly distributed between 0.0 and 0.01? Uniform distributions seldom happen in the real world. Did AMEX just add uniform noise to the 0's and noise to the 1's?</li>\n<li>If we use an NN, do we need to process this column differently than standard scaler for the NN to access all these tightly packed values?</li>\n</ul>\n<h1>UPDATE</h1>\n<p>At the same time I posted this discussion, Raddar posted a similar topic. Raddar hypothesizes that random uniform noise has been injected into the data. His discussion is <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\" target=\"_blank\">here</a> and notebook <a href=\"https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\" target=\"_blank\">here</a>. He points out that columns related to money such as <code>B_8</code> (which is a <code>balance variable</code>) should have less unique values than  <code>2_948_222</code>. The existence of many unique values supports the argument that random noise has been added.</p>\n<h2>UPDATE2 - XGB Starter Notebook - LB 794 - without Noise</h2>\n<p>Raddar has converted the entire train and test dataset to remove this noise <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>. His Kaggle dataset shrinks the original data from 50GB to 5GB! He posted a discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a>. </p>\n<p>Afterward, I published an XGBoost starter notebook using his Kaggle dataset <a href=\"https://tinyurl.com/yb769r9c\" target=\"_blank\">here</a> which achieves CV 0.792 and LB 0.794. This is the same accuracy as other public notebooks using similar features. Therefore it appears that this noise does not contain signal.</p>",
      "rawMarkdown": "# Strange Histograms\nI noticed that many features have strange histograms. For example, at first glance, it appears that column `B_8` is a binary feature that has values 0 and 1.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b1.png)\n\nHowever if we zoom in, we see that there are many values between 0 and 0.01. And many values between 1.0 and 1.01\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b2.png)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b3.png)\n\nSpecifically in the train data with `5_531_451` rows, there are `3_053_857` values between 0.0 and 0.1. And there are `2_455_326` values between 1.0 and 1.1.\n\n# Questions\nThis prompts questions. \n* If we convert this column into two values, namely 0 and 1, do we lose information?\n* Note when we convert this column from `float32` to `float16` we decrease the number of unique values from `2_948_222` to `8_501`. Does this lose important information? \n* Why are so many values uniformly distributed between 0.0 and 0.01? Uniform distributions seldom happen in the real world. Did AMEX just add uniform noise to the 0's and noise to the 1's?\n* If we use an NN, do we need to process this column differently than standard scaler for the NN to access all these tightly packed values?\n\n# UPDATE\nAt the same time I posted this discussion, Raddar posted a similar topic. Raddar hypothesizes that random uniform noise has been injected into the data. His discussion is [here][1] and notebook [here][2]. He points out that columns related to money such as `B_8` (which is a `balance variable`) should have less unique values than  `2_948_222`. The existence of many unique values supports the argument that random noise has been added.\n\n## UPDATE2 - XGB Starter Notebook - LB 794 - without Noise\nRaddar has converted the entire train and test dataset to remove this noise [here][5]. His Kaggle dataset shrinks the original data from 50GB to 5GB! He posted a discussion [here][3]. \n\nAfterward, I published an XGBoost starter notebook using his Kaggle dataset [here][4] which achieves CV 0.792 and LB 0.794. This is the same accuracy as other public notebooks using similar features. Therefore it appears that this noise does not contain signal.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\n[2]: https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\n[3]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n[4]: https://tinyurl.com/yb769r9c\n[5]: https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
      "votes": 71
    },
    {
      "id": 1809384,
      "postDate": "2022-06-02T16:44:22.320Z",
      "content": "<p>UPDATE: Raddar has converted the entire train and test dataset to remove this noise <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>. His Kaggle dataset shrinks the original data from 50GB to 5GB! He posted a discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a>. </p>\n<p>Here is another example of random uniform noise. The variable R_26 has 548160 unique values below 0.2 and is provided as float64 (8 bytes), but really from the plot we see it is only 6 discrete values with (random uniform) noise added and can be converted to int8 (1 byte) - without information loss.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/r_26.png\" alt=\"\"></p>\n<p>Afterward, I published an XGBoost starter notebook using his Kaggle dataset <a href=\"https://tinyurl.com/yb769r9c\" target=\"_blank\">here</a> which achieves CV 0.792 and LB 0.794. This is the same accuracy as other public notebooks using similar features. Therefore it appears that this noise does not contain signal.</p>",
      "rawMarkdown": "UPDATE: Raddar has converted the entire train and test dataset to remove this noise [here][5]. His Kaggle dataset shrinks the original data from 50GB to 5GB! He posted a discussion [here][3]. \n\nHere is another example of random uniform noise. The variable R_26 has 548160 unique values below 0.2 and is provided as float64 (8 bytes), but really from the plot we see it is only 6 discrete values with (random uniform) noise added and can be converted to int8 (1 byte) - without information loss.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/r_26.png)\n\nAfterward, I published an XGBoost starter notebook using his Kaggle dataset [here][4] which achieves CV 0.792 and LB 0.794. This is the same accuracy as other public notebooks using similar features. Therefore it appears that this noise does not contain signal.\n\n[3]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n[4]: https://tinyurl.com/yb769r9c\n[5]: https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
      "votes": 6
    },
    {
      "id": 1806932,
      "postDate": "2022-05-31T15:26:49.087Z",
      "content": "<p>These kind of artifacts don't have good record on Kaggle. Most of the time, people somehow reverse engineer them and gain benefit.</p>",
      "rawMarkdown": "These kind of artifacts don't have good record on Kaggle. Most of the time, people somehow reverse engineer them and gain benefit.",
      "votes": 3,
      "replies": [
        {
          "id": 1807045,
          "postDate": "2022-05-31T16:55:25.147Z",
          "content": "<p>haha good point</p>",
          "rawMarkdown": "haha good point"
        }
      ]
    },
    {
      "id": 1804053,
      "postDate": "2022-05-28T13:37:54.040Z",
      "content": "<p>Interesting observation. I agree that noise is a possible explanation. Another possibility is that this is a composite feature of a binary 0 or 1 summed with another feature whose typical values might be in the range 0 to 0.1 but predominantly at or close to 0. If that is the case then there could be important information that is lost with converting to float16. It does make sense that this is noise as 2_948_222 unique values for money sounds like a lot, but perhaps it isn't if this includes payments in foreign currencies with constantly moving fx rates being translated back into the currency of the cardholder before rounding to 2 decimal places for the lowest unit of currency. It would be nice if there was more information on the meaning of the features.</p>",
      "rawMarkdown": "Interesting observation. I agree that noise is a possible explanation. Another possibility is that this is a composite feature of a binary 0 or 1 summed with another feature whose typical values might be in the range 0 to 0.1 but predominantly at or close to 0. If that is the case then there could be important information that is lost with converting to float16. It does make sense that this is noise as 2_948_222 unique values for money sounds like a lot, but perhaps it isn't if this includes payments in foreign currencies with constantly moving fx rates being translated back into the currency of the cardholder before rounding to 2 decimal places for the lowest unit of currency. It would be nice if there was more information on the meaning of the features.",
      "votes": 4
    },
    {
      "id": 1838032,
      "postDate": "2022-06-30T06:33:42.220Z",
      "content": "<p>If you consider correlation, binary features give you a singular matrix. Adding a bit of noise prevents this. Maybe a reason?</p>",
      "rawMarkdown": "If you consider correlation, binary features give you a singular matrix. Adding a bit of noise prevents this. Maybe a reason?",
      "votes": 1,
      "replies": [
        {
          "id": 1840858,
          "postDate": "2022-07-02T15:37:24.933Z",
          "content": "<p>I think the reason that Fidelity added noise was to make the data more anonymous. </p>",
          "rawMarkdown": "I think the reason that Fidelity added noise was to make the data more anonymous. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1804496,
      "postDate": "2022-05-29T04:39:43.177Z",
      "content": "<p>I also found some variables with strange histograms. I have 2 questions though:</p>\n<ul>\n<li>Why is it too much to have 2.9M unique values for balance? I think in terms of dollar value, it could be a range from 0.00 to 30,000.00. It would be even more probable for a broader range, e.g. 0.00 - 35,000.00</li>\n<li>Why would the host add noises to the original data? </li>\n</ul>",
      "rawMarkdown": "I also found some variables with strange histograms. I have 2 questions though:\n- Why is it too much to have 2.9M unique values for balance? I think in terms of dollar value, it could be a range from 0.00 to 30,000.00. It would be even more probable for a broader range, e.g. 0.00 - 35,000.00\n- Why would the host add noises to the original data? ",
      "votes": 1,
      "replies": [
        {
          "id": 1804570,
          "postDate": "2022-05-29T07:06:28.980Z",
          "content": "<p>Without noise, it would be a lot easier to de-anonymize the data and find the right scale.</p>",
          "rawMarkdown": "Without noise, it would be a lot easier to de-anonymize the data and find the right scale.",
          "votes": 3
        },
        {
          "id": 1812970,
          "postDate": "2022-06-06T12:10:20.823Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1805617,
      "postDate": "2022-05-30T11:36:26.497Z",
      "content": "<p>Interesting!</p>",
      "rawMarkdown": "Interesting!",
      "votes": 2
    },
    {
      "id": 1805389,
      "postDate": "2022-05-30T05:49:47.310Z",
      "content": "<p>Thanks for sharing, this is very useful <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Maybe for the final submission, we can submit one solution with noise removal and one without noise removal. It would be interesting to see which one holds the \"test\" of time.</p>",
      "rawMarkdown": "Thanks for sharing, this is very useful @cdeotte. Maybe for the final submission, we can submit one solution with noise removal and one without noise removal. It would be interesting to see which one holds the \"test\" of time.",
      "votes": 2
    },
    {
      "id": 1804021,
      "postDate": "2022-05-28T13:05:51.707Z",
      "content": "<p>Seems like a lot of features become binary or categorical once you remove that noise… we might be loosing some info from this on continuous features but for the moment I am mainly happy as this could save a lot of memory. Thanks for sharing.</p>",
      "rawMarkdown": "Seems like a lot of features become binary or categorical once you remove that noise... we might be loosing some info from this on continuous features but for the moment I am mainly happy as this could save a lot of memory. Thanks for sharing.",
      "votes": 2
    },
    {
      "id": 1804007,
      "postDate": "2022-05-28T12:34:13.223Z",
      "content": "<p>Agreed -<br>\nValues between 0 and 0.01. and between 1.0 and 1.01, surely seem to be the case of random noise as pointed by Radar<br>\nAt least for variables like this, it should be easy to denoise</p>",
      "rawMarkdown": "Agreed -\nValues between 0 and 0.01. and between 1.0 and 1.01, surely seem to be the case of random noise as pointed by Radar\nAt least for variables like this, it should be easy to denoise\n\n",
      "votes": 2
    },
    {
      "id": 1806907,
      "postDate": "2022-05-31T15:00:11.197Z",
      "content": "<p>That's an interesting consideration. It's a situation where random noise is being put on the features on a dare.<br>\nThanks for sharing.👐</p>",
      "rawMarkdown": "That's an interesting consideration. It's a situation where random noise is being put on the features on a dare.\nThanks for sharing.👐"
    }
  ],
  "comments": [
    {
      "id": 1809384,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-06-02T16:44:22.320000",
      "content": "<p>UPDATE: Raddar has converted the entire train and test dataset to remove this noise <a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">here</a>. His Kaggle dataset shrinks the original data from 50GB to 5GB! He posted a discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\" target=\"_blank\">here</a>. </p>\n<p>Here is another example of random uniform noise. The variable R_26 has 548160 unique values below 0.2 and is provided as float64 (8 bytes), but really from the plot we see it is only 6 discrete values with (random uniform) noise added and can be converted to int8 (1 byte) - without information loss.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/r_26.png\" alt=\"\"></p>\n<p>Afterward, I published an XGBoost starter notebook using his Kaggle dataset <a href=\"https://tinyurl.com/yb769r9c\" target=\"_blank\">here</a> which achieves CV 0.792 and LB 0.794. This is the same accuracy as other public notebooks using similar features. Therefore it appears that this noise does not contain signal.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1806932,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-05-31T15:26:49.087000",
      "content": "<p>These kind of artifacts don't have good record on Kaggle. Most of the time, people somehow reverse engineer them and gain benefit.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1807045,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-31T16:55:25.147000",
          "content": "<p>haha good point</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1804053,
      "author_name": "MichaelP",
      "author_url": "",
      "post_date": "2022-05-28T13:37:54.040000",
      "content": "<p>Interesting observation. I agree that noise is a possible explanation. Another possibility is that this is a composite feature of a binary 0 or 1 summed with another feature whose typical values might be in the range 0 to 0.1 but predominantly at or close to 0. If that is the case then there could be important information that is lost with converting to float16. It does make sense that this is noise as 2_948_222 unique values for money sounds like a lot, but perhaps it isn't if this includes payments in foreign currencies with constantly moving fx rates being translated back into the currency of the cardholder before rounding to 2 decimal places for the lowest unit of currency. It would be nice if there was more information on the meaning of the features.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1838032,
      "author_name": "Tord Malmgren",
      "author_url": "",
      "post_date": "2022-06-30T06:33:42.220000",
      "content": "<p>If you consider correlation, binary features give you a singular matrix. Adding a bit of noise prevents this. Maybe a reason?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1840858,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-02T15:37:24.933000",
          "content": "<p>I think the reason that Fidelity added noise was to make the data more anonymous. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1804496,
      "author_name": "Hoa Le",
      "author_url": "",
      "post_date": "2022-05-29T04:39:43.177000",
      "content": "<p>I also found some variables with strange histograms. I have 2 questions though:</p>\n<ul>\n<li>Why is it too much to have 2.9M unique values for balance? I think in terms of dollar value, it could be a range from 0.00 to 30,000.00. It would be even more probable for a broader range, e.g. 0.00 - 35,000.00</li>\n<li>Why would the host add noises to the original data? </li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 1804570,
          "author_name": "Cihat Emre Çeliker",
          "author_url": "",
          "post_date": "2022-05-29T07:06:28.980000",
          "content": "<p>Without noise, it would be a lot easier to de-anonymize the data and find the right scale.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1812970,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-06T12:10:20.823000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1805617,
      "author_name": "SHARE at FAU",
      "author_url": "",
      "post_date": "2022-05-30T11:36:26.497000",
      "content": "<p>Interesting!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1805389,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2022-05-30T05:49:47.310000",
      "content": "<p>Thanks for sharing, this is very useful <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Maybe for the final submission, we can submit one solution with noise removal and one without noise removal. It would be interesting to see which one holds the \"test\" of time.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1804021,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-05-28T13:05:51.707000",
      "content": "<p>Seems like a lot of features become binary or categorical once you remove that noise… we might be loosing some info from this on continuous features but for the moment I am mainly happy as this could save a lot of memory. Thanks for sharing.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1804007,
      "author_name": "Seeker",
      "author_url": "",
      "post_date": "2022-05-28T12:34:13.223000",
      "content": "<p>Agreed -<br>\nValues between 0 and 0.01. and between 1.0 and 1.01, surely seem to be the case of random noise as pointed by Radar<br>\nAt least for variables like this, it should be easy to denoise</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1806907,
      "author_name": "chinchilla",
      "author_url": "",
      "post_date": "2022-05-31T15:00:11.197000",
      "content": "<p>That's an interesting consideration. It's a situation where random noise is being put on the features on a dare.<br>\nThanks for sharing.👐</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1803973": "# Strange Histograms\nI noticed that many features have strange histograms. For example, at first glance, it appears that column `B_8` is a binary feature that has values 0 and 1.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b1.png)\n\nHowever if we zoom in, we see that there are many values between 0 and 0.01. And many values between 1.0 and 1.01\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b2.png)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/b3.png)\n\nSpecifically in the train data with `5_531_451` rows, there are `3_053_857` values between 0.0 and 0.1. And there are `2_455_326` values between 1.0 and 1.1.\n\n# Questions\nThis prompts questions. \n* If we convert this column into two values, namely 0 and 1, do we lose information?\n* Note when we convert this column from `float32` to `float16` we decrease the number of unique values from `2_948_222` to `8_501`. Does this lose important information? \n* Why are so many values uniformly distributed between 0.0 and 0.01? Uniform distributions seldom happen in the real world. Did AMEX just add uniform noise to the 0's and noise to the 1's?\n* If we use an NN, do we need to process this column differently than standard scaler for the NN to access all these tightly packed values?\n\n# UPDATE\nAt the same time I posted this discussion, Raddar posted a similar topic. Raddar hypothesizes that random uniform noise has been injected into the data. His discussion is [here][1] and notebook [here][2]. He points out that columns related to money such as `B_8` (which is a `balance variable`) should have less unique values than  `2_948_222`. The existence of many unique values supports the argument that random noise has been added.\n\n## UPDATE2 - XGB Starter Notebook - LB 794 - without Noise\nRaddar has converted the entire train and test dataset to remove this noise [here][5]. His Kaggle dataset shrinks the original data from 50GB to 5GB! He posted a discussion [here][3]. \n\nAfterward, I published an XGBoost starter notebook using his Kaggle dataset [here][4] which achieves CV 0.792 and LB 0.794. This is the same accuracy as other public notebooks using similar features. Therefore it appears that this noise does not contain signal.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327649\n[2]: https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\n[3]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n[4]: https://tinyurl.com/yb769r9c\n[5]: https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
    "1809384": "UPDATE: Raddar has converted the entire train and test dataset to remove this noise [here][5]. His Kaggle dataset shrinks the original data from 50GB to 5GB! He posted a discussion [here][3]. \n\nHere is another example of random uniform noise. The variable R_26 has 548160 unique values below 0.2 and is provided as float64 (8 bytes), but really from the plot we see it is only 6 discrete values with (random uniform) noise added and can be converted to int8 (1 byte) - without information loss.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/r_26.png)\n\nAfterward, I published an XGBoost starter notebook using his Kaggle dataset [here][4] which achieves CV 0.792 and LB 0.794. This is the same accuracy as other public notebooks using similar features. Therefore it appears that this noise does not contain signal.\n\n[3]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/328514\n[4]: https://tinyurl.com/yb769r9c\n[5]: https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format",
    "1806932": "These kind of artifacts don't have good record on Kaggle. Most of the time, people somehow reverse engineer them and gain benefit.",
    "1804053": "Interesting observation. I agree that noise is a possible explanation. Another possibility is that this is a composite feature of a binary 0 or 1 summed with another feature whose typical values might be in the range 0 to 0.1 but predominantly at or close to 0. If that is the case then there could be important information that is lost with converting to float16. It does make sense that this is noise as 2_948_222 unique values for money sounds like a lot, but perhaps it isn't if this includes payments in foreign currencies with constantly moving fx rates being translated back into the currency of the cardholder before rounding to 2 decimal places for the lowest unit of currency. It would be nice if there was more information on the meaning of the features.",
    "1838032": "If you consider correlation, binary features give you a singular matrix. Adding a bit of noise prevents this. Maybe a reason?",
    "1804496": "I also found some variables with strange histograms. I have 2 questions though:\n- Why is it too much to have 2.9M unique values for balance? I think in terms of dollar value, it could be a range from 0.00 to 30,000.00. It would be even more probable for a broader range, e.g. 0.00 - 35,000.00\n- Why would the host add noises to the original data? ",
    "1805617": "Interesting!",
    "1805389": "Thanks for sharing, this is very useful @cdeotte. Maybe for the final submission, we can submit one solution with noise removal and one without noise removal. It would be interesting to see which one holds the \"test\" of time.",
    "1804021": "Seems like a lot of features become binary or categorical once you remove that noise... we might be loosing some info from this on continuous features but for the moment I am mainly happy as this could save a lot of memory. Thanks for sharing.",
    "1804007": "Agreed -\nValues between 0 and 0.01. and between 1.0 and 1.01, surely seem to be the case of random noise as pointed by Radar\nAt least for variables like this, it should be easy to denoise\n\n",
    "1806907": "That's an interesting consideration. It's a situation where random noise is being put on the features on a dare.\nThanks for sharing.👐"
  }
}