{
  "id": 328514,
  "title": "Integer columns in the data - here you go!",
  "url": "/competitions/amex-default-prediction/discussion/328514",
  "author_name": "raddar",
  "post_date": "2022-06-01T16:04:05.623000",
  "votes": 834,
  "comment_count": 116,
  "views": 0,
  "content": "<p>As it is already clear - all float type columns have random uniform noise of [0,0.01] added to each column. Having this information it is clear that it is very easy to spot integer type columns in the raw data. This calls for detective work, which I really love and have done many times in previous competitions :) </p>\n<p>It took me a couple of days, but I have it ready for you all to use!</p>\n<p>So, what's in it?</p>\n<ul>\n<li>originally we had 188 float/categorical type features. These were transformed into<ul>\n<li>95 np.int8/np.int16 types</li>\n<li>93 np.float32 types</li></ul></li>\n<li>Most float columns with [0, 0.01] and [1, 1.01] have these values rounded up at 0 and 1 respectively. This was done to ensure no data loss, as not all features could be rounded up safely.</li>\n<li>saved in parquet format (only 1.7GB training data!)</li>\n</ul>\n<p>Hopefully, this opens the door for many people to access this competition, as less RAM is required. </p>\n<p>Also more interesting feature engineering will be available as it is much easier to work with integers (my experience).</p>\n<p>The cleaned dataset can be found as a dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></p>\n<p>Notebooks used to transform the data:<br>\n<a href=\"https://www.kaggle.com/code/raddar/amex-data-int-types-train\" target=\"_blank\">https://www.kaggle.com/code/raddar/amex-data-int-types-train</a><br>\n<a href=\"https://www.kaggle.com/code/raddar/amex-data-int-types-test\" target=\"_blank\">https://www.kaggle.com/code/raddar/amex-data-int-types-test</a></p>\n<p>As many youtubers would say - please like and subscribe! This helps me to be motivated to share stuff with you all :) More things to come.</p>\n<p>UPDATE: </p>\n<p>dataset was updated with 3 extra int conversions</p>",
  "messages": [
    {
      "id": 1808215,
      "postDate": "2022-06-01T16:04:05.623Z",
      "content": "<p>As it is already clear - all float type columns have random uniform noise of [0,0.01] added to each column. Having this information it is clear that it is very easy to spot integer type columns in the raw data. This calls for detective work, which I really love and have done many times in previous competitions :) </p>\n<p>It took me a couple of days, but I have it ready for you all to use!</p>\n<p>So, what's in it?</p>\n<ul>\n<li>originally we had 188 float/categorical type features. These were transformed into<ul>\n<li>95 np.int8/np.int16 types</li>\n<li>93 np.float32 types</li></ul></li>\n<li>Most float columns with [0, 0.01] and [1, 1.01] have these values rounded up at 0 and 1 respectively. This was done to ensure no data loss, as not all features could be rounded up safely.</li>\n<li>saved in parquet format (only 1.7GB training data!)</li>\n</ul>\n<p>Hopefully, this opens the door for many people to access this competition, as less RAM is required. </p>\n<p>Also more interesting feature engineering will be available as it is much easier to work with integers (my experience).</p>\n<p>The cleaned dataset can be found as a dataset:<br>\n<a href=\"https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\" target=\"_blank\">https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format</a></p>\n<p>Notebooks used to transform the data:<br>\n<a href=\"https://www.kaggle.com/code/raddar/amex-data-int-types-train\" target=\"_blank\">https://www.kaggle.com/code/raddar/amex-data-int-types-train</a><br>\n<a href=\"https://www.kaggle.com/code/raddar/amex-data-int-types-test\" target=\"_blank\">https://www.kaggle.com/code/raddar/amex-data-int-types-test</a></p>\n<p>As many youtubers would say - please like and subscribe! This helps me to be motivated to share stuff with you all :) More things to come.</p>\n<p>UPDATE: </p>\n<p>dataset was updated with 3 extra int conversions</p>",
      "rawMarkdown": "As it is already clear - all float type columns have random uniform noise of [0,0.01] added to each column. Having this information it is clear that it is very easy to spot integer type columns in the raw data. This calls for detective work, which I really love and have done many times in previous competitions :) \n\nIt took me a couple of days, but I have it ready for you all to use!\n\nSo, what's in it?\n \n- originally we had 188 float/categorical type features. These were transformed into\n   - 95 np.int8/np.int16 types\n   - 93 np.float32 types\n- Most float columns with [0, 0.01] and [1, 1.01] have these values rounded up at 0 and 1 respectively. This was done to ensure no data loss, as not all features could be rounded up safely.\n- saved in parquet format (only 1.7GB training data!)\n\nHopefully, this opens the door for many people to access this competition, as less RAM is required. \n\nAlso more interesting feature engineering will be available as it is much easier to work with integers (my experience).\n\nThe cleaned dataset can be found as a dataset:\nhttps://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n\nNotebooks used to transform the data:\nhttps://www.kaggle.com/code/raddar/amex-data-int-types-train\nhttps://www.kaggle.com/code/raddar/amex-data-int-types-test\n\n\nAs many youtubers would say - please like and subscribe! This helps me to be motivated to share stuff with you all :) More things to come.\n\nUPDATE: \n\ndataset was updated with 3 extra int conversions",
      "votes": 832
    },
    {
      "id": 1826150,
      "postDate": "2022-06-20T04:57:04.827Z",
      "content": "<p><img src=\"https://i.ibb.co/6Fyt3Mk/Selection-905.png\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.ibb.co/6Fyt3Mk/Selection-905.png)",
      "votes": 68
    },
    {
      "id": 1808521,
      "postDate": "2022-06-01T23:49:45.287Z",
      "content": "<p>I published an XGB starter notebook using your data with CV 0.792 and LB 0.794 <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a>. This data is great. It is so small, that both the feature engineering and model training can easily occur within one Kaggle notebook. The total notebook time (reading data, feature engineering, training 5 folds, inferring test) is only 13 minutes! Everything is done on GPU!</p>",
      "rawMarkdown": "I published an XGB starter notebook using your data with CV 0.792 and LB 0.794 [here][1]. This data is great. It is so small, that both the feature engineering and model training can easily occur within one Kaggle notebook. The total notebook time (reading data, feature engineering, training 5 folds, inferring test) is only 13 minutes! Everything is done on GPU!\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
      "votes": 44
    },
    {
      "id": 1808256,
      "postDate": "2022-06-01T16:54:35.137Z",
      "content": "<p>Fantastic job raddar. I was just doing this myself today. When we zoom in on some variables, we see that without noise they are actually discrete values. </p>\n<p>For example <code>R_26</code>. Below the histogram is displayed with noise. The variable <code>R_26</code> has <code>548160</code> unique values below 0.2 and is provided as <code>float64</code> (8 bytes), but really from the plot we see it is only 6 discrete values with (random uniform) noise added and can be converted to <code>int8</code> (1 byte) - without information loss. </p>\n<p>It's funny how the data is actually 5GB with 45GB of noise added lol 😄</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/r_26.png\" alt=\"\"></p>",
      "rawMarkdown": "Fantastic job raddar. I was just doing this myself today. When we zoom in on some variables, we see that without noise they are actually discrete values. \n\nFor example `R_26`. Below the histogram is displayed with noise. The variable `R_26` has `548160` unique values below 0.2 and is provided as `float64` (8 bytes), but really from the plot we see it is only 6 discrete values with (random uniform) noise added and can be converted to `int8` (1 byte) - without information loss. \n\nIt's funny how the data is actually 5GB with 45GB of noise added lol 😄\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/r_26.png)",
      "votes": 41,
      "replies": [
        {
          "id": 1808273,
          "postDate": "2022-06-01T17:12:43.163Z",
          "content": "<p>makes you question why they added the noise in the first place. It was obvious this would happen :)</p>",
          "rawMarkdown": "makes you question why they added the noise in the first place. It was obvious this would happen :)",
          "votes": 19
        },
        {
          "id": 1808343,
          "postDate": "2022-06-01T18:11:49.853Z",
          "content": "<p>My best  theory is that they had nans in their integer columns (using another langage maybe) and python didn't allow them to keep those columns as integers. There still is a question about adding noise to floats.</p>",
          "rawMarkdown": "My best  theory is that they had nans in their integer columns (using another langage maybe) and python didn't allow them to keep those columns as integers. There still is a question about adding noise to floats.",
          "votes": 4
        },
        {
          "id": 1816727,
          "postDate": "2022-06-10T13:56:12.303Z",
          "content": "<p>In association football, the extra time is usually in whole minutes (human ref), but it seldom lasts exactly on the minutes (human ref), so there are narrow distributions around the whole minutes. But, should they be rounded without loss of information?</p>",
          "rawMarkdown": "In association football, the extra time is usually in whole minutes (human ref), but it seldom lasts exactly on the minutes (human ref), so there are narrow distributions around the whole minutes. But, should they be rounded without loss of information?",
          "votes": 2,
          "replies": [
            {
              "id": 2936737,
              "postDate": "2024-07-26T12:22:26.717Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 1808994,
      "postDate": "2022-06-02T09:57:16.710Z",
      "content": "<p>raddar - You are our Sherlock Holmes 🕵️</p>",
      "rawMarkdown": "raddar - You are our Sherlock Holmes 🕵️",
      "votes": 24
    },
    {
      "id": 1810193,
      "postDate": "2022-06-03T10:36:45.123Z",
      "content": "<p>I just made an update to the dataset, with 3 new features converted to int: <code>B_19</code>, <code>S_8</code> and <code>S_13</code>. notebooks updated as well :)</p>",
      "rawMarkdown": "I just made an update to the dataset, with 3 new features converted to int: `B_19`, `S_8` and `S_13`. notebooks updated as well :)",
      "votes": 12,
      "replies": [
        {
          "id": 1810446,
          "postDate": "2022-06-03T15:11:38.037Z",
          "content": "<p>Great work, i need to update my XGB notebook version. After each update, it would be interesting to take note of the XGB CV score. I wonder if the CV LB will keep improving as you remove noise from more columns.</p>\n<p>FYI for everyone. If you want an older version of Raddar's dataset, change the URL below to get whichever version that you want</p>\n<p><code>https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format/versions/2</code></p>\n<p>The number at the end is the version number. Currently Raddar has versions 1 and 2.</p>",
          "rawMarkdown": "Great work, i need to update my XGB notebook version. After each update, it would be interesting to take note of the XGB CV score. I wonder if the CV LB will keep improving as you remove noise from more columns.\n\nFYI for everyone. If you want an older version of Raddar's dataset, change the URL below to get whichever version that you want\n\n`https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format/versions/2`\n\nThe number at the end is the version number. Currently Raddar has versions 1 and 2.",
          "votes": 11
        }
      ]
    },
    {
      "id": 1812195,
      "postDate": "2022-06-05T15:14:20.360Z",
      "content": "<p>Hey there! Thank a lot <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> . </p>\n<p>I aggregated your parquet datasets with <a href=\"https://www.kaggle.com/huseyincot\" target=\"_blank\">@huseyincot</a> 's method! In case its useful for anyone, it improved my score a little bit.</p>\n<p>You can find the aggregated datasets <a href=\"https://www.kaggle.com/datasets/heyspaceturtle/agg-amex-data-without-noise\" target=\"_blank\">here</a><br>\nAnd my notebook <a href=\"https://www.kaggle.com/code/heyspaceturtle/data-aggregation?scriptVersionId=97546977\" target=\"_blank\">here</a></p>\n<p>good vibes! </p>",
      "rawMarkdown": "Hey there! Thank a lot @raddar . \n\nI aggregated your parquet datasets with @huseyincot 's method! In case its useful for anyone, it improved my score a little bit.\n\nYou can find the aggregated datasets [here](https://www.kaggle.com/datasets/heyspaceturtle/agg-amex-data-without-noise)\nAnd my notebook [here](https://www.kaggle.com/code/heyspaceturtle/data-aggregation?scriptVersionId=97546977)\n\ngood vibes! ",
      "votes": 8
    },
    {
      "id": 1837347,
      "postDate": "2022-06-29T14:00:27.293Z",
      "content": "<p>Thanks a lot, I still do not understand the meaning of \"random uniform noise\"?<br>\n please if anyone has a notebook that illustrates this idea, I would be glad if you share it with me.</p>",
      "rawMarkdown": "Thanks a lot, I still do not understand the meaning of \"random uniform noise\"?\n please if anyone has a notebook that illustrates this idea, I would be glad if you share it with me.",
      "votes": 5,
      "replies": [
        {
          "id": 1837480,
          "postDate": "2022-06-29T16:18:05.053Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\" target=\"_blank\">https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added</a></p>",
          "rawMarkdown": "https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added",
          "votes": 13
        }
      ]
    },
    {
      "id": 1847129,
      "postDate": "2022-07-07T16:46:18.600Z",
      "content": "<p>RADDAR: Read All Data, Denoise And Release</p>",
      "rawMarkdown": "RADDAR: Read All Data, Denoise And Release",
      "votes": 6
    },
    {
      "id": 1810347,
      "postDate": "2022-06-03T13:35:50.690Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, I'm still trying to figure out how did you set those intervals for floorify_frac?</p>\n<p>Let's say we have this feature B_4 histogram.</p>\n<p><img src=\"https://i.ibb.co/MSmCp6J/Screenshot-from-2022-06-03-16-13-15.png\" alt=\"1\"></p>\n<p>When we zoom in, we see those blocks.</p>\n<p><img src=\"https://i.ibb.co/znBSkst/Screenshot-from-2022-06-03-16-14-15.png\" alt=\"2\"></p>\n<p>When we zoom even more, we can see that +-0.01 random uniform noise added to them.</p>\n<p><img src=\"https://i.ibb.co/7KnyY8j/Screenshot-from-2022-06-03-16-21-08.png\" alt=\"3\"></p>\n<p>As far as I understand, they scaled raw features with a coefficient so that the injected noise won't be insignificant and won't cause any information loss at the same time.</p>\n<p>I guess any scaling factor that can project those blocks into the same integer value can be used here. 1 / 78 doesn't necessarily have to be the only valid value. I was wondering how did you empirically come up with those values?</p>",
      "rawMarkdown": "Hello @raddar, I'm still trying to figure out how did you set those intervals for floorify_frac?\n\nLet's say we have this feature B_4 histogram.\n\n![1](https://i.ibb.co/MSmCp6J/Screenshot-from-2022-06-03-16-13-15.png)\n\nWhen we zoom in, we see those blocks.\n\n![2](https://i.ibb.co/znBSkst/Screenshot-from-2022-06-03-16-14-15.png)\n\nWhen we zoom even more, we can see that +-0.01 random uniform noise added to them.\n\n![3](https://i.ibb.co/7KnyY8j/Screenshot-from-2022-06-03-16-21-08.png)\n\nAs far as I understand, they scaled raw features with a coefficient so that the injected noise won't be insignificant and won't cause any information loss at the same time.\n\nI guess any scaling factor that can project those blocks into the same integer value can be used here. 1 / 78 doesn't necessarily have to be the only valid value. I was wondering how did you empirically come up with those values?",
      "votes": 6,
      "replies": [
        {
          "id": 1810434,
          "postDate": "2022-06-03T15:00:41.280Z",
          "content": "<p>I did the very same thing you described - I zoomed in into very small interval space to identify these bins and then worked on that.</p>\n<p>We know that each bin is exactly 0.01 wide (due to uniform [0,0.01] random). This means each bin is represented by its minimum value - which is its original value before random noise. Then let's take 2 consecutive bins and look at the difference - this represents one or more steps between bins. Find the smallest difference (lowest step) - which is indicative to the increment by 1 in the integer space. then simpy 1/X of that value represents the common denominator..</p>\n<p>Other way to think is - we do simple optimization, like a pseudo code:</p>\n<pre><code>for i in range(N):\n  if x['B_8'] * i ~ integer for all values:\n    return i\n</code></pre>\n<p>So simply scanning for smallest value <code>i</code> to have all values as smallest possible integers.</p>\n<p>Of course 1/78 is not the only value. it is as valid as 1/(2 x 78), 1/(3 x 78)… however 2x, 3x multiplier would just mean that integers would be multiplied by 2x, 3x…</p>",
          "rawMarkdown": "I did the very same thing you described - I zoomed in into very small interval space to identify these bins and then worked on that.\n\nWe know that each bin is exactly 0.01 wide (due to uniform [0,0.01] random). This means each bin is represented by its minimum value - which is its original value before random noise. Then let's take 2 consecutive bins and look at the difference - this represents one or more steps between bins. Find the smallest difference (lowest step) - which is indicative to the increment by 1 in the integer space. then simpy 1/X of that value represents the common denominator..\n\nOther way to think is - we do simple optimization, like a pseudo code:\n```\nfor i in range(N):\n  if x['B_8'] * i ~ integer for all values:\n    return i\n ```\n\nSo simply scanning for smallest value `i` to have all values as smallest possible integers.\n\nOf course 1/78 is not the only value. it is as valid as 1/(2 x 78), 1/(3 x 78)... however 2x, 3x multiplier would just mean that integers would be multiplied by 2x, 3x...",
          "votes": 14
        },
        {
          "id": 1810460,
          "postDate": "2022-06-03T15:26:00.100Z",
          "content": "<p>Thanks for the explanation. I thought the random uniform noise was between -0.01 and 0.01. Now it makes sense why you pick bins' minimum values.</p>\n<p>I checked steps of multiple bins in B_4 and I found that the step size is 0.013. My multiplier is slightly different than yours. Do you think the step size might not be fixed for some bins? I'll try to calculate those step sizes programmatically to see how many unique values are there.</p>",
          "rawMarkdown": "Thanks for the explanation. I thought the random uniform noise was between -0.01 and 0.01. Now it makes sense why you pick bins' minimum values.\n\nI checked steps of multiple bins in B_4 and I found that the step size is 0.013. My multiplier is slightly different than yours. Do you think the step size might not be fixed for some bins? I'll try to calculate those step sizes programmatically to see how many unique values are there.",
          "votes": 2
        },
        {
          "id": 1810483,
          "postDate": "2022-06-03T15:40:49.917Z",
          "content": "<p>I guess you had <code>B_4</code> in mind?</p>",
          "rawMarkdown": "I guess you had `B_4` in mind?"
        },
        {
          "id": 1810489,
          "postDate": "2022-06-03T15:42:50.290Z",
          "content": "<p>Yes, I meant B_4. I made a typo.</p>",
          "rawMarkdown": "Yes, I meant B_4. I made a typo."
        },
        {
          "id": 1810527,
          "postDate": "2022-06-03T16:16:10.777Z",
          "content": "<p>Maybe you are working with float16? they do lose a lot of information.</p>\n<pre><code>z = pd.read_csv('train_data.csv',usecols=['B_4'])\nv = z.loc[(z.B_4&gt;0.615)].min() - z.loc[(z.B_4&gt;0.6)].min()\n1/v\n\n&gt; B_4    78.002257\n</code></pre>",
          "rawMarkdown": "Maybe you are working with float16? they do lose a lot of information.\n\n```\nz = pd.read_csv('train_data.csv',usecols=['B_4'])\nv = z.loc[(z.B_4>0.615)].min() - z.loc[(z.B_4>0.6)].min()\n1/v\n\n> B_4    78.002257\n```",
          "votes": 2
        },
        {
          "id": 1810553,
          "postDate": "2022-06-03T16:53:58.467Z",
          "content": "<p>I tried with 32 and 64 bits but there wasn't any difference. I think the difference arises from the histogram function. Check the image below. I was using 0.6145 as the bin edge and step size was different for those particular bins. I think your calculation is correct. I have find the optimal value by searching the dataframe, not the visualization.</p>\n<p><img src=\"https://i.ibb.co/Gn59cwn/Screenshot-from-2022-06-03-19-49-13.png\" alt=\"1\"></p>",
          "rawMarkdown": "I tried with 32 and 64 bits but there wasn't any difference. I think the difference arises from the histogram function. Check the image below. I was using 0.6145 as the bin edge and step size was different for those particular bins. I think your calculation is correct. I have find the optimal value by searching the dataframe, not the visualization.\n\n![1](https://i.ibb.co/Gn59cwn/Screenshot-from-2022-06-03-19-49-13.png)",
          "votes": 1
        },
        {
          "id": 1839869,
          "postDate": "2022-07-01T18:49:05.760Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1837124,
      "postDate": "2022-06-29T11:06:04.197Z",
      "content": "<p>hello <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, very brilliant reverse engineering (\"DeAnonymization\")</p>\n<p>However there are some ambiguous passages for me:</p>\n<ol>\n<li>what do you mean by \"<strong>each bin has AUC of 0.5</strong>\" for the the <code>B_19</code> feature ?  Do you have a related sample code ? thx..</li>\n<li><code>x['D_59'] = floorify_frac(x['D_59']+5/48,1/48)</code> why did you have to add some steps before reconstructing the feature (same for <code>D_124</code>)</li>\n<li>I did'nt understand the cross referencing with <code>S_13</code> of <code>S_11</code>  : because I didn't find any data in : <code>train.S_11.isin([15,16,17])</code> </li>\n</ol>\n<p>Thx for considering my questions😊</p>",
      "rawMarkdown": "hello @raddar, very brilliant reverse engineering (\"DeAnonymization\")\n\nHowever there are some ambiguous passages for me:\n1. what do you mean by \"**each bin has AUC of 0.5**\" for the the `B_19` feature ?  Do you have a related sample code ? thx..\n1. `x['D_59'] = floorify_frac(x['D_59']+5/48,1/48)` why did you have to add some steps before reconstructing the feature (same for `D_124`)\n1. I did'nt understand the cross referencing with `S_13` of `S_11`  : because I didn't find any data in : `train.S_11.isin([15,16,17])` \n\n\nThx for considering my questions😊",
      "votes": 4,
      "replies": [
        {
          "id": 1837493,
          "postDate": "2022-06-29T16:22:10.497Z",
          "content": "<ol>\n<li><p>I do not have code. The idea is: take a bin with values in range [0,0.01]. calculate AUC of that range. Then calculate AUC for range [0.01,0.02], etc. smth like <code>roc_auc_score(x.loc[(x.B_19&gt;=0.01) &amp; (x.B_19&lt;=0.02),'target'], x.loc[(x.B_19&gt;=0.01) &amp; (x.B_19&lt;=0.02),'B_19'])</code>. </p></li>\n<li><p>I wanted to make -1 to consistently represent NA. D_59 originally had some negative values. This should have no effect to models. This helps to reverse -1 to NA by just simply <code>x[x==-1] = np.nan</code></p></li>\n<li><p>S_13 and S_11 were already converted to integers, so <code>train.S_11.isin([15,16,17])</code> exists.</p></li>\n</ol>",
          "rawMarkdown": "1. I do not have code. The idea is: take a bin with values in range [0,0.01]. calculate AUC of that range. Then calculate AUC for range [0.01,0.02], etc. smth like `roc_auc_score(x.loc[(x.B_19>=0.01) & (x.B_19<=0.02),'target'], x.loc[(x.B_19>=0.01) & (x.B_19<=0.02),'B_19'])`. \n\n2. I wanted to make -1 to consistently represent NA. D_59 originally had some negative values. This should have no effect to models. This helps to reverse -1 to NA by just simply `x[x==-1] = np.nan`\n\n3. S_13 and S_11 were already converted to integers, so `train.S_11.isin([15,16,17])` exists.",
          "votes": 6
        },
        {
          "id": 1838441,
          "postDate": "2022-06-30T14:20:58.683Z",
          "content": "<p>Manyyy thanks it is pretty clear 🙏😊</p>",
          "rawMarkdown": "Manyyy thanks it is pretty clear 🙏😊"
        },
        {
          "id": 1840845,
          "postDate": "2022-07-02T15:28:14.727Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1840859,
          "postDate": "2022-07-02T15:38:31.673Z",
          "content": "<p>If there was a signal (as in increasing/decreasing target rate) within each bin ([0,0.01], [0.01,0.02], etc.), then AUC!=0.5. </p>\n<p>So by testing for AUC=0.5 we test the hypothesis that there is no target signal in each bin. </p>\n<p>If there is no signal - there is no mix of many original values in the bin, meaning that the bin is represented by only one number - the min value of the bin.</p>",
          "rawMarkdown": "If there was a signal (as in increasing/decreasing target rate) within each bin ([0,0.01], [0.01,0.02], etc.), then AUC!=0.5. \n\nSo by testing for AUC=0.5 we test the hypothesis that there is no target signal in each bin. \n\nIf there is no signal - there is no mix of many original values in the bin, meaning that the bin is represented by only one number - the min value of the bin.",
          "votes": 13
        },
        {
          "id": 1840880,
          "postDate": "2022-07-02T15:51:56.250Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1808471,
      "postDate": "2022-06-01T21:09:28.613Z",
      "content": "<p>HI <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> It seems that <code>B_19</code> would profit from <code>floorify_frac(x['B_19'])</code> as well. And the original values of <code>S_13</code> are all multiples of 1/1034, but as 1/1034 &lt; 0.01, the added noise cannot be removed (diagrams are at the end of <a href=\"https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense\" target=\"_blank\">my EDA</a>).</p>",
      "rawMarkdown": "HI @raddar It seems that `B_19` would profit from `floorify_frac(x['B_19'])` as well. And the original values of `S_13` are all multiples of 1/1034, but as 1/1034 < 0.01, the added noise cannot be removed (diagrams are at the end of [my EDA](https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense)).",
      "votes": 3,
      "replies": [
        {
          "id": 1808482,
          "postDate": "2022-06-01T21:29:20.917Z",
          "content": "<p>They are both inconclusive to me as intervals may overlap, so I left them as is.</p>",
          "rawMarkdown": "They are both inconclusive to me as intervals may overlap, so I left them as is.",
          "votes": 3
        },
        {
          "id": 1809184,
          "postDate": "2022-06-02T13:31:09.013Z",
          "content": "<p>I revisited your both var suggestions. I was able to convert them to int:)</p>\n<p><code>B_19</code> can be easily rounded every 0.01 (if you check AUC scores within each bin, AUC is ~0.5, so no target variance within each bin)</p>\n<p>as for S_13 I used another trick, cross referencing S_11</p>\n<pre><code>def floorify(x, lo):\n    \"\"\"example: x in [0, 0.01] -&gt; x := 0\"\"\"\n    return lo if x &lt;= lo+0.01 and x &gt;= lo else x\n\n### S_13 has these weird ordinal values\n# one value overlaps, but can be split by S_11\nx.loc[(x.S_13&gt;=0.67) &amp; (x.S_13&lt;=0.7) &amp; (x.S_11.isin([15,16,17])),'S_13'] = 0.67891681\nx.loc[(x.S_13&gt;=0.67) &amp; (x.S_13&lt;=0.7) &amp; ~(x.S_11.isin([15,16,17])),'S_13'] = 0.68762086\n\nfor c in (0.03771764, 0.28046423, 0.40135398, 0.42069634, 0.50676983, 0.52611219, 0.55512583, 0.62185686, 0.84332692):\n    x['S_13'] = x['S_13'].apply(lambda t: floorify(t,c))\n\nx['S_13'] = np.round(x['S_13']*1034)\n</code></pre>\n<p>so 2 more int features in the stack :) </p>",
          "rawMarkdown": "I revisited your both var suggestions. I was able to convert them to int:)\n\n`B_19` can be easily rounded every 0.01 (if you check AUC scores within each bin, AUC is ~0.5, so no target variance within each bin)\n\nas for S_13 I used another trick, cross referencing S_11\n\n```\ndef floorify(x, lo):\n    \"\"\"example: x in [0, 0.01] -> x := 0\"\"\"\n    return lo if x <= lo+0.01 and x >= lo else x\n\n### S_13 has these weird ordinal values\n# one value overlaps, but can be split by S_11\nx.loc[(x.S_13>=0.67) & (x.S_13<=0.7) & (x.S_11.isin([15,16,17])),'S_13'] = 0.67891681\nx.loc[(x.S_13>=0.67) & (x.S_13<=0.7) & ~(x.S_11.isin([15,16,17])),'S_13'] = 0.68762086\n\nfor c in (0.03771764, 0.28046423, 0.40135398, 0.42069634, 0.50676983, 0.52611219, 0.55512583, 0.62185686, 0.84332692):\n    x['S_13'] = x['S_13'].apply(lambda t: floorify(t,c))\n    \nx['S_13'] = np.round(x['S_13']*1034)\n```\n\nso 2 more int features in the stack :) \n\n\n\n",
          "votes": 4
        },
        {
          "id": 1809208,
          "postDate": "2022-06-02T14:02:41.170Z",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> What you found out about S_13 is impressive! How did I you figure out S_13 can be split using S_11? </p>",
          "rawMarkdown": "@raddar What you found out about S_13 is impressive! How did I you figure out S_13 can be split using S_11? ",
          "votes": 1
        },
        {
          "id": 1809242,
          "postDate": "2022-06-02T14:41:08.870Z",
          "content": "<p>Simply slicing that S_13 range, sorting, and opening that in excel. obvious pattern emerged</p>",
          "rawMarkdown": "Simply slicing that S_13 range, sorting, and opening that in excel. obvious pattern emerged",
          "votes": 12
        }
      ]
    },
    {
      "id": 1808410,
      "postDate": "2022-06-01T19:28:28.427Z",
      "content": "<p>There are also many small tweaks that can be made on this dataset, for example:</p>\n<pre><code>x.loc[(x.R_13==0) &amp; (x.R_17==0) &amp; (x.R_20==0) &amp; (x.R_8==0), 'R_6'] = 0\nx.loc[x.B_39==-1, 'B_36'] = 0\n</code></pre>\n<p>Do not know if it is worth the effort though to look for more</p>",
      "rawMarkdown": "There are also many small tweaks that can be made on this dataset, for example:\n\n```\nx.loc[(x.R_13==0) & (x.R_17==0) & (x.R_20==0) & (x.R_8==0), 'R_6'] = 0\nx.loc[x.B_39==-1, 'B_36'] = 0\n```\n\nDo not know if it is worth the effort though to look for more\n",
      "votes": 4
    },
    {
      "id": 1847883,
      "postDate": "2022-07-08T08:19:06.680Z",
      "content": "<p>Thank you for sharing this. It seems that D_128, D_130, B_8, and R_1 are also hidden int features.</p>",
      "rawMarkdown": "Thank you for sharing this. It seems that D_128, D_130, B_8, and R_1 are also hidden int features.",
      "votes": 1,
      "replies": [
        {
          "id": 1851080,
          "postDate": "2022-07-11T02:30:19.823Z",
          "content": "<p>Sorry, it was not true. ROCs are not 0.5.</p>",
          "rawMarkdown": "Sorry, it was not true. ROCs are not 0.5.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1809629,
      "postDate": "2022-06-02T21:43:26.367Z",
      "content": "<p>Great insight! Even <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> worked on immensely reducing the size of data so that it can fit into our poor RAMs.</p>",
      "rawMarkdown": "Great insight! Even @cdeotte worked on immensely reducing the size of data so that it can fit into our poor RAMs.",
      "votes": 1
    },
    {
      "id": 1895681,
      "postDate": "2022-08-12T10:10:09.643Z",
      "content": "<p>Hi Radar. Kudos to you. This is not only good work but also saves us a lot of time for everyone.</p>",
      "rawMarkdown": "Hi Radar. Kudos to you. This is not only good work but also saves us a lot of time for everyone.",
      "votes": 2
    },
    {
      "id": 1811436,
      "postDate": "2022-06-04T17:15:36.567Z",
      "content": "<p>Looks like these 93 float32 features hold the majority of the signal. I get CV 0.78 easy by just using them.<br>\nThanks!</p>",
      "rawMarkdown": "Looks like these 93 float32 features hold the majority of the signal. I get CV 0.78 easy by just using them.\nThanks!",
      "votes": 2
    },
    {
      "id": 1809575,
      "postDate": "2022-06-02T20:11:36.813Z",
      "content": "<p>Next step is to know what that features may be? 👀</p>",
      "rawMarkdown": "Next step is to know what that features may be? 👀",
      "votes": 2
    },
    {
      "id": 1809157,
      "postDate": "2022-06-02T13:02:51.480Z",
      "content": "<p>Wow ! You just saved for all of us a lot of hours finding the hidden denominator for every single variable and preprocess it for us in a super clean light dataset.</p>\n<p>Thanks a lot <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> ! </p>",
      "rawMarkdown": "Wow ! You just saved for all of us a lot of hours finding the hidden denominator for every single variable and preprocess it for us in a super clean light dataset.\n\nThanks a lot @raddar ! ",
      "votes": 2
    },
    {
      "id": 1808399,
      "postDate": "2022-06-01T19:17:30.627Z",
      "content": "<p>Nice catch <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, I was worried about converting some of these variables into float16 since we were losing precision in the process, but at the same time thinking about models might converge faster without that noise. I was planning to compare both with and w/o noise data models in cross validation, have you tried that yourself, seen any noticeable differences?</p>",
      "rawMarkdown": "Nice catch @raddar, I was worried about converting some of these variables into float16 since we were losing precision in the process, but at the same time thinking about models might converge faster without that noise. I was planning to compare both with and w/o noise data models in cross validation, have you tried that yourself, seen any noticeable differences?",
      "votes": 2,
      "replies": [
        {
          "id": 1808407,
          "postDate": "2022-06-01T19:25:21.143Z",
          "content": "<p>My current model is using the dataset I have provided in post. At this point it is similar to the model based on original dataset - but I haven't experimented with something specific about what integers can provide.</p>",
          "rawMarkdown": "My current model is using the dataset I have provided in post. At this point it is similar to the model based on original dataset - but I haven't experimented with something specific about what integers can provide.",
          "votes": 2
        },
        {
          "id": 1808412,
          "postDate": "2022-06-01T19:29:47.417Z",
          "content": "<p>Btw, float16 would not have allowed me to capture the integer types well - float32 was necessary. But after all int types have been discovered, float 16 is ok I think</p>",
          "rawMarkdown": "Btw, float16 would not have allowed me to capture the integer types well - float32 was necessary. But after all int types have been discovered, float 16 is ok I think",
          "votes": 2
        },
        {
          "id": 1808437,
          "postDate": "2022-06-01T20:02:52.660Z",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> Thank you for sharing! Just curious, if you had looked at float16, what would have got in the way of capturing the integers? </p>",
          "rawMarkdown": "@raddar Thank you for sharing! Just curious, if you had looked at float16, what would have got in the way of capturing the integers? ",
          "votes": 1
        },
        {
          "id": 1808472,
          "postDate": "2022-06-01T21:10:27.200Z",
          "content": "<p>float16 values got rounded too much. although not impossible, but it was harder to work with</p>",
          "rawMarkdown": "float16 values got rounded too much. although not impossible, but it was harder to work with",
          "votes": 2
        },
        {
          "id": 1882724,
          "postDate": "2022-08-03T11:51:38.633Z",
          "content": "<p>I have been thinking the same. Were you able to obtain any insights?</p>",
          "rawMarkdown": "I have been thinking the same. Were you able to obtain any insights?"
        }
      ]
    },
    {
      "id": 2872889,
      "postDate": "2024-06-15T07:00:22.193Z",
      "content": "<p>Hello!<br>\nThis is my first comment.<br>\nI need it to try find a team(on this you have inspired me, MANY THX!!)<br>\nAnd also thank u million times for showing me the way to work with less RAM))</p>",
      "rawMarkdown": "Hello!\nThis is my first comment.\nI need it to try find a team(on this you have inspired me, MANY THX!!)\nAnd also thank u million times for showing me the way to work with less RAM))"
    },
    {
      "id": 2854030,
      "postDate": "2024-06-04T05:08:28.547Z",
      "content": "<p>kudo for u</p>",
      "rawMarkdown": "kudo for u"
    },
    {
      "id": 2733293,
      "postDate": "2024-04-03T15:14:23.447Z",
      "content": "<p>You just saved for all of us a lot of hours finding the hidden denominator for every single variable and preprocess it for us in a super clean light dataset.</p>\n<p>Thanks a lot <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> !</p>",
      "rawMarkdown": "You just saved for all of us a lot of hours finding the hidden denominator for every single variable and preprocess it for us in a super clean light dataset.\n\nThanks a lot @raddar !\n\n\n"
    },
    {
      "id": 2018821,
      "postDate": "2022-11-06T03:21:09.777Z",
      "content": "<p>Oh wow! That's great! How did you figure this out? I'm a newb and want to know how you were able to learn there was random uniform noise of [0,0.01] added to each column.</p>",
      "rawMarkdown": "Oh wow! That's great! How did you figure this out? I'm a newb and want to know how you were able to learn there was random uniform noise of [0,0.01] added to each column."
    },
    {
      "id": 1941008,
      "postDate": "2022-09-15T17:44:15.093Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!"
    },
    {
      "id": 1919178,
      "postDate": "2022-08-30T07:31:59.003Z",
      "content": "<p>thanks a lot. saved me so much time. up vote.</p>",
      "rawMarkdown": "thanks a lot. saved me so much time. up vote."
    },
    {
      "id": 1889437,
      "postDate": "2022-08-08T07:28:48.723Z",
      "content": "<p>Thanks a lot. It is always good to learn something :)</p>",
      "rawMarkdown": "Thanks a lot. It is always good to learn something :)"
    },
    {
      "id": 1860813,
      "postDate": "2022-07-18T15:26:31.030Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> just wanted to know, how you were able to deal with columns like D_87 which has mostly null values. Do you think converting the D_87 columns to an integer (1 == values exist and 0 == null), would be beneficial for GBDT models? My assumption is that if a value is present in a feature with mostly null columns, the customer is more likely to default. Is it an ok assumption to make, or no?</p>",
      "rawMarkdown": "Hi @raddar just wanted to know, how you were able to deal with columns like D_87 which has mostly null values. Do you think converting the D_87 columns to an integer (1 == values exist and 0 == null), would be beneficial for GBDT models? My assumption is that if a value is present in a feature with mostly null columns, the customer is more likely to default. Is it an ok assumption to make, or no?",
      "replies": [
        {
          "id": 1860854,
          "postDate": "2022-07-18T16:05:14.500Z",
          "content": "<p>Rule of thumb - treat null just as any other number. That's what GBDT models do by default.</p>",
          "rawMarkdown": "Rule of thumb - treat null just as any other number. That's what GBDT models do by default.",
          "votes": 1
        },
        {
          "id": 1861950,
          "postDate": "2022-07-19T11:08:17.867Z",
          "content": "<p>Thanks, thought would be interesting to try</p>",
          "rawMarkdown": "Thanks, thought would be interesting to try"
        }
      ]
    },
    {
      "id": 1859946,
      "postDate": "2022-07-18T05:08:42.750Z",
      "content": "<p>great-grand master</p>",
      "rawMarkdown": "great-grand master"
    },
    {
      "id": 1854531,
      "postDate": "2022-07-13T19:02:46.567Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> : do either of you happen to know if there's an efficient way to do the below (-1 -&gt; nan) for cudf? It complains about data type, for cudf only.</p>\n<p>If not, I guess the workaround is to load the data in pandas, reverse back to na using snippet below, and <em>then</em> call cudf.from_pandas()</p>\n<p>raddar: \"I wanted to make -1 to consistently represent NA. D_59 originally had some negative values. This should have no effect to models. This helps to reverse -1 to NA by just simply x[x==-1] = np.nan.\"</p>\n<p>I suspect it could be a big improvement to get mean, std, etc with nan as nan, and thus excluded from the algorithm. Of course probably even better if first handling some nan's as per raddar's cluster analysis. But even the quick 1 line change might see improvement and is worth testing.</p>",
      "rawMarkdown": "Hi @raddar and @cdeotte : do either of you happen to know if there's an efficient way to do the below (-1 -> nan) for cudf? It complains about data type, for cudf only.\n\nIf not, I guess the workaround is to load the data in pandas, reverse back to na using snippet below, and *then* call cudf.from_pandas()\n\nraddar: \"I wanted to make -1 to consistently represent NA. D_59 originally had some negative values. This should have no effect to models. This helps to reverse -1 to NA by just simply x[x==-1] = np.nan.\"\n\nI suspect it could be a big improvement to get mean, std, etc with nan as nan, and thus excluded from the algorithm. Of course probably even better if first handling some nan's as per raddar's cluster analysis. But even the quick 1 line change might see improvement and is worth testing.",
      "replies": [
        {
          "id": 1854579,
          "postDate": "2022-07-13T20:08:30.667Z",
          "content": "<p>I don't use cudf - useless for this competition.</p>\n<p>as for how to assign NA is actually the problem itself. I am using <code>np.nanmean</code> and other similar numpy aggregate funcs </p>",
          "rawMarkdown": "I don't use cudf - useless for this competition.\n\nas for how to assign NA is actually the problem itself. I am using `np.nanmean` and other similar numpy aggregate funcs ",
          "votes": 1
        },
        {
          "id": 1854640,
          "postDate": "2022-07-13T21:26:07.823Z",
          "content": "<p>The following code will convert all <code>-1</code> to <code>NA</code> in cudf</p>\n<pre><code>for c in df.columns:\n    df.loc[df[c]==-1,c] = cudf.NA\n</code></pre>",
          "rawMarkdown": "The following code will convert all `-1` to `NA` in cudf\n\n    for c in df.columns:\n        df.loc[df[c]==-1,c] = cudf.NA",
          "votes": 4
        },
        {
          "id": 1854643,
          "postDate": "2022-07-13T21:35:33.697Z",
          "content": "<p>Awesome, thanks!</p>",
          "rawMarkdown": "Awesome, thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1854519,
      "postDate": "2022-07-13T18:55:05.323Z",
      "content": "<p>Great detective work! 👍</p>",
      "rawMarkdown": "Great detective work! 👍"
    },
    {
      "id": 1853093,
      "postDate": "2022-07-12T15:38:12.910Z",
      "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>: Thanks for this. I am also wondering why you prefer parquet format over pickle. I was reading <a href=\"https://medium.com/@u.praneel.nihar/improving-read-write-store-performance-by-changing-file-formats-serialization-protocols-bfdb13114004\" target=\"_blank\">this blog</a> which compares parquet to pickle to CSV and concludes pickle is the best of all format. Am I missing anything? </p>",
      "rawMarkdown": "@raddar: Thanks for this. I am also wondering why you prefer parquet format over pickle. I was reading [this blog](https://medium.com/@u.praneel.nihar/improving-read-write-store-performance-by-changing-file-formats-serialization-protocols-bfdb13114004) which compares parquet to pickle to CSV and concludes pickle is the best of all format. Am I missing anything? ",
      "replies": [
        {
          "id": 1853250,
          "postDate": "2022-07-12T18:03:18.277Z",
          "content": "<p>your link somehow forgets to mention how pickle performs in storage;) pickle may be fast, but not necessarily efficient in size</p>",
          "rawMarkdown": "your link somehow forgets to mention how pickle performs in storage;) pickle may be fast, but not necessarily efficient in size",
          "votes": 2
        }
      ]
    },
    {
      "id": 1849141,
      "postDate": "2022-07-09T08:49:23.137Z",
      "content": "<p>5GB -&gt; 50GB with noise.<br>\n\"Default Detection Comp\" with default dataset , lol 😄</p>",
      "rawMarkdown": "5GB -> 50GB with noise.\n\"Default Detection Comp\" with default dataset , lol 😄"
    },
    {
      "id": 1849131,
      "postDate": "2022-07-09T08:43:56.600Z",
      "content": "<p>subscribe + upvote !</p>\n<p>start this competition from here 🎉</p>",
      "rawMarkdown": "subscribe + upvote !\n\nstart this competition from here 🎉"
    },
    {
      "id": 1840438,
      "postDate": "2022-07-02T08:36:17.363Z",
      "content": "<p>Great work! Thx!!</p>",
      "rawMarkdown": "Great work! Thx!!"
    },
    {
      "id": 1839437,
      "postDate": "2022-07-01T12:16:10.043Z",
      "content": "<p>Thanks for posting this. I would be interested in seeing how you broke down the data work into sections small enough to work with in a notebook. </p>",
      "rawMarkdown": "Thanks for posting this. I would be interested in seeing how you broke down the data work into sections small enough to work with in a notebook. "
    },
    {
      "id": 1830673,
      "postDate": "2022-06-23T15:17:02.050Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">raddar</a>! I think that another 4 variables could be squeezed into the <strong>int</strong> type as well: **D104,D_112,B_8 and B_18 ** (look at the histograms of Version 16 of <a href=\"https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex\" target=\"_blank\">this notebook</a>.</p>\n<p>B_18 seems like a counting variable (Versions 14 and 15 of the same notebook above), while D_104,D_112 and B_8 seems like binary variables.</p>\n<p>For D_104 and D_112,the noise seems to be random, although not in an uniform or normal way (distribution). For B_8, the noise seems to be pretty uniform around 1.</p>",
      "rawMarkdown": "Hello [raddar](https://www.kaggle.com/raddar)! I think that another 4 variables could be squeezed into the **int** type as well: **D104,D_112,B_8 and B_18 ** (look at the histograms of Version 16 of [this notebook](https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex).\n\nB_18 seems like a counting variable (Versions 14 and 15 of the same notebook above), while D_104,D_112 and B_8 seems like binary variables.\n\nFor D_104 and D_112,the noise seems to be random, although not in an uniform or normal way (distribution). For B_8, the noise seems to be pretty uniform around 1.",
      "replies": [
        {
          "id": 1830710,
          "postDate": "2022-06-23T16:06:13.733Z",
          "content": "<p>Nice analysis. However, all your listed variables are not 100% certainly integer types. For example D_104 over 0.9 shows \"normal\" like distribution, which indicates that the values are definetly not integers. same goes for other variables - there are ranges where non-uniform distribution can be seen.</p>\n<p>B_18 is most likely integer type - however for me it was impossible to reverse engineer it.</p>",
          "rawMarkdown": "Nice analysis. However, all your listed variables are not 100% certainly integer types. For example D_104 over 0.9 shows \"normal\" like distribution, which indicates that the values are definetly not integers. same goes for other variables - there are ranges where non-uniform distribution can be seen.\n\nB_18 is most likely integer type - however for me it was impossible to reverse engineer it.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1830195,
      "postDate": "2022-06-23T08:53:01.207Z",
      "content": "<p>waw very ingenious ! You are aptly nicknamed Mr \"raddar\"🔥</p>",
      "rawMarkdown": "waw very ingenious ! You are aptly nicknamed Mr \"raddar\"🔥"
    },
    {
      "id": 1829848,
      "postDate": "2022-06-23T01:56:30.640Z",
      "content": "<p>1.7 GB , very smooth!!!  cheers, duh</p>",
      "rawMarkdown": "1.7 GB , very smooth!!!  cheers, duh"
    },
    {
      "id": 1825340,
      "postDate": "2022-06-19T09:02:39.023Z",
      "content": "<p>I have seen that just a moment ago and thanks for your helping hand <a href=\"https://www.kaggle.com/naiborhujosua\" target=\"_blank\">@naiborhujosua</a> 👍</p>",
      "rawMarkdown": "I have seen that just a moment ago and thanks for your helping hand @naiborhujosua 👍"
    },
    {
      "id": 1817024,
      "postDate": "2022-06-10T19:10:05.070Z",
      "content": "<p>Impressive, great work thannks.</p>",
      "rawMarkdown": "Impressive, great work thannks."
    },
    {
      "id": 1815587,
      "postDate": "2022-06-09T07:30:44.857Z",
      "content": "<p>is the target variable provided in that American express comepetion<br>\nwhat is the target variable</p>",
      "rawMarkdown": "is the target variable provided in that American express comepetion\nwhat is the target variable",
      "replies": [
        {
          "id": 1815603,
          "postDate": "2022-06-09T07:54:10.817Z",
          "content": "<p>There are three files in the original dataset: train, test and labels. All in .csv format.</p>",
          "rawMarkdown": "There are three files in the original dataset: train, test and labels. All in .csv format.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1814915,
      "postDate": "2022-06-08T13:20:22.703Z",
      "content": "<p>This is a really useful \"compression\" for using this data for fun and learning! Thanks!</p>",
      "rawMarkdown": "This is a really useful \"compression\" for using this data for fun and learning! Thanks!"
    },
    {
      "id": 1814379,
      "postDate": "2022-06-07T20:06:02.633Z",
      "content": "<p>Thanks a lot for this dataset! Great job.</p>",
      "rawMarkdown": "Thanks a lot for this dataset! Great job."
    },
    {
      "id": 1812062,
      "postDate": "2022-06-05T12:55:19.677Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. I was facing memory issue and this helps a lot</p>",
      "rawMarkdown": "Thanks @raddar. I was facing memory issue and this helps a lot"
    },
    {
      "id": 1811897,
      "postDate": "2022-06-05T09:14:49.390Z",
      "content": "<p><code>x['D_59'] = floorify_frac(x['D_59']+5/48,1/48)</code></p>\n<p>Another question. Did you add 5 / 48 in order to shift minimum value to 0 so floor operation works consistently?</p>",
      "rawMarkdown": "`x['D_59'] = floorify_frac(x['D_59']+5/48,1/48)`\n\nAnother question. Did you add 5 / 48 in order to shift minimum value to 0 so floor operation works consistently?",
      "replies": [
        {
          "id": 1812010,
          "postDate": "2022-06-05T11:43:11.320Z",
          "content": "<p>Cannot remember really but probably I wanted to make -1 consistently represent missing value</p>",
          "rawMarkdown": "Cannot remember really but probably I wanted to make -1 consistently represent missing value",
          "votes": 3
        }
      ]
    },
    {
      "id": 1811645,
      "postDate": "2022-06-05T00:08:02.823Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>. The test data is actually two separate datasets that can be distinguished based on max date. Would you consider separating public and private LB in two separate parquet ?</p>",
      "rawMarkdown": "Hey @raddar. The test data is actually two separate datasets that can be distinguished based on max date. Would you consider separating public and private LB in two separate parquet ?"
    },
    {
      "id": 1811085,
      "postDate": "2022-06-04T09:51:09.180Z",
      "content": "<p>thx! I'm going to use this right now!</p>",
      "rawMarkdown": "thx! I'm going to use this right now!"
    },
    {
      "id": 1810941,
      "postDate": "2022-06-04T05:06:54.283Z",
      "content": "<p>Fantastic detective work! 🔎</p>",
      "rawMarkdown": "Fantastic detective work! 🔎"
    },
    {
      "id": 2105294,
      "postDate": "2023-01-18T12:09:58.020Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2107149,
          "postDate": "2023-01-19T15:49:40.680Z",
          "content": "<p>Thank you. However, I doubt it would be useful in real world. It's a very specific competition problem. </p>",
          "rawMarkdown": "Thank you. However, I doubt it would be useful in real world. It's a very specific competition problem. "
        }
      ]
    },
    {
      "id": 1872835,
      "postDate": "2022-07-27T09:09:07.333Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1838152,
      "postDate": "2022-06-30T09:01:33.650Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1838171,
          "postDate": "2022-06-30T09:34:40.830Z",
          "content": "<p>Look at this: <a href=\"https://www.kaggle.com/code/raddar/target-true-meaning-revealed\" target=\"_blank\">https://www.kaggle.com/code/raddar/target-true-meaning-revealed</a></p>\n<p>Very similar problem, but a bit different approach</p>",
          "rawMarkdown": "Look at this: https://www.kaggle.com/code/raddar/target-true-meaning-revealed\n\nVery similar problem, but a bit different approach",
          "votes": 1
        }
      ]
    },
    {
      "id": 1819755,
      "postDate": "2022-06-14T05:41:18.377Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1819777,
          "postDate": "2022-06-14T06:00:05.143Z",
          "content": "<p>This information was not provided. it's rather an insight just by looking at the data - it is obvious.</p>",
          "rawMarkdown": "This information was not provided. it's rather an insight just by looking at the data - it is obvious.",
          "votes": 3
        },
        {
          "id": 1819779,
          "postDate": "2022-06-14T06:05:15.710Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1812681,
      "postDate": "2022-06-06T05:32:22.200Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 1812005,
      "postDate": "2022-06-05T11:39:21.120Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1812019,
          "postDate": "2022-06-05T11:51:21.087Z",
          "content": "<ol>\n<li><p><code>D_63</code> mapping order does not really matter as it is categorical anyway - tree models are good at it. and if you were to build neural networks, you would do this kind of sorting probably anyway.</p></li>\n<li><p>As for <code>S_13</code> I think I did not apply <code>floor_frac</code> as <code>1/min_step</code> is larger than the actual step, which caused my function to not work properly. Let's say you have [1,1.01] and min_step = 3000. if you apply <code>1/min_step</code> to this, you end up with [3000,3030] - which is not a single integer. my function only works if <code>min_step&lt;100</code>. You can test it yorself:)</p></li>\n</ol>",
          "rawMarkdown": "1. `D_63` mapping order does not really matter as it is categorical anyway - tree models are good at it. and if you were to build neural networks, you would do this kind of sorting probably anyway.\n\n2. As for `S_13` I think I did not apply `floor_frac` as `1/min_step` is larger than the actual step, which caused my function to not work properly. Let's say you have [1,1.01] and min_step = 3000. if you apply `1/min_step` to this, you end up with [3000,3030] - which is not a single integer. my function only works if `min_step<100`. You can test it yorself:)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1811123,
      "postDate": "2022-06-04T10:37:12.123Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1905666,
      "postDate": "2022-08-19T07:58:39.940Z",
      "content": "<p>Thanks a lot. very useful.</p>",
      "rawMarkdown": "Thanks a lot. very useful."
    },
    {
      "id": 1901372,
      "postDate": "2022-08-16T15:54:45.320Z",
      "content": "<p>Thanks a lot, great</p>",
      "rawMarkdown": "Thanks a lot, great"
    },
    {
      "id": 1897656,
      "postDate": "2022-08-14T01:30:27.537Z",
      "content": "<p>Thanks a lot. well done</p>",
      "rawMarkdown": "Thanks a lot. well done"
    },
    {
      "id": 1896040,
      "postDate": "2022-08-12T14:39:55.263Z",
      "content": "<p>Thank you for your information!</p>",
      "rawMarkdown": "Thank you for your information!"
    },
    {
      "id": 1880768,
      "postDate": "2022-08-02T03:35:20.087Z",
      "content": "<p>Thanks a lot,great</p>",
      "rawMarkdown": "Thanks a lot,great"
    },
    {
      "id": 1872383,
      "postDate": "2022-07-26T22:18:26.210Z",
      "content": "<p>Thanks for contribution</p>",
      "rawMarkdown": "Thanks for contribution"
    },
    {
      "id": 1861651,
      "postDate": "2022-07-19T06:31:51.777Z",
      "content": "<p>THANKS FOR SHARING！</p>",
      "rawMarkdown": "THANKS FOR SHARING！"
    },
    {
      "id": 1860727,
      "postDate": "2022-07-18T13:53:20.143Z",
      "content": "<p>Thanks for sharing your great work!</p>",
      "rawMarkdown": "Thanks for sharing your great work!"
    },
    {
      "id": 1853103,
      "postDate": "2022-07-12T15:47:02.630Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!"
    },
    {
      "id": 1850949,
      "postDate": "2022-07-11T00:18:12.197Z",
      "content": "<p>Thanks a lot !!!</p>",
      "rawMarkdown": "Thanks a lot !!!"
    },
    {
      "id": 1840436,
      "postDate": "2022-07-02T08:35:57.707Z",
      "content": "<p>Thanks a lot.</p>",
      "rawMarkdown": "Thanks a lot."
    },
    {
      "id": 1834578,
      "postDate": "2022-06-27T04:12:21.860Z",
      "content": "<p>Thanks for sharing greate techniques :)</p>",
      "rawMarkdown": "Thanks for sharing greate techniques :)"
    },
    {
      "id": 1824177,
      "postDate": "2022-06-18T04:02:33.610Z",
      "content": "<p>Just awesome! Thanks a lot!</p>",
      "rawMarkdown": "Just awesome! Thanks a lot!"
    },
    {
      "id": 1817390,
      "postDate": "2022-06-11T08:54:53.980Z",
      "content": "<p>Thanks a lot. It as really a good read.</p>",
      "rawMarkdown": "Thanks a lot. It as really a good read."
    },
    {
      "id": 1816171,
      "postDate": "2022-06-09T22:35:12.590Z",
      "content": "<p>Great content!! Thanks</p>",
      "rawMarkdown": "Great content!! Thanks"
    },
    {
      "id": 1813505,
      "postDate": "2022-06-06T22:52:23.530Z",
      "content": "<p>Thanks for the teaching</p>",
      "rawMarkdown": "Thanks for the teaching"
    },
    {
      "id": 1812145,
      "postDate": "2022-06-05T14:23:10.013Z",
      "content": "<p>Thanks for contribution</p>",
      "rawMarkdown": "Thanks for contribution"
    },
    {
      "id": 1809658,
      "postDate": "2022-06-02T23:25:19.820Z",
      "content": "<p>Thanks for contribution</p>",
      "rawMarkdown": "Thanks for contribution"
    },
    {
      "id": 1809434,
      "postDate": "2022-06-02T17:05:43.877Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for these contributions. </p>",
      "rawMarkdown": "Thanks @raddar for these contributions. "
    }
  ],
  "comments": [
    {
      "id": 1826150,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-06-20T04:57:04.827000",
      "content": "<p><img src=\"https://i.ibb.co/6Fyt3Mk/Selection-905.png\" alt=\"\"></p>",
      "votes": 68,
      "replies": []
    },
    {
      "id": 1808521,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-06-01T23:49:45.287000",
      "content": "<p>I published an XGB starter notebook using your data with CV 0.792 and LB 0.794 <a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">here</a>. This data is great. It is so small, that both the feature engineering and model training can easily occur within one Kaggle notebook. The total notebook time (reading data, feature engineering, training 5 folds, inferring test) is only 13 minutes! Everything is done on GPU!</p>",
      "votes": 44,
      "replies": []
    },
    {
      "id": 1808256,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-06-01T16:54:35.137000",
      "content": "<p>Fantastic job raddar. I was just doing this myself today. When we zoom in on some variables, we see that without noise they are actually discrete values. </p>\n<p>For example <code>R_26</code>. Below the histogram is displayed with noise. The variable <code>R_26</code> has <code>548160</code> unique values below 0.2 and is provided as <code>float64</code> (8 bytes), but really from the plot we see it is only 6 discrete values with (random uniform) noise added and can be converted to <code>int8</code> (1 byte) - without information loss. </p>\n<p>It's funny how the data is actually 5GB with 45GB of noise added lol 😄</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/r_26.png\" alt=\"\"></p>",
      "votes": 41,
      "replies": [
        {
          "id": 1808273,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-01T17:12:43.163000",
          "content": "<p>makes you question why they added the noise in the first place. It was obvious this would happen :)</p>",
          "votes": 19,
          "replies": []
        },
        {
          "id": 1808343,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-06-01T18:11:49.853000",
          "content": "<p>My best  theory is that they had nans in their integer columns (using another langage maybe) and python didn't allow them to keep those columns as integers. There still is a question about adding noise to floats.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1816727,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2022-06-10T13:56:12.303000",
          "content": "<p>In association football, the extra time is usually in whole minutes (human ref), but it seldom lasts exactly on the minutes (human ref), so there are narrow distributions around the whole minutes. But, should they be rounded without loss of information?</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2936737,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-07-26T12:22:26.717000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1808994,
      "author_name": "SRK",
      "author_url": "",
      "post_date": "2022-06-02T09:57:16.710000",
      "content": "<p>raddar - You are our Sherlock Holmes 🕵️</p>",
      "votes": 24,
      "replies": []
    },
    {
      "id": 1810193,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-06-03T10:36:45.123000",
      "content": "<p>I just made an update to the dataset, with 3 new features converted to int: <code>B_19</code>, <code>S_8</code> and <code>S_13</code>. notebooks updated as well :)</p>",
      "votes": 12,
      "replies": [
        {
          "id": 1810446,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-03T15:11:38.037000",
          "content": "<p>Great work, i need to update my XGB notebook version. After each update, it would be interesting to take note of the XGB CV score. I wonder if the CV LB will keep improving as you remove noise from more columns.</p>\n<p>FYI for everyone. If you want an older version of Raddar's dataset, change the URL below to get whichever version that you want</p>\n<p><code>https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format/versions/2</code></p>\n<p>The number at the end is the version number. Currently Raddar has versions 1 and 2.</p>",
          "votes": 11,
          "replies": []
        }
      ]
    },
    {
      "id": 1812195,
      "author_name": "Fernando Delgado",
      "author_url": "",
      "post_date": "2022-06-05T15:14:20.360000",
      "content": "<p>Hey there! Thank a lot <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> . </p>\n<p>I aggregated your parquet datasets with <a href=\"https://www.kaggle.com/huseyincot\" target=\"_blank\">@huseyincot</a> 's method! In case its useful for anyone, it improved my score a little bit.</p>\n<p>You can find the aggregated datasets <a href=\"https://www.kaggle.com/datasets/heyspaceturtle/agg-amex-data-without-noise\" target=\"_blank\">here</a><br>\nAnd my notebook <a href=\"https://www.kaggle.com/code/heyspaceturtle/data-aggregation?scriptVersionId=97546977\" target=\"_blank\">here</a></p>\n<p>good vibes! </p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1837347,
      "author_name": "Fatima HABIB",
      "author_url": "",
      "post_date": "2022-06-29T14:00:27.293000",
      "content": "<p>Thanks a lot, I still do not understand the meaning of \"random uniform noise\"?<br>\n please if anyone has a notebook that illustrates this idea, I would be glad if you share it with me.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1837480,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-29T16:18:05.053000",
          "content": "<p><a href=\"https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added\" target=\"_blank\">https://www.kaggle.com/code/raddar/the-data-has-random-uniform-noise-added</a></p>",
          "votes": 13,
          "replies": []
        }
      ]
    },
    {
      "id": 1847129,
      "author_name": "Felipe Loque",
      "author_url": "",
      "post_date": "2022-07-07T16:46:18.600000",
      "content": "<p>RADDAR: Read All Data, Denoise And Release</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1810347,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-06-03T13:35:50.690000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, I'm still trying to figure out how did you set those intervals for floorify_frac?</p>\n<p>Let's say we have this feature B_4 histogram.</p>\n<p><img src=\"https://i.ibb.co/MSmCp6J/Screenshot-from-2022-06-03-16-13-15.png\" alt=\"1\"></p>\n<p>When we zoom in, we see those blocks.</p>\n<p><img src=\"https://i.ibb.co/znBSkst/Screenshot-from-2022-06-03-16-14-15.png\" alt=\"2\"></p>\n<p>When we zoom even more, we can see that +-0.01 random uniform noise added to them.</p>\n<p><img src=\"https://i.ibb.co/7KnyY8j/Screenshot-from-2022-06-03-16-21-08.png\" alt=\"3\"></p>\n<p>As far as I understand, they scaled raw features with a coefficient so that the injected noise won't be insignificant and won't cause any information loss at the same time.</p>\n<p>I guess any scaling factor that can project those blocks into the same integer value can be used here. 1 / 78 doesn't necessarily have to be the only valid value. I was wondering how did you empirically come up with those values?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1810434,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-03T15:00:41.280000",
          "content": "<p>I did the very same thing you described - I zoomed in into very small interval space to identify these bins and then worked on that.</p>\n<p>We know that each bin is exactly 0.01 wide (due to uniform [0,0.01] random). This means each bin is represented by its minimum value - which is its original value before random noise. Then let's take 2 consecutive bins and look at the difference - this represents one or more steps between bins. Find the smallest difference (lowest step) - which is indicative to the increment by 1 in the integer space. then simpy 1/X of that value represents the common denominator..</p>\n<p>Other way to think is - we do simple optimization, like a pseudo code:</p>\n<pre><code>for i in range(N):\n  if x['B_8'] * i ~ integer for all values:\n    return i\n</code></pre>\n<p>So simply scanning for smallest value <code>i</code> to have all values as smallest possible integers.</p>\n<p>Of course 1/78 is not the only value. it is as valid as 1/(2 x 78), 1/(3 x 78)… however 2x, 3x multiplier would just mean that integers would be multiplied by 2x, 3x…</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1810460,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-06-03T15:26:00.100000",
          "content": "<p>Thanks for the explanation. I thought the random uniform noise was between -0.01 and 0.01. Now it makes sense why you pick bins' minimum values.</p>\n<p>I checked steps of multiple bins in B_4 and I found that the step size is 0.013. My multiplier is slightly different than yours. Do you think the step size might not be fixed for some bins? I'll try to calculate those step sizes programmatically to see how many unique values are there.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1810483,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-03T15:40:49.917000",
          "content": "<p>I guess you had <code>B_4</code> in mind?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1810489,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-06-03T15:42:50.290000",
          "content": "<p>Yes, I meant B_4. I made a typo.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1810527,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-03T16:16:10.777000",
          "content": "<p>Maybe you are working with float16? they do lose a lot of information.</p>\n<pre><code>z = pd.read_csv('train_data.csv',usecols=['B_4'])\nv = z.loc[(z.B_4&gt;0.615)].min() - z.loc[(z.B_4&gt;0.6)].min()\n1/v\n\n&gt; B_4    78.002257\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1810553,
          "author_name": "Gunes Evitan",
          "author_url": "",
          "post_date": "2022-06-03T16:53:58.467000",
          "content": "<p>I tried with 32 and 64 bits but there wasn't any difference. I think the difference arises from the histogram function. Check the image below. I was using 0.6145 as the bin edge and step size was different for those particular bins. I think your calculation is correct. I have find the optimal value by searching the dataframe, not the visualization.</p>\n<p><img src=\"https://i.ibb.co/Gn59cwn/Screenshot-from-2022-06-03-19-49-13.png\" alt=\"1\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1839869,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-01T18:49:05.760000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1837124,
      "author_name": "SchopenHacker75",
      "author_url": "",
      "post_date": "2022-06-29T11:06:04.197000",
      "content": "<p>hello <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, very brilliant reverse engineering (\"DeAnonymization\")</p>\n<p>However there are some ambiguous passages for me:</p>\n<ol>\n<li>what do you mean by \"<strong>each bin has AUC of 0.5</strong>\" for the the <code>B_19</code> feature ?  Do you have a related sample code ? thx..</li>\n<li><code>x['D_59'] = floorify_frac(x['D_59']+5/48,1/48)</code> why did you have to add some steps before reconstructing the feature (same for <code>D_124</code>)</li>\n<li>I did'nt understand the cross referencing with <code>S_13</code> of <code>S_11</code>  : because I didn't find any data in : <code>train.S_11.isin([15,16,17])</code> </li>\n</ol>\n<p>Thx for considering my questions😊</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1837493,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-29T16:22:10.497000",
          "content": "<ol>\n<li><p>I do not have code. The idea is: take a bin with values in range [0,0.01]. calculate AUC of that range. Then calculate AUC for range [0.01,0.02], etc. smth like <code>roc_auc_score(x.loc[(x.B_19&gt;=0.01) &amp; (x.B_19&lt;=0.02),'target'], x.loc[(x.B_19&gt;=0.01) &amp; (x.B_19&lt;=0.02),'B_19'])</code>. </p></li>\n<li><p>I wanted to make -1 to consistently represent NA. D_59 originally had some negative values. This should have no effect to models. This helps to reverse -1 to NA by just simply <code>x[x==-1] = np.nan</code></p></li>\n<li><p>S_13 and S_11 were already converted to integers, so <code>train.S_11.isin([15,16,17])</code> exists.</p></li>\n</ol>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1838441,
          "author_name": "SchopenHacker75",
          "author_url": "",
          "post_date": "2022-06-30T14:20:58.683000",
          "content": "<p>Manyyy thanks it is pretty clear 🙏😊</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1840845,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-02T15:28:14.727000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1840859,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-02T15:38:31.673000",
          "content": "<p>If there was a signal (as in increasing/decreasing target rate) within each bin ([0,0.01], [0.01,0.02], etc.), then AUC!=0.5. </p>\n<p>So by testing for AUC=0.5 we test the hypothesis that there is no target signal in each bin. </p>\n<p>If there is no signal - there is no mix of many original values in the bin, meaning that the bin is represented by only one number - the min value of the bin.</p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 1840880,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-02T15:51:56.250000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1808471,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2022-06-01T21:09:28.613000",
      "content": "<p>HI <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> It seems that <code>B_19</code> would profit from <code>floorify_frac(x['B_19'])</code> as well. And the original values of <code>S_13</code> are all multiples of 1/1034, but as 1/1034 &lt; 0.01, the added noise cannot be removed (diagrams are at the end of <a href=\"https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense\" target=\"_blank\">my EDA</a>).</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1808482,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-01T21:29:20.917000",
          "content": "<p>They are both inconclusive to me as intervals may overlap, so I left them as is.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1809184,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-02T13:31:09.013000",
          "content": "<p>I revisited your both var suggestions. I was able to convert them to int:)</p>\n<p><code>B_19</code> can be easily rounded every 0.01 (if you check AUC scores within each bin, AUC is ~0.5, so no target variance within each bin)</p>\n<p>as for S_13 I used another trick, cross referencing S_11</p>\n<pre><code>def floorify(x, lo):\n    \"\"\"example: x in [0, 0.01] -&gt; x := 0\"\"\"\n    return lo if x &lt;= lo+0.01 and x &gt;= lo else x\n\n### S_13 has these weird ordinal values\n# one value overlaps, but can be split by S_11\nx.loc[(x.S_13&gt;=0.67) &amp; (x.S_13&lt;=0.7) &amp; (x.S_11.isin([15,16,17])),'S_13'] = 0.67891681\nx.loc[(x.S_13&gt;=0.67) &amp; (x.S_13&lt;=0.7) &amp; ~(x.S_11.isin([15,16,17])),'S_13'] = 0.68762086\n\nfor c in (0.03771764, 0.28046423, 0.40135398, 0.42069634, 0.50676983, 0.52611219, 0.55512583, 0.62185686, 0.84332692):\n    x['S_13'] = x['S_13'].apply(lambda t: floorify(t,c))\n\nx['S_13'] = np.round(x['S_13']*1034)\n</code></pre>\n<p>so 2 more int features in the stack :) </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1809208,
          "author_name": "Tonghui Li",
          "author_url": "",
          "post_date": "2022-06-02T14:02:41.170000",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> What you found out about S_13 is impressive! How did I you figure out S_13 can be split using S_11? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1809242,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-02T14:41:08.870000",
          "content": "<p>Simply slicing that S_13 range, sorting, and opening that in excel. obvious pattern emerged</p>",
          "votes": 12,
          "replies": []
        }
      ]
    },
    {
      "id": 1808410,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2022-06-01T19:28:28.427000",
      "content": "<p>There are also many small tweaks that can be made on this dataset, for example:</p>\n<pre><code>x.loc[(x.R_13==0) &amp; (x.R_17==0) &amp; (x.R_20==0) &amp; (x.R_8==0), 'R_6'] = 0\nx.loc[x.B_39==-1, 'B_36'] = 0\n</code></pre>\n<p>Do not know if it is worth the effort though to look for more</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1847883,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2022-07-08T08:19:06.680000",
      "content": "<p>Thank you for sharing this. It seems that D_128, D_130, B_8, and R_1 are also hidden int features.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1851080,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2022-07-11T02:30:19.823000",
          "content": "<p>Sorry, it was not true. ROCs are not 0.5.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1809629,
      "author_name": "Anubhav Chhabra",
      "author_url": "",
      "post_date": "2022-06-02T21:43:26.367000",
      "content": "<p>Great insight! Even <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> worked on immensely reducing the size of data so that it can fit into our poor RAMs.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1895681,
      "author_name": "sc",
      "author_url": "",
      "post_date": "2022-08-12T10:10:09.643000",
      "content": "<p>Hi Radar. Kudos to you. This is not only good work but also saves us a lot of time for everyone.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1811436,
      "author_name": "Nikola Bacic",
      "author_url": "",
      "post_date": "2022-06-04T17:15:36.567000",
      "content": "<p>Looks like these 93 float32 features hold the majority of the signal. I get CV 0.78 easy by just using them.<br>\nThanks!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1809575,
      "author_name": "huseyincotel",
      "author_url": "",
      "post_date": "2022-06-02T20:11:36.813000",
      "content": "<p>Next step is to know what that features may be? 👀</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1809157,
      "author_name": "Wassim",
      "author_url": "",
      "post_date": "2022-06-02T13:02:51.480000",
      "content": "<p>Wow ! You just saved for all of us a lot of hours finding the hidden denominator for every single variable and preprocess it for us in a super clean light dataset.</p>\n<p>Thanks a lot <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> ! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1808399,
      "author_name": "Ertuğrul Demir",
      "author_url": "",
      "post_date": "2022-06-01T19:17:30.627000",
      "content": "<p>Nice catch <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, I was worried about converting some of these variables into float16 since we were losing precision in the process, but at the same time thinking about models might converge faster without that noise. I was planning to compare both with and w/o noise data models in cross validation, have you tried that yourself, seen any noticeable differences?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1808407,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-01T19:25:21.143000",
          "content": "<p>My current model is using the dataset I have provided in post. At this point it is similar to the model based on original dataset - but I haven't experimented with something specific about what integers can provide.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1808412,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-01T19:29:47.417000",
          "content": "<p>Btw, float16 would not have allowed me to capture the integer types well - float32 was necessary. But after all int types have been discovered, float 16 is ok I think</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1808437,
          "author_name": "Tonghui Li",
          "author_url": "",
          "post_date": "2022-06-01T20:02:52.660000",
          "content": "<p><a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> Thank you for sharing! Just curious, if you had looked at float16, what would have got in the way of capturing the integers? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1808472,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-06-01T21:10:27.200000",
          "content": "<p>float16 values got rounded too much. although not impossible, but it was harder to work with</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1882724,
          "author_name": "Luca Lanzendörfer",
          "author_url": "",
          "post_date": "2022-08-03T11:51:38.633000",
          "content": "<p>I have been thinking the same. Were you able to obtain any insights?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2872889,
      "author_name": "Nikolai Malkamalov",
      "author_url": "",
      "post_date": "2024-06-15T07:00:22.193000",
      "content": "<p>Hello!<br>\nThis is my first comment.<br>\nI need it to try find a team(on this you have inspired me, MANY THX!!)<br>\nAnd also thank u million times for showing me the way to work with less RAM))</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2854030,
      "author_name": "JubilLee",
      "author_url": "",
      "post_date": "2024-06-04T05:08:28.547000",
      "content": "<p>kudo for u</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2733293,
      "author_name": "Aaron Frias",
      "author_url": "",
      "post_date": "2024-04-03T15:14:23.447000",
      "content": "<p>You just saved for all of us a lot of hours finding the hidden denominator for every single variable and preprocess it for us in a super clean light dataset.</p>\n<p>Thanks a lot <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2018821,
      "author_name": "Rukshar Alam",
      "author_url": "",
      "post_date": "2022-11-06T03:21:09.777000",
      "content": "<p>Oh wow! That's great! How did you figure this out? I'm a newb and want to know how you were able to learn there was random uniform noise of [0,0.01] added to each column.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1941008,
      "author_name": "fzefii",
      "author_url": "",
      "post_date": "2022-09-15T17:44:15.093000",
      "content": "<p>Great work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1919178,
      "author_name": "meciwo",
      "author_url": "",
      "post_date": "2022-08-30T07:31:59.003000",
      "content": "<p>thanks a lot. saved me so much time. up vote.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1889437,
      "author_name": "NSDQ",
      "author_url": "",
      "post_date": "2022-08-08T07:28:48.723000",
      "content": "<p>Thanks a lot. It is always good to learn something :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1860813,
      "author_name": "Tarrasque9",
      "author_url": "",
      "post_date": "2022-07-18T15:26:31.030000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> just wanted to know, how you were able to deal with columns like D_87 which has mostly null values. Do you think converting the D_87 columns to an integer (1 == values exist and 0 == null), would be beneficial for GBDT models? My assumption is that if a value is present in a feature with mostly null columns, the customer is more likely to default. Is it an ok assumption to make, or no?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1860854,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-18T16:05:14.500000",
          "content": "<p>Rule of thumb - treat null just as any other number. That's what GBDT models do by default.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1861950,
          "author_name": "Tarrasque9",
          "author_url": "",
          "post_date": "2022-07-19T11:08:17.867000",
          "content": "<p>Thanks, thought would be interesting to try</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1859946,
      "author_name": "yuuniee",
      "author_url": "",
      "post_date": "2022-07-18T05:08:42.750000",
      "content": "<p>great-grand master</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1854531,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-07-13T19:02:46.567000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> : do either of you happen to know if there's an efficient way to do the below (-1 -&gt; nan) for cudf? It complains about data type, for cudf only.</p>\n<p>If not, I guess the workaround is to load the data in pandas, reverse back to na using snippet below, and <em>then</em> call cudf.from_pandas()</p>\n<p>raddar: \"I wanted to make -1 to consistently represent NA. D_59 originally had some negative values. This should have no effect to models. This helps to reverse -1 to NA by just simply x[x==-1] = np.nan.\"</p>\n<p>I suspect it could be a big improvement to get mean, std, etc with nan as nan, and thus excluded from the algorithm. Of course probably even better if first handling some nan's as per raddar's cluster analysis. But even the quick 1 line change might see improvement and is worth testing.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1854579,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2022-07-13T20:08:30.667000",
          "content": "<p>I don't use cudf - useless for this competition.</p>\n<p>as for how to assign NA is actually the problem itself. I am using <code>np.nanmean</code> and other similar numpy aggregate funcs </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1854640,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-13T21:26:07.823000",
          "content": "<p>The following code will convert all <code>-1</code> to <code>NA</code> in cudf</p>\n<pre><code>for c in df.columns:\n    df.loc[df[c]==-1,c] = cudf.NA\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1854643,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-07-13T21:35:33.697000",
          "content": "<p>Awesome, thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1854519,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-13T18:55:05.323000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1853093,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-12T15:38:12.910000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1853250,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-07-12T18:03:18.277000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1849141,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-09T08:49:23.137000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1849131,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-09T08:43:56.600000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1840438,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-02T08:36:17.363000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1839437,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-01T12:16:10.043000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1830673,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-23T15:17:02.050000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1830710,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-23T16:06:13.733000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1830195,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-23T08:53:01.207000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1829848,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-23T01:56:30.640000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1825340,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-19T09:02:39.023000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1817024,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-10T19:10:05.070000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1815587,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-09T07:30:44.857000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1815603,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-09T07:54:10.817000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1814915,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-08T13:20:22.703000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1814379,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-07T20:06:02.633000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1812062,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T12:55:19.677000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1811897,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T09:14:49.390000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1812010,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-05T11:43:11.320000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1811645,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T00:08:02.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1811085,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-04T09:51:09.180000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1810941,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-04T05:06:54.283000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2105294,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-18T12:09:58.020000",
      "content": "",
      "votes": -1,
      "replies": [
        {
          "id": 2107149,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-01-19T15:49:40.680000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1872835,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-27T09:09:07.333000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1838152,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-30T09:01:33.650000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1838171,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-30T09:34:40.830000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1819755,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-14T05:41:18.377000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1819777,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-14T06:00:05.143000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1819779,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-14T06:05:15.710000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1812681,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-06T05:32:22.200000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1812005,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T11:39:21.120000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1812019,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-06-05T11:51:21.087000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1811123,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-04T10:37:12.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1905666,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-19T07:58:39.940000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1901372,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-16T15:54:45.320000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1897656,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-14T01:30:27.537000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1896040,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-12T14:39:55.263000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1880768,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-02T03:35:20.087000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1872383,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-26T22:18:26.210000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1861651,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-19T06:31:51.777000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1860727,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-18T13:53:20.143000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1853103,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-12T15:47:02.630000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1850949,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-11T00:18:12.197000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1840436,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-02T08:35:57.707000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1834578,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-27T04:12:21.860000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1824177,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-18T04:02:33.610000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1817390,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-11T08:54:53.980000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1816171,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-09T22:35:12.590000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1813505,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-06T22:52:23.530000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1812145,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-05T14:23:10.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1809658,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T23:25:19.820000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1809434,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-02T17:05:43.877000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1808215": "As it is already clear - all float type columns have random uniform noise of [0,0.01] added to each column. Having this information it is clear that it is very easy to spot integer type columns in the raw data. This calls for detective work, which I really love and have done many times in previous competitions :) \n\nIt took me a couple of days, but I have it ready for you all to use!\n\nSo, what's in it?\n \n- originally we had 188 float/categorical type features. These were transformed into\n   - 95 np.int8/np.int16 types\n   - 93 np.float32 types\n- Most float columns with [0, 0.01] and [1, 1.01] have these values rounded up at 0 and 1 respectively. This was done to ensure no data loss, as not all features could be rounded up safely.\n- saved in parquet format (only 1.7GB training data!)\n\nHopefully, this opens the door for many people to access this competition, as less RAM is required. \n\nAlso more interesting feature engineering will be available as it is much easier to work with integers (my experience).\n\nThe cleaned dataset can be found as a dataset:\nhttps://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n\nNotebooks used to transform the data:\nhttps://www.kaggle.com/code/raddar/amex-data-int-types-train\nhttps://www.kaggle.com/code/raddar/amex-data-int-types-test\n\n\nAs many youtubers would say - please like and subscribe! This helps me to be motivated to share stuff with you all :) More things to come.\n\nUPDATE: \n\ndataset was updated with 3 extra int conversions",
    "1826150": "![](https://i.ibb.co/6Fyt3Mk/Selection-905.png)",
    "1808521": "I published an XGB starter notebook using your data with CV 0.792 and LB 0.794 [here][1]. This data is great. It is so small, that both the feature engineering and model training can easily occur within one Kaggle notebook. The total notebook time (reading data, feature engineering, training 5 folds, inferring test) is only 13 minutes! Everything is done on GPU!\n\n[1]: https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
    "1808256": "Fantastic job raddar. I was just doing this myself today. When we zoom in on some variables, we see that without noise they are actually discrete values. \n\nFor example `R_26`. Below the histogram is displayed with noise. The variable `R_26` has `548160` unique values below 0.2 and is provided as `float64` (8 bytes), but really from the plot we see it is only 6 discrete values with (random uniform) noise added and can be converted to `int8` (1 byte) - without information loss. \n\nIt's funny how the data is actually 5GB with 45GB of noise added lol 😄\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/r_26.png)",
    "1808994": "raddar - You are our Sherlock Holmes 🕵️",
    "1810193": "I just made an update to the dataset, with 3 new features converted to int: `B_19`, `S_8` and `S_13`. notebooks updated as well :)",
    "1812195": "Hey there! Thank a lot @raddar . \n\nI aggregated your parquet datasets with @huseyincot 's method! In case its useful for anyone, it improved my score a little bit.\n\nYou can find the aggregated datasets [here](https://www.kaggle.com/datasets/heyspaceturtle/agg-amex-data-without-noise)\nAnd my notebook [here](https://www.kaggle.com/code/heyspaceturtle/data-aggregation?scriptVersionId=97546977)\n\ngood vibes! ",
    "1837347": "Thanks a lot, I still do not understand the meaning of \"random uniform noise\"?\n please if anyone has a notebook that illustrates this idea, I would be glad if you share it with me.",
    "1847129": "RADDAR: Read All Data, Denoise And Release",
    "1810347": "Hello @raddar, I'm still trying to figure out how did you set those intervals for floorify_frac?\n\nLet's say we have this feature B_4 histogram.\n\n![1](https://i.ibb.co/MSmCp6J/Screenshot-from-2022-06-03-16-13-15.png)\n\nWhen we zoom in, we see those blocks.\n\n![2](https://i.ibb.co/znBSkst/Screenshot-from-2022-06-03-16-14-15.png)\n\nWhen we zoom even more, we can see that +-0.01 random uniform noise added to them.\n\n![3](https://i.ibb.co/7KnyY8j/Screenshot-from-2022-06-03-16-21-08.png)\n\nAs far as I understand, they scaled raw features with a coefficient so that the injected noise won't be insignificant and won't cause any information loss at the same time.\n\nI guess any scaling factor that can project those blocks into the same integer value can be used here. 1 / 78 doesn't necessarily have to be the only valid value. I was wondering how did you empirically come up with those values?",
    "1837124": "hello @raddar, very brilliant reverse engineering (\"DeAnonymization\")\n\nHowever there are some ambiguous passages for me:\n1. what do you mean by \"**each bin has AUC of 0.5**\" for the the `B_19` feature ?  Do you have a related sample code ? thx..\n1. `x['D_59'] = floorify_frac(x['D_59']+5/48,1/48)` why did you have to add some steps before reconstructing the feature (same for `D_124`)\n1. I did'nt understand the cross referencing with `S_13` of `S_11`  : because I didn't find any data in : `train.S_11.isin([15,16,17])` \n\n\nThx for considering my questions😊",
    "1808471": "HI @raddar It seems that `B_19` would profit from `floorify_frac(x['B_19'])` as well. And the original values of `S_13` are all multiples of 1/1034, but as 1/1034 < 0.01, the added noise cannot be removed (diagrams are at the end of [my EDA](https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense)).",
    "1808410": "There are also many small tweaks that can be made on this dataset, for example:\n\n```\nx.loc[(x.R_13==0) & (x.R_17==0) & (x.R_20==0) & (x.R_8==0), 'R_6'] = 0\nx.loc[x.B_39==-1, 'B_36'] = 0\n```\n\nDo not know if it is worth the effort though to look for more\n",
    "1847883": "Thank you for sharing this. It seems that D_128, D_130, B_8, and R_1 are also hidden int features.",
    "1809629": "Great insight! Even @cdeotte worked on immensely reducing the size of data so that it can fit into our poor RAMs.",
    "1895681": "Hi Radar. Kudos to you. This is not only good work but also saves us a lot of time for everyone.",
    "1811436": "Looks like these 93 float32 features hold the majority of the signal. I get CV 0.78 easy by just using them.\nThanks!",
    "1809575": "Next step is to know what that features may be? 👀",
    "1809157": "Wow ! You just saved for all of us a lot of hours finding the hidden denominator for every single variable and preprocess it for us in a super clean light dataset.\n\nThanks a lot @raddar ! ",
    "1808399": "Nice catch @raddar, I was worried about converting some of these variables into float16 since we were losing precision in the process, but at the same time thinking about models might converge faster without that noise. I was planning to compare both with and w/o noise data models in cross validation, have you tried that yourself, seen any noticeable differences?",
    "2872889": "Hello!\nThis is my first comment.\nI need it to try find a team(on this you have inspired me, MANY THX!!)\nAnd also thank u million times for showing me the way to work with less RAM))",
    "2854030": "kudo for u",
    "2733293": "You just saved for all of us a lot of hours finding the hidden denominator for every single variable and preprocess it for us in a super clean light dataset.\n\nThanks a lot @raddar !\n\n\n",
    "2018821": "Oh wow! That's great! How did you figure this out? I'm a newb and want to know how you were able to learn there was random uniform noise of [0,0.01] added to each column.",
    "1941008": "Great work!",
    "1919178": "thanks a lot. saved me so much time. up vote.",
    "1889437": "Thanks a lot. It is always good to learn something :)",
    "1860813": "Hi @raddar just wanted to know, how you were able to deal with columns like D_87 which has mostly null values. Do you think converting the D_87 columns to an integer (1 == values exist and 0 == null), would be beneficial for GBDT models? My assumption is that if a value is present in a feature with mostly null columns, the customer is more likely to default. Is it an ok assumption to make, or no?",
    "1859946": "great-grand master",
    "1854531": "Hi @raddar and @cdeotte : do either of you happen to know if there's an efficient way to do the below (-1 -> nan) for cudf? It complains about data type, for cudf only.\n\nIf not, I guess the workaround is to load the data in pandas, reverse back to na using snippet below, and *then* call cudf.from_pandas()\n\nraddar: \"I wanted to make -1 to consistently represent NA. D_59 originally had some negative values. This should have no effect to models. This helps to reverse -1 to NA by just simply x[x==-1] = np.nan.\"\n\nI suspect it could be a big improvement to get mean, std, etc with nan as nan, and thus excluded from the algorithm. Of course probably even better if first handling some nan's as per raddar's cluster analysis. But even the quick 1 line change might see improvement and is worth testing.",
    "1854519": "Great detective work! 👍",
    "1853093": "@raddar: Thanks for this. I am also wondering why you prefer parquet format over pickle. I was reading [this blog](https://medium.com/@u.praneel.nihar/improving-read-write-store-performance-by-changing-file-formats-serialization-protocols-bfdb13114004) which compares parquet to pickle to CSV and concludes pickle is the best of all format. Am I missing anything? ",
    "1849141": "5GB -> 50GB with noise.\n\"Default Detection Comp\" with default dataset , lol 😄",
    "1849131": "subscribe + upvote !\n\nstart this competition from here 🎉",
    "1840438": "Great work! Thx!!",
    "1839437": "Thanks for posting this. I would be interested in seeing how you broke down the data work into sections small enough to work with in a notebook. ",
    "1830673": "Hello [raddar](https://www.kaggle.com/raddar)! I think that another 4 variables could be squeezed into the **int** type as well: **D104,D_112,B_8 and B_18 ** (look at the histograms of Version 16 of [this notebook](https://www.kaggle.com/code/carlosasdesouza/eda-carlos-amex).\n\nB_18 seems like a counting variable (Versions 14 and 15 of the same notebook above), while D_104,D_112 and B_8 seems like binary variables.\n\nFor D_104 and D_112,the noise seems to be random, although not in an uniform or normal way (distribution). For B_8, the noise seems to be pretty uniform around 1.",
    "1830195": "waw very ingenious ! You are aptly nicknamed Mr \"raddar\"🔥",
    "1829848": "1.7 GB , very smooth!!!  cheers, duh",
    "1825340": "I have seen that just a moment ago and thanks for your helping hand @naiborhujosua 👍",
    "1817024": "Impressive, great work thannks.",
    "1815587": "is the target variable provided in that American express comepetion\nwhat is the target variable",
    "1814915": "This is a really useful \"compression\" for using this data for fun and learning! Thanks!",
    "1814379": "Thanks a lot for this dataset! Great job.",
    "1812062": "Thanks @raddar. I was facing memory issue and this helps a lot",
    "1811897": "`x['D_59'] = floorify_frac(x['D_59']+5/48,1/48)`\n\nAnother question. Did you add 5 / 48 in order to shift minimum value to 0 so floor operation works consistently?",
    "1811645": "Hey @raddar. The test data is actually two separate datasets that can be distinguished based on max date. Would you consider separating public and private LB in two separate parquet ?",
    "1811085": "thx! I'm going to use this right now!",
    "1810941": "Fantastic detective work! 🔎",
    "2105294": "",
    "1872835": "",
    "1838152": "",
    "1819755": "",
    "1812681": "",
    "1812005": "",
    "1811123": "",
    "1905666": "Thanks a lot. very useful.",
    "1901372": "Thanks a lot, great",
    "1897656": "Thanks a lot. well done",
    "1896040": "Thank you for your information!",
    "1880768": "Thanks a lot,great",
    "1872383": "Thanks for contribution",
    "1861651": "THANKS FOR SHARING！",
    "1860727": "Thanks for sharing your great work!",
    "1853103": "Thanks a lot!",
    "1850949": "Thanks a lot !!!",
    "1840436": "Thanks a lot.",
    "1834578": "Thanks for sharing greate techniques :)",
    "1824177": "Just awesome! Thanks a lot!",
    "1817390": "Thanks a lot. It as really a good read.",
    "1816171": "Great content!! Thanks",
    "1813505": "Thanks for the teaching",
    "1812145": "Thanks for contribution",
    "1809658": "Thanks for contribution",
    "1809434": "Thanks @raddar for these contributions. "
  }
}