{
  "id": 331059,
  "title": "RAPIDS cuDF aggregation operation numeric stability problem",
  "url": "/competitions/amex-default-prediction/discussion/331059",
  "author_name": "zhuoyin94",
  "post_date": "2022-06-15T15:49:12.002000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello Kagglers !</p>\n<p>When I do feature engineering with NVIDIA RAPIDS cuDF tool on the AMEX dataset, I found that the same aggregation function will have a slightly different aggregation result. Let's use 'D_44' feature for example:</p>\n<pre><code>feat_name = 'D_44'\n\ntmp_cudf = total_cudf[['customer_ID', feat_name]].copy()\n\nfeat_df = tmp_cudf.groupby(['customer_ID'])[feat_name].mean().reset_index()\nfeat_df.rename({feat_name: '{}_x'.format(feat_name)}, axis=1, inplace=True)\n\ntmp_feat_df = tmp_cudf.groupby(['customer_ID'])[feat_name].mean().reset_index()\ntmp_feat_df.rename({feat_name: '{}_y'.format(feat_name)}, axis=1, inplace=True)\n\nfeat_df = cudf.merge(feat_df, tmp_feat_df, on=['customer_ID'], how='left')\n\nfeat_df['diff'] = feat_df['{}_x'.format(feat_name)] != feat_df['{}_y'.format(feat_name)]\n\nprint('\\nBasic DataFrame information')\nprint('+++++++++++++++')\nprint(tmp_cudf.info())\n\nprint('\\nNumeric difference')\nprint('+++++++++++++++')\nprint(\n    feat_df.shape,\n    feat_df['diff'].astype(int).sum(),\n    feat_df['diff'].astype(int).sum() / len(feat_df) * 100\n)\n</code></pre>\n<p>The sample of the dataset:</p>\n<pre><code>tmp_cudf.sample(10)\n\ncustomer_ID    D_44\n812026    0.005080824\n740528    0.001454785\n391819    &lt;NA&gt;\n19165    0.002080851\n144817    0.002630037\n570573    0.006430216\n725572    0.00341707\n1365101    9.863440209e-05\n661881    0.009342765\n305162    0.134882405\n</code></pre>\n<p>The outcome of the code snippet:</p>\n<pre><code>Basic DataFrame information\n+++++++++++++++\n&lt;class 'cudf.core.dataframe.DataFrame'&gt;\nRangeIndex: 16895213 entries, 0 to 16895212\nData columns (total 2 columns):\n #   Column       Dtype\n---  ------       -----\n 0   customer_ID  int32\n 1   D_44         float32\ndtypes: float32(1), int32(1)\nmemory usage: 130.9 MB\nNone\n\nNumeric difference\n+++++++++++++++\n(1383534, 4) 85277 6.163708300627235\n</code></pre>\n<p>You can run the code snippet multiple times, you will find that the difference precent varies each time. Though the difference is small, but when you do these kind of aggregation across thousands of features, it will cause the numeric stability problem. In my experiments, I use xgboost-gpu as my learner, when the random seed of xgboost-gpu is fixed, I run the same feature script multiple times, and each time I train my xgboost-gpu learner. I compare the offline 5-fold cv score and average cv score each time. And the cv results shake range from 7920 +- 15, which is a huge difference.</p>\n<p>How can I fix this problem? Did you guys face the same problems?</p>",
  "messages": [
    {
      "id": 1821509,
      "postDate": "2022-06-15T15:49:12.003Z",
      "content": "<p>Hello Kagglers !</p>\n<p>When I do feature engineering with NVIDIA RAPIDS cuDF tool on the AMEX dataset, I found that the same aggregation function will have a slightly different aggregation result. Let's use 'D_44' feature for example:</p>\n<pre><code>feat_name = 'D_44'\n\ntmp_cudf = total_cudf[['customer_ID', feat_name]].copy()\n\nfeat_df = tmp_cudf.groupby(['customer_ID'])[feat_name].mean().reset_index()\nfeat_df.rename({feat_name: '{}_x'.format(feat_name)}, axis=1, inplace=True)\n\ntmp_feat_df = tmp_cudf.groupby(['customer_ID'])[feat_name].mean().reset_index()\ntmp_feat_df.rename({feat_name: '{}_y'.format(feat_name)}, axis=1, inplace=True)\n\nfeat_df = cudf.merge(feat_df, tmp_feat_df, on=['customer_ID'], how='left')\n\nfeat_df['diff'] = feat_df['{}_x'.format(feat_name)] != feat_df['{}_y'.format(feat_name)]\n\nprint('\\nBasic DataFrame information')\nprint('+++++++++++++++')\nprint(tmp_cudf.info())\n\nprint('\\nNumeric difference')\nprint('+++++++++++++++')\nprint(\n    feat_df.shape,\n    feat_df['diff'].astype(int).sum(),\n    feat_df['diff'].astype(int).sum() / len(feat_df) * 100\n)\n</code></pre>\n<p>The sample of the dataset:</p>\n<pre><code>tmp_cudf.sample(10)\n\ncustomer_ID    D_44\n812026    0.005080824\n740528    0.001454785\n391819    &lt;NA&gt;\n19165    0.002080851\n144817    0.002630037\n570573    0.006430216\n725572    0.00341707\n1365101    9.863440209e-05\n661881    0.009342765\n305162    0.134882405\n</code></pre>\n<p>The outcome of the code snippet:</p>\n<pre><code>Basic DataFrame information\n+++++++++++++++\n&lt;class 'cudf.core.dataframe.DataFrame'&gt;\nRangeIndex: 16895213 entries, 0 to 16895212\nData columns (total 2 columns):\n #   Column       Dtype\n---  ------       -----\n 0   customer_ID  int32\n 1   D_44         float32\ndtypes: float32(1), int32(1)\nmemory usage: 130.9 MB\nNone\n\nNumeric difference\n+++++++++++++++\n(1383534, 4) 85277 6.163708300627235\n</code></pre>\n<p>You can run the code snippet multiple times, you will find that the difference precent varies each time. Though the difference is small, but when you do these kind of aggregation across thousands of features, it will cause the numeric stability problem. In my experiments, I use xgboost-gpu as my learner, when the random seed of xgboost-gpu is fixed, I run the same feature script multiple times, and each time I train my xgboost-gpu learner. I compare the offline 5-fold cv score and average cv score each time. And the cv results shake range from 7920 +- 15, which is a huge difference.</p>\n<p>How can I fix this problem? Did you guys face the same problems?</p>",
      "rawMarkdown": "Hello Kagglers !\n\nWhen I do feature engineering with NVIDIA RAPIDS cuDF tool on the AMEX dataset, I found that the same aggregation function will have a slightly different aggregation result. Let's use 'D_44' feature for example:\n\n```python\nfeat_name = 'D_44'\n\ntmp_cudf = total_cudf[['customer_ID', feat_name]].copy()\n\nfeat_df = tmp_cudf.groupby(['customer_ID'])[feat_name].mean().reset_index()\nfeat_df.rename({feat_name: '{}_x'.format(feat_name)}, axis=1, inplace=True)\n\ntmp_feat_df = tmp_cudf.groupby(['customer_ID'])[feat_name].mean().reset_index()\ntmp_feat_df.rename({feat_name: '{}_y'.format(feat_name)}, axis=1, inplace=True)\n\nfeat_df = cudf.merge(feat_df, tmp_feat_df, on=['customer_ID'], how='left')\n\nfeat_df['diff'] = feat_df['{}_x'.format(feat_name)] != feat_df['{}_y'.format(feat_name)]\n\nprint('\\nBasic DataFrame information')\nprint('+++++++++++++++')\nprint(tmp_cudf.info())\n\nprint('\\nNumeric difference')\nprint('+++++++++++++++')\nprint(\n    feat_df.shape,\n    feat_df['diff'].astype(int).sum(),\n    feat_df['diff'].astype(int).sum() / len(feat_df) * 100\n)\n```\n\nThe sample of the dataset:\n```\ntmp_cudf.sample(10)\n\ncustomer_ID\tD_44\n812026\t0.005080824\n740528\t0.001454785\n391819\t<NA>\n19165\t0.002080851\n144817\t0.002630037\n570573\t0.006430216\n725572\t0.00341707\n1365101\t9.863440209e-05\n661881\t0.009342765\n305162\t0.134882405\n```\n\nThe outcome of the code snippet:\n```\nBasic DataFrame information\n+++++++++++++++\n<class 'cudf.core.dataframe.DataFrame'>\nRangeIndex: 16895213 entries, 0 to 16895212\nData columns (total 2 columns):\n #   Column       Dtype\n---  ------       -----\n 0   customer_ID  int32\n 1   D_44         float32\ndtypes: float32(1), int32(1)\nmemory usage: 130.9 MB\nNone\n\nNumeric difference\n+++++++++++++++\n(1383534, 4) 85277 6.163708300627235\n```\n\nYou can run the code snippet multiple times, you will find that the difference precent varies each time. Though the difference is small, but when you do these kind of aggregation across thousands of features, it will cause the numeric stability problem. In my experiments, I use xgboost-gpu as my learner, when the random seed of xgboost-gpu is fixed, I run the same feature script multiple times, and each time I train my xgboost-gpu learner. I compare the offline 5-fold cv score and average cv score each time. And the cv results shake range from 7920 +- 15, which is a huge difference.\n\nHow can I fix this problem? Did you guys face the same problems?",
      "votes": 3
    },
    {
      "id": 1822041,
      "postDate": "2022-06-16T03:14:33.683Z",
      "content": "<p>Could I please ask how you are able to set the seed for your xgboost?</p>",
      "rawMarkdown": "Could I please ask how you are able to set the seed for your xgboost?",
      "replies": [
        {
          "id": 1822073,
          "postDate": "2022-06-16T03:56:11.580Z",
          "content": "<p>Thanks for responding.</p>\n<p>Firstly, I use the following code snippet to set the overall random seed:</p>\n<pre><code>def seed_everything(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n</code></pre>\n<p>Secondly, for xgboost-gpu training params, I specify the random seed with the same random seed int the above code snippet:</p>\n<pre><code>xgb_params = {\n    'n_estimators': n_estimators,\n    'max_depth': 4,\n    'max_bin': 64,\n    'learning_rate': 0.05,\n    'verbosity': 0,\n    'objective': 'binary:logistic',\n    'booster': 'gbtree',\n    'colsample_bytree': 0.9,\n    'colsample_bylevel': 0.9,\n    'subsample': 0.95,\n    'disable_default_eval_metric': 1,\n    'missing': nan_token,\n    'enable_categorical': True,\n    'reg_alpha': 1.5,\n    'reg_lambda': 70,\n    'random_state': seed\n}\n</code></pre>\n<p>Thirdly, for the same feature engineering result that the cuDF feature engineering code script output, I ran the xgboost-gpu training script multiple times. The training and validation 5-fold cv results are extractly the same, I even tested the training script across different GPU equipment and got the same result.</p>\n<p>The above experiments show that the xgboost-gpu is not the reason of the problem. The problem is caused the cuDF aggregation operation. I guess that the numeric difference cased by cuDF infulenced the xgboost-gpu split gain computation.</p>",
          "rawMarkdown": "Thanks for responding.\n\nFirstly, I use the following code snippet to set the overall random seed:\n\n```python\ndef seed_everything(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n```\n\nSecondly, for xgboost-gpu training params, I specify the random seed with the same random seed int the above code snippet:\n\n```python\nxgb_params = {\n    'n_estimators': n_estimators,\n    'max_depth': 4,\n    'max_bin': 64,\n    'learning_rate': 0.05,\n    'verbosity': 0,\n    'objective': 'binary:logistic',\n    'booster': 'gbtree',\n    'colsample_bytree': 0.9,\n    'colsample_bylevel': 0.9,\n    'subsample': 0.95,\n    'disable_default_eval_metric': 1,\n    'missing': nan_token,\n    'enable_categorical': True,\n    'reg_alpha': 1.5,\n    'reg_lambda': 70,\n    'random_state': seed\n}\n```\n\nThirdly, for the same feature engineering result that the cuDF feature engineering code script output, I ran the xgboost-gpu training script multiple times. The training and validation 5-fold cv results are extractly the same, I even tested the training script across different GPU equipment and got the same result.\n\nThe above experiments show that the xgboost-gpu is not the reason of the problem. The problem is caused the cuDF aggregation operation. I guess that the numeric difference cased by cuDF infulenced the xgboost-gpu split gain computation."
        }
      ]
    },
    {
      "id": 1821513,
      "postDate": "2022-06-15T15:55:17.277Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1822041,
      "author_name": "Radek Osmulski",
      "author_url": "",
      "post_date": "2022-06-16T03:14:33.683000",
      "content": "<p>Could I please ask how you are able to set the seed for your xgboost?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1822073,
          "author_name": "zhuoyin94",
          "author_url": "",
          "post_date": "2022-06-16T03:56:11.580000",
          "content": "<p>Thanks for responding.</p>\n<p>Firstly, I use the following code snippet to set the overall random seed:</p>\n<pre><code>def seed_everything(seed):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n</code></pre>\n<p>Secondly, for xgboost-gpu training params, I specify the random seed with the same random seed int the above code snippet:</p>\n<pre><code>xgb_params = {\n    'n_estimators': n_estimators,\n    'max_depth': 4,\n    'max_bin': 64,\n    'learning_rate': 0.05,\n    'verbosity': 0,\n    'objective': 'binary:logistic',\n    'booster': 'gbtree',\n    'colsample_bytree': 0.9,\n    'colsample_bylevel': 0.9,\n    'subsample': 0.95,\n    'disable_default_eval_metric': 1,\n    'missing': nan_token,\n    'enable_categorical': True,\n    'reg_alpha': 1.5,\n    'reg_lambda': 70,\n    'random_state': seed\n}\n</code></pre>\n<p>Thirdly, for the same feature engineering result that the cuDF feature engineering code script output, I ran the xgboost-gpu training script multiple times. The training and validation 5-fold cv results are extractly the same, I even tested the training script across different GPU equipment and got the same result.</p>\n<p>The above experiments show that the xgboost-gpu is not the reason of the problem. The problem is caused the cuDF aggregation operation. I guess that the numeric difference cased by cuDF infulenced the xgboost-gpu split gain computation.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1821513,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-15T15:55:17.277000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1821509": "Hello Kagglers !\n\nWhen I do feature engineering with NVIDIA RAPIDS cuDF tool on the AMEX dataset, I found that the same aggregation function will have a slightly different aggregation result. Let's use 'D_44' feature for example:\n\n```python\nfeat_name = 'D_44'\n\ntmp_cudf = total_cudf[['customer_ID', feat_name]].copy()\n\nfeat_df = tmp_cudf.groupby(['customer_ID'])[feat_name].mean().reset_index()\nfeat_df.rename({feat_name: '{}_x'.format(feat_name)}, axis=1, inplace=True)\n\ntmp_feat_df = tmp_cudf.groupby(['customer_ID'])[feat_name].mean().reset_index()\ntmp_feat_df.rename({feat_name: '{}_y'.format(feat_name)}, axis=1, inplace=True)\n\nfeat_df = cudf.merge(feat_df, tmp_feat_df, on=['customer_ID'], how='left')\n\nfeat_df['diff'] = feat_df['{}_x'.format(feat_name)] != feat_df['{}_y'.format(feat_name)]\n\nprint('\\nBasic DataFrame information')\nprint('+++++++++++++++')\nprint(tmp_cudf.info())\n\nprint('\\nNumeric difference')\nprint('+++++++++++++++')\nprint(\n    feat_df.shape,\n    feat_df['diff'].astype(int).sum(),\n    feat_df['diff'].astype(int).sum() / len(feat_df) * 100\n)\n```\n\nThe sample of the dataset:\n```\ntmp_cudf.sample(10)\n\ncustomer_ID\tD_44\n812026\t0.005080824\n740528\t0.001454785\n391819\t<NA>\n19165\t0.002080851\n144817\t0.002630037\n570573\t0.006430216\n725572\t0.00341707\n1365101\t9.863440209e-05\n661881\t0.009342765\n305162\t0.134882405\n```\n\nThe outcome of the code snippet:\n```\nBasic DataFrame information\n+++++++++++++++\n<class 'cudf.core.dataframe.DataFrame'>\nRangeIndex: 16895213 entries, 0 to 16895212\nData columns (total 2 columns):\n #   Column       Dtype\n---  ------       -----\n 0   customer_ID  int32\n 1   D_44         float32\ndtypes: float32(1), int32(1)\nmemory usage: 130.9 MB\nNone\n\nNumeric difference\n+++++++++++++++\n(1383534, 4) 85277 6.163708300627235\n```\n\nYou can run the code snippet multiple times, you will find that the difference precent varies each time. Though the difference is small, but when you do these kind of aggregation across thousands of features, it will cause the numeric stability problem. In my experiments, I use xgboost-gpu as my learner, when the random seed of xgboost-gpu is fixed, I run the same feature script multiple times, and each time I train my xgboost-gpu learner. I compare the offline 5-fold cv score and average cv score each time. And the cv results shake range from 7920 +- 15, which is a huge difference.\n\nHow can I fix this problem? Did you guys face the same problems?",
    "1822041": "Could I please ask how you are able to set the seed for your xgboost?",
    "1821513": ""
  }
}