{
  "id": 332880,
  "title": "Help Needed!",
  "url": "/competitions/amex-default-prediction/discussion/332880",
  "author_name": "Varun Dutt",
  "post_date": "2022-06-23T18:40:47.519000",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Can someone suggest a computationally efficient method to group the data frames on 'costomer_ID' and then flatten all the columns into a single row?</p>",
  "messages": [
    {
      "id": 1830894,
      "postDate": "2022-06-23T18:51:09.800Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/varundutt9213\" target=\"_blank\">@varundutt9213</a> </p>\n<p>To make  \"wide\" panel data (I think this is what you are asking?) this is what I did, but it is far from elegant, nor computationally efficient:</p>\n<pre><code>    from functools import reduce\n\n    df1 = df.groupby(\"customer_ID\", as_index=False).nth([-1]).add_suffix('_1')\n    df1 = df1.rename(columns = {'customer_ID_1':'customer_ID'})\n    df2 = df.groupby(\"customer_ID\", as_index=False).nth([-2]).add_suffix('_2')\n    df2 = df2.rename(columns = {'customer_ID_2':'customer_ID'})\n    df3 = df.groupby(\"customer_ID\", as_index=False).nth([-3]).add_suffix('_3')\n    df3 = df3.rename(columns = {'customer_ID_3':'customer_ID'})\n    ...etc.\n\n    dfs = [df1, df2, df3, ... , df13]\n    # merge all DataFrames into one\n    df_wide = reduce(lambda left,right: pd.merge(left,right,on=['customer_ID'], how='left'), dfs)\n</code></pre>\n<p>It makes use of <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.core.groupby.GroupBy.nth.html\" target=\"_blank\"><code>pandas.core.groupby.GroupBy.nth</code></a> to create dataframes of the last statement, the next to last statement, <em>etc.</em> (note that non-existent statements are still created and are automatically  padded with NaN so there are no spaces/sparsity). Then merge these dataframes laterally into the final \"wide\"  file.</p>\n<p>That said, the file was huge as one is using all of the training data (and later, the test data) rather than just the last statement, and is thus unmanageable within a kaggle notebook memory (also, NaN are stored as floats, and this takes up space too). </p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @varundutt9213 \n\nTo make  \"wide\" panel data (I think this is what you are asking?) this is what I did, but it is far from elegant, nor computationally efficient:\n\n```\n    from functools import reduce\n\n    df1 = df.groupby(\"customer_ID\", as_index=False).nth([-1]).add_suffix('_1')\n    df1 = df1.rename(columns = {'customer_ID_1':'customer_ID'})\n    df2 = df.groupby(\"customer_ID\", as_index=False).nth([-2]).add_suffix('_2')\n    df2 = df2.rename(columns = {'customer_ID_2':'customer_ID'})\n    df3 = df.groupby(\"customer_ID\", as_index=False).nth([-3]).add_suffix('_3')\n    df3 = df3.rename(columns = {'customer_ID_3':'customer_ID'})\n    ...etc.\n\n    dfs = [df1, df2, df3, ... , df13]\n    # merge all DataFrames into one\n    df_wide = reduce(lambda left,right: pd.merge(left,right,on=['customer_ID'], how='left'), dfs)\n```\n\nIt makes use of [`pandas.core.groupby.GroupBy.nth`](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.core.groupby.GroupBy.nth.html) to create dataframes of the last statement, the next to last statement, *etc.* (note that non-existent statements are still created and are automatically  padded with NaN so there are no spaces/sparsity). Then merge these dataframes laterally into the final \"wide\"  file.\n\nThat said, the file was huge as one is using all of the training data (and later, the test data) rather than just the last statement, and is thus unmanageable within a kaggle notebook memory (also, NaN are stored as floats, and this takes up space too). \n\nAll the best,\ncarl",
      "votes": 6,
      "replies": [
        {
          "id": 1830933,
          "postDate": "2022-06-23T19:41:34.800Z",
          "content": "<p>Thank You, Carl a lot faster and computationally feasible way than the brute force for loop code I was trying!</p>",
          "rawMarkdown": "Thank You, Carl a lot faster and computationally feasible way than the brute force for loop code I was trying!",
          "votes": 1
        },
        {
          "id": 1831426,
          "postDate": "2022-06-24T07:00:01.867Z",
          "content": "<p>Assuming you have your training data in a dataframe with a multiindex, where level=0 is your <code>customer_ID</code> and level=1 is your month number, this works well:</p>\n<pre><code>train = train.unstack(level=-1)\ntrain.columns = [f\"{x[0]}_{x[1]}\" for x in train.columns]\n</code></pre>\n<p>I've not tried it on the full dataset though, so you may have to split into chunks and concatenate the results if you run out of RAM. </p>\n<p>Ref: <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.unstack.html#pandas.DataFrame.unstack\" target=\"_blank\"><code>pandas.DataFrame.unstack</code></a></p>",
          "rawMarkdown": "Assuming you have your training data in a dataframe with a multiindex, where level=0 is your `customer_ID` and level=1 is your month number, this works well:\n\n```\ntrain = train.unstack(level=-1)\ntrain.columns = [f\"{x[0]}_{x[1]}\" for x in train.columns]\n```\n\nI've not tried it on the full dataset though, so you may have to split into chunks and concatenate the results if you run out of RAM. \n\nRef: [`pandas.DataFrame.unstack`](https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.unstack.html#pandas.DataFrame.unstack)"
        }
      ]
    },
    {
      "id": 1830887,
      "postDate": "2022-06-23T18:40:47.520Z",
      "content": "<p>Can someone suggest a computationally efficient method to group the data frames on 'costomer_ID' and then flatten all the columns into a single row?</p>",
      "rawMarkdown": "Can someone suggest a computationally efficient method to group the data frames on 'costomer_ID' and then flatten all the columns into a single row?",
      "votes": 6
    },
    {
      "id": 1830934,
      "postDate": "2022-06-23T19:43:07.203Z",
      "content": "<p>Note that we can't flatten the data as is because not all customers have 13 rows of statements. First you must add rows so that every customer has 13 statements, next you can flatten. Personally, instead of flattening into 2D (two dimensional) data, I reshaped into 3D (three dimensional data). I posted discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\" target=\"_blank\">here</a> and notebook with code <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a></p>\n<p>Another approach is to aggregate the data OR only use the last statement for each customer. This will only use about <code>1/13th</code> of the data but it is what the top public notebook which scores LB 797 does. </p>\n<p>If you plan to use all 13 statements for each customer, note that the data structure will be large, and you will most likely need to read the original CSV file in as chunks, then process chunks, then save as separate files (i.e. 1 per chunk). This is what i do in my notebook with code above.</p>",
      "rawMarkdown": "Note that we can't flatten the data as is because not all customers have 13 rows of statements. First you must add rows so that every customer has 13 statements, next you can flatten. Personally, instead of flattening into 2D (two dimensional) data, I reshaped into 3D (three dimensional data). I posted discussion [here][1] and notebook with code [here][2]\n\nAnother approach is to aggregate the data OR only use the last statement for each customer. This will only use about `1/13th` of the data but it is what the top public notebook which scores LB 797 does. \n  \nIf you plan to use all 13 statements for each customer, note that the data structure will be large, and you will most likely need to read the original CSV file in as chunks, then process chunks, then save as separate files (i.e. 1 per chunk). This is what i do in my notebook with code above.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790",
      "votes": 4,
      "replies": [
        {
          "id": 1830972,
          "postDate": "2022-06-23T20:15:03.860Z",
          "content": "<p>Hey Chris,</p>\n<p>Thank you for the detailed explanation I did code my data preprocessing and modeling scripts in such a way so as to have the option to process and combine the frames in chunks preempting the ROR issue while processing the entire thing, for the issue of not all customers having 13 entries I padded the smaller ones with the value I used for NaNs but the 3D reshaping sounds interesting I'll definitely have a look. Using the last statement makes a lot of sense actually bcoz I have seen in my few experiments that despite having a big dataset there is a huge overfitting problem and I suspect there are some features that are mainly contributing to overfitting and not to generalization. </p>\n<p>Thank you for the code I wanted to ask for help because I wanted to enquire about an efficient way to do this for future use as well.</p>\n<p>Thank you <br>\nVarun</p>",
          "rawMarkdown": "Hey Chris,\n\nThank you for the detailed explanation I did code my data preprocessing and modeling scripts in such a way so as to have the option to process and combine the frames in chunks preempting the ROR issue while processing the entire thing, for the issue of not all customers having 13 entries I padded the smaller ones with the value I used for NaNs but the 3D reshaping sounds interesting I'll definitely have a look. Using the last statement makes a lot of sense actually bcoz I have seen in my few experiments that despite having a big dataset there is a huge overfitting problem and I suspect there are some features that are mainly contributing to overfitting and not to generalization. \n\nThank you for the code I wanted to ask for help because I wanted to enquire about an efficient way to do this for future use as well.\n\nThank you \nVarun",
          "votes": 1
        },
        {
          "id": 1831019,
          "postDate": "2022-06-23T21:27:22.977Z",
          "content": "<p>In my provided code notebook, the \"3D data magic\" happens in hidden code cell #6 with</p>\n<pre><code>data = train.iloc[:,1:-1].values.reshape((-1,13,188))\n</code></pre>\n<p>Once you have a dataframe where every consecutive 13 rows are one customer with the rows sorted in time order, then you can extract the data with <code>df[COLS].values</code> and reshape it to 3D with <code>df[COLS].values.reshape((-1,13,188))</code>. Then train your models with the 3D NumPy array. </p>",
          "rawMarkdown": "In my provided code notebook, the \"3D data magic\" happens in hidden code cell #6 with\n\n    data = train.iloc[:,1:-1].values.reshape((-1,13,188))\n\nOnce you have a dataframe where every consecutive 13 rows are one customer with the rows sorted in time order, then you can extract the data with `df[COLS].values` and reshape it to 3D with `df[COLS].values.reshape((-1,13,188))`. Then train your models with the 3D NumPy array. ",
          "votes": 2
        },
        {
          "id": 1838938,
          "postDate": "2022-07-01T02:35:33.937Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for sharing detailed explanation and insights. ☀️</p>",
          "rawMarkdown": "@cdeotte Thanks for sharing detailed explanation and insights. ☀️"
        }
      ]
    },
    {
      "id": 1853534,
      "postDate": "2022-07-13T00:10:32.747Z",
      "content": "<p>1) sorted the rows by S_2 for each customer_ID<br>\n2) assign a rank value for each rows, such as 1, 2,3,4,5…13<br>\n3) then loop, mege<br>\n  df_1 = df.loc[df.rk==1]<br>\n   for k in range(2, 13):<br>\n       # remember to rename columns<br>\n       df = pd.merge(df_1,  df.loc[df.rk==k])….</p>",
      "rawMarkdown": "1) sorted the rows by S_2 for each customer_ID\n2) assign a rank value for each rows, such as 1, 2,3,4,5...13\n3) then loop, mege\n  df_1 = df.loc[df.rk==1]\n   for k in range(2, 13):\n       # remember to rename columns\n       df = pd.merge(df_1,  df.loc[df.rk==k])....\n      "
    },
    {
      "id": 1838930,
      "postDate": "2022-07-01T02:21:58.603Z",
      "content": "<p>I think this may not improve the performance of your model since there are some customer_IDs with less than 13 records. This could add useless information to your model. The best approach is to aggregate the data for each customer. For example, you may want to aggregate mean, min, max, last of numerical features and number of unique and last value of categorical features.</p>",
      "rawMarkdown": "I think this may not improve the performance of your model since there are some customer_IDs with less than 13 records. This could add useless information to your model. The best approach is to aggregate the data for each customer. For example, you may want to aggregate mean, min, max, last of numerical features and number of unique and last value of categorical features."
    },
    {
      "id": 1830900,
      "postDate": "2022-06-23T19:09:19.237Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1830894,
      "author_name": "Carl McBride Ellis",
      "author_url": "",
      "post_date": "2022-06-23T18:51:09.800000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/varundutt9213\" target=\"_blank\">@varundutt9213</a> </p>\n<p>To make  \"wide\" panel data (I think this is what you are asking?) this is what I did, but it is far from elegant, nor computationally efficient:</p>\n<pre><code>    from functools import reduce\n\n    df1 = df.groupby(\"customer_ID\", as_index=False).nth([-1]).add_suffix('_1')\n    df1 = df1.rename(columns = {'customer_ID_1':'customer_ID'})\n    df2 = df.groupby(\"customer_ID\", as_index=False).nth([-2]).add_suffix('_2')\n    df2 = df2.rename(columns = {'customer_ID_2':'customer_ID'})\n    df3 = df.groupby(\"customer_ID\", as_index=False).nth([-3]).add_suffix('_3')\n    df3 = df3.rename(columns = {'customer_ID_3':'customer_ID'})\n    ...etc.\n\n    dfs = [df1, df2, df3, ... , df13]\n    # merge all DataFrames into one\n    df_wide = reduce(lambda left,right: pd.merge(left,right,on=['customer_ID'], how='left'), dfs)\n</code></pre>\n<p>It makes use of <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.core.groupby.GroupBy.nth.html\" target=\"_blank\"><code>pandas.core.groupby.GroupBy.nth</code></a> to create dataframes of the last statement, the next to last statement, <em>etc.</em> (note that non-existent statements are still created and are automatically  padded with NaN so there are no spaces/sparsity). Then merge these dataframes laterally into the final \"wide\"  file.</p>\n<p>That said, the file was huge as one is using all of the training data (and later, the test data) rather than just the last statement, and is thus unmanageable within a kaggle notebook memory (also, NaN are stored as floats, and this takes up space too). </p>\n<p>All the best,<br>\ncarl</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1830933,
          "author_name": "Varun Dutt",
          "author_url": "",
          "post_date": "2022-06-23T19:41:34.800000",
          "content": "<p>Thank You, Carl a lot faster and computationally feasible way than the brute force for loop code I was trying!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1831426,
          "author_name": "Burrito Dan",
          "author_url": "",
          "post_date": "2022-06-24T07:00:01.867000",
          "content": "<p>Assuming you have your training data in a dataframe with a multiindex, where level=0 is your <code>customer_ID</code> and level=1 is your month number, this works well:</p>\n<pre><code>train = train.unstack(level=-1)\ntrain.columns = [f\"{x[0]}_{x[1]}\" for x in train.columns]\n</code></pre>\n<p>I've not tried it on the full dataset though, so you may have to split into chunks and concatenate the results if you run out of RAM. </p>\n<p>Ref: <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.unstack.html#pandas.DataFrame.unstack\" target=\"_blank\"><code>pandas.DataFrame.unstack</code></a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1830934,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-06-23T19:43:07.203000",
      "content": "<p>Note that we can't flatten the data as is because not all customers have 13 rows of statements. First you must add rows so that every customer has 13 statements, next you can flatten. Personally, instead of flattening into 2D (two dimensional) data, I reshaped into 3D (three dimensional data). I posted discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\" target=\"_blank\">here</a> and notebook with code <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a></p>\n<p>Another approach is to aggregate the data OR only use the last statement for each customer. This will only use about <code>1/13th</code> of the data but it is what the top public notebook which scores LB 797 does. </p>\n<p>If you plan to use all 13 statements for each customer, note that the data structure will be large, and you will most likely need to read the original CSV file in as chunks, then process chunks, then save as separate files (i.e. 1 per chunk). This is what i do in my notebook with code above.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1830972,
          "author_name": "Varun Dutt",
          "author_url": "",
          "post_date": "2022-06-23T20:15:03.860000",
          "content": "<p>Hey Chris,</p>\n<p>Thank you for the detailed explanation I did code my data preprocessing and modeling scripts in such a way so as to have the option to process and combine the frames in chunks preempting the ROR issue while processing the entire thing, for the issue of not all customers having 13 entries I padded the smaller ones with the value I used for NaNs but the 3D reshaping sounds interesting I'll definitely have a look. Using the last statement makes a lot of sense actually bcoz I have seen in my few experiments that despite having a big dataset there is a huge overfitting problem and I suspect there are some features that are mainly contributing to overfitting and not to generalization. </p>\n<p>Thank you for the code I wanted to ask for help because I wanted to enquire about an efficient way to do this for future use as well.</p>\n<p>Thank you <br>\nVarun</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1831019,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-23T21:27:22.977000",
          "content": "<p>In my provided code notebook, the \"3D data magic\" happens in hidden code cell #6 with</p>\n<pre><code>data = train.iloc[:,1:-1].values.reshape((-1,13,188))\n</code></pre>\n<p>Once you have a dataframe where every consecutive 13 rows are one customer with the rows sorted in time order, then you can extract the data with <code>df[COLS].values</code> and reshape it to 3D with <code>df[COLS].values.reshape((-1,13,188))</code>. Then train your models with the 3D NumPy array. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1838938,
          "author_name": "Muhammad Irfan Azam",
          "author_url": "",
          "post_date": "2022-07-01T02:35:33.937000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for sharing detailed explanation and insights. ☀️</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1853534,
      "author_name": "jxlijunhao",
      "author_url": "",
      "post_date": "2022-07-13T00:10:32.747000",
      "content": "<p>1) sorted the rows by S_2 for each customer_ID<br>\n2) assign a rank value for each rows, such as 1, 2,3,4,5…13<br>\n3) then loop, mege<br>\n  df_1 = df.loc[df.rk==1]<br>\n   for k in range(2, 13):<br>\n       # remember to rename columns<br>\n       df = pd.merge(df_1,  df.loc[df.rk==k])….</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1838930,
      "author_name": "1110Ra",
      "author_url": "",
      "post_date": "2022-07-01T02:21:58.603000",
      "content": "<p>I think this may not improve the performance of your model since there are some customer_IDs with less than 13 records. This could add useless information to your model. The best approach is to aggregate the data for each customer. For example, you may want to aggregate mean, min, max, last of numerical features and number of unique and last value of categorical features.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1830900,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-23T19:09:19.237000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1830894": "Dear @varundutt9213 \n\nTo make  \"wide\" panel data (I think this is what you are asking?) this is what I did, but it is far from elegant, nor computationally efficient:\n\n```\n    from functools import reduce\n\n    df1 = df.groupby(\"customer_ID\", as_index=False).nth([-1]).add_suffix('_1')\n    df1 = df1.rename(columns = {'customer_ID_1':'customer_ID'})\n    df2 = df.groupby(\"customer_ID\", as_index=False).nth([-2]).add_suffix('_2')\n    df2 = df2.rename(columns = {'customer_ID_2':'customer_ID'})\n    df3 = df.groupby(\"customer_ID\", as_index=False).nth([-3]).add_suffix('_3')\n    df3 = df3.rename(columns = {'customer_ID_3':'customer_ID'})\n    ...etc.\n\n    dfs = [df1, df2, df3, ... , df13]\n    # merge all DataFrames into one\n    df_wide = reduce(lambda left,right: pd.merge(left,right,on=['customer_ID'], how='left'), dfs)\n```\n\nIt makes use of [`pandas.core.groupby.GroupBy.nth`](https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.core.groupby.GroupBy.nth.html) to create dataframes of the last statement, the next to last statement, *etc.* (note that non-existent statements are still created and are automatically  padded with NaN so there are no spaces/sparsity). Then merge these dataframes laterally into the final \"wide\"  file.\n\nThat said, the file was huge as one is using all of the training data (and later, the test data) rather than just the last statement, and is thus unmanageable within a kaggle notebook memory (also, NaN are stored as floats, and this takes up space too). \n\nAll the best,\ncarl",
    "1830887": "Can someone suggest a computationally efficient method to group the data frames on 'costomer_ID' and then flatten all the columns into a single row?",
    "1830934": "Note that we can't flatten the data as is because not all customers have 13 rows of statements. First you must add rows so that every customer has 13 statements, next you can flatten. Personally, instead of flattening into 2D (two dimensional) data, I reshaped into 3D (three dimensional data). I posted discussion [here][1] and notebook with code [here][2]\n\nAnother approach is to aggregate the data OR only use the last statement for each customer. This will only use about `1/13th` of the data but it is what the top public notebook which scores LB 797 does. \n  \nIf you plan to use all 13 statements for each customer, note that the data structure will be large, and you will most likely need to read the original CSV file in as chunks, then process chunks, then save as separate files (i.e. 1 per chunk). This is what i do in my notebook with code above.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790",
    "1853534": "1) sorted the rows by S_2 for each customer_ID\n2) assign a rank value for each rows, such as 1, 2,3,4,5...13\n3) then loop, mege\n  df_1 = df.loc[df.rk==1]\n   for k in range(2, 13):\n       # remember to rename columns\n       df = pd.merge(df_1,  df.loc[df.rk==k])....\n      ",
    "1838930": "I think this may not improve the performance of your model since there are some customer_IDs with less than 13 records. This could add useless information to your model. The best approach is to aggregate the data for each customer. For example, you may want to aggregate mean, min, max, last of numerical features and number of unique and last value of categorical features.",
    "1830900": ""
  }
}