{
  "id": 327828,
  "title": "Kaggle Dataset for Transformers and RNNs",
  "url": "/competitions/amex-default-prediction/discussion/327828",
  "author_name": "Chris Deotte",
  "post_date": "2022-05-29T13:29:49.977000",
  "votes": 159,
  "comment_count": 28,
  "views": 0,
  "content": "<h1>Competition Data CSV</h1>\n<p>This competition provides data about credit card customers in the form of <code>CSV</code> files. We cannot use these files without modification to train Transformers and RNNs because each customer has a different number of credit card statements. </p>\n<p>When inputting time series data into a Transformer and RNN, each customer needs the same sequence length (of 13 in this comp).</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/trans.png\" alt=\"\"></p>\n<h1>Competition Data NumPy Array</h1>\n<p>I have converted the competition data into <code>NumPy array</code> files where each customer has sequence length 13 and posted a Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\" target=\"_blank\">here</a> (This data is created in my notebook <a href=\"https://tinyurl.com/4sjwms6v\" target=\"_blank\">here</a>)</p>\n<p>For each customer with less than 13 credit card statements, I have padded their data with <code>pad = -1</code>. See example below where the customer has been padded with four -1's:</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/data_block.png\" alt=\"\"></p>\n<h1>Kaggle Dataset Details</h1>\n<p>My Kaggle dataset is <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\" target=\"_blank\">here</a>. The train data has been split into 10 NumPy arrays named <code>data_1.npy</code> thru <code>data_10.npy</code>. Each array has dimension <code>(45891, 13, 188)</code> which is <code>customer x statement x feature</code>. </p>\n<p>The associated targets are contained in the files <code>targets_1.pqt</code> thru <code>targets_10.pqt</code>. These are parquet files with columns <code>customer_ID</code> and <code>target</code> and have dimension <code>(45891, 2)</code>. An example training with these files is <a href=\"https://tinyurl.com/4sjwms6v\" target=\"_blank\">here</a>. Note that the <code>customer_ID</code> is a <code>int64</code> instead of the provided <code>string512</code>. It was compressed using the code shown <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">here</a>.</p>\n<p>The test data is split in 20 NumPy arrays. Each array has <code>46231</code> customers. <strong>Important Note</strong>: If you concatenate all these files, the row order of the <code>924621</code> test customers is <strong>not the same</strong> as the file <code>sample_submission.csv</code>. The order is contained in my <code>submission.csv</code> <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns?select=submission.csv\" target=\"_blank\">here</a>. So infer the 20 NumPy arrays in their order and then overwrite my prediction column with <code>submission['prediction'] = preds</code>.</p>\n<h1>Transformers and RNNs</h1>\n<p>It will be exciting to see whether GBT (gradient boosted trees) or NN will achieve the best accuracy in this competition. I posted a RNN notebook starter <a href=\"https://tinyurl.com/4sjwms6v\" target=\"_blank\">here</a> and discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327761\" target=\"_blank\">here</a>. If you want to build a Transformer, you can use my Transformer notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-112\" target=\"_blank\">here</a> from Ventilator competition and input Amex competition data. Additionally there are many great RNN notebooks in Ventilator competition for inspiration.</p>\n<h1>UPDATE: Transformer Starter - LB 0.790</h1>\n<p>I created and published a Transformer starter notebook <a href=\"https://tinyurl.com/45uu256r\" target=\"_blank\">here</a> today that uses my Kaggle dataset. It achieves 5-Fold CV 0.787 and LB 0.789. Enjoy!</p>\n<h1>Have Fun!</h1>",
  "messages": [
    {
      "id": 1804818,
      "postDate": "2022-05-29T13:29:49.977Z",
      "content": "<h1>Competition Data CSV</h1>\n<p>This competition provides data about credit card customers in the form of <code>CSV</code> files. We cannot use these files without modification to train Transformers and RNNs because each customer has a different number of credit card statements. </p>\n<p>When inputting time series data into a Transformer and RNN, each customer needs the same sequence length (of 13 in this comp).</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/trans.png\" alt=\"\"></p>\n<h1>Competition Data NumPy Array</h1>\n<p>I have converted the competition data into <code>NumPy array</code> files where each customer has sequence length 13 and posted a Kaggle dataset <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\" target=\"_blank\">here</a> (This data is created in my notebook <a href=\"https://tinyurl.com/4sjwms6v\" target=\"_blank\">here</a>)</p>\n<p>For each customer with less than 13 credit card statements, I have padded their data with <code>pad = -1</code>. See example below where the customer has been padded with four -1's:</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/data_block.png\" alt=\"\"></p>\n<h1>Kaggle Dataset Details</h1>\n<p>My Kaggle dataset is <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\" target=\"_blank\">here</a>. The train data has been split into 10 NumPy arrays named <code>data_1.npy</code> thru <code>data_10.npy</code>. Each array has dimension <code>(45891, 13, 188)</code> which is <code>customer x statement x feature</code>. </p>\n<p>The associated targets are contained in the files <code>targets_1.pqt</code> thru <code>targets_10.pqt</code>. These are parquet files with columns <code>customer_ID</code> and <code>target</code> and have dimension <code>(45891, 2)</code>. An example training with these files is <a href=\"https://tinyurl.com/4sjwms6v\" target=\"_blank\">here</a>. Note that the <code>customer_ID</code> is a <code>int64</code> instead of the provided <code>string512</code>. It was compressed using the code shown <a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\" target=\"_blank\">here</a>.</p>\n<p>The test data is split in 20 NumPy arrays. Each array has <code>46231</code> customers. <strong>Important Note</strong>: If you concatenate all these files, the row order of the <code>924621</code> test customers is <strong>not the same</strong> as the file <code>sample_submission.csv</code>. The order is contained in my <code>submission.csv</code> <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns?select=submission.csv\" target=\"_blank\">here</a>. So infer the 20 NumPy arrays in their order and then overwrite my prediction column with <code>submission['prediction'] = preds</code>.</p>\n<h1>Transformers and RNNs</h1>\n<p>It will be exciting to see whether GBT (gradient boosted trees) or NN will achieve the best accuracy in this competition. I posted a RNN notebook starter <a href=\"https://tinyurl.com/4sjwms6v\" target=\"_blank\">here</a> and discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327761\" target=\"_blank\">here</a>. If you want to build a Transformer, you can use my Transformer notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-112\" target=\"_blank\">here</a> from Ventilator competition and input Amex competition data. Additionally there are many great RNN notebooks in Ventilator competition for inspiration.</p>\n<h1>UPDATE: Transformer Starter - LB 0.790</h1>\n<p>I created and published a Transformer starter notebook <a href=\"https://tinyurl.com/45uu256r\" target=\"_blank\">here</a> today that uses my Kaggle dataset. It achieves 5-Fold CV 0.787 and LB 0.789. Enjoy!</p>\n<h1>Have Fun!</h1>",
      "rawMarkdown": "# Competition Data CSV \nThis competition provides data about credit card customers in the form of `CSV` files. We cannot use these files without modification to train Transformers and RNNs because each customer has a different number of credit card statements. \n\nWhen inputting time series data into a Transformer and RNN, each customer needs the same sequence length (of 13 in this comp).\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/trans.png)\n\n# Competition Data NumPy Array\nI have converted the competition data into `NumPy array` files where each customer has sequence length 13 and posted a Kaggle dataset [here][2] (This data is created in my notebook [here][1])\n\nFor each customer with less than 13 credit card statements, I have padded their data with `pad = -1`. See example below where the customer has been padded with four -1's:\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/data_block.png)\n\n# Kaggle Dataset Details\nMy Kaggle dataset is [here][2]. The train data has been split into 10 NumPy arrays named `data_1.npy` thru `data_10.npy`. Each array has dimension `(45891, 13, 188)` which is `customer x statement x feature`. \n\nThe associated targets are contained in the files `targets_1.pqt` thru `targets_10.pqt`. These are parquet files with columns `customer_ID` and `target` and have dimension `(45891, 2)`. An example training with these files is [here][1]. Note that the `customer_ID` is a `int64` instead of the provided `string512`. It was compressed using the code shown [here][3].\n\nThe test data is split in 20 NumPy arrays. Each array has `46231` customers. **Important Note**: If you concatenate all these files, the row order of the `924621` test customers is **not the same** as the file `sample_submission.csv`. The order is contained in my `submission.csv` [here][4]. So infer the 20 NumPy arrays in their order and then overwrite my prediction column with `submission['prediction'] = preds`.\n\n# Transformers and RNNs\nIt will be exciting to see whether GBT (gradient boosted trees) or NN will achieve the best accuracy in this competition. I posted a RNN notebook starter [here][1] and discussion [here][6]. If you want to build a Transformer, you can use my Transformer notebook [here][5] from Ventilator competition and input Amex competition data. Additionally there are many great RNN notebooks in Ventilator competition for inspiration.\n\n# UPDATE: Transformer Starter - LB 0.790\nI created and published a Transformer starter notebook [here][7] today that uses my Kaggle dataset. It achieves 5-Fold CV 0.787 and LB 0.789. Enjoy!\n\n# Have Fun!\n\n[1]: https://tinyurl.com/4sjwms6v\n[2]: https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\n[3]: https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\n[4]: https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns?select=submission.csv\n[5]: https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-112\n[6]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327761\n[7]: https://tinyurl.com/45uu256r",
      "votes": 159
    },
    {
      "id": 1852080,
      "postDate": "2022-07-11T18:32:46.627Z",
      "content": "<p>Really nice work. Thanks!</p>",
      "rawMarkdown": "Really nice work. Thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 1873783,
          "postDate": "2022-07-27T22:42:49.697Z",
          "content": "<p>Thanks Jerry</p>",
          "rawMarkdown": "Thanks Jerry"
        }
      ]
    },
    {
      "id": 1846843,
      "postDate": "2022-07-07T11:49:03.280Z",
      "content": "<p>Its awsome</p>",
      "rawMarkdown": "Its awsome",
      "votes": 1
    },
    {
      "id": 1844857,
      "postDate": "2022-07-05T21:23:33.127Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thanks so much. Your work is so clear. </p>",
      "rawMarkdown": "Hi @cdeotte , thanks so much. Your work is so clear. ",
      "votes": 1,
      "replies": [
        {
          "id": 1844873,
          "postDate": "2022-07-05T21:52:33.670Z",
          "content": "<p>Thanks Kevin</p>",
          "rawMarkdown": "Thanks Kevin",
          "votes": 1
        }
      ]
    },
    {
      "id": 1813795,
      "postDate": "2022-06-07T08:46:09.917Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thanks for the great kernel.<br>\nUnfortunately, I didn't see the full code so far but I'm missing something for sure: how can you create the padded sequence for the customers in test data?<br>\nThey don't have any previous credit statement, so their sequence would be something like [-1, -1, -1…..-1, -1, statement] (basically, no history provided in the sequence).</p>\n<p>What am I missing?<br>\nMany thanks</p>",
      "rawMarkdown": "Hi @cdeotte , thanks for the great kernel.\nUnfortunately, I didn't see the full code so far but I'm missing something for sure: how can you create the padded sequence for the customers in test data?\nThey don't have any previous credit statement, so their sequence would be something like [-1, -1, -1.....-1, -1, statement] (basically, no history provided in the sequence).\n\nWhat am I missing?\nMany thanks",
      "votes": 1,
      "replies": [
        {
          "id": 1813823,
          "postDate": "2022-06-07T09:11:33.753Z",
          "content": "<p>Most customers in test data have 13 statements. Perhaps you are using a Kaggle dataset that does not contain the full test data. For example, here is one customer in test data<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/june.png\" alt=\"\"></p>",
          "rawMarkdown": "Most customers in test data have 13 statements. Perhaps you are using a Kaggle dataset that does not contain the full test data. For example, here is one customer in test data\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/june.png)",
          "votes": 4
        }
      ]
    },
    {
      "id": 1809249,
      "postDate": "2022-06-02T14:44:15.893Z",
      "content": "<p>Lots of Learning!<br>\nThanks, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for the explanation!</p>",
      "rawMarkdown": "Lots of Learning!\nThanks, @cdeotte for the explanation!",
      "votes": 2
    },
    {
      "id": 1807205,
      "postDate": "2022-05-31T20:12:23.393Z",
      "content": "<p>UPDATE: I posted TensorFlow Transformer starter code using this dataset <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "UPDATE: I posted TensorFlow Transformer starter code using this dataset [here][1]\n\n[1]: https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790",
      "votes": 2
    },
    {
      "id": 1871637,
      "postDate": "2022-07-26T11:48:03.243Z",
      "content": "<p>Hi is it Time series classification problem ?<br>\nCan we submit csv with 1 and 0 only not by prediction in decimal like i'm thinking as classification problem or they are asking about **\"probability of target\" ** , that's why our submissions should be in probability of 0 or 1. ??? kindly support </p>",
      "rawMarkdown": "Hi is it Time series classification problem ?\nCan we submit csv with 1 and 0 only not by prediction in decimal like i'm thinking as classification problem or they are asking about **\"probability of target\" ** , that's why our submissions should be in probability of 0 or 1. ??? kindly support ",
      "replies": [
        {
          "id": 1872018,
          "postDate": "2022-07-26T15:44:33.953Z",
          "content": "<p>For each test customer we submit the probability of default which is a number between 0 and 1. For each test customer we have up to 13 observations in time (i.e. credit card statements), so we have a time series (of length 13) for each test customer and we must make 1 classification probability predicton per customer.</p>",
          "rawMarkdown": "For each test customer we submit the probability of default which is a number between 0 and 1. For each test customer we have up to 13 observations in time (i.e. credit card statements), so we have a time series (of length 13) for each test customer and we must make 1 classification probability predicton per customer.",
          "votes": 4
        },
        {
          "id": 1872613,
          "postDate": "2022-07-27T06:02:01.550Z",
          "content": "<p>Noted. great support thanks</p>",
          "rawMarkdown": "Noted. great support thanks",
          "votes": 1
        }
      ]
    },
    {
      "id": 1806068,
      "postDate": "2022-05-30T18:50:48.927Z",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for the explanation. I was looking to learn from the notebook that creates this data but the link was not provided. for example, the [here] doesn't have a link to the Notebook</p>\n<p>Below the Section of your post that reference the Link</p>\n<p><strong>Competition Data NumPy Array</strong><br>\nI have converted the competition data into NumPy array files where each customer has a sequence length 13. <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\" target=\"_blank\">https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns</a> (This data is created in my notebook [here][1])</p>",
      "rawMarkdown": "Hello, @cdeotte thanks for the explanation. I was looking to learn from the notebook that creates this data but the link was not provided. for example, the [here] doesn't have a link to the Notebook\n\nBelow the Section of your post that reference the Link\n\n**Competition Data NumPy Array**\nI have converted the competition data into NumPy array files where each customer has a sequence length 13. https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns (This data is created in my notebook [here][1])",
      "replies": [
        {
          "id": 1806071,
          "postDate": "2022-05-30T18:53:53.093Z",
          "content": "<p>Thanks. Kaggle has a new spam filter that prevents me from posting the URL because i already posted it in another thread. The link is <code>www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790</code></p>",
          "rawMarkdown": "Thanks. Kaggle has a new spam filter that prevents me from posting the URL because i already posted it in another thread. The link is `www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790`",
          "votes": 1
        },
        {
          "id": 1806073,
          "postDate": "2022-05-30T18:55:54.123Z",
          "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> This new Kaggle spam filter is very annoying. I cannot post links to cross reference my posts. Also i cannot edit any of my old posts because when i try to edit and save, i get the error message <code>You've been sharing this link too often. We prevent redundant posts to reduce spam.</code>. This occurs even if i do not change any URLs in the old posts.</p>",
          "rawMarkdown": "@addisonhoward This new Kaggle spam filter is very annoying. I cannot post links to cross reference my posts. Also i cannot edit any of my old posts because when i try to edit and save, i get the error message `You've been sharing this link too often. We prevent redundant posts to reduce spam.`. This occurs even if i do not change any URLs in the old posts.",
          "votes": 5
        },
        {
          "id": 1806099,
          "postDate": "2022-05-30T19:33:50.790Z",
          "content": "<p>Thanks for the link, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>; the new Kaggle spam filter sounds like a headache, especially if you want to create more connected posts. </p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I was trying to understand how to configure the dataset into a 3D version and created the *.npy files myself; I haven't been able to find this on any of your Notebooks. Did you make them locally? </p>",
          "rawMarkdown": "Thanks for the link, @cdeotte; the new Kaggle spam filter sounds like a headache, especially if you want to create more connected posts. \n\n@cdeotte I was trying to understand how to configure the dataset into a 3D version and created the *.npy files myself; I haven't been able to find this on any of your Notebooks. Did you make them locally? ",
          "votes": 1
        },
        {
          "id": 1806102,
          "postDate": "2022-05-30T19:36:29.410Z",
          "content": "<p>My Bad, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I should have read the code deeper is all in the link provided, thanks </p>",
          "rawMarkdown": "My Bad, @cdeotte I should have read the code deeper is all in the link provided, thanks ",
          "votes": 1
        },
        {
          "id": 1806138,
          "postDate": "2022-05-30T20:19:28.413Z",
          "content": "<p>No worries. The code cell that saves the NumPy files is hidden. It is code cell 6 (in the provided link above). And the dataframe processing occurs in code cell 5.</p>\n<p>Note that my NumPy arrays are front padded. Whenever a customer has less than 13 statements, the code pads in the beginning. In actuality, every train customer has a last statement in March 2018 and the missing statements can be anywhere in the previous 12 months. So code for inplace padding is posted below. So far the GRU model with front padding achieves the same CV as inplace padding. (And test data code for inplace padding needs a few changes from below)</p>\n<pre><code># PAD SO EACH CUSTOMER HAS 13 ROWS    \ntrain = train.sort_values(['customer_ID','year','month','day'],\n                         ascending=[True, False, False, False])\ntrain['i'] = train.groupby('customer_ID').year.agg('cumcount')\nfinal_statement = train.loc[train.i == 0].drop(['i'],axis=1)\ntrain = train.loc[train.i != 0].drop(['i'],axis=1)\n# PAD WITHIN FIRST 12 STATEMENTS\ncust = final_statement.customer_ID.unique()\nmonths = cupy.tile( cupy.arange(12)+1,len(cust) )\ncust = cupy.repeat( cust,12 )\nnew_train = cudf.DataFrame({'customer_ID':cust,'month':months})    \nnew_train = new_train.merge(train, on=['customer_ID','month'], how='left')\nnew_train = new_train.fillna(-1)\nnew_train[CATS] = cupy.clip( new_train[CATS].values, a_min=0)\ntime_map = {3:17,4:17,5:17,6:17,7:17,8:17,9:17,10:17,11:17,12:17,1:18,2:18}\nnew_train['year'] = new_train.month.map(time_map)\n# COMBINE FIRST 12 STATEMENTS WITH FINAL STATEMENT FOR 13\ntrain = cudf.concat([new_train,final_statement],axis=0,ignore_index=True)\ndel new_train, cust, months, final_statement\n</code></pre>",
          "rawMarkdown": "No worries. The code cell that saves the NumPy files is hidden. It is code cell 6 (in the provided link above). And the dataframe processing occurs in code cell 5.\n\nNote that my NumPy arrays are front padded. Whenever a customer has less than 13 statements, the code pads in the beginning. In actuality, every train customer has a last statement in March 2018 and the missing statements can be anywhere in the previous 12 months. So code for inplace padding is posted below. So far the GRU model with front padding achieves the same CV as inplace padding. (And test data code for inplace padding needs a few changes from below)\n\n    # PAD SO EACH CUSTOMER HAS 13 ROWS    \n    train = train.sort_values(['customer_ID','year','month','day'],\n                             ascending=[True, False, False, False])\n    train['i'] = train.groupby('customer_ID').year.agg('cumcount')\n    final_statement = train.loc[train.i == 0].drop(['i'],axis=1)\n    train = train.loc[train.i != 0].drop(['i'],axis=1)\n    # PAD WITHIN FIRST 12 STATEMENTS\n    cust = final_statement.customer_ID.unique()\n    months = cupy.tile( cupy.arange(12)+1,len(cust) )\n    cust = cupy.repeat( cust,12 )\n    new_train = cudf.DataFrame({'customer_ID':cust,'month':months})    \n    new_train = new_train.merge(train, on=['customer_ID','month'], how='left')\n    new_train = new_train.fillna(-1)\n    new_train[CATS] = cupy.clip( new_train[CATS].values, a_min=0)\n    time_map = {3:17,4:17,5:17,6:17,7:17,8:17,9:17,10:17,11:17,12:17,1:18,2:18}\n    new_train['year'] = new_train.month.map(time_map)\n    # COMBINE FIRST 12 STATEMENTS WITH FINAL STATEMENT FOR 13\n    train = cudf.concat([new_train,final_statement],axis=0,ignore_index=True)\n    del new_train, cust, months, final_statement",
          "votes": 3
        },
        {
          "id": 1806199,
          "postDate": "2022-05-30T21:47:41.840Z",
          "content": "<p>Thanks for the detailed explanation, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, and the code snippet; I will experiment with this approach. I believe time base feature engineering and model will have some advantages.</p>",
          "rawMarkdown": "Thanks for the detailed explanation, @cdeotte, and the code snippet; I will experiment with this approach. I believe time base feature engineering and model will have some advantages.",
          "votes": 1
        },
        {
          "id": 1807033,
          "postDate": "2022-05-31T16:49:51.497Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  - I've notified our engineering team who developed this filter to provide some more information. </p>",
          "rawMarkdown": "Hey @cdeotte  - I've notified our engineering team who developed this filter to provide some more information. ",
          "votes": 2
        },
        {
          "id": 1808313,
          "postDate": "2022-06-01T17:50:01.847Z",
          "content": "<p>Hi Chris and others, the spam filter has been helpful in preventing rampant link spam by some users to unfairly gain medals. That said we hear you and recognize there are legitimate reasons to share resources that are good for others. We are working on tweaking the filter now. Thank you for your feedback!</p>",
          "rawMarkdown": "Hi Chris and others, the spam filter has been helpful in preventing rampant link spam by some users to unfairly gain medals. That said we hear you and recognize there are legitimate reasons to share resources that are good for others. We are working on tweaking the filter now. Thank you for your feedback!",
          "votes": 4
        }
      ]
    },
    {
      "id": 1881182,
      "postDate": "2022-08-02T10:25:01.143Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1816043,
      "postDate": "2022-06-09T18:37:08.890Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1847179,
      "postDate": "2022-07-07T17:30:25.177Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks @cdeotte ",
      "votes": 1
    },
    {
      "id": 1816246,
      "postDate": "2022-06-10T02:49:52.640Z",
      "content": "<p>Nice work ! Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Nice work ! Thanks @cdeotte ",
      "votes": 1
    },
    {
      "id": 1809386,
      "postDate": "2022-06-02T16:44:44.063Z",
      "content": "<p>Wow!<br>\nThank you so much!</p>",
      "rawMarkdown": "Wow!\nThank you so much!",
      "votes": 1
    },
    {
      "id": 1809158,
      "postDate": "2022-06-02T13:03:01.390Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": 1
    },
    {
      "id": 2663852,
      "postDate": "2024-02-22T17:46:12.093Z",
      "content": "<p>thanks's man</p>",
      "rawMarkdown": "thanks's man"
    }
  ],
  "comments": [
    {
      "id": 1852080,
      "author_name": "Jerry",
      "author_url": "",
      "post_date": "2022-07-11T18:32:46.627000",
      "content": "<p>Really nice work. Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1873783,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-27T22:42:49.697000",
          "content": "<p>Thanks Jerry</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1846843,
      "author_name": "Ammara Jabbar",
      "author_url": "",
      "post_date": "2022-07-07T11:49:03.280000",
      "content": "<p>Its awsome</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1844857,
      "author_name": "Kevin Meates",
      "author_url": "",
      "post_date": "2022-07-05T21:23:33.127000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thanks so much. Your work is so clear. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1844873,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-05T21:52:33.670000",
          "content": "<p>Thanks Kevin</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1813795,
      "author_name": "Luca Sala",
      "author_url": "",
      "post_date": "2022-06-07T08:46:09.917000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thanks for the great kernel.<br>\nUnfortunately, I didn't see the full code so far but I'm missing something for sure: how can you create the padded sequence for the customers in test data?<br>\nThey don't have any previous credit statement, so their sequence would be something like [-1, -1, -1…..-1, -1, statement] (basically, no history provided in the sequence).</p>\n<p>What am I missing?<br>\nMany thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1813823,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-07T09:11:33.753000",
          "content": "<p>Most customers in test data have 13 statements. Perhaps you are using a Kaggle dataset that does not contain the full test data. For example, here is one customer in test data<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Jun-2022/june.png\" alt=\"\"></p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1809249,
      "author_name": "Pawan Sharma",
      "author_url": "",
      "post_date": "2022-06-02T14:44:15.893000",
      "content": "<p>Lots of Learning!<br>\nThanks, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for the explanation!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1807205,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-05-31T20:12:23.393000",
      "content": "<p>UPDATE: I posted TensorFlow Transformer starter code using this dataset <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">here</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1871637,
      "author_name": "Chirag Valani",
      "author_url": "",
      "post_date": "2022-07-26T11:48:03.243000",
      "content": "<p>Hi is it Time series classification problem ?<br>\nCan we submit csv with 1 and 0 only not by prediction in decimal like i'm thinking as classification problem or they are asking about **\"probability of target\" ** , that's why our submissions should be in probability of 0 or 1. ??? kindly support </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1872018,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-07-26T15:44:33.953000",
          "content": "<p>For each test customer we submit the probability of default which is a number between 0 and 1. For each test customer we have up to 13 observations in time (i.e. credit card statements), so we have a time series (of length 13) for each test customer and we must make 1 classification probability predicton per customer.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1872613,
          "author_name": "Chirag Valani",
          "author_url": "",
          "post_date": "2022-07-27T06:02:01.550000",
          "content": "<p>Noted. great support thanks</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1806068,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2022-05-30T18:50:48.927000",
      "content": "<p>Hello, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks for the explanation. I was looking to learn from the notebook that creates this data but the link was not provided. for example, the [here] doesn't have a link to the Notebook</p>\n<p>Below the Section of your post that reference the Link</p>\n<p><strong>Competition Data NumPy Array</strong><br>\nI have converted the competition data into NumPy array files where each customer has a sequence length 13. <a href=\"https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\" target=\"_blank\">https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns</a> (This data is created in my notebook [here][1])</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1806071,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-30T18:53:53.093000",
          "content": "<p>Thanks. Kaggle has a new spam filter that prevents me from posting the URL because i already posted it in another thread. The link is <code>www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1806073,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-30T18:55:54.123000",
          "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> This new Kaggle spam filter is very annoying. I cannot post links to cross reference my posts. Also i cannot edit any of my old posts because when i try to edit and save, i get the error message <code>You've been sharing this link too often. We prevent redundant posts to reduce spam.</code>. This occurs even if i do not change any URLs in the old posts.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1806099,
          "author_name": "C4rl05/V",
          "author_url": "",
          "post_date": "2022-05-30T19:33:50.790000",
          "content": "<p>Thanks for the link, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>; the new Kaggle spam filter sounds like a headache, especially if you want to create more connected posts. </p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I was trying to understand how to configure the dataset into a 3D version and created the *.npy files myself; I haven't been able to find this on any of your Notebooks. Did you make them locally? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1806102,
          "author_name": "C4rl05/V",
          "author_url": "",
          "post_date": "2022-05-30T19:36:29.410000",
          "content": "<p>My Bad, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I should have read the code deeper is all in the link provided, thanks </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1806138,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-30T20:19:28.413000",
          "content": "<p>No worries. The code cell that saves the NumPy files is hidden. It is code cell 6 (in the provided link above). And the dataframe processing occurs in code cell 5.</p>\n<p>Note that my NumPy arrays are front padded. Whenever a customer has less than 13 statements, the code pads in the beginning. In actuality, every train customer has a last statement in March 2018 and the missing statements can be anywhere in the previous 12 months. So code for inplace padding is posted below. So far the GRU model with front padding achieves the same CV as inplace padding. (And test data code for inplace padding needs a few changes from below)</p>\n<pre><code># PAD SO EACH CUSTOMER HAS 13 ROWS    \ntrain = train.sort_values(['customer_ID','year','month','day'],\n                         ascending=[True, False, False, False])\ntrain['i'] = train.groupby('customer_ID').year.agg('cumcount')\nfinal_statement = train.loc[train.i == 0].drop(['i'],axis=1)\ntrain = train.loc[train.i != 0].drop(['i'],axis=1)\n# PAD WITHIN FIRST 12 STATEMENTS\ncust = final_statement.customer_ID.unique()\nmonths = cupy.tile( cupy.arange(12)+1,len(cust) )\ncust = cupy.repeat( cust,12 )\nnew_train = cudf.DataFrame({'customer_ID':cust,'month':months})    \nnew_train = new_train.merge(train, on=['customer_ID','month'], how='left')\nnew_train = new_train.fillna(-1)\nnew_train[CATS] = cupy.clip( new_train[CATS].values, a_min=0)\ntime_map = {3:17,4:17,5:17,6:17,7:17,8:17,9:17,10:17,11:17,12:17,1:18,2:18}\nnew_train['year'] = new_train.month.map(time_map)\n# COMBINE FIRST 12 STATEMENTS WITH FINAL STATEMENT FOR 13\ntrain = cudf.concat([new_train,final_statement],axis=0,ignore_index=True)\ndel new_train, cust, months, final_statement\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1806199,
          "author_name": "C4rl05/V",
          "author_url": "",
          "post_date": "2022-05-30T21:47:41.840000",
          "content": "<p>Thanks for the detailed explanation, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, and the code snippet; I will experiment with this approach. I believe time base feature engineering and model will have some advantages.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1807033,
          "author_name": "Addison Howard",
          "author_url": "",
          "post_date": "2022-05-31T16:49:51.497000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>  - I've notified our engineering team who developed this filter to provide some more information. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1808313,
          "author_name": "Jessica Li",
          "author_url": "",
          "post_date": "2022-06-01T17:50:01.847000",
          "content": "<p>Hi Chris and others, the spam filter has been helpful in preventing rampant link spam by some users to unfairly gain medals. That said we hear you and recognize there are legitimate reasons to share resources that are good for others. We are working on tweaking the filter now. Thank you for your feedback!</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1881182,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-02T10:25:01.143000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1816043,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-06-09T18:37:08.890000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1847179,
      "author_name": "Pankaj Kumar",
      "author_url": "",
      "post_date": "2022-07-07T17:30:25.177000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1816246,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2022-06-10T02:49:52.640000",
      "content": "<p>Nice work ! Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809386,
      "author_name": "Agung Aldevando",
      "author_url": "",
      "post_date": "2022-06-02T16:44:44.063000",
      "content": "<p>Wow!<br>\nThank you so much!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1809158,
      "author_name": "Selbinyyaz Sultanova",
      "author_url": "",
      "post_date": "2022-06-02T13:03:01.390000",
      "content": "<p>Thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2663852,
      "author_name": "Sadik Sikder",
      "author_url": "",
      "post_date": "2024-02-22T17:46:12.093000",
      "content": "<p>thanks's man</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1804818": "# Competition Data CSV \nThis competition provides data about credit card customers in the form of `CSV` files. We cannot use these files without modification to train Transformers and RNNs because each customer has a different number of credit card statements. \n\nWhen inputting time series data into a Transformer and RNN, each customer needs the same sequence length (of 13 in this comp).\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/trans.png)\n\n# Competition Data NumPy Array\nI have converted the competition data into `NumPy array` files where each customer has sequence length 13 and posted a Kaggle dataset [here][2] (This data is created in my notebook [here][1])\n\nFor each customer with less than 13 credit card statements, I have padded their data with `pad = -1`. See example below where the customer has been padded with four -1's:\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/data_block.png)\n\n# Kaggle Dataset Details\nMy Kaggle dataset is [here][2]. The train data has been split into 10 NumPy arrays named `data_1.npy` thru `data_10.npy`. Each array has dimension `(45891, 13, 188)` which is `customer x statement x feature`. \n\nThe associated targets are contained in the files `targets_1.pqt` thru `targets_10.pqt`. These are parquet files with columns `customer_ID` and `target` and have dimension `(45891, 2)`. An example training with these files is [here][1]. Note that the `customer_ID` is a `int64` instead of the provided `string512`. It was compressed using the code shown [here][3].\n\nThe test data is split in 20 NumPy arrays. Each array has `46231` customers. **Important Note**: If you concatenate all these files, the row order of the `924621` test customers is **not the same** as the file `sample_submission.csv`. The order is contained in my `submission.csv` [here][4]. So infer the 20 NumPy arrays in their order and then overwrite my prediction column with `submission['prediction'] = preds`.\n\n# Transformers and RNNs\nIt will be exciting to see whether GBT (gradient boosted trees) or NN will achieve the best accuracy in this competition. I posted a RNN notebook starter [here][1] and discussion [here][6]. If you want to build a Transformer, you can use my Transformer notebook [here][5] from Ventilator competition and input Amex competition data. Additionally there are many great RNN notebooks in Ventilator competition for inspiration.\n\n# UPDATE: Transformer Starter - LB 0.790\nI created and published a Transformer starter notebook [here][7] today that uses my Kaggle dataset. It achieves 5-Fold CV 0.787 and LB 0.789. Enjoy!\n\n# Have Fun!\n\n[1]: https://tinyurl.com/4sjwms6v\n[2]: https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns\n[3]: https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/308635\n[4]: https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns?select=submission.csv\n[5]: https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-112\n[6]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327761\n[7]: https://tinyurl.com/45uu256r",
    "1852080": "Really nice work. Thanks!",
    "1846843": "Its awsome",
    "1844857": "Hi @cdeotte , thanks so much. Your work is so clear. ",
    "1813795": "Hi @cdeotte , thanks for the great kernel.\nUnfortunately, I didn't see the full code so far but I'm missing something for sure: how can you create the padded sequence for the customers in test data?\nThey don't have any previous credit statement, so their sequence would be something like [-1, -1, -1.....-1, -1, statement] (basically, no history provided in the sequence).\n\nWhat am I missing?\nMany thanks",
    "1809249": "Lots of Learning!\nThanks, @cdeotte for the explanation!",
    "1807205": "UPDATE: I posted TensorFlow Transformer starter code using this dataset [here][1]\n\n[1]: https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790",
    "1871637": "Hi is it Time series classification problem ?\nCan we submit csv with 1 and 0 only not by prediction in decimal like i'm thinking as classification problem or they are asking about **\"probability of target\" ** , that's why our submissions should be in probability of 0 or 1. ??? kindly support ",
    "1806068": "Hello, @cdeotte thanks for the explanation. I was looking to learn from the notebook that creates this data but the link was not provided. for example, the [here] doesn't have a link to the Notebook\n\nBelow the Section of your post that reference the Link\n\n**Competition Data NumPy Array**\nI have converted the competition data into NumPy array files where each customer has a sequence length 13. https://www.kaggle.com/datasets/cdeotte/amex-data-for-transformers-and-rnns (This data is created in my notebook [here][1])",
    "1881182": "",
    "1816043": "",
    "1847179": "Thanks @cdeotte ",
    "1816246": "Nice work ! Thanks @cdeotte ",
    "1809386": "Wow!\nThank you so much!",
    "1809158": "Thank you!",
    "2663852": "thanks's man"
  }
}