{
  "id": 327761,
  "title": "Time Series EDA and GRU Starter - LB 0.790",
  "url": "/competitions/amex-default-prediction/discussion/327761",
  "author_name": "Chris Deotte",
  "post_date": "2022-05-29T02:41:39.714000",
  "votes": 85,
  "comment_count": 14,
  "views": 0,
  "content": "<h1>TensorFlow GRU Starter - LB 0.790</h1>\n<p>I published an RNN starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a> which achieves 5-Fold CV 0.787 and LB 0.789. This model uses the provided Kaggle train data as is without new features and without any scaling and without data augmentation. It provides a simple RNN baseline.</p>\n<p>An RNN uses all of the provided data which includes each customers' 13 credit card statements. It has more information than just using the last statement and it can detect patterns that occur over time.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/rnn.png\" alt=\"\"></p>\n<h1>Time Series EDA</h1>\n<p>I published time series plots <a href=\"https://www.kaggle.com/code/cdeotte/time-series-eda\" target=\"_blank\">here</a> of some variables to show us how variables change over time. In the example below, we see that variable <code>D_102</code> generally increases over time.</p>\n<p>Each line is one customer. Blue lines are customers with <code>target=1</code> (defaulters) and orange lines are customers with <code>target=0</code> (non-dafaulters). In addition to the line plots, the histogram on the right shows that both customers have low <code>D_102</code> values, but the larger values of <code>D_102</code> generally belong to non-defaulters.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/t_plot.png\" alt=\"\"></p>\n<h1>Kaggle Dataset for Transformers and RNNs</h1>\n<p>I published a Kaggle dataset with competition data in 3D format for training transformers and RNNs <a href=\"https://tinyurl.com/mrxm4hsp\" target=\"_blank\">here</a> and discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\" target=\"_blank\">here</a>. These are the NumPy arrays created from the GRU starter notebook. Training transformers and RNNs requires 3D array data instead of 2D CSV data.</p>\n<h1>UPDATE: Transformer Starter - LB 0.790</h1>\n<p>I created and published a Transformer starter notebook <a href=\"https://tinyurl.com/yc68vzk9\" target=\"_blank\">here</a> today. It also achieves 5-Fold CV 0.787 and LB 0.789. Enjoy!</p>",
  "messages": [
    {
      "id": 1804459,
      "postDate": "2022-05-29T02:41:39.713Z",
      "content": "<h1>TensorFlow GRU Starter - LB 0.790</h1>\n<p>I published an RNN starter notebook <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\" target=\"_blank\">here</a> which achieves 5-Fold CV 0.787 and LB 0.789. This model uses the provided Kaggle train data as is without new features and without any scaling and without data augmentation. It provides a simple RNN baseline.</p>\n<p>An RNN uses all of the provided data which includes each customers' 13 credit card statements. It has more information than just using the last statement and it can detect patterns that occur over time.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/rnn.png\" alt=\"\"></p>\n<h1>Time Series EDA</h1>\n<p>I published time series plots <a href=\"https://www.kaggle.com/code/cdeotte/time-series-eda\" target=\"_blank\">here</a> of some variables to show us how variables change over time. In the example below, we see that variable <code>D_102</code> generally increases over time.</p>\n<p>Each line is one customer. Blue lines are customers with <code>target=1</code> (defaulters) and orange lines are customers with <code>target=0</code> (non-dafaulters). In addition to the line plots, the histogram on the right shows that both customers have low <code>D_102</code> values, but the larger values of <code>D_102</code> generally belong to non-defaulters.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/t_plot.png\" alt=\"\"></p>\n<h1>Kaggle Dataset for Transformers and RNNs</h1>\n<p>I published a Kaggle dataset with competition data in 3D format for training transformers and RNNs <a href=\"https://tinyurl.com/mrxm4hsp\" target=\"_blank\">here</a> and discussion <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\" target=\"_blank\">here</a>. These are the NumPy arrays created from the GRU starter notebook. Training transformers and RNNs requires 3D array data instead of 2D CSV data.</p>\n<h1>UPDATE: Transformer Starter - LB 0.790</h1>\n<p>I created and published a Transformer starter notebook <a href=\"https://tinyurl.com/yc68vzk9\" target=\"_blank\">here</a> today. It also achieves 5-Fold CV 0.787 and LB 0.789. Enjoy!</p>",
      "rawMarkdown": "# TensorFlow GRU Starter - LB 0.790\nI published an RNN starter notebook [here][1] which achieves 5-Fold CV 0.787 and LB 0.789. This model uses the provided Kaggle train data as is without new features and without any scaling and without data augmentation. It provides a simple RNN baseline.\n\nAn RNN uses all of the provided data which includes each customers' 13 credit card statements. It has more information than just using the last statement and it can detect patterns that occur over time.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/rnn.png)\n\n# Time Series EDA\nI published time series plots [here][2] of some variables to show us how variables change over time. In the example below, we see that variable `D_102` generally increases over time.\n\nEach line is one customer. Blue lines are customers with `target=1` (defaulters) and orange lines are customers with `target=0` (non-dafaulters). In addition to the line plots, the histogram on the right shows that both customers have low `D_102` values, but the larger values of `D_102` generally belong to non-defaulters.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/t_plot.png)\n\n# Kaggle Dataset for Transformers and RNNs\nI published a Kaggle dataset with competition data in 3D format for training transformers and RNNs [here][3] and discussion [here][4]. These are the NumPy arrays created from the GRU starter notebook. Training transformers and RNNs requires 3D array data instead of 2D CSV data.\n\n# UPDATE: Transformer Starter - LB 0.790\nI created and published a Transformer starter notebook [here][5] today. It also achieves 5-Fold CV 0.787 and LB 0.789. Enjoy!\n\n[1]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\n[2]: https://www.kaggle.com/code/cdeotte/time-series-eda\n[3]: https://tinyurl.com/mrxm4hsp\n[4]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\n[5]: https://tinyurl.com/yc68vzk9",
      "votes": 85
    },
    {
      "id": 1808964,
      "postDate": "2022-06-02T09:27:02.267Z",
      "content": "<p>Hi Chris thanks for sharing this.</p>\n<p>My question is why did you prefer -0.5 for numerical imputing?</p>",
      "rawMarkdown": "Hi Chris thanks for sharing this.\n\nMy question is why did you prefer -0.5 for numerical imputing?",
      "votes": 1,
      "replies": [
        {
          "id": 1809143,
          "postDate": "2022-06-02T12:46:57.110Z",
          "content": "<p>My intuition is the following. The NAN information in the dataset contains signal. If you convert the entire dataset with <code>df[df.notna()] = 0</code> and <code>df[df.isna()] = 1</code>, this will achieve CV AUC 0.55. So there is information hiding in the NAN values.</p>\n<p>I choose <code>df[NUMERIC] = df[NUMERIC].fillna(-0.5)</code> so that the model can still recognize the NAN. If we use <code>fillna(0)</code> then the NANs become camouflage and the model cannot tell the difference between what was originally 0 and what NAN was converted to 0.</p>\n<p>The numeric data is basically in the range <code>0 to 1</code>. So I use <code>-1</code> for padding and <code>-0.5</code> for NAN. Then the model can recognize both padding and NAN. Note that NN like data with mean 0 and std 1. So the choices -1 and -0.5 help shift the dataset's mean and std toward this (without the need for standard scaler). </p>\n<p>For categorical variables, i use <code>0</code> for padding and <code>1</code> for NAN. Then I shift all the label encoding to <code>2, 3, 4, etc, etc</code>.</p>\n<p>Of course the best way to know if what i did helps is to try other <code>fillna()</code> and compute the CV LB score. Perhaps there are better strategies. I have not searched yet.</p>",
          "rawMarkdown": "My intuition is the following. The NAN information in the dataset contains signal. If you convert the entire dataset with `df[df.notna()] = 0` and `df[df.isna()] = 1`, this will achieve CV AUC 0.55. So there is information hiding in the NAN values.\n\nI choose `df[NUMERIC] = df[NUMERIC].fillna(-0.5)` so that the model can still recognize the NAN. If we use `fillna(0)` then the NANs become camouflage and the model cannot tell the difference between what was originally 0 and what NAN was converted to 0.\n\nThe numeric data is basically in the range `0 to 1`. So I use `-1` for padding and `-0.5` for NAN. Then the model can recognize both padding and NAN. Note that NN like data with mean 0 and std 1. So the choices -1 and -0.5 help shift the dataset's mean and std toward this (without the need for standard scaler). \n\nFor categorical variables, i use `0` for padding and `1` for NAN. Then I shift all the label encoding to `2, 3, 4, etc, etc`.\n\nOf course the best way to know if what i did helps is to try other `fillna()` and compute the CV LB score. Perhaps there are better strategies. I have not searched yet.",
          "votes": 10
        },
        {
          "id": 1809502,
          "postDate": "2022-06-02T18:08:25.347Z",
          "content": "<p>Thank you that was what i needed</p>",
          "rawMarkdown": "Thank you that was what i needed",
          "votes": 1
        }
      ]
    },
    {
      "id": 1806251,
      "postDate": "2022-05-31T00:39:51.547Z",
      "content": "<p>Thanks for your insight. I'll try it too :)</p>",
      "rawMarkdown": "Thanks for your insight. I'll try it too :)",
      "votes": 1
    },
    {
      "id": 1805536,
      "postDate": "2022-05-30T09:20:43.657Z",
      "content": "<p>Thanks for the wonderful contribution and explanation. My gratitude.</p>",
      "rawMarkdown": "Thanks for the wonderful contribution and explanation. My gratitude.",
      "votes": 1
    },
    {
      "id": 1804573,
      "postDate": "2022-05-29T07:16:09.637Z",
      "content": "<p>Very nice TS EDA and baseline model. Score isn’t really higher than baseline gbdt though. Do you feel it has a lot of room for improvement ?</p>",
      "rawMarkdown": "Very nice TS EDA and baseline model. Score isn’t really higher than baseline gbdt though. Do you feel it has a lot of room for improvement ?",
      "votes": 1,
      "replies": [
        {
          "id": 1804760,
          "postDate": "2022-05-29T12:26:53.227Z",
          "content": "<p>Note that the GBT starters and MLP starters only achieved LB 0.793. So already the GRU starter is beating GBT and MLP with LB 0.789. </p>\n<p>The newest public versions of GBT and MLP used aggregated features describing the mean of all 13 statements in addition to using the last statement feature. So the current GBT and MLP public notebooks have added more features. </p>\n<p>My GRU starter uses no additional features besides what Kaggle has provided. If we improve the architecture and/or create new features and/or use data augmentation, it can beat the current public GBT and MLP. </p>\n<p>(Note that we can also input my processed data NumPy files into a Transformer instead of GRU too. And tune the transformer to beat the GRU).</p>",
          "rawMarkdown": "Note that the GBT starters and MLP starters only achieved LB 0.793. So already the GRU starter is beating GBT and MLP with LB 0.789. \n\nThe newest public versions of GBT and MLP used aggregated features describing the mean of all 13 statements in addition to using the last statement feature. So the current GBT and MLP public notebooks have added more features. \n\nMy GRU starter uses no additional features besides what Kaggle has provided. If we improve the architecture and/or create new features and/or use data augmentation, it can beat the current public GBT and MLP. \n\n(Note that we can also input my processed data NumPy files into a Transformer instead of GRU too. And tune the transformer to beat the GRU).",
          "votes": 3
        },
        {
          "id": 1804794,
          "postDate": "2022-05-29T13:13:52.993Z",
          "content": "<p>Ok thanks. I'd argue that the GRU get 13 times more features than a gbt that use 'last' only, but I see your point. Given that we only have 'short' time series (thus limiting the feature engineering that needs to be done to grasp the full signal) I was thinking that gbt could stay competitive. </p>",
          "rawMarkdown": "Ok thanks. I'd argue that the GRU get 13 times more features than a gbt that use 'last' only, but I see your point. Given that we only have 'short' time series (thus limiting the feature engineering that needs to be done to grasp the full signal) I was thinking that gbt could stay competitive. ",
          "votes": 1
        },
        {
          "id": 1804802,
          "postDate": "2022-05-29T13:22:41.813Z",
          "content": "<p>Yeah. It's too early to know where the secret signal is hiding. As the months move forward we will discover the magic features. If the magic features involve trends in time (like quick moves up and down for example) then RNN will capture them. However maybe the magic feature is just the 3rd decimal position of the difference of two columns. In which case GBT with feature engineering will find them. We'll have to wait and see where the secret is hiding :-)</p>",
          "rawMarkdown": "Yeah. It's too early to know where the secret signal is hiding. As the months move forward we will discover the magic features. If the magic features involve trends in time (like quick moves up and down for example) then RNN will capture them. However maybe the magic feature is just the 3rd decimal position of the difference of two columns. In which case GBT with feature engineering will find them. We'll have to wait and see where the secret is hiding :-)",
          "votes": 4
        },
        {
          "id": 1807086,
          "postDate": "2022-05-31T17:52:00.160Z",
          "content": "<p><a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> I did a quick experiment. In my GRU, i set all data in the first 12 statements equal to zero for all customers. This way the GRU only has access to the last statement. When i do this, the CV dropped by <code>-0.006</code>. So I suspect the LB would be around <code>0.783</code>. </p>\n<p>This demonstates that the GRU is indeed using information from the first 12 statements. And by looking at public notebooks, i see that the LGBM using last only score LB 0.783 and LGBM using last plus aggregations of first 12 statements achieves LB 0.790. So there is information hiding in the first 12 statements.</p>\n<p>Currently it seems that GRU and LGBM (with agg features) are performing similarly. The best public notebook is CatBoost which creates new features from the category features under the hood. So perhaps engineering features using the category columns is the next step to boost CV LB for LGBM and GRU.</p>",
          "rawMarkdown": "@lucasmorin I did a quick experiment. In my GRU, i set all data in the first 12 statements equal to zero for all customers. This way the GRU only has access to the last statement. When i do this, the CV dropped by `-0.006`. So I suspect the LB would be around `0.783`. \n\nThis demonstates that the GRU is indeed using information from the first 12 statements. And by looking at public notebooks, i see that the LGBM using last only score LB 0.783 and LGBM using last plus aggregations of first 12 statements achieves LB 0.790. So there is information hiding in the first 12 statements.\n\nCurrently it seems that GRU and LGBM (with agg features) are performing similarly. The best public notebook is CatBoost which creates new features from the category features under the hood. So perhaps engineering features using the category columns is the next step to boost CV LB for LGBM and GRU.",
          "votes": 7
        }
      ]
    },
    {
      "id": 1804806,
      "postDate": "2022-05-29T13:24:38.993Z",
      "content": "<p>Awesome work! I also started with NN on 13*188 (2D inputs), in detail I applied GRU + Transformers for the raw data without any features engineering, but only got 0.781 (local 5 fold cv)….thanks for your sharing!</p>",
      "rawMarkdown": "Awesome work! I also started with NN on 13*188 (2D inputs), in detail I applied GRU + Transformers for the raw data without any features engineering, but only got 0.781 (local 5 fold cv)....thanks for your sharing!",
      "votes": 2
    },
    {
      "id": 1804561,
      "postDate": "2022-05-29T06:44:17.397Z",
      "content": "<p>Thanks for sharing Chris. You instantly make this competition more interesting. Models that can capture temporal information makes more sense than GBDT models with aggregations.</p>\n<p>By looking at your time series eda, I can say that default trends are less stable and continuous. Diff-like features could be useful here. What do you think about the upwards trend though?</p>",
      "rawMarkdown": "Thanks for sharing Chris. You instantly make this competition more interesting. Models that can capture temporal information makes more sense than GBDT models with aggregations.\n\nBy looking at your time series eda, I can say that default trends are less stable and continuous. Diff-like features could be useful here. What do you think about the upwards trend though?",
      "votes": 2,
      "replies": [
        {
          "id": 1804761,
          "postDate": "2022-05-29T12:29:55.477Z",
          "content": "<p>There are lots of great feature ideas from Kaggle's Ventilator competition. The best public GRU in that competition <a href=\"https://www.kaggle.com/code/dlaststark/gb-vpp-pulp-fiction\" target=\"_blank\">here</a> used difference features.</p>",
          "rawMarkdown": "There are lots of great feature ideas from Kaggle's Ventilator competition. The best public GRU in that competition [here][1] used difference features.\n\n[1]: https://www.kaggle.com/code/dlaststark/gb-vpp-pulp-fiction",
          "votes": 7
        }
      ]
    },
    {
      "id": 1870602,
      "postDate": "2022-07-25T17:06:47.107Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1808964,
      "author_name": "huseyincotel",
      "author_url": "",
      "post_date": "2022-06-02T09:27:02.267000",
      "content": "<p>Hi Chris thanks for sharing this.</p>\n<p>My question is why did you prefer -0.5 for numerical imputing?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1809143,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-06-02T12:46:57.110000",
          "content": "<p>My intuition is the following. The NAN information in the dataset contains signal. If you convert the entire dataset with <code>df[df.notna()] = 0</code> and <code>df[df.isna()] = 1</code>, this will achieve CV AUC 0.55. So there is information hiding in the NAN values.</p>\n<p>I choose <code>df[NUMERIC] = df[NUMERIC].fillna(-0.5)</code> so that the model can still recognize the NAN. If we use <code>fillna(0)</code> then the NANs become camouflage and the model cannot tell the difference between what was originally 0 and what NAN was converted to 0.</p>\n<p>The numeric data is basically in the range <code>0 to 1</code>. So I use <code>-1</code> for padding and <code>-0.5</code> for NAN. Then the model can recognize both padding and NAN. Note that NN like data with mean 0 and std 1. So the choices -1 and -0.5 help shift the dataset's mean and std toward this (without the need for standard scaler). </p>\n<p>For categorical variables, i use <code>0</code> for padding and <code>1</code> for NAN. Then I shift all the label encoding to <code>2, 3, 4, etc, etc</code>.</p>\n<p>Of course the best way to know if what i did helps is to try other <code>fillna()</code> and compute the CV LB score. Perhaps there are better strategies. I have not searched yet.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1809502,
          "author_name": "huseyincotel",
          "author_url": "",
          "post_date": "2022-06-02T18:08:25.347000",
          "content": "<p>Thank you that was what i needed</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1806251,
      "author_name": "Neukgu",
      "author_url": "",
      "post_date": "2022-05-31T00:39:51.547000",
      "content": "<p>Thanks for your insight. I'll try it too :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1805536,
      "author_name": "Aninda Goswamy",
      "author_url": "",
      "post_date": "2022-05-30T09:20:43.657000",
      "content": "<p>Thanks for the wonderful contribution and explanation. My gratitude.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1804573,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-05-29T07:16:09.637000",
      "content": "<p>Very nice TS EDA and baseline model. Score isn’t really higher than baseline gbdt though. Do you feel it has a lot of room for improvement ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1804760,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-29T12:26:53.227000",
          "content": "<p>Note that the GBT starters and MLP starters only achieved LB 0.793. So already the GRU starter is beating GBT and MLP with LB 0.789. </p>\n<p>The newest public versions of GBT and MLP used aggregated features describing the mean of all 13 statements in addition to using the last statement feature. So the current GBT and MLP public notebooks have added more features. </p>\n<p>My GRU starter uses no additional features besides what Kaggle has provided. If we improve the architecture and/or create new features and/or use data augmentation, it can beat the current public GBT and MLP. </p>\n<p>(Note that we can also input my processed data NumPy files into a Transformer instead of GRU too. And tune the transformer to beat the GRU).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1804794,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-05-29T13:13:52.993000",
          "content": "<p>Ok thanks. I'd argue that the GRU get 13 times more features than a gbt that use 'last' only, but I see your point. Given that we only have 'short' time series (thus limiting the feature engineering that needs to be done to grasp the full signal) I was thinking that gbt could stay competitive. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1804802,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-29T13:22:41.813000",
          "content": "<p>Yeah. It's too early to know where the secret signal is hiding. As the months move forward we will discover the magic features. If the magic features involve trends in time (like quick moves up and down for example) then RNN will capture them. However maybe the magic feature is just the 3rd decimal position of the difference of two columns. In which case GBT with feature engineering will find them. We'll have to wait and see where the secret is hiding :-)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1807086,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-31T17:52:00.160000",
          "content": "<p><a href=\"https://www.kaggle.com/lucasmorin\" target=\"_blank\">@lucasmorin</a> I did a quick experiment. In my GRU, i set all data in the first 12 statements equal to zero for all customers. This way the GRU only has access to the last statement. When i do this, the CV dropped by <code>-0.006</code>. So I suspect the LB would be around <code>0.783</code>. </p>\n<p>This demonstates that the GRU is indeed using information from the first 12 statements. And by looking at public notebooks, i see that the LGBM using last only score LB 0.783 and LGBM using last plus aggregations of first 12 statements achieves LB 0.790. So there is information hiding in the first 12 statements.</p>\n<p>Currently it seems that GRU and LGBM (with agg features) are performing similarly. The best public notebook is CatBoost which creates new features from the category features under the hood. So perhaps engineering features using the category columns is the next step to boost CV LB for LGBM and GRU.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1804806,
      "author_name": "Neo Zhao",
      "author_url": "",
      "post_date": "2022-05-29T13:24:38.993000",
      "content": "<p>Awesome work! I also started with NN on 13*188 (2D inputs), in detail I applied GRU + Transformers for the raw data without any features engineering, but only got 0.781 (local 5 fold cv)….thanks for your sharing!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1804561,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-05-29T06:44:17.397000",
      "content": "<p>Thanks for sharing Chris. You instantly make this competition more interesting. Models that can capture temporal information makes more sense than GBDT models with aggregations.</p>\n<p>By looking at your time series eda, I can say that default trends are less stable and continuous. Diff-like features could be useful here. What do you think about the upwards trend though?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1804761,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-05-29T12:29:55.477000",
          "content": "<p>There are lots of great feature ideas from Kaggle's Ventilator competition. The best public GRU in that competition <a href=\"https://www.kaggle.com/code/dlaststark/gb-vpp-pulp-fiction\" target=\"_blank\">here</a> used difference features.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 1870602,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-25T17:06:47.107000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1804459": "# TensorFlow GRU Starter - LB 0.790\nI published an RNN starter notebook [here][1] which achieves 5-Fold CV 0.787 and LB 0.789. This model uses the provided Kaggle train data as is without new features and without any scaling and without data augmentation. It provides a simple RNN baseline.\n\nAn RNN uses all of the provided data which includes each customers' 13 credit card statements. It has more information than just using the last statement and it can detect patterns that occur over time.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/rnn.png)\n\n# Time Series EDA\nI published time series plots [here][2] of some variables to show us how variables change over time. In the example below, we see that variable `D_102` generally increases over time.\n\nEach line is one customer. Blue lines are customers with `target=1` (defaulters) and orange lines are customers with `target=0` (non-dafaulters). In addition to the line plots, the histogram on the right shows that both customers have low `D_102` values, but the larger values of `D_102` generally belong to non-defaulters.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/May-2022/t_plot.png)\n\n# Kaggle Dataset for Transformers and RNNs\nI published a Kaggle dataset with competition data in 3D format for training transformers and RNNs [here][3] and discussion [here][4]. These are the NumPy arrays created from the GRU starter notebook. Training transformers and RNNs requires 3D array data instead of 2D CSV data.\n\n# UPDATE: Transformer Starter - LB 0.790\nI created and published a Transformer starter notebook [here][5] today. It also achieves 5-Fold CV 0.787 and LB 0.789. Enjoy!\n\n[1]: https://www.kaggle.com/code/cdeotte/tensorflow-gru-starter-0-790\n[2]: https://www.kaggle.com/code/cdeotte/time-series-eda\n[3]: https://tinyurl.com/mrxm4hsp\n[4]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/327828\n[5]: https://tinyurl.com/yc68vzk9",
    "1808964": "Hi Chris thanks for sharing this.\n\nMy question is why did you prefer -0.5 for numerical imputing?",
    "1806251": "Thanks for your insight. I'll try it too :)",
    "1805536": "Thanks for the wonderful contribution and explanation. My gratitude.",
    "1804573": "Very nice TS EDA and baseline model. Score isn’t really higher than baseline gbdt though. Do you feel it has a lot of room for improvement ?",
    "1804806": "Awesome work! I also started with NN on 13*188 (2D inputs), in detail I applied GRU + Transformers for the raw data without any features engineering, but only got 0.781 (local 5 fold cv)....thanks for your sharing!",
    "1804561": "Thanks for sharing Chris. You instantly make this competition more interesting. Models that can capture temporal information makes more sense than GBDT models with aggregations.\n\nBy looking at your time series eda, I can say that default trends are less stable and continuous. Diff-like features could be useful here. What do you think about the upwards trend though?",
    "1870602": ""
  }
}