{
  "id": 347668,
  "title": "10th Place Solution: XGB with Autoregressive RNN features",
  "url": "/competitions/amex-default-prediction/discussion/347668",
  "author_name": "Jiwei Liu",
  "post_date": "2022-08-25T04:12:09.095000",
  "votes": 93,
  "comment_count": 29,
  "views": 0,
  "content": "<p>What a competition! I really enjoyed it and only hope I could have found more time. First of all, I would like to thank Raddar <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, Martin <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>, and many others who generously shared codes and datasets! The public solutions are of amazing quality. And it also determines my game plan: <strong>create something original and blend it with the best public solution.</strong> My solution is based on RAPIDS cudf for dataframe processing, XGB for training, and pytorch lightning for feature extraction.</p>\n<p>There are two vital observations of this dataset:</p>\n<ul>\n<li>the big test data which is from the future is available.</li>\n<li>short sequences (&lt;13) are culprits of bad performance.</li>\n</ul>\n<p>Let's start with the latter observation:<strong>the sequence length of each customer profiles</strong> plays a critical role in the model performance:</p>\n<pre><code>import cudf\npath = '/raid/amex'\ntrain = cudf.read_parquet(f'{path}/train.parquet',columns=['customer_ID'])\ntrainl = cudf.read_csv(f'{path}/train_labels.csv')\ntrain = train.merge(trainl,on='customer_ID',how='left')\ntrain['seq_len'] = train.groupby('customer_ID')['target'].transform('count')\ntrain = train.drop_duplicates('customer_ID',keep='last')\ntrain.groupby('seq_len').agg({'target':['mean','count']}).sort_index(ascending=False)\n</code></pre>\n<p>output:</p>\n<pre><code>           target\n       mean    count\nseq_len        \n13    0.231788    386034\n12    0.389344    10623\n11    0.446737    5961\n10    0.462282    6721\n9     0.450164    6411\n8     0.447300    6110\n7     0.418430    5198\n6     0.387670    5515\n5     0.392635    4671\n4     0.416221    4673\n3     0.358602    5778\n2     0.318465    6098\n1     0.335742    5120\n</code></pre>\n<p>It is obvious that <strong>sequence length 13</strong> is the most common but also with a significantly lower mean default rate. At first glance, I thought it meant shorter sequences are easier to predict since they have more positive samples. But I'm quickly proven wrong when checking my cross-validation results:</p>\n<pre><code>Fold 0 amex 0.7990 logloss 0.2144\nFold 0 L13 amex 0.8214 logloss 0.1928\nFold 0 Other amex 0.6724  logloss 0.3289\n</code></pre>\n<p>The 1st line is the overall score. The 2nd line is the score of sequences of length 13 and the 3rd line is the score of all the rest sequences. Apparently, shorter sequences have a much worse score than the full sequences of length 13. This is also an implication of how the short sequences are truncated: the more recent profiles are deleted, which could explain the big degradation of the score because more recent profiles have more predicting power in general.  For example, let's say for 13 consecutive months (M1~M13) and sequence A is of length 13 and sequence B is of length 8:</p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  1  1  1  1  1  1  1  1  1  0  0  0  0  0 \n</code></pre>\n<p><br>\nwhere <code>1</code> means features exist and <code>0</code> means features missing. Of course, there is another possibility:</p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1 \n</code></pre>\n<p><br>\nWe can actually find out which one is more plausible by <code>unstacking</code> dataframes in the above two ways and run xgboost with them, respectively. As expected, the former has a better CV score which indicates it is likely how the truncation of short sequences is done.</p>\n<p>If we could somehow predict the missing profiles of sequence <code>B</code>, the life of the downstream XGB models would be made much easier. An intuitive choice is to generate the missing profiles using a one-dimension auto-regressive RNN. Bascially we want to predict the features of the next month based on the feature values of the current month and all previous months. And when we have the prediction for the next month, we can use it as part of the input and predict again and so on so forth. This is also where the availability of the big test data really shines. Since we are predicting features, not <code>target</code>, we can train our models using both <code>train</code> and <code>test</code> data. The RNN structure is very simple:  just one GRU layer and some FC layers. The RNN performance is pretty decent. In terms of RMSE of all 178 numerical features, the GRU achieves validation <code>RMSE 0.019</code>. For simplicity, all features are log-transformed and <code>fillna(0)</code>. You might wonder how good is <code>RMSE 0.019</code>. We can simply compare it with the naive baseline: just repeat the last available month. For example, if I'm asked to predict features of M2, the naive baseline is just output features of M1. The RMSE of this naive baseline is <code>0.03</code> so our RNN actually learns something and could be useful.</p>\n<p>The rest would be straightforward, after predicting missing months, now every sequence is of length 13 so I just unstack the dataframe to increase the number of features 13x. For example, instead of having one feature <code>P_2</code> of the last month, now we have 13 features <code>P_2_M_1</code> to <code>P_2_M_13</code>. These are the most useful features I created. For downstream classifiers I only use XGB so that it is <em>not similar</em> to <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">the great LGB DART notebook</a>. By varying RNN hyperparameters, XGB hyperparameters and different combination of features, I end up with 7 XGB models, whose ensemble is  0.7993 CV and 0.799 public LB. Averaging it with the best public solution and with extraordinary luck, my final submission ended up in the gold zone.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F100236%2Facf9b55d595c8c030770291e3dd69844%2FWechatIMG5474.png?generation=1661401124767084&amp;alt=media\" alt=\"\"></p>\n<p>The final thought is my best auto-regressive features are generated 1 hour before the deadline. I'm very happy it worked!</p>",
  "messages": [
    {
      "id": 1912952,
      "postDate": "2022-08-25T04:12:09.097Z",
      "content": "<p>What a competition! I really enjoyed it and only hope I could have found more time. First of all, I would like to thank Raddar <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a>, Martin <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a>, and many others who generously shared codes and datasets! The public solutions are of amazing quality. And it also determines my game plan: <strong>create something original and blend it with the best public solution.</strong> My solution is based on RAPIDS cudf for dataframe processing, XGB for training, and pytorch lightning for feature extraction.</p>\n<p>There are two vital observations of this dataset:</p>\n<ul>\n<li>the big test data which is from the future is available.</li>\n<li>short sequences (&lt;13) are culprits of bad performance.</li>\n</ul>\n<p>Let's start with the latter observation:<strong>the sequence length of each customer profiles</strong> plays a critical role in the model performance:</p>\n<pre><code>import cudf\npath = '/raid/amex'\ntrain = cudf.read_parquet(f'{path}/train.parquet',columns=['customer_ID'])\ntrainl = cudf.read_csv(f'{path}/train_labels.csv')\ntrain = train.merge(trainl,on='customer_ID',how='left')\ntrain['seq_len'] = train.groupby('customer_ID')['target'].transform('count')\ntrain = train.drop_duplicates('customer_ID',keep='last')\ntrain.groupby('seq_len').agg({'target':['mean','count']}).sort_index(ascending=False)\n</code></pre>\n<p>output:</p>\n<pre><code>           target\n       mean    count\nseq_len        \n13    0.231788    386034\n12    0.389344    10623\n11    0.446737    5961\n10    0.462282    6721\n9     0.450164    6411\n8     0.447300    6110\n7     0.418430    5198\n6     0.387670    5515\n5     0.392635    4671\n4     0.416221    4673\n3     0.358602    5778\n2     0.318465    6098\n1     0.335742    5120\n</code></pre>\n<p>It is obvious that <strong>sequence length 13</strong> is the most common but also with a significantly lower mean default rate. At first glance, I thought it meant shorter sequences are easier to predict since they have more positive samples. But I'm quickly proven wrong when checking my cross-validation results:</p>\n<pre><code>Fold 0 amex 0.7990 logloss 0.2144\nFold 0 L13 amex 0.8214 logloss 0.1928\nFold 0 Other amex 0.6724  logloss 0.3289\n</code></pre>\n<p>The 1st line is the overall score. The 2nd line is the score of sequences of length 13 and the 3rd line is the score of all the rest sequences. Apparently, shorter sequences have a much worse score than the full sequences of length 13. This is also an implication of how the short sequences are truncated: the more recent profiles are deleted, which could explain the big degradation of the score because more recent profiles have more predicting power in general.  For example, let's say for 13 consecutive months (M1~M13) and sequence A is of length 13 and sequence B is of length 8:</p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  1  1  1  1  1  1  1  1  1  0  0  0  0  0 \n</code></pre>\n<p><br>\nwhere <code>1</code> means features exist and <code>0</code> means features missing. Of course, there is another possibility:</p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1 \n</code></pre>\n<p><br>\nWe can actually find out which one is more plausible by <code>unstacking</code> dataframes in the above two ways and run xgboost with them, respectively. As expected, the former has a better CV score which indicates it is likely how the truncation of short sequences is done.</p>\n<p>If we could somehow predict the missing profiles of sequence <code>B</code>, the life of the downstream XGB models would be made much easier. An intuitive choice is to generate the missing profiles using a one-dimension auto-regressive RNN. Bascially we want to predict the features of the next month based on the feature values of the current month and all previous months. And when we have the prediction for the next month, we can use it as part of the input and predict again and so on so forth. This is also where the availability of the big test data really shines. Since we are predicting features, not <code>target</code>, we can train our models using both <code>train</code> and <code>test</code> data. The RNN structure is very simple:  just one GRU layer and some FC layers. The RNN performance is pretty decent. In terms of RMSE of all 178 numerical features, the GRU achieves validation <code>RMSE 0.019</code>. For simplicity, all features are log-transformed and <code>fillna(0)</code>. You might wonder how good is <code>RMSE 0.019</code>. We can simply compare it with the naive baseline: just repeat the last available month. For example, if I'm asked to predict features of M2, the naive baseline is just output features of M1. The RMSE of this naive baseline is <code>0.03</code> so our RNN actually learns something and could be useful.</p>\n<p>The rest would be straightforward, after predicting missing months, now every sequence is of length 13 so I just unstack the dataframe to increase the number of features 13x. For example, instead of having one feature <code>P_2</code> of the last month, now we have 13 features <code>P_2_M_1</code> to <code>P_2_M_13</code>. These are the most useful features I created. For downstream classifiers I only use XGB so that it is <em>not similar</em> to <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">the great LGB DART notebook</a>. By varying RNN hyperparameters, XGB hyperparameters and different combination of features, I end up with 7 XGB models, whose ensemble is  0.7993 CV and 0.799 public LB. Averaging it with the best public solution and with extraordinary luck, my final submission ended up in the gold zone.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F100236%2Facf9b55d595c8c030770291e3dd69844%2FWechatIMG5474.png?generation=1661401124767084&amp;alt=media\" alt=\"\"></p>\n<p>The final thought is my best auto-regressive features are generated 1 hour before the deadline. I'm very happy it worked!</p>",
      "rawMarkdown": "What a competition! I really enjoyed it and only hope I could have found more time. First of all, I would like to thank Raddar @raddar, Martin @ragnar123, and many others who generously shared codes and datasets! The public solutions are of amazing quality. And it also determines my game plan: **create something original and blend it with the best public solution.** My solution is based on RAPIDS cudf for dataframe processing, XGB for training, and pytorch lightning for feature extraction.\n\nThere are two vital observations of this dataset:\n- the big test data which is from the future is available.\n- short sequences (<13) are culprits of bad performance.\n\nLet's start with the latter observation:**the sequence length of each customer profiles** plays a critical role in the model performance:\n```\nimport cudf\npath = '/raid/amex'\ntrain = cudf.read_parquet(f'{path}/train.parquet',columns=['customer_ID'])\ntrainl = cudf.read_csv(f'{path}/train_labels.csv')\ntrain = train.merge(trainl,on='customer_ID',how='left')\ntrain['seq_len'] = train.groupby('customer_ID')['target'].transform('count')\ntrain = train.drop_duplicates('customer_ID',keep='last')\ntrain.groupby('seq_len').agg({'target':['mean','count']}).sort_index(ascending=False)\n```\noutput:\n```\n           target\n       mean\tcount\nseq_len\t\t\n13\t0.231788\t386034\n12\t0.389344\t10623\n11\t0.446737\t5961\n10\t0.462282\t6721\n9\t 0.450164\t 6411\n8\t 0.447300\t 6110\n7\t 0.418430\t 5198\n6\t 0.387670\t 5515\n5\t 0.392635\t 4671\n4\t 0.416221\t 4673\n3\t 0.358602\t 5778\n2\t 0.318465\t 6098\n1\t 0.335742\t 5120\n```\nIt is obvious that **sequence length 13** is the most common but also with a significantly lower mean default rate. At first glance, I thought it meant shorter sequences are easier to predict since they have more positive samples. But I'm quickly proven wrong when checking my cross-validation results:\n```\nFold 0 amex 0.7990 logloss 0.2144\nFold 0 L13 amex 0.8214 logloss 0.1928\nFold 0 Other amex 0.6724  logloss 0.3289\n```\nThe 1st line is the overall score. The 2nd line is the score of sequences of length 13 and the 3rd line is the score of all the rest sequences. Apparently, shorter sequences have a much worse score than the full sequences of length 13. This is also an implication of how the short sequences are truncated: the more recent profiles are deleted, which could explain the big degradation of the score because more recent profiles have more predicting power in general.  For example, let's say for 13 consecutive months (M1~M13) and sequence A is of length 13 and sequence B is of length 8:\n```\n   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  1  1  1  1  1  1  1  1  1  0  0  0  0  0 \n``` \nwhere `1` means features exist and `0` means features missing. Of course, there is another possibility:\n ```\n   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1 \n``` \nWe can actually find out which one is more plausible by `unstacking` dataframes in the above two ways and run xgboost with them, respectively. As expected, the former has a better CV score which indicates it is likely how the truncation of short sequences is done.\n\nIf we could somehow predict the missing profiles of sequence `B`, the life of the downstream XGB models would be made much easier. An intuitive choice is to generate the missing profiles using a one-dimension auto-regressive RNN. Bascially we want to predict the features of the next month based on the feature values of the current month and all previous months. And when we have the prediction for the next month, we can use it as part of the input and predict again and so on so forth. This is also where the availability of the big test data really shines. Since we are predicting features, not `target`, we can train our models using both `train` and `test` data. The RNN structure is very simple:  just one GRU layer and some FC layers. The RNN performance is pretty decent. In terms of RMSE of all 178 numerical features, the GRU achieves validation `RMSE 0.019`. For simplicity, all features are log-transformed and `fillna(0)`. You might wonder how good is `RMSE 0.019`. We can simply compare it with the naive baseline: just repeat the last available month. For example, if I'm asked to predict features of M2, the naive baseline is just output features of M1. The RMSE of this naive baseline is `0.03` so our RNN actually learns something and could be useful.\n\nThe rest would be straightforward, after predicting missing months, now every sequence is of length 13 so I just unstack the dataframe to increase the number of features 13x. For example, instead of having one feature `P_2` of the last month, now we have 13 features `P_2_M_1` to `P_2_M_13`. These are the most useful features I created. For downstream classifiers I only use XGB so that it is *not similar* to [the great LGB DART notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977). By varying RNN hyperparameters, XGB hyperparameters and different combination of features, I end up with 7 XGB models, whose ensemble is  0.7993 CV and 0.799 public LB. Averaging it with the best public solution and with extraordinary luck, my final submission ended up in the gold zone.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F100236%2Facf9b55d595c8c030770291e3dd69844%2FWechatIMG5474.png?generation=1661401124767084&alt=media)\n\nThe final thought is my best auto-regressive features are generated 1 hour before the deadline. I'm very happy it worked!",
      "votes": 93
    },
    {
      "id": 1915127,
      "postDate": "2022-08-26T17:24:46.920Z",
      "content": "<p>I forgot to thank <a href=\"https://www.kaggle.com/pyagoubi\" target=\"_blank\">@pyagoubi</a> Art Vandelay for his great kernel: <a href=\"https://www.kaggle.com/code/pyagoubi/amex-eda-evolvement-of-numeric-features-over-time\" target=\"_blank\">https://www.kaggle.com/code/pyagoubi/amex-eda-evolvement-of-numeric-features-over-time</a><br>\nhere is a screenshot of his drawing:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F100236%2F3364ddfbcdfea257e2c047971aaaf0c2%2FScreen%20Shot%202022-08-26%20at%201.24.11%20PM.png?generation=1661534676162068&amp;alt=media\" alt=\"\"><br>\nThe pattern shows a clear and predictable trend, which motivates me to go down this path. </p>",
      "rawMarkdown": "I forgot to thank @pyagoubi Art Vandelay for his great kernel: https://www.kaggle.com/code/pyagoubi/amex-eda-evolvement-of-numeric-features-over-time\nhere is a screenshot of his drawing:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F100236%2F3364ddfbcdfea257e2c047971aaaf0c2%2FScreen%20Shot%202022-08-26%20at%201.24.11%20PM.png?generation=1661534676162068&alt=media)\nThe pattern shows a clear and predictable trend, which motivates me to go down this path. ",
      "votes": 3,
      "replies": [
        {
          "id": 1916368,
          "postDate": "2022-08-27T19:54:50.567Z",
          "content": "<p>wow thanks!</p>",
          "rawMarkdown": "wow thanks!"
        }
      ]
    },
    {
      "id": 1914643,
      "postDate": "2022-08-26T09:04:53.403Z",
      "content": "<p>Really great idea to generate features for the missing periods!</p>\n<p>Therefore just to understand, every cid has 13 records in the end, right?<br>\nAfter generating the features, did you also predict the target for these new records, or what value for target did you use?</p>\n<p>Can you also give more info on the one-dimension auto-regressive RNN? Does the RNN predicts all the features at once or loop / predict for every feature?</p>\n<p>Thanks</p>",
      "rawMarkdown": "Really great idea to generate features for the missing periods!\n\nTherefore just to understand, every cid has 13 records in the end, right?\nAfter generating the features, did you also predict the target for these new records, or what value for target did you use?\n\nCan you also give more info on the one-dimension auto-regressive RNN? Does the RNN predicts all the features at once or loop / predict for every feature?\n\nThanks",
      "votes": 1,
      "replies": [
        {
          "id": 1914875,
          "postDate": "2022-08-26T13:49:48.363Z",
          "content": "<p>Thank you for the question!</p>\n<blockquote>\n  <p>After generating the features, did you also predict the target for these new records, or what value for target did you use?</p>\n</blockquote>\n<p>I just used their original targets. The purpose is to generate more recent features with respect to the original targets.</p>\n<blockquote>\n  <p>Does the RNN predicts all the features at once?</p>\n</blockquote>\n<p>Yes. please also refer to my answer <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347668#1914873\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "Thank you for the question!\n\n> After generating the features, did you also predict the target for these new records, or what value for target did you use?\n\nI just used their original targets. The purpose is to generate more recent features with respect to the original targets.\n\n> Does the RNN predicts all the features at once?\n\nYes. please also refer to my answer [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347668#1914873)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1914122,
      "postDate": "2022-08-25T18:47:15.057Z",
      "content": "<p>that is nice keep it up!</p>",
      "rawMarkdown": "that is nice keep it up!",
      "votes": 1
    },
    {
      "id": 1913131,
      "postDate": "2022-08-25T06:56:24.130Z",
      "content": "<p>Thanks for sharing, I've learned a lot! Congratulations!</p>",
      "rawMarkdown": "Thanks for sharing, I've learned a lot! Congratulations!",
      "votes": 1
    },
    {
      "id": 1913087,
      "postDate": "2022-08-25T06:13:14.513Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a>,  that's a great insight.</p>",
      "rawMarkdown": "Congratulations @jiweiliu,  that's a great insight.",
      "votes": 1
    },
    {
      "id": 1912988,
      "postDate": "2022-08-25T04:50:49.160Z",
      "content": "<p>Big Congratulations <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> . <br>\nI was trying something similar with KNN regressors, couldn't get a good CV LB to make it work. Feels good that this approach worked like magic. </p>",
      "rawMarkdown": "Big Congratulations @jiweiliu . \nI was trying something similar with KNN regressors, couldn't get a good CV LB to make it work. Feels good that this approach worked like magic. ",
      "votes": 1
    },
    {
      "id": 1913582,
      "postDate": "2022-08-25T11:37:32.033Z",
      "content": "<p>Congratulations Jiwei! Great solution. I love your use of RNN and test data. </p>\n<p>Did you try giving the customers with 13 statements more statements? Like predict them a 14th, 15th, 16th etc etc statement? I wonder if that would boost CV LB</p>",
      "rawMarkdown": "Congratulations Jiwei! Great solution. I love your use of RNN and test data. \n\nDid you try giving the customers with 13 statements more statements? Like predict them a 14th, 15th, 16th etc etc statement? I wonder if that would boost CV LB",
      "votes": 2,
      "replies": [
        {
          "id": 1913613,
          "postDate": "2022-08-25T12:01:39.623Z",
          "content": "<p>Yes, I did. I only tried to predict the 14th statement of customers with 13 statements and it does help. Maybe predicting more is better. I stopped because the prediction of customers with 13 statements is already very accurate, above 0.82 AMEX metric. So I focused more on shorter sequences.</p>",
          "rawMarkdown": "Yes, I did. I only tried to predict the 14th statement of customers with 13 statements and it does help. Maybe predicting more is better. I stopped because the prediction of customers with 13 statements is already very accurate, above 0.82 AMEX metric. So I focused more on shorter sequences.",
          "votes": 4
        },
        {
          "id": 1913621,
          "postDate": "2022-08-25T12:07:14.073Z",
          "content": "<p>Another reason I stopped at 14th is I found that to improve overall AMEX metric, it is more important to make the poor better instead of making the rich richer.</p>\n<pre><code>Fold 0 amex 0.7990 logloss 0.2144\nFold 0 L13 amex 0.8214 logloss 0.1928\nFold 0 Other amex 0.6724  logloss 0.3289\n</code></pre>\n<p>Many times I found <code>L13 amex</code> is improved with the new features but <code>Other amex</code> with short sequences are worse or not improved, and the overall <code>amex</code> is not improved. If the metric is <code>log loss</code>, I would spend more time optimizing <code>L13</code>.</p>",
          "rawMarkdown": "Another reason I stopped at 14th is I found that to improve overall AMEX metric, it is more important to make the poor better instead of making the rich richer.\n```\nFold 0 amex 0.7990 logloss 0.2144\nFold 0 L13 amex 0.8214 logloss 0.1928\nFold 0 Other amex 0.6724  logloss 0.3289\n```\nMany times I found `L13 amex` is improved with the new features but `Other amex` with short sequences are worse or not improved, and the overall `amex` is not improved. If the metric is `log loss`, I would spend more time optimizing `L13`.",
          "votes": 6
        },
        {
          "id": 1913920,
          "postDate": "2022-08-25T15:22:06.790Z",
          "content": "<p>Did you try combining, for example predicted 14th - actual 13th (aka last)? Or last minus predicted last?</p>\n<p>It was on my list of things to try I didn't get to</p>",
          "rawMarkdown": "Did you try combining, for example predicted 14th - actual 13th (aka last)? Or last minus predicted last?\n\nIt was on my list of things to try I didn't get to",
          "votes": 1
        },
        {
          "id": 1914844,
          "postDate": "2022-08-26T13:24:31.057Z",
          "content": "<p>yes, I tried both. Keeping both predicted 14th and actual 13th (aka last) works best with xgb. Last minus predicted last doesn't add value.</p>",
          "rawMarkdown": "yes, I tried both. Keeping both predicted 14th and actual 13th (aka last) works best with xgb. Last minus predicted last doesn't add value.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1913400,
      "postDate": "2022-08-25T10:05:19.020Z",
      "content": "<p><a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> Thanks for sharing and congrats on the solo gold! <br>\nOne thing I wanted to learn more is your construction of short sequences. According to <code>S_2</code>, the sequence alignment should be the second possibility. Am I thinking about something different here? </p>",
      "rawMarkdown": "@jiweiliu Thanks for sharing and congrats on the solo gold! \nOne thing I wanted to learn more is your construction of short sequences. According to `S_2`, the sequence alignment should be the second possibility. Am I thinking about something different here? ",
      "votes": 2,
      "replies": [
        {
          "id": 1913603,
          "postDate": "2022-08-25T11:52:26.943Z",
          "content": "<p>Thank you for the comment! Yes, I think you are right and I wasn't clear in my post. If we align two sequences with the date <code>S_2</code> like Feb/2018, it is the 2nd scenario as you suggested. But if we consider aligning two sequences relative to when the defaults actually happened, it might become the 1st scenario. </p>\n<p>For example, if <code>M1</code> to <code>M13</code> represent <code>Feb/2018</code> to <code>Mar/2019</code>, it is aligned as the 2nd case. </p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1 \n</code></pre>\n<p><br>\nso when do the defaults happen? My assumption is default of <code>B</code> happened sometime after default of <code>A</code>, something like this:</p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13 M14 M15 M16 M17 M18\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1         x \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1                 x \n</code></pre>\n<p><br>\nwhere <code>x</code> indicates default. The point is sequence B is harder to predict which indicates its more recent profiles are missing. So if we align the default dates <code>x</code>, it looks like scenario 1 in my original post:</p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13 M14 M15 M16 M17 M18\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1           x\nB  1  1  1  1  1  1  1  1  1  0  0  0  0  0           x \n</code></pre>\n<p>Of course, this is all my hypothesis and by no means a precise alignment since the true default dates are unknown. However, the idea might be roughly correct and useful. Eventually, it all comes down to what the new feature should mean. <code>P_2_M_13</code> could mean the value of <code>P_2</code> in <code>Mar/2019</code> or it could mean the value of <code>P_2</code> in the most recent month relative to when the default happens. I think the downstream XGB learns better with the latter.</p>",
          "rawMarkdown": "Thank you for the comment! Yes, I think you are right and I wasn't clear in my post. If we align two sequences with the date `S_2` like Feb/2018, it is the 2nd scenario as you suggested. But if we consider aligning two sequences relative to when the defaults actually happened, it might become the 1st scenario. \n\nFor example, if `M1` to `M13` represent `Feb/2018` to `Mar/2019`, it is aligned as the 2nd case. \n ```\n   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1 \n``` \nso when do the defaults happen? My assumption is default of `B` happened sometime after default of `A`, something like this:\n ```\n   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13 M14 M15 M16 M17 M18\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1         x \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1                 x \n``` \nwhere `x` indicates default. The point is sequence B is harder to predict which indicates its more recent profiles are missing. So if we align the default dates `x`, it looks like scenario 1 in my original post:\n```\n   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13 M14 M15 M16 M17 M18\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1           x\nB  1  1  1  1  1  1  1  1  1  0  0  0  0  0           x \n``` \n\nOf course, this is all my hypothesis and by no means a precise alignment since the true default dates are unknown. However, the idea might be roughly correct and useful. Eventually, it all comes down to what the new feature should mean. `P_2_M_13` could mean the value of `P_2` in `Mar/2019` or it could mean the value of `P_2` in the most recent month relative to when the default happens. I think the downstream XGB learns better with the latter.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1928754,
      "postDate": "2022-09-06T15:40:38.713Z",
      "content": "<p>Congrats and thanks for sharing such a great solution! </p>\n<p>I read that you trained only one RNN with length 8 and only predicted 1 step ahead each time, but I didn't get which statements covers that length 8 sequence (you mentioned truncating/padding and right alignment). Following your example:</p>\n<table>\n<thead>\n<tr>\n<th>customer_ID</th>\n<th>M1</th>\n<th>M2</th>\n<th>M3</th>\n<th>M4</th>\n<th>M5</th>\n<th>M6</th>\n<th>M7</th>\n<th>M8</th>\n<th>M9</th>\n<th>M10</th>\n<th>M11</th>\n<th>M12</th>\n<th>M13</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n</tr>\n<tr>\n<td>B</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>C</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>The data used for training would be the following? (X's are features, Y the target):</p>\n<table>\n<thead>\n<tr>\n<th>customer_ID</th>\n<th>M1</th>\n<th>M2</th>\n<th>M3</th>\n<th>M4</th>\n<th>M5</th>\n<th>M6</th>\n<th>M7</th>\n<th>M8</th>\n<th>M9</th>\n<th>M10</th>\n<th>M11</th>\n<th>M12</th>\n<th>M13</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>X1</td>\n<td>X2</td>\n<td>X3</td>\n<td>X4</td>\n<td>X5</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n</tr>\n<tr>\n<td>B</td>\n<td>X2</td>\n<td>X3</td>\n<td>X4</td>\n<td>X5</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>C</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n</tr>\n</tbody>\n</table>\n<p>And in the end:</p>\n<table>\n<thead>\n<tr>\n<th>customer_ID</th>\n<th>S1</th>\n<th>S2</th>\n<th>S3</th>\n<th>S4</th>\n<th>S5</th>\n<th>S6</th>\n<th>S7</th>\n<th>S8</th>\n<th>Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>X1</td>\n<td>X2</td>\n<td>X3</td>\n<td>X4</td>\n<td>X5</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n</tr>\n<tr>\n<td>B</td>\n<td>Pad</td>\n<td>X2</td>\n<td>X3</td>\n<td>X4</td>\n<td>X5</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n</tr>\n<tr>\n<td>C</td>\n<td>Pad</td>\n<td>Pad</td>\n<td>Pad</td>\n<td>Pad</td>\n<td>Pad</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n</tr>\n</tbody>\n</table>\n<p>Is that correct?</p>",
      "rawMarkdown": "Congrats and thanks for sharing such a great solution! \n\nI read that you trained only one RNN with length 8 and only predicted 1 step ahead each time, but I didn't get which statements covers that length 8 sequence (you mentioned truncating/padding and right alignment). Following your example:\n\n\n| customer_ID | M1 | M2 | M3 | M4 | M5 | M6 | M7 | M8 | M9 | M10 | M11 | M12 | M13 |\n| --- | --- |\n| A | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |\n| B | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 |\n| C | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |\n\nThe data used for training would be the following? (X's are features, Y the target):\n\n| customer_ID | M1 | M2 | M3 | M4 | M5 | M6 | M7 | M8 | M9 | M10 | M11 | M12 | M13 | \n| --- | --- |\n| A | - | - | - | - | X1 | X2 | X3 | X4 | X5 | X6 | X7 | X8 | Y |\n| B | X2 | X3 | X4 | X5 | X6 | X7 | X8 | Y | - | - | - | - | - |\n| C | X6 | X7 | X8 | Y | - | - | - | - | - | - | - | - | - |\n\nAnd in the end:\n\n| customer_ID | S1 | S2 | S3 | S4 | S5 | S6 | S7 | S8 | Target\n| --- | --- |\n| A | X1 | X2 | X3 | X4 | X5 | X6 | X7 | X8 | Y |\n| B | Pad | X2 | X3 | X4 | X5 | X6 | X7 | X8 | Y | \n| C | Pad | Pad | Pad | Pad | Pad | X6 | X7 | X8 | Y |\n\nIs that correct?\n"
    },
    {
      "id": 1913032,
      "postDate": "2022-08-25T05:23:24.300Z",
      "content": "<p>Thank you very much for the well written, detailed explanation. I have some questions, but if you don't have time or desire to answer, I will not take offense. This is already very enlightening so feel free to ignore.</p>\n<h3>RNN:</h3>\n<p>1) Is the below impression nearly correct? <br>\nMy impression from <code>we have the prediction for the next month, we can use it as part of the input and predict again and so on so forth</code> is that the model predicts only 1 month ahead. Do you train different models for different number of previous months? If not, it would seem to me that the model would need features for up to 12 previous months, and instances with fewer than 12 would have NaN for those features. Also, you could have a large amount of training data, since an instance with no missing months provides 12 points of training data for this RNN model; the second month (with 1 previous month of know data), the third (with 2 previous months …) and so on. </p>\n<p>2) Did you provide a flag to the XGB models that the feature sequence of a row had filled data?</p>\n<h3>CV</h3>\n<p>3) Do you do a simple k-fold cv (early stop on the validation set and count the score on the same set for your CV score), or did you use nested cv, early stopping on the inner valid set. Or maybe not early stopping, but retraining with an average of best number of iterations. Is there ever a time in which you don't use any early stopping with XGBoost? Sorry for the detail seeking question about CV. I am just really curious as to best practices around cv are. I always figured nested cv was the only \"true\" CV, but I don't find people differentiating which scheme they are using for some reason.</p>",
      "rawMarkdown": "Thank you very much for the well written, detailed explanation. I have some questions, but if you don't have time or desire to answer, I will not take offense. This is already very enlightening so feel free to ignore.\n### RNN:  \n1) Is the below impression nearly correct? \nMy impression from `we have the prediction for the next month, we can use it as part of the input and predict again and so on so forth` is that the model predicts only 1 month ahead. Do you train different models for different number of previous months? If not, it would seem to me that the model would need features for up to 12 previous months, and instances with fewer than 12 would have NaN for those features. Also, you could have a large amount of training data, since an instance with no missing months provides 12 points of training data for this RNN model; the second month (with 1 previous month of know data), the third (with 2 previous months ...) and so on. \n\n2) Did you provide a flag to the XGB models that the feature sequence of a row had filled data?\n\n### CV\n3) Do you do a simple k-fold cv (early stop on the validation set and count the score on the same set for your CV score), or did you use nested cv, early stopping on the inner valid set. Or maybe not early stopping, but retraining with an average of best number of iterations. Is there ever a time in which you don't use any early stopping with XGBoost? Sorry for the detail seeking question about CV. I am just really curious as to best practices around cv are. I always figured nested cv was the only \"true\" CV, but I don't find people differentiating which scheme they are using for some reason.",
      "replies": [
        {
          "id": 1913688,
          "postDate": "2022-08-25T12:54:30.117Z",
          "content": "<p>Hi Chris, these are great questions! I'll come back to this and give more details. The short answers are:</p>\n<ul>\n<li>I only trained one RNN model of length <code>8</code> to generate future profiles. The input data are truncated or padded to make length 8 and they are right aligned.</li>\n<li>No, I don't give a flag to xgb to indicate the features are generated or real. But this is an interesting idea. might work.</li>\n<li>Yes, I used simple k-fold. but I found very late in the competition that a stratified k-fold on sequence-length might be better.</li>\n<li>Yes, I used xgb early stopping but very large <code>early_stopping_rounds</code>.</li>\n</ul>",
          "rawMarkdown": "Hi Chris, these are great questions! I'll come back to this and give more details. The short answers are:\n- I only trained one RNN model of length `8` to generate future profiles. The input data are truncated or padded to make length 8 and they are right aligned.\n- No, I don't give a flag to xgb to indicate the features are generated or real. But this is an interesting idea. might work.\n- Yes, I used simple k-fold. but I found very late in the competition that a stratified k-fold on sequence-length might be better.\n- Yes, I used xgb early stopping but very large `early_stopping_rounds`.",
          "votes": 4
        },
        {
          "id": 1913967,
          "postDate": "2022-08-25T15:57:32.277Z",
          "content": "<p>Thank you very much. The answers are very informative. </p>\n<p>I forgot to ask, but does the RNN predict all features at once? So the input is 8x188 output 1x188</p>",
          "rawMarkdown": "Thank you very much. The answers are very informative. \n\nI forgot to ask, but does the RNN predict all features at once? So the input is 8x188 output 1x188"
        },
        {
          "id": 1914850,
          "postDate": "2022-08-26T13:28:06.733Z",
          "content": "<p>Yes the RNN predicts all features at once. Both input and output are 8*188 but the output is one time step ahead.</p>",
          "rawMarkdown": "Yes the RNN predicts all features at once. Both input and output are 8*188 but the output is one time step ahead.",
          "votes": 2
        },
        {
          "id": 1914873,
          "postDate": "2022-08-26T13:49:19.650Z",
          "content": "<p>To be precise, RNN maps an input <code>[B,S,N]</code> tensor to an output <code>[B,S,N]</code> tensor, where <code>B</code> is the batch size, <code>S</code> is the sequence length and <code>N</code> is the number of all numerical features. The output tensor is one step ahead of the input tensor (output is for the future).  In inference, we only take the last step of the output tensor and attach it to the input tensor and predict again. </p>",
          "rawMarkdown": "To be precise, RNN maps an input `[B,S,N]` tensor to an output `[B,S,N]` tensor, where `B` is the batch size, `S` is the sequence length and `N` is the number of all numerical features. The output tensor is one step ahead of the input tensor (output is for the future).  In inference, we only take the last step of the output tensor and attach it to the input tensor and predict again. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1913022,
      "postDate": "2022-08-25T05:18:21.877Z",
      "content": "<p>Beatiful and original idea. Well done!</p>",
      "rawMarkdown": "Beatiful and original idea. Well done!"
    },
    {
      "id": 1913000,
      "postDate": "2022-08-25T04:59:18.513Z",
      "content": "<p>Nice Job! Congrats!</p>",
      "rawMarkdown": "Nice Job! Congrats!"
    },
    {
      "id": 1912974,
      "postDate": "2022-08-25T04:43:28.423Z",
      "content": "<p>Amazing way to analyze data and generate features!<br>\nThanks a lot!</p>",
      "rawMarkdown": "Amazing way to analyze data and generate features!\nThanks a lot!"
    },
    {
      "id": 1913082,
      "postDate": "2022-08-25T06:06:13.753Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1913611,
          "postDate": "2022-08-25T11:56:56.170Z",
          "content": "<p>Thank you for the question. The short answer is I ignored the gaps in the sequence. I think what's important is if the more recent profiles of a customer are available, not what's missing in the past. Please also refer to my response to Tonghui Li <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347668#1913603\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "Thank you for the question. The short answer is I ignored the gaps in the sequence. I think what's important is if the more recent profiles of a customer are available, not what's missing in the past. Please also refer to my response to Tonghui Li [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/347668#1913603).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1913211,
      "postDate": "2022-08-25T08:19:43.210Z",
      "content": "<p>Thanks for the explanation and congrats</p>",
      "rawMarkdown": "Thanks for the explanation and congrats",
      "votes": 1
    },
    {
      "id": 1913166,
      "postDate": "2022-08-25T07:35:32.273Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing~~~~~",
      "votes": 1
    },
    {
      "id": 1912957,
      "postDate": "2022-08-25T04:19:43.447Z",
      "content": "<p>Great job! Thanks for sharing</p>",
      "rawMarkdown": "Great job! Thanks for sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1915127,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2022-08-26T17:24:46.920000",
      "content": "<p>I forgot to thank <a href=\"https://www.kaggle.com/pyagoubi\" target=\"_blank\">@pyagoubi</a> Art Vandelay for his great kernel: <a href=\"https://www.kaggle.com/code/pyagoubi/amex-eda-evolvement-of-numeric-features-over-time\" target=\"_blank\">https://www.kaggle.com/code/pyagoubi/amex-eda-evolvement-of-numeric-features-over-time</a><br>\nhere is a screenshot of his drawing:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F100236%2F3364ddfbcdfea257e2c047971aaaf0c2%2FScreen%20Shot%202022-08-26%20at%201.24.11%20PM.png?generation=1661534676162068&amp;alt=media\" alt=\"\"><br>\nThe pattern shows a clear and predictable trend, which motivates me to go down this path. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1916368,
          "author_name": "Art Vandelay",
          "author_url": "",
          "post_date": "2022-08-27T19:54:50.567000",
          "content": "<p>wow thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1914643,
      "author_name": "Jonathan Mallia",
      "author_url": "",
      "post_date": "2022-08-26T09:04:53.403000",
      "content": "<p>Really great idea to generate features for the missing periods!</p>\n<p>Therefore just to understand, every cid has 13 records in the end, right?<br>\nAfter generating the features, did you also predict the target for these new records, or what value for target did you use?</p>\n<p>Can you also give more info on the one-dimension auto-regressive RNN? Does the RNN predicts all the features at once or loop / predict for every feature?</p>\n<p>Thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1914875,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2022-08-26T13:49:48.363000",
          "content": "<p>Thank you for the question!</p>\n<blockquote>\n  <p>After generating the features, did you also predict the target for these new records, or what value for target did you use?</p>\n</blockquote>\n<p>I just used their original targets. The purpose is to generate more recent features with respect to the original targets.</p>\n<blockquote>\n  <p>Does the RNN predicts all the features at once?</p>\n</blockquote>\n<p>Yes. please also refer to my answer <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347668#1914873\" target=\"_blank\">here</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1914122,
      "author_name": "rajashekar vt",
      "author_url": "",
      "post_date": "2022-08-25T18:47:15.057000",
      "content": "<p>that is nice keep it up!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913131,
      "author_name": "XiaFire",
      "author_url": "",
      "post_date": "2022-08-25T06:56:24.130000",
      "content": "<p>Thanks for sharing, I've learned a lot! Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913087,
      "author_name": "Priyanshu Chaudhary",
      "author_url": "",
      "post_date": "2022-08-25T06:13:14.513000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a>,  that's a great insight.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912988,
      "author_name": "tarick.morty",
      "author_url": "",
      "post_date": "2022-08-25T04:50:49.160000",
      "content": "<p>Big Congratulations <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> . <br>\nI was trying something similar with KNN regressors, couldn't get a good CV LB to make it work. Feels good that this approach worked like magic. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913582,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-08-25T11:37:32.033000",
      "content": "<p>Congratulations Jiwei! Great solution. I love your use of RNN and test data. </p>\n<p>Did you try giving the customers with 13 statements more statements? Like predict them a 14th, 15th, 16th etc etc statement? I wonder if that would boost CV LB</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1913613,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2022-08-25T12:01:39.623000",
          "content": "<p>Yes, I did. I only tried to predict the 14th statement of customers with 13 statements and it does help. Maybe predicting more is better. I stopped because the prediction of customers with 13 statements is already very accurate, above 0.82 AMEX metric. So I focused more on shorter sequences.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1913621,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2022-08-25T12:07:14.073000",
          "content": "<p>Another reason I stopped at 14th is I found that to improve overall AMEX metric, it is more important to make the poor better instead of making the rich richer.</p>\n<pre><code>Fold 0 amex 0.7990 logloss 0.2144\nFold 0 L13 amex 0.8214 logloss 0.1928\nFold 0 Other amex 0.6724  logloss 0.3289\n</code></pre>\n<p>Many times I found <code>L13 amex</code> is improved with the new features but <code>Other amex</code> with short sequences are worse or not improved, and the overall <code>amex</code> is not improved. If the metric is <code>log loss</code>, I would spend more time optimizing <code>L13</code>.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1913920,
          "author_name": "Robert Hatch",
          "author_url": "",
          "post_date": "2022-08-25T15:22:06.790000",
          "content": "<p>Did you try combining, for example predicted 14th - actual 13th (aka last)? Or last minus predicted last?</p>\n<p>It was on my list of things to try I didn't get to</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1914844,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2022-08-26T13:24:31.057000",
          "content": "<p>yes, I tried both. Keeping both predicted 14th and actual 13th (aka last) works best with xgb. Last minus predicted last doesn't add value.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1913400,
      "author_name": "Tonghui Li",
      "author_url": "",
      "post_date": "2022-08-25T10:05:19.020000",
      "content": "<p><a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> Thanks for sharing and congrats on the solo gold! <br>\nOne thing I wanted to learn more is your construction of short sequences. According to <code>S_2</code>, the sequence alignment should be the second possibility. Am I thinking about something different here? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1913603,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2022-08-25T11:52:26.943000",
          "content": "<p>Thank you for the comment! Yes, I think you are right and I wasn't clear in my post. If we align two sequences with the date <code>S_2</code> like Feb/2018, it is the 2nd scenario as you suggested. But if we consider aligning two sequences relative to when the defaults actually happened, it might become the 1st scenario. </p>\n<p>For example, if <code>M1</code> to <code>M13</code> represent <code>Feb/2018</code> to <code>Mar/2019</code>, it is aligned as the 2nd case. </p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1 \n</code></pre>\n<p><br>\nso when do the defaults happen? My assumption is default of <code>B</code> happened sometime after default of <code>A</code>, something like this:</p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13 M14 M15 M16 M17 M18\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1         x \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1                 x \n</code></pre>\n<p><br>\nwhere <code>x</code> indicates default. The point is sequence B is harder to predict which indicates its more recent profiles are missing. So if we align the default dates <code>x</code>, it looks like scenario 1 in my original post:</p>\n<pre><code>   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13 M14 M15 M16 M17 M18\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1           x\nB  1  1  1  1  1  1  1  1  1  0  0  0  0  0           x \n</code></pre>\n<p>Of course, this is all my hypothesis and by no means a precise alignment since the true default dates are unknown. However, the idea might be roughly correct and useful. Eventually, it all comes down to what the new feature should mean. <code>P_2_M_13</code> could mean the value of <code>P_2</code> in <code>Mar/2019</code> or it could mean the value of <code>P_2</code> in the most recent month relative to when the default happens. I think the downstream XGB learns better with the latter.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1928754,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "2022-09-06T15:40:38.713000",
      "content": "<p>Congrats and thanks for sharing such a great solution! </p>\n<p>I read that you trained only one RNN with length 8 and only predicted 1 step ahead each time, but I didn't get which statements covers that length 8 sequence (you mentioned truncating/padding and right alignment). Following your example:</p>\n<table>\n<thead>\n<tr>\n<th>customer_ID</th>\n<th>M1</th>\n<th>M2</th>\n<th>M3</th>\n<th>M4</th>\n<th>M5</th>\n<th>M6</th>\n<th>M7</th>\n<th>M8</th>\n<th>M9</th>\n<th>M10</th>\n<th>M11</th>\n<th>M12</th>\n<th>M13</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n</tr>\n<tr>\n<td>B</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n<tr>\n<td>C</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>The data used for training would be the following? (X's are features, Y the target):</p>\n<table>\n<thead>\n<tr>\n<th>customer_ID</th>\n<th>M1</th>\n<th>M2</th>\n<th>M3</th>\n<th>M4</th>\n<th>M5</th>\n<th>M6</th>\n<th>M7</th>\n<th>M8</th>\n<th>M9</th>\n<th>M10</th>\n<th>M11</th>\n<th>M12</th>\n<th>M13</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>X1</td>\n<td>X2</td>\n<td>X3</td>\n<td>X4</td>\n<td>X5</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n</tr>\n<tr>\n<td>B</td>\n<td>X2</td>\n<td>X3</td>\n<td>X4</td>\n<td>X5</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n</tr>\n<tr>\n<td>C</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n</tr>\n</tbody>\n</table>\n<p>And in the end:</p>\n<table>\n<thead>\n<tr>\n<th>customer_ID</th>\n<th>S1</th>\n<th>S2</th>\n<th>S3</th>\n<th>S4</th>\n<th>S5</th>\n<th>S6</th>\n<th>S7</th>\n<th>S8</th>\n<th>Target</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>A</td>\n<td>X1</td>\n<td>X2</td>\n<td>X3</td>\n<td>X4</td>\n<td>X5</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n</tr>\n<tr>\n<td>B</td>\n<td>Pad</td>\n<td>X2</td>\n<td>X3</td>\n<td>X4</td>\n<td>X5</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n</tr>\n<tr>\n<td>C</td>\n<td>Pad</td>\n<td>Pad</td>\n<td>Pad</td>\n<td>Pad</td>\n<td>Pad</td>\n<td>X6</td>\n<td>X7</td>\n<td>X8</td>\n<td>Y</td>\n</tr>\n</tbody>\n</table>\n<p>Is that correct?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913032,
      "author_name": "Chris Miles",
      "author_url": "",
      "post_date": "2022-08-25T05:23:24.300000",
      "content": "<p>Thank you very much for the well written, detailed explanation. I have some questions, but if you don't have time or desire to answer, I will not take offense. This is already very enlightening so feel free to ignore.</p>\n<h3>RNN:</h3>\n<p>1) Is the below impression nearly correct? <br>\nMy impression from <code>we have the prediction for the next month, we can use it as part of the input and predict again and so on so forth</code> is that the model predicts only 1 month ahead. Do you train different models for different number of previous months? If not, it would seem to me that the model would need features for up to 12 previous months, and instances with fewer than 12 would have NaN for those features. Also, you could have a large amount of training data, since an instance with no missing months provides 12 points of training data for this RNN model; the second month (with 1 previous month of know data), the third (with 2 previous months …) and so on. </p>\n<p>2) Did you provide a flag to the XGB models that the feature sequence of a row had filled data?</p>\n<h3>CV</h3>\n<p>3) Do you do a simple k-fold cv (early stop on the validation set and count the score on the same set for your CV score), or did you use nested cv, early stopping on the inner valid set. Or maybe not early stopping, but retraining with an average of best number of iterations. Is there ever a time in which you don't use any early stopping with XGBoost? Sorry for the detail seeking question about CV. I am just really curious as to best practices around cv are. I always figured nested cv was the only \"true\" CV, but I don't find people differentiating which scheme they are using for some reason.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1913688,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2022-08-25T12:54:30.117000",
          "content": "<p>Hi Chris, these are great questions! I'll come back to this and give more details. The short answers are:</p>\n<ul>\n<li>I only trained one RNN model of length <code>8</code> to generate future profiles. The input data are truncated or padded to make length 8 and they are right aligned.</li>\n<li>No, I don't give a flag to xgb to indicate the features are generated or real. But this is an interesting idea. might work.</li>\n<li>Yes, I used simple k-fold. but I found very late in the competition that a stratified k-fold on sequence-length might be better.</li>\n<li>Yes, I used xgb early stopping but very large <code>early_stopping_rounds</code>.</li>\n</ul>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1913967,
          "author_name": "Chris Miles",
          "author_url": "",
          "post_date": "2022-08-25T15:57:32.277000",
          "content": "<p>Thank you very much. The answers are very informative. </p>\n<p>I forgot to ask, but does the RNN predict all features at once? So the input is 8x188 output 1x188</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1914850,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2022-08-26T13:28:06.733000",
          "content": "<p>Yes the RNN predicts all features at once. Both input and output are 8*188 but the output is one time step ahead.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1914873,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2022-08-26T13:49:19.650000",
          "content": "<p>To be precise, RNN maps an input <code>[B,S,N]</code> tensor to an output <code>[B,S,N]</code> tensor, where <code>B</code> is the batch size, <code>S</code> is the sequence length and <code>N</code> is the number of all numerical features. The output tensor is one step ahead of the input tensor (output is for the future).  In inference, we only take the last step of the output tensor and attach it to the input tensor and predict again. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1913022,
      "author_name": "Elias",
      "author_url": "",
      "post_date": "2022-08-25T05:18:21.877000",
      "content": "<p>Beatiful and original idea. Well done!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913000,
      "author_name": "iliiiiiili",
      "author_url": "",
      "post_date": "2022-08-25T04:59:18.513000",
      "content": "<p>Nice Job! Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1912974,
      "author_name": "Oleg Khudyakov",
      "author_url": "",
      "post_date": "2022-08-25T04:43:28.423000",
      "content": "<p>Amazing way to analyze data and generate features!<br>\nThanks a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1913082,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T06:06:13.753000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1913611,
          "author_name": "Jiwei Liu",
          "author_url": "",
          "post_date": "2022-08-25T11:56:56.170000",
          "content": "<p>Thank you for the question. The short answer is I ignored the gaps in the sequence. I think what's important is if the more recent profiles of a customer are available, not what's missing in the past. Please also refer to my response to Tonghui Li <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347668#1913603\" target=\"_blank\">here</a>.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1913211,
      "author_name": "Santiago Mota",
      "author_url": "",
      "post_date": "2022-08-25T08:19:43.210000",
      "content": "<p>Thanks for the explanation and congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913166,
      "author_name": "SgangX",
      "author_url": "",
      "post_date": "2022-08-25T07:35:32.273000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912957,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2022-08-25T04:19:43.447000",
      "content": "<p>Great job! Thanks for sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1912952": "What a competition! I really enjoyed it and only hope I could have found more time. First of all, I would like to thank Raddar @raddar, Martin @ragnar123, and many others who generously shared codes and datasets! The public solutions are of amazing quality. And it also determines my game plan: **create something original and blend it with the best public solution.** My solution is based on RAPIDS cudf for dataframe processing, XGB for training, and pytorch lightning for feature extraction.\n\nThere are two vital observations of this dataset:\n- the big test data which is from the future is available.\n- short sequences (<13) are culprits of bad performance.\n\nLet's start with the latter observation:**the sequence length of each customer profiles** plays a critical role in the model performance:\n```\nimport cudf\npath = '/raid/amex'\ntrain = cudf.read_parquet(f'{path}/train.parquet',columns=['customer_ID'])\ntrainl = cudf.read_csv(f'{path}/train_labels.csv')\ntrain = train.merge(trainl,on='customer_ID',how='left')\ntrain['seq_len'] = train.groupby('customer_ID')['target'].transform('count')\ntrain = train.drop_duplicates('customer_ID',keep='last')\ntrain.groupby('seq_len').agg({'target':['mean','count']}).sort_index(ascending=False)\n```\noutput:\n```\n           target\n       mean\tcount\nseq_len\t\t\n13\t0.231788\t386034\n12\t0.389344\t10623\n11\t0.446737\t5961\n10\t0.462282\t6721\n9\t 0.450164\t 6411\n8\t 0.447300\t 6110\n7\t 0.418430\t 5198\n6\t 0.387670\t 5515\n5\t 0.392635\t 4671\n4\t 0.416221\t 4673\n3\t 0.358602\t 5778\n2\t 0.318465\t 6098\n1\t 0.335742\t 5120\n```\nIt is obvious that **sequence length 13** is the most common but also with a significantly lower mean default rate. At first glance, I thought it meant shorter sequences are easier to predict since they have more positive samples. But I'm quickly proven wrong when checking my cross-validation results:\n```\nFold 0 amex 0.7990 logloss 0.2144\nFold 0 L13 amex 0.8214 logloss 0.1928\nFold 0 Other amex 0.6724  logloss 0.3289\n```\nThe 1st line is the overall score. The 2nd line is the score of sequences of length 13 and the 3rd line is the score of all the rest sequences. Apparently, shorter sequences have a much worse score than the full sequences of length 13. This is also an implication of how the short sequences are truncated: the more recent profiles are deleted, which could explain the big degradation of the score because more recent profiles have more predicting power in general.  For example, let's say for 13 consecutive months (M1~M13) and sequence A is of length 13 and sequence B is of length 8:\n```\n   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  1  1  1  1  1  1  1  1  1  0  0  0  0  0 \n``` \nwhere `1` means features exist and `0` means features missing. Of course, there is another possibility:\n ```\n   M1 M2 M3 M4 M5 M6 M7 M8 M9 M10 M11 M12 M13\nA  1  1  1  1  1  1  1  1  1  1  1  1  1  1 \nB  0  0  0  0  0  1  1  1  1  1  1  1  1  1 \n``` \nWe can actually find out which one is more plausible by `unstacking` dataframes in the above two ways and run xgboost with them, respectively. As expected, the former has a better CV score which indicates it is likely how the truncation of short sequences is done.\n\nIf we could somehow predict the missing profiles of sequence `B`, the life of the downstream XGB models would be made much easier. An intuitive choice is to generate the missing profiles using a one-dimension auto-regressive RNN. Bascially we want to predict the features of the next month based on the feature values of the current month and all previous months. And when we have the prediction for the next month, we can use it as part of the input and predict again and so on so forth. This is also where the availability of the big test data really shines. Since we are predicting features, not `target`, we can train our models using both `train` and `test` data. The RNN structure is very simple:  just one GRU layer and some FC layers. The RNN performance is pretty decent. In terms of RMSE of all 178 numerical features, the GRU achieves validation `RMSE 0.019`. For simplicity, all features are log-transformed and `fillna(0)`. You might wonder how good is `RMSE 0.019`. We can simply compare it with the naive baseline: just repeat the last available month. For example, if I'm asked to predict features of M2, the naive baseline is just output features of M1. The RMSE of this naive baseline is `0.03` so our RNN actually learns something and could be useful.\n\nThe rest would be straightforward, after predicting missing months, now every sequence is of length 13 so I just unstack the dataframe to increase the number of features 13x. For example, instead of having one feature `P_2` of the last month, now we have 13 features `P_2_M_1` to `P_2_M_13`. These are the most useful features I created. For downstream classifiers I only use XGB so that it is *not similar* to [the great LGB DART notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977). By varying RNN hyperparameters, XGB hyperparameters and different combination of features, I end up with 7 XGB models, whose ensemble is  0.7993 CV and 0.799 public LB. Averaging it with the best public solution and with extraordinary luck, my final submission ended up in the gold zone.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F100236%2Facf9b55d595c8c030770291e3dd69844%2FWechatIMG5474.png?generation=1661401124767084&alt=media)\n\nThe final thought is my best auto-regressive features are generated 1 hour before the deadline. I'm very happy it worked!",
    "1915127": "I forgot to thank @pyagoubi Art Vandelay for his great kernel: https://www.kaggle.com/code/pyagoubi/amex-eda-evolvement-of-numeric-features-over-time\nhere is a screenshot of his drawing:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F100236%2F3364ddfbcdfea257e2c047971aaaf0c2%2FScreen%20Shot%202022-08-26%20at%201.24.11%20PM.png?generation=1661534676162068&alt=media)\nThe pattern shows a clear and predictable trend, which motivates me to go down this path. ",
    "1914643": "Really great idea to generate features for the missing periods!\n\nTherefore just to understand, every cid has 13 records in the end, right?\nAfter generating the features, did you also predict the target for these new records, or what value for target did you use?\n\nCan you also give more info on the one-dimension auto-regressive RNN? Does the RNN predicts all the features at once or loop / predict for every feature?\n\nThanks",
    "1914122": "that is nice keep it up!",
    "1913131": "Thanks for sharing, I've learned a lot! Congratulations!",
    "1913087": "Congratulations @jiweiliu,  that's a great insight.",
    "1912988": "Big Congratulations @jiweiliu . \nI was trying something similar with KNN regressors, couldn't get a good CV LB to make it work. Feels good that this approach worked like magic. ",
    "1913582": "Congratulations Jiwei! Great solution. I love your use of RNN and test data. \n\nDid you try giving the customers with 13 statements more statements? Like predict them a 14th, 15th, 16th etc etc statement? I wonder if that would boost CV LB",
    "1913400": "@jiweiliu Thanks for sharing and congrats on the solo gold! \nOne thing I wanted to learn more is your construction of short sequences. According to `S_2`, the sequence alignment should be the second possibility. Am I thinking about something different here? ",
    "1928754": "Congrats and thanks for sharing such a great solution! \n\nI read that you trained only one RNN with length 8 and only predicted 1 step ahead each time, but I didn't get which statements covers that length 8 sequence (you mentioned truncating/padding and right alignment). Following your example:\n\n\n| customer_ID | M1 | M2 | M3 | M4 | M5 | M6 | M7 | M8 | M9 | M10 | M11 | M12 | M13 |\n| --- | --- |\n| A | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |\n| B | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 |\n| C | 1 | 1 | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |\n\nThe data used for training would be the following? (X's are features, Y the target):\n\n| customer_ID | M1 | M2 | M3 | M4 | M5 | M6 | M7 | M8 | M9 | M10 | M11 | M12 | M13 | \n| --- | --- |\n| A | - | - | - | - | X1 | X2 | X3 | X4 | X5 | X6 | X7 | X8 | Y |\n| B | X2 | X3 | X4 | X5 | X6 | X7 | X8 | Y | - | - | - | - | - |\n| C | X6 | X7 | X8 | Y | - | - | - | - | - | - | - | - | - |\n\nAnd in the end:\n\n| customer_ID | S1 | S2 | S3 | S4 | S5 | S6 | S7 | S8 | Target\n| --- | --- |\n| A | X1 | X2 | X3 | X4 | X5 | X6 | X7 | X8 | Y |\n| B | Pad | X2 | X3 | X4 | X5 | X6 | X7 | X8 | Y | \n| C | Pad | Pad | Pad | Pad | Pad | X6 | X7 | X8 | Y |\n\nIs that correct?\n",
    "1913032": "Thank you very much for the well written, detailed explanation. I have some questions, but if you don't have time or desire to answer, I will not take offense. This is already very enlightening so feel free to ignore.\n### RNN:  \n1) Is the below impression nearly correct? \nMy impression from `we have the prediction for the next month, we can use it as part of the input and predict again and so on so forth` is that the model predicts only 1 month ahead. Do you train different models for different number of previous months? If not, it would seem to me that the model would need features for up to 12 previous months, and instances with fewer than 12 would have NaN for those features. Also, you could have a large amount of training data, since an instance with no missing months provides 12 points of training data for this RNN model; the second month (with 1 previous month of know data), the third (with 2 previous months ...) and so on. \n\n2) Did you provide a flag to the XGB models that the feature sequence of a row had filled data?\n\n### CV\n3) Do you do a simple k-fold cv (early stop on the validation set and count the score on the same set for your CV score), or did you use nested cv, early stopping on the inner valid set. Or maybe not early stopping, but retraining with an average of best number of iterations. Is there ever a time in which you don't use any early stopping with XGBoost? Sorry for the detail seeking question about CV. I am just really curious as to best practices around cv are. I always figured nested cv was the only \"true\" CV, but I don't find people differentiating which scheme they are using for some reason.",
    "1913022": "Beatiful and original idea. Well done!",
    "1913000": "Nice Job! Congrats!",
    "1912974": "Amazing way to analyze data and generate features!\nThanks a lot!",
    "1913082": "",
    "1913211": "Thanks for the explanation and congrats",
    "1913166": "Thanks for sharing~~~~~",
    "1912957": "Great job! Thanks for sharing"
  }
}