{
  "id": 347641,
  "title": "14th Place Gold – NN Transformer using LGBM Knowledge Distillation",
  "url": "/competitions/amex-default-prediction/discussion/347641",
  "author_name": "Chris Deotte",
  "post_date": "2022-08-25T00:23:33.130000",
  "votes": 272,
  "comment_count": 145,
  "views": 0,
  "content": "<p>Thank you Amex for sharing your data and hosting a fun tabular competition. Thank you Kaggle. Thank you Kagglers for sharing many helpful discussions and notebooks. Thank you Raddar and Martin for your contributions.</p>\n<h1>Solution Overview</h1>\n<p>My solution is a 50%/50% ensemble of LGBM and NN Transformer. The LGBM is based on Martin’s amazing public LGBM <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">here</a> and the NN Transformer is based on my public Transformer <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">here</a>. </p>\n<p>The secret sauce is how we train the Transformer. We first use knowledge distillation from our trained LGBM before fine tuning with the train targets. Furthermore, both train and test data are used for knowledge distillation which helps the Transformer learn the test data distribution.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/summary.png\" alt=\"\"></p>\n<h1>NN Transformer Training</h1>\n<p>My public notebook transformer has 2 layers with skip connections (to make training easy). When using knowledge distillation, we can train a deeper transformer successfully. My final solution uses a 4 layer transformer without skip connections. We also added a GRU layer after transformer blocks and before final classification layers.</p>\n<p>Training is done using 4 cycles of cosine learning schedule. In the first cold start cosine cycle, we pretrain (i.e. Knowldege Distillation) the Transformer using concatenated rows of both LGBM OOF preds and LGBM test preds and leave probabilities between 0 and 1 (i.e. soft labels). During the second cosine cycle, we use a warm start, reduce the learning rate and train with the hard (0 or 1) train targets. For the third and fourth cycle, we repeat cycles one and two.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/schedule.png\" alt=\"\"></p>\n<h1>Model Performance</h1>\n<p>When creating our submission.csv file from our two models, we can use the normal <strong>K-Fold</strong> LGBM OOF preds and normal LGBM test preds. So making a submission is fast an easy. Additionally, we average 5 seeds per model (and slight model variations) for improved performance.</p>\n<p>To tune our two models and compute optimal hyperparameters, we need a leak free reliable CV score. Leak-free CV score is created using <strong>Nested K-Fold</strong> CV. We divide each of 10 outer folds into 10 inner folds. We then train 100 models using GBT. Then each 1 of 10 outer folds has its unique OOF preds and unique test preds. These individualized OOF and test preds are created using only train targets from within the corresponding outer fold train data.</p>\n<p>When computing leak free CV score, we find that our NN Transformer has <strong>CV 0.798 / LB 0.799</strong>, our LGBM has <strong>CV 0.799 / LB 0.799</strong> and our 50%/50% ensemble has <strong>CV 0.800 / LB 0.801</strong>.</p>\n<h1>Fast Experimentation</h1>\n<p>Fast experimention was done using GPU. Thank you <a href=\"https://www.nvidia.com/en-us/\" target=\"_blank\">Nvidia</a> for providing me compute resources for this competition. Experiments were accelerated using 4xV100 32GB.</p>\n<p>Feature engineering was explored using <a href=\"https://rapids.ai/\" target=\"_blank\">RAPIDS cuDF</a> which performs operations like dataframe groupby aggregation on GPU 10-100x faster than using CPU. Many GBT experimental models were trained and evaluated using fast GPU XGB. With 1xV100 GPU, XGB can train 100 models for Nested 10 in 10 K-Fold (i.e. 100 models) on full data in only 2 hours. </p>\n<p>Feature selection was performed using both XGB feature importance and permutation importance. Using <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\" target=\"_blank\">RAPIDS FIL</a>, we can perform permutation importance where we randomly shuffle each feature column 10 times for each of 10 folds (and average 100 results) in blazing speed! </p>\n<p>Each of 1000s of feature columns, we infer 100 times. This is a total of 100,000s of model inferences where each model is 1000s of individual trees! Using <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\" target=\"_blank\">RAPIDs FIL</a>, we can perform this quickly! Note that we can even take an existing CPU LGBM Dart model and convert it into a GPU <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\" target=\"_blank\">RAPIDS FIL</a> inference model and perform permutation importance on existing LGBM Dart Models in blazing speed!</p>",
  "messages": [
    {
      "id": 1912731,
      "postDate": "2022-08-25T00:23:33.130Z",
      "content": "<p>Thank you Amex for sharing your data and hosting a fun tabular competition. Thank you Kaggle. Thank you Kagglers for sharing many helpful discussions and notebooks. Thank you Raddar and Martin for your contributions.</p>\n<h1>Solution Overview</h1>\n<p>My solution is a 50%/50% ensemble of LGBM and NN Transformer. The LGBM is based on Martin’s amazing public LGBM <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\" target=\"_blank\">here</a> and the NN Transformer is based on my public Transformer <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\" target=\"_blank\">here</a>. </p>\n<p>The secret sauce is how we train the Transformer. We first use knowledge distillation from our trained LGBM before fine tuning with the train targets. Furthermore, both train and test data are used for knowledge distillation which helps the Transformer learn the test data distribution.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/summary.png\" alt=\"\"></p>\n<h1>NN Transformer Training</h1>\n<p>My public notebook transformer has 2 layers with skip connections (to make training easy). When using knowledge distillation, we can train a deeper transformer successfully. My final solution uses a 4 layer transformer without skip connections. We also added a GRU layer after transformer blocks and before final classification layers.</p>\n<p>Training is done using 4 cycles of cosine learning schedule. In the first cold start cosine cycle, we pretrain (i.e. Knowldege Distillation) the Transformer using concatenated rows of both LGBM OOF preds and LGBM test preds and leave probabilities between 0 and 1 (i.e. soft labels). During the second cosine cycle, we use a warm start, reduce the learning rate and train with the hard (0 or 1) train targets. For the third and fourth cycle, we repeat cycles one and two.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/schedule.png\" alt=\"\"></p>\n<h1>Model Performance</h1>\n<p>When creating our submission.csv file from our two models, we can use the normal <strong>K-Fold</strong> LGBM OOF preds and normal LGBM test preds. So making a submission is fast an easy. Additionally, we average 5 seeds per model (and slight model variations) for improved performance.</p>\n<p>To tune our two models and compute optimal hyperparameters, we need a leak free reliable CV score. Leak-free CV score is created using <strong>Nested K-Fold</strong> CV. We divide each of 10 outer folds into 10 inner folds. We then train 100 models using GBT. Then each 1 of 10 outer folds has its unique OOF preds and unique test preds. These individualized OOF and test preds are created using only train targets from within the corresponding outer fold train data.</p>\n<p>When computing leak free CV score, we find that our NN Transformer has <strong>CV 0.798 / LB 0.799</strong>, our LGBM has <strong>CV 0.799 / LB 0.799</strong> and our 50%/50% ensemble has <strong>CV 0.800 / LB 0.801</strong>.</p>\n<h1>Fast Experimentation</h1>\n<p>Fast experimention was done using GPU. Thank you <a href=\"https://www.nvidia.com/en-us/\" target=\"_blank\">Nvidia</a> for providing me compute resources for this competition. Experiments were accelerated using 4xV100 32GB.</p>\n<p>Feature engineering was explored using <a href=\"https://rapids.ai/\" target=\"_blank\">RAPIDS cuDF</a> which performs operations like dataframe groupby aggregation on GPU 10-100x faster than using CPU. Many GBT experimental models were trained and evaluated using fast GPU XGB. With 1xV100 GPU, XGB can train 100 models for Nested 10 in 10 K-Fold (i.e. 100 models) on full data in only 2 hours. </p>\n<p>Feature selection was performed using both XGB feature importance and permutation importance. Using <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\" target=\"_blank\">RAPIDS FIL</a>, we can perform permutation importance where we randomly shuffle each feature column 10 times for each of 10 folds (and average 100 results) in blazing speed! </p>\n<p>Each of 1000s of feature columns, we infer 100 times. This is a total of 100,000s of model inferences where each model is 1000s of individual trees! Using <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\" target=\"_blank\">RAPIDs FIL</a>, we can perform this quickly! Note that we can even take an existing CPU LGBM Dart model and convert it into a GPU <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\" target=\"_blank\">RAPIDS FIL</a> inference model and perform permutation importance on existing LGBM Dart Models in blazing speed!</p>",
      "rawMarkdown": "Thank you Amex for sharing your data and hosting a fun tabular competition. Thank you Kaggle. Thank you Kagglers for sharing many helpful discussions and notebooks. Thank you Raddar and Martin for your contributions.\n\n# Solution Overview\nMy solution is a 50%/50% ensemble of LGBM and NN Transformer. The LGBM is based on Martin’s amazing public LGBM [here][1] and the NN Transformer is based on my public Transformer [here][2]. \n\nThe secret sauce is how we train the Transformer. We first use knowledge distillation from our trained LGBM before fine tuning with the train targets. Furthermore, both train and test data are used for knowledge distillation which helps the Transformer learn the test data distribution.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/summary.png)\n\n# NN Transformer Training\nMy public notebook transformer has 2 layers with skip connections (to make training easy). When using knowledge distillation, we can train a deeper transformer successfully. My final solution uses a 4 layer transformer without skip connections. We also added a GRU layer after transformer blocks and before final classification layers.\n\nTraining is done using 4 cycles of cosine learning schedule. In the first cold start cosine cycle, we pretrain (i.e. Knowldege Distillation) the Transformer using concatenated rows of both LGBM OOF preds and LGBM test preds and leave probabilities between 0 and 1 (i.e. soft labels). During the second cosine cycle, we use a warm start, reduce the learning rate and train with the hard (0 or 1) train targets. For the third and fourth cycle, we repeat cycles one and two.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/schedule.png)\n\n# Model Performance\nWhen creating our submission.csv file from our two models, we can use the normal **K-Fold** LGBM OOF preds and normal LGBM test preds. So making a submission is fast an easy. Additionally, we average 5 seeds per model (and slight model variations) for improved performance.\n\nTo tune our two models and compute optimal hyperparameters, we need a leak free reliable CV score. Leak-free CV score is created using **Nested K-Fold** CV. We divide each of 10 outer folds into 10 inner folds. We then train 100 models using GBT. Then each 1 of 10 outer folds has its unique OOF preds and unique test preds. These individualized OOF and test preds are created using only train targets from within the corresponding outer fold train data.\n\nWhen computing leak free CV score, we find that our NN Transformer has **CV 0.798 / LB 0.799**, our LGBM has **CV 0.799 / LB 0.799** and our 50%/50% ensemble has **CV 0.800 / LB 0.801**.\n\n# Fast Experimentation\nFast experimention was done using GPU. Thank you [Nvidia][5] for providing me compute resources for this competition. Experiments were accelerated using 4xV100 32GB.\n\nFeature engineering was explored using [RAPIDS cuDF][3] which performs operations like dataframe groupby aggregation on GPU 10-100x faster than using CPU. Many GBT experimental models were trained and evaluated using fast GPU XGB. With 1xV100 GPU, XGB can train 100 models for Nested 10 in 10 K-Fold (i.e. 100 models) on full data in only 2 hours. \n\nFeature selection was performed using both XGB feature importance and permutation importance. Using [RAPIDS FIL][4], we can perform permutation importance where we randomly shuffle each feature column 10 times for each of 10 folds (and average 100 results) in blazing speed! \n\nEach of 1000s of feature columns, we infer 100 times. This is a total of 100,000s of model inferences where each model is 1000s of individual trees! Using [RAPIDs FIL][4], we can perform this quickly! Note that we can even take an existing CPU LGBM Dart model and convert it into a GPU [RAPIDS FIL][4] inference model and perform permutation importance on existing LGBM Dart Models in blazing speed!\n\n[1]: https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\n[3]: https://rapids.ai/\n[4]: https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\n[5]: https://www.nvidia.com/en-us/",
      "votes": 272
    },
    {
      "id": 1912765,
      "postDate": "2022-08-25T00:49:53.133Z",
      "content": "<p>Awesome solution Chris, congrats!</p>\n<p>I'm curious about the data preparation you did for the transformer - was it the same as what you use in your public notebook? I spent a lot of time trying to trying to figure out the best way to represent missing statements, whether masking would help, etc. Got some boosts from imputing all the nulls with lgbm and using data augmentation (shift statements forward 1), but could never do better than .793 CV.</p>",
      "rawMarkdown": "Awesome solution Chris, congrats!\n\nI'm curious about the data preparation you did for the transformer - was it the same as what you use in your public notebook? I spent a lot of time trying to trying to figure out the best way to represent missing statements, whether masking would help, etc. Got some boosts from imputing all the nulls with lgbm and using data augmentation (shift statements forward 1), but could never do better than .793 CV.",
      "votes": 5,
      "replies": [
        {
          "id": 1912801,
          "postDate": "2022-08-25T01:16:44.143Z",
          "content": "<p>Hi Joe, congratulations to you and team for achieving 24th out of 5000 teams. That's great.</p>\n<p>I prepared data exactly like i did in my public notebook (with new targets of test preds and oof preds). I also tried many alternatives like adding padding correctly in time. (My public notebook adds all pads in front of sequence for customer with missing statements). I also tried different ways to represent NA and padding. I also tried adding position embedding. I did lots of things but the data representation didn't make much difference.</p>\n<p>I only made 4 change to my public notebook</p>\n<ul>\n<li><code>emb = layers.Embedding(10,4)</code> changed to <code>(10,8)</code></li>\n<li>after transformer blocks added <code>x = tf.keras.layers.GRU(units=128, return_sequences=False)(x)</code> before dense layers</li>\n<li>for submission.csv, average multiple random seed trained transformers with 2, 3, 4 blocks and remove skipped connections</li>\n<li>trained for multiple cosine cycles with varying learning rates.</li>\n</ul>\n<p>The big help was pretraining with test predictions. This allowed us to use all the future test data features that Amex gave us and helped the model train it's attention and other stuff. This boosted CV score from 0.790 to 0.798</p>\n<p>(It's like NLP unsupervised pretraining. Even without labels, just training a transformer on lots text helps the model understand language. And in this comp, our transformer will learn more about credit cards by seeing all the test features even without correct test labels)</p>",
          "rawMarkdown": "Hi Joe, congratulations to you and team for achieving 24th out of 5000 teams. That's great.\n\nI prepared data exactly like i did in my public notebook (with new targets of test preds and oof preds). I also tried many alternatives like adding padding correctly in time. (My public notebook adds all pads in front of sequence for customer with missing statements). I also tried different ways to represent NA and padding. I also tried adding position embedding. I did lots of things but the data representation didn't make much difference.\n\nI only made 4 change to my public notebook\n* `emb = layers.Embedding(10,4)` changed to `(10,8)`\n* after transformer blocks added `x = tf.keras.layers.GRU(units=128, return_sequences=False)(x)` before dense layers\n* for submission.csv, average multiple random seed trained transformers with 2, 3, 4 blocks and remove skipped connections\n* trained for multiple cosine cycles with varying learning rates.\n\nThe big help was pretraining with test predictions. This allowed us to use all the future test data features that Amex gave us and helped the model train it's attention and other stuff. This boosted CV score from 0.790 to 0.798\n\n(It's like NLP unsupervised pretraining. Even without labels, just training a transformer on lots text helps the model understand language. And in this comp, our transformer will learn more about credit cards by seeing all the test features even without correct test labels)",
          "votes": 11
        },
        {
          "id": 1912804,
          "postDate": "2022-08-25T01:19:17.463Z",
          "content": "<p>Maybe, if you change the percentage of the components of the target such as 70% (0) and 30% (1) before using splitting cv by 5 folds, you will get 0.795. this is <a href=\"https://www.kaggle.com/youneseloiarm/xgboost-starter-30-70\" target=\"_blank\">XGBoost Starter -30%-70%</a></p>",
          "rawMarkdown": "Maybe, if you change the percentage of the components of the target such as 70% (0) and 30% (1) before using splitting cv by 5 folds, you will get 0.795. this is [XGBoost Starter -30%-70%](https://www.kaggle.com/youneseloiarm/xgboost-starter-30-70)",
          "votes": 1
        },
        {
          "id": 1912842,
          "postDate": "2022-08-25T01:56:24.473Z",
          "content": "<p>Thanks for such a detailed (and fast!) response. Using the test data for a performance edge makes a ton of sense. A few days ago I had the idea to try transfer learning on P2 - pretrain a transformer on train+test to predict P2 before fine tuning on train - couldn't get a boost out of it, but wish I had had the idea earlier on to experiment with more thoroughly and possibly push to a better approach like this one.</p>",
          "rawMarkdown": "Thanks for such a detailed (and fast!) response. Using the test data for a performance edge makes a ton of sense. A few days ago I had the idea to try transfer learning on P2 - pretrain a transformer on train+test to predict P2 before fine tuning on train - couldn't get a boost out of it, but wish I had had the idea earlier on to experiment with more thoroughly and possibly push to a better approach like this one.",
          "votes": 3
        },
        {
          "id": 1912890,
          "postDate": "2022-08-25T02:48:52.590Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2212438,
      "postDate": "2023-04-06T19:06:47.257Z",
      "content": "<p>Thanks for sharing this  solution. It is very much helpfull for fresh data sientist as me. </p>\n<p>I couldnt understand what NN is? I got that lgbm is a machine learning model and ı also apply this for many competitions. But ı couldnt understand NN and how you can combine with lgbm ?</p>",
      "rawMarkdown": "Thanks for sharing this  solution. It is very much helpfull for fresh data sientist as me. \n\nI couldnt understand what NN is? I got that lgbm is a machine learning model and ı also apply this for many competitions. But ı couldnt understand NN and how you can combine with lgbm ?",
      "votes": 1,
      "replies": [
        {
          "id": 2212461,
          "postDate": "2023-04-06T19:18:41.820Z",
          "content": "<p>NN is an abbreviation for neural network. It refers to transformers, CNN (convolution neural networks), RNN (recursive neural networks like LSTM and GRU), MLP (multi-layer perceptron), etc</p>\n<p>People also use NN to refer to DL (i.e. deep learning solutions)</p>",
          "rawMarkdown": "NN is an abbreviation for neural network. It refers to transformers, CNN (convolution neural networks), RNN (recursive neural networks like LSTM and GRU), MLP (multi-layer perceptron), etc\n\nPeople also use NN to refer to DL (i.e. deep learning solutions)",
          "votes": 1,
          "replies": [
            {
              "id": 2212470,
              "postDate": "2023-04-06T19:29:27.463Z",
              "content": "<p>I couldn't quite get some things in my head. Thank you very much for your quick and descriptive reply Chris!<br>\nBut it is confusing for me that how did you combine lightgbm and nn in one?</p>",
              "rawMarkdown": "I couldn't quite get some things in my head. Thank you very much for your quick and descriptive reply Chris!\nBut it is confusing for me that how did you combine lightgbm and nn in one?",
              "votes": 1
            },
            {
              "id": 2212488,
              "postDate": "2023-04-06T19:44:41.923Z",
              "content": "<p>I stacked the two models. I used the predictions from LGBM as inputs to NN</p>",
              "rawMarkdown": "I stacked the two models. I used the predictions from LGBM as inputs to NN",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 1914862,
      "postDate": "2022-08-26T13:40:55.723Z",
      "content": "<p>Congrats with 17th place and thanks for sharing your work. If I understand correctly we use distillation of knowledge from larger model to smaller while you did from smaller to larger (from LGBM to Transformer) or am I wrong somewhere? And what exactly gave distillation in this case, why couldn't you just make an ensemble of two models without distillation?</p>",
      "rawMarkdown": "Congrats with 17th place and thanks for sharing your work. If I understand correctly we use distillation of knowledge from larger model to smaller while you did from smaller to larger (from LGBM to Transformer) or am I wrong somewhere? And what exactly gave distillation in this case, why couldn't you just make an ensemble of two models without distillation?",
      "votes": 3,
      "replies": [
        {
          "id": 1916064,
          "postDate": "2022-08-27T15:46:08.570Z",
          "content": "<p>Transformers are hard to train from scratch. The model needs to learn weights for its self attention and weights for all its layers. When we train with soft targets from another trained model, it provides more information than just using the original train hard targets.</p>\n<p>The original hard targets are 0's and 1's. But the other model's soft targets are 0.1, 0.2, 0.3, …, 0.9, 1.0 and all numbers in between. This contains much more information for the NN Transformer to learn with. So the knowledge distillation from LGBM helps the NN learn. Furthermore, NN is fundamentally different than LGBM, so once the knowledge is transferred, it will be \"represented differently\" and add it's own \"personality\" to the predictions and make it effective in ensemble.</p>\n<p>The biggest reason why this approach worked so well is that it utilizes the 11 million rows of unlabeled test data. Even though we don't have labels for test data, we still have 1 million time series of length 13 for the different features in test data. This teaches our models how the features change over time and helps the model predict what the features will be in the future. If we never use test data, we never gain access to these 1 million time series with all its information.</p>\n<p>Take feature <code>P_2</code> for example. This is like a customer's credit score. If we knew this value for 18 months into the future after the last credit card statement, we could probably predict every customer's default perfectly. What better way to guess what <code>P_2</code> will be in the future than watching the 1 million time series of length 13 in the test data to learn how it changes over time. And using that to train our model.</p>",
          "rawMarkdown": "Transformers are hard to train from scratch. The model needs to learn weights for its self attention and weights for all its layers. When we train with soft targets from another trained model, it provides more information than just using the original train hard targets.\n\nThe original hard targets are 0's and 1's. But the other model's soft targets are 0.1, 0.2, 0.3, ..., 0.9, 1.0 and all numbers in between. This contains much more information for the NN Transformer to learn with. So the knowledge distillation from LGBM helps the NN learn. Furthermore, NN is fundamentally different than LGBM, so once the knowledge is transferred, it will be \"represented differently\" and add it's own \"personality\" to the predictions and make it effective in ensemble.\n\nThe biggest reason why this approach worked so well is that it utilizes the 11 million rows of unlabeled test data. Even though we don't have labels for test data, we still have 1 million time series of length 13 for the different features in test data. This teaches our models how the features change over time and helps the model predict what the features will be in the future. If we never use test data, we never gain access to these 1 million time series with all its information.\n\nTake feature `P_2` for example. This is like a customer's credit score. If we knew this value for 18 months into the future after the last credit card statement, we could probably predict every customer's default perfectly. What better way to guess what `P_2` will be in the future than watching the 1 million time series of length 13 in the test data to learn how it changes over time. And using that to train our model.",
          "votes": 16
        }
      ]
    },
    {
      "id": 1913464,
      "postDate": "2022-08-25T10:33:46.613Z",
      "content": "<p>Congrats for the medal and thanks for sharing all that insightful code.</p>",
      "rawMarkdown": "Congrats for the medal and thanks for sharing all that insightful code.",
      "votes": 3,
      "replies": [
        {
          "id": 1916449,
          "postDate": "2022-08-27T22:44:11.500Z",
          "content": "<p>Thanks Lucas. Thanks for your helpful notebooks and discussions.</p>",
          "rawMarkdown": "Thanks Lucas. Thanks for your helpful notebooks and discussions.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1913394,
      "postDate": "2022-08-25T10:03:14.207Z",
      "content": "<p>What exactly do you mean by knowledge distillation? How do you extract it from the LGBM? Is it simply stacking its prediction?<br>\nAmazing work btw and congrats! </p>",
      "rawMarkdown": "What exactly do you mean by knowledge distillation? How do you extract it from the LGBM? Is it simply stacking its prediction?\nAmazing work btw and congrats! ",
      "votes": 3,
      "replies": [
        {
          "id": 1913527,
          "postDate": "2022-08-25T10:53:18.960Z",
          "content": "<p>Knowledge distillation is the transfer of knowledge from 1 model to another model. First we train an LGBM. Next we perform teacher student. We use the predictions (both OOF and test preds) of LGBM to train an NN Transformer. Finally we finetune NN Transformer on the true train targets.</p>",
          "rawMarkdown": "Knowledge distillation is the transfer of knowledge from 1 model to another model. First we train an LGBM. Next we perform teacher student. We use the predictions (both OOF and test preds) of LGBM to train an NN Transformer. Finally we finetune NN Transformer on the true train targets.",
          "votes": 7
        },
        {
          "id": 1913800,
          "postDate": "2022-08-25T14:08:48.057Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> interested to know more how this can work. <br>\nBy any chance are you publishing your code? That would be great</p>",
          "rawMarkdown": "Thanks @cdeotte interested to know more how this can work. \nBy any chance are you publishing your code? That would be great",
          "votes": 2
        },
        {
          "id": 1914130,
          "postDate": "2022-08-25T18:56:30.163Z",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, first of all, congrats and thanks for sharing your result, the second I have a silly question about <code>knowledge distillation</code>, when you said, <code>We use the predictions (both OOF and test preds) of LGBM to train an NN Transformer</code> do we use the test data directly into the training pipeline(and use the test preds from LGBM as labels)? If possible could you please refer to some materials where I can learn about this technique more?</p>",
          "rawMarkdown": "Hi, @cdeotte, first of all, congrats and thanks for sharing your result, the second I have a silly question about `knowledge distillation`, when you said, `We use the predictions (both OOF and test preds) of LGBM to train an NN Transformer` do we use the test data directly into the training pipeline(and use the test preds from LGBM as labels)? If possible could you please refer to some materials where I can learn about this technique more?",
          "votes": 3
        },
        {
          "id": 1917651,
          "postDate": "2022-08-29T00:28:35.233Z",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> , we use test data directly into the training pipeline (and use test preds from LGBM as labels). Note that we leave the test preds as is (as continuous numbers between 0 and 1) without converting them to 0's and 1's. This is a combination of both \"pseudo labeling\" and \"knowledge distillation\". You can google these two terms to learn more.</p>\n<p>Knowledge disllation helps the Transformer learn more accurately and quickly (by gaining the knowledge that LGBM has already learned but storing it in a new transformer way). And pseudo labeling allows us to access all the information in the unlabeled test data. Namely the features which are 1 million time series of length 13 from the 188 features. This is very useful information that Kaggle provided us!</p>",
          "rawMarkdown": "Yes @susnato , we use test data directly into the training pipeline (and use test preds from LGBM as labels). Note that we leave the test preds as is (as continuous numbers between 0 and 1) without converting them to 0's and 1's. This is a combination of both \"pseudo labeling\" and \"knowledge distillation\". You can google these two terms to learn more.\n\nKnowledge disllation helps the Transformer learn more accurately and quickly (by gaining the knowledge that LGBM has already learned but storing it in a new transformer way). And pseudo labeling allows us to access all the information in the unlabeled test data. Namely the features which are 1 million time series of length 13 from the 188 features. This is very useful information that Kaggle provided us!",
          "votes": 3
        }
      ]
    },
    {
      "id": 1912927,
      "postDate": "2022-08-25T03:47:17.970Z",
      "content": "<p>Congratulations on great finish and sharing the solution Chris. I wonder how well was your NN transformer cv-lb aligned after using Knowledge Distillation based on soft preds? I usually see a dramatic change in CV, but lesser improvement in Leaderboard, couldn't think of any possible way of leakage. </p>",
      "rawMarkdown": "Congratulations on great finish and sharing the solution Chris. I wonder how well was your NN transformer cv-lb aligned after using Knowledge Distillation based on soft preds? I usually see a dramatic change in CV, but lesser improvement in Leaderboard, couldn't think of any possible way of leakage. ",
      "votes": 3,
      "replies": [
        {
          "id": 1913576,
          "postDate": "2022-08-25T11:31:51.267Z",
          "content": "<p>Congratulations Nischay on your fantastic solo performance.</p>\n<p>My NN Transformer has a better private LB than LGBM below are the stats</p>\n<ul>\n<li>NN Transformer, leak-free-CV 0.7980, Public LB 0.7988, Private LB 0.8076</li>\n<li>LGBM, CV 0.7990, Public LB 0.7992, Private LB 0.8070</li>\n</ul>\n<p>===== Below is more info =====</p>\n<p>The NN Transformer CV is accurate and leak free. Because the LGBM is trained using 10 (inner) folds within 10 (outer) folds (i.e. Nested K-Fold). The outer folds are typical K-Fold. When training outer fold 1 we use train targets from folds 2 thru 10. We then split this group of 9 folds into its own 10 folds. Each of these inner 10 folds we train 1 model. These inner models never see the targets from validation outer fold 1. Using these inner models, we create OOF for outer fold 1 which gives us leak free predictions for outer folds 2-10. We also use these inner models to predict test preds. (Next we train 10 inner folds for outer fold 2, etc etc. In total we train 100 models)</p>\n<p>This gives us a unique <code>lstm_OOF_outer_fold_1.csv</code> and <code>lstm_Test_preds_outer_fold_1.csv</code> for each of 10 outer folds for a total of 20 CSV. When we train fold 1 of our NN Transformer, we use these two CSV. When we train fold 2 of our NN Transformer we use <code>lstm_OOF_outer_fold_2.csv</code> and <code>lstm_Test_preds_outer_fold_2.csv</code> etc etc. This prevents CV leaks.</p>",
          "rawMarkdown": "Congratulations Nischay on your fantastic solo performance.\n\nMy NN Transformer has a better private LB than LGBM below are the stats\n* NN Transformer, leak-free-CV 0.7980, Public LB 0.7988, Private LB 0.8076\n* LGBM, CV 0.7990, Public LB 0.7992, Private LB 0.8070\n\n===== Below is more info =====\n\nThe NN Transformer CV is accurate and leak free. Because the LGBM is trained using 10 (inner) folds within 10 (outer) folds (i.e. Nested K-Fold). The outer folds are typical K-Fold. When training outer fold 1 we use train targets from folds 2 thru 10. We then split this group of 9 folds into its own 10 folds. Each of these inner 10 folds we train 1 model. These inner models never see the targets from validation outer fold 1. Using these inner models, we create OOF for outer fold 1 which gives us leak free predictions for outer folds 2-10. We also use these inner models to predict test preds. (Next we train 10 inner folds for outer fold 2, etc etc. In total we train 100 models)\n\nThis gives us a unique `lstm_OOF_outer_fold_1.csv` and `lstm_Test_preds_outer_fold_1.csv` for each of 10 outer folds for a total of 20 CSV. When we train fold 1 of our NN Transformer, we use these two CSV. When we train fold 2 of our NN Transformer we use `lstm_OOF_outer_fold_2.csv` and `lstm_Test_preds_outer_fold_2.csv` etc etc. This prevents CV leaks.",
          "votes": 14
        },
        {
          "id": 1914330,
          "postDate": "2022-08-26T02:22:32.277Z",
          "content": "<p>I never thought of this possible leak free approach, that's something extraordinary. Thanks a lot for clearing up everything  🙇🙇</p>",
          "rawMarkdown": "I never thought of this possible leak free approach, that's something extraordinary. Thanks a lot for clearing up everything  🙇🙇",
          "votes": 3
        }
      ]
    },
    {
      "id": 1912919,
      "postDate": "2022-08-25T03:36:28.050Z",
      "content": "<p>Thank you for sharing this. There has been knowledge distillation in the Feedback competition as well (ending yesterday), something I really need to catch up.  Am I right to say that it's a student teacher model , where the student is trained with the prediction result of the teacher instead of the hard label ? <br>\nCongratulations for your result, and thanks again for all the sharings. </p>",
      "rawMarkdown": "Thank you for sharing this. There has been knowledge distillation in the Feedback competition as well (ending yesterday), something I really need to catch up.  Am I right to say that it's a student teacher model , where the student is trained with the prediction result of the teacher instead of the hard label ? \nCongratulations for your result, and thanks again for all the sharings. ",
      "votes": 3,
      "replies": [
        {
          "id": 1912925,
          "postDate": "2022-08-25T03:45:14.560Z",
          "content": "<p>Yes, exactly. It's very simple. It's like pseudo labels. Just save your OOF and test predictions. Then train your new model using the probabilities predictions from OOF and test preds using cross entropy. The loss cross entropy works when the targets are continuous preditcions between 0 and 1 (i.e. the targets targets don't need to be zeros and ones). </p>\n<p>Imagine you have two models, model A and model B. First train model A. Then makes predictions with model A on some dataset. Next train model B using the probabilities from model A and cross entropy (on that dataset). Then model B has learned knowledge distillation from model A.</p>\n<p>Afterward, you can either stop there or fine tune model B on more data. In this competition, I further trained on train data targets which are zeros and ones.</p>",
          "rawMarkdown": "Yes, exactly. It's very simple. It's like pseudo labels. Just save your OOF and test predictions. Then train your new model using the probabilities predictions from OOF and test preds using cross entropy. The loss cross entropy works when the targets are continuous preditcions between 0 and 1 (i.e. the targets targets don't need to be zeros and ones). \n\nImagine you have two models, model A and model B. First train model A. Then makes predictions with model A on some dataset. Next train model B using the probabilities from model A and cross entropy (on that dataset). Then model B has learned knowledge distillation from model A.\n\nAfterward, you can either stop there or fine tune model B on more data. In this competition, I further trained on train data targets which are zeros and ones.",
          "votes": 17
        },
        {
          "id": 1913006,
          "postDate": "2022-08-25T05:06:18.980Z",
          "content": "<p>Thanks a lot for explaining. Can we use similar approach for regression tasks as well.</p>",
          "rawMarkdown": "Thanks a lot for explaining. Can we use similar approach for regression tasks as well.",
          "votes": 1
        },
        {
          "id": 1913173,
          "postDate": "2022-08-25T07:42:45.963Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I get it what are you trying to explain, but isn't Pseudo labelling exactly this? then what's the difference between distillation and pseudo labelling?</p>",
          "rawMarkdown": "@cdeotte I get it what are you trying to explain, but isn't Pseudo labelling exactly this? then what's the difference between distillation and pseudo labelling?",
          "votes": 1
        },
        {
          "id": 1913763,
          "postDate": "2022-08-25T13:47:57.253Z",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> Below is my opinion, it might not be fully correct.</p>\n<p>Pseudo labeling is the process of assigning labels to unlabeled data (so that we can train with and extract info from the unlabeled data). Generally pseudo labels are converted into hard targets, i.e. 0's and 1's and only confident predictions are used. For example, after making predictions on unlabeled data we keep all samples with confident predictions less than 0.05 or greater than 0.95. Then round the predictions to 0 and 1. This allows us to perform supervised training on the previously unlabeled data which gives the model the benefit of using the unlabeled data's feature values.</p>\n<p>Knowledge distillation is the process of transferring one model's (or ensemble's) learning to another model (frequently used to transfer the performance of a complicated ensemble into a simple single model for purpose of efficient production inference). To facilitate the transfer, we can use any data whether it was originally labeled or not labeled. The original labels are discarded and the teacher model predicts new labels. These predicted labels are not converted to 0's and 1's but rather kept as probability values between 0 and 1. This allows for maximum transfer of the teacher's knowledge to the student model.</p>\n<p>We can also perform hybrids of the two. Like keeping pseudo labels soft. Then we simultaneously transfer knowledge from a teacher to a student and extract info from unlabeled data's features. My solution here is probably a combination of both.</p>",
          "rawMarkdown": "@mrinath Below is my opinion, it might not be fully correct.\n\nPseudo labeling is the process of assigning labels to unlabeled data (so that we can train with and extract info from the unlabeled data). Generally pseudo labels are converted into hard targets, i.e. 0's and 1's and only confident predictions are used. For example, after making predictions on unlabeled data we keep all samples with confident predictions less than 0.05 or greater than 0.95. Then round the predictions to 0 and 1. This allows us to perform supervised training on the previously unlabeled data which gives the model the benefit of using the unlabeled data's feature values.\n\nKnowledge distillation is the process of transferring one model's (or ensemble's) learning to another model (frequently used to transfer the performance of a complicated ensemble into a simple single model for purpose of efficient production inference). To facilitate the transfer, we can use any data whether it was originally labeled or not labeled. The original labels are discarded and the teacher model predicts new labels. These predicted labels are not converted to 0's and 1's but rather kept as probability values between 0 and 1. This allows for maximum transfer of the teacher's knowledge to the student model.\n\nWe can also perform hybrids of the two. Like keeping pseudo labels soft. Then we simultaneously transfer knowledge from a teacher to a student and extract info from unlabeled data's features. My solution here is probably a combination of both.",
          "votes": 10
        },
        {
          "id": 1913989,
          "postDate": "2022-08-25T16:13:57.130Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> it makes sense, thanks</p>",
          "rawMarkdown": "@cdeotte it makes sense, thanks",
          "votes": 1
        }
      ]
    },
    {
      "id": 1912747,
      "postDate": "2022-08-25T00:32:34.440Z",
      "content": "<p>Congratulations, Chris, you are the real MVP! I wonder which features had most importance</p>",
      "rawMarkdown": "Congratulations, Chris, you are the real MVP! I wonder which features had most importance",
      "votes": 3
    },
    {
      "id": 1914213,
      "postDate": "2022-08-25T20:45:45.463Z",
      "content": "<p>Awesome approach really Amazing solution</p>",
      "rawMarkdown": "Awesome approach really Amazing solution",
      "votes": 4
    },
    {
      "id": 1932845,
      "postDate": "2022-09-10T03:09:11.483Z",
      "content": "<p>Thank you for sharing! Looks like using NN Transformer will exploit the information from historical data. </p>",
      "rawMarkdown": "Thank you for sharing! Looks like using NN Transformer will exploit the information from historical data. ",
      "votes": 1
    },
    {
      "id": 1928924,
      "postDate": "2022-09-06T17:43:51.033Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for sharing and for detailed explanation!</p>",
      "rawMarkdown": "Thank you @cdeotte for sharing and for detailed explanation!",
      "votes": 1
    },
    {
      "id": 1928319,
      "postDate": "2022-09-06T12:18:14.833Z",
      "content": "<p>Great Very Helpful</p>",
      "rawMarkdown": "Great Very Helpful",
      "votes": 1
    },
    {
      "id": 1928308,
      "postDate": "2022-09-06T12:12:01.890Z",
      "content": "<p><strong>Congratulations</strong> and thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>",
      "rawMarkdown": "**Congratulations** and thanks for sharing @cdeotte",
      "votes": 1
    },
    {
      "id": 1924679,
      "postDate": "2022-09-03T09:39:22.900Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> first of all congratulations of the win, can you tell me some resources for learning more about knowledge distillation, I tried looking it up myself and all I could understand that it is a form of model compression.</p>",
      "rawMarkdown": "Hi, @cdeotte first of all congratulations of the win, can you tell me some resources for learning more about knowledge distillation, I tried looking it up myself and all I could understand that it is a form of model compression.",
      "votes": 1
    },
    {
      "id": 1923445,
      "postDate": "2022-09-02T08:29:24.910Z",
      "content": "<p>Great solution, I wasn't aware we can do an ensemble of LightGBM and NNs, thanks</p>",
      "rawMarkdown": "Great solution, I wasn't aware we can do an ensemble of LightGBM and NNs, thanks",
      "votes": 1,
      "replies": [
        {
          "id": 1927728,
          "postDate": "2022-09-05T23:28:11.910Z",
          "content": "<p>The power of ensemble is using diverse models. Using both GBM and NN makes a great high performing ensemble!</p>",
          "rawMarkdown": "The power of ensemble is using diverse models. Using both GBM and NN makes a great high performing ensemble!"
        }
      ]
    },
    {
      "id": 1922339,
      "postDate": "2022-09-01T12:18:27.167Z",
      "content": "<p>Congratulations and thank you for sharing!<br>\nThis helps a lot!</p>",
      "rawMarkdown": "Congratulations and thank you for sharing!\nThis helps a lot!",
      "votes": 1
    },
    {
      "id": 1920709,
      "postDate": "2022-08-31T10:41:29.600Z",
      "content": "<p>great job.</p>",
      "rawMarkdown": "great job.",
      "votes": 1
    },
    {
      "id": 1919925,
      "postDate": "2022-08-30T18:40:20.373Z",
      "content": "<p>this is great</p>",
      "rawMarkdown": "this is great",
      "votes": 1
    },
    {
      "id": 1918821,
      "postDate": "2022-08-29T21:50:05.217Z",
      "content": "<p>Chris, ins't your rank 14th? Title says 15th place solution. I just realized that after final announcement of cheaters removal a couple of days ago, today again rank was gone up by 1. Not sure if they will keep changing the ranks every day.</p>",
      "rawMarkdown": "Chris, ins't your rank 14th? Title says 15th place solution. I just realized that after final announcement of cheaters removal a couple of days ago, today again rank was gone up by 1. Not sure if they will keep changing the ranks every day.",
      "votes": 1,
      "replies": [
        {
          "id": 1918950,
          "postDate": "2022-08-30T02:15:31.903Z",
          "content": "<p>Thank you, you are right. My rank is 14th now. I'll update the title soon.</p>",
          "rawMarkdown": "Thank you, you are right. My rank is 14th now. I'll update the title soon.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1918294,
      "postDate": "2022-08-29T13:23:50.233Z",
      "content": "<p>Amazing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> its really creative to use Knowledge distillation from models like Lightgbm to a Deep Learning model like Transformer !</p>",
      "rawMarkdown": "Amazing @cdeotte its really creative to use Knowledge distillation from models like Lightgbm to a Deep Learning model like Transformer !",
      "votes": 1,
      "replies": [
        {
          "id": 1923119,
          "postDate": "2022-09-02T02:37:16.910Z",
          "content": "<p>Thanks Athar!</p>",
          "rawMarkdown": "Thanks Athar!"
        }
      ]
    },
    {
      "id": 1917748,
      "postDate": "2022-08-29T03:23:17.030Z",
      "content": "<p>Difference between performance in private and public leaderboard for NN is really large，are codes completely the same？</p>",
      "rawMarkdown": "Difference between performance in private and public leaderboard for NN is really large，are codes completely the same？\n",
      "votes": 1,
      "replies": [
        {
          "id": 1917757,
          "postDate": "2022-08-29T03:31:18.503Z",
          "content": "<p>yes code is the same. It appears that all teams' models did <code>+0.007</code> better in private LB than public LB.</p>",
          "rawMarkdown": "yes code is the same. It appears that all teams' models did `+0.007` better in private LB than public LB.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1917648,
      "postDate": "2022-08-29T00:24:51.773Z",
      "content": "<p>Learned a lot! Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Learned a lot! Thanks for sharing @cdeotte ",
      "votes": 1
    },
    {
      "id": 1917148,
      "postDate": "2022-08-28T13:13:16.013Z",
      "content": "<p>Very Helpful 😃</p>",
      "rawMarkdown": "Very Helpful 😃",
      "votes": 1
    },
    {
      "id": 1914760,
      "postDate": "2022-08-26T11:48:37.260Z",
      "content": "<p>Great Job!</p>",
      "rawMarkdown": "Great Job!",
      "votes": 1
    },
    {
      "id": 1914649,
      "postDate": "2022-08-26T09:16:35.800Z",
      "content": "<p>Pro, thanks for sharing, i learned very much from this notebook!</p>",
      "rawMarkdown": "Pro, thanks for sharing, i learned very much from this notebook!",
      "votes": 1
    },
    {
      "id": 1914491,
      "postDate": "2022-08-26T06:20:38.957Z",
      "content": "<p>Congrats! Great approach and thanks for sharing</p>",
      "rawMarkdown": "Congrats! Great approach and thanks for sharing",
      "votes": 1
    },
    {
      "id": 1914489,
      "postDate": "2022-08-26T06:13:02.403Z",
      "content": "<p>Awesome ****</p>",
      "rawMarkdown": "Awesome ****",
      "votes": 1
    },
    {
      "id": 1914483,
      "postDate": "2022-08-26T06:01:07.447Z",
      "content": "<p>Awesome approach</p>",
      "rawMarkdown": "Awesome approach\n",
      "votes": 1
    },
    {
      "id": 1913875,
      "postDate": "2022-08-25T14:54:43.147Z",
      "content": "<p>Congrats and thank you for sharing the this. This is useful. :)</p>",
      "rawMarkdown": "Congrats and thank you for sharing the this. This is useful. :)",
      "votes": 1
    },
    {
      "id": 1913289,
      "postDate": "2022-08-25T09:13:43.023Z",
      "content": "<p>Truly impressive solution.</p>",
      "rawMarkdown": "Truly impressive solution.",
      "votes": 1
    },
    {
      "id": 1913253,
      "postDate": "2022-08-25T08:57:53.900Z",
      "content": "<p>Congrats and thanks for all the resource </p>",
      "rawMarkdown": "Congrats and thanks for all the resource ",
      "votes": 1
    },
    {
      "id": 1913189,
      "postDate": "2022-08-25T07:56:46.253Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Congratulations , combination of NN Transformer and Knowledge Distillation is really nice :D</p>",
      "rawMarkdown": "@cdeotte Congratulations , combination of NN Transformer and Knowledge Distillation is really nice :D",
      "votes": 1,
      "replies": [
        {
          "id": 1914063,
          "postDate": "2022-08-25T17:32:50.330Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> !</p>",
          "rawMarkdown": "Thanks @awsaf49 !",
          "votes": 2
        }
      ]
    },
    {
      "id": 1913185,
      "postDate": "2022-08-25T07:52:30.693Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! This was my first Kaggle competition and I learned a ton from your work. Thank you!</p>",
      "rawMarkdown": "Congrats @cdeotte! This was my first Kaggle competition and I learned a ton from your work. Thank you!",
      "votes": 1
    },
    {
      "id": 1912940,
      "postDate": "2022-08-25T04:07:01.907Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> on great solo gold medal. Ensemble NN (with Knowledge Distillation) + tree-based is powerful</p>",
      "rawMarkdown": "Congrats @cdeotte on great solo gold medal. Ensemble NN (with Knowledge Distillation) + tree-based is powerful",
      "votes": 1,
      "replies": [
        {
          "id": 1918729,
          "postDate": "2022-08-29T19:31:39.317Z",
          "content": "<p>Thanks KhanhVD!</p>",
          "rawMarkdown": "Thanks KhanhVD!"
        }
      ]
    },
    {
      "id": 1912889,
      "postDate": "2022-08-25T02:48:32.793Z",
      "content": "<p>haha，I trained LGBM using  NN Transformer Knowledge Distillation, but got a poor result</p>",
      "rawMarkdown": "haha，I trained LGBM using  NN Transformer Knowledge Distillation, but got a poor result",
      "votes": 1,
      "replies": [
        {
          "id": 1913523,
          "postDate": "2022-08-25T10:51:50.077Z",
          "content": "<p>prolly need more GPU :p</p>",
          "rawMarkdown": "prolly need more GPU :p",
          "votes": 1
        },
        {
          "id": 1913766,
          "postDate": "2022-08-25T13:52:21.510Z",
          "content": "<p>Cool idea <a href=\"https://www.kaggle.com/bjjiang\" target=\"_blank\">@bjjiang</a> </p>",
          "rawMarkdown": "Cool idea @bjjiang "
        }
      ]
    },
    {
      "id": 1912885,
      "postDate": "2022-08-25T02:46:16.313Z",
      "content": "<p>Truly impressive solution. I wonder if I will ever be able to do something like that. For now I'm ok if I just keep learning from you. Congrats, Chris.</p>",
      "rawMarkdown": "Truly impressive solution. I wonder if I will ever be able to do something like that. For now I'm ok if I just keep learning from you. Congrats, Chris.",
      "votes": 1
    },
    {
      "id": 1912824,
      "postDate": "2022-08-25T01:38:21.270Z",
      "content": "<p>Congratulations and a really Nice innovative approach <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Chris and thanks for sharing! </p>",
      "rawMarkdown": "Congratulations and a really Nice innovative approach @cdeotte Chris and thanks for sharing! ",
      "votes": 1
    },
    {
      "id": 1912813,
      "postDate": "2022-08-25T01:24:51.713Z",
      "content": "<p>Congrat <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! Thanks for sharing your approach. Very insightful for someone like myself who is starting his kaggle career.</p>",
      "rawMarkdown": "Congrat @cdeotte! Thanks for sharing your approach. Very insightful for someone like myself who is starting his kaggle career.",
      "votes": 1
    },
    {
      "id": 1912812,
      "postDate": "2022-08-25T01:24:46.177Z",
      "content": "<p>That's so <br>\nAmazing!</p>",
      "rawMarkdown": "That's so \nAmazing!",
      "votes": 1
    },
    {
      "id": 1912777,
      "postDate": "2022-08-25T00:59:13.013Z",
      "content": "<p>congrats～～～</p>",
      "rawMarkdown": "congrats～～～",
      "votes": 1
    },
    {
      "id": 1912770,
      "postDate": "2022-08-25T00:53:21.617Z",
      "content": "<p>🎉 Woot woot! Congratulations, and thanks for all the great notebooks, learned a lot </p>",
      "rawMarkdown": "🎉 Woot woot! Congratulations, and thanks for all the great notebooks, learned a lot ",
      "votes": 1
    },
    {
      "id": 1912763,
      "postDate": "2022-08-25T00:46:15.267Z",
      "content": "<p>Congratulations Chris!</p>",
      "rawMarkdown": "Congratulations Chris!",
      "votes": 1
    },
    {
      "id": 1912761,
      "postDate": "2022-08-25T00:45:24.253Z",
      "content": "<p>Fantastic write-up, thanks for sharing. I used a somewhat similar process, just not as effectively: </p>\n<ul>\n<li>XG Boost on GPU for quick prototyping and feature importance </li>\n<li>A Variety of seeds and folds blended via the GBM Model</li>\n<li>I only tried GBDT once…. it did \"okay\" blended with DART on the public, but was my 2nd best model on the private. </li>\n</ul>\n<p>Didn't do the permutation importance, or knowledge distillation or leverage a completely different model like the NN Transformer for my ensembles, nor train 1000s of models. Will definitely add those into my bag of tricks for next time. Knowledge distillation will definitely be part of my weekend reading.   </p>",
      "rawMarkdown": "Fantastic write-up, thanks for sharing. I used a somewhat similar process, just not as effectively: \n* XG Boost on GPU for quick prototyping and feature importance \n* A Variety of seeds and folds blended via the GBM Model\n* I only tried GBDT once.... it did \"okay\" blended with DART on the public, but was my 2nd best model on the private. \n\nDidn't do the permutation importance, or knowledge distillation or leverage a completely different model like the NN Transformer for my ensembles, nor train 1000s of models. Will definitely add those into my bag of tricks for next time. Knowledge distillation will definitely be part of my weekend reading.   \n\n",
      "votes": 1
    },
    {
      "id": 1912756,
      "postDate": "2022-08-25T00:38:43.770Z",
      "content": "<p>Congratulations Chris! Amazing as ever! <br>\nI had much difficulty performing permutation importance with Dart LGBM. It is amazing to learn there is actually a way to make it work. How much improvement did <code>RAPIDS FIL</code> provide? Is it compatible with all sorts of tree models?</p>",
      "rawMarkdown": "Congratulations Chris! Amazing as ever! \nI had much difficulty performing permutation importance with Dart LGBM. It is amazing to learn there is actually a way to make it work. How much improvement did `RAPIDS FIL` provide? Is it compatible with all sorts of tree models?",
      "votes": 1,
      "replies": [
        {
          "id": 1912819,
          "postDate": "2022-08-25T01:33:38.957Z",
          "content": "<p>Hi Tonghu, Congratuations to you and team on your great finish!</p>\n<p><code>RAPIDS FIL</code> is a library that speeds up inference. The idea is that you can train your forest wtih XGB or LGBM or Sklearn Random Forest etc. Then we load the saved model with <code>RAPIDS FIL</code> and RAPIDS FIL will infer 10x-100x faster than using XGB, LGBM, Sklearn Random Forest.</p>\n<p>This helps if a company deploys a model in production. And in this competition, it helps us perform permutation importance. To compute permutation importance, we must repeatedly infer a model over and over. If we load a saved LGBM Dart on CPU then inferring over and over will takes minutes or hours. If we load a saved LGBM Dart using RAPIDS FIL unto GPU, then we can infer over and over in seconds!</p>",
          "rawMarkdown": "Hi Tonghu, Congratuations to you and team on your great finish!\n\n`RAPIDS FIL` is a library that speeds up inference. The idea is that you can train your forest wtih XGB or LGBM or Sklearn Random Forest etc. Then we load the saved model with `RAPIDS FIL` and RAPIDS FIL will infer 10x-100x faster than using XGB, LGBM, Sklearn Random Forest.\n\nThis helps if a company deploys a model in production. And in this competition, it helps us perform permutation importance. To compute permutation importance, we must repeatedly infer a model over and over. If we load a saved LGBM Dart on CPU then inferring over and over will takes minutes or hours. If we load a saved LGBM Dart using RAPIDS FIL unto GPU, then we can infer over and over in seconds!",
          "votes": 8
        },
        {
          "id": 1912823,
          "postDate": "2022-08-25T01:38:18.830Z",
          "content": "<p>After saving an LGBM Dart model (with LGBM's save feature not joblib nor pickle), we can use </p>\n<pre><code>import cuml\ncuml.ForestInference.load('model.txt', model_type='lightgbm')\nval_pred = model.predict(x_val)\n</code></pre>\n<p>And <code>x_val</code> can be a RAPIDS cuDF dataframe on GPU. The documentation for FIL is <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "After saving an LGBM Dart model (with LGBM's save feature not joblib nor pickle), we can use \n\n    import cuml\n    cuml.ForestInference.load('model.txt', model_type='lightgbm')\n    val_pred = model.predict(x_val)\n\nAnd `x_val` can be a RAPIDS cuDF dataframe on GPU. The documentation for FIL is [here][1]\n\n[1]: https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing",
          "votes": 10
        },
        {
          "id": 1912899,
          "postDate": "2022-08-25T03:07:14.053Z",
          "content": "<p>Sounds great! I will definitely try it out in following competitions! By the way, thanks for this great sharing, learned a lot from it.</p>",
          "rawMarkdown": "Sounds great! I will definitely try it out in following competitions! By the way, thanks for this great sharing, learned a lot from it.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1912749,
      "postDate": "2022-08-25T00:34:20.040Z",
      "content": "<p>Congrats Chris! And thanks for your GRU starter. I applied NN first time in this competition.<br>\nI did ensemble my own LGBs, XGBs and a GRU. My GRU was very poor, so not a great boost.</p>",
      "rawMarkdown": "Congrats Chris! And thanks for your GRU starter. I applied NN first time in this competition.\nI did ensemble my own LGBs, XGBs and a GRU. My GRU was very poor, so not a great boost.",
      "votes": 1,
      "replies": [
        {
          "id": 1912757,
          "postDate": "2022-08-25T00:40:17.340Z",
          "content": "<p>I observe that a best GRU can get a great boost in public and a best transformer can do it in private.</p>",
          "rawMarkdown": "I observe that a best GRU can get a great boost in public and a best transformer can do it in private.",
          "votes": 1
        },
        {
          "id": 1912809,
          "postDate": "2022-08-25T01:23:48.847Z",
          "content": "<p>Congrats <a href=\"https://www.kaggle.com/gogogopp\" target=\"_blank\">@gogogopp</a> and <a href=\"https://www.kaggle.com/jacksonyou\" target=\"_blank\">@jacksonyou</a> for achieving Silver! You guys did great!</p>\n<p>Yes ensembling with GRU and/or Transformer adds a lot of diversity and helps public and private LB. </p>",
          "rawMarkdown": "Congrats @gogogopp and @jacksonyou for achieving Silver! You guys did great!\n\nYes ensembling with GRU and/or Transformer adds a lot of diversity and helps public and private LB. ",
          "votes": 3
        },
        {
          "id": 1912832,
          "postDate": "2022-08-25T01:47:26.023Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I've been following you from Malware, Santander, IEEE-fraud and in Amex now. </p>\n<p>Got 2 solo silvers after learning from your fantastic starters. But it is hard to achieve gold in solo, need 1 to become Master ;)</p>",
          "rawMarkdown": "@cdeotte I've been following you from Malware, Santander, IEEE-fraud and in Amex now. \n\nGot 2 solo silvers after learning from your fantastic starters. But it is hard to achieve gold in solo, need 1 to become Master ;)",
          "votes": 2
        },
        {
          "id": 1912847,
          "postDate": "2022-08-25T02:06:20.950Z",
          "content": "<p><a href=\"https://www.kaggle.com/gogogopp\" target=\"_blank\">@gogogopp</a> You'll get Kaggle competition master soon! I checked out your achievements. You are doing fantastic solo. you got 2 solo silvers and 1 solo bronze. That is very difficult. Great job!</p>\n<p>To become Kaggle competition master, you will need 1 gold. You should consider teaming up in your next competition. That will help you get gold because you can ensemble your model with your teammates' models. For Kaggle competition master, you don't need a solo gold.</p>",
          "rawMarkdown": "@gogogopp You'll get Kaggle competition master soon! I checked out your achievements. You are doing fantastic solo. you got 2 solo silvers and 1 solo bronze. That is very difficult. Great job!\n\nTo become Kaggle competition master, you will need 1 gold. You should consider teaming up in your next competition. That will help you get gold because you can ensemble your model with your teammates' models. For Kaggle competition master, you don't need a solo gold.\n\n",
          "votes": 5
        },
        {
          "id": 1912854,
          "postDate": "2022-08-25T02:16:14.337Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks. Yes, teaming up would likely expedite achieving a gold. But I first wanted to build atleast a genuine profile by attaining a few solo medals. Not many strong candidates would team up with me otherwise.</p>",
          "rawMarkdown": "@cdeotte Thanks. Yes, teaming up would likely expedite achieving a gold. But I first wanted to build atleast a genuine profile by attaining a few solo medals. Not many strong candidates would team up with me otherwise.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1912744,
      "postDate": "2022-08-25T00:32:05.393Z",
      "content": "<p>Congratulation! Could you recommend me some papers like \"knowledge distillation\"? I feel that fun.</p>",
      "rawMarkdown": "Congratulation! Could you recommend me some papers like \"knowledge distillation\"? I feel that fun.",
      "votes": 1
    },
    {
      "id": 1912743,
      "postDate": "2022-08-25T00:31:35.793Z",
      "content": "<p>Wow! Well this is an elegant solution, you put that hardware to great use.</p>",
      "rawMarkdown": "Wow! Well this is an elegant solution, you put that hardware to great use.",
      "votes": 1
    },
    {
      "id": 2001917,
      "postDate": "2022-10-24T11:51:23.633Z",
      "content": "<p>Looks good</p>",
      "rawMarkdown": "Looks good",
      "votes": 2
    },
    {
      "id": 1923491,
      "postDate": "2022-09-02T09:05:59.083Z",
      "content": "<p>A solution with class as always, thanks for sharing and congratulations <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> ! A couple of questions:</p>\n<p><strong>1.</strong> In which cases do you think is worth to try knowledge distillation in order to increase the performance of a model?<br>\n<strong>2.</strong> How did you figure out the learning rate schedule showed in the picture?<br>\n<strong>3.</strong> To generate the test preds you averaged (bagging) the 10 folds models (also with several seeds), correct? I mean, you didn't retrain with all the available data.</p>",
      "rawMarkdown": "A solution with class as always, thanks for sharing and congratulations @cdeotte ! A couple of questions:\n\n**1.** In which cases do you think is worth to try knowledge distillation in order to increase the performance of a model?\n**2.** How did you figure out the learning rate schedule showed in the picture?\n**3.** To generate the test preds you averaged (bagging) the 10 folds models (also with several seeds), correct? I mean, you didn't retrain with all the available data.",
      "votes": 2,
      "replies": [
        {
          "id": 1926177,
          "postDate": "2022-09-04T15:59:51.840Z",
          "content": "<ol>\n<li><p>Knowledge distillation is commonly used to accelerate production in real life model deployment. i.e. we can convert a many model ensemble into a single simple model. In competitions, it is often helpful to distill one model (or ensemble) into another model that is diverse from the first because the new model will store the information differently and add boost when added to an ensemble. Furthermore when used in combination with unlabeled data it is powerful. (Then it is a combination of pseudo label and knowledge distillation).</p></li>\n<li><p>Trial and error. I first found the best schedule using just train data (without pretrain). Next I trained using this schedule with pretrain and this same schedule using finetune. Next I started adjusting the finetune schedule to maximize CV</p></li>\n<li><p>I did not train with 100% data. For each model, i trained and predicted 10 fold models and averaged them. Then I changed the seed and changed some architecture like perhaps changing layers from 3 to 4. I trained another 10 fold models and averaged those new 10 folds. Lastly I took all my averages and averaged those equally. Then this was my \"NN model\". Next I did this procedure with my \"LGBM model\". And then 50% 50% those two.</p></li>\n</ol>",
          "rawMarkdown": "1. Knowledge distillation is commonly used to accelerate production in real life model deployment. i.e. we can convert a many model ensemble into a single simple model. In competitions, it is often helpful to distill one model (or ensemble) into another model that is diverse from the first because the new model will store the information differently and add boost when added to an ensemble. Furthermore when used in combination with unlabeled data it is powerful. (Then it is a combination of pseudo label and knowledge distillation).\n\n2. Trial and error. I first found the best schedule using just train data (without pretrain). Next I trained using this schedule with pretrain and this same schedule using finetune. Next I started adjusting the finetune schedule to maximize CV\n\n3. I did not train with 100% data. For each model, i trained and predicted 10 fold models and averaged them. Then I changed the seed and changed some architecture like perhaps changing layers from 3 to 4. I trained another 10 fold models and averaged those new 10 folds. Lastly I took all my averages and averaged those equally. Then this was my \"NN model\". Next I did this procedure with my \"LGBM model\". And then 50% 50% those two.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1923287,
      "postDate": "2022-09-02T05:40:06.743Z",
      "content": "<p>An interesting solution. Thank you so much for the detailed description!</p>",
      "rawMarkdown": "An interesting solution. Thank you so much for the detailed description!",
      "votes": 2
    },
    {
      "id": 1919936,
      "postDate": "2022-08-30T18:53:05.317Z",
      "content": "<p>i am done with this</p>",
      "rawMarkdown": "i am done with this",
      "votes": 2
    },
    {
      "id": 1915172,
      "postDate": "2022-08-26T18:10:32.083Z",
      "content": "<p>Congratulations Chris, I have tried RAPIDs FIL, and it's pretty fast. Although with model explanatory libraries like dalex, eli5 or scikit  I needed to do some tweaks to get permutation importance working with FIL, and the performance gets reduced (mostly by sending to GPU per iteration) do you know any library that can be easily integrated?</p>",
      "rawMarkdown": "Congratulations Chris, I have tried RAPIDs FIL, and it's pretty fast. Although with model explanatory libraries like dalex, eli5 or scikit  I needed to do some tweaks to get permutation importance working with FIL, and the performance gets reduced (mostly by sending to GPU per iteration) do you know any library that can be easily integrated?",
      "votes": 2,
      "replies": [
        {
          "id": 1915189,
          "postDate": "2022-08-26T18:21:04.197Z",
          "content": "<p>For permutation importance, i just wrote a simple for-loop. Iterate over all the features. Then inside the for-loop, randomly shuffle the chosen feature, infer model with RAPIDS FIL, compute metric. For each feature, i would shuffle the column with 10 different seeds for each of 10 folds. I would infer all 100 times and average the result. Then save the result into a dataframe.</p>\n<p>Also I kept everything on GPU. The dataframe to infer is on cuDF GPU, the predictions are kept on GPU, and metric is computed on GPU. And result stored on GPU. The entire for-loop exectutes entirely on GPU and was unbelievably fast like a few seconds for each feature. (If I remember correctly, I think using CPU LGBM to infer took minutes per feature).</p>\n<p>Permutation importance worked well. It allowed me to drop hundreds of features without affecting CV LB but it never boosted the CV nor LB. So in the end, I just kept all the features anyway. In other comps, I have seen CV LB boosts with permutation importance feature selection but not here.</p>",
          "rawMarkdown": "For permutation importance, i just wrote a simple for-loop. Iterate over all the features. Then inside the for-loop, randomly shuffle the chosen feature, infer model with RAPIDS FIL, compute metric. For each feature, i would shuffle the column with 10 different seeds for each of 10 folds. I would infer all 100 times and average the result. Then save the result into a dataframe.\n\nAlso I kept everything on GPU. The dataframe to infer is on cuDF GPU, the predictions are kept on GPU, and metric is computed on GPU. And result stored on GPU. The entire for-loop exectutes entirely on GPU and was unbelievably fast like a few seconds for each feature. (If I remember correctly, I think using CPU LGBM to infer took minutes per feature).\n\nPermutation importance worked well. It allowed me to drop hundreds of features without affecting CV LB but it never boosted the CV nor LB. So in the end, I just kept all the features anyway. In other comps, I have seen CV LB boosts with permutation importance feature selection but not here.",
          "votes": 7
        },
        {
          "id": 1915207,
          "postDate": "2022-08-26T18:36:17.863Z",
          "content": "<p>Understood, wanted to take the easy path with something pre-build. But yes make sense FIL is just really in another level for inference, will build the loop!<br>\nthank you </p>",
          "rawMarkdown": "Understood, wanted to take the easy path with something pre-build. But yes make sense FIL is just really in another level for inference, will build the loop!\nthank you ",
          "votes": 2
        },
        {
          "id": 1915238,
          "postDate": "2022-08-26T19:30:28.810Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I didn't know about Rapids FIL library so I didn't do perm importance at all; the large number of features caused the process to annoyingly run for hours/days on CPU.</p>\n<p>What could be the reason for not improving the scores after dropping bad features obtained from permutation importance?</p>\n<p>Could it be that there were several direct or indirect correlated features so that shuffling process of permutation importance gave misleading decisions? If this is true, then if we see two highly correlated features, we can shuffle their values at the same time, so permutation scores won't be misleading. Are there any other sensible things to opt for? Thanks.</p>",
          "rawMarkdown": "@cdeotte I didn't know about Rapids FIL library so I didn't do perm importance at all; the large number of features caused the process to annoyingly run for hours/days on CPU.\n\nWhat could be the reason for not improving the scores after dropping bad features obtained from permutation importance?\n\nCould it be that there were several direct or indirect correlated features so that shuffling process of permutation importance gave misleading decisions? If this is true, then if we see two highly correlated features, we can shuffle their values at the same time, so permutation scores won't be misleading. Are there any other sensible things to opt for? Thanks.",
          "votes": 2
        },
        {
          "id": 1915259,
          "postDate": "2022-08-26T20:04:10.323Z",
          "content": "<p>Now that I think about it, i think i tried to remove too many features. I tried removing a few hundred. I didn't consider statistical significance. I just checked, only 35 out of 1366 features are statistically significant bad. I will try training a new LGBM Dart with these 35 removed.</p>\n<p>The reason many Kagglers had trouble with permutation importance is because you need to infer all 10 folds and you need to shuffle each column multiple times (like 10) and average. Therefore each feature needs to be inferred 100 times. If you tried using CPU this would take literally a month or more. </p>\n<p>If you do not infer each column 100 times, then you will only see random noise results because the metric has so much variance. When we do infer 100 times, then each average metric score has standard deviation around <code>0.00003</code> which means that if we observe an increase in metric score of <code>0.0001</code> then we are 99% confident that feature is bad. There are only 35 features that are statistically significant hurt the model.</p>\n<p>The orange line is the baseline metric score of 0.7977. The model is 10 folds of public notebook LGBM Dart by Martin. The dataframe images show the best 10 and worst 10 features. The last column <code>z</code> shows the statistical z-test score. If <code>z&lt;-3</code> that means we are 99% confident this feature helps. If <code>z&gt;3</code>, we are 99% confidence this feature is bad.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/feat1.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/feat2.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/df_h.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/df_t.png\" alt=\"\"></p>",
          "rawMarkdown": "Now that I think about it, i think i tried to remove too many features. I tried removing a few hundred. I didn't consider statistical significance. I just checked, only 35 out of 1366 features are statistically significant bad. I will try training a new LGBM Dart with these 35 removed.\n\nThe reason many Kagglers had trouble with permutation importance is because you need to infer all 10 folds and you need to shuffle each column multiple times (like 10) and average. Therefore each feature needs to be inferred 100 times. If you tried using CPU this would take literally a month or more. \n\nIf you do not infer each column 100 times, then you will only see random noise results because the metric has so much variance. When we do infer 100 times, then each average metric score has standard deviation around `0.00003` which means that if we observe an increase in metric score of `0.0001` then we are 99% confident that feature is bad. There are only 35 features that are statistically significant hurt the model.\n\nThe orange line is the baseline metric score of 0.7977. The model is 10 folds of public notebook LGBM Dart by Martin. The dataframe images show the best 10 and worst 10 features. The last column `z` shows the statistical z-test score. If `z<-3` that means we are 99% confident this feature helps. If `z>3`, we are 99% confidence this feature is bad.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/feat1.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/feat2.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/df_h.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/df_t.png)",
          "votes": 10
        },
        {
          "id": 1915289,
          "postDate": "2022-08-26T20:30:53.407Z",
          "content": "<p>Indeed. Aware of the random noise. Usually I do at least 3 runs and average but in this competition even 1 run for 1 fold was taking long time. I was checking just 1 run for a fold (that too for a subset (<em>_last</em> type) of features only); the least possible thing haha. </p>\n<p>Even though results will be noisy, I thought to consider only highly bad performing features in a hope to avoid the noise issue to an extent. But of course better ways to measure those \"highly\" bad or good performing features is to measure z-score as you computed above (this requires &gt; 1 run).</p>\n<p>Btw, I made a typo by saying CPU. I used GPU machines and cudf until features creation. Before training the model, cudf was converted back to pandas. Not sure if that affects permutation importance timing during for-loop if cudf was instead thrown into model training and until inference.</p>",
          "rawMarkdown": "Indeed. Aware of the random noise. Usually I do at least 3 runs and average but in this competition even 1 run for 1 fold was taking long time. I was checking just 1 run for a fold (that too for a subset (*_last* type) of features only); the least possible thing haha. \n\nEven though results will be noisy, I thought to consider only highly bad performing features in a hope to avoid the noise issue to an extent. But of course better ways to measure those \"highly\" bad or good performing features is to measure z-score as you computed above (this requires > 1 run).\n\nBtw, I made a typo by saying CPU. I used GPU machines and cudf until features creation. Before training the model, cudf was converted back to pandas. Not sure if that affects permutation importance timing during for-loop if cudf was instead thrown into model training and until inference.",
          "votes": 1
        },
        {
          "id": 1915426,
          "postDate": "2022-08-27T01:04:32.583Z",
          "content": "<p>Suggest to strongly consider evaluating against auc not amex score. Better chance of clearly seeing which features are statistically hurting the model. </p>\n<p>Or do both ways, try removing based on one vs the other and see what scores best. </p>\n<p>Another thought I had: if someone really wants to be sure to optimize the competition metric, probably better to rewrite a custom \"smoothed\" metric that averages scores at like 3.5%, 3.75%, 4, 4.25, and 4.5. Or something similar. Can test what gives good results but much smoother and smaller variance. </p>\n<p>Just idle speculation, but overfitting the competition metric on train and public could be one difference between public and private scores. Consider if more people defaulted in test set (covid). Then more people would default in top 4% too. But due to the 1/20th sampling of negative cases, it means there would be MORE CUSTOMERS in the 4% zone in private test! So instead of, say, top 20% of visible data is within the 4%, now the higher scores (807 instead of 800) include more like 25% of visible data within the 4% zone. </p>\n<p>It was a really weird metric. It didn't combine well with the undersampled data.</p>",
          "rawMarkdown": "Suggest to strongly consider evaluating against auc not amex score. Better chance of clearly seeing which features are statistically hurting the model. \n\nOr do both ways, try removing based on one vs the other and see what scores best. \n\nAnother thought I had: if someone really wants to be sure to optimize the competition metric, probably better to rewrite a custom \"smoothed\" metric that averages scores at like 3.5%, 3.75%, 4, 4.25, and 4.5. Or something similar. Can test what gives good results but much smoother and smaller variance. \n\nJust idle speculation, but overfitting the competition metric on train and public could be one difference between public and private scores. Consider if more people defaulted in test set (covid). Then more people would default in top 4% too. But due to the 1/20th sampling of negative cases, it means there would be MORE CUSTOMERS in the 4% zone in private test! So instead of, say, top 20% of visible data is within the 4%, now the higher scores (807 instead of 800) include more like 25% of visible data within the 4% zone. \n\nIt was a really weird metric. It didn't combine well with the undersampled data.",
          "votes": 2
        },
        {
          "id": 2007961,
          "postDate": "2022-10-28T15:51:46.423Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . First of all, congratulations on your insights. Second, would you mind sharing the code you used to create this dataframe with Z-Score? I'm using this competition a lot for studying but I'm still a newbie in Stats with Python</p>",
          "rawMarkdown": "Hello @cdeotte . First of all, congratulations on your insights. Second, would you mind sharing the code you used to create this dataframe with Z-Score? I'm using this competition a lot for studying but I'm still a newbie in Stats with Python"
        },
        {
          "id": 2007973,
          "postDate": "2022-10-28T16:06:53.353Z",
          "content": "<p>The following notebook <a href=\"https://www.kaggle.com/code/cdeotte/lstm-feature-importance\" target=\"_blank\">here</a> shows how to do permutation importance with a NN. This isn't the exact code to create the dataframe above, but its the foundation.</p>",
          "rawMarkdown": "The following notebook [here][1] shows how to do permutation importance with a NN. This isn't the exact code to create the dataframe above, but its the foundation.\n\n[1]: https://www.kaggle.com/code/cdeotte/lstm-feature-importance"
        }
      ]
    },
    {
      "id": 1913524,
      "postDate": "2022-08-25T10:52:00.047Z",
      "content": "<p>Thanks a lot for the great sharing! I really learned a lot! <br>\nMay I ask how much the feature selection improved your single-lgbm-model CV score? how many features did you generate and how many features did you keep in your final model? </p>",
      "rawMarkdown": "Thanks a lot for the great sharing! I really learned a lot! \nMay I ask how much the feature selection improved your single-lgbm-model CV score? how many features did you generate and how many features did you keep in your final model? ",
      "votes": 2,
      "replies": [
        {
          "id": 1913646,
          "postDate": "2022-08-25T12:32:33.973Z",
          "content": "<p>Feature selection/creation did not help me nor improve Martin's LGBM model. I explored adding many new groupby, diff, and product features. I also located hundreds of features that could be dropped without affecting CV. </p>\n<p>But in the end, my final submission is Martin's model as is with his features (none added nor removed). I trained his model locally with 10 folds and various seeds. My solution's LB boost comes from the NN Transformer.</p>",
          "rawMarkdown": "Feature selection/creation did not help me nor improve Martin's LGBM model. I explored adding many new groupby, diff, and product features. I also located hundreds of features that could be dropped without affecting CV. \n\nBut in the end, my final submission is Martin's model as is with his features (none added nor removed). I trained his model locally with 10 folds and various seeds. My solution's LB boost comes from the NN Transformer.",
          "votes": 4
        },
        {
          "id": 1913671,
          "postDate": "2022-08-25T12:43:28.020Z",
          "content": "<p>Thanks a lot!</p>",
          "rawMarkdown": "Thanks a lot!",
          "votes": 2
        }
      ]
    },
    {
      "id": 1913209,
      "postDate": "2022-08-25T08:17:44.117Z",
      "content": "<p>Congrats and thans for all the post during the competition</p>",
      "rawMarkdown": "Congrats and thans for all the post during the competition",
      "votes": 2
    },
    {
      "id": 1913199,
      "postDate": "2022-08-25T08:08:59.647Z",
      "content": "<p>Thanks for the write-up! The pseudo label/distilation seems like a tight-rope walk for leakage and over-fitting (if you don't know what you're doing like me).</p>",
      "rawMarkdown": "Thanks for the write-up! The pseudo label/distilation seems like a tight-rope walk for leakage and over-fitting (if you don't know what you're doing like me).",
      "votes": 2
    },
    {
      "id": 1912846,
      "postDate": "2022-08-25T02:04:34.060Z",
      "content": "<p>Thank you and congratulations Chris! Your public models helped me get started.</p>",
      "rawMarkdown": "Thank you and congratulations Chris! Your public models helped me get started.",
      "votes": 2
    },
    {
      "id": 1912844,
      "postDate": "2022-08-25T02:01:15.543Z",
      "content": "<p>Love the solution, how elegant it was, wow 😊 Huge congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, and thank you very much for the write-up!</p>\n<p>I am also super impressed by the software engineering that must have gone into this, both to train and do feature selection with nested CV, and to run training and inference at scale.</p>\n<p>Huge congrats! 🥳</p>",
      "rawMarkdown": "Love the solution, how elegant it was, wow 😊 Huge congrats @cdeotte, and thank you very much for the write-up!\n\nI am also super impressed by the software engineering that must have gone into this, both to train and do feature selection with nested CV, and to run training and inference at scale.\n\nHuge congrats! 🥳",
      "votes": 2
    },
    {
      "id": 1912738,
      "postDate": "2022-08-25T00:26:31.007Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 2
    },
    {
      "id": 1912924,
      "postDate": "2022-08-25T03:44:37.663Z",
      "content": "<p>I'm really interested in transformer model recently. Someone, please instruct me. By the way I envy cooperation with Nvidia.</p>",
      "rawMarkdown": "I'm really interested in transformer model recently. Someone, please instruct me. By the way I envy cooperation with Nvidia.",
      "votes": 1
    },
    {
      "id": 3115974,
      "postDate": "2025-02-05T13:40:20.943Z",
      "content": "<p>Thanks for sharing such a detailed approach! It helped me understand the concepts as a newbie data science enthusiast. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>",
      "rawMarkdown": "Thanks for sharing such a detailed approach! It helped me understand the concepts as a newbie data science enthusiast. @cdeotte"
    },
    {
      "id": 3113218,
      "postDate": "2025-02-02T11:19:45.390Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, Can you Please public the knowledge distillation code if available now.</p>",
      "rawMarkdown": "@cdeotte, Can you Please public the knowledge distillation code if available now."
    },
    {
      "id": 1915234,
      "postDate": "2022-08-26T19:26:40.610Z",
      "content": "<p><a href=\"https://www.kaggle.com/datasets/ashishmotwani/happyface\" target=\"_blank\">https://www.kaggle.com/datasets/ashishmotwani/happyface</a></p>",
      "rawMarkdown": "https://www.kaggle.com/datasets/ashishmotwani/happyface"
    },
    {
      "id": 1912752,
      "postDate": "2022-08-25T00:34:56.180Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> and Idea. The transformer we were not able to get much juice from though we went here with tabtransformer (0.795) <a href=\"https://www.kaggle.com/code/gauravbrills/tabtransformer-training\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/tabtransformer-training</a> ., I think knowledge distillation was the trick here ,definitely the secret sauce here is just awesome  </p>\n<p>Also one curious question I had as u had tried the GRU model with 13 seq length .. Were these model successfully for you, as we did try a similar one with TCN wavenet but couldn't score well.</p>",
      "rawMarkdown": "Great work @cdeotte and Idea. The transformer we were not able to get much juice from though we went here with tabtransformer (0.795) https://www.kaggle.com/code/gauravbrills/tabtransformer-training ., I think knowledge distillation was the trick here ,definitely the secret sauce here is just awesome  \n\n Also one curious question I had as u had tried the GRU model with 13 seq length .. Were these model successfully for you, as we did try a similar one with TCN wavenet but couldn't score well.",
      "replies": [
        {
          "id": 1912755,
          "postDate": "2022-08-25T00:38:29.633Z",
          "content": "<p>Congratulations Gaurav and team! I forget to mention that I added a GRU layer after my transformer before the final classification layers. Both transformer and GRU helped. I will update my post.</p>",
          "rawMarkdown": "Congratulations Gaurav and team! I forget to mention that I added a GRU layer after my transformer before the final classification layers. Both transformer and GRU helped. I will update my post.",
          "votes": 3
        },
        {
          "id": 1912762,
          "postDate": "2022-08-25T00:45:57.697Z",
          "content": "<p>Thanks Chris .. always great to learn from you .. One of our ensemble strategy was using your Forward selection algo (Not sure u still use that :) ).. interesting will be great to know how GRU was used as we did went into the path of using wavenet like in amex paper and combining with transformers , though just left it .</p>",
          "rawMarkdown": "Thanks Chris .. always great to learn from you .. One of our ensemble strategy was using your Forward selection algo (Not sure u still use that :) ).. interesting will be great to know how GRU was used as we did went into the path of using wavenet like in amex paper and combining with transformers , though just left it .",
          "votes": 1
        },
        {
          "id": 1912807,
          "postDate": "2022-08-25T01:22:04.757Z",
          "content": "<p>Below is my updated model architecture with GRU layer added: (also embeddings improved from <code>(10,4)</code> to <code>(10,8)</code> and skip connections removed):</p>\n<pre><code>def build_model():\n\n    # INPUT - FIRST 11 COLUMNS ARE CAT, NEXT 177 ARE NUMERIC\n    inp = layers.Input(shape=(13,188))\n    embeddings = []\n    for k in range(11):\n        emb = layers.Embedding(10,8)\n        embeddings.append( emb(inp[:,:,k]) )\n    x = layers.Concatenate()([inp[:,:,11:]]+embeddings)\n\n    # \"EMBEDDING LAYER\"\n    x = layers.Dense(feat_dim)(x)\n\n    # TRANSFORMER BLOCKS\n    for k in range(num_blocks):\n        transformer_block = TransformerBlock(embed_dim, feat_dim, num_heads, ff_dim, dropout_rate)\n        x = transformer_block(x)\n\n    # REGRESSION HEAD\n    x = tf.keras.layers.GRU(units=128, return_sequences=False)(x)\n    x = layers.Dense(64, activation=\"relu\")(x)\n    x = layers.Dense(32, activation=\"relu\")(x)\n    outputs = layers.Dense(1, activation=\"sigmoid\")(x)\n\n    model = keras.Model(inputs=inp, outputs=outputs)\n    opt = tf.keras.optimizers.Adam(learning_rate=0.001)\n    loss = tf.keras.losses.BinaryCrossentropy()\n    model.compile(loss=loss, optimizer = opt)\n\n    return model\n</code></pre>",
          "rawMarkdown": "Below is my updated model architecture with GRU layer added: (also embeddings improved from `(10,4)` to `(10,8)` and skip connections removed):\n\n    def build_model():\n    \n        # INPUT - FIRST 11 COLUMNS ARE CAT, NEXT 177 ARE NUMERIC\n        inp = layers.Input(shape=(13,188))\n        embeddings = []\n        for k in range(11):\n            emb = layers.Embedding(10,8)\n            embeddings.append( emb(inp[:,:,k]) )\n        x = layers.Concatenate()([inp[:,:,11:]]+embeddings)\n        \n        # \"EMBEDDING LAYER\"\n        x = layers.Dense(feat_dim)(x)\n    \n        # TRANSFORMER BLOCKS\n        for k in range(num_blocks):\n            transformer_block = TransformerBlock(embed_dim, feat_dim, num_heads, ff_dim, dropout_rate)\n            x = transformer_block(x)\n    \n        # REGRESSION HEAD\n        x = tf.keras.layers.GRU(units=128, return_sequences=False)(x)\n        x = layers.Dense(64, activation=\"relu\")(x)\n        x = layers.Dense(32, activation=\"relu\")(x)\n        outputs = layers.Dense(1, activation=\"sigmoid\")(x)\n    \n        model = keras.Model(inputs=inp, outputs=outputs)\n        opt = tf.keras.optimizers.Adam(learning_rate=0.001)\n        loss = tf.keras.losses.BinaryCrossentropy()\n        model.compile(loss=loss, optimizer = opt)\n        \n        return model",
          "votes": 14
        },
        {
          "id": 1913041,
          "postDate": "2022-08-25T05:27:29.723Z",
          "content": "<p>Thanks a lot for sharing</p>",
          "rawMarkdown": "Thanks a lot for sharing",
          "votes": 2
        },
        {
          "id": 1913266,
          "postDate": "2022-08-25T09:04:33.497Z",
          "content": "<p>Congratulations and thanks for guidance.<br>\nSince this topics about transformer, inspired by your works another try of tabtransformer <a href=\"https://www.kaggle.com/code/yekenot/amex-pytorch-tabtransformer\" target=\"_blank\">https://www.kaggle.com/code/yekenot/amex-pytorch-tabtransformer</a>. Not very good CV, but enough for learn.</p>",
          "rawMarkdown": "Congratulations and thanks for guidance.\nSince this topics about transformer, inspired by your works another try of tabtransformer https://www.kaggle.com/code/yekenot/amex-pytorch-tabtransformer. Not very good CV, but enough for learn.",
          "votes": 3
        },
        {
          "id": 1914845,
          "postDate": "2022-08-26T13:24:46.957Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I also had one curious newbie question , We did not have capacity to train on 10 folds bur did training NN on more folds (more data) helped . If yes why was this the case specially with Transformer type networks  (could a large batch size have helped here ) just curious 😳?</p>",
          "rawMarkdown": "@cdeotte I also had one curious newbie question , We did not have capacity to train on 10 folds bur did training NN on more folds (more data) helped . If yes why was this the case specially with Transformer type networks  (could a large batch size have helped here ) just curious 😳?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1918806,
      "postDate": "2022-08-29T21:20:55.863Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1913841,
      "postDate": "2022-08-25T14:33:46.177Z",
      "content": "<p>Congrats and thanks for sharing</p>",
      "rawMarkdown": "Congrats and thanks for sharing",
      "votes": 3
    },
    {
      "id": 1930180,
      "postDate": "2022-09-07T16:15:44.677Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing.",
      "votes": 1
    },
    {
      "id": 1929633,
      "postDate": "2022-09-07T08:27:26.320Z",
      "content": "<p>Thanks, great explanation</p>",
      "rawMarkdown": "Thanks, great explanation\n",
      "votes": 1
    },
    {
      "id": 1928429,
      "postDate": "2022-09-06T13:19:18Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1928301,
      "postDate": "2022-09-06T12:04:40.233Z",
      "content": "<p>thanks, dude</p>",
      "rawMarkdown": "thanks, dude\n",
      "votes": 1
    },
    {
      "id": 1927744,
      "postDate": "2022-09-05T23:56:21.847Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Thanks for sharing @cdeotte ",
      "votes": 1
    },
    {
      "id": 1923739,
      "postDate": "2022-09-02T13:46:48.913Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": 1
    },
    {
      "id": 1922034,
      "postDate": "2022-09-01T08:27:41.490Z",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)",
      "votes": 1
    },
    {
      "id": 1921405,
      "postDate": "2022-08-31T20:10:31.217Z",
      "content": "<p>Very useful. Thank you.</p>",
      "rawMarkdown": "Very useful. Thank you.",
      "votes": 1
    },
    {
      "id": 1921333,
      "postDate": "2022-08-31T18:18:38.810Z",
      "content": "<p>Great job! Learned a lot Thanks!🔥</p>",
      "rawMarkdown": "Great job! Learned a lot Thanks!🔥",
      "votes": 1
    },
    {
      "id": 1921188,
      "postDate": "2022-08-31T16:23:26.557Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": 1
    },
    {
      "id": 1920117,
      "postDate": "2022-08-30T23:13:11.357Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1919935,
      "postDate": "2022-08-30T18:52:36.293Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 1919862,
      "postDate": "2022-08-30T17:54:33.760Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1919390,
      "postDate": "2022-08-30T11:15:04.187Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1917815,
      "postDate": "2022-08-29T05:09:27.730Z",
      "content": "<p>Great. Thanks, Sir </p>",
      "rawMarkdown": "Great. Thanks, Sir ",
      "votes": 1
    },
    {
      "id": 1917762,
      "postDate": "2022-08-29T03:39:17.373Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1916709,
      "postDate": "2022-08-28T05:43:34.583Z",
      "content": "<p>Thank You Chris :)</p>",
      "rawMarkdown": "Thank You Chris :)",
      "votes": 1
    },
    {
      "id": 1916708,
      "postDate": "2022-08-28T05:43:17.270Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1915801,
      "postDate": "2022-08-27T11:39:08.647Z",
      "content": "<p>Thank You :)</p>",
      "rawMarkdown": "Thank You :)",
      "votes": 1
    },
    {
      "id": 1915736,
      "postDate": "2022-08-27T09:14:49.973Z",
      "content": "<p>Thank You :)</p>",
      "rawMarkdown": "Thank You :)\n",
      "votes": 1
    },
    {
      "id": 1915664,
      "postDate": "2022-08-27T07:04:19.957Z",
      "content": "<p>Great, thanks</p>",
      "rawMarkdown": "Great, thanks",
      "votes": 1
    },
    {
      "id": 1915474,
      "postDate": "2022-08-27T02:29:14.883Z",
      "content": "<p>Well explained Sir, Thanks.</p>",
      "rawMarkdown": "Well explained Sir, Thanks.",
      "votes": 1
    },
    {
      "id": 1915408,
      "postDate": "2022-08-27T00:28:27.453Z",
      "content": "<p>understood, thanks</p>",
      "rawMarkdown": "understood, thanks",
      "votes": 1
    },
    {
      "id": 1914853,
      "postDate": "2022-08-26T13:29:00.257Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1914640,
      "postDate": "2022-08-26T08:57:00.113Z",
      "content": "<p>Congrats and thanks for sharing</p>",
      "rawMarkdown": "Congrats and thanks for sharing",
      "votes": 1
    },
    {
      "id": 1914389,
      "postDate": "2022-08-26T04:15:24.160Z",
      "content": "<p>Thanks for sharing this approach</p>",
      "rawMarkdown": "Thanks for sharing this approach\n",
      "votes": 1
    },
    {
      "id": 1913871,
      "postDate": "2022-08-25T14:53:26.087Z",
      "content": "<p>Congrats and thanks for sharing.</p>",
      "rawMarkdown": "Congrats and thanks for sharing.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1912765,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2022-08-25T00:49:53.133000",
      "content": "<p>Awesome solution Chris, congrats!</p>\n<p>I'm curious about the data preparation you did for the transformer - was it the same as what you use in your public notebook? I spent a lot of time trying to trying to figure out the best way to represent missing statements, whether masking would help, etc. Got some boosts from imputing all the nulls with lgbm and using data augmentation (shift statements forward 1), but could never do better than .793 CV.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1912801,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-25T01:16:44.143000",
          "content": "<p>Hi Joe, congratulations to you and team for achieving 24th out of 5000 teams. That's great.</p>\n<p>I prepared data exactly like i did in my public notebook (with new targets of test preds and oof preds). I also tried many alternatives like adding padding correctly in time. (My public notebook adds all pads in front of sequence for customer with missing statements). I also tried different ways to represent NA and padding. I also tried adding position embedding. I did lots of things but the data representation didn't make much difference.</p>\n<p>I only made 4 change to my public notebook</p>\n<ul>\n<li><code>emb = layers.Embedding(10,4)</code> changed to <code>(10,8)</code></li>\n<li>after transformer blocks added <code>x = tf.keras.layers.GRU(units=128, return_sequences=False)(x)</code> before dense layers</li>\n<li>for submission.csv, average multiple random seed trained transformers with 2, 3, 4 blocks and remove skipped connections</li>\n<li>trained for multiple cosine cycles with varying learning rates.</li>\n</ul>\n<p>The big help was pretraining with test predictions. This allowed us to use all the future test data features that Amex gave us and helped the model train it's attention and other stuff. This boosted CV score from 0.790 to 0.798</p>\n<p>(It's like NLP unsupervised pretraining. Even without labels, just training a transformer on lots text helps the model understand language. And in this comp, our transformer will learn more about credit cards by seeing all the test features even without correct test labels)</p>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 1912804,
          "author_name": "EL Younes",
          "author_url": "",
          "post_date": "2022-08-25T01:19:17.463000",
          "content": "<p>Maybe, if you change the percentage of the components of the target such as 70% (0) and 30% (1) before using splitting cv by 5 folds, you will get 0.795. this is <a href=\"https://www.kaggle.com/youneseloiarm/xgboost-starter-30-70\" target=\"_blank\">XGBoost Starter -30%-70%</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1912842,
          "author_name": "Joe Eddy",
          "author_url": "",
          "post_date": "2022-08-25T01:56:24.473000",
          "content": "<p>Thanks for such a detailed (and fast!) response. Using the test data for a performance edge makes a ton of sense. A few days ago I had the idea to try transfer learning on P2 - pretrain a transformer on train+test to predict P2 before fine tuning on train - couldn't get a boost out of it, but wish I had had the idea earlier on to experiment with more thoroughly and possibly push to a better approach like this one.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1912890,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T02:48:52.590000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2212438,
      "author_name": "Çağatay Kılınç",
      "author_url": "",
      "post_date": "2023-04-06T19:06:47.257000",
      "content": "<p>Thanks for sharing this  solution. It is very much helpfull for fresh data sientist as me. </p>\n<p>I couldnt understand what NN is? I got that lgbm is a machine learning model and ı also apply this for many competitions. But ı couldnt understand NN and how you can combine with lgbm ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2212461,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2023-04-06T19:18:41.820000",
          "content": "<p>NN is an abbreviation for neural network. It refers to transformers, CNN (convolution neural networks), RNN (recursive neural networks like LSTM and GRU), MLP (multi-layer perceptron), etc</p>\n<p>People also use NN to refer to DL (i.e. deep learning solutions)</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2212470,
              "author_name": "Çağatay Kılınç",
              "author_url": "",
              "post_date": "2023-04-06T19:29:27.463000",
              "content": "<p>I couldn't quite get some things in my head. Thank you very much for your quick and descriptive reply Chris!<br>\nBut it is confusing for me that how did you combine lightgbm and nn in one?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2212488,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-04-06T19:44:41.923000",
              "content": "<p>I stacked the two models. I used the predictions from LGBM as inputs to NN</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1914862,
      "author_name": "Mike Mazurov",
      "author_url": "",
      "post_date": "2022-08-26T13:40:55.723000",
      "content": "<p>Congrats with 17th place and thanks for sharing your work. If I understand correctly we use distillation of knowledge from larger model to smaller while you did from smaller to larger (from LGBM to Transformer) or am I wrong somewhere? And what exactly gave distillation in this case, why couldn't you just make an ensemble of two models without distillation?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1916064,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-27T15:46:08.570000",
          "content": "<p>Transformers are hard to train from scratch. The model needs to learn weights for its self attention and weights for all its layers. When we train with soft targets from another trained model, it provides more information than just using the original train hard targets.</p>\n<p>The original hard targets are 0's and 1's. But the other model's soft targets are 0.1, 0.2, 0.3, …, 0.9, 1.0 and all numbers in between. This contains much more information for the NN Transformer to learn with. So the knowledge distillation from LGBM helps the NN learn. Furthermore, NN is fundamentally different than LGBM, so once the knowledge is transferred, it will be \"represented differently\" and add it's own \"personality\" to the predictions and make it effective in ensemble.</p>\n<p>The biggest reason why this approach worked so well is that it utilizes the 11 million rows of unlabeled test data. Even though we don't have labels for test data, we still have 1 million time series of length 13 for the different features in test data. This teaches our models how the features change over time and helps the model predict what the features will be in the future. If we never use test data, we never gain access to these 1 million time series with all its information.</p>\n<p>Take feature <code>P_2</code> for example. This is like a customer's credit score. If we knew this value for 18 months into the future after the last credit card statement, we could probably predict every customer's default perfectly. What better way to guess what <code>P_2</code> will be in the future than watching the 1 million time series of length 13 in the test data to learn how it changes over time. And using that to train our model.</p>",
          "votes": 16,
          "replies": []
        }
      ]
    },
    {
      "id": 1913464,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-08-25T10:33:46.613000",
      "content": "<p>Congrats for the medal and thanks for sharing all that insightful code.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1916449,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-27T22:44:11.500000",
          "content": "<p>Thanks Lucas. Thanks for your helpful notebooks and discussions.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1913394,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-08-25T10:03:14.207000",
      "content": "<p>What exactly do you mean by knowledge distillation? How do you extract it from the LGBM? Is it simply stacking its prediction?<br>\nAmazing work btw and congrats! </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1913527,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-25T10:53:18.960000",
          "content": "<p>Knowledge distillation is the transfer of knowledge from 1 model to another model. First we train an LGBM. Next we perform teacher student. We use the predictions (both OOF and test preds) of LGBM to train an NN Transformer. Finally we finetune NN Transformer on the true train targets.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1913800,
          "author_name": "Jonathan Mallia",
          "author_url": "",
          "post_date": "2022-08-25T14:08:48.057000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> interested to know more how this can work. <br>\nBy any chance are you publishing your code? That would be great</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1914130,
          "author_name": "Susnato Dhar",
          "author_url": "",
          "post_date": "2022-08-25T18:56:30.163000",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, first of all, congrats and thanks for sharing your result, the second I have a silly question about <code>knowledge distillation</code>, when you said, <code>We use the predictions (both OOF and test preds) of LGBM to train an NN Transformer</code> do we use the test data directly into the training pipeline(and use the test preds from LGBM as labels)? If possible could you please refer to some materials where I can learn about this technique more?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1917651,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-29T00:28:35.233000",
          "content": "<p>Yes <a href=\"https://www.kaggle.com/susnato\" target=\"_blank\">@susnato</a> , we use test data directly into the training pipeline (and use test preds from LGBM as labels). Note that we leave the test preds as is (as continuous numbers between 0 and 1) without converting them to 0's and 1's. This is a combination of both \"pseudo labeling\" and \"knowledge distillation\". You can google these two terms to learn more.</p>\n<p>Knowledge disllation helps the Transformer learn more accurately and quickly (by gaining the knowledge that LGBM has already learned but storing it in a new transformer way). And pseudo labeling allows us to access all the information in the unlabeled test data. Namely the features which are 1 million time series of length 13 from the 188 features. This is very useful information that Kaggle provided us!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1912927,
      "author_name": "Nischay Dhankhar",
      "author_url": "",
      "post_date": "2022-08-25T03:47:17.970000",
      "content": "<p>Congratulations on great finish and sharing the solution Chris. I wonder how well was your NN transformer cv-lb aligned after using Knowledge Distillation based on soft preds? I usually see a dramatic change in CV, but lesser improvement in Leaderboard, couldn't think of any possible way of leakage. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1913576,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-25T11:31:51.267000",
          "content": "<p>Congratulations Nischay on your fantastic solo performance.</p>\n<p>My NN Transformer has a better private LB than LGBM below are the stats</p>\n<ul>\n<li>NN Transformer, leak-free-CV 0.7980, Public LB 0.7988, Private LB 0.8076</li>\n<li>LGBM, CV 0.7990, Public LB 0.7992, Private LB 0.8070</li>\n</ul>\n<p>===== Below is more info =====</p>\n<p>The NN Transformer CV is accurate and leak free. Because the LGBM is trained using 10 (inner) folds within 10 (outer) folds (i.e. Nested K-Fold). The outer folds are typical K-Fold. When training outer fold 1 we use train targets from folds 2 thru 10. We then split this group of 9 folds into its own 10 folds. Each of these inner 10 folds we train 1 model. These inner models never see the targets from validation outer fold 1. Using these inner models, we create OOF for outer fold 1 which gives us leak free predictions for outer folds 2-10. We also use these inner models to predict test preds. (Next we train 10 inner folds for outer fold 2, etc etc. In total we train 100 models)</p>\n<p>This gives us a unique <code>lstm_OOF_outer_fold_1.csv</code> and <code>lstm_Test_preds_outer_fold_1.csv</code> for each of 10 outer folds for a total of 20 CSV. When we train fold 1 of our NN Transformer, we use these two CSV. When we train fold 2 of our NN Transformer we use <code>lstm_OOF_outer_fold_2.csv</code> and <code>lstm_Test_preds_outer_fold_2.csv</code> etc etc. This prevents CV leaks.</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1914330,
          "author_name": "Nischay Dhankhar",
          "author_url": "",
          "post_date": "2022-08-26T02:22:32.277000",
          "content": "<p>I never thought of this possible leak free approach, that's something extraordinary. Thanks a lot for clearing up everything  🙇🙇</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1912919,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2022-08-25T03:36:28.050000",
      "content": "<p>Thank you for sharing this. There has been knowledge distillation in the Feedback competition as well (ending yesterday), something I really need to catch up.  Am I right to say that it's a student teacher model , where the student is trained with the prediction result of the teacher instead of the hard label ? <br>\nCongratulations for your result, and thanks again for all the sharings. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 1912925,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-25T03:45:14.560000",
          "content": "<p>Yes, exactly. It's very simple. It's like pseudo labels. Just save your OOF and test predictions. Then train your new model using the probabilities predictions from OOF and test preds using cross entropy. The loss cross entropy works when the targets are continuous preditcions between 0 and 1 (i.e. the targets targets don't need to be zeros and ones). </p>\n<p>Imagine you have two models, model A and model B. First train model A. Then makes predictions with model A on some dataset. Next train model B using the probabilities from model A and cross entropy (on that dataset). Then model B has learned knowledge distillation from model A.</p>\n<p>Afterward, you can either stop there or fine tune model B on more data. In this competition, I further trained on train data targets which are zeros and ones.</p>",
          "votes": 17,
          "replies": []
        },
        {
          "id": 1913006,
          "author_name": "Aninda Goswamy",
          "author_url": "",
          "post_date": "2022-08-25T05:06:18.980000",
          "content": "<p>Thanks a lot for explaining. Can we use similar approach for regression tasks as well.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1913173,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-08-25T07:42:45.963000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I get it what are you trying to explain, but isn't Pseudo labelling exactly this? then what's the difference between distillation and pseudo labelling?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1913763,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-25T13:47:57.253000",
          "content": "<p><a href=\"https://www.kaggle.com/mrinath\" target=\"_blank\">@mrinath</a> Below is my opinion, it might not be fully correct.</p>\n<p>Pseudo labeling is the process of assigning labels to unlabeled data (so that we can train with and extract info from the unlabeled data). Generally pseudo labels are converted into hard targets, i.e. 0's and 1's and only confident predictions are used. For example, after making predictions on unlabeled data we keep all samples with confident predictions less than 0.05 or greater than 0.95. Then round the predictions to 0 and 1. This allows us to perform supervised training on the previously unlabeled data which gives the model the benefit of using the unlabeled data's feature values.</p>\n<p>Knowledge distillation is the process of transferring one model's (or ensemble's) learning to another model (frequently used to transfer the performance of a complicated ensemble into a simple single model for purpose of efficient production inference). To facilitate the transfer, we can use any data whether it was originally labeled or not labeled. The original labels are discarded and the teacher model predicts new labels. These predicted labels are not converted to 0's and 1's but rather kept as probability values between 0 and 1. This allows for maximum transfer of the teacher's knowledge to the student model.</p>\n<p>We can also perform hybrids of the two. Like keeping pseudo labels soft. Then we simultaneously transfer knowledge from a teacher to a student and extract info from unlabeled data's features. My solution here is probably a combination of both.</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1913989,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-08-25T16:13:57.130000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> it makes sense, thanks</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1912747,
      "author_name": "Leandro Destefani",
      "author_url": "",
      "post_date": "2022-08-25T00:32:34.440000",
      "content": "<p>Congratulations, Chris, you are the real MVP! I wonder which features had most importance</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1914213,
      "author_name": "Ganesh Borkar",
      "author_url": "",
      "post_date": "2022-08-25T20:45:45.463000",
      "content": "<p>Awesome approach really Amazing solution</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1932845,
      "author_name": "Jeremy Zhang",
      "author_url": "",
      "post_date": "2022-09-10T03:09:11.483000",
      "content": "<p>Thank you for sharing! Looks like using NN Transformer will exploit the information from historical data. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1928924,
      "author_name": "Josmar Augusto Fonseca Barbosa",
      "author_url": "",
      "post_date": "2022-09-06T17:43:51.033000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for sharing and for detailed explanation!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1928319,
      "author_name": "Rana Zaryab",
      "author_url": "",
      "post_date": "2022-09-06T12:18:14.833000",
      "content": "<p>Great Very Helpful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1928308,
      "author_name": "ACHOURI",
      "author_url": "",
      "post_date": "2022-09-06T12:12:01.890000",
      "content": "<p><strong>Congratulations</strong> and thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1924679,
      "author_name": "Saksham Parashar",
      "author_url": "",
      "post_date": "2022-09-03T09:39:22.900000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> first of all congratulations of the win, can you tell me some resources for learning more about knowledge distillation, I tried looking it up myself and all I could understand that it is a form of model compression.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1923445,
      "author_name": "Nikunj Prajapati",
      "author_url": "",
      "post_date": "2022-09-02T08:29:24.910000",
      "content": "<p>Great solution, I wasn't aware we can do an ensemble of LightGBM and NNs, thanks</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1927728,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-09-05T23:28:11.910000",
          "content": "<p>The power of ensemble is using diverse models. Using both GBM and NN makes a great high performing ensemble!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1922339,
      "author_name": "Jupytor1",
      "author_url": "",
      "post_date": "2022-09-01T12:18:27.167000",
      "content": "<p>Congratulations and thank you for sharing!<br>\nThis helps a lot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1920709,
      "author_name": "moon",
      "author_url": "",
      "post_date": "2022-08-31T10:41:29.600000",
      "content": "<p>great job.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1919925,
      "author_name": "Pratyush Ingale",
      "author_url": "",
      "post_date": "2022-08-30T18:40:20.373000",
      "content": "<p>this is great</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1918821,
      "author_name": "Piyush Paliwal",
      "author_url": "",
      "post_date": "2022-08-29T21:50:05.217000",
      "content": "<p>Chris, ins't your rank 14th? Title says 15th place solution. I just realized that after final announcement of cheaters removal a couple of days ago, today again rank was gone up by 1. Not sure if they will keep changing the ranks every day.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1918950,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-30T02:15:31.903000",
          "content": "<p>Thank you, you are right. My rank is 14th now. I'll update the title soon.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1918294,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2022-08-29T13:23:50.233000",
      "content": "<p>Amazing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> its really creative to use Knowledge distillation from models like Lightgbm to a Deep Learning model like Transformer !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1923119,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-09-02T02:37:16.910000",
          "content": "<p>Thanks Athar!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1917748,
      "author_name": "GodGod3",
      "author_url": "",
      "post_date": "2022-08-29T03:23:17.030000",
      "content": "<p>Difference between performance in private and public leaderboard for NN is really large，are codes completely the same？</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1917757,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-08-29T03:31:18.503000",
          "content": "<p>yes code is the same. It appears that all teams' models did <code>+0.007</code> better in private LB than public LB.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1917648,
      "author_name": "Making TARS",
      "author_url": "",
      "post_date": "2022-08-29T00:24:51.773000",
      "content": "<p>Learned a lot! Thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1917148,
      "author_name": "Pasha Khan",
      "author_url": "",
      "post_date": "2022-08-28T13:13:16.013000",
      "content": "<p>Very Helpful 😃</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914760,
      "author_name": "Sai Siddhaarth",
      "author_url": "",
      "post_date": "2022-08-26T11:48:37.260000",
      "content": "<p>Great Job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914649,
      "author_name": "llllllkkkkkk",
      "author_url": "",
      "post_date": "2022-08-26T09:16:35.800000",
      "content": "<p>Pro, thanks for sharing, i learned very much from this notebook!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914491,
      "author_name": "Michael Hoon",
      "author_url": "",
      "post_date": "2022-08-26T06:20:38.957000",
      "content": "<p>Congrats! Great approach and thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914489,
      "author_name": "Shahidul Islam Zahid",
      "author_url": "",
      "post_date": "2022-08-26T06:13:02.403000",
      "content": "<p>Awesome ****</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914483,
      "author_name": "Ashish Motwani",
      "author_url": "",
      "post_date": "2022-08-26T06:01:07.447000",
      "content": "<p>Awesome approach</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913875,
      "author_name": "Lohit Kamble",
      "author_url": "",
      "post_date": "2022-08-25T14:54:43.147000",
      "content": "<p>Congrats and thank you for sharing the this. This is useful. :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913289,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T09:13:43.023000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913253,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T08:57:53.900000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913189,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T07:56:46.253000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1914063,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T17:32:50.330000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1913185,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T07:52:30.693000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912940,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T04:07:01.907000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1918729,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-29T19:31:39.317000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1912889,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T02:48:32.793000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1913523,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T10:51:50.077000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1913766,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T13:52:21.510000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1912885,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T02:46:16.313000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912824,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T01:38:21.270000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912813,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T01:24:51.713000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912812,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T01:24:46.177000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912777,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:59:13.013000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912770,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:53:21.617000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912763,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:46:15.267000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912761,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:45:24.253000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912756,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:38:43.770000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1912819,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T01:33:38.957000",
          "content": "",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1912823,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T01:38:18.830000",
          "content": "",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1912899,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T03:07:14.053000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1912749,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:34:20.040000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1912757,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T00:40:17.340000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1912809,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T01:23:48.847000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1912832,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T01:47:26.023000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1912847,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T02:06:20.950000",
          "content": "",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1912854,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T02:16:14.337000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1912744,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:32:05.393000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1912743,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:31:35.793000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2001917,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-10-24T11:51:23.633000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1923491,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-02T09:05:59.083000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1926177,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-09-04T15:59:51.840000",
          "content": "",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1923287,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-02T05:40:06.743000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1919936,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-30T18:53:05.317000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1915172,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-26T18:10:32.083000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1915189,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-26T18:21:04.197000",
          "content": "",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1915207,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-26T18:36:17.863000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1915238,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-26T19:30:28.810000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1915259,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-26T20:04:10.323000",
          "content": "",
          "votes": 10,
          "replies": []
        },
        {
          "id": 1915289,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-26T20:30:53.407000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1915426,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-27T01:04:32.583000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2007961,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-28T15:51:46.423000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2007973,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-10-28T16:06:53.353000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1913524,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T10:52:00.047000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1913646,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T12:32:33.973000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1913671,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T12:43:28.020000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1913209,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T08:17:44.117000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1913199,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T08:08:59.647000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1912846,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T02:04:34.060000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1912844,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T02:01:15.543000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1912738,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:26:31.007000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1912924,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T03:44:37.663000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3115974,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-02-05T13:40:20.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3113218,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-02-02T11:19:45.390000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1915234,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-26T19:26:40.610000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1912752,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T00:34:56.180000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1912755,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T00:38:29.633000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1912762,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T00:45:57.697000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1912807,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T01:22:04.757000",
          "content": "",
          "votes": 14,
          "replies": []
        },
        {
          "id": 1913041,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T05:27:29.723000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1913266,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-25T09:04:33.497000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1914845,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-08-26T13:24:46.957000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1918806,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-29T21:20:55.863000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913841,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T14:33:46.177000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1930180,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-07T16:15:44.677000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1929633,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-07T08:27:26.320000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1928429,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-06T13:19:18",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1928301,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-06T12:04:40.233000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1927744,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-05T23:56:21.847000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1923739,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-02T13:46:48.913000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1922034,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-09-01T08:27:41.490000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1921405,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-31T20:10:31.217000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1921333,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-31T18:18:38.810000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1921188,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-31T16:23:26.557000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1920117,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-30T23:13:11.357000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1919935,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-30T18:52:36.293000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1919862,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-30T17:54:33.760000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1919390,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-30T11:15:04.187000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1917815,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-29T05:09:27.730000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1917762,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-29T03:39:17.373000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1916709,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-28T05:43:34.583000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1916708,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-28T05:43:17.270000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1915801,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-27T11:39:08.647000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1915736,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-27T09:14:49.973000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1915664,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-27T07:04:19.957000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1915474,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-27T02:29:14.883000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1915408,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-27T00:28:27.453000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914853,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-26T13:29:00.257000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914640,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-26T08:57:00.113000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1914389,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-26T04:15:24.160000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1913871,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-25T14:53:26.087000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1912731": "Thank you Amex for sharing your data and hosting a fun tabular competition. Thank you Kaggle. Thank you Kagglers for sharing many helpful discussions and notebooks. Thank you Raddar and Martin for your contributions.\n\n# Solution Overview\nMy solution is a 50%/50% ensemble of LGBM and NN Transformer. The LGBM is based on Martin’s amazing public LGBM [here][1] and the NN Transformer is based on my public Transformer [here][2]. \n\nThe secret sauce is how we train the Transformer. We first use knowledge distillation from our trained LGBM before fine tuning with the train targets. Furthermore, both train and test data are used for knowledge distillation which helps the Transformer learn the test data distribution.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/summary.png)\n\n# NN Transformer Training\nMy public notebook transformer has 2 layers with skip connections (to make training easy). When using knowledge distillation, we can train a deeper transformer successfully. My final solution uses a 4 layer transformer without skip connections. We also added a GRU layer after transformer blocks and before final classification layers.\n\nTraining is done using 4 cycles of cosine learning schedule. In the first cold start cosine cycle, we pretrain (i.e. Knowldege Distillation) the Transformer using concatenated rows of both LGBM OOF preds and LGBM test preds and leave probabilities between 0 and 1 (i.e. soft labels). During the second cosine cycle, we use a warm start, reduce the learning rate and train with the hard (0 or 1) train targets. For the third and fourth cycle, we repeat cycles one and two.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2022/schedule.png)\n\n# Model Performance\nWhen creating our submission.csv file from our two models, we can use the normal **K-Fold** LGBM OOF preds and normal LGBM test preds. So making a submission is fast an easy. Additionally, we average 5 seeds per model (and slight model variations) for improved performance.\n\nTo tune our two models and compute optimal hyperparameters, we need a leak free reliable CV score. Leak-free CV score is created using **Nested K-Fold** CV. We divide each of 10 outer folds into 10 inner folds. We then train 100 models using GBT. Then each 1 of 10 outer folds has its unique OOF preds and unique test preds. These individualized OOF and test preds are created using only train targets from within the corresponding outer fold train data.\n\nWhen computing leak free CV score, we find that our NN Transformer has **CV 0.798 / LB 0.799**, our LGBM has **CV 0.799 / LB 0.799** and our 50%/50% ensemble has **CV 0.800 / LB 0.801**.\n\n# Fast Experimentation\nFast experimention was done using GPU. Thank you [Nvidia][5] for providing me compute resources for this competition. Experiments were accelerated using 4xV100 32GB.\n\nFeature engineering was explored using [RAPIDS cuDF][3] which performs operations like dataframe groupby aggregation on GPU 10-100x faster than using CPU. Many GBT experimental models were trained and evaluated using fast GPU XGB. With 1xV100 GPU, XGB can train 100 models for Nested 10 in 10 K-Fold (i.e. 100 models) on full data in only 2 hours. \n\nFeature selection was performed using both XGB feature importance and permutation importance. Using [RAPIDS FIL][4], we can perform permutation importance where we randomly shuffle each feature column 10 times for each of 10 folds (and average 100 results) in blazing speed! \n\nEach of 1000s of feature columns, we infer 100 times. This is a total of 100,000s of model inferences where each model is 1000s of individual trees! Using [RAPIDs FIL][4], we can perform this quickly! Note that we can even take an existing CPU LGBM Dart model and convert it into a GPU [RAPIDS FIL][4] inference model and perform permutation importance on existing LGBM Dart Models in blazing speed!\n\n[1]: https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7977\n[2]: https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-790\n[3]: https://rapids.ai/\n[4]: https://docs.rapids.ai/api/cuml/stable/api.html#forest-inferencing\n[5]: https://www.nvidia.com/en-us/",
    "1912765": "Awesome solution Chris, congrats!\n\nI'm curious about the data preparation you did for the transformer - was it the same as what you use in your public notebook? I spent a lot of time trying to trying to figure out the best way to represent missing statements, whether masking would help, etc. Got some boosts from imputing all the nulls with lgbm and using data augmentation (shift statements forward 1), but could never do better than .793 CV.",
    "2212438": "Thanks for sharing this  solution. It is very much helpfull for fresh data sientist as me. \n\nI couldnt understand what NN is? I got that lgbm is a machine learning model and ı also apply this for many competitions. But ı couldnt understand NN and how you can combine with lgbm ?",
    "1914862": "Congrats with 17th place and thanks for sharing your work. If I understand correctly we use distillation of knowledge from larger model to smaller while you did from smaller to larger (from LGBM to Transformer) or am I wrong somewhere? And what exactly gave distillation in this case, why couldn't you just make an ensemble of two models without distillation?",
    "1913464": "Congrats for the medal and thanks for sharing all that insightful code.",
    "1913394": "What exactly do you mean by knowledge distillation? How do you extract it from the LGBM? Is it simply stacking its prediction?\nAmazing work btw and congrats! ",
    "1912927": "Congratulations on great finish and sharing the solution Chris. I wonder how well was your NN transformer cv-lb aligned after using Knowledge Distillation based on soft preds? I usually see a dramatic change in CV, but lesser improvement in Leaderboard, couldn't think of any possible way of leakage. ",
    "1912919": "Thank you for sharing this. There has been knowledge distillation in the Feedback competition as well (ending yesterday), something I really need to catch up.  Am I right to say that it's a student teacher model , where the student is trained with the prediction result of the teacher instead of the hard label ? \nCongratulations for your result, and thanks again for all the sharings. ",
    "1912747": "Congratulations, Chris, you are the real MVP! I wonder which features had most importance",
    "1914213": "Awesome approach really Amazing solution",
    "1932845": "Thank you for sharing! Looks like using NN Transformer will exploit the information from historical data. ",
    "1928924": "Thank you @cdeotte for sharing and for detailed explanation!",
    "1928319": "Great Very Helpful",
    "1928308": "**Congratulations** and thanks for sharing @cdeotte",
    "1924679": "Hi, @cdeotte first of all congratulations of the win, can you tell me some resources for learning more about knowledge distillation, I tried looking it up myself and all I could understand that it is a form of model compression.",
    "1923445": "Great solution, I wasn't aware we can do an ensemble of LightGBM and NNs, thanks",
    "1922339": "Congratulations and thank you for sharing!\nThis helps a lot!",
    "1920709": "great job.",
    "1919925": "this is great",
    "1918821": "Chris, ins't your rank 14th? Title says 15th place solution. I just realized that after final announcement of cheaters removal a couple of days ago, today again rank was gone up by 1. Not sure if they will keep changing the ranks every day.",
    "1918294": "Amazing @cdeotte its really creative to use Knowledge distillation from models like Lightgbm to a Deep Learning model like Transformer !",
    "1917748": "Difference between performance in private and public leaderboard for NN is really large，are codes completely the same？\n",
    "1917648": "Learned a lot! Thanks for sharing @cdeotte ",
    "1917148": "Very Helpful 😃",
    "1914760": "Great Job!",
    "1914649": "Pro, thanks for sharing, i learned very much from this notebook!",
    "1914491": "Congrats! Great approach and thanks for sharing",
    "1914489": "Awesome ****",
    "1914483": "Awesome approach\n",
    "1913875": "Congrats and thank you for sharing the this. This is useful. :)",
    "1913289": "Truly impressive solution.",
    "1913253": "Congrats and thanks for all the resource ",
    "1913189": "@cdeotte Congratulations , combination of NN Transformer and Knowledge Distillation is really nice :D",
    "1913185": "Congrats @cdeotte! This was my first Kaggle competition and I learned a ton from your work. Thank you!",
    "1912940": "Congrats @cdeotte on great solo gold medal. Ensemble NN (with Knowledge Distillation) + tree-based is powerful",
    "1912889": "haha，I trained LGBM using  NN Transformer Knowledge Distillation, but got a poor result",
    "1912885": "Truly impressive solution. I wonder if I will ever be able to do something like that. For now I'm ok if I just keep learning from you. Congrats, Chris.",
    "1912824": "Congratulations and a really Nice innovative approach @cdeotte Chris and thanks for sharing! ",
    "1912813": "Congrat @cdeotte! Thanks for sharing your approach. Very insightful for someone like myself who is starting his kaggle career.",
    "1912812": "That's so \nAmazing!",
    "1912777": "congrats～～～",
    "1912770": "🎉 Woot woot! Congratulations, and thanks for all the great notebooks, learned a lot ",
    "1912763": "Congratulations Chris!",
    "1912761": "Fantastic write-up, thanks for sharing. I used a somewhat similar process, just not as effectively: \n* XG Boost on GPU for quick prototyping and feature importance \n* A Variety of seeds and folds blended via the GBM Model\n* I only tried GBDT once.... it did \"okay\" blended with DART on the public, but was my 2nd best model on the private. \n\nDidn't do the permutation importance, or knowledge distillation or leverage a completely different model like the NN Transformer for my ensembles, nor train 1000s of models. Will definitely add those into my bag of tricks for next time. Knowledge distillation will definitely be part of my weekend reading.   \n\n",
    "1912756": "Congratulations Chris! Amazing as ever! \nI had much difficulty performing permutation importance with Dart LGBM. It is amazing to learn there is actually a way to make it work. How much improvement did `RAPIDS FIL` provide? Is it compatible with all sorts of tree models?",
    "1912749": "Congrats Chris! And thanks for your GRU starter. I applied NN first time in this competition.\nI did ensemble my own LGBs, XGBs and a GRU. My GRU was very poor, so not a great boost.",
    "1912744": "Congratulation! Could you recommend me some papers like \"knowledge distillation\"? I feel that fun.",
    "1912743": "Wow! Well this is an elegant solution, you put that hardware to great use.",
    "2001917": "Looks good",
    "1923491": "A solution with class as always, thanks for sharing and congratulations @cdeotte ! A couple of questions:\n\n**1.** In which cases do you think is worth to try knowledge distillation in order to increase the performance of a model?\n**2.** How did you figure out the learning rate schedule showed in the picture?\n**3.** To generate the test preds you averaged (bagging) the 10 folds models (also with several seeds), correct? I mean, you didn't retrain with all the available data.",
    "1923287": "An interesting solution. Thank you so much for the detailed description!",
    "1919936": "i am done with this",
    "1915172": "Congratulations Chris, I have tried RAPIDs FIL, and it's pretty fast. Although with model explanatory libraries like dalex, eli5 or scikit  I needed to do some tweaks to get permutation importance working with FIL, and the performance gets reduced (mostly by sending to GPU per iteration) do you know any library that can be easily integrated?",
    "1913524": "Thanks a lot for the great sharing! I really learned a lot! \nMay I ask how much the feature selection improved your single-lgbm-model CV score? how many features did you generate and how many features did you keep in your final model? ",
    "1913209": "Congrats and thans for all the post during the competition",
    "1913199": "Thanks for the write-up! The pseudo label/distilation seems like a tight-rope walk for leakage and over-fitting (if you don't know what you're doing like me).",
    "1912846": "Thank you and congratulations Chris! Your public models helped me get started.",
    "1912844": "Love the solution, how elegant it was, wow 😊 Huge congrats @cdeotte, and thank you very much for the write-up!\n\nI am also super impressed by the software engineering that must have gone into this, both to train and do feature selection with nested CV, and to run training and inference at scale.\n\nHuge congrats! 🥳",
    "1912738": "Congratulations!",
    "1912924": "I'm really interested in transformer model recently. Someone, please instruct me. By the way I envy cooperation with Nvidia.",
    "3115974": "Thanks for sharing such a detailed approach! It helped me understand the concepts as a newbie data science enthusiast. @cdeotte",
    "3113218": "@cdeotte, Can you Please public the knowledge distillation code if available now.",
    "1915234": "https://www.kaggle.com/datasets/ashishmotwani/happyface",
    "1912752": "Great work @cdeotte and Idea. The transformer we were not able to get much juice from though we went here with tabtransformer (0.795) https://www.kaggle.com/code/gauravbrills/tabtransformer-training ., I think knowledge distillation was the trick here ,definitely the secret sauce here is just awesome  \n\n Also one curious question I had as u had tried the GRU model with 13 seq length .. Were these model successfully for you, as we did try a similar one with TCN wavenet but couldn't score well.",
    "1918806": "",
    "1913841": "Congrats and thanks for sharing",
    "1930180": "Thank you for sharing.",
    "1929633": "Thanks, great explanation\n",
    "1928429": "Thanks for sharing!",
    "1928301": "thanks, dude\n",
    "1927744": "Thanks for sharing @cdeotte ",
    "1923739": "Thank you for sharing",
    "1922034": "Thanks for sharing :)",
    "1921405": "Very useful. Thank you.",
    "1921333": "Great job! Learned a lot Thanks!🔥",
    "1921188": "thanks for sharing",
    "1920117": "Thanks for sharing!",
    "1919935": "Thanks for sharing",
    "1919862": "Congrats and thanks for sharing!",
    "1919390": "Thanks for sharing!",
    "1917815": "Great. Thanks, Sir ",
    "1917762": "Thanks for sharing!",
    "1916709": "Thank You Chris :)",
    "1916708": "Congrats and thanks for sharing!",
    "1915801": "Thank You :)",
    "1915736": "Thank You :)\n",
    "1915664": "Great, thanks",
    "1915474": "Well explained Sir, Thanks.",
    "1915408": "understood, thanks",
    "1914853": "Congratulations and thanks for sharing!",
    "1914640": "Congrats and thanks for sharing",
    "1914389": "Thanks for sharing this approach\n",
    "1913871": "Congrats and thanks for sharing."
  }
}