{
  "id": 508122,
  "title": "Techniques for Better Stability",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508122",
  "author_name": "Evan",
  "post_date": "2024-05-28T10:33:41.831000",
  "votes": 16,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Aside from how to do a better metric hacking, I would like to share some techniques that have helped me achieve better stability.</p>\n<p>The first one is adversarial training. I didn't train models offline, so during training I have access to the real test data. For categorical variables, I drop values that have very different distribution between train and test. For numeric variables, I clip train values based on the test data. After that I do adversrial validation. I gradually drop important features, since adversarial validation model tends to focus on very few features. This process is done multiple times. The importance of features is evaluated using permutation importance. Once the auc of the adversarial model reaches a certain threshold (I use multiple threholds), I start training the main model. So the final model is an ensemble of models trained at different thresholds.</p>\n<p>The second technique is pseudo lableing. I belive this helps the model learn test distribution, which is very different from train. Although its effect on public LB is not very significant, it shows a much better effect on private LB.</p>\n<p>One area I spent a lot time but didn't work out was training a neural network model that creates an indistinguishable feature space between train and test sets to classify defualt customers. This concept is from the domain of domain adaptation.</p>\n<p>Finally, I want to thank the hosts for bringing stability into the picture. This is a very intertesing topic. In my humble opinion, I think a better metirc would be just the auc of data after one or two years. The cuurent metric penalizes the downtrend of Gini so much that the auc of the recent test data doesn't even matter.</p>",
  "messages": [
    {
      "id": 2840968,
      "postDate": "2024-05-28T10:33:41.830Z",
      "content": "<p>Aside from how to do a better metric hacking, I would like to share some techniques that have helped me achieve better stability.</p>\n<p>The first one is adversarial training. I didn't train models offline, so during training I have access to the real test data. For categorical variables, I drop values that have very different distribution between train and test. For numeric variables, I clip train values based on the test data. After that I do adversrial validation. I gradually drop important features, since adversarial validation model tends to focus on very few features. This process is done multiple times. The importance of features is evaluated using permutation importance. Once the auc of the adversarial model reaches a certain threshold (I use multiple threholds), I start training the main model. So the final model is an ensemble of models trained at different thresholds.</p>\n<p>The second technique is pseudo lableing. I belive this helps the model learn test distribution, which is very different from train. Although its effect on public LB is not very significant, it shows a much better effect on private LB.</p>\n<p>One area I spent a lot time but didn't work out was training a neural network model that creates an indistinguishable feature space between train and test sets to classify defualt customers. This concept is from the domain of domain adaptation.</p>\n<p>Finally, I want to thank the hosts for bringing stability into the picture. This is a very intertesing topic. In my humble opinion, I think a better metirc would be just the auc of data after one or two years. The cuurent metric penalizes the downtrend of Gini so much that the auc of the recent test data doesn't even matter.</p>",
      "rawMarkdown": "Aside from how to do a better metric hacking, I would like to share some techniques that have helped me achieve better stability.\n\nThe first one is adversarial training. I didn't train models offline, so during training I have access to the real test data. For categorical variables, I drop values that have very different distribution between train and test. For numeric variables, I clip train values based on the test data. After that I do adversrial validation. I gradually drop important features, since adversarial validation model tends to focus on very few features. This process is done multiple times. The importance of features is evaluated using permutation importance. Once the auc of the adversarial model reaches a certain threshold (I use multiple threholds), I start training the main model. So the final model is an ensemble of models trained at different thresholds.\n\nThe second technique is pseudo lableing. I belive this helps the model learn test distribution, which is very different from train. Although its effect on public LB is not very significant, it shows a much better effect on private LB.\n\nOne area I spent a lot time but didn't work out was training a neural network model that creates an indistinguishable feature space between train and test sets to classify defualt customers. This concept is from the domain of domain adaptation.\n\nFinally, I want to thank the hosts for bringing stability into the picture. This is a very intertesing topic. In my humble opinion, I think a better metirc would be just the auc of data after one or two years. The cuurent metric penalizes the downtrend of Gini so much that the auc of the recent test data doesn't even matter.",
      "votes": 16
    },
    {
      "id": 2849009,
      "postDate": "2024-06-01T10:31:38.157Z",
      "content": "<blockquote>\n  <p>I didn't train models offline, so during training I have access to the real test data.</p>\n</blockquote>\n<p>Can you elaborate how you get access to real test data?</p>",
      "rawMarkdown": "> I didn't train models offline, so during training I have access to the real test data.\n\nCan you elaborate how you get access to real test data?",
      "votes": 1,
      "replies": [
        {
          "id": 2851896,
          "postDate": "2024-06-03T02:31:08.273Z",
          "content": "<p>When you click the submission button, the sample test files would be replaced by real test data. I train my models after I click the submit button, so my model can get access to real test data. </p>",
          "rawMarkdown": "When you click the submission button, the sample test files would be replaced by real test data. I train my models after I click the submit button, so my model can get access to real test data. "
        }
      ]
    },
    {
      "id": 2841929,
      "postDate": "2024-05-28T18:55:23.777Z",
      "content": "<p><a href=\"https://www.kaggle.com/evan918\" target=\"_blank\">@evan918</a></p>\n<p>Could you elaborate on the concept of <code>adversarial training</code>? From a quick search, I understood that it's a process of manipulating a model to make worse predictions. How does this help to achieve stability on future data?</p>",
      "rawMarkdown": "@evan918\n\nCould you elaborate on the concept of `adversarial training`? From a quick search, I understood that it's a process of manipulating a model to make worse predictions. How does this help to achieve stability on future data?",
      "votes": 1,
      "replies": [
        {
          "id": 2842203,
          "postDate": "2024-05-29T00:50:42.163Z",
          "content": "<p>I think what he said is this.<a href=\"https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation\" target=\"_blank\">https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation</a></p>",
          "rawMarkdown": "I think what he said is this.https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation",
          "votes": 2,
          "replies": [
            {
              "id": 2875530,
              "postDate": "2024-06-17T08:07:10.363Z",
              "content": "<p>Can I just say, it's nice seeing you here :D still following the competition, I thought you had disappeared after the linear regression notebook. This is my first Kaggle competition, while I can't complete on time since I can't get the submission right, your first few notebooks had been of great help so I distinctively remember you alot.</p>",
              "rawMarkdown": "Can I just say, it's nice seeing you here :D still following the competition, I thought you had disappeared after the linear regression notebook. This is my first Kaggle competition, while I can't complete on time since I can't get the submission right, your first few notebooks had been of great help so I distinctively remember you alot."
            }
          ]
        },
        {
          "id": 2842261,
          "postDate": "2024-05-29T02:05:16.313Z",
          "content": "<p>The point is to improve generalization of model by decresing the distribution difference between train and test. The auc the adversarial validation between train and real test is above 0.9 or even higher, so there are some variables like refresh_date that are very important features in adversarial validatioin model. By removing these variables, you could decrease the distribution difference between tran and test.</p>",
          "rawMarkdown": "The point is to improve generalization of model by decresing the distribution difference between train and test. The auc the adversarial validation between train and real test is above 0.9 or even higher, so there are some variables like refresh_date that are very important features in adversarial validatioin model. By removing these variables, you could decrease the distribution difference between tran and test.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2841568,
      "postDate": "2024-05-28T16:07:05.683Z",
      "content": "<p>Very interesting, thank you for sharing that. If you want to share it with more details you could create some example notebook.</p>",
      "rawMarkdown": "Very interesting, thank you for sharing that. If you want to share it with more details you could create some example notebook."
    },
    {
      "id": 2843670,
      "postDate": "2024-05-29T17:00:42.423Z",
      "content": "<p>Thank you for this post Yan. I have learned a lot from this particular thread of discussion. I will eventually try out these with my code as well so I learn these two techniques. It is very interesting. I have been trying to look up some articles and references to better understand how it works. This is for reference for others who may want to know about the intuition behind these.</p>\n<p>Adversarial validation (same as shared by yunsuxiaozi ): <a href=\"https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation\" target=\"_blank\">https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation</a></p>\n<p>Psuedo-labeling: <a href=\"https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook</a></p>\n<p>Thank you once again for the code snippets as well. Helps a lot! 🙏</p>",
      "rawMarkdown": "Thank you for this post Yan. I have learned a lot from this particular thread of discussion. I will eventually try out these with my code as well so I learn these two techniques. It is very interesting. I have been trying to look up some articles and references to better understand how it works. This is for reference for others who may want to know about the intuition behind these.\n\nAdversarial validation (same as shared by yunsuxiaozi ): https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation\n\nPsuedo-labeling: https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook\n\nThank you once again for the code snippets as well. Helps a lot! 🙏"
    },
    {
      "id": 2841028,
      "postDate": "2024-05-28T11:15:27.810Z",
      "content": "<p>Please tell me about the following questions </p>\n<ol>\n<li>\"I didn't train models offline, so during training I have access to the real test data. \", <br>\n\"For categorical variables, I drop values that have very different distribution between train and test.\"</li>\n</ol>\n<p>What a metric did you use?</p>\n<p>2-1. \"The second technique is pseudo lableing.\"<br>\nWhat models did you use? Because of the limitation of running time, we cannot use this technique. (e.g. catboost)<br>\n2-2. \"I belive this helps the model learn test distribution, which is very different from train. \"<br>\n However you remove some columns for the reason, \"very different distribution between train and test.\".  please tell me how you decided the standard whether drop a column or not.</p>\n<ol>\n<li>Did you do your original feature engineering other than usual aggregation such as max and last?</li>\n</ol>",
      "rawMarkdown": "Please tell me about the following questions \n\n\n1. \"I didn't train models offline, so during training I have access to the real test data. \", \n\"For categorical variables, I drop values that have very different distribution between train and test.\"\n\nWhat a metric did you use?\n\n2-1. \"The second technique is pseudo lableing.\"\nWhat models did you use? Because of the limitation of running time, we cannot use this technique. (e.g. catboost)\n2-2. \"I belive this helps the model learn test distribution, which is very different from train. \"\n However you remove some columns for the reason, \"very different distribution between train and test.\".  please tell me how you decided the standard whether drop a column or not.\n\n3. Did you do your original feature engineering other than usual aggregation such as max and last?",
      "replies": [
        {
          "id": 2842251,
          "postDate": "2024-05-29T01:43:14.097Z",
          "content": "<p>Hi Gita, </p>\n<p>Thanks for your interest. </p>\n<ol>\n<li>In interms of removing categorical values, I remove values that are 4 times more frequent or 1/4 lest frequent in train than test. And also values appear less than 5% frequent in train or test</li>\n</ol>\n<pre><code>col_dist = df.loc[df[] == , col].value_counts(normalize=, dropna=).to_frame()\ncol_dist[] = df.loc[df[] == , col].value_counts(normalize=, dropna=)\ncol_dist[] = col_dist[].fillna() / (col_dist[].fillna() + )\ncond0 = (col_dist[] &lt; ) | (col_dist[] &lt; )\ncond1 = (col_dist[] &lt;  / ) | (col_dist[] &gt; )\n</code></pre>\n<p>2-1. You are correct. I only use lightgbm.</p>\n<p>2-2. Removing values that are different and adding pseudo labled test data into train is not contradictory. The goal is to decrese the distribution difference between train and test.</p>\n<p>Lastly, I build upon the public engineering methods. The main changes I made are 1) OHE categorical variables and aggreate into the mean of non-null values(excluding a55475b1); 2) combining tax data source into one; 3) extract client info from person data. Num_group1 means client. </p>",
          "rawMarkdown": "Hi Gita, \n\nThanks for your interest. \n\n1. In interms of removing categorical values, I remove values that are 4 times more frequent or 1/4 lest frequent in train than test. And also values appear less than 5% frequent in train or test\n\n```python\ncol_dist = df.loc[df[\"is_test\"] == 0, col].value_counts(normalize=True, dropna=False).to_frame(\"train_dist\")\ncol_dist[\"test_dist\"] = df.loc[df[\"is_test\"] == 1, col].value_counts(normalize=True, dropna=False)\ncol_dist[\"train_test_ratio\"] = col_dist[\"train_dist\"].fillna(0) / (col_dist[\"test_dist\"].fillna(0) + 1e-5)\ncond0 = (col_dist[\"train_dist\"] < 0.05) | (col_dist[\"test_dist\"] < 0.05)\ncond1 = (col_dist[\"train_test_ratio\"] < 1 / 4) | (col_dist[\"train_test_ratio\"] > 4)\n\n```\n\n2-1. You are correct. I only use lightgbm.\n\n2-2. Removing values that are different and adding pseudo labled test data into train is not contradictory. The goal is to decrese the distribution difference between train and test.\n\nLastly, I build upon the public engineering methods. The main changes I made are 1) OHE categorical variables and aggreate into the mean of non-null values(excluding a55475b1); 2) combining tax data source into one; 3) extract client info from person data. Num_group1 means client. ",
          "votes": 1,
          "replies": [
            {
              "id": 2844233,
              "postDate": "2024-05-30T02:35:07.443Z",
              "content": "<p>Hello Yan,</p>\n<p>Thank you for answering my question! It was very helpful!</p>",
              "rawMarkdown": "Hello Yan,\n\nThank you for answering my question! It was very helpful!"
            }
          ]
        }
      ]
    },
    {
      "id": 2840972,
      "postDate": "2024-05-28T10:35:24.167Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2841237,
          "postDate": "2024-05-28T13:24:56.847Z",
          "content": "<p>May be get model refit with the pseudo label on test data？</p>",
          "rawMarkdown": "May be get model refit with the pseudo label on test data？",
          "isDeleted": true,
          "replies": [
            {
              "id": 2841267,
              "postDate": "2024-05-28T13:46:52.237Z",
              "content": "<p>Yes, the important thing is get model to learn new distribution.</p>",
              "rawMarkdown": "Yes, the important thing is get model to learn new distribution."
            },
            {
              "id": 2841834,
              "postDate": "2024-05-28T17:52:26.743Z",
              "content": "<p>Would really appreciate if you could give more details about how to achieve this. If you could please share snippets of code if not the whole notebook that helps us understand how this can be done, it will be great for a beginner like me to learn.</p>",
              "rawMarkdown": "Would really appreciate if you could give more details about how to achieve this. If you could please share snippets of code if not the whole notebook that helps us understand how this can be done, it will be great for a beginner like me to learn."
            },
            {
              "id": 2842819,
              "postDate": "2024-05-29T09:18:56.237Z",
              "content": "<p>This code snipt should be able to explain how to do pseudo labeling. Just transform the most confident predition on test data into target column and add it to train data. </p>\n<pre><code>bad_cut = df_test[].quantile()\ngood_cut = df_test[].quantile()\ndf_test = df_test[(df_test[] &gt; bad_cut) | (df_test[] &lt; good_cut)]\ndf_test[] = (df_test[] &gt; ).astype()\ndf_test.drop(, axis=, inplace=)\n\ndf_train = pd.read_parquet()\ndf_train = pd.concat([df_train, df_test[df_train.columns]], ignore_index=, copy=).reset_index(drop=)\n</code></pre>",
              "rawMarkdown": "This code snipt should be able to explain how to do pseudo labeling. Just transform the most confident predition on test data into target column and add it to train data. \n\n```python\nbad_cut = df_test[\"preds\"].quantile(0.9995)\ngood_cut = df_test[\"preds\"].quantile(0.015)\ndf_test = df_test[(df_test[\"preds\"] > bad_cut) | (df_test[\"preds\"] < good_cut)]\ndf_test[\"target\"] = (df_test[\"preds\"] > 0.5).astype(int)\ndf_test.drop(\"preds\", axis=1, inplace=True)\n\ndf_train = pd.read_parquet(\"df_train.parquet\")\ndf_train = pd.concat([df_train, df_test[df_train.columns]], ignore_index=True, copy=False).reset_index(drop=True)\n```",
              "votes": 2
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2849009,
      "author_name": "Josef Švenda",
      "author_url": "",
      "post_date": "2024-06-01T10:31:38.157000",
      "content": "<blockquote>\n  <p>I didn't train models offline, so during training I have access to the real test data.</p>\n</blockquote>\n<p>Can you elaborate how you get access to real test data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2851896,
          "author_name": "Evan",
          "author_url": "",
          "post_date": "2024-06-03T02:31:08.273000",
          "content": "<p>When you click the submission button, the sample test files would be replaced by real test data. I train my models after I click the submit button, so my model can get access to real test data. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2841929,
      "author_name": "Andreas Bisiadis",
      "author_url": "",
      "post_date": "2024-05-28T18:55:23.777000",
      "content": "<p><a href=\"https://www.kaggle.com/evan918\" target=\"_blank\">@evan918</a></p>\n<p>Could you elaborate on the concept of <code>adversarial training</code>? From a quick search, I understood that it's a process of manipulating a model to make worse predictions. How does this help to achieve stability on future data?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2842203,
          "author_name": "yunsuxiaozi",
          "author_url": "",
          "post_date": "2024-05-29T00:50:42.163000",
          "content": "<p>I think what he said is this.<a href=\"https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation\" target=\"_blank\">https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation</a></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2875530,
              "author_name": "FlyingCamel",
              "author_url": "",
              "post_date": "2024-06-17T08:07:10.363000",
              "content": "<p>Can I just say, it's nice seeing you here :D still following the competition, I thought you had disappeared after the linear regression notebook. This is my first Kaggle competition, while I can't complete on time since I can't get the submission right, your first few notebooks had been of great help so I distinctively remember you alot.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2842261,
          "author_name": "Evan",
          "author_url": "",
          "post_date": "2024-05-29T02:05:16.313000",
          "content": "<p>The point is to improve generalization of model by decresing the distribution difference between train and test. The auc the adversarial validation between train and real test is above 0.9 or even higher, so there are some variables like refresh_date that are very important features in adversarial validatioin model. By removing these variables, you could decrease the distribution difference between tran and test.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2841568,
      "author_name": "Daniel Herman",
      "author_url": "",
      "post_date": "2024-05-28T16:07:05.683000",
      "content": "<p>Very interesting, thank you for sharing that. If you want to share it with more details you could create some example notebook.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2843670,
      "author_name": "Varuni Rao",
      "author_url": "",
      "post_date": "2024-05-29T17:00:42.423000",
      "content": "<p>Thank you for this post Yan. I have learned a lot from this particular thread of discussion. I will eventually try out these with my code as well so I learn these two techniques. It is very interesting. I have been trying to look up some articles and references to better understand how it works. This is for reference for others who may want to know about the intuition behind these.</p>\n<p>Adversarial validation (same as shared by yunsuxiaozi ): <a href=\"https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation\" target=\"_blank\">https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation</a></p>\n<p>Psuedo-labeling: <a href=\"https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook</a></p>\n<p>Thank you once again for the code snippets as well. Helps a lot! 🙏</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2841028,
      "author_name": "Aria",
      "author_url": "",
      "post_date": "2024-05-28T11:15:27.810000",
      "content": "<p>Please tell me about the following questions </p>\n<ol>\n<li>\"I didn't train models offline, so during training I have access to the real test data. \", <br>\n\"For categorical variables, I drop values that have very different distribution between train and test.\"</li>\n</ol>\n<p>What a metric did you use?</p>\n<p>2-1. \"The second technique is pseudo lableing.\"<br>\nWhat models did you use? Because of the limitation of running time, we cannot use this technique. (e.g. catboost)<br>\n2-2. \"I belive this helps the model learn test distribution, which is very different from train. \"<br>\n However you remove some columns for the reason, \"very different distribution between train and test.\".  please tell me how you decided the standard whether drop a column or not.</p>\n<ol>\n<li>Did you do your original feature engineering other than usual aggregation such as max and last?</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 2842251,
          "author_name": "Evan",
          "author_url": "",
          "post_date": "2024-05-29T01:43:14.097000",
          "content": "<p>Hi Gita, </p>\n<p>Thanks for your interest. </p>\n<ol>\n<li>In interms of removing categorical values, I remove values that are 4 times more frequent or 1/4 lest frequent in train than test. And also values appear less than 5% frequent in train or test</li>\n</ol>\n<pre><code>col_dist = df.loc[df[] == , col].value_counts(normalize=, dropna=).to_frame()\ncol_dist[] = df.loc[df[] == , col].value_counts(normalize=, dropna=)\ncol_dist[] = col_dist[].fillna() / (col_dist[].fillna() + )\ncond0 = (col_dist[] &lt; ) | (col_dist[] &lt; )\ncond1 = (col_dist[] &lt;  / ) | (col_dist[] &gt; )\n</code></pre>\n<p>2-1. You are correct. I only use lightgbm.</p>\n<p>2-2. Removing values that are different and adding pseudo labled test data into train is not contradictory. The goal is to decrese the distribution difference between train and test.</p>\n<p>Lastly, I build upon the public engineering methods. The main changes I made are 1) OHE categorical variables and aggreate into the mean of non-null values(excluding a55475b1); 2) combining tax data source into one; 3) extract client info from person data. Num_group1 means client. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2844233,
              "author_name": "Aria",
              "author_url": "",
              "post_date": "2024-05-30T02:35:07.443000",
              "content": "<p>Hello Yan,</p>\n<p>Thank you for answering my question! It was very helpful!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2840972,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-28T10:35:24.167000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2841237,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-28T13:24:56.847000",
          "content": "<p>May be get model refit with the pseudo label on test data？</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2841267,
              "author_name": "Evan",
              "author_url": "",
              "post_date": "2024-05-28T13:46:52.237000",
              "content": "<p>Yes, the important thing is get model to learn new distribution.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2841834,
              "author_name": "Varuni Rao",
              "author_url": "",
              "post_date": "2024-05-28T17:52:26.743000",
              "content": "<p>Would really appreciate if you could give more details about how to achieve this. If you could please share snippets of code if not the whole notebook that helps us understand how this can be done, it will be great for a beginner like me to learn.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2842819,
              "author_name": "Evan",
              "author_url": "",
              "post_date": "2024-05-29T09:18:56.237000",
              "content": "<p>This code snipt should be able to explain how to do pseudo labeling. Just transform the most confident predition on test data into target column and add it to train data. </p>\n<pre><code>bad_cut = df_test[].quantile()\ngood_cut = df_test[].quantile()\ndf_test = df_test[(df_test[] &gt; bad_cut) | (df_test[] &lt; good_cut)]\ndf_test[] = (df_test[] &gt; ).astype()\ndf_test.drop(, axis=, inplace=)\n\ndf_train = pd.read_parquet()\ndf_train = pd.concat([df_train, df_test[df_train.columns]], ignore_index=, copy=).reset_index(drop=)\n</code></pre>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2840968": "Aside from how to do a better metric hacking, I would like to share some techniques that have helped me achieve better stability.\n\nThe first one is adversarial training. I didn't train models offline, so during training I have access to the real test data. For categorical variables, I drop values that have very different distribution between train and test. For numeric variables, I clip train values based on the test data. After that I do adversrial validation. I gradually drop important features, since adversarial validation model tends to focus on very few features. This process is done multiple times. The importance of features is evaluated using permutation importance. Once the auc of the adversarial model reaches a certain threshold (I use multiple threholds), I start training the main model. So the final model is an ensemble of models trained at different thresholds.\n\nThe second technique is pseudo lableing. I belive this helps the model learn test distribution, which is very different from train. Although its effect on public LB is not very significant, it shows a much better effect on private LB.\n\nOne area I spent a lot time but didn't work out was training a neural network model that creates an indistinguishable feature space between train and test sets to classify defualt customers. This concept is from the domain of domain adaptation.\n\nFinally, I want to thank the hosts for bringing stability into the picture. This is a very intertesing topic. In my humble opinion, I think a better metirc would be just the auc of data after one or two years. The cuurent metric penalizes the downtrend of Gini so much that the auc of the recent test data doesn't even matter.",
    "2849009": "> I didn't train models offline, so during training I have access to the real test data.\n\nCan you elaborate how you get access to real test data?",
    "2841929": "@evan918\n\nCould you elaborate on the concept of `adversarial training`? From a quick search, I understood that it's a process of manipulating a model to make worse predictions. How does this help to achieve stability on future data?",
    "2841568": "Very interesting, thank you for sharing that. If you want to share it with more details you could create some example notebook.",
    "2843670": "Thank you for this post Yan. I have learned a lot from this particular thread of discussion. I will eventually try out these with my code as well so I learn these two techniques. It is very interesting. I have been trying to look up some articles and references to better understand how it works. This is for reference for others who may want to know about the intuition behind these.\n\nAdversarial validation (same as shared by yunsuxiaozi ): https://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation\n\nPsuedo-labeling: https://www.kaggle.com/code/cdeotte/pseudo-labeling-qda-0-969/notebook\n\nThank you once again for the code snippets as well. Helps a lot! 🙏",
    "2841028": "Please tell me about the following questions \n\n\n1. \"I didn't train models offline, so during training I have access to the real test data. \", \n\"For categorical variables, I drop values that have very different distribution between train and test.\"\n\nWhat a metric did you use?\n\n2-1. \"The second technique is pseudo lableing.\"\nWhat models did you use? Because of the limitation of running time, we cannot use this technique. (e.g. catboost)\n2-2. \"I belive this helps the model learn test distribution, which is very different from train. \"\n However you remove some columns for the reason, \"very different distribution between train and test.\".  please tell me how you decided the standard whether drop a column or not.\n\n3. Did you do your original feature engineering other than usual aggregation such as max and last?",
    "2840972": ""
  }
}