{
  "id": 508357,
  "title": "Use of adversarial validation for future competitions",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/508357",
  "author_name": "",
  "post_date": "2024-05-29T09:49:52.753243Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>After reading the writeups of many top solutions, a couple of them used the test data during training - <code>adversarial validation</code> - so that the models learn the test set distribution. </p>\n<p>For future competitions, is this practice endorsed or discouraged?</p>",
  "messages": [
    {
      "id": "2842872",
      "postDate": "05/29/2024 09:49:52",
      "content": "<p>After reading the writeups of many top solutions, a couple of them used the test data during training - <code>adversarial validation</code> - so that the models learn the test set distribution. </p>\n<p>For future competitions, is this practice endorsed or discouraged?</p>",
      "rawMarkdown": "After reading the writeups of many top solutions, a couple of them used the test data during training - `adversarial validation` - so that the models learn the test set distribution. \n\nFor future competitions, is this practice endorsed or discouraged?",
      "votes": null
    },
    {
      "id": "2843006",
      "postDate": "05/29/2024 11:37:30",
      "content": "<p>It doesn't make sense from the perspective of how the models are deployed in real world. When you deploy the model you should not retrain it later when in production (unless it is monitored substantially).</p>\n<p>However, it was shown that this can provide a benefit in terms of performance, so maybe it is something that could be used in the production with regular adversarial validation even if the target is not yet visible (we don't know immediately after approval if the client will be unable to pay us back in future).</p>\n<p>Then there is a kaggle competition reality and in my opinion, it doesn't make that big of difference from the host's perspective. In this competition, it was definitely wise to leverage this knowledge from kaggler's perspective. </p>",
      "rawMarkdown": "It doesn't make sense from the perspective of how the models are deployed in real world. When you deploy the model you should not retrain it later when in production (unless it is monitored substantially).\n\nHowever, it was shown that this can provide a benefit in terms of performance, so maybe it is something that could be used in the production with regular adversarial validation even if the target is not yet visible (we don't know immediately after approval if the client will be unable to pay us back in future).\n\nThen there is a kaggle competition reality and in my opinion, it doesn't make that big of difference from the host's perspective. In this competition, it was definitely wise to leverage this knowledge from kaggler's perspective.",
      "votes": null
    },
    {
      "id": "2844297",
      "postDate": "05/30/2024 03:20:29",
      "content": "<p>In my view, adversarial validation could be a useful tool even in real world. In this competition, the test data (customers who get approval of loan application) is used for adversarial validation. In real world, all the application could be used for adversarial validation, cause we don't need target. The point is to decrease the distribution difference between train datasets(approved customers one or two years ago) and the datasets in real production(recent applicant). </p>",
      "rawMarkdown": "In my view, adversarial validation could be a useful tool even in real world. In this competition, the test data (customers who get approval of loan application) is used for adversarial validation. In real world, all the application could be used for adversarial validation, cause we don't need target. The point is to decrease the distribution difference between train datasets(approved customers one or two years ago) and the datasets in real production(recent applicant).",
      "votes": null
    },
    {
      "id": "2845005",
      "postDate": "05/30/2024 10:40:50",
      "content": "<p>Can anyone briefly explain what \"adversarial validation\" means in this context? From what I have found on the internet, it means merging train and dataset and then train a model to see if it can predict whether an observation comes from the train or test datasets. I see that this can be used for the metric hacking. But what else apart from that? maybe feature selection (remove features that change a lot between datasets)?</p>",
      "rawMarkdown": "Can anyone briefly explain what \"adversarial validation\" means in this context? From what I have found on the internet, it means merging train and dataset and then train a model to see if it can predict whether an observation comes from the train or test datasets. I see that this can be used for the metric hacking. But what else apart from that? maybe feature selection (remove features that change a lot between datasets)?",
      "votes": null
    },
    {
      "id": "2845339",
      "postDate": "05/30/2024 14:14:21",
      "content": "<p><a href=\"https://www.kaggle.com/diegoiglesias\" target=\"_blank\">@diegoiglesias</a></p>\n<p>I'm not an expert, but from what I understand, a model learns based on the distributions of both the training and test data. For instance, imagine you're predicting stock prices over a two year period. The train data is the first year and the test data the second year. Suppose in the first year the stock price remains the same, but in the second year it soars by 300% (like the Nvidia stock these days). If you train on the first year and test on the second year, the model will not anticipate an exponential increase, but a linear one. </p>\n<p>By using adversarial validation, you include both the training and test data during training, thus the model can better adapt to shifts in the data distribution, ultimately improving its ability to generalize to new data and make more accurate predictions in real-world scenarios.</p>\n<blockquote>\n  <p>But what else apart from that? maybe feature selection (remove features that change a lot between datasets)?</p>\n</blockquote>\n<p>Adversarial validation can aid in identifying features that contribute to discrepancies in the distributions of the train and test data. By removing such features, the model learns more robust and generalizable patterns from the data.</p>\n<p>This is an example I thought of in the past few days, when I first saw the term <code>Adversarial validation</code>.</p>",
      "rawMarkdown": "diegoiglesias\n\nI'm not an expert, but from what I understand, a model learns based on the distributions of both the training and test data. For instance, imagine you're predicting stock prices over a two year period. The train data is the first year and the test data the second year. Suppose in the first year the stock price remains the same, but in the second year it soars by 300% (like the Nvidia stock these days). If you train on the first year and test on the second year, the model will not anticipate an exponential increase, but a linear one. \n\nBy using adversarial validation, you include both the training and test data during training, thus the model can better adapt to shifts in the data distribution, ultimately improving its ability to generalize to new data and make more accurate predictions in real-world scenarios.\n\n> But what else apart from that? maybe feature selection (remove features that change a lot between datasets)?\n\nAdversarial validation can aid in identifying features that contribute to discrepancies in the distributions of the train and test data. By removing such features, the model learns more robust and generalizable patterns from the data.\n\nThis is an example I thought of in the past few days, when I first saw the term `Adversarial validation `.",
      "votes": null
    },
    {
      "id": "2846076",
      "postDate": "05/30/2024 21:13:18",
      "content": "<p>Could still be used to detect drift from the training propulation etc. And there are (imo) probably other potential use cases as well.</p>",
      "rawMarkdown": "Could still be used to detect drift from the training propulation etc. And there are (imo) probably other potential use cases as well.",
      "votes": null
    },
    {
      "id": "2846084",
      "postDate": "05/30/2024 21:29:14",
      "content": "<p>When you don't have enough data, in practice you can use all the data available for the final fit. Again, there is kaggle reality and then the reality. For the problems Home Credit deals inhouse this is not the case, because the number of observations is great enough. Using 80% of train or 90% of train usually makes little of a difference, so adversarial validation would improve the reported results and maybe even the results in production, but if it takes the same amount of time as to prepare normal simple model, it is not worth it usually. You can't spend 4 months just on model preparation. Also, please note that we consider the FE as the most important step and usually better cleaner features boost your performance by a lot more than perfecting the model training. </p>\n<p>What I mean with my first comment is that it would be interesting to use adversarial validation as a means to prepare more stable model in time. </p>",
      "rawMarkdown": "When you don't have enough data, in practice you can use all the data available for the final fit. Again, there is kaggle reality and then the reality. For the problems Home Credit deals inhouse this is not the case, because the number of observations is great enough. Using 80% of train or 90% of train usually makes little of a difference, so adversarial validation would improve the reported results and maybe even the results in production, but if it takes the same amount of time as to prepare normal simple model, it is not worth it usually. You can't spend 4 months just on model preparation. Also, please note that we consider the FE as the most important step and usually better cleaner features boost your performance by a lot more than perfecting the model training. \n\nWhat I mean with my first comment is that it would be interesting to use adversarial validation as a means to prepare more stable model in time.",
      "votes": null
    },
    {
      "id": "2846901",
      "postDate": "05/31/2024 10:32:48",
      "content": "<p>hey Andreas, thanks for the detailed response. So, if I am getting it right, in the stocks example, adversarial validation can help us identify differences in the distributions of the train / test datasets. And to minimize the impact that can have over the model, you may end up training with a dataset composed of the train and test datasets. Would this imply doing pseudo-labeling (train first with train data, make preds on test data, and train again with whole dataset (train part using true labels, test part using preds)?</p>",
      "rawMarkdown": "hey Andreas, thanks for the detailed response. So, if I am getting it right, in the stocks example, adversarial validation can help us identify differences in the distributions of the train / test datasets. And to minimize the impact that can have over the model, you may end up training with a dataset composed of the train and test datasets. Would this imply doing pseudo-labeling (train first with train data, make preds on test data, and train again with whole dataset (train part using true labels, test part using preds)?",
      "votes": null
    },
    {
      "id": "2847767",
      "postDate": "05/31/2024 17:14:57",
      "content": "<p>In adversarial validation, you train two models jointly - the main model for the original task, and an adversarial model that tries to distinguish between the representations learned by the main model on the train vs test data. The goal is to train the main model such that the adversarial model cannot reliably discriminate between the train and test representations. </p>\n<p>So no, pseudo labeling is not involved. However, you do use the test data itself during this adversarial training process to make the main model more robust to distribution shifts. The adversarial validation helps the main model learn more generalizable representations to handle distribution changes in the test set.</p>\n<p>(take my response with a grain of salt since the term is new to me as well)</p>",
      "rawMarkdown": "In adversarial validation, you train two models jointly - the main model for the original task, and an adversarial model that tries to distinguish between the representations learned by the main model on the train vs test data. The goal is to train the main model such that the adversarial model cannot reliably discriminate between the train and test representations. \n\nSo no, pseudo labeling is not involved. However, you do use the test data itself during this adversarial training process to make the main model more robust to distribution shifts. The adversarial validation helps the main model learn more generalizable representations to handle distribution changes in the test set.\n\n(take my response with a grain of salt since the term is new to me as well)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2843006,
      "author_name": "jetakow",
      "author_url": "",
      "post_date": "05/29/2024 11:37:30",
      "content": "<p>It doesn't make sense from the perspective of how the models are deployed in real world. When you deploy the model you should not retrain it later when in production (unless it is monitored substantially).</p>\n<p>However, it was shown that this can provide a benefit in terms of performance, so maybe it is something that could be used in the production with regular adversarial validation even if the target is not yet visible (we don't know immediately after approval if the client will be unable to pay us back in future).</p>\n<p>Then there is a kaggle competition reality and in my opinion, it doesn't make that big of difference from the host's perspective. In this competition, it was definitely wise to leverage this knowledge from kaggler's perspective. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2844297,
          "author_name": "evan918",
          "author_url": "",
          "post_date": "05/30/2024 03:20:29",
          "content": "<p>In my view, adversarial validation could be a useful tool even in real world. In this competition, the test data (customers who get approval of loan application) is used for adversarial validation. In real world, all the application could be used for adversarial validation, cause we don't need target. The point is to decrease the distribution difference between train datasets(approved customers one or two years ago) and the datasets in real production(recent applicant). </p>",
          "votes": null,
          "replies": [
            {
              "id": 2846084,
              "author_name": "jetakow",
              "author_url": "",
              "post_date": "05/30/2024 21:29:14",
              "content": "<p>When you don't have enough data, in practice you can use all the data available for the final fit. Again, there is kaggle reality and then the reality. For the problems Home Credit deals inhouse this is not the case, because the number of observations is great enough. Using 80% of train or 90% of train usually makes little of a difference, so adversarial validation would improve the reported results and maybe even the results in production, but if it takes the same amount of time as to prepare normal simple model, it is not worth it usually. You can't spend 4 months just on model preparation. Also, please note that we consider the FE as the most important step and usually better cleaner features boost your performance by a lot more than perfecting the model training. </p>\n<p>What I mean with my first comment is that it would be interesting to use adversarial validation as a means to prepare more stable model in time. </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2846076,
          "author_name": "ern711",
          "author_url": "",
          "post_date": "05/30/2024 21:13:18",
          "content": "<p>Could still be used to detect drift from the training propulation etc. And there are (imo) probably other potential use cases as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2845005,
      "author_name": "diegoiglesias",
      "author_url": "",
      "post_date": "05/30/2024 10:40:50",
      "content": "<p>Can anyone briefly explain what \"adversarial validation\" means in this context? From what I have found on the internet, it means merging train and dataset and then train a model to see if it can predict whether an observation comes from the train or test datasets. I see that this can be used for the metric hacking. But what else apart from that? maybe feature selection (remove features that change a lot between datasets)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2845339,
          "author_name": "andreasbis",
          "author_url": "",
          "post_date": "05/30/2024 14:14:21",
          "content": "<p><a href=\"https://www.kaggle.com/diegoiglesias\" target=\"_blank\">@diegoiglesias</a></p>\n<p>I'm not an expert, but from what I understand, a model learns based on the distributions of both the training and test data. For instance, imagine you're predicting stock prices over a two year period. The train data is the first year and the test data the second year. Suppose in the first year the stock price remains the same, but in the second year it soars by 300% (like the Nvidia stock these days). If you train on the first year and test on the second year, the model will not anticipate an exponential increase, but a linear one. </p>\n<p>By using adversarial validation, you include both the training and test data during training, thus the model can better adapt to shifts in the data distribution, ultimately improving its ability to generalize to new data and make more accurate predictions in real-world scenarios.</p>\n<blockquote>\n  <p>But what else apart from that? maybe feature selection (remove features that change a lot between datasets)?</p>\n</blockquote>\n<p>Adversarial validation can aid in identifying features that contribute to discrepancies in the distributions of the train and test data. By removing such features, the model learns more robust and generalizable patterns from the data.</p>\n<p>This is an example I thought of in the past few days, when I first saw the term <code>Adversarial validation</code>.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2846901,
              "author_name": "diegoiglesias",
              "author_url": "",
              "post_date": "05/31/2024 10:32:48",
              "content": "<p>hey Andreas, thanks for the detailed response. So, if I am getting it right, in the stocks example, adversarial validation can help us identify differences in the distributions of the train / test datasets. And to minimize the impact that can have over the model, you may end up training with a dataset composed of the train and test datasets. Would this imply doing pseudo-labeling (train first with train data, make preds on test data, and train again with whole dataset (train part using true labels, test part using preds)?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2847767,
                  "author_name": "andreasbis",
                  "author_url": "",
                  "post_date": "05/31/2024 17:14:57",
                  "content": "<p>In adversarial validation, you train two models jointly - the main model for the original task, and an adversarial model that tries to distinguish between the representations learned by the main model on the train vs test data. The goal is to train the main model such that the adversarial model cannot reliably discriminate between the train and test representations. </p>\n<p>So no, pseudo labeling is not involved. However, you do use the test data itself during this adversarial training process to make the main model more robust to distribution shifts. The adversarial validation helps the main model learn more generalizable representations to handle distribution changes in the test set.</p>\n<p>(take my response with a grain of salt since the term is new to me as well)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2842872": "After reading the writeups of many top solutions, a couple of them used the test data during training - `adversarial validation` - so that the models learn the test set distribution. \n\nFor future competitions, is this practice endorsed or discouraged?",
    "2843006": "It doesn't make sense from the perspective of how the models are deployed in real world. When you deploy the model you should not retrain it later when in production (unless it is monitored substantially).\n\nHowever, it was shown that this can provide a benefit in terms of performance, so maybe it is something that could be used in the production with regular adversarial validation even if the target is not yet visible (we don't know immediately after approval if the client will be unable to pay us back in future).\n\nThen there is a kaggle competition reality and in my opinion, it doesn't make that big of difference from the host's perspective. In this competition, it was definitely wise to leverage this knowledge from kaggler's perspective.",
    "2844297": "In my view, adversarial validation could be a useful tool even in real world. In this competition, the test data (customers who get approval of loan application) is used for adversarial validation. In real world, all the application could be used for adversarial validation, cause we don't need target. The point is to decrease the distribution difference between train datasets(approved customers one or two years ago) and the datasets in real production(recent applicant).",
    "2845005": "Can anyone briefly explain what \"adversarial validation\" means in this context? From what I have found on the internet, it means merging train and dataset and then train a model to see if it can predict whether an observation comes from the train or test datasets. I see that this can be used for the metric hacking. But what else apart from that? maybe feature selection (remove features that change a lot between datasets)?",
    "2845339": "diegoiglesias\n\nI'm not an expert, but from what I understand, a model learns based on the distributions of both the training and test data. For instance, imagine you're predicting stock prices over a two year period. The train data is the first year and the test data the second year. Suppose in the first year the stock price remains the same, but in the second year it soars by 300% (like the Nvidia stock these days). If you train on the first year and test on the second year, the model will not anticipate an exponential increase, but a linear one. \n\nBy using adversarial validation, you include both the training and test data during training, thus the model can better adapt to shifts in the data distribution, ultimately improving its ability to generalize to new data and make more accurate predictions in real-world scenarios.\n\n> But what else apart from that? maybe feature selection (remove features that change a lot between datasets)?\n\nAdversarial validation can aid in identifying features that contribute to discrepancies in the distributions of the train and test data. By removing such features, the model learns more robust and generalizable patterns from the data.\n\nThis is an example I thought of in the past few days, when I first saw the term `Adversarial validation `.",
    "2846076": "Could still be used to detect drift from the training propulation etc. And there are (imo) probably other potential use cases as well.",
    "2846084": "When you don't have enough data, in practice you can use all the data available for the final fit. Again, there is kaggle reality and then the reality. For the problems Home Credit deals inhouse this is not the case, because the number of observations is great enough. Using 80% of train or 90% of train usually makes little of a difference, so adversarial validation would improve the reported results and maybe even the results in production, but if it takes the same amount of time as to prepare normal simple model, it is not worth it usually. You can't spend 4 months just on model preparation. Also, please note that we consider the FE as the most important step and usually better cleaner features boost your performance by a lot more than perfecting the model training. \n\nWhat I mean with my first comment is that it would be interesting to use adversarial validation as a means to prepare more stable model in time.",
    "2846901": "hey Andreas, thanks for the detailed response. So, if I am getting it right, in the stocks example, adversarial validation can help us identify differences in the distributions of the train / test datasets. And to minimize the impact that can have over the model, you may end up training with a dataset composed of the train and test datasets. Would this imply doing pseudo-labeling (train first with train data, make preds on test data, and train again with whole dataset (train part using true labels, test part using preds)?",
    "2847767": "In adversarial validation, you train two models jointly - the main model for the original task, and an adversarial model that tries to distinguish between the representations learned by the main model on the train vs test data. The goal is to train the main model such that the adversarial model cannot reliably discriminate between the train and test representations. \n\nSo no, pseudo labeling is not involved. However, you do use the test data itself during this adversarial training process to make the main model more robust to distribution shifts. The adversarial validation helps the main model learn more generalizable representations to handle distribution changes in the test set.\n\n(take my response with a grain of salt since the term is new to me as well)"
  },
  "source": "meta"
}