{
  "id": 335398,
  "title": "Does Adversarial Validation Always help in Improving scores on unseen test data ?",
  "url": "/competitions/amex-default-prediction/discussion/335398",
  "author_name": "",
  "post_date": "2022-07-06T01:40:01.505734200Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Adversarial Validation- we build a model with unseen data(test unseen) and data used for Model training(train), with label as 1 for test unseen and 0 for train. And model gini gives us a rough idea that are the data distribution among train and unseen test is similar or not. and top importance variables tell which variables distribution is different in test unseen vs train.<br>\nAccording we can remove such high importance variables from my final model(used for default prediction) to make our Model more robust.<br>\nwe can do this analysis between test seen and train or test seen and test unseen as well.</p>\n<p>My question is doing so is making your model more stable and reliable on unseen test. But does it also improving your final scores on unseen test. for eg (Hypothetically) without Adversarial my score on test seen vs test unseen can be - 0.799 and 0.798. after dropping some variables due to adversarial it can be  0.797 and 0.797. so My model has become more robust but its overall score on unseen test has decreased.   <br>\nin above example I have shown 3rd decimal place difference but it can be 4th or 5th decimal place difference in actual so we never know does our score on Unseen test are improving or not.</p>",
  "messages": [
    {
      "id": "1845005",
      "postDate": "07/06/2022 01:40:01",
      "content": "<p>Adversarial Validation- we build a model with unseen data(test unseen) and data used for Model training(train), with label as 1 for test unseen and 0 for train. And model gini gives us a rough idea that are the data distribution among train and unseen test is similar or not. and top importance variables tell which variables distribution is different in test unseen vs train.<br>\nAccording we can remove such high importance variables from my final model(used for default prediction) to make our Model more robust.<br>\nwe can do this analysis between test seen and train or test seen and test unseen as well.</p>\n<p>My question is doing so is making your model more stable and reliable on unseen test. But does it also improving your final scores on unseen test. for eg (Hypothetically) without Adversarial my score on test seen vs test unseen can be - 0.799 and 0.798. after dropping some variables due to adversarial it can be  0.797 and 0.797. so My model has become more robust but its overall score on unseen test has decreased.   <br>\nin above example I have shown 3rd decimal place difference but it can be 4th or 5th decimal place difference in actual so we never know does our score on Unseen test are improving or not.</p>",
      "rawMarkdown": "Adversarial Validation- we build a model with unseen data(test unseen) and data used for Model training(train), with label as 1 for test unseen and 0 for train. And model gini gives us a rough idea that are the data distribution among train and unseen test is similar or not. and top importance variables tell which variables distribution is different in test unseen vs train.\nAccording we can remove such high importance variables from my final model(used for default prediction) to make our Model more robust.\nwe can do this analysis between test seen and train or test seen and test unseen as well.\n\nMy question is doing so is making your model more stable and reliable on unseen test. But does it also improving your final scores on unseen test. for eg (Hypothetically) without Adversarial my score on test seen vs test unseen can be - 0.799 and 0.798. after dropping some variables due to adversarial it can be  0.797 and 0.797. so My model has become more robust but its overall score on unseen test has decreased.   \nin above example I have shown 3rd decimal place difference but it can be 4th or 5th decimal place difference in actual so we never know does our score on Unseen test are improving or not.",
      "votes": null
    },
    {
      "id": "1846329",
      "postDate": "07/07/2022 03:00:44",
      "content": "<p>This is a really interesting question. I'm not sure if there is a definitive answer, but I'll share my thoughts.<br>\nIt seems to me that the adversarial validation is helpful in making the model more robust, but it may not necessarily lead to an improvement in the score on the unseen test data.<br>\nIt's possible that the score on the unseen test data would eventually improve as the model continues to learn from more data, but that would always be the case.<br>\nUltimately, I think it depends on the data and the model. For some data and some models, adversarial validation may help to improve the score on unseen test data. For other data and models, it may not have much impact.<br>\nI hope that helps.</p>",
      "rawMarkdown": "This is a really interesting question. I'm not sure if there is a definitive answer, but I'll share my thoughts.\nIt seems to me that the adversarial validation is helpful in making the model more robust, but it may not necessarily lead to an improvement in the score on the unseen test data.\nIt's possible that the score on the unseen test data would eventually improve as the model continues to learn from more data, but that would always be the case.\nUltimately, I think it depends on the data and the model. For some data and some models, adversarial validation may help to improve the score on unseen test data. For other data and models, it may not have much impact.\nI hope that helps.",
      "votes": null
    },
    {
      "id": "1847962",
      "postDate": "07/08/2022 09:24:05",
      "content": "<p>Try placing your features on 2d coordinates, where x-axis is adversarial score and y-axis is model importance.<br>\nYou will see visually what features help to the model and which features harm the model. If feature helps more than harms (upper-left corner) then feature is good for the model. if its low-right corner of the scatterplot, then it harms and does not help. and so on.<br>\nPlease notice, that you do not really know what are the scale of which dimension of the coordinates.<br>\nAnother point, adversarial validation shows which features are having potential to harm (it may not affect the final score).<br>\nThose and other issues comes from the fact, that data may have complex non-linear relationships and it may disturb the expected results.</p>",
      "rawMarkdown": "Try placing your features on 2d coordinates, where x-axis is adversarial score and y-axis is model importance.\nYou will see visually what features help to the model and which features harm the model. If feature helps more than harms (upper-left corner) then feature is good for the model. if its low-right corner of the scatterplot, then it harms and does not help. and so on.\nPlease notice, that you do not really know what are the scale of which dimension of the coordinates.\nAnother point, adversarial validation shows which features are having potential to harm (it may not affect the final score).\nThose and other issues comes from the fact, that data may have complex non-linear relationships and it may disturb the expected results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1846329,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/07/2022 03:00:44",
      "content": "<p>This is a really interesting question. I'm not sure if there is a definitive answer, but I'll share my thoughts.<br>\nIt seems to me that the adversarial validation is helpful in making the model more robust, but it may not necessarily lead to an improvement in the score on the unseen test data.<br>\nIt's possible that the score on the unseen test data would eventually improve as the model continues to learn from more data, but that would always be the case.<br>\nUltimately, I think it depends on the data and the model. For some data and some models, adversarial validation may help to improve the score on unseen test data. For other data and models, it may not have much impact.<br>\nI hope that helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1847962,
      "author_name": "pavelvod",
      "author_url": "",
      "post_date": "07/08/2022 09:24:05",
      "content": "<p>Try placing your features on 2d coordinates, where x-axis is adversarial score and y-axis is model importance.<br>\nYou will see visually what features help to the model and which features harm the model. If feature helps more than harms (upper-left corner) then feature is good for the model. if its low-right corner of the scatterplot, then it harms and does not help. and so on.<br>\nPlease notice, that you do not really know what are the scale of which dimension of the coordinates.<br>\nAnother point, adversarial validation shows which features are having potential to harm (it may not affect the final score).<br>\nThose and other issues comes from the fact, that data may have complex non-linear relationships and it may disturb the expected results.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1845005": "Adversarial Validation- we build a model with unseen data(test unseen) and data used for Model training(train), with label as 1 for test unseen and 0 for train. And model gini gives us a rough idea that are the data distribution among train and unseen test is similar or not. and top importance variables tell which variables distribution is different in test unseen vs train.\nAccording we can remove such high importance variables from my final model(used for default prediction) to make our Model more robust.\nwe can do this analysis between test seen and train or test seen and test unseen as well.\n\nMy question is doing so is making your model more stable and reliable on unseen test. But does it also improving your final scores on unseen test. for eg (Hypothetically) without Adversarial my score on test seen vs test unseen can be - 0.799 and 0.798. after dropping some variables due to adversarial it can be  0.797 and 0.797. so My model has become more robust but its overall score on unseen test has decreased.   \nin above example I have shown 3rd decimal place difference but it can be 4th or 5th decimal place difference in actual so we never know does our score on Unseen test are improving or not.",
    "1846329": "This is a really interesting question. I'm not sure if there is a definitive answer, but I'll share my thoughts.\nIt seems to me that the adversarial validation is helpful in making the model more robust, but it may not necessarily lead to an improvement in the score on the unseen test data.\nIt's possible that the score on the unseen test data would eventually improve as the model continues to learn from more data, but that would always be the case.\nUltimately, I think it depends on the data and the model. For some data and some models, adversarial validation may help to improve the score on unseen test data. For other data and models, it may not have much impact.\nI hope that helps.",
    "1847962": "Try placing your features on 2d coordinates, where x-axis is adversarial score and y-axis is model importance.\nYou will see visually what features help to the model and which features harm the model. If feature helps more than harms (upper-left corner) then feature is good for the model. if its low-right corner of the scatterplot, then it harms and does not help. and so on.\nPlease notice, that you do not really know what are the scale of which dimension of the coordinates.\nAnother point, adversarial validation shows which features are having potential to harm (it may not affect the final score).\nThose and other issues comes from the fact, that data may have complex non-linear relationships and it may disturb the expected results."
  },
  "source": "meta"
}