{
  "id": 83523,
  "title": "How does adversarial validation remove features and reduce AUC",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/83523",
  "author_name": "",
  "post_date": "2019-03-11T00:20:11.689744Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I do adversarial validation follow <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317\">paul</a>, but the auc of model is quite high, about 0.97-0.99 after 15 epochs. The MCC can get 0.77-0.85, Unfortunately， the LB is extremely unsatisfactory（0.4-0.5）.The model is overfitting, i try to reduce the scale of the lstm model, but it does not work. \nThe idea of adversarial validation, i understand, is to select the validation samples from raw training samples that most resemble the test set. How does adversarial validation remove features and reduce AUC?</p>",
  "messages": [
    {
      "id": "487473",
      "postDate": "03/11/2019 00:20:11",
      "content": "<p>I do adversarial validation follow <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317\">paul</a>, but the auc of model is quite high, about 0.97-0.99 after 15 epochs. The MCC can get 0.77-0.85, Unfortunately， the LB is extremely unsatisfactory（0.4-0.5）.The model is overfitting, i try to reduce the scale of the lstm model, but it does not work. \nThe idea of adversarial validation, i understand, is to select the validation samples from raw training samples that most resemble the test set. How does adversarial validation remove features and reduce AUC?</p>",
      "rawMarkdown": "I do adversarial validation follow [paul](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317), but the auc of model is quite high, about 0.97-0.99 after 15 epochs. The MCC can get 0.77-0.85, Unfortunately， the LB is extremely unsatisfactory（0.4-0.5）.The model is overfitting, i try to reduce the scale of the lstm model, but it does not work. \nThe idea of adversarial validation, i understand, is to select the validation samples from raw training samples that most resemble the test set. How does adversarial validation remove features and reduce AUC?",
      "votes": null
    },
    {
      "id": "488577",
      "postDate": "03/12/2019 17:43:16",
      "content": "<p>From the adversarial Validation, look for important features and try to remove those features, before training the model. You can also try to analyze each feature in train and test set and look for the variation in distribution, if variation is quite significant, then you can also drop that features. </p>",
      "rawMarkdown": "From the adversarial Validation, look for important features and try to remove those features, before training the model. You can also try to analyze each feature in train and test set and look for the variation in distribution, if variation is quite significant, then you can also drop that features.",
      "votes": null
    },
    {
      "id": "488806",
      "postDate": "03/13/2019 02:33:37",
      "content": "<p>Thank you for your insight, it is very impressive. would you please tell me why we remove those important features? such as <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166\">Max Halford</a> do in his work, the signal_entropy and mean are the most important features. </p>",
      "rawMarkdown": "Thank you for your insight, it is very impressive. would you please tell me why we remove those important features? such as [Max Halford](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166) do in his work, the signal_entropy and mean are the most important features.",
      "votes": null
    },
    {
      "id": "488814",
      "postDate": "03/13/2019 03:05:29",
      "content": "<p>There is some misunderstanding, <strong>Max</strong> kernel is building a binary classifier to predict the fault. In Adversarial validation, we will predict whether the data point is from train or test set.  We combine train and test set and add new target variable, for example: \"IS_TEST\", which will be 0 for all the training set and 1 for test set. for this problem, we will find the important set of features, so important features in adversarial validation will be responsible for higher auc score, which imply, that train and test are from different distribution, since to make train and test same distribution, we will remove important features of adversarial modeling which are boosting the auc score. </p>\n\n<p>for quick test:\njust concatenate the meta data of train and test and do adversarial validation, you will get high auc score, because \"id_measurement\" would be a significant feature in adversarial validation, remove the \"id_measurement\" and repeat the same experiment using phase only. Why \"id_measurement\" is causing high auc, because there is no intersection of values between train and test \"id_measurement\".</p>",
      "rawMarkdown": "There is some misunderstanding, **Max** kernel is building a binary classifier to predict the fault. In Adversarial validation, we will predict whether the data point is from train or test set.  We combine train and test set and add new target variable, for example: \"IS_TEST\", which will be 0 for all the training set and 1 for test set. for this problem, we will find the important set of features, so important features in adversarial validation will be responsible for higher auc score, which imply, that train and test are from different distribution, since to make train and test same distribution, we will remove important features of adversarial modeling which are boosting the auc score. \n\nfor quick test:\njust concatenate the meta data of train and test and do adversarial validation, you will get high auc score, because \"id_measurement\" would be a significant feature in adversarial validation, remove the \"id_measurement\" and repeat the same experiment using phase only. Why \"id_measurement\" is causing high auc, because there is no intersection of values between train and test \"id_measurement\".",
      "votes": null
    },
    {
      "id": "488913",
      "postDate": "03/13/2019 07:16:52",
      "content": "<p>Thank you so so so so much! you teach me what adversarial validation actually do! I hope you will win the gold medal. thank you!</p>",
      "rawMarkdown": "Thank you so so so so much! you teach me what adversarial validation actually do! I hope you will win the gold medal. thank you!",
      "votes": null
    },
    {
      "id": "489323",
      "postDate": "03/13/2019 19:18:13",
      "content": "<p>Hehe. Glad that help. Thank you :)</p>",
      "rawMarkdown": "Hehe. Glad that help. Thank you :)",
      "votes": null
    },
    {
      "id": "489349",
      "postDate": "03/13/2019 20:00:14",
      "content": "<p>Thanks Harshit for this great explanation on adversarial modeling.</p>",
      "rawMarkdown": "Thanks Harshit for this great explanation on adversarial modeling.",
      "votes": null
    },
    {
      "id": "493173",
      "postDate": "03/18/2019 11:25:42",
      "content": "<p><a href=\"/junyun1002\">@junyun1002</a> Have a look at this kernel in a recently ended competition by a grandmaster <a href=\"https://www.kaggle.com/tunguz/elo-adversarial-validation\">https://www.kaggle.com/tunguz/elo-adversarial-validation</a> in this competition look at this discussion topic <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317</a></p>",
      "rawMarkdown": "junyun1002 Have a look at this kernel in a recently ended competition by a grandmaster https://www.kaggle.com/tunguz/elo-adversarial-validation in this competition look at this discussion topic https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 488577,
      "author_name": "harshit92",
      "author_url": "",
      "post_date": "03/12/2019 17:43:16",
      "content": "<p>From the adversarial Validation, look for important features and try to remove those features, before training the model. You can also try to analyze each feature in train and test set and look for the variation in distribution, if variation is quite significant, then you can also drop that features. </p>",
      "votes": null,
      "replies": [
        {
          "id": 488806,
          "author_name": "junyun1002",
          "author_url": "",
          "post_date": "03/13/2019 02:33:37",
          "content": "<p>Thank you for your insight, it is very impressive. would you please tell me why we remove those important features? such as <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166\">Max Halford</a> do in his work, the signal_entropy and mean are the most important features. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 488814,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "03/13/2019 03:05:29",
          "content": "<p>There is some misunderstanding, <strong>Max</strong> kernel is building a binary classifier to predict the fault. In Adversarial validation, we will predict whether the data point is from train or test set.  We combine train and test set and add new target variable, for example: \"IS_TEST\", which will be 0 for all the training set and 1 for test set. for this problem, we will find the important set of features, so important features in adversarial validation will be responsible for higher auc score, which imply, that train and test are from different distribution, since to make train and test same distribution, we will remove important features of adversarial modeling which are boosting the auc score. </p>\n\n<p>for quick test:\njust concatenate the meta data of train and test and do adversarial validation, you will get high auc score, because \"id_measurement\" would be a significant feature in adversarial validation, remove the \"id_measurement\" and repeat the same experiment using phase only. Why \"id_measurement\" is causing high auc, because there is no intersection of values between train and test \"id_measurement\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 488913,
          "author_name": "junyun1002",
          "author_url": "",
          "post_date": "03/13/2019 07:16:52",
          "content": "<p>Thank you so so so so much! you teach me what adversarial validation actually do! I hope you will win the gold medal. thank you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 489323,
          "author_name": "harshit92",
          "author_url": "",
          "post_date": "03/13/2019 19:18:13",
          "content": "<p>Hehe. Glad that help. Thank you :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 489349,
          "author_name": "arihantkumarjain",
          "author_url": "",
          "post_date": "03/13/2019 20:00:14",
          "content": "<p>Thanks Harshit for this great explanation on adversarial modeling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 493173,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "03/18/2019 11:25:42",
      "content": "<p><a href=\"/junyun1002\">@junyun1002</a> Have a look at this kernel in a recently ended competition by a grandmaster <a href=\"https://www.kaggle.com/tunguz/elo-adversarial-validation\">https://www.kaggle.com/tunguz/elo-adversarial-validation</a> in this competition look at this discussion topic <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "487473": "I do adversarial validation follow [paul](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317), but the auc of model is quite high, about 0.97-0.99 after 15 epochs. The MCC can get 0.77-0.85, Unfortunately， the LB is extremely unsatisfactory（0.4-0.5）.The model is overfitting, i try to reduce the scale of the lstm model, but it does not work. \nThe idea of adversarial validation, i understand, is to select the validation samples from raw training samples that most resemble the test set. How does adversarial validation remove features and reduce AUC?",
    "488577": "From the adversarial Validation, look for important features and try to remove those features, before training the model. You can also try to analyze each feature in train and test set and look for the variation in distribution, if variation is quite significant, then you can also drop that features.",
    "488806": "Thank you for your insight, it is very impressive. would you please tell me why we remove those important features? such as [Max Halford](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/80166) do in his work, the signal_entropy and mean are the most important features.",
    "488814": "There is some misunderstanding, **Max** kernel is building a binary classifier to predict the fault. In Adversarial validation, we will predict whether the data point is from train or test set.  We combine train and test set and add new target variable, for example: \"IS_TEST\", which will be 0 for all the training set and 1 for test set. for this problem, we will find the important set of features, so important features in adversarial validation will be responsible for higher auc score, which imply, that train and test are from different distribution, since to make train and test same distribution, we will remove important features of adversarial modeling which are boosting the auc score. \n\nfor quick test:\njust concatenate the meta data of train and test and do adversarial validation, you will get high auc score, because \"id_measurement\" would be a significant feature in adversarial validation, remove the \"id_measurement\" and repeat the same experiment using phase only. Why \"id_measurement\" is causing high auc, because there is no intersection of values between train and test \"id_measurement\".",
    "488913": "Thank you so so so so much! you teach me what adversarial validation actually do! I hope you will win the gold medal. thank you!",
    "489323": "Hehe. Glad that help. Thank you :)",
    "489349": "Thanks Harshit for this great explanation on adversarial modeling.",
    "493173": "junyun1002 Have a look at this kernel in a recently ended competition by a grandmaster https://www.kaggle.com/tunguz/elo-adversarial-validation in this competition look at this discussion topic https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317"
  },
  "source": "meta"
}