{
  "id": 346251,
  "title": "Adversarial validation on (Aggregated Data)",
  "url": "/competitions/amex-default-prediction/discussion/346251",
  "author_name": "",
  "post_date": "2022-08-18T15:42:17.543796700Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I performed the adversarial validation on the data generated in this notebook.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks</a><br>\nI used a simple cat boost and performed it in two ways:</p>\n<ol>\n<li>Train and Public Test: <strong>AUC:  0.9999</strong> </li>\n<li>Train and Private Test:  <strong>AUC: 0.999999</strong></li>\n</ol>\n<p><strong>The most distinguishing features were R_1 and B_29</strong> out of all the features.</p>\n<p>**<br>\nThis suggests that both public test and private test data are different, however, still there is a good correlation between public test and train data.**</p>\n<p>for the 2nd set:<br>\nI removed those features and performed it again and still found the AUC of 0.99999. I think the reason for the shakeup (if it happened) can be only two features just because they are significantly different not the other features as other notebooks are trying to tell.<br>\n<strong>That's my hypothesis, this forum is open to discussion and I will be glad to hear others' input.</strong>😄</p>",
  "messages": [
    {
      "id": "1904924",
      "postDate": "08/18/2022 15:42:17",
      "content": "<p>I performed the adversarial validation on the data generated in this notebook.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks</a><br>\nI used a simple cat boost and performed it in two ways:</p>\n<ol>\n<li>Train and Public Test: <strong>AUC:  0.9999</strong> </li>\n<li>Train and Private Test:  <strong>AUC: 0.999999</strong></li>\n</ol>\n<p><strong>The most distinguishing features were R_1 and B_29</strong> out of all the features.</p>\n<p>**<br>\nThis suggests that both public test and private test data are different, however, still there is a good correlation between public test and train data.**</p>\n<p>for the 2nd set:<br>\nI removed those features and performed it again and still found the AUC of 0.99999. I think the reason for the shakeup (if it happened) can be only two features just because they are significantly different not the other features as other notebooks are trying to tell.<br>\n<strong>That's my hypothesis, this forum is open to discussion and I will be glad to hear others' input.</strong>😄</p>",
      "rawMarkdown": "I performed the adversarial validation on the data generated in this notebook.\n[https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks](url)\nI used a simple cat boost and performed it in two ways:\n\n1.  Train and Public Test: **AUC:  0.9999** \n2.  Train and Private Test:  **AUC: 0.999999**\n\n**The most distinguishing features were R_1 and B_29** out of all the features.\n\n\n\n\n**\nThis suggests that both public test and private test data are different, however, still there is a good correlation between public test and train data.**\n\nfor the 2nd set:\nI removed those features and performed it again and still found the AUC of 0.99999. I think the reason for the shakeup (if it happened) can be only two features just because they are significantly different not the other features as other notebooks are trying to tell.\n**That's my hypothesis, this forum is open to discussion and I will be glad to hear others' input.**😄",
      "votes": null
    },
    {
      "id": "1906147",
      "postDate": "08/19/2022 16:11:32",
      "content": "<p>Do I understand it correctly, you merged the training and test data sets and labeled them 0 and 1. Then you trained a classifier and got AUC of 0.999.</p>\n<p>As I understand it, this means that your dataset contains features whose distribution in the training dataset is very different from the distribution in the test dataset. If the features of the two datasets were similarly distributed, you would get an AUC around 0.5.</p>\n<p>Or have I misunderstood something? </p>",
      "rawMarkdown": "Do I understand it correctly, you merged the training and test data sets and labeled them 0 and 1. Then you trained a classifier and got AUC of 0.999.\n\nAs I understand it, this means that your dataset contains features whose distribution in the training dataset is very different from the distribution in the test dataset. If the features of the two datasets were similarly distributed, you would get an AUC around 0.5.\n\nOr have I misunderstood something?",
      "votes": null
    },
    {
      "id": "1906155",
      "postDate": "08/19/2022 16:23:06",
      "content": "<p>yes exactly, i have taken the help from this notebook<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public</a><br>\nas a baseline and perform these experiments on the above-mentioned data.</p>\n<p>Of all the features the most differently distributed features were min , std, which if i am not wrong every data is going to use. still, CV/LB is correlating.</p>",
      "rawMarkdown": "yes exactly, i have taken the help from this notebook\n[https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public](url)\nas a baseline and perform these experiments on the above-mentioned data.\n\nOf all the features the most differently distributed features were min , std, which if i am not wrong every data is going to use. still, CV/LB is correlating.",
      "votes": null
    },
    {
      "id": "1908621",
      "postDate": "08/21/2022 20:32:06",
      "content": "<p>Thanks for the clarification and for your interesting post! </p>",
      "rawMarkdown": "Thanks for the clarification and for your interesting post!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1906147,
      "author_name": "viktorfairuschin",
      "author_url": "",
      "post_date": "08/19/2022 16:11:32",
      "content": "<p>Do I understand it correctly, you merged the training and test data sets and labeled them 0 and 1. Then you trained a classifier and got AUC of 0.999.</p>\n<p>As I understand it, this means that your dataset contains features whose distribution in the training dataset is very different from the distribution in the test dataset. If the features of the two datasets were similarly distributed, you would get an AUC around 0.5.</p>\n<p>Or have I misunderstood something? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1906155,
          "author_name": "chaudharypriyanshu",
          "author_url": "",
          "post_date": "08/19/2022 16:23:06",
          "content": "<p>yes exactly, i have taken the help from this notebook<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public</a><br>\nas a baseline and perform these experiments on the above-mentioned data.</p>\n<p>Of all the features the most differently distributed features were min , std, which if i am not wrong every data is going to use. still, CV/LB is correlating.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1908621,
          "author_name": "viktorfairuschin",
          "author_url": "",
          "post_date": "08/21/2022 20:32:06",
          "content": "<p>Thanks for the clarification and for your interesting post! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1904924": "I performed the adversarial validation on the data generated in this notebook.\n[https://www.kaggle.com/code/roberthatch/amex-feature-engg-gpu-or-cpu-process-in-chunks](url)\nI used a simple cat boost and performed it in two ways:\n\n1.  Train and Public Test: **AUC:  0.9999** \n2.  Train and Private Test:  **AUC: 0.999999**\n\n**The most distinguishing features were R_1 and B_29** out of all the features.\n\n\n\n\n**\nThis suggests that both public test and private test data are different, however, still there is a good correlation between public test and train data.**\n\nfor the 2nd set:\nI removed those features and performed it again and still found the AUC of 0.99999. I think the reason for the shakeup (if it happened) can be only two features just because they are significantly different not the other features as other notebooks are trying to tell.\n**That's my hypothesis, this forum is open to discussion and I will be glad to hear others' input.**😄",
    "1906147": "Do I understand it correctly, you merged the training and test data sets and labeled them 0 and 1. Then you trained a classifier and got AUC of 0.999.\n\nAs I understand it, this means that your dataset contains features whose distribution in the training dataset is very different from the distribution in the test dataset. If the features of the two datasets were similarly distributed, you would get an AUC around 0.5.\n\nOr have I misunderstood something?",
    "1906155": "yes exactly, i have taken the help from this notebook\n[https://www.kaggle.com/code/zakopur0/adversarial-validation-private-vs-public](url)\nas a baseline and perform these experiments on the above-mentioned data.\n\nOf all the features the most differently distributed features were min , std, which if i am not wrong every data is going to use. still, CV/LB is correlating.",
    "1908621": "Thanks for the clarification and for your interesting post!"
  },
  "source": "meta"
}