{
  "id": 357542,
  "title": "Adversarial Validation: When statistics won't capture the whole picture.",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/357542",
  "author_name": "Jose Cáliz",
  "post_date": "2022-10-04T17:06:26.352000",
  "votes": 13,
  "comment_count": 0,
  "views": 0,
  "content": "<p>TL;DR train and test share the same distribution.</p>\n<p>As a safety measure, I did some adversarial validation using KS-statistic. I was expecting a low statistics with a low p-value considering the huge dataset, so you can imagine my surprise when I saw the following:</p>\n<table>\n<thead>\n<tr>\n<th>variable</th>\n<th>ks_statistic</th>\n<th>p-value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>p4_pos_z</td>\n<td>0.291371</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p3_pos_z</td>\n<td>0.289708</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p1_pos_z</td>\n<td>0.287073</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p5_pos_z</td>\n<td>0.286655</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p0_pos_z</td>\n<td>0.283748</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p2_pos_z</td>\n<td>0.283301</td>\n<td>0.00</td>\n</tr>\n</tbody>\n</table>\n<p>KS suggests that all <code>z</code> features don't follow the same distribution. Further visual analyzes confirm (see plot below).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3145414%2Fd410707dd10cc81a2e0c05240b6f14f4%2FScreen%20Shot%202022-10-04%20at%209.58.26%20AM.png?generation=1664895685645853&amp;alt=media\" alt=\"\"></p>\n<p>Here is the fun part, if you strip precision the KS-statistic drops. And if we see the entire cumulative distribution, do they \"really\" look different to you?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3145414%2Fbce3972377fcfe1c8e7f936be0af463f%2FScreen%20Shot%202022-10-04%20at%209.58.52%20AM.png?generation=1664898006875489&amp;alt=media\" alt=\"\"></p>\n<p>For me, they are the same. I'm open to discussion as always.</p>\n<p><a href=\"https://www.kaggle.com/code/jcaliz/tps-oct22-animation-eda-keras-baseline#Adversal-Validation\" target=\"_blank\">Code &amp; orginal plots here</a></p>",
  "messages": [
    {
      "id": 1971555,
      "postDate": "2022-10-04T17:06:26.353Z",
      "content": "<p>TL;DR train and test share the same distribution.</p>\n<p>As a safety measure, I did some adversarial validation using KS-statistic. I was expecting a low statistics with a low p-value considering the huge dataset, so you can imagine my surprise when I saw the following:</p>\n<table>\n<thead>\n<tr>\n<th>variable</th>\n<th>ks_statistic</th>\n<th>p-value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>p4_pos_z</td>\n<td>0.291371</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p3_pos_z</td>\n<td>0.289708</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p1_pos_z</td>\n<td>0.287073</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p5_pos_z</td>\n<td>0.286655</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p0_pos_z</td>\n<td>0.283748</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>p2_pos_z</td>\n<td>0.283301</td>\n<td>0.00</td>\n</tr>\n</tbody>\n</table>\n<p>KS suggests that all <code>z</code> features don't follow the same distribution. Further visual analyzes confirm (see plot below).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3145414%2Fd410707dd10cc81a2e0c05240b6f14f4%2FScreen%20Shot%202022-10-04%20at%209.58.26%20AM.png?generation=1664895685645853&amp;alt=media\" alt=\"\"></p>\n<p>Here is the fun part, if you strip precision the KS-statistic drops. And if we see the entire cumulative distribution, do they \"really\" look different to you?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3145414%2Fbce3972377fcfe1c8e7f936be0af463f%2FScreen%20Shot%202022-10-04%20at%209.58.52%20AM.png?generation=1664898006875489&amp;alt=media\" alt=\"\"></p>\n<p>For me, they are the same. I'm open to discussion as always.</p>\n<p><a href=\"https://www.kaggle.com/code/jcaliz/tps-oct22-animation-eda-keras-baseline#Adversal-Validation\" target=\"_blank\">Code &amp; orginal plots here</a></p>",
      "rawMarkdown": "TL;DR train and test share the same distribution.\n\nAs a safety measure, I did some adversarial validation using KS-statistic. I was expecting a low statistics with a low p-value considering the huge dataset, so you can imagine my surprise when I saw the following:\n\n| variable | ks_statistic | p-value |\n|  ------- | ------------ | ------- |\n| p4_pos_z | \t0.291371 \t|0.00\n| p3_pos_z |  0.289708 \t| 0.00\n| p1_pos_z  |\t0.287073 \t|0.00\n| p5_pos_z  | 0.286655 \t|0.00\n| p0_pos_z  | 0.283748 \t|0.00\n| p2_pos_z  | 0.283301 \t|0.00\n\nKS suggests that all `z` features don't follow the same distribution. Further visual analyzes confirm (see plot below).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3145414%2Fd410707dd10cc81a2e0c05240b6f14f4%2FScreen%20Shot%202022-10-04%20at%209.58.26%20AM.png?generation=1664895685645853&alt=media)\n\nHere is the fun part, if you strip precision the KS-statistic drops. And if we see the entire cumulative distribution, do they \"really\" look different to you?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3145414%2Fbce3972377fcfe1c8e7f936be0af463f%2FScreen%20Shot%202022-10-04%20at%209.58.52%20AM.png?generation=1664898006875489&alt=media)\n\nFor me, they are the same. I'm open to discussion as always.\n\n[Code & orginal plots here](https://www.kaggle.com/code/jcaliz/tps-oct22-animation-eda-keras-baseline#Adversal-Validation)",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1971555": "TL;DR train and test share the same distribution.\n\nAs a safety measure, I did some adversarial validation using KS-statistic. I was expecting a low statistics with a low p-value considering the huge dataset, so you can imagine my surprise when I saw the following:\n\n| variable | ks_statistic | p-value |\n|  ------- | ------------ | ------- |\n| p4_pos_z | \t0.291371 \t|0.00\n| p3_pos_z |  0.289708 \t| 0.00\n| p1_pos_z  |\t0.287073 \t|0.00\n| p5_pos_z  | 0.286655 \t|0.00\n| p0_pos_z  | 0.283748 \t|0.00\n| p2_pos_z  | 0.283301 \t|0.00\n\nKS suggests that all `z` features don't follow the same distribution. Further visual analyzes confirm (see plot below).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3145414%2Fd410707dd10cc81a2e0c05240b6f14f4%2FScreen%20Shot%202022-10-04%20at%209.58.26%20AM.png?generation=1664895685645853&alt=media)\n\nHere is the fun part, if you strip precision the KS-statistic drops. And if we see the entire cumulative distribution, do they \"really\" look different to you?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3145414%2Fbce3972377fcfe1c8e7f936be0af463f%2FScreen%20Shot%202022-10-04%20at%209.58.52%20AM.png?generation=1664898006875489&alt=media)\n\nFor me, they are the same. I'm open to discussion as always.\n\n[Code & orginal plots here](https://www.kaggle.com/code/jcaliz/tps-oct22-animation-eda-keras-baseline#Adversal-Validation)"
  }
}