{
  "id": 198366,
  "title": "Training and test data distributions",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/198366",
  "author_name": "",
  "post_date": "2020-11-20T21:04:26.513598Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have noticed that on all my submissions there was quite a big mismatch between validation and test scores, although doing kFold cross validation.</p>\n<p>I assume that this is due to different distributions of training and test data and I also did a quick experiment to check this using T-SNE, i.e. after extracting hundreds of features, projecting them to a two dimensional space and then visualizing the features.</p>\n<p>The attached scatter plot shows the 2D embedding of train data (blue) and test data (red), which show at least differences in some regions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6121501%2F804f687cbef9526a47dfadb0c16e2969%2Findex.png?generation=1605906259134873&amp;alt=media\" alt=\"\"></p>\n<p>Having said that, does anyone have recommendations how to deal with the differences of the data distributions?</p>",
  "messages": [
    {
      "id": "1085342",
      "postDate": "11/20/2020 21:04:26",
      "content": "<p>I have noticed that on all my submissions there was quite a big mismatch between validation and test scores, although doing kFold cross validation.</p>\n<p>I assume that this is due to different distributions of training and test data and I also did a quick experiment to check this using T-SNE, i.e. after extracting hundreds of features, projecting them to a two dimensional space and then visualizing the features.</p>\n<p>The attached scatter plot shows the 2D embedding of train data (blue) and test data (red), which show at least differences in some regions.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6121501%2F804f687cbef9526a47dfadb0c16e2969%2Findex.png?generation=1605906259134873&amp;alt=media\" alt=\"\"></p>\n<p>Having said that, does anyone have recommendations how to deal with the differences of the data distributions?</p>",
      "rawMarkdown": "I have noticed that on all my submissions there was quite a big mismatch between validation and test scores, although doing kFold cross validation.\n\nI assume that this is due to different distributions of training and test data and I also did a quick experiment to check this using T-SNE, i.e. after extracting hundreds of features, projecting them to a two dimensional space and then visualizing the features.\n\nThe attached scatter plot shows the 2D embedding of train data (blue) and test data (red), which show at least differences in some regions.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6121501%2F804f687cbef9526a47dfadb0c16e2969%2Findex.png?generation=1605906259134873&alt=media)\n\nHaving said that, does anyone have recommendations how to deal with the differences of the data distributions?",
      "votes": null
    },
    {
      "id": "1085534",
      "postDate": "11/21/2020 00:38:46",
      "content": "<p>There have been a couple discussions about training and test differences. E.g., <a href=\"https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194300\" target=\"_blank\">Adversarial Validation - There is something fishy about the test data</a></p>",
      "rawMarkdown": "There have been a couple discussions about training and test differences. E.g., [Adversarial Validation - There is something fishy about the test data](https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194300)",
      "votes": null
    },
    {
      "id": "1085639",
      "postDate": "11/21/2020 03:37:57",
      "content": "<p>Nice work! Just a suggestion but the plot might benefit by first using PCA and increasing the perplexity of the TSNE </p>",
      "rawMarkdown": "Nice work! Just a suggestion but the plot might benefit by first using PCA and increasing the perplexity of the TSNE",
      "votes": null
    },
    {
      "id": "1085957",
      "postDate": "11/21/2020 09:52:03",
      "content": "<p>Just found a page on TowardsDataScience about that: <a href=\"https://towardsdatascience.com/visualising-high-dimensional-datasets-using-pca-and-t-sne-in-python-8ef87e7915b\" target=\"_blank\">https://towardsdatascience.com/visualising-high-dimensional-datasets-using-pca-and-t-sne-in-python-8ef87e7915b</a> <br>\nThanks for your suggestion, Adam!</p>",
      "rawMarkdown": "Just found a page on TowardsDataScience about that: https://towardsdatascience.com/visualising-high-dimensional-datasets-using-pca-and-t-sne-in-python-8ef87e7915b \nThanks for your suggestion, Adam!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1085534,
      "author_name": "bengutierrez",
      "author_url": "",
      "post_date": "11/21/2020 00:38:46",
      "content": "<p>There have been a couple discussions about training and test differences. E.g., <a href=\"https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194300\" target=\"_blank\">Adversarial Validation - There is something fishy about the test data</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1085639,
      "author_name": "ajcostarino",
      "author_url": "",
      "post_date": "11/21/2020 03:37:57",
      "content": "<p>Nice work! Just a suggestion but the plot might benefit by first using PCA and increasing the perplexity of the TSNE </p>",
      "votes": null,
      "replies": [
        {
          "id": 1085957,
          "author_name": "rhinopithecus",
          "author_url": "",
          "post_date": "11/21/2020 09:52:03",
          "content": "<p>Just found a page on TowardsDataScience about that: <a href=\"https://towardsdatascience.com/visualising-high-dimensional-datasets-using-pca-and-t-sne-in-python-8ef87e7915b\" target=\"_blank\">https://towardsdatascience.com/visualising-high-dimensional-datasets-using-pca-and-t-sne-in-python-8ef87e7915b</a> <br>\nThanks for your suggestion, Adam!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1085342": "I have noticed that on all my submissions there was quite a big mismatch between validation and test scores, although doing kFold cross validation.\n\nI assume that this is due to different distributions of training and test data and I also did a quick experiment to check this using T-SNE, i.e. after extracting hundreds of features, projecting them to a two dimensional space and then visualizing the features.\n\nThe attached scatter plot shows the 2D embedding of train data (blue) and test data (red), which show at least differences in some regions.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6121501%2F804f687cbef9526a47dfadb0c16e2969%2Findex.png?generation=1605906259134873&alt=media)\n\nHaving said that, does anyone have recommendations how to deal with the differences of the data distributions?",
    "1085534": "There have been a couple discussions about training and test differences. E.g., [Adversarial Validation - There is something fishy about the test data](https://www.kaggle.com/c/predict-volcanic-eruptions-ingv-oe/discussion/194300)",
    "1085639": "Nice work! Just a suggestion but the plot might benefit by first using PCA and increasing the perplexity of the TSNE",
    "1085957": "Just found a page on TowardsDataScience about that: https://towardsdatascience.com/visualising-high-dimensional-datasets-using-pca-and-t-sne-in-python-8ef87e7915b \nThanks for your suggestion, Adam!"
  },
  "source": "meta"
}