{
  "id": 52936,
  "title": "Data Distribution in Train and Test",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52936",
  "author_name": "",
  "post_date": "2018-03-25T04:04:27.055560500Z",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>How to find out whether data distribution in train and test are similar? I used correlation matrix but it is not of great use.</p>",
  "messages": [
    {
      "id": "302950",
      "postDate": "03/25/2018 04:04:27",
      "content": "<p>How to find out whether data distribution in train and test are similar? I used correlation matrix but it is not of great use.</p>",
      "rawMarkdown": "How to find out whether data distribution in train and test are similar? I used correlation matrix but it is not of great use.",
      "votes": null
    },
    {
      "id": "302969",
      "postDate": "03/25/2018 05:58:00",
      "content": "<p>Distribution of what? Distribution of is_attributed you won't be able to know as we have no knowledge for that for the test data (it has to be predicted). Distribution of the hour of the day is pretty much similiar, note however that in test there are only a few hours out there and not a whole day (on top of my head it was something like 4,5,9,10,14,15 hours)</p>",
      "rawMarkdown": "Distribution of what? Distribution of is_attributed you won't be able to know as we have no knowledge for that for the test data (it has to be predicted). Distribution of the hour of the day is pretty much similiar, note however that in test there are only a few hours out there and not a whole day (on top of my head it was something like 4,5,9,10,14,15 hours)",
      "votes": null
    },
    {
      "id": "302975",
      "postDate": "03/25/2018 06:03:47",
      "content": "<p>I am referrring to similarity in train and test features data.</p>",
      "rawMarkdown": "I am referrring to similarity in train and test features data.",
      "votes": null
    },
    {
      "id": "302995",
      "postDate": "03/25/2018 07:35:32",
      "content": "<p>I would go with adversarial validation - along the lines of </p>\n\n<p><a href=\"http://fastml.com/adversarial-validation-part-one/\">http://fastml.com/adversarial-validation-part-one/</a></p>",
      "rawMarkdown": "I would go with adversarial validation - along the lines of \n\nhttp://fastml.com/adversarial-validation-part-one/",
      "votes": null
    },
    {
      "id": "303023",
      "postDate": "03/25/2018 09:26:21",
      "content": "<p>thank you.</p>",
      "rawMarkdown": "thank you.",
      "votes": null
    },
    {
      "id": "303035",
      "postDate": "03/25/2018 10:13:54",
      "content": "<p>Check this cool kernel on adversarial validation :</p>\n\n<p><a href=\"https://www.kaggle.com/ogrellier/adversarial-validation-and-lb-shakeup\">https://www.kaggle.com/ogrellier/adversarial-validation-and-lb-shakeup</a></p>",
      "rawMarkdown": "Check this cool kernel on adversarial validation :\n\n https://www.kaggle.com/ogrellier/adversarial-validation-and-lb-shakeup",
      "votes": null
    },
    {
      "id": "303040",
      "postDate": "03/25/2018 10:34:27",
      "content": "<p>There is one difference that was identified by several people, see for instance: <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates/notebook\">https://www.kaggle.com/cpmpml/ip-download-rates/notebook</a></p>",
      "rawMarkdown": "There is one difference that was identified by several people, see for instance: https://www.kaggle.com/cpmpml/ip-download-rates/notebook",
      "votes": null
    },
    {
      "id": "303113",
      "postDate": "03/25/2018 14:55:23",
      "content": "<p>sure ...thank you</p>",
      "rawMarkdown": "sure ...thank you",
      "votes": null
    },
    {
      "id": "303400",
      "postDate": "03/26/2018 05:30:22",
      "content": "<p>I, first of all, looked at the individual feature distributions.\nFor example, ip and click_time have very different distributions in train vs test set. See my <a href=\"https://www.kaggle.com/araksstepanyan/feature-distributions-10mm-row-train-set\">kernel of distributions</a> where I use a 10,000,000-row training set.</p>",
      "rawMarkdown": "I, first of all, looked at the individual feature distributions.\nFor example, ip and click_time have very different distributions in train vs test set. See my [kernel of distributions][1] where I use a 10,000,000-row training set.\n\n\n  [1]: https://www.kaggle.com/araksstepanyan/feature-distributions-10mm-row-train-set",
      "votes": null
    },
    {
      "id": "855976",
      "postDate": "05/21/2020 11:35:50",
      "content": "<p>I know 2 ways of checking this:</p>\n\n<ol>\n<li><p>Build a classifier that tries to predict whether a data point belongs to the training or test set. If it succeeds, the distributions are different ;). <a href=\"https://towardsdatascience.com/why-you-may-be-getting-low-test-accuracy-try-this-quick-way-of-comparing-the-distribution-of-the-9f06f5a72cfc\">Here</a>'s a tutorial.</p></li>\n<li><p>Do a statistical test like Kolmogorov-Smirnov. This needs to be done in a per column level. If you are working with column-dependant data (like work embeddings, for example), there are some workarounds - basically, you transform the data so that it is uni dimensional, losing some information in the process, of course. <a href=\"https://towardsdatascience.com/why-you-may-be-getting-low-test-accuracy-try-this-simpstatistical-tests-30585b7ee4fa\">Here</a>'s a tutorial for both the statistical test and dealing with word embeddings. </p></li>\n</ol>",
      "rawMarkdown": "I know 2 ways of checking this:\n\n1. Build a classifier that tries to predict whether a data point belongs to the training or test set. If it succeeds, the distributions are different ;). [Here](https://towardsdatascience.com/why-you-may-be-getting-low-test-accuracy-try-this-quick-way-of-comparing-the-distribution-of-the-9f06f5a72cfc)'s a tutorial.\n\n2. Do a statistical test like Kolmogorov-Smirnov. This needs to be done in a per column level. If you are working with column-dependant data (like work embeddings, for example), there are some workarounds - basically, you transform the data so that it is uni dimensional, losing some information in the process, of course. [Here](https://towardsdatascience.com/why-you-may-be-getting-low-test-accuracy-try-this-simpstatistical-tests-30585b7ee4fa)'s a tutorial for both the statistical test and dealing with word embeddings.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 302969,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/25/2018 05:58:00",
      "content": "<p>Distribution of what? Distribution of is_attributed you won't be able to know as we have no knowledge for that for the test data (it has to be predicted). Distribution of the hour of the day is pretty much similiar, note however that in test there are only a few hours out there and not a whole day (on top of my head it was something like 4,5,9,10,14,15 hours)</p>",
      "votes": null,
      "replies": [
        {
          "id": 302975,
          "author_name": "maheshak04",
          "author_url": "",
          "post_date": "03/25/2018 06:03:47",
          "content": "<p>I am referrring to similarity in train and test features data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 302995,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "03/25/2018 07:35:32",
      "content": "<p>I would go with adversarial validation - along the lines of </p>\n\n<p><a href=\"http://fastml.com/adversarial-validation-part-one/\">http://fastml.com/adversarial-validation-part-one/</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 303023,
          "author_name": "maheshak04",
          "author_url": "",
          "post_date": "03/25/2018 09:26:21",
          "content": "<p>thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 303035,
      "author_name": "actuben",
      "author_url": "",
      "post_date": "03/25/2018 10:13:54",
      "content": "<p>Check this cool kernel on adversarial validation :</p>\n\n<p><a href=\"https://www.kaggle.com/ogrellier/adversarial-validation-and-lb-shakeup\">https://www.kaggle.com/ogrellier/adversarial-validation-and-lb-shakeup</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 303113,
          "author_name": "maheshak04",
          "author_url": "",
          "post_date": "03/25/2018 14:55:23",
          "content": "<p>sure ...thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 303040,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/25/2018 10:34:27",
      "content": "<p>There is one difference that was identified by several people, see for instance: <a href=\"https://www.kaggle.com/cpmpml/ip-download-rates/notebook\">https://www.kaggle.com/cpmpml/ip-download-rates/notebook</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 303400,
      "author_name": "araksstepanyan",
      "author_url": "",
      "post_date": "03/26/2018 05:30:22",
      "content": "<p>I, first of all, looked at the individual feature distributions.\nFor example, ip and click_time have very different distributions in train vs test set. See my <a href=\"https://www.kaggle.com/araksstepanyan/feature-distributions-10mm-row-train-set\">kernel of distributions</a> where I use a 10,000,000-row training set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 855976,
      "author_name": "gmosse",
      "author_url": "",
      "post_date": "05/21/2020 11:35:50",
      "content": "<p>I know 2 ways of checking this:</p>\n\n<ol>\n<li><p>Build a classifier that tries to predict whether a data point belongs to the training or test set. If it succeeds, the distributions are different ;). <a href=\"https://towardsdatascience.com/why-you-may-be-getting-low-test-accuracy-try-this-quick-way-of-comparing-the-distribution-of-the-9f06f5a72cfc\">Here</a>'s a tutorial.</p></li>\n<li><p>Do a statistical test like Kolmogorov-Smirnov. This needs to be done in a per column level. If you are working with column-dependant data (like work embeddings, for example), there are some workarounds - basically, you transform the data so that it is uni dimensional, losing some information in the process, of course. <a href=\"https://towardsdatascience.com/why-you-may-be-getting-low-test-accuracy-try-this-simpstatistical-tests-30585b7ee4fa\">Here</a>'s a tutorial for both the statistical test and dealing with word embeddings. </p></li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "302950": "How to find out whether data distribution in train and test are similar? I used correlation matrix but it is not of great use.",
    "302969": "Distribution of what? Distribution of is_attributed you won't be able to know as we have no knowledge for that for the test data (it has to be predicted). Distribution of the hour of the day is pretty much similiar, note however that in test there are only a few hours out there and not a whole day (on top of my head it was something like 4,5,9,10,14,15 hours)",
    "302975": "I am referrring to similarity in train and test features data.",
    "302995": "I would go with adversarial validation - along the lines of \n\nhttp://fastml.com/adversarial-validation-part-one/",
    "303023": "thank you.",
    "303035": "Check this cool kernel on adversarial validation :\n\n https://www.kaggle.com/ogrellier/adversarial-validation-and-lb-shakeup",
    "303040": "There is one difference that was identified by several people, see for instance: https://www.kaggle.com/cpmpml/ip-download-rates/notebook",
    "303113": "sure ...thank you",
    "303400": "I, first of all, looked at the individual feature distributions.\nFor example, ip and click_time have very different distributions in train vs test set. See my [kernel of distributions][1] where I use a 10,000,000-row training set.\n\n\n  [1]: https://www.kaggle.com/araksstepanyan/feature-distributions-10mm-row-train-set",
    "855976": "I know 2 ways of checking this:\n\n1. Build a classifier that tries to predict whether a data point belongs to the training or test set. If it succeeds, the distributions are different ;). [Here](https://towardsdatascience.com/why-you-may-be-getting-low-test-accuracy-try-this-quick-way-of-comparing-the-distribution-of-the-9f06f5a72cfc)'s a tutorial.\n\n2. Do a statistical test like Kolmogorov-Smirnov. This needs to be done in a per column level. If you are working with column-dependant data (like work embeddings, for example), there are some workarounds - basically, you transform the data so that it is uni dimensional, losing some information in the process, of course. [Here](https://towardsdatascience.com/why-you-may-be-getting-low-test-accuracy-try-this-simpstatistical-tests-30585b7ee4fa)'s a tutorial for both the statistical test and dealing with word embeddings."
  },
  "source": "meta"
}