{
  "id": 84887,
  "title": "Why train/test distributions so different?",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/84887",
  "author_name": "",
  "post_date": "2019-03-20T10:08:24.655626Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Competition Host Tomas Vantuc wrote \"Signals are properly mixed to possess similar distributions across train/test/etc.\" but didn't explain what does \"properly\" mean and also admitted that he \"did no investigation to know [how sets are separated]\"\n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771\">source</a></p>\n\n<p>It's obvious the distributions of test and train data are very different (for example <a href=\"https://www.kaggle.com/its7171/train-vs-test-analisys\">here</a> ). I personally couldn't prepare any features that cannot distinguish the sets and detect the fault.</p>\n\n<p>The question I have is why? I can suspect that the devices used for the measurement aren't calibrated well, or the signals are very different depending on the localization of the measurement. It's hard to address these potential issues. Has anyone figured something out? </p>\n\n<p>I tried to filter the signal, leaving only consecutive million of frequencies (like 1000000-2000000, 2000000-3000000 etc.) The signals are very different in all the bandwidths. </p>",
  "messages": [
    {
      "id": "494848",
      "postDate": "03/20/2019 10:08:24",
      "content": "<p>Competition Host Tomas Vantuc wrote \"Signals are properly mixed to possess similar distributions across train/test/etc.\" but didn't explain what does \"properly\" mean and also admitted that he \"did no investigation to know [how sets are separated]\"\n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771\">source</a></p>\n\n<p>It's obvious the distributions of test and train data are very different (for example <a href=\"https://www.kaggle.com/its7171/train-vs-test-analisys\">here</a> ). I personally couldn't prepare any features that cannot distinguish the sets and detect the fault.</p>\n\n<p>The question I have is why? I can suspect that the devices used for the measurement aren't calibrated well, or the signals are very different depending on the localization of the measurement. It's hard to address these potential issues. Has anyone figured something out? </p>\n\n<p>I tried to filter the signal, leaving only consecutive million of frequencies (like 1000000-2000000, 2000000-3000000 etc.) The signals are very different in all the bandwidths. </p>",
      "rawMarkdown": "Competition Host Tomas Vantuc wrote \"Signals are properly mixed to possess similar distributions across train/test/etc.\" but didn't explain what does \"properly\" mean and also admitted that he \"did no investigation to know [how sets are separated]\"\n[source](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771)\n\n\nIt's obvious the distributions of test and train data are very different (for example [here](https://www.kaggle.com/its7171/train-vs-test-analisys) ). I personally couldn't prepare any features that cannot distinguish the sets and detect the fault.\n\n\nThe question I have is why? I can suspect that the devices used for the measurement aren't calibrated well, or the signals are very different depending on the localization of the measurement. It's hard to address these potential issues. Has anyone figured something out? \n\nI tried to filter the signal, leaving only consecutive million of frequencies (like 1000000-2000000, 2000000-3000000 etc.) The signals are very different in all the bandwidths.",
      "votes": null
    },
    {
      "id": "495885",
      "postDate": "03/21/2019 16:32:45",
      "content": "<p>This is what makes this challenge so difficult. I wonder how the people with more than 0.85 MCC did it !</p>",
      "rawMarkdown": "This is what makes this challenge so difficult. I wonder how the people with more than 0.85 MCC did it !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 495885,
      "author_name": "tarunpaparaju",
      "author_url": "",
      "post_date": "03/21/2019 16:32:45",
      "content": "<p>This is what makes this challenge so difficult. I wonder how the people with more than 0.85 MCC did it !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "494848": "Competition Host Tomas Vantuc wrote \"Signals are properly mixed to possess similar distributions across train/test/etc.\" but didn't explain what does \"properly\" mean and also admitted that he \"did no investigation to know [how sets are separated]\"\n[source](https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/75771)\n\n\nIt's obvious the distributions of test and train data are very different (for example [here](https://www.kaggle.com/its7171/train-vs-test-analisys) ). I personally couldn't prepare any features that cannot distinguish the sets and detect the fault.\n\n\nThe question I have is why? I can suspect that the devices used for the measurement aren't calibrated well, or the signals are very different depending on the localization of the measurement. It's hard to address these potential issues. Has anyone figured something out? \n\nI tried to filter the signal, leaving only consecutive million of frequencies (like 1000000-2000000, 2000000-3000000 etc.) The signals are very different in all the bandwidths.",
    "495885": "This is what makes this challenge so difficult. I wonder how the people with more than 0.85 MCC did it !"
  },
  "source": "meta"
}