{
  "id": 79471,
  "title": "Test set distribution significantly different from train set",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/79471",
  "author_name": "",
  "post_date": "2019-02-04T13:17:24.678559Z",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi guys! An analysis of the acoustic signal in the train and the test sets may lead you to the conlusion that their distributions are completely different across the majority of features used in the publicly available kernels. Even if you first separately standardize each 150 000 chunk of data, then many features are still distributed quite differently. This means that traditional supervised learning technics may fail in this competition because they are based on the idea that data-generating procedure is the same for train and test. I think because of these differences between the train and test sets, this competition is more about predicting time_to_failure for the particular given test set rather than about predicting time_to_failure in general [which I suppose was an initial goal of this competition].</p>",
  "messages": [
    {
      "id": "465982",
      "postDate": "02/04/2019 13:17:24",
      "content": "<p>Hi guys! An analysis of the acoustic signal in the train and the test sets may lead you to the conlusion that their distributions are completely different across the majority of features used in the publicly available kernels. Even if you first separately standardize each 150 000 chunk of data, then many features are still distributed quite differently. This means that traditional supervised learning technics may fail in this competition because they are based on the idea that data-generating procedure is the same for train and test. I think because of these differences between the train and test sets, this competition is more about predicting time_to_failure for the particular given test set rather than about predicting time_to_failure in general [which I suppose was an initial goal of this competition].</p>",
      "rawMarkdown": "Hi guys! An analysis of the acoustic signal in the train and the test sets may lead you to the conlusion that their distributions are completely different across the majority of features used in the publicly available kernels. Even if you first separately standardize each 150 000 chunk of data, then many features are still distributed quite differently. This means that traditional supervised learning technics may fail in this competition because they are based on the idea that data-generating procedure is the same for train and test. I think because of these differences between the train and test sets, this competition is more about predicting time_to_failure for the particular given test set rather than about predicting time_to_failure in general [which I suppose was an initial goal of this competition].",
      "votes": null
    },
    {
      "id": "467402",
      "postDate": "02/07/2019 02:59:38",
      "content": "<p>thanks for sharing.enlighten</p>",
      "rawMarkdown": "thanks for sharing.enlighten",
      "votes": null
    },
    {
      "id": "467575",
      "postDate": "02/07/2019 11:00:50",
      "content": "<p>I reaffirmed public kernels and now I have the same concern.\nWhat I discovered through my experiment with the public kernel is:\n1) Most models show good predictions for highly periodic parts and bad predictive graphs for non-periodic parts.\n2) Overall, LB scores tend to be significantly better than CV scores (More than 0.5 sec).</p>\n\n<p>Considering that the distribution of 'time_to_failure' values ​​in the training set is not constant, the first one looks very natural.\nThe concern is that the second symptom is very severe.</p>\n\n<p>It is well known that the distribution of the training set makes the model bias. The problem is that the distribution of training sets and test sets seems to be significantly different.</p>\n\n<p>I am worried that the winner model would not  be 'good model' but the model that predicts the distribution of the final(private) test set.\nI would like to hear from other participants.</p>",
      "rawMarkdown": "I reaffirmed public kernels and now I have the same concern.\nWhat I discovered through my experiment with the public kernel is:\n1) Most models show good predictions for highly periodic parts and bad predictive graphs for non-periodic parts.\n2) Overall, LB scores tend to be significantly better than CV scores (More than 0.5 sec).\n\nConsidering that the distribution of 'time_to_failure' values ​​in the training set is not constant, the first one looks very natural.\nThe concern is that the second symptom is very severe.\n\nIt is well known that the distribution of the training set makes the model bias. The problem is that the distribution of training sets and test sets seems to be significantly different.\n\nI am worried that the winner model would not  be 'good model' but the model that predicts the distribution of the final(private) test set.\nI would like to hear from other participants.",
      "votes": null
    },
    {
      "id": "467635",
      "postDate": "02/07/2019 13:03:25",
      "content": "<p>Thanks Kostya for enlighting that, I hadn't noticed before.</p>\n\n<p>There are two different concerns according to me : \n- the distribution of time to failure, in the public LB, is very different to the distribution in train. That's probably because there are just one or two earthquakes in the public LB, which have a low time to failure. Predicting well on this public LB doesn't mean that we predict well on the private LB. But predicting well on training data should mean that we predict well the private LB: there are more earthquakes in private LB, which are likely to be close to the train earthquakes.\n- the distribution of the features: for example the 90th percentile of signal: the distribution is quite different between train and test data. This can lead to a big deviation on both public and private LB. And public LB is too small to help us to evaluate this deviation. Some features are very different (more than one standard deviation) between train and test data, I think these features should be removed from the model.</p>",
      "rawMarkdown": "Thanks Kostya for enlighting that, I hadn't noticed before.\n\nThere are two different concerns according to me : \n- the distribution of time to failure, in the public LB, is very different to the distribution in train. That's probably because there are just one or two earthquakes in the public LB, which have a low time to failure. Predicting well on this public LB doesn't mean that we predict well on the private LB. But predicting well on training data should mean that we predict well the private LB: there are more earthquakes in private LB, which are likely to be close to the train earthquakes.\n- the distribution of the features: for example the 90th percentile of signal: the distribution is quite different between train and test data. This can lead to a big deviation on both public and private LB. And public LB is too small to help us to evaluate this deviation. Some features are very different (more than one standard deviation) between train and test data, I think these features should be removed from the model.",
      "votes": null
    },
    {
      "id": "473909",
      "postDate": "02/18/2019 17:06:25",
      "content": "<p>The competition hosts may have been thinking along the same lines as <a href=\"/ultragamza\">@ultragamza</a> when they created the public data set.  The samples for the public data set may have been deliberately chosen to be outliers in an attempt to separate good (general) solutions from bad (overfit) solutions.  It is an unfortunate but real fact that in other competitions some competitors have interrogated the public data set with preliminary submissions designed to estimate the public data set's distribution.  The information gained by these interrogation submissions are then used to create a solution that can overfit the data very accurately.  If the competition hosts have chosen the public data set unwisely and the private data set has the same distribution as the public data set, the overfit solution wins.  The hosts for this competition could have chosen the public data set specifically to combat the interrogate-and-overfit approach, such that if a solution fits the public data set too well, it will perform poorly on the public data set.  This would ensure that only general solutions win the competition.  In other words, the competition hosts could have chosen the public data set such that its results provide no insight into actual performance.  This forces competitors to make blind submissions based on solid machine learning fundamentals.  That's how I would do it if I were them.</p>",
      "rawMarkdown": "The competition hosts may have been thinking along the same lines as @ultragamza when they created the public data set.  The samples for the public data set may have been deliberately chosen to be outliers in an attempt to separate good (general) solutions from bad (overfit) solutions.  It is an unfortunate but real fact that in other competitions some competitors have interrogated the public data set with preliminary submissions designed to estimate the public data set's distribution.  The information gained by these interrogation submissions are then used to create a solution that can overfit the data very accurately.  If the competition hosts have chosen the public data set unwisely and the private data set has the same distribution as the public data set, the overfit solution wins.  The hosts for this competition could have chosen the public data set specifically to combat the interrogate-and-overfit approach, such that if a solution fits the public data set too well, it will perform poorly on the public data set.  This would ensure that only general solutions win the competition.  In other words, the competition hosts could have chosen the public data set such that its results provide no insight into actual performance.  This forces competitors to make blind submissions based on solid machine learning fundamentals.  That's how I would do it if I were them.",
      "votes": null
    },
    {
      "id": "476294",
      "postDate": "02/21/2019 23:11:10",
      "content": "<p>great post, what do you think about the LB portion of the Test set is it contignous, or randomly sampled? Based on your comment it is a contignuous outlying portion.</p>",
      "rawMarkdown": "great post, what do you think about the LB portion of the Test set is it contignous, or randomly sampled? Based on your comment it is a contignuous outlying portion.",
      "votes": null
    },
    {
      "id": "476714",
      "postDate": "02/22/2019 15:37:14",
      "content": "<p>The hosts have indicated that the public test samples are contiguous data within a sample, but the samples are not contiguous with one another.  Each sample could be an anomaly (of varying degrees).  However, problems become very consistent with the right choice of features.  The search for a minimal feature set that cleanly separates the data is the holy grail of data scientists.  Perhaps the public test samples are not outliers when viewed with the proper feature set.  But so far it doesn't appear to be that way.</p>",
      "rawMarkdown": "The hosts have indicated that the public test samples are contiguous data within a sample, but the samples are not contiguous with one another.  Each sample could be an anomaly (of varying degrees).  However, problems become very consistent with the right choice of features.  The search for a minimal feature set that cleanly separates the data is the holy grail of data scientists.  Perhaps the public test samples are not outliers when viewed with the proper feature set.  But so far it doesn't appear to be that way.",
      "votes": null
    },
    {
      "id": "522135",
      "postDate": "04/23/2019 22:38:25",
      "content": "<p><a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda\">This kernel</a> mentioned and showed the problem, too. However, as I've also pointed out there, the mean of acoustic data should physically be zero. There seems to be some systematic measurement error we must correct for (i.e. subtract the mean from the data) - apparently separately for the training and testing data (and maybe also the different segment between quakes int the training data). Most discrepancies go away then (as other features depend on the mean, such as the quantiles).</p>",
      "rawMarkdown": "[This kernel](https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda) mentioned and showed the problem, too. However, as I've also pointed out there, the mean of acoustic data should physically be zero. There seems to be some systematic measurement error we must correct for (i.e. subtract the mean from the data) - apparently separately for the training and testing data (and maybe also the different segment between quakes int the training data). Most discrepancies go away then (as other features depend on the mean, such as the quantiles).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 467402,
      "author_name": "soysouce",
      "author_url": "",
      "post_date": "02/07/2019 02:59:38",
      "content": "<p>thanks for sharing.enlighten</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 467575,
      "author_name": "ultragamza",
      "author_url": "",
      "post_date": "02/07/2019 11:00:50",
      "content": "<p>I reaffirmed public kernels and now I have the same concern.\nWhat I discovered through my experiment with the public kernel is:\n1) Most models show good predictions for highly periodic parts and bad predictive graphs for non-periodic parts.\n2) Overall, LB scores tend to be significantly better than CV scores (More than 0.5 sec).</p>\n\n<p>Considering that the distribution of 'time_to_failure' values ​​in the training set is not constant, the first one looks very natural.\nThe concern is that the second symptom is very severe.</p>\n\n<p>It is well known that the distribution of the training set makes the model bias. The problem is that the distribution of training sets and test sets seems to be significantly different.</p>\n\n<p>I am worried that the winner model would not  be 'good model' but the model that predicts the distribution of the final(private) test set.\nI would like to hear from other participants.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 467635,
      "author_name": "zidmie",
      "author_url": "",
      "post_date": "02/07/2019 13:03:25",
      "content": "<p>Thanks Kostya for enlighting that, I hadn't noticed before.</p>\n\n<p>There are two different concerns according to me : \n- the distribution of time to failure, in the public LB, is very different to the distribution in train. That's probably because there are just one or two earthquakes in the public LB, which have a low time to failure. Predicting well on this public LB doesn't mean that we predict well on the private LB. But predicting well on training data should mean that we predict well the private LB: there are more earthquakes in private LB, which are likely to be close to the train earthquakes.\n- the distribution of the features: for example the 90th percentile of signal: the distribution is quite different between train and test data. This can lead to a big deviation on both public and private LB. And public LB is too small to help us to evaluate this deviation. Some features are very different (more than one standard deviation) between train and test data, I think these features should be removed from the model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 473909,
      "author_name": "rockymountaineli",
      "author_url": "",
      "post_date": "02/18/2019 17:06:25",
      "content": "<p>The competition hosts may have been thinking along the same lines as <a href=\"/ultragamza\">@ultragamza</a> when they created the public data set.  The samples for the public data set may have been deliberately chosen to be outliers in an attempt to separate good (general) solutions from bad (overfit) solutions.  It is an unfortunate but real fact that in other competitions some competitors have interrogated the public data set with preliminary submissions designed to estimate the public data set's distribution.  The information gained by these interrogation submissions are then used to create a solution that can overfit the data very accurately.  If the competition hosts have chosen the public data set unwisely and the private data set has the same distribution as the public data set, the overfit solution wins.  The hosts for this competition could have chosen the public data set specifically to combat the interrogate-and-overfit approach, such that if a solution fits the public data set too well, it will perform poorly on the public data set.  This would ensure that only general solutions win the competition.  In other words, the competition hosts could have chosen the public data set such that its results provide no insight into actual performance.  This forces competitors to make blind submissions based on solid machine learning fundamentals.  That's how I would do it if I were them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 476294,
          "author_name": "petersoltesz",
          "author_url": "",
          "post_date": "02/21/2019 23:11:10",
          "content": "<p>great post, what do you think about the LB portion of the Test set is it contignous, or randomly sampled? Based on your comment it is a contignuous outlying portion.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 476714,
          "author_name": "rockymountaineli",
          "author_url": "",
          "post_date": "02/22/2019 15:37:14",
          "content": "<p>The hosts have indicated that the public test samples are contiguous data within a sample, but the samples are not contiguous with one another.  Each sample could be an anomaly (of varying degrees).  However, problems become very consistent with the right choice of features.  The search for a minimal feature set that cleanly separates the data is the holy grail of data scientists.  Perhaps the public test samples are not outliers when viewed with the proper feature set.  But so far it doesn't appear to be that way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 522135,
      "author_name": "bernir",
      "author_url": "",
      "post_date": "04/23/2019 22:38:25",
      "content": "<p><a href=\"https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda\">This kernel</a> mentioned and showed the problem, too. However, as I've also pointed out there, the mean of acoustic data should physically be zero. There seems to be some systematic measurement error we must correct for (i.e. subtract the mean from the data) - apparently separately for the training and testing data (and maybe also the different segment between quakes int the training data). Most discrepancies go away then (as other features depend on the mean, such as the quantiles).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "465982": "Hi guys! An analysis of the acoustic signal in the train and the test sets may lead you to the conlusion that their distributions are completely different across the majority of features used in the publicly available kernels. Even if you first separately standardize each 150 000 chunk of data, then many features are still distributed quite differently. This means that traditional supervised learning technics may fail in this competition because they are based on the idea that data-generating procedure is the same for train and test. I think because of these differences between the train and test sets, this competition is more about predicting time_to_failure for the particular given test set rather than about predicting time_to_failure in general [which I suppose was an initial goal of this competition].",
    "467402": "thanks for sharing.enlighten",
    "467575": "I reaffirmed public kernels and now I have the same concern.\nWhat I discovered through my experiment with the public kernel is:\n1) Most models show good predictions for highly periodic parts and bad predictive graphs for non-periodic parts.\n2) Overall, LB scores tend to be significantly better than CV scores (More than 0.5 sec).\n\nConsidering that the distribution of 'time_to_failure' values ​​in the training set is not constant, the first one looks very natural.\nThe concern is that the second symptom is very severe.\n\nIt is well known that the distribution of the training set makes the model bias. The problem is that the distribution of training sets and test sets seems to be significantly different.\n\nI am worried that the winner model would not  be 'good model' but the model that predicts the distribution of the final(private) test set.\nI would like to hear from other participants.",
    "467635": "Thanks Kostya for enlighting that, I hadn't noticed before.\n\nThere are two different concerns according to me : \n- the distribution of time to failure, in the public LB, is very different to the distribution in train. That's probably because there are just one or two earthquakes in the public LB, which have a low time to failure. Predicting well on this public LB doesn't mean that we predict well on the private LB. But predicting well on training data should mean that we predict well the private LB: there are more earthquakes in private LB, which are likely to be close to the train earthquakes.\n- the distribution of the features: for example the 90th percentile of signal: the distribution is quite different between train and test data. This can lead to a big deviation on both public and private LB. And public LB is too small to help us to evaluate this deviation. Some features are very different (more than one standard deviation) between train and test data, I think these features should be removed from the model.",
    "473909": "The competition hosts may have been thinking along the same lines as @ultragamza when they created the public data set.  The samples for the public data set may have been deliberately chosen to be outliers in an attempt to separate good (general) solutions from bad (overfit) solutions.  It is an unfortunate but real fact that in other competitions some competitors have interrogated the public data set with preliminary submissions designed to estimate the public data set's distribution.  The information gained by these interrogation submissions are then used to create a solution that can overfit the data very accurately.  If the competition hosts have chosen the public data set unwisely and the private data set has the same distribution as the public data set, the overfit solution wins.  The hosts for this competition could have chosen the public data set specifically to combat the interrogate-and-overfit approach, such that if a solution fits the public data set too well, it will perform poorly on the public data set.  This would ensure that only general solutions win the competition.  In other words, the competition hosts could have chosen the public data set such that its results provide no insight into actual performance.  This forces competitors to make blind submissions based on solid machine learning fundamentals.  That's how I would do it if I were them.",
    "476294": "great post, what do you think about the LB portion of the Test set is it contignous, or randomly sampled? Based on your comment it is a contignuous outlying portion.",
    "476714": "The hosts have indicated that the public test samples are contiguous data within a sample, but the samples are not contiguous with one another.  Each sample could be an anomaly (of varying degrees).  However, problems become very consistent with the right choice of features.  The search for a minimal feature set that cleanly separates the data is the holy grail of data scientists.  Perhaps the public test samples are not outliers when viewed with the proper feature set.  But so far it doesn't appear to be that way.",
    "522135": "[This kernel](https://www.kaggle.com/gpreda/lanl-earthquake-new-approach-eda) mentioned and showed the problem, too. However, as I've also pointed out there, the mean of acoustic data should physically be zero. There seems to be some systematic measurement error we must correct for (i.e. subtract the mean from the data) - apparently separately for the training and testing data (and maybe also the different segment between quakes int the training data). Most discrepancies go away then (as other features depend on the mean, such as the quantiles)."
  },
  "source": "meta"
}