{
  "id": 89542,
  "title": "Consensus regarding test set...",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/89542",
  "author_name": "",
  "post_date": "2019-04-15T13:39:17.324880200Z",
  "votes": 12,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Just want to clarify what can we expect regarding test set (private and public).</p>\n\n<p>Officially Bertrand RL said \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\"</p>\n\n<p>On couple of occasions I saw different claims on forum discussions regarding following question:\nAre the test samples continuous in time in a sence that they were constructed the same way we split the training set. There was an continuous experiment and competition host looped sequentially and made cuts at the 150 000 mark. Or are the test sample rather random samples of these cuts (no overlapping is meant in anycase!)?</p>\n\n<p>I think it is important to clarify it because of various reasons:\n1. Do rNN make sence, in the case of random samples not really...\n2. How do we construct strong CV, again this information helps\n3. Feature engineering, if some features depend/spill the information in the next 150k sample than, random samples kill the effect of the feature\n4. etc...</p>\n\n<p>Assumptions: \n<a href=\"/olivier\">@olivier</a> pointed out that <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-516997\">Shuffling</a>, i.e. removing time increases CV and LB scores. Hence there is possiblity that test samples are also suffled.\nThey will be definately shuffled when the private scoring takes place (atleast to an extent of 13% versus 87%) but <strong>I am wondering will the 87% of private test data be time dependent?</strong></p>",
  "messages": [
    {
      "id": "517079",
      "postDate": "04/15/2019 13:39:17",
      "content": "<p>Just want to clarify what can we expect regarding test set (private and public).</p>\n\n<p>Officially Bertrand RL said \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\"</p>\n\n<p>On couple of occasions I saw different claims on forum discussions regarding following question:\nAre the test samples continuous in time in a sence that they were constructed the same way we split the training set. There was an continuous experiment and competition host looped sequentially and made cuts at the 150 000 mark. Or are the test sample rather random samples of these cuts (no overlapping is meant in anycase!)?</p>\n\n<p>I think it is important to clarify it because of various reasons:\n1. Do rNN make sence, in the case of random samples not really...\n2. How do we construct strong CV, again this information helps\n3. Feature engineering, if some features depend/spill the information in the next 150k sample than, random samples kill the effect of the feature\n4. etc...</p>\n\n<p>Assumptions: \n<a href=\"/olivier\">@olivier</a> pointed out that <a href=\"https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-516997\">Shuffling</a>, i.e. removing time increases CV and LB scores. Hence there is possiblity that test samples are also suffled.\nThey will be definately shuffled when the private scoring takes place (atleast to an extent of 13% versus 87%) but <strong>I am wondering will the 87% of private test data be time dependent?</strong></p>",
      "rawMarkdown": "Just want to clarify what can we expect regarding test set (private and public).\n\nOfficially Bertrand RL said \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\"\n\nOn couple of occasions I saw different claims on forum discussions regarding following question:\nAre the test samples continuous in time in a sence that they were constructed the same way we split the training set. There was an continuous experiment and competition host looped sequentially and made cuts at the 150 000 mark. Or are the test sample rather random samples of these cuts (no overlapping is meant in anycase!)?\n\nI think it is important to clarify it because of various reasons:\n1. Do rNN make sence, in the case of random samples not really...\n2. How do we construct strong CV, again this information helps\n3. Feature engineering, if some features depend/spill the information in the next 150k sample than, random samples kill the effect of the feature\n4. etc...\n\nAssumptions: \n@olivier pointed out that [Shuffling](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-516997), i.e. removing time increases CV and LB scores. Hence there is possiblity that test samples are also suffled.\nThey will be definately shuffled when the private scoring takes place (atleast to an extent of 13% versus 87%) but **I am wondering will the 87% of private test data be time dependent?**",
      "votes": null
    },
    {
      "id": "517129",
      "postDate": "04/15/2019 15:15:57",
      "content": "<p>i vote for the 2nd</p>",
      "rawMarkdown": "i vote for the 2nd",
      "votes": null
    },
    {
      "id": "517153",
      "postDate": "04/15/2019 16:14:36",
      "content": "<p>Answer is in the data page:</p>\n\n<blockquote>\n  <p>The training data is a single, continuous segment of experimental data. The test data consists of a folder containing many small segments. The data within each test file is continuous, but the test files do not represent a continuous segment of the experiment; thus, the predictions cannot be assumed to follow the same regular pattern seen in the training file. </p>\n</blockquote>\n\n<p>Disclaimer: I may purposely mislead people each time I share info.  This disclaimer has been made mandatory because of Santander competition where I was accused of misleading people when I shared.</p>",
      "rawMarkdown": "Answer is in the data page:\n\n&gt; The training data is a single, continuous segment of experimental data. The test data consists of a folder containing many small segments. The data within each test file is continuous, but the test files do not represent a continuous segment of the experiment; thus, the predictions cannot be assumed to follow the same regular pattern seen in the training file. \n\nDisclaimer: I may purposely mislead people each time I share info.  This disclaimer has been made mandatory because of Santander competition where I was accused of misleading people when I shared.",
      "votes": null
    },
    {
      "id": "517169",
      "postDate": "04/15/2019 16:45:31",
      "content": "<p>haha! The disclaimer is gold. It is also unnecessary, I think I read all your comments in santander and I would not call any misleading. Many were (properly) vague so I suppose one could mislead themselves depending on their interpretation. The main problem with the disclaimer is that it may actually be misleading to warn that a post may be misleading when it is not, or am I falling for some grandmaster-level misdirection...?</p>\n\n<p>In any case, cv is a challenge here. I have not had much success when taking segments from train of length other than 150k rows, certainly segments under 75k rows have been terrible but maybe I'm not looking at the big picture.</p>",
      "rawMarkdown": "haha! The disclaimer is gold. It is also unnecessary, I think I read all your comments in santander and I would not call any misleading. Many were (properly) vague so I suppose one could mislead themselves depending on their interpretation. The main problem with the disclaimer is that it may actually be misleading to warn that a post may be misleading when it is not, or am I falling for some grandmaster-level misdirection...?\n\nIn any case, cv is a challenge here. I have not had much success when taking segments from train of length other than 150k rows, certainly segments under 75k rows have been terrible but maybe I'm not looking at the big picture.",
      "votes": null
    },
    {
      "id": "517206",
      "postDate": "04/15/2019 17:58:53",
      "content": "<p>I won't repeat the disclaimer each time don't worry ;)</p>",
      "rawMarkdown": "I won't repeat the disclaimer each time don't worry ;)",
      "votes": null
    },
    {
      "id": "517210",
      "postDate": "04/15/2019 18:00:40",
      "content": "<p>@CPMP\nI am grateful beeing \"mislead\" by you, thanks in any case.</p>\n\n<p><a href=\"/interneuron\">@interneuron</a> <a href=\"https://www.imdb.com/title/tt1375666/\">this?</a></p>",
      "rawMarkdown": "CPMP\nI am grateful beeing \"mislead\" by you, thanks in any case.\n\n@interneuron [this?](https://www.imdb.com/title/tt1375666/)",
      "votes": null
    },
    {
      "id": "517217",
      "postDate": "04/15/2019 18:12:29",
      "content": "<p>That. </p>\n\n<p>Perhaps I've gone too deep already.</p>",
      "rawMarkdown": "That. \n\nPerhaps I've gone too deep already.",
      "votes": null
    },
    {
      "id": "517597",
      "postDate": "04/16/2019 08:44:22",
      "content": "<p>Some of my observations and line of thinking:</p>\n\n<p>1) rNN can make sense within a single sample (so the 150_000 sample points). However it seems that decision tree based solutions are doing better (as is often the case with regression problems). Of course in the end it is the huge ensemble that will win ;)</p>\n\n<p>2) I think the real challenge is that the test set seems to have a different distribution than the training set.  This shows up as the LB scores are very different from the CV scores. But also when I plot the test predictions it seems very unlikely that the test set is \"randomly uniform\" drawn. So perhaps making an educated guess how they sampled the test set and use this somehow in the predictions.  </p>\n\n<p>3) You would expect that the experiment changes over time. So the outcome of the test set given that the training set has been observed, might add additional insights if modelled correctly.</p>",
      "rawMarkdown": "Some of my observations and line of thinking:\n\n1) rNN can make sense within a single sample (so the 150_000 sample points). However it seems that decision tree based solutions are doing better (as is often the case with regression problems). Of course in the end it is the huge ensemble that will win ;)\n\n2) I think the real challenge is that the test set seems to have a different distribution than the training set.  This shows up as the LB scores are very different from the CV scores. But also when I plot the test predictions it seems very unlikely that the test set is \"randomly uniform\" drawn. So perhaps making an educated guess how they sampled the test set and use this somehow in the predictions.  \n\n3) You would expect that the experiment changes over time. So the outcome of the test set given that the training set has been observed, might add additional insights if modelled correctly.",
      "votes": null
    },
    {
      "id": "519010",
      "postDate": "04/18/2019 08:18:53",
      "content": "<p>Is there any evidence / info whether the segments in test data can overlap with each other?</p>",
      "rawMarkdown": "Is there any evidence / info whether the segments in test data can overlap with each other?",
      "votes": null
    },
    {
      "id": "519164",
      "postDate": "04/18/2019 13:36:45",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> </p>\n\n<p>Well organisator said:</p>\n\n<ol>\n<li><p>\"The data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\"\nand</p></li>\n<li><p>\"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\"</p></li>\n</ol>\n\n<p>So if they are from the same experiment (test and train data) and train data is not overlapping within itself, I think we could conclude no overlapping.</p>",
      "rawMarkdown": "philippsinger \n\nWell organisator said:\n\n1. \"The data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\"\nand\n\n2. \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\"\n\nSo if they are from the same experiment (test and train data) and train data is not overlapping within itself, I think we could conclude no overlapping.",
      "votes": null
    },
    {
      "id": "519211",
      "postDate": "04/18/2019 14:46:33",
      "content": "<p>tbh I always thought it was common sense to be careful when taking advice from other competitors. Not that I've ever found others to be purposefully misleading... people are incredibly helpful here. Still, it is your responsibility to understand and think for yourself. </p>",
      "rawMarkdown": "tbh I always thought it was common sense to be careful when taking advice from other competitors. Not that I've ever found others to be purposefully misleading... people are incredibly helpful here. Still, it is your responsibility to understand and think for yourself.",
      "votes": null
    },
    {
      "id": "519614",
      "postDate": "04/19/2019 09:36:01",
      "content": "<p>The people who claim you mislead them are the type of people who leave a restaurant without paying and then put a bad review on tripadvisor if they get food poisoning!</p>",
      "rawMarkdown": "The people who claim you mislead them are the type of people who leave a restaurant without paying and then put a bad review on tripadvisor if they get food poisoning!",
      "votes": null
    },
    {
      "id": "525285",
      "postDate": "04/30/2019 16:12:24",
      "content": "<p>I'm not sure that the train and test set have different distributions. The difference between LB and CV scores just shows that the train and the public portion of test have different distributions.</p>",
      "rawMarkdown": "I'm not sure that the train and test set have different distributions. The difference between LB and CV scores just shows that the train and the public portion of test have different distributions.",
      "votes": null
    },
    {
      "id": "525350",
      "postDate": "04/30/2019 19:06:49",
      "content": "<p>Not really, you can easily get a validation score similar to the private LB score in one fold. Thus, it may just be an easy to model part of the test set. </p>",
      "rawMarkdown": "Not really, you can easily get a validation score similar to the private LB score in one fold. Thus, it may just be an easy to model part of the test set.",
      "votes": null
    },
    {
      "id": "525381",
      "postDate": "04/30/2019 21:11:40",
      "content": "<p>The difference in distribution was based on a quick investigation of plotting the histogram of the values of the audio. So the histogram of the training set audio was (slightly) different from the test set.</p>\n\n<p>However doing some more investigation showed that the difference in the training set itself between two different quakes is even more different. So I guess indeed the difference in distribution is indeed not too significant.</p>",
      "rawMarkdown": "The difference in distribution was based on a quick investigation of plotting the histogram of the values of the audio. So the histogram of the training set audio was (slightly) different from the test set.\n\nHowever doing some more investigation showed that the difference in the training set itself between two different quakes is even more different. So I guess indeed the difference in distribution is indeed not too significant.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 517129,
      "author_name": "dkaraflos",
      "author_url": "",
      "post_date": "04/15/2019 15:15:57",
      "content": "<p>i vote for the 2nd</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 517153,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/15/2019 16:14:36",
      "content": "<p>Answer is in the data page:</p>\n\n<blockquote>\n  <p>The training data is a single, continuous segment of experimental data. The test data consists of a folder containing many small segments. The data within each test file is continuous, but the test files do not represent a continuous segment of the experiment; thus, the predictions cannot be assumed to follow the same regular pattern seen in the training file. </p>\n</blockquote>\n\n<p>Disclaimer: I may purposely mislead people each time I share info.  This disclaimer has been made mandatory because of Santander competition where I was accused of misleading people when I shared.</p>",
      "votes": null,
      "replies": [
        {
          "id": 517169,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "04/15/2019 16:45:31",
          "content": "<p>haha! The disclaimer is gold. It is also unnecessary, I think I read all your comments in santander and I would not call any misleading. Many were (properly) vague so I suppose one could mislead themselves depending on their interpretation. The main problem with the disclaimer is that it may actually be misleading to warn that a post may be misleading when it is not, or am I falling for some grandmaster-level misdirection...?</p>\n\n<p>In any case, cv is a challenge here. I have not had much success when taking segments from train of length other than 150k rows, certainly segments under 75k rows have been terrible but maybe I'm not looking at the big picture.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517206,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/15/2019 17:58:53",
          "content": "<p>I won't repeat the disclaimer each time don't worry ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517210,
          "author_name": "zikazika",
          "author_url": "",
          "post_date": "04/15/2019 18:00:40",
          "content": "<p>@CPMP\nI am grateful beeing \"mislead\" by you, thanks in any case.</p>\n\n<p><a href=\"/interneuron\">@interneuron</a> <a href=\"https://www.imdb.com/title/tt1375666/\">this?</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 517217,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "04/15/2019 18:12:29",
          "content": "<p>That. </p>\n\n<p>Perhaps I've gone too deep already.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519211,
          "author_name": "kennethhesse",
          "author_url": "",
          "post_date": "04/18/2019 14:46:33",
          "content": "<p>tbh I always thought it was common sense to be careful when taking advice from other competitors. Not that I've ever found others to be purposefully misleading... people are incredibly helpful here. Still, it is your responsibility to understand and think for yourself. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 519614,
          "author_name": "scirpus",
          "author_url": "",
          "post_date": "04/19/2019 09:36:01",
          "content": "<p>The people who claim you mislead them are the type of people who leave a restaurant without paying and then put a bad review on tripadvisor if they get food poisoning!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 517597,
      "author_name": "peterdekkers101",
      "author_url": "",
      "post_date": "04/16/2019 08:44:22",
      "content": "<p>Some of my observations and line of thinking:</p>\n\n<p>1) rNN can make sense within a single sample (so the 150_000 sample points). However it seems that decision tree based solutions are doing better (as is often the case with regression problems). Of course in the end it is the huge ensemble that will win ;)</p>\n\n<p>2) I think the real challenge is that the test set seems to have a different distribution than the training set.  This shows up as the LB scores are very different from the CV scores. But also when I plot the test predictions it seems very unlikely that the test set is \"randomly uniform\" drawn. So perhaps making an educated guess how they sampled the test set and use this somehow in the predictions.  </p>\n\n<p>3) You would expect that the experiment changes over time. So the outcome of the test set given that the training set has been observed, might add additional insights if modelled correctly.</p>",
      "votes": null,
      "replies": [
        {
          "id": 525285,
          "author_name": "mateiionita",
          "author_url": "",
          "post_date": "04/30/2019 16:12:24",
          "content": "<p>I'm not sure that the train and test set have different distributions. The difference between LB and CV scores just shows that the train and the public portion of test have different distributions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525350,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/30/2019 19:06:49",
          "content": "<p>Not really, you can easily get a validation score similar to the private LB score in one fold. Thus, it may just be an easy to model part of the test set. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 525381,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "04/30/2019 21:11:40",
          "content": "<p>The difference in distribution was based on a quick investigation of plotting the histogram of the values of the audio. So the histogram of the training set audio was (slightly) different from the test set.</p>\n\n<p>However doing some more investigation showed that the difference in the training set itself between two different quakes is even more different. So I guess indeed the difference in distribution is indeed not too significant.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 519010,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "04/18/2019 08:18:53",
      "content": "<p>Is there any evidence / info whether the segments in test data can overlap with each other?</p>",
      "votes": null,
      "replies": [
        {
          "id": 519164,
          "author_name": "zikazika",
          "author_url": "",
          "post_date": "04/18/2019 13:36:45",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> </p>\n\n<p>Well organisator said:</p>\n\n<ol>\n<li><p>\"The data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\"\nand</p></li>\n<li><p>\"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\"</p></li>\n</ol>\n\n<p>So if they are from the same experiment (test and train data) and train data is not overlapping within itself, I think we could conclude no overlapping.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "517079": "Just want to clarify what can we expect regarding test set (private and public).\n\nOfficially Bertrand RL said \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\"\n\nOn couple of occasions I saw different claims on forum discussions regarding following question:\nAre the test samples continuous in time in a sence that they were constructed the same way we split the training set. There was an continuous experiment and competition host looped sequentially and made cuts at the 150 000 mark. Or are the test sample rather random samples of these cuts (no overlapping is meant in anycase!)?\n\nI think it is important to clarify it because of various reasons:\n1. Do rNN make sence, in the case of random samples not really...\n2. How do we construct strong CV, again this information helps\n3. Feature engineering, if some features depend/spill the information in the next 150k sample than, random samples kill the effect of the feature\n4. etc...\n\nAssumptions: \n@olivier pointed out that [Shuffling](https://www.kaggle.com/c/LANL-Earthquake-Prediction/discussion/89366#latest-516997), i.e. removing time increases CV and LB scores. Hence there is possiblity that test samples are also suffled.\nThey will be definately shuffled when the private scoring takes place (atleast to an extent of 13% versus 87%) but **I am wondering will the 87% of private test data be time dependent?**",
    "517129": "i vote for the 2nd",
    "517153": "Answer is in the data page:\n\n&gt; The training data is a single, continuous segment of experimental data. The test data consists of a folder containing many small segments. The data within each test file is continuous, but the test files do not represent a continuous segment of the experiment; thus, the predictions cannot be assumed to follow the same regular pattern seen in the training file. \n\nDisclaimer: I may purposely mislead people each time I share info.  This disclaimer has been made mandatory because of Santander competition where I was accused of misleading people when I shared.",
    "517169": "haha! The disclaimer is gold. It is also unnecessary, I think I read all your comments in santander and I would not call any misleading. Many were (properly) vague so I suppose one could mislead themselves depending on their interpretation. The main problem with the disclaimer is that it may actually be misleading to warn that a post may be misleading when it is not, or am I falling for some grandmaster-level misdirection...?\n\nIn any case, cv is a challenge here. I have not had much success when taking segments from train of length other than 150k rows, certainly segments under 75k rows have been terrible but maybe I'm not looking at the big picture.",
    "517206": "I won't repeat the disclaimer each time don't worry ;)",
    "517210": "CPMP\nI am grateful beeing \"mislead\" by you, thanks in any case.\n\n@interneuron [this?](https://www.imdb.com/title/tt1375666/)",
    "517217": "That. \n\nPerhaps I've gone too deep already.",
    "517597": "Some of my observations and line of thinking:\n\n1) rNN can make sense within a single sample (so the 150_000 sample points). However it seems that decision tree based solutions are doing better (as is often the case with regression problems). Of course in the end it is the huge ensemble that will win ;)\n\n2) I think the real challenge is that the test set seems to have a different distribution than the training set.  This shows up as the LB scores are very different from the CV scores. But also when I plot the test predictions it seems very unlikely that the test set is \"randomly uniform\" drawn. So perhaps making an educated guess how they sampled the test set and use this somehow in the predictions.  \n\n3) You would expect that the experiment changes over time. So the outcome of the test set given that the training set has been observed, might add additional insights if modelled correctly.",
    "519010": "Is there any evidence / info whether the segments in test data can overlap with each other?",
    "519164": "philippsinger \n\nWell organisator said:\n\n1. \"The data is recorded in bins of 4096 samples. Withing those bins seismic data is recorded at 4MHz, but there is a 12 microseconds gap between each bin, an artifact of the recording device.\"\nand\n\n2. \"Both the training and the testing set come from the same experiment. There is no overlap between the training and testing sets, that are contiguous in time.\"\n\nSo if they are from the same experiment (test and train data) and train data is not overlapping within itself, I think we could conclude no overlapping.",
    "519211": "tbh I always thought it was common sense to be careful when taking advice from other competitors. Not that I've ever found others to be purposefully misleading... people are incredibly helpful here. Still, it is your responsibility to understand and think for yourself.",
    "519614": "The people who claim you mislead them are the type of people who leave a restaurant without paying and then put a bad review on tripadvisor if they get food poisoning!",
    "525285": "I'm not sure that the train and test set have different distributions. The difference between LB and CV scores just shows that the train and the public portion of test have different distributions.",
    "525350": "Not really, you can easily get a validation score similar to the private LB score in one fold. Thus, it may just be an easy to model part of the test set.",
    "525381": "The difference in distribution was based on a quick investigation of plotting the histogram of the values of the audio. So the histogram of the training set audio was (slightly) different from the test set.\n\nHowever doing some more investigation showed that the difference in the training set itself between two different quakes is even more different. So I guess indeed the difference in distribution is indeed not too significant."
  },
  "source": "meta"
}