{
  "id": 77447,
  "title": "A simple question about the data files",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/77447",
  "author_name": "",
  "post_date": "2019-01-12T22:14:52.212334400Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Sorry I am a little confused about how the information is arranged. </p>\n\n<p>My understanding is that data pieces can be easily arranged in the following ways:</p>\n\n<p>outcome predictor1 predictor2 ... predictorN</p>\n\n<p>However, the file train.csv has two columns like the following:</p>\n\n<p>acoustic_data,time_to_failure\n12,1.4690999832\n6,1.4690999821\n8,1.469099981\n...</p>\n\n<p>Which one is the outcome and which one is the predictor? It looks like neither. </p>\n\n<p>The test has only one column called \"acoustic_data\". I bet this column contains a time series as the input predictor and the outcome is to be provided by the competition participants, one prediction for one file. A series of numbers corresponds to one outcome. This makes sense. But how come the train.csv file has only one time point (acoustic_data) corresponding to an outcome?</p>\n\n<p>Thanks for any enlightenment!</p>",
  "messages": [
    {
      "id": "455070",
      "postDate": "01/12/2019 22:14:52",
      "content": "<p>Sorry I am a little confused about how the information is arranged. </p>\n\n<p>My understanding is that data pieces can be easily arranged in the following ways:</p>\n\n<p>outcome predictor1 predictor2 ... predictorN</p>\n\n<p>However, the file train.csv has two columns like the following:</p>\n\n<p>acoustic_data,time_to_failure\n12,1.4690999832\n6,1.4690999821\n8,1.469099981\n...</p>\n\n<p>Which one is the outcome and which one is the predictor? It looks like neither. </p>\n\n<p>The test has only one column called \"acoustic_data\". I bet this column contains a time series as the input predictor and the outcome is to be provided by the competition participants, one prediction for one file. A series of numbers corresponds to one outcome. This makes sense. But how come the train.csv file has only one time point (acoustic_data) corresponding to an outcome?</p>\n\n<p>Thanks for any enlightenment!</p>",
      "rawMarkdown": "Sorry I am a little confused about how the information is arranged. \n\nMy understanding is that data pieces can be easily arranged in the following ways:\n\noutcome predictor1 predictor2 ... predictorN\n\nHowever, the file train.csv has two columns like the following:\n\nacoustic_data,time_to_failure\n12,1.4690999832\n6,1.4690999821\n8,1.469099981\n...\n\nWhich one is the outcome and which one is the predictor? It looks like neither. \n\nThe test has only one column called \"acoustic_data\". I bet this column contains a time series as the input predictor and the outcome is to be provided by the competition participants, one prediction for one file. A series of numbers corresponds to one outcome. This makes sense. But how come the train.csv file has only one time point (acoustic_data) corresponding to an outcome?\n\nThanks for any enlightenment!",
      "votes": null
    },
    {
      "id": "455074",
      "postDate": "01/12/2019 22:28:24",
      "content": "<p>I think:\nThe first column is the \"predictor\" - it is one point in a massive time series. But one can't predict the time to failure from one sample in the series - one would use a bunch of neighbouring samples, like each of the test files where the entire file is used to predict, basically, one time-to-failure (given that one such time can automatically give the prediction for all the samples in the regular time series).</p>",
      "rawMarkdown": "I think:\nThe first column is the \"predictor\" - it is one point in a massive time series. But one can't predict the time to failure from one sample in the series - one would use a bunch of neighbouring samples, like each of the test files where the entire file is used to predict, basically, one time-to-failure (given that one such time can automatically give the prediction for all the samples in the regular time series).",
      "votes": null
    },
    {
      "id": "455097",
      "postDate": "01/13/2019 01:20:08",
      "content": "<p>OK, got it. Thanks much! Hope the data points are ordered in time from past to future.</p>",
      "rawMarkdown": "OK, got it. Thanks much! Hope the data points are ordered in time from past to future.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 455074,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "01/12/2019 22:28:24",
      "content": "<p>I think:\nThe first column is the \"predictor\" - it is one point in a massive time series. But one can't predict the time to failure from one sample in the series - one would use a bunch of neighbouring samples, like each of the test files where the entire file is used to predict, basically, one time-to-failure (given that one such time can automatically give the prediction for all the samples in the regular time series).</p>",
      "votes": null,
      "replies": [
        {
          "id": 455097,
          "author_name": "moushengxu",
          "author_url": "",
          "post_date": "01/13/2019 01:20:08",
          "content": "<p>OK, got it. Thanks much! Hope the data points are ordered in time from past to future.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "455070": "Sorry I am a little confused about how the information is arranged. \n\nMy understanding is that data pieces can be easily arranged in the following ways:\n\noutcome predictor1 predictor2 ... predictorN\n\nHowever, the file train.csv has two columns like the following:\n\nacoustic_data,time_to_failure\n12,1.4690999832\n6,1.4690999821\n8,1.469099981\n...\n\nWhich one is the outcome and which one is the predictor? It looks like neither. \n\nThe test has only one column called \"acoustic_data\". I bet this column contains a time series as the input predictor and the outcome is to be provided by the competition participants, one prediction for one file. A series of numbers corresponds to one outcome. This makes sense. But how come the train.csv file has only one time point (acoustic_data) corresponding to an outcome?\n\nThanks for any enlightenment!",
    "455074": "I think:\nThe first column is the \"predictor\" - it is one point in a massive time series. But one can't predict the time to failure from one sample in the series - one would use a bunch of neighbouring samples, like each of the test files where the entire file is used to predict, basically, one time-to-failure (given that one such time can automatically give the prediction for all the samples in the regular time series).",
    "455097": "OK, got it. Thanks much! Hope the data points are ordered in time from past to future."
  },
  "source": "meta"
}