{
  "id": 191145,
  "title": "Few questions",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/191145",
  "author_name": "",
  "post_date": "2020-10-14T21:51:44.670030200Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>First of all, sorry by advance for my english, not my mother language, probably i will not be cleared in all my questions.</p>\n<p>The subject is very interesting with a huge objective, so i just started looking the dataset to see if i can do something with that, but there are some points i don't understand in the way the datas are presented:</p>\n<ul>\n<li>I loaded the train.csv with the segment_id and the time_to_eruption column. Then I sorted ascending the time_to_eruption column and saw that the segment_ids are not continuous. Is it correct to sort by the time to eruption and ignoring the value of the segment_id in order to get the starting and end points of observations?</li>\n<li>are the datas in each segment sorted by time?</li>\n<li>if so, does the time_to_eruption in the train.csv file corresponds to the begin or the end of the underlying data in segment csv files?</li>\n<li>Did you split the data for generating the train and test set randomly or did you follow some rules. for instance did you splitted the data in order to make sure that two continuous segments ( in a time-series vision) never exist in the test dataset?</li>\n<li>i'm sorry to ask the next question, because you asked to not reclaim for others metadata, but are the characterics of the sensors the same (technical configuration) and do they have the same implementation around the volcano ( distance from the center, altitude,…)?  From my point of view, it can have an influence for the data preprocessing operations.</li>\n</ul>\n<p>Best regards</p>",
  "messages": [
    {
      "id": "1049900",
      "postDate": "10/14/2020 21:51:44",
      "content": "<p>Hi,</p>\n<p>First of all, sorry by advance for my english, not my mother language, probably i will not be cleared in all my questions.</p>\n<p>The subject is very interesting with a huge objective, so i just started looking the dataset to see if i can do something with that, but there are some points i don't understand in the way the datas are presented:</p>\n<ul>\n<li>I loaded the train.csv with the segment_id and the time_to_eruption column. Then I sorted ascending the time_to_eruption column and saw that the segment_ids are not continuous. Is it correct to sort by the time to eruption and ignoring the value of the segment_id in order to get the starting and end points of observations?</li>\n<li>are the datas in each segment sorted by time?</li>\n<li>if so, does the time_to_eruption in the train.csv file corresponds to the begin or the end of the underlying data in segment csv files?</li>\n<li>Did you split the data for generating the train and test set randomly or did you follow some rules. for instance did you splitted the data in order to make sure that two continuous segments ( in a time-series vision) never exist in the test dataset?</li>\n<li>i'm sorry to ask the next question, because you asked to not reclaim for others metadata, but are the characterics of the sensors the same (technical configuration) and do they have the same implementation around the volcano ( distance from the center, altitude,…)?  From my point of view, it can have an influence for the data preprocessing operations.</li>\n</ul>\n<p>Best regards</p>",
      "rawMarkdown": "Hi,\n\nFirst of all, sorry by advance for my english, not my mother language, probably i will not be cleared in all my questions.\n\nThe subject is very interesting with a huge objective, so i just started looking the dataset to see if i can do something with that, but there are some points i don't understand in the way the datas are presented:\n- I loaded the train.csv with the segment_id and the time_to_eruption column. Then I sorted ascending the time_to_eruption column and saw that the segment_ids are not continuous. Is it correct to sort by the time to eruption and ignoring the value of the segment_id in order to get the starting and end points of observations?\n- are the datas in each segment sorted by time?\n- if so, does the time_to_eruption in the train.csv file corresponds to the begin or the end of the underlying data in segment csv files?\n- Did you split the data for generating the train and test set randomly or did you follow some rules. for instance did you splitted the data in order to make sure that two continuous segments ( in a time-series vision) never exist in the test dataset?\n- i'm sorry to ask the next question, because you asked to not reclaim for others metadata, but are the characterics of the sensors the same (technical configuration) and do they have the same implementation around the volcano ( distance from the center, altitude,...)?  From my point of view, it can have an influence for the data preprocessing operations.\n\nBest regards",
      "votes": null
    },
    {
      "id": "1049918",
      "postDate": "10/14/2020 22:10:35",
      "content": "<p>Please see the data description for as much information as we can provide on these topics.</p>",
      "rawMarkdown": "Please see the data description for as much information as we can provide on these topics.",
      "votes": null
    },
    {
      "id": "1050192",
      "postDate": "10/15/2020 06:54:32",
      "content": "<p>Hi, ok.</p>\n<p>Last question : the description indicates that the data were normalized. Which methods were used ?</p>\n<p>Best regards</p>",
      "rawMarkdown": "Hi, ok.\n\nLast question : the description indicates that the data were normalized. Which methods were used ?\n\nBest regards",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1049918,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "10/14/2020 22:10:35",
      "content": "<p>Please see the data description for as much information as we can provide on these topics.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1050192,
      "author_name": "jace2005",
      "author_url": "",
      "post_date": "10/15/2020 06:54:32",
      "content": "<p>Hi, ok.</p>\n<p>Last question : the description indicates that the data were normalized. Which methods were used ?</p>\n<p>Best regards</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1049900": "Hi,\n\nFirst of all, sorry by advance for my english, not my mother language, probably i will not be cleared in all my questions.\n\nThe subject is very interesting with a huge objective, so i just started looking the dataset to see if i can do something with that, but there are some points i don't understand in the way the datas are presented:\n- I loaded the train.csv with the segment_id and the time_to_eruption column. Then I sorted ascending the time_to_eruption column and saw that the segment_ids are not continuous. Is it correct to sort by the time to eruption and ignoring the value of the segment_id in order to get the starting and end points of observations?\n- are the datas in each segment sorted by time?\n- if so, does the time_to_eruption in the train.csv file corresponds to the begin or the end of the underlying data in segment csv files?\n- Did you split the data for generating the train and test set randomly or did you follow some rules. for instance did you splitted the data in order to make sure that two continuous segments ( in a time-series vision) never exist in the test dataset?\n- i'm sorry to ask the next question, because you asked to not reclaim for others metadata, but are the characterics of the sensors the same (technical configuration) and do they have the same implementation around the volcano ( distance from the center, altitude,...)?  From my point of view, it can have an influence for the data preprocessing operations.\n\nBest regards",
    "1049918": "Please see the data description for as much information as we can provide on these topics.",
    "1050192": "Hi, ok.\n\nLast question : the description indicates that the data were normalized. Which methods were used ?\n\nBest regards"
  },
  "source": "meta"
}