{
  "id": 194414,
  "title": "I'm assuming the right ?",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/194414",
  "author_name": "",
  "post_date": "2020-11-01T17:29:17.973812Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I would like to know if this make sense. Let's take the first 3 value in the train.csv:</p>\n<pre><code>segment_id  time_to_eruption\n1136037770 12262005\n1969647810 32739612\n1895879680 14965999\n</code></pre>\n<p>Since, each file represent 10 minuts of registration of the activity, does it make sense take, for each file, the time_to_eruption and try to understand if it is bigger or lower than 10 minuts, in this way it should be possible to apply some time to events statistical methods? <br>\nFor example, let's assume that for the segment 1136037770 the correspondent time_to_eruption is lower than 10 minuts, for second segment_id it is bigger and so on… <br>\nDoes it make sense this ? </p>",
  "messages": [
    {
      "id": "1066381",
      "postDate": "11/01/2020 17:29:17",
      "content": "<p>I would like to know if this make sense. Let's take the first 3 value in the train.csv:</p>\n<pre><code>segment_id  time_to_eruption\n1136037770 12262005\n1969647810 32739612\n1895879680 14965999\n</code></pre>\n<p>Since, each file represent 10 minuts of registration of the activity, does it make sense take, for each file, the time_to_eruption and try to understand if it is bigger or lower than 10 minuts, in this way it should be possible to apply some time to events statistical methods? <br>\nFor example, let's assume that for the segment 1136037770 the correspondent time_to_eruption is lower than 10 minuts, for second segment_id it is bigger and so on… <br>\nDoes it make sense this ? </p>",
      "rawMarkdown": "I would like to know if this make sense. Let's take the first 3 value in the train.csv:\n```\nsegment_id  time_to_eruption\n1136037770 12262005\n1969647810 32739612\n1895879680 14965999\n```\nSince, each file represent 10 minuts of registration of the activity, does it make sense take, for each file, the time_to_eruption and try to understand if it is bigger or lower than 10 minuts, in this way it should be possible to apply some time to events statistical methods? \nFor example, let's assume that for the segment 1136037770 the correspondent time_to_eruption is lower than 10 minuts, for second segment_id it is bigger and so on... \nDoes it make sense this ?",
      "votes": null
    },
    {
      "id": "1066386",
      "postDate": "11/01/2020 17:33:38",
      "content": "<p>So at the end of the day, if my reasoning is right, there are 7 eruptions ?? </p>",
      "rawMarkdown": "So at the end of the day, if my reasoning is right, there are 7 eruptions ??",
      "votes": null
    },
    {
      "id": "1078125",
      "postDate": "11/14/2020 11:47:46",
      "content": "<p>I think I see your point. You are saying that there have been 7 eruptions, since there are 7 data traces where time_to_eruption is less than 10 minutes, correct?</p>\n<p>I am nevertheless not sure if this reasoning holds, because we are not given any information whether the time_to_eruption value is measured from the start of the data trace or from the end of the data trace. I.e. it could as well be the case that if time_to_eruption = 62.5 seconds, the eruption happens 62500 data samples (62.5 seconds) after the last sample of our trace.<br>\nAt least I haven't found any information about the reference of the time_to_eruption value. Have you?</p>\n<p>Could you give assistance here, <a href=\"https://www.kaggle.com/flaxio\" target=\"_blank\">@flaxio</a>, since this is a piece of information you would have in a real life scenario, wouldn't you?</p>",
      "rawMarkdown": "I think I see your point. You are saying that there have been 7 eruptions, since there are 7 data traces where time_to_eruption is less than 10 minutes, correct?\n\nI am nevertheless not sure if this reasoning holds, because we are not given any information whether the time_to_eruption value is measured from the start of the data trace or from the end of the data trace. I.e. it could as well be the case that if time_to_eruption = 62.5 seconds, the eruption happens 62500 data samples (62.5 seconds) after the last sample of our trace.\nAt least I haven't found any information about the reference of the time_to_eruption value. Have you?\n\nCould you give assistance here, @flaxio, since this is a piece of information you would have in a real life scenario, wouldn't you?",
      "votes": null
    },
    {
      "id": "1078919",
      "postDate": "11/15/2020 12:32:30",
      "content": "<p>I dont think we can say there are 7 eruptions in the dataset. We can say there are 7 samples that have a measured time to eruption less than 10 minutes, 39 samples with eruptions within 1 hour,  438 within 12 hours, 856 within 24 hours and so on…</p>\n<p>I understand that this is because the dataset we are working with has already been curated to create samples of X features with 60000 samples each with the corresponding to single value in Y target. So <a href=\"https://www.kaggle.com/flaxio\" target=\"_blank\">@flaxio</a> and the INGV team have already extracted the event information from the larger set of contious measurements. As such we can't be sure that these samples are even contiuous in time (i.e. sample 1 starts at 10:00 with Y = 1234, and sample 2 starts at 10:10 with Y = 1254). This is also why we see CV works and we are not using Time Series CV.</p>",
      "rawMarkdown": "I dont think we can say there are 7 eruptions in the dataset. We can say there are 7 samples that have a measured time to eruption less than 10 minutes, 39 samples with eruptions within 1 hour,  438 within 12 hours, 856 within 24 hours and so on...\n\nI understand that this is because the dataset we are working with has already been curated to create samples of X features with 60000 samples each with the corresponding to single value in Y target. So @flaxio and the INGV team have already extracted the event information from the larger set of contious measurements. As such we can't be sure that these samples are even contiuous in time (i.e. sample 1 starts at 10:00 with Y = 1234, and sample 2 starts at 10:10 with Y = 1254). This is also why we see CV works and we are not using Time Series CV.",
      "votes": null
    },
    {
      "id": "1100784",
      "postDate": "12/03/2020 10:38:52",
      "content": "<p>time_to_eruption is measured from the end of the data trace</p>",
      "rawMarkdown": "time_to_eruption is measured from the end of the data trace",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1066386,
      "author_name": "domizianostingi",
      "author_url": "",
      "post_date": "11/01/2020 17:33:38",
      "content": "<p>So at the end of the day, if my reasoning is right, there are 7 eruptions ?? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1078919,
          "author_name": "nicholasjhana",
          "author_url": "",
          "post_date": "11/15/2020 12:32:30",
          "content": "<p>I dont think we can say there are 7 eruptions in the dataset. We can say there are 7 samples that have a measured time to eruption less than 10 minutes, 39 samples with eruptions within 1 hour,  438 within 12 hours, 856 within 24 hours and so on…</p>\n<p>I understand that this is because the dataset we are working with has already been curated to create samples of X features with 60000 samples each with the corresponding to single value in Y target. So <a href=\"https://www.kaggle.com/flaxio\" target=\"_blank\">@flaxio</a> and the INGV team have already extracted the event information from the larger set of contious measurements. As such we can't be sure that these samples are even contiuous in time (i.e. sample 1 starts at 10:00 with Y = 1234, and sample 2 starts at 10:10 with Y = 1254). This is also why we see CV works and we are not using Time Series CV.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1078125,
      "author_name": "rhinopithecus",
      "author_url": "",
      "post_date": "11/14/2020 11:47:46",
      "content": "<p>I think I see your point. You are saying that there have been 7 eruptions, since there are 7 data traces where time_to_eruption is less than 10 minutes, correct?</p>\n<p>I am nevertheless not sure if this reasoning holds, because we are not given any information whether the time_to_eruption value is measured from the start of the data trace or from the end of the data trace. I.e. it could as well be the case that if time_to_eruption = 62.5 seconds, the eruption happens 62500 data samples (62.5 seconds) after the last sample of our trace.<br>\nAt least I haven't found any information about the reference of the time_to_eruption value. Have you?</p>\n<p>Could you give assistance here, <a href=\"https://www.kaggle.com/flaxio\" target=\"_blank\">@flaxio</a>, since this is a piece of information you would have in a real life scenario, wouldn't you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1100784,
          "author_name": "flaxio",
          "author_url": "",
          "post_date": "12/03/2020 10:38:52",
          "content": "<p>time_to_eruption is measured from the end of the data trace</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1066381": "I would like to know if this make sense. Let's take the first 3 value in the train.csv:\n```\nsegment_id  time_to_eruption\n1136037770 12262005\n1969647810 32739612\n1895879680 14965999\n```\nSince, each file represent 10 minuts of registration of the activity, does it make sense take, for each file, the time_to_eruption and try to understand if it is bigger or lower than 10 minuts, in this way it should be possible to apply some time to events statistical methods? \nFor example, let's assume that for the segment 1136037770 the correspondent time_to_eruption is lower than 10 minuts, for second segment_id it is bigger and so on... \nDoes it make sense this ?",
    "1066386": "So at the end of the day, if my reasoning is right, there are 7 eruptions ??",
    "1078125": "I think I see your point. You are saying that there have been 7 eruptions, since there are 7 data traces where time_to_eruption is less than 10 minutes, correct?\n\nI am nevertheless not sure if this reasoning holds, because we are not given any information whether the time_to_eruption value is measured from the start of the data trace or from the end of the data trace. I.e. it could as well be the case that if time_to_eruption = 62.5 seconds, the eruption happens 62500 data samples (62.5 seconds) after the last sample of our trace.\nAt least I haven't found any information about the reference of the time_to_eruption value. Have you?\n\nCould you give assistance here, @flaxio, since this is a piece of information you would have in a real life scenario, wouldn't you?",
    "1078919": "I dont think we can say there are 7 eruptions in the dataset. We can say there are 7 samples that have a measured time to eruption less than 10 minutes, 39 samples with eruptions within 1 hour,  438 within 12 hours, 856 within 24 hours and so on...\n\nI understand that this is because the dataset we are working with has already been curated to create samples of X features with 60000 samples each with the corresponding to single value in Y target. So @flaxio and the INGV team have already extracted the event information from the larger set of contious measurements. As such we can't be sure that these samples are even contiuous in time (i.e. sample 1 starts at 10:00 with Y = 1234, and sample 2 starts at 10:10 with Y = 1254). This is also why we see CV works and we are not using Time Series CV.",
    "1100784": "time_to_eruption is measured from the end of the data trace"
  },
  "source": "meta"
}