{
  "id": 215925,
  "title": "Test set timestamps",
  "url": "/competitions/indoor-location-navigation/discussion/215925",
  "author_name": "",
  "post_date": "2021-01-31T21:13:02.903592100Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi guys,</p>\n<p>Thanks for all the useful notebooks already published here and there, it's been really useful to understand the data.</p>\n<p>I am quite confused about how we are supposed to work with timestamps.<br>\nI understand that the different features are measured at different timestamps and frequencies. The way point are split by seconds (from ~5sec to ~15seconds) while gyroscope points for example are split by 20 milliseconds.</p>\n<p>I assume that modifying the raw timeseries to have matching timestamps is one of the challenge of this competition, which is something that can be dealt with.</p>\n<p>I am struggling however to understand on which timestamps we must predict the way points. On the training set we are given with the TYPE_WAYPOINT the timestamp of measure. We can't find that information on the test set (which is logical as it is the feature we are trying to predict), however we thus have no idea of the timestamps we need to consider for predictions.</p>\n<p>The only place where we can find the information of the test timestamps to consider is on the sample_submission.csv. Which means we need to look up the timestamp in the file (called sample) for each trace just to get the timestamps at which we need to predict the x y position.</p>\n<p>Although it seems feasible it sounds a bit weird to me. </p>\n<blockquote>\n  <p>Are the timestamps in the sample_submission.csv the real timestamps we need to used or are there just examples ?<br>\n  Are we to assume that we need to be able to predict prediction for every single timestep ? (20ms delta step for ex) or shall we expect similar deltas than the sample_submission.csv file (from ~5sec to ~15seconds).</p>\n</blockquote>\n<p>Regarding predictions using future. It seems that we can definitely do that, can't see any technical limits with the current provided data. I am not sure that is the point of the challenge though, as sampling frequency is pretty high and use-cases are mostly live notifications.</p>\n<p>Thanks a bunch in advance and good luck yall :)</p>",
  "messages": [
    {
      "id": "1179821",
      "postDate": "01/31/2021 21:13:02",
      "content": "<p>Hi guys,</p>\n<p>Thanks for all the useful notebooks already published here and there, it's been really useful to understand the data.</p>\n<p>I am quite confused about how we are supposed to work with timestamps.<br>\nI understand that the different features are measured at different timestamps and frequencies. The way point are split by seconds (from ~5sec to ~15seconds) while gyroscope points for example are split by 20 milliseconds.</p>\n<p>I assume that modifying the raw timeseries to have matching timestamps is one of the challenge of this competition, which is something that can be dealt with.</p>\n<p>I am struggling however to understand on which timestamps we must predict the way points. On the training set we are given with the TYPE_WAYPOINT the timestamp of measure. We can't find that information on the test set (which is logical as it is the feature we are trying to predict), however we thus have no idea of the timestamps we need to consider for predictions.</p>\n<p>The only place where we can find the information of the test timestamps to consider is on the sample_submission.csv. Which means we need to look up the timestamp in the file (called sample) for each trace just to get the timestamps at which we need to predict the x y position.</p>\n<p>Although it seems feasible it sounds a bit weird to me. </p>\n<blockquote>\n  <p>Are the timestamps in the sample_submission.csv the real timestamps we need to used or are there just examples ?<br>\n  Are we to assume that we need to be able to predict prediction for every single timestep ? (20ms delta step for ex) or shall we expect similar deltas than the sample_submission.csv file (from ~5sec to ~15seconds).</p>\n</blockquote>\n<p>Regarding predictions using future. It seems that we can definitely do that, can't see any technical limits with the current provided data. I am not sure that is the point of the challenge though, as sampling frequency is pretty high and use-cases are mostly live notifications.</p>\n<p>Thanks a bunch in advance and good luck yall :)</p>",
      "rawMarkdown": "Hi guys,\n\nThanks for all the useful notebooks already published here and there, it's been really useful to understand the data.\n\nI am quite confused about how we are supposed to work with timestamps.\nI understand that the different features are measured at different timestamps and frequencies. The way point are split by seconds (from ~5sec to ~15seconds) while gyroscope points for example are split by 20 milliseconds.\n\nI assume that modifying the raw timeseries to have matching timestamps is one of the challenge of this competition, which is something that can be dealt with.\n\nI am struggling however to understand on which timestamps we must predict the way points. On the training set we are given with the TYPE_WAYPOINT the timestamp of measure. We can't find that information on the test set (which is logical as it is the feature we are trying to predict), however we thus have no idea of the timestamps we need to consider for predictions.\n\nThe only place where we can find the information of the test timestamps to consider is on the sample_submission.csv. Which means we need to look up the timestamp in the file (called sample) for each trace just to get the timestamps at which we need to predict the x y position.\n\nAlthough it seems feasible it sounds a bit weird to me. \n\n> Are the timestamps in the sample_submission.csv the real timestamps we need to used or are there just examples ?\n> Are we to assume that we need to be able to predict prediction for every single timestep ? (20ms delta step for ex) or shall we expect similar deltas than the sample_submission.csv file (from ~5sec to ~15seconds).\n\n\nRegarding predictions using future. It seems that we can definitely do that, can't see any technical limits with the current provided data. I am not sure that is the point of the challenge though, as sampling frequency is pretty high and use-cases are mostly live notifications.\n\n\nThanks a bunch in advance and good luck yall :)",
      "votes": null
    },
    {
      "id": "1255768",
      "postDate": "03/29/2021 07:15:21",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/rafaelcartenet\" target=\"_blank\">@rafaelcartenet</a> for raising this point. In my opinion, this is a subtle but critical part of the competition. As far as different frequencies of different sensors, I guess that is the practical way it works. In the conventional approach, you have different mapping functions for different sensor data and the particular function is called when a new data from the sensor is got. </p>\n<p>I suppose, the competition metric is based on the predictions for the timestamps in the sample submission file only. <br>\nOne solution is you can build the path for the entire duration and fit some polynomial curve and sample at the required timestamps. </p>",
      "rawMarkdown": "Thanks @rafaelcartenet for raising this point. In my opinion, this is a subtle but critical part of the competition. As far as different frequencies of different sensors, I guess that is the practical way it works. In the conventional approach, you have different mapping functions for different sensor data and the particular function is called when a new data from the sensor is got. \n\nI suppose, the competition metric is based on the predictions for the timestamps in the sample submission file only. \nOne solution is you can build the path for the entire duration and fit some polynomial curve and sample at the required timestamps.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1255768,
      "author_name": "suryajrrafl",
      "author_url": "",
      "post_date": "03/29/2021 07:15:21",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/rafaelcartenet\" target=\"_blank\">@rafaelcartenet</a> for raising this point. In my opinion, this is a subtle but critical part of the competition. As far as different frequencies of different sensors, I guess that is the practical way it works. In the conventional approach, you have different mapping functions for different sensor data and the particular function is called when a new data from the sensor is got. </p>\n<p>I suppose, the competition metric is based on the predictions for the timestamps in the sample submission file only. <br>\nOne solution is you can build the path for the entire duration and fit some polynomial curve and sample at the required timestamps. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1179821": "Hi guys,\n\nThanks for all the useful notebooks already published here and there, it's been really useful to understand the data.\n\nI am quite confused about how we are supposed to work with timestamps.\nI understand that the different features are measured at different timestamps and frequencies. The way point are split by seconds (from ~5sec to ~15seconds) while gyroscope points for example are split by 20 milliseconds.\n\nI assume that modifying the raw timeseries to have matching timestamps is one of the challenge of this competition, which is something that can be dealt with.\n\nI am struggling however to understand on which timestamps we must predict the way points. On the training set we are given with the TYPE_WAYPOINT the timestamp of measure. We can't find that information on the test set (which is logical as it is the feature we are trying to predict), however we thus have no idea of the timestamps we need to consider for predictions.\n\nThe only place where we can find the information of the test timestamps to consider is on the sample_submission.csv. Which means we need to look up the timestamp in the file (called sample) for each trace just to get the timestamps at which we need to predict the x y position.\n\nAlthough it seems feasible it sounds a bit weird to me. \n\n> Are the timestamps in the sample_submission.csv the real timestamps we need to used or are there just examples ?\n> Are we to assume that we need to be able to predict prediction for every single timestep ? (20ms delta step for ex) or shall we expect similar deltas than the sample_submission.csv file (from ~5sec to ~15seconds).\n\n\nRegarding predictions using future. It seems that we can definitely do that, can't see any technical limits with the current provided data. I am not sure that is the point of the challenge though, as sampling frequency is pretty high and use-cases are mostly live notifications.\n\n\nThanks a bunch in advance and good luck yall :)",
    "1255768": "Thanks @rafaelcartenet for raising this point. In my opinion, this is a subtle but critical part of the competition. As far as different frequencies of different sensors, I guess that is the practical way it works. In the conventional approach, you have different mapping functions for different sensor data and the particular function is called when a new data from the sensor is got. \n\nI suppose, the competition metric is based on the predictions for the timestamps in the sample submission file only. \nOne solution is you can build the path for the entire duration and fit some polynomial curve and sample at the required timestamps."
  },
  "source": "meta"
}