{
  "id": 358021,
  "title": "timeseries",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/358021",
  "author_name": "",
  "post_date": "2022-10-06T11:03:32.727979400Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>i thought about tackling this competition as a timeseries problem with the events as the sequences.<br>\nThe problem here is that the test dataset doesn't have the event time column. (here i though about trying to cluster the events)</p>\n<p>What do you guys think?</p>",
  "messages": [
    {
      "id": "1974613",
      "postDate": "10/06/2022 11:03:32",
      "content": "<p>i thought about tackling this competition as a timeseries problem with the events as the sequences.<br>\nThe problem here is that the test dataset doesn't have the event time column. (here i though about trying to cluster the events)</p>\n<p>What do you guys think?</p>",
      "rawMarkdown": "i thought about tackling this competition as a timeseries problem with the events as the sequences.\nThe problem here is that the test dataset doesn't have the event time column. (here i though about trying to cluster the events)\n\nWhat do you guys think?",
      "votes": null
    },
    {
      "id": "1974690",
      "postDate": "10/06/2022 11:49:23",
      "content": "<p>HI <a href=\"https://www.kaggle.com/bayremabdellaoui\" target=\"_blank\">@bayremabdellaoui</a> , I think that while difficult the problem can be approached as a timeseries problem, in the sense that you learn to predict how the game evolves in the next 10 seconds and use this prediction to see if there was a goal or not.</p>\n<p>Approached in this way the competition is really hard and getting a correct prediction for test set given only the starting step is hard but can give insights that other models can't catch.</p>\n<p>Regarding clustering I think it is an interesting idea, but I give you 2 things that you should check out:</p>\n<ol>\n<li>2 games may result far apart to a clustering algorithm while being identical, just player 1 and 2 are in swapped positions, this can be solved probably by ordering players according to the distance from a given point ( use angle if there are ties)</li>\n<li>distance on z axis may have a different impact compared to distances on x and y axis,  clustering algorithms based on euclidean distance may not be the best to use.</li>\n</ol>\n<p>I hope to see soon your approach to this challenge using clustering or timeseries analysis!</p>",
      "rawMarkdown": "HI @bayremabdellaoui , I think that while difficult the problem can be approached as a timeseries problem, in the sense that you learn to predict how the game evolves in the next 10 seconds and use this prediction to see if there was a goal or not.\n\nApproached in this way the competition is really hard and getting a correct prediction for test set given only the starting step is hard but can give insights that other models can't catch.\n\nRegarding clustering I think it is an interesting idea, but I give you 2 things that you should check out:\n1. 2 games may result far apart to a clustering algorithm while being identical, just player 1 and 2 are in swapped positions, this can be solved probably by ordering players according to the distance from a given point ( use angle if there are ties)\n2. distance on z axis may have a different impact compared to distances on x and y axis,  clustering algorithms based on euclidean distance may not be the best to use.\n\nI hope to see soon your approach to this challenge using clustering or timeseries analysis!",
      "votes": null
    },
    {
      "id": "1974765",
      "postDate": "10/06/2022 12:46:03",
      "content": "<p>I have been trying to cluster the test data and unfortunately without any good results to show for. I was using only the positions, and boost timers(easy to forecast, can help differentiate between close examples of different matches). The 90th percentile distance between two consecutive frames in the training data was around 27(using RMS/Euclidean distance as a metric), hence I used that as a threshold.</p>\n<p>Applying agglomerative clustering directly gave very bad results(over 300k tiny clusters, not very useful), but calculating the next state using the velocities etc. did some better. About 100k clusters, average cluster size was about 5. I'm trying to figure out collisions, I'm assuming if I can get them right, then it'll help calculate the next possible state more precisely.</p>\n<p>The problem is that the balls velocity after a collision is very dependent on un-knowable variables like spin of the ball, angle between the point of contact and the center of masses of the car and the ball, both of which we have no idea about. Second, neither do the the players disappear right at impact and nor does the ball reset right after being scored, hence making forecasting them even more difficult as it's impossible to know how many exact frames later this is processed. The ball specially, hangs around a lot longer. I have a few tunings and approaches left to try, then I'll publish the notebook. Hopefully someone can pick something up, or maybe find a bug causing the poor performance.</p>",
      "rawMarkdown": "I have been trying to cluster the test data and unfortunately without any good results to show for. I was using only the positions, and boost timers(easy to forecast, can help differentiate between close examples of different matches). The 90th percentile distance between two consecutive frames in the training data was around 27(using RMS/Euclidean distance as a metric), hence I used that as a threshold.\n\nApplying agglomerative clustering directly gave very bad results(over 300k tiny clusters, not very useful), but calculating the next state using the velocities etc. did some better. About 100k clusters, average cluster size was about 5. I'm trying to figure out collisions, I'm assuming if I can get them right, then it'll help calculate the next possible state more precisely.\n\nThe problem is that the balls velocity after a collision is very dependent on un-knowable variables like spin of the ball, angle between the point of contact and the center of masses of the car and the ball, both of which we have no idea about. Second, neither do the the players disappear right at impact and nor does the ball reset right after being scored, hence making forecasting them even more difficult as it's impossible to know how many exact frames later this is processed. The ball specially, hangs around a lot longer. I have a few tunings and approaches left to try, then I'll publish the notebook. Hopefully someone can pick something up, or maybe find a bug causing the poor performance.",
      "votes": null
    },
    {
      "id": "1975151",
      "postDate": "10/06/2022 16:14:23",
      "content": "<p>These look like different approaches! I will be eager to look at the results, especially with the test data structure. Good luck <a href=\"https://www.kaggle.com/bayremabdellaoui\" target=\"_blank\">@bayremabdellaoui</a>!!</p>",
      "rawMarkdown": "These look like different approaches! I will be eager to look at the results, especially with the test data structure. Good luck @bayremabdellaoui!!",
      "votes": null
    },
    {
      "id": "1977360",
      "postDate": "10/08/2022 01:22:03",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/bayremabdellaoui\" target=\"_blank\">@bayremabdellaoui</a>, I have a hard time understanding how it's possible to utilize time series; at the end of the day, the data that you will be predicting has only one value per set of events, and there is no way to identify events previous to the row that you will be working, maybe there is a way to identify time events on the test dataset.</p>",
      "rawMarkdown": "Hello @bayremabdellaoui, I have a hard time understanding how it's possible to utilize time series; at the end of the day, the data that you will be predicting has only one value per set of events, and there is no way to identify events previous to the row that you will be working, maybe there is a way to identify time events on the test dataset.",
      "votes": null
    },
    {
      "id": "1978271",
      "postDate": "10/08/2022 16:48:53",
      "content": "<p>Update: Published the notebook, <a href=\"https://www.kaggle.com/code/aatiffraz/episodification-of-the-test-dataset?scriptVersionId=107489607\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Update: Published the notebook, [here](https://www.kaggle.com/code/aatiffraz/episodification-of-the-test-dataset?scriptVersionId=107489607).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1974690,
      "author_name": "pietromaldini1",
      "author_url": "",
      "post_date": "10/06/2022 11:49:23",
      "content": "<p>HI <a href=\"https://www.kaggle.com/bayremabdellaoui\" target=\"_blank\">@bayremabdellaoui</a> , I think that while difficult the problem can be approached as a timeseries problem, in the sense that you learn to predict how the game evolves in the next 10 seconds and use this prediction to see if there was a goal or not.</p>\n<p>Approached in this way the competition is really hard and getting a correct prediction for test set given only the starting step is hard but can give insights that other models can't catch.</p>\n<p>Regarding clustering I think it is an interesting idea, but I give you 2 things that you should check out:</p>\n<ol>\n<li>2 games may result far apart to a clustering algorithm while being identical, just player 1 and 2 are in swapped positions, this can be solved probably by ordering players according to the distance from a given point ( use angle if there are ties)</li>\n<li>distance on z axis may have a different impact compared to distances on x and y axis,  clustering algorithms based on euclidean distance may not be the best to use.</li>\n</ol>\n<p>I hope to see soon your approach to this challenge using clustering or timeseries analysis!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1974765,
          "author_name": "aatiffraz",
          "author_url": "",
          "post_date": "10/06/2022 12:46:03",
          "content": "<p>I have been trying to cluster the test data and unfortunately without any good results to show for. I was using only the positions, and boost timers(easy to forecast, can help differentiate between close examples of different matches). The 90th percentile distance between two consecutive frames in the training data was around 27(using RMS/Euclidean distance as a metric), hence I used that as a threshold.</p>\n<p>Applying agglomerative clustering directly gave very bad results(over 300k tiny clusters, not very useful), but calculating the next state using the velocities etc. did some better. About 100k clusters, average cluster size was about 5. I'm trying to figure out collisions, I'm assuming if I can get them right, then it'll help calculate the next possible state more precisely.</p>\n<p>The problem is that the balls velocity after a collision is very dependent on un-knowable variables like spin of the ball, angle between the point of contact and the center of masses of the car and the ball, both of which we have no idea about. Second, neither do the the players disappear right at impact and nor does the ball reset right after being scored, hence making forecasting them even more difficult as it's impossible to know how many exact frames later this is processed. The ball specially, hangs around a lot longer. I have a few tunings and approaches left to try, then I'll publish the notebook. Hopefully someone can pick something up, or maybe find a bug causing the poor performance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1978271,
          "author_name": "aatiffraz",
          "author_url": "",
          "post_date": "10/08/2022 16:48:53",
          "content": "<p>Update: Published the notebook, <a href=\"https://www.kaggle.com/code/aatiffraz/episodification-of-the-test-dataset?scriptVersionId=107489607\" target=\"_blank\">here</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1975151,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "10/06/2022 16:14:23",
      "content": "<p>These look like different approaches! I will be eager to look at the results, especially with the test data structure. Good luck <a href=\"https://www.kaggle.com/bayremabdellaoui\" target=\"_blank\">@bayremabdellaoui</a>!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1977360,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "10/08/2022 01:22:03",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/bayremabdellaoui\" target=\"_blank\">@bayremabdellaoui</a>, I have a hard time understanding how it's possible to utilize time series; at the end of the day, the data that you will be predicting has only one value per set of events, and there is no way to identify events previous to the row that you will be working, maybe there is a way to identify time events on the test dataset.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1974613": "i thought about tackling this competition as a timeseries problem with the events as the sequences.\nThe problem here is that the test dataset doesn't have the event time column. (here i though about trying to cluster the events)\n\nWhat do you guys think?",
    "1974690": "HI @bayremabdellaoui , I think that while difficult the problem can be approached as a timeseries problem, in the sense that you learn to predict how the game evolves in the next 10 seconds and use this prediction to see if there was a goal or not.\n\nApproached in this way the competition is really hard and getting a correct prediction for test set given only the starting step is hard but can give insights that other models can't catch.\n\nRegarding clustering I think it is an interesting idea, but I give you 2 things that you should check out:\n1. 2 games may result far apart to a clustering algorithm while being identical, just player 1 and 2 are in swapped positions, this can be solved probably by ordering players according to the distance from a given point ( use angle if there are ties)\n2. distance on z axis may have a different impact compared to distances on x and y axis,  clustering algorithms based on euclidean distance may not be the best to use.\n\nI hope to see soon your approach to this challenge using clustering or timeseries analysis!",
    "1974765": "I have been trying to cluster the test data and unfortunately without any good results to show for. I was using only the positions, and boost timers(easy to forecast, can help differentiate between close examples of different matches). The 90th percentile distance between two consecutive frames in the training data was around 27(using RMS/Euclidean distance as a metric), hence I used that as a threshold.\n\nApplying agglomerative clustering directly gave very bad results(over 300k tiny clusters, not very useful), but calculating the next state using the velocities etc. did some better. About 100k clusters, average cluster size was about 5. I'm trying to figure out collisions, I'm assuming if I can get them right, then it'll help calculate the next possible state more precisely.\n\nThe problem is that the balls velocity after a collision is very dependent on un-knowable variables like spin of the ball, angle between the point of contact and the center of masses of the car and the ball, both of which we have no idea about. Second, neither do the the players disappear right at impact and nor does the ball reset right after being scored, hence making forecasting them even more difficult as it's impossible to know how many exact frames later this is processed. The ball specially, hangs around a lot longer. I have a few tunings and approaches left to try, then I'll publish the notebook. Hopefully someone can pick something up, or maybe find a bug causing the poor performance.",
    "1975151": "These look like different approaches! I will be eager to look at the results, especially with the test data structure. Good luck @bayremabdellaoui!!",
    "1977360": "Hello @bayremabdellaoui, I have a hard time understanding how it's possible to utilize time series; at the end of the day, the data that you will be predicting has only one value per set of events, and there is no way to identify events previous to the row that you will be working, maybe there is a way to identify time events on the test dataset.",
    "1978271": "Update: Published the notebook, [here](https://www.kaggle.com/code/aatiffraz/episodification-of-the-test-dataset?scriptVersionId=107489607)."
  },
  "source": "meta"
}