{
  "id": 361817,
  "title": "3D CNN networks",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/361817",
  "author_name": "",
  "post_date": "2022-10-23T22:47:05.167509400Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>3 years ago there was a competition called <a href=\"https://www.kaggle.com/competitions/nfl-big-data-bowl-2020\" target=\"_blank\">NFL Big Data Bowl 2020</a> where the participants were asked to predict the target like we have in our current TPS competition using the data of players movements in 2D space.</p>\n<p>The winning team \"The Zoo\" have created the <a href=\"https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119400\" target=\"_blank\">2D CNN based on relative location and speed features only</a>. Can we do the same here but in the 3D space? For train of course yes as we have all the players and ball movements, but for the test data during prediction time we obviously will receive some problems - we don't exactly know the <code>game_num</code>, <code>event_id</code> and <code>event_time</code> for the test rows to create the sequence of players and ball movements.</p>\n<p>I hope that some time in the future <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> and <a href=\"https://www.kaggle.com/dster\" target=\"_blank\">@dster</a> will add these columns to the current dataset and we can improve our scores using the deep convolutional neural networks. That will be a great place to try them out and push the current limits 🙃</p>\n<p>Alex</p>",
  "messages": [
    {
      "id": "2001297",
      "postDate": "10/23/2022 22:47:05",
      "content": "<p>3 years ago there was a competition called <a href=\"https://www.kaggle.com/competitions/nfl-big-data-bowl-2020\" target=\"_blank\">NFL Big Data Bowl 2020</a> where the participants were asked to predict the target like we have in our current TPS competition using the data of players movements in 2D space.</p>\n<p>The winning team \"The Zoo\" have created the <a href=\"https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119400\" target=\"_blank\">2D CNN based on relative location and speed features only</a>. Can we do the same here but in the 3D space? For train of course yes as we have all the players and ball movements, but for the test data during prediction time we obviously will receive some problems - we don't exactly know the <code>game_num</code>, <code>event_id</code> and <code>event_time</code> for the test rows to create the sequence of players and ball movements.</p>\n<p>I hope that some time in the future <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> and <a href=\"https://www.kaggle.com/dster\" target=\"_blank\">@dster</a> will add these columns to the current dataset and we can improve our scores using the deep convolutional neural networks. That will be a great place to try them out and push the current limits 🙃</p>\n<p>Alex</p>",
      "rawMarkdown": "3 years ago there was a competition called [NFL Big Data Bowl 2020](https://www.kaggle.com/competitions/nfl-big-data-bowl-2020) where the participants were asked to predict the target like we have in our current TPS competition using the data of players movements in 2D space.\n\nThe winning team \"The Zoo\" have created the [2D CNN based on relative location and speed features only](https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119400). Can we do the same here but in the 3D space? For train of course yes as we have all the players and ball movements, but for the test data during prediction time we obviously will receive some problems - we don't exactly know the `game_num`, `event_id` and `event_time` for the test rows to create the sequence of players and ball movements.\n\nI hope that some time in the future @inversion and @dster will add these columns to the current dataset and we can improve our scores using the deep convolutional neural networks. That will be a great place to try them out and push the current limits 🙃\n\nAlex",
      "votes": null
    },
    {
      "id": "2002147",
      "postDate": "10/24/2022 14:36:08",
      "content": "<p>Hey Alexander, unfortunately I'm not sure I understand what you're asking.  I believe the NFL competition did not have timeseries data at all (train or test) within the context of one play.  You don't need those three timeseries columns to build a model for this competition based on The Zoo's NFL model, and I've done so myself 🙂</p>\n<p>Here we have timeseries data in train but not in test.  We didn't include it in test for a few reasons:</p>\n<ol>\n<li>It would become easy to figure out future data in the test set and infer target values, so we'd have to use a custom Python module like in the NFL competition to access the test set only in sequential order, and that would be a lot of complexity for a Tabular Playground Series competition.</li>\n<li>It's nice to imagine this problem being Markovian and not needing historical data to make an accurate prediction.  I think this is a reasonable assumption, except unfortunately Rocket League replays AFAIK do not capture whether a player still has their \"flip\" move available, so advanced moves such as ceiling shots and flip resets cannot be accurately predicted without the previous several seconds of data, and even then only imperfectly.</li>\n</ol>\n<p>We did include timeseries data in the train set so that one could train on it and take advantage of the richer information there, eg with a RNN.  I imagine there are useful methods in that direction but that's a bit beyond my expertise so I could be wrong 🙃 It also allows for some nice EDAs as we've seen.</p>",
      "rawMarkdown": "Hey Alexander, unfortunately I'm not sure I understand what you're asking.  I believe the NFL competition did not have timeseries data at all (train or test) within the context of one play.  You don't need those three timeseries columns to build a model for this competition based on The Zoo's NFL model, and I've done so myself 🙂\n\nHere we have timeseries data in train but not in test.  We didn't include it in test for a few reasons:\n\n1. It would become easy to figure out future data in the test set and infer target values, so we'd have to use a custom Python module like in the NFL competition to access the test set only in sequential order, and that would be a lot of complexity for a Tabular Playground Series competition.\n2. It's nice to imagine this problem being Markovian and not needing historical data to make an accurate prediction.  I think this is a reasonable assumption, except unfortunately Rocket League replays AFAIK do not capture whether a player still has their \"flip\" move available, so advanced moves such as ceiling shots and flip resets cannot be accurately predicted without the previous several seconds of data, and even then only imperfectly.\n\nWe did include timeseries data in the train set so that one could train on it and take advantage of the richer information there, eg with a RNN.  I imagine there are useful methods in that direction but that's a bit beyond my expertise so I could be wrong 🙃 It also allows for some nice EDAs as we've seen.",
      "votes": null
    },
    {
      "id": "2024487",
      "postDate": "11/10/2022 15:05:30",
      "content": "<p>Hi Alexander,</p>\n<p>I've already been here (before this comment).  And now, after reading DJ's talking about the Markovian problem I thought I'm so screwed. The only thing I was able to make was a naive, stupid meme Notebook (as always).  Though I enjoy the crazy, little things I publish. The important (in my case) is to participate and keep learning. Thanks to the generosity of users like you.</p>\n<p>You're doing a substantial, uncomparable work on TPS competitions publishing your high-quality material (code and topics).<br>\nNot every top Kaggler publish with that regularity.</p>",
      "rawMarkdown": "Hi Alexander,\n\nI've already been here (before this comment).  And now, after reading DJ's talking about the Markovian problem I thought I'm so screwed. The only thing I was able to make was a naive, stupid meme Notebook (as always).  Though I enjoy the crazy, little things I publish. The important (in my case) is to participate and keep learning. Thanks to the generosity of users like you.\n\nYou're doing a substantial, uncomparable work on TPS competitions publishing your high-quality material (code and topics).\nNot every top Kaggler publish with that regularity.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2002147,
      "author_name": "dster",
      "author_url": "",
      "post_date": "10/24/2022 14:36:08",
      "content": "<p>Hey Alexander, unfortunately I'm not sure I understand what you're asking.  I believe the NFL competition did not have timeseries data at all (train or test) within the context of one play.  You don't need those three timeseries columns to build a model for this competition based on The Zoo's NFL model, and I've done so myself 🙂</p>\n<p>Here we have timeseries data in train but not in test.  We didn't include it in test for a few reasons:</p>\n<ol>\n<li>It would become easy to figure out future data in the test set and infer target values, so we'd have to use a custom Python module like in the NFL competition to access the test set only in sequential order, and that would be a lot of complexity for a Tabular Playground Series competition.</li>\n<li>It's nice to imagine this problem being Markovian and not needing historical data to make an accurate prediction.  I think this is a reasonable assumption, except unfortunately Rocket League replays AFAIK do not capture whether a player still has their \"flip\" move available, so advanced moves such as ceiling shots and flip resets cannot be accurately predicted without the previous several seconds of data, and even then only imperfectly.</li>\n</ol>\n<p>We did include timeseries data in the train set so that one could train on it and take advantage of the richer information there, eg with a RNN.  I imagine there are useful methods in that direction but that's a bit beyond my expertise so I could be wrong 🙃 It also allows for some nice EDAs as we've seen.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2024487,
      "author_name": "mpwolke",
      "author_url": "",
      "post_date": "11/10/2022 15:05:30",
      "content": "<p>Hi Alexander,</p>\n<p>I've already been here (before this comment).  And now, after reading DJ's talking about the Markovian problem I thought I'm so screwed. The only thing I was able to make was a naive, stupid meme Notebook (as always).  Though I enjoy the crazy, little things I publish. The important (in my case) is to participate and keep learning. Thanks to the generosity of users like you.</p>\n<p>You're doing a substantial, uncomparable work on TPS competitions publishing your high-quality material (code and topics).<br>\nNot every top Kaggler publish with that regularity.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2001297": "3 years ago there was a competition called [NFL Big Data Bowl 2020](https://www.kaggle.com/competitions/nfl-big-data-bowl-2020) where the participants were asked to predict the target like we have in our current TPS competition using the data of players movements in 2D space.\n\nThe winning team \"The Zoo\" have created the [2D CNN based on relative location and speed features only](https://www.kaggle.com/competitions/nfl-big-data-bowl-2020/discussion/119400). Can we do the same here but in the 3D space? For train of course yes as we have all the players and ball movements, but for the test data during prediction time we obviously will receive some problems - we don't exactly know the `game_num`, `event_id` and `event_time` for the test rows to create the sequence of players and ball movements.\n\nI hope that some time in the future @inversion and @dster will add these columns to the current dataset and we can improve our scores using the deep convolutional neural networks. That will be a great place to try them out and push the current limits 🙃\n\nAlex",
    "2002147": "Hey Alexander, unfortunately I'm not sure I understand what you're asking.  I believe the NFL competition did not have timeseries data at all (train or test) within the context of one play.  You don't need those three timeseries columns to build a model for this competition based on The Zoo's NFL model, and I've done so myself 🙂\n\nHere we have timeseries data in train but not in test.  We didn't include it in test for a few reasons:\n\n1. It would become easy to figure out future data in the test set and infer target values, so we'd have to use a custom Python module like in the NFL competition to access the test set only in sequential order, and that would be a lot of complexity for a Tabular Playground Series competition.\n2. It's nice to imagine this problem being Markovian and not needing historical data to make an accurate prediction.  I think this is a reasonable assumption, except unfortunately Rocket League replays AFAIK do not capture whether a player still has their \"flip\" move available, so advanced moves such as ceiling shots and flip resets cannot be accurately predicted without the previous several seconds of data, and even then only imperfectly.\n\nWe did include timeseries data in the train set so that one could train on it and take advantage of the richer information there, eg with a RNN.  I imagine there are useful methods in that direction but that's a bit beyond my expertise so I could be wrong 🙃 It also allows for some nice EDAs as we've seen.",
    "2024487": "Hi Alexander,\n\nI've already been here (before this comment).  And now, after reading DJ's talking about the Markovian problem I thought I'm so screwed. The only thing I was able to make was a naive, stupid meme Notebook (as always).  Though I enjoy the crazy, little things I publish. The important (in my case) is to participate and keep learning. Thanks to the generosity of users like you.\n\nYou're doing a substantial, uncomparable work on TPS competitions publishing your high-quality material (code and topics).\nNot every top Kaggler publish with that regularity."
  },
  "source": "meta"
}