{
  "id": 358590,
  "title": "📔 Things learned on the first week - Second week plan for beginners 💡",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/358590",
  "author_name": "Pastor Soto",
  "post_date": "2022-10-08T15:56:53.673000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>The first week was an exploratory week to understand the competition, the data, the initial suggestions and the game itself.</p>\n<p>My previous week plan was:</p>\n<ol>\n<li>Understand online learning</li>\n<li>Check the discussion forums for ideas</li>\n<li>Understand the game</li>\n<li>Check NFL notebooks</li>\n</ol>\n<p>Let's go over one by one with the findings:</p>\n<h3>Understand online learning</h3>\n<p>Online learning is a method that allow to make predictions in live environment, it's a useful tool with a lot of applications such as stock market, sports events, weather and many others. It consumes less memory than batch prediction and tools like <a href=\"https://riverml.xyz/0.13.0/examples/batch-to-online/\" target=\"_blank\">river</a> mimics the sklearn syntax which make it really easy to implement. </p>\n<p>The cons, is that the resources are limited, and I couldn't find an example in a kaggle competition, also the performance is lower than batch prediction according to the river documentation. How to handle missing values can also be a problem if you perform some type of calculation on historical data.</p>\n<p>I am also not sure how does it work internally, maybe is a loop over every row in the model, if so it sounds computational expensive, and maybe a better approach will be using a sample of the training data with a batch prediction. I saw some results here that might support that path. </p>\n<h3>Check discussion forum for ideas</h3>\n<p>Key ideas I found were:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356540\" target=\"_blank\">Reduce the size of the dataset</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357158\" target=\"_blank\">Handle missing values</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356584\" target=\"_blank\">Online learning resources</a></li>\n<li>Some visualizations</li>\n</ul>\n<blockquote>\n  <p>Missing values needs to be explored a little further, which could start by checking similar approaches for other competitions, what happened when a player leaves the field? My guess is that we care to place in a position when the other team increase their chances to score and the team of the player decrease their chances to score. We might want to find this position as a naive imputation. </p>\n  <p>I need to explore more visualization, using things such as the one use in other competitions and in football. </p>\n</blockquote>\n<h3>Understand the game</h3>\n<p>There are several discussions that allow to understand the game, and watching some minutes can also give you a better intuition on what's going on. </p>\n<p>The key things that I found were:</p>\n<ul>\n<li>The game is 3 vs 3</li>\n<li>The game is similar to football, where the objective is to score a ball in the opposite field</li>\n<li>Players can be destroyed by other player for 3 seconds and reaper in their own field</li>\n</ul>\n<p>Things I still don't know:</p>\n<ul>\n<li>Why each game has more than one event?</li>\n<li>What's the size of the field?</li>\n</ul>\n<h3>Check NFL notebook</h3>\n<ul>\n<li>The winning solution was a neural network</li>\n<li>They simplified the game by using a runner and defenders and remove teammates</li>\n<li>The size of the dataset was significantly lower that this competition</li>\n<li>They had to optimize their memory to came up with the winning solution</li>\n</ul>\n<p>Things I still don't know:<br>\nHow do they handle missing values?</p>\n<h3>Second week plan</h3>\n<p><strong>Find a way to handle missing values</strong></p>\n<ul>\n<li>Check the NFL notebook to see how did they handle those values</li>\n<li>Check discussion forum and notebooks to see how people are handling missing values</li>\n<li>Understand the field and the position for, remove players that make more sense (lower probability of their team and increase the probability of the other team)</li>\n</ul>\n<p><strong>Create a visualization notebook using football approaches</strong></p>\n<ul>\n<li>Check this <a href=\"https://www.youtube.com/channel/UCUBFJYcag8j2rm_9HkrrA7w\" target=\"_blank\">YouTube channel</a> that has tracking approaches for football to draw a field and applied speed and position of the players</li>\n<li>Check the discussion forum and notebooks to see EDA ideas</li>\n<li>Create an EDA notebook that allow to understand the features and came up with a modeling plan</li>\n<li>From the EDA perform some feature engineering on the dataset</li>\n</ul>\n<p><strong>Prepare the third week plan</strong></p>\n<ul>\n<li>Create baseline model</li>\n</ul>",
  "messages": [
    {
      "id": 1978221,
      "postDate": "2022-10-08T15:56:53.673Z",
      "content": "<p>The first week was an exploratory week to understand the competition, the data, the initial suggestions and the game itself.</p>\n<p>My previous week plan was:</p>\n<ol>\n<li>Understand online learning</li>\n<li>Check the discussion forums for ideas</li>\n<li>Understand the game</li>\n<li>Check NFL notebooks</li>\n</ol>\n<p>Let's go over one by one with the findings:</p>\n<h3>Understand online learning</h3>\n<p>Online learning is a method that allow to make predictions in live environment, it's a useful tool with a lot of applications such as stock market, sports events, weather and many others. It consumes less memory than batch prediction and tools like <a href=\"https://riverml.xyz/0.13.0/examples/batch-to-online/\" target=\"_blank\">river</a> mimics the sklearn syntax which make it really easy to implement. </p>\n<p>The cons, is that the resources are limited, and I couldn't find an example in a kaggle competition, also the performance is lower than batch prediction according to the river documentation. How to handle missing values can also be a problem if you perform some type of calculation on historical data.</p>\n<p>I am also not sure how does it work internally, maybe is a loop over every row in the model, if so it sounds computational expensive, and maybe a better approach will be using a sample of the training data with a batch prediction. I saw some results here that might support that path. </p>\n<h3>Check discussion forum for ideas</h3>\n<p>Key ideas I found were:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356540\" target=\"_blank\">Reduce the size of the dataset</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357158\" target=\"_blank\">Handle missing values</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356584\" target=\"_blank\">Online learning resources</a></li>\n<li>Some visualizations</li>\n</ul>\n<blockquote>\n  <p>Missing values needs to be explored a little further, which could start by checking similar approaches for other competitions, what happened when a player leaves the field? My guess is that we care to place in a position when the other team increase their chances to score and the team of the player decrease their chances to score. We might want to find this position as a naive imputation. </p>\n  <p>I need to explore more visualization, using things such as the one use in other competitions and in football. </p>\n</blockquote>\n<h3>Understand the game</h3>\n<p>There are several discussions that allow to understand the game, and watching some minutes can also give you a better intuition on what's going on. </p>\n<p>The key things that I found were:</p>\n<ul>\n<li>The game is 3 vs 3</li>\n<li>The game is similar to football, where the objective is to score a ball in the opposite field</li>\n<li>Players can be destroyed by other player for 3 seconds and reaper in their own field</li>\n</ul>\n<p>Things I still don't know:</p>\n<ul>\n<li>Why each game has more than one event?</li>\n<li>What's the size of the field?</li>\n</ul>\n<h3>Check NFL notebook</h3>\n<ul>\n<li>The winning solution was a neural network</li>\n<li>They simplified the game by using a runner and defenders and remove teammates</li>\n<li>The size of the dataset was significantly lower that this competition</li>\n<li>They had to optimize their memory to came up with the winning solution</li>\n</ul>\n<p>Things I still don't know:<br>\nHow do they handle missing values?</p>\n<h3>Second week plan</h3>\n<p><strong>Find a way to handle missing values</strong></p>\n<ul>\n<li>Check the NFL notebook to see how did they handle those values</li>\n<li>Check discussion forum and notebooks to see how people are handling missing values</li>\n<li>Understand the field and the position for, remove players that make more sense (lower probability of their team and increase the probability of the other team)</li>\n</ul>\n<p><strong>Create a visualization notebook using football approaches</strong></p>\n<ul>\n<li>Check this <a href=\"https://www.youtube.com/channel/UCUBFJYcag8j2rm_9HkrrA7w\" target=\"_blank\">YouTube channel</a> that has tracking approaches for football to draw a field and applied speed and position of the players</li>\n<li>Check the discussion forum and notebooks to see EDA ideas</li>\n<li>Create an EDA notebook that allow to understand the features and came up with a modeling plan</li>\n<li>From the EDA perform some feature engineering on the dataset</li>\n</ul>\n<p><strong>Prepare the third week plan</strong></p>\n<ul>\n<li>Create baseline model</li>\n</ul>",
      "rawMarkdown": "The first week was an exploratory week to understand the competition, the data, the initial suggestions and the game itself.\n\nMy previous week plan was:\n\n1. Understand online learning\n2. Check the discussion forums for ideas\n3. Understand the game\n4. Check NFL notebooks\n\nLet's go over one by one with the findings:\n\n### Understand online learning\n\nOnline learning is a method that allow to make predictions in live environment, it's a useful tool with a lot of applications such as stock market, sports events, weather and many others. It consumes less memory than batch prediction and tools like [river](https://riverml.xyz/0.13.0/examples/batch-to-online/) mimics the sklearn syntax which make it really easy to implement. \n\nThe cons, is that the resources are limited, and I couldn't find an example in a kaggle competition, also the performance is lower than batch prediction according to the river documentation. How to handle missing values can also be a problem if you perform some type of calculation on historical data.\n\nI am also not sure how does it work internally, maybe is a loop over every row in the model, if so it sounds computational expensive, and maybe a better approach will be using a sample of the training data with a batch prediction. I saw some results here that might support that path. \n\n### Check discussion forum for ideas\n\nKey ideas I found were:\n\n- [Reduce the size of the dataset](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356540)\n- [Handle missing values](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357158)\n- [Online learning resources](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356584)\n- Some visualizations\n\n> Missing values needs to be explored a little further, which could start by checking similar approaches for other competitions, what happened when a player leaves the field? My guess is that we care to place in a position when the other team increase their chances to score and the team of the player decrease their chances to score. We might want to find this position as a naive imputation. \n\n> I need to explore more visualization, using things such as the one use in other competitions and in football. \n\n### Understand the game\n\nThere are several discussions that allow to understand the game, and watching some minutes can also give you a better intuition on what's going on. \n\nThe key things that I found were:\n\n- The game is 3 vs 3\n- The game is similar to football, where the objective is to score a ball in the opposite field\n- Players can be destroyed by other player for 3 seconds and reaper in their own field\n\nThings I still don't know:\n- Why each game has more than one event?\n- What's the size of the field?\n\n### Check NFL notebook\n\n- The winning solution was a neural network\n- They simplified the game by using a runner and defenders and remove teammates\n- The size of the dataset was significantly lower that this competition\n- They had to optimize their memory to came up with the winning solution\n\nThings I still don't know:\nHow do they handle missing values?\n\n### Second week plan\n\n**Find a way to handle missing values**\n\n- Check the NFL notebook to see how did they handle those values\n- Check discussion forum and notebooks to see how people are handling missing values\n- Understand the field and the position for, remove players that make more sense (lower probability of their team and increase the probability of the other team)\n\n**Create a visualization notebook using football approaches**\n\n- Check this [YouTube channel](https://www.youtube.com/channel/UCUBFJYcag8j2rm_9HkrrA7w) that has tracking approaches for football to draw a field and applied speed and position of the players\n- Check the discussion forum and notebooks to see EDA ideas\n- Create an EDA notebook that allow to understand the features and came up with a modeling plan\n- From the EDA perform some feature engineering on the dataset\n\n**Prepare the third week plan**\n\n- Create baseline model",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1978221": "The first week was an exploratory week to understand the competition, the data, the initial suggestions and the game itself.\n\nMy previous week plan was:\n\n1. Understand online learning\n2. Check the discussion forums for ideas\n3. Understand the game\n4. Check NFL notebooks\n\nLet's go over one by one with the findings:\n\n### Understand online learning\n\nOnline learning is a method that allow to make predictions in live environment, it's a useful tool with a lot of applications such as stock market, sports events, weather and many others. It consumes less memory than batch prediction and tools like [river](https://riverml.xyz/0.13.0/examples/batch-to-online/) mimics the sklearn syntax which make it really easy to implement. \n\nThe cons, is that the resources are limited, and I couldn't find an example in a kaggle competition, also the performance is lower than batch prediction according to the river documentation. How to handle missing values can also be a problem if you perform some type of calculation on historical data.\n\nI am also not sure how does it work internally, maybe is a loop over every row in the model, if so it sounds computational expensive, and maybe a better approach will be using a sample of the training data with a batch prediction. I saw some results here that might support that path. \n\n### Check discussion forum for ideas\n\nKey ideas I found were:\n\n- [Reduce the size of the dataset](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356540)\n- [Handle missing values](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/357158)\n- [Online learning resources](https://www.kaggle.com/competitions/tabular-playground-series-oct-2022/discussion/356584)\n- Some visualizations\n\n> Missing values needs to be explored a little further, which could start by checking similar approaches for other competitions, what happened when a player leaves the field? My guess is that we care to place in a position when the other team increase their chances to score and the team of the player decrease their chances to score. We might want to find this position as a naive imputation. \n\n> I need to explore more visualization, using things such as the one use in other competitions and in football. \n\n### Understand the game\n\nThere are several discussions that allow to understand the game, and watching some minutes can also give you a better intuition on what's going on. \n\nThe key things that I found were:\n\n- The game is 3 vs 3\n- The game is similar to football, where the objective is to score a ball in the opposite field\n- Players can be destroyed by other player for 3 seconds and reaper in their own field\n\nThings I still don't know:\n- Why each game has more than one event?\n- What's the size of the field?\n\n### Check NFL notebook\n\n- The winning solution was a neural network\n- They simplified the game by using a runner and defenders and remove teammates\n- The size of the dataset was significantly lower that this competition\n- They had to optimize their memory to came up with the winning solution\n\nThings I still don't know:\nHow do they handle missing values?\n\n### Second week plan\n\n**Find a way to handle missing values**\n\n- Check the NFL notebook to see how did they handle those values\n- Check discussion forum and notebooks to see how people are handling missing values\n- Understand the field and the position for, remove players that make more sense (lower probability of their team and increase the probability of the other team)\n\n**Create a visualization notebook using football approaches**\n\n- Check this [YouTube channel](https://www.youtube.com/channel/UCUBFJYcag8j2rm_9HkrrA7w) that has tracking approaches for football to draw a field and applied speed and position of the players\n- Check the discussion forum and notebooks to see EDA ideas\n- Create an EDA notebook that allow to understand the features and came up with a modeling plan\n- From the EDA perform some feature engineering on the dataset\n\n**Prepare the third week plan**\n\n- Create baseline model"
  }
}