{
  "id": 362325,
  "title": "Feature engineering recap",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/362325",
  "author_name": "Mateus Coelho",
  "post_date": "2022-10-26T17:29:30.199000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I know the competition is not over yet, but here is a compilation of feature ideas for this competition. This is a really nice competition to refine our feature engineering habilities. Not only having ideias and computing them, but also making them fast to calculate.</p>\n<p>Most of the ideas below are from discussions, kernels and the competition NFL Big Data Bowl 2018. A lot of people had the same ideas so I'm not going to provide names because I may be unfair to someone.</p>\n<p><strong>Starting with the basics</strong></p>\n<p>Raw position and velocity functions are useful, but player's ID are random, so models may have a hard time learning to use this features. So a lot of people is suggesting / suggested sorting the players by their distance to the ball, for example. </p>\n<p>For each of the features below you can compute aggregations over each team and all players.</p>\n<ol>\n<li>Compute the norm (both 3d and 2d) of the velocity vector.</li>\n<li>Compute the distance (2d and 3d) between each player and the ball. </li>\n<li>Number of players currently playing (remeber that cars can explode)</li>\n<li>Number of orbs available </li>\n<li>Distance of each player to each orb</li>\n<li>Average, minimum, maximum, variance of X, Y and Z positions.</li>\n<li>Average, minimum, maximum, variance of X, Y and Z velocities.</li>\n<li>Number of players on the offensive or deffensive area.</li>\n<li>Number of players elevated.</li>\n<li>Number of oponents between the ball and their goal. </li>\n</ol>\n<p><strong>More advanced stuff</strong></p>\n<ol>\n<li>Compute the time necessary for each player to reach the ball.</li>\n<li>Compute the angle between the vector that connects two players (one for each team) and the ball.</li>\n<li>Compute the angle between the ball velocity vector and each player's valocity vector.</li>\n<li>Features based on voronoi diagrams (as pointed out <a href=\"https://www.kaggle.com/code/mateuscco/voronoi-diagrams-computing-player-s-influence\" target=\"_blank\">here</a>).</li>\n<li><a href=\"https://www.kaggle.com/code/pednt9/vip-hint-coded/notebook#Little-Standardization-Step\" target=\"_blank\">Pitch control</a> can be probably adapted to this competition.</li>\n<li>This may seem nonsense, but computing the positions of each player and the ball <em>n</em> seconds in the future (assuming the velocities don't change) was very useful in NFL Big Data Bowl 2018. My intuition is that LGBM and NNs by themselves can't learn fully learn how the plays will be in the future. So giving these hints (which are mere approximations because players can change their velocities all the time) is useful.</li>\n</ol>\n<p>If you have more ideas to add, please write them below.</p>",
  "messages": [
    {
      "id": 2005043,
      "postDate": "2022-10-26T17:29:30.200Z",
      "content": "<p>I know the competition is not over yet, but here is a compilation of feature ideas for this competition. This is a really nice competition to refine our feature engineering habilities. Not only having ideias and computing them, but also making them fast to calculate.</p>\n<p>Most of the ideas below are from discussions, kernels and the competition NFL Big Data Bowl 2018. A lot of people had the same ideas so I'm not going to provide names because I may be unfair to someone.</p>\n<p><strong>Starting with the basics</strong></p>\n<p>Raw position and velocity functions are useful, but player's ID are random, so models may have a hard time learning to use this features. So a lot of people is suggesting / suggested sorting the players by their distance to the ball, for example. </p>\n<p>For each of the features below you can compute aggregations over each team and all players.</p>\n<ol>\n<li>Compute the norm (both 3d and 2d) of the velocity vector.</li>\n<li>Compute the distance (2d and 3d) between each player and the ball. </li>\n<li>Number of players currently playing (remeber that cars can explode)</li>\n<li>Number of orbs available </li>\n<li>Distance of each player to each orb</li>\n<li>Average, minimum, maximum, variance of X, Y and Z positions.</li>\n<li>Average, minimum, maximum, variance of X, Y and Z velocities.</li>\n<li>Number of players on the offensive or deffensive area.</li>\n<li>Number of players elevated.</li>\n<li>Number of oponents between the ball and their goal. </li>\n</ol>\n<p><strong>More advanced stuff</strong></p>\n<ol>\n<li>Compute the time necessary for each player to reach the ball.</li>\n<li>Compute the angle between the vector that connects two players (one for each team) and the ball.</li>\n<li>Compute the angle between the ball velocity vector and each player's valocity vector.</li>\n<li>Features based on voronoi diagrams (as pointed out <a href=\"https://www.kaggle.com/code/mateuscco/voronoi-diagrams-computing-player-s-influence\" target=\"_blank\">here</a>).</li>\n<li><a href=\"https://www.kaggle.com/code/pednt9/vip-hint-coded/notebook#Little-Standardization-Step\" target=\"_blank\">Pitch control</a> can be probably adapted to this competition.</li>\n<li>This may seem nonsense, but computing the positions of each player and the ball <em>n</em> seconds in the future (assuming the velocities don't change) was very useful in NFL Big Data Bowl 2018. My intuition is that LGBM and NNs by themselves can't learn fully learn how the plays will be in the future. So giving these hints (which are mere approximations because players can change their velocities all the time) is useful.</li>\n</ol>\n<p>If you have more ideas to add, please write them below.</p>",
      "rawMarkdown": "I know the competition is not over yet, but here is a compilation of feature ideas for this competition. This is a really nice competition to refine our feature engineering habilities. Not only having ideias and computing them, but also making them fast to calculate.\n\nMost of the ideas below are from discussions, kernels and the competition NFL Big Data Bowl 2018. A lot of people had the same ideas so I'm not going to provide names because I may be unfair to someone.\n\n**Starting with the basics**\n\nRaw position and velocity functions are useful, but player's ID are random, so models may have a hard time learning to use this features. So a lot of people is suggesting / suggested sorting the players by their distance to the ball, for example. \n\nFor each of the features below you can compute aggregations over each team and all players.\n\n1. Compute the norm (both 3d and 2d) of the velocity vector.\n2. Compute the distance (2d and 3d) between each player and the ball. \n3. Number of players currently playing (remeber that cars can explode)\n4. Number of orbs available \n5. Distance of each player to each orb\n6. Average, minimum, maximum, variance of X, Y and Z positions.\n7. Average, minimum, maximum, variance of X, Y and Z velocities.\n8. Number of players on the offensive or deffensive area.\n9. Number of players elevated.\n10. Number of oponents between the ball and their goal. \n\n\n**More advanced stuff**\n\n1. Compute the time necessary for each player to reach the ball.\n2. Compute the angle between the vector that connects two players (one for each team) and the ball.\n3. Compute the angle between the ball velocity vector and each player's valocity vector.\n4. Features based on voronoi diagrams (as pointed out [here](https://www.kaggle.com/code/mateuscco/voronoi-diagrams-computing-player-s-influence)).\n5. [Pitch control](https://www.kaggle.com/code/pednt9/vip-hint-coded/notebook#Little-Standardization-Step) can be probably adapted to this competition.\n6. This may seem nonsense, but computing the positions of each player and the ball *n* seconds in the future (assuming the velocities don't change) was very useful in NFL Big Data Bowl 2018. My intuition is that LGBM and NNs by themselves can't learn fully learn how the plays will be in the future. So giving these hints (which are mere approximations because players can change their velocities all the time) is useful.\n\nIf you have more ideas to add, please write them below.",
      "votes": 14
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2005043": "I know the competition is not over yet, but here is a compilation of feature ideas for this competition. This is a really nice competition to refine our feature engineering habilities. Not only having ideias and computing them, but also making them fast to calculate.\n\nMost of the ideas below are from discussions, kernels and the competition NFL Big Data Bowl 2018. A lot of people had the same ideas so I'm not going to provide names because I may be unfair to someone.\n\n**Starting with the basics**\n\nRaw position and velocity functions are useful, but player's ID are random, so models may have a hard time learning to use this features. So a lot of people is suggesting / suggested sorting the players by their distance to the ball, for example. \n\nFor each of the features below you can compute aggregations over each team and all players.\n\n1. Compute the norm (both 3d and 2d) of the velocity vector.\n2. Compute the distance (2d and 3d) between each player and the ball. \n3. Number of players currently playing (remeber that cars can explode)\n4. Number of orbs available \n5. Distance of each player to each orb\n6. Average, minimum, maximum, variance of X, Y and Z positions.\n7. Average, minimum, maximum, variance of X, Y and Z velocities.\n8. Number of players on the offensive or deffensive area.\n9. Number of players elevated.\n10. Number of oponents between the ball and their goal. \n\n\n**More advanced stuff**\n\n1. Compute the time necessary for each player to reach the ball.\n2. Compute the angle between the vector that connects two players (one for each team) and the ball.\n3. Compute the angle between the ball velocity vector and each player's valocity vector.\n4. Features based on voronoi diagrams (as pointed out [here](https://www.kaggle.com/code/mateuscco/voronoi-diagrams-computing-player-s-influence)).\n5. [Pitch control](https://www.kaggle.com/code/pednt9/vip-hint-coded/notebook#Little-Standardization-Step) can be probably adapted to this competition.\n6. This may seem nonsense, but computing the positions of each player and the ball *n* seconds in the future (assuming the velocities don't change) was very useful in NFL Big Data Bowl 2018. My intuition is that LGBM and NNs by themselves can't learn fully learn how the plays will be in the future. So giving these hints (which are mere approximations because players can change their velocities all the time) is useful.\n\nIf you have more ideas to add, please write them below."
  }
}