{
  "id": 356545,
  "title": "Player Position Null Values",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/356545",
  "author_name": "",
  "post_date": "2022-10-01T01:49:16.199317500Z",
  "votes": 24,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I created a 3D animation of a player's car position and noticed some null values during the animation. Then I remembered a core game mechanic of Rocket League when I used to play.</p>\n<h2>What is Demolition in Rocket League?</h2>\n<p>Demolition is referred to the bumps occur when a player runs into another player at supersonic speed. The victim of a demolition will explode and combust, but will respawn near their team's goal after a few seconds. If two players run into each other both at supersonic speed and their fenders collide, they will both be demolished. In higher-level gameplay, demolitions are sometimes used tactically to prevent an opponent from scoring, or to clear the way for a teammate to score.</p>\n<p>You can view the notebook with the animation <a href=\"https://www.kaggle.com/code/mattop/rocket-league-tps-eda/notebook\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": "1964906",
      "postDate": "10/01/2022 01:49:16",
      "content": "<p>I created a 3D animation of a player's car position and noticed some null values during the animation. Then I remembered a core game mechanic of Rocket League when I used to play.</p>\n<h2>What is Demolition in Rocket League?</h2>\n<p>Demolition is referred to the bumps occur when a player runs into another player at supersonic speed. The victim of a demolition will explode and combust, but will respawn near their team's goal after a few seconds. If two players run into each other both at supersonic speed and their fenders collide, they will both be demolished. In higher-level gameplay, demolitions are sometimes used tactically to prevent an opponent from scoring, or to clear the way for a teammate to score.</p>\n<p>You can view the notebook with the animation <a href=\"https://www.kaggle.com/code/mattop/rocket-league-tps-eda/notebook\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I created a 3D animation of a player's car position and noticed some null values during the animation. Then I remembered a core game mechanic of Rocket League when I used to play.\n\n##What is Demolition in Rocket League?\n\nDemolition is referred to the bumps occur when a player runs into another player at supersonic speed. The victim of a demolition will explode and combust, but will respawn near their team's goal after a few seconds. If two players run into each other both at supersonic speed and their fenders collide, they will both be demolished. In higher-level gameplay, demolitions are sometimes used tactically to prevent an opponent from scoring, or to clear the way for a teammate to score.\n\nYou can view the notebook with the animation [here](https://www.kaggle.com/code/mattop/rocket-league-tps-eda/notebook)",
      "votes": null
    },
    {
      "id": "1966450",
      "postDate": "10/02/2022 02:45:42",
      "content": "<p>Could we potentially engineer a feature from these null values that would utilize the lack of player(s) on the field? Also, I just added a 3D line plot to the notebook visualizing a segment of each player's race paths from both teams in train_0.</p>",
      "rawMarkdown": "Could we potentially engineer a feature from these null values that would utilize the lack of player(s) on the field? Also, I just added a 3D line plot to the notebook visualizing a segment of each player's race paths from both teams in train_0.",
      "votes": null
    },
    {
      "id": "1966524",
      "postDate": "10/02/2022 04:31:40",
      "content": "<p>I tried the obvious binary features; correlations are in the right directions albeit weak.😐</p>\n<pre><code>X = pd.DataFrame([],index=train_0.index)\nfor i in range(6):\n    X[F'p{i}_isnull'] = train_0[F'p{i}_pos_x'].isnull().astype(np.int8)\n\ndisplay((X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).corr(train_0['team_A_scoring_within_10sec']))\ndisplay((X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).corr(train_0['team_B_scoring_within_10sec']))\ndisplay((X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).corr(train_0['team_A_scoring_within_10sec']))\ndisplay((X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).corr(train_0['team_B_scoring_within_10sec']))\n</code></pre>\n<pre><code>-0.009979106983235795\n0.024060698532368793\n0.01614721452926487\n-0.007115157504775068\n</code></pre>",
      "rawMarkdown": "I tried the obvious binary features; correlations are in the right directions albeit weak.😐\n```\nX = pd.DataFrame([],index=train_0.index)\nfor i in range(6):\n    X[F'p{i}_isnull'] = train_0[F'p{i}_pos_x'].isnull().astype(np.int8)\n\ndisplay((X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).corr(train_0['team_A_scoring_within_10sec']))\ndisplay((X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).corr(train_0['team_B_scoring_within_10sec']))\ndisplay((X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).corr(train_0['team_A_scoring_within_10sec']))\ndisplay((X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).corr(train_0['team_B_scoring_within_10sec']))\n```\n```\n-0.009979106983235795\n0.024060698532368793\n0.01614721452926487\n-0.007115157504775068\n```",
      "votes": null
    },
    {
      "id": "1966623",
      "postDate": "10/02/2022 05:22:36",
      "content": "<p>Interesting… I tried adding a small <code>.shift()</code> in an attempt to improve the correlation (in Rocket League you respawn in one of your team's corners after a demolition) but had no success.</p>\n<p><code>X[F'p{i}_isnull'] = train_0[F'p{i}_pos_x'].shift().isnull().astype(np.int8)</code></p>",
      "rawMarkdown": "Interesting... I tried adding a small `.shift()` in an attempt to improve the correlation (in Rocket League you respawn in one of your team's corners after a demolition) but had no success.\n\n`X[F'p{i}_isnull'] = train_0[F'p{i}_pos_x'].shift().isnull().astype(np.int8)`",
      "votes": null
    },
    {
      "id": "1967478",
      "postDate": "10/02/2022 14:21:59",
      "content": "<p>Hi Matt. 😄Thanks for sharing !<br>\nI know the coordinate is meaningful in this contest, so do we still need to normalize the data?👀</p>",
      "rawMarkdown": "Hi Matt. 😄Thanks for sharing !\nI know the coordinate is meaningful in this contest, so do we still need to normalize the data?👀",
      "votes": null
    },
    {
      "id": "1967504",
      "postDate": "10/02/2022 14:36:46",
      "content": "<p>Now this is the kind of <em>domain knowledge</em> that I value 😆</p>",
      "rawMarkdown": "Now this is the kind of *domain knowledge* that I value 😆",
      "votes": null
    },
    {
      "id": "1967775",
      "postDate": "10/02/2022 17:11:56",
      "content": "<p>That's a good question <a href=\"https://www.kaggle.com/chal1ce\" target=\"_blank\">@chal1ce</a>! I think creating a solid CV strategy and doing some testing would be helpful in finding the answer. I see you are using StandardScaler in your notebook.</p>",
      "rawMarkdown": "That's a good question @chal1ce! I think creating a solid CV strategy and doing some testing would be helpful in finding the answer. I see you are using StandardScaler in your notebook.",
      "votes": null
    },
    {
      "id": "1967776",
      "postDate": "10/02/2022 17:12:18",
      "content": "<p>I forgot that demolitions were even part of the game haha</p>",
      "rawMarkdown": "I forgot that demolitions were even part of the game haha",
      "votes": null
    },
    {
      "id": "1968169",
      "postDate": "10/03/2022 00:13:34",
      "content": "<p>Yes, because either too high or too low data will have an impact on the model. <br>\nAfter I use StandardScaler to make them conform to the normal distribution, the accuracy of the model will be 0.0004 to 0.0005 higher than the original.</p>",
      "rawMarkdown": "Yes, because either too high or too low data will have an impact on the model. \nAfter I use StandardScaler to make them conform to the normal distribution, the accuracy of the model will be 0.0004 to 0.0005 higher than the original.",
      "votes": null
    },
    {
      "id": "1968242",
      "postDate": "10/03/2022 02:14:07",
      "content": "<p>Nevertheless the effects of these binary features on the target variables are statistically significant (very small p-values from the proportions z-test) so including them should be beneficial to the classification task.</p>\n<pre><code>from statsmodels.stats.proportion import proportions_ztest\n\ndf = train_0.loc[(X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).astype(bool)]\nfor team in 'AB':\n    prob = train_0[F'team_{team}_scoring_within_10sec'].mean()\n    display(proportions_ztest(df[F'team_{team}_scoring_within_10sec'].sum(),\\\n                                                len(df),prob,prop_var=prob))\n\ndf = train_0.loc[(X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).astype(bool)]\nfor team in 'AB':\n    prob = train_0[F'team_{team}_scoring_within_10sec'].mean()\n    display(proportions_ztest(df[F'team_{team}_scoring_within_10sec'].sum(),\\\n                                                len(df),prob,prop_var=prob))\n</code></pre>\n<pre><code>(-14.436060507821042, 3.06888721253238e-47)\n(34.806892085356104, 1.9131495080932972e-265)\n(23.35804952768753, 1.1415816957073028e-120)\n(-10.292561673261257, 7.612373796890326e-25)\n</code></pre>",
      "rawMarkdown": "Nevertheless the effects of these binary features on the target variables are statistically significant (very small p-values from the proportions z-test) so including them should be beneficial to the classification task.\n```\nfrom statsmodels.stats.proportion import proportions_ztest\n\ndf = train_0.loc[(X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).astype(bool)]\nfor team in 'AB':\n    prob = train_0[F'team_{team}_scoring_within_10sec'].mean()\n    display(proportions_ztest(df[F'team_{team}_scoring_within_10sec'].sum(),\\\n                                                len(df),prob,prop_var=prob))\n    \ndf = train_0.loc[(X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).astype(bool)]\nfor team in 'AB':\n    prob = train_0[F'team_{team}_scoring_within_10sec'].mean()\n    display(proportions_ztest(df[F'team_{team}_scoring_within_10sec'].sum(),\\\n                                                len(df),prob,prop_var=prob))\n```\n```\n(-14.436060507821042, 3.06888721253238e-47)\n(34.806892085356104, 1.9131495080932972e-265)\n(23.35804952768753, 1.1415816957073028e-120)\n(-10.292561673261257, 7.612373796890326e-25)\n```",
      "votes": null
    },
    {
      "id": "1969621",
      "postDate": "10/03/2022 15:38:47",
      "content": "<p>Very nice insight, those p-values are really small!</p>",
      "rawMarkdown": "Very nice insight, those p-values are really small!",
      "votes": null
    },
    {
      "id": "1990637",
      "postDate": "10/16/2022 17:04:46",
      "content": "<p>That's a good question <a href=\"https://www.kaggle.com/chal1ce\" target=\"_blank\">@chal1ce</a>! I think creating a solid CV strategy and doing some testing would be helpful in finding the answer. I see you are using StandardScaler in your notebook.</p>",
      "rawMarkdown": "That's a good question @chal1ce! I think creating a solid CV strategy and doing some testing would be helpful in finding the answer. I see you are using StandardScaler in your notebook.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1966450,
      "author_name": "mattop",
      "author_url": "",
      "post_date": "10/02/2022 02:45:42",
      "content": "<p>Could we potentially engineer a feature from these null values that would utilize the lack of player(s) on the field? Also, I just added a 3D line plot to the notebook visualizing a segment of each player's race paths from both teams in train_0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1966524,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "10/02/2022 04:31:40",
          "content": "<p>I tried the obvious binary features; correlations are in the right directions albeit weak.😐</p>\n<pre><code>X = pd.DataFrame([],index=train_0.index)\nfor i in range(6):\n    X[F'p{i}_isnull'] = train_0[F'p{i}_pos_x'].isnull().astype(np.int8)\n\ndisplay((X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).corr(train_0['team_A_scoring_within_10sec']))\ndisplay((X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).corr(train_0['team_B_scoring_within_10sec']))\ndisplay((X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).corr(train_0['team_A_scoring_within_10sec']))\ndisplay((X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).corr(train_0['team_B_scoring_within_10sec']))\n</code></pre>\n<pre><code>-0.009979106983235795\n0.024060698532368793\n0.01614721452926487\n-0.007115157504775068\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1966623,
          "author_name": "mattop",
          "author_url": "",
          "post_date": "10/02/2022 05:22:36",
          "content": "<p>Interesting… I tried adding a small <code>.shift()</code> in an attempt to improve the correlation (in Rocket League you respawn in one of your team's corners after a demolition) but had no success.</p>\n<p><code>X[F'p{i}_isnull'] = train_0[F'p{i}_pos_x'].shift().isnull().astype(np.int8)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1968242,
          "author_name": "siukeitin",
          "author_url": "",
          "post_date": "10/03/2022 02:14:07",
          "content": "<p>Nevertheless the effects of these binary features on the target variables are statistically significant (very small p-values from the proportions z-test) so including them should be beneficial to the classification task.</p>\n<pre><code>from statsmodels.stats.proportion import proportions_ztest\n\ndf = train_0.loc[(X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).astype(bool)]\nfor team in 'AB':\n    prob = train_0[F'team_{team}_scoring_within_10sec'].mean()\n    display(proportions_ztest(df[F'team_{team}_scoring_within_10sec'].sum(),\\\n                                                len(df),prob,prop_var=prob))\n\ndf = train_0.loc[(X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).astype(bool)]\nfor team in 'AB':\n    prob = train_0[F'team_{team}_scoring_within_10sec'].mean()\n    display(proportions_ztest(df[F'team_{team}_scoring_within_10sec'].sum(),\\\n                                                len(df),prob,prop_var=prob))\n</code></pre>\n<pre><code>(-14.436060507821042, 3.06888721253238e-47)\n(34.806892085356104, 1.9131495080932972e-265)\n(23.35804952768753, 1.1415816957073028e-120)\n(-10.292561673261257, 7.612373796890326e-25)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1969621,
          "author_name": "mattop",
          "author_url": "",
          "post_date": "10/03/2022 15:38:47",
          "content": "<p>Very nice insight, those p-values are really small!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1967478,
      "author_name": "",
      "author_url": "",
      "post_date": "10/02/2022 14:21:59",
      "content": "<p>Hi Matt. 😄Thanks for sharing !<br>\nI know the coordinate is meaningful in this contest, so do we still need to normalize the data?👀</p>",
      "votes": null,
      "replies": [
        {
          "id": 1967775,
          "author_name": "mattop",
          "author_url": "",
          "post_date": "10/02/2022 17:11:56",
          "content": "<p>That's a good question <a href=\"https://www.kaggle.com/chal1ce\" target=\"_blank\">@chal1ce</a>! I think creating a solid CV strategy and doing some testing would be helpful in finding the answer. I see you are using StandardScaler in your notebook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1968169,
          "author_name": "",
          "author_url": "",
          "post_date": "10/03/2022 00:13:34",
          "content": "<p>Yes, because either too high or too low data will have an impact on the model. <br>\nAfter I use StandardScaler to make them conform to the normal distribution, the accuracy of the model will be 0.0004 to 0.0005 higher than the original.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1967504,
      "author_name": "samuelcortinhas",
      "author_url": "",
      "post_date": "10/02/2022 14:36:46",
      "content": "<p>Now this is the kind of <em>domain knowledge</em> that I value 😆</p>",
      "votes": null,
      "replies": [
        {
          "id": 1967776,
          "author_name": "mattop",
          "author_url": "",
          "post_date": "10/02/2022 17:12:18",
          "content": "<p>I forgot that demolitions were even part of the game haha</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1990637,
      "author_name": "naveedakram1",
      "author_url": "",
      "post_date": "10/16/2022 17:04:46",
      "content": "<p>That's a good question <a href=\"https://www.kaggle.com/chal1ce\" target=\"_blank\">@chal1ce</a>! I think creating a solid CV strategy and doing some testing would be helpful in finding the answer. I see you are using StandardScaler in your notebook.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1964906": "I created a 3D animation of a player's car position and noticed some null values during the animation. Then I remembered a core game mechanic of Rocket League when I used to play.\n\n##What is Demolition in Rocket League?\n\nDemolition is referred to the bumps occur when a player runs into another player at supersonic speed. The victim of a demolition will explode and combust, but will respawn near their team's goal after a few seconds. If two players run into each other both at supersonic speed and their fenders collide, they will both be demolished. In higher-level gameplay, demolitions are sometimes used tactically to prevent an opponent from scoring, or to clear the way for a teammate to score.\n\nYou can view the notebook with the animation [here](https://www.kaggle.com/code/mattop/rocket-league-tps-eda/notebook)",
    "1966450": "Could we potentially engineer a feature from these null values that would utilize the lack of player(s) on the field? Also, I just added a 3D line plot to the notebook visualizing a segment of each player's race paths from both teams in train_0.",
    "1966524": "I tried the obvious binary features; correlations are in the right directions albeit weak.😐\n```\nX = pd.DataFrame([],index=train_0.index)\nfor i in range(6):\n    X[F'p{i}_isnull'] = train_0[F'p{i}_pos_x'].isnull().astype(np.int8)\n\ndisplay((X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).corr(train_0['team_A_scoring_within_10sec']))\ndisplay((X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).corr(train_0['team_B_scoring_within_10sec']))\ndisplay((X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).corr(train_0['team_A_scoring_within_10sec']))\ndisplay((X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).corr(train_0['team_B_scoring_within_10sec']))\n```\n```\n-0.009979106983235795\n0.024060698532368793\n0.01614721452926487\n-0.007115157504775068\n```",
    "1966623": "Interesting... I tried adding a small `.shift()` in an attempt to improve the correlation (in Rocket League you respawn in one of your team's corners after a demolition) but had no success.\n\n`X[F'p{i}_isnull'] = train_0[F'p{i}_pos_x'].shift().isnull().astype(np.int8)`",
    "1967478": "Hi Matt. 😄Thanks for sharing !\nI know the coordinate is meaningful in this contest, so do we still need to normalize the data?👀",
    "1967504": "Now this is the kind of *domain knowledge* that I value 😆",
    "1967775": "That's a good question @chal1ce! I think creating a solid CV strategy and doing some testing would be helpful in finding the answer. I see you are using StandardScaler in your notebook.",
    "1967776": "I forgot that demolitions were even part of the game haha",
    "1968169": "Yes, because either too high or too low data will have an impact on the model. \nAfter I use StandardScaler to make them conform to the normal distribution, the accuracy of the model will be 0.0004 to 0.0005 higher than the original.",
    "1968242": "Nevertheless the effects of these binary features on the target variables are statistically significant (very small p-values from the proportions z-test) so including them should be beneficial to the classification task.\n```\nfrom statsmodels.stats.proportion import proportions_ztest\n\ndf = train_0.loc[(X['p0_isnull']|X['p1_isnull']|X['p2_isnull']).astype(bool)]\nfor team in 'AB':\n    prob = train_0[F'team_{team}_scoring_within_10sec'].mean()\n    display(proportions_ztest(df[F'team_{team}_scoring_within_10sec'].sum(),\\\n                                                len(df),prob,prop_var=prob))\n    \ndf = train_0.loc[(X['p3_isnull']|X['p4_isnull']|X['p5_isnull']).astype(bool)]\nfor team in 'AB':\n    prob = train_0[F'team_{team}_scoring_within_10sec'].mean()\n    display(proportions_ztest(df[F'team_{team}_scoring_within_10sec'].sum(),\\\n                                                len(df),prob,prop_var=prob))\n```\n```\n(-14.436060507821042, 3.06888721253238e-47)\n(34.806892085356104, 1.9131495080932972e-265)\n(23.35804952768753, 1.1415816957073028e-120)\n(-10.292561673261257, 7.612373796890326e-25)\n```",
    "1969621": "Very nice insight, those p-values are really small!",
    "1990637": "That's a good question @chal1ce! I think creating a solid CV strategy and doing some testing would be helpful in finding the answer. I see you are using StandardScaler in your notebook."
  },
  "source": "meta"
}