{
  "id": 356996,
  "title": "Null values are useful",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/356996",
  "author_name": "",
  "post_date": "2022-10-02T20:27:09.542315500Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>It is safe to assume that null values are \"demolitions\".</p>\n<p>In rocket-league words, demolition is when a player crashes another with supersonic speed, causing the hit player to explode and to wait 3 seconds to re-spawn.</p>\n<p>Rocket League is a fast game, three seconds it's like an eternity. I wanted to check whether the proportion of scores vary given that one or more players are gone.</p>\n<p>The code for Hypothesis testing for population proportions it's available <a href=\"https://www.kaggle.com/code/jcaliz/tps-oct22-quickstart-eda-keras-baseline?scriptVersionId=107050788#Agregate-counts-and-do-the-test\" target=\"_blank\">here</a>,  and the results are:</p>\n<p>Global mean: 0.112</p>\n<table>\n<thead>\n<tr>\n<th>label</th>\n<th>z_score</th>\n<th>count/nobs (mean)</th>\n<th>p_value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>demolition_yes</td>\n<td>71.03</td>\n<td>148155/1108079 : 0.134</td>\n<td>0.0000</td>\n</tr>\n<tr>\n<td>demolition_no</td>\n<td>-16.68</td>\n<td>2234349/20089957 : 0.111</td>\n<td>0.0000</td>\n</tr>\n</tbody>\n</table>\n<p>The means are different by tiny amounts, since <code>N</code> is so large statistically speaking, the means are different.</p>\n<p>What can we do with that?</p>\n<ol>\n<li>During imputation, add an indicator column with <code>1</code> if one player is out. Good for Neural Networks</li>\n<li>Impute based on locations that are less probable for a player to score.</li>\n<li>Nothing, and let the GBDT do the work.</li>\n</ol>",
  "messages": [
    {
      "id": "1968035",
      "postDate": "10/02/2022 20:27:09",
      "content": "<p>It is safe to assume that null values are \"demolitions\".</p>\n<p>In rocket-league words, demolition is when a player crashes another with supersonic speed, causing the hit player to explode and to wait 3 seconds to re-spawn.</p>\n<p>Rocket League is a fast game, three seconds it's like an eternity. I wanted to check whether the proportion of scores vary given that one or more players are gone.</p>\n<p>The code for Hypothesis testing for population proportions it's available <a href=\"https://www.kaggle.com/code/jcaliz/tps-oct22-quickstart-eda-keras-baseline?scriptVersionId=107050788#Agregate-counts-and-do-the-test\" target=\"_blank\">here</a>,  and the results are:</p>\n<p>Global mean: 0.112</p>\n<table>\n<thead>\n<tr>\n<th>label</th>\n<th>z_score</th>\n<th>count/nobs (mean)</th>\n<th>p_value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>demolition_yes</td>\n<td>71.03</td>\n<td>148155/1108079 : 0.134</td>\n<td>0.0000</td>\n</tr>\n<tr>\n<td>demolition_no</td>\n<td>-16.68</td>\n<td>2234349/20089957 : 0.111</td>\n<td>0.0000</td>\n</tr>\n</tbody>\n</table>\n<p>The means are different by tiny amounts, since <code>N</code> is so large statistically speaking, the means are different.</p>\n<p>What can we do with that?</p>\n<ol>\n<li>During imputation, add an indicator column with <code>1</code> if one player is out. Good for Neural Networks</li>\n<li>Impute based on locations that are less probable for a player to score.</li>\n<li>Nothing, and let the GBDT do the work.</li>\n</ol>",
      "rawMarkdown": "It is safe to assume that null values are \"demolitions\".\n\nIn rocket-league words, demolition is when a player crashes another with supersonic speed, causing the hit player to explode and to wait 3 seconds to re-spawn.\n\nRocket League is a fast game, three seconds it's like an eternity. I wanted to check whether the proportion of scores vary given that one or more players are gone.\n\nThe code for Hypothesis testing for population proportions it's available [here](https://www.kaggle.com/code/jcaliz/tps-oct22-quickstart-eda-keras-baseline?scriptVersionId=107050788#Agregate-counts-and-do-the-test),  and the results are:\n\nGlobal mean: 0.112\n\n| label | z_score  | count/nobs (mean)   | p_value |\n| ----- | -------- | --------------------- | -------- |\n|demolition_yes |          71.03   |        148155/1108079 : 0.134    |      0.0000 |\n|demolition_no  |        -16.68    |     2234349/20089957 : 0.111     |   0.0000 |\n\n\nThe means are different by tiny amounts, since `N` is so large statistically speaking, the means are different.\n\nWhat can we do with that?\n1. During imputation, add an indicator column with `1` if one player is out. Good for Neural Networks\n2. Impute based on locations that are less probable for a player to score.\n3. Nothing, and let the GBDT do the work.",
      "votes": null
    },
    {
      "id": "1968185",
      "postDate": "10/03/2022 00:43:16",
      "content": "<p>Thanks for the insight, <a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a> .</p>",
      "rawMarkdown": "Thanks for the insight, @jcaliz .",
      "votes": null
    },
    {
      "id": "1968188",
      "postDate": "10/03/2022 00:49:46",
      "content": "<p>No problem.</p>",
      "rawMarkdown": "No problem.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1968185,
      "author_name": "mahdeemushfiquekamal",
      "author_url": "",
      "post_date": "10/03/2022 00:43:16",
      "content": "<p>Thanks for the insight, <a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a> .</p>",
      "votes": null,
      "replies": [
        {
          "id": 1968188,
          "author_name": "jcaliz",
          "author_url": "",
          "post_date": "10/03/2022 00:49:46",
          "content": "<p>No problem.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1968035": "It is safe to assume that null values are \"demolitions\".\n\nIn rocket-league words, demolition is when a player crashes another with supersonic speed, causing the hit player to explode and to wait 3 seconds to re-spawn.\n\nRocket League is a fast game, three seconds it's like an eternity. I wanted to check whether the proportion of scores vary given that one or more players are gone.\n\nThe code for Hypothesis testing for population proportions it's available [here](https://www.kaggle.com/code/jcaliz/tps-oct22-quickstart-eda-keras-baseline?scriptVersionId=107050788#Agregate-counts-and-do-the-test),  and the results are:\n\nGlobal mean: 0.112\n\n| label | z_score  | count/nobs (mean)   | p_value |\n| ----- | -------- | --------------------- | -------- |\n|demolition_yes |          71.03   |        148155/1108079 : 0.134    |      0.0000 |\n|demolition_no  |        -16.68    |     2234349/20089957 : 0.111     |   0.0000 |\n\n\nThe means are different by tiny amounts, since `N` is so large statistically speaking, the means are different.\n\nWhat can we do with that?\n1. During imputation, add an indicator column with `1` if one player is out. Good for Neural Networks\n2. Impute based on locations that are less probable for a player to score.\n3. Nothing, and let the GBDT do the work.",
    "1968185": "Thanks for the insight, @jcaliz .",
    "1968188": "No problem."
  },
  "source": "meta"
}