{
  "id": 357158,
  "title": "📌 Best Missing Value Handling Method 🔥",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/357158",
  "author_name": "",
  "post_date": "2022-10-03T11:13:13.661334Z",
  "votes": 13,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>When I worked on missing values, I realized that missing values are related to the players. You can see these issues below figure too. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381067%2Fa3e4857a6a5439fcecc63ccbc5806126%2F__results___21_1.png?generation=1664794528222509&amp;alt=media\" alt=\"\"></p>\n<p>Player's placement and velocity change over time. Therefore the best missing value handling method is interpolation which is the process of estimating unknown values that fall between known values.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381067%2F475d7352e42c8b48afdcaa41d998d4d5%2FZNERC_r.png?generation=1664795700035054&amp;alt=media\" alt=\"\"><br>\nsource: <a href=\"https://en.wikipedia.org/wiki/Linear_interpolation\" target=\"_blank\">Wikipedia</a></p>\n<p>You can use the below code for imputation in numeric features or you can find the clean dataset <a href=\"https://www.kaggle.com/datasets/hasanbasriakcay/tpsoct22-traintest-parquet\" target=\"_blank\">here</a>.</p>\n<p><code>train_cl =  train.groupby([\"game_num\",\"event_id\"]).apply(lambda group: group.interpolate(method='index', limit_direction='both'))</code></p>\n<p>Kindly upvote if you find it useful 👍</p>",
  "messages": [
    {
      "id": "1969173",
      "postDate": "10/03/2022 11:13:13",
      "content": "<p>Hi all,</p>\n<p>When I worked on missing values, I realized that missing values are related to the players. You can see these issues below figure too. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381067%2Fa3e4857a6a5439fcecc63ccbc5806126%2F__results___21_1.png?generation=1664794528222509&amp;alt=media\" alt=\"\"></p>\n<p>Player's placement and velocity change over time. Therefore the best missing value handling method is interpolation which is the process of estimating unknown values that fall between known values.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381067%2F475d7352e42c8b48afdcaa41d998d4d5%2FZNERC_r.png?generation=1664795700035054&amp;alt=media\" alt=\"\"><br>\nsource: <a href=\"https://en.wikipedia.org/wiki/Linear_interpolation\" target=\"_blank\">Wikipedia</a></p>\n<p>You can use the below code for imputation in numeric features or you can find the clean dataset <a href=\"https://www.kaggle.com/datasets/hasanbasriakcay/tpsoct22-traintest-parquet\" target=\"_blank\">here</a>.</p>\n<p><code>train_cl =  train.groupby([\"game_num\",\"event_id\"]).apply(lambda group: group.interpolate(method='index', limit_direction='both'))</code></p>\n<p>Kindly upvote if you find it useful 👍</p>",
      "rawMarkdown": "Hi all,\n\nWhen I worked on missing values, I realized that missing values are related to the players. You can see these issues below figure too. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381067%2Fa3e4857a6a5439fcecc63ccbc5806126%2F__results___21_1.png?generation=1664794528222509&alt=media)\n\nPlayer's placement and velocity change over time. Therefore the best missing value handling method is interpolation which is the process of estimating unknown values that fall between known values.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381067%2F475d7352e42c8b48afdcaa41d998d4d5%2FZNERC_r.png?generation=1664795700035054&alt=media)\nsource: [Wikipedia](https://en.wikipedia.org/wiki/Linear_interpolation)\n\n\nYou can use the below code for imputation in numeric features or you can find the clean dataset [here](https://www.kaggle.com/datasets/hasanbasriakcay/tpsoct22-traintest-parquet).\n\n`train_cl =  train.groupby([\"game_num\",\"event_id\"]).apply(lambda group: group.interpolate(method='index', limit_direction='both'))`\n\nKindly upvote if you find it useful 👍",
      "votes": null
    },
    {
      "id": "1969421",
      "postDate": "10/03/2022 13:26:02",
      "content": "<p>Short with proper representation.. please keep it up <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a> </p>",
      "rawMarkdown": "Short with proper representation.. please keep it up @hasanbasriakcay",
      "votes": null
    },
    {
      "id": "1969499",
      "postDate": "10/03/2022 14:21:15",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a>,</p>\n<p>Thank you for sharing your Nan values imputation methodology. I agree with you that this imputation strategy is one of the most appropriate to impute players' positions, but it requires to have the chronology of players' positions. We do not have this information in the test set (it's scrambled). Any thoughts about what could be done at the test set?</p>\n<p>Maher</p>",
      "rawMarkdown": "Dear @hasanbasriakcay,\n\nThank you for sharing your Nan values imputation methodology. I agree with you that this imputation strategy is one of the most appropriate to impute players' positions, but it requires to have the chronology of players' positions. We do not have this information in the test set (it's scrambled). Any thoughts about what could be done at the test set?\n\nMaher",
      "votes": null
    },
    {
      "id": "1970178",
      "postDate": "10/03/2022 23:08:35",
      "content": "<p>Nice explanation. Definitely will be useful. Thanks for sharing this.</p>",
      "rawMarkdown": "Nice explanation. Definitely will be useful. Thanks for sharing this.",
      "votes": null
    },
    {
      "id": "1970582",
      "postDate": "10/04/2022 06:41:03",
      "content": "<p>Thanks for your comment <a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a>. It is a good question. The positions of the players are so important especially for defending team because a goalkeeper can change the target. I think we can use machine learning models that ignore the nan features like HistGradientBoostingClassifier or we can impute missing values with two values which are mean coordinates and goalkeeper coordinates then average the preds.</p>",
      "rawMarkdown": "Thanks for your comment @maherelouahabi. It is a good question. The positions of the players are so important especially for defending team because a goalkeeper can change the target. I think we can use machine learning models that ignore the nan features like HistGradientBoostingClassifier or we can impute missing values with two values which are mean coordinates and goalkeeper coordinates then average the preds.",
      "votes": null
    },
    {
      "id": "1970980",
      "postDate": "10/04/2022 11:55:26",
      "content": "<p>Good finding <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a> </p>",
      "rawMarkdown": "Good finding @hasanbasriakcay",
      "votes": null
    },
    {
      "id": "1971112",
      "postDate": "10/04/2022 13:20:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a>, I don't think this is a good imputation method. Why? Because interpolation is implying that the car is moving, it has velocity and it has position.  Null values are the result of a demolition, the player was taken out of the game.</p>\n<p>However; null count is so low % and with the amount of that data we have, I don't think it will be critical to improve the scores. We can discuss about it, in the end this is the place to do so.</p>",
      "rawMarkdown": "Hi @hasanbasriakcay, I don't think this is a good imputation method. Why? Because interpolation is implying that the car is moving, it has velocity and it has position.  Null values are the result of a demolition, the player was taken out of the game.\n\nHowever; null count is so low % and with the amount of that data we have, I don't think it will be critical to improve the scores. We can discuss about it, in the end this is the place to do so.",
      "votes": null
    },
    {
      "id": "1971420",
      "postDate": "10/04/2022 15:39:30",
      "content": "<p>Thanks for your response</p>\n<p>Maher</p>",
      "rawMarkdown": "Thanks for your response\n\nMaher",
      "votes": null
    },
    {
      "id": "1971453",
      "postDate": "10/04/2022 16:13:47",
      "content": "<p>Thanks a lot for sharing! </p>",
      "rawMarkdown": "Thanks a lot for sharing!",
      "votes": null
    },
    {
      "id": "1972450",
      "postDate": "10/05/2022 06:53:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a>, thanks for your kind comment. Panda's interpolation method accepts speed of the car is constant. Of course, acceleration can change the player's position. However, there is no missing value imputer like that so we have to write this function on our own. I think the pandas' interpolate is a good option.</p>\n<p>I agree with you. Missing value percentages are so low therefore, missing values imputing have not much affect o on lb but little effects sometimes can make you a winner.</p>\n<p>I never played Rocket League. When a player was taken out of the game, is the player's car still in the arena or not?</p>",
      "rawMarkdown": "Hi @jcaliz, thanks for your kind comment. Panda's interpolation method accepts speed of the car is constant. Of course, acceleration can change the player's position. However, there is no missing value imputer like that so we have to write this function on our own. I think the pandas' interpolate is a good option.\n\nI agree with you. Missing value percentages are so low therefore, missing values imputing have not much affect o on lb but little effects sometimes can make you a winner.\n\nI never played Rocket League. When a player was taken out of the game, is the player's car still in the arena or not?",
      "votes": null
    },
    {
      "id": "1973025",
      "postDate": "10/05/2022 12:34:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a>, the car is not in the arena is on inter-dimensional state and it respawns after three seconds. O the other hand great to know that pandas interpolate also have constant-filling I'll play around with it.</p>",
      "rawMarkdown": "Hi @hasanbasriakcay, the car is not in the arena is on inter-dimensional state and it respawns after three seconds. O the other hand great to know that pandas interpolate also have constant-filling I'll play around with it.",
      "votes": null
    },
    {
      "id": "1973205",
      "postDate": "10/05/2022 14:16:12",
      "content": "<p>Descriptive and nice explanantion</p>",
      "rawMarkdown": "Descriptive and nice explanantion",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1969421,
      "author_name": "abrafey",
      "author_url": "",
      "post_date": "10/03/2022 13:26:02",
      "content": "<p>Short with proper representation.. please keep it up <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1969499,
      "author_name": "maherelouahabi",
      "author_url": "",
      "post_date": "10/03/2022 14:21:15",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a>,</p>\n<p>Thank you for sharing your Nan values imputation methodology. I agree with you that this imputation strategy is one of the most appropriate to impute players' positions, but it requires to have the chronology of players' positions. We do not have this information in the test set (it's scrambled). Any thoughts about what could be done at the test set?</p>\n<p>Maher</p>",
      "votes": null,
      "replies": [
        {
          "id": 1970582,
          "author_name": "hasanbasriakcay",
          "author_url": "",
          "post_date": "10/04/2022 06:41:03",
          "content": "<p>Thanks for your comment <a href=\"https://www.kaggle.com/maherelouahabi\" target=\"_blank\">@maherelouahabi</a>. It is a good question. The positions of the players are so important especially for defending team because a goalkeeper can change the target. I think we can use machine learning models that ignore the nan features like HistGradientBoostingClassifier or we can impute missing values with two values which are mean coordinates and goalkeeper coordinates then average the preds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1971420,
          "author_name": "maherelouahabi",
          "author_url": "",
          "post_date": "10/04/2022 15:39:30",
          "content": "<p>Thanks for your response</p>\n<p>Maher</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1970178,
      "author_name": "cid007",
      "author_url": "",
      "post_date": "10/03/2022 23:08:35",
      "content": "<p>Nice explanation. Definitely will be useful. Thanks for sharing this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1970980,
      "author_name": "landfallmotto",
      "author_url": "",
      "post_date": "10/04/2022 11:55:26",
      "content": "<p>Good finding <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1971112,
      "author_name": "jcaliz",
      "author_url": "",
      "post_date": "10/04/2022 13:20:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a>, I don't think this is a good imputation method. Why? Because interpolation is implying that the car is moving, it has velocity and it has position.  Null values are the result of a demolition, the player was taken out of the game.</p>\n<p>However; null count is so low % and with the amount of that data we have, I don't think it will be critical to improve the scores. We can discuss about it, in the end this is the place to do so.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1972450,
          "author_name": "hasanbasriakcay",
          "author_url": "",
          "post_date": "10/05/2022 06:53:23",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jcaliz\" target=\"_blank\">@jcaliz</a>, thanks for your kind comment. Panda's interpolation method accepts speed of the car is constant. Of course, acceleration can change the player's position. However, there is no missing value imputer like that so we have to write this function on our own. I think the pandas' interpolate is a good option.</p>\n<p>I agree with you. Missing value percentages are so low therefore, missing values imputing have not much affect o on lb but little effects sometimes can make you a winner.</p>\n<p>I never played Rocket League. When a player was taken out of the game, is the player's car still in the arena or not?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1973025,
          "author_name": "jcaliz",
          "author_url": "",
          "post_date": "10/05/2022 12:34:44",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hasanbasriakcay\" target=\"_blank\">@hasanbasriakcay</a>, the car is not in the arena is on inter-dimensional state and it respawns after three seconds. O the other hand great to know that pandas interpolate also have constant-filling I'll play around with it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1971453,
      "author_name": "akshay9508",
      "author_url": "",
      "post_date": "10/04/2022 16:13:47",
      "content": "<p>Thanks a lot for sharing! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1973205,
      "author_name": "priyanshuk14",
      "author_url": "",
      "post_date": "10/05/2022 14:16:12",
      "content": "<p>Descriptive and nice explanantion</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1969173": "Hi all,\n\nWhen I worked on missing values, I realized that missing values are related to the players. You can see these issues below figure too. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381067%2Fa3e4857a6a5439fcecc63ccbc5806126%2F__results___21_1.png?generation=1664794528222509&alt=media)\n\nPlayer's placement and velocity change over time. Therefore the best missing value handling method is interpolation which is the process of estimating unknown values that fall between known values.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381067%2F475d7352e42c8b48afdcaa41d998d4d5%2FZNERC_r.png?generation=1664795700035054&alt=media)\nsource: [Wikipedia](https://en.wikipedia.org/wiki/Linear_interpolation)\n\n\nYou can use the below code for imputation in numeric features or you can find the clean dataset [here](https://www.kaggle.com/datasets/hasanbasriakcay/tpsoct22-traintest-parquet).\n\n`train_cl =  train.groupby([\"game_num\",\"event_id\"]).apply(lambda group: group.interpolate(method='index', limit_direction='both'))`\n\nKindly upvote if you find it useful 👍",
    "1969421": "Short with proper representation.. please keep it up @hasanbasriakcay",
    "1969499": "Dear @hasanbasriakcay,\n\nThank you for sharing your Nan values imputation methodology. I agree with you that this imputation strategy is one of the most appropriate to impute players' positions, but it requires to have the chronology of players' positions. We do not have this information in the test set (it's scrambled). Any thoughts about what could be done at the test set?\n\nMaher",
    "1970178": "Nice explanation. Definitely will be useful. Thanks for sharing this.",
    "1970582": "Thanks for your comment @maherelouahabi. It is a good question. The positions of the players are so important especially for defending team because a goalkeeper can change the target. I think we can use machine learning models that ignore the nan features like HistGradientBoostingClassifier or we can impute missing values with two values which are mean coordinates and goalkeeper coordinates then average the preds.",
    "1970980": "Good finding @hasanbasriakcay",
    "1971112": "Hi @hasanbasriakcay, I don't think this is a good imputation method. Why? Because interpolation is implying that the car is moving, it has velocity and it has position.  Null values are the result of a demolition, the player was taken out of the game.\n\nHowever; null count is so low % and with the amount of that data we have, I don't think it will be critical to improve the scores. We can discuss about it, in the end this is the place to do so.",
    "1971420": "Thanks for your response\n\nMaher",
    "1971453": "Thanks a lot for sharing!",
    "1972450": "Hi @jcaliz, thanks for your kind comment. Panda's interpolation method accepts speed of the car is constant. Of course, acceleration can change the player's position. However, there is no missing value imputer like that so we have to write this function on our own. I think the pandas' interpolate is a good option.\n\nI agree with you. Missing value percentages are so low therefore, missing values imputing have not much affect o on lb but little effects sometimes can make you a winner.\n\nI never played Rocket League. When a player was taken out of the game, is the player's car still in the arena or not?",
    "1973025": "Hi @hasanbasriakcay, the car is not in the arena is on inter-dimensional state and it respawns after three seconds. O the other hand great to know that pandas interpolate also have constant-filling I'll play around with it.",
    "1973205": "Descriptive and nice explanantion"
  },
  "source": "meta"
}