{
  "id": 41025,
  "title": "Missing Members Data",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/41025",
  "author_name": "",
  "post_date": "2017-10-11T17:48:12.549279300Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi! I'm currently compiting on this challenge and came up with the following two problems on the data, both with the same importance:</p>\n\n<p><strong>1. MEMBERS DATASET</strong> is missing a lot of members that are explicitly defined on the train and predict datasets. The train dataset field can be removed and we cold train our model with the members that do exist, but...</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>when it comes to making the predictions, the SCORE will be hard to predict since most of the features are built on the members dataset, and there are a lot of missing members defined on the predict dataset.</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p><strong>2. TRANSACTIONS DATASET</strong> The same is true for the transactions dataset. Features are built using data from this dataset also.</p>\n\n<p>Could anyone comment regarding the missing data on the PREDICT dataset?</p>",
  "messages": [
    {
      "id": "230315",
      "postDate": "10/11/2017 17:48:12",
      "content": "<p>Hi! I'm currently compiting on this challenge and came up with the following two problems on the data, both with the same importance:</p>\n\n<p><strong>1. MEMBERS DATASET</strong> is missing a lot of members that are explicitly defined on the train and predict datasets. The train dataset field can be removed and we cold train our model with the members that do exist, but...</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>when it comes to making the predictions, the SCORE will be hard to predict since most of the features are built on the members dataset, and there are a lot of missing members defined on the predict dataset.</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p><strong>2. TRANSACTIONS DATASET</strong> The same is true for the transactions dataset. Features are built using data from this dataset also.</p>\n\n<p>Could anyone comment regarding the missing data on the PREDICT dataset?</p>",
      "rawMarkdown": "Hi! I'm currently compiting on this challenge and came up with the following two problems on the data, both with the same importance:\n\n **1. MEMBERS DATASET** is missing a lot of members that are explicitly defined on the train and predict datasets. The train dataset field can be removed and we cold train our model with the members that do exist, but...\n&gt;&gt;&gt; when it comes to making the predictions, the SCORE will be hard to predict since most of the features are built on the members dataset, and there are a lot of missing members defined on the predict dataset.\n\n**2. TRANSACTIONS DATASET** The same is true for the transactions dataset. Features are built using data from this dataset also.\n\nCould anyone comment regarding the missing data on the PREDICT dataset?",
      "votes": null
    },
    {
      "id": "230710",
      "postDate": "10/12/2017 15:02:26",
      "content": "<p>Regarding 1 they explicitly say this:\n\"user information. Note that not every user in the dataset is available.\"</p>\n\n<p>You can try creating a simplified model for these users or do some imputation.</p>",
      "rawMarkdown": "Regarding 1 they explicitly say this:\n\"user information. Note that not every user in the dataset is available.\"\n\nYou can try creating a simplified model for these users or do some imputation.",
      "votes": null
    },
    {
      "id": "230818",
      "postDate": "10/12/2017 21:52:02",
      "content": "<p>Thank you. What I'm trying to figure out is how can we achieve a better prediction if they included missing members on the prediction dataset.<br>\nI'm actually imputing data on the missing members, but I just wanted to clarify that this is a key data to have included.</p>\n\n<p>Thanks for the feedback!</p>",
      "rawMarkdown": "Thank you. What I'm trying to figure out is how can we achieve a better prediction if they included missing members on the prediction dataset.<br>\nI'm actually imputing data on the missing members, but I just wanted to clarify that this is a key data to have included.\n\nThanks for the feedback!",
      "votes": null
    },
    {
      "id": "347489",
      "postDate": "06/24/2018 13:53:16",
      "content": "<p>hi even we had this question, can you please tell the approach on how to deal with it</p>",
      "rawMarkdown": "hi even we had this question, can you please tell the approach on how to deal with it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 230710,
      "author_name": "rolandhazy",
      "author_url": "",
      "post_date": "10/12/2017 15:02:26",
      "content": "<p>Regarding 1 they explicitly say this:\n\"user information. Note that not every user in the dataset is available.\"</p>\n\n<p>You can try creating a simplified model for these users or do some imputation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 230818,
          "author_name": "juanumusic",
          "author_url": "",
          "post_date": "10/12/2017 21:52:02",
          "content": "<p>Thank you. What I'm trying to figure out is how can we achieve a better prediction if they included missing members on the prediction dataset.<br>\nI'm actually imputing data on the missing members, but I just wanted to clarify that this is a key data to have included.</p>\n\n<p>Thanks for the feedback!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 347489,
      "author_name": "vegisetti",
      "author_url": "",
      "post_date": "06/24/2018 13:53:16",
      "content": "<p>hi even we had this question, can you please tell the approach on how to deal with it</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "230315": "Hi! I'm currently compiting on this challenge and came up with the following two problems on the data, both with the same importance:\n\n **1. MEMBERS DATASET** is missing a lot of members that are explicitly defined on the train and predict datasets. The train dataset field can be removed and we cold train our model with the members that do exist, but...\n&gt;&gt;&gt; when it comes to making the predictions, the SCORE will be hard to predict since most of the features are built on the members dataset, and there are a lot of missing members defined on the predict dataset.\n\n**2. TRANSACTIONS DATASET** The same is true for the transactions dataset. Features are built using data from this dataset also.\n\nCould anyone comment regarding the missing data on the PREDICT dataset?",
    "230710": "Regarding 1 they explicitly say this:\n\"user information. Note that not every user in the dataset is available.\"\n\nYou can try creating a simplified model for these users or do some imputation.",
    "230818": "Thank you. What I'm trying to figure out is how can we achieve a better prediction if they included missing members on the prediction dataset.<br>\nI'm actually imputing data on the missing members, but I just wanted to clarify that this is a key data to have included.\n\nThanks for the feedback!",
    "347489": "hi even we had this question, can you please tell the approach on how to deal with it"
  },
  "source": "meta"
}