{
  "id": 44791,
  "title": "count of members vs submission",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/44791",
  "author_name": "",
  "post_date": "2017-12-02T14:43:52.552946900Z",
  "votes": null,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I realize I am trying to jump in to the competition late but I am confused about the number of lines in the members file vs. the number in the sample submission files.\nmembers_v3.csv seems to have over 6 million lines (using wc from cygwin) but the submission files are just under 1 million.  Also the zeros submission file and the other submission file differ in line count. <br>\nCan someone explain this?\nThanks</p>",
  "messages": [
    {
      "id": "252232",
      "postDate": "12/02/2017 14:43:52",
      "content": "<p>I realize I am trying to jump in to the competition late but I am confused about the number of lines in the members file vs. the number in the sample submission files.\nmembers_v3.csv seems to have over 6 million lines (using wc from cygwin) but the submission files are just under 1 million.  Also the zeros submission file and the other submission file differ in line count. <br>\nCan someone explain this?\nThanks</p>",
      "rawMarkdown": "I realize I am trying to jump in to the competition late but I am confused about the number of lines in the members file vs. the number in the sample submission files.\nmembers_v3.csv seems to have over 6 million lines (using wc from cygwin) but the submission files are just under 1 million.  Also the zeros submission file and the other submission file differ in line count.  \nCan someone explain this?\nThanks",
      "votes": null
    },
    {
      "id": "252401",
      "postDate": "12/02/2017 21:02:49",
      "content": "<p>Hi John, regarding your questions, here are my thoughts.  I am also new to the competition. </p>\n\n<p>1) The members_v3.csv has over 6 million unique msno's (customer ID's) so you have to only select the msno's that are used for training and for testing (submission to the contest).</p>\n\n<p>2) The submission_v2.csv file replaces submission_zero.csv</p>",
      "rawMarkdown": "Hi John, regarding your questions, here are my thoughts.  I am also new to the competition. \n\n1) The members_v3.csv has over 6 million unique msno's (customer ID's) so you have to only select the msno's that are used for training and for testing (submission to the contest).\n\n2) The submission_v2.csv file replaces submission_zero.csv",
      "votes": null
    },
    {
      "id": "254078",
      "postDate": "12/06/2017 06:23:35",
      "content": "<p>Submission file contains only the test group. member files contain data for train and test groups</p>",
      "rawMarkdown": "Submission file contains only the test group. member files contain data for train and test groups",
      "votes": null
    },
    {
      "id": "255306",
      "postDate": "12/08/2017 20:18:30",
      "content": "<p>Yes, but to make a submission file we need to do prediction for the list of users that are in the submission file. And the users that are in the submission (some of them) are not included in other sets that we need features like in members (to know bd, gender or whatever) and also in user_logs (to know total, date or whatever). In these case how to do the prediction for them or just to say they churn for sure because they ended up somehow in submission set but not in the member which I think they must be there!  </p>",
      "rawMarkdown": "Yes, but to make a submission file we need to do prediction for the list of users that are in the submission file. And the users that are in the submission (some of them) are not included in other sets that we need features like in members (to know bd, gender or whatever) and also in user_logs (to know total, date or whatever). In these case how to do the prediction for them or just to say they churn for sure because they ended up somehow in submission set but not in the member which I think they must be there!",
      "votes": null
    },
    {
      "id": "255356",
      "postDate": "12/08/2017 21:46:47",
      "content": "<p>The strategy you use to address missing data in the training set is the same one you use for the submission set.  Of course, this approach will not work if you dropped customers with missing data in the training set, since you can't drop customers in the submission set.</p>",
      "rawMarkdown": "The strategy you use to address missing data in the training set is the same one you use for the submission set.  Of course, this approach will not work if you dropped customers with missing data in the training set, since you can't drop customers in the submission set.",
      "votes": null
    },
    {
      "id": "255357",
      "postDate": "12/08/2017 21:50:05",
      "content": "<p>Yes, that's exactly what I'm also saying that we cannot just ignore them. Do you have any idea how to treat them?  </p>",
      "rawMarkdown": "Yes, that's exactly what I'm also saying that we cannot just ignore them. Do you have any idea how to treat them?",
      "votes": null
    },
    {
      "id": "255415",
      "postDate": "12/09/2017 01:13:07",
      "content": "<p>Hi, there are many ways of dealing with the missing values.  Here are a few thoughts.</p>\n\n<p>1) You can use a technique that is robust to missing values.  Decision trees, and all of its variants such as XGBoost, are examples. </p>\n\n<p>2) You can replace the missing values using standard approaches such as median replacement.</p>\n\n<p>3) You can use a dummy variable to represent the missing value.  </p>\n\n<p>4) You can drop variables, such as gender, which have a large amount of missing values.  </p>\n\n<p>and so on.....</p>\n\n<p>Moreover, in the world of machine learning, you don't have to choose a single approach.  You can try 10 different approaches and average the results.  </p>\n\n<p>The one point I would caution you on is making sure you identify what is a true missing value.  You mentioned missing log data.  Is the data missing, or did a user not use the service during the time period that you are looking at?  If it is the latter, then the data is not missing.  </p>",
      "rawMarkdown": "Hi, there are many ways of dealing with the missing values.  Here are a few thoughts.\n\n1) You can use a technique that is robust to missing values.  Decision trees, and all of its variants such as XGBoost, are examples. \n\n2) You can replace the missing values using standard approaches such as median replacement.\n\n3) You can use a dummy variable to represent the missing value.  \n\n4) You can drop variables, such as gender, which have a large amount of missing values.  \n\nand so on.....\n\nMoreover, in the world of machine learning, you don't have to choose a single approach.  You can try 10 different approaches and average the results.  \n\nThe one point I would caution you on is making sure you identify what is a true missing value.  You mentioned missing log data.  Is the data missing, or did a user not use the service during the time period that you are looking at?  If it is the latter, then the data is not missing.",
      "votes": null
    },
    {
      "id": "255514",
      "postDate": "12/09/2017 09:20:13",
      "content": "<p>Hi, I want to be more clear with an example of this data, and I really appreciate what you have describe how to treat missing values.</p>\n\n<p>Here is the example:</p>\n\n<p>Members table looks like (I'll take just some non real example to describe my problem in a simple way):</p>\n\n<pre><code>Msno       bd        ...       registration_init_time\n1          21        ...       20170321\n2          21        ...       20170320\n3          20        ...       20150104\n4          25        ...       20160203\n</code></pre>\n\n<p>Submission table looks like:</p>\n\n<pre><code>msno         is_churn\n1            0\n2            0\n5            0\n6            0\n</code></pre>\n\n<p>So what I can say here it is that when I want to merge members table with submission table on=\"msno\" it reduce the number of instances because msno 5 and 6 doesnt appear in the member table also this happens with other tables. So for example the instance in submission table with msno 5 is missing in all other tables (members, transactions, user_logs).</p>",
      "rawMarkdown": "Hi, I want to be more clear with an example of this data, and I really appreciate what you have describe how to treat missing values.\n\nHere is the example:\n\nMembers table looks like (I'll take just some non real example to describe my problem in a simple way):\n\n    Msno       bd        ...       registration_init_time\n    1          21        ...       20170321\n    2          21        ...       20170320\n    3          20        ...       20150104\n    4          25        ...       20160203\n\n\nSubmission table looks like:\n\n    msno         is_churn\n    1            0\n    2            0\n    5            0\n    6            0\n\nSo what I can say here it is that when I want to merge members table with submission table on=\"msno\" it reduce the number of instances because msno 5 and 6 doesnt appear in the member table also this happens with other tables. So for example the instance in submission table with msno 5 is missing in all other tables (members, transactions, user_logs).",
      "votes": null
    },
    {
      "id": "255525",
      "postDate": "12/09/2017 10:00:32",
      "content": "<p>Hi, in your code you have to specify what kind of merge you want to perform.  </p>\n\n<p>1) A left join will return a data set with only the msno's in the Member file.\n2) A right join will return a data set with only the msno's in the Submission file.\n3) An inner join will return a data set with msno's common to both data sets.\n4) An outer join will return a data set with msno's in either data sets.</p>\n\n<p>Thus, to build a data set for the submission data you want to do (2)</p>",
      "rawMarkdown": "Hi, in your code you have to specify what kind of merge you want to perform.  \n\n1) A left join will return a data set with only the msno's in the Member file.\n2) A right join will return a data set with only the msno's in the Submission file.\n3) An inner join will return a data set with msno's common to both data sets.\n4) An outer join will return a data set with msno's in either data sets.\n\nThus, to build a data set for the submission data you want to do (2)",
      "votes": null
    },
    {
      "id": "255536",
      "postDate": "12/09/2017 10:36:07",
      "content": "<p>Hi, yes I know that I can do right join to get all msno's in the submission set, but I can't specifically for those msno's to get their features for example I want to know bd, gender and so on from the members table and total secs from the user_logs table, so I'm missing these instance for some msno and I think to do the inner join is the right way to get all instances and a complete dataset which is merged with other to do a data integration in order than to do feature engineering. </p>",
      "rawMarkdown": "Hi, yes I know that I can do right join to get all msno's in the submission set, but I can't specifically for those msno's to get their features for example I want to know bd, gender and so on from the members table and total secs from the user_logs table, so I'm missing these instance for some msno and I think to do the inner join is the right way to get all instances and a complete dataset which is merged with other to do a data integration in order than to do feature engineering.",
      "votes": null
    },
    {
      "id": "255674",
      "postDate": "12/09/2017 19:18:55",
      "content": "<p>If you do the right join you will merge in all of the non-missing submission data fields.   For the data that is missing (for example age data for msnos 5 and 6 in your example above) you apply the missing value technique that you used for the training data.  So, as an example, if you decided to replace all the missing ages in the training data with 24, then you do the same for the submission dataset.</p>\n\n<p>Sorry if I am not getting your question.</p>",
      "rawMarkdown": "If you do the right join you will merge in all of the non-missing submission data fields.   For the data that is missing (for example age data for msnos 5 and 6 in your example above) you apply the missing value technique that you used for the training data.  So, as an example, if you decided to replace all the missing ages in the training data with 24, then you do the same for the submission dataset.\n\nSorry if I am not getting your question.",
      "votes": null
    },
    {
      "id": "255678",
      "postDate": "12/09/2017 19:36:52",
      "content": "<p>Yes, I think that is a misunderstanding, but I'll try to describe again what I'm finding a problem. </p>\n\n<p>1) I have to submit the submission set.\n2) This set have two columns (msno and is_churn)\n3) For some of the msno's in the submission set I don't have any features from other datasets that describe anything for those msno's.</p>\n\n<p>Tip: Try to find this user in members_v3 set which is in the submission set with msno: puL2A+Pe6eqOM6D+RKnqmdiJPaWrlKYQcCrJXKCXIIM= </p>\n\n<p>It doesn't appear in members_v3 set but for example my model for prediction need to know registration_init_time features and also it doesn't appear in user_logs_v2.</p>",
      "rawMarkdown": "Yes, I think that is a misunderstanding, but I'll try to describe again what I'm finding a problem. \n\n1) I have to submit the submission set.\n2) This set have two columns (msno and is_churn)\n3) For some of the msno's in the submission set I don't have any features from other datasets that describe anything for those msno's.\n\nTip: Try to find this user in members_v3 set which is in the submission set with msno: puL2A+Pe6eqOM6D+RKnqmdiJPaWrlKYQcCrJXKCXIIM= \n\nIt doesn't appear in members_v3 set but for example my model for prediction need to know registration_init_time features and also it doesn't appear in user_logs_v2.",
      "votes": null
    },
    {
      "id": "255689",
      "postDate": "12/09/2017 20:23:11",
      "content": "<p>Thanks for the detailed explanation.  I now understand your question.  There are two cases.  </p>\n\n<p>Case 1:  Your training set has msno's that don't have any membership, transactions, and log data - in other words, they are missing all data.  If so, then what did you do there?  In my experience, the most reasonable approach is to compute the churn rate for these mono's and used that as your prediction.  That would be what many decision tree's would do.  You then apply the same rule to your submission data.  </p>\n\n<p>Case 2: All of your training data have at least some data, so your training model is encountering a case that it has never seen before.  One choice is to use the overall churn rate on the training data as your prediction.  In many statistical approaches, that is the optimal prediction when you have no predictive information available.  Another choice is to use a Naive Bayes Model.  In your training data set, you probably have some data with missing age data, some data with missing membership data, some data with missing log data, and so on.  The model can then predict the churn rate if all data is missing, even though that exact case does not exist in the training data.  (That is one of the advantages of the Naive Bayes Model.)</p>",
      "rawMarkdown": "Thanks for the detailed explanation.  I now understand your question.  There are two cases.  \n\nCase 1:  Your training set has msno's that don't have any membership, transactions, and log data - in other words, they are missing all data.  If so, then what did you do there?  In my experience, the most reasonable approach is to compute the churn rate for these mono's and used that as your prediction.  That would be what many decision tree's would do.  You then apply the same rule to your submission data.  \n\nCase 2: All of your training data have at least some data, so your training model is encountering a case that it has never seen before.  One choice is to use the overall churn rate on the training data as your prediction.  In many statistical approaches, that is the optimal prediction when you have no predictive information available.  Another choice is to use a Naive Bayes Model.  In your training data set, you probably have some data with missing age data, some data with missing membership data, some data with missing log data, and so on.  The model can then predict the churn rate if all data is missing, even though that exact case does not exist in the training data.  (That is one of the advantages of the Naive Bayes Model.)",
      "votes": null
    },
    {
      "id": "255690",
      "postDate": "12/09/2017 20:27:10",
      "content": "<p>Thanks for your very useful explanation! It's just lovely!  </p>",
      "rawMarkdown": "Thanks for your very useful explanation! It's just lovely!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 252401,
      "author_name": "",
      "author_url": "",
      "post_date": "12/02/2017 21:02:49",
      "content": "<p>Hi John, regarding your questions, here are my thoughts.  I am also new to the competition. </p>\n\n<p>1) The members_v3.csv has over 6 million unique msno's (customer ID's) so you have to only select the msno's that are used for training and for testing (submission to the contest).</p>\n\n<p>2) The submission_v2.csv file replaces submission_zero.csv</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 254078,
      "author_name": "ripcurl",
      "author_url": "",
      "post_date": "12/06/2017 06:23:35",
      "content": "<p>Submission file contains only the test group. member files contain data for train and test groups</p>",
      "votes": null,
      "replies": [
        {
          "id": 255306,
          "author_name": "ademkikaj",
          "author_url": "",
          "post_date": "12/08/2017 20:18:30",
          "content": "<p>Yes, but to make a submission file we need to do prediction for the list of users that are in the submission file. And the users that are in the submission (some of them) are not included in other sets that we need features like in members (to know bd, gender or whatever) and also in user_logs (to know total, date or whatever). In these case how to do the prediction for them or just to say they churn for sure because they ended up somehow in submission set but not in the member which I think they must be there!  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255356,
          "author_name": "",
          "author_url": "",
          "post_date": "12/08/2017 21:46:47",
          "content": "<p>The strategy you use to address missing data in the training set is the same one you use for the submission set.  Of course, this approach will not work if you dropped customers with missing data in the training set, since you can't drop customers in the submission set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255357,
          "author_name": "ademkikaj",
          "author_url": "",
          "post_date": "12/08/2017 21:50:05",
          "content": "<p>Yes, that's exactly what I'm also saying that we cannot just ignore them. Do you have any idea how to treat them?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255415,
          "author_name": "",
          "author_url": "",
          "post_date": "12/09/2017 01:13:07",
          "content": "<p>Hi, there are many ways of dealing with the missing values.  Here are a few thoughts.</p>\n\n<p>1) You can use a technique that is robust to missing values.  Decision trees, and all of its variants such as XGBoost, are examples. </p>\n\n<p>2) You can replace the missing values using standard approaches such as median replacement.</p>\n\n<p>3) You can use a dummy variable to represent the missing value.  </p>\n\n<p>4) You can drop variables, such as gender, which have a large amount of missing values.  </p>\n\n<p>and so on.....</p>\n\n<p>Moreover, in the world of machine learning, you don't have to choose a single approach.  You can try 10 different approaches and average the results.  </p>\n\n<p>The one point I would caution you on is making sure you identify what is a true missing value.  You mentioned missing log data.  Is the data missing, or did a user not use the service during the time period that you are looking at?  If it is the latter, then the data is not missing.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255514,
          "author_name": "ademkikaj",
          "author_url": "",
          "post_date": "12/09/2017 09:20:13",
          "content": "<p>Hi, I want to be more clear with an example of this data, and I really appreciate what you have describe how to treat missing values.</p>\n\n<p>Here is the example:</p>\n\n<p>Members table looks like (I'll take just some non real example to describe my problem in a simple way):</p>\n\n<pre><code>Msno       bd        ...       registration_init_time\n1          21        ...       20170321\n2          21        ...       20170320\n3          20        ...       20150104\n4          25        ...       20160203\n</code></pre>\n\n<p>Submission table looks like:</p>\n\n<pre><code>msno         is_churn\n1            0\n2            0\n5            0\n6            0\n</code></pre>\n\n<p>So what I can say here it is that when I want to merge members table with submission table on=\"msno\" it reduce the number of instances because msno 5 and 6 doesnt appear in the member table also this happens with other tables. So for example the instance in submission table with msno 5 is missing in all other tables (members, transactions, user_logs).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255525,
          "author_name": "",
          "author_url": "",
          "post_date": "12/09/2017 10:00:32",
          "content": "<p>Hi, in your code you have to specify what kind of merge you want to perform.  </p>\n\n<p>1) A left join will return a data set with only the msno's in the Member file.\n2) A right join will return a data set with only the msno's in the Submission file.\n3) An inner join will return a data set with msno's common to both data sets.\n4) An outer join will return a data set with msno's in either data sets.</p>\n\n<p>Thus, to build a data set for the submission data you want to do (2)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255536,
          "author_name": "ademkikaj",
          "author_url": "",
          "post_date": "12/09/2017 10:36:07",
          "content": "<p>Hi, yes I know that I can do right join to get all msno's in the submission set, but I can't specifically for those msno's to get their features for example I want to know bd, gender and so on from the members table and total secs from the user_logs table, so I'm missing these instance for some msno and I think to do the inner join is the right way to get all instances and a complete dataset which is merged with other to do a data integration in order than to do feature engineering. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255674,
          "author_name": "",
          "author_url": "",
          "post_date": "12/09/2017 19:18:55",
          "content": "<p>If you do the right join you will merge in all of the non-missing submission data fields.   For the data that is missing (for example age data for msnos 5 and 6 in your example above) you apply the missing value technique that you used for the training data.  So, as an example, if you decided to replace all the missing ages in the training data with 24, then you do the same for the submission dataset.</p>\n\n<p>Sorry if I am not getting your question.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255678,
          "author_name": "ademkikaj",
          "author_url": "",
          "post_date": "12/09/2017 19:36:52",
          "content": "<p>Yes, I think that is a misunderstanding, but I'll try to describe again what I'm finding a problem. </p>\n\n<p>1) I have to submit the submission set.\n2) This set have two columns (msno and is_churn)\n3) For some of the msno's in the submission set I don't have any features from other datasets that describe anything for those msno's.</p>\n\n<p>Tip: Try to find this user in members_v3 set which is in the submission set with msno: puL2A+Pe6eqOM6D+RKnqmdiJPaWrlKYQcCrJXKCXIIM= </p>\n\n<p>It doesn't appear in members_v3 set but for example my model for prediction need to know registration_init_time features and also it doesn't appear in user_logs_v2.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255689,
          "author_name": "",
          "author_url": "",
          "post_date": "12/09/2017 20:23:11",
          "content": "<p>Thanks for the detailed explanation.  I now understand your question.  There are two cases.  </p>\n\n<p>Case 1:  Your training set has msno's that don't have any membership, transactions, and log data - in other words, they are missing all data.  If so, then what did you do there?  In my experience, the most reasonable approach is to compute the churn rate for these mono's and used that as your prediction.  That would be what many decision tree's would do.  You then apply the same rule to your submission data.  </p>\n\n<p>Case 2: All of your training data have at least some data, so your training model is encountering a case that it has never seen before.  One choice is to use the overall churn rate on the training data as your prediction.  In many statistical approaches, that is the optimal prediction when you have no predictive information available.  Another choice is to use a Naive Bayes Model.  In your training data set, you probably have some data with missing age data, some data with missing membership data, some data with missing log data, and so on.  The model can then predict the churn rate if all data is missing, even though that exact case does not exist in the training data.  (That is one of the advantages of the Naive Bayes Model.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 255690,
          "author_name": "ademkikaj",
          "author_url": "",
          "post_date": "12/09/2017 20:27:10",
          "content": "<p>Thanks for your very useful explanation! It's just lovely!  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "252232": "I realize I am trying to jump in to the competition late but I am confused about the number of lines in the members file vs. the number in the sample submission files.\nmembers_v3.csv seems to have over 6 million lines (using wc from cygwin) but the submission files are just under 1 million.  Also the zeros submission file and the other submission file differ in line count.  \nCan someone explain this?\nThanks",
    "252401": "Hi John, regarding your questions, here are my thoughts.  I am also new to the competition. \n\n1) The members_v3.csv has over 6 million unique msno's (customer ID's) so you have to only select the msno's that are used for training and for testing (submission to the contest).\n\n2) The submission_v2.csv file replaces submission_zero.csv",
    "254078": "Submission file contains only the test group. member files contain data for train and test groups",
    "255306": "Yes, but to make a submission file we need to do prediction for the list of users that are in the submission file. And the users that are in the submission (some of them) are not included in other sets that we need features like in members (to know bd, gender or whatever) and also in user_logs (to know total, date or whatever). In these case how to do the prediction for them or just to say they churn for sure because they ended up somehow in submission set but not in the member which I think they must be there!",
    "255356": "The strategy you use to address missing data in the training set is the same one you use for the submission set.  Of course, this approach will not work if you dropped customers with missing data in the training set, since you can't drop customers in the submission set.",
    "255357": "Yes, that's exactly what I'm also saying that we cannot just ignore them. Do you have any idea how to treat them?",
    "255415": "Hi, there are many ways of dealing with the missing values.  Here are a few thoughts.\n\n1) You can use a technique that is robust to missing values.  Decision trees, and all of its variants such as XGBoost, are examples. \n\n2) You can replace the missing values using standard approaches such as median replacement.\n\n3) You can use a dummy variable to represent the missing value.  \n\n4) You can drop variables, such as gender, which have a large amount of missing values.  \n\nand so on.....\n\nMoreover, in the world of machine learning, you don't have to choose a single approach.  You can try 10 different approaches and average the results.  \n\nThe one point I would caution you on is making sure you identify what is a true missing value.  You mentioned missing log data.  Is the data missing, or did a user not use the service during the time period that you are looking at?  If it is the latter, then the data is not missing.",
    "255514": "Hi, I want to be more clear with an example of this data, and I really appreciate what you have describe how to treat missing values.\n\nHere is the example:\n\nMembers table looks like (I'll take just some non real example to describe my problem in a simple way):\n\n    Msno       bd        ...       registration_init_time\n    1          21        ...       20170321\n    2          21        ...       20170320\n    3          20        ...       20150104\n    4          25        ...       20160203\n\n\nSubmission table looks like:\n\n    msno         is_churn\n    1            0\n    2            0\n    5            0\n    6            0\n\nSo what I can say here it is that when I want to merge members table with submission table on=\"msno\" it reduce the number of instances because msno 5 and 6 doesnt appear in the member table also this happens with other tables. So for example the instance in submission table with msno 5 is missing in all other tables (members, transactions, user_logs).",
    "255525": "Hi, in your code you have to specify what kind of merge you want to perform.  \n\n1) A left join will return a data set with only the msno's in the Member file.\n2) A right join will return a data set with only the msno's in the Submission file.\n3) An inner join will return a data set with msno's common to both data sets.\n4) An outer join will return a data set with msno's in either data sets.\n\nThus, to build a data set for the submission data you want to do (2)",
    "255536": "Hi, yes I know that I can do right join to get all msno's in the submission set, but I can't specifically for those msno's to get their features for example I want to know bd, gender and so on from the members table and total secs from the user_logs table, so I'm missing these instance for some msno and I think to do the inner join is the right way to get all instances and a complete dataset which is merged with other to do a data integration in order than to do feature engineering.",
    "255674": "If you do the right join you will merge in all of the non-missing submission data fields.   For the data that is missing (for example age data for msnos 5 and 6 in your example above) you apply the missing value technique that you used for the training data.  So, as an example, if you decided to replace all the missing ages in the training data with 24, then you do the same for the submission dataset.\n\nSorry if I am not getting your question.",
    "255678": "Yes, I think that is a misunderstanding, but I'll try to describe again what I'm finding a problem. \n\n1) I have to submit the submission set.\n2) This set have two columns (msno and is_churn)\n3) For some of the msno's in the submission set I don't have any features from other datasets that describe anything for those msno's.\n\nTip: Try to find this user in members_v3 set which is in the submission set with msno: puL2A+Pe6eqOM6D+RKnqmdiJPaWrlKYQcCrJXKCXIIM= \n\nIt doesn't appear in members_v3 set but for example my model for prediction need to know registration_init_time features and also it doesn't appear in user_logs_v2.",
    "255689": "Thanks for the detailed explanation.  I now understand your question.  There are two cases.  \n\nCase 1:  Your training set has msno's that don't have any membership, transactions, and log data - in other words, they are missing all data.  If so, then what did you do there?  In my experience, the most reasonable approach is to compute the churn rate for these mono's and used that as your prediction.  That would be what many decision tree's would do.  You then apply the same rule to your submission data.  \n\nCase 2: All of your training data have at least some data, so your training model is encountering a case that it has never seen before.  One choice is to use the overall churn rate on the training data as your prediction.  In many statistical approaches, that is the optimal prediction when you have no predictive information available.  Another choice is to use a Naive Bayes Model.  In your training data set, you probably have some data with missing age data, some data with missing membership data, some data with missing log data, and so on.  The model can then predict the churn rate if all data is missing, even though that exact case does not exist in the training data.  (That is one of the advantages of the Naive Bayes Model.)",
    "255690": "Thanks for your very useful explanation! It's just lovely!"
  },
  "source": "meta"
}