{
  "id": 39962,
  "title": "Understanding the Data and Your Opinion",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/39962",
  "author_name": "",
  "post_date": "2017-09-24T23:50:38.601763400Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Greetings! </p>\n\n<p>I have several questions regarding the <strong>relationships</strong> between data frames that were given for this contest. I have managed to load the data, took a quick peek to see the descriptive statistics of all data frames, created some charts and also read the description many times. However, I still have some things that don't add up completely. </p>\n\n<ol>\n<li><p>Why do both training and test  have only 2 columns each? </p></li>\n<li><p>The training data set has 992931 rows, while test set has 970960 rows; is it common to have a 51:49  data split?  </p></li>\n<li><p>What strategy would you propose to tackle the problem of having your data split on 5 different data frames? </p></li>\n</ol>\n\n<p>Here is what I did partially and plan to do: </p>\n\n<ul>\n<li>load the packages</li>\n<li>load the data of all the data frames</li>\n<li>check the data integrity, meaning are the classes set properly (the example would be if one column has <em>character</em> class, but it should be <em>numeric</em>)</li>\n<li>descriptive statistics for each data frame</li>\n<li><p>univariate and multivariate EDA</p>\n\n<ol><li>At this step I am not quite sure what to do next. Should I somehow use dplyr/tidyr and connect the ID's in the train data frame with some columns in the member, user_logs and transaction data frames? </li></ol></li>\n</ul>\n\n<p>Most of the time I have worked on relatively clean and straight forward data sets, where I didn't have these challenges. And now, I am going above my comfort zone. I hope the questions aren't too basic, it's just I am new to Machine Learning and eager to learn more. </p>\n\n<p>Please do share your opinions and thoughts. I am most definitely interested what you have to say. </p>\n\n<p>Luka </p>",
  "messages": [
    {
      "id": "224054",
      "postDate": "09/24/2017 23:50:38",
      "content": "<p>Greetings! </p>\n\n<p>I have several questions regarding the <strong>relationships</strong> between data frames that were given for this contest. I have managed to load the data, took a quick peek to see the descriptive statistics of all data frames, created some charts and also read the description many times. However, I still have some things that don't add up completely. </p>\n\n<ol>\n<li><p>Why do both training and test  have only 2 columns each? </p></li>\n<li><p>The training data set has 992931 rows, while test set has 970960 rows; is it common to have a 51:49  data split?  </p></li>\n<li><p>What strategy would you propose to tackle the problem of having your data split on 5 different data frames? </p></li>\n</ol>\n\n<p>Here is what I did partially and plan to do: </p>\n\n<ul>\n<li>load the packages</li>\n<li>load the data of all the data frames</li>\n<li>check the data integrity, meaning are the classes set properly (the example would be if one column has <em>character</em> class, but it should be <em>numeric</em>)</li>\n<li>descriptive statistics for each data frame</li>\n<li><p>univariate and multivariate EDA</p>\n\n<ol><li>At this step I am not quite sure what to do next. Should I somehow use dplyr/tidyr and connect the ID's in the train data frame with some columns in the member, user_logs and transaction data frames? </li></ol></li>\n</ul>\n\n<p>Most of the time I have worked on relatively clean and straight forward data sets, where I didn't have these challenges. And now, I am going above my comfort zone. I hope the questions aren't too basic, it's just I am new to Machine Learning and eager to learn more. </p>\n\n<p>Please do share your opinions and thoughts. I am most definitely interested what you have to say. </p>\n\n<p>Luka </p>",
      "rawMarkdown": "Greetings! \n\nI have several questions regarding the **relationships** between data frames that were given for this contest. I have managed to load the data, took a quick peek to see the descriptive statistics of all data frames, created some charts and also read the description many times. However, I still have some things that don't add up completely. \n\n1.  Why do both training and test  have only 2 columns each? \n\n2. The training data set has 992931 rows, while test set has 970960 rows; is it common to have a 51:49  data split?  \n\n3. What strategy would you propose to tackle the problem of having your data split on 5 different data frames? \n\nHere is what I did partially and plan to do: \n\n- load the packages\n- load the data of all the data frames\n- check the data integrity, meaning are the classes set properly (the example would be if one column has *character* class, but it should be *numeric*)\n- descriptive statistics for each data frame\n- univariate and multivariate EDA\n\n4. At this step I am not quite sure what to do next. Should I somehow use dplyr/tidyr and connect the ID's in the train data frame with some columns in the member, user_logs and transaction data frames? \n\nMost of the time I have worked on relatively clean and straight forward data sets, where I didn't have these challenges. And now, I am going above my comfort zone. I hope the questions aren't too basic, it's just I am new to Machine Learning and eager to learn more. \n\nPlease do share your opinions and thoughts. I am most definitely interested what you have to say. \n\nLuka",
      "votes": null
    },
    {
      "id": "224280",
      "postDate": "09/25/2017 19:14:11",
      "content": "<p>Using dplyr sounds like a good idea.</p>\n\n<p>Have a look at some of the kernels from the recently finished Instacart Market Basket Competition, this should give you an idea about joining, cleaning, feature engineering, efficient handling of memory issues and similar. </p>",
      "rawMarkdown": "Using dplyr sounds like a good idea.\n\nHave a look at some of the kernels from the recently finished Instacart Market Basket Competition, this should give you an idea about joining, cleaning, feature engineering, efficient handling of memory issues and similar.",
      "votes": null
    },
    {
      "id": "224286",
      "postDate": "09/25/2017 19:55:25",
      "content": "<ol>\n<li>Both of these datasets contain two columns: an identifier column (msno) and the target column (is_churn). You'll have to add your own features using the 3 other datasets, by joining them.</li>\n<li>The train dataset contains data from February 2017, the test dataset contains data from March 2017. The 51:49 split indicates that the amount of users whose subscription expires in February is about the same as the amount of users whose subscription expires in March, which makes sense since KKbox's subscription length is generally 30 days. In other competitions, the training set is generally much bigger than the test set, but due to the nature of this competition these sets are about the same size.</li>\n<li>As mentioned in the first point, you'll have to merge your datasets in order to create predictions. You can join the datasets on the msno-column, which exists in each of the 5 datasets. As I see you are using R, have a look at its <a href=\"https://stat.ethz.ch/R-manual/R-devel/library/base/html/merge.html\">merge-function</a>. I also published a <a href=\"https://www.kaggle.com/kevinbonnes/r-churn-prediction-baseline\">kernel</a> in which I joined some of the datasets, feel free to use any of the code for your own predictions.</li>\n</ol>",
      "rawMarkdown": "1. Both of these datasets contain two columns: an identifier column (msno) and the target column (is_churn). You'll have to add your own features using the 3 other datasets, by joining them.\n 2. The train dataset contains data from February 2017, the test dataset contains data from March 2017. The 51:49 split indicates that the amount of users whose subscription expires in February is about the same as the amount of users whose subscription expires in March, which makes sense since KKbox's subscription length is generally 30 days. In other competitions, the training set is generally much bigger than the test set, but due to the nature of this competition these sets are about the same size.\n 3. As mentioned in the first point, you'll have to merge your datasets in order to create predictions. You can join the datasets on the msno-column, which exists in each of the 5 datasets. As I see you are using R, have a look at its [merge-function][1]. I also published a [kernel][2] in which I joined some of the datasets, feel free to use any of the code for your own predictions.\n\n\n  [1]: https://stat.ethz.ch/R-manual/R-devel/library/base/html/merge.html\n  [2]: https://www.kaggle.com/kevinbonnes/r-churn-prediction-baseline",
      "votes": null
    },
    {
      "id": "224551",
      "postDate": "09/26/2017 18:41:07",
      "content": "<p>First of all, thank you for your input. </p>\n\n<ol>\n<li><p>Yes that makes sense now, since the size of the overall data is relatively large. Now I understand the logic behind the the overall division of the data sets. A person who organized the data sets gave us a wide range of options to tackle the prediction. That's nice. </p></li>\n<li><p>Well said. </p></li>\n<li><p>I saw your work and if I am not mistaken it was one the first that were published in Kernals for this project, at least regarding R Kernals. At first, I was drawn to see what you did in detail, but then I stopped my self and promised my self that I'll first tackle the problem on my own and then look at yours in more detail. I will most certainly give you feedback as soon as I go through it.   </p></li>\n</ol>",
      "rawMarkdown": "First of all, thank you for your input. \n\n1. Yes that makes sense now, since the size of the overall data is relatively large. Now I understand the logic behind the the overall division of the data sets. A person who organized the data sets gave us a wide range of options to tackle the prediction. That's nice. \n\n2. Well said. \n\n3. I saw your work and if I am not mistaken it was one the first that were published in Kernals for this project, at least regarding R Kernals. At first, I was drawn to see what you did in detail, but then I stopped my self and promised my self that I'll first tackle the problem on my own and then look at yours in more detail. I will most certainly give you feedback as soon as I go through it.",
      "votes": null
    },
    {
      "id": "224555",
      "postDate": "09/26/2017 18:48:36",
      "content": "<p>Thanks for the heads up, I'll definitely take a look at the Market basket Competition. </p>\n\n<p>I am intrigued by the whole tidyverse and it sparked my imagination. Especially this video: <a href=\"https://www.youtube.com/watch?v=rz3_FDVt9eg\">https://www.youtube.com/watch?v=rz3_FDVt9eg</a></p>\n\n<ol>\n<li>it talks about tidyr and how that package enables you to nest data frames, where one column is a list of data frames, quite interesting. </li>\n<li>Using purr instead of for loops, for easier data manipulation. </li>\n<li>Cleaning models with broom, before using ggplot visualization package. It just makes life easier. </li>\n</ol>",
      "rawMarkdown": "Thanks for the heads up, I'll definitely take a look at the Market basket Competition. \n\nI am intrigued by the whole tidyverse and it sparked my imagination. Especially this video: https://www.youtube.com/watch?v=rz3_FDVt9eg\n\n1. it talks about tidyr and how that package enables you to nest data frames, where one column is a list of data frames, quite interesting. \n2. Using purr instead of for loops, for easier data manipulation. \n3. Cleaning models with broom, before using ggplot visualization package. It just makes life easier.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 224280,
      "author_name": "aljaz91",
      "author_url": "",
      "post_date": "09/25/2017 19:14:11",
      "content": "<p>Using dplyr sounds like a good idea.</p>\n\n<p>Have a look at some of the kernels from the recently finished Instacart Market Basket Competition, this should give you an idea about joining, cleaning, feature engineering, efficient handling of memory issues and similar. </p>",
      "votes": null,
      "replies": [
        {
          "id": 224555,
          "author_name": "loncar5",
          "author_url": "",
          "post_date": "09/26/2017 18:48:36",
          "content": "<p>Thanks for the heads up, I'll definitely take a look at the Market basket Competition. </p>\n\n<p>I am intrigued by the whole tidyverse and it sparked my imagination. Especially this video: <a href=\"https://www.youtube.com/watch?v=rz3_FDVt9eg\">https://www.youtube.com/watch?v=rz3_FDVt9eg</a></p>\n\n<ol>\n<li>it talks about tidyr and how that package enables you to nest data frames, where one column is a list of data frames, quite interesting. </li>\n<li>Using purr instead of for loops, for easier data manipulation. </li>\n<li>Cleaning models with broom, before using ggplot visualization package. It just makes life easier. </li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 224286,
      "author_name": "kevinbonnes",
      "author_url": "",
      "post_date": "09/25/2017 19:55:25",
      "content": "<ol>\n<li>Both of these datasets contain two columns: an identifier column (msno) and the target column (is_churn). You'll have to add your own features using the 3 other datasets, by joining them.</li>\n<li>The train dataset contains data from February 2017, the test dataset contains data from March 2017. The 51:49 split indicates that the amount of users whose subscription expires in February is about the same as the amount of users whose subscription expires in March, which makes sense since KKbox's subscription length is generally 30 days. In other competitions, the training set is generally much bigger than the test set, but due to the nature of this competition these sets are about the same size.</li>\n<li>As mentioned in the first point, you'll have to merge your datasets in order to create predictions. You can join the datasets on the msno-column, which exists in each of the 5 datasets. As I see you are using R, have a look at its <a href=\"https://stat.ethz.ch/R-manual/R-devel/library/base/html/merge.html\">merge-function</a>. I also published a <a href=\"https://www.kaggle.com/kevinbonnes/r-churn-prediction-baseline\">kernel</a> in which I joined some of the datasets, feel free to use any of the code for your own predictions.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 224551,
          "author_name": "loncar5",
          "author_url": "",
          "post_date": "09/26/2017 18:41:07",
          "content": "<p>First of all, thank you for your input. </p>\n\n<ol>\n<li><p>Yes that makes sense now, since the size of the overall data is relatively large. Now I understand the logic behind the the overall division of the data sets. A person who organized the data sets gave us a wide range of options to tackle the prediction. That's nice. </p></li>\n<li><p>Well said. </p></li>\n<li><p>I saw your work and if I am not mistaken it was one the first that were published in Kernals for this project, at least regarding R Kernals. At first, I was drawn to see what you did in detail, but then I stopped my self and promised my self that I'll first tackle the problem on my own and then look at yours in more detail. I will most certainly give you feedback as soon as I go through it.   </p></li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "224054": "Greetings! \n\nI have several questions regarding the **relationships** between data frames that were given for this contest. I have managed to load the data, took a quick peek to see the descriptive statistics of all data frames, created some charts and also read the description many times. However, I still have some things that don't add up completely. \n\n1.  Why do both training and test  have only 2 columns each? \n\n2. The training data set has 992931 rows, while test set has 970960 rows; is it common to have a 51:49  data split?  \n\n3. What strategy would you propose to tackle the problem of having your data split on 5 different data frames? \n\nHere is what I did partially and plan to do: \n\n- load the packages\n- load the data of all the data frames\n- check the data integrity, meaning are the classes set properly (the example would be if one column has *character* class, but it should be *numeric*)\n- descriptive statistics for each data frame\n- univariate and multivariate EDA\n\n4. At this step I am not quite sure what to do next. Should I somehow use dplyr/tidyr and connect the ID's in the train data frame with some columns in the member, user_logs and transaction data frames? \n\nMost of the time I have worked on relatively clean and straight forward data sets, where I didn't have these challenges. And now, I am going above my comfort zone. I hope the questions aren't too basic, it's just I am new to Machine Learning and eager to learn more. \n\nPlease do share your opinions and thoughts. I am most definitely interested what you have to say. \n\nLuka",
    "224280": "Using dplyr sounds like a good idea.\n\nHave a look at some of the kernels from the recently finished Instacart Market Basket Competition, this should give you an idea about joining, cleaning, feature engineering, efficient handling of memory issues and similar.",
    "224286": "1. Both of these datasets contain two columns: an identifier column (msno) and the target column (is_churn). You'll have to add your own features using the 3 other datasets, by joining them.\n 2. The train dataset contains data from February 2017, the test dataset contains data from March 2017. The 51:49 split indicates that the amount of users whose subscription expires in February is about the same as the amount of users whose subscription expires in March, which makes sense since KKbox's subscription length is generally 30 days. In other competitions, the training set is generally much bigger than the test set, but due to the nature of this competition these sets are about the same size.\n 3. As mentioned in the first point, you'll have to merge your datasets in order to create predictions. You can join the datasets on the msno-column, which exists in each of the 5 datasets. As I see you are using R, have a look at its [merge-function][1]. I also published a [kernel][2] in which I joined some of the datasets, feel free to use any of the code for your own predictions.\n\n\n  [1]: https://stat.ethz.ch/R-manual/R-devel/library/base/html/merge.html\n  [2]: https://www.kaggle.com/kevinbonnes/r-churn-prediction-baseline",
    "224551": "First of all, thank you for your input. \n\n1. Yes that makes sense now, since the size of the overall data is relatively large. Now I understand the logic behind the the overall division of the data sets. A person who organized the data sets gave us a wide range of options to tackle the prediction. That's nice. \n\n2. Well said. \n\n3. I saw your work and if I am not mistaken it was one the first that were published in Kernals for this project, at least regarding R Kernals. At first, I was drawn to see what you did in detail, but then I stopped my self and promised my self that I'll first tackle the problem on my own and then look at yours in more detail. I will most certainly give you feedback as soon as I go through it.",
    "224555": "Thanks for the heads up, I'll definitely take a look at the Market basket Competition. \n\nI am intrigued by the whole tidyverse and it sparked my imagination. Especially this video: https://www.youtube.com/watch?v=rz3_FDVt9eg\n\n1. it talks about tidyr and how that package enables you to nest data frames, where one column is a list of data frames, quite interesting. \n2. Using purr instead of for loops, for easier data manipulation. \n3. Cleaning models with broom, before using ggplot visualization package. It just makes life easier."
  },
  "source": "meta"
}