{
  "id": 45047,
  "title": "Multiple sets (tables)",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/45047",
  "author_name": "",
  "post_date": "2017-12-05T20:51:06.778135700Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Since there is more than a set (table) with data should in the preprocessing step make a data integration between sets (tables) in order to have them all together or there is another specific way how to work with separated sets (tables)? I would like to have some suggestions about this topic or if anyone has already solved it and is working on missing values or other steps in the Data Cleaning or Data Preprocessing step. </p>",
  "messages": [
    {
      "id": "253907",
      "postDate": "12/05/2017 20:51:06",
      "content": "<p>Since there is more than a set (table) with data should in the preprocessing step make a data integration between sets (tables) in order to have them all together or there is another specific way how to work with separated sets (tables)? I would like to have some suggestions about this topic or if anyone has already solved it and is working on missing values or other steps in the Data Cleaning or Data Preprocessing step. </p>",
      "rawMarkdown": "Since there is more than a set (table) with data should in the preprocessing step make a data integration between sets (tables) in order to have them all together or there is another specific way how to work with separated sets (tables)? I would like to have some suggestions about this topic or if anyone has already solved it and is working on missing values or other steps in the Data Cleaning or Data Preprocessing step.",
      "votes": null
    },
    {
      "id": "254747",
      "postDate": "12/07/2017 14:38:28",
      "content": "<p>I've found that whether it is good to join the sets depends on which features you are training.  I.e.-some features are more sensitive to seasonality/temporal elements (and other elements such as maybe how the data was labeled in each set), and others are not.  So the answer to you question is you need have at least three training sets:  train1, train2, and train_combined, and then train different sets of features/models on each one, then combine the results.  How you combine them should be based on CV trial and error.</p>\n\n<p>If you really want to get advanced, you can then create more training data from the transactions.csv file using the Scala labeller that was provided, or in a similar script of your choice (SQL, Python, etc.).  Then the sky is the limit on your training sets and model/feature combinations, and you have LOTS of trial and error to do :)   </p>\n\n<p>Best of luck!</p>",
      "rawMarkdown": "I've found that whether it is good to join the sets depends on which features you are training.  I.e.-some features are more sensitive to seasonality/temporal elements (and other elements such as maybe how the data was labeled in each set), and others are not.  So the answer to you question is you need have at least three training sets:  train1, train2, and train_combined, and then train different sets of features/models on each one, then combine the results.  How you combine them should be based on CV trial and error.\n\nIf you really want to get advanced, you can then create more training data from the transactions.csv file using the Scala labeller that was provided, or in a similar script of your choice (SQL, Python, etc.).  Then the sky is the limit on your training sets and model/feature combinations, and you have LOTS of trial and error to do :)   \n\nBest of luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 254747,
      "author_name": "bryangregory",
      "author_url": "",
      "post_date": "12/07/2017 14:38:28",
      "content": "<p>I've found that whether it is good to join the sets depends on which features you are training.  I.e.-some features are more sensitive to seasonality/temporal elements (and other elements such as maybe how the data was labeled in each set), and others are not.  So the answer to you question is you need have at least three training sets:  train1, train2, and train_combined, and then train different sets of features/models on each one, then combine the results.  How you combine them should be based on CV trial and error.</p>\n\n<p>If you really want to get advanced, you can then create more training data from the transactions.csv file using the Scala labeller that was provided, or in a similar script of your choice (SQL, Python, etc.).  Then the sky is the limit on your training sets and model/feature combinations, and you have LOTS of trial and error to do :)   </p>\n\n<p>Best of luck!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "253907": "Since there is more than a set (table) with data should in the preprocessing step make a data integration between sets (tables) in order to have them all together or there is another specific way how to work with separated sets (tables)? I would like to have some suggestions about this topic or if anyone has already solved it and is working on missing values or other steps in the Data Cleaning or Data Preprocessing step.",
    "254747": "I've found that whether it is good to join the sets depends on which features you are training.  I.e.-some features are more sensitive to seasonality/temporal elements (and other elements such as maybe how the data was labeled in each set), and others are not.  So the answer to you question is you need have at least three training sets:  train1, train2, and train_combined, and then train different sets of features/models on each one, then combine the results.  How you combine them should be based on CV trial and error.\n\nIf you really want to get advanced, you can then create more training data from the transactions.csv file using the Scala labeller that was provided, or in a similar script of your choice (SQL, Python, etc.).  Then the sky is the limit on your training sets and model/feature combinations, and you have LOTS of trial and error to do :)   \n\nBest of luck!"
  },
  "source": "meta"
}