{
  "id": 356530,
  "title": "Welcome to the Rocket League Tabular Playground!",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/356530",
  "author_name": "",
  "post_date": "2022-09-30T23:52:03.880670900Z",
  "votes": 33,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hey there, my name is DJ and normally I'm a Software Engineer here at Kaggle, but this month I'll be your Data Scientist* for the Tabular Playground competition!  I'm a longtime Rocket League player and I'm excited to finally bring it to Kaggle -- I think this is a really interesting dataset with a noisy target and a lot of subtle signal to extract from the features.</p>\n<p>Some ideas:</p>\n<ol>\n<li>You could shuffle the training data and train a model as if the rows were independent, or you could try to use the timeseries information, for example with a recurrent network.</li>\n<li>Be careful with <code>NaN</code> values on the player columns when they get demolished -- this gives the other team an advantage.</li>\n<li>You can use data augmentation to get even more training data.</li>\n<li>Check out videos of professional Rocket League matches to learn more about the sport.</li>\n<li>Our <a href=\"https://www.kaggle.com/c/nfl-big-data-bowl-2020\" target=\"_blank\">2020 NFL Big Data Bowl competition</a> had a similar setup with predicting how a sports play would unfold based on a snapshot in time, there could be useful hints from the discussions there.</li>\n</ol>\n<p>Let me know if you have any questions or feedback.  Have fun and see you on the leaderboard!</p>\n<p>*<em>Not a real Data Scientist</em></p>",
  "messages": [
    {
      "id": "1964853",
      "postDate": "09/30/2022 23:52:03",
      "content": "<p>Hey there, my name is DJ and normally I'm a Software Engineer here at Kaggle, but this month I'll be your Data Scientist* for the Tabular Playground competition!  I'm a longtime Rocket League player and I'm excited to finally bring it to Kaggle -- I think this is a really interesting dataset with a noisy target and a lot of subtle signal to extract from the features.</p>\n<p>Some ideas:</p>\n<ol>\n<li>You could shuffle the training data and train a model as if the rows were independent, or you could try to use the timeseries information, for example with a recurrent network.</li>\n<li>Be careful with <code>NaN</code> values on the player columns when they get demolished -- this gives the other team an advantage.</li>\n<li>You can use data augmentation to get even more training data.</li>\n<li>Check out videos of professional Rocket League matches to learn more about the sport.</li>\n<li>Our <a href=\"https://www.kaggle.com/c/nfl-big-data-bowl-2020\" target=\"_blank\">2020 NFL Big Data Bowl competition</a> had a similar setup with predicting how a sports play would unfold based on a snapshot in time, there could be useful hints from the discussions there.</li>\n</ol>\n<p>Let me know if you have any questions or feedback.  Have fun and see you on the leaderboard!</p>\n<p>*<em>Not a real Data Scientist</em></p>",
      "rawMarkdown": "Hey there, my name is DJ and normally I'm a Software Engineer here at Kaggle, but this month I'll be your Data Scientist\\* for the Tabular Playground competition!  I'm a longtime Rocket League player and I'm excited to finally bring it to Kaggle -- I think this is a really interesting dataset with a noisy target and a lot of subtle signal to extract from the features.\n\nSome ideas:\n\n1. You could shuffle the training data and train a model as if the rows were independent, or you could try to use the timeseries information, for example with a recurrent network.\n2. Be careful with `NaN` values on the player columns when they get demolished -- this gives the other team an advantage.\n3. You can use data augmentation to get even more training data.\n4. Check out videos of professional Rocket League matches to learn more about the sport.\n5. Our [2020 NFL Big Data Bowl competition](https://www.kaggle.com/c/nfl-big-data-bowl-2020) had a similar setup with predicting how a sports play would unfold based on a snapshot in time, there could be useful hints from the discussions there.\n\nLet me know if you have any questions or feedback.  Have fun and see you on the leaderboard!\n\n\\**Not a real Data Scientist*",
      "votes": null
    },
    {
      "id": "1970890",
      "postDate": "10/04/2022 10:46:15",
      "content": "<p>this was helpful, thank you</p>",
      "rawMarkdown": "this was helpful, thank you",
      "votes": null
    },
    {
      "id": "1970892",
      "postDate": "10/04/2022 10:47:05",
      "content": "<p>thank you for this</p>",
      "rawMarkdown": "thank you for this",
      "votes": null
    },
    {
      "id": "1976947",
      "postDate": "10/07/2022 16:54:36",
      "content": "<p>I have a question!!!<br>\nIt says <strong><em>\"test.csv: Test set. Unlike the train set, the rows are scrambled.\"</em></strong><br>\nwhat does it mean, does is mean that all the rows in test set are shuffled so that each row can belongs to any game, any event randomly without the time sequence preserved?<br>\nthe 'event_time' is not given in the test set, and we are expected to predict team_[A|B]_scoring_within_10sec, how do we know at what point does the last 10sec begins without having 'event_time' in test set, because I assume the predictions for both teams should be 0 (or close to 0) before 10sec begin as it is in the train set</p>",
      "rawMarkdown": "I have a question!!!\nIt says ***\"test.csv: Test set. Unlike the train set, the rows are scrambled.\"***\nwhat does it mean, does is mean that all the rows in test set are shuffled so that each row can belongs to any game, any event randomly without the time sequence preserved?\nthe 'event_time' is not given in the test set, and we are expected to predict team_[A|B]_scoring_within_10sec, how do we know at what point does the last 10sec begins without having 'event_time' in test set, because I assume the predictions for both teams should be 0 (or close to 0) before 10sec begin as it is in the train set",
      "votes": null
    },
    {
      "id": "1981886",
      "postDate": "10/11/2022 05:57:58",
      "content": "<p>There is showimg memory error in my jupyter nootbook while data visualizing due to large data.is there any solution?</p>",
      "rawMarkdown": "There is showimg memory error in my jupyter nootbook while data visualizing due to large data.is there any solution?",
      "votes": null
    },
    {
      "id": "1982597",
      "postDate": "10/11/2022 14:09:30",
      "content": "<p>Hey Charitha, yes we have removed the time sequence information from the test set, so that predictions are made only using game state information.  The time sequence information, such as <code>event_time</code>, is provided because it could be useful in creating a better model, even though that information is not available at inference time.  You are correct that in the majority of cases the predictions for both teams will be close to 0, since goals are not scored very often.</p>",
      "rawMarkdown": "Hey Charitha, yes we have removed the time sequence information from the test set, so that predictions are made only using game state information.  The time sequence information, such as `event_time`, is provided because it could be useful in creating a better model, even though that information is not available at inference time.  You are correct that in the majority of cases the predictions for both teams will be close to 0, since goals are not scored very often.",
      "votes": null
    },
    {
      "id": "1982600",
      "postDate": "10/11/2022 14:11:26",
      "content": "<p>Hey there, many Kagglers have shared ways to reduce memory consumption in Notebooks when reading this competition's dataset, check out the Discussion tab and Code tab to learn more.</p>",
      "rawMarkdown": "Hey there, many Kagglers have shared ways to reduce memory consumption in Notebooks when reading this competition's dataset, check out the Discussion tab and Code tab to learn more.",
      "votes": null
    },
    {
      "id": "1995358",
      "postDate": "10/19/2022 15:50:17",
      "content": "<p>HI, quick question : <br>\nI do not understand why most columns are not present in the test set. <br>\nWhat is the point of integrating these columns in my model if I am not able to use the model on the test set ? <br>\nThank you very much :)</p>",
      "rawMarkdown": "HI, quick question : \nI do not understand why most columns are not present in the test set. \nWhat is the point of integrating these columns in my model if I am not able to use the model on the test set ? \nThank you very much :)",
      "votes": null
    },
    {
      "id": "1995624",
      "postDate": "10/19/2022 17:58:07",
      "content": "<p>Hey Valentin,<br>\nBesides the target columns, the only columns in train but not in test are <code>game_num</code>, <code>event_id</code>, and <code>event_time</code>, which are just metadata providing more information about rows which come from the same timeseries, but do not provide any in-game information about the game state, which is what we want to make predictions based on.  Personally I am not using those columns in my own model, except for train / validation split, but I could imagine that the timeseries correlations could help build a smarter model, for example using a Recurrent Neural Network.  I'd like to explore that, though it's possible it does not end up being helpful.</p>",
      "rawMarkdown": "Hey Valentin,\nBesides the target columns, the only columns in train but not in test are `game_num`, `event_id`, and `event_time`, which are just metadata providing more information about rows which come from the same timeseries, but do not provide any in-game information about the game state, which is what we want to make predictions based on.  Personally I am not using those columns in my own model, except for train / validation split, but I could imagine that the timeseries correlations could help build a smarter model, for example using a Recurrent Neural Network.  I'd like to explore that, though it's possible it does not end up being helpful.",
      "votes": null
    },
    {
      "id": "1997693",
      "postDate": "10/21/2022 05:27:26",
      "content": "<p>I am beginner here and  I have some  Questions:<br>\nwhy train data set is from train_0 to train_9 ?<br>\ncan I train one of these whole data set like train_0 enough !!</p>",
      "rawMarkdown": "I am beginner here and  I have some  Questions:\nwhy train data set is from train_0 to train_9 ?\ncan I train one of these whole data set like train_0 enough !!",
      "votes": null
    },
    {
      "id": "1998429",
      "postDate": "10/21/2022 15:46:44",
      "content": "<p>Hey there, the full train set is quite large, so we split it into 10 files.  You can probably get a good result using only <code>train_0.csv</code>, but generally it will be better to utilize all the training files.  One strategy would be to use only <code>train_0.csv</code> and get started with that, and once you have built a good pipeline, you can consider adding more train files, or utilizing data augmentation as others have mentioned in the forums, etc.</p>",
      "rawMarkdown": "Hey there, the full train set is quite large, so we split it into 10 files.  You can probably get a good result using only `train_0.csv`, but generally it will be better to utilize all the training files.  One strategy would be to use only `train_0.csv` and get started with that, and once you have built a good pipeline, you can consider adding more train files, or utilizing data augmentation as others have mentioned in the forums, etc.",
      "votes": null
    },
    {
      "id": "2000254",
      "postDate": "10/23/2022 07:10:28",
      "content": "<p>Thank you for your reply ! <br>\nI just saw that I did not read the test.csv correctly ! <br>\nGood luck !!</p>",
      "rawMarkdown": "Thank you for your reply ! \nI just saw that I did not read the test.csv correctly ! \nGood luck !!",
      "votes": null
    },
    {
      "id": "2000327",
      "postDate": "10/23/2022 07:38:37",
      "content": "<p>Hi DJ Sterling I hope u r well<br>\nwhat benefit of the feature in training set only &amp; If it  bad choice if I excepted this data in model?<br>\nI beginner here and I'm I do not want be stuck very much here</p>",
      "rawMarkdown": "Hi DJ Sterling I hope u r well\nwhat benefit of the feature in training set only & If it  bad choice if I excepted this data in model?\nI beginner here and I'm I do not want be stuck very much here",
      "votes": null
    },
    {
      "id": "2002191",
      "postDate": "10/24/2022 15:21:28",
      "content": "<p>Hey, there are three non-target columns which are in train and not test -- <code>game_num</code>, <code>event_id</code>, and <code>event_time</code> -- and yes, you can ignore this data in the train set and still produce a good model.</p>",
      "rawMarkdown": "Hey, there are three non-target columns which are in train and not test -- `game_num`, `event_id`, and `event_time` -- and yes, you can ignore this data in the train set and still produce a good model.",
      "votes": null
    },
    {
      "id": "2003742",
      "postDate": "10/25/2022 18:46:40",
      "content": "<p>As I am new to the data science field , this is my second problem statement on kaggle .I have some questions<br>\n1-How to deal with train_0 to train_9 and make it one training dataset?<br>\n2-some features are missing in testing dataset ,how to deal with it?</p>",
      "rawMarkdown": "As I am new to the data science field , this is my second problem statement on kaggle .I have some questions\n1-How to deal with train_0 to train_9 and make it one training dataset?\n2-some features are missing in testing dataset ,how to deal with it?",
      "votes": null
    },
    {
      "id": "2003914",
      "postDate": "10/25/2022 21:41:01",
      "content": "<p>Hey Alika,</p>\n<ol>\n<li>Most data frameworks will let you concatenate datasets with the same schema, for example with <code>import pandas as pd</code> in Python you can use <code>pd.read_csv()</code> to read each individual file, and <code>pd.concat()</code> to combine them together.  This may fail if you try to combine all the files at once and your machine doesn't have enough RAM -- there are many good ideas in the Discussions tab to reduce RAM cost and combine the data.</li>\n<li>The columns which are in the train set but not in the test set are metadata which should not be part of your model, but could be useful when deciding how to train your model.  The simplest approach is to remove those columns and shuffle the order of the rows in your train set before training your model.</li>\n</ol>\n<p>Hope this helps!</p>",
      "rawMarkdown": "Hey Alika,\n\n1. Most data frameworks will let you concatenate datasets with the same schema, for example with `import pandas as pd` in Python you can use `pd.read_csv()` to read each individual file, and `pd.concat()` to combine them together.  This may fail if you try to combine all the files at once and your machine doesn't have enough RAM -- there are many good ideas in the Discussions tab to reduce RAM cost and combine the data.\n2. The columns which are in the train set but not in the test set are metadata which should not be part of your model, but could be useful when deciding how to train your model.  The simplest approach is to remove those columns and shuffle the order of the rows in your train set before training your model.\n\nHope this helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1970890,
      "author_name": "chigozie01",
      "author_url": "",
      "post_date": "10/04/2022 10:46:15",
      "content": "<p>this was helpful, thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1970892,
      "author_name": "chigozie01",
      "author_url": "",
      "post_date": "10/04/2022 10:47:05",
      "content": "<p>thank you for this</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1976947,
      "author_name": "charithamaduranga",
      "author_url": "",
      "post_date": "10/07/2022 16:54:36",
      "content": "<p>I have a question!!!<br>\nIt says <strong><em>\"test.csv: Test set. Unlike the train set, the rows are scrambled.\"</em></strong><br>\nwhat does it mean, does is mean that all the rows in test set are shuffled so that each row can belongs to any game, any event randomly without the time sequence preserved?<br>\nthe 'event_time' is not given in the test set, and we are expected to predict team_[A|B]_scoring_within_10sec, how do we know at what point does the last 10sec begins without having 'event_time' in test set, because I assume the predictions for both teams should be 0 (or close to 0) before 10sec begin as it is in the train set</p>",
      "votes": null,
      "replies": [
        {
          "id": 1982597,
          "author_name": "dster",
          "author_url": "",
          "post_date": "10/11/2022 14:09:30",
          "content": "<p>Hey Charitha, yes we have removed the time sequence information from the test set, so that predictions are made only using game state information.  The time sequence information, such as <code>event_time</code>, is provided because it could be useful in creating a better model, even though that information is not available at inference time.  You are correct that in the majority of cases the predictions for both teams will be close to 0, since goals are not scored very often.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1981886,
      "author_name": "anmolstha1",
      "author_url": "",
      "post_date": "10/11/2022 05:57:58",
      "content": "<p>There is showimg memory error in my jupyter nootbook while data visualizing due to large data.is there any solution?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1982600,
          "author_name": "dster",
          "author_url": "",
          "post_date": "10/11/2022 14:11:26",
          "content": "<p>Hey there, many Kagglers have shared ways to reduce memory consumption in Notebooks when reading this competition's dataset, check out the Discussion tab and Code tab to learn more.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1995358,
      "author_name": "valentinfontanger",
      "author_url": "",
      "post_date": "10/19/2022 15:50:17",
      "content": "<p>HI, quick question : <br>\nI do not understand why most columns are not present in the test set. <br>\nWhat is the point of integrating these columns in my model if I am not able to use the model on the test set ? <br>\nThank you very much :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1995624,
          "author_name": "dster",
          "author_url": "",
          "post_date": "10/19/2022 17:58:07",
          "content": "<p>Hey Valentin,<br>\nBesides the target columns, the only columns in train but not in test are <code>game_num</code>, <code>event_id</code>, and <code>event_time</code>, which are just metadata providing more information about rows which come from the same timeseries, but do not provide any in-game information about the game state, which is what we want to make predictions based on.  Personally I am not using those columns in my own model, except for train / validation split, but I could imagine that the timeseries correlations could help build a smarter model, for example using a Recurrent Neural Network.  I'd like to explore that, though it's possible it does not end up being helpful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2000254,
          "author_name": "valentinfontanger",
          "author_url": "",
          "post_date": "10/23/2022 07:10:28",
          "content": "<p>Thank you for your reply ! <br>\nI just saw that I did not read the test.csv correctly ! <br>\nGood luck !!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1997693,
      "author_name": "aliibrahimali",
      "author_url": "",
      "post_date": "10/21/2022 05:27:26",
      "content": "<p>I am beginner here and  I have some  Questions:<br>\nwhy train data set is from train_0 to train_9 ?<br>\ncan I train one of these whole data set like train_0 enough !!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1998429,
          "author_name": "dster",
          "author_url": "",
          "post_date": "10/21/2022 15:46:44",
          "content": "<p>Hey there, the full train set is quite large, so we split it into 10 files.  You can probably get a good result using only <code>train_0.csv</code>, but generally it will be better to utilize all the training files.  One strategy would be to use only <code>train_0.csv</code> and get started with that, and once you have built a good pipeline, you can consider adding more train files, or utilizing data augmentation as others have mentioned in the forums, etc.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2000327,
      "author_name": "aliibrahimali",
      "author_url": "",
      "post_date": "10/23/2022 07:38:37",
      "content": "<p>Hi DJ Sterling I hope u r well<br>\nwhat benefit of the feature in training set only &amp; If it  bad choice if I excepted this data in model?<br>\nI beginner here and I'm I do not want be stuck very much here</p>",
      "votes": null,
      "replies": [
        {
          "id": 2002191,
          "author_name": "dster",
          "author_url": "",
          "post_date": "10/24/2022 15:21:28",
          "content": "<p>Hey, there are three non-target columns which are in train and not test -- <code>game_num</code>, <code>event_id</code>, and <code>event_time</code> -- and yes, you can ignore this data in the train set and still produce a good model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2003742,
      "author_name": "alikarizvi",
      "author_url": "",
      "post_date": "10/25/2022 18:46:40",
      "content": "<p>As I am new to the data science field , this is my second problem statement on kaggle .I have some questions<br>\n1-How to deal with train_0 to train_9 and make it one training dataset?<br>\n2-some features are missing in testing dataset ,how to deal with it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2003914,
          "author_name": "dster",
          "author_url": "",
          "post_date": "10/25/2022 21:41:01",
          "content": "<p>Hey Alika,</p>\n<ol>\n<li>Most data frameworks will let you concatenate datasets with the same schema, for example with <code>import pandas as pd</code> in Python you can use <code>pd.read_csv()</code> to read each individual file, and <code>pd.concat()</code> to combine them together.  This may fail if you try to combine all the files at once and your machine doesn't have enough RAM -- there are many good ideas in the Discussions tab to reduce RAM cost and combine the data.</li>\n<li>The columns which are in the train set but not in the test set are metadata which should not be part of your model, but could be useful when deciding how to train your model.  The simplest approach is to remove those columns and shuffle the order of the rows in your train set before training your model.</li>\n</ol>\n<p>Hope this helps!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1964853": "Hey there, my name is DJ and normally I'm a Software Engineer here at Kaggle, but this month I'll be your Data Scientist\\* for the Tabular Playground competition!  I'm a longtime Rocket League player and I'm excited to finally bring it to Kaggle -- I think this is a really interesting dataset with a noisy target and a lot of subtle signal to extract from the features.\n\nSome ideas:\n\n1. You could shuffle the training data and train a model as if the rows were independent, or you could try to use the timeseries information, for example with a recurrent network.\n2. Be careful with `NaN` values on the player columns when they get demolished -- this gives the other team an advantage.\n3. You can use data augmentation to get even more training data.\n4. Check out videos of professional Rocket League matches to learn more about the sport.\n5. Our [2020 NFL Big Data Bowl competition](https://www.kaggle.com/c/nfl-big-data-bowl-2020) had a similar setup with predicting how a sports play would unfold based on a snapshot in time, there could be useful hints from the discussions there.\n\nLet me know if you have any questions or feedback.  Have fun and see you on the leaderboard!\n\n\\**Not a real Data Scientist*",
    "1970890": "this was helpful, thank you",
    "1970892": "thank you for this",
    "1976947": "I have a question!!!\nIt says ***\"test.csv: Test set. Unlike the train set, the rows are scrambled.\"***\nwhat does it mean, does is mean that all the rows in test set are shuffled so that each row can belongs to any game, any event randomly without the time sequence preserved?\nthe 'event_time' is not given in the test set, and we are expected to predict team_[A|B]_scoring_within_10sec, how do we know at what point does the last 10sec begins without having 'event_time' in test set, because I assume the predictions for both teams should be 0 (or close to 0) before 10sec begin as it is in the train set",
    "1981886": "There is showimg memory error in my jupyter nootbook while data visualizing due to large data.is there any solution?",
    "1982597": "Hey Charitha, yes we have removed the time sequence information from the test set, so that predictions are made only using game state information.  The time sequence information, such as `event_time`, is provided because it could be useful in creating a better model, even though that information is not available at inference time.  You are correct that in the majority of cases the predictions for both teams will be close to 0, since goals are not scored very often.",
    "1982600": "Hey there, many Kagglers have shared ways to reduce memory consumption in Notebooks when reading this competition's dataset, check out the Discussion tab and Code tab to learn more.",
    "1995358": "HI, quick question : \nI do not understand why most columns are not present in the test set. \nWhat is the point of integrating these columns in my model if I am not able to use the model on the test set ? \nThank you very much :)",
    "1995624": "Hey Valentin,\nBesides the target columns, the only columns in train but not in test are `game_num`, `event_id`, and `event_time`, which are just metadata providing more information about rows which come from the same timeseries, but do not provide any in-game information about the game state, which is what we want to make predictions based on.  Personally I am not using those columns in my own model, except for train / validation split, but I could imagine that the timeseries correlations could help build a smarter model, for example using a Recurrent Neural Network.  I'd like to explore that, though it's possible it does not end up being helpful.",
    "1997693": "I am beginner here and  I have some  Questions:\nwhy train data set is from train_0 to train_9 ?\ncan I train one of these whole data set like train_0 enough !!",
    "1998429": "Hey there, the full train set is quite large, so we split it into 10 files.  You can probably get a good result using only `train_0.csv`, but generally it will be better to utilize all the training files.  One strategy would be to use only `train_0.csv` and get started with that, and once you have built a good pipeline, you can consider adding more train files, or utilizing data augmentation as others have mentioned in the forums, etc.",
    "2000254": "Thank you for your reply ! \nI just saw that I did not read the test.csv correctly ! \nGood luck !!",
    "2000327": "Hi DJ Sterling I hope u r well\nwhat benefit of the feature in training set only & If it  bad choice if I excepted this data in model?\nI beginner here and I'm I do not want be stuck very much here",
    "2002191": "Hey, there are three non-target columns which are in train and not test -- `game_num`, `event_id`, and `event_time` -- and yes, you can ignore this data in the train set and still produce a good model.",
    "2003742": "As I am new to the data science field , this is my second problem statement on kaggle .I have some questions\n1-How to deal with train_0 to train_9 and make it one training dataset?\n2-some features are missing in testing dataset ,how to deal with it?",
    "2003914": "Hey Alika,\n\n1. Most data frameworks will let you concatenate datasets with the same schema, for example with `import pandas as pd` in Python you can use `pd.read_csv()` to read each individual file, and `pd.concat()` to combine them together.  This may fail if you try to combine all the files at once and your machine doesn't have enough RAM -- there are many good ideas in the Discussions tab to reduce RAM cost and combine the data.\n2. The columns which are in the train set but not in the test set are metadata which should not be part of your model, but could be useful when deciding how to train your model.  The simplest approach is to remove those columns and shuffle the order of the rows in your train set before training your model.\n\nHope this helps!"
  },
  "source": "meta"
}