{
  "id": 53397,
  "title": "More data or More features?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53397",
  "author_name": "",
  "post_date": "2018-03-30T03:34:51.369640600Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello all,\nI do not have much experience in this field. I am making XGBoost machine with 8 million data points and around 30 features. This is giving me an AUC of 96.3! I believe, if I am able to use more data points, I can go around 0.97+ or maybe not. </p>\n\n<p>I do not want to use blending as of now to increase me score. I am having problem with loading my dataframe into dmatrix!</p>\n\n<p>Any solution? I do not have resources other than kaggle kernel!</p>",
  "messages": [
    {
      "id": "306212",
      "postDate": "03/30/2018 03:34:51",
      "content": "<p>Hello all,\nI do not have much experience in this field. I am making XGBoost machine with 8 million data points and around 30 features. This is giving me an AUC of 96.3! I believe, if I am able to use more data points, I can go around 0.97+ or maybe not. </p>\n\n<p>I do not want to use blending as of now to increase me score. I am having problem with loading my dataframe into dmatrix!</p>\n\n<p>Any solution? I do not have resources other than kaggle kernel!</p>",
      "rawMarkdown": "Hello all,\nI do not have much experience in this field. I am making XGBoost machine with 8 million data points and around 30 features. This is giving me an AUC of 96.3! I believe, if I am able to use more data points, I can go around 0.97+ or maybe not. \n\nI do not want to use blending as of now to increase me score. I am having problem with loading my dataframe into dmatrix!\n\nAny solution? I do not have resources other than kaggle kernel!",
      "votes": null
    },
    {
      "id": "306502",
      "postDate": "03/30/2018 14:22:01",
      "content": "<p>Given the fact that you are restricted to kaggle kernels, my best advice is to check out which of your 30 features really contribute to the model. I'm quite sure that only a hand full of your features is important. So try to reduce the number of features and increase the number of observations instead. This will surely improve your score.</p>",
      "rawMarkdown": "Given the fact that you are restricted to kaggle kernels, my best advice is to check out which of your 30 features really contribute to the model. I'm quite sure that only a hand full of your features is important. So try to reduce the number of features and increase the number of observations instead. This will surely improve your score.",
      "votes": null
    },
    {
      "id": "306747",
      "postDate": "03/30/2018 23:33:44",
      "content": "<p>OKay, I will create new features and delete useless features.\nThankyou danijel Kivaranovic</p>",
      "rawMarkdown": "OKay, I will create new features and delete useless features.\nThankyou danijel Kivaranovic",
      "votes": null
    },
    {
      "id": "306853",
      "postDate": "03/31/2018 07:06:25",
      "content": "<p>@ Gaurav - What was your logic behind picking 8 million records? Was it spread across different days ? Also - if it was across multiple days, did you maintain an equal balance in ratios between 1s and 0's across each data slice you picked. \nFor ex: This ratio is an average of say 0.24% (Approx) for every 10 million records </p>",
      "rawMarkdown": "Gaurav - What was your logic behind picking 8 million records? Was it spread across different days ? Also - if it was across multiple days, did you maintain an equal balance in ratios between 1s and 0's across each data slice you picked. \nFor ex: This ratio is an average of say 0.24% (Approx) for every 10 million records",
      "votes": null
    },
    {
      "id": "307015",
      "postDate": "03/31/2018 16:03:03",
      "content": "<p>I would concentrate on finding more features! Pick a small and relevant sample for training (you can read a lot in the forum what would be a \"relevant\" sample (8th and 9th, the same hours as the test data (10th)) and focus on coming up with good features! </p>",
      "rawMarkdown": "I would concentrate on finding more features! Pick a small and relevant sample for training (you can read a lot in the forum what would be a \"relevant\" sample (8th and 9th, the same hours as the test data (10th)) and focus on coming up with good features!",
      "votes": null
    },
    {
      "id": "307049",
      "postDate": "03/31/2018 17:54:17",
      "content": "<p>@Shanth\nI just used the maximum data points I could process with 30 features without hitting the memory on Kaggle server. </p>",
      "rawMarkdown": "Shanth\nI just used the maximum data points I could process with 30 features without hitting the memory on Kaggle server.",
      "votes": null
    },
    {
      "id": "307051",
      "postDate": "03/31/2018 17:56:19",
      "content": "<p>@Asparuh\nYes, finding a small but relevant data makes sense. How would you eliminate irrelevant features?</p>\n\n<p>I do that using VIF factors and feature importance plot.</p>",
      "rawMarkdown": "Asparuh\nYes, finding a small but relevant data makes sense. How would you eliminate irrelevant features?\n\nI do that using VIF factors and feature importance plot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 306502,
      "author_name": "danijelk",
      "author_url": "",
      "post_date": "03/30/2018 14:22:01",
      "content": "<p>Given the fact that you are restricted to kaggle kernels, my best advice is to check out which of your 30 features really contribute to the model. I'm quite sure that only a hand full of your features is important. So try to reduce the number of features and increase the number of observations instead. This will surely improve your score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 306747,
      "author_name": "aroragaurav",
      "author_url": "",
      "post_date": "03/30/2018 23:33:44",
      "content": "<p>OKay, I will create new features and delete useless features.\nThankyou danijel Kivaranovic</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 306853,
      "author_name": "shanth84",
      "author_url": "",
      "post_date": "03/31/2018 07:06:25",
      "content": "<p>@ Gaurav - What was your logic behind picking 8 million records? Was it spread across different days ? Also - if it was across multiple days, did you maintain an equal balance in ratios between 1s and 0's across each data slice you picked. \nFor ex: This ratio is an average of say 0.24% (Approx) for every 10 million records </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 307015,
      "author_name": "asparuhhristov",
      "author_url": "",
      "post_date": "03/31/2018 16:03:03",
      "content": "<p>I would concentrate on finding more features! Pick a small and relevant sample for training (you can read a lot in the forum what would be a \"relevant\" sample (8th and 9th, the same hours as the test data (10th)) and focus on coming up with good features! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 307049,
      "author_name": "aroragaurav",
      "author_url": "",
      "post_date": "03/31/2018 17:54:17",
      "content": "<p>@Shanth\nI just used the maximum data points I could process with 30 features without hitting the memory on Kaggle server. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 307051,
      "author_name": "aroragaurav",
      "author_url": "",
      "post_date": "03/31/2018 17:56:19",
      "content": "<p>@Asparuh\nYes, finding a small but relevant data makes sense. How would you eliminate irrelevant features?</p>\n\n<p>I do that using VIF factors and feature importance plot.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "306212": "Hello all,\nI do not have much experience in this field. I am making XGBoost machine with 8 million data points and around 30 features. This is giving me an AUC of 96.3! I believe, if I am able to use more data points, I can go around 0.97+ or maybe not. \n\nI do not want to use blending as of now to increase me score. I am having problem with loading my dataframe into dmatrix!\n\nAny solution? I do not have resources other than kaggle kernel!",
    "306502": "Given the fact that you are restricted to kaggle kernels, my best advice is to check out which of your 30 features really contribute to the model. I'm quite sure that only a hand full of your features is important. So try to reduce the number of features and increase the number of observations instead. This will surely improve your score.",
    "306747": "OKay, I will create new features and delete useless features.\nThankyou danijel Kivaranovic",
    "306853": "Gaurav - What was your logic behind picking 8 million records? Was it spread across different days ? Also - if it was across multiple days, did you maintain an equal balance in ratios between 1s and 0's across each data slice you picked. \nFor ex: This ratio is an average of say 0.24% (Approx) for every 10 million records",
    "307015": "I would concentrate on finding more features! Pick a small and relevant sample for training (you can read a lot in the forum what would be a \"relevant\" sample (8th and 9th, the same hours as the test data (10th)) and focus on coming up with good features!",
    "307049": "Shanth\nI just used the maximum data points I could process with 30 features without hitting the memory on Kaggle server.",
    "307051": "Asparuh\nYes, finding a small but relevant data makes sense. How would you eliminate irrelevant features?\n\nI do that using VIF factors and feature importance plot."
  },
  "source": "meta"
}