{
  "id": 53574,
  "title": "Extra data hurting LB score",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53574",
  "author_name": "",
  "post_date": "2018-04-01T22:26:18.472514900Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I've been experimenting with an LGBM model recently and, for some reason, the public LB score of the predictions decreases when I run the model on the entire training set as opposed to ~70M training rows. I'm trying to understand how that can happen since I'm passing in extra data. Could it be because some of the features I'm using are spurious and are actually hurting the model performance with more data on them?</p>\n\n<p>Another thought is: could it be because I have to tune the model parameters again with the extra data? Maybe if the right parameters are set, the performance will surpass that of one with ~70M rows?</p>\n\n<p>The same goes with extra features; I'm assuming it's possible for adding additional features to actually decrease a model's performance. So quality really is better than quantity in this case?</p>\n\n<p>Just want to check my understanding.</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "307515",
      "postDate": "04/01/2018 22:26:18",
      "content": "<p>I've been experimenting with an LGBM model recently and, for some reason, the public LB score of the predictions decreases when I run the model on the entire training set as opposed to ~70M training rows. I'm trying to understand how that can happen since I'm passing in extra data. Could it be because some of the features I'm using are spurious and are actually hurting the model performance with more data on them?</p>\n\n<p>Another thought is: could it be because I have to tune the model parameters again with the extra data? Maybe if the right parameters are set, the performance will surpass that of one with ~70M rows?</p>\n\n<p>The same goes with extra features; I'm assuming it's possible for adding additional features to actually decrease a model's performance. So quality really is better than quantity in this case?</p>\n\n<p>Just want to check my understanding.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "I've been experimenting with an LGBM model recently and, for some reason, the public LB score of the predictions decreases when I run the model on the entire training set as opposed to ~70M training rows. I'm trying to understand how that can happen since I'm passing in extra data. Could it be because some of the features I'm using are spurious and are actually hurting the model performance with more data on them?\n\nAnother thought is: could it be because I have to tune the model parameters again with the extra data? Maybe if the right parameters are set, the performance will surpass that of one with ~70M rows?\n\nThe same goes with extra features; I'm assuming it's possible for adding additional features to actually decrease a model's performance. So quality really is better than quantity in this case?\n\nJust want to check my understanding.\n\nThanks",
      "votes": null
    },
    {
      "id": "307542",
      "postDate": "04/02/2018 00:48:57",
      "content": "<p>You give too little information about what exactly 70M data you use, what kind of features you use, what is the lb score with extra data and what is the lb score with 70M data...So my answer is I don't know, sorry.</p>",
      "rawMarkdown": "You give too little information about what exactly 70M data you use, what kind of features you use, what is the lb score with extra data and what is the lb score with 70M data...So my answer is I don't know, sorry.",
      "votes": null
    },
    {
      "id": "307546",
      "postDate": "04/02/2018 01:03:37",
      "content": "<p>I'm sorry about not providing enough information. One of the cases is with this kernel: <a href=\"https://www.kaggle.com/aharless/try-pranav-s-r-lgbm-in-python/code\">https://www.kaggle.com/aharless/try-pranav-s-r-lgbm-in-python/code</a> </p>\n\n<p>It's able to score .9691 LB score with 100M rows, but when I train it on the full training set it goes down to roughly .9684 LB score. </p>",
      "rawMarkdown": "I'm sorry about not providing enough information. One of the cases is with this kernel: https://www.kaggle.com/aharless/try-pranav-s-r-lgbm-in-python/code \n\nIt's able to score .9691 LB score with 100M rows, but when I train it on the full training set it goes down to roughly .9684 LB score.",
      "votes": null
    },
    {
      "id": "307547",
      "postDate": "04/02/2018 01:11:31",
      "content": "<p>think about two things that may be helpful: parameter values and number of rounds you need to use when the data set changes</p>",
      "rawMarkdown": "think about two things that may be helpful: parameter values and number of rounds you need to use when the data set changes",
      "votes": null
    },
    {
      "id": "307575",
      "postDate": "04/02/2018 03:14:11",
      "content": "<p>Thanks for the suggestions! I'll try them out. It seems that you were able to get a higher LB score with the full training set? </p>",
      "rawMarkdown": "Thanks for the suggestions! I'll try them out. It seems that you were able to get a higher LB score with the full training set?",
      "votes": null
    },
    {
      "id": "307702",
      "postDate": "04/02/2018 09:01:35",
      "content": "<p>It may be because the kernel you started from is overfiting to the public LB.  Any change to it will degrade LB score .</p>",
      "rawMarkdown": "It may be because the kernel you started from is overfiting to the public LB.  Any change to it will degrade LB score .",
      "votes": null
    },
    {
      "id": "307706",
      "postDate": "04/02/2018 09:10:06",
      "content": "<p>@CPMP does increasing training data requires to change model parameters too?</p>",
      "rawMarkdown": "CPMP does increasing training data requires to change model parameters too?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 307542,
      "author_name": "wythhh",
      "author_url": "",
      "post_date": "04/02/2018 00:48:57",
      "content": "<p>You give too little information about what exactly 70M data you use, what kind of features you use, what is the lb score with extra data and what is the lb score with 70M data...So my answer is I don't know, sorry.</p>",
      "votes": null,
      "replies": [
        {
          "id": 307546,
          "author_name": "edchen8",
          "author_url": "",
          "post_date": "04/02/2018 01:03:37",
          "content": "<p>I'm sorry about not providing enough information. One of the cases is with this kernel: <a href=\"https://www.kaggle.com/aharless/try-pranav-s-r-lgbm-in-python/code\">https://www.kaggle.com/aharless/try-pranav-s-r-lgbm-in-python/code</a> </p>\n\n<p>It's able to score .9691 LB score with 100M rows, but when I train it on the full training set it goes down to roughly .9684 LB score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307547,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/02/2018 01:11:31",
          "content": "<p>think about two things that may be helpful: parameter values and number of rounds you need to use when the data set changes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 307575,
          "author_name": "edchen8",
          "author_url": "",
          "post_date": "04/02/2018 03:14:11",
          "content": "<p>Thanks for the suggestions! I'll try them out. It seems that you were able to get a higher LB score with the full training set? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 307702,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/02/2018 09:01:35",
      "content": "<p>It may be because the kernel you started from is overfiting to the public LB.  Any change to it will degrade LB score .</p>",
      "votes": null,
      "replies": [
        {
          "id": 307706,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/02/2018 09:10:06",
          "content": "<p>@CPMP does increasing training data requires to change model parameters too?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "307515": "I've been experimenting with an LGBM model recently and, for some reason, the public LB score of the predictions decreases when I run the model on the entire training set as opposed to ~70M training rows. I'm trying to understand how that can happen since I'm passing in extra data. Could it be because some of the features I'm using are spurious and are actually hurting the model performance with more data on them?\n\nAnother thought is: could it be because I have to tune the model parameters again with the extra data? Maybe if the right parameters are set, the performance will surpass that of one with ~70M rows?\n\nThe same goes with extra features; I'm assuming it's possible for adding additional features to actually decrease a model's performance. So quality really is better than quantity in this case?\n\nJust want to check my understanding.\n\nThanks",
    "307542": "You give too little information about what exactly 70M data you use, what kind of features you use, what is the lb score with extra data and what is the lb score with 70M data...So my answer is I don't know, sorry.",
    "307546": "I'm sorry about not providing enough information. One of the cases is with this kernel: https://www.kaggle.com/aharless/try-pranav-s-r-lgbm-in-python/code \n\nIt's able to score .9691 LB score with 100M rows, but when I train it on the full training set it goes down to roughly .9684 LB score.",
    "307547": "think about two things that may be helpful: parameter values and number of rounds you need to use when the data set changes",
    "307575": "Thanks for the suggestions! I'll try them out. It seems that you were able to get a higher LB score with the full training set?",
    "307702": "It may be because the kernel you started from is overfiting to the public LB.  Any change to it will degrade LB score .",
    "307706": "CPMP does increasing training data requires to change model parameters too?"
  },
  "source": "meta"
}