{
  "id": 59446,
  "title": "Add more features,but score and CV not change",
  "url": "/competitions/avito-demand-prediction/discussion/59446",
  "author_name": "",
  "post_date": "2018-06-22T08:56:35.817115100Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I add more features , but my lgb model leader board score and CV did not change, what should I do? \nfeature selection?</p>",
  "messages": [
    {
      "id": "346729",
      "postDate": "06/22/2018 08:56:35",
      "content": "<p>I add more features , but my lgb model leader board score and CV did not change, what should I do? \nfeature selection?</p>",
      "rawMarkdown": "I add more features , but my lgb model leader board score and CV did not change, what should I do? \nfeature selection?",
      "votes": null
    },
    {
      "id": "346746",
      "postDate": "06/22/2018 10:12:09",
      "content": "<p>take a break and read this <a href=\"http://www2.math.uu.se/~thulin/mm/breiman.pdf\">http://www2.math.uu.se/~thulin/mm/breiman.pdf</a></p>",
      "rawMarkdown": "take a break and read this http://www2.math.uu.se/~thulin/mm/breiman.pdf",
      "votes": null
    },
    {
      "id": "346753",
      "postDate": "06/22/2018 10:21:42",
      "content": "<p>Maybe ensemble with some different models? stacking several lgb models(0.2207~0.2212) gives 0.2201. </p>\n\n<p>And just adding a weak nn(~0.2277) the stacking lgb model gives 0.2196. (emmm, just hope it is not overfitting too much...</p>\n\n<p>And pay attention to some of features with quite high lgb's feature_importance, which may lead to faster 'overfitting'. For example, the VGG16 ImageNet top classes and public-shared location clusters are with really high feature_importance but gives me worse CV scores. So I deleted such features.</p>",
      "rawMarkdown": "Maybe ensemble with some different models? stacking several lgb models(0.2207~0.2212) gives 0.2201. \n\nAnd just adding a weak nn(~0.2277) the stacking lgb model gives 0.2196. (emmm, just hope it is not overfitting too much...\n\nAnd pay attention to some of features with quite high lgb's feature_importance, which may lead to faster 'overfitting'. For example, the VGG16 ImageNet top classes and public-shared location clusters are with really high feature_importance but gives me worse CV scores. So I deleted such features.",
      "votes": null
    },
    {
      "id": "346759",
      "postDate": "06/22/2018 10:41:29",
      "content": "<p>thanks your awesome link~</p>",
      "rawMarkdown": "thanks your awesome link~",
      "votes": null
    },
    {
      "id": "346761",
      "postDate": "06/22/2018 10:44:56",
      "content": "<p>Thanks your advice.I add image features like rgb mean,yuv mean,When I use both rgb and yuv，my score get worse.so I use yuv features only.I have use stacking,and I will try more models stacking.Thanks a lot</p>",
      "rawMarkdown": "Thanks your advice.I add image features like rgb mean,yuv mean,When I use both rgb and yuv，my score get worse.so I use yuv features only.I have use stacking,and I will try more models stacking.Thanks a lot",
      "votes": null
    },
    {
      "id": "347200",
      "postDate": "06/23/2018 15:11:41",
      "content": "<p>@Jiazhen Xi, can I ask you what model do you use for stacking/ensembling? lgb?</p>",
      "rawMarkdown": "Jiazhen Xi, can I ask you what model do you use for stacking/ensembling? lgb?",
      "votes": null
    },
    {
      "id": "347239",
      "postDate": "06/23/2018 17:14:46",
      "content": "<p>I only used lgb and ridge for stacking on validation set (10% holdout) and ridge seems slightly better on LB but has a weaker validation score, compared to lgb. </p>",
      "rawMarkdown": "I only used lgb and ridge for stacking on validation set (10% holdout) and ridge seems slightly better on LB but has a weaker validation score, compared to lgb.",
      "votes": null
    },
    {
      "id": "347247",
      "postDate": "06/23/2018 17:41:15",
      "content": "<p>Thanks for answer. I confirm that Ridge works fine here. At least on cv )</p>",
      "rawMarkdown": "Thanks for answer. I confirm that Ridge works fine here. At least on cv )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 346746,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "06/22/2018 10:12:09",
      "content": "<p>take a break and read this <a href=\"http://www2.math.uu.se/~thulin/mm/breiman.pdf\">http://www2.math.uu.se/~thulin/mm/breiman.pdf</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 346759,
          "author_name": "liuhdsgoal",
          "author_url": "",
          "post_date": "06/22/2018 10:41:29",
          "content": "<p>thanks your awesome link~</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 346753,
      "author_name": "johnfarrell",
      "author_url": "",
      "post_date": "06/22/2018 10:21:42",
      "content": "<p>Maybe ensemble with some different models? stacking several lgb models(0.2207~0.2212) gives 0.2201. </p>\n\n<p>And just adding a weak nn(~0.2277) the stacking lgb model gives 0.2196. (emmm, just hope it is not overfitting too much...</p>\n\n<p>And pay attention to some of features with quite high lgb's feature_importance, which may lead to faster 'overfitting'. For example, the VGG16 ImageNet top classes and public-shared location clusters are with really high feature_importance but gives me worse CV scores. So I deleted such features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 346761,
          "author_name": "liuhdsgoal",
          "author_url": "",
          "post_date": "06/22/2018 10:44:56",
          "content": "<p>Thanks your advice.I add image features like rgb mean,yuv mean,When I use both rgb and yuv，my score get worse.so I use yuv features only.I have use stacking,and I will try more models stacking.Thanks a lot</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347200,
          "author_name": "kruegger",
          "author_url": "",
          "post_date": "06/23/2018 15:11:41",
          "content": "<p>@Jiazhen Xi, can I ask you what model do you use for stacking/ensembling? lgb?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347239,
          "author_name": "johnfarrell",
          "author_url": "",
          "post_date": "06/23/2018 17:14:46",
          "content": "<p>I only used lgb and ridge for stacking on validation set (10% holdout) and ridge seems slightly better on LB but has a weaker validation score, compared to lgb. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347247,
          "author_name": "kruegger",
          "author_url": "",
          "post_date": "06/23/2018 17:41:15",
          "content": "<p>Thanks for answer. I confirm that Ridge works fine here. At least on cv )</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "346729": "I add more features , but my lgb model leader board score and CV did not change, what should I do? \nfeature selection?",
    "346746": "take a break and read this http://www2.math.uu.se/~thulin/mm/breiman.pdf",
    "346753": "Maybe ensemble with some different models? stacking several lgb models(0.2207~0.2212) gives 0.2201. \n\nAnd just adding a weak nn(~0.2277) the stacking lgb model gives 0.2196. (emmm, just hope it is not overfitting too much...\n\nAnd pay attention to some of features with quite high lgb's feature_importance, which may lead to faster 'overfitting'. For example, the VGG16 ImageNet top classes and public-shared location clusters are with really high feature_importance but gives me worse CV scores. So I deleted such features.",
    "346759": "thanks your awesome link~",
    "346761": "Thanks your advice.I add image features like rgb mean,yuv mean,When I use both rgb and yuv，my score get worse.so I use yuv features only.I have use stacking,and I will try more models stacking.Thanks a lot",
    "347200": "Jiazhen Xi, can I ask you what model do you use for stacking/ensembling? lgb?",
    "347239": "I only used lgb and ridge for stacking on validation set (10% holdout) and ridge seems slightly better on LB but has a weaker validation score, compared to lgb.",
    "347247": "Thanks for answer. I confirm that Ridge works fine here. At least on cv )"
  },
  "source": "meta"
}