{
  "id": 58232,
  "title": "TFIDF makes LGBM  worse?",
  "url": "/competitions/avito-demand-prediction/discussion/58232",
  "author_name": "",
  "post_date": "2018-06-04T19:19:16.708079Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi,\nI have a weird issue.\nI have 42 numeric+categorical features in LGBM that give CV of 0.218\nHowever, when adding 9000 columns of TFIDF values and running the LGBM (with more leaves) I get much worse result.</p>\n\n<p>Any ideas?</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "338296",
      "postDate": "06/04/2018 19:19:16",
      "content": "<p>Hi,\nI have a weird issue.\nI have 42 numeric+categorical features in LGBM that give CV of 0.218\nHowever, when adding 9000 columns of TFIDF values and running the LGBM (with more leaves) I get much worse result.</p>\n\n<p>Any ideas?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi,\nI have a weird issue.\nI have 42 numeric+categorical features in LGBM that give CV of 0.218\nHowever, when adding 9000 columns of TFIDF values and running the LGBM (with more leaves) I get much worse result.\n\nAny ideas?\n\nThanks!",
      "votes": null
    },
    {
      "id": "338305",
      "postDate": "06/04/2018 19:29:13",
      "content": "<p>Are you sure that all the features are ordered consistently and that you're partitioning the rows properly for train/validation splits? It's hard to give specific advice without knowing what your setup looks like, but it sounds pretty likely that there's some sort of mismatch - whether it's between numeric/categorical and tf-idf features within a training set, or between all features across train/validation etc.  If your training error gets worse with the same model parameters, that points pretty explicitly to a formatting problem - it shouldn't really ever happen that train performance degrades significantly when adding more proper features.</p>\n\n<p>I would revisit your feature engineering setup and see if you can identify where misalignment may have happened. For example, let's say you processed the tf-idf features into a sparse matrix, and afterwards merged the main training dataframe with some feature files. The merge will change the order of the df, and your tf-idf features will be totally out of sync. I would bet it's something like that happening, or a problem with the way you're indexing when dividing tabular features + tf-idf into train/val.</p>",
      "rawMarkdown": "Are you sure that all the features are ordered consistently and that you're partitioning the rows properly for train/validation splits? It's hard to give specific advice without knowing what your setup looks like, but it sounds pretty likely that there's some sort of mismatch - whether it's between numeric/categorical and tf-idf features within a training set, or between all features across train/validation etc.  If your training error gets worse with the same model parameters, that points pretty explicitly to a formatting problem - it shouldn't really ever happen that train performance degrades significantly when adding more proper features.\n\nI would revisit your feature engineering setup and see if you can identify where misalignment may have happened. For example, let's say you processed the tf-idf features into a sparse matrix, and afterwards merged the main training dataframe with some feature files. The merge will change the order of the df, and your tf-idf features will be totally out of sync. I would bet it's something like that happening, or a problem with the way you're indexing when dividing tabular features + tf-idf into train/val.",
      "votes": null
    },
    {
      "id": "338323",
      "postDate": "06/04/2018 20:41:31",
      "content": "<p>This doesn't seem impossible to me. More leaves and more features could invite more opportunities for overfitting and thus worse results. For example, I have some LGB models where adding TFIDF columns created a worse score on my CV than not having them at all.</p>",
      "rawMarkdown": "This doesn't seem impossible to me. More leaves and more features could invite more opportunities for overfitting and thus worse results. For example, I have some LGB models where adding TFIDF columns created a worse score on my CV than not having them at all.",
      "votes": null
    },
    {
      "id": "338329",
      "postDate": "06/04/2018 20:51:13",
      "content": "<p>Maybe you can try dimensionality reduction on 9000 features, and test again?</p>",
      "rawMarkdown": "Maybe you can try dimensionality reduction on 9000 features, and test again?",
      "votes": null
    },
    {
      "id": "338357",
      "postDate": "06/04/2018 22:19:52",
      "content": "<p>Peter and Joes comments are good. Another possibility is you are sampling features in LGB and it hardly gets non TFIDF features.</p>",
      "rawMarkdown": "Peter and Joes comments are good. Another possibility is you are sampling features in LGB and it hardly gets non TFIDF features.",
      "votes": null
    },
    {
      "id": "338367",
      "postDate": "06/04/2018 22:48:45",
      "content": "<p>I agree with both @Joe Eddy and @Peter Hurford here. Make sure that however you are splitting your data is consistent with the transformation for TFIDF. Also, just from what I have observed before it is very easy that with more features that LGBM can overfit. I would experiment with the number of features that you have in your TFIDF model.</p>",
      "rawMarkdown": "I agree with both @Joe Eddy and @Peter Hurford here. Make sure that however you are splitting your data is consistent with the transformation for TFIDF. Also, just from what I have observed before it is very easy that with more features that LGBM can overfit. I would experiment with the number of features that you have in your TFIDF model.",
      "votes": null
    },
    {
      "id": "338417",
      "postDate": "06/05/2018 02:35:03",
      "content": "<p>I agree. Especially 0.218 seems very very low... You may be overfitting your training data... If your LB for this LGB without sparse is too far from CV it might be interesting to investigate this difference between your LB and CV ?</p>",
      "rawMarkdown": "I agree. Especially 0.218 seems very very low... You may be overfitting your training data... If your LB for this LGB without sparse is too far from CV it might be interesting to investigate this difference between your LB and CV ?",
      "votes": null
    },
    {
      "id": "338656",
      "postDate": "06/05/2018 14:24:30",
      "content": "<p>Maybe you should increase num_leaves</p>",
      "rawMarkdown": "Maybe you should increase num_leaves",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 338305,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "06/04/2018 19:29:13",
      "content": "<p>Are you sure that all the features are ordered consistently and that you're partitioning the rows properly for train/validation splits? It's hard to give specific advice without knowing what your setup looks like, but it sounds pretty likely that there's some sort of mismatch - whether it's between numeric/categorical and tf-idf features within a training set, or between all features across train/validation etc.  If your training error gets worse with the same model parameters, that points pretty explicitly to a formatting problem - it shouldn't really ever happen that train performance degrades significantly when adding more proper features.</p>\n\n<p>I would revisit your feature engineering setup and see if you can identify where misalignment may have happened. For example, let's say you processed the tf-idf features into a sparse matrix, and afterwards merged the main training dataframe with some feature files. The merge will change the order of the df, and your tf-idf features will be totally out of sync. I would bet it's something like that happening, or a problem with the way you're indexing when dividing tabular features + tf-idf into train/val.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 338323,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "06/04/2018 20:41:31",
      "content": "<p>This doesn't seem impossible to me. More leaves and more features could invite more opportunities for overfitting and thus worse results. For example, I have some LGB models where adding TFIDF columns created a worse score on my CV than not having them at all.</p>",
      "votes": null,
      "replies": [
        {
          "id": 338417,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "06/05/2018 02:35:03",
          "content": "<p>I agree. Especially 0.218 seems very very low... You may be overfitting your training data... If your LB for this LGB without sparse is too far from CV it might be interesting to investigate this difference between your LB and CV ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 338329,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "06/04/2018 20:51:13",
      "content": "<p>Maybe you can try dimensionality reduction on 9000 features, and test again?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 338357,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "06/04/2018 22:19:52",
      "content": "<p>Peter and Joes comments are good. Another possibility is you are sampling features in LGB and it hardly gets non TFIDF features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 338367,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "06/04/2018 22:48:45",
      "content": "<p>I agree with both @Joe Eddy and @Peter Hurford here. Make sure that however you are splitting your data is consistent with the transformation for TFIDF. Also, just from what I have observed before it is very easy that with more features that LGBM can overfit. I would experiment with the number of features that you have in your TFIDF model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 338656,
      "author_name": "liujilong",
      "author_url": "",
      "post_date": "06/05/2018 14:24:30",
      "content": "<p>Maybe you should increase num_leaves</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "338296": "Hi,\nI have a weird issue.\nI have 42 numeric+categorical features in LGBM that give CV of 0.218\nHowever, when adding 9000 columns of TFIDF values and running the LGBM (with more leaves) I get much worse result.\n\nAny ideas?\n\nThanks!",
    "338305": "Are you sure that all the features are ordered consistently and that you're partitioning the rows properly for train/validation splits? It's hard to give specific advice without knowing what your setup looks like, but it sounds pretty likely that there's some sort of mismatch - whether it's between numeric/categorical and tf-idf features within a training set, or between all features across train/validation etc.  If your training error gets worse with the same model parameters, that points pretty explicitly to a formatting problem - it shouldn't really ever happen that train performance degrades significantly when adding more proper features.\n\nI would revisit your feature engineering setup and see if you can identify where misalignment may have happened. For example, let's say you processed the tf-idf features into a sparse matrix, and afterwards merged the main training dataframe with some feature files. The merge will change the order of the df, and your tf-idf features will be totally out of sync. I would bet it's something like that happening, or a problem with the way you're indexing when dividing tabular features + tf-idf into train/val.",
    "338323": "This doesn't seem impossible to me. More leaves and more features could invite more opportunities for overfitting and thus worse results. For example, I have some LGB models where adding TFIDF columns created a worse score on my CV than not having them at all.",
    "338329": "Maybe you can try dimensionality reduction on 9000 features, and test again?",
    "338357": "Peter and Joes comments are good. Another possibility is you are sampling features in LGB and it hardly gets non TFIDF features.",
    "338367": "I agree with both @Joe Eddy and @Peter Hurford here. Make sure that however you are splitting your data is consistent with the transformation for TFIDF. Also, just from what I have observed before it is very easy that with more features that LGBM can overfit. I would experiment with the number of features that you have in your TFIDF model.",
    "338417": "I agree. Especially 0.218 seems very very low... You may be overfitting your training data... If your LB for this LGB without sparse is too far from CV it might be interesting to investigate this difference between your LB and CV ?",
    "338656": "Maybe you should increase num_leaves"
  },
  "source": "meta"
}