{
  "id": 59104,
  "title": "Dense vs Sparse vs lightgbm.Dataset vs xgb.DMatrix. Could someone clarify some things?",
  "url": "/competitions/avito-demand-prediction/discussion/59104",
  "author_name": "",
  "post_date": "2018-06-18T15:25:53.990643100Z",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In this competition, I finally switch to sparse matrices after my giant TFIDFs were unable to fit into memory. Should have done that a long time ago. So much more memory efficient! And faster too. </p>\n\n<p>But now I have some new problems. I wander how more experienced kagglers deal with them.</p>\n\n<p>1) Sparse matrices don't store feature names. Store names separately? That’s not very convenient. </p>\n\n<p>2) When I was still using dense features (pandas DataFrame), my LightGBM showed some improvement when I explicitly specified that certain columns are categorical via astype('category'). I cannot do that with sparse matrix. :( Any workaround?</p>\n\n<p>3) What are lightgbm.Dataset and xgb.DMatrix for? Both xgboost and lightgbm can work with DataFrames, numpy arrays, sparse matrices. Why another data format? What are the benefits?</p>\n\n<p>4) Is there any way to use dense and sparse features simultaneously? All I can think of is hstacking dense features to sparse features last thing before fitting, but, in that case, problems 1 and 2 aren't solved. </p>",
  "messages": [
    {
      "id": "344730",
      "postDate": "06/18/2018 15:25:53",
      "content": "<p>In this competition, I finally switch to sparse matrices after my giant TFIDFs were unable to fit into memory. Should have done that a long time ago. So much more memory efficient! And faster too. </p>\n\n<p>But now I have some new problems. I wander how more experienced kagglers deal with them.</p>\n\n<p>1) Sparse matrices don't store feature names. Store names separately? That’s not very convenient. </p>\n\n<p>2) When I was still using dense features (pandas DataFrame), my LightGBM showed some improvement when I explicitly specified that certain columns are categorical via astype('category'). I cannot do that with sparse matrix. :( Any workaround?</p>\n\n<p>3) What are lightgbm.Dataset and xgb.DMatrix for? Both xgboost and lightgbm can work with DataFrames, numpy arrays, sparse matrices. Why another data format? What are the benefits?</p>\n\n<p>4) Is there any way to use dense and sparse features simultaneously? All I can think of is hstacking dense features to sparse features last thing before fitting, but, in that case, problems 1 and 2 aren't solved. </p>",
      "rawMarkdown": "In this competition, I finally switch to sparse matrices after my giant TFIDFs were unable to fit into memory. Should have done that a long time ago. So much more memory efficient! And faster too. \n\nBut now I have some new problems. I wander how more experienced kagglers deal with them.\n\n1) Sparse matrices don't store feature names. Store names separately? That’s not very convenient. \n\n2) When I was still using dense features (pandas DataFrame), my LightGBM showed some improvement when I explicitly specified that certain columns are categorical via astype('category'). I cannot do that with sparse matrix. :( Any workaround?\n\n3) What are lightgbm.Dataset and xgb.DMatrix for? Both xgboost and lightgbm can work with DataFrames, numpy arrays, sparse matrices. Why another data format? What are the benefits?\n\n4) Is there any way to use dense and sparse features simultaneously? All I can think of is hstacking dense features to sparse features last thing before fitting, but, in that case, problems 1 and 2 aren't solved.",
      "votes": null
    },
    {
      "id": "344779",
      "postDate": "06/18/2018 16:58:27",
      "content": "<p>Well, I am not an experienced Kaggler but I would like to suggest few ideas: <br>\n2) you can replace category features with dummy features (less optimal, but you can use with sparse matrix); actually, some claims that using dummy features is more efficient (in terms of performance) than using the categorical features list with lgb;     </p>\n\n<p>3) lightgbm.Dataset   (see: <a href=\"https://media.readthedocs.org/pdf/lightgbm/latest/lightgbm.pdf\">https://media.readthedocs.org/pdf/lightgbm/latest/lightgbm.pdf</a>, pag 12):  </p>\n\n<p>&gt; Memory efficent usage:   The Dataset object in LightGBM is very\n&gt; memory-efficient, due to it only need to save discrete bins. However,\n&gt; Numpy/Array/Pandas object is memory cost. If you concern about your\n&gt; memory consumption, you can save memory according to following: <br>\n&gt; 1. Let free_raw_data=True (default is True) when constructing the Dataset <br>\n&gt; 2. Explicit set raw_data=None after the Dataset has been constructed <br>\n&gt; 3. Call gc</p>",
      "rawMarkdown": "Well, I am not an experienced Kaggler but I would like to suggest few ideas:  \n2) you can replace category features with dummy features (less optimal, but you can use with sparse matrix); actually, some claims that using dummy features is more efficient (in terms of performance) than using the categorical features list with lgb;     \n\n3) lightgbm.Dataset   (see: https://media.readthedocs.org/pdf/lightgbm/latest/lightgbm.pdf, pag 12):  \n\n&gt; Memory efficent usage:   The Dataset object in LightGBM is very\n&gt; memory-efficient, due to it only need to save discrete bins. However,\n&gt; Numpy/Array/Pandas object is memory cost. If you concern about your\n&gt; memory consumption, you can save memory according to following:  \n&gt; 1. Let free_raw_data=True (default is True) when constructing the Dataset  \n&gt; 2. Explicit set raw_data=None after the Dataset has been constructed  \n&gt; 3. Call gc",
      "votes": null
    },
    {
      "id": "344782",
      "postDate": "06/18/2018 17:03:15",
      "content": "<blockquote>\n  <p>Sparse matrices don't store feature names. Store names separately</p>\n</blockquote>\n\n<p>Storing names separately is the only way I know of.</p>\n\n<blockquote>\n  <p>my LightGBM showed some improvement when I explicitly specified that certain columns are categorical via astype('category'). I cannot do that with sparse matrix. :( Any workaround?</p>\n</blockquote>\n\n<p>You'll have to manually implement sort of categorical-&gt;numeric scheme that captures this. I'd recommend label encoding, target encoding, or one hot encoding.</p>\n\n<blockquote>\n  <p>Is there any way to use dense and sparse features simultaneously? All I can think of is hstacking dense features to sparse features last thing before fitting, but, in that case, problems 1 and 2 aren't solved.</p>\n</blockquote>\n\n<p>Hstacking is the only way I know of. The features are still the same features (though they do have to be numeric) but the result is sparse. So you're right you can't fix your problems this way.</p>",
      "rawMarkdown": "&gt; Sparse matrices don't store feature names. Store names separately\n\nStoring names separately is the only way I know of.\n\n&gt; my LightGBM showed some improvement when I explicitly specified that certain columns are categorical via astype('category'). I cannot do that with sparse matrix. :( Any workaround?\n\nYou'll have to manually implement sort of categorical-&gt;numeric scheme that captures this. I'd recommend label encoding, target encoding, or one hot encoding.\n\n&gt; Is there any way to use dense and sparse features simultaneously? All I can think of is hstacking dense features to sparse features last thing before fitting, but, in that case, problems 1 and 2 aren't solved.\n\nHstacking is the only way I know of. The features are still the same features (though they do have to be numeric) but the result is sparse. So you're right you can't fix your problems this way.",
      "votes": null
    },
    {
      "id": "344803",
      "postDate": "06/18/2018 17:32:52",
      "content": "<blockquote>\n  <p>2) When I was still using dense features (pandas DataFrame), my\n  LightGBM showed some improvement when I explicitly specified that\n  certain columns are categorical via astype('category'). I cannot do\n  that with sparse matrix. :( Any workaround?</p>\n</blockquote>\n\n<p>As far as I know, LightGBM still supports categorical feature handling when passing a sparse matrix to construct a lgb dataset. You just have to make sure that you get your feature name list correctly ordered, and then pass a sublist to the categorical_feature argument, as below:</p>\n\n<pre><code>lgb_train = lgb.Dataset(X_train, \n                        label=y_train,\n                        feature_name=feature_names, \n                        categorical_feature=categorical_features)\n</code></pre>",
      "rawMarkdown": "&gt; 2) When I was still using dense features (pandas DataFrame), my\n&gt; LightGBM showed some improvement when I explicitly specified that\n&gt; certain columns are categorical via astype('category'). I cannot do\n&gt; that with sparse matrix. :( Any workaround?\n\nAs far as I know, LightGBM still supports categorical feature handling when passing a sparse matrix to construct a lgb dataset. You just have to make sure that you get your feature name list correctly ordered, and then pass a sublist to the categorical_feature argument, as below:\n\n    lgb_train = lgb.Dataset(X_train, \n                            label=y_train,\n                            feature_name=feature_names, \n                            categorical_feature=categorical_features)",
      "votes": null
    },
    {
      "id": "344816",
      "postDate": "06/18/2018 18:12:54",
      "content": "<p>Joe is right. In this competition the way I handle sparse matrix is simply to first build a list of all categoricals and numericals. Call that feature_names and then append to it the sparse names from the vocabulary of the vectorizer. Then in the Dataset object in lightgbm use what Joe has written.</p>",
      "rawMarkdown": "Joe is right. In this competition the way I handle sparse matrix is simply to first build a list of all categoricals and numericals. Call that feature_names and then append to it the sparse names from the vocabulary of the vectorizer. Then in the Dataset object in lightgbm use what Joe has written.",
      "votes": null
    },
    {
      "id": "344854",
      "postDate": "06/18/2018 19:30:36",
      "content": "<p>Yes, lgb datasets are very efficient. But so are sparse matrices  ¯_(ツ)_/¯. \nAnd they are also pretty universal. I doubt that lgb.dataset would work in sklearn, for example. </p>\n\n<p>Maybe lgb.dataset and xgboost.Dmatrix have some really cool advantages that I don't know about. Googling didn't help, so I asked here.</p>",
      "rawMarkdown": "Yes, lgb datasets are very efficient. But so are sparse matrices  ¯\\_(ツ)_/¯. \nAnd they are also pretty universal. I doubt that lgb.dataset would work in sklearn, for example. \n\nMaybe lgb.dataset and xgboost.Dmatrix have some really cool advantages that I don't know about. Googling didn't help, so I asked here.",
      "votes": null
    },
    {
      "id": "344855",
      "postDate": "06/18/2018 19:33:45",
      "content": "<p>I tried label encoding and one hot encoding. It works. But it seems that lightGBM know better than me how to handle those cat features. </p>",
      "rawMarkdown": "I tried label encoding and one hot encoding. It works. But it seems that lightGBM know better than me how to handle those cat features.",
      "votes": null
    },
    {
      "id": "345169",
      "postDate": "06/19/2018 10:34:01",
      "content": "<p>\"use number for index, e.g. categorical_feature=0,1,2 means column_0, column_1 and column_2 are categorical features\"</p>",
      "rawMarkdown": "\"use number for index, e.g. categorical_feature=0,1,2 means column_0, column_1 and column_2 are categorical features\"",
      "votes": null
    },
    {
      "id": "488813",
      "postDate": "03/13/2019 03:01:09",
      "content": "<p>I know this is an old post but... to answer question 3, what are Dataset and DMatrix for? For lightgbm at least, I found that when you make a model with mdl=LGBMRegressor(...), this sets some parameters up.  With your DataFrames (numpy arrays, whatever) x and y, if you run mdl.fit(x, y, ...) it in fact creates a Dataset from x and y, sets a few parameters based on what type of model mdl is (LGBMRegressor, Classifier, or Ranker), passes your parameters into param, and runs lightgbm.train(Dataset, params, ...)  lightGBM and XGBoost both ultimately do much of their work C-side (and optionally CUDA), I assume the Dataset and DMatrix match the format their C-side code expects.  </p>",
      "rawMarkdown": "I know this is an old post but... to answer question 3, what are Dataset and DMatrix for? For lightgbm at least, I found that when you make a model with mdl=LGBMRegressor(...), this sets some parameters up.  With your DataFrames (numpy arrays, whatever) x and y, if you run mdl.fit(x, y, ...) it in fact creates a Dataset from x and y, sets a few parameters based on what type of model mdl is (LGBMRegressor, Classifier, or Ranker), passes your parameters into param, and runs lightgbm.train(Dataset, params, ...)  lightGBM and XGBoost both ultimately do much of their work C-side (and optionally CUDA), I assume the Dataset and DMatrix match the format their C-side code expects.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 344779,
      "author_name": "gpreda",
      "author_url": "",
      "post_date": "06/18/2018 16:58:27",
      "content": "<p>Well, I am not an experienced Kaggler but I would like to suggest few ideas: <br>\n2) you can replace category features with dummy features (less optimal, but you can use with sparse matrix); actually, some claims that using dummy features is more efficient (in terms of performance) than using the categorical features list with lgb;     </p>\n\n<p>3) lightgbm.Dataset   (see: <a href=\"https://media.readthedocs.org/pdf/lightgbm/latest/lightgbm.pdf\">https://media.readthedocs.org/pdf/lightgbm/latest/lightgbm.pdf</a>, pag 12):  </p>\n\n<p>&gt; Memory efficent usage:   The Dataset object in LightGBM is very\n&gt; memory-efficient, due to it only need to save discrete bins. However,\n&gt; Numpy/Array/Pandas object is memory cost. If you concern about your\n&gt; memory consumption, you can save memory according to following: <br>\n&gt; 1. Let free_raw_data=True (default is True) when constructing the Dataset <br>\n&gt; 2. Explicit set raw_data=None after the Dataset has been constructed <br>\n&gt; 3. Call gc</p>",
      "votes": null,
      "replies": [
        {
          "id": 344854,
          "author_name": "kuliksv",
          "author_url": "",
          "post_date": "06/18/2018 19:30:36",
          "content": "<p>Yes, lgb datasets are very efficient. But so are sparse matrices  ¯_(ツ)_/¯. \nAnd they are also pretty universal. I doubt that lgb.dataset would work in sklearn, for example. </p>\n\n<p>Maybe lgb.dataset and xgboost.Dmatrix have some really cool advantages that I don't know about. Googling didn't help, so I asked here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 344782,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "06/18/2018 17:03:15",
      "content": "<blockquote>\n  <p>Sparse matrices don't store feature names. Store names separately</p>\n</blockquote>\n\n<p>Storing names separately is the only way I know of.</p>\n\n<blockquote>\n  <p>my LightGBM showed some improvement when I explicitly specified that certain columns are categorical via astype('category'). I cannot do that with sparse matrix. :( Any workaround?</p>\n</blockquote>\n\n<p>You'll have to manually implement sort of categorical-&gt;numeric scheme that captures this. I'd recommend label encoding, target encoding, or one hot encoding.</p>\n\n<blockquote>\n  <p>Is there any way to use dense and sparse features simultaneously? All I can think of is hstacking dense features to sparse features last thing before fitting, but, in that case, problems 1 and 2 aren't solved.</p>\n</blockquote>\n\n<p>Hstacking is the only way I know of. The features are still the same features (though they do have to be numeric) but the result is sparse. So you're right you can't fix your problems this way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 344855,
          "author_name": "kuliksv",
          "author_url": "",
          "post_date": "06/18/2018 19:33:45",
          "content": "<p>I tried label encoding and one hot encoding. It works. But it seems that lightGBM know better than me how to handle those cat features. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 345169,
          "author_name": "sagol79",
          "author_url": "",
          "post_date": "06/19/2018 10:34:01",
          "content": "<p>\"use number for index, e.g. categorical_feature=0,1,2 means column_0, column_1 and column_2 are categorical features\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 344803,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "06/18/2018 17:32:52",
      "content": "<blockquote>\n  <p>2) When I was still using dense features (pandas DataFrame), my\n  LightGBM showed some improvement when I explicitly specified that\n  certain columns are categorical via astype('category'). I cannot do\n  that with sparse matrix. :( Any workaround?</p>\n</blockquote>\n\n<p>As far as I know, LightGBM still supports categorical feature handling when passing a sparse matrix to construct a lgb dataset. You just have to make sure that you get your feature name list correctly ordered, and then pass a sublist to the categorical_feature argument, as below:</p>\n\n<pre><code>lgb_train = lgb.Dataset(X_train, \n                        label=y_train,\n                        feature_name=feature_names, \n                        categorical_feature=categorical_features)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 344816,
          "author_name": "arroqc",
          "author_url": "",
          "post_date": "06/18/2018 18:12:54",
          "content": "<p>Joe is right. In this competition the way I handle sparse matrix is simply to first build a list of all categoricals and numericals. Call that feature_names and then append to it the sparse names from the vocabulary of the vectorizer. Then in the Dataset object in lightgbm use what Joe has written.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 488813,
      "author_name": "hwertz",
      "author_url": "",
      "post_date": "03/13/2019 03:01:09",
      "content": "<p>I know this is an old post but... to answer question 3, what are Dataset and DMatrix for? For lightgbm at least, I found that when you make a model with mdl=LGBMRegressor(...), this sets some parameters up.  With your DataFrames (numpy arrays, whatever) x and y, if you run mdl.fit(x, y, ...) it in fact creates a Dataset from x and y, sets a few parameters based on what type of model mdl is (LGBMRegressor, Classifier, or Ranker), passes your parameters into param, and runs lightgbm.train(Dataset, params, ...)  lightGBM and XGBoost both ultimately do much of their work C-side (and optionally CUDA), I assume the Dataset and DMatrix match the format their C-side code expects.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "344730": "In this competition, I finally switch to sparse matrices after my giant TFIDFs were unable to fit into memory. Should have done that a long time ago. So much more memory efficient! And faster too. \n\nBut now I have some new problems. I wander how more experienced kagglers deal with them.\n\n1) Sparse matrices don't store feature names. Store names separately? That’s not very convenient. \n\n2) When I was still using dense features (pandas DataFrame), my LightGBM showed some improvement when I explicitly specified that certain columns are categorical via astype('category'). I cannot do that with sparse matrix. :( Any workaround?\n\n3) What are lightgbm.Dataset and xgb.DMatrix for? Both xgboost and lightgbm can work with DataFrames, numpy arrays, sparse matrices. Why another data format? What are the benefits?\n\n4) Is there any way to use dense and sparse features simultaneously? All I can think of is hstacking dense features to sparse features last thing before fitting, but, in that case, problems 1 and 2 aren't solved.",
    "344779": "Well, I am not an experienced Kaggler but I would like to suggest few ideas:  \n2) you can replace category features with dummy features (less optimal, but you can use with sparse matrix); actually, some claims that using dummy features is more efficient (in terms of performance) than using the categorical features list with lgb;     \n\n3) lightgbm.Dataset   (see: https://media.readthedocs.org/pdf/lightgbm/latest/lightgbm.pdf, pag 12):  \n\n&gt; Memory efficent usage:   The Dataset object in LightGBM is very\n&gt; memory-efficient, due to it only need to save discrete bins. However,\n&gt; Numpy/Array/Pandas object is memory cost. If you concern about your\n&gt; memory consumption, you can save memory according to following:  \n&gt; 1. Let free_raw_data=True (default is True) when constructing the Dataset  \n&gt; 2. Explicit set raw_data=None after the Dataset has been constructed  \n&gt; 3. Call gc",
    "344782": "&gt; Sparse matrices don't store feature names. Store names separately\n\nStoring names separately is the only way I know of.\n\n&gt; my LightGBM showed some improvement when I explicitly specified that certain columns are categorical via astype('category'). I cannot do that with sparse matrix. :( Any workaround?\n\nYou'll have to manually implement sort of categorical-&gt;numeric scheme that captures this. I'd recommend label encoding, target encoding, or one hot encoding.\n\n&gt; Is there any way to use dense and sparse features simultaneously? All I can think of is hstacking dense features to sparse features last thing before fitting, but, in that case, problems 1 and 2 aren't solved.\n\nHstacking is the only way I know of. The features are still the same features (though they do have to be numeric) but the result is sparse. So you're right you can't fix your problems this way.",
    "344803": "&gt; 2) When I was still using dense features (pandas DataFrame), my\n&gt; LightGBM showed some improvement when I explicitly specified that\n&gt; certain columns are categorical via astype('category'). I cannot do\n&gt; that with sparse matrix. :( Any workaround?\n\nAs far as I know, LightGBM still supports categorical feature handling when passing a sparse matrix to construct a lgb dataset. You just have to make sure that you get your feature name list correctly ordered, and then pass a sublist to the categorical_feature argument, as below:\n\n    lgb_train = lgb.Dataset(X_train, \n                            label=y_train,\n                            feature_name=feature_names, \n                            categorical_feature=categorical_features)",
    "344816": "Joe is right. In this competition the way I handle sparse matrix is simply to first build a list of all categoricals and numericals. Call that feature_names and then append to it the sparse names from the vocabulary of the vectorizer. Then in the Dataset object in lightgbm use what Joe has written.",
    "344854": "Yes, lgb datasets are very efficient. But so are sparse matrices  ¯\\_(ツ)_/¯. \nAnd they are also pretty universal. I doubt that lgb.dataset would work in sklearn, for example. \n\nMaybe lgb.dataset and xgboost.Dmatrix have some really cool advantages that I don't know about. Googling didn't help, so I asked here.",
    "344855": "I tried label encoding and one hot encoding. It works. But it seems that lightGBM know better than me how to handle those cat features.",
    "345169": "\"use number for index, e.g. categorical_feature=0,1,2 means column_0, column_1 and column_2 are categorical features\"",
    "488813": "I know this is an old post but... to answer question 3, what are Dataset and DMatrix for? For lightgbm at least, I found that when you make a model with mdl=LGBMRegressor(...), this sets some parameters up.  With your DataFrames (numpy arrays, whatever) x and y, if you run mdl.fit(x, y, ...) it in fact creates a Dataset from x and y, sets a few parameters based on what type of model mdl is (LGBMRegressor, Classifier, or Ranker), passes your parameters into param, and runs lightgbm.train(Dataset, params, ...)  lightGBM and XGBoost both ultimately do much of their work C-side (and optionally CUDA), I assume the Dataset and DMatrix match the format their C-side code expects."
  },
  "source": "meta"
}