{
  "id": 56798,
  "title": "tfidf dense feature vs tfidf SVD component feature",
  "url": "/competitions/avito-demand-prediction/discussion/56798",
  "author_name": "",
  "post_date": "2018-05-15T06:14:03.581297400Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, mates:</p>\n\n<p>I'm a starter at kaggle, </p>\n\n<p>I tried two feature engineering with tfidf dense feature and svd with 200 components fit_transform tfifd features to 200 features,</p>\n\n<p>What happened my LB score decreased 0.007~0.01 after do that.</p>\n\n<p>I create this thread for discuss some trick method to reduce the tfidf dense feature dimensions and don't loss hight accuracy.</p>",
  "messages": [
    {
      "id": "328818",
      "postDate": "05/15/2018 06:14:03",
      "content": "<p>Hi, mates:</p>\n\n<p>I'm a starter at kaggle, </p>\n\n<p>I tried two feature engineering with tfidf dense feature and svd with 200 components fit_transform tfifd features to 200 features,</p>\n\n<p>What happened my LB score decreased 0.007~0.01 after do that.</p>\n\n<p>I create this thread for discuss some trick method to reduce the tfidf dense feature dimensions and don't loss hight accuracy.</p>",
      "rawMarkdown": "Hi, mates:\n\nI'm a starter at kaggle, \n\nI tried two feature engineering with tfidf dense feature and svd with 200 components fit_transform tfifd features to 200 features,\n\nWhat happened my LB score decreased 0.007~0.01 after do that.\n\nI create this thread for discuss some trick method to reduce the tfidf dense feature dimensions and don't loss hight accuracy.",
      "votes": null
    },
    {
      "id": "328907",
      "postDate": "05/15/2018 10:50:26",
      "content": "<p>If the memory is a concern, why not use sparse tf-idf representation?</p>",
      "rawMarkdown": "If the memory is a concern, why not use sparse tf-idf representation?",
      "votes": null
    },
    {
      "id": "328923",
      "postDate": "05/15/2018 11:41:01",
      "content": "<p>I traind a model with dense features, there are some many vocabulary have no any contribution to my lgbm model. And it took a long time to fit. I o really want more quickly verify my other feature with a shorter time</p>",
      "rawMarkdown": "I traind a model with dense features, there are some many vocabulary have no any contribution to my lgbm model. And it took a long time to fit. I o really want more quickly verify my other feature with a shorter time",
      "votes": null
    },
    {
      "id": "328934",
      "postDate": "05/15/2018 11:58:55",
      "content": "<p>I don't think reducing the data set ( O(100000) -&gt; 200 features in this case) will ever make the score go up...</p>",
      "rawMarkdown": "I don't think reducing the data set ( O(100000) -&gt; 200 features in this case) will ever make the score go up...",
      "votes": null
    },
    {
      "id": "329097",
      "postDate": "05/15/2018 17:55:54",
      "content": "<p>The potential problem with SVD feature reduction is that it's unsupervised. It projects the original features in a way that can express a good deal of the variance of the original features in fewer dimensions, but directions of maximum variance do not necessarily capture the feature information with the most predictive signal. In contrast, using raw tf-idf features allows the model to discover for itself in a supervised manner where predictive information can be found. Supervised information extraction will often outperform unsupervised methods for the purposes of prediction.    </p>\n\n<p>Here are four possible approaches that may help you still get dimensionality reduction but with less loss of signal:</p>\n\n<ol>\n<li><p>Adjust your SVD parameters - it may work better with significantly more than 200 components. You can use the percent of variance explained ratio as a guide for selecting number of components.</p></li>\n<li><p>Adjust your tf-idf parameters / use the raw features more selectively. You can use term frequency cutoffs (min_df and max_df in sklearn) and max_features to reduce the number of word features extracted. You can also use feature importances as you've already done to figure out which tf-idf features the model finds useless, and simply remove those columns from your training data. E.g. if you use this to reduce to the 200 word features with highest feature importance, you may find that your results are significantly better than using 200 SVD components.</p></li>\n<li><p>Use an alternate unsupervised dimensionality reduction technique. It may be worth trying a different style of concept/topic model than SVD, for example NMF or LDA.</p></li>\n<li><p>Use neural-network based (supervised) dimensionality reduction techniques. There are a ton of different ways you could try to do this. You could run raw tf-idf features through a (reduced dimension) dense layer as input to a neural network then extract the features from that dense layer (this is in essence actually exactly what SVD does - linear combination of original features, except that here you can introduce a non-linearity and learn features with backpropagation instead of matrix decomposition). You could train a recurrent model on the text fields that takes word embeddings as inputs, extract the embeddings, and average the embedding values across a description to get description-level vectors. Lots of different creative avenues for approaching this - leaving it as the last approach since it's harder and more open-ended.</p></li>\n</ol>",
      "rawMarkdown": "The potential problem with SVD feature reduction is that it's unsupervised. It projects the original features in a way that can express a good deal of the variance of the original features in fewer dimensions, but directions of maximum variance do not necessarily capture the feature information with the most predictive signal. In contrast, using raw tf-idf features allows the model to discover for itself in a supervised manner where predictive information can be found. Supervised information extraction will often outperform unsupervised methods for the purposes of prediction.    \n\nHere are four possible approaches that may help you still get dimensionality reduction but with less loss of signal:\n\n1.  Adjust your SVD parameters - it may work better with significantly more than 200 components. You can use the percent of variance explained ratio as a guide for selecting number of components.\n\n2. Adjust your tf-idf parameters / use the raw features more selectively. You can use term frequency cutoffs (min_df and max_df in sklearn) and max_features to reduce the number of word features extracted. You can also use feature importances as you've already done to figure out which tf-idf features the model finds useless, and simply remove those columns from your training data. E.g. if you use this to reduce to the 200 word features with highest feature importance, you may find that your results are significantly better than using 200 SVD components.\n\n3. Use an alternate unsupervised dimensionality reduction technique. It may be worth trying a different style of concept/topic model than SVD, for example NMF or LDA.\n\n4. Use neural-network based (supervised) dimensionality reduction techniques. There are a ton of different ways you could try to do this. You could run raw tf-idf features through a (reduced dimension) dense layer as input to a neural network then extract the features from that dense layer (this is in essence actually exactly what SVD does - linear combination of original features, except that here you can introduce a non-linearity and learn features with backpropagation instead of matrix decomposition). You could train a recurrent model on the text fields that takes word embeddings as inputs, extract the embeddings, and average the embedding values across a description to get description-level vectors. Lots of different creative avenues for approaching this - leaving it as the last approach since it's harder and more open-ended.",
      "votes": null
    },
    {
      "id": "329186",
      "postDate": "05/16/2018 00:12:54",
      "content": "<p>When I used both  dense feature and svd feature with lgbm, my local validation score got worse. (I did not submit, so I did not know LB score.) Why did this thing happend?</p>",
      "rawMarkdown": "When I used both  dense feature and svd feature with lgbm, my local validation score got worse. (I did not submit, so I did not know LB score.) Why did this thing happend?",
      "votes": null
    },
    {
      "id": "330712",
      "postDate": "05/19/2018 14:09:54",
      "content": "<p>I thank you very much for your careful and patient answer, which made me great.</p>\n\n<p>Now I tried merge all textual features (raw tf-idf feature, SDV, NMF, LDA), then it improved 0.0002</p>",
      "rawMarkdown": "I thank you very much for your careful and patient answer, which made me great.\n\nNow I tried merge all textual features (raw tf-idf feature, SDV, NMF, LDA), then it improved 0.0002",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 328907,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "05/15/2018 10:50:26",
      "content": "<p>If the memory is a concern, why not use sparse tf-idf representation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 328923,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "05/15/2018 11:41:01",
          "content": "<p>I traind a model with dense features, there are some many vocabulary have no any contribution to my lgbm model. And it took a long time to fit. I o really want more quickly verify my other feature with a shorter time</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 328934,
      "author_name": "metadist",
      "author_url": "",
      "post_date": "05/15/2018 11:58:55",
      "content": "<p>I don't think reducing the data set ( O(100000) -&gt; 200 features in this case) will ever make the score go up...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 329097,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "05/15/2018 17:55:54",
      "content": "<p>The potential problem with SVD feature reduction is that it's unsupervised. It projects the original features in a way that can express a good deal of the variance of the original features in fewer dimensions, but directions of maximum variance do not necessarily capture the feature information with the most predictive signal. In contrast, using raw tf-idf features allows the model to discover for itself in a supervised manner where predictive information can be found. Supervised information extraction will often outperform unsupervised methods for the purposes of prediction.    </p>\n\n<p>Here are four possible approaches that may help you still get dimensionality reduction but with less loss of signal:</p>\n\n<ol>\n<li><p>Adjust your SVD parameters - it may work better with significantly more than 200 components. You can use the percent of variance explained ratio as a guide for selecting number of components.</p></li>\n<li><p>Adjust your tf-idf parameters / use the raw features more selectively. You can use term frequency cutoffs (min_df and max_df in sklearn) and max_features to reduce the number of word features extracted. You can also use feature importances as you've already done to figure out which tf-idf features the model finds useless, and simply remove those columns from your training data. E.g. if you use this to reduce to the 200 word features with highest feature importance, you may find that your results are significantly better than using 200 SVD components.</p></li>\n<li><p>Use an alternate unsupervised dimensionality reduction technique. It may be worth trying a different style of concept/topic model than SVD, for example NMF or LDA.</p></li>\n<li><p>Use neural-network based (supervised) dimensionality reduction techniques. There are a ton of different ways you could try to do this. You could run raw tf-idf features through a (reduced dimension) dense layer as input to a neural network then extract the features from that dense layer (this is in essence actually exactly what SVD does - linear combination of original features, except that here you can introduce a non-linearity and learn features with backpropagation instead of matrix decomposition). You could train a recurrent model on the text fields that takes word embeddings as inputs, extract the embeddings, and average the embedding values across a description to get description-level vectors. Lots of different creative avenues for approaching this - leaving it as the last approach since it's harder and more open-ended.</p></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 330712,
          "author_name": "classtag",
          "author_url": "",
          "post_date": "05/19/2018 14:09:54",
          "content": "<p>I thank you very much for your careful and patient answer, which made me great.</p>\n\n<p>Now I tried merge all textual features (raw tf-idf feature, SDV, NMF, LDA), then it improved 0.0002</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 329186,
      "author_name": "takuok",
      "author_url": "",
      "post_date": "05/16/2018 00:12:54",
      "content": "<p>When I used both  dense feature and svd feature with lgbm, my local validation score got worse. (I did not submit, so I did not know LB score.) Why did this thing happend?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "328818": "Hi, mates:\n\nI'm a starter at kaggle, \n\nI tried two feature engineering with tfidf dense feature and svd with 200 components fit_transform tfifd features to 200 features,\n\nWhat happened my LB score decreased 0.007~0.01 after do that.\n\nI create this thread for discuss some trick method to reduce the tfidf dense feature dimensions and don't loss hight accuracy.",
    "328907": "If the memory is a concern, why not use sparse tf-idf representation?",
    "328923": "I traind a model with dense features, there are some many vocabulary have no any contribution to my lgbm model. And it took a long time to fit. I o really want more quickly verify my other feature with a shorter time",
    "328934": "I don't think reducing the data set ( O(100000) -&gt; 200 features in this case) will ever make the score go up...",
    "329097": "The potential problem with SVD feature reduction is that it's unsupervised. It projects the original features in a way that can express a good deal of the variance of the original features in fewer dimensions, but directions of maximum variance do not necessarily capture the feature information with the most predictive signal. In contrast, using raw tf-idf features allows the model to discover for itself in a supervised manner where predictive information can be found. Supervised information extraction will often outperform unsupervised methods for the purposes of prediction.    \n\nHere are four possible approaches that may help you still get dimensionality reduction but with less loss of signal:\n\n1.  Adjust your SVD parameters - it may work better with significantly more than 200 components. You can use the percent of variance explained ratio as a guide for selecting number of components.\n\n2. Adjust your tf-idf parameters / use the raw features more selectively. You can use term frequency cutoffs (min_df and max_df in sklearn) and max_features to reduce the number of word features extracted. You can also use feature importances as you've already done to figure out which tf-idf features the model finds useless, and simply remove those columns from your training data. E.g. if you use this to reduce to the 200 word features with highest feature importance, you may find that your results are significantly better than using 200 SVD components.\n\n3. Use an alternate unsupervised dimensionality reduction technique. It may be worth trying a different style of concept/topic model than SVD, for example NMF or LDA.\n\n4. Use neural-network based (supervised) dimensionality reduction techniques. There are a ton of different ways you could try to do this. You could run raw tf-idf features through a (reduced dimension) dense layer as input to a neural network then extract the features from that dense layer (this is in essence actually exactly what SVD does - linear combination of original features, except that here you can introduce a non-linearity and learn features with backpropagation instead of matrix decomposition). You could train a recurrent model on the text fields that takes word embeddings as inputs, extract the embeddings, and average the embedding values across a description to get description-level vectors. Lots of different creative avenues for approaching this - leaving it as the last approach since it's harder and more open-ended.",
    "329186": "When I used both  dense feature and svd feature with lgbm, my local validation score got worse. (I did not submit, so I did not know LB score.) Why did this thing happend?",
    "330712": "I thank you very much for your careful and patient answer, which made me great.\n\nNow I tried merge all textual features (raw tf-idf feature, SDV, NMF, LDA), then it improved 0.0002"
  },
  "source": "meta"
}