{
  "id": 75319,
  "title": "What is the number of max_features you use?Should we use max_features to limit the size of embedding matrix?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75319",
  "author_name": "little_snail",
  "post_date": "2018-12-20T14:02:55.013000",
  "votes": 0,
  "comment_count": 4,
  "views": 0,
  "content": "<p>There are many public kernels use max_features to limit the size of embedding matrix.\nIs it useful?Why?</p>",
  "messages": [
    {
      "id": 443052,
      "postDate": "2018-12-20T23:58:27.397Z",
      "content": "<p><code>max_features</code> is a hyperparameter pretty much like most others that can be tuned based on validation feedback and memory/time complexity considerations. </p>\n\n<p>From a model fitting perspective, more words might mean more useful information, but including increasingly rare words as you raise <code>max_features</code> might cause overfitting - you can find the right balance by seeing how your validation scheme responds.</p>\n\n<p>The time/memory perspective is arguably trickier. Larger embedding matrices are naturally more expensive to store in RAM. But I think that if you freeze the pre-trained embeddings your model should train about as fast no matter what <code>max_features</code> is, since the pretrained embedding matrix just acts like a lookup table. If you update the embeddings, increasing max features should definitely mean slower training because you're obviously learning a lot more weights.</p>",
      "rawMarkdown": "`max_features` is a hyperparameter pretty much like most others that can be tuned based on validation feedback and memory/time complexity considerations. \n\nFrom a model fitting perspective, more words might mean more useful information, but including increasingly rare words as you raise `max_features` might cause overfitting - you can find the right balance by seeing how your validation scheme responds.\n\nThe time/memory perspective is arguably trickier. Larger embedding matrices are naturally more expensive to store in RAM. But I think that if you freeze the pre-trained embeddings your model should train about as fast no matter what `max_features` is, since the pretrained embedding matrix just acts like a lookup table. If you update the embeddings, increasing max features should definitely mean slower training because you're obviously learning a lot more weights.",
      "votes": 7
    },
    {
      "id": 442960,
      "postDate": "2018-12-20T19:40:19.443Z",
      "content": "<p>I tried running my model with all the features which gave me a lower val loss but also a lower val F1 and public lb score. Did anyone else see results like this?</p>",
      "rawMarkdown": "I tried running my model with all the features which gave me a lower val loss but also a lower val F1 and public lb score. Did anyone else see results like this?",
      "votes": 1
    },
    {
      "id": 442787,
      "postDate": "2018-12-20T14:02:55.013Z",
      "content": "<p>There are many public kernels use max_features to limit the size of embedding matrix.\nIs it useful?Why?</p>",
      "rawMarkdown": "There are many public kernels use max_features to limit the size of embedding matrix.\nIs it useful?Why?"
    },
    {
      "id": 442827,
      "postDate": "2018-12-20T15:22:27.207Z",
      "content": "<p>.</p>",
      "rawMarkdown": ".",
      "votes": -1
    },
    {
      "id": 442879,
      "postDate": "2018-12-20T16:23:23.777Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 443052,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2018-12-20T23:58:27.397000",
      "content": "<p><code>max_features</code> is a hyperparameter pretty much like most others that can be tuned based on validation feedback and memory/time complexity considerations. </p>\n\n<p>From a model fitting perspective, more words might mean more useful information, but including increasingly rare words as you raise <code>max_features</code> might cause overfitting - you can find the right balance by seeing how your validation scheme responds.</p>\n\n<p>The time/memory perspective is arguably trickier. Larger embedding matrices are naturally more expensive to store in RAM. But I think that if you freeze the pre-trained embeddings your model should train about as fast no matter what <code>max_features</code> is, since the pretrained embedding matrix just acts like a lookup table. If you update the embeddings, increasing max features should definitely mean slower training because you're obviously learning a lot more weights.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 442960,
      "author_name": "bilal2vec",
      "author_url": "",
      "post_date": "2018-12-20T19:40:19.443000",
      "content": "<p>I tried running my model with all the features which gave me a lower val loss but also a lower val F1 and public lb score. Did anyone else see results like this?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 442827,
      "author_name": "ms",
      "author_url": "",
      "post_date": "2018-12-20T15:22:27.207000",
      "content": "<p>.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 442879,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-20T16:23:23.777000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "443052": "`max_features` is a hyperparameter pretty much like most others that can be tuned based on validation feedback and memory/time complexity considerations. \n\nFrom a model fitting perspective, more words might mean more useful information, but including increasingly rare words as you raise `max_features` might cause overfitting - you can find the right balance by seeing how your validation scheme responds.\n\nThe time/memory perspective is arguably trickier. Larger embedding matrices are naturally more expensive to store in RAM. But I think that if you freeze the pre-trained embeddings your model should train about as fast no matter what `max_features` is, since the pretrained embedding matrix just acts like a lookup table. If you update the embeddings, increasing max features should definitely mean slower training because you're obviously learning a lot more weights.",
    "442960": "I tried running my model with all the features which gave me a lower val loss but also a lower val F1 and public lb score. Did anyone else see results like this?",
    "442787": "There are many public kernels use max_features to limit the size of embedding matrix.\nIs it useful?Why?",
    "442827": ".",
    "442879": ""
  }
}