{
  "id": 58819,
  "title": "TFidfVectorizer: ngram_range",
  "url": "/competitions/avito-demand-prediction/discussion/58819",
  "author_name": "",
  "post_date": "2018-06-14T06:00:22.532914600Z",
  "votes": 10,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Here is what I understand:-</p>\n\n<ul>\n<li><p>TFidfVectorizer (in sklearn python) converts the text documents to a matrix of \"tfidf\" features</p>\n\n<ul><li>tfidf features are the method to convert the textual information into the vector space, they are a \nmeasure of how important a word in a text is</li></ul></li>\n<li><p>ngram is the number of words in a sequence</p></li>\n</ul>\n\n<p>Here is what I dont understand:-</p>\n\n<ul>\n<li>What is the significance \"ngram_range\" parameter that we pass while creating an instance of TFidfVectorizer  </li>\n</ul>\n\n<p>Any help on this will be appreciated, I looked into the official documentation but it did not help much.</p>\n\n<p>Thank you! </p>",
  "messages": [
    {
      "id": "342799",
      "postDate": "06/14/2018 06:00:22",
      "content": "<p>Here is what I understand:-</p>\n\n<ul>\n<li><p>TFidfVectorizer (in sklearn python) converts the text documents to a matrix of \"tfidf\" features</p>\n\n<ul><li>tfidf features are the method to convert the textual information into the vector space, they are a \nmeasure of how important a word in a text is</li></ul></li>\n<li><p>ngram is the number of words in a sequence</p></li>\n</ul>\n\n<p>Here is what I dont understand:-</p>\n\n<ul>\n<li>What is the significance \"ngram_range\" parameter that we pass while creating an instance of TFidfVectorizer  </li>\n</ul>\n\n<p>Any help on this will be appreciated, I looked into the official documentation but it did not help much.</p>\n\n<p>Thank you! </p>",
      "rawMarkdown": "Here is what I understand:-\n\n  - TFidfVectorizer (in sklearn python) converts the text documents to a matrix of \"tfidf\" features\n\n   -  tfidf features are the method to convert the textual information into the vector space, they are a \n      measure of how important a word in a text is\n\n  - ngram is the number of words in a sequence\n\nHere is what I dont understand:-\n\n -  What is the significance \"ngram_range\" parameter that we pass while creating an instance of TFidfVectorizer  \n\n\nAny help on this will be appreciated, I looked into the official documentation but it did not help much.\n\nThank you!",
      "votes": null
    },
    {
      "id": "342805",
      "postDate": "06/14/2018 06:32:50",
      "content": "<p>Its about treating ngrams additionally. Let me illustrate with a simple example. \"very expensive\" is a 2-gram that is considered as an extra feature separatly from \"very\" and \"expensive\" when you have a n-gram range of (1,2) </p>",
      "rawMarkdown": "Its about treating ngrams additionally. Let me illustrate with a simple example. \"very expensive\" is a 2-gram that is considered as an extra feature separatly from \"very\" and \"expensive\" when you have a n-gram range of (1,2)",
      "votes": null
    },
    {
      "id": "342819",
      "postDate": "06/14/2018 07:14:54",
      "content": "<p>okay!</p>\n\n<p>so does it mean that if we have \"very expensive watch\" and the ngram_range(1,3), then the features will be\n\"very\", \"expensive\" and \"watch\" along with an additional feature of \"very expensive watch\".</p>\n\n<p>But if we have ngram_range(1,2) then the features will be \"very\", \"expensive\", \"watch\" and additional features with two grams e.g. \"very expensive\" or \"expensive watch\".</p>\n\n<p>Did I get that right?</p>",
      "rawMarkdown": "okay!\n\nso does it mean that if we have \"very expensive watch\" and the ngram_range(1,3), then the features will be\n\"very\", \"expensive\" and \"watch\" along with an additional feature of \"very expensive watch\".\n\nBut if we have ngram_range(1,2) then the features will be \"very\", \"expensive\", \"watch\" and additional features with two grams e.g. \"very expensive\" or \"expensive watch\".\n\nDid I get that right?",
      "votes": null
    },
    {
      "id": "342996",
      "postDate": "06/14/2018 14:10:02",
      "content": "<p>Here is a quick example that I could come up with to explain how the ngram range works in sklearn:</p>\n\n<p><strong>Example: range(1,2)</strong>\nv=text.CountVectorizer(ngram_range=(1,2))\nprint(v.fit([\"an apple a day keeps the doctor away\"]).vocabulary_)</p>\n\n<p>Result:\n{'an': 0, 'apple': 2, 'day': 5, 'keeps': 9, 'the': 11, 'doctor': 7, 'away': 4, 'an apple': 1, 'apple day': 3, 'day keeps': 6, 'keeps the': 10, 'the doctor': 12, 'doctor away': 8}</p>\n\n<p><strong>Example: range(1,3)</strong>\nv=text.CountVectorizer(ngram_range=(1,3))\nprint(v.fit([\"an apple a day keeps the doctor away\"]).vocabulary_)</p>\n\n<p>Result:\n{'an': 0, 'apple': 3, 'day': 7, 'keeps': 12, 'the': 15, 'doctor': 10, 'away': 6, 'an apple': 1, 'apple day': 4, 'day keeps': 8, 'keeps the': 13, 'the doctor': 16, 'doctor away': 11, 'an apple day': 2, 'apple day keeps': 5, 'day keeps the': 9, 'keeps the doctor': 14, 'the doctor away': 17}</p>\n\n<p>Hope it explains...</p>",
      "rawMarkdown": "Here is a quick example that I could come up with to explain how the ngram range works in sklearn:\n\n**Example: range(1,2)**\nv=text.CountVectorizer(ngram_range=(1,2))\nprint(v.fit([\"an apple a day keeps the doctor away\"]).vocabulary_)\n\nResult:\n{'an': 0, 'apple': 2, 'day': 5, 'keeps': 9, 'the': 11, 'doctor': 7, 'away': 4, 'an apple': 1, 'apple day': 3, 'day keeps': 6, 'keeps the': 10, 'the doctor': 12, 'doctor away': 8}\n\n**Example: range(1,3)**\nv=text.CountVectorizer(ngram_range=(1,3))\nprint(v.fit([\"an apple a day keeps the doctor away\"]).vocabulary_)\n\nResult:\n{'an': 0, 'apple': 3, 'day': 7, 'keeps': 12, 'the': 15, 'doctor': 10, 'away': 6, 'an apple': 1, 'apple day': 4, 'day keeps': 8, 'keeps the': 13, 'the doctor': 16, 'doctor away': 11, 'an apple day': 2, 'apple day keeps': 5, 'day keeps the': 9, 'keeps the doctor': 14, 'the doctor away': 17}\n\nHope it explains...",
      "votes": null
    },
    {
      "id": "343012",
      "postDate": "06/14/2018 14:49:30",
      "content": "<p>Yes ! it explained a lot, but I was wondering why it did not take into account \"a\"?</p>",
      "rawMarkdown": "Yes ! it explained a lot, but I was wondering why it did not take into account \"a\"?",
      "votes": null
    },
    {
      "id": "343247",
      "postDate": "06/15/2018 00:41:52",
      "content": "<p>Yes that is because of the default token pattern in Count Vectorizer which ignores uni-character words. In a sense it is doing it correct because it has no information. You could alter this by passing the token pattern as an additional language.</p>\n\n<p>v=CountVectorizer(ngram_range=(1,2),token_pattern='\\b\\w+\\b')</p>",
      "rawMarkdown": "Yes that is because of the default token pattern in Count Vectorizer which ignores uni-character words. In a sense it is doing it correct because it has no information. You could alter this by passing the token pattern as an additional language.\n\nv=CountVectorizer(ngram_range=(1,2),token_pattern='\\\\b\\\\w+\\\\b')",
      "votes": null
    },
    {
      "id": "343252",
      "postDate": "06/15/2018 00:58:54",
      "content": "<p>Thanks  @Vishy, you helped a lot! :)</p>",
      "rawMarkdown": "Thanks  @Vishy, you helped a lot! :)",
      "votes": null
    },
    {
      "id": "599545",
      "postDate": "08/15/2019 05:22:55",
      "content": "<p>by adding n-gram range how is it helping the model for better performance? I did not get the idea of getting this additional feature?</p>",
      "rawMarkdown": "by adding n-gram range how is it helping the model for better performance? I did not get the idea of getting this additional feature?",
      "votes": null
    },
    {
      "id": "689000",
      "postDate": "12/06/2019 10:28:01",
      "content": "<p><a href=\"/viswanathravindran\">@viswanathravindran</a>  (1,2) stands for minimum and maximum values the ngram_range can take? \nif above statement is true i can't understand your example.\n<a href=\"/ashukr\">@ashukr</a> <a href=\"/christofhenkel\">@christofhenkel</a> </p>",
      "rawMarkdown": "viswanathravindran  (1,2) stands for minimum and maximum values the ngram_range can take? \nif above statement is true i can't understand your example.\n@ashukr @christofhenkel",
      "votes": null
    },
    {
      "id": "887915",
      "postDate": "06/16/2020 01:58:12",
      "content": "<p>Yes ;)</p>",
      "rawMarkdown": "Yes ;)",
      "votes": null
    },
    {
      "id": "919648",
      "postDate": "07/08/2020 03:18:09",
      "content": "<p>1) ngram_range=(1,1)</p>\n\n<p>我爱你\n你爱我</p>\n\n<p>我: 1\n爱: 1\n你: 1</p>\n\n<p>我爱你 ＝ 你爱我 ??? </p>\n\n<h2>当然不是</h2>\n\n<p>1) ngram_range=(1,2)</p>\n\n<p>我爱你\n你爱我\n↓\n我, 爱, 你, 我爱, 爱你, 你爱, 爱我</p>\n\n<p>我爱你:\n我:1, 爱:1, 你:1, 我爱:1, 爱你:1, 你爱: 0, 爱我: 0</p>\n\n<p>你爱我:\n你,1, 爱:1, 我:1, 你爱: 1, 爱我: 1, 我爱:0, 爱你:0</p>\n\n<p>我爱你 ≠ 你爱我</p>",
      "rawMarkdown": "1) ngram_range=(1,1)\n\n我爱你\n你爱我\n\n我: 1\n爱: 1\n你: 1\n\n我爱你 ＝ 你爱我 ??? \n当然不是\n--------------------------------------\n1) ngram_range=(1,2)\n\n我爱你\n你爱我\n↓\n我, 爱, 你, 我爱, 爱你, 你爱, 爱我\n\n我爱你:\n我:1, 爱:1, 你:1, 我爱:1, 爱你:1, 你爱: 0, 爱我: 0\n\n你爱我:\n你,1, 爱:1, 我:1, 你爱: 1, 爱我: 1, 我爱:0, 爱你:0\n\n我爱你 ≠ 你爱我",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 342805,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "06/14/2018 06:32:50",
      "content": "<p>Its about treating ngrams additionally. Let me illustrate with a simple example. \"very expensive\" is a 2-gram that is considered as an extra feature separatly from \"very\" and \"expensive\" when you have a n-gram range of (1,2) </p>",
      "votes": null,
      "replies": [
        {
          "id": 342819,
          "author_name": "ashukr",
          "author_url": "",
          "post_date": "06/14/2018 07:14:54",
          "content": "<p>okay!</p>\n\n<p>so does it mean that if we have \"very expensive watch\" and the ngram_range(1,3), then the features will be\n\"very\", \"expensive\" and \"watch\" along with an additional feature of \"very expensive watch\".</p>\n\n<p>But if we have ngram_range(1,2) then the features will be \"very\", \"expensive\", \"watch\" and additional features with two grams e.g. \"very expensive\" or \"expensive watch\".</p>\n\n<p>Did I get that right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 342996,
          "author_name": "viswanathravindran",
          "author_url": "",
          "post_date": "06/14/2018 14:10:02",
          "content": "<p>Here is a quick example that I could come up with to explain how the ngram range works in sklearn:</p>\n\n<p><strong>Example: range(1,2)</strong>\nv=text.CountVectorizer(ngram_range=(1,2))\nprint(v.fit([\"an apple a day keeps the doctor away\"]).vocabulary_)</p>\n\n<p>Result:\n{'an': 0, 'apple': 2, 'day': 5, 'keeps': 9, 'the': 11, 'doctor': 7, 'away': 4, 'an apple': 1, 'apple day': 3, 'day keeps': 6, 'keeps the': 10, 'the doctor': 12, 'doctor away': 8}</p>\n\n<p><strong>Example: range(1,3)</strong>\nv=text.CountVectorizer(ngram_range=(1,3))\nprint(v.fit([\"an apple a day keeps the doctor away\"]).vocabulary_)</p>\n\n<p>Result:\n{'an': 0, 'apple': 3, 'day': 7, 'keeps': 12, 'the': 15, 'doctor': 10, 'away': 6, 'an apple': 1, 'apple day': 4, 'day keeps': 8, 'keeps the': 13, 'the doctor': 16, 'doctor away': 11, 'an apple day': 2, 'apple day keeps': 5, 'day keeps the': 9, 'keeps the doctor': 14, 'the doctor away': 17}</p>\n\n<p>Hope it explains...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 343012,
          "author_name": "ashukr",
          "author_url": "",
          "post_date": "06/14/2018 14:49:30",
          "content": "<p>Yes ! it explained a lot, but I was wondering why it did not take into account \"a\"?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 343247,
          "author_name": "viswanathravindran",
          "author_url": "",
          "post_date": "06/15/2018 00:41:52",
          "content": "<p>Yes that is because of the default token pattern in Count Vectorizer which ignores uni-character words. In a sense it is doing it correct because it has no information. You could alter this by passing the token pattern as an additional language.</p>\n\n<p>v=CountVectorizer(ngram_range=(1,2),token_pattern='\\b\\w+\\b')</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 343252,
          "author_name": "ashukr",
          "author_url": "",
          "post_date": "06/15/2018 00:58:54",
          "content": "<p>Thanks  @Vishy, you helped a lot! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 599545,
      "author_name": "rgalkar",
      "author_url": "",
      "post_date": "08/15/2019 05:22:55",
      "content": "<p>by adding n-gram range how is it helping the model for better performance? I did not get the idea of getting this additional feature?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 689000,
      "author_name": "kumaranu7",
      "author_url": "",
      "post_date": "12/06/2019 10:28:01",
      "content": "<p><a href=\"/viswanathravindran\">@viswanathravindran</a>  (1,2) stands for minimum and maximum values the ngram_range can take? \nif above statement is true i can't understand your example.\n<a href=\"/ashukr\">@ashukr</a> <a href=\"/christofhenkel\">@christofhenkel</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 887915,
          "author_name": "aianshay",
          "author_url": "",
          "post_date": "06/16/2020 01:58:12",
          "content": "<p>Yes ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 919648,
      "author_name": "chenyulong",
      "author_url": "",
      "post_date": "07/08/2020 03:18:09",
      "content": "<p>1) ngram_range=(1,1)</p>\n\n<p>我爱你\n你爱我</p>\n\n<p>我: 1\n爱: 1\n你: 1</p>\n\n<p>我爱你 ＝ 你爱我 ??? </p>\n\n<h2>当然不是</h2>\n\n<p>1) ngram_range=(1,2)</p>\n\n<p>我爱你\n你爱我\n↓\n我, 爱, 你, 我爱, 爱你, 你爱, 爱我</p>\n\n<p>我爱你:\n我:1, 爱:1, 你:1, 我爱:1, 爱你:1, 你爱: 0, 爱我: 0</p>\n\n<p>你爱我:\n你,1, 爱:1, 我:1, 你爱: 1, 爱我: 1, 我爱:0, 爱你:0</p>\n\n<p>我爱你 ≠ 你爱我</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "342799": "Here is what I understand:-\n\n  - TFidfVectorizer (in sklearn python) converts the text documents to a matrix of \"tfidf\" features\n\n   -  tfidf features are the method to convert the textual information into the vector space, they are a \n      measure of how important a word in a text is\n\n  - ngram is the number of words in a sequence\n\nHere is what I dont understand:-\n\n -  What is the significance \"ngram_range\" parameter that we pass while creating an instance of TFidfVectorizer  \n\n\nAny help on this will be appreciated, I looked into the official documentation but it did not help much.\n\nThank you!",
    "342805": "Its about treating ngrams additionally. Let me illustrate with a simple example. \"very expensive\" is a 2-gram that is considered as an extra feature separatly from \"very\" and \"expensive\" when you have a n-gram range of (1,2)",
    "342819": "okay!\n\nso does it mean that if we have \"very expensive watch\" and the ngram_range(1,3), then the features will be\n\"very\", \"expensive\" and \"watch\" along with an additional feature of \"very expensive watch\".\n\nBut if we have ngram_range(1,2) then the features will be \"very\", \"expensive\", \"watch\" and additional features with two grams e.g. \"very expensive\" or \"expensive watch\".\n\nDid I get that right?",
    "342996": "Here is a quick example that I could come up with to explain how the ngram range works in sklearn:\n\n**Example: range(1,2)**\nv=text.CountVectorizer(ngram_range=(1,2))\nprint(v.fit([\"an apple a day keeps the doctor away\"]).vocabulary_)\n\nResult:\n{'an': 0, 'apple': 2, 'day': 5, 'keeps': 9, 'the': 11, 'doctor': 7, 'away': 4, 'an apple': 1, 'apple day': 3, 'day keeps': 6, 'keeps the': 10, 'the doctor': 12, 'doctor away': 8}\n\n**Example: range(1,3)**\nv=text.CountVectorizer(ngram_range=(1,3))\nprint(v.fit([\"an apple a day keeps the doctor away\"]).vocabulary_)\n\nResult:\n{'an': 0, 'apple': 3, 'day': 7, 'keeps': 12, 'the': 15, 'doctor': 10, 'away': 6, 'an apple': 1, 'apple day': 4, 'day keeps': 8, 'keeps the': 13, 'the doctor': 16, 'doctor away': 11, 'an apple day': 2, 'apple day keeps': 5, 'day keeps the': 9, 'keeps the doctor': 14, 'the doctor away': 17}\n\nHope it explains...",
    "343012": "Yes ! it explained a lot, but I was wondering why it did not take into account \"a\"?",
    "343247": "Yes that is because of the default token pattern in Count Vectorizer which ignores uni-character words. In a sense it is doing it correct because it has no information. You could alter this by passing the token pattern as an additional language.\n\nv=CountVectorizer(ngram_range=(1,2),token_pattern='\\\\b\\\\w+\\\\b')",
    "343252": "Thanks  @Vishy, you helped a lot! :)",
    "599545": "by adding n-gram range how is it helping the model for better performance? I did not get the idea of getting this additional feature?",
    "689000": "viswanathravindran  (1,2) stands for minimum and maximum values the ngram_range can take? \nif above statement is true i can't understand your example.\n@ashukr @christofhenkel",
    "887915": "Yes ;)",
    "919648": "1) ngram_range=(1,1)\n\n我爱你\n你爱我\n\n我: 1\n爱: 1\n你: 1\n\n我爱你 ＝ 你爱我 ??? \n当然不是\n--------------------------------------\n1) ngram_range=(1,2)\n\n我爱你\n你爱我\n↓\n我, 爱, 你, 我爱, 爱你, 你爱, 爱我\n\n我爱你:\n我:1, 爱:1, 你:1, 我爱:1, 爱你:1, 你爱: 0, 爱我: 0\n\n你爱我:\n你,1, 爱:1, 我:1, 你爱: 1, 爱我: 1, 我爱:0, 爱你:0\n\n我爱你 ≠ 你爱我"
  },
  "source": "meta"
}