{
  "id": 76590,
  "title": "NLP feature extraction",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76590",
  "author_name": "",
  "post_date": "2019-01-04T13:25:14.485787300Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>How can we extract the features, the sentence holds. <br>\nEx: gud performance gud low ground clearance compared suv counted delicate suvs fuel economy gud poweful suv</p>\n\n<p>can we extract : Features - performance, low ground clearance, poweful suv\nthanks</p>",
  "messages": [
    {
      "id": "450192",
      "postDate": "01/04/2019 13:25:14",
      "content": "<p>How can we extract the features, the sentence holds. <br>\nEx: gud performance gud low ground clearance compared suv counted delicate suvs fuel economy gud poweful suv</p>\n\n<p>can we extract : Features - performance, low ground clearance, poweful suv\nthanks</p>",
      "rawMarkdown": "How can we extract the features, the sentence holds.  \nEx: gud performance gud low ground clearance compared suv counted delicate suvs fuel economy gud poweful suv\n\ncan we extract : Features - performance, low ground clearance, poweful suv\nthanks",
      "votes": null
    },
    {
      "id": "450208",
      "postDate": "01/04/2019 13:48:06",
      "content": "<p>It kind of depends what task you're working on and how you plan to process those features</p>\n\n<p>For example, many people in the Quora Insincere Questions Classification competition are using keras and RNNs. When you use an RNN like this it's typical to use the whole sentence (under the assumption that it's structure and the frequency/order of words contains a lot of information)</p>\n\n<p>RNNs take an input of tensors of numbers, not words so we need some way to replace sentences with a tensor. Keras has a class called <code>Tokenizer</code> which helps with this. Here is how you might use it</p>\n\n<p>```</p>\n\n<h1>create the tokeniser and declare how many words you want to keep track of</h1>\n\n<h1>this is something you may have to play with, in the Quora competition you see 95000 a lot</h1>\n\n<p>tokenizer = Tokenizer(num_words=max_features)</p>\n\n<h1>train the tokeniser on the data we have</h1>\n\n<h1>this will create a dictionary mapping ints -&gt; words, for example <code>1 -&gt; the</code> or <code>2133 -&gt; potato</code></h1>\n\n<p>tokenizer.fit_on_texts(list(train_X))</p>\n\n<h1>then use the tokenizer on our input data to replace sentences with an array of ints</h1>\n\n<p>train_X = tokenizer.texts_to_sequences(train_X)</p>\n\n<h1>finally we like to pad the array with zeros, so that everything is the same length</h1>\n\n<h1>again, the length is something you can play with</h1>\n\n<p>train_X = pad_sequences(train_X, maxlen=maxlen)\n```</p>\n\n<p>we can then use train_X as the input to our RNN</p>\n\n<p>If you'd like to see a full example, check out <a href=\"https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr\">this kernel</a> by <strong>shujian</strong></p>",
      "rawMarkdown": "It kind of depends what task you're working on and how you plan to process those features\n\nFor example, many people in the Quora Insincere Questions Classification competition are using keras and RNNs. When you use an RNN like this it's typical to use the whole sentence (under the assumption that it's structure and the frequency/order of words contains a lot of information)\n\nRNNs take an input of tensors of numbers, not words so we need some way to replace sentences with a tensor. Keras has a class called `Tokenizer` which helps with this. Here is how you might use it\n\n```\n# create the tokeniser and declare how many words you want to keep track of\n# this is something you may have to play with, in the Quora competition you see 95000 a lot\ntokenizer = Tokenizer(num_words=max_features)\n\n# train the tokeniser on the data we have\n# this will create a dictionary mapping ints -&gt; words, for example `1 -&gt; the` or `2133 -&gt; potato`\ntokenizer.fit_on_texts(list(train_X))\n\n# then use the tokenizer on our input data to replace sentences with an array of ints\ntrain_X = tokenizer.texts_to_sequences(train_X)\n\n# finally we like to pad the array with zeros, so that everything is the same length\n# again, the length is something you can play with\ntrain_X = pad_sequences(train_X, maxlen=maxlen)\n```\n\nwe can then use train_X as the input to our RNN\n\nIf you'd like to see a full example, check out [this kernel](https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr) by **shujian**",
      "votes": null
    },
    {
      "id": "456765",
      "postDate": "01/16/2019 13:48:06",
      "content": "<p>thnks for the knowledge mate </p>",
      "rawMarkdown": "thnks for the knowledge mate",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 450208,
      "author_name": "hamishdickson",
      "author_url": "",
      "post_date": "01/04/2019 13:48:06",
      "content": "<p>It kind of depends what task you're working on and how you plan to process those features</p>\n\n<p>For example, many people in the Quora Insincere Questions Classification competition are using keras and RNNs. When you use an RNN like this it's typical to use the whole sentence (under the assumption that it's structure and the frequency/order of words contains a lot of information)</p>\n\n<p>RNNs take an input of tensors of numbers, not words so we need some way to replace sentences with a tensor. Keras has a class called <code>Tokenizer</code> which helps with this. Here is how you might use it</p>\n\n<p>```</p>\n\n<h1>create the tokeniser and declare how many words you want to keep track of</h1>\n\n<h1>this is something you may have to play with, in the Quora competition you see 95000 a lot</h1>\n\n<p>tokenizer = Tokenizer(num_words=max_features)</p>\n\n<h1>train the tokeniser on the data we have</h1>\n\n<h1>this will create a dictionary mapping ints -&gt; words, for example <code>1 -&gt; the</code> or <code>2133 -&gt; potato</code></h1>\n\n<p>tokenizer.fit_on_texts(list(train_X))</p>\n\n<h1>then use the tokenizer on our input data to replace sentences with an array of ints</h1>\n\n<p>train_X = tokenizer.texts_to_sequences(train_X)</p>\n\n<h1>finally we like to pad the array with zeros, so that everything is the same length</h1>\n\n<h1>again, the length is something you can play with</h1>\n\n<p>train_X = pad_sequences(train_X, maxlen=maxlen)\n```</p>\n\n<p>we can then use train_X as the input to our RNN</p>\n\n<p>If you'd like to see a full example, check out <a href=\"https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr\">this kernel</a> by <strong>shujian</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 456765,
          "author_name": "vijay09",
          "author_url": "",
          "post_date": "01/16/2019 13:48:06",
          "content": "<p>thnks for the knowledge mate </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "450192": "How can we extract the features, the sentence holds.  \nEx: gud performance gud low ground clearance compared suv counted delicate suvs fuel economy gud poweful suv\n\ncan we extract : Features - performance, low ground clearance, poweful suv\nthanks",
    "450208": "It kind of depends what task you're working on and how you plan to process those features\n\nFor example, many people in the Quora Insincere Questions Classification competition are using keras and RNNs. When you use an RNN like this it's typical to use the whole sentence (under the assumption that it's structure and the frequency/order of words contains a lot of information)\n\nRNNs take an input of tensors of numbers, not words so we need some way to replace sentences with a tensor. Keras has a class called `Tokenizer` which helps with this. Here is how you might use it\n\n```\n# create the tokeniser and declare how many words you want to keep track of\n# this is something you may have to play with, in the Quora competition you see 95000 a lot\ntokenizer = Tokenizer(num_words=max_features)\n\n# train the tokeniser on the data we have\n# this will create a dictionary mapping ints -&gt; words, for example `1 -&gt; the` or `2133 -&gt; potato`\ntokenizer.fit_on_texts(list(train_X))\n\n# then use the tokenizer on our input data to replace sentences with an array of ints\ntrain_X = tokenizer.texts_to_sequences(train_X)\n\n# finally we like to pad the array with zeros, so that everything is the same length\n# again, the length is something you can play with\ntrain_X = pad_sequences(train_X, maxlen=maxlen)\n```\n\nwe can then use train_X as the input to our RNN\n\nIf you'd like to see a full example, check out [this kernel](https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr) by **shujian**",
    "456765": "thnks for the knowledge mate"
  },
  "source": "meta"
}