{
  "id": 55391,
  "title": "Joining Text and numeric",
  "url": "/competitions/avito-demand-prediction/discussion/55391",
  "author_name": "",
  "post_date": "2018-04-26T02:49:57.521058400Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>What are possible and best ways in combining numeric and text data into single model.</p>",
  "messages": [
    {
      "id": "319422",
      "postDate": "04/26/2018 02:49:57",
      "content": "<p>What are possible and best ways in combining numeric and text data into single model.</p>",
      "rawMarkdown": "What are possible and best ways in combining numeric and text data into single model.",
      "votes": null
    },
    {
      "id": "319442",
      "postDate": "04/26/2018 03:37:38",
      "content": "<p>First vectorize your text data using TFIDF Vectorizor (or CountVectorizer etc), Then use Hstack to combine text features and numeric features. </p>\n\n<pre><code>train_data = pd.read_csv('data.csv')\n\nPreparing yoru text data\ntext_data = train_data['text_feature']\ntfidf_vectorizer = TfidfVectorizer() \ntfidf_vectorizer.fit(text_data)\nX_tfidf = vect_word.transform(df[text_data) \n\n# Preparing your numerical data \nX_numerical = train_data[[' numerical_feature1', 'numerical_feature2']]\n\nfrom scipy.sparse import hstack, csr_matrix\nX_train = hstack([X_tfidf, csr_matrix(X_numerical)], 'csr')\n</code></pre>\n\n<p>And now you can use X_train in your model. Hope it Helps </p>",
      "rawMarkdown": "First vectorize your text data using TFIDF Vectorizor (or CountVectorizer etc), Then use Hstack to combine text features and numeric features. \n\n    train_data = pd.read_csv('data.csv')\n    \n    Preparing yoru text data\n    text_data = train_data['text_feature']\n    tfidf_vectorizer = TfidfVectorizer() \n    tfidf_vectorizer.fit(text_data)\n    X_tfidf = vect_word.transform(df[text_data) \n    \n    # Preparing your numerical data \n    X_numerical = train_data[[' numerical_feature1', 'numerical_feature2']]\n    \n    from scipy.sparse import hstack, csr_matrix\n    X_train = hstack([X_tfidf, csr_matrix(X_numerical)], 'csr')\n\nAnd now you can use X_train in your model. Hope it Helps",
      "votes": null
    },
    {
      "id": "319443",
      "postDate": "04/26/2018 03:40:46",
      "content": "<p>I have tried tfidf feature on this competition dataset. Kernel <a href=\"https://www.kaggle.com/classtag/russian-nlp-in-avito-demand-prediction\">https://www.kaggle.com/classtag/russian-nlp-in-avito-demand-prediction</a></p>",
      "rawMarkdown": "I have tried tfidf feature on this competition dataset. Kernel https://www.kaggle.com/classtag/russian-nlp-in-avito-demand-prediction",
      "votes": null
    },
    {
      "id": "320091",
      "postDate": "04/27/2018 13:01:11",
      "content": "<p>How about converting the text data to document term matrix and use less sparcity? say around 30%? We can merge that tdm with other columns. Or we can convert it into tf-idf matrix and merge next. \nI am stuck at converting Russian to English. For some reason, my R studio doesnt recognize Russian</p>",
      "rawMarkdown": "How about converting the text data to document term matrix and use less sparcity? say around 30%? We can merge that tdm with other columns. Or we can convert it into tf-idf matrix and merge next. \nI am stuck at converting Russian to English. For some reason, my R studio doesnt recognize Russian",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 319442,
      "author_name": "shivamb",
      "author_url": "",
      "post_date": "04/26/2018 03:37:38",
      "content": "<p>First vectorize your text data using TFIDF Vectorizor (or CountVectorizer etc), Then use Hstack to combine text features and numeric features. </p>\n\n<pre><code>train_data = pd.read_csv('data.csv')\n\nPreparing yoru text data\ntext_data = train_data['text_feature']\ntfidf_vectorizer = TfidfVectorizer() \ntfidf_vectorizer.fit(text_data)\nX_tfidf = vect_word.transform(df[text_data) \n\n# Preparing your numerical data \nX_numerical = train_data[[' numerical_feature1', 'numerical_feature2']]\n\nfrom scipy.sparse import hstack, csr_matrix\nX_train = hstack([X_tfidf, csr_matrix(X_numerical)], 'csr')\n</code></pre>\n\n<p>And now you can use X_train in your model. Hope it Helps </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 319443,
      "author_name": "classtag",
      "author_url": "",
      "post_date": "04/26/2018 03:40:46",
      "content": "<p>I have tried tfidf feature on this competition dataset. Kernel <a href=\"https://www.kaggle.com/classtag/russian-nlp-in-avito-demand-prediction\">https://www.kaggle.com/classtag/russian-nlp-in-avito-demand-prediction</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320091,
      "author_name": "vjkadekar",
      "author_url": "",
      "post_date": "04/27/2018 13:01:11",
      "content": "<p>How about converting the text data to document term matrix and use less sparcity? say around 30%? We can merge that tdm with other columns. Or we can convert it into tf-idf matrix and merge next. \nI am stuck at converting Russian to English. For some reason, my R studio doesnt recognize Russian</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "319422": "What are possible and best ways in combining numeric and text data into single model.",
    "319442": "First vectorize your text data using TFIDF Vectorizor (or CountVectorizer etc), Then use Hstack to combine text features and numeric features. \n\n    train_data = pd.read_csv('data.csv')\n    \n    Preparing yoru text data\n    text_data = train_data['text_feature']\n    tfidf_vectorizer = TfidfVectorizer() \n    tfidf_vectorizer.fit(text_data)\n    X_tfidf = vect_word.transform(df[text_data) \n    \n    # Preparing your numerical data \n    X_numerical = train_data[[' numerical_feature1', 'numerical_feature2']]\n    \n    from scipy.sparse import hstack, csr_matrix\n    X_train = hstack([X_tfidf, csr_matrix(X_numerical)], 'csr')\n\nAnd now you can use X_train in your model. Hope it Helps",
    "319443": "I have tried tfidf feature on this competition dataset. Kernel https://www.kaggle.com/classtag/russian-nlp-in-avito-demand-prediction",
    "320091": "How about converting the text data to document term matrix and use less sparcity? say around 30%? We can merge that tdm with other columns. Or we can convert it into tf-idf matrix and merge next. \nI am stuck at converting Russian to English. For some reason, my R studio doesnt recognize Russian"
  },
  "source": "meta"
}