{
  "id": 55931,
  "title": "How to use text data and numerical data together as input features?",
  "url": "/competitions/avito-demand-prediction/discussion/55931",
  "author_name": "",
  "post_date": "2018-05-03T09:00:15.797612600Z",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>After vectorization of text data, how can i combine the vector output along with numerical data</p>",
  "messages": [
    {
      "id": "322597",
      "postDate": "05/03/2018 09:00:15",
      "content": "<p>After vectorization of text data, how can i combine the vector output along with numerical data</p>",
      "rawMarkdown": "After vectorization of text data, how can i combine the vector output along with numerical data",
      "votes": null
    },
    {
      "id": "322627",
      "postDate": "05/03/2018 10:21:36",
      "content": "<p>For example like this:</p>\n\n<pre><code>X = np.hstack([X_vectorized, train])\n</code></pre>",
      "rawMarkdown": "For example like this:\n\n    X = np.hstack([X_vectorized, train])",
      "votes": null
    },
    {
      "id": "322652",
      "postDate": "05/03/2018 11:33:01",
      "content": "<pre><code>`tf = HashingVectorizer(analyzer='word',stop_words=stopwords.words('russian'), encoding='KOI8R',n_features=1000)\ntxt_transformed = tf.fit_transform(train_df['description'])\ntxt_transformed=np.asarray(txt_transformed)\nX_train = np.hstack([txt_transformed, train_X])`\n</code></pre>\n\n<p>This is returning</p>\n\n<pre><code>---------------------------------------------------------------------------\n</code></pre>\n\n<p>ValueError                                Traceback (most recent call last)\n in ()\n----&gt; 1 X_train = np.hstack([txt_transformed, train_X])</p>\n\n<p>/opt/conda/lib/python3.6/site-packages/numpy/core/shape_base.py in hstack(tup)\n    284     # As a special case, dimension 0 of 1-dimensional arrays is \"horizontal\"\n    285     if arrs and arrs[0].ndim == 1:\n--&gt; 286         return _nx.concatenate(arrs, 0)\n    287     else:\n    288         return _nx.concatenate(arrs, 1)</p>\n\n<p>ValueError: all the input arrays must have same number of dimensions</p>",
      "rawMarkdown": "`tf = HashingVectorizer(analyzer='word',stop_words=stopwords.words('russian'), encoding='KOI8R',n_features=1000)\n    txt_transformed = tf.fit_transform(train_df['description'])\n    txt_transformed=np.asarray(txt_transformed)\n    X_train = np.hstack([txt_transformed, train_X])`\n\nThis is returning\n\n    ---------------------------------------------------------------------------\nValueError                                Traceback (most recent call last)",
      "votes": null
    },
    {
      "id": "322659",
      "postDate": "05/03/2018 11:42:46",
      "content": "<p>Check dimensions of txt_transformed and train_X.</p>",
      "rawMarkdown": "Check dimensions of txt_transformed and train_X.",
      "votes": null
    },
    {
      "id": "322661",
      "postDate": "05/03/2018 11:45:14",
      "content": "<pre><code>train_X.shape\n</code></pre>\n\n<p>(1503424, 12)</p>\n\n<pre><code>txt_transformed\n</code></pre>\n\n<p>&lt;1503424x1000 sparse matrix of type ''\n    with 26251135 stored elements in Compressed Sparse Row format&gt;</p>\n\n<p>The number of rows are the same. Initially, I thought the issue was with the difference in number of rows. But that seems to be fine.</p>",
      "rawMarkdown": "train_X.shape\n(1503424, 12)\n\n    txt_transformed\n&lt;1503424x1000 sparse matrix of type '",
      "votes": null
    },
    {
      "id": "322685",
      "postDate": "05/03/2018 12:35:29",
      "content": "<p>This is quite strange... could you show an example of data in \"train_X\"?</p>",
      "rawMarkdown": "This is quite strange... could you show an example of data in \"train_X\"?",
      "votes": null
    },
    {
      "id": "322992",
      "postDate": "05/04/2018 05:07:07",
      "content": "<p><a href=\"https://drive.google.com/file/d/1mSk2v3-Z2jYW80G8btjQims5zF4PYGwk/view?usp=drivesdk\">screenshot of train_X</a></p>",
      "rawMarkdown": "[screenshot of train_X][1]\n\n\n  [1]: https://drive.google.com/file/d/1mSk2v3-Z2jYW80G8btjQims5zF4PYGwk/view?usp=drivesdk",
      "votes": null
    },
    {
      "id": "323041",
      "postDate": "05/04/2018 07:50:28",
      "content": "<p>I have found the problem, in fact there are two of them:</p>\n\n<ol>\n<li><p>You don't need to do this step:</p>\n\n<p>txt_transformed=np.asarray(txt_transformed)</p></li>\n</ol>\n\n<p>Data is already in sparse format, converting it to array isn't a good idea.</p>\n\n<ol>\n<li><p>If one of matrices is sparse, you need to to stacking in a different way:</p>\n\n<p>from scipy.sparse import csr_matrix, vstack, hstack\nX_train  = csr_matrix(hstack([np.asarray(txt_transformed), train_X]))</p></li>\n</ol>",
      "rawMarkdown": "I have found the problem, in fact there are two of them:\n\n1. You don't need to do this step:\n\n    txt_transformed=np.asarray(txt_transformed)\n\nData is already in sparse format, converting it to array isn't a good idea.\n\n2. If one of matrices is sparse, you need to to stacking in a different way:\n\n    from scipy.sparse import csr_matrix, vstack, hstack\n    X_train  = csr_matrix(hstack([np.asarray(txt_transformed), train_X]))",
      "votes": null
    },
    {
      "id": "323060",
      "postDate": "05/04/2018 09:00:50",
      "content": "<p>Lemme try this</p>",
      "rawMarkdown": "Lemme try this",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 322627,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "05/03/2018 10:21:36",
      "content": "<p>For example like this:</p>\n\n<pre><code>X = np.hstack([X_vectorized, train])\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 322652,
          "author_name": "rohitanil",
          "author_url": "",
          "post_date": "05/03/2018 11:33:01",
          "content": "<pre><code>`tf = HashingVectorizer(analyzer='word',stop_words=stopwords.words('russian'), encoding='KOI8R',n_features=1000)\ntxt_transformed = tf.fit_transform(train_df['description'])\ntxt_transformed=np.asarray(txt_transformed)\nX_train = np.hstack([txt_transformed, train_X])`\n</code></pre>\n\n<p>This is returning</p>\n\n<pre><code>---------------------------------------------------------------------------\n</code></pre>\n\n<p>ValueError                                Traceback (most recent call last)\n in ()\n----&gt; 1 X_train = np.hstack([txt_transformed, train_X])</p>\n\n<p>/opt/conda/lib/python3.6/site-packages/numpy/core/shape_base.py in hstack(tup)\n    284     # As a special case, dimension 0 of 1-dimensional arrays is \"horizontal\"\n    285     if arrs and arrs[0].ndim == 1:\n--&gt; 286         return _nx.concatenate(arrs, 0)\n    287     else:\n    288         return _nx.concatenate(arrs, 1)</p>\n\n<p>ValueError: all the input arrays must have same number of dimensions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322659,
          "author_name": "artgor",
          "author_url": "",
          "post_date": "05/03/2018 11:42:46",
          "content": "<p>Check dimensions of txt_transformed and train_X.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322661,
          "author_name": "rohitanil",
          "author_url": "",
          "post_date": "05/03/2018 11:45:14",
          "content": "<pre><code>train_X.shape\n</code></pre>\n\n<p>(1503424, 12)</p>\n\n<pre><code>txt_transformed\n</code></pre>\n\n<p>&lt;1503424x1000 sparse matrix of type ''\n    with 26251135 stored elements in Compressed Sparse Row format&gt;</p>\n\n<p>The number of rows are the same. Initially, I thought the issue was with the difference in number of rows. But that seems to be fine.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322685,
          "author_name": "artgor",
          "author_url": "",
          "post_date": "05/03/2018 12:35:29",
          "content": "<p>This is quite strange... could you show an example of data in \"train_X\"?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322992,
          "author_name": "rohitanil",
          "author_url": "",
          "post_date": "05/04/2018 05:07:07",
          "content": "<p><a href=\"https://drive.google.com/file/d/1mSk2v3-Z2jYW80G8btjQims5zF4PYGwk/view?usp=drivesdk\">screenshot of train_X</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323041,
          "author_name": "artgor",
          "author_url": "",
          "post_date": "05/04/2018 07:50:28",
          "content": "<p>I have found the problem, in fact there are two of them:</p>\n\n<ol>\n<li><p>You don't need to do this step:</p>\n\n<p>txt_transformed=np.asarray(txt_transformed)</p></li>\n</ol>\n\n<p>Data is already in sparse format, converting it to array isn't a good idea.</p>\n\n<ol>\n<li><p>If one of matrices is sparse, you need to to stacking in a different way:</p>\n\n<p>from scipy.sparse import csr_matrix, vstack, hstack\nX_train  = csr_matrix(hstack([np.asarray(txt_transformed), train_X]))</p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323060,
          "author_name": "rohitanil",
          "author_url": "",
          "post_date": "05/04/2018 09:00:50",
          "content": "<p>Lemme try this</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "322597": "After vectorization of text data, how can i combine the vector output along with numerical data",
    "322627": "For example like this:\n\n    X = np.hstack([X_vectorized, train])",
    "322652": "`tf = HashingVectorizer(analyzer='word',stop_words=stopwords.words('russian'), encoding='KOI8R',n_features=1000)\n    txt_transformed = tf.fit_transform(train_df['description'])\n    txt_transformed=np.asarray(txt_transformed)\n    X_train = np.hstack([txt_transformed, train_X])`\n\nThis is returning\n\n    ---------------------------------------------------------------------------\nValueError                                Traceback (most recent call last)",
    "322659": "Check dimensions of txt_transformed and train_X.",
    "322661": "train_X.shape\n(1503424, 12)\n\n    txt_transformed\n&lt;1503424x1000 sparse matrix of type '",
    "322685": "This is quite strange... could you show an example of data in \"train_X\"?",
    "322992": "[screenshot of train_X][1]\n\n\n  [1]: https://drive.google.com/file/d/1mSk2v3-Z2jYW80G8btjQims5zF4PYGwk/view?usp=drivesdk",
    "323041": "I have found the problem, in fact there are two of them:\n\n1. You don't need to do this step:\n\n    txt_transformed=np.asarray(txt_transformed)\n\nData is already in sparse format, converting it to array isn't a good idea.\n\n2. If one of matrices is sparse, you need to to stacking in a different way:\n\n    from scipy.sparse import csr_matrix, vstack, hstack\n    X_train  = csr_matrix(hstack([np.asarray(txt_transformed), train_X]))",
    "323060": "Lemme try this"
  },
  "source": "meta"
}