{
  "id": 57809,
  "title": "Need help with Wordbatch and FM_FTRL",
  "url": "/competitions/avito-demand-prediction/discussion/57809",
  "author_name": "",
  "post_date": "2018-05-29T11:57:40.921727Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I had made a public notebook here and could use Wordbatch FM_FTRL using the text columns like description and title. Also I had feature engineered few new columns like 'num_stop-words' etc. But I'm struggling hard to merge these new features using, say hstack(), and call FM_FTRL..I always get a nan value when using these new features. Appreciate if anyone can help me on this.</p>\n\n<p>Notebook Link - <a href=\"https://www.kaggle.com/samratp/wordbatch-ridge-fm-frtl-target-encoding-lgbm\">https://www.kaggle.com/samratp/wordbatch-ridge-fm-frtl-target-encoding-lgbm</a></p>",
  "messages": [
    {
      "id": "335227",
      "postDate": "05/29/2018 11:57:40",
      "content": "<p>I had made a public notebook here and could use Wordbatch FM_FTRL using the text columns like description and title. Also I had feature engineered few new columns like 'num_stop-words' etc. But I'm struggling hard to merge these new features using, say hstack(), and call FM_FTRL..I always get a nan value when using these new features. Appreciate if anyone can help me on this.</p>\n\n<p>Notebook Link - <a href=\"https://www.kaggle.com/samratp/wordbatch-ridge-fm-frtl-target-encoding-lgbm\">https://www.kaggle.com/samratp/wordbatch-ridge-fm-frtl-target-encoding-lgbm</a></p>",
      "rawMarkdown": "I had made a public notebook here and could use Wordbatch FM_FTRL using the text columns like description and title. Also I had feature engineered few new columns like 'num_stop-words' etc. But I'm struggling hard to merge these new features using, say hstack(), and call FM_FTRL..I always get a nan value when using these new features. Appreciate if anyone can help me on this.\n\nNotebook Link - https://www.kaggle.com/samratp/wordbatch-ridge-fm-frtl-target-encoding-lgbm",
      "votes": null
    },
    {
      "id": "335228",
      "postDate": "05/29/2018 11:58:20",
      "content": "<p><a href=\"/anttip\">@anttip</a> <a href=\"/peterhurford\">@peterhurford</a> <a href=\"/tunguz\">@tunguz</a> <a href=\"/cpmpml\">@cpmpml</a> - Any help....</p>",
      "rawMarkdown": "anttip @peterhurford @tunguz @cpmpml - Any help....",
      "votes": null
    },
    {
      "id": "335244",
      "postDate": "05/29/2018 12:34:42",
      "content": "<p>I had this issue before, numerical features should be normalized, you can check this thread:</p>\n\n<p><a href=\"https://www.kaggle.com/ogrellier/wordbatch-fm-ftrl-using-mse-lb-0-9804/comments#292457\">https://www.kaggle.com/ogrellier/wordbatch-fm-ftrl-using-mse-lb-0-9804/comments#292457</a></p>",
      "rawMarkdown": "I had this issue before, numerical features should be normalized, you can check this thread:\n\nhttps://www.kaggle.com/ogrellier/wordbatch-fm-ftrl-using-mse-lb-0-9804/comments#292457",
      "votes": null
    },
    {
      "id": "335263",
      "postDate": "05/29/2018 13:12:34",
      "content": "<p>Thanks a lot <a href=\"/bangdasun\">@bangdasun</a></p>",
      "rawMarkdown": "Thanks a lot @bangdasun",
      "votes": null
    },
    {
      "id": "335342",
      "postDate": "05/29/2018 15:59:14",
      "content": "<p>@baggdasun OK.. I used MinMaxScaler() for the numerical features.. Then how do I add it to the original sparse matrix? The below code did not work for me...</p>\n\n<pre><code>scaler = MinMaxScaler()\ntrain_num_features = scaler.fit_transform(df_train[num_features])\ntest_num_features = scaler.fit_transform(df_test[num_features])\ntrain_num_features = csr_matrix(train_num_features)\ntest_num_features = csr_matrix(test_num_features)\n\nsparse_merge_train = hstack((X_description_train, X_param1_train, X_name_train, X_region_train, X_city_train, X_imagetop1_train, X_user_type_train)).tocsr()\nsparse_merge_train = hstack(sparse_merge_train, train_num_features).tocsr()\n</code></pre>",
      "rawMarkdown": "baggdasun OK.. I used MinMaxScaler() for the numerical features.. Then how do I add it to the original sparse matrix? The below code did not work for me...\n\n    scaler = MinMaxScaler()\n    train_num_features = scaler.fit_transform(df_train[num_features])\n    test_num_features = scaler.fit_transform(df_test[num_features])\n    train_num_features = csr_matrix(train_num_features)\n    test_num_features = csr_matrix(test_num_features)\n    \n    sparse_merge_train = hstack((X_description_train, X_param1_train, X_name_train, X_region_train, X_city_train, X_imagetop1_train, X_user_type_train)).tocsr()\n    sparse_merge_train = hstack(sparse_merge_train, train_num_features).tocsr()",
      "votes": null
    },
    {
      "id": "335537",
      "postDate": "05/29/2018 22:07:32",
      "content": "<p>What kind of error did you get? I see that last row should be </p>\n\n<p><code>sparse_merge_train = hstack((sparse_merge_train, train_num_features)).tocsr()</code></p>\n\n<p>and also make sure each stacked feature column is a 2D array (apply <code>.reshape(-1, 1)</code> on them).</p>",
      "rawMarkdown": "What kind of error did you get? I see that last row should be \n\n`sparse_merge_train = hstack((sparse_merge_train, train_num_features)).tocsr()`\n\nand also make sure each stacked feature column is a 2D array (apply `.reshape(-1, 1)` on them).",
      "votes": null
    },
    {
      "id": "335800",
      "postDate": "05/30/2018 12:01:18",
      "content": "<p>I left some comments on the kernel.</p>",
      "rawMarkdown": "I left some comments on the kernel.",
      "votes": null
    },
    {
      "id": "335844",
      "postDate": "05/30/2018 13:32:26",
      "content": "<p><a href=\"/bangdasun\">@bangdasun</a> Thanks for the help. Finally figured it out.</p>",
      "rawMarkdown": "bangdasun Thanks for the help. Finally figured it out.",
      "votes": null
    },
    {
      "id": "337517",
      "postDate": "06/03/2018 04:21:33",
      "content": "<p>Sorry I haven't had a chance to look into this competition in detail yet.</p>\n\n<p>Overall numerical features are better preprocessed for linear models and extensions like FMs and feedforward NNs. Binning is the basic strategy, which allows the model to deal both with input variables of different scales, and model a non-linear response for the variable. For example, I used heavily log2 binning of count values to get close to LGB results in the Talkingdata competition with a FM. And Michael Jahrer's winning autoencoder-based solution for PortoSeguro used rank-based normalization for the input features.</p>\n\n<p>I intend to expand WordBatch more in the provided transformer classes, so this type of data munging is easier to do, by chaining one-line transformers into pipelines. Not as simple as throwing the data at a GBM, but easy enough to be convenient and allow rapid prototyping.</p>",
      "rawMarkdown": "Sorry I haven't had a chance to look into this competition in detail yet.\n\nOverall numerical features are better preprocessed for linear models and extensions like FMs and feedforward NNs. Binning is the basic strategy, which allows the model to deal both with input variables of different scales, and model a non-linear response for the variable. For example, I used heavily log2 binning of count values to get close to LGB results in the Talkingdata competition with a FM. And Michael Jahrer's winning autoencoder-based solution for PortoSeguro used rank-based normalization for the input features.\n\nI intend to expand WordBatch more in the provided transformer classes, so this type of data munging is easier to do, by chaining one-line transformers into pipelines. Not as simple as throwing the data at a GBM, but easy enough to be convenient and allow rapid prototyping.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 335228,
      "author_name": "samratp",
      "author_url": "",
      "post_date": "05/29/2018 11:58:20",
      "content": "<p><a href=\"/anttip\">@anttip</a> <a href=\"/peterhurford\">@peterhurford</a> <a href=\"/tunguz\">@tunguz</a> <a href=\"/cpmpml\">@cpmpml</a> - Any help....</p>",
      "votes": null,
      "replies": [
        {
          "id": 335800,
          "author_name": "peterhurford",
          "author_url": "",
          "post_date": "05/30/2018 12:01:18",
          "content": "<p>I left some comments on the kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 337517,
          "author_name": "anttip",
          "author_url": "",
          "post_date": "06/03/2018 04:21:33",
          "content": "<p>Sorry I haven't had a chance to look into this competition in detail yet.</p>\n\n<p>Overall numerical features are better preprocessed for linear models and extensions like FMs and feedforward NNs. Binning is the basic strategy, which allows the model to deal both with input variables of different scales, and model a non-linear response for the variable. For example, I used heavily log2 binning of count values to get close to LGB results in the Talkingdata competition with a FM. And Michael Jahrer's winning autoencoder-based solution for PortoSeguro used rank-based normalization for the input features.</p>\n\n<p>I intend to expand WordBatch more in the provided transformer classes, so this type of data munging is easier to do, by chaining one-line transformers into pipelines. Not as simple as throwing the data at a GBM, but easy enough to be convenient and allow rapid prototyping.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 335244,
      "author_name": "bangdasun",
      "author_url": "",
      "post_date": "05/29/2018 12:34:42",
      "content": "<p>I had this issue before, numerical features should be normalized, you can check this thread:</p>\n\n<p><a href=\"https://www.kaggle.com/ogrellier/wordbatch-fm-ftrl-using-mse-lb-0-9804/comments#292457\">https://www.kaggle.com/ogrellier/wordbatch-fm-ftrl-using-mse-lb-0-9804/comments#292457</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 335263,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "05/29/2018 13:12:34",
          "content": "<p>Thanks a lot <a href=\"/bangdasun\">@bangdasun</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335342,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "05/29/2018 15:59:14",
          "content": "<p>@baggdasun OK.. I used MinMaxScaler() for the numerical features.. Then how do I add it to the original sparse matrix? The below code did not work for me...</p>\n\n<pre><code>scaler = MinMaxScaler()\ntrain_num_features = scaler.fit_transform(df_train[num_features])\ntest_num_features = scaler.fit_transform(df_test[num_features])\ntrain_num_features = csr_matrix(train_num_features)\ntest_num_features = csr_matrix(test_num_features)\n\nsparse_merge_train = hstack((X_description_train, X_param1_train, X_name_train, X_region_train, X_city_train, X_imagetop1_train, X_user_type_train)).tocsr()\nsparse_merge_train = hstack(sparse_merge_train, train_num_features).tocsr()\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335537,
          "author_name": "bangdasun",
          "author_url": "",
          "post_date": "05/29/2018 22:07:32",
          "content": "<p>What kind of error did you get? I see that last row should be </p>\n\n<p><code>sparse_merge_train = hstack((sparse_merge_train, train_num_features)).tocsr()</code></p>\n\n<p>and also make sure each stacked feature column is a 2D array (apply <code>.reshape(-1, 1)</code> on them).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335844,
          "author_name": "samratp",
          "author_url": "",
          "post_date": "05/30/2018 13:32:26",
          "content": "<p><a href=\"/bangdasun\">@bangdasun</a> Thanks for the help. Finally figured it out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "335227": "I had made a public notebook here and could use Wordbatch FM_FTRL using the text columns like description and title. Also I had feature engineered few new columns like 'num_stop-words' etc. But I'm struggling hard to merge these new features using, say hstack(), and call FM_FTRL..I always get a nan value when using these new features. Appreciate if anyone can help me on this.\n\nNotebook Link - https://www.kaggle.com/samratp/wordbatch-ridge-fm-frtl-target-encoding-lgbm",
    "335228": "anttip @peterhurford @tunguz @cpmpml - Any help....",
    "335244": "I had this issue before, numerical features should be normalized, you can check this thread:\n\nhttps://www.kaggle.com/ogrellier/wordbatch-fm-ftrl-using-mse-lb-0-9804/comments#292457",
    "335263": "Thanks a lot @bangdasun",
    "335342": "baggdasun OK.. I used MinMaxScaler() for the numerical features.. Then how do I add it to the original sparse matrix? The below code did not work for me...\n\n    scaler = MinMaxScaler()\n    train_num_features = scaler.fit_transform(df_train[num_features])\n    test_num_features = scaler.fit_transform(df_test[num_features])\n    train_num_features = csr_matrix(train_num_features)\n    test_num_features = csr_matrix(test_num_features)\n    \n    sparse_merge_train = hstack((X_description_train, X_param1_train, X_name_train, X_region_train, X_city_train, X_imagetop1_train, X_user_type_train)).tocsr()\n    sparse_merge_train = hstack(sparse_merge_train, train_num_features).tocsr()",
    "335537": "What kind of error did you get? I see that last row should be \n\n`sparse_merge_train = hstack((sparse_merge_train, train_num_features)).tocsr()`\n\nand also make sure each stacked feature column is a 2D array (apply `.reshape(-1, 1)` on them).",
    "335800": "I left some comments on the kernel.",
    "335844": "bangdasun Thanks for the help. Finally figured it out.",
    "337517": "Sorry I haven't had a chance to look into this competition in detail yet.\n\nOverall numerical features are better preprocessed for linear models and extensions like FMs and feedforward NNs. Binning is the basic strategy, which allows the model to deal both with input variables of different scales, and model a non-linear response for the variable. For example, I used heavily log2 binning of count values to get close to LGB results in the Talkingdata competition with a FM. And Michael Jahrer's winning autoencoder-based solution for PortoSeguro used rank-based normalization for the input features.\n\nI intend to expand WordBatch more in the provided transformer classes, so this type of data munging is easier to do, by chaining one-line transformers into pipelines. Not as simple as throwing the data at a GBM, but easy enough to be convenient and allow rapid prototyping."
  },
  "source": "meta"
}