{
  "id": 21439,
  "title": "XGboost for multi class problem?",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21439",
  "author_name": "",
  "post_date": "2016-06-05T06:42:39.593Z",
  "votes": null,
  "comment_count": 4,
  "views": 1790,
  "content": "<p>Hi, does xgboost  (python), knows how to deal directly with 100 classification?</p>\n\n<p>I have used the follwoing:</p>\n\n<p>alg = xgb.XGBClassifier(learning_rate=0.1,\n                        n_estimators=28,\n                        max_depth=7,\n                        min_child_weight=1,\n                        gamma=0,\n                        subsample=0.9,\n                        colsample_bytree=0.9,\n                        objective='binary:logistic',\n                        nthread=4,\n                        scale_pos_weight=1,\n                        seed=27,\n                        reg_alpha=0.01)</p>\n\n<p>alg.fit(train[predictos], traint[&quot;hotel_cluster&quot;])</p>\n\n<p>traint[&quot;hotel_cluster&quot;] - has values in the range [1,100].</p>\n\n<p>Does xgboost knows how to deal directly with 100 classifation problem?\nOr must i divide the problem to 100 1-vs all?</p>",
  "messages": [
    {
      "id": "122566",
      "postDate": "06/05/2016 06:42:39",
      "content": "<p>Hi, does xgboost  (python), knows how to deal directly with 100 classification?</p>\n\n<p>I have used the follwoing:</p>\n\n<p>alg = xgb.XGBClassifier(learning_rate=0.1,\n                        n_estimators=28,\n                        max_depth=7,\n                        min_child_weight=1,\n                        gamma=0,\n                        subsample=0.9,\n                        colsample_bytree=0.9,\n                        objective='binary:logistic',\n                        nthread=4,\n                        scale_pos_weight=1,\n                        seed=27,\n                        reg_alpha=0.01)</p>\n\n<p>alg.fit(train[predictos], traint[&quot;hotel_cluster&quot;])</p>\n\n<p>traint[&quot;hotel_cluster&quot;] - has values in the range [1,100].</p>\n\n<p>Does xgboost knows how to deal directly with 100 classifation problem?\nOr must i divide the problem to 100 1-vs all?</p>",
      "rawMarkdown": "Hi, does xgboost  (python), knows how to deal directly with 100 classification?\r\n\r\nI have used the follwoing:\r\n\r\nalg = xgb.XGBClassifier(learning_rate=0.1,\r\n                        n_estimators=28,\r\n                        max_depth=7,\r\n                        min_child_weight=1,\r\n                        gamma=0,\r\n                        subsample=0.9,\r\n                        colsample_bytree=0.9,\r\n                        objective='binary:logistic',\r\n                        nthread=4,\r\n                        scale_pos_weight=1,\r\n                        seed=27,\r\n                        reg_alpha=0.01)\r\n \r\n\r\nalg.fit(train[predictos], traint[\"hotel_cluster\"])\r\n\r\ntraint[\"hotel_cluster\"] - has values in the range [1,100].\r\n\r\nDoes xgboost knows how to deal directly with 100 classifation problem?\r\nOr must i divide the problem to 100 1-vs all?",
      "votes": null
    },
    {
      "id": "122570",
      "postDate": "06/05/2016 08:20:30",
      "content": "<p>xgboost can be used for multi class.  Here is my code, adapted from a xgboost examples, where map5eval is ghe function I gave previously in this topic.  </p>\n\n<pre><code>watchlist = [ (Xtrain,'train'), (Xtest, 'test') ]\nparam['objective'] = 'multi:softprob'\nbst = xgb.train(param, Xtrain, num_round, watchlist, feval=map5eval, early_stopping_rounds=50, maximize=True);\nyprob = bst.predict( Xtest ).reshape( ytest.shape[0], 100 )\nylabel = np.argmax(yprob, axis=1)\n\nprint ('predicting, classification error=%f' % \\\n       (np.sum( ylabel != ytest) / float(ytest.shape[0]) ))\n</code></pre>",
      "rawMarkdown": "xgboost can be used for multi class.  Here is my code, adapted from a xgboost examples, where map5eval is ghe function I gave previously in this topic.  \r\n\r\n    watchlist = [ (Xtrain,'train'), (Xtest, 'test') ]\r\n    param['objective'] = 'multi:softprob'\r\n    bst = xgb.train(param, Xtrain, num_round, watchlist, feval=map5eval, early_stopping_rounds=50, maximize=True);\r\n    yprob = bst.predict( Xtest ).reshape( ytest.shape[0], 100 )\r\n    ylabel = np.argmax(yprob, axis=1)\r\n    \r\n    print ('predicting, classification error=%f' % \\\r\n           (np.sum( ylabel != ytest) / float(ytest.shape[0]) ))",
      "votes": null
    },
    {
      "id": "122572",
      "postDate": "06/05/2016 08:58:08",
      "content": "<p>what exactly the function that the xgboost is optimizing?\nWhat the meaning of the objective and feval?\nWhat is the difference between the'm?</p>\n\n<p>I used </p>\n\n<p>param_test7 = {\n 'n_estimators':[20]\n}</p>\n\n<p>clf = GridSearchCV(xgb.XGBClassifier(learning_rate =0.1,\n                                     n_estimators=250,\n                                     max_depth=9,\n                                     min_child_weight=1,\n                                     gamma=0.1,\n                                     subsample=0.85,\n                                     colsample_bytree=0.75,\n                                     objective= 'multi:softprob',\n                                     nthread=4,\n                                     scale_pos_weight=1,\n                                     seed=27,\n                                     reg_alpha=0.01),\n                   param_test7, \n                   verbose=1, \n                   cv=5, \n                   scoring='log_loss',\n                   n_jobs=4,\n                   iid=False)</p>\n\n<p>clf.fit(train[predictors],train[&quot;hotel_cluster&quot;]</p>\n\n<p>and i get the following:</p>\n\n<p>ValueError: y_true and y_pred have different number of classes 59, 53</p>\n\n<p>any one knows why?\nI guess the train predicts less classes than those in the test, which can happen.</p>\n\n<p>How did you resolve this?</p>",
      "rawMarkdown": "what exactly the function that the xgboost is optimizing?\r\nWhat the meaning of the objective and feval?\r\nWhat is the difference between the'm?\r\n\r\nI used \r\n\r\nparam_test7 = {\r\n 'n_estimators':[20]\r\n}\r\n\r\nclf = GridSearchCV(xgb.XGBClassifier(learning_rate =0.1,\r\n                                     n_estimators=250,\r\n                                     max_depth=9,\r\n                                     min_child_weight=1,\r\n                                     gamma=0.1,\r\n                                     subsample=0.85,\r\n                                     colsample_bytree=0.75,\r\n                                     objective= 'multi:softprob',\r\n                                     nthread=4,\r\n                                     scale_pos_weight=1,\r\n                                     seed=27,\r\n                                     reg_alpha=0.01),\r\n                   param_test7, \r\n                   verbose=1, \r\n                   cv=5, \r\n                   scoring='log_loss',\r\n                   n_jobs=4,\r\n                   iid=False)\r\n\r\n\r\nclf.fit(train[predictors],train[\"hotel_cluster\"]\r\n\r\nand i get the following:\r\n\r\nValueError: y_true and y_pred have different number of classes 59, 53\r\n\r\nany one knows why?\r\nI guess the train predicts less classes than those in the test, which can happen.\r\n\r\nHow did you resolve this?",
      "votes": null
    },
    {
      "id": "122661",
      "postDate": "06/06/2016 07:38:31",
      "content": "<p>XGBoost requires a differentiable objective function (i.e. a function that has a well defined gradient).  MAP@5 is not differentiable, hence it cannot be used directly as the objective. <br>\nWe therefore use a different objective, here we use softmax.  There are two flavors of softmax in xgboost, one that outputs a class, and another one that outputs probabilities.  i selected the latter to be able to blend results with other methods.</p>\n\n<p>Given we don't use MAP@5 as the objective, and given this is what we care about, we use MAP@5 as the output metric.  It is computed on what we pass as watch list, i.e. our training data and our test data.</p>\n\n<p>To avoid your class number issue, you must pass the number of classes as a parameter:</p>\n\n<pre><code>param['num_class'] = 100 \n</code></pre>",
      "rawMarkdown": "XGBoost requires a differentiable objective function (i.e. a function that has a well defined gradient).  MAP@5 is not differentiable, hence it cannot be used directly as the objective.  \r\nWe therefore use a different objective, here we use softmax.  There are two flavors of softmax in xgboost, one that outputs a class, and another one that outputs probabilities.  i selected the latter to be able to blend results with other methods.\r\n\r\nGiven we don't use MAP@5 as the objective, and given this is what we care about, we use MAP@5 as the output metric.  It is computed on what we pass as watch list, i.e. our training data and our test data.\r\n\r\nTo avoid your class number issue, you must pass the number of classes as a parameter:\r\n\r\n    param['num_class'] = 100",
      "votes": null
    },
    {
      "id": "667394",
      "postDate": "11/07/2019 07:24:21",
      "content": "<p>where can I get the <strong>map5eval</strong> ??</p>",
      "rawMarkdown": "where can I get the **map5eval** ??",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122570,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/05/2016 08:20:30",
      "content": "<p>xgboost can be used for multi class.  Here is my code, adapted from a xgboost examples, where map5eval is ghe function I gave previously in this topic.  </p>\n\n<pre><code>watchlist = [ (Xtrain,'train'), (Xtest, 'test') ]\nparam['objective'] = 'multi:softprob'\nbst = xgb.train(param, Xtrain, num_round, watchlist, feval=map5eval, early_stopping_rounds=50, maximize=True);\nyprob = bst.predict( Xtest ).reshape( ytest.shape[0], 100 )\nylabel = np.argmax(yprob, axis=1)\n\nprint ('predicting, classification error=%f' % \\\n       (np.sum( ylabel != ytest) / float(ytest.shape[0]) ))\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 667394,
          "author_name": "jeetranjeet619",
          "author_url": "",
          "post_date": "11/07/2019 07:24:21",
          "content": "<p>where can I get the <strong>map5eval</strong> ??</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 122572,
      "author_name": "dot277",
      "author_url": "",
      "post_date": "06/05/2016 08:58:08",
      "content": "<p>what exactly the function that the xgboost is optimizing?\nWhat the meaning of the objective and feval?\nWhat is the difference between the'm?</p>\n\n<p>I used </p>\n\n<p>param_test7 = {\n 'n_estimators':[20]\n}</p>\n\n<p>clf = GridSearchCV(xgb.XGBClassifier(learning_rate =0.1,\n                                     n_estimators=250,\n                                     max_depth=9,\n                                     min_child_weight=1,\n                                     gamma=0.1,\n                                     subsample=0.85,\n                                     colsample_bytree=0.75,\n                                     objective= 'multi:softprob',\n                                     nthread=4,\n                                     scale_pos_weight=1,\n                                     seed=27,\n                                     reg_alpha=0.01),\n                   param_test7, \n                   verbose=1, \n                   cv=5, \n                   scoring='log_loss',\n                   n_jobs=4,\n                   iid=False)</p>\n\n<p>clf.fit(train[predictors],train[&quot;hotel_cluster&quot;]</p>\n\n<p>and i get the following:</p>\n\n<p>ValueError: y_true and y_pred have different number of classes 59, 53</p>\n\n<p>any one knows why?\nI guess the train predicts less classes than those in the test, which can happen.</p>\n\n<p>How did you resolve this?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122661,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/06/2016 07:38:31",
      "content": "<p>XGBoost requires a differentiable objective function (i.e. a function that has a well defined gradient).  MAP@5 is not differentiable, hence it cannot be used directly as the objective. <br>\nWe therefore use a different objective, here we use softmax.  There are two flavors of softmax in xgboost, one that outputs a class, and another one that outputs probabilities.  i selected the latter to be able to blend results with other methods.</p>\n\n<p>Given we don't use MAP@5 as the objective, and given this is what we care about, we use MAP@5 as the output metric.  It is computed on what we pass as watch list, i.e. our training data and our test data.</p>\n\n<p>To avoid your class number issue, you must pass the number of classes as a parameter:</p>\n\n<pre><code>param['num_class'] = 100 \n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "122566": "Hi, does xgboost  (python), knows how to deal directly with 100 classification?\r\n\r\nI have used the follwoing:\r\n\r\nalg = xgb.XGBClassifier(learning_rate=0.1,\r\n                        n_estimators=28,\r\n                        max_depth=7,\r\n                        min_child_weight=1,\r\n                        gamma=0,\r\n                        subsample=0.9,\r\n                        colsample_bytree=0.9,\r\n                        objective='binary:logistic',\r\n                        nthread=4,\r\n                        scale_pos_weight=1,\r\n                        seed=27,\r\n                        reg_alpha=0.01)\r\n \r\n\r\nalg.fit(train[predictos], traint[\"hotel_cluster\"])\r\n\r\ntraint[\"hotel_cluster\"] - has values in the range [1,100].\r\n\r\nDoes xgboost knows how to deal directly with 100 classifation problem?\r\nOr must i divide the problem to 100 1-vs all?",
    "122570": "xgboost can be used for multi class.  Here is my code, adapted from a xgboost examples, where map5eval is ghe function I gave previously in this topic.  \r\n\r\n    watchlist = [ (Xtrain,'train'), (Xtest, 'test') ]\r\n    param['objective'] = 'multi:softprob'\r\n    bst = xgb.train(param, Xtrain, num_round, watchlist, feval=map5eval, early_stopping_rounds=50, maximize=True);\r\n    yprob = bst.predict( Xtest ).reshape( ytest.shape[0], 100 )\r\n    ylabel = np.argmax(yprob, axis=1)\r\n    \r\n    print ('predicting, classification error=%f' % \\\r\n           (np.sum( ylabel != ytest) / float(ytest.shape[0]) ))",
    "122572": "what exactly the function that the xgboost is optimizing?\r\nWhat the meaning of the objective and feval?\r\nWhat is the difference between the'm?\r\n\r\nI used \r\n\r\nparam_test7 = {\r\n 'n_estimators':[20]\r\n}\r\n\r\nclf = GridSearchCV(xgb.XGBClassifier(learning_rate =0.1,\r\n                                     n_estimators=250,\r\n                                     max_depth=9,\r\n                                     min_child_weight=1,\r\n                                     gamma=0.1,\r\n                                     subsample=0.85,\r\n                                     colsample_bytree=0.75,\r\n                                     objective= 'multi:softprob',\r\n                                     nthread=4,\r\n                                     scale_pos_weight=1,\r\n                                     seed=27,\r\n                                     reg_alpha=0.01),\r\n                   param_test7, \r\n                   verbose=1, \r\n                   cv=5, \r\n                   scoring='log_loss',\r\n                   n_jobs=4,\r\n                   iid=False)\r\n\r\n\r\nclf.fit(train[predictors],train[\"hotel_cluster\"]\r\n\r\nand i get the following:\r\n\r\nValueError: y_true and y_pred have different number of classes 59, 53\r\n\r\nany one knows why?\r\nI guess the train predicts less classes than those in the test, which can happen.\r\n\r\nHow did you resolve this?",
    "122661": "XGBoost requires a differentiable objective function (i.e. a function that has a well defined gradient).  MAP@5 is not differentiable, hence it cannot be used directly as the objective.  \r\nWe therefore use a different objective, here we use softmax.  There are two flavors of softmax in xgboost, one that outputs a class, and another one that outputs probabilities.  i selected the latter to be able to blend results with other methods.\r\n\r\nGiven we don't use MAP@5 as the objective, and given this is what we care about, we use MAP@5 as the output metric.  It is computed on what we pass as watch list, i.e. our training data and our test data.\r\n\r\nTo avoid your class number issue, you must pass the number of classes as a parameter:\r\n\r\n    param['num_class'] = 100",
    "667394": "where can I get the **map5eval** ??"
  },
  "source": "meta"
}