{
  "id": 55164,
  "title": "how to use stacking？",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55164",
  "author_name": "",
  "post_date": "2018-04-23T01:50:57.780984600Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I use lgb as first model，and xgb as second model, and <strong>mean</strong> the first model results,<strong>but my outputs (second layer xgboost )are all ZEROs</strong>,this is badly hurt me....</p>\n\n<p>I use code get Out-of-Fold Predictions below.</p>\n\n<pre><code> def get_oof(clf,x_train,y_train,x_test,ntrain,ntest,kf):\n        oof_train = np.zeros((ntrain,))\n        oof_test = np.zeros((ntest,))\n        oof_test_skf = np.empty((NFOLDS, ntest))\n\n        for i, (train_index, test_index) in enumerate(kf.split(x_train)):\n            x_tr = x_train.iloc[train_index]\n            y_tr = y_train.iloc[train_index]\n            x_te = x_train.iloc[test_index]\n\n            clf.train(x_tr, y_tr)\n\n            oof_train[test_index] = clf.predict(x_te)\n            print(oof_train[test_index][:20])\n            oof_test_skf[i, :] = clf.predict(x_test)\n\n        oof_test[:] = oof_test_skf.mean(axis=0)\n        return oof_train.reshape(-1, 1), oof_test.reshape(-1, 1)\n</code></pre>\n\n<p>my lgb parameters are below:</p>\n\n<pre><code>lgb_params = {\n    'boosting_type': 'gbdt',\n    'objective': objective,\n    'metric':metrics,\n    'learning_rate': 0.2,\n    #'is_unbalance': 'true',  #because training data is unbalance (replaced with scale_pos_weight)\n    'num_leaves': 31,  # we should let it be smaller than 2^(max_depth)\n    'max_depth': -1,  # -1 means no limit\n    'min_child_samples': 20,  # Minimum number of data need in a child(min_data_in_leaf)\n    'max_bin': 255,  # Number of bucketed bin for feature values\n    'subsample': 0.6,  # Subsample ratio of the training instance.\n    'subsample_freq': 0,  # frequence of subsample, &amp;lt;=0 means no enable\n    'colsample_bytree': 0.3,  # Subsample ratio of columns when constructing each tree.\n    'min_child_weight': 5,  # Minimum sum of instance weight(hessian) needed in a child(leaf)\n    'subsample_for_bin': 200000,  # Number of samples for constructing bin\n    'min_split_gain': 0,  # lambda_l1, lambda_l2 and min_gain_to_split to regularization\n    'reg_alpha': 0,  # L1 regularization term on weights\n    'reg_lambda': 0,  # L2 regularization term on weights\n    'nthread': 4,\n    'verbose': 0,\n    'metric':metrics\n}\n</code></pre>\n\n<p>my xgboost parameters are below:</p>\n\n<pre><code>params = {'eta': 0.3,\n      'tree_method': \"hist\",\n      'grow_policy': \"lossguide\",\n      'max_leaves': 1400,  \n      'max_depth': 0, \n      'subsample': 0.9, \n      'colsample_bytree': 0.7, \n      'colsample_bylevel':0.7,\n      'min_child_weight':0,\n      'alpha':4,\n      'objective': 'binary:logitraw', \n      'scale_pos_weight':9,\n      'eval_metric': 'auc', \n      'nthread':8,\n      'random_state': 99, \n      'silent': False}\n</code></pre>\n\n<p><strong>can anybody give me some suggestion?\nthanks</strong></p>",
  "messages": [
    {
      "id": "317982",
      "postDate": "04/23/2018 01:50:57",
      "content": "<p>I use lgb as first model，and xgb as second model, and <strong>mean</strong> the first model results,<strong>but my outputs (second layer xgboost )are all ZEROs</strong>,this is badly hurt me....</p>\n\n<p>I use code get Out-of-Fold Predictions below.</p>\n\n<pre><code> def get_oof(clf,x_train,y_train,x_test,ntrain,ntest,kf):\n        oof_train = np.zeros((ntrain,))\n        oof_test = np.zeros((ntest,))\n        oof_test_skf = np.empty((NFOLDS, ntest))\n\n        for i, (train_index, test_index) in enumerate(kf.split(x_train)):\n            x_tr = x_train.iloc[train_index]\n            y_tr = y_train.iloc[train_index]\n            x_te = x_train.iloc[test_index]\n\n            clf.train(x_tr, y_tr)\n\n            oof_train[test_index] = clf.predict(x_te)\n            print(oof_train[test_index][:20])\n            oof_test_skf[i, :] = clf.predict(x_test)\n\n        oof_test[:] = oof_test_skf.mean(axis=0)\n        return oof_train.reshape(-1, 1), oof_test.reshape(-1, 1)\n</code></pre>\n\n<p>my lgb parameters are below:</p>\n\n<pre><code>lgb_params = {\n    'boosting_type': 'gbdt',\n    'objective': objective,\n    'metric':metrics,\n    'learning_rate': 0.2,\n    #'is_unbalance': 'true',  #because training data is unbalance (replaced with scale_pos_weight)\n    'num_leaves': 31,  # we should let it be smaller than 2^(max_depth)\n    'max_depth': -1,  # -1 means no limit\n    'min_child_samples': 20,  # Minimum number of data need in a child(min_data_in_leaf)\n    'max_bin': 255,  # Number of bucketed bin for feature values\n    'subsample': 0.6,  # Subsample ratio of the training instance.\n    'subsample_freq': 0,  # frequence of subsample, &amp;lt;=0 means no enable\n    'colsample_bytree': 0.3,  # Subsample ratio of columns when constructing each tree.\n    'min_child_weight': 5,  # Minimum sum of instance weight(hessian) needed in a child(leaf)\n    'subsample_for_bin': 200000,  # Number of samples for constructing bin\n    'min_split_gain': 0,  # lambda_l1, lambda_l2 and min_gain_to_split to regularization\n    'reg_alpha': 0,  # L1 regularization term on weights\n    'reg_lambda': 0,  # L2 regularization term on weights\n    'nthread': 4,\n    'verbose': 0,\n    'metric':metrics\n}\n</code></pre>\n\n<p>my xgboost parameters are below:</p>\n\n<pre><code>params = {'eta': 0.3,\n      'tree_method': \"hist\",\n      'grow_policy': \"lossguide\",\n      'max_leaves': 1400,  \n      'max_depth': 0, \n      'subsample': 0.9, \n      'colsample_bytree': 0.7, \n      'colsample_bylevel':0.7,\n      'min_child_weight':0,\n      'alpha':4,\n      'objective': 'binary:logitraw', \n      'scale_pos_weight':9,\n      'eval_metric': 'auc', \n      'nthread':8,\n      'random_state': 99, \n      'silent': False}\n</code></pre>\n\n<p><strong>can anybody give me some suggestion?\nthanks</strong></p>",
      "rawMarkdown": "I use lgb as first model，and xgb as second model, and **mean** the first model results,**but my outputs (second layer xgboost )are all ZEROs**,this is badly hurt me....\n\nI use code get Out-of-Fold Predictions below.\n   \n\n     def get_oof(clf,x_train,y_train,x_test,ntrain,ntest,kf):\n            oof_train = np.zeros((ntrain,))\n            oof_test = np.zeros((ntest,))\n            oof_test_skf = np.empty((NFOLDS, ntest))\n        \n            for i, (train_index, test_index) in enumerate(kf.split(x_train)):\n                x_tr = x_train.iloc[train_index]\n                y_tr = y_train.iloc[train_index]\n                x_te = x_train.iloc[test_index]\n        \n                clf.train(x_tr, y_tr)\n        \n                oof_train[test_index] = clf.predict(x_te)\n                print(oof_train[test_index][:20])\n                oof_test_skf[i, :] = clf.predict(x_test)\n        \n            oof_test[:] = oof_test_skf.mean(axis=0)\n            return oof_train.reshape(-1, 1), oof_test.reshape(-1, 1)\n\nmy lgb parameters are below:\n\n    lgb_params = {\n        'boosting_type': 'gbdt',\n        'objective': objective,\n        'metric':metrics,\n        'learning_rate': 0.2,\n        #'is_unbalance': 'true',  #because training data is unbalance (replaced with scale_pos_weight)\n        'num_leaves': 31,  # we should let it be smaller than 2^(max_depth)\n        'max_depth': -1,  # -1 means no limit\n        'min_child_samples': 20,  # Minimum number of data need in a child(min_data_in_leaf)\n        'max_bin': 255,  # Number of bucketed bin for feature values\n        'subsample': 0.6,  # Subsample ratio of the training instance.\n        'subsample_freq': 0,  # frequence of subsample, &lt;=0 means no enable\n        'colsample_bytree': 0.3,  # Subsample ratio of columns when constructing each tree.\n        'min_child_weight': 5,  # Minimum sum of instance weight(hessian) needed in a child(leaf)\n        'subsample_for_bin': 200000,  # Number of samples for constructing bin\n        'min_split_gain': 0,  # lambda_l1, lambda_l2 and min_gain_to_split to regularization\n        'reg_alpha': 0,  # L1 regularization term on weights\n        'reg_lambda': 0,  # L2 regularization term on weights\n        'nthread': 4,\n        'verbose': 0,\n        'metric':metrics\n    }\n\nmy xgboost parameters are below:\n\n    params = {'eta': 0.3,\n          'tree_method': \"hist\",\n          'grow_policy': \"lossguide\",\n          'max_leaves': 1400,  \n          'max_depth': 0, \n          'subsample': 0.9, \n          'colsample_bytree': 0.7, \n          'colsample_bylevel':0.7,\n          'min_child_weight':0,\n          'alpha':4,\n          'objective': 'binary:logitraw', \n          'scale_pos_weight':9,\n          'eval_metric': 'auc', \n          'nthread':8,\n          'random_state': 99, \n          'silent': False}\n\n**can anybody give me some suggestion?\nthanks**",
      "votes": null
    },
    {
      "id": "317994",
      "postDate": "04/23/2018 02:26:23",
      "content": "<p>Maybe the problem is \"oof_train[test_index] = clf.predict(x_te)\", you can change the clf.predict(x_te) to clf.predict_proba(x_te)[:, 1], same as the the \"oof_test_skf[i, :] = clf.predict(x_test)\"  below. Hope this can help you:)</p>",
      "rawMarkdown": "Maybe the problem is \"oof_train[test_index] = clf.predict(x_te)\", you can change the clf.predict(x_te) to clf.predict_proba(x_te)[:, 1], same as the the \"oof_test_skf[i, :] = clf.predict(x_test)\"  below. Hope this can help you:)",
      "votes": null
    },
    {
      "id": "318004",
      "postDate": "04/23/2018 03:03:14",
      "content": "<p>thanks for repplying ,I have try it. but output of  clf.predict(x_te) is one dimension array. this code will throw exception</p>",
      "rawMarkdown": "thanks for repplying ,I have try it. but output of  clf.predict(x_te) is one dimension array. this code will throw exception",
      "votes": null
    },
    {
      "id": "318014",
      "postDate": "04/23/2018 03:23:44",
      "content": "<p>I have question about using OOF. This data has timestamps, is it appropriate to split the data randomly into k folds? Very appreciate if LZ or anyone could clarify this ;) Thanks.</p>",
      "rawMarkdown": "I have question about using OOF. This data has timestamps, is it appropriate to split the data randomly into k folds? Very appreciate if LZ or anyone could clarify this ;) Thanks.",
      "votes": null
    },
    {
      "id": "318025",
      "postDate": "04/23/2018 03:51:45",
      "content": "<p>Sorry....What I wrote is incorrect. You method is right. Just forget about what I wrote before. It seemed like the code you showed is right. You mentioned the output is all zero. The output here stand for the output from the lightgbm or xgboost?</p>",
      "rawMarkdown": "Sorry....What I wrote is incorrect. You method is right. Just forget about what I wrote before. It seemed like the code you showed is right. You mentioned the output is all zero. The output here stand for the output from the lightgbm or xgboost?",
      "votes": null
    },
    {
      "id": "318026",
      "postDate": "04/23/2018 03:54:13",
      "content": "<p>Actually I think I will related to the feature you make.</p>",
      "rawMarkdown": "Actually I think I will related to the feature you make.",
      "votes": null
    },
    {
      "id": "318029",
      "postDate": "04/23/2018 04:07:03",
      "content": "<p>the output of xgboost (second layer) is all zero</p>",
      "rawMarkdown": "the output of xgboost (second layer) is all zero",
      "votes": null
    },
    {
      "id": "318030",
      "postDate": "04/23/2018 04:14:54",
      "content": "<p>Is the input of the xgboost only depends on the output of lightgbm or you combined with the original features? </p>",
      "rawMarkdown": "Is the input of the xgboost only depends on the output of lightgbm or you combined with the original features?",
      "votes": null
    },
    {
      "id": "318059",
      "postDate": "04/23/2018 05:14:41",
      "content": "<p>just depends on lgb output</p>",
      "rawMarkdown": "just depends on lgb output",
      "votes": null
    },
    {
      "id": "318062",
      "postDate": "04/23/2018 05:19:46",
      "content": "<p>Maybe you can combine with the original feature. I think that may help:)</p>",
      "rawMarkdown": "Maybe you can combine with the original feature. I think that may help:)",
      "votes": null
    },
    {
      "id": "320687",
      "postDate": "04/29/2018 13:57:08",
      "content": "<p>yes i have the same worry..</p>",
      "rawMarkdown": "yes i have the same worry..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 317994,
      "author_name": "a1339398734",
      "author_url": "",
      "post_date": "04/23/2018 02:26:23",
      "content": "<p>Maybe the problem is \"oof_train[test_index] = clf.predict(x_te)\", you can change the clf.predict(x_te) to clf.predict_proba(x_te)[:, 1], same as the the \"oof_test_skf[i, :] = clf.predict(x_test)\"  below. Hope this can help you:)</p>",
      "votes": null,
      "replies": [
        {
          "id": 318004,
          "author_name": "liuhdsgoal",
          "author_url": "",
          "post_date": "04/23/2018 03:03:14",
          "content": "<p>thanks for repplying ,I have try it. but output of  clf.predict(x_te) is one dimension array. this code will throw exception</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318025,
          "author_name": "a1339398734",
          "author_url": "",
          "post_date": "04/23/2018 03:51:45",
          "content": "<p>Sorry....What I wrote is incorrect. You method is right. Just forget about what I wrote before. It seemed like the code you showed is right. You mentioned the output is all zero. The output here stand for the output from the lightgbm or xgboost?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318029,
          "author_name": "liuhdsgoal",
          "author_url": "",
          "post_date": "04/23/2018 04:07:03",
          "content": "<p>the output of xgboost (second layer) is all zero</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318030,
          "author_name": "a1339398734",
          "author_url": "",
          "post_date": "04/23/2018 04:14:54",
          "content": "<p>Is the input of the xgboost only depends on the output of lightgbm or you combined with the original features? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318059,
          "author_name": "liuhdsgoal",
          "author_url": "",
          "post_date": "04/23/2018 05:14:41",
          "content": "<p>just depends on lgb output</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318062,
          "author_name": "a1339398734",
          "author_url": "",
          "post_date": "04/23/2018 05:19:46",
          "content": "<p>Maybe you can combine with the original feature. I think that may help:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 318014,
      "author_name": "bangdasun",
      "author_url": "",
      "post_date": "04/23/2018 03:23:44",
      "content": "<p>I have question about using OOF. This data has timestamps, is it appropriate to split the data randomly into k folds? Very appreciate if LZ or anyone could clarify this ;) Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 318026,
          "author_name": "a1339398734",
          "author_url": "",
          "post_date": "04/23/2018 03:54:13",
          "content": "<p>Actually I think I will related to the feature you make.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 320687,
          "author_name": "jinmahkust",
          "author_url": "",
          "post_date": "04/29/2018 13:57:08",
          "content": "<p>yes i have the same worry..</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "317982": "I use lgb as first model，and xgb as second model, and **mean** the first model results,**but my outputs (second layer xgboost )are all ZEROs**,this is badly hurt me....\n\nI use code get Out-of-Fold Predictions below.\n   \n\n     def get_oof(clf,x_train,y_train,x_test,ntrain,ntest,kf):\n            oof_train = np.zeros((ntrain,))\n            oof_test = np.zeros((ntest,))\n            oof_test_skf = np.empty((NFOLDS, ntest))\n        \n            for i, (train_index, test_index) in enumerate(kf.split(x_train)):\n                x_tr = x_train.iloc[train_index]\n                y_tr = y_train.iloc[train_index]\n                x_te = x_train.iloc[test_index]\n        \n                clf.train(x_tr, y_tr)\n        \n                oof_train[test_index] = clf.predict(x_te)\n                print(oof_train[test_index][:20])\n                oof_test_skf[i, :] = clf.predict(x_test)\n        \n            oof_test[:] = oof_test_skf.mean(axis=0)\n            return oof_train.reshape(-1, 1), oof_test.reshape(-1, 1)\n\nmy lgb parameters are below:\n\n    lgb_params = {\n        'boosting_type': 'gbdt',\n        'objective': objective,\n        'metric':metrics,\n        'learning_rate': 0.2,\n        #'is_unbalance': 'true',  #because training data is unbalance (replaced with scale_pos_weight)\n        'num_leaves': 31,  # we should let it be smaller than 2^(max_depth)\n        'max_depth': -1,  # -1 means no limit\n        'min_child_samples': 20,  # Minimum number of data need in a child(min_data_in_leaf)\n        'max_bin': 255,  # Number of bucketed bin for feature values\n        'subsample': 0.6,  # Subsample ratio of the training instance.\n        'subsample_freq': 0,  # frequence of subsample, &lt;=0 means no enable\n        'colsample_bytree': 0.3,  # Subsample ratio of columns when constructing each tree.\n        'min_child_weight': 5,  # Minimum sum of instance weight(hessian) needed in a child(leaf)\n        'subsample_for_bin': 200000,  # Number of samples for constructing bin\n        'min_split_gain': 0,  # lambda_l1, lambda_l2 and min_gain_to_split to regularization\n        'reg_alpha': 0,  # L1 regularization term on weights\n        'reg_lambda': 0,  # L2 regularization term on weights\n        'nthread': 4,\n        'verbose': 0,\n        'metric':metrics\n    }\n\nmy xgboost parameters are below:\n\n    params = {'eta': 0.3,\n          'tree_method': \"hist\",\n          'grow_policy': \"lossguide\",\n          'max_leaves': 1400,  \n          'max_depth': 0, \n          'subsample': 0.9, \n          'colsample_bytree': 0.7, \n          'colsample_bylevel':0.7,\n          'min_child_weight':0,\n          'alpha':4,\n          'objective': 'binary:logitraw', \n          'scale_pos_weight':9,\n          'eval_metric': 'auc', \n          'nthread':8,\n          'random_state': 99, \n          'silent': False}\n\n**can anybody give me some suggestion?\nthanks**",
    "317994": "Maybe the problem is \"oof_train[test_index] = clf.predict(x_te)\", you can change the clf.predict(x_te) to clf.predict_proba(x_te)[:, 1], same as the the \"oof_test_skf[i, :] = clf.predict(x_test)\"  below. Hope this can help you:)",
    "318004": "thanks for repplying ,I have try it. but output of  clf.predict(x_te) is one dimension array. this code will throw exception",
    "318014": "I have question about using OOF. This data has timestamps, is it appropriate to split the data randomly into k folds? Very appreciate if LZ or anyone could clarify this ;) Thanks.",
    "318025": "Sorry....What I wrote is incorrect. You method is right. Just forget about what I wrote before. It seemed like the code you showed is right. You mentioned the output is all zero. The output here stand for the output from the lightgbm or xgboost?",
    "318026": "Actually I think I will related to the feature you make.",
    "318029": "the output of xgboost (second layer) is all zero",
    "318030": "Is the input of the xgboost only depends on the output of lightgbm or you combined with the original features?",
    "318059": "just depends on lgb output",
    "318062": "Maybe you can combine with the original feature. I think that may help:)",
    "320687": "yes i have the same worry.."
  },
  "source": "meta"
}