{
  "id": 205856,
  "title": "Xgboost nan outputs",
  "url": "/competitions/rfcx-species-audio-detection/discussion/205856",
  "author_name": "",
  "post_date": "2020-12-22T07:18:13.986282700Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I tried to use xgboost in the competition but i think god did not like it .</p>\n<p>I tried to predict outputs in probability as i needed them in probability not normal output everywhere there was nan , i used predict function i got all the answers as 0</p>\n<p>Note there was no nan in y_train </p>\n<p>Please help 🤒</p>",
  "messages": [
    {
      "id": "1122087",
      "postDate": "12/22/2020 07:18:13",
      "content": "<p>I tried to use xgboost in the competition but i think god did not like it .</p>\n<p>I tried to predict outputs in probability as i needed them in probability not normal output everywhere there was nan , i used predict function i got all the answers as 0</p>\n<p>Note there was no nan in y_train </p>\n<p>Please help 🤒</p>",
      "rawMarkdown": "I tried to use xgboost in the competition but i think god did not like it .\n\nI tried to predict outputs in probability as i needed them in probability not normal output everywhere there was nan , i used predict function i got all the answers as 0\n\nNote there was no nan in y_train \n\nPlease help 🤒",
      "votes": null
    },
    {
      "id": "1123835",
      "postDate": "12/23/2020 14:33:27",
      "content": "<p>Hello there,</p>\n<p>you can try to use <code>predict_proba</code> instead of <code>predict</code>.  <code>predict_proba</code>  predicts the probability of each sample, <code>predict</code>  simply returns 0 or 1 as the classified label, not the probabilities. </p>",
      "rawMarkdown": "Hello there,\n\nyou can try to use `predict_proba` instead of `predict`.  `predict_proba`  predicts the probability of each sample, `predict`  simply returns 0 or 1 as the classified label, not the probabilities.",
      "votes": null
    },
    {
      "id": "1123851",
      "postDate": "12/23/2020 14:42:46",
      "content": "<p>Sorry but you completely miss understood it I have used <code>predict_proba</code> all the values are coming nan when I used <code>predict</code> all the values come 0</p>",
      "rawMarkdown": "Sorry but you completely miss understood it I have used `predict_proba ` all the values are coming nan when I used `predict ` all the values come 0",
      "votes": null
    },
    {
      "id": "1123867",
      "postDate": "12/23/2020 14:49:19",
      "content": "<p>Can you post the code here that causes this behaviour   or  publish your notebook?  Otherwise it's hard to tell what could cause this problem.</p>",
      "rawMarkdown": "Can you post the code here that causes this behaviour   or  publish your notebook?  Otherwise it's hard to tell what could cause this problem.",
      "votes": null
    },
    {
      "id": "1123971",
      "postDate": "12/23/2020 16:08:01",
      "content": "<p>here is the code :-<br>\n'''<br>\nimport xgboost<br>\ndef get_model(Y):<br>\n    return xgboost.XGBClassifier(n_estimators=1000,<br>\n                                   max_depth=4,<br>\n                                   learning_rate=0.05,<br>\n                                   verbosity=0,<br>\n                                   objective=\"binary:logistic\",<br>\n                                   subsample=0.95,<br>\n                                   colsample_bytree=0.95,<br>\n                                   random_state=2021,<br>\n                                   n_jobs=2,<br>\n                                   scale_pos_weight = np.sum(Y==0) / np.sum(Y==1),<br>\n                                  ) </p>\n<p>'''<br>\n'''<br>\nfrom sklearn.model_selection import KFold<br>\nimport pickle<br>\nt=Y.T<br>\nmodels=[]<br>\naverages=[]<br>\ndef extract(indexes,i):<br>\n    global X<br>\n    global train</p>\n<pre><code>s_X=[train[ind] for ind in indexes]\ns_Y=[int(Y[ind][i]) for ind in indexes]\nreturn (s_X,s_Y)\n</code></pre>\n<p>kf = KFold(n_splits=3)</p>\n<p>for s in range(0,24):<br>\n    temp=[0,None]# accuracy , model<br>\n    ac=[]# scores<br>\n    g=0<br>\n    for ind_train,ind_test in kf.split(train):<br>\n        train2=extract(ind_train,s)<br>\n        test2=extract(ind_test,s)<br>\n        X_train,y_train=train2<br>\n        X_test,y_test=test2<br>\n        try:<br>\n             g=list(y_test).index(np.nan)<br>\n             print('y')<br>\n        except:<br>\n            g=1</p>\n<pre><code>    print(X_train,X_test)\n    model.fit(np.array(X_train),np.array(y_train))\n    score=model.score(np.array(X_test),np.array(y_test))\n    ac.append(score)\n    if score&gt;temp[0]:\n        temp[0]=score\n        temp[1]=model\n    print(g)\n    g+=1\n\n\npickle.dump(temp[1],open(str(s)+'xgboost.pkl','wb'))\n\nprint('------------------------------------')\nprint('MaxScore :',' ',s,temp[0])\nprint('------------------------------------')\nprint('Average :',' ',s,sum(ac)/len(ac))\naverages.append(sum(ac)/len(ac))\nmodels.append(temp[1])\n</code></pre>\n<p>print('average score:',sum(averages)/len(averages))<br>\n'''</p>\n<p>Sorry , you will some  problem reading it</p>",
      "rawMarkdown": "here is the code :-\n'''\nimport xgboost\ndef get_model(Y):\n    return xgboost.XGBClassifier(n_estimators=1000,\n                                   max_depth=4,\n                                   learning_rate=0.05,\n                                   verbosity=0,\n                                   objective=\"binary:logistic\",\n                                   subsample=0.95,\n                                   colsample_bytree=0.95,\n                                   random_state=2021,\n                                   n_jobs=2,\n                                   scale_pos_weight = np.sum(Y==0) / np.sum(Y==1),\n                                  ) \n                                   \n                                   \n'''\n'''\nfrom sklearn.model_selection import KFold\nimport pickle\nt=Y.T\nmodels=[]\naverages=[]\ndef extract(indexes,i):\n    global X\n    global train\n    \n    s_X=[train[ind] for ind in indexes]\n    s_Y=[int(Y[ind][i]) for ind in indexes]\n    return (s_X,s_Y)\nkf = KFold(n_splits=3)\n\nfor s in range(0,24):\n    temp=[0,None]# accuracy , model\n    ac=[]# scores\n    g=0\n    for ind_train,ind_test in kf.split(train):\n        train2=extract(ind_train,s)\n        test2=extract(ind_test,s)\n        X_train,y_train=train2\n        X_test,y_test=test2\n        try:\n             g=list(y_test).index(np.nan)\n             print('y')\n        except:\n            g=1\n        \n        print(X_train,X_test)\n        model.fit(np.array(X_train),np.array(y_train))\n        score=model.score(np.array(X_test),np.array(y_test))\n        ac.append(score)\n        if score>temp[0]:\n            temp[0]=score\n            temp[1]=model\n        print(g)\n        g+=1\n        \n    \n    pickle.dump(temp[1],open(str(s)+'xgboost.pkl','wb'))\n            \n    print('------------------------------------')\n    print('MaxScore :',' ',s,temp[0])\n    print('------------------------------------')\n    print('Average :',' ',s,sum(ac)/len(ac))\n    averages.append(sum(ac)/len(ac))\n    models.append(temp[1])\n    \nprint('average score:',sum(averages)/len(averages))\n'''\n\nSorry , you will some  problem reading it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1123835,
      "author_name": "jonas0",
      "author_url": "",
      "post_date": "12/23/2020 14:33:27",
      "content": "<p>Hello there,</p>\n<p>you can try to use <code>predict_proba</code> instead of <code>predict</code>.  <code>predict_proba</code>  predicts the probability of each sample, <code>predict</code>  simply returns 0 or 1 as the classified label, not the probabilities. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1123851,
          "author_name": "swaralipibose",
          "author_url": "",
          "post_date": "12/23/2020 14:42:46",
          "content": "<p>Sorry but you completely miss understood it I have used <code>predict_proba</code> all the values are coming nan when I used <code>predict</code> all the values come 0</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123867,
          "author_name": "jonas0",
          "author_url": "",
          "post_date": "12/23/2020 14:49:19",
          "content": "<p>Can you post the code here that causes this behaviour   or  publish your notebook?  Otherwise it's hard to tell what could cause this problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123971,
          "author_name": "swaralipibose",
          "author_url": "",
          "post_date": "12/23/2020 16:08:01",
          "content": "<p>here is the code :-<br>\n'''<br>\nimport xgboost<br>\ndef get_model(Y):<br>\n    return xgboost.XGBClassifier(n_estimators=1000,<br>\n                                   max_depth=4,<br>\n                                   learning_rate=0.05,<br>\n                                   verbosity=0,<br>\n                                   objective=\"binary:logistic\",<br>\n                                   subsample=0.95,<br>\n                                   colsample_bytree=0.95,<br>\n                                   random_state=2021,<br>\n                                   n_jobs=2,<br>\n                                   scale_pos_weight = np.sum(Y==0) / np.sum(Y==1),<br>\n                                  ) </p>\n<p>'''<br>\n'''<br>\nfrom sklearn.model_selection import KFold<br>\nimport pickle<br>\nt=Y.T<br>\nmodels=[]<br>\naverages=[]<br>\ndef extract(indexes,i):<br>\n    global X<br>\n    global train</p>\n<pre><code>s_X=[train[ind] for ind in indexes]\ns_Y=[int(Y[ind][i]) for ind in indexes]\nreturn (s_X,s_Y)\n</code></pre>\n<p>kf = KFold(n_splits=3)</p>\n<p>for s in range(0,24):<br>\n    temp=[0,None]# accuracy , model<br>\n    ac=[]# scores<br>\n    g=0<br>\n    for ind_train,ind_test in kf.split(train):<br>\n        train2=extract(ind_train,s)<br>\n        test2=extract(ind_test,s)<br>\n        X_train,y_train=train2<br>\n        X_test,y_test=test2<br>\n        try:<br>\n             g=list(y_test).index(np.nan)<br>\n             print('y')<br>\n        except:<br>\n            g=1</p>\n<pre><code>    print(X_train,X_test)\n    model.fit(np.array(X_train),np.array(y_train))\n    score=model.score(np.array(X_test),np.array(y_test))\n    ac.append(score)\n    if score&gt;temp[0]:\n        temp[0]=score\n        temp[1]=model\n    print(g)\n    g+=1\n\n\npickle.dump(temp[1],open(str(s)+'xgboost.pkl','wb'))\n\nprint('------------------------------------')\nprint('MaxScore :',' ',s,temp[0])\nprint('------------------------------------')\nprint('Average :',' ',s,sum(ac)/len(ac))\naverages.append(sum(ac)/len(ac))\nmodels.append(temp[1])\n</code></pre>\n<p>print('average score:',sum(averages)/len(averages))<br>\n'''</p>\n<p>Sorry , you will some  problem reading it</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1122087": "I tried to use xgboost in the competition but i think god did not like it .\n\nI tried to predict outputs in probability as i needed them in probability not normal output everywhere there was nan , i used predict function i got all the answers as 0\n\nNote there was no nan in y_train \n\nPlease help 🤒",
    "1123835": "Hello there,\n\nyou can try to use `predict_proba` instead of `predict`.  `predict_proba`  predicts the probability of each sample, `predict`  simply returns 0 or 1 as the classified label, not the probabilities.",
    "1123851": "Sorry but you completely miss understood it I have used `predict_proba ` all the values are coming nan when I used `predict ` all the values come 0",
    "1123867": "Can you post the code here that causes this behaviour   or  publish your notebook?  Otherwise it's hard to tell what could cause this problem.",
    "1123971": "here is the code :-\n'''\nimport xgboost\ndef get_model(Y):\n    return xgboost.XGBClassifier(n_estimators=1000,\n                                   max_depth=4,\n                                   learning_rate=0.05,\n                                   verbosity=0,\n                                   objective=\"binary:logistic\",\n                                   subsample=0.95,\n                                   colsample_bytree=0.95,\n                                   random_state=2021,\n                                   n_jobs=2,\n                                   scale_pos_weight = np.sum(Y==0) / np.sum(Y==1),\n                                  ) \n                                   \n                                   \n'''\n'''\nfrom sklearn.model_selection import KFold\nimport pickle\nt=Y.T\nmodels=[]\naverages=[]\ndef extract(indexes,i):\n    global X\n    global train\n    \n    s_X=[train[ind] for ind in indexes]\n    s_Y=[int(Y[ind][i]) for ind in indexes]\n    return (s_X,s_Y)\nkf = KFold(n_splits=3)\n\nfor s in range(0,24):\n    temp=[0,None]# accuracy , model\n    ac=[]# scores\n    g=0\n    for ind_train,ind_test in kf.split(train):\n        train2=extract(ind_train,s)\n        test2=extract(ind_test,s)\n        X_train,y_train=train2\n        X_test,y_test=test2\n        try:\n             g=list(y_test).index(np.nan)\n             print('y')\n        except:\n            g=1\n        \n        print(X_train,X_test)\n        model.fit(np.array(X_train),np.array(y_train))\n        score=model.score(np.array(X_test),np.array(y_test))\n        ac.append(score)\n        if score>temp[0]:\n            temp[0]=score\n            temp[1]=model\n        print(g)\n        g+=1\n        \n    \n    pickle.dump(temp[1],open(str(s)+'xgboost.pkl','wb'))\n            \n    print('------------------------------------')\n    print('MaxScore :',' ',s,temp[0])\n    print('------------------------------------')\n    print('Average :',' ',s,sum(ac)/len(ac))\n    averages.append(sum(ac)/len(ac))\n    models.append(temp[1])\n    \nprint('average score:',sum(averages)/len(averages))\n'''\n\nSorry , you will some  problem reading it"
  },
  "source": "meta"
}