{
  "id": 500226,
  "title": "LIghtGBM MOdel train vs fit",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/500226",
  "author_name": "Anil Kumar Reddy",
  "post_date": "2024-05-04T17:21:48.219000",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>When I was experimenting with multiple model to check which model is able to perform well &amp; which is giving random guess , got to found LightGBM model train method was giving above 0.8 ROC AUC score same param with LGBMClassifier fit method could not able to get score above 0.5 of ROC AUC is their any difference!! </p>\n<p>Please have a look at below notebook in Light GBM Model Section : <br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/kunduruanil/credit-risk-model-feature-selection</a></p>",
  "messages": [
    {
      "id": 2793969,
      "postDate": "2024-05-05T05:10:33.143Z",
      "content": "<p>Please check if you are using <strong>predict</strong>/ <strong>predict_proba</strong> for your scikit-learn interface <a href=\"https://www.kaggle.com/kunduruanil\" target=\"_blank\">@kunduruanil</a> <br>\nUsing <strong>predict</strong> usually produces low scores. <br>\nIn the native booster syntax, you can freely use predict for regression and classification problems <a href=\"https://www.kaggle.com/kunduruanil\" target=\"_blank\">@kunduruanil</a> </p>",
      "rawMarkdown": "Please check if you are using **predict**/ **predict_proba** for your scikit-learn interface @kunduruanil \nUsing **predict** usually produces low scores. \nIn the native booster syntax, you can freely use predict for regression and classification problems @kunduruanil ",
      "votes": 1,
      "replies": [
        {
          "id": 2793975,
          "postDate": "2024-05-05T05:18:38.587Z",
          "content": "<p>Currently I am using predict , try with predict_proba !! </p>\n<p>I think you are tagging haileynah not mee !! </p>",
          "rawMarkdown": "Currently I am using predict , try with predict_proba !! \n\nI think you are tagging haileynah not mee !! ",
          "votes": 1,
          "replies": [
            {
              "id": 2793994,
              "postDate": "2024-05-05T05:27:09.857Z",
              "content": "<p>So sorry for the wrong tagging, rectified it <a href=\"https://www.kaggle.com/kunduruanil\" target=\"_blank\">@kunduruanil</a>  😀<br>\nGood that the suggestion helped! I prefer the native booster syntax for this very reason. You can use the same syntax for classification and regression and use model.predict() for any type of predictions as well</p>",
              "rawMarkdown": "So sorry for the wrong tagging, rectified it @kunduruanil  😀\nGood that the suggestion helped! I prefer the native booster syntax for this very reason. You can use the same syntax for classification and regression and use model.predict() for any type of predictions as well",
              "votes": 1
            },
            {
              "id": 2794901,
              "postDate": "2024-05-05T15:06:24.933Z",
              "content": "<p>I am using below code to evaluate could not found any improvement :</p>\n<p>def evalvate(model,X,y):<br>\n    y_prob = model.predict_proba(X)<br>\n    threshold = 0.5<br>\n    y_pred = np.where(y_prob[:, 1] &gt;= threshold, 1, 0)<br>\n    print(f\"ROC AUC Score : {roc_auc_score(y,y_pred)}\")<br>\n    y_pred = np.where(y_pred&gt;=0.5,1,0)<br>\n    print(f\"Precision Score : {precision_score(y,y_pred)}\")<br>\n    print(\"<em>\"</em>50)</p>\n<p>please check notebook as well : <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/kunduruanil/credit-risk-model-feature-selection</a></p>",
              "rawMarkdown": "I am using below code to evaluate could not found any improvement :\n\ndef evalvate(model,X,y):\n    y_prob = model.predict_proba(X)\n    threshold = 0.5\n    y_pred = np.where(y_prob[:, 1] >= threshold, 1, 0)\n    print(f\"ROC AUC Score : {roc_auc_score(y,y_pred)}\")\n    y_pred = np.where(y_pred>=0.5,1,0)\n    print(f\"Precision Score : {precision_score(y,y_pred)}\")\n    print(\"*\"*50)\n\nplease check notebook as well : [https://www.kaggle.com/code/kunduruanil/credit-risk-model-feature-selection](url)"
            },
            {
              "id": 2797015,
              "postDate": "2024-05-06T14:00:13.930Z",
              "content": "<p>Check your embedded links… a click on these links don't take people to where you think they should. </p>",
              "rawMarkdown": "Check your embedded links... a click on these links don't take people to where you think they should. "
            }
          ]
        }
      ]
    },
    {
      "id": 2793378,
      "postDate": "2024-05-04T17:21:48.220Z",
      "content": "<p>When I was experimenting with multiple model to check which model is able to perform well &amp; which is giving random guess , got to found LightGBM model train method was giving above 0.8 ROC AUC score same param with LGBMClassifier fit method could not able to get score above 0.5 of ROC AUC is their any difference!! </p>\n<p>Please have a look at below notebook in Light GBM Model Section : <br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/kunduruanil/credit-risk-model-feature-selection</a></p>",
      "rawMarkdown": "When I was experimenting with multiple model to check which model is able to perform well & which is giving random guess , got to found LightGBM model train method was giving above 0.8 ROC AUC score same param with LGBMClassifier fit method could not able to get score above 0.5 of ROC AUC is their any difference!! \n\n\nPlease have a look at below notebook in Light GBM Model Section : \n[https://www.kaggle.com/code/kunduruanil/credit-risk-model-feature-selection](url)",
      "votes": 2
    },
    {
      "id": 2799899,
      "postDate": "2024-05-08T03:04:32.393Z",
      "content": "<p>lgb train also seems to provide easier ways to use custom metrics. By comparison, using custom metric in LGBMClassifier requires writing callback functions and such</p>",
      "rawMarkdown": "lgb train also seems to provide easier ways to use custom metrics. By comparison, using custom metric in LGBMClassifier requires writing callback functions and such",
      "replies": [
        {
          "id": 2813553,
          "postDate": "2024-05-14T20:23:11.470Z",
          "content": "<p>With fit method it's implemented in this notebook <a href=\"https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric\" target=\"_blank\">https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric</a></p>",
          "rawMarkdown": "With fit method it's implemented in this notebook https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric"
        }
      ]
    },
    {
      "id": 2797038,
      "postDate": "2024-05-06T14:06:50.923Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2793504,
      "postDate": "2024-05-04T18:52:35.253Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2793969,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2024-05-05T05:10:33.143000",
      "content": "<p>Please check if you are using <strong>predict</strong>/ <strong>predict_proba</strong> for your scikit-learn interface <a href=\"https://www.kaggle.com/kunduruanil\" target=\"_blank\">@kunduruanil</a> <br>\nUsing <strong>predict</strong> usually produces low scores. <br>\nIn the native booster syntax, you can freely use predict for regression and classification problems <a href=\"https://www.kaggle.com/kunduruanil\" target=\"_blank\">@kunduruanil</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2793975,
          "author_name": "Anil Kumar Reddy",
          "author_url": "",
          "post_date": "2024-05-05T05:18:38.587000",
          "content": "<p>Currently I am using predict , try with predict_proba !! </p>\n<p>I think you are tagging haileynah not mee !! </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2793994,
              "author_name": "Ravi Ramakrishnan",
              "author_url": "",
              "post_date": "2024-05-05T05:27:09.857000",
              "content": "<p>So sorry for the wrong tagging, rectified it <a href=\"https://www.kaggle.com/kunduruanil\" target=\"_blank\">@kunduruanil</a>  😀<br>\nGood that the suggestion helped! I prefer the native booster syntax for this very reason. You can use the same syntax for classification and regression and use model.predict() for any type of predictions as well</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2794901,
              "author_name": "Anil Kumar Reddy",
              "author_url": "",
              "post_date": "2024-05-05T15:06:24.933000",
              "content": "<p>I am using below code to evaluate could not found any improvement :</p>\n<p>def evalvate(model,X,y):<br>\n    y_prob = model.predict_proba(X)<br>\n    threshold = 0.5<br>\n    y_pred = np.where(y_prob[:, 1] &gt;= threshold, 1, 0)<br>\n    print(f\"ROC AUC Score : {roc_auc_score(y,y_pred)}\")<br>\n    y_pred = np.where(y_pred&gt;=0.5,1,0)<br>\n    print(f\"Precision Score : {precision_score(y,y_pred)}\")<br>\n    print(\"<em>\"</em>50)</p>\n<p>please check notebook as well : <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/kunduruanil/credit-risk-model-feature-selection</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2797015,
              "author_name": "Eduard Stefanescu",
              "author_url": "",
              "post_date": "2024-05-06T14:00:13.930000",
              "content": "<p>Check your embedded links… a click on these links don't take people to where you think they should. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2799899,
      "author_name": "TkiJoestar",
      "author_url": "",
      "post_date": "2024-05-08T03:04:32.393000",
      "content": "<p>lgb train also seems to provide easier ways to use custom metrics. By comparison, using custom metric in LGBMClassifier requires writing callback functions and such</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2813553,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-05-14T20:23:11.470000",
          "content": "<p>With fit method it's implemented in this notebook <a href=\"https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric\" target=\"_blank\">https://www.kaggle.com/code/eu1234/creditrisk-sklearn-pipeline-integrated-metric</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2797038,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-05-06T14:06:50.923000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2793504,
      "author_name": "haileynah",
      "author_url": "",
      "post_date": "2024-05-04T18:52:35.253000",
      "content": "<p>Thank you for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2793969": "Please check if you are using **predict**/ **predict_proba** for your scikit-learn interface @kunduruanil \nUsing **predict** usually produces low scores. \nIn the native booster syntax, you can freely use predict for regression and classification problems @kunduruanil ",
    "2793378": "When I was experimenting with multiple model to check which model is able to perform well & which is giving random guess , got to found LightGBM model train method was giving above 0.8 ROC AUC score same param with LGBMClassifier fit method could not able to get score above 0.5 of ROC AUC is their any difference!! \n\n\nPlease have a look at below notebook in Light GBM Model Section : \n[https://www.kaggle.com/code/kunduruanil/credit-risk-model-feature-selection](url)",
    "2799899": "lgb train also seems to provide easier ways to use custom metrics. By comparison, using custom metric in LGBMClassifier requires writing callback functions and such",
    "2797038": "",
    "2793504": "Thank you for sharing!"
  }
}