{
  "id": 416587,
  "title": "Newbie question || What am I doing wrong with XGBoost?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/416587",
  "author_name": "",
  "post_date": "2023-06-12T08:27:11.406882200Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hello 👀</p>\n<p>trying to tune XGBoost with Optuna, it takes way less time than LightGBM (about 3-10 seconds vs 30-100 using LightGBM). Beforehands only for XGBoost I do vanilla preprocessing filling nans and replacing infs with median in numerics (otherwise it didn't start).<br>\nThe score is also not too big comparing to LightGBM and CatBoost: 0.8485 (CB) vs 0.8493 (LightGBM) vs 0.8467 (XGB).</p>\n<p>Here are params I'm tuning:</p>\n<pre><code>    xgb_params = {\n        : trial(, e-, , log=True),\n        : trial(, e-, , log=True),\n        : trial(, , ),\n        : trial(, , ),\n        : trial(, , , log=True),\n        : trial(, , ),\n        : trial(, , ),\n        : trial(, , ),\n        : trial(, , ),\n        : trial(, ),\n\n        : trial(, ),\n        : ,\n        : trial(, ),\n        : trial(, ),\n        : trial(, ),\n    }\n</code></pre>\n<p>Here's my additional preprocessing:</p>\n<pre><code>    df_question = agg_dataframe()\n\n    numerics = \n    num_cols = df_question(include=numerics)\n     num_col  num_cols:\n        median_col = df_question()\n        median_col = median_col  median_col not    \n        df_question = df_question(, median_col)\n</code></pre>\n<p>What might be wrong?</p>",
  "messages": [
    {
      "id": "2296896",
      "postDate": "06/12/2023 08:27:11",
      "content": "<p>Hello 👀</p>\n<p>trying to tune XGBoost with Optuna, it takes way less time than LightGBM (about 3-10 seconds vs 30-100 using LightGBM). Beforehands only for XGBoost I do vanilla preprocessing filling nans and replacing infs with median in numerics (otherwise it didn't start).<br>\nThe score is also not too big comparing to LightGBM and CatBoost: 0.8485 (CB) vs 0.8493 (LightGBM) vs 0.8467 (XGB).</p>\n<p>Here are params I'm tuning:</p>\n<pre><code>    xgb_params = {\n        : trial(, e-, , log=True),\n        : trial(, e-, , log=True),\n        : trial(, , ),\n        : trial(, , ),\n        : trial(, , , log=True),\n        : trial(, , ),\n        : trial(, , ),\n        : trial(, , ),\n        : trial(, , ),\n        : trial(, ),\n\n        : trial(, ),\n        : ,\n        : trial(, ),\n        : trial(, ),\n        : trial(, ),\n    }\n</code></pre>\n<p>Here's my additional preprocessing:</p>\n<pre><code>    df_question = agg_dataframe()\n\n    numerics = \n    num_cols = df_question(include=numerics)\n     num_col  num_cols:\n        median_col = df_question()\n        median_col = median_col  median_col not    \n        df_question = df_question(, median_col)\n</code></pre>\n<p>What might be wrong?</p>",
      "rawMarkdown": "Hello 👀\n\ntrying to tune XGBoost with Optuna, it takes way less time than LightGBM (about 3-10 seconds vs 30-100 using LightGBM). Beforehands only for XGBoost I do vanilla preprocessing filling nans and replacing infs with median in numerics (otherwise it didn't start).\nThe score is also not too big comparing to LightGBM and CatBoost: 0.8485 (CB) vs 0.8493 (LightGBM) vs 0.8467 (XGB).\n\nHere are params I'm tuning:\n```\n    xgb_params = {\n        'lambda': trial.suggest_float('lambda', 1e-3, 1.0, log=True),\n        'alpha': trial.suggest_float('alpha', 1e-3, 1.0, log=True),\n        'colsample_bytree': trial.suggest_float('colsample_bytree', 0.05, 1.0),\n        'subsample': trial.suggest_float('subsample', 0.05, 1.0),\n        'learning_rate': trial.suggest_float(\"learning_rate\", 0.001, 0.05, log=True),\n        \"n_estimators\": trial.suggest_int(\"n_estimators\", 1000, 10000),\n        'max_depth': trial.suggest_int(\"max_depth\", 6, 12),\n        'num_round': trial.suggest_int(\"num_round\", 500, 10000),\n        'min_child_weight': trial.suggest_int('min_child_weight', 1, 300),\n        'use_label_encoder': trial.suggest_categorical('use_label_encoder', [True, False]),\n     \n        \"objective\": trial.suggest_categorical(\"objective\", [\"binary:logistic\"]),\n        \"verbosity\": 0,\n        \"boosting_type\": trial.suggest_categorical(\"boosting_type\", [\"gbdt\"]),\n        'random_state': trial.suggest_categorical('random_state', [42]),\n        'tree_method': trial.suggest_categorical('tree_method', ['hist']),\n    }\n```\nHere's my additional preprocessing:\n```\n    df_question = agg_dataframe.to_pandas().loc[:, features]\n\n    numerics = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    num_cols = df_question.select_dtypes(include=numerics).columns\n    for num_col in num_cols:\n        median_col = df_question[num_col].median()\n        median_col = median_col if median_col not in [np.inf, -np.inf, np.nan] else 0\n        df_question[num_col] = df_question[num_col].replace([np.inf, -np.inf, np.nan], median_col)\n```\n\nWhat might be wrong?",
      "votes": null
    },
    {
      "id": "2296913",
      "postDate": "06/12/2023 08:51:13",
      "content": "<p>What evaluation metric are you using when you mentioned 'score'? If you are using the same metric with the submission (which is Marco f1), then 0.848 is impossibly high compared with the current leaderboard. The most possible mistake is data leakage somewhere, which is more related to how you set your train/test dataset rather than tuning.</p>",
      "rawMarkdown": "What evaluation metric are you using when you mentioned 'score'? If you are using the same metric with the submission (which is Marco f1), then 0.848 is impossibly high compared with the current leaderboard. The most possible mistake is data leakage somewhere, which is more related to how you set your train/test dataset rather than tuning.",
      "votes": null
    },
    {
      "id": "2296926",
      "postDate": "06/12/2023 09:04:24",
      "content": "<p>No, it's F1 just for the 1st question, and it wasn't submitted yet.</p>",
      "rawMarkdown": "No, it's F1 just for the 1st question, and it wasn't submitted yet.",
      "votes": null
    },
    {
      "id": "2296940",
      "postDate": "06/12/2023 09:14:10",
      "content": "<blockquote>\n  <p>No, it's F1 just for the 1st question, and it wasn't submitted yet.</p>\n</blockquote>\n<p>Then why do you think XGBoost is 'wrong'? 0.002 is a sensible shake when you using a different model.</p>",
      "rawMarkdown": "> No, it's F1 just for the 1st question, and it wasn't submitted yet.\n\nThen why do you think XGBoost is 'wrong'? 0.002 is a sensible shake when you using a different model.",
      "votes": null
    },
    {
      "id": "2296970",
      "postDate": "06/12/2023 09:35:53",
      "content": "<p><a href=\"https://www.kaggle.com/yk4r22\" target=\"_blank\">@yk4r22</a>, I am not sure if your search space for the optuna tuning step is appropriate. You could perhaps try a smaller search space and assess the impact. <br>\nAlso, do you need to tune so many parameters?</p>",
      "rawMarkdown": "yk4r22, I am not sure if your search space for the optuna tuning step is appropriate. You could perhaps try a smaller search space and assess the impact. \nAlso, do you need to tune so many parameters?",
      "votes": null
    },
    {
      "id": "2297229",
      "postDate": "06/12/2023 12:34:45",
      "content": "<p>Because it takes 10 seconds for it to cross-validate over aggregates, while for CB and LightGBM it takes 30-100s and 100-800s respectively. Kinda weird :) 🤔</p>",
      "rawMarkdown": "Because it takes 10 seconds for it to cross-validate over aggregates, while for CB and LightGBM it takes 30-100s and 100-800s respectively. Kinda weird :) 🤔",
      "votes": null
    },
    {
      "id": "2297231",
      "postDate": "06/12/2023 12:35:43",
      "content": "<p>I see, this is just the preparation, I'll decrease the search space in the future, thank you!</p>",
      "rawMarkdown": "I see, this is just the preparation, I'll decrease the search space in the future, thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2296913,
      "author_name": "ataraxian",
      "author_url": "",
      "post_date": "06/12/2023 08:51:13",
      "content": "<p>What evaluation metric are you using when you mentioned 'score'? If you are using the same metric with the submission (which is Marco f1), then 0.848 is impossibly high compared with the current leaderboard. The most possible mistake is data leakage somewhere, which is more related to how you set your train/test dataset rather than tuning.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2296926,
          "author_name": "yk4r22",
          "author_url": "",
          "post_date": "06/12/2023 09:04:24",
          "content": "<p>No, it's F1 just for the 1st question, and it wasn't submitted yet.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2296940,
              "author_name": "ataraxian",
              "author_url": "",
              "post_date": "06/12/2023 09:14:10",
              "content": "<blockquote>\n  <p>No, it's F1 just for the 1st question, and it wasn't submitted yet.</p>\n</blockquote>\n<p>Then why do you think XGBoost is 'wrong'? 0.002 is a sensible shake when you using a different model.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2297229,
                  "author_name": "yk4r22",
                  "author_url": "",
                  "post_date": "06/12/2023 12:34:45",
                  "content": "<p>Because it takes 10 seconds for it to cross-validate over aggregates, while for CB and LightGBM it takes 30-100s and 100-800s respectively. Kinda weird :) 🤔</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2296970,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/12/2023 09:35:53",
      "content": "<p><a href=\"https://www.kaggle.com/yk4r22\" target=\"_blank\">@yk4r22</a>, I am not sure if your search space for the optuna tuning step is appropriate. You could perhaps try a smaller search space and assess the impact. <br>\nAlso, do you need to tune so many parameters?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2297231,
          "author_name": "yk4r22",
          "author_url": "",
          "post_date": "06/12/2023 12:35:43",
          "content": "<p>I see, this is just the preparation, I'll decrease the search space in the future, thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2296896": "Hello 👀\n\ntrying to tune XGBoost with Optuna, it takes way less time than LightGBM (about 3-10 seconds vs 30-100 using LightGBM). Beforehands only for XGBoost I do vanilla preprocessing filling nans and replacing infs with median in numerics (otherwise it didn't start).\nThe score is also not too big comparing to LightGBM and CatBoost: 0.8485 (CB) vs 0.8493 (LightGBM) vs 0.8467 (XGB).\n\nHere are params I'm tuning:\n```\n    xgb_params = {\n        'lambda': trial.suggest_float('lambda', 1e-3, 1.0, log=True),\n        'alpha': trial.suggest_float('alpha', 1e-3, 1.0, log=True),\n        'colsample_bytree': trial.suggest_float('colsample_bytree', 0.05, 1.0),\n        'subsample': trial.suggest_float('subsample', 0.05, 1.0),\n        'learning_rate': trial.suggest_float(\"learning_rate\", 0.001, 0.05, log=True),\n        \"n_estimators\": trial.suggest_int(\"n_estimators\", 1000, 10000),\n        'max_depth': trial.suggest_int(\"max_depth\", 6, 12),\n        'num_round': trial.suggest_int(\"num_round\", 500, 10000),\n        'min_child_weight': trial.suggest_int('min_child_weight', 1, 300),\n        'use_label_encoder': trial.suggest_categorical('use_label_encoder', [True, False]),\n     \n        \"objective\": trial.suggest_categorical(\"objective\", [\"binary:logistic\"]),\n        \"verbosity\": 0,\n        \"boosting_type\": trial.suggest_categorical(\"boosting_type\", [\"gbdt\"]),\n        'random_state': trial.suggest_categorical('random_state', [42]),\n        'tree_method': trial.suggest_categorical('tree_method', ['hist']),\n    }\n```\nHere's my additional preprocessing:\n```\n    df_question = agg_dataframe.to_pandas().loc[:, features]\n\n    numerics = ['int16', 'int32', 'int64', 'float16', 'float32', 'float64']\n    num_cols = df_question.select_dtypes(include=numerics).columns\n    for num_col in num_cols:\n        median_col = df_question[num_col].median()\n        median_col = median_col if median_col not in [np.inf, -np.inf, np.nan] else 0\n        df_question[num_col] = df_question[num_col].replace([np.inf, -np.inf, np.nan], median_col)\n```\n\nWhat might be wrong?",
    "2296913": "What evaluation metric are you using when you mentioned 'score'? If you are using the same metric with the submission (which is Marco f1), then 0.848 is impossibly high compared with the current leaderboard. The most possible mistake is data leakage somewhere, which is more related to how you set your train/test dataset rather than tuning.",
    "2296926": "No, it's F1 just for the 1st question, and it wasn't submitted yet.",
    "2296940": "> No, it's F1 just for the 1st question, and it wasn't submitted yet.\n\nThen why do you think XGBoost is 'wrong'? 0.002 is a sensible shake when you using a different model.",
    "2296970": "yk4r22, I am not sure if your search space for the optuna tuning step is appropriate. You could perhaps try a smaller search space and assess the impact. \nAlso, do you need to tune so many parameters?",
    "2297229": "Because it takes 10 seconds for it to cross-validate over aggregates, while for CB and LightGBM it takes 30-100s and 100-800s respectively. Kinda weird :) 🤔",
    "2297231": "I see, this is just the preparation, I'll decrease the search space in the future, thank you!"
  },
  "source": "meta"
}