{
  "id": 400706,
  "title": "Let's talk about the quality of the model before blending",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/400706",
  "author_name": "",
  "post_date": "2023-04-09T21:16:17.557822600Z",
  "votes": 10,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I wonder how high of a quality can be achieved using a single model. In my case, using CatBoost, I was able to achieve 0.7 (almost 0.701).</p>\n<p>If it's not a secret, could you please share the result</p>",
  "messages": [
    {
      "id": "2216201",
      "postDate": "04/09/2023 21:16:17",
      "content": "<p>I wonder how high of a quality can be achieved using a single model. In my case, using CatBoost, I was able to achieve 0.7 (almost 0.701).</p>\n<p>If it's not a secret, could you please share the result</p>",
      "rawMarkdown": "I wonder how high of a quality can be achieved using a single model. In my case, using CatBoost, I was able to achieve 0.7 (almost 0.701).\n\nIf it's not a secret, could you please share the result",
      "votes": null
    },
    {
      "id": "2216267",
      "postDate": "04/09/2023 23:59:11",
      "content": "<p>my single model, is about 0.578</p>",
      "rawMarkdown": "my single model, is about 0.578",
      "votes": null
    },
    {
      "id": "2216305",
      "postDate": "04/10/2023 01:37:16",
      "content": "<p>single XGBoost, CV 0.696, LB 0.697</p>",
      "rawMarkdown": "single XGBoost, CV 0.696, LB 0.697",
      "votes": null
    },
    {
      "id": "2217331",
      "postDate": "04/10/2023 19:17:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mikhaildonskoy\" target=\"_blank\">@mikhaildonskoy</a> , I couldn't get a boost from blending. Blending simply gives me the average F1-score. My models are too correlated. </p>",
      "rawMarkdown": "Hi @mikhaildonskoy , I couldn't get a boost from blending. Blending simply gives me the average F1-score. My models are too correlated.",
      "votes": null
    },
    {
      "id": "2217362",
      "postDate": "04/10/2023 20:05:55",
      "content": "<p>single model:<br>\nCV: 0.69952 LB: 0.702<br>\nCV: 0.70025 LB: 0.703<br>\nCV: 0.70033 LB: 0.704</p>",
      "rawMarkdown": "single model:\nCV: 0.69952 LB: 0.702\nCV: 0.70025 LB: 0.703\nCV: 0.70033 LB: 0.704",
      "votes": null
    },
    {
      "id": "2217549",
      "postDate": "04/11/2023 02:20:17",
      "content": "<p>That's a good question. I think it's important to evaluate the quality of the single models before blending them, because blending can sometimes hide the weaknesses of the individual models or introduce new problems.</p>\n<p>I'm curious about your CatBoost model. What hyperparameters and features did you use? How did you deal with overfitting? Did you use any other models besides CatBoost?</p>",
      "rawMarkdown": "That's a good question. I think it's important to evaluate the quality of the single models before blending them, because blending can sometimes hide the weaknesses of the individual models or introduce new problems.\n\nI'm curious about your CatBoost model. What hyperparameters and features did you use? How did you deal with overfitting? Did you use any other models besides CatBoost?",
      "votes": null
    },
    {
      "id": "2217562",
      "postDate": "04/11/2023 02:54:32",
      "content": "<p>Thanks for sharing! Do you frequently see your LB higher than your CV? Does it happen before or after the new hidden test set, or both?</p>",
      "rawMarkdown": "Thanks for sharing! Do you frequently see your LB higher than your CV? Does it happen before or after the new hidden test set, or both?",
      "votes": null
    },
    {
      "id": "2217633",
      "postDate": "04/11/2023 04:24:00",
      "content": "<p>After the new hidden test set, LB is higher than CV.</p>",
      "rawMarkdown": "After the new hidden test set, LB is higher than CV.",
      "votes": null
    },
    {
      "id": "2217856",
      "postDate": "04/11/2023 08:09:26",
      "content": "<p>Thanks! My LB has been either equal to or lower than CV, even after the new hidden test set. I suppose I'm overfitting…</p>",
      "rawMarkdown": "Thanks! My LB has been either equal to or lower than CV, even after the new hidden test set. I suppose I'm overfitting...",
      "votes": null
    },
    {
      "id": "2217893",
      "postDate": "04/11/2023 08:47:22",
      "content": "<p>I reckon it's overfitting, too. Maybe you could check:</p>\n<blockquote>\n  <p>previous target probability features<br>\n  target encoding features<br>\n  too many features</p>\n</blockquote>\n<p>These are what I think may augur overfitting.</p>",
      "rawMarkdown": "I reckon it's overfitting, too. Maybe you could check:\n> previous target probability features\n> target encoding features\n> too many features\n\nThese are what I think may augur overfitting.",
      "votes": null
    },
    {
      "id": "2217949",
      "postDate": "04/11/2023 09:39:10",
      "content": "<p>Thanks for the suggestions. I have two questions</p>\n<ul>\n<li>Re <code>previous target probability features</code>, why do you think adding previous question's target probability (e.g. question 5's probability to data fitted in question 6) can cause overfitting?</li>\n<li>Would you mind clarifying what you mean by <code>target encoding features</code>?</li>\n</ul>",
      "rawMarkdown": "Thanks for the suggestions. I have two questions\n- Re `previous target probability features`, why do you think adding previous question's target probability (e.g. question 5's probability to data fitted in question 6) can cause overfitting?\n- Would you mind clarifying what you mean by `target encoding features`?",
      "votes": null
    },
    {
      "id": "2217981",
      "postDate": "04/11/2023 10:06:40",
      "content": "<p>I used over 1200 features with very strict regularization. I tried xgboost, but the result was worse</p>",
      "rawMarkdown": "I used over 1200 features with very strict regularization. I tried xgboost, but the result was worse",
      "votes": null
    },
    {
      "id": "2218098",
      "postDate": "04/11/2023 12:36:46",
      "content": "<p>I'm not sure <code>previous correct probability features</code> whether cause overfitting. But I think we can try both with it and not. If we have higher CV but much lower LB with it, it may overfit.<br>\n<code>target encoding features</code> means the average correct probability of categorical features, such as <code>the correct probability</code> of each <strong>month</strong>, each <strong>hour</strong>, or each <strong>group</strong> you cluster, etc.</p>",
      "rawMarkdown": "I'm not sure `previous correct probability features` whether cause overfitting. But I think we can try both with it and not. If we have higher CV but much lower LB with it, it may overfit.\n`target encoding features` means the average correct probability of categorical features, such as `the correct probability` of each **month**, each **hour**, or each **group** you cluster, etc.",
      "votes": null
    },
    {
      "id": "2218196",
      "postDate": "04/11/2023 13:49:54",
      "content": "<p>For regularization, do you use both <code>alpha</code> and <code>lambda</code> or only the former?</p>",
      "rawMarkdown": "For regularization, do you use both `alpha` and `lambda` or only the former?",
      "votes": null
    },
    {
      "id": "2218215",
      "postDate": "04/11/2023 13:58:34",
      "content": "<p>I used both options</p>",
      "rawMarkdown": "I used both options",
      "votes": null
    },
    {
      "id": "2218260",
      "postDate": "04/11/2023 14:28:48",
      "content": "<p>Mine is about 0.698 *Xgboost </p>",
      "rawMarkdown": "Mine is about 0.698 *Xgboost",
      "votes": null
    },
    {
      "id": "2218940",
      "postDate": "04/12/2023 06:34:44",
      "content": "<p>How to chose the value between <code>alpha</code> and <code>lambda</code> ?</p>",
      "rawMarkdown": "How to chose the value between `alpha` and `lambda` ?",
      "votes": null
    },
    {
      "id": "2219360",
      "postDate": "04/12/2023 14:12:48",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/takanashihumbert\" target=\"_blank\">@takanashihumbert</a>. Could you tell me the fold number you used to get cv score?<br>\nI think the gap between cv and lb also depends on the folds because the number of data that is available while training changes.<br>\nThank you!</p>",
      "rawMarkdown": "Hi, @takanashihumbert. Could you tell me the fold number you used to get cv score?\nI think the gap between cv and lb also depends on the folds because the number of data that is available while training changes.\nThank you!",
      "votes": null
    },
    {
      "id": "2219362",
      "postDate": "04/12/2023 14:17:23",
      "content": "<p>xgboost, cv: 0.698 </p>",
      "rawMarkdown": "xgboost, cv: 0.698",
      "votes": null
    },
    {
      "id": "2219576",
      "postDate": "04/12/2023 17:13:30",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/nynyny67\" target=\"_blank\">@nynyny67</a>, 5 folds for cv.</p>",
      "rawMarkdown": "Hi @nynyny67, 5 folds for cv.",
      "votes": null
    },
    {
      "id": "2219890",
      "postDate": "04/13/2023 01:18:03",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2220238",
      "postDate": "04/13/2023 08:00:57",
      "content": "<p>Mine XGB: 0.698, hit a wall after using all 2000+ features. Am I missing something?</p>",
      "rawMarkdown": "Mine XGB: 0.698, hit a wall after using all 2000+ features. Am I missing something?",
      "votes": null
    },
    {
      "id": "2220444",
      "postDate": "04/13/2023 11:52:52",
      "content": "<p>Same. I guess it's time for feature pruning &amp; hyper-parameter tuning</p>",
      "rawMarkdown": "Same. I guess it's time for feature pruning & hyper-parameter tuning",
      "votes": null
    },
    {
      "id": "2220447",
      "postDate": "04/13/2023 11:57:04",
      "content": "<p><a href=\"https://www.kaggle.com/takanashihumbert\" target=\"_blank\">@takanashihumbert</a> if you don't mind sharing, do you use all 5 output model instances from the 5 folds for inference, or only one of the 5 folds, or do you refit and use the refitted one? If you've tried multiple methods, do you see a large gap between them? I usually saw a 0.001 or less.</p>",
      "rawMarkdown": "takanashihumbert if you don't mind sharing, do you use all 5 output model instances from the 5 folds for inference, or only one of the 5 folds, or do you refit and use the refitted one? If you've tried multiple methods, do you see a large gap between them? I usually saw a 0.001 or less.",
      "votes": null
    },
    {
      "id": "2220458",
      "postDate": "04/13/2023 12:10:45",
      "content": "<p>Yes, I have setup OPTUNA for hyperparameter tuning will be sharing the code soon</p>",
      "rawMarkdown": "Yes, I have setup OPTUNA for hyperparameter tuning will be sharing the code soon",
      "votes": null
    },
    {
      "id": "2220479",
      "postDate": "04/13/2023 12:29:20",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>, we use 5 folds(5 models) and then get averaged output.</p>",
      "rawMarkdown": "Hi @hoangnguyen719, we use 5 folds(5 models) and then get averaged output.",
      "votes": null
    },
    {
      "id": "2221595",
      "postDate": "04/14/2023 11:53:01",
      "content": "<p>Single XGB with 2k+ features LB:0.697, but very very slow. Inference time is ~9h where I use pre-trained model.</p>",
      "rawMarkdown": "Single XGB with 2k+ features LB:0.697, but very very slow. Inference time is ~9h where I use pre-trained model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2216267,
      "author_name": "tyeestudio",
      "author_url": "",
      "post_date": "04/09/2023 23:59:11",
      "content": "<p>my single model, is about 0.578</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2216305,
      "author_name": "dianxu",
      "author_url": "",
      "post_date": "04/10/2023 01:37:16",
      "content": "<p>single XGBoost, CV 0.696, LB 0.697</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2217331,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "04/10/2023 19:17:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mikhaildonskoy\" target=\"_blank\">@mikhaildonskoy</a> , I couldn't get a boost from blending. Blending simply gives me the average F1-score. My models are too correlated. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2217362,
      "author_name": "takanashihumbert",
      "author_url": "",
      "post_date": "04/10/2023 20:05:55",
      "content": "<p>single model:<br>\nCV: 0.69952 LB: 0.702<br>\nCV: 0.70025 LB: 0.703<br>\nCV: 0.70033 LB: 0.704</p>",
      "votes": null,
      "replies": [
        {
          "id": 2217562,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "04/11/2023 02:54:32",
          "content": "<p>Thanks for sharing! Do you frequently see your LB higher than your CV? Does it happen before or after the new hidden test set, or both?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2217633,
              "author_name": "takanashihumbert",
              "author_url": "",
              "post_date": "04/11/2023 04:24:00",
              "content": "<p>After the new hidden test set, LB is higher than CV.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2217856,
                  "author_name": "hoangnguyen719",
                  "author_url": "",
                  "post_date": "04/11/2023 08:09:26",
                  "content": "<p>Thanks! My LB has been either equal to or lower than CV, even after the new hidden test set. I suppose I'm overfitting…</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2217893,
                      "author_name": "takanashihumbert",
                      "author_url": "",
                      "post_date": "04/11/2023 08:47:22",
                      "content": "<p>I reckon it's overfitting, too. Maybe you could check:</p>\n<blockquote>\n  <p>previous target probability features<br>\n  target encoding features<br>\n  too many features</p>\n</blockquote>\n<p>These are what I think may augur overfitting.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2217949,
                          "author_name": "hoangnguyen719",
                          "author_url": "",
                          "post_date": "04/11/2023 09:39:10",
                          "content": "<p>Thanks for the suggestions. I have two questions</p>\n<ul>\n<li>Re <code>previous target probability features</code>, why do you think adding previous question's target probability (e.g. question 5's probability to data fitted in question 6) can cause overfitting?</li>\n<li>Would you mind clarifying what you mean by <code>target encoding features</code>?</li>\n</ul>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2218098,
                              "author_name": "takanashihumbert",
                              "author_url": "",
                              "post_date": "04/11/2023 12:36:46",
                              "content": "<p>I'm not sure <code>previous correct probability features</code> whether cause overfitting. But I think we can try both with it and not. If we have higher CV but much lower LB with it, it may overfit.<br>\n<code>target encoding features</code> means the average correct probability of categorical features, such as <code>the correct probability</code> of each <strong>month</strong>, each <strong>hour</strong>, or each <strong>group</strong> you cluster, etc.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2219360,
          "author_name": "nynyny67",
          "author_url": "",
          "post_date": "04/12/2023 14:12:48",
          "content": "<p>Hi, <a href=\"https://www.kaggle.com/takanashihumbert\" target=\"_blank\">@takanashihumbert</a>. Could you tell me the fold number you used to get cv score?<br>\nI think the gap between cv and lb also depends on the folds because the number of data that is available while training changes.<br>\nThank you!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2219576,
              "author_name": "takanashihumbert",
              "author_url": "",
              "post_date": "04/12/2023 17:13:30",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/nynyny67\" target=\"_blank\">@nynyny67</a>, 5 folds for cv.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2219890,
                  "author_name": "nynyny67",
                  "author_url": "",
                  "post_date": "04/13/2023 01:18:03",
                  "content": "<p>Thanks for sharing!</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2220447,
                  "author_name": "hoangnguyen719",
                  "author_url": "",
                  "post_date": "04/13/2023 11:57:04",
                  "content": "<p><a href=\"https://www.kaggle.com/takanashihumbert\" target=\"_blank\">@takanashihumbert</a> if you don't mind sharing, do you use all 5 output model instances from the 5 folds for inference, or only one of the 5 folds, or do you refit and use the refitted one? If you've tried multiple methods, do you see a large gap between them? I usually saw a 0.001 or less.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2220479,
                      "author_name": "takanashihumbert",
                      "author_url": "",
                      "post_date": "04/13/2023 12:29:20",
                      "content": "<p>Hi <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>, we use 5 folds(5 models) and then get averaged output.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2217549,
      "author_name": "yus002",
      "author_url": "",
      "post_date": "04/11/2023 02:20:17",
      "content": "<p>That's a good question. I think it's important to evaluate the quality of the single models before blending them, because blending can sometimes hide the weaknesses of the individual models or introduce new problems.</p>\n<p>I'm curious about your CatBoost model. What hyperparameters and features did you use? How did you deal with overfitting? Did you use any other models besides CatBoost?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2217981,
          "author_name": "mikhaildonskoy",
          "author_url": "",
          "post_date": "04/11/2023 10:06:40",
          "content": "<p>I used over 1200 features with very strict regularization. I tried xgboost, but the result was worse</p>",
          "votes": null,
          "replies": [
            {
              "id": 2218196,
              "author_name": "hoangnguyen719",
              "author_url": "",
              "post_date": "04/11/2023 13:49:54",
              "content": "<p>For regularization, do you use both <code>alpha</code> and <code>lambda</code> or only the former?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2218215,
                  "author_name": "mikhaildonskoy",
                  "author_url": "",
                  "post_date": "04/11/2023 13:58:34",
                  "content": "<p>I used both options</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2218940,
                  "author_name": "thitiwat",
                  "author_url": "",
                  "post_date": "04/12/2023 06:34:44",
                  "content": "<p>How to chose the value between <code>alpha</code> and <code>lambda</code> ?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2218260,
      "author_name": "tumpanjawat",
      "author_url": "",
      "post_date": "04/11/2023 14:28:48",
      "content": "<p>Mine is about 0.698 *Xgboost </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2219362,
      "author_name": "nynyny67",
      "author_url": "",
      "post_date": "04/12/2023 14:17:23",
      "content": "<p>xgboost, cv: 0.698 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2220238,
      "author_name": "zahidfaiz",
      "author_url": "",
      "post_date": "04/13/2023 08:00:57",
      "content": "<p>Mine XGB: 0.698, hit a wall after using all 2000+ features. Am I missing something?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2220444,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "04/13/2023 11:52:52",
          "content": "<p>Same. I guess it's time for feature pruning &amp; hyper-parameter tuning</p>",
          "votes": null,
          "replies": [
            {
              "id": 2220458,
              "author_name": "zahidfaiz",
              "author_url": "",
              "post_date": "04/13/2023 12:10:45",
              "content": "<p>Yes, I have setup OPTUNA for hyperparameter tuning will be sharing the code soon</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2221595,
      "author_name": "ngocuong",
      "author_url": "",
      "post_date": "04/14/2023 11:53:01",
      "content": "<p>Single XGB with 2k+ features LB:0.697, but very very slow. Inference time is ~9h where I use pre-trained model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2216201": "I wonder how high of a quality can be achieved using a single model. In my case, using CatBoost, I was able to achieve 0.7 (almost 0.701).\n\nIf it's not a secret, could you please share the result",
    "2216267": "my single model, is about 0.578",
    "2216305": "single XGBoost, CV 0.696, LB 0.697",
    "2217331": "Hi @mikhaildonskoy , I couldn't get a boost from blending. Blending simply gives me the average F1-score. My models are too correlated.",
    "2217362": "single model:\nCV: 0.69952 LB: 0.702\nCV: 0.70025 LB: 0.703\nCV: 0.70033 LB: 0.704",
    "2217549": "That's a good question. I think it's important to evaluate the quality of the single models before blending them, because blending can sometimes hide the weaknesses of the individual models or introduce new problems.\n\nI'm curious about your CatBoost model. What hyperparameters and features did you use? How did you deal with overfitting? Did you use any other models besides CatBoost?",
    "2217562": "Thanks for sharing! Do you frequently see your LB higher than your CV? Does it happen before or after the new hidden test set, or both?",
    "2217633": "After the new hidden test set, LB is higher than CV.",
    "2217856": "Thanks! My LB has been either equal to or lower than CV, even after the new hidden test set. I suppose I'm overfitting...",
    "2217893": "I reckon it's overfitting, too. Maybe you could check:\n> previous target probability features\n> target encoding features\n> too many features\n\nThese are what I think may augur overfitting.",
    "2217949": "Thanks for the suggestions. I have two questions\n- Re `previous target probability features`, why do you think adding previous question's target probability (e.g. question 5's probability to data fitted in question 6) can cause overfitting?\n- Would you mind clarifying what you mean by `target encoding features`?",
    "2217981": "I used over 1200 features with very strict regularization. I tried xgboost, but the result was worse",
    "2218098": "I'm not sure `previous correct probability features` whether cause overfitting. But I think we can try both with it and not. If we have higher CV but much lower LB with it, it may overfit.\n`target encoding features` means the average correct probability of categorical features, such as `the correct probability` of each **month**, each **hour**, or each **group** you cluster, etc.",
    "2218196": "For regularization, do you use both `alpha` and `lambda` or only the former?",
    "2218215": "I used both options",
    "2218260": "Mine is about 0.698 *Xgboost",
    "2218940": "How to chose the value between `alpha` and `lambda` ?",
    "2219360": "Hi, @takanashihumbert. Could you tell me the fold number you used to get cv score?\nI think the gap between cv and lb also depends on the folds because the number of data that is available while training changes.\nThank you!",
    "2219362": "xgboost, cv: 0.698",
    "2219576": "Hi @nynyny67, 5 folds for cv.",
    "2219890": "Thanks for sharing!",
    "2220238": "Mine XGB: 0.698, hit a wall after using all 2000+ features. Am I missing something?",
    "2220444": "Same. I guess it's time for feature pruning & hyper-parameter tuning",
    "2220447": "takanashihumbert if you don't mind sharing, do you use all 5 output model instances from the 5 folds for inference, or only one of the 5 folds, or do you refit and use the refitted one? If you've tried multiple methods, do you see a large gap between them? I usually saw a 0.001 or less.",
    "2220458": "Yes, I have setup OPTUNA for hyperparameter tuning will be sharing the code soon",
    "2220479": "Hi @hoangnguyen719, we use 5 folds(5 models) and then get averaged output.",
    "2221595": "Single XGB with 2k+ features LB:0.697, but very very slow. Inference time is ~9h where I use pre-trained model."
  },
  "source": "meta"
}