{
  "id": 546396,
  "title": "Do we really need complex models here?",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/546396",
  "author_name": "",
  "post_date": "2024-11-15T13:54:42.467660500Z",
  "votes": 21,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I presume a lot of us are highly inspired by a lot of good public work in this competition and are perhaps aligning our approaches in line with these kernels to some extent at least. I wish to discuss a few points I observed in these kernels and wish to seek your inputs on the same- </p>\n<ol>\n<li>Models like boosted trees, auto-encoder, tabnet regressor necessitate larger data sizes while the concurrent data is quite small. Up on slicing the data into CV folds, the data becomes even smaller, based on the number of folds one wishes to use. In such cases, do you think such complex approaches will work?</li>\n<li>The dataset is inherently very highly noisy and a lot of the tabular features are not highly predictive. Do you think using a NN-like architecture will overfit to the noise in the data in such cases and create an unwanted churn here?</li>\n<li>Do you have a reasonable CV-LB relation here? </li>\n</ol>\n<p>Thoughts? Comments?</p>",
  "messages": [
    {
      "id": "3046469",
      "postDate": "11/15/2024 13:54:42",
      "content": "<p>Hello all,</p>\n<p>I presume a lot of us are highly inspired by a lot of good public work in this competition and are perhaps aligning our approaches in line with these kernels to some extent at least. I wish to discuss a few points I observed in these kernels and wish to seek your inputs on the same- </p>\n<ol>\n<li>Models like boosted trees, auto-encoder, tabnet regressor necessitate larger data sizes while the concurrent data is quite small. Up on slicing the data into CV folds, the data becomes even smaller, based on the number of folds one wishes to use. In such cases, do you think such complex approaches will work?</li>\n<li>The dataset is inherently very highly noisy and a lot of the tabular features are not highly predictive. Do you think using a NN-like architecture will overfit to the noise in the data in such cases and create an unwanted churn here?</li>\n<li>Do you have a reasonable CV-LB relation here? </li>\n</ol>\n<p>Thoughts? Comments?</p>",
      "rawMarkdown": "Hello all,\n\nI presume a lot of us are highly inspired by a lot of good public work in this competition and are perhaps aligning our approaches in line with these kernels to some extent at least. I wish to discuss a few points I observed in these kernels and wish to seek your inputs on the same- \n\n1. Models like boosted trees, auto-encoder, tabnet regressor necessitate larger data sizes while the concurrent data is quite small. Up on slicing the data into CV folds, the data becomes even smaller, based on the number of folds one wishes to use. In such cases, do you think such complex approaches will work?\n2. The dataset is inherently very highly noisy and a lot of the tabular features are not highly predictive. Do you think using a NN-like architecture will overfit to the noise in the data in such cases and create an unwanted churn here?\n3. Do you have a reasonable CV-LB relation here? \n\nThoughts? Comments?",
      "votes": null
    },
    {
      "id": "3046749",
      "postDate": "11/15/2024 20:34:44",
      "content": "<p>3- What's a reasonable/good CV-LB relation? My best score currently is CV 0.475 - LB 0.469.<br>\nI think that's reasonable but also not very good, given the placement. The good relation is probably due to the fact, that I didnt use any imputation so far and only use a single modle (LightGBM) with feature engineering and data cleaning. I tried implementing the things mentioned in \"1.\" to little success so far. But maybe I have to try everything at once and not step by step. But then I will probably end up at the same notebooks, which are already public…</p>\n<p>I also made a post about the common notebooks, which use 3 models. I can't quite follow, why the models are chosen this way.</p>",
      "rawMarkdown": "3- What's a reasonable/good CV-LB relation? My best score currently is CV 0.475 - LB 0.469.\nI think that's reasonable but also not very good, given the placement. The good relation is probably due to the fact, that I didnt use any imputation so far and only use a single modle (LightGBM) with feature engineering and data cleaning. I tried implementing the things mentioned in \"1.\" to little success so far. But maybe I have to try everything at once and not step by step. But then I will probably end up at the same notebooks, which are already public...\n\nI also made a post about the common notebooks, which use 3 models. I can't quite follow, why the models are chosen this way.",
      "votes": null
    },
    {
      "id": "3046756",
      "postDate": "11/15/2024 20:39:41",
      "content": "<p><a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a> thanks for the rejoinder. I think we are in for a big shakeup as many public kernel elements are not sustainable. </p>",
      "rawMarkdown": "mariusheuser thanks for the rejoinder. I think we are in for a big shakeup as many public kernel elements are not sustainable.",
      "votes": null
    },
    {
      "id": "3046927",
      "postDate": "11/16/2024 04:00:51",
      "content": "<p>I agree with your thoughts and in addition to the 3 models, it can be noticed that KNN Imputer has been applied with target variable SII which will lead to data leakage. Normally KNN Imputer shouldn't be applied with the target variable so, some how that wrong use of KNN Imputer is working well with 38% of test data and pushing the LB score high. In addition to this I read some people in discussion said that auto encoder has also been applied incorrectly, I didn't got much time to dig down in the auto encoder implementation of those notebooks. If you know what mistake is in them then do share, thank you.<br>\nOverall we can say if these mistakes are leading to a better score with 38% of data then we can't say how it will perform with remaining 62% of data. If that 62% is almost similar to this 38% that results could be similar but on the other hand if that 62% is different then there can be a significant shakeup. Let's see what happens. </p>",
      "rawMarkdown": "I agree with your thoughts and in addition to the 3 models, it can be noticed that KNN Imputer has been applied with target variable SII which will lead to data leakage. Normally KNN Imputer shouldn't be applied with the target variable so, some how that wrong use of KNN Imputer is working well with 38% of test data and pushing the LB score high. In addition to this I read some people in discussion said that auto encoder has also been applied incorrectly, I didn't got much time to dig down in the auto encoder implementation of those notebooks. If you know what mistake is in them then do share, thank you.\nOverall we can say if these mistakes are leading to a better score with 38% of data then we can't say how it will perform with remaining 62% of data. If that 62% is almost similar to this 38% that results could be similar but on the other hand if that 62% is different then there can be a significant shakeup. Let's see what happens.",
      "votes": null
    },
    {
      "id": "3047047",
      "postDate": "11/16/2024 07:23:03",
      "content": "<blockquote>\n  <p>Do you have a reasonable CV-LB relation here?</p>\n</blockquote>\n<p><strong>for me at least, no, <br>\nnot relating by any mean.</strong></p>",
      "rawMarkdown": "> Do you have a reasonable CV-LB relation here?\n\n **for me at least, no, <br>\nnot relating by any mean.**",
      "votes": null
    },
    {
      "id": "3047075",
      "postDate": "11/16/2024 08:12:27",
      "content": "<p>Same here - no alignment is an indicator of a shakeup <a href=\"https://www.kaggle.com/letemoin\" target=\"_blank\">@letemoin</a> </p>",
      "rawMarkdown": "Same here - no alignment is an indicator of a shakeup @letemoin",
      "votes": null
    },
    {
      "id": "3047381",
      "postDate": "11/16/2024 15:58:29",
      "content": "<p>Totally agree. According to Kaggle competitions, there are often inconsistencies between the distribution of public and private LBs, so overfitting to the public LB usually does not lead to better (or even worse) results. For this reason, CV is usually more trustworthy and robust than LB.</p>",
      "rawMarkdown": "Totally agree. According to Kaggle competitions, there are often inconsistencies between the distribution of public and private LBs, so overfitting to the public LB usually does not lead to better (or even worse) results. For this reason, CV is usually more trustworthy and robust than LB.",
      "votes": null
    },
    {
      "id": "3047453",
      "postDate": "11/16/2024 17:34:49",
      "content": "<p>Agree with you <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "Agree with you @ravi20076",
      "votes": null
    },
    {
      "id": "3047475",
      "postDate": "11/16/2024 18:09:04",
      "content": "<p>negative values for BIA-BIA_BMC, BIA-BIA_FMI, BIA-BIA_Fat! Could it be due to an error or a calibration issue?</p>",
      "rawMarkdown": "negative values for BIA-BIA_BMC, BIA-BIA_FMI, BIA-BIA_Fat! Could it be due to an error or a calibration issue?",
      "votes": null
    },
    {
      "id": "3047482",
      "postDate": "11/16/2024 18:16:23",
      "content": "<p>I agree with you on this point.</p>",
      "rawMarkdown": "I agree with you on this point.",
      "votes": null
    },
    {
      "id": "3047549",
      "postDate": "11/16/2024 20:20:30",
      "content": "<p>Probably it's a better idea to follow the simplicity path rather then the complexity path,There's alot of risk to overfit here and a high LB score doesn't look like it reflects how robust a model is.</p>",
      "rawMarkdown": "Probably it's a better idea to follow the simplicity path rather then the complexity path,There's alot of risk to overfit here and a high LB score doesn't look like it reflects how robust a model is.",
      "votes": null
    },
    {
      "id": "3047692",
      "postDate": "11/17/2024 04:12:14",
      "content": "<p>I think you are right</p>",
      "rawMarkdown": "I think you are right",
      "votes": null
    },
    {
      "id": "3048319",
      "postDate": "11/17/2024 18:41:03",
      "content": "<p>For me, there are some reasonable CV-LB relation points, but they lie between 0.445 and 0.465 on the public leaderboard. However, I still believe we are just overfitting to the public leaderboard. I haven't seen any reproducible work among the high scores on the public leaderboard. <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "For me, there are some reasonable CV-LB relation points, but they lie between 0.445 and 0.465 on the public leaderboard. However, I still believe we are just overfitting to the public leaderboard. I haven't seen any reproducible work among the high scores on the public leaderboard. @ravi20076",
      "votes": null
    },
    {
      "id": "3048322",
      "postDate": "11/17/2024 18:45:40",
      "content": "<p>Yes, I tried experimenting with this. After seeing many public solutions, I noticed that they impute the target variable using KNN. After doing that, both the LB score and CV score improve, but it is clearly a case of data leakage/Overfitting.</p>\n<p>The Main Problem is Copy-Paste , People Didn't Read Solutions , They Start Copying it Without Understanding it. Let See What Happen In the Private !</p>",
      "rawMarkdown": "Yes, I tried experimenting with this. After seeing many public solutions, I noticed that they impute the target variable using KNN. After doing that, both the LB score and CV score improve, but it is clearly a case of data leakage/Overfitting.\n\nThe Main Problem is Copy-Paste , People Didn't Read Solutions , They Start Copying it Without Understanding it. Let See What Happen In the Private !",
      "votes": null
    },
    {
      "id": "3048406",
      "postDate": "11/17/2024 21:11:08",
      "content": "<p>True, same here - scores between 0.45 and 0.46 are correlating with CV. Anything more than that is all off-place <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> </p>",
      "rawMarkdown": "True, same here - scores between 0.45 and 0.46 are correlating with CV. Anything more than that is all off-place @abdmental01",
      "votes": null
    },
    {
      "id": "3048505",
      "postDate": "11/18/2024 03:30:12",
      "content": "<p>Yes exactly, that leads to data leakage/over-fitting. The only thing that surprises me is that this overfitting is working well with 38% of data. That is more than 1/4th of the private dataset. </p>",
      "rawMarkdown": "Yes exactly, that leads to data leakage/over-fitting. The only thing that surprises me is that this overfitting is working well with 38% of data. That is more than 1/4th of the private dataset.",
      "votes": null
    },
    {
      "id": "3048599",
      "postDate": "11/18/2024 06:32:46",
      "content": "<p>My two cents:</p>\n<ol>\n<li>With highly noisy data like this one, simple models with good ML basics would work, leading to CV-LB relation over 0.44-0.46. Letting boosted trees to handle missing values and preprocessing is nice, and fitting is better compared to NN, but I don't think it makes any sense to spend too much time on tuning the parameters.</li>\n<li>The NN models (mostly MLPs) I tried are worse compared to LGBMs. But I think they correlate better to LB (less overfitting but lower LB score) and I know exactly when I'm overfitting.</li>\n<li>With NNs, it's over 0.42-0.44 range. With LGBMs no.</li>\n</ol>\n<p>BTW, this comp is ICR all over again. Expect shakeup at the end.</p>",
      "rawMarkdown": "My two cents:\n\n1. With highly noisy data like this one, simple models with good ML basics would work, leading to CV-LB relation over 0.44-0.46. Letting boosted trees to handle missing values and preprocessing is nice, and fitting is better compared to NN, but I don't think it makes any sense to spend too much time on tuning the parameters.\n2. The NN models (mostly MLPs) I tried are worse compared to LGBMs. But I think they correlate better to LB (less overfitting but lower LB score) and I know exactly when I'm overfitting.\n3. With NNs, it's over 0.42-0.44 range. With LGBMs no.\n\nBTW, this comp is ICR all over again. Expect shakeup at the end.",
      "votes": null
    },
    {
      "id": "3048994",
      "postDate": "11/18/2024 15:33:33",
      "content": "<p>I think this is ICR competition (# ﾟДﾟ)(# ﾟДﾟ)</p>",
      "rawMarkdown": "I think this is ICR competition (# ﾟДﾟ)(# ﾟДﾟ)",
      "votes": null
    },
    {
      "id": "3049049",
      "postDate": "11/18/2024 16:29:39",
      "content": "<p>To be honest I've almost given up in this competition, which is very sad to me, after spending over 150 hours in the first month.<br>\nI had a bad strategy in this game, and pushed too hard that I burnout too quickly.<br>\nMe being a beginner in ML, I guess the tech you mention won't work in this competition. If someone is courageous enough, maybe they should drop some features. I can't recall it anymore, but in one test in this game, I dropped quite many columns and the score didn't drop.<br>\nI guess I'll submit my 0.472 as final submission, basically a LGBM and a bit of data cleaning + featuring. I don't even bother blending it with catboost + xgboost anymore.<br>\nThis is my first participation in competition with prize and to be honest, I enjoyed much more in the playground game more. I should have manged my expectation better on Day 1.</p>\n<p>Good luck and may I wish everyone all the best!</p>",
      "rawMarkdown": "To be honest I've almost given up in this competition, which is very sad to me, after spending over 150 hours in the first month.\nI had a bad strategy in this game, and pushed too hard that I burnout too quickly.\nMe being a beginner in ML, I guess the tech you mention won't work in this competition. If someone is courageous enough, maybe they should drop some features. I can't recall it anymore, but in one test in this game, I dropped quite many columns and the score didn't drop.\nI guess I'll submit my 0.472 as final submission, basically a LGBM and a bit of data cleaning + featuring. I don't even bother blending it with catboost + xgboost anymore.\nThis is my first participation in competition with prize and to be honest, I enjoyed much more in the playground game more. I should have manged my expectation better on Day 1.\n\nGood luck and may I wish everyone all the best!",
      "votes": null
    },
    {
      "id": "3049702",
      "postDate": "11/19/2024 12:31:44",
      "content": "<p>I also feel so, ICR part 2 <a href=\"https://www.kaggle.com/shin9915\" target=\"_blank\">@shin9915</a> </p>",
      "rawMarkdown": "I also feel so, ICR part 2 @shin9915",
      "votes": null
    },
    {
      "id": "3050492",
      "postDate": "11/20/2024 09:05:10",
      "content": "<p>Makes sense! In fact, one of my 0.495 scores was achieved purely using machine learning without incorporating complex models like TabNet, and the result turned out to be better than when I included it.</p>",
      "rawMarkdown": "Makes sense! In fact, one of my 0.495 scores was achieved purely using machine learning without incorporating complex models like TabNet, and the result turned out to be better than when I included it.",
      "votes": null
    },
    {
      "id": "3051758",
      "postDate": "11/21/2024 15:57:28",
      "content": "<p>I treated this competition as a regression problem; and my score was 0.366. Funny thing is that I did not even include data from the parquet files! I messed up, and never returned back to continue in this competition. It's been 2 months since my last submission haha</p>",
      "rawMarkdown": "I treated this competition as a regression problem; and my score was 0.366. Funny thing is that I did not even include data from the parquet files! I messed up, and never returned back to continue in this competition. It's been 2 months since my last submission haha",
      "votes": null
    },
    {
      "id": "3053393",
      "postDate": "11/23/2024 12:44:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/taimour\" target=\"_blank\">@taimour</a> ,</p>\n<p>About <code>auto encoder has also been applied incorrectly</code> : the autoencoder is trained on train timeseries and trained again on test timeseries, Instead of using a <code>transform</code> or <code>predict</code> method for test dataset.</p>",
      "rawMarkdown": "Hi @taimour ,\n\nAbout ``auto encoder has also been applied incorrectly`` : the autoencoder is trained on train timeseries and trained again on test timeseries, Instead of using a ``transform`` or ``predict`` method for test dataset.",
      "votes": null
    },
    {
      "id": "3053436",
      "postDate": "11/23/2024 13:35:16",
      "content": "<p>Thank you for details.</p>",
      "rawMarkdown": "Thank you for details.",
      "votes": null
    },
    {
      "id": "3057191",
      "postDate": "11/27/2024 20:34:45",
      "content": "<p>I tried out a simple CatBoost and was able to get 0.372. But the challenge here is to deal with both questionaire as well as the continuous time series data</p>",
      "rawMarkdown": "I tried out a simple CatBoost and was able to get 0.372. But the challenge here is to deal with both questionaire as well as the continuous time series data",
      "votes": null
    },
    {
      "id": "3061512",
      "postDate": "12/02/2024 17:50:10",
      "content": "<p>That's great! I think your direction is right, and I'm really looking forward to further exchanging ideas with you. I'm also trying to improve the model's score using simple models and appropriate data preprocessing methods.</p>",
      "rawMarkdown": "That's great! I think your direction is right, and I'm really looking forward to further exchanging ideas with you. I'm also trying to improve the model's score using simple models and appropriate data preprocessing methods.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3046749,
      "author_name": "mariusheuser",
      "author_url": "",
      "post_date": "11/15/2024 20:34:44",
      "content": "<p>3- What's a reasonable/good CV-LB relation? My best score currently is CV 0.475 - LB 0.469.<br>\nI think that's reasonable but also not very good, given the placement. The good relation is probably due to the fact, that I didnt use any imputation so far and only use a single modle (LightGBM) with feature engineering and data cleaning. I tried implementing the things mentioned in \"1.\" to little success so far. But maybe I have to try everything at once and not step by step. But then I will probably end up at the same notebooks, which are already public…</p>\n<p>I also made a post about the common notebooks, which use 3 models. I can't quite follow, why the models are chosen this way.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3046756,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "11/15/2024 20:39:41",
          "content": "<p><a href=\"https://www.kaggle.com/mariusheuser\" target=\"_blank\">@mariusheuser</a> thanks for the rejoinder. I think we are in for a big shakeup as many public kernel elements are not sustainable. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3047453,
              "author_name": "nancyalaswad90",
              "author_url": "",
              "post_date": "11/16/2024 17:34:49",
              "content": "<p>Agree with you <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3046927,
          "author_name": "taimour",
          "author_url": "",
          "post_date": "11/16/2024 04:00:51",
          "content": "<p>I agree with your thoughts and in addition to the 3 models, it can be noticed that KNN Imputer has been applied with target variable SII which will lead to data leakage. Normally KNN Imputer shouldn't be applied with the target variable so, some how that wrong use of KNN Imputer is working well with 38% of test data and pushing the LB score high. In addition to this I read some people in discussion said that auto encoder has also been applied incorrectly, I didn't got much time to dig down in the auto encoder implementation of those notebooks. If you know what mistake is in them then do share, thank you.<br>\nOverall we can say if these mistakes are leading to a better score with 38% of data then we can't say how it will perform with remaining 62% of data. If that 62% is almost similar to this 38% that results could be similar but on the other hand if that 62% is different then there can be a significant shakeup. Let's see what happens. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3047381,
              "author_name": "fangzitao",
              "author_url": "",
              "post_date": "11/16/2024 15:58:29",
              "content": "<p>Totally agree. According to Kaggle competitions, there are often inconsistencies between the distribution of public and private LBs, so overfitting to the public LB usually does not lead to better (or even worse) results. For this reason, CV is usually more trustworthy and robust than LB.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3047482,
                  "author_name": "taimour",
                  "author_url": "",
                  "post_date": "11/16/2024 18:16:23",
                  "content": "<p>I agree with you on this point.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3048322,
              "author_name": "abdmental01",
              "author_url": "",
              "post_date": "11/17/2024 18:45:40",
              "content": "<p>Yes, I tried experimenting with this. After seeing many public solutions, I noticed that they impute the target variable using KNN. After doing that, both the LB score and CV score improve, but it is clearly a case of data leakage/Overfitting.</p>\n<p>The Main Problem is Copy-Paste , People Didn't Read Solutions , They Start Copying it Without Understanding it. Let See What Happen In the Private !</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3048505,
                  "author_name": "taimour",
                  "author_url": "",
                  "post_date": "11/18/2024 03:30:12",
                  "content": "<p>Yes exactly, that leads to data leakage/over-fitting. The only thing that surprises me is that this overfitting is working well with 38% of data. That is more than 1/4th of the private dataset. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3053393,
              "author_name": "adaubas",
              "author_url": "",
              "post_date": "11/23/2024 12:44:40",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/taimour\" target=\"_blank\">@taimour</a> ,</p>\n<p>About <code>auto encoder has also been applied incorrectly</code> : the autoencoder is trained on train timeseries and trained again on test timeseries, Instead of using a <code>transform</code> or <code>predict</code> method for test dataset.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3053436,
                  "author_name": "taimour",
                  "author_url": "",
                  "post_date": "11/23/2024 13:35:16",
                  "content": "<p>Thank you for details.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3047047,
      "author_name": "letemoin",
      "author_url": "",
      "post_date": "11/16/2024 07:23:03",
      "content": "<blockquote>\n  <p>Do you have a reasonable CV-LB relation here?</p>\n</blockquote>\n<p><strong>for me at least, no, <br>\nnot relating by any mean.</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 3047075,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "11/16/2024 08:12:27",
          "content": "<p>Same here - no alignment is an indicator of a shakeup <a href=\"https://www.kaggle.com/letemoin\" target=\"_blank\">@letemoin</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3047475,
      "author_name": "jamalsaeedi",
      "author_url": "",
      "post_date": "11/16/2024 18:09:04",
      "content": "<p>negative values for BIA-BIA_BMC, BIA-BIA_FMI, BIA-BIA_Fat! Could it be due to an error or a calibration issue?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3047549,
      "author_name": "mohammedahmedxx12",
      "author_url": "",
      "post_date": "11/16/2024 20:20:30",
      "content": "<p>Probably it's a better idea to follow the simplicity path rather then the complexity path,There's alot of risk to overfit here and a high LB score doesn't look like it reflects how robust a model is.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3047692,
      "author_name": "chuanchengshi",
      "author_url": "",
      "post_date": "11/17/2024 04:12:14",
      "content": "<p>I think you are right</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3048319,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "11/17/2024 18:41:03",
      "content": "<p>For me, there are some reasonable CV-LB relation points, but they lie between 0.445 and 0.465 on the public leaderboard. However, I still believe we are just overfitting to the public leaderboard. I haven't seen any reproducible work among the high scores on the public leaderboard. <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3048406,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "11/17/2024 21:11:08",
          "content": "<p>True, same here - scores between 0.45 and 0.46 are correlating with CV. Anything more than that is all off-place <a href=\"https://www.kaggle.com/abdmental01\" target=\"_blank\">@abdmental01</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3048599,
      "author_name": "passengerc07",
      "author_url": "",
      "post_date": "11/18/2024 06:32:46",
      "content": "<p>My two cents:</p>\n<ol>\n<li>With highly noisy data like this one, simple models with good ML basics would work, leading to CV-LB relation over 0.44-0.46. Letting boosted trees to handle missing values and preprocessing is nice, and fitting is better compared to NN, but I don't think it makes any sense to spend too much time on tuning the parameters.</li>\n<li>The NN models (mostly MLPs) I tried are worse compared to LGBMs. But I think they correlate better to LB (less overfitting but lower LB score) and I know exactly when I'm overfitting.</li>\n<li>With NNs, it's over 0.42-0.44 range. With LGBMs no.</li>\n</ol>\n<p>BTW, this comp is ICR all over again. Expect shakeup at the end.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3048994,
      "author_name": "shin9915",
      "author_url": "",
      "post_date": "11/18/2024 15:33:33",
      "content": "<p>I think this is ICR competition (# ﾟДﾟ)(# ﾟДﾟ)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3049702,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "11/19/2024 12:31:44",
          "content": "<p>I also feel so, ICR part 2 <a href=\"https://www.kaggle.com/shin9915\" target=\"_blank\">@shin9915</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3049049,
      "author_name": "tomyuen",
      "author_url": "",
      "post_date": "11/18/2024 16:29:39",
      "content": "<p>To be honest I've almost given up in this competition, which is very sad to me, after spending over 150 hours in the first month.<br>\nI had a bad strategy in this game, and pushed too hard that I burnout too quickly.<br>\nMe being a beginner in ML, I guess the tech you mention won't work in this competition. If someone is courageous enough, maybe they should drop some features. I can't recall it anymore, but in one test in this game, I dropped quite many columns and the score didn't drop.<br>\nI guess I'll submit my 0.472 as final submission, basically a LGBM and a bit of data cleaning + featuring. I don't even bother blending it with catboost + xgboost anymore.<br>\nThis is my first participation in competition with prize and to be honest, I enjoyed much more in the playground game more. I should have manged my expectation better on Day 1.</p>\n<p>Good luck and may I wish everyone all the best!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3051758,
          "author_name": "eraakash",
          "author_url": "",
          "post_date": "11/21/2024 15:57:28",
          "content": "<p>I treated this competition as a regression problem; and my score was 0.366. Funny thing is that I did not even include data from the parquet files! I messed up, and never returned back to continue in this competition. It's been 2 months since my last submission haha</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3050492,
      "author_name": "xiaoling98899",
      "author_url": "",
      "post_date": "11/20/2024 09:05:10",
      "content": "<p>Makes sense! In fact, one of my 0.495 scores was achieved purely using machine learning without incorporating complex models like TabNet, and the result turned out to be better than when I included it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3061512,
          "author_name": "fanmingchen",
          "author_url": "",
          "post_date": "12/02/2024 17:50:10",
          "content": "<p>That's great! I think your direction is right, and I'm really looking forward to further exchanging ideas with you. I'm also trying to improve the model's score using simple models and appropriate data preprocessing methods.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3057191,
      "author_name": "kasuresh",
      "author_url": "",
      "post_date": "11/27/2024 20:34:45",
      "content": "<p>I tried out a simple CatBoost and was able to get 0.372. But the challenge here is to deal with both questionaire as well as the continuous time series data</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3046469": "Hello all,\n\nI presume a lot of us are highly inspired by a lot of good public work in this competition and are perhaps aligning our approaches in line with these kernels to some extent at least. I wish to discuss a few points I observed in these kernels and wish to seek your inputs on the same- \n\n1. Models like boosted trees, auto-encoder, tabnet regressor necessitate larger data sizes while the concurrent data is quite small. Up on slicing the data into CV folds, the data becomes even smaller, based on the number of folds one wishes to use. In such cases, do you think such complex approaches will work?\n2. The dataset is inherently very highly noisy and a lot of the tabular features are not highly predictive. Do you think using a NN-like architecture will overfit to the noise in the data in such cases and create an unwanted churn here?\n3. Do you have a reasonable CV-LB relation here? \n\nThoughts? Comments?",
    "3046749": "3- What's a reasonable/good CV-LB relation? My best score currently is CV 0.475 - LB 0.469.\nI think that's reasonable but also not very good, given the placement. The good relation is probably due to the fact, that I didnt use any imputation so far and only use a single modle (LightGBM) with feature engineering and data cleaning. I tried implementing the things mentioned in \"1.\" to little success so far. But maybe I have to try everything at once and not step by step. But then I will probably end up at the same notebooks, which are already public...\n\nI also made a post about the common notebooks, which use 3 models. I can't quite follow, why the models are chosen this way.",
    "3046756": "mariusheuser thanks for the rejoinder. I think we are in for a big shakeup as many public kernel elements are not sustainable.",
    "3046927": "I agree with your thoughts and in addition to the 3 models, it can be noticed that KNN Imputer has been applied with target variable SII which will lead to data leakage. Normally KNN Imputer shouldn't be applied with the target variable so, some how that wrong use of KNN Imputer is working well with 38% of test data and pushing the LB score high. In addition to this I read some people in discussion said that auto encoder has also been applied incorrectly, I didn't got much time to dig down in the auto encoder implementation of those notebooks. If you know what mistake is in them then do share, thank you.\nOverall we can say if these mistakes are leading to a better score with 38% of data then we can't say how it will perform with remaining 62% of data. If that 62% is almost similar to this 38% that results could be similar but on the other hand if that 62% is different then there can be a significant shakeup. Let's see what happens.",
    "3047047": "> Do you have a reasonable CV-LB relation here?\n\n **for me at least, no, <br>\nnot relating by any mean.**",
    "3047075": "Same here - no alignment is an indicator of a shakeup @letemoin",
    "3047381": "Totally agree. According to Kaggle competitions, there are often inconsistencies between the distribution of public and private LBs, so overfitting to the public LB usually does not lead to better (or even worse) results. For this reason, CV is usually more trustworthy and robust than LB.",
    "3047453": "Agree with you @ravi20076",
    "3047475": "negative values for BIA-BIA_BMC, BIA-BIA_FMI, BIA-BIA_Fat! Could it be due to an error or a calibration issue?",
    "3047482": "I agree with you on this point.",
    "3047549": "Probably it's a better idea to follow the simplicity path rather then the complexity path,There's alot of risk to overfit here and a high LB score doesn't look like it reflects how robust a model is.",
    "3047692": "I think you are right",
    "3048319": "For me, there are some reasonable CV-LB relation points, but they lie between 0.445 and 0.465 on the public leaderboard. However, I still believe we are just overfitting to the public leaderboard. I haven't seen any reproducible work among the high scores on the public leaderboard. @ravi20076",
    "3048322": "Yes, I tried experimenting with this. After seeing many public solutions, I noticed that they impute the target variable using KNN. After doing that, both the LB score and CV score improve, but it is clearly a case of data leakage/Overfitting.\n\nThe Main Problem is Copy-Paste , People Didn't Read Solutions , They Start Copying it Without Understanding it. Let See What Happen In the Private !",
    "3048406": "True, same here - scores between 0.45 and 0.46 are correlating with CV. Anything more than that is all off-place @abdmental01",
    "3048505": "Yes exactly, that leads to data leakage/over-fitting. The only thing that surprises me is that this overfitting is working well with 38% of data. That is more than 1/4th of the private dataset.",
    "3048599": "My two cents:\n\n1. With highly noisy data like this one, simple models with good ML basics would work, leading to CV-LB relation over 0.44-0.46. Letting boosted trees to handle missing values and preprocessing is nice, and fitting is better compared to NN, but I don't think it makes any sense to spend too much time on tuning the parameters.\n2. The NN models (mostly MLPs) I tried are worse compared to LGBMs. But I think they correlate better to LB (less overfitting but lower LB score) and I know exactly when I'm overfitting.\n3. With NNs, it's over 0.42-0.44 range. With LGBMs no.\n\nBTW, this comp is ICR all over again. Expect shakeup at the end.",
    "3048994": "I think this is ICR competition (# ﾟДﾟ)(# ﾟДﾟ)",
    "3049049": "To be honest I've almost given up in this competition, which is very sad to me, after spending over 150 hours in the first month.\nI had a bad strategy in this game, and pushed too hard that I burnout too quickly.\nMe being a beginner in ML, I guess the tech you mention won't work in this competition. If someone is courageous enough, maybe they should drop some features. I can't recall it anymore, but in one test in this game, I dropped quite many columns and the score didn't drop.\nI guess I'll submit my 0.472 as final submission, basically a LGBM and a bit of data cleaning + featuring. I don't even bother blending it with catboost + xgboost anymore.\nThis is my first participation in competition with prize and to be honest, I enjoyed much more in the playground game more. I should have manged my expectation better on Day 1.\n\nGood luck and may I wish everyone all the best!",
    "3049702": "I also feel so, ICR part 2 @shin9915",
    "3050492": "Makes sense! In fact, one of my 0.495 scores was achieved purely using machine learning without incorporating complex models like TabNet, and the result turned out to be better than when I included it.",
    "3051758": "I treated this competition as a regression problem; and my score was 0.366. Funny thing is that I did not even include data from the parquet files! I messed up, and never returned back to continue in this competition. It's been 2 months since my last submission haha",
    "3053393": "Hi @taimour ,\n\nAbout ``auto encoder has also been applied incorrectly`` : the autoencoder is trained on train timeseries and trained again on test timeseries, Instead of using a ``transform`` or ``predict`` method for test dataset.",
    "3053436": "Thank you for details.",
    "3057191": "I tried out a simple CatBoost and was able to get 0.372. But the challenge here is to deal with both questionaire as well as the continuous time series data",
    "3061512": "That's great! I think your direction is right, and I'm really looking forward to further exchanging ideas with you. I'm also trying to improve the model's score using simple models and appropriate data preprocessing methods."
  },
  "source": "meta"
}