{
  "id": 552501,
  "title": "Short advice for lottery competitors (ICR, CMI)",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/552501",
  "author_name": "",
  "post_date": "2024-12-20T02:08:17.792523800Z",
  "votes": 19,
  "comment_count": 13,
  "views": 0,
  "content": "<p>First of all, I would like to express my gratitude to all the organizers and related parties for their hard work on this competition. I will give a short advice to Kagglers for the upcoming lottery competition. (I apologize, but it's not a complete solution.)</p>\n<p>First, my history in the lottery competition:</p>\n<ul>\n<li>ICR 1773 -&gt; 151</li>\n<li>CMI 2425 (base on highest PB) -&gt; 60</li>\n</ul>\n<p>I am likely to win two silver medals, which is quite fortunate. The submissions I chose for this competition were a public notebook with LB 0.497 and my own notebook. Even in competitions where luck seems to play a big role, if you set good criteria for choosing your submission, you can achieve robust results.</p>\n<h3>[ 1 ] Trust your CV.</h3>\n<p>In most other competitions, the CV, Public, and Private scores tend to be similar, but in this particular competition, the Public is just for fun. You should trust the CV as much as possible without violating next items. Even if you look at my submission history below, the difference between the CV and Private is almost constant. And make sure there is no data leakage in the CV.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.4096</td>\n<td>0.454</td>\n<td>0.427</td>\n</tr>\n<tr>\n<td>0.4478</td>\n<td>0.452</td>\n<td>0.460</td>\n</tr>\n</tbody>\n</table>\n<h3>[ 2 ] Stay away from optimization.</h3>\n<p>I did not apply model parameter tuning and threshold tuning in this competition. (Only one model used publicly available tuned parameters, while others used randomly integerized or rounded off decimal values for public parameters.) These two techniques can be useful in normal competitions but may torture CV when labels are sparse. In particular, if you try to replace threshold tuning with performing it on each fold, you will see that there is no improvement in CV. I simply averaged individual fold models and performed round(0).</p>\n<h3>[ 3 ] Apply as simple an ensemble as possible.</h3>\n<p>I have ensembled 4 lgb models and 2 xgb models. Here, we can choose the weight of each model, but determining this weight can also lead to overfitting, so I simply obtained the mode value and took the smaller label value in case of a tie. Although the ensemble's CV was lower than that of individual models as shown below, I prioritized the ensemble result considering the influence of mode and minimizing variance.</p>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>model 0</td>\n<td>0.438611</td>\n</tr>\n<tr>\n<td>model 1</td>\n<td>0.410161</td>\n</tr>\n<tr>\n<td>model 2</td>\n<td>0.431642</td>\n</tr>\n<tr>\n<td>model 3</td>\n<td>0.420303</td>\n</tr>\n<tr>\n<td>model 4</td>\n<td>0.449363</td>\n</tr>\n<tr>\n<td>model 5</td>\n<td>0.427404</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>0.447887</td>\n</tr>\n</tbody>\n</table>\n<h3>[ 4 ] Don't try to do anything for a public score.</h3>\n<p>There are many good materials in the public notebooks. All these materials come together to get a good public score. However, some of them have low CVs and seem to be incorrect ways. If you want to add any element, make sure to judge based on solid CVs.</p>\n<ul>\n<li>feature_engineering: Features such as BMI_Age ~ ICW_TBW lowered the CV (I couldn't check individually).</li>\n<li>autoencoder: There seemed to be a data leakage in the published method, and I didn't want to apply it to time data with many empty values.</li>\n<li>tabnet and tasktype (multiclass…): Did not record high CV.</li>\n</ul>\n<h3>Thank you, Kagglers</h3>\n<p>One of the good things about Kaggle is that you get a chance to test your psychological factors that don't get swayed by LBs. Please aim for a robust model and stick to the basics. Lastly, I participated in this competition briefly, but I was inspired by the notebooks and discussions below. Also, thanks to everyone else who made their notebooks public.</p>\n<ul>\n<li>CMI | Best Single Model : <a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-model?scriptVersionId=198691885\" target=\"_blank\">https://www.kaggle.com/code/abdmental01/cmi-best-single-model?scriptVersionId=198691885</a></li>\n<li>Results of verifying the effectiveness of the custom objective for LGBM with Quadratic Weighted Kappa (QWK) : <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535052\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535052</a></li>\n<li>The Optimisers Curse : <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/550074\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/550074</a></li>\n<li>Optimized QWK effect on the LB score : <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/540738\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/540738</a></li>\n</ul>",
  "messages": [
    {
      "id": "3076494",
      "postDate": "12/20/2024 02:08:17",
      "content": "<p>First of all, I would like to express my gratitude to all the organizers and related parties for their hard work on this competition. I will give a short advice to Kagglers for the upcoming lottery competition. (I apologize, but it's not a complete solution.)</p>\n<p>First, my history in the lottery competition:</p>\n<ul>\n<li>ICR 1773 -&gt; 151</li>\n<li>CMI 2425 (base on highest PB) -&gt; 60</li>\n</ul>\n<p>I am likely to win two silver medals, which is quite fortunate. The submissions I chose for this competition were a public notebook with LB 0.497 and my own notebook. Even in competitions where luck seems to play a big role, if you set good criteria for choosing your submission, you can achieve robust results.</p>\n<h3>[ 1 ] Trust your CV.</h3>\n<p>In most other competitions, the CV, Public, and Private scores tend to be similar, but in this particular competition, the Public is just for fun. You should trust the CV as much as possible without violating next items. Even if you look at my submission history below, the difference between the CV and Private is almost constant. And make sure there is no data leakage in the CV.</p>\n<table>\n<thead>\n<tr>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.4096</td>\n<td>0.454</td>\n<td>0.427</td>\n</tr>\n<tr>\n<td>0.4478</td>\n<td>0.452</td>\n<td>0.460</td>\n</tr>\n</tbody>\n</table>\n<h3>[ 2 ] Stay away from optimization.</h3>\n<p>I did not apply model parameter tuning and threshold tuning in this competition. (Only one model used publicly available tuned parameters, while others used randomly integerized or rounded off decimal values for public parameters.) These two techniques can be useful in normal competitions but may torture CV when labels are sparse. In particular, if you try to replace threshold tuning with performing it on each fold, you will see that there is no improvement in CV. I simply averaged individual fold models and performed round(0).</p>\n<h3>[ 3 ] Apply as simple an ensemble as possible.</h3>\n<p>I have ensembled 4 lgb models and 2 xgb models. Here, we can choose the weight of each model, but determining this weight can also lead to overfitting, so I simply obtained the mode value and took the smaller label value in case of a tie. Although the ensemble's CV was lower than that of individual models as shown below, I prioritized the ensemble result considering the influence of mode and minimizing variance.</p>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>model 0</td>\n<td>0.438611</td>\n</tr>\n<tr>\n<td>model 1</td>\n<td>0.410161</td>\n</tr>\n<tr>\n<td>model 2</td>\n<td>0.431642</td>\n</tr>\n<tr>\n<td>model 3</td>\n<td>0.420303</td>\n</tr>\n<tr>\n<td>model 4</td>\n<td>0.449363</td>\n</tr>\n<tr>\n<td>model 5</td>\n<td>0.427404</td>\n</tr>\n<tr>\n<td>ensemble</td>\n<td>0.447887</td>\n</tr>\n</tbody>\n</table>\n<h3>[ 4 ] Don't try to do anything for a public score.</h3>\n<p>There are many good materials in the public notebooks. All these materials come together to get a good public score. However, some of them have low CVs and seem to be incorrect ways. If you want to add any element, make sure to judge based on solid CVs.</p>\n<ul>\n<li>feature_engineering: Features such as BMI_Age ~ ICW_TBW lowered the CV (I couldn't check individually).</li>\n<li>autoencoder: There seemed to be a data leakage in the published method, and I didn't want to apply it to time data with many empty values.</li>\n<li>tabnet and tasktype (multiclass…): Did not record high CV.</li>\n</ul>\n<h3>Thank you, Kagglers</h3>\n<p>One of the good things about Kaggle is that you get a chance to test your psychological factors that don't get swayed by LBs. Please aim for a robust model and stick to the basics. Lastly, I participated in this competition briefly, but I was inspired by the notebooks and discussions below. Also, thanks to everyone else who made their notebooks public.</p>\n<ul>\n<li>CMI | Best Single Model : <a href=\"https://www.kaggle.com/code/abdmental01/cmi-best-single-model?scriptVersionId=198691885\" target=\"_blank\">https://www.kaggle.com/code/abdmental01/cmi-best-single-model?scriptVersionId=198691885</a></li>\n<li>Results of verifying the effectiveness of the custom objective for LGBM with Quadratic Weighted Kappa (QWK) : <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535052\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535052</a></li>\n<li>The Optimisers Curse : <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/550074\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/550074</a></li>\n<li>Optimized QWK effect on the LB score : <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/540738\" target=\"_blank\">https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/540738</a></li>\n</ul>",
      "rawMarkdown": "First of all, I would like to express my gratitude to all the organizers and related parties for their hard work on this competition. I will give a short advice to Kagglers for the upcoming lottery competition. (I apologize, but it's not a complete solution.)\n\nFirst, my history in the lottery competition:\n- ICR 1773 -> 151\n- CMI 2425 (base on highest PB) -> 60\n\nI am likely to win two silver medals, which is quite fortunate. The submissions I chose for this competition were a public notebook with LB 0.497 and my own notebook. Even in competitions where luck seems to play a big role, if you set good criteria for choosing your submission, you can achieve robust results.\n\n### [ 1 ] Trust your CV.\nIn most other competitions, the CV, Public, and Private scores tend to be similar, but in this particular competition, the Public is just for fun. You should trust the CV as much as possible without violating next items. Even if you look at my submission history below, the difference between the CV and Private is almost constant. And make sure there is no data leakage in the CV.\n\n| CV | Public  | Private |\n| --- | --- | --- |\n| 0.4096 | 0.454 | 0.427 |\n| 0.4478 | 0.452 | 0.460 |  \n  \n  \n### [ 2 ] Stay away from optimization.\nI did not apply model parameter tuning and threshold tuning in this competition. (Only one model used publicly available tuned parameters, while others used randomly integerized or rounded off decimal values for public parameters.) These two techniques can be useful in normal competitions but may torture CV when labels are sparse. In particular, if you try to replace threshold tuning with performing it on each fold, you will see that there is no improvement in CV. I simply averaged individual fold models and performed round(0).\n  \n\n### [ 3 ] Apply as simple an ensemble as possible.\nI have ensembled 4 lgb models and 2 xgb models. Here, we can choose the weight of each model, but determining this weight can also lead to overfitting, so I simply obtained the mode value and took the smaller label value in case of a tie. Although the ensemble's CV was lower than that of individual models as shown below, I prioritized the ensemble result considering the influence of mode and minimizing variance.\n| No | CV |\n| --- | --- |\n|  model 0 |  0.438611 | \n|  model 1 |  0.410161 | \n|  model 2 |  0.431642 | \n|  model 3 |  0.420303 | \n|  model 4 |  0.449363 | \n|  model 5 |  0.427404 | \n|  ensemble |  0.447887| \n  \n\n### [ 4 ] Don't try to do anything for a public score.\nThere are many good materials in the public notebooks. All these materials come together to get a good public score. However, some of them have low CVs and seem to be incorrect ways. If you want to add any element, make sure to judge based on solid CVs.\n- feature_engineering: Features such as BMI_Age ~ ICW_TBW lowered the CV (I couldn't check individually).\n- autoencoder: There seemed to be a data leakage in the published method, and I didn't want to apply it to time data with many empty values.\n- tabnet and tasktype (multiclass...): Did not record high CV.\n  \n### Thank you, Kagglers\nOne of the good things about Kaggle is that you get a chance to test your psychological factors that don't get swayed by LBs. Please aim for a robust model and stick to the basics. Lastly, I participated in this competition briefly, but I was inspired by the notebooks and discussions below. Also, thanks to everyone else who made their notebooks public.\n\n- CMI | Best Single Model : https://www.kaggle.com/code/abdmental01/cmi-best-single-model?scriptVersionId=198691885\n- Results of verifying the effectiveness of the custom objective for LGBM with Quadratic Weighted Kappa (QWK) : https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535052\n- The Optimisers Curse : https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/550074\n- Optimized QWK effect on the LB score : https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/540738",
      "votes": null
    },
    {
      "id": "3076593",
      "postDate": "12/20/2024 05:24:07",
      "content": "<p><a href=\"https://www.kaggle.com/jaewook704\" target=\"_blank\">@jaewook704</a> I have a slightly different viewpoint.</p>\n<p>In the ICR competition, even CV was not reliable and a no-cv no-fe catboost with default parameters gave 4th place finish. This competition has a better structure and better approaches are at least appearing. <br>\nI think such competitions are 99% luck (rebranded as trust your cv) and 1% everything else. One may trust the cv only is one can build a trustworthy cv scheme first. In such competitions, such trustworthy cv schemes cannot be built, so one needs to rely on luck alone. </p>",
      "rawMarkdown": "jaewook704 I have a slightly different viewpoint.\n\nIn the ICR competition, even CV was not reliable and a no-cv no-fe catboost with default parameters gave 4th place finish. This competition has a better structure and better approaches are at least appearing. \nI think such competitions are 99% luck (rebranded as trust your cv) and 1% everything else. One may trust the cv only is one can build a trustworthy cv scheme first. In such competitions, such trustworthy cv schemes cannot be built, so one needs to rely on luck alone.",
      "votes": null
    },
    {
      "id": "3076627",
      "postDate": "12/20/2024 06:02:58",
      "content": "<p>I agree. In both competitions, I trusted my CV cuz I didn't have anything else to do. My single model oof score was 0.4882 here, but it scored 0.39x on private lol</p>",
      "rawMarkdown": "I agree. In both competitions, I trusted my CV cuz I didn't have anything else to do. My single model oof score was 0.4882 here, but it scored 0.39x on private lol",
      "votes": null
    },
    {
      "id": "3076655",
      "postDate": "12/20/2024 06:42:58",
      "content": "<p>That's right. There is a better approach to every competition, and 99% of me would have been lucky.<br>\nHowever, if there is a 1% difference, I think it was an attitude not to repeat the mistakes of the past by creating a random and simple model as much as possible. (When it is very difficult to know a better way.. like this competition)</p>\n<p>The CV configuration used simple StratifiedKFold. Personally, I often use StratifiedKFold and I see it on public notebook.<br>\nICR : StratifiedKFold(n_splits=5, shuffle=True, random_state=123)<br>\nCMI : StratifiedKFold(n_splits=5, shuffle=True, random_state=42)</p>\n<p>My history in ICR, too much parameter adjustment and model selection showed overoptimization to CV. Therefore, I did not perform parameter tuning and model selection in this competition.</p>\n<table>\n<thead>\n<tr>\n<th>Selected</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td>0.209863 ± 0.054387</td>\n<td>0.18261</td>\n<td>0.38702</td>\n<td></td>\n</tr>\n<tr>\n<td></td>\n<td>0.205530 ± 0.055983</td>\n<td>0.17007</td>\n<td>0.38525</td>\n<td>Best</td>\n</tr>\n<tr>\n<td>v</td>\n<td>0.192759 ± 0.042104</td>\n<td>0.17794</td>\n<td>0.39483</td>\n<td>+ parameter tuning start</td>\n</tr>\n<tr>\n<td>v</td>\n<td>0.174996 ± 0.058490</td>\n<td>0.18514</td>\n<td>0.42498</td>\n<td>+ manually model selection</td>\n</tr>\n</tbody>\n</table>\n<p>If there is such a competition in the future, I will choose a similar approach.</p>",
      "rawMarkdown": "That's right. There is a better approach to every competition, and 99% of me would have been lucky.\nHowever, if there is a 1% difference, I think it was an attitude not to repeat the mistakes of the past by creating a random and simple model as much as possible. (When it is very difficult to know a better way.. like this competition)\n\nThe CV configuration used simple StratifiedKFold. Personally, I often use StratifiedKFold and I see it on public notebook.\nICR : StratifiedKFold(n_splits=5, shuffle=True, random_state=123)\nCMI : StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\n\nMy history in ICR, too much parameter adjustment and model selection showed overoptimization to CV. Therefore, I did not perform parameter tuning and model selection in this competition.\n|Selected| CV | Public  | Private | |\n| --- | --- | --- | --- | --- |\n|| 0.209863 ± 0.054387 | 0.18261 | 0.38702 | |\n|| 0.205530 ± 0.055983 | 0.17007 | 0.38525 | Best |\n|v| 0.192759 ± 0.042104 | 0.17794 | 0.39483 | + parameter tuning start |\n|v| 0.174996 ± 0.058490 | 0.18514 | 0.42498 | + manually model selection |  \n  \n  \nIf there is such a competition in the future, I will choose a similar approach.",
      "votes": null
    },
    {
      "id": "3076663",
      "postDate": "12/20/2024 06:55:45",
      "content": "<p>I think it depends… Sometimes overfitting to CV works on the private test set and sometimes it doesn't. I wouldn't say anything I did on this competition was a mistake. It was my best bet and I gave it a shot, but it didn't work. It doesn't necessarily mean that it won't work on another lottery competition. I don't like safe solutions in lottery competitions like this since I don't have anything to lose, I just go all in on one direction. Maybe I could use another approach in my second submission though.</p>",
      "rawMarkdown": "I think it depends... Sometimes overfitting to CV works on the private test set and sometimes it doesn't. I wouldn't say anything I did on this competition was a mistake. It was my best bet and I gave it a shot, but it didn't work. It doesn't necessarily mean that it won't work on another lottery competition. I don't like safe solutions in lottery competitions like this since I don't have anything to lose, I just go all in on one direction. Maybe I could use another approach in my second submission though.",
      "votes": null
    },
    {
      "id": "3076671",
      "postDate": "12/20/2024 07:10:17",
      "content": "<p>Of course. It depends on the competition.<br>\nUnfortunately, if I had more time, I would have chosen an aggressive solution with another submission. I'm sure all attempts are good and will be the foundation for the next competition. Anyway, I'm always paying attention to your posts <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> and they're very helpful. Thank you for your good opinion and I look forward to seeing you again in other competitions.</p>",
      "rawMarkdown": "Of course. It depends on the competition.\nUnfortunately, if I had more time, I would have chosen an aggressive solution with another submission. I'm sure all attempts are good and will be the foundation for the next competition. Anyway, I'm always paying attention to your posts @ravi20076 @gunesevitan and they're very helpful. Thank you for your good opinion and I look forward to seeing you again in other competitions.",
      "votes": null
    },
    {
      "id": "3076695",
      "postDate": "12/20/2024 07:35:53",
      "content": "<p>Good job! Thanks! A lot</p>",
      "rawMarkdown": "Good job! Thanks! A lot",
      "votes": null
    },
    {
      "id": "3076781",
      "postDate": "12/20/2024 09:01:44",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/jaewook704\" target=\"_blank\">@jaewook704</a>, This is total luck game, many people included me Selected their best Cv but Best CV didn't work sometimes </p>",
      "rawMarkdown": "Congratulations @jaewook704, This is total luck game, many people included me Selected their best Cv but Best CV didn't work sometimes",
      "votes": null
    },
    {
      "id": "3076801",
      "postDate": "12/20/2024 09:23:29",
      "content": "<p>There is no valid advice for \"Lottery Competition\". If the dataset is noisy then it's just complete luck. You can't fit a model to random noise :)</p>",
      "rawMarkdown": "There is no valid advice for \"Lottery Competition\". If the dataset is noisy then it's just complete luck. You can't fit a model to random noise :)",
      "votes": null
    },
    {
      "id": "3077462",
      "postDate": "12/21/2024 02:03:53",
      "content": "<p>If the data were random noise, wouldn't I have submitted zero or random value? Did you?</p>",
      "rawMarkdown": "If the data were random noise, wouldn't I have submitted zero or random value? Did you?",
      "votes": null
    },
    {
      "id": "3077463",
      "postDate": "12/21/2024 02:06:43",
      "content": "<p>That's right. CVs often make us get lost. So I think it's important to construct a very solid CV. And thank you again. Your notebook helped a lot.</p>",
      "rawMarkdown": "That's right. CVs often make us get lost. So I think it's important to construct a very solid CV. And thank you again. Your notebook helped a lot.",
      "votes": null
    },
    {
      "id": "3077745",
      "postDate": "12/21/2024 11:00:59",
      "content": "<p>For this competition, the input feature matrix X' = X + \\epsilon_1. Where X is the original signal and \\epsilon_1 is the noise in the training set. It is possible to fit a model to X' but the noise part  is going to be different in the test set X_test. So, it's pretty much impossible to say which model will perform the best. </p>",
      "rawMarkdown": "For this competition, the input feature matrix X' = X + \\epsilon_1. Where X is the original signal and \\epsilon_1 is the noise in the training set. It is possible to fit a model to X' but the noise part  is going to be different in the test set X_test. So, it's pretty much impossible to say which model will perform the best.",
      "votes": null
    },
    {
      "id": "3077829",
      "postDate": "12/21/2024 12:53:01",
      "content": "<p>Focus on clarity and precision in your responses every word counts. Stay calm, trust your preparation, and manage your time wisely. Good luck!</p>",
      "rawMarkdown": "Focus on clarity and precision in your responses every word counts. Stay calm, trust your preparation, and manage your time wisely. Good luck!",
      "votes": null
    },
    {
      "id": "3078210",
      "postDate": "12/22/2024 00:04:08",
      "content": "<p>Right, that's why we need robust CV verification (other seeds, different fold split, etc.) and we need to take a simple approach. The more complicated the method, the more rigorous it should have been. Even though it's hard to find the best model, we'll be able to choose a good model from our many submissions.</p>",
      "rawMarkdown": "Right, that's why we need robust CV verification (other seeds, different fold split, etc.) and we need to take a simple approach. The more complicated the method, the more rigorous it should have been. Even though it's hard to find the best model, we'll be able to choose a good model from our many submissions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076593,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "12/20/2024 05:24:07",
      "content": "<p><a href=\"https://www.kaggle.com/jaewook704\" target=\"_blank\">@jaewook704</a> I have a slightly different viewpoint.</p>\n<p>In the ICR competition, even CV was not reliable and a no-cv no-fe catboost with default parameters gave 4th place finish. This competition has a better structure and better approaches are at least appearing. <br>\nI think such competitions are 99% luck (rebranded as trust your cv) and 1% everything else. One may trust the cv only is one can build a trustworthy cv scheme first. In such competitions, such trustworthy cv schemes cannot be built, so one needs to rely on luck alone. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3076627,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/20/2024 06:02:58",
          "content": "<p>I agree. In both competitions, I trusted my CV cuz I didn't have anything else to do. My single model oof score was 0.4882 here, but it scored 0.39x on private lol</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3076655,
          "author_name": "jaewook704",
          "author_url": "",
          "post_date": "12/20/2024 06:42:58",
          "content": "<p>That's right. There is a better approach to every competition, and 99% of me would have been lucky.<br>\nHowever, if there is a 1% difference, I think it was an attitude not to repeat the mistakes of the past by creating a random and simple model as much as possible. (When it is very difficult to know a better way.. like this competition)</p>\n<p>The CV configuration used simple StratifiedKFold. Personally, I often use StratifiedKFold and I see it on public notebook.<br>\nICR : StratifiedKFold(n_splits=5, shuffle=True, random_state=123)<br>\nCMI : StratifiedKFold(n_splits=5, shuffle=True, random_state=42)</p>\n<p>My history in ICR, too much parameter adjustment and model selection showed overoptimization to CV. Therefore, I did not perform parameter tuning and model selection in this competition.</p>\n<table>\n<thead>\n<tr>\n<th>Selected</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td></td>\n<td>0.209863 ± 0.054387</td>\n<td>0.18261</td>\n<td>0.38702</td>\n<td></td>\n</tr>\n<tr>\n<td></td>\n<td>0.205530 ± 0.055983</td>\n<td>0.17007</td>\n<td>0.38525</td>\n<td>Best</td>\n</tr>\n<tr>\n<td>v</td>\n<td>0.192759 ± 0.042104</td>\n<td>0.17794</td>\n<td>0.39483</td>\n<td>+ parameter tuning start</td>\n</tr>\n<tr>\n<td>v</td>\n<td>0.174996 ± 0.058490</td>\n<td>0.18514</td>\n<td>0.42498</td>\n<td>+ manually model selection</td>\n</tr>\n</tbody>\n</table>\n<p>If there is such a competition in the future, I will choose a similar approach.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3076663,
              "author_name": "gunesevitan",
              "author_url": "",
              "post_date": "12/20/2024 06:55:45",
              "content": "<p>I think it depends… Sometimes overfitting to CV works on the private test set and sometimes it doesn't. I wouldn't say anything I did on this competition was a mistake. It was my best bet and I gave it a shot, but it didn't work. It doesn't necessarily mean that it won't work on another lottery competition. I don't like safe solutions in lottery competitions like this since I don't have anything to lose, I just go all in on one direction. Maybe I could use another approach in my second submission though.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3076671,
                  "author_name": "jaewook704",
                  "author_url": "",
                  "post_date": "12/20/2024 07:10:17",
                  "content": "<p>Of course. It depends on the competition.<br>\nUnfortunately, if I had more time, I would have chosen an aggressive solution with another submission. I'm sure all attempts are good and will be the foundation for the next competition. Anyway, I'm always paying attention to your posts <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> and they're very helpful. Thank you for your good opinion and I look forward to seeing you again in other competitions.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3076695,
      "author_name": "mrsimple07",
      "author_url": "",
      "post_date": "12/20/2024 07:35:53",
      "content": "<p>Good job! Thanks! A lot</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3076781,
      "author_name": "abdmental01",
      "author_url": "",
      "post_date": "12/20/2024 09:01:44",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/jaewook704\" target=\"_blank\">@jaewook704</a>, This is total luck game, many people included me Selected their best Cv but Best CV didn't work sometimes </p>",
      "votes": null,
      "replies": [
        {
          "id": 3077463,
          "author_name": "jaewook704",
          "author_url": "",
          "post_date": "12/21/2024 02:06:43",
          "content": "<p>That's right. CVs often make us get lost. So I think it's important to construct a very solid CV. And thank you again. Your notebook helped a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3076801,
      "author_name": "basu1999",
      "author_url": "",
      "post_date": "12/20/2024 09:23:29",
      "content": "<p>There is no valid advice for \"Lottery Competition\". If the dataset is noisy then it's just complete luck. You can't fit a model to random noise :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3077462,
          "author_name": "jaewook704",
          "author_url": "",
          "post_date": "12/21/2024 02:03:53",
          "content": "<p>If the data were random noise, wouldn't I have submitted zero or random value? Did you?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3077745,
              "author_name": "basu1999",
              "author_url": "",
              "post_date": "12/21/2024 11:00:59",
              "content": "<p>For this competition, the input feature matrix X' = X + \\epsilon_1. Where X is the original signal and \\epsilon_1 is the noise in the training set. It is possible to fit a model to X' but the noise part  is going to be different in the test set X_test. So, it's pretty much impossible to say which model will perform the best. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3078210,
                  "author_name": "jaewook704",
                  "author_url": "",
                  "post_date": "12/22/2024 00:04:08",
                  "content": "<p>Right, that's why we need robust CV verification (other seeds, different fold split, etc.) and we need to take a simple approach. The more complicated the method, the more rigorous it should have been. Even though it's hard to find the best model, we'll be able to choose a good model from our many submissions.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3077829,
      "author_name": "talhaanjum0",
      "author_url": "",
      "post_date": "12/21/2024 12:53:01",
      "content": "<p>Focus on clarity and precision in your responses every word counts. Stay calm, trust your preparation, and manage your time wisely. Good luck!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3076494": "First of all, I would like to express my gratitude to all the organizers and related parties for their hard work on this competition. I will give a short advice to Kagglers for the upcoming lottery competition. (I apologize, but it's not a complete solution.)\n\nFirst, my history in the lottery competition:\n- ICR 1773 -> 151\n- CMI 2425 (base on highest PB) -> 60\n\nI am likely to win two silver medals, which is quite fortunate. The submissions I chose for this competition were a public notebook with LB 0.497 and my own notebook. Even in competitions where luck seems to play a big role, if you set good criteria for choosing your submission, you can achieve robust results.\n\n### [ 1 ] Trust your CV.\nIn most other competitions, the CV, Public, and Private scores tend to be similar, but in this particular competition, the Public is just for fun. You should trust the CV as much as possible without violating next items. Even if you look at my submission history below, the difference between the CV and Private is almost constant. And make sure there is no data leakage in the CV.\n\n| CV | Public  | Private |\n| --- | --- | --- |\n| 0.4096 | 0.454 | 0.427 |\n| 0.4478 | 0.452 | 0.460 |  \n  \n  \n### [ 2 ] Stay away from optimization.\nI did not apply model parameter tuning and threshold tuning in this competition. (Only one model used publicly available tuned parameters, while others used randomly integerized or rounded off decimal values for public parameters.) These two techniques can be useful in normal competitions but may torture CV when labels are sparse. In particular, if you try to replace threshold tuning with performing it on each fold, you will see that there is no improvement in CV. I simply averaged individual fold models and performed round(0).\n  \n\n### [ 3 ] Apply as simple an ensemble as possible.\nI have ensembled 4 lgb models and 2 xgb models. Here, we can choose the weight of each model, but determining this weight can also lead to overfitting, so I simply obtained the mode value and took the smaller label value in case of a tie. Although the ensemble's CV was lower than that of individual models as shown below, I prioritized the ensemble result considering the influence of mode and minimizing variance.\n| No | CV |\n| --- | --- |\n|  model 0 |  0.438611 | \n|  model 1 |  0.410161 | \n|  model 2 |  0.431642 | \n|  model 3 |  0.420303 | \n|  model 4 |  0.449363 | \n|  model 5 |  0.427404 | \n|  ensemble |  0.447887| \n  \n\n### [ 4 ] Don't try to do anything for a public score.\nThere are many good materials in the public notebooks. All these materials come together to get a good public score. However, some of them have low CVs and seem to be incorrect ways. If you want to add any element, make sure to judge based on solid CVs.\n- feature_engineering: Features such as BMI_Age ~ ICW_TBW lowered the CV (I couldn't check individually).\n- autoencoder: There seemed to be a data leakage in the published method, and I didn't want to apply it to time data with many empty values.\n- tabnet and tasktype (multiclass...): Did not record high CV.\n  \n### Thank you, Kagglers\nOne of the good things about Kaggle is that you get a chance to test your psychological factors that don't get swayed by LBs. Please aim for a robust model and stick to the basics. Lastly, I participated in this competition briefly, but I was inspired by the notebooks and discussions below. Also, thanks to everyone else who made their notebooks public.\n\n- CMI | Best Single Model : https://www.kaggle.com/code/abdmental01/cmi-best-single-model?scriptVersionId=198691885\n- Results of verifying the effectiveness of the custom objective for LGBM with Quadratic Weighted Kappa (QWK) : https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/535052\n- The Optimisers Curse : https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/550074\n- Optimized QWK effect on the LB score : https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/540738",
    "3076593": "jaewook704 I have a slightly different viewpoint.\n\nIn the ICR competition, even CV was not reliable and a no-cv no-fe catboost with default parameters gave 4th place finish. This competition has a better structure and better approaches are at least appearing. \nI think such competitions are 99% luck (rebranded as trust your cv) and 1% everything else. One may trust the cv only is one can build a trustworthy cv scheme first. In such competitions, such trustworthy cv schemes cannot be built, so one needs to rely on luck alone.",
    "3076627": "I agree. In both competitions, I trusted my CV cuz I didn't have anything else to do. My single model oof score was 0.4882 here, but it scored 0.39x on private lol",
    "3076655": "That's right. There is a better approach to every competition, and 99% of me would have been lucky.\nHowever, if there is a 1% difference, I think it was an attitude not to repeat the mistakes of the past by creating a random and simple model as much as possible. (When it is very difficult to know a better way.. like this competition)\n\nThe CV configuration used simple StratifiedKFold. Personally, I often use StratifiedKFold and I see it on public notebook.\nICR : StratifiedKFold(n_splits=5, shuffle=True, random_state=123)\nCMI : StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\n\nMy history in ICR, too much parameter adjustment and model selection showed overoptimization to CV. Therefore, I did not perform parameter tuning and model selection in this competition.\n|Selected| CV | Public  | Private | |\n| --- | --- | --- | --- | --- |\n|| 0.209863 ± 0.054387 | 0.18261 | 0.38702 | |\n|| 0.205530 ± 0.055983 | 0.17007 | 0.38525 | Best |\n|v| 0.192759 ± 0.042104 | 0.17794 | 0.39483 | + parameter tuning start |\n|v| 0.174996 ± 0.058490 | 0.18514 | 0.42498 | + manually model selection |  \n  \n  \nIf there is such a competition in the future, I will choose a similar approach.",
    "3076663": "I think it depends... Sometimes overfitting to CV works on the private test set and sometimes it doesn't. I wouldn't say anything I did on this competition was a mistake. It was my best bet and I gave it a shot, but it didn't work. It doesn't necessarily mean that it won't work on another lottery competition. I don't like safe solutions in lottery competitions like this since I don't have anything to lose, I just go all in on one direction. Maybe I could use another approach in my second submission though.",
    "3076671": "Of course. It depends on the competition.\nUnfortunately, if I had more time, I would have chosen an aggressive solution with another submission. I'm sure all attempts are good and will be the foundation for the next competition. Anyway, I'm always paying attention to your posts @ravi20076 @gunesevitan and they're very helpful. Thank you for your good opinion and I look forward to seeing you again in other competitions.",
    "3076695": "Good job! Thanks! A lot",
    "3076781": "Congratulations @jaewook704, This is total luck game, many people included me Selected their best Cv but Best CV didn't work sometimes",
    "3076801": "There is no valid advice for \"Lottery Competition\". If the dataset is noisy then it's just complete luck. You can't fit a model to random noise :)",
    "3077462": "If the data were random noise, wouldn't I have submitted zero or random value? Did you?",
    "3077463": "That's right. CVs often make us get lost. So I think it's important to construct a very solid CV. And thank you again. Your notebook helped a lot.",
    "3077745": "For this competition, the input feature matrix X' = X + \\epsilon_1. Where X is the original signal and \\epsilon_1 is the noise in the training set. It is possible to fit a model to X' but the noise part  is going to be different in the test set X_test. So, it's pretty much impossible to say which model will perform the best.",
    "3077829": "Focus on clarity and precision in your responses every word counts. Stay calm, trust your preparation, and manage your time wisely. Good luck!",
    "3078210": "Right, that's why we need robust CV verification (other seeds, different fold split, etc.) and we need to take a simple approach. The more complicated the method, the more rigorous it should have been. Even though it's hard to find the best model, we'll be able to choose a good model from our many submissions."
  },
  "source": "meta"
}