{
  "id": 133746,
  "title": "I made a mistake...",
  "url": "/competitions/bengaliai-cv19/discussion/133746",
  "author_name": "",
  "post_date": "2020-03-04T04:23:29.935229300Z",
  "votes": 27,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Well, I made a silly mistake that I hard-coded <code>fold = 0</code> in my training code.\nWhich means, I thought I trained <code>5-fold</code> models, but in fact it was <code>5-fold0</code> models 😑 </p>\n\n<p>Yesterday, I see some people are able to get around LB 0.9905 while their best single model is around 0.9890 or something, but in my case the boost was much smaller when I do 5-fold ensemble, so I checked my code and found the silly mistake......</p>\n\n<p>It turns out that I don't really know how 5-fold models can boost the LB score against to single fold 😑 But luckily I have enough time to retrain 80% of my models....</p>\n\n<p>For those guys who didn't get a boost around 0.001~ by 5-fold models, please check your code to make sure that you didn't make a same mistake like me. 👀 </p>",
  "messages": [
    {
      "id": "763043",
      "postDate": "03/04/2020 04:23:29",
      "content": "<p>Well, I made a silly mistake that I hard-coded <code>fold = 0</code> in my training code.\nWhich means, I thought I trained <code>5-fold</code> models, but in fact it was <code>5-fold0</code> models 😑 </p>\n\n<p>Yesterday, I see some people are able to get around LB 0.9905 while their best single model is around 0.9890 or something, but in my case the boost was much smaller when I do 5-fold ensemble, so I checked my code and found the silly mistake......</p>\n\n<p>It turns out that I don't really know how 5-fold models can boost the LB score against to single fold 😑 But luckily I have enough time to retrain 80% of my models....</p>\n\n<p>For those guys who didn't get a boost around 0.001~ by 5-fold models, please check your code to make sure that you didn't make a same mistake like me. 👀 </p>",
      "rawMarkdown": "Well, I made a silly mistake that I hard-coded `fold = 0` in my training code.\nWhich means, I thought I trained `5-fold` models, but in fact it was `5-fold0` models 😑 \n\nYesterday, I see some people are able to get around LB 0.9905 while their best single model is around 0.9890 or something, but in my case the boost was much smaller when I do 5-fold ensemble, so I checked my code and found the silly mistake......\n\nIt turns out that I don't really know how 5-fold models can boost the LB score against to single fold 😑 But luckily I have enough time to retrain 80% of my models....\n\nFor those guys who didn't get a boost around 0.001~ by 5-fold models, please check your code to make sure that you didn't make a same mistake like me. 👀",
      "votes": null
    },
    {
      "id": "763050",
      "postDate": "03/04/2020 04:36:43",
      "content": "<p>0.993x is coming😏 </p>",
      "rawMarkdown": "0.993x is coming😏",
      "votes": null
    },
    {
      "id": "763124",
      "postDate": "03/04/2020 06:53:11",
      "content": "<p>I have a question, is it really useful to train 5 models? I mean, the bootstrapping might help a lot when you ensemble 2 or 3 models, but adding additional folds might make a long time and only have marginal (possibly) improvements. Isn't it better when you e.g. train 2 models on different folds? (like seresnext50 on fold-0, fold1, and seresnext101 on fold2, fold3)</p>",
      "rawMarkdown": "I have a question, is it really useful to train 5 models? I mean, the bootstrapping might help a lot when you ensemble 2 or 3 models, but adding additional folds might make a long time and only have marginal (possibly) improvements. Isn't it better when you e.g. train 2 models on different folds? (like seresnext50 on fold-0, fold1, and seresnext101 on fold2, fold3)",
      "votes": null
    },
    {
      "id": "763129",
      "postDate": "03/04/2020 06:57:01",
      "content": "<p>I am also curious if anyone would be kind enough to share the benefits of 5-fold over e.g. 2-fold. I have not trained any model beyond 2 folds myself. Wondering if it is worth the while to train 5-folds instead of 2 or three different architectures, each 2 folds</p>",
      "rawMarkdown": "I am also curious if anyone would be kind enough to share the benefits of 5-fold over e.g. 2-fold. I have not trained any model beyond 2 folds myself. Wondering if it is worth the while to train 5-folds instead of 2 or three different architectures, each 2 folds",
      "votes": null
    },
    {
      "id": "763131",
      "postDate": "03/04/2020 07:04:22",
      "content": "<p>The gap widens</p>",
      "rawMarkdown": "The gap widens",
      "votes": null
    },
    {
      "id": "763139",
      "postDate": "03/04/2020 07:12:54",
      "content": "<p>Well, you use 5 runs, which still yields a boost over a single run.  Not sure you'll see lot sof upside, but hopefully you'll see some.</p>\n\n<p>If you need gpus I can team with you and run your models ;)</p>",
      "rawMarkdown": "Well, you use 5 runs, which still yields a boost over a single run.  Not sure you'll see lot sof upside, but hopefully you'll see some.\n\nIf you need gpus I can team with you and run your models ;)",
      "votes": null
    },
    {
      "id": "763182",
      "postDate": "03/04/2020 08:10:51",
      "content": "<p>One 5 fold ensembling experiment of mine:\nmy best fold was around 98.2 validation : LB 97.51\nother folds where around 97.7~98.0 validation : ensembling them 97.68\nstacking less folds yielded scores from 97.63 to 97.66</p>\n\n<p>But since my scores are not around 99 I guess the uplift would decrease with better score, I still believe you can get better results with ensembling different folds</p>",
      "rawMarkdown": "One 5 fold ensembling experiment of mine:\nmy best fold was around 98.2 validation : LB 97.51\nother folds where around 97.7~98.0 validation : ensembling them 97.68\nstacking less folds yielded scores from 97.63 to 97.66\n\nBut since my scores are not around 99 I guess the uplift would decrease with better score, I still believe you can get better results with ensembling different folds",
      "votes": null
    },
    {
      "id": "763186",
      "postDate": "03/04/2020 08:14:58",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> I can do a full linting of your code and make sure that everything is clean when we share it to the world if you wanna team! ^^\nOtherwise, <a href=\"/cpmpml\">@cpmpml</a> wanna heat some GPUs with me?</p>",
      "rawMarkdown": "haqishen I can do a full linting of your code and make sure that everything is clean when we share it to the world if you wanna team! ^^\nOtherwise, @cpmpml wanna heat some GPUs with me?",
      "votes": null
    },
    {
      "id": "763206",
      "postDate": "03/04/2020 08:41:01",
      "content": "<p>whatever works I guess, there is no truth regarding blending/stacking</p>",
      "rawMarkdown": "whatever works I guess, there is no truth regarding blending/stacking",
      "votes": null
    },
    {
      "id": "763220",
      "postDate": "03/04/2020 09:03:03",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Yes, I'll need some luck to make 5-fold works. </p>\n\n<p><a href=\"/cpmpml\">@cpmpml</a> <a href=\"/optimo\">@optimo</a> Thanks for offering! Maybe in the next competition, we can have a chance to burst our gpus or review the code together ;)</p>",
      "rawMarkdown": "cpmpml Yes, I'll need some luck to make 5-fold works. \n\n@cpmpml @optimo Thanks for offering! Maybe in the next competition, we can have a chance to burst our gpus or review the code together ;)",
      "votes": null
    },
    {
      "id": "763222",
      "postDate": "03/04/2020 09:05:29",
      "content": "<p>just like <a href=\"/optimo\">@optimo</a> said, in most of the case, running 5-fold over 2-fold is kind of... trying to do things perfectly, but no necessary to have a significant improvement.</p>",
      "rawMarkdown": "just like @optimo said, in most of the case, running 5-fold over 2-fold is kind of... trying to do things perfectly, but no necessary to have a significant improvement.",
      "votes": null
    },
    {
      "id": "763227",
      "postDate": "03/04/2020 09:12:14",
      "content": "<p>&gt; e.g. train 2 models on different folds? (like seresnext50 on fold-0, fold1, and seresnext101 on fold2, fold3)</p>\n\n<p>Indeed, this is LGTM. But I didn't try it so I can't tell anything.</p>",
      "rawMarkdown": "&gt; e.g. train 2 models on different folds? (like seresnext50 on fold-0, fold1, and seresnext101 on fold2, fold3)\n\nIndeed, this is LGTM. But I didn't try it so I can't tell anything.",
      "votes": null
    },
    {
      "id": "763232",
      "postDate": "03/04/2020 09:19:03",
      "content": "<p>I need some luck on it 🤓 </p>",
      "rawMarkdown": "I need some luck on it 🤓",
      "votes": null
    },
    {
      "id": "763480",
      "postDate": "03/04/2020 14:20:34",
      "content": "<p>Currently training 5fold seresnext50s, in total it takes like 7-10 days straight of training. So far 2 have finished but ensembling them didn't give a boost over single model.\nWith my limited ensembling experience it seems you get the most boost if you ensemble drastically different models (whether different architecture or aug, etc).\nEDIT: Finished 5fold, ensembling doesn't give any boost over just using one model, wonder why...</p>",
      "rawMarkdown": "Currently training 5fold seresnext50s, in total it takes like 7-10 days straight of training. So far 2 have finished but ensembling them didn't give a boost over single model.\nWith my limited ensembling experience it seems you get the most boost if you ensemble drastically different models (whether different architecture or aug, etc).\nEDIT: Finished 5fold, ensembling doesn't give any boost over just using one model, wonder why...",
      "votes": null
    },
    {
      "id": "763508",
      "postDate": "03/04/2020 15:04:34",
      "content": "<p>For me, 2 folds ensemble is better than 5 folds. 🙈 </p>",
      "rawMarkdown": "For me, 2 folds ensemble is better than 5 folds. 🙈",
      "votes": null
    },
    {
      "id": "763576",
      "postDate": "03/04/2020 16:07:06",
      "content": "<p>Well, you should believe that it will make a difference in private LB ;)</p>",
      "rawMarkdown": "Well, you should believe that it will make a difference in private LB ;)",
      "votes": null
    },
    {
      "id": "763664",
      "postDate": "03/04/2020 18:03:04",
      "content": "<p>\"Well, I made a silly mistake that I hard-coded fold = 0 in my training code.\nWhich means, I thought I trained 5-fold models, but in fact it was 5-fold0 \"</p>\n\n<p>average the weights into single model</p>",
      "rawMarkdown": "\"Well, I made a silly mistake that I hard-coded fold = 0 in my training code.\nWhich means, I thought I trained 5-fold models, but in fact it was 5-fold0 \"\n\naverage the weights into single model",
      "votes": null
    },
    {
      "id": "763965",
      "postDate": "03/05/2020 03:11:59",
      "content": "<p>It's 5 models running separately for over 200 epochs... 😂 \nI think simply average the weights won't work. It's not like SWA.</p>",
      "rawMarkdown": "It's 5 models running separately for over 200 epochs... 😂 \nI think simply average the weights won't work. It's not like SWA.",
      "votes": null
    },
    {
      "id": "763994",
      "postDate": "03/05/2020 04:01:12",
      "content": "<p>did you try to perform model ensembling via LightGBM?</p>",
      "rawMarkdown": "did you try to perform model ensembling via LightGBM?",
      "votes": null
    },
    {
      "id": "764018",
      "postDate": "03/05/2020 04:39:58",
      "content": "<p>pretty wasteful if they are not used in submission. i am thinking that since their train data is the same, maybe it can work. alternatively, you can finetune them once all using the same exact same sampling, e.g. shuffle the train sample index and fix the sample and use it to finetune them all. </p>\n\n<p>but again, it may take too much efforts and you might just choose the best model out of all.</p>",
      "rawMarkdown": "pretty wasteful if they are not used in submission. i am thinking that since their train data is the same, maybe it can work. alternatively, you can finetune them once all using the same exact same sampling, e.g. shuffle the train sample index and fix the sample and use it to finetune them all. \n\nbut again, it may take too much efforts and you might just choose the best model out of all.",
      "votes": null
    },
    {
      "id": "764043",
      "postDate": "03/05/2020 05:08:59",
      "content": "<p>\"For me, 2 folds ensemble is better than 5 folds\"</p>\n\n<p>my experience tells me that the private and public LB will disagree on the statement above</p>",
      "rawMarkdown": "\"For me, 2 folds ensemble is better than 5 folds\"\n\nmy experience tells me that the private and public LB will disagree on the statement above",
      "votes": null
    },
    {
      "id": "764284",
      "postDate": "03/05/2020 09:50:31",
      "content": "<p>Thus any boosts from 1 fold-0 model into 5 fold-0 models? It is a bagging ensemble.</p>",
      "rawMarkdown": "Thus any boosts from 1 fold-0 model into 5 fold-0 models? It is a bagging ensemble.",
      "votes": null
    },
    {
      "id": "764297",
      "postDate": "03/05/2020 10:05:55",
      "content": "<p>A tiny improvement can be observed...</p>",
      "rawMarkdown": "A tiny improvement can be observed...",
      "votes": null
    },
    {
      "id": "764504",
      "postDate": "03/05/2020 14:36:58",
      "content": "<p>:D</p>",
      "rawMarkdown": ":D",
      "votes": null
    },
    {
      "id": "764565",
      "postDate": "03/05/2020 15:58:17",
      "content": "<p>In a similar vein, I just realized I've been feeding <code>0-255</code> values into my net. been really struggling to understand why i couldn't break 0.949 regardless of architecture (ENb3,ENb4,SERNX,etc). I hope this is the reason.... off to another 15 hours of training.</p>",
      "rawMarkdown": "In a similar vein, I just realized I've been feeding `0-255` values into my net. been really struggling to understand why i couldn't break 0.949 regardless of architecture (ENb3,ENb4,SERNX,etc). I hope this is the reason.... off to another 15 hours of training.",
      "votes": null
    },
    {
      "id": "764775",
      "postDate": "03/05/2020 22:03:27",
      "content": "<p>My results:\n1 fold: 0.99721 local, 0.9881 LB\n5 folds: ? CV, 0.9892 LB</p>",
      "rawMarkdown": "My results:\n1 fold: 0.99721 local, 0.9881 LB\n5 folds: ? CV, 0.9892 LB",
      "votes": null
    },
    {
      "id": "764791",
      "postDate": "03/05/2020 22:45:19",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "764804",
      "postDate": "03/05/2020 23:24:04",
      "content": "<p>i find it fascinating that just using 80% of the data, model is only 0.001 worse</p>",
      "rawMarkdown": "i find it fascinating that just using 80% of the data, model is only 0.001 worse",
      "votes": null
    },
    {
      "id": "765010",
      "postDate": "03/06/2020 07:20:21",
      "content": "<p><a href=\"/drn01z3\">@drn01z3</a> Thank you very much! This should motivate me to do some hardcore ensembling...</p>",
      "rawMarkdown": "drn01z3 Thank you very much! This should motivate me to do some hardcore ensembling...",
      "votes": null
    },
    {
      "id": "765045",
      "postDate": "03/06/2020 07:55:12",
      "content": "<p>Thanks for the information! Hope I can get to 0.993+ with my 5 fold models 😄</p>",
      "rawMarkdown": "Thanks for the information! Hope I can get to 0.993+ with my 5 fold models 😄",
      "votes": null
    },
    {
      "id": "765195",
      "postDate": "03/06/2020 11:04:43",
      "content": "<p>You should sub again :) We are all curious.</p>",
      "rawMarkdown": "You should sub again :) We are all curious.",
      "votes": null
    },
    {
      "id": "765222",
      "postDate": "03/06/2020 11:53:18",
      "content": "<p>I'm the one who most curious, but it still need 50~60h to run. 😇</p>",
      "rawMarkdown": "I'm the one who most curious, but it still need 50~60h to run. 😇",
      "votes": null
    },
    {
      "id": "767253",
      "postDate": "03/09/2020 11:23:17",
      "content": "<p>Seems you fixed your mistake!</p>",
      "rawMarkdown": "Seems you fixed your mistake!",
      "votes": null
    },
    {
      "id": "767256",
      "postDate": "03/09/2020 11:28:04",
      "content": "<p>Yea, by proper ensemble the boost of LB can definitely be 0.001~ 😃</p>",
      "rawMarkdown": "Yea, by proper ensemble the boost of LB can definitely be 0.001~ 😃",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 763050,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "03/04/2020 04:36:43",
      "content": "<p>0.993x is coming😏 </p>",
      "votes": null,
      "replies": [
        {
          "id": 763131,
          "author_name": "authman",
          "author_url": "",
          "post_date": "03/04/2020 07:04:22",
          "content": "<p>The gap widens</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763232,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/04/2020 09:19:03",
          "content": "<p>I need some luck on it 🤓 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764504,
          "author_name": "tuhinadribanerjee",
          "author_url": "",
          "post_date": "03/05/2020 14:36:58",
          "content": "<p>:D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 763124,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "03/04/2020 06:53:11",
      "content": "<p>I have a question, is it really useful to train 5 models? I mean, the bootstrapping might help a lot when you ensemble 2 or 3 models, but adding additional folds might make a long time and only have marginal (possibly) improvements. Isn't it better when you e.g. train 2 models on different folds? (like seresnext50 on fold-0, fold1, and seresnext101 on fold2, fold3)</p>",
      "votes": null,
      "replies": [
        {
          "id": 763206,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/04/2020 08:41:01",
          "content": "<p>whatever works I guess, there is no truth regarding blending/stacking</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763227,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/04/2020 09:12:14",
          "content": "<p>&gt; e.g. train 2 models on different folds? (like seresnext50 on fold-0, fold1, and seresnext101 on fold2, fold3)</p>\n\n<p>Indeed, this is LGTM. But I didn't try it so I can't tell anything.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 763129,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "03/04/2020 06:57:01",
      "content": "<p>I am also curious if anyone would be kind enough to share the benefits of 5-fold over e.g. 2-fold. I have not trained any model beyond 2 folds myself. Wondering if it is worth the while to train 5-folds instead of 2 or three different architectures, each 2 folds</p>",
      "votes": null,
      "replies": [
        {
          "id": 763182,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "03/04/2020 08:10:51",
          "content": "<p>One 5 fold ensembling experiment of mine:\nmy best fold was around 98.2 validation : LB 97.51\nother folds where around 97.7~98.0 validation : ensembling them 97.68\nstacking less folds yielded scores from 97.63 to 97.66</p>\n\n<p>But since my scores are not around 99 I guess the uplift would decrease with better score, I still believe you can get better results with ensembling different folds</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763222,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/04/2020 09:05:29",
          "content": "<p>just like <a href=\"/optimo\">@optimo</a> said, in most of the case, running 5-fold over 2-fold is kind of... trying to do things perfectly, but no necessary to have a significant improvement.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763480,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "03/04/2020 14:20:34",
          "content": "<p>Currently training 5fold seresnext50s, in total it takes like 7-10 days straight of training. So far 2 have finished but ensembling them didn't give a boost over single model.\nWith my limited ensembling experience it seems you get the most boost if you ensemble drastically different models (whether different architecture or aug, etc).\nEDIT: Finished 5fold, ensembling doesn't give any boost over just using one model, wonder why...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763994,
          "author_name": "projdev",
          "author_url": "",
          "post_date": "03/05/2020 04:01:12",
          "content": "<p>did you try to perform model ensembling via LightGBM?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 763139,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/04/2020 07:12:54",
      "content": "<p>Well, you use 5 runs, which still yields a boost over a single run.  Not sure you'll see lot sof upside, but hopefully you'll see some.</p>\n\n<p>If you need gpus I can team with you and run your models ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 763186,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "03/04/2020 08:14:58",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> I can do a full linting of your code and make sure that everything is clean when we share it to the world if you wanna team! ^^\nOtherwise, <a href=\"/cpmpml\">@cpmpml</a> wanna heat some GPUs with me?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763220,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/04/2020 09:03:03",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> Yes, I'll need some luck to make 5-fold works. </p>\n\n<p><a href=\"/cpmpml\">@cpmpml</a> <a href=\"/optimo\">@optimo</a> Thanks for offering! Maybe in the next competition, we can have a chance to burst our gpus or review the code together ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 763508,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "03/04/2020 15:04:34",
      "content": "<p>For me, 2 folds ensemble is better than 5 folds. 🙈 </p>",
      "votes": null,
      "replies": [
        {
          "id": 763576,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/04/2020 16:07:06",
          "content": "<p>Well, you should believe that it will make a difference in private LB ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764043,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/05/2020 05:08:59",
          "content": "<p>\"For me, 2 folds ensemble is better than 5 folds\"</p>\n\n<p>my experience tells me that the private and public LB will disagree on the statement above</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764791,
          "author_name": "nonhermitian",
          "author_url": "",
          "post_date": "03/05/2020 22:45:19",
          "content": "",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 763664,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/04/2020 18:03:04",
      "content": "<p>\"Well, I made a silly mistake that I hard-coded fold = 0 in my training code.\nWhich means, I thought I trained 5-fold models, but in fact it was 5-fold0 \"</p>\n\n<p>average the weights into single model</p>",
      "votes": null,
      "replies": [
        {
          "id": 763965,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/05/2020 03:11:59",
          "content": "<p>It's 5 models running separately for over 200 epochs... 😂 \nI think simply average the weights won't work. It's not like SWA.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764018,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "03/05/2020 04:39:58",
          "content": "<p>pretty wasteful if they are not used in submission. i am thinking that since their train data is the same, maybe it can work. alternatively, you can finetune them once all using the same exact same sampling, e.g. shuffle the train sample index and fix the sample and use it to finetune them all. </p>\n\n<p>but again, it may take too much efforts and you might just choose the best model out of all.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 764284,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "03/05/2020 09:50:31",
      "content": "<p>Thus any boosts from 1 fold-0 model into 5 fold-0 models? It is a bagging ensemble.</p>",
      "votes": null,
      "replies": [
        {
          "id": 764297,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/05/2020 10:05:55",
          "content": "<p>A tiny improvement can be observed...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 764565,
      "author_name": "authman",
      "author_url": "",
      "post_date": "03/05/2020 15:58:17",
      "content": "<p>In a similar vein, I just realized I've been feeding <code>0-255</code> values into my net. been really struggling to understand why i couldn't break 0.949 regardless of architecture (ENb3,ENb4,SERNX,etc). I hope this is the reason.... off to another 15 hours of training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 764775,
      "author_name": "drn01z3",
      "author_url": "",
      "post_date": "03/05/2020 22:03:27",
      "content": "<p>My results:\n1 fold: 0.99721 local, 0.9881 LB\n5 folds: ? CV, 0.9892 LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 764804,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "03/05/2020 23:24:04",
          "content": "<p>i find it fascinating that just using 80% of the data, model is only 0.001 worse</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765010,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "03/06/2020 07:20:21",
          "content": "<p><a href=\"/drn01z3\">@drn01z3</a> Thank you very much! This should motivate me to do some hardcore ensembling...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765045,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/06/2020 07:55:12",
          "content": "<p>Thanks for the information! Hope I can get to 0.993+ with my 5 fold models 😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765195,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/06/2020 11:04:43",
          "content": "<p>You should sub again :) We are all curious.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765222,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/06/2020 11:53:18",
          "content": "<p>I'm the one who most curious, but it still need 50~60h to run. 😇</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 767253,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/09/2020 11:23:17",
      "content": "<p>Seems you fixed your mistake!</p>",
      "votes": null,
      "replies": [
        {
          "id": 767256,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/09/2020 11:28:04",
          "content": "<p>Yea, by proper ensemble the boost of LB can definitely be 0.001~ 😃</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "763043": "Well, I made a silly mistake that I hard-coded `fold = 0` in my training code.\nWhich means, I thought I trained `5-fold` models, but in fact it was `5-fold0` models 😑 \n\nYesterday, I see some people are able to get around LB 0.9905 while their best single model is around 0.9890 or something, but in my case the boost was much smaller when I do 5-fold ensemble, so I checked my code and found the silly mistake......\n\nIt turns out that I don't really know how 5-fold models can boost the LB score against to single fold 😑 But luckily I have enough time to retrain 80% of my models....\n\nFor those guys who didn't get a boost around 0.001~ by 5-fold models, please check your code to make sure that you didn't make a same mistake like me. 👀",
    "763050": "0.993x is coming😏",
    "763124": "I have a question, is it really useful to train 5 models? I mean, the bootstrapping might help a lot when you ensemble 2 or 3 models, but adding additional folds might make a long time and only have marginal (possibly) improvements. Isn't it better when you e.g. train 2 models on different folds? (like seresnext50 on fold-0, fold1, and seresnext101 on fold2, fold3)",
    "763129": "I am also curious if anyone would be kind enough to share the benefits of 5-fold over e.g. 2-fold. I have not trained any model beyond 2 folds myself. Wondering if it is worth the while to train 5-folds instead of 2 or three different architectures, each 2 folds",
    "763131": "The gap widens",
    "763139": "Well, you use 5 runs, which still yields a boost over a single run.  Not sure you'll see lot sof upside, but hopefully you'll see some.\n\nIf you need gpus I can team with you and run your models ;)",
    "763182": "One 5 fold ensembling experiment of mine:\nmy best fold was around 98.2 validation : LB 97.51\nother folds where around 97.7~98.0 validation : ensembling them 97.68\nstacking less folds yielded scores from 97.63 to 97.66\n\nBut since my scores are not around 99 I guess the uplift would decrease with better score, I still believe you can get better results with ensembling different folds",
    "763186": "haqishen I can do a full linting of your code and make sure that everything is clean when we share it to the world if you wanna team! ^^\nOtherwise, @cpmpml wanna heat some GPUs with me?",
    "763206": "whatever works I guess, there is no truth regarding blending/stacking",
    "763220": "cpmpml Yes, I'll need some luck to make 5-fold works. \n\n@cpmpml @optimo Thanks for offering! Maybe in the next competition, we can have a chance to burst our gpus or review the code together ;)",
    "763222": "just like @optimo said, in most of the case, running 5-fold over 2-fold is kind of... trying to do things perfectly, but no necessary to have a significant improvement.",
    "763227": "&gt; e.g. train 2 models on different folds? (like seresnext50 on fold-0, fold1, and seresnext101 on fold2, fold3)\n\nIndeed, this is LGTM. But I didn't try it so I can't tell anything.",
    "763232": "I need some luck on it 🤓",
    "763480": "Currently training 5fold seresnext50s, in total it takes like 7-10 days straight of training. So far 2 have finished but ensembling them didn't give a boost over single model.\nWith my limited ensembling experience it seems you get the most boost if you ensemble drastically different models (whether different architecture or aug, etc).\nEDIT: Finished 5fold, ensembling doesn't give any boost over just using one model, wonder why...",
    "763508": "For me, 2 folds ensemble is better than 5 folds. 🙈",
    "763576": "Well, you should believe that it will make a difference in private LB ;)",
    "763664": "\"Well, I made a silly mistake that I hard-coded fold = 0 in my training code.\nWhich means, I thought I trained 5-fold models, but in fact it was 5-fold0 \"\n\naverage the weights into single model",
    "763965": "It's 5 models running separately for over 200 epochs... 😂 \nI think simply average the weights won't work. It's not like SWA.",
    "763994": "did you try to perform model ensembling via LightGBM?",
    "764018": "pretty wasteful if they are not used in submission. i am thinking that since their train data is the same, maybe it can work. alternatively, you can finetune them once all using the same exact same sampling, e.g. shuffle the train sample index and fix the sample and use it to finetune them all. \n\nbut again, it may take too much efforts and you might just choose the best model out of all.",
    "764043": "\"For me, 2 folds ensemble is better than 5 folds\"\n\nmy experience tells me that the private and public LB will disagree on the statement above",
    "764284": "Thus any boosts from 1 fold-0 model into 5 fold-0 models? It is a bagging ensemble.",
    "764297": "A tiny improvement can be observed...",
    "764504": ":D",
    "764565": "In a similar vein, I just realized I've been feeding `0-255` values into my net. been really struggling to understand why i couldn't break 0.949 regardless of architecture (ENb3,ENb4,SERNX,etc). I hope this is the reason.... off to another 15 hours of training.",
    "764775": "My results:\n1 fold: 0.99721 local, 0.9881 LB\n5 folds: ? CV, 0.9892 LB",
    "764791": "",
    "764804": "i find it fascinating that just using 80% of the data, model is only 0.001 worse",
    "765010": "drn01z3 Thank you very much! This should motivate me to do some hardcore ensembling...",
    "765045": "Thanks for the information! Hope I can get to 0.993+ with my 5 fold models 😄",
    "765195": "You should sub again :) We are all curious.",
    "765222": "I'm the one who most curious, but it still need 50~60h to run. 😇",
    "767253": "Seems you fixed your mistake!",
    "767256": "Yea, by proper ensemble the boost of LB can definitely be 0.001~ 😃"
  },
  "source": "meta"
}