{
  "id": 220752,
  "title": "123rd place (Silver) solution",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/georgi-pamukov-123rd-place-silver-solution",
  "author_name": "",
  "post_date": "2021-02-19T21:34:46.013Z",
  "votes": 10,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hey Dear Kagglers :)</p>\n<p>First of all, congratulations to all winners and participants!<br>\nAnd many thanks to all contributors for their amazing kernels and topics!</p>\n<p>Wanted to put my 2 cents in - so you can find a brief overview of my solution below:</p>\n<h3>General approach</h3>\n<p>I noticed the CV results was quite unstable overall - I was observing very significant discrepancies in validation and LB scores between the different folds of the same models (mostly because it is not easy to stratify the noise (probably :) ) ). When blending 2-3 best epochs from each fold it was kind of stabilizing - but it was not a very efficient way to do so in my opinion (especially considering the submission time limitations of the code competitions). I discovered though, I was able to produce better/equivalent results by blending the 3-4 best epochs of a model trained on the full set (stabilizing the models \"horizontally\" in a way). This also allowed me to include way more (diverse) models in my submissions. <br>\nAnd so I did:</p>\n<ul>\n<li>Final models were trained on the FULL data.</li>\n<li>Still used CV for validation/evaluation purposes only.</li>\n</ul>\n<p>Was this the best approach? Well on one side I perfectly survived the shakeup. On the other - I failed to identify my best submission (which would have been in the money/gold medal zone: <a href=\"https://www.kaggle.com/gpamoukoff/private-lb-0-902-blend-effnet-b3-b4-vit-resnext50)\" target=\"_blank\">https://www.kaggle.com/gpamoukoff/private-lb-0-902-blend-effnet-b3-b4-vit-resnext50)</a>.</p>\n<h3>Models</h3>\n<p>Trained with the aim to be as diverse as possible:</p>\n<ul>\n<li>effnet b3/b4, vit, xception, resnext50 architectures</li>\n<li>Label Smoothing</li>\n<li>Various loss functions: CE, Symmetric CE, Taylor CE</li>\n<li>Various schedulers: Gradual Warmup/CosineAnnealingWarmRestarts</li>\n<li>Augmentation: played mostly around crops/cutmix apart from the standard techniques.</li>\n<li>Data:<ul>\n<li>only 2020</li>\n<li>pretrain on 2019,  train on 2020</li>\n<li>train on both (one pass)</li></ul></li>\n</ul>\n<h3>Ensemble:</h3>\n<p>Hierarchical blend:</p>\n<ul>\n<li>Blended different epoch/best of kind models first</li>\n<li>Then blended the already blended “super” models</li>\n<li>Used CV to optimise the weights</li>\n</ul>\n<p>In short - could have been way better. Next time. Keep walking…</p>",
  "messages": [
    {
      "id": "1210394",
      "postDate": "02/19/2021 11:57:44",
      "content": "<p>Hey Dear Kagglers :)</p>\n<p>First of all, congratulations to all winners and participants!<br>\nAnd many thanks to all contributors for their amazing kernels and topics!</p>\n<p>Wanted to put my 2 cents in - so you can find a brief overview of my solution below:</p>\n<h3>General approach</h3>\n<p>I noticed the CV results was quite unstable overall - I was observing very significant discrepancies in validation and LB scores between the different folds of the same models (mostly because it is not easy to stratify the noise (probably :) ) ). When blending 2-3 best epochs from each fold it was kind of stabilizing - but it was not a very efficient way to do so in my opinion (especially considering the submission time limitations of the code competitions). I discovered though, I was able to produce better/equivalent results by blending the 3-4 best epochs of a model trained on the full set (stabilizing the models \"horizontally\" in a way). This also allowed me to include way more (diverse) models in my submissions. <br>\nAnd so I did:</p>\n<ul>\n<li>Final models were trained on the FULL data.</li>\n<li>Still used CV for validation/evaluation purposes only.</li>\n</ul>\n<p>Was this the best approach? Well on one side I perfectly survived the shakeup. On the other - I failed to identify my best submission (which would have been in the money/gold medal zone: <a href=\"https://www.kaggle.com/gpamoukoff/private-lb-0-902-blend-effnet-b3-b4-vit-resnext50)\" target=\"_blank\">https://www.kaggle.com/gpamoukoff/private-lb-0-902-blend-effnet-b3-b4-vit-resnext50)</a>.</p>\n<h3>Models</h3>\n<p>Trained with the aim to be as diverse as possible:</p>\n<ul>\n<li>effnet b3/b4, vit, xception, resnext50 architectures</li>\n<li>Label Smoothing</li>\n<li>Various loss functions: CE, Symmetric CE, Taylor CE</li>\n<li>Various schedulers: Gradual Warmup/CosineAnnealingWarmRestarts</li>\n<li>Augmentation: played mostly around crops/cutmix apart from the standard techniques.</li>\n<li>Data:<ul>\n<li>only 2020</li>\n<li>pretrain on 2019,  train on 2020</li>\n<li>train on both (one pass)</li></ul></li>\n</ul>\n<h3>Ensemble:</h3>\n<p>Hierarchical blend:</p>\n<ul>\n<li>Blended different epoch/best of kind models first</li>\n<li>Then blended the already blended “super” models</li>\n<li>Used CV to optimise the weights</li>\n</ul>\n<p>In short - could have been way better. Next time. Keep walking…</p>",
      "rawMarkdown": "Hey Dear Kagglers :)\n\nFirst of all, congratulations to all winners and participants!\nAnd many thanks to all contributors for their amazing kernels and topics!\n\nWanted to put my 2 cents in - so you can find a brief overview of my solution below:\n\n### General approach\n\nI noticed the CV results was quite unstable overall - I was observing very significant discrepancies in validation and LB scores between the different folds of the same models (mostly because it is not easy to stratify the noise (probably :) ) ). When blending 2-3 best epochs from each fold it was kind of stabilizing - but it was not a very efficient way to do so in my opinion (especially considering the submission time limitations of the code competitions). I discovered though, I was able to produce better/equivalent results by blending the 3-4 best epochs of a model trained on the full set (stabilizing the models \"horizontally\" in a way). This also allowed me to include way more (diverse) models in my submissions. \nAnd so I did:\n- Final models were trained on the FULL data.\n- Still used CV for validation/evaluation purposes only.\n\nWas this the best approach? Well on one side I perfectly survived the shakeup. On the other - I failed to identify my best submission (which would have been in the money/gold medal zone: https://www.kaggle.com/gpamoukoff/private-lb-0-902-blend-effnet-b3-b4-vit-resnext50).\n\n### Models\nTrained with the aim to be as diverse as possible:\n- effnet b3/b4, vit, xception, resnext50 architectures\n- Label Smoothing\n- Various loss functions: CE, Symmetric CE, Taylor CE\n- Various schedulers: Gradual Warmup/CosineAnnealingWarmRestarts\n- Augmentation: played mostly around crops/cutmix apart from the standard techniques.\n- Data:\n    - only 2020\n    - pretrain on 2019,  train on 2020\n    - train on both (one pass)\n\n### Ensemble:\nHierarchical blend:\n- Blended different epoch/best of kind models first\n- Then blended the already blended “super” models\n- Used CV to optimise the weights\n\nIn short - could have been way better. Next time. Keep walking...",
      "votes": null
    },
    {
      "id": "1210402",
      "postDate": "02/19/2021 12:07:31",
      "content": "<p>Nice summary. I have a doubt though, what do you mean by \"blending  2-3 best epochs\"?<br>\nI know blending is an ensemble technique but I can't understand blending epochs.<br>\nI presume you used pytorch. correct me if I am wrong.</p>",
      "rawMarkdown": "Nice summary. I have a doubt though, what do you mean by \"blending  2-3 best epochs\"?\nI know blending is an ensemble technique but I can't understand blending epochs.\nI presume you used pytorch. correct me if I am wrong.",
      "votes": null
    },
    {
      "id": "1210423",
      "postDate": "02/19/2021 12:22:09",
      "content": "<p>Nice work, congrats 🎉</p>\n<p>I have blended several epochs also, mainly from fold2 since seed 42 gave really good results on fold2. Fold2: Epoch 8 - &gt; 0.901, epoch  10 -&gt; 0.905 , Fold1: epoch 11: 0.900 , epoch 13: 0.901. This approach costed me shake down for 1100+ places 😅. Ensemble of this model gave LB of 0.904, with light TTA, light augmentations, ELR loss (noise 0.4, beta: 2), EfficientNetB4_ns, , 512 image size,  2020 data, Adam optimizer, step lr (step size 10) for 15 epochs. I used pixel rescaling of the images, it gave me the best results tho. Where did I failed I don't know. I trusted my cv also…</p>",
      "rawMarkdown": "Nice work, congrats 🎉\n\nI have blended several epochs also, mainly from fold2 since seed 42 gave really good results on fold2. Fold2: Epoch 8 - > 0.901, epoch  10 -> 0.905 , Fold1: epoch 11: 0.900 , epoch 13: 0.901. This approach costed me shake down for 1100+ places 😅. Ensemble of this model gave LB of 0.904, with light TTA, light augmentations, ELR loss (noise 0.4, beta: 2), EfficientNetB4_ns, , 512 image size,  2020 data, Adam optimizer, step lr (step size 10) for 15 epochs. I used pixel rescaling of the images, it gave me the best results tho. Where did I failed I don't know. I trusted my cv also...",
      "votes": null
    },
    {
      "id": "1210452",
      "postDate": "02/19/2021 12:54:59",
      "content": "<p>Продължаваме напред! 😄( мен много ме удари шейкъпа, прайвъта беше като биткоин ) / Keep walking, strong approach! 😀 ( in contrast I was heavily hit by the shake-up )</p>",
      "rawMarkdown": "Продължаваме напред! 😄( мен много ме удари шейкъпа, прайвъта беше като биткоин ) / Keep walking, strong approach! 😀 ( in contrast I was heavily hit by the shake-up )",
      "votes": null
    },
    {
      "id": "1210460",
      "postDate": "02/19/2021 13:02:07",
      "content": "<p>Hey, usually when you predict from cross validated model, you take the best epoch(weights) for each fold, predict from it - and then average the results to create the final prediction. You can make more predictions from each fold though - like instead of one(best) - take the 2-3 best epochs(weights), predict from them and average at the end. This is kind of \"self blending\" the same model. For example if you have 5 fold CV - and you choose and predict from 3 best epochs for each fold - you will end up with averaging 15 predictions at the end. That will be of course more stable than averaging just 5 predictions (one (best epoch) from each fold). I achieved pretty much the same result though by training the model on the full data (not using CV) - and then averaging the predictions from the 3-4 best epochs. Basically achieving stability with 3-4 predictions rather than 15. I still used CV to guess the performance of the model and choose the (probably) best epochs to blend. Is this a good idea - depends on the particular case of course. But is a valid approach. Please let me know if something is still unclear :)</p>",
      "rawMarkdown": "Hey, usually when you predict from cross validated model, you take the best epoch(weights) for each fold, predict from it - and then average the results to create the final prediction. You can make more predictions from each fold though - like instead of one(best) - take the 2-3 best epochs(weights), predict from them and average at the end. This is kind of \"self blending\" the same model. For example if you have 5 fold CV - and you choose and predict from 3 best epochs for each fold - you will end up with averaging 15 predictions at the end. That will be of course more stable than averaging just 5 predictions (one (best epoch) from each fold). I achieved pretty much the same result though by training the model on the full data (not using CV) - and then averaging the predictions from the 3-4 best epochs. Basically achieving stability with 3-4 predictions rather than 15. I still used CV to guess the performance of the model and choose the (probably) best epochs to blend. Is this a good idea - depends on the particular case of course. But is a valid approach. Please let me know if something is still unclear :)",
      "votes": null
    },
    {
      "id": "1210466",
      "postDate": "02/19/2021 13:06:21",
      "content": "<p>And yes I use pytorch</p>",
      "rawMarkdown": "And yes I use pytorch",
      "votes": null
    },
    {
      "id": "1210484",
      "postDate": "02/19/2021 13:25:24",
      "content": "<p>Определено беше доста … нестабилно (за да не използвам по-силни думи :D ). Не, че е нещо необичайно напоследък. За мен беше (за съжаление) по-скоро пропусната възможност :( Моделът, който щеше да ми донесе злато беше сред финалните ми кандидати (и се представяше най-добре сред тях). Аз обаче не му повярвах - все ми се струваше, че следващият най-добър е по-стабилен. И така приключих със сребро … :( Keep walking…</p>",
      "rawMarkdown": "Определено беше доста ... нестабилно (за да не използвам по-силни думи :D ). Не, че е нещо необичайно напоследък. За мен беше (за съжаление) по-скоро пропусната възможност :( Моделът, който щеше да ми донесе злато беше сред финалните ми кандидати (и се представяше най-добре сред тях). Аз обаче не му повярвах - все ми се струваше, че следващият най-добър е по-стабилен. И така приключих със сребро ... :( Keep walking...",
      "votes": null
    },
    {
      "id": "1210486",
      "postDate": "02/19/2021 13:28:50",
      "content": "<p>Don't think you failed mate - you did it all correct. It was the super noisy (and unstable) data unfortunately.</p>",
      "rawMarkdown": "Don't think you failed mate - you did it all correct. It was the super noisy (and unstable) data unfortunately.",
      "votes": null
    },
    {
      "id": "1210492",
      "postDate": "02/19/2021 13:34:23",
      "content": "<p>Thank you mate for the insights.  First competition so far, I really learn a lot anyways. Lessons learned, for many more competitions !</p>",
      "rawMarkdown": "Thank you mate for the insights.  First competition so far, I really learn a lot anyways. Lessons learned, for many more competitions !",
      "votes": null
    },
    {
      "id": "1210729",
      "postDate": "02/19/2021 16:30:16",
      "content": "<p>Did not have an idea we can do this kind of self blending. Thank you very much for the thorough explanation , this was very informative, something to try out in the future I guess.</p>",
      "rawMarkdown": "Did not have an idea we can do this kind of self blending. Thank you very much for the thorough explanation , this was very informative, something to try out in the future I guess.",
      "votes": null
    },
    {
      "id": "1210730",
      "postDate": "02/19/2021 16:31:28",
      "content": "<p>I need to seriously switch to pytorch, keras just doesn't seem good enough, every new paper implementation is in pytorch now, even the google guys won't use TF2. </p>",
      "rawMarkdown": "I need to seriously switch to pytorch, keras just doesn't seem good enough, every new paper implementation is in pytorch now, even the google guys won't use TF2.",
      "votes": null
    },
    {
      "id": "1211013",
      "postDate": "02/19/2021 21:57:43",
      "content": "<p>This is for sure the best part of the Kaggle journey - learning new stuff :) Keep it up and good luck!</p>",
      "rawMarkdown": "This is for sure the best part of the Kaggle journey - learning new stuff :) Keep it up and good luck!",
      "votes": null
    },
    {
      "id": "1211015",
      "postDate": "02/19/2021 22:11:02",
      "content": "<p>Same mate, wish you all the best also ^_^ !</p>",
      "rawMarkdown": "Same mate, wish you all the best also ^_^ !",
      "votes": null
    },
    {
      "id": "1211034",
      "postDate": "02/19/2021 22:55:37",
      "content": "<p>Now honestly it depends (re Keras/Tensorflow). Pytorch is very good for research/experimentation/quick&amp;dirty iterations - which makes it of course super fit for Kaggle. But is it that good for production/real-world applications (where performance/stability/scalability are crucial) - this is arguable (in my opinion). For example, I use ONLY TF in my job and also for my side projects. My point is - the picture you see here on Kaggle is a bit biased (re technologies in use/popularity) - and doesn't necessarily reflect the real world situation. TF is dominating the industry with a huge margin - and for a good reason :) (just like pytorch is dominating the research/academic fields). Good luck with your next competition - cheers!      </p>",
      "rawMarkdown": "Now honestly it depends (re Keras/Tensorflow). Pytorch is very good for research/experimentation/quick&dirty iterations - which makes it of course super fit for Kaggle. But is it that good for production/real-world applications (where performance/stability/scalability are crucial) - this is arguable (in my opinion). For example, I use ONLY TF in my job and also for my side projects. My point is - the picture you see here on Kaggle is a bit biased (re technologies in use/popularity) - and doesn't necessarily reflect the real world situation. TF is dominating the industry with a huge margin - and for a good reason :) (just like pytorch is dominating the research/academic fields). Good luck with your next competition - cheers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1210402,
      "author_name": "mohneesh7",
      "author_url": "",
      "post_date": "02/19/2021 12:07:31",
      "content": "<p>Nice summary. I have a doubt though, what do you mean by \"blending  2-3 best epochs\"?<br>\nI know blending is an ensemble technique but I can't understand blending epochs.<br>\nI presume you used pytorch. correct me if I am wrong.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210460,
          "author_name": "gpamoukoff",
          "author_url": "",
          "post_date": "02/19/2021 13:02:07",
          "content": "<p>Hey, usually when you predict from cross validated model, you take the best epoch(weights) for each fold, predict from it - and then average the results to create the final prediction. You can make more predictions from each fold though - like instead of one(best) - take the 2-3 best epochs(weights), predict from them and average at the end. This is kind of \"self blending\" the same model. For example if you have 5 fold CV - and you choose and predict from 3 best epochs for each fold - you will end up with averaging 15 predictions at the end. That will be of course more stable than averaging just 5 predictions (one (best epoch) from each fold). I achieved pretty much the same result though by training the model on the full data (not using CV) - and then averaging the predictions from the 3-4 best epochs. Basically achieving stability with 3-4 predictions rather than 15. I still used CV to guess the performance of the model and choose the (probably) best epochs to blend. Is this a good idea - depends on the particular case of course. But is a valid approach. Please let me know if something is still unclear :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210466,
          "author_name": "gpamoukoff",
          "author_url": "",
          "post_date": "02/19/2021 13:06:21",
          "content": "<p>And yes I use pytorch</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210729,
          "author_name": "mohneesh7",
          "author_url": "",
          "post_date": "02/19/2021 16:30:16",
          "content": "<p>Did not have an idea we can do this kind of self blending. Thank you very much for the thorough explanation , this was very informative, something to try out in the future I guess.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210730,
          "author_name": "mohneesh7",
          "author_url": "",
          "post_date": "02/19/2021 16:31:28",
          "content": "<p>I need to seriously switch to pytorch, keras just doesn't seem good enough, every new paper implementation is in pytorch now, even the google guys won't use TF2. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211034,
          "author_name": "gpamoukoff",
          "author_url": "",
          "post_date": "02/19/2021 22:55:37",
          "content": "<p>Now honestly it depends (re Keras/Tensorflow). Pytorch is very good for research/experimentation/quick&amp;dirty iterations - which makes it of course super fit for Kaggle. But is it that good for production/real-world applications (where performance/stability/scalability are crucial) - this is arguable (in my opinion). For example, I use ONLY TF in my job and also for my side projects. My point is - the picture you see here on Kaggle is a bit biased (re technologies in use/popularity) - and doesn't necessarily reflect the real world situation. TF is dominating the industry with a huge margin - and for a good reason :) (just like pytorch is dominating the research/academic fields). Good luck with your next competition - cheers!      </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210423,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "02/19/2021 12:22:09",
      "content": "<p>Nice work, congrats 🎉</p>\n<p>I have blended several epochs also, mainly from fold2 since seed 42 gave really good results on fold2. Fold2: Epoch 8 - &gt; 0.901, epoch  10 -&gt; 0.905 , Fold1: epoch 11: 0.900 , epoch 13: 0.901. This approach costed me shake down for 1100+ places 😅. Ensemble of this model gave LB of 0.904, with light TTA, light augmentations, ELR loss (noise 0.4, beta: 2), EfficientNetB4_ns, , 512 image size,  2020 data, Adam optimizer, step lr (step size 10) for 15 epochs. I used pixel rescaling of the images, it gave me the best results tho. Where did I failed I don't know. I trusted my cv also…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210486,
          "author_name": "gpamoukoff",
          "author_url": "",
          "post_date": "02/19/2021 13:28:50",
          "content": "<p>Don't think you failed mate - you did it all correct. It was the super noisy (and unstable) data unfortunately.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210492,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "02/19/2021 13:34:23",
          "content": "<p>Thank you mate for the insights.  First competition so far, I really learn a lot anyways. Lessons learned, for many more competitions !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211013,
          "author_name": "gpamoukoff",
          "author_url": "",
          "post_date": "02/19/2021 21:57:43",
          "content": "<p>This is for sure the best part of the Kaggle journey - learning new stuff :) Keep it up and good luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211015,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "02/19/2021 22:11:02",
          "content": "<p>Same mate, wish you all the best also ^_^ !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210452,
      "author_name": "atanasova",
      "author_url": "",
      "post_date": "02/19/2021 12:54:59",
      "content": "<p>Продължаваме напред! 😄( мен много ме удари шейкъпа, прайвъта беше като биткоин ) / Keep walking, strong approach! 😀 ( in contrast I was heavily hit by the shake-up )</p>",
      "votes": null,
      "replies": [
        {
          "id": 1210484,
          "author_name": "gpamoukoff",
          "author_url": "",
          "post_date": "02/19/2021 13:25:24",
          "content": "<p>Определено беше доста … нестабилно (за да не използвам по-силни думи :D ). Не, че е нещо необичайно напоследък. За мен беше (за съжаление) по-скоро пропусната възможност :( Моделът, който щеше да ми донесе злато беше сред финалните ми кандидати (и се представяше най-добре сред тях). Аз обаче не му повярвах - все ми се струваше, че следващият най-добър е по-стабилен. И така приключих със сребро … :( Keep walking…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1210394": "Hey Dear Kagglers :)\n\nFirst of all, congratulations to all winners and participants!\nAnd many thanks to all contributors for their amazing kernels and topics!\n\nWanted to put my 2 cents in - so you can find a brief overview of my solution below:\n\n### General approach\n\nI noticed the CV results was quite unstable overall - I was observing very significant discrepancies in validation and LB scores between the different folds of the same models (mostly because it is not easy to stratify the noise (probably :) ) ). When blending 2-3 best epochs from each fold it was kind of stabilizing - but it was not a very efficient way to do so in my opinion (especially considering the submission time limitations of the code competitions). I discovered though, I was able to produce better/equivalent results by blending the 3-4 best epochs of a model trained on the full set (stabilizing the models \"horizontally\" in a way). This also allowed me to include way more (diverse) models in my submissions. \nAnd so I did:\n- Final models were trained on the FULL data.\n- Still used CV for validation/evaluation purposes only.\n\nWas this the best approach? Well on one side I perfectly survived the shakeup. On the other - I failed to identify my best submission (which would have been in the money/gold medal zone: https://www.kaggle.com/gpamoukoff/private-lb-0-902-blend-effnet-b3-b4-vit-resnext50).\n\n### Models\nTrained with the aim to be as diverse as possible:\n- effnet b3/b4, vit, xception, resnext50 architectures\n- Label Smoothing\n- Various loss functions: CE, Symmetric CE, Taylor CE\n- Various schedulers: Gradual Warmup/CosineAnnealingWarmRestarts\n- Augmentation: played mostly around crops/cutmix apart from the standard techniques.\n- Data:\n    - only 2020\n    - pretrain on 2019,  train on 2020\n    - train on both (one pass)\n\n### Ensemble:\nHierarchical blend:\n- Blended different epoch/best of kind models first\n- Then blended the already blended “super” models\n- Used CV to optimise the weights\n\nIn short - could have been way better. Next time. Keep walking...",
    "1210402": "Nice summary. I have a doubt though, what do you mean by \"blending  2-3 best epochs\"?\nI know blending is an ensemble technique but I can't understand blending epochs.\nI presume you used pytorch. correct me if I am wrong.",
    "1210423": "Nice work, congrats 🎉\n\nI have blended several epochs also, mainly from fold2 since seed 42 gave really good results on fold2. Fold2: Epoch 8 - > 0.901, epoch  10 -> 0.905 , Fold1: epoch 11: 0.900 , epoch 13: 0.901. This approach costed me shake down for 1100+ places 😅. Ensemble of this model gave LB of 0.904, with light TTA, light augmentations, ELR loss (noise 0.4, beta: 2), EfficientNetB4_ns, , 512 image size,  2020 data, Adam optimizer, step lr (step size 10) for 15 epochs. I used pixel rescaling of the images, it gave me the best results tho. Where did I failed I don't know. I trusted my cv also...",
    "1210452": "Продължаваме напред! 😄( мен много ме удари шейкъпа, прайвъта беше като биткоин ) / Keep walking, strong approach! 😀 ( in contrast I was heavily hit by the shake-up )",
    "1210460": "Hey, usually when you predict from cross validated model, you take the best epoch(weights) for each fold, predict from it - and then average the results to create the final prediction. You can make more predictions from each fold though - like instead of one(best) - take the 2-3 best epochs(weights), predict from them and average at the end. This is kind of \"self blending\" the same model. For example if you have 5 fold CV - and you choose and predict from 3 best epochs for each fold - you will end up with averaging 15 predictions at the end. That will be of course more stable than averaging just 5 predictions (one (best epoch) from each fold). I achieved pretty much the same result though by training the model on the full data (not using CV) - and then averaging the predictions from the 3-4 best epochs. Basically achieving stability with 3-4 predictions rather than 15. I still used CV to guess the performance of the model and choose the (probably) best epochs to blend. Is this a good idea - depends on the particular case of course. But is a valid approach. Please let me know if something is still unclear :)",
    "1210466": "And yes I use pytorch",
    "1210484": "Определено беше доста ... нестабилно (за да не използвам по-силни думи :D ). Не, че е нещо необичайно напоследък. За мен беше (за съжаление) по-скоро пропусната възможност :( Моделът, който щеше да ми донесе злато беше сред финалните ми кандидати (и се представяше най-добре сред тях). Аз обаче не му повярвах - все ми се струваше, че следващият най-добър е по-стабилен. И така приключих със сребро ... :( Keep walking...",
    "1210486": "Don't think you failed mate - you did it all correct. It was the super noisy (and unstable) data unfortunately.",
    "1210492": "Thank you mate for the insights.  First competition so far, I really learn a lot anyways. Lessons learned, for many more competitions !",
    "1210729": "Did not have an idea we can do this kind of self blending. Thank you very much for the thorough explanation , this was very informative, something to try out in the future I guess.",
    "1210730": "I need to seriously switch to pytorch, keras just doesn't seem good enough, every new paper implementation is in pytorch now, even the google guys won't use TF2.",
    "1211013": "This is for sure the best part of the Kaggle journey - learning new stuff :) Keep it up and good luck!",
    "1211015": "Same mate, wish you all the best also ^_^ !",
    "1211034": "Now honestly it depends (re Keras/Tensorflow). Pytorch is very good for research/experimentation/quick&dirty iterations - which makes it of course super fit for Kaggle. But is it that good for production/real-world applications (where performance/stability/scalability are crucial) - this is arguable (in my opinion). For example, I use ONLY TF in my job and also for my side projects. My point is - the picture you see here on Kaggle is a bit biased (re technologies in use/popularity) - and doesn't necessarily reflect the real world situation. TF is dominating the industry with a huge margin - and for a good reason :) (just like pytorch is dominating the research/academic fields). Good luck with your next competition - cheers!"
  },
  "source": "meta"
}