{
  "id": 214959,
  "title": "random seeding and ensemble tricks that every kagglers should know",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/214959",
  "author_name": "",
  "post_date": "2021-01-28T06:17:10.209791Z",
  "votes": 144,
  "comment_count": 27,
  "views": 0,
  "content": "<p>you show know WHY …</p>\n<p><img src=\"https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Ensemble_Figre2_updated-1024x532.jpg\" alt=\"\"></p>\n<p><img src=\"https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure1_esemble-blog-1024x388.jpg\" alt=\"\"></p>\n<p><a href=\"https://www.microsoft.com/en-us/research/blog/three-mysteries-in-deep-learning-ensemble-knowledge-distillation-and-self-distillation/\" target=\"_blank\">https://www.microsoft.com/en-us/research/blog/three-mysteries-in-deep-learning-ensemble-knowledge-distillation-and-self-distillation/</a></p>\n<p>Three mysteries in deep learning: Ensemble, knowledge distillation, and self-distillation</p>",
  "messages": [
    {
      "id": "1173864",
      "postDate": "01/28/2021 06:17:10",
      "content": "<p>you show know WHY …</p>\n<p><img src=\"https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Ensemble_Figre2_updated-1024x532.jpg\" alt=\"\"></p>\n<p><img src=\"https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure1_esemble-blog-1024x388.jpg\" alt=\"\"></p>\n<p><a href=\"https://www.microsoft.com/en-us/research/blog/three-mysteries-in-deep-learning-ensemble-knowledge-distillation-and-self-distillation/\" target=\"_blank\">https://www.microsoft.com/en-us/research/blog/three-mysteries-in-deep-learning-ensemble-knowledge-distillation-and-self-distillation/</a></p>\n<p>Three mysteries in deep learning: Ensemble, knowledge distillation, and self-distillation</p>",
      "rawMarkdown": "you show know WHY ...\n\n![](https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Ensemble_Figre2_updated-1024x532.jpg)\n\n![](https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure1_esemble-blog-1024x388.jpg)\n\nhttps://www.microsoft.com/en-us/research/blog/three-mysteries-in-deep-learning-ensemble-knowledge-distillation-and-self-distillation/\n\nThree mysteries in deep learning: Ensemble, knowledge distillation, and self-distillation",
      "votes": null
    },
    {
      "id": "1173874",
      "postDate": "01/28/2021 06:22:16",
      "content": "<p>another trick that i common used is to train a resnet to match results of efficient-net (i.e. different architecture).<br>\nactually, I used to explain this phenomenon as bias shifting (my own unproven theory)</p>\n<ul>\n<li><p>additional minimizing anything else than the data loss is good (regularisation). best parameters for best test loss is close to that of best train loss but these two sets of parameters are not equal (unless perfect distribution).</p></li>\n<li><p>if there is domain shift, then there is also parameters shift</p></li>\n</ul>",
      "rawMarkdown": "another trick that i common used is to train a resnet to match results of efficient-net (i.e. different architecture).\nactually, I used to explain this phenomenon as bias shifting (my own unproven theory)\n\n- additional minimizing anything else than the data loss is good (regularisation). best parameters for best test loss is close to that of best train loss but these two sets of parameters are not equal (unless perfect distribution).\n\n- if there is domain shift, then there is also parameters shift",
      "votes": null
    },
    {
      "id": "1173947",
      "postDate": "01/28/2021 07:32:11",
      "content": "<p>Interesting, thanks for sharing.</p>",
      "rawMarkdown": "Interesting, thanks for sharing.",
      "votes": null
    },
    {
      "id": "1174098",
      "postDate": "01/28/2021 09:11:16",
      "content": "<p>Very interesting reading!</p>\n<p>Regarding mistery-1, it doesn't make sense at all (for me)</p>\n<p>Could it be that they…?:</p>\n<ul>\n<li><p>ensured the same different seed for each network between experiments</p></li>\n<li><p>ensured the same image order in DataLoaders</p></li>\n<li><p>forgot to set the seed for augmentation</p></li>\n</ul>\n<p>So, in the \"altogether\" experiments, the 10 networks with different intilizations, are being trained on the same images and augmentations.</p>\n<p>But the in the \"separately\" experiments, the 10 networks with different intilizations, are taking advantage of being trained on the same images but different augmentations.</p>",
      "rawMarkdown": "Very interesting reading!\n\nRegarding mistery-1, it doesn't make sense at all (for me)\n\nCould it be that they...?:\n\n  - ensured the same different seed for each network between experiments\n\n  - ensured the same image order in DataLoaders\n\n  - forgot to set the seed for augmentation\n\nSo, in the \"altogether\" experiments, the 10 networks with different intilizations, are being trained on the same images and augmentations.\n\nBut the in the \"separately\" experiments, the 10 networks with different intilizations, are taking advantage of being trained on the same images but different augmentations.",
      "votes": null
    },
    {
      "id": "1174181",
      "postDate": "01/28/2021 10:20:58",
      "content": "<p>It's not about augmentations, they (probably) didn't use any and if they did, it's unimportant to the outcome. </p>",
      "rawMarkdown": "It's not about augmentations, they (probably) didn't use any and if they did, it's unimportant to the outcome.",
      "votes": null
    },
    {
      "id": "1174323",
      "postDate": "01/28/2021 12:08:04",
      "content": "<p>Sorry what did you meant by mack ?</p>",
      "rawMarkdown": "Sorry what did you meant by mack ?",
      "votes": null
    },
    {
      "id": "1174326",
      "postDate": "01/28/2021 12:09:31",
      "content": "<p>sorry for the typo, it should have been \"match\"</p>",
      "rawMarkdown": "sorry for the typo, it should have been \"match\"",
      "votes": null
    },
    {
      "id": "1174411",
      "postDate": "01/28/2021 13:06:34",
      "content": "<p>Thanks, it was to be sure x) </p>\n<p>So, you train the resnet on the label output of the efficient net ? (edit : training the resnet on x : the images and y not beeing the reel labels but the prediction of the effnet).</p>",
      "rawMarkdown": "Thanks, it was to be sure x) \n\nSo, you train the resnet on the label output of the efficient net ? (edit : training the resnet on x : the images and y not beeing the reel labels but the prediction of the effnet).",
      "votes": null
    },
    {
      "id": "1174461",
      "postDate": "01/28/2021 13:44:46",
      "content": "<p>Btw, thanks for the paper, I wanted to go look for what was \"distillation knowledge\", after reading the solutions from 1st to 4th on the Plant Pathology Competition.</p>",
      "rawMarkdown": "Btw, thanks for the paper, I wanted to go look for what was \"distillation knowledge\", after reading the solutions from 1st to 4th on the Plant Pathology Competition.",
      "votes": null
    },
    {
      "id": "1175160",
      "postDate": "01/29/2021 01:07:18",
      "content": "<p>cool cool cool</p>",
      "rawMarkdown": "cool cool cool",
      "votes": null
    },
    {
      "id": "1175244",
      "postDate": "01/29/2021 03:09:19",
      "content": "<p>In ensemble learning is all about Synergy and Redundancy (terms taken from information theory). If base models have Synergy then lead to performance boosts. However is very hard to measure and define this term here. Its just easier to try a trial on test data and make statistics. </p>\n<p>In other simble words, in ensembes we just need the base models to see diferent things into the dataset (they must be different characters in which each one fill the weakneses of the other). By training independnly each model, then the model is free to \"think out of the box\" and take its own unique individual path. But training all model together, then we lose the independency of learning and each model affect the other and learn same things. </p>\n<p>This is my qualititave analysis of ensemble learning. </p>",
      "rawMarkdown": "In ensemble learning is all about Synergy and Redundancy (terms taken from information theory). If base models have Synergy then lead to performance boosts. However is very hard to measure and define this term here. Its just easier to try a trial on test data and make statistics. \n\nIn other simble words, in ensembes we just need the base models to see diferent things into the dataset (they must be different characters in which each one fill the weakneses of the other). By training independnly each model, then the model is free to \"think out of the box\" and take its own unique individual path. But training all model together, then we lose the independency of learning and each model affect the other and learn same things. \n\nThis is my qualititave analysis of ensemble learning.",
      "votes": null
    },
    {
      "id": "1175759",
      "postDate": "01/29/2021 10:30:05",
      "content": "<p>Nice discuss</p>",
      "rawMarkdown": "Nice discuss",
      "votes": null
    },
    {
      "id": "1175844",
      "postDate": "01/29/2021 11:26:35",
      "content": "<p>Hey, can you give a little more detail on how the idependent training is different from training together?</p>",
      "rawMarkdown": "Hey, can you give a little more detail on how the idependent training is different from training together?",
      "votes": null
    },
    {
      "id": "1176661",
      "postDate": "01/29/2021 18:00:57",
      "content": "<p>This is all about ensemble learning. In ensemble learning you need <strong>diferent</strong> models which will provide something <strong>new</strong>  to the ensemble; just like when a coach builds a team. You need each player provide something new to the team, that all the other players cannot provide. Its all about synergy. In idependent training, each model learns its own unique patterns and so they become diferent. Training together, each model affects the other and so they conclude to learn almost the same patterns. Thus in this case the synergy will be less since their diference factor will be also low.  </p>",
      "rawMarkdown": "This is all about ensemble learning. In ensemble learning you need **diferent** models which will provide something **new**  to the ensemble; just like when a coach builds a team. You need each player provide something new to the team, that all the other players cannot provide. Its all about synergy. In idependent training, each model learns its own unique patterns and so they become diferent. Training together, each model affects the other and so they conclude to learn almost the same patterns. Thus in this case the synergy will be less since their diference factor will be also low.",
      "votes": null
    },
    {
      "id": "1176796",
      "postDate": "01/29/2021 19:37:14",
      "content": "<p>I'm asking a bit more specific: you don't actually need different models but just have to train the same model many times separately as the article suggests. Now what do you input in these separate trainings other than training models \"together\"? (i'm really asking about what is stated in the article)</p>",
      "rawMarkdown": "I'm asking a bit more specific: you don't actually need different models but just have to train the same model many times separately as the article suggests. Now what do you input in these separate trainings other than training models \"together\"? (i'm really asking about what is stated in the article)",
      "votes": null
    },
    {
      "id": "1176805",
      "postDate": "01/29/2021 19:43:22",
      "content": "<p>When i say \"diferrent\" i dont mean diferent models like resnet and densenet, but it could be resnet1 and resnet2 which just trained on diferent image resolutions or diferent seeds. The term \"diferent\" means that each model learned diferent patterns somthing very essential when build an ensemble.</p>",
      "rawMarkdown": "When i say \"diferrent\" i dont mean diferent models like resnet and densenet, but it could be resnet1 and resnet2 which just trained on diferent image resolutions or diferent seeds. The term \"diferent\" means that each model learned diferent patterns somthing very essential when build an ensemble.",
      "votes": null
    },
    {
      "id": "1176824",
      "postDate": "01/29/2021 20:18:58",
      "content": "<p>Yes but what is actually the difference here between green and orange? <img src=\"https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure1_esemble-blog-1024x388.jpg\" alt=\"\"></p>\n<p>they're both trained on different seed each time</p>",
      "rawMarkdown": "Yes but what is actually the difference here between green and orange? ![](https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure1_esemble-blog-1024x388.jpg)\n\nthey're both trained on different seed each time",
      "votes": null
    },
    {
      "id": "1176863",
      "postDate": "01/29/2021 21:04:10",
      "content": "<p>I guess this is something like multi-parallel training. In this concept an activation function (e.g. sigmoid) will be in the ouput and update its weigths with respect to the base CNN models outputs which will be trained in this time too.</p>",
      "rawMarkdown": "I guess this is something like multi-parallel training. In this concept an activation function (e.g. sigmoid) will be in the ouput and update its weigths with respect to the base CNN models outputs which will be trained in this time too.",
      "votes": null
    },
    {
      "id": "1176897",
      "postDate": "01/29/2021 21:44:29",
      "content": "<p>Really interesting, regarding mystery 1. Assuming that the models 1…10 are not identical, training 10 models next to each other is a little bit like increasing the number of parameters of the model, right? So its logical that the model can learn more complex functions which increases predictive performance.</p>",
      "rawMarkdown": "Really interesting, regarding mystery 1. Assuming that the models 1...10 are not identical, training 10 models next to each other is a little bit like increasing the number of parameters of the model, right? So its logical that the model can learn more complex functions which increases predictive performance.",
      "votes": null
    },
    {
      "id": "1177216",
      "postDate": "01/30/2021 06:27:00",
      "content": "<p>IMHO, the difference between green and orange is</p>\n<p>green:</p>\n<ul>\n<li>You have 10 models with different initializations (i.e. different seeds) at the same time. It means that you define 10 models in a script.</li>\n<li>train these models together by using loss((F1+F2+…+F10)/10, gt).</li>\n<li>output: (F1+F2+…+F10)/10</li>\n</ul>\n<p>orange (I think this is a normal ensemble method):</p>\n<ul>\n<li>You have 10 models with different initializations (i.e. different seeds) separately. It means that you have a model in a script and execute it 10 times.</li>\n<li>train each model by using loss(Fi, gt).</li>\n<li>output (by taking the average): (F1+F2+…+F10)/10</li>\n</ul>",
      "rawMarkdown": "IMHO, the difference between green and orange is\n\ngreen:\n- You have 10 models with different initializations (i.e. different seeds) at the same time. It means that you define 10 models in a script.\n- train these models together by using loss((F1+F2+...+F10)/10, gt).\n- output: (F1+F2+...+F10)/10\n\norange (I think this is a normal ensemble method):\n- You have 10 models with different initializations (i.e. different seeds) separately. It means that you have a model in a script and execute it 10 times.\n- train each model by using loss(Fi, gt).\n- output (by taking the average): (F1+F2+...+F10)/10",
      "votes": null
    },
    {
      "id": "1180353",
      "postDate": "02/01/2021 08:14:49",
      "content": "<p>Insightful discuss</p>",
      "rawMarkdown": "Insightful discuss",
      "votes": null
    },
    {
      "id": "1180485",
      "postDate": "02/01/2021 09:35:46",
      "content": "<p>OK so it's basically about a distinct loss function for each model</p>",
      "rawMarkdown": "OK so it's basically about a distinct loss function for each model",
      "votes": null
    },
    {
      "id": "1180575",
      "postDate": "02/01/2021 10:49:06",
      "content": "<p>I am preparing the code to do the knowledge distillation. I found one implementation on the website of keras itself : <a href=\"https://keras.io/examples/vision/knowledge_distillation/#construct-distiller-class\" target=\"_blank\">https://keras.io/examples/vision/knowledge_distillation/#construct-distiller-class</a></p>\n<p>But, the code is meant for only one teacher. And since there is a loss that propagate from the teacher to the student, I don't understand how you can do this with ensembling.</p>",
      "rawMarkdown": "I am preparing the code to do the knowledge distillation. I found one implementation on the website of keras itself : https://keras.io/examples/vision/knowledge_distillation/#construct-distiller-class\n\nBut, the code is meant for only one teacher. And since there is a loss that propagate from the teacher to the student, I don't understand how you can do this with ensembling.",
      "votes": null
    },
    {
      "id": "1183454",
      "postDate": "02/03/2021 02:29:30",
      "content": "<p>Good one 🙌</p>",
      "rawMarkdown": "Good one 🙌",
      "votes": null
    },
    {
      "id": "1183482",
      "postDate": "02/03/2021 03:24:32",
      "content": "<p>\"OK so it's basically about a distinct loss function for each model\"</p>\n<p>No. It is not the loss function.</p>\n<p>green:</p>\n<ul>\n<li>train together and each model see the same training data in one batch as the other.</li>\n</ul>\n<p>orange:</p>\n<ul>\n<li>train separately and each model see different data in one batch. (different seed)</li>\n</ul>",
      "rawMarkdown": "\"OK so it's basically about a distinct loss function for each model\"\n\nNo. It is not the loss function.\n\ngreen:\n- train together and each model see the same training data in one batch as the other.\n\norange:\n- train separately and each model see different data in one batch. (different seed)",
      "votes": null
    },
    {
      "id": "1184619",
      "postDate": "02/03/2021 16:08:19",
      "content": "<p>Where could I read about the use of information theory on neural networks? Thanks</p>",
      "rawMarkdown": "Where could I read about the use of information theory on neural networks? Thanks",
      "votes": null
    },
    {
      "id": "1188101",
      "postDate": "02/06/2021 00:47:47",
      "content": "<p>For mystery 2, do you need it to be the exact same model ? How do it work when performing cross validation ? should you ensemble model from different folds ?</p>",
      "rawMarkdown": "For mystery 2, do you need it to be the exact same model ? How do it work when performing cross validation ? should you ensemble model from different folds ?",
      "votes": null
    },
    {
      "id": "1189023",
      "postDate": "02/06/2021 16:46:06",
      "content": "<p>Neat, thanks 👍</p>",
      "rawMarkdown": "Neat, thanks 👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1173874,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/28/2021 06:22:16",
      "content": "<p>another trick that i common used is to train a resnet to match results of efficient-net (i.e. different architecture).<br>\nactually, I used to explain this phenomenon as bias shifting (my own unproven theory)</p>\n<ul>\n<li><p>additional minimizing anything else than the data loss is good (regularisation). best parameters for best test loss is close to that of best train loss but these two sets of parameters are not equal (unless perfect distribution).</p></li>\n<li><p>if there is domain shift, then there is also parameters shift</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1174323,
          "author_name": "cdk292",
          "author_url": "",
          "post_date": "01/28/2021 12:08:04",
          "content": "<p>Sorry what did you meant by mack ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1174326,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/28/2021 12:09:31",
          "content": "<p>sorry for the typo, it should have been \"match\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1174411,
          "author_name": "cdk292",
          "author_url": "",
          "post_date": "01/28/2021 13:06:34",
          "content": "<p>Thanks, it was to be sure x) </p>\n<p>So, you train the resnet on the label output of the efficient net ? (edit : training the resnet on x : the images and y not beeing the reel labels but the prediction of the effnet).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1174461,
          "author_name": "cdk292",
          "author_url": "",
          "post_date": "01/28/2021 13:44:46",
          "content": "<p>Btw, thanks for the paper, I wanted to go look for what was \"distillation knowledge\", after reading the solutions from 1st to 4th on the Plant Pathology Competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1173947,
      "author_name": "yizhitao",
      "author_url": "",
      "post_date": "01/28/2021 07:32:11",
      "content": "<p>Interesting, thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1174098,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "01/28/2021 09:11:16",
      "content": "<p>Very interesting reading!</p>\n<p>Regarding mistery-1, it doesn't make sense at all (for me)</p>\n<p>Could it be that they…?:</p>\n<ul>\n<li><p>ensured the same different seed for each network between experiments</p></li>\n<li><p>ensured the same image order in DataLoaders</p></li>\n<li><p>forgot to set the seed for augmentation</p></li>\n</ul>\n<p>So, in the \"altogether\" experiments, the 10 networks with different intilizations, are being trained on the same images and augmentations.</p>\n<p>But the in the \"separately\" experiments, the 10 networks with different intilizations, are taking advantage of being trained on the same images but different augmentations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1174181,
          "author_name": "alexanderriedel",
          "author_url": "",
          "post_date": "01/28/2021 10:20:58",
          "content": "<p>It's not about augmentations, they (probably) didn't use any and if they did, it's unimportant to the outcome. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1175160,
      "author_name": "alincijov",
      "author_url": "",
      "post_date": "01/29/2021 01:07:18",
      "content": "<p>cool cool cool</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1175244,
      "author_name": "",
      "author_url": "",
      "post_date": "01/29/2021 03:09:19",
      "content": "<p>In ensemble learning is all about Synergy and Redundancy (terms taken from information theory). If base models have Synergy then lead to performance boosts. However is very hard to measure and define this term here. Its just easier to try a trial on test data and make statistics. </p>\n<p>In other simble words, in ensembes we just need the base models to see diferent things into the dataset (they must be different characters in which each one fill the weakneses of the other). By training independnly each model, then the model is free to \"think out of the box\" and take its own unique individual path. But training all model together, then we lose the independency of learning and each model affect the other and learn same things. </p>\n<p>This is my qualititave analysis of ensemble learning. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1175844,
          "author_name": "alexanderriedel",
          "author_url": "",
          "post_date": "01/29/2021 11:26:35",
          "content": "<p>Hey, can you give a little more detail on how the idependent training is different from training together?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1176661,
          "author_name": "",
          "author_url": "",
          "post_date": "01/29/2021 18:00:57",
          "content": "<p>This is all about ensemble learning. In ensemble learning you need <strong>diferent</strong> models which will provide something <strong>new</strong>  to the ensemble; just like when a coach builds a team. You need each player provide something new to the team, that all the other players cannot provide. Its all about synergy. In idependent training, each model learns its own unique patterns and so they become diferent. Training together, each model affects the other and so they conclude to learn almost the same patterns. Thus in this case the synergy will be less since their diference factor will be also low.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1176796,
          "author_name": "alexanderriedel",
          "author_url": "",
          "post_date": "01/29/2021 19:37:14",
          "content": "<p>I'm asking a bit more specific: you don't actually need different models but just have to train the same model many times separately as the article suggests. Now what do you input in these separate trainings other than training models \"together\"? (i'm really asking about what is stated in the article)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1176805,
          "author_name": "",
          "author_url": "",
          "post_date": "01/29/2021 19:43:22",
          "content": "<p>When i say \"diferrent\" i dont mean diferent models like resnet and densenet, but it could be resnet1 and resnet2 which just trained on diferent image resolutions or diferent seeds. The term \"diferent\" means that each model learned diferent patterns somthing very essential when build an ensemble.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1176824,
          "author_name": "alexanderriedel",
          "author_url": "",
          "post_date": "01/29/2021 20:18:58",
          "content": "<p>Yes but what is actually the difference here between green and orange? <img src=\"https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure1_esemble-blog-1024x388.jpg\" alt=\"\"></p>\n<p>they're both trained on different seed each time</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1176863,
          "author_name": "",
          "author_url": "",
          "post_date": "01/29/2021 21:04:10",
          "content": "<p>I guess this is something like multi-parallel training. In this concept an activation function (e.g. sigmoid) will be in the ouput and update its weigths with respect to the base CNN models outputs which will be trained in this time too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1177216,
          "author_name": "yukia18",
          "author_url": "",
          "post_date": "01/30/2021 06:27:00",
          "content": "<p>IMHO, the difference between green and orange is</p>\n<p>green:</p>\n<ul>\n<li>You have 10 models with different initializations (i.e. different seeds) at the same time. It means that you define 10 models in a script.</li>\n<li>train these models together by using loss((F1+F2+…+F10)/10, gt).</li>\n<li>output: (F1+F2+…+F10)/10</li>\n</ul>\n<p>orange (I think this is a normal ensemble method):</p>\n<ul>\n<li>You have 10 models with different initializations (i.e. different seeds) separately. It means that you have a model in a script and execute it 10 times.</li>\n<li>train each model by using loss(Fi, gt).</li>\n<li>output (by taking the average): (F1+F2+…+F10)/10</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1180485,
          "author_name": "alexanderriedel",
          "author_url": "",
          "post_date": "02/01/2021 09:35:46",
          "content": "<p>OK so it's basically about a distinct loss function for each model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1183482,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/03/2021 03:24:32",
          "content": "<p>\"OK so it's basically about a distinct loss function for each model\"</p>\n<p>No. It is not the loss function.</p>\n<p>green:</p>\n<ul>\n<li>train together and each model see the same training data in one batch as the other.</li>\n</ul>\n<p>orange:</p>\n<ul>\n<li>train separately and each model see different data in one batch. (different seed)</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1184619,
          "author_name": "emmanuelmess",
          "author_url": "",
          "post_date": "02/03/2021 16:08:19",
          "content": "<p>Where could I read about the use of information theory on neural networks? Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1175759,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "01/29/2021 10:30:05",
      "content": "<p>Nice discuss</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1176897,
      "author_name": "dunky11",
      "author_url": "",
      "post_date": "01/29/2021 21:44:29",
      "content": "<p>Really interesting, regarding mystery 1. Assuming that the models 1…10 are not identical, training 10 models next to each other is a little bit like increasing the number of parameters of the model, right? So its logical that the model can learn more complex functions which increases predictive performance.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1180353,
      "author_name": "yihengpeng",
      "author_url": "",
      "post_date": "02/01/2021 08:14:49",
      "content": "<p>Insightful discuss</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1180575,
      "author_name": "cdk292",
      "author_url": "",
      "post_date": "02/01/2021 10:49:06",
      "content": "<p>I am preparing the code to do the knowledge distillation. I found one implementation on the website of keras itself : <a href=\"https://keras.io/examples/vision/knowledge_distillation/#construct-distiller-class\" target=\"_blank\">https://keras.io/examples/vision/knowledge_distillation/#construct-distiller-class</a></p>\n<p>But, the code is meant for only one teacher. And since there is a loss that propagate from the teacher to the student, I don't understand how you can do this with ensembling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1183454,
      "author_name": "mightyak",
      "author_url": "",
      "post_date": "02/03/2021 02:29:30",
      "content": "<p>Good one 🙌</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1188101,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "02/06/2021 00:47:47",
      "content": "<p>For mystery 2, do you need it to be the exact same model ? How do it work when performing cross validation ? should you ensemble model from different folds ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1189023,
      "author_name": "haydn56",
      "author_url": "",
      "post_date": "02/06/2021 16:46:06",
      "content": "<p>Neat, thanks 👍</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1173864": "you show know WHY ...\n\n![](https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Ensemble_Figre2_updated-1024x532.jpg)\n\n![](https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure1_esemble-blog-1024x388.jpg)\n\nhttps://www.microsoft.com/en-us/research/blog/three-mysteries-in-deep-learning-ensemble-knowledge-distillation-and-self-distillation/\n\nThree mysteries in deep learning: Ensemble, knowledge distillation, and self-distillation",
    "1173874": "another trick that i common used is to train a resnet to match results of efficient-net (i.e. different architecture).\nactually, I used to explain this phenomenon as bias shifting (my own unproven theory)\n\n- additional minimizing anything else than the data loss is good (regularisation). best parameters for best test loss is close to that of best train loss but these two sets of parameters are not equal (unless perfect distribution).\n\n- if there is domain shift, then there is also parameters shift",
    "1173947": "Interesting, thanks for sharing.",
    "1174098": "Very interesting reading!\n\nRegarding mistery-1, it doesn't make sense at all (for me)\n\nCould it be that they...?:\n\n  - ensured the same different seed for each network between experiments\n\n  - ensured the same image order in DataLoaders\n\n  - forgot to set the seed for augmentation\n\nSo, in the \"altogether\" experiments, the 10 networks with different intilizations, are being trained on the same images and augmentations.\n\nBut the in the \"separately\" experiments, the 10 networks with different intilizations, are taking advantage of being trained on the same images but different augmentations.",
    "1174181": "It's not about augmentations, they (probably) didn't use any and if they did, it's unimportant to the outcome.",
    "1174323": "Sorry what did you meant by mack ?",
    "1174326": "sorry for the typo, it should have been \"match\"",
    "1174411": "Thanks, it was to be sure x) \n\nSo, you train the resnet on the label output of the efficient net ? (edit : training the resnet on x : the images and y not beeing the reel labels but the prediction of the effnet).",
    "1174461": "Btw, thanks for the paper, I wanted to go look for what was \"distillation knowledge\", after reading the solutions from 1st to 4th on the Plant Pathology Competition.",
    "1175160": "cool cool cool",
    "1175244": "In ensemble learning is all about Synergy and Redundancy (terms taken from information theory). If base models have Synergy then lead to performance boosts. However is very hard to measure and define this term here. Its just easier to try a trial on test data and make statistics. \n\nIn other simble words, in ensembes we just need the base models to see diferent things into the dataset (they must be different characters in which each one fill the weakneses of the other). By training independnly each model, then the model is free to \"think out of the box\" and take its own unique individual path. But training all model together, then we lose the independency of learning and each model affect the other and learn same things. \n\nThis is my qualititave analysis of ensemble learning.",
    "1175759": "Nice discuss",
    "1175844": "Hey, can you give a little more detail on how the idependent training is different from training together?",
    "1176661": "This is all about ensemble learning. In ensemble learning you need **diferent** models which will provide something **new**  to the ensemble; just like when a coach builds a team. You need each player provide something new to the team, that all the other players cannot provide. Its all about synergy. In idependent training, each model learns its own unique patterns and so they become diferent. Training together, each model affects the other and so they conclude to learn almost the same patterns. Thus in this case the synergy will be less since their diference factor will be also low.",
    "1176796": "I'm asking a bit more specific: you don't actually need different models but just have to train the same model many times separately as the article suggests. Now what do you input in these separate trainings other than training models \"together\"? (i'm really asking about what is stated in the article)",
    "1176805": "When i say \"diferrent\" i dont mean diferent models like resnet and densenet, but it could be resnet1 and resnet2 which just trained on diferent image resolutions or diferent seeds. The term \"diferent\" means that each model learned diferent patterns somthing very essential when build an ensemble.",
    "1176824": "Yes but what is actually the difference here between green and orange? ![](https://www.microsoft.com/en-us/research/uploads/prod/2021/01/Figure1_esemble-blog-1024x388.jpg)\n\nthey're both trained on different seed each time",
    "1176863": "I guess this is something like multi-parallel training. In this concept an activation function (e.g. sigmoid) will be in the ouput and update its weigths with respect to the base CNN models outputs which will be trained in this time too.",
    "1176897": "Really interesting, regarding mystery 1. Assuming that the models 1...10 are not identical, training 10 models next to each other is a little bit like increasing the number of parameters of the model, right? So its logical that the model can learn more complex functions which increases predictive performance.",
    "1177216": "IMHO, the difference between green and orange is\n\ngreen:\n- You have 10 models with different initializations (i.e. different seeds) at the same time. It means that you define 10 models in a script.\n- train these models together by using loss((F1+F2+...+F10)/10, gt).\n- output: (F1+F2+...+F10)/10\n\norange (I think this is a normal ensemble method):\n- You have 10 models with different initializations (i.e. different seeds) separately. It means that you have a model in a script and execute it 10 times.\n- train each model by using loss(Fi, gt).\n- output (by taking the average): (F1+F2+...+F10)/10",
    "1180353": "Insightful discuss",
    "1180485": "OK so it's basically about a distinct loss function for each model",
    "1180575": "I am preparing the code to do the knowledge distillation. I found one implementation on the website of keras itself : https://keras.io/examples/vision/knowledge_distillation/#construct-distiller-class\n\nBut, the code is meant for only one teacher. And since there is a loss that propagate from the teacher to the student, I don't understand how you can do this with ensembling.",
    "1183454": "Good one 🙌",
    "1183482": "\"OK so it's basically about a distinct loss function for each model\"\n\nNo. It is not the loss function.\n\ngreen:\n- train together and each model see the same training data in one batch as the other.\n\norange:\n- train separately and each model see different data in one batch. (different seed)",
    "1184619": "Where could I read about the use of information theory on neural networks? Thanks",
    "1188101": "For mystery 2, do you need it to be the exact same model ? How do it work when performing cross validation ? should you ensemble model from different folds ?",
    "1189023": "Neat, thanks 👍"
  },
  "source": "meta"
}