{
  "id": 166399,
  "title": "Ensemble is the key",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/166399",
  "author_name": "",
  "post_date": "2020-07-12T19:34:48.317974500Z",
  "votes": 11,
  "comment_count": 17,
  "views": 0,
  "content": "<p>In this phase, everyone is thinking about how to get the most from his/her models.</p>\n\n<p>In my own experience, the most significant results are coming from ensembles. Today I was able to go from 270 to 134 place with a small improvement in one of my models. </p>\n\n<p>I went from 0.924 to 0.931. Seems irrelevant. But this brought the final ensemble score to 0.949.</p>\n\n<p>These are the score for my single models:\n- 0.936, 0.931, 0.929, 0.920 </p>\n\n<p>and the weighted ensemble score is 0.949. Much higher than single score. \nI think I can improve the 0.920 (it is the one related to 768x768, so longer training and difficult to stay inside 3 hours).</p>\n\n<p>In my working experience, few customers think about ensembling. They have a \"Theoretical approach\" and think that ensembling is a trick. This is really interesting experience.</p>",
  "messages": [
    {
      "id": "926569",
      "postDate": "07/12/2020 19:34:48",
      "content": "<p>In this phase, everyone is thinking about how to get the most from his/her models.</p>\n\n<p>In my own experience, the most significant results are coming from ensembles. Today I was able to go from 270 to 134 place with a small improvement in one of my models. </p>\n\n<p>I went from 0.924 to 0.931. Seems irrelevant. But this brought the final ensemble score to 0.949.</p>\n\n<p>These are the score for my single models:\n- 0.936, 0.931, 0.929, 0.920 </p>\n\n<p>and the weighted ensemble score is 0.949. Much higher than single score. \nI think I can improve the 0.920 (it is the one related to 768x768, so longer training and difficult to stay inside 3 hours).</p>\n\n<p>In my working experience, few customers think about ensembling. They have a \"Theoretical approach\" and think that ensembling is a trick. This is really interesting experience.</p>",
      "rawMarkdown": "In this phase, everyone is thinking about how to get the most from his/her models.\n\nIn my own experience, the most significant results are coming from ensembles. Today I was able to go from 270 to 134 place with a small improvement in one of my models. \n\nI went from 0.924 to 0.931. Seems irrelevant. But this brought the final ensemble score to 0.949.\n\nThese are the score for my single models:\n- 0.936, 0.931, 0.929, 0.920 \n\nand the weighted ensemble score is 0.949. Much higher than single score. \nI think I can improve the 0.920 (it is the one related to 768x768, so longer training and difficult to stay inside 3 hours).\n\nIn my working experience, few customers think about ensembling. They have a \"Theoretical approach\" and think that ensembling is a trick. This is really interesting experience.",
      "votes": null
    },
    {
      "id": "926583",
      "postDate": "07/12/2020 19:43:55",
      "content": "<p>Wow! nice score, didnt know ensembling increases the score that much, you give me hope haha</p>",
      "rawMarkdown": "Wow! nice score, didnt know ensembling increases the score that much, you give me hope haha",
      "votes": null
    },
    {
      "id": "926591",
      "postDate": "07/12/2020 19:48:19",
      "content": "<p>I use semble too and have 950. But I think that's not good for  Private LB. because you can see if only predict 0.2 % can get 94.9 and the same if can predict 2 %. FInal LB maybe the diff</p>",
      "rawMarkdown": "I use semble too and have 950. But I think that's not good for  Private LB. because you can see if only predict 0.2 % can get 94.9 and the same if can predict 2 %. FInal LB maybe the diff",
      "votes": null
    },
    {
      "id": "926604",
      "postDate": "07/12/2020 19:57:46",
      "content": "<p>Yeah gotta make sure not to overfitt on public lb, gotta check if ensembling helps local CV</p>",
      "rawMarkdown": "Yeah gotta make sure not to overfitt on public lb, gotta check if ensembling helps local CV",
      "votes": null
    },
    {
      "id": "926811",
      "postDate": "07/13/2020 02:52:19",
      "content": "<p>Yes, I am using it too!</p>",
      "rawMarkdown": "Yes, I am using it too!",
      "votes": null
    },
    {
      "id": "927943",
      "postDate": "07/13/2020 16:47:23",
      "content": "<p>Why ensemble and not a nn model to add them together ?</p>",
      "rawMarkdown": "Why ensemble and not a nn model to add them together ?",
      "votes": null
    },
    {
      "id": "928186",
      "postDate": "07/13/2020 19:36:17",
      "content": "<p>well, that's too an ensemble. We enter here in what is called Stacked Ensembles. Combine base learners with a meta-learner. SInce I want to avoid to overfit to the public LB I have decided to avoid things too complex.</p>",
      "rawMarkdown": "well, that's too an ensemble. We enter here in what is called Stacked Ensembles. Combine base learners with a meta-learner. SInce I want to avoid to overfit to the public LB I have decided to avoid things too complex.",
      "votes": null
    },
    {
      "id": "928191",
      "postDate": "07/13/2020 19:38:26",
      "content": "<p>Use Power Averaging: maybe helpfully</p>",
      "rawMarkdown": "Use Power Averaging: maybe helpfully",
      "votes": null
    },
    {
      "id": "929220",
      "postDate": "07/14/2020 14:38:59",
      "content": "<p>Can you please explain this more!</p>",
      "rawMarkdown": "Can you please explain this more!",
      "votes": null
    },
    {
      "id": "929675",
      "postDate": "07/14/2020 20:52:15",
      "content": "<p>Hi <a href=\"/fireheart7\">@fireheart7</a> you can learn more about power averaging from <a href=\"/sirishks\">@sirishks</a> great explanation:</p>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653</a></p>",
      "rawMarkdown": "Hi @fireheart7 you can learn more about power averaging from @sirishks great explanation:\n\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653",
      "votes": null
    },
    {
      "id": "929687",
      "postDate": "07/14/2020 21:04:29",
      "content": "<p>Especially in the private leaderboard, the more generalized solution you have, the better rank you will have. Ensemble is the best way to generalize your solution.</p>",
      "rawMarkdown": "Especially in the private leaderboard, the more generalized solution you have, the better rank you will have. Ensemble is the best way to generalize your solution.",
      "votes": null
    },
    {
      "id": "932802",
      "postDate": "07/17/2020 09:45:34",
      "content": "<p>Hi <a href=\"/luigisaetta\">@luigisaetta</a>  can you help me, create an ensemble ?\ni Have created one model for the tabular data and a second model for the vision data and i would like to add concatenate them via a hard voting or similar.</p>",
      "rawMarkdown": "Hi @luigisaetta  can you help me, create an ensemble ?\ni Have created one model for the tabular data and a second model for the vision data and i would like to add concatenate them via a hard voting or similar.",
      "votes": null
    },
    {
      "id": "934276",
      "postDate": "07/18/2020 10:55:29",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    },
    {
      "id": "934955",
      "postDate": "07/19/2020 00:30:48",
      "content": "<p>Ensemble is the key for most of the competitions if you want to win a higher rank.</p>",
      "rawMarkdown": "Ensemble is the key for most of the competitions if you want to win a higher rank.",
      "votes": null
    },
    {
      "id": "935025",
      "postDate": "07/19/2020 03:10:58",
      "content": "<p>Since I don't do this for a living I know nothing about customers :)</p>\n\n<p>If you do ensemble as the best solution how can that be turned into a working model that gets new input to solve each day?  </p>\n\n<p>Makes sense to ensemble for a competition like Kaggle - but is it  a tool that is only useful on Kaggle?</p>",
      "rawMarkdown": "Since I don't do this for a living I know nothing about customers :)\n\nIf you do ensemble as the best solution how can that be turned into a working model that gets new input to solve each day?  \n\nMakes sense to ensemble for a competition like Kaggle - but is it  a tool that is only useful on Kaggle?",
      "votes": null
    },
    {
      "id": "935071",
      "postDate": "07/19/2020 04:12:21",
      "content": "<p>wait until they over-train and than the shake-up happens :) it's always a surprise when that happens</p>",
      "rawMarkdown": "wait until they over-train and than the shake-up happens :) it's always a surprise when that happens",
      "votes": null
    },
    {
      "id": "935076",
      "postDate": "07/19/2020 04:16:40",
      "content": "<p>You could also use blending (multiply the results from each datasets by weights, than either take the sum with means, median or std deviation)</p>",
      "rawMarkdown": "You could also use blending (multiply the results from each datasets by weights, than either take the sum with means, median or std deviation)",
      "votes": null
    },
    {
      "id": "938789",
      "postDate": "07/21/2020 18:39:12",
      "content": "<p>Hi, to start, you could do a simple weighted average of the predictions given by the single models. It is up to you to decide the weights (AUC from VC,  score from LB).</p>",
      "rawMarkdown": "Hi, to start, you could do a simple weighted average of the predictions given by the single models. It is up to you to decide the weights (AUC from VC,  score from LB).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 926583,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "07/12/2020 19:43:55",
      "content": "<p>Wow! nice score, didnt know ensembling increases the score that much, you give me hope haha</p>",
      "votes": null,
      "replies": [
        {
          "id": 935071,
          "author_name": "alincijov",
          "author_url": "",
          "post_date": "07/19/2020 04:12:21",
          "content": "<p>wait until they over-train and than the shake-up happens :) it's always a surprise when that happens</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 926591,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "07/12/2020 19:48:19",
      "content": "<p>I use semble too and have 950. But I think that's not good for  Private LB. because you can see if only predict 0.2 % can get 94.9 and the same if can predict 2 %. FInal LB maybe the diff</p>",
      "votes": null,
      "replies": [
        {
          "id": 926604,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "07/12/2020 19:57:46",
          "content": "<p>Yeah gotta make sure not to overfitt on public lb, gotta check if ensembling helps local CV</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 926811,
      "author_name": "",
      "author_url": "",
      "post_date": "07/13/2020 02:52:19",
      "content": "<p>Yes, I am using it too!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 927943,
      "author_name": "",
      "author_url": "",
      "post_date": "07/13/2020 16:47:23",
      "content": "<p>Why ensemble and not a nn model to add them together ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 928186,
          "author_name": "luigisaetta",
          "author_url": "",
          "post_date": "07/13/2020 19:36:17",
          "content": "<p>well, that's too an ensemble. We enter here in what is called Stacked Ensembles. Combine base learners with a meta-learner. SInce I want to avoid to overfit to the public LB I have decided to avoid things too complex.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928191,
          "author_name": "doanquanvietnamca",
          "author_url": "",
          "post_date": "07/13/2020 19:38:26",
          "content": "<p>Use Power Averaging: maybe helpfully</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929220,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "07/14/2020 14:38:59",
          "content": "<p>Can you please explain this more!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929675,
          "author_name": "santiviquez",
          "author_url": "",
          "post_date": "07/14/2020 20:52:15",
          "content": "<p>Hi <a href=\"/fireheart7\">@fireheart7</a> you can learn more about power averaging from <a href=\"/sirishks\">@sirishks</a> great explanation:</p>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932802,
          "author_name": "",
          "author_url": "",
          "post_date": "07/17/2020 09:45:34",
          "content": "<p>Hi <a href=\"/luigisaetta\">@luigisaetta</a>  can you help me, create an ensemble ?\ni Have created one model for the tabular data and a second model for the vision data and i would like to add concatenate them via a hard voting or similar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 934276,
          "author_name": "fireheart7",
          "author_url": "",
          "post_date": "07/18/2020 10:55:29",
          "content": "<p>Thank you so much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 935076,
          "author_name": "alincijov",
          "author_url": "",
          "post_date": "07/19/2020 04:16:40",
          "content": "<p>You could also use blending (multiply the results from each datasets by weights, than either take the sum with means, median or std deviation)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 938789,
          "author_name": "luigisaetta",
          "author_url": "",
          "post_date": "07/21/2020 18:39:12",
          "content": "<p>Hi, to start, you could do a simple weighted average of the predictions given by the single models. It is up to you to decide the weights (AUC from VC,  score from LB).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 929687,
      "author_name": "umeyrkl",
      "author_url": "",
      "post_date": "07/14/2020 21:04:29",
      "content": "<p>Especially in the private leaderboard, the more generalized solution you have, the better rank you will have. Ensemble is the best way to generalize your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 934955,
      "author_name": "alincijov",
      "author_url": "",
      "post_date": "07/19/2020 00:30:48",
      "content": "<p>Ensemble is the key for most of the competitions if you want to win a higher rank.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 935025,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/19/2020 03:10:58",
      "content": "<p>Since I don't do this for a living I know nothing about customers :)</p>\n\n<p>If you do ensemble as the best solution how can that be turned into a working model that gets new input to solve each day?  </p>\n\n<p>Makes sense to ensemble for a competition like Kaggle - but is it  a tool that is only useful on Kaggle?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "926569": "In this phase, everyone is thinking about how to get the most from his/her models.\n\nIn my own experience, the most significant results are coming from ensembles. Today I was able to go from 270 to 134 place with a small improvement in one of my models. \n\nI went from 0.924 to 0.931. Seems irrelevant. But this brought the final ensemble score to 0.949.\n\nThese are the score for my single models:\n- 0.936, 0.931, 0.929, 0.920 \n\nand the weighted ensemble score is 0.949. Much higher than single score. \nI think I can improve the 0.920 (it is the one related to 768x768, so longer training and difficult to stay inside 3 hours).\n\nIn my working experience, few customers think about ensembling. They have a \"Theoretical approach\" and think that ensembling is a trick. This is really interesting experience.",
    "926583": "Wow! nice score, didnt know ensembling increases the score that much, you give me hope haha",
    "926591": "I use semble too and have 950. But I think that's not good for  Private LB. because you can see if only predict 0.2 % can get 94.9 and the same if can predict 2 %. FInal LB maybe the diff",
    "926604": "Yeah gotta make sure not to overfitt on public lb, gotta check if ensembling helps local CV",
    "926811": "Yes, I am using it too!",
    "927943": "Why ensemble and not a nn model to add them together ?",
    "928186": "well, that's too an ensemble. We enter here in what is called Stacked Ensembles. Combine base learners with a meta-learner. SInce I want to avoid to overfit to the public LB I have decided to avoid things too complex.",
    "928191": "Use Power Averaging: maybe helpfully",
    "929220": "Can you please explain this more!",
    "929675": "Hi @fireheart7 you can learn more about power averaging from @sirishks great explanation:\n\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165653",
    "929687": "Especially in the private leaderboard, the more generalized solution you have, the better rank you will have. Ensemble is the best way to generalize your solution.",
    "932802": "Hi @luigisaetta  can you help me, create an ensemble ?\ni Have created one model for the tabular data and a second model for the vision data and i would like to add concatenate them via a hard voting or similar.",
    "934276": "Thank you so much!",
    "934955": "Ensemble is the key for most of the competitions if you want to win a higher rank.",
    "935025": "Since I don't do this for a living I know nothing about customers :)\n\nIf you do ensemble as the best solution how can that be turned into a working model that gets new input to solve each day?  \n\nMakes sense to ensemble for a competition like Kaggle - but is it  a tool that is only useful on Kaggle?",
    "935071": "wait until they over-train and than the shake-up happens :) it's always a surprise when that happens",
    "935076": "You could also use blending (multiply the results from each datasets by weights, than either take the sum with means, median or std deviation)",
    "938789": "Hi, to start, you could do a simple weighted average of the predictions given by the single models. It is up to you to decide the weights (AUC from VC,  score from LB)."
  },
  "source": "meta"
}