{
  "id": 175390,
  "title": "Did we start ensembling too early?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175390",
  "author_name": "",
  "post_date": "2020-08-18T04:29:41.971148100Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In the few competitions that I have entered, generally, the bulk of the ensembling efforts are done in the last week or two of the competitions.</p>\n<p>This competition was different. Before I'd even started work on this competition there were lots of discussions and kernels suggesting that ensembling was the main way forward (for the public LB at least).</p>\n<p>I feel that this direction somewhat stifled the development of interesting solutions and discussions (and therefore learning opportunities). Ian Pan's great solution had a single model that would have taken the top spot on its own, and I'm sure there are many more that were not selected.</p>\n<p>What are your thoughts?</p>",
  "messages": [
    {
      "id": "974855",
      "postDate": "08/18/2020 04:29:41",
      "content": "<p>In the few competitions that I have entered, generally, the bulk of the ensembling efforts are done in the last week or two of the competitions.</p>\n<p>This competition was different. Before I'd even started work on this competition there were lots of discussions and kernels suggesting that ensembling was the main way forward (for the public LB at least).</p>\n<p>I feel that this direction somewhat stifled the development of interesting solutions and discussions (and therefore learning opportunities). Ian Pan's great solution had a single model that would have taken the top spot on its own, and I'm sure there are many more that were not selected.</p>\n<p>What are your thoughts?</p>",
      "rawMarkdown": "In the few competitions that I have entered, generally, the bulk of the ensembling efforts are done in the last week or two of the competitions.\n\nThis competition was different. Before I'd even started work on this competition there were lots of discussions and kernels suggesting that ensembling was the main way forward (for the public LB at least).\n\nI feel that this direction somewhat stifled the development of interesting solutions and discussions (and therefore learning opportunities). Ian Pan's great solution had a single model that would have taken the top spot on its own, and I'm sure there are many more that were not selected.\n\nWhat are your thoughts?",
      "votes": null
    },
    {
      "id": "974876",
      "postDate": "08/18/2020 04:38:18",
      "content": "<p>Ensembling a bit earlier was not a bad strategy here. Our team found the feedback from the ensemble/stacking model much more informative (and stable) than single models. </p>",
      "rawMarkdown": "Ensembling a bit earlier was not a bad strategy here. Our team found the feedback from the ensemble/stacking model much more informative (and stable) than single models.",
      "votes": null
    },
    {
      "id": "974914",
      "postDate": "08/18/2020 04:53:52",
      "content": "<p>I think it depends. I have in past competitions successfully ensembled the same model on different seeds with a good degree of success on private LB without much effort given to improving the model. Bengali competition for example, I didn't have much time to investigate the model deeper. In that competition, ensembling seeds of a basic model was enough to get silver but I didn't even have time to log into give much consideration to my final selection. Sometimes decent solutions are the least interesting.</p>",
      "rawMarkdown": "I think it depends. I have in past competitions successfully ensembled the same model on different seeds with a good degree of success on private LB without much effort given to improving the model. Bengali competition for example, I didn't have much time to investigate the model deeper. In that competition, ensembling seeds of a basic model was enough to get silver but I didn't even have time to log into give much consideration to my final selection. Sometimes decent solutions are the least interesting.",
      "votes": null
    },
    {
      "id": "974958",
      "postDate": "08/18/2020 05:12:38",
      "content": "<p>I ensembled from the beginning because I was having trouble achieving consistent high CV score with a single model. However in retrospect, if I spent a little more time building a stronger single model and then using that single model with different image sizes etc, I think my final ensemble would have scored higher.</p>",
      "rawMarkdown": "I ensembled from the beginning because I was having trouble achieving consistent high CV score with a single model. However in retrospect, if I spent a little more time building a stronger single model and then using that single model with different image sizes etc, I think my final ensemble would have scored higher.",
      "votes": null
    },
    {
      "id": "975194",
      "postDate": "08/18/2020 07:33:40",
      "content": "<p>To be honest I've been working on training the models for most of the time and started ensembling only around a week before the end. </p>\n<p>I had thought about class imbalance and never believed in the MinMax (I saw it is very unstable). We mainly decided to use quantile based ensemble (which helped with single model overfitting). </p>\n<p>I spend 40 hours + looking at the patterns in data (scores), before coming up with the final solution. </p>\n<p>Funny story, I created solution and sent the result 6 hours before the end of the competition, then I couldn't sleep and when I was in bed I thought of better approach, sent it and it improved Public LB by 0.003 (0.9601) and Private LB by 0.003 -&gt; (0.9451)  🙈</p>",
      "rawMarkdown": "To be honest I've been working on training the models for most of the time and started ensembling only around a week before the end. \n\nI had thought about class imbalance and never believed in the MinMax (I saw it is very unstable). We mainly decided to use quantile based ensemble (which helped with single model overfitting). \n\nI spend 40 hours + looking at the patterns in data (scores), before coming up with the final solution. \n\nFunny story, I created solution and sent the result 6 hours before the end of the competition, then I couldn't sleep and when I was in bed I thought of better approach, sent it and it improved Public LB by 0.003 (0.9601) and Private LB by 0.003 -> (0.9451)  🙈",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 974876,
      "author_name": "kazanova",
      "author_url": "",
      "post_date": "08/18/2020 04:38:18",
      "content": "<p>Ensembling a bit earlier was not a bad strategy here. Our team found the feedback from the ensemble/stacking model much more informative (and stable) than single models. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974914,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "08/18/2020 04:53:52",
      "content": "<p>I think it depends. I have in past competitions successfully ensembled the same model on different seeds with a good degree of success on private LB without much effort given to improving the model. Bengali competition for example, I didn't have much time to investigate the model deeper. In that competition, ensembling seeds of a basic model was enough to get silver but I didn't even have time to log into give much consideration to my final selection. Sometimes decent solutions are the least interesting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 974958,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/18/2020 05:12:38",
      "content": "<p>I ensembled from the beginning because I was having trouble achieving consistent high CV score with a single model. However in retrospect, if I spent a little more time building a stronger single model and then using that single model with different image sizes etc, I think my final ensemble would have scored higher.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 975194,
      "author_name": "janidziak",
      "author_url": "",
      "post_date": "08/18/2020 07:33:40",
      "content": "<p>To be honest I've been working on training the models for most of the time and started ensembling only around a week before the end. </p>\n<p>I had thought about class imbalance and never believed in the MinMax (I saw it is very unstable). We mainly decided to use quantile based ensemble (which helped with single model overfitting). </p>\n<p>I spend 40 hours + looking at the patterns in data (scores), before coming up with the final solution. </p>\n<p>Funny story, I created solution and sent the result 6 hours before the end of the competition, then I couldn't sleep and when I was in bed I thought of better approach, sent it and it improved Public LB by 0.003 (0.9601) and Private LB by 0.003 -&gt; (0.9451)  🙈</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "974855": "In the few competitions that I have entered, generally, the bulk of the ensembling efforts are done in the last week or two of the competitions.\n\nThis competition was different. Before I'd even started work on this competition there were lots of discussions and kernels suggesting that ensembling was the main way forward (for the public LB at least).\n\nI feel that this direction somewhat stifled the development of interesting solutions and discussions (and therefore learning opportunities). Ian Pan's great solution had a single model that would have taken the top spot on its own, and I'm sure there are many more that were not selected.\n\nWhat are your thoughts?",
    "974876": "Ensembling a bit earlier was not a bad strategy here. Our team found the feedback from the ensemble/stacking model much more informative (and stable) than single models.",
    "974914": "I think it depends. I have in past competitions successfully ensembled the same model on different seeds with a good degree of success on private LB without much effort given to improving the model. Bengali competition for example, I didn't have much time to investigate the model deeper. In that competition, ensembling seeds of a basic model was enough to get silver but I didn't even have time to log into give much consideration to my final selection. Sometimes decent solutions are the least interesting.",
    "974958": "I ensembled from the beginning because I was having trouble achieving consistent high CV score with a single model. However in retrospect, if I spent a little more time building a stronger single model and then using that single model with different image sizes etc, I think my final ensemble would have scored higher.",
    "975194": "To be honest I've been working on training the models for most of the time and started ensembling only around a week before the end. \n\nI had thought about class imbalance and never believed in the MinMax (I saw it is very unstable). We mainly decided to use quantile based ensemble (which helped with single model overfitting). \n\nI spend 40 hours + looking at the patterns in data (scores), before coming up with the final solution. \n\nFunny story, I created solution and sent the result 6 hours before the end of the competition, then I couldn't sleep and when I was in bed I thought of better approach, sent it and it improved Public LB by 0.003 (0.9601) and Private LB by 0.003 -> (0.9451)  🙈"
  },
  "source": "meta"
}