{
  "id": 137153,
  "title": "Unaccounted competition points & learned lessons",
  "url": "/competitions/bengaliai-cv19/discussion/137153",
  "author_name": "",
  "post_date": "2020-03-19T11:37:40.340110500Z",
  "votes": 7,
  "comment_count": 3,
  "views": 0,
  "content": "<p>It was my first competition experience and I would like to share some conclusions that I've made during this competition.</p>\n\n<ol>\n<li>A good way to understand the solution problems is to make an error analysis. Consider model losses separately. Which labels predicted poorly? Which classes and class combinations predicted worse?</li>\n<li>Pay attention to the underlying competition task problem. Models in this task should be robust to new class combinations, i.e. they must be able to generalize prediction to unseen graphemes.</li>\n<li>Don’t get stuck with some certain solution. Try as many as possible approaches.</li>\n<li>Be aware of possible public/private dataset distribution discrepancy.</li>\n<li><p>Models ensemble is one of the most powerful approaches which especially fits kaggle competitions application. Models within an ensemble are preferable to be different by some of the following points or all of them:</p>\n\n<ul><li>Training data folds</li>\n<li>Architecture, training settings, hyper-params</li>\n<li>Augmentation, data preprocessing</li>\n<li>Biasing towards the specific data criteria (like in this competition seen &amp; unseen graphemes)</li></ul></li>\n<li><p>Pay attention to algorithms performance and memory usage. Check your code for unnecessary loops, run speed tests and find an optimal tuning for parameters like batch size,  image size. Step by step small code speed improvements lead to overall experimenting speed boost.</p></li>\n</ol>\n\n<p>Please, feel free to correct me if I'm wrong and to add some points to this list. Thanks to the competition hosts, all participants and kaggle contributors for experience sharing! :\")</p>",
  "messages": [
    {
      "id": "779469",
      "postDate": "03/19/2020 11:37:40",
      "content": "<p>It was my first competition experience and I would like to share some conclusions that I've made during this competition.</p>\n\n<ol>\n<li>A good way to understand the solution problems is to make an error analysis. Consider model losses separately. Which labels predicted poorly? Which classes and class combinations predicted worse?</li>\n<li>Pay attention to the underlying competition task problem. Models in this task should be robust to new class combinations, i.e. they must be able to generalize prediction to unseen graphemes.</li>\n<li>Don’t get stuck with some certain solution. Try as many as possible approaches.</li>\n<li>Be aware of possible public/private dataset distribution discrepancy.</li>\n<li><p>Models ensemble is one of the most powerful approaches which especially fits kaggle competitions application. Models within an ensemble are preferable to be different by some of the following points or all of them:</p>\n\n<ul><li>Training data folds</li>\n<li>Architecture, training settings, hyper-params</li>\n<li>Augmentation, data preprocessing</li>\n<li>Biasing towards the specific data criteria (like in this competition seen &amp; unseen graphemes)</li></ul></li>\n<li><p>Pay attention to algorithms performance and memory usage. Check your code for unnecessary loops, run speed tests and find an optimal tuning for parameters like batch size,  image size. Step by step small code speed improvements lead to overall experimenting speed boost.</p></li>\n</ol>\n\n<p>Please, feel free to correct me if I'm wrong and to add some points to this list. Thanks to the competition hosts, all participants and kaggle contributors for experience sharing! :\")</p>",
      "rawMarkdown": "It was my first competition experience and I would like to share some conclusions that I've made during this competition.\n\n1. A good way to understand the solution problems is to make an error analysis. Consider model losses separately. Which labels predicted poorly? Which classes and class combinations predicted worse?\n2. Pay attention to the underlying competition task problem. Models in this task should be robust to new class combinations, i.e. they must be able to generalize prediction to unseen graphemes.\n3. Don’t get stuck with some certain solution. Try as many as possible approaches.\n4. Be aware of possible public/private dataset distribution discrepancy.\n5. Models ensemble is one of the most powerful approaches which especially fits kaggle competitions application. Models within an ensemble are preferable to be different by some of the following points or all of them:\n  - Training data folds\n  - Architecture, training settings, hyper-params\n  - Augmentation, data preprocessing\n  - Biasing towards the specific data criteria (like in this competition seen &amp; unseen graphemes)\n\n6. Pay attention to algorithms performance and memory usage. Check your code for unnecessary loops, run speed tests and find an optimal tuning for parameters like batch size,  image size. Step by step small code speed improvements lead to overall experimenting speed boost.\n\nPlease, feel free to correct me if I'm wrong and to add some points to this list. Thanks to the competition hosts, all participants and kaggle contributors for experience sharing! :\")",
      "votes": null
    },
    {
      "id": "779525",
      "postDate": "03/19/2020 12:50:45",
      "content": "<p>Really like reading people's thoughts after a competition.</p>\n\n<p>I think one great way to make ensembles that also links to your first point is also to verify what types of errors each model does. Ideally you want as you mentioned models that has different architectures, augmentations, etc but also different errors I think. If you have 2 models that get mistakes in different places, that means you can harvest the best (or worst, depending on how you do it :P) of each model to make a ensemble that would be better than the sum of the two.</p>",
      "rawMarkdown": "Really like reading people's thoughts after a competition.\n\nI think one great way to make ensembles that also links to your first point is also to verify what types of errors each model does. Ideally you want as you mentioned models that has different architectures, augmentations, etc but also different errors I think. If you have 2 models that get mistakes in different places, that means you can harvest the best (or worst, depending on how you do it :P) of each model to make a ensemble that would be better than the sum of the two.",
      "votes": null
    },
    {
      "id": "779548",
      "postDate": "03/19/2020 13:17:47",
      "content": "<p>Thanks! Indeed, it's another good criterion for choosing a model to an ensemble set.</p>",
      "rawMarkdown": "Thanks! Indeed, it's another good criterion for choosing a model to an ensemble set.",
      "votes": null
    },
    {
      "id": "780784",
      "postDate": "03/20/2020 15:39:04",
      "content": "<p>Good points <a href=\"/arturdatascientist\">@arturdatascientist</a> . We all learn more and more with each competition we participate</p>",
      "rawMarkdown": "Good points @arturdatascientist . We all learn more and more with each competition we participate",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 779525,
      "author_name": "maxlenormand",
      "author_url": "",
      "post_date": "03/19/2020 12:50:45",
      "content": "<p>Really like reading people's thoughts after a competition.</p>\n\n<p>I think one great way to make ensembles that also links to your first point is also to verify what types of errors each model does. Ideally you want as you mentioned models that has different architectures, augmentations, etc but also different errors I think. If you have 2 models that get mistakes in different places, that means you can harvest the best (or worst, depending on how you do it :P) of each model to make a ensemble that would be better than the sum of the two.</p>",
      "votes": null,
      "replies": [
        {
          "id": 779548,
          "author_name": "arturdatascientist",
          "author_url": "",
          "post_date": "03/19/2020 13:17:47",
          "content": "<p>Thanks! Indeed, it's another good criterion for choosing a model to an ensemble set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 780784,
      "author_name": "vladvdv",
      "author_url": "",
      "post_date": "03/20/2020 15:39:04",
      "content": "<p>Good points <a href=\"/arturdatascientist\">@arturdatascientist</a> . We all learn more and more with each competition we participate</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "779469": "It was my first competition experience and I would like to share some conclusions that I've made during this competition.\n\n1. A good way to understand the solution problems is to make an error analysis. Consider model losses separately. Which labels predicted poorly? Which classes and class combinations predicted worse?\n2. Pay attention to the underlying competition task problem. Models in this task should be robust to new class combinations, i.e. they must be able to generalize prediction to unseen graphemes.\n3. Don’t get stuck with some certain solution. Try as many as possible approaches.\n4. Be aware of possible public/private dataset distribution discrepancy.\n5. Models ensemble is one of the most powerful approaches which especially fits kaggle competitions application. Models within an ensemble are preferable to be different by some of the following points or all of them:\n  - Training data folds\n  - Architecture, training settings, hyper-params\n  - Augmentation, data preprocessing\n  - Biasing towards the specific data criteria (like in this competition seen &amp; unseen graphemes)\n\n6. Pay attention to algorithms performance and memory usage. Check your code for unnecessary loops, run speed tests and find an optimal tuning for parameters like batch size,  image size. Step by step small code speed improvements lead to overall experimenting speed boost.\n\nPlease, feel free to correct me if I'm wrong and to add some points to this list. Thanks to the competition hosts, all participants and kaggle contributors for experience sharing! :\")",
    "779525": "Really like reading people's thoughts after a competition.\n\nI think one great way to make ensembles that also links to your first point is also to verify what types of errors each model does. Ideally you want as you mentioned models that has different architectures, augmentations, etc but also different errors I think. If you have 2 models that get mistakes in different places, that means you can harvest the best (or worst, depending on how you do it :P) of each model to make a ensemble that would be better than the sum of the two.",
    "779548": "Thanks! Indeed, it's another good criterion for choosing a model to an ensemble set.",
    "780784": "Good points @arturdatascientist . We all learn more and more with each competition we participate"
  },
  "source": "meta"
}