{
  "id": 73856,
  "title": "What is the right way to compete?",
  "url": "/competitions/humpback-whale-identification/discussion/73856",
  "author_name": "",
  "post_date": "2018-12-06T06:47:06.759149900Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi Everyone, \nThis is my second ever competition that I plan to sincerely participate in, and I seek your advice on best practises for competing. </p>\n\n<p>Here are a few things I'm doing:</p>\n\n<ul>\n<li>Teaming up with people with similar experience: @pnvijay and @init927 (We're all almost new to kaggle and slightly familiar with Deep Learning via the fastai MOOC)</li>\n<li>Using Public Kernels and hacking the public resources into fastai machinery (I'm a fastai student and am very familiar with the framework). Ex: Using the kernel by @suicaokhoailang</li>\n<li>Searching Past Competitions for similar traces/ideas: Ex: My Teammate @pnvijay found MAP Code in an older kernel </li>\n<li>Running constant experiments based on public kernels and going forth with the ones that are showing an increased score. </li>\n<li>Having a good local validation and cross checking if it works by comparing our estimates against the public LB. </li>\n<li>Long Term Startegy: Aiming for a medal, making a submission and trying to maintaing the position achieved on the LB in the future. </li>\n</ul>\n\n<p>Do you think this is a good plan? \nMy main question is what happens once I've done all experiments with fastai. For ex, the lib supports ResNet34,50,101,152. What should we consider doing once our experiments with these are over? </p>\n\n<p>Also, any suggestions on how should we generally distribute work amongst the team-mates?</p>\n\n<p>Thanks in Advance!</p>",
  "messages": [
    {
      "id": "434286",
      "postDate": "12/06/2018 06:47:06",
      "content": "<p>Hi Everyone, \nThis is my second ever competition that I plan to sincerely participate in, and I seek your advice on best practises for competing. </p>\n\n<p>Here are a few things I'm doing:</p>\n\n<ul>\n<li>Teaming up with people with similar experience: @pnvijay and @init927 (We're all almost new to kaggle and slightly familiar with Deep Learning via the fastai MOOC)</li>\n<li>Using Public Kernels and hacking the public resources into fastai machinery (I'm a fastai student and am very familiar with the framework). Ex: Using the kernel by @suicaokhoailang</li>\n<li>Searching Past Competitions for similar traces/ideas: Ex: My Teammate @pnvijay found MAP Code in an older kernel </li>\n<li>Running constant experiments based on public kernels and going forth with the ones that are showing an increased score. </li>\n<li>Having a good local validation and cross checking if it works by comparing our estimates against the public LB. </li>\n<li>Long Term Startegy: Aiming for a medal, making a submission and trying to maintaing the position achieved on the LB in the future. </li>\n</ul>\n\n<p>Do you think this is a good plan? \nMy main question is what happens once I've done all experiments with fastai. For ex, the lib supports ResNet34,50,101,152. What should we consider doing once our experiments with these are over? </p>\n\n<p>Also, any suggestions on how should we generally distribute work amongst the team-mates?</p>\n\n<p>Thanks in Advance!</p>",
      "rawMarkdown": "Hi Everyone, \nThis is my second ever competition that I plan to sincerely participate in, and I seek your advice on best practises for competing. \n\nHere are a few things I'm doing:\n\n - Teaming up with people with similar experience: @pnvijay and @init927 (We're all almost new to kaggle and slightly familiar with Deep Learning via the fastai MOOC)\n - Using Public Kernels and hacking the public resources into fastai machinery (I'm a fastai student and am very familiar with the framework). Ex: Using the kernel by @suicaokhoailang\n - Searching Past Competitions for similar traces/ideas: Ex: My Teammate @pnvijay found MAP Code in an older kernel \n - Running constant experiments based on public kernels and going forth with the ones that are showing an increased score. \n - Having a good local validation and cross checking if it works by comparing our estimates against the public LB. \n - Long Term Startegy: Aiming for a medal, making a submission and trying to maintaing the position achieved on the LB in the future. \n\nDo you think this is a good plan? \nMy main question is what happens once I've done all experiments with fastai. For ex, the lib supports ResNet34,50,101,152. What should we consider doing once our experiments with these are over? \n\nAlso, any suggestions on how should we generally distribute work amongst the team-mates?\n\nThanks in Advance!",
      "votes": null
    },
    {
      "id": "434657",
      "postDate": "12/06/2018 18:33:14",
      "content": "<p>I haven't been in many competitions either, but I guess one way to go is to look for useful kernels <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">like this one</a>. For the dataset have so many classes with just only one observation each and one class (new_whale) with lots of observations, overfitting is your number 1 enemy. Therefore a good preprocessing is highly necessary- perhaps even more important than to select the model itself. Check how the training data preprocessing have been done in this kernel to have an idea on what to do and try your own implementations :)  </p>",
      "rawMarkdown": "I haven't been in many competitions either, but I guess one way to go is to look for useful kernels [like this one](https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563). For the dataset have so many classes with just only one observation each and one class (new_whale) with lots of observations, overfitting is your number 1 enemy. Therefore a good preprocessing is highly necessary- perhaps even more important than to select the model itself. Check how the training data preprocessing have been done in this kernel to have an idea on what to do and try your own implementations :)",
      "votes": null
    },
    {
      "id": "434665",
      "postDate": "12/06/2018 18:42:40",
      "content": "<p><a href=\"/init27\">@init27</a> even I am stuck in same experimentation of various Resnet architectures</p>",
      "rawMarkdown": "init27 even I am stuck in same experimentation of various Resnet architectures",
      "votes": null
    },
    {
      "id": "435141",
      "postDate": "12/07/2018 15:20:22",
      "content": "<p>With regards to @Henrique's comment and from Piotte's Notebook:</p>\n\n<blockquote>\n  <p>Image selection\n  To begin, we reduce the number of images from the training set:</p>\n  \n  <ul>\n  <li>Images from the blacklist are removed;</li>\n  <li>Duplicate images are removed;</li>\n  <li>All 'new_whale' images are removed;</li>\n  <li>All whales with a single image are removed.</li>\n  </ul>\n  \n  <p>The blacklist was constructed manually by spotting images unhelpful to training. Reasons could be the underside of the fluke is not visible, or we see only dead fluke fragments on the a beach, there are two whales in the picture, etc. The list is in no way exhaustive.</p>\n</blockquote>",
      "rawMarkdown": "With regards to @Henrique's comment and from Piotte's Notebook:\n\n&gt; Image selection\n&gt; To begin, we reduce the number of images from the training set:\n&gt;\n&gt; * Images from the blacklist are removed;\n&gt; * Duplicate images are removed;\n&gt; * All 'new_whale' images are removed;\n&gt; * All whales with a single image are removed.\n&gt;\n&gt; The blacklist was constructed manually by spotting images unhelpful to training. Reasons could be the underside of the fluke is not visible, or we see only dead fluke fragments on the a beach, there are two whales in the picture, etc. The list is in no way exhaustive.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 434657,
      "author_name": "hrmello",
      "author_url": "",
      "post_date": "12/06/2018 18:33:14",
      "content": "<p>I haven't been in many competitions either, but I guess one way to go is to look for useful kernels <a href=\"https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563\">like this one</a>. For the dataset have so many classes with just only one observation each and one class (new_whale) with lots of observations, overfitting is your number 1 enemy. Therefore a good preprocessing is highly necessary- perhaps even more important than to select the model itself. Check how the training data preprocessing have been done in this kernel to have an idea on what to do and try your own implementations :)  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 434665,
      "author_name": "kurianbenoy",
      "author_url": "",
      "post_date": "12/06/2018 18:42:40",
      "content": "<p><a href=\"/init27\">@init27</a> even I am stuck in same experimentation of various Resnet architectures</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 435141,
      "author_name": "badtyprr",
      "author_url": "",
      "post_date": "12/07/2018 15:20:22",
      "content": "<p>With regards to @Henrique's comment and from Piotte's Notebook:</p>\n\n<blockquote>\n  <p>Image selection\n  To begin, we reduce the number of images from the training set:</p>\n  \n  <ul>\n  <li>Images from the blacklist are removed;</li>\n  <li>Duplicate images are removed;</li>\n  <li>All 'new_whale' images are removed;</li>\n  <li>All whales with a single image are removed.</li>\n  </ul>\n  \n  <p>The blacklist was constructed manually by spotting images unhelpful to training. Reasons could be the underside of the fluke is not visible, or we see only dead fluke fragments on the a beach, there are two whales in the picture, etc. The list is in no way exhaustive.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "434286": "Hi Everyone, \nThis is my second ever competition that I plan to sincerely participate in, and I seek your advice on best practises for competing. \n\nHere are a few things I'm doing:\n\n - Teaming up with people with similar experience: @pnvijay and @init927 (We're all almost new to kaggle and slightly familiar with Deep Learning via the fastai MOOC)\n - Using Public Kernels and hacking the public resources into fastai machinery (I'm a fastai student and am very familiar with the framework). Ex: Using the kernel by @suicaokhoailang\n - Searching Past Competitions for similar traces/ideas: Ex: My Teammate @pnvijay found MAP Code in an older kernel \n - Running constant experiments based on public kernels and going forth with the ones that are showing an increased score. \n - Having a good local validation and cross checking if it works by comparing our estimates against the public LB. \n - Long Term Startegy: Aiming for a medal, making a submission and trying to maintaing the position achieved on the LB in the future. \n\nDo you think this is a good plan? \nMy main question is what happens once I've done all experiments with fastai. For ex, the lib supports ResNet34,50,101,152. What should we consider doing once our experiments with these are over? \n\nAlso, any suggestions on how should we generally distribute work amongst the team-mates?\n\nThanks in Advance!",
    "434657": "I haven't been in many competitions either, but I guess one way to go is to look for useful kernels [like this one](https://www.kaggle.com/martinpiotte/whale-recognition-model-with-score-0-78563). For the dataset have so many classes with just only one observation each and one class (new_whale) with lots of observations, overfitting is your number 1 enemy. Therefore a good preprocessing is highly necessary- perhaps even more important than to select the model itself. Check how the training data preprocessing have been done in this kernel to have an idea on what to do and try your own implementations :)",
    "434665": "init27 even I am stuck in same experimentation of various Resnet architectures",
    "435141": "With regards to @Henrique's comment and from Piotte's Notebook:\n\n&gt; Image selection\n&gt; To begin, we reduce the number of images from the training set:\n&gt;\n&gt; * Images from the blacklist are removed;\n&gt; * Duplicate images are removed;\n&gt; * All 'new_whale' images are removed;\n&gt; * All whales with a single image are removed.\n&gt;\n&gt; The blacklist was constructed manually by spotting images unhelpful to training. Reasons could be the underside of the fluke is not visible, or we see only dead fluke fragments on the a beach, there are two whales in the picture, etc. The list is in no way exhaustive."
  },
  "source": "meta"
}