{
  "id": 75286,
  "title": "Retraining on full dataset",
  "url": "/competitions/humpback-whale-identification/discussion/75286",
  "author_name": "",
  "post_date": "2018-12-20T08:01:38.335402Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In my <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/74647#latest-442351\">starter pack</a> I retrain on full the full dataset. This is a double edged sword at best and it is quite easy to mess up with this.</p>\n\n<p>Anyhow, to stay true to not sharing outside of Kaggle, here is a <a href=\"https://twitter.com/radekosmulski/status/1075656371532689408\">tweet</a> with some thoughts on this subject. I would recommend checking out the <a href=\"https://arxiv.org/abs/1302.4389\">paper</a> by Ian Goodfellow et al.</p>\n\n<p>In some sense relying on train loss to learn something about how our model will generalize sounds like a blasphemy. As it turns out in a controlled situation this approach can be leveraged to good extent and depending on the approach you take, might be very useful for this competition due to the many whales in the dataset having so few images.</p>",
  "messages": [
    {
      "id": "442618",
      "postDate": "12/20/2018 08:01:38",
      "content": "<p>In my <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/74647#latest-442351\">starter pack</a> I retrain on full the full dataset. This is a double edged sword at best and it is quite easy to mess up with this.</p>\n\n<p>Anyhow, to stay true to not sharing outside of Kaggle, here is a <a href=\"https://twitter.com/radekosmulski/status/1075656371532689408\">tweet</a> with some thoughts on this subject. I would recommend checking out the <a href=\"https://arxiv.org/abs/1302.4389\">paper</a> by Ian Goodfellow et al.</p>\n\n<p>In some sense relying on train loss to learn something about how our model will generalize sounds like a blasphemy. As it turns out in a controlled situation this approach can be leveraged to good extent and depending on the approach you take, might be very useful for this competition due to the many whales in the dataset having so few images.</p>",
      "rawMarkdown": "In my [starter pack](https://www.kaggle.com/c/humpback-whale-identification/discussion/74647#latest-442351) I retrain on full the full dataset. This is a double edged sword at best and it is quite easy to mess up with this.\n\nAnyhow, to stay true to not sharing outside of Kaggle, here is a [tweet](https://twitter.com/radekosmulski/status/1075656371532689408) with some thoughts on this subject. I would recommend checking out the [paper](https://arxiv.org/abs/1302.4389) by Ian Goodfellow et al.\n\nIn some sense relying on train loss to learn something about how our model will generalize sounds like a blasphemy. As it turns out in a controlled situation this approach can be leveraged to good extent and depending on the approach you take, might be very useful for this competition due to the many whales in the dataset having so few images.",
      "votes": null
    },
    {
      "id": "442839",
      "postDate": "12/20/2018 15:37:49",
      "content": "<p>FWIW, after playing around with your code, I'm pretty convinced that the main bump I got from it was because of your core validation strategy/upsampling strategy which (I think) vibes with what you're saying in that tweet. Looking forward to digging into the paper.</p>",
      "rawMarkdown": "FWIW, after playing around with your code, I'm pretty convinced that the main bump I got from it was because of your core validation strategy/upsampling strategy which (I think) vibes with what you're saying in that tweet. Looking forward to digging into the paper.",
      "votes": null
    },
    {
      "id": "443044",
      "postDate": "12/20/2018 23:22:18",
      "content": "<p>Thanks for sharing, kernel version is working now...</p>\n\n<p><a href=\"https://www.kaggle.com/dromosys/radek-fast-ai-whale/\">https://www.kaggle.com/dromosys/radek-fast-ai-whale/</a></p>",
      "rawMarkdown": "Thanks for sharing, kernel version is working now...\n\nhttps://www.kaggle.com/dromosys/radek-fast-ai-whale/",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 442839,
      "author_name": "larcat",
      "author_url": "",
      "post_date": "12/20/2018 15:37:49",
      "content": "<p>FWIW, after playing around with your code, I'm pretty convinced that the main bump I got from it was because of your core validation strategy/upsampling strategy which (I think) vibes with what you're saying in that tweet. Looking forward to digging into the paper.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 443044,
      "author_name": "dromosys",
      "author_url": "",
      "post_date": "12/20/2018 23:22:18",
      "content": "<p>Thanks for sharing, kernel version is working now...</p>\n\n<p><a href=\"https://www.kaggle.com/dromosys/radek-fast-ai-whale/\">https://www.kaggle.com/dromosys/radek-fast-ai-whale/</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "442618": "In my [starter pack](https://www.kaggle.com/c/humpback-whale-identification/discussion/74647#latest-442351) I retrain on full the full dataset. This is a double edged sword at best and it is quite easy to mess up with this.\n\nAnyhow, to stay true to not sharing outside of Kaggle, here is a [tweet](https://twitter.com/radekosmulski/status/1075656371532689408) with some thoughts on this subject. I would recommend checking out the [paper](https://arxiv.org/abs/1302.4389) by Ian Goodfellow et al.\n\nIn some sense relying on train loss to learn something about how our model will generalize sounds like a blasphemy. As it turns out in a controlled situation this approach can be leveraged to good extent and depending on the approach you take, might be very useful for this competition due to the many whales in the dataset having so few images.",
    "442839": "FWIW, after playing around with your code, I'm pretty convinced that the main bump I got from it was because of your core validation strategy/upsampling strategy which (I think) vibes with what you're saying in that tweet. Looking forward to digging into the paper.",
    "443044": "Thanks for sharing, kernel version is working now...\n\nhttps://www.kaggle.com/dromosys/radek-fast-ai-whale/"
  },
  "source": "meta"
}