{
  "id": 40124,
  "title": "Did it overfit?",
  "url": "/competitions/carvana-image-masking-challenge/discussion/40124",
  "author_name": "",
  "post_date": "2017-09-28T01:20:16.760287500Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>you can compare your public and private score. here are my observations:</p>\n\n<ul>\n<li>pseudo-labelling(unet1024) : public LB 0.9968, private 0.9967. But note that I use pseudo-labelled data for pretraining only. later i use train images for fine tunning.</li>\n</ul>\n\n<p>yet another example:</p>\n\n<p>unet512 (baseline):  public 0.9961, private 0.9958</p>\n\n<p>unet512 (pseudo-labelling with baseline): public 0.9962, private 0.9960</p>\n\n<hr>\n\n<ul>\n<li>effect of express-van (see: <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39948\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39948</a>)  </li>\n</ul>\n\n<p>public LB (with and without quick fix) 0.9970, </p>\n\n<p>private(with) 0.9970, (without) 0.9968</p>\n\n<p>.</p>\n\n<p>i double check my results, it seems correct. other kaggler may want to verify this. </p>\n\n<hr>\n\n<ul>\n<li>effects of thresholding. for an ensemble of 9 models: </li>\n</ul>\n\n<p>public LB (majority vote threshold=0.4) 0.9969, (0.5) 0.9969, (0.6) 0.9969,</p>\n\n<p>note: by ranking, (0.5) is best &gt; (0.6) &gt; 0.4  is worst</p>\n\n<p>.</p>\n\n<p>private LB (majority vote threshold=0.4) 0.9968, (0.5) 0.9967, (0.6) 0.9966,</p>",
  "messages": [
    {
      "id": "224960",
      "postDate": "09/28/2017 01:20:16",
      "content": "<p>you can compare your public and private score. here are my observations:</p>\n\n<ul>\n<li>pseudo-labelling(unet1024) : public LB 0.9968, private 0.9967. But note that I use pseudo-labelled data for pretraining only. later i use train images for fine tunning.</li>\n</ul>\n\n<p>yet another example:</p>\n\n<p>unet512 (baseline):  public 0.9961, private 0.9958</p>\n\n<p>unet512 (pseudo-labelling with baseline): public 0.9962, private 0.9960</p>\n\n<hr>\n\n<ul>\n<li>effect of express-van (see: <a href=\"https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39948\">https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39948</a>)  </li>\n</ul>\n\n<p>public LB (with and without quick fix) 0.9970, </p>\n\n<p>private(with) 0.9970, (without) 0.9968</p>\n\n<p>.</p>\n\n<p>i double check my results, it seems correct. other kaggler may want to verify this. </p>\n\n<hr>\n\n<ul>\n<li>effects of thresholding. for an ensemble of 9 models: </li>\n</ul>\n\n<p>public LB (majority vote threshold=0.4) 0.9969, (0.5) 0.9969, (0.6) 0.9969,</p>\n\n<p>note: by ranking, (0.5) is best &gt; (0.6) &gt; 0.4  is worst</p>\n\n<p>.</p>\n\n<p>private LB (majority vote threshold=0.4) 0.9968, (0.5) 0.9967, (0.6) 0.9966,</p>",
      "rawMarkdown": "you can compare your public and private score. here are my observations:\n\n- pseudo-labelling(unet1024) : public LB 0.9968, private 0.9967. But note that I use pseudo-labelled data for pretraining only. later i use train images for fine tunning.\n\nyet another example:\n\nunet512 (baseline):  public 0.9961, private 0.9958\n\nunet512 (pseudo-labelling with baseline): public 0.9962, private 0.9960\n\n ----------------------\n\n\n - effect of express-van (see: https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39948)  \n\npublic LB (with and without quick fix) 0.9970, \n\nprivate(with) 0.9970, (without) 0.9968\n\n.\n\n i double check my results, it seems correct. other kaggler may want to verify this. \n\n ----------------------\n\n- effects of thresholding. for an ensemble of 9 models: \n\npublic LB (majority vote threshold=0.4) 0.9969, (0.5) 0.9969, (0.6) 0.9969,\n\n  note: by ranking, (0.5) is best &gt; (0.6) &gt; 0.4  is worst\n\n.\n\n\nprivate LB (majority vote threshold=0.4) 0.9968, (0.5) 0.9967, (0.6) 0.9966,",
      "votes": null
    },
    {
      "id": "224997",
      "postDate": "09/28/2017 03:09:48",
      "content": "<p>We ended up overfitting badly with 0.9968 on public LB (87th place) drop down to 0.9964 on private LB (129th place). I think our main issue is that we should have been more careful with our train/valid split by car instead of random split.  </p>\n\n<p>In the end, we actually had a better performing model that we didn't pick because it performed worse on public LB. </p>",
      "rawMarkdown": "We ended up overfitting badly with 0.9968 on public LB (87th place) drop down to 0.9964 on private LB (129th place). I think our main issue is that we should have been more careful with our train/valid split by car instead of random split.  \n\nIn the end, we actually had a better performing model that we didn't pick because it performed worse on public LB.",
      "votes": null
    },
    {
      "id": "225090",
      "postDate": "09/28/2017 08:08:42",
      "content": "<p>I did the same mistake of random splitting.  Feels bad man. Optimezed threshold (0.61) was giving around +1.2 on cv and lb, but now Im curious whether it would improve score if I splitted properly by car.</p>",
      "rawMarkdown": "I did the same mistake of random splitting.  Feels bad man. Optimezed threshold (0.61) was giving around +1.2 on cv and lb, but now Im curious whether it would improve score if I splitted properly by car.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 224997,
      "author_name": "jamesrequa",
      "author_url": "",
      "post_date": "09/28/2017 03:09:48",
      "content": "<p>We ended up overfitting badly with 0.9968 on public LB (87th place) drop down to 0.9964 on private LB (129th place). I think our main issue is that we should have been more careful with our train/valid split by car instead of random split.  </p>\n\n<p>In the end, we actually had a better performing model that we didn't pick because it performed worse on public LB. </p>",
      "votes": null,
      "replies": [
        {
          "id": 225090,
          "author_name": "heyt0ny",
          "author_url": "",
          "post_date": "09/28/2017 08:08:42",
          "content": "<p>I did the same mistake of random splitting.  Feels bad man. Optimezed threshold (0.61) was giving around +1.2 on cv and lb, but now Im curious whether it would improve score if I splitted properly by car.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "224960": "you can compare your public and private score. here are my observations:\n\n- pseudo-labelling(unet1024) : public LB 0.9968, private 0.9967. But note that I use pseudo-labelled data for pretraining only. later i use train images for fine tunning.\n\nyet another example:\n\nunet512 (baseline):  public 0.9961, private 0.9958\n\nunet512 (pseudo-labelling with baseline): public 0.9962, private 0.9960\n\n ----------------------\n\n\n - effect of express-van (see: https://www.kaggle.com/c/carvana-image-masking-challenge/discussion/39948)  \n\npublic LB (with and without quick fix) 0.9970, \n\nprivate(with) 0.9970, (without) 0.9968\n\n.\n\n i double check my results, it seems correct. other kaggler may want to verify this. \n\n ----------------------\n\n- effects of thresholding. for an ensemble of 9 models: \n\npublic LB (majority vote threshold=0.4) 0.9969, (0.5) 0.9969, (0.6) 0.9969,\n\n  note: by ranking, (0.5) is best &gt; (0.6) &gt; 0.4  is worst\n\n.\n\n\nprivate LB (majority vote threshold=0.4) 0.9968, (0.5) 0.9967, (0.6) 0.9966,",
    "224997": "We ended up overfitting badly with 0.9968 on public LB (87th place) drop down to 0.9964 on private LB (129th place). I think our main issue is that we should have been more careful with our train/valid split by car instead of random split.  \n\nIn the end, we actually had a better performing model that we didn't pick because it performed worse on public LB.",
    "225090": "I did the same mistake of random splitting.  Feels bad man. Optimezed threshold (0.61) was giving around +1.2 on cv and lb, but now Im curious whether it would improve score if I splitted properly by car."
  },
  "source": "meta"
}