{
  "id": 118142,
  "title": "Blending using np.logical_or works",
  "url": "/competitions/understanding_cloud_organization/discussion/118142",
  "author_name": "",
  "post_date": "2019-11-19T18:11:17.115989200Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Blending multiple submissions using numpy.logical_or function can significantly increase the public score but not the private score. It seems this approach is overfitting. I have published a kernel <a href=\"https://www.kaggle.com/gogo827jz/blending-submissions-using-np-logical-or\">Blending submissions using np.logical or</a> to show an example how it works.</p>\n\n<p>Wellcome to discuss here if you have used it. Thanks.</p>",
  "messages": [
    {
      "id": "677031",
      "postDate": "11/19/2019 18:11:17",
      "content": "<p>Blending multiple submissions using numpy.logical_or function can significantly increase the public score but not the private score. It seems this approach is overfitting. I have published a kernel <a href=\"https://www.kaggle.com/gogo827jz/blending-submissions-using-np-logical-or\">Blending submissions using np.logical or</a> to show an example how it works.</p>\n\n<p>Wellcome to discuss here if you have used it. Thanks.</p>",
      "rawMarkdown": "Blending multiple submissions using numpy.logical_or function can significantly increase the public score but not the private score. It seems this approach is overfitting. I have published a kernel [Blending submissions using np.logical or][1] to show an example how it works.\n\nWellcome to discuss here if you have used it. Thanks.\n\n[1]: https://www.kaggle.com/gogo827jz/blending-submissions-using-np-logical-or",
      "votes": null
    },
    {
      "id": "677153",
      "postDate": "11/19/2019 21:26:43",
      "content": "<p>Interesting. Before I made my own models, I experimented blending public kernels and reached Public LB 0.661 and LB 0.663 when I added another weak <code>submission.csv</code> file.</p>\n\n<p>If you have submission files where the masks have already been converted to 0 and 1 and thus you only have <code>rle</code> (so you don't have the pixel probabilities anymore), then you can ensemble the mask pixels and empty masks separately. For mask pixels you can use</p>\n\n<pre><code>for k in range(len(subs[0]))\n    mask = rle2mask( sub.EncodedPixels[k] )\n    for sub in subs[1:]: mask += rle2mask( sub.EncodedPixels[k] )\n    subs[0].EncodedPixels[k] = mask2rle( (mask &gt;= threshold_pixel).astype(int) )\n</code></pre>\n\n<p>Then to ensemble empty masks (i.e. remove false positives) you can use</p>\n\n<pre><code>for k in range(len(subs[0]))\n    empty = sub.EncodedPixels[k] == ''\n    for sub in subs[1:]: empty += sub.EncodedPixels[k] == ''\n    if empty &gt;= threshold_empty: subs[0].EncodedPixels[k] = ''\n</code></pre>\n\n<p>You choose both a <code>threshold_pixel</code> and <code>threshold_empty</code>. And there are other variations like using one model as the \"main model\" and using lots of weak models with method 2 above to remove false positives from the \"main model\". And then once you've decided which masks (among 14792 rows of submission.csv) you will predict, you can put all models into method 1 and find the best masks to replace the ones you've decided to submit.</p>",
      "rawMarkdown": "Interesting. Before I made my own models, I experimented blending public kernels and reached Public LB 0.661 and LB 0.663 when I added another weak `submission.csv` file.\n\nIf you have submission files where the masks have already been converted to 0 and 1 and thus you only have `rle` (so you don't have the pixel probabilities anymore), then you can ensemble the mask pixels and empty masks separately. For mask pixels you can use\n\n    for k in range(len(subs[0]))\n        mask = rle2mask( sub.EncodedPixels[k] )\n        for sub in subs[1:]: mask += rle2mask( sub.EncodedPixels[k] )\n        subs[0].EncodedPixels[k] = mask2rle( (mask &gt;= threshold_pixel).astype(int) )\n\nThen to ensemble empty masks (i.e. remove false positives) you can use\n\n    for k in range(len(subs[0]))\n        empty = sub.EncodedPixels[k] == ''\n        for sub in subs[1:]: empty += sub.EncodedPixels[k] == ''\n        if empty &gt;= threshold_empty: subs[0].EncodedPixels[k] = ''\n\nYou choose both a `threshold_pixel` and `threshold_empty`. And there are other variations like using one model as the \"main model\" and using lots of weak models with method 2 above to remove false positives from the \"main model\". And then once you've decided which masks (among 14792 rows of submission.csv) you will predict, you can put all models into method 1 and find the best masks to replace the ones you've decided to submit.",
      "votes": null
    },
    {
      "id": "677156",
      "postDate": "11/19/2019 21:32:53",
      "content": "<p>Yes, I have tried some similar ideas as you have mentioned here. Some of them has better private scores but worse public scores. Unfortunately, I picked the one with highest public score but low private score and dropped 40 places...</p>",
      "rawMarkdown": "Yes, I have tried some similar ideas as you have mentioned here. Some of them has better private scores but worse public scores. Unfortunately, I picked the one with highest public score but low private score and dropped 40 places...",
      "votes": null
    },
    {
      "id": "677170",
      "postDate": "11/19/2019 21:48:28",
      "content": "<p>Ah, that's why you dropped. Sorry to hear that. I did see a similar thing.</p>\n\n<p>I was combining submission files in weird ways (using both <code>logical_and</code> and <code>logical_or</code>) and I got my public LB to jump from 0.670 to 0.672. However I applied the same technique to my CV and the score did not increase so I was suspicious. After the comp, I see that the corresponding private LB decreased.</p>\n\n<p>When I applied the techniques I posted above, both my public LB and CV increased so I used that for one of my final submissions which turned out to be my best private LB.</p>",
      "rawMarkdown": "Ah, that's why you dropped. Sorry to hear that. I did see a similar thing.\n\nI was combining submission files in weird ways (using both `logical_and` and `logical_or`) and I got my public LB to jump from 0.670 to 0.672. However I applied the same technique to my CV and the score did not increase so I was suspicious. After the comp, I see that the corresponding private LB decreased.\n\nWhen I applied the techniques I posted above, both my public LB and CV increased so I used that for one of my final submissions which turned out to be my best private LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 677153,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/19/2019 21:26:43",
      "content": "<p>Interesting. Before I made my own models, I experimented blending public kernels and reached Public LB 0.661 and LB 0.663 when I added another weak <code>submission.csv</code> file.</p>\n\n<p>If you have submission files where the masks have already been converted to 0 and 1 and thus you only have <code>rle</code> (so you don't have the pixel probabilities anymore), then you can ensemble the mask pixels and empty masks separately. For mask pixels you can use</p>\n\n<pre><code>for k in range(len(subs[0]))\n    mask = rle2mask( sub.EncodedPixels[k] )\n    for sub in subs[1:]: mask += rle2mask( sub.EncodedPixels[k] )\n    subs[0].EncodedPixels[k] = mask2rle( (mask &gt;= threshold_pixel).astype(int) )\n</code></pre>\n\n<p>Then to ensemble empty masks (i.e. remove false positives) you can use</p>\n\n<pre><code>for k in range(len(subs[0]))\n    empty = sub.EncodedPixels[k] == ''\n    for sub in subs[1:]: empty += sub.EncodedPixels[k] == ''\n    if empty &gt;= threshold_empty: subs[0].EncodedPixels[k] = ''\n</code></pre>\n\n<p>You choose both a <code>threshold_pixel</code> and <code>threshold_empty</code>. And there are other variations like using one model as the \"main model\" and using lots of weak models with method 2 above to remove false positives from the \"main model\". And then once you've decided which masks (among 14792 rows of submission.csv) you will predict, you can put all models into method 1 and find the best masks to replace the ones you've decided to submit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 677156,
          "author_name": "gogo827jz",
          "author_url": "",
          "post_date": "11/19/2019 21:32:53",
          "content": "<p>Yes, I have tried some similar ideas as you have mentioned here. Some of them has better private scores but worse public scores. Unfortunately, I picked the one with highest public score but low private score and dropped 40 places...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 677170,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/19/2019 21:48:28",
          "content": "<p>Ah, that's why you dropped. Sorry to hear that. I did see a similar thing.</p>\n\n<p>I was combining submission files in weird ways (using both <code>logical_and</code> and <code>logical_or</code>) and I got my public LB to jump from 0.670 to 0.672. However I applied the same technique to my CV and the score did not increase so I was suspicious. After the comp, I see that the corresponding private LB decreased.</p>\n\n<p>When I applied the techniques I posted above, both my public LB and CV increased so I used that for one of my final submissions which turned out to be my best private LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "677031": "Blending multiple submissions using numpy.logical_or function can significantly increase the public score but not the private score. It seems this approach is overfitting. I have published a kernel [Blending submissions using np.logical or][1] to show an example how it works.\n\nWellcome to discuss here if you have used it. Thanks.\n\n[1]: https://www.kaggle.com/gogo827jz/blending-submissions-using-np-logical-or",
    "677153": "Interesting. Before I made my own models, I experimented blending public kernels and reached Public LB 0.661 and LB 0.663 when I added another weak `submission.csv` file.\n\nIf you have submission files where the masks have already been converted to 0 and 1 and thus you only have `rle` (so you don't have the pixel probabilities anymore), then you can ensemble the mask pixels and empty masks separately. For mask pixels you can use\n\n    for k in range(len(subs[0]))\n        mask = rle2mask( sub.EncodedPixels[k] )\n        for sub in subs[1:]: mask += rle2mask( sub.EncodedPixels[k] )\n        subs[0].EncodedPixels[k] = mask2rle( (mask &gt;= threshold_pixel).astype(int) )\n\nThen to ensemble empty masks (i.e. remove false positives) you can use\n\n    for k in range(len(subs[0]))\n        empty = sub.EncodedPixels[k] == ''\n        for sub in subs[1:]: empty += sub.EncodedPixels[k] == ''\n        if empty &gt;= threshold_empty: subs[0].EncodedPixels[k] = ''\n\nYou choose both a `threshold_pixel` and `threshold_empty`. And there are other variations like using one model as the \"main model\" and using lots of weak models with method 2 above to remove false positives from the \"main model\". And then once you've decided which masks (among 14792 rows of submission.csv) you will predict, you can put all models into method 1 and find the best masks to replace the ones you've decided to submit.",
    "677156": "Yes, I have tried some similar ideas as you have mentioned here. Some of them has better private scores but worse public scores. Unfortunately, I picked the one with highest public score but low private score and dropped 40 places...",
    "677170": "Ah, that's why you dropped. Sorry to hear that. I did see a similar thing.\n\nI was combining submission files in weird ways (using both `logical_and` and `logical_or`) and I got my public LB to jump from 0.670 to 0.672. However I applied the same technique to my CV and the score did not increase so I was suspicious. After the comp, I see that the corresponding private LB decreased.\n\nWhen I applied the techniques I posted above, both my public LB and CV increased so I used that for one of my final submissions which turned out to be my best private LB."
  },
  "source": "meta"
}