{
  "id": 230923,
  "title": "The gap between CV and LB",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/230923",
  "author_name": "",
  "post_date": "2021-04-06T07:20:20.049077800Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>My 5-fold CV is 0.917, but LB is 0.870. Is this gap reasonable? How to reduce it?</p>",
  "messages": [
    {
      "id": "1264458",
      "postDate": "04/06/2021 07:20:20",
      "content": "<p>My 5-fold CV is 0.917, but LB is 0.870. Is this gap reasonable? How to reduce it?</p>",
      "rawMarkdown": "My 5-fold CV is 0.917, but LB is 0.870. Is this gap reasonable? How to reduce it?",
      "votes": null
    },
    {
      "id": "1264467",
      "postDate": "04/06/2021 07:25:06",
      "content": "<p>My is 0.9 vs 0.84. I don't know why.</p>",
      "rawMarkdown": "My is 0.9 vs 0.84. I don't know why.",
      "votes": null
    },
    {
      "id": "1264503",
      "postDate": "04/06/2021 07:54:44",
      "content": "<p>My LB increased from 0.848 to 0.870 after changing the way to read the data for data.count==1.<br>\nold:</p>\n<pre><code>if self.data.count == 1:\nimg=  data.read([1,1,1], window=Window.from_slices((x1, x2), (y1, y2)))\nimg = np.moveaxis(img, 0, -1)\n</code></pre>\n<p>new：</p>\n<pre><code>if self.data.count == 1:\nimg = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n       for i, layer in enumerate(self.layers):\n             img[:,:,i] = layer.read(window=Window.from_slices((x1, x2),(y1, y2)))\n</code></pre>\n<p>I don't know how to explain it, but you can try. But there are still some problems with 0.87, and I want to find out.</p>",
      "rawMarkdown": "My LB increased from 0.848 to 0.870 after changing the way to read the data for data.count==1.\nold:\n```\nif self.data.count == 1:\nimg=  data.read([1,1,1], window=Window.from_slices((x1, x2), (y1, y2)))\nimg = np.moveaxis(img, 0, -1)\n```\nnew：\n```\nif self.data.count == 1:\nimg = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n       for i, layer in enumerate(self.layers):\n             img[:,:,i] = layer.read(window=Window.from_slices((x1, x2),(y1, y2)))\n```\nI don't know how to explain it, but you can try. But there are still some problems with 0.87, and I want to find out.",
      "votes": null
    },
    {
      "id": "1264560",
      "postDate": "04/06/2021 08:51:09",
      "content": "<p>I guess you got it from here: <a href=\"https://www.kaggle.com/iafoss/256x256-images\" target=\"_blank\">https://www.kaggle.com/iafoss/256x256-images</a><br>\ndata.read([1,1,1] - is strange, you read red channel 3 times.<br>\nDo you have \"else:\" ?</p>",
      "rawMarkdown": "I guess you got it from here: https://www.kaggle.com/iafoss/256x256-images\ndata.read([1,1,1] - is strange, you read red channel 3 times.\nDo you have \"else:\" ?",
      "votes": null
    },
    {
      "id": "1265234",
      "postDate": "04/06/2021 17:44:40",
      "content": "<p>For LB it looks like they weight each test image equally so when doing submissions with just one test image my score ranges from 0.16-0.19. Also for me that is pretty consistent for doing cross validation. You might want to try this to see if your predictions are good for some of the images and just really bad for one of the images (thats the case for me :))</p>\n<p>I had a similar issue where I wasn't looking at the out of fold predictions for each file just the average of all the predictions so some of the files that are larger would make my CV look a lot higher or lower than LB score. For training my cv can range from 0.6-0.9 depending on the file. </p>\n<p>Hope this helps and makes sense.</p>",
      "rawMarkdown": "For LB it looks like they weight each test image equally so when doing submissions with just one test image my score ranges from 0.16-0.19. Also for me that is pretty consistent for doing cross validation. You might want to try this to see if your predictions are good for some of the images and just really bad for one of the images (thats the case for me :))\n\nI had a similar issue where I wasn't looking at the out of fold predictions for each file just the average of all the predictions so some of the files that are larger would make my CV look a lot higher or lower than LB score. For training my cv can range from 0.6-0.9 depending on the file. \n\nHope this helps and makes sense.",
      "votes": null
    },
    {
      "id": "1265534",
      "postDate": "04/07/2021 01:58:55",
      "content": "<p>Thank you very much. But I still wonder how to do submissions with just one test image. Does it mean that only one test image is predicted and the other images are submitted with the RLE encoding of an all-zero array? According your comment, i think i need to make out-of-fold predictions for every single image in training cv. How do you solve the case of bad predictions for some images in cross-validation or testing? </p>",
      "rawMarkdown": "Thank you very much. But I still wonder how to do submissions with just one test image. Does it mean that only one test image is predicted and the other images are submitted with the RLE encoding of an all-zero array? According your comment, i think i need to make out-of-fold predictions for every single image in training cv. How do you solve the case of bad predictions for some images in cross-validation or testing?",
      "votes": null
    },
    {
      "id": "1267165",
      "postDate": "04/08/2021 11:02:38",
      "content": "<blockquote>\n  <p>Does it mean that only one test image is predicted and the other images are submitted with the RLE encoding of an all-zero array?&lt;</p>\n</blockquote>\n<p>Here is an example of the start of my prediction loop where I uncomment/comment the ids that I want to predict or not predict.</p>\n<pre><code>df = pd.read_csv(\"../input/hubmap-kidney-segmentation/sample_submission.csv\")\ndf = df.replace(np.nan, '', regex=True)\n\npublic_test_ids = [\n    # \"aa05346ff\",\n    # \"2ec3f1bb9\",\n    # \"3589adb90\",\n    \"d488c759a\",\n    # \"57512b7f1\",\n]\n\nfor idx,row in tqdm(df.iterrows(),total=len(df)):\n     if patient_id not in public_test_ids:\n             continue\n    # normal stuff\n\ndf.to_csv(\"submission.csv\", index=False)\n</code></pre>\n<blockquote>\n  <p>How do you solve the case of bad predictions for some images in cross-validation or testing? &lt;</p>\n</blockquote>\n<p>That is the million dollar question  😊  I think identifying the difficult ones from the easy ones is important in handling all of the possible test cases.</p>",
      "rawMarkdown": ">  Does it mean that only one test image is predicted and the other images are submitted with the RLE encoding of an all-zero array?<\n\nHere is an example of the start of my prediction loop where I uncomment/comment the ids that I want to predict or not predict.\n\n```\ndf = pd.read_csv(\"../input/hubmap-kidney-segmentation/sample_submission.csv\")\ndf = df.replace(np.nan, '', regex=True)\n\npublic_test_ids = [\n    # \"aa05346ff\",\n    # \"2ec3f1bb9\",\n    # \"3589adb90\",\n    \"d488c759a\",\n    # \"57512b7f1\",\n]\n\nfor idx,row in tqdm(df.iterrows(),total=len(df)):\n     if patient_id not in public_test_ids:\n             continue\n    # normal stuff\n\ndf.to_csv(\"submission.csv\", index=False)\n```\n\n> How do you solve the case of bad predictions for some images in cross-validation or testing? <\n\nThat is the million dollar question  😊  I think identifying the difficult ones from the easy ones is important in handling all of the possible test cases.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1264467,
      "author_name": "vedenev",
      "author_url": "",
      "post_date": "04/06/2021 07:25:06",
      "content": "<p>My is 0.9 vs 0.84. I don't know why.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1264503,
          "author_name": "ruansj",
          "author_url": "",
          "post_date": "04/06/2021 07:54:44",
          "content": "<p>My LB increased from 0.848 to 0.870 after changing the way to read the data for data.count==1.<br>\nold:</p>\n<pre><code>if self.data.count == 1:\nimg=  data.read([1,1,1], window=Window.from_slices((x1, x2), (y1, y2)))\nimg = np.moveaxis(img, 0, -1)\n</code></pre>\n<p>new：</p>\n<pre><code>if self.data.count == 1:\nimg = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n       for i, layer in enumerate(self.layers):\n             img[:,:,i] = layer.read(window=Window.from_slices((x1, x2),(y1, y2)))\n</code></pre>\n<p>I don't know how to explain it, but you can try. But there are still some problems with 0.87, and I want to find out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1264560,
          "author_name": "vedenev",
          "author_url": "",
          "post_date": "04/06/2021 08:51:09",
          "content": "<p>I guess you got it from here: <a href=\"https://www.kaggle.com/iafoss/256x256-images\" target=\"_blank\">https://www.kaggle.com/iafoss/256x256-images</a><br>\ndata.read([1,1,1] - is strange, you read red channel 3 times.<br>\nDo you have \"else:\" ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1265234,
      "author_name": "governor",
      "author_url": "",
      "post_date": "04/06/2021 17:44:40",
      "content": "<p>For LB it looks like they weight each test image equally so when doing submissions with just one test image my score ranges from 0.16-0.19. Also for me that is pretty consistent for doing cross validation. You might want to try this to see if your predictions are good for some of the images and just really bad for one of the images (thats the case for me :))</p>\n<p>I had a similar issue where I wasn't looking at the out of fold predictions for each file just the average of all the predictions so some of the files that are larger would make my CV look a lot higher or lower than LB score. For training my cv can range from 0.6-0.9 depending on the file. </p>\n<p>Hope this helps and makes sense.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1265534,
          "author_name": "ruansj",
          "author_url": "",
          "post_date": "04/07/2021 01:58:55",
          "content": "<p>Thank you very much. But I still wonder how to do submissions with just one test image. Does it mean that only one test image is predicted and the other images are submitted with the RLE encoding of an all-zero array? According your comment, i think i need to make out-of-fold predictions for every single image in training cv. How do you solve the case of bad predictions for some images in cross-validation or testing? </p>",
          "votes": null,
          "replies": [
            {
              "id": 1267165,
              "author_name": "governor",
              "author_url": "",
              "post_date": "04/08/2021 11:02:38",
              "content": "<blockquote>\n  <p>Does it mean that only one test image is predicted and the other images are submitted with the RLE encoding of an all-zero array?&lt;</p>\n</blockquote>\n<p>Here is an example of the start of my prediction loop where I uncomment/comment the ids that I want to predict or not predict.</p>\n<pre><code>df = pd.read_csv(\"../input/hubmap-kidney-segmentation/sample_submission.csv\")\ndf = df.replace(np.nan, '', regex=True)\n\npublic_test_ids = [\n    # \"aa05346ff\",\n    # \"2ec3f1bb9\",\n    # \"3589adb90\",\n    \"d488c759a\",\n    # \"57512b7f1\",\n]\n\nfor idx,row in tqdm(df.iterrows(),total=len(df)):\n     if patient_id not in public_test_ids:\n             continue\n    # normal stuff\n\ndf.to_csv(\"submission.csv\", index=False)\n</code></pre>\n<blockquote>\n  <p>How do you solve the case of bad predictions for some images in cross-validation or testing? &lt;</p>\n</blockquote>\n<p>That is the million dollar question  😊  I think identifying the difficult ones from the easy ones is important in handling all of the possible test cases.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1264458": "My 5-fold CV is 0.917, but LB is 0.870. Is this gap reasonable? How to reduce it?",
    "1264467": "My is 0.9 vs 0.84. I don't know why.",
    "1264503": "My LB increased from 0.848 to 0.870 after changing the way to read the data for data.count==1.\nold:\n```\nif self.data.count == 1:\nimg=  data.read([1,1,1], window=Window.from_slices((x1, x2), (y1, y2)))\nimg = np.moveaxis(img, 0, -1)\n```\nnew：\n```\nif self.data.count == 1:\nimg = np.zeros((WINDOW, WINDOW, 3), dtype=np.uint8)\n       for i, layer in enumerate(self.layers):\n             img[:,:,i] = layer.read(window=Window.from_slices((x1, x2),(y1, y2)))\n```\nI don't know how to explain it, but you can try. But there are still some problems with 0.87, and I want to find out.",
    "1264560": "I guess you got it from here: https://www.kaggle.com/iafoss/256x256-images\ndata.read([1,1,1] - is strange, you read red channel 3 times.\nDo you have \"else:\" ?",
    "1265234": "For LB it looks like they weight each test image equally so when doing submissions with just one test image my score ranges from 0.16-0.19. Also for me that is pretty consistent for doing cross validation. You might want to try this to see if your predictions are good for some of the images and just really bad for one of the images (thats the case for me :))\n\nI had a similar issue where I wasn't looking at the out of fold predictions for each file just the average of all the predictions so some of the files that are larger would make my CV look a lot higher or lower than LB score. For training my cv can range from 0.6-0.9 depending on the file. \n\nHope this helps and makes sense.",
    "1265534": "Thank you very much. But I still wonder how to do submissions with just one test image. Does it mean that only one test image is predicted and the other images are submitted with the RLE encoding of an all-zero array? According your comment, i think i need to make out-of-fold predictions for every single image in training cv. How do you solve the case of bad predictions for some images in cross-validation or testing?",
    "1267165": ">  Does it mean that only one test image is predicted and the other images are submitted with the RLE encoding of an all-zero array?<\n\nHere is an example of the start of my prediction loop where I uncomment/comment the ids that I want to predict or not predict.\n\n```\ndf = pd.read_csv(\"../input/hubmap-kidney-segmentation/sample_submission.csv\")\ndf = df.replace(np.nan, '', regex=True)\n\npublic_test_ids = [\n    # \"aa05346ff\",\n    # \"2ec3f1bb9\",\n    # \"3589adb90\",\n    \"d488c759a\",\n    # \"57512b7f1\",\n]\n\nfor idx,row in tqdm(df.iterrows(),total=len(df)):\n     if patient_id not in public_test_ids:\n             continue\n    # normal stuff\n\ndf.to_csv(\"submission.csv\", index=False)\n```\n\n> How do you solve the case of bad predictions for some images in cross-validation or testing? <\n\nThat is the million dollar question  😊  I think identifying the difficult ones from the easy ones is important in handling all of the possible test cases."
  },
  "source": "meta"
}