{
  "id": 231143,
  "title": "rasterio bug on private dataset?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/231143",
  "author_name": "",
  "post_date": "2021-04-07T05:10:07.415164900Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I've made a code that tests if rasterio has all 3 channels (RGB) for all images. So I measured mean value of each channel. I know that this value has minimal possible value is about 120. So I put 40 threshold to be sure. If not all channels pass the threshold then dumb submission.csv will not be created and  it actually it is not created and I got submission scoring error (or may be I just has another error in the code). Do you know something about this? How to solve this? Is any alternatives to rasterio that works on the private dataset?<br>\nHere is the code:<br>\n`import numpy as np<br>\nimport matplotlib.pyplot as plt<br>\nimport cv2<br>\nimport csv<br>\nimport os<br>\nimport gc<br>\nimport glob<br>\nimport json<br>\nimport rasterio</p>\n<p>IMAGES_DIR = '../input/hubmap-kidney-segmentation/test'</p>\n<p>OUTPUT_FILE = 'submission.csv'</p>\n<p>def int_round(inp):<br>\n    return int(np.round(inp))</p>\n<p>files = glob.glob(IMAGES_DIR + '/*.tiff')</p>\n<p>all_ok = True<br>\nRESIZE_FACTOR_TMP = 0.05</p>\n<p>for file_index in range(len(files)):</p>\n<pre><code>file = files[file_index]\nfile_base = os.path.basename(file)\nfile_base_no_ext = os.path.splitext(file_base)[0]\n\n\nimage_id = file_base_no_ext\n\n\nprint(file_index, len(files) - 1)\n\n\n\n\n\n\n\n\nfid = rasterio.open(IMAGES_DIR + '/' + image_id + '.tiff', 'r', num_threads='all_cpus')\n\nwidth = fid.shape[1]\nheight = fid.shape[0]\n\nwidth_small = int_round(width * RESIZE_FACTOR_TMP)\nheight_small = int_round(height * RESIZE_FACTOR_TMP)\n\nsize_small = (width_small, height_small)\n\nimage_small = np.zeros((height_small, width_small, 3), np.uint8)\n\nif fid.count == 1:\n\n    for subdataset_index in range(len(fid.subdatasets)):\n        subdataset = fid.subdatasets[subdataset_index]\n        layer = rasterio.open(subdataset)\n        channel = layer.read(1)\n        channel_small = cv2.resize(channel, size_small, interpolation=cv2.INTER_CUBIC)\n        image_small[:, :, subdataset_index] = channel_small\n        del channel_small\n        gc.collect()\n\nelse:\n    for channel_index in range(3):\n        channel = fid.read(channel_index + 1)\n        channel_small = cv2.resize(channel, size_small, interpolation=cv2.INTER_CUBIC)\n        image_small[:, :, channel_index] = channel_small\n        del channel_small\n        gc.collect()\n\nfid.close()\ndel fid\n\nfor channel_index in range(3):\n    mean_tmp = np.mean(image_small[:, :, channel_index])\n    if mean_tmp &lt; 40.0:\n        all_ok = False\n        del image_small\n        gc.collect()\n        break\n\n\ndel image_small\ngc.collect()\n</code></pre>\n<p>if all_ok:<br>\n    output_file_fid = open(OUTPUT_FILE, 'w')<br>\n    line = 'id,predicted\\n'<br>\n    for file_index in range(len(files)):<br>\n        file = files[file_index]<br>\n        file_base = os.path.basename(file)<br>\n        file_base_no_ext = os.path.splitext(file_base)[0]<br>\n        image_id = file_base_no_ext<br>\n        line = image_id + ',10 10' + '\\n'<br>\n        output_file_fid.write(line)<br>\n    output_file_fid.close()`</p>",
  "messages": [
    {
      "id": "1265649",
      "postDate": "04/07/2021 05:10:07",
      "content": "<p>I've made a code that tests if rasterio has all 3 channels (RGB) for all images. So I measured mean value of each channel. I know that this value has minimal possible value is about 120. So I put 40 threshold to be sure. If not all channels pass the threshold then dumb submission.csv will not be created and  it actually it is not created and I got submission scoring error (or may be I just has another error in the code). Do you know something about this? How to solve this? Is any alternatives to rasterio that works on the private dataset?<br>\nHere is the code:<br>\n`import numpy as np<br>\nimport matplotlib.pyplot as plt<br>\nimport cv2<br>\nimport csv<br>\nimport os<br>\nimport gc<br>\nimport glob<br>\nimport json<br>\nimport rasterio</p>\n<p>IMAGES_DIR = '../input/hubmap-kidney-segmentation/test'</p>\n<p>OUTPUT_FILE = 'submission.csv'</p>\n<p>def int_round(inp):<br>\n    return int(np.round(inp))</p>\n<p>files = glob.glob(IMAGES_DIR + '/*.tiff')</p>\n<p>all_ok = True<br>\nRESIZE_FACTOR_TMP = 0.05</p>\n<p>for file_index in range(len(files)):</p>\n<pre><code>file = files[file_index]\nfile_base = os.path.basename(file)\nfile_base_no_ext = os.path.splitext(file_base)[0]\n\n\nimage_id = file_base_no_ext\n\n\nprint(file_index, len(files) - 1)\n\n\n\n\n\n\n\n\nfid = rasterio.open(IMAGES_DIR + '/' + image_id + '.tiff', 'r', num_threads='all_cpus')\n\nwidth = fid.shape[1]\nheight = fid.shape[0]\n\nwidth_small = int_round(width * RESIZE_FACTOR_TMP)\nheight_small = int_round(height * RESIZE_FACTOR_TMP)\n\nsize_small = (width_small, height_small)\n\nimage_small = np.zeros((height_small, width_small, 3), np.uint8)\n\nif fid.count == 1:\n\n    for subdataset_index in range(len(fid.subdatasets)):\n        subdataset = fid.subdatasets[subdataset_index]\n        layer = rasterio.open(subdataset)\n        channel = layer.read(1)\n        channel_small = cv2.resize(channel, size_small, interpolation=cv2.INTER_CUBIC)\n        image_small[:, :, subdataset_index] = channel_small\n        del channel_small\n        gc.collect()\n\nelse:\n    for channel_index in range(3):\n        channel = fid.read(channel_index + 1)\n        channel_small = cv2.resize(channel, size_small, interpolation=cv2.INTER_CUBIC)\n        image_small[:, :, channel_index] = channel_small\n        del channel_small\n        gc.collect()\n\nfid.close()\ndel fid\n\nfor channel_index in range(3):\n    mean_tmp = np.mean(image_small[:, :, channel_index])\n    if mean_tmp &lt; 40.0:\n        all_ok = False\n        del image_small\n        gc.collect()\n        break\n\n\ndel image_small\ngc.collect()\n</code></pre>\n<p>if all_ok:<br>\n    output_file_fid = open(OUTPUT_FILE, 'w')<br>\n    line = 'id,predicted\\n'<br>\n    for file_index in range(len(files)):<br>\n        file = files[file_index]<br>\n        file_base = os.path.basename(file)<br>\n        file_base_no_ext = os.path.splitext(file_base)[0]<br>\n        image_id = file_base_no_ext<br>\n        line = image_id + ',10 10' + '\\n'<br>\n        output_file_fid.write(line)<br>\n    output_file_fid.close()`</p>",
      "rawMarkdown": "I've made a code that tests if rasterio has all 3 channels (RGB) for all images. So I measured mean value of each channel. I know that this value has minimal possible value is about 120. So I put 40 threshold to be sure. If not all channels pass the threshold then dumb submission.csv will not be created and  it actually it is not created and I got submission scoring error (or may be I just has another error in the code). Do you know something about this? How to solve this? Is any alternatives to rasterio that works on the private dataset?\nHere is the code:\n`import numpy as np\nimport matplotlib.pyplot as plt\nimport cv2\nimport csv\nimport os\nimport gc\nimport glob\nimport json\nimport rasterio\n\n\n\nIMAGES_DIR = '../input/hubmap-kidney-segmentation/test'\n\n\nOUTPUT_FILE = 'submission.csv'\n\ndef int_round(inp):\n    return int(np.round(inp))\n\nfiles = glob.glob(IMAGES_DIR + '/*.tiff')\n\n\n\nall_ok = True\nRESIZE_FACTOR_TMP = 0.05\n\nfor file_index in range(len(files)):\n\n    \n    file = files[file_index]\n    file_base = os.path.basename(file)\n    file_base_no_ext = os.path.splitext(file_base)[0]\n    \n    \n    image_id = file_base_no_ext\n    \n\n    print(file_index, len(files) - 1)\n    \n\n    \n    \n    \n    \n    \n    \n    fid = rasterio.open(IMAGES_DIR + '/' + image_id + '.tiff', 'r', num_threads='all_cpus')\n    \n    width = fid.shape[1]\n    height = fid.shape[0]\n    \n    width_small = int_round(width * RESIZE_FACTOR_TMP)\n    height_small = int_round(height * RESIZE_FACTOR_TMP)\n    \n    size_small = (width_small, height_small)\n    \n    image_small = np.zeros((height_small, width_small, 3), np.uint8)\n    \n    if fid.count == 1:\n        \n        for subdataset_index in range(len(fid.subdatasets)):\n            subdataset = fid.subdatasets[subdataset_index]\n            layer = rasterio.open(subdataset)\n            channel = layer.read(1)\n            channel_small = cv2.resize(channel, size_small, interpolation=cv2.INTER_CUBIC)\n            image_small[:, :, subdataset_index] = channel_small\n            del channel_small\n            gc.collect()\n        \n    else:\n        for channel_index in range(3):\n            channel = fid.read(channel_index + 1)\n            channel_small = cv2.resize(channel, size_small, interpolation=cv2.INTER_CUBIC)\n            image_small[:, :, channel_index] = channel_small\n            del channel_small\n            gc.collect()\n    \n    fid.close()\n    del fid\n    \n    for channel_index in range(3):\n        mean_tmp = np.mean(image_small[:, :, channel_index])\n        if mean_tmp < 40.0:\n            all_ok = False\n            del image_small\n            gc.collect()\n            break\n    \n    \n    del image_small\n    gc.collect()\n    \n\nif all_ok:\n    output_file_fid = open(OUTPUT_FILE, 'w')\n    line = 'id,predicted\\n'\n    for file_index in range(len(files)):\n        file = files[file_index]\n        file_base = os.path.basename(file)\n        file_base_no_ext = os.path.splitext(file_base)[0]\n        image_id = file_base_no_ext\n        line = image_id + ',10 10' + '\\n'\n        output_file_fid.write(line)\n    output_file_fid.close()`",
      "votes": null
    },
    {
      "id": "1266714",
      "postDate": "04/08/2021 02:38:32",
      "content": "<p>Actually it works. It was another error. I fixed it. So at least each channel mean value is more then 40</p>",
      "rawMarkdown": "Actually it works. It was another error. I fixed it. So at least each channel mean value is more then 40",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1266714,
      "author_name": "vedenev",
      "author_url": "",
      "post_date": "04/08/2021 02:38:32",
      "content": "<p>Actually it works. It was another error. I fixed it. So at least each channel mean value is more then 40</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1265649": "I've made a code that tests if rasterio has all 3 channels (RGB) for all images. So I measured mean value of each channel. I know that this value has minimal possible value is about 120. So I put 40 threshold to be sure. If not all channels pass the threshold then dumb submission.csv will not be created and  it actually it is not created and I got submission scoring error (or may be I just has another error in the code). Do you know something about this? How to solve this? Is any alternatives to rasterio that works on the private dataset?\nHere is the code:\n`import numpy as np\nimport matplotlib.pyplot as plt\nimport cv2\nimport csv\nimport os\nimport gc\nimport glob\nimport json\nimport rasterio\n\n\n\nIMAGES_DIR = '../input/hubmap-kidney-segmentation/test'\n\n\nOUTPUT_FILE = 'submission.csv'\n\ndef int_round(inp):\n    return int(np.round(inp))\n\nfiles = glob.glob(IMAGES_DIR + '/*.tiff')\n\n\n\nall_ok = True\nRESIZE_FACTOR_TMP = 0.05\n\nfor file_index in range(len(files)):\n\n    \n    file = files[file_index]\n    file_base = os.path.basename(file)\n    file_base_no_ext = os.path.splitext(file_base)[0]\n    \n    \n    image_id = file_base_no_ext\n    \n\n    print(file_index, len(files) - 1)\n    \n\n    \n    \n    \n    \n    \n    \n    fid = rasterio.open(IMAGES_DIR + '/' + image_id + '.tiff', 'r', num_threads='all_cpus')\n    \n    width = fid.shape[1]\n    height = fid.shape[0]\n    \n    width_small = int_round(width * RESIZE_FACTOR_TMP)\n    height_small = int_round(height * RESIZE_FACTOR_TMP)\n    \n    size_small = (width_small, height_small)\n    \n    image_small = np.zeros((height_small, width_small, 3), np.uint8)\n    \n    if fid.count == 1:\n        \n        for subdataset_index in range(len(fid.subdatasets)):\n            subdataset = fid.subdatasets[subdataset_index]\n            layer = rasterio.open(subdataset)\n            channel = layer.read(1)\n            channel_small = cv2.resize(channel, size_small, interpolation=cv2.INTER_CUBIC)\n            image_small[:, :, subdataset_index] = channel_small\n            del channel_small\n            gc.collect()\n        \n    else:\n        for channel_index in range(3):\n            channel = fid.read(channel_index + 1)\n            channel_small = cv2.resize(channel, size_small, interpolation=cv2.INTER_CUBIC)\n            image_small[:, :, channel_index] = channel_small\n            del channel_small\n            gc.collect()\n    \n    fid.close()\n    del fid\n    \n    for channel_index in range(3):\n        mean_tmp = np.mean(image_small[:, :, channel_index])\n        if mean_tmp < 40.0:\n            all_ok = False\n            del image_small\n            gc.collect()\n            break\n    \n    \n    del image_small\n    gc.collect()\n    \n\nif all_ok:\n    output_file_fid = open(OUTPUT_FILE, 'w')\n    line = 'id,predicted\\n'\n    for file_index in range(len(files)):\n        file = files[file_index]\n        file_base = os.path.basename(file)\n        file_base_no_ext = os.path.splitext(file_base)[0]\n        image_id = file_base_no_ext\n        line = image_id + ',10 10' + '\\n'\n        output_file_fid.write(line)\n    output_file_fid.close()`",
    "1266714": "Actually it works. It was another error. I fixed it. So at least each channel mean value is more then 40"
  },
  "source": "meta"
}