{
  "id": 456126,
  "title": "Begineer question about the actual test files, folder and format",
  "url": "/competitions/blood-vessel-segmentation/discussion/456126",
  "author_name": "",
  "post_date": "2023-11-18T07:01:57.323959400Z",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<ol>\n<li>Does the actual test set maintain the same folder structure as shown in the example set?</li>\n<li>Is there a possibility of encountering non-TIFF images in the test set?</li>\n<li>Are the dimensions of the images in the test set relatively similar to those in the training set?</li>\n<li>Assuming the test set follows the same folder structure, should we anticipate the same number of slices per subject or kidney as in the training set?</li>\n</ol>",
  "messages": [
    {
      "id": "2529372",
      "postDate": "11/18/2023 07:01:57",
      "content": "<ol>\n<li>Does the actual test set maintain the same folder structure as shown in the example set?</li>\n<li>Is there a possibility of encountering non-TIFF images in the test set?</li>\n<li>Are the dimensions of the images in the test set relatively similar to those in the training set?</li>\n<li>Assuming the test set follows the same folder structure, should we anticipate the same number of slices per subject or kidney as in the training set?</li>\n</ol>",
      "rawMarkdown": "1. Does the actual test set maintain the same folder structure as shown in the example set?\n2. Is there a possibility of encountering non-TIFF images in the test set?\n3. Are the dimensions of the images in the test set relatively similar to those in the training set?\n4. Assuming the test set follows the same folder structure, should we anticipate the same number of slices per subject or kidney as in the training set?",
      "votes": null
    },
    {
      "id": "2529603",
      "postDate": "11/18/2023 11:22:23",
      "content": "<p>the best way is to refer to some submission codes that are successful in giving public score.</p>\n<p>1) Does the actual test set maintain the same folder structure as shown in the example set?</p>\n<ul>\n<li>yes</li>\n</ul>\n<p>2) Is there a possibility of encountering non-TIFF images in the test set?</p>\n<ul>\n<li>No. but this is not an issue because we will usually glob for a list of files and use  e.g. opencv to read file:</li>\n</ul>\n<pre><code> = glob(f  \n f  :\n     image = cv2.imread(f. ...)\n</code></pre>\n<p>3) Are the dimensions of the images in the test set relatively similar to those in the training set?</p>\n<ul>\n<li>not the same, but should be similar. you can write submission code to test.</li>\n</ul>\n<pre><code>file = glob(  \n f  file:\n     image = cv2.imread(f. ...)\n     H,W = image.shape\n      H  3x larger than train image:\n             notImplementError \n</code></pre>\n<p>4) Assuming the test set follows the same folder structure, should we anticipate the same number of slices per - subject or kidney as in the training set?<br>\nnumber of slices should be different.</p>\n<hr>\n<p>suggestion:</p>\n<p>write a dummy submission code to predict all mask image.<br>\nif it pass, then you know how to access the test images.</p>\n<pre><code> \nprediction = np.zeros_like(image)\n \n</code></pre>\n<hr>\n<p>e.g.<br>\n<a href=\"https://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline\" target=\"_blank\">https://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline</a></p>\n<pre><code>DATASET_FOLDER = \nls_images = (os(DATASET_FOLDER, , , , ))\n\n</code></pre>",
      "rawMarkdown": "the best way is to refer to some submission codes that are successful in giving public score.\n\n\n1) Does the actual test set maintain the same folder structure as shown in the example set?\n- yes\n\n2) Is there a possibility of encountering non-TIFF images in the test set?\n- No. but this is not an issue because we will usually glob for a list of files and use  e.g. opencv to read file:\n\n```\nfile = glob(f'{test_folder}\\*'  # glob for any format\nfor f in file:\n     image = cv2.imread(f. ...)\n\n```\n\n3) Are the dimensions of the images in the test set relatively similar to those in the training set?\n- not the same, but should be similar. you can write submission code to test.\n\n```\nfile = glob(f'{test_folder}\\*'  # glob for any format\nfor f in file:\n     image = cv2.imread(f. ...)\n     H,W = image.shape\n     if H is 3x larger than train image:\n           raise  notImplementError #this will cause submission failure\n\n```\n\n\n4) Assuming the test set follows the same folder structure, should we anticipate the same number of slices per - subject or kidney as in the training set?\nnumber of slices should be different.\n\n----\nsuggestion:\n\nwrite a dummy submission code to predict all mask image.\nif it pass, then you know how to access the test images.\n\n```\n... read all test images ...\nprediction = np.zeros_like(image)\n... encode in rle and make submission csv.\n\n```\n---\n\ne.g.\nhttps://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline\n```\nDATASET_FOLDER = \"/kaggle/input/blood-vessel-segmentation\"\nls_images = glob(os.path.join(DATASET_FOLDER, \"test\", \"*\", \"*\", \"*.tif\"))\nprint(f\"found images: {len(ls_images)}\")\n```",
      "votes": null
    },
    {
      "id": "2530285",
      "postDate": "11/19/2023 01:10:21",
      "content": "<p>These are great suggestions. I tried a submission with random 1s and 0s for each image. The code is below. But this failed after 5 hours. Any suggestions?</p>\n<pre><code> os\n numpy  np\n pandas  pd\n cv2\n glob  glob\n tqdm  tqdm\n torch\n torch.nn  nn\n torch.utils.data  Dataset, DataLoader\n albumentations  A\n\n numpy  np\n\n ():\n    \n    random_numbers = np.random.random()\n\n    \n    p = random_numbers / np.(random_numbers)\n\n    \n    random_data = np.random.choice([, ], size=img.shape, p=p)\n\n     random_data\n\n ():\n    \n    pixels = img.flatten()\n    pixels = np.concatenate([[], pixels, []])\n    runs = np.where(pixels[:] != pixels[:-])[] + \n    runs[::] -= runs[::]\n    rle = .join((x)  x  runs)\n     rle==:\n        rle = \n     rle\n\n\ntest_images = glob()\ntest_images.sort()\n=[]\nrle=[]\n image  test_images:\n    name = image.split()[]\n    slice_number = image.split()[].split()[]\n    rle_encoded = rle_encode(get_mask(cv2.imread(image)))\n    .append(name++slice_number)\n    rle.append(rle_encoded)\n\ndf = pd.DataFrame({: , : rle})\ndf.to_csv(, index=)\n</code></pre>",
      "rawMarkdown": "These are great suggestions. I tried a submission with random 1s and 0s for each image. The code is below. But this failed after 5 hours. Any suggestions?\n\n```python\nimport os\nimport numpy as np\nimport pandas as pd\nimport cv2\nfrom glob import glob\nfrom tqdm import tqdm\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nimport albumentations as A\n\nimport numpy as np\n\ndef get_mask(img):\n    # Generate two random numbers\n    random_numbers = np.random.random(2)\n\n    # Normalize them to sum up to 1\n    p = random_numbers / np.sum(random_numbers)\n\n    # Use this distribution in np.random.choice\n    random_data = np.random.choice([0, 1], size=img.shape, p=p)\n    \n    return random_data\n\ndef rle_encode(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels = img.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    rle = ' '.join(str(x) for x in runs)\n    if rle=='':\n        rle = '1 1'\n    return rle\n\n\ntest_images = glob(\"/kaggle/input/blood-vessel-segmentation/test/*/*/*.tif\")\ntest_images.sort()\nid=[]\nrle=[]\nfor image in test_images:\n    name = image.split('/')[5]\n    slice_number = image.split('/')[7].split('.')[0]\n    rle_encoded = rle_encode(get_mask(cv2.imread(image)))\n    id.append(name+'_'+slice_number)\n    rle.append(rle_encoded)\n\ndf = pd.DataFrame({'id': id, 'rle': rle})\ndf.to_csv('submission.csv', index=False)\n```",
      "votes": null
    },
    {
      "id": "2530565",
      "postDate": "11/19/2023 09:08:08",
      "content": "<p>So your code here is correct. I Modified 2 things and here is the outcome:</p>\n<ol>\n<li>Fixed rle_encoded to \"1 0\" and the submission competed with a score of 0 in less than 10 minutes.</li>\n<li>Fixed rle_encoded to \"1 1\" and the submission competed with a score of 0 in around 30 minutes.</li>\n</ol>\n<p>So as noted in other discussions (e.g. <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455894#2528861\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455894#2528861</a> ), your error is most likely caused  due to a problem with the metric score needing a lot of memory from the small randomly generated masks.</p>",
      "rawMarkdown": "So your code here is correct. I Modified 2 things and here is the outcome:\n1. Fixed rle_encoded to \"1 0\" and the submission competed with a score of 0 in less than 10 minutes.\n1. Fixed rle_encoded to \"1 1\" and the submission competed with a score of 0 in around 30 minutes.\n\nSo as noted in other discussions (e.g. https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455894#2528861 ), your error is most likely caused  due to a problem with the metric score needing a lot of memory from the small randomly generated masks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2529603,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/18/2023 11:22:23",
      "content": "<p>the best way is to refer to some submission codes that are successful in giving public score.</p>\n<p>1) Does the actual test set maintain the same folder structure as shown in the example set?</p>\n<ul>\n<li>yes</li>\n</ul>\n<p>2) Is there a possibility of encountering non-TIFF images in the test set?</p>\n<ul>\n<li>No. but this is not an issue because we will usually glob for a list of files and use  e.g. opencv to read file:</li>\n</ul>\n<pre><code> = glob(f  \n f  :\n     image = cv2.imread(f. ...)\n</code></pre>\n<p>3) Are the dimensions of the images in the test set relatively similar to those in the training set?</p>\n<ul>\n<li>not the same, but should be similar. you can write submission code to test.</li>\n</ul>\n<pre><code>file = glob(  \n f  file:\n     image = cv2.imread(f. ...)\n     H,W = image.shape\n      H  3x larger than train image:\n             notImplementError \n</code></pre>\n<p>4) Assuming the test set follows the same folder structure, should we anticipate the same number of slices per - subject or kidney as in the training set?<br>\nnumber of slices should be different.</p>\n<hr>\n<p>suggestion:</p>\n<p>write a dummy submission code to predict all mask image.<br>\nif it pass, then you know how to access the test images.</p>\n<pre><code> \nprediction = np.zeros_like(image)\n \n</code></pre>\n<hr>\n<p>e.g.<br>\n<a href=\"https://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline\" target=\"_blank\">https://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline</a></p>\n<pre><code>DATASET_FOLDER = \nls_images = (os(DATASET_FOLDER, , , , ))\n\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2530285,
          "author_name": "dhinkris",
          "author_url": "",
          "post_date": "11/19/2023 01:10:21",
          "content": "<p>These are great suggestions. I tried a submission with random 1s and 0s for each image. The code is below. But this failed after 5 hours. Any suggestions?</p>\n<pre><code> os\n numpy  np\n pandas  pd\n cv2\n glob  glob\n tqdm  tqdm\n torch\n torch.nn  nn\n torch.utils.data  Dataset, DataLoader\n albumentations  A\n\n numpy  np\n\n ():\n    \n    random_numbers = np.random.random()\n\n    \n    p = random_numbers / np.(random_numbers)\n\n    \n    random_data = np.random.choice([, ], size=img.shape, p=p)\n\n     random_data\n\n ():\n    \n    pixels = img.flatten()\n    pixels = np.concatenate([[], pixels, []])\n    runs = np.where(pixels[:] != pixels[:-])[] + \n    runs[::] -= runs[::]\n    rle = .join((x)  x  runs)\n     rle==:\n        rle = \n     rle\n\n\ntest_images = glob()\ntest_images.sort()\n=[]\nrle=[]\n image  test_images:\n    name = image.split()[]\n    slice_number = image.split()[].split()[]\n    rle_encoded = rle_encode(get_mask(cv2.imread(image)))\n    .append(name++slice_number)\n    rle.append(rle_encoded)\n\ndf = pd.DataFrame({: , : rle})\ndf.to_csv(, index=)\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2530565,
              "author_name": "coderrkj",
              "author_url": "",
              "post_date": "11/19/2023 09:08:08",
              "content": "<p>So your code here is correct. I Modified 2 things and here is the outcome:</p>\n<ol>\n<li>Fixed rle_encoded to \"1 0\" and the submission competed with a score of 0 in less than 10 minutes.</li>\n<li>Fixed rle_encoded to \"1 1\" and the submission competed with a score of 0 in around 30 minutes.</li>\n</ol>\n<p>So as noted in other discussions (e.g. <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455894#2528861\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455894#2528861</a> ), your error is most likely caused  due to a problem with the metric score needing a lot of memory from the small randomly generated masks.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2529372": "1. Does the actual test set maintain the same folder structure as shown in the example set?\n2. Is there a possibility of encountering non-TIFF images in the test set?\n3. Are the dimensions of the images in the test set relatively similar to those in the training set?\n4. Assuming the test set follows the same folder structure, should we anticipate the same number of slices per subject or kidney as in the training set?",
    "2529603": "the best way is to refer to some submission codes that are successful in giving public score.\n\n\n1) Does the actual test set maintain the same folder structure as shown in the example set?\n- yes\n\n2) Is there a possibility of encountering non-TIFF images in the test set?\n- No. but this is not an issue because we will usually glob for a list of files and use  e.g. opencv to read file:\n\n```\nfile = glob(f'{test_folder}\\*'  # glob for any format\nfor f in file:\n     image = cv2.imread(f. ...)\n\n```\n\n3) Are the dimensions of the images in the test set relatively similar to those in the training set?\n- not the same, but should be similar. you can write submission code to test.\n\n```\nfile = glob(f'{test_folder}\\*'  # glob for any format\nfor f in file:\n     image = cv2.imread(f. ...)\n     H,W = image.shape\n     if H is 3x larger than train image:\n           raise  notImplementError #this will cause submission failure\n\n```\n\n\n4) Assuming the test set follows the same folder structure, should we anticipate the same number of slices per - subject or kidney as in the training set?\nnumber of slices should be different.\n\n----\nsuggestion:\n\nwrite a dummy submission code to predict all mask image.\nif it pass, then you know how to access the test images.\n\n```\n... read all test images ...\nprediction = np.zeros_like(image)\n... encode in rle and make submission csv.\n\n```\n---\n\ne.g.\nhttps://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline\n```\nDATASET_FOLDER = \"/kaggle/input/blood-vessel-segmentation\"\nls_images = glob(os.path.join(DATASET_FOLDER, \"test\", \"*\", \"*\", \"*.tif\"))\nprint(f\"found images: {len(ls_images)}\")\n```",
    "2530285": "These are great suggestions. I tried a submission with random 1s and 0s for each image. The code is below. But this failed after 5 hours. Any suggestions?\n\n```python\nimport os\nimport numpy as np\nimport pandas as pd\nimport cv2\nfrom glob import glob\nfrom tqdm import tqdm\nimport torch\nimport torch.nn as nn\nfrom torch.utils.data import Dataset, DataLoader\nimport albumentations as A\n\nimport numpy as np\n\ndef get_mask(img):\n    # Generate two random numbers\n    random_numbers = np.random.random(2)\n\n    # Normalize them to sum up to 1\n    p = random_numbers / np.sum(random_numbers)\n\n    # Use this distribution in np.random.choice\n    random_data = np.random.choice([0, 1], size=img.shape, p=p)\n    \n    return random_data\n\ndef rle_encode(img):\n    '''\n    img: numpy array, 1 - mask, 0 - background\n    Returns run length as string formated\n    '''\n    pixels = img.flatten()\n    pixels = np.concatenate([[0], pixels, [0]])\n    runs = np.where(pixels[1:] != pixels[:-1])[0] + 1\n    runs[1::2] -= runs[::2]\n    rle = ' '.join(str(x) for x in runs)\n    if rle=='':\n        rle = '1 1'\n    return rle\n\n\ntest_images = glob(\"/kaggle/input/blood-vessel-segmentation/test/*/*/*.tif\")\ntest_images.sort()\nid=[]\nrle=[]\nfor image in test_images:\n    name = image.split('/')[5]\n    slice_number = image.split('/')[7].split('.')[0]\n    rle_encoded = rle_encode(get_mask(cv2.imread(image)))\n    id.append(name+'_'+slice_number)\n    rle.append(rle_encoded)\n\ndf = pd.DataFrame({'id': id, 'rle': rle})\ndf.to_csv('submission.csv', index=False)\n```",
    "2530565": "So your code here is correct. I Modified 2 things and here is the outcome:\n1. Fixed rle_encoded to \"1 0\" and the submission competed with a score of 0 in less than 10 minutes.\n1. Fixed rle_encoded to \"1 1\" and the submission competed with a score of 0 in around 30 minutes.\n\nSo as noted in other discussions (e.g. https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455894#2528861 ), your error is most likely caused  due to a problem with the metric score needing a lot of memory from the small randomly generated masks."
  },
  "source": "meta"
}