{
  "id": 475260,
  "title": "14th Place Solution",
  "url": "/competitions/blood-vessel-segmentation/writeups/yumeneko-14th-place-solution",
  "author_name": "",
  "post_date": "2024-02-08T02:39:12.930Z",
  "votes": 17,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F6a7d79a9543e55064e881b6961b14061%2Foverview.jpg?generation=1707324869303897&amp;alt=media\" alt=\"overview\"></p>\n<ul>\n<li>2.5D segmentation model that inputs N consecutive slices stacked in ch direction and outputs corresponding Nch masks.  </li>\n<li>Input images are cropped to the kidney area only and then resized.  </li>\n</ul>\n<h1>Pipeline Detail　　</h1>\n<h2>1. Preprocess</h2>\n<h3>1-1. Normalization</h3>\n<ul>\n<li>A histogram of luminance values is calculated for the entire kidney and normalized based on minimum and maximum values.  </li>\n<li>Normalization based on maximum and minimum values per image unit could cause variations in the appearance of images, resulting in unnatural switching of inference results. To counteract this, normalization based on the luminance distribution of the entire kidney was employed.  </li>\n<li>The code is as follows.  </li>\n</ul>\n<pre><code> ():\n    img_paths = (glob(os.path.join(image_dir, )))\n\n    pixels = np.zeros((,), dtype=np.int64)\n     img_path  tqdm(img_paths):\n        img = cv2.imread(img_path, cv2.IMREAD_UNCHANGED)\n        _pixels = np.bincount(img.flatten(), minlength=)\n        pixels += _pixels\n\n     = \n    hist = []\n    bins = []\n     i  (, +, ):\n        hist.append(pixels[i:i+].())\n        bins.append(i)\n    hist = np.array(hist)\n    hist_rate = hist/hist.()\n    idxes = np.where(hist_rate&gt;)[]\n    min_idx = idxes[]\n    max_idx = idxes[-]\n     bins[min_idx]-, bins[max_idx]+  \n\nmin_val, max_val = get_min_max_val(inference_img_dir)\nimg = cv2.imread(path, cv2.IMREAD_UNCHANGED)\nimg = img.astype()\nimg = np.clip(img, min_value, max_value)\nimg = (img-min_value)/max_value\n</code></pre>\n<h3>1-2. Crop</h3>\n<ul>\n<li>Obtain a rectangle of the kidney region using a segmentation model that infers a mask of the entire kidney.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F68c25e3b97c72b5b599b5607e8bcff8f%2Fcrop.jpg?generation=1707324937523266&amp;alt=media\" alt=\"crop\"></li>\n<li>The kidney segmentation model used a single model of Unet (backbone: efficientnet_b4) and only kidney_1_dense was used as training data.  </li>\n<li>Cropping only the kidney region eliminates wasted areas in the image and greatly improves the accuracy of vessel segmentation.  </li>\n<li>During training, the height and width of the rectangle were stochastically increased or decreased by ±5% as part of augmentation.</li>\n</ul>\n<h3>1-3. Resize</h3>\n<ul>\n<li>Because of the strict masking requirements of this competition metric, it was important to resize the image to a larger image size.  </li>\n<li>In my solution, I trained the model by resizing the image to as large as GPU memory would allow, in the range of 1536~1920.  </li>\n<li>If the mask is resized by OpenCV's resize function and then resized back to the original size again, the mask pixels are shifted to the lower right, resulting in a significant loss of accuracy. Therefore, care should be taken in resizing.<ul>\n<li>In my solution, I used an affine transformation that simultaneously translates by 0.5 pixel and scales the image to prevent pixel misalignment.　　</li>\n<li>Incidentally, this idea is strongly influenced by the contrail competition solution.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fe326568835894e0923ac98172873977f%2Fresize.jpg?generation=1707324969473901&amp;alt=media\"></li></ul></li>\n</ul>\n<h2>2. Vessel Segmentation</h2>\n<h3>Model</h3>\n<ul>\n<li>Model was Unet (using smp implementation), resnest14d, resnest50d, maxvit_tiny were used for backbone, and ensemble with equal weights was used as final sub.  </li>\n<li>Since I wanted to use depth direction information as well, we employed a 2.5D model that takes an input image consisting of n (5 or 7) consecutive slices stacked in the ch direction and outputs the corresponding n ch masks.  </li>\n</ul>\n<h3>Data</h3>\n<ul>\n<li>The data was kidney_1_dense as training data and kidney_3_dense as validation data. Some models used kidney_2 and Pseudo Labeled data for external data as training data.  </li>\n</ul>\n<h3>Augmentation</h3>\n<ul>\n<li>Use augmentation on rotation, flipping, and brightness (using the albumentations implementation).</li>\n<li>The shape-changing type augmentaion (e.g., Distortion) was tried but was not used because it worsens the accuracy of both cv/lb.</li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>Inference in each view in XY, XZ, and ZY directions.  </li>\n<li>The accuracy of both cv/lb was increased by inputting a larger size than the image size used for training during inference.  </li>\n<li>The threshold was determined based on CV and used 0.25.  </li>\n</ul>\n<h3>Summary</h3>\n<p>The final scores are as follows.  </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>N(ch)</th>\n<th>train_data</th>\n<th>validation_data</th>\n<th>input_size(train)</th>\n<th>input_size(inference)</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>resnest14d</td>\n<td>7</td>\n<td>kidney_1_dense</td>\n<td>kidney_3_dense</td>\n<td>1920x1920</td>\n<td>2304x2304</td>\n<td>0.909</td>\n<td>0.835</td>\n<td>0.659</td>\n</tr>\n<tr>\n<td>resnest50d</td>\n<td>5</td>\n<td>kideny_1_dense</td>\n<td>kidney_3_dense</td>\n<td>1536x1536</td>\n<td>1920x1920</td>\n<td>0.903</td>\n<td>0.818</td>\n<td>0.599</td>\n</tr>\n<tr>\n<td>maxvit_tiny</td>\n<td>7</td>\n<td>kidney_1_dense, kidney_2(pseudo label), extra_data(pseudo label)</td>\n<td>kidney_3_dense</td>\n<td>1536x1536</td>\n<td>2048x2048</td>\n<td>0.901</td>\n<td>0.810</td>\n<td>0.623</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.913</td>\n<td>0.824</td>\n<td>0.645</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2641779",
      "postDate": "02/07/2024 16:59:45",
      "content": "<h1>Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F6a7d79a9543e55064e881b6961b14061%2Foverview.jpg?generation=1707324869303897&amp;alt=media\" alt=\"overview\"></p>\n<ul>\n<li>2.5D segmentation model that inputs N consecutive slices stacked in ch direction and outputs corresponding Nch masks.  </li>\n<li>Input images are cropped to the kidney area only and then resized.  </li>\n</ul>\n<h1>Pipeline Detail　　</h1>\n<h2>1. Preprocess</h2>\n<h3>1-1. Normalization</h3>\n<ul>\n<li>A histogram of luminance values is calculated for the entire kidney and normalized based on minimum and maximum values.  </li>\n<li>Normalization based on maximum and minimum values per image unit could cause variations in the appearance of images, resulting in unnatural switching of inference results. To counteract this, normalization based on the luminance distribution of the entire kidney was employed.  </li>\n<li>The code is as follows.  </li>\n</ul>\n<pre><code> ():\n    img_paths = (glob(os.path.join(image_dir, )))\n\n    pixels = np.zeros((,), dtype=np.int64)\n     img_path  tqdm(img_paths):\n        img = cv2.imread(img_path, cv2.IMREAD_UNCHANGED)\n        _pixels = np.bincount(img.flatten(), minlength=)\n        pixels += _pixels\n\n     = \n    hist = []\n    bins = []\n     i  (, +, ):\n        hist.append(pixels[i:i+].())\n        bins.append(i)\n    hist = np.array(hist)\n    hist_rate = hist/hist.()\n    idxes = np.where(hist_rate&gt;)[]\n    min_idx = idxes[]\n    max_idx = idxes[-]\n     bins[min_idx]-, bins[max_idx]+  \n\nmin_val, max_val = get_min_max_val(inference_img_dir)\nimg = cv2.imread(path, cv2.IMREAD_UNCHANGED)\nimg = img.astype()\nimg = np.clip(img, min_value, max_value)\nimg = (img-min_value)/max_value\n</code></pre>\n<h3>1-2. Crop</h3>\n<ul>\n<li>Obtain a rectangle of the kidney region using a segmentation model that infers a mask of the entire kidney.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F68c25e3b97c72b5b599b5607e8bcff8f%2Fcrop.jpg?generation=1707324937523266&amp;alt=media\" alt=\"crop\"></li>\n<li>The kidney segmentation model used a single model of Unet (backbone: efficientnet_b4) and only kidney_1_dense was used as training data.  </li>\n<li>Cropping only the kidney region eliminates wasted areas in the image and greatly improves the accuracy of vessel segmentation.  </li>\n<li>During training, the height and width of the rectangle were stochastically increased or decreased by ±5% as part of augmentation.</li>\n</ul>\n<h3>1-3. Resize</h3>\n<ul>\n<li>Because of the strict masking requirements of this competition metric, it was important to resize the image to a larger image size.  </li>\n<li>In my solution, I trained the model by resizing the image to as large as GPU memory would allow, in the range of 1536~1920.  </li>\n<li>If the mask is resized by OpenCV's resize function and then resized back to the original size again, the mask pixels are shifted to the lower right, resulting in a significant loss of accuracy. Therefore, care should be taken in resizing.<ul>\n<li>In my solution, I used an affine transformation that simultaneously translates by 0.5 pixel and scales the image to prevent pixel misalignment.　　</li>\n<li>Incidentally, this idea is strongly influenced by the contrail competition solution.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fe326568835894e0923ac98172873977f%2Fresize.jpg?generation=1707324969473901&amp;alt=media\"></li></ul></li>\n</ul>\n<h2>2. Vessel Segmentation</h2>\n<h3>Model</h3>\n<ul>\n<li>Model was Unet (using smp implementation), resnest14d, resnest50d, maxvit_tiny were used for backbone, and ensemble with equal weights was used as final sub.  </li>\n<li>Since I wanted to use depth direction information as well, we employed a 2.5D model that takes an input image consisting of n (5 or 7) consecutive slices stacked in the ch direction and outputs the corresponding n ch masks.  </li>\n</ul>\n<h3>Data</h3>\n<ul>\n<li>The data was kidney_1_dense as training data and kidney_3_dense as validation data. Some models used kidney_2 and Pseudo Labeled data for external data as training data.  </li>\n</ul>\n<h3>Augmentation</h3>\n<ul>\n<li>Use augmentation on rotation, flipping, and brightness (using the albumentations implementation).</li>\n<li>The shape-changing type augmentaion (e.g., Distortion) was tried but was not used because it worsens the accuracy of both cv/lb.</li>\n</ul>\n<h3>Inference</h3>\n<ul>\n<li>Inference in each view in XY, XZ, and ZY directions.  </li>\n<li>The accuracy of both cv/lb was increased by inputting a larger size than the image size used for training during inference.  </li>\n<li>The threshold was determined based on CV and used 0.25.  </li>\n</ul>\n<h3>Summary</h3>\n<p>The final scores are as follows.  </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>N(ch)</th>\n<th>train_data</th>\n<th>validation_data</th>\n<th>input_size(train)</th>\n<th>input_size(inference)</th>\n<th>CV</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>resnest14d</td>\n<td>7</td>\n<td>kidney_1_dense</td>\n<td>kidney_3_dense</td>\n<td>1920x1920</td>\n<td>2304x2304</td>\n<td>0.909</td>\n<td>0.835</td>\n<td>0.659</td>\n</tr>\n<tr>\n<td>resnest50d</td>\n<td>5</td>\n<td>kideny_1_dense</td>\n<td>kidney_3_dense</td>\n<td>1536x1536</td>\n<td>1920x1920</td>\n<td>0.903</td>\n<td>0.818</td>\n<td>0.599</td>\n</tr>\n<tr>\n<td>maxvit_tiny</td>\n<td>7</td>\n<td>kidney_1_dense, kidney_2(pseudo label), extra_data(pseudo label)</td>\n<td>kidney_3_dense</td>\n<td>1536x1536</td>\n<td>2048x2048</td>\n<td>0.901</td>\n<td>0.810</td>\n<td>0.623</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>-</td>\n<td>0.913</td>\n<td>0.824</td>\n<td>0.645</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# Overview  \n![overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F6a7d79a9543e55064e881b6961b14061%2Foverview.jpg?generation=1707324869303897&alt=media)\n* 2.5D segmentation model that inputs N consecutive slices stacked in ch direction and outputs corresponding Nch masks.  \n* Input images are cropped to the kidney area only and then resized.  \n\n# Pipeline Detail　　\n## 1. Preprocess  \n### 1-1. Normalization  \n* A histogram of luminance values is calculated for the entire kidney and normalized based on minimum and maximum values.  \n* Normalization based on maximum and minimum values per image unit could cause variations in the appearance of images, resulting in unnatural switching of inference results. To counteract this, normalization based on the luminance distribution of the entire kidney was employed.  \n* The code is as follows.  \n```python\ndef get_min_max_val(image_dir):\n    img_paths = sorted(glob(os.path.join(image_dir, \"*.tif\")))\n    \n    pixels = np.zeros((65536,), dtype=np.int64)\n    for img_path in tqdm(img_paths):\n        img = cv2.imread(img_path, cv2.IMREAD_UNCHANGED)\n        _pixels = np.bincount(img.flatten(), minlength=65536)\n        pixels += _pixels\n\n    bin = 1000\n    hist = []\n    bins = []\n    for i in range(0, 65535+bin, bin):\n        hist.append(pixels[i:i+bin].sum())\n        bins.append(i)\n    hist = np.array(hist)\n    hist_rate = hist/hist.max()\n    idxes = np.where(hist_rate>0.01)[0]\n    min_idx = idxes[0]\n    max_idx = idxes[-1]\n    return bins[min_idx]-5000, bins[max_idx]+5000  \n\nmin_val, max_val = get_min_max_val(inference_img_dir)\nimg = cv2.imread(path, cv2.IMREAD_UNCHANGED)\nimg = img.astype('float32')\nimg = np.clip(img, min_value, max_value)\nimg = (img-min_value)/max_value\n```\n\n### 1-2. Crop  \n* Obtain a rectangle of the kidney region using a segmentation model that infers a mask of the entire kidney.  \n![crop](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F68c25e3b97c72b5b599b5607e8bcff8f%2Fcrop.jpg?generation=1707324937523266&alt=media)\n* The kidney segmentation model used a single model of Unet (backbone: efficientnet_b4) and only kidney_1_dense was used as training data.  \n* Cropping only the kidney region eliminates wasted areas in the image and greatly improves the accuracy of vessel segmentation.  \n* During training, the height and width of the rectangle were stochastically increased or decreased by ±5% as part of augmentation.\n\n### 1-3. Resize  \n* Because of the strict masking requirements of this competition metric, it was important to resize the image to a larger image size.  \n* In my solution, I trained the model by resizing the image to as large as GPU memory would allow, in the range of 1536~1920.  \n* If the mask is resized by OpenCV's resize function and then resized back to the original size again, the mask pixels are shifted to the lower right, resulting in a significant loss of accuracy. Therefore, care should be taken in resizing.\n  * In my solution, I used an affine transformation that simultaneously translates by 0.5 pixel and scales the image to prevent pixel misalignment.　　\n  * Incidentally, this idea is strongly influenced by the contrail competition solution.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fe326568835894e0923ac98172873977f%2Fresize.jpg?generation=1707324969473901&alt=media)\n\n## 2. Vessel Segmentation  \n### Model  \n* Model was Unet (using smp implementation), resnest14d, resnest50d, maxvit_tiny were used for backbone, and ensemble with equal weights was used as final sub.  \n* Since I wanted to use depth direction information as well, we employed a 2.5D model that takes an input image consisting of n (5 or 7) consecutive slices stacked in the ch direction and outputs the corresponding n ch masks.  \n\n### Data\n* The data was kidney_1_dense as training data and kidney_3_dense as validation data. Some models used kidney_2 and Pseudo Labeled data for external data as training data.  \n\n### Augmentation\n* Use augmentation on rotation, flipping, and brightness (using the albumentations implementation).\n* The shape-changing type augmentaion (e.g., Distortion) was tried but was not used because it worsens the accuracy of both cv/lb.\n\n### Inference\n* Inference in each view in XY, XZ, and ZY directions.  \n* The accuracy of both cv/lb was increased by inputting a larger size than the image size used for training during inference.  \n* The threshold was determined based on CV and used 0.25.  \n\n### Summary  \nThe final scores are as follows.  \n| Model| N(ch) | train_data | validation_data | input_size(train) | input_size(inference) | CV | Public | Private |\n| ---  | --- | ---  | ---  | ---  | ---  |  --- | ---  | --- |\n| resnest14d | 7 | kidney_1_dense | kidney_3_dense | 1920x1920 | 2304x2304 | 0.909 | 0.835 | 0.659 |\n| resnest50d | 5 | kideny_1_dense | kidney_3_dense | 1536x1536 | 1920x1920 | 0.903 | 0.818 | 0.599 |\n| maxvit_tiny| 7 | kidney_1_dense, kidney_2(pseudo label), extra_data(pseudo label) | kidney_3_dense | 1536x1536 | 2048x2048 | 0.901 | 0.810 | 0.623 |\n| Ensemble | - | - | - | - | - | 0.913 | 0.824 | 0.645 |",
      "votes": null
    },
    {
      "id": "2642387",
      "postDate": "02/08/2024 06:17:55",
      "content": "<p>Thank you for the useful solution.</p>\n<blockquote>\n  <p>1-2. Crop<br>\n  The kidney segmentation model used a single model of Unet (backbone: efficientnet_b4) and only kidney_1_dense was used as training data.</p>\n</blockquote>\n<p>Did you make the kidney region labels yourself?</p>",
      "rawMarkdown": "Thank you for the useful solution.\n\n>1-2. Crop\n>The kidney segmentation model used a single model of Unet (backbone: efficientnet_b4) and only kidney_1_dense was used as training data.\n\nDid you make the kidney region labels yourself?",
      "votes": null
    },
    {
      "id": "2642888",
      "postDate": "02/08/2024 13:46:17",
      "content": "<p>Thanks for reading solution.</p>\n<blockquote>\n  <p>Did you make the kidney region labels yourself?</p>\n</blockquote>\n<p>No, I used the dataset published by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> .<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/blood-vessel-segmentation-kidney-mask\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/blood-vessel-segmentation-kidney-mask</a></p>",
      "rawMarkdown": "Thanks for reading solution.\n\n> Did you make the kidney region labels yourself?\n\nNo, I used the dataset published by @hengck23 .\nhttps://www.kaggle.com/datasets/hengck23/blood-vessel-segmentation-kidney-mask",
      "votes": null
    },
    {
      "id": "2642984",
      "postDate": "02/08/2024 14:40:14",
      "content": "<p>Congratulations on your silver medal! Very clear explanation!</p>",
      "rawMarkdown": "Congratulations on your silver medal! Very clear explanation!",
      "votes": null
    },
    {
      "id": "2643509",
      "postDate": "02/08/2024 21:44:54",
      "content": "<p>Congratulations, nicely explained</p>",
      "rawMarkdown": "Congratulations, nicely explained",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2642387,
      "author_name": "minfuka",
      "author_url": "",
      "post_date": "02/08/2024 06:17:55",
      "content": "<p>Thank you for the useful solution.</p>\n<blockquote>\n  <p>1-2. Crop<br>\n  The kidney segmentation model used a single model of Unet (backbone: efficientnet_b4) and only kidney_1_dense was used as training data.</p>\n</blockquote>\n<p>Did you make the kidney region labels yourself?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2642888,
          "author_name": "kashiwaba",
          "author_url": "",
          "post_date": "02/08/2024 13:46:17",
          "content": "<p>Thanks for reading solution.</p>\n<blockquote>\n  <p>Did you make the kidney region labels yourself?</p>\n</blockquote>\n<p>No, I used the dataset published by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> .<br>\n<a href=\"https://www.kaggle.com/datasets/hengck23/blood-vessel-segmentation-kidney-mask\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/blood-vessel-segmentation-kidney-mask</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2642984,
      "author_name": "garfield2021",
      "author_url": "",
      "post_date": "02/08/2024 14:40:14",
      "content": "<p>Congratulations on your silver medal! Very clear explanation!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2643509,
      "author_name": "humaperveen",
      "author_url": "",
      "post_date": "02/08/2024 21:44:54",
      "content": "<p>Congratulations, nicely explained</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2641779": "# Overview  \n![overview](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F6a7d79a9543e55064e881b6961b14061%2Foverview.jpg?generation=1707324869303897&alt=media)\n* 2.5D segmentation model that inputs N consecutive slices stacked in ch direction and outputs corresponding Nch masks.  \n* Input images are cropped to the kidney area only and then resized.  \n\n# Pipeline Detail　　\n## 1. Preprocess  \n### 1-1. Normalization  \n* A histogram of luminance values is calculated for the entire kidney and normalized based on minimum and maximum values.  \n* Normalization based on maximum and minimum values per image unit could cause variations in the appearance of images, resulting in unnatural switching of inference results. To counteract this, normalization based on the luminance distribution of the entire kidney was employed.  \n* The code is as follows.  \n```python\ndef get_min_max_val(image_dir):\n    img_paths = sorted(glob(os.path.join(image_dir, \"*.tif\")))\n    \n    pixels = np.zeros((65536,), dtype=np.int64)\n    for img_path in tqdm(img_paths):\n        img = cv2.imread(img_path, cv2.IMREAD_UNCHANGED)\n        _pixels = np.bincount(img.flatten(), minlength=65536)\n        pixels += _pixels\n\n    bin = 1000\n    hist = []\n    bins = []\n    for i in range(0, 65535+bin, bin):\n        hist.append(pixels[i:i+bin].sum())\n        bins.append(i)\n    hist = np.array(hist)\n    hist_rate = hist/hist.max()\n    idxes = np.where(hist_rate>0.01)[0]\n    min_idx = idxes[0]\n    max_idx = idxes[-1]\n    return bins[min_idx]-5000, bins[max_idx]+5000  \n\nmin_val, max_val = get_min_max_val(inference_img_dir)\nimg = cv2.imread(path, cv2.IMREAD_UNCHANGED)\nimg = img.astype('float32')\nimg = np.clip(img, min_value, max_value)\nimg = (img-min_value)/max_value\n```\n\n### 1-2. Crop  \n* Obtain a rectangle of the kidney region using a segmentation model that infers a mask of the entire kidney.  \n![crop](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2F68c25e3b97c72b5b599b5607e8bcff8f%2Fcrop.jpg?generation=1707324937523266&alt=media)\n* The kidney segmentation model used a single model of Unet (backbone: efficientnet_b4) and only kidney_1_dense was used as training data.  \n* Cropping only the kidney region eliminates wasted areas in the image and greatly improves the accuracy of vessel segmentation.  \n* During training, the height and width of the rectangle were stochastically increased or decreased by ±5% as part of augmentation.\n\n### 1-3. Resize  \n* Because of the strict masking requirements of this competition metric, it was important to resize the image to a larger image size.  \n* In my solution, I trained the model by resizing the image to as large as GPU memory would allow, in the range of 1536~1920.  \n* If the mask is resized by OpenCV's resize function and then resized back to the original size again, the mask pixels are shifted to the lower right, resulting in a significant loss of accuracy. Therefore, care should be taken in resizing.\n  * In my solution, I used an affine transformation that simultaneously translates by 0.5 pixel and scales the image to prevent pixel misalignment.　　\n  * Incidentally, this idea is strongly influenced by the contrail competition solution.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3823496%2Fe326568835894e0923ac98172873977f%2Fresize.jpg?generation=1707324969473901&alt=media)\n\n## 2. Vessel Segmentation  \n### Model  \n* Model was Unet (using smp implementation), resnest14d, resnest50d, maxvit_tiny were used for backbone, and ensemble with equal weights was used as final sub.  \n* Since I wanted to use depth direction information as well, we employed a 2.5D model that takes an input image consisting of n (5 or 7) consecutive slices stacked in the ch direction and outputs the corresponding n ch masks.  \n\n### Data\n* The data was kidney_1_dense as training data and kidney_3_dense as validation data. Some models used kidney_2 and Pseudo Labeled data for external data as training data.  \n\n### Augmentation\n* Use augmentation on rotation, flipping, and brightness (using the albumentations implementation).\n* The shape-changing type augmentaion (e.g., Distortion) was tried but was not used because it worsens the accuracy of both cv/lb.\n\n### Inference\n* Inference in each view in XY, XZ, and ZY directions.  \n* The accuracy of both cv/lb was increased by inputting a larger size than the image size used for training during inference.  \n* The threshold was determined based on CV and used 0.25.  \n\n### Summary  \nThe final scores are as follows.  \n| Model| N(ch) | train_data | validation_data | input_size(train) | input_size(inference) | CV | Public | Private |\n| ---  | --- | ---  | ---  | ---  | ---  |  --- | ---  | --- |\n| resnest14d | 7 | kidney_1_dense | kidney_3_dense | 1920x1920 | 2304x2304 | 0.909 | 0.835 | 0.659 |\n| resnest50d | 5 | kideny_1_dense | kidney_3_dense | 1536x1536 | 1920x1920 | 0.903 | 0.818 | 0.599 |\n| maxvit_tiny| 7 | kidney_1_dense, kidney_2(pseudo label), extra_data(pseudo label) | kidney_3_dense | 1536x1536 | 2048x2048 | 0.901 | 0.810 | 0.623 |\n| Ensemble | - | - | - | - | - | 0.913 | 0.824 | 0.645 |",
    "2642387": "Thank you for the useful solution.\n\n>1-2. Crop\n>The kidney segmentation model used a single model of Unet (backbone: efficientnet_b4) and only kidney_1_dense was used as training data.\n\nDid you make the kidney region labels yourself?",
    "2642888": "Thanks for reading solution.\n\n> Did you make the kidney region labels yourself?\n\nNo, I used the dataset published by @hengck23 .\nhttps://www.kaggle.com/datasets/hengck23/blood-vessel-segmentation-kidney-mask",
    "2642984": "Congratulations on your silver medal! Very clear explanation!",
    "2643509": "Congratulations, nicely explained"
  },
  "source": "meta"
}