{
  "id": 476457,
  "title": "12st Place Solution for the  SenNet + HOA - Hacking the Human Vasculature in 3D",
  "url": "/competitions/blood-vessel-segmentation/discussion/476457",
  "author_name": "Igor PI",
  "post_date": "2024-02-12T13:07:56.745000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Context</h1>\n<p>This solution was implemented as part of a blood vessel segmentation competition organized by the Common Fund’s Cellular Senescence Network (SenNet) Programm in cooperation with the Human Organ Atlas (HOA). <br>\nCompetition overview page: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation\" target=\"_blank\">SenNet + HOA - Hacking the Human Vasculature in 3D</a><br>\nCompetition dataset is <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/data\" target=\"_blank\">here</a><br>\nMany thanks to the organizers for the opportunity!</p>\n<h1>Overview</h1>\n<p>Framework — <strong>TensorFlow</strong><br>\nData pipeline — 2d, roi, resize (<strong>1024x704</strong>), <strong>tfrecord</strong><br>\nModel — almost classic <strong>U-net</strong> (details below)<br>\nThe solution is presented in two notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/pib73nl/sennet-hoa-bvs-12th-place-solution-train\" target=\"_blank\">train</a></li>\n<li><a href=\"https://www.kaggle.com/code/pib73nl/sennet-hoa-bvs-12th-place-solution-infer\" target=\"_blank\">inference</a></li>\n</ul>\n<h1>Disclamer</h1>\n<p>This solution was developed in December 2023, before Santa's New Year's gift, which ultimately helped more than seven hundred participants to jump above 0.8. I left for the holidays in 238th place with a score of 0.567. When I next opened the leaderboard a week later, I had dropped over 150 positions! And that was just the beginning! ;) I worked on this approach for another week. I must say that I did this without much enthusiasm, since by improving the metric I found myself lower and lower in the ranking.<br>\nFinaly, with a result of 0.636, which was achieved by increasing the image size and minor architecture changes, I began to look for other approaches (see below in chapter <em>„Fruitless attempts“</em>).</p>\n<h1>Data preparation</h1>\n<p>All data (except for the kidney_3_dense labels) were used as training data. The images have large fields that contain no useful information. To reduce these fields, the images were preprocessed to extract roi using statistical methods. </p>\n<pre><code> ():\n    \n    \n    row_mask = image.std(axis=)&gt;\n    clmn_mask = image.std(axis=)&gt;\n\n    \n    row_mask = cleaning_mask(row_mask)\n    clmn_mask = cleaning_mask(clmn_mask)\n\n    image = image[row_mask,:][:, clmn_mask]\n    label = label[row_mask,:][:, clmn_mask]  (label, np.ndarray)   \n\n    \n    row_pad = (row_mask.argmax(), row_mask[::-].argmax())\n    clmn_pad = (clmn_mask.argmax(), clmn_mask[::-].argmax())\n\n     image, label, (row_pad, clmn_pad)\n\n ():\n    \n    \n    mask[] = \n    mask[-] = \n\n    \n    frames = np.nonzero(mask[:-]!=mask[:])[]\n    \n    delta = frames[:]-frames[:-]\n    \n    max_solid_block_begin = np.argmax(delta)\n    \n    garbage = np.delete(frames, [max_solid_block_begin, max_solid_block_begin+])\n    \n     a, b  (garbage[::], garbage[::]):\n        mask[a+:b+] = \n\n     mask\n</code></pre>\n<p>Next, all images were reduced to a single size of 1024x704. The experiments started with a size of 384x256, and as the size increased, the result expectedly improved. 1024x704 is the maximum size that did not result in an OOM error. An example of the processed image is below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fdc967a133b4763004a3d347af9c0af4e%2Fprepr_img.png?generation=1707740226468394&amp;alt=media\" alt=\"Example of a processed image\"><br>\nEvery 25 images (4%) were used for validation, since the density of the labels varies greatly along the z-axis.<br>\nThe resulting images and tags were packed into tfrecord files to organize a multi-threaded pipeline (total files - 92, 162 MB each). The maximum possible batch size for the 1024x704 shape turned out to be 32. Augmentation was not used - I just couldn’t get around to it!</p>\n<h1>Model</h1>\n<p>The more or less classical <strong>U-net</strong> architecture was used as a model.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fc84d61b6e342a86e18e47abafb575cb5%2Funet%20arcitect.png?generation=1707741082913852&amp;alt=media\" alt=\"U-net architect\"><br>\nLosses were estimated using <strong>binary crosentropy</strong>. The <strong>Adam</strong> optimizer was used for optimization. The <strong>learning rate</strong> was changed according to the <strong>cosine decay</strong> schedule with warmup.<br>\nSince there is a significant class imbalance, weights were used. The idea was to set the weights at the instance level, since the class ratios vary greatly as we move from the center of the kidney to the edges (along the z-axis). But to begin with, I hardcode the weights, and it worked tolerably well. I didn't return to this issue later, so there is room for improvement.<br>\nThe model was created from scratch and trained for 60 epochs. For prediction, epochs with a minimum value of validation losses were taken.<br>\nThe prediction result looked something like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F7325fd2a8223112484047b8a02408aed%2Fpredict_res.png?generation=1707740910547839&amp;alt=media\" alt=\"Prediction result\"><br>\nThere are quite a lot of FP here… However, it makes sense to work on the sample weights 🤔</p>\n<h1>Fruitless attempts</h1>\n<p>Obviously, given the large number of small details, any resizing harms the result. I tried to solve this problem by dividing the image into fragments (intersecting tiles of 256x256 size). I used the same model architecture. But the labels turned out to be exclusively in the places where the tiles overlapped, and having assembled a mask from the tiles, I got a blank sheet! I haven't had time to figure this out.<br>\nSecond. I tried to solve the problem of label resizing by changing the architecture - I added another “kinda u-net” to the end of decoder - 2 convolution layers and two reconvolution ones. Didn't do well here either, but would have been in 67th place on the private leaderboard 😉<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F5d62e9402997d2db29eff231c65447af%2Fscore_of_ext_model.png?generation=1707741415352143&amp;alt=media\"></p>\n<h1>Some observations</h1>\n<p>Yes, yes… There was a big quake… For some reason, most of the solutions failed in suspiciously similar ways ;) This is clearly noticeable in the interval of about 100-600 places.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fbf0eec25d430cbc29190c3b83d5e5dfb%2FLB_quake.png?generation=1707741692231072&amp;alt=media\" alt=\"LB shake plot\"><br>\nMy main solutions, similar in network architecture and image size to the winning one, gave stable results on a public and private dataset. On a private dataset - even a little better!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F3673a9e41dfdde1120b98b7f01a75017%2Fscore_of_win_model.png?generation=1707741845733027&amp;alt=media\" alt=\"Score of the winning model\"><br>\nThe difference between the public and private data sets was 0.012 points. Such stable results in the first thousand can be counted on the fingers of one hand. In general, the variance is already normal, all that remains is to work on the bias 😁<br>\nThanks to everyone who worked on the problem! It was interesting with you! Good luck! ✋</p>",
  "messages": [
    {
      "id": 2648831,
      "postDate": "2024-02-12T13:07:56.747Z",
      "content": "<h1>Context</h1>\n<p>This solution was implemented as part of a blood vessel segmentation competition organized by the Common Fund’s Cellular Senescence Network (SenNet) Programm in cooperation with the Human Organ Atlas (HOA). <br>\nCompetition overview page: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation\" target=\"_blank\">SenNet + HOA - Hacking the Human Vasculature in 3D</a><br>\nCompetition dataset is <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/data\" target=\"_blank\">here</a><br>\nMany thanks to the organizers for the opportunity!</p>\n<h1>Overview</h1>\n<p>Framework — <strong>TensorFlow</strong><br>\nData pipeline — 2d, roi, resize (<strong>1024x704</strong>), <strong>tfrecord</strong><br>\nModel — almost classic <strong>U-net</strong> (details below)<br>\nThe solution is presented in two notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/pib73nl/sennet-hoa-bvs-12th-place-solution-train\" target=\"_blank\">train</a></li>\n<li><a href=\"https://www.kaggle.com/code/pib73nl/sennet-hoa-bvs-12th-place-solution-infer\" target=\"_blank\">inference</a></li>\n</ul>\n<h1>Disclamer</h1>\n<p>This solution was developed in December 2023, before Santa's New Year's gift, which ultimately helped more than seven hundred participants to jump above 0.8. I left for the holidays in 238th place with a score of 0.567. When I next opened the leaderboard a week later, I had dropped over 150 positions! And that was just the beginning! ;) I worked on this approach for another week. I must say that I did this without much enthusiasm, since by improving the metric I found myself lower and lower in the ranking.<br>\nFinaly, with a result of 0.636, which was achieved by increasing the image size and minor architecture changes, I began to look for other approaches (see below in chapter <em>„Fruitless attempts“</em>).</p>\n<h1>Data preparation</h1>\n<p>All data (except for the kidney_3_dense labels) were used as training data. The images have large fields that contain no useful information. To reduce these fields, the images were preprocessed to extract roi using statistical methods. </p>\n<pre><code> ():\n    \n    \n    row_mask = image.std(axis=)&gt;\n    clmn_mask = image.std(axis=)&gt;\n\n    \n    row_mask = cleaning_mask(row_mask)\n    clmn_mask = cleaning_mask(clmn_mask)\n\n    image = image[row_mask,:][:, clmn_mask]\n    label = label[row_mask,:][:, clmn_mask]  (label, np.ndarray)   \n\n    \n    row_pad = (row_mask.argmax(), row_mask[::-].argmax())\n    clmn_pad = (clmn_mask.argmax(), clmn_mask[::-].argmax())\n\n     image, label, (row_pad, clmn_pad)\n\n ():\n    \n    \n    mask[] = \n    mask[-] = \n\n    \n    frames = np.nonzero(mask[:-]!=mask[:])[]\n    \n    delta = frames[:]-frames[:-]\n    \n    max_solid_block_begin = np.argmax(delta)\n    \n    garbage = np.delete(frames, [max_solid_block_begin, max_solid_block_begin+])\n    \n     a, b  (garbage[::], garbage[::]):\n        mask[a+:b+] = \n\n     mask\n</code></pre>\n<p>Next, all images were reduced to a single size of 1024x704. The experiments started with a size of 384x256, and as the size increased, the result expectedly improved. 1024x704 is the maximum size that did not result in an OOM error. An example of the processed image is below.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fdc967a133b4763004a3d347af9c0af4e%2Fprepr_img.png?generation=1707740226468394&amp;alt=media\" alt=\"Example of a processed image\"><br>\nEvery 25 images (4%) were used for validation, since the density of the labels varies greatly along the z-axis.<br>\nThe resulting images and tags were packed into tfrecord files to organize a multi-threaded pipeline (total files - 92, 162 MB each). The maximum possible batch size for the 1024x704 shape turned out to be 32. Augmentation was not used - I just couldn’t get around to it!</p>\n<h1>Model</h1>\n<p>The more or less classical <strong>U-net</strong> architecture was used as a model.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fc84d61b6e342a86e18e47abafb575cb5%2Funet%20arcitect.png?generation=1707741082913852&amp;alt=media\" alt=\"U-net architect\"><br>\nLosses were estimated using <strong>binary crosentropy</strong>. The <strong>Adam</strong> optimizer was used for optimization. The <strong>learning rate</strong> was changed according to the <strong>cosine decay</strong> schedule with warmup.<br>\nSince there is a significant class imbalance, weights were used. The idea was to set the weights at the instance level, since the class ratios vary greatly as we move from the center of the kidney to the edges (along the z-axis). But to begin with, I hardcode the weights, and it worked tolerably well. I didn't return to this issue later, so there is room for improvement.<br>\nThe model was created from scratch and trained for 60 epochs. For prediction, epochs with a minimum value of validation losses were taken.<br>\nThe prediction result looked something like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F7325fd2a8223112484047b8a02408aed%2Fpredict_res.png?generation=1707740910547839&amp;alt=media\" alt=\"Prediction result\"><br>\nThere are quite a lot of FP here… However, it makes sense to work on the sample weights 🤔</p>\n<h1>Fruitless attempts</h1>\n<p>Obviously, given the large number of small details, any resizing harms the result. I tried to solve this problem by dividing the image into fragments (intersecting tiles of 256x256 size). I used the same model architecture. But the labels turned out to be exclusively in the places where the tiles overlapped, and having assembled a mask from the tiles, I got a blank sheet! I haven't had time to figure this out.<br>\nSecond. I tried to solve the problem of label resizing by changing the architecture - I added another “kinda u-net” to the end of decoder - 2 convolution layers and two reconvolution ones. Didn't do well here either, but would have been in 67th place on the private leaderboard 😉<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F5d62e9402997d2db29eff231c65447af%2Fscore_of_ext_model.png?generation=1707741415352143&amp;alt=media\"></p>\n<h1>Some observations</h1>\n<p>Yes, yes… There was a big quake… For some reason, most of the solutions failed in suspiciously similar ways ;) This is clearly noticeable in the interval of about 100-600 places.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fbf0eec25d430cbc29190c3b83d5e5dfb%2FLB_quake.png?generation=1707741692231072&amp;alt=media\" alt=\"LB shake plot\"><br>\nMy main solutions, similar in network architecture and image size to the winning one, gave stable results on a public and private dataset. On a private dataset - even a little better!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F3673a9e41dfdde1120b98b7f01a75017%2Fscore_of_win_model.png?generation=1707741845733027&amp;alt=media\" alt=\"Score of the winning model\"><br>\nThe difference between the public and private data sets was 0.012 points. Such stable results in the first thousand can be counted on the fingers of one hand. In general, the variance is already normal, all that remains is to work on the bias 😁<br>\nThanks to everyone who worked on the problem! It was interesting with you! Good luck! ✋</p>",
      "rawMarkdown": "# Context\n\nThis solution was implemented as part of a blood vessel segmentation competition organized by the Common Fund’s Cellular Senescence Network (SenNet) Programm in cooperation with the Human Organ Atlas (HOA). \nCompetition overview page: [SenNet + HOA - Hacking the Human Vasculature in 3D](https://www.kaggle.com/competitions/blood-vessel-segmentation)\nCompetition dataset is [here](https://www.kaggle.com/competitions/blood-vessel-segmentation/data)\nMany thanks to the organizers for the opportunity!\n# Overview\nFramework — **TensorFlow**\nData pipeline — 2d, roi, resize (**1024x704**), **tfrecord**\nModel — almost classic **U-net** (details below)\nThe solution is presented in two notebooks:\n - [train](https://www.kaggle.com/code/pib73nl/sennet-hoa-bvs-12th-place-solution-train)\n - [inference](https://www.kaggle.com/code/pib73nl/sennet-hoa-bvs-12th-place-solution-infer)\n# Disclamer\nThis solution was developed in December 2023, before Santa's New Year's gift, which ultimately helped more than seven hundred participants to jump above 0.8. I left for the holidays in 238th place with a score of 0.567. When I next opened the leaderboard a week later, I had dropped over 150 positions! And that was just the beginning! ;) I worked on this approach for another week. I must say that I did this without much enthusiasm, since by improving the metric I found myself lower and lower in the ranking.\nFinaly, with a result of 0.636, which was achieved by increasing the image size and minor architecture changes, I began to look for other approaches (see below in chapter *„Fruitless attempts“*).\n# Data preparation\nAll data (except for the kidney_3_dense labels) were used as training data. The images have large fields that contain no useful information. To reduce these fields, the images were preprocessed to extract roi using statistical methods. \n```python\ndef apply_roi(image, label=None):\n    \"\"\"\n    Exclusion of uninformative image fields \n    \"\"\"\n    # just throw out rows and columnt with low std\n    row_mask = image.std(axis=1)>0.22\n    clmn_mask = image.std(axis=0)>0.22\n    \n    # cleaning up the of noize of this approach and taking a solid region\n    row_mask = cleaning_mask(row_mask)\n    clmn_mask = cleaning_mask(clmn_mask)\n    \n    image = image[row_mask,:][:, clmn_mask]\n    label = label[row_mask,:][:, clmn_mask] if isinstance(label, np.ndarray) else None \n    \n    # remember the size of the pads for subsequent correct restoration\n    row_pad = (row_mask.argmax(), row_mask[::-1].argmax())\n    clmn_pad = (clmn_mask.argmax(), clmn_mask[::-1].argmax())\n    \n    return image, label, (row_pad, clmn_pad)\n\ndef cleaning_mask(mask):\n    \"\"\"\n    Selecting a solid region from a noisy mask\n    \"\"\"\n    # if frame starts from the first element or finishes at the last\n    mask[0] = False\n    mask[-1] = False\n    \n    # taking edges of frames\n    frames = np.nonzero(mask[:-1]!=mask[1:])[0]\n    # taking length of frames\n    delta = frames[1:]-frames[:-1]\n    # taking index of max len frame\n    max_solid_block_begin = np.argmax(delta)\n    # other is garbage\n    garbage = np.delete(frames, [max_solid_block_begin, max_solid_block_begin+1])\n    # clearing the mask\n    for a, b in zip(garbage[::2], garbage[1::2]):\n        mask[a+1:b+1] = False\n    \n    return mask\n```\nNext, all images were reduced to a single size of 1024x704. The experiments started with a size of 384x256, and as the size increased, the result expectedly improved. 1024x704 is the maximum size that did not result in an OOM error. An example of the processed image is below.\n![Example of a processed image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fdc967a133b4763004a3d347af9c0af4e%2Fprepr_img.png?generation=1707740226468394&alt=media)\nEvery 25 images (4%) were used for validation, since the density of the labels varies greatly along the z-axis.\nThe resulting images and tags were packed into tfrecord files to organize a multi-threaded pipeline (total files - 92, 162 MB each). The maximum possible batch size for the 1024x704 shape turned out to be 32. Augmentation was not used - I just couldn’t get around to it!\n# Model\nThe more or less classical **U-net** architecture was used as a model.\n![U-net architect](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fc84d61b6e342a86e18e47abafb575cb5%2Funet%20arcitect.png?generation=1707741082913852&alt=media)\nLosses were estimated using **binary crosentropy**. The **Adam** optimizer was used for optimization. The **learning rate** was changed according to the **cosine decay** schedule with warmup.\nSince there is a significant class imbalance, weights were used. The idea was to set the weights at the instance level, since the class ratios vary greatly as we move from the center of the kidney to the edges (along the z-axis). But to begin with, I hardcode the weights, and it worked tolerably well. I didn't return to this issue later, so there is room for improvement.\nThe model was created from scratch and trained for 60 epochs. For prediction, epochs with a minimum value of validation losses were taken.\nThe prediction result looked something like this:\n![Prediction result](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F7325fd2a8223112484047b8a02408aed%2Fpredict_res.png?generation=1707740910547839&alt=media)\nThere are quite a lot of FP here… However, it makes sense to work on the sample weights 🤔\n# Fruitless attempts\nObviously, given the large number of small details, any resizing harms the result. I tried to solve this problem by dividing the image into fragments (intersecting tiles of 256x256 size). I used the same model architecture. But the labels turned out to be exclusively in the places where the tiles overlapped, and having assembled a mask from the tiles, I got a blank sheet! I haven't had time to figure this out.\nSecond. I tried to solve the problem of label resizing by changing the architecture - I added another “kinda u-net” to the end of decoder - 2 convolution layers and two reconvolution ones. Didn't do well here either, but would have been in 67th place on the private leaderboard 😉\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F5d62e9402997d2db29eff231c65447af%2Fscore_of_ext_model.png?generation=1707741415352143&alt=media)\n# Some observations\nYes, yes… There was a big quake… For some reason, most of the solutions failed in suspiciously similar ways ;) This is clearly noticeable in the interval of about 100-600 places.\n![LB shake plot](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fbf0eec25d430cbc29190c3b83d5e5dfb%2FLB_quake.png?generation=1707741692231072&alt=media)\nMy main solutions, similar in network architecture and image size to the winning one, gave stable results on a public and private dataset. On a private dataset - even a little better!\n![Score of the winning model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F3673a9e41dfdde1120b98b7f01a75017%2Fscore_of_win_model.png?generation=1707741845733027&alt=media)\nThe difference between the public and private data sets was 0.012 points. Such stable results in the first thousand can be counted on the fingers of one hand. In general, the variance is already normal, all that remains is to work on the bias 😁\nThanks to everyone who worked on the problem! It was interesting with you! Good luck! ✋",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2648831": "# Context\n\nThis solution was implemented as part of a blood vessel segmentation competition organized by the Common Fund’s Cellular Senescence Network (SenNet) Programm in cooperation with the Human Organ Atlas (HOA). \nCompetition overview page: [SenNet + HOA - Hacking the Human Vasculature in 3D](https://www.kaggle.com/competitions/blood-vessel-segmentation)\nCompetition dataset is [here](https://www.kaggle.com/competitions/blood-vessel-segmentation/data)\nMany thanks to the organizers for the opportunity!\n# Overview\nFramework — **TensorFlow**\nData pipeline — 2d, roi, resize (**1024x704**), **tfrecord**\nModel — almost classic **U-net** (details below)\nThe solution is presented in two notebooks:\n - [train](https://www.kaggle.com/code/pib73nl/sennet-hoa-bvs-12th-place-solution-train)\n - [inference](https://www.kaggle.com/code/pib73nl/sennet-hoa-bvs-12th-place-solution-infer)\n# Disclamer\nThis solution was developed in December 2023, before Santa's New Year's gift, which ultimately helped more than seven hundred participants to jump above 0.8. I left for the holidays in 238th place with a score of 0.567. When I next opened the leaderboard a week later, I had dropped over 150 positions! And that was just the beginning! ;) I worked on this approach for another week. I must say that I did this without much enthusiasm, since by improving the metric I found myself lower and lower in the ranking.\nFinaly, with a result of 0.636, which was achieved by increasing the image size and minor architecture changes, I began to look for other approaches (see below in chapter *„Fruitless attempts“*).\n# Data preparation\nAll data (except for the kidney_3_dense labels) were used as training data. The images have large fields that contain no useful information. To reduce these fields, the images were preprocessed to extract roi using statistical methods. \n```python\ndef apply_roi(image, label=None):\n    \"\"\"\n    Exclusion of uninformative image fields \n    \"\"\"\n    # just throw out rows and columnt with low std\n    row_mask = image.std(axis=1)>0.22\n    clmn_mask = image.std(axis=0)>0.22\n    \n    # cleaning up the of noize of this approach and taking a solid region\n    row_mask = cleaning_mask(row_mask)\n    clmn_mask = cleaning_mask(clmn_mask)\n    \n    image = image[row_mask,:][:, clmn_mask]\n    label = label[row_mask,:][:, clmn_mask] if isinstance(label, np.ndarray) else None \n    \n    # remember the size of the pads for subsequent correct restoration\n    row_pad = (row_mask.argmax(), row_mask[::-1].argmax())\n    clmn_pad = (clmn_mask.argmax(), clmn_mask[::-1].argmax())\n    \n    return image, label, (row_pad, clmn_pad)\n\ndef cleaning_mask(mask):\n    \"\"\"\n    Selecting a solid region from a noisy mask\n    \"\"\"\n    # if frame starts from the first element or finishes at the last\n    mask[0] = False\n    mask[-1] = False\n    \n    # taking edges of frames\n    frames = np.nonzero(mask[:-1]!=mask[1:])[0]\n    # taking length of frames\n    delta = frames[1:]-frames[:-1]\n    # taking index of max len frame\n    max_solid_block_begin = np.argmax(delta)\n    # other is garbage\n    garbage = np.delete(frames, [max_solid_block_begin, max_solid_block_begin+1])\n    # clearing the mask\n    for a, b in zip(garbage[::2], garbage[1::2]):\n        mask[a+1:b+1] = False\n    \n    return mask\n```\nNext, all images were reduced to a single size of 1024x704. The experiments started with a size of 384x256, and as the size increased, the result expectedly improved. 1024x704 is the maximum size that did not result in an OOM error. An example of the processed image is below.\n![Example of a processed image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fdc967a133b4763004a3d347af9c0af4e%2Fprepr_img.png?generation=1707740226468394&alt=media)\nEvery 25 images (4%) were used for validation, since the density of the labels varies greatly along the z-axis.\nThe resulting images and tags were packed into tfrecord files to organize a multi-threaded pipeline (total files - 92, 162 MB each). The maximum possible batch size for the 1024x704 shape turned out to be 32. Augmentation was not used - I just couldn’t get around to it!\n# Model\nThe more or less classical **U-net** architecture was used as a model.\n![U-net architect](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fc84d61b6e342a86e18e47abafb575cb5%2Funet%20arcitect.png?generation=1707741082913852&alt=media)\nLosses were estimated using **binary crosentropy**. The **Adam** optimizer was used for optimization. The **learning rate** was changed according to the **cosine decay** schedule with warmup.\nSince there is a significant class imbalance, weights were used. The idea was to set the weights at the instance level, since the class ratios vary greatly as we move from the center of the kidney to the edges (along the z-axis). But to begin with, I hardcode the weights, and it worked tolerably well. I didn't return to this issue later, so there is room for improvement.\nThe model was created from scratch and trained for 60 epochs. For prediction, epochs with a minimum value of validation losses were taken.\nThe prediction result looked something like this:\n![Prediction result](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F7325fd2a8223112484047b8a02408aed%2Fpredict_res.png?generation=1707740910547839&alt=media)\nThere are quite a lot of FP here… However, it makes sense to work on the sample weights 🤔\n# Fruitless attempts\nObviously, given the large number of small details, any resizing harms the result. I tried to solve this problem by dividing the image into fragments (intersecting tiles of 256x256 size). I used the same model architecture. But the labels turned out to be exclusively in the places where the tiles overlapped, and having assembled a mask from the tiles, I got a blank sheet! I haven't had time to figure this out.\nSecond. I tried to solve the problem of label resizing by changing the architecture - I added another “kinda u-net” to the end of decoder - 2 convolution layers and two reconvolution ones. Didn't do well here either, but would have been in 67th place on the private leaderboard 😉\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F5d62e9402997d2db29eff231c65447af%2Fscore_of_ext_model.png?generation=1707741415352143&alt=media)\n# Some observations\nYes, yes… There was a big quake… For some reason, most of the solutions failed in suspiciously similar ways ;) This is clearly noticeable in the interval of about 100-600 places.\n![LB shake plot](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2Fbf0eec25d430cbc29190c3b83d5e5dfb%2FLB_quake.png?generation=1707741692231072&alt=media)\nMy main solutions, similar in network architecture and image size to the winning one, gave stable results on a public and private dataset. On a private dataset - even a little better!\n![Score of the winning model](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F937321%2F3673a9e41dfdde1120b98b7f01a75017%2Fscore_of_win_model.png?generation=1707741845733027&alt=media)\nThe difference between the public and private data sets was 0.012 points. Such stable results in the first thousand can be counted on the fingers of one hand. In general, the variance is already normal, all that remains is to work on the bias 😁\nThanks to everyone who worked on the problem! It was interesting with you! Good luck! ✋"
  }
}