{
  "id": 417444,
  "title": "13th place solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/417444",
  "author_name": "Mikhail Kotyushev",
  "post_date": "2023-06-15T18:50:15.517000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello everyone! Here is a write-up of my best private leaderboard solution (top 13 with 0.641814 score, 0.723345 on public LB). </p>\n<p>I am very grateful to organizers and participants for this competition, it was very interesting and challenging. I've learned a lot about segmentation and had a lot of fun! Some of the ideas I have used in the solution came from code / discussion tabs on Kaggle, hopefully I have mentioned them all.</p>\n<p>If you have any questions, please do not hesitate to reach me in comments.</p>\n<h1>Overview</h1>\n<p>Data:</p>\n<ul>\n<li>Z shift and scale pre-processing based on mean layer intensity curve fitting to fragment 3's mean intensity curve (inspired by <a href=\"https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis\" target=\"_blank\">this</a> notebook)</li>\n<li>5 fold training (1, 2a, 2b, 2c, 3 fragments, 2nd fragment is split by scroll mask area: top part + 2 bottom parts)</li>\n<li>weighted sampling from fragments according to scroll mask area</li>\n<li>patch size is 384, on test overlap is 192 via spline-weighted averaging (inspired by <a href=\"https://github.com/bnsreenu/python_for_microscopists/blob/master/229_smooth_predictions_by_blending_patches/smooth_tiled_predictions.py\" target=\"_blank\">this</a> approach)</li>\n<li>24 slices with indices in [20, 44), random crop of 18 of them, each 3 slices window is fed to model and 6 resulting logits aggregated via linear layer</li>\n</ul>\n<p>Training:</p>\n<ul>\n<li>model is SMP Unet with 2D pre-trained encoder + linear layer aggregation</li>\n<li>random crop inside scroll mask + standard vision augmentations</li>\n<li>BCE loss with 0.5 weight of positive examples, selection by F0.5 score on validation fold</li>\n<li>64 epoch training, epoch length is ~ number of patches in train set (it is not constant because of random cropping)</li>\n<li>LR is 1e-4, no layers freezing / LR decay, schedule is cosine annealing with warmup (1e-1, 1, 1e-2 factors, 10% warmup), optimizer is AdamW</li>\n</ul>\n<p>Inference:</p>\n<ul>\n<li>TTA: cartesian product of 4 rotations (0, 90, 180, 270) and 3 flips (no flip, horizontal, vertical) -&gt; 12 predictions per volume (brought this idea from <a href=\"https://www.kaggle.com/code/yoyobar/2-5d-segmentaion-model-with-rotate-tta\" target=\"_blank\">here</a>)</li>\n<li>TTA probabilities are averaged, writen to image file, then 5 models images averaged</li>\n</ul>\n<p>Environment &amp; tools:</p>\n<ul>\n<li>docker, pytorch lightning CLI, wandb</li>\n<li>1 x 3090, 1 x 3080Ti mobile</li>\n</ul>\n<h1>Some details</h1>\n<h2>Data</h2>\n<p>Inspired by <a href=\"https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis\" target=\"_blank\">this</a> analysis, I plot the mean intensity of each layer of each fragment for full image, only for ink and only for backgound.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F011cb02384ff04722ab4cdd093d4a888%2Fintencity_curve.png?generation=1686855075215824&amp;alt=media\" alt=\"Mean intensity of each layer of each fragment\"></p>\n<p>Here it could be observed that fragments 1 and 3 are more or less similar in terms of both intencity values and minimum and maximum positions, while fragment 2 is different. </p>\n<p>My physical intuition behind this is following: papirus seems not being aligned by depth inside 3D volume, moreover papirus thinknes seems to vary between fragments 1, 3 and fragment 2. </p>\n<p>So, I've decided to align and stretch / compress papirus in z dimension by shift and scale operation:</p>\n<ol>\n<li>Split the volume on overlapping patches in H and W dimentions</li>\n<li>For each patch </li>\n</ol>\n<ul>\n<li>calculate mean intensity of each layer</li>\n<li>fit <code>z_shift</code> parameter to minimize following loss (<code>y_target</code> here is the intensity curve of full fragment 3)</li>\n</ul>\n<pre><code>z_shifted = z + z_shift\nf = interpolate.interp1d(z_shifted, y, =, =)\nreturn f(z_target) - y_target\n</code></pre>\n<ol>\n<li>For full volume fit single <code>z_scale</code> (assuming pairus thickness is same for same fragment) in a similar manner.</li>\n<li>Linearly upsample resulting <code>z_shift</code> and <code>z_scale</code> maps to full volume size and save them on disk.</li>\n<li>Apply <code>z_shift</code> and <code>z_scale</code> to full volume via <code>scipy.ndimage.geometric_transform</code> with transform writen as C module for speed.</li>\n<li>For training save it on disk, for inference apply on the fly.</li>\n</ol>\n<p>Resulting <code>z_shift</code> maps and <code>z_scale</code> values are following:</p>\n<table>\n<thead>\n<tr>\n<th>Fragment 1</th>\n<th>Fragment 2</th>\n<th>Fragment 3</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F387734e4ccdf846560e4f61421ffb52c%2F1_z_shift.png?generation=1686854266142975&amp;alt=media\" alt=\"Z shift map and its histogram\"></td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F4a66d91e17b7a2332b5574defb0d1b37%2F2_z_shift.png?generation=1686854298329163&amp;alt=media\" alt=\"Z scale map and its histogram\"></td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3510608c738e619389b036e2d1d871ef%2F3_z_shift.png?generation=1686854366312885&amp;alt=media\" alt=\"Z scale map and its histogram\"></td>\n</tr>\n</tbody>\n</table>\n<p>Z scale is ~ 1.0 for fragments 1 and 3 and ~ 0.6 for fragment 2 which could indicate the fragment 2 being ~ twice thicker that others.</p>\n<p>Such operation should be beneficial for 2D approaches. It seems to improve CV on ~0.03-0.05 F05, but I did not run full comparison due to time constraints.</p>\n<p>Also, I've tried to fit / calculate normalization for each fragment separately to completely fit the intencity curve, but it did not work well.</p>\n<h2>Model</h2>\n<p>Pre-trained maxvit encoder + SMP 2D unet with aggregation has shown the best CV score, so <code>maxvit_rmlp_base_rw_384.sw_in12k_ft_in1k</code> is the one used in best solution. </p>\n<p>2D unet with aggregation works as follows: unet model is applied on each slice of the input with size = step = 3 along z dimention (total 6 slices), then 6 resulting logits are aggregated via linear layer.</p>\n<p>Usage of pre-trained models is yielding far better results that random initialization. Unfortunately I did not came up with better architecture allowing to handle 3D inputs and 2D outputs while allowing to use pre-trained models.</p>\n<p>Best CV scores of ths model and corresponding predictions are following:</p>\n<table>\n<thead>\n<tr>\n<th>Fragment</th>\n<th>F05</th>\n<th>Prediction</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.6544</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fc9e6f42f6cf07b913a23b6fd989076d0%2F1_3jzUWpdZJR.png?generation=1686854422630598&amp;alt=media\" alt=\"Fragment 1 probabilities\"></td>\n</tr>\n<tr>\n<td>2a</td>\n<td>0.6643</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Feeef4afaf95c79def86b96014d5a7b5b%2F2a_jP08265cbJ.png?generation=1686854447976196&amp;alt=media\" alt=\"Fragment 2a probabilities\"></td>\n</tr>\n<tr>\n<td>2b</td>\n<td>0.7489</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fe70babdbeb33a34a0976195453bf995d%2F2b_whqg8ITAI7.png?generation=1686854469983013&amp;alt=media\" alt=\"Fragment 2b probabilities\"></td>\n</tr>\n<tr>\n<td>2c</td>\n<td>0.6583</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fff2531343b00439ebdb81f310894b61b%2F2c_Aomw3d44sk.png?generation=1686854491195589&amp;alt=media\" alt=\"Fragment 2c probabilities\"></td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.7027</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F626278b3fa70ceb963b517c5069f1d3f%2F3_AM8P9eim4r.png?generation=1686854513349728&amp;alt=media\" alt=\"Fragment 3 probabilities\"></td>\n</tr>\n</tbody>\n</table>\n<p><strong>Tried, but did not work well</strong></p>\n<p>Multiple backbone models:</p>\n<ul>\n<li>convnext v2</li>\n<li>swin transformer v2</li>\n<li>caformer</li>\n<li>eve02</li>\n<li>maxvit </li>\n</ul>\n<p>And multiple full model approaches: full 3D with aggregation on the last layer, 2.5D, 2D with aggregation, custom 3D models with pre-trained weights conversion from 2D to 3D.</p>\n<h2>Transforms</h2>\n<p>Following augmentations are used for training</p>\n<pre><code>train_transform = A.Compose(\n    [\n        RandomCropVolumeInside2dMask(\n            =self.hparams.img_size, \n            =self.hparams.img_size_z,\n            scale=(0.5, 2.0),\n            ratio=(0.9, 1.1),\n            scale_z=(1 / self.hparams.z_scale_limit, self.hparams.z_scale_limit),\n            =,\n            =0,\n        ),\n        A.Rotate(\n            =0.5, \n            =rotate_limit_degrees_xy, \n            =,\n        ),\n        ResizeVolume(\n            =self.hparams.img_size, \n            =self.hparams.img_size,\n            =self.hparams.img_size_z,\n            =,\n        ),\n        A.HorizontalFlip(=0.5),\n        A.VerticalFlip(=0.5),\n        A.RandomRotate90(=0.5),\n        A.RandomBrightnessContrast(=0.5, =0.1, =0.1),\n        A.OneOf(\n            [\n                A.GaussNoise(var_limit=[10, 50]),\n                A.GaussianBlur(),\n                A.MotionBlur(),\n            ], \n            =0.4\n        ),\n        A.GridDistortion(=5, =0.3, =0.5),\n        A.CoarseDropout(\n            =1, \n            =int(self.hparams.img_size * 0.3), \n            =int(self.hparams.img_size * 0.3), \n            =0, =0.5\n        ),\n        A.Normalize(\n            =MAX_PIXEL_VALUE,\n            =self.train_volume_mean,\n            =self.train_volume_std,\n            =,\n        ),\n        ToTensorV2(),\n        ToCHWD(=),\n    ],\n)\n</code></pre>\n<p>Here, <code>RandomCropVolumeInside2dMask</code> crops random 3D patch from volume so that its center's (x, y) are inside mask, <code>ResizeVolume</code> is trilinear 3D resize and <code>ToCHWD</code> simply permutes dimentions. <code>img_size</code> = 384, <code>img_size_z</code> = 18, <code>z_scale_limit</code> =  1.33, <code>rotate_limit_degrees_xy</code> = 45, <code>train_volume_mean</code> and <code>train_volume_std</code> are average of imagenet stats.</p>\n<p>For inference, following transforms are used</p>\n<pre><code>test_transform = A.Compose(\n    [\n        CenterCropVolume(\n            =None, \n            =None,\n            =math.ceil(self.hparams.img_size_z * self.hparams.z_scale_limit),\n            =,\n            =,\n            =1.0,\n        ),\n        ResizeVolume(\n            =self.hparams.img_size, \n            =self.hparams.img_size,\n            =self.hparams.img_size_z,\n            =,\n        ),\n        A.Normalize(\n            =MAX_PIXEL_VALUE,\n            =self.train_volume_mean,\n            =self.train_volume_std,\n            =,\n        ),\n        ToTensorV2(),\n        ToCHWD(=),\n    ],\n)\n</code></pre>\n<p>Here, <code>CenterCropVolume</code> is center crop of 3D volume.</p>\n<p>Mixing approaches (cutmix, mixup, custom stuff like copy-paste the positive class voxels) also were tried but shown contradicting results, so were not included.</p>\n<p>Cartesian product of <code>no flip / H flip / V flip</code> and <code>no 90 rotation / 90 rotation / 180 rotation / 270 rotation</code> is used for TTA yielding total 12 predictions.</p>\n<p><strong>Update</strong>: project source code for training could be found <a href=\"https://github.com/mkotyushev/scrolls\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": 2304186,
      "postDate": "2023-06-15T18:50:15.517Z",
      "content": "<p>Hello everyone! Here is a write-up of my best private leaderboard solution (top 13 with 0.641814 score, 0.723345 on public LB). </p>\n<p>I am very grateful to organizers and participants for this competition, it was very interesting and challenging. I've learned a lot about segmentation and had a lot of fun! Some of the ideas I have used in the solution came from code / discussion tabs on Kaggle, hopefully I have mentioned them all.</p>\n<p>If you have any questions, please do not hesitate to reach me in comments.</p>\n<h1>Overview</h1>\n<p>Data:</p>\n<ul>\n<li>Z shift and scale pre-processing based on mean layer intensity curve fitting to fragment 3's mean intensity curve (inspired by <a href=\"https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis\" target=\"_blank\">this</a> notebook)</li>\n<li>5 fold training (1, 2a, 2b, 2c, 3 fragments, 2nd fragment is split by scroll mask area: top part + 2 bottom parts)</li>\n<li>weighted sampling from fragments according to scroll mask area</li>\n<li>patch size is 384, on test overlap is 192 via spline-weighted averaging (inspired by <a href=\"https://github.com/bnsreenu/python_for_microscopists/blob/master/229_smooth_predictions_by_blending_patches/smooth_tiled_predictions.py\" target=\"_blank\">this</a> approach)</li>\n<li>24 slices with indices in [20, 44), random crop of 18 of them, each 3 slices window is fed to model and 6 resulting logits aggregated via linear layer</li>\n</ul>\n<p>Training:</p>\n<ul>\n<li>model is SMP Unet with 2D pre-trained encoder + linear layer aggregation</li>\n<li>random crop inside scroll mask + standard vision augmentations</li>\n<li>BCE loss with 0.5 weight of positive examples, selection by F0.5 score on validation fold</li>\n<li>64 epoch training, epoch length is ~ number of patches in train set (it is not constant because of random cropping)</li>\n<li>LR is 1e-4, no layers freezing / LR decay, schedule is cosine annealing with warmup (1e-1, 1, 1e-2 factors, 10% warmup), optimizer is AdamW</li>\n</ul>\n<p>Inference:</p>\n<ul>\n<li>TTA: cartesian product of 4 rotations (0, 90, 180, 270) and 3 flips (no flip, horizontal, vertical) -&gt; 12 predictions per volume (brought this idea from <a href=\"https://www.kaggle.com/code/yoyobar/2-5d-segmentaion-model-with-rotate-tta\" target=\"_blank\">here</a>)</li>\n<li>TTA probabilities are averaged, writen to image file, then 5 models images averaged</li>\n</ul>\n<p>Environment &amp; tools:</p>\n<ul>\n<li>docker, pytorch lightning CLI, wandb</li>\n<li>1 x 3090, 1 x 3080Ti mobile</li>\n</ul>\n<h1>Some details</h1>\n<h2>Data</h2>\n<p>Inspired by <a href=\"https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis\" target=\"_blank\">this</a> analysis, I plot the mean intensity of each layer of each fragment for full image, only for ink and only for backgound.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F011cb02384ff04722ab4cdd093d4a888%2Fintencity_curve.png?generation=1686855075215824&amp;alt=media\" alt=\"Mean intensity of each layer of each fragment\"></p>\n<p>Here it could be observed that fragments 1 and 3 are more or less similar in terms of both intencity values and minimum and maximum positions, while fragment 2 is different. </p>\n<p>My physical intuition behind this is following: papirus seems not being aligned by depth inside 3D volume, moreover papirus thinknes seems to vary between fragments 1, 3 and fragment 2. </p>\n<p>So, I've decided to align and stretch / compress papirus in z dimension by shift and scale operation:</p>\n<ol>\n<li>Split the volume on overlapping patches in H and W dimentions</li>\n<li>For each patch </li>\n</ol>\n<ul>\n<li>calculate mean intensity of each layer</li>\n<li>fit <code>z_shift</code> parameter to minimize following loss (<code>y_target</code> here is the intensity curve of full fragment 3)</li>\n</ul>\n<pre><code>z_shifted = z + z_shift\nf = interpolate.interp1d(z_shifted, y, =, =)\nreturn f(z_target) - y_target\n</code></pre>\n<ol>\n<li>For full volume fit single <code>z_scale</code> (assuming pairus thickness is same for same fragment) in a similar manner.</li>\n<li>Linearly upsample resulting <code>z_shift</code> and <code>z_scale</code> maps to full volume size and save them on disk.</li>\n<li>Apply <code>z_shift</code> and <code>z_scale</code> to full volume via <code>scipy.ndimage.geometric_transform</code> with transform writen as C module for speed.</li>\n<li>For training save it on disk, for inference apply on the fly.</li>\n</ol>\n<p>Resulting <code>z_shift</code> maps and <code>z_scale</code> values are following:</p>\n<table>\n<thead>\n<tr>\n<th>Fragment 1</th>\n<th>Fragment 2</th>\n<th>Fragment 3</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F387734e4ccdf846560e4f61421ffb52c%2F1_z_shift.png?generation=1686854266142975&amp;alt=media\" alt=\"Z shift map and its histogram\"></td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F4a66d91e17b7a2332b5574defb0d1b37%2F2_z_shift.png?generation=1686854298329163&amp;alt=media\" alt=\"Z scale map and its histogram\"></td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3510608c738e619389b036e2d1d871ef%2F3_z_shift.png?generation=1686854366312885&amp;alt=media\" alt=\"Z scale map and its histogram\"></td>\n</tr>\n</tbody>\n</table>\n<p>Z scale is ~ 1.0 for fragments 1 and 3 and ~ 0.6 for fragment 2 which could indicate the fragment 2 being ~ twice thicker that others.</p>\n<p>Such operation should be beneficial for 2D approaches. It seems to improve CV on ~0.03-0.05 F05, but I did not run full comparison due to time constraints.</p>\n<p>Also, I've tried to fit / calculate normalization for each fragment separately to completely fit the intencity curve, but it did not work well.</p>\n<h2>Model</h2>\n<p>Pre-trained maxvit encoder + SMP 2D unet with aggregation has shown the best CV score, so <code>maxvit_rmlp_base_rw_384.sw_in12k_ft_in1k</code> is the one used in best solution. </p>\n<p>2D unet with aggregation works as follows: unet model is applied on each slice of the input with size = step = 3 along z dimention (total 6 slices), then 6 resulting logits are aggregated via linear layer.</p>\n<p>Usage of pre-trained models is yielding far better results that random initialization. Unfortunately I did not came up with better architecture allowing to handle 3D inputs and 2D outputs while allowing to use pre-trained models.</p>\n<p>Best CV scores of ths model and corresponding predictions are following:</p>\n<table>\n<thead>\n<tr>\n<th>Fragment</th>\n<th>F05</th>\n<th>Prediction</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.6544</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fc9e6f42f6cf07b913a23b6fd989076d0%2F1_3jzUWpdZJR.png?generation=1686854422630598&amp;alt=media\" alt=\"Fragment 1 probabilities\"></td>\n</tr>\n<tr>\n<td>2a</td>\n<td>0.6643</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Feeef4afaf95c79def86b96014d5a7b5b%2F2a_jP08265cbJ.png?generation=1686854447976196&amp;alt=media\" alt=\"Fragment 2a probabilities\"></td>\n</tr>\n<tr>\n<td>2b</td>\n<td>0.7489</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fe70babdbeb33a34a0976195453bf995d%2F2b_whqg8ITAI7.png?generation=1686854469983013&amp;alt=media\" alt=\"Fragment 2b probabilities\"></td>\n</tr>\n<tr>\n<td>2c</td>\n<td>0.6583</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fff2531343b00439ebdb81f310894b61b%2F2c_Aomw3d44sk.png?generation=1686854491195589&amp;alt=media\" alt=\"Fragment 2c probabilities\"></td>\n</tr>\n<tr>\n<td>3</td>\n<td>0.7027</td>\n<td><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F626278b3fa70ceb963b517c5069f1d3f%2F3_AM8P9eim4r.png?generation=1686854513349728&amp;alt=media\" alt=\"Fragment 3 probabilities\"></td>\n</tr>\n</tbody>\n</table>\n<p><strong>Tried, but did not work well</strong></p>\n<p>Multiple backbone models:</p>\n<ul>\n<li>convnext v2</li>\n<li>swin transformer v2</li>\n<li>caformer</li>\n<li>eve02</li>\n<li>maxvit </li>\n</ul>\n<p>And multiple full model approaches: full 3D with aggregation on the last layer, 2.5D, 2D with aggregation, custom 3D models with pre-trained weights conversion from 2D to 3D.</p>\n<h2>Transforms</h2>\n<p>Following augmentations are used for training</p>\n<pre><code>train_transform = A.Compose(\n    [\n        RandomCropVolumeInside2dMask(\n            =self.hparams.img_size, \n            =self.hparams.img_size_z,\n            scale=(0.5, 2.0),\n            ratio=(0.9, 1.1),\n            scale_z=(1 / self.hparams.z_scale_limit, self.hparams.z_scale_limit),\n            =,\n            =0,\n        ),\n        A.Rotate(\n            =0.5, \n            =rotate_limit_degrees_xy, \n            =,\n        ),\n        ResizeVolume(\n            =self.hparams.img_size, \n            =self.hparams.img_size,\n            =self.hparams.img_size_z,\n            =,\n        ),\n        A.HorizontalFlip(=0.5),\n        A.VerticalFlip(=0.5),\n        A.RandomRotate90(=0.5),\n        A.RandomBrightnessContrast(=0.5, =0.1, =0.1),\n        A.OneOf(\n            [\n                A.GaussNoise(var_limit=[10, 50]),\n                A.GaussianBlur(),\n                A.MotionBlur(),\n            ], \n            =0.4\n        ),\n        A.GridDistortion(=5, =0.3, =0.5),\n        A.CoarseDropout(\n            =1, \n            =int(self.hparams.img_size * 0.3), \n            =int(self.hparams.img_size * 0.3), \n            =0, =0.5\n        ),\n        A.Normalize(\n            =MAX_PIXEL_VALUE,\n            =self.train_volume_mean,\n            =self.train_volume_std,\n            =,\n        ),\n        ToTensorV2(),\n        ToCHWD(=),\n    ],\n)\n</code></pre>\n<p>Here, <code>RandomCropVolumeInside2dMask</code> crops random 3D patch from volume so that its center's (x, y) are inside mask, <code>ResizeVolume</code> is trilinear 3D resize and <code>ToCHWD</code> simply permutes dimentions. <code>img_size</code> = 384, <code>img_size_z</code> = 18, <code>z_scale_limit</code> =  1.33, <code>rotate_limit_degrees_xy</code> = 45, <code>train_volume_mean</code> and <code>train_volume_std</code> are average of imagenet stats.</p>\n<p>For inference, following transforms are used</p>\n<pre><code>test_transform = A.Compose(\n    [\n        CenterCropVolume(\n            =None, \n            =None,\n            =math.ceil(self.hparams.img_size_z * self.hparams.z_scale_limit),\n            =,\n            =,\n            =1.0,\n        ),\n        ResizeVolume(\n            =self.hparams.img_size, \n            =self.hparams.img_size,\n            =self.hparams.img_size_z,\n            =,\n        ),\n        A.Normalize(\n            =MAX_PIXEL_VALUE,\n            =self.train_volume_mean,\n            =self.train_volume_std,\n            =,\n        ),\n        ToTensorV2(),\n        ToCHWD(=),\n    ],\n)\n</code></pre>\n<p>Here, <code>CenterCropVolume</code> is center crop of 3D volume.</p>\n<p>Mixing approaches (cutmix, mixup, custom stuff like copy-paste the positive class voxels) also were tried but shown contradicting results, so were not included.</p>\n<p>Cartesian product of <code>no flip / H flip / V flip</code> and <code>no 90 rotation / 90 rotation / 180 rotation / 270 rotation</code> is used for TTA yielding total 12 predictions.</p>\n<p><strong>Update</strong>: project source code for training could be found <a href=\"https://github.com/mkotyushev/scrolls\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Hello everyone! Here is a write-up of my best private leaderboard solution (top 13 with 0.641814 score, 0.723345 on public LB). \n\nI am very grateful to organizers and participants for this competition, it was very interesting and challenging. I've learned a lot about segmentation and had a lot of fun! Some of the ideas I have used in the solution came from code / discussion tabs on Kaggle, hopefully I have mentioned them all.\n\nIf you have any questions, please do not hesitate to reach me in comments.\n\n# Overview\n\nData:\n- Z shift and scale pre-processing based on mean layer intensity curve fitting to fragment 3's mean intensity curve (inspired by [this](https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis) notebook)\n- 5 fold training (1, 2a, 2b, 2c, 3 fragments, 2nd fragment is split by scroll mask area: top part + 2 bottom parts)\n- weighted sampling from fragments according to scroll mask area\n- patch size is 384, on test overlap is 192 via spline-weighted averaging (inspired by [this](https://github.com/bnsreenu/python_for_microscopists/blob/master/229_smooth_predictions_by_blending_patches/smooth_tiled_predictions.py) approach)\n- 24 slices with indices in [20, 44), random crop of 18 of them, each 3 slices window is fed to model and 6 resulting logits aggregated via linear layer\n\nTraining:\n- model is SMP Unet with 2D pre-trained encoder + linear layer aggregation\n- random crop inside scroll mask + standard vision augmentations\n- BCE loss with 0.5 weight of positive examples, selection by F0.5 score on validation fold\n- 64 epoch training, epoch length is ~ number of patches in train set (it is not constant because of random cropping)\n- LR is 1e-4, no layers freezing / LR decay, schedule is cosine annealing with warmup (1e-1, 1, 1e-2 factors, 10% warmup), optimizer is AdamW\n\nInference:\n- TTA: cartesian product of 4 rotations (0, 90, 180, 270) and 3 flips (no flip, horizontal, vertical) -> 12 predictions per volume (brought this idea from [here](https://www.kaggle.com/code/yoyobar/2-5d-segmentaion-model-with-rotate-tta))\n- TTA probabilities are averaged, writen to image file, then 5 models images averaged\n\nEnvironment & tools:\n- docker, pytorch lightning CLI, wandb\n- 1 x 3090, 1 x 3080Ti mobile\n\n# Some details\n## Data\n\nInspired by [this](https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis) analysis, I plot the mean intensity of each layer of each fragment for full image, only for ink and only for backgound.\n\n![Mean intensity of each layer of each fragment](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F011cb02384ff04722ab4cdd093d4a888%2Fintencity_curve.png?generation=1686855075215824&alt=media)\n\nHere it could be observed that fragments 1 and 3 are more or less similar in terms of both intencity values and minimum and maximum positions, while fragment 2 is different. \n\nMy physical intuition behind this is following: papirus seems not being aligned by depth inside 3D volume, moreover papirus thinknes seems to vary between fragments 1, 3 and fragment 2. \n\nSo, I've decided to align and stretch / compress papirus in z dimension by shift and scale operation:\n\n1. Split the volume on overlapping patches in H and W dimentions\n2. For each patch \n- calculate mean intensity of each layer\n- fit `z_shift` parameter to minimize following loss (`y_target` here is the intensity curve of full fragment 3)\n```\nz_shifted = z + z_shift\nf = interpolate.interp1d(z_shifted, y, bounds_error=False, fill_value='extrapolate')\nreturn f(z_target) - y_target\n```\n3. For full volume fit single `z_scale` (assuming pairus thickness is same for same fragment) in a similar manner.\n4. Linearly upsample resulting `z_shift` and `z_scale` maps to full volume size and save them on disk.\n5. Apply `z_shift` and `z_scale` to full volume via `scipy.ndimage.geometric_transform` with transform writen as C module for speed.\n6. For training save it on disk, for inference apply on the fly.\n\nResulting `z_shift` maps and `z_scale` values are following:\n|Fragment 1|Fragment 2|Fragment 3|\n| -------- | ------ | ---------- |\n|![Z shift map and its histogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F387734e4ccdf846560e4f61421ffb52c%2F1_z_shift.png?generation=1686854266142975&alt=media \"Z shift map and its histogram\")|![Z scale map and its histogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F4a66d91e17b7a2332b5574defb0d1b37%2F2_z_shift.png?generation=1686854298329163&alt=media)|![Z scale map and its histogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3510608c738e619389b036e2d1d871ef%2F3_z_shift.png?generation=1686854366312885&alt=media)|\n\nZ scale is ~ 1.0 for fragments 1 and 3 and ~ 0.6 for fragment 2 which could indicate the fragment 2 being ~ twice thicker that others.\n\nSuch operation should be beneficial for 2D approaches. It seems to improve CV on ~0.03-0.05 F05, but I did not run full comparison due to time constraints.\n\nAlso, I've tried to fit / calculate normalization for each fragment separately to completely fit the intencity curve, but it did not work well.\n\n## Model\n\nPre-trained maxvit encoder + SMP 2D unet with aggregation has shown the best CV score, so `maxvit_rmlp_base_rw_384.sw_in12k_ft_in1k` is the one used in best solution. \n\n2D unet with aggregation works as follows: unet model is applied on each slice of the input with size = step = 3 along z dimention (total 6 slices), then 6 resulting logits are aggregated via linear layer.\n\nUsage of pre-trained models is yielding far better results that random initialization. Unfortunately I did not came up with better architecture allowing to handle 3D inputs and 2D outputs while allowing to use pre-trained models.\n\nBest CV scores of ths model and corresponding predictions are following:\n\n| Fragment | F05    | Prediction |\n| -------- | ------ | ---------- |\n|1         | 0.6544 |![Fragment 1 probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fc9e6f42f6cf07b913a23b6fd989076d0%2F1_3jzUWpdZJR.png?generation=1686854422630598&alt=media)|\n|2a        | 0.6643 |![Fragment 2a probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Feeef4afaf95c79def86b96014d5a7b5b%2F2a_jP08265cbJ.png?generation=1686854447976196&alt=media)|\n|2b        | 0.7489 |![Fragment 2b probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fe70babdbeb33a34a0976195453bf995d%2F2b_whqg8ITAI7.png?generation=1686854469983013&alt=media)|\n|2c        | 0.6583 |![Fragment 2c probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fff2531343b00439ebdb81f310894b61b%2F2c_Aomw3d44sk.png?generation=1686854491195589&alt=media)|\n|3         | 0.7027 |![Fragment 3 probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F626278b3fa70ceb963b517c5069f1d3f%2F3_AM8P9eim4r.png?generation=1686854513349728&alt=media)|\n\n**Tried, but did not work well**\n\nMultiple backbone models:\n- convnext v2\n- swin transformer v2\n- caformer\n- eve02\n- maxvit \n\nAnd multiple full model approaches: full 3D with aggregation on the last layer, 2.5D, 2D with aggregation, custom 3D models with pre-trained weights conversion from 2D to 3D.\n\n## Transforms\n\nFollowing augmentations are used for training\n```\ntrain_transform = A.Compose(\n    [\n        RandomCropVolumeInside2dMask(\n            base_size=self.hparams.img_size, \n            base_depth=self.hparams.img_size_z,\n            scale=(0.5, 2.0),\n            ratio=(0.9, 1.1),\n            scale_z=(1 / self.hparams.z_scale_limit, self.hparams.z_scale_limit),\n            always_apply=True,\n            crop_mask_index=0,\n        ),\n        A.Rotate(\n            p=0.5, \n            limit=rotate_limit_degrees_xy, \n            crop_border=False,\n        ),\n        ResizeVolume(\n            height=self.hparams.img_size, \n            width=self.hparams.img_size,\n            depth=self.hparams.img_size_z,\n            always_apply=True,\n        ),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.RandomRotate90(p=0.5),\n        A.RandomBrightnessContrast(p=0.5, brightness_limit=0.1, contrast_limit=0.1),\n        A.OneOf(\n            [\n                A.GaussNoise(var_limit=[10, 50]),\n                A.GaussianBlur(),\n                A.MotionBlur(),\n            ], \n            p=0.4\n        ),\n        A.GridDistortion(num_steps=5, distort_limit=0.3, p=0.5),\n        A.CoarseDropout(\n            max_holes=1, \n            max_width=int(self.hparams.img_size * 0.3), \n            max_height=int(self.hparams.img_size * 0.3), \n            mask_fill_value=0, p=0.5\n        ),\n        A.Normalize(\n            max_pixel_value=MAX_PIXEL_VALUE,\n            mean=self.train_volume_mean,\n            std=self.train_volume_std,\n            always_apply=True,\n        ),\n        ToTensorV2(),\n        ToCHWD(always_apply=True),\n    ],\n)\n```\n\nHere, `RandomCropVolumeInside2dMask` crops random 3D patch from volume so that its center's (x, y) are inside mask, `ResizeVolume` is trilinear 3D resize and `ToCHWD` simply permutes dimentions. `img_size` = 384, `img_size_z` = 18, `z_scale_limit` =  1.33, `rotate_limit_degrees_xy` = 45, `train_volume_mean` and `train_volume_std` are average of imagenet stats.\n\nFor inference, following transforms are used\n\n```\ntest_transform = A.Compose(\n    [\n        CenterCropVolume(\n            height=None, \n            width=None,\n            depth=math.ceil(self.hparams.img_size_z * self.hparams.z_scale_limit),\n            strict=True,\n            always_apply=True,\n            p=1.0,\n        ),\n        ResizeVolume(\n            height=self.hparams.img_size, \n            width=self.hparams.img_size,\n            depth=self.hparams.img_size_z,\n            always_apply=True,\n        ),\n        A.Normalize(\n            max_pixel_value=MAX_PIXEL_VALUE,\n            mean=self.train_volume_mean,\n            std=self.train_volume_std,\n            always_apply=True,\n        ),\n        ToTensorV2(),\n        ToCHWD(always_apply=True),\n    ],\n)\n```\n\nHere, `CenterCropVolume` is center crop of 3D volume.\n\nMixing approaches (cutmix, mixup, custom stuff like copy-paste the positive class voxels) also were tried but shown contradicting results, so were not included.\n\nCartesian product of `no flip / H flip / V flip` and `no 90 rotation / 90 rotation / 180 rotation / 270 rotation` is used for TTA yielding total 12 predictions.\n\n**Update**: project source code for training could be found [here](https://github.com/mkotyushev/scrolls)",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2304186": "Hello everyone! Here is a write-up of my best private leaderboard solution (top 13 with 0.641814 score, 0.723345 on public LB). \n\nI am very grateful to organizers and participants for this competition, it was very interesting and challenging. I've learned a lot about segmentation and had a lot of fun! Some of the ideas I have used in the solution came from code / discussion tabs on Kaggle, hopefully I have mentioned them all.\n\nIf you have any questions, please do not hesitate to reach me in comments.\n\n# Overview\n\nData:\n- Z shift and scale pre-processing based on mean layer intensity curve fitting to fragment 3's mean intensity curve (inspired by [this](https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis) notebook)\n- 5 fold training (1, 2a, 2b, 2c, 3 fragments, 2nd fragment is split by scroll mask area: top part + 2 bottom parts)\n- weighted sampling from fragments according to scroll mask area\n- patch size is 384, on test overlap is 192 via spline-weighted averaging (inspired by [this](https://github.com/bnsreenu/python_for_microscopists/blob/master/229_smooth_predictions_by_blending_patches/smooth_tiled_predictions.py) approach)\n- 24 slices with indices in [20, 44), random crop of 18 of them, each 3 slices window is fed to model and 6 resulting logits aggregated via linear layer\n\nTraining:\n- model is SMP Unet with 2D pre-trained encoder + linear layer aggregation\n- random crop inside scroll mask + standard vision augmentations\n- BCE loss with 0.5 weight of positive examples, selection by F0.5 score on validation fold\n- 64 epoch training, epoch length is ~ number of patches in train set (it is not constant because of random cropping)\n- LR is 1e-4, no layers freezing / LR decay, schedule is cosine annealing with warmup (1e-1, 1, 1e-2 factors, 10% warmup), optimizer is AdamW\n\nInference:\n- TTA: cartesian product of 4 rotations (0, 90, 180, 270) and 3 flips (no flip, horizontal, vertical) -> 12 predictions per volume (brought this idea from [here](https://www.kaggle.com/code/yoyobar/2-5d-segmentaion-model-with-rotate-tta))\n- TTA probabilities are averaged, writen to image file, then 5 models images averaged\n\nEnvironment & tools:\n- docker, pytorch lightning CLI, wandb\n- 1 x 3090, 1 x 3080Ti mobile\n\n# Some details\n## Data\n\nInspired by [this](https://www.kaggle.com/code/ajland/eda-a-slice-by-slice-analysis) analysis, I plot the mean intensity of each layer of each fragment for full image, only for ink and only for backgound.\n\n![Mean intensity of each layer of each fragment](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F011cb02384ff04722ab4cdd093d4a888%2Fintencity_curve.png?generation=1686855075215824&alt=media)\n\nHere it could be observed that fragments 1 and 3 are more or less similar in terms of both intencity values and minimum and maximum positions, while fragment 2 is different. \n\nMy physical intuition behind this is following: papirus seems not being aligned by depth inside 3D volume, moreover papirus thinknes seems to vary between fragments 1, 3 and fragment 2. \n\nSo, I've decided to align and stretch / compress papirus in z dimension by shift and scale operation:\n\n1. Split the volume on overlapping patches in H and W dimentions\n2. For each patch \n- calculate mean intensity of each layer\n- fit `z_shift` parameter to minimize following loss (`y_target` here is the intensity curve of full fragment 3)\n```\nz_shifted = z + z_shift\nf = interpolate.interp1d(z_shifted, y, bounds_error=False, fill_value='extrapolate')\nreturn f(z_target) - y_target\n```\n3. For full volume fit single `z_scale` (assuming pairus thickness is same for same fragment) in a similar manner.\n4. Linearly upsample resulting `z_shift` and `z_scale` maps to full volume size and save them on disk.\n5. Apply `z_shift` and `z_scale` to full volume via `scipy.ndimage.geometric_transform` with transform writen as C module for speed.\n6. For training save it on disk, for inference apply on the fly.\n\nResulting `z_shift` maps and `z_scale` values are following:\n|Fragment 1|Fragment 2|Fragment 3|\n| -------- | ------ | ---------- |\n|![Z shift map and its histogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F387734e4ccdf846560e4f61421ffb52c%2F1_z_shift.png?generation=1686854266142975&alt=media \"Z shift map and its histogram\")|![Z scale map and its histogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F4a66d91e17b7a2332b5574defb0d1b37%2F2_z_shift.png?generation=1686854298329163&alt=media)|![Z scale map and its histogram](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3510608c738e619389b036e2d1d871ef%2F3_z_shift.png?generation=1686854366312885&alt=media)|\n\nZ scale is ~ 1.0 for fragments 1 and 3 and ~ 0.6 for fragment 2 which could indicate the fragment 2 being ~ twice thicker that others.\n\nSuch operation should be beneficial for 2D approaches. It seems to improve CV on ~0.03-0.05 F05, but I did not run full comparison due to time constraints.\n\nAlso, I've tried to fit / calculate normalization for each fragment separately to completely fit the intencity curve, but it did not work well.\n\n## Model\n\nPre-trained maxvit encoder + SMP 2D unet with aggregation has shown the best CV score, so `maxvit_rmlp_base_rw_384.sw_in12k_ft_in1k` is the one used in best solution. \n\n2D unet with aggregation works as follows: unet model is applied on each slice of the input with size = step = 3 along z dimention (total 6 slices), then 6 resulting logits are aggregated via linear layer.\n\nUsage of pre-trained models is yielding far better results that random initialization. Unfortunately I did not came up with better architecture allowing to handle 3D inputs and 2D outputs while allowing to use pre-trained models.\n\nBest CV scores of ths model and corresponding predictions are following:\n\n| Fragment | F05    | Prediction |\n| -------- | ------ | ---------- |\n|1         | 0.6544 |![Fragment 1 probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fc9e6f42f6cf07b913a23b6fd989076d0%2F1_3jzUWpdZJR.png?generation=1686854422630598&alt=media)|\n|2a        | 0.6643 |![Fragment 2a probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Feeef4afaf95c79def86b96014d5a7b5b%2F2a_jP08265cbJ.png?generation=1686854447976196&alt=media)|\n|2b        | 0.7489 |![Fragment 2b probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fe70babdbeb33a34a0976195453bf995d%2F2b_whqg8ITAI7.png?generation=1686854469983013&alt=media)|\n|2c        | 0.6583 |![Fragment 2c probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fff2531343b00439ebdb81f310894b61b%2F2c_Aomw3d44sk.png?generation=1686854491195589&alt=media)|\n|3         | 0.7027 |![Fragment 3 probabilities](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F626278b3fa70ceb963b517c5069f1d3f%2F3_AM8P9eim4r.png?generation=1686854513349728&alt=media)|\n\n**Tried, but did not work well**\n\nMultiple backbone models:\n- convnext v2\n- swin transformer v2\n- caformer\n- eve02\n- maxvit \n\nAnd multiple full model approaches: full 3D with aggregation on the last layer, 2.5D, 2D with aggregation, custom 3D models with pre-trained weights conversion from 2D to 3D.\n\n## Transforms\n\nFollowing augmentations are used for training\n```\ntrain_transform = A.Compose(\n    [\n        RandomCropVolumeInside2dMask(\n            base_size=self.hparams.img_size, \n            base_depth=self.hparams.img_size_z,\n            scale=(0.5, 2.0),\n            ratio=(0.9, 1.1),\n            scale_z=(1 / self.hparams.z_scale_limit, self.hparams.z_scale_limit),\n            always_apply=True,\n            crop_mask_index=0,\n        ),\n        A.Rotate(\n            p=0.5, \n            limit=rotate_limit_degrees_xy, \n            crop_border=False,\n        ),\n        ResizeVolume(\n            height=self.hparams.img_size, \n            width=self.hparams.img_size,\n            depth=self.hparams.img_size_z,\n            always_apply=True,\n        ),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.RandomRotate90(p=0.5),\n        A.RandomBrightnessContrast(p=0.5, brightness_limit=0.1, contrast_limit=0.1),\n        A.OneOf(\n            [\n                A.GaussNoise(var_limit=[10, 50]),\n                A.GaussianBlur(),\n                A.MotionBlur(),\n            ], \n            p=0.4\n        ),\n        A.GridDistortion(num_steps=5, distort_limit=0.3, p=0.5),\n        A.CoarseDropout(\n            max_holes=1, \n            max_width=int(self.hparams.img_size * 0.3), \n            max_height=int(self.hparams.img_size * 0.3), \n            mask_fill_value=0, p=0.5\n        ),\n        A.Normalize(\n            max_pixel_value=MAX_PIXEL_VALUE,\n            mean=self.train_volume_mean,\n            std=self.train_volume_std,\n            always_apply=True,\n        ),\n        ToTensorV2(),\n        ToCHWD(always_apply=True),\n    ],\n)\n```\n\nHere, `RandomCropVolumeInside2dMask` crops random 3D patch from volume so that its center's (x, y) are inside mask, `ResizeVolume` is trilinear 3D resize and `ToCHWD` simply permutes dimentions. `img_size` = 384, `img_size_z` = 18, `z_scale_limit` =  1.33, `rotate_limit_degrees_xy` = 45, `train_volume_mean` and `train_volume_std` are average of imagenet stats.\n\nFor inference, following transforms are used\n\n```\ntest_transform = A.Compose(\n    [\n        CenterCropVolume(\n            height=None, \n            width=None,\n            depth=math.ceil(self.hparams.img_size_z * self.hparams.z_scale_limit),\n            strict=True,\n            always_apply=True,\n            p=1.0,\n        ),\n        ResizeVolume(\n            height=self.hparams.img_size, \n            width=self.hparams.img_size,\n            depth=self.hparams.img_size_z,\n            always_apply=True,\n        ),\n        A.Normalize(\n            max_pixel_value=MAX_PIXEL_VALUE,\n            mean=self.train_volume_mean,\n            std=self.train_volume_std,\n            always_apply=True,\n        ),\n        ToTensorV2(),\n        ToCHWD(always_apply=True),\n    ],\n)\n```\n\nHere, `CenterCropVolume` is center crop of 3D volume.\n\nMixing approaches (cutmix, mixup, custom stuff like copy-paste the positive class voxels) also were tried but shown contradicting results, so were not included.\n\nCartesian product of `no flip / H flip / V flip` and `no 90 rotation / 90 rotation / 180 rotation / 270 rotation` is used for TTA yielding total 12 predictions.\n\n**Update**: project source code for training could be found [here](https://github.com/mkotyushev/scrolls)"
  }
}