{
  "id": 417496,
  "title": "1st place solution",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/417496",
  "author_name": "ryches",
  "post_date": "2023-06-16T00:27:07.284000",
  "votes": 112,
  "comment_count": 46,
  "views": 0,
  "content": "<p>It was a very nervy ending to this competition, we held first for a long time but I am always scared of a shake-up. We tried to play it very defensively to ensure that we weren't doing anything to over-optimize for the leaderboard. Still, there is always a certain level of uncertainty when the public leaderboard is so small and your local validation is not extremely stable.</p>\n<p><strong>tl;dr</strong><br>\nAt a high level, I would attribute most of our success to:</p>\n<ul>\n<li>Larger crops </li>\n<li>Strong depth-invariant models</li>\n<li>Averaging several models to give us better calibration</li>\n<li>Training against all available data after validating against fragment 1 rigorously</li>\n</ul>\n<p><strong>Data prep</strong><br>\nWe initially started with the <a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\" target=\"_blank\">2.5d starter code</a> and evolved it over time. What we ended up settling on was just taking the middle 16 layers, but we tried plenty of other things that didn't work. I will save that for another post. Initially, the code was configured to train against all of the crops of the image, but we filtered this very simply based on if a crop was blank. This greatly reduced our training time because many of the crops had no data in them at all. </p>\n<p>Early on there was very little progress being made and then we ran simple ablations to see how performance was impacted by the crop size and found that there was good scaling potential there. Seemed like one of the clearest patterns we saw early on. 128 got beat by 512 which got beaten by 1024. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fadddcf67a8c6ed4f68716e050215eb2d%2FWB%20Chart%206_15_2023%203_13_17%20PM.png?generation=1686867276107142&amp;alt=media\" alt=\"\"></p>\n<p>It seemed that the x,y context was very important here because the crop could look at a whole letter or even a couple letters at a time and draw them out more completely. Smaller crops yielded much patchier-looking outputs. One nice thing about this approach was it actually didn't increase training time at all. if you ran it with 1024x1024 crops you had to train against way less crops than if you did 128x128 crops so the epoch time and wall-clock convergence time was virtually identical, unlike some other competitions where increasing the resize makes the computation take longer. </p>\n<p><strong>Models</strong><br>\nWe started with the 2.5D approach but it became apparent that it wasn't optimal because the model was learning which layers had ink in them and we knew that this varied between fragments. We wanted a solution that was depth-invariant, that could detect ink in any layer. One approach would be training a simple 2d model on each slice of the 3d volume, but with the way the data was labeled, only 2d to start with, we would be giving bad signal to the model if we told it ink was in layers that it actually wasnt and vice versa. We had many discussions about how we handle this, 1D convolutions, max pooling, 3d convolutions with size (8, 1, 1) so it was only truly looking across the depth patterns. What we found to be the most performant was using strong 3d models that would output a new 3d volume with many channels and then we could flatten them along the depth axis with a max. </p>\n<p>So for example, our first approach in this vein was a simple 3dcnn. input shape was (batch_size, 1, 16, 1024, 1024). We applied 4 layers of 3d convolutions on top of this with progressively more filters until finally, we had an output volume of (batch_size, 16, 16, 1024, 1024). At this point we just took the max across the z-dimension to squish it down to (batch_size, 16, 1024, 1024) so our 16 depth dimensions were now replaced with 16 feature dimensions. This alone was a decent approach but then passing this through a strong 2d segmentation model made it much better. We heavily relied on segformer for this, as others have mentioned b3 backbone worked great, but we also found that directionally the even bigger models performed better on the leaderboard even if they didnt perform measurably better on local validation. </p>\n<p>We iterated on this design because it seemed to satisfy the qualities we wanted, a strong 2d segmenter applied to a 3d volume that was invariant to depth. We tried different first-stage 3d methods and found that even stronger 3d models yielded better results. Evolving to 3d unet's and then eventually 3d unetr. Sometimes the results were unclear if they were better but we found a clear trend that the on-paper better models also did better on the leaderboard. Our best individual model was a UNETR first stage that passed 32 channels into a b5 segformer which scored 0.82 on the public leaderboard and 0.67 on the private. We applied a small amount of dropout on the channels in between the 3d and 2d stage. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fb59a1b2b524ffdbea3882fd6883e87ad%2FScreen%20Shot%202023-06-15%20at%204.12.57%20PM.png?generation=1686870810392089&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd3566ee53b477abe6e4026db29d35531%2FScreen%20Shot%202023-06-15%20at%204.03.32%20PM.png?generation=1686870239615259&amp;alt=media\" alt=\"\"></p>\n<p>One thing I always wondered about was if the strength of the segformer wasnt actually that it was the best model but that it was making predictions at a lower resolution. We pass it the 1024x1024 and it returns back a 256x256 segmentation map. We simply upscaled it with a very simple conv2d transpose. I believe the second place solution verifies this that the lower resolution actually helps because it is doing much coarser classification which makes it easier than trying to more precisely get every pixel. </p>\n<p>Our final best solution was 9 different models</p>\n<table>\n<thead>\n<tr>\n<th>1st stage</th>\n<th>2nd stage</th>\n<th>Resolution</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3d unet(16 channels)</td>\n<td>segformer b3</td>\n<td>1024</td>\n<td>.78</td>\n</tr>\n<tr>\n<td>3d unet(16 channels)</td>\n<td>segformer b3</td>\n<td>512</td>\n<td>.77</td>\n</tr>\n<tr>\n<td>3d unet(16 channels with SWA)</td>\n<td>segformer b3</td>\n<td>512</td>\n<td>.78</td>\n</tr>\n<tr>\n<td>3d cnn(32 channels)</td>\n<td>segformer b3</td>\n<td>1024</td>\n<td>.78</td>\n</tr>\n<tr>\n<td>3d cnn(32 channels)</td>\n<td>segformer b5</td>\n<td>1024</td>\n<td>.77</td>\n</tr>\n<tr>\n<td>3d cnn(64 channels)</td>\n<td>segformer b3</td>\n<td>1024</td>\n<td>.78</td>\n</tr>\n<tr>\n<td>3d unet(32 channels)</td>\n<td>segformer b5</td>\n<td>1024</td>\n<td>.79</td>\n</tr>\n<tr>\n<td>3d unetr(32 channels)</td>\n<td>segformer b5</td>\n<td>512</td>\n<td>.82</td>\n</tr>\n<tr>\n<td>3d unetr multiclass(32 channels)</td>\n<td>segformer b5</td>\n<td>512</td>\n<td>?</td>\n</tr>\n</tbody>\n</table>\n<p>One thing we tried late-on was multi-class output. Instead of binary we tried to predict nothing-mask-ink. This ended up yielding a much cleaner looking output actually, in most of our models we had some amount of noise anywhere there was papyrus and this seemed to quiet that down a lot. We did not get to explore it thoroughly enough to confirm if this really worked or not though. </p>\n<p><strong>Training procedure</strong><br>\nOur training procedure was fairly standard, we used the existing adamW optimizer, dice+bce loss and hyperparameters, fixing some small bugs in setting the min-lr and continuing to use the gradual warmup learning rate scheduling. We added on stochastic weight averaging to get wider optima instead of needing to pick a specific checkpoint because we found these pretty inconsistent. We found that validating against fragment 1 wasnt perfect but was at least directionally useful. Once we found something that worked on local validation against fold 1 we would submit that to the leaderboard to double-check its validity and we would continue training with all folds for several epochs after. We would submit and evaluate the new checkpoint trained against all folds and it was typically about 0.04 better. Sometimes we would train against all fragments from the start instead of training fragments 2, 3 and latter adding in 1, but we did not thoroughly evaluate this. </p>\n<p><strong>Augmentations</strong><br>\nWe tried many permutations of augmentation, mostly from albumentations, some custom, and some from 3d packages, but ultimately couldn't find strong alpha there with anything fancy. Our best model was trained with:</p>\n<ul>\n<li>50% horizontal and vertical flips</li>\n<li>75% 90-degree rotations</li>\n<li>50% brightness contrast</li>\n<li>25% 1-2 channel dropout(in this case our channels was actually depth)</li>\n<li>10% shift scale rotate</li>\n<li>10% noise and blur</li>\n<li>10% coarse dropout</li>\n<li>10% grid distortion</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F44380bdf9ac40366cfd12f6db39f1bef%2FScreen%20Shot%202023-06-15%20at%204.29.24%20PM.png?generation=1686871786148868&amp;alt=media\" alt=\"\"></p>\n<p>Overall seemed like the rotations and flips were crucial and everything else was a non-factor. Rotations seemed important but we did not ascertain that the test set itself was actually rotated, we just knew it was important so tried to make the model as invariant as possible. </p>\n<p><strong>Ensembling</strong><br>\nWe ended up with a big pile of model checkpoints and had to whittle them down to what we believed to be the most performant trading off runtime vs throughput. We tried many different combinations, heavier tta with all rotations and flips, more models, smaller strides. The winner seemed to be 4x rotation TTA with 1/4 crop strided windows and as many good models as we could fit in. Halving the stride helped an extra .01, but that was trumped by being able to add way more models, flips didnt seem to add anything at all. It's possible with just the corrected orientation instead of TTA we could have done better but a model invariant to rotations seemed just as strong.</p>\n<p>One thing we struggled with a lot was how to best combine predictions. For each model we predicted on the same pixel 4x because of our strided approach and 4x of that because of TTA. With many models this actually gave us a ton of options for aggregating the predictions per pixel. We tinkered with a lot of stuff locally but what seemed to work best was averaging per pixel for each model all ~16x predictions and then applying the sigmoid to that averaged signal and then average those probabilities together. This posed a tough memory constraint on us on kaggles system so we had to be a bit efficient with putting things away and accumulating them as densely as possible instead of just creating large arrays and averaging at the end. </p>\n<p><strong>Thresholding</strong><br>\nOne thing we went back and forth on for a long time was the calibration of predictions. As discussed in other posts, deciding a threshold was critical to getting good results and choosing the wrong threshold could give you very misleading signal on your models performance. During training and evaluation we would constantly be monitoring the AUC, Precision, Recall and f0.5 at a sweep of thresholds. Some models were well-calibrated with an optimal threshold at 0.5, but many were not. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffee3f2b5c722068974d291af29fc7049%2FScreen%20Shot%202023-06-15%20at%204.58.40%20PM.png?generation=1686873556582356&amp;alt=media\" alt=\"\"></p>\n<p>Because of this, we knew that even a great model could show up as terrible on the leaderboard if its threshold wasn't set correctly. We considered using the percentile method that some people used but did not even make a submission for it because it seemed too risky if the distribution of ink was not what we expected. Ultimately what we relied on was that averaged out our predictions they would end up calibrated. We found this to be true on our local validation and held true for the leaderboard as well. Individual models would have wide optimal threshold ranges but after averaging many predictions from many models it was almost universally centered on 0.5. We actually took our last submission to be brave and try 0.55 but it did worse on the public leaderboard and finished with only 30 minutes left to spare so we didnt pick it. It did end up performing slightly higher on the private leaderboard though. I think we actually left a lot on the table in terms of ensembling and optimal thresholding. </p>\n<p><strong>Postprocessing</strong><br>\nOne little magic function we borrowed from the old cloud segmentation competition was used to clean up our predictions. After everything was done and already binarized we would use cv2.connectedcomponents to find the masses and remove them if they were beyond a certain size. This would clean up anything that was too small, just little speckles and noise. We found it locally optimal to set our threshold a little lower and then clean up the extras, anything under 25000 pixels, but we only cleaned up things under 10k on our submissions and didnt try more aggressive cleaning. </p>\n<p>Before</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F25116a2dfc769b5784abb12704df8401%2F__results___27_30.png?generation=1686874373183834&amp;alt=media\" alt=\"\"></p>\n<p>After<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F9a6ac60974f49a6ecfbb80f47fcccbb0%2F__results___27_31.png?generation=1686874392858212&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2304372,
      "postDate": "2023-06-16T00:27:07.283Z",
      "content": "<p>It was a very nervy ending to this competition, we held first for a long time but I am always scared of a shake-up. We tried to play it very defensively to ensure that we weren't doing anything to over-optimize for the leaderboard. Still, there is always a certain level of uncertainty when the public leaderboard is so small and your local validation is not extremely stable.</p>\n<p><strong>tl;dr</strong><br>\nAt a high level, I would attribute most of our success to:</p>\n<ul>\n<li>Larger crops </li>\n<li>Strong depth-invariant models</li>\n<li>Averaging several models to give us better calibration</li>\n<li>Training against all available data after validating against fragment 1 rigorously</li>\n</ul>\n<p><strong>Data prep</strong><br>\nWe initially started with the <a href=\"https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training\" target=\"_blank\">2.5d starter code</a> and evolved it over time. What we ended up settling on was just taking the middle 16 layers, but we tried plenty of other things that didn't work. I will save that for another post. Initially, the code was configured to train against all of the crops of the image, but we filtered this very simply based on if a crop was blank. This greatly reduced our training time because many of the crops had no data in them at all. </p>\n<p>Early on there was very little progress being made and then we ran simple ablations to see how performance was impacted by the crop size and found that there was good scaling potential there. Seemed like one of the clearest patterns we saw early on. 128 got beat by 512 which got beaten by 1024. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fadddcf67a8c6ed4f68716e050215eb2d%2FWB%20Chart%206_15_2023%203_13_17%20PM.png?generation=1686867276107142&amp;alt=media\" alt=\"\"></p>\n<p>It seemed that the x,y context was very important here because the crop could look at a whole letter or even a couple letters at a time and draw them out more completely. Smaller crops yielded much patchier-looking outputs. One nice thing about this approach was it actually didn't increase training time at all. if you ran it with 1024x1024 crops you had to train against way less crops than if you did 128x128 crops so the epoch time and wall-clock convergence time was virtually identical, unlike some other competitions where increasing the resize makes the computation take longer. </p>\n<p><strong>Models</strong><br>\nWe started with the 2.5D approach but it became apparent that it wasn't optimal because the model was learning which layers had ink in them and we knew that this varied between fragments. We wanted a solution that was depth-invariant, that could detect ink in any layer. One approach would be training a simple 2d model on each slice of the 3d volume, but with the way the data was labeled, only 2d to start with, we would be giving bad signal to the model if we told it ink was in layers that it actually wasnt and vice versa. We had many discussions about how we handle this, 1D convolutions, max pooling, 3d convolutions with size (8, 1, 1) so it was only truly looking across the depth patterns. What we found to be the most performant was using strong 3d models that would output a new 3d volume with many channels and then we could flatten them along the depth axis with a max. </p>\n<p>So for example, our first approach in this vein was a simple 3dcnn. input shape was (batch_size, 1, 16, 1024, 1024). We applied 4 layers of 3d convolutions on top of this with progressively more filters until finally, we had an output volume of (batch_size, 16, 16, 1024, 1024). At this point we just took the max across the z-dimension to squish it down to (batch_size, 16, 1024, 1024) so our 16 depth dimensions were now replaced with 16 feature dimensions. This alone was a decent approach but then passing this through a strong 2d segmentation model made it much better. We heavily relied on segformer for this, as others have mentioned b3 backbone worked great, but we also found that directionally the even bigger models performed better on the leaderboard even if they didnt perform measurably better on local validation. </p>\n<p>We iterated on this design because it seemed to satisfy the qualities we wanted, a strong 2d segmenter applied to a 3d volume that was invariant to depth. We tried different first-stage 3d methods and found that even stronger 3d models yielded better results. Evolving to 3d unet's and then eventually 3d unetr. Sometimes the results were unclear if they were better but we found a clear trend that the on-paper better models also did better on the leaderboard. Our best individual model was a UNETR first stage that passed 32 channels into a b5 segformer which scored 0.82 on the public leaderboard and 0.67 on the private. We applied a small amount of dropout on the channels in between the 3d and 2d stage. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fb59a1b2b524ffdbea3882fd6883e87ad%2FScreen%20Shot%202023-06-15%20at%204.12.57%20PM.png?generation=1686870810392089&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd3566ee53b477abe6e4026db29d35531%2FScreen%20Shot%202023-06-15%20at%204.03.32%20PM.png?generation=1686870239615259&amp;alt=media\" alt=\"\"></p>\n<p>One thing I always wondered about was if the strength of the segformer wasnt actually that it was the best model but that it was making predictions at a lower resolution. We pass it the 1024x1024 and it returns back a 256x256 segmentation map. We simply upscaled it with a very simple conv2d transpose. I believe the second place solution verifies this that the lower resolution actually helps because it is doing much coarser classification which makes it easier than trying to more precisely get every pixel. </p>\n<p>Our final best solution was 9 different models</p>\n<table>\n<thead>\n<tr>\n<th>1st stage</th>\n<th>2nd stage</th>\n<th>Resolution</th>\n<th>Public Score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>3d unet(16 channels)</td>\n<td>segformer b3</td>\n<td>1024</td>\n<td>.78</td>\n</tr>\n<tr>\n<td>3d unet(16 channels)</td>\n<td>segformer b3</td>\n<td>512</td>\n<td>.77</td>\n</tr>\n<tr>\n<td>3d unet(16 channels with SWA)</td>\n<td>segformer b3</td>\n<td>512</td>\n<td>.78</td>\n</tr>\n<tr>\n<td>3d cnn(32 channels)</td>\n<td>segformer b3</td>\n<td>1024</td>\n<td>.78</td>\n</tr>\n<tr>\n<td>3d cnn(32 channels)</td>\n<td>segformer b5</td>\n<td>1024</td>\n<td>.77</td>\n</tr>\n<tr>\n<td>3d cnn(64 channels)</td>\n<td>segformer b3</td>\n<td>1024</td>\n<td>.78</td>\n</tr>\n<tr>\n<td>3d unet(32 channels)</td>\n<td>segformer b5</td>\n<td>1024</td>\n<td>.79</td>\n</tr>\n<tr>\n<td>3d unetr(32 channels)</td>\n<td>segformer b5</td>\n<td>512</td>\n<td>.82</td>\n</tr>\n<tr>\n<td>3d unetr multiclass(32 channels)</td>\n<td>segformer b5</td>\n<td>512</td>\n<td>?</td>\n</tr>\n</tbody>\n</table>\n<p>One thing we tried late-on was multi-class output. Instead of binary we tried to predict nothing-mask-ink. This ended up yielding a much cleaner looking output actually, in most of our models we had some amount of noise anywhere there was papyrus and this seemed to quiet that down a lot. We did not get to explore it thoroughly enough to confirm if this really worked or not though. </p>\n<p><strong>Training procedure</strong><br>\nOur training procedure was fairly standard, we used the existing adamW optimizer, dice+bce loss and hyperparameters, fixing some small bugs in setting the min-lr and continuing to use the gradual warmup learning rate scheduling. We added on stochastic weight averaging to get wider optima instead of needing to pick a specific checkpoint because we found these pretty inconsistent. We found that validating against fragment 1 wasnt perfect but was at least directionally useful. Once we found something that worked on local validation against fold 1 we would submit that to the leaderboard to double-check its validity and we would continue training with all folds for several epochs after. We would submit and evaluate the new checkpoint trained against all folds and it was typically about 0.04 better. Sometimes we would train against all fragments from the start instead of training fragments 2, 3 and latter adding in 1, but we did not thoroughly evaluate this. </p>\n<p><strong>Augmentations</strong><br>\nWe tried many permutations of augmentation, mostly from albumentations, some custom, and some from 3d packages, but ultimately couldn't find strong alpha there with anything fancy. Our best model was trained with:</p>\n<ul>\n<li>50% horizontal and vertical flips</li>\n<li>75% 90-degree rotations</li>\n<li>50% brightness contrast</li>\n<li>25% 1-2 channel dropout(in this case our channels was actually depth)</li>\n<li>10% shift scale rotate</li>\n<li>10% noise and blur</li>\n<li>10% coarse dropout</li>\n<li>10% grid distortion</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F44380bdf9ac40366cfd12f6db39f1bef%2FScreen%20Shot%202023-06-15%20at%204.29.24%20PM.png?generation=1686871786148868&amp;alt=media\" alt=\"\"></p>\n<p>Overall seemed like the rotations and flips were crucial and everything else was a non-factor. Rotations seemed important but we did not ascertain that the test set itself was actually rotated, we just knew it was important so tried to make the model as invariant as possible. </p>\n<p><strong>Ensembling</strong><br>\nWe ended up with a big pile of model checkpoints and had to whittle them down to what we believed to be the most performant trading off runtime vs throughput. We tried many different combinations, heavier tta with all rotations and flips, more models, smaller strides. The winner seemed to be 4x rotation TTA with 1/4 crop strided windows and as many good models as we could fit in. Halving the stride helped an extra .01, but that was trumped by being able to add way more models, flips didnt seem to add anything at all. It's possible with just the corrected orientation instead of TTA we could have done better but a model invariant to rotations seemed just as strong.</p>\n<p>One thing we struggled with a lot was how to best combine predictions. For each model we predicted on the same pixel 4x because of our strided approach and 4x of that because of TTA. With many models this actually gave us a ton of options for aggregating the predictions per pixel. We tinkered with a lot of stuff locally but what seemed to work best was averaging per pixel for each model all ~16x predictions and then applying the sigmoid to that averaged signal and then average those probabilities together. This posed a tough memory constraint on us on kaggles system so we had to be a bit efficient with putting things away and accumulating them as densely as possible instead of just creating large arrays and averaging at the end. </p>\n<p><strong>Thresholding</strong><br>\nOne thing we went back and forth on for a long time was the calibration of predictions. As discussed in other posts, deciding a threshold was critical to getting good results and choosing the wrong threshold could give you very misleading signal on your models performance. During training and evaluation we would constantly be monitoring the AUC, Precision, Recall and f0.5 at a sweep of thresholds. Some models were well-calibrated with an optimal threshold at 0.5, but many were not. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffee3f2b5c722068974d291af29fc7049%2FScreen%20Shot%202023-06-15%20at%204.58.40%20PM.png?generation=1686873556582356&amp;alt=media\" alt=\"\"></p>\n<p>Because of this, we knew that even a great model could show up as terrible on the leaderboard if its threshold wasn't set correctly. We considered using the percentile method that some people used but did not even make a submission for it because it seemed too risky if the distribution of ink was not what we expected. Ultimately what we relied on was that averaged out our predictions they would end up calibrated. We found this to be true on our local validation and held true for the leaderboard as well. Individual models would have wide optimal threshold ranges but after averaging many predictions from many models it was almost universally centered on 0.5. We actually took our last submission to be brave and try 0.55 but it did worse on the public leaderboard and finished with only 30 minutes left to spare so we didnt pick it. It did end up performing slightly higher on the private leaderboard though. I think we actually left a lot on the table in terms of ensembling and optimal thresholding. </p>\n<p><strong>Postprocessing</strong><br>\nOne little magic function we borrowed from the old cloud segmentation competition was used to clean up our predictions. After everything was done and already binarized we would use cv2.connectedcomponents to find the masses and remove them if they were beyond a certain size. This would clean up anything that was too small, just little speckles and noise. We found it locally optimal to set our threshold a little lower and then clean up the extras, anything under 25000 pixels, but we only cleaned up things under 10k on our submissions and didnt try more aggressive cleaning. </p>\n<p>Before</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F25116a2dfc769b5784abb12704df8401%2F__results___27_30.png?generation=1686874373183834&amp;alt=media\" alt=\"\"></p>\n<p>After<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F9a6ac60974f49a6ecfbb80f47fcccbb0%2F__results___27_31.png?generation=1686874392858212&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "It was a very nervy ending to this competition, we held first for a long time but I am always scared of a shake-up. We tried to play it very defensively to ensure that we weren't doing anything to over-optimize for the leaderboard. Still, there is always a certain level of uncertainty when the public leaderboard is so small and your local validation is not extremely stable.\n\n**tl;dr**\nAt a high level, I would attribute most of our success to:\n- Larger crops \n- Strong depth-invariant models\n- Averaging several models to give us better calibration\n- Training against all available data after validating against fragment 1 rigorously\n\n**Data prep**\nWe initially started with the [2.5d starter code](https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training) and evolved it over time. What we ended up settling on was just taking the middle 16 layers, but we tried plenty of other things that didn't work. I will save that for another post. Initially, the code was configured to train against all of the crops of the image, but we filtered this very simply based on if a crop was blank. This greatly reduced our training time because many of the crops had no data in them at all. \n\nEarly on there was very little progress being made and then we ran simple ablations to see how performance was impacted by the crop size and found that there was good scaling potential there. Seemed like one of the clearest patterns we saw early on. 128 got beat by 512 which got beaten by 1024. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fadddcf67a8c6ed4f68716e050215eb2d%2FWB%20Chart%206_15_2023%203_13_17%20PM.png?generation=1686867276107142&alt=media)\n\nIt seemed that the x,y context was very important here because the crop could look at a whole letter or even a couple letters at a time and draw them out more completely. Smaller crops yielded much patchier-looking outputs. One nice thing about this approach was it actually didn't increase training time at all. if you ran it with 1024x1024 crops you had to train against way less crops than if you did 128x128 crops so the epoch time and wall-clock convergence time was virtually identical, unlike some other competitions where increasing the resize makes the computation take longer. \n\n**Models**\nWe started with the 2.5D approach but it became apparent that it wasn't optimal because the model was learning which layers had ink in them and we knew that this varied between fragments. We wanted a solution that was depth-invariant, that could detect ink in any layer. One approach would be training a simple 2d model on each slice of the 3d volume, but with the way the data was labeled, only 2d to start with, we would be giving bad signal to the model if we told it ink was in layers that it actually wasnt and vice versa. We had many discussions about how we handle this, 1D convolutions, max pooling, 3d convolutions with size (8, 1, 1) so it was only truly looking across the depth patterns. What we found to be the most performant was using strong 3d models that would output a new 3d volume with many channels and then we could flatten them along the depth axis with a max. \n\nSo for example, our first approach in this vein was a simple 3dcnn. input shape was (batch_size, 1, 16, 1024, 1024). We applied 4 layers of 3d convolutions on top of this with progressively more filters until finally, we had an output volume of (batch_size, 16, 16, 1024, 1024). At this point we just took the max across the z-dimension to squish it down to (batch_size, 16, 1024, 1024) so our 16 depth dimensions were now replaced with 16 feature dimensions. This alone was a decent approach but then passing this through a strong 2d segmentation model made it much better. We heavily relied on segformer for this, as others have mentioned b3 backbone worked great, but we also found that directionally the even bigger models performed better on the leaderboard even if they didnt perform measurably better on local validation. \n\nWe iterated on this design because it seemed to satisfy the qualities we wanted, a strong 2d segmenter applied to a 3d volume that was invariant to depth. We tried different first-stage 3d methods and found that even stronger 3d models yielded better results. Evolving to 3d unet's and then eventually 3d unetr. Sometimes the results were unclear if they were better but we found a clear trend that the on-paper better models also did better on the leaderboard. Our best individual model was a UNETR first stage that passed 32 channels into a b5 segformer which scored 0.82 on the public leaderboard and 0.67 on the private. We applied a small amount of dropout on the channels in between the 3d and 2d stage. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fb59a1b2b524ffdbea3882fd6883e87ad%2FScreen%20Shot%202023-06-15%20at%204.12.57%20PM.png?generation=1686870810392089&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd3566ee53b477abe6e4026db29d35531%2FScreen%20Shot%202023-06-15%20at%204.03.32%20PM.png?generation=1686870239615259&alt=media)\n\nOne thing I always wondered about was if the strength of the segformer wasnt actually that it was the best model but that it was making predictions at a lower resolution. We pass it the 1024x1024 and it returns back a 256x256 segmentation map. We simply upscaled it with a very simple conv2d transpose. I believe the second place solution verifies this that the lower resolution actually helps because it is doing much coarser classification which makes it easier than trying to more precisely get every pixel. \n\nOur final best solution was 9 different models\n\n\n| 1st stage \t| 2nd stage \t| Resolution \t| Public Score \t|\n|---\t|---\t|---\t|---\t|\n| 3d unet(16 channels) \t| segformer b3 \t| 1024 \t| .78 \t|\n| 3d unet(16 channels) \t| segformer b3 \t| 512 \t| .77 \t|\n| 3d unet(16 channels with SWA) \t| segformer b3 \t| 512 \t| .78 \t|\n| 3d cnn(32 channels) \t| segformer b3 \t| 1024 \t| .78 \t|\n| 3d cnn(32 channels) \t| segformer b5 \t| 1024 \t| .77 \t|\n| 3d cnn(64 channels) \t| segformer b3 \t| 1024 \t| .78 \t|\n| 3d unet(32 channels) \t| segformer b5 \t| 1024 \t| .79 \t|\n| 3d unetr(32 channels) \t| segformer b5 \t| 512 \t| .82 \t|\n| 3d unetr multiclass(32 channels) \t| segformer b5 \t| 512 \t| ? \t|\n\n\nOne thing we tried late-on was multi-class output. Instead of binary we tried to predict nothing-mask-ink. This ended up yielding a much cleaner looking output actually, in most of our models we had some amount of noise anywhere there was papyrus and this seemed to quiet that down a lot. We did not get to explore it thoroughly enough to confirm if this really worked or not though. \n\n\n**Training procedure**\nOur training procedure was fairly standard, we used the existing adamW optimizer, dice+bce loss and hyperparameters, fixing some small bugs in setting the min-lr and continuing to use the gradual warmup learning rate scheduling. We added on stochastic weight averaging to get wider optima instead of needing to pick a specific checkpoint because we found these pretty inconsistent. We found that validating against fragment 1 wasnt perfect but was at least directionally useful. Once we found something that worked on local validation against fold 1 we would submit that to the leaderboard to double-check its validity and we would continue training with all folds for several epochs after. We would submit and evaluate the new checkpoint trained against all folds and it was typically about 0.04 better. Sometimes we would train against all fragments from the start instead of training fragments 2, 3 and latter adding in 1, but we did not thoroughly evaluate this. \n\n**Augmentations**\nWe tried many permutations of augmentation, mostly from albumentations, some custom, and some from 3d packages, but ultimately couldn't find strong alpha there with anything fancy. Our best model was trained with:\n- 50% horizontal and vertical flips\n- 75% 90-degree rotations\n- 50% brightness contrast\n- 25% 1-2 channel dropout(in this case our channels was actually depth)\n- 10% shift scale rotate\n- 10% noise and blur\n- 10% coarse dropout\n- 10% grid distortion\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F44380bdf9ac40366cfd12f6db39f1bef%2FScreen%20Shot%202023-06-15%20at%204.29.24%20PM.png?generation=1686871786148868&alt=media)\n\nOverall seemed like the rotations and flips were crucial and everything else was a non-factor. Rotations seemed important but we did not ascertain that the test set itself was actually rotated, we just knew it was important so tried to make the model as invariant as possible. \n\n\n**Ensembling**\nWe ended up with a big pile of model checkpoints and had to whittle them down to what we believed to be the most performant trading off runtime vs throughput. We tried many different combinations, heavier tta with all rotations and flips, more models, smaller strides. The winner seemed to be 4x rotation TTA with 1/4 crop strided windows and as many good models as we could fit in. Halving the stride helped an extra .01, but that was trumped by being able to add way more models, flips didnt seem to add anything at all. It's possible with just the corrected orientation instead of TTA we could have done better but a model invariant to rotations seemed just as strong.\n\nOne thing we struggled with a lot was how to best combine predictions. For each model we predicted on the same pixel 4x because of our strided approach and 4x of that because of TTA. With many models this actually gave us a ton of options for aggregating the predictions per pixel. We tinkered with a lot of stuff locally but what seemed to work best was averaging per pixel for each model all ~16x predictions and then applying the sigmoid to that averaged signal and then average those probabilities together. This posed a tough memory constraint on us on kaggles system so we had to be a bit efficient with putting things away and accumulating them as densely as possible instead of just creating large arrays and averaging at the end. \n\n**Thresholding**\nOne thing we went back and forth on for a long time was the calibration of predictions. As discussed in other posts, deciding a threshold was critical to getting good results and choosing the wrong threshold could give you very misleading signal on your models performance. During training and evaluation we would constantly be monitoring the AUC, Precision, Recall and f0.5 at a sweep of thresholds. Some models were well-calibrated with an optimal threshold at 0.5, but many were not. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffee3f2b5c722068974d291af29fc7049%2FScreen%20Shot%202023-06-15%20at%204.58.40%20PM.png?generation=1686873556582356&alt=media)\n\nBecause of this, we knew that even a great model could show up as terrible on the leaderboard if its threshold wasn't set correctly. We considered using the percentile method that some people used but did not even make a submission for it because it seemed too risky if the distribution of ink was not what we expected. Ultimately what we relied on was that averaged out our predictions they would end up calibrated. We found this to be true on our local validation and held true for the leaderboard as well. Individual models would have wide optimal threshold ranges but after averaging many predictions from many models it was almost universally centered on 0.5. We actually took our last submission to be brave and try 0.55 but it did worse on the public leaderboard and finished with only 30 minutes left to spare so we didnt pick it. It did end up performing slightly higher on the private leaderboard though. I think we actually left a lot on the table in terms of ensembling and optimal thresholding. \n\n**Postprocessing**\nOne little magic function we borrowed from the old cloud segmentation competition was used to clean up our predictions. After everything was done and already binarized we would use cv2.connectedcomponents to find the masses and remove them if they were beyond a certain size. This would clean up anything that was too small, just little speckles and noise. We found it locally optimal to set our threshold a little lower and then clean up the extras, anything under 25000 pixels, but we only cleaned up things under 10k on our submissions and didnt try more aggressive cleaning. \n\nBefore\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F25116a2dfc769b5784abb12704df8401%2F__results___27_30.png?generation=1686874373183834&alt=media)\n\nAfter\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F9a6ac60974f49a6ecfbb80f47fcccbb0%2F__results___27_31.png?generation=1686874392858212&alt=media)\n",
      "votes": 112
    },
    {
      "id": 2304375,
      "postDate": "2023-06-16T00:32:32.733Z",
      "content": "<p>Hoping to aggregate a list of all the things we tried that didn't pan out as well but will have to add those later. One ironic thing with all of this is that I was visiting nvidia for work while the competition was closing so I got to see the news at their hq. Being there and winning a competition with one of their models dawned on me part way through the day. </p>",
      "rawMarkdown": "Hoping to aggregate a list of all the things we tried that didn't pan out as well but will have to add those later. One ironic thing with all of this is that I was visiting nvidia for work while the competition was closing so I got to see the news at their hq. Being there and winning a competition with one of their models dawned on me part way through the day. ",
      "votes": 11
    },
    {
      "id": 2304386,
      "postDate": "2023-06-16T00:51:18.253Z",
      "content": "<p>One thing we tried that didnt work was instead of filtering the crops based on if they were empty or not, filter them in case they had any ink or not. This massively reduced the number of crops and looked pretty good on paper but it taught the model that every single crop must have some ink and it didnt learn to handle blank papyrus well. Became overexcited and greatly over predicted. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F3e6b629a11dc8824e2d9b9e54a818ade%2FScreen%20Shot%202023-06-15%20at%205.50.22%20PM.png?generation=1686876640893243&amp;alt=media\" alt=\"\"></p>\n<p>maybe this would have worked with bigger crop sizes. This made it converge extremely fast but you can see the result isnt that great especially in reference to precision. </p>",
      "rawMarkdown": "One thing we tried that didnt work was instead of filtering the crops based on if they were empty or not, filter them in case they had any ink or not. This massively reduced the number of crops and looked pretty good on paper but it taught the model that every single crop must have some ink and it didnt learn to handle blank papyrus well. Became overexcited and greatly over predicted. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F3e6b629a11dc8824e2d9b9e54a818ade%2FScreen%20Shot%202023-06-15%20at%205.50.22%20PM.png?generation=1686876640893243&alt=media)\n\nmaybe this would have worked with bigger crop sizes. This made it converge extremely fast but you can see the result isnt that great especially in reference to precision. ",
      "votes": 5,
      "replies": [
        {
          "id": 2895021,
          "postDate": "2024-06-28T20:50:56.683Z",
          "content": "<p>Thanks for this explanation! </p>",
          "rawMarkdown": "Thanks for this explanation! "
        }
      ]
    },
    {
      "id": 2305432,
      "postDate": "2023-06-16T17:29:50.650Z",
      "content": "<p>Many congrats on the win, and thanks for the fantastic write-up! Lots of valuable insight in there.</p>\n<p>I like the focus on depth invariance, and the ChannelDropout. One idea I had, which I didn't get round to testing, was to use ChannelShuffle in train &amp; tta. Did you happen you try that, and would you reckon it might help with the depth invariance?</p>",
      "rawMarkdown": "Many congrats on the win, and thanks for the fantastic write-up! Lots of valuable insight in there.\n\nI like the focus on depth invariance, and the ChannelDropout. One idea I had, which I didn't get round to testing, was to use ChannelShuffle in train & tta. Did you happen you try that, and would you reckon it might help with the depth invariance?",
      "votes": 6,
      "replies": [
        {
          "id": 2305456,
          "postDate": "2023-06-16T17:52:12.493Z",
          "content": "<p>I didn't feel we had time to explore shuffling channels.  If the signal we're chasing is exclusively located within the XY plane of each slice, then this would help with regularization.  But to the extent that some of the signal lies in detecting where there is a transition from ink to no ink as you get deeper into the papyrus, then channel shuffle would have destroyed that.  It's an interesting ablation to try!</p>",
          "rawMarkdown": "I didn't feel we had time to explore shuffling channels.  If the signal we're chasing is exclusively located within the XY plane of each slice, then this would help with regularization.  But to the extent that some of the signal lies in detecting where there is a transition from ink to no ink as you get deeper into the papyrus, then channel shuffle would have destroyed that.  It's an interesting ablation to try!",
          "votes": 3
        },
        {
          "id": 2305582,
          "postDate": "2023-06-16T19:26:59.923Z",
          "content": "<p>I actually tried a channel shuffled trial. It failed to converge at all. I specifically wanted to see how a sort of bag of depths approach worked and it completely failed with our models and a couple variations on them. I'll have to see if I can find the training graph for that one. </p>",
          "rawMarkdown": "I actually tried a channel shuffled trial. It failed to converge at all. I specifically wanted to see how a sort of bag of depths approach worked and it completely failed with our models and a couple variations on them. I'll have to see if I can find the training graph for that one. ",
          "votes": 3,
          "replies": [
            {
              "id": 2305590,
              "postDate": "2023-06-16T19:31:37.607Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F642627376e3d046484bec6b82112a907%2FWB%20Chart%206_16_2023%2012_30_02%20PM.png?generation=1686943850364553&amp;alt=media\" alt=\"\"></p>\n<p>The one way down at the bottom is what happened when we shuffled the channels on the way into the model. Didnt seem to converge. We wouldve been better off just always grabbing the center channel than having them shuffled. </p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F642627376e3d046484bec6b82112a907%2FWB%20Chart%206_16_2023%2012_30_02%20PM.png?generation=1686943850364553&alt=media)\n\nThe one way down at the bottom is what happened when we shuffled the channels on the way into the model. Didnt seem to converge. We wouldve been better off just always grabbing the center channel than having them shuffled. ",
              "votes": 3
            },
            {
              "id": 2305632,
              "postDate": "2023-06-16T20:07:12.490Z",
              "content": "<p>That's very interesting, thank you! Is that with 16 or 32 channels? We did some analysis, which I'm hoping to post in our write-up soon, on the amount of information contained in each z-axis slice. From that, I'd guess that shuffling 32 channels would include a significant number of channels that just have noise, and maybe the shuffling could help for the few channels that contain the most ink signal.</p>",
              "rawMarkdown": "That's very interesting, thank you! Is that with 16 or 32 channels? We did some analysis, which I'm hoping to post in our write-up soon, on the amount of information contained in each z-axis slice. From that, I'd guess that shuffling 32 channels would include a significant number of channels that just have noise, and maybe the shuffling could help for the few channels that contain the most ink signal.",
              "votes": 1
            },
            {
              "id": 2305650,
              "postDate": "2023-06-16T20:17:03.330Z",
              "content": "<p>That was with the center 16 I believe. I also did a pretty thorough analysis of the info of ink in each layer of each fragment. I trained a single-layer model for each depth and then applied it to the held out fragment. You can see for each series there is a rise and fall of information density and at the ends there is nothing left. Just base guessing. </p>\n<p>I think we left a little performance on the table by not honing in better on which layers had ink. We tried some approachs to move the depth per fragment or per crop and couldnt find anything that did consistently better than just taking the middle 16. We couldve maybe done with middle 32 and gotten even more signal. Early on I was seeing that 32 depth beat 16, but for double the vram usage it didnt seem worth it. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Faabd32462a191a9b9e092ee7be888983%2FScreen%20Shot%202023-06-16%20at%201.14.27%20PM.png?generation=1686946620916433&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "That was with the center 16 I believe. I also did a pretty thorough analysis of the info of ink in each layer of each fragment. I trained a single-layer model for each depth and then applied it to the held out fragment. You can see for each series there is a rise and fall of information density and at the ends there is nothing left. Just base guessing. \n\nI think we left a little performance on the table by not honing in better on which layers had ink. We tried some approachs to move the depth per fragment or per crop and couldnt find anything that did consistently better than just taking the middle 16. We couldve maybe done with middle 32 and gotten even more signal. Early on I was seeing that 32 depth beat 16, but for double the vram usage it didnt seem worth it. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Faabd32462a191a9b9e092ee7be888983%2FScreen%20Shot%202023-06-16%20at%201.14.27%20PM.png?generation=1686946620916433&alt=media)",
              "votes": 2
            },
            {
              "id": 2305663,
              "postDate": "2023-06-16T20:32:22.627Z",
              "content": "<p>Nice! Very consistent with my findings. Did you ever try models for just the top 4 or 5 layers in your plot that seem to contain the most signal? I thought that I might get some additional speed / larger batch sizes out of that but the performance drop was significant.</p>\n<p>For my depth scans, I also looked at whether different layers might contain more ink in different x, y regions of a given fragment, but didn't find any indications for that. Did you see anything along those lines?</p>",
              "rawMarkdown": "Nice! Very consistent with my findings. Did you ever try models for just the top 4 or 5 layers in your plot that seem to contain the most signal? I thought that I might get some additional speed / larger batch sizes out of that but the performance drop was significant.\n\nFor my depth scans, I also looked at whether different layers might contain more ink in different x, y regions of a given fragment, but didn't find any indications for that. Did you see anything along those lines?",
              "votes": 1
            },
            {
              "id": 2305670,
              "postDate": "2023-06-16T20:39:14.847Z",
              "content": "<p>I tried varying placements and windows of depth but center 16 seemed sufficient. Couldnt find anything better. I tried to analytically find the best for each fragment individually. Taking the peak of average values per layer as the center point. I tried also doing that per crop, so for every crop I measured the average intensity per layer and then used the one with the highest value as the center and then went 8 up and down from there. </p>\n<p>Was extremely surprising to find that it did pretty comparably even though many of the crops had way different values than the middle 16, for some crops it would pick like the 15th layer as the center instead of the 32nd. Didnt pursue it too far. Maybe there was some better method for finding the regions most likely to have ink and then hone into them or broadcast our 2d label into 3d in a smart analytical way and then train against our generated 3d labels but we weren't able to come up with something that worked</p>",
              "rawMarkdown": "I tried varying placements and windows of depth but center 16 seemed sufficient. Couldnt find anything better. I tried to analytically find the best for each fragment individually. Taking the peak of average values per layer as the center point. I tried also doing that per crop, so for every crop I measured the average intensity per layer and then used the one with the highest value as the center and then went 8 up and down from there. \n\nWas extremely surprising to find that it did pretty comparably even though many of the crops had way different values than the middle 16, for some crops it would pick like the 15th layer as the center instead of the 32nd. Didnt pursue it too far. Maybe there was some better method for finding the regions most likely to have ink and then hone into them or broadcast our 2d label into 3d in a smart analytical way and then train against our generated 3d labels but we weren't able to come up with something that worked",
              "votes": 1
            },
            {
              "id": 2305672,
              "postDate": "2023-06-16T20:44:44.367Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F4ffa67c0437055c8f3256673f1a20b6c%2FWB%20Chart%206_16_2023%201_42_18%20PM.png?generation=1686948170284985&amp;alt=media\" alt=\"\"></p>\n<p>this was trying to find the max response per layer per tile first. Maybe shouldve explored it more, but didnt immediately give any value. This was with old iterations of our model. Maybe worth revisiting. </p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F4ffa67c0437055c8f3256673f1a20b6c%2FWB%20Chart%206_16_2023%201_42_18%20PM.png?generation=1686948170284985&alt=media)\n\nthis was trying to find the max response per layer per tile first. Maybe shouldve explored it more, but didnt immediately give any value. This was with old iterations of our model. Maybe worth revisiting. ",
              "votes": 1
            },
            {
              "id": 2305673,
              "postDate": "2023-06-16T20:46:19.493Z",
              "content": "<p>One of the things we tried early on was backing into the 3d labels by having our 3d unet output the 3d volume and then just take the max of the depth axis while training and then we could inspect its 3d volume without the max and see where it was finding the ink from. We could clearly see that it was picking things up from different layers</p>",
              "rawMarkdown": "One of the things we tried early on was backing into the 3d labels by having our 3d unet output the 3d volume and then just take the max of the depth axis while training and then we could inspect its 3d volume without the max and see where it was finding the ink from. We could clearly see that it was picking things up from different layers",
              "votes": 1
            },
            {
              "id": 2305685,
              "postDate": "2023-06-16T21:00:36.590Z",
              "content": "<p>That's super interesting! If you have some sensitivity maps for different layers based on your 3d volumes then that would be a fascinating insight.</p>",
              "rawMarkdown": "That's super interesting! If you have some sensitivity maps for different layers based on your 3d volumes then that would be a fascinating insight.",
              "votes": 1
            },
            {
              "id": 2305687,
              "postDate": "2023-06-16T21:04:16.807Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fe772fcafe4f29d22d7068ff6b0a57886%2F2-0_MdgQnL3V.gif?generation=1686949374516762&amp;alt=media\" alt=\"\"></p>\n<p>here is an example of our 3d models output. Sweeping through the 3d volume that it output before we crushed it with a max across the depth. You can see that it is seeing different stuff on different layers in different x,y regions. This could happen from the ink soaking in to different depths or it could be from the segmentation model used to flatten the fragments centering slightly differently. </p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fe772fcafe4f29d22d7068ff6b0a57886%2F2-0_MdgQnL3V.gif?generation=1686949374516762&alt=media)\n\nhere is an example of our 3d models output. Sweeping through the 3d volume that it output before we crushed it with a max across the depth. You can see that it is seeing different stuff on different layers in different x,y regions. This could happen from the ink soaking in to different depths or it could be from the segmentation model used to flatten the fragments centering slightly differently. ",
              "votes": 5
            },
            {
              "id": 2305691,
              "postDate": "2023-06-16T21:16:37.157Z",
              "content": "<p>Wow! This is mesmerizing. I think it would be very interesting to figure out how much of this are intrinsic differences in depth, and how robustly you could detect them at inference time. Probably less important for a depth-invariant model, but could still help to really tease out the ink from the noise on a local level.</p>",
              "rawMarkdown": "Wow! This is mesmerizing. I think it would be very interesting to figure out how much of this are intrinsic differences in depth, and how robustly you could detect them at inference time. Probably less important for a depth-invariant model, but could still help to really tease out the ink from the noise on a local level.",
              "votes": 1
            },
            {
              "id": 2305717,
              "postDate": "2023-06-16T21:38:48.173Z",
              "content": "<p>I think there is some potential for strong pseudo labeling still. First pass do the 3d volume to 2d label and then back out the 3d label from the model and train against that. Maybe the extra 3d localization helps the model be more precise and understand where things are coming from. Or maybe you just teach the second model to mimic the first. Hard to say without experiment. </p>",
              "rawMarkdown": "I think there is some potential for strong pseudo labeling still. First pass do the 3d volume to 2d label and then back out the 3d label from the model and train against that. Maybe the extra 3d localization helps the model be more precise and understand where things are coming from. Or maybe you just teach the second model to mimic the first. Hard to say without experiment. ",
              "votes": 1
            },
            {
              "id": 2305786,
              "postDate": "2023-06-16T23:01:37.077Z",
              "content": "<p>Agreed; this could produce a significant boost. This is all great stuff; thank you Ryan for all the detail! Much appreciated. I'm gonna think more about those 3d features and I'm happy to discuss further. I'll see if I can make it to your meetup event tomorrow.</p>",
              "rawMarkdown": "Agreed; this could produce a significant boost. This is all great stuff; thank you Ryan for all the detail! Much appreciated. I'm gonna think more about those 3d features and I'm happy to discuss further. I'll see if I can make it to your meetup event tomorrow.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2335550,
      "postDate": "2023-07-08T16:07:51.210Z",
      "content": "<p>Congrats on the win! And thanks for sharing. I am a beginner when it comes to data science and I am trying to reproduce the results of one of your models (to prove to myself that I can make it work end to end). I have been unable to match the CV results that you shared for your models. I am running on a A100 in Google Colab, but I have to use 6 transformer heads for the UNETR, \"nvidia/mit-b0\" for the Segformer and set precision=\"16-mixed\" for Pytorch lightning and run with a resolution of 512 in order to not run the GPU out of memory (of which it has 40 GB). I also hit NaN loss values (likely due to mixed precision) when using AdamW and therefore I have to use SGD. I am finding myself struggling a lot more with setting up a training harness that can run a big model and reproduce the results than with testing different architectures.</p>\n<p>What hardware did you train on? How long did each epoch take? How many epochs did you run for? How did you tune the hyperparams as the epochs progressed?</p>\n<p>Additionally, how did you load the images from disk and did you pre-process them to change the numeric range/normalize (or did you rely on the Normalize transformation from albumentations entirely?</p>\n<p>How did you define your folds and did you retrain the model N times assuming you had N-folds? How long did all this training take? Did you then use your best trained fold, did you average the weights from all folds, or did you train without holding back one of the folds for validation?</p>\n<p>Thanks in advance for the clarification. My goal with these questions is to closely reproduce the performance of one of your models. If you have exact code to share, I would love to learn from that.</p>",
      "rawMarkdown": "Congrats on the win! And thanks for sharing. I am a beginner when it comes to data science and I am trying to reproduce the results of one of your models (to prove to myself that I can make it work end to end). I have been unable to match the CV results that you shared for your models. I am running on a A100 in Google Colab, but I have to use 6 transformer heads for the UNETR, \"nvidia/mit-b0\" for the Segformer and set precision=\"16-mixed\" for Pytorch lightning and run with a resolution of 512 in order to not run the GPU out of memory (of which it has 40 GB). I also hit NaN loss values (likely due to mixed precision) when using AdamW and therefore I have to use SGD. I am finding myself struggling a lot more with setting up a training harness that can run a big model and reproduce the results than with testing different architectures.\n\nWhat hardware did you train on? How long did each epoch take? How many epochs did you run for? How did you tune the hyperparams as the epochs progressed?\n\nAdditionally, how did you load the images from disk and did you pre-process them to change the numeric range/normalize (or did you rely on the Normalize transformation from albumentations entirely?\n\nHow did you define your folds and did you retrain the model N times assuming you had N-folds? How long did all this training take? Did you then use your best trained fold, did you average the weights from all folds, or did you train without holding back one of the folds for validation?\n\nThanks in advance for the clarification. My goal with these questions is to closely reproduce the performance of one of your models. If you have exact code to share, I would love to learn from that.",
      "votes": 1
    },
    {
      "id": 2310560,
      "postDate": "2023-06-20T13:14:42.080Z",
      "content": "<p>Thanks for sharing your thoughts! It's always important to know what someone who achieves success thinks! Does your team plan to publish your complete solution on Github or similar platform, or not?</p>",
      "rawMarkdown": "Thanks for sharing your thoughts! It's always important to know what someone who achieves success thinks! Does your team plan to publish your complete solution on Github or similar platform, or not?",
      "votes": 1
    },
    {
      "id": 2304410,
      "postDate": "2023-06-16T01:36:10.717Z",
      "content": "<p>Congratulations for the 1st prize.</p>\n<p>I was also impressed of the thorough research.</p>\n<p>I have questions, </p>\n<ul>\n<li>I once used SegFormer for 2D encoder (only 1st stage), but it seemed to take much epochs than CNN encoder to converge. However, as shown in the performance graph, the conversion seems after &lt;25 epochs. Do you think the second stage strategy (3D-2D) helps to converge faster?</li>\n<li>I was also observed very fluctuated metrics even if closed to convergence. It makes me hard to which model is strong is very difficult. So the most of experiments I used seed-average of the models. How do you evaluate the local model's performance?</li>\n</ul>",
      "rawMarkdown": "Congratulations for the 1st prize.\n\nI was also impressed of the thorough research.\n\nI have questions, \n\n- I once used SegFormer for 2D encoder (only 1st stage), but it seemed to take much epochs than CNN encoder to converge. However, as shown in the performance graph, the conversion seems after <25 epochs. Do you think the second stage strategy (3D-2D) helps to converge faster?\n- I was also observed very fluctuated metrics even if closed to convergence. It makes me hard to which model is strong is very difficult. So the most of experiments I used seed-average of the models. How do you evaluate the local model's performance?",
      "votes": 1,
      "replies": [
        {
          "id": 2304431,
          "postDate": "2023-06-16T02:00:55Z",
          "content": "<p>generally I felt like segformer converged more quickly than other stage 2 models we tried like unet, upernet with various backbones. not sure exactly why you saw slow convergence but I never tried it on its own. </p>\n<p>I mostly just looked at the general trend of the model. If I measured performance with the swa model performance movement was very slow and smooth curves, but we saw a similar effect just using the running average of the charts on wandb</p>",
          "rawMarkdown": "generally I felt like segformer converged more quickly than other stage 2 models we tried like unet, upernet with various backbones. not sure exactly why you saw slow convergence but I never tried it on its own. \n\nI mostly just looked at the general trend of the model. If I measured performance with the swa model performance movement was very slow and smooth curves, but we saw a similar effect just using the running average of the charts on wandb",
          "votes": 3
        }
      ]
    },
    {
      "id": 2305095,
      "postDate": "2023-06-16T12:44:46.533Z",
      "content": "<p>nice writeup, congratulations for the 1st place!</p>\n<p>One question: did you train the 3dcnn and 2d segmentation models jointly or separately?</p>",
      "rawMarkdown": "nice writeup, congratulations for the 1st place!\n\nOne question: did you train the 3dcnn and 2d segmentation models jointly or separately?",
      "votes": 2,
      "replies": [
        {
          "id": 2305234,
          "postDate": "2023-06-16T14:21:01.533Z",
          "content": "<p>It was all trained end to end. My belief is training end to end and optimizing directly for your goal is almost always universally better. Sometimes you have to do things in stages just because of memory/compute constraints but in general I will always prefer end to end</p>",
          "rawMarkdown": "It was all trained end to end. My belief is training end to end and optimizing directly for your goal is almost always universally better. Sometimes you have to do things in stages just because of memory/compute constraints but in general I will always prefer end to end",
          "votes": 4,
          "replies": [
            {
              "id": 2305250,
              "postDate": "2023-06-16T14:29:41.893Z",
              "content": "<p>Great, thanks for the insight! </p>",
              "rawMarkdown": "Great, thanks for the insight! ",
              "votes": 1
            },
            {
              "id": 2305459,
              "postDate": "2023-06-16T17:55:45.930Z",
              "content": "<p>I'll also add that I was a big proponent of end-to-end training.  The downside is you typically need more data for the deeper pipeline.  We only had three fragments to train on, but at least we had many crops per fragment and all of the image augmentations.  In the end, we were able to get end-to-end to converge pretty well, so we stuck with it.</p>",
              "rawMarkdown": "I'll also add that I was a big proponent of end-to-end training.  The downside is you typically need more data for the deeper pipeline.  We only had three fragments to train on, but at least we had many crops per fragment and all of the image augmentations.  In the end, we were able to get end-to-end to converge pretty well, so we stuck with it.",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2541483,
      "postDate": "2023-11-28T14:34:01.977Z",
      "content": "<p>Congrats on winning this competition! I'd like to play around with the models and training / fine-tuning on the actual scroll data sets. Did you share the training code somewhere? Would you be open to share the training code?</p>",
      "rawMarkdown": "Congrats on winning this competition! I'd like to play around with the models and training / fine-tuning on the actual scroll data sets. Did you share the training code somewhere? Would you be open to share the training code?",
      "replies": [
        {
          "id": 2541737,
          "postDate": "2023-11-28T18:29:17.767Z",
          "content": "<p>All the code for the 10 winners is public, links are at <a href=\"https://scrollprize.org/community_projects#fragment-based-ink-detection\" target=\"_blank\">https://scrollprize.org/community_projects#fragment-based-ink-detection</a></p>",
          "rawMarkdown": "All the code for the 10 winners is public, links are at https://scrollprize.org/community_projects#fragment-based-ink-detection",
          "votes": 1,
          "replies": [
            {
              "id": 2541802,
              "postDate": "2023-11-28T19:52:24.170Z",
              "content": "<p>Ah, perfect, missed that. I looked in Ryan's Github but missed Aina's. Not the best premise for an aspiring treasure hunter…</p>",
              "rawMarkdown": "Ah, perfect, missed that. I looked in Ryan's Github but missed Aina's. Not the best premise for an aspiring treasure hunter...",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2319206,
      "postDate": "2023-06-27T01:36:34.337Z",
      "content": "<p>Where can I find the segformerForSemanticSegmentation function? I would like to see the performance of segfomer.</p>",
      "rawMarkdown": "Where can I find the segformerForSemanticSegmentation function? I would like to see the performance of segfomer.",
      "replies": [
        {
          "id": 2319232,
          "postDate": "2023-06-27T02:07:41.947Z",
          "content": "<p>It comes from the transformers package</p>",
          "rawMarkdown": "It comes from the transformers package",
          "votes": 2,
          "replies": [
            {
              "id": 2319280,
              "postDate": "2023-06-27T03:30:51.570Z",
              "content": "<p>Thank you! I find it.</p>",
              "rawMarkdown": "Thank you! I find it."
            },
            {
              "id": 2319282,
              "postDate": "2023-06-27T03:31:10.387Z",
              "content": "<p>And how many epochs did you train for segformer?</p>",
              "rawMarkdown": "And how many epochs did you train for segformer?"
            }
          ]
        },
        {
          "id": 2319279,
          "postDate": "2023-06-27T03:30:25.430Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2310125,
      "postDate": "2023-06-20T06:04:34.990Z",
      "content": "<p>Nice thoughts and ideas from you and this is helpful for learners to improve our skills and gain insights.</p>",
      "rawMarkdown": "Nice thoughts and ideas from you and this is helpful for learners to improve our skills and gain insights."
    },
    {
      "id": 2310107,
      "postDate": "2023-06-20T05:44:34.713Z",
      "content": "<p>Informative, Congratulations 🎉</p>",
      "rawMarkdown": "Informative, Congratulations 🎉"
    },
    {
      "id": 2309989,
      "postDate": "2023-06-20T03:38:11.840Z",
      "content": "<p>Amazing job!!! Congrats!</p>",
      "rawMarkdown": "Amazing job!!! Congrats!"
    },
    {
      "id": 2309944,
      "postDate": "2023-06-20T02:29:26.550Z",
      "content": "<p>so cool! congrats!</p>",
      "rawMarkdown": "so cool! congrats!"
    },
    {
      "id": 2308738,
      "postDate": "2023-06-19T07:10:38.967Z",
      "content": "<p>I want to know how you calculate the loss? The mask size output by segfomer is 1/4 of the input size. If the input size is 512 * 512, then the predicted mask size is 128 * 128. When calculating losses, should the predicted mask be upsampled to 512 * 512, or should the ground truth be downsampled to 128 * 128?</p>",
      "rawMarkdown": "I want to know how you calculate the loss? The mask size output by segfomer is 1/4 of the input size. If the input size is 512 * 512, then the predicted mask size is 128 * 128. When calculating losses, should the predicted mask be upsampled to 512 * 512, or should the ground truth be downsampled to 128 * 128?",
      "replies": [
        {
          "id": 2308790,
          "postDate": "2023-06-19T08:06:53.443Z",
          "content": "<p>Yes, we upscale the predictions back to the original resolution.  If you look in the models section above, Ryan shows in the forward pass that the output of the Segformer goes through two 2x upscaling layers.  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F909432%2Fec8e2e11a7484f580f25262e8c8ef5c3%2Fupscale.png?generation=1687161955698991&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Yes, we upscale the predictions back to the original resolution.  If you look in the models section above, Ryan shows in the forward pass that the output of the Segformer goes through two 2x upscaling layers.  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F909432%2Fec8e2e11a7484f580f25262e8c8ef5c3%2Fupscale.png?generation=1687161955698991&alt=media)",
          "votes": 1,
          "replies": [
            {
              "id": 2308802,
              "postDate": "2023-06-19T08:19:11.287Z",
              "content": "<p>I feel that this method will introduce some errors, but it seems that everyone is doing it this way. Have you ever tried downsampling ground truth and calculating loss?</p>",
              "rawMarkdown": "I feel that this method will introduce some errors, but it seems that everyone is doing it this way. Have you ever tried downsampling ground truth and calculating loss?"
            }
          ]
        }
      ]
    },
    {
      "id": 2308584,
      "postDate": "2023-06-19T05:01:01.743Z",
      "content": "<p>it's awesome<br>\n😲</p>",
      "rawMarkdown": "it's awesome\n😲"
    },
    {
      "id": 2308290,
      "postDate": "2023-06-18T19:17:46.690Z",
      "content": "<p>Congratulations for the win, thanks for the writeup !  <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> </p>",
      "rawMarkdown": "Congratulations for the win, thanks for the writeup !  @ryches ",
      "replies": [
        {
          "id": 2308586,
          "postDate": "2023-06-19T05:01:22.900Z",
          "content": "<p>Congratulations to you</p>",
          "rawMarkdown": "Congratulations to you"
        }
      ]
    },
    {
      "id": 2307969,
      "postDate": "2023-06-18T14:54:09.960Z",
      "content": "<p>Great work Thanks for Sharing</p>",
      "rawMarkdown": "Great work Thanks for Sharing"
    },
    {
      "id": 2306834,
      "postDate": "2023-06-17T16:14:02.873Z",
      "content": "<p>great! thank for sharing.</p>",
      "rawMarkdown": "great! thank for sharing."
    }
  ],
  "comments": [
    {
      "id": 2304375,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2023-06-16T00:32:32.733000",
      "content": "<p>Hoping to aggregate a list of all the things we tried that didn't pan out as well but will have to add those later. One ironic thing with all of this is that I was visiting nvidia for work while the competition was closing so I got to see the news at their hq. Being there and winning a competition with one of their models dawned on me part way through the day. </p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 2304386,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2023-06-16T00:51:18.253000",
      "content": "<p>One thing we tried that didnt work was instead of filtering the crops based on if they were empty or not, filter them in case they had any ink or not. This massively reduced the number of crops and looked pretty good on paper but it taught the model that every single crop must have some ink and it didnt learn to handle blank papyrus well. Became overexcited and greatly over predicted. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F3e6b629a11dc8824e2d9b9e54a818ade%2FScreen%20Shot%202023-06-15%20at%205.50.22%20PM.png?generation=1686876640893243&amp;alt=media\" alt=\"\"></p>\n<p>maybe this would have worked with bigger crop sizes. This made it converge extremely fast but you can see the result isnt that great especially in reference to precision. </p>",
      "votes": 5,
      "replies": [
        {
          "id": 2895021,
          "author_name": "Vanshita Verma",
          "author_url": "",
          "post_date": "2024-06-28T20:50:56.683000",
          "content": "<p>Thanks for this explanation! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2305432,
      "author_name": "Heads or Tails",
      "author_url": "",
      "post_date": "2023-06-16T17:29:50.650000",
      "content": "<p>Many congrats on the win, and thanks for the fantastic write-up! Lots of valuable insight in there.</p>\n<p>I like the focus on depth invariance, and the ChannelDropout. One idea I had, which I didn't get round to testing, was to use ChannelShuffle in train &amp; tta. Did you happen you try that, and would you reckon it might help with the depth invariance?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2305456,
          "author_name": "Ted K",
          "author_url": "",
          "post_date": "2023-06-16T17:52:12.493000",
          "content": "<p>I didn't feel we had time to explore shuffling channels.  If the signal we're chasing is exclusively located within the XY plane of each slice, then this would help with regularization.  But to the extent that some of the signal lies in detecting where there is a transition from ink to no ink as you get deeper into the papyrus, then channel shuffle would have destroyed that.  It's an interesting ablation to try!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2305582,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2023-06-16T19:26:59.923000",
          "content": "<p>I actually tried a channel shuffled trial. It failed to converge at all. I specifically wanted to see how a sort of bag of depths approach worked and it completely failed with our models and a couple variations on them. I'll have to see if I can find the training graph for that one. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 2305590,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2023-06-16T19:31:37.607000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F642627376e3d046484bec6b82112a907%2FWB%20Chart%206_16_2023%2012_30_02%20PM.png?generation=1686943850364553&amp;alt=media\" alt=\"\"></p>\n<p>The one way down at the bottom is what happened when we shuffled the channels on the way into the model. Didnt seem to converge. We wouldve been better off just always grabbing the center channel than having them shuffled. </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2305632,
              "author_name": "Heads or Tails",
              "author_url": "",
              "post_date": "2023-06-16T20:07:12.490000",
              "content": "<p>That's very interesting, thank you! Is that with 16 or 32 channels? We did some analysis, which I'm hoping to post in our write-up soon, on the amount of information contained in each z-axis slice. From that, I'd guess that shuffling 32 channels would include a significant number of channels that just have noise, and maybe the shuffling could help for the few channels that contain the most ink signal.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305650,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2023-06-16T20:17:03.330000",
              "content": "<p>That was with the center 16 I believe. I also did a pretty thorough analysis of the info of ink in each layer of each fragment. I trained a single-layer model for each depth and then applied it to the held out fragment. You can see for each series there is a rise and fall of information density and at the ends there is nothing left. Just base guessing. </p>\n<p>I think we left a little performance on the table by not honing in better on which layers had ink. We tried some approachs to move the depth per fragment or per crop and couldnt find anything that did consistently better than just taking the middle 16. We couldve maybe done with middle 32 and gotten even more signal. Early on I was seeing that 32 depth beat 16, but for double the vram usage it didnt seem worth it. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Faabd32462a191a9b9e092ee7be888983%2FScreen%20Shot%202023-06-16%20at%201.14.27%20PM.png?generation=1686946620916433&amp;alt=media\" alt=\"\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2305663,
              "author_name": "Heads or Tails",
              "author_url": "",
              "post_date": "2023-06-16T20:32:22.627000",
              "content": "<p>Nice! Very consistent with my findings. Did you ever try models for just the top 4 or 5 layers in your plot that seem to contain the most signal? I thought that I might get some additional speed / larger batch sizes out of that but the performance drop was significant.</p>\n<p>For my depth scans, I also looked at whether different layers might contain more ink in different x, y regions of a given fragment, but didn't find any indications for that. Did you see anything along those lines?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305670,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2023-06-16T20:39:14.847000",
              "content": "<p>I tried varying placements and windows of depth but center 16 seemed sufficient. Couldnt find anything better. I tried to analytically find the best for each fragment individually. Taking the peak of average values per layer as the center point. I tried also doing that per crop, so for every crop I measured the average intensity per layer and then used the one with the highest value as the center and then went 8 up and down from there. </p>\n<p>Was extremely surprising to find that it did pretty comparably even though many of the crops had way different values than the middle 16, for some crops it would pick like the 15th layer as the center instead of the 32nd. Didnt pursue it too far. Maybe there was some better method for finding the regions most likely to have ink and then hone into them or broadcast our 2d label into 3d in a smart analytical way and then train against our generated 3d labels but we weren't able to come up with something that worked</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305672,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2023-06-16T20:44:44.367000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F4ffa67c0437055c8f3256673f1a20b6c%2FWB%20Chart%206_16_2023%201_42_18%20PM.png?generation=1686948170284985&amp;alt=media\" alt=\"\"></p>\n<p>this was trying to find the max response per layer per tile first. Maybe shouldve explored it more, but didnt immediately give any value. This was with old iterations of our model. Maybe worth revisiting. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305673,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2023-06-16T20:46:19.493000",
              "content": "<p>One of the things we tried early on was backing into the 3d labels by having our 3d unet output the 3d volume and then just take the max of the depth axis while training and then we could inspect its 3d volume without the max and see where it was finding the ink from. We could clearly see that it was picking things up from different layers</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305685,
              "author_name": "Heads or Tails",
              "author_url": "",
              "post_date": "2023-06-16T21:00:36.590000",
              "content": "<p>That's super interesting! If you have some sensitivity maps for different layers based on your 3d volumes then that would be a fascinating insight.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305687,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2023-06-16T21:04:16.807000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fe772fcafe4f29d22d7068ff6b0a57886%2F2-0_MdgQnL3V.gif?generation=1686949374516762&amp;alt=media\" alt=\"\"></p>\n<p>here is an example of our 3d models output. Sweeping through the 3d volume that it output before we crushed it with a max across the depth. You can see that it is seeing different stuff on different layers in different x,y regions. This could happen from the ink soaking in to different depths or it could be from the segmentation model used to flatten the fragments centering slightly differently. </p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2305691,
              "author_name": "Heads or Tails",
              "author_url": "",
              "post_date": "2023-06-16T21:16:37.157000",
              "content": "<p>Wow! This is mesmerizing. I think it would be very interesting to figure out how much of this are intrinsic differences in depth, and how robustly you could detect them at inference time. Probably less important for a depth-invariant model, but could still help to really tease out the ink from the noise on a local level.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305717,
              "author_name": "ryches",
              "author_url": "",
              "post_date": "2023-06-16T21:38:48.173000",
              "content": "<p>I think there is some potential for strong pseudo labeling still. First pass do the 3d volume to 2d label and then back out the 3d label from the model and train against that. Maybe the extra 3d localization helps the model be more precise and understand where things are coming from. Or maybe you just teach the second model to mimic the first. Hard to say without experiment. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305786,
              "author_name": "Heads or Tails",
              "author_url": "",
              "post_date": "2023-06-16T23:01:37.077000",
              "content": "<p>Agreed; this could produce a significant boost. This is all great stuff; thank you Ryan for all the detail! Much appreciated. I'm gonna think more about those 3d features and I'm happy to discuss further. I'll see if I can make it to your meetup event tomorrow.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2335550,
      "author_name": "Stephen Demjanenko",
      "author_url": "",
      "post_date": "2023-07-08T16:07:51.210000",
      "content": "<p>Congrats on the win! And thanks for sharing. I am a beginner when it comes to data science and I am trying to reproduce the results of one of your models (to prove to myself that I can make it work end to end). I have been unable to match the CV results that you shared for your models. I am running on a A100 in Google Colab, but I have to use 6 transformer heads for the UNETR, \"nvidia/mit-b0\" for the Segformer and set precision=\"16-mixed\" for Pytorch lightning and run with a resolution of 512 in order to not run the GPU out of memory (of which it has 40 GB). I also hit NaN loss values (likely due to mixed precision) when using AdamW and therefore I have to use SGD. I am finding myself struggling a lot more with setting up a training harness that can run a big model and reproduce the results than with testing different architectures.</p>\n<p>What hardware did you train on? How long did each epoch take? How many epochs did you run for? How did you tune the hyperparams as the epochs progressed?</p>\n<p>Additionally, how did you load the images from disk and did you pre-process them to change the numeric range/normalize (or did you rely on the Normalize transformation from albumentations entirely?</p>\n<p>How did you define your folds and did you retrain the model N times assuming you had N-folds? How long did all this training take? Did you then use your best trained fold, did you average the weights from all folds, or did you train without holding back one of the folds for validation?</p>\n<p>Thanks in advance for the clarification. My goal with these questions is to closely reproduce the performance of one of your models. If you have exact code to share, I would love to learn from that.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2310560,
      "author_name": "Guillermo Perez G",
      "author_url": "",
      "post_date": "2023-06-20T13:14:42.080000",
      "content": "<p>Thanks for sharing your thoughts! It's always important to know what someone who achieves success thinks! Does your team plan to publish your complete solution on Github or similar platform, or not?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2304410,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2023-06-16T01:36:10.717000",
      "content": "<p>Congratulations for the 1st prize.</p>\n<p>I was also impressed of the thorough research.</p>\n<p>I have questions, </p>\n<ul>\n<li>I once used SegFormer for 2D encoder (only 1st stage), but it seemed to take much epochs than CNN encoder to converge. However, as shown in the performance graph, the conversion seems after &lt;25 epochs. Do you think the second stage strategy (3D-2D) helps to converge faster?</li>\n<li>I was also observed very fluctuated metrics even if closed to convergence. It makes me hard to which model is strong is very difficult. So the most of experiments I used seed-average of the models. How do you evaluate the local model's performance?</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 2304431,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2023-06-16T02:00:55",
          "content": "<p>generally I felt like segformer converged more quickly than other stage 2 models we tried like unet, upernet with various backbones. not sure exactly why you saw slow convergence but I never tried it on its own. </p>\n<p>I mostly just looked at the general trend of the model. If I measured performance with the swa model performance movement was very slow and smooth curves, but we saw a similar effect just using the running average of the charts on wandb</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2305095,
      "author_name": "Viktor Cikojevic",
      "author_url": "",
      "post_date": "2023-06-16T12:44:46.533000",
      "content": "<p>nice writeup, congratulations for the 1st place!</p>\n<p>One question: did you train the 3dcnn and 2d segmentation models jointly or separately?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2305234,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2023-06-16T14:21:01.533000",
          "content": "<p>It was all trained end to end. My belief is training end to end and optimizing directly for your goal is almost always universally better. Sometimes you have to do things in stages just because of memory/compute constraints but in general I will always prefer end to end</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2305250,
              "author_name": "Viktor Cikojevic",
              "author_url": "",
              "post_date": "2023-06-16T14:29:41.893000",
              "content": "<p>Great, thanks for the insight! </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2305459,
              "author_name": "Ted K",
              "author_url": "",
              "post_date": "2023-06-16T17:55:45.930000",
              "content": "<p>I'll also add that I was a big proponent of end-to-end training.  The downside is you typically need more data for the deeper pipeline.  We only had three fragments to train on, but at least we had many crops per fragment and all of the image augmentations.  In the end, we were able to get end-to-end to converge pretty well, so we stuck with it.</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2541483,
      "author_name": "jrudolph",
      "author_url": "",
      "post_date": "2023-11-28T14:34:01.977000",
      "content": "<p>Congrats on winning this competition! I'd like to play around with the models and training / fine-tuning on the actual scroll data sets. Did you share the training code somewhere? Would you be open to share the training code?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2541737,
          "author_name": "JP Posma",
          "author_url": "",
          "post_date": "2023-11-28T18:29:17.767000",
          "content": "<p>All the code for the 10 winners is public, links are at <a href=\"https://scrollprize.org/community_projects#fragment-based-ink-detection\" target=\"_blank\">https://scrollprize.org/community_projects#fragment-based-ink-detection</a></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2541802,
              "author_name": "jrudolph",
              "author_url": "",
              "post_date": "2023-11-28T19:52:24.170000",
              "content": "<p>Ah, perfect, missed that. I looked in Ryan's Github but missed Aina's. Not the best premise for an aspiring treasure hunter…</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2319206,
      "author_name": "WangXuC",
      "author_url": "",
      "post_date": "2023-06-27T01:36:34.337000",
      "content": "<p>Where can I find the segformerForSemanticSegmentation function? I would like to see the performance of segfomer.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2319232,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2023-06-27T02:07:41.947000",
          "content": "<p>It comes from the transformers package</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2319280,
              "author_name": "WangXuC",
              "author_url": "",
              "post_date": "2023-06-27T03:30:51.570000",
              "content": "<p>Thank you! I find it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2319282,
              "author_name": "WangXuC",
              "author_url": "",
              "post_date": "2023-06-27T03:31:10.387000",
              "content": "<p>And how many epochs did you train for segformer?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2319279,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-27T03:30:25.430000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2310125,
      "author_name": "zhang jiaxun",
      "author_url": "",
      "post_date": "2023-06-20T06:04:34.990000",
      "content": "<p>Nice thoughts and ideas from you and this is helpful for learners to improve our skills and gain insights.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2310107,
      "author_name": "PavanKalyan",
      "author_url": "",
      "post_date": "2023-06-20T05:44:34.713000",
      "content": "<p>Informative, Congratulations 🎉</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2309989,
      "author_name": "Diesel X",
      "author_url": "",
      "post_date": "2023-06-20T03:38:11.840000",
      "content": "<p>Amazing job!!! Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2309944,
      "author_name": "Rakha Abid Bangsawan",
      "author_url": "",
      "post_date": "2023-06-20T02:29:26.550000",
      "content": "<p>so cool! congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2308738,
      "author_name": "WangXuC",
      "author_url": "",
      "post_date": "2023-06-19T07:10:38.967000",
      "content": "<p>I want to know how you calculate the loss? The mask size output by segfomer is 1/4 of the input size. If the input size is 512 * 512, then the predicted mask size is 128 * 128. When calculating losses, should the predicted mask be upsampled to 512 * 512, or should the ground truth be downsampled to 128 * 128?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2308790,
          "author_name": "Ted K",
          "author_url": "",
          "post_date": "2023-06-19T08:06:53.443000",
          "content": "<p>Yes, we upscale the predictions back to the original resolution.  If you look in the models section above, Ryan shows in the forward pass that the output of the Segformer goes through two 2x upscaling layers.  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F909432%2Fec8e2e11a7484f580f25262e8c8ef5c3%2Fupscale.png?generation=1687161955698991&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2308802,
              "author_name": "WangXuC",
              "author_url": "",
              "post_date": "2023-06-19T08:19:11.287000",
              "content": "<p>I feel that this method will introduce some errors, but it seems that everyone is doing it this way. Have you ever tried downsampling ground truth and calculating loss?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2308584,
      "author_name": "Jerjis Balamjan",
      "author_url": "",
      "post_date": "2023-06-19T05:01:01.743000",
      "content": "<p>it's awesome<br>\n😲</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2308290,
      "author_name": "Akshit Sharma",
      "author_url": "",
      "post_date": "2023-06-18T19:17:46.690000",
      "content": "<p>Congratulations for the win, thanks for the writeup !  <a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2308586,
          "author_name": "Jerjis Balamjan",
          "author_url": "",
          "post_date": "2023-06-19T05:01:22.900000",
          "content": "<p>Congratulations to you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2307969,
      "author_name": "Denslin Nunes",
      "author_url": "",
      "post_date": "2023-06-18T14:54:09.960000",
      "content": "<p>Great work Thanks for Sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2306834,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-06-17T16:14:02.873000",
      "content": "<p>great! thank for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2304372": "It was a very nervy ending to this competition, we held first for a long time but I am always scared of a shake-up. We tried to play it very defensively to ensure that we weren't doing anything to over-optimize for the leaderboard. Still, there is always a certain level of uncertainty when the public leaderboard is so small and your local validation is not extremely stable.\n\n**tl;dr**\nAt a high level, I would attribute most of our success to:\n- Larger crops \n- Strong depth-invariant models\n- Averaging several models to give us better calibration\n- Training against all available data after validating against fragment 1 rigorously\n\n**Data prep**\nWe initially started with the [2.5d starter code](https://www.kaggle.com/code/tanakar/2-5d-segmentaion-baseline-training) and evolved it over time. What we ended up settling on was just taking the middle 16 layers, but we tried plenty of other things that didn't work. I will save that for another post. Initially, the code was configured to train against all of the crops of the image, but we filtered this very simply based on if a crop was blank. This greatly reduced our training time because many of the crops had no data in them at all. \n\nEarly on there was very little progress being made and then we ran simple ablations to see how performance was impacted by the crop size and found that there was good scaling potential there. Seemed like one of the clearest patterns we saw early on. 128 got beat by 512 which got beaten by 1024. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fadddcf67a8c6ed4f68716e050215eb2d%2FWB%20Chart%206_15_2023%203_13_17%20PM.png?generation=1686867276107142&alt=media)\n\nIt seemed that the x,y context was very important here because the crop could look at a whole letter or even a couple letters at a time and draw them out more completely. Smaller crops yielded much patchier-looking outputs. One nice thing about this approach was it actually didn't increase training time at all. if you ran it with 1024x1024 crops you had to train against way less crops than if you did 128x128 crops so the epoch time and wall-clock convergence time was virtually identical, unlike some other competitions where increasing the resize makes the computation take longer. \n\n**Models**\nWe started with the 2.5D approach but it became apparent that it wasn't optimal because the model was learning which layers had ink in them and we knew that this varied between fragments. We wanted a solution that was depth-invariant, that could detect ink in any layer. One approach would be training a simple 2d model on each slice of the 3d volume, but with the way the data was labeled, only 2d to start with, we would be giving bad signal to the model if we told it ink was in layers that it actually wasnt and vice versa. We had many discussions about how we handle this, 1D convolutions, max pooling, 3d convolutions with size (8, 1, 1) so it was only truly looking across the depth patterns. What we found to be the most performant was using strong 3d models that would output a new 3d volume with many channels and then we could flatten them along the depth axis with a max. \n\nSo for example, our first approach in this vein was a simple 3dcnn. input shape was (batch_size, 1, 16, 1024, 1024). We applied 4 layers of 3d convolutions on top of this with progressively more filters until finally, we had an output volume of (batch_size, 16, 16, 1024, 1024). At this point we just took the max across the z-dimension to squish it down to (batch_size, 16, 1024, 1024) so our 16 depth dimensions were now replaced with 16 feature dimensions. This alone was a decent approach but then passing this through a strong 2d segmentation model made it much better. We heavily relied on segformer for this, as others have mentioned b3 backbone worked great, but we also found that directionally the even bigger models performed better on the leaderboard even if they didnt perform measurably better on local validation. \n\nWe iterated on this design because it seemed to satisfy the qualities we wanted, a strong 2d segmenter applied to a 3d volume that was invariant to depth. We tried different first-stage 3d methods and found that even stronger 3d models yielded better results. Evolving to 3d unet's and then eventually 3d unetr. Sometimes the results were unclear if they were better but we found a clear trend that the on-paper better models also did better on the leaderboard. Our best individual model was a UNETR first stage that passed 32 channels into a b5 segformer which scored 0.82 on the public leaderboard and 0.67 on the private. We applied a small amount of dropout on the channels in between the 3d and 2d stage. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fb59a1b2b524ffdbea3882fd6883e87ad%2FScreen%20Shot%202023-06-15%20at%204.12.57%20PM.png?generation=1686870810392089&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Fd3566ee53b477abe6e4026db29d35531%2FScreen%20Shot%202023-06-15%20at%204.03.32%20PM.png?generation=1686870239615259&alt=media)\n\nOne thing I always wondered about was if the strength of the segformer wasnt actually that it was the best model but that it was making predictions at a lower resolution. We pass it the 1024x1024 and it returns back a 256x256 segmentation map. We simply upscaled it with a very simple conv2d transpose. I believe the second place solution verifies this that the lower resolution actually helps because it is doing much coarser classification which makes it easier than trying to more precisely get every pixel. \n\nOur final best solution was 9 different models\n\n\n| 1st stage \t| 2nd stage \t| Resolution \t| Public Score \t|\n|---\t|---\t|---\t|---\t|\n| 3d unet(16 channels) \t| segformer b3 \t| 1024 \t| .78 \t|\n| 3d unet(16 channels) \t| segformer b3 \t| 512 \t| .77 \t|\n| 3d unet(16 channels with SWA) \t| segformer b3 \t| 512 \t| .78 \t|\n| 3d cnn(32 channels) \t| segformer b3 \t| 1024 \t| .78 \t|\n| 3d cnn(32 channels) \t| segformer b5 \t| 1024 \t| .77 \t|\n| 3d cnn(64 channels) \t| segformer b3 \t| 1024 \t| .78 \t|\n| 3d unet(32 channels) \t| segformer b5 \t| 1024 \t| .79 \t|\n| 3d unetr(32 channels) \t| segformer b5 \t| 512 \t| .82 \t|\n| 3d unetr multiclass(32 channels) \t| segformer b5 \t| 512 \t| ? \t|\n\n\nOne thing we tried late-on was multi-class output. Instead of binary we tried to predict nothing-mask-ink. This ended up yielding a much cleaner looking output actually, in most of our models we had some amount of noise anywhere there was papyrus and this seemed to quiet that down a lot. We did not get to explore it thoroughly enough to confirm if this really worked or not though. \n\n\n**Training procedure**\nOur training procedure was fairly standard, we used the existing adamW optimizer, dice+bce loss and hyperparameters, fixing some small bugs in setting the min-lr and continuing to use the gradual warmup learning rate scheduling. We added on stochastic weight averaging to get wider optima instead of needing to pick a specific checkpoint because we found these pretty inconsistent. We found that validating against fragment 1 wasnt perfect but was at least directionally useful. Once we found something that worked on local validation against fold 1 we would submit that to the leaderboard to double-check its validity and we would continue training with all folds for several epochs after. We would submit and evaluate the new checkpoint trained against all folds and it was typically about 0.04 better. Sometimes we would train against all fragments from the start instead of training fragments 2, 3 and latter adding in 1, but we did not thoroughly evaluate this. \n\n**Augmentations**\nWe tried many permutations of augmentation, mostly from albumentations, some custom, and some from 3d packages, but ultimately couldn't find strong alpha there with anything fancy. Our best model was trained with:\n- 50% horizontal and vertical flips\n- 75% 90-degree rotations\n- 50% brightness contrast\n- 25% 1-2 channel dropout(in this case our channels was actually depth)\n- 10% shift scale rotate\n- 10% noise and blur\n- 10% coarse dropout\n- 10% grid distortion\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F44380bdf9ac40366cfd12f6db39f1bef%2FScreen%20Shot%202023-06-15%20at%204.29.24%20PM.png?generation=1686871786148868&alt=media)\n\nOverall seemed like the rotations and flips were crucial and everything else was a non-factor. Rotations seemed important but we did not ascertain that the test set itself was actually rotated, we just knew it was important so tried to make the model as invariant as possible. \n\n\n**Ensembling**\nWe ended up with a big pile of model checkpoints and had to whittle them down to what we believed to be the most performant trading off runtime vs throughput. We tried many different combinations, heavier tta with all rotations and flips, more models, smaller strides. The winner seemed to be 4x rotation TTA with 1/4 crop strided windows and as many good models as we could fit in. Halving the stride helped an extra .01, but that was trumped by being able to add way more models, flips didnt seem to add anything at all. It's possible with just the corrected orientation instead of TTA we could have done better but a model invariant to rotations seemed just as strong.\n\nOne thing we struggled with a lot was how to best combine predictions. For each model we predicted on the same pixel 4x because of our strided approach and 4x of that because of TTA. With many models this actually gave us a ton of options for aggregating the predictions per pixel. We tinkered with a lot of stuff locally but what seemed to work best was averaging per pixel for each model all ~16x predictions and then applying the sigmoid to that averaged signal and then average those probabilities together. This posed a tough memory constraint on us on kaggles system so we had to be a bit efficient with putting things away and accumulating them as densely as possible instead of just creating large arrays and averaging at the end. \n\n**Thresholding**\nOne thing we went back and forth on for a long time was the calibration of predictions. As discussed in other posts, deciding a threshold was critical to getting good results and choosing the wrong threshold could give you very misleading signal on your models performance. During training and evaluation we would constantly be monitoring the AUC, Precision, Recall and f0.5 at a sweep of thresholds. Some models were well-calibrated with an optimal threshold at 0.5, but many were not. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffee3f2b5c722068974d291af29fc7049%2FScreen%20Shot%202023-06-15%20at%204.58.40%20PM.png?generation=1686873556582356&alt=media)\n\nBecause of this, we knew that even a great model could show up as terrible on the leaderboard if its threshold wasn't set correctly. We considered using the percentile method that some people used but did not even make a submission for it because it seemed too risky if the distribution of ink was not what we expected. Ultimately what we relied on was that averaged out our predictions they would end up calibrated. We found this to be true on our local validation and held true for the leaderboard as well. Individual models would have wide optimal threshold ranges but after averaging many predictions from many models it was almost universally centered on 0.5. We actually took our last submission to be brave and try 0.55 but it did worse on the public leaderboard and finished with only 30 minutes left to spare so we didnt pick it. It did end up performing slightly higher on the private leaderboard though. I think we actually left a lot on the table in terms of ensembling and optimal thresholding. \n\n**Postprocessing**\nOne little magic function we borrowed from the old cloud segmentation competition was used to clean up our predictions. After everything was done and already binarized we would use cv2.connectedcomponents to find the masses and remove them if they were beyond a certain size. This would clean up anything that was too small, just little speckles and noise. We found it locally optimal to set our threshold a little lower and then clean up the extras, anything under 25000 pixels, but we only cleaned up things under 10k on our submissions and didnt try more aggressive cleaning. \n\nBefore\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F25116a2dfc769b5784abb12704df8401%2F__results___27_30.png?generation=1686874373183834&alt=media)\n\nAfter\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F9a6ac60974f49a6ecfbb80f47fcccbb0%2F__results___27_31.png?generation=1686874392858212&alt=media)\n",
    "2304375": "Hoping to aggregate a list of all the things we tried that didn't pan out as well but will have to add those later. One ironic thing with all of this is that I was visiting nvidia for work while the competition was closing so I got to see the news at their hq. Being there and winning a competition with one of their models dawned on me part way through the day. ",
    "2304386": "One thing we tried that didnt work was instead of filtering the crops based on if they were empty or not, filter them in case they had any ink or not. This massively reduced the number of crops and looked pretty good on paper but it taught the model that every single crop must have some ink and it didnt learn to handle blank papyrus well. Became overexcited and greatly over predicted. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F3e6b629a11dc8824e2d9b9e54a818ade%2FScreen%20Shot%202023-06-15%20at%205.50.22%20PM.png?generation=1686876640893243&alt=media)\n\nmaybe this would have worked with bigger crop sizes. This made it converge extremely fast but you can see the result isnt that great especially in reference to precision. ",
    "2305432": "Many congrats on the win, and thanks for the fantastic write-up! Lots of valuable insight in there.\n\nI like the focus on depth invariance, and the ChannelDropout. One idea I had, which I didn't get round to testing, was to use ChannelShuffle in train & tta. Did you happen you try that, and would you reckon it might help with the depth invariance?",
    "2335550": "Congrats on the win! And thanks for sharing. I am a beginner when it comes to data science and I am trying to reproduce the results of one of your models (to prove to myself that I can make it work end to end). I have been unable to match the CV results that you shared for your models. I am running on a A100 in Google Colab, but I have to use 6 transformer heads for the UNETR, \"nvidia/mit-b0\" for the Segformer and set precision=\"16-mixed\" for Pytorch lightning and run with a resolution of 512 in order to not run the GPU out of memory (of which it has 40 GB). I also hit NaN loss values (likely due to mixed precision) when using AdamW and therefore I have to use SGD. I am finding myself struggling a lot more with setting up a training harness that can run a big model and reproduce the results than with testing different architectures.\n\nWhat hardware did you train on? How long did each epoch take? How many epochs did you run for? How did you tune the hyperparams as the epochs progressed?\n\nAdditionally, how did you load the images from disk and did you pre-process them to change the numeric range/normalize (or did you rely on the Normalize transformation from albumentations entirely?\n\nHow did you define your folds and did you retrain the model N times assuming you had N-folds? How long did all this training take? Did you then use your best trained fold, did you average the weights from all folds, or did you train without holding back one of the folds for validation?\n\nThanks in advance for the clarification. My goal with these questions is to closely reproduce the performance of one of your models. If you have exact code to share, I would love to learn from that.",
    "2310560": "Thanks for sharing your thoughts! It's always important to know what someone who achieves success thinks! Does your team plan to publish your complete solution on Github or similar platform, or not?",
    "2304410": "Congratulations for the 1st prize.\n\nI was also impressed of the thorough research.\n\nI have questions, \n\n- I once used SegFormer for 2D encoder (only 1st stage), but it seemed to take much epochs than CNN encoder to converge. However, as shown in the performance graph, the conversion seems after <25 epochs. Do you think the second stage strategy (3D-2D) helps to converge faster?\n- I was also observed very fluctuated metrics even if closed to convergence. It makes me hard to which model is strong is very difficult. So the most of experiments I used seed-average of the models. How do you evaluate the local model's performance?",
    "2305095": "nice writeup, congratulations for the 1st place!\n\nOne question: did you train the 3dcnn and 2d segmentation models jointly or separately?",
    "2541483": "Congrats on winning this competition! I'd like to play around with the models and training / fine-tuning on the actual scroll data sets. Did you share the training code somewhere? Would you be open to share the training code?",
    "2319206": "Where can I find the segformerForSemanticSegmentation function? I would like to see the performance of segfomer.",
    "2310125": "Nice thoughts and ideas from you and this is helpful for learners to improve our skills and gain insights.",
    "2310107": "Informative, Congratulations 🎉",
    "2309989": "Amazing job!!! Congrats!",
    "2309944": "so cool! congrats!",
    "2308738": "I want to know how you calculate the loss? The mask size output by segfomer is 1/4 of the input size. If the input size is 512 * 512, then the predicted mask size is 128 * 128. When calculating losses, should the predicted mask be upsampled to 512 * 512, or should the ground truth be downsampled to 128 * 128?",
    "2308584": "it's awesome\n😲",
    "2308290": "Congratulations for the win, thanks for the writeup !  @ryches ",
    "2307969": "Great work Thanks for Sharing",
    "2306834": "great! thank for sharing."
  }
}