{
  "id": 456787,
  "title": "[LB 0.726] Experimental Configuration and Results",
  "url": "/competitions/blood-vessel-segmentation/discussion/456787",
  "author_name": "kcetskcaz",
  "post_date": "2023-11-21T16:57:22.719000",
  "votes": 40,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>Key Takeaways So Far (For Me)</h1>\n<p>Will update this thread as I continue to try new things. A few high level takeaways (in order of percentage improvement provided):</p>\n<ol>\n<li><strong>Combine masks intelligently</strong> - If you're performing any tile/sub-image processing, make sure that you recombine masks intelligently. I'm tiling using the code I posted <a href=\"https://www.kaggle.com/code/squidinator/sennet-hoa-in-memory-tiled-dataset-pytorch\" target=\"_blank\">here</a>. Critically, if you're tiling the scenes, you must project the masks back to the scene coordinate system at some point before submission. Originally, I was naively saving the results for a tile using <code>=</code> assignment, which overwrote detections in overlap regions. Once I changed this to write back with a logical \"or\", I saw a ~3% increase in LB performance. This is likely an artifact of using tile level statistics for dynamic range compression rather than scene level statistics, which has introduced artifacts in geospatial applications in the past (e.g., the network doesn't detect the target in one tile, but then detects it in the adjacent tile. If the target lies in the overlap region, then combining with a logical \"or\" takes the best (and worst) of both tiles).</li>\n<li><strong>Operating point is important</strong> - Although some are encountering issues with their selected score threshold and OOM errors, I'd like to emphasize that selecting the appropriate score threshold for your respective model matters. See the below Table 1 for an example.</li>\n<li><strong>Subsample the data to iterate experiments quickly</strong> - In combination with prototyping a small model (e.g., efficientnet-b1), you can iterate extremely quickly while making the assumption that findings on the subset and small model <em>should</em> generalize to the full train/val sets + larger model.</li>\n</ol>\n<hr>\n<h1>Experiment Configuration</h1>\n<p><strong>Architecture</strong>: UNet/UNet++ (investigating others)<br>\n<strong>Backbone</strong>: </p>\n<ul>\n<li>seresnext26d (best results so far)</li>\n<li>efficientnet-b4</li>\n<li>efficientnet-b1</li>\n<li>Investigating other backbones</li>\n</ul>\n<p><strong>Loss</strong>: DiceLoss (Investigating others)<br>\n<strong>Model Selection Metric</strong>: DiceLoss<br>\n<strong>Optimizer</strong>: AdamW (Investigating others)<br>\n<strong>Scheduler</strong>: CosineAnnealingLR<br>\n<strong>Transforms</strong>: </p>\n<ul>\n<li>Train Transforms = HFlip, VFlip, RRotate90, ElasticTransform, RandomBrightness/Contrast</li>\n<li>Validation Transforms = None</li>\n</ul>\n<p><strong>Tile Size</strong>:</p>\n<ul>\n<li>Train: 512x512 w/ 20% overlap</li>\n<li>Inference: 800x800 w/ 20% overlap (improves performance, needs investigated)</li>\n</ul>\n<p><strong>Epochs</strong>: 20 (no model best checkpoints have occurred after 20 epochs)</p>\n<hr>\n<h1>Experiment Results (Ongoing)</h1>\n<h3>Model Descriptions:</h3>\n<table>\n<thead>\n<tr>\n<th>Experiment Name</th>\n<th>Meta Architecture</th>\n<th>Encoder</th>\n<th>Epochs Trained</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B1</td>\n<td>UNet</td>\n<td>efficientnet-b4</td>\n<td>1</td>\n</tr>\n<tr>\n<td>B2</td>\n<td>UNet</td>\n<td>efficientnet-b4</td>\n<td>5</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>UNet++</td>\n<td>efficientnet-b4</td>\n<td>5</td>\n</tr>\n<tr>\n<td>B4</td>\n<td>UNet++</td>\n<td>efficientnet-b6</td>\n<td>8</td>\n</tr>\n<tr>\n<td>B5</td>\n<td>UNet++</td>\n<td>seresnext26d</td>\n<td>11</td>\n</tr>\n</tbody>\n</table>\n<h3>Naive random 80:20 w/ fixed seed=42</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>LB</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.67</td>\n<td>0.44</td>\n<td>B1</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.56</td>\n<td>0.49</td>\n<td>B2</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.56</td>\n<td>0.54</td>\n<td>B2 + Logical \"OR\" Mask Combination (LOMC)</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td><strong>0.58</strong></td>\n<td><strong>0.57</strong></td>\n<td>B2 + LOMC+ Reduced Score Threshold</td>\n</tr>\n</tbody>\n</table>\n<p>*All models moving forward will use LOMC.<br>\n*Unless otherwise noted, all inference runs use 800x800 tile size</p>\n<h3>Train on all, CV on Kidney_3_dense</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>LB</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.53</td>\n<td>0.652</td>\n<td>B3</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.55</td>\n<td><strong>0.655</strong></td>\n<td>B3d</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td><strong>0.55</strong></td>\n<td>OOM</td>\n<td>B3 + 512x512 Tiles</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td><strong>0.55</strong></td>\n<td>OOM</td>\n<td>B3 + 512x512 Tiles</td>\n</tr>\n</tbody>\n</table>\n<h3>Train on Kidney 1 Dense, CV on Kidney 3 Dense</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>LB</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.05</td>\n<td>0.74</td>\n<td>0.658</td>\n<td>B4 + 384x384 Tiles</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.74</td>\n<td>0.663</td>\n<td>B4 + 800x800 Tiles</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.74</td>\n<td>0.673</td>\n<td>B4 + 800x800 Tiles + TTA</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td><strong>0.89</strong></td>\n<td>0.72</td>\n<td>B5 + 800x800 Tiles + TTA</td>\n</tr>\n<tr>\n<td>0.01</td>\n<td><strong>0.89</strong></td>\n<td><strong>0.726</strong></td>\n<td>B5 + 800x800 Tiles + TTA</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2533206,
      "postDate": "2023-11-21T16:57:22.720Z",
      "content": "<h1>Key Takeaways So Far (For Me)</h1>\n<p>Will update this thread as I continue to try new things. A few high level takeaways (in order of percentage improvement provided):</p>\n<ol>\n<li><strong>Combine masks intelligently</strong> - If you're performing any tile/sub-image processing, make sure that you recombine masks intelligently. I'm tiling using the code I posted <a href=\"https://www.kaggle.com/code/squidinator/sennet-hoa-in-memory-tiled-dataset-pytorch\" target=\"_blank\">here</a>. Critically, if you're tiling the scenes, you must project the masks back to the scene coordinate system at some point before submission. Originally, I was naively saving the results for a tile using <code>=</code> assignment, which overwrote detections in overlap regions. Once I changed this to write back with a logical \"or\", I saw a ~3% increase in LB performance. This is likely an artifact of using tile level statistics for dynamic range compression rather than scene level statistics, which has introduced artifacts in geospatial applications in the past (e.g., the network doesn't detect the target in one tile, but then detects it in the adjacent tile. If the target lies in the overlap region, then combining with a logical \"or\" takes the best (and worst) of both tiles).</li>\n<li><strong>Operating point is important</strong> - Although some are encountering issues with their selected score threshold and OOM errors, I'd like to emphasize that selecting the appropriate score threshold for your respective model matters. See the below Table 1 for an example.</li>\n<li><strong>Subsample the data to iterate experiments quickly</strong> - In combination with prototyping a small model (e.g., efficientnet-b1), you can iterate extremely quickly while making the assumption that findings on the subset and small model <em>should</em> generalize to the full train/val sets + larger model.</li>\n</ol>\n<hr>\n<h1>Experiment Configuration</h1>\n<p><strong>Architecture</strong>: UNet/UNet++ (investigating others)<br>\n<strong>Backbone</strong>: </p>\n<ul>\n<li>seresnext26d (best results so far)</li>\n<li>efficientnet-b4</li>\n<li>efficientnet-b1</li>\n<li>Investigating other backbones</li>\n</ul>\n<p><strong>Loss</strong>: DiceLoss (Investigating others)<br>\n<strong>Model Selection Metric</strong>: DiceLoss<br>\n<strong>Optimizer</strong>: AdamW (Investigating others)<br>\n<strong>Scheduler</strong>: CosineAnnealingLR<br>\n<strong>Transforms</strong>: </p>\n<ul>\n<li>Train Transforms = HFlip, VFlip, RRotate90, ElasticTransform, RandomBrightness/Contrast</li>\n<li>Validation Transforms = None</li>\n</ul>\n<p><strong>Tile Size</strong>:</p>\n<ul>\n<li>Train: 512x512 w/ 20% overlap</li>\n<li>Inference: 800x800 w/ 20% overlap (improves performance, needs investigated)</li>\n</ul>\n<p><strong>Epochs</strong>: 20 (no model best checkpoints have occurred after 20 epochs)</p>\n<hr>\n<h1>Experiment Results (Ongoing)</h1>\n<h3>Model Descriptions:</h3>\n<table>\n<thead>\n<tr>\n<th>Experiment Name</th>\n<th>Meta Architecture</th>\n<th>Encoder</th>\n<th>Epochs Trained</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>B1</td>\n<td>UNet</td>\n<td>efficientnet-b4</td>\n<td>1</td>\n</tr>\n<tr>\n<td>B2</td>\n<td>UNet</td>\n<td>efficientnet-b4</td>\n<td>5</td>\n</tr>\n<tr>\n<td>B3</td>\n<td>UNet++</td>\n<td>efficientnet-b4</td>\n<td>5</td>\n</tr>\n<tr>\n<td>B4</td>\n<td>UNet++</td>\n<td>efficientnet-b6</td>\n<td>8</td>\n</tr>\n<tr>\n<td>B5</td>\n<td>UNet++</td>\n<td>seresnext26d</td>\n<td>11</td>\n</tr>\n</tbody>\n</table>\n<h3>Naive random 80:20 w/ fixed seed=42</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>LB</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.67</td>\n<td>0.44</td>\n<td>B1</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.56</td>\n<td>0.49</td>\n<td>B2</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.56</td>\n<td>0.54</td>\n<td>B2 + Logical \"OR\" Mask Combination (LOMC)</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td><strong>0.58</strong></td>\n<td><strong>0.57</strong></td>\n<td>B2 + LOMC+ Reduced Score Threshold</td>\n</tr>\n</tbody>\n</table>\n<p>*All models moving forward will use LOMC.<br>\n*Unless otherwise noted, all inference runs use 800x800 tile size</p>\n<h3>Train on all, CV on Kidney_3_dense</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>LB</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.53</td>\n<td>0.652</td>\n<td>B3</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.55</td>\n<td><strong>0.655</strong></td>\n<td>B3d</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td><strong>0.55</strong></td>\n<td>OOM</td>\n<td>B3 + 512x512 Tiles</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td><strong>0.55</strong></td>\n<td>OOM</td>\n<td>B3 + 512x512 Tiles</td>\n</tr>\n</tbody>\n</table>\n<h3>Train on Kidney 1 Dense, CV on Kidney 3 Dense</h3>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV</th>\n<th>LB</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.05</td>\n<td>0.74</td>\n<td>0.658</td>\n<td>B4 + 384x384 Tiles</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.74</td>\n<td>0.663</td>\n<td>B4 + 800x800 Tiles</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.74</td>\n<td>0.673</td>\n<td>B4 + 800x800 Tiles + TTA</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td><strong>0.89</strong></td>\n<td>0.72</td>\n<td>B5 + 800x800 Tiles + TTA</td>\n</tr>\n<tr>\n<td>0.01</td>\n<td><strong>0.89</strong></td>\n<td><strong>0.726</strong></td>\n<td>B5 + 800x800 Tiles + TTA</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# Key Takeaways So Far (For Me)\n\nWill update this thread as I continue to try new things. A few high level takeaways (in order of percentage improvement provided):\n\n1. **Combine masks intelligently** - If you're performing any tile/sub-image processing, make sure that you recombine masks intelligently. I'm tiling using the code I posted [here](https://www.kaggle.com/code/squidinator/sennet-hoa-in-memory-tiled-dataset-pytorch). Critically, if you're tiling the scenes, you must project the masks back to the scene coordinate system at some point before submission. Originally, I was naively saving the results for a tile using `=` assignment, which overwrote detections in overlap regions. Once I changed this to write back with a logical \"or\", I saw a ~3% increase in LB performance. This is likely an artifact of using tile level statistics for dynamic range compression rather than scene level statistics, which has introduced artifacts in geospatial applications in the past (e.g., the network doesn't detect the target in one tile, but then detects it in the adjacent tile. If the target lies in the overlap region, then combining with a logical \"or\" takes the best (and worst) of both tiles).\n2. **Operating point is important** - Although some are encountering issues with their selected score threshold and OOM errors, I'd like to emphasize that selecting the appropriate score threshold for your respective model matters. See the below Table 1 for an example.\n3. **Subsample the data to iterate experiments quickly** - In combination with prototyping a small model (e.g., efficientnet-b1), you can iterate extremely quickly while making the assumption that findings on the subset and small model *should* generalize to the full train/val sets + larger model.\n____\n\n# Experiment Configuration\n\n**Architecture**: UNet/UNet++ (investigating others)\n**Backbone**: \n- seresnext26d (best results so far)\n- efficientnet-b4\n- efficientnet-b1\n- Investigating other backbones\n\n**Loss**: DiceLoss (Investigating others)\n**Model Selection Metric**: DiceLoss\n**Optimizer**: AdamW (Investigating others)\n**Scheduler**: CosineAnnealingLR\n**Transforms**: \n- Train Transforms = HFlip, VFlip, RRotate90, ElasticTransform, RandomBrightness/Contrast\n- Validation Transforms = None\n\n**Tile Size**:\n- Train: 512x512 w/ 20% overlap\n- Inference: 800x800 w/ 20% overlap (improves performance, needs investigated)\n\n**Epochs**: 20 (no model best checkpoints have occurred after 20 epochs)\n____\n\n# Experiment Results (Ongoing)\n\n### Model Descriptions:\n\n| Experiment Name | Meta Architecture | Encoder | Epochs Trained |\n| --- | --- | --- | --- |\n| B1 |  UNet | efficientnet-b4 | 1 |\n| B2 |  UNet | efficientnet-b4 | 5 |\n| B3 |  UNet++ | efficientnet-b4 | 5 |\n| B4 |  UNet++ | efficientnet-b6 | 8 |\n| B5 |  UNet++ | seresnext26d | 11 |\n\n### Naive random 80:20 w/ fixed seed=42\n\n| Threshold | CV  | LB | Notes |\n| --- | --- | --- | --- |\n| 0.5 |  0.67 | 0.44 | B1 |\n| 0.5 |  0.56 | 0.49 | B2 |\n| 0.5 |  0.56 | 0.54 | B2 + Logical \"OR\" Mask Combination (LOMC) |\n| 0.05 |  **0.58** | **0.57** | B2 + LOMC+ Reduced Score Threshold |\n\n*All models moving forward will use LOMC.\n*Unless otherwise noted, all inference runs use 800x800 tile size\n\n### Train on all, CV on Kidney_3_dense\n\n| Threshold | CV  | LB | Notes |\n| --- | --- | --- | --- |\n| 0.5 |  0.53 | 0.652 | B3 |\n| 0.05 |  0.55 | **0.655** | B3d |\n| 0.05 |  **0.55** | OOM | B3 + 512x512 Tiles |\n| 0.05 |  **0.55** | OOM | B3 + 512x512 Tiles |\n\n### Train on Kidney 1 Dense, CV on Kidney 3 Dense\n\n| Threshold | CV  | LB | Notes |\n| --- | --- | --- | --- |\n| 0.05 |  0.74 | 0.658 | B4 + 384x384 Tiles |\n| 0.05 |  0.74 | 0.663 | B4 + 800x800 Tiles |\n| 0.05 |  0.74 | 0.673 | B4 + 800x800 Tiles + TTA |\n| 0.05 |  **0.89** | 0.72 | B5 + 800x800 Tiles + TTA |\n| 0.01 |  **0.89** | **0.726**| B5 + 800x800 Tiles + TTA |",
      "votes": 40
    },
    {
      "id": 2534985,
      "postDate": "2023-11-23T03:22:53.747Z",
      "content": "<p>thx for sharing.</p>",
      "rawMarkdown": "thx for sharing.",
      "votes": 1
    },
    {
      "id": 2533565,
      "postDate": "2023-11-22T03:39:54.400Z",
      "content": "<p>what is your image dimensions</p>",
      "rawMarkdown": "what is your image dimensions",
      "votes": 1,
      "replies": [
        {
          "id": 2534382,
          "postDate": "2023-11-22T15:24:23.647Z",
          "content": "<p>I'm tiling the full scenes in to square sub-images for processing (basically a sliding window). This avoids downsampling the whole image, which introduces distortions and reduces the resolution of the target regions. It also avoids any form of cropping that my accidentally remove target regions from the image. </p>\n<p>I have some helpful logic that subsamples \"background\" tiles to keep the class ratio somewhat balanced. This is all done online during training to avoid writing mass quantites of data to disk (I don't yet know what tile size is optimal, but 512x512 has worked well). I'll release the code shortly.</p>",
          "rawMarkdown": "I'm tiling the full scenes in to square sub-images for processing (basically a sliding window). This avoids downsampling the whole image, which introduces distortions and reduces the resolution of the target regions. It also avoids any form of cropping that my accidentally remove target regions from the image. \n\nI have some helpful logic that subsamples \"background\" tiles to keep the class ratio somewhat balanced. This is all done online during training to avoid writing mass quantites of data to disk (I don't yet know what tile size is optimal, but 512x512 has worked well). I'll release the code shortly.",
          "votes": 2
        },
        {
          "id": 2534500,
          "postDate": "2023-11-22T17:08:13.777Z",
          "content": "<p>Hi Arunodhayan, I posted the code for in-memory tiling here: <a href=\"https://www.kaggle.com/code/squidinator/sennet-hoa-in-memory-tiled-dataset-pytorch\" target=\"_blank\">https://www.kaggle.com/code/squidinator/sennet-hoa-in-memory-tiled-dataset-pytorch</a></p>\n<p>If you have any questions or suggestions, don't hesititate to reach out!</p>",
          "rawMarkdown": "Hi Arunodhayan, I posted the code for in-memory tiling here: https://www.kaggle.com/code/squidinator/sennet-hoa-in-memory-tiled-dataset-pytorch\n\nIf you have any questions or suggestions, don't hesititate to reach out!",
          "votes": 7,
          "replies": [
            {
              "id": 2538293,
              "postDate": "2023-11-26T01:50:29.990Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2535052,
      "postDate": "2023-11-23T04:40:47.383Z",
      "content": "<p>Still in many cases OOM error …hope it gets solved soon</p>",
      "rawMarkdown": "Still in many cases OOM error ...hope it gets solved soon",
      "votes": 2
    },
    {
      "id": 2546087,
      "postDate": "2023-12-02T05:23:26.820Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/squidinator\" target=\"_blank\">@squidinator</a>. What's your best model's DiceLoss on validation set? I see you use DiceLoss as model selection function.</p>",
      "rawMarkdown": "Hi, @squidinator. What's your best model's DiceLoss on validation set? I see you use DiceLoss as model selection function."
    },
    {
      "id": 2536180,
      "postDate": "2023-11-24T01:53:50.963Z",
      "content": "<p>I added  remove_small ， but still scoring error for forked public baseline.</p>",
      "rawMarkdown": "I added  remove_small ， but still scoring error for forked public baseline."
    }
  ],
  "comments": [
    {
      "id": 2534985,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-11-23T03:22:53.747000",
      "content": "<p>thx for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2533565,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2023-11-22T03:39:54.400000",
      "content": "<p>what is your image dimensions</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2534382,
          "author_name": "kcetskcaz",
          "author_url": "",
          "post_date": "2023-11-22T15:24:23.647000",
          "content": "<p>I'm tiling the full scenes in to square sub-images for processing (basically a sliding window). This avoids downsampling the whole image, which introduces distortions and reduces the resolution of the target regions. It also avoids any form of cropping that my accidentally remove target regions from the image. </p>\n<p>I have some helpful logic that subsamples \"background\" tiles to keep the class ratio somewhat balanced. This is all done online during training to avoid writing mass quantites of data to disk (I don't yet know what tile size is optimal, but 512x512 has worked well). I'll release the code shortly.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2534500,
          "author_name": "kcetskcaz",
          "author_url": "",
          "post_date": "2023-11-22T17:08:13.777000",
          "content": "<p>Hi Arunodhayan, I posted the code for in-memory tiling here: <a href=\"https://www.kaggle.com/code/squidinator/sennet-hoa-in-memory-tiled-dataset-pytorch\" target=\"_blank\">https://www.kaggle.com/code/squidinator/sennet-hoa-in-memory-tiled-dataset-pytorch</a></p>\n<p>If you have any questions or suggestions, don't hesititate to reach out!</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2538293,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-11-26T01:50:29.990000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2535052,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2023-11-23T04:40:47.383000",
      "content": "<p>Still in many cases OOM error …hope it gets solved soon</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2546087,
      "author_name": "Snorf",
      "author_url": "",
      "post_date": "2023-12-02T05:23:26.820000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/squidinator\" target=\"_blank\">@squidinator</a>. What's your best model's DiceLoss on validation set? I see you use DiceLoss as model selection function.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2536180,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-11-24T01:53:50.963000",
      "content": "<p>I added  remove_small ， but still scoring error for forked public baseline.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2533206": "# Key Takeaways So Far (For Me)\n\nWill update this thread as I continue to try new things. A few high level takeaways (in order of percentage improvement provided):\n\n1. **Combine masks intelligently** - If you're performing any tile/sub-image processing, make sure that you recombine masks intelligently. I'm tiling using the code I posted [here](https://www.kaggle.com/code/squidinator/sennet-hoa-in-memory-tiled-dataset-pytorch). Critically, if you're tiling the scenes, you must project the masks back to the scene coordinate system at some point before submission. Originally, I was naively saving the results for a tile using `=` assignment, which overwrote detections in overlap regions. Once I changed this to write back with a logical \"or\", I saw a ~3% increase in LB performance. This is likely an artifact of using tile level statistics for dynamic range compression rather than scene level statistics, which has introduced artifacts in geospatial applications in the past (e.g., the network doesn't detect the target in one tile, but then detects it in the adjacent tile. If the target lies in the overlap region, then combining with a logical \"or\" takes the best (and worst) of both tiles).\n2. **Operating point is important** - Although some are encountering issues with their selected score threshold and OOM errors, I'd like to emphasize that selecting the appropriate score threshold for your respective model matters. See the below Table 1 for an example.\n3. **Subsample the data to iterate experiments quickly** - In combination with prototyping a small model (e.g., efficientnet-b1), you can iterate extremely quickly while making the assumption that findings on the subset and small model *should* generalize to the full train/val sets + larger model.\n____\n\n# Experiment Configuration\n\n**Architecture**: UNet/UNet++ (investigating others)\n**Backbone**: \n- seresnext26d (best results so far)\n- efficientnet-b4\n- efficientnet-b1\n- Investigating other backbones\n\n**Loss**: DiceLoss (Investigating others)\n**Model Selection Metric**: DiceLoss\n**Optimizer**: AdamW (Investigating others)\n**Scheduler**: CosineAnnealingLR\n**Transforms**: \n- Train Transforms = HFlip, VFlip, RRotate90, ElasticTransform, RandomBrightness/Contrast\n- Validation Transforms = None\n\n**Tile Size**:\n- Train: 512x512 w/ 20% overlap\n- Inference: 800x800 w/ 20% overlap (improves performance, needs investigated)\n\n**Epochs**: 20 (no model best checkpoints have occurred after 20 epochs)\n____\n\n# Experiment Results (Ongoing)\n\n### Model Descriptions:\n\n| Experiment Name | Meta Architecture | Encoder | Epochs Trained |\n| --- | --- | --- | --- |\n| B1 |  UNet | efficientnet-b4 | 1 |\n| B2 |  UNet | efficientnet-b4 | 5 |\n| B3 |  UNet++ | efficientnet-b4 | 5 |\n| B4 |  UNet++ | efficientnet-b6 | 8 |\n| B5 |  UNet++ | seresnext26d | 11 |\n\n### Naive random 80:20 w/ fixed seed=42\n\n| Threshold | CV  | LB | Notes |\n| --- | --- | --- | --- |\n| 0.5 |  0.67 | 0.44 | B1 |\n| 0.5 |  0.56 | 0.49 | B2 |\n| 0.5 |  0.56 | 0.54 | B2 + Logical \"OR\" Mask Combination (LOMC) |\n| 0.05 |  **0.58** | **0.57** | B2 + LOMC+ Reduced Score Threshold |\n\n*All models moving forward will use LOMC.\n*Unless otherwise noted, all inference runs use 800x800 tile size\n\n### Train on all, CV on Kidney_3_dense\n\n| Threshold | CV  | LB | Notes |\n| --- | --- | --- | --- |\n| 0.5 |  0.53 | 0.652 | B3 |\n| 0.05 |  0.55 | **0.655** | B3d |\n| 0.05 |  **0.55** | OOM | B3 + 512x512 Tiles |\n| 0.05 |  **0.55** | OOM | B3 + 512x512 Tiles |\n\n### Train on Kidney 1 Dense, CV on Kidney 3 Dense\n\n| Threshold | CV  | LB | Notes |\n| --- | --- | --- | --- |\n| 0.05 |  0.74 | 0.658 | B4 + 384x384 Tiles |\n| 0.05 |  0.74 | 0.663 | B4 + 800x800 Tiles |\n| 0.05 |  0.74 | 0.673 | B4 + 800x800 Tiles + TTA |\n| 0.05 |  **0.89** | 0.72 | B5 + 800x800 Tiles + TTA |\n| 0.01 |  **0.89** | **0.726**| B5 + 800x800 Tiles + TTA |",
    "2534985": "thx for sharing.",
    "2533565": "what is your image dimensions",
    "2535052": "Still in many cases OOM error ...hope it gets solved soon",
    "2546087": "Hi, @squidinator. What's your best model's DiceLoss on validation set? I see you use DiceLoss as model selection function.",
    "2536180": "I added  remove_small ， but still scoring error for forked public baseline."
  }
}