{
  "id": 407972,
  "title": "[lb0.71-one-fold-fragment_id_1 !!!??? ] my experimental results ... the trick of getting good results?",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/407972",
  "author_name": "hengck23",
  "post_date": "2023-05-08T22:42:45.540000",
  "votes": 80,
  "comment_count": 97,
  "views": 0,
  "content": "<p>[url=<a href=\"https://ibb.co/N6Yxp27][img]https://i.ibb.co/Fb76J4n/Selection-999-1974.png[/img][/url\" target=\"_blank\">https://ibb.co/N6Yxp27][img]https://i.ibb.co/Fb76J4n/Selection-999-1974.png[/img][/url</a>]<br>\n<img src=\"https://i.ibb.co/Fb76J4n/Selection-999-1974.png\" alt=\"https://i.ibb.co/Fb76J4n/Selection-999-1974.png\"></p>\n<p>I conducted preliminary experiments:</p>\n<ol>\n<li><p>3dCNN encoder is modified from [1],[1a]. replace Conv3d(stride=2) with Conv3d(stride=1) + AvgPool3d(stride=2). Use gobal pool instead of vectorizing feature volume before feeding into final linear ink classifier.</p></li>\n<li><p>Follow experiment protocols from [2]. divide fragments into upper and lower sub-fragments. validation = one of the sub-fragments. I have the same results as the paper[2], table.1: recall=0.41, FPR=0.051</p></li>\n<li><p>The trick is not to over-train your 3dCNN subvolume encoder. Since it doesn't use context, train for high recall (but high FPR is ok). We will reduce FPR using segmentation (which will has its own encoder and decoder) where we have larger context. Segmentation net is more like a denoiser and superresolution net.</p></li>\n</ol>\n<p>The digram show are real results from my experiments for 3dcnn encoder, which are pretty good and surprises myself. the results include tta.</p>\n<p>[1]  <a href=\"https://github.com/educelab/ink-id\" target=\"_blank\">https://github.com/educelab/ink-id</a>  <br>\n[1a] <a href=\"https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\" target=\"_blank\">https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial</a></p>\n<p>[2] EduceLab-Scrolls: Verifiable Recovery of Text from Herculaneum Papyri using X-ray CT  <br>\n<a href=\"https://arxiv.org/pdf/2304.02084.pdf\" target=\"_blank\">https://arxiv.org/pdf/2304.02084.pdf</a>  </p>\n<hr>\n<p>no sure if this would work, but we can:</p>\n<ol>\n<li>learn a noise generator to approximate results of 3dCNNencoder.</li>\n<li>used synthetic (+real) images to train segmetation. </li>\n</ol>\n<p>then we have infinite train data. we can even train segmention to be alphabet detector, i.e. different class for each alphabet.<br>\nif you undertsand the language in the scrolls, you can filter out invidual words ,etc … </p>\n<hr>\n<p>to be updated ….<br>\nnotebook link:  to be updated ….</p>\n<p>(1) baseline : attentioned pool at encoder (resnet34d) unet, one fold validated on fragement_2a <br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d</a></p>\n<p><img src=\"https://i.ibb.co/HVzz5Rw/Selection-999-2032.png\" alt=\"https://i.ibb.co/HVzz5Rw/Selection-999-2032.png\"><br>\n<a href=\"https://docs.google.com/spreadsheets/d/17t_FFAJNQ7s23NeJ0tHwbBoGJYhqVcUfjAX3-3GNPaE\" target=\"_blank\">https://docs.google.com/spreadsheets/d/17t_FFAJNQ7s23NeJ0tHwbBoGJYhqVcUfjAX3-3GNPaE</a></p>\n<p><img src=\"https://i.ibb.co/1bWxjLp/Selection-999-2048.png\" alt=\"https://i.ibb.co/1bWxjLp/Selection-999-2048.png\"></p>\n<hr>\n<h2>\"We extend our thanks to HP for providing the Z8-G4 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"</h2>",
  "messages": [
    {
      "id": 2250913,
      "postDate": "2023-05-08T22:42:45.540Z",
      "content": "<p>[url=<a href=\"https://ibb.co/N6Yxp27][img]https://i.ibb.co/Fb76J4n/Selection-999-1974.png[/img][/url\" target=\"_blank\">https://ibb.co/N6Yxp27][img]https://i.ibb.co/Fb76J4n/Selection-999-1974.png[/img][/url</a>]<br>\n<img src=\"https://i.ibb.co/Fb76J4n/Selection-999-1974.png\" alt=\"https://i.ibb.co/Fb76J4n/Selection-999-1974.png\"></p>\n<p>I conducted preliminary experiments:</p>\n<ol>\n<li><p>3dCNN encoder is modified from [1],[1a]. replace Conv3d(stride=2) with Conv3d(stride=1) + AvgPool3d(stride=2). Use gobal pool instead of vectorizing feature volume before feeding into final linear ink classifier.</p></li>\n<li><p>Follow experiment protocols from [2]. divide fragments into upper and lower sub-fragments. validation = one of the sub-fragments. I have the same results as the paper[2], table.1: recall=0.41, FPR=0.051</p></li>\n<li><p>The trick is not to over-train your 3dCNN subvolume encoder. Since it doesn't use context, train for high recall (but high FPR is ok). We will reduce FPR using segmentation (which will has its own encoder and decoder) where we have larger context. Segmentation net is more like a denoiser and superresolution net.</p></li>\n</ol>\n<p>The digram show are real results from my experiments for 3dcnn encoder, which are pretty good and surprises myself. the results include tta.</p>\n<p>[1]  <a href=\"https://github.com/educelab/ink-id\" target=\"_blank\">https://github.com/educelab/ink-id</a>  <br>\n[1a] <a href=\"https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\" target=\"_blank\">https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial</a></p>\n<p>[2] EduceLab-Scrolls: Verifiable Recovery of Text from Herculaneum Papyri using X-ray CT  <br>\n<a href=\"https://arxiv.org/pdf/2304.02084.pdf\" target=\"_blank\">https://arxiv.org/pdf/2304.02084.pdf</a>  </p>\n<hr>\n<p>no sure if this would work, but we can:</p>\n<ol>\n<li>learn a noise generator to approximate results of 3dCNNencoder.</li>\n<li>used synthetic (+real) images to train segmetation. </li>\n</ol>\n<p>then we have infinite train data. we can even train segmention to be alphabet detector, i.e. different class for each alphabet.<br>\nif you undertsand the language in the scrolls, you can filter out invidual words ,etc … </p>\n<hr>\n<p>to be updated ….<br>\nnotebook link:  to be updated ….</p>\n<p>(1) baseline : attentioned pool at encoder (resnet34d) unet, one fold validated on fragement_2a <br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d</a></p>\n<p><img src=\"https://i.ibb.co/HVzz5Rw/Selection-999-2032.png\" alt=\"https://i.ibb.co/HVzz5Rw/Selection-999-2032.png\"><br>\n<a href=\"https://docs.google.com/spreadsheets/d/17t_FFAJNQ7s23NeJ0tHwbBoGJYhqVcUfjAX3-3GNPaE\" target=\"_blank\">https://docs.google.com/spreadsheets/d/17t_FFAJNQ7s23NeJ0tHwbBoGJYhqVcUfjAX3-3GNPaE</a></p>\n<p><img src=\"https://i.ibb.co/1bWxjLp/Selection-999-2048.png\" alt=\"https://i.ibb.co/1bWxjLp/Selection-999-2048.png\"></p>\n<hr>\n<h2>\"We extend our thanks to HP for providing the Z8-G4 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"</h2>",
      "rawMarkdown": "[url=https://ibb.co/N6Yxp27][img]https://i.ibb.co/Fb76J4n/Selection-999-1974.png[/img][/url]\n![https://i.ibb.co/Fb76J4n/Selection-999-1974.png](https://i.ibb.co/Fb76J4n/Selection-999-1974.png)\n\n\nI conducted preliminary experiments:\n1. 3dCNN encoder is modified from [1],[1a]. replace Conv3d(stride=2) with Conv3d(stride=1) + AvgPool3d(stride=2). Use gobal pool instead of vectorizing feature volume before feeding into final linear ink classifier.\n\n2. Follow experiment protocols from [2]. divide fragments into upper and lower sub-fragments. validation = one of the sub-fragments. I have the same results as the paper[2], table.1: recall=0.41, FPR=0.051\n\n3. The trick is not to over-train your 3dCNN subvolume encoder. Since it doesn't use context, train for high recall (but high FPR is ok). We will reduce FPR using segmentation (which will has its own encoder and decoder) where we have larger context. Segmentation net is more like a denoiser and superresolution net.\n\nThe digram show are real results from my experiments for 3dcnn encoder, which are pretty good and surprises myself. the results include tta.\n\n\n[1]  https://github.com/educelab/ink-id  \n[1a] https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\n\n[2] EduceLab-Scrolls: Verifiable Recovery of Text from Herculaneum Papyri using X-ray CT  \nhttps://arxiv.org/pdf/2304.02084.pdf  \n\n\n----\nno sure if this would work, but we can:\n1. learn a noise generator to approximate results of 3dCNNencoder.\n2. used synthetic (+real) images to train segmetation. \n\nthen we have infinite train data. we can even train segmention to be alphabet detector, i.e. different class for each alphabet.\nif you undertsand the language in the scrolls, you can filter out invidual words ,etc ... \n\n\n\n----\n\n\nto be updated ....\nnotebook link:  to be updated ....\n\n(1) baseline : attentioned pool at encoder (resnet34d) unet, one fold validated on fragement\\_2a \nhttps://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d\n\n![https://i.ibb.co/HVzz5Rw/Selection-999-2032.png](https://i.ibb.co/HVzz5Rw/Selection-999-2032.png)\nhttps://docs.google.com/spreadsheets/d/17t_FFAJNQ7s23NeJ0tHwbBoGJYhqVcUfjAX3-3GNPaE\n\n![https://i.ibb.co/1bWxjLp/Selection-999-2048.png](https://i.ibb.co/1bWxjLp/Selection-999-2048.png)\n\n---\n\n##\"We extend our thanks to HP for providing the Z8-G4 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"",
      "votes": 80
    },
    {
      "id": 2266159,
      "postDate": "2023-05-19T18:55:05.797Z",
      "content": "<p>how to pool with pos encoding (use conv3d as convolutional positional encoding in z direction)</p>\n<pre><code>        #-- pool attention weight\n        self.weight = nn.ModuleList([\n            nn.Sequential(\n                nn.Conv3d(dim, dim, kernel_size=3, padding=1),\n                nn.ReLU(inplace=True),\n            ) for dim in encoder_dim\n        ])\n\n    def forward(self, batch):\n        v = batch['volume']\n        B,C,H,W = v.shape\n        vv = [\n            v[:,i:i+CFG.crop_depth] for i in [0,2,4,]\n        ]\n        K = len(vv)\n        x = torch.cat(vv,0)\n\n        # ---------------------------------\n        # encoder = self.encoder.forward_features(x)\n        encoder = []\n        x = self.encoder.conv1(x)\n        x = self.encoder.bn1(x)\n        x = self.encoder.act1(x)  ; encoder.append(x)\n        x = F.avg_pool2d(x,kernel_size=2,stride=2)\n        x = self.encoder.layer1(x); encoder.append(x)\n        x = self.encoder.layer2(x); encoder.append(x)\n        x = self.encoder.layer3(x); encoder.append(x)\n        x = self.encoder.layer4(x); encoder.append(x)\n        #print('encoder', [f.shape for f in encoder])\n        # ---------------------------------\n        for i in range(len(encoder)):\n            e = encoder[i]\n            _, c, h, w = e.shape\n            e = rearrange(e, '(K B) c h w -&gt; B c h w K', K=K, B=B, h=h, w=w) #\n            f = self.weight[i](e)\n            w = F.softmax(f, -1)\n            e = (w * e).sum(-1)\n            encoder[i] = e\n</code></pre>",
      "rawMarkdown": "how to pool with pos encoding (use conv3d as convolutional positional encoding in z direction)\n\n```\n\n\t\t#-- pool attention weight\n\t\tself.weight = nn.ModuleList([\n\t\t\tnn.Sequential(\n\t\t\t\tnn.Conv3d(dim, dim, kernel_size=3, padding=1),\n\t\t\t\tnn.ReLU(inplace=True),\n\t\t\t) for dim in encoder_dim\n\t\t])\n\n\tdef forward(self, batch):\n\t\tv = batch['volume']\n\t\tB,C,H,W = v.shape\n\t\tvv = [\n\t\t\tv[:,i:i+CFG.crop_depth] for i in [0,2,4,]\n\t\t]\n\t\tK = len(vv)\n\t\tx = torch.cat(vv,0)\n\n\t\t# ---------------------------------\n\t\t# encoder = self.encoder.forward_features(x)\n\t\tencoder = []\n\t\tx = self.encoder.conv1(x)\n\t\tx = self.encoder.bn1(x)\n\t\tx = self.encoder.act1(x)  ; encoder.append(x)\n\t\tx = F.avg_pool2d(x,kernel_size=2,stride=2)\n\t\tx = self.encoder.layer1(x); encoder.append(x)\n\t\tx = self.encoder.layer2(x); encoder.append(x)\n\t\tx = self.encoder.layer3(x); encoder.append(x)\n\t\tx = self.encoder.layer4(x); encoder.append(x)\n\t\t#print('encoder', [f.shape for f in encoder])\n\t\t# ---------------------------------\n\t\tfor i in range(len(encoder)):\n\t\t\te = encoder[i]\n\t\t\t_, c, h, w = e.shape\n\t\t\te = rearrange(e, '(K B) c h w -> B c h w K', K=K, B=B, h=h, w=w) #\n\t\t\tf = self.weight[i](e)\n\t\t\tw = F.softmax(f, -1)\n\t\t\te = (w * e).sum(-1)\n\t\t\tencoder[i] = e\n\n\n```",
      "votes": 10
    },
    {
      "id": 2272189,
      "postDate": "2023-05-24T10:44:05.173Z",
      "content": "<p>how to get lb 0.68 :</p>\n<ul>\n<li>augmentation: label noise</li>\n<li>model: stacked Unets</li>\n<li>validation: fragement1</li>\n<li>inference 4x rotate TTA</li>\n</ul>\n<p>example notebook: <br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb0-68-one-fold-stacked-unet\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-68-one-fold-stacked-unet</a></p>",
      "rawMarkdown": "how to get lb 0.68 :\n-  augmentation: label noise\n-  model: stacked Unets\n- validation: fragement1\n- inference 4x rotate TTA\n\n\nexample notebook: \nhttps://www.kaggle.com/code/hengck23/lb0-68-one-fold-stacked-unet",
      "votes": 7,
      "replies": [
        {
          "id": 2272273,
          "postDate": "2023-05-24T11:36:00.950Z",
          "content": "<p>architure used for lb0.68. <br>\nresnet34-pool-unet +  resnet10t-unet.<br>\nThis is very fast, only 10 min for 4x rotate TTA</p>\n<p><img src=\"https://i.ibb.co/RTPwsxv/Selection-999-2116.png\" alt=\"https://i.ibb.co/RTPwsxv/Selection-999-2116.png\"><br>\n<img src=\"https://i.ibb.co/8xJzn8R/Selection-999-2115.png\" alt=\"https://i.ibb.co/8xJzn8R/Selection-999-2115.png\"></p>",
          "rawMarkdown": "architure used for lb0.68. \nresnet34-pool-unet +  resnet10t-unet.\nThis is very fast, only 10 min for 4x rotate TTA\n\n![https://i.ibb.co/RTPwsxv/Selection-999-2116.png](https://i.ibb.co/RTPwsxv/Selection-999-2116.png)\n![https://i.ibb.co/8xJzn8R/Selection-999-2115.png](https://i.ibb.co/8xJzn8R/Selection-999-2115.png)",
          "votes": 12,
          "replies": [
            {
              "id": 2272308,
              "postDate": "2023-05-24T12:00:41.143Z",
              "content": "<p>label noise augmentation used:</p>\n<p><img src=\"https://i.ibb.co/p3G5fK0/Selection-999-2120.png\" alt=\"https://i.ibb.co/p3G5fK0/Selection-999-2120.png\"></p>",
              "rawMarkdown": "label noise augmentation used:\n\n![https://i.ibb.co/p3G5fK0/Selection-999-2120.png](https://i.ibb.co/p3G5fK0/Selection-999-2120.png)",
              "votes": 5
            },
            {
              "id": 2272409,
              "postDate": "2023-05-24T13:15:03.823Z",
              "content": "<p><img src=\"https://i.ibb.co/4m5QFHx/Selection-999-2123.png\" alt=\"https://i.ibb.co/4m5QFHx/Selection-999-2123.png\"></p>",
              "rawMarkdown": "![https://i.ibb.co/4m5QFHx/Selection-999-2123.png](https://i.ibb.co/4m5QFHx/Selection-999-2123.png)",
              "votes": 2
            },
            {
              "id": 2272510,
              "postDate": "2023-05-24T14:46:22.777Z",
              "content": "<p>what you can do next step:</p>\n<ul>\n<li>different way to stack … what is the input/output scale of each unet? what is the inter-connecting scale?</li>\n<li>different encoder/decoder for stacking … e.g. VIT, CNN …</li>\n</ul>\n<p>now the inference speed is very fast, you can consider  larger ensemble solution</p>",
              "rawMarkdown": "what you can do next step:\n- different way to stack ... what is the input/output scale of each unet? what is the inter-connecting scale?\n- different encoder/decoder for stacking ... e.g. VIT, CNN ...\n\nnow the inference speed is very fast, you can consider  larger ensemble solution",
              "votes": 2
            },
            {
              "id": 2272598,
              "postDate": "2023-05-24T15:46:01.010Z",
              "content": "<p>why you just rotate label in \"add label noise \" augment？Won't this confusing the model👀</p>",
              "rawMarkdown": "why you just rotate label in \"add label noise \" augment？Won't this confusing the model👀",
              "votes": 4
            },
            {
              "id": 2274380,
              "postDate": "2023-05-25T23:38:05.677Z",
              "content": "<p>can you elaborate more on \"loss1\" please? <br>\nis it <code>BCE(logit1, mask_192x192)</code> ? </p>\n<p>also the overall loss you optimize is: <code>loss = 0.5*loss1 + 0.5*loss2</code> ?</p>",
              "rawMarkdown": "can you elaborate more on \"loss1\" please? \nis it `BCE(logit1, mask_192x192)` ? \n\nalso the overall loss you optimize is: `loss = 0.5*loss1 + 0.5*loss2` ?"
            },
            {
              "id": 2274406,
              "postDate": "2023-05-26T00:57:32.297Z",
              "content": "<p>BCE(logit1, mask_192x192) is correct.</p>\n<p>loss is simply   loss1 + loss2</p>",
              "rawMarkdown": "BCE(logit1, mask_192x192) is correct.\n\nloss is simply   loss1 + loss2",
              "votes": 2
            },
            {
              "id": 2274407,
              "postDate": "2023-05-26T01:01:46.447Z",
              "content": "<p>model is stable:<br>\nfold1 = lb 0.68 /cv0.571<br>\nfold2aa = lb 0.67 /cv0.628<br>\nfold2bb = ? /cv0.647<br>\nfold2cc = ? /cv0.680<br>\nensemble  fold1+fold2aa+fold2bb+fold2cc = lb0.69 (12 min)</p>\n<p>all threshold are 0.5</p>",
              "rawMarkdown": "model is stable:\nfold1 = lb 0.68 /cv0.571\nfold2aa = lb 0.67 /cv0.628\nfold2bb = ? /cv0.647\nfold2cc = ? /cv0.680\nensemble  fold1+fold2aa+fold2bb+fold2cc = lb0.69 (12 min)\n\nall threshold are 0.5",
              "votes": 3
            },
            {
              "id": 2276606,
              "postDate": "2023-05-27T04:03:30.810Z",
              "content": "<p>I am a beginner，I'm trying to train with your model, but the metrics are always the same value, confused</p>",
              "rawMarkdown": "I am a beginner，I'm trying to train with your model, but the metrics are always the same value, confused"
            },
            {
              "id": 2277032,
              "postDate": "2023-05-27T11:36:37.070Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2278271,
              "postDate": "2023-05-28T14:01:59.213Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  I'm not clear about the role of <code>valid_threshold</code>. Could you provide more details about its implementation?</p>\n<p>The following is my implementation, is it similar to yours implementation?</p>\n<pre><code> ():\n    \n    height, width, _ = volume.shape\n     :\n        x = np.random.choice(width - crop_size)  width &gt; crop_size  \n        y = np.random.choice(height - crop_size)  height &gt; crop_size  \n        cropped_volume = volume[y:y+crop_size, x:x+crop_size]\n        cropped_label = label[y:y+crop_size, x:x+crop_size]\n        cropped_mask = mask[y:y+crop_size, x:x+crop_size]\n         np.mean(cropped_mask) &gt; valid_threshold:\n             cropped_volume, cropped_label\n</code></pre>",
              "rawMarkdown": "@hengck23  I'm not clear about the role of `valid_threshold`. Could you provide more details about its implementation?\n\nThe following is my implementation, is it similar to yours implementation?\n```python\ndef do_random_crop(volume, label, mask, crop_size, valid_threshold):\n    \"\"\"Randomly crop the volume and labels.\"\"\"\n    height, width, _ = volume.shape\n    while True:\n        x = np.random.choice(width - crop_size) if width > crop_size else 0\n        y = np.random.choice(height - crop_size) if height > crop_size else 0\n        cropped_volume = volume[y:y+crop_size, x:x+crop_size]\n        cropped_label = label[y:y+crop_size, x:x+crop_size]\n        cropped_mask = mask[y:y+crop_size, x:x+crop_size]\n        if np.mean(cropped_mask) > valid_threshold:\n            return cropped_volume, cropped_label\n```"
            },
            {
              "id": 2278286,
              "postDate": "2023-05-28T14:30:30.427Z",
              "content": "<pre><code>## geometric --------------------------------\ndef do_random_crop(\n    image,\n    mask,\n    valid,\n    crop_size=224,\n    valid_threshold=0.5,\n):\n    height, width = image.shape[:2]\n    while (1):\n        y = np.random.randint(0, height - crop_size+1)\n        x = np.random.randint(0, width  - crop_size+1)\n        crop_image = image[y:y + crop_size, x:x + crop_size]\n        crop_mask  = mask[y:y + crop_size, x:x + crop_size]\n        crop_valid  = valid[y:y + crop_size, x:x + crop_size]\n        if crop_valid.mean() &gt; valid_threshold:\n            break\n    return crop_image, crop_mask\n\n\ndef do_random_affine_crop(\n    image,\n    mask,\n    valid,\n    crop_size=224,\n    scale  = (0.5,2.0),\n    aspect = (0.9,1/0.9),\n    degree = (-45,45),\n    image_fill=0,\n    mask_fill=0,\n    valid_threshold=0.5,\n):\n    height, width = image.shape[:2]\n    s = crop_size\n    point = np.array([\n        [0,0],\n        [0,s],\n        [s,s],\n        [s,0],\n    ])\n    point1 = do_random_affine_on_point(\n        point,\n        scale  = scale,\n        aspect = aspect,\n        degree = degree,\n    )\n    point1 = point1.astype(int)\n    x0, y0 = point1[:,0].min(), point1[:,1].min()\n    point1 = point1 - [[x0, y0]]\n    w, h = point1[:,0].max(), point1[:,1].max()\n\n    while (1):\n        y = np.random.randint(0, height - w+1)\n        x = np.random.randint(0, width  - h+1)\n        crop_valid  = valid[y:y + h, x:x + w]\n        if crop_valid.mean() &gt; valid_threshold:\n            break\n\n    point2 = point1 + [[x, y]]\n    #----\n    matrix = cv2.getAffineTransform(\n        point2[:3].astype(np.float32),\n        point[:3].astype(np.float32),\n    )\n    crop_image = cv2.warpAffine(\n        image, matrix, (crop_size,crop_size),\n        flags=cv2.INTER_LINEAR,borderMode=cv2.BORDER_CONSTANT, borderValue=image_fill,\n    )\n    crop_mask = cv2.warpAffine(\n        mask, matrix, (crop_size,crop_size),\n        flags=cv2.INTER_LINEAR,borderMode=cv2.BORDER_CONSTANT, borderValue=mask_fill,\n    )\n\n\n    return crop_image, crop_mask\n</code></pre>",
              "rawMarkdown": "```\n## geometric --------------------------------\ndef do_random_crop(\n\timage,\n\tmask,\n\tvalid,\n\tcrop_size=224,\n\tvalid_threshold=0.5,\n):\n\theight, width = image.shape[:2]\n\twhile (1):\n\t\ty = np.random.randint(0, height - crop_size+1)\n\t\tx = np.random.randint(0, width  - crop_size+1)\n\t\tcrop_image = image[y:y + crop_size, x:x + crop_size]\n\t\tcrop_mask  = mask[y:y + crop_size, x:x + crop_size]\n\t\tcrop_valid  = valid[y:y + crop_size, x:x + crop_size]\n\t\tif crop_valid.mean() > valid_threshold:\n\t\t\tbreak\n\treturn crop_image, crop_mask\n\n \ndef do_random_affine_crop(\n\timage,\n\tmask,\n\tvalid,\n\tcrop_size=224,\n\tscale  = (0.5,2.0),\n\taspect = (0.9,1/0.9),\n\tdegree = (-45,45),\n\timage_fill=0,\n\tmask_fill=0,\n\tvalid_threshold=0.5,\n):\n\theight, width = image.shape[:2]\n\ts = crop_size\n\tpoint = np.array([\n\t\t[0,0],\n\t\t[0,s],\n\t\t[s,s],\n\t\t[s,0],\n\t])\n\tpoint1 = do_random_affine_on_point(\n\t\tpoint,\n\t\tscale  = scale,\n\t\taspect = aspect,\n\t\tdegree = degree,\n\t)\n\tpoint1 = point1.astype(int)\n\tx0, y0 = point1[:,0].min(), point1[:,1].min()\n\tpoint1 = point1 - [[x0, y0]]\n\tw, h = point1[:,0].max(), point1[:,1].max()\n\n\twhile (1):\n\t\ty = np.random.randint(0, height - w+1)\n\t\tx = np.random.randint(0, width  - h+1)\n\t\tcrop_valid  = valid[y:y + h, x:x + w]\n\t\tif crop_valid.mean() > valid_threshold:\n\t\t\tbreak\n\n\tpoint2 = point1 + [[x, y]]\n\t#----\n\tmatrix = cv2.getAffineTransform(\n\t\tpoint2[:3].astype(np.float32),\n\t\tpoint[:3].astype(np.float32),\n\t)\n\tcrop_image = cv2.warpAffine(\n\t\timage, matrix, (crop_size,crop_size),\n\t\tflags=cv2.INTER_LINEAR,borderMode=cv2.BORDER_CONSTANT, borderValue=image_fill,\n\t)\n\tcrop_mask = cv2.warpAffine(\n\t\tmask, matrix, (crop_size,crop_size),\n\t\tflags=cv2.INTER_LINEAR,borderMode=cv2.BORDER_CONSTANT, borderValue=mask_fill,\n\t)\n\t \n\n\treturn crop_image, crop_mask\n\n```",
              "votes": 1
            },
            {
              "id": 2278297,
              "postDate": "2023-05-28T14:43:57.843Z",
              "content": "<p>May I ask what's the image size before you random cropp it</p>",
              "rawMarkdown": "May I ask what's the image size before you random cropp it"
            },
            {
              "id": 2278400,
              "postDate": "2023-05-28T16:42:27.613Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for your very clear implementation. If I understood correctly, the <code>label</code> in my code is equivalent to the <code>mask</code> in your code. I entered this competition late and I'm currently trying my best to reproduce your experiments and writing a training script. Did you publish any training code that I haven't found? As always, your experimental thread is a wonderful learning material. 👍👍👍</p>",
              "rawMarkdown": "@hengck23 Thanks for your very clear implementation. If I understood correctly, the `label` in my code is equivalent to the `mask` in your code. I entered this competition late and I'm currently trying my best to reproduce your experiments and writing a training script. Did you publish any training code that I haven't found? As always, your experimental thread is a wonderful learning material. 👍👍👍"
            },
            {
              "id": 2278856,
              "postDate": "2023-05-29T02:58:19.520Z",
              "content": "<pre><code>def read_data(fragment_id, z0=Z0, z1=Z1):\n    # volume = np.load(f'{train_dir}/{i}/surface_volume{i}.8.npy')\n    volume = []\n    start_timer = timer()\n    for j in range(z0, z1):\n        v = np.array(Image.open(f'{train_dir}/{fragment_id}/surface_volume/{j:02d}.tif'), dtype=np.uint16)\n        v = (v &gt;&gt; 8).astype(np.uint8)\n        # v = (v / 65535.0 * 255).astype(np.uint8)\n        volume.append(v)\n        print(f'\\r @ read volume{j}  {time_to_str(timer() - start_timer, \"sec\")}', end='', flush=True)\n    print('')\n    volume = np.stack(volume, -1)\n    height, width, depth = volume.shape\n    print(f'fragment_id={fragment_id} volume: {volume.shape}')\n\n    # ---\n    ir = cv2.imread(f'{train_dir}/{fragment_id}/ir.png', cv2.IMREAD_GRAYSCALE)\n    label = cv2.imread(f'{train_dir}/{fragment_id}/inklabels.png', cv2.IMREAD_GRAYSCALE)\n    mask = cv2.imread(f'{train_dir}/{fragment_id}/mask.png', cv2.IMREAD_GRAYSCALE)\n    volume[mask == 0] = 0\n\n    # ----\n    ir = ir / 255\n    label = do_binarise(label)\n    mask = do_binarise(mask)\n    if 0:\n        image_show_norm(f'ir', ir, min=0, max=1, resize=0.1)\n        image_show_norm(f'label', label, min=0, max=1, resize=0.1)\n        cv2.waitKey(0)\n\n    d = dotdict(\n        fragment_id=fragment_id,\n        crop_id=fragment_id,\n        volume=volume,\n        ir=ir,\n        label=label,\n        mask=mask,\n    )\n    return d\n</code></pre>\n<pre><code>volume, label = do_random_affine_crop(\n            d.volume[..., z0:z1],\n            d.label,\n            d.mask,\n            crop_size=crop_size,\n</code></pre>",
              "rawMarkdown": "```\ndef read_data(fragment_id, z0=Z0, z1=Z1):\n\t# volume = np.load(f'{train_dir}/{i}/surface_volume{i}.8.npy')\n\tvolume = []\n\tstart_timer = timer()\n\tfor j in range(z0, z1):\n\t\tv = np.array(Image.open(f'{train_dir}/{fragment_id}/surface_volume/{j:02d}.tif'), dtype=np.uint16)\n\t\tv = (v >> 8).astype(np.uint8)\n\t\t# v = (v / 65535.0 * 255).astype(np.uint8)\n\t\tvolume.append(v)\n\t\tprint(f'\\r @ read volume{j}  {time_to_str(timer() - start_timer, \"sec\")}', end='', flush=True)\n\tprint('')\n\tvolume = np.stack(volume, -1)\n\theight, width, depth = volume.shape\n\tprint(f'fragment_id={fragment_id} volume: {volume.shape}')\n\n\t# ---\n\tir = cv2.imread(f'{train_dir}/{fragment_id}/ir.png', cv2.IMREAD_GRAYSCALE)\n\tlabel = cv2.imread(f'{train_dir}/{fragment_id}/inklabels.png', cv2.IMREAD_GRAYSCALE)\n\tmask = cv2.imread(f'{train_dir}/{fragment_id}/mask.png', cv2.IMREAD_GRAYSCALE)\n\tvolume[mask == 0] = 0\n\n\t# ----\n\tir = ir / 255\n\tlabel = do_binarise(label)\n\tmask = do_binarise(mask)\n\tif 0:\n\t\timage_show_norm(f'ir', ir, min=0, max=1, resize=0.1)\n\t\timage_show_norm(f'label', label, min=0, max=1, resize=0.1)\n\t\tcv2.waitKey(0)\n\n\td = dotdict(\n\t\tfragment_id=fragment_id,\n\t\tcrop_id=fragment_id,\n\t\tvolume=volume,\n\t\tir=ir,\n\t\tlabel=label,\n\t\tmask=mask,\n\t)\n\treturn d\n\n```\n\n\n```\nvolume, label = do_random_affine_crop(\n\t\t\td.volume[..., z0:z1],\n\t\t\td.label,\n\t\t\td.mask,\n\t\t\tcrop_size=crop_size,\n```"
            },
            {
              "id": 2280358,
              "postDate": "2023-05-30T04:37:37.503Z",
              "content": "<p>Why the input is 384, mask is 192, how to get mask？</p>",
              "rawMarkdown": "Why the input is 384, mask is 192, how to get mask？"
            }
          ]
        },
        {
          "id": 2277768,
          "postDate": "2023-05-28T03:52:58.960Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2281715,
          "postDate": "2023-05-31T04:51:30.050Z",
          "content": "<p>Thanks for your great work! I am reading through your code. Do you mind explaining what infer_mask do and whether it makes a significant difference in model performance? Thanks!</p>",
          "rawMarkdown": "Thanks for your great work! I am reading through your code. Do you mind explaining what infer_mask do and whether it makes a significant difference in model performance? Thanks!"
        }
      ]
    },
    {
      "id": 2260220,
      "postDate": "2023-05-15T14:37:17.087Z",
      "content": "<p>Which to use: 3D, 2.5D or pool-2.5D? </p>\n<p>it actually depends how different is the z location of  target signal in train and hidden test data.</p>\n<p>1) Assume that both train,test has signal in about the same z location, then 2.5D is enough. 2d convolution is only invariant to changes in (x,y) location and not  invariant to channel (z). But target has same \"z image characteristics\", so 2d convolution is enough.</p>\n<p>2) Assume train,test has signal in very different z locations. Then you need 3d conv, invariant to changes in (x,y,z) location. But 3d convolution is costly. But this may not be true in the competition because the true resolution of the ink target is low (the inklabels ong images are enlarged version of smaller resolution)</p>\n<hr>\n<p>pool-2.5D is essentially a factorised version of 3d convolution. the pooling is 1d convolution in z direction. attention-pool uses dynamic weight in z direction. if also use fixed weights (via nn parameters) or equal weight(mean pool)</p>",
      "rawMarkdown": "Which to use: 3D, 2.5D or pool-2.5D? \n\nit actually depends how different is the z location of  target signal in train and hidden test data.\n\n1) Assume that both train,test has signal in about the same z location, then 2.5D is enough. 2d convolution is only invariant to changes in (x,y) location and not  invariant to channel (z). But target has same \"z image characteristics\", so 2d convolution is enough.\n\n2) Assume train,test has signal in very different z locations. Then you need 3d conv, invariant to changes in (x,y,z) location. But 3d convolution is costly. But this may not be true in the competition because the true resolution of the ink target is low (the inklabels ong images are enlarged version of smaller resolution)\n\n---\n\npool-2.5D is essentially a factorised version of 3d convolution. the pooling is 1d convolution in z direction. attention-pool uses dynamic weight in z direction. if also use fixed weights (via nn parameters) or equal weight(mean pool)\n\n",
      "votes": 7,
      "replies": [
        {
          "id": 2263103,
          "postDate": "2023-05-17T10:41:18.373Z",
          "content": "<p>Very useful thoughts, thanks!</p>",
          "rawMarkdown": "Very useful thoughts, thanks!"
        }
      ]
    },
    {
      "id": 2258613,
      "postDate": "2023-05-14T10:49:56.153Z",
      "content": "<p>i have a feeling that the test images are rotated</p>",
      "rawMarkdown": "i have a feeling that the test images are rotated",
      "votes": 8,
      "replies": [
        {
          "id": 2258715,
          "postDate": "2023-05-14T12:35:22.230Z",
          "content": "<p>i realised that one can train a rotation classifier to prob the test</p>",
          "rawMarkdown": "i realised that one can train a rotation classifier to prob the test",
          "votes": 1
        },
        {
          "id": 2259804,
          "postDate": "2023-05-15T08:36:25.433Z",
          "content": "<p>hmm I thought the same. </p>\n<p>BTW looking at scores from your excel sheet I see that no.3 <code>rot90(k=2)</code> and no.5 <code>rot(k=1,2,3)</code> are exactly same  </p>",
          "rawMarkdown": "hmm I thought the same. \n\nBTW looking at scores from your excel sheet I see that no.3 `rot90(k=2)` and no.5 `rot(k=1,2,3)` are exactly same  ",
          "votes": 1,
          "replies": [
            {
              "id": 2259838,
              "postDate": "2023-05-15T09:09:36.743Z",
              "content": "<p>thanks, i updated for \"5)TTA: rot(k=0,1,2,3)\"</p>",
              "rawMarkdown": "thanks, i updated for \"5)TTA: rot(k=0,1,2,3)\"",
              "votes": 2
            }
          ]
        },
        {
          "id": 2260144,
          "postDate": "2023-05-15T13:39:41.747Z",
          "content": "<p>I tried looking into this more, TTA with mean of 4 rot improved LB form 53 to 67 for me. <br>\nI tried optimizing TTA on val set (fold 1) but did not see improvements in val score, until I noticed that I only used horizontal and vertical flips in augmentation and no rot90, I added rot90 as train augmentation and suspected that now that this kind of rotation is part of the training data distribution, I would see much higher score on the test set (given that test images are rotated) even without TTA, but I noticed almost no difference in LB to the 53 LB (no TTA). If someone has an idea what might be going on, I would really appreciate your feedback!</p>",
          "rawMarkdown": "I tried looking into this more, TTA with mean of 4 rot improved LB form 53 to 67 for me. \nI tried optimizing TTA on val set (fold 1) but did not see improvements in val score, until I noticed that I only used horizontal and vertical flips in augmentation and no rot90, I added rot90 as train augmentation and suspected that now that this kind of rotation is part of the training data distribution, I would see much higher score on the test set (given that test images are rotated) even without TTA, but I noticed almost no difference in LB to the 53 LB (no TTA). If someone has an idea what might be going on, I would really appreciate your feedback!",
          "replies": [
            {
              "id": 2260243,
              "postDate": "2023-05-15T14:52:37.470Z",
              "content": "<ol>\n<li><p>it depends on the  probability rotated augmentation in training. If there is equal chance that the training input is rotated, then TTA in validation and test will not show any difference in performance. That is also provided that rotated image will not results in label noise (i.e. there are no cases that rotated positive = negative  or rotated negative = positive)</p></li>\n<li><p>at inference, background pixel may be falsely labelled as positive, but it is unlikely that the same perturbed pixel will still falsely lablled as positive. so TTA has an effect of rejecting  false positive (rejecting more  false positive at the expense of rejecting some weak true positive)</p></li>\n</ol>\n<p>in theory you can apply TTA at some selective pixel (instead of all) to get even better results.</p>",
              "rawMarkdown": "1. it depends on the  probability rotated augmentation in training. If there is equal chance that the training input is rotated, then TTA in validation and test will not show any difference in performance. That is also provided that rotated image will not results in label noise (i.e. there are no cases that rotated positive = negative  or rotated negative = positive)\n\n2. at inference, background pixel may be falsely labelled as positive, but it is unlikely that the same perturbed pixel will still falsely lablled as positive. so TTA has an effect of rejecting  false positive (rejecting more  false positive at the expense of rejecting some weak true positive)\n\nin theory you can apply TTA at some selective pixel (instead of all) to get even better results.\n\n\n",
              "votes": 2
            },
            {
              "id": 2260250,
              "postDate": "2023-05-15T14:55:43.723Z",
              "content": "<p>i wonder if anyone is worried that rotation only works for public test and actually degrade resulst for hidden private data?</p>\n<p>what does cross validation show?  the issue for aligning cross validation with lb score is that (i think) the FP rate of test is higher (see the need for higher threshold). it is very hard to draw confident conclusion.</p>",
              "rawMarkdown": "i wonder if anyone is worried that rotation only works for public test and actually degrade resulst for hidden private data?\n\nwhat does cross validation show?  the issue for aligning cross validation with lb score is that (i think) the FP rate of test is higher (see the need for higher threshold). it is very hard to draw confident conclusion.",
              "votes": 2
            },
            {
              "id": 2260311,
              "postDate": "2023-05-15T15:45:06.323Z",
              "content": "<p>hmm, my score is actually not getting better at all with TTA.. very strange</p>",
              "rawMarkdown": "hmm, my score is actually not getting better at all with TTA.. very strange"
            },
            {
              "id": 2261053,
              "postDate": "2023-05-16T04:33:20.273Z",
              "content": "<p>I get no LB improvement from TTA (hv-flips + 90-rots). I think because I trained with albu augs hv-flips and -360+360 range rots.</p>",
              "rawMarkdown": "I get no LB improvement from TTA (hv-flips + 90-rots). I think because I trained with albu augs hv-flips and -360+360 range rots."
            }
          ]
        }
      ]
    },
    {
      "id": 2301888,
      "postDate": "2023-06-14T07:49:13.020Z",
      "content": "<p>\"人人有机会， 个个没把握\" (everyone has a chance but no one is sure to win)</p>\n<p>Be sure to read the data paper. \"Imagine\" how the train and public / private test fragment would look like …</p>\n<p>estimate your predicted ink pixel precision and number of truth pixels carefully.<br>\nis it more prediction  = more fp ?<br>\nor more prdiction = more tp?</p>\n<p>don't be shaken down by wrong threshold.</p>\n<p>best threshold for public  = best threshold for private ???? </p>\n<ul>\n<li><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F83f199f9403710ea8f63bb15eba0bbf4%2FSelection_999(2223).png?generation=1686729121339846&amp;alt=media\" alt=\"\"></li>\n</ul>",
      "rawMarkdown": "\"人人有机会， 个个没把握\" (everyone has a chance but no one is sure to win)\n\nBe sure to read the data paper. \"Imagine\" how the train and public / private test fragment would look like ...\n\nestimate your predicted ink pixel precision and number of truth pixels carefully.\nis it more prediction  = more fp ?\nor more prdiction = more tp?\n\ndon't be shaken down by wrong threshold.\n\nbest threshold for public  = best threshold for private ???? \n\n- ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F83f199f9403710ea8f63bb15eba0bbf4%2FSelection_999(2223).png?generation=1686729121339846&alt=media)\n ",
      "votes": 5,
      "replies": [
        {
          "id": 2302891,
          "postDate": "2023-06-15T00:13:24.947Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd42e66f572884975cc1d23fb36b986ad%2FSelection_999(2227).png?generation=1686787992964081&amp;alt=media\" alt=\"\"></p>\n<p>did you get my hint?</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd42e66f572884975cc1d23fb36b986ad%2FSelection_999(2227).png?generation=1686787992964081&alt=media)\n\ndid you get my hint?",
          "votes": 6,
          "replies": [
            {
              "id": 2302894,
              "postDate": "2023-06-15T00:17:21.137Z",
              "content": "<p>Congrats on your solo gold!!! </p>",
              "rawMarkdown": "Congrats on your solo gold!!! "
            }
          ]
        }
      ]
    },
    {
      "id": 2283935,
      "postDate": "2023-06-01T15:45:07.650Z",
      "content": "<p>did i just find some magic?<br>\nfrom <a href=\"https://www.kaggle.com/hughsando\" target=\"_blank\">@hughsando</a>, <a href=\"https://www.kaggle.com/code/hughsando/visualize-the-3d-fragment-height\" target=\"_blank\">https://www.kaggle.com/code/hughsando/visualize-the-3d-fragment-height</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc2c92845f59f35f8c4de32e5a23de2b6%2FSelection_999(2175).png?generation=1685634284049911&amp;alt=media\" alt=\"\"></p>\n<p>height image and label superimposed</p>",
      "rawMarkdown": "did i just find some magic?\nfrom @hughsando, https://www.kaggle.com/code/hughsando/visualize-the-3d-fragment-height\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc2c92845f59f35f8c4de32e5a23de2b6%2FSelection_999(2175).png?generation=1685634284049911&alt=media)\n\nheight image and label superimposed",
      "votes": 5,
      "replies": [
        {
          "id": 2283936,
          "postDate": "2023-06-01T15:45:49.273Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff5e1e63504811b49ad344a0b58f87baa%2FSelection_999(2174).png?generation=1685634346837324&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff5e1e63504811b49ad344a0b58f87baa%2FSelection_999(2174).png?generation=1685634346837324&alt=media)",
          "votes": 1,
          "replies": [
            {
              "id": 2289546,
              "postDate": "2023-06-06T08:20:52.773Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2284364,
          "postDate": "2023-06-02T01:26:23.860Z",
          "content": "<p>This is very interesting. From the looks of the notebook, it seems as though the threshold is experimentally found. However, if an optimal threshold can be found algorithmically, the segmentation could become much more obvious and simple.</p>",
          "rawMarkdown": "This is very interesting. From the looks of the notebook, it seems as though the threshold is experimentally found. However, if an optimal threshold can be found algorithmically, the segmentation could become much more obvious and simple."
        },
        {
          "id": 2293244,
          "postDate": "2023-06-09T04:00:16.170Z",
          "content": "<p>linking height image and activation CAM heatmap</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1550c0c29ce587631a4aaf4454523dc4%2FSelection_999(2218).png?generation=1686283189708144&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "linking height image and activation CAM heatmap\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1550c0c29ce587631a4aaf4454523dc4%2FSelection_999(2218).png?generation=1686283189708144&alt=media)",
          "votes": 2,
          "replies": [
            {
              "id": 2294333,
              "postDate": "2023-06-10T01:44:16.657Z",
              "content": "<p>Can you expand more on what you are doing here? I am a beginner, so please mind the simple questions. Thank you!</p>",
              "rawMarkdown": "Can you expand more on what you are doing here? I am a beginner, so please mind the simple questions. Thank you!"
            }
          ]
        }
      ]
    },
    {
      "id": 2277577,
      "postDate": "2023-05-27T22:04:44.323Z",
      "content": "<p>the power of context!!!!<br>\n<a href=\"https://www.kaggle.com/fengqilong\" target=\"_blank\">@fengqilong</a> check this! freeze and scale …<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F37dd82554ad4e12377b86900f16cc541%2FSelection_999(2146).png?generation=1685247621602263&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "the power of context!!!!\n@fengqilong check this! freeze and scale ...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F37dd82554ad4e12377b86900f16cc541%2FSelection_999(2146).png?generation=1685247621602263&alt=media)\n \n\n \n",
      "votes": 6,
      "replies": [
        {
          "id": 2277732,
          "postDate": "2023-05-28T02:30:49.050Z",
          "content": "<p>nice improvement! What do you mean by scaling the encoder tho? is it just the mean pooling?</p>",
          "rawMarkdown": "nice improvement! What do you mean by scaling the encoder tho? is it just the mean pooling?",
          "votes": 1
        },
        {
          "id": 2277779,
          "postDate": "2023-05-28T04:20:58.847Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2ba87f0a1d2314657c1a83ca28f09187%2FSelection_999(2139).png?generation=1685247643963019&amp;alt=media\" alt=\"\"><br>\nvisual results</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2ba87f0a1d2314657c1a83ca28f09187%2FSelection_999(2139).png?generation=1685247643963019&alt=media)\nvisual results",
          "votes": 1
        },
        {
          "id": 2278161,
          "postDate": "2023-05-28T12:02:51.093Z",
          "content": "<p>its wonderful!<br>\nbut after average pool，the feature map can‘t decoder to 768x768.<br>\nHow you did it？</p>",
          "rawMarkdown": "its wonderful!\nbut after average pool，the feature map can‘t decoder to 768x768.\nHow you did it？"
        },
        {
          "id": 2284888,
          "postDate": "2023-06-02T10:27:01.537Z",
          "content": "<p>inspiring, I will give it a try.</p>",
          "rawMarkdown": "inspiring, I will give it a try."
        }
      ]
    },
    {
      "id": 2263872,
      "postDate": "2023-05-18T01:47:59.477Z",
      "content": "<p>paper:<br>\nDeciphering Ancient Papyrus Texts: A Machine Learning Approach with Varied Architectures and Encoders<br>\n<a href=\"https://www.researchgate.net/publication/370729553_Deciphering_Ancient_Papyrus_Texts_A_Machine_Learning_Approach_with_Varied_Architectures_and_Encoders\" target=\"_blank\">https://www.researchgate.net/publication/370729553_Deciphering_Ancient_Papyrus_Texts_A_Machine_Learning_Approach_with_Varied_Architectures_and_Encoders</a></p>",
      "rawMarkdown": "paper:\nDeciphering Ancient Papyrus Texts: A Machine Learning Approach with Varied Architectures and Encoders\nhttps://www.researchgate.net/publication/370729553_Deciphering_Ancient_Papyrus_Texts_A_Machine_Learning_Approach_with_Varied_Architectures_and_Encoders",
      "votes": 6
    },
    {
      "id": 2250988,
      "postDate": "2023-05-09T01:07:16.580Z",
      "content": "<p><img src=\"https://i.ibb.co/ZVYqcMQ/Selection-999-1975.png\" alt=\"https://i.ibb.co/ZVYqcMQ/Selection-999-1975.png\"> <img src=\"https://i.ibb.co/mzBHh89/Selection-999-1976.png\" alt=\"https://i.ibb.co/mzBHh89/Selection-999-1976.png\"></p>\n<p>4th fragment . or ?</p>",
      "rawMarkdown": "![https://i.ibb.co/ZVYqcMQ/Selection-999-1975.png](https://i.ibb.co/ZVYqcMQ/Selection-999-1975.png) ![https://i.ibb.co/mzBHh89/Selection-999-1976.png](https://i.ibb.co/mzBHh89/Selection-999-1976.png)\n \n4th fragment . or ?\n",
      "votes": 6,
      "replies": [
        {
          "id": 2251195,
          "postDate": "2023-05-09T06:47:39.400Z",
          "content": "<p>How did you obtain the fourth fragment?</p>",
          "rawMarkdown": "How did you obtain the fourth fragment?",
          "replies": [
            {
              "id": 2251461,
              "postDate": "2023-05-09T11:11:29.597Z",
              "content": "<p>one of youtube presentation videos in the early days by the orangizer<br>\n<a href=\"https://www.youtube.com/watch?v=T0mWqsFrJpk&amp;t=3507s\" target=\"_blank\">https://www.youtube.com/watch?v=T0mWqsFrJpk&amp;t=3507s</a><br>\n<a href=\"https://www.youtube.com/watch?v=YWfIx5ggeYI&amp;t=121s\" target=\"_blank\">https://www.youtube.com/watch?v=YWfIx5ggeYI&amp;t=121s</a></p>",
              "rawMarkdown": "one of youtube presentation videos in the early days by the orangizer\nhttps://www.youtube.com/watch?v=T0mWqsFrJpk&t=3507s\nhttps://www.youtube.com/watch?v=YWfIx5ggeYI&t=121s",
              "votes": 4
            },
            {
              "id": 2251636,
              "postDate": "2023-05-09T14:28:37.667Z",
              "content": "<p>according to the video, some kaggle train fragment are actually overlapping two fragments?</p>",
              "rawMarkdown": "according to the video, some kaggle train fragment are actually overlapping two fragments?",
              "votes": 2
            }
          ]
        },
        {
          "id": 2251983,
          "postDate": "2023-05-09T19:41:56.040Z",
          "content": "<p>Interesting find! We obviously can't talk about the origins of the secret fragment. However, this is a good opportunity to remind everyone of the rule prohibiting use of external Herculaneum papyri sources, such as this fragment, or others that might be floating around. Submissions must be a legitimate ink-detection method reproducible by the contest organizers, including training.</p>",
          "rawMarkdown": "Interesting find! We obviously can't talk about the origins of the secret fragment. However, this is a good opportunity to remind everyone of the rule prohibiting use of external Herculaneum papyri sources, such as this fragment, or others that might be floating around. Submissions must be a legitimate ink-detection method reproducible by the contest organizers, including training.",
          "votes": 5,
          "replies": [
            {
              "id": 2252018,
              "postDate": "2023-05-09T20:10:24.683Z",
              "content": "<p>🙂 I was saying to my wife Sally last night, after reading this thread: what would be the odds that I enter my second Kaggle comp in 11 years, and <em><a href=\"https://www.kaggle.com/c/FacebookRecruiting/discussion/2082#11936\" target=\"_blank\">it gets blown up again</a>?</em> 🙂</p>\n<p>In this case, if this <em>is</em> the hidden fragment (and I can see why it would be sensible to slice it horizontally, like the dummy test fragments supplied, so … sadness 😔), then some entries on the leaderboards (private and public) could potentially be meaningless, but as <a href=\"https://www.kaggle.com/jpposma\" target=\"_blank\">@jpposma</a> notes, it's not going to affect the integrity of the final judgments by the organizers, which will discard those submissions.</p>",
              "rawMarkdown": "🙂 I was saying to my wife Sally last night, after reading this thread: what would be the odds that I enter my second Kaggle comp in 11 years, and *[it gets blown up again](https://www.kaggle.com/c/FacebookRecruiting/discussion/2082#11936)?* 🙂\n\nIn this case, if this *is* the hidden fragment (and I can see why it would be sensible to slice it horizontally, like the dummy test fragments supplied, so … sadness 😔), then some entries on the leaderboards (private and public) could potentially be meaningless, but as @jpposma notes, it's not going to affect the integrity of the final judgments by the organizers, which will discard those submissions.",
              "votes": 1
            },
            {
              "id": 2252025,
              "postDate": "2023-05-09T20:22:59.117Z",
              "content": "<p>The hidden fragment is in fragment3. (I think the scrolls get folded in)<br>\nThis could be why some kagglers reported that if you use too many slices results are not as good?</p>\n<p>(hence the model should be detecting the ink at the surface volume, rather than ink in the volume.)</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/jfQs9JV/Selection-999-1981.png\" alt=\"Selection-999-1981\"></a></p>\n<hr>\n<p>fragment4 is a new fragment not in kaggle train dataset. <br>\nIt shows how the writings can vary and look different from the training set.</p>\n<p>(if you use too much context, it may overfit and not generalisd)</p>",
              "rawMarkdown": "The hidden fragment is in fragment3. (I think the scrolls get folded in)\nThis could be why some kagglers reported that if you use too many slices results are not as good?\n\n(hence the model should be detecting the ink at the surface volume, rather than ink in the volume.)\n\n<a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/jfQs9JV/Selection-999-1981.png\" alt=\"Selection-999-1981\" border=\"0\"></a>\n\n---\n\nfragment4 is a new fragment not in kaggle train dataset. \nIt shows how the writings can vary and look different from the training set.\n\n(if you use too much context, it may overfit and not generalisd)\n\n",
              "votes": 2
            },
            {
              "id": 2252044,
              "postDate": "2023-05-09T20:45:20.853Z",
              "content": "<p>Sorry, yes, <code>s/hidden/fourth/</code>.</p>",
              "rawMarkdown": "Sorry, yes, `s/hidden/fourth/`."
            },
            {
              "id": 2258958,
              "postDate": "2023-05-14T15:09:35.207Z",
              "content": "<p><a href=\"https://www.kaggle.com/jpposma\" target=\"_blank\">@jpposma</a> </p>\n<p>refer to the dicussion at <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/395233#2184750\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/395233#2184750</a></p>\n<p>you mentioned \"You are allowed to use any data from <a href=\"https://scrollprize.org/data\" target=\"_blank\">https://scrollprize.org/data</a> in your training\"</p>\n<hr>\n<p>Now I am starting experiment on self-supervised learning unlabelled data.<br>\ni want to clarify again both data below can be used:</p>\n<ul>\n<li>\"/full-scrolls\"  (non-kaggle,  Grand Prize data )</li>\n<li>\"/fragments\"  (kaggle data)</li>\n</ul>\n<p>Note that i may may recover more hidden unlabelled fragements from the raw \"/fragments\" if i re-segment or segment deeper (instead of the surface data for kaggle) the volume myself.</p>\n<p>Thanks!</p>\n<p><a href=\"https://scrollprize.org/data_organization\" target=\"_blank\">https://scrollprize.org/data_organization</a><br>\n<a href=\"https://scrollprize.org/data\" target=\"_blank\">https://scrollprize.org/data</a></p>",
              "rawMarkdown": "@jpposma \n\nrefer to the dicussion at https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/395233#2184750\n\nyou mentioned \"You are allowed to use any data from https://scrollprize.org/data in your training\"\n\n---\n\nNow I am starting experiment on self-supervised learning unlabelled data.\ni want to clarify again both data below can be used:\n- \"/full-scrolls\"  (non-kaggle,  Grand Prize data )\n- \"/fragments\"  (kaggle data)\n\nNote that i may may recover more hidden unlabelled fragements from the raw \"/fragments\" if i re-segment or segment deeper (instead of the surface data for kaggle) the volume myself.\n\nThanks!\n\nhttps://scrollprize.org/data_organization\nhttps://scrollprize.org/data\n",
              "votes": 1
            },
            {
              "id": 2260563,
              "postDate": "2023-05-15T18:11:27.613Z",
              "content": "<p>Yes, both \"/full-scrolls\" and \"/fragments\" can be used. Just not any external data pertaining to the Herculaneum scrolls, like the Youtube videos you found.</p>",
              "rawMarkdown": "Yes, both \"/full-scrolls\" and \"/fragments\" can be used. Just not any external data pertaining to the Herculaneum scrolls, like the Youtube videos you found.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2265521,
      "postDate": "2023-05-19T09:59:09.877Z",
      "content": "<p>very important details of the hidden test data:<br>\n[1] EduceLab-Scrolls: Verifiable Recovery of Text from Herculaneum Papyri using X-ray CT<br>\n<a href=\"https://arxiv.org/pdf/2304.02084.pdf\" target=\"_blank\">https://arxiv.org/pdf/2304.02084.pdf</a></p>\n<p>Figure 1. </p>\n<p>\" Fragments 1 and 2 are from a philosophical work authored by Philodemus, as<br>\nare many others in the Herculaneum collection. This book, titled On slander, comes from his greater work On Vices and the opposite virtues.\"</p>\n<p>\"Fragments 3 and 4 come from a scroll about Hellenistic Dynastic history\"</p>\n<hr>\n<p>\"hidden, subsurface layers of the scroll fragments for which there is no ground truth.\"</p>\n<hr>\n<p>\"in which a trained model is applied to hidden layers extracted from the X-ray CT under the exposed surface. In this case, there is no risk of memorization, so a model is trained on the entire training dataset, or all four fragment surfaces.\"</p>\n<hr>\n<p>Table 2.</p>\n<pre><code>Fragment Characters Recall FPR\n4 25 0.64 0.12\n</code></pre>\n<p>\"Importantly, the misidentified characters are not invented “whole cloth.” They are at least identified in the correct location, and could be accurately identified with marginal improvements in ink detection\"</p>\n<hr>\n<p>Table 3.</p>\n<pre><code>Size (GB) Surface volume (GB)\nFragment 4 198 7.3\n</code></pre>\n<p>Use GB size to estimate image size</p>",
      "rawMarkdown": "very important details of the hidden test data:\n[1] EduceLab-Scrolls: Verifiable Recovery of Text from Herculaneum Papyri using X-ray CT\nhttps://arxiv.org/pdf/2304.02084.pdf\n\nFigure 1. \n\n\n\" Fragments 1 and 2 are from a philosophical work authored by Philodemus, as\nare many others in the Herculaneum collection. This book, titled On slander, comes from his greater work On Vices and the opposite virtues.\"\n\n\"Fragments 3 and 4 come from a scroll about Hellenistic Dynastic history\"\n\n---\n\"hidden, subsurface layers of the scroll fragments for which there is no ground truth.\"\n\n\n---\n\n\"in which a trained model is applied to hidden layers extracted from the X-ray CT under the exposed surface. In this case, there is no risk of memorization, so a model is trained on the entire training dataset, or all four fragment surfaces.\"\n \n---\n\nTable 2.\n```\nFragment Characters Recall FPR\n4 25 0.64 0.12\n```\n\"Importantly, the misidentified characters are not invented “whole cloth.” They are at least identified in the correct location, and could be accurately identified with marginal improvements in ink detection\"\n\n---\nTable 3.\n\n```\n\nSize (GB) Surface volume (GB)\nFragment 4 198 7.3\n```\n\nUse GB size to estimate image size",
      "votes": 3
    },
    {
      "id": 2265133,
      "postDate": "2023-05-19T02:30:02.850Z",
      "content": "<p>lb0.65 recipe </p>\n<p>encoder = resnext26d<br>\nvalidation = fragement 1<br>\nuse pool resnet unet </p>\n<p>(same configure is 0.54 for fragement 2b validation)</p>",
      "rawMarkdown": "lb0.65 recipe \n\nencoder = resnext26d\nvalidation = fragement 1\nuse pool resnet unet \n\n(same configure is 0.54 for fragement 2b validation)",
      "votes": 4,
      "replies": [
        {
          "id": 2265192,
          "postDate": "2023-05-19T04:11:54.037Z",
          "content": "<pre><code>        conv_dim = 64\n        encoder_dim  = [conv_dim, 256, 512, 1024, 2048]#64, 128, 256, 512, ] #[\n        decoder_dim  = [256, 128, 64, 32, 16]\n        self.encoder = seresnext26t_32x4d(pretrained=True, in_chans=CFG.crop_depth)\n</code></pre>\n<pre><code>** start training here! **\n   batch_size = 64 \n   experiment = ['004_pool_resnet26_224_unet_00', 'run_train_fold1.py']\n                           |----------------------- VALID---------------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   thr    recall  fpr   p_sum  score   | loss                 | time           \n----------------------------------------------------------------------------------------------------------------\n0.00e+0   00000000*   0.00 | 0.663  0.100  1.000  1.000  1.777  0.126  | 0.000  0.000  0.000  |  0 hr 00 min\n1.00e-3   00000216*   1.00 | 0.526  0.400  0.973  0.431  0.866  0.244  | 0.581  0.000  0.000  |  0 hr 08 min\n1.00e-3   00000432*   2.00 | 0.449  0.500  0.298  0.034  0.109  0.442  | 0.525  0.000  0.000  |  0 hr 15 min\n1.00e-3   00000648*   3.00 | 0.387  0.500  0.234  0.023  0.080  0.425  | 0.461  0.000  0.000  |  0 hr 22 min\n1.00e-3   00000864*   4.00 | 0.314  0.400  0.317  0.025  0.098  0.505  | 0.404  0.000  0.000  |  0 hr 29 min\n1.00e-3   00001080*   5.00 | 0.292  0.400  0.335  0.023  0.098  0.533  | 0.361  0.000  0.000  |  0 hr 36 min\n1.00e-3   00001296*   6.00 | 0.260  0.500  0.284  0.018  0.080  0.517  | 0.328  0.000  0.000  |  0 hr 43 min\n1.00e-3   00001512*   7.00 | 0.242  0.500  0.289  0.015  0.077  0.537  | 0.312  0.000  0.000  |  0 hr 50 min\n1.00e-3   00001728*   8.00 | 0.243  0.600  0.388  0.030  0.119  0.539  | 0.290  0.000  0.000  |  0 hr 57 min\n1.00e-3   00001944*   9.00 | 0.270  0.400  0.301  0.022  0.091  0.505  | 0.275  0.000  0.000  |  1 hr 04 min\n</code></pre>\n<pre><code>model_path=/home/titanx/hengck/share1/kaggle/2022/ink-detect/result/run004/pool_resnet26_224_unet_00/fold-1/checkpoint/00001728.model.pth\nfragment_id=1\nCFG.stride=56\n\nbce=0.24270\np_sum  th   prec   recall   fpr   dice   score\n----------------------------------------------\n0.33, 0.10, 0.374, 0.683, 0.131,  0.483,  0.411\n0.15, 0.20, 0.547, 0.445, 0.042,  0.491,  0.523\n0.09, 0.30, 0.647, 0.327, 0.021,  0.434,  0.541\n0.06, 0.40, 0.730, 0.243, 0.010,  0.364,  0.521\n0.04, 0.50, 0.792, 0.178, 0.005,  0.291,  0.468\n0.03, 0.60, 0.854, 0.128, 0.003,  0.223,  0.401\n0.02, 0.70, 0.906, 0.081, 0.001,  0.148,  0.297\n0.01, 0.80, 0.947, 0.035, 0.000,  0.068,  0.153\n0.00, 0.90, 0.992, 0.002, 0.000,  0.004,  0.009\n</code></pre>",
          "rawMarkdown": "```\n\t\tconv_dim = 64\n\t\tencoder_dim  = [conv_dim, 256, 512, 1024, 2048]#64, 128, 256, 512, ] #[\n\t\tdecoder_dim  = [256, 128, 64, 32, 16]\n\t\tself.encoder = seresnext26t_32x4d(pretrained=True, in_chans=CFG.crop_depth)\n```\n\n```\n** start training here! **\n   batch_size = 64 \n   experiment = ['004_pool_resnet26_224_unet_00', 'run_train_fold1.py']\n                           |----------------------- VALID---------------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   thr    recall  fpr   p_sum  score   | loss                 | time           \n----------------------------------------------------------------------------------------------------------------\n0.00e+0   00000000*   0.00 | 0.663  0.100  1.000  1.000  1.777  0.126  | 0.000  0.000  0.000  |  0 hr 00 min\n1.00e-3   00000216*   1.00 | 0.526  0.400  0.973  0.431  0.866  0.244  | 0.581  0.000  0.000  |  0 hr 08 min\n1.00e-3   00000432*   2.00 | 0.449  0.500  0.298  0.034  0.109  0.442  | 0.525  0.000  0.000  |  0 hr 15 min\n1.00e-3   00000648*   3.00 | 0.387  0.500  0.234  0.023  0.080  0.425  | 0.461  0.000  0.000  |  0 hr 22 min\n1.00e-3   00000864*   4.00 | 0.314  0.400  0.317  0.025  0.098  0.505  | 0.404  0.000  0.000  |  0 hr 29 min\n1.00e-3   00001080*   5.00 | 0.292  0.400  0.335  0.023  0.098  0.533  | 0.361  0.000  0.000  |  0 hr 36 min\n1.00e-3   00001296*   6.00 | 0.260  0.500  0.284  0.018  0.080  0.517  | 0.328  0.000  0.000  |  0 hr 43 min\n1.00e-3   00001512*   7.00 | 0.242  0.500  0.289  0.015  0.077  0.537  | 0.312  0.000  0.000  |  0 hr 50 min\n1.00e-3   00001728*   8.00 | 0.243  0.600  0.388  0.030  0.119  0.539  | 0.290  0.000  0.000  |  0 hr 57 min\n1.00e-3   00001944*   9.00 | 0.270  0.400  0.301  0.022  0.091  0.505  | 0.275  0.000  0.000  |  1 hr 04 min\n\n```\n\n\n```\nmodel_path=/home/titanx/hengck/share1/kaggle/2022/ink-detect/result/run004/pool_resnet26_224_unet_00/fold-1/checkpoint/00001728.model.pth\nfragment_id=1\nCFG.stride=56\n\nbce=0.24270\np_sum  th   prec   recall   fpr   dice   score\n----------------------------------------------\n0.33, 0.10, 0.374, 0.683, 0.131,  0.483,  0.411\n0.15, 0.20, 0.547, 0.445, 0.042,  0.491,  0.523\n0.09, 0.30, 0.647, 0.327, 0.021,  0.434,  0.541\n0.06, 0.40, 0.730, 0.243, 0.010,  0.364,  0.521\n0.04, 0.50, 0.792, 0.178, 0.005,  0.291,  0.468\n0.03, 0.60, 0.854, 0.128, 0.003,  0.223,  0.401\n0.02, 0.70, 0.906, 0.081, 0.001,  0.148,  0.297\n0.01, 0.80, 0.947, 0.035, 0.000,  0.068,  0.153\n0.00, 0.90, 0.992, 0.002, 0.000,  0.004,  0.009\n\n```\n",
          "votes": 4,
          "replies": [
            {
              "id": 2280517,
              "postDate": "2023-05-30T06:37:36.690Z",
              "content": "<p>It seems that there are only three training samples here. How to use the batch size of 64. Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
              "rawMarkdown": "It seems that there are only three training samples here. How to use the batch size of 64. Thanks @hengck23 "
            }
          ]
        }
      ]
    },
    {
      "id": 2256006,
      "postDate": "2023-05-12T07:37:24.237Z",
      "content": "<p>how to select the best depth crop:</p>\n<p><a href=\"https://ibb.co/RzwbWtb\"><img src=\"https://i.ibb.co/wMmg18g/Selection-999-2027.png\" alt=\"Selection-999-2027\"></a></p>\n<p>as reference:<br>\nvesuvius_2d_slide_exp002/vesuvius-models/Unet_fold1_best.pth<br>\nfragment 1 validation 0.572</p>",
      "rawMarkdown": "how to select the best depth crop:\n\n<a href=\"https://ibb.co/RzwbWtb\"><img src=\"https://i.ibb.co/wMmg18g/Selection-999-2027.png\" alt=\"Selection-999-2027\" border=\"0\"></a>\n\nas reference:\nvesuvius_2d_slide_exp002/vesuvius-models/Unet_fold1_best.pth\nfragment 1 validation 0.572",
      "votes": 4,
      "replies": [
        {
          "id": 2256016,
          "postDate": "2023-05-12T07:46:26.640Z",
          "content": "<p>you can treat it as multiple instance segmentation (for each pixel).<br>\ndifferent pixel location uses different depth (assume the surface is not flat)</p>\n<p>you can pool:<br>\n1) at the end  of each scale of encoder <br>\n2) at the end of last scale decoder,<br>\netc …</p>",
          "rawMarkdown": "you can treat it as multiple instance segmentation (for each pixel).\ndifferent pixel location uses different depth (assume the surface is not flat)\n\nyou can pool:\n1) at the end  of each scale of encoder \n2) at the end of last scale decoder,\netc ...",
          "votes": 1,
          "replies": [
            {
              "id": 2258379,
              "postDate": "2023-05-14T06:44:51.330Z",
              "content": "<p>example notebook at: <a href=\"https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d</a></p>",
              "rawMarkdown": "example notebook at: https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d",
              "votes": 2
            },
            {
              "id": 2259590,
              "postDate": "2023-05-15T05:38:01.613Z",
              "content": "<p>one fold: fragment2a: LB 0.56 , threshold=0.5<br>\ntwo fold: fragment2a,2b: LB 0.59, threshold=0.5</p>\n<hr>\n<p>LB result (falls to 0.4x range) not that good if i add positional encoding in z-axis. maybe the test data are different?</p>\n<p>but CV result results improve by 0.01 for the same positional encoding</p>",
              "rawMarkdown": "one fold: fragment2a: LB 0.56 , threshold=0.5\ntwo fold: fragment2a,2b: LB 0.59, threshold=0.5\n\n---\nLB result (falls to 0.4x range) not that good if i add positional encoding in z-axis. maybe the test data are different?\n\nbut CV result results improve by 0.01 for the same positional encoding",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2287963,
      "postDate": "2023-06-05T04:31:21.147Z",
      "content": "<p>better threshold method?</p>\n<ul>\n<li><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F14c90169c92f5db5e2c02a2b23ccc237%2FSelection_999(2194).png?generation=1685967907793639&amp;alt=media\" alt=\"\"></li>\n</ul>",
      "rawMarkdown": "better threshold method?\n- ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F14c90169c92f5db5e2c02a2b23ccc237%2FSelection_999(2194).png?generation=1685967907793639&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 2288497,
          "postDate": "2023-06-05T12:26:49.290Z",
          "content": "<p>it is local contrast equalization</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F448f98a9d7ca6ef3c45417f05bc01071%2FSelection_999(2195).png?generation=1685967973387298&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "it is local contrast equalization\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F448f98a9d7ca6ef3c45417f05bc01071%2FSelection_999(2195).png?generation=1685967973387298&alt=media)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2281576,
      "postDate": "2023-05-31T01:55:25.057Z",
      "content": "<p>surprise!!!!!<br>\ni am surprise that i can train will crop depth of 32</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb3ae8a9d4cb24ac90713e416253ee775%2FSelection_999(2170).png?generation=1685498117144286&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "surprise!!!!!\ni am surprise that i can train will crop depth of 32\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb3ae8a9d4cb24ac90713e416253ee775%2FSelection_999(2170).png?generation=1685498117144286&alt=media)\n",
      "votes": 1,
      "replies": [
        {
          "id": 2286191,
          "postDate": "2023-06-03T09:40:43.697Z",
          "content": "<p>simplified using mean pooling:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/meanpool-resnet34d-unet\" target=\"_blank\">https://www.kaggle.com/code/hengck23/meanpool-resnet34d-unet</a></p>\n<pre><code></code></pre>",
          "rawMarkdown": "simplified using mean pooling:\nhttps://www.kaggle.com/code/hengck23/meanpool-resnet34d-unet\n\n```\n\nclass Config(object):\n\tvalid_threshold = 0.80\n\tbeta = 1\n\tcrop_fade  = 32\n\tcrop_size  = 128 #256 \n\tcrop_depth = 5\n\tinfer_fragment_z = [\n\t\t32-16,\n\t\t32+16,\n\t]#32 slices\n\tdz = 0\nCFG1 = Config()\n\n\nclass Net(nn.Module):\n\tdef __init__(self, ):\n\t\tsuper().__init__()\n\t\tself.output_type = ['inference', 'loss']\n\n\t\t# --------------------------------\n\t\tCFG = CFG1\n\t\tself.crop_depth = CFG.crop_depth\n\n\t\tconv_dim = 64\n\t\tencoder_dim = [conv_dim, 64, 128, 256, 512, ]\n\t\tdecoder_dim = [256, 128, 64, 32, 16]\n\n\t\tself.encoder = resnet34d(pretrained=True, in_chans=self.crop_depth)\n\n\t\tself.decoder = SmpUnetDecoder(\n\t\t\tin_channel=encoder_dim[-1],\n\t\t\tskip_channel=encoder_dim[:-1][::-1] + [0],\n\t\t\tout_channel=decoder_dim,\n\t\t)\n\t\tself.logit = nn.Conv2d(decoder_dim[-1], 1, kernel_size=1)\n\n\t\t# --------------------------------\n\t\tself.aux = nn.ModuleList([\n\t\t\tnn.Conv2d(encoder_dim[i], 1, kernel_size=1, padding=0) for i in range(len(encoder_dim))\n\t\t])\n\n\n\tdef forward(self, batch):\n\t\tv = batch['volume']\n\t\tB, C, H, W = v.shape\n\t\tvv = [\n\t\t\tv[:, i:i + self.crop_depth] for i in range(0,C-self.crop_depth+1,2)\n\t\t]\n\t\tK = len(vv)\n\t\tx = torch.cat(vv, 0)\n\n\t\t# ---------------------------------\n\n\t\tencoder = []\n\t\te = self.encoder\n\n\t\tx = e.conv1(x)\n\t\tx = e.bn1(x)\n\t\tx = e.act1(x); encoder.append(x)\n\t\tx = F.avg_pool2d(x, kernel_size=2, stride=2)\n\t\tx = e.layer1(x); encoder.append(x)\n\t\tx = e.layer2(x); encoder.append(x)\n\t\tx = e.layer3(x); encoder.append(x)\n\t\tx = e.layer4(x); encoder.append(x)\n\t\t##[print('encoder',i,f.shape) for i,f in enumerate(encoder)]\n\n\t\tfor i in range(len(encoder)):\n\t\t\te = encoder[i]\n\t\t\t_, c, h, w = e.shape\n\t\t\te = rearrange(e, '(K B) c h w -> K B c h w', K=K, B=B, h=h, w=w)\n\t\t\tencoder[i] = e.mean(0)\n\n\t\tlast, decoder = self.decoder(feature = encoder[-1], skip = encoder[:-1][::-1]  + [None])\n\n\n\t\t# ---------------------------------\n\t\tlogit = self.logit(last)\n\n\t\toutput = {}\n\t\tif 1:\n\t\t\tif logit.shape[2:]!=(H, W):\n\t\t\t\tlogit = F.interpolate(logit, size=(H, W), mode='bilinear', align_corners=False, antialias=True)\n\t\t\toutput['ink'] = torch.sigmoid(logit)\n\n\t\treturn output\n\n```",
          "votes": 4
        }
      ]
    },
    {
      "id": 2265586,
      "postDate": "2023-05-19T11:10:34.710Z",
      "content": "<p><a href=\"https://stackoverflow.com/questions/71690251/binary-cross-entropy-with-logits-weight-vs-pos-weight-what-are-the-differences\" target=\"_blank\">https://stackoverflow.com/questions/71690251/binary-cross-entropy-with-logits-weight-vs-pos-weight-what-are-the-differences</a></p>\n<p>The pos_weight parameter allows you to balance the positive example thus controlling the tradeoff between recall and precision (see also). A detailed explanation can be found on this thread along with the explicit math expression. </p>",
      "rawMarkdown": "https://stackoverflow.com/questions/71690251/binary-cross-entropy-with-logits-weight-vs-pos-weight-what-are-the-differences\n\nThe pos_weight parameter allows you to balance the positive example thus controlling the tradeoff between recall and precision (see also). A detailed explanation can be found on this thread along with the explicit math expression. \n",
      "votes": 1
    },
    {
      "id": 2265153,
      "postDate": "2023-05-19T02:58:09.640Z",
      "content": "<p>it is probably difficult to augment positive samples …<br>\nbut it is easy to augment negative samples !!!!</p>\n<p>since fbeta favours high precision, if you can create \"infinite negative samples\" , then ….<br>\nfurther, there are so many samples in the external dataset at the scrollprize website</p>",
      "rawMarkdown": "it is probably difficult to augment positive samples ...\nbut it is easy to augment negative samples !!!!\n\nsince fbeta favours high precision, if you can create \"infinite negative samples\" , then ....\nfurther, there are so many samples in the external dataset at the scrollprize website",
      "votes": 1
    },
    {
      "id": 2260347,
      "postDate": "2023-05-15T16:02:06.260Z",
      "content": "<p>how not to shake up:<br>\nsince there is only two hidden test image (a and b), it is easy to \"get\" \"these information\".</p>\n<p>what you need is not really CV/LB alignment. What you need is:<br>\n1) number of mask pixels in a and b (assume larger mask is private).<br>\n2) for each of a and b, triplet : (LB score, threshold used, number of detected pixels)</p>\n<p>(2) is a very good indicative of fpr, recall and precision values</p>",
      "rawMarkdown": "how not to shake up:\nsince there is only two hidden test image (a and b), it is easy to \"get\" \"these information\".\n\nwhat you need is not really CV/LB alignment. What you need is:\n1) number of mask pixels in a and b (assume larger mask is private).\n2) for each of a and b, triplet : (LB score, threshold used, number of detected pixels)\n\n(2) is a very good indicative of fpr, recall and precision values\n",
      "votes": 1
    },
    {
      "id": 2258434,
      "postDate": "2023-05-14T07:50:24.593Z",
      "content": "<p>if we treat frame id 1,2,3 as class 1,2,3, we can train a fragement classifier.<br>\nthen we can see if the fragment are the same or not from the confusion matrix.</p>\n<p>we can use these fragment classfiier to probe if the public test and private test are smiliar to fragment1,2,3</p>",
      "rawMarkdown": "if we treat frame id 1,2,3 as class 1,2,3, we can train a fragement classifier.\nthen we can see if the fragment are the same or not from the confusion matrix.\n\nwe can use these fragment classfiier to probe if the public test and private test are smiliar to fragment1,2,3",
      "votes": 1
    },
    {
      "id": 2254508,
      "postDate": "2023-05-11T03:53:58.797Z",
      "content": "<p>so interesting find, i was curious abt it</p>",
      "rawMarkdown": "so interesting find, i was curious abt it",
      "votes": 1
    },
    {
      "id": 2273044,
      "postDate": "2023-05-25T01:27:08.390Z",
      "content": "<p>Another approach to train a denoiser<br>\n<a href=\"url\" target=\"_blank\">https://arxiv.org/abs/2205.11423</a></p>",
      "rawMarkdown": "Another approach to train a denoiser\n[https://arxiv.org/abs/2205.11423](url)",
      "votes": 2,
      "replies": [
        {
          "id": 2273203,
          "postDate": "2023-05-25T04:31:16.873Z",
          "content": "<p>thanks, since we only have one class, step #1 of figure.1 in the paper can be modified to predict  e.g. number of ink pixels, patch (32x32) classifiers, etc</p>",
          "rawMarkdown": "thanks, since we only have one class, step #1 of figure.1 in the paper can be modified to predict  e.g. number of ink pixels, patch (32x32) classifiers, etc",
          "votes": 1,
          "replies": [
            {
              "id": 2273226,
              "postDate": "2023-05-25T04:46:08.517Z",
              "content": "<p>what I did was pretraining the 3d encoder using an architecture similar to the ink-id inkclassifier, it worked well.</p>",
              "rawMarkdown": "what I did was pretraining the 3d encoder using an architecture similar to the ink-id inkclassifier, it worked well.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2262130,
      "postDate": "2023-05-16T18:43:14.170Z",
      "content": "<p>from the table of experimental results:</p>\n<ol>\n<li>VIT like pvt-v2 is much better encoder than resnet</li>\n<li>pool as 2.5+1d  improved results </li>\n</ol>",
      "rawMarkdown": "from the table of experimental results:\n1. VIT like pvt-v2 is much better encoder than resnet\n2. pool as 2.5+1d  improved results ",
      "votes": 2,
      "replies": [
        {
          "id": 2272044,
          "postDate": "2023-05-24T09:01:43.167Z",
          "content": "<p>What does 2.5d+1d mean？I just use 2.5d model</p>",
          "rawMarkdown": "What does 2.5d+1d mean？I just use 2.5d model",
          "votes": 1
        }
      ]
    },
    {
      "id": 2251914,
      "postDate": "2023-05-09T18:06:27.107Z",
      "content": "<p>a useful paper for my work:</p>\n<p>Noise2Noise: Learning Image Restoration without Clean Data<br>\n<a href=\"https://arxiv.org/pdf/1803.04189.pdf\" target=\"_blank\">https://arxiv.org/pdf/1803.04189.pdf</a></p>",
      "rawMarkdown": "a useful paper for my work:\n\nNoise2Noise: Learning Image Restoration without Clean Data\nhttps://arxiv.org/pdf/1803.04189.pdf\n\n",
      "votes": 2
    },
    {
      "id": 2296668,
      "postDate": "2023-06-12T03:57:47.863Z",
      "content": "<p>speedup yout thresholding testing by submitting public test only<br>\n(warning !!!! remember to resubmit for both public+private test after finding the best threshold)</p>\n<pre><code>  submission = defaultdict()\n     fragment_id  valid_id:\n        d = read\n\n\n         fragment_id==:\n            rle =' '  #\n        : \n            probability = \n            count = \n             i, cfg  enumerate(configure):\n\n                print\n                net = cfg.\n                f = torch.load(cfg.checkpoint, map_location=lambda storage, loc: storage)\n                print(net.load)  # True\n                net.cuda\n                net.eval\n\n                p = infer</code></pre>",
      "rawMarkdown": "speedup yout thresholding testing by submitting public test only\n(warning !!!! remember to resubmit for both public+private test after finding the best threshold)\n\n```\n\n  submission = defaultdict(list)\n    for fragment_id in valid_id:\n        d = read_data1(fragment_id, z0=32-16, z1=32+16)\n \n        \n        if fragment_id=='b':\n            rle ='1 2'  #private\n        else: \n            probability = 0\n            count = 0\n            for i, cfg in enumerate(configure):\n         \n                print_cfg(cfg)\n                net = cfg.Net()\n                f = torch.load(cfg.checkpoint, map_location=lambda storage, loc: storage)\n                print(net.load_state_dict(f['state_dict'], strict=True))  # True\n                net.cuda()\n                net.eval()\n\n                p = infer_one_ms(net, d, cfg)\n                ...\n\n```\n"
    },
    {
      "id": 2296582,
      "postDate": "2023-06-12T01:25:58.027Z",
      "content": "<p>Great post!</p>",
      "rawMarkdown": "Great post!"
    },
    {
      "id": 2279236,
      "postDate": "2023-05-29T07:48:49.943Z",
      "content": "<p>deleted message</p>",
      "rawMarkdown": "deleted message"
    },
    {
      "id": 2277603,
      "postDate": "2023-05-27T22:38:46.897Z",
      "content": "<p>Is the training code released?, Anyone can share the link to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's training code for any model.</p>",
      "rawMarkdown": "Is the training code released?, Anyone can share the link to @hengck23 's training code for any model.",
      "replies": [
        {
          "id": 2277698,
          "postDate": "2023-05-28T01:30:36.237Z",
          "content": "<p>There is no training code. But it is easy to reproduce the results with your own training code</p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2262506\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2262506</a></p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2274502\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2274502</a></p>",
          "rawMarkdown": "There is no training code. But it is easy to reproduce the results with your own training code\n\nhttps://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2262506\n\nhttps://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2274502\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2261180,
      "postDate": "2023-05-16T06:52:44.123Z",
      "content": "<p>It turns out that you don't need decoder after all?</p>",
      "rawMarkdown": "It turns out that you don't need decoder after all?"
    },
    {
      "id": 2258470,
      "postDate": "2023-05-14T08:20:42.397Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, your idea is very inspiring. I'm interested in attention pool you mentioned, do you have any suggested paper or blog about this technique?<br>\nMany thanks to your sharing!</p>",
      "rawMarkdown": "Hi, @hengck23, your idea is very inspiring. I'm interested in attention pool you mentioned, do you have any suggested paper or blog about this technique?\nMany thanks to your sharing!",
      "replies": [
        {
          "id": 2258558,
          "postDate": "2023-05-14T10:16:03.667Z",
          "content": "<p>google for video unet or spatiotemporal unet (e.g. satellite image segmentation with input over 12 months).<br>\nthe time axis can be our z axis here.</p>",
          "rawMarkdown": "google for video unet or spatiotemporal unet (e.g. satellite image segmentation with input over 12 months).\nthe time axis can be our z axis here.",
          "votes": 5,
          "replies": [
            {
              "id": 2258567,
              "postDate": "2023-05-14T10:22:09Z",
              "content": "<p>e.g. <a href=\"https://zhuanlan.zhihu.com/p/421147308\" target=\"_blank\">https://zhuanlan.zhihu.com/p/421147308</a><br>\n<a href=\"https://github.com/VSainteuf/utae-paps\" target=\"_blank\">https://github.com/VSainteuf/utae-paps</a><br>\n<a href=\"https://github.com/VSainteuf/utae-paps/blob/main/gfx/utae.png\" target=\"_blank\">https://github.com/VSainteuf/utae-paps/blob/main/gfx/utae.png</a> </p>",
              "rawMarkdown": "e.g. https://zhuanlan.zhihu.com/p/421147308\nhttps://github.com/VSainteuf/utae-paps\nhttps://github.com/VSainteuf/utae-paps/blob/main/gfx/utae.png ",
              "votes": 3
            },
            {
              "id": 2259449,
              "postDate": "2023-05-15T02:39:01.400Z",
              "content": "<p>thanks for sharing!</p>",
              "rawMarkdown": "thanks for sharing!"
            }
          ]
        }
      ]
    },
    {
      "id": 2256383,
      "postDate": "2023-05-12T12:56:20.120Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, </p>\n<p>Thanks for sharing your ideas and thoughts. I would like to try out something similar, and have a few questions: </p>\n<ul>\n<li>what exactly do you mean with \"no context\" for the 3DCnnEncoder? And what do the multiple 3DCnnEncoders represent, the same model trained on different folds?</li>\n<li>on what voxel size are you training the 3DCNNEncoder? I guess at least 16x16 otherwise you don't need the AdaptivePooling layer you mention in between the 3DCnnEncoder and the InkDetector to average the feature maps.</li>\n</ul>",
      "rawMarkdown": "Hi @hengck23, \n\nThanks for sharing your ideas and thoughts. I would like to try out something similar, and have a few questions: \n\n- what exactly do you mean with \"no context\" for the 3DCnnEncoder? And what do the multiple 3DCnnEncoders represent, the same model trained on different folds?\n- on what voxel size are you training the 3DCNNEncoder? I guess at least 16x16 otherwise you don't need the AdaptivePooling layer you mention in between the 3DCnnEncoder and the InkDetector to average the feature maps."
    },
    {
      "id": 2254055,
      "postDate": "2023-05-10T16:16:37.307Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>! </p>\n<blockquote>\n  <p>Follow experiment protocols from [2]. divide fragments into upper and lower sub-fragments. validation = one of the sub-fragments. I have the same results as the paper[2], table.1: recall=0.41, FPR=0.051</p>\n</blockquote>\n<p>I have few questions if you like to answer</p>\n<p>1) what is the sub-voxel shape as input to the model? did you use a 6-fold validation scheme (instead of 8 mentioned in the paper) ? <br>\n2) how many epochs/iters train in your experiment and how long takes per epoch approx?<br>\n3) the net output was a 2d mask or the center pixel xy-coords class? </p>\n<blockquote>\n  <p>This could be why some kagglers reported that if you use too many slices results are not as good?</p>\n</blockquote>\n<p>BTW I can confirm this as well, whenever tried to include more channels/slices metrics were worst</p>",
      "rawMarkdown": "Thanks for sharing @hengck23! \n\n> Follow experiment protocols from [2]. divide fragments into upper and lower sub-fragments. validation = one of the sub-fragments. I have the same results as the paper[2], table.1: recall=0.41, FPR=0.051\n\nI have few questions if you like to answer\n\n1) what is the sub-voxel shape as input to the model? did you use a 6-fold validation scheme (instead of 8 mentioned in the paper) ? \n2) how many epochs/iters train in your experiment and how long takes per epoch approx?\n3) the net output was a 2d mask or the center pixel xy-coords class? \n\n\n> This could be why some kagglers reported that if you use too many slices results are not as good?\n\nBTW I can confirm this as well, whenever tried to include more channels/slices metrics were worst",
      "replies": [
        {
          "id": 2286473,
          "postDate": "2023-06-03T13:35:38.980Z",
          "content": "<p>My preliminary experiments showed that slices from 28 to 33 achieve a better cross-validation. Merging all training fragments and splitting a 5-fold cross validation is also a good way to go.</p>",
          "rawMarkdown": "My preliminary experiments showed that slices from 28 to 33 achieve a better cross-validation. Merging all training fragments and splitting a 5-fold cross validation is also a good way to go."
        }
      ]
    },
    {
      "id": 2251189,
      "postDate": "2023-05-09T06:41:45.253Z",
      "content": "<p>You will use a 3D encoder-decoder for segmentation?</p>",
      "rawMarkdown": "You will use a 3D encoder-decoder for segmentation?",
      "replies": [
        {
          "id": 2251282,
          "postDate": "2023-05-09T08:28:08.430Z",
          "content": "<p>no. 2d segmentation. input channels are output of previous 3dcnn, which are already 3d pooled.</p>\n<p>this is one of the reason for 3d pooling. another is that the cam class activation map of 3d poolong tells you which of z slice get activated for different location</p>",
          "rawMarkdown": "no. 2d segmentation. input channels are output of previous 3dcnn, which are already 3d pooled.\n\nthis is one of the reason for 3d pooling. another is that the cam class activation map of 3d poolong tells you which of z slice get activated for different location",
          "votes": 4
        }
      ]
    },
    {
      "id": 2283490,
      "postDate": "2023-06-01T09:38:08.497Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2282052,
      "postDate": "2023-05-31T10:23:49.123Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2264637,
      "postDate": "2023-05-18T15:38:48.303Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2266159,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-19T18:55:05.797000",
      "content": "<p>how to pool with pos encoding (use conv3d as convolutional positional encoding in z direction)</p>\n<pre><code>        #-- pool attention weight\n        self.weight = nn.ModuleList([\n            nn.Sequential(\n                nn.Conv3d(dim, dim, kernel_size=3, padding=1),\n                nn.ReLU(inplace=True),\n            ) for dim in encoder_dim\n        ])\n\n    def forward(self, batch):\n        v = batch['volume']\n        B,C,H,W = v.shape\n        vv = [\n            v[:,i:i+CFG.crop_depth] for i in [0,2,4,]\n        ]\n        K = len(vv)\n        x = torch.cat(vv,0)\n\n        # ---------------------------------\n        # encoder = self.encoder.forward_features(x)\n        encoder = []\n        x = self.encoder.conv1(x)\n        x = self.encoder.bn1(x)\n        x = self.encoder.act1(x)  ; encoder.append(x)\n        x = F.avg_pool2d(x,kernel_size=2,stride=2)\n        x = self.encoder.layer1(x); encoder.append(x)\n        x = self.encoder.layer2(x); encoder.append(x)\n        x = self.encoder.layer3(x); encoder.append(x)\n        x = self.encoder.layer4(x); encoder.append(x)\n        #print('encoder', [f.shape for f in encoder])\n        # ---------------------------------\n        for i in range(len(encoder)):\n            e = encoder[i]\n            _, c, h, w = e.shape\n            e = rearrange(e, '(K B) c h w -&gt; B c h w K', K=K, B=B, h=h, w=w) #\n            f = self.weight[i](e)\n            w = F.softmax(f, -1)\n            e = (w * e).sum(-1)\n            encoder[i] = e\n</code></pre>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 2272189,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-24T10:44:05.173000",
      "content": "<p>how to get lb 0.68 :</p>\n<ul>\n<li>augmentation: label noise</li>\n<li>model: stacked Unets</li>\n<li>validation: fragement1</li>\n<li>inference 4x rotate TTA</li>\n</ul>\n<p>example notebook: <br>\n<a href=\"https://www.kaggle.com/code/hengck23/lb0-68-one-fold-stacked-unet\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-68-one-fold-stacked-unet</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 2272273,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-05-24T11:36:00.950000",
          "content": "<p>architure used for lb0.68. <br>\nresnet34-pool-unet +  resnet10t-unet.<br>\nThis is very fast, only 10 min for 4x rotate TTA</p>\n<p><img src=\"https://i.ibb.co/RTPwsxv/Selection-999-2116.png\" alt=\"https://i.ibb.co/RTPwsxv/Selection-999-2116.png\"><br>\n<img src=\"https://i.ibb.co/8xJzn8R/Selection-999-2115.png\" alt=\"https://i.ibb.co/8xJzn8R/Selection-999-2115.png\"></p>",
          "votes": 12,
          "replies": [
            {
              "id": 2272308,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-24T12:00:41.143000",
              "content": "<p>label noise augmentation used:</p>\n<p><img src=\"https://i.ibb.co/p3G5fK0/Selection-999-2120.png\" alt=\"https://i.ibb.co/p3G5fK0/Selection-999-2120.png\"></p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2272409,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-24T13:15:03.823000",
              "content": "<p><img src=\"https://i.ibb.co/4m5QFHx/Selection-999-2123.png\" alt=\"https://i.ibb.co/4m5QFHx/Selection-999-2123.png\"></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2272510,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-24T14:46:22.777000",
              "content": "<p>what you can do next step:</p>\n<ul>\n<li>different way to stack … what is the input/output scale of each unet? what is the inter-connecting scale?</li>\n<li>different encoder/decoder for stacking … e.g. VIT, CNN …</li>\n</ul>\n<p>now the inference speed is very fast, you can consider  larger ensemble solution</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2272598,
              "author_name": "zw",
              "author_url": "",
              "post_date": "2023-05-24T15:46:01.010000",
              "content": "<p>why you just rotate label in \"add label noise \" augment？Won't this confusing the model👀</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2274380,
              "author_name": "Ioannis M",
              "author_url": "",
              "post_date": "2023-05-25T23:38:05.677000",
              "content": "<p>can you elaborate more on \"loss1\" please? <br>\nis it <code>BCE(logit1, mask_192x192)</code> ? </p>\n<p>also the overall loss you optimize is: <code>loss = 0.5*loss1 + 0.5*loss2</code> ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2274406,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-26T00:57:32.297000",
              "content": "<p>BCE(logit1, mask_192x192) is correct.</p>\n<p>loss is simply   loss1 + loss2</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2274407,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-26T01:01:46.447000",
              "content": "<p>model is stable:<br>\nfold1 = lb 0.68 /cv0.571<br>\nfold2aa = lb 0.67 /cv0.628<br>\nfold2bb = ? /cv0.647<br>\nfold2cc = ? /cv0.680<br>\nensemble  fold1+fold2aa+fold2bb+fold2cc = lb0.69 (12 min)</p>\n<p>all threshold are 0.5</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2276606,
              "author_name": "answer-qzd2",
              "author_url": "",
              "post_date": "2023-05-27T04:03:30.810000",
              "content": "<p>I am a beginner，I'm trying to train with your model, but the metrics are always the same value, confused</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2277032,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-05-27T11:36:37.070000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2278271,
              "author_name": "Givan",
              "author_url": "",
              "post_date": "2023-05-28T14:01:59.213000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  I'm not clear about the role of <code>valid_threshold</code>. Could you provide more details about its implementation?</p>\n<p>The following is my implementation, is it similar to yours implementation?</p>\n<pre><code> ():\n    \n    height, width, _ = volume.shape\n     :\n        x = np.random.choice(width - crop_size)  width &gt; crop_size  \n        y = np.random.choice(height - crop_size)  height &gt; crop_size  \n        cropped_volume = volume[y:y+crop_size, x:x+crop_size]\n        cropped_label = label[y:y+crop_size, x:x+crop_size]\n        cropped_mask = mask[y:y+crop_size, x:x+crop_size]\n         np.mean(cropped_mask) &gt; valid_threshold:\n             cropped_volume, cropped_label\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2278286,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-28T14:30:30.427000",
              "content": "<pre><code>## geometric --------------------------------\ndef do_random_crop(\n    image,\n    mask,\n    valid,\n    crop_size=224,\n    valid_threshold=0.5,\n):\n    height, width = image.shape[:2]\n    while (1):\n        y = np.random.randint(0, height - crop_size+1)\n        x = np.random.randint(0, width  - crop_size+1)\n        crop_image = image[y:y + crop_size, x:x + crop_size]\n        crop_mask  = mask[y:y + crop_size, x:x + crop_size]\n        crop_valid  = valid[y:y + crop_size, x:x + crop_size]\n        if crop_valid.mean() &gt; valid_threshold:\n            break\n    return crop_image, crop_mask\n\n\ndef do_random_affine_crop(\n    image,\n    mask,\n    valid,\n    crop_size=224,\n    scale  = (0.5,2.0),\n    aspect = (0.9,1/0.9),\n    degree = (-45,45),\n    image_fill=0,\n    mask_fill=0,\n    valid_threshold=0.5,\n):\n    height, width = image.shape[:2]\n    s = crop_size\n    point = np.array([\n        [0,0],\n        [0,s],\n        [s,s],\n        [s,0],\n    ])\n    point1 = do_random_affine_on_point(\n        point,\n        scale  = scale,\n        aspect = aspect,\n        degree = degree,\n    )\n    point1 = point1.astype(int)\n    x0, y0 = point1[:,0].min(), point1[:,1].min()\n    point1 = point1 - [[x0, y0]]\n    w, h = point1[:,0].max(), point1[:,1].max()\n\n    while (1):\n        y = np.random.randint(0, height - w+1)\n        x = np.random.randint(0, width  - h+1)\n        crop_valid  = valid[y:y + h, x:x + w]\n        if crop_valid.mean() &gt; valid_threshold:\n            break\n\n    point2 = point1 + [[x, y]]\n    #----\n    matrix = cv2.getAffineTransform(\n        point2[:3].astype(np.float32),\n        point[:3].astype(np.float32),\n    )\n    crop_image = cv2.warpAffine(\n        image, matrix, (crop_size,crop_size),\n        flags=cv2.INTER_LINEAR,borderMode=cv2.BORDER_CONSTANT, borderValue=image_fill,\n    )\n    crop_mask = cv2.warpAffine(\n        mask, matrix, (crop_size,crop_size),\n        flags=cv2.INTER_LINEAR,borderMode=cv2.BORDER_CONSTANT, borderValue=mask_fill,\n    )\n\n\n    return crop_image, crop_mask\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2278297,
              "author_name": "zw",
              "author_url": "",
              "post_date": "2023-05-28T14:43:57.843000",
              "content": "<p>May I ask what's the image size before you random cropp it</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2278400,
              "author_name": "Givan",
              "author_url": "",
              "post_date": "2023-05-28T16:42:27.613000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for your very clear implementation. If I understood correctly, the <code>label</code> in my code is equivalent to the <code>mask</code> in your code. I entered this competition late and I'm currently trying my best to reproduce your experiments and writing a training script. Did you publish any training code that I haven't found? As always, your experimental thread is a wonderful learning material. 👍👍👍</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2278856,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-29T02:58:19.520000",
              "content": "<pre><code>def read_data(fragment_id, z0=Z0, z1=Z1):\n    # volume = np.load(f'{train_dir}/{i}/surface_volume{i}.8.npy')\n    volume = []\n    start_timer = timer()\n    for j in range(z0, z1):\n        v = np.array(Image.open(f'{train_dir}/{fragment_id}/surface_volume/{j:02d}.tif'), dtype=np.uint16)\n        v = (v &gt;&gt; 8).astype(np.uint8)\n        # v = (v / 65535.0 * 255).astype(np.uint8)\n        volume.append(v)\n        print(f'\\r @ read volume{j}  {time_to_str(timer() - start_timer, \"sec\")}', end='', flush=True)\n    print('')\n    volume = np.stack(volume, -1)\n    height, width, depth = volume.shape\n    print(f'fragment_id={fragment_id} volume: {volume.shape}')\n\n    # ---\n    ir = cv2.imread(f'{train_dir}/{fragment_id}/ir.png', cv2.IMREAD_GRAYSCALE)\n    label = cv2.imread(f'{train_dir}/{fragment_id}/inklabels.png', cv2.IMREAD_GRAYSCALE)\n    mask = cv2.imread(f'{train_dir}/{fragment_id}/mask.png', cv2.IMREAD_GRAYSCALE)\n    volume[mask == 0] = 0\n\n    # ----\n    ir = ir / 255\n    label = do_binarise(label)\n    mask = do_binarise(mask)\n    if 0:\n        image_show_norm(f'ir', ir, min=0, max=1, resize=0.1)\n        image_show_norm(f'label', label, min=0, max=1, resize=0.1)\n        cv2.waitKey(0)\n\n    d = dotdict(\n        fragment_id=fragment_id,\n        crop_id=fragment_id,\n        volume=volume,\n        ir=ir,\n        label=label,\n        mask=mask,\n    )\n    return d\n</code></pre>\n<pre><code>volume, label = do_random_affine_crop(\n            d.volume[..., z0:z1],\n            d.label,\n            d.mask,\n            crop_size=crop_size,\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2280358,
              "author_name": "kranz",
              "author_url": "",
              "post_date": "2023-05-30T04:37:37.503000",
              "content": "<p>Why the input is 384, mask is 192, how to get mask？</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2277768,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-05-28T03:52:58.960000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2281715,
          "author_name": "lionfishy",
          "author_url": "",
          "post_date": "2023-05-31T04:51:30.050000",
          "content": "<p>Thanks for your great work! I am reading through your code. Do you mind explaining what infer_mask do and whether it makes a significant difference in model performance? Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2260220,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-15T14:37:17.087000",
      "content": "<p>Which to use: 3D, 2.5D or pool-2.5D? </p>\n<p>it actually depends how different is the z location of  target signal in train and hidden test data.</p>\n<p>1) Assume that both train,test has signal in about the same z location, then 2.5D is enough. 2d convolution is only invariant to changes in (x,y) location and not  invariant to channel (z). But target has same \"z image characteristics\", so 2d convolution is enough.</p>\n<p>2) Assume train,test has signal in very different z locations. Then you need 3d conv, invariant to changes in (x,y,z) location. But 3d convolution is costly. But this may not be true in the competition because the true resolution of the ink target is low (the inklabels ong images are enlarged version of smaller resolution)</p>\n<hr>\n<p>pool-2.5D is essentially a factorised version of 3d convolution. the pooling is 1d convolution in z direction. attention-pool uses dynamic weight in z direction. if also use fixed weights (via nn parameters) or equal weight(mean pool)</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2263103,
          "author_name": "Zien Huang",
          "author_url": "",
          "post_date": "2023-05-17T10:41:18.373000",
          "content": "<p>Very useful thoughts, thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2258613,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-14T10:49:56.153000",
      "content": "<p>i have a feeling that the test images are rotated</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2258715,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-05-14T12:35:22.230000",
          "content": "<p>i realised that one can train a rotation classifier to prob the test</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2259804,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2023-05-15T08:36:25.433000",
          "content": "<p>hmm I thought the same. </p>\n<p>BTW looking at scores from your excel sheet I see that no.3 <code>rot90(k=2)</code> and no.5 <code>rot(k=1,2,3)</code> are exactly same  </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2259838,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-15T09:09:36.743000",
              "content": "<p>thanks, i updated for \"5)TTA: rot(k=0,1,2,3)\"</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2260144,
          "author_name": "Raki",
          "author_url": "",
          "post_date": "2023-05-15T13:39:41.747000",
          "content": "<p>I tried looking into this more, TTA with mean of 4 rot improved LB form 53 to 67 for me. <br>\nI tried optimizing TTA on val set (fold 1) but did not see improvements in val score, until I noticed that I only used horizontal and vertical flips in augmentation and no rot90, I added rot90 as train augmentation and suspected that now that this kind of rotation is part of the training data distribution, I would see much higher score on the test set (given that test images are rotated) even without TTA, but I noticed almost no difference in LB to the 53 LB (no TTA). If someone has an idea what might be going on, I would really appreciate your feedback!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2260243,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-15T14:52:37.470000",
              "content": "<ol>\n<li><p>it depends on the  probability rotated augmentation in training. If there is equal chance that the training input is rotated, then TTA in validation and test will not show any difference in performance. That is also provided that rotated image will not results in label noise (i.e. there are no cases that rotated positive = negative  or rotated negative = positive)</p></li>\n<li><p>at inference, background pixel may be falsely labelled as positive, but it is unlikely that the same perturbed pixel will still falsely lablled as positive. so TTA has an effect of rejecting  false positive (rejecting more  false positive at the expense of rejecting some weak true positive)</p></li>\n</ol>\n<p>in theory you can apply TTA at some selective pixel (instead of all) to get even better results.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2260250,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-15T14:55:43.723000",
              "content": "<p>i wonder if anyone is worried that rotation only works for public test and actually degrade resulst for hidden private data?</p>\n<p>what does cross validation show?  the issue for aligning cross validation with lb score is that (i think) the FP rate of test is higher (see the need for higher threshold). it is very hard to draw confident conclusion.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2260311,
              "author_name": "Lucas",
              "author_url": "",
              "post_date": "2023-05-15T15:45:06.323000",
              "content": "<p>hmm, my score is actually not getting better at all with TTA.. very strange</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2261053,
              "author_name": "dmitrykonovalov",
              "author_url": "",
              "post_date": "2023-05-16T04:33:20.273000",
              "content": "<p>I get no LB improvement from TTA (hv-flips + 90-rots). I think because I trained with albu augs hv-flips and -360+360 range rots.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2301888,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-06-14T07:49:13.020000",
      "content": "<p>\"人人有机会， 个个没把握\" (everyone has a chance but no one is sure to win)</p>\n<p>Be sure to read the data paper. \"Imagine\" how the train and public / private test fragment would look like …</p>\n<p>estimate your predicted ink pixel precision and number of truth pixels carefully.<br>\nis it more prediction  = more fp ?<br>\nor more prdiction = more tp?</p>\n<p>don't be shaken down by wrong threshold.</p>\n<p>best threshold for public  = best threshold for private ???? </p>\n<ul>\n<li><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F83f199f9403710ea8f63bb15eba0bbf4%2FSelection_999(2223).png?generation=1686729121339846&amp;alt=media\" alt=\"\"></li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 2302891,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-06-15T00:13:24.947000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd42e66f572884975cc1d23fb36b986ad%2FSelection_999(2227).png?generation=1686787992964081&amp;alt=media\" alt=\"\"></p>\n<p>did you get my hint?</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2302894,
              "author_name": "june",
              "author_url": "",
              "post_date": "2023-06-15T00:17:21.137000",
              "content": "<p>Congrats on your solo gold!!! </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2283935,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-06-01T15:45:07.650000",
      "content": "<p>did i just find some magic?<br>\nfrom <a href=\"https://www.kaggle.com/hughsando\" target=\"_blank\">@hughsando</a>, <a href=\"https://www.kaggle.com/code/hughsando/visualize-the-3d-fragment-height\" target=\"_blank\">https://www.kaggle.com/code/hughsando/visualize-the-3d-fragment-height</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc2c92845f59f35f8c4de32e5a23de2b6%2FSelection_999(2175).png?generation=1685634284049911&amp;alt=media\" alt=\"\"></p>\n<p>height image and label superimposed</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2283936,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-06-01T15:45:49.273000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff5e1e63504811b49ad344a0b58f87baa%2FSelection_999(2174).png?generation=1685634346837324&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2289546,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-06T08:20:52.773000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2284364,
          "author_name": "ML Student06",
          "author_url": "",
          "post_date": "2023-06-02T01:26:23.860000",
          "content": "<p>This is very interesting. From the looks of the notebook, it seems as though the threshold is experimentally found. However, if an optimal threshold can be found algorithmically, the segmentation could become much more obvious and simple.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2293244,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-06-09T04:00:16.170000",
          "content": "<p>linking height image and activation CAM heatmap</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1550c0c29ce587631a4aaf4454523dc4%2FSelection_999(2218).png?generation=1686283189708144&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2294333,
              "author_name": "ML Student06",
              "author_url": "",
              "post_date": "2023-06-10T01:44:16.657000",
              "content": "<p>Can you expand more on what you are doing here? I am a beginner, so please mind the simple questions. Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2277577,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-27T22:04:44.323000",
      "content": "<p>the power of context!!!!<br>\n<a href=\"https://www.kaggle.com/fengqilong\" target=\"_blank\">@fengqilong</a> check this! freeze and scale …<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F37dd82554ad4e12377b86900f16cc541%2FSelection_999(2146).png?generation=1685247621602263&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 2277732,
          "author_name": "Feng Qilong",
          "author_url": "",
          "post_date": "2023-05-28T02:30:49.050000",
          "content": "<p>nice improvement! What do you mean by scaling the encoder tho? is it just the mean pooling?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2277779,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-05-28T04:20:58.847000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2ba87f0a1d2314657c1a83ca28f09187%2FSelection_999(2139).png?generation=1685247643963019&amp;alt=media\" alt=\"\"><br>\nvisual results</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2278161,
          "author_name": "kranz",
          "author_url": "",
          "post_date": "2023-05-28T12:02:51.093000",
          "content": "<p>its wonderful!<br>\nbut after average pool，the feature map can‘t decoder to 768x768.<br>\nHow you did it？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2284888,
          "author_name": "Sally2015",
          "author_url": "",
          "post_date": "2023-06-02T10:27:01.537000",
          "content": "<p>inspiring, I will give it a try.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2263872,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-18T01:47:59.477000",
      "content": "<p>paper:<br>\nDeciphering Ancient Papyrus Texts: A Machine Learning Approach with Varied Architectures and Encoders<br>\n<a href=\"https://www.researchgate.net/publication/370729553_Deciphering_Ancient_Papyrus_Texts_A_Machine_Learning_Approach_with_Varied_Architectures_and_Encoders\" target=\"_blank\">https://www.researchgate.net/publication/370729553_Deciphering_Ancient_Papyrus_Texts_A_Machine_Learning_Approach_with_Varied_Architectures_and_Encoders</a></p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2250988,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-09T01:07:16.580000",
      "content": "<p><img src=\"https://i.ibb.co/ZVYqcMQ/Selection-999-1975.png\" alt=\"https://i.ibb.co/ZVYqcMQ/Selection-999-1975.png\"> <img src=\"https://i.ibb.co/mzBHh89/Selection-999-1976.png\" alt=\"https://i.ibb.co/mzBHh89/Selection-999-1976.png\"></p>\n<p>4th fragment . or ?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2251195,
          "author_name": "Kyle Peters",
          "author_url": "",
          "post_date": "2023-05-09T06:47:39.400000",
          "content": "<p>How did you obtain the fourth fragment?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2251461,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-09T11:11:29.597000",
              "content": "<p>one of youtube presentation videos in the early days by the orangizer<br>\n<a href=\"https://www.youtube.com/watch?v=T0mWqsFrJpk&amp;t=3507s\" target=\"_blank\">https://www.youtube.com/watch?v=T0mWqsFrJpk&amp;t=3507s</a><br>\n<a href=\"https://www.youtube.com/watch?v=YWfIx5ggeYI&amp;t=121s\" target=\"_blank\">https://www.youtube.com/watch?v=YWfIx5ggeYI&amp;t=121s</a></p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2251636,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-09T14:28:37.667000",
              "content": "<p>according to the video, some kaggle train fragment are actually overlapping two fragments?</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2251983,
          "author_name": "JP Posma",
          "author_url": "",
          "post_date": "2023-05-09T19:41:56.040000",
          "content": "<p>Interesting find! We obviously can't talk about the origins of the secret fragment. However, this is a good opportunity to remind everyone of the rule prohibiting use of external Herculaneum papyri sources, such as this fragment, or others that might be floating around. Submissions must be a legitimate ink-detection method reproducible by the contest organizers, including training.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2252018,
              "author_name": "John Costella",
              "author_url": "",
              "post_date": "2023-05-09T20:10:24.683000",
              "content": "<p>🙂 I was saying to my wife Sally last night, after reading this thread: what would be the odds that I enter my second Kaggle comp in 11 years, and <em><a href=\"https://www.kaggle.com/c/FacebookRecruiting/discussion/2082#11936\" target=\"_blank\">it gets blown up again</a>?</em> 🙂</p>\n<p>In this case, if this <em>is</em> the hidden fragment (and I can see why it would be sensible to slice it horizontally, like the dummy test fragments supplied, so … sadness 😔), then some entries on the leaderboards (private and public) could potentially be meaningless, but as <a href=\"https://www.kaggle.com/jpposma\" target=\"_blank\">@jpposma</a> notes, it's not going to affect the integrity of the final judgments by the organizers, which will discard those submissions.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2252025,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-09T20:22:59.117000",
              "content": "<p>The hidden fragment is in fragment3. (I think the scrolls get folded in)<br>\nThis could be why some kagglers reported that if you use too many slices results are not as good?</p>\n<p>(hence the model should be detecting the ink at the surface volume, rather than ink in the volume.)</p>\n<p><a href=\"https://imgbb.com/\"><img src=\"https://i.ibb.co/jfQs9JV/Selection-999-1981.png\" alt=\"Selection-999-1981\"></a></p>\n<hr>\n<p>fragment4 is a new fragment not in kaggle train dataset. <br>\nIt shows how the writings can vary and look different from the training set.</p>\n<p>(if you use too much context, it may overfit and not generalisd)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2252044,
              "author_name": "John Costella",
              "author_url": "",
              "post_date": "2023-05-09T20:45:20.853000",
              "content": "<p>Sorry, yes, <code>s/hidden/fourth/</code>.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2258958,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-14T15:09:35.207000",
              "content": "<p><a href=\"https://www.kaggle.com/jpposma\" target=\"_blank\">@jpposma</a> </p>\n<p>refer to the dicussion at <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/395233#2184750\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-ink-detection/discussion/395233#2184750</a></p>\n<p>you mentioned \"You are allowed to use any data from <a href=\"https://scrollprize.org/data\" target=\"_blank\">https://scrollprize.org/data</a> in your training\"</p>\n<hr>\n<p>Now I am starting experiment on self-supervised learning unlabelled data.<br>\ni want to clarify again both data below can be used:</p>\n<ul>\n<li>\"/full-scrolls\"  (non-kaggle,  Grand Prize data )</li>\n<li>\"/fragments\"  (kaggle data)</li>\n</ul>\n<p>Note that i may may recover more hidden unlabelled fragements from the raw \"/fragments\" if i re-segment or segment deeper (instead of the surface data for kaggle) the volume myself.</p>\n<p>Thanks!</p>\n<p><a href=\"https://scrollprize.org/data_organization\" target=\"_blank\">https://scrollprize.org/data_organization</a><br>\n<a href=\"https://scrollprize.org/data\" target=\"_blank\">https://scrollprize.org/data</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2260563,
              "author_name": "JP Posma",
              "author_url": "",
              "post_date": "2023-05-15T18:11:27.613000",
              "content": "<p>Yes, both \"/full-scrolls\" and \"/fragments\" can be used. Just not any external data pertaining to the Herculaneum scrolls, like the Youtube videos you found.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2265521,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-19T09:59:09.877000",
      "content": "<p>very important details of the hidden test data:<br>\n[1] EduceLab-Scrolls: Verifiable Recovery of Text from Herculaneum Papyri using X-ray CT<br>\n<a href=\"https://arxiv.org/pdf/2304.02084.pdf\" target=\"_blank\">https://arxiv.org/pdf/2304.02084.pdf</a></p>\n<p>Figure 1. </p>\n<p>\" Fragments 1 and 2 are from a philosophical work authored by Philodemus, as<br>\nare many others in the Herculaneum collection. This book, titled On slander, comes from his greater work On Vices and the opposite virtues.\"</p>\n<p>\"Fragments 3 and 4 come from a scroll about Hellenistic Dynastic history\"</p>\n<hr>\n<p>\"hidden, subsurface layers of the scroll fragments for which there is no ground truth.\"</p>\n<hr>\n<p>\"in which a trained model is applied to hidden layers extracted from the X-ray CT under the exposed surface. In this case, there is no risk of memorization, so a model is trained on the entire training dataset, or all four fragment surfaces.\"</p>\n<hr>\n<p>Table 2.</p>\n<pre><code>Fragment Characters Recall FPR\n4 25 0.64 0.12\n</code></pre>\n<p>\"Importantly, the misidentified characters are not invented “whole cloth.” They are at least identified in the correct location, and could be accurately identified with marginal improvements in ink detection\"</p>\n<hr>\n<p>Table 3.</p>\n<pre><code>Size (GB) Surface volume (GB)\nFragment 4 198 7.3\n</code></pre>\n<p>Use GB size to estimate image size</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2265133,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-19T02:30:02.850000",
      "content": "<p>lb0.65 recipe </p>\n<p>encoder = resnext26d<br>\nvalidation = fragement 1<br>\nuse pool resnet unet </p>\n<p>(same configure is 0.54 for fragement 2b validation)</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2265192,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-05-19T04:11:54.037000",
          "content": "<pre><code>        conv_dim = 64\n        encoder_dim  = [conv_dim, 256, 512, 1024, 2048]#64, 128, 256, 512, ] #[\n        decoder_dim  = [256, 128, 64, 32, 16]\n        self.encoder = seresnext26t_32x4d(pretrained=True, in_chans=CFG.crop_depth)\n</code></pre>\n<pre><code>** start training here! **\n   batch_size = 64 \n   experiment = ['004_pool_resnet26_224_unet_00', 'run_train_fold1.py']\n                           |----------------------- VALID---------------|---- TRAIN/BATCH ----------------------\nrate      iter       epoch | loss   thr    recall  fpr   p_sum  score   | loss                 | time           \n----------------------------------------------------------------------------------------------------------------\n0.00e+0   00000000*   0.00 | 0.663  0.100  1.000  1.000  1.777  0.126  | 0.000  0.000  0.000  |  0 hr 00 min\n1.00e-3   00000216*   1.00 | 0.526  0.400  0.973  0.431  0.866  0.244  | 0.581  0.000  0.000  |  0 hr 08 min\n1.00e-3   00000432*   2.00 | 0.449  0.500  0.298  0.034  0.109  0.442  | 0.525  0.000  0.000  |  0 hr 15 min\n1.00e-3   00000648*   3.00 | 0.387  0.500  0.234  0.023  0.080  0.425  | 0.461  0.000  0.000  |  0 hr 22 min\n1.00e-3   00000864*   4.00 | 0.314  0.400  0.317  0.025  0.098  0.505  | 0.404  0.000  0.000  |  0 hr 29 min\n1.00e-3   00001080*   5.00 | 0.292  0.400  0.335  0.023  0.098  0.533  | 0.361  0.000  0.000  |  0 hr 36 min\n1.00e-3   00001296*   6.00 | 0.260  0.500  0.284  0.018  0.080  0.517  | 0.328  0.000  0.000  |  0 hr 43 min\n1.00e-3   00001512*   7.00 | 0.242  0.500  0.289  0.015  0.077  0.537  | 0.312  0.000  0.000  |  0 hr 50 min\n1.00e-3   00001728*   8.00 | 0.243  0.600  0.388  0.030  0.119  0.539  | 0.290  0.000  0.000  |  0 hr 57 min\n1.00e-3   00001944*   9.00 | 0.270  0.400  0.301  0.022  0.091  0.505  | 0.275  0.000  0.000  |  1 hr 04 min\n</code></pre>\n<pre><code>model_path=/home/titanx/hengck/share1/kaggle/2022/ink-detect/result/run004/pool_resnet26_224_unet_00/fold-1/checkpoint/00001728.model.pth\nfragment_id=1\nCFG.stride=56\n\nbce=0.24270\np_sum  th   prec   recall   fpr   dice   score\n----------------------------------------------\n0.33, 0.10, 0.374, 0.683, 0.131,  0.483,  0.411\n0.15, 0.20, 0.547, 0.445, 0.042,  0.491,  0.523\n0.09, 0.30, 0.647, 0.327, 0.021,  0.434,  0.541\n0.06, 0.40, 0.730, 0.243, 0.010,  0.364,  0.521\n0.04, 0.50, 0.792, 0.178, 0.005,  0.291,  0.468\n0.03, 0.60, 0.854, 0.128, 0.003,  0.223,  0.401\n0.02, 0.70, 0.906, 0.081, 0.001,  0.148,  0.297\n0.01, 0.80, 0.947, 0.035, 0.000,  0.068,  0.153\n0.00, 0.90, 0.992, 0.002, 0.000,  0.004,  0.009\n</code></pre>",
          "votes": 4,
          "replies": [
            {
              "id": 2280517,
              "author_name": "Monojito",
              "author_url": "",
              "post_date": "2023-05-30T06:37:36.690000",
              "content": "<p>It seems that there are only three training samples here. How to use the batch size of 64. Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2256006,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-12T07:37:24.237000",
      "content": "<p>how to select the best depth crop:</p>\n<p><a href=\"https://ibb.co/RzwbWtb\"><img src=\"https://i.ibb.co/wMmg18g/Selection-999-2027.png\" alt=\"Selection-999-2027\"></a></p>\n<p>as reference:<br>\nvesuvius_2d_slide_exp002/vesuvius-models/Unet_fold1_best.pth<br>\nfragment 1 validation 0.572</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2256016,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-05-12T07:46:26.640000",
          "content": "<p>you can treat it as multiple instance segmentation (for each pixel).<br>\ndifferent pixel location uses different depth (assume the surface is not flat)</p>\n<p>you can pool:<br>\n1) at the end  of each scale of encoder <br>\n2) at the end of last scale decoder,<br>\netc …</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2258379,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-14T06:44:51.330000",
              "content": "<p>example notebook at: <a href=\"https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d</a></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2259590,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-15T05:38:01.613000",
              "content": "<p>one fold: fragment2a: LB 0.56 , threshold=0.5<br>\ntwo fold: fragment2a,2b: LB 0.59, threshold=0.5</p>\n<hr>\n<p>LB result (falls to 0.4x range) not that good if i add positional encoding in z-axis. maybe the test data are different?</p>\n<p>but CV result results improve by 0.01 for the same positional encoding</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2287963,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-06-05T04:31:21.147000",
      "content": "<p>better threshold method?</p>\n<ul>\n<li><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F14c90169c92f5db5e2c02a2b23ccc237%2FSelection_999(2194).png?generation=1685967907793639&amp;alt=media\" alt=\"\"></li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 2288497,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-06-05T12:26:49.290000",
          "content": "<p>it is local contrast equalization</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F448f98a9d7ca6ef3c45417f05bc01071%2FSelection_999(2195).png?generation=1685967973387298&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2281576,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-31T01:55:25.057000",
      "content": "<p>surprise!!!!!<br>\ni am surprise that i can train will crop depth of 32</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb3ae8a9d4cb24ac90713e416253ee775%2FSelection_999(2170).png?generation=1685498117144286&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 2286191,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-06-03T09:40:43.697000",
          "content": "<p>simplified using mean pooling:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/meanpool-resnet34d-unet\" target=\"_blank\">https://www.kaggle.com/code/hengck23/meanpool-resnet34d-unet</a></p>\n<pre><code></code></pre>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2265586,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-19T11:10:34.710000",
      "content": "<p><a href=\"https://stackoverflow.com/questions/71690251/binary-cross-entropy-with-logits-weight-vs-pos-weight-what-are-the-differences\" target=\"_blank\">https://stackoverflow.com/questions/71690251/binary-cross-entropy-with-logits-weight-vs-pos-weight-what-are-the-differences</a></p>\n<p>The pos_weight parameter allows you to balance the positive example thus controlling the tradeoff between recall and precision (see also). A detailed explanation can be found on this thread along with the explicit math expression. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2265153,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-19T02:58:09.640000",
      "content": "<p>it is probably difficult to augment positive samples …<br>\nbut it is easy to augment negative samples !!!!</p>\n<p>since fbeta favours high precision, if you can create \"infinite negative samples\" , then ….<br>\nfurther, there are so many samples in the external dataset at the scrollprize website</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2260347,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-15T16:02:06.260000",
      "content": "<p>how not to shake up:<br>\nsince there is only two hidden test image (a and b), it is easy to \"get\" \"these information\".</p>\n<p>what you need is not really CV/LB alignment. What you need is:<br>\n1) number of mask pixels in a and b (assume larger mask is private).<br>\n2) for each of a and b, triplet : (LB score, threshold used, number of detected pixels)</p>\n<p>(2) is a very good indicative of fpr, recall and precision values</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2258434,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-14T07:50:24.593000",
      "content": "<p>if we treat frame id 1,2,3 as class 1,2,3, we can train a fragement classifier.<br>\nthen we can see if the fragment are the same or not from the confusion matrix.</p>\n<p>we can use these fragment classfiier to probe if the public test and private test are smiliar to fragment1,2,3</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2254508,
      "author_name": "Aisuluu Ulan kyzy",
      "author_url": "",
      "post_date": "2023-05-11T03:53:58.797000",
      "content": "<p>so interesting find, i was curious abt it</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2273044,
      "author_name": "Feng Qilong",
      "author_url": "",
      "post_date": "2023-05-25T01:27:08.390000",
      "content": "<p>Another approach to train a denoiser<br>\n<a href=\"url\" target=\"_blank\">https://arxiv.org/abs/2205.11423</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2273203,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-05-25T04:31:16.873000",
          "content": "<p>thanks, since we only have one class, step #1 of figure.1 in the paper can be modified to predict  e.g. number of ink pixels, patch (32x32) classifiers, etc</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2273226,
              "author_name": "Feng Qilong",
              "author_url": "",
              "post_date": "2023-05-25T04:46:08.517000",
              "content": "<p>what I did was pretraining the 3d encoder using an architecture similar to the ink-id inkclassifier, it worked well.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2262130,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-16T18:43:14.170000",
      "content": "<p>from the table of experimental results:</p>\n<ol>\n<li>VIT like pvt-v2 is much better encoder than resnet</li>\n<li>pool as 2.5+1d  improved results </li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 2272044,
          "author_name": "WangXuC",
          "author_url": "",
          "post_date": "2023-05-24T09:01:43.167000",
          "content": "<p>What does 2.5d+1d mean？I just use 2.5d model</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2251914,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-09T18:06:27.107000",
      "content": "<p>a useful paper for my work:</p>\n<p>Noise2Noise: Learning Image Restoration without Clean Data<br>\n<a href=\"https://arxiv.org/pdf/1803.04189.pdf\" target=\"_blank\">https://arxiv.org/pdf/1803.04189.pdf</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2296668,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-06-12T03:57:47.863000",
      "content": "<p>speedup yout thresholding testing by submitting public test only<br>\n(warning !!!! remember to resubmit for both public+private test after finding the best threshold)</p>\n<pre><code>  submission = defaultdict()\n     fragment_id  valid_id:\n        d = read\n\n\n         fragment_id==:\n            rle =' '  #\n        : \n            probability = \n            count = \n             i, cfg  enumerate(configure):\n\n                print\n                net = cfg.\n                f = torch.load(cfg.checkpoint, map_location=lambda storage, loc: storage)\n                print(net.load)  # True\n                net.cuda\n                net.eval\n\n                p = infer</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2296582,
      "author_name": "Tamzid Ullah",
      "author_url": "",
      "post_date": "2023-06-12T01:25:58.027000",
      "content": "<p>Great post!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2279236,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-29T07:48:49.943000",
      "content": "<p>deleted message</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2277603,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2023-05-27T22:38:46.897000",
      "content": "<p>Is the training code released?, Anyone can share the link to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's training code for any model.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2277698,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-05-28T01:30:36.237000",
          "content": "<p>There is no training code. But it is easy to reproduce the results with your own training code</p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2262506\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2262506</a></p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2274502\" target=\"_blank\">https://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d/comments#2274502</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2261180,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-05-16T06:52:44.123000",
      "content": "<p>It turns out that you don't need decoder after all?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2258470,
      "author_name": "Cliche",
      "author_url": "",
      "post_date": "2023-05-14T08:20:42.397000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, your idea is very inspiring. I'm interested in attention pool you mentioned, do you have any suggested paper or blog about this technique?<br>\nMany thanks to your sharing!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2258558,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-05-14T10:16:03.667000",
          "content": "<p>google for video unet or spatiotemporal unet (e.g. satellite image segmentation with input over 12 months).<br>\nthe time axis can be our z axis here.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2258567,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-05-14T10:22:09",
              "content": "<p>e.g. <a href=\"https://zhuanlan.zhihu.com/p/421147308\" target=\"_blank\">https://zhuanlan.zhihu.com/p/421147308</a><br>\n<a href=\"https://github.com/VSainteuf/utae-paps\" target=\"_blank\">https://github.com/VSainteuf/utae-paps</a><br>\n<a href=\"https://github.com/VSainteuf/utae-paps/blob/main/gfx/utae.png\" target=\"_blank\">https://github.com/VSainteuf/utae-paps/blob/main/gfx/utae.png</a> </p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2259449,
              "author_name": "Cliche",
              "author_url": "",
              "post_date": "2023-05-15T02:39:01.400000",
              "content": "<p>thanks for sharing!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2256383,
      "author_name": "Lucas",
      "author_url": "",
      "post_date": "2023-05-12T12:56:20.120000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, </p>\n<p>Thanks for sharing your ideas and thoughts. I would like to try out something similar, and have a few questions: </p>\n<ul>\n<li>what exactly do you mean with \"no context\" for the 3DCnnEncoder? And what do the multiple 3DCnnEncoders represent, the same model trained on different folds?</li>\n<li>on what voxel size are you training the 3DCNNEncoder? I guess at least 16x16 otherwise you don't need the AdaptivePooling layer you mention in between the 3DCnnEncoder and the InkDetector to average the feature maps.</li>\n</ul>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2254055,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-10T16:16:37.307000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2286473,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-03T13:35:38.980000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2251189,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-09T06:41:45.253000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2251282,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-05-09T08:28:08.430000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2283490,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-01T09:38:08.497000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2282052,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-31T10:23:49.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2264637,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-18T15:38:48.303000",
      "content": "",
      "votes": -2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2250913": "[url=https://ibb.co/N6Yxp27][img]https://i.ibb.co/Fb76J4n/Selection-999-1974.png[/img][/url]\n![https://i.ibb.co/Fb76J4n/Selection-999-1974.png](https://i.ibb.co/Fb76J4n/Selection-999-1974.png)\n\n\nI conducted preliminary experiments:\n1. 3dCNN encoder is modified from [1],[1a]. replace Conv3d(stride=2) with Conv3d(stride=1) + AvgPool3d(stride=2). Use gobal pool instead of vectorizing feature volume before feeding into final linear ink classifier.\n\n2. Follow experiment protocols from [2]. divide fragments into upper and lower sub-fragments. validation = one of the sub-fragments. I have the same results as the paper[2], table.1: recall=0.41, FPR=0.051\n\n3. The trick is not to over-train your 3dCNN subvolume encoder. Since it doesn't use context, train for high recall (but high FPR is ok). We will reduce FPR using segmentation (which will has its own encoder and decoder) where we have larger context. Segmentation net is more like a denoiser and superresolution net.\n\nThe digram show are real results from my experiments for 3dcnn encoder, which are pretty good and surprises myself. the results include tta.\n\n\n[1]  https://github.com/educelab/ink-id  \n[1a] https://www.kaggle.com/code/jpposma/vesuvius-challenge-ink-detection-tutorial\n\n[2] EduceLab-Scrolls: Verifiable Recovery of Text from Herculaneum Papyri using X-ray CT  \nhttps://arxiv.org/pdf/2304.02084.pdf  \n\n\n----\nno sure if this would work, but we can:\n1. learn a noise generator to approximate results of 3dCNNencoder.\n2. used synthetic (+real) images to train segmetation. \n\nthen we have infinite train data. we can even train segmention to be alphabet detector, i.e. different class for each alphabet.\nif you undertsand the language in the scrolls, you can filter out invidual words ,etc ... \n\n\n\n----\n\n\nto be updated ....\nnotebook link:  to be updated ....\n\n(1) baseline : attentioned pool at encoder (resnet34d) unet, one fold validated on fragement\\_2a \nhttps://www.kaggle.com/code/hengck23/lb0-56-one-fold-tta-encoder-pooled-uet-resnet34d\n\n![https://i.ibb.co/HVzz5Rw/Selection-999-2032.png](https://i.ibb.co/HVzz5Rw/Selection-999-2032.png)\nhttps://docs.google.com/spreadsheets/d/17t_FFAJNQ7s23NeJ0tHwbBoGJYhqVcUfjAX3-3GNPaE\n\n![https://i.ibb.co/1bWxjLp/Selection-999-2048.png](https://i.ibb.co/1bWxjLp/Selection-999-2048.png)\n\n---\n\n##\"We extend our thanks to HP for providing the Z8-G4 Data Science Workstation, which empowered our deep learning experiments. The high computational power and large GPU memory enabled us to design our models swiftly.\"",
    "2266159": "how to pool with pos encoding (use conv3d as convolutional positional encoding in z direction)\n\n```\n\n\t\t#-- pool attention weight\n\t\tself.weight = nn.ModuleList([\n\t\t\tnn.Sequential(\n\t\t\t\tnn.Conv3d(dim, dim, kernel_size=3, padding=1),\n\t\t\t\tnn.ReLU(inplace=True),\n\t\t\t) for dim in encoder_dim\n\t\t])\n\n\tdef forward(self, batch):\n\t\tv = batch['volume']\n\t\tB,C,H,W = v.shape\n\t\tvv = [\n\t\t\tv[:,i:i+CFG.crop_depth] for i in [0,2,4,]\n\t\t]\n\t\tK = len(vv)\n\t\tx = torch.cat(vv,0)\n\n\t\t# ---------------------------------\n\t\t# encoder = self.encoder.forward_features(x)\n\t\tencoder = []\n\t\tx = self.encoder.conv1(x)\n\t\tx = self.encoder.bn1(x)\n\t\tx = self.encoder.act1(x)  ; encoder.append(x)\n\t\tx = F.avg_pool2d(x,kernel_size=2,stride=2)\n\t\tx = self.encoder.layer1(x); encoder.append(x)\n\t\tx = self.encoder.layer2(x); encoder.append(x)\n\t\tx = self.encoder.layer3(x); encoder.append(x)\n\t\tx = self.encoder.layer4(x); encoder.append(x)\n\t\t#print('encoder', [f.shape for f in encoder])\n\t\t# ---------------------------------\n\t\tfor i in range(len(encoder)):\n\t\t\te = encoder[i]\n\t\t\t_, c, h, w = e.shape\n\t\t\te = rearrange(e, '(K B) c h w -> B c h w K', K=K, B=B, h=h, w=w) #\n\t\t\tf = self.weight[i](e)\n\t\t\tw = F.softmax(f, -1)\n\t\t\te = (w * e).sum(-1)\n\t\t\tencoder[i] = e\n\n\n```",
    "2272189": "how to get lb 0.68 :\n-  augmentation: label noise\n-  model: stacked Unets\n- validation: fragement1\n- inference 4x rotate TTA\n\n\nexample notebook: \nhttps://www.kaggle.com/code/hengck23/lb0-68-one-fold-stacked-unet",
    "2260220": "Which to use: 3D, 2.5D or pool-2.5D? \n\nit actually depends how different is the z location of  target signal in train and hidden test data.\n\n1) Assume that both train,test has signal in about the same z location, then 2.5D is enough. 2d convolution is only invariant to changes in (x,y) location and not  invariant to channel (z). But target has same \"z image characteristics\", so 2d convolution is enough.\n\n2) Assume train,test has signal in very different z locations. Then you need 3d conv, invariant to changes in (x,y,z) location. But 3d convolution is costly. But this may not be true in the competition because the true resolution of the ink target is low (the inklabels ong images are enlarged version of smaller resolution)\n\n---\n\npool-2.5D is essentially a factorised version of 3d convolution. the pooling is 1d convolution in z direction. attention-pool uses dynamic weight in z direction. if also use fixed weights (via nn parameters) or equal weight(mean pool)\n\n",
    "2258613": "i have a feeling that the test images are rotated",
    "2301888": "\"人人有机会， 个个没把握\" (everyone has a chance but no one is sure to win)\n\nBe sure to read the data paper. \"Imagine\" how the train and public / private test fragment would look like ...\n\nestimate your predicted ink pixel precision and number of truth pixels carefully.\nis it more prediction  = more fp ?\nor more prdiction = more tp?\n\ndon't be shaken down by wrong threshold.\n\nbest threshold for public  = best threshold for private ???? \n\n- ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F83f199f9403710ea8f63bb15eba0bbf4%2FSelection_999(2223).png?generation=1686729121339846&alt=media)\n ",
    "2283935": "did i just find some magic?\nfrom @hughsando, https://www.kaggle.com/code/hughsando/visualize-the-3d-fragment-height\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc2c92845f59f35f8c4de32e5a23de2b6%2FSelection_999(2175).png?generation=1685634284049911&alt=media)\n\nheight image and label superimposed",
    "2277577": "the power of context!!!!\n@fengqilong check this! freeze and scale ...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F37dd82554ad4e12377b86900f16cc541%2FSelection_999(2146).png?generation=1685247621602263&alt=media)\n \n\n \n",
    "2263872": "paper:\nDeciphering Ancient Papyrus Texts: A Machine Learning Approach with Varied Architectures and Encoders\nhttps://www.researchgate.net/publication/370729553_Deciphering_Ancient_Papyrus_Texts_A_Machine_Learning_Approach_with_Varied_Architectures_and_Encoders",
    "2250988": "![https://i.ibb.co/ZVYqcMQ/Selection-999-1975.png](https://i.ibb.co/ZVYqcMQ/Selection-999-1975.png) ![https://i.ibb.co/mzBHh89/Selection-999-1976.png](https://i.ibb.co/mzBHh89/Selection-999-1976.png)\n \n4th fragment . or ?\n",
    "2265521": "very important details of the hidden test data:\n[1] EduceLab-Scrolls: Verifiable Recovery of Text from Herculaneum Papyri using X-ray CT\nhttps://arxiv.org/pdf/2304.02084.pdf\n\nFigure 1. \n\n\n\" Fragments 1 and 2 are from a philosophical work authored by Philodemus, as\nare many others in the Herculaneum collection. This book, titled On slander, comes from his greater work On Vices and the opposite virtues.\"\n\n\"Fragments 3 and 4 come from a scroll about Hellenistic Dynastic history\"\n\n---\n\"hidden, subsurface layers of the scroll fragments for which there is no ground truth.\"\n\n\n---\n\n\"in which a trained model is applied to hidden layers extracted from the X-ray CT under the exposed surface. In this case, there is no risk of memorization, so a model is trained on the entire training dataset, or all four fragment surfaces.\"\n \n---\n\nTable 2.\n```\nFragment Characters Recall FPR\n4 25 0.64 0.12\n```\n\"Importantly, the misidentified characters are not invented “whole cloth.” They are at least identified in the correct location, and could be accurately identified with marginal improvements in ink detection\"\n\n---\nTable 3.\n\n```\n\nSize (GB) Surface volume (GB)\nFragment 4 198 7.3\n```\n\nUse GB size to estimate image size",
    "2265133": "lb0.65 recipe \n\nencoder = resnext26d\nvalidation = fragement 1\nuse pool resnet unet \n\n(same configure is 0.54 for fragement 2b validation)",
    "2256006": "how to select the best depth crop:\n\n<a href=\"https://ibb.co/RzwbWtb\"><img src=\"https://i.ibb.co/wMmg18g/Selection-999-2027.png\" alt=\"Selection-999-2027\" border=\"0\"></a>\n\nas reference:\nvesuvius_2d_slide_exp002/vesuvius-models/Unet_fold1_best.pth\nfragment 1 validation 0.572",
    "2287963": "better threshold method?\n- ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F14c90169c92f5db5e2c02a2b23ccc237%2FSelection_999(2194).png?generation=1685967907793639&alt=media)",
    "2281576": "surprise!!!!!\ni am surprise that i can train will crop depth of 32\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fb3ae8a9d4cb24ac90713e416253ee775%2FSelection_999(2170).png?generation=1685498117144286&alt=media)\n",
    "2265586": "https://stackoverflow.com/questions/71690251/binary-cross-entropy-with-logits-weight-vs-pos-weight-what-are-the-differences\n\nThe pos_weight parameter allows you to balance the positive example thus controlling the tradeoff between recall and precision (see also). A detailed explanation can be found on this thread along with the explicit math expression. \n",
    "2265153": "it is probably difficult to augment positive samples ...\nbut it is easy to augment negative samples !!!!\n\nsince fbeta favours high precision, if you can create \"infinite negative samples\" , then ....\nfurther, there are so many samples in the external dataset at the scrollprize website",
    "2260347": "how not to shake up:\nsince there is only two hidden test image (a and b), it is easy to \"get\" \"these information\".\n\nwhat you need is not really CV/LB alignment. What you need is:\n1) number of mask pixels in a and b (assume larger mask is private).\n2) for each of a and b, triplet : (LB score, threshold used, number of detected pixels)\n\n(2) is a very good indicative of fpr, recall and precision values\n",
    "2258434": "if we treat frame id 1,2,3 as class 1,2,3, we can train a fragement classifier.\nthen we can see if the fragment are the same or not from the confusion matrix.\n\nwe can use these fragment classfiier to probe if the public test and private test are smiliar to fragment1,2,3",
    "2254508": "so interesting find, i was curious abt it",
    "2273044": "Another approach to train a denoiser\n[https://arxiv.org/abs/2205.11423](url)",
    "2262130": "from the table of experimental results:\n1. VIT like pvt-v2 is much better encoder than resnet\n2. pool as 2.5+1d  improved results ",
    "2251914": "a useful paper for my work:\n\nNoise2Noise: Learning Image Restoration without Clean Data\nhttps://arxiv.org/pdf/1803.04189.pdf\n\n",
    "2296668": "speedup yout thresholding testing by submitting public test only\n(warning !!!! remember to resubmit for both public+private test after finding the best threshold)\n\n```\n\n  submission = defaultdict(list)\n    for fragment_id in valid_id:\n        d = read_data1(fragment_id, z0=32-16, z1=32+16)\n \n        \n        if fragment_id=='b':\n            rle ='1 2'  #private\n        else: \n            probability = 0\n            count = 0\n            for i, cfg in enumerate(configure):\n         \n                print_cfg(cfg)\n                net = cfg.Net()\n                f = torch.load(cfg.checkpoint, map_location=lambda storage, loc: storage)\n                print(net.load_state_dict(f['state_dict'], strict=True))  # True\n                net.cuda()\n                net.eval()\n\n                p = infer_one_ms(net, d, cfg)\n                ...\n\n```\n",
    "2296582": "Great post!",
    "2279236": "deleted message",
    "2277603": "Is the training code released?, Anyone can share the link to @hengck23 's training code for any model.",
    "2261180": "It turns out that you don't need decoder after all?",
    "2258470": "Hi, @hengck23, your idea is very inspiring. I'm interested in attention pool you mentioned, do you have any suggested paper or blog about this technique?\nMany thanks to your sharing!",
    "2256383": "Hi @hengck23, \n\nThanks for sharing your ideas and thoughts. I would like to try out something similar, and have a few questions: \n\n- what exactly do you mean with \"no context\" for the 3DCnnEncoder? And what do the multiple 3DCnnEncoders represent, the same model trained on different folds?\n- on what voxel size are you training the 3DCNNEncoder? I guess at least 16x16 otherwise you don't need the AdaptivePooling layer you mention in between the 3DCnnEncoder and the InkDetector to average the feature maps.",
    "2254055": "Thanks for sharing @hengck23! \n\n> Follow experiment protocols from [2]. divide fragments into upper and lower sub-fragments. validation = one of the sub-fragments. I have the same results as the paper[2], table.1: recall=0.41, FPR=0.051\n\nI have few questions if you like to answer\n\n1) what is the sub-voxel shape as input to the model? did you use a 6-fold validation scheme (instead of 8 mentioned in the paper) ? \n2) how many epochs/iters train in your experiment and how long takes per epoch approx?\n3) the net output was a 2d mask or the center pixel xy-coords class? \n\n\n> This could be why some kagglers reported that if you use too many slices results are not as good?\n\nBTW I can confirm this as well, whenever tried to include more channels/slices metrics were worst",
    "2251189": "You will use a 3D encoder-decoder for segmentation?",
    "2283490": "",
    "2282052": "",
    "2264637": ""
  }
}