{
  "id": 663127,
  "title": "Affinity Feature Strenghtening Network",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/663127",
  "author_name": "Manas Choudhary",
  "post_date": "2025-12-16T15:11:24.535000",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I am trying to enforce topological information into UNETS, while doing this I came across this <a href=\"https://arxiv.org/pdf/2211.06578\" target=\"_blank\">paper</a>. It's an architecture in 2d, in which two parallel outputs are produced, one stream tries to predict affinities at various scales (in simple words, it predicts if some pixel belongs to the same category for all it's neighbors in all 8 directions (2d) at varying distances(scales)) and other the segmentation results, at each step they pass information to each other, this forces the model to learn neighbor relations not just a binary value for each voxel. In the image below you can see, how the model has learn  which neighbors are of same category and which are different.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2Ff2bb158a7660baf87af095ac047e1b14%2FScreenshot%202025-12-16%20203643.png?generation=1765897627031374&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2F5bd551bbfddc71366e5610a4318f6025%2FScreenshot%202025-12-16%20201152.png?generation=1765896177504490&amp;alt=media\" alt=\"img\">\nI implemented a 3d version of it, with two building blocks the UAFS and MAFS as given in the paper. If someone want's to build some models with these two blocks, here is the <a href=\"https://github.com/ManasChoudhary1/AFN-3D\" target=\"_blank\">repo</a>. There are only two changes that I have made in this from the original paper, first of all I removed the Gated merging of skipped connections, and in the mafs block I have made the weights of scales fixed unlike the original paper where they where they depended on the input.</p>\n<p>I also am thinking about training this network in an alternate manner, like in expectation maximization, This idea comes to my mind because affinity and segmentation both depend on each other, and losses of both might conflict.</p>\n<p>One major downside of this architecture that I caught early on is that it requires a lot of memory, even 128^3 inputs are infiesable, because for each input 26<em>(num of different scales in mafs = 3(standard))</em>(number of input voxels) voxels are outputed, which makes it equivalent to a standard unet with 512^3 input, so i am planning to try 2.5D or for a starting point 64^3.</p>",
  "messages": [
    {
      "id": 3377563,
      "postDate": "2025-12-16T15:11:24.537Z",
      "content": "<p>I am trying to enforce topological information into UNETS, while doing this I came across this <a href=\"https://arxiv.org/pdf/2211.06578\" target=\"_blank\">paper</a>. It's an architecture in 2d, in which two parallel outputs are produced, one stream tries to predict affinities at various scales (in simple words, it predicts if some pixel belongs to the same category for all it's neighbors in all 8 directions (2d) at varying distances(scales)) and other the segmentation results, at each step they pass information to each other, this forces the model to learn neighbor relations not just a binary value for each voxel. In the image below you can see, how the model has learn  which neighbors are of same category and which are different.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2Ff2bb158a7660baf87af095ac047e1b14%2FScreenshot%202025-12-16%20203643.png?generation=1765897627031374&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2F5bd551bbfddc71366e5610a4318f6025%2FScreenshot%202025-12-16%20201152.png?generation=1765896177504490&amp;alt=media\" alt=\"img\">\nI implemented a 3d version of it, with two building blocks the UAFS and MAFS as given in the paper. If someone want's to build some models with these two blocks, here is the <a href=\"https://github.com/ManasChoudhary1/AFN-3D\" target=\"_blank\">repo</a>. There are only two changes that I have made in this from the original paper, first of all I removed the Gated merging of skipped connections, and in the mafs block I have made the weights of scales fixed unlike the original paper where they where they depended on the input.</p>\n<p>I also am thinking about training this network in an alternate manner, like in expectation maximization, This idea comes to my mind because affinity and segmentation both depend on each other, and losses of both might conflict.</p>\n<p>One major downside of this architecture that I caught early on is that it requires a lot of memory, even 128^3 inputs are infiesable, because for each input 26<em>(num of different scales in mafs = 3(standard))</em>(number of input voxels) voxels are outputed, which makes it equivalent to a standard unet with 512^3 input, so i am planning to try 2.5D or for a starting point 64^3.</p>",
      "rawMarkdown": "I am trying to enforce topological information into UNETS, while doing this I came across this [paper](https://arxiv.org/pdf/2211.06578). It's an architecture in 2d, in which two parallel outputs are produced, one stream tries to predict affinities at various scales (in simple words, it predicts if some pixel belongs to the same category for all it's neighbors in all 8 directions (2d) at varying distances(scales)) and other the segmentation results, at each step they pass information to each other, this forces the model to learn neighbor relations not just a binary value for each voxel. In the image below you can see, how the model has learn  which neighbors are of same category and which are different.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2Ff2bb158a7660baf87af095ac047e1b14%2FScreenshot%202025-12-16%20203643.png?generation=1765897627031374&alt=media)\n\n\n![img](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2F5bd551bbfddc71366e5610a4318f6025%2FScreenshot%202025-12-16%20201152.png?generation=1765896177504490&alt=media)\nI implemented a 3d version of it, with two building blocks the UAFS and MAFS as given in the paper. If someone want's to build some models with these two blocks, here is the [repo](https://github.com/ManasChoudhary1/AFN-3D). There are only two changes that I have made in this from the original paper, first of all I removed the Gated merging of skipped connections, and in the mafs block I have made the weights of scales fixed unlike the original paper where they where they depended on the input.\n\nI also am thinking about training this network in an alternate manner, like in expectation maximization, This idea comes to my mind because affinity and segmentation both depend on each other, and losses of both might conflict.\n\nOne major downside of this architecture that I caught early on is that it requires a lot of memory, even 128^3 inputs are infiesable, because for each input 26*(num of different scales in mafs = 3(standard))*(number of input voxels) voxels are outputed, which makes it equivalent to a standard unet with 512^3 input, so i am planning to try 2.5D or for a starting point 64^3.",
      "votes": 5
    },
    {
      "id": 3378136,
      "postDate": "2025-12-17T15:30:30.053Z",
      "content": "<p>To combat the memory problem, I am thinking about removing the affinity prediction in the mafs head (78 predictions per voxel) to 8-16 dim embeddings which I can dot and try to predict near 1 for similar voxels and 0 for different ones. </p>",
      "rawMarkdown": "To combat the memory problem, I am thinking about removing the affinity prediction in the mafs head (78 predictions per voxel) to 8-16 dim embeddings which I can dot and try to predict near 1 for similar voxels and 0 for different ones. ",
      "votes": 1,
      "replies": [
        {
          "id": 3378174,
          "postDate": "2025-12-17T17:12:40.493Z",
          "content": "<p>somewhat similar to instance segmentation?</p>",
          "rawMarkdown": "somewhat similar to instance segmentation?",
          "votes": 1,
          "replies": [
            {
              "id": 3378213,
              "postDate": "2025-12-17T18:41:25.027Z",
              "content": "<p>Yes, the idea is somewhat similar, Trying to force the model to learn other things and not just fg-bg classification.</p>",
              "rawMarkdown": "Yes, the idea is somewhat similar, Trying to force the model to learn other things and not just fg-bg classification.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3377969,
      "postDate": "2025-12-17T08:51:10.540Z",
      "content": "<p>This looks promising, just keep in mind the annotation specificity, the sheats have almost constant thickness regardless the supposedly thickness observed in the image… </p>",
      "rawMarkdown": "This looks promising, just keep in mind the annotation specificity, the sheats have almost constant thickness regardless the supposedly thickness observed in the image... ",
      "votes": 1
    },
    {
      "id": 3378038,
      "postDate": "2025-12-17T12:25:02.330Z",
      "content": "<p>there is a trick to detect local line pattern.\ni show for 3x3 conv (and the same idea can be extended to 5x5, etc)</p>\n<pre><code>- consider  w = x3 slide window  prob map.\n- we binarise : w = w&gt; (other processing can be ranking   values,  detecting  pattern, etc )\n- our conv kernel is\n            kernel =[\n                  ,  ,  ,\n                ,  ,  ,\n                 , , \n            ]\n- now conv = (binary * kernel).() is  pattern code            \n-    current pixel x,y  is   conv==+,  is vertical .  conv ==+  is  horizontal , etc ...\n\nthis is also  trick used  detect coarse  angles  computer vision engineer.\n</code></pre>",
      "rawMarkdown": "there is a trick to detect local line pattern.\ni show for 3x3 conv (and the same idea can be extended to 5x5, etc)\n\n```\n- consider a w = 3x3 slide window of prob map.\n- we binarise it: w = w>0.5 (other processing can be ranking the 9 values, for detecting local pattern, etc )\n- our conv kernel is\n            kernel =[\n                  1,  2,  4,\n                128,  0,  8,\n                 64, 32, 16\n            ]\n- now conv = (binary * kernel).sum() is a pattern code            \n-  if the current pixel x,y value is 1 and conv==2+32, it is vertical line. if conv ==128+8. it is a horizontal line, etc ...\n\nthis is also a trick used to detect coarse line angles by computer vision engineer.\n\n```\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 3378060,
          "postDate": "2025-12-17T13:01:07.013Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3378068,
          "postDate": "2025-12-17T13:20:41.997Z",
          "content": "<p>This is a good trick, I am was also thinking of trying something alternative(like make the model predict such a code for each scale) in-place of the affinity prediction stream because it takes a lot of memory, and also compute time.</p>",
          "rawMarkdown": "This is a good trick, I am was also thinking of trying something alternative(like make the model predict such a code for each scale) in-place of the affinity prediction stream because it takes a lot of memory, and also compute time.",
          "votes": 1,
          "replies": [
            {
              "id": 3378384,
              "postDate": "2025-12-18T02:06:05Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a12f9f85bbc7d99044d340258d44844%2FSelection_1762.png?generation=1766023399606535&amp;alt=media\" alt=\"\"></p>\n<p>the trick to local convolutional regression is<br>\n````</p>\n<p>1) for a window say 5x5. binarise probability.<br>\n2) compute L2 or L1 loss (like Chamfer Distance) for all template pattern.<br>\n3) output the least loss pattern in prediction  </p>\n<p>alternatively,</p>\n<p>the loss compute be compute as\n1) predict p=binary output (e.g. via softmax or max … in a 5x5 window, there is only one value column or rowise)\n2) coordinate then = ((mx or my)*p).sum(x or y wise)\n4) loss for back propagate = L1 or l2 (coordinate, ground truth in x,y)</p>\n<p>```</p>\n<p><a href=\"https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221yD-OtPNCVRaQVskosaDPkhRdsuXatSSm%22%5D,%22action%22:%22open%22,%22userId%22:%22108640649683022616856%22,%22resourceKeys%22:%7B%7D%7D&amp;usp=sharing\" target=\"_blank\">https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221yD-OtPNCVRaQVskosaDPkhRdsuXatSSm%22%5D,%22action%22:%22open%22,%22userId%22:%22108640649683022616856%22,%22resourceKeys%22:%7B%7D%7D&amp;usp=sharing</a></p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a12f9f85bbc7d99044d340258d44844%2FSelection_1762.png?generation=1766023399606535&alt=media)\n\n\nthe trick to local convolutional regression is  \n````\n\n1) for a window say 5x5. binarise probability.  \n2) compute L2 or L1 loss (like Chamfer Distance) for all template pattern.  \n3) output the least loss pattern in prediction  \n\n\nalternatively,\n\nthe loss compute be compute as\n1) predict p=binary output (e.g. via softmax or max ... in a 5x5 window, there is only one value column or rowise)\n2) coordinate then = ((mx or my)*p).sum(x or y wise)\n4) loss for back propagate = L1 or l2 (coordinate, ground truth in x,y)\n\n```\n\nhttps://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221yD-OtPNCVRaQVskosaDPkhRdsuXatSSm%22%5D,%22action%22:%22open%22,%22userId%22:%22108640649683022616856%22,%22resourceKeys%22:%7B%7D%7D&usp=sharing",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3378136,
      "author_name": "Manas Choudhary",
      "author_url": "",
      "post_date": "2025-12-17T15:30:30.053000",
      "content": "<p>To combat the memory problem, I am thinking about removing the affinity prediction in the mafs head (78 predictions per voxel) to 8-16 dim embeddings which I can dot and try to predict near 1 for similar voxels and 0 for different ones. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3378174,
          "author_name": "Wang Zhiyao (王致尧)",
          "author_url": "",
          "post_date": "2025-12-17T17:12:40.493000",
          "content": "<p>somewhat similar to instance segmentation?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3378213,
              "author_name": "Manas Choudhary",
              "author_url": "",
              "post_date": "2025-12-17T18:41:25.027000",
              "content": "<p>Yes, the idea is somewhat similar, Trying to force the model to learn other things and not just fg-bg classification.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3377969,
      "author_name": "Jirka",
      "author_url": "",
      "post_date": "2025-12-17T08:51:10.540000",
      "content": "<p>This looks promising, just keep in mind the annotation specificity, the sheats have almost constant thickness regardless the supposedly thickness observed in the image… </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3378038,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-12-17T12:25:02.330000",
      "content": "<p>there is a trick to detect local line pattern.\ni show for 3x3 conv (and the same idea can be extended to 5x5, etc)</p>\n<pre><code>- consider  w = x3 slide window  prob map.\n- we binarise : w = w&gt; (other processing can be ranking   values,  detecting  pattern, etc )\n- our conv kernel is\n            kernel =[\n                  ,  ,  ,\n                ,  ,  ,\n                 , , \n            ]\n- now conv = (binary * kernel).() is  pattern code            \n-    current pixel x,y  is   conv==+,  is vertical .  conv ==+  is  horizontal , etc ...\n\nthis is also  trick used  detect coarse  angles  computer vision engineer.\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 3378060,
          "author_name": "",
          "author_url": "",
          "post_date": "2025-12-17T13:01:07.013000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3378068,
          "author_name": "Manas Choudhary",
          "author_url": "",
          "post_date": "2025-12-17T13:20:41.997000",
          "content": "<p>This is a good trick, I am was also thinking of trying something alternative(like make the model predict such a code for each scale) in-place of the affinity prediction stream because it takes a lot of memory, and also compute time.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3378384,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2025-12-18T02:06:05",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2a12f9f85bbc7d99044d340258d44844%2FSelection_1762.png?generation=1766023399606535&amp;alt=media\" alt=\"\"></p>\n<p>the trick to local convolutional regression is<br>\n````</p>\n<p>1) for a window say 5x5. binarise probability.<br>\n2) compute L2 or L1 loss (like Chamfer Distance) for all template pattern.<br>\n3) output the least loss pattern in prediction  </p>\n<p>alternatively,</p>\n<p>the loss compute be compute as\n1) predict p=binary output (e.g. via softmax or max … in a 5x5 window, there is only one value column or rowise)\n2) coordinate then = ((mx or my)*p).sum(x or y wise)\n4) loss for back propagate = L1 or l2 (coordinate, ground truth in x,y)</p>\n<p>```</p>\n<p><a href=\"https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221yD-OtPNCVRaQVskosaDPkhRdsuXatSSm%22%5D,%22action%22:%22open%22,%22userId%22:%22108640649683022616856%22,%22resourceKeys%22:%7B%7D%7D&amp;usp=sharing\" target=\"_blank\">https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221yD-OtPNCVRaQVskosaDPkhRdsuXatSSm%22%5D,%22action%22:%22open%22,%22userId%22:%22108640649683022616856%22,%22resourceKeys%22:%7B%7D%7D&amp;usp=sharing</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3377563": "I am trying to enforce topological information into UNETS, while doing this I came across this [paper](https://arxiv.org/pdf/2211.06578). It's an architecture in 2d, in which two parallel outputs are produced, one stream tries to predict affinities at various scales (in simple words, it predicts if some pixel belongs to the same category for all it's neighbors in all 8 directions (2d) at varying distances(scales)) and other the segmentation results, at each step they pass information to each other, this forces the model to learn neighbor relations not just a binary value for each voxel. In the image below you can see, how the model has learn  which neighbors are of same category and which are different.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2Ff2bb158a7660baf87af095ac047e1b14%2FScreenshot%202025-12-16%20203643.png?generation=1765897627031374&alt=media)\n\n\n![img](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2F5bd551bbfddc71366e5610a4318f6025%2FScreenshot%202025-12-16%20201152.png?generation=1765896177504490&alt=media)\nI implemented a 3d version of it, with two building blocks the UAFS and MAFS as given in the paper. If someone want's to build some models with these two blocks, here is the [repo](https://github.com/ManasChoudhary1/AFN-3D). There are only two changes that I have made in this from the original paper, first of all I removed the Gated merging of skipped connections, and in the mafs block I have made the weights of scales fixed unlike the original paper where they where they depended on the input.\n\nI also am thinking about training this network in an alternate manner, like in expectation maximization, This idea comes to my mind because affinity and segmentation both depend on each other, and losses of both might conflict.\n\nOne major downside of this architecture that I caught early on is that it requires a lot of memory, even 128^3 inputs are infiesable, because for each input 26*(num of different scales in mafs = 3(standard))*(number of input voxels) voxels are outputed, which makes it equivalent to a standard unet with 512^3 input, so i am planning to try 2.5D or for a starting point 64^3.",
    "3378136": "To combat the memory problem, I am thinking about removing the affinity prediction in the mafs head (78 predictions per voxel) to 8-16 dim embeddings which I can dot and try to predict near 1 for similar voxels and 0 for different ones. ",
    "3377969": "This looks promising, just keep in mind the annotation specificity, the sheats have almost constant thickness regardless the supposedly thickness observed in the image... ",
    "3378038": "there is a trick to detect local line pattern.\ni show for 3x3 conv (and the same idea can be extended to 5x5, etc)\n\n```\n- consider a w = 3x3 slide window of prob map.\n- we binarise it: w = w>0.5 (other processing can be ranking the 9 values, for detecting local pattern, etc )\n- our conv kernel is\n            kernel =[\n                  1,  2,  4,\n                128,  0,  8,\n                 64, 32, 16\n            ]\n- now conv = (binary * kernel).sum() is a pattern code            \n-  if the current pixel x,y value is 1 and conv==2+32, it is vertical line. if conv ==128+8. it is a horizontal line, etc ...\n\nthis is also a trick used to detect coarse line angles by computer vision engineer.\n\n```\n\n"
  }
}