{
  "id": 524500,
  "title": "Coordinate Pretraining Dataset",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/524500",
  "author_name": "Bartley",
  "post_date": "2024-08-06T16:06:19.631000",
  "votes": 52,
  "comment_count": 36,
  "views": 0,
  "content": "<p>I wanted to share a dataset that we have been using to pretrain models in this competition. This dataset has sped up the convergence time during training and is used in our current pipeline.</p>\n<p>The <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">Lumbar Coordinate Pretraining Dataset</a> was created by collecting T2 MRIs, T1 MRIs, and CT scans of the lower lumbar vertebrae, and then manually annotating the key points of the 5 lower lumbar vertebrae. All the sources are available under CC BY 4.0, and were labelled using the VGG Image Annotator <a href=\"https://www.robots.ox.ac.uk/~vgg/software/via/\" target=\"_blank\">here</a>. </p>\n<p>See a simple pipeline for pretraining with this dataset <a href=\"https://www.kaggle.com/code/brendanartley/lumbar-coordinate-pretraining-example\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fb982e9d30d93d1ee27b582cf861fe178%2Fdataset-cover.png?generation=1722960362507287&amp;alt=media\" alt=\"Image\"></p>\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": 2949392,
      "postDate": "2024-08-06T16:06:19.630Z",
      "content": "<p>I wanted to share a dataset that we have been using to pretrain models in this competition. This dataset has sped up the convergence time during training and is used in our current pipeline.</p>\n<p>The <a href=\"https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset\" target=\"_blank\">Lumbar Coordinate Pretraining Dataset</a> was created by collecting T2 MRIs, T1 MRIs, and CT scans of the lower lumbar vertebrae, and then manually annotating the key points of the 5 lower lumbar vertebrae. All the sources are available under CC BY 4.0, and were labelled using the VGG Image Annotator <a href=\"https://www.robots.ox.ac.uk/~vgg/software/via/\" target=\"_blank\">here</a>. </p>\n<p>See a simple pipeline for pretraining with this dataset <a href=\"https://www.kaggle.com/code/brendanartley/lumbar-coordinate-pretraining-example\" target=\"_blank\">here</a>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fb982e9d30d93d1ee27b582cf861fe178%2Fdataset-cover.png?generation=1722960362507287&amp;alt=media\" alt=\"Image\"></p>\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "I wanted to share a dataset that we have been using to pretrain models in this competition. This dataset has sped up the convergence time during training and is used in our current pipeline.\n\nThe [Lumbar Coordinate Pretraining Dataset](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset) was created by collecting T2 MRIs, T1 MRIs, and CT scans of the lower lumbar vertebrae, and then manually annotating the key points of the 5 lower lumbar vertebrae. All the sources are available under CC BY 4.0, and were labelled using the VGG Image Annotator [here](https://www.robots.ox.ac.uk/~vgg/software/via/). \n\nSee a simple pipeline for pretraining with this dataset [here](https://www.kaggle.com/code/brendanartley/lumbar-coordinate-pretraining-example).\n\n![Image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fb982e9d30d93d1ee27b582cf861fe178%2Fdataset-cover.png?generation=1722960362507287&alt=media)\n\n\n\nHappy Kaggling!",
      "votes": 52
    },
    {
      "id": 2957276,
      "postDate": "2024-08-13T01:12:50.947Z",
      "content": "<p>just see see.</p>",
      "rawMarkdown": "just see see.",
      "votes": 11
    },
    {
      "id": 2950672,
      "postDate": "2024-08-07T19:06:26.837Z",
      "content": "<p>thanks for the data. I have one question. are you annotating \"spinal canal stenosis\" points?<br>\nOr are you using your own set of rules (e.g. corners of spine vertebrae)?</p>\n<p>Actually kaggle label coords work well (and of course i expect additional external data would work even better). The trick to use kaggle data is:</p>\n<ol>\n<li>select a saggital t2 series.</li>\n<li>from the ground truth label coords csv, get the xy coords and z(instance number) of 5 \"spinal canal stenosis\" points for 5 levels </li>\n<li>compute mz = median of z </li>\n<li>during training random select a slice from the rangem z +/- limit, then set groud truth as (x,y). Now, each slice has 5 points.</li>\n</ol>\n<p>I use the above method to train. With proper image size, correct circle radius in target mask and good model to cover the context, convergent is very fast. My results are at: </p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approac\" target=\"_blank\">https://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approac</a></p>",
      "rawMarkdown": "thanks for the data. I have one question. are you annotating \"spinal canal stenosis\" points?\nOr are you using your own set of rules (e.g. corners of spine vertebrae)?\n\nActually kaggle label coords work well (and of course i expect additional external data would work even better). The trick to use kaggle data is:\n1. select a saggital t2 series.\n2. from the ground truth label coords csv, get the xy coords and z(instance number) of 5 \"spinal canal stenosis\" points for 5 levels \n3. compute mz = median of z \n4. during training random select a slice from the rangem z +/- limit, then set groud truth as (x,y). Now, each slice has 5 points.\n\nI use the above method to train. With proper image size, correct circle radius in target mask and good model to cover the context, convergent is very fast. My results are at: \n\nhttps://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approac",
      "votes": 7,
      "replies": [
        {
          "id": 2950681,
          "postDate": "2024-08-07T19:25:38.917Z",
          "content": "<p>Great job. Actually, you don't need even an explicit mask.</p>\n<p><a href=\"https://www.kaggle.com/code/sacuscreed/foraminal-unet-vit-train\" target=\"_blank\">https://www.kaggle.com/code/sacuscreed/foraminal-unet-vit-train</a></p>\n<p>I've replicated it for all the volumes, and I use it actively in my pipeline.</p>",
          "rawMarkdown": "Great job. Actually, you don't need even an explicit mask.\n\nhttps://www.kaggle.com/code/sacuscreed/foraminal-unet-vit-train\n\nI've replicated it for all the volumes, and I use it actively in my pipeline.",
          "votes": 4,
          "replies": [
            {
              "id": 2950691,
              "postDate": "2024-08-07T19:48:16.530Z",
              "content": "<p>yes you are correct. there are direct regression, heatmap based, coord/positinoal emcoding and the hybrid of these. the important consideration is the case of degeneration:</p>\n<ul>\n<li><p>e.g. for the input image, purposely erase (or occluded) a keypoint, e.g. using photoshop.<br>\nfor mask based method you will see no segmentation. direct regression may work. </p></li>\n<li><p>another degeneration is to copy the top point and paste on top of it (i.e. 5 keypoint + 1),<br>\nin this case will direct regression give the mean value (becuase it can only give one output)? heatmap segmentation will give two ouput and we can use post processing to select one.</p></li>\n</ul>\n<p>I am testing these cases now with different models, including heatmap and direct regression, etc…</p>",
              "rawMarkdown": "yes you are correct. there are direct regression, heatmap based, coord/positinoal emcoding and the hybrid of these. the important consideration is the case of degeneration:\n\n- e.g. for the input image, purposely erase (or occluded) a keypoint, e.g. using photoshop.\nfor mask based method you will see no segmentation. direct regression may work. \n\n- another degeneration is to copy the top point and paste on top of it (i.e. 5 keypoint + 1),\nin this case will direct regression give the mean value (becuase it can only give one output)? heatmap segmentation will give two ouput and we can use post processing to select one.\n\nI am testing these cases now with different models, including heatmap and direct regression, etc...",
              "votes": 3
            },
            {
              "id": 2951490,
              "postDate": "2024-08-08T17:05:29.420Z",
              "content": "<p>some axial volume have missing level. be careful if you are using direct regression on axial for level point</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbd810bb883d5a923b0797c270a379e90%2FSelection_999(5740).png?generation=1723136726974605&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "some axial volume have missing level. be careful if you are using direct regression on axial for level point\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbd810bb883d5a923b0797c270a379e90%2FSelection_999(5740).png?generation=1723136726974605&alt=media)",
              "votes": 5
            },
            {
              "id": 2957783,
              "postDate": "2024-08-13T12:33:40.803Z",
              "content": "<p>its actually true…..</p>",
              "rawMarkdown": "its actually true....."
            }
          ]
        },
        {
          "id": 2950693,
          "postDate": "2024-08-07T19:55:53.097Z",
          "content": "<p>Yes, that is correct, I did try and mimic the \"spinal canal stenosis\" labels. </p>\n<p>There may be slight differences in this dataset, as we do not know how the organizers defined the keypoints..</p>",
          "rawMarkdown": "Yes, that is correct, I did try and mimic the \"spinal canal stenosis\" labels. \n\nThere may be slight differences in this dataset, as we do not know how the organizers defined the keypoints..",
          "votes": 2,
          "replies": [
            {
              "id": 2950706,
              "postDate": "2024-08-07T20:12:27.070Z",
              "content": "<p>\" in this dataset, as we do not know how the organizers defined the keypoints.\"</p>\n<p>you can try to label kaggle dataset from scratch. </p>\n<p>then you can measure your human error/bias/adjustment with the ground truth in label coord csv file.<br>\nhere I think it is not so important. we just need to crop bigger (if you are using 2 stage approach) for the calssifier later.</p>",
              "rawMarkdown": "\" in this dataset, as we do not know how the organizers defined the keypoints.\"\n\nyou can try to label kaggle dataset from scratch. \n\nthen you can measure your human error/bias/adjustment with the ground truth in label coord csv file.\nhere I think it is not so important. we just need to crop bigger (if you are using 2 stage approach) for the calssifier later.",
              "votes": 4
            },
            {
              "id": 2950749,
              "postDate": "2024-08-07T21:32:42.657Z",
              "content": "<p>Good idea! We are actually using a <strong>single stage approach</strong> (but maybe we should shift to 2-stage haha).</p>\n<p>We tried adding the competition coordinates as an auxiliary task, but the performance did not improve.</p>",
              "rawMarkdown": "Good idea! We are actually using a **single stage approach** (but maybe we should shift to 2-stage haha).\n\n We tried adding the competition coordinates as an auxiliary task, but the performance did not improve.",
              "votes": 4
            },
            {
              "id": 2957741,
              "postDate": "2024-08-13T11:57:33.180Z",
              "rawMarkdown": "",
              "votes": -2,
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2962231,
      "postDate": "2024-08-17T10:46:03.160Z",
      "content": "<p>i confirm double detection is the result of different annotations.<br>\ni.e. even if you change your image encoder, you will still get the \"same results\"</p>\n<p>below are examples of ambiguous level labeling which have been discussed before in other posts.<br>\nand if I make a clean dataset without ambiguous labels, the trained model  would not give double detection in my experiments.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F625ebfca7037c1257302f9fd68b9a533%2FSelection_381.png?generation=1723891591633551&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "i confirm double detection is the result of different annotations.\ni.e. even if you change your image encoder, you will still get the \"same results\"\n\nbelow are examples of ambiguous level labeling which have been discussed before in other posts.\nand if I make a clean dataset without ambiguous labels, the trained model  would not give double detection in my experiments.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F625ebfca7037c1257302f9fd68b9a533%2FSelection_381.png?generation=1723891591633551&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 2962236,
          "postDate": "2024-08-17T10:56:22.320Z",
          "content": "<p>Now, we have assumed that one input image should have only one possible set of level points.<br>\nThis could be wrong. A model can (and maybe should) give a few sets of possible level points!</p>\n<p>It all depends on the distribution of error annotations for the grade class.<br>\ne.g. say for the mild class, \"wrong/ambiguous\" ground truth annotation only accounts for less than 1%, we can ignore them. If that is 50%, then the model can output multiple points per level. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fca2b0af0a10a8b462177c80f9caf33c7%2FSelection_380.png?generation=1723891967192459&amp;alt=media\" alt=\"\"></p>\n<p>what should you do next?</p>\n<ul>\n<li>you can carry out experiments to measure the cv and lb performance of single versus multiple-output models.</li>\n<li>decide if you want to use this for submission.</li>\n</ul>\n<p>obtaining the highest LB score is not about making the \"correct\" prediction. it is about making the prediction that best fits the hidden test distribution (i.e. we have to learn the error as well)</p>",
          "rawMarkdown": "Now, we have assumed that one input image should have only one possible set of level points.\nThis could be wrong. A model can (and maybe should) give a few sets of possible level points!\n\nIt all depends on the distribution of error annotations for the grade class.\ne.g. say for the mild class, \"wrong/ambiguous\" ground truth annotation only accounts for less than 1%, we can ignore them. If that is 50%, then the model can output multiple points per level. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fca2b0af0a10a8b462177c80f9caf33c7%2FSelection_380.png?generation=1723891967192459&alt=media)\n\nwhat should you do next?\n- you can carry out experiments to measure the cv and lb performance of single versus multiple-output models.\n- decide if you want to use this for submission.\n\nobtaining the highest LB score is not about making the \"correct\" prediction. it is about making the prediction that best fits the hidden test distribution (i.e. we have to learn the error as well)",
          "votes": 5
        }
      ]
    },
    {
      "id": 2961869,
      "postDate": "2024-08-17T01:52:07.330Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa156bc0d69b9fd919ac410408338a212%2FSelection_999(5861).png?generation=1723859260030201&amp;alt=media\" alt=\"\"></p>\n<p>now for spinal canal stenosis level point detection, i have quite a number of double detection (e.g. above). i find it puzzling.i suspect it is due to wrong label from kaggle dataset.<br>\n(i do see some wrong annotations, but it is difficult to check all kaggle annotations)</p>\n<p>my question:</p>\n<ul>\n<li>have you compared models  trained using your annotated dataset only (which i presume is 100% clean) and kaggle datset only?</li>\n<li>is there any differences in the results?</li>\n</ul>\n<hr>\n<p>on a side note, i have read some papers that can identify the training images that causes the model to make the current prediction in segmentation/classification. but i cannot recall the title of the paper. if anyone knows these paper or what keyword to google, please put it here. thanks! </p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa156bc0d69b9fd919ac410408338a212%2FSelection_999(5861).png?generation=1723859260030201&alt=media)\n\nnow for spinal canal stenosis level point detection, i have quite a number of double detection (e.g. above). i find it puzzling.i suspect it is due to wrong label from kaggle dataset.\n(i do see some wrong annotations, but it is difficult to check all kaggle annotations)\n\nmy question:\n- have you compared models  trained using your annotated dataset only (which i presume is 100% clean) and kaggle datset only?\n- is there any differences in the results?\n\n---\n\non a side note, i have read some papers that can identify the training images that causes the model to make the current prediction in segmentation/classification. but i cannot recall the title of the paper. if anyone knows these paper or what keyword to google, please put it here. thanks! ",
      "votes": 5,
      "replies": [
        {
          "id": 2961886,
          "postDate": "2024-08-17T02:32:40.907Z",
          "content": "<p>I went through all the sagittal images (T1 + T2) and corrected coordinates. I just uploaded these clean coordinates today in version 2 of the dataset. I only use the external data for pretraining at the moment, and therefore haven't directly compared results.</p>\n<p>In regards to double detections. I am seeing the same thing with a segmentation approach. I think our models are basically reproducing our discussion on the difficulty of identifying the bottom vertebrae <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/520279#2923289\" target=\"_blank\">here</a>, which is effecting predictions at other levels. Still not sure the best way to overcome this..</p>",
          "rawMarkdown": "I went through all the sagittal images (T1 + T2) and corrected coordinates. I just uploaded these clean coordinates today in version 2 of the dataset. I only use the external data for pretraining at the moment, and therefore haven't directly compared results.\n\nIn regards to double detections. I am seeing the same thing with a segmentation approach. I think our models are basically reproducing our discussion on the difficulty of identifying the bottom vertebrae [here](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/520279#2923289), which is effecting predictions at other levels. Still not sure the best way to overcome this..",
          "votes": 5,
          "replies": [
            {
              "id": 2961921,
              "postDate": "2024-08-17T03:26:20.490Z",
              "content": "<p>\"Still not sure the best way to overcome this..\"</p>\n<p>my approach:</p>\n<ol>\n<li>ouput a set of points y-sorted p1=(x1,y1) …. p2=(x6,y6) without considering the level. there are 6 points in my example.</li>\n<li>ouput level points only if there is no double detection. hence i will have l2,l3,l4,l5, i.e. 4 confirmed points and missing l1 points(red).</li>\n<li>alignment:<br>\ndistnce1 = [p1,p2,p3,p4]-[l2,l3,l4,l5]<br>\ndistnce2 = [p2,p3,p4,p5]-[l2,l3,l4,l5]<br>\ndistnce3 = [p3,p4,p5,p6]-[l2,l3,l4,l5]<br>\n…<br>\nchoose the least distance, i.e.<br>\nbest aligned is :<br>\n[p3,p4,p5,p6]-[l2,l3,l4,l5]<br>\nthen we can deduce l1=p2</li>\n</ol>\n<hr>\n<p>you can also google for \"multi Person pose estimation, overlapping\". most method are based on non-max suppression, or making multiple candidate graphs/poses and choosing the max-scored ones.</p>",
              "rawMarkdown": "\"Still not sure the best way to overcome this..\"\n\nmy approach:\n1. ouput a set of points y-sorted p1=(x1,y1) .... p2=(x6,y6) without considering the level. there are 6 points in my example.\n2. ouput level points only if there is no double detection. hence i will have l2,l3,l4,l5, i.e. 4 confirmed points and missing l1 points(red).\n3. alignment:\ndistnce1 = [p1,p2,p3,p4]-[l2,l3,l4,l5]\ndistnce2 = [p2,p3,p4,p5]-[l2,l3,l4,l5]\ndistnce3 = [p3,p4,p5,p6]-[l2,l3,l4,l5]\n...\nchoose the least distance, i.e.\nbest aligned is :\n [p3,p4,p5,p6]-[l2,l3,l4,l5]\nthen we can deduce l1=p2\n\n---\nyou can also google for \"multi Person pose estimation, overlapping\". most method are based on non-max suppression, or making multiple candidate graphs/poses and choosing the max-scored ones.",
              "votes": 5
            }
          ]
        }
      ]
    },
    {
      "id": 2970047,
      "postDate": "2024-08-25T16:48:16.690Z",
      "content": "<p>i mentioned before wrong label is cuasing wrong detection results.<br>\nin particular, if you have missing point, this could be the reason.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0cbd474368efac9e1a29b238805b6ad1%2FSelection_999(5950).png?generation=1724604475838911&amp;alt=media\" alt=\"\"></p>\n<p>top: validation image (at the predict, it miss the red point)<br>\nbottom: training image (at the truth, it has wrong label red point)</p>\n<p>both image are similar. the wrong label cuase the missing detection</p>",
      "rawMarkdown": "i mentioned before wrong label is cuasing wrong detection results.\nin particular, if you have missing point, this could be the reason.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0cbd474368efac9e1a29b238805b6ad1%2FSelection_999(5950).png?generation=1724604475838911&alt=media)\n\ntop: validation image (at the predict, it miss the red point)\nbottom: training image (at the truth, it has wrong label red point)\n\nboth image are similar. the wrong label cuase the missing detection",
      "votes": 4
    },
    {
      "id": 2964094,
      "postDate": "2024-08-19T14:29:23.673Z",
      "content": "<p>instead of heatmap detection and direct regression, here comes the third method:</p>\n<ul>\n<li>softmax over all possible y-coordinate (and x-coordinate)<br>\nit doesn't have heatmap and takes care of double detection </li>\n</ul>\n<hr>\n<p>left to right: input image, prediction, truth</p>\n<p>normal prediction</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe10d284129505ca02212c93a6d5a0de2%2FSelection_999(5884).png?generation=1724077798492469&amp;alt=media\" alt=\"\"></p>\n<p>double prediction<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8b0c4f01d484b58052aa6162b4f51550%2FSelection_999(5883).png?generation=1724077760199360&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbf593db3d80fa3a65f202db332df96f6%2FSelection_999(5885).png?generation=1724077717548340&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "instead of heatmap detection and direct regression, here comes the third method:\n- softmax over all possible y-coordinate (and x-coordinate)\nit doesn't have heatmap and takes care of double detection \n\n\n---\nleft to right: input image, prediction, truth\n\n\nnormal prediction\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe10d284129505ca02212c93a6d5a0de2%2FSelection_999(5884).png?generation=1724077798492469&alt=media)\n\ndouble prediction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8b0c4f01d484b58052aa6162b4f51550%2FSelection_999(5883).png?generation=1724077760199360&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbf593db3d80fa3a65f202db332df96f6%2FSelection_999(5885).png?generation=1724077717548340&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 2964182,
          "postDate": "2024-08-19T15:57:24.100Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2964435,
          "postDate": "2024-08-19T19:10:52.760Z",
          "content": "<p>in theory, you can project the image y into 3d space by affine trasnform.<br>\nso you can use all 3 views and softmax over \"the z of common 3d world\"</p>",
          "rawMarkdown": "in theory, you can project the image y into 3d space by affine trasnform.\nso you can use all 3 views and softmax over \"the z of common 3d world\"",
          "votes": 1,
          "replies": [
            {
              "id": 2966896,
              "postDate": "2024-08-22T10:43:22.707Z",
              "content": "<p>Hi. I've deleted my las comment because I've realized I wasn't understanding correctly your insights. The thing is that recently I made some changes to my segmentation model. And, although still I need to check why, I've ended to this kind of predicted masks, the usual were circular:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F4de5f59a9a1e3c7577afbda2de9dbb70%2Faxial.png?generation=1724323314045588&amp;alt=media\" alt=\"\"></p>\n<p>Perhaps I wrote wrong some sum axis and I accidentally ended to the third possibility you commented? Now I'm not sure if I should \"fix\" it or test it.</p>",
              "rawMarkdown": "Hi. I've deleted my las comment because I've realized I wasn't understanding correctly your insights. The thing is that recently I made some changes to my segmentation model. And, although still I need to check why, I've ended to this kind of predicted masks, the usual were circular:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F4de5f59a9a1e3c7577afbda2de9dbb70%2Faxial.png?generation=1724323314045588&alt=media)\n\nPerhaps I wrote wrong some sum axis and I accidentally ended to the third possibility you commented? Now I'm not sure if I should \"fix\" it or test it.",
              "votes": 2
            },
            {
              "id": 2966903,
              "postDate": "2024-08-22T10:55:44.630Z",
              "content": "<p>my inplementation is:</p>\n<pre><code>truth_y = Bx5 (B=batch size , 5= 5 level points)\nmodel ouput:\nfeature = BxCxHxW \nfeature_y = BxCxH (pool over W) \nlogit_y = Bx5xH  \n\nloss = F.cross_entropy(logit_y,truth_y)\n</code></pre>\n<p>but now i have a better implement, that has now 99% accuracy. i will make a post later.<br>\nit is based on <br>\n\"Numerical Coordinate Regression with Convolutional Neural Networks\"<br>\n<a href=\"https://arxiv.org/abs/1801.07372\" target=\"_blank\">https://arxiv.org/abs/1801.07372</a></p>\n<pre><code>in summary:\n\ntruth_xy = (size , =  lavel points)\n\ntruth_map_xy= \nlogit_xy = prob_xy =   \ncoord_map = meshgrid(over WH)\nxy = (prob_xy * coord_map).sum(\n\n= (prob_xy,truth_map_xy)\nMSE loss(xy,truth_xy)\n\nyou output xy heatmap !!!\n</code></pre>",
              "rawMarkdown": "my inplementation is:\n\n```\ntruth_y = Bx5 (B=batch size , 5= 5 level points)\nmodel ouput:\nfeature = BxCxHxW \nfeature_y = BxCxH (pool over W) \nlogit_y = Bx5xH  \n\nloss = F.cross_entropy(logit_y,truth_y)\n\n\n```\n\nbut now i have a better implement, that has now 99% accuracy. i will make a post later.\nit is based on \n\"Numerical Coordinate Regression with Convolutional Neural Networks\"\nhttps://arxiv.org/abs/1801.07372\n\n```\nin summary:\n\ntruth_xy = Bx5x2 (B=batch size , 5= 5 lavel points)\n#make gaussian map from point\ntruth_map_xy= Bx5xHxW\n\nlogit_xy = Bx5xHxW\nprob_xy = Bx5xHxW   #softmax over xy (i.e prob_xy.sum(dim=(2,3)) = all ones\ncoord_map = meshgrid(over WH)\nxy = (prob_xy * coord_map).sum(dim=(2,3)\n\n#back prob\nJD_divergence_loss = (prob_xy,truth_map_xy)\nMSE loss(xy,truth_xy)\n\nyou output both xy and heatmap !!!\n\n```",
              "votes": 4
            },
            {
              "id": 2966934,
              "postDate": "2024-08-22T11:22:48.250Z",
              "content": "<p>Great. About my code, I'm definitely generating the wrong ideal heatmaps, that's why.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fe9582aa37ea201fe91e5c604b1138e06%2Fheatmap.png?generation=1724325944835986&amp;alt=media\" alt=\"\"></p>\n<p>Fixed:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F05c2c674cca6ed2ece0c9a5c01071d78%2Fheatmap.png?generation=1724327144714933&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Great. About my code, I'm definitely generating the wrong ideal heatmaps, that's why.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fe9582aa37ea201fe91e5c604b1138e06%2Fheatmap.png?generation=1724325944835986&alt=media)\n\nFixed:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F05c2c674cca6ed2ece0c9a5c01071d78%2Fheatmap.png?generation=1724327144714933&alt=media)"
            }
          ]
        }
      ]
    },
    {
      "id": 2950098,
      "postDate": "2024-08-07T08:21:14.277Z",
      "content": "<p>Can we use this in our solution if they're manually annotated? </p>",
      "rawMarkdown": "Can we use this in our solution if they're manually annotated? ",
      "votes": 1,
      "replies": [
        {
          "id": 2950260,
          "postDate": "2024-08-07T12:14:47.177Z",
          "content": "<p>May be I'm wrong. But I'm pretty sure you can use anything that's public.</p>",
          "rawMarkdown": "May be I'm wrong. But I'm pretty sure you can use anything that's public.",
          "votes": 2
        },
        {
          "id": 2950316,
          "postDate": "2024-08-07T13:42:20.133Z",
          "content": "<p>Yep, I believe this data should be good for use. I believe that the rules only apply to pseudolabelling the LB data..</p>\n<p><em>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records</em></p>",
          "rawMarkdown": "Yep, I believe this data should be good for use. I believe that the rules only apply to pseudolabelling the LB data..\n\n*Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records*",
          "votes": 3,
          "replies": [
            {
              "id": 2950520,
              "postDate": "2024-08-07T16:23:13.980Z",
              "content": "<p>So, just out of interest would you use this for RoI or spine curve? </p>",
              "rawMarkdown": "So, just out of interest would you use this for RoI or spine curve? ",
              "votes": 2
            },
            {
              "id": 2950554,
              "postDate": "2024-08-07T16:54:43.030Z",
              "content": "<p>Up to you! </p>\n<p>For now we are only pretrain on the exact coordinate values. We use this pipeline <a href=\"https://www.kaggle.com/code/brendanartley/lumbar-coordinate-pretraining-example\" target=\"_blank\">here</a> with heaver augmentations and bigger backbones. </p>",
              "rawMarkdown": "Up to you! \n\nFor now we are only pretrain on the exact coordinate values. We use this pipeline [here](https://www.kaggle.com/code/brendanartley/lumbar-coordinate-pretraining-example) with heaver augmentations and bigger backbones. ",
              "votes": 2
            },
            {
              "id": 2950559,
              "postDate": "2024-08-07T16:56:04.083Z",
              "content": "<p>Interesting thank you </p>",
              "rawMarkdown": "Interesting thank you ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2970315,
      "postDate": "2024-08-26T01:49:38.070Z",
      "content": "<p>there is yet another way to formulate the problem.<br>\nbecuase of different annotation, we predict the top2 solution (instead of the best solution).<br>\nwe acknowledge that it is not one-to-one regression, but a possible one-to-two.</p>\n<p>here is how kaggle did it in the past:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdcfe734eb2d6f3457c6a295687958ba8%2FSelection_999(5952).png?generation=1724636755026822&amp;alt=media\" alt=\"\"></p>\n<p>implementation of loss<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F85e5fb18fe78f7faef774da3ffa7920f%2FSelection_999(5953).png?generation=1724636778270518&amp;alt=media\" alt=\"\"></p>\n<pre><code> summy:\n- predict  solution\n- compute  loss\n- back propagative ONLY  best loss\n- we also need  predict  probability  correctness    solution.  this ground truth   selected best index above (online labeling)\n\n\n is like  DETR (transformer object detection paper) where you have  online matching algorithm  find  ground truth   predicted box label)\n</code></pre>",
      "rawMarkdown": "there is yet another way to formulate the problem.\nbecuase of different annotation, we predict the top2 solution (instead of the best solution).\nwe acknowledge that it is not one-to-one regression, but a possible one-to-two.\n\nhere is how kaggle did it in the past:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdcfe734eb2d6f3457c6a295687958ba8%2FSelection_999(5952).png?generation=1724636755026822&alt=media)\n\nimplementation of loss\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F85e5fb18fe78f7faef774da3ffa7920f%2FSelection_999(5953).png?generation=1724636778270518&alt=media)\n\n```\nin summy:\n- predict 2 solution\n- compute two loss\n- back propagative ONLY the best loss\n- we also need to predict the probability of correctness of the two solution. set this ground truth to the selected best index above (online labeling)\n\n\nit is like the DETR (transformer object detection paper) where you have an online matching algorithm to find the ground truth of the predicted box label)\n```",
      "votes": 2,
      "replies": [
        {
          "id": 2970673,
          "postDate": "2024-08-26T11:32:27.703Z",
          "content": "<h1>Potential issues:</h1>\n<h1>1. Overfitting to Best Predictions: If the model consistently selects the same type of prediction as the \"best\" one, it might develop a bias toward certain patterns or features, leading to overfitting. This is particularly concerning if there's limited diversity between the two solutions.</h1>\n<h1>2. Noise Sensitivity: The process of selecting the best prediction might introduce noise, especially when the difference between losses is small or the data is noisy. This could result in instability during training and reduced model performance.</h1>",
          "rawMarkdown": "# Potential issues:\n# 1. Overfitting to Best Predictions: If the model consistently selects the same type of prediction as the \"best\" one, it might develop a bias toward certain patterns or features, leading to overfitting. This is particularly concerning if there's limited diversity between the two solutions.\n# 2. Noise Sensitivity: The process of selecting the best prediction might introduce noise, especially when the difference between losses is small or the data is noisy. This could result in instability during training and reduced model performance.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2960130,
      "postDate": "2024-08-15T15:55:59.180Z",
      "content": "<p>Hello, could you help me with a few questions? I would greatly appreciate it. Are you using 3D convolution? Also, do you have any advice on processing sagittal and axial images?</p>",
      "rawMarkdown": "Hello, could you help me with a few questions? I would greatly appreciate it. Are you using 3D convolution? Also, do you have any advice on processing sagittal and axial images?",
      "replies": [
        {
          "id": 2960140,
          "postDate": "2024-08-15T16:05:19.647Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/theexaltedone\" target=\"_blank\">@theexaltedone</a>. I would recommend reading through the discussions and sample notebooks, as lots of good info has been shared. Here are a couple relevant ones.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/523859\" target=\"_blank\">Sagittal Processing</a> by llleeeoooh</li>\n<li><a href=\"https://www.kaggle.com/code/vaillant/cross-reference-images-in-different-mri-planes\" target=\"_blank\">Cross-Reference Images in Different MRI Planes</a> by vaillant</li>\n</ul>",
          "rawMarkdown": "Hi @theexaltedone. I would recommend reading through the discussions and sample notebooks, as lots of good info has been shared. Here are a couple relevant ones.\n\n- [Sagittal Processing](https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/523859) by llleeeoooh\n- [Cross-Reference Images in Different MRI Planes](https://www.kaggle.com/code/vaillant/cross-reference-images-in-different-mri-planes) by vaillant",
          "votes": 3,
          "replies": [
            {
              "id": 2960162,
              "postDate": "2024-08-15T16:22:15.500Z",
              "content": "<p>So how will I handle missing data?  conditions that are missing…</p>",
              "rawMarkdown": "So how will I handle missing data?  conditions that are missing...\n"
            }
          ]
        }
      ]
    },
    {
      "id": 2976910,
      "postDate": "2024-09-02T12:33:05.030Z",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you.",
      "votes": 1
    },
    {
      "id": 2965445,
      "postDate": "2024-08-20T21:38:49.153Z",
      "content": "<p>thanks, really cool stuff in here :D</p>",
      "rawMarkdown": "thanks, really cool stuff in here :D",
      "votes": 1
    },
    {
      "id": 2949917,
      "postDate": "2024-08-07T04:25:13.933Z",
      "content": "<p>Oh thank you 🙏 </p>",
      "rawMarkdown": "Oh thank you 🙏 ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2957276,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-08-13T01:12:50.947000",
      "content": "<p>just see see.</p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 2950672,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-08-07T19:06:26.837000",
      "content": "<p>thanks for the data. I have one question. are you annotating \"spinal canal stenosis\" points?<br>\nOr are you using your own set of rules (e.g. corners of spine vertebrae)?</p>\n<p>Actually kaggle label coords work well (and of course i expect additional external data would work even better). The trick to use kaggle data is:</p>\n<ol>\n<li>select a saggital t2 series.</li>\n<li>from the ground truth label coords csv, get the xy coords and z(instance number) of 5 \"spinal canal stenosis\" points for 5 levels </li>\n<li>compute mz = median of z </li>\n<li>during training random select a slice from the rangem z +/- limit, then set groud truth as (x,y). Now, each slice has 5 points.</li>\n</ol>\n<p>I use the above method to train. With proper image size, correct circle radius in target mask and good model to cover the context, convergent is very fast. My results are at: </p>\n<p><a href=\"https://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approac\" target=\"_blank\">https://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approac</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 2950681,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2024-08-07T19:25:38.917000",
          "content": "<p>Great job. Actually, you don't need even an explicit mask.</p>\n<p><a href=\"https://www.kaggle.com/code/sacuscreed/foraminal-unet-vit-train\" target=\"_blank\">https://www.kaggle.com/code/sacuscreed/foraminal-unet-vit-train</a></p>\n<p>I've replicated it for all the volumes, and I use it actively in my pipeline.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2950691,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-08-07T19:48:16.530000",
              "content": "<p>yes you are correct. there are direct regression, heatmap based, coord/positinoal emcoding and the hybrid of these. the important consideration is the case of degeneration:</p>\n<ul>\n<li><p>e.g. for the input image, purposely erase (or occluded) a keypoint, e.g. using photoshop.<br>\nfor mask based method you will see no segmentation. direct regression may work. </p></li>\n<li><p>another degeneration is to copy the top point and paste on top of it (i.e. 5 keypoint + 1),<br>\nin this case will direct regression give the mean value (becuase it can only give one output)? heatmap segmentation will give two ouput and we can use post processing to select one.</p></li>\n</ul>\n<p>I am testing these cases now with different models, including heatmap and direct regression, etc…</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2951490,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-08-08T17:05:29.420000",
              "content": "<p>some axial volume have missing level. be careful if you are using direct regression on axial for level point</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbd810bb883d5a923b0797c270a379e90%2FSelection_999(5740).png?generation=1723136726974605&amp;alt=media\" alt=\"\"></p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2957783,
              "author_name": "Exalted Joseph",
              "author_url": "",
              "post_date": "2024-08-13T12:33:40.803000",
              "content": "<p>its actually true…..</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2950693,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-08-07T19:55:53.097000",
          "content": "<p>Yes, that is correct, I did try and mimic the \"spinal canal stenosis\" labels. </p>\n<p>There may be slight differences in this dataset, as we do not know how the organizers defined the keypoints..</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2950706,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-08-07T20:12:27.070000",
              "content": "<p>\" in this dataset, as we do not know how the organizers defined the keypoints.\"</p>\n<p>you can try to label kaggle dataset from scratch. </p>\n<p>then you can measure your human error/bias/adjustment with the ground truth in label coord csv file.<br>\nhere I think it is not so important. we just need to crop bigger (if you are using 2 stage approach) for the calssifier later.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2950749,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-08-07T21:32:42.657000",
              "content": "<p>Good idea! We are actually using a <strong>single stage approach</strong> (but maybe we should shift to 2-stage haha).</p>\n<p>We tried adding the competition coordinates as an auxiliary task, but the performance did not improve.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2957741,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-08-13T11:57:33.180000",
              "content": "",
              "votes": -2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2962231,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-08-17T10:46:03.160000",
      "content": "<p>i confirm double detection is the result of different annotations.<br>\ni.e. even if you change your image encoder, you will still get the \"same results\"</p>\n<p>below are examples of ambiguous level labeling which have been discussed before in other posts.<br>\nand if I make a clean dataset without ambiguous labels, the trained model  would not give double detection in my experiments.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F625ebfca7037c1257302f9fd68b9a533%2FSelection_381.png?generation=1723891591633551&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2962236,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-08-17T10:56:22.320000",
          "content": "<p>Now, we have assumed that one input image should have only one possible set of level points.<br>\nThis could be wrong. A model can (and maybe should) give a few sets of possible level points!</p>\n<p>It all depends on the distribution of error annotations for the grade class.<br>\ne.g. say for the mild class, \"wrong/ambiguous\" ground truth annotation only accounts for less than 1%, we can ignore them. If that is 50%, then the model can output multiple points per level. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fca2b0af0a10a8b462177c80f9caf33c7%2FSelection_380.png?generation=1723891967192459&amp;alt=media\" alt=\"\"></p>\n<p>what should you do next?</p>\n<ul>\n<li>you can carry out experiments to measure the cv and lb performance of single versus multiple-output models.</li>\n<li>decide if you want to use this for submission.</li>\n</ul>\n<p>obtaining the highest LB score is not about making the \"correct\" prediction. it is about making the prediction that best fits the hidden test distribution (i.e. we have to learn the error as well)</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 2961869,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-08-17T01:52:07.330000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa156bc0d69b9fd919ac410408338a212%2FSelection_999(5861).png?generation=1723859260030201&amp;alt=media\" alt=\"\"></p>\n<p>now for spinal canal stenosis level point detection, i have quite a number of double detection (e.g. above). i find it puzzling.i suspect it is due to wrong label from kaggle dataset.<br>\n(i do see some wrong annotations, but it is difficult to check all kaggle annotations)</p>\n<p>my question:</p>\n<ul>\n<li>have you compared models  trained using your annotated dataset only (which i presume is 100% clean) and kaggle datset only?</li>\n<li>is there any differences in the results?</li>\n</ul>\n<hr>\n<p>on a side note, i have read some papers that can identify the training images that causes the model to make the current prediction in segmentation/classification. but i cannot recall the title of the paper. if anyone knows these paper or what keyword to google, please put it here. thanks! </p>",
      "votes": 5,
      "replies": [
        {
          "id": 2961886,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-08-17T02:32:40.907000",
          "content": "<p>I went through all the sagittal images (T1 + T2) and corrected coordinates. I just uploaded these clean coordinates today in version 2 of the dataset. I only use the external data for pretraining at the moment, and therefore haven't directly compared results.</p>\n<p>In regards to double detections. I am seeing the same thing with a segmentation approach. I think our models are basically reproducing our discussion on the difficulty of identifying the bottom vertebrae <a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/520279#2923289\" target=\"_blank\">here</a>, which is effecting predictions at other levels. Still not sure the best way to overcome this..</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2961921,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-08-17T03:26:20.490000",
              "content": "<p>\"Still not sure the best way to overcome this..\"</p>\n<p>my approach:</p>\n<ol>\n<li>ouput a set of points y-sorted p1=(x1,y1) …. p2=(x6,y6) without considering the level. there are 6 points in my example.</li>\n<li>ouput level points only if there is no double detection. hence i will have l2,l3,l4,l5, i.e. 4 confirmed points and missing l1 points(red).</li>\n<li>alignment:<br>\ndistnce1 = [p1,p2,p3,p4]-[l2,l3,l4,l5]<br>\ndistnce2 = [p2,p3,p4,p5]-[l2,l3,l4,l5]<br>\ndistnce3 = [p3,p4,p5,p6]-[l2,l3,l4,l5]<br>\n…<br>\nchoose the least distance, i.e.<br>\nbest aligned is :<br>\n[p3,p4,p5,p6]-[l2,l3,l4,l5]<br>\nthen we can deduce l1=p2</li>\n</ol>\n<hr>\n<p>you can also google for \"multi Person pose estimation, overlapping\". most method are based on non-max suppression, or making multiple candidate graphs/poses and choosing the max-scored ones.</p>",
              "votes": 5,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2970047,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-08-25T16:48:16.690000",
      "content": "<p>i mentioned before wrong label is cuasing wrong detection results.<br>\nin particular, if you have missing point, this could be the reason.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0cbd474368efac9e1a29b238805b6ad1%2FSelection_999(5950).png?generation=1724604475838911&amp;alt=media\" alt=\"\"></p>\n<p>top: validation image (at the predict, it miss the red point)<br>\nbottom: training image (at the truth, it has wrong label red point)</p>\n<p>both image are similar. the wrong label cuase the missing detection</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2964094,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-08-19T14:29:23.673000",
      "content": "<p>instead of heatmap detection and direct regression, here comes the third method:</p>\n<ul>\n<li>softmax over all possible y-coordinate (and x-coordinate)<br>\nit doesn't have heatmap and takes care of double detection </li>\n</ul>\n<hr>\n<p>left to right: input image, prediction, truth</p>\n<p>normal prediction</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe10d284129505ca02212c93a6d5a0de2%2FSelection_999(5884).png?generation=1724077798492469&amp;alt=media\" alt=\"\"></p>\n<p>double prediction<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8b0c4f01d484b58052aa6162b4f51550%2FSelection_999(5883).png?generation=1724077760199360&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbf593db3d80fa3a65f202db332df96f6%2FSelection_999(5885).png?generation=1724077717548340&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 2964182,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-08-19T15:57:24.100000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2964435,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-08-19T19:10:52.760000",
          "content": "<p>in theory, you can project the image y into 3d space by affine trasnform.<br>\nso you can use all 3 views and softmax over \"the z of common 3d world\"</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2966896,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-08-22T10:43:22.707000",
              "content": "<p>Hi. I've deleted my las comment because I've realized I wasn't understanding correctly your insights. The thing is that recently I made some changes to my segmentation model. And, although still I need to check why, I've ended to this kind of predicted masks, the usual were circular:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F4de5f59a9a1e3c7577afbda2de9dbb70%2Faxial.png?generation=1724323314045588&amp;alt=media\" alt=\"\"></p>\n<p>Perhaps I wrote wrong some sum axis and I accidentally ended to the third possibility you commented? Now I'm not sure if I should \"fix\" it or test it.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2966903,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-08-22T10:55:44.630000",
              "content": "<p>my inplementation is:</p>\n<pre><code>truth_y = Bx5 (B=batch size , 5= 5 level points)\nmodel ouput:\nfeature = BxCxHxW \nfeature_y = BxCxH (pool over W) \nlogit_y = Bx5xH  \n\nloss = F.cross_entropy(logit_y,truth_y)\n</code></pre>\n<p>but now i have a better implement, that has now 99% accuracy. i will make a post later.<br>\nit is based on <br>\n\"Numerical Coordinate Regression with Convolutional Neural Networks\"<br>\n<a href=\"https://arxiv.org/abs/1801.07372\" target=\"_blank\">https://arxiv.org/abs/1801.07372</a></p>\n<pre><code>in summary:\n\ntruth_xy = (size , =  lavel points)\n\ntruth_map_xy= \nlogit_xy = prob_xy =   \ncoord_map = meshgrid(over WH)\nxy = (prob_xy * coord_map).sum(\n\n= (prob_xy,truth_map_xy)\nMSE loss(xy,truth_xy)\n\nyou output xy heatmap !!!\n</code></pre>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2966934,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-08-22T11:22:48.250000",
              "content": "<p>Great. About my code, I'm definitely generating the wrong ideal heatmaps, that's why.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fe9582aa37ea201fe91e5c604b1138e06%2Fheatmap.png?generation=1724325944835986&amp;alt=media\" alt=\"\"></p>\n<p>Fixed:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2F05c2c674cca6ed2ece0c9a5c01071d78%2Fheatmap.png?generation=1724327144714933&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2950098,
      "author_name": "Harry4463",
      "author_url": "",
      "post_date": "2024-08-07T08:21:14.277000",
      "content": "<p>Can we use this in our solution if they're manually annotated? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2950260,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2024-08-07T12:14:47.177000",
          "content": "<p>May be I'm wrong. But I'm pretty sure you can use anything that's public.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2950316,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-08-07T13:42:20.133000",
          "content": "<p>Yep, I believe this data should be good for use. I believe that the rules only apply to pseudolabelling the LB data..</p>\n<p><em>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records</em></p>",
          "votes": 3,
          "replies": [
            {
              "id": 2950520,
              "author_name": "Harry4463",
              "author_url": "",
              "post_date": "2024-08-07T16:23:13.980000",
              "content": "<p>So, just out of interest would you use this for RoI or spine curve? </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2950554,
              "author_name": "Bartley",
              "author_url": "",
              "post_date": "2024-08-07T16:54:43.030000",
              "content": "<p>Up to you! </p>\n<p>For now we are only pretrain on the exact coordinate values. We use this pipeline <a href=\"https://www.kaggle.com/code/brendanartley/lumbar-coordinate-pretraining-example\" target=\"_blank\">here</a> with heaver augmentations and bigger backbones. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2950559,
              "author_name": "Harry4463",
              "author_url": "",
              "post_date": "2024-08-07T16:56:04.083000",
              "content": "<p>Interesting thank you </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2970315,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-08-26T01:49:38.070000",
      "content": "<p>there is yet another way to formulate the problem.<br>\nbecuase of different annotation, we predict the top2 solution (instead of the best solution).<br>\nwe acknowledge that it is not one-to-one regression, but a possible one-to-two.</p>\n<p>here is how kaggle did it in the past:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdcfe734eb2d6f3457c6a295687958ba8%2FSelection_999(5952).png?generation=1724636755026822&amp;alt=media\" alt=\"\"></p>\n<p>implementation of loss<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F85e5fb18fe78f7faef774da3ffa7920f%2FSelection_999(5953).png?generation=1724636778270518&amp;alt=media\" alt=\"\"></p>\n<pre><code> summy:\n- predict  solution\n- compute  loss\n- back propagative ONLY  best loss\n- we also need  predict  probability  correctness    solution.  this ground truth   selected best index above (online labeling)\n\n\n is like  DETR (transformer object detection paper) where you have  online matching algorithm  find  ground truth   predicted box label)\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 2970673,
          "author_name": "Exalted Joseph",
          "author_url": "",
          "post_date": "2024-08-26T11:32:27.703000",
          "content": "<h1>Potential issues:</h1>\n<h1>1. Overfitting to Best Predictions: If the model consistently selects the same type of prediction as the \"best\" one, it might develop a bias toward certain patterns or features, leading to overfitting. This is particularly concerning if there's limited diversity between the two solutions.</h1>\n<h1>2. Noise Sensitivity: The process of selecting the best prediction might introduce noise, especially when the difference between losses is small or the data is noisy. This could result in instability during training and reduced model performance.</h1>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2960130,
      "author_name": "Exalted Joseph",
      "author_url": "",
      "post_date": "2024-08-15T15:55:59.180000",
      "content": "<p>Hello, could you help me with a few questions? I would greatly appreciate it. Are you using 3D convolution? Also, do you have any advice on processing sagittal and axial images?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2960140,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-08-15T16:05:19.647000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/theexaltedone\" target=\"_blank\">@theexaltedone</a>. I would recommend reading through the discussions and sample notebooks, as lots of good info has been shared. Here are a couple relevant ones.</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/523859\" target=\"_blank\">Sagittal Processing</a> by llleeeoooh</li>\n<li><a href=\"https://www.kaggle.com/code/vaillant/cross-reference-images-in-different-mri-planes\" target=\"_blank\">Cross-Reference Images in Different MRI Planes</a> by vaillant</li>\n</ul>",
          "votes": 3,
          "replies": [
            {
              "id": 2960162,
              "author_name": "Exalted Joseph",
              "author_url": "",
              "post_date": "2024-08-15T16:22:15.500000",
              "content": "<p>So how will I handle missing data?  conditions that are missing…</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2976910,
      "author_name": "AyFukushima",
      "author_url": "",
      "post_date": "2024-09-02T12:33:05.030000",
      "content": "<p>Thank you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2965445,
      "author_name": "Albert Garcia",
      "author_url": "",
      "post_date": "2024-08-20T21:38:49.153000",
      "content": "<p>thanks, really cool stuff in here :D</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2949917,
      "author_name": "samu2505",
      "author_url": "",
      "post_date": "2024-08-07T04:25:13.933000",
      "content": "<p>Oh thank you 🙏 </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2949392": "I wanted to share a dataset that we have been using to pretrain models in this competition. This dataset has sped up the convergence time during training and is used in our current pipeline.\n\nThe [Lumbar Coordinate Pretraining Dataset](https://www.kaggle.com/datasets/brendanartley/lumbar-coordinate-pretraining-dataset) was created by collecting T2 MRIs, T1 MRIs, and CT scans of the lower lumbar vertebrae, and then manually annotating the key points of the 5 lower lumbar vertebrae. All the sources are available under CC BY 4.0, and were labelled using the VGG Image Annotator [here](https://www.robots.ox.ac.uk/~vgg/software/via/). \n\nSee a simple pipeline for pretraining with this dataset [here](https://www.kaggle.com/code/brendanartley/lumbar-coordinate-pretraining-example).\n\n![Image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2Fb982e9d30d93d1ee27b582cf861fe178%2Fdataset-cover.png?generation=1722960362507287&alt=media)\n\n\n\nHappy Kaggling!",
    "2957276": "just see see.",
    "2950672": "thanks for the data. I have one question. are you annotating \"spinal canal stenosis\" points?\nOr are you using your own set of rules (e.g. corners of spine vertebrae)?\n\nActually kaggle label coords work well (and of course i expect additional external data would work even better). The trick to use kaggle data is:\n1. select a saggital t2 series.\n2. from the ground truth label coords csv, get the xy coords and z(instance number) of 5 \"spinal canal stenosis\" points for 5 levels \n3. compute mz = median of z \n4. during training random select a slice from the rangem z +/- limit, then set groud truth as (x,y). Now, each slice has 5 points.\n\nI use the above method to train. With proper image size, correct circle radius in target mask and good model to cover the context, convergent is very fast. My results are at: \n\nhttps://www.kaggle.com/code/hengck23/ver-1-demo-workflow-2-stage-approac",
    "2962231": "i confirm double detection is the result of different annotations.\ni.e. even if you change your image encoder, you will still get the \"same results\"\n\nbelow are examples of ambiguous level labeling which have been discussed before in other posts.\nand if I make a clean dataset without ambiguous labels, the trained model  would not give double detection in my experiments.\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F625ebfca7037c1257302f9fd68b9a533%2FSelection_381.png?generation=1723891591633551&alt=media)",
    "2961869": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa156bc0d69b9fd919ac410408338a212%2FSelection_999(5861).png?generation=1723859260030201&alt=media)\n\nnow for spinal canal stenosis level point detection, i have quite a number of double detection (e.g. above). i find it puzzling.i suspect it is due to wrong label from kaggle dataset.\n(i do see some wrong annotations, but it is difficult to check all kaggle annotations)\n\nmy question:\n- have you compared models  trained using your annotated dataset only (which i presume is 100% clean) and kaggle datset only?\n- is there any differences in the results?\n\n---\n\non a side note, i have read some papers that can identify the training images that causes the model to make the current prediction in segmentation/classification. but i cannot recall the title of the paper. if anyone knows these paper or what keyword to google, please put it here. thanks! ",
    "2970047": "i mentioned before wrong label is cuasing wrong detection results.\nin particular, if you have missing point, this could be the reason.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F0cbd474368efac9e1a29b238805b6ad1%2FSelection_999(5950).png?generation=1724604475838911&alt=media)\n\ntop: validation image (at the predict, it miss the red point)\nbottom: training image (at the truth, it has wrong label red point)\n\nboth image are similar. the wrong label cuase the missing detection",
    "2964094": "instead of heatmap detection and direct regression, here comes the third method:\n- softmax over all possible y-coordinate (and x-coordinate)\nit doesn't have heatmap and takes care of double detection \n\n\n---\nleft to right: input image, prediction, truth\n\n\nnormal prediction\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe10d284129505ca02212c93a6d5a0de2%2FSelection_999(5884).png?generation=1724077798492469&alt=media)\n\ndouble prediction\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F8b0c4f01d484b58052aa6162b4f51550%2FSelection_999(5883).png?generation=1724077760199360&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbf593db3d80fa3a65f202db332df96f6%2FSelection_999(5885).png?generation=1724077717548340&alt=media)",
    "2950098": "Can we use this in our solution if they're manually annotated? ",
    "2970315": "there is yet another way to formulate the problem.\nbecuase of different annotation, we predict the top2 solution (instead of the best solution).\nwe acknowledge that it is not one-to-one regression, but a possible one-to-two.\n\nhere is how kaggle did it in the past:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdcfe734eb2d6f3457c6a295687958ba8%2FSelection_999(5952).png?generation=1724636755026822&alt=media)\n\nimplementation of loss\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F85e5fb18fe78f7faef774da3ffa7920f%2FSelection_999(5953).png?generation=1724636778270518&alt=media)\n\n```\nin summy:\n- predict 2 solution\n- compute two loss\n- back propagative ONLY the best loss\n- we also need to predict the probability of correctness of the two solution. set this ground truth to the selected best index above (online labeling)\n\n\nit is like the DETR (transformer object detection paper) where you have an online matching algorithm to find the ground truth of the predicted box label)\n```",
    "2960130": "Hello, could you help me with a few questions? I would greatly appreciate it. Are you using 3D convolution? Also, do you have any advice on processing sagittal and axial images?",
    "2976910": "Thank you.",
    "2965445": "thanks, really cool stuff in here :D",
    "2949917": "Oh thank you 🙏 "
  }
}