{
  "id": 535827,
  "title": "[Help]Questions about image data",
  "url": "/competitions/rsna-2024-lumbar-spine-degenerative-classification/discussion/535827",
  "author_name": "",
  "post_date": "2024-09-24T13:23:17.290700400Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, everyone, I am new to the computer vision task, but I am trying to learn something from this competition.<br>\ni met some puzzles in this notebook, These are the following two questions:</p>\n<ol>\n<li>do you think that the three types of MRI images (Sagittal T1, Sagittal T2/STIR, and Axial T2) in different channels of the same image? </li>\n</ol>\n<pre><code>     ():\n\n        x = np.zeros((Config.image_size, Config.image_size, Config.in_chans), dtype=np.uint8)\n        t = self.df.iloc[idx]\n        st_id = (t[])\n        label = t[:].values.astype(np.int64)\n\n\n        \n         i  (, , ):\n            :\n                p = \n                img = Image.(p).convert()\n                img = np.array(img)\n                x[..., i] = img.astype(np.uint8)\n            :\n                ()\n                \n\n        \n         i  (, , ):\n            :\n                p = \n                img = Image.(p).convert()\n                img = np.array(img)\n                x[..., i+] = img.astype(np.uint8)\n            :\n                ()\n                \n\n        \n        axt2 = glob()\n        axt2 = (axt2)\n\n        step = (axt2) / \n        st = (axt2)/ - *step\n        end = (axt2)+\n\n         i, j  (np.arange(st, end, step)):\n            :\n                p = axt2[(, ((j-).()))]\n                img = Image.(p).convert()\n                img = np.array(img)\n                x[..., i+] = img.astype(np.uint8)\n            :\n                ()\n                  \n\n         np.(x)&gt;\n\n         self.transform   :\n            x = self.transform(image=x)[]\n\n        x = x.transpose(, , )\n\n         x, label\n</code></pre>\n<ol>\n<li>why is it possible to learn and infer 75 prediction labels at once? Because the labels' shape is  (25, )</li>\n</ol>\n<pre><code>     epoch  (, Config.epochs + ):\n        ()\n        model.train()\n        total_loss = \n         tqdm(train_dataloader, leave=)  pbar:\n            optimizer.zero_grad()\n             idx, (x, t)  (pbar):\n                x = x.to(Config.device)\n                t = t.to(Config.device)\n\n                 autocast:\n                    loss = \n                    y = model(x)\n                     col  (Config.n_labels):\n                        pred = y[:,col*:col*+]\n                        gt = t[:,col]\n                        loss = loss + criterion(pred, gt) / Config.n_labels\n</code></pre>\n<p>Thank you very much for your help!</p>",
  "messages": [
    {
      "id": "2997348",
      "postDate": "09/24/2024 13:23:17",
      "content": "<p>Hi, everyone, I am new to the computer vision task, but I am trying to learn something from this competition.<br>\ni met some puzzles in this notebook, These are the following two questions:</p>\n<ol>\n<li>do you think that the three types of MRI images (Sagittal T1, Sagittal T2/STIR, and Axial T2) in different channels of the same image? </li>\n</ol>\n<pre><code>     ():\n\n        x = np.zeros((Config.image_size, Config.image_size, Config.in_chans), dtype=np.uint8)\n        t = self.df.iloc[idx]\n        st_id = (t[])\n        label = t[:].values.astype(np.int64)\n\n\n        \n         i  (, , ):\n            :\n                p = \n                img = Image.(p).convert()\n                img = np.array(img)\n                x[..., i] = img.astype(np.uint8)\n            :\n                ()\n                \n\n        \n         i  (, , ):\n            :\n                p = \n                img = Image.(p).convert()\n                img = np.array(img)\n                x[..., i+] = img.astype(np.uint8)\n            :\n                ()\n                \n\n        \n        axt2 = glob()\n        axt2 = (axt2)\n\n        step = (axt2) / \n        st = (axt2)/ - *step\n        end = (axt2)+\n\n         i, j  (np.arange(st, end, step)):\n            :\n                p = axt2[(, ((j-).()))]\n                img = Image.(p).convert()\n                img = np.array(img)\n                x[..., i+] = img.astype(np.uint8)\n            :\n                ()\n                  \n\n         np.(x)&gt;\n\n         self.transform   :\n            x = self.transform(image=x)[]\n\n        x = x.transpose(, , )\n\n         x, label\n</code></pre>\n<ol>\n<li>why is it possible to learn and infer 75 prediction labels at once? Because the labels' shape is  (25, )</li>\n</ol>\n<pre><code>     epoch  (, Config.epochs + ):\n        ()\n        model.train()\n        total_loss = \n         tqdm(train_dataloader, leave=)  pbar:\n            optimizer.zero_grad()\n             idx, (x, t)  (pbar):\n                x = x.to(Config.device)\n                t = t.to(Config.device)\n\n                 autocast:\n                    loss = \n                    y = model(x)\n                     col  (Config.n_labels):\n                        pred = y[:,col*:col*+]\n                        gt = t[:,col]\n                        loss = loss + criterion(pred, gt) / Config.n_labels\n</code></pre>\n<p>Thank you very much for your help!</p>",
      "rawMarkdown": "Hi, everyone, I am new to the computer vision task, but I am trying to learn something from this competition.\ni met some puzzles in this notebook, These are the following two questions:\n1. do you think that the three types of MRI images (Sagittal T1, Sagittal T2/STIR, and Axial T2) in different channels of the same image? \n```python\n    def __getitem__(self, idx):\n\n        x = np.zeros((Config.image_size, Config.image_size, Config.in_chans), dtype=np.uint8)\n        t = self.df.iloc[idx]\n        st_id = int(t['study_id'])\n        label = t[1:].values.astype(np.int64)\n        \n\n        # Sagittal T1:\n        for i in range(0, 10, 1):\n            try:\n                p = f'../data/cvt_png/{st_id}/Sagittal T1/{i:03d}.png'\n                img = Image.open(p).convert('L')\n                img = np.array(img)\n                x[..., i] = img.astype(np.uint8)\n            except:\n                print(f'failed to load on {st_id}, Sagittal T1')\n                pass\n\n        # Sagittal T2/STIR:\n        for i in range(0, 10, 1):\n            try:\n                p = f'../data/cvt_png/{st_id}/Sagittal T2_STIR/{i:03d}.png'\n                img = Image.open(p).convert('L')\n                img = np.array(img)\n                x[..., i+10] = img.astype(np.uint8)\n            except:\n                print(f'failed to load on {st_id}, Sagittal T2/STIR')\n                pass\n            \n        # Axial T2\n        axt2 = glob(f'../data/cvt_png/{st_id}/Axial T2/*.png')\n        axt2 = sorted(axt2)\n    \n        step = len(axt2) / 10.0\n        st = len(axt2)/2.0 - 4.0*step\n        end = len(axt2)+0.0001\n                \n        for i, j in enumerate(np.arange(st, end, step)):\n            try:\n                p = axt2[max(0, int((j-0.5001).round()))]\n                img = Image.open(p).convert('L')\n                img = np.array(img)\n                x[..., i+20] = img.astype(np.uint8)\n            except:\n                print(f'failed to load on {st_id}, Sagittal T2/STIR')\n                pass  \n\n        assert np.sum(x)>0\n\n        if self.transform is not None:\n            x = self.transform(image=x)['image']\n\n        x = x.transpose(2, 0, 1)\n\n        return x, label\n```\n2. why is it possible to learn and infer 75 prediction labels at once? Because the labels' shape is  (25, )\n```python\n    for epoch in range(1, Config.epochs + 1):\n        print(f'start epoch {epoch}')\n        model.train()\n        total_loss = 0\n        with tqdm(train_dataloader, leave=True) as pbar:\n            optimizer.zero_grad()\n            for idx, (x, t) in enumerate(pbar):\n                x = x.to(Config.device)\n                t = t.to(Config.device)\n\n                with autocast:\n                    loss = 0\n                    y = model(x)\n                    for col in range(Config.n_labels):\n                        pred = y[:,col*3:col*3+3]\n                        gt = t[:,col]\n                        loss = loss + criterion(pred, gt) / Config.n_labels\n\n\n```\n\nThank you very much for your help!",
      "votes": null
    },
    {
      "id": "2997842",
      "postDate": "09/25/2024 00:29:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/alannikos\" target=\"_blank\">@alannikos</a> and welcome, I don't quite get the first question but as for the second part:</p>\n<ol>\n<li><code>y[:, col*3:col*3+3]</code> slices out the predictions for each label from a range of 3 logits. <br>\nSo it is just fancy ways to reshape (75, ) into 25 by 3</li>\n<li><code>gt = t[:, col]</code> grabs the corresponding ground truth (<code>gt</code>) for the same label.</li>\n</ol>\n<p>Now you pass everything to BCE, I believe, and it works just fine. 3 logits and one corresponding labels (0, 1 or 2).</p>\n<p>Hope it makes sense.</p>",
      "rawMarkdown": "Hi @alannikos and welcome, I don't quite get the first question but as for the second part:\n\n1. `y[:, col*3:col*3+3]` slices out the predictions for each label from a range of 3 logits. \n    So it is just fancy ways to reshape (75, ) into 25 by 3\n2. `gt = t[:, col]` grabs the corresponding ground truth (`gt`) for the same label.\n\nNow you pass everything to BCE, I believe, and it works just fine. 3 logits and one corresponding labels (0, 1 or 2).\n\nHope it makes sense.",
      "votes": null
    },
    {
      "id": "2999088",
      "postDate": "09/26/2024 10:50:49",
      "content": "<p>Thank you for your answer! I upvoted the answer. My first question is that during the establishment of this dataset, it seems that 30 png images were loaded into one sample, that is, in_channels=30. However, I wonder if the general image data is 3channels, is it because it is medical image data?</p>",
      "rawMarkdown": "Thank you for your answer! I upvoted the answer. My first question is that during the establishment of this dataset, it seems that 30 png images were loaded into one sample, that is, in_channels=30. However, I wonder if the general image data is 3channels, is it because it is medical image data?",
      "votes": null
    },
    {
      "id": "2999917",
      "postDate": "09/27/2024 04:20:22",
      "content": "<p>Yes, it is because the DICOM slices are gray scale. There are variable number of slices so the code is taking an average guess of 30 which may contain the relevant data.</p>",
      "rawMarkdown": "Yes, it is because the DICOM slices are gray scale. There are variable number of slices so the code is taking an average guess of 30 which may contain the relevant data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2997842,
      "author_name": "sergiosaharovskiy",
      "author_url": "",
      "post_date": "09/25/2024 00:29:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/alannikos\" target=\"_blank\">@alannikos</a> and welcome, I don't quite get the first question but as for the second part:</p>\n<ol>\n<li><code>y[:, col*3:col*3+3]</code> slices out the predictions for each label from a range of 3 logits. <br>\nSo it is just fancy ways to reshape (75, ) into 25 by 3</li>\n<li><code>gt = t[:, col]</code> grabs the corresponding ground truth (<code>gt</code>) for the same label.</li>\n</ol>\n<p>Now you pass everything to BCE, I believe, and it works just fine. 3 logits and one corresponding labels (0, 1 or 2).</p>\n<p>Hope it makes sense.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2999088,
          "author_name": "alannikos",
          "author_url": "",
          "post_date": "09/26/2024 10:50:49",
          "content": "<p>Thank you for your answer! I upvoted the answer. My first question is that during the establishment of this dataset, it seems that 30 png images were loaded into one sample, that is, in_channels=30. However, I wonder if the general image data is 3channels, is it because it is medical image data?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2999917,
              "author_name": "coderrkj",
              "author_url": "",
              "post_date": "09/27/2024 04:20:22",
              "content": "<p>Yes, it is because the DICOM slices are gray scale. There are variable number of slices so the code is taking an average guess of 30 which may contain the relevant data.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2997348": "Hi, everyone, I am new to the computer vision task, but I am trying to learn something from this competition.\ni met some puzzles in this notebook, These are the following two questions:\n1. do you think that the three types of MRI images (Sagittal T1, Sagittal T2/STIR, and Axial T2) in different channels of the same image? \n```python\n    def __getitem__(self, idx):\n\n        x = np.zeros((Config.image_size, Config.image_size, Config.in_chans), dtype=np.uint8)\n        t = self.df.iloc[idx]\n        st_id = int(t['study_id'])\n        label = t[1:].values.astype(np.int64)\n        \n\n        # Sagittal T1:\n        for i in range(0, 10, 1):\n            try:\n                p = f'../data/cvt_png/{st_id}/Sagittal T1/{i:03d}.png'\n                img = Image.open(p).convert('L')\n                img = np.array(img)\n                x[..., i] = img.astype(np.uint8)\n            except:\n                print(f'failed to load on {st_id}, Sagittal T1')\n                pass\n\n        # Sagittal T2/STIR:\n        for i in range(0, 10, 1):\n            try:\n                p = f'../data/cvt_png/{st_id}/Sagittal T2_STIR/{i:03d}.png'\n                img = Image.open(p).convert('L')\n                img = np.array(img)\n                x[..., i+10] = img.astype(np.uint8)\n            except:\n                print(f'failed to load on {st_id}, Sagittal T2/STIR')\n                pass\n            \n        # Axial T2\n        axt2 = glob(f'../data/cvt_png/{st_id}/Axial T2/*.png')\n        axt2 = sorted(axt2)\n    \n        step = len(axt2) / 10.0\n        st = len(axt2)/2.0 - 4.0*step\n        end = len(axt2)+0.0001\n                \n        for i, j in enumerate(np.arange(st, end, step)):\n            try:\n                p = axt2[max(0, int((j-0.5001).round()))]\n                img = Image.open(p).convert('L')\n                img = np.array(img)\n                x[..., i+20] = img.astype(np.uint8)\n            except:\n                print(f'failed to load on {st_id}, Sagittal T2/STIR')\n                pass  \n\n        assert np.sum(x)>0\n\n        if self.transform is not None:\n            x = self.transform(image=x)['image']\n\n        x = x.transpose(2, 0, 1)\n\n        return x, label\n```\n2. why is it possible to learn and infer 75 prediction labels at once? Because the labels' shape is  (25, )\n```python\n    for epoch in range(1, Config.epochs + 1):\n        print(f'start epoch {epoch}')\n        model.train()\n        total_loss = 0\n        with tqdm(train_dataloader, leave=True) as pbar:\n            optimizer.zero_grad()\n            for idx, (x, t) in enumerate(pbar):\n                x = x.to(Config.device)\n                t = t.to(Config.device)\n\n                with autocast:\n                    loss = 0\n                    y = model(x)\n                    for col in range(Config.n_labels):\n                        pred = y[:,col*3:col*3+3]\n                        gt = t[:,col]\n                        loss = loss + criterion(pred, gt) / Config.n_labels\n\n\n```\n\nThank you very much for your help!",
    "2997842": "Hi @alannikos and welcome, I don't quite get the first question but as for the second part:\n\n1. `y[:, col*3:col*3+3]` slices out the predictions for each label from a range of 3 logits. \n    So it is just fancy ways to reshape (75, ) into 25 by 3\n2. `gt = t[:, col]` grabs the corresponding ground truth (`gt`) for the same label.\n\nNow you pass everything to BCE, I believe, and it works just fine. 3 logits and one corresponding labels (0, 1 or 2).\n\nHope it makes sense.",
    "2999088": "Thank you for your answer! I upvoted the answer. My first question is that during the establishment of this dataset, it seems that 30 png images were loaded into one sample, that is, in_channels=30. However, I wonder if the general image data is 3channels, is it because it is medical image data?",
    "2999917": "Yes, it is because the DICOM slices are gray scale. There are variable number of slices so the code is taking an average guess of 30 which may contain the relevant data."
  },
  "source": "meta"
}