{
  "id": 214880,
  "title": "Are CNN's really suited to this task?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/214880",
  "author_name": "",
  "post_date": "2021-01-28T00:39:41.710036800Z",
  "votes": 16,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This is more of a philosophical post than actual competition advice. Just something to consider. Clearly CNNs are working quite well in this competition.</p>\n<p>One of the things I have been reminded of for this competition is some research I did for a previous competition. <a href=\"https://www.kaggle.com/ryches/what-can-a-cnn-do\" target=\"_blank\">https://www.kaggle.com/ryches/what-can-a-cnn-do</a> I spent some time evaluating the toy problem of just measuring the distance between two points. The model does surprisingly poorly and is very fragile to various different perturbations. It takes more epochs to train than I would expect and does not hold robust to different scaled images, which doesn't entirely make sense to me. </p>\n<p>I believe this is fundamental to this problem because the placement of the catheter in reference to various key points in the body. Similar to the measuring task, but even harder because the model also must detect the end of the catheter and the key points in the body, these are not as simple as just single pixels. </p>\n<p>A CNN is translationally invariant so if it is just detecting catheter and a body keypoint anywhere in the body it will not behave particularly differently. Where it will start detecting the various classifications is if the receptive field is wide enough that it can see the catheter and the keypoint within the same field and look at their interaction. I believe that this might be why we see the resnet200d doing so well in reference to shallower network architectures. The repeated convolutions might allow it to see a larger field. </p>",
  "messages": [
    {
      "id": "1173558",
      "postDate": "01/28/2021 00:39:41",
      "content": "<p>This is more of a philosophical post than actual competition advice. Just something to consider. Clearly CNNs are working quite well in this competition.</p>\n<p>One of the things I have been reminded of for this competition is some research I did for a previous competition. <a href=\"https://www.kaggle.com/ryches/what-can-a-cnn-do\" target=\"_blank\">https://www.kaggle.com/ryches/what-can-a-cnn-do</a> I spent some time evaluating the toy problem of just measuring the distance between two points. The model does surprisingly poorly and is very fragile to various different perturbations. It takes more epochs to train than I would expect and does not hold robust to different scaled images, which doesn't entirely make sense to me. </p>\n<p>I believe this is fundamental to this problem because the placement of the catheter in reference to various key points in the body. Similar to the measuring task, but even harder because the model also must detect the end of the catheter and the key points in the body, these are not as simple as just single pixels. </p>\n<p>A CNN is translationally invariant so if it is just detecting catheter and a body keypoint anywhere in the body it will not behave particularly differently. Where it will start detecting the various classifications is if the receptive field is wide enough that it can see the catheter and the keypoint within the same field and look at their interaction. I believe that this might be why we see the resnet200d doing so well in reference to shallower network architectures. The repeated convolutions might allow it to see a larger field. </p>",
      "rawMarkdown": "This is more of a philosophical post than actual competition advice. Just something to consider. Clearly CNNs are working quite well in this competition.\n\nOne of the things I have been reminded of for this competition is some research I did for a previous competition. https://www.kaggle.com/ryches/what-can-a-cnn-do I spent some time evaluating the toy problem of just measuring the distance between two points. The model does surprisingly poorly and is very fragile to various different perturbations. It takes more epochs to train than I would expect and does not hold robust to different scaled images, which doesn't entirely make sense to me. \n\nI believe this is fundamental to this problem because the placement of the catheter in reference to various key points in the body. Similar to the measuring task, but even harder because the model also must detect the end of the catheter and the key points in the body, these are not as simple as just single pixels. \n\nA CNN is translationally invariant so if it is just detecting catheter and a body keypoint anywhere in the body it will not behave particularly differently. Where it will start detecting the various classifications is if the receptive field is wide enough that it can see the catheter and the keypoint within the same field and look at their interaction. I believe that this might be why we see the resnet200d doing so well in reference to shallower network architectures. The repeated convolutions might allow it to see a larger field.",
      "votes": null
    },
    {
      "id": "1173576",
      "postDate": "01/28/2021 01:20:38",
      "content": "<p>I think that CNN models are not ideal, but very well suited for the task. Most of the competition tasks can be solved just by counting grids where object (CVC/NGT/ETT) is present in penultimate layers of the CNN model. The tip of CVC can also be seen as detecting if the tip is in the specific cell in the grid: below is CVC-normal CVC tip end point visualized across all annotations - it is highly concentrated:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F405318%2F7941c78df4971e6b12b0c2659a944b48%2Fnormal.png?generation=1610201668234131&amp;alt=media\" alt=\"\"></p>\n<p>So on average CNN are a possible approach. </p>\n<p>However -</p>\n<p>I tested simple xgboost model on train_annotations by making bunch of features with relation to lung masks. Assuming you have a perfect UNet segmentation model, xgboost on top can provide +5% on CVC class alone. However, segmentation task is also as challenging as classification models.</p>",
      "rawMarkdown": "I think that CNN models are not ideal, but very well suited for the task. Most of the competition tasks can be solved just by counting grids where object (CVC/NGT/ETT) is present in penultimate layers of the CNN model. The tip of CVC can also be seen as detecting if the tip is in the specific cell in the grid: below is CVC-normal CVC tip end point visualized across all annotations - it is highly concentrated:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F405318%2F7941c78df4971e6b12b0c2659a944b48%2Fnormal.png?generation=1610201668234131&alt=media)\n\nSo on average CNN are a possible approach. \n\nHowever -\n\nI tested simple xgboost model on train_annotations by making bunch of features with relation to lung masks. Assuming you have a perfect UNet segmentation model, xgboost on top can provide +5% on CVC class alone. However, segmentation task is also as challenging as classification models.",
      "votes": null
    },
    {
      "id": "1173643",
      "postDate": "01/28/2021 03:09:44",
      "content": "<p>I think the question I have about these CNN models is how much are they truly learning the relative position of the catheter to the appropriate body part, and how much are they learning the position on the image?  For the latter, that means that they are dependent upon the fact that hospitals position the x-rays consistently, so the heart and lungs are at the approximate same (x,y) coordinates on different images.  If they are somewhat dependent on positioning, then the models could give wrong predictions for x-rays that are slightly higher/lower than usual.  If this is happening, then the question would be can we better train them to understand relative position to body parts and less depend on (x,y) position on the image?</p>",
      "rawMarkdown": "I think the question I have about these CNN models is how much are they truly learning the relative position of the catheter to the appropriate body part, and how much are they learning the position on the image?  For the latter, that means that they are dependent upon the fact that hospitals position the x-rays consistently, so the heart and lungs are at the approximate same (x,y) coordinates on different images.  If they are somewhat dependent on positioning, then the models could give wrong predictions for x-rays that are slightly higher/lower than usual.  If this is happening, then the question would be can we better train them to understand relative position to body parts and less depend on (x,y) position on the image?",
      "votes": null
    },
    {
      "id": "1174143",
      "postDate": "01/28/2021 09:48:42",
      "content": "<p>I think this is mostly solved with shift/crop/rotate augmentations </p>",
      "rawMarkdown": "I think this is mostly solved with shift/crop/rotate augmentations",
      "votes": null
    },
    {
      "id": "1174153",
      "postDate": "01/28/2021 10:03:07",
      "content": "<p>Agreed.  I do wonder if we're fighting the translational invariance of CNNs, and CVC is the hardest because it is the most positionally sensitive one.</p>",
      "rawMarkdown": "Agreed.  I do wonder if we're fighting the translational invariance of CNNs, and CVC is the hardest because it is the most positionally sensitive one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1173576,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "01/28/2021 01:20:38",
      "content": "<p>I think that CNN models are not ideal, but very well suited for the task. Most of the competition tasks can be solved just by counting grids where object (CVC/NGT/ETT) is present in penultimate layers of the CNN model. The tip of CVC can also be seen as detecting if the tip is in the specific cell in the grid: below is CVC-normal CVC tip end point visualized across all annotations - it is highly concentrated:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F405318%2F7941c78df4971e6b12b0c2659a944b48%2Fnormal.png?generation=1610201668234131&amp;alt=media\" alt=\"\"></p>\n<p>So on average CNN are a possible approach. </p>\n<p>However -</p>\n<p>I tested simple xgboost model on train_annotations by making bunch of features with relation to lung masks. Assuming you have a perfect UNet segmentation model, xgboost on top can provide +5% on CVC class alone. However, segmentation task is also as challenging as classification models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1173643,
      "author_name": "socated",
      "author_url": "",
      "post_date": "01/28/2021 03:09:44",
      "content": "<p>I think the question I have about these CNN models is how much are they truly learning the relative position of the catheter to the appropriate body part, and how much are they learning the position on the image?  For the latter, that means that they are dependent upon the fact that hospitals position the x-rays consistently, so the heart and lungs are at the approximate same (x,y) coordinates on different images.  If they are somewhat dependent on positioning, then the models could give wrong predictions for x-rays that are slightly higher/lower than usual.  If this is happening, then the question would be can we better train them to understand relative position to body parts and less depend on (x,y) position on the image?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1174143,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "01/28/2021 09:48:42",
          "content": "<p>I think this is mostly solved with shift/crop/rotate augmentations </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1174153,
          "author_name": "socated",
          "author_url": "",
          "post_date": "01/28/2021 10:03:07",
          "content": "<p>Agreed.  I do wonder if we're fighting the translational invariance of CNNs, and CVC is the hardest because it is the most positionally sensitive one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1173558": "This is more of a philosophical post than actual competition advice. Just something to consider. Clearly CNNs are working quite well in this competition.\n\nOne of the things I have been reminded of for this competition is some research I did for a previous competition. https://www.kaggle.com/ryches/what-can-a-cnn-do I spent some time evaluating the toy problem of just measuring the distance between two points. The model does surprisingly poorly and is very fragile to various different perturbations. It takes more epochs to train than I would expect and does not hold robust to different scaled images, which doesn't entirely make sense to me. \n\nI believe this is fundamental to this problem because the placement of the catheter in reference to various key points in the body. Similar to the measuring task, but even harder because the model also must detect the end of the catheter and the key points in the body, these are not as simple as just single pixels. \n\nA CNN is translationally invariant so if it is just detecting catheter and a body keypoint anywhere in the body it will not behave particularly differently. Where it will start detecting the various classifications is if the receptive field is wide enough that it can see the catheter and the keypoint within the same field and look at their interaction. I believe that this might be why we see the resnet200d doing so well in reference to shallower network architectures. The repeated convolutions might allow it to see a larger field.",
    "1173576": "I think that CNN models are not ideal, but very well suited for the task. Most of the competition tasks can be solved just by counting grids where object (CVC/NGT/ETT) is present in penultimate layers of the CNN model. The tip of CVC can also be seen as detecting if the tip is in the specific cell in the grid: below is CVC-normal CVC tip end point visualized across all annotations - it is highly concentrated:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F405318%2F7941c78df4971e6b12b0c2659a944b48%2Fnormal.png?generation=1610201668234131&alt=media)\n\nSo on average CNN are a possible approach. \n\nHowever -\n\nI tested simple xgboost model on train_annotations by making bunch of features with relation to lung masks. Assuming you have a perfect UNet segmentation model, xgboost on top can provide +5% on CVC class alone. However, segmentation task is also as challenging as classification models.",
    "1173643": "I think the question I have about these CNN models is how much are they truly learning the relative position of the catheter to the appropriate body part, and how much are they learning the position on the image?  For the latter, that means that they are dependent upon the fact that hospitals position the x-rays consistently, so the heart and lungs are at the approximate same (x,y) coordinates on different images.  If they are somewhat dependent on positioning, then the models could give wrong predictions for x-rays that are slightly higher/lower than usual.  If this is happening, then the question would be can we better train them to understand relative position to body parts and less depend on (x,y) position on the image?",
    "1174143": "I think this is mostly solved with shift/crop/rotate augmentations",
    "1174153": "Agreed.  I do wonder if we're fighting the translational invariance of CNNs, and CVC is the hardest because it is the most positionally sensitive one."
  },
  "source": "meta"
}