{
  "id": 175560,
  "title": "GNN for Ugly Duckling Effect ?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175560",
  "author_name": "",
  "post_date": "2020-08-18T15:12:13.144985100Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>How did you deal with \"Ugly Duckling Effect\"?</h1>\n<p>As it has already been mentioned, dermatologist use the \"Ugly Duckling Effect\" to detect melanomas : they look at all the moles of the patient, if one looks weird and very different from other moles then they get suspicious. This is based on the fact that it's impossible (or very very unlikely) that some one would get 20 melanomas at the same time, so if you have 20 moles looking alike then there is probably nothing to worry about. But knowing this, have you been able to use it successfully in your models?</p>\n<h1>Graph Neural Networks</h1>\n<p>I personally tried a quite unusual approach without much success but I thought it would still be nice to share it publicly and create this thread so that people can discuss about their approach here.</p>\n<h2>The Graph</h2>\n<p>I created a graph for each patient :</p>\n<ul>\n<li><strong>each node</strong>: consisted in meta features related to one image, like age or location AND the final head of pretrained model (about 2048 features). I also added the OOF prediction score from the pretrained model. With this every node has information about one specific image.</li>\n<li><strong>each edge</strong> : every node was connected to any other node, so the graph was fully connected and undirected. Edges contained only two information (but more could have been added) : L2 distance between pertained embeddings  and Age difference between the two nodes</li>\n<li><strong>global features</strong> : each graph has patient specific features like sex.</li>\n</ul>\n<h2>The Model</h2>\n<p>I tried to use a Message Passing Neural network approach using <code>MetaLayer</code> from the excellent pytorch-geometric library (<a href=\"https://github.com/rusty1s/pytorch_geometric)\" target=\"_blank\">https://github.com/rusty1s/pytorch_geometric)</a>. This consist of 3 models for Edges, Nodes and GlobalFeatures that learn how to exchange information from each other in order to do Node classification.</p>\n<p>So I stacked 3 Metalayers and then a simple MLP in order to get the class for each node (each image). The idea was that you are using the entire graph in order to take a decision, and the model had access to enough information to learn the \"Ugly Duckling\" concept, if you had a lot of edges with small distances maybe you're not very different from others so you are not a melanoma.</p>\n<h2>The results</h2>\n<p>My model trained and I could get something like CV score of 0.935~0.94 (but this was used with multiple oof predictions so a simple averaging would have done better ~0.95). So I did not manage to push this idea to a useful submission but I think there is some potential into it. As I'm no expert in GNNs I would be happy to receive advice on this!</p>\n<p>Maybe this approach could work but we would need to find a better way to encode the image than with pretrained embeddings, maybe some smarter distances would help…</p>\n<h3>What was your approach ?</h3>",
  "messages": [
    {
      "id": "975993",
      "postDate": "08/18/2020 15:12:13",
      "content": "<h1>How did you deal with \"Ugly Duckling Effect\"?</h1>\n<p>As it has already been mentioned, dermatologist use the \"Ugly Duckling Effect\" to detect melanomas : they look at all the moles of the patient, if one looks weird and very different from other moles then they get suspicious. This is based on the fact that it's impossible (or very very unlikely) that some one would get 20 melanomas at the same time, so if you have 20 moles looking alike then there is probably nothing to worry about. But knowing this, have you been able to use it successfully in your models?</p>\n<h1>Graph Neural Networks</h1>\n<p>I personally tried a quite unusual approach without much success but I thought it would still be nice to share it publicly and create this thread so that people can discuss about their approach here.</p>\n<h2>The Graph</h2>\n<p>I created a graph for each patient :</p>\n<ul>\n<li><strong>each node</strong>: consisted in meta features related to one image, like age or location AND the final head of pretrained model (about 2048 features). I also added the OOF prediction score from the pretrained model. With this every node has information about one specific image.</li>\n<li><strong>each edge</strong> : every node was connected to any other node, so the graph was fully connected and undirected. Edges contained only two information (but more could have been added) : L2 distance between pertained embeddings  and Age difference between the two nodes</li>\n<li><strong>global features</strong> : each graph has patient specific features like sex.</li>\n</ul>\n<h2>The Model</h2>\n<p>I tried to use a Message Passing Neural network approach using <code>MetaLayer</code> from the excellent pytorch-geometric library (<a href=\"https://github.com/rusty1s/pytorch_geometric)\" target=\"_blank\">https://github.com/rusty1s/pytorch_geometric)</a>. This consist of 3 models for Edges, Nodes and GlobalFeatures that learn how to exchange information from each other in order to do Node classification.</p>\n<p>So I stacked 3 Metalayers and then a simple MLP in order to get the class for each node (each image). The idea was that you are using the entire graph in order to take a decision, and the model had access to enough information to learn the \"Ugly Duckling\" concept, if you had a lot of edges with small distances maybe you're not very different from others so you are not a melanoma.</p>\n<h2>The results</h2>\n<p>My model trained and I could get something like CV score of 0.935~0.94 (but this was used with multiple oof predictions so a simple averaging would have done better ~0.95). So I did not manage to push this idea to a useful submission but I think there is some potential into it. As I'm no expert in GNNs I would be happy to receive advice on this!</p>\n<p>Maybe this approach could work but we would need to find a better way to encode the image than with pretrained embeddings, maybe some smarter distances would help…</p>\n<h3>What was your approach ?</h3>",
      "rawMarkdown": "# How did you deal with \"Ugly Duckling Effect\"?\n\nAs it has already been mentioned, dermatologist use the \"Ugly Duckling Effect\" to detect melanomas : they look at all the moles of the patient, if one looks weird and very different from other moles then they get suspicious. This is based on the fact that it's impossible (or very very unlikely) that some one would get 20 melanomas at the same time, so if you have 20 moles looking alike then there is probably nothing to worry about. But knowing this, have you been able to use it successfully in your models?\n\n\n# Graph Neural Networks\n\nI personally tried a quite unusual approach without much success but I thought it would still be nice to share it publicly and create this thread so that people can discuss about their approach here.\n\n## The Graph\n\nI created a graph for each patient :\n- **each node**: consisted in meta features related to one image, like age or location AND the final head of pretrained model (about 2048 features). I also added the OOF prediction score from the pretrained model. With this every node has information about one specific image.\n- **each edge** : every node was connected to any other node, so the graph was fully connected and undirected. Edges contained only two information (but more could have been added) : L2 distance between pertained embeddings  and Age difference between the two nodes\n- **global features** : each graph has patient specific features like sex.\n\n## The Model\n\nI tried to use a Message Passing Neural network approach using `MetaLayer` from the excellent pytorch-geometric library (https://github.com/rusty1s/pytorch_geometric). This consist of 3 models for Edges, Nodes and GlobalFeatures that learn how to exchange information from each other in order to do Node classification.\n\nSo I stacked 3 Metalayers and then a simple MLP in order to get the class for each node (each image). The idea was that you are using the entire graph in order to take a decision, and the model had access to enough information to learn the \"Ugly Duckling\" concept, if you had a lot of edges with small distances maybe you're not very different from others so you are not a melanoma.\n\n\n## The results\n\nMy model trained and I could get something like CV score of 0.935~0.94 (but this was used with multiple oof predictions so a simple averaging would have done better ~0.95). So I did not manage to push this idea to a useful submission but I think there is some potential into it. As I'm no expert in GNNs I would be happy to receive advice on this!\n\nMaybe this approach could work but we would need to find a better way to encode the image than with pretrained embeddings, maybe some smarter distances would help...\n\n\n### What was your approach ?",
      "votes": null
    },
    {
      "id": "976009",
      "postDate": "08/18/2020 15:24:59",
      "content": "<p>We had a similar idea, but didnt execute it in the end. Couldn't you perhaps train it end-to-end such that you learn both representations of your images (CNN) and how the images and metadata of a same patient can be combined/aggregated (GNN)?</p>",
      "rawMarkdown": "We had a similar idea, but didnt execute it in the end. Couldn't you perhaps train it end-to-end such that you learn both representations of your images (CNN) and how the images and metadata of a same patient can be combined/aggregated (GNN)?",
      "votes": null
    },
    {
      "id": "976014",
      "postDate": "08/18/2020 15:28:35",
      "content": "<p>Yeah that seems a good approach but it would be a huge model as you'll have to provide the entire image for each node. If you want to have a batch size of say 8 : you would need to have 8xnum_images in your model seems huge.</p>\n<p>It would be very interesting to try though!</p>",
      "rawMarkdown": "Yeah that seems a good approach but it would be a huge model as you'll have to provide the entire image for each node. If you want to have a batch size of say 8 : you would need to have 8xnum_images in your model seems huge.\n\nIt would be very interesting to try though!",
      "votes": null
    },
    {
      "id": "976969",
      "postDate": "08/19/2020 07:53:06",
      "content": "<p>This is a very interesting approach! Did you try checking performance on LB with a late submission?</p>\n<p>We did some experiments on “ugly duckling” using a simpler strategy: based on within-patient image distances extracted from CNN embeddings and based on predictions distribution within patients but did not succeed as well. I posted a brief overview of our experiments <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175500\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "This is a very interesting approach! Did you try checking performance on LB with a late submission?\n\nWe did some experiments on “ugly duckling” using a simpler strategy: based on within-patient image distances extracted from CNN embeddings and based on predictions distribution within patients but did not succeed as well. I posted a brief overview of our experiments [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175500).",
      "votes": null
    },
    {
      "id": "977138",
      "postDate": "08/19/2020 10:11:03",
      "content": "<p>I'll try a late submission and update you with the results!</p>",
      "rawMarkdown": "I'll try a late submission and update you with the results!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 976009,
      "author_name": "group16",
      "author_url": "",
      "post_date": "08/18/2020 15:24:59",
      "content": "<p>We had a similar idea, but didnt execute it in the end. Couldn't you perhaps train it end-to-end such that you learn both representations of your images (CNN) and how the images and metadata of a same patient can be combined/aggregated (GNN)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 976014,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "08/18/2020 15:28:35",
          "content": "<p>Yeah that seems a good approach but it would be a huge model as you'll have to provide the entire image for each node. If you want to have a batch size of say 8 : you would need to have 8xnum_images in your model seems huge.</p>\n<p>It would be very interesting to try though!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 976969,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "08/19/2020 07:53:06",
      "content": "<p>This is a very interesting approach! Did you try checking performance on LB with a late submission?</p>\n<p>We did some experiments on “ugly duckling” using a simpler strategy: based on within-patient image distances extracted from CNN embeddings and based on predictions distribution within patients but did not succeed as well. I posted a brief overview of our experiments <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175500\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 977138,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "08/19/2020 10:11:03",
          "content": "<p>I'll try a late submission and update you with the results!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "975993": "# How did you deal with \"Ugly Duckling Effect\"?\n\nAs it has already been mentioned, dermatologist use the \"Ugly Duckling Effect\" to detect melanomas : they look at all the moles of the patient, if one looks weird and very different from other moles then they get suspicious. This is based on the fact that it's impossible (or very very unlikely) that some one would get 20 melanomas at the same time, so if you have 20 moles looking alike then there is probably nothing to worry about. But knowing this, have you been able to use it successfully in your models?\n\n\n# Graph Neural Networks\n\nI personally tried a quite unusual approach without much success but I thought it would still be nice to share it publicly and create this thread so that people can discuss about their approach here.\n\n## The Graph\n\nI created a graph for each patient :\n- **each node**: consisted in meta features related to one image, like age or location AND the final head of pretrained model (about 2048 features). I also added the OOF prediction score from the pretrained model. With this every node has information about one specific image.\n- **each edge** : every node was connected to any other node, so the graph was fully connected and undirected. Edges contained only two information (but more could have been added) : L2 distance between pertained embeddings  and Age difference between the two nodes\n- **global features** : each graph has patient specific features like sex.\n\n## The Model\n\nI tried to use a Message Passing Neural network approach using `MetaLayer` from the excellent pytorch-geometric library (https://github.com/rusty1s/pytorch_geometric). This consist of 3 models for Edges, Nodes and GlobalFeatures that learn how to exchange information from each other in order to do Node classification.\n\nSo I stacked 3 Metalayers and then a simple MLP in order to get the class for each node (each image). The idea was that you are using the entire graph in order to take a decision, and the model had access to enough information to learn the \"Ugly Duckling\" concept, if you had a lot of edges with small distances maybe you're not very different from others so you are not a melanoma.\n\n\n## The results\n\nMy model trained and I could get something like CV score of 0.935~0.94 (but this was used with multiple oof predictions so a simple averaging would have done better ~0.95). So I did not manage to push this idea to a useful submission but I think there is some potential into it. As I'm no expert in GNNs I would be happy to receive advice on this!\n\nMaybe this approach could work but we would need to find a better way to encode the image than with pretrained embeddings, maybe some smarter distances would help...\n\n\n### What was your approach ?",
    "976009": "We had a similar idea, but didnt execute it in the end. Couldn't you perhaps train it end-to-end such that you learn both representations of your images (CNN) and how the images and metadata of a same patient can be combined/aggregated (GNN)?",
    "976014": "Yeah that seems a good approach but it would be a huge model as you'll have to provide the entire image for each node. If you want to have a batch size of say 8 : you would need to have 8xnum_images in your model seems huge.\n\nIt would be very interesting to try though!",
    "976969": "This is a very interesting approach! Did you try checking performance on LB with a late submission?\n\nWe did some experiments on “ugly duckling” using a simpler strategy: based on within-patient image distances extracted from CNN embeddings and based on predictions distribution within patients but did not succeed as well. I posted a brief overview of our experiments [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175500).",
    "977138": "I'll try a late submission and update you with the results!"
  },
  "source": "meta"
}