{
  "id": 166733,
  "title": "Using diagnosis as secondary output",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/166733",
  "author_name": "",
  "post_date": "2020-07-13T19:38:25.219661200Z",
  "votes": 2,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Are you using diagnosis column as a secondary output for you model? </p>\n\n<p>If you think about it, maybee this target provides more information so your model can differentiate melanoma.</p>\n\n<p>Here is a stackoverflow discussion were they explain how this work:</p>\n\n<p><a href=\"https://stackoverflow.com/questions/57149476/how-is-a-multiple-outputs-deep-learning-model-trained\">https://stackoverflow.com/questions/57149476/how-is-a-multiple-outputs-deep-learning-model-trained</a></p>\n\n<p>Cheers and good luck in the competition.</p>",
  "messages": [
    {
      "id": "928190",
      "postDate": "07/13/2020 19:38:25",
      "content": "<p>Are you using diagnosis column as a secondary output for you model? </p>\n\n<p>If you think about it, maybee this target provides more information so your model can differentiate melanoma.</p>\n\n<p>Here is a stackoverflow discussion were they explain how this work:</p>\n\n<p><a href=\"https://stackoverflow.com/questions/57149476/how-is-a-multiple-outputs-deep-learning-model-trained\">https://stackoverflow.com/questions/57149476/how-is-a-multiple-outputs-deep-learning-model-trained</a></p>\n\n<p>Cheers and good luck in the competition.</p>",
      "rawMarkdown": "Are you using diagnosis column as a secondary output for you model? \n\nIf you think about it, maybee this target provides more information so your model can differentiate melanoma.\n\nHere is a stackoverflow discussion were they explain how this work:\n\nhttps://stackoverflow.com/questions/57149476/how-is-a-multiple-outputs-deep-learning-model-trained\n\nCheers and good luck in the competition.",
      "votes": null
    },
    {
      "id": "928220",
      "postDate": "07/13/2020 19:51:24",
      "content": "<p>Deploy it on PyTorch look easy! But didn't try.</p>",
      "rawMarkdown": "Deploy it on PyTorch look easy! But didn't try.",
      "votes": null
    },
    {
      "id": "928379",
      "postDate": "07/13/2020 23:58:26",
      "content": "<p>```\ntrain_df[\"diagnosis\"].value_counts()\nunknown                               27124\nnevus                                  5193\nmelanoma                                584\nseborrheic keratosis                    135\nlentigo NOS                              44\nlichenoid keratosis                      37\nsolar lentigo                             7\ncafe-au-lait macule                       1\natypical melanocytic proliferation        1\nName: diagnosis, dtype: int64</p>\n\n<p>train_df[train_df[\"target\"] == 1][\"diagnosis\"].value_counts()\nmelanoma    584\nName: diagnosis, dtype: int64\n```</p>\n\n<p>The problem is that 80% is uknown. I think the useful value would be nevus. What's your idea?</p>",
      "rawMarkdown": "```\ntrain_df[\"diagnosis\"].value_counts()\nunknown                               27124\nnevus                                  5193\nmelanoma                                584\nseborrheic keratosis                    135\nlentigo NOS                              44\nlichenoid keratosis                      37\nsolar lentigo                             7\ncafe-au-lait macule                       1\natypical melanocytic proliferation        1\nName: diagnosis, dtype: int64\n\ntrain_df[train_df[\"target\"] == 1][\"diagnosis\"].value_counts()\nmelanoma    584\nName: diagnosis, dtype: int64\n```\n\nThe problem is that 80% is uknown. I think the useful value would be nevus. What's your idea?",
      "votes": null
    },
    {
      "id": "928400",
      "postDate": "07/14/2020 00:37:34",
      "content": "<p>Yep i was thinking about nevus, maybee that can help the model. It's just and idea 👊 </p>",
      "rawMarkdown": "Yep i was thinking about nevus, maybee that can help the model. It's just and idea 👊",
      "votes": null
    },
    {
      "id": "928417",
      "postDate": "07/14/2020 01:02:30",
      "content": "<p>it may be, please share your findings :)</p>",
      "rawMarkdown": "it may be, please share your findings :)",
      "votes": null
    },
    {
      "id": "928421",
      "postDate": "07/14/2020 01:15:48",
      "content": "<p>I just trained a model and the cv did not improve 👎 </p>",
      "rawMarkdown": "I just trained a model and the cv did not improve 👎",
      "votes": null
    },
    {
      "id": "929202",
      "postDate": "07/14/2020 14:20:46",
      "content": "<p><a href=\"/ragnar123\">@ragnar123</a> can you describe what you did? I think to improve score you need multiple models here</p>",
      "rawMarkdown": "ragnar123 can you describe what you did? I think to improve score you need multiple models here",
      "votes": null
    },
    {
      "id": "929214",
      "postDate": "07/14/2020 14:36:09",
      "content": "<p>Sure, i just used a secondary output using diagnosis with a categorical crossentropy loss.  In other words the model will have 2 outputs, 1 for the binary (target column) and another for the multiclass part (diagnosis).</p>",
      "rawMarkdown": "Sure, i just used a secondary output using diagnosis with a categorical crossentropy loss.  In other words the model will have 2 outputs, 1 for the binary (target column) and another for the multiclass part (diagnosis).",
      "votes": null
    },
    {
      "id": "929240",
      "postDate": "07/14/2020 14:51:04",
      "content": "<p>so basically you train one model to learn two outputs instead one, right? and then you use only one output for the competition, second one is just to make training better?</p>",
      "rawMarkdown": "so basically you train one model to learn two outputs instead one, right? and then you use only one output for the competition, second one is just to make training better?",
      "votes": null
    },
    {
      "id": "929384",
      "postDate": "07/14/2020 16:30:12",
      "content": "<p>Yep exactly that. The main idea is that if you have two outputs the loss of the model is calculated with all the weights of the model (output1 + output2) so basically the secondary output can help the binary output. The main problem is that i still don't have a stable cv scheme so im not sure if it helps or not</p>",
      "rawMarkdown": "Yep exactly that. The main idea is that if you have two outputs the loss of the model is calculated with all the weights of the model (output1 + output2) so basically the secondary output can help the binary output. The main problem is that i still don't have a stable cv scheme so im not sure if it helps or not",
      "votes": null
    },
    {
      "id": "929506",
      "postDate": "07/14/2020 17:57:36",
      "content": "<p>Thanks so much, I was trying to implement it for a long time after seeing Andrej's post about getting a improvement in the score, but couldn't implement it yet, maybe its time! :)</p>",
      "rawMarkdown": "Thanks so much, I was trying to implement it for a long time after seeing Andrej's post about getting a improvement in the score, but couldn't implement it yet, maybe its time! :)",
      "votes": null
    },
    {
      "id": "933673",
      "postDate": "07/17/2020 22:12:13",
      "content": "<p>Before you take the \"raw\" diagnoses, I would personally merge the following classes so you can reuse them with the HAM10000 / ISIC2019 dataset:</p>\n\n<p>unknown --&gt; nv (probably not perfect, but probably the \"best guess\" - pseudolabeling the train-set also seems interesting)\nnevus --&gt; nv\nmelanoma --&gt; mel\nseborrheic keratosis --&gt; bkl (HAM10000-abbreviation for \"seborrheic keratosis\")\nlentigo NOS --&gt; bkl (or nv, hard to tell)\nlichenoid keratosis --&gt; bkl (this is an \"inflamed\" variant of bkl)\nsolar lentigo --&gt; bkl (flat variant of bkl)\ncafe-au-lait macule --&gt; bkl (not really biologically, but probably \"close enough\")\natypical melanocytic proliferation --&gt; either nv or mel, hard to tell</p>\n\n<p>Weighting the second loss class-wise should probably also happen</p>",
      "rawMarkdown": "Before you take the \"raw\" diagnoses, I would personally merge the following classes so you can reuse them with the HAM10000 / ISIC2019 dataset:\n\nunknown --&gt; nv (probably not perfect, but probably the \"best guess\" - pseudolabeling the train-set also seems interesting)\nnevus --&gt; nv\nmelanoma --&gt; mel\nseborrheic keratosis --&gt; bkl (HAM10000-abbreviation for \"seborrheic keratosis\")\nlentigo NOS --&gt; bkl (or nv, hard to tell)\nlichenoid keratosis --&gt; bkl (this is an \"inflamed\" variant of bkl)\nsolar lentigo --&gt; bkl (flat variant of bkl)\ncafe-au-lait macule --&gt; bkl (not really biologically, but probably \"close enough\")\natypical melanocytic proliferation --&gt; either nv or mel, hard to tell\n\nWeighting the second loss class-wise should probably also happen",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 928220,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "07/13/2020 19:51:24",
      "content": "<p>Deploy it on PyTorch look easy! But didn't try.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 928379,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "07/13/2020 23:58:26",
      "content": "<p>```\ntrain_df[\"diagnosis\"].value_counts()\nunknown                               27124\nnevus                                  5193\nmelanoma                                584\nseborrheic keratosis                    135\nlentigo NOS                              44\nlichenoid keratosis                      37\nsolar lentigo                             7\ncafe-au-lait macule                       1\natypical melanocytic proliferation        1\nName: diagnosis, dtype: int64</p>\n\n<p>train_df[train_df[\"target\"] == 1][\"diagnosis\"].value_counts()\nmelanoma    584\nName: diagnosis, dtype: int64\n```</p>\n\n<p>The problem is that 80% is uknown. I think the useful value would be nevus. What's your idea?</p>",
      "votes": null,
      "replies": [
        {
          "id": 928400,
          "author_name": "ragnar123",
          "author_url": "",
          "post_date": "07/14/2020 00:37:34",
          "content": "<p>Yep i was thinking about nevus, maybee that can help the model. It's just and idea 👊 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928417,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/14/2020 01:02:30",
          "content": "<p>it may be, please share your findings :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 928421,
          "author_name": "ragnar123",
          "author_url": "",
          "post_date": "07/14/2020 01:15:48",
          "content": "<p>I just trained a model and the cv did not improve 👎 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929202,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/14/2020 14:20:46",
          "content": "<p><a href=\"/ragnar123\">@ragnar123</a> can you describe what you did? I think to improve score you need multiple models here</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929214,
          "author_name": "ragnar123",
          "author_url": "",
          "post_date": "07/14/2020 14:36:09",
          "content": "<p>Sure, i just used a secondary output using diagnosis with a categorical crossentropy loss.  In other words the model will have 2 outputs, 1 for the binary (target column) and another for the multiclass part (diagnosis).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929240,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/14/2020 14:51:04",
          "content": "<p>so basically you train one model to learn two outputs instead one, right? and then you use only one output for the competition, second one is just to make training better?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 929384,
          "author_name": "ragnar123",
          "author_url": "",
          "post_date": "07/14/2020 16:30:12",
          "content": "<p>Yep exactly that. The main idea is that if you have two outputs the loss of the model is calculated with all the weights of the model (output1 + output2) so basically the secondary output can help the binary output. The main problem is that i still don't have a stable cv scheme so im not sure if it helps or not</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933673,
          "author_name": "tschandl",
          "author_url": "",
          "post_date": "07/17/2020 22:12:13",
          "content": "<p>Before you take the \"raw\" diagnoses, I would personally merge the following classes so you can reuse them with the HAM10000 / ISIC2019 dataset:</p>\n\n<p>unknown --&gt; nv (probably not perfect, but probably the \"best guess\" - pseudolabeling the train-set also seems interesting)\nnevus --&gt; nv\nmelanoma --&gt; mel\nseborrheic keratosis --&gt; bkl (HAM10000-abbreviation for \"seborrheic keratosis\")\nlentigo NOS --&gt; bkl (or nv, hard to tell)\nlichenoid keratosis --&gt; bkl (this is an \"inflamed\" variant of bkl)\nsolar lentigo --&gt; bkl (flat variant of bkl)\ncafe-au-lait macule --&gt; bkl (not really biologically, but probably \"close enough\")\natypical melanocytic proliferation --&gt; either nv or mel, hard to tell</p>\n\n<p>Weighting the second loss class-wise should probably also happen</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 929506,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "07/14/2020 17:57:36",
      "content": "<p>Thanks so much, I was trying to implement it for a long time after seeing Andrej's post about getting a improvement in the score, but couldn't implement it yet, maybe its time! :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "928190": "Are you using diagnosis column as a secondary output for you model? \n\nIf you think about it, maybee this target provides more information so your model can differentiate melanoma.\n\nHere is a stackoverflow discussion were they explain how this work:\n\nhttps://stackoverflow.com/questions/57149476/how-is-a-multiple-outputs-deep-learning-model-trained\n\nCheers and good luck in the competition.",
    "928220": "Deploy it on PyTorch look easy! But didn't try.",
    "928379": "```\ntrain_df[\"diagnosis\"].value_counts()\nunknown                               27124\nnevus                                  5193\nmelanoma                                584\nseborrheic keratosis                    135\nlentigo NOS                              44\nlichenoid keratosis                      37\nsolar lentigo                             7\ncafe-au-lait macule                       1\natypical melanocytic proliferation        1\nName: diagnosis, dtype: int64\n\ntrain_df[train_df[\"target\"] == 1][\"diagnosis\"].value_counts()\nmelanoma    584\nName: diagnosis, dtype: int64\n```\n\nThe problem is that 80% is uknown. I think the useful value would be nevus. What's your idea?",
    "928400": "Yep i was thinking about nevus, maybee that can help the model. It's just and idea 👊",
    "928417": "it may be, please share your findings :)",
    "928421": "I just trained a model and the cv did not improve 👎",
    "929202": "ragnar123 can you describe what you did? I think to improve score you need multiple models here",
    "929214": "Sure, i just used a secondary output using diagnosis with a categorical crossentropy loss.  In other words the model will have 2 outputs, 1 for the binary (target column) and another for the multiclass part (diagnosis).",
    "929240": "so basically you train one model to learn two outputs instead one, right? and then you use only one output for the competition, second one is just to make training better?",
    "929384": "Yep exactly that. The main idea is that if you have two outputs the loss of the model is calculated with all the weights of the model (output1 + output2) so basically the secondary output can help the binary output. The main problem is that i still don't have a stable cv scheme so im not sure if it helps or not",
    "929506": "Thanks so much, I was trying to implement it for a long time after seeing Andrej's post about getting a improvement in the score, but couldn't implement it yet, maybe its time! :)",
    "933673": "Before you take the \"raw\" diagnoses, I would personally merge the following classes so you can reuse them with the HAM10000 / ISIC2019 dataset:\n\nunknown --&gt; nv (probably not perfect, but probably the \"best guess\" - pseudolabeling the train-set also seems interesting)\nnevus --&gt; nv\nmelanoma --&gt; mel\nseborrheic keratosis --&gt; bkl (HAM10000-abbreviation for \"seborrheic keratosis\")\nlentigo NOS --&gt; bkl (or nv, hard to tell)\nlichenoid keratosis --&gt; bkl (this is an \"inflamed\" variant of bkl)\nsolar lentigo --&gt; bkl (flat variant of bkl)\ncafe-au-lait macule --&gt; bkl (not really biologically, but probably \"close enough\")\natypical melanocytic proliferation --&gt; either nv or mel, hard to tell\n\nWeighting the second loss class-wise should probably also happen"
  },
  "source": "meta"
}