{
  "id": 170499,
  "title": "Are you opting for learning from image + metadata at together or seperatley ?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/170499",
  "author_name": "",
  "post_date": "2020-07-28T02:08:15.045416Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>So as I see it there are two ways to go about it. 🤓 \n1. Use metadata and CNN embedding concatenated in the dense layer. Leverage back prop to learn better(?) intermediate representations. 👀 \n2. Use output of CNN embedding + Metadata and train an ensemble of model (With NN weights frozen).  👀 </p>\n\n<p>Which method are you leaning towards ? And why ?</p>\n\n<p>I got a late start in the competition, so trying me some transfer learning from you all 😄 </p>",
  "messages": [
    {
      "id": "948479",
      "postDate": "07/28/2020 02:08:15",
      "content": "<p>So as I see it there are two ways to go about it. 🤓 \n1. Use metadata and CNN embedding concatenated in the dense layer. Leverage back prop to learn better(?) intermediate representations. 👀 \n2. Use output of CNN embedding + Metadata and train an ensemble of model (With NN weights frozen).  👀 </p>\n\n<p>Which method are you leaning towards ? And why ?</p>\n\n<p>I got a late start in the competition, so trying me some transfer learning from you all 😄 </p>",
      "rawMarkdown": "So as I see it there are two ways to go about it. 🤓 \n1. Use metadata and CNN embedding concatenated in the dense layer. Leverage back prop to learn better(?) intermediate representations. 👀 \n2. Use output of CNN embedding + Metadata and train an ensemble of model (With NN weights frozen).  👀 \n\nWhich method are you leaning towards ? And why ?\n\nI got a late start in the competition, so trying me some transfer learning from you all 😄",
      "votes": null
    },
    {
      "id": "948538",
      "postDate": "07/28/2020 03:45:18",
      "content": "<p>While experimentation I found that an ensemble of image + metadata which is averaged for final prediction works better then image+meta multi-input model.\n&gt; Why ?</p>\n\n<p>I do not say that the multi-input model cannot get higher scores it's just you need a good NN architecture that can learn the metadata effectively (And finding a good one was taking much time).</p>",
      "rawMarkdown": "While experimentation I found that an ensemble of image + metadata which is averaged for final prediction works better then image+meta multi-input model.\n&gt; Why ?\n\nI do not say that the multi-input model cannot get higher scores it's just you need a good NN architecture that can learn the metadata effectively (And finding a good one was taking much time).",
      "votes": null
    },
    {
      "id": "952560",
      "postDate": "07/31/2020 03:44:13",
      "content": "<p>Yes, I too feel that it is better than the other option.</p>",
      "rawMarkdown": "Yes, I too feel that it is better than the other option.",
      "votes": null
    },
    {
      "id": "954696",
      "postDate": "08/02/2020 02:21:53",
      "content": "<p>Both are working well. Both have increased LB. However so far, I have only increased CV with method (1). What have others' observed?</p>",
      "rawMarkdown": "Both are working well. Both have increased LB. However so far, I have only increased CV with method (1). What have others' observed?",
      "votes": null
    },
    {
      "id": "955513",
      "postDate": "08/02/2020 17:14:17",
      "content": "<p>Hey ! Chris. Currently I am experimenting with the triple stratified TFRecord  that you  created 👍 🙌 . I will go with method (1) as I want focus/learn more about how I can improve my CV score by tweaking different aspects of the pipeline (LR schedule, backbone, post processing etc). 🤓 </p>",
      "rawMarkdown": "Hey ! Chris. Currently I am experimenting with the triple stratified TFRecord  that you  created 👍 🙌 . I will go with method (1) as I want focus/learn more about how I can improve my CV score by tweaking different aspects of the pipeline (LR schedule, backbone, post processing etc). 🤓",
      "votes": null
    },
    {
      "id": "955529",
      "postDate": "08/02/2020 17:19:24",
      "content": "<p>Sounds great RealSid. There are many notebook examples showing how to use meta data in your CNN. The notebook <a href=\"https://www.kaggle.com/rajnishe/rc-fork-siim-isic-melanoma-384x384?scriptVersionId=39612412\">here</a> (version 3) does a good job. If that notebook doesn't include meta data then the notebook CV decreases, so it looks like <a href=\"/rajnishe\">@rajnishe</a> is successfully using meta data in his CNN</p>\n\n<p>(If you review that notebook link, note that it includes external data in the validation folds, so the CV isn't actually 0.946 but the validation score is still good if you remove external data from validation folds. Also note that they get this great validation score without TTA).</p>",
      "rawMarkdown": "Sounds great RealSid. There are many notebook examples showing how to use meta data in your CNN. The notebook [here][1] (version 3) does a good job. If that notebook doesn't include meta data then the notebook CV decreases, so it looks like @rajnishe is successfully using meta data in his CNN\n\n(If you review that notebook link, note that it includes external data in the validation folds, so the CV isn't actually 0.946 but the validation score is still good if you remove external data from validation folds. Also note that they get this great validation score without TTA).\n\n[1]: https://www.kaggle.com/rajnishe/rc-fork-siim-isic-melanoma-384x384?scriptVersionId=39612412",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 948538,
      "author_name": "prateek0x",
      "author_url": "",
      "post_date": "07/28/2020 03:45:18",
      "content": "<p>While experimentation I found that an ensemble of image + metadata which is averaged for final prediction works better then image+meta multi-input model.\n&gt; Why ?</p>\n\n<p>I do not say that the multi-input model cannot get higher scores it's just you need a good NN architecture that can learn the metadata effectively (And finding a good one was taking much time).</p>",
      "votes": null,
      "replies": [
        {
          "id": 952560,
          "author_name": "jaseemck",
          "author_url": "",
          "post_date": "07/31/2020 03:44:13",
          "content": "<p>Yes, I too feel that it is better than the other option.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 954696,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/02/2020 02:21:53",
      "content": "<p>Both are working well. Both have increased LB. However so far, I have only increased CV with method (1). What have others' observed?</p>",
      "votes": null,
      "replies": [
        {
          "id": 955513,
          "author_name": "realsid",
          "author_url": "",
          "post_date": "08/02/2020 17:14:17",
          "content": "<p>Hey ! Chris. Currently I am experimenting with the triple stratified TFRecord  that you  created 👍 🙌 . I will go with method (1) as I want focus/learn more about how I can improve my CV score by tweaking different aspects of the pipeline (LR schedule, backbone, post processing etc). 🤓 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955529,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/02/2020 17:19:24",
          "content": "<p>Sounds great RealSid. There are many notebook examples showing how to use meta data in your CNN. The notebook <a href=\"https://www.kaggle.com/rajnishe/rc-fork-siim-isic-melanoma-384x384?scriptVersionId=39612412\">here</a> (version 3) does a good job. If that notebook doesn't include meta data then the notebook CV decreases, so it looks like <a href=\"/rajnishe\">@rajnishe</a> is successfully using meta data in his CNN</p>\n\n<p>(If you review that notebook link, note that it includes external data in the validation folds, so the CV isn't actually 0.946 but the validation score is still good if you remove external data from validation folds. Also note that they get this great validation score without TTA).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "948479": "So as I see it there are two ways to go about it. 🤓 \n1. Use metadata and CNN embedding concatenated in the dense layer. Leverage back prop to learn better(?) intermediate representations. 👀 \n2. Use output of CNN embedding + Metadata and train an ensemble of model (With NN weights frozen).  👀 \n\nWhich method are you leaning towards ? And why ?\n\nI got a late start in the competition, so trying me some transfer learning from you all 😄",
    "948538": "While experimentation I found that an ensemble of image + metadata which is averaged for final prediction works better then image+meta multi-input model.\n&gt; Why ?\n\nI do not say that the multi-input model cannot get higher scores it's just you need a good NN architecture that can learn the metadata effectively (And finding a good one was taking much time).",
    "952560": "Yes, I too feel that it is better than the other option.",
    "954696": "Both are working well. Both have increased LB. However so far, I have only increased CV with method (1). What have others' observed?",
    "955513": "Hey ! Chris. Currently I am experimenting with the triple stratified TFRecord  that you  created 👍 🙌 . I will go with method (1) as I want focus/learn more about how I can improve my CV score by tweaking different aspects of the pipeline (LR schedule, backbone, post processing etc). 🤓",
    "955529": "Sounds great RealSid. There are many notebook examples showing how to use meta data in your CNN. The notebook [here][1] (version 3) does a good job. If that notebook doesn't include meta data then the notebook CV decreases, so it looks like @rajnishe is successfully using meta data in his CNN\n\n(If you review that notebook link, note that it includes external data in the validation folds, so the CV isn't actually 0.946 but the validation score is still good if you remove external data from validation folds. Also note that they get this great validation score without TTA).\n\n[1]: https://www.kaggle.com/rajnishe/rc-fork-siim-isic-melanoma-384x384?scriptVersionId=39612412"
  },
  "source": "meta"
}