{
  "id": 155251,
  "title": "Use Both Image Data and Tabular Data",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155251",
  "author_name": "Chris Deotte",
  "post_date": "2020-06-01T00:27:37.485000",
  "votes": 141,
  "comment_count": 106,
  "views": 0,
  "content": "<p>This competition provides both <strong>image data</strong> and <strong>tabular data</strong>. This is exciting! We have seen from this notebook <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">here</a> that using just <strong>images</strong> can score LB 0.910. And we have seen from this notebook <a href=\"https://www.kaggle.com/titericz/simple-baseline\">here</a> that using only the provided <strong>tabular data</strong> can score LB 0.700. Let's discuss ideas how build a model that uses both <strong>images</strong> and <strong>tabular data</strong>.</p>\n\n<p>Three ideas come to mind.\n* Build a CNN image model and find a way to input the tabular data into the CNN image model\n* Build a Tabular data model and find a way to extract image embeddings and input into a Tabular data model\n* Build 2 separate models and ensemble</p>\n\n<p>What other ideas do people have? Has anyone tried one of the above 3 with success?</p>\n\n<p>========================</p>\n\n<p>UPDATE: I did a quick experiment <a href=\"https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\">here</a>. Even though the tabular data predictions have a low LB of 0.700 and the image only predictions have a high LB 0.910, if you add the tabular data predictions to the image data predictions, the LB increases. Thus the tabular data contains new information.</p>\n\n<p>========================</p>\n\n<p>UPDATE: I created TFRecords that contain both the images and tabular data. So you can easily build TensorFlow models that utilize both. Size 768x768 data is <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a>, 512x512 dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">here</a>, 384x384 is <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">here</a>, and 256x256 dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>. A description of the TFRecords fields is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>.</p>",
  "messages": [
    {
      "id": 869339,
      "postDate": "2020-06-01T00:27:37.487Z",
      "content": "<p>This competition provides both <strong>image data</strong> and <strong>tabular data</strong>. This is exciting! We have seen from this notebook <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">here</a> that using just <strong>images</strong> can score LB 0.910. And we have seen from this notebook <a href=\"https://www.kaggle.com/titericz/simple-baseline\">here</a> that using only the provided <strong>tabular data</strong> can score LB 0.700. Let's discuss ideas how build a model that uses both <strong>images</strong> and <strong>tabular data</strong>.</p>\n\n<p>Three ideas come to mind.\n* Build a CNN image model and find a way to input the tabular data into the CNN image model\n* Build a Tabular data model and find a way to extract image embeddings and input into a Tabular data model\n* Build 2 separate models and ensemble</p>\n\n<p>What other ideas do people have? Has anyone tried one of the above 3 with success?</p>\n\n<p>========================</p>\n\n<p>UPDATE: I did a quick experiment <a href=\"https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\">here</a>. Even though the tabular data predictions have a low LB of 0.700 and the image only predictions have a high LB 0.910, if you add the tabular data predictions to the image data predictions, the LB increases. Thus the tabular data contains new information.</p>\n\n<p>========================</p>\n\n<p>UPDATE: I created TFRecords that contain both the images and tabular data. So you can easily build TensorFlow models that utilize both. Size 768x768 data is <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a>, 512x512 dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">here</a>, 384x384 is <a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">here</a>, and 256x256 dataset is <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a>. A description of the TFRecords fields is <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a>.</p>",
      "rawMarkdown": "This competition provides both **image data** and **tabular data**. This is exciting! We have seen from this notebook [here][1] that using just **images** can score LB 0.910. And we have seen from this notebook [here][2] that using only the provided **tabular data** can score LB 0.700. Let's discuss ideas how build a model that uses both **images** and **tabular data**.\n\nThree ideas come to mind.\n* Build a CNN image model and find a way to input the tabular data into the CNN image model\n* Build a Tabular data model and find a way to extract image embeddings and input into a Tabular data model\n* Build 2 separate models and ensemble\n\nWhat other ideas do people have? Has anyone tried one of the above 3 with success?\n\n========================\n\nUPDATE: I did a quick experiment [here][3]. Even though the tabular data predictions have a low LB of 0.700 and the image only predictions have a high LB 0.910, if you add the tabular data predictions to the image data predictions, the LB increases. Thus the tabular data contains new information.\n\n========================\n\nUPDATE: I created TFRecords that contain both the images and tabular data. So you can easily build TensorFlow models that utilize both. Size 768x768 data is [here][6], 512x512 dataset is [here][4], 384x384 is [here][7], and 256x256 dataset is [here][5]. A description of the TFRecords fields is [here][8].\n\n[1]: https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\n[2]: https://www.kaggle.com/titericz/simple-baseline\n[3]: https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\n[4]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[5]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[6]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[7]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[8]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579",
      "votes": 140
    },
    {
      "id": 869369,
      "postDate": "2020-06-01T01:04:59.523Z",
      "content": "<p>the past isic leaderboard website has a list of methods that uses image or image+table:\n<a href=\"https://challenge2019.isic-archive.com/leaderboard.html\">https://challenge2019.isic-archive.com/leaderboard.html</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F62811b28d3e8c6c43b8564c09e5f6ff8%2FSelection_041.png?generation=1590973485561188&amp;alt=media\" alt=\"\"></p>\n\n<p>i am implementing transformer based methods</p>",
      "rawMarkdown": "the past isic leaderboard website has a list of methods that uses image or image+table:\nhttps://challenge2019.isic-archive.com/leaderboard.html\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F62811b28d3e8c6c43b8564c09e5f6ff8%2FSelection_041.png?generation=1590973485561188&amp;alt=media)\n\n\ni am implementing transformer based methods",
      "votes": 17,
      "replies": [
        {
          "id": 869393,
          "postDate": "2020-06-01T01:50:48.083Z",
          "content": "<p>Transformers aren't more adequate for NLP problems?</p>",
          "rawMarkdown": "Transformers aren't more adequate for NLP problems?"
        },
        {
          "id": 869504,
          "postDate": "2020-06-01T04:03:43.197Z",
          "content": "<p>Transformers are not for NLP only ..And if you are someone like <a href=\"/limerobot\">@limerobot</a> , you can use them whenever you have the opportunity : See <a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891\">here</a> and  <a href=\"https://www.kaggle.com/c/champs-scalar-coupling/discussion/106572\">here</a></p>",
          "rawMarkdown": "Transformers are not for NLP only ..And if you are someone like @limerobot , you can use them whenever you have the opportunity : See [here](https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891) and  [here](https://www.kaggle.com/c/champs-scalar-coupling/discussion/106572)",
          "votes": 3
        },
        {
          "id": 869507,
          "postDate": "2020-06-01T04:05:39.730Z",
          "rawMarkdown": ""
        },
        {
          "id": 869512,
          "postDate": "2020-06-01T04:11:41.177Z",
          "content": "<p>i am trying to modify this : \"DETR: End-to-End Object Detection with Transformers\"</p>\n\n<p>i am more concerned with \"multiple images per patient\". so we have a sequence of images rather then one images. i think this will have significant impact on the score.</p>",
          "rawMarkdown": "i am trying to modify this : \"DETR: End-to-End Object Detection with Transformers\"\n\ni am more concerned with \"multiple images per patient\". so we have a sequence of images rather then one images. i think this will have significant impact on the score.\n\n",
          "votes": 9
        },
        {
          "id": 869655,
          "postDate": "2020-06-01T07:25:21.437Z",
          "content": "<p>@Heng that is very interesting thing for me also . But  I am not sure how I can a relate a patient with a skin image taken from neck at the age of 35 as Melanoma = 0 , for the same patient if i take another skin image at another bodypart at the age of 40 . and find melanoma =1 . May be this timeseries relationship is little complex , need to figure out .</p>",
          "rawMarkdown": "@Heng that is very interesting thing for me also . But  I am not sure how I can a relate a patient with a skin image taken from neck at the age of 35 as Melanoma = 0 , for the same patient if i take another skin image at another bodypart at the age of 40 . and find melanoma =1 . May be this timeseries relationship is little complex , need to figure out ."
        },
        {
          "id": 869670,
          "postDate": "2020-06-01T07:33:01.447Z",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> </p>\n\n<p>it also have to do with the way data is collected. i think the data has been processed (e.g. images filtered by doctors for recording purpose? ). in particular, from pure data analysis point of view:</p>\n\n<ul>\n<li><p>for a given patient and location (e.g.  upper extremity, torso, etc), there at at most few +ve images. the rest are negative. i.e. the presence some high score images will lower posterior scores of others.</p></li>\n<li><p>for a given patient, there are at most one or two locations with +ve  images. similarly, the presence some high score at some location will lower posterior scores of others.</p></li>\n</ul>\n\n<p>if a model learn such biases, LB score can be improved a lot. but it doesn't make more accurate melanoma detection in real life application. </p>\n\n<p>this is a problem and dilemma  in data science competition. should you exploit data biases for better scores or build a useful model (possibly with lower score)?</p>",
          "rawMarkdown": "@phoenix9032 \n\n it also have to do with the way data is collected. i think the data has been processed (e.g. images filtered by doctors for recording purpose? ). in particular, from pure data analysis point of view:\n\n- for a given patient and location (e.g.  upper extremity, torso, etc), there at at most few +ve images. the rest are negative. i.e. the presence some high score images will lower posterior scores of others.\n\n- for a given patient, there are at most one or two locations with +ve  images. similarly, the presence some high score at some location will lower posterior scores of others.\n\nif a model learn such biases, LB score can be improved a lot. but it doesn't make more accurate melanoma detection in real life application. \n\nthis is a problem and dilemma  in data science competition. should you exploit data biases for better scores or build a useful model (possibly with lower score)?",
          "votes": 1
        },
        {
          "id": 870749,
          "postDate": "2020-06-01T22:04:00.370Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> your actual submission was using DETR? </p>",
          "rawMarkdown": "@hengck23 your actual submission was using DETR? "
        },
        {
          "id": 873482,
          "postDate": "2020-06-04T07:28:52.593Z",
          "content": "<p><a href=\"/hiramcho\">@hiramcho</a> \n\"<a href=\"/hengck23\">@hengck23</a> your actual submission was using DETR?\"\nnot yet</p>",
          "rawMarkdown": "@hiramcho \n\"@hengck23 your actual submission was using DETR?\"\nnot yet",
          "votes": 1
        }
      ]
    },
    {
      "id": 875367,
      "postDate": "2020-06-05T17:51:47.783Z",
      "content": "<p>this is the prior probability from <a href=\"https://en.wikipedia.org/wiki/Melanoma\">https://en.wikipedia.org/wiki/Melanoma</a>\n(instead from the train.csv file)</p>\n\n<p><img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/3/37/Diagram_showing_where_melanoma_is_most_likely_to_develop_CRUK_383.svg/800px-Diagram_showing_where_melanoma_is_most_likely_to_develop_CRUK_383.svg.png\" alt=\"\"></p>",
      "rawMarkdown": "this is the prior probability from https://en.wikipedia.org/wiki/Melanoma\n(instead from the train.csv file)\n\n![](https://upload.wikimedia.org/wikipedia/commons/thumb/3/37/Diagram_showing_where_melanoma_is_most_likely_to_develop_CRUK_383.svg/800px-Diagram_showing_where_melanoma_is_most_likely_to_develop_CRUK_383.svg.png)",
      "votes": 13
    },
    {
      "id": 871604,
      "postDate": "2020-06-02T13:57:07.667Z",
      "content": "<p><a href=\"https://ai.googleblog.com/2019/09/using-deep-learning-to-inform.html\">https://ai.googleblog.com/2019/09/using-deep-learning-to-inform.html</a>\n“A Deep Learning System for Differential Diagnosis of Skin Diseases” \n<img src=\"https://1.bp.blogspot.com/-b3HBRhjUZs0/XXoeACfbsmI/AAAAAAAAEoc/f4aL1TH8J2g6Ix5Tlfj-rYXemzRgPmVXwCEwYBhgL/s640/image1.png\" alt=\"\"></p>\n\n<p><img src=\"https://1.bp.blogspot.com/-wMnHrj5PPys/XXoeAKVc2xI/AAAAAAAAEoI/64V1-8jUT6s-BuBRjc-EmjIpWvR7W2yZQCLcBGAsYHQ/s640/image3.png\" alt=\"\"></p>",
      "rawMarkdown": "https://ai.googleblog.com/2019/09/using-deep-learning-to-inform.html\n“A Deep Learning System for Differential Diagnosis of Skin Diseases” \n![](https://1.bp.blogspot.com/-b3HBRhjUZs0/XXoeACfbsmI/AAAAAAAAEoc/f4aL1TH8J2g6Ix5Tlfj-rYXemzRgPmVXwCEwYBhgL/s640/image1.png)\n\n![](https://1.bp.blogspot.com/-wMnHrj5PPys/XXoeAKVc2xI/AAAAAAAAEoI/64V1-8jUT6s-BuBRjc-EmjIpWvR7W2yZQCLcBGAsYHQ/s640/image3.png)",
      "votes": 11
    },
    {
      "id": 952589,
      "postDate": "2020-07-31T04:18:06.783Z",
      "content": "<p>Took tabular data from <a href=\"https://www.kaggle.com/awsaf49/xgboost-tabular-data-ml-cv-85-lb-787\">this</a> kernel and made images from the scaled (0..1) outer product of the features.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1828058%2Fc55ad02be5a4ca92ea13029b6955ef78%2FTabImage.png?generation=1596168989462514&amp;alt=media\" alt=\"\">\nAdded to a simple 4 channel cnn model and boosted LB score... wonder how is best to mingle this in EfficientNet?</p>",
      "rawMarkdown": "Took tabular data from [this](https://www.kaggle.com/awsaf49/xgboost-tabular-data-ml-cv-85-lb-787) kernel and made images from the scaled (0..1) outer product of the features.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1828058%2Fc55ad02be5a4ca92ea13029b6955ef78%2FTabImage.png?generation=1596168989462514&amp;alt=media)\nAdded to a simple 4 channel cnn model and boosted LB score... wonder how is best to mingle this in EfficientNet?",
      "votes": 10,
      "replies": [
        {
          "id": 952596,
          "postDate": "2020-07-31T04:30:48.403Z",
          "content": "<p>this is awesome , any resource to teach how to do that ??</p>",
          "rawMarkdown": "this is awesome , any resource to teach how to do that ??",
          "votes": 1
        },
        {
          "id": 952626,
          "postDate": "2020-07-31T05:14:21.500Z",
          "content": "<p>Wow, this is awesome <a href=\"/kittlein\">@kittlein</a> , great idea. Each of these images is 8x8. Perhaps you can tile these into 256x256 and then concatenate that with images of size 256x256x3. Then if you use <code>tf.keras.layers.Conv2D(3,8)</code> you could incorporate all the info into each pixel of the original images. And feed that into pretrained EfficientNet.</p>\n\n<p><a href=\"/phoenix9032\">@phoenix9032</a> for each row of train data, let's say you have 8 meta features. First scale them to between 0 and 1. Next create a vector <code>x</code> of size <code>8x1</code>. Lastly if you if multiply <code>x</code> by <code>x.T</code> the transpose of <code>x</code> the result is an 8x8 matrix. These 8x8 matrices are displayed above.</p>",
          "rawMarkdown": "Wow, this is awesome @kittlein , great idea. Each of these images is 8x8. Perhaps you can tile these into 256x256 and then concatenate that with images of size 256x256x3. Then if you use `tf.keras.layers.Conv2D(3,8)` you could incorporate all the info into each pixel of the original images. And feed that into pretrained EfficientNet.\n\n@phoenix9032 for each row of train data, let's say you have 8 meta features. First scale them to between 0 and 1. Next create a vector `x` of size `8x1`. Lastly if you if multiply `x` by `x.T` the transpose of `x` the result is an 8x8 matrix. These 8x8 matrices are displayed above.\n\n",
          "votes": 4
        }
      ]
    },
    {
      "id": 871389,
      "postDate": "2020-06-02T10:15:07.797Z",
      "content": "<p>Lol, what did you do! :)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2070357%2Fa0d0dc316725c83ae5fd05dd16f7fa13%2FScreenshot%20from%202020-06-02%2015-42-31.png?generation=1591092841763019&amp;alt=media\" alt=\"\"></p>\n\n<p>Some sort of a benchmark. :/</p>",
      "rawMarkdown": "Lol, what did you do! :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2070357%2Fa0d0dc316725c83ae5fd05dd16f7fa13%2FScreenshot%20from%202020-06-02%2015-42-31.png?generation=1591092841763019&amp;alt=media)\n\n\nSome sort of a benchmark. :/",
      "votes": 7,
      "replies": [
        {
          "id": 871720,
          "postDate": "2020-06-02T15:40:17.493Z",
          "content": "<p>haha, maybe its just a coincidence that everyone got the same LB of 0.914 as my notebook 😄  </p>",
          "rawMarkdown": "haha, maybe its just a coincidence that everyone got the same LB of 0.914 as my notebook 😄  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 869748,
      "postDate": "2020-06-01T08:59:11.683Z",
      "content": "<p>A picture from the topic right next to yours\n<img src=\"https://ars.els-cdn.com/content/image/1-s2.0-S2215016120300832-gr2.jpg\" alt=\"\"></p>",
      "rawMarkdown": "A picture from the topic right next to yours\n![](https://ars.els-cdn.com/content/image/1-s2.0-S2215016120300832-gr2.jpg)",
      "votes": 7
    },
    {
      "id": 962846,
      "postDate": "2020-08-08T13:35:45.760Z",
      "content": "<p>I just recently shared the notebook for using dual-input CNNs for this problem. Please check it out !</p>\n<p><a href=\"https://www.kaggle.com/niteshx2/pipeline-dual-input-cnn?scriptVersionId=40368875\" target=\"_blank\">https://www.kaggle.com/niteshx2/pipeline-dual-input-cnn?scriptVersionId=40368875</a></p>\n<p><img src=\"https://ars.els-cdn.com/content/image/1-s2.0-S2215016120300832-gr2.jpg\" alt=\"\"></p>",
      "rawMarkdown": "I just recently shared the notebook for using dual-input CNNs for this problem. Please check it out !\n\nhttps://www.kaggle.com/niteshx2/pipeline-dual-input-cnn?scriptVersionId=40368875\n\n![](https://ars.els-cdn.com/content/image/1-s2.0-S2215016120300832-gr2.jpg)",
      "votes": 5,
      "replies": [
        {
          "id": 962871,
          "postDate": "2020-08-08T14:02:25.727Z",
          "content": "<p>Great job Nitesh, thanks for posting. Your notebook is a helpful example on how to combine meta into CNN.</p>",
          "rawMarkdown": "Great job Nitesh, thanks for posting. Your notebook is a helpful example on how to combine meta into CNN."
        }
      ]
    },
    {
      "id": 875901,
      "postDate": "2020-06-06T08:49:06.017Z",
      "content": "<p><a href=\"https://www.youtube.com/watch?v=ysBaZO8YmX8\">https://www.youtube.com/watch?v=ysBaZO8YmX8</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F08f597f592772bbbab337cb3f8d54804%2FSelection_098.png?generation=1591433343089407&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "https://www.youtube.com/watch?v=ysBaZO8YmX8\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F08f597f592772bbbab337cb3f8d54804%2FSelection_098.png?generation=1591433343089407&amp;alt=media)\n",
      "votes": 6,
      "replies": [
        {
          "id": 876325,
          "postDate": "2020-06-06T16:16:06.140Z",
          "content": "<p>Interesting. Thanks for sharing. The slides for that video are posted <a href=\"https://www.slideshare.net/SebastienFischman/tab-netpresentation/SebastienFischman/tab-netpresentation\">here</a></p>",
          "rawMarkdown": "Interesting. Thanks for sharing. The slides for that video are posted [here][1]\n\n[1]: https://www.slideshare.net/SebastienFischman/tab-netpresentation/SebastienFischman/tab-netpresentation",
          "votes": 5
        },
        {
          "id": 876337,
          "postDate": "2020-06-06T16:20:16.533Z",
          "content": "<p><a href=\"/optimo\">@optimo</a>  tagging you once more , while I am trying the new version in other comp.</p>",
          "rawMarkdown": "@optimo  tagging you once more , while I am trying the new version in other comp.",
          "votes": 1
        },
        {
          "id": 877060,
          "postDate": "2020-06-07T09:33:13.493Z",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> I don't know anything about the other comp but there is a new kernel using multi target regression if you want <a href=\"https://www.kaggle.com/tanulsingh077/achieving-sota-results-with-tabnet/\">https://www.kaggle.com/tanulsingh077/achieving-sota-results-with-tabnet/</a></p>",
          "rawMarkdown": "@phoenix9032 I don't know anything about the other comp but there is a new kernel using multi target regression if you want https://www.kaggle.com/tanulsingh077/achieving-sota-results-with-tabnet/",
          "votes": 4
        },
        {
          "id": 877359,
          "postDate": "2020-06-07T14:30:07.383Z",
          "content": "<p>Thank you <a href=\"/optimo\">@optimo</a> I believe that Tabnet has got a lot potential. <a href=\"/cdeotte\">@cdeotte</a> I am. Planning on doing some research with Tabnet for this competition as well, after tweet competition ends, only if you don't do ut earlier😜</p>",
          "rawMarkdown": "Thank you @optimo I believe that Tabnet has got a lot potential. @cdeotte I am. Planning on doing some research with Tabnet for this competition as well, after tweet competition ends, only if you don't do ut earlier😜"
        },
        {
          "id": 877362,
          "postDate": "2020-06-07T14:33:54.610Z",
          "content": "<p>Yes  <a href=\"/optimo\">@optimo</a>  I started using multitarget in the TReNDs neuroimaging comp and its indeed giving me better score , I have not made my kernel public as <a href=\"/tanulsingh077\">@tanulsingh077</a>  has already done it well ..</p>",
          "rawMarkdown": "Yes  @optimo  I started using multitarget in the TReNDs neuroimaging comp and its indeed giving me better score , I have not made my kernel public as @tanulsingh077  has already done it well ..",
          "votes": 1
        }
      ]
    },
    {
      "id": 871997,
      "postDate": "2020-06-02T19:48:05.167Z",
      "content": "<p>Meta_data definitely helps.  My TPU Quota is almost exhausted (I use it for Jigsaw competition too ) but the first fold with Image+meta gives me 0.909 on LB </p>\n\n<p>No Augmentation, No external data for now</p>",
      "rawMarkdown": "Meta_data definitely helps.  My TPU Quota is almost exhausted (I use it for Jigsaw competition too ) but the first fold with Image+meta gives me 0.909 on LB \n\nNo Augmentation, No external data for now\n",
      "votes": 6,
      "replies": [
        {
          "id": 872051,
          "postDate": "2020-06-02T21:15:11.847Z",
          "content": "<p>Nice</p>",
          "rawMarkdown": "Nice"
        },
        {
          "id": 872061,
          "postDate": "2020-06-02T21:31:18.083Z",
          "content": "<p>Thanks !</p>\n\n<p>LB  0.912 for the second fold...TPU Quota completely exhausted :(</p>",
          "rawMarkdown": "Thanks !\n\nLB  0.912 for the second fold...TPU Quota completely exhausted :(",
          "votes": 3
        },
        {
          "id": 872770,
          "postDate": "2020-06-03T13:41:04.663Z",
          "content": "<p>Cool!</p>",
          "rawMarkdown": "Cool!"
        },
        {
          "id": 872780,
          "postDate": "2020-06-03T13:52:34.903Z",
          "content": "<p>have you checked without meta? i get 0.919 without meta and 512 input</p>",
          "rawMarkdown": "have you checked without meta? i get 0.919 without meta and 512 input\n",
          "votes": 2
        },
        {
          "id": 873075,
          "postDate": "2020-06-03T18:52:58.887Z",
          "content": "<p>I got 0.905 for all 5 folds  and 768 as input size</p>\n\n<p>But it was first or second day of the competition . I changed the loss function in the last experiment.</p>",
          "rawMarkdown": "I got 0.905 for all 5 folds  and 768 as input size\n\nBut it was first or second day of the competition . I changed the loss function in the last experiment.\n\n",
          "votes": 2
        },
        {
          "id": 874240,
          "postDate": "2020-06-04T18:24:02.120Z",
          "content": "<p>thx</p>",
          "rawMarkdown": "thx"
        }
      ]
    },
    {
      "id": 871264,
      "postDate": "2020-06-02T08:18:18.173Z",
      "content": "<p>UPDATE: I created TFRecords that contain both the meta data and image data <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a> for 256x256 and <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">here</a> for 512x512. All of the tabular data is inside as fields in addition to the images.</p>\n\n<pre><code>  feature = {\n      'image': _bytes_feature,\n      'image_name': _bytes_feature,\n      'patient_id': _int64_feature,\n      'sex': _int64_feature,\n      'age_approx': _int64_feature,\n      'anatom_site_general_challenge': _int64_feature,\n      'diagnosis': _int64_feature,\n      'target': _int64_feature\n  }\n</code></pre>",
      "rawMarkdown": "UPDATE: I created TFRecords that contain both the meta data and image data [here][1] for 256x256 and [here][2] for 512x512. All of the tabular data is inside as fields in addition to the images.\n\n      feature = {\n          'image': _bytes_feature,\n          'image_name': _bytes_feature,\n          'patient_id': _int64_feature,\n          'sex': _int64_feature,\n          'age_approx': _int64_feature,\n          'anatom_site_general_challenge': _int64_feature,\n          'diagnosis': _int64_feature,\n          'target': _int64_feature\n      }\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[2]: https://www.kaggle.com/cdeotte/melanoma-512x512",
      "votes": 6,
      "replies": [
        {
          "id": 871273,
          "postDate": "2020-06-02T08:31:20.187Z",
          "content": "<p>Fantastic thanks :)</p>",
          "rawMarkdown": "Fantastic thanks :)",
          "votes": 1
        },
        {
          "id": 871725,
          "postDate": "2020-06-02T15:41:44.403Z",
          "content": "<p>I remember 512x512 being a nice size in Flower Comp, so I thought it would be good here too. I'm also uploading 768x768 and 384x384 now. And i plan to upload 192x192 and 128x128 TFRecords</p>",
          "rawMarkdown": "I remember 512x512 being a nice size in Flower Comp, so I thought it would be good here too. I'm also uploading 768x768 and 384x384 now. And i plan to upload 192x192 and 128x128 TFRecords"
        },
        {
          "id": 872947,
          "postDate": "2020-06-03T16:12:26.733Z",
          "content": "<p>Very interesting ... any chance of releasnig the code for creating these datasets?</p>\n\n<p>I wonder ... why stop only with those meta data.\nIf you check you can see there is a difference between distribution of other image variables for 'target' 1 and 0 ... (i.e. 'standard deviation' by channel/without channels) and some other parameters. \nSome of them could be very useful.</p>\n\n<p>And one other question ... why did you put 'diagnose' in features? \nThere is no 'diagnose' for test images.</p>",
          "rawMarkdown": "Very interesting ... any chance of releasnig the code for creating these datasets?\n\nI wonder ... why stop only with those meta data.\nIf you check you can see there is a difference between distribution of other image variables for 'target' 1 and 0 ... (i.e. 'standard deviation' by channel/without channels) and some other parameters. \nSome of them could be very useful.\n\nAnd one other question ... why did you put 'diagnose' in features? \nThere is no 'diagnose' for test images.",
          "votes": 1
        },
        {
          "id": 872960,
          "postDate": "2020-06-03T16:38:46.403Z",
          "content": "<p>You can use <code>diagnosis</code> as an auxiliary target. For example you train your model to predict both <code>target</code> and <code>diagnosis</code>. Then when you infer test, you just ignore your model's <code>diagnosis</code> prediction and keep your model's <code>target</code> prediction and submit that to this comp.</p>\n\n<p>That is called auxiliary training with an auxiliary head and can help your model make better predictions for target.</p>",
          "rawMarkdown": "You can use `diagnosis` as an auxiliary target. For example you train your model to predict both `target` and `diagnosis`. Then when you infer test, you just ignore your model's `diagnosis` prediction and keep your model's `target` prediction and submit that to this comp.\n\nThat is called auxiliary training with an auxiliary head and can help your model make better predictions for target.",
          "votes": 4
        },
        {
          "id": 873072,
          "postDate": "2020-06-03T18:49:39.200Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> how would you treat \"unknown\" diagnosis?</p>\n\n<p>Theoretically, it can be any of the remaining classes or something else. Here are my thoughts:\n0. unknown as a separate class  (probably sub-optimal)\n1. filtering - discard loss from unknown\n2. pseudo labels / knowledge distillation - train on examples from known diagnosis and predict unknown</p>",
          "rawMarkdown": "@cdeotte how would you treat \"unknown\" diagnosis?\n \nTheoretically, it can be any of the remaining classes or something else. Here are my thoughts:\n0. unknown as a separate class  (probably sub-optimal)\n1. filtering - discard loss from unknown\n2. pseudo labels / knowledge distillation - train on examples from known diagnosis and predict unknown",
          "votes": 2
        },
        {
          "id": 874113,
          "postDate": "2020-06-04T16:35:40.977Z",
          "content": "<p>how good is progressive resizing in this competition?</p>",
          "rawMarkdown": "how good is progressive resizing in this competition?"
        },
        {
          "id": 882971,
          "postDate": "2020-06-12T09:12:37.190Z",
          "content": "<p>any more suggestions? thanks.</p>",
          "rawMarkdown": "any more suggestions? thanks."
        }
      ]
    },
    {
      "id": 948466,
      "postDate": "2020-07-28T01:32:14.817Z",
      "content": "<p>I created this one which could potentially help \n<a href=\"https://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta\">https://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta</a></p>",
      "rawMarkdown": "I created this one which could potentially help \nhttps://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta",
      "votes": 3,
      "replies": [
        {
          "id": 948469,
          "postDate": "2020-07-28T01:39:02.527Z",
          "content": "<p>Great job. Stacking image OOF with meta features is a good idea.</p>",
          "rawMarkdown": "Great job. Stacking image OOF with meta features is a good idea."
        }
      ]
    },
    {
      "id": 872894,
      "postDate": "2020-06-03T15:36:44.640Z",
      "content": "<p>got me from 0.927 to 0.932 thx</p>\n\n<p>TTA works well either - but gain variates a lot - I got +0.009 once, but usually more around +0.004</p>",
      "rawMarkdown": "got me from 0.927 to 0.932 thx\n\nTTA works well either - but gain variates a lot - I got +0.009 once, but usually more around +0.004",
      "votes": 3
    },
    {
      "id": 870795,
      "postDate": "2020-06-01T23:29:07.717Z",
      "content": "<p>We can apply different image model(like-efficientnet B0, B6,resnet, etc.) and apply stacking using the predicted values.\nIn one of the Kaggle meet up, the winner team from the previously held competition on \"skin cancer diagnosis\" discussed their approach, where they used stacking like this too.\n<a href=\"https://www.youtube.com/watch?v=meg8I1GdtUg\">https://www.youtube.com/watch?v=meg8I1GdtUg</a></p>",
      "rawMarkdown": "We can apply different image model(like-efficientnet B0, B6,resnet, etc.) and apply stacking using the predicted values.\nIn one of the Kaggle meet up, the winner team from the previously held competition on \"skin cancer diagnosis\" discussed their approach, where they used stacking like this too.\nhttps://www.youtube.com/watch?v=meg8I1GdtUg",
      "votes": 3
    },
    {
      "id": 869701,
      "postDate": "2020-06-01T08:15:05.153Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I saw Dieter's kernel: \"Extract Image Features from Pretrained NN\" and I will keep people posted on my implementation of \"feature extraction + XGB\" combo here.</p>",
      "rawMarkdown": "@cdeotte I saw Dieter's kernel: \"Extract Image Features from Pretrained NN\" and I will keep people posted on my implementation of \"feature extraction + XGB\" combo here.",
      "votes": 3,
      "replies": [
        {
          "id": 869715,
          "postDate": "2020-06-01T08:29:15.573Z",
          "content": "<p>Isnt that the same <a href=\"/tunguz\">@tunguz</a>  did here ?</p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling\">https://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling</a></p>",
          "rawMarkdown": "Isnt that the same @tunguz  did here ?\n\nhttps://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling"
        },
        {
          "id": 870618,
          "postDate": "2020-06-01T19:09:31.897Z",
          "content": "<p>Great idea Trigram. That is my option 2 above \"Build a Tabular data model and find a way to extract image embeddings and input into a Tabular data model\". </p>\n\n<p><a href=\"/phoenix9032\">@phoenix9032</a> What Bojan did is a very simple form of it. He has converted each image into a 1024 dimension vector. But his vector are just the pixel values from a 32x32 resized image. </p>\n\n<p>You can get a better 1024 dimension vector (embedding) by inputting the original full size image into a pretrained CNN and taking the output from the last convolutional layer and applying a <code>GlobalAveragePooling2D</code>. That will also be a 1024 dimension vector but that vector will contain much more information about the image.</p>\n\n<p>Each element of this better 1024 dimension vector will represent whether a certain feature is present in the image or not because these elements are results of convolutions. (As opposed to the elements in the basic 1024 dimension vector which just represent whether a certain pixel is present or not).</p>",
          "rawMarkdown": "Great idea Trigram. That is my option 2 above \"Build a Tabular data model and find a way to extract image embeddings and input into a Tabular data model\". \n\n@phoenix9032 What Bojan did is a very simple form of it. He has converted each image into a 1024 dimension vector. But his vector are just the pixel values from a 32x32 resized image. \n  \nYou can get a better 1024 dimension vector (embedding) by inputting the original full size image into a pretrained CNN and taking the output from the last convolutional layer and applying a `GlobalAveragePooling2D`. That will also be a 1024 dimension vector but that vector will contain much more information about the image.\n\nEach element of this better 1024 dimension vector will represent whether a certain feature is present in the image or not because these elements are results of convolutions. (As opposed to the elements in the basic 1024 dimension vector which just represent whether a certain pixel is present or not).",
          "votes": 3
        },
        {
          "id": 870667,
          "postDate": "2020-06-01T19:52:59.537Z",
          "content": "<p>Got it . </p>\n\n<ol>\n<li>effnet --&gt; feature --&gt; GAP --&gt; (Embedding + Tabular) --&gt; LGBM </li>\n<li>effnet --&gt; feature --&gt; GAP --&gt; (Embedding + Tabular)  --&gt; NN </li>\n</ol>\n\n<p>I am trying 2 , Trigram trying 1.</p>",
          "rawMarkdown": "Got it . \n\n1. effnet --&gt; feature --&gt; GAP --&gt; (Embedding + Tabular) --&gt; LGBM \n2.  effnet --&gt; feature --&gt; GAP --&gt; (Embedding + Tabular)  --&gt; NN \n\nI am trying 2 , Trigram trying 1.",
          "votes": 1
        }
      ]
    },
    {
      "id": 871685,
      "postDate": "2020-06-02T15:12:18.707Z",
      "content": "<p>Of course metadata is important. Simply looking at the age/malignancy correlation (remember: <strong>correlation does not mean causation</strong>) and using naive Bayes you can have ~70% accuracy.</p>",
      "rawMarkdown": "Of course metadata is important. Simply looking at the age/malignancy correlation (remember: **correlation does not mean causation**) and using naive Bayes you can have ~70% accuracy.",
      "votes": 4,
      "replies": [
        {
          "id": 871732,
          "postDate": "2020-06-02T15:47:51.093Z",
          "content": "<p>Great point. Yes when i look at the images myself they all look the same, but when age is high, then i would more likely guess malignant. (So if that's what i do, our models would probably benefit from knowing age too).</p>",
          "rawMarkdown": "Great point. Yes when i look at the images myself they all look the same, but when age is high, then i would more likely guess malignant. (So if that's what i do, our models would probably benefit from knowing age too).",
          "votes": 3
        }
      ]
    },
    {
      "id": 870175,
      "postDate": "2020-06-01T14:41:32.193Z",
      "content": "<p>UPDATE: I did a quick experiment <a href=\"https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\">here</a>. Even though the tabular data predictions have a low LB of 0.700 <a href=\"https://www.kaggle.com/titericz/simple-baseline\">here</a> and the image only predictions have a high LB 0.910 <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">here</a>, if you add the tabular data predictions to the image data predictions, the LB increases. Thus the tabular data contains new information.</p>",
      "rawMarkdown": "UPDATE: I did a quick experiment [here][3]. Even though the tabular data predictions have a low LB of 0.700 [here][2] and the image only predictions have a high LB 0.910 [here][1], if you add the tabular data predictions to the image data predictions, the LB increases. Thus the tabular data contains new information.\n\n[1]: https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\n[2]: https://www.kaggle.com/titericz/simple-baseline\n[3]: https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\n",
      "votes": 4,
      "replies": [
        {
          "id": 871358,
          "postDate": "2020-06-02T09:41:24.490Z",
          "content": "<p>just confirming that my score also increased 0.926 =&gt; 0.930. A lot more juice to squeeze here....</p>",
          "rawMarkdown": "just confirming that my score also increased 0.926 =&gt; 0.930. A lot more juice to squeeze here....",
          "votes": 4
        },
        {
          "id": 871952,
          "postDate": "2020-06-02T18:48:14.597Z",
          "content": "<p>Yup, even for me..score increased from 0.919 to 0.923</p>",
          "rawMarkdown": "Yup, even for me..score increased from 0.919 to 0.923",
          "votes": 2
        }
      ]
    },
    {
      "id": 869340,
      "postDate": "2020-06-01T00:32:18.960Z",
      "content": "<p>what do you think about this approach <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880\">\"Dance with Ensembl\"</a> ? I'm going spend the rest of the competition building something based on that. </p>",
      "rawMarkdown": "what do you think about this approach [\"Dance with Ensembl\"](https://www.kaggle.com/c/avito-demand-prediction/discussion/59880) ? I'm going spend the rest of the competition building something based on that. ",
      "votes": 4,
      "replies": [
        {
          "id": 869346,
          "postDate": "2020-06-01T00:39:25.017Z",
          "content": "<p>Wow, awesome model and it won 1st place in another competition. I think that has great potential here.</p>",
          "rawMarkdown": "Wow, awesome model and it won 1st place in another competition. I think that has great potential here."
        },
        {
          "id": 869352,
          "postDate": "2020-06-01T00:47:18.813Z",
          "content": "<p>There are complex models like TabNet, specially design for tabular data and i wonder if, it's possible to make and ensembl with?</p>",
          "rawMarkdown": "There are complex models like TabNet, specially design for tabular data and i wonder if, it's possible to make and ensembl with?",
          "votes": 5
        },
        {
          "id": 869657,
          "postDate": "2020-06-01T07:26:22.610Z",
          "content": "<p>You can also see solution for Petfinder competition . It had multi-modal input .</p>",
          "rawMarkdown": "You can also see solution for Petfinder competition . It had multi-modal input .",
          "votes": 3
        }
      ]
    },
    {
      "id": 937269,
      "postDate": "2020-07-20T22:04:39.113Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  Is it possible to use pseudo labels from a csv file (for pseudo labelling the test data) without recreating the tfrecord ? For instance let's say we have a dictionnary such as the keys are the image_id and the value are the pseudo labelling. Is it possible to  link these pseudo labels to an item in the tfrecord with tf.dataset and map function ? I did not succeed to do that</p>",
      "rawMarkdown": "@cdeotte  Is it possible to use pseudo labels from a csv file (for pseudo labelling the test data) without recreating the tfrecord ? For instance let's say we have a dictionnary such as the keys are the image_id and the value are the pseudo labelling. Is it possible to  link these pseudo labels to an item in the tfrecord with tf.dataset and map function ? I did not succeed to do that",
      "votes": 1,
      "replies": [
        {
          "id": 937272,
          "postDate": "2020-07-20T22:11:24.627Z",
          "content": "<p>Yes, i think there is. I have not worked out the details. I think you first read the test TFRecords in order (shuffle = False) and extract the test file names in order. Then you reorder the rows of a Pandas dataframe with the order of those names. Next you convert any meta features from the Pandas dataframe into a <code>tf.data.Dataset</code>. Lastly you concatenate that <code>tf.data.Dataset</code> with the original <code>tf.data.Dataset</code> with (shuffle = False), then lastly, you can do <code>ds = ds.shuffle(2048)</code> on the new merged dataset and shuffle it all up. I think that could work but i haven't tried yet.</p>\n\n<p>If this worked it would allow for engineering features easily and adding them to an existing <code>tf.data.Dataset</code>.</p>",
          "rawMarkdown": "Yes, i think there is. I have not worked out the details. I think you first read the test TFRecords in order (shuffle = False) and extract the test file names in order. Then you reorder the rows of a Pandas dataframe with the order of those names. Next you convert any meta features from the Pandas dataframe into a `tf.data.Dataset`. Lastly you concatenate that `tf.data.Dataset` with the original `tf.data.Dataset` with (shuffle = False), then lastly, you can do `ds = ds.shuffle(2048)` on the new merged dataset and shuffle it all up. I think that could work but i haven't tried yet.\n\nIf this worked it would allow for engineering features easily and adding them to an existing `tf.data.Dataset`.",
          "votes": 2
        }
      ]
    },
    {
      "id": 885370,
      "postDate": "2020-06-14T06:21:13.767Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> at the rescue as always! Thanks for sharing the data. One thing, I want to know is are the datasets you shared have been converted to grayscale or any other format for better pattern catching? And can you please refer to a notebook on the methodology used for creating these datasets? If there aren't any would you mind creating a notebook yourself to demonstrate how we can create a similar kind of data?? Thanks a lot.</p>",
      "rawMarkdown": "@cdeotte at the rescue as always! Thanks for sharing the data. One thing, I want to know is are the datasets you shared have been converted to grayscale or any other format for better pattern catching? And can you please refer to a notebook on the methodology used for creating these datasets? If there aren't any would you mind creating a notebook yourself to demonstrate how we can create a similar kind of data?? Thanks a lot.",
      "votes": 1,
      "replies": [
        {
          "id": 885381,
          "postDate": "2020-06-14T06:42:21.130Z",
          "content": "<p>I post some notes on how to use my Kaggle datasets (that include both images and meta data) with Keras models <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579#872250\">here</a>. My datasets are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\">here</a>.</p>",
          "rawMarkdown": "I post some notes on how to use my Kaggle datasets (that include both images and meta data) with Keras models [here][1]. My datasets are [here][2] and [here][3].\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579#872250\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\n[3]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245"
        }
      ]
    },
    {
      "id": 882059,
      "postDate": "2020-06-11T14:56:01.197Z",
      "content": "<p>Glad that you shared datasets too. Saved me a ton of time.</p>",
      "rawMarkdown": "Glad that you shared datasets too. Saved me a ton of time.",
      "votes": 1
    },
    {
      "id": 876292,
      "postDate": "2020-06-06T15:50:17.373Z",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> That's informative! Will implement and use it.</p>",
      "rawMarkdown": "@cdeotte That's informative! Will implement and use it.",
      "votes": 1,
      "replies": [
        {
          "id": 881975,
          "postDate": "2020-06-11T14:08:59.043Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 875421,
      "postDate": "2020-06-05T19:04:48.067Z",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!",
      "votes": 1
    },
    {
      "id": 874461,
      "postDate": "2020-06-05T02:26:30.590Z",
      "content": "<p>Great!</p>",
      "rawMarkdown": "Great!",
      "votes": 1
    },
    {
      "id": 874237,
      "postDate": "2020-06-04T18:21:52.757Z",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!",
      "votes": 1
    },
    {
      "id": 874149,
      "postDate": "2020-06-04T17:00:57.857Z",
      "content": "<p>Really a nice one!!!</p>",
      "rawMarkdown": "Really a nice one!!!",
      "votes": 1
    },
    {
      "id": 873373,
      "postDate": "2020-06-04T04:57:15.743Z",
      "content": "<p>Cool</p>",
      "rawMarkdown": "Cool",
      "votes": 1
    },
    {
      "id": 872409,
      "postDate": "2020-06-03T07:02:56.087Z",
      "content": "<p>Good Work</p>",
      "rawMarkdown": "Good Work",
      "votes": 1
    },
    {
      "id": 869646,
      "postDate": "2020-06-01T07:14:15.317Z",
      "content": "<p>Hi Chris , </p>\n\n<p>I tried to build a pipeline here using both image and tabular data together .  I pass the tabular data as a dense layer and concat with the image features . Kind of how we add  MLP to the Transformer models .</p>\n\n<p>I am working on fixing the metrics though .Currently its just accuracy . </p>\n\n<p><a href=\"https://www.kaggle.com/phoenix9032/siim-melanomas-pytorch-simple-multi-input-method\">https://www.kaggle.com/phoenix9032/siim-melanomas-pytorch-simple-multi-input-method</a></p>\n\n<p>Please see if this is what you meant . </p>",
      "rawMarkdown": "Hi Chris , \n\nI tried to build a pipeline here using both image and tabular data together .  I pass the tabular data as a dense layer and concat with the image features . Kind of how we add  MLP to the Transformer models .\n\n I am working on fixing the metrics though .Currently its just accuracy . \n\nhttps://www.kaggle.com/phoenix9032/siim-melanomas-pytorch-simple-multi-input-method\n\nPlease see if this is what you meant . ",
      "votes": 1,
      "replies": [
        {
          "id": 870596,
          "postDate": "2020-06-01T18:57:01.263Z",
          "content": "<p>That's nice Nirjhar. That's one way to do it. This is new to me too, so I need to research all the options myself. </p>\n\n<p>I remember from Dog Comp, that you could add info in the dense layer like you're doing, but you can also add info earlier in the conv layers like this. In the picture below imagine that <code>class</code> is the tabular data.</p>\n\n<p>If your model has the tabular data sooner (than the final dense layer), it can use that information when it is doing convolutions. In dog comp, i found option 1 to be better than option 2.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F55ad0b22bfa47b03e9bcb57ec5c206e1%2Fdisc3b.jpg?generation=1591037586805801&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "That's nice Nirjhar. That's one way to do it. This is new to me too, so I need to research all the options myself. \n\nI remember from Dog Comp, that you could add info in the dense layer like you're doing, but you can also add info earlier in the conv layers like this. In the picture below imagine that `class` is the tabular data.\n\nIf your model has the tabular data sooner (than the final dense layer), it can use that information when it is doing convolutions. In dog comp, i found option 1 to be better than option 2.\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F55ad0b22bfa47b03e9bcb57ec5c206e1%2Fdisc3b.jpg?generation=1591037586805801&amp;alt=media)\n",
          "votes": 2
        },
        {
          "id": 870606,
          "postDate": "2020-06-01T19:04:18.083Z",
          "content": "<p>That's a really nice idea . But might need to play with shape .  Probably need to repeat the array. Features are (bsxrowxcol) and conv layer (bs*ch*h*w) . May be feature need to be changed to (bsx1xrowxcol) . Will try this out .</p>",
          "rawMarkdown": "That's a really nice idea . But might need to play with shape .  Probably need to repeat the array. Features are (bsxrowxcol) and conv layer (bs*ch*h*w) . May be feature need to be changed to (bsx1xrowxcol) . Will try this out ."
        },
        {
          "id": 870641,
          "postDate": "2020-06-01T19:21:57.763Z",
          "content": "<p>Yes, you need to duplicate and/or reshape stuff so concatenation works out. Keep in mind there are two ways to add info</p>\n\n<h3>Add info to height and width of image</h3>\n\n<p>You can add a blue bar to the right side of all images for male and a red bar to the right side of all images for female. (And green bar for nan).</p>\n\n<h3>Add info to channel dimension of image</h3>\n\n<p>Images are 3 channels of red, green, blue. You can add a fourth channel that is all 1's for female and all 0's for male.</p>\n\n<p>(Note: the option 1 in the figure above shows adding information as a new channel.)</p>",
          "rawMarkdown": "Yes, you need to duplicate and/or reshape stuff so concatenation works out. Keep in mind there are two ways to add info\n\n### Add info to height and width of image\nYou can add a blue bar to the right side of all images for male and a red bar to the right side of all images for female. (And green bar for nan).\n\n### Add info to channel dimension of image\nImages are 3 channels of red, green, blue. You can add a fourth channel that is all 1's for female and all 0's for male.\n\n(Note: the option 1 in the figure above shows adding information as a new channel.)",
          "votes": 2
        },
        {
          "id": 870689,
          "postDate": "2020-06-01T20:30:00.443Z",
          "content": "<p>That is a smart idea. But how about features with bigger cardinality?</p>",
          "rawMarkdown": "That is a smart idea. But how about features with bigger cardinality?",
          "votes": 2
        },
        {
          "id": 870698,
          "postDate": "2020-06-01T20:37:50.793Z",
          "content": "<p>It works the same way. Let's say you have a feature like <code>age</code>. Then you can normalize it by subtracting mean and dividing standard deviation. Then it goes between -1 and 1.</p>\n\n<p>Next either\n* Increase height and/or width of image with a new long bar of color down the right side of the image where you set the color as follows: <code>red = age</code>, <code>green = age</code>, <code>blue = age</code> \n* Or, Add a new channel to image which has dimensions <code>height x width x 1</code> and set all values equal to <code>age</code></p>",
          "rawMarkdown": "It works the same way. Let's say you have a feature like `age`. Then you can normalize it by subtracting mean and dividing standard deviation. Then it goes between -1 and 1.\n\nNext either\n* Increase height and/or width of image with a new long bar of color down the right side of the image where you set the color as follows: `red = age`, `green = age`, `blue = age ` \n* Or, Add a new channel to image which has dimensions `height x width x 1` and set all values equal to `age`",
          "votes": 6
        },
        {
          "id": 870728,
          "postDate": "2020-06-01T21:06:59.770Z",
          "content": "<p>And add a new layer for each feature. That is definitely worth trying. Thank you for the bright idea, once again.</p>",
          "rawMarkdown": "And add a new layer for each feature. That is definitely worth trying. Thank you for the bright idea, once again.",
          "votes": 1
        },
        {
          "id": 884485,
          "postDate": "2020-06-13T12:10:58.443Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 886230,
      "postDate": "2020-06-14T20:13:53.350Z",
      "content": "<p>I made available <a href=\"https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512\">this public notebook</a> demonstrating how to implement 5-fold cross validation with both image and tabular data in Tensor Flow. See <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158395\">this discussion topic</a> for more details.</p>",
      "rawMarkdown": "I made available [this public notebook](https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512) demonstrating how to implement 5-fold cross validation with both image and tabular data in Tensor Flow. See [this discussion topic](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158395) for more details.",
      "votes": 2
    },
    {
      "id": 873235,
      "postDate": "2020-06-04T00:14:55.260Z",
      "content": "<p>weighting the table data with 0.15 instead of 0.1 helped in my case</p>",
      "rawMarkdown": "weighting the table data with 0.15 instead of 0.1 helped in my case",
      "votes": 2,
      "replies": [
        {
          "id": 873380,
          "postDate": "2020-06-04T05:03:40.303Z",
          "content": "<p>I tried giving .15 to table and .85 to image submission but it reduced score . Best score came for .1 to table and .9 to image data .</p>",
          "rawMarkdown": "I tried giving .15 to table and .85 to image submission but it reduced score . Best score came for .1 to table and .9 to image data .",
          "votes": 2
        },
        {
          "id": 873589,
          "postDate": "2020-06-04T09:34:04.960Z",
          "content": "<p>interesting - I would say it is still worth trying, as depending on the model it might help...</p>",
          "rawMarkdown": "interesting - I would say it is still worth trying, as depending on the model it might help...",
          "votes": 2
        },
        {
          "id": 873595,
          "postDate": "2020-06-04T09:37:12.077Z",
          "content": "<p>but you are right Roman .. for high image solution , tweaking it is giving better results .. so worth trying tweaking it ..... for high scoring image solution .. less weightage to table is giving boost .. e.g. current highest public lb ..0.95 and 0.05 tabular will give some boost . </p>",
          "rawMarkdown": "but you are right Roman .. for high image solution , tweaking it is giving better results .. so worth trying tweaking it ..... for high scoring image solution .. less weightage to table is giving boost .. e.g. current highest public lb ..0.95 and 0.05 tabular will give some boost . ",
          "votes": 2
        },
        {
          "id": 874014,
          "postDate": "2020-06-04T15:25:26.903Z",
          "content": "<p>got another 0.002 boost with 25% weight on tabular data - a bit strange hmm -higher weighting did not help...</p>",
          "rawMarkdown": "got another 0.002 boost with 25% weight on tabular data - a bit strange hmm -higher weighting did not help...",
          "votes": 2
        },
        {
          "id": 874020,
          "postDate": "2020-06-04T15:29:20.533Z",
          "content": "<p>Yes Roman\n.I used the files in Chris s kernel..so .91 image and .7 for tabular...I think you may be ensembling some other files..</p>",
          "rawMarkdown": "Yes Roman\n.I used the files in Chris s kernel..so .91 image and .7 for tabular...I think you may be ensembling some other files.."
        },
        {
          "id": 874080,
          "postDate": "2020-06-04T16:19:40.683Z",
          "content": "<p>sorry for not being exact. I use my best training result csv to combine it with the tabular data CSV (not sure anymore where I got this from... prob chris)</p>",
          "rawMarkdown": "sorry for not being exact. I use my best training result csv to combine it with the tabular data CSV (not sure anymore where I got this from... prob chris)",
          "votes": 1
        }
      ]
    },
    {
      "id": 882904,
      "postDate": "2020-06-12T08:18:43.870Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2198511%2Fc1bacd87cbe2a4cf1816bc6ad07e2e45%2FScreenshot%20from%202020-06-11%2020-42-27.png?generation=1591949431030175&amp;alt=media\" alt=\"\"></p>\n\n<p>How to prepare the model based on the tfrec data along with the metadata ? I tried above model it is not working.Thanks for creating the tfrec files</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2198511%2Fc1bacd87cbe2a4cf1816bc6ad07e2e45%2FScreenshot%20from%202020-06-11%2020-42-27.png?generation=1591949431030175&amp;alt=media)\n\nHow to prepare the model based on the tfrec data along with the metadata ? I tried above model it is not working.Thanks for creating the tfrec files",
      "replies": [
        {
          "id": 882927,
          "postDate": "2020-06-12T08:32:43.247Z",
          "content": "<p>In what way does it not work?</p>",
          "rawMarkdown": "In what way does it not work?"
        },
        {
          "id": 883001,
          "postDate": "2020-06-12T09:41:24.943Z",
          "content": "<p>I suggest you make your notebook public and then folks can help you a lot easier. \nAlso you get a bunch of additional input which is even more valuable.</p>\n\n<p>PS: I haven't done it myself, hence my learning curve was very slow :)</p>",
          "rawMarkdown": "I suggest you make your notebook public and then folks can help you a lot easier. \nAlso you get a bunch of additional input which is even more valuable.\n\nPS: I haven't done it myself, hence my learning curve was very slow :)"
        },
        {
          "id": 883816,
          "postDate": "2020-06-13T00:56:22.517Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2198511%2F30273c6d44d764ea7d828839dedfc2e2%2FScreenshot%20from%202020-06-13%2006-23-26.png?generation=1592009679271068&amp;alt=media\" alt=\"\">\nIt was the model I'm trying to create . It is throwing this error.\nI don't know how to tackle it.\n` in create_model()\n      4     inp2=Input(shape=(3))\n      5     X=base_model(inp1,training=False)\n----&gt; 6     X=Dense(3,input_shape=inp2)(X)\n      7     X=GlobalAveragePooling2D()(X)\n      8     X=Dense(1024,activation='relu')(X)</p>\n\n<p>'\n'TypeError: Cannot iterate over a tensor with unknown first dimension.'</p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2198511%2F30273c6d44d764ea7d828839dedfc2e2%2FScreenshot%20from%202020-06-13%2006-23-26.png?generation=1592009679271068&amp;alt=media)\nIt was the model I'm trying to create . It is throwing this error.\nI don't know how to tackle it.\n`"
        },
        {
          "id": 883882,
          "postDate": "2020-06-13T03:19:27.250Z",
          "content": "<p>After your GlobalAveragePooling2D you can concatenate the other input</p>\n\n<pre><code>from tf.keras.layers import Concatenate\n\nX = base_model(inp1, training=False)\nX = GlobalAveragePooling2D()(X)\nZ = Dense(16, activation='relu')(inp2)\nX = Concatenate()([X,Z]))\nX = Dense(1024, activation='relu')(X)\n</code></pre>",
          "rawMarkdown": "After your GlobalAveragePooling2D you can concatenate the other input\n\n    from tf.keras.layers import Concatenate\n\n    X = base_model(inp1, training=False)\n    X = GlobalAveragePooling2D()(X)\n    Z = Dense(16, activation='relu')(inp2)\n    X = Concatenate()([X,Z]))\n    X = Dense(1024, activation='relu')(X)"
        },
        {
          "id": 884097,
          "postDate": "2020-06-13T07:32:03.377Z",
          "content": "<p>Thank you very much 👍 </p>",
          "rawMarkdown": "Thank you very much 👍 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 874613,
      "postDate": "2020-06-05T06:20:20.120Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!"
    },
    {
      "id": 874566,
      "postDate": "2020-06-05T05:35:10.947Z",
      "content": "<p>fantastic!</p>",
      "rawMarkdown": "fantastic!"
    },
    {
      "id": 869659,
      "postDate": "2020-06-01T07:26:59.360Z",
      "content": "<p>What about Multitask ? Predicting the Diagnosis as well as Target ?</p>",
      "rawMarkdown": "What about Multitask ? Predicting the Diagnosis as well as Target ?",
      "replies": [
        {
          "id": 870597,
          "postDate": "2020-06-01T18:59:03.270Z",
          "content": "<p>Interesting idea. Yes, you could add an auxiliary head that predicts the tabular features from the images at the same time that it predicts malignancy probability. Auxiliary heads are a way of making a model aware of additional data.</p>",
          "rawMarkdown": "Interesting idea. Yes, you could add an auxiliary head that predicts the tabular features from the images at the same time that it predicts malignancy probability. Auxiliary heads are a way of making a model aware of additional data.",
          "votes": 7
        },
        {
          "id": 872031,
          "postDate": "2020-06-02T20:40:42.120Z",
          "content": "<p>Pls Chris can you explain this in more detail?</p>",
          "rawMarkdown": "Pls Chris can you explain this in more detail?"
        }
      ]
    },
    {
      "id": 890679,
      "postDate": "2020-06-17T16:28:20.780Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 874695,
      "postDate": "2020-06-05T08:28:42.813Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 870126,
      "postDate": "2020-06-01T14:05:34.067Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 870628,
          "postDate": "2020-06-01T19:12:21.423Z",
          "content": "<p>Great idea. That is a 4th option not in my list above. You are stacking two models (whereas my 3rd option is ensembling two models). You are building one model to predict melanoma from images. Then you are using those predictions plus additional tabular data to build another model to predict your final predictions. </p>",
          "rawMarkdown": "Great idea. That is a 4th option not in my list above. You are stacking two models (whereas my 3rd option is ensembling two models). You are building one model to predict melanoma from images. Then you are using those predictions plus additional tabular data to build another model to predict your final predictions. ",
          "votes": 1
        },
        {
          "id": 870918,
          "postDate": "2020-06-02T03:12:03.727Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 871729,
          "postDate": "2020-06-02T15:44:43.020Z",
          "content": "<p><a href=\"/justayaan\">@justayaan</a>\nThe secret to Kaggle image competitions is first you must set up a reliable CV that estimates LB score well. Next make sure to write your code as fast as possible because you need to do lots of experiments. Also consider using small image sizes while you experiment (and then switch to large after you've found a good setup). Experiment with the following:</p>\n\n<ul>\n<li>preprocess</li>\n<li>data augmentation</li>\n<li>model architecture, loss, etc</li>\n<li>training schedule, optimizer, etc</li>\n<li>postprocess</li>\n</ul>\n\n<p>Keep any idea that increases your CV and discard others</p>",
          "rawMarkdown": "@justayaan\nThe secret to Kaggle image competitions is first you must set up a reliable CV that estimates LB score well. Next make sure to write your code as fast as possible because you need to do lots of experiments. Also consider using small image sizes while you experiment (and then switch to large after you've found a good setup). Experiment with the following:\n\n* preprocess\n* data augmentation\n* model architecture, loss, etc\n* training schedule, optimizer, etc\n* postprocess\n\nKeep any idea that increases your CV and discard others",
          "votes": 8
        },
        {
          "id": 872501,
          "postDate": "2020-06-03T08:55:27.540Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 873060,
          "postDate": "2020-06-03T18:30:08.523Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 877511,
      "postDate": "2020-06-07T17:05:48.457Z",
      "content": "<p>Interesting suggestion . Thanks!</p>",
      "rawMarkdown": "Interesting suggestion . Thanks!"
    }
  ],
  "comments": [
    {
      "id": 869369,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-06-01T01:04:59.523000",
      "content": "<p>the past isic leaderboard website has a list of methods that uses image or image+table:\n<a href=\"https://challenge2019.isic-archive.com/leaderboard.html\">https://challenge2019.isic-archive.com/leaderboard.html</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F62811b28d3e8c6c43b8564c09e5f6ff8%2FSelection_041.png?generation=1590973485561188&amp;alt=media\" alt=\"\"></p>\n\n<p>i am implementing transformer based methods</p>",
      "votes": 17,
      "replies": [
        {
          "id": 869393,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-06-01T01:50:48.083000",
          "content": "<p>Transformers aren't more adequate for NLP problems?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 869504,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-06-01T04:03:43.197000",
          "content": "<p>Transformers are not for NLP only ..And if you are someone like <a href=\"/limerobot\">@limerobot</a> , you can use them whenever you have the opportunity : See <a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891\">here</a> and  <a href=\"https://www.kaggle.com/c/champs-scalar-coupling/discussion/106572\">here</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 869507,
          "author_name": "Leon",
          "author_url": "",
          "post_date": "2020-06-01T04:05:39.730000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 869512,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-01T04:11:41.177000",
          "content": "<p>i am trying to modify this : \"DETR: End-to-End Object Detection with Transformers\"</p>\n\n<p>i am more concerned with \"multiple images per patient\". so we have a sequence of images rather then one images. i think this will have significant impact on the score.</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 869655,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-06-01T07:25:21.437000",
          "content": "<p>@Heng that is very interesting thing for me also . But  I am not sure how I can a relate a patient with a skin image taken from neck at the age of 35 as Melanoma = 0 , for the same patient if i take another skin image at another bodypart at the age of 40 . and find melanoma =1 . May be this timeseries relationship is little complex , need to figure out .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 869670,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-01T07:33:01.447000",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> </p>\n\n<p>it also have to do with the way data is collected. i think the data has been processed (e.g. images filtered by doctors for recording purpose? ). in particular, from pure data analysis point of view:</p>\n\n<ul>\n<li><p>for a given patient and location (e.g.  upper extremity, torso, etc), there at at most few +ve images. the rest are negative. i.e. the presence some high score images will lower posterior scores of others.</p></li>\n<li><p>for a given patient, there are at most one or two locations with +ve  images. similarly, the presence some high score at some location will lower posterior scores of others.</p></li>\n</ul>\n\n<p>if a model learn such biases, LB score can be improved a lot. but it doesn't make more accurate melanoma detection in real life application. </p>\n\n<p>this is a problem and dilemma  in data science competition. should you exploit data biases for better scores or build a useful model (possibly with lower score)?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 870749,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-06-01T22:04:00.370000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> your actual submission was using DETR? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 873482,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-06-04T07:28:52.593000",
          "content": "<p><a href=\"/hiramcho\">@hiramcho</a> \n\"<a href=\"/hengck23\">@hengck23</a> your actual submission was using DETR?\"\nnot yet</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 875367,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-06-05T17:51:47.783000",
      "content": "<p>this is the prior probability from <a href=\"https://en.wikipedia.org/wiki/Melanoma\">https://en.wikipedia.org/wiki/Melanoma</a>\n(instead from the train.csv file)</p>\n\n<p><img src=\"https://upload.wikimedia.org/wikipedia/commons/thumb/3/37/Diagram_showing_where_melanoma_is_most_likely_to_develop_CRUK_383.svg/800px-Diagram_showing_where_melanoma_is_most_likely_to_develop_CRUK_383.svg.png\" alt=\"\"></p>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 871604,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-06-02T13:57:07.667000",
      "content": "<p><a href=\"https://ai.googleblog.com/2019/09/using-deep-learning-to-inform.html\">https://ai.googleblog.com/2019/09/using-deep-learning-to-inform.html</a>\n“A Deep Learning System for Differential Diagnosis of Skin Diseases” \n<img src=\"https://1.bp.blogspot.com/-b3HBRhjUZs0/XXoeACfbsmI/AAAAAAAAEoc/f4aL1TH8J2g6Ix5Tlfj-rYXemzRgPmVXwCEwYBhgL/s640/image1.png\" alt=\"\"></p>\n\n<p><img src=\"https://1.bp.blogspot.com/-wMnHrj5PPys/XXoeAKVc2xI/AAAAAAAAEoI/64V1-8jUT6s-BuBRjc-EmjIpWvR7W2yZQCLcBGAsYHQ/s640/image3.png\" alt=\"\"></p>",
      "votes": 11,
      "replies": []
    },
    {
      "id": 952589,
      "author_name": "Marcelo Kittlein",
      "author_url": "",
      "post_date": "2020-07-31T04:18:06.783000",
      "content": "<p>Took tabular data from <a href=\"https://www.kaggle.com/awsaf49/xgboost-tabular-data-ml-cv-85-lb-787\">this</a> kernel and made images from the scaled (0..1) outer product of the features.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1828058%2Fc55ad02be5a4ca92ea13029b6955ef78%2FTabImage.png?generation=1596168989462514&amp;alt=media\" alt=\"\">\nAdded to a simple 4 channel cnn model and boosted LB score... wonder how is best to mingle this in EfficientNet?</p>",
      "votes": 10,
      "replies": [
        {
          "id": 952596,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-07-31T04:30:48.403000",
          "content": "<p>this is awesome , any resource to teach how to do that ??</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 952626,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-31T05:14:21.500000",
          "content": "<p>Wow, this is awesome <a href=\"/kittlein\">@kittlein</a> , great idea. Each of these images is 8x8. Perhaps you can tile these into 256x256 and then concatenate that with images of size 256x256x3. Then if you use <code>tf.keras.layers.Conv2D(3,8)</code> you could incorporate all the info into each pixel of the original images. And feed that into pretrained EfficientNet.</p>\n\n<p><a href=\"/phoenix9032\">@phoenix9032</a> for each row of train data, let's say you have 8 meta features. First scale them to between 0 and 1. Next create a vector <code>x</code> of size <code>8x1</code>. Lastly if you if multiply <code>x</code> by <code>x.T</code> the transpose of <code>x</code> the result is an 8x8 matrix. These 8x8 matrices are displayed above.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 871389,
      "author_name": "Gajendra Saraswat",
      "author_url": "",
      "post_date": "2020-06-02T10:15:07.797000",
      "content": "<p>Lol, what did you do! :)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2070357%2Fa0d0dc316725c83ae5fd05dd16f7fa13%2FScreenshot%20from%202020-06-02%2015-42-31.png?generation=1591092841763019&amp;alt=media\" alt=\"\"></p>\n\n<p>Some sort of a benchmark. :/</p>",
      "votes": 7,
      "replies": [
        {
          "id": 871720,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-02T15:40:17.493000",
          "content": "<p>haha, maybe its just a coincidence that everyone got the same LB of 0.914 as my notebook 😄  </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 869748,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2020-06-01T08:59:11.683000",
      "content": "<p>A picture from the topic right next to yours\n<img src=\"https://ars.els-cdn.com/content/image/1-s2.0-S2215016120300832-gr2.jpg\" alt=\"\"></p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 962846,
      "author_name": "Nitesh Chaudhry",
      "author_url": "",
      "post_date": "2020-08-08T13:35:45.760000",
      "content": "<p>I just recently shared the notebook for using dual-input CNNs for this problem. Please check it out !</p>\n<p><a href=\"https://www.kaggle.com/niteshx2/pipeline-dual-input-cnn?scriptVersionId=40368875\" target=\"_blank\">https://www.kaggle.com/niteshx2/pipeline-dual-input-cnn?scriptVersionId=40368875</a></p>\n<p><img src=\"https://ars.els-cdn.com/content/image/1-s2.0-S2215016120300832-gr2.jpg\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 962871,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-08-08T14:02:25.727000",
          "content": "<p>Great job Nitesh, thanks for posting. Your notebook is a helpful example on how to combine meta into CNN.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 875901,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-06-06T08:49:06.017000",
      "content": "<p><a href=\"https://www.youtube.com/watch?v=ysBaZO8YmX8\">https://www.youtube.com/watch?v=ysBaZO8YmX8</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F08f597f592772bbbab337cb3f8d54804%2FSelection_098.png?generation=1591433343089407&amp;alt=media\" alt=\"\"></p>",
      "votes": 6,
      "replies": [
        {
          "id": 876325,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-06T16:16:06.140000",
          "content": "<p>Interesting. Thanks for sharing. The slides for that video are posted <a href=\"https://www.slideshare.net/SebastienFischman/tab-netpresentation/SebastienFischman/tab-netpresentation\">here</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 876337,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-06-06T16:20:16.533000",
          "content": "<p><a href=\"/optimo\">@optimo</a>  tagging you once more , while I am trying the new version in other comp.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 877060,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-06-07T09:33:13.493000",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> I don't know anything about the other comp but there is a new kernel using multi target regression if you want <a href=\"https://www.kaggle.com/tanulsingh077/achieving-sota-results-with-tabnet/\">https://www.kaggle.com/tanulsingh077/achieving-sota-results-with-tabnet/</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 877359,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2020-06-07T14:30:07.383000",
          "content": "<p>Thank you <a href=\"/optimo\">@optimo</a> I believe that Tabnet has got a lot potential. <a href=\"/cdeotte\">@cdeotte</a> I am. Planning on doing some research with Tabnet for this competition as well, after tweet competition ends, only if you don't do ut earlier😜</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 877362,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-06-07T14:33:54.610000",
          "content": "<p>Yes  <a href=\"/optimo\">@optimo</a>  I started using multitarget in the TReNDs neuroimaging comp and its indeed giving me better score , I have not made my kernel public as <a href=\"/tanulsingh077\">@tanulsingh077</a>  has already done it well ..</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 871997,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-06-02T19:48:05.167000",
      "content": "<p>Meta_data definitely helps.  My TPU Quota is almost exhausted (I use it for Jigsaw competition too ) but the first fold with Image+meta gives me 0.909 on LB </p>\n\n<p>No Augmentation, No external data for now</p>",
      "votes": 6,
      "replies": [
        {
          "id": 872051,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-02T21:15:11.847000",
          "content": "<p>Nice</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 872061,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-06-02T21:31:18.083000",
          "content": "<p>Thanks !</p>\n\n<p>LB  0.912 for the second fold...TPU Quota completely exhausted :(</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 872770,
          "author_name": "Camilo Cuevas",
          "author_url": "",
          "post_date": "2020-06-03T13:41:04.663000",
          "content": "<p>Cool!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 872780,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-03T13:52:34.903000",
          "content": "<p>have you checked without meta? i get 0.919 without meta and 512 input</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 873075,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-06-03T18:52:58.887000",
          "content": "<p>I got 0.905 for all 5 folds  and 768 as input size</p>\n\n<p>But it was first or second day of the competition . I changed the loss function in the last experiment.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 874240,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-04T18:24:02.120000",
          "content": "<p>thx</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 871264,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-02T08:18:18.173000",
      "content": "<p>UPDATE: I created TFRecords that contain both the meta data and image data <a href=\"https://www.kaggle.com/cdeotte/melanoma-256x256\">here</a> for 256x256 and <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">here</a> for 512x512. All of the tabular data is inside as fields in addition to the images.</p>\n\n<pre><code>  feature = {\n      'image': _bytes_feature,\n      'image_name': _bytes_feature,\n      'patient_id': _int64_feature,\n      'sex': _int64_feature,\n      'age_approx': _int64_feature,\n      'anatom_site_general_challenge': _int64_feature,\n      'diagnosis': _int64_feature,\n      'target': _int64_feature\n  }\n</code></pre>",
      "votes": 6,
      "replies": [
        {
          "id": 871273,
          "author_name": "Bruce Young",
          "author_url": "",
          "post_date": "2020-06-02T08:31:20.187000",
          "content": "<p>Fantastic thanks :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 871725,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-02T15:41:44.403000",
          "content": "<p>I remember 512x512 being a nice size in Flower Comp, so I thought it would be good here too. I'm also uploading 768x768 and 384x384 now. And i plan to upload 192x192 and 128x128 TFRecords</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 872947,
          "author_name": "Mladen Radošević",
          "author_url": "",
          "post_date": "2020-06-03T16:12:26.733000",
          "content": "<p>Very interesting ... any chance of releasnig the code for creating these datasets?</p>\n\n<p>I wonder ... why stop only with those meta data.\nIf you check you can see there is a difference between distribution of other image variables for 'target' 1 and 0 ... (i.e. 'standard deviation' by channel/without channels) and some other parameters. \nSome of them could be very useful.</p>\n\n<p>And one other question ... why did you put 'diagnose' in features? \nThere is no 'diagnose' for test images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 872960,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-03T16:38:46.403000",
          "content": "<p>You can use <code>diagnosis</code> as an auxiliary target. For example you train your model to predict both <code>target</code> and <code>diagnosis</code>. Then when you infer test, you just ignore your model's <code>diagnosis</code> prediction and keep your model's <code>target</code> prediction and submit that to this comp.</p>\n\n<p>That is called auxiliary training with an auxiliary head and can help your model make better predictions for target.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 873072,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-03T18:49:39.200000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> how would you treat \"unknown\" diagnosis?</p>\n\n<p>Theoretically, it can be any of the remaining classes or something else. Here are my thoughts:\n0. unknown as a separate class  (probably sub-optimal)\n1. filtering - discard loss from unknown\n2. pseudo labels / knowledge distillation - train on examples from known diagnosis and predict unknown</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 874113,
          "author_name": "darthgera123",
          "author_url": "",
          "post_date": "2020-06-04T16:35:40.977000",
          "content": "<p>how good is progressive resizing in this competition?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 882971,
          "author_name": "ucanmakeit",
          "author_url": "",
          "post_date": "2020-06-12T09:12:37.190000",
          "content": "<p>any more suggestions? thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 948466,
      "author_name": "FC",
      "author_url": "",
      "post_date": "2020-07-28T01:32:14.817000",
      "content": "<p>I created this one which could potentially help \n<a href=\"https://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta\">https://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 948469,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-28T01:39:02.527000",
          "content": "<p>Great job. Stacking image OOF with meta features is a good idea.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 872894,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-06-03T15:36:44.640000",
      "content": "<p>got me from 0.927 to 0.932 thx</p>\n\n<p>TTA works well either - but gain variates a lot - I got +0.009 once, but usually more around +0.004</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 870795,
      "author_name": "somuSan",
      "author_url": "",
      "post_date": "2020-06-01T23:29:07.717000",
      "content": "<p>We can apply different image model(like-efficientnet B0, B6,resnet, etc.) and apply stacking using the predicted values.\nIn one of the Kaggle meet up, the winner team from the previously held competition on \"skin cancer diagnosis\" discussed their approach, where they used stacking like this too.\n<a href=\"https://www.youtube.com/watch?v=meg8I1GdtUg\">https://www.youtube.com/watch?v=meg8I1GdtUg</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 869701,
      "author_name": "Trigram",
      "author_url": "",
      "post_date": "2020-06-01T08:15:05.153000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I saw Dieter's kernel: \"Extract Image Features from Pretrained NN\" and I will keep people posted on my implementation of \"feature extraction + XGB\" combo here.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 869715,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-06-01T08:29:15.573000",
          "content": "<p>Isnt that the same <a href=\"/tunguz\">@tunguz</a>  did here ?</p>\n\n<p><a href=\"https://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling\">https://www.kaggle.com/tunguz/melanoma-classification-eda-and-modeling</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 870618,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-01T19:09:31.897000",
          "content": "<p>Great idea Trigram. That is my option 2 above \"Build a Tabular data model and find a way to extract image embeddings and input into a Tabular data model\". </p>\n\n<p><a href=\"/phoenix9032\">@phoenix9032</a> What Bojan did is a very simple form of it. He has converted each image into a 1024 dimension vector. But his vector are just the pixel values from a 32x32 resized image. </p>\n\n<p>You can get a better 1024 dimension vector (embedding) by inputting the original full size image into a pretrained CNN and taking the output from the last convolutional layer and applying a <code>GlobalAveragePooling2D</code>. That will also be a 1024 dimension vector but that vector will contain much more information about the image.</p>\n\n<p>Each element of this better 1024 dimension vector will represent whether a certain feature is present in the image or not because these elements are results of convolutions. (As opposed to the elements in the basic 1024 dimension vector which just represent whether a certain pixel is present or not).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 870667,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-06-01T19:52:59.537000",
          "content": "<p>Got it . </p>\n\n<ol>\n<li>effnet --&gt; feature --&gt; GAP --&gt; (Embedding + Tabular) --&gt; LGBM </li>\n<li>effnet --&gt; feature --&gt; GAP --&gt; (Embedding + Tabular)  --&gt; NN </li>\n</ol>\n\n<p>I am trying 2 , Trigram trying 1.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 871685,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2020-06-02T15:12:18.707000",
      "content": "<p>Of course metadata is important. Simply looking at the age/malignancy correlation (remember: <strong>correlation does not mean causation</strong>) and using naive Bayes you can have ~70% accuracy.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 871732,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-02T15:47:51.093000",
          "content": "<p>Great point. Yes when i look at the images myself they all look the same, but when age is high, then i would more likely guess malignant. (So if that's what i do, our models would probably benefit from knowing age too).</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 870175,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-01T14:41:32.193000",
      "content": "<p>UPDATE: I did a quick experiment <a href=\"https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\">here</a>. Even though the tabular data predictions have a low LB of 0.700 <a href=\"https://www.kaggle.com/titericz/simple-baseline\">here</a> and the image only predictions have a high LB 0.910 <a href=\"https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\">here</a>, if you add the tabular data predictions to the image data predictions, the LB increases. Thus the tabular data contains new information.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 871358,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2020-06-02T09:41:24.490000",
          "content": "<p>just confirming that my score also increased 0.926 =&gt; 0.930. A lot more juice to squeeze here....</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 871952,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-02T18:48:14.597000",
          "content": "<p>Yup, even for me..score increased from 0.919 to 0.923</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 869340,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2020-06-01T00:32:18.960000",
      "content": "<p>what do you think about this approach <a href=\"https://www.kaggle.com/c/avito-demand-prediction/discussion/59880\">\"Dance with Ensembl\"</a> ? I'm going spend the rest of the competition building something based on that. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 869346,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-01T00:39:25.017000",
          "content": "<p>Wow, awesome model and it won 1st place in another competition. I think that has great potential here.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 869352,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2020-06-01T00:47:18.813000",
          "content": "<p>There are complex models like TabNet, specially design for tabular data and i wonder if, it's possible to make and ensembl with?</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 869657,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-06-01T07:26:22.610000",
          "content": "<p>You can also see solution for Petfinder competition . It had multi-modal input .</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 937269,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2020-07-20T22:04:39.113000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  Is it possible to use pseudo labels from a csv file (for pseudo labelling the test data) without recreating the tfrecord ? For instance let's say we have a dictionnary such as the keys are the image_id and the value are the pseudo labelling. Is it possible to  link these pseudo labels to an item in the tfrecord with tf.dataset and map function ? I did not succeed to do that</p>",
      "votes": 1,
      "replies": [
        {
          "id": 937272,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-07-20T22:11:24.627000",
          "content": "<p>Yes, i think there is. I have not worked out the details. I think you first read the test TFRecords in order (shuffle = False) and extract the test file names in order. Then you reorder the rows of a Pandas dataframe with the order of those names. Next you convert any meta features from the Pandas dataframe into a <code>tf.data.Dataset</code>. Lastly you concatenate that <code>tf.data.Dataset</code> with the original <code>tf.data.Dataset</code> with (shuffle = False), then lastly, you can do <code>ds = ds.shuffle(2048)</code> on the new merged dataset and shuffle it all up. I think that could work but i haven't tried yet.</p>\n\n<p>If this worked it would allow for engineering features easily and adding them to an existing <code>tf.data.Dataset</code>.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 885370,
      "author_name": "Divyansh Agrawal",
      "author_url": "",
      "post_date": "2020-06-14T06:21:13.767000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> at the rescue as always! Thanks for sharing the data. One thing, I want to know is are the datasets you shared have been converted to grayscale or any other format for better pattern catching? And can you please refer to a notebook on the methodology used for creating these datasets? If there aren't any would you mind creating a notebook yourself to demonstrate how we can create a similar kind of data?? Thanks a lot.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 885381,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-14T06:42:21.130000",
          "content": "<p>I post some notes on how to use my Kaggle datasets (that include both images and meta data) with Keras models <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579#872250\">here</a>. My datasets are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">here</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156245\">here</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 882059,
      "author_name": "Nitesh Chaudhry",
      "author_url": "",
      "post_date": "2020-06-11T14:56:01.197000",
      "content": "<p>Glad that you shared datasets too. Saved me a ton of time.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 876292,
      "author_name": "Dhruv Aggarwal",
      "author_url": "",
      "post_date": "2020-06-06T15:50:17.373000",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> That's informative! Will implement and use it.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 881975,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-11T14:08:59.043000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 875421,
      "author_name": "arya2001",
      "author_url": "",
      "post_date": "2020-06-05T19:04:48.067000",
      "content": "<p>Great Work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 874461,
      "author_name": "Kazutaka Ota",
      "author_url": "",
      "post_date": "2020-06-05T02:26:30.590000",
      "content": "<p>Great!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 874237,
      "author_name": "Sibil Sarjam Soren",
      "author_url": "",
      "post_date": "2020-06-04T18:21:52.757000",
      "content": "<p>Great Work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 874149,
      "author_name": "Debanjan Ghosh",
      "author_url": "",
      "post_date": "2020-06-04T17:00:57.857000",
      "content": "<p>Really a nice one!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 873373,
      "author_name": "Zakka Izzatur Rahman Noor",
      "author_url": "",
      "post_date": "2020-06-04T04:57:15.743000",
      "content": "<p>Cool</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 872409,
      "author_name": "Jayachandiran Udayakumar",
      "author_url": "",
      "post_date": "2020-06-03T07:02:56.087000",
      "content": "<p>Good Work</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 869646,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2020-06-01T07:14:15.317000",
      "content": "<p>Hi Chris , </p>\n\n<p>I tried to build a pipeline here using both image and tabular data together .  I pass the tabular data as a dense layer and concat with the image features . Kind of how we add  MLP to the Transformer models .</p>\n\n<p>I am working on fixing the metrics though .Currently its just accuracy . </p>\n\n<p><a href=\"https://www.kaggle.com/phoenix9032/siim-melanomas-pytorch-simple-multi-input-method\">https://www.kaggle.com/phoenix9032/siim-melanomas-pytorch-simple-multi-input-method</a></p>\n\n<p>Please see if this is what you meant . </p>",
      "votes": 1,
      "replies": [
        {
          "id": 870596,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-01T18:57:01.263000",
          "content": "<p>That's nice Nirjhar. That's one way to do it. This is new to me too, so I need to research all the options myself. </p>\n\n<p>I remember from Dog Comp, that you could add info in the dense layer like you're doing, but you can also add info earlier in the conv layers like this. In the picture below imagine that <code>class</code> is the tabular data.</p>\n\n<p>If your model has the tabular data sooner (than the final dense layer), it can use that information when it is doing convolutions. In dog comp, i found option 1 to be better than option 2.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F55ad0b22bfa47b03e9bcb57ec5c206e1%2Fdisc3b.jpg?generation=1591037586805801&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 870606,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-06-01T19:04:18.083000",
          "content": "<p>That's a really nice idea . But might need to play with shape .  Probably need to repeat the array. Features are (bsxrowxcol) and conv layer (bs*ch*h*w) . May be feature need to be changed to (bsx1xrowxcol) . Will try this out .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 870641,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-01T19:21:57.763000",
          "content": "<p>Yes, you need to duplicate and/or reshape stuff so concatenation works out. Keep in mind there are two ways to add info</p>\n\n<h3>Add info to height and width of image</h3>\n\n<p>You can add a blue bar to the right side of all images for male and a red bar to the right side of all images for female. (And green bar for nan).</p>\n\n<h3>Add info to channel dimension of image</h3>\n\n<p>Images are 3 channels of red, green, blue. You can add a fourth channel that is all 1's for female and all 0's for male.</p>\n\n<p>(Note: the option 1 in the figure above shows adding information as a new channel.)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 870689,
          "author_name": "Roman",
          "author_url": "",
          "post_date": "2020-06-01T20:30:00.443000",
          "content": "<p>That is a smart idea. But how about features with bigger cardinality?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 870698,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-01T20:37:50.793000",
          "content": "<p>It works the same way. Let's say you have a feature like <code>age</code>. Then you can normalize it by subtracting mean and dividing standard deviation. Then it goes between -1 and 1.</p>\n\n<p>Next either\n* Increase height and/or width of image with a new long bar of color down the right side of the image where you set the color as follows: <code>red = age</code>, <code>green = age</code>, <code>blue = age</code> \n* Or, Add a new channel to image which has dimensions <code>height x width x 1</code> and set all values equal to <code>age</code></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 870728,
          "author_name": "Roman",
          "author_url": "",
          "post_date": "2020-06-01T21:06:59.770000",
          "content": "<p>And add a new layer for each feature. That is definitely worth trying. Thank you for the bright idea, once again.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 884485,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-13T12:10:58.443000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 886230,
      "author_name": "Alexey Pronin",
      "author_url": "",
      "post_date": "2020-06-14T20:13:53.350000",
      "content": "<p>I made available <a href=\"https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512\">this public notebook</a> demonstrating how to implement 5-fold cross validation with both image and tabular data in Tensor Flow. See <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158395\">this discussion topic</a> for more details.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 873235,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-04T00:14:55.260000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 873380,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T05:03:40.303000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 873589,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T09:34:04.960000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 873595,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T09:37:12.077000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 874014,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T15:25:26.903000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 874020,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T15:29:20.533000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 874080,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-04T16:19:40.683000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 882904,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-12T08:18:43.870000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 882927,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-12T08:32:43.247000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 883001,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-12T09:41:24.943000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 883816,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-13T00:56:22.517000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 883882,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-13T03:19:27.250000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 884097,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-13T07:32:03.377000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 874613,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-05T06:20:20.120000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 874566,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-05T05:35:10.947000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 869659,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-01T07:26:59.360000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 870597,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-01T18:59:03.270000",
          "content": "",
          "votes": 7,
          "replies": []
        },
        {
          "id": 872031,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-02T20:40:42.120000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 890679,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-17T16:28:20.780000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 874695,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-05T08:28:42.813000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 870126,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-01T14:05:34.067000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 870628,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-01T19:12:21.423000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 870918,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-02T03:12:03.727000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 871729,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-02T15:44:43.020000",
          "content": "",
          "votes": 8,
          "replies": []
        },
        {
          "id": 872501,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-03T08:55:27.540000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 873060,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-03T18:30:08.523000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 877511,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-07T17:05:48.457000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "869339": "This competition provides both **image data** and **tabular data**. This is exciting! We have seen from this notebook [here][1] that using just **images** can score LB 0.910. And we have seen from this notebook [here][2] that using only the provided **tabular data** can score LB 0.700. Let's discuss ideas how build a model that uses both **images** and **tabular data**.\n\nThree ideas come to mind.\n* Build a CNN image model and find a way to input the tabular data into the CNN image model\n* Build a Tabular data model and find a way to extract image embeddings and input into a Tabular data model\n* Build 2 separate models and ensemble\n\nWhat other ideas do people have? Has anyone tried one of the above 3 with success?\n\n========================\n\nUPDATE: I did a quick experiment [here][3]. Even though the tabular data predictions have a low LB of 0.700 and the image only predictions have a high LB 0.910, if you add the tabular data predictions to the image data predictions, the LB increases. Thus the tabular data contains new information.\n\n========================\n\nUPDATE: I created TFRecords that contain both the images and tabular data. So you can easily build TensorFlow models that utilize both. Size 768x768 data is [here][6], 512x512 dataset is [here][4], 384x384 is [here][7], and 256x256 dataset is [here][5]. A description of the TFRecords fields is [here][8].\n\n[1]: https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\n[2]: https://www.kaggle.com/titericz/simple-baseline\n[3]: https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\n[4]: https://www.kaggle.com/cdeotte/melanoma-512x512\n[5]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[6]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[7]: https://www.kaggle.com/cdeotte/melanoma-384x384\n[8]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579",
    "869369": "the past isic leaderboard website has a list of methods that uses image or image+table:\nhttps://challenge2019.isic-archive.com/leaderboard.html\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F62811b28d3e8c6c43b8564c09e5f6ff8%2FSelection_041.png?generation=1590973485561188&amp;alt=media)\n\n\ni am implementing transformer based methods",
    "875367": "this is the prior probability from https://en.wikipedia.org/wiki/Melanoma\n(instead from the train.csv file)\n\n![](https://upload.wikimedia.org/wikipedia/commons/thumb/3/37/Diagram_showing_where_melanoma_is_most_likely_to_develop_CRUK_383.svg/800px-Diagram_showing_where_melanoma_is_most_likely_to_develop_CRUK_383.svg.png)",
    "871604": "https://ai.googleblog.com/2019/09/using-deep-learning-to-inform.html\n“A Deep Learning System for Differential Diagnosis of Skin Diseases” \n![](https://1.bp.blogspot.com/-b3HBRhjUZs0/XXoeACfbsmI/AAAAAAAAEoc/f4aL1TH8J2g6Ix5Tlfj-rYXemzRgPmVXwCEwYBhgL/s640/image1.png)\n\n![](https://1.bp.blogspot.com/-wMnHrj5PPys/XXoeAKVc2xI/AAAAAAAAEoI/64V1-8jUT6s-BuBRjc-EmjIpWvR7W2yZQCLcBGAsYHQ/s640/image3.png)",
    "952589": "Took tabular data from [this](https://www.kaggle.com/awsaf49/xgboost-tabular-data-ml-cv-85-lb-787) kernel and made images from the scaled (0..1) outer product of the features.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1828058%2Fc55ad02be5a4ca92ea13029b6955ef78%2FTabImage.png?generation=1596168989462514&amp;alt=media)\nAdded to a simple 4 channel cnn model and boosted LB score... wonder how is best to mingle this in EfficientNet?",
    "871389": "Lol, what did you do! :)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2070357%2Fa0d0dc316725c83ae5fd05dd16f7fa13%2FScreenshot%20from%202020-06-02%2015-42-31.png?generation=1591092841763019&amp;alt=media)\n\n\nSome sort of a benchmark. :/",
    "869748": "A picture from the topic right next to yours\n![](https://ars.els-cdn.com/content/image/1-s2.0-S2215016120300832-gr2.jpg)",
    "962846": "I just recently shared the notebook for using dual-input CNNs for this problem. Please check it out !\n\nhttps://www.kaggle.com/niteshx2/pipeline-dual-input-cnn?scriptVersionId=40368875\n\n![](https://ars.els-cdn.com/content/image/1-s2.0-S2215016120300832-gr2.jpg)",
    "875901": "https://www.youtube.com/watch?v=ysBaZO8YmX8\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F08f597f592772bbbab337cb3f8d54804%2FSelection_098.png?generation=1591433343089407&amp;alt=media)\n",
    "871997": "Meta_data definitely helps.  My TPU Quota is almost exhausted (I use it for Jigsaw competition too ) but the first fold with Image+meta gives me 0.909 on LB \n\nNo Augmentation, No external data for now\n",
    "871264": "UPDATE: I created TFRecords that contain both the meta data and image data [here][1] for 256x256 and [here][2] for 512x512. All of the tabular data is inside as fields in addition to the images.\n\n      feature = {\n          'image': _bytes_feature,\n          'image_name': _bytes_feature,\n          'patient_id': _int64_feature,\n          'sex': _int64_feature,\n          'age_approx': _int64_feature,\n          'anatom_site_general_challenge': _int64_feature,\n          'diagnosis': _int64_feature,\n          'target': _int64_feature\n      }\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-256x256\n[2]: https://www.kaggle.com/cdeotte/melanoma-512x512",
    "948466": "I created this one which could potentially help \nhttps://www.kaggle.com/chen2222/lgb-mlp-ridge-stacking-effnet-oof-with-meta",
    "872894": "got me from 0.927 to 0.932 thx\n\nTTA works well either - but gain variates a lot - I got +0.009 once, but usually more around +0.004",
    "870795": "We can apply different image model(like-efficientnet B0, B6,resnet, etc.) and apply stacking using the predicted values.\nIn one of the Kaggle meet up, the winner team from the previously held competition on \"skin cancer diagnosis\" discussed their approach, where they used stacking like this too.\nhttps://www.youtube.com/watch?v=meg8I1GdtUg",
    "869701": "@cdeotte I saw Dieter's kernel: \"Extract Image Features from Pretrained NN\" and I will keep people posted on my implementation of \"feature extraction + XGB\" combo here.",
    "871685": "Of course metadata is important. Simply looking at the age/malignancy correlation (remember: **correlation does not mean causation**) and using naive Bayes you can have ~70% accuracy.",
    "870175": "UPDATE: I did a quick experiment [here][3]. Even though the tabular data predictions have a low LB of 0.700 [here][2] and the image only predictions have a high LB 0.910 [here][1], if you add the tabular data predictions to the image data predictions, the LB increases. Thus the tabular data contains new information.\n\n[1]: https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\n[2]: https://www.kaggle.com/titericz/simple-baseline\n[3]: https://www.kaggle.com/cdeotte/image-and-tabular-data-0-915\n",
    "869340": "what do you think about this approach [\"Dance with Ensembl\"](https://www.kaggle.com/c/avito-demand-prediction/discussion/59880) ? I'm going spend the rest of the competition building something based on that. ",
    "937269": "@cdeotte  Is it possible to use pseudo labels from a csv file (for pseudo labelling the test data) without recreating the tfrecord ? For instance let's say we have a dictionnary such as the keys are the image_id and the value are the pseudo labelling. Is it possible to  link these pseudo labels to an item in the tfrecord with tf.dataset and map function ? I did not succeed to do that",
    "885370": "@cdeotte at the rescue as always! Thanks for sharing the data. One thing, I want to know is are the datasets you shared have been converted to grayscale or any other format for better pattern catching? And can you please refer to a notebook on the methodology used for creating these datasets? If there aren't any would you mind creating a notebook yourself to demonstrate how we can create a similar kind of data?? Thanks a lot.",
    "882059": "Glad that you shared datasets too. Saved me a ton of time.",
    "876292": "@cdeotte That's informative! Will implement and use it.",
    "875421": "Great Work!",
    "874461": "Great!",
    "874237": "Great Work!",
    "874149": "Really a nice one!!!",
    "873373": "Cool",
    "872409": "Good Work",
    "869646": "Hi Chris , \n\nI tried to build a pipeline here using both image and tabular data together .  I pass the tabular data as a dense layer and concat with the image features . Kind of how we add  MLP to the Transformer models .\n\n I am working on fixing the metrics though .Currently its just accuracy . \n\nhttps://www.kaggle.com/phoenix9032/siim-melanomas-pytorch-simple-multi-input-method\n\nPlease see if this is what you meant . ",
    "886230": "I made available [this public notebook](https://www.kaggle.com/graf10a/effnb0-tabular-features-tf-cv5-512x512) demonstrating how to implement 5-fold cross validation with both image and tabular data in Tensor Flow. See [this discussion topic](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158395) for more details.",
    "873235": "weighting the table data with 0.15 instead of 0.1 helped in my case",
    "882904": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2198511%2Fc1bacd87cbe2a4cf1816bc6ad07e2e45%2FScreenshot%20from%202020-06-11%2020-42-27.png?generation=1591949431030175&amp;alt=media)\n\nHow to prepare the model based on the tfrec data along with the metadata ? I tried above model it is not working.Thanks for creating the tfrec files",
    "874613": "Great work!",
    "874566": "fantastic!",
    "869659": "What about Multitask ? Predicting the Diagnosis as well as Target ?",
    "890679": "",
    "874695": "",
    "870126": "",
    "877511": "Interesting suggestion . Thanks!"
  }
}