{
  "id": 169728,
  "title": "metadata don't want to cooperate",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/169728",
  "author_name": "",
  "post_date": "2020-07-25T01:28:38.767368200Z",
  "votes": 12,
  "comment_count": 42,
  "views": 0,
  "content": "<p>In metadata we have age, sex and area.\nWe can train model on metadata only and we can see some score larger than 0.5, so yes, data is valid and it is useful.</p>\n\n<p>However, in public kernel metadata is concatenated with CNN output, then there is one last Linear/Dense layer and we have single output.</p>\n\n<p>It means that two models (metadata and CNN) are just blended. There is no combination of \"this output from metadata and this output from CNN  work together\". We have just some number of outputs from CNN and some number of outputs from metadata we multiply each with weight, sum them up and that's all.</p>\n\n<p>I would assume that metadata will create outputs like \"this specific area of woman aged about 60\" and then CNN will create output to combine with this metadata output. But adding additional Linear layer never helps, I did lots of tests today.</p>\n\n<p>Do you have any other experiences or do you know why it works like that?</p>",
  "messages": [
    {
      "id": "944265",
      "postDate": "07/25/2020 01:28:38",
      "content": "<p>In metadata we have age, sex and area.\nWe can train model on metadata only and we can see some score larger than 0.5, so yes, data is valid and it is useful.</p>\n\n<p>However, in public kernel metadata is concatenated with CNN output, then there is one last Linear/Dense layer and we have single output.</p>\n\n<p>It means that two models (metadata and CNN) are just blended. There is no combination of \"this output from metadata and this output from CNN  work together\". We have just some number of outputs from CNN and some number of outputs from metadata we multiply each with weight, sum them up and that's all.</p>\n\n<p>I would assume that metadata will create outputs like \"this specific area of woman aged about 60\" and then CNN will create output to combine with this metadata output. But adding additional Linear layer never helps, I did lots of tests today.</p>\n\n<p>Do you have any other experiences or do you know why it works like that?</p>",
      "rawMarkdown": "In metadata we have age, sex and area.\nWe can train model on metadata only and we can see some score larger than 0.5, so yes, data is valid and it is useful.\n\nHowever, in public kernel metadata is concatenated with CNN output, then there is one last Linear/Dense layer and we have single output.\n\nIt means that two models (metadata and CNN) are just blended. There is no combination of \"this output from metadata and this output from CNN  work together\". We have just some number of outputs from CNN and some number of outputs from metadata we multiply each with weight, sum them up and that's all.\n\nI would assume that metadata will create outputs like \"this specific area of woman aged about 60\" and then CNN will create output to combine with this metadata output. But adding additional Linear layer never helps, I did lots of tests today.\n\nDo you have any other experiences or do you know why it works like that?",
      "votes": null
    },
    {
      "id": "944339",
      "postDate": "07/25/2020 03:10:37",
      "content": "<p>I agree that if we only concatenate the last layers then the combination isn't much better than simple ensemble. If you want our CNN to extract different features, I think we need to add the meta features to the bottom layers. I think we need to try option 1 below. Imagine that the word \"class\" says \"meta data\"\n<img src=\"http://playagricola.com/Kaggle/disc3b.jpg\" alt=\"image\"></p>",
      "rawMarkdown": "I agree that if we only concatenate the last layers then the combination isn't much better than simple ensemble. If you want our CNN to extract different features, I think we need to add the meta features to the bottom layers. I think we need to try option 1 below. Imagine that the word \"class\" says \"meta data\"\n![image](http://playagricola.com/Kaggle/disc3b.jpg)",
      "votes": null
    },
    {
      "id": "944346",
      "postDate": "07/25/2020 03:25:58",
      "content": "<p>What I was using until now is this:</p>\n\n<p>For image data:</p>\n\n<pre><code>arch._fc = nn.Linear(in_features = 1280, out_features = 500, bias = True)\n</code></pre>\n\n<p>For meta features:</p>\n\n<pre><code>self.meta = nn.Sequential(nn.Linear(n_meta_features, 500), \n                             nn.BatchNorm1d(500), \n                             nn.ReLU(), \n                             nn.Dropout(p = 0.25), \n                             nn.Linear(500, 250), \n                             nn.BatchNorm1d(250), \n                             nn.ReLU(), \n                             nn.Dropout(p = 0.2))\n</code></pre>\n\n<p>Now, Classification:</p>\n\n<pre><code>self.classifier = nn.Linear(500 + 250, 2)\n</code></pre>\n\n<p>It may not be much different from simple blending, but of course it should work much better than that. </p>\n\n<p>*I could not test it on heavy models, due to some difficulty implementing TPU. :/</p>",
      "rawMarkdown": "What I was using until now is this:\n\nFor image data:\n\n    arch._fc = nn.Linear(in_features = 1280, out_features = 500, bias = True)\n\nFor meta features:\n    \n    self.meta = nn.Sequential(nn.Linear(n_meta_features, 500), \n                                 nn.BatchNorm1d(500), \n                                 nn.ReLU(), \n                                 nn.Dropout(p = 0.25), \n                                 nn.Linear(500, 250), \n                                 nn.BatchNorm1d(250), \n                                 nn.ReLU(), \n                                 nn.Dropout(p = 0.2))\n\nNow, Classification:\n\n    self.classifier = nn.Linear(500 + 250, 2)\n\nIt may not be much different from simple blending, but of course it should work much better than that. \n\n*I could not test it on heavy models, due to some difficulty implementing TPU. :/",
      "votes": null
    },
    {
      "id": "944365",
      "postDate": "07/25/2020 03:56:44",
      "content": "<p>Since this is the first serious attempt I have made for multiple inputs I am commenting as a complete novice. <br>\nI saw improvement over blending when I did a b0 and densenet169 model that concatenated those two models.  Likewise, I saw improvement over blending when I added a vgg16 model and concat the three.  I did not look super hard but concatenate was the only tf method I saw for mutliple inputs.   </p>\n\n<p>That experience suggests to me that it works over blending.  The issue is that it's very expensive.  I could not run large image sizes on my local Ubuntu machine even at batch size of 2.  It's also likely that my \"blend\" was not optimized.  SO - who knows??  But I believe it is at least an optimized version of blending? </p>\n\n<p>I added a fourth input for the meta data.  I had 16 features.  I started with simple single layer and did not really see an improvement.  But if I was using tf on only the meta data I sure would not settle for a simple single layer.  So playing around I got a tf metadata model.  It was not as good on LB as an XGboost or catboost at 0.73LB.  Adding that model in did show a small improvement, but the improvement was smaller than the std dev among the 5 folds so I don't believe it significant.</p>\n\n<p>My conclusions - garbage in - garbage out.   My 16 features are better than flipping a coin, but nice models may be smart enough to ignore this useless data.  Guess after the competition ends will see if others found big value in the meta data, but I suspect it did not work for you because it is none value added information that models are smart enough to ignore.</p>",
      "rawMarkdown": "Since this is the first serious attempt I have made for multiple inputs I am commenting as a complete novice.  \nI saw improvement over blending when I did a b0 and densenet169 model that concatenated those two models.  Likewise, I saw improvement over blending when I added a vgg16 model and concat the three.  I did not look super hard but concatenate was the only tf method I saw for mutliple inputs.   \n\nThat experience suggests to me that it works over blending.  The issue is that it's very expensive.  I could not run large image sizes on my local Ubuntu machine even at batch size of 2.  It's also likely that my \"blend\" was not optimized.  SO - who knows??  But I believe it is at least an optimized version of blending? \n\nI added a fourth input for the meta data.  I had 16 features.  I started with simple single layer and did not really see an improvement.  But if I was using tf on only the meta data I sure would not settle for a simple single layer.  So playing around I got a tf metadata model.  It was not as good on LB as an XGboost or catboost at 0.73LB.  Adding that model in did show a small improvement, but the improvement was smaller than the std dev among the 5 folds so I don't believe it significant.\n\nMy conclusions - garbage in - garbage out.   My 16 features are better than flipping a coin, but nice models may be smart enough to ignore this useless data.  Guess after the competition ends will see if others found big value in the meta data, but I suspect it did not work for you because it is none value added information that models are smart enough to ignore.",
      "votes": null
    },
    {
      "id": "944366",
      "postDate": "07/25/2020 03:58:33",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1062256%2F93d8fca78e4274be85161545b71802cf%2F4inputmodel.png?generation=1595649510770080&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1062256%2F93d8fca78e4274be85161545b71802cf%2F4inputmodel.png?generation=1595649510770080&amp;alt=media)",
      "votes": null
    },
    {
      "id": "944511",
      "postDate": "07/25/2020 06:58:40",
      "content": "<p>If it is not a secret. Did such approach help you improve CV score?</p>",
      "rawMarkdown": "If it is not a secret. Did such approach help you improve CV score?",
      "votes": null
    },
    {
      "id": "944549",
      "postDate": "07/25/2020 07:24:24",
      "content": "<p>hello <a href=\"/sarques\">@sarques</a> \nthis examply shows exactly what i mean:\n- you create 500 classes from efficientnet (I don't understand why people replace _fc instead just constructing efficientnet with this number of classes - what's the difference?)\n- you create 9 -&gt; 500 -&gt; 250 net for meta features\n- you concat 500 with 250 then you do 750 -&gt; 1 net</p>",
      "rawMarkdown": "hello @sarques \nthis examply shows exactly what i mean:\n- you create 500 classes from efficientnet (I don't understand why people replace _fc instead just constructing efficientnet with this number of classes - what's the difference?)\n- you create 9 -&gt; 500 -&gt; 250 net for meta features\n- you concat 500 with 250 then you do 750 -&gt; 1 net",
      "votes": null
    },
    {
      "id": "944554",
      "postDate": "07/25/2020 07:26:39",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I was thinking about this approach too, but it means we need to add numeric values with some magic to the image itself, because we use transfer learning so the input must be an image.</p>",
      "rawMarkdown": "cdeotte I was thinking about this approach too, but it means we need to add numeric values with some magic to the image itself, because we use transfer learning so the input must be an image.",
      "votes": null
    },
    {
      "id": "944559",
      "postDate": "07/25/2020 07:28:48",
      "content": "<p>IMHO they are not useless because training with meta only works, but combining them with CNN suggest they are, I am not really uderstanding what test you did, I tried like 20 different combinations yesterday</p>",
      "rawMarkdown": "IMHO they are not useless because training with meta only works, but combining them with CNN suggest they are, I am not really uderstanding what test you did, I tried like 20 different combinations yesterday",
      "votes": null
    },
    {
      "id": "944569",
      "postDate": "07/25/2020 07:33:27",
      "content": "<p>There is example TensorFlow code <a href=\"https://www.kaggle.com/cdeotte/dog-breed-acgan-lb-52\">here</a>. I used this in Dog Comp. Option 1 did better than Option 2. By inputting the meta data earlier in the CNN, it allows the CNN to adjust its feature generation as opposed to only adjusting classification.</p>",
      "rawMarkdown": "There is example TensorFlow code [here][1]. I used this in Dog Comp. Option 1 did better than Option 2. By inputting the meta data earlier in the CNN, it allows the CNN to adjust its feature generation as opposed to only adjusting classification.\n\n[1]: https://www.kaggle.com/cdeotte/dog-breed-acgan-lb-52",
      "votes": null
    },
    {
      "id": "944612",
      "postDate": "07/25/2020 07:56:44",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> by quickly looking - you are not using transfer learning in this example but your own net, right?</p>",
      "rawMarkdown": "cdeotte by quickly looking - you are not using transfer learning in this example but your own net, right?",
      "votes": null
    },
    {
      "id": "944621",
      "postDate": "07/25/2020 08:04:03",
      "content": "<p>Correct. To use transfer learning, you should apply convolution to convert <code>256x256x4</code> into <code>256x256x3</code></p>\n\n<pre><code>    x = layers.Embedding(IN_SIZE, DIM*DIM, input_length=1)(label)\n    x = layers.Reshape((DIM,DIM,1))(x)\n    x = layers.concatenate([image,x])\n    # NOW WE ARE DIM x DIM x 4\n    x = layers.Conv2D(3)(x)\n    # NOW WE ARE DIM x DIM x 3\n    x = EfficientNetB0()(x)\n</code></pre>",
      "rawMarkdown": "Correct. To use transfer learning, you should apply convolution to convert `256x256x4` into `256x256x3`\n\n        x = layers.Embedding(IN_SIZE, DIM*DIM, input_length=1)(label)\n        x = layers.Reshape((DIM,DIM,1))(x)\n        x = layers.concatenate([image,x])\n        # NOW WE ARE DIM x DIM x 4\n        x = layers.Conv2D(3)(x)\n        # NOW WE ARE DIM x DIM x 3\n        x = EfficientNetB0()(x)",
      "votes": null
    },
    {
      "id": "944646",
      "postDate": "07/25/2020 08:20:24",
      "content": "<p>you are at 3 channels but your data is not real image anymore, right? so learned weights may be incorrect</p>",
      "rawMarkdown": "you are at 3 channels but your data is not real image anymore, right? so learned weights may be incorrect",
      "votes": null
    },
    {
      "id": "944675",
      "postDate": "07/25/2020 08:42:45",
      "content": "<p>True but this should work and transfer learning should still help. CNNs are intelligent and will learn how to use this new convolutional layer. Perhaps it will add red to all patients over 60 and remove red from all patients under 60, then the CNN can still use it's transfer learning.</p>",
      "rawMarkdown": "True but this should work and transfer learning should still help. CNNs are intelligent and will learn how to use this new convolutional layer. Perhaps it will add red to all patients over 60 and remove red from all patients under 60, then the CNN can still use it's transfer learning.",
      "votes": null
    },
    {
      "id": "946234",
      "postDate": "07/26/2020 13:00:41",
      "content": "<p>Tested another about 20-30 models, without any effect.\nLooks like this is idea for another competition, and for this one I just need to blend meta with CNN in one Linear layer.</p>",
      "rawMarkdown": "Tested another about 20-30 models, without any effect.\nLooks like this is idea for another competition, and for this one I just need to blend meta with CNN in one Linear layer.",
      "votes": null
    },
    {
      "id": "946436",
      "postDate": "07/26/2020 15:19:25",
      "content": "<blockquote>\n  <p>However, in public kernel metadata is concatenated with CNN output, then there is one last Linear/Dense layer and we have single output.</p>\n  \n  <p>It means that two models (metadata and CNN) are just blended. There is no combination of \"this output from metadata and this output from CNN work together\". We have just some number of outputs from CNN and some number of outputs from metadata we multiply each with weight, sum them up and that's all.</p>\n</blockquote>\n\n<p>This is not completely true. When you concatenate meta and image embeddings in the last layer of a CNN it will affect the model's back propagation learning and thus affect which convolution filters are created.</p>\n\n<p>For example, a model <strong>without</strong> meta features may learn to create convolution filters that differentiate skin color because male and female have different skin color. But a model <strong>with</strong> meta features already has access to gender meta feature. So that model may not create convolution filters that differentiate skin color (to differentiate male and female) because it already has access to the information.</p>\n\n<p>So, in the same way that dropout forces your model to look at new parts of the image. Adding meta features also forces your model to look at new parts of the image.</p>",
      "rawMarkdown": "&gt; However, in public kernel metadata is concatenated with CNN output, then there is one last Linear/Dense layer and we have single output.\n  \n&gt; It means that two models (metadata and CNN) are just blended. There is no combination of \"this output from metadata and this output from CNN work together\". We have just some number of outputs from CNN and some number of outputs from metadata we multiply each with weight, sum them up and that's all.\n\nThis is not completely true. When you concatenate meta and image embeddings in the last layer of a CNN it will affect the model's back propagation learning and thus affect which convolution filters are created.\n\nFor example, a model **without** meta features may learn to create convolution filters that differentiate skin color because male and female have different skin color. But a model **with** meta features already has access to gender meta feature. So that model may not create convolution filters that differentiate skin color (to differentiate male and female) because it already has access to the information.\n\nSo, in the same way that dropout forces your model to look at new parts of the image. Adding meta features also forces your model to look at new parts of the image.",
      "votes": null
    },
    {
      "id": "946603",
      "postDate": "07/26/2020 17:10:31",
      "content": "<p>I disagree but probably we think about same thing.</p>\n\n<p>When you have one Linear layer you only add probabilities, there is no way to communicate between CNN and meta.</p>\n\n<p>With one more Linear layer backpropagation could send back information to CNN to look at images more needed for specific meta.</p>\n\n<p>What I can agree is that when model will learn that larger age means more probability for 1 then CNN can spend less resources on learning images from people with larger age. But that's all.</p>",
      "rawMarkdown": "I disagree but probably we think about same thing.\n\nWhen you have one Linear layer you only add probabilities, there is no way to communicate between CNN and meta.\n\nWith one more Linear layer backpropagation could send back information to CNN to look at images more needed for specific meta.\n\nWhat I can agree is that when model will learn that larger age means more probability for 1 then CNN can spend less resources on learning images from people with larger age. But that's all.",
      "votes": null
    },
    {
      "id": "946630",
      "postDate": "07/26/2020 17:28:14",
      "content": "<p>Communication occurs. For example if you concatenate the true labels as in </p>\n\n<pre><code>x = EfficientNetB0(input)\nx = GlobalAveragePooling2D()(x)\nx = Concatenate()([x, TRUE_TARGETS])\nx = Dense(1, activation='sigmoid')(x)\n</code></pre>\n\n<p>Then the EfficientNetB0 backbone will not get any learning. Learning is driven by <code>error</code>. When you concatenate another source of features such as meta data, you are affecting the <code>error</code> and thus you are affecting the backpropagation of learning (which affects all proceeding input feature pipelines).</p>",
      "rawMarkdown": "Communication occurs. For example if you concatenate the true labels as in \n\n    x = EfficientNetB0(input)\n    x = GlobalAveragePooling2D()(x)\n    x = Concatenate()([x, TRUE_TARGETS])\n    x = Dense(1, activation='sigmoid')(x)\n\nThen the EfficientNetB0 backbone will not get any learning. Learning is driven by `error`. When you concatenate another source of features such as meta data, you are affecting the `error` and thus you are affecting the backpropagation of learning (which affects all proceeding input feature pipelines).",
      "votes": null
    },
    {
      "id": "947490",
      "postDate": "07/27/2020 09:52:25",
      "content": "<p><a href=\"/sarques\">@sarques</a> Why do you take your classifier to an output of 2 instead of 1? Are you using CrossEntropy instead of BinaryCrossEntropy?</p>",
      "rawMarkdown": "sarques Why do you take your classifier to an output of 2 instead of 1? Are you using CrossEntropy instead of BinaryCrossEntropy?",
      "votes": null
    },
    {
      "id": "947512",
      "postDate": "07/27/2020 10:08:05",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Interesting net you have there.  So you were using 3 pre-trained nets at the same time, and then added in the Meta data as well.  What kind of hardware are you using to handle that?  And you say that you found this way of doing things with concatenating, better than say just taking prediction outputs from the separate models and combining them?</p>",
      "rawMarkdown": "pcjimmmy Interesting net you have there.  So you were using 3 pre-trained nets at the same time, and then added in the Meta data as well.  What kind of hardware are you using to handle that?  And you say that you found this way of doing things with concatenating, better than say just taking prediction outputs from the separate models and combining them?",
      "votes": null
    },
    {
      "id": "947532",
      "postDate": "07/27/2020 10:19:49",
      "content": "<p>Yes, I am using CrossEntropy! :)</p>",
      "rawMarkdown": "Yes, I am using CrossEntropy! :)",
      "votes": null
    },
    {
      "id": "947640",
      "postDate": "07/27/2020 11:50:57",
      "content": "<p>Is there an advantage to using CrossEntropy for a binary problem vs BinaryCrossEntropy?  What's your motivation to use Cross-Entropy instead of Binary Cross Entropy?</p>",
      "rawMarkdown": "Is there an advantage to using CrossEntropy for a binary problem vs BinaryCrossEntropy?  What's your motivation to use Cross-Entropy instead of Binary Cross Entropy?",
      "votes": null
    },
    {
      "id": "947747",
      "postDate": "07/27/2020 13:11:07",
      "content": "<p>There is not much about this, it just happened that I started with CrossEntropy, but lately I was using Focal loss for CrossEntropy, also, I can not test my code because there is some problem in implementing TPU with PyTorch XLA. :/ As of now, I am using Chris's notebook of Tensorflow! :)</p>",
      "rawMarkdown": "There is not much about this, it just happened that I started with CrossEntropy, but lately I was using Focal loss for CrossEntropy, also, I can not test my code because there is some problem in implementing TPU with PyTorch XLA. :/ As of now, I am using Chris's notebook of Tensorflow! :)",
      "votes": null
    },
    {
      "id": "948470",
      "postDate": "07/28/2020 01:40:53",
      "content": "<p>4 machines all - Ubuntu 20.04 - Intel® Core™ i7-9700K CPU @ 3.60GHz  dual gpu's on all the machines - either 1070 and 1080.  Pretty sure it would not run on a kaggle kernel but ok on my setups.</p>\n\n<p>My single pre-trained blends were not weighted so I don't know if a little bit of work could have weighted the blends to match the triple model.  I was doing a different augmentation of the image for each of the three.  Gave up on the approach as it's expensive and biggest image size was not going to be very large.  The expensive part just really needed me to have more patience - but letting things run for two days before you get a hint if your on the right track is beyond me.  But not getting a decent image size seemed like the real road block.   </p>\n\n<p>Been playing with mixed precision on single models and might return.  But this discussion post makes me worry that concatenating this way will not perform better than separate and properly weighted blends.</p>",
      "rawMarkdown": "4 machines all - Ubuntu 20.04 - Intel® Core™ i7-9700K CPU @ 3.60GHz  dual gpu's on all the machines - either 1070 and 1080.  Pretty sure it would not run on a kaggle kernel but ok on my setups.\n\nMy single pre-trained blends were not weighted so I don't know if a little bit of work could have weighted the blends to match the triple model.  I was doing a different augmentation of the image for each of the three.  Gave up on the approach as it's expensive and biggest image size was not going to be very large.  The expensive part just really needed me to have more patience - but letting things run for two days before you get a hint if your on the right track is beyond me.  But not getting a decent image size seemed like the real road block.   \n\nBeen playing with mixed precision on single models and might return.  But this discussion post makes me worry that concatenating this way will not perform better than separate and properly weighted blends.",
      "votes": null
    },
    {
      "id": "948498",
      "postDate": "07/28/2020 02:42:34",
      "content": "<p>Mixed precision may not be stable on a 1080/1070.  Realize the 16bit performance of both of those cards is terrible.  Are you using Distributed Data Paralllel (DDP)?  That would seem like a good approach given you have 4 separate nodes..........the faster the connectivity between the machines the better though, as that could be a source of bottleneck.</p>",
      "rawMarkdown": "Mixed precision may not be stable on a 1080/1070.  Realize the 16bit performance of both of those cards is terrible.  Are you using Distributed Data Paralllel (DDP)?  That would seem like a good approach given you have 4 separate nodes..........the faster the connectivity between the machines the better though, as that could be a source of bottleneck.",
      "votes": null
    },
    {
      "id": "949211",
      "postDate": "07/28/2020 13:21:26",
      "content": "<p>I think some of the issues come from the train &amp; test datasets being slightly different. This makes it possible that a model would perform well on CV but not so good on test. Raddar did some experiments here:\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155885\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155885</a></p>\n\n<p>While writing this post, I started writing down reasons why this might be, for example, 62% of the positive cases are male which seemed way too high to me (gut feeling). But then I Googled \"melanoma male vs female\" and was surprised with what I saw.</p>\n\n<p>I think the trick is choosing which metadata features to use and which to discard.</p>",
      "rawMarkdown": "I think some of the issues come from the train &amp; test datasets being slightly different. This makes it possible that a model would perform well on CV but not so good on test. Raddar did some experiments here:\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155885\n\nWhile writing this post, I started writing down reasons why this might be, for example, 62% of the positive cases are male which seemed way too high to me (gut feeling). But then I Googled \"melanoma male vs female\" and was surprised with what I saw.\n\nI think the trick is choosing which metadata features to use and which to discard.",
      "votes": null
    },
    {
      "id": "949535",
      "postDate": "07/28/2020 17:37:17",
      "content": "<p><a href=\"/anjum48\">@anjum48</a> What features are people using besides Gender, Age and Site?  Those all seem valid..........not sure it would make sense to use anything else.</p>",
      "rawMarkdown": "anjum48 What features are people using besides Gender, Age and Site?  Those all seem valid..........not sure it would make sense to use anything else.",
      "votes": null
    },
    {
      "id": "949556",
      "postDate": "07/28/2020 17:56:48",
      "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> If you use count of patient id, you can push a metadata only model up to 0.74. This might make sense because if a patient has lots of benign moles that look similar, then the model can use this information to adjust its predictions (and vice-versa for malignant using the Ugly Duckling concept).</p>\n\n<p>Also, age is probably a risk factor, but the number of people above 70 is rare, so you could use frequency encoded features. This can get a meta only model above 0.77 (CV) when combined with the feature above.</p>\n\n<p>But if the test set is different from train then this approach can create issues in having a model that will generalise well.</p>",
      "rawMarkdown": "brianfeeny If you use count of patient id, you can push a metadata only model up to 0.74. This might make sense because if a patient has lots of benign moles that look similar, then the model can use this information to adjust its predictions (and vice-versa for malignant using the Ugly Duckling concept).\n\nAlso, age is probably a risk factor, but the number of people above 70 is rare, so you could use frequency encoded features. This can get a meta only model above 0.77 (CV) when combined with the feature above.\n\nBut if the test set is different from train then this approach can create issues in having a model that will generalise well.",
      "votes": null
    },
    {
      "id": "949574",
      "postDate": "07/28/2020 18:07:34",
      "content": "<p>Yeah, using things like patent id has to be well thought through.  My personal gut feeling is that is going to be overfitting and not a good feature.</p>",
      "rawMarkdown": "Yeah, using things like patent id has to be well thought through.  My personal gut feeling is that is going to be overfitting and not a good feature.",
      "votes": null
    },
    {
      "id": "949595",
      "postDate": "07/28/2020 18:27:12",
      "content": "<p>Using meta features is frustrating and risky because the data was not randomly split between <code>train.csv</code> and <code>test.csv</code>. There are many artificially created correlations (i.e. false patterns in train data).</p>",
      "rawMarkdown": "Using meta features is frustrating and risky because the data was not randomly split between `train.csv` and `test.csv`. There are many artificially created correlations (i.e. false patterns in train data).",
      "votes": null
    },
    {
      "id": "949911",
      "postDate": "07/29/2020 04:22:39",
      "content": "<p>Sorry I confused you - I have 4 machines but running different versions of the 4 input models separately.  Was thinking about running all 4 on the same script but decided to wait until a future competition - doing that is outside my skill set so it would be a full on learning experience but the tensorflow docs make it seem something learnable.</p>\n\n<p>Mixed precision on single model seems to be running ok - nice boost in the size of the batch that can be run - I also started doing a short warmup with batch size of 4 and than continuing with a new fit and much larger batch size.  That really allowed for a huge batch size increase.  Tensorflow apparently a huge RAM hog for the first epoch.   Between mixed and the warmup I am going to retry the 4 input model again to see how big I can get the image. </p>",
      "rawMarkdown": "Sorry I confused you - I have 4 machines but running different versions of the 4 input models separately.  Was thinking about running all 4 on the same script but decided to wait until a future competition - doing that is outside my skill set so it would be a full on learning experience but the tensorflow docs make it seem something learnable.\n\nMixed precision on single model seems to be running ok - nice boost in the size of the batch that can be run - I also started doing a short warmup with batch size of 4 and than continuing with a new fit and much larger batch size.  That really allowed for a huge batch size increase.  Tensorflow apparently a huge RAM hog for the first epoch.   Between mixed and the warmup I am going to retry the 4 input model again to see how big I can get the image.",
      "votes": null
    },
    {
      "id": "950142",
      "postDate": "07/29/2020 07:53:04",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> are you using NVIDIA's AMP for mixed-precision?  If so, what optimization level do you have it on?  I just don't understand how a 1070 and 1080 can do true 16 bit at any sort of appreciable speed.</p>\n\n<p>1070 is 5783 GFLOPS @ 32bit, 90 GFLOPS @ 16bit (yes 90, not a typo)\n1080 is 8228 GFLOPS @ 32bit, 128 GFLOPS @ 16bit</p>\n\n<p>I have no experience in mixed precision.  Soon I am swapping my 1080 Ti's out for 2080 Ti's, and those have TensorCores and can actually do halfway decent 16-bit precision.  I am looking forward to be abl able to take on larger models and/or speed up my training times.</p>",
      "rawMarkdown": "pcjimmmy are you using NVIDIA's AMP for mixed-precision?  If so, what optimization level do you have it on?  I just don't understand how a 1070 and 1080 can do true 16 bit at any sort of appreciable speed.\n\n1070 is 5783 GFLOPS @ 32bit, 90 GFLOPS @ 16bit (yes 90, not a typo)\n1080 is 8228 GFLOPS @ 32bit, 128 GFLOPS @ 16bit\n\nI have no experience in mixed precision.  Soon I am swapping my 1080 Ti's out for 2080 Ti's, and those have TensorCores and can actually do halfway decent 16-bit precision.  I am looking forward to be abl able to take on larger models and/or speed up my training times.",
      "votes": null
    },
    {
      "id": "951169",
      "postDate": "07/30/2020 00:39:40",
      "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> </p>\n\n<p>Running very standard Ubuntu 20.04 - my 1070's on one of the machines are more specifically as shown on settings info:\nGeForce GTX 1070/PCIe/SSE2 / GeForce GTX 1070/PCIe/SSE2</p>\n\n<p>With mixed precision I am not quite able to double the batch size and I am pretty sure I would have noticed a speed drop of the order your numbers showing and bailed out on that approach.  When one of the machines ends it's current script will do a simple timing evaluation with and with the tensorflow mixed precision.  Seems like I did the fastai equivalent of these several months ago and likewise don't recall a huge speed price for the almost doubling of the batch size.  </p>",
      "rawMarkdown": "brianfeeny \n\nRunning very standard Ubuntu 20.04 - my 1070's on one of the machines are more specifically as shown on settings info:\nGeForce GTX 1070/PCIe/SSE2 / GeForce GTX 1070/PCIe/SSE2\n\nWith mixed precision I am not quite able to double the batch size and I am pretty sure I would have noticed a speed drop of the order your numbers showing and bailed out on that approach.  When one of the machines ends it's current script will do a simple timing evaluation with and with the tensorflow mixed precision.  Seems like I did the fastai equivalent of these several months ago and likewise don't recall a huge speed price for the almost doubling of the batch size.",
      "votes": null
    },
    {
      "id": "951238",
      "postDate": "07/30/2020 02:25:15",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> But how are you doing mixed precision? Are you leveraging NVIDIA amp? if so what optimization level?  </p>",
      "rawMarkdown": "pcjimmmy But how are you doing mixed precision? Are you leveraging NVIDIA amp? if so what optimization level?",
      "votes": null
    },
    {
      "id": "951263",
      "postDate": "07/30/2020 02:51:19",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> I'm not aware of this trick. How do you do it? Do you call <code>model.fit()</code> twice, or change the batchsize dynamically at the beginning of the second epoch. Any code you can provide would help, thanks.</p>\n\n<blockquote>\n  <p>a short warmup with batch size of 4 and than continuing with a new fit and much larger batch size. That really allowed for a huge batch size increase. Tensorflow apparently a huge RAM hog for the first epoch. </p>\n</blockquote>",
      "rawMarkdown": "pcjimmmy I'm not aware of this trick. How do you do it? Do you call `model.fit()` twice, or change the batchsize dynamically at the beginning of the second epoch. Any code you can provide would help, thanks.\n\n&gt; a short warmup with batch size of 4 and than continuing with a new fit and much larger batch size. That really allowed for a huge batch size increase. Tensorflow apparently a huge RAM hog for the first epoch.",
      "votes": null
    },
    {
      "id": "951325",
      "postDate": "07/30/2020 04:14:28",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> </p>\n\n<p>Chris - thinking I might know something that you don't makes me worried that what I am doing does not really work:  I would classify it as a work around.</p>\n\n<p>In the fastai course on one of the early lessons I found option during his fit for a report out on GPU memory.  The first epoch reports out a number near the max of the GPU if you have attempted to maximize the number of batches.  The epochs after the first report out numbers much smaller - memory from 73 year man here for 6 months ago - but I recall around 1/4 of the full GPU memory in use for all the following epochs was typical.</p>\n\n<p>Tensorflow grabs around 96% of available GPU unless you code for memory_growth.  Trying to run large image sizes on GPU for this competition gets you in the 2 or 4 batch size on my system.</p>\n\n<p>I have NOT done this on a Kaggle GPU kernel - only my local Ubuntu machines - have 4 with dual GPU but no two are the same with half dual 11GB and the other two dual 8GB.  I started to test it on Kaggle with TPU.   But the 3 hour time quota is an issue - I burn so much time warming up that no advantage for larger batch - so I killed the run before the warmup ever finished.  So no idea if it also works on a TPU when your not watching the quota clock.</p>\n\n<p>To replicate what I have done:\n1.  Code for memory growth so you can see the GPU memory usage in the warm up vs the real fit.  For whatever reasons my pasting into these posts NEVER looks right - so just sticking in the key line.</p>\n\n<p><code>tf.config.experimental.set_memory_growth(gpu, True)</code></p>\n\n<p>Open a window with nvidia or your fav tool for watching GPU \n<code>watch -n 1 nvidia-smi</code></p>\n\n<p>I put the warmup inside your triple kernel just before your fit, and as you can see it is your fit, with batch size and steps per epoch the only changes.</p>\n\n<p>`# warmup</p>\n\n<pre><code>history = model.fit(\n    get_dataset(files_train, augment=True, shuffle=True, repeat=True,\n            dim=IMG_SIZES[fold], batch_size = 4), \n    epochs=2, callbacks = [sv,get_lr_callback(BATCH_SIZES[fold]), tensorboard_callback], \n    steps_per_epoch=50,\n    validation_data=get_dataset(files_valid,augment=False,shuffle=False,\n            repeat=False,dim=IMG_SIZES[fold]), #class_weight = {0:1,1:2},\n    verbose=VERBOSE\n)`\n</code></pre>\n\n<p>Have run with full number of steps per epoch but that's not needed.  50 or smaller works.  I do two epochs - more is a waste of time - only 1 did not always seem to work - but I am normally  exploring for the right batch size so death when I tried 1 might have been just too big of a batch.</p>\n\n<p>This lets you at least double batch size - since the whole memory vs image not a simple linear - on smaller image sizes it's often 4 times larger batch or more.</p>",
      "rawMarkdown": "cdeotte \n\nChris - thinking I might know something that you don't makes me worried that what I am doing does not really work:  I would classify it as a work around.\n\nIn the fastai course on one of the early lessons I found option during his fit for a report out on GPU memory.  The first epoch reports out a number near the max of the GPU if you have attempted to maximize the number of batches.  The epochs after the first report out numbers much smaller - memory from 73 year man here for 6 months ago - but I recall around 1/4 of the full GPU memory in use for all the following epochs was typical.\n\nTensorflow grabs around 96% of available GPU unless you code for memory_growth.  Trying to run large image sizes on GPU for this competition gets you in the 2 or 4 batch size on my system.\n\nI have NOT done this on a Kaggle GPU kernel - only my local Ubuntu machines - have 4 with dual GPU but no two are the same with half dual 11GB and the other two dual 8GB.  I started to test it on Kaggle with TPU.   But the 3 hour time quota is an issue - I burn so much time warming up that no advantage for larger batch - so I killed the run before the warmup ever finished.  So no idea if it also works on a TPU when your not watching the quota clock.\n\nTo replicate what I have done:\n1.  Code for memory growth so you can see the GPU memory usage in the warm up vs the real fit.  For whatever reasons my pasting into these posts NEVER looks right - so just sticking in the key line.\n\n`tf.config.experimental.set_memory_growth(gpu, True)`\n\nOpen a window with nvidia or your fav tool for watching GPU \n`watch -n 1 nvidia-smi`\n\nI put the warmup inside your triple kernel just before your fit, and as you can see it is your fit, with batch size and steps per epoch the only changes.\n\n`# warmup\n\n    history = model.fit(\n        get_dataset(files_train, augment=True, shuffle=True, repeat=True,\n                dim=IMG_SIZES[fold], batch_size = 4), \n        epochs=2, callbacks = [sv,get_lr_callback(BATCH_SIZES[fold]), tensorboard_callback], \n        steps_per_epoch=50,\n        validation_data=get_dataset(files_valid,augment=False,shuffle=False,\n                repeat=False,dim=IMG_SIZES[fold]), #class_weight = {0:1,1:2},\n        verbose=VERBOSE\n    )`\n\n\nHave run with full number of steps per epoch but that's not needed.  50 or smaller works.  I do two epochs - more is a waste of time - only 1 did not always seem to work - but I am normally  exploring for the right batch size so death when I tried 1 might have been just too big of a batch.\n\nThis lets you at least double batch size - since the whole memory vs image not a simple linear - on smaller image sizes it's often 4 times larger batch or more.",
      "votes": null
    },
    {
      "id": "951329",
      "postDate": "07/30/2020 04:26:35",
      "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> </p>\n\n<p>I make 3 changes - two lines added and a modification of the third.  Tensorflow does all the rest.</p>\n\n<p><code>\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n</code></p>\n\n<p>and set dtype=32 for last layer built.</p>\n\n<p><code>x = tf.keras.layers.Dense(1,activation='sigmoid',dtype='float32')(concat)</code></p>",
      "rawMarkdown": "brianfeeny \n    \nI make 3 changes - two lines added and a modification of the third.  Tensorflow does all the rest.\n\n```\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n```\n\nand set dtype=32 for last layer built.\n\n`x = tf.keras.layers.Dense(1,activation='sigmoid',dtype='float32')(concat)`",
      "votes": null
    },
    {
      "id": "954700",
      "postDate": "08/02/2020 02:25:43",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Great tips Jimmy! Thanks. I will try all this out. </p>\n\n<p>Have you found <code>mixed_float16</code> to work well in TensorFlow? When i use it simultaneously with multiple GPUs <code>tf.distribute.MirroredStrategy()</code>, i thought I noticed that model accuracy decreases. </p>",
      "rawMarkdown": "pcjimmmy Great tips Jimmy! Thanks. I will try all this out. \n\nHave you found `mixed_float16` to work well in TensorFlow? When i use it simultaneously with multiple GPUs `tf.distribute.MirroredStrategy()`, i thought I noticed that model accuracy decreases.",
      "votes": null
    },
    {
      "id": "954772",
      "postDate": "08/02/2020 04:49:09",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> It's a mixed bag. Care should be taken for 20% decrease in training time. </p>",
      "rawMarkdown": "cdeotte It's a mixed bag. Care should be taken for 20% decrease in training time.",
      "votes": null
    },
    {
      "id": "955311",
      "postDate": "08/02/2020 14:40:45",
      "content": "<p>hmm - did not experience this. Can the decreased acc be a random effect or did you see it in most experiments?</p>",
      "rawMarkdown": "hmm - did not experience this. Can the decreased acc be a random effect or did you see it in most experiments?",
      "votes": null
    },
    {
      "id": "955426",
      "postDate": "08/02/2020 16:19:35",
      "content": "<p>ok Roman, perhaps it was a random effect. I was using both multiple GPUs ( <code>tf.distribute.MirroredStrategy()</code> ) and mixed precision, and I noticed that a few experiments had lower accuracy than they usually do. I'll try it some more.</p>",
      "rawMarkdown": "ok Roman, perhaps it was a random effect. I was using both multiple GPUs ( `tf.distribute.MirroredStrategy()` ) and mixed precision, and I noticed that a few experiments had lower accuracy than they usually do. I'll try it some more.",
      "votes": null
    },
    {
      "id": "955654",
      "postDate": "08/02/2020 19:29:52",
      "content": "<p>Chris - accuracy question with mixed ...</p>\n\n<p>Using MirroredStrategy as all 4 of my machines are dual GPU.  Even with your triple stratified tfrecords the few \"formal\" looks at accuracy I have attempted are hampered by the large fold to fold variation for this data set.  It seems like it takes a pretty huge response for anything to get outside that fold to fold std deviation.</p>\n\n<p>Since I was interested in larger model / large image with mixed I have not attempted a look at accuracy with mixed and likely not going to get there for this competition.  Did add it to my to do / wish list to at least do a simple with/without mixed and two different seeds for a 224 image size - that should not burn up too much time on one of my machines.  Will holler here if I get that done.</p>",
      "rawMarkdown": "Chris - accuracy question with mixed ...\n\nUsing MirroredStrategy as all 4 of my machines are dual GPU.  Even with your triple stratified tfrecords the few \"formal\" looks at accuracy I have attempted are hampered by the large fold to fold variation for this data set.  It seems like it takes a pretty huge response for anything to get outside that fold to fold std deviation.\n\nSince I was interested in larger model / large image with mixed I have not attempted a look at accuracy with mixed and likely not going to get there for this competition.  Did add it to my to do / wish list to at least do a simple with/without mixed and two different seeds for a 224 image size - that should not burn up too much time on one of my machines.  Will holler here if I get that done.",
      "votes": null
    },
    {
      "id": "958975",
      "postDate": "08/05/2020 08:54:58",
      "content": "<p>Chris - does this not create a massive embedding for each variable of the meta data? I understand why there may be value in adding at the beginning but adding say sex, age and location to the model would create effectively 3 layers equal to the input image. I am giving it an experiment but it seems like a massive embedding.</p>",
      "rawMarkdown": "Chris - does this not create a massive embedding for each variable of the meta data? I understand why there may be value in adding at the beginning but adding say sex, age and location to the model would create effectively 3 layers equal to the input image. I am giving it an experiment but it seems like a massive embedding.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 944339,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/25/2020 03:10:37",
      "content": "<p>I agree that if we only concatenate the last layers then the combination isn't much better than simple ensemble. If you want our CNN to extract different features, I think we need to add the meta features to the bottom layers. I think we need to try option 1 below. Imagine that the word \"class\" says \"meta data\"\n<img src=\"http://playagricola.com/Kaggle/disc3b.jpg\" alt=\"image\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 944511,
          "author_name": "vladimirsydor",
          "author_url": "",
          "post_date": "07/25/2020 06:58:40",
          "content": "<p>If it is not a secret. Did such approach help you improve CV score?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944554,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/25/2020 07:26:39",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I was thinking about this approach too, but it means we need to add numeric values with some magic to the image itself, because we use transfer learning so the input must be an image.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944569,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/25/2020 07:33:27",
          "content": "<p>There is example TensorFlow code <a href=\"https://www.kaggle.com/cdeotte/dog-breed-acgan-lb-52\">here</a>. I used this in Dog Comp. Option 1 did better than Option 2. By inputting the meta data earlier in the CNN, it allows the CNN to adjust its feature generation as opposed to only adjusting classification.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944612,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/25/2020 07:56:44",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> by quickly looking - you are not using transfer learning in this example but your own net, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944621,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/25/2020 08:04:03",
          "content": "<p>Correct. To use transfer learning, you should apply convolution to convert <code>256x256x4</code> into <code>256x256x3</code></p>\n\n<pre><code>    x = layers.Embedding(IN_SIZE, DIM*DIM, input_length=1)(label)\n    x = layers.Reshape((DIM,DIM,1))(x)\n    x = layers.concatenate([image,x])\n    # NOW WE ARE DIM x DIM x 4\n    x = layers.Conv2D(3)(x)\n    # NOW WE ARE DIM x DIM x 3\n    x = EfficientNetB0()(x)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944646,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/25/2020 08:20:24",
          "content": "<p>you are at 3 channels but your data is not real image anymore, right? so learned weights may be incorrect</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944675,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/25/2020 08:42:45",
          "content": "<p>True but this should work and transfer learning should still help. CNNs are intelligent and will learn how to use this new convolutional layer. Perhaps it will add red to all patients over 60 and remove red from all patients under 60, then the CNN can still use it's transfer learning.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 958975,
          "author_name": "clivep",
          "author_url": "",
          "post_date": "08/05/2020 08:54:58",
          "content": "<p>Chris - does this not create a massive embedding for each variable of the meta data? I understand why there may be value in adding at the beginning but adding say sex, age and location to the model would create effectively 3 layers equal to the input image. I am giving it an experiment but it seems like a massive embedding.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 944346,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "07/25/2020 03:25:58",
      "content": "<p>What I was using until now is this:</p>\n\n<p>For image data:</p>\n\n<pre><code>arch._fc = nn.Linear(in_features = 1280, out_features = 500, bias = True)\n</code></pre>\n\n<p>For meta features:</p>\n\n<pre><code>self.meta = nn.Sequential(nn.Linear(n_meta_features, 500), \n                             nn.BatchNorm1d(500), \n                             nn.ReLU(), \n                             nn.Dropout(p = 0.25), \n                             nn.Linear(500, 250), \n                             nn.BatchNorm1d(250), \n                             nn.ReLU(), \n                             nn.Dropout(p = 0.2))\n</code></pre>\n\n<p>Now, Classification:</p>\n\n<pre><code>self.classifier = nn.Linear(500 + 250, 2)\n</code></pre>\n\n<p>It may not be much different from simple blending, but of course it should work much better than that. </p>\n\n<p>*I could not test it on heavy models, due to some difficulty implementing TPU. :/</p>",
      "votes": null,
      "replies": [
        {
          "id": 944549,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/25/2020 07:24:24",
          "content": "<p>hello <a href=\"/sarques\">@sarques</a> \nthis examply shows exactly what i mean:\n- you create 500 classes from efficientnet (I don't understand why people replace _fc instead just constructing efficientnet with this number of classes - what's the difference?)\n- you create 9 -&gt; 500 -&gt; 250 net for meta features\n- you concat 500 with 250 then you do 750 -&gt; 1 net</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947490,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/27/2020 09:52:25",
          "content": "<p><a href=\"/sarques\">@sarques</a> Why do you take your classifier to an output of 2 instead of 1? Are you using CrossEntropy instead of BinaryCrossEntropy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947532,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "07/27/2020 10:19:49",
          "content": "<p>Yes, I am using CrossEntropy! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947640,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/27/2020 11:50:57",
          "content": "<p>Is there an advantage to using CrossEntropy for a binary problem vs BinaryCrossEntropy?  What's your motivation to use Cross-Entropy instead of Binary Cross Entropy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947747,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "07/27/2020 13:11:07",
          "content": "<p>There is not much about this, it just happened that I started with CrossEntropy, but lately I was using Focal loss for CrossEntropy, also, I can not test my code because there is some problem in implementing TPU with PyTorch XLA. :/ As of now, I am using Chris's notebook of Tensorflow! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 944365,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/25/2020 03:56:44",
      "content": "<p>Since this is the first serious attempt I have made for multiple inputs I am commenting as a complete novice. <br>\nI saw improvement over blending when I did a b0 and densenet169 model that concatenated those two models.  Likewise, I saw improvement over blending when I added a vgg16 model and concat the three.  I did not look super hard but concatenate was the only tf method I saw for mutliple inputs.   </p>\n\n<p>That experience suggests to me that it works over blending.  The issue is that it's very expensive.  I could not run large image sizes on my local Ubuntu machine even at batch size of 2.  It's also likely that my \"blend\" was not optimized.  SO - who knows??  But I believe it is at least an optimized version of blending? </p>\n\n<p>I added a fourth input for the meta data.  I had 16 features.  I started with simple single layer and did not really see an improvement.  But if I was using tf on only the meta data I sure would not settle for a simple single layer.  So playing around I got a tf metadata model.  It was not as good on LB as an XGboost or catboost at 0.73LB.  Adding that model in did show a small improvement, but the improvement was smaller than the std dev among the 5 folds so I don't believe it significant.</p>\n\n<p>My conclusions - garbage in - garbage out.   My 16 features are better than flipping a coin, but nice models may be smart enough to ignore this useless data.  Guess after the competition ends will see if others found big value in the meta data, but I suspect it did not work for you because it is none value added information that models are smart enough to ignore.</p>",
      "votes": null,
      "replies": [
        {
          "id": 944366,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/25/2020 03:58:33",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1062256%2F93d8fca78e4274be85161545b71802cf%2F4inputmodel.png?generation=1595649510770080&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944559,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/25/2020 07:28:48",
          "content": "<p>IMHO they are not useless because training with meta only works, but combining them with CNN suggest they are, I am not really uderstanding what test you did, I tried like 20 different combinations yesterday</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947512,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/27/2020 10:08:05",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Interesting net you have there.  So you were using 3 pre-trained nets at the same time, and then added in the Meta data as well.  What kind of hardware are you using to handle that?  And you say that you found this way of doing things with concatenating, better than say just taking prediction outputs from the separate models and combining them?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948470,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/28/2020 01:40:53",
          "content": "<p>4 machines all - Ubuntu 20.04 - Intel® Core™ i7-9700K CPU @ 3.60GHz  dual gpu's on all the machines - either 1070 and 1080.  Pretty sure it would not run on a kaggle kernel but ok on my setups.</p>\n\n<p>My single pre-trained blends were not weighted so I don't know if a little bit of work could have weighted the blends to match the triple model.  I was doing a different augmentation of the image for each of the three.  Gave up on the approach as it's expensive and biggest image size was not going to be very large.  The expensive part just really needed me to have more patience - but letting things run for two days before you get a hint if your on the right track is beyond me.  But not getting a decent image size seemed like the real road block.   </p>\n\n<p>Been playing with mixed precision on single models and might return.  But this discussion post makes me worry that concatenating this way will not perform better than separate and properly weighted blends.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948498,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/28/2020 02:42:34",
          "content": "<p>Mixed precision may not be stable on a 1080/1070.  Realize the 16bit performance of both of those cards is terrible.  Are you using Distributed Data Paralllel (DDP)?  That would seem like a good approach given you have 4 separate nodes..........the faster the connectivity between the machines the better though, as that could be a source of bottleneck.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949911,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/29/2020 04:22:39",
          "content": "<p>Sorry I confused you - I have 4 machines but running different versions of the 4 input models separately.  Was thinking about running all 4 on the same script but decided to wait until a future competition - doing that is outside my skill set so it would be a full on learning experience but the tensorflow docs make it seem something learnable.</p>\n\n<p>Mixed precision on single model seems to be running ok - nice boost in the size of the batch that can be run - I also started doing a short warmup with batch size of 4 and than continuing with a new fit and much larger batch size.  That really allowed for a huge batch size increase.  Tensorflow apparently a huge RAM hog for the first epoch.   Between mixed and the warmup I am going to retry the 4 input model again to see how big I can get the image. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 950142,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/29/2020 07:53:04",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> are you using NVIDIA's AMP for mixed-precision?  If so, what optimization level do you have it on?  I just don't understand how a 1070 and 1080 can do true 16 bit at any sort of appreciable speed.</p>\n\n<p>1070 is 5783 GFLOPS @ 32bit, 90 GFLOPS @ 16bit (yes 90, not a typo)\n1080 is 8228 GFLOPS @ 32bit, 128 GFLOPS @ 16bit</p>\n\n<p>I have no experience in mixed precision.  Soon I am swapping my 1080 Ti's out for 2080 Ti's, and those have TensorCores and can actually do halfway decent 16-bit precision.  I am looking forward to be abl able to take on larger models and/or speed up my training times.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951169,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/30/2020 00:39:40",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> </p>\n\n<p>Running very standard Ubuntu 20.04 - my 1070's on one of the machines are more specifically as shown on settings info:\nGeForce GTX 1070/PCIe/SSE2 / GeForce GTX 1070/PCIe/SSE2</p>\n\n<p>With mixed precision I am not quite able to double the batch size and I am pretty sure I would have noticed a speed drop of the order your numbers showing and bailed out on that approach.  When one of the machines ends it's current script will do a simple timing evaluation with and with the tensorflow mixed precision.  Seems like I did the fastai equivalent of these several months ago and likewise don't recall a huge speed price for the almost doubling of the batch size.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951238,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/30/2020 02:25:15",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> But how are you doing mixed precision? Are you leveraging NVIDIA amp? if so what optimization level?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951263,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/30/2020 02:51:19",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> I'm not aware of this trick. How do you do it? Do you call <code>model.fit()</code> twice, or change the batchsize dynamically at the beginning of the second epoch. Any code you can provide would help, thanks.</p>\n\n<blockquote>\n  <p>a short warmup with batch size of 4 and than continuing with a new fit and much larger batch size. That really allowed for a huge batch size increase. Tensorflow apparently a huge RAM hog for the first epoch. </p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951325,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/30/2020 04:14:28",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> </p>\n\n<p>Chris - thinking I might know something that you don't makes me worried that what I am doing does not really work:  I would classify it as a work around.</p>\n\n<p>In the fastai course on one of the early lessons I found option during his fit for a report out on GPU memory.  The first epoch reports out a number near the max of the GPU if you have attempted to maximize the number of batches.  The epochs after the first report out numbers much smaller - memory from 73 year man here for 6 months ago - but I recall around 1/4 of the full GPU memory in use for all the following epochs was typical.</p>\n\n<p>Tensorflow grabs around 96% of available GPU unless you code for memory_growth.  Trying to run large image sizes on GPU for this competition gets you in the 2 or 4 batch size on my system.</p>\n\n<p>I have NOT done this on a Kaggle GPU kernel - only my local Ubuntu machines - have 4 with dual GPU but no two are the same with half dual 11GB and the other two dual 8GB.  I started to test it on Kaggle with TPU.   But the 3 hour time quota is an issue - I burn so much time warming up that no advantage for larger batch - so I killed the run before the warmup ever finished.  So no idea if it also works on a TPU when your not watching the quota clock.</p>\n\n<p>To replicate what I have done:\n1.  Code for memory growth so you can see the GPU memory usage in the warm up vs the real fit.  For whatever reasons my pasting into these posts NEVER looks right - so just sticking in the key line.</p>\n\n<p><code>tf.config.experimental.set_memory_growth(gpu, True)</code></p>\n\n<p>Open a window with nvidia or your fav tool for watching GPU \n<code>watch -n 1 nvidia-smi</code></p>\n\n<p>I put the warmup inside your triple kernel just before your fit, and as you can see it is your fit, with batch size and steps per epoch the only changes.</p>\n\n<p>`# warmup</p>\n\n<pre><code>history = model.fit(\n    get_dataset(files_train, augment=True, shuffle=True, repeat=True,\n            dim=IMG_SIZES[fold], batch_size = 4), \n    epochs=2, callbacks = [sv,get_lr_callback(BATCH_SIZES[fold]), tensorboard_callback], \n    steps_per_epoch=50,\n    validation_data=get_dataset(files_valid,augment=False,shuffle=False,\n            repeat=False,dim=IMG_SIZES[fold]), #class_weight = {0:1,1:2},\n    verbose=VERBOSE\n)`\n</code></pre>\n\n<p>Have run with full number of steps per epoch but that's not needed.  50 or smaller works.  I do two epochs - more is a waste of time - only 1 did not always seem to work - but I am normally  exploring for the right batch size so death when I tried 1 might have been just too big of a batch.</p>\n\n<p>This lets you at least double batch size - since the whole memory vs image not a simple linear - on smaller image sizes it's often 4 times larger batch or more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951329,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "07/30/2020 04:26:35",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> </p>\n\n<p>I make 3 changes - two lines added and a modification of the third.  Tensorflow does all the rest.</p>\n\n<p><code>\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n</code></p>\n\n<p>and set dtype=32 for last layer built.</p>\n\n<p><code>x = tf.keras.layers.Dense(1,activation='sigmoid',dtype='float32')(concat)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 954700,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/02/2020 02:25:43",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Great tips Jimmy! Thanks. I will try all this out. </p>\n\n<p>Have you found <code>mixed_float16</code> to work well in TensorFlow? When i use it simultaneously with multiple GPUs <code>tf.distribute.MirroredStrategy()</code>, i thought I noticed that model accuracy decreases. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 954772,
          "author_name": "teeyee314",
          "author_url": "",
          "post_date": "08/02/2020 04:49:09",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> It's a mixed bag. Care should be taken for 20% decrease in training time. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955311,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "08/02/2020 14:40:45",
          "content": "<p>hmm - did not experience this. Can the decreased acc be a random effect or did you see it in most experiments?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955426,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/02/2020 16:19:35",
          "content": "<p>ok Roman, perhaps it was a random effect. I was using both multiple GPUs ( <code>tf.distribute.MirroredStrategy()</code> ) and mixed precision, and I noticed that a few experiments had lower accuracy than they usually do. I'll try it some more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955654,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "08/02/2020 19:29:52",
          "content": "<p>Chris - accuracy question with mixed ...</p>\n\n<p>Using MirroredStrategy as all 4 of my machines are dual GPU.  Even with your triple stratified tfrecords the few \"formal\" looks at accuracy I have attempted are hampered by the large fold to fold variation for this data set.  It seems like it takes a pretty huge response for anything to get outside that fold to fold std deviation.</p>\n\n<p>Since I was interested in larger model / large image with mixed I have not attempted a look at accuracy with mixed and likely not going to get there for this competition.  Did add it to my to do / wish list to at least do a simple with/without mixed and two different seeds for a 224 image size - that should not burn up too much time on one of my machines.  Will holler here if I get that done.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 946234,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "07/26/2020 13:00:41",
      "content": "<p>Tested another about 20-30 models, without any effect.\nLooks like this is idea for another competition, and for this one I just need to blend meta with CNN in one Linear layer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 946436,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/26/2020 15:19:25",
      "content": "<blockquote>\n  <p>However, in public kernel metadata is concatenated with CNN output, then there is one last Linear/Dense layer and we have single output.</p>\n  \n  <p>It means that two models (metadata and CNN) are just blended. There is no combination of \"this output from metadata and this output from CNN work together\". We have just some number of outputs from CNN and some number of outputs from metadata we multiply each with weight, sum them up and that's all.</p>\n</blockquote>\n\n<p>This is not completely true. When you concatenate meta and image embeddings in the last layer of a CNN it will affect the model's back propagation learning and thus affect which convolution filters are created.</p>\n\n<p>For example, a model <strong>without</strong> meta features may learn to create convolution filters that differentiate skin color because male and female have different skin color. But a model <strong>with</strong> meta features already has access to gender meta feature. So that model may not create convolution filters that differentiate skin color (to differentiate male and female) because it already has access to the information.</p>\n\n<p>So, in the same way that dropout forces your model to look at new parts of the image. Adding meta features also forces your model to look at new parts of the image.</p>",
      "votes": null,
      "replies": [
        {
          "id": 946603,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/26/2020 17:10:31",
          "content": "<p>I disagree but probably we think about same thing.</p>\n\n<p>When you have one Linear layer you only add probabilities, there is no way to communicate between CNN and meta.</p>\n\n<p>With one more Linear layer backpropagation could send back information to CNN to look at images more needed for specific meta.</p>\n\n<p>What I can agree is that when model will learn that larger age means more probability for 1 then CNN can spend less resources on learning images from people with larger age. But that's all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 946630,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/26/2020 17:28:14",
          "content": "<p>Communication occurs. For example if you concatenate the true labels as in </p>\n\n<pre><code>x = EfficientNetB0(input)\nx = GlobalAveragePooling2D()(x)\nx = Concatenate()([x, TRUE_TARGETS])\nx = Dense(1, activation='sigmoid')(x)\n</code></pre>\n\n<p>Then the EfficientNetB0 backbone will not get any learning. Learning is driven by <code>error</code>. When you concatenate another source of features such as meta data, you are affecting the <code>error</code> and thus you are affecting the backpropagation of learning (which affects all proceeding input feature pipelines).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 949211,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "07/28/2020 13:21:26",
      "content": "<p>I think some of the issues come from the train &amp; test datasets being slightly different. This makes it possible that a model would perform well on CV but not so good on test. Raddar did some experiments here:\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155885\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155885</a></p>\n\n<p>While writing this post, I started writing down reasons why this might be, for example, 62% of the positive cases are male which seemed way too high to me (gut feeling). But then I Googled \"melanoma male vs female\" and was surprised with what I saw.</p>\n\n<p>I think the trick is choosing which metadata features to use and which to discard.</p>",
      "votes": null,
      "replies": [
        {
          "id": 949535,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/28/2020 17:37:17",
          "content": "<p><a href=\"/anjum48\">@anjum48</a> What features are people using besides Gender, Age and Site?  Those all seem valid..........not sure it would make sense to use anything else.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949556,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "07/28/2020 17:56:48",
          "content": "<p><a href=\"/brianfeeny\">@brianfeeny</a> If you use count of patient id, you can push a metadata only model up to 0.74. This might make sense because if a patient has lots of benign moles that look similar, then the model can use this information to adjust its predictions (and vice-versa for malignant using the Ugly Duckling concept).</p>\n\n<p>Also, age is probably a risk factor, but the number of people above 70 is rare, so you could use frequency encoded features. This can get a meta only model above 0.77 (CV) when combined with the feature above.</p>\n\n<p>But if the test set is different from train then this approach can create issues in having a model that will generalise well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949574,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "07/28/2020 18:07:34",
          "content": "<p>Yeah, using things like patent id has to be well thought through.  My personal gut feeling is that is going to be overfitting and not a good feature.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949595,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/28/2020 18:27:12",
          "content": "<p>Using meta features is frustrating and risky because the data was not randomly split between <code>train.csv</code> and <code>test.csv</code>. There are many artificially created correlations (i.e. false patterns in train data).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "944265": "In metadata we have age, sex and area.\nWe can train model on metadata only and we can see some score larger than 0.5, so yes, data is valid and it is useful.\n\nHowever, in public kernel metadata is concatenated with CNN output, then there is one last Linear/Dense layer and we have single output.\n\nIt means that two models (metadata and CNN) are just blended. There is no combination of \"this output from metadata and this output from CNN  work together\". We have just some number of outputs from CNN and some number of outputs from metadata we multiply each with weight, sum them up and that's all.\n\nI would assume that metadata will create outputs like \"this specific area of woman aged about 60\" and then CNN will create output to combine with this metadata output. But adding additional Linear layer never helps, I did lots of tests today.\n\nDo you have any other experiences or do you know why it works like that?",
    "944339": "I agree that if we only concatenate the last layers then the combination isn't much better than simple ensemble. If you want our CNN to extract different features, I think we need to add the meta features to the bottom layers. I think we need to try option 1 below. Imagine that the word \"class\" says \"meta data\"\n![image](http://playagricola.com/Kaggle/disc3b.jpg)",
    "944346": "What I was using until now is this:\n\nFor image data:\n\n    arch._fc = nn.Linear(in_features = 1280, out_features = 500, bias = True)\n\nFor meta features:\n    \n    self.meta = nn.Sequential(nn.Linear(n_meta_features, 500), \n                                 nn.BatchNorm1d(500), \n                                 nn.ReLU(), \n                                 nn.Dropout(p = 0.25), \n                                 nn.Linear(500, 250), \n                                 nn.BatchNorm1d(250), \n                                 nn.ReLU(), \n                                 nn.Dropout(p = 0.2))\n\nNow, Classification:\n\n    self.classifier = nn.Linear(500 + 250, 2)\n\nIt may not be much different from simple blending, but of course it should work much better than that. \n\n*I could not test it on heavy models, due to some difficulty implementing TPU. :/",
    "944365": "Since this is the first serious attempt I have made for multiple inputs I am commenting as a complete novice.  \nI saw improvement over blending when I did a b0 and densenet169 model that concatenated those two models.  Likewise, I saw improvement over blending when I added a vgg16 model and concat the three.  I did not look super hard but concatenate was the only tf method I saw for mutliple inputs.   \n\nThat experience suggests to me that it works over blending.  The issue is that it's very expensive.  I could not run large image sizes on my local Ubuntu machine even at batch size of 2.  It's also likely that my \"blend\" was not optimized.  SO - who knows??  But I believe it is at least an optimized version of blending? \n\nI added a fourth input for the meta data.  I had 16 features.  I started with simple single layer and did not really see an improvement.  But if I was using tf on only the meta data I sure would not settle for a simple single layer.  So playing around I got a tf metadata model.  It was not as good on LB as an XGboost or catboost at 0.73LB.  Adding that model in did show a small improvement, but the improvement was smaller than the std dev among the 5 folds so I don't believe it significant.\n\nMy conclusions - garbage in - garbage out.   My 16 features are better than flipping a coin, but nice models may be smart enough to ignore this useless data.  Guess after the competition ends will see if others found big value in the meta data, but I suspect it did not work for you because it is none value added information that models are smart enough to ignore.",
    "944366": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1062256%2F93d8fca78e4274be85161545b71802cf%2F4inputmodel.png?generation=1595649510770080&amp;alt=media)",
    "944511": "If it is not a secret. Did such approach help you improve CV score?",
    "944549": "hello @sarques \nthis examply shows exactly what i mean:\n- you create 500 classes from efficientnet (I don't understand why people replace _fc instead just constructing efficientnet with this number of classes - what's the difference?)\n- you create 9 -&gt; 500 -&gt; 250 net for meta features\n- you concat 500 with 250 then you do 750 -&gt; 1 net",
    "944554": "cdeotte I was thinking about this approach too, but it means we need to add numeric values with some magic to the image itself, because we use transfer learning so the input must be an image.",
    "944559": "IMHO they are not useless because training with meta only works, but combining them with CNN suggest they are, I am not really uderstanding what test you did, I tried like 20 different combinations yesterday",
    "944569": "There is example TensorFlow code [here][1]. I used this in Dog Comp. Option 1 did better than Option 2. By inputting the meta data earlier in the CNN, it allows the CNN to adjust its feature generation as opposed to only adjusting classification.\n\n[1]: https://www.kaggle.com/cdeotte/dog-breed-acgan-lb-52",
    "944612": "cdeotte by quickly looking - you are not using transfer learning in this example but your own net, right?",
    "944621": "Correct. To use transfer learning, you should apply convolution to convert `256x256x4` into `256x256x3`\n\n        x = layers.Embedding(IN_SIZE, DIM*DIM, input_length=1)(label)\n        x = layers.Reshape((DIM,DIM,1))(x)\n        x = layers.concatenate([image,x])\n        # NOW WE ARE DIM x DIM x 4\n        x = layers.Conv2D(3)(x)\n        # NOW WE ARE DIM x DIM x 3\n        x = EfficientNetB0()(x)",
    "944646": "you are at 3 channels but your data is not real image anymore, right? so learned weights may be incorrect",
    "944675": "True but this should work and transfer learning should still help. CNNs are intelligent and will learn how to use this new convolutional layer. Perhaps it will add red to all patients over 60 and remove red from all patients under 60, then the CNN can still use it's transfer learning.",
    "946234": "Tested another about 20-30 models, without any effect.\nLooks like this is idea for another competition, and for this one I just need to blend meta with CNN in one Linear layer.",
    "946436": "&gt; However, in public kernel metadata is concatenated with CNN output, then there is one last Linear/Dense layer and we have single output.\n  \n&gt; It means that two models (metadata and CNN) are just blended. There is no combination of \"this output from metadata and this output from CNN work together\". We have just some number of outputs from CNN and some number of outputs from metadata we multiply each with weight, sum them up and that's all.\n\nThis is not completely true. When you concatenate meta and image embeddings in the last layer of a CNN it will affect the model's back propagation learning and thus affect which convolution filters are created.\n\nFor example, a model **without** meta features may learn to create convolution filters that differentiate skin color because male and female have different skin color. But a model **with** meta features already has access to gender meta feature. So that model may not create convolution filters that differentiate skin color (to differentiate male and female) because it already has access to the information.\n\nSo, in the same way that dropout forces your model to look at new parts of the image. Adding meta features also forces your model to look at new parts of the image.",
    "946603": "I disagree but probably we think about same thing.\n\nWhen you have one Linear layer you only add probabilities, there is no way to communicate between CNN and meta.\n\nWith one more Linear layer backpropagation could send back information to CNN to look at images more needed for specific meta.\n\nWhat I can agree is that when model will learn that larger age means more probability for 1 then CNN can spend less resources on learning images from people with larger age. But that's all.",
    "946630": "Communication occurs. For example if you concatenate the true labels as in \n\n    x = EfficientNetB0(input)\n    x = GlobalAveragePooling2D()(x)\n    x = Concatenate()([x, TRUE_TARGETS])\n    x = Dense(1, activation='sigmoid')(x)\n\nThen the EfficientNetB0 backbone will not get any learning. Learning is driven by `error`. When you concatenate another source of features such as meta data, you are affecting the `error` and thus you are affecting the backpropagation of learning (which affects all proceeding input feature pipelines).",
    "947490": "sarques Why do you take your classifier to an output of 2 instead of 1? Are you using CrossEntropy instead of BinaryCrossEntropy?",
    "947512": "pcjimmmy Interesting net you have there.  So you were using 3 pre-trained nets at the same time, and then added in the Meta data as well.  What kind of hardware are you using to handle that?  And you say that you found this way of doing things with concatenating, better than say just taking prediction outputs from the separate models and combining them?",
    "947532": "Yes, I am using CrossEntropy! :)",
    "947640": "Is there an advantage to using CrossEntropy for a binary problem vs BinaryCrossEntropy?  What's your motivation to use Cross-Entropy instead of Binary Cross Entropy?",
    "947747": "There is not much about this, it just happened that I started with CrossEntropy, but lately I was using Focal loss for CrossEntropy, also, I can not test my code because there is some problem in implementing TPU with PyTorch XLA. :/ As of now, I am using Chris's notebook of Tensorflow! :)",
    "948470": "4 machines all - Ubuntu 20.04 - Intel® Core™ i7-9700K CPU @ 3.60GHz  dual gpu's on all the machines - either 1070 and 1080.  Pretty sure it would not run on a kaggle kernel but ok on my setups.\n\nMy single pre-trained blends were not weighted so I don't know if a little bit of work could have weighted the blends to match the triple model.  I was doing a different augmentation of the image for each of the three.  Gave up on the approach as it's expensive and biggest image size was not going to be very large.  The expensive part just really needed me to have more patience - but letting things run for two days before you get a hint if your on the right track is beyond me.  But not getting a decent image size seemed like the real road block.   \n\nBeen playing with mixed precision on single models and might return.  But this discussion post makes me worry that concatenating this way will not perform better than separate and properly weighted blends.",
    "948498": "Mixed precision may not be stable on a 1080/1070.  Realize the 16bit performance of both of those cards is terrible.  Are you using Distributed Data Paralllel (DDP)?  That would seem like a good approach given you have 4 separate nodes..........the faster the connectivity between the machines the better though, as that could be a source of bottleneck.",
    "949211": "I think some of the issues come from the train &amp; test datasets being slightly different. This makes it possible that a model would perform well on CV but not so good on test. Raddar did some experiments here:\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155885\n\nWhile writing this post, I started writing down reasons why this might be, for example, 62% of the positive cases are male which seemed way too high to me (gut feeling). But then I Googled \"melanoma male vs female\" and was surprised with what I saw.\n\nI think the trick is choosing which metadata features to use and which to discard.",
    "949535": "anjum48 What features are people using besides Gender, Age and Site?  Those all seem valid..........not sure it would make sense to use anything else.",
    "949556": "brianfeeny If you use count of patient id, you can push a metadata only model up to 0.74. This might make sense because if a patient has lots of benign moles that look similar, then the model can use this information to adjust its predictions (and vice-versa for malignant using the Ugly Duckling concept).\n\nAlso, age is probably a risk factor, but the number of people above 70 is rare, so you could use frequency encoded features. This can get a meta only model above 0.77 (CV) when combined with the feature above.\n\nBut if the test set is different from train then this approach can create issues in having a model that will generalise well.",
    "949574": "Yeah, using things like patent id has to be well thought through.  My personal gut feeling is that is going to be overfitting and not a good feature.",
    "949595": "Using meta features is frustrating and risky because the data was not randomly split between `train.csv` and `test.csv`. There are many artificially created correlations (i.e. false patterns in train data).",
    "949911": "Sorry I confused you - I have 4 machines but running different versions of the 4 input models separately.  Was thinking about running all 4 on the same script but decided to wait until a future competition - doing that is outside my skill set so it would be a full on learning experience but the tensorflow docs make it seem something learnable.\n\nMixed precision on single model seems to be running ok - nice boost in the size of the batch that can be run - I also started doing a short warmup with batch size of 4 and than continuing with a new fit and much larger batch size.  That really allowed for a huge batch size increase.  Tensorflow apparently a huge RAM hog for the first epoch.   Between mixed and the warmup I am going to retry the 4 input model again to see how big I can get the image.",
    "950142": "pcjimmmy are you using NVIDIA's AMP for mixed-precision?  If so, what optimization level do you have it on?  I just don't understand how a 1070 and 1080 can do true 16 bit at any sort of appreciable speed.\n\n1070 is 5783 GFLOPS @ 32bit, 90 GFLOPS @ 16bit (yes 90, not a typo)\n1080 is 8228 GFLOPS @ 32bit, 128 GFLOPS @ 16bit\n\nI have no experience in mixed precision.  Soon I am swapping my 1080 Ti's out for 2080 Ti's, and those have TensorCores and can actually do halfway decent 16-bit precision.  I am looking forward to be abl able to take on larger models and/or speed up my training times.",
    "951169": "brianfeeny \n\nRunning very standard Ubuntu 20.04 - my 1070's on one of the machines are more specifically as shown on settings info:\nGeForce GTX 1070/PCIe/SSE2 / GeForce GTX 1070/PCIe/SSE2\n\nWith mixed precision I am not quite able to double the batch size and I am pretty sure I would have noticed a speed drop of the order your numbers showing and bailed out on that approach.  When one of the machines ends it's current script will do a simple timing evaluation with and with the tensorflow mixed precision.  Seems like I did the fastai equivalent of these several months ago and likewise don't recall a huge speed price for the almost doubling of the batch size.",
    "951238": "pcjimmmy But how are you doing mixed precision? Are you leveraging NVIDIA amp? if so what optimization level?",
    "951263": "pcjimmmy I'm not aware of this trick. How do you do it? Do you call `model.fit()` twice, or change the batchsize dynamically at the beginning of the second epoch. Any code you can provide would help, thanks.\n\n&gt; a short warmup with batch size of 4 and than continuing with a new fit and much larger batch size. That really allowed for a huge batch size increase. Tensorflow apparently a huge RAM hog for the first epoch.",
    "951325": "cdeotte \n\nChris - thinking I might know something that you don't makes me worried that what I am doing does not really work:  I would classify it as a work around.\n\nIn the fastai course on one of the early lessons I found option during his fit for a report out on GPU memory.  The first epoch reports out a number near the max of the GPU if you have attempted to maximize the number of batches.  The epochs after the first report out numbers much smaller - memory from 73 year man here for 6 months ago - but I recall around 1/4 of the full GPU memory in use for all the following epochs was typical.\n\nTensorflow grabs around 96% of available GPU unless you code for memory_growth.  Trying to run large image sizes on GPU for this competition gets you in the 2 or 4 batch size on my system.\n\nI have NOT done this on a Kaggle GPU kernel - only my local Ubuntu machines - have 4 with dual GPU but no two are the same with half dual 11GB and the other two dual 8GB.  I started to test it on Kaggle with TPU.   But the 3 hour time quota is an issue - I burn so much time warming up that no advantage for larger batch - so I killed the run before the warmup ever finished.  So no idea if it also works on a TPU when your not watching the quota clock.\n\nTo replicate what I have done:\n1.  Code for memory growth so you can see the GPU memory usage in the warm up vs the real fit.  For whatever reasons my pasting into these posts NEVER looks right - so just sticking in the key line.\n\n`tf.config.experimental.set_memory_growth(gpu, True)`\n\nOpen a window with nvidia or your fav tool for watching GPU \n`watch -n 1 nvidia-smi`\n\nI put the warmup inside your triple kernel just before your fit, and as you can see it is your fit, with batch size and steps per epoch the only changes.\n\n`# warmup\n\n    history = model.fit(\n        get_dataset(files_train, augment=True, shuffle=True, repeat=True,\n                dim=IMG_SIZES[fold], batch_size = 4), \n        epochs=2, callbacks = [sv,get_lr_callback(BATCH_SIZES[fold]), tensorboard_callback], \n        steps_per_epoch=50,\n        validation_data=get_dataset(files_valid,augment=False,shuffle=False,\n                repeat=False,dim=IMG_SIZES[fold]), #class_weight = {0:1,1:2},\n        verbose=VERBOSE\n    )`\n\n\nHave run with full number of steps per epoch but that's not needed.  50 or smaller works.  I do two epochs - more is a waste of time - only 1 did not always seem to work - but I am normally  exploring for the right batch size so death when I tried 1 might have been just too big of a batch.\n\nThis lets you at least double batch size - since the whole memory vs image not a simple linear - on smaller image sizes it's often 4 times larger batch or more.",
    "951329": "brianfeeny \n    \nI make 3 changes - two lines added and a modification of the third.  Tensorflow does all the rest.\n\n```\npolicy = tf.keras.mixed_precision.experimental.Policy('mixed_float16')\nmixed_precision.set_policy(policy)\n```\n\nand set dtype=32 for last layer built.\n\n`x = tf.keras.layers.Dense(1,activation='sigmoid',dtype='float32')(concat)`",
    "954700": "pcjimmmy Great tips Jimmy! Thanks. I will try all this out. \n\nHave you found `mixed_float16` to work well in TensorFlow? When i use it simultaneously with multiple GPUs `tf.distribute.MirroredStrategy()`, i thought I noticed that model accuracy decreases.",
    "954772": "cdeotte It's a mixed bag. Care should be taken for 20% decrease in training time.",
    "955311": "hmm - did not experience this. Can the decreased acc be a random effect or did you see it in most experiments?",
    "955426": "ok Roman, perhaps it was a random effect. I was using both multiple GPUs ( `tf.distribute.MirroredStrategy()` ) and mixed precision, and I noticed that a few experiments had lower accuracy than they usually do. I'll try it some more.",
    "955654": "Chris - accuracy question with mixed ...\n\nUsing MirroredStrategy as all 4 of my machines are dual GPU.  Even with your triple stratified tfrecords the few \"formal\" looks at accuracy I have attempted are hampered by the large fold to fold variation for this data set.  It seems like it takes a pretty huge response for anything to get outside that fold to fold std deviation.\n\nSince I was interested in larger model / large image with mixed I have not attempted a look at accuracy with mixed and likely not going to get there for this competition.  Did add it to my to do / wish list to at least do a simple with/without mixed and two different seeds for a 224 image size - that should not burn up too much time on one of my machines.  Will holler here if I get that done.",
    "958975": "Chris - does this not create a massive embedding for each variable of the meta data? I understand why there may be value in adding at the beginning but adding say sex, age and location to the model would create effectively 3 layers equal to the input image. I am giving it an experiment but it seems like a massive embedding."
  },
  "source": "meta"
}