{
  "id": 168954,
  "title": "Exploring The limits of Meta-Features using Custom Tabnet",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168954",
  "author_name": "",
  "post_date": "2020-07-22T13:13:09.579355900Z",
  "votes": 10,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In this competition we have been given two types of data , one is the images of skin lesions of patients , other is the tabular meta data . Now there are three ways of combining these two information together :-</p>\n\n<ul>\n<li>Build a CNN image model and find a way to input the tabular data into the CNN image model</li>\n<li>Build a Tabular data model and find a way to extract image embeddings or image features and input into the Tabular data model</li>\n<li>Build 2 separate models for images and metadata and ensemble</li>\n</ul>\n\n<p>We have tried all three and the third option works the best and gives significant amount of boost.Another question which comes to mind is, what models can we use for modelling with tabular data , my notebook tries to answer this question. It can be found here <a href=\"https://www.kaggle.com/tanulsingh077/exploring-limits-of-meta-features-using-tabnet\">https://www.kaggle.com/tanulsingh077/exploring-limits-of-meta-features-using-tabnet</a></p>\n\n<p>I present how to use Tabnet as a custom Model instead of the scikit-learn type interface provided by Pytorch-Tabnet . I show how anyone can use Tabnet just like a torchvision or any torch-hub model for any downstream task . My notebook is written keeping this specific melanoma competition in mind but the code can be easily used for your own tasks by changing the dataloader</p>\n\n<p>Below is the resulting Graph of one fold training which is done without any fine-tuning/feature engg . It easily gets 0.68 valid_auc without much efforts , I have been able to get to 0.715 local auc , still have to check lb , I have shared it with the commnunity so that it's full potential can be reached\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779944%2Fa6a518ffb74c6fb512d377fdcd0e0348%2Ftab.PNG?generation=1595423398069884&amp;alt=media\" alt=\"\"></p>\n\n<p>I hope you find my efforts useful, Thanks for reading</p>",
  "messages": [
    {
      "id": "939786",
      "postDate": "07/22/2020 13:13:09",
      "content": "<p>In this competition we have been given two types of data , one is the images of skin lesions of patients , other is the tabular meta data . Now there are three ways of combining these two information together :-</p>\n\n<ul>\n<li>Build a CNN image model and find a way to input the tabular data into the CNN image model</li>\n<li>Build a Tabular data model and find a way to extract image embeddings or image features and input into the Tabular data model</li>\n<li>Build 2 separate models for images and metadata and ensemble</li>\n</ul>\n\n<p>We have tried all three and the third option works the best and gives significant amount of boost.Another question which comes to mind is, what models can we use for modelling with tabular data , my notebook tries to answer this question. It can be found here <a href=\"https://www.kaggle.com/tanulsingh077/exploring-limits-of-meta-features-using-tabnet\">https://www.kaggle.com/tanulsingh077/exploring-limits-of-meta-features-using-tabnet</a></p>\n\n<p>I present how to use Tabnet as a custom Model instead of the scikit-learn type interface provided by Pytorch-Tabnet . I show how anyone can use Tabnet just like a torchvision or any torch-hub model for any downstream task . My notebook is written keeping this specific melanoma competition in mind but the code can be easily used for your own tasks by changing the dataloader</p>\n\n<p>Below is the resulting Graph of one fold training which is done without any fine-tuning/feature engg . It easily gets 0.68 valid_auc without much efforts , I have been able to get to 0.715 local auc , still have to check lb , I have shared it with the commnunity so that it's full potential can be reached\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779944%2Fa6a518ffb74c6fb512d377fdcd0e0348%2Ftab.PNG?generation=1595423398069884&amp;alt=media\" alt=\"\"></p>\n\n<p>I hope you find my efforts useful, Thanks for reading</p>",
      "rawMarkdown": "In this competition we have been given two types of data , one is the images of skin lesions of patients , other is the tabular meta data . Now there are three ways of combining these two information together :-\n\n* Build a CNN image model and find a way to input the tabular data into the CNN image model\n* Build a Tabular data model and find a way to extract image embeddings or image features and input into the Tabular data model\n* Build 2 separate models for images and metadata and ensemble\n\nWe have tried all three and the third option works the best and gives significant amount of boost.Another question which comes to mind is, what models can we use for modelling with tabular data , my notebook tries to answer this question. It can be found here https://www.kaggle.com/tanulsingh077/exploring-limits-of-meta-features-using-tabnet\n\nI present how to use Tabnet as a custom Model instead of the scikit-learn type interface provided by Pytorch-Tabnet . I show how anyone can use Tabnet just like a torchvision or any torch-hub model for any downstream task . My notebook is written keeping this specific melanoma competition in mind but the code can be easily used for your own tasks by changing the dataloader\n\nBelow is the resulting Graph of one fold training which is done without any fine-tuning/feature engg . It easily gets 0.68 valid_auc without much efforts , I have been able to get to 0.715 local auc , still have to check lb , I have shared it with the commnunity so that it's full potential can be reached\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779944%2Fa6a518ffb74c6fb512d377fdcd0e0348%2Ftab.PNG?generation=1595423398069884&amp;alt=media)\n\n\nI hope you find my efforts useful, Thanks for reading",
      "votes": null
    },
    {
      "id": "939972",
      "postDate": "07/22/2020 15:48:45",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a>,</p>\n<p>Thanks for sharing this and all the citings. I'm trying to add a comment directly on the notebook but it seems that I can't add the first comment, UX problem from Kaggle, so I'll leave it here.</p>\n<p>Honestly I'm a bit confused by this notebook that seems incredibly complicated for something that should be simple, I know that some simple customization are not available yet on <code>pytorch-tabnet</code> and I'll try to make things easier, but I guess some things here are over complicated.</p>\n<p>I'll leave here a few comments/questions if you don't mind, please correct me if I got something wrong:</p>\n<ul>\n<li>all this customization is actually used for one reason: changing the loss function (and the early stopping function but tabnet early stops on AUC which seems reasonable here)? But did you know that there is a <code>loss_fn</code> parameter in TabNet's <code>fit</code> method that would allow you to switch from logloss to <code>SoftMarginFocalLoss</code> with a one liner? (Let me know if you encounter any bug when trying this - I'll try to be responsive on the repo).</li>\n<li>About the CustomTabNet class, is this just to give an example on how on could use tabnet inside a bigger class? Because at the moment it looks like a perfect duplicate of <code>TabNet</code>, so you could have just used <code>TabNet</code> everywhere instead of <code>CustomTabNet</code> and it would work the same. Or am I missing something?</li>\n<li>be careful on the feature you are using, even if patient<em>id means something here adding this feature in the model can't work on the test set, because all patient</em>ids are completely new (the drop should be visible in your train/validation since the CV scheme separates patients) I would definitely not use patient_id as a feature.</li>\n<li>Since there are very few features (3) in metadata here, I honestly can't see how attention is going to help. So I doubt that TabNet is the best choice in this specific situation : if you add more features why not! (But again, I've not been very good at using metadata in this competition so far so…)</li>\n</ul>",
      "rawMarkdown": "Hey @tanulsingh077,\n\nThanks for sharing this and all the citings. I'm trying to add a comment directly on the notebook but it seems that I can't add the first comment, UX problem from Kaggle, so I'll leave it here.\n\n Honestly I'm a bit confused by this notebook that seems incredibly complicated for something that should be simple, I know that some simple customization are not available yet on `pytorch-tabnet` and I'll try to make things easier, but I guess some things here are over complicated.\n\nI'll leave here a few comments/questions if you don't mind, please correct me if I got something wrong:\n- all this customization is actually used for one reason: changing the loss function (and the early stopping function but tabnet early stops on AUC which seems reasonable here)? But did you know that there is a `loss_fn` parameter in TabNet's `fit` method that would allow you to switch from logloss to `SoftMarginFocalLoss` with a one liner? (Let me know if you encounter any bug when trying this - I'll try to be responsive on the repo).\n- About the CustomTabNet class, is this just to give an example on how on could use tabnet inside a bigger class? Because at the moment it looks like a perfect duplicate of `TabNet`, so you could have just used `TabNet` everywhere instead of `CustomTabNet` and it would work the same. Or am I missing something?\n- be careful on the feature you are using, even if patient_id means something here adding this feature in the model can't work on the test set, because all patient_ids are completely new (the drop should be visible in your train/validation since the CV scheme separates patients) I would definitely not use patient_id as a feature.\n- Since there are very few features (3) in metadata here, I honestly can't see how attention is going to help. So I doubt that TabNet is the best choice in this specific situation : if you add more features why not! (But again, I've not been very good at using metadata in this competition so far so...)",
      "votes": null
    },
    {
      "id": "940003",
      "postDate": "07/22/2020 16:06:05",
      "content": "<p><a href=\"/optimo\">@optimo</a> Thanks for your advices and time . I will try to answer each question here and please correct me if I am wrong :- \n* I am totally aware that loss_fn can easily take softmargin_loss , but for using that loss we need to one-hot encode the targets that's something I do in Dataset , also that's not the only thing , I needed to include balanced class sampler which cannot be added directly if I am not wrong . Another reason includes custom LR's , if someone wants to use lr which chris deotte uses in his kernel then it can be done .The plotting function added is also a result of the custom usage .Also i have done this to serve as a starting point of any more customization if anyone wants and provides a greater control over anything . I am sorry if I have over-complicated things , the point here was to provide greater control and customizations\n* CustomTabnet looks exactly like Tabnet because I didn't modify it right now but by using this we can modify even the outputs other heads , etc . Also if not using this we cannot use differential LR's etc techniques if we wanted to . Again as I said the point was not over-complication the point was to provide a way that can lead to a lot of experimentations etc\n* True , I guess I did a mistake using Patient_id thanks for the heads_up and I totally agree with such less features we can't do anything , But again the idea was to provide code for usage anywhere . Also I Have few ideas for Feature engg that I can try</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "optimo Thanks for your advices and time . I will try to answer each question here and please correct me if I am wrong :- \n* I am totally aware that loss_fn can easily take softmargin_loss , but for using that loss we need to one-hot encode the targets that's something I do in Dataset , also that's not the only thing , I needed to include balanced class sampler which cannot be added directly if I am not wrong . Another reason includes custom LR's , if someone wants to use lr which chris deotte uses in his kernel then it can be done .The plotting function added is also a result of the custom usage .Also i have done this to serve as a starting point of any more customization if anyone wants and provides a greater control over anything . I am sorry if I have over-complicated things , the point here was to provide greater control and customizations\n* CustomTabnet looks exactly like Tabnet because I didn't modify it right now but by using this we can modify even the outputs other heads , etc . Also if not using this we cannot use differential LR's etc techniques if we wanted to . Again as I said the point was not over-complication the point was to provide a way that can lead to a lot of experimentations etc\n* True , I guess I did a mistake using Patient_id thanks for the heads_up and I totally agree with such less features we can't do anything , But again the idea was to provide code for usage anywhere . Also I Have few ideas for Feature engg that I can try\n\nThanks",
      "votes": null
    },
    {
      "id": "940136",
      "postDate": "07/22/2020 17:32:55",
      "content": "<p>Just to add with <a href=\"/optimo\">@optimo</a>, there is some redundant self implementation that could easily achieve by default settings. </p>\n\n<p>In case you're more interested to use <code>TabNet</code> in this problem, We would encourage you to see <a href=\"https://www.kaggle.com/ipythonx/tf-keras-melanoma-classification-starter-tabnet/notebook#Tabular-Modeling\">here</a>. We've shown how to use <code>TabNet</code> for this problem. However, by this setup, we didn't get much than <strong>.70</strong> in LB but after incorporating some feature engineering, we've managed to achieve near <strong>.94</strong></p>",
      "rawMarkdown": "Just to add with @optimo, there is some redundant self implementation that could easily achieve by default settings. \n\nIn case you're more interested to use `TabNet` in this problem, We would encourage you to see [here](https://www.kaggle.com/ipythonx/tf-keras-melanoma-classification-starter-tabnet/notebook#Tabular-Modeling). We've shown how to use `TabNet` for this problem. However, by this setup, we didn't get much than **.70** in LB but after incorporating some feature engineering, we've managed to achieve near **.94**",
      "votes": null
    },
    {
      "id": "940210",
      "postDate": "07/22/2020 18:33:10",
      "content": "<p>Yes, great work guys!\nI’m just trying to find the pain points when using pytorch-tabnet to try to make it as complete and as easy to use as possible. But allowing flexibility while keeping things simple is quite a challenge!</p>\n\n<p>That’s why I’m so peaky with your work, how can we allow easy and simple customization for these kind of experiments?</p>\n\n<p>I guess we’ll need to discuss more in order to find the best way/trade off! Very happy to see you guys manage to use TabNet in such competitive settings!</p>\n\n<p>Keep going!</p>",
      "rawMarkdown": "Yes, great work guys!\nI’m just trying to find the pain points when using pytorch-tabnet to try to make it as complete and as easy to use as possible. But allowing flexibility while keeping things simple is quite a challenge!\n\nThat’s why I’m so peaky with your work, how can we allow easy and simple customization for these kind of experiments?\n\nI guess we’ll need to discuss more in order to find the best way/trade off! Very happy to see you guys manage to use TabNet in such competitive settings!\n\nKeep going!",
      "votes": null
    },
    {
      "id": "940268",
      "postDate": "07/22/2020 19:39:35",
      "content": "<p>You achieve 0.94 just with meta features only?? Or you incorporate the images as <a href=\"/ipythonx\">@ipythonx</a>\nI agree <a href=\"/optimo\">@optimo</a> its difficult to do both things but IMO Pytorch tabnet is awesome piece of work and if someone needs customization, complexity automatically kicks in</p>",
      "rawMarkdown": "You achieve 0.94 just with meta features only?? Or you incorporate the images as @ipythonx\nI agree @optimo its difficult to do both things but IMO Pytorch tabnet is awesome piece of work and if someone needs customization, complexity automatically kicks in",
      "votes": null
    },
    {
      "id": "941425",
      "postDate": "07/23/2020 08:04:40",
      "content": "<p><a href=\"/tanulsingh077\">@tanulsingh077</a> still thinking out loud how to make things easier with tabnet:</p>\n\n<ul>\n<li>about the plots : Tabnet is saving some information about how the training went that you can access in <code>clf.history</code> potentially anything could be stored here in order to make nice plots later, need to either add more info or find a way of allowing people to ask for specific history</li>\n<li>about the balanced class sampler : there is already one included in TabNet that you can easily access during fit method by giving either <code>weights</code>=1 for class balanced sampling or giving your own dictionary of class balances. This will perform a <code>WeightedRandomSampler</code> so if you want to use an other sampling method it will start to be difficult. I'll try to see if something could be done to allow more samplers.</li>\n</ul>\n\n<p>Not related to tabnet, <code>SoftMarginFocalLoss</code> does not have to use OHE targets right? it's just the implementation you are using? Did you have good results with this?</p>\n\n<p>I did an experiment with this implementation : \n```\ndef criterion_margin_focal_binary_cross_entropy(logit, truth):\n    \"\"\"\n    adapted from <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880993\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880993</a>\n    \"\"\"\n    weight_pos=2\n    weight_neg=1\n    gamma=2\n    margin=0.2\n    em = np.exp(margin)</p>\n\n<pre><code>logit = torch.nn.Softmax(dim=-1)(logit)[:, 1].view(-1)\ntruth = truth.view(-1)\nlog_pos = -F.logsigmoid( logit)\nlog_neg = -F.logsigmoid(-logit)\n\nlog_prob = truth*log_pos + (1-truth)*log_neg\nprob = torch.exp(-log_prob)\nmargin = torch.log(em +(1-em)*prob)\n\nweight = truth*weight_pos + (1-truth)*weight_neg\nloss = margin + weight*(1 - prob) ** gamma * log_prob\n\nloss = loss.mean()\nreturn loss\n</code></pre>\n\n<p>```</p>\n\n<p>But did not have good results (there might be something wrong).\nThis example should work directly with TabNet <code>loss_fn</code> parameter.</p>",
      "rawMarkdown": "tanulsingh077 still thinking out loud how to make things easier with tabnet:\n\n- about the plots : Tabnet is saving some information about how the training went that you can access in `clf.history` potentially anything could be stored here in order to make nice plots later, need to either add more info or find a way of allowing people to ask for specific history\n- about the balanced class sampler : there is already one included in TabNet that you can easily access during fit method by giving either `weights`=1 for class balanced sampling or giving your own dictionary of class balances. This will perform a `WeightedRandomSampler` so if you want to use an other sampling method it will start to be difficult. I'll try to see if something could be done to allow more samplers.\n\nNot related to tabnet, `SoftMarginFocalLoss` does not have to use OHE targets right? it's just the implementation you are using? Did you have good results with this?\n\nI did an experiment with this implementation : \n```\ndef criterion_margin_focal_binary_cross_entropy(logit, truth):\n    \"\"\"\n    adapted from https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880993\n    \"\"\"\n    weight_pos=2\n    weight_neg=1\n    gamma=2\n    margin=0.2\n    em = np.exp(margin)\n\n    logit = torch.nn.Softmax(dim=-1)(logit)[:, 1].view(-1)\n    truth = truth.view(-1)\n    log_pos = -F.logsigmoid( logit)\n    log_neg = -F.logsigmoid(-logit)\n\n    log_prob = truth*log_pos + (1-truth)*log_neg\n    prob = torch.exp(-log_prob)\n    margin = torch.log(em +(1-em)*prob)\n\n    weight = truth*weight_pos + (1-truth)*weight_neg\n    loss = margin + weight*(1 - prob) ** gamma * log_prob\n\n    loss = loss.mean()\n    return loss\n```\n\nBut did not have good results (there might be something wrong).\nThis example should work directly with TabNet `loss_fn` parameter.",
      "votes": null
    },
    {
      "id": "942125",
      "postDate": "07/23/2020 15:40:35",
      "content": "<p>Thanks <a href=\"/optimo\">@optimo</a> maybe definitely something can be done with clf.history , but as I said Pytorch-tabnet is awesome on its own , if someone wants to use all the customizations they can by the methods I showed or if they it simpler they can use the interface as I have shown already in a different kernel , it's just dependent about the amount of experimentation one needs\nAs for extending tabnet , maybe I can contribute to the repo if allowed and we can do it together</p>",
      "rawMarkdown": "Thanks @optimo maybe definitely something can be done with clf.history , but as I said Pytorch-tabnet is awesome on its own , if someone wants to use all the customizations they can by the methods I showed or if they it simpler they can use the interface as I have shown already in a different kernel , it's just dependent about the amount of experimentation one needs\nAs for extending tabnet , maybe I can contribute to the repo if allowed and we can do it together",
      "votes": null
    },
    {
      "id": "942142",
      "postDate": "07/23/2020 15:51:46",
      "content": "<p>Of course you are welcome to contribute! Open a PR regarding an existing issue or create a new one and let's roll!</p>",
      "rawMarkdown": "Of course you are welcome to contribute! Open a PR regarding an existing issue or create a new one and let's roll!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 939972,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "07/22/2020 15:48:45",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a>,</p>\n<p>Thanks for sharing this and all the citings. I'm trying to add a comment directly on the notebook but it seems that I can't add the first comment, UX problem from Kaggle, so I'll leave it here.</p>\n<p>Honestly I'm a bit confused by this notebook that seems incredibly complicated for something that should be simple, I know that some simple customization are not available yet on <code>pytorch-tabnet</code> and I'll try to make things easier, but I guess some things here are over complicated.</p>\n<p>I'll leave here a few comments/questions if you don't mind, please correct me if I got something wrong:</p>\n<ul>\n<li>all this customization is actually used for one reason: changing the loss function (and the early stopping function but tabnet early stops on AUC which seems reasonable here)? But did you know that there is a <code>loss_fn</code> parameter in TabNet's <code>fit</code> method that would allow you to switch from logloss to <code>SoftMarginFocalLoss</code> with a one liner? (Let me know if you encounter any bug when trying this - I'll try to be responsive on the repo).</li>\n<li>About the CustomTabNet class, is this just to give an example on how on could use tabnet inside a bigger class? Because at the moment it looks like a perfect duplicate of <code>TabNet</code>, so you could have just used <code>TabNet</code> everywhere instead of <code>CustomTabNet</code> and it would work the same. Or am I missing something?</li>\n<li>be careful on the feature you are using, even if patient<em>id means something here adding this feature in the model can't work on the test set, because all patient</em>ids are completely new (the drop should be visible in your train/validation since the CV scheme separates patients) I would definitely not use patient_id as a feature.</li>\n<li>Since there are very few features (3) in metadata here, I honestly can't see how attention is going to help. So I doubt that TabNet is the best choice in this specific situation : if you add more features why not! (But again, I've not been very good at using metadata in this competition so far so…)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 940003,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "07/22/2020 16:06:05",
          "content": "<p><a href=\"/optimo\">@optimo</a> Thanks for your advices and time . I will try to answer each question here and please correct me if I am wrong :- \n* I am totally aware that loss_fn can easily take softmargin_loss , but for using that loss we need to one-hot encode the targets that's something I do in Dataset , also that's not the only thing , I needed to include balanced class sampler which cannot be added directly if I am not wrong . Another reason includes custom LR's , if someone wants to use lr which chris deotte uses in his kernel then it can be done .The plotting function added is also a result of the custom usage .Also i have done this to serve as a starting point of any more customization if anyone wants and provides a greater control over anything . I am sorry if I have over-complicated things , the point here was to provide greater control and customizations\n* CustomTabnet looks exactly like Tabnet because I didn't modify it right now but by using this we can modify even the outputs other heads , etc . Also if not using this we cannot use differential LR's etc techniques if we wanted to . Again as I said the point was not over-complication the point was to provide a way that can lead to a lot of experimentations etc\n* True , I guess I did a mistake using Patient_id thanks for the heads_up and I totally agree with such less features we can't do anything , But again the idea was to provide code for usage anywhere . Also I Have few ideas for Feature engg that I can try</p>\n\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 940136,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "07/22/2020 17:32:55",
          "content": "<p>Just to add with <a href=\"/optimo\">@optimo</a>, there is some redundant self implementation that could easily achieve by default settings. </p>\n\n<p>In case you're more interested to use <code>TabNet</code> in this problem, We would encourage you to see <a href=\"https://www.kaggle.com/ipythonx/tf-keras-melanoma-classification-starter-tabnet/notebook#Tabular-Modeling\">here</a>. We've shown how to use <code>TabNet</code> for this problem. However, by this setup, we didn't get much than <strong>.70</strong> in LB but after incorporating some feature engineering, we've managed to achieve near <strong>.94</strong></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 940210,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "07/22/2020 18:33:10",
          "content": "<p>Yes, great work guys!\nI’m just trying to find the pain points when using pytorch-tabnet to try to make it as complete and as easy to use as possible. But allowing flexibility while keeping things simple is quite a challenge!</p>\n\n<p>That’s why I’m so peaky with your work, how can we allow easy and simple customization for these kind of experiments?</p>\n\n<p>I guess we’ll need to discuss more in order to find the best way/trade off! Very happy to see you guys manage to use TabNet in such competitive settings!</p>\n\n<p>Keep going!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 940268,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "07/22/2020 19:39:35",
          "content": "<p>You achieve 0.94 just with meta features only?? Or you incorporate the images as <a href=\"/ipythonx\">@ipythonx</a>\nI agree <a href=\"/optimo\">@optimo</a> its difficult to do both things but IMO Pytorch tabnet is awesome piece of work and if someone needs customization, complexity automatically kicks in</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 941425,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "07/23/2020 08:04:40",
          "content": "<p><a href=\"/tanulsingh077\">@tanulsingh077</a> still thinking out loud how to make things easier with tabnet:</p>\n\n<ul>\n<li>about the plots : Tabnet is saving some information about how the training went that you can access in <code>clf.history</code> potentially anything could be stored here in order to make nice plots later, need to either add more info or find a way of allowing people to ask for specific history</li>\n<li>about the balanced class sampler : there is already one included in TabNet that you can easily access during fit method by giving either <code>weights</code>=1 for class balanced sampling or giving your own dictionary of class balances. This will perform a <code>WeightedRandomSampler</code> so if you want to use an other sampling method it will start to be difficult. I'll try to see if something could be done to allow more samplers.</li>\n</ul>\n\n<p>Not related to tabnet, <code>SoftMarginFocalLoss</code> does not have to use OHE targets right? it's just the implementation you are using? Did you have good results with this?</p>\n\n<p>I did an experiment with this implementation : \n```\ndef criterion_margin_focal_binary_cross_entropy(logit, truth):\n    \"\"\"\n    adapted from <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880993\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880993</a>\n    \"\"\"\n    weight_pos=2\n    weight_neg=1\n    gamma=2\n    margin=0.2\n    em = np.exp(margin)</p>\n\n<pre><code>logit = torch.nn.Softmax(dim=-1)(logit)[:, 1].view(-1)\ntruth = truth.view(-1)\nlog_pos = -F.logsigmoid( logit)\nlog_neg = -F.logsigmoid(-logit)\n\nlog_prob = truth*log_pos + (1-truth)*log_neg\nprob = torch.exp(-log_prob)\nmargin = torch.log(em +(1-em)*prob)\n\nweight = truth*weight_pos + (1-truth)*weight_neg\nloss = margin + weight*(1 - prob) ** gamma * log_prob\n\nloss = loss.mean()\nreturn loss\n</code></pre>\n\n<p>```</p>\n\n<p>But did not have good results (there might be something wrong).\nThis example should work directly with TabNet <code>loss_fn</code> parameter.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 942125,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "07/23/2020 15:40:35",
          "content": "<p>Thanks <a href=\"/optimo\">@optimo</a> maybe definitely something can be done with clf.history , but as I said Pytorch-tabnet is awesome on its own , if someone wants to use all the customizations they can by the methods I showed or if they it simpler they can use the interface as I have shown already in a different kernel , it's just dependent about the amount of experimentation one needs\nAs for extending tabnet , maybe I can contribute to the repo if allowed and we can do it together</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 942142,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "07/23/2020 15:51:46",
          "content": "<p>Of course you are welcome to contribute! Open a PR regarding an existing issue or create a new one and let's roll!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "939786": "In this competition we have been given two types of data , one is the images of skin lesions of patients , other is the tabular meta data . Now there are three ways of combining these two information together :-\n\n* Build a CNN image model and find a way to input the tabular data into the CNN image model\n* Build a Tabular data model and find a way to extract image embeddings or image features and input into the Tabular data model\n* Build 2 separate models for images and metadata and ensemble\n\nWe have tried all three and the third option works the best and gives significant amount of boost.Another question which comes to mind is, what models can we use for modelling with tabular data , my notebook tries to answer this question. It can be found here https://www.kaggle.com/tanulsingh077/exploring-limits-of-meta-features-using-tabnet\n\nI present how to use Tabnet as a custom Model instead of the scikit-learn type interface provided by Pytorch-Tabnet . I show how anyone can use Tabnet just like a torchvision or any torch-hub model for any downstream task . My notebook is written keeping this specific melanoma competition in mind but the code can be easily used for your own tasks by changing the dataloader\n\nBelow is the resulting Graph of one fold training which is done without any fine-tuning/feature engg . It easily gets 0.68 valid_auc without much efforts , I have been able to get to 0.715 local auc , still have to check lb , I have shared it with the commnunity so that it's full potential can be reached\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779944%2Fa6a518ffb74c6fb512d377fdcd0e0348%2Ftab.PNG?generation=1595423398069884&amp;alt=media)\n\n\nI hope you find my efforts useful, Thanks for reading",
    "939972": "Hey @tanulsingh077,\n\nThanks for sharing this and all the citings. I'm trying to add a comment directly on the notebook but it seems that I can't add the first comment, UX problem from Kaggle, so I'll leave it here.\n\n Honestly I'm a bit confused by this notebook that seems incredibly complicated for something that should be simple, I know that some simple customization are not available yet on `pytorch-tabnet` and I'll try to make things easier, but I guess some things here are over complicated.\n\nI'll leave here a few comments/questions if you don't mind, please correct me if I got something wrong:\n- all this customization is actually used for one reason: changing the loss function (and the early stopping function but tabnet early stops on AUC which seems reasonable here)? But did you know that there is a `loss_fn` parameter in TabNet's `fit` method that would allow you to switch from logloss to `SoftMarginFocalLoss` with a one liner? (Let me know if you encounter any bug when trying this - I'll try to be responsive on the repo).\n- About the CustomTabNet class, is this just to give an example on how on could use tabnet inside a bigger class? Because at the moment it looks like a perfect duplicate of `TabNet`, so you could have just used `TabNet` everywhere instead of `CustomTabNet` and it would work the same. Or am I missing something?\n- be careful on the feature you are using, even if patient_id means something here adding this feature in the model can't work on the test set, because all patient_ids are completely new (the drop should be visible in your train/validation since the CV scheme separates patients) I would definitely not use patient_id as a feature.\n- Since there are very few features (3) in metadata here, I honestly can't see how attention is going to help. So I doubt that TabNet is the best choice in this specific situation : if you add more features why not! (But again, I've not been very good at using metadata in this competition so far so...)",
    "940003": "optimo Thanks for your advices and time . I will try to answer each question here and please correct me if I am wrong :- \n* I am totally aware that loss_fn can easily take softmargin_loss , but for using that loss we need to one-hot encode the targets that's something I do in Dataset , also that's not the only thing , I needed to include balanced class sampler which cannot be added directly if I am not wrong . Another reason includes custom LR's , if someone wants to use lr which chris deotte uses in his kernel then it can be done .The plotting function added is also a result of the custom usage .Also i have done this to serve as a starting point of any more customization if anyone wants and provides a greater control over anything . I am sorry if I have over-complicated things , the point here was to provide greater control and customizations\n* CustomTabnet looks exactly like Tabnet because I didn't modify it right now but by using this we can modify even the outputs other heads , etc . Also if not using this we cannot use differential LR's etc techniques if we wanted to . Again as I said the point was not over-complication the point was to provide a way that can lead to a lot of experimentations etc\n* True , I guess I did a mistake using Patient_id thanks for the heads_up and I totally agree with such less features we can't do anything , But again the idea was to provide code for usage anywhere . Also I Have few ideas for Feature engg that I can try\n\nThanks",
    "940136": "Just to add with @optimo, there is some redundant self implementation that could easily achieve by default settings. \n\nIn case you're more interested to use `TabNet` in this problem, We would encourage you to see [here](https://www.kaggle.com/ipythonx/tf-keras-melanoma-classification-starter-tabnet/notebook#Tabular-Modeling). We've shown how to use `TabNet` for this problem. However, by this setup, we didn't get much than **.70** in LB but after incorporating some feature engineering, we've managed to achieve near **.94**",
    "940210": "Yes, great work guys!\nI’m just trying to find the pain points when using pytorch-tabnet to try to make it as complete and as easy to use as possible. But allowing flexibility while keeping things simple is quite a challenge!\n\nThat’s why I’m so peaky with your work, how can we allow easy and simple customization for these kind of experiments?\n\nI guess we’ll need to discuss more in order to find the best way/trade off! Very happy to see you guys manage to use TabNet in such competitive settings!\n\nKeep going!",
    "940268": "You achieve 0.94 just with meta features only?? Or you incorporate the images as @ipythonx\nI agree @optimo its difficult to do both things but IMO Pytorch tabnet is awesome piece of work and if someone needs customization, complexity automatically kicks in",
    "941425": "tanulsingh077 still thinking out loud how to make things easier with tabnet:\n\n- about the plots : Tabnet is saving some information about how the training went that you can access in `clf.history` potentially anything could be stored here in order to make nice plots later, need to either add more info or find a way of allowing people to ask for specific history\n- about the balanced class sampler : there is already one included in TabNet that you can easily access during fit method by giving either `weights`=1 for class balanced sampling or giving your own dictionary of class balances. This will perform a `WeightedRandomSampler` so if you want to use an other sampling method it will start to be difficult. I'll try to see if something could be done to allow more samplers.\n\nNot related to tabnet, `SoftMarginFocalLoss` does not have to use OHE targets right? it's just the implementation you are using? Did you have good results with this?\n\nI did an experiment with this implementation : \n```\ndef criterion_margin_focal_binary_cross_entropy(logit, truth):\n    \"\"\"\n    adapted from https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880993\n    \"\"\"\n    weight_pos=2\n    weight_neg=1\n    gamma=2\n    margin=0.2\n    em = np.exp(margin)\n\n    logit = torch.nn.Softmax(dim=-1)(logit)[:, 1].view(-1)\n    truth = truth.view(-1)\n    log_pos = -F.logsigmoid( logit)\n    log_neg = -F.logsigmoid(-logit)\n\n    log_prob = truth*log_pos + (1-truth)*log_neg\n    prob = torch.exp(-log_prob)\n    margin = torch.log(em +(1-em)*prob)\n\n    weight = truth*weight_pos + (1-truth)*weight_neg\n    loss = margin + weight*(1 - prob) ** gamma * log_prob\n\n    loss = loss.mean()\n    return loss\n```\n\nBut did not have good results (there might be something wrong).\nThis example should work directly with TabNet `loss_fn` parameter.",
    "942125": "Thanks @optimo maybe definitely something can be done with clf.history , but as I said Pytorch-tabnet is awesome on its own , if someone wants to use all the customizations they can by the methods I showed or if they it simpler they can use the interface as I have shown already in a different kernel , it's just dependent about the amount of experimentation one needs\nAs for extending tabnet , maybe I can contribute to the repo if allowed and we can do it together",
    "942142": "Of course you are welcome to contribute! Open a PR regarding an existing issue or create a new one and let's roll!"
  },
  "source": "meta"
}