{
  "id": 155885,
  "title": "Feature engineering on metadata ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/155885",
  "author_name": "",
  "post_date": "2020-06-03T12:14:39.696179900Z",
  "votes": 23,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Tried to beat Giba's simple metadata-only baseline (<a href=\"https://www.kaggle.com/titericz/simple-baseline\">https://www.kaggle.com/titericz/simple-baseline</a>) with ML approach. So tried some straightforward feature engineering and tested that on the leaderboard:</p>\n\n<p><code>features = ['sex','age_approx','anatom_site_general_challenge']</code>\nvalidation: 0.66 LB:0.67</p>\n\n<p><code>features = ['sex','age_approx','anatom_site_general_challenge','age_min','age_max']</code>\nvalidation: 0.69 LB:0.64</p>\n\n<p><code>features = ['sex','age_approx','anatom_site_general_challenge','img_count','img_age_combo_count','img_anatomy_combo_count','img_age_anatomy_combo_count','age_count','age_min','age_max','anatomy_count']</code>\nvalidation: 0.77 LB:0.66</p>\n\n<p>It seems that train and test is different enough to stay away from any time-dependent feature engineering.</p>",
  "messages": [
    {
      "id": "872670",
      "postDate": "06/03/2020 12:14:39",
      "content": "<p>Tried to beat Giba's simple metadata-only baseline (<a href=\"https://www.kaggle.com/titericz/simple-baseline\">https://www.kaggle.com/titericz/simple-baseline</a>) with ML approach. So tried some straightforward feature engineering and tested that on the leaderboard:</p>\n\n<p><code>features = ['sex','age_approx','anatom_site_general_challenge']</code>\nvalidation: 0.66 LB:0.67</p>\n\n<p><code>features = ['sex','age_approx','anatom_site_general_challenge','age_min','age_max']</code>\nvalidation: 0.69 LB:0.64</p>\n\n<p><code>features = ['sex','age_approx','anatom_site_general_challenge','img_count','img_age_combo_count','img_anatomy_combo_count','img_age_anatomy_combo_count','age_count','age_min','age_max','anatomy_count']</code>\nvalidation: 0.77 LB:0.66</p>\n\n<p>It seems that train and test is different enough to stay away from any time-dependent feature engineering.</p>",
      "rawMarkdown": "Tried to beat Giba's simple metadata-only baseline (https://www.kaggle.com/titericz/simple-baseline) with ML approach. So tried some straightforward feature engineering and tested that on the leaderboard:\n\n`features = ['sex','age_approx','anatom_site_general_challenge']`\nvalidation: 0.66 LB:0.67\n\n\n`features = ['sex','age_approx','anatom_site_general_challenge','age_min','age_max']`\nvalidation: 0.69 LB:0.64\n\n\n`features = ['sex','age_approx','anatom_site_general_challenge','img_count','img_age_combo_count','img_anatomy_combo_count','img_age_anatomy_combo_count','age_count','age_min','age_max','anatomy_count']`\nvalidation: 0.77 LB:0.66\n\nIt seems that train and test is different enough to stay away from any time-dependent feature engineering.",
      "votes": null
    },
    {
      "id": "873191",
      "postDate": "06/03/2020 22:17:29",
      "content": "<p>Raddar, have you included <code>height</code> and <code>width</code> of original image as meta feature? I have not yet but was wondering if that had any correlation with target.</p>",
      "rawMarkdown": "Raddar, have you included `height` and `width` of original image as meta feature? I have not yet but was wondering if that had any correlation with target.",
      "votes": null
    },
    {
      "id": "873197",
      "postDate": "06/03/2020 22:22:30",
      "content": "<p>Also are there any other meta features hiding around? Sometimes images have meta features stored in their compression or perhaps the DICOM files have more meta. It would be nice to expose all meta features now.</p>",
      "rawMarkdown": "Also are there any other meta features hiding around? Sometimes images have meta features stored in their compression or perhaps the DICOM files have more meta. It would be nice to expose all meta features now.",
      "votes": null
    },
    {
      "id": "873202",
      "postDate": "06/03/2020 22:33:06",
      "content": "<p>I can imagine that features like 'is_one_malign_for patient_id' or 'no_of_malign_for patient_id' could do something</p>",
      "rawMarkdown": "I can imagine that features like 'is_one_malign_for patient_id' or 'no_of_malign_for patient_id' could do something",
      "votes": null
    },
    {
      "id": "873230",
      "postDate": "06/04/2020 00:05:04",
      "content": "<p>This information is not available for test set as the split was based on patient_id</p>",
      "rawMarkdown": "This information is not available for test set as the split was based on patient_id",
      "votes": null
    },
    {
      "id": "873231",
      "postDate": "06/04/2020 00:06:29",
      "content": "<p>I haven't tested yet as I checked if my CNN models output correlates with mean target based on dimension categories - they were pretty close</p>",
      "rawMarkdown": "I haven't tested yet as I checked if my CNN models output correlates with mean target based on dimension categories - they were pretty close",
      "votes": null
    },
    {
      "id": "873246",
      "postDate": "06/04/2020 00:43:08",
      "content": "<p>does this matter? I mean I am talking about derived features in the training data or I don't understand someting..</p>",
      "rawMarkdown": "does this matter? I mean I am talking about derived features in the training data or I don't understand someting..",
      "votes": null
    },
    {
      "id": "873563",
      "postDate": "06/04/2020 09:14:38",
      "content": "<p>in order to reuse the model with its all features you have in a training set, you need to have same features in a test set. and for test set we do not have \"malignant\" features (because that is something we are trying to predict in the first place!)</p>",
      "rawMarkdown": "in order to reuse the model with its all features you have in a training set, you need to have same features in a test set. and for test set we do not have \"malignant\" features (because that is something we are trying to predict in the first place!)",
      "votes": null
    },
    {
      "id": "873670",
      "postDate": "06/04/2020 10:41:25",
      "content": "<p>ok, I do understand now.\nmaybe one can pseudo label the test data and look if that does something</p>",
      "rawMarkdown": "ok, I do understand now.\nmaybe one can pseudo label the test data and look if that does something",
      "votes": null
    },
    {
      "id": "873852",
      "postDate": "06/04/2020 13:32:32",
      "content": "<p>I also tried using more features and observed something very similar... Got 0.76 CV and 0.63 LB. What is interesting, my features give only 0.53 AUC on the adversarial validation test. </p>",
      "rawMarkdown": "I also tried using more features and observed something very similar... Got 0.76 CV and 0.63 LB. What is interesting, my features give only 0.53 AUC on the adversarial validation test.",
      "votes": null
    },
    {
      "id": "875971",
      "postDate": "06/06/2020 10:24:32",
      "content": "<p>I'm still working on getting a stable CV vs LB. Since I started in the competition with the tabular data, the image count per patient was one of my first features. The difference between CV and LB is very big for this feature.\n-&gt; features = ['PicCntOfPatient']\n   LB = 0.584 and CV = 0.719 (using Grouped 3-Fold, i.e. grouped by patient_id)</p>\n\n<p>-&gt; features = ['age_approx', 'anatom_site_general_challenge', 'sex']\n   LB = 0.670 and CV = 0.682 (also using Grouped 3-Fold, i.e. grouped by patient_id)</p>\n\n<p><code>\n    df_train[\"PicCntOfPatient\"] = df_train[\"patient_id\"].map(df_train.groupby(\"patient_id\").size())\n    df_test[\"PicCntOfPatient\"] = df_test[\"patient_id\"].map(df_test.groupby(\"patient_id\").size())\n</code>\nWhy is the 'PicCntOfPatient' so strong on train and useless for test?</p>",
      "rawMarkdown": "I'm still working on getting a stable CV vs LB. Since I started in the competition with the tabular data, the image count per patient was one of my first features. The difference between CV and LB is very big for this feature.\n-&gt; features = ['PicCntOfPatient']\n   LB = 0.584 and CV = 0.719 (using Grouped 3-Fold, i.e. grouped by patient_id)\n\n-&gt; features = ['age_approx', 'anatom_site_general_challenge', 'sex']\n   LB = 0.670 and CV = 0.682 (also using Grouped 3-Fold, i.e. grouped by patient_id)\n\n```\n    df_train[\"PicCntOfPatient\"] = df_train[\"patient_id\"].map(df_train.groupby(\"patient_id\").size())\n    df_test[\"PicCntOfPatient\"] = df_test[\"patient_id\"].map(df_test.groupby(\"patient_id\").size())\n```\nWhy is the 'PicCntOfPatient' so strong on train and useless for test?",
      "votes": null
    },
    {
      "id": "876093",
      "postDate": "06/06/2020 12:36:58",
      "content": "<p>because the test is small.</p>",
      "rawMarkdown": "because the test is small.",
      "votes": null
    },
    {
      "id": "876141",
      "postDate": "06/06/2020 13:26:48",
      "content": "<p><a href=\"/agentauers\">@agentauers</a> there is no answer for that yet. There could be a lot of things - maybe organizers made test set a little different to train set. Maybe test set patients are from different clinic, etc..</p>",
      "rawMarkdown": "agentauers there is no answer for that yet. There could be a lot of things - maybe organizers made test set a little different to train set. Maybe test set patients are from different clinic, etc..",
      "votes": null
    },
    {
      "id": "876206",
      "postDate": "06/06/2020 14:27:47",
      "content": "<p>I was wondering how image height and width can help us.They are not deciding factors for melanoma.Has anyone tested it?</p>",
      "rawMarkdown": "I was wondering how image height and width can help us.They are not deciding factors for melanoma.Has anyone tested it?",
      "votes": null
    },
    {
      "id": "877036",
      "postDate": "06/07/2020 09:10:03",
      "content": "<p>I tried this : <code>features = ['sex', 'age_approx', 'anatom_site_general_challenge', 'img_count', 'age_count', 'age_min', 'age_max', 'anatomy_count', 'age_diff']</code></p>\n\n<p>CV = 0.794, LB = 0.623 :/</p>\n\n<p>Seems like test set don't have any interest in feature engineering. 😄 </p>",
      "rawMarkdown": "I tried this : `features = ['sex', 'age_approx', 'anatom_site_general_challenge', 'img_count', 'age_count', 'age_min', 'age_max', 'anatomy_count', 'age_diff']`\n\nCV = 0.794, LB = 0.623 :/\n\nSeems like test set don't have any interest in feature engineering. 😄",
      "votes": null
    },
    {
      "id": "895784",
      "postDate": "06/21/2020 15:50:29",
      "content": "<p>How is the metadata mining and modeling going on so far? I tried some feature engineering and different models, (XGBoost, LGBM) some of them score 0.70 - 0.75 AUC in LB but they all do worst than <a href=\"https://www.kaggle.com/titericz/simple-baseline\">giba's 0.70 AUC notebook</a> when I ensemble them with the CNN. </p>",
      "rawMarkdown": "How is the metadata mining and modeling going on so far? I tried some feature engineering and different models, (XGBoost, LGBM) some of them score 0.70 - 0.75 AUC in LB but they all do worst than [giba's 0.70 AUC notebook](https://www.kaggle.com/titericz/simple-baseline) when I ensemble them with the CNN.",
      "votes": null
    },
    {
      "id": "896704",
      "postDate": "06/22/2020 11:44:19",
      "content": "<p>One feature that presumably does correlate with melanoma is lesion size, which we are not given directly.  I would think that resizing all images to a standard width and height for processing throws away some scale information that could be retrieved, to a limited extent, from the original image dimensions.\nDoes this make sense?</p>",
      "rawMarkdown": "One feature that presumably does correlate with melanoma is lesion size, which we are not given directly.  I would think that resizing all images to a standard width and height for processing throws away some scale information that could be retrieved, to a limited extent, from the original image dimensions.\nDoes this make sense?",
      "votes": null
    },
    {
      "id": "949968",
      "postDate": "07/29/2020 05:16:00",
      "content": "<p>That's a revelation...😃 </p>",
      "rawMarkdown": "That's a revelation...😃",
      "votes": null
    },
    {
      "id": "951075",
      "postDate": "07/29/2020 21:18:46",
      "content": "<p>Well, <a href=\"/romanweilguny\">@romanweilguny</a> , if I was you I would try to avoid that.</p>\n\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n\n<p>ref: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\">https://www.kaggle.com/c/siim-isic-melanoma-classification/rules</a></p>",
      "rawMarkdown": "Well, @romanweilguny , if I was you I would try to avoid that.\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nref: https://www.kaggle.com/c/siim-isic-melanoma-classification/rules",
      "votes": null
    },
    {
      "id": "951077",
      "postDate": "07/29/2020 21:20:22",
      "content": "<p>How do you plan to retrieve this information? A higher dimension image could just be a closer look at the melanoma, hence might not be indicative of a larger (or smaller) lesion.\n<a href=\"/dslate\">@dslate</a> </p>",
      "rawMarkdown": "How do you plan to retrieve this information? A higher dimension image could just be a closer look at the melanoma, hence might not be indicative of a larger (or smaller) lesion.\n@dslate",
      "votes": null
    },
    {
      "id": "951132",
      "postDate": "07/29/2020 22:55:39",
      "content": "<p>As far as I know - pseudo labelling is not considered as hand labelling or human prediction</p>",
      "rawMarkdown": "As far as I know - pseudo labelling is not considered as hand labelling or human prediction",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 873191,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/03/2020 22:17:29",
      "content": "<p>Raddar, have you included <code>height</code> and <code>width</code> of original image as meta feature? I have not yet but was wondering if that had any correlation with target.</p>",
      "votes": null,
      "replies": [
        {
          "id": 873197,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/03/2020 22:22:30",
          "content": "<p>Also are there any other meta features hiding around? Sometimes images have meta features stored in their compression or perhaps the DICOM files have more meta. It would be nice to expose all meta features now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 873231,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/04/2020 00:06:29",
          "content": "<p>I haven't tested yet as I checked if my CNN models output correlates with mean target based on dimension categories - they were pretty close</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 876206,
          "author_name": "",
          "author_url": "",
          "post_date": "06/06/2020 14:27:47",
          "content": "<p>I was wondering how image height and width can help us.They are not deciding factors for melanoma.Has anyone tested it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 896704,
          "author_name": "dslate",
          "author_url": "",
          "post_date": "06/22/2020 11:44:19",
          "content": "<p>One feature that presumably does correlate with melanoma is lesion size, which we are not given directly.  I would think that resizing all images to a standard width and height for processing throws away some scale information that could be retrieved, to a limited extent, from the original image dimensions.\nDoes this make sense?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951077,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "07/29/2020 21:20:22",
          "content": "<p>How do you plan to retrieve this information? A higher dimension image could just be a closer look at the melanoma, hence might not be indicative of a larger (or smaller) lesion.\n<a href=\"/dslate\">@dslate</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 873202,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "06/03/2020 22:33:06",
      "content": "<p>I can imagine that features like 'is_one_malign_for patient_id' or 'no_of_malign_for patient_id' could do something</p>",
      "votes": null,
      "replies": [
        {
          "id": 873230,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/04/2020 00:05:04",
          "content": "<p>This information is not available for test set as the split was based on patient_id</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 873246,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "06/04/2020 00:43:08",
          "content": "<p>does this matter? I mean I am talking about derived features in the training data or I don't understand someting..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 873563,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/04/2020 09:14:38",
          "content": "<p>in order to reuse the model with its all features you have in a training set, you need to have same features in a test set. and for test set we do not have \"malignant\" features (because that is something we are trying to predict in the first place!)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 873670,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "06/04/2020 10:41:25",
          "content": "<p>ok, I do understand now.\nmaybe one can pseudo label the test data and look if that does something</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951075,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "07/29/2020 21:18:46",
          "content": "<p>Well, <a href=\"/romanweilguny\">@romanweilguny</a> , if I was you I would try to avoid that.</p>\n\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n\n<p>ref: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\">https://www.kaggle.com/c/siim-isic-melanoma-classification/rules</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951132,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "07/29/2020 22:55:39",
          "content": "<p>As far as I know - pseudo labelling is not considered as hand labelling or human prediction</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 873852,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "06/04/2020 13:32:32",
      "content": "<p>I also tried using more features and observed something very similar... Got 0.76 CV and 0.63 LB. What is interesting, my features give only 0.53 AUC on the adversarial validation test. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 875971,
      "author_name": "agentauers",
      "author_url": "",
      "post_date": "06/06/2020 10:24:32",
      "content": "<p>I'm still working on getting a stable CV vs LB. Since I started in the competition with the tabular data, the image count per patient was one of my first features. The difference between CV and LB is very big for this feature.\n-&gt; features = ['PicCntOfPatient']\n   LB = 0.584 and CV = 0.719 (using Grouped 3-Fold, i.e. grouped by patient_id)</p>\n\n<p>-&gt; features = ['age_approx', 'anatom_site_general_challenge', 'sex']\n   LB = 0.670 and CV = 0.682 (also using Grouped 3-Fold, i.e. grouped by patient_id)</p>\n\n<p><code>\n    df_train[\"PicCntOfPatient\"] = df_train[\"patient_id\"].map(df_train.groupby(\"patient_id\").size())\n    df_test[\"PicCntOfPatient\"] = df_test[\"patient_id\"].map(df_test.groupby(\"patient_id\").size())\n</code>\nWhy is the 'PicCntOfPatient' so strong on train and useless for test?</p>",
      "votes": null,
      "replies": [
        {
          "id": 876093,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "06/06/2020 12:36:58",
          "content": "<p>because the test is small.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 876141,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/06/2020 13:26:48",
          "content": "<p><a href=\"/agentauers\">@agentauers</a> there is no answer for that yet. There could be a lot of things - maybe organizers made test set a little different to train set. Maybe test set patients are from different clinic, etc..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 877036,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "06/07/2020 09:10:03",
      "content": "<p>I tried this : <code>features = ['sex', 'age_approx', 'anatom_site_general_challenge', 'img_count', 'age_count', 'age_min', 'age_max', 'anatomy_count', 'age_diff']</code></p>\n\n<p>CV = 0.794, LB = 0.623 :/</p>\n\n<p>Seems like test set don't have any interest in feature engineering. 😄 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 895784,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "06/21/2020 15:50:29",
      "content": "<p>How is the metadata mining and modeling going on so far? I tried some feature engineering and different models, (XGBoost, LGBM) some of them score 0.70 - 0.75 AUC in LB but they all do worst than <a href=\"https://www.kaggle.com/titericz/simple-baseline\">giba's 0.70 AUC notebook</a> when I ensemble them with the CNN. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 949968,
      "author_name": "jaseemck",
      "author_url": "",
      "post_date": "07/29/2020 05:16:00",
      "content": "<p>That's a revelation...😃 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "872670": "Tried to beat Giba's simple metadata-only baseline (https://www.kaggle.com/titericz/simple-baseline) with ML approach. So tried some straightforward feature engineering and tested that on the leaderboard:\n\n`features = ['sex','age_approx','anatom_site_general_challenge']`\nvalidation: 0.66 LB:0.67\n\n\n`features = ['sex','age_approx','anatom_site_general_challenge','age_min','age_max']`\nvalidation: 0.69 LB:0.64\n\n\n`features = ['sex','age_approx','anatom_site_general_challenge','img_count','img_age_combo_count','img_anatomy_combo_count','img_age_anatomy_combo_count','age_count','age_min','age_max','anatomy_count']`\nvalidation: 0.77 LB:0.66\n\nIt seems that train and test is different enough to stay away from any time-dependent feature engineering.",
    "873191": "Raddar, have you included `height` and `width` of original image as meta feature? I have not yet but was wondering if that had any correlation with target.",
    "873197": "Also are there any other meta features hiding around? Sometimes images have meta features stored in their compression or perhaps the DICOM files have more meta. It would be nice to expose all meta features now.",
    "873202": "I can imagine that features like 'is_one_malign_for patient_id' or 'no_of_malign_for patient_id' could do something",
    "873230": "This information is not available for test set as the split was based on patient_id",
    "873231": "I haven't tested yet as I checked if my CNN models output correlates with mean target based on dimension categories - they were pretty close",
    "873246": "does this matter? I mean I am talking about derived features in the training data or I don't understand someting..",
    "873563": "in order to reuse the model with its all features you have in a training set, you need to have same features in a test set. and for test set we do not have \"malignant\" features (because that is something we are trying to predict in the first place!)",
    "873670": "ok, I do understand now.\nmaybe one can pseudo label the test data and look if that does something",
    "873852": "I also tried using more features and observed something very similar... Got 0.76 CV and 0.63 LB. What is interesting, my features give only 0.53 AUC on the adversarial validation test.",
    "875971": "I'm still working on getting a stable CV vs LB. Since I started in the competition with the tabular data, the image count per patient was one of my first features. The difference between CV and LB is very big for this feature.\n-&gt; features = ['PicCntOfPatient']\n   LB = 0.584 and CV = 0.719 (using Grouped 3-Fold, i.e. grouped by patient_id)\n\n-&gt; features = ['age_approx', 'anatom_site_general_challenge', 'sex']\n   LB = 0.670 and CV = 0.682 (also using Grouped 3-Fold, i.e. grouped by patient_id)\n\n```\n    df_train[\"PicCntOfPatient\"] = df_train[\"patient_id\"].map(df_train.groupby(\"patient_id\").size())\n    df_test[\"PicCntOfPatient\"] = df_test[\"patient_id\"].map(df_test.groupby(\"patient_id\").size())\n```\nWhy is the 'PicCntOfPatient' so strong on train and useless for test?",
    "876093": "because the test is small.",
    "876141": "agentauers there is no answer for that yet. There could be a lot of things - maybe organizers made test set a little different to train set. Maybe test set patients are from different clinic, etc..",
    "876206": "I was wondering how image height and width can help us.They are not deciding factors for melanoma.Has anyone tested it?",
    "877036": "I tried this : `features = ['sex', 'age_approx', 'anatom_site_general_challenge', 'img_count', 'age_count', 'age_min', 'age_max', 'anatomy_count', 'age_diff']`\n\nCV = 0.794, LB = 0.623 :/\n\nSeems like test set don't have any interest in feature engineering. 😄",
    "895784": "How is the metadata mining and modeling going on so far? I tried some feature engineering and different models, (XGBoost, LGBM) some of them score 0.70 - 0.75 AUC in LB but they all do worst than [giba's 0.70 AUC notebook](https://www.kaggle.com/titericz/simple-baseline) when I ensemble them with the CNN.",
    "896704": "One feature that presumably does correlate with melanoma is lesion size, which we are not given directly.  I would think that resizing all images to a standard width and height for processing throws away some scale information that could be retrieved, to a limited extent, from the original image dimensions.\nDoes this make sense?",
    "949968": "That's a revelation...😃",
    "951075": "Well, @romanweilguny , if I was you I would try to avoid that.\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nref: https://www.kaggle.com/c/siim-isic-melanoma-classification/rules",
    "951077": "How do you plan to retrieve this information? A higher dimension image could just be a closer look at the melanoma, hence might not be indicative of a larger (or smaller) lesion.\n@dslate",
    "951132": "As far as I know - pseudo labelling is not considered as hand labelling or human prediction"
  },
  "source": "meta"
}