{
  "id": 175467,
  "title": "With Patient context",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175467",
  "author_name": "",
  "post_date": "2020-08-18T09:11:15.960001700Z",
  "votes": 5,
  "comment_count": 10,
  "views": 0,
  "content": "<p>As described in the competition prizes tab there are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175313\" target=\"_blank\">now prizes up for grabs</a> for the top-scoring model making use of patient-level contextual information:</p>\n<blockquote>\n  <p>Special Prizes: Awarded to the top scoring models using or not using patient-level contextual information.</p>\n  <p>With Context - $5,000 (Top-scoring model making use of patient-level contextual information)<br>\n  Without Context - $5,000 (Top-scoring model without using any patient-level contextual information)</p>\n</blockquote>\n<p>From the competition description, this looks like one of the main motivations of the SIIM &amp; ISIC 2020  challenge. </p>\n<p>I'm curious, did people make use of patient level contextual information? I know Chris discussed clustering and I <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168456\" target=\"_blank\">used this</a> to show you could use patient embeddings to build a classifier, but so far it doesn't look like top competitors used patient level contextual information. Does anyone have anything to add?</p>",
  "messages": [
    {
      "id": "975375",
      "postDate": "08/18/2020 09:11:15",
      "content": "<p>As described in the competition prizes tab there are <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175313\" target=\"_blank\">now prizes up for grabs</a> for the top-scoring model making use of patient-level contextual information:</p>\n<blockquote>\n  <p>Special Prizes: Awarded to the top scoring models using or not using patient-level contextual information.</p>\n  <p>With Context - $5,000 (Top-scoring model making use of patient-level contextual information)<br>\n  Without Context - $5,000 (Top-scoring model without using any patient-level contextual information)</p>\n</blockquote>\n<p>From the competition description, this looks like one of the main motivations of the SIIM &amp; ISIC 2020  challenge. </p>\n<p>I'm curious, did people make use of patient level contextual information? I know Chris discussed clustering and I <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168456\" target=\"_blank\">used this</a> to show you could use patient embeddings to build a classifier, but so far it doesn't look like top competitors used patient level contextual information. Does anyone have anything to add?</p>",
      "rawMarkdown": "As described in the competition prizes tab there are [now prizes up for grabs](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175313) for the top-scoring model making use of patient-level contextual information:\n\n> Special Prizes: Awarded to the top scoring models using or not using patient-level contextual information.\n\n> With Context - $5,000 (Top-scoring model making use of patient-level contextual information)\nWithout Context - $5,000 (Top-scoring model without using any patient-level contextual information)\n\nFrom the competition description, this looks like one of the main motivations of the SIIM & ISIC 2020  challenge. \n\nI'm curious, did people make use of patient level contextual information? I know Chris discussed clustering and I [used this](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168456) to show you could use patient embeddings to build a classifier, but so far it doesn't look like top competitors used patient level contextual information. Does anyone have anything to add?",
      "votes": null
    },
    {
      "id": "975383",
      "postDate": "08/18/2020 09:14:33",
      "content": "<p>We haven't used it. We had some wild idea of generating embeddings with efficientnet for each of a patient's images and then using some graph convolutional layers (which can deal with a \"variable amount of features\") to combine these images and the metadata together in a smart way. As such, the model would incorporate the other images of the patient as well. </p>\n<p>Unforatuntely, the idea seemed to wild and too much work to actually try it out :p.</p>",
      "rawMarkdown": "We haven't used it. We had some wild idea of generating embeddings with efficientnet for each of a patient's images and then using some graph convolutional layers (which can deal with a \"variable amount of features\") to combine these images and the metadata together in a smart way. As such, the model would incorporate the other images of the patient as well. \n\nUnforatuntely, the idea seemed to wild and too much work to actually try it out :p.",
      "votes": null
    },
    {
      "id": "975392",
      "postDate": "08/18/2020 09:18:26",
      "content": "<p>Thanks Gilles. I also thought using patient context seemed like a gamble when it was difficult to exhaust all 'standard' techniques in the length of the competition. </p>\n<p>Maybe someone will share something concrete. I'm now wishing I submitted my classifier which used patient-level context on the off-chance no one else did!</p>",
      "rawMarkdown": "Thanks Gilles. I also thought using patient context seemed like a gamble when it was difficult to exhaust all 'standard' techniques in the length of the competition. \n\nMaybe someone will share something concrete. I'm now wishing I submitted my classifier which used patient-level context on the off-chance no one else did!",
      "votes": null
    },
    {
      "id": "975397",
      "postDate": "08/18/2020 09:23:12",
      "content": "<p>Here's a simple trick that worked in CV for us that somewhat uses \"patient-level\" information. I would not recommend this trick in production/clinical setting though…</p>\n<p>If a patient has many images (let's say over 50), then in the training set, only a maximum of 8 of these images could be positive. If we assume that these positive images will be within the top 50% of the predictions of a patient, then we can put the predictions of the other 50% of images on 0. This boosted our CV with 0.002. We never made a submission with it though since these crude assumptions may not have held on the private LB.</p>",
      "rawMarkdown": "Here's a simple trick that worked in CV for us that somewhat uses \"patient-level\" information. I would not recommend this trick in production/clinical setting though...\n\nIf a patient has many images (let's say over 50), then in the training set, only a maximum of 8 of these images could be positive. If we assume that these positive images will be within the top 50% of the predictions of a patient, then we can put the predictions of the other 50% of images on 0. This boosted our CV with 0.002. We never made a submission with it though since these crude assumptions may not have held on the private LB.",
      "votes": null
    },
    {
      "id": "975403",
      "postDate": "08/18/2020 09:28:07",
      "content": "<p>We tried a somewhat similar approach: after grouping by <code>patient_id</code>, multiply predictions outside of the top-10 predictions with the highest probability by <code>k &lt; 1</code>, where <code>k</code> can be optimized on CV. We managed to get a boost on CV but none of the experiments translated to the LB improvement.</p>",
      "rawMarkdown": "We tried a somewhat similar approach: after grouping by `patient_id`, multiply predictions outside of the top-10 predictions with the highest probability by `k < 1`, where `k` can be optimized on CV. We managed to get a boost on CV but none of the experiments translated to the LB improvement.",
      "votes": null
    },
    {
      "id": "975406",
      "postDate": "08/18/2020 09:28:40",
      "content": "<p>Yes a risky assumption to make in a clinical setting! Effectively gaming the metric, I think the impact on the AUC is determined by the number of false negatives (for patients with low number of images). It seems reasonable that your CV score would translate to private LB if your model is making the 'same' kind of mistakes, but agree its a weird trick!</p>",
      "rawMarkdown": "Yes a risky assumption to make in a clinical setting! Effectively gaming the metric, I think the impact on the AUC is determined by the number of false negatives (for patients with low number of images). It seems reasonable that your CV score would translate to private LB if your model is making the 'same' kind of mistakes, but agree its a weird trick!",
      "votes": null
    },
    {
      "id": "975409",
      "postDate": "08/18/2020 09:30:29",
      "content": "<p>I tried early on to put the average embeddings of a patient into a successive fit, but also saw no boosts, also any PP on groupby patient_id did not really work. It only worked in LGB models if you for example group by patient and then do aggregates, but we did not use it in the end.</p>",
      "rawMarkdown": "I tried early on to put the average embeddings of a patient into a successive fit, but also saw no boosts, also any PP on groupby patient_id did not really work. It only worked in LGB models if you for example group by patient and then do aggregates, but we did not use it in the end.",
      "votes": null
    },
    {
      "id": "977693",
      "postDate": "08/19/2020 16:43:21",
      "content": "<p>Exactly, we can make use of patient specific information to extract deep semantics and this could help to build 100 percent effective detection based on individual patient features! Did any one made use of patient centric features to build a segmentation model or local features information to isolate diseased region of the skin! </p>",
      "rawMarkdown": "Exactly, we can make use of patient specific information to extract deep semantics and this could help to build 100 percent effective detection based on individual patient features! Did any one made use of patient centric features to build a segmentation model or local features information to isolate diseased region of the skin!",
      "votes": null
    },
    {
      "id": "977744",
      "postDate": "08/19/2020 17:22:53",
      "content": "<p>One of my final sub was with this: compute oof predictions of my best blend, then rank predicted score by patient in descending order, and clip the rank to 4.  Then use this ans an additional feature to each individual model oof.  Goal was to have the ensembling model (logistic regression or xgboost) learn that the number of melanoma doe snot increase with the number of images per patient.  It improved my CV by 0.003 and I submitted with it.  Given I entered very late I  didn't had enough subs to submit without it.</p>\n<p>I checked with some late subs to see if this could have boosted other, previous ensemble subs I had, and it did not.    Jury is still out on the merit of this idea.</p>",
      "rawMarkdown": "One of my final sub was with this: compute oof predictions of my best blend, then rank predicted score by patient in descending order, and clip the rank to 4.  Then use this ans an additional feature to each individual model oof.  Goal was to have the ensembling model (logistic regression or xgboost) learn that the number of melanoma doe snot increase with the number of images per patient.  It improved my CV by 0.003 and I submitted with it.  Given I entered very late I  didn't had enough subs to submit without it.\n\nI checked with some late subs to see if this could have boosted other, previous ensemble subs I had, and it did not.    Jury is still out on the merit of this idea.",
      "votes": null
    },
    {
      "id": "977780",
      "postDate": "08/19/2020 17:44:53",
      "content": "<p>The number of images per patient is predictive. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fb4ccd25b3bd603d39757294a890a2cc6%2FScreen%20Shot%202020-08-19%20at%2010.34.16%20AM.png?generation=1597858492272494&amp;alt=media\" alt=\"\"></p>\n<p>The AUC of this single feature in train is 0.7207 </p>\n<pre><code>train['patient_ct'] = -1 * train.groupby('patient_id').patient_id.transform('count')\nprint( roc_auc_score(train.target,train.patient_ct) )\n# THIS PRINTS 0.7207\n</code></pre>\n<p>And this single feature scores private LB 0.6146, public LB 0.5852</p>\n<pre><code>test['patient_ct'] = -1 * test.groupby('patient_id').patient_id.transform('count')\ntest[['image_name','patient_ct']].rename({'patient_ct':'target'},axis=1)\\\n        .to_csv('submission_patient_ct.csv',index=False)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Ffd8ef98ecbbf92616c613f8e7f60d605%2Fplb.png?generation=1597859153888453&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The number of images per patient is predictive. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fb4ccd25b3bd603d39757294a890a2cc6%2FScreen%20Shot%202020-08-19%20at%2010.34.16%20AM.png?generation=1597858492272494&alt=media)\n\nThe AUC of this single feature in train is 0.7207 \n\n    train['patient_ct'] = -1 * train.groupby('patient_id').patient_id.transform('count')\n    print( roc_auc_score(train.target,train.patient_ct) )\n    # THIS PRINTS 0.7207\n\nAnd this single feature scores private LB 0.6146, public LB 0.5852\n\n    test['patient_ct'] = -1 * test.groupby('patient_id').patient_id.transform('count')\n    test[['image_name','patient_ct']].rename({'patient_ct':'target'},axis=1)\\\n            .to_csv('submission_patient_ct.csv',index=False)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Ffd8ef98ecbbf92616c613f8e7f60d605%2Fplb.png?generation=1597859153888453&alt=media)",
      "votes": null
    },
    {
      "id": "977843",
      "postDate": "08/19/2020 18:35:43",
      "content": "<p>I started with that, but it is extremely weak and did not help my ensemble.  I ended up with the rank I describe in my other comment.  It has a roc_auc of 0.83.</p>",
      "rawMarkdown": "I started with that, but it is extremely weak and did not help my ensemble.  I ended up with the rank I describe in my other comment.  It has a roc_auc of 0.83.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 975383,
      "author_name": "group16",
      "author_url": "",
      "post_date": "08/18/2020 09:14:33",
      "content": "<p>We haven't used it. We had some wild idea of generating embeddings with efficientnet for each of a patient's images and then using some graph convolutional layers (which can deal with a \"variable amount of features\") to combine these images and the metadata together in a smart way. As such, the model would incorporate the other images of the patient as well. </p>\n<p>Unforatuntely, the idea seemed to wild and too much work to actually try it out :p.</p>",
      "votes": null,
      "replies": [
        {
          "id": 975392,
          "author_name": "fchmiel",
          "author_url": "",
          "post_date": "08/18/2020 09:18:26",
          "content": "<p>Thanks Gilles. I also thought using patient context seemed like a gamble when it was difficult to exhaust all 'standard' techniques in the length of the competition. </p>\n<p>Maybe someone will share something concrete. I'm now wishing I submitted my classifier which used patient-level context on the off-chance no one else did!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 975397,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/18/2020 09:23:12",
          "content": "<p>Here's a simple trick that worked in CV for us that somewhat uses \"patient-level\" information. I would not recommend this trick in production/clinical setting though…</p>\n<p>If a patient has many images (let's say over 50), then in the training set, only a maximum of 8 of these images could be positive. If we assume that these positive images will be within the top 50% of the predictions of a patient, then we can put the predictions of the other 50% of images on 0. This boosted our CV with 0.002. We never made a submission with it though since these crude assumptions may not have held on the private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 975403,
          "author_name": "kozodoi",
          "author_url": "",
          "post_date": "08/18/2020 09:28:07",
          "content": "<p>We tried a somewhat similar approach: after grouping by <code>patient_id</code>, multiply predictions outside of the top-10 predictions with the highest probability by <code>k &lt; 1</code>, where <code>k</code> can be optimized on CV. We managed to get a boost on CV but none of the experiments translated to the LB improvement.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 975406,
          "author_name": "fchmiel",
          "author_url": "",
          "post_date": "08/18/2020 09:28:40",
          "content": "<p>Yes a risky assumption to make in a clinical setting! Effectively gaming the metric, I think the impact on the AUC is determined by the number of false negatives (for patients with low number of images). It seems reasonable that your CV score would translate to private LB if your model is making the 'same' kind of mistakes, but agree its a weird trick!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 975409,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/18/2020 09:30:29",
          "content": "<p>I tried early on to put the average embeddings of a patient into a successive fit, but also saw no boosts, also any PP on groupby patient_id did not really work. It only worked in LGB models if you for example group by patient and then do aggregates, but we did not use it in the end.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 977693,
      "author_name": "loveall",
      "author_url": "",
      "post_date": "08/19/2020 16:43:21",
      "content": "<p>Exactly, we can make use of patient specific information to extract deep semantics and this could help to build 100 percent effective detection based on individual patient features! Did any one made use of patient centric features to build a segmentation model or local features information to isolate diseased region of the skin! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 977744,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/19/2020 17:22:53",
      "content": "<p>One of my final sub was with this: compute oof predictions of my best blend, then rank predicted score by patient in descending order, and clip the rank to 4.  Then use this ans an additional feature to each individual model oof.  Goal was to have the ensembling model (logistic regression or xgboost) learn that the number of melanoma doe snot increase with the number of images per patient.  It improved my CV by 0.003 and I submitted with it.  Given I entered very late I  didn't had enough subs to submit without it.</p>\n<p>I checked with some late subs to see if this could have boosted other, previous ensemble subs I had, and it did not.    Jury is still out on the merit of this idea.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 977780,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/19/2020 17:44:53",
      "content": "<p>The number of images per patient is predictive. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fb4ccd25b3bd603d39757294a890a2cc6%2FScreen%20Shot%202020-08-19%20at%2010.34.16%20AM.png?generation=1597858492272494&amp;alt=media\" alt=\"\"></p>\n<p>The AUC of this single feature in train is 0.7207 </p>\n<pre><code>train['patient_ct'] = -1 * train.groupby('patient_id').patient_id.transform('count')\nprint( roc_auc_score(train.target,train.patient_ct) )\n# THIS PRINTS 0.7207\n</code></pre>\n<p>And this single feature scores private LB 0.6146, public LB 0.5852</p>\n<pre><code>test['patient_ct'] = -1 * test.groupby('patient_id').patient_id.transform('count')\ntest[['image_name','patient_ct']].rename({'patient_ct':'target'},axis=1)\\\n        .to_csv('submission_patient_ct.csv',index=False)\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Ffd8ef98ecbbf92616c613f8e7f60d605%2Fplb.png?generation=1597859153888453&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 977843,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/19/2020 18:35:43",
          "content": "<p>I started with that, but it is extremely weak and did not help my ensemble.  I ended up with the rank I describe in my other comment.  It has a roc_auc of 0.83.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "975375": "As described in the competition prizes tab there are [now prizes up for grabs](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/175313) for the top-scoring model making use of patient-level contextual information:\n\n> Special Prizes: Awarded to the top scoring models using or not using patient-level contextual information.\n\n> With Context - $5,000 (Top-scoring model making use of patient-level contextual information)\nWithout Context - $5,000 (Top-scoring model without using any patient-level contextual information)\n\nFrom the competition description, this looks like one of the main motivations of the SIIM & ISIC 2020  challenge. \n\nI'm curious, did people make use of patient level contextual information? I know Chris discussed clustering and I [used this](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168456) to show you could use patient embeddings to build a classifier, but so far it doesn't look like top competitors used patient level contextual information. Does anyone have anything to add?",
    "975383": "We haven't used it. We had some wild idea of generating embeddings with efficientnet for each of a patient's images and then using some graph convolutional layers (which can deal with a \"variable amount of features\") to combine these images and the metadata together in a smart way. As such, the model would incorporate the other images of the patient as well. \n\nUnforatuntely, the idea seemed to wild and too much work to actually try it out :p.",
    "975392": "Thanks Gilles. I also thought using patient context seemed like a gamble when it was difficult to exhaust all 'standard' techniques in the length of the competition. \n\nMaybe someone will share something concrete. I'm now wishing I submitted my classifier which used patient-level context on the off-chance no one else did!",
    "975397": "Here's a simple trick that worked in CV for us that somewhat uses \"patient-level\" information. I would not recommend this trick in production/clinical setting though...\n\nIf a patient has many images (let's say over 50), then in the training set, only a maximum of 8 of these images could be positive. If we assume that these positive images will be within the top 50% of the predictions of a patient, then we can put the predictions of the other 50% of images on 0. This boosted our CV with 0.002. We never made a submission with it though since these crude assumptions may not have held on the private LB.",
    "975403": "We tried a somewhat similar approach: after grouping by `patient_id`, multiply predictions outside of the top-10 predictions with the highest probability by `k < 1`, where `k` can be optimized on CV. We managed to get a boost on CV but none of the experiments translated to the LB improvement.",
    "975406": "Yes a risky assumption to make in a clinical setting! Effectively gaming the metric, I think the impact on the AUC is determined by the number of false negatives (for patients with low number of images). It seems reasonable that your CV score would translate to private LB if your model is making the 'same' kind of mistakes, but agree its a weird trick!",
    "975409": "I tried early on to put the average embeddings of a patient into a successive fit, but also saw no boosts, also any PP on groupby patient_id did not really work. It only worked in LGB models if you for example group by patient and then do aggregates, but we did not use it in the end.",
    "977693": "Exactly, we can make use of patient specific information to extract deep semantics and this could help to build 100 percent effective detection based on individual patient features! Did any one made use of patient centric features to build a segmentation model or local features information to isolate diseased region of the skin!",
    "977744": "One of my final sub was with this: compute oof predictions of my best blend, then rank predicted score by patient in descending order, and clip the rank to 4.  Then use this ans an additional feature to each individual model oof.  Goal was to have the ensembling model (logistic regression or xgboost) learn that the number of melanoma doe snot increase with the number of images per patient.  It improved my CV by 0.003 and I submitted with it.  Given I entered very late I  didn't had enough subs to submit without it.\n\nI checked with some late subs to see if this could have boosted other, previous ensemble subs I had, and it did not.    Jury is still out on the merit of this idea.",
    "977780": "The number of images per patient is predictive. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fb4ccd25b3bd603d39757294a890a2cc6%2FScreen%20Shot%202020-08-19%20at%2010.34.16%20AM.png?generation=1597858492272494&alt=media)\n\nThe AUC of this single feature in train is 0.7207 \n\n    train['patient_ct'] = -1 * train.groupby('patient_id').patient_id.transform('count')\n    print( roc_auc_score(train.target,train.patient_ct) )\n    # THIS PRINTS 0.7207\n\nAnd this single feature scores private LB 0.6146, public LB 0.5852\n\n    test['patient_ct'] = -1 * test.groupby('patient_id').patient_id.transform('count')\n    test[['image_name','patient_ct']].rename({'patient_ct':'target'},axis=1)\\\n            .to_csv('submission_patient_ct.csv',index=False)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Ffd8ef98ecbbf92616c613f8e7f60d605%2Fplb.png?generation=1597859153888453&alt=media)",
    "977843": "I started with that, but it is extremely weak and did not help my ensemble.  I ended up with the rank I describe in my other comment.  It has a roc_auc of 0.83."
  },
  "source": "meta"
}