{
  "id": 156284,
  "title": "Dixon's Q-test to identify \"Ugly duckling\" outliers",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156284",
  "author_name": "",
  "post_date": "2020-06-05T10:08:09.917063800Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Please see the post by @hengck23 for understanding <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\">Ugly Duckling Concept</a></p>\n\n<p>Since <strong>outliers</strong> among images of a given patient are more likely to be Malignant Melanomas, but similar looking lesions in another patient may not be malignant, we can improve accuracy of predictions by considering the <strong>context</strong> (i.e. patient_id).</p>\n\n<p>However, the <strong>numbers per patient</strong> are very small. Please see below contextual analysis:\n   2056 unique patient_id in train data 33126 images\n   428 patients have malignant melanoma, out of which\n     - 319 have only 1 malignant lesion and no benign images\n     - 109 have both types (benign and malignant)\n   These 109 patient images with both types:\n     - have 5 to 78 total images (min, max) per patient\n     - where 2 to 8 are malignant AND\n     - another 3 to 76 are benign images per patient</p>\n\n<p>For such <strong>small</strong> groups of data, we can use <a href=\"https://www.statisticshowto.com/dixons-q-test/#%3a~%3atext=Dixon%27s%20Q%20test%2C%20\">Dixon's Q test</a> (see <a href=\"https://en.wikipedia.org/wiki/Dixon%27s_Q_test\">wiki</a> here), where\n   <code>Q = abs(gap) / range</code>\ni.e. after obtaining <strong>sorted</strong> model predictions, we can group them by patient_id, look for <strong>gaps</strong> (in the sorted list) and identify <strong>outliers</strong>.</p>",
  "messages": [
    {
      "id": "874808",
      "postDate": "06/05/2020 10:08:09",
      "content": "<p>Please see the post by @hengck23 for understanding <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348\">Ugly Duckling Concept</a></p>\n\n<p>Since <strong>outliers</strong> among images of a given patient are more likely to be Malignant Melanomas, but similar looking lesions in another patient may not be malignant, we can improve accuracy of predictions by considering the <strong>context</strong> (i.e. patient_id).</p>\n\n<p>However, the <strong>numbers per patient</strong> are very small. Please see below contextual analysis:\n   2056 unique patient_id in train data 33126 images\n   428 patients have malignant melanoma, out of which\n     - 319 have only 1 malignant lesion and no benign images\n     - 109 have both types (benign and malignant)\n   These 109 patient images with both types:\n     - have 5 to 78 total images (min, max) per patient\n     - where 2 to 8 are malignant AND\n     - another 3 to 76 are benign images per patient</p>\n\n<p>For such <strong>small</strong> groups of data, we can use <a href=\"https://www.statisticshowto.com/dixons-q-test/#%3a~%3atext=Dixon%27s%20Q%20test%2C%20\">Dixon's Q test</a> (see <a href=\"https://en.wikipedia.org/wiki/Dixon%27s_Q_test\">wiki</a> here), where\n   <code>Q = abs(gap) / range</code>\ni.e. after obtaining <strong>sorted</strong> model predictions, we can group them by patient_id, look for <strong>gaps</strong> (in the sorted list) and identify <strong>outliers</strong>.</p>",
      "rawMarkdown": "Please see the post by @hengck23 for understanding [Ugly Duckling Concept](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348)\n\nSince **outliers** among images of a given patient are more likely to be Malignant Melanomas, but similar looking lesions in another patient may not be malignant, we can improve accuracy of predictions by considering the **context** (i.e. patient_id).\n\nHowever, the **numbers per patient** are very small. Please see below contextual analysis:\n   2056 unique patient_id in train data 33126 images\n   428 patients have malignant melanoma, out of which\n     - 319 have only 1 malignant lesion and no benign images\n     - 109 have both types (benign and malignant)\n   These 109 patient images with both types:\n     - have 5 to 78 total images (min, max) per patient\n     - where 2 to 8 are malignant AND\n     - another 3 to 76 are benign images per patient\n\nFor such **small** groups of data, we can use [Dixon's Q test](https://www.statisticshowto.com/dixons-q-test/#:~:text=Dixon's%20Q%20test%2C%20) (see [wiki](https://en.wikipedia.org/wiki/Dixon%27s_Q_test) here), where\n   `Q = abs(gap) / range`\ni.e. after obtaining **sorted** model predictions, we can group them by patient_id, look for **gaps** (in the sorted list) and identify **outliers**.",
      "votes": null
    },
    {
      "id": "885971",
      "postDate": "06/14/2020 15:55:17",
      "content": "<p>Thank you for the pointers, your observations definitely lead us in the right direction.</p>\n\n<p>As a sort of sanity check, however, I wanted to point out that I am getting different numbers than you: out of those 2056 patients, I get 428 patients with positive (malignant melanoma) samples, but <strong>427</strong> with both positive and negative samples. In other words, I find there is only one patient with <em>just positive</em> sample(s).</p>\n\n<p>I was wondering if you wouldn't mind sharing your logic. Here is mine:</p>\n\n<p><code>\npatient_ids = sorted(data.patient_id.unique())\ndata_pat_tar_gb = data.groupby(['patient_id', 'target'])\ndata_pat_tar_dict = data_pat_tar_gb.groups\npatients_pos = [p for p in patient_ids if (p, 1) in data_pat_tar_dict]\npatients_neg = [p for p in patient_ids if (p, 0) in data_pat_tar_dict]\npatients_both = set(patients_pos).intersection(patients_neg)\n</code></p>",
      "rawMarkdown": "Thank you for the pointers, your observations definitely lead us in the right direction.\n\nAs a sort of sanity check, however, I wanted to point out that I am getting different numbers than you: out of those 2056 patients, I get 428 patients with positive (malignant melanoma) samples, but **427** with both positive and negative samples. In other words, I find there is only one patient with *just positive* sample(s).\n\nI was wondering if you wouldn't mind sharing your logic. Here is mine:\n\n```\npatient_ids = sorted(data.patient_id.unique())\ndata_pat_tar_gb = data.groupby(['patient_id', 'target'])\ndata_pat_tar_dict = data_pat_tar_gb.groups\npatients_pos = [p for p in patient_ids if (p, 1) in data_pat_tar_dict]\npatients_neg = [p for p in patient_ids if (p, 0) in data_pat_tar_dict]\npatients_both = set(patients_pos).intersection(patients_neg)\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 885971,
      "author_name": "vitala",
      "author_url": "",
      "post_date": "06/14/2020 15:55:17",
      "content": "<p>Thank you for the pointers, your observations definitely lead us in the right direction.</p>\n\n<p>As a sort of sanity check, however, I wanted to point out that I am getting different numbers than you: out of those 2056 patients, I get 428 patients with positive (malignant melanoma) samples, but <strong>427</strong> with both positive and negative samples. In other words, I find there is only one patient with <em>just positive</em> sample(s).</p>\n\n<p>I was wondering if you wouldn't mind sharing your logic. Here is mine:</p>\n\n<p><code>\npatient_ids = sorted(data.patient_id.unique())\ndata_pat_tar_gb = data.groupby(['patient_id', 'target'])\ndata_pat_tar_dict = data_pat_tar_gb.groups\npatients_pos = [p for p in patient_ids if (p, 1) in data_pat_tar_dict]\npatients_neg = [p for p in patient_ids if (p, 0) in data_pat_tar_dict]\npatients_both = set(patients_pos).intersection(patients_neg)\n</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "874808": "Please see the post by @hengck23 for understanding [Ugly Duckling Concept](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155348)\n\nSince **outliers** among images of a given patient are more likely to be Malignant Melanomas, but similar looking lesions in another patient may not be malignant, we can improve accuracy of predictions by considering the **context** (i.e. patient_id).\n\nHowever, the **numbers per patient** are very small. Please see below contextual analysis:\n   2056 unique patient_id in train data 33126 images\n   428 patients have malignant melanoma, out of which\n     - 319 have only 1 malignant lesion and no benign images\n     - 109 have both types (benign and malignant)\n   These 109 patient images with both types:\n     - have 5 to 78 total images (min, max) per patient\n     - where 2 to 8 are malignant AND\n     - another 3 to 76 are benign images per patient\n\nFor such **small** groups of data, we can use [Dixon's Q test](https://www.statisticshowto.com/dixons-q-test/#:~:text=Dixon's%20Q%20test%2C%20) (see [wiki](https://en.wikipedia.org/wiki/Dixon%27s_Q_test) here), where\n   `Q = abs(gap) / range`\ni.e. after obtaining **sorted** model predictions, we can group them by patient_id, look for **gaps** (in the sorted list) and identify **outliers**.",
    "885971": "Thank you for the pointers, your observations definitely lead us in the right direction.\n\nAs a sort of sanity check, however, I wanted to point out that I am getting different numbers than you: out of those 2056 patients, I get 428 patients with positive (malignant melanoma) samples, but **427** with both positive and negative samples. In other words, I find there is only one patient with *just positive* sample(s).\n\nI was wondering if you wouldn't mind sharing your logic. Here is mine:\n\n```\npatient_ids = sorted(data.patient_id.unique())\ndata_pat_tar_gb = data.groupby(['patient_id', 'target'])\ndata_pat_tar_dict = data_pat_tar_gb.groups\npatients_pos = [p for p in patient_ids if (p, 1) in data_pat_tar_dict]\npatients_neg = [p for p in patient_ids if (p, 0) in data_pat_tar_dict]\npatients_both = set(patients_pos).intersection(patients_neg)\n```"
  },
  "source": "meta"
}