{
  "id": 154906,
  "title": "66% chance of correct classifying based only on patient info",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154906",
  "author_name": "",
  "post_date": "2020-05-30T11:33:51.161780300Z",
  "votes": 13,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone !</p>\n\n<p>I have published <a href=\"https://www.kaggle.com/louise2001/beginner-lightgbm-on-patient-only-information\">this kernel</a> that makes prediction <strong>based only on patient-level information</strong> (no image processing !)</p>\n\n<p>It can be a good baseline to add to your image processing models.</p>\n\n<p>Note that with only patient-level information, we still achieve an AUC of 0.66, which of course isn't top leaderboard, but is interesting insofar as we have 2 chances out of three of correctly classifying a melanoma based only on patient information...</p>\n\n<p>Hope you find it interesting. Happy kaggling to all of you !</p>",
  "messages": [
    {
      "id": "867513",
      "postDate": "05/30/2020 11:33:51",
      "content": "<p>Hi everyone !</p>\n\n<p>I have published <a href=\"https://www.kaggle.com/louise2001/beginner-lightgbm-on-patient-only-information\">this kernel</a> that makes prediction <strong>based only on patient-level information</strong> (no image processing !)</p>\n\n<p>It can be a good baseline to add to your image processing models.</p>\n\n<p>Note that with only patient-level information, we still achieve an AUC of 0.66, which of course isn't top leaderboard, but is interesting insofar as we have 2 chances out of three of correctly classifying a melanoma based only on patient information...</p>\n\n<p>Hope you find it interesting. Happy kaggling to all of you !</p>",
      "rawMarkdown": "Hi everyone !\n\nI have published [this kernel](https://www.kaggle.com/louise2001/beginner-lightgbm-on-patient-only-information) that makes prediction **based only on patient-level information** (no image processing !)\n\nIt can be a good baseline to add to your image processing models.\n\nNote that with only patient-level information, we still achieve an AUC of 0.66, which of course isn't top leaderboard, but is interesting insofar as we have 2 chances out of three of correctly classifying a melanoma based only on patient information...\n\nHope you find it interesting. Happy kaggling to all of you !",
      "votes": null
    },
    {
      "id": "867624",
      "postDate": "05/30/2020 13:17:21",
      "content": "<p>You did a good analysis, thank you for the insight. But your seemingly arbitrary usage of terms is disturbing. </p>\n\n<ol>\n<li>First, when you say \"we have 2 chances out of three\" it seems you are talking in accuracy terms. But accuracy is barely related to AUC. For example, by submitting all zeros I will get accuracy 0.98 (as only about 0.02 are melanomas), but AUC is 0.5</li>\n<li>In your kernel you say that you prefer to optimize \"Mean Average Precision\", but the latest version (version 5) of the kernel is actually optimizes log-loss. It is specified in objective equals binary. The different metric and feval parameters that you specify are used only for early stopping. </li>\n</ol>",
      "rawMarkdown": "You did a good analysis, thank you for the insight. But your seemingly arbitrary usage of terms is disturbing. \n\n1. First, when you say \"we have 2 chances out of three\" it seems you are talking in accuracy terms. But accuracy is barely related to AUC. For example, by submitting all zeros I will get accuracy 0.98 (as only about 0.02 are melanomas), but AUC is 0.5\n2. In your kernel you say that you prefer to optimize \"Mean Average Precision\", but the latest version (version 5) of the kernel is actually optimizes log-loss. It is specified in objective equals binary. The different metric and feval parameters that you specify are used only for early stopping.",
      "votes": null
    },
    {
      "id": "867667",
      "postDate": "05/30/2020 13:59:39",
      "content": "<p>Yes, you are right. If you take a look at the precision and recall I get on 1 values, you can also see that it is nearly always 0. That means that the model is indeed predicting nearly always 0 : I am not at all saying it is a good model, on the contrary. </p>\n\n<p>Also, when I say two chances out of 3 of correctly classifying, I should indeed be more precise : it is rather that between a benign and a malignant melanoma, the model has 2 chances out of 3 of correctly putting a higher danger score to the malignant one. You must excuse my approximation : there isn't much place in a post's title 😉</p>",
      "rawMarkdown": "Yes, you are right. If you take a look at the precision and recall I get on 1 values, you can also see that it is nearly always 0. That means that the model is indeed predicting nearly always 0 : I am not at all saying it is a good model, on the contrary. \n\nAlso, when I say two chances out of 3 of correctly classifying, I should indeed be more precise : it is rather that between a benign and a malignant melanoma, the model has 2 chances out of 3 of correctly putting a higher danger score to the malignant one. You must excuse my approximation : there isn't much place in a post's title 😉",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 867624,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "05/30/2020 13:17:21",
      "content": "<p>You did a good analysis, thank you for the insight. But your seemingly arbitrary usage of terms is disturbing. </p>\n\n<ol>\n<li>First, when you say \"we have 2 chances out of three\" it seems you are talking in accuracy terms. But accuracy is barely related to AUC. For example, by submitting all zeros I will get accuracy 0.98 (as only about 0.02 are melanomas), but AUC is 0.5</li>\n<li>In your kernel you say that you prefer to optimize \"Mean Average Precision\", but the latest version (version 5) of the kernel is actually optimizes log-loss. It is specified in objective equals binary. The different metric and feval parameters that you specify are used only for early stopping. </li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 867667,
          "author_name": "louise2001",
          "author_url": "",
          "post_date": "05/30/2020 13:59:39",
          "content": "<p>Yes, you are right. If you take a look at the precision and recall I get on 1 values, you can also see that it is nearly always 0. That means that the model is indeed predicting nearly always 0 : I am not at all saying it is a good model, on the contrary. </p>\n\n<p>Also, when I say two chances out of 3 of correctly classifying, I should indeed be more precise : it is rather that between a benign and a malignant melanoma, the model has 2 chances out of 3 of correctly putting a higher danger score to the malignant one. You must excuse my approximation : there isn't much place in a post's title 😉</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "867513": "Hi everyone !\n\nI have published [this kernel](https://www.kaggle.com/louise2001/beginner-lightgbm-on-patient-only-information) that makes prediction **based only on patient-level information** (no image processing !)\n\nIt can be a good baseline to add to your image processing models.\n\nNote that with only patient-level information, we still achieve an AUC of 0.66, which of course isn't top leaderboard, but is interesting insofar as we have 2 chances out of three of correctly classifying a melanoma based only on patient information...\n\nHope you find it interesting. Happy kaggling to all of you !",
    "867624": "You did a good analysis, thank you for the insight. But your seemingly arbitrary usage of terms is disturbing. \n\n1. First, when you say \"we have 2 chances out of three\" it seems you are talking in accuracy terms. But accuracy is barely related to AUC. For example, by submitting all zeros I will get accuracy 0.98 (as only about 0.02 are melanomas), but AUC is 0.5\n2. In your kernel you say that you prefer to optimize \"Mean Average Precision\", but the latest version (version 5) of the kernel is actually optimizes log-loss. It is specified in objective equals binary. The different metric and feval parameters that you specify are used only for early stopping.",
    "867667": "Yes, you are right. If you take a look at the precision and recall I get on 1 values, you can also see that it is nearly always 0. That means that the model is indeed predicting nearly always 0 : I am not at all saying it is a good model, on the contrary. \n\nAlso, when I say two chances out of 3 of correctly classifying, I should indeed be more precise : it is rather that between a benign and a malignant melanoma, the model has 2 chances out of 3 of correctly putting a higher danger score to the malignant one. You must excuse my approximation : there isn't much place in a post's title 😉"
  },
  "source": "meta"
}