{
  "id": 169340,
  "title": "Predictions Scatter Plot",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/169340",
  "author_name": "",
  "post_date": "2020-07-23T15:21:08.944097300Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am using a custom BCELossWithLogits which upweights the loss of each sample by a costant factor and a model made of an efficientnet-b3 for the images and a simple MLP for the tabular data. </p>\n\n<p>I trained the model for 1 epoch using the folds 1-4 and fold 0 for validation from <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">melanoma-merged-external-data-512x512-jpeg</a>.</p>\n\n<p>To analyze the feature vectors that I get from the efficientnet I used PCA for dimensionality reduction and I was expecting to get two well defined clusters with a few misplaced predictions, but this is what I get.</p>\n\n<p>Predictions.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F7718372830d0e2e29ae4c98e886c8d95%2Fscatter_pred.png?generation=1595514602930125&amp;alt=media\" alt=\"\"></p>\n\n<p>True labels.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F3b3828b98f9609c2bcd830177154784d%2Fscatter_target.png?generation=1595514670804409&amp;alt=media\" alt=\"\"></p>\n\n<p>Does anyone else get something like this with different loss functions? Given that the malign samples are so closely clustered, do you think the model is overfitting???</p>\n\n<p>PS. This one gets a validation AUC of 0.93523 and a LB score of 0.9345</p>",
  "messages": [
    {
      "id": "942093",
      "postDate": "07/23/2020 15:21:08",
      "content": "<p>I am using a custom BCELossWithLogits which upweights the loss of each sample by a costant factor and a model made of an efficientnet-b3 for the images and a simple MLP for the tabular data. </p>\n\n<p>I trained the model for 1 epoch using the folds 1-4 and fold 0 for validation from <a href=\"https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg\">melanoma-merged-external-data-512x512-jpeg</a>.</p>\n\n<p>To analyze the feature vectors that I get from the efficientnet I used PCA for dimensionality reduction and I was expecting to get two well defined clusters with a few misplaced predictions, but this is what I get.</p>\n\n<p>Predictions.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F7718372830d0e2e29ae4c98e886c8d95%2Fscatter_pred.png?generation=1595514602930125&amp;alt=media\" alt=\"\"></p>\n\n<p>True labels.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F3b3828b98f9609c2bcd830177154784d%2Fscatter_target.png?generation=1595514670804409&amp;alt=media\" alt=\"\"></p>\n\n<p>Does anyone else get something like this with different loss functions? Given that the malign samples are so closely clustered, do you think the model is overfitting???</p>\n\n<p>PS. This one gets a validation AUC of 0.93523 and a LB score of 0.9345</p>",
      "rawMarkdown": "I am using a custom BCELossWithLogits which upweights the loss of each sample by a costant factor and a model made of an efficientnet-b3 for the images and a simple MLP for the tabular data. \n\nI trained the model for 1 epoch using the folds 1-4 and fold 0 for validation from [melanoma-merged-external-data-512x512-jpeg](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg).\n\nTo analyze the feature vectors that I get from the efficientnet I used PCA for dimensionality reduction and I was expecting to get two well defined clusters with a few misplaced predictions, but this is what I get.\n\nPredictions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F7718372830d0e2e29ae4c98e886c8d95%2Fscatter_pred.png?generation=1595514602930125&amp;alt=media)\n\nTrue labels.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F3b3828b98f9609c2bcd830177154784d%2Fscatter_target.png?generation=1595514670804409&amp;alt=media)\n\nDoes anyone else get something like this with different loss functions? Given that the malign samples are so closely clustered, do you think the model is overfitting???\n\nPS. This one gets a validation AUC of 0.93523 and a LB score of 0.9345",
      "votes": null
    },
    {
      "id": "942694",
      "postDate": "07/24/2020 00:01:43",
      "content": "<p>Cool plot Paolo. When i get time, i will do these plots on my models. It looks like your model is doing a good job of separating malignant.</p>\n\n<p>In the bottom plot, if you plot all the <code>target=0</code> blue points first and then all the <code>target=1</code> yellow points second, we could better see how much they overlap. It looks like there is more overlap than is indicated. (And using a smaller dot, may help distinguish density).</p>",
      "rawMarkdown": "Cool plot Paolo. When i get time, i will do these plots on my models. It looks like your model is doing a good job of separating malignant.\n\nIn the bottom plot, if you plot all the `target=0` blue points first and then all the `target=1` yellow points second, we could better see how much they overlap. It looks like there is more overlap than is indicated. (And using a smaller dot, may help distinguish density).",
      "votes": null
    },
    {
      "id": "943551",
      "postDate": "07/24/2020 12:33:50",
      "content": "<p>Hey <a href=\"/meraxes10\">@meraxes10</a> \nI think you shouldn't be surprised that PCA could get most of the variance in 1 component, as the last layer on the BCE is just a simple linear combination of the feature vector (assuming you used the last layer?) </p>\n\n<p>What's the explained variance of your 1st PCA component?</p>",
      "rawMarkdown": "Hey @meraxes10 \nI think you shouldn't be surprised that PCA could get most of the variance in 1 component, as the last layer on the BCE is just a simple linear combination of the feature vector (assuming you used the last layer?) \n\nWhat's the explained variance of your 1st PCA component?",
      "votes": null
    },
    {
      "id": "944689",
      "postDate": "07/25/2020 08:56:36",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F261ca3dc7ecaad67d8854ba7fc0c7de9%2Fscatter_1.png?generation=1595667120673991&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F9f593b9f5aba02c944a2383e47de4cc7%2Fscatter_0.png?generation=1595667122005563&amp;alt=media\" alt=\"\"></p>\n\n<p>Thank you for the comment. There is too much overlap. In the outlined region there are 1587 benign (out of 10396) and 712 malignant (out of 1048)</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F261ca3dc7ecaad67d8854ba7fc0c7de9%2Fscatter_1.png?generation=1595667120673991&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F9f593b9f5aba02c944a2383e47de4cc7%2Fscatter_0.png?generation=1595667122005563&amp;alt=media)\n\nThank you for the comment. There is too much overlap. In the outlined region there are 1587 benign (out of 10396) and 712 malignant (out of 1048)",
      "votes": null
    },
    {
      "id": "944695",
      "postDate": "07/25/2020 09:06:55",
      "content": "<p>This is a helpful enlightening plot. We can look at (EDA) the images in the outlined region and consider what new features we can add to our models to help them distinguish samples in this overlapping region.</p>",
      "rawMarkdown": "This is a helpful enlightening plot. We can look at (EDA) the images in the outlined region and consider what new features we can add to our models to help them distinguish samples in this overlapping region.",
      "votes": null
    },
    {
      "id": "944702",
      "postDate": "07/25/2020 09:14:21",
      "content": "<p>Thank you for the comment. Yes, I am using the last layer of the efficientnet, which is followed just by a linear layer before the sigmoid. The explained variance ratio of the first component is 0.6759. I am thinking to do some pretraining in order to get a better representation.</p>",
      "rawMarkdown": "Thank you for the comment. Yes, I am using the last layer of the efficientnet, which is followed just by a linear layer before the sigmoid. The explained variance ratio of the first component is 0.6759. I am thinking to do some pretraining in order to get a better representation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 942694,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/24/2020 00:01:43",
      "content": "<p>Cool plot Paolo. When i get time, i will do these plots on my models. It looks like your model is doing a good job of separating malignant.</p>\n\n<p>In the bottom plot, if you plot all the <code>target=0</code> blue points first and then all the <code>target=1</code> yellow points second, we could better see how much they overlap. It looks like there is more overlap than is indicated. (And using a smaller dot, may help distinguish density).</p>",
      "votes": null,
      "replies": [
        {
          "id": 944689,
          "author_name": "meraxes10",
          "author_url": "",
          "post_date": "07/25/2020 08:56:36",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F261ca3dc7ecaad67d8854ba7fc0c7de9%2Fscatter_1.png?generation=1595667120673991&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F9f593b9f5aba02c944a2383e47de4cc7%2Fscatter_0.png?generation=1595667122005563&amp;alt=media\" alt=\"\"></p>\n\n<p>Thank you for the comment. There is too much overlap. In the outlined region there are 1587 benign (out of 10396) and 712 malignant (out of 1048)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944695,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/25/2020 09:06:55",
          "content": "<p>This is a helpful enlightening plot. We can look at (EDA) the images in the outlined region and consider what new features we can add to our models to help them distinguish samples in this overlapping region.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 943551,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "07/24/2020 12:33:50",
      "content": "<p>Hey <a href=\"/meraxes10\">@meraxes10</a> \nI think you shouldn't be surprised that PCA could get most of the variance in 1 component, as the last layer on the BCE is just a simple linear combination of the feature vector (assuming you used the last layer?) </p>\n\n<p>What's the explained variance of your 1st PCA component?</p>",
      "votes": null,
      "replies": [
        {
          "id": 944702,
          "author_name": "meraxes10",
          "author_url": "",
          "post_date": "07/25/2020 09:14:21",
          "content": "<p>Thank you for the comment. Yes, I am using the last layer of the efficientnet, which is followed just by a linear layer before the sigmoid. The explained variance ratio of the first component is 0.6759. I am thinking to do some pretraining in order to get a better representation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "942093": "I am using a custom BCELossWithLogits which upweights the loss of each sample by a costant factor and a model made of an efficientnet-b3 for the images and a simple MLP for the tabular data. \n\nI trained the model for 1 epoch using the folds 1-4 and fold 0 for validation from [melanoma-merged-external-data-512x512-jpeg](https://www.kaggle.com/shonenkov/melanoma-merged-external-data-512x512-jpeg).\n\nTo analyze the feature vectors that I get from the efficientnet I used PCA for dimensionality reduction and I was expecting to get two well defined clusters with a few misplaced predictions, but this is what I get.\n\nPredictions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F7718372830d0e2e29ae4c98e886c8d95%2Fscatter_pred.png?generation=1595514602930125&amp;alt=media)\n\nTrue labels.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F3b3828b98f9609c2bcd830177154784d%2Fscatter_target.png?generation=1595514670804409&amp;alt=media)\n\nDoes anyone else get something like this with different loss functions? Given that the malign samples are so closely clustered, do you think the model is overfitting???\n\nPS. This one gets a validation AUC of 0.93523 and a LB score of 0.9345",
    "942694": "Cool plot Paolo. When i get time, i will do these plots on my models. It looks like your model is doing a good job of separating malignant.\n\nIn the bottom plot, if you plot all the `target=0` blue points first and then all the `target=1` yellow points second, we could better see how much they overlap. It looks like there is more overlap than is indicated. (And using a smaller dot, may help distinguish density).",
    "943551": "Hey @meraxes10 \nI think you shouldn't be surprised that PCA could get most of the variance in 1 component, as the last layer on the BCE is just a simple linear combination of the feature vector (assuming you used the last layer?) \n\nWhat's the explained variance of your 1st PCA component?",
    "944689": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F261ca3dc7ecaad67d8854ba7fc0c7de9%2Fscatter_1.png?generation=1595667120673991&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3739757%2F9f593b9f5aba02c944a2383e47de4cc7%2Fscatter_0.png?generation=1595667122005563&amp;alt=media)\n\nThank you for the comment. There is too much overlap. In the outlined region there are 1587 benign (out of 10396) and 712 malignant (out of 1048)",
    "944695": "This is a helpful enlightening plot. We can look at (EDA) the images in the outlined region and consider what new features we can add to our models to help them distinguish samples in this overlapping region.",
    "944702": "Thank you for the comment. Yes, I am using the last layer of the efficientnet, which is followed just by a linear layer before the sigmoid. The explained variance ratio of the first component is 0.6759. I am thinking to do some pretraining in order to get a better representation."
  },
  "source": "meta"
}