{
  "id": 173130,
  "title": "Upsampling Malignant Improves AUC 0.003!",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173130",
  "author_name": "",
  "post_date": "2020-08-08T00:23:22.317276500Z",
  "votes": 53,
  "comment_count": 34,
  "views": 0,
  "content": "<p>People are wondering whether upsample can improve AUC, so I ran an experiment…</p>\n<h1>Upsampling Malignant Can Improve AUC for kNN !</h1>\n<p>The following experiment shows that upsample can increase AUC for kNN in Melanoma Comp by at least 0.003. The y-axis is triple stratified CV AUC score for kNN. The x-axis is malignant upsample ratio. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fee9a2d33de8e3904f032b69a56949406%2Faucc.png?generation=1596844592049793&amp;alt=media\" alt=\"\"></p>\n<p>This model uses just the 2020 comp data which (after removing duplicates) has 32692 images with 584 malignant for a proportion of 1.79% (when upsample=0). Then upsample 25x means that we add 25 additional copies of each malignant image. (i.e. we add <code>14600 = 25*584</code> images). Using 25x upsample increases the proportion to 32.1% malignant.</p>\n<h1>Upsampling Malignant Can Improve AUC for CNN !</h1>\n<p>Note that every model is different. So the AUC gain you observe in your model and the optimal upsample rate for your model may be different. In order to confirm that upsampling increases AUC for CNN. I also conducting a paired t-test for CNN and found that upsample increases AUC for CNN. I will post the code and CNN experiment details after the comp finishes. Below I post the code and experiment details for my kNN experiments.</p>\n<h1>Statistically Significant</h1>\n<p>Just to be sure this wasn't luck, I ran this kNN experiment 100 times with 100 different triple stratified 5 KFold seeds. The average AUC increase is 0.002 with STD 0.0012. We are 99.99999% confident that upsample increases AUC for kNN. (And based on my CNN experiments, we are 99.99999% confident that upsample increases AUC for CNN). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fdf48581148ee4265602dbd89b4f4ff63%2Fhist.png?generation=1596895000816762&amp;alt=media\" alt=\"\"></p>\n<h1>kNN Experiment</h1>\n<p>We can add upsampling to my triple stratified RAPIDS cuML kNN notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">here</a>. The model uses image embeddings from EfficientNetB0 and 256x256 images. We modify the below code to the below below code. To generate the plot above, we use the for-loop <code>for UP in np.arange(0,65,5):</code></p>\n<pre><code># CREATE TRAIN AND VALIDATION SUBSETS\nidxT2 = train.loc[train.tfrecord.isin(idxT)].index.values #2020 train\nidxV2 = train.loc[train.tfrecord.isin(idxV)].index.values #2020 valid\n\n# MODEL\nmodel = cuml.neighbors.KNeighborsClassifier(n_neighbors=299)\nmodel.fit(embed[idxT2,],train.target.values[idxT2])\n</code></pre>\n<p>We change the above code to the below code</p>\n<pre><code># HOW MANY TIMES TO UPSAMPLE\nUP = 25\n\n# CREATE TRAIN AND VALIDATION SUBSETS\nidxT2 = train.loc[train.tfrecord.isin(idxT)].index.values #2020 train\nidxV2 = train.loc[train.tfrecord.isin(idxV)].index.values #2020 valid\n\n# UPSAMPLE\nidxM = np.where( train.target.values[idxT2]==1 )[0]; idxM2 = idxT2[idxM]\nX_more = np.zeros((len(idxT2)+len(idxM2)*UP,embed.shape[1])) \nX_more[:len(idxT2),] = embed[idxT2,]\ny_more = np.zeros((len(idxT2)+len(idxM2)*UP))\ny_more[:len(idxT2)] = train.target.values[idxT2]\nfor k in range(UP): \n    X_more[len(idxT2)+len(idxM2)*k:len(idxT2)+len(idxM2)*(k+1),] = embed[idxM2,]\n    y_more[len(idxT2)+len(idxM2)*k:len(idxT2)+len(idxM2)*(k+1)] = train.target.values[idxM2]\n\n# MODEL\nmodel = cuml.neighbors.KNeighborsClassifier(n_neighbors=299)\nmodel.fit(X_more,y_more)\n</code></pre>\n<h1>How Upsample Improves AUC</h1>\n<p>Upsample affects AUC by increasing some predictions more than others. Here's an example showing upsample improving AUC. In the plot below, there are 4 unknown data points labeled <code>A, B, C, D</code>. What would you guess are their true labels? Yellow denotes <code>target=1</code> and blue denotes <code>target=0</code>. </p>\n<h2>Without Upsample - AUC = 0.75</h2>\n<p>If you use kNN with <code>k=5</code>, then we will predict probabilities <code>0.0, 0.4, 0.6, 1.0</code> for points <code>A, B, C, D</code> respectively. It turns out that the true labels are <code>0, 1, 0, 1</code> so our <code>AUC = 0.75</code> because two predictions are out of ranked order.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fc897605c98e78b26c634e6659d87f8cb%2Fknn1.jpg?generation=1596844755349238&amp;alt=media\" alt=\"\"></p>\n<h2>With Upsample - AUC = 1.00</h2>\n<p>If we upsample the malignant samples by 2x and then use kNN with <code>k=5</code>, then we predict <code>0.0, 0.8, 0.6, 1.0</code>. Since the true labels are <code>0, 1, 0, 1</code> our <code>AUC = 1.0</code> because no predictions are out of ranked order.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F4389b112dbdef9497ef7fab2b6312177%2Fknn2.jpg?generation=1596845180779774&amp;alt=media\" alt=\"\"></p>\n<h1>How To Upsample Discussion</h1>\n<p>There is a discussion explaining how to upsample <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a></p>\n<h1>How To Upsample Notebook</h1>\n<p>There is a starter notebook demonstrating upsample <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\" target=\"_blank\">here</a></p>",
  "messages": [
    {
      "id": "962242",
      "postDate": "08/08/2020 00:23:22",
      "content": "<p>People are wondering whether upsample can improve AUC, so I ran an experiment…</p>\n<h1>Upsampling Malignant Can Improve AUC for kNN !</h1>\n<p>The following experiment shows that upsample can increase AUC for kNN in Melanoma Comp by at least 0.003. The y-axis is triple stratified CV AUC score for kNN. The x-axis is malignant upsample ratio. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fee9a2d33de8e3904f032b69a56949406%2Faucc.png?generation=1596844592049793&amp;alt=media\" alt=\"\"></p>\n<p>This model uses just the 2020 comp data which (after removing duplicates) has 32692 images with 584 malignant for a proportion of 1.79% (when upsample=0). Then upsample 25x means that we add 25 additional copies of each malignant image. (i.e. we add <code>14600 = 25*584</code> images). Using 25x upsample increases the proportion to 32.1% malignant.</p>\n<h1>Upsampling Malignant Can Improve AUC for CNN !</h1>\n<p>Note that every model is different. So the AUC gain you observe in your model and the optimal upsample rate for your model may be different. In order to confirm that upsampling increases AUC for CNN. I also conducting a paired t-test for CNN and found that upsample increases AUC for CNN. I will post the code and CNN experiment details after the comp finishes. Below I post the code and experiment details for my kNN experiments.</p>\n<h1>Statistically Significant</h1>\n<p>Just to be sure this wasn't luck, I ran this kNN experiment 100 times with 100 different triple stratified 5 KFold seeds. The average AUC increase is 0.002 with STD 0.0012. We are 99.99999% confident that upsample increases AUC for kNN. (And based on my CNN experiments, we are 99.99999% confident that upsample increases AUC for CNN). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fdf48581148ee4265602dbd89b4f4ff63%2Fhist.png?generation=1596895000816762&amp;alt=media\" alt=\"\"></p>\n<h1>kNN Experiment</h1>\n<p>We can add upsampling to my triple stratified RAPIDS cuML kNN notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">here</a>. The model uses image embeddings from EfficientNetB0 and 256x256 images. We modify the below code to the below below code. To generate the plot above, we use the for-loop <code>for UP in np.arange(0,65,5):</code></p>\n<pre><code># CREATE TRAIN AND VALIDATION SUBSETS\nidxT2 = train.loc[train.tfrecord.isin(idxT)].index.values #2020 train\nidxV2 = train.loc[train.tfrecord.isin(idxV)].index.values #2020 valid\n\n# MODEL\nmodel = cuml.neighbors.KNeighborsClassifier(n_neighbors=299)\nmodel.fit(embed[idxT2,],train.target.values[idxT2])\n</code></pre>\n<p>We change the above code to the below code</p>\n<pre><code># HOW MANY TIMES TO UPSAMPLE\nUP = 25\n\n# CREATE TRAIN AND VALIDATION SUBSETS\nidxT2 = train.loc[train.tfrecord.isin(idxT)].index.values #2020 train\nidxV2 = train.loc[train.tfrecord.isin(idxV)].index.values #2020 valid\n\n# UPSAMPLE\nidxM = np.where( train.target.values[idxT2]==1 )[0]; idxM2 = idxT2[idxM]\nX_more = np.zeros((len(idxT2)+len(idxM2)*UP,embed.shape[1])) \nX_more[:len(idxT2),] = embed[idxT2,]\ny_more = np.zeros((len(idxT2)+len(idxM2)*UP))\ny_more[:len(idxT2)] = train.target.values[idxT2]\nfor k in range(UP): \n    X_more[len(idxT2)+len(idxM2)*k:len(idxT2)+len(idxM2)*(k+1),] = embed[idxM2,]\n    y_more[len(idxT2)+len(idxM2)*k:len(idxT2)+len(idxM2)*(k+1)] = train.target.values[idxM2]\n\n# MODEL\nmodel = cuml.neighbors.KNeighborsClassifier(n_neighbors=299)\nmodel.fit(X_more,y_more)\n</code></pre>\n<h1>How Upsample Improves AUC</h1>\n<p>Upsample affects AUC by increasing some predictions more than others. Here's an example showing upsample improving AUC. In the plot below, there are 4 unknown data points labeled <code>A, B, C, D</code>. What would you guess are their true labels? Yellow denotes <code>target=1</code> and blue denotes <code>target=0</code>. </p>\n<h2>Without Upsample - AUC = 0.75</h2>\n<p>If you use kNN with <code>k=5</code>, then we will predict probabilities <code>0.0, 0.4, 0.6, 1.0</code> for points <code>A, B, C, D</code> respectively. It turns out that the true labels are <code>0, 1, 0, 1</code> so our <code>AUC = 0.75</code> because two predictions are out of ranked order.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fc897605c98e78b26c634e6659d87f8cb%2Fknn1.jpg?generation=1596844755349238&amp;alt=media\" alt=\"\"></p>\n<h2>With Upsample - AUC = 1.00</h2>\n<p>If we upsample the malignant samples by 2x and then use kNN with <code>k=5</code>, then we predict <code>0.0, 0.8, 0.6, 1.0</code>. Since the true labels are <code>0, 1, 0, 1</code> our <code>AUC = 1.0</code> because no predictions are out of ranked order.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F4389b112dbdef9497ef7fab2b6312177%2Fknn2.jpg?generation=1596845180779774&amp;alt=media\" alt=\"\"></p>\n<h1>How To Upsample Discussion</h1>\n<p>There is a discussion explaining how to upsample <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\" target=\"_blank\">here</a></p>\n<h1>How To Upsample Notebook</h1>\n<p>There is a starter notebook demonstrating upsample <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "People are wondering whether upsample can improve AUC, so I ran an experiment...\n# Upsampling Malignant Can Improve AUC for kNN !\nThe following experiment shows that upsample can increase AUC for kNN in Melanoma Comp by at least 0.003. The y-axis is triple stratified CV AUC score for kNN. The x-axis is malignant upsample ratio. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fee9a2d33de8e3904f032b69a56949406%2Faucc.png?generation=1596844592049793&amp;alt=media)\n\nThis model uses just the 2020 comp data which (after removing duplicates) has 32692 images with 584 malignant for a proportion of 1.79% (when upsample=0). Then upsample 25x means that we add 25 additional copies of each malignant image. (i.e. we add `14600 = 25*584` images). Using 25x upsample increases the proportion to 32.1% malignant.\n\n# Upsampling Malignant Can Improve AUC for CNN !\nNote that every model is different. So the AUC gain you observe in your model and the optimal upsample rate for your model may be different. In order to confirm that upsampling increases AUC for CNN. I also conducting a paired t-test for CNN and found that upsample increases AUC for CNN. I will post the code and CNN experiment details after the comp finishes. Below I post the code and experiment details for my kNN experiments.\n\n# Statistically Significant\nJust to be sure this wasn't luck, I ran this kNN experiment 100 times with 100 different triple stratified 5 KFold seeds. The average AUC increase is 0.002 with STD 0.0012. We are 99.99999% confident that upsample increases AUC for kNN. (And based on my CNN experiments, we are 99.99999% confident that upsample increases AUC for CNN). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fdf48581148ee4265602dbd89b4f4ff63%2Fhist.png?generation=1596895000816762&amp;alt=media)\n\n# kNN Experiment\nWe can add upsampling to my triple stratified RAPIDS cuML kNN notebook [here][1]. The model uses image embeddings from EfficientNetB0 and 256x256 images. We modify the below code to the below below code. To generate the plot above, we use the for-loop `for UP in np.arange(0,65,5):`\n\n    # CREATE TRAIN AND VALIDATION SUBSETS\n    idxT2 = train.loc[train.tfrecord.isin(idxT)].index.values #2020 train\n    idxV2 = train.loc[train.tfrecord.isin(idxV)].index.values #2020 valid\n    \n    # MODEL\n    model = cuml.neighbors.KNeighborsClassifier(n_neighbors=299)\n    model.fit(embed[idxT2,],train.target.values[idxT2])\n\nWe change the above code to the below code\n\n    # HOW MANY TIMES TO UPSAMPLE\n    UP = 25\n\n    # CREATE TRAIN AND VALIDATION SUBSETS\n    idxT2 = train.loc[train.tfrecord.isin(idxT)].index.values #2020 train\n    idxV2 = train.loc[train.tfrecord.isin(idxV)].index.values #2020 valid\n    \n    # UPSAMPLE\n    idxM = np.where( train.target.values[idxT2]==1 )[0]; idxM2 = idxT2[idxM]\n    X_more = np.zeros((len(idxT2)+len(idxM2)*UP,embed.shape[1])) \n    X_more[:len(idxT2),] = embed[idxT2,]\n    y_more = np.zeros((len(idxT2)+len(idxM2)*UP))\n    y_more[:len(idxT2)] = train.target.values[idxT2]\n    for k in range(UP): \n        X_more[len(idxT2)+len(idxM2)*k:len(idxT2)+len(idxM2)*(k+1),] = embed[idxM2,]\n        y_more[len(idxT2)+len(idxM2)*k:len(idxT2)+len(idxM2)*(k+1)] = train.target.values[idxM2]\n    \n    # MODEL\n    model = cuml.neighbors.KNeighborsClassifier(n_neighbors=299)\n    model.fit(X_more,y_more)\n\n# How Upsample Improves AUC\nUpsample affects AUC by increasing some predictions more than others. Here's an example showing upsample improving AUC. In the plot below, there are 4 unknown data points labeled `A, B, C, D`. What would you guess are their true labels? Yellow denotes `target=1` and blue denotes `target=0`. \n## Without Upsample - AUC = 0.75\nIf you use kNN with `k=5`, then we will predict probabilities `0.0, 0.4, 0.6, 1.0` for points `A, B, C, D` respectively. It turns out that the true labels are `0, 1, 0, 1` so our `AUC = 0.75` because two predictions are out of ranked order.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fc897605c98e78b26c634e6659d87f8cb%2Fknn1.jpg?generation=1596844755349238&amp;alt=media)\n\n## With Upsample - AUC = 1.00\nIf we upsample the malignant samples by 2x and then use kNN with `k=5`, then we predict `0.0, 0.8, 0.6, 1.0`. Since the true labels are `0, 1, 0, 1` our `AUC = 1.0` because no predictions are out of ranked order.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F4389b112dbdef9497ef7fab2b6312177%2Fknn2.jpg?generation=1596845180779774&amp;alt=media)\n\n# How To Upsample Discussion\nThere is a discussion explaining how to upsample [here][3]\n# How To Upsample Notebook\nThere is a starter notebook demonstrating upsample [here][4]\n\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n[3]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n[4]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
      "votes": null
    },
    {
      "id": "962467",
      "postDate": "08/08/2020 06:29:37",
      "content": "<p>Interesting result, thank you for sharing! Would not have thought that such a strong upsample (25x) would work best.</p>",
      "rawMarkdown": "Interesting result, thank you for sharing! Would not have thought that such a strong upsample (25x) would work best.",
      "votes": null
    },
    {
      "id": "962486",
      "postDate": "08/08/2020 06:52:50",
      "content": "<p>Thanks Chris! Wouldn't setting <code>class_weights</code> achieve the same result?</p>",
      "rawMarkdown": "Thanks Chris! Wouldn't setting `class_weights` achieve the same result?",
      "votes": null
    },
    {
      "id": "962537",
      "postDate": "08/08/2020 08:06:17",
      "content": "<p>Would be nice to test two approaches. Think when we upsample then model see the same melanoma images after different augmentations and thats could lead to better result than weights.</p>",
      "rawMarkdown": "Would be nice to test two approaches. Think when we upsample then model see the same melanoma images after different augmentations and thats could lead to better result than weights.",
      "votes": null
    },
    {
      "id": "962573",
      "postDate": "08/08/2020 08:45:26",
      "content": "<p>This is interesting, but it does not mean upsampling would improve CNN roc-auc.  </p>",
      "rawMarkdown": "This is interesting, but it does not mean upsampling would improve CNN roc-auc.",
      "votes": null
    },
    {
      "id": "962623",
      "postDate": "08/08/2020 09:34:58",
      "content": "<p>The datasets are also the hyper-parameters in this competition. 😭😭😭</p>",
      "rawMarkdown": "The datasets are also the hyper-parameters in this competition. 😭😭😭",
      "votes": null
    },
    {
      "id": "962634",
      "postDate": "08/08/2020 09:45:40",
      "content": "<p>I may be wrong, but improvements like 0.003 AUC are well inside confidence intervals for this competition and most likely occur by chance. At least for me, standard deviations for model scores are around 0.01 for 5-fold CV and I would not consider 0.003 as a definite improvement.</p>",
      "rawMarkdown": "I may be wrong, but improvements like 0.003 AUC are well inside confidence intervals for this competition and most likely occur by chance. At least for me, standard deviations for model scores are around 0.01 for 5-fold CV and I would not consider 0.003 as a definite improvement.",
      "votes": null
    },
    {
      "id": "962637",
      "postDate": "08/08/2020 09:58:25",
      "content": "<p>Neither RAPIDS nor scikit-learn KNN support class weights.  </p>",
      "rawMarkdown": "Neither RAPIDS nor scikit-learn KNN support class weights.",
      "votes": null
    },
    {
      "id": "962639",
      "postDate": "08/08/2020 09:59:18",
      "content": "<p>It depend son your CV setting.  For mine, a 0.003 CV improvement translates to a LB improvement so far.</p>",
      "rawMarkdown": "It depend son your CV setting.  For mine, a 0.003 CV improvement translates to a LB improvement so far.",
      "votes": null
    },
    {
      "id": "962643",
      "postDate": "08/08/2020 10:04:32",
      "content": "<p>I was referring to the graph that plots the CNN performance in function of the upsampling ratio. Not the artificial KNN example.\nEDIT: i mis-interpreted the plot. Looking at the y-axis, it might have been KNN there as well.</p>\n\n<p>After some reading, they are not equivalent. By duplicating images, the probability that one or more are included in a batch increases. This is not the case for custom class weights.</p>",
      "rawMarkdown": "I was referring to the graph that plots the CNN performance in function of the upsampling ratio. Not the artificial KNN example.\nEDIT: i mis-interpreted the plot. Looking at the y-axis, it might have been KNN there as well.\n\nAfter some reading, they are not equivalent. By duplicating images, the probability that one or more are included in a batch increases. This is not the case for custom class weights.",
      "votes": null
    },
    {
      "id": "962656",
      "postDate": "08/08/2020 10:11:44",
      "content": "<p>The graph plots KNN performance, not CNN.  </p>",
      "rawMarkdown": "The graph plots KNN performance, not CNN.",
      "votes": null
    },
    {
      "id": "962732",
      "postDate": "08/08/2020 11:47:29",
      "content": "<p>Best work</p>",
      "rawMarkdown": "Best work",
      "votes": null
    },
    {
      "id": "962825",
      "postDate": "08/08/2020 13:13:32",
      "content": "<blockquote>\n  <p>Wouldn't setting class_weights achieve the same result?</p>\n</blockquote>\n<p>No. Using <code>class_weights</code> will not change the rank of kNN predictions and therefore will not change the AUC. In my example above, using <code>class_weights = {0:1, 1:2}</code> will make predictions <code>0.0, 4/7, 6/8, 1.0</code> and obtain <code>AUC = 0.75</code> (the same AUC as using <code>class_weights = {0:1, 1:X} for any X</code>) . </p>",
      "rawMarkdown": "&gt; Wouldn't setting class_weights achieve the same result?\n\nNo. Using `class_weights` will not change the rank of kNN predictions and therefore will not change the AUC. In my example above, using `class_weights = {0:1, 1:2}` will make predictions `0.0, 4/7, 6/8, 1.0` and obtain `AUC = 0.75` (the same AUC as using `class_weights = {0:1, 1:X} for any X`) .",
      "votes": null
    },
    {
      "id": "962865",
      "postDate": "08/08/2020 13:55:00",
      "content": "<p>I ran this experiment 100 times with 100 different seeds. The average AUC increase is 0.002 with STD 0.0012. These results are statistically significant with p&lt;0.0001 (paired one sided t test). We are 99.99999% confident that upsampling increases AUC.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F64fff89f470211965a6c6d5c265200e3%2Fhist.png?generation=1596894830615660&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I ran this experiment 100 times with 100 different seeds. The average AUC increase is 0.002 with STD 0.0012. These results are statistically significant with p&lt;0.0001 (paired one sided t test). We are 99.99999% confident that upsampling increases AUC.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F64fff89f470211965a6c6d5c265200e3%2Fhist.png?generation=1596894830615660&amp;alt=media)",
      "votes": null
    },
    {
      "id": "963638",
      "postDate": "08/09/2020 07:27:03",
      "content": "<p>Thanks for the experiments. Still, I think it may not be  completely relevant to CNNs</p>",
      "rawMarkdown": "Thanks for the experiments. Still, I think it may not be  completely relevant to CNNs",
      "votes": null
    },
    {
      "id": "964399",
      "postDate": "08/09/2020 20:09:08",
      "content": "<p>Hi Chris, thanks for running this analysis! What statistical test did you use to compute your p-value? Given an observed increase in AUC of 0.002 with standard deviation 0.0012, what was the null hypothesis that this was compared to? </p>",
      "rawMarkdown": "Hi Chris, thanks for running this analysis! What statistical test did you use to compute your p-value? Given an observed increase in AUC of 0.002 with standard deviation 0.0012, what was the null hypothesis that this was compared to?",
      "votes": null
    },
    {
      "id": "964504",
      "postDate": "08/10/2020 00:33:28",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": null
    },
    {
      "id": "964517",
      "postDate": "08/10/2020 00:50:36",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/jaronthompson\" target=\"_blank\">@jaronthompson</a> , the null hypothesis is <code>H0 : mean AUC difference = 0</code> (i.e AUC stays the same) and alternative hypothesis is <code>HA : mean AUC difference &gt; 0</code> (i.e. AUC increases). We then perform a one-sided <a href=\"https://online.stat.psu.edu/stat415/lesson/10/10.3\" target=\"_blank\">paired t-test</a> and reject the null hypothesis with <code>p = &lt; .00001</code></p>\n<p>There's a greater than 99.99999% chance that upsample increases AUC!</p>\n<p>For each of 100 pairs of experiments, i choose a different seed for triple stratified KFold. Then I trained with and without upsampling and calculated the difference in AUC score. The mean difference was <code>0.00207</code> with standard deviation <code>0.00122</code>. Using the online t-test calulator <a href=\"https://www.usablestats.com/calcs/1samplet&amp;summary=1\" target=\"_blank\">here</a>, we see that the <code>p = &lt; .00001</code>.</p>",
      "rawMarkdown": "Hi @jaronthompson , the null hypothesis is `H0 : mean AUC difference = 0` (i.e AUC stays the same) and alternative hypothesis is `HA : mean AUC difference &gt; 0` (i.e. AUC increases). We then perform a one-sided [paired t-test][1] and reject the null hypothesis with `p = &lt; .00001`\n\nThere's a greater than 99.99999% chance that upsample increases AUC!\n\nFor each of 100 pairs of experiments, i choose a different seed for triple stratified KFold. Then I trained with and without upsampling and calculated the difference in AUC score. The mean difference was `0.00207` with standard deviation `0.00122`. Using the online t-test calulator [here][2], we see that the `p = &lt; .00001`.\n\n\n\n[1]: https://online.stat.psu.edu/stat415/lesson/10/10.3\n[2]: https://www.usablestats.com/calcs/1samplet&amp;summary=1",
      "votes": null
    },
    {
      "id": "964531",
      "postDate": "08/10/2020 01:22:47",
      "content": "<p>How do you use class weight with  rapids KNeighborsClassifier?</p>",
      "rawMarkdown": "How do you use class weight with  rapids KNeighborsClassifier?",
      "votes": null
    },
    {
      "id": "964575",
      "postDate": "08/10/2020 02:56:57",
      "content": "<blockquote>\n  <p>Still, I think it may not be completely relevant to CNNs</p>\n</blockquote>\n<p>I also was curious about CNNs, so I repeated the paired t test on CNNs and observed an AUC increase with p&lt;0.0001. We are 99.99999% confident that upsampling increases AUC for CNNs in Melanoma comp! I will post the experiment results after the comp finishes. </p>",
      "rawMarkdown": "&gt; Still, I think it may not be completely relevant to CNNs\n\nI also was curious about CNNs, so I repeated the paired t test on CNNs and observed an AUC increase with p&lt;0.0001. We are 99.99999% confident that upsampling increases AUC for CNNs in Melanoma comp! I will post the experiment results after the comp finishes.",
      "votes": null
    },
    {
      "id": "965256",
      "postDate": "08/10/2020 13:54:10",
      "content": "<p>Amazing Job! </p>",
      "rawMarkdown": "Amazing Job!",
      "votes": null
    },
    {
      "id": "965260",
      "postDate": "08/10/2020 13:57:44",
      "content": "<p>Great, thanks!</p>",
      "rawMarkdown": "Great, thanks!",
      "votes": null
    },
    {
      "id": "965304",
      "postDate": "08/10/2020 14:32:32",
      "content": "<p>Chris, thanks, this is the first time someone says upsampling improves roc-auc for CNN.  What improvement did you see in your experiment?</p>",
      "rawMarkdown": "Chris, thanks, this is the first time someone says upsampling improves roc-auc for CNN.  What improvement did you see in your experiment?",
      "votes": null
    },
    {
      "id": "965724",
      "postDate": "08/10/2020 19:48:07",
      "content": "<p>With kNN, (if i understand class weight correctly) you can change class weight using post process formula. Given <code>prediction = p</code>, with <code>kNN parameter = k</code>, and class weight <code>{0:1, 1:w}</code>. Then you can apply class weight with</p>\n<pre><code>new pred = w*p*k / (w*p*k + (1-p)*k)\n</code></pre>\n<p>This is a rank preserving transformation, so it will not alter AUC. Regarding kNN, you need to add multiple copies of malignant to affect AUC.</p>",
      "rawMarkdown": "With kNN, (if i understand class weight correctly) you can change class weight using post process formula. Given `prediction = p`, with `kNN parameter = k`, and class weight `{0:1, 1:w}`. Then you can apply class weight with\n\n    new pred = w*p*k / (w*p*k + (1-p)*k)\n\nThis is a rank preserving transformation, so it will not alter AUC. Regarding kNN, you need to add multiple copies of malignant to affect AUC.",
      "votes": null
    },
    {
      "id": "965727",
      "postDate": "08/10/2020 19:50:02",
      "content": "<p><a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> I'm not sure if you saw my answer to your question above. It is interesting to note that <code>class weight</code> and adding more copies of malignant is not the same for kNN. </p>",
      "rawMarkdown": "group16 I'm not sure if you saw my answer to your question above. It is interesting to note that `class weight` and adding more copies of malignant is not the same for kNN.",
      "votes": null
    },
    {
      "id": "965730",
      "postDate": "08/10/2020 19:53:18",
      "content": "<p>I saw it <a href=\"/cdeotte\">@cdeotte</a> (but forgot to upvote, shame on me)! Our team is on it ;)</p>",
      "rawMarkdown": "I saw it @cdeotte (but forgot to upvote, shame on me)! Our team is on it ;)",
      "votes": null
    },
    {
      "id": "965879",
      "postDate": "08/10/2020 23:34:35",
      "content": "<p>Is the 0.003 worth it?</p>",
      "rawMarkdown": "Is the 0.003 worth it?",
      "votes": null
    },
    {
      "id": "966315",
      "postDate": "08/11/2020 10:23:36",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Your math seems correct to me indeed:</p>\n<p><code>new pred = w*p*k / (w*p*k + (1-p)*k)</code></p>\n<p>Thanks for checking.</p>",
      "rawMarkdown": "cdeotte Your math seems correct to me indeed:\n\n `new pred = w*p*k / (w*p*k + (1-p)*k)`\n\nThanks for checking.",
      "votes": null
    },
    {
      "id": "966699",
      "postDate": "08/11/2020 16:05:59",
      "content": "<p>This is a very good analysis. However, I am still debating on using upsampling (or undersampling) to deal with the imbalance. For me, the focal loss strategy has been yielding some interesting results, but the reason I am refraining from upsampling is to match the real-world scenario as accurately as possible. In the real world, there are, of course, significantly higher non-malignant (benign) cases than those of malignant. So if the model is now learning more about malignant cases, it might as well (due to small factors in images) detect (falsely…?) non-malignant images as malignant. Being a beginner, I am flooded with questions, and I might be very wrong too. </p>",
      "rawMarkdown": "This is a very good analysis. However, I am still debating on using upsampling (or undersampling) to deal with the imbalance. For me, the focal loss strategy has been yielding some interesting results, but the reason I am refraining from upsampling is to match the real-world scenario as accurately as possible. In the real world, there are, of course, significantly higher non-malignant (benign) cases than those of malignant. So if the model is now learning more about malignant cases, it might as well (due to small factors in images) detect (falsely...?) non-malignant images as malignant. Being a beginner, I am flooded with questions, and I might be very wrong too.",
      "votes": null
    },
    {
      "id": "966739",
      "postDate": "08/11/2020 16:26:50",
      "content": "<p>These are good concerns. You are correct that in every real world situation we must adjust our model to maximize some metric which will lower unacceptable errors (reduce false positive, reduce false negative, etc). Each situation is different and our employer will tell us what to maximize.</p>\n<p>In this competition, the leaderboard is our \"employer\", and our \"employer\" is asking us to maximize AUC</p>",
      "rawMarkdown": "These are good concerns. You are correct that in every real world situation we must adjust our model to maximize some metric which will lower unacceptable errors (reduce false positive, reduce false negative, etc). Each situation is different and our employer will tell us what to maximize.\n\nIn this competition, the leaderboard is our \"employer\", and our \"employer\" is asking us to maximize AUC",
      "votes": null
    },
    {
      "id": "966753",
      "postDate": "08/11/2020 16:40:20",
      "content": "<p>False positives are (in most medical cases) often better than false negatives. Detecting non-malignant images as malignant will require the doctors to do an unnecessary investigation.</p>\n<p>Detecting a malignant image as benign will cost a human life.</p>",
      "rawMarkdown": "False positives are (in most medical cases) often better than false negatives. Detecting non-malignant images as malignant will require the doctors to do an unnecessary investigation.\n\nDetecting a malignant image as benign will cost a human life.",
      "votes": null
    },
    {
      "id": "966761",
      "postDate": "08/11/2020 16:43:55",
      "content": "<p>Of course. That is a margin of around 25 places on the top of the LB.  And probably a few 100 places lower on the LB.</p>",
      "rawMarkdown": "Of course. That is a margin of around 25 places on the top of the LB.  And probably a few 100 places lower on the LB.",
      "votes": null
    },
    {
      "id": "966774",
      "postDate": "08/11/2020 16:49:36",
      "content": "<p>Yes. And in real life an increase in 0.003 can help us save more lives.</p>",
      "rawMarkdown": "Yes. And in real life an increase in 0.003 can help us save more lives.",
      "votes": null
    },
    {
      "id": "970907",
      "postDate": "08/15/2020 01:54:41",
      "content": "<p>I see your point. In the end, we're just trying to do anything we can to improve in the race! </p>",
      "rawMarkdown": "I see your point. In the end, we're just trying to do anything we can to improve in the race!",
      "votes": null
    },
    {
      "id": "976105",
      "postDate": "08/18/2020 16:31:44",
      "content": "<p>Below are the results of my upsample experiments on CNN. Once again, I used my embedding extraction notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">here</a>. After extracting embeddings, we can model them with an MLP. This simulates modeling with a CNN and freezing the backbone. The MLP head is as follows</p>\n<pre><code>def MakeModel():\n    inp = tf.keras.layers.Input((1280,))\n    x = tf.keras.layers.Dense(1,activation='sigmoid')(inp)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    model.compile(optimizer=tf.keras.optimizers.SGD(),\n           loss=tf.keras.losses.BinaryCrossentropy(),\n           metrics=['AUC'])\n    return model\n</code></pre>\n<p>Next we train with the same code posted above for upsample, and we use early stopping on the MLP</p>\n<pre><code>        model = MakeModel()\n\n        cc = tf.keras.callbacks.EarlyStopping(\n            monitor='val_auc', min_delta=0, patience=2, verbose=0, mode='max',\n            baseline=None, restore_best_weights=True\n        )\n\n        model.fit(X_more, y_more, epochs=30, verbose=0,\n             validation_data = (embed[idxV2,],train.target.values[idxV2]),callbacks=[cc] )\n</code></pre>\n<p>Here is the affect of upsample</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8a784af724af627812c6c997645e23c8%2Fa1.png?generation=1597768115192486&amp;alt=media\" alt=\"\"></p>\n<p>And here is the result of 50 experiments. Mean AUC increase is 0.026 with std 0.0036</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcd0388046eac1e5ef4286e855fcc5cda%2Fa2.png?generation=1597768137923981&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Below are the results of my upsample experiments on CNN. Once again, I used my embedding extraction notebook [here][1]. After extracting embeddings, we can model them with an MLP. This simulates modeling with a CNN and freezing the backbone. The MLP head is as follows\n\n    def MakeModel():\n        inp = tf.keras.layers.Input((1280,))\n        x = tf.keras.layers.Dense(1,activation='sigmoid')(inp)\n        model = tf.keras.Model(inputs=inp,outputs=x)\n        model.compile(optimizer=tf.keras.optimizers.SGD(),\n               loss=tf.keras.losses.BinaryCrossentropy(),\n               metrics=['AUC'])\n        return model\n\nNext we train with the same code posted above for upsample, and we use early stopping on the MLP\n\n           \n            model = MakeModel()\n\n            cc = tf.keras.callbacks.EarlyStopping(\n                monitor='val_auc', min_delta=0, patience=2, verbose=0, mode='max',\n                baseline=None, restore_best_weights=True\n            )\n        \n            model.fit(X_more, y_more, epochs=30, verbose=0,\n                 validation_data = (embed[idxV2,],train.target.values[idxV2]),callbacks=[cc] )\n\nHere is the affect of upsample\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8a784af724af627812c6c997645e23c8%2Fa1.png?generation=1597768115192486&alt=media)\n\nAnd here is the result of 50 experiments. Mean AUC increase is 0.026 with std 0.0036\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcd0388046eac1e5ef4286e855fcc5cda%2Fa2.png?generation=1597768137923981&alt=media)\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 962573,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/08/2020 08:45:26",
      "content": "<p>This is interesting, but it does not mean upsampling would improve CNN roc-auc.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962623,
      "author_name": "apthagowda",
      "author_url": "",
      "post_date": "08/08/2020 09:34:58",
      "content": "<p>The datasets are also the hyper-parameters in this competition. 😭😭😭</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962634,
      "author_name": "ademyanchuk",
      "author_url": "",
      "post_date": "08/08/2020 09:45:40",
      "content": "<p>I may be wrong, but improvements like 0.003 AUC are well inside confidence intervals for this competition and most likely occur by chance. At least for me, standard deviations for model scores are around 0.01 for 5-fold CV and I would not consider 0.003 as a definite improvement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 962639,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2020 09:59:18",
          "content": "<p>It depend son your CV setting.  For mine, a 0.003 CV improvement translates to a LB improvement so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962865,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/08/2020 13:55:00",
          "content": "<p>I ran this experiment 100 times with 100 different seeds. The average AUC increase is 0.002 with STD 0.0012. These results are statistically significant with p&lt;0.0001 (paired one sided t test). We are 99.99999% confident that upsampling increases AUC.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F64fff89f470211965a6c6d5c265200e3%2Fhist.png?generation=1596894830615660&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 963638,
          "author_name": "ademyanchuk",
          "author_url": "",
          "post_date": "08/09/2020 07:27:03",
          "content": "<p>Thanks for the experiments. Still, I think it may not be  completely relevant to CNNs</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964575,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/10/2020 02:56:57",
          "content": "<blockquote>\n  <p>Still, I think it may not be completely relevant to CNNs</p>\n</blockquote>\n<p>I also was curious about CNNs, so I repeated the paired t test on CNNs and observed an AUC increase with p&lt;0.0001. We are 99.99999% confident that upsampling increases AUC for CNNs in Melanoma comp! I will post the experiment results after the comp finishes. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965260,
          "author_name": "ademyanchuk",
          "author_url": "",
          "post_date": "08/10/2020 13:57:44",
          "content": "<p>Great, thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965304,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2020 14:32:32",
          "content": "<p>Chris, thanks, this is the first time someone says upsampling improves roc-auc for CNN.  What improvement did you see in your experiment?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 964399,
      "author_name": "jaronthompson",
      "author_url": "",
      "post_date": "08/09/2020 20:09:08",
      "content": "<p>Hi Chris, thanks for running this analysis! What statistical test did you use to compute your p-value? Given an observed increase in AUC of 0.002 with standard deviation 0.0012, what was the null hypothesis that this was compared to? </p>",
      "votes": null,
      "replies": [
        {
          "id": 964517,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/10/2020 00:50:36",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/jaronthompson\" target=\"_blank\">@jaronthompson</a> , the null hypothesis is <code>H0 : mean AUC difference = 0</code> (i.e AUC stays the same) and alternative hypothesis is <code>HA : mean AUC difference &gt; 0</code> (i.e. AUC increases). We then perform a one-sided <a href=\"https://online.stat.psu.edu/stat415/lesson/10/10.3\" target=\"_blank\">paired t-test</a> and reject the null hypothesis with <code>p = &lt; .00001</code></p>\n<p>There's a greater than 99.99999% chance that upsample increases AUC!</p>\n<p>For each of 100 pairs of experiments, i choose a different seed for triple stratified KFold. Then I trained with and without upsampling and calculated the difference in AUC score. The mean difference was <code>0.00207</code> with standard deviation <code>0.00122</code>. Using the online t-test calulator <a href=\"https://www.usablestats.com/calcs/1samplet&amp;summary=1\" target=\"_blank\">here</a>, we see that the <code>p = &lt; .00001</code>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 966699,
      "author_name": "harshnagouda",
      "author_url": "",
      "post_date": "08/11/2020 16:05:59",
      "content": "<p>This is a very good analysis. However, I am still debating on using upsampling (or undersampling) to deal with the imbalance. For me, the focal loss strategy has been yielding some interesting results, but the reason I am refraining from upsampling is to match the real-world scenario as accurately as possible. In the real world, there are, of course, significantly higher non-malignant (benign) cases than those of malignant. So if the model is now learning more about malignant cases, it might as well (due to small factors in images) detect (falsely…?) non-malignant images as malignant. Being a beginner, I am flooded with questions, and I might be very wrong too. </p>",
      "votes": null,
      "replies": [
        {
          "id": 966739,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/11/2020 16:26:50",
          "content": "<p>These are good concerns. You are correct that in every real world situation we must adjust our model to maximize some metric which will lower unacceptable errors (reduce false positive, reduce false negative, etc). Each situation is different and our employer will tell us what to maximize.</p>\n<p>In this competition, the leaderboard is our \"employer\", and our \"employer\" is asking us to maximize AUC</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 966753,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/11/2020 16:40:20",
          "content": "<p>False positives are (in most medical cases) often better than false negatives. Detecting non-malignant images as malignant will require the doctors to do an unnecessary investigation.</p>\n<p>Detecting a malignant image as benign will cost a human life.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970907,
          "author_name": "harshnagouda",
          "author_url": "",
          "post_date": "08/15/2020 01:54:41",
          "content": "<p>I see your point. In the end, we're just trying to do anything we can to improve in the race! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 976105,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/18/2020 16:31:44",
      "content": "<p>Below are the results of my upsample experiments on CNN. Once again, I used my embedding extraction notebook <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">here</a>. After extracting embeddings, we can model them with an MLP. This simulates modeling with a CNN and freezing the backbone. The MLP head is as follows</p>\n<pre><code>def MakeModel():\n    inp = tf.keras.layers.Input((1280,))\n    x = tf.keras.layers.Dense(1,activation='sigmoid')(inp)\n    model = tf.keras.Model(inputs=inp,outputs=x)\n    model.compile(optimizer=tf.keras.optimizers.SGD(),\n           loss=tf.keras.losses.BinaryCrossentropy(),\n           metrics=['AUC'])\n    return model\n</code></pre>\n<p>Next we train with the same code posted above for upsample, and we use early stopping on the MLP</p>\n<pre><code>        model = MakeModel()\n\n        cc = tf.keras.callbacks.EarlyStopping(\n            monitor='val_auc', min_delta=0, patience=2, verbose=0, mode='max',\n            baseline=None, restore_best_weights=True\n        )\n\n        model.fit(X_more, y_more, epochs=30, verbose=0,\n             validation_data = (embed[idxV2,],train.target.values[idxV2]),callbacks=[cc] )\n</code></pre>\n<p>Here is the affect of upsample</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8a784af724af627812c6c997645e23c8%2Fa1.png?generation=1597768115192486&amp;alt=media\" alt=\"\"></p>\n<p>And here is the result of 50 experiments. Mean AUC increase is 0.026 with std 0.0036</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcd0388046eac1e5ef4286e855fcc5cda%2Fa2.png?generation=1597768137923981&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962467,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "08/08/2020 06:29:37",
      "content": "<p>Interesting result, thank you for sharing! Would not have thought that such a strong upsample (25x) would work best.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 962486,
      "author_name": "group16",
      "author_url": "",
      "post_date": "08/08/2020 06:52:50",
      "content": "<p>Thanks Chris! Wouldn't setting <code>class_weights</code> achieve the same result?</p>",
      "votes": null,
      "replies": [
        {
          "id": 962537,
          "author_name": "aybatov",
          "author_url": "",
          "post_date": "08/08/2020 08:06:17",
          "content": "<p>Would be nice to test two approaches. Think when we upsample then model see the same melanoma images after different augmentations and thats could lead to better result than weights.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962637,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2020 09:58:25",
          "content": "<p>Neither RAPIDS nor scikit-learn KNN support class weights.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962643,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/08/2020 10:04:32",
          "content": "<p>I was referring to the graph that plots the CNN performance in function of the upsampling ratio. Not the artificial KNN example.\nEDIT: i mis-interpreted the plot. Looking at the y-axis, it might have been KNN there as well.</p>\n\n<p>After some reading, they are not equivalent. By duplicating images, the probability that one or more are included in a batch increases. This is not the case for custom class weights.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962656,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/08/2020 10:11:44",
          "content": "<p>The graph plots KNN performance, not CNN.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962825,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/08/2020 13:13:32",
          "content": "<blockquote>\n  <p>Wouldn't setting class_weights achieve the same result?</p>\n</blockquote>\n<p>No. Using <code>class_weights</code> will not change the rank of kNN predictions and therefore will not change the AUC. In my example above, using <code>class_weights = {0:1, 1:2}</code> will make predictions <code>0.0, 4/7, 6/8, 1.0</code> and obtain <code>AUC = 0.75</code> (the same AUC as using <code>class_weights = {0:1, 1:X} for any X</code>) . </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 964531,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/10/2020 01:22:47",
          "content": "<p>How do you use class weight with  rapids KNeighborsClassifier?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965724,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/10/2020 19:48:07",
          "content": "<p>With kNN, (if i understand class weight correctly) you can change class weight using post process formula. Given <code>prediction = p</code>, with <code>kNN parameter = k</code>, and class weight <code>{0:1, 1:w}</code>. Then you can apply class weight with</p>\n<pre><code>new pred = w*p*k / (w*p*k + (1-p)*k)\n</code></pre>\n<p>This is a rank preserving transformation, so it will not alter AUC. Regarding kNN, you need to add multiple copies of malignant to affect AUC.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965727,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/10/2020 19:50:02",
          "content": "<p><a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a> I'm not sure if you saw my answer to your question above. It is interesting to note that <code>class weight</code> and adding more copies of malignant is not the same for kNN. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 965730,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/10/2020 19:53:18",
          "content": "<p>I saw it <a href=\"/cdeotte\">@cdeotte</a> (but forgot to upvote, shame on me)! Our team is on it ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 966315,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/11/2020 10:23:36",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Your math seems correct to me indeed:</p>\n<p><code>new pred = w*p*k / (w*p*k + (1-p)*k)</code></p>\n<p>Thanks for checking.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 962732,
      "author_name": "faizaislam",
      "author_url": "",
      "post_date": "08/08/2020 11:47:29",
      "content": "<p>Best work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 964504,
      "author_name": "",
      "author_url": "",
      "post_date": "08/10/2020 00:33:28",
      "content": "<p>nice</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 965256,
      "author_name": "fantasticbobo",
      "author_url": "",
      "post_date": "08/10/2020 13:54:10",
      "content": "<p>Amazing Job! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 965879,
      "author_name": "srikanthpotukuchi",
      "author_url": "",
      "post_date": "08/10/2020 23:34:35",
      "content": "<p>Is the 0.003 worth it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 966761,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/11/2020 16:43:55",
          "content": "<p>Of course. That is a margin of around 25 places on the top of the LB.  And probably a few 100 places lower on the LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 966774,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/11/2020 16:49:36",
          "content": "<p>Yes. And in real life an increase in 0.003 can help us save more lives.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "962242": "People are wondering whether upsample can improve AUC, so I ran an experiment...\n# Upsampling Malignant Can Improve AUC for kNN !\nThe following experiment shows that upsample can increase AUC for kNN in Melanoma Comp by at least 0.003. The y-axis is triple stratified CV AUC score for kNN. The x-axis is malignant upsample ratio. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fee9a2d33de8e3904f032b69a56949406%2Faucc.png?generation=1596844592049793&amp;alt=media)\n\nThis model uses just the 2020 comp data which (after removing duplicates) has 32692 images with 584 malignant for a proportion of 1.79% (when upsample=0). Then upsample 25x means that we add 25 additional copies of each malignant image. (i.e. we add `14600 = 25*584` images). Using 25x upsample increases the proportion to 32.1% malignant.\n\n# Upsampling Malignant Can Improve AUC for CNN !\nNote that every model is different. So the AUC gain you observe in your model and the optimal upsample rate for your model may be different. In order to confirm that upsampling increases AUC for CNN. I also conducting a paired t-test for CNN and found that upsample increases AUC for CNN. I will post the code and CNN experiment details after the comp finishes. Below I post the code and experiment details for my kNN experiments.\n\n# Statistically Significant\nJust to be sure this wasn't luck, I ran this kNN experiment 100 times with 100 different triple stratified 5 KFold seeds. The average AUC increase is 0.002 with STD 0.0012. We are 99.99999% confident that upsample increases AUC for kNN. (And based on my CNN experiments, we are 99.99999% confident that upsample increases AUC for CNN). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fdf48581148ee4265602dbd89b4f4ff63%2Fhist.png?generation=1596895000816762&amp;alt=media)\n\n# kNN Experiment\nWe can add upsampling to my triple stratified RAPIDS cuML kNN notebook [here][1]. The model uses image embeddings from EfficientNetB0 and 256x256 images. We modify the below code to the below below code. To generate the plot above, we use the for-loop `for UP in np.arange(0,65,5):`\n\n    # CREATE TRAIN AND VALIDATION SUBSETS\n    idxT2 = train.loc[train.tfrecord.isin(idxT)].index.values #2020 train\n    idxV2 = train.loc[train.tfrecord.isin(idxV)].index.values #2020 valid\n    \n    # MODEL\n    model = cuml.neighbors.KNeighborsClassifier(n_neighbors=299)\n    model.fit(embed[idxT2,],train.target.values[idxT2])\n\nWe change the above code to the below code\n\n    # HOW MANY TIMES TO UPSAMPLE\n    UP = 25\n\n    # CREATE TRAIN AND VALIDATION SUBSETS\n    idxT2 = train.loc[train.tfrecord.isin(idxT)].index.values #2020 train\n    idxV2 = train.loc[train.tfrecord.isin(idxV)].index.values #2020 valid\n    \n    # UPSAMPLE\n    idxM = np.where( train.target.values[idxT2]==1 )[0]; idxM2 = idxT2[idxM]\n    X_more = np.zeros((len(idxT2)+len(idxM2)*UP,embed.shape[1])) \n    X_more[:len(idxT2),] = embed[idxT2,]\n    y_more = np.zeros((len(idxT2)+len(idxM2)*UP))\n    y_more[:len(idxT2)] = train.target.values[idxT2]\n    for k in range(UP): \n        X_more[len(idxT2)+len(idxM2)*k:len(idxT2)+len(idxM2)*(k+1),] = embed[idxM2,]\n        y_more[len(idxT2)+len(idxM2)*k:len(idxT2)+len(idxM2)*(k+1)] = train.target.values[idxM2]\n    \n    # MODEL\n    model = cuml.neighbors.KNeighborsClassifier(n_neighbors=299)\n    model.fit(X_more,y_more)\n\n# How Upsample Improves AUC\nUpsample affects AUC by increasing some predictions more than others. Here's an example showing upsample improving AUC. In the plot below, there are 4 unknown data points labeled `A, B, C, D`. What would you guess are their true labels? Yellow denotes `target=1` and blue denotes `target=0`. \n## Without Upsample - AUC = 0.75\nIf you use kNN with `k=5`, then we will predict probabilities `0.0, 0.4, 0.6, 1.0` for points `A, B, C, D` respectively. It turns out that the true labels are `0, 1, 0, 1` so our `AUC = 0.75` because two predictions are out of ranked order.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fc897605c98e78b26c634e6659d87f8cb%2Fknn1.jpg?generation=1596844755349238&amp;alt=media)\n\n## With Upsample - AUC = 1.00\nIf we upsample the malignant samples by 2x and then use kNN with `k=5`, then we predict `0.0, 0.8, 0.6, 1.0`. Since the true labels are `0, 1, 0, 1` our `AUC = 1.0` because no predictions are out of ranked order.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F4389b112dbdef9497ef7fab2b6312177%2Fknn2.jpg?generation=1596845180779774&amp;alt=media)\n\n# How To Upsample Discussion\nThere is a discussion explaining how to upsample [here][3]\n# How To Upsample Notebook\nThere is a starter notebook demonstrating upsample [here][4]\n\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n[3]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/169139\n[4]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout",
    "962467": "Interesting result, thank you for sharing! Would not have thought that such a strong upsample (25x) would work best.",
    "962486": "Thanks Chris! Wouldn't setting `class_weights` achieve the same result?",
    "962537": "Would be nice to test two approaches. Think when we upsample then model see the same melanoma images after different augmentations and thats could lead to better result than weights.",
    "962573": "This is interesting, but it does not mean upsampling would improve CNN roc-auc.",
    "962623": "The datasets are also the hyper-parameters in this competition. 😭😭😭",
    "962634": "I may be wrong, but improvements like 0.003 AUC are well inside confidence intervals for this competition and most likely occur by chance. At least for me, standard deviations for model scores are around 0.01 for 5-fold CV and I would not consider 0.003 as a definite improvement.",
    "962637": "Neither RAPIDS nor scikit-learn KNN support class weights.",
    "962639": "It depend son your CV setting.  For mine, a 0.003 CV improvement translates to a LB improvement so far.",
    "962643": "I was referring to the graph that plots the CNN performance in function of the upsampling ratio. Not the artificial KNN example.\nEDIT: i mis-interpreted the plot. Looking at the y-axis, it might have been KNN there as well.\n\nAfter some reading, they are not equivalent. By duplicating images, the probability that one or more are included in a batch increases. This is not the case for custom class weights.",
    "962656": "The graph plots KNN performance, not CNN.",
    "962732": "Best work",
    "962825": "&gt; Wouldn't setting class_weights achieve the same result?\n\nNo. Using `class_weights` will not change the rank of kNN predictions and therefore will not change the AUC. In my example above, using `class_weights = {0:1, 1:2}` will make predictions `0.0, 4/7, 6/8, 1.0` and obtain `AUC = 0.75` (the same AUC as using `class_weights = {0:1, 1:X} for any X`) .",
    "962865": "I ran this experiment 100 times with 100 different seeds. The average AUC increase is 0.002 with STD 0.0012. These results are statistically significant with p&lt;0.0001 (paired one sided t test). We are 99.99999% confident that upsampling increases AUC.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F64fff89f470211965a6c6d5c265200e3%2Fhist.png?generation=1596894830615660&amp;alt=media)",
    "963638": "Thanks for the experiments. Still, I think it may not be  completely relevant to CNNs",
    "964399": "Hi Chris, thanks for running this analysis! What statistical test did you use to compute your p-value? Given an observed increase in AUC of 0.002 with standard deviation 0.0012, what was the null hypothesis that this was compared to?",
    "964504": "nice",
    "964517": "Hi @jaronthompson , the null hypothesis is `H0 : mean AUC difference = 0` (i.e AUC stays the same) and alternative hypothesis is `HA : mean AUC difference &gt; 0` (i.e. AUC increases). We then perform a one-sided [paired t-test][1] and reject the null hypothesis with `p = &lt; .00001`\n\nThere's a greater than 99.99999% chance that upsample increases AUC!\n\nFor each of 100 pairs of experiments, i choose a different seed for triple stratified KFold. Then I trained with and without upsampling and calculated the difference in AUC score. The mean difference was `0.00207` with standard deviation `0.00122`. Using the online t-test calulator [here][2], we see that the `p = &lt; .00001`.\n\n\n\n[1]: https://online.stat.psu.edu/stat415/lesson/10/10.3\n[2]: https://www.usablestats.com/calcs/1samplet&amp;summary=1",
    "964531": "How do you use class weight with  rapids KNeighborsClassifier?",
    "964575": "&gt; Still, I think it may not be completely relevant to CNNs\n\nI also was curious about CNNs, so I repeated the paired t test on CNNs and observed an AUC increase with p&lt;0.0001. We are 99.99999% confident that upsampling increases AUC for CNNs in Melanoma comp! I will post the experiment results after the comp finishes.",
    "965256": "Amazing Job!",
    "965260": "Great, thanks!",
    "965304": "Chris, thanks, this is the first time someone says upsampling improves roc-auc for CNN.  What improvement did you see in your experiment?",
    "965724": "With kNN, (if i understand class weight correctly) you can change class weight using post process formula. Given `prediction = p`, with `kNN parameter = k`, and class weight `{0:1, 1:w}`. Then you can apply class weight with\n\n    new pred = w*p*k / (w*p*k + (1-p)*k)\n\nThis is a rank preserving transformation, so it will not alter AUC. Regarding kNN, you need to add multiple copies of malignant to affect AUC.",
    "965727": "group16 I'm not sure if you saw my answer to your question above. It is interesting to note that `class weight` and adding more copies of malignant is not the same for kNN.",
    "965730": "I saw it @cdeotte (but forgot to upvote, shame on me)! Our team is on it ;)",
    "965879": "Is the 0.003 worth it?",
    "966315": "cdeotte Your math seems correct to me indeed:\n\n `new pred = w*p*k / (w*p*k + (1-p)*k)`\n\nThanks for checking.",
    "966699": "This is a very good analysis. However, I am still debating on using upsampling (or undersampling) to deal with the imbalance. For me, the focal loss strategy has been yielding some interesting results, but the reason I am refraining from upsampling is to match the real-world scenario as accurately as possible. In the real world, there are, of course, significantly higher non-malignant (benign) cases than those of malignant. So if the model is now learning more about malignant cases, it might as well (due to small factors in images) detect (falsely...?) non-malignant images as malignant. Being a beginner, I am flooded with questions, and I might be very wrong too.",
    "966739": "These are good concerns. You are correct that in every real world situation we must adjust our model to maximize some metric which will lower unacceptable errors (reduce false positive, reduce false negative, etc). Each situation is different and our employer will tell us what to maximize.\n\nIn this competition, the leaderboard is our \"employer\", and our \"employer\" is asking us to maximize AUC",
    "966753": "False positives are (in most medical cases) often better than false negatives. Detecting non-malignant images as malignant will require the doctors to do an unnecessary investigation.\n\nDetecting a malignant image as benign will cost a human life.",
    "966761": "Of course. That is a margin of around 25 places on the top of the LB.  And probably a few 100 places lower on the LB.",
    "966774": "Yes. And in real life an increase in 0.003 can help us save more lives.",
    "970907": "I see your point. In the end, we're just trying to do anything we can to improve in the race!",
    "976105": "Below are the results of my upsample experiments on CNN. Once again, I used my embedding extraction notebook [here][1]. After extracting embeddings, we can model them with an MLP. This simulates modeling with a CNN and freezing the backbone. The MLP head is as follows\n\n    def MakeModel():\n        inp = tf.keras.layers.Input((1280,))\n        x = tf.keras.layers.Dense(1,activation='sigmoid')(inp)\n        model = tf.keras.Model(inputs=inp,outputs=x)\n        model.compile(optimizer=tf.keras.optimizers.SGD(),\n               loss=tf.keras.losses.BinaryCrossentropy(),\n               metrics=['AUC'])\n        return model\n\nNext we train with the same code posted above for upsample, and we use early stopping on the MLP\n\n           \n            model = MakeModel()\n\n            cc = tf.keras.callbacks.EarlyStopping(\n                monitor='val_auc', min_delta=0, patience=2, verbose=0, mode='max',\n                baseline=None, restore_best_weights=True\n            )\n        \n            model.fit(X_more, y_more, epochs=30, verbose=0,\n                 validation_data = (embed[idxV2,],train.target.values[idxV2]),callbacks=[cc] )\n\nHere is the affect of upsample\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2F8a784af724af627812c6c997645e23c8%2Fa1.png?generation=1597768115192486&alt=media)\n\nAnd here is the result of 50 experiments. Mean AUC increase is 0.026 with std 0.0036\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1723677%2Fcd0388046eac1e5ef4286e855fcc5cda%2Fa2.png?generation=1597768137923981&alt=media)\n\n[1]: https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates"
  },
  "source": "meta"
}