{
  "id": 331389,
  "title": "Interesting paper: SCARF: SELF-SUPERVISED CONTRASTIVE LEARNING USING RANDOM FEATURE CORRUPTION",
  "url": "/competitions/amex-default-prediction/discussion/331389",
  "author_name": "",
  "post_date": "2022-06-17T04:05:37.052504200Z",
  "votes": 20,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Link paper: <a href=\"https://arxiv.org/pdf/2106.15147.pdf\" target=\"_blank\">SCARF: SELF-SUPERVISED CONTRASTIVE LEARNING\nUSING RANDOM FEATURE CORRUPTION - Published as a conference paper at ICLR 2022</a></p>\n<p><strong>Abstract</strong>: Self-supervised contrastive representation learning has proved incredibly successful in the vision and natural language domains, enabling state-of-the-art performance with orders of magnitude less labeled data. However, such methods are domain-specific and little has been done to leverage this technique on real-world tabular datasets. We propose SCARF, a simple, widely-applicable technique for contrastive learning, where views are formed by corrupting a random subset of features. When applied to pre-train deep neural networks on the 69 real-world, tabular classification<br>\ndatasets from the OpenML-CC18 benchmark, SCARF not only improves classification accuracy in the fully-supervised setting but does so also in the presence of label noise and in the semi-supervised setting where only a fraction of the available training data is labeled.</p>\n<p><a href=\"https://postimg.cc/V0LGzXPQ\" target=\"_blank\"><img src=\"https://i.postimg.cc/d3Gzm9W1/abc.png\" alt=\"abc.png\"></a></p>\n<p><strong>Summary</strong>: SimCLR for tabular data</p>\n<ol>\n<li>Corruption of the data is really simple</li>\n<li>Replace random features with any value from the distribution of that feature</li>\n</ol>\n<p><strong>Baselines. We use the following baselines</strong></p>\n<p>We can see many commonly used techniques in computer vision and NLP listed here, you may also have seen or used in contests on Kaggle</p>\n<ol>\n<li>Label smoothing: We use a weight of 0.1 on the smoothing term.</li>\n<li>Dropout. We use standard dropout (Srivastava et al., 2014) using rate 0.04 on all layers.</li>\n<li>Mixup (Zhang et al., 2017), using α = 0.2.</li>\n<li>Autoencoders (Rumelhart et al., 1985). We use this as our key ablative pre-training baseline.<br>\nWe use the classical autoencoder (“no noise AE”), the denoising autoencoder (Vincent<br>\net al., 2008; 2010) using Gaussian additive noise (“add. noise AE”) as well as SCARF’s<br>\ncorruption method (“SCARF AE”). We use MSE for the reconstruction loss. We try both<br>\npre-training and co-training with the supervised task, and when co-training, we add 0.1 times<br>\nthe autoencoder loss to the supervised objective. We discuss co-training in the Appendix as<br>\nit is less effective than pre-training.</li>\n<li>SCARF data-augmentation. In order to isolate the effect of our proposed feature corruption<br>\ntechnique, we skip pre-training and instead train on the corrupted inputs during supervised<br>\nfine-tuning. We discuss results for this baseline in the Appendix as it is less effective than<br>\nthe others.</li>\n<li>Discriminative SCARF. Here, our pre-training objective is to discriminate between original<br>\ninput features and their counterparts that have been corrupted using our proposed technique.<br>\nTo this end, we update our pre-training head network to include a final linear projection and<br>\nswap the InfoNCE with a binary logistic loss. We use classification error, not logistic loss,<br>\nas the validation metric for early stopping, as we found it to perform slightly better.</li>\n<li>Self-distillation (Hinton et al., 2015; Zhang et al., 2019a). We first train the model on the<br>\nlabeled data and then train the final model on both the labeled and unlabeled data using the<br>\nfirst models’ predictions as soft labels for both sets.</li>\n<li>Deep k-NN (Bahri et al., 2020), a recently proposed method for label noise. We set k = 50.</li>\n<li>Bi-tempered loss (Amid et al., 2019), a recently proposed method for label noise. We use 5<br>\niterations, t1 = 0.8, and t2 = 1.2.</li>\n<li>Self-training: A classical semi-supervised method<br>\n– each iteration, we train on pseudo-labeled data (initialized to be the original labeled dataset)<br>\nand add highly confident predictions to the training set using the prediction as the label. We<br>\nthen train our final model on the final dataset. We use a softmax prediction threshold of 0.75<br>\nand run for 10 iterations.</li>\n<li>Tri-training (Zhou &amp; Li, 2005). Like self-training, but using three models with different<br>\ninitial labeled data via bootstrap sampling. Each iteration, every model’s training set is<br>\nupdated by adding only unlabeled points whose predictions made by the other two models<br>\nagree. It was shown to be competitive in modern semi-supervised NLP tasks (Ruder &amp;<br>\nPlank, 2018). We use same hyperparameters as self-training</li>\n</ol>\n<p><strong>SCARF PRE-TRAINING IMPROVES PERFORMANCE IN THE PRESENCE TO LABEL NOISE</strong></p>\n<p><a href=\"https://postimg.cc/phK4s2y3\" target=\"_blank\"><img src=\"https://i.postimg.cc/vBKyHTy8/9-Table1-1.png\" alt=\"9-Table1-1.png\"></a></p>\n<p>We can use some ideas from this paper for the competition, any comments are welcome!</p>",
  "messages": [
    {
      "id": "1823109",
      "postDate": "06/17/2022 04:05:37",
      "content": "<p>Link paper: <a href=\"https://arxiv.org/pdf/2106.15147.pdf\" target=\"_blank\">SCARF: SELF-SUPERVISED CONTRASTIVE LEARNING\nUSING RANDOM FEATURE CORRUPTION - Published as a conference paper at ICLR 2022</a></p>\n<p><strong>Abstract</strong>: Self-supervised contrastive representation learning has proved incredibly successful in the vision and natural language domains, enabling state-of-the-art performance with orders of magnitude less labeled data. However, such methods are domain-specific and little has been done to leverage this technique on real-world tabular datasets. We propose SCARF, a simple, widely-applicable technique for contrastive learning, where views are formed by corrupting a random subset of features. When applied to pre-train deep neural networks on the 69 real-world, tabular classification<br>\ndatasets from the OpenML-CC18 benchmark, SCARF not only improves classification accuracy in the fully-supervised setting but does so also in the presence of label noise and in the semi-supervised setting where only a fraction of the available training data is labeled.</p>\n<p><a href=\"https://postimg.cc/V0LGzXPQ\" target=\"_blank\"><img src=\"https://i.postimg.cc/d3Gzm9W1/abc.png\" alt=\"abc.png\"></a></p>\n<p><strong>Summary</strong>: SimCLR for tabular data</p>\n<ol>\n<li>Corruption of the data is really simple</li>\n<li>Replace random features with any value from the distribution of that feature</li>\n</ol>\n<p><strong>Baselines. We use the following baselines</strong></p>\n<p>We can see many commonly used techniques in computer vision and NLP listed here, you may also have seen or used in contests on Kaggle</p>\n<ol>\n<li>Label smoothing: We use a weight of 0.1 on the smoothing term.</li>\n<li>Dropout. We use standard dropout (Srivastava et al., 2014) using rate 0.04 on all layers.</li>\n<li>Mixup (Zhang et al., 2017), using α = 0.2.</li>\n<li>Autoencoders (Rumelhart et al., 1985). We use this as our key ablative pre-training baseline.<br>\nWe use the classical autoencoder (“no noise AE”), the denoising autoencoder (Vincent<br>\net al., 2008; 2010) using Gaussian additive noise (“add. noise AE”) as well as SCARF’s<br>\ncorruption method (“SCARF AE”). We use MSE for the reconstruction loss. We try both<br>\npre-training and co-training with the supervised task, and when co-training, we add 0.1 times<br>\nthe autoencoder loss to the supervised objective. We discuss co-training in the Appendix as<br>\nit is less effective than pre-training.</li>\n<li>SCARF data-augmentation. In order to isolate the effect of our proposed feature corruption<br>\ntechnique, we skip pre-training and instead train on the corrupted inputs during supervised<br>\nfine-tuning. We discuss results for this baseline in the Appendix as it is less effective than<br>\nthe others.</li>\n<li>Discriminative SCARF. Here, our pre-training objective is to discriminate between original<br>\ninput features and their counterparts that have been corrupted using our proposed technique.<br>\nTo this end, we update our pre-training head network to include a final linear projection and<br>\nswap the InfoNCE with a binary logistic loss. We use classification error, not logistic loss,<br>\nas the validation metric for early stopping, as we found it to perform slightly better.</li>\n<li>Self-distillation (Hinton et al., 2015; Zhang et al., 2019a). We first train the model on the<br>\nlabeled data and then train the final model on both the labeled and unlabeled data using the<br>\nfirst models’ predictions as soft labels for both sets.</li>\n<li>Deep k-NN (Bahri et al., 2020), a recently proposed method for label noise. We set k = 50.</li>\n<li>Bi-tempered loss (Amid et al., 2019), a recently proposed method for label noise. We use 5<br>\niterations, t1 = 0.8, and t2 = 1.2.</li>\n<li>Self-training: A classical semi-supervised method<br>\n– each iteration, we train on pseudo-labeled data (initialized to be the original labeled dataset)<br>\nand add highly confident predictions to the training set using the prediction as the label. We<br>\nthen train our final model on the final dataset. We use a softmax prediction threshold of 0.75<br>\nand run for 10 iterations.</li>\n<li>Tri-training (Zhou &amp; Li, 2005). Like self-training, but using three models with different<br>\ninitial labeled data via bootstrap sampling. Each iteration, every model’s training set is<br>\nupdated by adding only unlabeled points whose predictions made by the other two models<br>\nagree. It was shown to be competitive in modern semi-supervised NLP tasks (Ruder &amp;<br>\nPlank, 2018). We use same hyperparameters as self-training</li>\n</ol>\n<p><strong>SCARF PRE-TRAINING IMPROVES PERFORMANCE IN THE PRESENCE TO LABEL NOISE</strong></p>\n<p><a href=\"https://postimg.cc/phK4s2y3\" target=\"_blank\"><img src=\"https://i.postimg.cc/vBKyHTy8/9-Table1-1.png\" alt=\"9-Table1-1.png\"></a></p>\n<p>We can use some ideas from this paper for the competition, any comments are welcome!</p>",
      "rawMarkdown": "Link paper: [SCARF: SELF-SUPERVISED CONTRASTIVE LEARNING\nUSING RANDOM FEATURE CORRUPTION - Published as a conference paper at ICLR 2022](https://arxiv.org/pdf/2106.15147.pdf)\n\n**Abstract**: Self-supervised contrastive representation learning has proved incredibly successful in the vision and natural language domains, enabling state-of-the-art performance with orders of magnitude less labeled data. However, such methods are domain-specific and little has been done to leverage this technique on real-world tabular datasets. We propose SCARF, a simple, widely-applicable technique for contrastive learning, where views are formed by corrupting a random subset of features. When applied to pre-train deep neural networks on the 69 real-world, tabular classification\ndatasets from the OpenML-CC18 benchmark, SCARF not only improves classification accuracy in the fully-supervised setting but does so also in the presence of label noise and in the semi-supervised setting where only a fraction of the available training data is labeled.\n\n[![abc.png](https://i.postimg.cc/d3Gzm9W1/abc.png)](https://postimg.cc/V0LGzXPQ)\n\n**Summary**: SimCLR for tabular data\n1. Corruption of the data is really simple\n2. Replace random features with any value from the distribution of that feature\n\n**Baselines. We use the following baselines**\n\nWe can see many commonly used techniques in computer vision and NLP listed here, you may also have seen or used in contests on Kaggle\n\n1. Label smoothing: We use a weight of 0.1 on the smoothing term.\n2. Dropout. We use standard dropout (Srivastava et al., 2014) using rate 0.04 on all layers.\n3. Mixup (Zhang et al., 2017), using α = 0.2.\n4. Autoencoders (Rumelhart et al., 1985). We use this as our key ablative pre-training baseline.\nWe use the classical autoencoder (“no noise AE”), the denoising autoencoder (Vincent\net al., 2008; 2010) using Gaussian additive noise (“add. noise AE”) as well as SCARF’s\ncorruption method (“SCARF AE”). We use MSE for the reconstruction loss. We try both\npre-training and co-training with the supervised task, and when co-training, we add 0.1 times\nthe autoencoder loss to the supervised objective. We discuss co-training in the Appendix as\nit is less effective than pre-training.\n5. SCARF data-augmentation. In order to isolate the effect of our proposed feature corruption\ntechnique, we skip pre-training and instead train on the corrupted inputs during supervised\nfine-tuning. We discuss results for this baseline in the Appendix as it is less effective than\nthe others.\n6. Discriminative SCARF. Here, our pre-training objective is to discriminate between original\ninput features and their counterparts that have been corrupted using our proposed technique.\nTo this end, we update our pre-training head network to include a final linear projection and\nswap the InfoNCE with a binary logistic loss. We use classification error, not logistic loss,\nas the validation metric for early stopping, as we found it to perform slightly better.\n7. Self-distillation (Hinton et al., 2015; Zhang et al., 2019a). We first train the model on the\nlabeled data and then train the final model on both the labeled and unlabeled data using the\nfirst models’ predictions as soft labels for both sets.\n8. Deep k-NN (Bahri et al., 2020), a recently proposed method for label noise. We set k = 50.\n9. Bi-tempered loss (Amid et al., 2019), a recently proposed method for label noise. We use 5\niterations, t1 = 0.8, and t2 = 1.2.\n10. Self-training: A classical semi-supervised method\n– each iteration, we train on pseudo-labeled data (initialized to be the original labeled dataset)\nand add highly confident predictions to the training set using the prediction as the label. We\nthen train our final model on the final dataset. We use a softmax prediction threshold of 0.75\nand run for 10 iterations.\n11. Tri-training (Zhou & Li, 2005). Like self-training, but using three models with different\ninitial labeled data via bootstrap sampling. Each iteration, every model’s training set is\nupdated by adding only unlabeled points whose predictions made by the other two models\nagree. It was shown to be competitive in modern semi-supervised NLP tasks (Ruder &\nPlank, 2018). We use same hyperparameters as self-training\n\n**SCARF PRE-TRAINING IMPROVES PERFORMANCE IN THE PRESENCE TO LABEL NOISE**\n\n[![9-Table1-1.png](https://i.postimg.cc/vBKyHTy8/9-Table1-1.png)](https://postimg.cc/phK4s2y3)\n\nWe can use some ideas from this paper for the competition, any comments are welcome!",
      "votes": null
    },
    {
      "id": "1823581",
      "postDate": "06/17/2022 13:58:44",
      "content": "<p>Great share, It might be a way to improve, we should try it</p>",
      "rawMarkdown": "Great share, It might be a way to improve, we should try it",
      "votes": null
    },
    {
      "id": "1824899",
      "postDate": "06/18/2022 17:57:21",
      "content": "<p>Good job, thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> </p>",
      "rawMarkdown": "Good job, thanks @duykhanh99",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1823581,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "06/17/2022 13:58:44",
      "content": "<p>Great share, It might be a way to improve, we should try it</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1824899,
      "author_name": "saberghaderi",
      "author_url": "",
      "post_date": "06/18/2022 17:57:21",
      "content": "<p>Good job, thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1823109": "Link paper: [SCARF: SELF-SUPERVISED CONTRASTIVE LEARNING\nUSING RANDOM FEATURE CORRUPTION - Published as a conference paper at ICLR 2022](https://arxiv.org/pdf/2106.15147.pdf)\n\n**Abstract**: Self-supervised contrastive representation learning has proved incredibly successful in the vision and natural language domains, enabling state-of-the-art performance with orders of magnitude less labeled data. However, such methods are domain-specific and little has been done to leverage this technique on real-world tabular datasets. We propose SCARF, a simple, widely-applicable technique for contrastive learning, where views are formed by corrupting a random subset of features. When applied to pre-train deep neural networks on the 69 real-world, tabular classification\ndatasets from the OpenML-CC18 benchmark, SCARF not only improves classification accuracy in the fully-supervised setting but does so also in the presence of label noise and in the semi-supervised setting where only a fraction of the available training data is labeled.\n\n[![abc.png](https://i.postimg.cc/d3Gzm9W1/abc.png)](https://postimg.cc/V0LGzXPQ)\n\n**Summary**: SimCLR for tabular data\n1. Corruption of the data is really simple\n2. Replace random features with any value from the distribution of that feature\n\n**Baselines. We use the following baselines**\n\nWe can see many commonly used techniques in computer vision and NLP listed here, you may also have seen or used in contests on Kaggle\n\n1. Label smoothing: We use a weight of 0.1 on the smoothing term.\n2. Dropout. We use standard dropout (Srivastava et al., 2014) using rate 0.04 on all layers.\n3. Mixup (Zhang et al., 2017), using α = 0.2.\n4. Autoencoders (Rumelhart et al., 1985). We use this as our key ablative pre-training baseline.\nWe use the classical autoencoder (“no noise AE”), the denoising autoencoder (Vincent\net al., 2008; 2010) using Gaussian additive noise (“add. noise AE”) as well as SCARF’s\ncorruption method (“SCARF AE”). We use MSE for the reconstruction loss. We try both\npre-training and co-training with the supervised task, and when co-training, we add 0.1 times\nthe autoencoder loss to the supervised objective. We discuss co-training in the Appendix as\nit is less effective than pre-training.\n5. SCARF data-augmentation. In order to isolate the effect of our proposed feature corruption\ntechnique, we skip pre-training and instead train on the corrupted inputs during supervised\nfine-tuning. We discuss results for this baseline in the Appendix as it is less effective than\nthe others.\n6. Discriminative SCARF. Here, our pre-training objective is to discriminate between original\ninput features and their counterparts that have been corrupted using our proposed technique.\nTo this end, we update our pre-training head network to include a final linear projection and\nswap the InfoNCE with a binary logistic loss. We use classification error, not logistic loss,\nas the validation metric for early stopping, as we found it to perform slightly better.\n7. Self-distillation (Hinton et al., 2015; Zhang et al., 2019a). We first train the model on the\nlabeled data and then train the final model on both the labeled and unlabeled data using the\nfirst models’ predictions as soft labels for both sets.\n8. Deep k-NN (Bahri et al., 2020), a recently proposed method for label noise. We set k = 50.\n9. Bi-tempered loss (Amid et al., 2019), a recently proposed method for label noise. We use 5\niterations, t1 = 0.8, and t2 = 1.2.\n10. Self-training: A classical semi-supervised method\n– each iteration, we train on pseudo-labeled data (initialized to be the original labeled dataset)\nand add highly confident predictions to the training set using the prediction as the label. We\nthen train our final model on the final dataset. We use a softmax prediction threshold of 0.75\nand run for 10 iterations.\n11. Tri-training (Zhou & Li, 2005). Like self-training, but using three models with different\ninitial labeled data via bootstrap sampling. Each iteration, every model’s training set is\nupdated by adding only unlabeled points whose predictions made by the other two models\nagree. It was shown to be competitive in modern semi-supervised NLP tasks (Ruder &\nPlank, 2018). We use same hyperparameters as self-training\n\n**SCARF PRE-TRAINING IMPROVES PERFORMANCE IN THE PRESENCE TO LABEL NOISE**\n\n[![9-Table1-1.png](https://i.postimg.cc/vBKyHTy8/9-Table1-1.png)](https://postimg.cc/phK4s2y3)\n\nWe can use some ideas from this paper for the competition, any comments are welcome!",
    "1823581": "Great share, It might be a way to improve, we should try it",
    "1824899": "Good job, thanks @duykhanh99"
  },
  "source": "meta"
}