{
  "id": 241770,
  "title": "How to create non-leaking folds",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/241770",
  "author_name": "Darek Kłeczek",
  "post_date": "2021-05-26T05:02:25.724000",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I'd like to share a methodology to create non-leaking folds. I released notebooks we used in this challenge, and will try to come up with a consolidated notebook. I think this approach should generalize well to any competition where we can obtain good embeddings for train examples. </p>\n<ol>\n<li><p>Obtain embeddings for each example in your training set. We used bestfitting's 2018 metric learning model for this.<br>\n<a href=\"https://www.kaggle.com/thedrcat/bestfitting-ml\" target=\"_blank\">https://www.kaggle.com/thedrcat/bestfitting-ml</a><br>\n<a href=\"https://www.kaggle.com/thedrcat/bestfitting-ml-public\" target=\"_blank\">https://www.kaggle.com/thedrcat/bestfitting-ml-public</a></p></li>\n<li><p>Cluster similar examples together. I use the approach from topic modeling (e.g. BERTopic) which first reduces the embedding dimensions using UMAP, and then clusters them with DBSCAN. <br>\n<a href=\"https://www.kaggle.com/thedrcat/knn-dataset\" target=\"_blank\">https://www.kaggle.com/thedrcat/knn-dataset</a></p></li>\n<li><p>Create folds ensuring each cluster is contained within the same fold. In this case, we also did iterative multilabel stratification to address class imbalance. <br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-kfold-split\" target=\"_blank\">https://www.kaggle.com/thedrcat/hpa-kfold-split</a></p></li>\n</ol>\n<p>I'd love to get feedback on this approach. Does it seem sound? Would you change or improve anything?</p>",
  "messages": [
    {
      "id": 1323219,
      "postDate": "2021-05-26T05:02:25.723Z",
      "content": "<p>I'd like to share a methodology to create non-leaking folds. I released notebooks we used in this challenge, and will try to come up with a consolidated notebook. I think this approach should generalize well to any competition where we can obtain good embeddings for train examples. </p>\n<ol>\n<li><p>Obtain embeddings for each example in your training set. We used bestfitting's 2018 metric learning model for this.<br>\n<a href=\"https://www.kaggle.com/thedrcat/bestfitting-ml\" target=\"_blank\">https://www.kaggle.com/thedrcat/bestfitting-ml</a><br>\n<a href=\"https://www.kaggle.com/thedrcat/bestfitting-ml-public\" target=\"_blank\">https://www.kaggle.com/thedrcat/bestfitting-ml-public</a></p></li>\n<li><p>Cluster similar examples together. I use the approach from topic modeling (e.g. BERTopic) which first reduces the embedding dimensions using UMAP, and then clusters them with DBSCAN. <br>\n<a href=\"https://www.kaggle.com/thedrcat/knn-dataset\" target=\"_blank\">https://www.kaggle.com/thedrcat/knn-dataset</a></p></li>\n<li><p>Create folds ensuring each cluster is contained within the same fold. In this case, we also did iterative multilabel stratification to address class imbalance. <br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-kfold-split\" target=\"_blank\">https://www.kaggle.com/thedrcat/hpa-kfold-split</a></p></li>\n</ol>\n<p>I'd love to get feedback on this approach. Does it seem sound? Would you change or improve anything?</p>",
      "rawMarkdown": "I'd like to share a methodology to create non-leaking folds. I released notebooks we used in this challenge, and will try to come up with a consolidated notebook. I think this approach should generalize well to any competition where we can obtain good embeddings for train examples. \n\n1. Obtain embeddings for each example in your training set. We used bestfitting's 2018 metric learning model for this.\nhttps://www.kaggle.com/thedrcat/bestfitting-ml\nhttps://www.kaggle.com/thedrcat/bestfitting-ml-public\n\n2. Cluster similar examples together. I use the approach from topic modeling (e.g. BERTopic) which first reduces the embedding dimensions using UMAP, and then clusters them with DBSCAN. \nhttps://www.kaggle.com/thedrcat/knn-dataset\n\n3. Create folds ensuring each cluster is contained within the same fold. In this case, we also did iterative multilabel stratification to address class imbalance. \nhttps://www.kaggle.com/thedrcat/hpa-kfold-split\n\nI'd love to get feedback on this approach. Does it seem sound? Would you change or improve anything?\n",
      "votes": 4
    },
    {
      "id": 1323342,
      "postDate": "2021-05-26T07:02:16.430Z",
      "content": "<p>I did the same as you. However in hindsight I would split by cell line, which looks like a leakfree yet effortless approach</p>",
      "rawMarkdown": "I did the same as you. However in hindsight I would split by cell line, which looks like a leakfree yet effortless approach",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1323342,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2021-05-26T07:02:16.430000",
      "content": "<p>I did the same as you. However in hindsight I would split by cell line, which looks like a leakfree yet effortless approach</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1323219": "I'd like to share a methodology to create non-leaking folds. I released notebooks we used in this challenge, and will try to come up with a consolidated notebook. I think this approach should generalize well to any competition where we can obtain good embeddings for train examples. \n\n1. Obtain embeddings for each example in your training set. We used bestfitting's 2018 metric learning model for this.\nhttps://www.kaggle.com/thedrcat/bestfitting-ml\nhttps://www.kaggle.com/thedrcat/bestfitting-ml-public\n\n2. Cluster similar examples together. I use the approach from topic modeling (e.g. BERTopic) which first reduces the embedding dimensions using UMAP, and then clusters them with DBSCAN. \nhttps://www.kaggle.com/thedrcat/knn-dataset\n\n3. Create folds ensuring each cluster is contained within the same fold. In this case, we also did iterative multilabel stratification to address class imbalance. \nhttps://www.kaggle.com/thedrcat/hpa-kfold-split\n\nI'd love to get feedback on this approach. Does it seem sound? Would you change or improve anything?\n",
    "1323342": "I did the same as you. However in hindsight I would split by cell line, which looks like a leakfree yet effortless approach"
  }
}