{
  "id": 35107,
  "title": "validation strategy that worked",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/writeups/towards-empirically-stable-training-validation-str",
  "author_name": "",
  "post_date": "2017-06-22T11:20:34.541457Z",
  "votes": 33,
  "comment_count": 19,
  "views": 0,
  "content": "<p>The competition was very challanging because of the quality and the origin of the data. The main challange for us was how to make use of additional data and have a solid validation framework. </p>\n\n<p>First thing we noticed, that additional dataset had strong correlation with train images. Looking thorougly at each image we found, that training set included the best quality image of a single patient, and all other frames were put into additional data file.</p>\n\n<p>So to avoid simillar image leakage to our validation set we did the following:</p>\n\n<p>1) calculate image colour histograms.\n2) run k-means clustering with k = 100\n3) take random 20 clusters and use that for validation</p>\n\n<p>Its quite easy solution, but it really helped to merge simillar images into a cluster and we avoided validation set contamination with simillar images in training set. As a result we could use additional dataset without too much worries (we did put less weight on additional data in our final models).</p>\n\n<p>On top of that, we removed (by hand) all the blurry or not clear images in additional dataset (roughly 15% of additional images).</p>\n\n<p>In the end we had decent validation set, with a mix of original and additional dataset images. Our models on validation set scored 0.75-0.8 - pretty close to our final result, which I am very happy with!</p>\n\n<p>Special thanks to this kernel which I used to calculate histograms and k-means for our validation:\n<a href=\"https://www.kaggle.com/vfdev5/type-1-clustering\">https://www.kaggle.com/vfdev5/type-1-clustering</a></p>\n\n<p>It really deserves more upvotes:)</p>\n\n<p>Very interesting to hear, how other teams did validaiton. Please share:)</p>",
  "messages": [
    {
      "id": "194957",
      "postDate": "06/22/2017 11:20:34",
      "content": "<p>The competition was very challanging because of the quality and the origin of the data. The main challange for us was how to make use of additional data and have a solid validation framework. </p>\n\n<p>First thing we noticed, that additional dataset had strong correlation with train images. Looking thorougly at each image we found, that training set included the best quality image of a single patient, and all other frames were put into additional data file.</p>\n\n<p>So to avoid simillar image leakage to our validation set we did the following:</p>\n\n<p>1) calculate image colour histograms.\n2) run k-means clustering with k = 100\n3) take random 20 clusters and use that for validation</p>\n\n<p>Its quite easy solution, but it really helped to merge simillar images into a cluster and we avoided validation set contamination with simillar images in training set. As a result we could use additional dataset without too much worries (we did put less weight on additional data in our final models).</p>\n\n<p>On top of that, we removed (by hand) all the blurry or not clear images in additional dataset (roughly 15% of additional images).</p>\n\n<p>In the end we had decent validation set, with a mix of original and additional dataset images. Our models on validation set scored 0.75-0.8 - pretty close to our final result, which I am very happy with!</p>\n\n<p>Special thanks to this kernel which I used to calculate histograms and k-means for our validation:\n<a href=\"https://www.kaggle.com/vfdev5/type-1-clustering\">https://www.kaggle.com/vfdev5/type-1-clustering</a></p>\n\n<p>It really deserves more upvotes:)</p>\n\n<p>Very interesting to hear, how other teams did validaiton. Please share:)</p>",
      "rawMarkdown": "The competition was very challanging because of the quality and the origin of the data. The main challange for us was how to make use of additional data and have a solid validation framework. \n\nFirst thing we noticed, that additional dataset had strong correlation with train images. Looking thorougly at each image we found, that training set included the best quality image of a single patient, and all other frames were put into additional data file.\n\nSo to avoid simillar image leakage to our validation set we did the following:\n\n1) calculate image colour histograms.\n2) run k-means clustering with k = 100\n3) take random 20 clusters and use that for validation\n\nIts quite easy solution, but it really helped to merge simillar images into a cluster and we avoided validation set contamination with simillar images in training set. As a result we could use additional dataset without too much worries (we did put less weight on additional data in our final models).\n\nOn top of that, we removed (by hand) all the blurry or not clear images in additional dataset (roughly 15% of additional images).\n\nIn the end we had decent validation set, with a mix of original and additional dataset images. Our models on validation set scored 0.75-0.8 - pretty close to our final result, which I am very happy with!\n\nSpecial thanks to this kernel which I used to calculate histograms and k-means for our validation:\nhttps://www.kaggle.com/vfdev5/type-1-clustering\n\nIt really deserves more upvotes:)\n\nVery interesting to hear, how other teams did validaiton. Please share:)",
      "votes": null
    },
    {
      "id": "194958",
      "postDate": "06/22/2017 11:33:11",
      "content": "<p>I was just about to ask on the forum whether any of the top teams used the additional data as it seemed so far that many who posted solutions did not. </p>\n\n<p>Good to see that it worked for you! I thought I was doing something similar by clustering the train/additional images and only using one image from each cluster... I thought I was being pretty pedantic about it but I must not have done this well enough as my private LB score was way higher than what I was getting in validation. I expect that if I had ignored the additional set completely I may have ended up with something more reliable...</p>\n\n<p>Congratulations, by the way. Good job!</p>",
      "rawMarkdown": "I was just about to ask on the forum whether any of the top teams used the additional data as it seemed so far that many who posted solutions did not. \n\nGood to see that it worked for you! I thought I was doing something similar by clustering the train/additional images and only using one image from each cluster... I thought I was being pretty pedantic about it but I must not have done this well enough as my private LB score was way higher than what I was getting in validation. I expect that if I had ignored the additional set completely I may have ended up with something more reliable...\n\nCongratulations, by the way. Good job!",
      "votes": null
    },
    {
      "id": "194959",
      "postDate": "06/22/2017 11:37:42",
      "content": "<p>@raddar, thanks for sharing your strategy and congratulations ! \nHowever I did not get how did you composed training and validations sets. You clustered all train and additional in 100 clusters and picked a training dataset as 80 random clusters and validation dataset 20 others ? </p>\n\n<p>ps. glad that the kernel served</p>",
      "rawMarkdown": "raddar, thanks for sharing your strategy and congratulations ! \nHowever I did not get how did you composed training and validations sets. You clustered all train and additional in 100 clusters and picked a training dataset as 80 random clusters and validation dataset 20 others ? \n\nps. glad that the kernel served",
      "votes": null
    },
    {
      "id": "194960",
      "postDate": "06/22/2017 11:39:54",
      "content": "<p>You did reverse what we did :) the point was to not pick one image from cluster, but use whole cluster as valdiation:)</p>",
      "rawMarkdown": "You did reverse what we did :) the point was to not pick one image from cluster, but use whole cluster as valdiation:)",
      "votes": null
    },
    {
      "id": "194961",
      "postDate": "06/22/2017 11:40:45",
      "content": "<p>Yes, that is indeed what I have done - merged train+additional and ran your k-means. Thanks again!</p>",
      "rawMarkdown": "Yes, that is indeed what I have done - merged train+additional and ran your k-means. Thanks again!",
      "votes": null
    },
    {
      "id": "194964",
      "postDate": "06/22/2017 11:55:28",
      "content": "<p>Hmmm, not sure what you mean by reverse? I did the following:</p>\n\n<ol>\n<li>Cluster the train + additional data, and reduce to one image per cluster</li>\n<li>Split into training and validation sets (as there was only 1 image per cluster it meant that training and validation images couldn't be from the same patient... well the aim was to eliminate leakage, but I probably didn't do this well enough!)</li>\n</ol>\n\n<p>From my understanding you guys kept all of the images that were similar to eachother, but ensured that all were either in the training set, or in the validation set (no cross over between the two)? Makes sense! I was worried that if I did this I'd end up with some bias around the fact that some images were replicated many times and others were not... But that obviously didn't affect your performance.</p>\n\n<p>Thanks for the pointers.</p>",
      "rawMarkdown": "Hmmm, not sure what you mean by reverse? I did the following:\n\n1. Cluster the train + additional data, and reduce to one image per cluster\n2. Split into training and validation sets (as there was only 1 image per cluster it meant that training and validation images couldn't be from the same patient... well the aim was to eliminate leakage, but I probably didn't do this well enough!)\n\nFrom my understanding you guys kept all of the images that were similar to eachother, but ensured that all were either in the training set, or in the validation set (no cross over between the two)? Makes sense! I was worried that if I did this I'd end up with some bias around the fact that some images were replicated many times and others were not... But that obviously didn't affect your performance.\n\nThanks for the pointers.",
      "votes": null
    },
    {
      "id": "194966",
      "postDate": "06/22/2017 12:04:47",
      "content": "<p>The point is not to have same patient both in training and validation. From what I understood you did not achieve that.</p>\n\n<p>having simillar images in training set does not hurt, because you would usually do some image augmentations anyway, so similar images could be considered as some kind of augmented data as well.</p>",
      "rawMarkdown": "The point is not to have same patient both in training and validation. From what I understood you did not achieve that.\n\nhaving simillar images in training set does not hurt, because you would usually do some image augmentations anyway, so similar images could be considered as some kind of augmented data as well.",
      "votes": null
    },
    {
      "id": "194968",
      "postDate": "06/22/2017 12:30:53",
      "content": "<p>Out of interest, how many clusters(patients) did you have across train + additional? I had ~6500 clusters across the ~8700 images. I'm thinking that I was being too stringent in trying to pick out identical patients, with the result that a lot slipped through the net...</p>",
      "rawMarkdown": "Out of interest, how many clusters(patients) did you have across train + additional? I had ~6500 clusters across the ~8700 images. I'm thinking that I was being too stringent in trying to pick out identical patients, with the result that a lot slipped through the net...",
      "votes": null
    },
    {
      "id": "194974",
      "postDate": "06/22/2017 12:44:15",
      "content": "<p>Thanks raddar, and congrats on another win--keep pushing us!    </p>\n\n<p>Since we probed the lb we ended up using a naive two-fold split between public training and test with no additional data.   In hindsight this did not perform super well, in the sense that our local cv scores were below 0.7 and chosen ensemble weights do not align as expected with private test set scores on single models.   We set up 10 folds at one point and were planning to use them more extensively but ran out of time.   </p>\n\n<p>A cluster-based approach to the additional data like you have done seems very reasonable.   I guess this would induce some bias towards patients who got more pictures, which could be due to a variety of reasons, e.g. technicians who took more because of perceived poor quality of previous pics, ones who used a green filter in addition to none, ones who snapped before/after iodine or surgery, &amp;etc.   </p>\n\n<p>Could you please provide a little more detail on how you down-weighted the additional data in your final models? </p>",
      "rawMarkdown": "Thanks raddar, and congrats on another win--keep pushing us!    \n\nSince we probed the lb we ended up using a naive two-fold split between public training and test with no additional data.   In hindsight this did not perform super well, in the sense that our local cv scores were below 0.7 and chosen ensemble weights do not align as expected with private test set scores on single models.   We set up 10 folds at one point and were planning to use them more extensively but ran out of time.   \n\nA cluster-based approach to the additional data like you have done seems very reasonable.   I guess this would induce some bias towards patients who got more pictures, which could be due to a variety of reasons, e.g. technicians who took more because of perceived poor quality of previous pics, ones who used a green filter in addition to none, ones who snapped before/after iodine or surgery, &amp;etc.   \n\nCould you please provide a little more detail on how you down-weighted the additional data in your final models?",
      "votes": null
    },
    {
      "id": "194975",
      "postDate": "06/22/2017 12:48:22",
      "content": "<p>Unique patient detection was unnecessary and you overcomplicated yourself, I think.</p>",
      "rawMarkdown": "Unique patient detection was unnecessary and you overcomplicated yourself, I think.",
      "votes": null
    },
    {
      "id": "194977",
      "postDate": "06/22/2017 12:52:48",
      "content": "<p>Yes! This was exactly the problem I saw as well. However I saw this two way to late in the competition, when I was not able to make any re-splitting of the CV folds. Congratulations to your team.</p>",
      "rawMarkdown": "Yes! This was exactly the problem I saw as well. However I saw this two way to late in the competition, when I was not able to make any re-splitting of the CV folds. Congratulations to your team.",
      "votes": null
    },
    {
      "id": "194978",
      "postDate": "06/22/2017 12:52:57",
      "content": "<p>Yeah, I agree! How simple things seem in hindsight :)</p>",
      "rawMarkdown": "Yeah, I agree! How simple things seem in hindsight :)",
      "votes": null
    },
    {
      "id": "194979",
      "postDate": "06/22/2017 12:55:22",
      "content": "<p>Thanks for sharing raddar! And a well deserved victory! Congrats</p>",
      "rawMarkdown": "Thanks for sharing raddar! And a well deserved victory! Congrats",
      "votes": null
    },
    {
      "id": "194983",
      "postDate": "06/22/2017 13:15:00",
      "content": "<p>haha, it does seem though that having anonymized and thoroughly randomized patient ids would have helped the competition.  </p>",
      "rawMarkdown": "haha, it does seem though that having anonymized and thoroughly randomized patient ids would have helped the competition.",
      "votes": null
    },
    {
      "id": "195111",
      "postDate": "06/22/2017 19:21:48",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "195169",
      "postDate": "06/22/2017 21:54:50",
      "content": "<p>Using additional gave me a big boost, too.  As you guys, duplicate detection for train/valid splitting was necessary for the gain.  sklearn has a handy function, GroupKFold, after grouping was done.</p>",
      "rawMarkdown": "Using additional gave me a big boost, too.  As you guys, duplicate detection for train/valid splitting was necessary for the gain.  sklearn has a handy function, GroupKFold, after grouping was done.",
      "votes": null
    },
    {
      "id": "196984",
      "postDate": "06/28/2017 14:29:55",
      "content": "<p>Thanks, good to know about GroupKFold in sklearn</p>",
      "rawMarkdown": "Thanks, good to know about GroupKFold in sklearn",
      "votes": null
    },
    {
      "id": "196986",
      "postDate": "06/28/2017 14:32:18",
      "content": "<p>Congratulations and thanks for sharing! This explains a lot of overfitting I saw in my models.</p>",
      "rawMarkdown": "Congratulations and thanks for sharing! This explains a lot of overfitting I saw in my models.",
      "votes": null
    },
    {
      "id": "199056",
      "postDate": "07/04/2017 12:39:45",
      "content": "<p>Congratulations! Really cool way to split the validation set !</p>",
      "rawMarkdown": "Congratulations! Really cool way to split the validation set !",
      "votes": null
    },
    {
      "id": "530564",
      "postDate": "05/13/2019 07:22:37",
      "content": "<p>Hey raddar, first of all congrats!\nCould you maybe comment on the method you used 2 years ago?</p>\n\n<p>We are a group of students from Heidelberg University and we're currently working on this challenge in a bioinformatics seminar. Would be great if you could give us some advice.</p>",
      "rawMarkdown": "Hey raddar, first of all congrats!\nCould you maybe comment on the method you used 2 years ago?\n\nWe are a group of students from Heidelberg University and we're currently working on this challenge in a bioinformatics seminar. Would be great if you could give us some advice.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 194958,
      "author_name": "fergusoci",
      "author_url": "",
      "post_date": "06/22/2017 11:33:11",
      "content": "<p>I was just about to ask on the forum whether any of the top teams used the additional data as it seemed so far that many who posted solutions did not. </p>\n\n<p>Good to see that it worked for you! I thought I was doing something similar by clustering the train/additional images and only using one image from each cluster... I thought I was being pretty pedantic about it but I must not have done this well enough as my private LB score was way higher than what I was getting in validation. I expect that if I had ignored the additional set completely I may have ended up with something more reliable...</p>\n\n<p>Congratulations, by the way. Good job!</p>",
      "votes": null,
      "replies": [
        {
          "id": 194960,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/22/2017 11:39:54",
          "content": "<p>You did reverse what we did :) the point was to not pick one image from cluster, but use whole cluster as valdiation:)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194964,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "06/22/2017 11:55:28",
          "content": "<p>Hmmm, not sure what you mean by reverse? I did the following:</p>\n\n<ol>\n<li>Cluster the train + additional data, and reduce to one image per cluster</li>\n<li>Split into training and validation sets (as there was only 1 image per cluster it meant that training and validation images couldn't be from the same patient... well the aim was to eliminate leakage, but I probably didn't do this well enough!)</li>\n</ol>\n\n<p>From my understanding you guys kept all of the images that were similar to eachother, but ensured that all were either in the training set, or in the validation set (no cross over between the two)? Makes sense! I was worried that if I did this I'd end up with some bias around the fact that some images were replicated many times and others were not... But that obviously didn't affect your performance.</p>\n\n<p>Thanks for the pointers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194966,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/22/2017 12:04:47",
          "content": "<p>The point is not to have same patient both in training and validation. From what I understood you did not achieve that.</p>\n\n<p>having simillar images in training set does not hurt, because you would usually do some image augmentations anyway, so similar images could be considered as some kind of augmented data as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194968,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "06/22/2017 12:30:53",
          "content": "<p>Out of interest, how many clusters(patients) did you have across train + additional? I had ~6500 clusters across the ~8700 images. I'm thinking that I was being too stringent in trying to pick out identical patients, with the result that a lot slipped through the net...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194975,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/22/2017 12:48:22",
          "content": "<p>Unique patient detection was unnecessary and you overcomplicated yourself, I think.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194978,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "06/22/2017 12:52:57",
          "content": "<p>Yeah, I agree! How simple things seem in hindsight :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194983,
          "author_name": "sasrdw",
          "author_url": "",
          "post_date": "06/22/2017 13:15:00",
          "content": "<p>haha, it does seem though that having anonymized and thoroughly randomized patient ids would have helped the competition.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 194959,
      "author_name": "vfdev5",
      "author_url": "",
      "post_date": "06/22/2017 11:37:42",
      "content": "<p>@raddar, thanks for sharing your strategy and congratulations ! \nHowever I did not get how did you composed training and validations sets. You clustered all train and additional in 100 clusters and picked a training dataset as 80 random clusters and validation dataset 20 others ? </p>\n\n<p>ps. glad that the kernel served</p>",
      "votes": null,
      "replies": [
        {
          "id": 194961,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "06/22/2017 11:40:45",
          "content": "<p>Yes, that is indeed what I have done - merged train+additional and ran your k-means. Thanks again!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 194974,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "06/22/2017 12:44:15",
      "content": "<p>Thanks raddar, and congrats on another win--keep pushing us!    </p>\n\n<p>Since we probed the lb we ended up using a naive two-fold split between public training and test with no additional data.   In hindsight this did not perform super well, in the sense that our local cv scores were below 0.7 and chosen ensemble weights do not align as expected with private test set scores on single models.   We set up 10 folds at one point and were planning to use them more extensively but ran out of time.   </p>\n\n<p>A cluster-based approach to the additional data like you have done seems very reasonable.   I guess this would induce some bias towards patients who got more pictures, which could be due to a variety of reasons, e.g. technicians who took more because of perceived poor quality of previous pics, ones who used a green filter in addition to none, ones who snapped before/after iodine or surgery, &amp;etc.   </p>\n\n<p>Could you please provide a little more detail on how you down-weighted the additional data in your final models? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 194977,
      "author_name": "oysteijo",
      "author_url": "",
      "post_date": "06/22/2017 12:52:48",
      "content": "<p>Yes! This was exactly the problem I saw as well. However I saw this two way to late in the competition, when I was not able to make any re-splitting of the CV folds. Congratulations to your team.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 194979,
      "author_name": "yadavsarthak",
      "author_url": "",
      "post_date": "06/22/2017 12:55:22",
      "content": "<p>Thanks for sharing raddar! And a well deserved victory! Congrats</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 195111,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/22/2017 19:21:48",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 195169,
      "author_name": "kubilai",
      "author_url": "",
      "post_date": "06/22/2017 21:54:50",
      "content": "<p>Using additional gave me a big boost, too.  As you guys, duplicate detection for train/valid splitting was necessary for the gain.  sklearn has a handy function, GroupKFold, after grouping was done.</p>",
      "votes": null,
      "replies": [
        {
          "id": 196984,
          "author_name": "lakshaykc",
          "author_url": "",
          "post_date": "06/28/2017 14:29:55",
          "content": "<p>Thanks, good to know about GroupKFold in sklearn</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 196986,
      "author_name": "lakshaykc",
      "author_url": "",
      "post_date": "06/28/2017 14:32:18",
      "content": "<p>Congratulations and thanks for sharing! This explains a lot of overfitting I saw in my models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 199056,
      "author_name": "scottykwok",
      "author_url": "",
      "post_date": "07/04/2017 12:39:45",
      "content": "<p>Congratulations! Really cool way to split the validation set !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 530564,
      "author_name": "aliceschmid",
      "author_url": "",
      "post_date": "05/13/2019 07:22:37",
      "content": "<p>Hey raddar, first of all congrats!\nCould you maybe comment on the method you used 2 years ago?</p>\n\n<p>We are a group of students from Heidelberg University and we're currently working on this challenge in a bioinformatics seminar. Would be great if you could give us some advice.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "194957": "The competition was very challanging because of the quality and the origin of the data. The main challange for us was how to make use of additional data and have a solid validation framework. \n\nFirst thing we noticed, that additional dataset had strong correlation with train images. Looking thorougly at each image we found, that training set included the best quality image of a single patient, and all other frames were put into additional data file.\n\nSo to avoid simillar image leakage to our validation set we did the following:\n\n1) calculate image colour histograms.\n2) run k-means clustering with k = 100\n3) take random 20 clusters and use that for validation\n\nIts quite easy solution, but it really helped to merge simillar images into a cluster and we avoided validation set contamination with simillar images in training set. As a result we could use additional dataset without too much worries (we did put less weight on additional data in our final models).\n\nOn top of that, we removed (by hand) all the blurry or not clear images in additional dataset (roughly 15% of additional images).\n\nIn the end we had decent validation set, with a mix of original and additional dataset images. Our models on validation set scored 0.75-0.8 - pretty close to our final result, which I am very happy with!\n\nSpecial thanks to this kernel which I used to calculate histograms and k-means for our validation:\nhttps://www.kaggle.com/vfdev5/type-1-clustering\n\nIt really deserves more upvotes:)\n\nVery interesting to hear, how other teams did validaiton. Please share:)",
    "194958": "I was just about to ask on the forum whether any of the top teams used the additional data as it seemed so far that many who posted solutions did not. \n\nGood to see that it worked for you! I thought I was doing something similar by clustering the train/additional images and only using one image from each cluster... I thought I was being pretty pedantic about it but I must not have done this well enough as my private LB score was way higher than what I was getting in validation. I expect that if I had ignored the additional set completely I may have ended up with something more reliable...\n\nCongratulations, by the way. Good job!",
    "194959": "raddar, thanks for sharing your strategy and congratulations ! \nHowever I did not get how did you composed training and validations sets. You clustered all train and additional in 100 clusters and picked a training dataset as 80 random clusters and validation dataset 20 others ? \n\nps. glad that the kernel served",
    "194960": "You did reverse what we did :) the point was to not pick one image from cluster, but use whole cluster as valdiation:)",
    "194961": "Yes, that is indeed what I have done - merged train+additional and ran your k-means. Thanks again!",
    "194964": "Hmmm, not sure what you mean by reverse? I did the following:\n\n1. Cluster the train + additional data, and reduce to one image per cluster\n2. Split into training and validation sets (as there was only 1 image per cluster it meant that training and validation images couldn't be from the same patient... well the aim was to eliminate leakage, but I probably didn't do this well enough!)\n\nFrom my understanding you guys kept all of the images that were similar to eachother, but ensured that all were either in the training set, or in the validation set (no cross over between the two)? Makes sense! I was worried that if I did this I'd end up with some bias around the fact that some images were replicated many times and others were not... But that obviously didn't affect your performance.\n\nThanks for the pointers.",
    "194966": "The point is not to have same patient both in training and validation. From what I understood you did not achieve that.\n\nhaving simillar images in training set does not hurt, because you would usually do some image augmentations anyway, so similar images could be considered as some kind of augmented data as well.",
    "194968": "Out of interest, how many clusters(patients) did you have across train + additional? I had ~6500 clusters across the ~8700 images. I'm thinking that I was being too stringent in trying to pick out identical patients, with the result that a lot slipped through the net...",
    "194974": "Thanks raddar, and congrats on another win--keep pushing us!    \n\nSince we probed the lb we ended up using a naive two-fold split between public training and test with no additional data.   In hindsight this did not perform super well, in the sense that our local cv scores were below 0.7 and chosen ensemble weights do not align as expected with private test set scores on single models.   We set up 10 folds at one point and were planning to use them more extensively but ran out of time.   \n\nA cluster-based approach to the additional data like you have done seems very reasonable.   I guess this would induce some bias towards patients who got more pictures, which could be due to a variety of reasons, e.g. technicians who took more because of perceived poor quality of previous pics, ones who used a green filter in addition to none, ones who snapped before/after iodine or surgery, &amp;etc.   \n\nCould you please provide a little more detail on how you down-weighted the additional data in your final models?",
    "194975": "Unique patient detection was unnecessary and you overcomplicated yourself, I think.",
    "194977": "Yes! This was exactly the problem I saw as well. However I saw this two way to late in the competition, when I was not able to make any re-splitting of the CV folds. Congratulations to your team.",
    "194978": "Yeah, I agree! How simple things seem in hindsight :)",
    "194979": "Thanks for sharing raddar! And a well deserved victory! Congrats",
    "194983": "haha, it does seem though that having anonymized and thoroughly randomized patient ids would have helped the competition.",
    "195111": "Congrats!",
    "195169": "Using additional gave me a big boost, too.  As you guys, duplicate detection for train/valid splitting was necessary for the gain.  sklearn has a handy function, GroupKFold, after grouping was done.",
    "196984": "Thanks, good to know about GroupKFold in sklearn",
    "196986": "Congratulations and thanks for sharing! This explains a lot of overfitting I saw in my models.",
    "199056": "Congratulations! Really cool way to split the validation set !",
    "530564": "Hey raddar, first of all congrats!\nCould you maybe comment on the method you used 2 years ago?\n\nWe are a group of students from Heidelberg University and we're currently working on this challenge in a bioinformatics seminar. Would be great if you could give us some advice."
  },
  "source": "meta"
}