{
  "id": 35111,
  "title": "26th place solution",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/writeups/odinn-26th-place-solution",
  "author_name": "",
  "post_date": "2017-06-22T22:05:05.740Z",
  "votes": 8,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First thanks to the organizers and the congratulations to winners, and those who are satisfied with their leaderboard placement and/or effort! Thanks to my team mate! As usual I learned a lot.</p>\n\n<p>Team OdiNN formed just hours before the team merger deadline. We (Graham and me) had a short conversation before the deadline and agreed to team up.</p>\n\n<p>We then started to discuss our individual models, and we realized that we had two totally different approaches to a solutions. Graham had used ROI on the train images and fully trained neural networks with VGG16 and VGG19 structures and then averaging over different seeds.However I had used train and additional dataset and done transfer learning from Keras imagenet pretrained neural networks. (vgg16, vgg19, resnet50, xception, inception_v3). The I trained many models and stacked them.</p>\n\n<p>Since we had this totally different approaches and had not agreed on any common train/test split or sharing a common seed, we basically made the following masterplan: We take the each of our best models (Grahams best and my best) and then we simply average the two based on Stage1 leaderboard feedback.</p>\n\n<p>My stacking model.\nI extracted the features for all images (train and additional) through pretrained vgg16, vgg19, resnet50, xception and inception_v3. Each image was flipped and rotate to give me 8 different flips and rotations. Each of these feature set was then trained in 5 fold cross validation in six different learning algorithms. (I safely kept flips and rotations of one image within the same fold to avoid leakage over the folds) The six learning algorithms where: KNN, XGBoost, Random Forest, ExtraTrees, Logistic Regression and fully connected neural network. I kept the same hyperparameters for these learners for all feature sets. This then became 30 models (5x6) where each was trained in 5 fold cv. Out-of-fold predictions where taken and saved of course.</p>\n\n<p>These 30 models where then stacked to a three second level models. An XGBoost model, Logistic Regression and another fully connected neural network. </p>\n\n<p>At the third level I tried to do weighted geometric averaging, however it then appeared to me that the there must have been a leak over my CV folds. I guess some of the images are actually from the same patient. :-(  So, the weighted geometric (nor arithmetic) averaging did not work and I ended up with the plain unweighted arithmetic mean. I didn't notice CV-fold leak until the second day of stage two, and training the first stage took about 50 hours, so I had no chance of solving this problem.</p>\n\n<p>The other model that was blended with my stacked model was based on ROI extraction of the images and then fully retrain neural networks of VGG16 and VGG19 neural network. He only used the original train set (not the additional data). Graham may elaborate on this. If I understand him correctly he is averaging each over five training sessions with different seeds. The problem he got was that the vgg19 neural net did not manage to make it to the deadline so one vgg19 net was used in the final mix.</p>\n\n<p>Thanks all!</p>",
  "messages": [
    {
      "id": "194976",
      "postDate": "06/22/2017 12:49:12",
      "content": "<p>First thanks to the organizers and the congratulations to winners, and those who are satisfied with their leaderboard placement and/or effort! Thanks to my team mate! As usual I learned a lot.</p>\n\n<p>Team OdiNN formed just hours before the team merger deadline. We (Graham and me) had a short conversation before the deadline and agreed to team up.</p>\n\n<p>We then started to discuss our individual models, and we realized that we had two totally different approaches to a solutions. Graham had used ROI on the train images and fully trained neural networks with VGG16 and VGG19 structures and then averaging over different seeds.However I had used train and additional dataset and done transfer learning from Keras imagenet pretrained neural networks. (vgg16, vgg19, resnet50, xception, inception_v3). The I trained many models and stacked them.</p>\n\n<p>Since we had this totally different approaches and had not agreed on any common train/test split or sharing a common seed, we basically made the following masterplan: We take the each of our best models (Grahams best and my best) and then we simply average the two based on Stage1 leaderboard feedback.</p>\n\n<p>My stacking model.\nI extracted the features for all images (train and additional) through pretrained vgg16, vgg19, resnet50, xception and inception_v3. Each image was flipped and rotate to give me 8 different flips and rotations. Each of these feature set was then trained in 5 fold cross validation in six different learning algorithms. (I safely kept flips and rotations of one image within the same fold to avoid leakage over the folds) The six learning algorithms where: KNN, XGBoost, Random Forest, ExtraTrees, Logistic Regression and fully connected neural network. I kept the same hyperparameters for these learners for all feature sets. This then became 30 models (5x6) where each was trained in 5 fold cv. Out-of-fold predictions where taken and saved of course.</p>\n\n<p>These 30 models where then stacked to a three second level models. An XGBoost model, Logistic Regression and another fully connected neural network. </p>\n\n<p>At the third level I tried to do weighted geometric averaging, however it then appeared to me that the there must have been a leak over my CV folds. I guess some of the images are actually from the same patient. :-(  So, the weighted geometric (nor arithmetic) averaging did not work and I ended up with the plain unweighted arithmetic mean. I didn't notice CV-fold leak until the second day of stage two, and training the first stage took about 50 hours, so I had no chance of solving this problem.</p>\n\n<p>The other model that was blended with my stacked model was based on ROI extraction of the images and then fully retrain neural networks of VGG16 and VGG19 neural network. He only used the original train set (not the additional data). Graham may elaborate on this. If I understand him correctly he is averaging each over five training sessions with different seeds. The problem he got was that the vgg19 neural net did not manage to make it to the deadline so one vgg19 net was used in the final mix.</p>\n\n<p>Thanks all!</p>",
      "rawMarkdown": "First thanks to the organizers and the congratulations to winners, and those who are satisfied with their leaderboard placement and/or effort! Thanks to my team mate! As usual I learned a lot.\n\nTeam OdiNN formed just hours before the team merger deadline. We (Graham and me) had a short conversation before the deadline and agreed to team up.\n\nWe then started to discuss our individual models, and we realized that we had two totally different approaches to a solutions. Graham had used ROI on the train images and fully trained neural networks with VGG16 and VGG19 structures and then averaging over different seeds.However I had used train and additional dataset and done transfer learning from Keras imagenet pretrained neural networks. (vgg16, vgg19, resnet50, xception, inception_v3). The I trained many models and stacked them.\n\nSince we had this totally different approaches and had not agreed on any common train/test split or sharing a common seed, we basically made the following masterplan: We take the each of our best models (Grahams best and my best) and then we simply average the two based on Stage1 leaderboard feedback.\n\nMy stacking model.\nI extracted the features for all images (train and additional) through pretrained vgg16, vgg19, resnet50, xception and inception_v3. Each image was flipped and rotate to give me 8 different flips and rotations. Each of these feature set was then trained in 5 fold cross validation in six different learning algorithms. (I safely kept flips and rotations of one image within the same fold to avoid leakage over the folds) The six learning algorithms where: KNN, XGBoost, Random Forest, ExtraTrees, Logistic Regression and fully connected neural network. I kept the same hyperparameters for these learners for all feature sets. This then became 30 models (5x6) where each was trained in 5 fold cv. Out-of-fold predictions where taken and saved of course.\n\nThese 30 models where then stacked to a three second level models. An XGBoost model, Logistic Regression and another fully connected neural network. \n\nAt the third level I tried to do weighted geometric averaging, however it then appeared to me that the there must have been a leak over my CV folds. I guess some of the images are actually from the same patient. :-(  So, the weighted geometric (nor arithmetic) averaging did not work and I ended up with the plain unweighted arithmetic mean. I didn't notice CV-fold leak until the second day of stage two, and training the first stage took about 50 hours, so I had no chance of solving this problem.\n\nThe other model that was blended with my stacked model was based on ROI extraction of the images and then fully retrain neural networks of VGG16 and VGG19 neural network. He only used the original train set (not the additional data). Graham may elaborate on this. If I understand him correctly he is averaging each over five training sessions with different seeds. The problem he got was that the vgg19 neural net did not manage to make it to the deadline so one vgg19 net was used in the final mix.\n\nThanks all!",
      "votes": null
    },
    {
      "id": "195025",
      "postDate": "06/22/2017 16:28:17",
      "content": "<p>Thanks for writing this up Øystein!</p>\n\n<p>Hi everyone. I am Øystein's teammate in this competition.</p>\n\n<p>As  Øystein mentioned we did our modelling completely separate and then took a weighted average of our predicted probability distributions at the end. As we perceived it, this came with pros and cons. On one hand, not coordinating our earlier efforts made is difficult/impossible to stack our models together. We probably could have increased our accuracy by ensembling in a more sophisticated way. The positive to our approach though is that we could mitigate the risk of going all in on a single strategy. Especially with all the confusion surrounding the data leak, it was unclear to us whether it was better or worse to use the additional data. That is why we stuck to having one model that used the additional data and one model that did not.</p>\n\n<p>My side of the submission used a two stage model - one model to crop out the ROI (region of interest) and a set of networks to do the actual predictions. The cropper was based on the VGG-16 architecture. I drew bounding boxes on the training data (and later on the stage 1 test data as well). My cropper was largely inspired by Felix Lau's submission in the right whale competition (<a href=\"http://felixlaumon.github.io/2015/01/08/kaggle-right-whale.html\">http://felixlaumon.github.io/2015/01/08/kaggle-right-whale.html</a>). The idea is to view the problem as a multi-label regression and used a CNN to estimate 4 values (the proportional height/width of the extreme corners of the bounding boxes). The network can then predict ROIs on new images. For the cropper I felt that something that worked intuitively (i.e. the cropped images looked visually informative) would be good enough, so I didn't spend too much time trying to optimize it. To evaluate changes I would manually go through the crops it produced and see if they looked informative to me.</p>\n\n<p>The classifier layer was a collection of four network structures. I used VGG-16 and VGG-19 with imagenet weights frozen up to layers 15 and 17 respectively, and VGG-16 and VGG-19 from scratch. Each model was averaged over 5 CV folds and a number of random seeds (24 for the fine-tuned models, 16 for VGG-16 from scratch and 1 for VGG-19 because I ran out of time). I noticed on the first stage public leaderboard that I had a huge jump when I moved from training a single model for each network to training cv_fold_num*random_seed_num models. Of course, using this sort of technique made my training time sky-rocket.</p>\n\n<p>My ensemble was just a flat average. I experimented with some simple stacking techniques but was finding that  the stacking was adding no additional value (my guess is my models were too homogenous). Training the whole pipeline took 7 days (I started it right when stage 1 ended and submitted right before the stage 2 deadline) and I have a 1080ti and a titan X pascal edition.</p>\n\n<p>I want to thank Intel and MobileODT for sponsoring this competition and Kaggle for hosting it. I work at the Women and Children's Health Research Institute in Canada and we were thrilled to learn about this competition. Our mandate is to promote health research in women and children's health and this competition was a wonderful and innovative way to advance technology for women's health.</p>",
      "rawMarkdown": "Thanks for writing this up Øystein!\n\nHi everyone. I am Øystein's teammate in this competition.\n\nAs  Øystein mentioned we did our modelling completely separate and then took a weighted average of our predicted probability distributions at the end. As we perceived it, this came with pros and cons. On one hand, not coordinating our earlier efforts made is difficult/impossible to stack our models together. We probably could have increased our accuracy by ensembling in a more sophisticated way. The positive to our approach though is that we could mitigate the risk of going all in on a single strategy. Especially with all the confusion surrounding the data leak, it was unclear to us whether it was better or worse to use the additional data. That is why we stuck to having one model that used the additional data and one model that did not.\n\nMy side of the submission used a two stage model - one model to crop out the ROI (region of interest) and a set of networks to do the actual predictions. The cropper was based on the VGG-16 architecture. I drew bounding boxes on the training data (and later on the stage 1 test data as well). My cropper was largely inspired by Felix Lau's submission in the right whale competition (http://felixlaumon.github.io/2015/01/08/kaggle-right-whale.html). The idea is to view the problem as a multi-label regression and used a CNN to estimate 4 values (the proportional height/width of the extreme corners of the bounding boxes). The network can then predict ROIs on new images. For the cropper I felt that something that worked intuitively (i.e. the cropped images looked visually informative) would be good enough, so I didn't spend too much time trying to optimize it. To evaluate changes I would manually go through the crops it produced and see if they looked informative to me.\n\nThe classifier layer was a collection of four network structures. I used VGG-16 and VGG-19 with imagenet weights frozen up to layers 15 and 17 respectively, and VGG-16 and VGG-19 from scratch. Each model was averaged over 5 CV folds and a number of random seeds (24 for the fine-tuned models, 16 for VGG-16 from scratch and 1 for VGG-19 because I ran out of time). I noticed on the first stage public leaderboard that I had a huge jump when I moved from training a single model for each network to training cv_fold_num*random_seed_num models. Of course, using this sort of technique made my training time sky-rocket.\n\nMy ensemble was just a flat average. I experimented with some simple stacking techniques but was finding that  the stacking was adding no additional value (my guess is my models were too homogenous). Training the whole pipeline took 7 days (I started it right when stage 1 ended and submitted right before the stage 2 deadline) and I have a 1080ti and a titan X pascal edition.\n\nI want to thank Intel and MobileODT for sponsoring this competition and Kaggle for hosting it. I work at the Women and Children's Health Research Institute in Canada and we were thrilled to learn about this competition. Our mandate is to promote health research in women and children's health and this competition was a wonderful and innovative way to advance technology for women's health.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 195025,
      "author_name": "gkericks",
      "author_url": "",
      "post_date": "06/22/2017 16:28:17",
      "content": "<p>Thanks for writing this up Øystein!</p>\n\n<p>Hi everyone. I am Øystein's teammate in this competition.</p>\n\n<p>As  Øystein mentioned we did our modelling completely separate and then took a weighted average of our predicted probability distributions at the end. As we perceived it, this came with pros and cons. On one hand, not coordinating our earlier efforts made is difficult/impossible to stack our models together. We probably could have increased our accuracy by ensembling in a more sophisticated way. The positive to our approach though is that we could mitigate the risk of going all in on a single strategy. Especially with all the confusion surrounding the data leak, it was unclear to us whether it was better or worse to use the additional data. That is why we stuck to having one model that used the additional data and one model that did not.</p>\n\n<p>My side of the submission used a two stage model - one model to crop out the ROI (region of interest) and a set of networks to do the actual predictions. The cropper was based on the VGG-16 architecture. I drew bounding boxes on the training data (and later on the stage 1 test data as well). My cropper was largely inspired by Felix Lau's submission in the right whale competition (<a href=\"http://felixlaumon.github.io/2015/01/08/kaggle-right-whale.html\">http://felixlaumon.github.io/2015/01/08/kaggle-right-whale.html</a>). The idea is to view the problem as a multi-label regression and used a CNN to estimate 4 values (the proportional height/width of the extreme corners of the bounding boxes). The network can then predict ROIs on new images. For the cropper I felt that something that worked intuitively (i.e. the cropped images looked visually informative) would be good enough, so I didn't spend too much time trying to optimize it. To evaluate changes I would manually go through the crops it produced and see if they looked informative to me.</p>\n\n<p>The classifier layer was a collection of four network structures. I used VGG-16 and VGG-19 with imagenet weights frozen up to layers 15 and 17 respectively, and VGG-16 and VGG-19 from scratch. Each model was averaged over 5 CV folds and a number of random seeds (24 for the fine-tuned models, 16 for VGG-16 from scratch and 1 for VGG-19 because I ran out of time). I noticed on the first stage public leaderboard that I had a huge jump when I moved from training a single model for each network to training cv_fold_num*random_seed_num models. Of course, using this sort of technique made my training time sky-rocket.</p>\n\n<p>My ensemble was just a flat average. I experimented with some simple stacking techniques but was finding that  the stacking was adding no additional value (my guess is my models were too homogenous). Training the whole pipeline took 7 days (I started it right when stage 1 ended and submitted right before the stage 2 deadline) and I have a 1080ti and a titan X pascal edition.</p>\n\n<p>I want to thank Intel and MobileODT for sponsoring this competition and Kaggle for hosting it. I work at the Women and Children's Health Research Institute in Canada and we were thrilled to learn about this competition. Our mandate is to promote health research in women and children's health and this competition was a wonderful and innovative way to advance technology for women's health.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "194976": "First thanks to the organizers and the congratulations to winners, and those who are satisfied with their leaderboard placement and/or effort! Thanks to my team mate! As usual I learned a lot.\n\nTeam OdiNN formed just hours before the team merger deadline. We (Graham and me) had a short conversation before the deadline and agreed to team up.\n\nWe then started to discuss our individual models, and we realized that we had two totally different approaches to a solutions. Graham had used ROI on the train images and fully trained neural networks with VGG16 and VGG19 structures and then averaging over different seeds.However I had used train and additional dataset and done transfer learning from Keras imagenet pretrained neural networks. (vgg16, vgg19, resnet50, xception, inception_v3). The I trained many models and stacked them.\n\nSince we had this totally different approaches and had not agreed on any common train/test split or sharing a common seed, we basically made the following masterplan: We take the each of our best models (Grahams best and my best) and then we simply average the two based on Stage1 leaderboard feedback.\n\nMy stacking model.\nI extracted the features for all images (train and additional) through pretrained vgg16, vgg19, resnet50, xception and inception_v3. Each image was flipped and rotate to give me 8 different flips and rotations. Each of these feature set was then trained in 5 fold cross validation in six different learning algorithms. (I safely kept flips and rotations of one image within the same fold to avoid leakage over the folds) The six learning algorithms where: KNN, XGBoost, Random Forest, ExtraTrees, Logistic Regression and fully connected neural network. I kept the same hyperparameters for these learners for all feature sets. This then became 30 models (5x6) where each was trained in 5 fold cv. Out-of-fold predictions where taken and saved of course.\n\nThese 30 models where then stacked to a three second level models. An XGBoost model, Logistic Regression and another fully connected neural network. \n\nAt the third level I tried to do weighted geometric averaging, however it then appeared to me that the there must have been a leak over my CV folds. I guess some of the images are actually from the same patient. :-(  So, the weighted geometric (nor arithmetic) averaging did not work and I ended up with the plain unweighted arithmetic mean. I didn't notice CV-fold leak until the second day of stage two, and training the first stage took about 50 hours, so I had no chance of solving this problem.\n\nThe other model that was blended with my stacked model was based on ROI extraction of the images and then fully retrain neural networks of VGG16 and VGG19 neural network. He only used the original train set (not the additional data). Graham may elaborate on this. If I understand him correctly he is averaging each over five training sessions with different seeds. The problem he got was that the vgg19 neural net did not manage to make it to the deadline so one vgg19 net was used in the final mix.\n\nThanks all!",
    "195025": "Thanks for writing this up Øystein!\n\nHi everyone. I am Øystein's teammate in this competition.\n\nAs  Øystein mentioned we did our modelling completely separate and then took a weighted average of our predicted probability distributions at the end. As we perceived it, this came with pros and cons. On one hand, not coordinating our earlier efforts made is difficult/impossible to stack our models together. We probably could have increased our accuracy by ensembling in a more sophisticated way. The positive to our approach though is that we could mitigate the risk of going all in on a single strategy. Especially with all the confusion surrounding the data leak, it was unclear to us whether it was better or worse to use the additional data. That is why we stuck to having one model that used the additional data and one model that did not.\n\nMy side of the submission used a two stage model - one model to crop out the ROI (region of interest) and a set of networks to do the actual predictions. The cropper was based on the VGG-16 architecture. I drew bounding boxes on the training data (and later on the stage 1 test data as well). My cropper was largely inspired by Felix Lau's submission in the right whale competition (http://felixlaumon.github.io/2015/01/08/kaggle-right-whale.html). The idea is to view the problem as a multi-label regression and used a CNN to estimate 4 values (the proportional height/width of the extreme corners of the bounding boxes). The network can then predict ROIs on new images. For the cropper I felt that something that worked intuitively (i.e. the cropped images looked visually informative) would be good enough, so I didn't spend too much time trying to optimize it. To evaluate changes I would manually go through the crops it produced and see if they looked informative to me.\n\nThe classifier layer was a collection of four network structures. I used VGG-16 and VGG-19 with imagenet weights frozen up to layers 15 and 17 respectively, and VGG-16 and VGG-19 from scratch. Each model was averaged over 5 CV folds and a number of random seeds (24 for the fine-tuned models, 16 for VGG-16 from scratch and 1 for VGG-19 because I ran out of time). I noticed on the first stage public leaderboard that I had a huge jump when I moved from training a single model for each network to training cv_fold_num*random_seed_num models. Of course, using this sort of technique made my training time sky-rocket.\n\nMy ensemble was just a flat average. I experimented with some simple stacking techniques but was finding that  the stacking was adding no additional value (my guess is my models were too homogenous). Training the whole pipeline took 7 days (I started it right when stage 1 ended and submitted right before the stage 2 deadline) and I have a 1080ti and a titan X pascal edition.\n\nI want to thank Intel and MobileODT for sponsoring this competition and Kaggle for hosting it. I work at the Women and Children's Health Research Institute in Canada and we were thrilled to learn about this competition. Our mandate is to promote health research in women and children's health and this competition was a wonderful and innovative way to advance technology for women's health."
  },
  "source": "meta"
}