{
  "id": 35104,
  "title": "9th place solution",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/writeups/zfturbo-9th-place-solution",
  "author_name": "",
  "post_date": "2017-06-22T09:13:59.675325800Z",
  "votes": 33,
  "comment_count": 16,
  "views": 0,
  "content": "<h2>Used tools</h2>\n\n<p>Windows 10 + Python 3.4 + Keras 1.2 + Theano 0.9</p>\n\n<h2>ZF_UNET_224 (version 1)</h2>\n\n<p>To find important parts of images I used <a href=\"https://github.com/ZFTurbo/ZF_UNET_224_Pretrained_Model\">my own version of UNET</a>. First UNET was trained on segmentation of train images, <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/194936/6708/polygons.json\">which I made by hands</a> with Sloth. In this segmentation I tried to cut all elements which are totally useless like speculum and everything outside it. The predictions look nice:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/194936/6709/ZF_UNET_v1_predictions.jpg\" alt=\"enter image description here\" title=\"\"></p>\n\n<h2>ZF_UNET_224 (version 2)</h2>\n\n<p>Second UNET was trained on rectangles <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565\">provided by Paul</a>. This net finds the main region of interest. The predictions look even better.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/194936/6710/ZF_UNET_v2_predictions.jpg\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>I didn’t use these predictions directly. I only extract bounding box rectangles for each image to form rectangles.csv file which I used in later training process.</p>\n\n<h2>Zoo</h2>\n\n<p>I used the following set of CNNs: VGG16, VGG19, RESNET50, INCEPTION_V3, SQUEEZE_NET, DENSENET_161, DENSENET_121. The training process for them was mostly similar. 5 KFold validation, vary only learning rate parameters and batch size. For each model I obtain: train OOF predictions with same length as number of training images and test predictions.</p>\n\n<h2>Augmentations</h2>\n\n<p>I think augmentation was important key for this problem, since we had very small amount of data. I used: </p>\n\n<ol>\n<li>Random crops based on rectangle.csv generated from UNETs. These crops were in very big range from UNET_v1 prediction to UNET_v2 predictions. </li>\n<li>Random perspective transformation </li>\n<li>Random rotations (mean color or border reflect at random) and mirroring </li>\n<li>Lightning change </li>\n<li>Rare random blur</li>\n</ol>\n\n<p>Basically neural nets never see the same pictures during training.</p>\n\n<h2>Batch generator</h2>\n\n<p>For learning process I used fit_generator function from Keras, which I recommend to use for everyone. You only need to create your own batch_generator function. No need to store many different images in memory or on HDD.\nIn my batch generator I add some fraction of images from “additional” folder ~25%. It greatly improves validation and leaderboard score. And we all know now it’s because of leakage of test images in additional folder.</p>\n\n<h2>Ensemble</h2>\n\n<p>For ensemble I used XGBoost blender. Final solution was average on 500 XGBoost iterations with different seed and random run parameters. </p>\n\n<h2>Submission</h2>\n\n<p>The only difference for my 2 final submissions was to train with usage of “additional” images or without them. My final models used “train” and “test_stg1” images for training.</p>\n\n<p>Submission v1 (with additional): Validation score: 0.57457 Private LB: 0.88856</p>\n\n<p>Submission v2 (without additional): Validation score: 0.64254 Private LB: 0.83209</p>\n\n<p>I believe low scores on private LB depends on bad data preparation, mostly because of mislabeling. We could clean it by hands, but in case private test set had the same labeling quality it would make everything worse.</p>\n\n<h2>Code</h2>\n\n<p>You can find my code as it was prepared for Kaggle <a href=\"https://github.com/ZFTurbo/KAGGLE_CERVICAL_CANCER_2017\">on GitHub</a></p>",
  "messages": [
    {
      "id": "194936",
      "postDate": "06/22/2017 09:13:59",
      "content": "<h2>Used tools</h2>\n\n<p>Windows 10 + Python 3.4 + Keras 1.2 + Theano 0.9</p>\n\n<h2>ZF_UNET_224 (version 1)</h2>\n\n<p>To find important parts of images I used <a href=\"https://github.com/ZFTurbo/ZF_UNET_224_Pretrained_Model\">my own version of UNET</a>. First UNET was trained on segmentation of train images, <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/194936/6708/polygons.json\">which I made by hands</a> with Sloth. In this segmentation I tried to cut all elements which are totally useless like speculum and everything outside it. The predictions look nice:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/194936/6709/ZF_UNET_v1_predictions.jpg\" alt=\"enter image description here\" title=\"\"></p>\n\n<h2>ZF_UNET_224 (version 2)</h2>\n\n<p>Second UNET was trained on rectangles <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565\">provided by Paul</a>. This net finds the main region of interest. The predictions look even better.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/194936/6710/ZF_UNET_v2_predictions.jpg\" alt=\"enter image description here\" title=\"\"></p>\n\n<p>I didn’t use these predictions directly. I only extract bounding box rectangles for each image to form rectangles.csv file which I used in later training process.</p>\n\n<h2>Zoo</h2>\n\n<p>I used the following set of CNNs: VGG16, VGG19, RESNET50, INCEPTION_V3, SQUEEZE_NET, DENSENET_161, DENSENET_121. The training process for them was mostly similar. 5 KFold validation, vary only learning rate parameters and batch size. For each model I obtain: train OOF predictions with same length as number of training images and test predictions.</p>\n\n<h2>Augmentations</h2>\n\n<p>I think augmentation was important key for this problem, since we had very small amount of data. I used: </p>\n\n<ol>\n<li>Random crops based on rectangle.csv generated from UNETs. These crops were in very big range from UNET_v1 prediction to UNET_v2 predictions. </li>\n<li>Random perspective transformation </li>\n<li>Random rotations (mean color or border reflect at random) and mirroring </li>\n<li>Lightning change </li>\n<li>Rare random blur</li>\n</ol>\n\n<p>Basically neural nets never see the same pictures during training.</p>\n\n<h2>Batch generator</h2>\n\n<p>For learning process I used fit_generator function from Keras, which I recommend to use for everyone. You only need to create your own batch_generator function. No need to store many different images in memory or on HDD.\nIn my batch generator I add some fraction of images from “additional” folder ~25%. It greatly improves validation and leaderboard score. And we all know now it’s because of leakage of test images in additional folder.</p>\n\n<h2>Ensemble</h2>\n\n<p>For ensemble I used XGBoost blender. Final solution was average on 500 XGBoost iterations with different seed and random run parameters. </p>\n\n<h2>Submission</h2>\n\n<p>The only difference for my 2 final submissions was to train with usage of “additional” images or without them. My final models used “train” and “test_stg1” images for training.</p>\n\n<p>Submission v1 (with additional): Validation score: 0.57457 Private LB: 0.88856</p>\n\n<p>Submission v2 (without additional): Validation score: 0.64254 Private LB: 0.83209</p>\n\n<p>I believe low scores on private LB depends on bad data preparation, mostly because of mislabeling. We could clean it by hands, but in case private test set had the same labeling quality it would make everything worse.</p>\n\n<h2>Code</h2>\n\n<p>You can find my code as it was prepared for Kaggle <a href=\"https://github.com/ZFTurbo/KAGGLE_CERVICAL_CANCER_2017\">on GitHub</a></p>",
      "rawMarkdown": "Used tools\n--------------\n\nWindows 10 + Python 3.4 + Keras 1.2 + Theano 0.9\n\nZF_UNET_224 (version 1)\n-----------------------\n\nTo find important parts of images I used [my own version of UNET][1]. First UNET was trained on segmentation of train images, [which I made by hands][2] with Sloth. In this segmentation I tried to cut all elements which are totally useless like speculum and everything outside it. The predictions look nice:\n\n![enter image description here][3]\n\nZF_UNET_224 (version 2)\n-----------------------\n\nSecond UNET was trained on rectangles [provided by Paul][4]. This net finds the main region of interest. The predictions look even better.\n\n![enter image description here][5]\n \nI didn’t use these predictions directly. I only extract bounding box rectangles for each image to form rectangles.csv file which I used in later training process.\n\nZoo\n---\n\nI used the following set of CNNs: VGG16, VGG19, RESNET50, INCEPTION_V3, SQUEEZE_NET, DENSENET_161, DENSENET_121. The training process for them was mostly similar. 5 KFold validation, vary only learning rate parameters and batch size. For each model I obtain: train OOF predictions with same length as number of training images and test predictions.\n\nAugmentations\n-------------\n\nI think augmentation was important key for this problem, since we had very small amount of data. I used: \n\n 1. Random crops based on rectangle.csv generated from UNETs. These crops were in very big range from UNET_v1 prediction to UNET_v2 predictions. \n 2. Random perspective transformation \n 3. Random rotations (mean color or border reflect at random) and mirroring \n 4. Lightning change \n 5. Rare random blur\n\nBasically neural nets never see the same pictures during training.\n\nBatch generator\n---------------\n\nFor learning process I used fit_generator function from Keras, which I recommend to use for everyone. You only need to create your own batch_generator function. No need to store many different images in memory or on HDD.\nIn my batch generator I add some fraction of images from “additional” folder ~25%. It greatly improves validation and leaderboard score. And we all know now it’s because of leakage of test images in additional folder.\n\nEnsemble\n--------\n\nFor ensemble I used XGBoost blender. Final solution was average on 500 XGBoost iterations with different seed and random run parameters. \n\nSubmission\n----------\n\nThe only difference for my 2 final submissions was to train with usage of “additional” images or without them. My final models used “train” and “test_stg1” images for training.\n\nSubmission v1 (with additional): Validation score: 0.57457 Private LB: 0.88856\n\nSubmission v2 (without additional): Validation score: 0.64254 Private LB: 0.83209\n\nI believe low scores on private LB depends on bad data preparation, mostly because of mislabeling. We could clean it by hands, but in case private test set had the same labeling quality it would make everything worse.\n\nCode\n----\n\nYou can find my code as it was prepared for Kaggle [on GitHub][6]\n\n\n  [1]: https://github.com/ZFTurbo/ZF_UNET_224_Pretrained_Model\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/194936/6708/polygons.json\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/194936/6709/ZF_UNET_v1_predictions.jpg\n  [4]: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565\n  [5]: https://kaggle2.blob.core.windows.net/forum-message-attachments/194936/6710/ZF_UNET_v2_predictions.jpg\n  [6]: https://github.com/ZFTurbo/KAGGLE_CERVICAL_CANCER_2017",
      "votes": null
    },
    {
      "id": "194941",
      "postDate": "06/22/2017 09:26:18",
      "content": "<p>Thanks for the insights and the codes</p>",
      "rawMarkdown": "Thanks for the insights and the codes",
      "votes": null
    },
    {
      "id": "194942",
      "postDate": "06/22/2017 09:34:14",
      "content": "<p>How important was the usage of bounding boxes? In my experiments i was not able to get any significant improvements from this information. \nIn the end, I just fine-tuned resnet and densenet=)</p>",
      "rawMarkdown": "How important was the usage of bounding boxes? In my experiments i was not able to get any significant improvements from this information. \nIn the end, I just fine-tuned resnet and densenet=)",
      "votes": null
    },
    {
      "id": "194944",
      "postDate": "06/22/2017 09:44:26",
      "content": "<p>With bounding boxes training process was more stable for me, without big jumps and lower speed of overfitting. It's because with boxes I was able to make larger crops without affect of area of interest. In the beginning I used just 10% random crops from all sides of image. </p>\n\n<p>Bounding boxes gave me improvements on stage 1 LB, but I don't remember exact value.</p>",
      "rawMarkdown": "With bounding boxes training process was more stable for me, without big jumps and lower speed of overfitting. It's because with boxes I was able to make larger crops without affect of area of interest. In the beginning I used just 10% random crops from all sides of image. \n\nBounding boxes gave me improvements on stage 1 LB, but I don't remember exact value.",
      "votes": null
    },
    {
      "id": "194946",
      "postDate": "06/22/2017 09:53:00",
      "content": "<p>@utility Yeah, it seems so that bounding boxes weren't much important. The only thing I did was to take a center crop assuming the cervix to be more or less at the center due to the MobileODT tool which has got a cylindrical tube that enters through the vaginal opening. And  I was able to train quite stably on green-channel-replicated-thrice images.</p>",
      "rawMarkdown": "utility Yeah, it seems so that bounding boxes weren't much important. The only thing I did was to take a center crop assuming the cervix to be more or less at the center due to the MobileODT tool which has got a cylindrical tube that enters through the vaginal opening. And  I was able to train quite stably on green-channel-replicated-thrice images.",
      "votes": null
    },
    {
      "id": "194950",
      "postDate": "06/22/2017 10:14:06",
      "content": "<p>My experience was the opposite - for some reason I had a really hard time training anything on crops. It wrecked test time performance every time, even when I was sure I wasn't cropping out anything of relevance. </p>",
      "rawMarkdown": "My experience was the opposite - for some reason I had a really hard time training anything on crops. It wrecked test time performance every time, even when I was sure I wasn't cropping out anything of relevance.",
      "votes": null
    },
    {
      "id": "194951",
      "postDate": "06/22/2017 10:22:04",
      "content": "<p>Congrats on grandmaster title! welcome to the club :)</p>",
      "rawMarkdown": "Congrats on grandmaster title! welcome to the club :)",
      "votes": null
    },
    {
      "id": "195001",
      "postDate": "06/22/2017 14:55:57",
      "content": "<p>Thanks! ) Congrats with another one 1st place :P</p>",
      "rawMarkdown": "Thanks! ) Congrats with another one 1st place :P",
      "votes": null
    },
    {
      "id": "195262",
      "postDate": "06/23/2017 03:50:00",
      "content": "<p>Thanks a lot ZFTurbo for sharing your approaches</p>",
      "rawMarkdown": "Thanks a lot ZFTurbo for sharing your approaches",
      "votes": null
    },
    {
      "id": "195296",
      "postDate": "06/23/2017 07:17:13",
      "content": "<p>Nice work! I learned a lot from the work you done.</p>",
      "rawMarkdown": "Nice work! I learned a lot from the work you done.",
      "votes": null
    },
    {
      "id": "195379",
      "postDate": "06/23/2017 12:45:15",
      "content": "<p>ZFTurbo,\nCongrats once again!\nI have a question about the xgb blender: what are the inputs for the model? </p>",
      "rawMarkdown": "ZFTurbo,\nCongrats once again!\nI have a question about the xgb blender: what are the inputs for the model?",
      "votes": null
    },
    {
      "id": "195383",
      "postDate": "06/23/2017 12:55:00",
      "content": "<p>The input for XGBoost blender are predictions on validation set from each neural net model. So neural net predictions used as features. Validation set has the same length as train since it combined from 5 Folds.</p>",
      "rawMarkdown": "The input for XGBoost blender are predictions on validation set from each neural net model. So neural net predictions used as features. Validation set has the same length as train since it combined from 5 Folds.",
      "votes": null
    },
    {
      "id": "195554",
      "postDate": "06/24/2017 00:57:59",
      "content": "<p>bounding boxes gained about 0.1 for me on the stage 1 LB</p>",
      "rawMarkdown": "bounding boxes gained about 0.1 for me on the stage 1 LB",
      "votes": null
    },
    {
      "id": "199109",
      "postDate": "07/04/2017 16:21:03",
      "content": "<p>Congrats and thanks for sharing! <br>\nJust curious, does XGBoost blender perform better than a simple average blender in your case?</p>",
      "rawMarkdown": "Congrats and thanks for sharing!  \nJust curious, does XGBoost blender perform better than a simple average blender in your case?",
      "votes": null
    },
    {
      "id": "199173",
      "postDate": "07/04/2017 19:57:18",
      "content": "<p>I stopped comparing in the early beginning of contest. In my first attempts XGBoost blender was better.</p>",
      "rawMarkdown": "I stopped comparing in the early beginning of contest. In my first attempts XGBoost blender was better.",
      "votes": null
    },
    {
      "id": "199836",
      "postDate": "07/06/2017 12:31:48",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "200878",
      "postDate": "07/09/2017 17:43:25",
      "content": "<p>Great work. Thanks for sharing! Congratulations!</p>",
      "rawMarkdown": "Great work. Thanks for sharing! Congratulations!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 194941,
      "author_name": "xiaomen",
      "author_url": "",
      "post_date": "06/22/2017 09:26:18",
      "content": "<p>Thanks for the insights and the codes</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 194942,
      "author_name": "utility",
      "author_url": "",
      "post_date": "06/22/2017 09:34:14",
      "content": "<p>How important was the usage of bounding boxes? In my experiments i was not able to get any significant improvements from this information. \nIn the end, I just fine-tuned resnet and densenet=)</p>",
      "votes": null,
      "replies": [
        {
          "id": 194944,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "06/22/2017 09:44:26",
          "content": "<p>With bounding boxes training process was more stable for me, without big jumps and lower speed of overfitting. It's because with boxes I was able to make larger crops without affect of area of interest. In the beginning I used just 10% random crops from all sides of image. </p>\n\n<p>Bounding boxes gave me improvements on stage 1 LB, but I don't remember exact value.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194946,
          "author_name": "bvineeth007",
          "author_url": "",
          "post_date": "06/22/2017 09:53:00",
          "content": "<p>@utility Yeah, it seems so that bounding boxes weren't much important. The only thing I did was to take a center crop assuming the cervix to be more or less at the center due to the MobileODT tool which has got a cylindrical tube that enters through the vaginal opening. And  I was able to train quite stably on green-channel-replicated-thrice images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194950,
          "author_name": "zadrras",
          "author_url": "",
          "post_date": "06/22/2017 10:14:06",
          "content": "<p>My experience was the opposite - for some reason I had a really hard time training anything on crops. It wrecked test time performance every time, even when I was sure I wasn't cropping out anything of relevance. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 195554,
          "author_name": "kubilai",
          "author_url": "",
          "post_date": "06/24/2017 00:57:59",
          "content": "<p>bounding boxes gained about 0.1 for me on the stage 1 LB</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 194951,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "06/22/2017 10:22:04",
      "content": "<p>Congrats on grandmaster title! welcome to the club :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 195001,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "06/22/2017 14:55:57",
          "content": "<p>Thanks! ) Congrats with another one 1st place :P</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 195262,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "06/23/2017 03:50:00",
      "content": "<p>Thanks a lot ZFTurbo for sharing your approaches</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 195296,
      "author_name": "badoun",
      "author_url": "",
      "post_date": "06/23/2017 07:17:13",
      "content": "<p>Nice work! I learned a lot from the work you done.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 195379,
      "author_name": "chabir",
      "author_url": "",
      "post_date": "06/23/2017 12:45:15",
      "content": "<p>ZFTurbo,\nCongrats once again!\nI have a question about the xgb blender: what are the inputs for the model? </p>",
      "votes": null,
      "replies": [
        {
          "id": 195383,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "06/23/2017 12:55:00",
          "content": "<p>The input for XGBoost blender are predictions on validation set from each neural net model. So neural net predictions used as features. Validation set has the same length as train since it combined from 5 Folds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 199109,
          "author_name": "scottykwok",
          "author_url": "",
          "post_date": "07/04/2017 16:21:03",
          "content": "<p>Congrats and thanks for sharing! <br>\nJust curious, does XGBoost blender perform better than a simple average blender in your case?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 199173,
          "author_name": "zfturbo",
          "author_url": "",
          "post_date": "07/04/2017 19:57:18",
          "content": "<p>I stopped comparing in the early beginning of contest. In my first attempts XGBoost blender was better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 199836,
      "author_name": "pavlin",
      "author_url": "",
      "post_date": "07/06/2017 12:31:48",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 200878,
      "author_name": "cpattekar",
      "author_url": "",
      "post_date": "07/09/2017 17:43:25",
      "content": "<p>Great work. Thanks for sharing! Congratulations!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "194936": "Used tools\n--------------\n\nWindows 10 + Python 3.4 + Keras 1.2 + Theano 0.9\n\nZF_UNET_224 (version 1)\n-----------------------\n\nTo find important parts of images I used [my own version of UNET][1]. First UNET was trained on segmentation of train images, [which I made by hands][2] with Sloth. In this segmentation I tried to cut all elements which are totally useless like speculum and everything outside it. The predictions look nice:\n\n![enter image description here][3]\n\nZF_UNET_224 (version 2)\n-----------------------\n\nSecond UNET was trained on rectangles [provided by Paul][4]. This net finds the main region of interest. The predictions look even better.\n\n![enter image description here][5]\n \nI didn’t use these predictions directly. I only extract bounding box rectangles for each image to form rectangles.csv file which I used in later training process.\n\nZoo\n---\n\nI used the following set of CNNs: VGG16, VGG19, RESNET50, INCEPTION_V3, SQUEEZE_NET, DENSENET_161, DENSENET_121. The training process for them was mostly similar. 5 KFold validation, vary only learning rate parameters and batch size. For each model I obtain: train OOF predictions with same length as number of training images and test predictions.\n\nAugmentations\n-------------\n\nI think augmentation was important key for this problem, since we had very small amount of data. I used: \n\n 1. Random crops based on rectangle.csv generated from UNETs. These crops were in very big range from UNET_v1 prediction to UNET_v2 predictions. \n 2. Random perspective transformation \n 3. Random rotations (mean color or border reflect at random) and mirroring \n 4. Lightning change \n 5. Rare random blur\n\nBasically neural nets never see the same pictures during training.\n\nBatch generator\n---------------\n\nFor learning process I used fit_generator function from Keras, which I recommend to use for everyone. You only need to create your own batch_generator function. No need to store many different images in memory or on HDD.\nIn my batch generator I add some fraction of images from “additional” folder ~25%. It greatly improves validation and leaderboard score. And we all know now it’s because of leakage of test images in additional folder.\n\nEnsemble\n--------\n\nFor ensemble I used XGBoost blender. Final solution was average on 500 XGBoost iterations with different seed and random run parameters. \n\nSubmission\n----------\n\nThe only difference for my 2 final submissions was to train with usage of “additional” images or without them. My final models used “train” and “test_stg1” images for training.\n\nSubmission v1 (with additional): Validation score: 0.57457 Private LB: 0.88856\n\nSubmission v2 (without additional): Validation score: 0.64254 Private LB: 0.83209\n\nI believe low scores on private LB depends on bad data preparation, mostly because of mislabeling. We could clean it by hands, but in case private test set had the same labeling quality it would make everything worse.\n\nCode\n----\n\nYou can find my code as it was prepared for Kaggle [on GitHub][6]\n\n\n  [1]: https://github.com/ZFTurbo/ZF_UNET_224_Pretrained_Model\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/194936/6708/polygons.json\n  [3]: https://kaggle2.blob.core.windows.net/forum-message-attachments/194936/6709/ZF_UNET_v1_predictions.jpg\n  [4]: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/31565\n  [5]: https://kaggle2.blob.core.windows.net/forum-message-attachments/194936/6710/ZF_UNET_v2_predictions.jpg\n  [6]: https://github.com/ZFTurbo/KAGGLE_CERVICAL_CANCER_2017",
    "194941": "Thanks for the insights and the codes",
    "194942": "How important was the usage of bounding boxes? In my experiments i was not able to get any significant improvements from this information. \nIn the end, I just fine-tuned resnet and densenet=)",
    "194944": "With bounding boxes training process was more stable for me, without big jumps and lower speed of overfitting. It's because with boxes I was able to make larger crops without affect of area of interest. In the beginning I used just 10% random crops from all sides of image. \n\nBounding boxes gave me improvements on stage 1 LB, but I don't remember exact value.",
    "194946": "utility Yeah, it seems so that bounding boxes weren't much important. The only thing I did was to take a center crop assuming the cervix to be more or less at the center due to the MobileODT tool which has got a cylindrical tube that enters through the vaginal opening. And  I was able to train quite stably on green-channel-replicated-thrice images.",
    "194950": "My experience was the opposite - for some reason I had a really hard time training anything on crops. It wrecked test time performance every time, even when I was sure I wasn't cropping out anything of relevance.",
    "194951": "Congrats on grandmaster title! welcome to the club :)",
    "195001": "Thanks! ) Congrats with another one 1st place :P",
    "195262": "Thanks a lot ZFTurbo for sharing your approaches",
    "195296": "Nice work! I learned a lot from the work you done.",
    "195379": "ZFTurbo,\nCongrats once again!\nI have a question about the xgb blender: what are the inputs for the model?",
    "195383": "The input for XGBoost blender are predictions on validation set from each neural net model. So neural net predictions used as features. Validation set has the same length as train since it combined from 5 Folds.",
    "195554": "bounding boxes gained about 0.1 for me on the stage 1 LB",
    "199109": "Congrats and thanks for sharing!  \nJust curious, does XGBoost blender perform better than a simple average blender in your case?",
    "199173": "I stopped comparing in the early beginning of contest. In my first attempts XGBoost blender was better.",
    "199836": "Thanks for sharing!",
    "200878": "Great work. Thanks for sharing! Congratulations!"
  },
  "source": "meta"
}