{
  "id": 35478,
  "title": "Brief overview of #2 solution",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/writeups/i-rustandi-brief-overview-of-2-solution",
  "author_name": "",
  "post_date": "2017-06-29T09:48:13.876067300Z",
  "votes": 19,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thanks for Intel and MobileODT for sponsoring this competition, and for Kaggle for hosting it. Hopefully, the results can indeed make a significant impact in early detection and treatment of cervical cancer.</p>\n\n<p>For this competition, I decided to trust my own validation. So I did not probe the test labels during stage 1; in fact, I did not make any submissions at all during stage 1. Also my final models did not incorporate any of the stage 1 test data. I did, however, use the additional data in addition to the training data.</p>\n\n<p>My validation set consists of 10% of the train data; the validation set did not include any additional data. Unlike the number 1 team's approach, I did not try to make sure that the patients in the validation data were not present in the remaining data; in hindsight, had I done this, it might have reduced overfitting in the models I trained. In any case, for the remaining train data + additional data, I went through each image to flag those that I deemed to be of bad quality (blurry, or just not relevant at all). This exercise left me with around 5000 images available to train my models.</p>\n\n<p>Out of these 5000 images I created two sets:</p>\n\n<ul>\n<li>256x256 image scaled from the original</li>\n<li>256x256 image cropped using the cervix segmentation kernel <a href=\"https://www.kaggle.com/chattob/cervix-segmentation-gmm\">here</a>.</li>\n</ul>\n\n<p>Several models using as base the densenet161 and resnet152 models (with custom classification layers, just basic fully connected layers with custom number of hidden layers) were trained on these two sets, some data augmentation were included during training. The final submission consisted of a simple blend of the predictions of each of these models, inspired by Felix Laumon's <a href=\"https://github.com/felixlaumon/kaggle-right-whale/blob/master/scripts/blend_submission_files.py\">script</a>. FWIW, I used PyTorch.</p>",
  "messages": [
    {
      "id": "197314",
      "postDate": "06/29/2017 09:48:13",
      "content": "<p>Thanks for Intel and MobileODT for sponsoring this competition, and for Kaggle for hosting it. Hopefully, the results can indeed make a significant impact in early detection and treatment of cervical cancer.</p>\n\n<p>For this competition, I decided to trust my own validation. So I did not probe the test labels during stage 1; in fact, I did not make any submissions at all during stage 1. Also my final models did not incorporate any of the stage 1 test data. I did, however, use the additional data in addition to the training data.</p>\n\n<p>My validation set consists of 10% of the train data; the validation set did not include any additional data. Unlike the number 1 team's approach, I did not try to make sure that the patients in the validation data were not present in the remaining data; in hindsight, had I done this, it might have reduced overfitting in the models I trained. In any case, for the remaining train data + additional data, I went through each image to flag those that I deemed to be of bad quality (blurry, or just not relevant at all). This exercise left me with around 5000 images available to train my models.</p>\n\n<p>Out of these 5000 images I created two sets:</p>\n\n<ul>\n<li>256x256 image scaled from the original</li>\n<li>256x256 image cropped using the cervix segmentation kernel <a href=\"https://www.kaggle.com/chattob/cervix-segmentation-gmm\">here</a>.</li>\n</ul>\n\n<p>Several models using as base the densenet161 and resnet152 models (with custom classification layers, just basic fully connected layers with custom number of hidden layers) were trained on these two sets, some data augmentation were included during training. The final submission consisted of a simple blend of the predictions of each of these models, inspired by Felix Laumon's <a href=\"https://github.com/felixlaumon/kaggle-right-whale/blob/master/scripts/blend_submission_files.py\">script</a>. FWIW, I used PyTorch.</p>",
      "rawMarkdown": "Thanks for Intel and MobileODT for sponsoring this competition, and for Kaggle for hosting it. Hopefully, the results can indeed make a significant impact in early detection and treatment of cervical cancer.\n\nFor this competition, I decided to trust my own validation. So I did not probe the test labels during stage 1; in fact, I did not make any submissions at all during stage 1. Also my final models did not incorporate any of the stage 1 test data. I did, however, use the additional data in addition to the training data.\n\nMy validation set consists of 10% of the train data; the validation set did not include any additional data. Unlike the number 1 team's approach, I did not try to make sure that the patients in the validation data were not present in the remaining data; in hindsight, had I done this, it might have reduced overfitting in the models I trained. In any case, for the remaining train data + additional data, I went through each image to flag those that I deemed to be of bad quality (blurry, or just not relevant at all). This exercise left me with around 5000 images available to train my models.\n\nOut of these 5000 images I created two sets:\n\n* 256x256 image scaled from the original\n* 256x256 image cropped using the cervix segmentation kernel [here](https://www.kaggle.com/chattob/cervix-segmentation-gmm).\n\nSeveral models using as base the densenet161 and resnet152 models (with custom classification layers, just basic fully connected layers with custom number of hidden layers) were trained on these two sets, some data augmentation were included during training. The final submission consisted of a simple blend of the predictions of each of these models, inspired by Felix Laumon's [script](https://github.com/felixlaumon/kaggle-right-whale/blob/master/scripts/blend_submission_files.py). FWIW, I used PyTorch.",
      "votes": null
    },
    {
      "id": "199623",
      "postDate": "07/06/2017 01:15:31",
      "content": "<p>Nice job, thanks for sharing</p>",
      "rawMarkdown": "Nice job, thanks for sharing",
      "votes": null
    },
    {
      "id": "832348",
      "postDate": "05/04/2020 04:32:15",
      "content": "<p>How do you work with two dataset?\nKeeping them in one file or ?</p>",
      "rawMarkdown": "How do you work with two dataset?\nKeeping them in one file or ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 199623,
      "author_name": "sontung1404",
      "author_url": "",
      "post_date": "07/06/2017 01:15:31",
      "content": "<p>Nice job, thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 832348,
      "author_name": "kitoislam",
      "author_url": "",
      "post_date": "05/04/2020 04:32:15",
      "content": "<p>How do you work with two dataset?\nKeeping them in one file or ?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "197314": "Thanks for Intel and MobileODT for sponsoring this competition, and for Kaggle for hosting it. Hopefully, the results can indeed make a significant impact in early detection and treatment of cervical cancer.\n\nFor this competition, I decided to trust my own validation. So I did not probe the test labels during stage 1; in fact, I did not make any submissions at all during stage 1. Also my final models did not incorporate any of the stage 1 test data. I did, however, use the additional data in addition to the training data.\n\nMy validation set consists of 10% of the train data; the validation set did not include any additional data. Unlike the number 1 team's approach, I did not try to make sure that the patients in the validation data were not present in the remaining data; in hindsight, had I done this, it might have reduced overfitting in the models I trained. In any case, for the remaining train data + additional data, I went through each image to flag those that I deemed to be of bad quality (blurry, or just not relevant at all). This exercise left me with around 5000 images available to train my models.\n\nOut of these 5000 images I created two sets:\n\n* 256x256 image scaled from the original\n* 256x256 image cropped using the cervix segmentation kernel [here](https://www.kaggle.com/chattob/cervix-segmentation-gmm).\n\nSeveral models using as base the densenet161 and resnet152 models (with custom classification layers, just basic fully connected layers with custom number of hidden layers) were trained on these two sets, some data augmentation were included during training. The final submission consisted of a simple blend of the predictions of each of these models, inspired by Felix Laumon's [script](https://github.com/felixlaumon/kaggle-right-whale/blob/master/scripts/blend_submission_files.py). FWIW, I used PyTorch.",
    "199623": "Nice job, thanks for sharing",
    "832348": "How do you work with two dataset?\nKeeping them in one file or ?"
  },
  "source": "meta"
}