{
  "id": 30489,
  "title": "This might save some of us some time! Also, help me make sense of my own observations",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/30489",
  "author_name": "",
  "post_date": "2017-03-22T05:47:51.247984Z",
  "votes": 6,
  "comment_count": 7,
  "views": 2,
  "content": "<p>As seen by my previous post <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30395\">here</a>, I have had quite a lot of trouble with training on the additional images.\nAfter throwing everything but the kitchen sink at the problem, here are a few observations I have made with regards to various \"experiments\":</p>\n\n<ol>\n<li><p><strong>Experiment 1:</strong> Using Additional data for training and validation</p>\n\n<ul><li>90% for training, 10% for validation</li>\n<li>Augmentation: Horizontal Flips and Zoom</li>\n<li>training on various models from scratch as well as fine-tuning Vgg16 shows good \nresults, achieving accuracy scores of &gt;70% on the validation set without a lot of \ntinkering.</li>\n<li>However Piss poor results when validating using the training data (I didnt waste a \nsubmission on this one).</li></ul></li>\n<li><p><strong>Experiment 2:</strong> Using Additional data for training and original train data</p>\n\n<ul><li>Training set: additional data with Horizontal flips and random zooms.</li>\n<li>validation on the entire original train data.</li>\n<li>Training on the same models as Experiment 1.</li>\n<li>WILD OVERFITTING, never jumps around 54% accuracy on the validation (original \ntrain data)</li>\n<li>1.43446 LB Score</li></ul></li>\n</ol>\n\n<p>PS: I have ran these experiments numerous times</p>\n\n<p>I believe the above is occurring due to</p>\n\n<ol>\n<li><p>Bad image quality in Additional set: A lot of images are very blurry.</p></li>\n<li><p>Additional Set comprises of multiple images from the same patient, whereas train set comprises of one image per person. Training on Additional set creates a \"person\" predictor rather than a disease predictor and this falls on it's face when ran on the train set. (thanks @<a href=\"https://www.kaggle.com/visoft\">visoft</a> for helping me)</p></li>\n</ol>\n\n<p>Please let me know what do you think about it and correct me where I might be wrong.</p>",
  "messages": [
    {
      "id": "169696",
      "postDate": "03/22/2017 05:47:51",
      "content": "<p>As seen by my previous post <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30395\">here</a>, I have had quite a lot of trouble with training on the additional images.\nAfter throwing everything but the kitchen sink at the problem, here are a few observations I have made with regards to various \"experiments\":</p>\n\n<ol>\n<li><p><strong>Experiment 1:</strong> Using Additional data for training and validation</p>\n\n<ul><li>90% for training, 10% for validation</li>\n<li>Augmentation: Horizontal Flips and Zoom</li>\n<li>training on various models from scratch as well as fine-tuning Vgg16 shows good \nresults, achieving accuracy scores of &gt;70% on the validation set without a lot of \ntinkering.</li>\n<li>However Piss poor results when validating using the training data (I didnt waste a \nsubmission on this one).</li></ul></li>\n<li><p><strong>Experiment 2:</strong> Using Additional data for training and original train data</p>\n\n<ul><li>Training set: additional data with Horizontal flips and random zooms.</li>\n<li>validation on the entire original train data.</li>\n<li>Training on the same models as Experiment 1.</li>\n<li>WILD OVERFITTING, never jumps around 54% accuracy on the validation (original \ntrain data)</li>\n<li>1.43446 LB Score</li></ul></li>\n</ol>\n\n<p>PS: I have ran these experiments numerous times</p>\n\n<p>I believe the above is occurring due to</p>\n\n<ol>\n<li><p>Bad image quality in Additional set: A lot of images are very blurry.</p></li>\n<li><p>Additional Set comprises of multiple images from the same patient, whereas train set comprises of one image per person. Training on Additional set creates a \"person\" predictor rather than a disease predictor and this falls on it's face when ran on the train set. (thanks @<a href=\"https://www.kaggle.com/visoft\">visoft</a> for helping me)</p></li>\n</ol>\n\n<p>Please let me know what do you think about it and correct me where I might be wrong.</p>",
      "rawMarkdown": "As seen by my previous post [here][1], I have had quite a lot of trouble with training on the additional images.\nAfter throwing everything but the kitchen sink at the problem, here are a few observations I have made with regards to various \"experiments\":\n\n1. **Experiment 1:** Using Additional data for training and validation\n     - 90% for training, 10% for validation\n     - Augmentation: Horizontal Flips and Zoom\n     - training on various models from scratch as well as fine-tuning Vgg16 shows good \n        results, achieving accuracy scores of >70% on the validation set without a lot of \n        tinkering.\n     - However Piss poor results when validating using the training data (I didnt waste a \n        submission on this one).\n\n2. **Experiment 2:** Using Additional data for training and original train data\n     - Training set: additional data with Horizontal flips and random zooms.\n     - validation on the entire original train data.\n     - Training on the same models as Experiment 1.\n     - WILD OVERFITTING, never jumps around 54% accuracy on the validation (original \n        train data)\n     - 1.43446 LB Score\n\nPS: I have ran these experiments numerous times\n\nI believe the above is occurring due to\n\n1. Bad image quality in Additional set: A lot of images are very blurry.\n\n2. Additional Set comprises of multiple images from the same patient, whereas train set comprises of one image per person. Training on Additional set creates a \"person\" predictor rather than a disease predictor and this falls on it's face when ran on the train set. (thanks @[visoft][2] for helping me)\n\nPlease let me know what do you think about it and correct me where I might be wrong.\n\n  [1]: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30395\n  [2]: https://www.kaggle.com/visoft",
      "votes": null
    },
    {
      "id": "169779",
      "postDate": "03/22/2017 15:11:47",
      "content": "<p>I think it learns the characteristics based on the average color and the setup (gloves, speculum position/color, etc...), so what about removing the blurry images in an automated way and then extracting the remaining cervix regions?\nIf it is well done, I am sure it could decrease the tendency to learn the patient's features.</p>",
      "rawMarkdown": "I think it learns the characteristics based on the average color and the setup (gloves, speculum position/color, etc...), so what about removing the blurry images in an automated way and then extracting the remaining cervix regions?\nIf it is well done, I am sure it could decrease the tendency to learn the patient's features.",
      "votes": null
    },
    {
      "id": "169789",
      "postDate": "03/22/2017 16:18:15",
      "content": "<p>I think this will be a project won by someone who can be bothered to do a lot of data cleaning, processing, manual annotation etc. It should be quite a straight forward classification task otherwise.</p>",
      "rawMarkdown": "I think this will be a project won by someone who can be bothered to do a lot of data cleaning, processing, manual annotation etc. It should be quite a straight forward classification task otherwise.",
      "votes": null
    },
    {
      "id": "169799",
      "postDate": "03/22/2017 17:08:22",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "175512",
      "postDate": "04/15/2017 21:05:13",
      "content": "<p>So, any progress on this? I haven't been able to train on the additional data either. My validation loss stays around 1.0 when I add the additional data to the training set (I split off 10% of training data for validation). I have tried a few different models, and getting the same poor results. </p>",
      "rawMarkdown": "So, any progress on this? I haven't been able to train on the additional data either. My validation loss stays around 1.0 when I add the additional data to the training set (I split off 10% of training data for validation). I have tried a few different models, and getting the same poor results.",
      "votes": null
    },
    {
      "id": "180664",
      "postDate": "05/06/2017 13:20:03",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "180665",
      "postDate": "05/06/2017 13:22:57",
      "content": "<p>Hi Sarthak,</p>\n\n<p>As you said, many images are duplicated in the additional datasets. Even worse, some images are absolutely unrelated to cervix. I show a few of them at the end of <a href=\"https://www.kaggle.com/deveaup/checking-bounding-boxes-and-additional-dataset\">this kernel</a> if it can help.</p>",
      "rawMarkdown": "Hi Sarthak,\n\nAs you said, many images are duplicated in the additional datasets. Even worse, some images are absolutely unrelated to cervix. I show a few of them at the end of [this kernel](https://www.kaggle.com/deveaup/checking-bounding-boxes-and-additional-dataset) if it can help.",
      "votes": null
    },
    {
      "id": "189239",
      "postDate": "06/05/2017 14:47:45",
      "content": "<p>Hi SarthakYadav,</p>\n\n<p>Its getting closer to the end of the competition and I am wondering if you ever got to a point where the additional data is useful? I won't push for details just trying to figure out how to prioritize my remaining time.</p>",
      "rawMarkdown": "Hi SarthakYadav,\n\nIts getting closer to the end of the competition and I am wondering if you ever got to a point where the additional data is useful? I won't push for details just trying to figure out how to prioritize my remaining time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 169779,
      "author_name": "npielawski",
      "author_url": "",
      "post_date": "03/22/2017 15:11:47",
      "content": "<p>I think it learns the characteristics based on the average color and the setup (gloves, speculum position/color, etc...), so what about removing the blurry images in an automated way and then extracting the remaining cervix regions?\nIf it is well done, I am sure it could decrease the tendency to learn the patient's features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 169789,
          "author_name": "craigglastonbury",
          "author_url": "",
          "post_date": "03/22/2017 16:18:15",
          "content": "<p>I think this will be a project won by someone who can be bothered to do a lot of data cleaning, processing, manual annotation etc. It should be quite a straight forward classification task otherwise.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 169799,
          "author_name": "yadavsarthak",
          "author_url": "",
          "post_date": "03/22/2017 17:08:22",
          "content": "",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 175512,
      "author_name": "theeasyway8",
      "author_url": "",
      "post_date": "04/15/2017 21:05:13",
      "content": "<p>So, any progress on this? I haven't been able to train on the additional data either. My validation loss stays around 1.0 when I add the additional data to the training set (I split off 10% of training data for validation). I have tried a few different models, and getting the same poor results. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 180664,
      "author_name": "deveaup",
      "author_url": "",
      "post_date": "05/06/2017 13:20:03",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 180665,
      "author_name": "deveaup",
      "author_url": "",
      "post_date": "05/06/2017 13:22:57",
      "content": "<p>Hi Sarthak,</p>\n\n<p>As you said, many images are duplicated in the additional datasets. Even worse, some images are absolutely unrelated to cervix. I show a few of them at the end of <a href=\"https://www.kaggle.com/deveaup/checking-bounding-boxes-and-additional-dataset\">this kernel</a> if it can help.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 189239,
      "author_name": "gkericks",
      "author_url": "",
      "post_date": "06/05/2017 14:47:45",
      "content": "<p>Hi SarthakYadav,</p>\n\n<p>Its getting closer to the end of the competition and I am wondering if you ever got to a point where the additional data is useful? I won't push for details just trying to figure out how to prioritize my remaining time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "169696": "As seen by my previous post [here][1], I have had quite a lot of trouble with training on the additional images.\nAfter throwing everything but the kitchen sink at the problem, here are a few observations I have made with regards to various \"experiments\":\n\n1. **Experiment 1:** Using Additional data for training and validation\n     - 90% for training, 10% for validation\n     - Augmentation: Horizontal Flips and Zoom\n     - training on various models from scratch as well as fine-tuning Vgg16 shows good \n        results, achieving accuracy scores of >70% on the validation set without a lot of \n        tinkering.\n     - However Piss poor results when validating using the training data (I didnt waste a \n        submission on this one).\n\n2. **Experiment 2:** Using Additional data for training and original train data\n     - Training set: additional data with Horizontal flips and random zooms.\n     - validation on the entire original train data.\n     - Training on the same models as Experiment 1.\n     - WILD OVERFITTING, never jumps around 54% accuracy on the validation (original \n        train data)\n     - 1.43446 LB Score\n\nPS: I have ran these experiments numerous times\n\nI believe the above is occurring due to\n\n1. Bad image quality in Additional set: A lot of images are very blurry.\n\n2. Additional Set comprises of multiple images from the same patient, whereas train set comprises of one image per person. Training on Additional set creates a \"person\" predictor rather than a disease predictor and this falls on it's face when ran on the train set. (thanks @[visoft][2] for helping me)\n\nPlease let me know what do you think about it and correct me where I might be wrong.\n\n  [1]: https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/30395\n  [2]: https://www.kaggle.com/visoft",
    "169779": "I think it learns the characteristics based on the average color and the setup (gloves, speculum position/color, etc...), so what about removing the blurry images in an automated way and then extracting the remaining cervix regions?\nIf it is well done, I am sure it could decrease the tendency to learn the patient's features.",
    "169789": "I think this will be a project won by someone who can be bothered to do a lot of data cleaning, processing, manual annotation etc. It should be quite a straight forward classification task otherwise.",
    "169799": "",
    "175512": "So, any progress on this? I haven't been able to train on the additional data either. My validation loss stays around 1.0 when I add the additional data to the training set (I split off 10% of training data for validation). I have tried a few different models, and getting the same poor results.",
    "180664": "",
    "180665": "Hi Sarthak,\n\nAs you said, many images are duplicated in the additional datasets. Even worse, some images are absolutely unrelated to cervix. I show a few of them at the end of [this kernel](https://www.kaggle.com/deveaup/checking-bounding-boxes-and-additional-dataset) if it can help.",
    "189239": "Hi SarthakYadav,\n\nIts getting closer to the end of the competition and I am wondering if you ever got to a point where the additional data is useful? I won't push for details just trying to figure out how to prioritize my remaining time."
  },
  "source": "meta"
}