{
  "id": 34177,
  "title": "Need some hints !",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/34177",
  "author_name": "",
  "post_date": "2017-06-05T11:23:07.646755Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi Kagglers,</p>\n\n<p>I am a bit confused about the approach that I am following. I have started to use several approaches for this task such as, Inception V3 pre-trained, Small CNN, VGG16 and etc but still could not get any good results. At first, I trained the data-set without doing GMM segmentation and I just used train and additional train images as validation and training sets.  Well, up to some extends, it could work but not satisfactory validation loss. \nIn the second step, I applied GMM segmentation, Data Augmentation, hyper-parameters tuning, removing corrupted images and feed the crop images to several networks but the maximum validation accuracy that I could get after 500 Epochs with IncpetionV3 fine tuned a validation accuracy of 64% only. </p>\n\n<p>Do you have any recommendation ?</p>",
  "messages": [
    {
      "id": "189151",
      "postDate": "06/05/2017 11:23:07",
      "content": "<p>Hi Kagglers,</p>\n\n<p>I am a bit confused about the approach that I am following. I have started to use several approaches for this task such as, Inception V3 pre-trained, Small CNN, VGG16 and etc but still could not get any good results. At first, I trained the data-set without doing GMM segmentation and I just used train and additional train images as validation and training sets.  Well, up to some extends, it could work but not satisfactory validation loss. \nIn the second step, I applied GMM segmentation, Data Augmentation, hyper-parameters tuning, removing corrupted images and feed the crop images to several networks but the maximum validation accuracy that I could get after 500 Epochs with IncpetionV3 fine tuned a validation accuracy of 64% only. </p>\n\n<p>Do you have any recommendation ?</p>",
      "rawMarkdown": "Hi Kagglers,\n\nI am a bit confused about the approach that I am following. I have started to use several approaches for this task such as, Inception V3 pre-trained, Small CNN, VGG16 and etc but still could not get any good results. At first, I trained the data-set without doing GMM segmentation and I just used train and additional train images as validation and training sets.  Well, up to some extends, it could work but not satisfactory validation loss. \nIn the second step, I applied GMM segmentation, Data Augmentation, hyper-parameters tuning, removing corrupted images and feed the crop images to several networks but the maximum validation accuracy that I could get after 500 Epochs with IncpetionV3 fine tuned a validation accuracy of 64% only. \n\nDo you have any recommendation ?",
      "votes": null
    },
    {
      "id": "189227",
      "postDate": "06/05/2017 14:21:07",
      "content": "<p>We're all in the same boat. People who have the highest accuracy on the LB say they feel their score is due to data leakage, but that the best <em>accuracy</em> we should probably expect to get is 60-80%. If you are +70 without data leakage, then you're likely 'in the money'. Wish I could offer more advice, but... just keep tinkering and best of luck(?) !</p>",
      "rawMarkdown": "We're all in the same boat. People who have the highest accuracy on the LB say they feel their score is due to data leakage, but that the best _accuracy_ we should probably expect to get is 60-80%. If you are +70 without data leakage, then you're likely 'in the money'. Wish I could offer more advice, but... just keep tinkering and best of luck(?) !",
      "votes": null
    },
    {
      "id": "189234",
      "postDate": "06/05/2017 14:39:17",
      "content": "<p>Hi authman,</p>\n\n<p>Is there another thread with more details on this the I might have missed? Especially the \"People who have the highest accuracy on the LB say they feel their score is due to data leakage\" part. I am aware that the 0 loss entries are taking advantage of the leak and that many top entries might have misleading results but I had also assumed that some of the high entries must be legitimate.  I'm also curious what you feel the +70 accuracy is in terms of log-loss. Thanks!</p>",
      "rawMarkdown": "Hi authman,\n\nIs there another thread with more details on this the I might have missed? Especially the \"People who have the highest accuracy on the LB say they feel their score is due to data leakage\" part. I am aware that the 0 loss entries are taking advantage of the leak and that many top entries might have misleading results but I had also assumed that some of the high entries must be legitimate.  I'm also curious what you feel the +70 accuracy is in terms of log-loss. Thanks!",
      "votes": null
    },
    {
      "id": "189235",
      "postDate": "06/05/2017 14:42:12",
      "content": "<p>Hi Aramis,</p>\n\n<p>I am also having trouble training a network that isn't fine-tuned (pre-trained). I can get a VGG-like network to train but not with satisfactory results. I'm curious if you or anyone else has any insights they are willing to share about what the problematic part of this dataset is. Are the networks people are trying too deep? Too wide? Not sufficiently regularized? Thanks.</p>",
      "rawMarkdown": "Hi Aramis,\n\nI am also having trouble training a network that isn't fine-tuned (pre-trained). I can get a VGG-like network to train but not with satisfactory results. I'm curious if you or anyone else has any insights they are willing to share about what the problematic part of this dataset is. Are the networks people are trying too deep? Too wide? Not sufficiently regularized? Thanks.",
      "votes": null
    },
    {
      "id": "189556",
      "postDate": "06/06/2017 10:40:08",
      "content": "<p>Hi gKericks, </p>\n\n<p>I have already tried with VGG16-like that isn't fine-tuned (pre-trained) but it worked really worse, then I started to do real time augmentation using Keras which helped a bit. Due to this, I just shift back to my lovely InceptionV3 pre-trained network which worked better for me. In your case, localization of cervical is an important part, as far as I know those who are in the top of LB used localization as well as getting advantage of data leakage.  </p>",
      "rawMarkdown": "Hi gKericks, \n\nI have already tried with VGG16-like that isn't fine-tuned (pre-trained) but it worked really worse, then I started to do real time augmentation using Keras which helped a bit. Due to this, I just shift back to my lovely InceptionV3 pre-trained network which worked better for me. In your case, localization of cervical is an important part, as far as I know those who are in the top of LB used localization as well as getting advantage of data leakage.",
      "votes": null
    },
    {
      "id": "189560",
      "postDate": "06/06/2017 10:45:38",
      "content": "<p>You should use an ensemble of a few nets, since single net produces pretty confident (to itself) results. And in case some image was mispredicted, your log loss will blow up. Ensembling softens results and deals with outliers. In case you've got only 1 trained net, try to soften the predictions to avoid very low scores for certain classes.</p>",
      "rawMarkdown": "You should use an ensemble of a few nets, since single net produces pretty confident (to itself) results. And in case some image was mispredicted, your log loss will blow up. Ensembling softens results and deals with outliers. In case you've got only 1 trained net, try to soften the predictions to avoid very low scores for certain classes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 189227,
      "author_name": "authman",
      "author_url": "",
      "post_date": "06/05/2017 14:21:07",
      "content": "<p>We're all in the same boat. People who have the highest accuracy on the LB say they feel their score is due to data leakage, but that the best <em>accuracy</em> we should probably expect to get is 60-80%. If you are +70 without data leakage, then you're likely 'in the money'. Wish I could offer more advice, but... just keep tinkering and best of luck(?) !</p>",
      "votes": null,
      "replies": [
        {
          "id": 189234,
          "author_name": "gkericks",
          "author_url": "",
          "post_date": "06/05/2017 14:39:17",
          "content": "<p>Hi authman,</p>\n\n<p>Is there another thread with more details on this the I might have missed? Especially the \"People who have the highest accuracy on the LB say they feel their score is due to data leakage\" part. I am aware that the 0 loss entries are taking advantage of the leak and that many top entries might have misleading results but I had also assumed that some of the high entries must be legitimate.  I'm also curious what you feel the +70 accuracy is in terms of log-loss. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 189235,
      "author_name": "gkericks",
      "author_url": "",
      "post_date": "06/05/2017 14:42:12",
      "content": "<p>Hi Aramis,</p>\n\n<p>I am also having trouble training a network that isn't fine-tuned (pre-trained). I can get a VGG-like network to train but not with satisfactory results. I'm curious if you or anyone else has any insights they are willing to share about what the problematic part of this dataset is. Are the networks people are trying too deep? Too wide? Not sufficiently regularized? Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 189556,
          "author_name": "svesal",
          "author_url": "",
          "post_date": "06/06/2017 10:40:08",
          "content": "<p>Hi gKericks, </p>\n\n<p>I have already tried with VGG16-like that isn't fine-tuned (pre-trained) but it worked really worse, then I started to do real time augmentation using Keras which helped a bit. Due to this, I just shift back to my lovely InceptionV3 pre-trained network which worked better for me. In your case, localization of cervical is an important part, as far as I know those who are in the top of LB used localization as well as getting advantage of data leakage.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 189560,
          "author_name": "oleggrinch",
          "author_url": "",
          "post_date": "06/06/2017 10:45:38",
          "content": "<p>You should use an ensemble of a few nets, since single net produces pretty confident (to itself) results. And in case some image was mispredicted, your log loss will blow up. Ensembling softens results and deals with outliers. In case you've got only 1 trained net, try to soften the predictions to avoid very low scores for certain classes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "189151": "Hi Kagglers,\n\nI am a bit confused about the approach that I am following. I have started to use several approaches for this task such as, Inception V3 pre-trained, Small CNN, VGG16 and etc but still could not get any good results. At first, I trained the data-set without doing GMM segmentation and I just used train and additional train images as validation and training sets.  Well, up to some extends, it could work but not satisfactory validation loss. \nIn the second step, I applied GMM segmentation, Data Augmentation, hyper-parameters tuning, removing corrupted images and feed the crop images to several networks but the maximum validation accuracy that I could get after 500 Epochs with IncpetionV3 fine tuned a validation accuracy of 64% only. \n\nDo you have any recommendation ?",
    "189227": "We're all in the same boat. People who have the highest accuracy on the LB say they feel their score is due to data leakage, but that the best _accuracy_ we should probably expect to get is 60-80%. If you are +70 without data leakage, then you're likely 'in the money'. Wish I could offer more advice, but... just keep tinkering and best of luck(?) !",
    "189234": "Hi authman,\n\nIs there another thread with more details on this the I might have missed? Especially the \"People who have the highest accuracy on the LB say they feel their score is due to data leakage\" part. I am aware that the 0 loss entries are taking advantage of the leak and that many top entries might have misleading results but I had also assumed that some of the high entries must be legitimate.  I'm also curious what you feel the +70 accuracy is in terms of log-loss. Thanks!",
    "189235": "Hi Aramis,\n\nI am also having trouble training a network that isn't fine-tuned (pre-trained). I can get a VGG-like network to train but not with satisfactory results. I'm curious if you or anyone else has any insights they are willing to share about what the problematic part of this dataset is. Are the networks people are trying too deep? Too wide? Not sufficiently regularized? Thanks.",
    "189556": "Hi gKericks, \n\nI have already tried with VGG16-like that isn't fine-tuned (pre-trained) but it worked really worse, then I started to do real time augmentation using Keras which helped a bit. Due to this, I just shift back to my lovely InceptionV3 pre-trained network which worked better for me. In your case, localization of cervical is an important part, as far as I know those who are in the top of LB used localization as well as getting advantage of data leakage.",
    "189560": "You should use an ensemble of a few nets, since single net produces pretty confident (to itself) results. And in case some image was mispredicted, your log loss will blow up. Ensembling softens results and deals with outliers. In case you've got only 1 trained net, try to soften the predictions to avoid very low scores for certain classes."
  },
  "source": "meta"
}