{
  "id": 32141,
  "title": "Pointers",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/32141",
  "author_name": "fergusoci",
  "post_date": "2017-04-26T20:12:14.724000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>If anyone else has been scratching their head for the past while as to why they are seeing funny behaviour, or why their models are not optimizing as they would like... I was in the same boat.</p>\n\n<p>If it is of interest, these are some of the items that have helped me:</p>\n\n<p>Cluster the images to identify groups of the same images. The clustering needs to be done over the training, additional and test sets.</p>\n\n<p>There are images in the test set that match those in the train/additional sets. The clustering should identify these and allow you to increase your LB score, if you cared about that at this first stage of the competition.</p>\n\n<p>There are further images in the train and additional sets that are duplicates. Some images in the train set match others in the train set, so even if you are not using the additional data it may help you to identify these. If you are using the train plus the additional set then you need to remove the duplications across both (i.e. don't just clean the additional set, as this will still leave you with duplicate images).</p>\n\n<p>I have only just started cleaning this data in earnest, so I'm not certain what the full impact of this will be. With a first run through it was enough to start making sense of models and see a LB increase of approx 0.1.</p>\n\n<p>Not ideal, but it feels like without this it would not be possible to produce very meaningful results outside this competition.</p>\n\n<p>DM me if anybody is interested in forming a team - looking forward to finally getting started on this...</p>",
  "messages": [
    {
      "id": 178155,
      "postDate": "2017-04-26T20:12:14.723Z",
      "content": "<p>Hi all,</p>\n\n<p>If anyone else has been scratching their head for the past while as to why they are seeing funny behaviour, or why their models are not optimizing as they would like... I was in the same boat.</p>\n\n<p>If it is of interest, these are some of the items that have helped me:</p>\n\n<p>Cluster the images to identify groups of the same images. The clustering needs to be done over the training, additional and test sets.</p>\n\n<p>There are images in the test set that match those in the train/additional sets. The clustering should identify these and allow you to increase your LB score, if you cared about that at this first stage of the competition.</p>\n\n<p>There are further images in the train and additional sets that are duplicates. Some images in the train set match others in the train set, so even if you are not using the additional data it may help you to identify these. If you are using the train plus the additional set then you need to remove the duplications across both (i.e. don't just clean the additional set, as this will still leave you with duplicate images).</p>\n\n<p>I have only just started cleaning this data in earnest, so I'm not certain what the full impact of this will be. With a first run through it was enough to start making sense of models and see a LB increase of approx 0.1.</p>\n\n<p>Not ideal, but it feels like without this it would not be possible to produce very meaningful results outside this competition.</p>\n\n<p>DM me if anybody is interested in forming a team - looking forward to finally getting started on this...</p>",
      "rawMarkdown": "Hi all,\n\nIf anyone else has been scratching their head for the past while as to why they are seeing funny behaviour, or why their models are not optimizing as they would like... I was in the same boat.\n\nIf it is of interest, these are some of the items that have helped me:\n\nCluster the images to identify groups of the same images. The clustering needs to be done over the training, additional and test sets.\n\nThere are images in the test set that match those in the train/additional sets. The clustering should identify these and allow you to increase your LB score, if you cared about that at this first stage of the competition.\n\nThere are further images in the train and additional sets that are duplicates. Some images in the train set match others in the train set, so even if you are not using the additional data it may help you to identify these. If you are using the train plus the additional set then you need to remove the duplications across both (i.e. don't just clean the additional set, as this will still leave you with duplicate images).\n\nI have only just started cleaning this data in earnest, so I'm not certain what the full impact of this will be. With a first run through it was enough to start making sense of models and see a LB increase of approx 0.1.\n\nNot ideal, but it feels like without this it would not be possible to produce very meaningful results outside this competition.\n\nDM me if anybody is interested in forming a team - looking forward to finally getting started on this...",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "178155": "Hi all,\n\nIf anyone else has been scratching their head for the past while as to why they are seeing funny behaviour, or why their models are not optimizing as they would like... I was in the same boat.\n\nIf it is of interest, these are some of the items that have helped me:\n\nCluster the images to identify groups of the same images. The clustering needs to be done over the training, additional and test sets.\n\nThere are images in the test set that match those in the train/additional sets. The clustering should identify these and allow you to increase your LB score, if you cared about that at this first stage of the competition.\n\nThere are further images in the train and additional sets that are duplicates. Some images in the train set match others in the train set, so even if you are not using the additional data it may help you to identify these. If you are using the train plus the additional set then you need to remove the duplications across both (i.e. don't just clean the additional set, as this will still leave you with duplicate images).\n\nI have only just started cleaning this data in earnest, so I'm not certain what the full impact of this will be. With a first run through it was enough to start making sense of models and see a LB increase of approx 0.1.\n\nNot ideal, but it feels like without this it would not be possible to produce very meaningful results outside this competition.\n\nDM me if anybody is interested in forming a team - looking forward to finally getting started on this..."
  }
}