{
  "id": 15704,
  "title": "Anyone using unsupervised techniques?",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/15704",
  "author_name": "",
  "post_date": "2015-08-01T22:29:48.137Z",
  "votes": 1,
  "comment_count": 4,
  "views": 1529,
  "content": "<p>So far 3 solutions have been posted on this forum and all of them are based on OxfordNet.</p>\n\n<p>Did anyone try using autoencoders, RBMs or other unsupervised techniques? </p>\n\n<p>We had an idea that it could be helpful to recognize the healthy parts of an eye, then classify by looking at the remaining parts, because in almost every case some parts of the images have no signs of illness. We thought we could recognize &quot;healthy&quot; parts of the eyes by training an autoencoder or RBM over the small tiles of 0-labelled images. If this works, we might be able to filter out healthy parts of the other images and get a set of small tiles corresponding to every level of the illness, and even train RBMs on each of these sets. During testing we could check if there are, for example, tiles that are best recognized by the level-4-RBM, and classify the image as 4...</p>\n\n<p>We couldn't check if the idea works because of technical issues [dealing with millions of tiles was really tough] and didn't go beyond ConvNets during the contest. [We finished at 82nd place, will post details about the mistakes we made in a few days]</p>",
  "messages": [
    {
      "id": "87906",
      "postDate": "08/01/2015 22:29:48",
      "content": "<p>So far 3 solutions have been posted on this forum and all of them are based on OxfordNet.</p>\n\n<p>Did anyone try using autoencoders, RBMs or other unsupervised techniques? </p>\n\n<p>We had an idea that it could be helpful to recognize the healthy parts of an eye, then classify by looking at the remaining parts, because in almost every case some parts of the images have no signs of illness. We thought we could recognize &quot;healthy&quot; parts of the eyes by training an autoencoder or RBM over the small tiles of 0-labelled images. If this works, we might be able to filter out healthy parts of the other images and get a set of small tiles corresponding to every level of the illness, and even train RBMs on each of these sets. During testing we could check if there are, for example, tiles that are best recognized by the level-4-RBM, and classify the image as 4...</p>\n\n<p>We couldn't check if the idea works because of technical issues [dealing with millions of tiles was really tough] and didn't go beyond ConvNets during the contest. [We finished at 82nd place, will post details about the mistakes we made in a few days]</p>",
      "rawMarkdown": "So far 3 solutions have been posted on this forum and all of them are based on OxfordNet.\r\n\r\nDid anyone try using autoencoders, RBMs or other unsupervised techniques? \r\n\r\nWe had an idea that it could be helpful to recognize the healthy parts of an eye, then classify by looking at the remaining parts, because in almost every case some parts of the images have no signs of illness. We thought we could recognize \"healthy\" parts of the eyes by training an autoencoder or RBM over the small tiles of 0-labelled images. If this works, we might be able to filter out healthy parts of the other images and get a set of small tiles corresponding to every level of the illness, and even train RBMs on each of these sets. During testing we could check if there are, for example, tiles that are best recognized by the level-4-RBM, and classify the image as 4...\r\n\r\nWe couldn't check if the idea works because of technical issues [dealing with millions of tiles was really tough] and didn't go beyond ConvNets during the contest. [We finished at 82nd place, will post details about the mistakes we made in a few days]",
      "votes": null
    },
    {
      "id": "87920",
      "postDate": "08/02/2015 01:37:24",
      "content": "<p>I implemented the technique in this paper: <a href=\"http://arxiv.org/abs/1406.6909\">http://arxiv.org/abs/1406.6909</a></p>\n\n<p>The performance was very disappointing. My guess is that the features necessary to discriminate arbitrary images aren't generally useful for distinguishing retinopathy levels. </p>",
      "rawMarkdown": "I implemented the technique in this paper: http://arxiv.org/abs/1406.6909\r\n\r\nThe performance was very disappointing. My guess is that the features necessary to discriminate arbitrary images aren't generally useful for distinguishing retinopathy levels.",
      "votes": null
    },
    {
      "id": "87946",
      "postDate": "08/02/2015 10:56:44",
      "content": "<p>[quote=Dan Nuffer;87920]</p>\n\n<p>I implemented the technique in this paper: <a href=\"http://arxiv.org/abs/1406.6909\">http://arxiv.org/abs/1406.6909</a></p>\n\n<p>The performance was very disappointing. My guess is that the features necessary to discriminate arbitrary images aren't generally useful for distinguishing retinopathy levels. \n[/quote]</p>\n\n<p>Interesting paper!</p>\n\n<p>Although its lack of performance here shouldn't be that surprising - even the abstract itself says &quot;While such generic features cannot compete with class specific features from supervised training on a classification task ...&quot;</p>",
      "rawMarkdown": "[quote=Dan Nuffer;87920]\r\n\r\nI implemented the technique in this paper: http://arxiv.org/abs/1406.6909\r\n\r\nThe performance was very disappointing. My guess is that the features necessary to discriminate arbitrary images aren't generally useful for distinguishing retinopathy levels. \r\n[/quote]\r\n\r\nInteresting paper!\r\n\r\nAlthough its lack of performance here shouldn't be that surprising - even the abstract itself says \"While such generic features cannot compete with class specific features from supervised training on a classification task ...\"",
      "votes": null
    },
    {
      "id": "87955",
      "postDate": "08/02/2015 14:32:25",
      "content": "<p>It's interesting to see how the authors have updated that paper. That sentence and the experiments with ImageNet were added in the most recent version from about a month ago. If you look at v1 in the arxiv, it just says &quot;We find that this simple feature learning algorithm is surprisingly successful when applied to visual object recognition. The feature representation learned by our algorithm achieves classification results matching or outperforming the current state-of-the-art for unsupervised learning on several popular datasets&quot;.</p>",
      "rawMarkdown": "It's interesting to see how the authors have updated that paper. That sentence and the experiments with ImageNet were added in the most recent version from about a month ago. If you look at v1 in the arxiv, it just says \"We find that this simple feature learning algorithm is surprisingly successful when applied to visual object recognition. The feature representation learned by our algorithm achieves classification results matching or outperforming the current state-of-the-art for unsupervised learning on several popular datasets\".",
      "votes": null
    },
    {
      "id": "88855",
      "postDate": "08/07/2015 20:34:28",
      "content": "<p>I used some unsupervised techniques in the sense of K-means clustering and sparse coding.</p>\n\n<p>The final LB score that I got with that was 0.69230 and 35th place.\nI was using regression (with various classifiers) and used hyperopt to pick the tresholds between the classes.\nI also always averaged the regression scores of the left and right eye before looking for tresholds. This always gave for me better scores than looking at each eye individually.</p>\n\n<p>My approach was first to find vessels and inpaint them using openCV. Then I used mostly the green channel and divide by the median of a 20x20 or 40x40 or 80x80 square. Then I extracted darker and lighter patches and used some features like elongation, sharpness of the boundary, (corrected) hue at the center etc.</p>\n\n<p>To this set of features I used K-means clustering but used soft encoding with gaussian rbf instead of hard encoding. These were a positive set of features for each light/dark patch.\nThen for each image I made a histogram of the patches and used xgboost for training.</p>\n\n<p>What worked better was when I tried to associate the relevance of each patch by training xgboost on the individual patches with image target. Then when I weighted the patches in the histogram by this relevance (cubed) this gave significantly better results.</p>\n\n<p>Another improvement was to train a CNN on the 31x31 patches I identifed earlier (with image targets). This gave improved relevance weights for the individual dark and light patches within an image.</p>\n\n<p>The final try was to use the values of the last hidden layer as new features for the patches, used nonnegative sparse coding for them, made histograms and used xgboost.</p>\n\n<p>Putting these things together, averaging, including some other features of the relevance scores etc. gave my final best score.</p>\n\n<p>I feel that the approach was saturated at that score. Perhaps the bottleneck was the initial choice of the dark/light patches...\nIn the final month of the competition I was looking at multiple instance learning (and even tried one approach which unfortuantely did not work well) but did not go anywhere with that. Essentially the weighting of the histogram with the instance weights already was within the multiple instance paradigm. </p>",
      "rawMarkdown": "I used some unsupervised techniques in the sense of K-means clustering and sparse coding.\r\n\r\nThe final LB score that I got with that was 0.69230 and 35th place.\r\nI was using regression (with various classifiers) and used hyperopt to pick the tresholds between the classes.\r\nI also always averaged the regression scores of the left and right eye before looking for tresholds. This always gave for me better scores than looking at each eye individually.\r\n\r\nMy approach was first to find vessels and inpaint them using openCV. Then I used mostly the green channel and divide by the median of a 20x20 or 40x40 or 80x80 square. Then I extracted darker and lighter patches and used some features like elongation, sharpness of the boundary, (corrected) hue at the center etc.\r\n\r\nTo this set of features I used K-means clustering but used soft encoding with gaussian rbf instead of hard encoding. These were a positive set of features for each light/dark patch.\r\nThen for each image I made a histogram of the patches and used xgboost for training.\r\n\r\nWhat worked better was when I tried to associate the relevance of each patch by training xgboost on the individual patches with image target. Then when I weighted the patches in the histogram by this relevance (cubed) this gave significantly better results.\r\n\r\nAnother improvement was to train a CNN on the 31x31 patches I identifed earlier (with image targets). This gave improved relevance weights for the individual dark and light patches within an image.\r\n\r\nThe final try was to use the values of the last hidden layer as new features for the patches, used nonnegative sparse coding for them, made histograms and used xgboost.\r\n\r\nPutting these things together, averaging, including some other features of the relevance scores etc. gave my final best score.\r\n\r\nI feel that the approach was saturated at that score. Perhaps the bottleneck was the initial choice of the dark/light patches...\r\nIn the final month of the competition I was looking at multiple instance learning (and even tried one approach which unfortuantely did not work well) but did not go anywhere with that. Essentially the weighting of the histogram with the instance weights already was within the multiple instance paradigm.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 87920,
      "author_name": "dannuffer",
      "author_url": "",
      "post_date": "08/02/2015 01:37:24",
      "content": "<p>I implemented the technique in this paper: <a href=\"http://arxiv.org/abs/1406.6909\">http://arxiv.org/abs/1406.6909</a></p>\n\n<p>The performance was very disappointing. My guess is that the features necessary to discriminate arbitrary images aren't generally useful for distinguishing retinopathy levels. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87946,
      "author_name": "nicolaykorslund",
      "author_url": "",
      "post_date": "08/02/2015 10:56:44",
      "content": "<p>[quote=Dan Nuffer;87920]</p>\n\n<p>I implemented the technique in this paper: <a href=\"http://arxiv.org/abs/1406.6909\">http://arxiv.org/abs/1406.6909</a></p>\n\n<p>The performance was very disappointing. My guess is that the features necessary to discriminate arbitrary images aren't generally useful for distinguishing retinopathy levels. \n[/quote]</p>\n\n<p>Interesting paper!</p>\n\n<p>Although its lack of performance here shouldn't be that surprising - even the abstract itself says &quot;While such generic features cannot compete with class specific features from supervised training on a classification task ...&quot;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87955,
      "author_name": "dannuffer",
      "author_url": "",
      "post_date": "08/02/2015 14:32:25",
      "content": "<p>It's interesting to see how the authors have updated that paper. That sentence and the experiments with ImageNet were added in the most recent version from about a month ago. If you look at v1 in the arxiv, it just says &quot;We find that this simple feature learning algorithm is surprisingly successful when applied to visual object recognition. The feature representation learned by our algorithm achieves classification results matching or outperforming the current state-of-the-art for unsupervised learning on several popular datasets&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 88855,
      "author_name": "romualdj",
      "author_url": "",
      "post_date": "08/07/2015 20:34:28",
      "content": "<p>I used some unsupervised techniques in the sense of K-means clustering and sparse coding.</p>\n\n<p>The final LB score that I got with that was 0.69230 and 35th place.\nI was using regression (with various classifiers) and used hyperopt to pick the tresholds between the classes.\nI also always averaged the regression scores of the left and right eye before looking for tresholds. This always gave for me better scores than looking at each eye individually.</p>\n\n<p>My approach was first to find vessels and inpaint them using openCV. Then I used mostly the green channel and divide by the median of a 20x20 or 40x40 or 80x80 square. Then I extracted darker and lighter patches and used some features like elongation, sharpness of the boundary, (corrected) hue at the center etc.</p>\n\n<p>To this set of features I used K-means clustering but used soft encoding with gaussian rbf instead of hard encoding. These were a positive set of features for each light/dark patch.\nThen for each image I made a histogram of the patches and used xgboost for training.</p>\n\n<p>What worked better was when I tried to associate the relevance of each patch by training xgboost on the individual patches with image target. Then when I weighted the patches in the histogram by this relevance (cubed) this gave significantly better results.</p>\n\n<p>Another improvement was to train a CNN on the 31x31 patches I identifed earlier (with image targets). This gave improved relevance weights for the individual dark and light patches within an image.</p>\n\n<p>The final try was to use the values of the last hidden layer as new features for the patches, used nonnegative sparse coding for them, made histograms and used xgboost.</p>\n\n<p>Putting these things together, averaging, including some other features of the relevance scores etc. gave my final best score.</p>\n\n<p>I feel that the approach was saturated at that score. Perhaps the bottleneck was the initial choice of the dark/light patches...\nIn the final month of the competition I was looking at multiple instance learning (and even tried one approach which unfortuantely did not work well) but did not go anywhere with that. Essentially the weighting of the histogram with the instance weights already was within the multiple instance paradigm. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "87906": "So far 3 solutions have been posted on this forum and all of them are based on OxfordNet.\r\n\r\nDid anyone try using autoencoders, RBMs or other unsupervised techniques? \r\n\r\nWe had an idea that it could be helpful to recognize the healthy parts of an eye, then classify by looking at the remaining parts, because in almost every case some parts of the images have no signs of illness. We thought we could recognize \"healthy\" parts of the eyes by training an autoencoder or RBM over the small tiles of 0-labelled images. If this works, we might be able to filter out healthy parts of the other images and get a set of small tiles corresponding to every level of the illness, and even train RBMs on each of these sets. During testing we could check if there are, for example, tiles that are best recognized by the level-4-RBM, and classify the image as 4...\r\n\r\nWe couldn't check if the idea works because of technical issues [dealing with millions of tiles was really tough] and didn't go beyond ConvNets during the contest. [We finished at 82nd place, will post details about the mistakes we made in a few days]",
    "87920": "I implemented the technique in this paper: http://arxiv.org/abs/1406.6909\r\n\r\nThe performance was very disappointing. My guess is that the features necessary to discriminate arbitrary images aren't generally useful for distinguishing retinopathy levels.",
    "87946": "[quote=Dan Nuffer;87920]\r\n\r\nI implemented the technique in this paper: http://arxiv.org/abs/1406.6909\r\n\r\nThe performance was very disappointing. My guess is that the features necessary to discriminate arbitrary images aren't generally useful for distinguishing retinopathy levels. \r\n[/quote]\r\n\r\nInteresting paper!\r\n\r\nAlthough its lack of performance here shouldn't be that surprising - even the abstract itself says \"While such generic features cannot compete with class specific features from supervised training on a classification task ...\"",
    "87955": "It's interesting to see how the authors have updated that paper. That sentence and the experiments with ImageNet were added in the most recent version from about a month ago. If you look at v1 in the arxiv, it just says \"We find that this simple feature learning algorithm is surprisingly successful when applied to visual object recognition. The feature representation learned by our algorithm achieves classification results matching or outperforming the current state-of-the-art for unsupervised learning on several popular datasets\".",
    "88855": "I used some unsupervised techniques in the sense of K-means clustering and sparse coding.\r\n\r\nThe final LB score that I got with that was 0.69230 and 35th place.\r\nI was using regression (with various classifiers) and used hyperopt to pick the tresholds between the classes.\r\nI also always averaged the regression scores of the left and right eye before looking for tresholds. This always gave for me better scores than looking at each eye individually.\r\n\r\nMy approach was first to find vessels and inpaint them using openCV. Then I used mostly the green channel and divide by the median of a 20x20 or 40x40 or 80x80 square. Then I extracted darker and lighter patches and used some features like elongation, sharpness of the boundary, (corrected) hue at the center etc.\r\n\r\nTo this set of features I used K-means clustering but used soft encoding with gaussian rbf instead of hard encoding. These were a positive set of features for each light/dark patch.\r\nThen for each image I made a histogram of the patches and used xgboost for training.\r\n\r\nWhat worked better was when I tried to associate the relevance of each patch by training xgboost on the individual patches with image target. Then when I weighted the patches in the histogram by this relevance (cubed) this gave significantly better results.\r\n\r\nAnother improvement was to train a CNN on the 31x31 patches I identifed earlier (with image targets). This gave improved relevance weights for the individual dark and light patches within an image.\r\n\r\nThe final try was to use the values of the last hidden layer as new features for the patches, used nonnegative sparse coding for them, made histograms and used xgboost.\r\n\r\nPutting these things together, averaging, including some other features of the relevance scores etc. gave my final best score.\r\n\r\nI feel that the approach was saturated at that score. Perhaps the bottleneck was the initial choice of the dark/light patches...\r\nIn the final month of the competition I was looking at multiple instance learning (and even tried one approach which unfortuantely did not work well) but did not go anywhere with that. Essentially the weighting of the histogram with the instance weights already was within the multiple instance paradigm."
  },
  "source": "meta"
}