{
  "id": 13248,
  "title": "Some results",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/13248",
  "author_name": "",
  "post_date": "2015-04-06T04:21:10.687Z",
  "votes": 7,
  "comment_count": 16,
  "views": 6322,
  "content": "<p>Hi guys - I thought I'd share some results.</p>\n<p>Confusion matrix, rows = human label, columns = predicted.</p>\n<p><code>[[ 1175 90 23 &nbsp;3 &nbsp;0 ]<br>&nbsp;[ &nbsp; 72 37 12 &nbsp;1 &nbsp;0 ]<br>&nbsp;[ &nbsp; 55 53 98 58 &nbsp;1 ]<br>&nbsp;[ &nbsp; &nbsp;0 &nbsp;1 &nbsp;4 38 &nbsp;1 ]<br>&nbsp;[ &nbsp; &nbsp;0 &nbsp;3 &nbsp;2 28 &nbsp;3 ]]</code></p>\n<p><br>Kappa ~0.781</p>\n<p>Binary classification, none or mild (0-1) vs moderate and above (2+):</p>\n<p>ROC AUC ~0.934<br>73% TPR at 5% FPR<br>84% TPR at 10% FPR</p>\n<p>ROC curve attached.</p>\n<p>I'm wondering if anyone has a model&nbsp;that would be good to combine with this?</p>\n<p>Regards,</p>\n<p>Alex</p>",
  "messages": [
    {
      "id": "69765",
      "postDate": "04/06/2015 04:21:10",
      "content": "<p>Hi guys - I thought I'd share some results.</p>\n<p>Confusion matrix, rows = human label, columns = predicted.</p>\n<p><code>[[ 1175 90 23 &nbsp;3 &nbsp;0 ]<br>&nbsp;[ &nbsp; 72 37 12 &nbsp;1 &nbsp;0 ]<br>&nbsp;[ &nbsp; 55 53 98 58 &nbsp;1 ]<br>&nbsp;[ &nbsp; &nbsp;0 &nbsp;1 &nbsp;4 38 &nbsp;1 ]<br>&nbsp;[ &nbsp; &nbsp;0 &nbsp;3 &nbsp;2 28 &nbsp;3 ]]</code></p>\n<p><br>Kappa ~0.781</p>\n<p>Binary classification, none or mild (0-1) vs moderate and above (2+):</p>\n<p>ROC AUC ~0.934<br>73% TPR at 5% FPR<br>84% TPR at 10% FPR</p>\n<p>ROC curve attached.</p>\n<p>I'm wondering if anyone has a model&nbsp;that would be good to combine with this?</p>\n<p>Regards,</p>\n<p>Alex</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69811",
      "postDate": "04/06/2015 16:56:31",
      "content": "<p>Dude, there's still 3 months to go and you are already 0.77. Take it easy :) jk! Awesome!</p>\n<p>What's the mystery&nbsp;though, seriously? Are you just doing&nbsp;a combination of binary classifications or ensemble them with multiclass?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69926",
      "postDate": "04/07/2015 19:06:37",
      "content": "<p>Alex,</p>\n<p>Congrats for these results,&nbsp;</p>\n<p>I'm surprised to know how you would expect one's model to be good to combine with yours?</p>\n<p>Maybe any model with a score of 0.50 would add you around 0.01 benefit, but it's difficult to predict the benefit with just some ROC curves, isn't it?</p>\n<p>I'd be pleased to combine my model with yours but I'm not sure it will be of any benefit for you :-)&nbsp;</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69935",
      "postDate": "04/07/2015 20:33:12",
      "content": "<p>@Deep: Thank you! &nbsp;I can't say much about what I'm doing yet, but&nbsp;I hope the way it works will be an interesting surprise :)</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69937",
      "postDate": "04/07/2015 20:54:57",
      "content": "<p>Hi Zidmie,</p>\n<p>I think it is possible: for example if there was someone who was really great at finding 4s, that could significantly improve my model. &nbsp;They would not necessarily have a high&nbsp;score: correct prediction of all 4s gives you 0.3-something. &nbsp;Anyway, I thought I'd offer. &nbsp;</p>\n<p>I admit I'm also just curious, esp whether&nbsp;anyone has a much better ROC curve for very low false positives area.</p>\n<p>Best,</p>\n<p>Alex</p>\n<p>[quote=Zidmie;69926]</p>\n<p>Alex,</p>\n<p>Congrats for these results,&nbsp;</p>\n<p>I'm surprised to know how you would expect one's model to be good to combine with yours?</p>\n<p>Maybe any model with a score of 0.50 would add you around 0.01 benefit, but it's difficult to predict the benefit with just some ROC curves, isn't it?</p>\n<p>I'd be pleased to combine my model with yours but I'm not sure it will be of any benefit for you :-)&nbsp;</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69944",
      "postDate": "04/07/2015 21:12:17",
      "content": "<p>[quote=Alexander Izvorski;69935]</p>\n<p>@Deep: Thank you! &nbsp;I can't say much about what I'm doing yet, but&nbsp;I hope the way it works will be an interesting surprise :)</p>\n<p>[/quote]</p>\n<p>Congratulations on your results! :)</p>\n<p>I am curious: do you combine different methods (neural networks, classic detection algorithms etc.) or do you have some other kind of magic way?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69946",
      "postDate": "04/07/2015 21:23:43",
      "content": "<p>I understand you may not feel comfortable to share your methods (even at very high level discussion) at this point.</p>\n<p>So, just curious, what is the nature of this collaboration say if someone has a model that is really good at predicting level 4 vs.&nbsp;level 3 really well (you seem to struggle with it :) )? Are you looking for team merger?</p>\n<p>Also, is the confusion matrix&nbsp;for a stratified sample?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "69949",
      "postDate": "04/07/2015 22:33:49",
      "content": "<p>Hi&nbsp;Deep:</p>\n<p>I wouldn't mind a team merger if that's what it takes. &nbsp;This is a really hard problem, model diversity is always a good thing. &nbsp;This might also spark a bit of discussion about which parts of it are hard and why.</p>\n<p>Yes, it's a stratified sample.</p>\n<p>-Alex</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "71262",
      "postDate": "04/11/2015 01:56:14",
      "content": "<p>Hi Zoltan - Thanks! &nbsp;I appreciate your&nbsp;curiosity! &nbsp;It's a hybrid approach. &nbsp;I guess I'd better keep the details &quot;top secret&quot; for now :)</p>\n<p>-Alex</p>\n\n<p>[quote=Zoltan Fegyver;69944]</p>\n<p>Congratulations on your results! :)</p>\n<p>I am curious: do you combine different methods (neural networks, classic detection algorithms etc.) or do you have some other kind of magic way?</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "72671",
      "postDate": "04/20/2015 17:03:41",
      "content": "<p>Alex,</p>\n<p>Thanks for posting this; it's interesting to see how another approach is doing. Just to reciprocate, here's my confusion matrix, good for a kappa of ~0.63</p>\n<p><code>[[1099 161 27 &nbsp;2 &nbsp;2]<br>&nbsp;[ 100 &nbsp;25 &nbsp;3 &nbsp;0 &nbsp;0]<br>&nbsp;[ &nbsp;93 &nbsp;89 51 27 &nbsp;0]<br>&nbsp;[ &nbsp; 1 &nbsp;11 17 20 &nbsp;0]<br>&nbsp;[ &nbsp; 0 &nbsp; 2 11 &nbsp;3 18]]</code></p>\n<p>Looks like I'm doing better with class 4, and considerably worse everywhere else. I'll be very curious to see what kind of secret sauce you are using when this is over.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "73010",
      "postDate": "04/21/2015 19:30:43",
      "content": "<p>Hi Tim,</p>\n<p>Thanks! &nbsp;That's really solid.</p>\n<p>It looks like the &quot;confusability&quot; of things in your model is about the same as in mine, qualitatively speaking. &nbsp;For example, there just doesn't seem to be a great way to distinguish&nbsp;1s from 0s, and real 2s get spread somewhat evenly between predicted 0-3s. &nbsp;That's what I was curious about.</p>\n<p>I wonder, how much of this is limitations in our models, and how much is inherent &quot;softness&quot; of the category boundaries? &nbsp;I mean in theory just <em>one</em> microaneurysm (MA) is enough to make a 0 into a 1, and I&nbsp;think&nbsp;humans are not much better than (say) 50% recall/80% precision on identifying those.</p>\n<p>Likewise :)</p>\n<p>Best,</p>\n<p>Alex</p>\n\n<p>[quote=Tim Hochberg;72671]</p>\n<p>Alex,</p>\n<p>Thanks for posting this; it's interesting to see how another approach is doing. Just to reciprocate, here's my confusion matrix, good for a kappa of ~0.63</p>\n<p>[[1099 161 27 2 2]<br> [ 100 25 3 0 0]<br> [ 93 89 51 27 0]<br> [ 1 11 17 20 0]<br> [ 0 2 11 3 18]]</p>\n<p>Looks like I'm doing better with class 4, and considerably worse everywhere else. I'll be very curious to see what kind of secret sauce you are using when this is over.</p>\n<p>[/quote]</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "73012",
      "postDate": "04/21/2015 19:41:52",
      "content": "<p>Hi Alex,</p>\n<p>Didn't you use test set for improving results?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74308",
      "postDate": "04/23/2015 15:34:13",
      "content": "<p>Hi Alex,</p>\n<p>I agree that the softness of the categories is going to be a limiting factor. Presumably the progression of the disease is more or less continuous. (In fact, different facets of the disease might proceed at different rates in different patients, but that's a whole 'nother level of complication) and mapping that into discrete bins is bound to result in problems near the bin borders. Add to that these are rated by different doctors using different equipment and there's bound to be a limit to how close of an agreement can be achieved.&nbsp;</p>\n<p>It doesn't add to my confidence any that a significant fraction of the images are either extremely over or under exposed and seem unlikely to actually have been used in the ratings of the patients!&nbsp;Still, it may be possible to reach a useful level of accuracy even&nbsp;given these limitations.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74865",
      "postDate": "04/26/2015 15:03:07",
      "content": "<p>Hi, I am a new member. I am just wondering about these results. Are these results obtained by applying your method on the test set or are these obtained from for e.g. cross validation on the training set?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "74875",
      "postDate": "04/26/2015 16:15:21",
      "content": "<p>I can't speak for Alex, but I'm reserving 5% of the data for cross validation. I compute the predicted values using the CV data and compute kappa and the confusion matrix based on the&nbsp;predictions and the provided labels.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "77402",
      "postDate": "05/06/2015 21:09:48",
      "content": "<p>How do you determine your ROC curve and other info when Kaggle doesn't give you any statistical information about the test set other than the kappa?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "87314",
      "postDate": "07/28/2015 15:54:49",
      "content": "<p>[quote=Alexander Izvorski;71262]</p>\n<p>Hi Zoltan - Thanks! &nbsp;I appreciate your&nbsp;curiosity! &nbsp;It's a hybrid approach. &nbsp;I guess I'd better keep the details &quot;top secret&quot; for now :)</p>\n<p>-Alex</p>\n<p>[/quote]</p>\n<p>Hi Alex.</p>\n<p>I am very interesting (hope not only me) , how you got the 19th final result 3 month ago. I am feel, you used rather simple and shallow solution (comparing with large deep convolution net)</p>\n<p>It's seems this is not &quot;top secret&quot; for now and you can provide us some more details about your solution, at last :-)</p>\n<p>Thanks</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 69811,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "04/06/2015 16:56:31",
      "content": "<p>Dude, there's still 3 months to go and you are already 0.77. Take it easy :) jk! Awesome!</p>\n<p>What's the mystery&nbsp;though, seriously? Are you just doing&nbsp;a combination of binary classifications or ensemble them with multiclass?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69926,
      "author_name": "zidmie",
      "author_url": "",
      "post_date": "04/07/2015 19:06:37",
      "content": "<p>Alex,</p>\n<p>Congrats for these results,&nbsp;</p>\n<p>I'm surprised to know how you would expect one's model to be good to combine with yours?</p>\n<p>Maybe any model with a score of 0.50 would add you around 0.01 benefit, but it's difficult to predict the benefit with just some ROC curves, isn't it?</p>\n<p>I'd be pleased to combine my model with yours but I'm not sure it will be of any benefit for you :-)&nbsp;</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69935,
      "author_name": "aizvorski",
      "author_url": "",
      "post_date": "04/07/2015 20:33:12",
      "content": "<p>@Deep: Thank you! &nbsp;I can't say much about what I'm doing yet, but&nbsp;I hope the way it works will be an interesting surprise :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69937,
      "author_name": "aizvorski",
      "author_url": "",
      "post_date": "04/07/2015 20:54:57",
      "content": "<p>Hi Zidmie,</p>\n<p>I think it is possible: for example if there was someone who was really great at finding 4s, that could significantly improve my model. &nbsp;They would not necessarily have a high&nbsp;score: correct prediction of all 4s gives you 0.3-something. &nbsp;Anyway, I thought I'd offer. &nbsp;</p>\n<p>I admit I'm also just curious, esp whether&nbsp;anyone has a much better ROC curve for very low false positives area.</p>\n<p>Best,</p>\n<p>Alex</p>\n<p>[quote=Zidmie;69926]</p>\n<p>Alex,</p>\n<p>Congrats for these results,&nbsp;</p>\n<p>I'm surprised to know how you would expect one's model to be good to combine with yours?</p>\n<p>Maybe any model with a score of 0.50 would add you around 0.01 benefit, but it's difficult to predict the benefit with just some ROC curves, isn't it?</p>\n<p>I'd be pleased to combine my model with yours but I'm not sure it will be of any benefit for you :-)&nbsp;</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69944,
      "author_name": "zoltanfegyver",
      "author_url": "",
      "post_date": "04/07/2015 21:12:17",
      "content": "<p>[quote=Alexander Izvorski;69935]</p>\n<p>@Deep: Thank you! &nbsp;I can't say much about what I'm doing yet, but&nbsp;I hope the way it works will be an interesting surprise :)</p>\n<p>[/quote]</p>\n<p>Congratulations on your results! :)</p>\n<p>I am curious: do you combine different methods (neural networks, classic detection algorithms etc.) or do you have some other kind of magic way?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69946,
      "author_name": "deepcnn",
      "author_url": "",
      "post_date": "04/07/2015 21:23:43",
      "content": "<p>I understand you may not feel comfortable to share your methods (even at very high level discussion) at this point.</p>\n<p>So, just curious, what is the nature of this collaboration say if someone has a model that is really good at predicting level 4 vs.&nbsp;level 3 really well (you seem to struggle with it :) )? Are you looking for team merger?</p>\n<p>Also, is the confusion matrix&nbsp;for a stratified sample?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 69949,
      "author_name": "aizvorski",
      "author_url": "",
      "post_date": "04/07/2015 22:33:49",
      "content": "<p>Hi&nbsp;Deep:</p>\n<p>I wouldn't mind a team merger if that's what it takes. &nbsp;This is a really hard problem, model diversity is always a good thing. &nbsp;This might also spark a bit of discussion about which parts of it are hard and why.</p>\n<p>Yes, it's a stratified sample.</p>\n<p>-Alex</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 71262,
      "author_name": "aizvorski",
      "author_url": "",
      "post_date": "04/11/2015 01:56:14",
      "content": "<p>Hi Zoltan - Thanks! &nbsp;I appreciate your&nbsp;curiosity! &nbsp;It's a hybrid approach. &nbsp;I guess I'd better keep the details &quot;top secret&quot; for now :)</p>\n<p>-Alex</p>\n\n<p>[quote=Zoltan Fegyver;69944]</p>\n<p>Congratulations on your results! :)</p>\n<p>I am curious: do you combine different methods (neural networks, classic detection algorithms etc.) or do you have some other kind of magic way?</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 72671,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "04/20/2015 17:03:41",
      "content": "<p>Alex,</p>\n<p>Thanks for posting this; it's interesting to see how another approach is doing. Just to reciprocate, here's my confusion matrix, good for a kappa of ~0.63</p>\n<p><code>[[1099 161 27 &nbsp;2 &nbsp;2]<br>&nbsp;[ 100 &nbsp;25 &nbsp;3 &nbsp;0 &nbsp;0]<br>&nbsp;[ &nbsp;93 &nbsp;89 51 27 &nbsp;0]<br>&nbsp;[ &nbsp; 1 &nbsp;11 17 20 &nbsp;0]<br>&nbsp;[ &nbsp; 0 &nbsp; 2 11 &nbsp;3 18]]</code></p>\n<p>Looks like I'm doing better with class 4, and considerably worse everywhere else. I'll be very curious to see what kind of secret sauce you are using when this is over.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 73010,
      "author_name": "aizvorski",
      "author_url": "",
      "post_date": "04/21/2015 19:30:43",
      "content": "<p>Hi Tim,</p>\n<p>Thanks! &nbsp;That's really solid.</p>\n<p>It looks like the &quot;confusability&quot; of things in your model is about the same as in mine, qualitatively speaking. &nbsp;For example, there just doesn't seem to be a great way to distinguish&nbsp;1s from 0s, and real 2s get spread somewhat evenly between predicted 0-3s. &nbsp;That's what I was curious about.</p>\n<p>I wonder, how much of this is limitations in our models, and how much is inherent &quot;softness&quot; of the category boundaries? &nbsp;I mean in theory just <em>one</em> microaneurysm (MA) is enough to make a 0 into a 1, and I&nbsp;think&nbsp;humans are not much better than (say) 50% recall/80% precision on identifying those.</p>\n<p>Likewise :)</p>\n<p>Best,</p>\n<p>Alex</p>\n\n<p>[quote=Tim Hochberg;72671]</p>\n<p>Alex,</p>\n<p>Thanks for posting this; it's interesting to see how another approach is doing. Just to reciprocate, here's my confusion matrix, good for a kappa of ~0.63</p>\n<p>[[1099 161 27 2 2]<br> [ 100 25 3 0 0]<br> [ 93 89 51 27 0]<br> [ 1 11 17 20 0]<br> [ 0 2 11 3 18]]</p>\n<p>Looks like I'm doing better with class 4, and considerably worse everywhere else. I'll be very curious to see what kind of secret sauce you are using when this is over.</p>\n<p>[/quote]</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 73012,
      "author_name": "",
      "author_url": "",
      "post_date": "04/21/2015 19:41:52",
      "content": "<p>Hi Alex,</p>\n<p>Didn't you use test set for improving results?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74308,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "04/23/2015 15:34:13",
      "content": "<p>Hi Alex,</p>\n<p>I agree that the softness of the categories is going to be a limiting factor. Presumably the progression of the disease is more or less continuous. (In fact, different facets of the disease might proceed at different rates in different patients, but that's a whole 'nother level of complication) and mapping that into discrete bins is bound to result in problems near the bin borders. Add to that these are rated by different doctors using different equipment and there's bound to be a limit to how close of an agreement can be achieved.&nbsp;</p>\n<p>It doesn't add to my confidence any that a significant fraction of the images are either extremely over or under exposed and seem unlikely to actually have been used in the ratings of the patients!&nbsp;Still, it may be possible to reach a useful level of accuracy even&nbsp;given these limitations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74865,
      "author_name": "thewinner",
      "author_url": "",
      "post_date": "04/26/2015 15:03:07",
      "content": "<p>Hi, I am a new member. I am just wondering about these results. Are these results obtained by applying your method on the test set or are these obtained from for e.g. cross validation on the training set?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 74875,
      "author_name": "bitsofbits",
      "author_url": "",
      "post_date": "04/26/2015 16:15:21",
      "content": "<p>I can't speak for Alex, but I'm reserving 5% of the data for cross validation. I compute the predicted values using the CV data and compute kappa and the confusion matrix based on the&nbsp;predictions and the provided labels.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 77402,
      "author_name": "dannn314472",
      "author_url": "",
      "post_date": "05/06/2015 21:09:48",
      "content": "<p>How do you determine your ROC curve and other info when Kaggle doesn't give you any statistical information about the test set other than the kappa?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 87314,
      "author_name": "goodok",
      "author_url": "",
      "post_date": "07/28/2015 15:54:49",
      "content": "<p>[quote=Alexander Izvorski;71262]</p>\n<p>Hi Zoltan - Thanks! &nbsp;I appreciate your&nbsp;curiosity! &nbsp;It's a hybrid approach. &nbsp;I guess I'd better keep the details &quot;top secret&quot; for now :)</p>\n<p>-Alex</p>\n<p>[/quote]</p>\n<p>Hi Alex.</p>\n<p>I am very interesting (hope not only me) , how you got the 19th final result 3 month ago. I am feel, you used rather simple and shallow solution (comparing with large deep convolution net)</p>\n<p>It's seems this is not &quot;top secret&quot; for now and you can provide us some more details about your solution, at last :-)</p>\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "69765": "",
    "69811": "",
    "69926": "",
    "69935": "",
    "69937": "",
    "69944": "",
    "69946": "",
    "69949": "",
    "71262": "",
    "72671": "",
    "73010": "",
    "73012": "",
    "74308": "",
    "74865": "",
    "74875": "",
    "77402": "",
    "87314": ""
  },
  "source": "meta"
}