{
  "id": 12702,
  "title": "Local evaluation method?",
  "url": "/competitions/diabetic-retinopathy-detection/discussion/12702",
  "author_name": "",
  "post_date": "2015-03-06T00:42:09.340Z",
  "votes": 2,
  "comment_count": 7,
  "views": 2971,
  "content": "<p>Hi, I'm wondering what local evaluation method or metric is suitable for approximating&nbsp;weighted kappa. I'm using mean square error. however it is not consistent with LB.</p>\n<p>for example, cv error&nbsp;(lower the better) 1.82, LB kappa 0.091; cv error&nbsp;1.77, LB kappa 0.076</p>\n<p>Any&nbsp;comments are appreciated.</p>",
  "messages": [
    {
      "id": "65555",
      "postDate": "03/06/2015 00:42:09",
      "content": "<p>Hi, I'm wondering what local evaluation method or metric is suitable for approximating&nbsp;weighted kappa. I'm using mean square error. however it is not consistent with LB.</p>\n<p>for example, cv error&nbsp;(lower the better) 1.82, LB kappa 0.091; cv error&nbsp;1.77, LB kappa 0.076</p>\n<p>Any&nbsp;comments are appreciated.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65579",
      "postDate": "03/06/2015 09:01:36",
      "content": "<p>Why don't you use ... weighted kappa?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65612",
      "postDate": "03/06/2015 16:39:03",
      "content": "<p>[quote=Tsakalis Kostas;65579]</p>\n<p>Why don't you use ... weighted kappa?</p>\n<p>[/quote]</p>\n<p>Thank you. I foolishly thought it depends on some unknown cost. Never mind.</p>\n<p>I found this <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/quadratic_weighted_kappa.py\">script</a>, is it the same with what we use here?</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65629",
      "postDate": "03/06/2015 19:57:22",
      "content": "<p>Weighted Kappa would add a little more accuracy, however the prevalence of greater than mild DR is about 40%, and greater than moderate is about 25% so the difference should not big.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65644",
      "postDate": "03/06/2015 22:31:45",
      "content": "<p><strong>Jorge: In the training set, the number of cases greater than mild is 17%, and greater than moderate is 4%. You estimate these proportions are about 40% and 25% respectively. Given the random nature of the file naming IDs, I would expect the training and test datasets to have very similar proportions - a true random test sample suitable for evaluation.</strong></p>\n<ul>\n<li><strong>Is the test set a truly (pseudo)random selection of the combined test+train dataset?<br></strong></li>\n<li><strong>Is the combined test+train dataset a subset of the global population?</strong></li>\n<li><strong>Why are your stated proportions so different from the observed values in the training set?</strong></li>\n</ul>\n<p><strong>Thanks, Alastair</strong></p>\n<p>Cases&nbsp;&nbsp;&nbsp;&nbsp; Level&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Proportion</p>\n<p><code>25810&nbsp; 0 - No DR&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 73%<br>&nbsp;2443&nbsp; 1 - Mild&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 7%<br>&nbsp;5292&nbsp; 2 - Moderate&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 15%<br>&nbsp; 873&nbsp; 3 - Severe&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 2%<br>&nbsp; 708&nbsp; 4 - Proliferative DR&nbsp;&nbsp; 2%<br>35126&nbsp; Total<br></code></p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "65647",
      "postDate": "03/06/2015 23:00:57",
      "content": "<p>Alastair:</p>\n<p>I was speaking generally and I'm not able to comment on the competition dataset.</p>\n<p>-Jorge</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "80141",
      "postDate": "05/27/2015 14:32:42",
      "content": "<p>Should I use like this:&nbsp;quadratic_weighted_kappa(y_true, y_pred) ?</p>\n\n<p>Regards,</p>\n<p>Emilio</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "80212",
      "postDate": "05/28/2015 03:22:59",
      "content": "<p>@Emilio Yes. That script has worked for me. I never bothered checking if it's 100%&nbsp;correct though ^^</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 65579,
      "author_name": "epinephelus",
      "author_url": "",
      "post_date": "03/06/2015 09:01:36",
      "content": "<p>Why don't you use ... weighted kappa?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65612,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "03/06/2015 16:39:03",
      "content": "<p>[quote=Tsakalis Kostas;65579]</p>\n<p>Why don't you use ... weighted kappa?</p>\n<p>[/quote]</p>\n<p>Thank you. I foolishly thought it depends on some unknown cost. Never mind.</p>\n<p>I found this <a href=\"https://github.com/benhamner/Metrics/blob/master/Python/ml_metrics/quadratic_weighted_kappa.py\">script</a>, is it the same with what we use here?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65629,
      "author_name": "jorge9",
      "author_url": "",
      "post_date": "03/06/2015 19:57:22",
      "content": "<p>Weighted Kappa would add a little more accuracy, however the prevalence of greater than mild DR is about 40%, and greater than moderate is about 25% so the difference should not big.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65644,
      "author_name": "alastairk",
      "author_url": "",
      "post_date": "03/06/2015 22:31:45",
      "content": "<p><strong>Jorge: In the training set, the number of cases greater than mild is 17%, and greater than moderate is 4%. You estimate these proportions are about 40% and 25% respectively. Given the random nature of the file naming IDs, I would expect the training and test datasets to have very similar proportions - a true random test sample suitable for evaluation.</strong></p>\n<ul>\n<li><strong>Is the test set a truly (pseudo)random selection of the combined test+train dataset?<br></strong></li>\n<li><strong>Is the combined test+train dataset a subset of the global population?</strong></li>\n<li><strong>Why are your stated proportions so different from the observed values in the training set?</strong></li>\n</ul>\n<p><strong>Thanks, Alastair</strong></p>\n<p>Cases&nbsp;&nbsp;&nbsp;&nbsp; Level&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; Proportion</p>\n<p><code>25810&nbsp; 0 - No DR&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 73%<br>&nbsp;2443&nbsp; 1 - Mild&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 7%<br>&nbsp;5292&nbsp; 2 - Moderate&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 15%<br>&nbsp; 873&nbsp; 3 - Severe&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 2%<br>&nbsp; 708&nbsp; 4 - Proliferative DR&nbsp;&nbsp; 2%<br>35126&nbsp; Total<br></code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 65647,
      "author_name": "jorge9",
      "author_url": "",
      "post_date": "03/06/2015 23:00:57",
      "content": "<p>Alastair:</p>\n<p>I was speaking generally and I'm not able to comment on the competition dataset.</p>\n<p>-Jorge</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 80141,
      "author_name": "emiliocs",
      "author_url": "",
      "post_date": "05/27/2015 14:32:42",
      "content": "<p>Should I use like this:&nbsp;quadratic_weighted_kappa(y_true, y_pred) ?</p>\n\n<p>Regards,</p>\n<p>Emilio</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 80212,
      "author_name": "sveitser",
      "author_url": "",
      "post_date": "05/28/2015 03:22:59",
      "content": "<p>@Emilio Yes. That script has worked for me. I never bothered checking if it's 100%&nbsp;correct though ^^</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "65555": "",
    "65579": "",
    "65612": "",
    "65629": "",
    "65644": "",
    "65647": "",
    "80141": "",
    "80212": ""
  },
  "source": "meta"
}