{
  "id": 4059,
  "title": "Hints to improve the results using only the features provided.",
  "url": "/competitions/icdar2013-gender-prediction-from-handwriting/discussion/4059",
  "author_name": "",
  "post_date": "2013-03-20T11:32:41.670Z",
  "votes": null,
  "comment_count": 6,
  "views": 3896,
  "content": "<p>Hi,</p>\r\n<p>I am new to data mining.<br>\r\nI don't want to download the images, since my network is slow.<br>\r\nAs I see all 3 standard benchmarks which use all the features provided, stand around 0.645.<br>\r\nAre there ways to reduce this number without using extracting features from the images provided?<br>\r\nSome general pointers would be appreciated ...</p>\r\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "21436",
      "postDate": "03/20/2013 11:32:41",
      "content": "<p>Hi,</p>\r\n<p>I am new to data mining.<br>\r\nI don't want to download the images, since my network is slow.<br>\r\nAs I see all 3 standard benchmarks which use all the features provided, stand around 0.645.<br>\r\nAre there ways to reduce this number without using extracting features from the images provided?<br>\r\nSome general pointers would be appreciated ...</p>\r\n<p>Thanks.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "21495",
      "postDate": "03/21/2013 10:16:20",
      "content": "<p>[quote=sck_t_ths;21436]</p>\r\n<p>Hi,</p>\r\n<p>I am new to data mining.<br>\r\nI don't want to download the images, since my network is slow.<br>\r\nAs I see all 3 standard benchmarks which use all the features provided, stand around 0.645.<br>\r\nAre there ways to reduce this number without using extracting features from the images provided?<br>\r\nSome general pointers would be appreciated ...</p>\r\n<p>Thanks.</p>\r\n<p>[/quote]</p>\r\n<p>It is possible to get into the Top-10 using the features given (provided you do the pre-processing including variable selection) and regularised logistic regression.</p>\r\n<p>One would be surprised by how much one can improve on the benchmark by doing a little bit of pre-processing and applying the same algorithms used in the benchmarks that were submitted (&#43; a little bit of tuning).</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "21527",
      "postDate": "03/22/2013 05:21:29",
      "content": "<p>Thanks Sashi for your insight... I was wondering how important variable selection is. Would not a linear model eliminate that variable anyways, if the variable is constant among all rows.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "21528",
      "postDate": "03/22/2013 07:33:58",
      "content": "<p>[quote=velociraptor;21527]</p>\r\n<p>Would not a linear model eliminate that variable anyways, if the variable is constant among all rows.</p>\r\n<p>[/quote]</p>\r\n<p>That suggests a couple of issues for me.</p>\r\n<p>Lets say you leave the variables with constant values in the model and the linear model sets their coefficients to zero. By doing this you are building a model on un-standardised variables.Anyone who attended Machine Learing classes on Coursera.org by Andrew\r\n Ng will remember that parametric models will perform well when variables are standardised.</p>\r\n<p>Proof? if the variables were standardised, then the variables with constant value would have a std. deviation of zero and whilst transforming then using (x-mean)/std those variables would be set to NaN.</p>\r\n<p>In R, if a variable is full of NaNs then most algorithms do not work; the only solution is to drop them. This step will get rid of 2415(~34%) of the 7068 variables given.</p>\r\n<p>Now, the main improvement comes from standardising your variables, yes simple standardisation &amp; regularised logistic regression with some tuning&nbsp; will improve your model by at least 28% on the logistic regression benchmark.</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "21544",
      "postDate": "03/22/2013 20:14:51",
      "content": "<p>Wow! Thanks Sashi, this is really helpful!<br>\r\nI will try out the standardization and variable selection.</p>\r\n<p>Hope it works!</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22041",
      "postDate": "04/03/2013 01:48:32",
      "content": "<p>Hey,</p>\r\n<p>I was interested in this question, so we wrote some example code (in MATLAB, using our toolbox:\r\n<a href=\"https://github.com/newfolder/prt\">https://github.com/newfolder/prt</a>&nbsp;) to do better than the classic off-the-shelf classifiers on the provided features. &nbsp;You can read about the simple approach taken here:</p>\r\n<p><a href=\"http://www.newfolderconsulting.com/node/571\">http://www.newfolderconsulting.com/node/571</a></p>\r\n<p>Of course, a lot of people are getting a lot better performance than our simple code will get you, but it should at least get you started!</p>\r\n<p>&nbsp;</p>\r\n<p>-Pete</p>",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "22046",
      "postDate": "04/03/2013 05:10:24",
      "content": "<p>Thanks Pete! That's awesome!<br>\r\nLiked the incremental approach... learnt a few new cool tricks :)<br>\r\nAnd thanks for a new tool, will try it out</p>",
      "rawMarkdown": "",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 21495,
      "author_name": "sashikanthdareddy",
      "author_url": "",
      "post_date": "03/21/2013 10:16:20",
      "content": "<p>[quote=sck_t_ths;21436]</p>\r\n<p>Hi,</p>\r\n<p>I am new to data mining.<br>\r\nI don't want to download the images, since my network is slow.<br>\r\nAs I see all 3 standard benchmarks which use all the features provided, stand around 0.645.<br>\r\nAre there ways to reduce this number without using extracting features from the images provided?<br>\r\nSome general pointers would be appreciated ...</p>\r\n<p>Thanks.</p>\r\n<p>[/quote]</p>\r\n<p>It is possible to get into the Top-10 using the features given (provided you do the pre-processing including variable selection) and regularised logistic regression.</p>\r\n<p>One would be surprised by how much one can improve on the benchmark by doing a little bit of pre-processing and applying the same algorithms used in the benchmarks that were submitted (&#43; a little bit of tuning).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 21527,
      "author_name": "velociraptor0",
      "author_url": "",
      "post_date": "03/22/2013 05:21:29",
      "content": "<p>Thanks Sashi for your insight... I was wondering how important variable selection is. Would not a linear model eliminate that variable anyways, if the variable is constant among all rows.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 21528,
      "author_name": "sashikanthdareddy",
      "author_url": "",
      "post_date": "03/22/2013 07:33:58",
      "content": "<p>[quote=velociraptor;21527]</p>\r\n<p>Would not a linear model eliminate that variable anyways, if the variable is constant among all rows.</p>\r\n<p>[/quote]</p>\r\n<p>That suggests a couple of issues for me.</p>\r\n<p>Lets say you leave the variables with constant values in the model and the linear model sets their coefficients to zero. By doing this you are building a model on un-standardised variables.Anyone who attended Machine Learing classes on Coursera.org by Andrew\r\n Ng will remember that parametric models will perform well when variables are standardised.</p>\r\n<p>Proof? if the variables were standardised, then the variables with constant value would have a std. deviation of zero and whilst transforming then using (x-mean)/std those variables would be set to NaN.</p>\r\n<p>In R, if a variable is full of NaNs then most algorithms do not work; the only solution is to drop them. This step will get rid of 2415(~34%) of the 7068 variables given.</p>\r\n<p>Now, the main improvement comes from standardising your variables, yes simple standardisation &amp; regularised logistic regression with some tuning&nbsp; will improve your model by at least 28% on the logistic regression benchmark.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 21544,
      "author_name": "vervenumen",
      "author_url": "",
      "post_date": "03/22/2013 20:14:51",
      "content": "<p>Wow! Thanks Sashi, this is really helpful!<br>\r\nI will try out the standardization and variable selection.</p>\r\n<p>Hope it works!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22041,
      "author_name": "petetorrione",
      "author_url": "",
      "post_date": "04/03/2013 01:48:32",
      "content": "<p>Hey,</p>\r\n<p>I was interested in this question, so we wrote some example code (in MATLAB, using our toolbox:\r\n<a href=\"https://github.com/newfolder/prt\">https://github.com/newfolder/prt</a>&nbsp;) to do better than the classic off-the-shelf classifiers on the provided features. &nbsp;You can read about the simple approach taken here:</p>\r\n<p><a href=\"http://www.newfolderconsulting.com/node/571\">http://www.newfolderconsulting.com/node/571</a></p>\r\n<p>Of course, a lot of people are getting a lot better performance than our simple code will get you, but it should at least get you started!</p>\r\n<p>&nbsp;</p>\r\n<p>-Pete</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 22046,
      "author_name": "vervenumen",
      "author_url": "",
      "post_date": "04/03/2013 05:10:24",
      "content": "<p>Thanks Pete! That's awesome!<br>\r\nLiked the incremental approach... learnt a few new cool tricks :)<br>\r\nAnd thanks for a new tool, will try it out</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "21436": "",
    "21495": "",
    "21527": "",
    "21528": "",
    "21544": "",
    "22041": "",
    "22046": ""
  },
  "source": "meta"
}