{
  "id": 16726,
  "title": "Difference between R and Python",
  "url": "/competitions/dato-native/discussion/16726",
  "author_name": "",
  "post_date": "2015-09-30T05:55:49.300Z",
  "votes": null,
  "comment_count": 2,
  "views": 623,
  "content": "<p>Hi kagglers,\ni am using R and randomForest, same as the proposed <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model\">python code</a> . I tried out various settings, ending up not even close to the benchmark of 90%. Do you think it depends on the implementation of the model in R? Or is my feature engineering just worse?</p>",
  "messages": [
    {
      "id": "93678",
      "postDate": "09/30/2015 05:55:49",
      "content": "<p>Hi kagglers,\ni am using R and randomForest, same as the proposed <a href=\"https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model\">python code</a> . I tried out various settings, ending up not even close to the benchmark of 90%. Do you think it depends on the implementation of the model in R? Or is my feature engineering just worse?</p>",
      "rawMarkdown": "Hi kagglers,\r\ni am using R and randomForest, same as the proposed [python code][1] . I tried out various settings, ending up not even close to the benchmark of 90%. Do you think it depends on the implementation of the model in R? Or is my feature engineering just worse?\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model",
      "votes": null
    },
    {
      "id": "93681",
      "postDate": "09/30/2015 06:57:02",
      "content": "<p>From my experience, R's randomForest is very close to python sklearn's randomForest, so it's fine with the implementation of the model in R.</p>\n\n<p>I've tried to manually extract some features almost the same with that post, and I can't get close to 0.9 either. So I think the result is very sensitive to features. </p>",
      "rawMarkdown": "From my experience, R's randomForest is very close to python sklearn's randomForest, so it's fine with the implementation of the model in R.\r\n\r\nI've tried to manually extract some features almost the same with that post, and I can't get close to 0.9 either. So I think the result is very sensitive to features.",
      "votes": null
    },
    {
      "id": "93687",
      "postDate": "09/30/2015 10:37:13",
      "content": "<p>I extracted the same feature points out via Perl (but using what should essentially be the same technique as the Python approach) and then ran them through R's glmnet, randomForest, and XGboost - I also did this for the straight feature counts as proposed (V1), and again with ratios and percents in there with the counts (V2).</p>\n\n<p>My local score is usually ~0.002 lower than the LB score from what I have seen in other tests.</p>\n\n<ul>\n<li>glmnet_v1: 0.6587269 (I didn't run it on v2)</li>\n<li>RF_v1: 0.9140512</li>\n<li>RF_v2: 0.9131562</li>\n<li>XGB_v1: 0.9054729</li>\n<li>XGB_v2: 0.9062069</li>\n</ul>\n\n<p>From what I have seen with other similar types of feature sets, ExtraTrees may perform better than all of those on the same data (or rather, same type of data - there are some feature sets it does poorly on), so maybe a higher 0.91xx or a low 0.92 - but I haven't tested it since these are not scoring better than my other models (and ET seems fantastic and taking a long time to run and then crashing with no warning and no results - although not quite on par with CBART in that respect), and when looking at how they contribute to ensembles - they seem to be lower contributions (so presumably more correlated with things I already have that are already better contributions).</p>\n\n<p>The idea is good though, particularly if you look at it from the perspective of &quot;those marked as sponsored have this value X% of the time, and those not marked as sponsored have this value Y% of the time&quot; and then start looking at other things that way.</p>\n\n<p>Also expanding on the idea, for example if tabs tend to show up more on non sponsored pages (or whatever the case may be), does their position in the document matter for that ratio? \nIf their position matters, do others as well?\nWhy would tabs show up more in non sponsored (if that is the case)? Is it because someone sat there and manually typed it out, hitting the tab key, but the sponsored stuff is more likely autogenerated? If so, are there are things that could be found that come from autogen stuff? etc</p>\n\n<p>But that is all further from &quot;load this table into R and dump it into a bunch of libraries and upload them all to see what happens&quot; and instead more data massage.  I think that relates to another post on here, as to why this contest isn't as popular as some of the others - no scripts to piggyback on, and no load and done type strategies.</p>",
      "rawMarkdown": "I extracted the same feature points out via Perl (but using what should essentially be the same technique as the Python approach) and then ran them through R's glmnet, randomForest, and XGboost - I also did this for the straight feature counts as proposed (V1), and again with ratios and percents in there with the counts (V2).\r\n\r\nMy local score is usually ~0.002 lower than the LB score from what I have seen in other tests.\r\n\r\n - glmnet_v1: 0.6587269 (I didn't run it on v2)\r\n - RF_v1: 0.9140512\r\n - RF_v2: 0.9131562\r\n - XGB_v1: 0.9054729\r\n - XGB_v2: 0.9062069\r\n\r\nFrom what I have seen with other similar types of feature sets, ExtraTrees may perform better than all of those on the same data (or rather, same type of data - there are some feature sets it does poorly on), so maybe a higher 0.91xx or a low 0.92 - but I haven't tested it since these are not scoring better than my other models (and ET seems fantastic and taking a long time to run and then crashing with no warning and no results - although not quite on par with CBART in that respect), and when looking at how they contribute to ensembles - they seem to be lower contributions (so presumably more correlated with things I already have that are already better contributions).\r\n\r\nThe idea is good though, particularly if you look at it from the perspective of \"those marked as sponsored have this value X% of the time, and those not marked as sponsored have this value Y% of the time\" and then start looking at other things that way.\r\n\r\nAlso expanding on the idea, for example if tabs tend to show up more on non sponsored pages (or whatever the case may be), does their position in the document matter for that ratio? \r\nIf their position matters, do others as well?\r\nWhy would tabs show up more in non sponsored (if that is the case)? Is it because someone sat there and manually typed it out, hitting the tab key, but the sponsored stuff is more likely autogenerated? If so, are there are things that could be found that come from autogen stuff? etc\r\n\r\nBut that is all further from \"load this table into R and dump it into a bunch of libraries and upload them all to see what happens\" and instead more data massage.  I think that relates to another post on here, as to why this contest isn't as popular as some of the others - no scripts to piggyback on, and no load and done type strategies.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 93681,
      "author_name": "yejiming",
      "author_url": "",
      "post_date": "09/30/2015 06:57:02",
      "content": "<p>From my experience, R's randomForest is very close to python sklearn's randomForest, so it's fine with the implementation of the model in R.</p>\n\n<p>I've tried to manually extract some features almost the same with that post, and I can't get close to 0.9 either. So I think the result is very sensitive to features. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 93687,
      "author_name": "omgponies",
      "author_url": "",
      "post_date": "09/30/2015 10:37:13",
      "content": "<p>I extracted the same feature points out via Perl (but using what should essentially be the same technique as the Python approach) and then ran them through R's glmnet, randomForest, and XGboost - I also did this for the straight feature counts as proposed (V1), and again with ratios and percents in there with the counts (V2).</p>\n\n<p>My local score is usually ~0.002 lower than the LB score from what I have seen in other tests.</p>\n\n<ul>\n<li>glmnet_v1: 0.6587269 (I didn't run it on v2)</li>\n<li>RF_v1: 0.9140512</li>\n<li>RF_v2: 0.9131562</li>\n<li>XGB_v1: 0.9054729</li>\n<li>XGB_v2: 0.9062069</li>\n</ul>\n\n<p>From what I have seen with other similar types of feature sets, ExtraTrees may perform better than all of those on the same data (or rather, same type of data - there are some feature sets it does poorly on), so maybe a higher 0.91xx or a low 0.92 - but I haven't tested it since these are not scoring better than my other models (and ET seems fantastic and taking a long time to run and then crashing with no warning and no results - although not quite on par with CBART in that respect), and when looking at how they contribute to ensembles - they seem to be lower contributions (so presumably more correlated with things I already have that are already better contributions).</p>\n\n<p>The idea is good though, particularly if you look at it from the perspective of &quot;those marked as sponsored have this value X% of the time, and those not marked as sponsored have this value Y% of the time&quot; and then start looking at other things that way.</p>\n\n<p>Also expanding on the idea, for example if tabs tend to show up more on non sponsored pages (or whatever the case may be), does their position in the document matter for that ratio? \nIf their position matters, do others as well?\nWhy would tabs show up more in non sponsored (if that is the case)? Is it because someone sat there and manually typed it out, hitting the tab key, but the sponsored stuff is more likely autogenerated? If so, are there are things that could be found that come from autogen stuff? etc</p>\n\n<p>But that is all further from &quot;load this table into R and dump it into a bunch of libraries and upload them all to see what happens&quot; and instead more data massage.  I think that relates to another post on here, as to why this contest isn't as popular as some of the others - no scripts to piggyback on, and no load and done type strategies.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "93678": "Hi kagglers,\r\ni am using R and randomForest, same as the proposed [python code][1] . I tried out various settings, ending up not even close to the benchmark of 90%. Do you think it depends on the implementation of the model in R? Or is my feature engineering just worse?\r\n\r\n\r\n  [1]: https://www.kaggle.com/c/dato-native/forums/t/16626/beat-the-benchmark-0-90388-with-simple-model",
    "93681": "From my experience, R's randomForest is very close to python sklearn's randomForest, so it's fine with the implementation of the model in R.\r\n\r\nI've tried to manually extract some features almost the same with that post, and I can't get close to 0.9 either. So I think the result is very sensitive to features.",
    "93687": "I extracted the same feature points out via Perl (but using what should essentially be the same technique as the Python approach) and then ran them through R's glmnet, randomForest, and XGboost - I also did this for the straight feature counts as proposed (V1), and again with ratios and percents in there with the counts (V2).\r\n\r\nMy local score is usually ~0.002 lower than the LB score from what I have seen in other tests.\r\n\r\n - glmnet_v1: 0.6587269 (I didn't run it on v2)\r\n - RF_v1: 0.9140512\r\n - RF_v2: 0.9131562\r\n - XGB_v1: 0.9054729\r\n - XGB_v2: 0.9062069\r\n\r\nFrom what I have seen with other similar types of feature sets, ExtraTrees may perform better than all of those on the same data (or rather, same type of data - there are some feature sets it does poorly on), so maybe a higher 0.91xx or a low 0.92 - but I haven't tested it since these are not scoring better than my other models (and ET seems fantastic and taking a long time to run and then crashing with no warning and no results - although not quite on par with CBART in that respect), and when looking at how they contribute to ensembles - they seem to be lower contributions (so presumably more correlated with things I already have that are already better contributions).\r\n\r\nThe idea is good though, particularly if you look at it from the perspective of \"those marked as sponsored have this value X% of the time, and those not marked as sponsored have this value Y% of the time\" and then start looking at other things that way.\r\n\r\nAlso expanding on the idea, for example if tabs tend to show up more on non sponsored pages (or whatever the case may be), does their position in the document matter for that ratio? \r\nIf their position matters, do others as well?\r\nWhy would tabs show up more in non sponsored (if that is the case)? Is it because someone sat there and manually typed it out, hitting the tab key, but the sponsored stuff is more likely autogenerated? If so, are there are things that could be found that come from autogen stuff? etc\r\n\r\nBut that is all further from \"load this table into R and dump it into a bunch of libraries and upload them all to see what happens\" and instead more data massage.  I think that relates to another post on here, as to why this contest isn't as popular as some of the others - no scripts to piggyback on, and no load and done type strategies."
  },
  "source": "meta"
}