{
  "id": 20328,
  "title": "Vowpal Wabbit",
  "url": "/competitions/expedia-hotel-recommendations/discussion/20328",
  "author_name": "",
  "post_date": "2016-04-22T07:52:54.320Z",
  "votes": 2,
  "comment_count": 4,
  "views": 1204,
  "content": "<p>Hey guys,\nto tackle the dataset I thought about using vowpal wabbit.\nHas someone an idea how to use this here?\nRegards,\nSterby</p>",
  "messages": [
    {
      "id": "116140",
      "postDate": "04/22/2016 07:52:54",
      "content": "<p>Hey guys,\nto tackle the dataset I thought about using vowpal wabbit.\nHas someone an idea how to use this here?\nRegards,\nSterby</p>",
      "rawMarkdown": "Hey guys,\r\nto tackle the dataset I thought about using vowpal wabbit.\r\nHas someone an idea how to use this here?\r\nRegards,\r\nSterby",
      "votes": null
    },
    {
      "id": "117960",
      "postDate": "05/02/2016 02:48:39",
      "content": "<p>Probably the same way as: <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20503/is-this-competition-a-100-classifications-issue\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20503/is-this-competition-a-100-classifications-issue</a></p>\n\n<p>Where you would make it into a binary classification problem with 100 different classes. Take the top 5 percentages. There's probably some other criterion you might want to use in conjunction with the &quot;top 5&quot; portion.</p>",
      "rawMarkdown": "Probably the same way as: https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20503/is-this-competition-a-100-classifications-issue\r\n\r\nWhere you would make it into a binary classification problem with 100 different classes. Take the top 5 percentages. There's probably some other criterion you might want to use in conjunction with the \"top 5\" portion.",
      "votes": null
    },
    {
      "id": "117994",
      "postDate": "05/02/2016 11:59:07",
      "content": "<p>Hi,</p>\n\n<p>I used vw once for the dataset here. And I got a very, very bad result. That is not to say that vw can't do it. Maybe just I can't. I ended up with a MAP@5 of 0.02 - yes, no typo. Logloss while running was around 5.x.</p>\n\n<p>Basically I engineered some features that come to mind fast, added destination info (although modified a bit - 10 ** x * 100 or something like that) and sent it off.</p>\n\n<p>With more or less the same features and 1% of train data or only the is_booking rows of data and xgboost I get around 0.28 MAP@5.</p>\n\n<p>On a side note: My input file was big: 40GB :-) Usually RAM is the bottleneck, not disk. And vw ran a few hours. But what is that compared to xgboost?</p>\n\n<p>I am looking forward to someone who has better results or who can hint me in the right direction!</p>\n\n<p>Other than -oaa I did not use any fancy parameters.</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nI used vw once for the dataset here. And I got a very, very bad result. That is not to say that vw can't do it. Maybe just I can't. I ended up with a MAP@5 of 0.02 - yes, no typo. Logloss while running was around 5.x.\r\n\r\nBasically I engineered some features that come to mind fast, added destination info (although modified a bit - 10 ** x * 100 or something like that) and sent it off.\r\n\r\nWith more or less the same features and 1% of train data or only the is_booking rows of data and xgboost I get around 0.28 MAP@5.\r\n\r\nOn a side note: My input file was big: 40GB :-) Usually RAM is the bottleneck, not disk. And vw ran a few hours. But what is that compared to xgboost?\r\n\r\nI am looking forward to someone who has better results or who can hint me in the right direction!\r\n\r\nOther than -oaa I did not use any fancy parameters.\r\n\r\nGerhard",
      "votes": null
    },
    {
      "id": "121004",
      "postDate": "05/22/2016 17:09:47",
      "content": "<p>[quote=MightyBird;117994]</p>\n\n<p>Hi,</p>\n\n<p>I used vw once for the dataset here. And I got a very, very bad result. That is not to say that vw can't do it. Maybe just I can't. I ended up with a MAP@5 of 0.02 - yes, no typo. Logloss while running was around 5.x.</p>\n\n<p>Basically I engineered some features that come to mind fast, added destination info (although modified a bit - 10 ** x * 100 or something like that) and sent it off.</p>\n\n<p>With more or less the same features and 1% of train data or only the is_booking rows of data and xgboost I get around 0.28 MAP@5.</p>\n\n<p>On a side note: My input file was big: 40GB :-) Usually RAM is the bottleneck, not disk. And vw ran a few hours. But what is that compared to xgboost?</p>\n\n<p>I am looking forward to someone who has better results or who can hint me in the right direction!</p>\n\n<p>Other than -oaa I did not use any fancy parameters.</p>\n\n<p>Gerhard</p>\n\n<p>[/quote]</p>\n\n<p>Gerhard, I understand that this would be a very stupid question, but anyway. How exactly did you manage to use vw here? Did you train 100 instances of it or did you use it with --oaa and --top 5 flags? I am trying to do the latter, but it only returns one prediction for each example.</p>\n\n<p>Thank you in advance.</p>\n\n<p>Mikhail.</p>",
      "rawMarkdown": "[quote=MightyBird;117994]\r\n\r\nHi,\r\n\r\nI used vw once for the dataset here. And I got a very, very bad result. That is not to say that vw can't do it. Maybe just I can't. I ended up with a MAP@5 of 0.02 - yes, no typo. Logloss while running was around 5.x.\r\n\r\nBasically I engineered some features that come to mind fast, added destination info (although modified a bit - 10 ** x * 100 or something like that) and sent it off.\r\n\r\nWith more or less the same features and 1% of train data or only the is_booking rows of data and xgboost I get around 0.28 MAP@5.\r\n\r\nOn a side note: My input file was big: 40GB :-) Usually RAM is the bottleneck, not disk. And vw ran a few hours. But what is that compared to xgboost?\r\n\r\nI am looking forward to someone who has better results or who can hint me in the right direction!\r\n\r\nOther than -oaa I did not use any fancy parameters.\r\n\r\nGerhard\r\n\r\n\r\n[/quote]\r\n\r\nGerhard, I understand that this would be a very stupid question, but anyway. How exactly did you manage to use vw here? Did you train 100 instances of it or did you use it with --oaa and --top 5 flags? I am trying to do the latter, but it only returns one prediction for each example.\r\n\r\nThank you in advance.\r\n\r\nMikhail.",
      "votes": null
    },
    {
      "id": "121006",
      "postDate": "05/22/2016 17:32:54",
      "content": "<p>Hi,</p>\n\n<p>I used -oaa (so one run training, one for predictions) and then told it to produce raw predictions. From there you can get top5 elements.</p>\n\n<p>By the way, I merged destination features into all instances. Maybe not a good idea...</p>\n\n<p>Gerhard</p>",
      "rawMarkdown": "Hi,\r\n\r\nI used -oaa (so one run training, one for predictions) and then told it to produce raw predictions. From there you can get top5 elements.\r\n\r\nBy the way, I merged destination features into all instances. Maybe not a good idea...\r\n\r\nGerhard",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 117960,
      "author_name": "",
      "author_url": "",
      "post_date": "05/02/2016 02:48:39",
      "content": "<p>Probably the same way as: <a href=\"https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20503/is-this-competition-a-100-classifications-issue\">https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20503/is-this-competition-a-100-classifications-issue</a></p>\n\n<p>Where you would make it into a binary classification problem with 100 different classes. Take the top 5 percentages. There's probably some other criterion you might want to use in conjunction with the &quot;top 5&quot; portion.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 117994,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "05/02/2016 11:59:07",
      "content": "<p>Hi,</p>\n\n<p>I used vw once for the dataset here. And I got a very, very bad result. That is not to say that vw can't do it. Maybe just I can't. I ended up with a MAP@5 of 0.02 - yes, no typo. Logloss while running was around 5.x.</p>\n\n<p>Basically I engineered some features that come to mind fast, added destination info (although modified a bit - 10 ** x * 100 or something like that) and sent it off.</p>\n\n<p>With more or less the same features and 1% of train data or only the is_booking rows of data and xgboost I get around 0.28 MAP@5.</p>\n\n<p>On a side note: My input file was big: 40GB :-) Usually RAM is the bottleneck, not disk. And vw ran a few hours. But what is that compared to xgboost?</p>\n\n<p>I am looking forward to someone who has better results or who can hint me in the right direction!</p>\n\n<p>Other than -oaa I did not use any fancy parameters.</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121004,
      "author_name": "akimovmike",
      "author_url": "",
      "post_date": "05/22/2016 17:09:47",
      "content": "<p>[quote=MightyBird;117994]</p>\n\n<p>Hi,</p>\n\n<p>I used vw once for the dataset here. And I got a very, very bad result. That is not to say that vw can't do it. Maybe just I can't. I ended up with a MAP@5 of 0.02 - yes, no typo. Logloss while running was around 5.x.</p>\n\n<p>Basically I engineered some features that come to mind fast, added destination info (although modified a bit - 10 ** x * 100 or something like that) and sent it off.</p>\n\n<p>With more or less the same features and 1% of train data or only the is_booking rows of data and xgboost I get around 0.28 MAP@5.</p>\n\n<p>On a side note: My input file was big: 40GB :-) Usually RAM is the bottleneck, not disk. And vw ran a few hours. But what is that compared to xgboost?</p>\n\n<p>I am looking forward to someone who has better results or who can hint me in the right direction!</p>\n\n<p>Other than -oaa I did not use any fancy parameters.</p>\n\n<p>Gerhard</p>\n\n<p>[/quote]</p>\n\n<p>Gerhard, I understand that this would be a very stupid question, but anyway. How exactly did you manage to use vw here? Did you train 100 instances of it or did you use it with --oaa and --top 5 flags? I am trying to do the latter, but it only returns one prediction for each example.</p>\n\n<p>Thank you in advance.</p>\n\n<p>Mikhail.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121006,
      "author_name": "mightybird",
      "author_url": "",
      "post_date": "05/22/2016 17:32:54",
      "content": "<p>Hi,</p>\n\n<p>I used -oaa (so one run training, one for predictions) and then told it to produce raw predictions. From there you can get top5 elements.</p>\n\n<p>By the way, I merged destination features into all instances. Maybe not a good idea...</p>\n\n<p>Gerhard</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "116140": "Hey guys,\r\nto tackle the dataset I thought about using vowpal wabbit.\r\nHas someone an idea how to use this here?\r\nRegards,\r\nSterby",
    "117960": "Probably the same way as: https://www.kaggle.com/c/expedia-hotel-recommendations/forums/t/20503/is-this-competition-a-100-classifications-issue\r\n\r\nWhere you would make it into a binary classification problem with 100 different classes. Take the top 5 percentages. There's probably some other criterion you might want to use in conjunction with the \"top 5\" portion.",
    "117994": "Hi,\r\n\r\nI used vw once for the dataset here. And I got a very, very bad result. That is not to say that vw can't do it. Maybe just I can't. I ended up with a MAP@5 of 0.02 - yes, no typo. Logloss while running was around 5.x.\r\n\r\nBasically I engineered some features that come to mind fast, added destination info (although modified a bit - 10 ** x * 100 or something like that) and sent it off.\r\n\r\nWith more or less the same features and 1% of train data or only the is_booking rows of data and xgboost I get around 0.28 MAP@5.\r\n\r\nOn a side note: My input file was big: 40GB :-) Usually RAM is the bottleneck, not disk. And vw ran a few hours. But what is that compared to xgboost?\r\n\r\nI am looking forward to someone who has better results or who can hint me in the right direction!\r\n\r\nOther than -oaa I did not use any fancy parameters.\r\n\r\nGerhard",
    "121004": "[quote=MightyBird;117994]\r\n\r\nHi,\r\n\r\nI used vw once for the dataset here. And I got a very, very bad result. That is not to say that vw can't do it. Maybe just I can't. I ended up with a MAP@5 of 0.02 - yes, no typo. Logloss while running was around 5.x.\r\n\r\nBasically I engineered some features that come to mind fast, added destination info (although modified a bit - 10 ** x * 100 or something like that) and sent it off.\r\n\r\nWith more or less the same features and 1% of train data or only the is_booking rows of data and xgboost I get around 0.28 MAP@5.\r\n\r\nOn a side note: My input file was big: 40GB :-) Usually RAM is the bottleneck, not disk. And vw ran a few hours. But what is that compared to xgboost?\r\n\r\nI am looking forward to someone who has better results or who can hint me in the right direction!\r\n\r\nOther than -oaa I did not use any fancy parameters.\r\n\r\nGerhard\r\n\r\n\r\n[/quote]\r\n\r\nGerhard, I understand that this would be a very stupid question, but anyway. How exactly did you manage to use vw here? Did you train 100 instances of it or did you use it with --oaa and --top 5 flags? I am trying to do the latter, but it only returns one prediction for each example.\r\n\r\nThank you in advance.\r\n\r\nMikhail.",
    "121006": "Hi,\r\n\r\nI used -oaa (so one run training, one for predictions) and then told it to produce raw predictions. From there you can get top5 elements.\r\n\r\nBy the way, I merged destination features into all instances. Maybe not a good idea...\r\n\r\nGerhard"
  },
  "source": "meta"
}