{
  "id": 21419,
  "title": "Machine Learning Final project (needs advice) ",
  "url": "/competitions/expedia-hotel-recommendations/discussion/21419",
  "author_name": "",
  "post_date": "2016-06-04T05:07:51.750Z",
  "votes": null,
  "comment_count": 5,
  "views": 1452,
  "content": "<p>Hi, </p>\n\n<p>I am Muratcan Cicek from Oregon. I am quite beginner in the Machine Learning. I don't know why but I have  chosen this problem for my final project. Now I have to do some experiments and report the results. </p>\n\n<p>I have actually submitted the leakage solution with 0.49 accuracy and I understood its logic, how it works. But I actually don't care the competition and my instructors wants me use some ML Algorithms and explain why they work or don't work. </p>\n\n<p>I can also do downsampling on data, currently i tried to reduce the frequencies of each cluster to same value and I have 2 millions rows now in total.  but I don't know it's a scientific way. </p>\n\n<p>So, I need find a correct way to downsampling and a couple preferable ML methods for this problem. I linear regression but it gives 0.0087 and I got memory error with random forests.</p>\n\n<p>By the way, I have <a href=\"https://papers.nips.cc/paper/5367-online-and-stochastic-gradient-methods-for-non-decomposable-loss-functions.pdf\" title=\"method\">this method</a> suggested. But I couldn't implement it. Maybe it works for you.</p>\n\n<p>Thank you,</p>\n\n<p>Muratcan </p>",
  "messages": [
    {
      "id": "122441",
      "postDate": "06/04/2016 05:07:51",
      "content": "<p>Hi, </p>\n\n<p>I am Muratcan Cicek from Oregon. I am quite beginner in the Machine Learning. I don't know why but I have  chosen this problem for my final project. Now I have to do some experiments and report the results. </p>\n\n<p>I have actually submitted the leakage solution with 0.49 accuracy and I understood its logic, how it works. But I actually don't care the competition and my instructors wants me use some ML Algorithms and explain why they work or don't work. </p>\n\n<p>I can also do downsampling on data, currently i tried to reduce the frequencies of each cluster to same value and I have 2 millions rows now in total.  but I don't know it's a scientific way. </p>\n\n<p>So, I need find a correct way to downsampling and a couple preferable ML methods for this problem. I linear regression but it gives 0.0087 and I got memory error with random forests.</p>\n\n<p>By the way, I have <a href=\"https://papers.nips.cc/paper/5367-online-and-stochastic-gradient-methods-for-non-decomposable-loss-functions.pdf\" title=\"method\">this method</a> suggested. But I couldn't implement it. Maybe it works for you.</p>\n\n<p>Thank you,</p>\n\n<p>Muratcan </p>",
      "rawMarkdown": "Hi, \r\n\r\nI am Muratcan Cicek from Oregon. I am quite beginner in the Machine Learning. I don't know why but I have  chosen this problem for my final project. Now I have to do some experiments and report the results. \r\n\r\nI have actually submitted the leakage solution with 0.49 accuracy and I understood its logic, how it works. But I actually don't care the competition and my instructors wants me use some ML Algorithms and explain why they work or don't work. \r\n\r\nI can also do downsampling on data, currently i tried to reduce the frequencies of each cluster to same value and I have 2 millions rows now in total.  but I don't know it's a scientific way. \r\n\r\nSo, I need find a correct way to downsampling and a couple preferable ML methods for this problem. I linear regression but it gives 0.0087 and I got memory error with random forests.\r\n\r\nBy the way, I have [this method][1] suggested. But I couldn't implement it. Maybe it works for you.\r\n\r\nThank you,\r\n\r\nMuratcan \r\n\r\n\r\n  [1]: https://papers.nips.cc/paper/5367-online-and-stochastic-gradient-methods-for-non-decomposable-loss-functions.pdf \"method\"",
      "votes": null
    },
    {
      "id": "122444",
      "postDate": "06/04/2016 05:28:22",
      "content": "<p>Firstly, you may want to check if the data here can be used outside of the competition, usually the rules state that you need permission from the host to use it elsewhere. Second, if you're a beginner in ML I honestly would not suggest using this data set for your project. The sheer size of the data set along with the large number of classes make this a pretty challenging problem even for those with experience and the computing resources. Kaggle has a bunch of <a href=\"https://www.kaggle.com/datasets\">datasets</a> which are interesting, much more manageable, and can be used for academic purposes.</p>",
      "rawMarkdown": "Firstly, you may want to check if the data here can be used outside of the competition, usually the rules state that you need permission from the host to use it elsewhere. Second, if you're a beginner in ML I honestly would not suggest using this data set for your project. The sheer size of the data set along with the large number of classes make this a pretty challenging problem even for those with experience and the computing resources. Kaggle has a bunch of [datasets][1] which are interesting, much more manageable, and can be used for academic purposes.\r\n\r\n\r\n  [1]: https://www.kaggle.com/datasets",
      "votes": null
    },
    {
      "id": "122448",
      "postDate": "06/04/2016 05:48:06",
      "content": "<p>Thanks for replying. You are right I have to check for the permission but it's just an introduction course and I will not publish the data or any results anywhere else. My instructor will just grade me on my report and that's all.</p>\n\n<p>It's one of my biggest mistakes that choosing this data but it's too late. </p>",
      "rawMarkdown": "Thanks for replying. You are right I have to check for the permission but it's just an introduction course and I will not publish the data or any results anywhere else. My instructor will just grade me on my report and that's all.\r\n\r\nIt's one of my biggest mistakes that choosing this data but it's too late.",
      "votes": null
    },
    {
      "id": "122479",
      "postDate": "06/04/2016 13:42:59",
      "content": "<p>Putting any use rules aside, I think you still have a lot of flexibility and have a lot to write about if you structure the problem differently.  This assumes that you are only committed to the data in general, and not the exact data or the exact evaluation metric.  If that assumption doesn't hold, then you are in really bad shape.  So, assuming the above holds:</p>\n\n<ol>\n<li>Forget about the leaderboard and simply make 2014 your validation set and 2013 your training set.  That gives you flexibility on the evaluation metric as you can measure it yourself.</li>\n<li>Pick the two most popular hotel metrics and turn this into a binary classification problem with logloss or AUC as your evaluation metric.  You won't have to down sample the data because there'll only be a few hundred thousand records remaining.</li>\n</ol>\n\n<p>From there, you'll have a million of things to experiment with and write about.  Why do some seemingly numeric features need to be treated as categorical for some algorithms?  What effect does having the distance feature have on overfitting (think about the data leak)?  Why is scaling some of your features important for certain algorithms?  What is the variance vs bias trade off across the algorithms you choose with learning curves of training and validation scores to support your analysis.</p>\n\n<p>Good luck.</p>",
      "rawMarkdown": "Putting any use rules aside, I think you still have a lot of flexibility and have a lot to write about if you structure the problem differently.  This assumes that you are only committed to the data in general, and not the exact data or the exact evaluation metric.  If that assumption doesn't hold, then you are in really bad shape.  So, assuming the above holds:\r\n\r\n 1. Forget about the leaderboard and simply make 2014 your validation set and 2013 your training set.  That gives you flexibility on the evaluation metric as you can measure it yourself.\r\n 2. Pick the two most popular hotel metrics and turn this into a binary classification problem with logloss or AUC as your evaluation metric.  You won't have to down sample the data because there'll only be a few hundred thousand records remaining.\r\n \r\nFrom there, you'll have a million of things to experiment with and write about.  Why do some seemingly numeric features need to be treated as categorical for some algorithms?  What effect does having the distance feature have on overfitting (think about the data leak)?  Why is scaling some of your features important for certain algorithms?  What is the variance vs bias trade off across the algorithms you choose with learning curves of training and validation scores to support your analysis.\r\n\r\nGood luck.",
      "votes": null
    },
    {
      "id": "122573",
      "postDate": "06/05/2016 09:08:52",
      "content": "<p>I talked my  instructor about your second suggestion and she accepted. I downsampled data to 2 clusters as you advised.  then, I use logistic regression and I got %52 accuracy. ,</p>\n\n<p>By the way, I am using Imputer for the missing data but are there any alternative ways to handle them you can suggest?</p>\n\n<p>Thank you again.</p>",
      "rawMarkdown": "I talked my  instructor about your second suggestion and she accepted. I downsampled data to 2 clusters as you advised.  then, I use logistic regression and I got %52 accuracy. ,\r\n\r\nBy the way, I am using Imputer for the missing data but are there any alternative ways to handle them you can suggest?\r\n\r\nThank you again.",
      "votes": null
    },
    {
      "id": "122621",
      "postDate": "06/05/2016 21:54:39",
      "content": "<p>There are a few ways to handle missing data. You already tried imputing. You could also try some different ML algorithms that handle missing data. For example, XGBoost is very popular now, and you can specify a value for missing data.</p>",
      "rawMarkdown": "There are a few ways to handle missing data. You already tried imputing. You could also try some different ML algorithms that handle missing data. For example, XGBoost is very popular now, and you can specify a value for missing data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 122444,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "06/04/2016 05:28:22",
      "content": "<p>Firstly, you may want to check if the data here can be used outside of the competition, usually the rules state that you need permission from the host to use it elsewhere. Second, if you're a beginner in ML I honestly would not suggest using this data set for your project. The sheer size of the data set along with the large number of classes make this a pretty challenging problem even for those with experience and the computing resources. Kaggle has a bunch of <a href=\"https://www.kaggle.com/datasets\">datasets</a> which are interesting, much more manageable, and can be used for academic purposes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122448,
      "author_name": "muratcancicek",
      "author_url": "",
      "post_date": "06/04/2016 05:48:06",
      "content": "<p>Thanks for replying. You are right I have to check for the permission but it's just an introduction course and I will not publish the data or any results anywhere else. My instructor will just grade me on my report and that's all.</p>\n\n<p>It's one of my biggest mistakes that choosing this data but it's too late. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122479,
      "author_name": "davidshinn",
      "author_url": "",
      "post_date": "06/04/2016 13:42:59",
      "content": "<p>Putting any use rules aside, I think you still have a lot of flexibility and have a lot to write about if you structure the problem differently.  This assumes that you are only committed to the data in general, and not the exact data or the exact evaluation metric.  If that assumption doesn't hold, then you are in really bad shape.  So, assuming the above holds:</p>\n\n<ol>\n<li>Forget about the leaderboard and simply make 2014 your validation set and 2013 your training set.  That gives you flexibility on the evaluation metric as you can measure it yourself.</li>\n<li>Pick the two most popular hotel metrics and turn this into a binary classification problem with logloss or AUC as your evaluation metric.  You won't have to down sample the data because there'll only be a few hundred thousand records remaining.</li>\n</ol>\n\n<p>From there, you'll have a million of things to experiment with and write about.  Why do some seemingly numeric features need to be treated as categorical for some algorithms?  What effect does having the distance feature have on overfitting (think about the data leak)?  Why is scaling some of your features important for certain algorithms?  What is the variance vs bias trade off across the algorithms you choose with learning curves of training and validation scores to support your analysis.</p>\n\n<p>Good luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122573,
      "author_name": "muratcancicek",
      "author_url": "",
      "post_date": "06/05/2016 09:08:52",
      "content": "<p>I talked my  instructor about your second suggestion and she accepted. I downsampled data to 2 clusters as you advised.  then, I use logistic regression and I got %52 accuracy. ,</p>\n\n<p>By the way, I am using Imputer for the missing data but are there any alternative ways to handle them you can suggest?</p>\n\n<p>Thank you again.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 122621,
      "author_name": "jeffhebert",
      "author_url": "",
      "post_date": "06/05/2016 21:54:39",
      "content": "<p>There are a few ways to handle missing data. You already tried imputing. You could also try some different ML algorithms that handle missing data. For example, XGBoost is very popular now, and you can specify a value for missing data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "122441": "Hi, \r\n\r\nI am Muratcan Cicek from Oregon. I am quite beginner in the Machine Learning. I don't know why but I have  chosen this problem for my final project. Now I have to do some experiments and report the results. \r\n\r\nI have actually submitted the leakage solution with 0.49 accuracy and I understood its logic, how it works. But I actually don't care the competition and my instructors wants me use some ML Algorithms and explain why they work or don't work. \r\n\r\nI can also do downsampling on data, currently i tried to reduce the frequencies of each cluster to same value and I have 2 millions rows now in total.  but I don't know it's a scientific way. \r\n\r\nSo, I need find a correct way to downsampling and a couple preferable ML methods for this problem. I linear regression but it gives 0.0087 and I got memory error with random forests.\r\n\r\nBy the way, I have [this method][1] suggested. But I couldn't implement it. Maybe it works for you.\r\n\r\nThank you,\r\n\r\nMuratcan \r\n\r\n\r\n  [1]: https://papers.nips.cc/paper/5367-online-and-stochastic-gradient-methods-for-non-decomposable-loss-functions.pdf \"method\"",
    "122444": "Firstly, you may want to check if the data here can be used outside of the competition, usually the rules state that you need permission from the host to use it elsewhere. Second, if you're a beginner in ML I honestly would not suggest using this data set for your project. The sheer size of the data set along with the large number of classes make this a pretty challenging problem even for those with experience and the computing resources. Kaggle has a bunch of [datasets][1] which are interesting, much more manageable, and can be used for academic purposes.\r\n\r\n\r\n  [1]: https://www.kaggle.com/datasets",
    "122448": "Thanks for replying. You are right I have to check for the permission but it's just an introduction course and I will not publish the data or any results anywhere else. My instructor will just grade me on my report and that's all.\r\n\r\nIt's one of my biggest mistakes that choosing this data but it's too late.",
    "122479": "Putting any use rules aside, I think you still have a lot of flexibility and have a lot to write about if you structure the problem differently.  This assumes that you are only committed to the data in general, and not the exact data or the exact evaluation metric.  If that assumption doesn't hold, then you are in really bad shape.  So, assuming the above holds:\r\n\r\n 1. Forget about the leaderboard and simply make 2014 your validation set and 2013 your training set.  That gives you flexibility on the evaluation metric as you can measure it yourself.\r\n 2. Pick the two most popular hotel metrics and turn this into a binary classification problem with logloss or AUC as your evaluation metric.  You won't have to down sample the data because there'll only be a few hundred thousand records remaining.\r\n \r\nFrom there, you'll have a million of things to experiment with and write about.  Why do some seemingly numeric features need to be treated as categorical for some algorithms?  What effect does having the distance feature have on overfitting (think about the data leak)?  Why is scaling some of your features important for certain algorithms?  What is the variance vs bias trade off across the algorithms you choose with learning curves of training and validation scores to support your analysis.\r\n\r\nGood luck.",
    "122573": "I talked my  instructor about your second suggestion and she accepted. I downsampled data to 2 clusters as you advised.  then, I use logistic regression and I got %52 accuracy. ,\r\n\r\nBy the way, I am using Imputer for the missing data but are there any alternative ways to handle them you can suggest?\r\n\r\nThank you again.",
    "122621": "There are a few ways to handle missing data. You already tried imputing. You could also try some different ML algorithms that handle missing data. For example, XGBoost is very popular now, and you can specify a value for missing data."
  },
  "source": "meta"
}