{
  "id": 58674,
  "title": "New user_id feature?",
  "url": "/competitions/avito-demand-prediction/discussion/58674",
  "author_name": "",
  "post_date": "2018-06-12T06:39:05.875023600Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In the Two Sigma Rental Inquiries Competition, manager features, especially manager_skill, drastically improved results. This guy here explains how to do it:\n<a href=\"https://www.kaggle.com/den3b81/improve-perfomances-using-manager-features\">https://www.kaggle.com/den3b81/improve-perfomances-using-manager-features</a></p>\n\n<p>Has anybody here considered making a feature which is going to represent the marketing ability of each user, something like user_id_skill, which could be computed using 'user_id' and 'deal_probability' columns?</p>",
  "messages": [
    {
      "id": "341727",
      "postDate": "06/12/2018 06:39:05",
      "content": "<p>In the Two Sigma Rental Inquiries Competition, manager features, especially manager_skill, drastically improved results. This guy here explains how to do it:\n<a href=\"https://www.kaggle.com/den3b81/improve-perfomances-using-manager-features\">https://www.kaggle.com/den3b81/improve-perfomances-using-manager-features</a></p>\n\n<p>Has anybody here considered making a feature which is going to represent the marketing ability of each user, something like user_id_skill, which could be computed using 'user_id' and 'deal_probability' columns?</p>",
      "rawMarkdown": "In the Two Sigma Rental Inquiries Competition, manager features, especially manager_skill, drastically improved results. This guy here explains how to do it:\nhttps://www.kaggle.com/den3b81/improve-perfomances-using-manager-features\n\nHas anybody here considered making a feature which is going to represent the marketing ability of each user, something like user_id_skill, which could be computed using 'user_id' and 'deal_probability' columns?",
      "votes": null
    },
    {
      "id": "342046",
      "postDate": "06/12/2018 18:47:51",
      "content": "<p>My guess is that this will be much harder to do in this competition because of the amount of users with very few listings and thus you'll be at high risk of overfitting. Though there still may be merit in this approach. Good luck!</p>",
      "rawMarkdown": "My guess is that this will be much harder to do in this competition because of the amount of users with very few listings and thus you'll be at high risk of overfitting. Though there still may be merit in this approach. Good luck!",
      "votes": null
    },
    {
      "id": "342211",
      "postDate": "06/13/2018 05:27:36",
      "content": "<p>That's a valid point. I've also found out that there are only 67 000 common users in the test_df and train_df,  and that test_df holds an additional 220 000 new users for which we could not compute the user_skill because of the absence of the deal_probability column. We would probably have to assign mean values to these users and I don't think that this has the potential to improve out predictions.</p>",
      "rawMarkdown": "That's a valid point. I've also found out that there are only 67 000 common users in the test_df and train_df,  and that test_df holds an additional 220 000 new users for which we could not compute the user_skill because of the absence of the deal_probability column. We would probably have to assign mean values to these users and I don't think that this has the potential to improve out predictions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 342046,
      "author_name": "peterhurford",
      "author_url": "",
      "post_date": "06/12/2018 18:47:51",
      "content": "<p>My guess is that this will be much harder to do in this competition because of the amount of users with very few listings and thus you'll be at high risk of overfitting. Though there still may be merit in this approach. Good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 342211,
          "author_name": "iandzindo",
          "author_url": "",
          "post_date": "06/13/2018 05:27:36",
          "content": "<p>That's a valid point. I've also found out that there are only 67 000 common users in the test_df and train_df,  and that test_df holds an additional 220 000 new users for which we could not compute the user_skill because of the absence of the deal_probability column. We would probably have to assign mean values to these users and I don't think that this has the potential to improve out predictions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "341727": "In the Two Sigma Rental Inquiries Competition, manager features, especially manager_skill, drastically improved results. This guy here explains how to do it:\nhttps://www.kaggle.com/den3b81/improve-perfomances-using-manager-features\n\nHas anybody here considered making a feature which is going to represent the marketing ability of each user, something like user_id_skill, which could be computed using 'user_id' and 'deal_probability' columns?",
    "342046": "My guess is that this will be much harder to do in this competition because of the amount of users with very few listings and thus you'll be at high risk of overfitting. Though there still may be merit in this approach. Good luck!",
    "342211": "That's a valid point. I've also found out that there are only 67 000 common users in the test_df and train_df,  and that test_df holds an additional 220 000 new users for which we could not compute the user_skill because of the absence of the deal_probability column. We would probably have to assign mean values to these users and I don't think that this has the potential to improve out predictions."
  },
  "source": "meta"
}