{
  "id": 189381,
  "title": "Are you going to use the user_id or not?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189381",
  "author_name": "",
  "post_date": "2020-10-07T12:50:01.504204100Z",
  "votes": 13,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I was checking current high score kernels so far, many kernels are using user_id for feature of the prediction model. That’s right, all of user_ids on public test data set are included on the user_id on train dataset. </p>\n<p>However how about private test dataset? I think this is risky. What do you think?</p>\n<p>[Edit]<br>\nI found the user_id on test data that is not contained on train data. Thank you for pointing out, <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> <br>\nEither way, I think user_id should be treated carefully to predict for unseen data.</p>",
  "messages": [
    {
      "id": "1040918",
      "postDate": "10/07/2020 12:50:01",
      "content": "<p>I was checking current high score kernels so far, many kernels are using user_id for feature of the prediction model. That’s right, all of user_ids on public test data set are included on the user_id on train dataset. </p>\n<p>However how about private test dataset? I think this is risky. What do you think?</p>\n<p>[Edit]<br>\nI found the user_id on test data that is not contained on train data. Thank you for pointing out, <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> <br>\nEither way, I think user_id should be treated carefully to predict for unseen data.</p>",
      "rawMarkdown": "I was checking current high score kernels so far, many kernels are using user_id for feature of the prediction model. That’s right, all of user_ids on public test data set are included on the user_id on train dataset. \n\nHowever how about private test dataset? I think this is risky. What do you think?\n\n[Edit]\nI found the user_id on test data that is not contained on train data. Thank you for pointing out, @aquatic \nEither way, I think user_id should be treated carefully to predict for unseen data.",
      "votes": null
    },
    {
      "id": "1041043",
      "postDate": "10/07/2020 14:13:48",
      "content": "<p>I think this is a baseline approach. Later people will use features based on user_id and not user_id itself</p>",
      "rawMarkdown": "I think this is a baseline approach. Later people will use features based on user_id and not user_id itself",
      "votes": null
    },
    {
      "id": "1041073",
      "postDate": "10/07/2020 14:36:09",
      "content": "<p>I think it  a good topic to discuss. In my LGBM model which uses features including <code>user_id</code>, the biggest feature importance is given by <code>user_id</code>. I, as well as you, am somewhat suspicious of it.</p>",
      "rawMarkdown": "I think it  a good topic to discuss. In my LGBM model which uses features including `user_id`, the biggest feature importance is given by `user_id`. I, as well as you, am somewhat suspicious of it.",
      "votes": null
    },
    {
      "id": "1041098",
      "postDate": "10/07/2020 15:02:12",
      "content": "<p>Yes,  I also think these are baseline at really first phase of this competition, however, my concern is if the user_id is useful  for only public test data, not   private test data, in that case public LB will be not helpful to evaluate models. (Should we truct local CV? lol)</p>",
      "rawMarkdown": "Yes,  I also think these are baseline at really first phase of this competition, however, my concern is if the user_id is useful  for only public test data, not   private test data, in that case public LB will be not helpful to evaluate models. (Should we truct local CV? lol)",
      "votes": null
    },
    {
      "id": "1041133",
      "postDate": "10/07/2020 15:20:17",
      "content": "<blockquote>\n  <p>That’s right, all of user_ids on public test data set are included on the user_id on train dataset. </p>\n</blockquote>\n<p>How do you know this -- was it tested or demonstrated somewhere, if you don't mind sharing? </p>\n<p>I wonder if all test users do actually show up in train. In the data description we get:</p>\n<blockquote>\n  <p>Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual</p>\n</blockquote>\n<p>They emphasize the presence of new questions, but not of new users. Why not both if that's really the case? </p>\n<p>Regardless, representing users with other features instead of their <code>user_id</code> is likely a good idea. Gradient boosted trees and neural nets both don't particularly like high cardinality category inputs and id type features seem particularly likely to cause overfitting. It may be much better to figure out how to encode the user behavior patterns that matter as features so the model can focus on more general patterns instead of memorizing users.</p>",
      "rawMarkdown": "> That’s right, all of user_ids on public test data set are included on the user_id on train dataset. \n\nHow do you know this -- was it tested or demonstrated somewhere, if you don't mind sharing? \n\nI wonder if all test users do actually show up in train. In the data description we get:\n\n> Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual\n\nThey emphasize the presence of new questions, but not of new users. Why not both if that's really the case? \n\nRegardless, representing users with other features instead of their `user_id` is likely a good idea. Gradient boosted trees and neural nets both don't particularly like high cardinality category inputs and id type features seem particularly likely to cause overfitting. It may be much better to figure out how to encode the user behavior patterns that matter as features so the model can focus on more general patterns instead of memorizing users.",
      "votes": null
    },
    {
      "id": "1041138",
      "postDate": "10/07/2020 15:22:51",
      "content": "<p>It's likely true that <code>user_id</code> helps a lot to predict on training data, but keep in mind that it's also likely to see spurious feature importances driven by a feature having a lot of distinct values, especially if you're looking at frequency importance. E.g. try using an index column as a feature and see what happens in feature importance.</p>",
      "rawMarkdown": "It's likely true that `user_id` helps a lot to predict on training data, but keep in mind that it's also likely to see spurious feature importances driven by a feature having a lot of distinct values, especially if you're looking at frequency importance. E.g. try using an index column as a feature and see what happens in feature importance.",
      "votes": null
    },
    {
      "id": "1041216",
      "postDate": "10/07/2020 16:11:27",
      "content": "<p>Oh, I'm wrong, I found the user_id on test data that is not contained on train data!<br>\nThank you for pointing out.</p>",
      "rawMarkdown": "Oh, I'm wrong, I found the user_id on test data that is not contained on train data!\nThank you for pointing out.",
      "votes": null
    },
    {
      "id": "1041236",
      "postDate": "10/07/2020 16:23:22",
      "content": "<p>Interesting, thanks for sharing!</p>",
      "rawMarkdown": "Interesting, thanks for sharing!",
      "votes": null
    },
    {
      "id": "1041255",
      "postDate": "10/07/2020 16:42:12",
      "content": "<p>I agree with you in that LGBM feature importance is affected by cardinalities. Thank you for your comments.</p>",
      "rawMarkdown": "I agree with you in that LGBM feature importance is affected by cardinalities. Thank you for your comments.",
      "votes": null
    },
    {
      "id": "1041555",
      "postDate": "10/07/2020 20:06:49",
      "content": "<p>in the example_test.csv, the same user id is seen at different groups. I thought we are supposed to track the user performance as the group number increases. hence the user id is supposed to be used. am I wrong on that?</p>",
      "rawMarkdown": "in the example_test.csv, the same user id is seen at different groups. I thought we are supposed to track the user performance as the group number increases. hence the user id is supposed to be used. am I wrong on that?",
      "votes": null
    },
    {
      "id": "1041866",
      "postDate": "10/08/2020 00:22:29",
      "content": "<p>Well, you can think of this competition as recommending items to users (here, the recommendable items are given to the <em>Agent</em>  at prediction time but the whole items set is fixed). Hence, the goal is to track the agent/content over time and that could hardly be done without using agent/content Ids.</p>\n<p>But, your concerns are legitimate because new contents or new users could appear over time, this is known as <strong>cold start problem</strong> in the the recommendation jargon and there are several techniques to deal with it (content based, user meta data, top items …).</p>",
      "rawMarkdown": "Well, you can think of this competition as recommending items to users (here, the recommendable items are given to the *Agent*  at prediction time but the whole items set is fixed). Hence, the goal is to track the agent/content over time and that could hardly be done without using agent/content Ids.\n\nBut, your concerns are legitimate because new contents or new users could appear over time, this is known as **cold start problem** in the the recommendation jargon and there are several techniques to deal with it (content based, user meta data, top items ...).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1041043,
      "author_name": "artgor",
      "author_url": "",
      "post_date": "10/07/2020 14:13:48",
      "content": "<p>I think this is a baseline approach. Later people will use features based on user_id and not user_id itself</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041098,
          "author_name": "kenmatsu4",
          "author_url": "",
          "post_date": "10/07/2020 15:02:12",
          "content": "<p>Yes,  I also think these are baseline at really first phase of this competition, however, my concern is if the user_id is useful  for only public test data, not   private test data, in that case public LB will be not helpful to evaluate models. (Should we truct local CV? lol)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041073,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "10/07/2020 14:36:09",
      "content": "<p>I think it  a good topic to discuss. In my LGBM model which uses features including <code>user_id</code>, the biggest feature importance is given by <code>user_id</code>. I, as well as you, am somewhat suspicious of it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041138,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/07/2020 15:22:51",
          "content": "<p>It's likely true that <code>user_id</code> helps a lot to predict on training data, but keep in mind that it's also likely to see spurious feature importances driven by a feature having a lot of distinct values, especially if you're looking at frequency importance. E.g. try using an index column as a feature and see what happens in feature importance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041255,
          "author_name": "sishihara",
          "author_url": "",
          "post_date": "10/07/2020 16:42:12",
          "content": "<p>I agree with you in that LGBM feature importance is affected by cardinalities. Thank you for your comments.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041133,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "10/07/2020 15:20:17",
      "content": "<blockquote>\n  <p>That’s right, all of user_ids on public test data set are included on the user_id on train dataset. </p>\n</blockquote>\n<p>How do you know this -- was it tested or demonstrated somewhere, if you don't mind sharing? </p>\n<p>I wonder if all test users do actually show up in train. In the data description we get:</p>\n<blockquote>\n  <p>Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual</p>\n</blockquote>\n<p>They emphasize the presence of new questions, but not of new users. Why not both if that's really the case? </p>\n<p>Regardless, representing users with other features instead of their <code>user_id</code> is likely a good idea. Gradient boosted trees and neural nets both don't particularly like high cardinality category inputs and id type features seem particularly likely to cause overfitting. It may be much better to figure out how to encode the user behavior patterns that matter as features so the model can focus on more general patterns instead of memorizing users.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041216,
          "author_name": "kenmatsu4",
          "author_url": "",
          "post_date": "10/07/2020 16:11:27",
          "content": "<p>Oh, I'm wrong, I found the user_id on test data that is not contained on train data!<br>\nThank you for pointing out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041236,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/07/2020 16:23:22",
          "content": "<p>Interesting, thanks for sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041555,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "10/07/2020 20:06:49",
      "content": "<p>in the example_test.csv, the same user id is seen at different groups. I thought we are supposed to track the user performance as the group number increases. hence the user id is supposed to be used. am I wrong on that?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1041866,
      "author_name": "kneroma",
      "author_url": "",
      "post_date": "10/08/2020 00:22:29",
      "content": "<p>Well, you can think of this competition as recommending items to users (here, the recommendable items are given to the <em>Agent</em>  at prediction time but the whole items set is fixed). Hence, the goal is to track the agent/content over time and that could hardly be done without using agent/content Ids.</p>\n<p>But, your concerns are legitimate because new contents or new users could appear over time, this is known as <strong>cold start problem</strong> in the the recommendation jargon and there are several techniques to deal with it (content based, user meta data, top items …).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1040918": "I was checking current high score kernels so far, many kernels are using user_id for feature of the prediction model. That’s right, all of user_ids on public test data set are included on the user_id on train dataset. \n\nHowever how about private test dataset? I think this is risky. What do you think?\n\n[Edit]\nI found the user_id on test data that is not contained on train data. Thank you for pointing out, @aquatic \nEither way, I think user_id should be treated carefully to predict for unseen data.",
    "1041043": "I think this is a baseline approach. Later people will use features based on user_id and not user_id itself",
    "1041073": "I think it  a good topic to discuss. In my LGBM model which uses features including `user_id`, the biggest feature importance is given by `user_id`. I, as well as you, am somewhat suspicious of it.",
    "1041098": "Yes,  I also think these are baseline at really first phase of this competition, however, my concern is if the user_id is useful  for only public test data, not   private test data, in that case public LB will be not helpful to evaluate models. (Should we truct local CV? lol)",
    "1041133": "> That’s right, all of user_ids on public test data set are included on the user_id on train dataset. \n\nHow do you know this -- was it tested or demonstrated somewhere, if you don't mind sharing? \n\nI wonder if all test users do actually show up in train. In the data description we get:\n\n> Some questions will appear in the hidden test set that have NOT been presented in the train set, emulating the challenge of quickly adapting to modeling newly introduced questions. Their metadata is still in question.csv as usual\n\nThey emphasize the presence of new questions, but not of new users. Why not both if that's really the case? \n\nRegardless, representing users with other features instead of their `user_id` is likely a good idea. Gradient boosted trees and neural nets both don't particularly like high cardinality category inputs and id type features seem particularly likely to cause overfitting. It may be much better to figure out how to encode the user behavior patterns that matter as features so the model can focus on more general patterns instead of memorizing users.",
    "1041138": "It's likely true that `user_id` helps a lot to predict on training data, but keep in mind that it's also likely to see spurious feature importances driven by a feature having a lot of distinct values, especially if you're looking at frequency importance. E.g. try using an index column as a feature and see what happens in feature importance.",
    "1041216": "Oh, I'm wrong, I found the user_id on test data that is not contained on train data!\nThank you for pointing out.",
    "1041236": "Interesting, thanks for sharing!",
    "1041255": "I agree with you in that LGBM feature importance is affected by cardinalities. Thank you for your comments.",
    "1041555": "in the example_test.csv, the same user id is seen at different groups. I thought we are supposed to track the user performance as the group number increases. hence the user id is supposed to be used. am I wrong on that?",
    "1041866": "Well, you can think of this competition as recommending items to users (here, the recommendable items are given to the *Agent*  at prediction time but the whole items set is fixed). Hence, the goal is to track the agent/content over time and that could hardly be done without using agent/content Ids.\n\nBut, your concerns are legitimate because new contents or new users could appear over time, this is known as **cold start problem** in the the recommendation jargon and there are several techniques to deal with it (content based, user meta data, top items ...)."
  },
  "source": "meta"
}