{
  "id": 197349,
  "title": "Question tags - what information can we get out of them?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/197349",
  "author_name": "",
  "post_date": "2020-11-15T20:49:39.526723600Z",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I just investigated the question tags a bit, see my kernel</p>\n<p><a href=\"https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags</a></p>\n<p>There are two clear tag clusters, a bunch of tags which aren't connected to anything and then there's tag 162 which seems to be the center of it all.</p>\n<p>Not sure what to make of it honestly. Any ideas?</p>",
  "messages": [
    {
      "id": "1079277",
      "postDate": "11/15/2020 20:49:39",
      "content": "<p>I just investigated the question tags a bit, see my kernel</p>\n<p><a href=\"https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags</a></p>\n<p>There are two clear tag clusters, a bunch of tags which aren't connected to anything and then there's tag 162 which seems to be the center of it all.</p>\n<p>Not sure what to make of it honestly. Any ideas?</p>",
      "rawMarkdown": "I just investigated the question tags a bit, see my kernel\n\n[https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags](https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags)\n\nThere are two clear tag clusters, a bunch of tags which aren't connected to anything and then there's tag 162 which seems to be the center of it all.\n\nNot sure what to make of it honestly. Any ideas?",
      "votes": null
    },
    {
      "id": "1082247",
      "postDate": "11/17/2020 18:24:09",
      "content": "<p>Nice analysis and visualization, although I'm not sure what we can do with this either lol. The question tags from what the data page said are supposedly a list of labels that can describe a question, so at the end of the day it gives an representation of the question content_id, then what can we do when we get an representation of the content_id? </p>\n<p>I guess it could be helpful to identify what kind of question/content_id the user is good at or bad at, which could provide some predictive power on the next question, but one thing I was wondering, did  the user pick what questions to work on or the questions were generated by their system in some heuristic ways (e.g. if the user has been doing very well in certain kind of questions, then the system would stop showing similar questions and move onto something else)</p>\n<p>I tried using the record of these questions tags in a binary represented way, e.g. question tag [1, 4, 99] -&gt; [0, 1, 0, 0, 1, …, 1, …0], then if user answers the question correctly then the value stays at 1 (since if the user gets the question correctly then supposedly the user knows all the tags well), and if user answers (single tag) question incorrectly, the binary value gets turned into -1, then for upcoming question, we can compute how many tags the user has seen, has gotten correctly and incorrectly. (an issue is when user gets a question wrong, its hard to tell if the user doesn't know all the tags or just one of the tag gave the user trouble). This does show up in my top feature importance but still nothing compares to target encoded feature or content_id when used at categorical value , and due to its engineering complexity, i stopped doing this and moved onto other method. (also, I saw there were cases where user got one question correct but then the next question with the exactly same tags the user got wrong….)</p>",
      "rawMarkdown": "Nice analysis and visualization, although I'm not sure what we can do with this either lol. The question tags from what the data page said are supposedly a list of labels that can describe a question, so at the end of the day it gives an representation of the question content_id, then what can we do when we get an representation of the content_id? \n\nI guess it could be helpful to identify what kind of question/content_id the user is good at or bad at, which could provide some predictive power on the next question, but one thing I was wondering, did  the user pick what questions to work on or the questions were generated by their system in some heuristic ways (e.g. if the user has been doing very well in certain kind of questions, then the system would stop showing similar questions and move onto something else)\n\nI tried using the record of these questions tags in a binary represented way, e.g. question tag [1, 4, 99] -> [0, 1, 0, 0, 1, ..., 1, ...0], then if user answers the question correctly then the value stays at 1 (since if the user gets the question correctly then supposedly the user knows all the tags well), and if user answers (single tag) question incorrectly, the binary value gets turned into -1, then for upcoming question, we can compute how many tags the user has seen, has gotten correctly and incorrectly. (an issue is when user gets a question wrong, its hard to tell if the user doesn't know all the tags or just one of the tag gave the user trouble). This does show up in my top feature importance but still nothing compares to target encoded feature or content\\_id when used at categorical value , and due to its engineering complexity, i stopped doing this and moved onto other method. (also, I saw there were cases where user got one question correct but then the next question with the exactly same tags the user got wrong....)",
      "votes": null
    },
    {
      "id": "1082259",
      "postDate": "11/17/2020 18:32:28",
      "content": "<p>Different question can have same set of tags. Just like you would expect them in an examination. e.g. Different levels of Algebraic Questions.</p>\n<blockquote>\n  <p>id the user pick what questions to work on or the questions were generated by their system in some heuristic ways </p>\n</blockquote>\n<p>It's decided by a system.</p>",
      "rawMarkdown": "Different question can have same set of tags. Just like you would expect them in an examination. e.g. Different levels of Algebraic Questions.\n\n>id the user pick what questions to work on or the questions were generated by their system in some heuristic ways \n\nIt's decided by a system.",
      "votes": null
    },
    {
      "id": "1082272",
      "postDate": "11/17/2020 18:44:21",
      "content": "<p>do we know how its decided?</p>",
      "rawMarkdown": "do we know how its decided?",
      "votes": null
    },
    {
      "id": "1082273",
      "postDate": "11/17/2020 18:45:48",
      "content": "<p>There's a paper by Riiid which states/clearly hints that it's decided by some system but we don't know the details.</p>",
      "rawMarkdown": "There's a paper by Riiid which states/clearly hints that it's decided by some system but we don't know the details.",
      "votes": null
    },
    {
      "id": "1082848",
      "postDate": "11/18/2020 09:46:22",
      "content": "<p>I also started using information on which question has been seen by each user as well as tag-specific user correctness, but somehow it didn't really increase my score…<br>\nAdding cluster-specific user correctness made absolutely no difference for me.</p>",
      "rawMarkdown": "I also started using information on which question has been seen by each user as well as tag-specific user correctness, but somehow it didn't really increase my score...\nAdding cluster-specific user correctness made absolutely no difference for me.",
      "votes": null
    },
    {
      "id": "1082963",
      "postDate": "11/18/2020 12:49:44",
      "content": "<p>It did help me though; I constructed it <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347\" target=\"_blank\">this</a> way. There's a working implementation below those comments in the same thread as well.</p>",
      "rawMarkdown": "It did help me though; I constructed it [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347) way. There's a working implementation below those comments in the same thread as well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1082247,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "11/17/2020 18:24:09",
      "content": "<p>Nice analysis and visualization, although I'm not sure what we can do with this either lol. The question tags from what the data page said are supposedly a list of labels that can describe a question, so at the end of the day it gives an representation of the question content_id, then what can we do when we get an representation of the content_id? </p>\n<p>I guess it could be helpful to identify what kind of question/content_id the user is good at or bad at, which could provide some predictive power on the next question, but one thing I was wondering, did  the user pick what questions to work on or the questions were generated by their system in some heuristic ways (e.g. if the user has been doing very well in certain kind of questions, then the system would stop showing similar questions and move onto something else)</p>\n<p>I tried using the record of these questions tags in a binary represented way, e.g. question tag [1, 4, 99] -&gt; [0, 1, 0, 0, 1, …, 1, …0], then if user answers the question correctly then the value stays at 1 (since if the user gets the question correctly then supposedly the user knows all the tags well), and if user answers (single tag) question incorrectly, the binary value gets turned into -1, then for upcoming question, we can compute how many tags the user has seen, has gotten correctly and incorrectly. (an issue is when user gets a question wrong, its hard to tell if the user doesn't know all the tags or just one of the tag gave the user trouble). This does show up in my top feature importance but still nothing compares to target encoded feature or content_id when used at categorical value , and due to its engineering complexity, i stopped doing this and moved onto other method. (also, I saw there were cases where user got one question correct but then the next question with the exactly same tags the user got wrong….)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1082259,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/17/2020 18:32:28",
          "content": "<p>Different question can have same set of tags. Just like you would expect them in an examination. e.g. Different levels of Algebraic Questions.</p>\n<blockquote>\n  <p>id the user pick what questions to work on or the questions were generated by their system in some heuristic ways </p>\n</blockquote>\n<p>It's decided by a system.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1082272,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "11/17/2020 18:44:21",
          "content": "<p>do we know how its decided?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1082273,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/17/2020 18:45:48",
          "content": "<p>There's a paper by Riiid which states/clearly hints that it's decided by some system but we don't know the details.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1082848,
          "author_name": "spacelx",
          "author_url": "",
          "post_date": "11/18/2020 09:46:22",
          "content": "<p>I also started using information on which question has been seen by each user as well as tag-specific user correctness, but somehow it didn't really increase my score…<br>\nAdding cluster-specific user correctness made absolutely no difference for me.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1082963,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "11/18/2020 12:49:44",
          "content": "<p>It did help me though; I constructed it <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347\" target=\"_blank\">this</a> way. There's a working implementation below those comments in the same thread as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1079277": "I just investigated the question tags a bit, see my kernel\n\n[https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags](https://www.kaggle.com/spacelx/2020-r3id-clustering-question-tags)\n\nThere are two clear tag clusters, a bunch of tags which aren't connected to anything and then there's tag 162 which seems to be the center of it all.\n\nNot sure what to make of it honestly. Any ideas?",
    "1082247": "Nice analysis and visualization, although I'm not sure what we can do with this either lol. The question tags from what the data page said are supposedly a list of labels that can describe a question, so at the end of the day it gives an representation of the question content_id, then what can we do when we get an representation of the content_id? \n\nI guess it could be helpful to identify what kind of question/content_id the user is good at or bad at, which could provide some predictive power on the next question, but one thing I was wondering, did  the user pick what questions to work on or the questions were generated by their system in some heuristic ways (e.g. if the user has been doing very well in certain kind of questions, then the system would stop showing similar questions and move onto something else)\n\nI tried using the record of these questions tags in a binary represented way, e.g. question tag [1, 4, 99] -> [0, 1, 0, 0, 1, ..., 1, ...0], then if user answers the question correctly then the value stays at 1 (since if the user gets the question correctly then supposedly the user knows all the tags well), and if user answers (single tag) question incorrectly, the binary value gets turned into -1, then for upcoming question, we can compute how many tags the user has seen, has gotten correctly and incorrectly. (an issue is when user gets a question wrong, its hard to tell if the user doesn't know all the tags or just one of the tag gave the user trouble). This does show up in my top feature importance but still nothing compares to target encoded feature or content\\_id when used at categorical value , and due to its engineering complexity, i stopped doing this and moved onto other method. (also, I saw there were cases where user got one question correct but then the next question with the exactly same tags the user got wrong....)",
    "1082259": "Different question can have same set of tags. Just like you would expect them in an examination. e.g. Different levels of Algebraic Questions.\n\n>id the user pick what questions to work on or the questions were generated by their system in some heuristic ways \n\nIt's decided by a system.",
    "1082272": "do we know how its decided?",
    "1082273": "There's a paper by Riiid which states/clearly hints that it's decided by some system but we don't know the details.",
    "1082848": "I also started using information on which question has been seen by each user as well as tag-specific user correctness, but somehow it didn't really increase my score...\nAdding cluster-specific user correctness made absolutely no difference for me.",
    "1082963": "It did help me though; I constructed it [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194266#1069347) way. There's a working implementation below those comments in the same thread as well."
  },
  "source": "meta"
}