{
  "id": 207148,
  "title": "Conditionnal probabilities using user answers",
  "url": "/competitions/riiid-test-answer-prediction/discussion/207148",
  "author_name": "",
  "post_date": "2020-12-28T13:18:11.993818800Z",
  "votes": 25,
  "comment_count": 15,
  "views": 0,
  "content": "<p>This is a topic that has not been discussed much yet here, but still, I think it could bring a lot of values if used wisely.</p>\n<p>I have the intuition that some questions are very similar to each other, and knowing that a user already answered correctly a question A could probably give a edge on the prediction on answer correctly on question B.</p>\n<p>I calculated the conditionnal probability P(B=1|A=1) and P(B=0|A=0) for each question B given the 500 most frequent questions A, and compared it to the \"difficulty\" feature that we all computed P(B=1).</p>\n<p>This showed me that the probability P(B=1|A=1) can be used sometimes as a nice way to extend the difficulty metric.</p>\n<p>Let's take an example to be a bit more clear, I computed below the probability P(B=1|A=1) and P(B=0|A=1) for B = 6528 and A the top 500 more frequent questions.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fcaf1838591a933baffd2e385b0a89b86%2FSans%20titre.png?generation=1609161066214477&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fe6471e266cf5aad06c9bfc9594f09c81%2FSans%20titre.png?generation=1609161207959498&amp;alt=media\" alt=\"\"></p>\n<p>The graphs above shows that a lot of questions could be used to leverage the difficulty features. It shows as well some unexpecting behaviour for some questions, probably due to the lack of samples available for those questions.</p>\n<p>I personnaly use this has one of my features using 2 threshold system to filter probabilities computed with too few samples and probabilities too close to the general difficulty feature. I then average the contribution of all the questions above the thresholds I set up.</p>\n<p>This method seems to give a nice boost to my score, nethertheless, I feel this is not really optimal and there is probably better things to do, and that's why I would like to share this ideas with you who might have a better mathematical background than me !</p>\n<p>What do you thing of this ? Do you see any idea to improve this ?</p>\n<p>Among the different questions I am asking myself are:</p>\n<ul>\n<li><p>How to determine probabilities that are statistically meaningfull for the problem ? Computing a T-stat is enough ?</p></li>\n<li><p>How to combine different probabilities ? Let's say that out of the last 10 questions, 3 answers give a particular good indication that answer will be 1, but 1 question tend to show the user might answer wrong. Making just an average of those probabilities is probably not the most efficient solution.</p></li>\n<li><p>What \"memory\" shall we consider for computing the probabilities ? For now, I don't consider questions A after a lag of 100 questions, but this is really a random rule of thumb, I am wondering if there is not something more robust. Also, the memory is probably not symetric between P(B=1|A=1) and P(B=1|A=0) given than a user can improve himself faster as he learn new things…</p></li>\n</ul>",
  "messages": [
    {
      "id": "1129616",
      "postDate": "12/28/2020 13:18:11",
      "content": "<p>This is a topic that has not been discussed much yet here, but still, I think it could bring a lot of values if used wisely.</p>\n<p>I have the intuition that some questions are very similar to each other, and knowing that a user already answered correctly a question A could probably give a edge on the prediction on answer correctly on question B.</p>\n<p>I calculated the conditionnal probability P(B=1|A=1) and P(B=0|A=0) for each question B given the 500 most frequent questions A, and compared it to the \"difficulty\" feature that we all computed P(B=1).</p>\n<p>This showed me that the probability P(B=1|A=1) can be used sometimes as a nice way to extend the difficulty metric.</p>\n<p>Let's take an example to be a bit more clear, I computed below the probability P(B=1|A=1) and P(B=0|A=1) for B = 6528 and A the top 500 more frequent questions.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fcaf1838591a933baffd2e385b0a89b86%2FSans%20titre.png?generation=1609161066214477&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fe6471e266cf5aad06c9bfc9594f09c81%2FSans%20titre.png?generation=1609161207959498&amp;alt=media\" alt=\"\"></p>\n<p>The graphs above shows that a lot of questions could be used to leverage the difficulty features. It shows as well some unexpecting behaviour for some questions, probably due to the lack of samples available for those questions.</p>\n<p>I personnaly use this has one of my features using 2 threshold system to filter probabilities computed with too few samples and probabilities too close to the general difficulty feature. I then average the contribution of all the questions above the thresholds I set up.</p>\n<p>This method seems to give a nice boost to my score, nethertheless, I feel this is not really optimal and there is probably better things to do, and that's why I would like to share this ideas with you who might have a better mathematical background than me !</p>\n<p>What do you thing of this ? Do you see any idea to improve this ?</p>\n<p>Among the different questions I am asking myself are:</p>\n<ul>\n<li><p>How to determine probabilities that are statistically meaningfull for the problem ? Computing a T-stat is enough ?</p></li>\n<li><p>How to combine different probabilities ? Let's say that out of the last 10 questions, 3 answers give a particular good indication that answer will be 1, but 1 question tend to show the user might answer wrong. Making just an average of those probabilities is probably not the most efficient solution.</p></li>\n<li><p>What \"memory\" shall we consider for computing the probabilities ? For now, I don't consider questions A after a lag of 100 questions, but this is really a random rule of thumb, I am wondering if there is not something more robust. Also, the memory is probably not symetric between P(B=1|A=1) and P(B=1|A=0) given than a user can improve himself faster as he learn new things…</p></li>\n</ul>",
      "rawMarkdown": "This is a topic that has not been discussed much yet here, but still, I think it could bring a lot of values if used wisely.\n\nI have the intuition that some questions are very similar to each other, and knowing that a user already answered correctly a question A could probably give a edge on the prediction on answer correctly on question B.\n\nI calculated the conditionnal probability P(B=1|A=1) and P(B=0|A=0) for each question B given the 500 most frequent questions A, and compared it to the \"difficulty\" feature that we all computed P(B=1).\n\nThis showed me that the probability P(B=1|A=1) can be used sometimes as a nice way to extend the difficulty metric.\n\nLet's take an example to be a bit more clear, I computed below the probability P(B=1|A=1) and P(B=0|A=1) for B = 6528 and A the top 500 more frequent questions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fcaf1838591a933baffd2e385b0a89b86%2FSans%20titre.png?generation=1609161066214477&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fe6471e266cf5aad06c9bfc9594f09c81%2FSans%20titre.png?generation=1609161207959498&alt=media)\n\nThe graphs above shows that a lot of questions could be used to leverage the difficulty features. It shows as well some unexpecting behaviour for some questions, probably due to the lack of samples available for those questions.\n\nI personnaly use this has one of my features using 2 threshold system to filter probabilities computed with too few samples and probabilities too close to the general difficulty feature. I then average the contribution of all the questions above the thresholds I set up.\n\nThis method seems to give a nice boost to my score, nethertheless, I feel this is not really optimal and there is probably better things to do, and that's why I would like to share this ideas with you who might have a better mathematical background than me !\n\nWhat do you thing of this ? Do you see any idea to improve this ?\n\nAmong the different questions I am asking myself are:\n- How to determine probabilities that are statistically meaningfull for the problem ? Computing a T-stat is enough ?\n- How to combine different probabilities ? Let's say that out of the last 10 questions, 3 answers give a particular good indication that answer will be 1, but 1 question tend to show the user might answer wrong. Making just an average of those probabilities is probably not the most efficient solution.\n\n- What \"memory\" shall we consider for computing the probabilities ? For now, I don't consider questions A after a lag of 100 questions, but this is really a random rule of thumb, I am wondering if there is not something more robust. Also, the memory is probably not symetric between P(B=1|A=1) and P(B=1|A=0) given than a user can improve himself faster as he learn new things...",
      "votes": null
    },
    {
      "id": "1129663",
      "postDate": "12/28/2020 13:55:34",
      "content": "<p>What a great post.</p>\n<p>I had similar idea some time ago but stopped short because:</p>\n<blockquote>\n  <p>I calculated the conditional probability P(B=1|A=1) and P(B=0|A=0) for each question B given the <strong>500 most frequent</strong> questions A</p>\n</blockquote>\n<p>Unlike you, I never thought to restrict to just the most frequent, and the idea of doing an N*N content_id comparison left a bad taste in my mouth, so I did not pursue. It makes sense though to limit to only most frequent questions, both to reduce noise by making the results more statistically sound, as well as to increase the chance the user has actually answered any of them, enabling the feature.</p>\n<p>At the end of the day, I consider this is a sort of clustering. Here you cluster on Bayesian probabilities and then have to choose your thresholds and groupings. Other way of clustering is simply on correlation, and again, you have to choose your thresholds for grouping similar content_ids. One way (I haven't tested) to avoid diluting your existing difficulty feature and to also circumvent some of the issues you posed at the end would be to include cluster association as a categorical variable in your model and have it do the hard lifting, rather than manually trying to figure out how to blend the probabilities in cases where you have a user who's answered 3/4 Q in a group correct, and 1/4 Q in a group incorrectly with different weighting.</p>",
      "rawMarkdown": "What a great post.\n\nI had similar idea some time ago but stopped short because:\n\n> I calculated the conditional probability P(B=1|A=1) and P(B=0|A=0) for each question B given the **500 most frequent** questions A\n\nUnlike you, I never thought to restrict to just the most frequent, and the idea of doing an N*N content_id comparison left a bad taste in my mouth, so I did not pursue. It makes sense though to limit to only most frequent questions, both to reduce noise by making the results more statistically sound, as well as to increase the chance the user has actually answered any of them, enabling the feature.\n\nAt the end of the day, I consider this is a sort of clustering. Here you cluster on Bayesian probabilities and then have to choose your thresholds and groupings. Other way of clustering is simply on correlation, and again, you have to choose your thresholds for grouping similar content_ids. One way (I haven't tested) to avoid diluting your existing difficulty feature and to also circumvent some of the issues you posed at the end would be to include cluster association as a categorical variable in your model and have it do the hard lifting, rather than manually trying to figure out how to blend the probabilities in cases where you have a user who's answered 3/4 Q in a group correct, and 1/4 Q in a group incorrectly with different weighting.",
      "votes": null
    },
    {
      "id": "1129893",
      "postDate": "12/28/2020 16:04:29",
      "content": "<p>Any other relationship between A and B? such as they share the same tags set? </p>",
      "rawMarkdown": "Any other relationship between A and B? such as they share the same tags set?",
      "votes": null
    },
    {
      "id": "1130765",
      "postDate": "12/29/2020 09:49:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>, and thanks for your detailed answer!</p>\n<blockquote>\n  <p>Unlike you, I never thought to restrict to just the most frequent, and the idea of doing an N*N content_id comparison left a bad taste in my mouth, so I did not pursue. It makes sense though to limit to only most frequent questions, both to reduce noise by making the results more statistically sound, as well as to increase the chance the user has actually answered any of them, enabling the feature.</p>\n</blockquote>\n<p>Yes, and it also reduce computationnal time, which is a big plus in that case ! Again, the question is where to put the threshold to select N questions…</p>\n<blockquote>\n  <p>At the end of the day, I consider this is a sort of clustering. Here you cluster on Bayesian probabilities and then have to choose your thresholds and groupings.</p>\n</blockquote>\n<p>For now, I am more trying to make a correction of the \"difficulty\" feature ( P(B=1) ), I take the intersection between the impactfull questions for the current row, and average the probability of all those conditionnal probabilities.</p>\n<p>Nethertheless, clustering is surely a great idea and thats something I'm gonna try out in the next coming days.</p>",
      "rawMarkdown": "Hi @authman, and thanks for your detailed answer!\n\n> Unlike you, I never thought to restrict to just the most frequent, and the idea of doing an N*N content_id comparison left a bad taste in my mouth, so I did not pursue. It makes sense though to limit to only most frequent questions, both to reduce noise by making the results more statistically sound, as well as to increase the chance the user has actually answered any of them, enabling the feature.\n\nYes, and it also reduce computationnal time, which is a big plus in that case ! Again, the question is where to put the threshold to select N questions...\n\n> At the end of the day, I consider this is a sort of clustering. Here you cluster on Bayesian probabilities and then have to choose your thresholds and groupings.\n\nFor now, I am more trying to make a correction of the \"difficulty\" feature ( P(B=1) ), I take the intersection between the impactfull questions for the current row, and average the probability of all those conditionnal probabilities.\n\nNethertheless, clustering is surely a great idea and thats something I'm gonna try out in the next coming days.",
      "votes": null
    },
    {
      "id": "1130779",
      "postDate": "12/29/2020 09:52:33",
      "content": "<p>Hi William, I didn't really dig this, my idea here was more to create a tag/part free metric. It would be actually quite difficult to check such relation as I built a different dictionnary for each question in the format:<br>\n{B1: {(A1,0): P10, (A1,1): P11, (A2,0):P20}, B2: {…}…}</p>",
      "rawMarkdown": "Hi William, I didn't really dig this, my idea here was more to create a tag/part free metric. It would be actually quite difficult to check such relation as I built a different dictionnary for each question in the format:\n{B1: {(A1,0): P10, (A1,1): P11, (A2,0):P20}, B2: {...}...}",
      "votes": null
    },
    {
      "id": "1130992",
      "postDate": "12/29/2020 13:34:48",
      "content": "<p>What I did as a first step is to limit this correlation to questions that lie in the same bundle (because they will be answered one after another. Therefore, you have &lt; 2 x 3 x N probabilities, which is quite manageable. But on the other hand this restriction seems too strict, as using the conditional probabilities sadly had no effect on my score and if both are part of the same evaluation block you won't know whether the user answered the previous question correctly.</p>",
      "rawMarkdown": "What I did as a first step is to limit this correlation to questions that lie in the same bundle (because they will be answered one after another. Therefore, you have < 2 x 3 x N probabilities, which is quite manageable. But on the other hand this restriction seems too strict, as using the conditional probabilities sadly had no effect on my score and if both are part of the same evaluation block you won't know whether the user answered the previous question correctly.",
      "votes": null
    },
    {
      "id": "1131298",
      "postDate": "12/29/2020 16:38:09",
      "content": "<blockquote>\n  <p>What I did as a first step is to limit this correlation to questions that lie in the same bundle</p>\n</blockquote>\n<p>Careful, <a href=\"https://www.kaggle.com/eugenkeil\" target=\"_blank\">@eugenkeil</a>. From the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/data\" target=\"_blank\">Data Description</a>:</p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.</p>\n</blockquote>\n<p>Keep in mind that <code>task_container_id</code> is a user-specific subset of <code>bundle</code>. That means in the dataset, for any given <code>{user_id, task_container_id}</code> tuple, you will only encounter a single <code>bundle_id</code>. So while we can calculate conditional probabilities across all users for a specific bundle, when it comes to inferring with that information (as long as the entire task_container_id is presented at once via the API as opposed to across multiple pulls—and I believe this to be a reasonable assumption since the groupings are small), then you won't have any completed in-bundle question_id priors.</p>",
      "rawMarkdown": "> What I did as a first step is to limit this correlation to questions that lie in the same bundle\n\nCareful, @eugenkeil. From the [Data Description](https://www.kaggle.com/c/riiid-test-answer-prediction/data):\n\n> The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.\n\nKeep in mind that `task_container_id` is a user-specific subset of `bundle`. That means in the dataset, for any given `{user_id, task_container_id}` tuple, you will only encounter a single `bundle_id`. So while we can calculate conditional probabilities across all users for a specific bundle, when it comes to inferring with that information (as long as the entire task_container_id is presented at once via the API as opposed to across multiple pulls—and I believe this to be a reasonable assumption since the groupings are small), then you won't have any completed in-bundle question_id priors.",
      "votes": null
    },
    {
      "id": "1131300",
      "postDate": "12/29/2020 16:38:54",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> , I made some very quick tests adding the clustering feature, and it seems to be improving my priliminary CV by ~0.001.</p>\n<p>edit: And a feature importance analysis place the feature at rank 11/31, so quite impactful !</p>",
      "rawMarkdown": "authman , I made some very quick tests adding the clustering feature, and it seems to be improving my priliminary CV by ~0.001.\n\nedit: And a feature importance analysis place the feature at rank 11/31, so quite impactful !",
      "votes": null
    },
    {
      "id": "1131313",
      "postDate": "12/29/2020 16:45:35",
      "content": "<p><a href=\"https://www.kaggle.com/eugenkeil\" target=\"_blank\">@eugenkeil</a> , I'd like to second <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> here. Please also check <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206279\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206279</a> , item 5, it seems to hold true that during test, bundles are served completely, leaving the correct answers for the next group. I first didn't believe it, but inspecting the file example_test.csv makes it quite obvious.</p>",
      "rawMarkdown": "eugenkeil , I'd like to second @authman here. Please also check https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206279 , item 5, it seems to hold true that during test, bundles are served completely, leaving the correct answers for the next group. I first didn't believe it, but inspecting the file example_test.csv makes it quite obvious.",
      "votes": null
    },
    {
      "id": "1131341",
      "postDate": "12/29/2020 17:06:13",
      "content": "<p>Awesome! Did you do the clustering on the correlations or the conditionals?</p>",
      "rawMarkdown": "Awesome! Did you do the clustering on the correlations or the conditionals?",
      "votes": null
    },
    {
      "id": "1131477",
      "postDate": "12/29/2020 18:39:13",
      "content": "<p>On the conditionnals, but I just did a Kmean using a random number of centroid… There is probably ways to improve this!</p>",
      "rawMarkdown": "On the conditionnals, but I just did a Kmean using a random number of centroid... There is probably ways to improve this!",
      "votes": null
    },
    {
      "id": "1131483",
      "postDate": "12/29/2020 18:46:14",
      "content": "<p>Thanks for sharing. I'll give it a try too with both the conditionals as well as the correlations and will update with results timely.</p>",
      "rawMarkdown": "Thanks for sharing. I'll give it a try too with both the conditionals as well as the correlations and will update with results timely.",
      "votes": null
    },
    {
      "id": "1132598",
      "postDate": "12/30/2020 14:27:41",
      "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> <a href=\"https://www.kaggle.com/stmandl\" target=\"_blank\">@stmandl</a> <br>\nThanks for the warning :-). As I said in the second part of my comment I was already suspecting that something like this might be going on. But good to have a confirmation.</p>\n<p>My second step will be to use only question pairs that share one tag, but I didn't get around to implementing it yet so I cannot say how many of these there are in total and whether it has a measurable impact. I am still struggling with memory issues, so adding more complexity doesn't help and the recent public notebooks decreased my motivation quite bit, when I dropped 700 ranks :-(.</p>",
      "rawMarkdown": "authman @stmandl \nThanks for the warning :-). As I said in the second part of my comment I was already suspecting that something like this might be going on. But good to have a confirmation.\n\nMy second step will be to use only question pairs that share one tag, but I didn't get around to implementing it yet so I cannot say how many of these there are in total and whether it has a measurable impact. I am still struggling with memory issues, so adding more complexity doesn't help and the recent public notebooks decreased my motivation quite bit, when I dropped 700 ranks :-(.",
      "votes": null
    },
    {
      "id": "1132619",
      "postDate": "12/30/2020 14:40:08",
      "content": "<p>you mean cluster tags is useful?  a good news ,may you can try svd, i learn it from pbulic notebook</p>",
      "rawMarkdown": "you mean cluster tags is useful?  a good news ,may you can try svd, i learn it from pbulic notebook",
      "votes": null
    },
    {
      "id": "1132627",
      "postDate": "12/30/2020 14:45:33",
      "content": "<p>Keep going, <a href=\"https://www.kaggle.com/eugenkeil\" target=\"_blank\">@eugenkeil</a>. It's not over until it's over. Just imagine how researchers or people in industry feel when their PhDs and years worth of development are over-turned in an instant due to a recent SoTA development.</p>\n<p>Anything happens in boxing, so keep fighting and you might connect a lucky left hook.</p>",
      "rawMarkdown": "Keep going, @eugenkeil. It's not over until it's over. Just imagine how researchers or people in industry feel when their PhDs and years worth of development are over-turned in an instant due to a recent SoTA development.\n\nAnything happens in boxing, so keep fighting and you might connect a lucky left hook.",
      "votes": null
    },
    {
      "id": "1132639",
      "postDate": "12/30/2020 14:59:13",
      "content": "<p>keep fight,I dropped about 500ranks twice….I'll be back👊</p>",
      "rawMarkdown": "keep fight,I dropped about 500ranks twice....I'll be back👊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1129663,
      "author_name": "authman",
      "author_url": "",
      "post_date": "12/28/2020 13:55:34",
      "content": "<p>What a great post.</p>\n<p>I had similar idea some time ago but stopped short because:</p>\n<blockquote>\n  <p>I calculated the conditional probability P(B=1|A=1) and P(B=0|A=0) for each question B given the <strong>500 most frequent</strong> questions A</p>\n</blockquote>\n<p>Unlike you, I never thought to restrict to just the most frequent, and the idea of doing an N*N content_id comparison left a bad taste in my mouth, so I did not pursue. It makes sense though to limit to only most frequent questions, both to reduce noise by making the results more statistically sound, as well as to increase the chance the user has actually answered any of them, enabling the feature.</p>\n<p>At the end of the day, I consider this is a sort of clustering. Here you cluster on Bayesian probabilities and then have to choose your thresholds and groupings. Other way of clustering is simply on correlation, and again, you have to choose your thresholds for grouping similar content_ids. One way (I haven't tested) to avoid diluting your existing difficulty feature and to also circumvent some of the issues you posed at the end would be to include cluster association as a categorical variable in your model and have it do the hard lifting, rather than manually trying to figure out how to blend the probabilities in cases where you have a user who's answered 3/4 Q in a group correct, and 1/4 Q in a group incorrectly with different weighting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1130765,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/29/2020 09:49:48",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a>, and thanks for your detailed answer!</p>\n<blockquote>\n  <p>Unlike you, I never thought to restrict to just the most frequent, and the idea of doing an N*N content_id comparison left a bad taste in my mouth, so I did not pursue. It makes sense though to limit to only most frequent questions, both to reduce noise by making the results more statistically sound, as well as to increase the chance the user has actually answered any of them, enabling the feature.</p>\n</blockquote>\n<p>Yes, and it also reduce computationnal time, which is a big plus in that case ! Again, the question is where to put the threshold to select N questions…</p>\n<blockquote>\n  <p>At the end of the day, I consider this is a sort of clustering. Here you cluster on Bayesian probabilities and then have to choose your thresholds and groupings.</p>\n</blockquote>\n<p>For now, I am more trying to make a correction of the \"difficulty\" feature ( P(B=1) ), I take the intersection between the impactfull questions for the current row, and average the probability of all those conditionnal probabilities.</p>\n<p>Nethertheless, clustering is surely a great idea and thats something I'm gonna try out in the next coming days.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130992,
          "author_name": "eugenkeil",
          "author_url": "",
          "post_date": "12/29/2020 13:34:48",
          "content": "<p>What I did as a first step is to limit this correlation to questions that lie in the same bundle (because they will be answered one after another. Therefore, you have &lt; 2 x 3 x N probabilities, which is quite manageable. But on the other hand this restriction seems too strict, as using the conditional probabilities sadly had no effect on my score and if both are part of the same evaluation block you won't know whether the user answered the previous question correctly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131298,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/29/2020 16:38:09",
          "content": "<blockquote>\n  <p>What I did as a first step is to limit this correlation to questions that lie in the same bundle</p>\n</blockquote>\n<p>Careful, <a href=\"https://www.kaggle.com/eugenkeil\" target=\"_blank\">@eugenkeil</a>. From the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/data\" target=\"_blank\">Data Description</a>:</p>\n<blockquote>\n  <p>The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.</p>\n</blockquote>\n<p>Keep in mind that <code>task_container_id</code> is a user-specific subset of <code>bundle</code>. That means in the dataset, for any given <code>{user_id, task_container_id}</code> tuple, you will only encounter a single <code>bundle_id</code>. So while we can calculate conditional probabilities across all users for a specific bundle, when it comes to inferring with that information (as long as the entire task_container_id is presented at once via the API as opposed to across multiple pulls—and I believe this to be a reasonable assumption since the groupings are small), then you won't have any completed in-bundle question_id priors.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131300,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/29/2020 16:38:54",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> , I made some very quick tests adding the clustering feature, and it seems to be improving my priliminary CV by ~0.001.</p>\n<p>edit: And a feature importance analysis place the feature at rank 11/31, so quite impactful !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131313,
          "author_name": "stmandl",
          "author_url": "",
          "post_date": "12/29/2020 16:45:35",
          "content": "<p><a href=\"https://www.kaggle.com/eugenkeil\" target=\"_blank\">@eugenkeil</a> , I'd like to second <a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> here. Please also check <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206279\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206279</a> , item 5, it seems to hold true that during test, bundles are served completely, leaving the correct answers for the next group. I first didn't believe it, but inspecting the file example_test.csv makes it quite obvious.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131341,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/29/2020 17:06:13",
          "content": "<p>Awesome! Did you do the clustering on the correlations or the conditionals?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131477,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/29/2020 18:39:13",
          "content": "<p>On the conditionnals, but I just did a Kmean using a random number of centroid… There is probably ways to improve this!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1131483,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/29/2020 18:46:14",
          "content": "<p>Thanks for sharing. I'll give it a try too with both the conditionals as well as the correlations and will update with results timely.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132598,
          "author_name": "eugenkeil",
          "author_url": "",
          "post_date": "12/30/2020 14:27:41",
          "content": "<p><a href=\"https://www.kaggle.com/authman\" target=\"_blank\">@authman</a> <a href=\"https://www.kaggle.com/stmandl\" target=\"_blank\">@stmandl</a> <br>\nThanks for the warning :-). As I said in the second part of my comment I was already suspecting that something like this might be going on. But good to have a confirmation.</p>\n<p>My second step will be to use only question pairs that share one tag, but I didn't get around to implementing it yet so I cannot say how many of these there are in total and whether it has a measurable impact. I am still struggling with memory issues, so adding more complexity doesn't help and the recent public notebooks decreased my motivation quite bit, when I dropped 700 ranks :-(.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132619,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "12/30/2020 14:40:08",
          "content": "<p>you mean cluster tags is useful?  a good news ,may you can try svd, i learn it from pbulic notebook</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132627,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/30/2020 14:45:33",
          "content": "<p>Keep going, <a href=\"https://www.kaggle.com/eugenkeil\" target=\"_blank\">@eugenkeil</a>. It's not over until it's over. Just imagine how researchers or people in industry feel when their PhDs and years worth of development are over-turned in an instant due to a recent SoTA development.</p>\n<p>Anything happens in boxing, so keep fighting and you might connect a lucky left hook.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132639,
          "author_name": "yangxiaoshuai",
          "author_url": "",
          "post_date": "12/30/2020 14:59:13",
          "content": "<p>keep fight,I dropped about 500ranks twice….I'll be back👊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1129893,
      "author_name": "wuwenmin",
      "author_url": "",
      "post_date": "12/28/2020 16:04:29",
      "content": "<p>Any other relationship between A and B? such as they share the same tags set? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1130779,
          "author_name": "bowaka",
          "author_url": "",
          "post_date": "12/29/2020 09:52:33",
          "content": "<p>Hi William, I didn't really dig this, my idea here was more to create a tag/part free metric. It would be actually quite difficult to check such relation as I built a different dictionnary for each question in the format:<br>\n{B1: {(A1,0): P10, (A1,1): P11, (A2,0):P20}, B2: {…}…}</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1129616": "This is a topic that has not been discussed much yet here, but still, I think it could bring a lot of values if used wisely.\n\nI have the intuition that some questions are very similar to each other, and knowing that a user already answered correctly a question A could probably give a edge on the prediction on answer correctly on question B.\n\nI calculated the conditionnal probability P(B=1|A=1) and P(B=0|A=0) for each question B given the 500 most frequent questions A, and compared it to the \"difficulty\" feature that we all computed P(B=1).\n\nThis showed me that the probability P(B=1|A=1) can be used sometimes as a nice way to extend the difficulty metric.\n\nLet's take an example to be a bit more clear, I computed below the probability P(B=1|A=1) and P(B=0|A=1) for B = 6528 and A the top 500 more frequent questions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fcaf1838591a933baffd2e385b0a89b86%2FSans%20titre.png?generation=1609161066214477&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1161354%2Fe6471e266cf5aad06c9bfc9594f09c81%2FSans%20titre.png?generation=1609161207959498&alt=media)\n\nThe graphs above shows that a lot of questions could be used to leverage the difficulty features. It shows as well some unexpecting behaviour for some questions, probably due to the lack of samples available for those questions.\n\nI personnaly use this has one of my features using 2 threshold system to filter probabilities computed with too few samples and probabilities too close to the general difficulty feature. I then average the contribution of all the questions above the thresholds I set up.\n\nThis method seems to give a nice boost to my score, nethertheless, I feel this is not really optimal and there is probably better things to do, and that's why I would like to share this ideas with you who might have a better mathematical background than me !\n\nWhat do you thing of this ? Do you see any idea to improve this ?\n\nAmong the different questions I am asking myself are:\n- How to determine probabilities that are statistically meaningfull for the problem ? Computing a T-stat is enough ?\n- How to combine different probabilities ? Let's say that out of the last 10 questions, 3 answers give a particular good indication that answer will be 1, but 1 question tend to show the user might answer wrong. Making just an average of those probabilities is probably not the most efficient solution.\n\n- What \"memory\" shall we consider for computing the probabilities ? For now, I don't consider questions A after a lag of 100 questions, but this is really a random rule of thumb, I am wondering if there is not something more robust. Also, the memory is probably not symetric between P(B=1|A=1) and P(B=1|A=0) given than a user can improve himself faster as he learn new things...",
    "1129663": "What a great post.\n\nI had similar idea some time ago but stopped short because:\n\n> I calculated the conditional probability P(B=1|A=1) and P(B=0|A=0) for each question B given the **500 most frequent** questions A\n\nUnlike you, I never thought to restrict to just the most frequent, and the idea of doing an N*N content_id comparison left a bad taste in my mouth, so I did not pursue. It makes sense though to limit to only most frequent questions, both to reduce noise by making the results more statistically sound, as well as to increase the chance the user has actually answered any of them, enabling the feature.\n\nAt the end of the day, I consider this is a sort of clustering. Here you cluster on Bayesian probabilities and then have to choose your thresholds and groupings. Other way of clustering is simply on correlation, and again, you have to choose your thresholds for grouping similar content_ids. One way (I haven't tested) to avoid diluting your existing difficulty feature and to also circumvent some of the issues you posed at the end would be to include cluster association as a categorical variable in your model and have it do the hard lifting, rather than manually trying to figure out how to blend the probabilities in cases where you have a user who's answered 3/4 Q in a group correct, and 1/4 Q in a group incorrectly with different weighting.",
    "1129893": "Any other relationship between A and B? such as they share the same tags set?",
    "1130765": "Hi @authman, and thanks for your detailed answer!\n\n> Unlike you, I never thought to restrict to just the most frequent, and the idea of doing an N*N content_id comparison left a bad taste in my mouth, so I did not pursue. It makes sense though to limit to only most frequent questions, both to reduce noise by making the results more statistically sound, as well as to increase the chance the user has actually answered any of them, enabling the feature.\n\nYes, and it also reduce computationnal time, which is a big plus in that case ! Again, the question is where to put the threshold to select N questions...\n\n> At the end of the day, I consider this is a sort of clustering. Here you cluster on Bayesian probabilities and then have to choose your thresholds and groupings.\n\nFor now, I am more trying to make a correction of the \"difficulty\" feature ( P(B=1) ), I take the intersection between the impactfull questions for the current row, and average the probability of all those conditionnal probabilities.\n\nNethertheless, clustering is surely a great idea and thats something I'm gonna try out in the next coming days.",
    "1130779": "Hi William, I didn't really dig this, my idea here was more to create a tag/part free metric. It would be actually quite difficult to check such relation as I built a different dictionnary for each question in the format:\n{B1: {(A1,0): P10, (A1,1): P11, (A2,0):P20}, B2: {...}...}",
    "1130992": "What I did as a first step is to limit this correlation to questions that lie in the same bundle (because they will be answered one after another. Therefore, you have < 2 x 3 x N probabilities, which is quite manageable. But on the other hand this restriction seems too strict, as using the conditional probabilities sadly had no effect on my score and if both are part of the same evaluation block you won't know whether the user answered the previous question correctly.",
    "1131298": "> What I did as a first step is to limit this correlation to questions that lie in the same bundle\n\nCareful, @eugenkeil. From the [Data Description](https://www.kaggle.com/c/riiid-test-answer-prediction/data):\n\n> The API provides user interactions groups in the order in which they occurred. Each group will contain interactions from many different users, but no more than one task_container_id of questions from any single user. Each group has between 1 and 1000 users.\n\nKeep in mind that `task_container_id` is a user-specific subset of `bundle`. That means in the dataset, for any given `{user_id, task_container_id}` tuple, you will only encounter a single `bundle_id`. So while we can calculate conditional probabilities across all users for a specific bundle, when it comes to inferring with that information (as long as the entire task_container_id is presented at once via the API as opposed to across multiple pulls—and I believe this to be a reasonable assumption since the groupings are small), then you won't have any completed in-bundle question_id priors.",
    "1131300": "authman , I made some very quick tests adding the clustering feature, and it seems to be improving my priliminary CV by ~0.001.\n\nedit: And a feature importance analysis place the feature at rank 11/31, so quite impactful !",
    "1131313": "eugenkeil , I'd like to second @authman here. Please also check https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206279 , item 5, it seems to hold true that during test, bundles are served completely, leaving the correct answers for the next group. I first didn't believe it, but inspecting the file example_test.csv makes it quite obvious.",
    "1131341": "Awesome! Did you do the clustering on the correlations or the conditionals?",
    "1131477": "On the conditionnals, but I just did a Kmean using a random number of centroid... There is probably ways to improve this!",
    "1131483": "Thanks for sharing. I'll give it a try too with both the conditionals as well as the correlations and will update with results timely.",
    "1132598": "authman @stmandl \nThanks for the warning :-). As I said in the second part of my comment I was already suspecting that something like this might be going on. But good to have a confirmation.\n\nMy second step will be to use only question pairs that share one tag, but I didn't get around to implementing it yet so I cannot say how many of these there are in total and whether it has a measurable impact. I am still struggling with memory issues, so adding more complexity doesn't help and the recent public notebooks decreased my motivation quite bit, when I dropped 700 ranks :-(.",
    "1132619": "you mean cluster tags is useful?  a good news ,may you can try svd, i learn it from pbulic notebook",
    "1132627": "Keep going, @eugenkeil. It's not over until it's over. Just imagine how researchers or people in industry feel when their PhDs and years worth of development are over-turned in an instant due to a recent SoTA development.\n\nAnything happens in boxing, so keep fighting and you might connect a lucky left hook.",
    "1132639": "keep fight,I dropped about 500ranks twice....I'll be back👊"
  },
  "source": "meta"
}