{
  "id": 211451,
  "title": "Riiid with no NNs, LGBMs, or CV (0.763 private)",
  "url": "/competitions/riiid-test-answer-prediction/discussion/211451",
  "author_name": "Yana Malysheva",
  "post_date": "2021-01-15T07:46:03.085000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, a couple of notes:</p>\n<ul>\n<li>No, I wasn't explictly <em>trying</em> to avoid NNs or LGBMs or doing anything the \"normal\" way. I just kind of prioritized what was interesting to me. Actually, I did try throwing some of the data in an LGBM at the last moment, but it did much <em>worse</em> than my more \"analytical\" solution.<ul>\n<li>I was more interested in regression/clustering kind of stuff because I was hoping to learn something from the data about how/whether people actually learn through the app and such. I think there are some interesting points to discuss there,  which I have now done <a href=\"https://www.kaggle.com/yanamal/riiid-dataset-what-can-we-learn-about-learning\" target=\"_blank\">here</a>.</li></ul></li>\n<li>My submissions consistently did slightly <em>better</em> on the private data set than the public leaderboard, which lends credence to the idea of not needing cross-validation (in this particular case), and using more analytical methods like the Bayesian Information Criterion instead.</li>\n<li>I'm happy to publish the whole actual pipeline of notebooks I used, too, if anyone's actually interested. Though I'm not sure how readable they are.</li>\n</ul>\n<h1>Overview</h1>\n<p>The model consists of 3 very simple (mathematically) building blocks:</p>\n<ol>\n<li><strong>Logistic regression</strong> which seeks to capture how different users seem to learn different concepts, based on Learning Factor Analysis/Performance Factor Analysis (see also my notebook discussing LFA <a href=\"https://www.kaggle.com/yanamal/learning-factor-analysis-are-tags-skills\" target=\"_blank\">here</a>)</li>\n<li><strong>Hierarchical clustering</strong> of tag content (questions and lectures) into \"sub-tags\", using the <a href=\"https://en.wikipedia.org/wiki/Bayesian_information_criterion\" target=\"_blank\">Baysian Information Criterion</a> to decide when and whether to split</li>\n<li>Per-question and per-user <strong>residuals</strong>, as demonstrated in <a href=\"https://www.kaggle.com/rar4dx/two-feature-model\" target=\"_blank\">this brilliantly simple notebook</a> by <a href=\"https://www.kaggle.com/rar4dx\" target=\"_blank\">@rar4dx</a> </li>\n</ol>\n<h1>Logistic regression of learning curves</h1>\n<p>Learning Factor Analysis tries to determine the odds of making an error based on the number of encounters the student has had with some relevant concept so far. Performance Factor Analysis is similar, but measures the odds of making an error as a function of number of correct answers and number of wrong answers so far. Both use the same general kind of power function:</p>\n<p>$$ \\mathrm{odds}(\\mathrm{error}) = \\alpha \\prod_i \\mathrm{count_i}^{\\gamma_i}$$</p>\n<p>where </p>\n<ul>\n<li>each <code>i</code> is some factor being taken into account (encounters in LFA, correct/wrong answers in PFA)</li>\n<li><code>alpha</code> and <code>gamma_i</code> are some parameters which define the \"shape\" of the the learning curve - the parameters to fit with ML</li>\n<li><code>count_i</code> is the count for that particular factor, for some student at some moment in time(e.g. \"7 encounters with this concept\" or \"4 correct answers\"</li>\n</ul>\n<p>In practice, this amounts to doing logistic regression with the <strong>logarithms</strong> of your <code>count_i</code>s as input. (Taking the log of both sides of the equation so far gives you a linear equation to find the log-odds, which is exactly what logistic regression does). The intercept and coefficients from your logistic regression are <code>log(alpha)</code> and <code>gamma_i</code>s in the equation above.</p>\n<p><em>Side note: The papers on LFA/PFA for some reason never explicitly mention that they take the log of the counts. Not sure if they neglected to take the log, or to mention it. But the evidence from studying real-world learning curves points to the specific power function above, so the log would be necessary. And empirically, it did work better with the log.</em></p>\n<p>I took this power law idea even further than PFA, and used five different counters for each \"concept\" (tag or, after the hierarchical clustering, sub-tag):</p>\n<ul>\n<li>Number of lectures viewed for the concept</li>\n<li>Number of \"diagnostic\" questions answered correctly for the concept</li>\n<li>Number of \"diagnostic\" questions answered incorrectly for the concept</li>\n<li>Number of \"feedback\" questions answered correctly for the concept</li>\n<li>Number of \"feedback\" questions answered incorrectly for the concept</li>\n</ul>\n<p>The diagnostic/feedback split was done based on whether that particular question instance had an explanation (given by <code>prior_question_had_explanation</code> and then painstakingly shifted back to the question it belongs to)</p>\n<p>I did also try combinations with fewer parameters (e.g. ignoring the feedback/diagnostic split; ignoring  lectures). But the Bayesian Information  Criterion (discussed in more detail in the next section) nearly always turned out better for the full split (except on one or two tags), so I went with that.</p>\n<p>There's the slight wrinkle that lots of questions actually have <em>several</em> tags, and so several predictions from each tag. For those, I found the most effective thing was to average the log-odds predictions for each tag and then use that as the log-odds value (which is conceptually equivalent to doing a geometric average of the odds).</p>\n<p>I trained each regression on all available training data for that tag/subtag, and only ever evaluated the output on the training data itself until I actually submitted (The AUC tended to be about 0.02 better on the training data than the test data).</p>\n<p><em>Side note: I also did a test run of a single regression over</em> all <em>the tags - or at least all the ones that actually intersected with each other - but it wouldn't converge, even with just two dimensions (correct/wrong). Surprisingly, even the unconverged version tended to do a tiny bit better than the equivalent per-tag average (also only correct/wrong, no sub-tag clustering). I say this is surprising because a single regression has a single intercept, which means that all users answering their very first question were lumped in the same exact group, and generally, it should be worse at distinguishing between beginner users.</em></p>\n<h1>Hierarchical Clustering</h1>\n<p>The goodness-of-fit for the regression varied quite a bit from tag to tag. So I decided to split the questions and lectures further into \"sub-tags\". </p>\n<p>I determined the \"similarity\" of questions within the same tag like this:</p>\n<ol>\n<li>Make a User x Question matrix; each cell in the matrix is the <em>average correctness</em> for how that user answered that question (1 = correct every time; 0 = wrong every time; 0.5 = half and half; etc.). This matrix is pretty sparse,<ul>\n<li>The lectures are also in this matrix, alongside the questions: the value is 1 if that user has seen that lecture.</li></ul></li>\n<li>Use that matrix to get a Question x Question (well, Content x Content) correlation matrix with pandas.corr and a similarity function I called \"Jacardish\" -Jacard similarity, but can also deal with non-binary values. (I think it basically ended up being a normalized manhattan distance).</li>\n<li>Use the similarity metric <em>again</em> on the Content x Content matrix to get a sort of \"similarity similarity\" measure - questions that are similar to the same questions, and different from the same questions, are considered more \"similarly-similar\" to each other.<ul>\n<li>This was kind of just  a happy accident. I stumbled upon this idea by reading some random blog post about how to make correlation matrices more readable by doing this kind of correlation on a correlation matrix. At first I thought it was kind of silly, but empirically, it produced better results - more balanced hierarchical trees in the next step</li></ul></li>\n</ol>\n<p>After that, I used <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.cluster.hierarchy.linkage.html\" target=\"_blank\">scipy's hierarchical clustering library</a> to derive a <a href=\"https://en.wikipedia.org/wiki/Hierarchical_clustering\" target=\"_blank\">hierachical tree</a> where each leaf is  a question/lecture, and they are progressively joined into clusters of similar questions.</p>\n<p>Now I could  split each tag's content into sub-tags by splitting the hierarchical tree at the root. For each tree, I decided whether to split it by re-doing the logistic regression from step one with the split, and comparing it to the quality of the regression without the split. Specifically, I calculated the <a href=\"https://en.wikipedia.org/wiki/Bayesian_information_criterion\" target=\"_blank\">Baysian Information Criterion (BIC)</a> for each version, and chose to split if the BIC was smaller. The BIC takes into account the log-likelihood of the model (log of the chance of seeing the exact data we see, if the model is right), the number of parameters in the model, and the (log of) number of total data points. In general, a more complex model is more prone to overfitting, since there are more \"knobs\" one can turn to fit the training data precisely. But if the BIC for a more complex model is smaller than the simpler version, then the increase in complexity is probably justified by the increase in how well the model explains the data. </p>\n<p>Even though the BIC  is considered to be pretty harsh and pessimistic about model complexity, it happily split the tags down into over 2500 subtags, sometimes with just one question in a \"sub-tag\". I considered dialing it back artificially a bit, by introducing a bigger threshold for splitting. But in the end, I trusted the math and it seems to have  worked out. </p>\n<p><em>Side note: I also tried an alternative method for choosing what splits to attempt: instead of doing the hierarchical clustering step, for bigger tags, I tried splitting out questions based on whether they had</em> some other tag <em>. However, this was much worse than the hierarchical clustering: I found no splits that actually reduced the BIC, whereas the hierarchical clustering went a bit crazy with the splitting, if anything. So it seems like tags might not actually be a great way of grouping questions together for the purpose of analyzing learning curves.</em></p>\n<h1>Residuals</h1>\n<p>This is quite a simple idea:</p>\n<ul>\n<li>For each data point, generate a prediction (probability of user answering correctly)</li>\n<li>Subtract that prediction from what actually happened (0 = incorrect or 1 = correct)</li>\n<li>average these residuals:<ul>\n<li>across each question</li>\n<li>across each user</li></ul></li>\n<li>When doing inference, for each data point, generate a prediction as normal; then add the per-user and per-question residuals to the predicted probability, and report that</li>\n<li>Update the running user averages with each new batch of dataa (I didn't bother updating the per-question one).</li>\n</ul>\n<p>It's a bit scary how well this works. I went from 0.728 with no residuals to 0.763 with both per-user and per-question residuals.</p>\n<h1>Things I didn't try</h1>\n<p>Lots of things I wanted to add, but ran out of time:</p>\n<ul>\n<li>Evaluate provided tags: As I mentioned above, the tags may not have actually been a good way of splitting the data for learning curve analysis. There's lots of interesting techniques in the literature, e.g. Q-matrix, rule space analysis, etc. that I just didn't get to. Part of the problem is that all the literature kind of assumes that I know a lot more about the nature of the problems/concepts/tags than just an integer label.</li>\n<li>clustering incorrect answers into \"misconceptions\": the same way that I clustered questions based on correctness, I could have clustered specific incorrect answers as seemingly related misconceptions</li>\n<li>incorporate question elapsed time: either into the logistic regression, or as its own separate predictor.</li>\n<li>incorporate timestamps (e.g. large gaps in time)</li>\n<li>dig into multiple-tag regression quirks: That regression assigned some interesting values to the \"learning curve\" slope of a lot of tags. Specifically, it seemed to treat a <em>lot</em> of tags as signals that the student is less likely too answer questions correctly if they've had a lot of \"experience\" with a tag.</li>\n<li>isolate specific types of 'bad' users: I think both the logistic regression and the clustering actually did a lot of that implicitly (e.g. capturing that users who seem to guess randomly will probably continue to guess randomly; or isolating groups of questions which seem to be given repeatedly to \"remedial\" users). But more can probably be done explcitly.</li>\n</ul>",
  "messages": [
    {
      "id": 1153854,
      "postDate": "2021-01-15T07:46:03.087Z",
      "content": "<p>First, a couple of notes:</p>\n<ul>\n<li>No, I wasn't explictly <em>trying</em> to avoid NNs or LGBMs or doing anything the \"normal\" way. I just kind of prioritized what was interesting to me. Actually, I did try throwing some of the data in an LGBM at the last moment, but it did much <em>worse</em> than my more \"analytical\" solution.<ul>\n<li>I was more interested in regression/clustering kind of stuff because I was hoping to learn something from the data about how/whether people actually learn through the app and such. I think there are some interesting points to discuss there,  which I have now done <a href=\"https://www.kaggle.com/yanamal/riiid-dataset-what-can-we-learn-about-learning\" target=\"_blank\">here</a>.</li></ul></li>\n<li>My submissions consistently did slightly <em>better</em> on the private data set than the public leaderboard, which lends credence to the idea of not needing cross-validation (in this particular case), and using more analytical methods like the Bayesian Information Criterion instead.</li>\n<li>I'm happy to publish the whole actual pipeline of notebooks I used, too, if anyone's actually interested. Though I'm not sure how readable they are.</li>\n</ul>\n<h1>Overview</h1>\n<p>The model consists of 3 very simple (mathematically) building blocks:</p>\n<ol>\n<li><strong>Logistic regression</strong> which seeks to capture how different users seem to learn different concepts, based on Learning Factor Analysis/Performance Factor Analysis (see also my notebook discussing LFA <a href=\"https://www.kaggle.com/yanamal/learning-factor-analysis-are-tags-skills\" target=\"_blank\">here</a>)</li>\n<li><strong>Hierarchical clustering</strong> of tag content (questions and lectures) into \"sub-tags\", using the <a href=\"https://en.wikipedia.org/wiki/Bayesian_information_criterion\" target=\"_blank\">Baysian Information Criterion</a> to decide when and whether to split</li>\n<li>Per-question and per-user <strong>residuals</strong>, as demonstrated in <a href=\"https://www.kaggle.com/rar4dx/two-feature-model\" target=\"_blank\">this brilliantly simple notebook</a> by <a href=\"https://www.kaggle.com/rar4dx\" target=\"_blank\">@rar4dx</a> </li>\n</ol>\n<h1>Logistic regression of learning curves</h1>\n<p>Learning Factor Analysis tries to determine the odds of making an error based on the number of encounters the student has had with some relevant concept so far. Performance Factor Analysis is similar, but measures the odds of making an error as a function of number of correct answers and number of wrong answers so far. Both use the same general kind of power function:</p>\n<p>$$ \\mathrm{odds}(\\mathrm{error}) = \\alpha \\prod_i \\mathrm{count_i}^{\\gamma_i}$$</p>\n<p>where </p>\n<ul>\n<li>each <code>i</code> is some factor being taken into account (encounters in LFA, correct/wrong answers in PFA)</li>\n<li><code>alpha</code> and <code>gamma_i</code> are some parameters which define the \"shape\" of the the learning curve - the parameters to fit with ML</li>\n<li><code>count_i</code> is the count for that particular factor, for some student at some moment in time(e.g. \"7 encounters with this concept\" or \"4 correct answers\"</li>\n</ul>\n<p>In practice, this amounts to doing logistic regression with the <strong>logarithms</strong> of your <code>count_i</code>s as input. (Taking the log of both sides of the equation so far gives you a linear equation to find the log-odds, which is exactly what logistic regression does). The intercept and coefficients from your logistic regression are <code>log(alpha)</code> and <code>gamma_i</code>s in the equation above.</p>\n<p><em>Side note: The papers on LFA/PFA for some reason never explicitly mention that they take the log of the counts. Not sure if they neglected to take the log, or to mention it. But the evidence from studying real-world learning curves points to the specific power function above, so the log would be necessary. And empirically, it did work better with the log.</em></p>\n<p>I took this power law idea even further than PFA, and used five different counters for each \"concept\" (tag or, after the hierarchical clustering, sub-tag):</p>\n<ul>\n<li>Number of lectures viewed for the concept</li>\n<li>Number of \"diagnostic\" questions answered correctly for the concept</li>\n<li>Number of \"diagnostic\" questions answered incorrectly for the concept</li>\n<li>Number of \"feedback\" questions answered correctly for the concept</li>\n<li>Number of \"feedback\" questions answered incorrectly for the concept</li>\n</ul>\n<p>The diagnostic/feedback split was done based on whether that particular question instance had an explanation (given by <code>prior_question_had_explanation</code> and then painstakingly shifted back to the question it belongs to)</p>\n<p>I did also try combinations with fewer parameters (e.g. ignoring the feedback/diagnostic split; ignoring  lectures). But the Bayesian Information  Criterion (discussed in more detail in the next section) nearly always turned out better for the full split (except on one or two tags), so I went with that.</p>\n<p>There's the slight wrinkle that lots of questions actually have <em>several</em> tags, and so several predictions from each tag. For those, I found the most effective thing was to average the log-odds predictions for each tag and then use that as the log-odds value (which is conceptually equivalent to doing a geometric average of the odds).</p>\n<p>I trained each regression on all available training data for that tag/subtag, and only ever evaluated the output on the training data itself until I actually submitted (The AUC tended to be about 0.02 better on the training data than the test data).</p>\n<p><em>Side note: I also did a test run of a single regression over</em> all <em>the tags - or at least all the ones that actually intersected with each other - but it wouldn't converge, even with just two dimensions (correct/wrong). Surprisingly, even the unconverged version tended to do a tiny bit better than the equivalent per-tag average (also only correct/wrong, no sub-tag clustering). I say this is surprising because a single regression has a single intercept, which means that all users answering their very first question were lumped in the same exact group, and generally, it should be worse at distinguishing between beginner users.</em></p>\n<h1>Hierarchical Clustering</h1>\n<p>The goodness-of-fit for the regression varied quite a bit from tag to tag. So I decided to split the questions and lectures further into \"sub-tags\". </p>\n<p>I determined the \"similarity\" of questions within the same tag like this:</p>\n<ol>\n<li>Make a User x Question matrix; each cell in the matrix is the <em>average correctness</em> for how that user answered that question (1 = correct every time; 0 = wrong every time; 0.5 = half and half; etc.). This matrix is pretty sparse,<ul>\n<li>The lectures are also in this matrix, alongside the questions: the value is 1 if that user has seen that lecture.</li></ul></li>\n<li>Use that matrix to get a Question x Question (well, Content x Content) correlation matrix with pandas.corr and a similarity function I called \"Jacardish\" -Jacard similarity, but can also deal with non-binary values. (I think it basically ended up being a normalized manhattan distance).</li>\n<li>Use the similarity metric <em>again</em> on the Content x Content matrix to get a sort of \"similarity similarity\" measure - questions that are similar to the same questions, and different from the same questions, are considered more \"similarly-similar\" to each other.<ul>\n<li>This was kind of just  a happy accident. I stumbled upon this idea by reading some random blog post about how to make correlation matrices more readable by doing this kind of correlation on a correlation matrix. At first I thought it was kind of silly, but empirically, it produced better results - more balanced hierarchical trees in the next step</li></ul></li>\n</ol>\n<p>After that, I used <a href=\"https://docs.scipy.org/doc/scipy/reference/generated/scipy.cluster.hierarchy.linkage.html\" target=\"_blank\">scipy's hierarchical clustering library</a> to derive a <a href=\"https://en.wikipedia.org/wiki/Hierarchical_clustering\" target=\"_blank\">hierachical tree</a> where each leaf is  a question/lecture, and they are progressively joined into clusters of similar questions.</p>\n<p>Now I could  split each tag's content into sub-tags by splitting the hierarchical tree at the root. For each tree, I decided whether to split it by re-doing the logistic regression from step one with the split, and comparing it to the quality of the regression without the split. Specifically, I calculated the <a href=\"https://en.wikipedia.org/wiki/Bayesian_information_criterion\" target=\"_blank\">Baysian Information Criterion (BIC)</a> for each version, and chose to split if the BIC was smaller. The BIC takes into account the log-likelihood of the model (log of the chance of seeing the exact data we see, if the model is right), the number of parameters in the model, and the (log of) number of total data points. In general, a more complex model is more prone to overfitting, since there are more \"knobs\" one can turn to fit the training data precisely. But if the BIC for a more complex model is smaller than the simpler version, then the increase in complexity is probably justified by the increase in how well the model explains the data. </p>\n<p>Even though the BIC  is considered to be pretty harsh and pessimistic about model complexity, it happily split the tags down into over 2500 subtags, sometimes with just one question in a \"sub-tag\". I considered dialing it back artificially a bit, by introducing a bigger threshold for splitting. But in the end, I trusted the math and it seems to have  worked out. </p>\n<p><em>Side note: I also tried an alternative method for choosing what splits to attempt: instead of doing the hierarchical clustering step, for bigger tags, I tried splitting out questions based on whether they had</em> some other tag <em>. However, this was much worse than the hierarchical clustering: I found no splits that actually reduced the BIC, whereas the hierarchical clustering went a bit crazy with the splitting, if anything. So it seems like tags might not actually be a great way of grouping questions together for the purpose of analyzing learning curves.</em></p>\n<h1>Residuals</h1>\n<p>This is quite a simple idea:</p>\n<ul>\n<li>For each data point, generate a prediction (probability of user answering correctly)</li>\n<li>Subtract that prediction from what actually happened (0 = incorrect or 1 = correct)</li>\n<li>average these residuals:<ul>\n<li>across each question</li>\n<li>across each user</li></ul></li>\n<li>When doing inference, for each data point, generate a prediction as normal; then add the per-user and per-question residuals to the predicted probability, and report that</li>\n<li>Update the running user averages with each new batch of dataa (I didn't bother updating the per-question one).</li>\n</ul>\n<p>It's a bit scary how well this works. I went from 0.728 with no residuals to 0.763 with both per-user and per-question residuals.</p>\n<h1>Things I didn't try</h1>\n<p>Lots of things I wanted to add, but ran out of time:</p>\n<ul>\n<li>Evaluate provided tags: As I mentioned above, the tags may not have actually been a good way of splitting the data for learning curve analysis. There's lots of interesting techniques in the literature, e.g. Q-matrix, rule space analysis, etc. that I just didn't get to. Part of the problem is that all the literature kind of assumes that I know a lot more about the nature of the problems/concepts/tags than just an integer label.</li>\n<li>clustering incorrect answers into \"misconceptions\": the same way that I clustered questions based on correctness, I could have clustered specific incorrect answers as seemingly related misconceptions</li>\n<li>incorporate question elapsed time: either into the logistic regression, or as its own separate predictor.</li>\n<li>incorporate timestamps (e.g. large gaps in time)</li>\n<li>dig into multiple-tag regression quirks: That regression assigned some interesting values to the \"learning curve\" slope of a lot of tags. Specifically, it seemed to treat a <em>lot</em> of tags as signals that the student is less likely too answer questions correctly if they've had a lot of \"experience\" with a tag.</li>\n<li>isolate specific types of 'bad' users: I think both the logistic regression and the clustering actually did a lot of that implicitly (e.g. capturing that users who seem to guess randomly will probably continue to guess randomly; or isolating groups of questions which seem to be given repeatedly to \"remedial\" users). But more can probably be done explcitly.</li>\n</ul>",
      "rawMarkdown": "First, a couple of notes:\n\n- No, I wasn't explictly *trying* to avoid NNs or LGBMs or doing anything the \"normal\" way. I just kind of prioritized what was interesting to me. Actually, I did try throwing some of the data in an LGBM at the last moment, but it did much *worse* than my more \"analytical\" solution.\n    - I was more interested in regression/clustering kind of stuff because I was hoping to learn something from the data about how/whether people actually learn through the app and such. I think there are some interesting points to discuss there,  which I have now done [here](https://www.kaggle.com/yanamal/riiid-dataset-what-can-we-learn-about-learning).\n- My submissions consistently did slightly *better* on the private data set than the public leaderboard, which lends credence to the idea of not needing cross-validation (in this particular case), and using more analytical methods like the Bayesian Information Criterion instead.\n- I'm happy to publish the whole actual pipeline of notebooks I used, too, if anyone's actually interested. Though I'm not sure how readable they are.\n\n# Overview\n\nThe model consists of 3 very simple (mathematically) building blocks:\n\n1. **Logistic regression** which seeks to capture how different users seem to learn different concepts, based on Learning Factor Analysis/Performance Factor Analysis (see also my notebook discussing LFA [here](https://www.kaggle.com/yanamal/learning-factor-analysis-are-tags-skills))\n2. **Hierarchical clustering** of tag content (questions and lectures) into \"sub-tags\", using the [Baysian Information Criterion](https://en.wikipedia.org/wiki/Bayesian_information_criterion) to decide when and whether to split\n3. Per-question and per-user **residuals**, as demonstrated in [this brilliantly simple notebook](https://www.kaggle.com/rar4dx/two-feature-model) by @rar4dx \n\n# Logistic regression of learning curves\n\nLearning Factor Analysis tries to determine the odds of making an error based on the number of encounters the student has had with some relevant concept so far. Performance Factor Analysis is similar, but measures the odds of making an error as a function of number of correct answers and number of wrong answers so far. Both use the same general kind of power function:\n\n$$ \\mathrm{odds}(\\mathrm{error}) = \\alpha \\prod_i \\mathrm{count_i}^{\\gamma_i}$$\n\nwhere \n- each `i` is some factor being taken into account (encounters in LFA, correct/wrong answers in PFA)\n- `alpha` and `gamma_i` are some parameters which define the \"shape\" of the the learning curve - the parameters to fit with ML\n- `count_i` is the count for that particular factor, for some student at some moment in time(e.g. \"7 encounters with this concept\" or \"4 correct answers\"\n\n\nIn practice, this amounts to doing logistic regression with the **logarithms** of your `count_i`s as input. (Taking the log of both sides of the equation so far gives you a linear equation to find the log-odds, which is exactly what logistic regression does). The intercept and coefficients from your logistic regression are `log(alpha)` and `gamma_i`s in the equation above.\n\n*Side note: The papers on LFA/PFA for some reason never explicitly mention that they take the log of the counts. Not sure if they neglected to take the log, or to mention it. But the evidence from studying real-world learning curves points to the specific power function above, so the log would be necessary. And empirically, it did work better with the log.*\n\nI took this power law idea even further than PFA, and used five different counters for each \"concept\" (tag or, after the hierarchical clustering, sub-tag):\n\n- Number of lectures viewed for the concept\n- Number of \"diagnostic\" questions answered correctly for the concept\n- Number of \"diagnostic\" questions answered incorrectly for the concept\n- Number of \"feedback\" questions answered correctly for the concept\n- Number of \"feedback\" questions answered incorrectly for the concept\n\nThe diagnostic/feedback split was done based on whether that particular question instance had an explanation (given by `prior_question_had_explanation` and then painstakingly shifted back to the question it belongs to)\n\nI did also try combinations with fewer parameters (e.g. ignoring the feedback/diagnostic split; ignoring  lectures). But the Bayesian Information  Criterion (discussed in more detail in the next section) nearly always turned out better for the full split (except on one or two tags), so I went with that.\n\nThere's the slight wrinkle that lots of questions actually have *several* tags, and so several predictions from each tag. For those, I found the most effective thing was to average the log-odds predictions for each tag and then use that as the log-odds value (which is conceptually equivalent to doing a geometric average of the odds).\n\nI trained each regression on all available training data for that tag/subtag, and only ever evaluated the output on the training data itself until I actually submitted (The AUC tended to be about 0.02 better on the training data than the test data).\n\n\n*Side note: I also did a test run of a single regression over* all *the tags - or at least all the ones that actually intersected with each other - but it wouldn't converge, even with just two dimensions (correct/wrong). Surprisingly, even the unconverged version tended to do a tiny bit better than the equivalent per-tag average (also only correct/wrong, no sub-tag clustering). I say this is surprising because a single regression has a single intercept, which means that all users answering their very first question were lumped in the same exact group, and generally, it should be worse at distinguishing between beginner users.*\n\n# Hierarchical Clustering\n\nThe goodness-of-fit for the regression varied quite a bit from tag to tag. So I decided to split the questions and lectures further into \"sub-tags\". \n\nI determined the \"similarity\" of questions within the same tag like this:\n\n1. Make a User x Question matrix; each cell in the matrix is the *average correctness* for how that user answered that question (1 = correct every time; 0 = wrong every time; 0.5 = half and half; etc.). This matrix is pretty sparse,\n    - The lectures are also in this matrix, alongside the questions: the value is 1 if that user has seen that lecture.\n2. Use that matrix to get a Question x Question (well, Content x Content) correlation matrix with pandas.corr and a similarity function I called \"Jacardish\" -Jacard similarity, but can also deal with non-binary values. (I think it basically ended up being a normalized manhattan distance).\n3. Use the similarity metric *again* on the Content x Content matrix to get a sort of \"similarity similarity\" measure - questions that are similar to the same questions, and different from the same questions, are considered more \"similarly-similar\" to each other.\n    - This was kind of just  a happy accident. I stumbled upon this idea by reading some random blog post about how to make correlation matrices more readable by doing this kind of correlation on a correlation matrix. At first I thought it was kind of silly, but empirically, it produced better results - more balanced hierarchical trees in the next step\n    \nAfter that, I used [scipy's hierarchical clustering library](https://docs.scipy.org/doc/scipy/reference/generated/scipy.cluster.hierarchy.linkage.html) to derive a [hierachical tree](https://en.wikipedia.org/wiki/Hierarchical_clustering) where each leaf is  a question/lecture, and they are progressively joined into clusters of similar questions.\n\nNow I could  split each tag's content into sub-tags by splitting the hierarchical tree at the root. For each tree, I decided whether to split it by re-doing the logistic regression from step one with the split, and comparing it to the quality of the regression without the split. Specifically, I calculated the [Baysian Information Criterion (BIC)](https://en.wikipedia.org/wiki/Bayesian_information_criterion) for each version, and chose to split if the BIC was smaller. The BIC takes into account the log-likelihood of the model (log of the chance of seeing the exact data we see, if the model is right), the number of parameters in the model, and the (log of) number of total data points. In general, a more complex model is more prone to overfitting, since there are more \"knobs\" one can turn to fit the training data precisely. But if the BIC for a more complex model is smaller than the simpler version, then the increase in complexity is probably justified by the increase in how well the model explains the data. \n\nEven though the BIC  is considered to be pretty harsh and pessimistic about model complexity, it happily split the tags down into over 2500 subtags, sometimes with just one question in a \"sub-tag\". I considered dialing it back artificially a bit, by introducing a bigger threshold for splitting. But in the end, I trusted the math and it seems to have  worked out. \n\n*Side note: I also tried an alternative method for choosing what splits to attempt: instead of doing the hierarchical clustering step, for bigger tags, I tried splitting out questions based on whether they had* some other tag *. However, this was much worse than the hierarchical clustering: I found no splits that actually reduced the BIC, whereas the hierarchical clustering went a bit crazy with the splitting, if anything. So it seems like tags might not actually be a great way of grouping questions together for the purpose of analyzing learning curves.*\n\n# Residuals\n\nThis is quite a simple idea:\n\n- For each data point, generate a prediction (probability of user answering correctly)\n- Subtract that prediction from what actually happened (0 = incorrect or 1 = correct)\n- average these residuals:\n    - across each question\n    - across each user\n- When doing inference, for each data point, generate a prediction as normal; then add the per-user and per-question residuals to the predicted probability, and report that\n- Update the running user averages with each new batch of dataa (I didn't bother updating the per-question one).\n\nIt's a bit scary how well this works. I went from 0.728 with no residuals to 0.763 with both per-user and per-question residuals.\n\n# Things I didn't try\n\nLots of things I wanted to add, but ran out of time:\n\n- Evaluate provided tags: As I mentioned above, the tags may not have actually been a good way of splitting the data for learning curve analysis. There's lots of interesting techniques in the literature, e.g. Q-matrix, rule space analysis, etc. that I just didn't get to. Part of the problem is that all the literature kind of assumes that I know a lot more about the nature of the problems/concepts/tags than just an integer label.\n- clustering incorrect answers into \"misconceptions\": the same way that I clustered questions based on correctness, I could have clustered specific incorrect answers as seemingly related misconceptions\n- incorporate question elapsed time: either into the logistic regression, or as its own separate predictor.\n- incorporate timestamps (e.g. large gaps in time)\n- dig into multiple-tag regression quirks: That regression assigned some interesting values to the \"learning curve\" slope of a lot of tags. Specifically, it seemed to treat a *lot* of tags as signals that the student is less likely too answer questions correctly if they've had a lot of \"experience\" with a tag.\n- isolate specific types of 'bad' users: I think both the logistic regression and the clustering actually did a lot of that implicitly (e.g. capturing that users who seem to guess randomly will probably continue to guess randomly; or isolating groups of questions which seem to be given repeatedly to \"remedial\" users). But more can probably be done explcitly.\n",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1153854": "First, a couple of notes:\n\n- No, I wasn't explictly *trying* to avoid NNs or LGBMs or doing anything the \"normal\" way. I just kind of prioritized what was interesting to me. Actually, I did try throwing some of the data in an LGBM at the last moment, but it did much *worse* than my more \"analytical\" solution.\n    - I was more interested in regression/clustering kind of stuff because I was hoping to learn something from the data about how/whether people actually learn through the app and such. I think there are some interesting points to discuss there,  which I have now done [here](https://www.kaggle.com/yanamal/riiid-dataset-what-can-we-learn-about-learning).\n- My submissions consistently did slightly *better* on the private data set than the public leaderboard, which lends credence to the idea of not needing cross-validation (in this particular case), and using more analytical methods like the Bayesian Information Criterion instead.\n- I'm happy to publish the whole actual pipeline of notebooks I used, too, if anyone's actually interested. Though I'm not sure how readable they are.\n\n# Overview\n\nThe model consists of 3 very simple (mathematically) building blocks:\n\n1. **Logistic regression** which seeks to capture how different users seem to learn different concepts, based on Learning Factor Analysis/Performance Factor Analysis (see also my notebook discussing LFA [here](https://www.kaggle.com/yanamal/learning-factor-analysis-are-tags-skills))\n2. **Hierarchical clustering** of tag content (questions and lectures) into \"sub-tags\", using the [Baysian Information Criterion](https://en.wikipedia.org/wiki/Bayesian_information_criterion) to decide when and whether to split\n3. Per-question and per-user **residuals**, as demonstrated in [this brilliantly simple notebook](https://www.kaggle.com/rar4dx/two-feature-model) by @rar4dx \n\n# Logistic regression of learning curves\n\nLearning Factor Analysis tries to determine the odds of making an error based on the number of encounters the student has had with some relevant concept so far. Performance Factor Analysis is similar, but measures the odds of making an error as a function of number of correct answers and number of wrong answers so far. Both use the same general kind of power function:\n\n$$ \\mathrm{odds}(\\mathrm{error}) = \\alpha \\prod_i \\mathrm{count_i}^{\\gamma_i}$$\n\nwhere \n- each `i` is some factor being taken into account (encounters in LFA, correct/wrong answers in PFA)\n- `alpha` and `gamma_i` are some parameters which define the \"shape\" of the the learning curve - the parameters to fit with ML\n- `count_i` is the count for that particular factor, for some student at some moment in time(e.g. \"7 encounters with this concept\" or \"4 correct answers\"\n\n\nIn practice, this amounts to doing logistic regression with the **logarithms** of your `count_i`s as input. (Taking the log of both sides of the equation so far gives you a linear equation to find the log-odds, which is exactly what logistic regression does). The intercept and coefficients from your logistic regression are `log(alpha)` and `gamma_i`s in the equation above.\n\n*Side note: The papers on LFA/PFA for some reason never explicitly mention that they take the log of the counts. Not sure if they neglected to take the log, or to mention it. But the evidence from studying real-world learning curves points to the specific power function above, so the log would be necessary. And empirically, it did work better with the log.*\n\nI took this power law idea even further than PFA, and used five different counters for each \"concept\" (tag or, after the hierarchical clustering, sub-tag):\n\n- Number of lectures viewed for the concept\n- Number of \"diagnostic\" questions answered correctly for the concept\n- Number of \"diagnostic\" questions answered incorrectly for the concept\n- Number of \"feedback\" questions answered correctly for the concept\n- Number of \"feedback\" questions answered incorrectly for the concept\n\nThe diagnostic/feedback split was done based on whether that particular question instance had an explanation (given by `prior_question_had_explanation` and then painstakingly shifted back to the question it belongs to)\n\nI did also try combinations with fewer parameters (e.g. ignoring the feedback/diagnostic split; ignoring  lectures). But the Bayesian Information  Criterion (discussed in more detail in the next section) nearly always turned out better for the full split (except on one or two tags), so I went with that.\n\nThere's the slight wrinkle that lots of questions actually have *several* tags, and so several predictions from each tag. For those, I found the most effective thing was to average the log-odds predictions for each tag and then use that as the log-odds value (which is conceptually equivalent to doing a geometric average of the odds).\n\nI trained each regression on all available training data for that tag/subtag, and only ever evaluated the output on the training data itself until I actually submitted (The AUC tended to be about 0.02 better on the training data than the test data).\n\n\n*Side note: I also did a test run of a single regression over* all *the tags - or at least all the ones that actually intersected with each other - but it wouldn't converge, even with just two dimensions (correct/wrong). Surprisingly, even the unconverged version tended to do a tiny bit better than the equivalent per-tag average (also only correct/wrong, no sub-tag clustering). I say this is surprising because a single regression has a single intercept, which means that all users answering their very first question were lumped in the same exact group, and generally, it should be worse at distinguishing between beginner users.*\n\n# Hierarchical Clustering\n\nThe goodness-of-fit for the regression varied quite a bit from tag to tag. So I decided to split the questions and lectures further into \"sub-tags\". \n\nI determined the \"similarity\" of questions within the same tag like this:\n\n1. Make a User x Question matrix; each cell in the matrix is the *average correctness* for how that user answered that question (1 = correct every time; 0 = wrong every time; 0.5 = half and half; etc.). This matrix is pretty sparse,\n    - The lectures are also in this matrix, alongside the questions: the value is 1 if that user has seen that lecture.\n2. Use that matrix to get a Question x Question (well, Content x Content) correlation matrix with pandas.corr and a similarity function I called \"Jacardish\" -Jacard similarity, but can also deal with non-binary values. (I think it basically ended up being a normalized manhattan distance).\n3. Use the similarity metric *again* on the Content x Content matrix to get a sort of \"similarity similarity\" measure - questions that are similar to the same questions, and different from the same questions, are considered more \"similarly-similar\" to each other.\n    - This was kind of just  a happy accident. I stumbled upon this idea by reading some random blog post about how to make correlation matrices more readable by doing this kind of correlation on a correlation matrix. At first I thought it was kind of silly, but empirically, it produced better results - more balanced hierarchical trees in the next step\n    \nAfter that, I used [scipy's hierarchical clustering library](https://docs.scipy.org/doc/scipy/reference/generated/scipy.cluster.hierarchy.linkage.html) to derive a [hierachical tree](https://en.wikipedia.org/wiki/Hierarchical_clustering) where each leaf is  a question/lecture, and they are progressively joined into clusters of similar questions.\n\nNow I could  split each tag's content into sub-tags by splitting the hierarchical tree at the root. For each tree, I decided whether to split it by re-doing the logistic regression from step one with the split, and comparing it to the quality of the regression without the split. Specifically, I calculated the [Baysian Information Criterion (BIC)](https://en.wikipedia.org/wiki/Bayesian_information_criterion) for each version, and chose to split if the BIC was smaller. The BIC takes into account the log-likelihood of the model (log of the chance of seeing the exact data we see, if the model is right), the number of parameters in the model, and the (log of) number of total data points. In general, a more complex model is more prone to overfitting, since there are more \"knobs\" one can turn to fit the training data precisely. But if the BIC for a more complex model is smaller than the simpler version, then the increase in complexity is probably justified by the increase in how well the model explains the data. \n\nEven though the BIC  is considered to be pretty harsh and pessimistic about model complexity, it happily split the tags down into over 2500 subtags, sometimes with just one question in a \"sub-tag\". I considered dialing it back artificially a bit, by introducing a bigger threshold for splitting. But in the end, I trusted the math and it seems to have  worked out. \n\n*Side note: I also tried an alternative method for choosing what splits to attempt: instead of doing the hierarchical clustering step, for bigger tags, I tried splitting out questions based on whether they had* some other tag *. However, this was much worse than the hierarchical clustering: I found no splits that actually reduced the BIC, whereas the hierarchical clustering went a bit crazy with the splitting, if anything. So it seems like tags might not actually be a great way of grouping questions together for the purpose of analyzing learning curves.*\n\n# Residuals\n\nThis is quite a simple idea:\n\n- For each data point, generate a prediction (probability of user answering correctly)\n- Subtract that prediction from what actually happened (0 = incorrect or 1 = correct)\n- average these residuals:\n    - across each question\n    - across each user\n- When doing inference, for each data point, generate a prediction as normal; then add the per-user and per-question residuals to the predicted probability, and report that\n- Update the running user averages with each new batch of dataa (I didn't bother updating the per-question one).\n\nIt's a bit scary how well this works. I went from 0.728 with no residuals to 0.763 with both per-user and per-question residuals.\n\n# Things I didn't try\n\nLots of things I wanted to add, but ran out of time:\n\n- Evaluate provided tags: As I mentioned above, the tags may not have actually been a good way of splitting the data for learning curve analysis. There's lots of interesting techniques in the literature, e.g. Q-matrix, rule space analysis, etc. that I just didn't get to. Part of the problem is that all the literature kind of assumes that I know a lot more about the nature of the problems/concepts/tags than just an integer label.\n- clustering incorrect answers into \"misconceptions\": the same way that I clustered questions based on correctness, I could have clustered specific incorrect answers as seemingly related misconceptions\n- incorporate question elapsed time: either into the logistic regression, or as its own separate predictor.\n- incorporate timestamps (e.g. large gaps in time)\n- dig into multiple-tag regression quirks: That regression assigned some interesting values to the \"learning curve\" slope of a lot of tags. Specifically, it seemed to treat a *lot* of tags as signals that the student is less likely too answer questions correctly if they've had a lot of \"experience\" with a tag.\n- isolate specific types of 'bad' users: I think both the logistic regression and the clustering actually did a lot of that implicitly (e.g. capturing that users who seem to guess randomly will probably continue to guess randomly; or isolating groups of questions which seem to be given repeatedly to \"remedial\" users). But more can probably be done explcitly.\n"
  }
}