{
  "id": 194032,
  "title": "The next question: Is it model determined?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/194032",
  "author_name": "",
  "post_date": "2020-10-30T09:53:45.342903400Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Is the next question in sequence (or next bundle of questions) to the user emitted by clever Riiid models (on their part likely also digesting user and question mileage somewhat like in this challenge)?</p>",
  "messages": [
    {
      "id": "1064600",
      "postDate": "10/30/2020 09:53:45",
      "content": "<p>Is the next question in sequence (or next bundle of questions) to the user emitted by clever Riiid models (on their part likely also digesting user and question mileage somewhat like in this challenge)?</p>",
      "rawMarkdown": "Is the next question in sequence (or next bundle of questions) to the user emitted by clever Riiid models (on their part likely also digesting user and question mileage somewhat like in this challenge)?",
      "votes": null
    },
    {
      "id": "1065256",
      "postDate": "10/31/2020 04:53:47",
      "content": "<p>It's an interesting question which only the hosts would be able to answer. My hunch is it's model determined.<br>\nI'm curious to know if getting the answer would help improve models for this competition?</p>",
      "rawMarkdown": "It's an interesting question which only the hosts would be able to answer. My hunch is it's model determined.\nI'm curious to know if getting the answer would help improve models for this competition?",
      "votes": null
    },
    {
      "id": "1065261",
      "postDate": "10/31/2020 05:04:04",
      "content": "<p>It's highly likely that it's model determined as they have  Self-Personalization and Self-Engagement (pre-determined based on user's metrics etc), Check this <a href=\"https://arxiv.org/pdf/2005.03818.pdf\" target=\"_blank\">paper</a> by Riiid!</p>",
      "rawMarkdown": "It's highly likely that it's model determined as they have  Self-Personalization and Self-Engagement (pre-determined based on user's metrics etc), Check this [paper](https://arxiv.org/pdf/2005.03818.pdf) by Riiid!",
      "votes": null
    },
    {
      "id": "1065263",
      "postDate": "10/31/2020 05:10:02",
      "content": "<p>The only little doubt I had is if the dataset was compiled 'before' the Riiid models were put into production.</p>\n<p>I couldn't find a statement that clarifies this. If you noticed it somewhere, let me know!</p>",
      "rawMarkdown": "The only little doubt I had is if the dataset was compiled 'before' the Riiid models were put into production.\n\nI couldn't find a statement that clarifies this. If you noticed it somewhere, let me know!",
      "votes": null
    },
    {
      "id": "1065539",
      "postDate": "10/31/2020 12:16:41",
      "content": "<p>EdNet <a href=\"https://github.com/riiid/ednet\" target=\"_blank\">description</a> says that this is a model (one of the sources!). </p>\n<blockquote>\n  <p>For each day, Santa recommends questions and lectures based on each student's current knowledge status, i.e. correctness probabilities predicted by the Collaborative Filtering model. Such source is called Today's Recommendation.</p>\n</blockquote>\n<p>another important point</p>\n<blockquote>\n  <p>Once the number of incorrect answers to questions with particular tags exceeds certain threshold, Santa suggests lectures and questions with corresponding tags. Such suggestion is recorded as adaptive_offer. It also offers lectures and questions if the average correctness rate of questions with particular tags decreased by more than a certain threshold.</p>\n</blockquote>",
      "rawMarkdown": "EdNet [description](https://github.com/riiid/ednet) says that this is a model (one of the sources!). \n> For each day, Santa recommends questions and lectures based on each student's current knowledge status, i.e. correctness probabilities predicted by the Collaborative Filtering model. Such source is called Today's Recommendation.\n\nanother important point\n> Once the number of incorrect answers to questions with particular tags exceeds certain threshold, Santa suggests lectures and questions with corresponding tags. Such suggestion is recorded as adaptive_offer. It also offers lectures and questions if the average correctness rate of questions with particular tags decreased by more than a certain threshold.",
      "votes": null
    },
    {
      "id": "1066080",
      "postDate": "11/01/2020 09:36:10",
      "content": "<p>How could knowledge of such a running question producing model help progress in this competition? I am unsure if at all, but have been thinking along these lines:</p>\n<p>Assumptions:<br>\n*This question model is indeed running emitting targeted questions, to which we, immediately after, are figuring out answer correctness.</p>\n<p>*The question model is purposed to pose questions at moderate difficulty level  - questions that are hard to predict answer correctness to.  I would reckon that both silly easy and stupid hard questions would be wastes of time and counter-examples of a tailored user journey.</p>\n<p>With this given, we might try to seperately setup and train a next-question model towards the set of questions, multiclass. The idea is then to bring this model's question predictions as data into the competion's answer model: Not only its best question guess, but say the best five, with measurements of prediction uncertainty et.c., <br>\nand/or a trained layer out of a neural net, if the model is of such a kind.</p>\n<p>The idea then, or notion or hope, is that our question model is mimicking the real but unseen riiiid! next question model so well that it is picking up on complex interactions originally trained on data that we are not given in this competition, which can be data recorded during other time periods, or different feature sets from what we are given here.</p>\n<p>That's a whole lot of if's and assumptions, and an implementation can amount to a time costly detour of exactly zero return. It also appears I have shuffled the sequence of:</p>\n<p>1.Read what's actually documented and 2.Start assuming and rambling. I am referring to Pavel Orlov's findings (and thanks!)</p>\n<p>I would in any case be interested to learn about logic errors in my reasoning above, of which there surely may be one or two. Thanks.</p>",
      "rawMarkdown": "How could knowledge of such a running question producing model help progress in this competition? I am unsure if at all, but have been thinking along these lines:\n\nAssumptions:\n*This question model is indeed running emitting targeted questions, to which we, immediately after, are figuring out answer correctness.\n\n*The question model is purposed to pose questions at moderate difficulty level  - questions that are hard to predict answer correctness to.  I would reckon that both silly easy and stupid hard questions would be wastes of time and counter-examples of a tailored user journey.\n           \nWith this given, we might try to seperately setup and train a next-question model towards the set of questions, multiclass. The idea is then to bring this model's question predictions as data into the competion's answer model: Not only its best question guess, but say the best five, with measurements of prediction uncertainty et.c., \nand/or a trained layer out of a neural net, if the model is of such a kind.\n\nThe idea then, or notion or hope, is that our question model is mimicking the real but unseen riiiid! next question model so well that it is picking up on complex interactions originally trained on data that we are not given in this competition, which can be data recorded during other time periods, or different feature sets from what we are given here.\n\nThat's a whole lot of if's and assumptions, and an implementation can amount to a time costly detour of exactly zero return. It also appears I have shuffled the sequence of:\n\n 1.Read what's actually documented and 2.Start assuming and rambling. I am referring to Pavel Orlov's findings (and thanks!)\n\nI would in any case be interested to learn about logic errors in my reasoning above, of which there surely may be one or two. Thanks.",
      "votes": null
    },
    {
      "id": "1066083",
      "postDate": "11/01/2020 09:45:30",
      "content": "<p>I understand the thought process because I had thought about it as well. The place where I was not convinced of it's use in this competition is:</p>\n<p>1 - We actually do know the next question that is asked to user because it is part of the test data (before scoring). So is there still value in predicting or knowing why a certain question was asked? Well, I don't know 🤔</p>\n<p>2 - Hypothetically even if we did end up building the perfect model to predict the next question, how/why would that help in predicting the answer correctness? Maybe if we knew that for a certain user the intention is to ask a question that the user is unlikely to solve, then yes maybe it is useful.</p>\n<p>Thoughts galore but unconvincing to me yet 😬</p>",
      "rawMarkdown": "I understand the thought process because I had thought about it as well. The place where I was not convinced of it's use in this competition is:\n\n1 - We actually do know the next question that is asked to user because it is part of the test data (before scoring). So is there still value in predicting or knowing why a certain question was asked? Well, I don't know 🤔\n\n2 - Hypothetically even if we did end up building the perfect model to predict the next question, how/why would that help in predicting the answer correctness? Maybe if we knew that for a certain user the intention is to ask a question that the user is unlikely to solve, then yes maybe it is useful.\n\nThoughts galore but unconvincing to me yet 😬",
      "votes": null
    },
    {
      "id": "1066111",
      "postDate": "11/01/2020 10:40:50",
      "content": "<p>For me, the program's tactics for changing the complexity of the proposed questions are more important</p>\n<blockquote>\n  <p>Once the number of incorrect answers to questions with particular tags exceeds certain threshold, then …</p>\n</blockquote>\n<p>and</p>\n<blockquote>\n  <p>… if the average correctness rate of questions with particular tags decreased by more than a certain threshold …</p>\n</blockquote>\n<p>So I see the benefit of predicting: a) is the question suggested by the program or selected by the user? b) the complexity of questions for the user : \"did not change\", \"increased\", \"decreased\".<br>\nIf you know this, you can somehow use it for prediction.</p>\n<p>If I understand correctly, these are the 'jumps' we see on the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194141\" target=\"_blank\">chart</a></p>",
      "rawMarkdown": "For me, the program's tactics for changing the complexity of the proposed questions are more important\n> Once the number of incorrect answers to questions with particular tags exceeds certain threshold, then ...\n\nand\n\n> ... if the average correctness rate of questions with particular tags decreased by more than a certain threshold ...\n\nSo I see the benefit of predicting: a) is the question suggested by the program or selected by the user? b) the complexity of questions for the user : \"did not change\", \"increased\", \"decreased\".\nIf you know this, you can somehow use it for prediction.\n\nIf I understand correctly, these are the 'jumps' we see on the [chart](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194141)",
      "votes": null
    },
    {
      "id": "1068218",
      "postDate": "11/03/2020 07:25:46",
      "content": "<p>The idea is that alongside the next question itself, which we already know, the model would also provide: Runner up best questions, certainty scores and more, and that these new features might contribute to a better understanding of the actually issued questions's fit to the situation.</p>\n<p>But that said, I am unconvinced, too, and may leave efforts in this direction for now. Thanks again for coming back.</p>",
      "rawMarkdown": "The idea is that alongside the next question itself, which we already know, the model would also provide: Runner up best questions, certainty scores and more, and that these new features might contribute to a better understanding of the actually issued questions's fit to the situation.\n\nBut that said, I am unconvinced, too, and may leave efforts in this direction for now. Thanks again for coming back.",
      "votes": null
    },
    {
      "id": "1068242",
      "postDate": "11/03/2020 07:48:45",
      "content": "<p>Thanks for bringing the existence of this rule-based engine into my knowledge, Pavel Orlov. I'll browse up on its likely actions with respect to various jumps and more. My own assumptions and attempted approach start to seem a poor fit to the reality of this competiion - and quite 'academic'.</p>",
      "rawMarkdown": "Thanks for bringing the existence of this rule-based engine into my knowledge, Pavel Orlov. I'll browse up on its likely actions with respect to various jumps and more. My own assumptions and attempted approach start to seem a poor fit to the reality of this competiion - and quite 'academic'.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1065256,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "10/31/2020 04:53:47",
      "content": "<p>It's an interesting question which only the hosts would be able to answer. My hunch is it's model determined.<br>\nI'm curious to know if getting the answer would help improve models for this competition?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1065261,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/31/2020 05:04:04",
          "content": "<p>It's highly likely that it's model determined as they have  Self-Personalization and Self-Engagement (pre-determined based on user's metrics etc), Check this <a href=\"https://arxiv.org/pdf/2005.03818.pdf\" target=\"_blank\">paper</a> by Riiid!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065263,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "10/31/2020 05:10:02",
          "content": "<p>The only little doubt I had is if the dataset was compiled 'before' the Riiid models were put into production.</p>\n<p>I couldn't find a statement that clarifies this. If you noticed it somewhere, let me know!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1065539,
      "author_name": "sapr3s",
      "author_url": "",
      "post_date": "10/31/2020 12:16:41",
      "content": "<p>EdNet <a href=\"https://github.com/riiid/ednet\" target=\"_blank\">description</a> says that this is a model (one of the sources!). </p>\n<blockquote>\n  <p>For each day, Santa recommends questions and lectures based on each student's current knowledge status, i.e. correctness probabilities predicted by the Collaborative Filtering model. Such source is called Today's Recommendation.</p>\n</blockquote>\n<p>another important point</p>\n<blockquote>\n  <p>Once the number of incorrect answers to questions with particular tags exceeds certain threshold, Santa suggests lectures and questions with corresponding tags. Such suggestion is recorded as adaptive_offer. It also offers lectures and questions if the average correctness rate of questions with particular tags decreased by more than a certain threshold.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1066080,
      "author_name": "nohitwonder",
      "author_url": "",
      "post_date": "11/01/2020 09:36:10",
      "content": "<p>How could knowledge of such a running question producing model help progress in this competition? I am unsure if at all, but have been thinking along these lines:</p>\n<p>Assumptions:<br>\n*This question model is indeed running emitting targeted questions, to which we, immediately after, are figuring out answer correctness.</p>\n<p>*The question model is purposed to pose questions at moderate difficulty level  - questions that are hard to predict answer correctness to.  I would reckon that both silly easy and stupid hard questions would be wastes of time and counter-examples of a tailored user journey.</p>\n<p>With this given, we might try to seperately setup and train a next-question model towards the set of questions, multiclass. The idea is then to bring this model's question predictions as data into the competion's answer model: Not only its best question guess, but say the best five, with measurements of prediction uncertainty et.c., <br>\nand/or a trained layer out of a neural net, if the model is of such a kind.</p>\n<p>The idea then, or notion or hope, is that our question model is mimicking the real but unseen riiiid! next question model so well that it is picking up on complex interactions originally trained on data that we are not given in this competition, which can be data recorded during other time periods, or different feature sets from what we are given here.</p>\n<p>That's a whole lot of if's and assumptions, and an implementation can amount to a time costly detour of exactly zero return. It also appears I have shuffled the sequence of:</p>\n<p>1.Read what's actually documented and 2.Start assuming and rambling. I am referring to Pavel Orlov's findings (and thanks!)</p>\n<p>I would in any case be interested to learn about logic errors in my reasoning above, of which there surely may be one or two. Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1066083,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "11/01/2020 09:45:30",
          "content": "<p>I understand the thought process because I had thought about it as well. The place where I was not convinced of it's use in this competition is:</p>\n<p>1 - We actually do know the next question that is asked to user because it is part of the test data (before scoring). So is there still value in predicting or knowing why a certain question was asked? Well, I don't know 🤔</p>\n<p>2 - Hypothetically even if we did end up building the perfect model to predict the next question, how/why would that help in predicting the answer correctness? Maybe if we knew that for a certain user the intention is to ask a question that the user is unlikely to solve, then yes maybe it is useful.</p>\n<p>Thoughts galore but unconvincing to me yet 😬</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066111,
          "author_name": "sapr3s",
          "author_url": "",
          "post_date": "11/01/2020 10:40:50",
          "content": "<p>For me, the program's tactics for changing the complexity of the proposed questions are more important</p>\n<blockquote>\n  <p>Once the number of incorrect answers to questions with particular tags exceeds certain threshold, then …</p>\n</blockquote>\n<p>and</p>\n<blockquote>\n  <p>… if the average correctness rate of questions with particular tags decreased by more than a certain threshold …</p>\n</blockquote>\n<p>So I see the benefit of predicting: a) is the question suggested by the program or selected by the user? b) the complexity of questions for the user : \"did not change\", \"increased\", \"decreased\".<br>\nIf you know this, you can somehow use it for prediction.</p>\n<p>If I understand correctly, these are the 'jumps' we see on the <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194141\" target=\"_blank\">chart</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068218,
          "author_name": "nohitwonder",
          "author_url": "",
          "post_date": "11/03/2020 07:25:46",
          "content": "<p>The idea is that alongside the next question itself, which we already know, the model would also provide: Runner up best questions, certainty scores and more, and that these new features might contribute to a better understanding of the actually issued questions's fit to the situation.</p>\n<p>But that said, I am unconvinced, too, and may leave efforts in this direction for now. Thanks again for coming back.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1068242,
          "author_name": "nohitwonder",
          "author_url": "",
          "post_date": "11/03/2020 07:48:45",
          "content": "<p>Thanks for bringing the existence of this rule-based engine into my knowledge, Pavel Orlov. I'll browse up on its likely actions with respect to various jumps and more. My own assumptions and attempted approach start to seem a poor fit to the reality of this competiion - and quite 'academic'.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1064600": "Is the next question in sequence (or next bundle of questions) to the user emitted by clever Riiid models (on their part likely also digesting user and question mileage somewhat like in this challenge)?",
    "1065256": "It's an interesting question which only the hosts would be able to answer. My hunch is it's model determined.\nI'm curious to know if getting the answer would help improve models for this competition?",
    "1065261": "It's highly likely that it's model determined as they have  Self-Personalization and Self-Engagement (pre-determined based on user's metrics etc), Check this [paper](https://arxiv.org/pdf/2005.03818.pdf) by Riiid!",
    "1065263": "The only little doubt I had is if the dataset was compiled 'before' the Riiid models were put into production.\n\nI couldn't find a statement that clarifies this. If you noticed it somewhere, let me know!",
    "1065539": "EdNet [description](https://github.com/riiid/ednet) says that this is a model (one of the sources!). \n> For each day, Santa recommends questions and lectures based on each student's current knowledge status, i.e. correctness probabilities predicted by the Collaborative Filtering model. Such source is called Today's Recommendation.\n\nanother important point\n> Once the number of incorrect answers to questions with particular tags exceeds certain threshold, Santa suggests lectures and questions with corresponding tags. Such suggestion is recorded as adaptive_offer. It also offers lectures and questions if the average correctness rate of questions with particular tags decreased by more than a certain threshold.",
    "1066080": "How could knowledge of such a running question producing model help progress in this competition? I am unsure if at all, but have been thinking along these lines:\n\nAssumptions:\n*This question model is indeed running emitting targeted questions, to which we, immediately after, are figuring out answer correctness.\n\n*The question model is purposed to pose questions at moderate difficulty level  - questions that are hard to predict answer correctness to.  I would reckon that both silly easy and stupid hard questions would be wastes of time and counter-examples of a tailored user journey.\n           \nWith this given, we might try to seperately setup and train a next-question model towards the set of questions, multiclass. The idea is then to bring this model's question predictions as data into the competion's answer model: Not only its best question guess, but say the best five, with measurements of prediction uncertainty et.c., \nand/or a trained layer out of a neural net, if the model is of such a kind.\n\nThe idea then, or notion or hope, is that our question model is mimicking the real but unseen riiiid! next question model so well that it is picking up on complex interactions originally trained on data that we are not given in this competition, which can be data recorded during other time periods, or different feature sets from what we are given here.\n\nThat's a whole lot of if's and assumptions, and an implementation can amount to a time costly detour of exactly zero return. It also appears I have shuffled the sequence of:\n\n 1.Read what's actually documented and 2.Start assuming and rambling. I am referring to Pavel Orlov's findings (and thanks!)\n\nI would in any case be interested to learn about logic errors in my reasoning above, of which there surely may be one or two. Thanks.",
    "1066083": "I understand the thought process because I had thought about it as well. The place where I was not convinced of it's use in this competition is:\n\n1 - We actually do know the next question that is asked to user because it is part of the test data (before scoring). So is there still value in predicting or knowing why a certain question was asked? Well, I don't know 🤔\n\n2 - Hypothetically even if we did end up building the perfect model to predict the next question, how/why would that help in predicting the answer correctness? Maybe if we knew that for a certain user the intention is to ask a question that the user is unlikely to solve, then yes maybe it is useful.\n\nThoughts galore but unconvincing to me yet 😬",
    "1066111": "For me, the program's tactics for changing the complexity of the proposed questions are more important\n> Once the number of incorrect answers to questions with particular tags exceeds certain threshold, then ...\n\nand\n\n> ... if the average correctness rate of questions with particular tags decreased by more than a certain threshold ...\n\nSo I see the benefit of predicting: a) is the question suggested by the program or selected by the user? b) the complexity of questions for the user : \"did not change\", \"increased\", \"decreased\".\nIf you know this, you can somehow use it for prediction.\n\nIf I understand correctly, these are the 'jumps' we see on the [chart](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194141)",
    "1068218": "The idea is that alongside the next question itself, which we already know, the model would also provide: Runner up best questions, certainty scores and more, and that these new features might contribute to a better understanding of the actually issued questions's fit to the situation.\n\nBut that said, I am unconvinced, too, and may leave efforts in this direction for now. Thanks again for coming back.",
    "1068242": "Thanks for bringing the existence of this rule-based engine into my knowledge, Pavel Orlov. I'll browse up on its likely actions with respect to various jumps and more. My own assumptions and attempted approach start to seem a poor fit to the reality of this competiion - and quite 'academic'."
  },
  "source": "meta"
}