{
  "id": 189409,
  "title": "Riiid & EdNet resources",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189409",
  "author_name": "",
  "post_date": "2020-10-07T14:06:59.631841800Z",
  "votes": 43,
  "comment_count": 5,
  "views": 0,
  "content": "<h2>Riiid company</h2>\n<p><img src=\"https://prnewswire2-a.akamaihd.net/p/1893751/sp/189375100/thumbnail/entry_id/0_u0ia3ust/def_height/840/def_width/1400/version/100012/type/1\"></p>\n<p>Riiid was founded in 2014. In 2016, the company pioneered the patented AI-based education improvement algorithm and launched 'Santa for TOEIC!' beta version (final version launched next year). Santa is a mobile test prep application for the popular English proficiency exam, Test of English for International Communication (TOEIC). Santa  reached No. 1 in sales among education apps in Japan and Korea. Riiid’s proprietary AI technology analyzes student data and content, predicts scores and user behavior, and recommends personalized study plans to help students optimize their learning potential. In 2019 launched Santa SAT and in 2020 EdNet, the largest open database for AI Education.</p>\n<h2>EdNet arXiv paper</h2>\n<p>The original paper on EdNet, published on arXiv: <a href=\"https://arxiv.org/pdf/1912.03072.pdf\" target=\"_blank\"><strong>EdNet: A Large-Scale Hierarchical Dataset in\nEducation</strong></a>; excerpt from the Abstract: \"EdNet, a large-scale hierarchical dataset of diverse student activities collected by Santa, a multi-platform self-study solution equipped with an artificial intelligence tutoring system. EdNet contains 131,417,236 interactions from 784,309 students collected over more than 2 years, making it the largest public IES dataset released to date\".</p>\n<h2>More papers from Riiid</h2>\n<p><a href=\"https://arxiv.org/abs/2005.03818\" target=\"_blank\">Choose Your Own Question: Encouraging Self-Personalization in Learning Path Construction</a> - In the case of the existing learning model in which AI recommends optimal customized learning content, students passively take on recommended questions or lecture content without judgment, which makes it difficult for students to identify their learning paths, and makes it difficult for them to evaluate each question or the content. Riiid give users control, thereby increasing their learning immersion and enables users to identify and select one’s path.</p>\n<p><a href=\"https://arxiv.org/abs/2002.05505\" target=\"_blank\">Assessment Modeling: Fundamental Pre-training Tasks for Interactive Educational Systems</a> - here Riiid proposes a deep learning Transformer-based assessment model to overcome this limitation and improve predictive accuracy with sparse data. A model is developed that pre-trains a user’s probability in making correct/incorrect answers and in-time problem solving. The model is then fine-tuned to match score predictions based on small amounts of score data.</p>\n<p><a href=\"https://arxiv.org/abs/2002.11624\" target=\"_blank\">Deep Attentive Study Session Dropout Prediction in Mobile Learning Environment</a> - Analyzing a relatively short session-based mobile learning environment, Riiid developed a deep learning Transformer-based predictive model called Deep Attributive Study Session Description (DAS). The model defines the concept of user ‘dropout’ for the first time and accurately predicts dropout rates by examining various learning behaviors of users.</p>\n<p><a href=\"https://arxiv.org/abs/2002.07033\" target=\"_blank\">Towards an Appropriate Query, Key, and Value Computation for Knowledge Tracing</a> - Applies the Transformer deep-learning model, mainly used in natural language processing, to Knowledge Tracing. The authors tuned  the Transformer model to fit the education domain with Riiid’s Separated Self-Attentive Neural Knowledge Tracing (SAINT) model.</p>\n<h2>EdNet github project</h2>\n<p>EdNet is made available through a GitHub project, here: <a href=\"https://github.com/riiid/ednet\" target=\"_blank\">https://github.com/riiid/ednet</a></p>\n<p>There are 4 datasets, as following:</p>\n<ul>\n<li>EdNet-KT1 -  students' question-solving logs;  </li>\n<li>EdNet-KT2 - action sequences of each user are compiled; action types can be enter, respond and submit;  </li>\n<li>EdNet-KT3 - include additional learning activities by students, such as reading through experts' commentary on a question or watching lectures provided by the system;  </li>\n<li>EdNet-KT4 - additional actions by students are added, including, besides what is existent in EdNet-KT3: erase_choice, undo_erase_choice, play_audio, pause_audio, play_video, pause_video, pay, refund, and enroll_coupon.</li>\n</ul>",
  "messages": [
    {
      "id": "1041031",
      "postDate": "10/07/2020 14:06:59",
      "content": "<h2>Riiid company</h2>\n<p><img src=\"https://prnewswire2-a.akamaihd.net/p/1893751/sp/189375100/thumbnail/entry_id/0_u0ia3ust/def_height/840/def_width/1400/version/100012/type/1\"></p>\n<p>Riiid was founded in 2014. In 2016, the company pioneered the patented AI-based education improvement algorithm and launched 'Santa for TOEIC!' beta version (final version launched next year). Santa is a mobile test prep application for the popular English proficiency exam, Test of English for International Communication (TOEIC). Santa  reached No. 1 in sales among education apps in Japan and Korea. Riiid’s proprietary AI technology analyzes student data and content, predicts scores and user behavior, and recommends personalized study plans to help students optimize their learning potential. In 2019 launched Santa SAT and in 2020 EdNet, the largest open database for AI Education.</p>\n<h2>EdNet arXiv paper</h2>\n<p>The original paper on EdNet, published on arXiv: <a href=\"https://arxiv.org/pdf/1912.03072.pdf\" target=\"_blank\"><strong>EdNet: A Large-Scale Hierarchical Dataset in\nEducation</strong></a>; excerpt from the Abstract: \"EdNet, a large-scale hierarchical dataset of diverse student activities collected by Santa, a multi-platform self-study solution equipped with an artificial intelligence tutoring system. EdNet contains 131,417,236 interactions from 784,309 students collected over more than 2 years, making it the largest public IES dataset released to date\".</p>\n<h2>More papers from Riiid</h2>\n<p><a href=\"https://arxiv.org/abs/2005.03818\" target=\"_blank\">Choose Your Own Question: Encouraging Self-Personalization in Learning Path Construction</a> - In the case of the existing learning model in which AI recommends optimal customized learning content, students passively take on recommended questions or lecture content without judgment, which makes it difficult for students to identify their learning paths, and makes it difficult for them to evaluate each question or the content. Riiid give users control, thereby increasing their learning immersion and enables users to identify and select one’s path.</p>\n<p><a href=\"https://arxiv.org/abs/2002.05505\" target=\"_blank\">Assessment Modeling: Fundamental Pre-training Tasks for Interactive Educational Systems</a> - here Riiid proposes a deep learning Transformer-based assessment model to overcome this limitation and improve predictive accuracy with sparse data. A model is developed that pre-trains a user’s probability in making correct/incorrect answers and in-time problem solving. The model is then fine-tuned to match score predictions based on small amounts of score data.</p>\n<p><a href=\"https://arxiv.org/abs/2002.11624\" target=\"_blank\">Deep Attentive Study Session Dropout Prediction in Mobile Learning Environment</a> - Analyzing a relatively short session-based mobile learning environment, Riiid developed a deep learning Transformer-based predictive model called Deep Attributive Study Session Description (DAS). The model defines the concept of user ‘dropout’ for the first time and accurately predicts dropout rates by examining various learning behaviors of users.</p>\n<p><a href=\"https://arxiv.org/abs/2002.07033\" target=\"_blank\">Towards an Appropriate Query, Key, and Value Computation for Knowledge Tracing</a> - Applies the Transformer deep-learning model, mainly used in natural language processing, to Knowledge Tracing. The authors tuned  the Transformer model to fit the education domain with Riiid’s Separated Self-Attentive Neural Knowledge Tracing (SAINT) model.</p>\n<h2>EdNet github project</h2>\n<p>EdNet is made available through a GitHub project, here: <a href=\"https://github.com/riiid/ednet\" target=\"_blank\">https://github.com/riiid/ednet</a></p>\n<p>There are 4 datasets, as following:</p>\n<ul>\n<li>EdNet-KT1 -  students' question-solving logs;  </li>\n<li>EdNet-KT2 - action sequences of each user are compiled; action types can be enter, respond and submit;  </li>\n<li>EdNet-KT3 - include additional learning activities by students, such as reading through experts' commentary on a question or watching lectures provided by the system;  </li>\n<li>EdNet-KT4 - additional actions by students are added, including, besides what is existent in EdNet-KT3: erase_choice, undo_erase_choice, play_audio, pause_audio, play_video, pause_video, pay, refund, and enroll_coupon.</li>\n</ul>",
      "rawMarkdown": "## Riiid company\n\n<img src=\"https://prnewswire2-a.akamaihd.net/p/1893751/sp/189375100/thumbnail/entry_id/0_u0ia3ust/def_height/840/def_width/1400/version/100012/type/1\" width=\"600\"></img>\n\n\nRiiid was founded in 2014. In 2016, the company pioneered the patented AI-based education improvement algorithm and launched 'Santa for TOEIC!' beta version (final version launched next year). Santa is a mobile test prep application for the popular English proficiency exam, Test of English for International Communication (TOEIC). Santa  reached No. 1 in sales among education apps in Japan and Korea. Riiid’s proprietary AI technology analyzes student data and content, predicts scores and user behavior, and recommends personalized study plans to help students optimize their learning potential. In 2019 launched Santa SAT and in 2020 EdNet, the largest open database for AI Education.\n\n\n## EdNet arXiv paper\n\nThe original paper on EdNet, published on arXiv: [**EdNet: A Large-Scale Hierarchical Dataset in\nEducation**](https://arxiv.org/pdf/1912.03072.pdf); excerpt from the Abstract: \"EdNet, a large-scale hierarchical dataset of diverse student activities collected by Santa, a multi-platform self-study solution equipped with an artificial intelligence tutoring system. EdNet contains 131,417,236 interactions from 784,309 students collected over more than 2 years, making it the largest public IES dataset released to date\".\n\n## More papers from Riiid\n\n[Choose Your Own Question: Encouraging Self-Personalization in Learning Path Construction](https://arxiv.org/abs/2005.03818) - In the case of the existing learning model in which AI recommends optimal customized learning content, students passively take on recommended questions or lecture content without judgment, which makes it difficult for students to identify their learning paths, and makes it difficult for them to evaluate each question or the content. Riiid give users control, thereby increasing their learning immersion and enables users to identify and select one’s path.\n\n[Assessment Modeling: Fundamental Pre-training Tasks for Interactive Educational Systems](https://arxiv.org/abs/2002.05505) - here Riiid proposes a deep learning Transformer-based assessment model to overcome this limitation and improve predictive accuracy with sparse data. A model is developed that pre-trains a user’s probability in making correct/incorrect answers and in-time problem solving. The model is then fine-tuned to match score predictions based on small amounts of score data.\n\n[Deep Attentive Study Session Dropout Prediction in Mobile Learning Environment](https://arxiv.org/abs/2002.11624) - Analyzing a relatively short session-based mobile learning environment, Riiid developed a deep learning Transformer-based predictive model called Deep Attributive Study Session Description (DAS). The model defines the concept of user ‘dropout’ for the first time and accurately predicts dropout rates by examining various learning behaviors of users.\n\n[Towards an Appropriate Query, Key, and Value Computation for Knowledge Tracing](https://arxiv.org/abs/2002.07033) - Applies the Transformer deep-learning model, mainly used in natural language processing, to Knowledge Tracing. The authors tuned  the Transformer model to fit the education domain with Riiid’s Separated Self-Attentive Neural Knowledge Tracing (SAINT) model.\n\n## EdNet github project \n\nEdNet is made available through a GitHub project, here: https://github.com/riiid/ednet\n\nThere are 4 datasets, as following:\n\n* EdNet-KT1 -  students' question-solving logs;  \n* EdNet-KT2 - action sequences of each user are compiled; action types can be enter, respond and submit;  \n* EdNet-KT3 - include additional learning activities by students, such as reading through experts' commentary on a question or watching lectures provided by the system;  \n* EdNet-KT4 - additional actions by students are added, including, besides what is existent in EdNet-KT3: erase_choice, undo_erase_choice, play_audio, pause_audio, play_video, pause_video, pay, refund, and enroll_coupon.",
      "votes": null
    },
    {
      "id": "1041078",
      "postDate": "10/07/2020 14:43:13",
      "content": "<p>Already shared here: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188901\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188901</a></p>",
      "rawMarkdown": "Already shared here: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188901",
      "votes": null
    },
    {
      "id": "1041081",
      "postDate": "10/07/2020 14:46:31",
      "content": "<p>Uhgh. I guess I better read the topics before posting one new. Thank you for pointing out.</p>",
      "rawMarkdown": "Uhgh. I guess I better read the topics before posting one new. Thank you for pointing out.",
      "votes": null
    },
    {
      "id": "1049225",
      "postDate": "10/14/2020 08:13:28",
      "content": "<p>Which one is of the above dataset is provided to us?</p>",
      "rawMarkdown": "Which one is of the above dataset is provided to us?",
      "votes": null
    },
    {
      "id": "1049492",
      "postDate": "10/14/2020 13:25:00",
      "content": "<p>EdNet represents the base for the data shared here.</p>",
      "rawMarkdown": "EdNet represents the base for the data shared here.",
      "votes": null
    },
    {
      "id": "1105765",
      "postDate": "12/08/2020 07:21:51",
      "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>  i got stopped here while searching for more data as i read in this kernel comments<br>\n<a href=\"https://www.kaggle.com/mpware/sakt-fork?scriptVersionId=48197140\" target=\"_blank\">https://www.kaggle.com/mpware/sakt-fork?scriptVersionId=48197140</a></p>\n<pre><code>That's good to know. I got my fork up to 0.7793 last night on validation. That's with 1422046 \"users\". It's possible that's leaking though. I'm going to change my validation set today to only include the most recent sequences. I'm also trying to change the TestDataset so it can update user sequences as testing progresses. I imagine that would have a huge impact at inference time.\n\nEdit: Indeed some leaky business going on there.\n</code></pre>\n<p>Was wondering from where have gotten these many users when  I get only 323456 no of users in total .</p>",
      "rawMarkdown": "rohanrao  i got stopped here while searching for more data as i read in this kernel comments\nhttps://www.kaggle.com/mpware/sakt-fork?scriptVersionId=48197140\n```\nThat's good to know. I got my fork up to 0.7793 last night on validation. That's with 1422046 \"users\". It's possible that's leaking though. I'm going to change my validation set today to only include the most recent sequences. I'm also trying to change the TestDataset so it can update user sequences as testing progresses. I imagine that would have a huge impact at inference time.\n\nEdit: Indeed some leaky business going on there.\n```\n\nWas wondering from where have gotten these many users when  I get only 323456 no of users in total .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1041078,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "10/07/2020 14:43:13",
      "content": "<p>Already shared here: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188901\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188901</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1041081,
          "author_name": "gpreda",
          "author_url": "",
          "post_date": "10/07/2020 14:46:31",
          "content": "<p>Uhgh. I guess I better read the topics before posting one new. Thank you for pointing out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105765,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/08/2020 07:21:51",
          "content": "<p><a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>  i got stopped here while searching for more data as i read in this kernel comments<br>\n<a href=\"https://www.kaggle.com/mpware/sakt-fork?scriptVersionId=48197140\" target=\"_blank\">https://www.kaggle.com/mpware/sakt-fork?scriptVersionId=48197140</a></p>\n<pre><code>That's good to know. I got my fork up to 0.7793 last night on validation. That's with 1422046 \"users\". It's possible that's leaking though. I'm going to change my validation set today to only include the most recent sequences. I'm also trying to change the TestDataset so it can update user sequences as testing progresses. I imagine that would have a huge impact at inference time.\n\nEdit: Indeed some leaky business going on there.\n</code></pre>\n<p>Was wondering from where have gotten these many users when  I get only 323456 no of users in total .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1049225,
      "author_name": "aravindpadman",
      "author_url": "",
      "post_date": "10/14/2020 08:13:28",
      "content": "<p>Which one is of the above dataset is provided to us?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1049492,
          "author_name": "gpreda",
          "author_url": "",
          "post_date": "10/14/2020 13:25:00",
          "content": "<p>EdNet represents the base for the data shared here.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1041031": "## Riiid company\n\n<img src=\"https://prnewswire2-a.akamaihd.net/p/1893751/sp/189375100/thumbnail/entry_id/0_u0ia3ust/def_height/840/def_width/1400/version/100012/type/1\" width=\"600\"></img>\n\n\nRiiid was founded in 2014. In 2016, the company pioneered the patented AI-based education improvement algorithm and launched 'Santa for TOEIC!' beta version (final version launched next year). Santa is a mobile test prep application for the popular English proficiency exam, Test of English for International Communication (TOEIC). Santa  reached No. 1 in sales among education apps in Japan and Korea. Riiid’s proprietary AI technology analyzes student data and content, predicts scores and user behavior, and recommends personalized study plans to help students optimize their learning potential. In 2019 launched Santa SAT and in 2020 EdNet, the largest open database for AI Education.\n\n\n## EdNet arXiv paper\n\nThe original paper on EdNet, published on arXiv: [**EdNet: A Large-Scale Hierarchical Dataset in\nEducation**](https://arxiv.org/pdf/1912.03072.pdf); excerpt from the Abstract: \"EdNet, a large-scale hierarchical dataset of diverse student activities collected by Santa, a multi-platform self-study solution equipped with an artificial intelligence tutoring system. EdNet contains 131,417,236 interactions from 784,309 students collected over more than 2 years, making it the largest public IES dataset released to date\".\n\n## More papers from Riiid\n\n[Choose Your Own Question: Encouraging Self-Personalization in Learning Path Construction](https://arxiv.org/abs/2005.03818) - In the case of the existing learning model in which AI recommends optimal customized learning content, students passively take on recommended questions or lecture content without judgment, which makes it difficult for students to identify their learning paths, and makes it difficult for them to evaluate each question or the content. Riiid give users control, thereby increasing their learning immersion and enables users to identify and select one’s path.\n\n[Assessment Modeling: Fundamental Pre-training Tasks for Interactive Educational Systems](https://arxiv.org/abs/2002.05505) - here Riiid proposes a deep learning Transformer-based assessment model to overcome this limitation and improve predictive accuracy with sparse data. A model is developed that pre-trains a user’s probability in making correct/incorrect answers and in-time problem solving. The model is then fine-tuned to match score predictions based on small amounts of score data.\n\n[Deep Attentive Study Session Dropout Prediction in Mobile Learning Environment](https://arxiv.org/abs/2002.11624) - Analyzing a relatively short session-based mobile learning environment, Riiid developed a deep learning Transformer-based predictive model called Deep Attributive Study Session Description (DAS). The model defines the concept of user ‘dropout’ for the first time and accurately predicts dropout rates by examining various learning behaviors of users.\n\n[Towards an Appropriate Query, Key, and Value Computation for Knowledge Tracing](https://arxiv.org/abs/2002.07033) - Applies the Transformer deep-learning model, mainly used in natural language processing, to Knowledge Tracing. The authors tuned  the Transformer model to fit the education domain with Riiid’s Separated Self-Attentive Neural Knowledge Tracing (SAINT) model.\n\n## EdNet github project \n\nEdNet is made available through a GitHub project, here: https://github.com/riiid/ednet\n\nThere are 4 datasets, as following:\n\n* EdNet-KT1 -  students' question-solving logs;  \n* EdNet-KT2 - action sequences of each user are compiled; action types can be enter, respond and submit;  \n* EdNet-KT3 - include additional learning activities by students, such as reading through experts' commentary on a question or watching lectures provided by the system;  \n* EdNet-KT4 - additional actions by students are added, including, besides what is existent in EdNet-KT3: erase_choice, undo_erase_choice, play_audio, pause_audio, play_video, pause_video, pay, refund, and enroll_coupon.",
    "1041078": "Already shared here: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188901",
    "1041081": "Uhgh. I guess I better read the topics before posting one new. Thank you for pointing out.",
    "1049225": "Which one is of the above dataset is provided to us?",
    "1049492": "EdNet represents the base for the data shared here.",
    "1105765": "rohanrao  i got stopped here while searching for more data as i read in this kernel comments\nhttps://www.kaggle.com/mpware/sakt-fork?scriptVersionId=48197140\n```\nThat's good to know. I got my fork up to 0.7793 last night on validation. That's with 1422046 \"users\". It's possible that's leaking though. I'm going to change my validation set today to only include the most recent sequences. I'm also trying to change the TestDataset so it can update user sequences as testing progresses. I imagine that would have a huge impact at inference time.\n\nEdit: Indeed some leaky business going on there.\n```\n\nWas wondering from where have gotten these many users when  I get only 323456 no of users in total ."
  },
  "source": "meta"
}