{
  "id": 188913,
  "title": "How to ensure the unbiased and ethical use of the dataset?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/188913",
  "author_name": "",
  "post_date": "2020-10-05T20:12:55.914319800Z",
  "votes": 47,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Thank you to Riiid labs and Kaggle for organizing this competition. AI for education is a very interesting subject and it's great to see more of this on Kaggle. I am also sure that the results of this competition will help many students across the world.</p>\n<p>However, I have a question and I'd love to hear the organizers' thoughts: <strong>How are you planning to handle potential biases that could arise when generating personalized curriculum?</strong> </p>\n<p>Students across the world have varying strengths and weaknesses. Some will be stronger in the arts, other in the sciences. Some will have learning disabilities and personal issues that could hinder their progression. My understanding of personalized learning (please correct me if I'm wrong) is that progressively more advanced content will be proposed based on the predicted \"knowledge\" level of the student. However, I'm curious whether models that were built using the dataset from <a href=\"https://arxiv.org/pdf/1912.03072.pdf\" target=\"_blank\">Choi et al., 2019</a> and the models proposed in <a href=\"https://arxiv.org/pdf/2002.05505.pdf\" target=\"_blank\">Choi et al., 2020</a> are able to provide an accurate estimate of the progress even when considering factors that could unintentionally bias the model, as well as recommend learning material based on the preferences of the student?</p>\n<p>A naive example would be that of learning English by watching short videos and answering questions. The subject of the video could be about popular music and art, or it could be about various scientific research across the world. Obviously, a student that is more interested by arts will pay more attention for the former than the latter and potentially score lower when given the latter. Would the model be able to differentiate a lack of progress with a lack of interest, and give an unbiased estimate of the student's abilities and recommend more interesting content?</p>\n<p>I think it'd be very interesting to learn more about how bias and ethical usage are built-in Riiid models and products. For example, a Google-style <a href=\"https://research.google/pubs/pub48120/\" target=\"_blank\">model card</a> would be an amazing way to approach this.</p>\n<p>Again, thanks for organizing this competition and hearing me out!</p>",
  "messages": [
    {
      "id": "1038484",
      "postDate": "10/05/2020 20:12:55",
      "content": "<p>Thank you to Riiid labs and Kaggle for organizing this competition. AI for education is a very interesting subject and it's great to see more of this on Kaggle. I am also sure that the results of this competition will help many students across the world.</p>\n<p>However, I have a question and I'd love to hear the organizers' thoughts: <strong>How are you planning to handle potential biases that could arise when generating personalized curriculum?</strong> </p>\n<p>Students across the world have varying strengths and weaknesses. Some will be stronger in the arts, other in the sciences. Some will have learning disabilities and personal issues that could hinder their progression. My understanding of personalized learning (please correct me if I'm wrong) is that progressively more advanced content will be proposed based on the predicted \"knowledge\" level of the student. However, I'm curious whether models that were built using the dataset from <a href=\"https://arxiv.org/pdf/1912.03072.pdf\" target=\"_blank\">Choi et al., 2019</a> and the models proposed in <a href=\"https://arxiv.org/pdf/2002.05505.pdf\" target=\"_blank\">Choi et al., 2020</a> are able to provide an accurate estimate of the progress even when considering factors that could unintentionally bias the model, as well as recommend learning material based on the preferences of the student?</p>\n<p>A naive example would be that of learning English by watching short videos and answering questions. The subject of the video could be about popular music and art, or it could be about various scientific research across the world. Obviously, a student that is more interested by arts will pay more attention for the former than the latter and potentially score lower when given the latter. Would the model be able to differentiate a lack of progress with a lack of interest, and give an unbiased estimate of the student's abilities and recommend more interesting content?</p>\n<p>I think it'd be very interesting to learn more about how bias and ethical usage are built-in Riiid models and products. For example, a Google-style <a href=\"https://research.google/pubs/pub48120/\" target=\"_blank\">model card</a> would be an amazing way to approach this.</p>\n<p>Again, thanks for organizing this competition and hearing me out!</p>",
      "rawMarkdown": "Thank you to Riiid labs and Kaggle for organizing this competition. AI for education is a very interesting subject and it's great to see more of this on Kaggle. I am also sure that the results of this competition will help many students across the world.\n\nHowever, I have a question and I'd love to hear the organizers' thoughts: **How are you planning to handle potential biases that could arise when generating personalized curriculum?** \n\nStudents across the world have varying strengths and weaknesses. Some will be stronger in the arts, other in the sciences. Some will have learning disabilities and personal issues that could hinder their progression. My understanding of personalized learning (please correct me if I'm wrong) is that progressively more advanced content will be proposed based on the predicted \"knowledge\" level of the student. However, I'm curious whether models that were built using the dataset from [Choi et al., 2019](https://arxiv.org/pdf/1912.03072.pdf) and the models proposed in [Choi et al., 2020](https://arxiv.org/pdf/2002.05505.pdf) are able to provide an accurate estimate of the progress even when considering factors that could unintentionally bias the model, as well as recommend learning material based on the preferences of the student?\n\nA naive example would be that of learning English by watching short videos and answering questions. The subject of the video could be about popular music and art, or it could be about various scientific research across the world. Obviously, a student that is more interested by arts will pay more attention for the former than the latter and potentially score lower when given the latter. Would the model be able to differentiate a lack of progress with a lack of interest, and give an unbiased estimate of the student's abilities and recommend more interesting content?\n\nI think it'd be very interesting to learn more about how bias and ethical usage are built-in Riiid models and products. For example, a Google-style [model card](https://research.google/pubs/pub48120/) would be an amazing way to approach this.\n\nAgain, thanks for organizing this competition and hearing me out!",
      "votes": null
    },
    {
      "id": "1039777",
      "postDate": "10/06/2020 19:31:04",
      "content": "<p>Well I'm guessing that there are a lot of complications with potentially personally-identifiable data for the contest, you know like with HIPAA / PIPEDA for healthcare stuff, specifically since the ids in this contest are kids.<br>\nBut for sure its really important that they take this stuff into account when they engineer their own stuff for commercial use! </p>",
      "rawMarkdown": "Well I'm guessing that there are a lot of complications with potentially personally-identifiable data for the contest, you know like with HIPAA / PIPEDA for healthcare stuff, specifically since the ids in this contest are kids.\nBut for sure its really important that they take this stuff into account when they engineer their own stuff for commercial use!",
      "votes": null
    },
    {
      "id": "1041781",
      "postDate": "10/07/2020 23:05:27",
      "content": "<p>This is a really good point!</p>",
      "rawMarkdown": "This is a really good point!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1039777,
      "author_name": "abhaykatoch",
      "author_url": "",
      "post_date": "10/06/2020 19:31:04",
      "content": "<p>Well I'm guessing that there are a lot of complications with potentially personally-identifiable data for the contest, you know like with HIPAA / PIPEDA for healthcare stuff, specifically since the ids in this contest are kids.<br>\nBut for sure its really important that they take this stuff into account when they engineer their own stuff for commercial use! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1041781,
      "author_name": "ashwathraj",
      "author_url": "",
      "post_date": "10/07/2020 23:05:27",
      "content": "<p>This is a really good point!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1038484": "Thank you to Riiid labs and Kaggle for organizing this competition. AI for education is a very interesting subject and it's great to see more of this on Kaggle. I am also sure that the results of this competition will help many students across the world.\n\nHowever, I have a question and I'd love to hear the organizers' thoughts: **How are you planning to handle potential biases that could arise when generating personalized curriculum?** \n\nStudents across the world have varying strengths and weaknesses. Some will be stronger in the arts, other in the sciences. Some will have learning disabilities and personal issues that could hinder their progression. My understanding of personalized learning (please correct me if I'm wrong) is that progressively more advanced content will be proposed based on the predicted \"knowledge\" level of the student. However, I'm curious whether models that were built using the dataset from [Choi et al., 2019](https://arxiv.org/pdf/1912.03072.pdf) and the models proposed in [Choi et al., 2020](https://arxiv.org/pdf/2002.05505.pdf) are able to provide an accurate estimate of the progress even when considering factors that could unintentionally bias the model, as well as recommend learning material based on the preferences of the student?\n\nA naive example would be that of learning English by watching short videos and answering questions. The subject of the video could be about popular music and art, or it could be about various scientific research across the world. Obviously, a student that is more interested by arts will pay more attention for the former than the latter and potentially score lower when given the latter. Would the model be able to differentiate a lack of progress with a lack of interest, and give an unbiased estimate of the student's abilities and recommend more interesting content?\n\nI think it'd be very interesting to learn more about how bias and ethical usage are built-in Riiid models and products. For example, a Google-style [model card](https://research.google/pubs/pub48120/) would be an amazing way to approach this.\n\nAgain, thanks for organizing this competition and hearing me out!",
    "1039777": "Well I'm guessing that there are a lot of complications with potentially personally-identifiable data for the contest, you know like with HIPAA / PIPEDA for healthcare stuff, specifically since the ids in this contest are kids.\nBut for sure its really important that they take this stuff into account when they engineer their own stuff for commercial use!",
    "1041781": "This is a really good point!"
  },
  "source": "meta"
}