{
  "id": 209586,
  "title": "[0.801 Private, 60th place] Single LGB using 1% of data",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209586",
  "author_name": "Ming Pan",
  "post_date": "2021-01-08T00:25:29.540000",
  "votes": 77,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Thanks to the organizers of this very interesting competition and congrats to the winners. My LB score (0.800 public / 0.801 private) was achieved using a single LGB model trained on 1m rows. It wasn't my intention to only use 1m rows, but I had some bugs that I fixed too late to retrain on more data in time. Nevertheless, I'm sharing my solution in case it ends up being useful to anyone.</p>\n<h1>Features</h1>\n<p>The final model used 220 features that fit into following categories in rough descending order of importance:<br>\nnote: AC = answered_correctly</p>\n<ul>\n<li>Question-Question matrix (QQM) features: out of all users who answered A on question X, how many of them get question Y right? I built a 13523 x 13523 x 4 matrix of all the probabilities, then calculated user stats on their last 3, 10, 30, 50, 100, 200 questions. I also created a separate set of features weighting by Mahalanobis distance between the past and current question embeddings (see section on PEBG embeddings). </li>\n<li>question's overall accuracy/difficulty across all users, number of responses</li>\n<li>user's accuracy, number of respones across their entire history &amp; current session (where a new session is created if the user is idle for more than 5 min)</li>\n<li>question metadata, i.e. tags, part, index in bundle</li>\n<li>elapsed time and lag time - based features. Ratios were good here, e.g. ratio of prior elapsed time to lag time, user's average elapsed time to average lag time. I also made a simple model to predict elapsed time of the current question</li>\n<li>PEBG embedding-based features based on this paper: <a href=\"https://arxiv.org/pdf/2012.05031.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.05031.pdf</a>. I took the current question embedding, the user's average embedding overall, and the average embedding over the user's correct and incorrect responses. Mahalanobis distance between the current question's embedding to the user's correct and incorrect embeddings and the ratio between correct and incorrect distance were all important. </li>\n<li>features from Ednet dataset (see below)</li>\n<li>stats on user's past attempts on the same question, e.g. most recent AC, average AC, number of past attempts, time since last attempt, whether user saw explanation on last attempt</li>\n<li>stats on user's past attempts on questions with the same tags</li>\n<li>user's accuracy normalized by question difficulty across different windows, e.g. average of user's residuals for their last 3, 10, 30, 50, 100, 200 questions; residual = AC - question accuracy</li>\n<li>user's accuracy trend over their last 4, 12, 30, 50, 100, 200 questions</li>\n<li>user's performance on diagnostic questions: the 30 question sequence starting with 7900, 7876, 175, … is usually asked to each user at the beginning. I made 30 features for the user's AC on each question</li>\n<li>lecture features: time since user watched a lecture with the same tag, time spent watching lecture compared to lecture duration (from Ednet data)</li>\n</ul>\n<p>I applied additive smoothing to target encoding features to guard against small samples (this noticeably helped the QQM features). </p>\n<h1>Ednet dataset</h1>\n<p>As you probably know, the competition data was taken from Riiid's Ednet database, from which a large public dataset has been open sourced at <a href=\"https://github.com/riiid/ednet\" target=\"_blank\">https://github.com/riiid/ednet</a>. At first glance, the overlap between the competition data and the open source data is not obvious - for example, the open source data has 784,309 users whereas the competition data only has 393,656 and the contents don't look the same either. But the Ednet data comes with question and lecture metadata files, and by comparing them to the metadata files we have in the competition, it's quite easy to map the questions, tags, and users between the datasets. The competition data seems to be a bit more recent, and has mostly filtered out users who only answered a few questions, but 350k+ out of ~393k users appear in both datasets. That's important because most of the private test data is from users appearing in the training set, so you can use the Ednet data to calculate some features on those users that you wouldn't otherwise have. Examples are:</p>\n<ul>\n<li>Ednet data has real timestamps, not shifted to start from zero. I calculated the hour, day, and week in KST</li>\n<li>how often the user switches answers</li>\n<li>Ground truth video lengths for lectures, which can be used in lecture features</li>\n<li>whether user is primarily on mobile or web</li>\n</ul>\n<p>Also, I extracted ~6m rows from the Ednet data that was filtered out of the competition data. These were mostly users with few interactions so I didn't add them to the training set, but I did use them to pre-populate the QQM which helped a bit. </p>\n<h1>Pipeline &amp; Remarks</h1>\n<p>I used a loop-based accumulation framework which updated user states every iteration. To strike a good balance between speed and memory usage, I used dictionaries to look up users and numpy arrays to store the accumulated features. This worked well in general and I could use the same code for inference, which took ~3 hours in the end with 220 features and used about 9GB RAM.</p>\n<p>In the end, I wasn't able to train on the entire dataset, or even 10m rows due to bugs that weren't fixed until the morning of the deadline. One particularly frustrating one was from the very first piece of code I ran in the competition: the code to read in the train.csv file. I had copied a line to specify the data types from a public notebook assuming it would be the same one from the Competition Starter, but the notebook creator had changed prior_question_elapsed_time to float16 instead of float32. This resulted in all values in the column &gt;=65536 being converted to np.inf. Lesson learned: never blindly copy from a public notebook! </p>\n<p>Despite this, I really enjoyed the competition as by taking the feature engineering approach, it felt like I was uncovering new things about the data all the time. </p>",
  "messages": [
    {
      "id": 1143557,
      "postDate": "2021-01-08T00:25:29.540Z",
      "content": "<p>Thanks to the organizers of this very interesting competition and congrats to the winners. My LB score (0.800 public / 0.801 private) was achieved using a single LGB model trained on 1m rows. It wasn't my intention to only use 1m rows, but I had some bugs that I fixed too late to retrain on more data in time. Nevertheless, I'm sharing my solution in case it ends up being useful to anyone.</p>\n<h1>Features</h1>\n<p>The final model used 220 features that fit into following categories in rough descending order of importance:<br>\nnote: AC = answered_correctly</p>\n<ul>\n<li>Question-Question matrix (QQM) features: out of all users who answered A on question X, how many of them get question Y right? I built a 13523 x 13523 x 4 matrix of all the probabilities, then calculated user stats on their last 3, 10, 30, 50, 100, 200 questions. I also created a separate set of features weighting by Mahalanobis distance between the past and current question embeddings (see section on PEBG embeddings). </li>\n<li>question's overall accuracy/difficulty across all users, number of responses</li>\n<li>user's accuracy, number of respones across their entire history &amp; current session (where a new session is created if the user is idle for more than 5 min)</li>\n<li>question metadata, i.e. tags, part, index in bundle</li>\n<li>elapsed time and lag time - based features. Ratios were good here, e.g. ratio of prior elapsed time to lag time, user's average elapsed time to average lag time. I also made a simple model to predict elapsed time of the current question</li>\n<li>PEBG embedding-based features based on this paper: <a href=\"https://arxiv.org/pdf/2012.05031.pdf\" target=\"_blank\">https://arxiv.org/pdf/2012.05031.pdf</a>. I took the current question embedding, the user's average embedding overall, and the average embedding over the user's correct and incorrect responses. Mahalanobis distance between the current question's embedding to the user's correct and incorrect embeddings and the ratio between correct and incorrect distance were all important. </li>\n<li>features from Ednet dataset (see below)</li>\n<li>stats on user's past attempts on the same question, e.g. most recent AC, average AC, number of past attempts, time since last attempt, whether user saw explanation on last attempt</li>\n<li>stats on user's past attempts on questions with the same tags</li>\n<li>user's accuracy normalized by question difficulty across different windows, e.g. average of user's residuals for their last 3, 10, 30, 50, 100, 200 questions; residual = AC - question accuracy</li>\n<li>user's accuracy trend over their last 4, 12, 30, 50, 100, 200 questions</li>\n<li>user's performance on diagnostic questions: the 30 question sequence starting with 7900, 7876, 175, … is usually asked to each user at the beginning. I made 30 features for the user's AC on each question</li>\n<li>lecture features: time since user watched a lecture with the same tag, time spent watching lecture compared to lecture duration (from Ednet data)</li>\n</ul>\n<p>I applied additive smoothing to target encoding features to guard against small samples (this noticeably helped the QQM features). </p>\n<h1>Ednet dataset</h1>\n<p>As you probably know, the competition data was taken from Riiid's Ednet database, from which a large public dataset has been open sourced at <a href=\"https://github.com/riiid/ednet\" target=\"_blank\">https://github.com/riiid/ednet</a>. At first glance, the overlap between the competition data and the open source data is not obvious - for example, the open source data has 784,309 users whereas the competition data only has 393,656 and the contents don't look the same either. But the Ednet data comes with question and lecture metadata files, and by comparing them to the metadata files we have in the competition, it's quite easy to map the questions, tags, and users between the datasets. The competition data seems to be a bit more recent, and has mostly filtered out users who only answered a few questions, but 350k+ out of ~393k users appear in both datasets. That's important because most of the private test data is from users appearing in the training set, so you can use the Ednet data to calculate some features on those users that you wouldn't otherwise have. Examples are:</p>\n<ul>\n<li>Ednet data has real timestamps, not shifted to start from zero. I calculated the hour, day, and week in KST</li>\n<li>how often the user switches answers</li>\n<li>Ground truth video lengths for lectures, which can be used in lecture features</li>\n<li>whether user is primarily on mobile or web</li>\n</ul>\n<p>Also, I extracted ~6m rows from the Ednet data that was filtered out of the competition data. These were mostly users with few interactions so I didn't add them to the training set, but I did use them to pre-populate the QQM which helped a bit. </p>\n<h1>Pipeline &amp; Remarks</h1>\n<p>I used a loop-based accumulation framework which updated user states every iteration. To strike a good balance between speed and memory usage, I used dictionaries to look up users and numpy arrays to store the accumulated features. This worked well in general and I could use the same code for inference, which took ~3 hours in the end with 220 features and used about 9GB RAM.</p>\n<p>In the end, I wasn't able to train on the entire dataset, or even 10m rows due to bugs that weren't fixed until the morning of the deadline. One particularly frustrating one was from the very first piece of code I ran in the competition: the code to read in the train.csv file. I had copied a line to specify the data types from a public notebook assuming it would be the same one from the Competition Starter, but the notebook creator had changed prior_question_elapsed_time to float16 instead of float32. This resulted in all values in the column &gt;=65536 being converted to np.inf. Lesson learned: never blindly copy from a public notebook! </p>\n<p>Despite this, I really enjoyed the competition as by taking the feature engineering approach, it felt like I was uncovering new things about the data all the time. </p>",
      "rawMarkdown": "Thanks to the organizers of this very interesting competition and congrats to the winners. My LB score (0.800 public / 0.801 private) was achieved using a single LGB model trained on 1m rows. It wasn't my intention to only use 1m rows, but I had some bugs that I fixed too late to retrain on more data in time. Nevertheless, I'm sharing my solution in case it ends up being useful to anyone.\n\n# Features\nThe final model used 220 features that fit into following categories in rough descending order of importance:\nnote: AC = answered_correctly\n- Question-Question matrix (QQM) features: out of all users who answered A on question X, how many of them get question Y right? I built a 13523 x 13523 x 4 matrix of all the probabilities, then calculated user stats on their last 3, 10, 30, 50, 100, 200 questions. I also created a separate set of features weighting by Mahalanobis distance between the past and current question embeddings (see section on PEBG embeddings). \n- question's overall accuracy/difficulty across all users, number of responses\n- user's accuracy, number of respones across their entire history & current session (where a new session is created if the user is idle for more than 5 min)\n- question metadata, i.e. tags, part, index in bundle\n- elapsed time and lag time - based features. Ratios were good here, e.g. ratio of prior elapsed time to lag time, user's average elapsed time to average lag time. I also made a simple model to predict elapsed time of the current question\n- PEBG embedding-based features based on this paper: https://arxiv.org/pdf/2012.05031.pdf. I took the current question embedding, the user's average embedding overall, and the average embedding over the user's correct and incorrect responses. Mahalanobis distance between the current question's embedding to the user's correct and incorrect embeddings and the ratio between correct and incorrect distance were all important. \n- features from Ednet dataset (see below)\n- stats on user's past attempts on the same question, e.g. most recent AC, average AC, number of past attempts, time since last attempt, whether user saw explanation on last attempt\n- stats on user's past attempts on questions with the same tags\n- user's accuracy normalized by question difficulty across different windows, e.g. average of user's residuals for their last 3, 10, 30, 50, 100, 200 questions; residual = AC - question accuracy\n- user's accuracy trend over their last 4, 12, 30, 50, 100, 200 questions\n- user's performance on diagnostic questions: the 30 question sequence starting with 7900, 7876, 175, ... is usually asked to each user at the beginning. I made 30 features for the user's AC on each question\n- lecture features: time since user watched a lecture with the same tag, time spent watching lecture compared to lecture duration (from Ednet data)\n\nI applied additive smoothing to target encoding features to guard against small samples (this noticeably helped the QQM features). \n\n# Ednet dataset\nAs you probably know, the competition data was taken from Riiid's Ednet database, from which a large public dataset has been open sourced at https://github.com/riiid/ednet. At first glance, the overlap between the competition data and the open source data is not obvious - for example, the open source data has 784,309 users whereas the competition data only has 393,656 and the contents don't look the same either. But the Ednet data comes with question and lecture metadata files, and by comparing them to the metadata files we have in the competition, it's quite easy to map the questions, tags, and users between the datasets. The competition data seems to be a bit more recent, and has mostly filtered out users who only answered a few questions, but 350k+ out of ~393k users appear in both datasets. That's important because most of the private test data is from users appearing in the training set, so you can use the Ednet data to calculate some features on those users that you wouldn't otherwise have. Examples are:\n- Ednet data has real timestamps, not shifted to start from zero. I calculated the hour, day, and week in KST\n- how often the user switches answers\n- Ground truth video lengths for lectures, which can be used in lecture features\n- whether user is primarily on mobile or web\n\nAlso, I extracted ~6m rows from the Ednet data that was filtered out of the competition data. These were mostly users with few interactions so I didn't add them to the training set, but I did use them to pre-populate the QQM which helped a bit. \n\n# Pipeline & Remarks\nI used a loop-based accumulation framework which updated user states every iteration. To strike a good balance between speed and memory usage, I used dictionaries to look up users and numpy arrays to store the accumulated features. This worked well in general and I could use the same code for inference, which took ~3 hours in the end with 220 features and used about 9GB RAM.\n\nIn the end, I wasn't able to train on the entire dataset, or even 10m rows due to bugs that weren't fixed until the morning of the deadline. One particularly frustrating one was from the very first piece of code I ran in the competition: the code to read in the train.csv file. I had copied a line to specify the data types from a public notebook assuming it would be the same one from the Competition Starter, but the notebook creator had changed prior_question_elapsed_time to float16 instead of float32. This resulted in all values in the column >=65536 being converted to np.inf. Lesson learned: never blindly copy from a public notebook! \n\nDespite this, I really enjoyed the competition as by taking the feature engineering approach, it felt like I was uncovering new things about the data all the time. ",
      "votes": 76
    },
    {
      "id": 1143567,
      "postDate": "2021-01-08T00:30:57.780Z",
      "content": "<p>Would you possibly be sharing your snip for the QQM matrix part? Thanks! Super impressive work! Also, What boost did your get when you trained on full data? Ty!</p>",
      "rawMarkdown": "Would you possibly be sharing your snip for the QQM matrix part? Thanks! Super impressive work! Also, What boost did your get when you trained on full data? Ty!",
      "votes": 1,
      "replies": [
        {
          "id": 1143579,
          "postDate": "2021-01-08T00:39:47.080Z",
          "content": "<p>I might share the code after I get a chance to clean it up. In past experiments I got ~0.006-0.01 boost going from 1m to 10m rows and others have reported a similar boost going from 10m rows to the full data. </p>",
          "rawMarkdown": "I might share the code after I get a chance to clean it up. In past experiments I got ~0.006-0.01 boost going from 1m to 10m rows and others have reported a similar boost going from 10m rows to the full data. ",
          "votes": 1
        },
        {
          "id": 1144098,
          "postDate": "2021-01-08T08:28:12.250Z",
          "content": "<p>Similar to i2i matrix in recommendation problem, but a little different here.</p>",
          "rawMarkdown": "Similar to i2i matrix in recommendation problem, but a little different here."
        }
      ]
    },
    {
      "id": 1143563,
      "postDate": "2021-01-08T00:28:54.387Z",
      "content": "<p>Good work with LGBM and small training data <a href=\"https://www.kaggle.com/mingpan07\" target=\"_blank\">@mingpan07</a> </p>",
      "rawMarkdown": "Good work with LGBM and small training data @mingpan07 ",
      "votes": 1
    },
    {
      "id": 1147518,
      "postDate": "2021-01-10T15:05:16.813Z",
      "content": "<p>unbelievable! lgb model with 1M data achieve 0.801 so high precisin! Using distance is challenging insights.</p>",
      "rawMarkdown": "unbelievable! lgb model with 1M data achieve 0.801 so high precisin! Using distance is challenging insights."
    },
    {
      "id": 1146440,
      "postDate": "2021-01-09T19:20:34.240Z",
      "content": "<p>0.801 only with 1% data is really surprising. I guess your LGBM is probably the best in all the participants, if trained with 100M data.</p>",
      "rawMarkdown": "0.801 only with 1% data is really surprising. I guess your LGBM is probably the best in all the participants, if trained with 100M data."
    },
    {
      "id": 1145905,
      "postDate": "2021-01-09T12:20:29.277Z",
      "content": "<p>Great work! Although kinda knowing the answer (If you don't fancy it, just don't join), I am wondering if it is fair that people with access to huge amounts of computing power have a big advantage in competitions like this?</p>",
      "rawMarkdown": "Great work! Although kinda knowing the answer (If you don't fancy it, just don't join), I am wondering if it is fair that people with access to huge amounts of computing power have a big advantage in competitions like this?"
    },
    {
      "id": 1144504,
      "postDate": "2021-01-08T13:50:11.727Z",
      "content": "<p>Congrats on the unique and cool solution!<br>\nCould you please share the details of how you mapped the ednet and kaggle dataset?<br>\nI was only able to map 30% of the questions, and the features didn't help my model (I tried mobile/web and \"deployed_at\")</p>",
      "rawMarkdown": "Congrats on the unique and cool solution!\nCould you please share the details of how you mapped the ednet and kaggle dataset?\nI was only able to map 30% of the questions, and the features didn't help my model (I tried mobile/web and \"deployed_at\")"
    },
    {
      "id": 1143895,
      "postDate": "2021-01-08T05:52:11.087Z",
      "content": "<p>Awesome! I used 1% of data reached .789. Learned a lot thx.</p>",
      "rawMarkdown": "Awesome! I used 1% of data reached .789. Learned a lot thx."
    },
    {
      "id": 1143780,
      "postDate": "2021-01-08T04:10:55.367Z",
      "content": "<p>Very nice work with such a small training dataset, and the residual is very impressive. Congrats!<br>\nI'm still new here and have a little questions.</p>\n<ol>\n<li>Could you detail more about the QQM ?</li>\n<li>How could you import the Ednet ? The notebook should off line, right ?</li>\n</ol>\n<p>Thanks and Congrats!</p>",
      "rawMarkdown": "Very nice work with such a small training dataset, and the residual is very impressive. Congrats!\nI'm still new here and have a little questions.\n1. Could you detail more about the QQM ?\n2. How could you import the Ednet ? The notebook should off line, right ?\n\nThanks and Congrats!",
      "replies": [
        {
          "id": 1144092,
          "postDate": "2021-01-08T08:24:07.963Z",
          "content": "<p>QQM:Question-Question matrix (QQM) </p>",
          "rawMarkdown": "QQM:Question-Question matrix (QQM) "
        }
      ]
    },
    {
      "id": 1143696,
      "postDate": "2021-01-08T02:24:46.347Z",
      "content": "<p>1% data..amazing..very insightful as it opened my eyes that so little data can do wonders..congrats !!! </p>",
      "rawMarkdown": "1% data..amazing..very insightful as it opened my eyes that so little data can do wonders..congrats !!! "
    },
    {
      "id": 1143695,
      "postDate": "2021-01-08T02:24:02.490Z",
      "content": "<p>0.801 with Single LGB from 1% of data is so awesome!</p>\n<blockquote>\n  <p>3 hours in the end with 220 features and used about 9GB RAM</p>\n</blockquote>\n<p>So, we could achieve 0.801 in kaggle notebook environment.<br>\nI am looking forward to reading your code and the result from full data if possible.</p>",
      "rawMarkdown": "0.801 with Single LGB from 1% of data is so awesome!\n> 3 hours in the end with 220 features and used about 9GB RAM\n\nSo, we could achieve 0.801 in kaggle notebook environment.\nI am looking forward to reading your code and the result from full data if possible."
    },
    {
      "id": 1143652,
      "postDate": "2021-01-08T01:44:54.293Z",
      "content": "<p>Well done.Very impressive.</p>",
      "rawMarkdown": "Well done.Very impressive."
    },
    {
      "id": 1143651,
      "postDate": "2021-01-08T01:43:19.153Z",
      "content": "<p>1% data? awesome!</p>",
      "rawMarkdown": "1% data? awesome!"
    },
    {
      "id": 1143585,
      "postDate": "2021-01-08T00:47:04.990Z",
      "content": "<p>Thank you for sharing. Your work on Ednet dataset is impressive.</p>",
      "rawMarkdown": "Thank you for sharing. Your work on Ednet dataset is impressive."
    },
    {
      "id": 1576910,
      "postDate": "2021-11-09T14:49:23.033Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1567010,
      "postDate": "2021-11-01T12:50:01.203Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1148028,
      "postDate": "2021-01-10T21:01:40.970Z",
      "content": "<p>Good work 👌, thanks for sharing 😊</p>",
      "rawMarkdown": "Good work 👌, thanks for sharing 😊"
    }
  ],
  "comments": [
    {
      "id": 1143567,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2021-01-08T00:30:57.780000",
      "content": "<p>Would you possibly be sharing your snip for the QQM matrix part? Thanks! Super impressive work! Also, What boost did your get when you trained on full data? Ty!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1143579,
          "author_name": "Ming Pan",
          "author_url": "",
          "post_date": "2021-01-08T00:39:47.080000",
          "content": "<p>I might share the code after I get a chance to clean it up. In past experiments I got ~0.006-0.01 boost going from 1m to 10m rows and others have reported a similar boost going from 10m rows to the full data. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1144098,
          "author_name": "LongYin/杰少",
          "author_url": "",
          "post_date": "2021-01-08T08:28:12.250000",
          "content": "<p>Similar to i2i matrix in recommendation problem, but a little different here.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1143563,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-01-08T00:28:54.387000",
      "content": "<p>Good work with LGBM and small training data <a href=\"https://www.kaggle.com/mingpan07\" target=\"_blank\">@mingpan07</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1147518,
      "author_name": "2981",
      "author_url": "",
      "post_date": "2021-01-10T15:05:16.813000",
      "content": "<p>unbelievable! lgb model with 1M data achieve 0.801 so high precisin! Using distance is challenging insights.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1146440,
      "author_name": "mamas",
      "author_url": "",
      "post_date": "2021-01-09T19:20:34.240000",
      "content": "<p>0.801 only with 1% data is really surprising. I guess your LGBM is probably the best in all the participants, if trained with 100M data.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1145905,
      "author_name": "Erik Bruin",
      "author_url": "",
      "post_date": "2021-01-09T12:20:29.277000",
      "content": "<p>Great work! Although kinda knowing the answer (If you don't fancy it, just don't join), I am wondering if it is fair that people with access to huge amounts of computing power have a big advantage in competitions like this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1144504,
      "author_name": "pocket",
      "author_url": "",
      "post_date": "2021-01-08T13:50:11.727000",
      "content": "<p>Congrats on the unique and cool solution!<br>\nCould you please share the details of how you mapped the ednet and kaggle dataset?<br>\nI was only able to map 30% of the questions, and the features didn't help my model (I tried mobile/web and \"deployed_at\")</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1143895,
      "author_name": "朴大福",
      "author_url": "",
      "post_date": "2021-01-08T05:52:11.087000",
      "content": "<p>Awesome! I used 1% of data reached .789. Learned a lot thx.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1143780,
      "author_name": "marksein07",
      "author_url": "",
      "post_date": "2021-01-08T04:10:55.367000",
      "content": "<p>Very nice work with such a small training dataset, and the residual is very impressive. Congrats!<br>\nI'm still new here and have a little questions.</p>\n<ol>\n<li>Could you detail more about the QQM ?</li>\n<li>How could you import the Ednet ? The notebook should off line, right ?</li>\n</ol>\n<p>Thanks and Congrats!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1144092,
          "author_name": "LongYin/杰少",
          "author_url": "",
          "post_date": "2021-01-08T08:24:07.963000",
          "content": "<p>QQM:Question-Question matrix (QQM) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1143696,
      "author_name": "Anurag Trivedi",
      "author_url": "",
      "post_date": "2021-01-08T02:24:46.347000",
      "content": "<p>1% data..amazing..very insightful as it opened my eyes that so little data can do wonders..congrats !!! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1143695,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-01-08T02:24:02.490000",
      "content": "<p>0.801 with Single LGB from 1% of data is so awesome!</p>\n<blockquote>\n  <p>3 hours in the end with 220 features and used about 9GB RAM</p>\n</blockquote>\n<p>So, we could achieve 0.801 in kaggle notebook environment.<br>\nI am looking forward to reading your code and the result from full data if possible.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1143652,
      "author_name": "kwang",
      "author_url": "",
      "post_date": "2021-01-08T01:44:54.293000",
      "content": "<p>Well done.Very impressive.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1143651,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2021-01-08T01:43:19.153000",
      "content": "<p>1% data? awesome!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1143585,
      "author_name": "u++",
      "author_url": "",
      "post_date": "2021-01-08T00:47:04.990000",
      "content": "<p>Thank you for sharing. Your work on Ednet dataset is impressive.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1576910,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-09T14:49:23.033000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1567010,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-01T12:50:01.203000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1148028,
      "author_name": "Miloud Belarebia",
      "author_url": "",
      "post_date": "2021-01-10T21:01:40.970000",
      "content": "<p>Good work 👌, thanks for sharing 😊</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143557": "Thanks to the organizers of this very interesting competition and congrats to the winners. My LB score (0.800 public / 0.801 private) was achieved using a single LGB model trained on 1m rows. It wasn't my intention to only use 1m rows, but I had some bugs that I fixed too late to retrain on more data in time. Nevertheless, I'm sharing my solution in case it ends up being useful to anyone.\n\n# Features\nThe final model used 220 features that fit into following categories in rough descending order of importance:\nnote: AC = answered_correctly\n- Question-Question matrix (QQM) features: out of all users who answered A on question X, how many of them get question Y right? I built a 13523 x 13523 x 4 matrix of all the probabilities, then calculated user stats on their last 3, 10, 30, 50, 100, 200 questions. I also created a separate set of features weighting by Mahalanobis distance between the past and current question embeddings (see section on PEBG embeddings). \n- question's overall accuracy/difficulty across all users, number of responses\n- user's accuracy, number of respones across their entire history & current session (where a new session is created if the user is idle for more than 5 min)\n- question metadata, i.e. tags, part, index in bundle\n- elapsed time and lag time - based features. Ratios were good here, e.g. ratio of prior elapsed time to lag time, user's average elapsed time to average lag time. I also made a simple model to predict elapsed time of the current question\n- PEBG embedding-based features based on this paper: https://arxiv.org/pdf/2012.05031.pdf. I took the current question embedding, the user's average embedding overall, and the average embedding over the user's correct and incorrect responses. Mahalanobis distance between the current question's embedding to the user's correct and incorrect embeddings and the ratio between correct and incorrect distance were all important. \n- features from Ednet dataset (see below)\n- stats on user's past attempts on the same question, e.g. most recent AC, average AC, number of past attempts, time since last attempt, whether user saw explanation on last attempt\n- stats on user's past attempts on questions with the same tags\n- user's accuracy normalized by question difficulty across different windows, e.g. average of user's residuals for their last 3, 10, 30, 50, 100, 200 questions; residual = AC - question accuracy\n- user's accuracy trend over their last 4, 12, 30, 50, 100, 200 questions\n- user's performance on diagnostic questions: the 30 question sequence starting with 7900, 7876, 175, ... is usually asked to each user at the beginning. I made 30 features for the user's AC on each question\n- lecture features: time since user watched a lecture with the same tag, time spent watching lecture compared to lecture duration (from Ednet data)\n\nI applied additive smoothing to target encoding features to guard against small samples (this noticeably helped the QQM features). \n\n# Ednet dataset\nAs you probably know, the competition data was taken from Riiid's Ednet database, from which a large public dataset has been open sourced at https://github.com/riiid/ednet. At first glance, the overlap between the competition data and the open source data is not obvious - for example, the open source data has 784,309 users whereas the competition data only has 393,656 and the contents don't look the same either. But the Ednet data comes with question and lecture metadata files, and by comparing them to the metadata files we have in the competition, it's quite easy to map the questions, tags, and users between the datasets. The competition data seems to be a bit more recent, and has mostly filtered out users who only answered a few questions, but 350k+ out of ~393k users appear in both datasets. That's important because most of the private test data is from users appearing in the training set, so you can use the Ednet data to calculate some features on those users that you wouldn't otherwise have. Examples are:\n- Ednet data has real timestamps, not shifted to start from zero. I calculated the hour, day, and week in KST\n- how often the user switches answers\n- Ground truth video lengths for lectures, which can be used in lecture features\n- whether user is primarily on mobile or web\n\nAlso, I extracted ~6m rows from the Ednet data that was filtered out of the competition data. These were mostly users with few interactions so I didn't add them to the training set, but I did use them to pre-populate the QQM which helped a bit. \n\n# Pipeline & Remarks\nI used a loop-based accumulation framework which updated user states every iteration. To strike a good balance between speed and memory usage, I used dictionaries to look up users and numpy arrays to store the accumulated features. This worked well in general and I could use the same code for inference, which took ~3 hours in the end with 220 features and used about 9GB RAM.\n\nIn the end, I wasn't able to train on the entire dataset, or even 10m rows due to bugs that weren't fixed until the morning of the deadline. One particularly frustrating one was from the very first piece of code I ran in the competition: the code to read in the train.csv file. I had copied a line to specify the data types from a public notebook assuming it would be the same one from the Competition Starter, but the notebook creator had changed prior_question_elapsed_time to float16 instead of float32. This resulted in all values in the column >=65536 being converted to np.inf. Lesson learned: never blindly copy from a public notebook! \n\nDespite this, I really enjoyed the competition as by taking the feature engineering approach, it felt like I was uncovering new things about the data all the time. ",
    "1143567": "Would you possibly be sharing your snip for the QQM matrix part? Thanks! Super impressive work! Also, What boost did your get when you trained on full data? Ty!",
    "1143563": "Good work with LGBM and small training data @mingpan07 ",
    "1147518": "unbelievable! lgb model with 1M data achieve 0.801 so high precisin! Using distance is challenging insights.",
    "1146440": "0.801 only with 1% data is really surprising. I guess your LGBM is probably the best in all the participants, if trained with 100M data.",
    "1145905": "Great work! Although kinda knowing the answer (If you don't fancy it, just don't join), I am wondering if it is fair that people with access to huge amounts of computing power have a big advantage in competitions like this?",
    "1144504": "Congrats on the unique and cool solution!\nCould you please share the details of how you mapped the ednet and kaggle dataset?\nI was only able to map 30% of the questions, and the features didn't help my model (I tried mobile/web and \"deployed_at\")",
    "1143895": "Awesome! I used 1% of data reached .789. Learned a lot thx.",
    "1143780": "Very nice work with such a small training dataset, and the residual is very impressive. Congrats!\nI'm still new here and have a little questions.\n1. Could you detail more about the QQM ?\n2. How could you import the Ednet ? The notebook should off line, right ?\n\nThanks and Congrats!",
    "1143696": "1% data..amazing..very insightful as it opened my eyes that so little data can do wonders..congrats !!! ",
    "1143695": "0.801 with Single LGB from 1% of data is so awesome!\n> 3 hours in the end with 220 features and used about 9GB RAM\n\nSo, we could achieve 0.801 in kaggle notebook environment.\nI am looking forward to reading your code and the result from full data if possible.",
    "1143652": "Well done.Very impressive.",
    "1143651": "1% data? awesome!",
    "1143585": "Thank you for sharing. Your work on Ednet dataset is impressive.",
    "1576910": "",
    "1567010": "",
    "1148028": "Good work 👌, thanks for sharing 😊"
  }
}