{
  "id": 210552,
  "title": "#8 solution: Ensemble of 15 same NN models",
  "url": "/competitions/riiid-test-answer-prediction/writeups/chlxyd-8-solution-ensemble-of-15-same-nn-models",
  "author_name": "",
  "post_date": "2021-01-11T09:13:20.776282500Z",
  "votes": 40,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi all,<br>\nLearned a lot from other's solutions! In this post, I would like to share some insights in my solution.</p>\n<p>In summary, my best single NN model could achieve 813/815 public/private score, by ensemble of 5 folds and 3 snapshots in each fold, finally 15 nn models achieve the 814/816 public/private score.</p>\n<p>With 15 nn models, the online inference cost <strong>less than 4 hours</strong>, thus it can ensemble at least 30 models in this pipeline.</p>\n<h2>Dataset split</h2>\n<p>I split the train and valid set by several steps:</p>\n<ol>\n<li>Calculate all unique users in dataset.</li>\n<li>Select 5% users totally in the valid set, and 45% users totally in the train set.</li>\n<li>For the least 50% users, random split the user data into train and valid set by time.</li>\n</ol>\n<p>Thus we have both individual user in train and valid set, and also have many users who appear in both train and valid set. This split can achieve less than 0.001 score difference compared with LB.</p>\n<p>By changing different random seed, we can get different folds.</p>\n<h2>Feature engineering</h2>\n<p>Since there are many detailed FE in other's post, I will just share some important points.</p>\n<h3>Basic features</h3>\n<ol>\n<li><p>Evaluate user ability</p>\n<ul>\n<li>We can evaluate user ability by his history action, including the correctness, time elapsed, lag time, and so on.</li>\n<li>These features can be calculated on not only content level, but also on same part, same tags, same content and so on.</li>\n<li>These features can also be extended based on time, for example, we can make a feature which is the user correctness in last 60 seconds.</li></ul></li>\n<li><p>Evaluate content difficulty</p>\n<ul>\n<li>We can evaluate content difficulty by calculate its global accuracy, std, average time elapsed and so on.</li></ul></li>\n<li><p>Evaluate User x Content features</p>\n<p>Even a content is difficult, the user may still skill enough and can solve it correctly. Thus we also need to describe how the user could perform on this content.</p>\n<ul>\n<li><p>user_acc_diff: For a content, if the content has low global acc, but the user answered it correctly, the user might above the average of all users. We can use the $logloss(content_global_average_acc, user_answer)$ to evaluate the difference.</p></li>\n<li><p>user_elapsed_diff: like the user_acc_diff, we can also evaluate user time elapsed difference in user history contents.</p></li></ul></li>\n</ol>\n<p>We can dig many features on above three fields, for example, user's average lag time on history content/ history same part content/ history same tag content.. could also be a useful features.</p>\n<h3>Other features/tricks</h3>\n<ol>\n<li>Time related features: last_content_timestamp_diff, last_lag_time and its statistical information in history.</li>\n<li>Abnormal Users: If a user answered every content less in 4 seconds, and all his chose answer are same (such as C), then if the correct answer of next content is C, we can believe he will correctly answered next content.</li>\n<li>Learned lectures for a specific content: If there is a content-lecture-content pattern in user history, and the two content are same contents. It might the user learned a specific lecture for this content, which means in the second time, the probability he answered it correctly is high.</li>\n<li>Wrong answer ratio: There might be a pattern like \"select C as default for hard contents\".  Thus calculating the ratio of user choice on his incorrectly content can tell as weather the user could lucky guess current content even he don't know he correct answer.</li>\n</ol>\n<p>Also, features such as current content id/current timestamp/current part are also used. Finally, I get 120 dim features, which can get public 0.806 by a single lgb model in single fold.</p>\n<p>P.S. The categorical feature (set on content id) of lightgbm can boost my score about 0.003.</p>\n<p>There are also many useful hint which could improve the speed and save the memory:</p>\n<ol>\n<li>Do not use pandas to calculate features, transfer it to numpy or just python.</li>\n<li>Using a class to store user information, and save each user information in individual file.</li>\n<li>In the inference phase, we can only read the user information for who appeared in the test set, the total number of users in test set are less than 10,000. This is the key to reduce memory. If the memory still not enough, the LRU-cache could be used to remove unused users.</li>\n<li>In my experiment, using numpy array to store information cost more disk space compared with python variable.</li>\n</ol>\n<h2>NN</h2>\n<p>Since it is very late when I notice the key to the top is NN model, I don't have much time analysis the NN models, especially design a specific features or structures. To save the time, I use the lightgbm features as the NN input (Time axis is added and seq len is 128).</p>\n<h4>Robust Standard Normalization</h4>\n<p>There are many outlier in the features from lightgbm, thus simply utilize standard normalization can hardly get desired results. I utilized the robust standard normalization to normalize all the features:</p>\n<pre><code>def robust_normalization(column):\n    cur_mean = np.nanmedian(column)\n    cur_qmin, cur_qmax = np.nanpercentile(cur,[2.5, 97.5])\n    cur_std = np.nanstd(column[(column&gt;=cur_qmin) &amp; (column&lt;=cur_qmax)])\n    column = np.clip(column, a_min = cur_qmin, a_max = cur_qmax)\n    column = (column-cur_mean)/cur_std\n    return column\n</code></pre>\n<h4>Models</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1458993%2F4ae543902ffae52e87dfaf0484477c7d%2FWeChat%20Image_20210111165953.png?generation=1610355676642083&amp;alt=media\" alt=\"\"></p>\n<p>My NN model is very simple, just a transformer encoder with a fc classifier. The encoder has 4 transformer layers and embed dim is 128, no other modifications. The size of model file is only 7M, that's why I can ensemble many models on inference phase.</p>\n<p>The model can achieve 813/815 on single model/single fold. and 814/816 on 5 folds ensemble. There are much time left on inference, I use 3 snapshot ensemble on each fold, which can also boost some scores. </p>",
  "messages": [
    {
      "id": "1148621",
      "postDate": "01/11/2021 09:13:20",
      "content": "<p>Hi all,<br>\nLearned a lot from other's solutions! In this post, I would like to share some insights in my solution.</p>\n<p>In summary, my best single NN model could achieve 813/815 public/private score, by ensemble of 5 folds and 3 snapshots in each fold, finally 15 nn models achieve the 814/816 public/private score.</p>\n<p>With 15 nn models, the online inference cost <strong>less than 4 hours</strong>, thus it can ensemble at least 30 models in this pipeline.</p>\n<h2>Dataset split</h2>\n<p>I split the train and valid set by several steps:</p>\n<ol>\n<li>Calculate all unique users in dataset.</li>\n<li>Select 5% users totally in the valid set, and 45% users totally in the train set.</li>\n<li>For the least 50% users, random split the user data into train and valid set by time.</li>\n</ol>\n<p>Thus we have both individual user in train and valid set, and also have many users who appear in both train and valid set. This split can achieve less than 0.001 score difference compared with LB.</p>\n<p>By changing different random seed, we can get different folds.</p>\n<h2>Feature engineering</h2>\n<p>Since there are many detailed FE in other's post, I will just share some important points.</p>\n<h3>Basic features</h3>\n<ol>\n<li><p>Evaluate user ability</p>\n<ul>\n<li>We can evaluate user ability by his history action, including the correctness, time elapsed, lag time, and so on.</li>\n<li>These features can be calculated on not only content level, but also on same part, same tags, same content and so on.</li>\n<li>These features can also be extended based on time, for example, we can make a feature which is the user correctness in last 60 seconds.</li></ul></li>\n<li><p>Evaluate content difficulty</p>\n<ul>\n<li>We can evaluate content difficulty by calculate its global accuracy, std, average time elapsed and so on.</li></ul></li>\n<li><p>Evaluate User x Content features</p>\n<p>Even a content is difficult, the user may still skill enough and can solve it correctly. Thus we also need to describe how the user could perform on this content.</p>\n<ul>\n<li><p>user_acc_diff: For a content, if the content has low global acc, but the user answered it correctly, the user might above the average of all users. We can use the $logloss(content_global_average_acc, user_answer)$ to evaluate the difference.</p></li>\n<li><p>user_elapsed_diff: like the user_acc_diff, we can also evaluate user time elapsed difference in user history contents.</p></li></ul></li>\n</ol>\n<p>We can dig many features on above three fields, for example, user's average lag time on history content/ history same part content/ history same tag content.. could also be a useful features.</p>\n<h3>Other features/tricks</h3>\n<ol>\n<li>Time related features: last_content_timestamp_diff, last_lag_time and its statistical information in history.</li>\n<li>Abnormal Users: If a user answered every content less in 4 seconds, and all his chose answer are same (such as C), then if the correct answer of next content is C, we can believe he will correctly answered next content.</li>\n<li>Learned lectures for a specific content: If there is a content-lecture-content pattern in user history, and the two content are same contents. It might the user learned a specific lecture for this content, which means in the second time, the probability he answered it correctly is high.</li>\n<li>Wrong answer ratio: There might be a pattern like \"select C as default for hard contents\".  Thus calculating the ratio of user choice on his incorrectly content can tell as weather the user could lucky guess current content even he don't know he correct answer.</li>\n</ol>\n<p>Also, features such as current content id/current timestamp/current part are also used. Finally, I get 120 dim features, which can get public 0.806 by a single lgb model in single fold.</p>\n<p>P.S. The categorical feature (set on content id) of lightgbm can boost my score about 0.003.</p>\n<p>There are also many useful hint which could improve the speed and save the memory:</p>\n<ol>\n<li>Do not use pandas to calculate features, transfer it to numpy or just python.</li>\n<li>Using a class to store user information, and save each user information in individual file.</li>\n<li>In the inference phase, we can only read the user information for who appeared in the test set, the total number of users in test set are less than 10,000. This is the key to reduce memory. If the memory still not enough, the LRU-cache could be used to remove unused users.</li>\n<li>In my experiment, using numpy array to store information cost more disk space compared with python variable.</li>\n</ol>\n<h2>NN</h2>\n<p>Since it is very late when I notice the key to the top is NN model, I don't have much time analysis the NN models, especially design a specific features or structures. To save the time, I use the lightgbm features as the NN input (Time axis is added and seq len is 128).</p>\n<h4>Robust Standard Normalization</h4>\n<p>There are many outlier in the features from lightgbm, thus simply utilize standard normalization can hardly get desired results. I utilized the robust standard normalization to normalize all the features:</p>\n<pre><code>def robust_normalization(column):\n    cur_mean = np.nanmedian(column)\n    cur_qmin, cur_qmax = np.nanpercentile(cur,[2.5, 97.5])\n    cur_std = np.nanstd(column[(column&gt;=cur_qmin) &amp; (column&lt;=cur_qmax)])\n    column = np.clip(column, a_min = cur_qmin, a_max = cur_qmax)\n    column = (column-cur_mean)/cur_std\n    return column\n</code></pre>\n<h4>Models</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1458993%2F4ae543902ffae52e87dfaf0484477c7d%2FWeChat%20Image_20210111165953.png?generation=1610355676642083&amp;alt=media\" alt=\"\"></p>\n<p>My NN model is very simple, just a transformer encoder with a fc classifier. The encoder has 4 transformer layers and embed dim is 128, no other modifications. The size of model file is only 7M, that's why I can ensemble many models on inference phase.</p>\n<p>The model can achieve 813/815 on single model/single fold. and 814/816 on 5 folds ensemble. There are much time left on inference, I use 3 snapshot ensemble on each fold, which can also boost some scores. </p>",
      "rawMarkdown": "Hi all,\nLearned a lot from other's solutions! In this post, I would like to share some insights in my solution.\n\nIn summary, my best single NN model could achieve 813/815 public/private score, by ensemble of 5 folds and 3 snapshots in each fold, finally 15 nn models achieve the 814/816 public/private score.\n\nWith 15 nn models, the online inference cost **less than 4 hours**, thus it can ensemble at least 30 models in this pipeline.\n\n## Dataset split\n\nI split the train and valid set by several steps:\n\n1. Calculate all unique users in dataset.\n2. Select 5% users totally in the valid set, and 45% users totally in the train set.\n3. For the least 50% users, random split the user data into train and valid set by time.\n\nThus we have both individual user in train and valid set, and also have many users who appear in both train and valid set. This split can achieve less than 0.001 score difference compared with LB.\n\nBy changing different random seed, we can get different folds.\n\n## Feature engineering\n\nSince there are many detailed FE in other's post, I will just share some important points.\n\n### Basic features\n\n1. Evaluate user ability\n\n   + We can evaluate user ability by his history action, including the correctness, time elapsed, lag time, and so on.\n   + These features can be calculated on not only content level, but also on same part, same tags, same content and so on.\n   + These features can also be extended based on time, for example, we can make a feature which is the user correctness in last 60 seconds.\n\n2. Evaluate content difficulty\n\n   + We can evaluate content difficulty by calculate its global accuracy, std, average time elapsed and so on.\n\n3. Evaluate User x Content features\n\n   Even a content is difficult, the user may still skill enough and can solve it correctly. Thus we also need to describe how the user could perform on this content.\n\n   + user_acc_diff: For a content, if the content has low global acc, but the user answered it correctly, the user might above the average of all users. We can use the $logloss(content_global_average_acc, user_answer)$ to evaluate the difference.\n\n   + user_elapsed_diff: like the user_acc_diff, we can also evaluate user time elapsed difference in user history contents.\n\nWe can dig many features on above three fields, for example, user's average lag time on history content/ history same part content/ history same tag content.. could also be a useful features.\n\n### Other features/tricks\n\n1. Time related features: last_content_timestamp_diff, last_lag_time and its statistical information in history.\n2. Abnormal Users: If a user answered every content less in 4 seconds, and all his chose answer are same (such as C), then if the correct answer of next content is C, we can believe he will correctly answered next content.\n3. Learned lectures for a specific content: If there is a content-lecture-content pattern in user history, and the two content are same contents. It might the user learned a specific lecture for this content, which means in the second time, the probability he answered it correctly is high.\n4. Wrong answer ratio: There might be a pattern like \"select C as default for hard contents\".  Thus calculating the ratio of user choice on his incorrectly content can tell as weather the user could lucky guess current content even he don't know he correct answer.\n\nAlso, features such as current content id/current timestamp/current part are also used. Finally, I get 120 dim features, which can get public 0.806 by a single lgb model in single fold.\n\nP.S. The categorical feature (set on content id) of lightgbm can boost my score about 0.003.\n\nThere are also many useful hint which could improve the speed and save the memory:\n\n1. Do not use pandas to calculate features, transfer it to numpy or just python.\n2. Using a class to store user information, and save each user information in individual file.\n3. In the inference phase, we can only read the user information for who appeared in the test set, the total number of users in test set are less than 10,000. This is the key to reduce memory. If the memory still not enough, the LRU-cache could be used to remove unused users.\n4. In my experiment, using numpy array to store information cost more disk space compared with python variable.\n\n## NN\n\nSince it is very late when I notice the key to the top is NN model, I don't have much time analysis the NN models, especially design a specific features or structures. To save the time, I use the lightgbm features as the NN input (Time axis is added and seq len is 128).\n\n#### Robust Standard Normalization\n\n There are many outlier in the features from lightgbm, thus simply utilize standard normalization can hardly get desired results. I utilized the robust standard normalization to normalize all the features:\n\n```\ndef robust_normalization(column):\n\tcur_mean = np.nanmedian(column)\n    cur_qmin, cur_qmax = np.nanpercentile(cur,[2.5, 97.5])\n    cur_std = np.nanstd(column[(column>=cur_qmin) & (column<=cur_qmax)])\n    column = np.clip(column, a_min = cur_qmin, a_max = cur_qmax)\n    column = (column-cur_mean)/cur_std\n    return column\n```\n\n#### Models\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1458993%2F4ae543902ffae52e87dfaf0484477c7d%2FWeChat%20Image_20210111165953.png?generation=1610355676642083&alt=media)\n\nMy NN model is very simple, just a transformer encoder with a fc classifier. The encoder has 4 transformer layers and embed dim is 128, no other modifications. The size of model file is only 7M, that's why I can ensemble many models on inference phase.\n\nThe model can achieve 813/815 on single model/single fold. and 814/816 on 5 folds ensemble. There are much time left on inference, I use 3 snapshot ensemble on each fold, which can also boost some scores.",
      "votes": null
    },
    {
      "id": "1148680",
      "postDate": "01/11/2021 09:53:17",
      "content": "<p>Congratulations on winning the solo gold medal~~， hange tql👍👍👍</p>",
      "rawMarkdown": "Congratulations on winning the solo gold medal~~， hange tql👍👍👍",
      "votes": null
    },
    {
      "id": "1148690",
      "postDate": "01/11/2021 10:18:52",
      "content": "<p>Great post! Learn a lot from it, thanks!</p>",
      "rawMarkdown": "Great post! Learn a lot from it, thanks!",
      "votes": null
    },
    {
      "id": "1148707",
      "postDate": "01/11/2021 10:30:25",
      "content": "<p>I have a question about the ensembling. Why don't you ensemble nn model with lgb model. For me, when ensembling a nn model(0.806) with a lgb model(0.798), the LB score is 0.809. I believe you must have a better lgb model than mein. </p>",
      "rawMarkdown": "I have a question about the ensembling. Why don't you ensemble nn model with lgb model. For me, when ensembling a nn model(0.806) with a lgb model(0.798), the LB score is 0.809. I believe you must have a better lgb model than mein.",
      "votes": null
    },
    {
      "id": "1148753",
      "postDate": "01/11/2021 11:12:09",
      "content": "<p>My LGB and NN are from the same feature group, thus I can get only 0.0002 boost after my NN score higher than 0.810, that's why I give up the LGB model in the later submission.</p>",
      "rawMarkdown": "My LGB and NN are from the same feature group, thus I can get only 0.0002 boost after my NN score higher than 0.810, that's why I give up the LGB model in the later submission.",
      "votes": null
    },
    {
      "id": "1148849",
      "postDate": "01/11/2021 12:28:41",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/chlxyd\" target=\"_blank\">@chlxyd</a> on 8th place and thanks for sharing details solution</p>",
      "rawMarkdown": "Congrats @chlxyd on 8th place and thanks for sharing details solution",
      "votes": null
    },
    {
      "id": "1149236",
      "postDate": "01/11/2021 17:41:04",
      "content": "<p>big big congrats bro!</p>",
      "rawMarkdown": "big big congrats bro!",
      "votes": null
    },
    {
      "id": "1149563",
      "postDate": "01/12/2021 00:36:00",
      "content": "<p>congrats bro!</p>",
      "rawMarkdown": "congrats bro!",
      "votes": null
    },
    {
      "id": "1149603",
      "postDate": "01/12/2021 02:09:36",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/chlxyd\" target=\"_blank\">@chlxyd</a> </p>",
      "rawMarkdown": "congrats @chlxyd",
      "votes": null
    },
    {
      "id": "1149678",
      "postDate": "01/12/2021 04:08:34",
      "content": "<p>congrats! ddw</p>",
      "rawMarkdown": "congrats! ddw",
      "votes": null
    },
    {
      "id": "1156044",
      "postDate": "01/16/2021 22:07:58",
      "content": "<p>Congrats on the great solution and nice finish!</p>\n<p>Could you elaborate on how to save each user's information separately, with 400K pickle files or h5py? Also would like to learn more about your implementation on <code>LRU-Cache</code>.<br>\nBesides, could you share a little bit more about your NN model structure and params, any code snippet would be helpful.</p>\n<p>Thanks~!</p>",
      "rawMarkdown": "Congrats on the great solution and nice finish!\n\nCould you elaborate on how to save each user's information separately, with 400K pickle files or h5py? Also would like to learn more about your implementation on `LRU-Cache`.\nBesides, could you share a little bit more about your NN model structure and params, any code snippet would be helpful.\n \nThanks~!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1148680,
      "author_name": "meistermorxrc",
      "author_url": "",
      "post_date": "01/11/2021 09:53:17",
      "content": "<p>Congratulations on winning the solo gold medal~~， hange tql👍👍👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1148690,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "01/11/2021 10:18:52",
      "content": "<p>Great post! Learn a lot from it, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1148707,
      "author_name": "lihaorocky",
      "author_url": "",
      "post_date": "01/11/2021 10:30:25",
      "content": "<p>I have a question about the ensembling. Why don't you ensemble nn model with lgb model. For me, when ensembling a nn model(0.806) with a lgb model(0.798), the LB score is 0.809. I believe you must have a better lgb model than mein. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1148753,
          "author_name": "chlxyd",
          "author_url": "",
          "post_date": "01/11/2021 11:12:09",
          "content": "<p>My LGB and NN are from the same feature group, thus I can get only 0.0002 boost after my NN score higher than 0.810, that's why I give up the LGB model in the later submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1148849,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/11/2021 12:28:41",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/chlxyd\" target=\"_blank\">@chlxyd</a> on 8th place and thanks for sharing details solution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1149236,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "01/11/2021 17:41:04",
      "content": "<p>big big congrats bro!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1149563,
      "author_name": "fengari",
      "author_url": "",
      "post_date": "01/12/2021 00:36:00",
      "content": "<p>congrats bro!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1149603,
      "author_name": "lovedm",
      "author_url": "",
      "post_date": "01/12/2021 02:09:36",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/chlxyd\" target=\"_blank\">@chlxyd</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1149678,
      "author_name": "chizhu2018",
      "author_url": "",
      "post_date": "01/12/2021 04:08:34",
      "content": "<p>congrats! ddw</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1156044,
      "author_name": "zonemercy",
      "author_url": "",
      "post_date": "01/16/2021 22:07:58",
      "content": "<p>Congrats on the great solution and nice finish!</p>\n<p>Could you elaborate on how to save each user's information separately, with 400K pickle files or h5py? Also would like to learn more about your implementation on <code>LRU-Cache</code>.<br>\nBesides, could you share a little bit more about your NN model structure and params, any code snippet would be helpful.</p>\n<p>Thanks~!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1148621": "Hi all,\nLearned a lot from other's solutions! In this post, I would like to share some insights in my solution.\n\nIn summary, my best single NN model could achieve 813/815 public/private score, by ensemble of 5 folds and 3 snapshots in each fold, finally 15 nn models achieve the 814/816 public/private score.\n\nWith 15 nn models, the online inference cost **less than 4 hours**, thus it can ensemble at least 30 models in this pipeline.\n\n## Dataset split\n\nI split the train and valid set by several steps:\n\n1. Calculate all unique users in dataset.\n2. Select 5% users totally in the valid set, and 45% users totally in the train set.\n3. For the least 50% users, random split the user data into train and valid set by time.\n\nThus we have both individual user in train and valid set, and also have many users who appear in both train and valid set. This split can achieve less than 0.001 score difference compared with LB.\n\nBy changing different random seed, we can get different folds.\n\n## Feature engineering\n\nSince there are many detailed FE in other's post, I will just share some important points.\n\n### Basic features\n\n1. Evaluate user ability\n\n   + We can evaluate user ability by his history action, including the correctness, time elapsed, lag time, and so on.\n   + These features can be calculated on not only content level, but also on same part, same tags, same content and so on.\n   + These features can also be extended based on time, for example, we can make a feature which is the user correctness in last 60 seconds.\n\n2. Evaluate content difficulty\n\n   + We can evaluate content difficulty by calculate its global accuracy, std, average time elapsed and so on.\n\n3. Evaluate User x Content features\n\n   Even a content is difficult, the user may still skill enough and can solve it correctly. Thus we also need to describe how the user could perform on this content.\n\n   + user_acc_diff: For a content, if the content has low global acc, but the user answered it correctly, the user might above the average of all users. We can use the $logloss(content_global_average_acc, user_answer)$ to evaluate the difference.\n\n   + user_elapsed_diff: like the user_acc_diff, we can also evaluate user time elapsed difference in user history contents.\n\nWe can dig many features on above three fields, for example, user's average lag time on history content/ history same part content/ history same tag content.. could also be a useful features.\n\n### Other features/tricks\n\n1. Time related features: last_content_timestamp_diff, last_lag_time and its statistical information in history.\n2. Abnormal Users: If a user answered every content less in 4 seconds, and all his chose answer are same (such as C), then if the correct answer of next content is C, we can believe he will correctly answered next content.\n3. Learned lectures for a specific content: If there is a content-lecture-content pattern in user history, and the two content are same contents. It might the user learned a specific lecture for this content, which means in the second time, the probability he answered it correctly is high.\n4. Wrong answer ratio: There might be a pattern like \"select C as default for hard contents\".  Thus calculating the ratio of user choice on his incorrectly content can tell as weather the user could lucky guess current content even he don't know he correct answer.\n\nAlso, features such as current content id/current timestamp/current part are also used. Finally, I get 120 dim features, which can get public 0.806 by a single lgb model in single fold.\n\nP.S. The categorical feature (set on content id) of lightgbm can boost my score about 0.003.\n\nThere are also many useful hint which could improve the speed and save the memory:\n\n1. Do not use pandas to calculate features, transfer it to numpy or just python.\n2. Using a class to store user information, and save each user information in individual file.\n3. In the inference phase, we can only read the user information for who appeared in the test set, the total number of users in test set are less than 10,000. This is the key to reduce memory. If the memory still not enough, the LRU-cache could be used to remove unused users.\n4. In my experiment, using numpy array to store information cost more disk space compared with python variable.\n\n## NN\n\nSince it is very late when I notice the key to the top is NN model, I don't have much time analysis the NN models, especially design a specific features or structures. To save the time, I use the lightgbm features as the NN input (Time axis is added and seq len is 128).\n\n#### Robust Standard Normalization\n\n There are many outlier in the features from lightgbm, thus simply utilize standard normalization can hardly get desired results. I utilized the robust standard normalization to normalize all the features:\n\n```\ndef robust_normalization(column):\n\tcur_mean = np.nanmedian(column)\n    cur_qmin, cur_qmax = np.nanpercentile(cur,[2.5, 97.5])\n    cur_std = np.nanstd(column[(column>=cur_qmin) & (column<=cur_qmax)])\n    column = np.clip(column, a_min = cur_qmin, a_max = cur_qmax)\n    column = (column-cur_mean)/cur_std\n    return column\n```\n\n#### Models\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1458993%2F4ae543902ffae52e87dfaf0484477c7d%2FWeChat%20Image_20210111165953.png?generation=1610355676642083&alt=media)\n\nMy NN model is very simple, just a transformer encoder with a fc classifier. The encoder has 4 transformer layers and embed dim is 128, no other modifications. The size of model file is only 7M, that's why I can ensemble many models on inference phase.\n\nThe model can achieve 813/815 on single model/single fold. and 814/816 on 5 folds ensemble. There are much time left on inference, I use 3 snapshot ensemble on each fold, which can also boost some scores.",
    "1148680": "Congratulations on winning the solo gold medal~~， hange tql👍👍👍",
    "1148690": "Great post! Learn a lot from it, thanks!",
    "1148707": "I have a question about the ensembling. Why don't you ensemble nn model with lgb model. For me, when ensembling a nn model(0.806) with a lgb model(0.798), the LB score is 0.809. I believe you must have a better lgb model than mein.",
    "1148753": "My LGB and NN are from the same feature group, thus I can get only 0.0002 boost after my NN score higher than 0.810, that's why I give up the LGB model in the later submission.",
    "1148849": "Congrats @chlxyd on 8th place and thanks for sharing details solution",
    "1149236": "big big congrats bro!",
    "1149563": "congrats bro!",
    "1149603": "congrats @chlxyd",
    "1149678": "congrats! ddw",
    "1156044": "Congrats on the great solution and nice finish!\n\nCould you elaborate on how to save each user's information separately, with 400K pickle files or h5py? Also would like to learn more about your implementation on `LRU-Cache`.\nBesides, could you share a little bit more about your NN model structure and params, any code snippet would be helpful.\n \nThanks~!"
  },
  "source": "meta"
}