{
  "id": 210041,
  "title": "28th place solution",
  "url": "/competitions/riiid-test-answer-prediction/writeups/mada-toberu-28th-place-solution",
  "author_name": "",
  "post_date": "2021-01-09T13:10:11.483Z",
  "votes": 28,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I want to thank the hosts for organizing a meaningful competition that reflected the critical constraints of real-world difficulties (time series API, limited RAM, and inference time). Also, I thanks my teammate NARI who worked hard with me until the end of the competition.</p>\n<h1>Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2F5cfa0a830e6e45584a9088a794720ef7%2FOverview.jpg?generation=1610195752423547&amp;alt=media\" alt=\"\"></p>\n<p>Our final solution is an ensemble of Catboost and Transformer. Although Transformer could not create a high CV model in time like the top teams, it contributes a lot to boosting the stacking model score. Our solution's validation strategy and feature engineering pipeline were heavily influenced by tito's notebooks (<a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">validation strategy</a>, <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">feature engineering</a>). We definitely would not have reached this score if he had not been shared early in the competition.</p>\n<h1>Features</h1>\n<p>We extracted 160 features for training Catboost. </p>\n<p>The main features are as follows. The detailed contribution to the score of each feature has not been confirmed, but the feature importance by CatBoost is available <a href=\"https://github.com/haradai1262/kaggle_riiid-test-answer-prediction/blob/main/notebook/check_feature_importance.ipynb\" target=\"_blank\">here</a></p>\n<ul>\n<li><p>Content-related features</p>\n<ul>\n<li><p>Aggregation features</p>\n<ul>\n<li>mean, std for answered_correctly</li>\n<li>mean for elapsed time</li>\n<li>We also used aggregation in records only for the second and subsequent answers to the same question by the user.</li></ul></li>\n<li><p>Word2vec features<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2Fa381717acbf8460946ef6455aef948c4%2Fword2vec%20figure.jpg?generation=1610196796564277&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>We extract Word2vec features using each user's answer history as sentences, each question_id and its associated part, tag_id, and lecture_id as words.</li>\n<li>We also extracted cases in which users' answer histories for only correct answers, and only incorrect answers were treated as separate sentences.</li>\n<li>The similarity and clustering by the word2vec feature was also used for other feature extraction.</li></ul></li>\n<li><p>Graph features<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2Fb4fd67c9bad9534451cf405bc27d04ef%2FGraph%20figure.jpg?generation=1610196761780811&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Based on the users' answer histories, we created a directed graph, whose nodes are question_ids and edges' weights are defined by the number of users' transitions, and used the following node feature as question features.<ul>\n<li>node metrics (eigenvector_centrality, betweenness_centrality, trophic_levels <a href=\"https://networkx.org/documentation/stable/reference/algorithms/centrality.html\" target=\"_blank\">NetworkX document</a>)</li>\n<li>node embedding (DeepWalk <a href=\"http://www.perozzi.net/publications/14_kdd_deepwalk.pdf\" target=\"_blank\">paper</a>, struc2vec <a href=\"https://arxiv.org/pdf/1704.03165.pdf\" target=\"_blank\">paper</a>, <a href=\"https://github.com/shenweichen/GraphEmbedding\" target=\"_blank\">implementation</a>)</li>\n<li>SVD for adjacency matrix</li></ul></li></ul></li></ul></li>\n<li><p>User history features</p>\n<ul>\n<li>Answer count,  Correct answer rate<ul>\n<li>We use the user's correct answer rate and the correct answer rate in the last N times.</li>\n<li>We also used the correct answer rate in the first ten times, one day and one week.</li></ul></li>\n<li>Time-related<ul>\n<li>difftime (timestamp - previous_timestamp) worked especially well</li>\n<li>We used some features related to difftime, including difftime with the most recent 5 step timestamp and their statistics.</li></ul></li>\n<li>Question-related<ul>\n<li>Whether the user has answered the target question in the past or not, and the number of times the user has answered the question in the past were worked.</li>\n<li>We also use correct answer rate for questions in the same cluster (clustered by k-means using similarities based on Word2vec features described above, with the number of clusters set to 100)</li>\n<li>We also used some features related to the similarity between the target question and the recent questions or these parts (the similarities were based on Word2vec features described above).</li></ul></li>\n<li>Part-related<ul>\n<li>The correct answer rate in the part of the question users are answering was worked.</li>\n<li>We also used the count and correct answer rate of each part.</li></ul></li>\n<li>Tag-related<ul>\n<li>The correct answer rate for each tag of the user is kept. The statistics (max, min, mean) of the user's correct answer rate for each tag in the question being answered are used.</li></ul></li>\n<li>Lecture-related<ul>\n<li>The number of answers from the user's most recent lecture worked best for lecture-related features.</li></ul></li></ul></li>\n</ul>\n<h1>Model</h1>\n<h3>Catboost</h3>\n<ul>\n<li>Data split: tito CV</li>\n<li>Input<ul>\n<li>160 features</li></ul></li>\n<li>CV: 0.803~0.804, PublicLB: 0.801~0.802</li>\n</ul>\n<h3>Transformer (SAINT-like model)</h3>\n<ul>\n<li>Data split: tito CV</li>\n<li>Input (sequence length 120)<ul>\n<li>Encoder<ul>\n<li>question_id, part, tag, difftime</li></ul></li>\n<li>Decoder<ul>\n<li>answered_correctly, elasped_time</li></ul></li></ul></li>\n<li>CV: 0.797~0.798</li>\n</ul>\n<h1>Stacking</h1>\n<ul>\n<li>Model: Catboost</li>\n<li>Data split: k-fold CV (k=9) of 2.5M records</li>\n<li>Input<ul>\n<li>Predictions of Catboost (6 models, random seeds) and Transformer (2 models, a slight variation of hyperparameters and  structure)</li>\n<li>Top 75 higher importance features based on feature importance of CatBoost</li></ul></li>\n<li>3 Seed average</li>\n<li>CV: 0.810, PublicLB: 0.807, PrivateLB: 0.809</li>\n</ul>\n<hr>\n<p>Thank you for your attention! Our code is available in</p>\n<ul>\n<li><a href=\"https://github.com/haradai1262/kaggle_riiid-test-answer-prediction\" target=\"_blank\">https://github.com/haradai1262/kaggle_riiid-test-answer-prediction</a></li>\n<li><a href=\"https://www.kaggle.com/haradataman/riiid-28th-solution-inference-only\" target=\"_blank\">Inference (kaggle notebook)</a></li>\n</ul>",
  "messages": [
    {
      "id": "1145964",
      "postDate": "01/09/2021 13:01:43",
      "content": "<p>I want to thank the hosts for organizing a meaningful competition that reflected the critical constraints of real-world difficulties (time series API, limited RAM, and inference time). Also, I thanks my teammate NARI who worked hard with me until the end of the competition.</p>\n<h1>Overview</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2F5cfa0a830e6e45584a9088a794720ef7%2FOverview.jpg?generation=1610195752423547&amp;alt=media\" alt=\"\"></p>\n<p>Our final solution is an ensemble of Catboost and Transformer. Although Transformer could not create a high CV model in time like the top teams, it contributes a lot to boosting the stacking model score. Our solution's validation strategy and feature engineering pipeline were heavily influenced by tito's notebooks (<a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">validation strategy</a>, <a href=\"https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering\" target=\"_blank\">feature engineering</a>). We definitely would not have reached this score if he had not been shared early in the competition.</p>\n<h1>Features</h1>\n<p>We extracted 160 features for training Catboost. </p>\n<p>The main features are as follows. The detailed contribution to the score of each feature has not been confirmed, but the feature importance by CatBoost is available <a href=\"https://github.com/haradai1262/kaggle_riiid-test-answer-prediction/blob/main/notebook/check_feature_importance.ipynb\" target=\"_blank\">here</a></p>\n<ul>\n<li><p>Content-related features</p>\n<ul>\n<li><p>Aggregation features</p>\n<ul>\n<li>mean, std for answered_correctly</li>\n<li>mean for elapsed time</li>\n<li>We also used aggregation in records only for the second and subsequent answers to the same question by the user.</li></ul></li>\n<li><p>Word2vec features<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2Fa381717acbf8460946ef6455aef948c4%2Fword2vec%20figure.jpg?generation=1610196796564277&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>We extract Word2vec features using each user's answer history as sentences, each question_id and its associated part, tag_id, and lecture_id as words.</li>\n<li>We also extracted cases in which users' answer histories for only correct answers, and only incorrect answers were treated as separate sentences.</li>\n<li>The similarity and clustering by the word2vec feature was also used for other feature extraction.</li></ul></li>\n<li><p>Graph features<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2Fb4fd67c9bad9534451cf405bc27d04ef%2FGraph%20figure.jpg?generation=1610196761780811&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>Based on the users' answer histories, we created a directed graph, whose nodes are question_ids and edges' weights are defined by the number of users' transitions, and used the following node feature as question features.<ul>\n<li>node metrics (eigenvector_centrality, betweenness_centrality, trophic_levels <a href=\"https://networkx.org/documentation/stable/reference/algorithms/centrality.html\" target=\"_blank\">NetworkX document</a>)</li>\n<li>node embedding (DeepWalk <a href=\"http://www.perozzi.net/publications/14_kdd_deepwalk.pdf\" target=\"_blank\">paper</a>, struc2vec <a href=\"https://arxiv.org/pdf/1704.03165.pdf\" target=\"_blank\">paper</a>, <a href=\"https://github.com/shenweichen/GraphEmbedding\" target=\"_blank\">implementation</a>)</li>\n<li>SVD for adjacency matrix</li></ul></li></ul></li></ul></li>\n<li><p>User history features</p>\n<ul>\n<li>Answer count,  Correct answer rate<ul>\n<li>We use the user's correct answer rate and the correct answer rate in the last N times.</li>\n<li>We also used the correct answer rate in the first ten times, one day and one week.</li></ul></li>\n<li>Time-related<ul>\n<li>difftime (timestamp - previous_timestamp) worked especially well</li>\n<li>We used some features related to difftime, including difftime with the most recent 5 step timestamp and their statistics.</li></ul></li>\n<li>Question-related<ul>\n<li>Whether the user has answered the target question in the past or not, and the number of times the user has answered the question in the past were worked.</li>\n<li>We also use correct answer rate for questions in the same cluster (clustered by k-means using similarities based on Word2vec features described above, with the number of clusters set to 100)</li>\n<li>We also used some features related to the similarity between the target question and the recent questions or these parts (the similarities were based on Word2vec features described above).</li></ul></li>\n<li>Part-related<ul>\n<li>The correct answer rate in the part of the question users are answering was worked.</li>\n<li>We also used the count and correct answer rate of each part.</li></ul></li>\n<li>Tag-related<ul>\n<li>The correct answer rate for each tag of the user is kept. The statistics (max, min, mean) of the user's correct answer rate for each tag in the question being answered are used.</li></ul></li>\n<li>Lecture-related<ul>\n<li>The number of answers from the user's most recent lecture worked best for lecture-related features.</li></ul></li></ul></li>\n</ul>\n<h1>Model</h1>\n<h3>Catboost</h3>\n<ul>\n<li>Data split: tito CV</li>\n<li>Input<ul>\n<li>160 features</li></ul></li>\n<li>CV: 0.803~0.804, PublicLB: 0.801~0.802</li>\n</ul>\n<h3>Transformer (SAINT-like model)</h3>\n<ul>\n<li>Data split: tito CV</li>\n<li>Input (sequence length 120)<ul>\n<li>Encoder<ul>\n<li>question_id, part, tag, difftime</li></ul></li>\n<li>Decoder<ul>\n<li>answered_correctly, elasped_time</li></ul></li></ul></li>\n<li>CV: 0.797~0.798</li>\n</ul>\n<h1>Stacking</h1>\n<ul>\n<li>Model: Catboost</li>\n<li>Data split: k-fold CV (k=9) of 2.5M records</li>\n<li>Input<ul>\n<li>Predictions of Catboost (6 models, random seeds) and Transformer (2 models, a slight variation of hyperparameters and  structure)</li>\n<li>Top 75 higher importance features based on feature importance of CatBoost</li></ul></li>\n<li>3 Seed average</li>\n<li>CV: 0.810, PublicLB: 0.807, PrivateLB: 0.809</li>\n</ul>\n<hr>\n<p>Thank you for your attention! Our code is available in</p>\n<ul>\n<li><a href=\"https://github.com/haradai1262/kaggle_riiid-test-answer-prediction\" target=\"_blank\">https://github.com/haradai1262/kaggle_riiid-test-answer-prediction</a></li>\n<li><a href=\"https://www.kaggle.com/haradataman/riiid-28th-solution-inference-only\" target=\"_blank\">Inference (kaggle notebook)</a></li>\n</ul>",
      "rawMarkdown": "I want to thank the hosts for organizing a meaningful competition that reflected the critical constraints of real-world difficulties (time series API, limited RAM, and inference time). Also, I thanks my teammate NARI who worked hard with me until the end of the competition.\n\n# Overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2F5cfa0a830e6e45584a9088a794720ef7%2FOverview.jpg?generation=1610195752423547&alt=media)\n\nOur final solution is an ensemble of Catboost and Transformer. Although Transformer could not create a high CV model in time like the top teams, it contributes a lot to boosting the stacking model score. Our solution's validation strategy and feature engineering pipeline were heavily influenced by tito's notebooks ([validation strategy](https://www.kaggle.com/its7171/cv-strategy), [feature engineering](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering)). We definitely would not have reached this score if he had not been shared early in the competition.\n\n# Features\n\nWe extracted 160 features for training Catboost. \n\nThe main features are as follows. The detailed contribution to the score of each feature has not been confirmed, but the feature importance by CatBoost is available [here](https://github.com/haradai1262/kaggle_riiid-test-answer-prediction/blob/main/notebook/check_feature_importance.ipynb)\n\n- Content-related features\n    - Aggregation features\n        - mean, std for answered_correctly\n        - mean for elapsed time\n        - We also used aggregation in records only for the second and subsequent answers to the same question by the user.\n    - Word2vec features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2Fa381717acbf8460946ef6455aef948c4%2Fword2vec%20figure.jpg?generation=1610196796564277&alt=media)\n        - We extract Word2vec features using each user's answer history as sentences, each question_id and its associated part, tag_id, and lecture_id as words.\n        - We also extracted cases in which users' answer histories for only correct answers, and only incorrect answers were treated as separate sentences.\n        - The similarity and clustering by the word2vec feature was also used for other feature extraction.\n\n    - Graph features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2Fb4fd67c9bad9534451cf405bc27d04ef%2FGraph%20figure.jpg?generation=1610196761780811&alt=media)\n        - Based on the users' answer histories, we created a directed graph, whose nodes are question_ids and edges' weights are defined by the number of users' transitions, and used the following node feature as question features.\n            - node metrics (eigenvector_centrality, betweenness_centrality, trophic_levels [NetworkX document](https://networkx.org/documentation/stable/reference/algorithms/centrality.html))\n            - node embedding (DeepWalk [paper](http://www.perozzi.net/publications/14_kdd_deepwalk.pdf), struc2vec [paper](https://arxiv.org/pdf/1704.03165.pdf), [implementation](https://github.com/shenweichen/GraphEmbedding))\n            - SVD for adjacency matrix\n- User history features\n    - Answer count,  Correct answer rate\n        - We use the user's correct answer rate and the correct answer rate in the last N times.\n        - We also used the correct answer rate in the first ten times, one day and one week.\n    - Time-related\n        - difftime (timestamp - previous_timestamp) worked especially well\n        - We used some features related to difftime, including difftime with the most recent 5 step timestamp and their statistics.\n    - Question-related\n        - Whether the user has answered the target question in the past or not, and the number of times the user has answered the question in the past were worked.\n        - We also use correct answer rate for questions in the same cluster (clustered by k-means using similarities based on Word2vec features described above, with the number of clusters set to 100)\n        - We also used some features related to the similarity between the target question and the recent questions or these parts (the similarities were based on Word2vec features described above).\n    - Part-related\n        - The correct answer rate in the part of the question users are answering was worked.\n        - We also used the count and correct answer rate of each part.\n    - Tag-related\n        - The correct answer rate for each tag of the user is kept. The statistics (max, min, mean) of the user's correct answer rate for each tag in the question being answered are used.\n    - Lecture-related\n        - The number of answers from the user's most recent lecture worked best for lecture-related features.\n\n\n# Model\n\n\n### Catboost\n\n- Data split: tito CV\n- Input\n    - 160 features\n- CV: 0.803~0.804, PublicLB: 0.801~0.802\n\n### Transformer (SAINT-like model)\n\n- Data split: tito CV\n- Input (sequence length 120)\n    - Encoder\n        - question_id, part, tag, difftime\n    - Decoder\n        - answered_correctly, elasped_time\n- CV: 0.797~0.798\n\n# Stacking\n\n- Model: Catboost\n- Data split: k-fold CV (k=9) of 2.5M records\n- Input\n    - Predictions of Catboost (6 models, random seeds) and Transformer (2 models, a slight variation of hyperparameters and  structure)\n    - Top 75 higher importance features based on feature importance of CatBoost\n- 3 Seed average\n- CV: 0.810, PublicLB: 0.807, PrivateLB: 0.809\n\n---\n\nThank you for your attention! Our code is available in\n\n- [https://github.com/haradai1262/kaggle_riiid-test-answer-prediction](https://github.com/haradai1262/kaggle_riiid-test-answer-prediction)\n- [Inference (kaggle notebook)](https://www.kaggle.com/haradataman/riiid-28th-solution-inference-only)",
      "votes": null
    },
    {
      "id": "1146133",
      "postDate": "01/09/2021 14:59:46",
      "content": "<p>Congrats and thanks for sharing solution and code <a href=\"https://www.kaggle.com/haradataman\" target=\"_blank\">@haradataman</a> </p>",
      "rawMarkdown": "Congrats and thanks for sharing solution and code @haradataman",
      "votes": null
    },
    {
      "id": "1146679",
      "postDate": "01/09/2021 23:56:02",
      "content": "<p>Thanks for sharing. Graph features seem to work, but I'm wondering if it would take longer time?</p>",
      "rawMarkdown": "Thanks for sharing. Graph features seem to work, but I'm wondering if it would take longer time?",
      "votes": null
    },
    {
      "id": "1146706",
      "postDate": "01/10/2021 00:53:49",
      "content": "<p>Thanks for the question! <br>\nThe cost of creating the graph is small, but calculating node metrics (networkX implementation, single thread CPU) and node embeddings (gensim-based implementation, multi-threads CPU) took several dozen hours (I don't remember exactly) in the implementation I used. I didn't care so much about the cost as it was only needed for the training phase.</p>",
      "rawMarkdown": "Thanks for the question! \nThe cost of creating the graph is small, but calculating node metrics (networkX implementation, single thread CPU) and node embeddings (gensim-based implementation, multi-threads CPU) took several dozen hours (I don't remember exactly) in the implementation I used. I didn't care so much about the cost as it was only needed for the training phase.",
      "votes": null
    },
    {
      "id": "1146710",
      "postDate": "01/10/2021 01:05:13",
      "content": "<p>I see. Thank you for the additional info.</p>",
      "rawMarkdown": "I see. Thank you for the additional info.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1146133,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "01/09/2021 14:59:46",
      "content": "<p>Congrats and thanks for sharing solution and code <a href=\"https://www.kaggle.com/haradataman\" target=\"_blank\">@haradataman</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1146679,
      "author_name": "sishihara",
      "author_url": "",
      "post_date": "01/09/2021 23:56:02",
      "content": "<p>Thanks for sharing. Graph features seem to work, but I'm wondering if it would take longer time?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1146706,
          "author_name": "haradataman",
          "author_url": "",
          "post_date": "01/10/2021 00:53:49",
          "content": "<p>Thanks for the question! <br>\nThe cost of creating the graph is small, but calculating node metrics (networkX implementation, single thread CPU) and node embeddings (gensim-based implementation, multi-threads CPU) took several dozen hours (I don't remember exactly) in the implementation I used. I didn't care so much about the cost as it was only needed for the training phase.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146710,
          "author_name": "sishihara",
          "author_url": "",
          "post_date": "01/10/2021 01:05:13",
          "content": "<p>I see. Thank you for the additional info.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1145964": "I want to thank the hosts for organizing a meaningful competition that reflected the critical constraints of real-world difficulties (time series API, limited RAM, and inference time). Also, I thanks my teammate NARI who worked hard with me until the end of the competition.\n\n# Overview\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2F5cfa0a830e6e45584a9088a794720ef7%2FOverview.jpg?generation=1610195752423547&alt=media)\n\nOur final solution is an ensemble of Catboost and Transformer. Although Transformer could not create a high CV model in time like the top teams, it contributes a lot to boosting the stacking model score. Our solution's validation strategy and feature engineering pipeline were heavily influenced by tito's notebooks ([validation strategy](https://www.kaggle.com/its7171/cv-strategy), [feature engineering](https://www.kaggle.com/its7171/lgbm-with-loop-feature-engineering)). We definitely would not have reached this score if he had not been shared early in the competition.\n\n# Features\n\nWe extracted 160 features for training Catboost. \n\nThe main features are as follows. The detailed contribution to the score of each feature has not been confirmed, but the feature importance by CatBoost is available [here](https://github.com/haradai1262/kaggle_riiid-test-answer-prediction/blob/main/notebook/check_feature_importance.ipynb)\n\n- Content-related features\n    - Aggregation features\n        - mean, std for answered_correctly\n        - mean for elapsed time\n        - We also used aggregation in records only for the second and subsequent answers to the same question by the user.\n    - Word2vec features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2Fa381717acbf8460946ef6455aef948c4%2Fword2vec%20figure.jpg?generation=1610196796564277&alt=media)\n        - We extract Word2vec features using each user's answer history as sentences, each question_id and its associated part, tag_id, and lecture_id as words.\n        - We also extracted cases in which users' answer histories for only correct answers, and only incorrect answers were treated as separate sentences.\n        - The similarity and clustering by the word2vec feature was also used for other feature extraction.\n\n    - Graph features\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1360119%2Fb4fd67c9bad9534451cf405bc27d04ef%2FGraph%20figure.jpg?generation=1610196761780811&alt=media)\n        - Based on the users' answer histories, we created a directed graph, whose nodes are question_ids and edges' weights are defined by the number of users' transitions, and used the following node feature as question features.\n            - node metrics (eigenvector_centrality, betweenness_centrality, trophic_levels [NetworkX document](https://networkx.org/documentation/stable/reference/algorithms/centrality.html))\n            - node embedding (DeepWalk [paper](http://www.perozzi.net/publications/14_kdd_deepwalk.pdf), struc2vec [paper](https://arxiv.org/pdf/1704.03165.pdf), [implementation](https://github.com/shenweichen/GraphEmbedding))\n            - SVD for adjacency matrix\n- User history features\n    - Answer count,  Correct answer rate\n        - We use the user's correct answer rate and the correct answer rate in the last N times.\n        - We also used the correct answer rate in the first ten times, one day and one week.\n    - Time-related\n        - difftime (timestamp - previous_timestamp) worked especially well\n        - We used some features related to difftime, including difftime with the most recent 5 step timestamp and their statistics.\n    - Question-related\n        - Whether the user has answered the target question in the past or not, and the number of times the user has answered the question in the past were worked.\n        - We also use correct answer rate for questions in the same cluster (clustered by k-means using similarities based on Word2vec features described above, with the number of clusters set to 100)\n        - We also used some features related to the similarity between the target question and the recent questions or these parts (the similarities were based on Word2vec features described above).\n    - Part-related\n        - The correct answer rate in the part of the question users are answering was worked.\n        - We also used the count and correct answer rate of each part.\n    - Tag-related\n        - The correct answer rate for each tag of the user is kept. The statistics (max, min, mean) of the user's correct answer rate for each tag in the question being answered are used.\n    - Lecture-related\n        - The number of answers from the user's most recent lecture worked best for lecture-related features.\n\n\n# Model\n\n\n### Catboost\n\n- Data split: tito CV\n- Input\n    - 160 features\n- CV: 0.803~0.804, PublicLB: 0.801~0.802\n\n### Transformer (SAINT-like model)\n\n- Data split: tito CV\n- Input (sequence length 120)\n    - Encoder\n        - question_id, part, tag, difftime\n    - Decoder\n        - answered_correctly, elasped_time\n- CV: 0.797~0.798\n\n# Stacking\n\n- Model: Catboost\n- Data split: k-fold CV (k=9) of 2.5M records\n- Input\n    - Predictions of Catboost (6 models, random seeds) and Transformer (2 models, a slight variation of hyperparameters and  structure)\n    - Top 75 higher importance features based on feature importance of CatBoost\n- 3 Seed average\n- CV: 0.810, PublicLB: 0.807, PrivateLB: 0.809\n\n---\n\nThank you for your attention! Our code is available in\n\n- [https://github.com/haradai1262/kaggle_riiid-test-answer-prediction](https://github.com/haradai1262/kaggle_riiid-test-answer-prediction)\n- [Inference (kaggle notebook)](https://www.kaggle.com/haradataman/riiid-28th-solution-inference-only)",
    "1146133": "Congrats and thanks for sharing solution and code @haradataman",
    "1146679": "Thanks for sharing. Graph features seem to work, but I'm wondering if it would take longer time?",
    "1146706": "Thanks for the question! \nThe cost of creating the graph is small, but calculating node metrics (networkX implementation, single thread CPU) and node embeddings (gensim-based implementation, multi-threads CPU) took several dozen hours (I don't remember exactly) in the implementation I used. I didn't care so much about the cost as it was only needed for the training phase.",
    "1146710": "I see. Thank you for the additional info."
  },
  "source": "meta"
}