{
  "id": 209580,
  "title": "403 place solution",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209580",
  "author_name": "Takamichi Toda",
  "post_date": "2021-01-08T00:12:49.155000",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<h1>403 place solution</h1>\n<p>My main strategy is training my model by using BigQuery ML because I wanted to learn it.<br>\nI think it was a good way for me. And one other good thing about BigQuery ML is I can use all data to train.</p>\n<h2>Model</h2>\n<h3>XGBoost</h3>\n<p>It looks that most people use LGBM but I use XGBoost.<br>\nI'm not familiar GBDT algorithm, I hear that LGBM is better than XGBoost.<br>\nUnfortunately, BigQuery ML only has XGBoost. So, I use XGBoost.</p>\n<p>BigQuery ML has other ML models. For example, Logistic Regression, ARIMA, and Neural Network.<br>\nI tried them but XGBoost was best.</p>\n<h3>Parameter</h3>\n<p>This is my BigQuery code to train my XGB model.<br>\nThe parameter can be seen here.</p>\n<pre><code>CREATE MODEL `&lt;My model PATH&gt;`\nOPTIONS(MODEL_TYPE='BOOSTED_TREE_CLASSIFIER',\n        BOOSTER_TYPE = 'GBTREE',\n        NUM_PARALLEL_TREE = 1,\n        MAX_ITERATIONS = 300,\n        TREE_METHOD = 'HIST',\n        EARLY_STOP = False,\n        MIN_REL_PROGRESS=0.0001,\n        LEARN_RATE =0.3,\n        MAX_TREE_DEPTH=11,\n        COLSAMPLE_BYTREE=1.0,\n        COLSAMPLE_BYLEVEL=0.4,\n        SUBSAMPLE = 0.9,\n        MIN_TREE_CHILD_WEIGHT=2,\n        L1_REG=0,\n        L2_REG=1.0,\n        INPUT_LABEL_COLS = ['answered_correctly'],\n        DATA_SPLIT_METHOD='CUSTOM',\n        DATA_SPLIT_COL='is_test')    \nAS \nSELECT `&lt;My Feature Tables&gt;`\n</code></pre>\n<h2>CV strategy</h2>\n<h3>leave-one-out</h3>\n<p>I use <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this strategy</a> by leave-one-out. It has looked good to me.</p>\n<h2>Ensemble</h2>\n<p>I use the best XGBoost model and two SAKT models.<br>\nSAKT models are made by forking <a href=\"https://www.kaggle.com/tarique7/v4-fork-of-riiid-sakt-model-full\" target=\"_blank\">this notebook</a>.<br>\nI little updated three in this notebook that:</p>\n<ol>\n<li>Randomly cut to a length of 200 in train data. (original code had used last 200 data)</li>\n<li>Save best score epoch model.</li>\n<li>Parameter tuning.</li>\n</ol>\n<p>Each score is here.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGB</td>\n<td>0.7784</td>\n<td>0.781</td>\n</tr>\n<tr>\n<td>SAKT1</td>\n<td>0.772</td>\n<td>0.774</td>\n</tr>\n<tr>\n<td>SAKT2</td>\n<td>0.772</td>\n<td>0.773</td>\n</tr>\n</tbody>\n</table>\n<p>I weighted averaging 3:1:1 this model and I got 0.787 in the public leaderboard.</p>\n<p>I tried iteration ensemble and CV fold ensemble in XGBoost, however, both got the same score as my single model.</p>\n<p>Final my submission notebook is <a href=\"https://www.kaggle.com/takamichitoda/riiid-infer-v8-xgb-sakt?scriptVersionId=51291197\" target=\"_blank\">here</a>.</p>\n<h2>Feature</h2>\n<p>My features is:</p>\n<ul>\n<li>aggregation by contents<ul>\n<li>answered_correctl: count/sum/avg/std</li>\n<li>timestamp: max/min/avg/std</li>\n<li>user_id: count/unique_count</li>\n<li>task_container_id: max/min/avg/std</li>\n<li>user_answer_0~3: sum/avg/std</li>\n<li>question_elapsed_time: max/min/avg/std</li>\n<li>question_had_explanation: sum/avg/std</li></ul></li>\n<li>each question tags cumulative sum<ul>\n<li>number of try</li>\n<li>number of correct answer</li>\n<li>correct answer rate</li></ul></li>\n<li>each lecture tag cumulative sum<ul>\n<li>number of try</li></ul></li>\n<li>time window features that size is 200 <ul>\n<li>answer correctry: sum/avg/std</li>\n<li>question_elapsed_time: avg/max</li>\n<li>question_had_explanation: sum/avg/std</li>\n<li>timestamp lag: avg/max</li>\n<li>number of content_id=0</li>\n<li>number of content_id=1</li></ul></li>\n<li>timestamp lag</li>\n</ul>\n<h2>Reference</h2>\n<h3>Notebook</h3>\n<ul>\n<li>Create SQL: <a href=\"https://www.kaggle.com/takamichitoda/riiid-create-sql\" target=\"_blank\">https://www.kaggle.com/takamichitoda/riiid-create-sql</a></li>\n<li>Create Tags Onehot Encode Table: <a href=\"https://www.kaggle.com/takamichitoda/riiid-create-feature-table-csv\" target=\"_blank\">https://www.kaggle.com/takamichitoda/riiid-create-feature-table-csv</a></li>\n<li>Create CV index: <a href=\"https://www.kaggle.com/takamichitoda/riiid-make-cv-index\" target=\"_blank\">https://www.kaggle.com/takamichitoda/riiid-make-cv-index</a></li>\n</ul>\n<p>EDIT:</p>\n<ul>\n<li><a href=\"https://www.ai-shift.co.jp/techblog/1505\" target=\"_blank\">my blog</a>(Japanese)</li>\n<li><a href=\"https://github.com/trtd56/Riiid/tree/master/work/sql\" target=\"_blank\">SQL to generate Window Features.</a></li>\n</ul>",
  "messages": [
    {
      "id": 1143533,
      "postDate": "2021-01-08T00:12:49.157Z",
      "content": "<h1>403 place solution</h1>\n<p>My main strategy is training my model by using BigQuery ML because I wanted to learn it.<br>\nI think it was a good way for me. And one other good thing about BigQuery ML is I can use all data to train.</p>\n<h2>Model</h2>\n<h3>XGBoost</h3>\n<p>It looks that most people use LGBM but I use XGBoost.<br>\nI'm not familiar GBDT algorithm, I hear that LGBM is better than XGBoost.<br>\nUnfortunately, BigQuery ML only has XGBoost. So, I use XGBoost.</p>\n<p>BigQuery ML has other ML models. For example, Logistic Regression, ARIMA, and Neural Network.<br>\nI tried them but XGBoost was best.</p>\n<h3>Parameter</h3>\n<p>This is my BigQuery code to train my XGB model.<br>\nThe parameter can be seen here.</p>\n<pre><code>CREATE MODEL `&lt;My model PATH&gt;`\nOPTIONS(MODEL_TYPE='BOOSTED_TREE_CLASSIFIER',\n        BOOSTER_TYPE = 'GBTREE',\n        NUM_PARALLEL_TREE = 1,\n        MAX_ITERATIONS = 300,\n        TREE_METHOD = 'HIST',\n        EARLY_STOP = False,\n        MIN_REL_PROGRESS=0.0001,\n        LEARN_RATE =0.3,\n        MAX_TREE_DEPTH=11,\n        COLSAMPLE_BYTREE=1.0,\n        COLSAMPLE_BYLEVEL=0.4,\n        SUBSAMPLE = 0.9,\n        MIN_TREE_CHILD_WEIGHT=2,\n        L1_REG=0,\n        L2_REG=1.0,\n        INPUT_LABEL_COLS = ['answered_correctly'],\n        DATA_SPLIT_METHOD='CUSTOM',\n        DATA_SPLIT_COL='is_test')    \nAS \nSELECT `&lt;My Feature Tables&gt;`\n</code></pre>\n<h2>CV strategy</h2>\n<h3>leave-one-out</h3>\n<p>I use <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">this strategy</a> by leave-one-out. It has looked good to me.</p>\n<h2>Ensemble</h2>\n<p>I use the best XGBoost model and two SAKT models.<br>\nSAKT models are made by forking <a href=\"https://www.kaggle.com/tarique7/v4-fork-of-riiid-sakt-model-full\" target=\"_blank\">this notebook</a>.<br>\nI little updated three in this notebook that:</p>\n<ol>\n<li>Randomly cut to a length of 200 in train data. (original code had used last 200 data)</li>\n<li>Save best score epoch model.</li>\n<li>Parameter tuning.</li>\n</ol>\n<p>Each score is here.</p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>CV</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGB</td>\n<td>0.7784</td>\n<td>0.781</td>\n</tr>\n<tr>\n<td>SAKT1</td>\n<td>0.772</td>\n<td>0.774</td>\n</tr>\n<tr>\n<td>SAKT2</td>\n<td>0.772</td>\n<td>0.773</td>\n</tr>\n</tbody>\n</table>\n<p>I weighted averaging 3:1:1 this model and I got 0.787 in the public leaderboard.</p>\n<p>I tried iteration ensemble and CV fold ensemble in XGBoost, however, both got the same score as my single model.</p>\n<p>Final my submission notebook is <a href=\"https://www.kaggle.com/takamichitoda/riiid-infer-v8-xgb-sakt?scriptVersionId=51291197\" target=\"_blank\">here</a>.</p>\n<h2>Feature</h2>\n<p>My features is:</p>\n<ul>\n<li>aggregation by contents<ul>\n<li>answered_correctl: count/sum/avg/std</li>\n<li>timestamp: max/min/avg/std</li>\n<li>user_id: count/unique_count</li>\n<li>task_container_id: max/min/avg/std</li>\n<li>user_answer_0~3: sum/avg/std</li>\n<li>question_elapsed_time: max/min/avg/std</li>\n<li>question_had_explanation: sum/avg/std</li></ul></li>\n<li>each question tags cumulative sum<ul>\n<li>number of try</li>\n<li>number of correct answer</li>\n<li>correct answer rate</li></ul></li>\n<li>each lecture tag cumulative sum<ul>\n<li>number of try</li></ul></li>\n<li>time window features that size is 200 <ul>\n<li>answer correctry: sum/avg/std</li>\n<li>question_elapsed_time: avg/max</li>\n<li>question_had_explanation: sum/avg/std</li>\n<li>timestamp lag: avg/max</li>\n<li>number of content_id=0</li>\n<li>number of content_id=1</li></ul></li>\n<li>timestamp lag</li>\n</ul>\n<h2>Reference</h2>\n<h3>Notebook</h3>\n<ul>\n<li>Create SQL: <a href=\"https://www.kaggle.com/takamichitoda/riiid-create-sql\" target=\"_blank\">https://www.kaggle.com/takamichitoda/riiid-create-sql</a></li>\n<li>Create Tags Onehot Encode Table: <a href=\"https://www.kaggle.com/takamichitoda/riiid-create-feature-table-csv\" target=\"_blank\">https://www.kaggle.com/takamichitoda/riiid-create-feature-table-csv</a></li>\n<li>Create CV index: <a href=\"https://www.kaggle.com/takamichitoda/riiid-make-cv-index\" target=\"_blank\">https://www.kaggle.com/takamichitoda/riiid-make-cv-index</a></li>\n</ul>\n<p>EDIT:</p>\n<ul>\n<li><a href=\"https://www.ai-shift.co.jp/techblog/1505\" target=\"_blank\">my blog</a>(Japanese)</li>\n<li><a href=\"https://github.com/trtd56/Riiid/tree/master/work/sql\" target=\"_blank\">SQL to generate Window Features.</a></li>\n</ul>",
      "rawMarkdown": "# 403 place solution\n\nMy main strategy is training my model by using BigQuery ML because I wanted to learn it.\nI think it was a good way for me. And one other good thing about BigQuery ML is I can use all data to train.\n\n## Model\n### XGBoost\nIt looks that most people use LGBM but I use XGBoost.\nI'm not familiar GBDT algorithm, I hear that LGBM is better than XGBoost.\nUnfortunately, BigQuery ML only has XGBoost. So, I use XGBoost.\n\nBigQuery ML has other ML models. For example, Logistic Regression, ARIMA, and Neural Network.\nI tried them but XGBoost was best.\n\n### Parameter\n\nThis is my BigQuery code to train my XGB model.\nThe parameter can be seen here.\n\n```sql\nCREATE MODEL `<My model PATH>`\nOPTIONS(MODEL_TYPE='BOOSTED_TREE_CLASSIFIER',\n        BOOSTER_TYPE = 'GBTREE',\n        NUM_PARALLEL_TREE = 1,\n        MAX_ITERATIONS = 300,\n        TREE_METHOD = 'HIST',\n        EARLY_STOP = False,\n        MIN_REL_PROGRESS=0.0001,\n        LEARN_RATE =0.3,\n        MAX_TREE_DEPTH=11,\n        COLSAMPLE_BYTREE=1.0,\n        COLSAMPLE_BYLEVEL=0.4,\n        SUBSAMPLE = 0.9,\n        MIN_TREE_CHILD_WEIGHT=2,\n        L1_REG=0,\n        L2_REG=1.0,\n        INPUT_LABEL_COLS = ['answered_correctly'],\n        DATA_SPLIT_METHOD='CUSTOM',\n        DATA_SPLIT_COL='is_test')    \nAS \nSELECT `<My Feature Tables>`\n```\n\n## CV strategy\n### leave-one-out\nI use [this strategy](https://www.kaggle.com/its7171/cv-strategy) by leave-one-out. It has looked good to me.\n\n## Ensemble\n\nI use the best XGBoost model and two SAKT models.\nSAKT models are made by forking [this notebook](https://www.kaggle.com/tarique7/v4-fork-of-riiid-sakt-model-full).\nI little updated three in this notebook that:\n1. Randomly cut to a length of 200 in train data. (original code had used last 200 data)\n2. Save best score epoch model.\n3. Parameter tuning.\n\nEach score is here.\n\n|model|CV|LB|\n| -- | -- | -- |\n|XGB|0.7784|0.781|\n|SAKT1|0.772|0.774|\n|SAKT2|0.772|0.773|\n\nI weighted averaging 3:1:1 this model and I got 0.787 in the public leaderboard.\n\nI tried iteration ensemble and CV fold ensemble in XGBoost, however, both got the same score as my single model.\n\nFinal my submission notebook is [here](https://www.kaggle.com/takamichitoda/riiid-infer-v8-xgb-sakt?scriptVersionId=51291197).\n\n## Feature\nMy features is:\n  - aggregation by contents\n    - answered_correctl: count/sum/avg/std\n    - timestamp: max/min/avg/std\n    - user_id: count/unique_count\n    - task_container_id: max/min/avg/std\n    - user_answer_0~3: sum/avg/std\n    - question_elapsed_time: max/min/avg/std\n    - question_had_explanation: sum/avg/std\n  - each question tags cumulative sum\n    - number of try\n    - number of correct answer\n    - correct answer rate\n  - each lecture tag cumulative sum\n    - number of try\n  - time window features that size is 200 \n    - answer correctry: sum/avg/std\n    - question_elapsed_time: avg/max\n    - question_had_explanation: sum/avg/std\n    - timestamp lag: avg/max\n    - number of content_id=0\n    - number of content_id=1\n  - timestamp lag\n\n## Reference\n### Notebook\n- Create SQL: https://www.kaggle.com/takamichitoda/riiid-create-sql\n- Create Tags Onehot Encode Table: https://www.kaggle.com/takamichitoda/riiid-create-feature-table-csv\n- Create CV index: https://www.kaggle.com/takamichitoda/riiid-make-cv-index\n\nEDIT:\n- [my blog](https://www.ai-shift.co.jp/techblog/1505)(Japanese)\n- [SQL to generate Window Features.](https://github.com/trtd56/Riiid/tree/master/work/sql)",
      "votes": 9
    },
    {
      "id": 1143565,
      "postDate": "2021-01-08T00:29:54.020Z",
      "content": "<p>Thanks for sharing. I'm a bit surprised to see you did almost everything in BigQuery.</p>",
      "rawMarkdown": "Thanks for sharing. I'm a bit surprised to see you did almost everything in BigQuery.",
      "votes": 1,
      "replies": [
        {
          "id": 1143675,
          "postDate": "2021-01-08T02:02:47.413Z",
          "content": "<p>Thank you comment!</p>\n<p>I have tried this idea by seen <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190270\" target=\"_blank\">tkm2261's discussion thread</a>.</p>\n<p>BigQuery ML merit is here:</p>\n<ul>\n<li>You can treat very large data easily.</li>\n<li>You can manage feature values with BigQuery</li>\n</ul>\n<p>BigQuery ML demerit is here:</p>\n<ul>\n<li>If you make a sloppy query, you will have to spend a lot of money.</li>\n<li>In the GBDT algorithm, it can only use XGB (looking at some solutions, LGBM is better.)</li>\n</ul>\n<p>I'm satisfied to learn many things.<br>\nBut it’s still regrettable and frustrating because I lost the medal since I shakedown.<br>\nI'm studying top solutions now for the next competition.</p>",
          "rawMarkdown": "Thank you comment!\n\nI have tried this idea by seen [tkm2261's discussion thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190270).\n\nBigQuery ML merit is here:\n- You can treat very large data easily.\n- You can manage feature values with BigQuery\n\nBigQuery ML demerit is here:\n- If you make a sloppy query, you will have to spend a lot of money.\n- In the GBDT algorithm, it can only use XGB (looking at some solutions, LGBM is better.)\n\nI'm satisfied to learn many things.\nBut it’s still regrettable and frustrating because I lost the medal since I shakedown.\nI'm studying top solutions now for the next competition.\n\n",
          "votes": 1
        },
        {
          "id": 1143682,
          "postDate": "2021-01-08T02:07:41.043Z",
          "content": "<p>I feel that BigQuery ML can be one of the powerful options (not only for competitions but particularly in the business use). Thank you again for sharing your experience on this competition.</p>",
          "rawMarkdown": "I feel that BigQuery ML can be one of the powerful options (not only for competitions but particularly in the business use). Thank you again for sharing your experience on this competition.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1585636,
      "postDate": "2021-11-17T11:59:38.453Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1143565,
      "author_name": "u++",
      "author_url": "",
      "post_date": "2021-01-08T00:29:54.020000",
      "content": "<p>Thanks for sharing. I'm a bit surprised to see you did almost everything in BigQuery.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1143675,
          "author_name": "Takamichi Toda",
          "author_url": "",
          "post_date": "2021-01-08T02:02:47.413000",
          "content": "<p>Thank you comment!</p>\n<p>I have tried this idea by seen <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/190270\" target=\"_blank\">tkm2261's discussion thread</a>.</p>\n<p>BigQuery ML merit is here:</p>\n<ul>\n<li>You can treat very large data easily.</li>\n<li>You can manage feature values with BigQuery</li>\n</ul>\n<p>BigQuery ML demerit is here:</p>\n<ul>\n<li>If you make a sloppy query, you will have to spend a lot of money.</li>\n<li>In the GBDT algorithm, it can only use XGB (looking at some solutions, LGBM is better.)</li>\n</ul>\n<p>I'm satisfied to learn many things.<br>\nBut it’s still regrettable and frustrating because I lost the medal since I shakedown.<br>\nI'm studying top solutions now for the next competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1143682,
          "author_name": "u++",
          "author_url": "",
          "post_date": "2021-01-08T02:07:41.043000",
          "content": "<p>I feel that BigQuery ML can be one of the powerful options (not only for competitions but particularly in the business use). Thank you again for sharing your experience on this competition.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1585636,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-17T11:59:38.453000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143533": "# 403 place solution\n\nMy main strategy is training my model by using BigQuery ML because I wanted to learn it.\nI think it was a good way for me. And one other good thing about BigQuery ML is I can use all data to train.\n\n## Model\n### XGBoost\nIt looks that most people use LGBM but I use XGBoost.\nI'm not familiar GBDT algorithm, I hear that LGBM is better than XGBoost.\nUnfortunately, BigQuery ML only has XGBoost. So, I use XGBoost.\n\nBigQuery ML has other ML models. For example, Logistic Regression, ARIMA, and Neural Network.\nI tried them but XGBoost was best.\n\n### Parameter\n\nThis is my BigQuery code to train my XGB model.\nThe parameter can be seen here.\n\n```sql\nCREATE MODEL `<My model PATH>`\nOPTIONS(MODEL_TYPE='BOOSTED_TREE_CLASSIFIER',\n        BOOSTER_TYPE = 'GBTREE',\n        NUM_PARALLEL_TREE = 1,\n        MAX_ITERATIONS = 300,\n        TREE_METHOD = 'HIST',\n        EARLY_STOP = False,\n        MIN_REL_PROGRESS=0.0001,\n        LEARN_RATE =0.3,\n        MAX_TREE_DEPTH=11,\n        COLSAMPLE_BYTREE=1.0,\n        COLSAMPLE_BYLEVEL=0.4,\n        SUBSAMPLE = 0.9,\n        MIN_TREE_CHILD_WEIGHT=2,\n        L1_REG=0,\n        L2_REG=1.0,\n        INPUT_LABEL_COLS = ['answered_correctly'],\n        DATA_SPLIT_METHOD='CUSTOM',\n        DATA_SPLIT_COL='is_test')    \nAS \nSELECT `<My Feature Tables>`\n```\n\n## CV strategy\n### leave-one-out\nI use [this strategy](https://www.kaggle.com/its7171/cv-strategy) by leave-one-out. It has looked good to me.\n\n## Ensemble\n\nI use the best XGBoost model and two SAKT models.\nSAKT models are made by forking [this notebook](https://www.kaggle.com/tarique7/v4-fork-of-riiid-sakt-model-full).\nI little updated three in this notebook that:\n1. Randomly cut to a length of 200 in train data. (original code had used last 200 data)\n2. Save best score epoch model.\n3. Parameter tuning.\n\nEach score is here.\n\n|model|CV|LB|\n| -- | -- | -- |\n|XGB|0.7784|0.781|\n|SAKT1|0.772|0.774|\n|SAKT2|0.772|0.773|\n\nI weighted averaging 3:1:1 this model and I got 0.787 in the public leaderboard.\n\nI tried iteration ensemble and CV fold ensemble in XGBoost, however, both got the same score as my single model.\n\nFinal my submission notebook is [here](https://www.kaggle.com/takamichitoda/riiid-infer-v8-xgb-sakt?scriptVersionId=51291197).\n\n## Feature\nMy features is:\n  - aggregation by contents\n    - answered_correctl: count/sum/avg/std\n    - timestamp: max/min/avg/std\n    - user_id: count/unique_count\n    - task_container_id: max/min/avg/std\n    - user_answer_0~3: sum/avg/std\n    - question_elapsed_time: max/min/avg/std\n    - question_had_explanation: sum/avg/std\n  - each question tags cumulative sum\n    - number of try\n    - number of correct answer\n    - correct answer rate\n  - each lecture tag cumulative sum\n    - number of try\n  - time window features that size is 200 \n    - answer correctry: sum/avg/std\n    - question_elapsed_time: avg/max\n    - question_had_explanation: sum/avg/std\n    - timestamp lag: avg/max\n    - number of content_id=0\n    - number of content_id=1\n  - timestamp lag\n\n## Reference\n### Notebook\n- Create SQL: https://www.kaggle.com/takamichitoda/riiid-create-sql\n- Create Tags Onehot Encode Table: https://www.kaggle.com/takamichitoda/riiid-create-feature-table-csv\n- Create CV index: https://www.kaggle.com/takamichitoda/riiid-make-cv-index\n\nEDIT:\n- [my blog](https://www.ai-shift.co.jp/techblog/1505)(Japanese)\n- [SQL to generate Window Features.](https://github.com/trtd56/Riiid/tree/master/work/sql)",
    "1143565": "Thanks for sharing. I'm a bit surprised to see you did almost everything in BigQuery.",
    "1585636": ""
  }
}