{
  "id": 209659,
  "title": "24th Place Solution saintpp + lgb private0.810",
  "url": "/competitions/riiid-test-answer-prediction/discussion/209659",
  "author_name": "cswwp",
  "post_date": "2021-01-08T06:54:09.614000",
  "votes": 28,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Firstly, thanks to Kaggle hold such a good competition, almost no shake. Congrats tops and very thanks to my teammates <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  <a href=\"https://www.kaggle.com/ekffar\" target=\"_blank\">@ekffar</a>. Finally our team finally public score 0.808, private 0.810, and final rank 24. we missed gold, it's ok, now i want to share our method</p>\n<p>because of jet lag, i have no read discuss in detail, so i just see my teammate <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> had shared our <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209625\" target=\"_blank\">solution</a> right now. As a complementary, not repeat, i just expand the detail based on  <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209625\" target=\"_blank\">jaideepvalani's solution</a></p>\n<h1><strong>Architecture</strong></h1>\n<pre><code>SaintPlusPlus, lgb\n</code></pre>\n<h1>Embeddings</h1>\n<h2>Exercise:</h2>\n<pre><code>Question_id\nPart\nCommunity id (tried, but no boosting, maybe need tuning)\nTask container id (tried, but no boosting, maybe need tuning)\n</code></pre>\n<h2>Interaction:</h2>\n<pre><code>lag time  (same container id, keep same with first one) (discrete embedding)\nprior question elapsed time(continuous embedding)\nprior had explain (discrete embedding)\nPrior lecture(discrete embedding)\n</code></pre>\n<h2>Response:</h2>\n<pre><code>Prior question correctness(discrete embedding)\n</code></pre>\n<h1>Data sample optimization</h1>\n<p>Data sampling is important in our experiment<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F347724%2F8884cd3f5fca04b3232adb038b131f0f%2Friiid.png?generation=1610088827710995&amp;alt=media\" alt=\"\"></p>\n<h1>optimization</h1>\n<p>## SAINTPP<br>\n   1 skip connection, we add skip connection for multi encoder layers inner, and multi decoder layers inner, the origin implement i use from github have no this skip connection <br>\n   2 add dropout for encoder, decoder, FFN<br>\n   3 try TCN causal conv after embedding, no boost, but now, i think it should tune and will work<br>\n   4 big seq len for saintpp give our team boosting, i think it maybe caused by too short seq make <br>\n       transformer link model struggle into local minest point </p>\n<h2>LGB</h2>\n<p>updating</p>\n<h1>Conclusion</h1>\n<p>Our best saintpp(seqlen 320, embdim224, 2 x encder+ 2 x decoder) get cv 0.8016 lb 0.806, and lgb near get 0.796. Finally submission is merge lgb(lb 796) + saintpp cv1(cv 8016) + saintpp cv3(cv 0.7992)  = public lb 0.808, and private 0.810.<br>\nshould do things:<br>\n  1 should continue optimize our net with casual conv(because transformer like is a global attention, some local correlation maybe ignored) and gru, but finally have no time<br>\n  2 should try big net continue</p>\n<p>Thanks again to my teammates, they give me a lots of help and let me learn a lot. Going  for gold next Compete, keep moving.</p>",
  "messages": [
    {
      "id": 1143978,
      "postDate": "2021-01-08T06:54:09.613Z",
      "content": "<p>Firstly, thanks to Kaggle hold such a good competition, almost no shake. Congrats tops and very thanks to my teammates <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a>  <a href=\"https://www.kaggle.com/ekffar\" target=\"_blank\">@ekffar</a>. Finally our team finally public score 0.808, private 0.810, and final rank 24. we missed gold, it's ok, now i want to share our method</p>\n<p>because of jet lag, i have no read discuss in detail, so i just see my teammate <a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> had shared our <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209625\" target=\"_blank\">solution</a> right now. As a complementary, not repeat, i just expand the detail based on  <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209625\" target=\"_blank\">jaideepvalani's solution</a></p>\n<h1><strong>Architecture</strong></h1>\n<pre><code>SaintPlusPlus, lgb\n</code></pre>\n<h1>Embeddings</h1>\n<h2>Exercise:</h2>\n<pre><code>Question_id\nPart\nCommunity id (tried, but no boosting, maybe need tuning)\nTask container id (tried, but no boosting, maybe need tuning)\n</code></pre>\n<h2>Interaction:</h2>\n<pre><code>lag time  (same container id, keep same with first one) (discrete embedding)\nprior question elapsed time(continuous embedding)\nprior had explain (discrete embedding)\nPrior lecture(discrete embedding)\n</code></pre>\n<h2>Response:</h2>\n<pre><code>Prior question correctness(discrete embedding)\n</code></pre>\n<h1>Data sample optimization</h1>\n<p>Data sampling is important in our experiment<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F347724%2F8884cd3f5fca04b3232adb038b131f0f%2Friiid.png?generation=1610088827710995&amp;alt=media\" alt=\"\"></p>\n<h1>optimization</h1>\n<p>## SAINTPP<br>\n   1 skip connection, we add skip connection for multi encoder layers inner, and multi decoder layers inner, the origin implement i use from github have no this skip connection <br>\n   2 add dropout for encoder, decoder, FFN<br>\n   3 try TCN causal conv after embedding, no boost, but now, i think it should tune and will work<br>\n   4 big seq len for saintpp give our team boosting, i think it maybe caused by too short seq make <br>\n       transformer link model struggle into local minest point </p>\n<h2>LGB</h2>\n<p>updating</p>\n<h1>Conclusion</h1>\n<p>Our best saintpp(seqlen 320, embdim224, 2 x encder+ 2 x decoder) get cv 0.8016 lb 0.806, and lgb near get 0.796. Finally submission is merge lgb(lb 796) + saintpp cv1(cv 8016) + saintpp cv3(cv 0.7992)  = public lb 0.808, and private 0.810.<br>\nshould do things:<br>\n  1 should continue optimize our net with casual conv(because transformer like is a global attention, some local correlation maybe ignored) and gru, but finally have no time<br>\n  2 should try big net continue</p>\n<p>Thanks again to my teammates, they give me a lots of help and let me learn a lot. Going  for gold next Compete, keep moving.</p>",
      "rawMarkdown": "Firstly, thanks to Kaggle hold such a good competition, almost no shake. Congrats tops and very thanks to my teammates @jaideepvalani  @ekffar. Finally our team finally public score 0.808, private 0.810, and final rank 24. we missed gold, it's ok, now i want to share our method\n\nbecause of jet lag, i have no read discuss in detail, so i just see my teammate @jaideepvalani had shared our [solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209625) right now. As a complementary, not repeat, i just expand the detail based on  [jaideepvalani's solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209625)\n\n\n# **Architecture**\n    SaintPlusPlus, lgb\n    \n\n# Embeddings\n \n## Exercise:\n    Question_id\n    Part\n    Community id (tried, but no boosting, maybe need tuning)\n    Task container id (tried, but no boosting, maybe need tuning)\n## Interaction:\n    lag time  (same container id, keep same with first one) (discrete embedding)\n    prior question elapsed time(continuous embedding)\n    prior had explain (discrete embedding)\n    Prior lecture(discrete embedding)\n## Response:\n    Prior question correctness(discrete embedding)\n\n# Data sample optimization\nData sampling is important in our experiment\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F347724%2F8884cd3f5fca04b3232adb038b131f0f%2Friiid.png?generation=1610088827710995&alt=media)\n\n\n# optimization\n ## SAINTPP\n   1 skip connection, we add skip connection for multi encoder layers inner, and multi decoder layers inner, the origin implement i use from github have no this skip connection \n   2 add dropout for encoder, decoder, FFN\n   3 try TCN causal conv after embedding, no boost, but now, i think it should tune and will work\n   4 big seq len for saintpp give our team boosting, i think it maybe caused by too short seq make \n       transformer link model struggle into local minest point \n\n## LGB\nupdating\n\n# Conclusion\n\nOur best saintpp(seqlen 320, embdim224, 2 x encder+ 2 x decoder) get cv 0.8016 lb 0.806, and lgb near get 0.796. Finally submission is merge lgb(lb 796) + saintpp cv1(cv 8016) + saintpp cv3(cv 0.7992)  = public lb 0.808, and private 0.810.\nshould do things:\n  1 should continue optimize our net with casual conv(because transformer like is a global attention, some local correlation maybe ignored) and gru, but finally have no time\n  2 should try big net continue\n\nThanks again to my teammates, they give me a lots of help and let me learn a lot. Going  for gold next Compete, keep moving.\n",
      "votes": 28
    },
    {
      "id": 1144275,
      "postDate": "2021-01-08T11:05:31.287Z",
      "content": "<p>There may be some common mistakes in calculation of below features for saint plus</p>\n<p>1) Lagtime needs to be same for all question of Container. While calculating it using shift  ,we get 0 lagtime for all questions except the first one ,so one needs to replace those zero using simple group by user/task container and transform first.  This gave us boost of atleast 0.0015 to 0.003. Note: Paper mentions prior question lag time but this info seem to be getting too stale for model to intepret some thing for current question,so some how just ts2-ts1 only worked best for us.</p>\n<p>2) Second most important mistake some people could be making in calculating lagtime during inference. Users Task containers are spreaded across multiple iterations unlike train where we have got all interactions at one place ,so using transformation as in step 1 works only partially ,as it misses on lagtime calculation for first question of same user in the other iteration, so one must maintain dict of last timestamp for each user across iterations.<br>\nThis correction gave us boost of 0.005 .</p>\n<p>3) Same goes Prior interaction as lecture, calculating it for train was similar to lagtime but for inference it was even more tricky. One had to take into account last interaction of user in the train if it was Lecture, similarly do same for users last interaction in every test iteration if its a lecture.</p>\n<p>We regret we couldnt do much experiment over using of prior lecture elapsed time,prios lecture part. </p>",
      "rawMarkdown": "There may be some common mistakes in calculation of below features for saint plus\n\n1) Lagtime needs to be same for all question of Container. While calculating it using shift  ,we get 0 lagtime for all questions except the first one ,so one needs to replace those zero using simple group by user/task container and transform first.  This gave us boost of atleast 0.0015 to 0.003. Note: Paper mentions prior question lag time but this info seem to be getting too stale for model to intepret some thing for current question,so some how just ts2-ts1 only worked best for us.\n\n2) Second most important mistake some people could be making in calculating lagtime during inference. Users Task containers are spreaded across multiple iterations unlike train where we have got all interactions at one place ,so using transformation as in step 1 works only partially ,as it misses on lagtime calculation for first question of same user in the other iteration, so one must maintain dict of last timestamp for each user across iterations.\nThis correction gave us boost of 0.005 .\n\n3) Same goes Prior interaction as lecture, calculating it for train was similar to lagtime but for inference it was even more tricky. One had to take into account last interaction of user in the train if it was Lecture, similarly do same for users last interaction in every test iteration if its a lecture.\n\nWe regret we couldnt do much experiment over using of prior lecture elapsed time,prios lecture part. \n\n",
      "votes": 3
    },
    {
      "id": 1146442,
      "postDate": "2021-01-09T19:22:50.407Z",
      "content": "<p>Congrats!!! Really impressive</p>",
      "rawMarkdown": "Congrats!!! Really impressive",
      "votes": 1,
      "replies": [
        {
          "id": 1150959,
          "postDate": "2021-01-13T02:13:20.653Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/tangshuyun\" target=\"_blank\">@tangshuyun</a> sorry for late reply</p>",
          "rawMarkdown": "Thank you @tangshuyun sorry for late reply",
          "votes": 1
        }
      ]
    },
    {
      "id": 1583523,
      "postDate": "2021-11-15T23:14:01.303Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1144275,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2021-01-08T11:05:31.287000",
      "content": "<p>There may be some common mistakes in calculation of below features for saint plus</p>\n<p>1) Lagtime needs to be same for all question of Container. While calculating it using shift  ,we get 0 lagtime for all questions except the first one ,so one needs to replace those zero using simple group by user/task container and transform first.  This gave us boost of atleast 0.0015 to 0.003. Note: Paper mentions prior question lag time but this info seem to be getting too stale for model to intepret some thing for current question,so some how just ts2-ts1 only worked best for us.</p>\n<p>2) Second most important mistake some people could be making in calculating lagtime during inference. Users Task containers are spreaded across multiple iterations unlike train where we have got all interactions at one place ,so using transformation as in step 1 works only partially ,as it misses on lagtime calculation for first question of same user in the other iteration, so one must maintain dict of last timestamp for each user across iterations.<br>\nThis correction gave us boost of 0.005 .</p>\n<p>3) Same goes Prior interaction as lecture, calculating it for train was similar to lagtime but for inference it was even more tricky. One had to take into account last interaction of user in the train if it was Lecture, similarly do same for users last interaction in every test iteration if its a lecture.</p>\n<p>We regret we couldnt do much experiment over using of prior lecture elapsed time,prios lecture part. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1146442,
      "author_name": "Shu",
      "author_url": "",
      "post_date": "2021-01-09T19:22:50.407000",
      "content": "<p>Congrats!!! Really impressive</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1150959,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2021-01-13T02:13:20.653000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/tangshuyun\" target=\"_blank\">@tangshuyun</a> sorry for late reply</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1583523,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-11-15T23:14:01.303000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143978": "Firstly, thanks to Kaggle hold such a good competition, almost no shake. Congrats tops and very thanks to my teammates @jaideepvalani  @ekffar. Finally our team finally public score 0.808, private 0.810, and final rank 24. we missed gold, it's ok, now i want to share our method\n\nbecause of jet lag, i have no read discuss in detail, so i just see my teammate @jaideepvalani had shared our [solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209625) right now. As a complementary, not repeat, i just expand the detail based on  [jaideepvalani's solution](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/209625)\n\n\n# **Architecture**\n    SaintPlusPlus, lgb\n    \n\n# Embeddings\n \n## Exercise:\n    Question_id\n    Part\n    Community id (tried, but no boosting, maybe need tuning)\n    Task container id (tried, but no boosting, maybe need tuning)\n## Interaction:\n    lag time  (same container id, keep same with first one) (discrete embedding)\n    prior question elapsed time(continuous embedding)\n    prior had explain (discrete embedding)\n    Prior lecture(discrete embedding)\n## Response:\n    Prior question correctness(discrete embedding)\n\n# Data sample optimization\nData sampling is important in our experiment\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F347724%2F8884cd3f5fca04b3232adb038b131f0f%2Friiid.png?generation=1610088827710995&alt=media)\n\n\n# optimization\n ## SAINTPP\n   1 skip connection, we add skip connection for multi encoder layers inner, and multi decoder layers inner, the origin implement i use from github have no this skip connection \n   2 add dropout for encoder, decoder, FFN\n   3 try TCN causal conv after embedding, no boost, but now, i think it should tune and will work\n   4 big seq len for saintpp give our team boosting, i think it maybe caused by too short seq make \n       transformer link model struggle into local minest point \n\n## LGB\nupdating\n\n# Conclusion\n\nOur best saintpp(seqlen 320, embdim224, 2 x encder+ 2 x decoder) get cv 0.8016 lb 0.806, and lgb near get 0.796. Finally submission is merge lgb(lb 796) + saintpp cv1(cv 8016) + saintpp cv3(cv 0.7992)  = public lb 0.808, and private 0.810.\nshould do things:\n  1 should continue optimize our net with casual conv(because transformer like is a global attention, some local correlation maybe ignored) and gru, but finally have no time\n  2 should try big net continue\n\nThanks again to my teammates, they give me a lots of help and let me learn a lot. Going  for gold next Compete, keep moving.\n",
    "1144275": "There may be some common mistakes in calculation of below features for saint plus\n\n1) Lagtime needs to be same for all question of Container. While calculating it using shift  ,we get 0 lagtime for all questions except the first one ,so one needs to replace those zero using simple group by user/task container and transform first.  This gave us boost of atleast 0.0015 to 0.003. Note: Paper mentions prior question lag time but this info seem to be getting too stale for model to intepret some thing for current question,so some how just ts2-ts1 only worked best for us.\n\n2) Second most important mistake some people could be making in calculating lagtime during inference. Users Task containers are spreaded across multiple iterations unlike train where we have got all interactions at one place ,so using transformation as in step 1 works only partially ,as it misses on lagtime calculation for first question of same user in the other iteration, so one must maintain dict of last timestamp for each user across iterations.\nThis correction gave us boost of 0.005 .\n\n3) Same goes Prior interaction as lecture, calculating it for train was similar to lagtime but for inference it was even more tricky. One had to take into account last interaction of user in the train if it was Lecture, similarly do same for users last interaction in every test iteration if its a lecture.\n\nWe regret we couldnt do much experiment over using of prior lecture elapsed time,prios lecture part. \n\n",
    "1146442": "Congrats!!! Really impressive",
    "1583523": ""
  }
}