{
  "id": 193282,
  "title": "Leaderboard Tracking",
  "url": "/competitions/riiid-test-answer-prediction/discussion/193282",
  "author_name": "",
  "post_date": "2020-10-26T10:06:25.908013400Z",
  "votes": 12,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I have released a notebook to track the leaderboard, which you can find <a href=\"https://www.kaggle.com/ghostskipper/tracking-the-leaderboard\" target=\"_blank\">here</a>.<br>\nTo get the history I am also publishing a dataset of the leaderboard, which you can find <a href=\"https://www.kaggle.com/ghostskipper/riiid-leaderboad\" target=\"_blank\">here</a>.</p>\n<p>This will be automatically updated everyday at 9am UK time.</p>\n<p>Let me know if you have any suggestions.</p>",
  "messages": [
    {
      "id": "1060553",
      "postDate": "10/26/2020 10:06:25",
      "content": "<p>Hello,</p>\n<p>I have released a notebook to track the leaderboard, which you can find <a href=\"https://www.kaggle.com/ghostskipper/tracking-the-leaderboard\" target=\"_blank\">here</a>.<br>\nTo get the history I am also publishing a dataset of the leaderboard, which you can find <a href=\"https://www.kaggle.com/ghostskipper/riiid-leaderboad\" target=\"_blank\">here</a>.</p>\n<p>This will be automatically updated everyday at 9am UK time.</p>\n<p>Let me know if you have any suggestions.</p>",
      "rawMarkdown": "Hello,\n\nI have released a notebook to track the leaderboard, which you can find [here](https://www.kaggle.com/ghostskipper/tracking-the-leaderboard).\nTo get the history I am also publishing a dataset of the leaderboard, which you can find [here](https://www.kaggle.com/ghostskipper/riiid-leaderboad).\n\nThis will be automatically updated everyday at 9am UK time.\n\nLet me know if you have any suggestions.",
      "votes": null
    },
    {
      "id": "1061524",
      "postDate": "10/27/2020 04:06:23",
      "content": "<p>Thanks G.S. - interesting stuff 😄</p>\n<p>In \"Scores Over Time\", how about stopping the line at the last date shown on the Leaderboard? Then we will get a sense of how many teams are actively improving.</p>",
      "rawMarkdown": "Thanks G.S. - interesting stuff 😄\n\nIn \"Scores Over Time\", how about stopping the line at the last date shown on the Leaderboard? Then we will get a sense of how many teams are actively improving.",
      "votes": null
    },
    {
      "id": "1061780",
      "postDate": "10/27/2020 10:17:06",
      "content": "<p>I've made some improvement. I don't entirely understand what you are asking with the last date shown in the leaderboard, but I have added plots for the last 7 days. Hopefully this shows how many teams are actively improving.</p>",
      "rawMarkdown": "I've made some improvement. I don't entirely understand what you are asking with the last date shown in the leaderboard, but I have added plots for the last 7 days. Hopefully this shows how many teams are actively improving.",
      "votes": null
    },
    {
      "id": "1065435",
      "postDate": "10/31/2020 09:30:01",
      "content": "<p>A bit of a shake up in the last few days with <a href=\"https://www.kaggle.com/lynnhhw\" target=\"_blank\">@lynnhhw</a> jumping to first place and the early leader Question the Answers dropping to 5th.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1263569%2Fc8fb6c7db3e3f21012b7d129233787ba%2FScreenshot%20from%202020-10-31%2009-11-53.png?generation=1604135566418977&amp;alt=media\" alt=\"\"></p>\n<p>I'm curious what models are being used at the top, feel free to share 😋.<br>\n<a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a>  has commented they have used a sequencial model to get to 17th place. <br>\nHas anyone has implemented the SAINT/SAINT+ models yet? </p>",
      "rawMarkdown": "A bit of a shake up in the last few days with @lynnhhw jumping to first place and the early leader Question the Answers dropping to 5th.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1263569%2Fc8fb6c7db3e3f21012b7d129233787ba%2FScreenshot%20from%202020-10-31%2009-11-53.png?generation=1604135566418977&alt=media)\n\nI'm curious what models are being used at the top, feel free to share 😋.\n@abhimanyud  has commented they have used a sequencial model to get to 17th place. \nHas anyone has implemented the SAINT/SAINT+ models yet?",
      "votes": null
    },
    {
      "id": "1065442",
      "postDate": "10/31/2020 09:36:46",
      "content": "<p>Guess there's some hidden magic haha or other's who are in lower LB positions are doing something wrong or have overlooked something…</p>",
      "rawMarkdown": "Guess there's some hidden magic haha or other's who are in lower LB positions are doing something wrong or have overlooked something...",
      "votes": null
    },
    {
      "id": "1065453",
      "postDate": "10/31/2020 09:56:29",
      "content": "<p>😁 still searching for this hidden magic! I'm quite enjoying this competition.</p>",
      "rawMarkdown": "😁 still searching for this hidden magic! I'm quite enjoying this competition.",
      "votes": null
    },
    {
      "id": "1065477",
      "postDate": "10/31/2020 10:28:56",
      "content": "<p>Using a sequential model. Not implemented SAINT or its improved variant yet. It seems it cites heavily from the  \"Attention is all you need\" paper by Vaswani et al. which introduced transformers. In the pipeline :)</p>",
      "rawMarkdown": "Using a sequential model. Not implemented SAINT or its improved variant yet. It seems it cites heavily from the  \"Attention is all you need\" paper by Vaswani et al. which introduced transformers. In the pipeline :)",
      "votes": null
    },
    {
      "id": "1065488",
      "postDate": "10/31/2020 10:40:03",
      "content": "<p>Sounds like I have some reading to do!</p>",
      "rawMarkdown": "Sounds like I have some reading to do!",
      "votes": null
    },
    {
      "id": "1065648",
      "postDate": "10/31/2020 14:46:41",
      "content": "<p>There's a notable gap between me and the top, but FWIW my take -- there's really no magic, just good features that aren't particularly surprising and work well so long as you get the test pipeline right. I just use lightgbm to model for now, which I'd guess is true for most people at the top too though <a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> has shown that neural nets are competitive. </p>\n<p>It's not a question of doing something wrong, it's more a question of just doing more to tap into the feature potential in the data. Think about the factors that characterize how a student would perform on a particular question, and also about the factors that characterize how a student interacts with the app (in particular, the factors that aren't directly captured by static features for a single row but require information derived across many rows, since this is what your model truly can't figure out on its own). Figure out how to quantify those factors on both train and test, and your model will improve.   </p>",
      "rawMarkdown": "There's a notable gap between me and the top, but FWIW my take -- there's really no magic, just good features that aren't particularly surprising and work well so long as you get the test pipeline right. I just use lightgbm to model for now, which I'd guess is true for most people at the top too though @abhimanyud has shown that neural nets are competitive. \n\nIt's not a question of doing something wrong, it's more a question of just doing more to tap into the feature potential in the data. Think about the factors that characterize how a student would perform on a particular question, and also about the factors that characterize how a student interacts with the app (in particular, the factors that aren't directly captured by static features for a single row but require information derived across many rows, since this is what your model truly can't figure out on its own). Figure out how to quantify those factors on both train and test, and your model will improve.",
      "votes": null
    },
    {
      "id": "1065651",
      "postDate": "10/31/2020 14:55:08",
      "content": "<p>Relatedly, I think a good general rule of thumb in tabular competitions is to always assume that there are more features you can engineer, and that focusing on that is almost always a better use of your time than tweaking model parameters. Even if you're going with a neural net, data/feature preparation is still critical and it's not just about model architecture (I would even argue that getting a strong gradient boosting model often makes it easier to build a neural net, since you get a good sense of what feature information really matters and then figure out how to capture that in your NN).</p>",
      "rawMarkdown": "Relatedly, I think a good general rule of thumb in tabular competitions is to always assume that there are more features you can engineer, and that focusing on that is almost always a better use of your time than tweaking model parameters. Even if you're going with a neural net, data/feature preparation is still critical and it's not just about model architecture (I would even argue that getting a strong gradient boosting model often makes it easier to build a neural net, since you get a good sense of what feature information really matters and then figure out how to capture that in your NN).",
      "votes": null
    },
    {
      "id": "1065681",
      "postDate": "10/31/2020 15:31:38",
      "content": "<p>Thanks Joe. I completely agree that your features are going to play a huge role as always and the way you can incorporate these into test set as well. And it's quite independent of which model's you choose to do so as well. Plus as you rightly highlighted, it's quite important to have an end to end set-up to test different things </p>\n<p>Thanks for the inspiration again! I am going to take a fresh look on my pipeline today.</p>\n<p>How do you validate that your features kinda behave the way they should?</p>",
      "rawMarkdown": "Thanks Joe. I completely agree that your features are going to play a huge role as always and the way you can incorporate these into test set as well. And it's quite independent of which model's you choose to do so as well. Plus as you rightly highlighted, it's quite important to have an end to end set-up to test different things \n\nThanks for the inspiration again! I am going to take a fresh look on my pipeline today.\n\nHow do you validate that your features kinda behave the way they should?",
      "votes": null
    },
    {
      "id": "1066405",
      "postDate": "11/01/2020 18:16:21",
      "content": "<p>Thanks for the advice 👍</p>",
      "rawMarkdown": "Thanks for the advice 👍",
      "votes": null
    },
    {
      "id": "1071780",
      "postDate": "11/07/2020 11:59:02",
      "content": "<p>I have discovered a bug in the data processing that didn't account for teams changing their names. Sorry about that, I've now fixed this. <br>\nThe top teams leaderboard for the last 7 days now looks like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1263569%2Faaa4b400a0dca11a34aabc6105d9b8f8%2FScreenshot%20from%202020-11-07%2011-50-18.png?generation=1604749835789112&amp;alt=media\" alt=\"\"><br>\nLooks like an exciting battle between <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> and <a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> for the top spot this week. It's all to play for with a large number of teams within striking distance of the top spot and a long way to go in the competition. </p>",
      "rawMarkdown": "I have discovered a bug in the data processing that didn't account for teams changing their names. Sorry about that, I've now fixed this. \nThe top teams leaderboard for the last 7 days now looks like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1263569%2Faaa4b400a0dca11a34aabc6105d9b8f8%2FScreenshot%20from%202020-11-07%2011-50-18.png?generation=1604749835789112&alt=media)\nLooks like an exciting battle between @mamasinkgs and @abhimanyud for the top spot this week. It's all to play for with a large number of teams within striking distance of the top spot and a long way to go in the competition.",
      "votes": null
    },
    {
      "id": "1072905",
      "postDate": "11/08/2020 21:00:25",
      "content": "<p>I have fixed the notebook so it excludes submissions &gt; 0.999</p>",
      "rawMarkdown": "I have fixed the notebook so it excludes submissions > 0.999",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1061524,
      "author_name": "mikel1",
      "author_url": "",
      "post_date": "10/27/2020 04:06:23",
      "content": "<p>Thanks G.S. - interesting stuff 😄</p>\n<p>In \"Scores Over Time\", how about stopping the line at the last date shown on the Leaderboard? Then we will get a sense of how many teams are actively improving.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1061780,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "10/27/2020 10:17:06",
          "content": "<p>I've made some improvement. I don't entirely understand what you are asking with the last date shown in the leaderboard, but I have added plots for the last 7 days. Hopefully this shows how many teams are actively improving.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1065435,
      "author_name": "ghostskipper",
      "author_url": "",
      "post_date": "10/31/2020 09:30:01",
      "content": "<p>A bit of a shake up in the last few days with <a href=\"https://www.kaggle.com/lynnhhw\" target=\"_blank\">@lynnhhw</a> jumping to first place and the early leader Question the Answers dropping to 5th.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1263569%2Fc8fb6c7db3e3f21012b7d129233787ba%2FScreenshot%20from%202020-10-31%2009-11-53.png?generation=1604135566418977&amp;alt=media\" alt=\"\"></p>\n<p>I'm curious what models are being used at the top, feel free to share 😋.<br>\n<a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a>  has commented they have used a sequencial model to get to 17th place. <br>\nHas anyone has implemented the SAINT/SAINT+ models yet? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1065442,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/31/2020 09:36:46",
          "content": "<p>Guess there's some hidden magic haha or other's who are in lower LB positions are doing something wrong or have overlooked something…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065453,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "10/31/2020 09:56:29",
          "content": "<p>😁 still searching for this hidden magic! I'm quite enjoying this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065477,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "10/31/2020 10:28:56",
          "content": "<p>Using a sequential model. Not implemented SAINT or its improved variant yet. It seems it cites heavily from the  \"Attention is all you need\" paper by Vaswani et al. which introduced transformers. In the pipeline :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065488,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "10/31/2020 10:40:03",
          "content": "<p>Sounds like I have some reading to do!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065648,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/31/2020 14:46:41",
          "content": "<p>There's a notable gap between me and the top, but FWIW my take -- there's really no magic, just good features that aren't particularly surprising and work well so long as you get the test pipeline right. I just use lightgbm to model for now, which I'd guess is true for most people at the top too though <a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> has shown that neural nets are competitive. </p>\n<p>It's not a question of doing something wrong, it's more a question of just doing more to tap into the feature potential in the data. Think about the factors that characterize how a student would perform on a particular question, and also about the factors that characterize how a student interacts with the app (in particular, the factors that aren't directly captured by static features for a single row but require information derived across many rows, since this is what your model truly can't figure out on its own). Figure out how to quantify those factors on both train and test, and your model will improve.   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065651,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "10/31/2020 14:55:08",
          "content": "<p>Relatedly, I think a good general rule of thumb in tabular competitions is to always assume that there are more features you can engineer, and that focusing on that is almost always a better use of your time than tweaking model parameters. Even if you're going with a neural net, data/feature preparation is still critical and it's not just about model architecture (I would even argue that getting a strong gradient boosting model often makes it easier to build a neural net, since you get a good sense of what feature information really matters and then figure out how to capture that in your NN).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065681,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/31/2020 15:31:38",
          "content": "<p>Thanks Joe. I completely agree that your features are going to play a huge role as always and the way you can incorporate these into test set as well. And it's quite independent of which model's you choose to do so as well. Plus as you rightly highlighted, it's quite important to have an end to end set-up to test different things </p>\n<p>Thanks for the inspiration again! I am going to take a fresh look on my pipeline today.</p>\n<p>How do you validate that your features kinda behave the way they should?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1066405,
          "author_name": "ghostskipper",
          "author_url": "",
          "post_date": "11/01/2020 18:16:21",
          "content": "<p>Thanks for the advice 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1071780,
      "author_name": "ghostskipper",
      "author_url": "",
      "post_date": "11/07/2020 11:59:02",
      "content": "<p>I have discovered a bug in the data processing that didn't account for teams changing their names. Sorry about that, I've now fixed this. <br>\nThe top teams leaderboard for the last 7 days now looks like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1263569%2Faaa4b400a0dca11a34aabc6105d9b8f8%2FScreenshot%20from%202020-11-07%2011-50-18.png?generation=1604749835789112&amp;alt=media\" alt=\"\"><br>\nLooks like an exciting battle between <a href=\"https://www.kaggle.com/mamasinkgs\" target=\"_blank\">@mamasinkgs</a> and <a href=\"https://www.kaggle.com/abhimanyud\" target=\"_blank\">@abhimanyud</a> for the top spot this week. It's all to play for with a large number of teams within striking distance of the top spot and a long way to go in the competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1072905,
      "author_name": "ghostskipper",
      "author_url": "",
      "post_date": "11/08/2020 21:00:25",
      "content": "<p>I have fixed the notebook so it excludes submissions &gt; 0.999</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1060553": "Hello,\n\nI have released a notebook to track the leaderboard, which you can find [here](https://www.kaggle.com/ghostskipper/tracking-the-leaderboard).\nTo get the history I am also publishing a dataset of the leaderboard, which you can find [here](https://www.kaggle.com/ghostskipper/riiid-leaderboad).\n\nThis will be automatically updated everyday at 9am UK time.\n\nLet me know if you have any suggestions.",
    "1061524": "Thanks G.S. - interesting stuff 😄\n\nIn \"Scores Over Time\", how about stopping the line at the last date shown on the Leaderboard? Then we will get a sense of how many teams are actively improving.",
    "1061780": "I've made some improvement. I don't entirely understand what you are asking with the last date shown in the leaderboard, but I have added plots for the last 7 days. Hopefully this shows how many teams are actively improving.",
    "1065435": "A bit of a shake up in the last few days with @lynnhhw jumping to first place and the early leader Question the Answers dropping to 5th.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1263569%2Fc8fb6c7db3e3f21012b7d129233787ba%2FScreenshot%20from%202020-10-31%2009-11-53.png?generation=1604135566418977&alt=media)\n\nI'm curious what models are being used at the top, feel free to share 😋.\n@abhimanyud  has commented they have used a sequencial model to get to 17th place. \nHas anyone has implemented the SAINT/SAINT+ models yet?",
    "1065442": "Guess there's some hidden magic haha or other's who are in lower LB positions are doing something wrong or have overlooked something...",
    "1065453": "😁 still searching for this hidden magic! I'm quite enjoying this competition.",
    "1065477": "Using a sequential model. Not implemented SAINT or its improved variant yet. It seems it cites heavily from the  \"Attention is all you need\" paper by Vaswani et al. which introduced transformers. In the pipeline :)",
    "1065488": "Sounds like I have some reading to do!",
    "1065648": "There's a notable gap between me and the top, but FWIW my take -- there's really no magic, just good features that aren't particularly surprising and work well so long as you get the test pipeline right. I just use lightgbm to model for now, which I'd guess is true for most people at the top too though @abhimanyud has shown that neural nets are competitive. \n\nIt's not a question of doing something wrong, it's more a question of just doing more to tap into the feature potential in the data. Think about the factors that characterize how a student would perform on a particular question, and also about the factors that characterize how a student interacts with the app (in particular, the factors that aren't directly captured by static features for a single row but require information derived across many rows, since this is what your model truly can't figure out on its own). Figure out how to quantify those factors on both train and test, and your model will improve.",
    "1065651": "Relatedly, I think a good general rule of thumb in tabular competitions is to always assume that there are more features you can engineer, and that focusing on that is almost always a better use of your time than tweaking model parameters. Even if you're going with a neural net, data/feature preparation is still critical and it's not just about model architecture (I would even argue that getting a strong gradient boosting model often makes it easier to build a neural net, since you get a good sense of what feature information really matters and then figure out how to capture that in your NN).",
    "1065681": "Thanks Joe. I completely agree that your features are going to play a huge role as always and the way you can incorporate these into test set as well. And it's quite independent of which model's you choose to do so as well. Plus as you rightly highlighted, it's quite important to have an end to end set-up to test different things \n\nThanks for the inspiration again! I am going to take a fresh look on my pipeline today.\n\nHow do you validate that your features kinda behave the way they should?",
    "1066405": "Thanks for the advice 👍",
    "1071780": "I have discovered a bug in the data processing that didn't account for teams changing their names. Sorry about that, I've now fixed this. \nThe top teams leaderboard for the last 7 days now looks like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1263569%2Faaa4b400a0dca11a34aabc6105d9b8f8%2FScreenshot%20from%202020-11-07%2011-50-18.png?generation=1604749835789112&alt=media)\nLooks like an exciting battle between @mamasinkgs and @abhimanyud for the top spot this week. It's all to play for with a large number of teams within striking distance of the top spot and a long way to go in the competition.",
    "1072905": "I have fixed the notebook so it excludes submissions > 0.999"
  },
  "source": "meta"
}