{
  "id": 195703,
  "title": "Are we supposed to train RNN's?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/195703",
  "author_name": "",
  "post_date": "2020-11-06T22:46:02.266265800Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, the competition title says \"Track knowledge states of 1M+ students in the wild\", which gives me the impression that we are supposed to create RNNs, since they do so well with time related data.  Many of the notebooks, though, use LGBM and other simple tabular data approaches.  Looking at the test data, I see that many of the values are not actually the first questions, which causes me to wonder if I am supposed to train only using simple tabular data approaches.  Am I correct?  Does the data simply have holes in it and that is part of the competition or should I go back to tabular regression approaches?  Have people created RNN approaches that I simply have no idea about?</p>",
  "messages": [
    {
      "id": "1071452",
      "postDate": "11/06/2020 22:46:02",
      "content": "<p>Hi, the competition title says \"Track knowledge states of 1M+ students in the wild\", which gives me the impression that we are supposed to create RNNs, since they do so well with time related data.  Many of the notebooks, though, use LGBM and other simple tabular data approaches.  Looking at the test data, I see that many of the values are not actually the first questions, which causes me to wonder if I am supposed to train only using simple tabular data approaches.  Am I correct?  Does the data simply have holes in it and that is part of the competition or should I go back to tabular regression approaches?  Have people created RNN approaches that I simply have no idea about?</p>",
      "rawMarkdown": "Hi, the competition title says \"Track knowledge states of 1M+ students in the wild\", which gives me the impression that we are supposed to create RNNs, since they do so well with time related data.  Many of the notebooks, though, use LGBM and other simple tabular data approaches.  Looking at the test data, I see that many of the values are not actually the first questions, which causes me to wonder if I am supposed to train only using simple tabular data approaches.  Am I correct?  Does the data simply have holes in it and that is part of the competition or should I go back to tabular regression approaches?  Have people created RNN approaches that I simply have no idea about?",
      "votes": null
    },
    {
      "id": "1071456",
      "postDate": "11/06/2020 22:50:55",
      "content": "<p>You can use whatever kind of model you want.</p>",
      "rawMarkdown": "You can use whatever kind of model you want.",
      "votes": null
    },
    {
      "id": "1071918",
      "postDate": "11/07/2020 15:03:23",
      "content": "<p>Yes, and I understand this, but is the previous data from the same students supposed to be used to try to predict the future data, or are the same students there to make data gathering easier?</p>",
      "rawMarkdown": "Yes, and I understand this, but is the previous data from the same students supposed to be used to try to predict the future data, or are the same students there to make data gathering easier?",
      "votes": null
    },
    {
      "id": "1071938",
      "postDate": "11/07/2020 15:43:40",
      "content": "<p>You can try any model you want. It does not require that \"You must train XYZ\": as long as it works we simply do not care which type is it.</p>",
      "rawMarkdown": "You can try any model you want. It does not require that \"You must train XYZ\": as long as it works we simply do not care which type is it.",
      "votes": null
    },
    {
      "id": "1071954",
      "postDate": "11/07/2020 16:04:34",
      "content": "<p>I thought the same when i read the title.</p>",
      "rawMarkdown": "I thought the same when i read the title.",
      "votes": null
    },
    {
      "id": "1071956",
      "postDate": "11/07/2020 16:07:27",
      "content": "<p>A typical approach in data science is not to try out only one model or one technique, the best thing is, when you try out multiple approaches on the same set of data. Afterwards you can compare all approaches and pick the best working onces for fine tuning/hyperparameter tuning. If you are done with that you would pick the best working method and try it on the whole data.</p>\n<p>So if you believe, that RNN is good then you should try it and compare it to your previous methods.</p>",
      "rawMarkdown": "A typical approach in data science is not to try out only one model or one technique, the best thing is, when you try out multiple approaches on the same set of data. Afterwards you can compare all approaches and pick the best working onces for fine tuning/hyperparameter tuning. If you are done with that you would pick the best working method and try it on the whole data.\n\nSo if you believe, that RNN is good then you should try it and compare it to your previous methods.",
      "votes": null
    },
    {
      "id": "1072447",
      "postDate": "11/08/2020 08:29:17",
      "content": "<p>The good thing about DL and ML is that you can use whatever model kind you want to use depending on the data and approaches you want to use and the domain knowledge.</p>",
      "rawMarkdown": "The good thing about DL and ML is that you can use whatever model kind you want to use depending on the data and approaches you want to use and the domain knowledge.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1071456,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "11/06/2020 22:50:55",
      "content": "<p>You can use whatever kind of model you want.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1071918,
          "author_name": "fatehaliyev",
          "author_url": "",
          "post_date": "11/07/2020 15:03:23",
          "content": "<p>Yes, and I understand this, but is the previous data from the same students supposed to be used to try to predict the future data, or are the same students there to make data gathering easier?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1071938,
      "author_name": "aeryss",
      "author_url": "",
      "post_date": "11/07/2020 15:43:40",
      "content": "<p>You can try any model you want. It does not require that \"You must train XYZ\": as long as it works we simply do not care which type is it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1071954,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "11/07/2020 16:04:34",
          "content": "<p>I thought the same when i read the title.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1071956,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "11/07/2020 16:07:27",
      "content": "<p>A typical approach in data science is not to try out only one model or one technique, the best thing is, when you try out multiple approaches on the same set of data. Afterwards you can compare all approaches and pick the best working onces for fine tuning/hyperparameter tuning. If you are done with that you would pick the best working method and try it on the whole data.</p>\n<p>So if you believe, that RNN is good then you should try it and compare it to your previous methods.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1072447,
      "author_name": "ahmedwagdy95",
      "author_url": "",
      "post_date": "11/08/2020 08:29:17",
      "content": "<p>The good thing about DL and ML is that you can use whatever model kind you want to use depending on the data and approaches you want to use and the domain knowledge.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1071452": "Hi, the competition title says \"Track knowledge states of 1M+ students in the wild\", which gives me the impression that we are supposed to create RNNs, since they do so well with time related data.  Many of the notebooks, though, use LGBM and other simple tabular data approaches.  Looking at the test data, I see that many of the values are not actually the first questions, which causes me to wonder if I am supposed to train only using simple tabular data approaches.  Am I correct?  Does the data simply have holes in it and that is part of the competition or should I go back to tabular regression approaches?  Have people created RNN approaches that I simply have no idea about?",
    "1071456": "You can use whatever kind of model you want.",
    "1071918": "Yes, and I understand this, but is the previous data from the same students supposed to be used to try to predict the future data, or are the same students there to make data gathering easier?",
    "1071938": "You can try any model you want. It does not require that \"You must train XYZ\": as long as it works we simply do not care which type is it.",
    "1071954": "I thought the same when i read the title.",
    "1071956": "A typical approach in data science is not to try out only one model or one technique, the best thing is, when you try out multiple approaches on the same set of data. Afterwards you can compare all approaches and pick the best working onces for fine tuning/hyperparameter tuning. If you are done with that you would pick the best working method and try it on the whole data.\n\nSo if you believe, that RNN is good then you should try it and compare it to your previous methods.",
    "1072447": "The good thing about DL and ML is that you can use whatever model kind you want to use depending on the data and approaches you want to use and the domain knowledge."
  },
  "source": "meta"
}