{
  "id": 196987,
  "title": "Some findings of data preparation and modeling",
  "url": "/competitions/riiid-test-answer-prediction/discussion/196987",
  "author_name": "",
  "post_date": "2020-11-13T17:47:22.407967700Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>For data splitting:<br>\nThis dataset seems a combination of time series and stationary data. So we have to train on complete histories of users, same as validation. (Random shuffle and split won't work! I have tested, it generate prediction score of ~0.5)</p>\n<p>For modeling:<br>\nFor this dataset, it is based on well engineered features which is not like high dimension vectors. So regular classification methods should and which is proved to perform better than <a href=\"https://www.kaggle.com/chenmingml/dnn-based-on-lgbm-iii-data-preparation\" target=\"_blank\">deep learning models</a>: same dataset but different modeling method. Output is 0.666 vs 0.756 of LGBM.<br>\nRef: <a href=\"https://www.manning.com/books/deep-learning-with-python\" target=\"_blank\">Deep Learning with Python</a>, Francois-chollet, </p>",
  "messages": [
    {
      "id": "1077561",
      "postDate": "11/13/2020 17:47:22",
      "content": "<p>For data splitting:<br>\nThis dataset seems a combination of time series and stationary data. So we have to train on complete histories of users, same as validation. (Random shuffle and split won't work! I have tested, it generate prediction score of ~0.5)</p>\n<p>For modeling:<br>\nFor this dataset, it is based on well engineered features which is not like high dimension vectors. So regular classification methods should and which is proved to perform better than <a href=\"https://www.kaggle.com/chenmingml/dnn-based-on-lgbm-iii-data-preparation\" target=\"_blank\">deep learning models</a>: same dataset but different modeling method. Output is 0.666 vs 0.756 of LGBM.<br>\nRef: <a href=\"https://www.manning.com/books/deep-learning-with-python\" target=\"_blank\">Deep Learning with Python</a>, Francois-chollet, </p>",
      "rawMarkdown": "For data splitting:\nThis dataset seems a combination of time series and stationary data. So we have to train on complete histories of users, same as validation. (Random shuffle and split won't work! I have tested, it generate prediction score of ~0.5)\n\nFor modeling:\nFor this dataset, it is based on well engineered features which is not like high dimension vectors. So regular classification methods should and which is proved to perform better than [deep learning models](https://www.kaggle.com/chenmingml/dnn-based-on-lgbm-iii-data-preparation): same dataset but different modeling method. Output is 0.666 vs 0.756 of LGBM.\nRef: [Deep Learning with Python](https://www.manning.com/books/deep-learning-with-python), Francois-chollet,",
      "votes": null
    },
    {
      "id": "1079159",
      "postDate": "11/15/2020 17:29:01",
      "content": "<p>Im also working on some deep learning models but during validation the score is just way lower and takes longer than all boosted classification methods, i have tested e.g. RandomForestClassifier etc.</p>",
      "rawMarkdown": "Im also working on some deep learning models but during validation the score is just way lower and takes longer than all boosted classification methods, i have tested e.g. RandomForestClassifier etc.",
      "votes": null
    },
    {
      "id": "1079387",
      "postDate": "11/16/2020 02:19:21",
      "content": "<p>It's true the DL model takes longer time than RF classification model. But with GPU it is OK. <br>\nUse Dense layers is not slow. I have tried LSTM which may not be useful but it gets stuck. :(<br>\nThe problem is in submission and scoring but it seems ppl have found better ways replacing pd.merge.</p>",
      "rawMarkdown": "It's true the DL model takes longer time than RF classification model. But with GPU it is OK. \nUse Dense layers is not slow. I have tried LSTM which may not be useful but it gets stuck. :(\nThe problem is in submission and scoring but it seems ppl have found better ways replacing pd.merge.",
      "votes": null
    },
    {
      "id": "1079766",
      "postDate": "11/16/2020 12:59:39",
      "content": "<p>I just tried a very simple NN with a couple of linear layers, nothing special and I got 0.5, so a randomized evaluation is as good as my model. I will try a more complex model and respond again.</p>\n<p>EDIT: You can see how bad/good the model peforms, if you just feed one batch to it and try to overfit it, my model is not even able to overfit on one batch so it is to narrow and has to be more complex.</p>",
      "rawMarkdown": "I just tried a very simple NN with a couple of linear layers, nothing special and I got 0.5, so a randomized evaluation is as good as my model. I will try a more complex model and respond again.\n\nEDIT: You can see how bad/good the model peforms, if you just feed one batch to it and try to overfit it, my model is not even able to overfit on one batch so it is to narrow and has to be more complex.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1079159,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "11/15/2020 17:29:01",
      "content": "<p>Im also working on some deep learning models but during validation the score is just way lower and takes longer than all boosted classification methods, i have tested e.g. RandomForestClassifier etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1079387,
          "author_name": "chenmingml",
          "author_url": "",
          "post_date": "11/16/2020 02:19:21",
          "content": "<p>It's true the DL model takes longer time than RF classification model. But with GPU it is OK. <br>\nUse Dense layers is not slow. I have tried LSTM which may not be useful but it gets stuck. :(<br>\nThe problem is in submission and scoring but it seems ppl have found better ways replacing pd.merge.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1079766,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "11/16/2020 12:59:39",
          "content": "<p>I just tried a very simple NN with a couple of linear layers, nothing special and I got 0.5, so a randomized evaluation is as good as my model. I will try a more complex model and respond again.</p>\n<p>EDIT: You can see how bad/good the model peforms, if you just feed one batch to it and try to overfit it, my model is not even able to overfit on one batch so it is to narrow and has to be more complex.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1077561": "For data splitting:\nThis dataset seems a combination of time series and stationary data. So we have to train on complete histories of users, same as validation. (Random shuffle and split won't work! I have tested, it generate prediction score of ~0.5)\n\nFor modeling:\nFor this dataset, it is based on well engineered features which is not like high dimension vectors. So regular classification methods should and which is proved to perform better than [deep learning models](https://www.kaggle.com/chenmingml/dnn-based-on-lgbm-iii-data-preparation): same dataset but different modeling method. Output is 0.666 vs 0.756 of LGBM.\nRef: [Deep Learning with Python](https://www.manning.com/books/deep-learning-with-python), Francois-chollet,",
    "1079159": "Im also working on some deep learning models but during validation the score is just way lower and takes longer than all boosted classification methods, i have tested e.g. RandomForestClassifier etc.",
    "1079387": "It's true the DL model takes longer time than RF classification model. But with GPU it is OK. \nUse Dense layers is not slow. I have tried LSTM which may not be useful but it gets stuck. :(\nThe problem is in submission and scoring but it seems ppl have found better ways replacing pd.merge.",
    "1079766": "I just tried a very simple NN with a couple of linear layers, nothing special and I got 0.5, so a randomized evaluation is as good as my model. I will try a more complex model and respond again.\n\nEDIT: You can see how bad/good the model peforms, if you just feed one batch to it and try to overfit it, my model is not even able to overfit on one batch so it is to narrow and has to be more complex."
  },
  "source": "meta"
}