{
  "id": 192625,
  "title": "Why are neural networks are underperforming for this task?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/192625",
  "author_name": "Abdessalem Boukil",
  "post_date": "2020-10-22T12:19:42.971000",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I am trying to create a baseline model to start with, and since I am more acquainted with deep learning rather than classical machine learning, I am attempting to do it with deep neural networks. However the neural nets are barely learning, played with the loss/optimizer/lr … nothing, the results seem to be more arbitrary rather than the model learning.<br>\nSeeing the notebooks here though, it seems that boosting methods are having more success with it, is it the architecture of my NN that prevents it from learning, or is that boosting is better equipped for these types of problems. Thanks!<br>\nAnd also sorry for my newbie question.</p>",
  "messages": [
    {
      "id": 1057150,
      "postDate": "2020-10-22T12:19:42.970Z",
      "content": "<p>I am trying to create a baseline model to start with, and since I am more acquainted with deep learning rather than classical machine learning, I am attempting to do it with deep neural networks. However the neural nets are barely learning, played with the loss/optimizer/lr … nothing, the results seem to be more arbitrary rather than the model learning.<br>\nSeeing the notebooks here though, it seems that boosting methods are having more success with it, is it the architecture of my NN that prevents it from learning, or is that boosting is better equipped for these types of problems. Thanks!<br>\nAnd also sorry for my newbie question.</p>",
      "rawMarkdown": "I am trying to create a baseline model to start with, and since I am more acquainted with deep learning rather than classical machine learning, I am attempting to do it with deep neural networks. However the neural nets are barely learning, played with the loss/optimizer/lr ... nothing, the results seem to be more arbitrary rather than the model learning.\nSeeing the notebooks here though, it seems that boosting methods are having more success with it, is it the architecture of my NN that prevents it from learning, or is that boosting is better equipped for these types of problems. Thanks!\nAnd also sorry for my newbie question.",
      "votes": 8
    },
    {
      "id": 1057584,
      "postDate": "2020-10-22T19:21:02.633Z",
      "content": "<p>The best model that I met in <a href=\"https://arxiv.org/abs/2002.07033\" target=\"_blank\">one of the articles</a> was just a neural network (transformer). And it gave a result (on the dataset from Riiid) slightly higher than the current leader. From this I conclude that the best result will be approximately 0.80-0.83</p>",
      "rawMarkdown": "The best model that I met in [one of the articles](https://arxiv.org/abs/2002.07033) was just a neural network (transformer). And it gave a result (on the dataset from Riiid) slightly higher than the current leader. From this I conclude that the best result will be approximately 0.80-0.83",
      "votes": 4,
      "replies": [
        {
          "id": 1057604,
          "postDate": "2020-10-22T19:43:43.137Z",
          "content": "<p>One thing to consider is that if such model can finish the prediction in 2 hours with Kaggle GPU or in 9 hours of CPU. I think this will be changeling. Of course, we can try to reduce the model size, but the question becomes if we can achieve the same result in this paper with smaller models that are fast to predict.</p>",
          "rawMarkdown": "One thing to consider is that if such model can finish the prediction in 2 hours with Kaggle GPU or in 9 hours of CPU. I think this will be changeling. Of course, we can try to reduce the model size, but the question becomes if we can achieve the same result in this paper with smaller models that are fast to predict.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1057193,
      "postDate": "2020-10-22T13:09:16.060Z",
      "content": "<p>What type of deep learning architecture did you use to confirm this??</p>",
      "rawMarkdown": "What type of deep learning architecture did you use to confirm this??",
      "votes": 1,
      "replies": [
        {
          "id": 1057247,
          "postDate": "2020-10-22T14:06:08.317Z",
          "content": "<p>I used a regular architecture. Features -&gt; (64 NN) -&gt; (Relu + Dropout + LayerNorm) -&gt; (64 NN) -&gt; (Relu + Dropout + LayerNorm) -&gt; Map it to two inputs. Trained on ADAM, and crossentropyloss.</p>",
          "rawMarkdown": "I used a regular architecture. Features -> (64 NN) -> (Relu + Dropout + LayerNorm) -> (64 NN) -> (Relu + Dropout + LayerNorm) -> Map it to two inputs. Trained on ADAM, and crossentropyloss."
        },
        {
          "id": 1057483,
          "postDate": "2020-10-22T17:27:28.103Z",
          "content": "<p>I guess the simple NN won't work well. Try sequential models like RNN or transformers. I didn't check Boosting techniques based kernels, but even if these don't use sequential properties, they might create features about users' statistics based on their interaction sequences. So it is not surprising that it works better than simple NN.</p>",
          "rawMarkdown": "I guess the simple NN won't work well. Try sequential models like RNN or transformers. I didn't check Boosting techniques based kernels, but even if these don't use sequential properties, they might create features about users' statistics based on their interaction sequences. So it is not surprising that it works better than simple NN.",
          "votes": 3
        },
        {
          "id": 1057513,
          "postDate": "2020-10-22T18:01:39.927Z",
          "content": "<p>My simple NN has a consistent ~.73[789]-.74 on LB, So it depends how you are using them… (and trained on &lt; 1M)</p>",
          "rawMarkdown": "My simple NN has a consistent ~.73[789]-.74 on LB, So it depends how you are using them... (and trained on < 1M)",
          "votes": 1
        },
        {
          "id": 1057524,
          "postDate": "2020-10-22T18:12:14.367Z",
          "content": "<p>Mine too, my public incremental learning kernel has a very simple one which gets some 0.73 or so after seeing just a few thousand data rows.<br>\nSo I guess you have a bug somewhere…</p>",
          "rawMarkdown": "Mine too, my public incremental learning kernel has a very simple one which gets some 0.73 or so after seeing just a few thousand data rows.\nSo I guess you have a bug somewhere...",
          "votes": 2
        },
        {
          "id": 1057566,
          "postDate": "2020-10-22T19:04:01.677Z",
          "content": "<p>Okay guys, thanks for the help!</p>",
          "rawMarkdown": "Okay guys, thanks for the help!"
        }
      ]
    },
    {
      "id": 1057209,
      "postDate": "2020-10-22T13:24:57.570Z",
      "content": "<p>Neural networks are certainly learning too (plenty of examples in public kernels)… just not as well as boosting methods. But I'm sure one could get a decent score with NNs too - maybe check your architecture and code, that sounds more like something is structurally wrong if it's not learning.</p>",
      "rawMarkdown": "Neural networks are certainly learning too (plenty of examples in public kernels)... just not as well as boosting methods. But I'm sure one could get a decent score with NNs too - maybe check your architecture and code, that sounds more like something is structurally wrong if it's not learning.",
      "votes": 2
    },
    {
      "id": 1057201,
      "postDate": "2020-10-22T13:19:02.390Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1057584,
      "author_name": "Pavel Orlov",
      "author_url": "",
      "post_date": "2020-10-22T19:21:02.633000",
      "content": "<p>The best model that I met in <a href=\"https://arxiv.org/abs/2002.07033\" target=\"_blank\">one of the articles</a> was just a neural network (transformer). And it gave a result (on the dataset from Riiid) slightly higher than the current leader. From this I conclude that the best result will be approximately 0.80-0.83</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1057604,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-10-22T19:43:43.137000",
          "content": "<p>One thing to consider is that if such model can finish the prediction in 2 hours with Kaggle GPU or in 9 hours of CPU. I think this will be changeling. Of course, we can try to reduce the model size, but the question becomes if we can achieve the same result in this paper with smaller models that are fast to predict.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1057193,
      "author_name": "MhdSharuk",
      "author_url": "",
      "post_date": "2020-10-22T13:09:16.060000",
      "content": "<p>What type of deep learning architecture did you use to confirm this??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1057247,
          "author_name": "Abdessalem Boukil",
          "author_url": "",
          "post_date": "2020-10-22T14:06:08.317000",
          "content": "<p>I used a regular architecture. Features -&gt; (64 NN) -&gt; (Relu + Dropout + LayerNorm) -&gt; (64 NN) -&gt; (Relu + Dropout + LayerNorm) -&gt; Map it to two inputs. Trained on ADAM, and crossentropyloss.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1057483,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2020-10-22T17:27:28.103000",
          "content": "<p>I guess the simple NN won't work well. Try sequential models like RNN or transformers. I didn't check Boosting techniques based kernels, but even if these don't use sequential properties, they might create features about users' statistics based on their interaction sequences. So it is not surprising that it works better than simple NN.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1057513,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-22T18:01:39.927000",
          "content": "<p>My simple NN has a consistent ~.73[789]-.74 on LB, So it depends how you are using them… (and trained on &lt; 1M)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1057524,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-22T18:12:14.367000",
          "content": "<p>Mine too, my public incremental learning kernel has a very simple one which gets some 0.73 or so after seeing just a few thousand data rows.<br>\nSo I guess you have a bug somewhere…</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1057566,
          "author_name": "Abdessalem Boukil",
          "author_url": "",
          "post_date": "2020-10-22T19:04:01.677000",
          "content": "<p>Okay guys, thanks for the help!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1057209,
      "author_name": "Alex Bader",
      "author_url": "",
      "post_date": "2020-10-22T13:24:57.570000",
      "content": "<p>Neural networks are certainly learning too (plenty of examples in public kernels)… just not as well as boosting methods. But I'm sure one could get a decent score with NNs too - maybe check your architecture and code, that sounds more like something is structurally wrong if it's not learning.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1057201,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-22T13:19:02.390000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1057150": "I am trying to create a baseline model to start with, and since I am more acquainted with deep learning rather than classical machine learning, I am attempting to do it with deep neural networks. However the neural nets are barely learning, played with the loss/optimizer/lr ... nothing, the results seem to be more arbitrary rather than the model learning.\nSeeing the notebooks here though, it seems that boosting methods are having more success with it, is it the architecture of my NN that prevents it from learning, or is that boosting is better equipped for these types of problems. Thanks!\nAnd also sorry for my newbie question.",
    "1057584": "The best model that I met in [one of the articles](https://arxiv.org/abs/2002.07033) was just a neural network (transformer). And it gave a result (on the dataset from Riiid) slightly higher than the current leader. From this I conclude that the best result will be approximately 0.80-0.83",
    "1057193": "What type of deep learning architecture did you use to confirm this??",
    "1057209": "Neural networks are certainly learning too (plenty of examples in public kernels)... just not as well as boosting methods. But I'm sure one could get a decent score with NNs too - maybe check your architecture and code, that sounds more like something is structurally wrong if it's not learning.",
    "1057201": ""
  }
}