{
  "id": 76003,
  "title": "Things that didn't work",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76003",
  "author_name": "",
  "post_date": "2018-12-28T11:42:15.376590600Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I wanted to provide a list of things that I tried and didn't work for me. My approach is mainly using many deep learning models with most of the training data and then blending with a linear model with the remaining training data. Following this approach, the following <em>didn't</em> work.</p>\n\n<ul>\n<li>Using f1 loss as the loss function in keras (as in <a href=\"https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric\">https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric</a>) didn't work for me. The regular cross-entropy loss function gave better f1 scores, ironically.</li>\n<li>Reducing the dimensionality (padding the sentence to 50 words, and using 70000 words as the vocabulary, in comparison with the usual 70 and 95000). To me, it looked like it would make the learning easier for the NN, but it fails to give f1 scores higher than 0.65. </li>\n<li>Using ridge regression for blending gave way worse scores than lasso.</li>\n<li>The results of any linear model are not comparable to neural nets. I'd like to work a little bit more in this direction, as the results of neural nets correlate a lot, which is bad for blending.</li>\n</ul>\n\n<p>I'll be updating more things that didn't work.</p>\n\n<p>Updates:</p>\n\n<ul>\n<li>The size of the training set used for blending should be quite small. I plotted validation curves and everything above 10% seemed to be a waste of data for the neural nets.</li>\n<li>When using Tfidf + Logistic Regression, using ngrams up to 3 words works well. However, all of them are 1000000 and it is very hard to train classifiers on them. Using a maximum of 50000 ngrams and regularizing the logistic regression has worked the best for me, as in <a href=\"https://www.kaggle.com/david26694/nb-svm-baseline-trigrams\">NBSVM with trigrams</a>.</li>\n<li>CNN seem to have a very erratic behaviour, as reported in other threads.</li>\n<li>When using my best linear model as an input for the stacking with the Lasso, the Lasso doesn't select it. For this reason, I would dismiss any linear model if there are neural nets for the stacking stage.</li>\n<li>When training the lasso, I did it using grid search of the parameter alpha and all the glmnet stuff. Glmnet usually chooses the estimator at 1 standard deviation of the best estimator to reduce overfitting. Here, if we do that, 1 standard deviation is too much and we end up with too much regularisation.</li>\n<li>Kind of stupid, but I did a lasso with the variables selected from the lasso and got the same coefficients. Now I know that re-lassoing doesn't make sense.</li>\n<li>Rules didn't work. Rules like: if there's the word 'trump', just set it to insencere didn't work. It is better for the model to learn the behaviours itself.</li>\n</ul>",
  "messages": [
    {
      "id": "446614",
      "postDate": "12/28/2018 11:42:15",
      "content": "<p>I wanted to provide a list of things that I tried and didn't work for me. My approach is mainly using many deep learning models with most of the training data and then blending with a linear model with the remaining training data. Following this approach, the following <em>didn't</em> work.</p>\n\n<ul>\n<li>Using f1 loss as the loss function in keras (as in <a href=\"https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric\">https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric</a>) didn't work for me. The regular cross-entropy loss function gave better f1 scores, ironically.</li>\n<li>Reducing the dimensionality (padding the sentence to 50 words, and using 70000 words as the vocabulary, in comparison with the usual 70 and 95000). To me, it looked like it would make the learning easier for the NN, but it fails to give f1 scores higher than 0.65. </li>\n<li>Using ridge regression for blending gave way worse scores than lasso.</li>\n<li>The results of any linear model are not comparable to neural nets. I'd like to work a little bit more in this direction, as the results of neural nets correlate a lot, which is bad for blending.</li>\n</ul>\n\n<p>I'll be updating more things that didn't work.</p>\n\n<p>Updates:</p>\n\n<ul>\n<li>The size of the training set used for blending should be quite small. I plotted validation curves and everything above 10% seemed to be a waste of data for the neural nets.</li>\n<li>When using Tfidf + Logistic Regression, using ngrams up to 3 words works well. However, all of them are 1000000 and it is very hard to train classifiers on them. Using a maximum of 50000 ngrams and regularizing the logistic regression has worked the best for me, as in <a href=\"https://www.kaggle.com/david26694/nb-svm-baseline-trigrams\">NBSVM with trigrams</a>.</li>\n<li>CNN seem to have a very erratic behaviour, as reported in other threads.</li>\n<li>When using my best linear model as an input for the stacking with the Lasso, the Lasso doesn't select it. For this reason, I would dismiss any linear model if there are neural nets for the stacking stage.</li>\n<li>When training the lasso, I did it using grid search of the parameter alpha and all the glmnet stuff. Glmnet usually chooses the estimator at 1 standard deviation of the best estimator to reduce overfitting. Here, if we do that, 1 standard deviation is too much and we end up with too much regularisation.</li>\n<li>Kind of stupid, but I did a lasso with the variables selected from the lasso and got the same coefficients. Now I know that re-lassoing doesn't make sense.</li>\n<li>Rules didn't work. Rules like: if there's the word 'trump', just set it to insencere didn't work. It is better for the model to learn the behaviours itself.</li>\n</ul>",
      "rawMarkdown": "I wanted to provide a list of things that I tried and didn't work for me. My approach is mainly using many deep learning models with most of the training data and then blending with a linear model with the remaining training data. Following this approach, the following *didn't* work.\n\n- Using f1 loss as the loss function in keras (as in https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric) didn't work for me. The regular cross-entropy loss function gave better f1 scores, ironically.\n- Reducing the dimensionality (padding the sentence to 50 words, and using 70000 words as the vocabulary, in comparison with the usual 70 and 95000). To me, it looked like it would make the learning easier for the NN, but it fails to give f1 scores higher than 0.65. \n- Using ridge regression for blending gave way worse scores than lasso.\n- The results of any linear model are not comparable to neural nets. I'd like to work a little bit more in this direction, as the results of neural nets correlate a lot, which is bad for blending.\n\nI'll be updating more things that didn't work.\n\nUpdates:\n\n- The size of the training set used for blending should be quite small. I plotted validation curves and everything above 10% seemed to be a waste of data for the neural nets.\n- When using Tfidf + Logistic Regression, using ngrams up to 3 words works well. However, all of them are 1000000 and it is very hard to train classifiers on them. Using a maximum of 50000 ngrams and regularizing the logistic regression has worked the best for me, as in [NBSVM with trigrams][1].\n- CNN seem to have a very erratic behaviour, as reported in other threads.\n- When using my best linear model as an input for the stacking with the Lasso, the Lasso doesn't select it. For this reason, I would dismiss any linear model if there are neural nets for the stacking stage.\n- When training the lasso, I did it using grid search of the parameter alpha and all the glmnet stuff. Glmnet usually chooses the estimator at 1 standard deviation of the best estimator to reduce overfitting. Here, if we do that, 1 standard deviation is too much and we end up with too much regularisation.\n- Kind of stupid, but I did a lasso with the variables selected from the lasso and got the same coefficients. Now I know that re-lassoing doesn't make sense.\n- Rules didn't work. Rules like: if there's the word 'trump', just set it to insencere didn't work. It is better for the model to learn the behaviours itself.\n\n  [1]: https://www.kaggle.com/david26694/nb-svm-baseline-trigrams",
      "votes": null
    },
    {
      "id": "446616",
      "postDate": "12/28/2018 11:43:49",
      "content": "<ul>\n<li>The size of the training set used for blending should be quite small. I plotted validation curves and everything above 10% seemed to be a waste of data for the neural nets. </li>\n</ul>",
      "rawMarkdown": "The size of the training set used for blending should be quite small. I plotted validation curves and everything above 10% seemed to be a waste of data for the neural nets.",
      "votes": null
    },
    {
      "id": "446682",
      "postDate": "12/28/2018 13:40:25",
      "content": "<p>Thanks for sharing! I'm still trying to get my head around some of the NN architectures but will after that also share some insights.</p>",
      "rawMarkdown": "Thanks for sharing! I'm still trying to get my head around some of the NN architectures but will after that also share some insights.",
      "votes": null
    },
    {
      "id": "446885",
      "postDate": "12/28/2018 20:34:59",
      "content": "<ul>\n<li>When using Tfidf + Logistic Regression, using ngrams up to 3 words works well.  However, all of them are 1000000 and it is very hard to train classifiers on them. Using a maximum of 50000 ngrams and regularizing the logistic regression has worked the best for me.</li>\n</ul>",
      "rawMarkdown": "When using Tfidf + Logistic Regression, using ngrams up to 3 words works well.  However, all of them are 1000000 and it is very hard to train classifiers on them. Using a maximum of 50000 ngrams and regularizing the logistic regression has worked the best for me.",
      "votes": null
    },
    {
      "id": "447173",
      "postDate": "12/29/2018 09:46:14",
      "content": "<ul>\n<li>CNN seem to have a very erratic behaviour, as reported in other threads</li>\n</ul>",
      "rawMarkdown": "CNN seem to have a very erratic behaviour, as reported in other threads",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 446616,
      "author_name": "david26694",
      "author_url": "",
      "post_date": "12/28/2018 11:43:49",
      "content": "<ul>\n<li>The size of the training set used for blending should be quite small. I plotted validation curves and everything above 10% seemed to be a waste of data for the neural nets. </li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 446682,
      "author_name": "timothylucas",
      "author_url": "",
      "post_date": "12/28/2018 13:40:25",
      "content": "<p>Thanks for sharing! I'm still trying to get my head around some of the NN architectures but will after that also share some insights.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 446885,
      "author_name": "david26694",
      "author_url": "",
      "post_date": "12/28/2018 20:34:59",
      "content": "<ul>\n<li>When using Tfidf + Logistic Regression, using ngrams up to 3 words works well.  However, all of them are 1000000 and it is very hard to train classifiers on them. Using a maximum of 50000 ngrams and regularizing the logistic regression has worked the best for me.</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 447173,
      "author_name": "david26694",
      "author_url": "",
      "post_date": "12/29/2018 09:46:14",
      "content": "<ul>\n<li>CNN seem to have a very erratic behaviour, as reported in other threads</li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "446614": "I wanted to provide a list of things that I tried and didn't work for me. My approach is mainly using many deep learning models with most of the training data and then blending with a linear model with the remaining training data. Following this approach, the following *didn't* work.\n\n- Using f1 loss as the loss function in keras (as in https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric) didn't work for me. The regular cross-entropy loss function gave better f1 scores, ironically.\n- Reducing the dimensionality (padding the sentence to 50 words, and using 70000 words as the vocabulary, in comparison with the usual 70 and 95000). To me, it looked like it would make the learning easier for the NN, but it fails to give f1 scores higher than 0.65. \n- Using ridge regression for blending gave way worse scores than lasso.\n- The results of any linear model are not comparable to neural nets. I'd like to work a little bit more in this direction, as the results of neural nets correlate a lot, which is bad for blending.\n\nI'll be updating more things that didn't work.\n\nUpdates:\n\n- The size of the training set used for blending should be quite small. I plotted validation curves and everything above 10% seemed to be a waste of data for the neural nets.\n- When using Tfidf + Logistic Regression, using ngrams up to 3 words works well. However, all of them are 1000000 and it is very hard to train classifiers on them. Using a maximum of 50000 ngrams and regularizing the logistic regression has worked the best for me, as in [NBSVM with trigrams][1].\n- CNN seem to have a very erratic behaviour, as reported in other threads.\n- When using my best linear model as an input for the stacking with the Lasso, the Lasso doesn't select it. For this reason, I would dismiss any linear model if there are neural nets for the stacking stage.\n- When training the lasso, I did it using grid search of the parameter alpha and all the glmnet stuff. Glmnet usually chooses the estimator at 1 standard deviation of the best estimator to reduce overfitting. Here, if we do that, 1 standard deviation is too much and we end up with too much regularisation.\n- Kind of stupid, but I did a lasso with the variables selected from the lasso and got the same coefficients. Now I know that re-lassoing doesn't make sense.\n- Rules didn't work. Rules like: if there's the word 'trump', just set it to insencere didn't work. It is better for the model to learn the behaviours itself.\n\n  [1]: https://www.kaggle.com/david26694/nb-svm-baseline-trigrams",
    "446616": "The size of the training set used for blending should be quite small. I plotted validation curves and everything above 10% seemed to be a waste of data for the neural nets.",
    "446682": "Thanks for sharing! I'm still trying to get my head around some of the NN architectures but will after that also share some insights.",
    "446885": "When using Tfidf + Logistic Regression, using ngrams up to 3 words works well.  However, all of them are 1000000 and it is very hard to train classifiers on them. Using a maximum of 50000 ngrams and regularizing the logistic regression has worked the best for me.",
    "447173": "CNN seem to have a very erratic behaviour, as reported in other threads"
  },
  "source": "meta"
}