{
  "id": 74466,
  "title": "[Beginer][Keras]Keep getting 0 as prediction for every input",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74466",
  "author_name": "",
  "post_date": "2018-12-12T13:43:04.811314800Z",
  "votes": -1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I was trying to run simple model architectures on this dataset, i cleaned text using few basic regexes, made an embedding layer and then tried below architectures,\n1) embedding + avgpool + dense + dense sigmoid\n2)embedding + lstm + dense sigmoid\nfor both types my model keeps on predicting 0's for each input. Is it becuase of class imbalance or i am making mistake in model i am unsure.Please give me few pointers to begin.</p>",
  "messages": [
    {
      "id": "437785",
      "postDate": "12/12/2018 13:43:04",
      "content": "<p>I was trying to run simple model architectures on this dataset, i cleaned text using few basic regexes, made an embedding layer and then tried below architectures,\n1) embedding + avgpool + dense + dense sigmoid\n2)embedding + lstm + dense sigmoid\nfor both types my model keeps on predicting 0's for each input. Is it becuase of class imbalance or i am making mistake in model i am unsure.Please give me few pointers to begin.</p>",
      "rawMarkdown": "I was trying to run simple model architectures on this dataset, i cleaned text using few basic regexes, made an embedding layer and then tried below architectures,\n1) embedding + avgpool + dense + dense sigmoid\n2)embedding + lstm + dense sigmoid\nfor both types my model keeps on predicting 0's for each input. Is it becuase of class imbalance or i am making mistake in model i am unsure.Please give me few pointers to begin.",
      "votes": null
    },
    {
      "id": "437842",
      "postDate": "12/12/2018 15:51:33",
      "content": "<p>I think a little more information is needed before knowing for sure. First, i would make sure you have all the needed base layers: input -&gt; embedding -&gt; pooling or flatten -&gt; dense (maybe use relu) -&gt; dense (sigmoid). Secondly, It could be coming from overfitting if you're using too small of a batch size or too big of a epoch size. Lastly, as shown in my kernel (<a href=\"https://www.kaggle.com/mgancita/logit-regression-detailed-building-of-lstm-model?scriptVersionId=8098690\">https://www.kaggle.com/mgancita/logit-regression-detailed-building-of-lstm-model?scriptVersionId=8098690</a> ) only about 6.2% of the questions are insincere so this imbalance may be causing your model to be overfit to always predict 0.</p>",
      "rawMarkdown": "I think a little more information is needed before knowing for sure. First, i would make sure you have all the needed base layers: input -&gt; embedding -&gt; pooling or flatten -&gt; dense (maybe use relu) -&gt; dense (sigmoid). Secondly, It could be coming from overfitting if you're using too small of a batch size or too big of a epoch size. Lastly, as shown in my kernel (https://www.kaggle.com/mgancita/logit-regression-detailed-building-of-lstm-model?scriptVersionId=8098690 ) only about 6.2% of the questions are insincere so this imbalance may be causing your model to be overfit to always predict 0.",
      "votes": null
    },
    {
      "id": "438135",
      "postDate": "12/13/2018 06:44:06",
      "content": "<p>i am exactly using input -&gt; embedding -&gt; pooling or flatten -&gt; dense (maybe use relu) -&gt; dense (sigmoid) type of structure.I tired varying batch size(256-512-1024) and epochs(1-10). \nAbout overfitting there are other kernels out there who also use simple lstm models and yet they get atleast 0.60 score not sure why i would get all 0's</p>",
      "rawMarkdown": "i am exactly using input -&gt; embedding -&gt; pooling or flatten -&gt; dense (maybe use relu) -&gt; dense (sigmoid) type of structure.I tired varying batch size(256-512-1024) and epochs(1-10). \nAbout overfitting there are other kernels out there who also use simple lstm models and yet they get atleast 0.60 score not sure why i would get all 0's",
      "votes": null
    },
    {
      "id": "438664",
      "postDate": "12/14/2018 02:07:16",
      "content": "<p>I have not gone through your code. But I believe a big part of this is the imbalance between labels 0 and 1. This is known as the imbalanced classification problem. Because one label is in a much larger number than the others, the model has failed to learn and simply predict the majority class label.</p>\n\n<p>What you can do is to resample the data.</p>",
      "rawMarkdown": "I have not gone through your code. But I believe a big part of this is the imbalance between labels 0 and 1. This is known as the imbalanced classification problem. Because one label is in a much larger number than the others, the model has failed to learn and simply predict the majority class label.\n\nWhat you can do is to resample the data.",
      "votes": null
    },
    {
      "id": "439292",
      "postDate": "12/15/2018 05:13:58",
      "content": "<p>i get the imbalanced classification.But there are other kernels out here , who use similar simple arcitecture and yet get somewhere in region of 60%,\ntake this kernel for example \n<a href=\"https://www.kaggle.com/nitinaggarwal008/basic-model\">https://www.kaggle.com/nitinaggarwal008/basic-model</a>\ni examined and found same imbalance in train label sets\n 0    1127298\n1      74334</p>",
      "rawMarkdown": "i get the imbalanced classification.But there are other kernels out here , who use similar simple arcitecture and yet get somewhere in region of 60%,\ntake this kernel for example \nhttps://www.kaggle.com/nitinaggarwal008/basic-model\ni examined and found same imbalance in train label sets\n 0    1127298\n1      74334",
      "votes": null
    },
    {
      "id": "440054",
      "postDate": "12/17/2018 00:58:55",
      "content": "<p>Perhaps you can post the link to your kernel for us to see?</p>",
      "rawMarkdown": "Perhaps you can post the link to your kernel for us to see?",
      "votes": null
    },
    {
      "id": "443693",
      "postDate": "12/22/2018 07:05:44",
      "content": "<p>Got it resolved.I was not properly reading glove vectors in my Embedding layer.My embedding layer was always np.zeros array and hence the all 0 predictions.Feeling very stupid now.</p>",
      "rawMarkdown": "Got it resolved.I was not properly reading glove vectors in my Embedding layer.My embedding layer was always np.zeros array and hence the all 0 predictions.Feeling very stupid now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 437842,
      "author_name": "mgancita",
      "author_url": "",
      "post_date": "12/12/2018 15:51:33",
      "content": "<p>I think a little more information is needed before knowing for sure. First, i would make sure you have all the needed base layers: input -&gt; embedding -&gt; pooling or flatten -&gt; dense (maybe use relu) -&gt; dense (sigmoid). Secondly, It could be coming from overfitting if you're using too small of a batch size or too big of a epoch size. Lastly, as shown in my kernel (<a href=\"https://www.kaggle.com/mgancita/logit-regression-detailed-building-of-lstm-model?scriptVersionId=8098690\">https://www.kaggle.com/mgancita/logit-regression-detailed-building-of-lstm-model?scriptVersionId=8098690</a> ) only about 6.2% of the questions are insincere so this imbalance may be causing your model to be overfit to always predict 0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 438135,
          "author_name": "chansbest93",
          "author_url": "",
          "post_date": "12/13/2018 06:44:06",
          "content": "<p>i am exactly using input -&gt; embedding -&gt; pooling or flatten -&gt; dense (maybe use relu) -&gt; dense (sigmoid) type of structure.I tired varying batch size(256-512-1024) and epochs(1-10). \nAbout overfitting there are other kernels out there who also use simple lstm models and yet they get atleast 0.60 score not sure why i would get all 0's</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 438664,
      "author_name": "gxkok21",
      "author_url": "",
      "post_date": "12/14/2018 02:07:16",
      "content": "<p>I have not gone through your code. But I believe a big part of this is the imbalance between labels 0 and 1. This is known as the imbalanced classification problem. Because one label is in a much larger number than the others, the model has failed to learn and simply predict the majority class label.</p>\n\n<p>What you can do is to resample the data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 439292,
          "author_name": "chansbest93",
          "author_url": "",
          "post_date": "12/15/2018 05:13:58",
          "content": "<p>i get the imbalanced classification.But there are other kernels out here , who use similar simple arcitecture and yet get somewhere in region of 60%,\ntake this kernel for example \n<a href=\"https://www.kaggle.com/nitinaggarwal008/basic-model\">https://www.kaggle.com/nitinaggarwal008/basic-model</a>\ni examined and found same imbalance in train label sets\n 0    1127298\n1      74334</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440054,
          "author_name": "gxkok21",
          "author_url": "",
          "post_date": "12/17/2018 00:58:55",
          "content": "<p>Perhaps you can post the link to your kernel for us to see?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443693,
          "author_name": "chansbest93",
          "author_url": "",
          "post_date": "12/22/2018 07:05:44",
          "content": "<p>Got it resolved.I was not properly reading glove vectors in my Embedding layer.My embedding layer was always np.zeros array and hence the all 0 predictions.Feeling very stupid now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "437785": "I was trying to run simple model architectures on this dataset, i cleaned text using few basic regexes, made an embedding layer and then tried below architectures,\n1) embedding + avgpool + dense + dense sigmoid\n2)embedding + lstm + dense sigmoid\nfor both types my model keeps on predicting 0's for each input. Is it becuase of class imbalance or i am making mistake in model i am unsure.Please give me few pointers to begin.",
    "437842": "I think a little more information is needed before knowing for sure. First, i would make sure you have all the needed base layers: input -&gt; embedding -&gt; pooling or flatten -&gt; dense (maybe use relu) -&gt; dense (sigmoid). Secondly, It could be coming from overfitting if you're using too small of a batch size or too big of a epoch size. Lastly, as shown in my kernel (https://www.kaggle.com/mgancita/logit-regression-detailed-building-of-lstm-model?scriptVersionId=8098690 ) only about 6.2% of the questions are insincere so this imbalance may be causing your model to be overfit to always predict 0.",
    "438135": "i am exactly using input -&gt; embedding -&gt; pooling or flatten -&gt; dense (maybe use relu) -&gt; dense (sigmoid) type of structure.I tired varying batch size(256-512-1024) and epochs(1-10). \nAbout overfitting there are other kernels out there who also use simple lstm models and yet they get atleast 0.60 score not sure why i would get all 0's",
    "438664": "I have not gone through your code. But I believe a big part of this is the imbalance between labels 0 and 1. This is known as the imbalanced classification problem. Because one label is in a much larger number than the others, the model has failed to learn and simply predict the majority class label.\n\nWhat you can do is to resample the data.",
    "439292": "i get the imbalanced classification.But there are other kernels out here , who use similar simple arcitecture and yet get somewhere in region of 60%,\ntake this kernel for example \nhttps://www.kaggle.com/nitinaggarwal008/basic-model\ni examined and found same imbalance in train label sets\n 0    1127298\n1      74334",
    "440054": "Perhaps you can post the link to your kernel for us to see?",
    "443693": "Got it resolved.I was not properly reading glove vectors in my Embedding layer.My embedding layer was always np.zeros array and hence the all 0 predictions.Feeling very stupid now."
  },
  "source": "meta"
}