{
  "id": 75180,
  "title": "how to find the index for missing questions?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75180",
  "author_name": "",
  "post_date": "2018-12-19T07:30:04.563167800Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello everyone!\nWhen I tried to find the \"text_questions\" with missing values using the following command</p>\n\n<pre><code>trainData[\"question_text\"].isnull().any()\n</code></pre>\n\n<p>it shows that all the values in <em>question_text</em> are present by giving the output <em>false</em> . </p>\n\n<p>Please help me with the piece of <em>code</em> to find the missing values in <em>question_text</em> column and also help me understand, why my method is not working.</p>\n\n<p>Thank you! </p>",
  "messages": [
    {
      "id": "441873",
      "postDate": "12/19/2018 07:30:04",
      "content": "<p>Hello everyone!\nWhen I tried to find the \"text_questions\" with missing values using the following command</p>\n\n<pre><code>trainData[\"question_text\"].isnull().any()\n</code></pre>\n\n<p>it shows that all the values in <em>question_text</em> are present by giving the output <em>false</em> . </p>\n\n<p>Please help me with the piece of <em>code</em> to find the missing values in <em>question_text</em> column and also help me understand, why my method is not working.</p>\n\n<p>Thank you! </p>",
      "rawMarkdown": "Hello everyone!\nWhen I tried to find the \"text_questions\" with missing values using the following command\n\n    trainData[\"question_text\"].isnull().any()\n\nit shows that all the values in *question\\_text* are present by giving the output *false* . \n\nPlease help me with the piece of *code* to find the missing values in *question\\_text* column and also help me understand, why my method is not working.\n\nThank you!",
      "votes": null
    },
    {
      "id": "442007",
      "postDate": "12/19/2018 10:54:15",
      "content": "<p>Hi Ashutosh,\nhow could the main column be missing?\nBasically the question_text is the only column on which you may perform feature engineering.</p>\n\n<blockquote>\n  <p>qTrain =  pd.read_csv(\"train.csv\")</p>\n  \n  <p>questions = qTrain[\"question_text\"]\n  questions[questions.isna()]</p>\n</blockquote>\n\n<p>Series([], Name: question_text, dtype: object)</p>",
      "rawMarkdown": "Hi Ashutosh,\nhow could the main column be missing?\nBasically the question_text is the only column on which you may perform feature engineering.\n\n&gt; qTrain =  pd.read_csv(\"train.csv\")\n\n&gt; questions = qTrain[\"question_text\"]\n&gt; questions[questions.isna()]\n\t\t\t\t\t     \nSeries([], Name: question_text, dtype: object)",
      "votes": null
    },
    {
      "id": "442597",
      "postDate": "12/20/2018 07:14:11",
      "content": "<p>in <a href=\"https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings/notebook\">this notebook</a> the author had tried to fill the missing values, which made me think twice, and i concluded that something was wrong with my piece of code.</p>\n\n<p>hope you can help me with the purpose of this code</p>\n\n<pre><code> ## fill up the missing values\n train_X = train_df[\"question_text\"].fillna(\"_na_\").values\n val_X = val_df[\"question_text\"].fillna(\"_na_\").values\n test_X = test_df[\"question_text\"].fillna(\"_na_\").values\n</code></pre>\n\n<p>what purpose does it serve?</p>",
      "rawMarkdown": "in [this notebook](https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings/notebook) the author had tried to fill the missing values, which made me think twice, and i concluded that something was wrong with my piece of code.\n\nhope you can help me with the purpose of this code\n\n     ## fill up the missing values\n     train_X = train_df[\"question_text\"].fillna(\"_na_\").values\n     val_X = val_df[\"question_text\"].fillna(\"_na_\").values\n     test_X = test_df[\"question_text\"].fillna(\"_na_\").values\n\nwhat purpose does it serve?",
      "votes": null
    },
    {
      "id": "442626",
      "postDate": "12/20/2018 08:27:01",
      "content": "<p>I think that was just a standard operation within preprocessing which SRK used to do - but in this case the calling of the function fillna(\"<em>na</em>\") does not make any change - here is the comparison of their results:</p>\n\n<blockquote>\n  <p>train_X = train_df[\"question_text\"].fillna(\"<em>na</em>\").values</p>\n  \n  <p>train_X2 = train_df[\"question_text\"].values</p>\n  \n  <p>ar2 = (train_X2 == train_X)</p>\n  \n  <p>np.all(ar2[:] == True)</p>\n</blockquote>\n\n<p>True</p>",
      "rawMarkdown": "I think that was just a standard operation within preprocessing which SRK used to do - but in this case the calling of the function fillna(\"_na_\") does not make any change - here is the comparison of their results:\n\n&gt; train_X = train_df[\"question_text\"].fillna(\"_na_\").values\n\n&gt; train_X2 = train_df[\"question_text\"].values\n\n&gt; ar2 = (train_X2 == train_X)\n\n&gt; np.all(ar2[:] == True)\n\nTrue",
      "votes": null
    },
    {
      "id": "445874",
      "postDate": "12/27/2018 06:40:42",
      "content": "<p>That's because it returns a Boolean pd.series(you requested that , so its returning that), just add train[your code here] and you will get what you are looking for probably!</p>",
      "rawMarkdown": "That's because it returns a Boolean pd.series(you requested that , so its returning that), just add train[your code here] and you will get what you are looking for probably!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 442007,
      "author_name": "akuropatwinski",
      "author_url": "",
      "post_date": "12/19/2018 10:54:15",
      "content": "<p>Hi Ashutosh,\nhow could the main column be missing?\nBasically the question_text is the only column on which you may perform feature engineering.</p>\n\n<blockquote>\n  <p>qTrain =  pd.read_csv(\"train.csv\")</p>\n  \n  <p>questions = qTrain[\"question_text\"]\n  questions[questions.isna()]</p>\n</blockquote>\n\n<p>Series([], Name: question_text, dtype: object)</p>",
      "votes": null,
      "replies": [
        {
          "id": 442597,
          "author_name": "ashukr",
          "author_url": "",
          "post_date": "12/20/2018 07:14:11",
          "content": "<p>in <a href=\"https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings/notebook\">this notebook</a> the author had tried to fill the missing values, which made me think twice, and i concluded that something was wrong with my piece of code.</p>\n\n<p>hope you can help me with the purpose of this code</p>\n\n<pre><code> ## fill up the missing values\n train_X = train_df[\"question_text\"].fillna(\"_na_\").values\n val_X = val_df[\"question_text\"].fillna(\"_na_\").values\n test_X = test_df[\"question_text\"].fillna(\"_na_\").values\n</code></pre>\n\n<p>what purpose does it serve?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442626,
          "author_name": "akuropatwinski",
          "author_url": "",
          "post_date": "12/20/2018 08:27:01",
          "content": "<p>I think that was just a standard operation within preprocessing which SRK used to do - but in this case the calling of the function fillna(\"<em>na</em>\") does not make any change - here is the comparison of their results:</p>\n\n<blockquote>\n  <p>train_X = train_df[\"question_text\"].fillna(\"<em>na</em>\").values</p>\n  \n  <p>train_X2 = train_df[\"question_text\"].values</p>\n  \n  <p>ar2 = (train_X2 == train_X)</p>\n  \n  <p>np.all(ar2[:] == True)</p>\n</blockquote>\n\n<p>True</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 445874,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "12/27/2018 06:40:42",
      "content": "<p>That's because it returns a Boolean pd.series(you requested that , so its returning that), just add train[your code here] and you will get what you are looking for probably!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "441873": "Hello everyone!\nWhen I tried to find the \"text_questions\" with missing values using the following command\n\n    trainData[\"question_text\"].isnull().any()\n\nit shows that all the values in *question\\_text* are present by giving the output *false* . \n\nPlease help me with the piece of *code* to find the missing values in *question\\_text* column and also help me understand, why my method is not working.\n\nThank you!",
    "442007": "Hi Ashutosh,\nhow could the main column be missing?\nBasically the question_text is the only column on which you may perform feature engineering.\n\n&gt; qTrain =  pd.read_csv(\"train.csv\")\n\n&gt; questions = qTrain[\"question_text\"]\n&gt; questions[questions.isna()]\n\t\t\t\t\t     \nSeries([], Name: question_text, dtype: object)",
    "442597": "in [this notebook](https://www.kaggle.com/sudalairajkumar/a-look-at-different-embeddings/notebook) the author had tried to fill the missing values, which made me think twice, and i concluded that something was wrong with my piece of code.\n\nhope you can help me with the purpose of this code\n\n     ## fill up the missing values\n     train_X = train_df[\"question_text\"].fillna(\"_na_\").values\n     val_X = val_df[\"question_text\"].fillna(\"_na_\").values\n     test_X = test_df[\"question_text\"].fillna(\"_na_\").values\n\nwhat purpose does it serve?",
    "442626": "I think that was just a standard operation within preprocessing which SRK used to do - but in this case the calling of the function fillna(\"_na_\") does not make any change - here is the comparison of their results:\n\n&gt; train_X = train_df[\"question_text\"].fillna(\"_na_\").values\n\n&gt; train_X2 = train_df[\"question_text\"].values\n\n&gt; ar2 = (train_X2 == train_X)\n\n&gt; np.all(ar2[:] == True)\n\nTrue",
    "445874": "That's because it returns a Boolean pd.series(you requested that , so its returning that), just add train[your code here] and you will get what you are looking for probably!"
  },
  "source": "meta"
}