{
  "id": 542758,
  "title": "Missing values vs correlations",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/542758",
  "author_name": "DavidHGuerrero",
  "post_date": "2024-10-26T16:35:08.684000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>in this case, in the correlations, it is necessary to consider that many indicators have NaN. I have generated this image (with Whiteboard and Paint of Windows 10) where these two variables can be more easily visualized.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F981311%2Fea4486b89148db5839afde07218be992%2Fnull%20(1).png?generation=1729959939460056&amp;alt=media\" alt=\"\"></p>\n<p>Source of Missing values</p>\n<pre><code>percent_missing = train_ex.isnull().() *  / (train_ex)\nmissing_value_df = pd.DataFrame({: train_ex.columns, : percent_missing})\nmissing_value_df.plot.bar(figsize=(, ))\n</code></pre>\n<p>Source to get correlations</p>\n<pre><code> sklearn.preprocessing  OrdinalEncoder\n\nlabel_train_ex = train_ex.copy()\n\nordinal_encoder = OrdinalEncoder()\nlabel_train_ex[object_cols] = ordinal_encoder.fit_transform(train_ex[object_cols])\nlabel_train_ex[object_cols].head()\n</code></pre>\n<pre><code>fig, ax = plt.subplots(figsize=(, ))\n\ncorr = label_train_ex[features_ex].corr()\nmask = np.zeros_like(corr)\nmask[np.triu_indices_from(mask)] = \nax = sns.heatmap(corr, mask=mask, vmin=-, vmax=,\n                 cmap=sns.diverging_palette(, , as_cmap=),\n                 annot=)\nplt.tight_layout()\nplt.show()\n</code></pre>\n<p>Happy kaggling !</p>\n<p>David =)</p>",
  "messages": [
    {
      "id": 3028924,
      "postDate": "2024-10-26T16:35:08.683Z",
      "content": "<p>Hello,</p>\n<p>in this case, in the correlations, it is necessary to consider that many indicators have NaN. I have generated this image (with Whiteboard and Paint of Windows 10) where these two variables can be more easily visualized.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F981311%2Fea4486b89148db5839afde07218be992%2Fnull%20(1).png?generation=1729959939460056&amp;alt=media\" alt=\"\"></p>\n<p>Source of Missing values</p>\n<pre><code>percent_missing = train_ex.isnull().() *  / (train_ex)\nmissing_value_df = pd.DataFrame({: train_ex.columns, : percent_missing})\nmissing_value_df.plot.bar(figsize=(, ))\n</code></pre>\n<p>Source to get correlations</p>\n<pre><code> sklearn.preprocessing  OrdinalEncoder\n\nlabel_train_ex = train_ex.copy()\n\nordinal_encoder = OrdinalEncoder()\nlabel_train_ex[object_cols] = ordinal_encoder.fit_transform(train_ex[object_cols])\nlabel_train_ex[object_cols].head()\n</code></pre>\n<pre><code>fig, ax = plt.subplots(figsize=(, ))\n\ncorr = label_train_ex[features_ex].corr()\nmask = np.zeros_like(corr)\nmask[np.triu_indices_from(mask)] = \nax = sns.heatmap(corr, mask=mask, vmin=-, vmax=,\n                 cmap=sns.diverging_palette(, , as_cmap=),\n                 annot=)\nplt.tight_layout()\nplt.show()\n</code></pre>\n<p>Happy kaggling !</p>\n<p>David =)</p>",
      "rawMarkdown": "Hello,\n\nin this case, in the correlations, it is necessary to consider that many indicators have NaN. I have generated this image (with Whiteboard and Paint of Windows 10) where these two variables can be more easily visualized.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F981311%2Fea4486b89148db5839afde07218be992%2Fnull%20(1).png?generation=1729959939460056&alt=media)\n\nSource of Missing values\n\n```python\npercent_missing = train_ex.isnull().sum() * 100 / len(train_ex)\nmissing_value_df = pd.DataFrame({'column_name': train_ex.columns, 'percent_missing': percent_missing})\nmissing_value_df.plot.bar(figsize=(18, 4))\n```\nSource to get correlations\n```python\nfrom sklearn.preprocessing import OrdinalEncoder\n# Make copy to avoid changing original data \nlabel_train_ex = train_ex.copy()\n# Apply ordinal encoder to each column with categorical data\nordinal_encoder = OrdinalEncoder()\nlabel_train_ex[object_cols] = ordinal_encoder.fit_transform(train_ex[object_cols])\nlabel_train_ex[object_cols].head()\n```\n\n```python\nfig, ax = plt.subplots(figsize=(28, 16))\n# https://www.kaggle.com/code/sergiosaharovskiy/ps-s3e11-2023-eda-and-submission?scriptVersionId=123851844&cellId=26\ncorr = label_train_ex[features_ex].corr()\nmask = np.zeros_like(corr)\nmask[np.triu_indices_from(mask)] = True\nax = sns.heatmap(corr, mask=mask, vmin=-1, vmax=1,\n                 cmap=sns.diverging_palette(20, 220, as_cmap=True),\n                 annot=False)\nplt.tight_layout()\nplt.show()\n```\n\nHappy kaggling !\n\nDavid =)",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3028924": "Hello,\n\nin this case, in the correlations, it is necessary to consider that many indicators have NaN. I have generated this image (with Whiteboard and Paint of Windows 10) where these two variables can be more easily visualized.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F981311%2Fea4486b89148db5839afde07218be992%2Fnull%20(1).png?generation=1729959939460056&alt=media)\n\nSource of Missing values\n\n```python\npercent_missing = train_ex.isnull().sum() * 100 / len(train_ex)\nmissing_value_df = pd.DataFrame({'column_name': train_ex.columns, 'percent_missing': percent_missing})\nmissing_value_df.plot.bar(figsize=(18, 4))\n```\nSource to get correlations\n```python\nfrom sklearn.preprocessing import OrdinalEncoder\n# Make copy to avoid changing original data \nlabel_train_ex = train_ex.copy()\n# Apply ordinal encoder to each column with categorical data\nordinal_encoder = OrdinalEncoder()\nlabel_train_ex[object_cols] = ordinal_encoder.fit_transform(train_ex[object_cols])\nlabel_train_ex[object_cols].head()\n```\n\n```python\nfig, ax = plt.subplots(figsize=(28, 16))\n# https://www.kaggle.com/code/sergiosaharovskiy/ps-s3e11-2023-eda-and-submission?scriptVersionId=123851844&cellId=26\ncorr = label_train_ex[features_ex].corr()\nmask = np.zeros_like(corr)\nmask[np.triu_indices_from(mask)] = True\nax = sns.heatmap(corr, mask=mask, vmin=-1, vmax=1,\n                 cmap=sns.diverging_palette(20, 220, as_cmap=True),\n                 annot=False)\nplt.tight_layout()\nplt.show()\n```\n\nHappy kaggling !\n\nDavid =)"
  }
}