{
  "id": 315407,
  "title": "Dataframe specification?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/315407",
  "author_name": "",
  "post_date": "2022-03-28T06:13:30.122126400Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<pre><code>kf = StratifiedKFold(n_splits = 5)\nfor fold, ( _, val_) in enumerate(kf.split(X=df, y=df['individual_id'])):\n    df.loc[val_ , \"kfold\"] = fold\n</code></pre>\n<p>When I try this code, df['kfold'] is a float.<br>\nHow can this happen even though 'fold' is int?</p>",
  "messages": [
    {
      "id": "1737129",
      "postDate": "03/28/2022 06:13:30",
      "content": "<pre><code>kf = StratifiedKFold(n_splits = 5)\nfor fold, ( _, val_) in enumerate(kf.split(X=df, y=df['individual_id'])):\n    df.loc[val_ , \"kfold\"] = fold\n</code></pre>\n<p>When I try this code, df['kfold'] is a float.<br>\nHow can this happen even though 'fold' is int?</p>",
      "rawMarkdown": "```\nkf = StratifiedKFold(n_splits = 5)\nfor fold, ( _, val_) in enumerate(kf.split(X=df, y=df['individual_id'])):\n    df.loc[val_ , \"kfold\"] = fold\n```\nWhen I try this code, df['kfold'] is a float.\nHow can this happen even though 'fold' is int?",
      "votes": null
    },
    {
      "id": "1737802",
      "postDate": "03/28/2022 18:09:14",
      "content": "<p>This is a case of upcasting to prevent data error and loss. </p>",
      "rawMarkdown": "This is a case of upcasting to prevent data error and loss.",
      "votes": null
    },
    {
      "id": "1737938",
      "postDate": "03/28/2022 21:57:19",
      "content": "<p>The first time you call <code>df.loc[val_ , \"kfold\"] = fold</code>, then 20% of the rows get an <code>int</code> and 80% of the rows get <code>NAN</code>. The value <code>NAN</code> is a <code>float</code>, so the column becomes <code>float</code>. If you want to avoid this, then before the for-loop, you can initialize the column like <code>df['kfold'] = -1</code>. Then the for-loop will not need to add <code>NAN</code></p>",
      "rawMarkdown": "The first time you call `df.loc[val_ , \"kfold\"] = fold`, then 20% of the rows get an `int` and 80% of the rows get `NAN`. The value `NAN` is a `float`, so the column becomes `float`. If you want to avoid this, then before the for-loop, you can initialize the column like `df['kfold'] = -1`. Then the for-loop will not need to add `NAN`",
      "votes": null
    },
    {
      "id": "1738181",
      "postDate": "03/29/2022 04:59:42",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>!</p>",
      "rawMarkdown": "Thanks, @ravi20076!",
      "votes": null
    },
    {
      "id": "1738182",
      "postDate": "03/29/2022 05:00:23",
      "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>!<br>\nI really understand.</p>",
      "rawMarkdown": "Thanks, @cdeotte!\nI really understand.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1737802,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "03/28/2022 18:09:14",
      "content": "<p>This is a case of upcasting to prevent data error and loss. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1738181,
          "author_name": "naokisugimura",
          "author_url": "",
          "post_date": "03/29/2022 04:59:42",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a>!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1737938,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/28/2022 21:57:19",
      "content": "<p>The first time you call <code>df.loc[val_ , \"kfold\"] = fold</code>, then 20% of the rows get an <code>int</code> and 80% of the rows get <code>NAN</code>. The value <code>NAN</code> is a <code>float</code>, so the column becomes <code>float</code>. If you want to avoid this, then before the for-loop, you can initialize the column like <code>df['kfold'] = -1</code>. Then the for-loop will not need to add <code>NAN</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1738182,
          "author_name": "naokisugimura",
          "author_url": "",
          "post_date": "03/29/2022 05:00:23",
          "content": "<p>Thanks, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>!<br>\nI really understand.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1737129": "```\nkf = StratifiedKFold(n_splits = 5)\nfor fold, ( _, val_) in enumerate(kf.split(X=df, y=df['individual_id'])):\n    df.loc[val_ , \"kfold\"] = fold\n```\nWhen I try this code, df['kfold'] is a float.\nHow can this happen even though 'fold' is int?",
    "1737802": "This is a case of upcasting to prevent data error and loss.",
    "1737938": "The first time you call `df.loc[val_ , \"kfold\"] = fold`, then 20% of the rows get an `int` and 80% of the rows get `NAN`. The value `NAN` is a `float`, so the column becomes `float`. If you want to avoid this, then before the for-loop, you can initialize the column like `df['kfold'] = -1`. Then the for-loop will not need to add `NAN`",
    "1738181": "Thanks, @ravi20076!",
    "1738182": "Thanks, @cdeotte!\nI really understand."
  },
  "source": "meta"
}