{
  "id": 341256,
  "title": "How to use df.loc in cudf???",
  "url": "/competitions/amex-default-prediction/discussion/341256",
  "author_name": "",
  "post_date": "2022-08-02T04:38:41.288596700Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>can I do something similar to this in cudf?<br>\n<code>df.loc[df['col'] == 10, 'col_is_10'] = 1</code></p>",
  "messages": [
    {
      "id": "1880827",
      "postDate": "08/02/2022 04:38:41",
      "content": "<p>can I do something similar to this in cudf?<br>\n<code>df.loc[df['col'] == 10, 'col_is_10'] = 1</code></p>",
      "rawMarkdown": "can I do something similar to this in cudf?\n`df.loc[df['col'] == 10, 'col_is_10'] = 1`",
      "votes": null
    },
    {
      "id": "1882136",
      "postDate": "08/03/2022 05:08:36",
      "content": "<p>Yes, you can do that in cudf. Is it not working?</p>\n<p>You may need to run the line <code>df['col_is_10'] = 0</code> before you run that command. I don't think you can use your line to create a new column if it doesn't exist already. (And I don't think you can create new column in Pandas that way either without setting to zero first).</p>",
      "rawMarkdown": "Yes, you can do that in cudf. Is it not working?\n\nYou may need to run the line `df['col_is_10'] = 0` before you run that command. I don't think you can use your line to create a new column if it doesn't exist already. (And I don't think you can create new column in Pandas that way either without setting to zero first).",
      "votes": null
    },
    {
      "id": "1884872",
      "postDate": "08/04/2022 17:52:24",
      "content": "<p>What about using <code>train_test_split</code> than <code>ShuffleSplit</code>, I tried before, but I got the error.</p>",
      "rawMarkdown": "What about using `train_test_split` than `ShuffleSplit`, I tried before, but I got the error.",
      "votes": null
    },
    {
      "id": "1884990",
      "postDate": "08/04/2022 19:26:17",
      "content": "<p>Try Chris's starter, and comment out the \"to_pandas\" line which I recall happens before the CV folds code. If that works, you have a template that might help fix any issue in your code. If it doesn't work, it gives concrete example of the error. :)</p>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></p>",
      "rawMarkdown": "Try Chris's starter, and comment out the \"to_pandas\" line which I recall happens before the CV folds code. If that works, you have a template that might help fix any issue in your code. If it doesn't work, it gives concrete example of the error. :)\n\nhttps://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
      "votes": null
    },
    {
      "id": "1885047",
      "postDate": "08/04/2022 20:13:06",
      "content": "<p>Yeah, it's possible when we use <code>to_pandas</code>, but, if there is any method with Cudf, it will better than using pandas, however I tried other method, it works well, Thanks</p>",
      "rawMarkdown": "Yeah, it's possible when we use `to_pandas`, but, if there is any method with Cudf, it will better than using pandas, however I tried other method, it works well, Thanks",
      "votes": null
    },
    {
      "id": "1885066",
      "postDate": "08/04/2022 20:58:26",
      "content": "<p>I'm not sure what you are trying to do BOOBA. The library RAPIDS does have a train test split if that's what you want. It's called </p>\n<pre><code>cuml.model_selection.train_test_split\n</code></pre>\n<p>The docs webpage is <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#model-selection-and-data-splitting\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I'm not sure what you are trying to do BOOBA. The library RAPIDS does have a train test split if that's what you want. It's called \n    \n    cuml.model_selection.train_test_split\n\nThe docs webpage is [here][1]\n\n[1]: https://docs.rapids.ai/api/cuml/stable/api.html#model-selection-and-data-splitting",
      "votes": null
    },
    {
      "id": "1885094",
      "postDate": "08/04/2022 21:43:12",
      "content": "<p>Thanks <strong>Chris Deotte</strong>, I have two ideas, I am working on them, and maybe those will help me to raise my score, the first is about using Target encoding with adding some noises, I know it's the overfitted method, and the second thing, is about using an special CV to use data more efficiently.</p>",
      "rawMarkdown": "Thanks **Chris Deotte**, I have two ideas, I am working on them, and maybe those will help me to raise my score, the first is about using Target encoding with adding some noises, I know it's the overfitted method, and the second thing, is about using an special CV to use data more efficiently.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1882136,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/03/2022 05:08:36",
      "content": "<p>Yes, you can do that in cudf. Is it not working?</p>\n<p>You may need to run the line <code>df['col_is_10'] = 0</code> before you run that command. I don't think you can use your line to create a new column if it doesn't exist already. (And I don't think you can create new column in Pandas that way either without setting to zero first).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1884872,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "08/04/2022 17:52:24",
          "content": "<p>What about using <code>train_test_split</code> than <code>ShuffleSplit</code>, I tried before, but I got the error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1884990,
          "author_name": "roberthatch",
          "author_url": "",
          "post_date": "08/04/2022 19:26:17",
          "content": "<p>Try Chris's starter, and comment out the \"to_pandas\" line which I recall happens before the CV folds code. If that works, you have a template that might help fix any issue in your code. If it doesn't work, it gives concrete example of the error. :)</p>\n<p><a href=\"https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\" target=\"_blank\">https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1885047,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "08/04/2022 20:13:06",
          "content": "<p>Yeah, it's possible when we use <code>to_pandas</code>, but, if there is any method with Cudf, it will better than using pandas, however I tried other method, it works well, Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1885066,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/04/2022 20:58:26",
          "content": "<p>I'm not sure what you are trying to do BOOBA. The library RAPIDS does have a train test split if that's what you want. It's called </p>\n<pre><code>cuml.model_selection.train_test_split\n</code></pre>\n<p>The docs webpage is <a href=\"https://docs.rapids.ai/api/cuml/stable/api.html#model-selection-and-data-splitting\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1885094,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "08/04/2022 21:43:12",
          "content": "<p>Thanks <strong>Chris Deotte</strong>, I have two ideas, I am working on them, and maybe those will help me to raise my score, the first is about using Target encoding with adding some noises, I know it's the overfitted method, and the second thing, is about using an special CV to use data more efficiently.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1880827": "can I do something similar to this in cudf?\n`df.loc[df['col'] == 10, 'col_is_10'] = 1`",
    "1882136": "Yes, you can do that in cudf. Is it not working?\n\nYou may need to run the line `df['col_is_10'] = 0` before you run that command. I don't think you can use your line to create a new column if it doesn't exist already. (And I don't think you can create new column in Pandas that way either without setting to zero first).",
    "1884872": "What about using `train_test_split` than `ShuffleSplit`, I tried before, but I got the error.",
    "1884990": "Try Chris's starter, and comment out the \"to_pandas\" line which I recall happens before the CV folds code. If that works, you have a template that might help fix any issue in your code. If it doesn't work, it gives concrete example of the error. :)\n\nhttps://www.kaggle.com/code/cdeotte/xgboost-starter-0-793",
    "1885047": "Yeah, it's possible when we use `to_pandas`, but, if there is any method with Cudf, it will better than using pandas, however I tried other method, it works well, Thanks",
    "1885066": "I'm not sure what you are trying to do BOOBA. The library RAPIDS does have a train test split if that's what you want. It's called \n    \n    cuml.model_selection.train_test_split\n\nThe docs webpage is [here][1]\n\n[1]: https://docs.rapids.ai/api/cuml/stable/api.html#model-selection-and-data-splitting",
    "1885094": "Thanks **Chris Deotte**, I have two ideas, I am working on them, and maybe those will help me to raise my score, the first is about using Target encoding with adding some noises, I know it's the overfitted method, and the second thing, is about using an special CV to use data more efficiently."
  },
  "source": "meta"
}