{
  "id": 329647,
  "title": "Need help: How to run in Colab  🙏",
  "url": "/competitions/amex-default-prediction/discussion/329647",
  "author_name": "",
  "post_date": "2022-06-07T21:50:11.297954800Z",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I have used up my weekly GPU quata since beginning of the week. 😯<br>\nI plan to imgrate my code to Colab.<br>\nIs anyone retrieving raw data, preparing feature and trainning model on Colab? <br>\nIt seems cupy is not available on all colab environment. <br>\nIs anyone running on Colab?  Could you give some quidance / hints? </p>\n<p>Many thanks in advance. 🙏</p>",
  "messages": [
    {
      "id": "1814428",
      "postDate": "06/07/2022 21:50:11",
      "content": "<p>I have used up my weekly GPU quata since beginning of the week. 😯<br>\nI plan to imgrate my code to Colab.<br>\nIs anyone retrieving raw data, preparing feature and trainning model on Colab? <br>\nIt seems cupy is not available on all colab environment. <br>\nIs anyone running on Colab?  Could you give some quidance / hints? </p>\n<p>Many thanks in advance. 🙏</p>",
      "rawMarkdown": "I have used up my weekly GPU quata since beginning of the week. 😯\nI plan to imgrate my code to Colab.\nIs anyone retrieving raw data, preparing feature and trainning model on Colab? \nIt seems cupy is not available on all colab environment. \nIs anyone running on Colab?  Could you give some quidance / hints? \n\nMany thanks in advance. 🙏",
      "votes": null
    },
    {
      "id": "1814487",
      "postDate": "06/08/2022 01:14:47",
      "content": "<p>I think this is where I'll be by tomorrow 🤣</p>",
      "rawMarkdown": "I think this is where I'll be by tomorrow 🤣",
      "votes": null
    },
    {
      "id": "1814542",
      "postDate": "06/08/2022 03:15:35",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/jingwora1\" target=\"_blank\">@jingwora1</a>, yes, I'm running Colab and a local environment; what are you looking for?<br>\nColab it's pretty straightforward. You need to connect to the Kaggle API, and it will be the same process as Kaggle. Once done training and estimating, you can upload the submission file to Kaggle. Let me know if you need an example. I can post some code that can get you started</p>",
      "rawMarkdown": "Hello @jingwora1, yes, I'm running Colab and a local environment; what are you looking for?\nColab it's pretty straightforward. You need to connect to the Kaggle API, and it will be the same process as Kaggle. Once done training and estimating, you can upload the submission file to Kaggle. Let me know if you need an example. I can post some code that can get you started",
      "votes": null
    },
    {
      "id": "1814816",
      "postDate": "06/08/2022 11:03:05",
      "content": "<p>Thank you for offering some helps. I have 2 problems right now.<br>\n1) I use kaggle.jason to login and load feather with pd.read_feather funcion. There is an error about credentail during loading. I used to do this way to load tfrecord and it is OK. Could you share how do you load your data? </p>\n<p>I load from this dataset.<br>\n<a href=\"https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction\" target=\"_blank\">https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction</a><br>\n(Ruchi Bhatia)</p>\n<p>2) I am working on feature engineering 2,000+ feature. I take a long time to process. Do you know how to speed up the process? Like chaging file format or using joblib?</p>\n<p>Many thanks in advance. 🙏</p>",
      "rawMarkdown": "Thank you for offering some helps. I have 2 problems right now.\n1) I use kaggle.jason to login and load feather with pd.read_feather funcion. There is an error about credentail during loading. I used to do this way to load tfrecord and it is OK. Could you share how do you load your data? \n\nI load from this dataset.\nhttps://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction\n(Ruchi Bhatia)\n\n2) I am working on feature engineering 2,000+ feature. I take a long time to process. Do you know how to speed up the process? Like chaging file format or using joblib?\n\nMany thanks in advance. 🙏",
      "votes": null
    },
    {
      "id": "1815629",
      "postDate": "06/09/2022 08:40:48",
      "content": "<p><a href=\"https://www.kaggle.com/cv13j0\" target=\"_blank\">@cv13j0</a> Do you use <code>cudf</code> on Colab? How did you set it up?</p>",
      "rawMarkdown": "cv13j0 Do you use `cudf` on Colab? How did you set it up?",
      "votes": null
    },
    {
      "id": "1815848",
      "postDate": "06/09/2022 14:41:18",
      "content": "<p>I haven't worked on this competition yet, but you could run the data cleaning scripts in a Kaggle notebook and then copy the results to your Google Drive or a Kaggle Dataset and then transfer that to colab.  That only needs to be done once, and it'll save a lot of space!</p>",
      "rawMarkdown": "I haven't worked on this competition yet, but you could run the data cleaning scripts in a Kaggle notebook and then copy the results to your Google Drive or a Kaggle Dataset and then transfer that to colab.  That only needs to be done once, and it'll save a lot of space!",
      "votes": null
    },
    {
      "id": "1816253",
      "postDate": "06/10/2022 03:10:22",
      "content": "<p>Hello, I haven't used cudf; I will try to run it on the example I will provide</p>",
      "rawMarkdown": "Hello, I haven't used cudf; I will try to run it on the example I will provide",
      "votes": null
    },
    {
      "id": "1816259",
      "postDate": "06/10/2022 03:28:29",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/jingwora1\" target=\"_blank\">@jingwora1</a>, I use the Kaggle API. and I download the data from Kaggle directly and perform all of the operations I need; I will share soon a link to my Colab Notebooks</p>",
      "rawMarkdown": "Hello @jingwora1, I use the Kaggle API. and I download the data from Kaggle directly and perform all of the operations I need; I will share soon a link to my Colab Notebooks",
      "votes": null
    },
    {
      "id": "1817549",
      "postDate": "06/11/2022 14:30:52",
      "content": "<p><a href=\"https://www.kaggle.com/itacdonev\" target=\"_blank\">@itacdonev</a> <br>\nI cannot use cudf on Colab. By the way, I read data by pd.read_feather.</p>",
      "rawMarkdown": "itacdonev \nI cannot use cudf on Colab. By the way, I read data by pd.read_feather.",
      "votes": null
    },
    {
      "id": "1817556",
      "postDate": "06/11/2022 14:36:00",
      "content": "<p><a href=\"https://www.kaggle.com/cv13j0\" target=\"_blank\">@cv13j0</a>  Thank you! <br>\nHere is my solution. I copy feather data to google drive. <br>\nThen I mount the drive and copy files to colab. <br>\nI read data by pd.read_feather.<br>\nI think it is not efficient way yet it work for me so far.</p>",
      "rawMarkdown": "cv13j0  Thank you! \nHere is my solution. I copy feather data to google drive. \nThen I mount the drive and copy files to colab. \nI read data by pd.read_feather.\nI think it is not efficient way yet it work for me so far.",
      "votes": null
    },
    {
      "id": "1818000",
      "postDate": "06/12/2022 05:54:52",
      "content": "<p>For me, I use a Kaggle Notebook to first preprocess the files(so that I don't have to do them each time I use colab…this reduces time and memory consumption) and then I download the files from Kaggle Notebook to Colab using the API.  <br>\nHere is my code -</p>\n<pre><code>!pip install -q kaggle\nfrom google.colab import files\nfiles.upload()\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json\n!kaggle kernels output susnato/amex-data-preprocesing-feature-engineering -p '/content/Data/'\n</code></pre>",
      "rawMarkdown": "For me, I use a Kaggle Notebook to first preprocess the files(so that I don't have to do them each time I use colab...this reduces time and memory consumption) and then I download the files from Kaggle Notebook to Colab using the API.  \nHere is my code -\n```\n\n!pip install -q kaggle\nfrom google.colab import files\nfiles.upload()\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json\n!kaggle kernels output susnato/amex-data-preprocesing-feature-engineering -p '/content/Data/'\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1814487,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "06/08/2022 01:14:47",
      "content": "<p>I think this is where I'll be by tomorrow 🤣</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1814542,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "06/08/2022 03:15:35",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/jingwora1\" target=\"_blank\">@jingwora1</a>, yes, I'm running Colab and a local environment; what are you looking for?<br>\nColab it's pretty straightforward. You need to connect to the Kaggle API, and it will be the same process as Kaggle. Once done training and estimating, you can upload the submission file to Kaggle. Let me know if you need an example. I can post some code that can get you started</p>",
      "votes": null,
      "replies": [
        {
          "id": 1814816,
          "author_name": "jingwora1",
          "author_url": "",
          "post_date": "06/08/2022 11:03:05",
          "content": "<p>Thank you for offering some helps. I have 2 problems right now.<br>\n1) I use kaggle.jason to login and load feather with pd.read_feather funcion. There is an error about credentail during loading. I used to do this way to load tfrecord and it is OK. Could you share how do you load your data? </p>\n<p>I load from this dataset.<br>\n<a href=\"https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction\" target=\"_blank\">https://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction</a><br>\n(Ruchi Bhatia)</p>\n<p>2) I am working on feature engineering 2,000+ feature. I take a long time to process. Do you know how to speed up the process? Like chaging file format or using joblib?</p>\n<p>Many thanks in advance. 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1815629,
          "author_name": "itacdonev",
          "author_url": "",
          "post_date": "06/09/2022 08:40:48",
          "content": "<p><a href=\"https://www.kaggle.com/cv13j0\" target=\"_blank\">@cv13j0</a> Do you use <code>cudf</code> on Colab? How did you set it up?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1816253,
          "author_name": "cv13j0",
          "author_url": "",
          "post_date": "06/10/2022 03:10:22",
          "content": "<p>Hello, I haven't used cudf; I will try to run it on the example I will provide</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1816259,
          "author_name": "cv13j0",
          "author_url": "",
          "post_date": "06/10/2022 03:28:29",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/jingwora1\" target=\"_blank\">@jingwora1</a>, I use the Kaggle API. and I download the data from Kaggle directly and perform all of the operations I need; I will share soon a link to my Colab Notebooks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1817549,
          "author_name": "jingwora1",
          "author_url": "",
          "post_date": "06/11/2022 14:30:52",
          "content": "<p><a href=\"https://www.kaggle.com/itacdonev\" target=\"_blank\">@itacdonev</a> <br>\nI cannot use cudf on Colab. By the way, I read data by pd.read_feather.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1817556,
          "author_name": "jingwora1",
          "author_url": "",
          "post_date": "06/11/2022 14:36:00",
          "content": "<p><a href=\"https://www.kaggle.com/cv13j0\" target=\"_blank\">@cv13j0</a>  Thank you! <br>\nHere is my solution. I copy feather data to google drive. <br>\nThen I mount the drive and copy files to colab. <br>\nI read data by pd.read_feather.<br>\nI think it is not efficient way yet it work for me so far.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1815848,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "06/09/2022 14:41:18",
      "content": "<p>I haven't worked on this competition yet, but you could run the data cleaning scripts in a Kaggle notebook and then copy the results to your Google Drive or a Kaggle Dataset and then transfer that to colab.  That only needs to be done once, and it'll save a lot of space!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1818000,
      "author_name": "susnato",
      "author_url": "",
      "post_date": "06/12/2022 05:54:52",
      "content": "<p>For me, I use a Kaggle Notebook to first preprocess the files(so that I don't have to do them each time I use colab…this reduces time and memory consumption) and then I download the files from Kaggle Notebook to Colab using the API.  <br>\nHere is my code -</p>\n<pre><code>!pip install -q kaggle\nfrom google.colab import files\nfiles.upload()\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json\n!kaggle kernels output susnato/amex-data-preprocesing-feature-engineering -p '/content/Data/'\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1814428": "I have used up my weekly GPU quata since beginning of the week. 😯\nI plan to imgrate my code to Colab.\nIs anyone retrieving raw data, preparing feature and trainning model on Colab? \nIt seems cupy is not available on all colab environment. \nIs anyone running on Colab?  Could you give some quidance / hints? \n\nMany thanks in advance. 🙏",
    "1814487": "I think this is where I'll be by tomorrow 🤣",
    "1814542": "Hello @jingwora1, yes, I'm running Colab and a local environment; what are you looking for?\nColab it's pretty straightforward. You need to connect to the Kaggle API, and it will be the same process as Kaggle. Once done training and estimating, you can upload the submission file to Kaggle. Let me know if you need an example. I can post some code that can get you started",
    "1814816": "Thank you for offering some helps. I have 2 problems right now.\n1) I use kaggle.jason to login and load feather with pd.read_feather funcion. There is an error about credentail during loading. I used to do this way to load tfrecord and it is OK. Could you share how do you load your data? \n\nI load from this dataset.\nhttps://www.kaggle.com/datasets/ruchi798/parquet-files-amexdefault-prediction\n(Ruchi Bhatia)\n\n2) I am working on feature engineering 2,000+ feature. I take a long time to process. Do you know how to speed up the process? Like chaging file format or using joblib?\n\nMany thanks in advance. 🙏",
    "1815629": "cv13j0 Do you use `cudf` on Colab? How did you set it up?",
    "1815848": "I haven't worked on this competition yet, but you could run the data cleaning scripts in a Kaggle notebook and then copy the results to your Google Drive or a Kaggle Dataset and then transfer that to colab.  That only needs to be done once, and it'll save a lot of space!",
    "1816253": "Hello, I haven't used cudf; I will try to run it on the example I will provide",
    "1816259": "Hello @jingwora1, I use the Kaggle API. and I download the data from Kaggle directly and perform all of the operations I need; I will share soon a link to my Colab Notebooks",
    "1817549": "itacdonev \nI cannot use cudf on Colab. By the way, I read data by pd.read_feather.",
    "1817556": "cv13j0  Thank you! \nHere is my solution. I copy feather data to google drive. \nThen I mount the drive and copy files to colab. \nI read data by pd.read_feather.\nI think it is not efficient way yet it work for me so far.",
    "1818000": "For me, I use a Kaggle Notebook to first preprocess the files(so that I don't have to do them each time I use colab...this reduces time and memory consumption) and then I download the files from Kaggle Notebook to Colab using the API.  \nHere is my code -\n```\n\n!pip install -q kaggle\nfrom google.colab import files\nfiles.upload()\n!mkdir ~/.kaggle\n!cp kaggle.json ~/.kaggle/\n!chmod 600 ~/.kaggle/kaggle.json\n!kaggle kernels output susnato/amex-data-preprocesing-feature-engineering -p '/content/Data/'\n```"
  },
  "source": "meta"
}