{
  "id": 405958,
  "title": "How to create a Kaggle Dataset from the Kaggle working directory",
  "url": "/competitions/asl-signs/discussion/405958",
  "author_name": "Henry Javier",
  "post_date": "2023-04-30T06:32:38.292000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello friends of the community, I need to create a dataset that allows me to take the data directly from ./ of Kaggle, and in this way I do not have to download it to my machine (because the files are very large). What should I do to achieve it?. </p>",
  "messages": [
    {
      "id": 2240017,
      "postDate": "2023-04-30T06:32:38.293Z",
      "content": "<p>Hello friends of the community, I need to create a dataset that allows me to take the data directly from ./ of Kaggle, and in this way I do not have to download it to my machine (because the files are very large). What should I do to achieve it?. </p>",
      "rawMarkdown": "Hello friends of the community, I need to create a dataset that allows me to take the data directly from ./ of Kaggle, and in this way I do not have to download it to my machine (because the files are very large). What should I do to achieve it?. ",
      "votes": 4
    },
    {
      "id": 2240673,
      "postDate": "2023-04-30T18:53:36.363Z",
      "content": "<p>Great work from you!</p>",
      "rawMarkdown": "Great work from you!",
      "votes": 2,
      "replies": [
        {
          "id": 2240712,
          "postDate": "2023-04-30T19:23:07.707Z",
          "content": "<p>Thank you, Hassan…</p>",
          "rawMarkdown": "Thank you, Hassan..."
        }
      ]
    },
    {
      "id": 2240121,
      "postDate": "2023-04-30T08:43:50.517Z",
      "content": "<p>You could use a kaggle dataset API:</p>\n<pre><code>### Create Kaggle Dataset if not exists \nDATASET_NAME = 'asl-tensorflow-record-dataset'\n\n!rm -rf /tmp/{DATASET_NAME}\n\nshutil.rmtree(f'/tmp/{DATASET_NAME}', ignore_errors=True)\nos.makedirs(f'/tmp/{DATASET_NAME}', exist_ok=True)\n\nwith open('../input/my-secret/kaggle.json') as f:\n    kaggle_creds = json.load(f)\n\nos.environ['KAGGLE_USERNAME'] = kaggle_creds['username']\nos.environ['KAGGLE_KEY'] = kaggle_creds['key']\n\n!kaggle datasets init -p /tmp/{DATASET_NAME}\n\nwith open(f'/tmp/{DATASET_NAME}/dataset-metadata.json') as f:\n    dataset_meta = json.load(f)\n\ndataset_meta['id'] = f'meowmeowmeowmeowmeow/{DATASET_NAME}'\ndataset_meta['title'] = DATASET_NAME\ndataset_meta['note'] = 'Stratified K-fold by clsses and keep one unqiue participant in validation set'\n\nwith open(f'/tmp/{DATASET_NAME}/dataset-metadata.json', \"w\") as outfile:\n    json.dump(dataset_meta, outfile)\n#print(dataset_meta)\n\n!cp /tmp/{DATASET_NAME}/dataset-metadata.json /tmp/{DATASET_NAME}/meta.json\n!ls /tmp/{DATASET_NAME}\n\n\n!kaggle datasets version -m {version_name} -p /tmp/{DATASET_NAME} -r zip\n</code></pre>\n<p>Note, you need to upload your kaggle secrets as a dataset and link it to the notebook in order to get access to your datasets from the code</p>",
      "rawMarkdown": "You could use a kaggle dataset API:\n\n```\n### Create Kaggle Dataset if not exists \nDATASET_NAME = 'asl-tensorflow-record-dataset'\n\n!rm -rf /tmp/{DATASET_NAME}\n\nshutil.rmtree(f'/tmp/{DATASET_NAME}', ignore_errors=True)\nos.makedirs(f'/tmp/{DATASET_NAME}', exist_ok=True)\n\nwith open('../input/my-secret/kaggle.json') as f:\n    kaggle_creds = json.load(f)\n    \nos.environ['KAGGLE_USERNAME'] = kaggle_creds['username']\nos.environ['KAGGLE_KEY'] = kaggle_creds['key']\n\n!kaggle datasets init -p /tmp/{DATASET_NAME}\n\nwith open(f'/tmp/{DATASET_NAME}/dataset-metadata.json') as f:\n    dataset_meta = json.load(f)\n\ndataset_meta['id'] = f'meowmeowmeowmeowmeow/{DATASET_NAME}'\ndataset_meta['title'] = DATASET_NAME\ndataset_meta['note'] = 'Stratified K-fold by clsses and keep one unqiue participant in validation set'\n\nwith open(f'/tmp/{DATASET_NAME}/dataset-metadata.json', \"w\") as outfile:\n    json.dump(dataset_meta, outfile)\n#print(dataset_meta)\n\n!cp /tmp/{DATASET_NAME}/dataset-metadata.json /tmp/{DATASET_NAME}/meta.json\n!ls /tmp/{DATASET_NAME}\n\n\n!kaggle datasets version -m {version_name} -p /tmp/{DATASET_NAME} -r zip\n```\n\nNote, you need to upload your kaggle secrets as a dataset and link it to the notebook in order to get access to your datasets from the code",
      "votes": 2,
      "replies": [
        {
          "id": 2240311,
          "postDate": "2023-04-30T12:31:15.933Z",
          "content": "<p>Thank you very much, Mykola, for your invaluable technical help. I'm sure this will help the whole community a lot.</p>",
          "rawMarkdown": "Thank you very much, Mykola, for your invaluable technical help. I'm sure this will help the whole community a lot.",
          "votes": 1,
          "replies": [
            {
              "id": 2240480,
              "postDate": "2023-04-30T15:14:06.950Z",
              "content": "<p>Btw, I've seen people using a notebook output as an input in another notebook. I don't know how it works, but maybe this also would help you</p>",
              "rawMarkdown": "Btw, I've seen people using a notebook output as an input in another notebook. I don't know how it works, but maybe this also would help you"
            },
            {
              "id": 2243079,
              "postDate": "2023-05-02T17:10:46.777Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2243082,
          "postDate": "2023-05-02T17:11:26.190Z",
          "content": "<p><a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a> thank you for the explanation!👍</p>",
          "rawMarkdown": "@meowmeowmeowmeowmeow thank you for the explanation!👍"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2240673,
      "author_name": "Hassan Ul Haq",
      "author_url": "",
      "post_date": "2023-04-30T18:53:36.363000",
      "content": "<p>Great work from you!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2240712,
          "author_name": "Henry Javier",
          "author_url": "",
          "post_date": "2023-04-30T19:23:07.707000",
          "content": "<p>Thank you, Hassan…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2240121,
      "author_name": "Mykola",
      "author_url": "",
      "post_date": "2023-04-30T08:43:50.517000",
      "content": "<p>You could use a kaggle dataset API:</p>\n<pre><code>### Create Kaggle Dataset if not exists \nDATASET_NAME = 'asl-tensorflow-record-dataset'\n\n!rm -rf /tmp/{DATASET_NAME}\n\nshutil.rmtree(f'/tmp/{DATASET_NAME}', ignore_errors=True)\nos.makedirs(f'/tmp/{DATASET_NAME}', exist_ok=True)\n\nwith open('../input/my-secret/kaggle.json') as f:\n    kaggle_creds = json.load(f)\n\nos.environ['KAGGLE_USERNAME'] = kaggle_creds['username']\nos.environ['KAGGLE_KEY'] = kaggle_creds['key']\n\n!kaggle datasets init -p /tmp/{DATASET_NAME}\n\nwith open(f'/tmp/{DATASET_NAME}/dataset-metadata.json') as f:\n    dataset_meta = json.load(f)\n\ndataset_meta['id'] = f'meowmeowmeowmeowmeow/{DATASET_NAME}'\ndataset_meta['title'] = DATASET_NAME\ndataset_meta['note'] = 'Stratified K-fold by clsses and keep one unqiue participant in validation set'\n\nwith open(f'/tmp/{DATASET_NAME}/dataset-metadata.json', \"w\") as outfile:\n    json.dump(dataset_meta, outfile)\n#print(dataset_meta)\n\n!cp /tmp/{DATASET_NAME}/dataset-metadata.json /tmp/{DATASET_NAME}/meta.json\n!ls /tmp/{DATASET_NAME}\n\n\n!kaggle datasets version -m {version_name} -p /tmp/{DATASET_NAME} -r zip\n</code></pre>\n<p>Note, you need to upload your kaggle secrets as a dataset and link it to the notebook in order to get access to your datasets from the code</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2240311,
          "author_name": "Henry Javier",
          "author_url": "",
          "post_date": "2023-04-30T12:31:15.933000",
          "content": "<p>Thank you very much, Mykola, for your invaluable technical help. I'm sure this will help the whole community a lot.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2240480,
              "author_name": "Mykola",
              "author_url": "",
              "post_date": "2023-04-30T15:14:06.950000",
              "content": "<p>Btw, I've seen people using a notebook output as an input in another notebook. I don't know how it works, but maybe this also would help you</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2243079,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-05-02T17:10:46.777000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2243082,
          "author_name": "Ivan Isaev",
          "author_url": "",
          "post_date": "2023-05-02T17:11:26.190000",
          "content": "<p><a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a> thank you for the explanation!👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2240017": "Hello friends of the community, I need to create a dataset that allows me to take the data directly from ./ of Kaggle, and in this way I do not have to download it to my machine (because the files are very large). What should I do to achieve it?. ",
    "2240673": "Great work from you!",
    "2240121": "You could use a kaggle dataset API:\n\n```\n### Create Kaggle Dataset if not exists \nDATASET_NAME = 'asl-tensorflow-record-dataset'\n\n!rm -rf /tmp/{DATASET_NAME}\n\nshutil.rmtree(f'/tmp/{DATASET_NAME}', ignore_errors=True)\nos.makedirs(f'/tmp/{DATASET_NAME}', exist_ok=True)\n\nwith open('../input/my-secret/kaggle.json') as f:\n    kaggle_creds = json.load(f)\n    \nos.environ['KAGGLE_USERNAME'] = kaggle_creds['username']\nos.environ['KAGGLE_KEY'] = kaggle_creds['key']\n\n!kaggle datasets init -p /tmp/{DATASET_NAME}\n\nwith open(f'/tmp/{DATASET_NAME}/dataset-metadata.json') as f:\n    dataset_meta = json.load(f)\n\ndataset_meta['id'] = f'meowmeowmeowmeowmeow/{DATASET_NAME}'\ndataset_meta['title'] = DATASET_NAME\ndataset_meta['note'] = 'Stratified K-fold by clsses and keep one unqiue participant in validation set'\n\nwith open(f'/tmp/{DATASET_NAME}/dataset-metadata.json', \"w\") as outfile:\n    json.dump(dataset_meta, outfile)\n#print(dataset_meta)\n\n!cp /tmp/{DATASET_NAME}/dataset-metadata.json /tmp/{DATASET_NAME}/meta.json\n!ls /tmp/{DATASET_NAME}\n\n\n!kaggle datasets version -m {version_name} -p /tmp/{DATASET_NAME} -r zip\n```\n\nNote, you need to upload your kaggle secrets as a dataset and link it to the notebook in order to get access to your datasets from the code"
  }
}