{
  "id": 363617,
  "title": "Processed dataset",
  "url": "/competitions/otto-recommender-system/discussion/363617",
  "author_name": "",
  "post_date": "2022-11-02T13:16:44.914423Z",
  "votes": 30,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Because we all love pandas, here's the training dataset in .csv and .parquet</p>\n<p><a href=\"https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\" target=\"_blank\">https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe</a></p>\n<p>Enjoy :-) </p>",
  "messages": [
    {
      "id": "2014266",
      "postDate": "11/02/2022 13:16:44",
      "content": "<p>Because we all love pandas, here's the training dataset in .csv and .parquet</p>\n<p><a href=\"https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\" target=\"_blank\">https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe</a></p>\n<p>Enjoy :-) </p>",
      "rawMarkdown": "Because we all love pandas, here's the training dataset in .csv and .parquet\n\nhttps://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\n\nEnjoy :-)",
      "votes": null
    },
    {
      "id": "2014273",
      "postDate": "11/02/2022 13:25:48",
      "content": "<p>Thanks! Could you share also the code?</p>",
      "rawMarkdown": "Thanks! Could you share also the code?",
      "votes": null
    },
    {
      "id": "2014287",
      "postDate": "11/02/2022 13:40:27",
      "content": "<p>Here is another dataset in pickle format<br>\n<a href=\"https://www.kaggle.com/datasets/gunesevitan/otto-mors-pickled-data\" target=\"_blank\">https://www.kaggle.com/datasets/gunesevitan/otto-mors-pickled-data</a></p>\n<p>and the code<br>\n<a href=\"https://www.kaggle.com/code/gunesevitan/otto-multi-objective-recommender-system-pickle\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/otto-multi-objective-recommender-system-pickle</a></p>",
      "rawMarkdown": "Here is another dataset in pickle format\nhttps://www.kaggle.com/datasets/gunesevitan/otto-mors-pickled-data\n\nand the code\nhttps://www.kaggle.com/code/gunesevitan/otto-multi-objective-recommender-system-pickle",
      "votes": null
    },
    {
      "id": "2014292",
      "postDate": "11/02/2022 13:42:47",
      "content": "<p><a href=\"https://www.kaggle.com/konradb/dataset-as-df\" target=\"_blank\">https://www.kaggle.com/konradb/dataset-as-df</a></p>",
      "rawMarkdown": "https://www.kaggle.com/konradb/dataset-as-df",
      "votes": null
    },
    {
      "id": "2014578",
      "postDate": "11/02/2022 16:57:52",
      "content": "<p>thanks! I was struggling a little with the JSON files</p>",
      "rawMarkdown": "thanks! I was struggling a little with the JSON files",
      "votes": null
    },
    {
      "id": "2015497",
      "postDate": "11/03/2022 10:18:44",
      "content": "<p>Thanks for the parquet/CSV!<br>\n(I have no idea why Kaggle doesn't just provide it as such. You can sample from those easily enough)</p>",
      "rawMarkdown": "Thanks for the parquet/CSV!\n(I have no idea why Kaggle doesn't just provide it as such. You can sample from those easily enough)",
      "votes": null
    },
    {
      "id": "2018135",
      "postDate": "11/05/2022 12:38:35",
      "content": "<p>Hello, just a friendly warning that it's <strong>not the entire dataset</strong>. It's about 1,6% of the train data sessions (~2/129 chunks) and about 11,8% (~2/17 chunks) of the test data sessions. </p>\n<p>I've created a full CSV dataset based on Konrad Banachewicz's and Gunes Evitan's notebooks.</p>\n<p>Dataset: <a href=\"https://www.kaggle.com/datasets/adamnarozniak/full-otto-dataset-in-csv\" target=\"_blank\">https://www.kaggle.com/datasets/adamnarozniak/full-otto-dataset-in-csv</a><br>\nCode: <a href=\"https://www.kaggle.com/code/adamnarozniak/full-otto-dataset-in-csv-for-pandas\" target=\"_blank\">https://www.kaggle.com/code/adamnarozniak/full-otto-dataset-in-csv-for-pandas</a></p>",
      "rawMarkdown": "Hello, just a friendly warning that it's **not the entire dataset**. It's about 1,6% of the train data sessions (~2/129 chunks) and about 11,8% (~2/17 chunks) of the test data sessions. \n\nI've created a full CSV dataset based on Konrad Banachewicz's and Gunes Evitan's notebooks.\n\nDataset: https://www.kaggle.com/datasets/adamnarozniak/full-otto-dataset-in-csv\nCode: https://www.kaggle.com/code/adamnarozniak/full-otto-dataset-in-csv-for-pandas",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2014273,
      "author_name": "enric1296",
      "author_url": "",
      "post_date": "11/02/2022 13:25:48",
      "content": "<p>Thanks! Could you share also the code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2014292,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "11/02/2022 13:42:47",
          "content": "<p><a href=\"https://www.kaggle.com/konradb/dataset-as-df\" target=\"_blank\">https://www.kaggle.com/konradb/dataset-as-df</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2014287,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "11/02/2022 13:40:27",
      "content": "<p>Here is another dataset in pickle format<br>\n<a href=\"https://www.kaggle.com/datasets/gunesevitan/otto-mors-pickled-data\" target=\"_blank\">https://www.kaggle.com/datasets/gunesevitan/otto-mors-pickled-data</a></p>\n<p>and the code<br>\n<a href=\"https://www.kaggle.com/code/gunesevitan/otto-multi-objective-recommender-system-pickle\" target=\"_blank\">https://www.kaggle.com/code/gunesevitan/otto-multi-objective-recommender-system-pickle</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2014578,
      "author_name": "mrgabrielblins",
      "author_url": "",
      "post_date": "11/02/2022 16:57:52",
      "content": "<p>thanks! I was struggling a little with the JSON files</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2015497,
      "author_name": "danofer",
      "author_url": "",
      "post_date": "11/03/2022 10:18:44",
      "content": "<p>Thanks for the parquet/CSV!<br>\n(I have no idea why Kaggle doesn't just provide it as such. You can sample from those easily enough)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2018135,
      "author_name": "adamnarozniak",
      "author_url": "",
      "post_date": "11/05/2022 12:38:35",
      "content": "<p>Hello, just a friendly warning that it's <strong>not the entire dataset</strong>. It's about 1,6% of the train data sessions (~2/129 chunks) and about 11,8% (~2/17 chunks) of the test data sessions. </p>\n<p>I've created a full CSV dataset based on Konrad Banachewicz's and Gunes Evitan's notebooks.</p>\n<p>Dataset: <a href=\"https://www.kaggle.com/datasets/adamnarozniak/full-otto-dataset-in-csv\" target=\"_blank\">https://www.kaggle.com/datasets/adamnarozniak/full-otto-dataset-in-csv</a><br>\nCode: <a href=\"https://www.kaggle.com/code/adamnarozniak/full-otto-dataset-in-csv-for-pandas\" target=\"_blank\">https://www.kaggle.com/code/adamnarozniak/full-otto-dataset-in-csv-for-pandas</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2014266": "Because we all love pandas, here's the training dataset in .csv and .parquet\n\nhttps://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\n\nEnjoy :-)",
    "2014273": "Thanks! Could you share also the code?",
    "2014287": "Here is another dataset in pickle format\nhttps://www.kaggle.com/datasets/gunesevitan/otto-mors-pickled-data\n\nand the code\nhttps://www.kaggle.com/code/gunesevitan/otto-multi-objective-recommender-system-pickle",
    "2014292": "https://www.kaggle.com/konradb/dataset-as-df",
    "2014578": "thanks! I was struggling a little with the JSON files",
    "2015497": "Thanks for the parquet/CSV!\n(I have no idea why Kaggle doesn't just provide it as such. You can sample from those easily enough)",
    "2018135": "Hello, just a friendly warning that it's **not the entire dataset**. It's about 1,6% of the train data sessions (~2/129 chunks) and about 11,8% (~2/17 chunks) of the test data sessions. \n\nI've created a full CSV dataset based on Konrad Banachewicz's and Gunes Evitan's notebooks.\n\nDataset: https://www.kaggle.com/datasets/adamnarozniak/full-otto-dataset-in-csv\nCode: https://www.kaggle.com/code/adamnarozniak/full-otto-dataset-in-csv-for-pandas"
  },
  "source": "meta"
}