{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"HI, \n\nIn the challenge presented to us we have a lot of data, and that's the issue. \n\nNO, too much data isn't the issue, reading and using it is. We have limited disk space allotted to us in the kernels we use, And we need to make sure we can fit the data we get into this pace. But for this challenge we have a bit too much of data! What do we do? \nWell, there are a lot of ways we can handle this issue. \n\n* The first one is we take a small sample of data, that sounds good but is losing data optimum? NO.\n* Datatable\n* Dask\n\nI will be showing how we can use Datatable for the task at hand and also providing the links to various notebooks and documentation that would help you as well! \n","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nimport gc\nimport cv2\nimport time\nimport requests\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom pathlib import Path\nimport matplotlib.pyplot as plt\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\nsns.set()\nseed = 1234\nimport datatable as dt\nnp.random.seed(seed)\nfrom datatable import dt, fread\n\nfrom datatable import dt, f, by, g, join, sort, update, ifelse\n\n%config IPCompleter.use_jedi = False\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-09-15T05:44:55.436012Z","iopub.execute_input":"2021-09-15T05:44:55.436373Z","iopub.status.idle":"2021-09-15T05:44:55.473217Z","shell.execute_reply.started":"2021-09-15T05:44:55.436340Z","shell.execute_reply":"2021-09-15T05:44:55.472571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Path to the directory where data is stored\ndata_path = Path(\"../input/wikipedia-image-caption/\")\n\n# Selecting only a subset of columns as there\n# many columns that we don't need.\ncolumns_to_select = [\"language\",\n                     \"page_url\",\n                     \"image_url\",\n                     \"caption_title_and_reference_description\"\n                    ]\n\n# Get the list of all tsv files we need to read for training\ntsvs = sorted(list(data_path.glob(\"*.tsv\")))\n\n# Remove test tsv file as we don't need it for now\ntsvs.remove(data_path / \"test.tsv\")\n\nprint(\"Number of TSV files found: \", len(tsvs))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-09-15T05:40:57.102355Z","iopub.execute_input":"2021-09-15T05:40:57.102626Z","iopub.status.idle":"2021-09-15T05:40:57.110074Z","shell.execute_reply.started":"2021-09-15T05:40:57.102598Z","shell.execute_reply":"2021-09-15T05:40:57.109041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# About Datatable\n\n\nIt is a Python package for manipulating 2-dimensional tabular data structures\n\nAs per the documentation :\n> This is a Python package for manipulating 2-dimensional tabular data structures (aka data frames). It is close in spirit to pandas or SFrame; however we put specific emphasis on speed and big data support. As the name suggests, the package is closely related to R's data.table and attempts to mimic its core algorithms and API.\n> \n> Requirements: Python 3.6+ (64 bit) and pip 20.3+.","metadata":{}},{"cell_type":"markdown","source":"# IMPORTING DATATABLE AND READING THE DATA","metadata":{"execution":{"iopub.status.busy":"2021-09-14T18:47:18.667404Z","iopub.execute_input":"2021-09-14T18:47:18.668066Z","iopub.status.idle":"2021-09-14T18:47:18.671801Z","shell.execute_reply.started":"2021-09-14T18:47:18.668017Z","shell.execute_reply":"2021-09-14T18:47:18.671223Z"}}},{"cell_type":"code","source":"# File path\nfilepath = \"/kaggle/input/wikipedia-image-caption/train-00000-of-00005.tsv\"\n\n# Read using pandas now\nstart_time = time.time()\ndf = dt.fread(filepath)\nprint(f\"Time taken to read {filepath} in datatable format: {time.time()-start_time:.2f} seconds\")\nprint(\"\")\n\ndf.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-09-15T05:40:58.913570Z","iopub.execute_input":"2021-09-15T05:40:58.913909Z","iopub.status.idle":"2021-09-15T05:42:06.655735Z","shell.execute_reply.started":"2021-09-15T05:40:58.913877Z","shell.execute_reply":"2021-09-15T05:42:06.654895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Train size:\", df.shape)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-09-15T05:42:06.657199Z","iopub.execute_input":"2021-09-15T05:42:06.657678Z","iopub.status.idle":"2021-09-15T05:42:06.666612Z","shell.execute_reply.started":"2021-09-15T05:42:06.657640Z","shell.execute_reply":"2021-09-15T05:42:06.665740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head(2)","metadata":{"execution":{"iopub.status.busy":"2021-09-15T05:42:06.668084Z","iopub.execute_input":"2021-09-15T05:42:06.668474Z","iopub.status.idle":"2021-09-15T05:42:06.682544Z","shell.execute_reply.started":"2021-09-15T05:42:06.668442Z","shell.execute_reply":"2021-09-15T05:42:06.681351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"displaying few rows\")\ndf[[2, 3, 4], :]","metadata":{"execution":{"iopub.status.busy":"2021-09-15T05:43:47.769712Z","iopub.execute_input":"2021-09-15T05:43:47.770017Z","iopub.status.idle":"2021-09-15T05:43:47.780922Z","shell.execute_reply.started":"2021-09-15T05:43:47.769987Z","shell.execute_reply":"2021-09-15T05:43:47.774877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Subset of a datafame based on a condition on one column\")\ndf[f.language==\"en\", :].head(1)","metadata":{"execution":{"iopub.status.busy":"2021-09-15T05:45:14.074483Z","iopub.execute_input":"2021-09-15T05:45:14.074764Z","iopub.status.idle":"2021-09-15T05:45:14.723183Z","shell.execute_reply.started":"2021-09-15T05:45:14.074735Z","shell.execute_reply":"2021-09-15T05:45:14.722318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Creating a new column\")\ndf['new_col'] = 2\ndf[:, update(new_col=2)]","metadata":{"execution":{"iopub.status.busy":"2021-09-15T05:45:45.405104Z","iopub.execute_input":"2021-09-15T05:45:45.405549Z","iopub.status.idle":"2021-09-15T05:45:45.410802Z","shell.execute_reply.started":"2021-09-15T05:45:45.405513Z","shell.execute_reply":"2021-09-15T05:45:45.409752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Updating column names\")\ndf.names = {\"new_col\": \"updated_new_col\"}","metadata":{"execution":{"iopub.status.busy":"2021-09-15T05:46:12.953821Z","iopub.execute_input":"2021-09-15T05:46:12.956026Z","iopub.status.idle":"2021-09-15T05:46:12.962052Z","shell.execute_reply.started":"2021-09-15T05:46:12.955977Z","shell.execute_reply":"2021-09-15T05:46:12.960953Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del df['updated_new_col']","metadata":{"execution":{"iopub.status.busy":"2021-09-15T05:46:35.175479Z","iopub.execute_input":"2021-09-15T05:46:35.176179Z","iopub.status.idle":"2021-09-15T05:46:35.180467Z","shell.execute_reply.started":"2021-09-15T05:46:35.176127Z","shell.execute_reply":"2021-09-15T05:46:35.179438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.sort('page_changed_recently').head(1)","metadata":{"execution":{"iopub.status.busy":"2021-09-15T05:47:35.891321Z","iopub.execute_input":"2021-09-15T05:47:35.892194Z","iopub.status.idle":"2021-09-15T05:47:35.955424Z","shell.execute_reply.started":"2021-09-15T05:47:35.892154Z","shell.execute_reply":"2021-09-15T05:47:35.954604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You can read the data this way! \nEnjoy exploring the data and also the learning :)\n\n\n for further help please visit : https://datatable.readthedocs.io/en/latest/manual/comparison_with_pandas.html","metadata":{}},{"cell_type":"markdown","source":"# Other Links and Notebooks! \n\n## **Links for datatable:**\n- https://www.kaggle.com/sudalairajkumar/getting-started-with-python-datatable\n- https://www.kaggle.com/santhraul/methods-to-read-large-datasets\n- https://docs.h2o.ai/driverless-ai/latest-stable/docs/userguide/datasets-import.html\n- https://github.com/h2oai/datatable\n\n\n## **For Dask:**\n- **Please use the notebook provided by [NAIN](https://www.kaggle.com/aakashnain/can-we-read-faster)**\n","metadata":{}},{"cell_type":"markdown","source":"# Endnote\n\nIf there are any suggesion for the notebook please comment, that would be helpful. Also please upvote if you liked it! Thank you!!\n\nMy work for this competition \n\n**[Notebook for EDA on the WIKIPEDIA IMAGE AND CAPTION MATCHING](https://www.kaggle.com/udbhavpangotra/eda-wikipedia-image-cap-matching)**\n\n### Some of my other works:\n\n* [TPS- APR](https://www.kaggle.com/udbhavpangotra/tps-apr21-eda-model) \n* [HEART ATTACKS](https://www.kaggle.com/udbhavpangotra/heart-attacks-extensive-eda-and-visualizations) \n* [YOUTUBE DATA EXPLORATION](https://www.kaggle.com/udbhavpangotra/what-do-people-use-youtube-for-in-great-britain)\n* [TPS MAY](https://www.kaggle.com/udbhavpangotra/tps-may-21-extensive-eda-catboost-shap)\n* [COVID-19 DIGITAL LEARNING](https://www.kaggle.com/udbhavpangotra/how-did-covid-19-impact-digital-learning-eda)\n* [TPS - SEPT](https://www.kaggle.com/udbhavpangotra/extensive-eda-baseline-shap)\n\n* [also try this dataset ReliefWeb Crisis Figures Data](https://www.kaggle.com/udbhavpangotra/reliefweb-crisis-figures-data)","metadata":{"execution":{"iopub.status.busy":"2021-09-14T19:07:43.226172Z","iopub.execute_input":"2021-09-14T19:07:43.226486Z","iopub.status.idle":"2021-09-14T19:07:43.236986Z","shell.execute_reply.started":"2021-09-14T19:07:43.226456Z","shell.execute_reply":"2021-09-14T19:07:43.235839Z"}}},{"cell_type":"code","source":"%%html\n<marquee style='width: 90% ;height:70%; color: #45B39D ;'>\n    <b>Do UPVOTE if you like my work, I will be adding some more content to this kernel post understanding the files :) </b></marquee>","metadata":{"execution":{"iopub.status.busy":"2021-09-14T19:10:04.24859Z","iopub.execute_input":"2021-09-14T19:10:04.249041Z","iopub.status.idle":"2021-09-14T19:10:04.256897Z","shell.execute_reply.started":"2021-09-14T19:10:04.249005Z","shell.execute_reply":"2021-09-14T19:10:04.256182Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]}]}