{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img src=\"https://storage.googleapis.com/kaggle-competitions/kaggle/36414/logos/header.png\" width=1500 class=\"center\">\n\n<h1 align=\"center\" style=\"background-color:#C8FFF8\">Create Your Own Dataset!</h1> \n\nCheckout the [dataset here](https://www.kaggle.com/datasets/alejopaullier/guie-toys-dataset).\n\nWelcome to this competition! 👋👋👋\n\nIn this competition you will create a model that extracts a feature embedding for the images and submit the model via Kaggle Notebooks. \n\n**No training data is provided in this competition!** According to the competition's description:\n```\nWe do not provide a training set. In accordance with the rules, any training data may be used, as long as it is disclosed in the forum by the relevant deadline. There exist many public datasets for different object types (e.g., artworks, landmarks, products, etc), and we encourage participants to experiment with them as needed.\n```\nHow can we get data?\n- Search for public datasets in the internet with public licenses.\n- Build your own dataset.\n\n<h1 align=\"center\" style=\"background-color:#C8FFF8\">Let's build a toys dataset!</h1> \n<img src=\"https://www.ripleys.com/wp-content/uploads/2020/05/shutterstock_1375929740-1024x768.jpg\" width=1000 class=\"center\">","metadata":{}},{"cell_type":"markdown","source":"### Install `icrawler`\n\nWe will web scrape Google images and extract images from queries using [icrawler](https://github.com/hellock/icrawler). Run the following cell to install it!","metadata":{}},{"cell_type":"code","source":"pip install icrawler","metadata":{"execution":{"iopub.status.busy":"2022-07-12T15:57:01.841489Z","iopub.execute_input":"2022-07-12T15:57:01.842555Z","iopub.status.idle":"2022-07-12T15:57:14.811810Z","shell.execute_reply.started":"2022-07-12T15:57:01.842453Z","shell.execute_reply":"2022-07-12T15:57:14.810594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Create a folder \n\nLet's create a folder to store our images","metadata":{}},{"cell_type":"code","source":"!mkdir toys","metadata":{"execution":{"iopub.status.busy":"2022-07-12T15:57:14.814252Z","iopub.execute_input":"2022-07-12T15:57:14.814584Z","iopub.status.idle":"2022-07-12T15:57:15.622848Z","shell.execute_reply.started":"2022-07-12T15:57:14.814553Z","shell.execute_reply":"2022-07-12T15:57:15.621337Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Download one image!\n\nRun the following cell. A troll will be downloaded into your `toys` folder.","metadata":{}},{"cell_type":"code","source":"from icrawler.builtin import GoogleImageCrawler\n\ngoogle_crawler = GoogleImageCrawler(storage={'root_dir': 'toys'})\ngoogle_crawler.crawl(keyword='troll toy', max_num=1)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-12T15:57:15.625577Z","iopub.execute_input":"2022-07-12T15:57:15.626086Z","iopub.status.idle":"2022-07-12T15:57:20.798346Z","shell.execute_reply.started":"2022-07-12T15:57:15.626019Z","shell.execute_reply":"2022-07-12T15:57:20.797400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Load CSV data\n\nWe will be downloading the most famous toys from several generations (from 1980 to 2010). These toy names were extracted mostly from Wikipedia:\n- [1970s toys](https://en.wikipedia.org/wiki/Category:1970s_toys)\n- [1980s toys](https://en.wikipedia.org/wiki/Category:1980s_toys)\n- [1990s toys](https://en.wikipedia.org/wiki/Category:1990s_toys)\n- [2000s toys](https://en.wikipedia.org/wiki/Category:2000s_toys)\n- [2010s toys](https://en.wikipedia.org/wiki/Category:2010s_toys)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ndf = pd.read_csv(\"../input/toys-dataset/toys_dataset.csv\", sep=',')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-12T15:57:20.800006Z","iopub.execute_input":"2022-07-12T15:57:20.800728Z","iopub.status.idle":"2022-07-12T15:57:20.843949Z","shell.execute_reply.started":"2022-07-12T15:57:20.800684Z","shell.execute_reply":"2022-07-12T15:57:20.842805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tqdm import tqdm\n\nfor idx, query in enumerate(tqdm(df[\"toy_query\"])):\n    google_crawler = GoogleImageCrawler(storage={'root_dir': 'toys/' + query})\n    google_crawler.crawl(keyword=query, max_num=1)","metadata":{"execution":{"iopub.status.busy":"2022-07-12T15:57:20.845465Z","iopub.execute_input":"2022-07-12T15:57:20.846017Z","iopub.status.idle":"2022-07-12T15:57:48.930744Z","shell.execute_reply.started":"2022-07-12T15:57:20.845972Z","shell.execute_reply":"2022-07-12T15:57:48.929770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 align=\"center\" style=\"background-color:#C8FFF8\">We have built our dataset!</h1> \n\nYou can try adding more images by modifying the `max_num` parameter (which now limits images to 1) to extract more images.","metadata":{}}]}