{"cells":[{"metadata":{"_uuid":"b311c024132c361a8b537c91dec69b2657180343"},"cell_type":"markdown","source":"Intro\n\nIn this kernel I parse the input text file in order to use all the available information in the dataset\nI propose the source code and the output result in this kernel, and will bind to the file as soon as they will be available\n   "},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"9bab0847903c8705d00606d0e466e7c34201f544"},"cell_type":"code","source":"import pandas as pd\nimport csv\nimport re\nimport os","execution_count":2,"outputs":[]},{"metadata":{"_uuid":"79cfad2006aff2d9f15f6631cc36581dbdd9b78c"},"cell_type":"markdown","source":"I was not able to access \"list files\" in the dataset environnement (don't know why the following files are not availables)\n* ../input/cvpr-2018-autonomous-driving/test_video_list_and_name_mapping/*\n* ../input/cvpr-2018-autonomous-driving/train_video_list/*"},{"metadata":{"_uuid":"21b6b5103a95f6723abec67df07a11ef7edd6a2c","trusted":true},"cell_type":"code","source":"print(os.listdir('../input/'))\nprint(os.listdir('../input/cvpr-2018-autonomous-driving/'))","execution_count":3,"outputs":[]},{"metadata":{"_uuid":"01f5d804fb1ba8301e7b701ce9ee73771663c3d3"},"cell_type":"markdown","source":"You will find bellow sample of what is genereted from this kernel"},{"metadata":{"trusted":true,"_uuid":"b295e8e73168460a75c64cb5f4b18b1024251021"},"cell_type":"code","source":"pd.read_csv(\"../input/cvpr-database-detail/train_database.csv\").head()","execution_count":4,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5f58c9146395c7706b5f3c0639b6d95676b8a63e"},"cell_type":"code","source":"pd.read_csv(\"../input/cvpr-database-detail/test_database.csv\").head()","execution_count":5,"outputs":[]},{"metadata":{"_uuid":"2e451649d89b842f5784f013a98313da03a3be7b"},"cell_type":"markdown","source":"In the next cell you will find the configuration path used to locate the files"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"5afa2011a007ca7530774cdab2856c22bc858a29"},"cell_type":"code","source":"# Base path where are stored the unzip dataset\ndatabase_path       = \"../inputs/\"\n\n# Images location subpath \ntrain_image_subpath = \"train_color/\"\ntrain_label_subpath = \"train_label/\"\ntest_image_subpath  = \"test/\"\n\n# Text description file sub path\nlist_train_subpath        = \"train_video_list/\"\nlist_test_subpath         = \"test_video_list_and_name_mapping/list_test/\"\nlist_test_mapping_subpath = \"test_video_list_and_name_mapping/list_test_mapping/\"\n\n","execution_count":6,"outputs":[]},{"metadata":{"_uuid":"f1c0200065f2ed763c25a08c40b25252ef44005b"},"cell_type":"markdown","source":"Building an index to associate the original file_name with the md5 signature used to name the test file images"},{"metadata":{"trusted":false,"collapsed":true,"_uuid":"d35b70d5b7552cc8bbfea2f882388f0f346f5ece"},"cell_type":"code","source":"\ndef extractMapping(filename, path):\n    df = pd.read_csv(path + filename ,sep=\"\\t\" , quoting=csv.QUOTE_NONE,header=None, names=[\"md5\", \"image_name\"])\n    return df\n\ntest_md5_files = os.listdir(database_path + list_test_mapping_subpath)\nmd5_info  = pd.concat([extractMapping(filename, database_path + list_test_mapping_subpath) for filename in test_md5_files], ignore_index=True, copy=False)    \nindex_md5 = {row[1][\"image_name\"] : row[1][\"md5\"] for row in md5_info.iterrows() }\n","execution_count":3,"outputs":[]},{"metadata":{"_uuid":"835fd67afb63a00f4296a2334425357a0c3a9bd1"},"cell_type":"markdown","source":"We implement a function to extract usefull information from list provided \nWe will use that function for both : \n- Train files \n- Test files\n\nThe extracted informations will be : \n- Camera ID (5 or 6) \n- Road ID ( the road used for that video, to be used as contextual information)\n- Record ( an id that can be used to regroupe image in a video sequence as well as video ID)\n- Car ID ( an id that can be used to associate to the car used for the video capture)\n- Timestamp (a number that can be used to order the images in a video"},{"metadata":{"trusted":false,"collapsed":true,"_uuid":"6f004f1cbb1302d085d108676a307f559054780d"},"cell_type":"code","source":"\ndef extractImageListInfo(filename, file_path):\n    filepath = file_path + filename\n\n    df = pd.read_csv(filepath ,sep=\"\\t\" , quoting=csv.QUOTE_NONE,header=None,names=[\"image_name\", \"ids_name\"])\n    \n    # We have an ID in the filename, probably the video ID\n    df['VideoID'] = filename.split('_')[4]    \n    \n    # We extract information from file name\n    df['Road'], _, df['Record'], df['Camera'], df['ImageFileName']  = df['image_name'].str.split('\\\\').str\n    df['Road'] = df['Road'].apply(lambda x: re.sub(\"road([0-9]*)_.*\",r'\\1',x))\n    df['Record'] = df['Record'].apply(lambda x: re.sub(\"Record([0-9]*)\",r'\\1',x))\n    df['Camera'] = df['Camera'].apply(lambda x: re.sub(\"Camera([0-9]*)\",r'\\1',x))\n    df['CarID'], df['TimeStamp'],_, _  = df['ImageFileName'].str.split('_').str\n    return df\n\n","execution_count":4,"outputs":[]},{"metadata":{"_uuid":"1b50c5fcbcb50cab5529ca68c9b43848cb2fe4d1"},"cell_type":"markdown","source":"We build an index CSV with all contextual information for training database\nWe build the relative path to access images"},{"metadata":{"trusted":false,"_uuid":"7ffb0820004f777883ef5c2d947706bab8801e64","collapsed":true},"cell_type":"code","source":"train_list_files = os.listdir(database_path + list_train_subpath)\ntrain_info = pd.concat([extractImageListInfo(filename, database_path + list_train_subpath) for filename in train_list_files], ignore_index=True, copy=False)    \n\n_, _, _, _, train_info['IdsFileName']  = train_info['ids_name'].str.split('\\\\').str\ntrain_info['ImageFileName'] = train_info['ImageFileName'].apply(lambda x: train_image_subpath + x)\ntrain_info['IdsFileName'] = train_info['IdsFileName'].apply(lambda x: train_label_subpath + x) \ndel train_info['ids_name']\ndel train_info['image_name']\n\ntrain_info.to_csv(\"train_database.csv\")\ntrain_info.head()\n","execution_count":5,"outputs":[]},{"metadata":{"_uuid":"bd196cfc43f25f14485a648a93d1563409507b3c"},"cell_type":"markdown","source":"We build an index CSV with all contextual information for testing database\nWe build the relative path to access images"},{"metadata":{"trusted":false,"_uuid":"45b575d50dc5d2920fad790e8afb23e8c5988724","collapsed":true},"cell_type":"code","source":"test_list_files = os.listdir(database_path + list_test_subpath)\ntest_info  = pd.concat([extractImageListInfo(filename, database_path + list_test_subpath) for filename in test_list_files], ignore_index=True, copy=False)    \n\ntest_info['ImageFileName'] = test_info['image_name'].apply(lambda x: test_image_subpath + index_md5[x] + \".jpg\")\ndel test_info['image_name']\ndel test_info['ids_name']\n\ntrain_info.to_csv(\"test_database.csv\")\ntest_info.head()","execution_count":6,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"a6b0c7fdbf3f3af9b78d1d51a7f98f68c82813f4"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}