{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":13836,"databundleVersionId":1718836,"sourceType":"competition"}],"dockerImageVersionId":30886,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Cassava Leaf Disease Project","metadata":{}},{"cell_type":"markdown","source":"### By: Kevin Roberts, Nathan Weber and Gabe Tonks","metadata":{}},{"cell_type":"markdown","source":"First we import all necessary libraries:","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames[:10]:\n        print(os.path.join(dirname, filename))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T16:03:21.685624Z","iopub.execute_input":"2025-03-10T16:03:21.686017Z","iopub.status.idle":"2025-03-10T16:03:22.073450Z","shell.execute_reply.started":"2025-03-10T16:03:21.685984Z","shell.execute_reply":"2025-03-10T16:03:22.072324Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Then we read in the train.csv file:","metadata":{}},{"cell_type":"code","source":"input_path = '/kaggle/input/cassava-leaf-disease-classification'\ntrain_csv_filename = 'train.csv'\n\ntrain_data = pd.read_csv(os.path.join(input_path, train_csv_filename))\ntrain_data","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T16:06:14.920043Z","iopub.execute_input":"2025-03-10T16:06:14.920366Z","iopub.status.idle":"2025-03-10T16:06:14.946877Z","shell.execute_reply.started":"2025-03-10T16:06:14.920332Z","shell.execute_reply":"2025-03-10T16:06:14.945869Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Each row of the data set train.csv represents a unique picture of a plant, and the first column is the name of the picture, and the second column is it's Cassava type. There are 5 different types: Cassava Bacterical Blight (CBB), Cassava Brown Streak Disease (CBSD), Cassava Green Mottle (CGM), Cassava Mosaic Disease (CMD), and Healthy. These can be put into a dictionary, identical to the one already given:","metadata":{}},{"cell_type":"code","source":"plant_dict = {\"0\": \"Cassava Bacterial Blight (CBB)\", \n              \"1\": \"Cassava Brown Streak Disease (CBSD)\", \n              \"2\": \"Cassava Green Mottle (CGM)\", \n              \"3\": \"Cassava Mosaic Disease (CMD)\", \n              \"4\": \"Healthy\"}","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T16:06:17.811101Z","iopub.execute_input":"2025-03-10T16:06:17.811444Z","iopub.status.idle":"2025-03-10T16:06:17.815983Z","shell.execute_reply.started":"2025-03-10T16:06:17.811406Z","shell.execute_reply":"2025-03-10T16:06:17.814731Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As a small test, we can get the plant description of say the 1st plant listed in the train.csv:","metadata":{}},{"cell_type":"code","source":"first_plant_name = plant_dict[str(train_data.iloc[0, 1])]\nprint(\"The plant in row 50 is classified as: \" + str(first_plant_name))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T16:06:19.549176Z","iopub.execute_input":"2025-03-10T16:06:19.549497Z","iopub.status.idle":"2025-03-10T16:06:19.555137Z","shell.execute_reply.started":"2025-03-10T16:06:19.549471Z","shell.execute_reply":"2025-03-10T16:06:19.553892Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"For the first dummy submission, we apply a \"most frequent\" approach. We simply predict that given any plant, regardless of how it looks, will be classified as the the plant type that appears most frequently in the train data.","metadata":{}},{"cell_type":"code","source":"most_frequent = train_data.iloc[:, 1].mode()[0] # the most frequently classified plant in the train.csv\nprint(\"The most frequently occurring plant is: \" + str(plant_dict[str(most_frequent)]))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T16:06:21.909052Z","iopub.execute_input":"2025-03-10T16:06:21.909363Z","iopub.status.idle":"2025-03-10T16:06:21.920875Z","shell.execute_reply.started":"2025-03-10T16:06:21.909339Z","shell.execute_reply":"2025-03-10T16:06:21.919592Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Now in the following lines of code we create a file to submit to Kaggle:","metadata":{}},{"cell_type":"code","source":"test_images_folder = \"test_images\"\ntest_image_names = os.listdir(os.path.join(input_path, test_images_folder))\n\n# Creating a submission Dataframe manually\nsubmission = pd.DataFrame({\n    \"image_id\": test_image_names,\n    \"label\": most_frequent\n})\n\n# Save the submission file\nsubmission.to_csv(\"most_frequent_submission.csv\", index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-10T16:08:20.112727Z","iopub.execute_input":"2025-03-10T16:08:20.113064Z","iopub.status.idle":"2025-03-10T16:08:20.139381Z","shell.execute_reply.started":"2025-03-10T16:08:20.113035Z","shell.execute_reply":"2025-03-10T16:08:20.138435Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}