{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":20270,"databundleVersionId":1222630,"sourceType":"competition"}],"dockerImageVersionId":30587,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Ok, here I am. AI Rookie with his first Kaggle notebook.\n\nI am working through the second chapter of fast.ai course and follow the hint of Jeremy Howard to work on a project with drives me.\n\n### Drivetrain Approach\n\nJeremy Howard and fellows described the [Drivetrain Approach](https://www.oreilly.com/radar/drivetrain-approach-data-products/) as a four step process to build great data products.\n\n#### First things first\n\n1. Define objective\n\nWhat outcome do I want to achive? In other words, what is the main goal of my project?\n\nUsing this solution, Users should be warned about malignant melanoms as early as possible. This could relieve suffering and save their lives. Even in developed countries like Germany it is not easy for everyone to organise a schedule by a dermatologist. It should be easy to use this solution for everyone by taking a photo of their skin mole and analyse it privately.   \n\n2. Identify the Levers of the project which we can control\n\nWhat input data to the project can we control to influence the outcome of the project?\n\nWe can control which dataset for the training of the model we want to use. The whole project must be able to run on an actual Android smartphone, without lacking any personal information to any service.\n\n3. Which new data can we collect?\n\nWe can use an users image of his skin mole and try to classificate it according to the general dataset AND his own personal trained dataset collected over the time.\n\n4. What model should we build?\n\nHow do the levers influence the outcome?\n\n5. (Added from me) How can we optimize the model?\n\n#### Questions to be solved\n\n1. How can I find out that the project I want to work on is not too difficult for me?\n\nJeremy's hint:\n***When selecting a project, the most important consideration is data availability.***\n\nThere exist a large dataset for melanom classification which I think I can use:\n\n[SIIM-ISIC melanom classification competition](https://www.kaggle.com/competitions/siim-isic-melanoma-classification/overview)\n\nSo let me try to get access and interpret the ***ISIC 2020 Challenge Dataset***.\n\n1.1 How can I link the dataset to my notebook? I don't want to copy all of the data everytime I try something new.\n\nI made a video about it at my youtube channel [Kaggle demo](https://www.youtube.com/watch?v=_6gU3mqOjOs)\n\n1.2 How can I programatically access this dataset?\n\nLet us try out some things.","metadata":{}},{"cell_type":"markdown","source":"First let us import the fastbook module to get access to the Image class.","metadata":{}},{"cell_type":"code","source":"#hide\n! pip install -Uqq fastbook\n","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:04:46.201559Z","iopub.execute_input":"2023-12-20T19:04:46.202465Z","iopub.status.idle":"2023-12-20T19:05:06.640861Z","shell.execute_reply.started":"2023-12-20T19:04:46.202416Z","shell.execute_reply":"2023-12-20T19:05:06.639431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Than import it.","metadata":{}},{"cell_type":"code","source":"from fastbook import *\n","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:05:15.896276Z","iopub.execute_input":"2023-12-20T19:05:15.896807Z","iopub.status.idle":"2023-12-20T19:05:15.914532Z","shell.execute_reply.started":"2023-12-20T19:05:15.896759Z","shell.execute_reply":"2023-12-20T19:05:15.913248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And define the path to the images for the training data in jpeg format.\nThe linked dataset is available through the input directory.","metadata":{}},{"cell_type":"code","source":"training_path = '../input/siim-isic-melanoma-classification/jpeg/train/'\nsource = training_path + 'ISIC_0015719.jpg'\nim = Image.open(source)\nim.to_thumb(256,256)\n\n     ","metadata":{"execution":{"iopub.status.busy":"2023-12-09T14:28:22.781979Z","iopub.execute_input":"2023-12-09T14:28:22.782745Z","iopub.status.idle":"2023-12-09T14:28:23.148056Z","shell.execute_reply.started":"2023-12-09T14:28:22.782694Z","shell.execute_reply":"2023-12-09T14:28:23.147156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that we have explored the dataset and can access it, we need to organise it in way that fastai can work with it.\nAs of now we only know that we have to put the different images classes into distict directories so that fastai can work with them.\n\nLet us have a look at the train.csv file.\n\nWe can see, that for every image we have an entry benign or malignant. All we have to do right now is to create one directory for each of the benign or malignant melanoma and store the coresspondig image in it.\n\nWe will use the pandas library to read the train.csv file.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\ndataset_path = '../input/siim-isic-melanoma-classification/'\nmetadata = dataset_path + 'train.csv'\n\ndf = pd.read_csv(metadata)\n","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:05:18.802739Z","iopub.execute_input":"2023-12-20T19:05:18.803192Z","iopub.status.idle":"2023-12-20T19:05:18.928817Z","shell.execute_reply.started":"2023-12-20T19:05:18.803159Z","shell.execute_reply":"2023-12-20T19:05:18.927411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have 33126 images of melanoma in total.","metadata":{}},{"cell_type":"code","source":"benign = (df[['benign_malignant']] == 'benign').sum()\nmalignant =(df[['benign_malignant']] == 'malignant').sum()\ntarget =(df[['target']] == 1).sum()\n\nprint('benign: {:d}, malignant: {:d}, target: {}'.format(benign[0], malignant[0], target[0]))\n","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:05:25.899573Z","iopub.execute_input":"2023-12-20T19:05:25.900120Z","iopub.status.idle":"2023-12-20T19:05:25.953439Z","shell.execute_reply.started":"2023-12-20T19:05:25.900076Z","shell.execute_reply":"2023-12-20T19:05:25.952180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From these 33126 images, 32542 images show benign melanoma and 584 show malignant melanoma.\n\nNow let us build a list of all filenames which are malignant melanoma.\n\n","metadata":{}},{"cell_type":"code","source":"malignant_images = df.loc[df['benign_malignant'] == 'malignant']\nmalignant_images = malignant_images['image_name'].tolist()\nnumbers_of_malignant_images = len(malignant_images)\n\nbenign_images = df.loc[df['benign_malignant'] == 'benign']\nbenign_images = benign_images['image_name'].tolist()\nnumbers_of_benign_images = len(benign_images)\n\nprint('benign images: {:d}'.format(numbers_of_benign_images))\nprint('malignant images: {:d}'.format(numbers_of_malignant_images))","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:05:34.010339Z","iopub.execute_input":"2023-12-20T19:05:34.010804Z","iopub.status.idle":"2023-12-20T19:05:34.043690Z","shell.execute_reply.started":"2023-12-20T19:05:34.010769Z","shell.execute_reply":"2023-12-20T19:05:34.042485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us implement a function wich stores a number of a specific type of melanoma images in a specific directory.","metadata":{}},{"cell_type":"code","source":"import os\n\ndef link_images(list_of_images, number, attribute, input_path, output_path):\n    attribute_path = output_path + attribute\n\n    if not os.path.exists(attribute_path):\n            os.mkdir(attribute_path)\n    \n    for i in list_of_images:\n        \n        filename = i + '.jpg'\n        src = input_path + filename\n        dest = attribute_path + '/' + filename\n        \n#        print('linking ' + src + ' to ' + dest)\n        if not os.path.exists(dest):\n            os.symlink(src, dest)\n            \n        number = number - 1\n        if number <= 0:\n            break\n\n    return\n            \n\ninput_path = '/kaggle/input/siim-isic-melanoma-classification/jpeg/train/'\noutput_path = '/kaggle/working/training/'\n\nif not os.path.exists(output_path):\n    os.mkdir(output_path)\n\n# let us prepare the malignant training images directory\nlink_images(malignant_images, 10, 'malignant', input_path, output_path)\n\n# let us prepare the benign training images directory\nlink_images(benign_images, 10, 'benign', input_path, output_path)\n   ","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:05:40.966591Z","iopub.execute_input":"2023-12-20T19:05:40.967072Z","iopub.status.idle":"2023-12-20T19:05:40.983304Z","shell.execute_reply.started":"2023-12-20T19:05:40.967018Z","shell.execute_reply":"2023-12-20T19:05:40.981322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# source = output_path + 'malignant/' + 'ISIC_9998682.jpg'\n# print(source)\n# im = Image.open(source)\n# im.to_thumb(256,256)","metadata":{"execution":{"iopub.status.busy":"2023-12-16T07:22:19.175738Z","iopub.execute_input":"2023-12-16T07:22:19.176151Z","iopub.status.idle":"2023-12-16T07:22:19.397477Z","shell.execute_reply.started":"2023-12-16T07:22:19.176115Z","shell.execute_reply":"2023-12-16T07:22:19.396475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training Session\n\nUsing fastai means that we have to create a *DataLoaders* class.\n\n*DataLoaders* is a thin python class which stores multiple DataLoader objects, normally a *train* and a associated *valid*\n\nTransfer our linked SIIM-ISIC Melanoma Dataset to a DataLoaders object we need to tell fastai at least four things (TGLV):\n\n## T.G.L.V (Type, Grab, Label, Validate) ##\n\n1) Type - What kind of data are we working with?\n2) Grab - How to get the list of items?\n3) Label - How to label these items\n4) Validate - How to create the validation set\n","metadata":{}},{"cell_type":"markdown","source":"Fastai provides a data block API. This API allows us to customize every stage which is needed to build a Dataloaders.\nAs a first try, we follow directly the instructions of the chapter 2 from fastai course.\n\nFirst, we create naivly a Datablock with our melanoma training data.","metadata":{}},{"cell_type":"code","source":"melanoms = DataBlock(\n    blocks=(ImageBlock, CategoryBlock), \n    get_items=get_image_files, \n    splitter=RandomSplitter(valid_pct=0.2, seed=42),\n    get_y=parent_label,\n    item_tfms=Resize(128))","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:09:19.392644Z","iopub.execute_input":"2023-12-20T19:09:19.393050Z","iopub.status.idle":"2023-12-20T19:09:19.405554Z","shell.execute_reply.started":"2023-12-20T19:09:19.393000Z","shell.execute_reply":"2023-12-20T19:09:19.404333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Testing different image transforming methods.","metadata":{}},{"cell_type":"code","source":"training_path = '/kaggle/working/training/'\nmelanoms_test = melanoms.new(item_tfms=Resize(128, ResizeMethod.Pad, pad_mode='zeros'))\ndls = melanoms.dataloaders(training_path)\ndls.valid.show_batch(max_n=4, nrows=1)","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:15:05.109216Z","iopub.execute_input":"2023-12-20T19:15:05.110927Z","iopub.status.idle":"2023-12-20T19:15:06.966649Z","shell.execute_reply.started":"2023-12-20T19:15:05.110871Z","shell.execute_reply":"2023-12-20T19:15:06.965773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_path = '/kaggle/working/training/'\nmelanoms_test = melanoms.new(item_tfms=Resize(128, ResizeMethod.Squish))\ndls = melanoms.dataloaders(training_path)\ndls.valid.show_batch(max_n=4, nrows=1)","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:15:11.846185Z","iopub.execute_input":"2023-12-20T19:15:11.846799Z","iopub.status.idle":"2023-12-20T19:15:13.767503Z","shell.execute_reply.started":"2023-12-20T19:15:11.846767Z","shell.execute_reply":"2023-12-20T19:15:13.766443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_path = '/kaggle/working/training/'\nmelanoms_test = melanoms.new(item_tfms=RandomResizedCrop(256, min_scale=0.1))\ndls = melanoms.dataloaders(training_path)\ndls.valid.show_batch(max_n=5, nrows=1, unique=True)","metadata":{"execution":{"iopub.status.busy":"2023-12-20T19:20:51.345791Z","iopub.execute_input":"2023-12-20T19:20:51.346253Z","iopub.status.idle":"2023-12-20T19:20:54.662023Z","shell.execute_reply.started":"2023-12-20T19:20:51.346220Z","shell.execute_reply":"2023-12-20T19:20:54.661170Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"training_path = '/kaggle/working/training/'\ndls = melanoms.dataloaders(training_path)","metadata":{"execution":{"iopub.status.busy":"2023-12-18T16:59:49.424326Z","iopub.execute_input":"2023-12-18T16:59:49.424749Z","iopub.status.idle":"2023-12-18T16:59:50.446207Z","shell.execute_reply.started":"2023-12-18T16:59:49.424720Z","shell.execute_reply":"2023-12-18T16:59:50.445175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dls.valid.show_batch(max_n=4, nrows=1)\n","metadata":{"execution":{"iopub.status.busy":"2023-12-18T16:59:54.624062Z","iopub.execute_input":"2023-12-18T16:59:54.624483Z","iopub.status.idle":"2023-12-18T16:59:56.677845Z","shell.execute_reply.started":"2023-12-18T16:59:54.624454Z","shell.execute_reply":"2023-12-18T16:59:56.676719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"But how do we have to organise the melanoma images to be able to train our modell with the specific labels?\nLet us examine the parameters of the Datablock() function.","metadata":{}}]}