{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"I preload all the necessary packages needed for ML, data loading, processing and visualization.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport cv2\nimport matplotlib.pyplot as plt\n#import imagehash\n\n%matplotlib inline \n#To plot within the notebook\nfrom sklearn.decomposition import PCA # Surmoning the Principal Component Gods\nimport skimage.transform as imtran # The wizard of Image transformation. They do a mean Image resize\nfrom PIL import Image # This one is a bit boring, Opens images, bla bla bla, the usual.\nfrom mpl_toolkits.mplot3d import Axes3D # 3D projections that are actually better than imax!\nimport glob #An earth wannabe, 'global environment' package.\n# Input data files are available in the \"../input/\" directory.\nfrom sklearn.preprocessing import StandardScaler\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\nfrom pylab import array, plot, show, axis, arange, figure, uint8\nimport os\nprint(os.listdir(\"../input/\"))\n\n# Any results you write to the current directory are saved as output.","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2022-12-20T02:34:26.250273Z","iopub.execute_input":"2022-12-20T02:34:26.250660Z","iopub.status.idle":"2022-12-20T02:34:28.467544Z","shell.execute_reply.started":"2022-12-20T02:34:26.250623Z","shell.execute_reply":"2022-12-20T02:34:28.466482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Load the training labels. Given that our dataset is limited, we will make use of the training label to both train and test this quick and dirty approach.","metadata":{}},{"cell_type":"code","source":"train_labels = pd.read_csv(\"../input/train.csv\")\nprint(len(train_labels))\ntrain_labels.head()","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.status.busy":"2022-12-20T02:34:32.877184Z","iopub.execute_input":"2022-12-20T02:34:32.877502Z","iopub.status.idle":"2022-12-20T02:34:32.915854Z","shell.execute_reply.started":"2022-12-20T02:34:32.877451Z","shell.execute_reply":"2022-12-20T02:34:32.914481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I use this to increase the intensity of the the black patches in the images. i noticed this helps with the accuracy as the level of severity of a person with blindness can be told from the level of intensity of the black patches and given that the contrast transformation is similar for all images, it really does not not distort the results. Make sure to do the same for the test set though. This might slow down your algorithm though.","metadata":{}},{"cell_type":"code","source":"def increase_contrast(image):\n    #Sourced:\n    #https://stackoverflow.com/questions/19363293/...\n    #...whats-the-fastest-way-to-increase-color-image-contrast-with-opencv-in-python-c\n    maxIntensity = 255.0 # depends on dtype of image data\n    x = arange(maxIntensity) \n\n    # Parameters for manipulating image data\n    phi = 1\n    theta = 1\n\n    # Increase intensity such that\n    # dark pixels become much brighter, \n    # bright pixels become slightly bright\n    newImage0 = (maxIntensity/phi)*(image/(maxIntensity/theta))**0.5\n    newImage0 = array(newImage0,dtype=uint8)\n\n   \n\n    y = (maxIntensity/phi)*(x/(maxIntensity/theta))**0.5\n\n    # Decrease intensity such that\n    # dark pixels become much darker, \n    # bright pixels become slightly dark \n    newImage1 = (maxIntensity/phi)*(image/(maxIntensity/theta))**2\n    newImage1 = array(newImage1,dtype=uint8)\n    return newImage1","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:34:38.122175Z","iopub.execute_input":"2022-12-20T02:34:38.122476Z","iopub.status.idle":"2022-12-20T02:34:38.128847Z","shell.execute_reply.started":"2022-12-20T02:34:38.122427Z","shell.execute_reply":"2022-12-20T02:34:38.127848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I use this function to load the images and standardize them. This is the first standardization step and I think it is well done durng preloading. I also do an incersion of the image which, while it distort the colors, it really helps emphasize the black patches within the eye ball scans. This I noticed helped improve the accuracy of detection.","metadata":{}},{"cell_type":"code","source":"size = (100, 100)\ndef load_images_from_folder(folder, labels, tr):\n    images, names, im_diagnostics = [],[],[]\n    kernel = np.ones((3,3),np.float32)/9\n    for filename in os.listdir(folder):\n        img = cv2.imread(os.path.join(folder,filename))\n        \n        if img is not None:\n            #Increase Contrast\n            #img = increase_contrast(img)\n            filename = filename.split(\".\")[0]\n            names.append(filename)\n            if tr:\n                im_diagnostics.append(int(labels[labels[\"id_code\"] == filename].diagnosis))\n            image_resized = cv2.resize(img, size)\n            image_filtered = cv2.filter2D(image_resized,-1,kernel)\n            # invert\n            image_inverted = cv2.bitwise_not(image_resized)\n            gray_image = cv2.cvtColor(image_inverted,cv2.COLOR_BGR2RGB)\n            images.append(gray_image)\n    return images, names, im_diagnostics\n\n","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:34:44.409019Z","iopub.execute_input":"2022-12-20T02:34:44.409313Z","iopub.status.idle":"2022-12-20T02:34:44.417963Z","shell.execute_reply.started":"2022-12-20T02:34:44.409265Z","shell.execute_reply":"2022-12-20T02:34:44.416888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Load the training images. Note, for this, you only need the training images given that we dont have the training labels. The competition does not give us these so our training dataset will have to suffice for cross validation and testing purposes. I never has a CV set though since I wanted to do a quick and dirty implementation so how the way the algorithm scales might be questionable. Nonetheless, I think I can mitigate this by fitting multiple algorithms then allowing them to vote and take the modal vote. In case of a tie, just trust the most accurate of the algorithms so far.","metadata":{}},{"cell_type":"code","source":"images, im_names, im_diagnostics = load_images_from_folder('../input/train_images', labels = train_labels, tr= True)","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:34:49.736292Z","iopub.execute_input":"2022-12-20T02:34:49.736574Z","iopub.status.idle":"2022-12-20T02:40:06.484016Z","shell.execute_reply.started":"2022-12-20T02:34:49.736526Z","shell.execute_reply":"2022-12-20T02:40:06.482550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"images_ts_main, im_names_ts_main, im_diagnostics_ts = load_images_from_folder('../input/test_images', labels = train_labels,tr=False)","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:41:08.054750Z","iopub.execute_input":"2022-12-20T02:41:08.055027Z","iopub.status.idle":"2022-12-20T02:41:47.988531Z","shell.execute_reply.started":"2022-12-20T02:41:08.054974Z","shell.execute_reply":"2022-12-20T02:41:47.987513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"Images Loaded: {len(images)} Labels Organized: {len(im_diagnostics)}\")","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:42:50.584676Z","iopub.execute_input":"2022-12-20T02:42:50.584986Z","iopub.status.idle":"2022-12-20T02:42:50.590649Z","shell.execute_reply.started":"2022-12-20T02:42:50.584943Z","shell.execute_reply":"2022-12-20T02:42:50.589534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A simple utility to display the images.","metadata":{}},{"cell_type":"code","source":"def display_image(images, labels, im_names, rows = 1,columns = 3):\n    count = 0\n    #You can visualize a random sample\n    sampleList = np.random.choice(range(len(images)), 12, replace = False)\n    \n    simages, sim_names = [], [] \n    for i in sampleList:\n        simages.append(images[i])\n        sim_names.append(im_names[i]+\"-->\"+str(labels[i]))\n    \n    for i in range(1, len(sampleList)//3):\n        plot_image = np.concatenate((simages[count],simages[count+1],simages[count+2]), axis=1)\n        \n        plt.imshow(plot_image)\n        print(f\"These are sample Images: {', '.join(sim_names[count:count+3])} Respectively\")\n        plt.show() \n        count+=3\n        \ndisplay_image(images,im_diagnostics, im_names=im_names, rows=4)","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:42:54.495033Z","iopub.execute_input":"2022-12-20T02:42:54.495352Z","iopub.status.idle":"2022-12-20T02:42:54.958811Z","shell.execute_reply.started":"2022-12-20T02:42:54.495301Z","shell.execute_reply":"2022-12-20T02:42:54.956659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Used this to verify the labels were correctly retrieved\nprint(int(train_labels[train_labels[\"id_code\"] == \"b3d12069e1c5\"].diagnosis))","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:43:00.870517Z","iopub.execute_input":"2022-12-20T02:43:00.870950Z","iopub.status.idle":"2022-12-20T02:43:00.878381Z","shell.execute_reply.started":"2022-12-20T02:43:00.870777Z","shell.execute_reply":"2022-12-20T02:43:00.877585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I perform PCA on this retaining 99% of the variance also partition the testing set. Ts main is the main label. The best fitting model is one trained over the whole data set so for submission purposes, I simply removed the partitioning on the training model and used this to make predictions on ts_images_main. I also performed standard scaling on this so that the mean is 0. This is especially beneficial for the neural network performance. ","metadata":{}},{"cell_type":"code","source":"#Run a basic neural network on the data\nfrom sklearn.neural_network import MLPClassifier\nimport random\n#I reduce the number of features with PCA\nfrom sklearn.decomposition import PCA\nscaler= StandardScaler()\n\npca = PCA(.99)\n\n#Stratification into training and testing data.\n#I will add the CV data later\ntr_images  = pca.fit_transform(scaler.fit_transform([img.flatten() for img in images]))\ntr_labels = im_diagnostics\nts_images  = pca.transform(scaler.transform([img.flatten()for img in images[3001:]]))\nts_labels = im_diagnostics[3001:]\nts_images_main=pca.transform(scaler.transform([img.flatten()for img in images_ts_main]))","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:43:06.801297Z","iopub.execute_input":"2022-12-20T02:43:06.801699Z","iopub.status.idle":"2022-12-20T02:45:56.728583Z","shell.execute_reply.started":"2022-12-20T02:43:06.801657Z","shell.execute_reply":"2022-12-20T02:45:56.727557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Algorithm one, Use a simple Multi-Layer Perceptron with 5 hidden layer of size (input)//2. This is borne out of testing. You may play around with the model but this configration wourked well for me so I stuck for it.","metadata":{}},{"cell_type":"code","source":"#Now Lets Fit a neural Network on a portion of the data just to keep things fast and for testing\nprint(\"Fitting Model\")\nlayers = [((pca.n_components_)//2)]*5\nclf = MLPClassifier(activation = \"relu\",solver = \"sgd\", hidden_layer_sizes=tuple(layers))\nclf.fit(tr_images, tr_labels)\nfrom sklearn.metrics import accuracy_score\n\nprediction_nn = clf.predict(ts_images)\nmain_prediction_nn = clf.predict(ts_images_main)\nprint(set(prediction_nn))\nprint(f\"First 10 labels: {ts_labels[:10]}\")\nprint(f\"First 10 prediction: {prediction_nn[:10]}\")\nacc_nn=accuracy_score(prediction_nn, ts_labels)\nprint(acc_nn)","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:46:49.471569Z","iopub.execute_input":"2022-12-20T02:46:49.471932Z","iopub.status.idle":"2022-12-20T02:52:41.085905Z","shell.execute_reply.started":"2022-12-20T02:46:49.471876Z","shell.execute_reply":"2022-12-20T02:52:41.085222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now to leverage Keras to see if we can achieve better accuracy with more advanced Neural networks.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers","metadata":{"execution":{"iopub.status.busy":"2022-12-20T02:52:48.290173Z","iopub.execute_input":"2022-12-20T02:52:48.290788Z","iopub.status.idle":"2022-12-20T02:52:50.866706Z","shell.execute_reply.started":"2022-12-20T02:52:48.290738Z","shell.execute_reply":"2022-12-20T02:52:50.866035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I am making use of the combined opinions of the different models to make a selection. Take the mode of the votes or if the is a tie, take the vote of the most accurate of the models.","metadata":{}},{"cell_type":"code","source":"\"\"\"\nfrom statistics import mode\ncombo = []\nmax_accs = [acc_svc, acc_lg,acc_nn].index(max([acc_svc, acc_lg,acc_nn]))\nfor votes in zip(prediction_nn):\n    try:\n        combo.append(mode(votes))\n    except:\n        combo.append(votes[max_accs])\nprint(f\"First 10 labels: {ts_labels[:100]}\")\nprint(f\"First 10 prediction: {combo[:100]}\")\nprint(accuracy_score(combo, ts_labels))\n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2022-12-20T03:09:32.569846Z","iopub.execute_input":"2022-12-20T03:09:32.570137Z","iopub.status.idle":"2022-12-20T03:09:32.576429Z","shell.execute_reply.started":"2022-12-20T03:09:32.570090Z","shell.execute_reply":"2022-12-20T03:09:32.575382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This was used for submission purposes. ","metadata":{}},{"cell_type":"code","source":"from statistics import mode\nmain_combo = []\n\n\nfor votes in zip(main_prediction_nn):\n    try:\n        main_combo.append(mode(votes))\n    except:\n        main_combo.append(votes[max_accs])\n\n\nsubmission = pd.DataFrame({\"id_code\":im_names_ts_main, \"diagnosis\":main_combo})\nsubmission.head()\nsubmission.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-12-20T03:10:04.814684Z","iopub.execute_input":"2022-12-20T03:10:04.815006Z","iopub.status.idle":"2022-12-20T03:10:04.847907Z","shell.execute_reply.started":"2022-12-20T03:10:04.814957Z","shell.execute_reply":"2022-12-20T03:10:04.846685Z"},"trusted":true},"execution_count":null,"outputs":[]}]}