{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"nbconvert_exporter":"python","mimetype":"text/x-python","codemirror_mode":{"version":3,"name":"ipython"},"version":"3.6.3","pygments_lexer":"ipython3","file_extension":".py","name":"python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":6775,"sourceType":"competition"},{"sourceId":8756,"sourceType":"datasetVersion","datasetId":5891},{"sourceId":8791,"sourceType":"datasetVersion","datasetId":5925},{"sourceId":8793,"sourceType":"datasetVersion","datasetId":5927},{"sourceId":8806,"sourceType":"datasetVersion","datasetId":5927},{"sourceId":10122,"sourceType":"datasetVersion","datasetId":5925},{"sourceId":10125,"sourceType":"datasetVersion","datasetId":7044}],"dockerImageVersionId":33,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Introduction**\n\nThe task is to predict whether a passenger in an airport could be a Threat because he could have any kind of dangerous object ( gun, knifes ...). This objects could be detected by a millimetre wave scanner called the High Definition-Advanced Imaging Technology (HD-AIT) system.\nThe scanner generates a complete set of body scanner images and we want to find any kind of anomalies in the image it could think us that the passenger could have an object and it could be a threat for the rest of passengers.\nWe are going to apply clustering techniques for detection of anomalies in the image, and previously we are going to generate a complete training dataset with image that we are viewed before checking that the passenger don't have any kind of internal object.","metadata":{"_cell_guid":"4fc38b3b-f51a-4fde-add8-ccda28357144","_uuid":"00b8832b3c051b1b4ea6e8c568c57c96fd66494f"}},{"cell_type":"markdown","source":"En el entorno actual de la aviación, las largas filas y el proceso de pasar el equipaje por los controles de seguridad no son experiencias agradables, pero la seguridad en los aeropuertos es un requisito crítico e indispensable para garantizar un viaje seguro. La Administración de Seguridad en el Transporte de los Estados Unidos (TSA), responsable de la seguridad en todos los aeropuertos del país, entiende mejor que nadie la necesidad de realizar revisiones exhaustivas sin que esto conlleve a tiempos de espera prolongados. Con más de dos millones de pasajeros revisados diariamente, la eficiencia es crucial.\n\nComo parte de su programa \"Apex Screening at Speed\", el Departamento de Seguridad Nacional (DHS) ha identificado que las altas tasas de falsas alarmas son un factor que crea cuellos de botella significativos en los puntos de control de los aeropuertos. Cuando los sensores y algoritmos de la TSA predicen una posible amenaza, el personal debe realizar una revisión manual secundaria, lo que ralentiza considerablemente el proceso. A medida que el número de viajeros aumenta cada año y surgen nuevas amenazas, los algoritmos de predicción necesitan mejorar continuamente para satisfacer la demanda creciente.\n\nActualmente, la TSA adquiere algoritmos actualizados exclusivamente de los fabricantes de los equipos de escaneo utilizados. Estos algoritmos son propietarios, costosos y suelen lanzarse en ciclos prolongados. En este contexto, la TSA está adoptando un enfoque innovador al desafiar a la comunidad de ciencia de datos a mejorar la precisión de sus algoritmos de predicción de amenazas. Usando un conjunto de datos de imágenes recogidas con la última generación de escáneres, los participantes deben identificar la presencia de amenazas simuladas en diversas condiciones de tipos de objetos, tipos de ropa y tipos de cuerpo. Incluso una reducción modesta en las falsas alarmas ayudará significativamente a la TSA a mejorar la experiencia de los pasajeros mientras se mantiene un alto nivel de seguridad.\n\nEste desafío se basa en el uso de un algoritmo de aprendizaje profundo basado en Redes Neuronales Convolucionales (CNN), que ha demostrado ser altamente eficaz en la clasificación y detección de imágenes. La competencia se desarrollará en dos etapas, y se invita a los participantes a consultar las FAQ de las dos etapas para comprender mejor el proceso.\n\nEs importante destacar que todas las personas incluidas en el conjunto de datos son voluntarias que han consentido en el uso de sus imágenes para esta competencia. Las imágenes pueden contener contenido sensible, por lo que se solicita a todos los participantes que se conduzcan con profesionalismo, respeto y madurez al trabajar con estos datos.","metadata":{}},{"cell_type":"code","source":"import sys\nimport os\nimport numpy as np\nfrom numpy import array, asarray, ma, zeros, sum\nfrom matplotlib import pyplot as plt\nimport cv2\nimport pandas as pd\nimport seaborn as sns\nimport scipy.stats as stats\nfrom scipy.cluster.vq import vq, kmeans, whiten, kmeans2\nfrom sklearn.cluster import DBSCAN\nfrom sklearn import metrics\nfrom sklearn.datasets.samples_generator import make_blobs\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.svm import LinearSVC\nfrom sklearn.datasets import make_classification","metadata":{"_cell_guid":"5dcc3732-bef3-44a9-a9f7-c0031bdf493c","_uuid":"49dcf09247abfef871c28bb767936b60baa3036f","collapsed":true,"jupyter":{"outputs_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nfrom subprocess import check_output\ntry:\n    print(check_output([\"ls\", \"../input\"]).decode(\"utf8\"))\n    os.system('mkdir ./output')\n    os.system('mkdir ./output/Threats')\n    os.system('mkdir ./output/NoThreats')\n    os.system('mkdir ./training')\n    os.system('mkdir ./results')\n    os.system('mkdir ./models')\n    os.system('mkdir ./input')\n    print(check_output([\"ls\", \"../input/datainput\"]).decode(\"utf8\"))\n    print(check_output([\"ls\", \"../input/modelstraining\"]).decode(\"utf8\"))\n    print(check_output([\"ls\", \"../input/highresolution\"]).decode(\"utf8\"))\n    print(check_output([\"ls\", \"../input/modeldbscan\"]).decode(\"utf8\"))\n\nexcept Exception as exc:\n    print(\"Exception: {0}\".format(exc))\n# Any results you write to the current directory are saved as output.","metadata":{"_cell_guid":"1ddb84fa-49c6-4b2e-ab83-5b9fa9be0ac8","_uuid":"4979653104b147671f429590b937b1c5309abd9a"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Analysis Approach**\nThis original dataset contains a large number of body scans acquired by a new generation of millimetre wave scanner called the High Definition-Advanced Imaging Technology (HD-AIT) system. I considered the application of clustering techniques such as k-means and DBSCAN to generate a csv files with data that could be useful to do detect possible threats and marking the body zones scanned with possibles anomalies across the the careful analysis of distortion parameters in k-means techniques and the statistics parameters produced by  the application of DBSCAN techniques","metadata":{"_cell_guid":"2c294650-ce0c-492c-9234-3ae8b32237ff","_uuid":"f8450046d92d639e21effdfaef01306108467310"}},{"cell_type":"code","source":"PATH_PHOTO_FILES = '.'\nNAME_FILE_MODEL_KMEANS = '../input/modelstraining/model_kmeans.csv'\nNAME_FILE_MODEL_DBSCAN = '../input/modeldbscan/model_dbscan.csv'\nDATA_FOR_PREDICTION_FILE = '../input/low-resolution/low_quality.csv'\nDATA_FOR_PREDICTION_FILE_FILTERED = '../input/highresolution/high_quality.csv'","metadata":{"_cell_guid":"d6022323-bc15-48e9-bb3d-478ed1657801","_uuid":"d4eac8087a5b06cd2fcf6a14cf9a6ef65544d93f","collapsed":true,"jupyter":{"outputs_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"############################################################################################################\n# Name: cluster_analisys\n# Autor: Ramiro Bueno Martínez\n# Date: 27/11/2017\n##########################################################################################################\ndef cluster_analisys(subset):\n\ttry:\n\t\twhitened = whiten(subset)\n\t\tcodebook,distortion = kmeans(whitened,5)\n\t\treturn distortion,whitened,codebook\n\texcept Exception as exception:\n\t\tprint (\"cluster_analisys: Excepcion {0}\".format(exception))","metadata":{"_cell_guid":"cc483fe8-8fe8-46ab-a997-133808119ba2","_uuid":"9b4ace7f17f9d1471b08539b68bc5a940cdefba1","collapsed":true,"jupyter":{"outputs_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"############################################################################################################\n# Name: creation_dataset_dbscan\n# Autor: Ramiro Bueno Martínez\n# Date: 30/11/2017\n# Description: The goal of fthis function is the dbscan analysis an generate a group of statistic datas\n# usuful in the building of a model training for finding possibles anomalies in the image \n############################################################################################################\ndef creation_dataset_dbscan(dataset, FILENAME,highcontrast=False):\n\ttry:\n\t\toutput_dir = PATH_PHOTO_FILES + '/' + 'output'\n\t\tinput_dir = PATH_PHOTO_FILES + '/' + 'input'\n\t\tname_of_photo_file =  FILENAME\n\t\tprint(\"Starting creaction dataset dbscan {0}\".format(name_of_photo_file))\n\t\trange_images = [0,4,8,12]\n\t\tm_homegeneity = 0\n\t\tfor nth in range_images:\n\t\t\tan_img = get_single_image(name_of_photo_file, nth)  \t\t\t\t#returns the nth=3 image from the image stack\n\t\t\tif highcontrast == True:\n\t\t\t\timg_rescaled = convert_to_grayscale(an_img)\n\t\t\t\tan_img = spread_spectrum(img_rescaled)\n\t\t\tdata_array = np.array(an_img)\n\t\t\twhitened = whiten(data_array)\n\t\t\trecord = pd.DataFrame({'x_axes': whitened[:, 0],'y_axes':whitened[:, 1]})\n\t\t\tcenters = np.array(record)\n\t\t\tX, labels_true = make_blobs(n_samples=750, centers=centers, cluster_std=0.4, random_state=0)\n\t\t\tX = StandardScaler().fit_transform(X)\n\t\t\tdb = DBSCAN(eps=0.3, min_samples=10).fit(X)\n\t\t\tcore_samples_mask = np.zeros_like(db.labels_, dtype=bool)\n\t\t\tcore_samples_mask[db.core_sample_indices_] = True\n\t\t\tlabels = db.labels_\n\t\t\tn_clusters_ = len(set(labels)) - (1 if -1 in labels else 0)\n\t\t\tName = FILENAME\n\t\t\tnClusters = n_clusters_\n\t\t\tHomegeneity = metrics.homogeneity_score(labels_true, labels)\n\t\t\tm_homegeneity = m_homegeneity + Homegeneity\n\t\t\tCompleteness = \tmetrics.completeness_score(labels_true, labels)\n\t\t\tVmeasure = metrics.v_measure_score(labels_true, labels)\n\t\t\tARIndex = metrics.adjusted_rand_score(labels_true, labels)\n\t\t\tAMInformation = \tmetrics.adjusted_mutual_info_score(labels_true, labels)\n\t\t\tSilCoefficient = metrics.silhouette_score(X, labels)\n\t\t\tThreat = 'nothreat'\n\t\t\trecord_dbscan = pd.DataFrame([[Name, nClusters, Homegeneity, Completeness, Vmeasure, ARIndex, AMInformation, SilCoefficient, Threat]],columns=['Name','nClusters','Homegeneity','Completeness','Vmeasure','ARIndex','AMInformation','SilCoefficient','Threat'])\n\t\t\tdataset = dataset.append(record_dbscan,ignore_index=True)\n\t\t\t\n\t\tmedia = 0\n\t\tmedia = m_homegeneity / 4\n\t\tpercent = 5  # 5% of the mean value we suppose that its a wrong value \n\t\tthreshold_thread = media * percent / 100\n\t\tfor i in range(len(dataset)):\n\t\t\tif abs(media - dataset.loc[i,('Homegeneity')]) > threshold_thread: \n\t\t\t\tprint(\"Finded a possible case\")\n\t\t\t\tdataset.loc[i,('Threat')] = 'threat'\n\t\t\tif abs(media - dataset.loc[i,('Homegeneity')]) <= threshold_thread:\n\t\t\t\tdataset.loc[i,('Threat')] = 'nothreat'\n\texcept Exception as ex:\n\t\tprint(\"Exception: {0}\".format(ex))\n\t\tsys.exit(-1)\n\tfinally:\n\t\treturn dataset\n","metadata":{"_cell_guid":"51e07080-d301-4b31-b9d4-2508eed14cdc","_uuid":"d0a02a559c3c35137b4059c389822b080188136f","collapsed":true,"jupyter":{"outputs_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"############################################################################################################\n# Name: creation_dataset_threats\n# Autor: Ramiro Bueno Martínez\n# Date: 27/11/2017\n# Description: The goal of this function is the finding of possible anomalies in the different kind of\n# images generated with the body scannes, using clustering techniques based in the analysis of different\n# part of the body and different possition imagees\n############################################################################################################\ndef creation_dataset_threats(dataset, FILENAME, highcontrast=False):\n\t\tm_distortion = 0\n\t\tcol = 0\n\t\trow = 0\n\t\tnth = 0\n\t\toutput_dir = PATH_PHOTO_FILES + '/' + 'output'\n\t\tinput_dir = PATH_PHOTO_FILES + '/' + 'input'\n\t\tname_of_photo_file = FILENAME\n\t\tprint (\"Starting clustering Analisys: {0}\".format(FILENAME))\n\t\trange_images = [0,4,8,12]\n\t\ttry:\n\t\t\tfor nth in range_images:\n\t\t\t\tan_img = get_single_image(name_of_photo_file, nth)  \t\t\t\t#returns the nth=3 image from the image stack\n\t\t\t\tif highcontrast == True:\n\t\t\t\t\timg_rescaled = convert_to_grayscale(an_img)\n\t\t\t\t\tan_img = spread_spectrum(img_rescaled)\n\t\t\t\tdata_array = np.array(an_img)\n\t\t\t\tdistortion_array[nth,1],whitened,codebook= cluster_analisys(data_array)\n\t\t\t\tm_distortion = m_distortion + distortion_array[nth,1]\n\t\t\t\trecord = pd.DataFrame([[FILENAME, nth, distortion_array[nth,1],'nothreat']],columns=['Name','Nth','Distortion','Threat'])\n\t\t\t\tdataset = dataset.append(record,ignore_index=True)\n\t\t\t\tcol = col + 1\n\t\t\t\tif col % 4 == 0:\n\t\t\t\t\trow = row + 1\n\t\t\t\t\tcol = 0\n\t\t\t\tprint(\".\")\n\t\t\tmedia = 0\n\t\t\tmedia = m_distortion / 4\n\t\t\tpercent = 17  # 17% of the mean value we suppose that its a wrong value \n\t\t\tthreshold_thread = media * percent / 100\n\t\t\tfor i in range(len(dataset)):\n\t\t\t\tif abs(media - dataset.loc[i,('Distortion')]) > threshold_thread: \n\t\t\t\t\tdataset.loc[i,('Threat')] = 'threat'\n\t\t\t\tif abs(media - dataset.loc[i,('Distortion')]) <= threshold_thread:\n\t\t\t\t\tdataset.loc[i,('Threat')] = 'nothreat'\n\t\texcept Exception as exception:\n\t\t\tprint (\"Function [creation_dataset_threats]: Excepcion {0}\".format(exception))\n\t\t\tsys.exit(1)\n\t\tfinally:\n\t\t\treturn dataset\n","metadata":{"_cell_guid":"fb9a9aba-4a68-499e-96e7-7509b5ccffbd","_uuid":"939e33644b539a34169d9a9c567ed82a18e0c753","collapsed":true,"jupyter":{"outputs_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"############################################################################################################\n# Name: creation_dataset_threats\n# Autor: Ramiro Bueno Martínez\n# Date: 27/11/2017\n# Description: The goal of this function is the finding of possible anomalies in the different kind of\n# images generated with the body scannes, using clustering techniques based in the analysis of different\n# part of the body and different possition imagees\n############################################################################################################\ndef creation_dataset_threats(dataset, FILENAME, highcontrast=False):\n\t\tm_distortion = 0\n\t\tcol = 0\n\t\trow = 0\n\t\tnth = 0\n\t\toutput_dir = PATH_PHOTO_FILES + '/' + 'output'\n\t\tinput_dir = PATH_PHOTO_FILES + '/' + 'input'\n\t\tname_of_photo_file = input_dir + '/' + FILENAME\n\t\tprint (\"Starting clustering Analisys: {0}\".format(FILENAME))\n\t\trange_images = [0,4,8,12]\n\t\ttry:\n\t\t\tfor nth in range_images:\n\t\t\t\tan_img = get_single_image(name_of_photo_file, nth)  \n\t\t\t\tif highcontrast == True:\n\t\t\t\t\timg_rescaled = convert_to_grayscale(an_img)\n\t\t\t\t\tan_img = spread_spectrum(img_rescaled)\n\t\t\t\tdata_array = np.array(an_img)\n\t\t\t\tdistortion_array[nth,1],whitened,codebook= cluster_analisys(data_array)\n\t\t\t\tm_distortion = m_distortion + distortion_array[nth,1]\n\t\t\t\trecord = pd.DataFrame([[FILENAME, nth, distortion_array[nth,1],'nothreat']],columns=['Name','Nth','Distortion','Threat'])\n\t\t\t\tdataset = dataset.append(record,ignore_index=True)\n\t\t\t\tcol = col + 1\n\t\t\t\tif col % 4 == 0:\n\t\t\t\t\trow = row + 1\n\t\t\t\t\tcol = 0\n\t\t\t\tprint(\".\")\n\t\t\tmedia = 0\n\t\t\tmedia = m_distortion / 4\n\t\t\tpercent = 17  # 17% of the mean value we suppose that its a wrong value \n\t\t\tthreshold_thread = media * percent / 100\n\t\t\tfor i in range(len(dataset)):\n\t\t\t\tif abs(media - dataset.loc[i,('Distortion')]) > threshold_thread: \n\t\t\t\t\tdataset.loc[i,('Threat')] = 'threat'\n\t\t\t\tif abs(media - dataset.loc[i,('Distortion')]) <= threshold_thread:\n\t\t\t\t\tdataset.loc[i,('Threat')] = 'nothreat'\n\t\texcept Exception as exception:\n\t\t\tprint (\"Function [creation_dataset_threats]: Excepcion {0}\".format(exception))\n\t\t\tsys.exit(1)\n\t\tfinally:\n\t\t\treturn dataset","metadata":{"_cell_guid":"ff3d8bf0-34f0-4900-a7ad-4e2f262735da","_uuid":"cc11b623c77913c54873338c39b5608d678ac7d2","collapsed":true,"jupyter":{"outputs_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"####################################################################################################\n# Name: print_screen_results\n# Autor: Ramiro Bueno Martínez\n# Date : 28/11/2017\n# Description: Function that plot the results obtained in 4 graphis\n# Graphic 1 - The body scan image of the people\n# Graphic 2 - Adaptative Histogram Equalization of the image\n# Graphic 3 - Graphic results after the image matrix of datas had been analysed with the k-means algorithm\n# Graphic 4 - Graphic results after the image matrix of datas had been analysed with the DBSCAN algorithm\n# Parameter: \n# Input: Image File\n# Output: \n#\tParameter1: Record k-means data analysis\n#\tParameter2: Record dbscan data analysis\n###########################################################################################################\n\ndef print_screen_results(image):\n\tCOLORMAP = 'pink'\n\tfig, axarr = plt.subplots(nrows=2, ncols=2, figsize=(50, 25))\n\taxarr[0,0].set_title('Body Scanner Image')\n\taxarr[0,0].imshow(image, cmap=COLORMAP)\n\taxarr[0,1].set_title('Adaptive Histogram Equalization')\n\taxarr[0,1].hist(image.flatten(), bins=256, color='c')\n\tdata_array = np.array(image)\n\twhitened = whiten(data_array)\n\tcodebook,distortion = kmeans(whitened,5)\n\taxarr[1,0].scatter(whitened[:, 0], whitened[:, 1], c='b')\n\taxarr[1,0].scatter(codebook[:, 0], codebook[:, 1], c='r')\n\tstream = 'k-means analyis: Distortion {0}%'.format(distortion)\n\taxarr[1,0].set_title(stream)\n\trecord_kmeans = pd.DataFrame([[' ', 0, distortion,'nothreat']],columns=['Name','Nth','Distortion','Threat'])\n\twhitened = whiten(data_array)\n\trecord = pd.DataFrame({'x_axes': whitened[:, 0],'y_axes':whitened[:, 1]})\n\tcenters = np.array(record)\n\tX, labels_true = make_blobs(n_samples=750, centers=centers, cluster_std=0.4, random_state=0)\n\tX = StandardScaler().fit_transform(X)\n\tdb = DBSCAN(eps=0.3, min_samples=10).fit(X)\n\tcore_samples_mask = np.zeros_like(db.labels_, dtype=bool)\n\tcore_samples_mask[db.core_sample_indices_] = True\n\tlabels = db.labels_\n\tn_clusters_ = len(set(labels)) - (1 if -1 in labels else 0)\n\tplt.subplot(224)\n\tunique_labels = set(labels)\n\tcolors = [plt.cm.Spectral(each)\tfor each in np.linspace(0, 1, len(unique_labels))]\n\tfor k, col in zip(unique_labels, colors):\n\t\tif k == -1:\n\t\t\tcol = [0, 0, 0, 1]\n\t\tclass_member_mask = (labels == k)\n\t\txy = X[class_member_mask & core_samples_mask]\n\t\tplt.plot(xy[:, 0], xy[:, 1], 'o', markerfacecolor=tuple(col),\n\t\t\tmarkeredgecolor='k', markersize=14)\n\t\txy = X[class_member_mask & ~core_samples_mask]\n\t\tplt.plot(xy[:, 0], xy[:, 1], 'o', markerfacecolor=tuple(col),\n\t\t\tmarkeredgecolor='k', markersize=6)\n\t\tName = ''\n\t\tnClusters = n_clusters_\n\t\tHomegeneity = metrics.homogeneity_score(labels_true, labels)\n\t\tCompleteness = \tmetrics.completeness_score(labels_true, labels)\n\t\tVmeasure = metrics.v_measure_score(labels_true, labels)\n\t\tARIndex = metrics.adjusted_rand_score(labels_true, labels)\n\t\tAMInformation = \tmetrics.adjusted_mutual_info_score(labels_true, labels)\n\t\tSilCoefficient = metrics.silhouette_score(X, labels)\n\t\tThreat = 'nothreat'\n\t\trecord_dbscan = pd.DataFrame([[Name, nClusters, Homegeneity, Completeness, Vmeasure, ARIndex, AMInformation, SilCoefficient, Threat]],columns=['Name','nClusters','Homegeneity','Completeness','Vmeasure','ARIndex','AMInformation','SilCoefficient','Threat'])\n\tplt.show()\n\treturn record_kmeans, record_dbscan\n","metadata":{"_cell_guid":"11fa2f6d-102f-47a1-9134-8f71b634b8a2","_uuid":"7028fa94cf47b61a7a0b7c57520bd8793cc5a909","collapsed":true,"jupyter":{"outputs_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def print_img(image,title='Image'):\n\t\n\tfig, axarr = plt.subplots(nrows=1, ncols=2, figsize=(20, 5))\n\taxarr[0].imshow(image, cmap='pink')\t\t\t#Put in the screen the image after adaptative histogram equialization\n\tplt.subplot(122)\n\tplt.hist(image.flatten(), bins=256, color='c')\n\tplt.title(title)\n\tplt.xlabel('Adaptive Histogram Equalization')\n\tplt.ylabel('Frequency')\n\tplt.show()\n\t\n\treturn image","metadata":{"_cell_guid":"ecf9458b-5fa7-47b5-9c9e-ea92d7ccdf7f","_uuid":"4d0de8b8f115bca17d281301d554ff64f0ce043e","collapsed":true,"jupyter":{"outputs_hidden":true}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Results SECTION 1: Prediction with svm algorithm**\n\nThe first stage of our study is the generation of a complete training dataset with information about the possible threats detected in image. We are going to generate different kind of dataset (csv files) with distortion of the image parameters with respect to the mean of typical parameters obtained in images that we are checked that it hasn't presented any kind of threat using the next field parameters: \n\n      ['Name','Nth','Difference','Distortion', 'Mean', 'Threat']\n \nWith the application of k-means algorithm we can obtain the distortion with respect to 5 cluster centroid in the measurement, this parameter it's analysed with respect to the sum of the mean value and an initial threshold, when the parameter it's out of this margins we are considered that we could have a possible anomaly that it's necessary to analyse deeply. \n\nTo check the training dataset we are going to use a classification method (support vector machine) to do a prediction about of possibles anomalies in the image, and detect a threat in the passenger. \n\nThis is a very simple approach with ncluster, 5 initial cluster\n","metadata":{"_cell_guid":"bc58a94e-3bff-4cc7-8e1f-2ca75cb49a1e","_uuid":"abcd9b1826a70c4dfc17173a9d9111bf093cf238"}},{"cell_type":"code","source":"if __name__ == \"__main__\":\n\tprint (\"DATA_ANALISIS_TOOL: Initialization of the system...\")\n\tdistortion_indicator1=0\n\ttry:\n\t\tdataset = pd.read_csv(NAME_FILE_MODEL_KMEANS, sep=',')\n\t\tn_samp = len(dataset)\n\t\tX, y = make_classification(n_samples=n_samp, n_features=4, random_state=0)\n\t\trow = 0\n\t\tfor index in range(len(dataset)):\n\t\t\tX[row][0]=dataset.loc[index,'Nth']\n\t\t\tX[row][1]=dataset.loc[index,'Distortion']\n\t\t\tX[row][2]=dataset.loc[index,'Difference']\n\t\t\tX[row][3]=dataset.loc[index,'Mean']\n\t\t\tif dataset.loc[index,'Threat'] == 'threat':\n\t\t\t\ty[index]=np.int32(1)\n\t\t\telse:\n\t\t\t\ty[index]=np.int32(0)\n\t\t\t\n\t\t\trow = row + 1\n\t\tclf = LinearSVC(random_state=0)\n\t\tclf.fit(X, y)\n\t\tdataset_new = pd.read_csv(DATA_FOR_PREDICTION_FILE, sep=',')\n\t\tdata_array = np.array(dataset_new)\n\t\tdistortion,whitened,codebook = cluster_analisys(data_array)\n\t\tmean_value = dataset['Distortion'].mean()\n\t\tdifference = abs(mean_value - distortion)\n\t\tdistortion_indicator1 = difference / mean_value * 100\n\t\tnth = 3\n\t\tresult_of_prediction = clf.predict([[nth,distortion,difference,mean_value]]) \t\n\t\tname_file = DATA_FOR_PREDICTION_FILE\n\t\tname_passenger = str(name_file).split('/')\n\t\tname_passenger = str(name_passenger[2]).split('.')\n\t\tname_passenger = name_passenger[0]\n\t\t\n\t\tif result_of_prediction == True:\n\t\t\tprint(\"Passenger {0} it could suppose a thread at Body Scanner {1}\".format(name_passenger,3))\n\t\t\tthreat_result = True\n\t\telse:\n\t\t\tprint(\"Passenger {0} it could suppose a thread at Body Scanner {1}\".format(namename_passenger,3))\n\texcept Exception as exception:\n\t\tprint (\"DATA_ANALISIS_TOOL: Excepcion {0}\".format(exception))\n\tfinally:\n\t\tprint(\"Image Distortion {0}\".format(distortion_indicator1))    ","metadata":{"_cell_guid":"4d228864-dabb-43dd-9c1e-2297e5d15614","_uuid":"591e1a1258f7a1700a99059875ffd1f5f721fa3f"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Results SECTION 2: Prediction with svm algorithm over image with high contrast**\n\nIn a second approach we are detect that any anomalies could be produced by images with less quality. At this point we are going to apply to cleaning the datainput initial, based in an increase of high quality image with a filtering of the image datainput (contrast filtering techniques)\n\nWhen we are tested the new results applying the same datainput and passed this input data across our support vector machine we can check in the output data that the results are better because we are reduced the distortion of the image and the model.","metadata":{"_cell_guid":"86f56207-514e-4a96-b1ed-12308801a9bd","_uuid":"eb6bed4954de36f5367be4a14fbfbe6da5e2cb63"}},{"cell_type":"code","source":"try:\n\tname_file = DATA_FOR_PREDICTION_FILE_FILTERED\n\tdataset_new = pd.read_csv(DATA_FOR_PREDICTION_FILE_FILTERED, sep=',')\n\tdata_array = np.array(dataset_new)\n\tdistortion,whitened,codebook = cluster_analisys(data_array)\n\tmean_value = dataset['Distortion'].mean()\n\tdifference = abs(mean_value - distortion)\n\tdistortion_indicator2 = difference / mean_value * 100\n\tnth = 3\n\tresult_of_prediction = clf.predict([[nth,distortion,difference,mean_value]]) \t\n\tname_file = DATA_FOR_PREDICTION_FILE_FILTERED\n\tname_passenger = str(name_file).split('/')\n\tname_passenger = str(name_passenger[2]).split('.')\n\tname_passenger = name_passenger[0]\n\tif result_of_prediction == True:\n\t\tprint(\"Passenger {0} it could suppose a thread at Body Scanner {1}\".format(name_passenger,3))\n\t\tthreat_result = True\n\telse:\n\t\tprint(\"Passenger {0} it could suppose a thread at Body Scanner {1}\".format(name_passenger,3))\n\t\n\nexcept Exception as exception:\n\tprint (\"DATA_ANALISIS_TOOL: Excepcion {0}\".format(exception))\nfinally:\n\tprint(\"Image Distortion {0}\".format(distortion_indicator2)) \n","metadata":{"_cell_guid":"6681d663-5eae-47be-b71a-03378efd5119","_uuid":"656d24466b0910ee405f36d37e96a60f6e6b2557"},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Improvement in detection images with filtering** \n\nAlso we are try to apply any techniques of DBSCAN for identify possibles case of anomalies in the image, but this tasks required more time and the purposes of this practice for me it's only the application of the knowledge acquired in the course, with a simple method of clustering to find mistakes and possible anomalies that it could think us that the passenger could have any kind of threat object. With a simple application of a method of classification for the prediction such as support vector machine I think we could achieve our initial goals joining the clustering analysis and support vector machine ( classification method) in a useful programme for detect possible threats of passengers in the control of the airports.   \n\nThe improvement is found in the development of the practice with the application of image's cleaning techniques filtering obtaining a better difference results with respect the initial results around the mean value of image distortion, with an improvement in the distortion of image around of 0,61 reflected in the precession of the possible prediction.\n\n","metadata":{"_cell_guid":"6a1c99c5-8c95-4e45-b220-8bd4ab4610b4","_uuid":"b13273cbbeb8389225492beffbb23aa9c5cb7812"}},{"cell_type":"code","source":"improvement = distortion_indicator2 - distortion_indicator1\nprint(\"Image distortion improvement with the filtering image {0}% \".format(improvement))","metadata":{"_cell_guid":"20e0deb5-34f0-44b5-bb47-bc110f370904","_uuid":"76a7a004212d222de5377b9a147c3f3b56c03cb4"},"execution_count":null,"outputs":[]}]}