{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"# Use Both Image and Tabular Data\n\nKaggle's Melanoma Classification competition provides both **image data** and **tabular data** about each sample. Our task is to use both types of data to predict the probability that a sample is malignant. How can we build a model that uses both **images** and **tabular data**?\n\nThree ideas come to mind.\n\n* Build a CNN image model and find a way to input the tabular data into the CNN image model\n* Build a Tabular data model and find a way to extract image embeddings and input into the Tabular data model\n* Build 2 separate models and ensemble\n\nIn this notebook, we explore the third idea. A model that uses only image data is [here][1] and scores LB 0.910. A model that uses only tabular data is [here][2] and scores LB 0.700. We will make a simple ensemble of the two and thus utilize both the provided image and provided tabular data. The resultant LB score is 0.915 demonstrating that each type of data adds additional information.\n\nTo learn more ideas about how to build models that utilize both image and tabular data or to share your ideas, please participate in the Kaggle discussion [here][3].\n\n\n[1]: https://www.kaggle.com/ajaykumar7778/melanoma-tpu-efficientnet-b5-dense-head\n[2]: https://www.kaggle.com/titericz/simple-baseline\n[3]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155251","execution_count":null},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_sub = pd.read_csv('../input/melanomapreds/submissionImage.csv')\ntabular_sub = pd.read_csv('../input/melanomapreds/submissionTabular.csv')\ntabular_sub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub = image_sub.copy()\nsub.target = 0.9 * image_sub.target.values + 0.1 * tabular_sub.target.values\nsub.to_csv('submission.csv',index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.hist(sub.target,bins=100)\nplt.ylim((0,100))\nplt.show()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}