{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":14774,"databundleVersionId":875431,"sourceType":"competition"},{"sourceId":46552,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":39011},{"sourceId":46553,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":39012},{"sourceId":46587,"sourceType":"modelInstanceVersion","isSourceIdPinned":true,"modelInstanceId":39041}],"dockerImageVersionId":30699,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Table Of Contents","metadata":{}},{"cell_type":"markdown","source":"## Preparation <a id=\"2\"></a>","metadata":{}},{"cell_type":"markdown","source":"## Metric (Quadratic Weighted Kappa) <a id=\"3\"></a>","metadata":{}},{"cell_type":"markdown","source":"## EDA (Exploratory Data Analysis) <a id=\"4\"></a>","metadata":{}},{"cell_type":"markdown","source":"For EDA on image datasets I think one should at least examine the label distribution, the images before preprocessing and the images after preprocessing. Through examining these three aspects we can get a good sense of the problem. Note that the distribution on the test set can still vary wildly from the training data.","metadata":{}},{"cell_type":"markdown","source":"## Preprocessing <a id=\"5\"></a>","metadata":{}},{"cell_type":"markdown","source":"## Modeling (EfficientNetB5) <a id=\"6\"></a>","metadata":{}},{"cell_type":"markdown","source":"Thanks to the amazing wrapper by [qubvel](https://github.com/qubvel/efficientnet) we can load in a model like the Keras API. We specify the input shape and that we want the model without the top (the final Dense layer). Then we load in the weights which are provided in [this Kaggle dataset](https://www.kaggle.com/ratthachat/efficientnet-keras-weights-b0b5). Note that we will use the [RAdam optimizer](https://arxiv.org/pdf/1908.03265v1.pdf) since it often yields better convergence than Vanilla Adam. Thanks to CyberZHG who implemented [RAdam for Keras](https://github.com/CyberZHG/keras-radam/blob/master/keras_radam/optimizers.py).","metadata":{}},{"cell_type":"markdown","source":"## Evaluation <a id=\"7\"></a>","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv('/kaggle/input/b5tens/tensorflow2/csv/1/submission.csv')","metadata":{"execution":{"iopub.status.busy":"2024-05-11T17:16:13.427440Z","iopub.execute_input":"2024-05-11T17:16:13.427878Z","iopub.status.idle":"2024-05-11T17:16:13.442714Z","shell.execute_reply.started":"2024-05-11T17:16:13.427842Z","shell.execute_reply":"2024-05-11T17:16:13.441662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df\n","metadata":{"execution":{"iopub.status.busy":"2024-05-11T17:16:13.891758Z","iopub.execute_input":"2024-05-11T17:16:13.892259Z","iopub.status.idle":"2024-05-11T17:16:13.905193Z","shell.execute_reply.started":"2024-05-11T17:16:13.892223Z","shell.execute_reply":"2024-05-11T17:16:13.904002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(df.shape[0]):\n    df.loc[i,'id_code'] = df.loc[i,'id_code'].split('.')[0]","metadata":{"execution":{"iopub.status.busy":"2024-05-11T17:16:14.770265Z","iopub.execute_input":"2024-05-11T17:16:14.770641Z","iopub.status.idle":"2024-05-11T17:16:15.243271Z","shell.execute_reply.started":"2024-05-11T17:16:14.770614Z","shell.execute_reply":"2024-05-11T17:16:15.242132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['diagnosis'] = df['diagnosis'].astype('int64')\ndf['id_code']=df['id_code'].astype(\"string\")","metadata":{"execution":{"iopub.status.busy":"2024-05-11T17:16:15.790395Z","iopub.execute_input":"2024-05-11T17:16:15.790813Z","iopub.status.idle":"2024-05-11T17:16:15.798783Z","shell.execute_reply.started":"2024-05-11T17:16:15.790779Z","shell.execute_reply":"2024-05-11T17:16:15.797524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check submission\nprint(\"Submission File\")\n\ndf.to_csv('submission.csv', index=False)\ndisplay(df.head())","metadata":{"execution":{"iopub.status.busy":"2024-05-11T17:16:18.436985Z","iopub.execute_input":"2024-05-11T17:16:18.437508Z","iopub.status.idle":"2024-05-11T17:16:18.462438Z","shell.execute_reply.started":"2024-05-11T17:16:18.437458Z","shell.execute_reply":"2024-05-11T17:16:18.461274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}