{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Acknowledgements\nhttps://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask\n- This provides the most simple submission possible, we use the mask on the test images as prediction.\n\nhttps://www.kaggle.com/code/yoyobar/2-5d-segmentaion-model-with-rotate-tta\n- Highest current puclic score\n\n# Why?\n\nAfter reviewing several public kernels + my own experiments with different NN configurations and their cross validation results, I'm leaning to a conclusion that none of them is learning an actual ink. \nWhat they rather learn are cracks and sides of papirus and other places where ink is unlikely to contain ink. By excluding such regions and including everything else you can get pretty good score. I'm not convinced based on just scores on public leaderboard that even top solutions learn to find an actual ink.\nIt is remarkable that public kernels (including baseline) usually select for cross validation parts of data where letters are almost visible by naked eye and easiest to learn. By changing only validation region and rerunning their results I was able to see that seemingly positive validation results are rather coinsidence than a result of useful knowledge gained.\n\n[https://www.kaggle.com/code/lucasvw/0-11-simplest-possible-solution-submit-testmask] - Thanks to this simple idea it becomes clear that one can figure out some pretty simple heuristics to get pretty good score not only without ML, but not even looking at CT scans!\nTo illustrate my point, I use one of many possible rules:\n- Areas near the border of valid mask are less likely to contain ink and more likely to contain various artifacts: cracks, etc.\n\nSo when you see your model getting 0.28 or even 0.52 cross validation score don't get too excited.\n\nTake a look at this!","metadata":{}},{"cell_type":"code","source":"from fastai.vision.all import *\nfrom tqdm.auto import tqdm, trange\nimport matplotlib.pyplot as plt\nfrom scipy.ndimage import distance_transform_edt\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2023-05-15T18:24:52.016815Z","iopub.execute_input":"2023-05-15T18:24:52.018246Z","iopub.status.idle":"2023-05-15T18:24:58.319022Z","shell.execute_reply.started":"2023-05-15T18:24:52.018190Z","shell.execute_reply":"2023-05-15T18:24:58.317625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"base_path = Path('/kaggle/input/vesuvius-challenge-ink-detection')\nPath.BASE_PATH = base_path\ntrain_path = base_path / 'train'\ntest_path = base_path / 'test'","metadata":{"execution":{"iopub.status.busy":"2023-05-15T18:24:58.321680Z","iopub.execute_input":"2023-05-15T18:24:58.322065Z","iopub.status.idle":"2023-05-15T18:24:58.329047Z","shell.execute_reply.started":"2023-05-15T18:24:58.322030Z","shell.execute_reply":"2023-05-15T18:24:58.327943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_files = get_files(test_path, extensions='.png')\nimage_files.sort()\nimage_files","metadata":{"execution":{"iopub.status.busy":"2023-05-15T18:24:58.330656Z","iopub.execute_input":"2023-05-15T18:24:58.331336Z","iopub.status.idle":"2023-05-15T18:24:58.401880Z","shell.execute_reply.started":"2023-05-15T18:24:58.331295Z","shell.execute_reply":"2023-05-15T18:24:58.400806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def rle(output):\n    flat_img = np.where(output.flatten() > 0.4, 1, 0).astype(np.uint8)\n    starts = np.array((flat_img[:-1] == 0) & (flat_img[1:] == 1))\n    ends = np.array((flat_img[:-1] == 1) & (flat_img[1:] == 0))\n    starts_ix = np.where(starts)[0] + 2\n    ends_ix = np.where(ends)[0] + 2\n    lengths = ends_ix - starts_ix\n    return \" \".join(map(str, sum(zip(starts_ix, lengths), ())))","metadata":{"execution":{"iopub.status.busy":"2023-05-15T18:24:58.403170Z","iopub.execute_input":"2023-05-15T18:24:58.403504Z","iopub.status.idle":"2023-05-15T18:24:58.413542Z","shell.execute_reply.started":"2023-05-15T18:24:58.403472Z","shell.execute_reply":"2023-05-15T18:24:58.412250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = defaultdict(list)\nfig, axs = plt.subplots(1,len(image_files), figsize=(10,5))\n\nfor i, fragment_name in enumerate(image_files):\n    submission[\"Id\"].append(fragment_name.parent.name)\n    res = np.array(Image.open(fragment_name))\n    res = distance_transform_edt(np.pad(res,1))[1:-1,1:-1]\n    res = res > 700\n    axs[i].imshow(res)\n    submission[\"Predicted\"].append(rle(res))","metadata":{"execution":{"iopub.status.busy":"2023-05-15T18:30:48.290681Z","iopub.execute_input":"2023-05-15T18:30:48.291970Z","iopub.status.idle":"2023-05-15T18:30:59.792453Z","shell.execute_reply.started":"2023-05-15T18:30:48.291904Z","shell.execute_reply":"2023-05-15T18:30:59.790876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.DataFrame.from_dict(submission)\nprint(df)\ndf.to_csv(\"/kaggle/working/submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-05-15T18:25:02.761732Z","iopub.status.idle":"2023-05-15T18:25:02.762340Z","shell.execute_reply.started":"2023-05-15T18:25:02.762090Z","shell.execute_reply":"2023-05-15T18:25:02.762119Z"},"trusted":true},"execution_count":null,"outputs":[]}]}