{"nbformat": 4, "nbformat_minor": 1, "cells": [{"cell_type": "markdown", "metadata": {"_cell_guid": "7b50a651-4800-4c8f-a1d2-1544dfd69a3a", "_uuid": "ea970f227684bce1899d1388e8844b5ac8019961"}, "source": ["# Open tensorflow kernel with CLI download, Multi-GPU support and much more\n", "\n", "[Code @GitHub](https://github.com/Cognexa/cdiscount-kernel)\n", "\n", "Just run a VGG-like convnet baseline while you analyze the data!\n", "\n", "Works on Linux with Python 3.5+.\n", "\n", "Features:\n", "- CLI data download\n", "- Data validation with SHA256 hash\n", "- Simple data visualization\n", "- Train-Valid splitting\n", "- Low memory footprint data streams\n", "- Base VGG-like convnet\n", "- Multi-GPU training with a single argument!\n", "- TensorBoard training tracking\n", "\n", "## Quick start\n", "Install tensorflow and 7z.\n", "\n", "Clone repo and install the requirements\n", "```\n", "git clone https://github.com/Cognexa/cdiscount-kernel && cd cdiscount-kernel\n", "pip3 install -r requirements.txt --user\n", "```\n", "\n", "Download dataset with kaggle-cli (this may take a while, 3 hours in my case)\n", "```\n", "# requires >57Gb of free space\n", "KG_USER=\"<YOUR KAGGLE USERNAME\" KG_PASS=\"<YOUR KAGGLE PASSWORD>\" cxflow dataset download cdc\n", "```\n", "\n", "Or if you have downloaded the data earlier:\n", "```\n", "mkdir data\n", "# mv/cp your etracted files to data directory\n", "```\n", "\n", "Validate your download and see the example data:\n", "```\n", "# in the root directory (cdiscount-kernel)\n", "cxflow dataset validate cdc\n", "cxflow dataset show cdc\n", "# now see the newly created visual directory\n", "```\n", "\n", "Create a random validation split with 10% of the data and start training:\n", "```\n", "cxflow dataset split cdc\n", "cxflow train cdc model.n_gpus=<NUMBER OF GPUS TO USE>\n", "```\n", "\n", "Observe the training with TensorBoard (note: a summary is written only after each epoch)\n", "```\n", "tensorboard --logdir=log\n", "```\n", "\n", "## UPDATE [LB 0.65]\n", "**important:** update **cxflow** and **cxflow-tensorflow** with `pip3 install cxflow cxflow-tensorflow --user --upgrade`\n", "\n", "Main features:\n", "- XCeption net (https://arxiv.org/abs/1610.02357)\n", "- Fast random data access\n", "\n", "Resize the data to `dataset.size` with (this may take a few hours)\n", "```\n", "cxflow dataset resize cdc/xception.yaml\n", "cxflow dataset split cdc/xception.yaml\n", "```\n", "\n", "Run the training with\n", "```\n", "cxflow train cdc/xception.yaml\n", "```\n", "\n", "Training procedure that reached 0.65:\n", "- Train with original size, LR 0.0001, 4 middle flow repeats until stalled\n", "- Fine-tune with 128x128, LR 0.0001, 0.5 dropout, 0.00001 weight decay until stalled\n", "- Fine-tune as above but with LR 0.00001 (10x smaller)\n", "\n", "Tips:\n", "- Use small images right away\n", "- The final GlobalAveragePooling may be a bottleneck\n", "- Net does not overfit so far, no augmentations needed\n", "\n", "## Example output:\n", "```\n", "2017-09-16 00:22:14.000262: INFO    @common         : Creating dataset\n", "2017-09-16 00:22:14.000776: INFO    @common         : \tCDCNaiveDataset created\n", "2017-09-16 00:22:14.000777: INFO    @common         : Creating a model\n", "2017-09-16 00:22:20.000724: INFO    @model          : \tCreating TF model on 2 GPU devices\n", "2017-09-16 00:22:21.000362: INFO    @cdc_net        : Flatten shape `(?, 8192)`\n", "2017-09-16 00:22:21.000387: INFO    @cdc_dataset    : Loading metadata\n", "2017-09-16 00:22:26.000826: INFO    @cdc_net        : Output shape `(?, 5270)`\n", "2017-09-16 00:22:26.000893: INFO    @cdc_net        : Flatten shape `(?, 8192)`\n", "2017-09-16 00:22:26.000901: INFO    @cdc_net        : Output shape `(?, 5270)`\n", "2017-09-16 00:22:29.000351: INFO    @common         : \tCDCNaiveNet created\n", "2017-09-16 00:22:29.000354: INFO    @common         : Creating hooks\n", "2017-09-16 00:22:29.000355: INFO    @common         : \tShowProgress created\n", "2017-09-16 00:22:29.000355: INFO    @common         : \tComputeStats created\n", "2017-09-16 00:22:29.000355: INFO    @common         : \tLogVariables created\n", "2017-09-16 00:22:29.000356: INFO    @common         : \tLogProfile created\n", "2017-09-16 00:22:29.000356: INFO    @common         : \tSaveEvery created\n", "2017-09-16 00:22:29.000356: INFO    @common         : \tSaveBest created\n", "2017-09-16 00:22:29.000357: INFO    @common         : \tCatchSigint created\n", "2017-09-16 00:22:29.000357: INFO    @common         : \tStopAfter created\n", "2017-09-16 00:22:30.000968: INFO    @common         : \tWriteTensorBoard created\n", "2017-09-16 00:22:30.000968: INFO    @common         : Creating main loop\n", "2017-09-16 00:22:30.000968: INFO    @common         : Running the main loop\n", "2017-09-16 03:13:01.000457: INFO    @log_variables  : After epoch 1\n", "2017-09-16 03:13:01.000457: INFO    @log_variables  : \ttrain loss mean: 4.243194\n", "2017-09-16 03:13:01.000457: INFO    @log_variables  : \ttrain accuracy mean: 0.320435\n", "2017-09-16 03:13:01.000457: INFO    @log_variables  : \tvalid loss mean: 3.313541\n", "2017-09-16 03:13:01.000457: INFO    @log_variables  : \tvalid accuracy mean: 0.434122\n", "2017-09-16 03:13:03.000486: INFO    @save           : Model saved to: ./log/CDCNaiveNet_2017-09-16-00-22-14_ngz6u4_b/model_1.ckpt\n", "2017-09-16 03:13:05.000217: INFO    @save           : Model saved to: ./log/CDCNaiveNet_2017-09-16-00-22-14_ngz6u4_b/model_best.ckpt\n", "2017-09-16 03:13:05.000219: INFO    @log_profile    : \tT read data:\t1594.948686\n", "2017-09-16 03:13:05.000219: INFO    @log_profile    : \tT train:\t8347.117133\n", "2017-09-16 03:13:05.000219: INFO    @log_profile    : \tT eval:\t282.549242\n", "2017-09-16 03:13:05.000219: INFO    @log_profile    : \tT hooks:\t8.592250\n", "2017-09-16 03:13:05.000219: INFO    @main_loop      : Epochs done: 1\n", "2017-09-16 06:03:17.000103: INFO    @log_variables  : After epoch 2\n", "2017-09-16 06:03:17.000103: INFO    @log_variables  : \ttrain loss mean: 2.952012\n", "2017-09-16 06:03:17.000103: INFO    @log_variables  : \ttrain accuracy mean: 0.480100\n", "2017-09-16 06:03:17.000103: INFO    @log_variables  : \tvalid loss mean: 2.863293\n", "2017-09-16 06:03:17.000104: INFO    @log_variables  : \tvalid accuracy mean: 0.496674\n", "2017-09-16 06:03:18.000840: INFO    @save           : Model saved to: ./log/CDCNaiveNet_2017-09-16-00-22-14_ngz6u4_b/model_2.ckpt\n", "2017-09-16 06:03:20.000762: INFO    @save           : Model saved to: ./log/CDCNaiveNet_2017-09-16-00-22-14_ngz6u4_b/model_best.ckpt\n", "2017-09-16 06:03:20.000764: INFO    @log_profile    : \tT read data:\t1581.478134\n", "2017-09-16 06:03:20.000764: INFO    @log_profile    : \tT train:\t8342.576470\n", "2017-09-16 06:03:20.000764: INFO    @log_profile    : \tT eval:\t281.916230\n", "2017-09-16 06:03:20.000764: INFO    @log_profile    : \tT hooks:\t8.502520\n", "2017-09-16 06:03:20.000764: INFO    @main_loop      : Epochs done: 2\n", "\n", "...```\n", "\n", "\n", "## About\n", "This kernel is written in [cxflow-tensorflow](https://github.com/Cognexa/cxflow-tensorflow), a plugin for [cxflow](https://github.com/Cognexa/cxflow) framework. Make sure you check it out!\n", "\n", "A simple submission script will be added soon, stay tuned!\n"]}, {"cell_type": "markdown", "metadata": {"_cell_guid": "435ac892-ee6a-485d-8fd4-c0ba83e69cac", "_uuid": "48f6f5607a4603b22bb1406967dee6caaac546bc"}, "source": []}], "metadata": {"language_info": {"pygments_lexer": "ipython3", "mimetype": "text/x-python", "codemirror_mode": {"name": "ipython", "version": 3}, "name": "python", "nbconvert_exporter": "python", "version": "3.6.3", "file_extension": ".py"}, "kernelspec": {"name": "python3", "language": "python", "display_name": "Python 3"}}}