{
  "id": 497545,
  "title": "Interpretable ML. Algorithm Unrolling. Whoosh (Python Search Library) on  Kaggle .",
  "url": "/competitions/uspto-explainable-ai/discussion/497545",
  "author_name": "",
  "post_date": "2024-04-25T01:44:41.887469100Z",
  "votes": 22,
  "comment_count": 5,
  "views": 0,
  "content": "<h1>Whoosh (Python Search Library) on Kaggle Notebooks</h1>\n<p>Sad fact, many of these Whoosh code were made during Covid19 times (4y ago) to retrieve Scientific Papers.</p>\n<p>By Daniel Wolffram <a href=\"https://www.kaggle.com/code/danielwolffram/whoosh-search\" target=\"_blank\">Whoosh Search</a></p>\n<p>By Max Feinberg <a href=\"https://www.kaggle.com/code/mxfeinberg/using-whoosh-for-indexing-and-querying\" target=\"_blank\">Using Whoosh for Indexing and Querying</a></p>\n<p>By Alexander Rubin <a href=\"https://www.kaggle.com/code/arubin/covid19-build-fulltext-indexes-with-whoosh\" target=\"_blank\">Covid19 - build fulltext indexes with Whoosh</a></p>\n<p>By Michael C Gold <a href=\"https://www.kaggle.com/code/michaelcgold/covid-19-origins-in-a-whoosh\" target=\"_blank\">https://www.kaggle.com/code/michaelcgold/covid-19-origins-in-a-whoosh</a></p>\n<p>By Leire <a href=\"https://www.kaggle.com/code/leireher/rag-for-triviaqa\" target=\"_blank\">RAG for TriviaQA</a></p>\n<p>By Sohier Dane <a href=\"https://www.kaggle.com/code/sohier/basic-whoosh-search-demo\" target=\"_blank\">Basic Whoosh search demo</a></p>\n<h1>On GitHub: ScispaCy</h1>\n<p>\"ScispaCy is a Python package containing spaCy models for processing biomedical, scientific or clinical text.\"</p>\n<h1>Interpretability: Definitions, Methods and Applications</h1>\n<p>Interpretable machine learning: definitions, methods, and applications</p>\n<p>Authors: W. James Murdocha, Chandan Singhb, Karl Kumbiera, Reza Abbasi-Aslb, and Bin Yua</p>\n<p>Defining interpretable machine learning.</p>\n<p>\"On its own, interpretability is a broad, poorly defined concept. Taken to its full generality, to interpret data means to extract information (of some form) from it. The set of methods falling under this umbrella spans everything from designing an initial experiment to visualizing final results. In this overly general form, interpretability is not substantially different from the established concepts of data science and applied statistics.\"</p>\n<p>Post hoc analysis</p>\n<p>\"Having fit a model (or models), the practitioner then analyzes it for answers to the original question. The process of analyzing the model often involves using interpretability methods to extract various (stable) forms of information from the model. The extracted information can then be analyzed and displayed using standard data analysis methods, such as scatter plots and histograms. The ability of the interpretations to properly describe what the model has learned is denoted by descriptive accuracy.\"</p>\n<p>Demonstrating relevancy to real-world problems.</p>\n<p>\"Another angle for developing improved interpretation methods is to improve the relevancy of interpretations for some audience or problem. This is normally done by introducing a novel form of output, such as feature heatmaps, rationales, feature hierarchies or identifying important elements in the training set.\"</p>\n<p><a href=\"https://arxiv.org/pdf/1901.04592.pdf\" target=\"_blank\">https://arxiv.org/pdf/1901.04592.pdf</a></p>\n<h1>The black-box Nature</h1>\n<p>Interpretability of Machine Learning: Recent Advances and Future Prospects</p>\n<p>Authors: Gao, Lei and Guan, Ling</p>\n<p>\"The proliferation of machine learning (ML) has drawn unprecedented interest in the study of various multimedia contents such as text, image, audio and video, among others. Consequently,understanding and learning ML-based representations have taken center stage in knowledge discovery in intelligent multimedia research and applications. Nevertheless, the black-box nature of contemporary ML, especially in deep neural networks (DNNs), has posed a primary challenge for ML-based representation learning. To address this black-box problem, the studies on interpretability of ML have attracted tremendous interests in recent years.\"</p>\n<p>Algorithm Unrolling</p>\n<p>\"Algorithm unrolling solves model interpretability by providing a concrete and systematic connection between iterative algorithms that are widely used in signal process ing and DNNs. Given an iterative algorithm, a corresponding deep network is generated by cascading its iterations h. Then, iteration step h is executed a number of times, resulting in different parameters h1, h2,… Each iteration h depends on algorithm parameters, which are transferred into network parameters 1, 2,… Instead of determining parameters through cross-validation or analytical derivations, the parameters 1, 2,… are learned from training datasets through end-to-end training. In this way, the network layers naturally inherit interpretability from the iteration procedure.\"</p>\n<p><a href=\"https://arxiv.org/pdf/2305.00537.pdf\" target=\"_blank\">https://arxiv.org/pdf/2305.00537.pdf</a></p>\n<h1>Algorithm Unrolling</h1>\n<p>Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing</p>\n<p>Authors: Vishal Monga, Senior Member, Yuelong Li and Yonina C. Eldar</p>\n<p>\"Deep neural networks provide unprecedented performance gains in many real world problems in signal and image processing. Despite these gains, future development and practical deployment of deep networks is hindered by their black box nature (i.e.lack of interpretability, and by the need for very large training sets).\"</p>\n<p>\"An emerging technique called algorithm unrolling or unfolding offers promise in eliminating these issues (i.e.lack of interpretability, and by the need for very large training sets) by providing a concrete and systematic connection between iterative algorithms that are used widely in signal processing and deep neural networks.\"</p>\n<p>\"Unrolling methods were first proposed to develop fast neural network approximations for sparse coding. More recently, this direction has attracted enormous attention and is rapidly growing both in theoretic investigations and practical applications. The growing popularity of unrolled deep networks is due in part to their potential in developing efficient, high-performance and yet interpretable network architectures from reasonable size training sets.\"</p>\n<p>\"In this article, the authors reviewed algorithm unrolling for signal and image processing. They extensively covered popular techniques for algorithm unrolling in various domains of signal and image processing including imaging, vision and recognition, and speech processing. By reviewing previous works, they revealed the connections between iterative algorithms and neural networks and present recent theoretical results. Finally, the authors provided a discussion on current limitations of unrolling and suggest possible future research directions.\"</p>\n<p>Unrolling Sparse Coding Algorithms into Deep Networks</p>\n<p>\"The earliest work in algorithm unrolling dates back to Gregor et al.’s paper (2010) on improving the computational efficiency of sparse coding algorithms through end-to-end training. In particular, they discussed how to improve the efficiency of the Iterative Shrinkage and Thresholding Algorithm (ISTA), one of the most popular approaches in sparse coding. Learned ISTA. Each iteration of ISTA comprises one linear operation followed by a non-linear soft-thresholding operation, which mimics the ReLU activation function. A diagram representation of one iteration step reveals its resemblance to network is dubbed Learned ISTA (LISTA).\"</p>\n<p>Bridging the Gap between Theory and Practice:</p>\n<p>\"While substantial progress has been achieved towards understanding the network behavior through unrolling, more works need to be done to thoroughly understand its mechanism. Although the effectiveness of some networks on image reconstruction tasks has been explained somehow by drawing parallels to sparse coding algorithms, it is still mysterious why state-of-the art networks perform well on various recognition tasks. Further more, unfolding itself is not uniquely defined. For instance, there are multiple ways to choose the underlying iterative algorithms, to decide what parameters become trainable and what parameters to fix, and more.\"</p>\n<p>\"Another interesting direction is to develop a theory that provides guidance for practical applications. For instance, it is interesting to perform analysis that guide practical network design choices, such as dimensions of parameters, network depth, etc. It is particularly interesting to identify factors that have high impact on network performance.\"</p>\n<p><a href=\"https://arxiv.org/pdf/1912.10557.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.10557.pdf</a></p>",
  "messages": [
    {
      "id": "2773953",
      "postDate": "04/25/2024 01:44:41",
      "content": "<h1>Whoosh (Python Search Library) on Kaggle Notebooks</h1>\n<p>Sad fact, many of these Whoosh code were made during Covid19 times (4y ago) to retrieve Scientific Papers.</p>\n<p>By Daniel Wolffram <a href=\"https://www.kaggle.com/code/danielwolffram/whoosh-search\" target=\"_blank\">Whoosh Search</a></p>\n<p>By Max Feinberg <a href=\"https://www.kaggle.com/code/mxfeinberg/using-whoosh-for-indexing-and-querying\" target=\"_blank\">Using Whoosh for Indexing and Querying</a></p>\n<p>By Alexander Rubin <a href=\"https://www.kaggle.com/code/arubin/covid19-build-fulltext-indexes-with-whoosh\" target=\"_blank\">Covid19 - build fulltext indexes with Whoosh</a></p>\n<p>By Michael C Gold <a href=\"https://www.kaggle.com/code/michaelcgold/covid-19-origins-in-a-whoosh\" target=\"_blank\">https://www.kaggle.com/code/michaelcgold/covid-19-origins-in-a-whoosh</a></p>\n<p>By Leire <a href=\"https://www.kaggle.com/code/leireher/rag-for-triviaqa\" target=\"_blank\">RAG for TriviaQA</a></p>\n<p>By Sohier Dane <a href=\"https://www.kaggle.com/code/sohier/basic-whoosh-search-demo\" target=\"_blank\">Basic Whoosh search demo</a></p>\n<h1>On GitHub: ScispaCy</h1>\n<p>\"ScispaCy is a Python package containing spaCy models for processing biomedical, scientific or clinical text.\"</p>\n<h1>Interpretability: Definitions, Methods and Applications</h1>\n<p>Interpretable machine learning: definitions, methods, and applications</p>\n<p>Authors: W. James Murdocha, Chandan Singhb, Karl Kumbiera, Reza Abbasi-Aslb, and Bin Yua</p>\n<p>Defining interpretable machine learning.</p>\n<p>\"On its own, interpretability is a broad, poorly defined concept. Taken to its full generality, to interpret data means to extract information (of some form) from it. The set of methods falling under this umbrella spans everything from designing an initial experiment to visualizing final results. In this overly general form, interpretability is not substantially different from the established concepts of data science and applied statistics.\"</p>\n<p>Post hoc analysis</p>\n<p>\"Having fit a model (or models), the practitioner then analyzes it for answers to the original question. The process of analyzing the model often involves using interpretability methods to extract various (stable) forms of information from the model. The extracted information can then be analyzed and displayed using standard data analysis methods, such as scatter plots and histograms. The ability of the interpretations to properly describe what the model has learned is denoted by descriptive accuracy.\"</p>\n<p>Demonstrating relevancy to real-world problems.</p>\n<p>\"Another angle for developing improved interpretation methods is to improve the relevancy of interpretations for some audience or problem. This is normally done by introducing a novel form of output, such as feature heatmaps, rationales, feature hierarchies or identifying important elements in the training set.\"</p>\n<p><a href=\"https://arxiv.org/pdf/1901.04592.pdf\" target=\"_blank\">https://arxiv.org/pdf/1901.04592.pdf</a></p>\n<h1>The black-box Nature</h1>\n<p>Interpretability of Machine Learning: Recent Advances and Future Prospects</p>\n<p>Authors: Gao, Lei and Guan, Ling</p>\n<p>\"The proliferation of machine learning (ML) has drawn unprecedented interest in the study of various multimedia contents such as text, image, audio and video, among others. Consequently,understanding and learning ML-based representations have taken center stage in knowledge discovery in intelligent multimedia research and applications. Nevertheless, the black-box nature of contemporary ML, especially in deep neural networks (DNNs), has posed a primary challenge for ML-based representation learning. To address this black-box problem, the studies on interpretability of ML have attracted tremendous interests in recent years.\"</p>\n<p>Algorithm Unrolling</p>\n<p>\"Algorithm unrolling solves model interpretability by providing a concrete and systematic connection between iterative algorithms that are widely used in signal process ing and DNNs. Given an iterative algorithm, a corresponding deep network is generated by cascading its iterations h. Then, iteration step h is executed a number of times, resulting in different parameters h1, h2,… Each iteration h depends on algorithm parameters, which are transferred into network parameters 1, 2,… Instead of determining parameters through cross-validation or analytical derivations, the parameters 1, 2,… are learned from training datasets through end-to-end training. In this way, the network layers naturally inherit interpretability from the iteration procedure.\"</p>\n<p><a href=\"https://arxiv.org/pdf/2305.00537.pdf\" target=\"_blank\">https://arxiv.org/pdf/2305.00537.pdf</a></p>\n<h1>Algorithm Unrolling</h1>\n<p>Algorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing</p>\n<p>Authors: Vishal Monga, Senior Member, Yuelong Li and Yonina C. Eldar</p>\n<p>\"Deep neural networks provide unprecedented performance gains in many real world problems in signal and image processing. Despite these gains, future development and practical deployment of deep networks is hindered by their black box nature (i.e.lack of interpretability, and by the need for very large training sets).\"</p>\n<p>\"An emerging technique called algorithm unrolling or unfolding offers promise in eliminating these issues (i.e.lack of interpretability, and by the need for very large training sets) by providing a concrete and systematic connection between iterative algorithms that are used widely in signal processing and deep neural networks.\"</p>\n<p>\"Unrolling methods were first proposed to develop fast neural network approximations for sparse coding. More recently, this direction has attracted enormous attention and is rapidly growing both in theoretic investigations and practical applications. The growing popularity of unrolled deep networks is due in part to their potential in developing efficient, high-performance and yet interpretable network architectures from reasonable size training sets.\"</p>\n<p>\"In this article, the authors reviewed algorithm unrolling for signal and image processing. They extensively covered popular techniques for algorithm unrolling in various domains of signal and image processing including imaging, vision and recognition, and speech processing. By reviewing previous works, they revealed the connections between iterative algorithms and neural networks and present recent theoretical results. Finally, the authors provided a discussion on current limitations of unrolling and suggest possible future research directions.\"</p>\n<p>Unrolling Sparse Coding Algorithms into Deep Networks</p>\n<p>\"The earliest work in algorithm unrolling dates back to Gregor et al.’s paper (2010) on improving the computational efficiency of sparse coding algorithms through end-to-end training. In particular, they discussed how to improve the efficiency of the Iterative Shrinkage and Thresholding Algorithm (ISTA), one of the most popular approaches in sparse coding. Learned ISTA. Each iteration of ISTA comprises one linear operation followed by a non-linear soft-thresholding operation, which mimics the ReLU activation function. A diagram representation of one iteration step reveals its resemblance to network is dubbed Learned ISTA (LISTA).\"</p>\n<p>Bridging the Gap between Theory and Practice:</p>\n<p>\"While substantial progress has been achieved towards understanding the network behavior through unrolling, more works need to be done to thoroughly understand its mechanism. Although the effectiveness of some networks on image reconstruction tasks has been explained somehow by drawing parallels to sparse coding algorithms, it is still mysterious why state-of-the art networks perform well on various recognition tasks. Further more, unfolding itself is not uniquely defined. For instance, there are multiple ways to choose the underlying iterative algorithms, to decide what parameters become trainable and what parameters to fix, and more.\"</p>\n<p>\"Another interesting direction is to develop a theory that provides guidance for practical applications. For instance, it is interesting to perform analysis that guide practical network design choices, such as dimensions of parameters, network depth, etc. It is particularly interesting to identify factors that have high impact on network performance.\"</p>\n<p><a href=\"https://arxiv.org/pdf/1912.10557.pdf\" target=\"_blank\">https://arxiv.org/pdf/1912.10557.pdf</a></p>",
      "rawMarkdown": "#Whoosh (Python Search Library) on Kaggle Notebooks\n\nSad fact, many of these Whoosh code were made during Covid19 times (4y ago) to retrieve Scientific Papers.\n\nBy Daniel Wolffram [Whoosh Search](https://www.kaggle.com/code/danielwolffram/whoosh-search)\n\nBy Max Feinberg [Using Whoosh for Indexing and Querying](https://www.kaggle.com/code/mxfeinberg/using-whoosh-for-indexing-and-querying)\n\nBy Alexander Rubin [Covid19 - build fulltext indexes with Whoosh](https://www.kaggle.com/code/arubin/covid19-build-fulltext-indexes-with-whoosh)\n\nBy Michael C Gold [https://www.kaggle.com/code/michaelcgold/covid-19-origins-in-a-whoosh](https://www.kaggle.com/code/michaelcgold/covid-19-origins-in-a-whoosh)\n\nBy Leire [RAG for TriviaQA](https://www.kaggle.com/code/leireher/rag-for-triviaqa)\n\nBy Sohier Dane [Basic Whoosh search demo](https://www.kaggle.com/code/sohier/basic-whoosh-search-demo)\n\n#On GitHub: ScispaCy \n\n\"ScispaCy is a Python package containing spaCy models for processing biomedical, scientific or clinical text.\"\n\n#Interpretability: Definitions, Methods and Applications\n\nInterpretable machine learning: definitions, methods, and applications\n\nAuthors: W. James Murdocha, Chandan Singhb, Karl Kumbiera, Reza Abbasi-Aslb, and Bin Yua\n\nDefining interpretable machine learning.\n\n\"On its own, interpretability is a broad, poorly defined concept. Taken to its full generality, to interpret data means to extract information (of some form) from it. The set of methods falling under this umbrella spans everything from designing an initial experiment to visualizing final results. In this overly general form, interpretability is not substantially different from the established concepts of data science and applied statistics.\"\n\nPost hoc analysis\n\n\"Having fit a model (or models), the practitioner then analyzes it for answers to the original question. The process of analyzing the model often involves using interpretability methods to extract various (stable) forms of information from the model. The extracted information can then be analyzed and displayed using standard data analysis methods, such as scatter plots and histograms. The ability of the interpretations to properly describe what the model has learned is denoted by descriptive accuracy.\"\n\nDemonstrating relevancy to real-world problems.\n\n\"Another angle for developing improved interpretation methods is to improve the relevancy of interpretations for some audience or problem. This is normally done by introducing a novel form of output, such as feature heatmaps, rationales, feature hierarchies or identifying important elements in the training set.\"\n\nhttps://arxiv.org/pdf/1901.04592.pdf\n\n\n#The black-box Nature\n\nInterpretability of Machine Learning: Recent Advances and Future Prospects\n\nAuthors: Gao, Lei and Guan, Ling\n\n\"The proliferation of machine learning (ML) has drawn unprecedented interest in the study of various multimedia contents such as text, image, audio and video, among others. Consequently,understanding and learning ML-based representations have taken center stage in knowledge discovery in intelligent multimedia research and applications. Nevertheless, the black-box nature of contemporary ML, especially in deep neural networks (DNNs), has posed a primary challenge for ML-based representation learning. To address this black-box problem, the studies on interpretability of ML have attracted tremendous interests in recent years.\"\n\nAlgorithm Unrolling\n\n\"Algorithm unrolling solves model interpretability by providing a concrete and systematic connection between iterative algorithms that are widely used in signal process ing and DNNs. Given an iterative algorithm, a corresponding deep network is generated by cascading its iterations h. Then, iteration step h is executed a number of times, resulting in different parameters h1, h2,... Each iteration h depends on algorithm parameters, which are transferred into network parameters 1, 2,... Instead of determining parameters through cross-validation or analytical derivations, the parameters 1, 2,... are learned from training datasets through end-to-end training. In this way, the network layers naturally inherit interpretability from the iteration procedure.\"\n\nhttps://arxiv.org/pdf/2305.00537.pdf\n\n#Algorithm Unrolling\n\nAlgorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing\n\nAuthors: Vishal Monga, Senior Member, Yuelong Li and Yonina C. Eldar\n\n\"Deep neural networks provide unprecedented performance gains in many real world problems in signal and image processing. Despite these gains, future development and practical deployment of deep networks is hindered by their black box nature (i.e.lack of interpretability, and by the need for very large training sets).\"\n\n\"An emerging technique called algorithm unrolling or unfolding offers promise in eliminating these issues (i.e.lack of interpretability, and by the need for very large training sets) by providing a concrete and systematic connection between iterative algorithms that are used widely in signal processing and deep neural networks.\"\n\n\"Unrolling methods were first proposed to develop fast neural network approximations for sparse coding. More recently, this direction has attracted enormous attention and is rapidly growing both in theoretic investigations and practical applications. The growing popularity of unrolled deep networks is due in part to their potential in developing efficient, high-performance and yet interpretable network architectures from reasonable size training sets.\"\n\n\"In this article, the authors reviewed algorithm unrolling for signal and image processing. They extensively covered popular techniques for algorithm unrolling in various domains of signal and image processing including imaging, vision and recognition, and speech processing. By reviewing previous works, they revealed the connections between iterative algorithms and neural networks and present recent theoretical results. Finally, the authors provided a discussion on current limitations of unrolling and suggest possible future research directions.\"\n\nUnrolling Sparse Coding Algorithms into Deep Networks\n\n\"The earliest work in algorithm unrolling dates back to Gregor et al.’s paper (2010) on improving the computational efficiency of sparse coding algorithms through end-to-end training. In particular, they discussed how to improve the efficiency of the Iterative Shrinkage and Thresholding Algorithm (ISTA), one of the most popular approaches in sparse coding. Learned ISTA. Each iteration of ISTA comprises one linear operation followed by a non-linear soft-thresholding operation, which mimics the ReLU activation function. A diagram representation of one iteration step reveals its resemblance to network is dubbed Learned ISTA (LISTA).\"\n\nBridging the Gap between Theory and Practice:\n\n\"While substantial progress has been achieved towards understanding the network behavior through unrolling, more works need to be done to thoroughly understand its mechanism. Although the effectiveness of some networks on image reconstruction tasks has been explained somehow by drawing parallels to sparse coding algorithms, it is still mysterious why state-of-the art networks perform well on various recognition tasks. Further more, unfolding itself is not uniquely defined. For instance, there are multiple ways to choose the underlying iterative algorithms, to decide what parameters become trainable and what parameters to fix, and more.\"\n\n\"Another interesting direction is to develop a theory that provides guidance for practical applications. For instance, it is interesting to perform analysis that guide practical network design choices, such as dimensions of parameters, network depth, etc. It is particularly interesting to identify factors that have high impact on network performance.\"\n\nhttps://arxiv.org/pdf/1912.10557.pdf",
      "votes": null
    },
    {
      "id": "2774006",
      "postDate": "04/25/2024 02:56:44",
      "content": "<p>Thank you for taking the time to review [Interpretable Machine Learning]. Your feedback is greatly appreciated! <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> </p>",
      "rawMarkdown": "Thank you for taking the time to review [Interpretable Machine Learning]. Your feedback is greatly appreciated! @mpwolke",
      "votes": null
    },
    {
      "id": "2774338",
      "postDate": "04/25/2024 06:40:56",
      "content": "<p>Wait, what are \"Whoosh\" notebooks? And what is their connection to… papers?</p>",
      "rawMarkdown": "Wait, what are \"Whoosh\" notebooks? And what is their connection to... papers?",
      "votes": null
    },
    {
      "id": "2775056",
      "postDate": "04/25/2024 13:08:04",
      "content": "<p>Hi TheItCrow,<br>\nWhen we had the Pandemic we had <a href=\"https://www.kaggle.com/datasets/allen-institute-for-ai/CORD-19-research-challenge\" target=\"_blank\">COVID-19 Open Research Dataset Challenge (CORD-19)</a> Challenge in which the objective was to find Scientific papers.</p>\n<p>The dataset was created by the Allen Institute for AI  and we have TASKS to find specific subjects such as: vaccines, diagnostics, non-pharmaceutical interventions.  Many Kagglers used Whoosh (FOUR YEARS AGO) that is a pure Python search engine library to perform those tasks and find the papers.<br>\n<a href=\"https://whoosh.readthedocs.io/en/latest/intro.html\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/intro.html</a></p>\n<blockquote>\n  <p>ScispaCy<br>\n   I added also on this topic ScispaCy, that is a Python package containing spaCy models for processing biomedical, scientific or clinical text.\"</p>\n</blockquote>\n<p>Besides, after reading this USPTO Competition Data Section (\"train_index A WHOOSH text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries.\" ), I decided to make something different, since I didn't apply Whoosh on the last Challenge. I searched by code Whoosh however I failed so many times cause all the Codes mentioned time/intervals and authors. And we don't have these features. And you know that I'm Not able to adapt \"if/else\" (i.e. any function snippet : )</p>\n<p>I was almost finishing my Notebook when I saw Sohier's (Competition Host) starter code.  However, even his code I wasn't able to deliver due to modules Not found.</p>\n<p>I added \"Python Search Library\" to this topic title  to avoid any misunderstanding. Since many Kagglers don't read, some could think that Whoosh is \"Wow\" and Not the name of a Python Search Library.</p>\n<p>Unfortunately, I wasn't able to adapt any of those Notebooks. That's my favorite \"Whoosh Code\" that won one of the tasks of the Competition:<br>\nBy Daniel Wolffram <a href=\"https://www.kaggle.com/code/danielwolffram/whoosh-search\" target=\"_blank\">Whoosh Search</a></p>",
      "rawMarkdown": "Hi TheItCrow,\nWhen we had the Pandemic we had [COVID-19 Open Research Dataset Challenge (CORD-19)](https://www.kaggle.com/datasets/allen-institute-for-ai/CORD-19-research-challenge) Challenge in which the objective was to find Scientific papers.\n\nThe dataset was created by the Allen Institute for AI  and we have TASKS to find specific subjects such as: vaccines, diagnostics, non-pharmaceutical interventions.  Many Kagglers used Whoosh (FOUR YEARS AGO) that is a pure Python search engine library to perform those tasks and find the papers.\nhttps://whoosh.readthedocs.io/en/latest/intro.html\n\n>ScispaCy\n I added also on this topic ScispaCy, that is a Python package containing spaCy models for processing biomedical, scientific or clinical text.\"\n\nBesides, after reading this USPTO Competition Data Section (\"train_index A WHOOSH text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries.\" ), I decided to make something different, since I didn't apply Whoosh on the last Challenge. I searched by code Whoosh however I failed so many times cause all the Codes mentioned time/intervals and authors. And we don't have these features. And you know that I'm Not able to adapt \"if/else\" (i.e. any function snippet : )\n\nI was almost finishing my Notebook when I saw Sohier's (Competition Host) starter code.  However, even his code I wasn't able to deliver due to modules Not found.\n\n I added \"Python Search Library\" to this topic title  to avoid any misunderstanding. Since many Kagglers don't read, some could think that Whoosh is \"Wow\" and Not the name of a Python Search Library.\n\n Unfortunately, I wasn't able to adapt any of those Notebooks. That's my favorite \"Whoosh Code\" that won one of the tasks of the Competition:\nBy Daniel Wolffram [Whoosh Search](https://www.kaggle.com/code/danielwolffram/whoosh-search)",
      "votes": null
    },
    {
      "id": "2775094",
      "postDate": "04/25/2024 13:30:01",
      "content": "<p>Hi Alaref,<br>\nwhat many beginners should read is that below, since many don't write anything after the model is done: <br>\n\"Post hoc analysis:\"Having fit a model (or models), the practitioner then analyzes it for answers to the original question.\"   We ready many Kaggle Notebooks without any conclusion after the model.  Thank you. </p>",
      "rawMarkdown": "Hi Alaref,\nwhat many beginners should read is that below, since many don't write anything after the model is done: \n\"Post hoc analysis:\"Having fit a model (or models), the practitioner then analyzes it for answers to the original question.\"   We ready many Kaggle Notebooks without any conclusion after the model.  Thank you.",
      "votes": null
    },
    {
      "id": "2775208",
      "postDate": "04/25/2024 14:44:02",
      "content": "<p>I see, thank you!</p>",
      "rawMarkdown": "I see, thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2774006,
      "author_name": "adnanalaref",
      "author_url": "",
      "post_date": "04/25/2024 02:56:44",
      "content": "<p>Thank you for taking the time to review [Interpretable Machine Learning]. Your feedback is greatly appreciated! <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2775094,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "04/25/2024 13:30:01",
          "content": "<p>Hi Alaref,<br>\nwhat many beginners should read is that below, since many don't write anything after the model is done: <br>\n\"Post hoc analysis:\"Having fit a model (or models), the practitioner then analyzes it for answers to the original question.\"   We ready many Kaggle Notebooks without any conclusion after the model.  Thank you. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2774338,
      "author_name": "kevinbnisch",
      "author_url": "",
      "post_date": "04/25/2024 06:40:56",
      "content": "<p>Wait, what are \"Whoosh\" notebooks? And what is their connection to… papers?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2775056,
          "author_name": "mpwolke",
          "author_url": "",
          "post_date": "04/25/2024 13:08:04",
          "content": "<p>Hi TheItCrow,<br>\nWhen we had the Pandemic we had <a href=\"https://www.kaggle.com/datasets/allen-institute-for-ai/CORD-19-research-challenge\" target=\"_blank\">COVID-19 Open Research Dataset Challenge (CORD-19)</a> Challenge in which the objective was to find Scientific papers.</p>\n<p>The dataset was created by the Allen Institute for AI  and we have TASKS to find specific subjects such as: vaccines, diagnostics, non-pharmaceutical interventions.  Many Kagglers used Whoosh (FOUR YEARS AGO) that is a pure Python search engine library to perform those tasks and find the papers.<br>\n<a href=\"https://whoosh.readthedocs.io/en/latest/intro.html\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/intro.html</a></p>\n<blockquote>\n  <p>ScispaCy<br>\n   I added also on this topic ScispaCy, that is a Python package containing spaCy models for processing biomedical, scientific or clinical text.\"</p>\n</blockquote>\n<p>Besides, after reading this USPTO Competition Data Section (\"train_index A WHOOSH text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries.\" ), I decided to make something different, since I didn't apply Whoosh on the last Challenge. I searched by code Whoosh however I failed so many times cause all the Codes mentioned time/intervals and authors. And we don't have these features. And you know that I'm Not able to adapt \"if/else\" (i.e. any function snippet : )</p>\n<p>I was almost finishing my Notebook when I saw Sohier's (Competition Host) starter code.  However, even his code I wasn't able to deliver due to modules Not found.</p>\n<p>I added \"Python Search Library\" to this topic title  to avoid any misunderstanding. Since many Kagglers don't read, some could think that Whoosh is \"Wow\" and Not the name of a Python Search Library.</p>\n<p>Unfortunately, I wasn't able to adapt any of those Notebooks. That's my favorite \"Whoosh Code\" that won one of the tasks of the Competition:<br>\nBy Daniel Wolffram <a href=\"https://www.kaggle.com/code/danielwolffram/whoosh-search\" target=\"_blank\">Whoosh Search</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2775208,
              "author_name": "kevinbnisch",
              "author_url": "",
              "post_date": "04/25/2024 14:44:02",
              "content": "<p>I see, thank you!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2773953": "#Whoosh (Python Search Library) on Kaggle Notebooks\n\nSad fact, many of these Whoosh code were made during Covid19 times (4y ago) to retrieve Scientific Papers.\n\nBy Daniel Wolffram [Whoosh Search](https://www.kaggle.com/code/danielwolffram/whoosh-search)\n\nBy Max Feinberg [Using Whoosh for Indexing and Querying](https://www.kaggle.com/code/mxfeinberg/using-whoosh-for-indexing-and-querying)\n\nBy Alexander Rubin [Covid19 - build fulltext indexes with Whoosh](https://www.kaggle.com/code/arubin/covid19-build-fulltext-indexes-with-whoosh)\n\nBy Michael C Gold [https://www.kaggle.com/code/michaelcgold/covid-19-origins-in-a-whoosh](https://www.kaggle.com/code/michaelcgold/covid-19-origins-in-a-whoosh)\n\nBy Leire [RAG for TriviaQA](https://www.kaggle.com/code/leireher/rag-for-triviaqa)\n\nBy Sohier Dane [Basic Whoosh search demo](https://www.kaggle.com/code/sohier/basic-whoosh-search-demo)\n\n#On GitHub: ScispaCy \n\n\"ScispaCy is a Python package containing spaCy models for processing biomedical, scientific or clinical text.\"\n\n#Interpretability: Definitions, Methods and Applications\n\nInterpretable machine learning: definitions, methods, and applications\n\nAuthors: W. James Murdocha, Chandan Singhb, Karl Kumbiera, Reza Abbasi-Aslb, and Bin Yua\n\nDefining interpretable machine learning.\n\n\"On its own, interpretability is a broad, poorly defined concept. Taken to its full generality, to interpret data means to extract information (of some form) from it. The set of methods falling under this umbrella spans everything from designing an initial experiment to visualizing final results. In this overly general form, interpretability is not substantially different from the established concepts of data science and applied statistics.\"\n\nPost hoc analysis\n\n\"Having fit a model (or models), the practitioner then analyzes it for answers to the original question. The process of analyzing the model often involves using interpretability methods to extract various (stable) forms of information from the model. The extracted information can then be analyzed and displayed using standard data analysis methods, such as scatter plots and histograms. The ability of the interpretations to properly describe what the model has learned is denoted by descriptive accuracy.\"\n\nDemonstrating relevancy to real-world problems.\n\n\"Another angle for developing improved interpretation methods is to improve the relevancy of interpretations for some audience or problem. This is normally done by introducing a novel form of output, such as feature heatmaps, rationales, feature hierarchies or identifying important elements in the training set.\"\n\nhttps://arxiv.org/pdf/1901.04592.pdf\n\n\n#The black-box Nature\n\nInterpretability of Machine Learning: Recent Advances and Future Prospects\n\nAuthors: Gao, Lei and Guan, Ling\n\n\"The proliferation of machine learning (ML) has drawn unprecedented interest in the study of various multimedia contents such as text, image, audio and video, among others. Consequently,understanding and learning ML-based representations have taken center stage in knowledge discovery in intelligent multimedia research and applications. Nevertheless, the black-box nature of contemporary ML, especially in deep neural networks (DNNs), has posed a primary challenge for ML-based representation learning. To address this black-box problem, the studies on interpretability of ML have attracted tremendous interests in recent years.\"\n\nAlgorithm Unrolling\n\n\"Algorithm unrolling solves model interpretability by providing a concrete and systematic connection between iterative algorithms that are widely used in signal process ing and DNNs. Given an iterative algorithm, a corresponding deep network is generated by cascading its iterations h. Then, iteration step h is executed a number of times, resulting in different parameters h1, h2,... Each iteration h depends on algorithm parameters, which are transferred into network parameters 1, 2,... Instead of determining parameters through cross-validation or analytical derivations, the parameters 1, 2,... are learned from training datasets through end-to-end training. In this way, the network layers naturally inherit interpretability from the iteration procedure.\"\n\nhttps://arxiv.org/pdf/2305.00537.pdf\n\n#Algorithm Unrolling\n\nAlgorithm Unrolling: Interpretable, Efficient Deep Learning for Signal and Image Processing\n\nAuthors: Vishal Monga, Senior Member, Yuelong Li and Yonina C. Eldar\n\n\"Deep neural networks provide unprecedented performance gains in many real world problems in signal and image processing. Despite these gains, future development and practical deployment of deep networks is hindered by their black box nature (i.e.lack of interpretability, and by the need for very large training sets).\"\n\n\"An emerging technique called algorithm unrolling or unfolding offers promise in eliminating these issues (i.e.lack of interpretability, and by the need for very large training sets) by providing a concrete and systematic connection between iterative algorithms that are used widely in signal processing and deep neural networks.\"\n\n\"Unrolling methods were first proposed to develop fast neural network approximations for sparse coding. More recently, this direction has attracted enormous attention and is rapidly growing both in theoretic investigations and practical applications. The growing popularity of unrolled deep networks is due in part to their potential in developing efficient, high-performance and yet interpretable network architectures from reasonable size training sets.\"\n\n\"In this article, the authors reviewed algorithm unrolling for signal and image processing. They extensively covered popular techniques for algorithm unrolling in various domains of signal and image processing including imaging, vision and recognition, and speech processing. By reviewing previous works, they revealed the connections between iterative algorithms and neural networks and present recent theoretical results. Finally, the authors provided a discussion on current limitations of unrolling and suggest possible future research directions.\"\n\nUnrolling Sparse Coding Algorithms into Deep Networks\n\n\"The earliest work in algorithm unrolling dates back to Gregor et al.’s paper (2010) on improving the computational efficiency of sparse coding algorithms through end-to-end training. In particular, they discussed how to improve the efficiency of the Iterative Shrinkage and Thresholding Algorithm (ISTA), one of the most popular approaches in sparse coding. Learned ISTA. Each iteration of ISTA comprises one linear operation followed by a non-linear soft-thresholding operation, which mimics the ReLU activation function. A diagram representation of one iteration step reveals its resemblance to network is dubbed Learned ISTA (LISTA).\"\n\nBridging the Gap between Theory and Practice:\n\n\"While substantial progress has been achieved towards understanding the network behavior through unrolling, more works need to be done to thoroughly understand its mechanism. Although the effectiveness of some networks on image reconstruction tasks has been explained somehow by drawing parallels to sparse coding algorithms, it is still mysterious why state-of-the art networks perform well on various recognition tasks. Further more, unfolding itself is not uniquely defined. For instance, there are multiple ways to choose the underlying iterative algorithms, to decide what parameters become trainable and what parameters to fix, and more.\"\n\n\"Another interesting direction is to develop a theory that provides guidance for practical applications. For instance, it is interesting to perform analysis that guide practical network design choices, such as dimensions of parameters, network depth, etc. It is particularly interesting to identify factors that have high impact on network performance.\"\n\nhttps://arxiv.org/pdf/1912.10557.pdf",
    "2774006": "Thank you for taking the time to review [Interpretable Machine Learning]. Your feedback is greatly appreciated! @mpwolke",
    "2774338": "Wait, what are \"Whoosh\" notebooks? And what is their connection to... papers?",
    "2775056": "Hi TheItCrow,\nWhen we had the Pandemic we had [COVID-19 Open Research Dataset Challenge (CORD-19)](https://www.kaggle.com/datasets/allen-institute-for-ai/CORD-19-research-challenge) Challenge in which the objective was to find Scientific papers.\n\nThe dataset was created by the Allen Institute for AI  and we have TASKS to find specific subjects such as: vaccines, diagnostics, non-pharmaceutical interventions.  Many Kagglers used Whoosh (FOUR YEARS AGO) that is a pure Python search engine library to perform those tasks and find the papers.\nhttps://whoosh.readthedocs.io/en/latest/intro.html\n\n>ScispaCy\n I added also on this topic ScispaCy, that is a Python package containing spaCy models for processing biomedical, scientific or clinical text.\"\n\nBesides, after reading this USPTO Competition Data Section (\"train_index A WHOOSH text search index equivalent in size and setup to the index the metric will use to evaluate submitted queries.\" ), I decided to make something different, since I didn't apply Whoosh on the last Challenge. I searched by code Whoosh however I failed so many times cause all the Codes mentioned time/intervals and authors. And we don't have these features. And you know that I'm Not able to adapt \"if/else\" (i.e. any function snippet : )\n\nI was almost finishing my Notebook when I saw Sohier's (Competition Host) starter code.  However, even his code I wasn't able to deliver due to modules Not found.\n\n I added \"Python Search Library\" to this topic title  to avoid any misunderstanding. Since many Kagglers don't read, some could think that Whoosh is \"Wow\" and Not the name of a Python Search Library.\n\n Unfortunately, I wasn't able to adapt any of those Notebooks. That's my favorite \"Whoosh Code\" that won one of the tasks of the Competition:\nBy Daniel Wolffram [Whoosh Search](https://www.kaggle.com/code/danielwolffram/whoosh-search)",
    "2775094": "Hi Alaref,\nwhat many beginners should read is that below, since many don't write anything after the model is done: \n\"Post hoc analysis:\"Having fit a model (or models), the practitioner then analyzes it for answers to the original question.\"   We ready many Kaggle Notebooks without any conclusion after the model.  Thank you.",
    "2775208": "I see, thank you!"
  },
  "source": "meta"
}