XAI - An eXplainability toolbox for machine learning

Last update: Dec 27, 2022

Overview

XAI - An eXplainability toolbox for machine learning

XAI is a Machine Learning library that is designed with AI explainability in its core. XAI contains various tools that enable for analysis and evaluation of data and models. The XAI library is maintained by The Institute for Ethical AI & ML, and it was developed based on the 8 principles for Responsible Machine Learning.

You can find the documentation at https://ethicalml.github.io/xai/index.html. You can also check out our talk at Tensorflow London where the idea was first conceived - the talk also contains an insight on the definitions and principles in this library.

YouTube video showing how to use XAI to mitigate undesired biases

This video of the talk presented at the PyData London 2019 Conference which provides an overview on the motivations for machine learning explainability as well as techniques to introduce explainability and mitigate undesired biases using the XAI Library.
Do you want to learn about more awesome machine learning explainability tools? Check out our community-built "Awesome Machine Learning Production & Operations" list which contains an extensive list of tools for explainability, privacy, orchestration and beyond.

0.1.0

If you want to see a fully functional demo in action clone this repo and run the Example Jupyter Notebook in the Examples folder.

What do we mean by eXplainable AI?

We see the challenge of explainability as more than just an algorithmic challenge, which requires a combination of data science best practices with domain-specific knowledge. The XAI library is designed to empower machine learning engineers and relevant domain experts to analyse the end-to-end solution and identify discrepancies that may result in sub-optimal performance relative to the objectives required. More broadly, the XAI library is designed using the 3-steps of explainable machine learning, which involve 1) data analysis, 2) model evaluation, and 3) production monitoring.

We provide a visual overview of these three steps mentioned above in this diagram:

XAI Quickstart

Installation

The XAI package is on PyPI. To install you can run:

pip install xai

Alternatively you can install from source by cloning the repo and running:

python setup.py install

Usage

You can find example usage in the examples folder.

1) Data Analysis

With XAI you can identify imbalances in the data. For this, we will load the census dataset from the XAI library.

import xai.data
df = xai.data.load_census()
df.head()

View class imbalances for all categories of one column

ims = xai.imbalance_plot(df, "gender")

View imbalances for all categories across multiple columns

im = xai.imbalance_plot(df, "gender", "loan")

Balance classes using upsampling and/or downsampling

bal_df = xai.balance(df, "gender", "loan", upsample=0.8)

Perform custom operations on groups

groups = xai.group_by_columns(df, ["gender", "loan"])
for group, group_df in groups:    
    print(group) 
    print(group_df["loan"].head(), "\n")

Visualise correlations as a matrix

_ = xai.correlations(df, include_categorical=True, plot_type="matrix")

Visualise correlations as a hierarchical dendogram

_ = xai.correlations(df, include_categorical=True)

Create a balanced validation and training split dataset

# Balanced train-test split with minimum 300 examples of 
#     the cross of the target y and the column gender
x_train, y_train, x_test, y_test, train_idx, test_idx = \
    xai.balanced_train_test_split(
            x, y, "gender", 
            min_per_group=300,
            max_per_group=300,
            categorical_cols=categorical_cols)

x_train_display = bal_df[train_idx]
x_test_display = bal_df[test_idx]

print("Total number of examples: ", x_test.shape[0])

df_test = x_test_display.copy()
df_test["loan"] = y_test

_= xai.imbalance_plot(df_test, "gender", "loan", categorical_cols=categorical_cols)

2) Model Evaluation

We are able to also analyse the interaction between inference results and input features. For this, we will train a single layer deep learning model.

= 0.5).astype(int).T[0]) ">

model = build_model(proc_df.drop("loan", axis=1))

model.fit(f_in(x_train), y_train, epochs=50, batch_size=512)

probabilities = model.predict(f_in(x_test))
predictions = list((probabilities >= 0.5).astype(int).T[0])

Visualise permutation feature importance

def get_avg(x, y):
    return model.evaluate(f_in(x), y, verbose=0)[1]

imp = xai.feature_importance(x_test, y_test, get_avg)

imp.head()

Identify metric imbalances against all test data

_= xai.metrics_plot(
        y_test, 
        probabilities)

Identify metric imbalances across a specific column

_ = xai.metrics_plot(
    y_test, 
    probabilities, 
    df=x_test_display, 
    cross_cols=["gender"],
    categorical_cols=categorical_cols)

Identify metric imbalances across multiple columns

_ = xai.metrics_plot(
    y_test, 
    probabilities, 
    df=x_test_display, 
    cross_cols=["gender", "ethnicity"],
    categorical_cols=categorical_cols)

Draw confusion matrix

xai.confusion_matrix_plot(y_test, pred)

Visualise the ROC curve against all test data

_ = xai.roc_plot(y_test, probabilities)

Visualise the ROC curves grouped by a protected column

protected = ["gender", "ethnicity", "age"]
_ = [xai.roc_plot(
    y_test, 
    probabilities, 
    df=x_test_display, 
    cross_cols=[p],
    categorical_cols=categorical_cols) for p in protected]

Visualise accuracy grouped by probability buckets

d = xai.smile_imbalance(
    y_test, 
    probabilities)

Visualise statistical metrics grouped by probability buckets

d = xai.smile_imbalance(
    y_test, 
    probabilities,
    display_breakdown=True)

Visualise benefits of adding manual review on probability thresholds

d = xai.smile_imbalance(
    y_test, 
    probabilities,
    bins=9,
    threshold=0.75,
    manual_review=0.375,
    display_breakdown=False)

Comments

matplotlib error while installing package

Collecting matplotlib==3.0.2

Using cached matplotlib-3.0.2.tar.gz (36.5 MB) ERROR: Command errored out with exit status 1: command: /opt/anaconda3/envs/ethicalml/bin/python -c 'import sys, setuptools, tokenize; sys.argv[0] = '"'"'/private/var/folders/wv/m62_p54d5bx1dnq_m07ck3l40000gn/T/pip-install-303disb7/matplotlib/setup.py'"'"'; file='"'"'/private/var/folders/wv/m62_p54d5bx1dnq_m07ck3l40000gn/T/pip-install-303disb7/matplotlib/setup.py'"'"';f=getattr(tokenize, '"'"'open'"'"', open)(file);code=f.read().replace('"'"'\r\n'"'"', '"'"'\n'"'"');f.close();exec(compile(code, file, '"'"'exec'"'"'))' egg_info --egg-base /private/var/folders/wv/m62_p54d5bx1dnq_m07ck3l40000gn/T/pip-install-303disb7/matplotlib/pip-egg-info

opened by ArpitSisodia 3
Requirements
your requirements are very restrictive. Can you please change it to >= instead of ==. for example:

numpy>=1.3 pandas>=0.23.0 matplotlib>2.02,<=3.0.3 scikit-learn>=0.19.0
opened by idanmoradarthas 3
converters the probs into np array if its already not
smile_imbalance() funciton argument "probs" does not specify that it is required to be numpy array, but it does so i have added that data type in the argument letting the user know if he/she is to refer to the docs and i have also added a line np.array() which is an idempotent operation(if the array passed is already numpy array then it does nothing but if its not it changes the list into numpy array)

Suggestion

If possible can you guys consider adding "save_plot_path" method to each function, so that when this package is used in production (which i am and people considering Continuous model delivery would use) all these plots could be saved to a particular directory for data scientists to look at later since in production, code would be used in scripts running on EC2 or other cloud servers and not on jupyter notebooks

My use case is I am retraining the model every week and XAI allows me to generate a evaluation report allowing me to remotely decide weather to push this weeks mode into production

I considered adding it myself but i was not sure if this is the direction you guys wanted to take

Thank you
opened by sai-krishna-msk 2
Unable to install package
Hello!

I've been trying to install this package and am unable to do so. I've tried both methods on my Ubuntu machine.

pip install xai

python setup.py install

What can I do to install this? Also, is this project active anymore at all?
opened by varunbanda 2
Can we explain BERT models using this package?

I'm working with text data and looking for ways to explain BERT models. Is there any workaround using XAI or any other package/resources if anyone can recommend?

opened by techwithshadab 1
Add a conda install option for `xai`
A conda installation option could be very helpful. I have already started working on this, to add xai to conda-forge.

Conda-forge PR:

https://github.com/conda-forge/staged-recipes/pull/17601

Once the conda-forge PR is merged, you will be able to install the library with conda as follows:

conda install -c conda-forge xai

:bulb: I will push a PR to update the docs once the package is available on conda-forge.
opened by sugatoray 0
Wrong series returned from _curve
There is some bug in https://github.com/EthicalML/xai/blob/master/xai/init.py#L962

it was written as

r1s = r2s = []

but should be instead

r1s, r2s = [], []

The impact is that if the user would like to us r1s and r2s returned to construct the the curve (e.g. for storing the data for later analysis), they would find that r1s and r2s are referring to the same instance which stores all the curve data that should have been separately stored in r1s and r2s
opened by chen0040 1

Releases(v0.1.0)

v0.1.0(Oct 30, 2021)

Release Version v0.1.0
Source code(tar.gz)
Source code(zip)

Owner

The Institute for Ethical Machine Learning

The Institute for Ethical Machine Learning is a think-tank that brings together with technology leaders, policymakers & academics to develop standards for ML.

GitHub Repository https://ethical.institute/principles.html#commitment-3

Summer: compartmental disease modelling in Python

Summer: compartmental disease modelling in Python Summer is a Python-based framework for the creation and execution of compartmental (or "state-based"

6 May 13, 2022

A repository of PyBullet utility functions for robotic motion planning, manipulation planning, and task and motion planning

pybullet-planning (previously ss-pybullet) A repository of PyBullet utility functions for robotic motion planning, manipulation planning, and task and

260 Dec 27, 2022

A collection of interactive machine-learning experiments: 🏋️models training + 🎨models demo

🤖 Interactive Machine Learning experiments: 🏋️models training + 🎨models demo

1.4k Jan 06, 2023

A game theoretic approach to explain the output of any machine learning model.

SHAP (SHapley Additive exPlanations) is a game theoretic approach to explain the output of any machine learning model. It connects optimal credit allo

18.2k Jan 02, 2023

Bayesian optimization in JAX

26 May 11, 2022

Class-imbalanced / Long-tailed ensemble learning in Python. Modular, flexible, and extensible

176 Jan 04, 2023

Uses WiFi signals :signal_strength: and machine learning to predict where you are

Uses WiFi signals and machine learning (sklearn's RandomForest) to predict where you are. Even works for small distances like 2-10 meters.

5k Jan 09, 2023

The Ultimate FREE Machine Learning Study Plan

2.5k Jan 05, 2023

A Multipurpose Library for Synthetic Time Series Generation in Python

TimeSynth Multipurpose Library for Synthetic Time Series Please cite as: J. R. Maat, A. Malali, and P. Protopapas, “TimeSynth: A Multipurpose Library

278 Dec 26, 2022

Tribuo - A Java machine learning library

Tribuo - A Java prediction library (v4.1) Tribuo is a machine learning library in Java that provides multi-class classification, regression, clusterin

1.1k Dec 28, 2022

This repository demonstrates the usage of hover to understand and supervise a machine learning task.

Hover Example Apps (works out-of-the-box on Binder) This repository demonstrates the usage of hover to understand and supervise a machine learning tas

43 Dec 03, 2021

Auto updating website that tracks closed & open issues/PRs on scikit-learn/scikit-learn.

Repository Status for Scikit-learn Live webpage Auto updating website that tracks closed & open issues/PRs on scikit-learn/scikit-learn. Running local

6 Dec 27, 2022

MICOM is a Python package for metabolic modeling of microbial communities

Welcome MICOM is a Python package for metabolic modeling of microbial communities currently developed in the Gibbons Lab at the Institute for Systems

57 Dec 21, 2022

PLUR is a collection of source code datasets suitable for graph-based machine learning.

PLUR (Programming-Language Understanding and Repair) is a collection of source code datasets suitable for graph-based machine learning. We provide scripts for downloading, processing, and loading the

76 Nov 25, 2022

pure-predict: Machine learning prediction in pure Python

pure-predict speeds up and slims down machine learning prediction applications. It is a foundational tool for serverless inference or small batch prediction with popular machine learning frameworks l

84 Dec 29, 2022

Provide an input CSV and a target field to predict, generate a model + code to run it.

automl-gs Give an input CSV file and a target field you want to predict to automl-gs, and get a trained high-performing machine learning or deep learn

1.8k Jan 04, 2023

DistML is a Ray extension library to support large-scale distributed ML training on heterogeneous multi-node multi-GPU clusters

27 Aug 19, 2022

XAI - An eXplainability toolbox for machine learning

Related tags

Overview

XAI - An eXplainability toolbox for machine learning

YouTube video showing how to use XAI to mitigate undesired biases

0.1.0

What do we mean by eXplainable AI?

XAI Quickstart

Installation

Usage

1) Data Analysis

View class imbalances for all categories of one column

View imbalances for all categories across multiple columns

Balance classes using upsampling and/or downsampling

Perform custom operations on groups

Visualise correlations as a matrix

Visualise correlations as a hierarchical dendogram

Create a balanced validation and training split dataset

2) Model Evaluation

Visualise permutation feature importance

Identify metric imbalances against all test data

Identify metric imbalances across a specific column

Identify metric imbalances across multiple columns

Draw confusion matrix

Visualise the ROC curve against all test data

Visualise the ROC curves grouped by a protected column

Visualise accuracy grouped by probability buckets

Visualise statistical metrics grouped by probability buckets

Visualise benefits of adding manual review on probability thresholds

Comments

matplotlib error while installing package

Requirements

converters the probs into np array if its already not

Unable to install package

Can we explain BERT models using this package?

Add a conda install option for `xai`

Wrong series returned from _curve

Releases(v0.1.0)

v0.1.0(Oct 30, 2021)

Release Version v0.1.0

Owner

The Institute for Ethical Machine Learning

Summer: compartmental disease modelling in Python

A repository of PyBullet utility functions for robotic motion planning, manipulation planning, and task and motion planning

A collection of interactive machine-learning experiments: 🏋️models training + 🎨models demo

A game theoretic approach to explain the output of any machine learning model.

Bayesian optimization in JAX

Class-imbalanced / Long-tailed ensemble learning in Python. Modular, flexible, and extensible

Uses WiFi signals :signal_strength: and machine learning to predict where you are

The Ultimate FREE Machine Learning Study Plan

A Multipurpose Library for Synthetic Time Series Generation in Python

Tribuo - A Java machine learning library

This repository demonstrates the usage of hover to understand and supervise a machine learning task.

Auto updating website that tracks closed & open issues/PRs on scikit-learn/scikit-learn.

MICOM is a Python package for metabolic modeling of microbial communities

PLUR is a collection of source code datasets suitable for graph-based machine learning.

pure-predict: Machine learning prediction in pure Python

Provide an input CSV and a target field to predict, generate a model + code to run it.

DistML is a Ray extension library to support large-scale distributed ML training on heterogeneous multi-node multi-GPU clusters

Arquivos do curso online sobre a estatística voltada para ciência de dados e aprendizado de máquina.

ML-powered Loan-Marketer Customer Filtering Engine

Data science, Data manipulation and Machine learning package.